A burst of output leaves a Shell pane with about one screen of browser
scrollback, because tmux repaints the burst instead of scrolling it,
while tmux itself keeps every line. Shell declined the scroll-to-top
re-pull other modes use, and the Load full history button renders only
once a replay was truncated, so a Shell tab under 1 MiB could not
scroll back at all.
The scroll gesture now pulls ?full=1&tail=TERMINAL_TAIL_SIZE, the same
bound a tab switch loads; the route's existing tail cut marks longer
histories 'tail', so the banner still offers the unbounded pull. A
window no longer than the browser's buffer is not rewritten.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The phone case picker had no way to narrow a long case list, so finding
one meant scrolling a sheet that showed about six rows at a time.
- A search field filters rows by name (every typed word must match, any
order, case-insensitive), with a "No matching cases" state. Enter picks
the case when exactly one row is left; Escape clears, then closes.
- The field is not auto-focused, so opening the picker does not raise the
keyboard. The list holds its unfiltered height while searching so the
sheet does not jump, and the input is 16px so iOS Safari does not zoom.
- Layout: the sheet padded the home-indicator inset on top of the footer
already doing so, leaving a dead band under Create New Case; the sheet
now grows to 80dvh and the list fills it instead of a separate 50vh cap.
- Opening scrolls the list (its own box, not scrollIntoView) to the
currently selected case.
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
The Run bar's case picker only loaded /api/cases at page load, so folders
deleted or created on disk stayed listed until a reload. It now refetches
on open and every 5 seconds while open, repainting only when the list
changed and falling back to another case if the selected one was removed.
The Manage tab of Add Case gains a search box filtering by name or path.
Reorder arrows are disabled while a filter is active so a swap cannot
involve a hidden case.
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(cleanup): keep .claude-images while a sibling session uses the same dir
cleanupSession() recursively removes {workingDir}/.claude-images. That
directory belongs to the working directory rather than to the session, and
several sessions routinely share one case directory, so closing one session
deleted the pasted images a live sibling still referred to.
The removal now runs only when no other live session has the same working
directory. A session that is itself being cleaned up does not count as live,
so two sessions of one case closed together still remove the dir.
Split out ahead of the exited-agent sweep for Ark0N/Codeman#446, which closes
sessions unattended and would otherwise make the loss routine.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(session): close sessions whose agent exited cleanly (#446)
Part 2 of Ark0N/Codeman#446. Part 1 records an exited agent as
SessionState.paneExit. A session whose agent the user ended with /exit is
now closed through cleanupSession(), the same path the X button takes, so
finished sessions stop piling up on the board. The lifecycle log records
the reason as "agent exited cleanly (status 0)", and the conversation stays
resumable from the Resume list.
shouldCloseCleanlyExitedSession() in the new pure module pane-exit-sweep.ts
holds the rule. It closes a session only when all of these hold:
- The exit status is an explicit numeric 0 with no signal. An absent status
is how a SIGKILL presents on tmux 3.2a, so it counts as unknown and the
row stays. A non-zero status or any signal also keeps the row, with the
exit code on the tab.
- Two authoritative pane reads agreed on that exit.
TmuxManager.getPaneExitReadCount() counts them, and a failed, empty or
skipped read neither confirms nor resets the count.
- No start, attach or relaunch is running for the pane.
Session.paneLifecycleInFlight covers _setupOrAttachMuxSession(), whose
dead-pane branch revives an exited pane on purpose, and restartCli().
setPaneExit() already scopes paneExit to local mux-backed sessions, so
remote, docker and direct-PTY sessions are never closed.
planRebootRestore() now refuses a record whose persisted paneExit is a
clean exit. That covers an agent that exited just before a reboot, before
the sweep reached it. A crashed agent's record stays eligible, like its row.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(web): show "exited" on the phone overview and desktop home rail (#446)
Part 1 of Ark0N/Codeman#446 taught the tab strip and the rich rail rows
to say that a session's agent has exited. The phone overview and the
desktop home rail still said "idle", beside a green or pulsing dot.
_mobileOverviewExit() in mobile-overview.js is now the one rule for all
three surfaces, and _sidebarRichRow() uses it as well. It changes what a
row shows and leaves the row's state alone, because the state still picks
the section and the sort order. An exited row gets an "exited" pill, a
neutral dot and row accent, and a duration measured from when the server
first saw the pane dead. A pending permission prompt or question still
wins, as it does on the tab.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(cleanup): close the gaps review found in the #446 sweep and image guard
Four fixes from a dual review of Ark0N/Codeman#446 part 2.
- The .claude-images guard compares canonical paths, so a sibling that
reaches the same directory through a symlink keeps it. Its comment used to
say that case only missed a deletion; it caused one.
- A detached session counts as a live sibling. DELETE ?killMux=false removes
it from the server's map while its pane keeps running, so the guard now
reads persisted records too, and exempts only sessions being killed rather
than every session in cleaningUp.
- A session being closed refuses startInteractive() and startShell(). The
/interactive route awaits listener setup before the start, and a start
that raced the close could launch a CLI in a tmux session whose record was
then deleted. A failed close clears the mark again.
- The clean-exit sweep tries each exit once, keyed by session id and the
exit's at stamp, so a close that fails is not retried and logged every
two seconds.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(session): keep a clean exit that lands within 10 s of a pane start (#446)
A CLI that prints a startup error ("not logged in", a bad profile, a config
error) and exits 0 used to lose its tab, and the error with it, about 4 s
after launch. The sweep now keeps any clean exit that lands within
CLEAN_EXIT_MIN_PANE_LIFETIME_MS (10 s) of the last start, attach or relaunch
finishing (Session.paneStartedAt, stamped when _withPaneLifecycle ends). The
row stays as "exited (0)" for the user to read and close.
Verified on an isolated instance: a shell that ran `exit 0` 2 s after start
kept its row, one that exited after 13 s was closed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A single input over MAX_INPUT_LENGTH (64 KiB) was queued for reliable
delivery, refused by both transports (the WebSocket silently, POST with a
400), and never dropped: the client treated the 400 as transient, so the
frame was re-sent every 2 s forever, blocked every later input for that
session, and came back from localStorage on every reload.
- Client: a paste over the frame limit is split into in-limit frames
(never cutting a surrogate pair) delivered in seq order; over 1 MiB, or
an oversized mux write, it is refused with a toast and never queued.
- Client: the POST drain drops a frame answered 400/413; a WS error ACK
drops it too; frames over the limit persisted by an older build are
pruned on load.
- Server: the WebSocket answers an oversized sequenced frame with
{t:'ia',seq,err:'too_large',max} instead of silence (an older client
reads that as a plain ACK and drops it); the POST schema uses
MAX_INPUT_LENGTH instead of a second 100000 limit.
Verified end to end on an isolated instance: a 110 KB paste reached the
PTY byte-identical over both the WebSocket and the POST path, a poisoned
120 KB persisted frame was pruned on load, and a 2 MB paste showed the
refusal toast with nothing queued.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Adds claude-opus-5-5 to the App Settings model picker (base option with
data-ctx="1" plus its [1m] companion row, since Opus 5.5 has a 1M window)
and to the five task-routing selects, mirroring how Fable 5.1 was added.
Co-authored-by: Claude <noreply@anthropic.com>
* fix(sessions): stop pinning the w1-myapp placeholder as Claude's /resume title
Local claude spawns passed the tab name as `--name`. That flag is not only the
cross-session peer name: it is also the prompt-box label, the `/resume` picker
entry and the terminal title, and a pinned title stops Claude generating its own
(`customTitle ?? aiTitle`). So every conversation of a case was listed in
`/resume` as the same `w1-myapp`, and none of them got a generated title. On one
workspace, 34 of 34 conversations spawned with `--name` had no ai-title, while
every conversation spawned without it had one.
Only a name the user chose is pinned now: `Session.cliPinnedName` is the name
when `nameSource === 'manual'`, carried to the builders as a separate `cliName`
so the tab/mux name is untouched. Placeholder and auto names let Claude title
the conversation again.
A rename in Codeman also reaches `/resume`: the new name is appended to the
conversation's transcript as the `custom-title` row `/rename` writes (never
creating the file, never writing an empty title). For a pane spawned without
`--name` this holds immediately; a pane spawned with one re-appends its own
title each turn, so there the new name holds from the next spawn.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(sessions): skip no-op renames and docker sessions when syncing the /resume title
A same-name PUT (the Session Options field saves on blur and recomposes the
unchanged placeholder) no longer flips nameSource to manual or appends a
custom-title row, and docker sessions skip the host transcript scan since their
transcript lives in the container. The skill pages no longer use a w<N>- name
as the peer-name example, and the changeset notes the re-append caveat.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs: record that nameSource decides --name and renames reach /resume
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: codeman-local <codeman@local>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
_loadSessionManagerList() re-projects each unified item into the
history-record shape _buildHistoryItem renders, and dropped these three
fields. The row's own onActivate still read them from the unified item, but
everything built from the record did not: the ⋯ menu's "Resume session"
relaunched a codex row as claude (no mode, no resumeId), a resumed session
lost its conversation id, and Cmd+K rows showed no mode badge. Same class
of bug as the worktree fields the re-projection already carries (#266).
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (da07b38c, db4557d9) with no
corresponding update to the plan doc itself — every checklist still read
Status: TODO and every box unchecked. Brings the doc in line with the tree:
- A new "Status as of 2026-09-22" section up top: what's actually
implemented (verified by grepping the routes/schema/UI, not just trusting
the commit messages), the availability-flag staleness bug found and fixed
in this session (commit 0c77dd0a) with its devbox verification record, and
one real outstanding gap.
- The outstanding gap: a custom CLI created via Phase 5's write API has no
way to actually be launched. The Run menu is static per-mode markup with
no consumer of window.__codemanCliCatalog, so Phase 6's own "create a
custom entry, confirm it can be launched" verify step was never actually
exercised against this. Documented with two candidate fixes, neither
started.
- Each phase's checklist flipped to [x] where confirmed present in the tree,
Status lines updated from TODO to DONE, and the two originally-open
questions (Phase 2's installed source, Phase 5's PUT endpoint shape)
marked resolved against what actually shipped.
No code changes in this commit — documentation only, so a future session
(or the one already mid-flight on a separate checkout of this same branch)
picks up accurate status instead of a stale plan.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs: add the CLI-registry deployment plan and the parked Copilot plan
Both were sitting as untracked scratch files in the master checkout,
never committed to any branch. Moving them here rather than leaving them
loose:
- DEPLOYMENT_PLAN.md is the live tracker for the CLI-registry follow-up
series (PR A #347 merged, PR B #380 merged, PR B2 merged as #458) and
is where PR C (this branch's own CLI-management work) belongs.
- docs/copilot-integration-plan.md is explicitly PARKED, referenced by
name in docs/cli-enable-disable-plan.md's own header as a sibling plan
tracked separately — kept for continuity, not active on this branch.
The other scratch files found alongside these (PRA.md, PRB.md, PR-B2.md
and their review-response counterparts) described PR A/B/B2, all now
merged — deleted from the master checkout as stale rather than committed
anywhere, since their content is superseded by the real merged PRs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* fix(cli-registry): render enabled CLIs in launch surfaces
* test(cli-registry): update frontend branch guard
* fix(test): isolate suite from deployment environment
* fix(cli-registry): revise Decision 4 - claude is toggleable, shell stays permanent
shell/claude were both structurally un-disableable in the original plan
(Decision 4). Revised: shell keeps the hard backend guarantee (it is the
one non-agent mode several code paths assume always exists as a raw-
terminal fallback), but claude is now a normal toggleable entry like any
other CLI.
Safe to do because internal session creation (tmux-manager.ts, session.ts,
Ralph, plan-orchestrator) resolves a CLI via getCli(), which does not
check `enabled` at all - only the Run menu and the HTTP-facing
sessionModeSchema() (new session requests through the normal API) key off
it. Disabling claude therefore behaves identically in kind to disabling
any other CLI: no internal fallback path breaks, it just stops being
offered for new sessions until re-enabled.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): hide shell's toggle entirely instead of greying it out
A permanently-disabled switch next to every other row's working toggle
read as broken rather than intentional. shell now renders no switch at
all - a plain "Always available" label - so there is nothing to click
that could look like it should work but doesn't. Backend guard is
unchanged (UNDISABLEABLE_IDS still refuses shell unconditionally); this
is UI-only.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): sort the Installed CLIs list, installed-first then alphabetical
renderCliList() previously rendered in registry order (each entry's fixed
order field). Now sorts installed CLIs first, then not-installed, each
group alphabetical by label - matches how a user actually scans the list
(what's ready to use, then what needs installing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* style: prettier fixes from the master merge
* fix(cli-registry): install/edit take effect immediately, confirm before install, phone labels
Four gaps found verifying #476 against the #343 review trail:
- Installed or edited CLIs kept reading as missing/stale. Every binary lookup
(the nine per-CLI resolvers and the generic registry one) caches in its own
closure, with a negative-cache backoff of up to 5 minutes, and nothing
cleared them. invalidateCliExecutableResolvers(binaries) now drops those
caches per binary; install (success or failure), create, edit and delete
call it plus invalidateCliResolverCache(id). Before this, a CLI installed
from Settings could fail to launch for minutes, and an edited custom entry
kept launching its old binary until a restart.
- The Settings "installed" badge for a custom entry used a private `which`,
ignoring the entry's searchDirs and the login-shell lookup that spawn and
the Run menu use; it now asks the same generic resolver they do.
- Install ran on a single click. The #343 review asked for auto-install to
sit behind an explicit confirm; the confirm now names the exact command,
which GET /api/clis returns for stock entries only (installCommand).
- The phone Run button showed the two-letter tab badge ("CC", "CX") instead
of the word ("Claude", "Codex"). It uses the registry label again, which is
identical to the old static table for every stock CLI (now pinned).
14 new tests; 9 of them fail against the previous head and pass here.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): address #476 review — safe serialized writes, no id branches, docs
Must-fix:
- registry-writer: start fresh only on ENOENT; refuse (409) a clis.json that
does not parse or has group/world permission bits instead of overwriting it
(isUnsafePermissions now exported from registry.ts)
- mutateRegistryFile(): one promise chain for every mutation, with the
existence/duplicate checks inside the serialized step, plus a unique tmp
name per write
- docs: CLAUDE.md, architecture-invariants, cli-registry (new Settings
section) and api-reference (the six /api/clis routes)
- drop DEPLOYMENT_PLAN.md and docs/copilot-integration-plan.md
Smaller:
- PUT /api/clis/custom/:id keeps the entry's current enabled state when the
body omits it
- runMode setter falls back to the first enabled catalogue entry, not 'claude'
- shell guard keyed on kind === 'shell' (routes + Settings list); stock probe
map shared with server.ts via utils/cli-installed-probes.ts
- stock claude label is now 'Claude Code', so the Run menu / phone overview
label rewrites are gone (doctor row keeps "Claude CLI" via its override)
- welcome buttons are translatable again and read "Run Claude Code" /
"Run Shell"; zh-CN gains "Run Codex" / "Run OMP"
- install: per-id in-flight guard (409) and CODEMAN_* stripped from its env
- fileoverview / CliEnableSchema comments no longer say stock-only
- test-env isolation changes moved to their own PR
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* test(cli-registry): pin the #343/#347 findings #476 makes reachable
A CLI toggled or created through the routes is accepted or rejected by
CreateSessionSchema with no restart (#343 finding 2), and a custom CLI created
through the API renders a real local, remote and docker launch command
(#347 finding 5: no more `cd <path> && undefined`).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(self-update): stop a stalled status from blocking every later update
A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.
- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
is older than the stale window; applied on every read (start + status
poll) and persisted. The live updater heartbeats every few seconds, so a
running update never trips it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down
On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.
- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
then SIGKILL it. tmux sessions live outside the server and survive.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
- docs/wiki/The-Dashboard.md: the tab-appearance table gains the exited
state (muted dot plus an `exited (137)` badge) and explains the bare
`exited` variant.
- The detailed sidebar and rail no longer pair the muted dot with an "idle"
pill: an exited session's pill reads "exited" (neutral styling) and its
since stamp measures from the observed exit. This is a label override on
the row model, not a new state, so SESSION_ACTIVITY_RANK and the home
screen order are untouched, and a pending alert still keeps its own pill.
The row signature includes the flag so the incremental path repaints it.
- The exited badge is aria-hidden like its sibling badges, and the exit is
appended to the tab's aria-label in both render paths through one helper.
- test/tmux-manager.test.ts re-adds the junk-trailing-field parser case
against parsePaneRows.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- test/setup.ts strips CLAUDE_CONFIG_DIR (pinned in test-env-isolation), so
transcript-fixture tests such as session-custom-model-restart no longer go
red on a machine that exports it for a separate Claude account (#255).
- The vanished-tmux-session branch of _setupOrAttachMuxSession() relaunches
the CLI through createSession() just like a failed respawn, so it now takes
the same resume pin. A genuinely new session is unaffected.
- After a dead-pane respawn of a fallback-chain CLI, _claudeSessionId names
the conversation the walk actually pinned instead of the chain tail, which
the walk may have passed over for lack of a transcript.
- _claudeConfigDir() trims the override like claudeProjectsDir() does.
- The remote-reattach test is labelled as documentation, since the pin
builder's own remote guard would make it pass either way.
- CLAUDE.md: the create-path pin persists through toState() as
resumeSessionId, and the end of the walk adds no pin rather than clearing
the launch seed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- The TERMINAL DROP crash-trail line moves behind the scheduler's debounce
guard, so it is written once per window rather than once per dropped
frame. At the server's 8ms batching, one second of drops evicted the whole
50-entry trail, including the recovery lines that explain it.
- A refresh that failed at the capture fetch deadline now returns
'deadline', and the scheduler does not retry it: that is a stalled link,
not contention, and each retry was another ?full=1 capture waiting out a
deadline of up to two minutes. The early-return retries are unchanged.
CLAUDE.md and the code comments no longer claim every skip reason is
transient contention.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- While another device holds the pane width (_paneWidthRefused), a resize
now asks for the container's width without applying it locally
(_geometryForResizeRequest: rows follow the container, columns stay at
the PTY's). Fitting first re-wrapped the whole buffer to the container
and back on every 30s mobile retry, and throttledResize ran the
scrollback clear for a resize that brings no redraw. selectSession
clears the flag, since it belongs to the previous pane. New unit tests
run the real mixin against a fake terminal and fail without the fix.
- Session seeds _ptyCols/_ptyRows at spawn (_notePtySpawnGeometry), so a
reattached pane reports its tmux window's real size through ptyGeometry.
- Session.resize's declined-branch comment names ptyGeometry, not the
deleted ptyCols/ptyRows getters.
- Delete the dead terminalGeometryAgrees() and its window export.
- test/xterm-private-api.test.ts header: it pins the exact locked version,
so any bump fails, not only a major.
- The main-terminal fit sweep also matches fitAddon?.fit?.(), and
CLAUDE.md names the modules it actually covers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Clone Repo clears the credential helpers for a non-admin, but a non-admin's
Docker case with credential seeding on still receives a copy of the server
account's gh/az sign-in when the agent-image switches are on, the same as
the Claude and Codex credentials. Say so in the multi-user notes so the docs
do not read as a stronger guarantee than they are.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose
label after a plain `down`, instead of `down --volumes` (which also takes
any volume an override file declares while the message named two).
`down --volumes` remains only as a warned fallback when the project name
cannot be resolved.
- Report a failing first `docker compose config --format json` call with a
clear error instead of exiting silently under `set -e`.
- Filter empty label lines in the collision guard so an unlabelled container
cannot hide a real collision; name the moved-checkout exit in its error.
- Comments no longer cite a guard or incident in Start-Codeman.sh that does
not exist; the README states the real gap (a Node base-image bump leaves
codeman-node-modules stale because the lockfile did not move).
- docs: Update-Codeman.sh in the docker-self-update.md short-version table
and a mention in docker-compose.md; "Major updates" moved under "Updating"
in docker/README.md.
- test: smoke test covers the new sequence, the config failure and the
empty-line case; quiet stdio; @fileoverview names the fourth concern.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- session.ts: a pane capture that fails now CLEARS the watching label
(and emits watchingChanged so pages drop the badge) instead of keeping
the last one, so a failed capture degrades toward an alert rather than
pre-acknowledging the next real idle prompt. Test updated; invariant
noted in architecture-invariants.
- approvals-ui.js: the header bell counts only unacknowledged items
(pendingApprovalsCount), matching codeman tui's pendingApprovalCount();
pinned in watching-no-alert.test.ts.
- mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back
onto MOBILE_OVERVIEW_PILL_LABEL.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- stock.ts: claude is no longer the only entry declaring transcriptGutter;
codex declares it too.
- architecture-invariants: the strip applies when the session's CLI declares
a margin (not detection), and a note that it keys on the session's launch
mode, not on what is running in the pane (a claude pane dropped to a shell
still loses up to two columns; copyStripMargin is the escape hatch).
- render-index-html test: the gutter map is injected for a solo
/session/:id render as well.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- mobile.css: gemini and antigravity run/gear rules get `!important` like
pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent
while the body takes the mode colour (two-tone button on the default skin).
- test/skin-themes.test.ts: static guard that every run mode with a base
`.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the
`html:not([data-skin="og"])` block; ids are derived from the stylesheet.
- stock.ts: grok's accent comment names zinc-300 (border/badge colour);
gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted
as the one exception to the border-colour method.
- types.ts: "(below)" -> "(above)".
- docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CLAUDE.md loads into every session, and its Architecture section had grown
feature write-ups (history, measurements, rationale) that belong in
docs/architecture-invariants.md per the file's own header. Each long block
now keeps what the feature is, where it lives, its setting/default and the
rules that prevent real bugs, and links to its invariants section. Everything
removed was moved there: 29 new sections, extra facts appended to the
existing ones.
Also: hard-coded counts (SSE events, route handlers, module/file counts,
device profiles) replaced by pointers to the source of truth, and the
Debugging commands fixed to use the codeman tmux socket and HTTPS for prod.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Addresses the review on #472.
- CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv`
(CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts
skips a store unless that variable is exactly `1`, read at container
create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL
cache no longer copies them into every case container. Tests: the default
environment seeds neither even with the files present, and each store
follows only its own switch.
- Multi-user mode: a non-admin's Clone Repo clone and preflight run with
`git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the
subcommand), so the server account's helpers are never lent to them.
Verified against a real private repo that it also clears the URL-scoped
credential.<url>.helper entries, and that public clones still work.
Tests: the argv in test/git-clone.test.ts, and the route decision
(non-admin cleared; admin and single-user kept) in
test/routes/case-clone-credential-helpers.test.ts.
- Docs: recreate the case container to pick up seeds (docker/README.md,
Docker-Cases wiki, docker-cases.md); the multi-user behaviour in
docker/README.md and security-architecture.md; "functionally unchanged"
instead of "unchanged" for an image built with both switches off
(server.Dockerfile comment, README, changeset).
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
CONTRIBUTING says releases are handled by the maintainer via changesets after
merge, and every `.changeset/*.md` on master was written by him or by the
release bot — including the ones covering other people's pull requests. The
summary this file carried moves to the pull-request description, where it is
the maintainer's to reuse or rewrite.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claimed after a manual test that a session comes back from a server restart
without its badge until it next produces output. Measured instead of assumed,
and it is wrong: a codex session with a background terminal still running had
its label back within about 20 seconds of the restart, with no input from
anyone. Reconciliation re-attaches the pane, the attach repaint carries the
composer glyph, the idle confirmation arms on it, and the probe re-reads the
label — the ordinary path, doing the ordinary thing.
What produced the false claim was a session whose monitor had simply expired
while it sat there. Its footer carries no chip, so `watching: null` was the
right answer and there was nothing missing to restore.
Recorded at the field and in the invariants, because the shape of this invites
exactly one wrong fix: a polling timer to keep a value fresh that the pane
already refreshes by itself.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Add Case -> Clone Repo could only reach public repositories in the Docker
deployment. This lets a deployment opt in to the GitHub CLI and the Azure
CLI (+ azure-devops extension) as git credential helpers. Codeman itself
still collects no credentials.
- server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH /
CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the
build). Off leaves no apt repository, package, extension, helper script
or credential entry, so a default build is unchanged. On installs from
the vendors' apt repositories and configures system gitconfig helpers:
github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com /
*.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID
token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT).
A helper whose CLI is not signed in prints nothing, so a private clone
still fails fast.
- The extension lives in AZURE_EXTENSION_DIR outside HOME
(/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0
group-writable in the agent image).
- Hosts turn them on in docker-compose.override.yml: `build: args:` for the
server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the
agent image. build-agent-image.mjs and the in-app auto-build share one
env -> ARG table (pinned by the parity test) and pass nothing when unset.
docker-compose.yaml is untouched; .env.example only gains a comment, so
the self-updater's environment gate sees no new keys.
- Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and
the az sign-in files from ~/.azure per file, read-only, like pi/grok.
- The Clone Repo AUTH_REQUIRED message says how to sign the server's git
in instead of claiming private repositories cannot be cloned.
- Docs: docker/README.md "Private repositories", docker-compose.md,
docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki
pages, security-architecture.md, architecture-invariants.md, changeset.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
Review fixes for #469.
The Ctrl+C branch cleaned the selection to decide whether to copy and then
passed that cleaned string to copyTerminalSelection(), which cleans again. The
trailing trim is a fixed point, so that was safe until this PR; the margin
strip is not, because it takes the lesser of the declared width and the run
every line shares, so a second pass takes up to `margin` columns more. The
branch now gates on the cleaned string and hands the raw one on. Verified in
chromium with a real drag, a real Ctrl+C and a real clipboard read on a live
claude pane: an on-screen ` fix(terminal): trim it` reaches the clipboard
as ` fix(terminal): trim it`, and reverting the branch reproduces the
reported ` fix(terminal): trim it`.
Pane B of a split resolves its own width. `_cliGutterColumns()` and
`_normalisedSelectionRange()` take the session and the terminal to read,
defaulting to the primary pane's, so Pane B looks its own run mode up instead
of keeping a margin Pane A drops on the same keystroke. Verified live with two
claude panes open side by side.
A detached session window (`/session/:id`) receives the gutter map. The
injection sat inside the block that skips the run menu's payloads for a solo
window, so the toggle worked in the main window and did nothing in the popup on
the same device. It needs no availability probe, so it moved below that block
and the solo window still carries none of the payloads it skipped before.
The settings description said the width is measured and named Codex as exempt.
Nothing is measured, and Codex is one of the two panes that are stripped.
docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle
carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing
else" one sentence before the leading-margin rule, and both it and
docs/architecture-invariants.md record that the strip is not idempotent.
Two round-trip tests run on a mode that declares a gutter, which the existing
copyTerminalSelection cases could not, since they all use the harness default
mode that declares none. The Ctrl+C branch itself is pinned at the source,
because it lives inside initTerminal's attachCustomKeyEventHandler closure over
a real xterm the vm harness cannot build. Both pins fail on the reintroduced
bug. Gate: 7865 passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`_onSessionTerminal` drops an incoming frame once the app-owned render queues
already hold 128KB. That is the right call — the alternative is an unbounded
backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced
cursor is muffled text (#464). The drop was only half of it.
The recovery was a fire-and-forget timer: it nulled its own handle and then
called `_onSessionNeedsRefresh()`, which opens with four early returns. Two of
them — a buffer load already in flight, a refresh already owning this session —
are MOST likely to be true during exactly the output burst that caused the
drop. So the recovery was skipped precisely when it was needed, with nothing
left to retry it, and the dropped bytes were never replayed.
`_onSessionNeedsRefresh` reports whether it actually repainted now, and
`_scheduleDroppedOutputRecovery` re-arms while it has not. Bounded by
`DROP_RECOVERY_MAX_ATTEMPTS`, because every reason the refresh can be skipped is
transient contention that clears in seconds and a permanently failing refresh
must not become a loop against the API; giving up at the cap leaves exactly what
the old code left, so the floor is no worse than before. The same 2s debounce
still collapses a burst of drops into one attempt.
This is the principle Ark0N established reviewing #431 for the WebSocket
output-gap marker — only a repaint that actually happened settles the recovery —
applied to the one recovery path that still trusted a timer having fired.
The retry decision is a pure function in constants.js so the gate can reach it,
and the scheduler itself is driven from app.js under a fake clock. The retry
case and the no-retry case only pin the fix AS A PAIR: either alone passes
against something wrong, one against the old fire-and-forget timer and the other
against retrying forever. Checked by reverting app.js to the old shape, where
three of the twelve fail.
Two harness details that would otherwise have made the tests lie. The vm context
baked in the real `setTimeout`, so `vi.useFakeTimers()` could not reach the
scheduler and every case reported zero calls; it delegates lazily now. And
app.js reached `CodemanDroppedOutput` as a bare global, which resolves in a
browser but not in the vm — worth fixing beyond the test, because that call sits
inside a timer where a ReferenceError is swallowed and would take the recovery
with it. It reads through `window.` like terminal-ui.js does with its own
constants.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Reported from a manual test: a Codex session went on showing the watching badge
after its background terminal had finished. The server was right and the page
was stale — `Session.watching` changes while the session's status does not, and
nothing broadcast it.
The label is usually SET on the idle transition, which broadcasts anyway, so the
badge always appeared correctly. It CLEARS when the work ends, and a CLI can end
background work without taking a turn: codex repaints its background-terminal
row away and stays idle, so `_confirmIdle()` concludes without emitting `idle`
(that emit is guarded by `wasWorking || isInitialReady`) and no other event
fires. Every open page kept drawing a badge the server had already dropped.
`_readWatching()` now emits `watchingChanged` when, and only when, the label
really changes, and the wiring pushes the session state on it. No new SSE event:
the badge reads off the session payload every surface already has.
A/B measured on an isolated beta with the page loaded and then left untouched.
Without this commit the server dropped the label at t+50s and the page still
showed the badge at t+100s; with it, page and server cleared in the same
ten-second window.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without
waiting outlives the turn exactly as a background terminal does — the sandboxed
process was still running — and codex shows nothing for it. The last rows of the
pane are the composer and the status line, and `Sub-agents running` belongs to
the on-demand `/subagents` panel rather than to the row above the composer.
So there is no second codex label to add. A codex session waiting on a sub-agent
reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and
simply leaves that one kind of quiet unexplained until codex pins a row for it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief)
found the trust boundary weaker than the comments around it claimed. Eleven
findings, all applied.
The two blockers were both about who can write the row the label is read from.
Claude's window covered two rows, and the second one is the status line, whose
command a session running with permissions bypassed can write into its own
`.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of
its own and silence its own idle alert. The default window is one row now, which
is the footer and nothing else, and the constant says why. Separately, the label
reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML,
which is an injection sink for any config-supplied pattern whose capture group is
permissive; it goes through escapeHtml() like every other untrusted string in
that file.
The Codex entry could not be fixed the same way, and now says so. Its row is
third from the bottom only while a terminal runs; with none running that slot
holds the last row of the transcript, so matching the complete row (with the
`/stop to close` tail, window narrowed to three) raises the bar without closing
it. What contains it is `hooks: 'none'`: no hook event from a codex session
reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an
alert. The registry comment, `docs/cli-registry.md` and the test all state that
rather than claiming a guarantee the code does not have.
Also from the review: the TUI header badge no longer counts an acknowledged
item, which was the same gate the classifier fix already went through and was
wrong for human acknowledgement too; the TUI approval card reads the quiet
reason and drops to a new `info` tone instead of asking for a reply; the badge
carries an aria-label, because the phone it was built for has no hover target;
the schema refuses `watchingLines` without a `watchingLine`; and the pattern and
its window are resolved together rather than one memoized and one not.
Documentation moved with it. The mechanism now lives in
`docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer,
`docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped
buzzing, and both that page and the changeset name the limitation neither did
before: a question asked in plain prose is not a dialog, so it is silenced along
with the false alarms while background work runs.
Verified live again after the narrowing, on an isolated beta: a Claude session
reported `1 monitor` and took its idle prompt acknowledged, and a Codex session
reported `1 background terminal` against the full-row anchor.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.
A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.
The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.
The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.
Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>