Compare commits

...
Author SHA1 Message Date
Codeman maintainer 00fb3b0908 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:18:23 +02:00
Codeman maintainer ffccde4f7d fix(cli): codeman status probes the running server (#230)
Reported by @mtiller.

`codeman status` runs in its own fresh process, and reported THAT process's
always-stopped Ralph loop under a bare "Status:", which reads as "the web server
is down" while the service is running fine and agents are reachable. It now probes
the real server first (`CODEMAN_API_URL`, else https then http on the local port,
overridable with `--url`) and reports reachability, version and live session
state. Any HTTP answer proves the server is up, including a 401 from a
password-protected install. The Ralph loop keeps its own `codeman ralph status`.

This complements `codeman web --status` from the daemon work: that answers "did I
start a daemon", this answers "is a server running at all", which is what the bare
command was already being used for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:24 +02:00
Codeman maintainer bec3da3d31 fix(ui): a described session tab shows just the description (#232)
Reported by @mtiller.

A session named `w2-foo-bar: some description` rendered both halves on the tab, so
the generated id ate the width that the part the user actually chose needed. The
tab now shows the description alone and the `w<n>-<case>` id moves to the tooltip,
where it stays available without being read every time. It is still shown in the
session settings modal. Undescribed tabs are unchanged.

`aria-label` deliberately keeps the FULL name, so screen readers still get the id.

Also fixes a re-render loop this exposed: the incremental update compared
`nameEl.textContent` against the full name, which for a described tab never
matched, so those tabs re-rendered on every pass. The compare now targets the
display label.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:24 +02:00
Codeman maintainer 6e89eb9ec1 fix(web-tabs): bound time-to-headers, not the whole proxied exchange (#237, #238)
Reported by @DodgyBadger.

#237: the proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which
bounded the ENTIRE exchange rather than the wait for response headers. A dashboard
endpoint doing model inference, and any actively streaming response, both died at
30s as a generic 502 that Codeman never logged, so it read as an intermittent
network error. The timeout now bounds time-to-headers only and is cleared the
moment headers arrive, so a slow endpoint and a long stream both survive. The
default moves to 300s because "the app is thinking" is normal for the dashboards
people proxy; abandoned upstreams are reclaimed by the client-hangup abort rather
than by this value.

A browser that navigates away mid-request now aborts the upstream fetch, guarded
by `writableFinished` for the same reason as `abortOnClientHangUp` in
session-routes: `close` also fires after a completed response and must not abort
anything. Header timeouts are logged as a warning with a sanitized identity
(method plus origin plus path, never the query string, which can carry the
dashboard's tokens), and a client hangup is deliberately not warned since nobody
is listening and it would read as the dashboard being broken.

The WebSocket handshake keeps its own 30s budget
(`CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`), decoupled from the request timeout:
a handshake is connection establishment, and waiting minutes on one only delays
the browser's reconnect logic.

#238: the web-tab guide covered sandboxed dashboards having no cookies, but not
cookie authentication in front of Codeman itself (Cloudflare Access and similar),
where a sandboxed frame's asset and API requests carry no auth cookie, bounce to
the login provider, and leave the embedded app looking unstyled or broken while
trusted mode works. Documented, and the Test button's result now says it probes
server-to-upstream reachability only, not how the page behaves in a sandboxed
frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:23 +02:00
Codeman maintainer 94aa53c65b chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:38:40 +02:00
Codeman maintainer e88b971bb7 feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.

Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
  Case names resolve through linked-cases.json first, mirroring the
  server's resolveCasePath(), so a case linked in from outside
  ~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
  hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
  skill is never touched, and a symlinked skill dir is refused (this
  repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
  ports/config-port.ts, server.ts, session-routes.ts (add-only injection
  on Claude session create and quick-start), plus the App Settings toggle.

Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
  Shell state does not survive between agent tool calls, and an undefined
  is_self exited 127, firing the `||` branch and deleting the caller's own
  session with the one guard bypassed. The request now lives inside the
  guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
  tool call, so the documented resend-identical-request loop stopped being
  a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
  workers. It returns clean transcript text; the terminal scrape it
  replaces returns a wall of TUI repaint noise. Its transcript flush lags
  the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
  yielded the literal session id "null" and burned the whole readiness
  budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
  corrected the hooks-config comment that claimed a toggle-off sweep
  exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
  and that caseName resolves linked cases, so a generic name can land a
  worker in a real repo.

Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:30:15 +02:00
Codeman maintainer 8406c497e2 fix(terminal): stop forwarding the wheel to codex, it ignores SGR reports
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.

Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.

_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.

Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:22:06 +02:00
Codeman maintainer 40b4aba043 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:35:59 +02:00
Codeman maintainer 4b44988bfc test: give daemon-control tests a unique port (3212 was already taken)
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:22:01 +02:00
Codeman maintainer 316d0a4c82 Merge pull request #233 from Lint111/feat/hooks-config
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:21:52 +02:00
Ark0N 1184720648 Merge pull request #239 from Ark0N/feat/daemon-mode
feat(cli): codeman web -d and codeman service install (#231)
2026-08-09 02:19:50 +02:00
Ark0N b067aad9b6 Merge pull request #235 from Lint111/feat/deferred-terminal-flush
fix(terminal): drain deferred output without a wake event
2026-08-09 02:19:32 +02:00
Ark0N 19a3d7c773 Merge pull request #234 from Lint111/feat/ai-checker-stderr
fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
2026-08-09 02:19:10 +02:00
Codeman maintainer 085f4acb60 feat(cli): codeman web -d and codeman service install (#231)
Two ways to keep the server running, split by how long it should last.

`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.

`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.

Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.

The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:34:55 +02:00
Codeman maintainer d26f26fe34 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:18:03 +02:00
lior 091df2b6d8 fix(terminal): drain deferred output without a wake event 2026-08-08 23:00:36 +03:00
lior 5f775b1ab1 fix(hooks): guard subagent stops and rewake from the parent transcript
Two defects in the background-task hook scripts.

SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.

The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.

The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.

12 tests fail on unmodified master, e.g.
  expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
  expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
2026-08-08 22:31:38 +03:00
lior da51193264 fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".

stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.

Two tests, both failing on master:
  expected 'export PATH="…' to contain ' 2> "'
  expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
2026-08-08 22:30:39 +03:00
Codeman maintainer fa18eeef35 feat: tab action icons on the active tab only, middle-click closes tabs
Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.

Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.

The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer a9f26bd03a feat: fixed-width tab hover with sliding title, pop-out button now opt-in
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).

The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer 8dc8b164a7 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:48:12 +02:00
Ark0N 2524759655 Merge pull request #229 from Lint111/feat/keyboard-viewport-settle
fix(mobile): coalesce keyboard viewport settling
2026-08-08 12:51:04 +02:00
Codeman maintainer 1f164bc8d2 fix(mobile): only arm the viewport settle on a real keyboard transition
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 12:06:08 +02:00
lior 0a1439b1e9 test(mobile): make the coalescing test actually exercise the settle path
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.

Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.

Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.

Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).
2026-08-08 08:41:04 +03:00
lior 66abe6c70a fix(mobile): coalesce keyboard viewport settling 2026-08-08 08:13:58 +03:00
Codeman maintainer fa1700da5b chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:38:35 +02:00
Ark0N aed1e59ee3 Merge pull request #227 from Ark0N/fix/scrollback-205-round2
fix(terminal): scrollback round 2 for #205 (re-pull downgrade guard, PageUp fallback, CLI version probe retry)
2026-08-08 01:36:03 +02:00
Ark0N 7f6d18b398 Merge pull request #226 from christianhaberl/fix/input-loss-on-failed-delivery
fix(api,ws): an input whose delivery fails can be retried instead of being lost
2026-08-08 01:31:10 +02:00
Ark0N 52571c7fd4 Merge pull request #225 from christianhaberl/fix/bound-the-process-tree-walk
fix(mux): bound the process-tree walk — unbounded pgrep recursion can take a machine down
2026-08-08 01:31:00 +02:00
Ark0N cb95a8562c Merge pull request #224 from christianhaberl/fix/raw-writehead-drops-security-headers
fix(http): raw writeHead routes drop every header the security hook set
2026-08-08 01:30:47 +02:00
Codeman maintainer 3cb7e30636 fix(ui): scope the wheel-opt-out tooltip's paging fallback to Claude
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:29:48 +02:00
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00
Claudia 9d27cc0bab docs: merge the stacked doc comments the previous commits left behind
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.

- write() had two: the original description with @param and @example, then a
  @returns-only block added on top, which dropped the params and examples from
  hover. Merged into one. The @returns wording is also honest now — write() still
  discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
  and its declaration, leaving that function undocumented on hover and the doc
  attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:29:46 +02:00
Claudia 84132d3025 fix(ws): pin the withheld ACK with a test, and correct the changeset
Two blockers from the pre-submission gate, both reproduced before fixing.

1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
   code says the opposite in as many words ("NOT an error response,
   deliberately"), the commit message says response codes are unchanged, and the
   test asserts the 200. It was a leftover sentence from an earlier iteration that
   would have shipped into the CHANGELOG announcing an API contract change that
   does not exist — and errorCode values are SemVer-relevant per
   docs/versioning-policy.md.

2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
   master left all 9 tests green, while the commit message sells "plus the whole
   WebSocket path" as part of the fix. Three tests added against the real WS
   route — ACK on delivery, ACK withheld and seq re-opened when the write did not
   land, and a deduplicated frame still ACKed so the client can drop it. Verified
   the other way round: with ws-routes.ts reverted, the middle one fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:52:44 +02:00
Codeman maintainer cc163792e5 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:43:47 +02:00
Ark0N 6f1ff17ccc Merge pull request #223 from Ark0N/fix/scrollback-shell-alt-screen
fix: terminal scrollback overhaul for shell and CLI sessions (#205)
2026-08-07 13:42:47 +02:00
Codeman maintainer f262b8cb69 feat(terminal): gentler glide start and fractional wheel accumulation
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:36:27 +02:00
Codeman maintainer 5f2b491d99 feat(terminal): ease-out smooth scrolling for the local wheel path
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:30:34 +02:00
Codeman maintainer c067167dbc fix(terminal): take the wheel in capture phase; xterm's scroller is deaf after reset
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.

The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.

Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:03:17 +02:00
Codeman maintainer ad2ca9b575 docs: record the #205 scrollback mechanisms and the shipped fix plan
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:42 +02:00
Codeman maintainer a7a1cef3d6 fix(session): probe the Claude CLI version over ssh for remote sessions
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:41 +02:00
Codeman maintainer a1d7ec02e9 fix(terminal): forward touch scrolls to the CLI transcript on mobile
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:25 +02:00
Codeman maintainer dfa43928af docs: record the scrollback analysis and its measurements for #205 2026-08-07 04:33:15 +02:00
Codeman maintainer adbb74cd5a fix(terminal): keep the CLI's input box pinned when scrolling with the wheel
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.

_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:

- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
  history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
  REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
  furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
  non-empty row, so any session with trailing blank rows was left off-bottom
  and every later wheel event went local without the user ever scrolling.

Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.

Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
2026-08-07 04:27:16 +02:00
Codeman maintainer eb8d11ffc3 fix(terminal): restore shell scrollback, recover history lost to tmux repaints
Four fixes for the scrollback reports in #205 (plus its follow-up comment).

1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
   ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
   (\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
   to claude/codex/gemini, so it reached the browser verbatim. In the alternate
   buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
   and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
   which readline receives as shell history navigation. Both reported symptoms,
   one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
   modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
   own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
   the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
   never forwards a pane's alt-screen toggles to its client, it repaints;
   captured from a real attach, vim/less/htop emit zero.

2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
   onto tmux's history, and tmux repaints the pane rectangle instead of emitting
   linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
   scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
   same 60 lines emitted slowly added all 60. Scrolling up at the top now
   re-pulls the full tmux scrollback and holds the user's place. Verified
   end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.

3. The full-scrollback replay was gated on a single "first load after page load"
   flag, which whichever session auto-selected consumed, so every other tab
   started with one visible frame. Now tracked per session.

4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
   per notch) scrolled one line where Chrome scrolls four or five, and capped
   the forwarded SGR report at one tick. Line and page deltas are now converted,
   and a pure horizontal swipe no longer falls through to a phantom -1.

Analysis and measurements: docs/scrollback-issues-analysis.md
2026-08-07 04:06:54 +02:00
Claudia ebfcac6ad1 fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.

When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.

- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
  is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
  client redelivers.
- `Session.write()` reports whether it reached a PTY.

Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.

What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.

9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:36:33 +02:00
Claudia 2e69e28e71 fix(mux): bound the process-tree walk — it can take a machine down
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.

Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.

- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
  with a single-flight guard. Async matters: under the same procfs pathology,
  `execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
  which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
  visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
  both reporting when they truncate. Pure so the regression tests can exercise the
  shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
  re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
  would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
  bounded, so a wedged `ps` cannot stop killSession from reaching its
  process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
  fresh; a truncated table would make whole subtrees invisible to the kill path.

13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:34:47 +02:00
Claudia 1a32e63765 fix(http): raw writeHead routes lost every header the security hook set
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.

The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.

Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.

Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:31:50 +02:00
Codeman maintainer d41f28bc14 docs(docker): warn that a plain agent-image rebuild keeps stale CLIs
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.

That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.

Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.

No changeset: docs-only, rides the next release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 09:02:28 +02:00
Codeman maintainer 322f21ef9f docs(extending): scope the "no sandbox" claim, point at Docker cases
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:

- Integration code cannot be sandboxed by Codeman because Codeman never
  launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
  documented isolation story.

Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:46:22 +02:00
Codeman maintainer c2d973cb2d docs: link the integration guide from both READMEs, fix three inaccuracies
Adds a pointer to docs/extending-codeman.md at the end of the API section in
README.md and README.zh-CN.md, so the guide is reachable from where people
read about endpoints rather than only from CLAUDE.md.

Reading the README's programmatic guide alongside the new page surfaced three
errors in it, all now fixed:

- POST /api/sessions/:id/input takes `useMux`, not `useScreen`. The latter is
  a legacy name that no longer appears in the schema.
- The page told integrators to send `\r` to submit. With `useMux: true` the
  server delivers text and Enter as two separate writes (writeViaMux does
  send-keys -l then send-keys Enter), so appending `\r` is wrong.
- "Unwrap the envelope" was incomplete: a few legacy GETs put the payload at
  the top level, so the advice is now `body.data ?? body`.

Also cross-references the README's programmatic guide, which covers the
in-session case (CODEMAN_MUX, CODEMAN_API_URL, CODEMAN_SESSION_ID,
CODEMAN_HOOK_SECRET_FILE) that the new page deliberately does not duplicate,
and documents the optional clientId/seq exactly-once fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:34:58 +02:00
Codeman maintainer 84e31c0ee1 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:31:02 +02:00
Codeman maintainer 0d0b772619 feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.

Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
  mode 'antigravity' died on command-not-found. agy is not on npm, so it
  gets its own installer step. --dir /usr/local/bin is load-bearing: the
  default $HOME/.local/bin resolves to root's home at build time and is
  unreachable by the `agent` user the container runs as. Verified inside
  codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
  ~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
  present like the other CLI buttons, with a cyan identity matching the
  toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
  it as a satisfying AI CLI, and recommends it over Gemini in the install
  hints. Detection only, no new auto-install path.

Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
  opencode/codex/gemini when the code has included antigravity for a
  while, said "all three modes", and omitted ANTIGRAVITY_ from the env
  prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
  remote-sessions' RemoteCommandMode were all stale.

Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.

test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.

Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.

isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:22:29 +02:00
Ark0N bfff20a093 Merge pull request #216 from shenlvkang-collab/fix/response-viewer-brief-format
fix(web): align brief Response Viewer formatting
2026-08-06 07:22:13 +02:00
codeman-local b982c5d0e0 fix(web): align brief response viewer formatting 2026-08-06 10:16:15 +08:00
Codeman maintainer f50c922240 docs: add extending-codeman.md, the third-party integration guide
Codeman has no plugin runtime by design: running third-party code inside
the process that spawns agents, on a server people expose over a tunnel,
would trade away the security posture that is a reason to use it. But it
already has four extension seams that work from any language with nothing
installed, and they were undocumented.

Documents web tabs (render your own UI as a tab), the SSE event channel
(react when an agent needs you), the HTTP API plus the codeman CLI (drive
it from a script), and hook events. Every endpoint, schema field, event
name and header in the page was read from source and then verified against
a running instance, including the localhost-only CORS behavior and the SSE
framing the example depends on.

Also corrects a stale line in CLAUDE.md: it claimed the HTTP/SSE API was
internal/unstable, which contradicts docs/versioning-policy.md, where the
API under /api/v1 was finalized as part of the stable surface for the 1.0
cut. No new stability commitment is made here; the page makes an existing
one discoverable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 02:03:39 +02:00
Codeman maintainer de5b048c3f docs(vm): VM cases plan + Apple virtualization stack reference
Two design/reference docs for the planned native-macOS VM isolation tier
("VM cases"), a location overlay on cases in the same shape as Docker and
remote-SSH cases, never a sixth SessionMode. Nothing is implemented; both
docs are marked PLANNED and are blocked on macOS 27 GA.

- vm-cases-plan.md: the Codeman-side design and phased plan. Swift helper
  CLI, DiskImageKit base + per-case overlay, sessions riding the existing
  remote-SSH machinery, VirtioFS workspace at the same absolute path, and
  seeded credentials, each mirroring an established Docker-cases rule.

- vm-subsystem-apple-stack.md: what the Apple stack actually provides,
  measured on the 27 beta rather than inferred from the WWDC session. Of
  note: the 2-concurrent-macOS-VM cap is a kernel quota (refused at 39%
  free RAM, so more hardware does not help), DiskImageKit has no flatten
  API so exports must ship the layer chain, and a macOS guest renders
  nothing without an attached view in an unlocked host session.

No credentials, hostnames, tailnet addresses or account names in either
file; every host/guest reference is a placeholder.

Also joins a table row that a stray blank line had split off into its own
malformed table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 01:28:17 +02:00
Codeman maintainer 12a5f5919e chore: version packages 2026-08-05 22:36:51 +02:00
Codeman maintainer ecd3f3f32a harden(history): exclude automated transcripts by SDK shape, not by "not cli"
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.

Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.

Test fails against the pre-fix line and passes after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 21:47:06 +02:00
Ark0N c19d884a51 Merge pull request #215 from timkjr/fix/past-sessions-history-quality
fix(history): three Past Sessions data-quality bugs (automated-session noise, cross-contaminated previews, blank restart-heavy rows)
2026-08-05 21:44:47 +02:00
Ark0N 22e77a1827 Merge pull request #214 from timkjr/fix/mobile-overview-run-gating
fix(mobile): gate the phone overview's run picker on CLI availability
2026-08-05 21:44:42 +02:00
Ark0N b641560040 Merge pull request #203 from shenlvkang-collab/contrib/claude-viewer-session-pin
fix(web): pin the Claude response viewer to the pane's own conversation
2026-08-05 21:44:37 +02:00
timkjr 8300c15cbd fix(history): entrypoint detection was first-field-wins, plus a two-tier head read
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).

Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 09f5f28017 docs(test): correct an overclaiming comment in the tail-fallback regression test
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 251706be3b harden: scope entrypoint detection to message lines, add fallback coverage
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:

1. extractTranscriptEntrypoint() scanned any line containing the
   substring "entrypoint", not specifically the first "type":"user"/
   "type":"assistant" message line (unlike its sibling
   extractFirstUserPrompt, which does scope to type). A transcript
   that started under an older Claude Code version (no entrypoint
   field) and got resumed under a newer one mid-conversation could
   pick up the field from a much later message than the true first
   one, misattributing the session's origin. Scoped it to match.

2. Added a regression test proving the tail-read fallback still
   engages correctly when bookkeeping accumulation exceeds even the
   new 128KB head window, not just the 16KB it previously blanked at.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 18b473f0e4 fix(history): raise the transcript head-read window to fit restart bookkeeping
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).

Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 a2aed38073 fix(unified-sessions): stop the firstPrompt workingDir backfill from cross-contaminating history rows
COD-140's backfill was meant to cover live/persisted rows whose Codeman
id doesn't match an on-disk transcript UUID, guessing from the newest
transcript in the same workingDir as a last resort. It was also firing
for pure history rows whose OWN transcript scan already ran (and
genuinely found nothing, e.g. an oversized first message) -- those got
silently backfilled with the newest OTHER session's opening line from
the same directory. Not a blank row, but actively wrong: old sessions
displayed a completely unrelated (often today's live) conversation's
first prompt as if it were their own.

Skip the workingDir guess for any item that already has its own
'history' source -- it already had a real, direct attempt. Rows with
no history source at all (their transcript isn't linked/scanned under
their own id yet) still get the guess, matching the original intent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 e888c65c52 fix(history): exclude non-interactive (SDK-driven) transcripts from Past Sessions
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.

Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 1ea39de650 fix(mobile): gate the phone overview's run picker on CLI availability
MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu is a separate,
hardcoded duplicate of the toolbar's #runModeMenu (mobile-overview.js
is a newer feature that mirrors the toolbar menu's look/behavior
rather than reusing its render), so it never picked up #201's
isCliAvailable() gating and offered every backend regardless of what
the server actually has installed.

Gate it the same way: skip an entry unless isCliAvailable(mode),
shell always exempt. Added functional + static regression tests
mirroring the toolbar menu's own test pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:15 -05:00
Codeman maintainer e2a644997e chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 09:01:45 +02:00
Ark0N cd5a101626 Merge pull request #213 from Ark0N/feat/file-viewer-edit-mode
File Viewer: edit mode for text files (edit + save in the viewer)
2026-08-05 09:00:22 +02:00
Codeman maintainer 4ea781c80f feat(file-viewer): edit mode for text files (edit + save in the viewer)
Closes #212. The file-preview overlay can now edit workspace text files in
place, phone-first: agent writes a file, you review it in the viewer, tweak
two lines, save, tell the agent to continue.

Backend (file-routes.ts, policy in src/config/file-editing.ts):
- GET file-content?edit=1: read-for-edit that never truncates (a truncated
  buffer must never become an edit buffer), 512KB cap (413 over it), and
  returns the sha256 hash + detected EOL the client echoes back on save.
- PUT /api/sessions/:id/file-content: edit-in-place only, with no O_CREAT
  anywhere in the handler. Confinement matches the read path (realpath +
  workspace boundary + ownership via findSessionOrFail), plus sensitive-path
  and attachment-guard blocklists, a .git subtree deny, and an extension
  allowlist (svg and env deliberately excluded). Optimistic concurrency via
  baseHash: mismatch is a 409 unless force. Writes are wx-temp + fchmod +
  fsync + rename, closing the validate-then-write TOCTOU window.
- Corruption guards: NUL sniff + UTF-8 round-trip compare (refuses binary
  and latin-1), and server-side EOL re-application so a textarea's LF
  normalization cannot rewrite every line of a CRLF file.
- Plain reads gain an additive editable flag the UI keys the button off.

Frontend (panels-ui.js + overlay markup/styles):
- Edit button on editable text previews; textarea editor with Save/Cancel,
  dirty indicator, discard-confirm on cancel/close, and a conflict dialog
  that offers overwrite (force) when the file changed on disk mid-edit.
- Phone: full-bleed window sized by --app-height so the editor and Save bar
  track the OS keyboard; 16px editor font (iOS zoom guard); no autofocus.
- zh-CN strings for the new chrome.

Tests: pure policy unit tests plus a route suite that deliberately does NOT
mock node:fs. It runs against a real temp workspace so symlink escapes,
write-through of in-workspace symlinks, mode preservation, CRLF round-trip,
409/force, and the no-create property are exercised for real. Also verified
end to end on an isolated beta instance: 39-check curl matrix, Playwright
desktop flow (real clicks and typing, bytes asserted on disk, live conflict
with an external rewrite), and a 393px phone profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 08:44:47 +02:00
Codeman maintainer d9123de9eb feat(terminal): Ctrl+C copies the selection, interrupts when nothing is selected
Closes #211. Copying from the terminal only worked through the browser
context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown,
so the muscle-memory copy failed silently and read as "no copy-paste at all".

With a selection, Ctrl+C now copies it, toasts, clears the selection and
sends nothing to the PTY. With no selection it falls through unchanged, so
the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never
falls through: an explicit copy that interrupts a running agent because the
selection happened to be empty would be a footgun.

Three details that keep the interrupt safe:

- The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the
  no-selection path returns true WITHOUT preventDefault. xterm calls the
  custom handler before its own cancel(), so returning false alone does not
  cancel the event; the copy path therefore calls preventDefault explicitly,
  or the browser would run its native copy on top of ours.
- copy-selection is a full registry entry (rebindable and disableable in App
  Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the
  same trick command-palette uses: the generic capture loop preventDefaults
  every match it dispatches, which would cost the user the interrupt key.
- The gate is keydown-only, since the custom handler also runs for keypress
  and keyup.

Copy goes through _copyText (Clipboard API, then hidden-textarea +
execCommand) rather than raw navigator.clipboard, because install.sh's LAN
option serves plain HTTP where navigator.clipboard is undefined; the
fallback steals focus, so the terminal is refocused afterwards.

Tests: test/terminal-copy-selection.test.ts pins the gate and the
SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real
key presses in chromium and asserts on the clipboard plus the bytes xterm
emitted (browser-driven, so excluded from test:ci like the other Playwright
suites). Verified manually on an isolated beta instance before landing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 02:38:57 +02:00
Codeman maintainer 1e5f6c8ee1 chore: version packages 2026-08-05 02:10:55 +02:00
Codeman maintainer 5d2899907e fix(cli-gating): gate the tunnel button instead of deleting it, and cover antigravity
Follow-up to #200 and #201, which gate the welcome buttons and the run-mode
dropdown on whether the CLI is actually installed. Four corrections:

1. #200 also DELETED the Cloudflare Tunnel welcome button and the QR widget
   outright. Its rationale is right (offering a tunnel where cloudflared is not
   installed is a bad default) but the conclusion overshoots: the welcome QR is
   the whole scan-to-connect-from-your-phone flow, and deleting it left a large
   block of live tunnel code in settings-ui.js driving elements that no longer
   existed. Both are restored and the button is gated on cloudflared, which is
   what the stated rationale actually asks for. New cloudflared-resolver.ts
   mirrors the CLI resolvers, and TunnelManager now shares its search path so
   the button and the spawn can never disagree about where cloudflared lives.

2. Antigravity was missing from the run-mode gating, the one run mode LEAST
   likely to be installed. It slipped past because #201 predates it. Covered
   now, plus a static test that fails if a sixth mode reaches the dropdown
   without being gated, so the next one cannot slip the same way.

3. The per-surface fetches are replaced by the injected availability object
   already used for the Codex settings tab, so the codebase has one mechanism
   rather than two. The status routes buy nothing as a gating source: every
   resolver memoizes its PATH probe server-side, so a fetch is exactly as stale
   as an injected value while costing a round trip every time the dropdown opens
   and leaving the welcome buttons to flicker in after paint. The routes
   themselves stay, including the /api/claude/status that #200 adds.

4. Unknown availability now reads as AVAILABLE for run buttons. Both PRs hid the
   button on a failed fetch, so a blip left a working install with nothing to
   click; a genuinely missing CLI only ever produced an error toast. The Codex
   settings TAB keeps the opposite default, since hiding it costs nothing.

The dropdown query is also scoped to the menu: `.run-mode-option` is the class
the saved-dashboard and history rows use too, and a document-wide querySelector
would have found whichever came first in the DOM.

Fixes a latent environment-sensitivity in 816d900 while here: the index-title
test asserted the template was untouched apart from the title, which held only
on a machine with no codex installed.

Verified end-to-end against a real server on an isolated instance+socket, with
Playwright: gemini/codex hidden and claude/opencode/antigravity/shell shown,
matching this host, tunnel button back, Codex settings tab still hidden, no
console errors. Full test:ci sweep green (3902 tests).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:53:43 +02:00
Codeman maintainer 8facd5e7e7 Merge pull request #201 from timkjr/pr/gate-run-mode-dropdown
fix(run-mode): gate dropdown entries on CLI availability
2026-08-05 01:40:47 +02:00
Codeman maintainer b54094a4c8 Merge pull request #200 from timkjr/pr/gate-gemini-drop-tunnel-button
fix(welcome): gate CLI welcome buttons on actual availability
2026-08-05 01:40:44 +02:00
Codeman maintainer 2b89f35599 fix(shell,remote-ssh): allowlist the login flags, and keep only CRASHED remote panes
Follow-up to #209 and #210. Both land a real fix (a pane that is a login shell
picks up /etc/profile and the per-user PATH entries an ssh remote command never
sees, which is what was failing agent CLIs with exit 127). Three corrections:

1. `-i -l` is no longer hardcoded onto the resolved shell. That path ultimately
   comes from the passwd entry, which is user data and can name anything, and a
   shell that rejects an unknown flag exits on the spot: nushell, elvish and xonsh
   take neither flag, so a user with one of those in passwd would have gotten a
   dead pane on arrival, which is exactly the #208 failure #209 builds on top of.
   loginShellArgs() applies them only to the POSIX-family shells verified to
   accept both, and a test really launches every allowlisted shell present on the
   machine rather than trusting the set. csh/tcsh are excluded deliberately: tcsh
   honors -l only when it is the ONLY flag.

2. `remain-on-exit on` -> `failed`, moved LAST in the tmux command chain. `on`
   keeps the pane after a CLEAN exit too, so typing `exit` in a remote shell
   stranded a dead pane, the session outlived it, and the next launch's `-A`
   reattached to that corpse: "Pane is dead (status 0)" instead of a shell,
   permanently, on the DEFAULT path. Verified against a real tmux, as was the
   fix: `failed` tears the session down on status 0 and keeps the pane on 127
   with the "command not found" still on screen, which is the case #210 wanted.
   It is last because tmux aborts the remaining commands of a `\;` sequence once
   one errors (also verified) and `failed` needs tmux >= 3.2 on the REMOTE host;
   leading, a rejection there would have silently dropped status/mouse/prefix/
   escape-time/window-size along with it.

3. `$SHELL` -> `"${SHELL:-/bin/sh}"`, via one shared remoteLoginShellCommand()
   helper instead of the string being rebuilt in tmux-manager as well.

Also corrects the rationale both PRs carried: a tmux pane already hands the shell
a tty, so it was interactive all along ($- contains i for a bare /bin/bash in a
pane) and ~/.bashrc was always being sourced. `-l` is the flag doing the work.

End-to-end verified, not just unit-tested: the emitted remote pane command was
run through all three quoting layers under a minimal sshd-style PATH with the
CLI installed only on a login-shell PATH entry, and it resolved and launched the
CLI with its arguments intact and a space-containing remote path preserved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:40:37 +02:00
Codeman maintainer ee670c38f6 Merge pull request #210 from timkjr/fix/remote-ssh-login-shell
fix(remote-ssh): route shell + agent CLIs through a real interactive login shell
2026-08-05 01:34:23 +02:00
Codeman maintainer ad57109dcf Merge pull request #209 from timkjr/fix/shell-login-shell
fix(shell): launch shell tabs as an interactive login shell
2026-08-05 01:34:22 +02:00
Codeman maintainer c15b8345b5 fix(history): never treat the empty split segment as a directory name
Follow-up to #202. The dotdir decode landed there was reachable only when
nothing else matched first, and in the greedy half it was not reachable at all.

decodeProjectKey() splits the project key on '-', so the '/.' that the encoder
collapses leaves an EMPTY segment behind. Both loops offered that empty string
as a candidate directory name, and isDir(current + '/' + '') stats current + '/',
which always succeeds. So the empty segment matched unconditionally:

  - backtracking half: ~/.sib resolved to "/home/x//sib" whenever a non-dot
    sibling ~/sib existed (wrong directory, and a doubled slash that then fails
    every string comparison against session.workingDir). Without a sibling it
    only backtracked out by luck.
  - greedy half: that loop is shortest-match-first, so the empty candidate
    matched on the FIRST iteration and set matched=true, leaving #202's dotdir
    branch permanently dead there.

An empty string is never a real path component, so skip it in both loops. The
unmatched tail then has to handle the empty segment too, or it would append a
bare '/' and re-introduce the '//' path it just stopped producing; it now emits
the dotdir guess instead, which is what the encoder implies.

Regression test asserts both halves: the dotdir wins over the non-dot sibling,
and the result never contains '//'. Verified it fails on #202 as merged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:34:16 +02:00
Codeman maintainer 45ae9f4064 Merge pull request #202 from timkjr/fix/dotdir-workingdir-decode
fix: decode dotdir working directories in history session scanning
2026-08-05 01:32:47 +02:00
Codeman maintainer 816d900857 feat(settings): show the Codex CLI tab only where codex is installed
Both settings on the App Settings "Codex CLI" tab (bypass approvals, animated
status effects) are handed to `codex` at launch, so on an instance where the
binary does not resolve the tab offers choices nothing can act on. Gate it on
availability instead.

renderIndexHtml injects window.__codemanCodexAvailable, mirroring the existing
gesture-availability flag, and settings-ui.js hides the tab button when it is
absent. Injected rather than fetched on modal open so the tab cannot flicker in
and back out; isCodexAvailable() memoizes its PATH probe, so the per-render cost
is nil. Installing codex later needs a restart, exactly like the
/api/codex/status route that already backs the Run menu. Solo popups skip the
probe since they have no settings modal.

Only the tab BUTTON is toggled. The panel already carries
.modal-tab-content.hidden unless it is the selected tab and openAppSettings()
always reopens on Display, so an unreachable button keeps the panel unreachable.
The inputs stay in the DOM and are still populated and read back on save, so a
user without codex cannot silently wipe the codex preferences of an instance
that has it. Animations stay off by default for new local Codex sessions.

Verified in a browser on this host, which has no codex: the flag is absent, the
Codex tab is hidden while the other tabs are unaffected, and saving App Settings
with the tab hidden leaves codexAnimationsEnabled/codexDangerouslyBypassApprovals
untouched. With the flag forced on, the tab appears, its panel opens, and
toggling the visible slider persists. The openAppSettings coupling test was
checked to fail when the call is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:14:53 +02:00
Ark0N ddc267c6ff Merge pull request #181 from Lint111/agent/split-codex-animations
feat(codex): make terminal animations configurable
2026-08-05 00:20:38 +02:00
timkjrandClaude Sonnet 5 d66007053b fix(shell): launch shell tabs as an interactive login shell
Shell-mode sessions resolve to an absolute shell path (issue #208's
fix) but launch it bare, with no -i/-l flags. Without those, the
spawned shell runs as a non-interactive child of the non-interactive
`bash -c` that launches the pane, so it never sources ~/.zshrc or
~/.bashrc — silently dropping aliases, PATH additions, and tool init
(zoxide, nvm, etc.).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:12:28 -05:00
timkjrandClaude Sonnet 5 f470f3a4e7 fix: decode dotdir working directories in history session scanning
decodeProjectKey() couldn't recover a dotdir path (e.g. ~/.codeman) from
Claude Code's encoded project-key names: the encoder maps both '/' and
'.' to '-', so the decoder's candidate joins never matched a hidden
directory on disk. It silently fell through to bare $HOME instead,
which corrupted workingDir for any resumed session under a dotdir case
(observed on ~/.codeman itself: history rows and state.json recorded
"/home/timkjr" instead of "/home/timkjr/.codeman").

Add a dot-prefixed candidate to both the backtracking decoder and its
greedy fallback so a leading empty split segment (the signature of a
literal '.' in the original path) is retried as a hidden directory.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:11:51 -05:00
Codeman maintainer cb3eecad9b Merge branch 'master' into pr181 2026-08-05 00:04:43 +02:00
Ark0N db24fc6d7e Merge pull request #180 from Lint111/agent/split-preserve-active-launch
fix(sessions): preserve active terminal during launches
2026-08-04 23:53:14 +02:00
Codeman maintainer 292ba2c775 fix(sessions): route antigravity launches through the ownership helpers
runAntigravity() landed on master after this branch was cut, so it kept the
exact pattern the rest of this PR removes: terminal.clear() plus direct
writeln into whatever session happened to be active. Merging master in
surfaced it, leaving one of six run modes still wiping the active session's
xterm on launch.

Also adds regression coverage that can actually see the bug. The existing
test drives the three helpers directly, so it stays green even when a run*()
function is reverted to writing at the terminal itself: reverting
runClaude()'s call site keeps all 16 tests passing. The new static guard
scans session-ui.js and fails if any run*() body touches
this.terminal.clear/writeln, which catches a regressed call site and would
have caught runAntigravity on its own. A second unit test covers the
home-screen path that nothing exercised: with no active session, launch
progress must still clear and render in the terminal.

Verified in a browser against a live instance. With a session active,
runShell() and runAntigravity() leave its terminal untouched (clear() calls:
0, writes: 0) and emit one info toast; on master the same run wipes the
session's marker text. The session-less home screen still clears and writes
exactly as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 23:47:17 +02:00
timkjrandClaude Sonnet 5 e803186dfe fix(remote-ssh): route claude/opencode/codex/gemini/antigravity through login shell
remain-on-exit (previous commit) preserved dead remote panes instead of
destroying them, which revealed the real failure: `exec claude`/`exec
opencode` ran under ssh's non-interactive, non-login remote-command
shell, which only sees sshd's minimal default PATH — not the ~/.zshrc
PATH entries where these CLIs actually live (e.g. ~/.local/bin,
~/.opencode/bin). Wrap them in `$SHELL -i -l -c '<cmd>'`, mirroring the
fix shell mode already had.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 16:41:27 -05:00
timkjrandClaude Sonnet 5 474efd9023 fix(remote-ssh): use remote user's real shell, keep dead panes alive
Remote shell-mode sessions hardcoded 'exec bash -l', ignoring the
remote user's actual login shell. sshd sets $SHELL from the remote
user's /etc/passwd entry, so 'exec $SHELL -i -l' launches their real
shell (zsh, fish, etc.) with rc files sourced, same fix as the local
shell-mode launch.

Also set remain-on-exit on the remote tmux session. It was only ever
set on the local socket, so if the remote command exited for any
reason -- even something transient -- tmux destroyed the pane, window,
and (being the only session) the whole remote server, tearing down the
local ssh attach along with it and leaving no trace to diagnose. The
local pane saw this as an instant clean exit, and reconnect's -A then
created a fresh session, which could repeat as a flap loop with no
evidence surviving between attempts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 16:39:29 -05:00
Codeman maintainer 03bb40c78a Merge branch 'master' into pr180 2026-08-04 23:37:59 +02:00
timkjr 3ea1ea28f0 fix(welcome): gate Claude and Opencode buttons on CLI availability too
Extends the Gemini gating from bb7fb9e to the other welcome-screen
buttons that had the same problem: shown unconditionally even when the
underlying CLI isn't installed.

- Add isClaudeAvailable() (claude-cli-resolver.ts) and GET
  /api/claude/status, mirroring the existing opencode/codex/gemini
  resolvers and status endpoints.
- Opencode already had a working /api/opencode/status the welcome
  screen just wasn't checking; wire it up the same way.
- Refactor loadGeminiAvailability() into a shared
  _loadCliAvailability(buttonId, statusUrl) helper instead of
  duplicating the fetch/try-catch three times.

Run-mode dropdown entries (Opencode/Codex) are intentionally left
unconditional here — follow-up PR.
2026-08-04 16:22:21 -05:00
timkjr 62008fb408 fix(welcome): gate Gemini button on availability, drop unconditional tunnel button
- Remove the always-visible Cloudflare Tunnel welcome button and QR
  widget; offering it regardless of whether cloudflared is installed
  is a bad default.
- Hide the "Run Gemini" welcome button by default and only show it
  when /api/gemini/status reports available:true, via new
  loadGeminiAvailability() called from showWelcome().
2026-08-04 16:21:55 -05:00
timkjr 660b320a67 fix(run-mode): gate dropdown entries on CLI availability
Follow-up to the welcome-screen gating (#200): the run-mode dropdown
(gear menu next to Run) had the same problem — Claude/Opencode/Codex/
Gemini entries were always shown regardless of whether the CLI is
actually installed, so picking one could spawn a session that
immediately errors out.

- Add _refreshRunModeAvailability() (session-ui.js), called each time
  the dropdown opens; hides entries whose /api/<cli>/status reports
  unavailable.
- Shell is intentionally never gated (no external CLI dependency).

Depends on isClaudeAvailable()/GET /api/claude/status, which don't
exist on upstream/master yet — duplicated here from #200 so this PR
is self-contained and independently mergeable. Once #200 lands this
branch should be rebased onto master, which will collapse the
duplicate cleanly.
2026-08-04 16:21:17 -05:00
Codeman maintainer 529d8fa8ea chore: version packages 2026-08-04 23:09:00 +02:00
Codeman maintainer 19af37977a fix(ui): stop dropping the session name typed in the options modal
Two independent ways a tab description could be typed in and silently lost.

1. Session Options modal (deterministic). The Session Name input saves on
   blur, and every autosave handler in the modal bails on a null
   editingSessionId. closeSessionOptions() cleared that id BEFORE hiding the
   modal, and hiding it is what blurs the input, so the save always ran too
   late and returned early. Escape and backdrop-click lost the name with no
   PUT at all; only the X button worked, because mousedown blurs the input
   before the click handler runs. Fix: blur the focused modal field first,
   then clear the id. That also covers the auto-compact prompt, which saves
   on change and had the same fate.

2. Right-click inline rename (racy). The _inlineRenameActive guard from #81
   sits in renderSessionTabs() (the scheduler) and _fullRenderSessionTabs(),
   but not in _renderSessionTabsImmediate() (the debounced executor). A
   render queued in the ~100ms before the rename opened still fires and the
   incremental branch rewrites .tab-name's innerHTML, destroying the input
   mid-keystroke: it commits a truncated name, or, if it lands before the
   first keystroke, closes the rename so everything typed after goes
   nowhere. Fix: guard the executor too. finishRename() re-renders on both
   commit and cancel, so a render dropped there is picked back up.

Verified end-to-end against a live server on an isolated instance: all three
modal close paths now persist the name, and the rename input survives a
render mid-typing. Both regression tests were checked to fail with their fix
reverted; the render one was vacuous at first because the synthetic tab sat
on <body> instead of inside #sessionTabs, so it now builds the tab in the
real container.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 17:41:02 +02:00
Codeman maintainer 23f258a85d chore: version packages
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).

Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.

Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 15:02:46 +02:00
Codeman maintainer aa4f423d8a chore(gitignore): ignore the root pr/ working dir
pr/ holds machine-local promo drafts that are never meant for git. Anchored with
a leading slash so it matches only the root dir, matching the /public entry below
it, rather than swallowing any nested pr/ elsewhere in the tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 13:56:52 +02:00
Codeman maintainer 1b1057d9e0 chore: version packages 2026-08-04 13:01:28 +02:00
Codeman maintainer 26cbbe0dcb feat(cli): Antigravity run mode
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.

- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
  resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
  env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
  injected via socket-scoped `tmux setenv`, never on the spawn command line, so
  the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
  colours. `runAntigravity()` routes remote/docker cases through
  `POST /api/quick-start` and skips the local status probe for them.

Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 12:59:40 +02:00
Codeman maintainer 1113d34ca8 feat(ui): opt-in entrance animations for tabs, terminal pane, agent windows and connection lines
All OFF by default (the `legacy` theme), so an untouched install behaves exactly
as before and every mark/apply hook short-circuits on its first line. Opt in via
App Settings > Appearance > Entrance Animations; per-surface control and a live
preview lab at ?animlab=1.

Surfaces and styles:
- Tabs: slide, pop, crt, unroll, boot, flip. A batch launched together cascades
  by a configurable stagger.
- Terminal pane: crt, boot, wipe, slide, fade.
- Agent windows: fly (the pre-existing tab-to-window flight, still the default),
  crt, materialize, unfold, beam, pop.
- Connection lines: draw, packet, fade.

Three constraints drove the design:

1. Tabs and connection lines are DESTROYED mid-animation on every re-render:
   _fullRenderSessionTabs() replaces the strip's innerHTML and
   _updateConnectionLinesImmediate() does `svg.innerHTML = ''`, both of which run
   constantly while sessions and agents spawn. Each is tracked by id and
   re-applied to the fresh element with a NEGATIVE animation-delay so it resumes
   at the same offset instead of restarting or snapping. Verified on the real
   path: a forced rebuild mid-draw resumed at -0.243s.

2. Terminal-pane styles animate transform/opacity/clip-path ONLY. xterm's
   FitAddon derives rows+cols from getComputedStyle(parent).width/height, the
   untransformed layout box, so transforms are invisible to it; animating
   width/height/padding would have resized the PTY. Verified by forcing
   fitAddon.fit() eight times mid-animation: dimensions held at 178x38.

3. A window entrance that transforms also moves the rect its connection line
   aims at (crt drifts it 109px, pop 81px). `beam` animates opacity/filter only
   (0px drift) so its line can draw toward a stable target; the others refresh
   the lines on animationend.

Also fixes: an agent window spawning hidden (its agent belongs to a background
tab) is display:none, so its animation never runs and animationend never fires,
which left the entrance class and its inline custom property stuck on the window
permanently. Hidden windows now skip the entrance entirely.

Styles persist to their own codeman:*Anim localStorage keys, keeping them
per-device without touching the .strict() SettingsUpdateSchema.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 12:31:52 +02:00
shenlvkang-collab ab7a703e90 fix(web): keep the viewer's conversation anchor across a Codeman restart
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.

Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.

Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
2026-08-03 21:22:33 +08:00
Codeman maintainer 8a31f10b7d chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 14:09:17 +02:00
Ark0N 2891ae0d6d Merge pull request #178 from Lint111/agent/split-notification-noise
fix(notifications): quiet lifecycle hook noise
2026-08-03 14:07:54 +02:00
Ark0N 17b86b1007 Merge pull request #177 from Lint111/agent/split-transcript-tool-results
fix(transcripts): complete tools from user results
2026-08-03 14:05:38 +02:00
shenlvkang-collabandClaude Opus 5 73315bc351 fix(web): pin the Claude response viewer to the pane's own conversation
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.

Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:53:42 +08:00
Codeman maintainer 7e357691af chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 12:41:26 +02:00
Codeman maintainer 80e7249a39 fix(hooks,test): harden background rewake, fix hook timeout units, stabilize CI teardown
Follow-ups from the PR #175/#176 reviews:

- Rewake helper self-terminates on its own 6h deadline and when orphaned,
  instead of relying on Claude Code to reap the poller
- Rewake marker versioned (V2) with a version-agnostic ownership prefix, so
  future script updates replace older handlers instead of duplicating them;
  regression test covers the V1 to V2 swap
- HOOK_TIMEOUT_MS renamed to HOOK_TIMEOUT_SECONDS = 10: the hook timeout
  field is seconds (the CLI multiplies by 1000), so the curl hooks have
  effectively had a ~2.8h timeout since COD-54
- Test echo PTY switches to raw mode: each input byte echoes exactly once
  (tty line discipline doubled every line and buffered until Enter)
- test/setup.ts: drain in-flight console-log rpc forwards before environment
  teardown (fixes the EnvironmentTeardownError that failed CI twice on the
  merge commit with all 3820 tests passing), clean the temp home on process
  exit (fully-skipped files leaked it), fix the Windows Playwright cache
  fallback path
- test/webview-proxy.test.ts: stop naming the vitest environment directive in
  prose; vitest matches it inside comments and silently ran the whole file
  under the jsdom environment while the comment claimed node
- CLAUDE.md: document the temp-HOME and echo-PTY test isolation

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-31 08:47:52 +02:00
Ark0N e0226f7186 Merge pull request #176 from Lint111/agent/split-hook-lifecycle
fix(hooks): reawaken jobs without replacing user hooks
2026-07-31 08:33:58 +02:00
Ark0N e8681f575f Merge pull request #175 from Lint111/agent/split-quick-start-fixture
test: isolate runtime state and PTY integration
2026-07-31 07:22:30 +02:00
Codeman maintainer 64be4e3029 ci(release): pin the Latest badge to the Codeman release
The workspace publishes two packages, changesets creates a GitHub release
for each, and GitHub awards "Latest" to whichever was published last. That
is a race: 1.9.2 kept the badge, 1.9.4 lost it to xterm-zerolag-input@0.1.7
by two seconds. Set make_latest in the rename PATCH, which runs after every
package release already exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:20:32 +02:00
Codeman maintainer cb7d0ba565 chore: version packages
PUT /api/settings service toggles now resolve from `merged` (persisted +
incoming) instead of the raw request body, so a partial PUT no longer
starts the subagent watcher and stops the workflow + image watchers by
treating every omitted key as "apply the default". Pinned by a 4-case
regression test verified to fail against the old handler.

Also trims the links line from the Codeman callout in the
xterm-zerolag-input README.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 16:11:25 +02:00
Codeman maintainer 22cb563f1e chore: version packages
Plan-usage chip defaults ON on desktop (handhelds stay OFF), resolved
through a single planUsageChipEnabled() helper so the checkbox, the chip
and the create-time statusLineTelemetry flag cannot disagree. Correct the
stale "Cron button defaults ON" comment (it is OFF in code, template and
CSS) and the styles.css comment claiming the server strips the chip's
hidden class.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 14:27:45 +02:00
Codeman maintainer f0e13f9fc3 docs(zerolag): Codeman promo up top, simpler 30-second graphic
Replace the misaligned 8-line keystroke-flow diagram (its branch sat two
columns off the junction it attached to) with a two-lane contrast that
makes the same point in two lines: stock xterm.js waiting 300ms vs the
overlay painting immediately. The mechanism detail it was annotating
moved into the following paragraph.

Add a Codeman callout between the badges and the demo GIF, with links to
getcodeman.com, the install one-liner and the repo, and rewrite the
bottom Origin section so it argues credibility instead of repeating the
promo.

Not released: the npm page updates only on publish, so the next COM
needs an "xterm-zerolag-input": patch changeset for this to ship.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 13:36:08 +02:00
Codeman maintainer 28c5b5c1eb chore: version packages
Rewrite the xterm-zerolag-input README (hero demo GIF, value-first
structure) and fix its drift against the source: 175 tests not 78,
CJK/emoji wide-char support documented instead of listed as a
limitation, setPrompt() documented.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 12:44:48 +02:00
Codeman maintainer af9db455ff chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 17:50:41 +02:00
Codeman maintainer a406aef2fa chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-29 09:01:18 +02:00
lior bba3d80971 test: isolate runtime state and PTY integration 2026-07-29 09:19:46 +03:00
lior 94e3aae57d feat(codex): make terminal animations configurable 2026-07-29 03:39:07 +03:00
lior 0a039239e4 fix(sessions): preserve active terminal during launches 2026-07-29 03:30:38 +03:00
lior 67eb5b43eb fix(notifications): quiet lifecycle hook noise 2026-07-28 23:20:01 +03:00
lior 4a4720cb62 fix(transcripts): complete tools from user results 2026-07-28 23:18:10 +03:00
lior 3c903b36ca fix(hooks): reawaken jobs without replacing user hooks 2026-07-28 23:16:37 +03:00
lior 7c07284b95 test: isolate quick-start case fixtures 2026-07-28 23:12:27 +03:00
Codeman maintainer 77bcbc9b94 docs(readme): move Zero-Lag Input Overlay to fourth section
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 11:46:59 +02:00
Codeman maintainer d4540c5ce6 docs(readme): move Mobile-Optimized Web UI right after Quick Start
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 11:42:53 +02:00
Codeman maintainer 4a83efcd48 Merge remote master (response viewer normalization) into local
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 11:29:42 +02:00
Codeman maintainer 473c57c7ca docs(readme): drop the static tests badge
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 11:28:40 +02:00
Ark0N b586007f14 Merge pull request #169 from shenlvkang-collab/contrib/claude-response-viewer-normalization
fix(web): normalize Claude response viewer turns at real human boundaries
2026-07-28 11:10:28 +02:00
Codeman maintainer b388b84cc2 merge master into claude-response-viewer-normalization
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
The response-viewer detail now lives in architecture-invariants, so the
Claude turn-grouping and restored-placeholder rebind notes moved there.
Changeset rewritten to record the measured effect on real transcripts.
2026-07-28 11:04:41 +02:00
Codeman maintainer d13642ebce docs(readme): merge touch-optimized content into the mobile section
- One compact Mobile-Optimized Web UI section: the two current screenshots
  (mobile-session-keyboard, mobile-toolbar-enter) side by side, comparison
  table, condensed feature bullets, then QR auth
- Drop the outdated black-background phone screenshots
  (mobile-landing-qr.png, mobile-session-active.png)
- Same restructure in README.zh-CN.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:57:14 +02:00
Codeman maintainer 3cff98fe56 fix(security): scope the filesystem path picker per user in multi-user mode
Both picker endpoints are a second file-serving surface, and they
inherited neither the attachment guard's confinement nor its ownership
scoping. Two separate holes:

1. `sessionId` contributes that session's workingDir as a browse root,
   but it was resolved straight off ctx.sessions/ctx.store with no owner
   check, unlike the nine other session-scoped handlers in this file. A
   non-admin could pin ANOTHER user's working directory as a root just
   by passing their session id, then list and preview underneath it. Now
   runs canAccessOwned and reports 404, which also avoids confirming
   that a session id exists.

2. `Home` and `CASES_DIR` were unconditional roots for every caller.
   Per-user spaces live at <USER_SPACES_DIR>/<username>, which is INSIDE
   homedir(), so the Home root alone exposed every other user's
   workspace. A multi-user non-admin now gets only their own
   userSpacePath plus anything explicitly listed in
   CODEMAN_FILE_PICKER_ROOTS. /mnt/d is dropped as well: a broad host
   mount should be an explicit operator decision in a multi-user
   deployment, and operators who want it can name it in that env var.

Admins and single-user mode keep the host-wide roots, so behavior is
unchanged unless CODEMAN_MULTIUSER is on (opt-in, off by default).

All three discriminating tests were verified to fail against the
previous code: browse and preview both returned 200 instead of 404, and
the roots came back as [Home, Codeman Cases, ...] instead of [My Space].
Full suite green, 3784 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 10:49:49 +02:00
Codeman maintainer bc232e5ff3 docs(readme): move zero-lag demo back below agent visualization
The zerolag composition renders on a pure black page background, which
read as an outdated screenshot when placed right under the hero. Top of
the README now shows only the current-skin visuals (subagent gif + tour).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:49:27 +02:00
Codeman maintainer 2a7e035d2b docs(readme): value-first overhaul with getcodeman.com install and new zerolag demo
- Move the install one-liner (curl getcodeman.com/install | bash) and value bullets to the top so the pitch and quick start fit in the first scrolls
- Promote Zero-Lag Input Overlay to right after the hero, with a new side-by-side phone demo gif generated from the current zerolag master
- Switch all install commands (incl. WSL) to the getcodeman.com short URL
- Remove outdated screenshots (multi-session-dashboard.png, ralph-tracker-8tasks-44percent.png) and the old zerolag-demo.gif
- Remove Ralph tracker content: tracking section, API table, CLI example, autonomy-table row, architecture-diagram node
- Mirror all changes in README.zh-CN.md

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:40:44 +02:00
Ark0N 80a88ea857 Merge pull request #168 from shenlvkang-collab/contrib/mobile-path-picker-preview
feat(mobile): add filesystem path picker and document/image previews
2026-07-28 10:25:17 +02:00
Codeman maintainer 5c45d434ac test(qr-auth): replace flaky max-deviation bias check with chi-square
The short-code distribution test asserted that no base62 character
deviated more than 15% from its expected count. That statistic is the
maximum over 62 correlated near-normal cells, so its tail is fat: at
n=36000 the per-cell relative SD is ~4.1%, which puts the 15% bound at
|z| ~ 3.65, and taken as a max over 62 cells it fires on a perfectly
uniform generator about 1.6% of the time. Measured directly: 48 spurious
failures in 3000 simulated runs. It had been rerun-to-green repeatedly
and most recently red-herringed a PR review.

Chi-square is the correct test for "is this multinomial uniform", and
unlike 0.15 its threshold is derivable. df=61, Wilson-Hilferty puts the
p=1e-6 critical value at ~129, so the bound is 130.

Power is unchanged. Removing rejection sampling from generateShortCode
reintroduces modulo bias (256 % 62 = 8, so eight characters draw five
chances per 256 instead of four) and was verified against the real code
in an isolated worktree: chi-square 243.06 against the 130 limit. The
threshold sits in a wide empty gap, 3000 clean runs peaked at 104 while
200 biased runs bottomed out at 174.5.

Also iterate the alphabet explicitly rather than the observed keys, so a
character that never appears counts as zero instead of being skipped.

Verified: 30/30 consecutive runs of the real test pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 04:15:30 +02:00
Codeman maintainer 390516ca3f merge master into mobile-path-picker-preview
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
Route counts reconciled against master's numbering (files 14 -> 16 for
the two new filesystem endpoints, total ~197 -> ~199) and the path
picker's detail moved into architecture-invariants under its own
section.
2026-07-28 01:13:32 +02:00
Codeman maintainer 57b6be1ed5 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 01:10:51 +02:00
Ark0N cbae989e02 Merge pull request #170 from shenlvkang-collab/contrib/light-skins-1.8.0
feat(ui): add four light skins (Paper Gray, Solarized Light, Catppuccin Latte, Rosé Pine Dawn)
2026-07-28 01:01:54 +02:00
Codeman maintainer 84f47e8ee0 fix(skins): keep tinted badges readable on light skins, pin OG modals
The light skins themed the app chrome, but a class of status badges and
accent-tinted pills still hardcode pale light-on-dark ink (#cdddff,
#9dc0ff, #ffc107, #fff) over a low-alpha tint. Measured on a rendered
page, that lands at 1.0 to 1.9:1 under all four light skins: the search
filter chips (Sessions / Events / Files) render as empty blue pills.
Re-point the ink at each skin's own dark tokens and keep the tint as the
category signal, which moves the same components to 3.2 to 14:1.

Also pin --floating-bg on the OG skin. The new :root default is slate
rgba(31,38,48,.96), which suits the Daylight palettes (their glass
header is already rgba(31,38,48,.85)) but repaints OG's modals, command
palette and floating windows away from the neutral near-black that skin
is built on.

Verified against a live instance across all seven skins, plus a real
shell session for terminal ANSI output. Full suite green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-28 00:56:15 +02:00
Codeman maintainer 541d9c8131 merge master into light-skins 2026-07-28 00:09:34 +02:00
Codeman maintainer e4ea785a28 docs: move the demo MP4s out to the private media archive
Both were unreferenced by either README and are now kept in Ark0N/gittrend
under assets/codeman-demos, alongside the source recordings they were cut from.

As with the GIF removal, this does not shrink the repository: the blobs remain
in history and only new checkouts stop carrying them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 22:59:53 +02:00
Codeman maintainer b7a6a189f9 docs: drop the superseded 29MB subagent-demo.gif
Neither README references it: both switched to the dated
subagent-demo-20260724.gif in 8e9f254, which kept this file only so external
hotlinks would keep resolving. Removing it now at the maintainer's request.

Note this does NOT shrink the repository. The blob stays in history, so clone
size is unchanged; only new checkouts stop carrying the 29MB file. Actually
reclaiming the space needs a history rewrite, which would invalidate every
existing clone and is a separate decision.

The three capture scripts that write docs/images/subagent-demo.gif are
unaffected: they create the file, they do not read it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 22:57:48 +02:00
Codeman maintainer e063222ac2 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 22:27:43 +02:00
Codeman maintainer 346bc8b173 fix(terminal): link whole URLs and paths instead of truncating them
Three separate truncations, each cutting a clickable link short so it opened the
wrong target (or nothing at all).

1. A single `&` ended the match. It is a query-parameter separator, so every real
   query string was cut: a WordPress edit link resolved to `?post=1479` and opened
   the post list instead of the editor, and Claude Code's own `/login` URL was not
   usable at all. `&` is now part of a URL; `&&` stays a boundary, since that is
   the shell operator and never appears inside one. A lone trailing `&` is still
   trimmed as punctuation.

2. Links longer than the terminal is wide were cut at the row boundary. xterm
   calls the link provider once per visible ROW and translateToString returns only
   that row, despite a comment here claiming it handled wrapping. The provider now
   stitches the continuation rows back into one logical line and maps match offsets
   back to (x, y), so a link can span rows.

   Two kinds of continuation exist and handling only the first is not enough. A
   SOFT wrap is the emulator running out of columns, which flags the next row
   `isWrapped`. A HARD wrap is the program wrapping its own output and emitting a
   real newline, which flags nothing: Ink does this, which is why the /login URL
   was cut at the window edge and why the clickable part grew when the window was
   widened. A row that fills the full width is now treated as continuing into the
   next, that being the only trace a hard wrap leaves behind. Bounded to 12 rows so
   a screenful of wide output cannot make every hover re-scan the viewport.

3. Image and PDF paths were not matched at all. `.claude-images/paste-*.png`, what
   Codeman writes for a pasted screenshot, rendered as plain text. Those extensions
   are now linked and open the file preview, which renders images inline, rather
   than the log viewer, which would show binary noise.

Verified in a real terminal: a 450-char /login URL hard-wrapped across 5 rows with
zero isWrapped flags in the buffer (so a genuine hard wrap, not the soft case)
links intact, as do soft-wrapped URLs and a wrapped attachment path. Regression
cases added to link-provider-regex.test.ts, which extracts the patterns from the
shipped source so they cannot drift. Its existing ReDoS guard still passes, which
matters because this changes a pattern that once froze the tab on hover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 22:25:29 +02:00
Codeman maintainer b34fcaf928 feat(web-tabs): open dashboard URLs as tabs beside agent sessions
Adds a "Web / URL" section to the Run dropdown. A saved URL renders as a tab in
the same strip as Claude/Codex/Gemini sessions, with the same Alt+1..9 numbering,
so Codeman is one mission control instead of Codeman plus a pile of browser tabs.

A webview is NOT a sixth SessionMode: no PTY, no tmux, no respawn, no idle
detection. It is a separate resource sharing only the tab strip and the main
content area, the same call that keeps Docker and remote-SSH as case overlays.

Dashboards are proxied through Codeman's own origin, because a direct iframe
fails three ways at once in the shipped deployment: prod serves HTTPS behind
tailscale serve, so http:// targets are hard-blocked as mixed content (with no
override at all on iOS Safari); Grafana/Portainer-class dashboards send
X-Frame-Options: DENY; and our own default-src 'self' CSP blocks cross-origin
frames. Proxying dissolves all three and leaves the production CSP byte-for-byte
unchanged, since /webview/... is already covered by 'self'. A useful side effect:
the fetch happens server-side, so a tailnet-only dashboard is reachable from a
phone that is not on the tailnet.

The proxy is not an API surface. It authenticates on a 192-bit capability in the
path (memory-only, rolling TTL, bound to the minting user, revoked on edit or
delete) and is correspondingly exempt from the cookie and Origin checks, because
a sandboxed iframe is opaque-origin: it sends no SameSite=lax cookie and its
writes arrive with Origin: null. The Host allowlist is never bypassed. A second
Referer-keyed form of the exemption exists for root-absolute assets and is fenced
to safe methods on non-/api, non-/ws, non-/q paths.

Iframes omit allow-same-origin unless a URL is explicitly marked trusted, since a
proxied page is served from Codeman's own origin and could otherwise read this
document and drive the agent-spawning API. Authorization and codeman_session are
stripped upstream in BOTH modes, so CODEMAN_PASSWORD cannot leak into a dashboard.

Two things only a real browser reveals, both presenting as the dashboard's own
"Failed to fetch" while the page itself renders fine:

- Runtime-built root-absolute URLs (fetch('/api/data')) escape <base href> and
  land on Codeman's root. Widening the Referer fallback into /api would trade
  security for it, so an injected shim patches fetch/XHR/WebSocket/EventSource
  inside the frame instead, removing the class rather than the guard.
- An opaque-origin document CORS-checks every request, including to the host it
  was served from. Script/css/img loads are not CORS-checked, which is why the
  page renders while its API calls die. The proxy now emits CORS headers and
  answers preflights itself. registerSecurityHeaders answered every OPTIONS with
  a bare 204 before routing, carrying no ACAO for Origin: null, so that
  short-circuit now exempts a valid capability.

Neither is reproducible with curl, which does not enforce CORS.

Also fixes a pre-existing bug found on the way: .toolbar has backdrop-filter,
making it a stacking context that trapped .run-mode-menu's z-index:1000, so
.welcome-overlay painted over the whole Run menu. With no session open, every
item in it (Claude Code included) was unclickable.

Verified end to end against a real tailnet dashboard: live data, WebSocket push,
no failed requests, and switching tabs does not reload the frame. 98 new tests
cover the pure rewrite helpers, the CORS helper, the shim's rewrite logic, route
CRUD, and every edge of the auth exemption.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 17:06:36 +02:00
Codeman maintainer ea4c935d51 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 15:06:15 +02:00
Codeman maintainer 716b7ccdbb docs(claude): record this session's changes and the traps found along the way
Feature + layout changes:
- Phone toolbar: Enter replaced Shell below 430px; shell launching moved into
  the Run dropdown. Documents the ordering (setRunMode -> run -> runShell) and
  that runMode is a loose string server-side so new modes need no schema edit.
- Repo root layout: config/ holds knip.json, Prettier config is the package.json
  "prettier" key, SECURITY.md is under .github/, and the list of files that must
  stay at the root with the reason each one is pinned there.
- Pointer to docs/SPEEDRUN.md, which nothing linked to after the move.

Traps worth not rediscovering:
- sendEnterKey MUST use triggerDataEvent, not sendInput or a raw POST. Local
  echo is on by default on touch devices, so typed text is buffered client-side
  and a bare CR submits an empty line while the text stays stranded. Cost me two
  wrong fixes before the real cause surfaced.
- styles.css nests skin overrides under html:not([data-skin="og"]), giving a
  bare .btn-toolbar rule (0,2,1) which outranks .btn-toolbar.btn-x (0,2,0) in
  mobile.css whatever the load order. Explains why mobile.css needs !important.
- Browser tests pass vacuously on mobile input: sendInput() bypasses the overlay,
  and headless Chromium reports isTouchDevice() false even with hasTouch, so the
  local-echo branch never runs. Assert on overlay state and the tmux pane.
- The working tree is shared with other agent sessions: check the branch before
  every commit (a commit silently landed on feat/web-tabs today and the push to
  master reported "Everything up-to-date"), push with HEAD:master rather than
  checking master out, and never git add -A.
- COM step 5 no longer tells you to git add -A, which has swept another
  session's WIP into a release before.

Verified: 33 relative links and 30 invariants anchors all resolve, and every
factual claim re-checked against the tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 14:58:22 +02:00
Codeman maintainer da7a095e33 chore: move knip config into config/ and Prettier config into package.json
Continues trimming the repo root so the README is reached with less scrolling.
Root files: 19 -> 15 across both passes.

- knip.json -> config/knip.json, joining eslint.config.js and the vitest
  configs. `npm run knip` now passes --config explicitly. Verified by A/B: the
  run from the new location produces byte-identical findings and the same five
  configuration hints as from the root, so knip resolves its globs relative to
  cwd rather than the config file. Those hints are pre-existing, not caused by
  the move.
- .prettierrc -> the "prettier" key in package.json, a config source Prettier
  reads natively, so editor format-on-save keeps working with no --config flag
  anywhere. Verified live: `npm run format:check` still passes across src/**,
  which it could not if the config had been lost (Prettier's defaults are
  double quotes at 80 columns and would flag nearly every file).

.prettierignore deliberately stays at the root: Prettier resolves it relative
to cwd, so moving it would require threading --ignore-path through every
script and would break editor integration.

Everything else in the root is load-bearing: .editorconfig (walks up from the
edited file), .nvmrc/.npmrc (read from the project root), tsconfig.json (bare
`tsc` discovers it), LICENSE (GitHub license detection), install.sh (its raw
URL is the published one-liner in the README and cannot move without breaking
every copy in the wild), plus the five documented .md files.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 14:44:02 +02:00
Codeman maintainer 149cee6bcd docs: move SECURITY.md to .github/ and SPEEDRUN.md to docs/
Trims the repo root listing so the README is reached with less scrolling.
Only these two were movable; the other five root .md files are load-bearing
and stay put:

- README.md / README.zh-CN.md — the landing page and the language-switcher
  entry point
- CLAUDE.md — Claude Code loads project instructions from the ROOT path;
  moving it silently breaks every future session in this repo
- AGENTS.md — the agent-convention file Codex reads from the root and injects
  as context (see the comments in session-routes.ts)
- CHANGELOG.md — the changesets default writer emits it next to package.json,
  so moving it breaks `npm run version-packages`

GitHub officially resolves .github/SECURITY.md, so the Security policy tab
keeps working. Inbound links updated in both READMEs, CLAUDE.md and
docs/versioning-policy.md. CHANGELOG.md also names SECURITY.md but is left
alone: it is a historical record, not a live reference.

The move broke a link the other direction too: SECURITY.md pointed at
docs/security-architecture.md, which from .github/ resolved to
.github/docs/... — repointed to ../docs/. All relative links in the six
touched files verified resolving (62 links, 0 broken).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 14:00:26 +02:00
Codeman maintainer 7cda2194c3 docs(readme): real phone screenshots of the new Enter button toolbar
Two captures from an actual phone, replacing the placeholder-ish shots:

- Mobile table, middle cell: the full-height capture with the keyboard open,
  showing the accessory bar and the new Enter button while answering a plan
  prompt. Supersedes mobile-session-question-20260727.png from this morning,
  which showed the pre-Enter toolbar.
- Touch-Optimized Interface: the cropped toolbar capture as a standalone
  560px figure, where a near-square crop reads better than it would squeezed
  into a 260px table cell.

Picking the tall capture for the table also fixes a row the earlier square
shot had left lopsided: the three cells now render 473 / 482 / 469px tall
instead of 473 / 263 / 469.

Adds a "Dedicated Enter button" bullet documenting the behaviour, including
why it replays the keypress (local-echo flush) rather than sending a bare
carriage return, and that shell launching moved into the Run dropdown.
Both READMEs updated so EN and zh-CN stay in sync.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 13:52:31 +02:00
Codeman maintainer eb8724bbf2 feat(mobile): replace the phone Shell button with Enter, move Shell into Run
On phones the toolbar slot held "Shell", which starts a rarely-needed session
type. Sending Enter is a constant need on a touch keyboard, so the slot now
holds a dark blue Enter button and shell launching moves into the expandable
Run dropdown (Terminal / Shell, label "Run SH"). Desktop and tablet are
unchanged: the green Run Shell button stays exactly where it was.

Enter goes through xterm's own input path:

  coreService.triggerDataEvent('\r', true)

NOT through sendInput() or a direct POST to /input. localEchoEnabled defaults
to MobileDetection.isTouchDevice(), so on a phone the characters you type are
buffered client-side in the LocalEchoOverlay and have never reached the PTY.
The onData Enter branch in terminal-ui.js is what flushes that buffer before
sending \r. A bare \r submits an empty line and leaves the typed text stranded
on screen, which presents as "the Enter button does nothing". Replaying the
keypress reuses the overlay flush, the flushed-offset cleanup and the 80ms
text-before-CR ordering instead of reimplementing them.

Verified with local echo forced on: before the fix the overlay still held
"echo OLD_WAY" after Enter; after it, pendingText is empty and the command
executes in the pane.

The !important on the Enter button's colors is required, not habit: styles.css
nests its skin overrides inside `html:not([data-skin="og"]) { … }`, so a plain
.btn-toolbar there resolves to (0,2,1) and outranks .btn-toolbar.btn-enter at
(0,2,0). Without it the button renders in generic toolbar grey.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 03:01:14 +02:00
Codeman maintainer cb6c25220f docs(readme): swap the mobile idle screenshot for an interactive prompt
Replaces the middle cell of the Mobile-Optimized Web UI table in both
READMEs. The new shot shows an agent's multiple-choice prompt being answered
on a phone, with the touch accessory bar and bottom toolbar visible, which
demonstrates more of the mobile UI than the old idle-session capture.

Uses a dated filename per the convention the other 2026-07 images follow.
That also avoids GitHub's image cache serving the old picture, which an
in-place overwrite of mobile-session-idle.png would have risked. The old
file is left on disk so any existing external link to it keeps working.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 02:30:44 +02:00
Codeman maintainer de87c4e315 docs(readme): add contributors and total-commits badges
Two live shields.io badges in the header block of both READMEs, linking to
the contributors graph and the commit history. Colors reuse the existing
palette (3b82f6, 1e3a5f) and keep the flat-square style.

Verified both endpoints render real data matching the GitHub API
(contributors: 13, commits: 1.5k against 1,460 on master) and that master
is the default branch, so the /commits/master link target is correct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 02:07:28 +02:00
Codeman maintainer 63710cf2c1 docs: split CLAUDE.md deep detail into architecture-invariants, ignore it in prettier
CLAUDE.md was 110KB (~27.5k tokens) loaded into every session, with 30 lines
carrying 49% of the bytes as single-paragraph walls (the Docker cases entry
alone was 9,388 chars). Extract the implementation detail verbatim into
docs/architecture-invariants.md (41 sections) and leave the rule plus a
pointer inline. Result: 59.5KB, ~14.9k tokens, 46% smaller.

Also:

- Add CLAUDE.md to .prettierignore. Prettier's markdown printer escapes
  underscores in the glob-heavy paths used throughout, which had already
  corrupted the Ultracode paragraph (agent-*.jsonl became agent-\_.jsonl,
  collapsing backtick spans). npm run format:check is unaffected; its globs
  are src/** only.
- Move version archaeology (PR numbers, ticket ids, commit shas, "was X now
  Y" lineage) into the invariants doc, keeping the rules and their reasoning
  inline.
- De-duplicate the Core Files table against Key Patterns.
- Document install.sh in Scripts, and why Prettier's scope is deliberately
  narrow (14 hand-formatted public JS modules are guarded by
  check:public-assets and check:frontend-syntax instead).

Two factual fixes found while verifying: displayKeys is a client-side merge
policy, not a wire filter, and showResponseViewer / showPlanUsageLimits /
language are declared in SettingsUpdateSchema and do persist server-side; and
the respawn route count is 7, not 18.

Verified: 30/30 cross-doc pointers resolve, 1,184 of 1,190 backticked
identifiers from the original survive (the 6 others are dropped archaeology
or the prettier-corrupted spellings), 59 table rows well-formed,
format:check and check:frontend-syntax clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-27 02:00:13 +02:00
codeman-local 8c089a4819 chore: add light skins changeset 2026-07-25 15:31:23 +08:00
codeman-local f812f65a33 feat(files): preview picker documents and images 2026-07-25 15:29:00 +08:00
codeman-local 2667150f33 feat(mobile): add filesystem path picker 2026-07-25 15:28:19 +08:00
codeman-local a842b091bf fix(ui): theme stateful light surfaces 2026-07-25 15:28:04 +08:00
codeman-local dae82388ed feat(ui): add light skin themes 2026-07-25 15:28:04 +08:00
codeman-local bca56b4273 fix(web): normalize Claude response viewer turns 2026-07-25 15:26:34 +08:00
Codeman maintainer 86c634959d docs: reword hero bullet to 'Self-hosted and private'
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 14:55:55 +02:00
Codeman maintainer fc5294e7c2 docs: fresh Live Agent Visualization images (subagent windows + ultracode run)
Replaces the dated subagent-spawn.png with the recaptured floating-windows
still (clean header, three haiku Explore agents, connector lines) and adds
the live ultracode workflow-run window below the feature bullets, in both
READMEs. Dated filenames so caches never serve a stale render.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 14:55:00 +02:00
Codeman maintainer 8e9f25482a docs: new README hero GIF (smooth subagent pop, 3MB) + annotated dashboard tour
Replaces the 29MB subagent-demo.gif reference with a recaptured 6s loop:
three haiku Explore agents pop as floating windows (25fps through the pop,
bayer dither, 1080px) on the new clean header. Adds the annotated dashboard
tour screenshot below the feature bullets in both READMEs. Old GIF file kept
on disk so external hotlinks stay alive; new files use dated names so caches
can never serve a stale render.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 14:50:50 +02:00
Codeman maintainer 211f3c07dd feat(ui): clean default header (usage chips, file viewer on; token chip, lifecycle log off)
The default desktop header right cluster is now: WS, CPU, MEM, File Viewer
folder button, 5H/7D plan-usage chips, gear. The token-count chip and the
lifecycle-log document button default OFF (both still honor stored prefs),
and the File Viewer button defaults ON (phones keep hiding it via mobile.css).
Templates ship the hidden/shown state so nothing flashes before settings load.
Capture scripts seed showTokenCount:false so screenshots match regardless of
server defaults.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 14:50:43 +02:00
Codeman maintainer 876f9a75b4 fix(install): preserve the existing network binding on updates and re-installs
Updating must never silently loosen security. The update path already never
rewrites service files; this covers the remaining gap, re-running the full
installer over an existing setup:

- read_existing_binding() parses the current systemd unit or launchd plist
  (a pre-1.8 service without our env lines counts as loopback).
- The network-access prompt defaults to the CURRENT setup instead of the
  network default, shows what that setup is, and Enter keeps it, including
  a custom non-loopback host and the existing password.
- Non-interactive re-installs adopt the existing binding wholesale.
- The update path's closing security notice now reflects the service's
  actual binding instead of the generic loopback text.

Round-trip escaping tested for both formats (quotes, backslashes, XML
specials) plus the legacy-unit, preserve, and Enter-keeps flows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 10:58:40 +02:00
Codeman maintainer d7bb726213 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 09:16:48 +02:00
Codeman maintainer 715aef2076 feat(install): ask for network binding, default to LAN access with password prompt
The loopback-only default was safe but left most installs unreachable from
the devices people actually use. The installer now asks at the end of setup:

1) Any device on your network (0.0.0.0), the default. Prompts for a
   dashboard password (confirmed twice); skipping it requires an explicit
   confirmation and prints a big red warning as the final output.
2) This machine only (127.0.0.1), the safer option, for tunnel/Tailscale
   setups.

The choice flows into the systemd unit, the launchd plist (values escaped
for both formats), the run-now exec path, and the printed URLs (LAN IP
detection included). Non-interactive installs keep the safe loopback
default unless CODEMAN_HOST is preset; the server binary's own default
binding is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 03:09:43 +02:00
Codeman maintainer 608ec8a10e chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 00:46:03 +02:00
Codeman maintainer 303afd7fe1 docs: add blog article images (dashboard tour, mobile shots)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-24 00:44:00 +02:00
Codeman maintainer 0ee268ba82 fix(ui): stop the centered voice button overlapping the case picker
.toolbar-center is absolutely centered (left: 50%), so on viewports below
~1500px, or with long case names widening the left toolbar group, the voice
button rendered on top of the case picker's chevron and the + button. Below
1500px it now falls back into normal flex flow where overlap is impossible;
wide viewports keep the centered layout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 23:36:38 +02:00
Codeman maintainer 1be98ff8a3 fix(mobile): collapse header brand to a C home button on phones
On <430px screens the full Codeman wordmark wasted header space; the brand
now renders a single C (same tap target, still app.goHome()). Desktop and
tablet keep the full wordmark. The compact letter lives in a separate
aria-hidden span so i18n custom branding keeps rewriting only the wordmark.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 22:47:09 +02:00
Codeman maintainer b710013add docs: README hero pitch, badges, star CTA
Add a short what-is-Codeman pitch block with deep links after the hero GIF,
npm version + GitHub stars badges, and a closing star/issues CTA.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 17:25:11 +02:00
Codeman maintainer 4343805672 docs: sync CLAUDE.md core-files table (Infra docker modules, app.js ~5K lines)
Adds src/docker-hosts.ts + src/docker-export.ts to the Infra row and
corrects the app.js size note (4906 lines), merged with the 1.7.0
i18n.js additions to the same rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 10:14:27 +02:00
Codeman maintainer 4f8471189e chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 09:43:22 +02:00
Ark0N fad7cdc1ab Merge pull request #165 from shenlvkang-collab/feat/custom-name-i18n
feat(ui): add custom branding and Chinese localization
2026-07-23 09:41:48 +02:00
Codeman maintainer 56db02412b Merge master into feat/custom-name-i18n; keep windowTitle stable on solo renders
Resolves the CLAUDE.md paragraph conflict with #162, skips the
windowTitle recompute for solo-session renders so a detached window
cannot reset the push-notification hostTitle prefix to the default,
and prettier-formats test/mobile/devices.ts (came in unformatted via
the #162 merge; CI format:check only covers src/**).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-23 09:17:04 +02:00
Ark0N 689d9fc5e5 Merge pull request #164 from shenlvkang-collab/fix/run-session-tab-dedup
fix(ui): show new run tabs immediately
2026-07-23 09:14:22 +02:00
Ark0N 50547a4e89 Merge pull request #163 from shenlvkang-collab/fix/unicode-working-directories
fix(paths): accept Unicode working directories
2026-07-23 09:13:24 +02:00
Ark0N 3c2a5bfef3 Merge pull request #162 from shenlvkang-collab/fix/foldable-mobile-settings
fix(mobile): preserve settings across foldable postures
2026-07-22 19:05:02 +02:00
Codeman maintainer bc66add7ed docs: sync READMEs with the 1.6.2 installer behavior
Quick Start now documents the consent-first flow (every system change is
prompted; the closing menu chooses terminal / background service / skip),
safe re-runs (finished installs update in place with local changes stashed
and the service restart verified; interrupted installs resume full setup;
install.sh update/uninstall), and the headless contract (system-changing
steps abort without CODEMAN_NONINTERACTIVE=1). The AI CLI note now says all
four CLIs are auto-detected with an install-or-skip choice when none exist,
and the background-service section points at installer menu option 2 before
the manual instructions. Same changes mirrored in README.zh-CN.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:52:47 +02:00
Codeman maintainer 24b5d8fa63 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-22 12:36:57 +02:00
codeman-local 8d9fc4195b feat(ui): add custom branding and Chinese localization 2026-07-21 02:49:25 +08:00
codeman-local 5abcae16b4 fix(ui): show new run tabs immediately 2026-07-21 00:20:34 +08:00
codeman-local 66ad681666 fix(paths): accept Unicode working directories 2026-07-21 00:04:40 +08:00
Codeman maintainer 6c8d4ca72f chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 17:35:40 +02:00
Codeman maintainer 2fdf7dabac docs: sync READMEs with 1.6.0 (remote SSH, session manager, permissions); fix installer prompts under curl|bash
README.md + README.zh-CN.md:
- New "Remote SSH Sessions" section (durable remote tmux, auto-reconnect,
  discover/attach with detach-not-kill, shared sessions, injection-safe ssh)
- New "Session Manager & Command Palette" subsection (pinning survives kill,
  name retention on resume, cross-device tab order sync)
- Multi-user quick start right after installation (users add + --multiuser),
  and the zh-CN README gains the full Multi-User Mode section it was missing
- Security: document the configurable startup permission mode (skip/auto/
  normal/allowedTools) and the multi-user auto downgrade
- Cron header button noted as opt-in (Header Displays); API section counts
  refreshed (~190 handlers / 20 route modules) with pin, session-order and
  unified endpoints; Development now recommends npm run test:ci

CLAUDE.md (/init audit): session-order.ts in the Session row, PR #157
session-manager polish appended to the unified-list pattern, opt-in Cron
button documented, route/SSE counts refreshed (20 modules, ~188 handlers,
~146 events)

install.sh: the post-install "How would you like to run Codeman?" menu (and
the CLI picker + yes/no prompts) read from stdin, which under curl | bash is
the pipe, so choices were impossible and the script silently fell through to
the default. New has_tty()/read_reply() helpers prompt via /dev/tty whenever
a real terminal exists (same approach the sudo path already used) and only
fall back to defaults when there is genuinely none, now with an info line
saying so. Verified both paths with a pty harness (script(1)) and setsid.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 16:17:32 +02:00
codeman-local 51cb3a7205 fix(mobile): preserve settings across foldable postures 2026-07-20 21:22:28 +08:00
Codeman maintainer b10e354936 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 14:53:48 +02:00
Codeman maintainer 64559b60d1 feat(ui): hide the Cron toolbar button by default (opt-in via Header Displays)
The Cron button now follows the same opt-in pattern as the Session Manager /
Away Digest / File Viewer buttons: the template ships the btn-cron--hidden
marker class and applyHeaderVisibilitySettings() removes it only when the
per-device showCronButton setting (App Settings -> Display -> Header Displays)
is enabled. Defaults flipped to false in the mobile defaults block and both
?? fallbacks. Cron jobs remain fully functional; only the launcher is opt-in.

Verified in a live browser: fresh profile hides the button + unchecked toggle,
enabling shows it immediately and persists across reload.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 14:43:26 +02:00
Ark0N 6351b4143f Merge pull request #157 from aakhter/cod-162-session-manager-polish
Session Manager: pinning, cross-device ordering, name/prompt retention
2026-07-20 14:43:11 +02:00
Codeman maintainer 3d6e3f3d6e Merge remote-tracking branch 'origin/master' into pr-157 2026-07-20 14:33:10 +02:00
Ark0N 5c20fcf464 Merge pull request #156 from aakhter/cod-114-remote-tmux-durability
Remote tmux durability: survive SSH drop, discover/attach, collaborative sessions
2026-07-20 14:32:50 +02:00
Codeman maintainer 683544a22e Merge master into PR #157 (session manager polish)
Resolutions (sse-events.ts / constants.js / app.js): unions of the docker/
multi-user event registrations from master with the session-order/pin events
from this branch.

Additions on top of the merge:
- POST /api/sessions/:id/pin now falls back to the persisted store record when
  no live session exists: COD-142 deliberately preserves pinned records after
  kill (and cleanupStaleSessions skips them), so without this a pinned-then-
  killed session could never be unpinned. Owner-scoped in multi-user mode.
- SessionOrderUpdateSchema bounds (id <= 100 chars, <= 500 entries) so a buggy
  client can't persist megabytes into state.json; empty strings still flow to
  normalizeSessionOrder which drops them.
- Route tests for the persisted-record pin fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 14:31:44 +02:00
Codeman maintainer 5181c9abb0 docs: sync remote-sessions.md launch section with the shipped socket/naming
The durable-launch section predated the #145 consolidation: owned launches use
the dedicated -L codeman-remote socket with codeman-ssh-<id8> names and
per-session set -t options (never -g). Discovery/attach (COD-105) genuinely
target the canonical -L codeman socket; the asymmetry is now called out.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 14:27:06 +02:00
Codeman maintainer 25c67f9415 Merge master into PR #156 (remote tmux durability)
Resolutions:
- session.ts: keep the extracted _buildRespawnPaneOptions() helper (COD-108)
  and add master's docker/owner fields to it
- tmux-manager.ts: docker branch first, then remote via buildRemoteSessionCommand
  (now an options object threading claudeMode/allowedTools into
  buildRemoteLaunchCommand, preserving the 6.3 multi-user permission downgrade)
- case-routes.ts: keep master's adminOnly helper; gate the new COD-105 discovery
  endpoint admin-only in multi-user mode (hosts are machine-level infra)
- settings-ui.js: union of remoteAutoReconnect + master's header-button defaults
- session-routes.ts: union of imports; session gets remote + owner

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 14:27:02 +02:00
Codeman maintainer d6917e3b21 chore: version packages
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 13:46:38 +02:00
Codeman maintainer 524a096e14 Merge feat/docker-session-mode into master (docker deep-review fixes)
Brings the docker session-mode deep-review work (intended for the skipped
1.4.2) onto the 1.5.x line: deterministic-conversation-id resume across
container stop/recreate, config-drift detection + POST /api/docker-cases/:name/recreate,
docker model-picker support, import-manifest hardening, remote-daemon (context/
daemonHost) correctness, comma-in-path rejection, and the zh-CN README re-translation.

Conflicts resolved to preserve BOTH the multi-user security scoping already on
master (ownership checks, workingDir confinement, permission downgrade) AND the
docker features. Version kept at master's 1.5.0 (the 1.4.2 bump is superseded;
a fresh changeset bumps to 1.5.1). tsc, eslint, and test:ci all green (3548 tests).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 13:45:37 +02:00
Codeman maintainer 8d9dd70b51 chore: version packages
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 13:14:10 +02:00
Ark0N ed47a599be Merge pull request #161 from Ark0N/feat/multiuser-mode
feat: opt-in multi-user mode (per-user spaces + admin panel)
2026-07-20 13:07:58 +02:00
Codeman maintainer ccb3afc9ee fix(multiuser): close cross-user web-layer scoping holes found in review
The opt-in multi-user feature's only enforcement is web-layer scoping
(all sessions share one OS account). An adversarial review found 8 critical
+ 7 high cross-user holes that defeated it, plus mediums; all fixed here.
Single-user (flag-off) behavior stays byte-identical apart from documented
consistency deltas.

Ownership / confinement:
- DELETE /api/sessions (bulk) + /:id now owner-scope / findSessionOrFail
- quick-start, cron (create+fire), scheduled runs confine workingDir to the
  owner's space; case link/docker-link/docker-import confine the host path
- resolveCasePath no longer resolves linked cases for non-admins; foreign
  remote/docker cases are skipped (fall through to the caller's own local case)
- history, subagents/workflows, mux-sessions, orchestrator, cron run-history,
  away-digest, and remote/docker host reads are owner- or admin-scoped

Permission policy (section 6.3):
- non-granted users are downgraded at every spawn site incl. legacy
  /api/scheduled, PlanOrchestrator one-shots, remote launch, and the cron-fire
  gemini/codex bypass switches; resolveClaudeModeForUsername now fails closed

Auth / store:
- verify-first login throttle (a correct password is never locked out),
  /ws terminal subject to the change-password lockbox, cookie fast-path
  re-validates identity live, role/grant changes revoke sessions, admin delete
  runs the last-admin guard before any teardown
- users.json: distinguish missing (ENOENT) from corrupt/unreadable so a bad
  read can't overwrite all accounts; unique per-process temp write path

Event streams:
- debounced session:updated + batched task:updated, clipboard, and push
  notifications route by owner (fail closed); getLightState hides machine-wide
  globalStats from non-admins

Tests: two suites updated to assert the fixed (secure) behavior. tsc, eslint,
and test:ci all green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 12:33:12 +02:00
Codeman maintainer c3b0dc345b fix(multiuser): wrap modal tabs so the injected Users tab is clickable
A Playwright browser pass found the injected 9th App Settings tab (Users)
overflowed the non-wrapping .modal-tabs flex row and landed under the modal
backdrop (elementFromPoint returned .modal-backdrop, not the button), so a real
mouse click was intercepted. flex-wrap:wrap lets the tabs wrap to a second row;
the built-in 8-tab modals still fit on one row (no visual change).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 08:38:06 +02:00
Codeman maintainer 0ab2416460 docs(multiuser): plan status, CLAUDE.md, security-architecture, README + changeset
Stamp the plan doc with shipped-by-phase status; add the multi-user Key Patterns
entry + State Files + case-spaces note to CLAUDE.md; add a multi-user section to
the security architecture (threat model: workspace separation, not a security
boundary) and a README opt-in section; add a minor changeset.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 04:34:44 +02:00
Codeman maintainer ac6fe6ef79 feat(multiuser): phase 5b, frontend (identity boot + admin panel)
- public/admin-ui.js (new, self-contained): on boot fetches GET /api/me and
  stores window.__codemanUser; installs a fetch interceptor that opens a
  change-password modal on any 403 PASSWORD_CHANGE_REQUIRED (and on boot when
  mustChangePassword is set); for a multi-user admin, injects a "Users" tab into
  the existing App Settings modal (create/reset/disable/enable/promote/demote/
  grant-bypass/delete with typed confirm + one-time-password reveal). No header
  button, so the mobile-header policy stays green; nothing renders in single-user
  mode.
- me-routes: GET /api/me returns a `multiUser` flag so the UI distinguishes a
  single-user admin (no admin UI) from a multi-user admin.
- index.html: load admin-ui.js after settings-ui.js, before session-ui.js.

Tests: test/admin-ui.test.ts (JSDOM: identity boot, Users-tab injection gating by
role/mode, forced change-password modal, script-order wiring). Backend verified
end-to-end by test/admin-routes.test.ts against a live server. A full Playwright
pass is recommended before merge.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 04:31:30 +02:00
Codeman maintainer dafe3de185 feat(multiuser): phase 5a, admin user-management API
- routes/admin-routes.ts: GET/POST /api/admin/users, PATCH/DELETE
  /api/admin/users/:username, reset-password, logout. Multi-user only (404
  otherwise), requireAdmin, last-admin invariants, one-time-password on create /
  reset (returned once + mustChangePassword), disable/reset/delete revoke cookie
  sessions, delete kills the user's live sessions first (normal teardown) and can
  delete their space (guarded). Per-user stats (live/active sessions, case count).
- web/admin-audit.ts: append-only ~/.codeman/admin-audit.jsonl (timestamp, acting
  admin, action, target, IP) for every user-management action.
- SSE admin:usersChanged + auth:passwordChangeRequired (sse-events.ts + constants.js).

fix(user-store): serialize users.json read-modify-write

touchLastLogin fires on every Basic auth (fire-and-forget) and was racing route
writes (create/update), clobbering records — a real corruption bug surfaced by
the admin tests. All mutators now run under a single write lock, and
touchLastLogin is throttled to once/minute per user to bound disk churn.

Tests: test/admin-routes.test.ts (8, live server) + user-store lock verified by
the existing user-store suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 04:26:44 +02:00
Codeman maintainer 2a06f7a5a8 feat(multiuser): phase 4, event fan-out + stream scoping
Scopes real-time streams and the init snapshot so a multi-user client only
receives what it owns. No-op in single-user mode (identity-less clients).

- WS terminal (ws-routes): owner gate after the session lookup. A non-admin may
  only attach to their own session (close 4003); the global auth hook already
  ran on the upgrade and decorated req.authUser, so an unauthenticated upgrade
  never reaches the handler.
- SSE (sse-stream-manager): per-client identity stored at addClient; broadcast()
  and the terminal-batch flush both enforce a routing hint via canDeliver().
  WebServer.broadcast auto-derives the hint (deriveSseHint): session-scoped event
  families resolve the owner from the payload's session id (fail closed when the
  owner can't be resolved), machine-level families (docker/tunnel/update/system/
  cron) + host-plan telemetry are admin-only, everything else stays global. Raw
  terminal bytes resolve the owner once and are withheld from non-owners.
- getLightState is filtered per connection AFTER the shared cache (sessions,
  respawnStatus, subagents, workflowRuns by owner; scheduledRuns + planUsage
  admin-only); applied to both the SSE init snapshot and GET /api/status.
- file-routes: getKnownSessionWorkingDir + getSessionAttachmentHistory (the
  preview/thumbnail/history helpers that bypass findSessionOrFail) now owner-check
  the session, closing a cross-user file-read path.
- GET /api/search: harvestSources is owner-scoped.

Deferred to a follow-up (documented in docs/multi-user-plan.md): away-digest +
subagent/workflow REST list scoping, push-subscription identity + routing,
per-user screenshot subdirs. The live-event versions of these are already routed
by the SSE hint; only the on-demand REST aggregates remain global for admins-only
follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 04:17:17 +02:00
Codeman maintainer 453605a58f feat(multiuser): phase 3, ownership threading + scoping
Threads per-user ownership through sessions, cases, cron, and the permission
policy. All scoping is a no-op in single-user mode (isMultiUserMode() guards).

Sessions
- Session.owner stamped at every create path from req.authUser / job.owner:
  POST /api/sessions, /api/run, /api/quick-start, ralph start, cron launch,
  plan generation. Round-trips through recovery (MuxSession.owner mirror, read
  muxSession.owner ?? savedState?.owner) and the mux layer.
- findSessionOrFail(ctx, id, req) now does a NOT_FOUND owner check (never 403, so
  other users' session existence is not leaked); wired at ~50 call sites.
- List endpoints filtered by owner: GET /api/sessions, /api/sessions/unified
  (live+persisted+lifecycle scoped, host-wide transcripts admin-only), cron jobs.

Permission policy (section 6.3)
- resolveClaudeModeForUsername wraps getClaudeModeConfig at every spawn site so a
  non-granted user is forced to --permission-mode auto (bypass -> auto), including
  recovery (or a reboot would un-downgrade). buildPromptArgs now respects the
  session's claudeMode, closing the one-shot (runPrompt) bypass hole.
- Shell mode and cron launchCommand require canBypassPermissions: 403 at
  POST /api/sessions, /api/quick-start create, cron job create, AND cron fire time
  (re-checked against the owner's current grant).

Cases
- resolveCasesDir(user): per-user ~/codeman-users/<name>/cases in multi-user, the
  shared ~/codeman-cases otherwise. All case CRUD + ralph + plan + quick-start
  resolve through it. resolveCasePath is owner-aware.
- GET /api/cases scoped per user (own folders; legacy linked cases admin-only;
  remote/docker cases owner-filtered). RemoteCase/DockerCase gain owner, stamped
  at link/quickcreate/import.
- Remote + Docker host CRUD is admin-only.
- Non-admin workingDir confinement (the linchpin): realpath must resolve inside the
  user's space, enforced at POST /api/sessions and /api/run BEFORE any disk write.

Limits
- sessionCapacityState / sessionCapacityMessage centralize the global + per-user
  cap (CODEMAN_MAX_SESSIONS_PER_USER, default global/2), replacing the 6 copy-pasted
  MAX_CONCURRENT_SESSIONS checks.

Tests: test/ownership-scoping.test.ts (case isolation, host-CRUD gate, workingDir +
shell gates, and the scoping helpers). Deferred to phase 4: WS owner gate, SSE
fan-out filtering, file-route preview/thumbnail helper scoping, push routing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 04:02:46 +02:00
Codeman maintainer 4d8857f72a feat(multiuser): phase 2, multi-user auth pipeline
Adds a parallel multi-user auth branch (the single-user Basic-auth path is
left byte-identical). Off unless CODEMAN_MULTIUSER/--multiuser.

- middleware/auth.ts: mode-selecting registerAuthMiddleware. New async
  multi-user hook verifies username:password against the user store (scrypt),
  mints identity-carrying cookies, decorates req.authUser, enforces a per-IP
  AND per-username failure bucket, and the mustChangePassword lockbox. The
  hook-secret loopback bypass is now a single shared helper used by both
  branches. FastifyRequest.authUser module augmentation.
- ports/auth-port.ts: AuthSessionRecord gains username/role/mustChangePassword.
- user-store.ts: verifyPassword (timing-equalized against user enumeration).
- route-helpers.ts: getAuthUser (synthetic admin fallback), canAccessOwned,
  requireAdmin, revokeUserSessions; findSessionOrFail gains an optional req for
  a NOT_FOUND owner check (dormant until phase 3 wires callers).
- routes/me-routes.ts: GET /api/me (synthetic admin in single-user) and
  POST /api/me/password (verify current, min 8, clear mustChangePassword,
  revoke other sessions).
- QR: QrTokenRecord + AuthSessionRecord carry a username; tunnel-manager
  mintUserToken / consumeTokenWithIdentity / getQrSvgForCode; /q/:code binds
  the cookie to the token's user (rejects identity-less tokens in multi-user);
  GET /api/tunnel/qr mints a per-user token. Single-user keeps the rotating token.
- server.ts: bootstrap the initial admin from CODEMAN_USERNAME/PASSWORD on first
  boot (refuse to start with no users); multi-user with >= 1 user satisfies the
  non-loopback auth requirement and the tunnel-enable guard; userFailures bucket
  disposal.
- types/api.ts: FORBIDDEN, PASSWORD_CHANGE_REQUIRED, USER_EXISTS, USER_NOT_FOUND,
  LAST_ADMIN error codes (message + status wired).
- Session.owner field + getter/setter, SessionState.owner, MuxSession.owner,
  CreateSessionOptions.owner (foundation for phase 3 ownership threading).

Tests: test/multiuser-auth.test.ts (10, live server on 3170/3171). Existing auth
suite (auth-security, qr-auth, cod54-hook-event, network-auth-policy) unchanged
and green; full test:ci sweep passes (3519 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 03:26:21 +02:00
Codeman maintainer f496e35d71 feat(multiuser): phase 1, user store, mode plumbing, CLI
Opt-in multi-user foundation (off by default; no behavior change without
CODEMAN_MULTIUSER/--multiuser):

- src/config/multiuser.ts: isMultiUserMode(), getUserSpacesDir()/userCasesDir(),
  maxUsers(), maxSessionsPerUser() (per-user fairness cap = global/2).
- src/types/user.ts: UserRecord/PasswordHash/AuthUser/PublicUser/UserRole.
- src/user-store.ts: ~/.codeman/users.json (atomic tmp+rename, mode 0600, short
  TTL cache). scrypt hashing with per-record params + timingSafeEqual verify plus
  rehash detection; createUser/setPassword/updateUser/deleteUser with last-admin
  invariants; guarded deleteUserSpace (symlink + realpath confinement, section 8);
  pure section-6.3 resolvers (resolveClaudeModeForUser downgrades bypass to auto
  for non-granted users; canRunPrivilegedCommands); bootstrapInitialAdmin.
- src/cli.ts: "codeman users add|passwd|list|rm" (hidden prompt or
  --password-stdin) operating directly on users.json; a --multiuser flag on the
  web command.

Tests: test/user-store.test.ts (29 tests: hashing/verify/rehash, username
validation, atomic 0600 write, last-admin invariants, 6.3 resolvers,
delete-space guards, bootstrap).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:58:51 +02:00
Codeman maintainer 91070f5dda feat(claude): add 'auto' startup permission mode
Adds Anthropic's classifier-guarded low-prompt mode (--permission-mode
auto) as a fourth ClaudeMode alongside skip-permissions/normal/allowedTools.
Wired through both spawn paths (buildPermissionArgs for direct PTY,
buildClaudePermissionFlags for tmux), the getClaudeModeConfig validator,
and the App Settings Startup Mode picker. Exports buildSpawnCommand for
test coverage.

This is the prerequisite for multi-user mode section 6.3, which downgrades
non-granted users' sessions to 'auto'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:48:14 +02:00
Codeman maintainer fdce57ce5a docs: multi-user mode design plan
Design plan for opt-in multi-user support (per-user case spaces, admin
panel, ownership scoping). Ported onto master as the base for the
feat/multiuser-mode implementation branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:48:07 +02:00
Codeman maintainer a21400614a chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:36:02 +02:00
Codeman maintainer 9046b95b7e docs: multi-user mode design plan (reviewed against code)
Design for opt-in --multiuser: per-user case spaces, scrypt-hashed
users.json, ownership threading across sessions/cases/SSE/push, admin
panel, and a per-user Claude permission-mode policy. Reviewed against
the actual auth/SSE/case/session code; the plan encodes verified
call-site inventories, the non-admin workingDir confinement rule,
WS-upgrade identity plumbing, per-user QR minting, and the linked-cases
v2 format migration.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 02:29:50 +02:00
Codeman maintainer 8b3fa5f37c docs(zh): full re-translation sync of README.zh-CN.md
Bring the Chinese README to 1:1 section parity with the English one.
Adds the three missing sections (Using Codeman: A Human's Guide,
Driving Codeman from an Agent: Programmatic Guide, and Versioning),
updates the keyboard-shortcut table to the current registry (session
palette chord, Option bindings, prev/next tab), refreshes the API
section (18 route modules / ~160 handlers, ApiResponse envelope note,
Sessions rows with clientId+seq, new Cron table), and adds the
SECURITY.md disclosure pointer to the Security intro.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 01:51:01 +02:00
Codeman maintainer 4c6f96a2ef docs: sync CLAUDE.md and READMEs with the 1.4.1 feature set
CLAUDE.md (the 1.4.1 release commit only reformatted it): document the
seeded credential-isolation model (resolveDockerClaudeArtifacts /
resolveDockerCredentialArtifacts, buildSeamlessClaudeConfig), the
auto-built agent base image (ensureAgentBaseImage + docker:imageBuild*
SSE events), the C.UTF-8 image locale, w<n>-<case> tab naming, and the
opt-in File Viewer header button; bump the SSE registry count to ~138.

README.md: add Gemini to every CLI enumeration (tagline, install, WSL,
Multi-CLI, security, architecture diagram), split the Docker section's
hardening bullet into hardening + seamless-auth/credential-isolation,
note the base image now auto-builds on first use, add a File Viewer
bullet to More Features.

README.zh-CN.md: mirror all of the above, add the previously missing
"Isolated Docker Sessions" section and a Docker bullet in More
Features, fix the Node badge to 22+, and run prettier over the file.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-20 01:46:59 +02:00
Ark0N d1928f300e Merge pull request #160 from Ark0N/feat/docker-session-mode
v1.4.1: Docker session mode hardening + File Viewer button
2026-07-20 01:41:58 +02:00
Codeman maintainer ca731c67b3 feat(docker): harden session mode + File Viewer button (v1.4.1)
Docker cases: seamless Claude auth (seed ~/.claude.json instead of the
corruption-prone single-file mount), full credential-store isolation for
claude + codex/gemini/gcloud/opencode (share only transcripts/rollouts,
seed the rest), auto-build the base image on first use, C.UTF-8 locale
(fixes box-drawing), collapsed/shortened Create-Case UI + short "(docker)"
case-menu tags, and w<n>-<case> tab naming for docker/remote sessions.
Also: opt-in File Viewer header button; fix a TZ-boundary flaky test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 01:36:21 +02:00
Ark0N a3fe0ae728 Merge pull request #159 from Ark0N/feat/docker-session-mode
docs: reflect shipped Docker session mode (1.4.0)
2026-07-19 22:28:41 +02:00
Codeman maintainer 82825cbfb3 docs: reflect shipped Docker session mode (1.4.0) across CLAUDE.md/README/security
- CLAUDE.md: rewrite the Docker cases Key Pattern to the shipped 1.4.0 state
  (removes the stale "not on master / Phases remaining" framing); add
  docker-quickcreate/templates/GPU/elastic-disk/export-import, the
  CODEMAN_DOCKER_BRIDGE_HOOKS listener, docker state files, env vars, route +
  SSE counts, and the build-agent-image command.
- README.md: new "Isolated Docker Sessions" section + a More Features bullet.
- docs/security-architecture.md: new §10 "Docker container isolation" (hardening,
  commit-safe creds, blast radius, untrusted-import safety, bridge-hooks) +
  Quick-reference env vars.

docs/docker-cases.md (user guide) and docs/docker-cases-plan.md (design) were
shipped with the feature.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 22:23:05 +02:00
Aamer Akhter 5ec71ace5a COD-145 show last (most recent) prompt alongside first in session manager
Building on COD-140's firstPrompt backfill, surface each session's most
recent user prompt too, so a long-running session is identifiable by both
where it started and where it is now.

- session-routes: add extractLastUserPrompt() (mirrors extractFirstUserPrompt
  with last-match semantics + same noise/secret/slash-command filters + 120
  cap); scanProjectDir computes lastPrompt from the file tail (reads a tail for
  large files; small files scan head); thread lastPrompt through HistorySession
  and the /api/sessions/unified history rows.
- unified-session-service: add lastPrompt to UnifiedSessionItem + HistoryInput,
  set it from history in the merge, and extend the backfill with parallel
  by-uuid / newest-by-workingDir indexes (never overwrites); add lastPrompt to
  the filterAndPaginate search haystack.
- terminal-ui: render a 'Last prompt' detail row, omitted when absent or equal
  to the first prompt (single-prompt sessions show one line).

Tests: unified-session-service.test.ts +5 (uuid-join, workingDir fallback,
newest-wins, no-overwrite, search). Beta-verified: /api/sessions/unified
populated firstPrompt+lastPrompt on all 200 rows (12 distinct); Playwright on
the session-manager modal rendered 12 'Last prompt' rows, 0 console errors.

(cherry picked from commit 115f4d397e91decc1a6381b47a99d74922e9055b)
2026-07-17 16:31:09 -04:00
Aamer Akhter b27a0e9188 COD-140 backfill firstPrompt for sessions whose id != transcript UUID
The unified session list only set firstPrompt from the transcript-history
view, keyed by the Claude transcript file's UUID. A live/persisted row
keyed by its Codeman id only inherited a prompt when that id happened to
equal an on-disk transcript filename; when it didn't (stale/wrong
claudeSessionId, post-/clear new uuid, resumed/attached/worktree session),
the session manager showed "(no prompt captured)" even though a real
transcript for that working dir existed under a different UUID.

Add a pure firstPrompt backfill pass in mergeUnifiedSessions (after the
merge loops, using the already-passed history source): for any row with no
firstPrompt, join by claudeSessionId first, then fall back to the newest
transcript in the same workingDir. Never overwrites a non-empty prompt, so
rows keyed to their own transcript are untouched; rows with genuinely no
transcript still show the placeholder. Pure, unit-tested (+5).

(cherry picked from commit 1f9f53ec64a61c9fa7f77d29efcbdd1d2794ec38)
2026-07-17 16:30:39 -04:00
Aamer Akhter 35a0217ccd COD-142 retain pin when a pinned session is killed
A killed session was full-deleted from state.json (removeSession),
dropping the COD-139 pinned/pinnedAt fields, so the session vanished
from the session-manager pinned group. cleanupStaleSessions also reaped
any persisted record with no live session on boot, which would have
wiped a preserved pin on the next restart.

Fix (state-store):
- demoteOrRemoveSession(id): on kill, demote a *pinned* record to a
  lightweight stopped record (status=stopped, pid=null, pin retained)
  instead of deleting; unpinned records are removed as before.
- cleanupStaleSessions skips pinned records so the pin survives restart.
- server _doCleanupSession calls demoteOrRemoveSession on the killMux
  path (shutdown path unchanged).

Restoration iterates live mux sessions, not state.json, so a stopped+
pinned record is never auto-revived. Unit-tested on the real StateStore
path (state-store.test.ts +4); session-cleanup/session-pin regress green.

(cherry picked from commit 86f183eacfc3f2f6ac28499fb1ae2d21eef2bbed)
2026-07-17 16:26:36 -04:00
Aamer Akhter 7a86cf87f7 COD-143 retain session name when resuming from the session manager
resumeHistorySession ignored the row's name and always synthesized a fresh
w<N>-<dir> name from the working dir, so resuming a custom-named session lost its
name. Thread the name through resumeHistorySession(sessionId, workingDir, name) and
extract the choice into a pure _resolveResumeName helper: prefer a non-empty existing
name, else generate the next free w<N>-<dir>. Forward s.name at all three call sites
(terminal-ui.js history-item + session-manager menu, session-ui.js run-mode history);
sessions without a name fall back to the generated name (unchanged behavior). The
unified session rows already carry name, so session-manager rows resume with it.
TDD: test/resume-name.test.ts drives the real _resolveResumeName via vm-harness.

(cherry picked from commit 56c7906a48d8b453ed55810a02ec70f72d34ed32)
2026-07-17 16:26:28 -04:00
Aamer Akhter 5792c2d62e COD-139 add session pinning (float pinned sessions to top of session manager list)
Pin/unpin a session via POST /api/sessions/:id/pin {pinned}; pinned sessions
sort above unpinned in the unified session manager list (COD-121), ordered by
pinnedAt descending. Pin state lives on SessionState, persists to state.json,
and survives reload/reconnect/restart (persisted-input carries pinned; the
merge skips undefined so a recovered live session can't clobber it). New SSE
event session:pinned re-sorts the open list live across clients. Pin/Unpin
affordance in the session-row kebab menu with a 📌 glyph + amber highlight.

(cherry picked from commit 82749747039afcd4a3104f6a97ce7d3c2ddd048d)
2026-07-17 16:25:29 -04:00
Aamer AkhterandClaude Opus 4.8 8807b3ff6d COD-131 sync tab order across devices via server state
Tab reordering (drag-and-drop + Ctrl+Shift+{/}) persisted only to
localStorage (codeman-session-order), so each device kept its own private
order. Add server-side persistence so the order follows the user across
devices, live. Takes the issue's recommended default (a): one global order,
server authoritative, localStorage as offline fallback.

- session-order.ts (new, pure + unit-tested): normalizeSessionOrder (coerce
  to string[], drop empty/non-string, dedup) and mergeSessionOrder (the
  pushing device's order wins; ids the device hadn't loaded fall to the end
  in their existing relative order, never dropped — graceful for
  closed/remote/parked sessions absent on that device).
- AppState.sessionOrder?: string[]; StateStore get/setSessionOrder + the field
  added to buildPartialJson() (the incremental serializer whitelists fields,
  so without this the value never reached disk / survived a restart).
- PUT /api/session-order (session-routes): parse -> merge -> persist ->
  broadcast session:orderChanged; getLightState() init snapshot now carries
  sessionOrder so a fresh load/reconnect restores it.
- SSE event session:orderChanged registered in sse-events.ts + constants.js.
- app.js: handleInit seeds localStorage from the server snapshot before
  syncSessionOrder(); saveSessionOrder() also PUTs to the server (debounced
  400ms, covers drag + both keyboard moves); _onSessionOrderChanged adopts a
  remote order and re-renders (no-op-guarded to avoid echo flicker).

Verified (orchestrator re-ran all gates): tsc 0, lint 0, frontend-syntax +
prettier clean, build ok; session-order + session-order-routes + state-store
56/56. Functional round-trip on an isolated beta: PUT {a,b,c} -> status
snapshot reflects it; merge PUT {c,a} vs {a,b,c} -> {c,a,b} (b preserved at
end); malformed payload rejected with a clean 400; sessionOrder persisted to
state.json and survived a restart.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

(cherry picked from commit 79415f2fdfbdf3fbe362a063534e7f84c553eefb)
2026-07-17 16:21:16 -04:00
Aamer Akhter 115ada1e9e docs: cover COD-105 remote discover/attach + detach-not-kill in remote-sessions.md
55f5ada (COD-105) added Phase 2 of the remote-tmux arc: discover codeman-*
sessions on a host and attach to non-owned ones, with detach-not-kill on close.

- Data model: SessionRemote.owned/remoteSessionName + RemoteSessionInfo;
  toSessionRemote (owned:true) vs toAttachedSessionRemote (owned:false).
- New Ownership section: discovery (listRemoteCodemanSessions, the literal-\t
  parse quirk, never-throws/VITEST), attach-vs-launch selection
  (buildRemoteSessionCommand), and the killSession detach-not-kill guarantee.
- API: GET /api/remote-hosts/:hostId/sessions + the attachRemoteSession
  create path.
- CLAUDE.md Remote Key Pattern notes discover/attach + detach-not-kill.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
(cherry picked from commit f321e1200a9a7c1e58c69ba936b680200fd53275)
2026-07-17 16:05:06 -04:00
Aamer AkhterandClaude Opus 4.8 b2ebdcbf47 COD-106 shared/collaborative remote tmux sessions (window-size latest + shared badge)
Two Codeman clients attaching the same durable remote tmux session at different
viewports would fight: tmux sizes a window to the SMALLEST attached client by
default. Push `window-size latest` to the remote session config so the window
tracks the most-recently-active client instead, letting concurrent clients
coexist; surface the client count for a "shared · N" badge.

Reconciled onto upstream PR #145: #145 moved the durable remote session onto the
dedicated `-L codeman-remote` socket under a `codeman-ssh-` name and scoped every
tmux set-option PER-SESSION (`set -t <name>`, never `-g`) so a shared remote tmux
server's OTHER sessions keep their own prefix/mouse/sizing. The original COD-106
commit added `set -g window-size latest` (GLOBAL) on the old `-L codeman` socket —
a regression against #145's hardening. This commit layers the window-size feature
onto #145's structure as `set -t <name> window-size latest` (per-session, on the
codeman-remote socket). Test assertions updated to the per-session form
(remote-shared-sessions.test.ts) and the byte-identical launch-command test
(remote-ssh-options.test.ts) extended with the window-size line — which supersedes
the separate f09323c9 assertion fix (dropped: it targeted the global form and also
carried unrelated CLAUDE.md doc changes).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 16:03:03 -04:00
Aamer Akhter 6dba8b5227 COD-108 auto-reconnect remote tmux sessions on SSH drop
Continuous remote-only reconnect watcher closing the COD-104 durability
arc: when a remote session's local ssh pane dies mid-run, re-establish it
automatically instead of leaving a dead pane until the user pokes it.

Design decisions (per cod108 design doc):
- D1 event->owner: TmuxManager watcher DETECTS a dead remote pane and emits
  `remoteSessionDropped`; the session owner (server) reassembles the same
  RespawnPaneOptions and calls Session.reattachRemote() -> respawnPane, which
  re-runs the idempotent remote command (owned new-session -A / non-owned
  attach) and REJOINS the still-running durable remote tmux session. The
  watcher never reassembles options itself, and never routes through the
  Claude-idle respawn-controller.
- D2 bounded backoff: per-session exponential backoff [5s,15s,45s,2m,5m,5m],
  reset on a successful reattach, `remoteReconnectExhausted` emitted once after
  the cap. Pure, unit-tested schedule + eligibility decision.
- D3 always-on + kill-switch: `remoteAutoReconnect` app setting (default ON),
  read each tick; when false the watcher does nothing.

Guards: killSession() (incl. the non-owned DETACH early-return) and shutdown
add the session to an intentional-teardown guard set + clear its backoff
BEFORE teardown, so a closed/killed tab is never auto-revived. Exactly one
reconnect in flight per session (inFlight guard prevents stacked respawns).
Per-session reconnect/guard state cleared on session removal.

New: src/remote-reconnect.ts (pure backoff + decideReconnect), TmuxManager
startRemoteReconnectWatcher/stop + runRemoteReconnectTick + noteRemoteReconnect
+ guardRemoteReconnect + clearRemoteReconnectState; Session.reattachRemote()
(+ extracted _buildRespawnPaneOptions, shared with interactive start); server
wiring + watcher start; 3 SSE events (sse-events.ts + constants.js in sync,
broadcast + app.js exhausted "Reconnect" affordance); remoteAutoReconnect
schema + settings-ui toggle.

Tests: test/remote-auto-reconnect.test.ts (21) - pure schedule, eligibility
(guarded never reconnects, non-remote/pane-alive/not-due skip, over-cap
exhaust), and manager-level integration (dead remote pane -> dropped ->
backoff -> exhausted; guarded emits nothing; reset-on-success; kill-switch
off; state-cleared-on-remove). Verified real-remote against aa-desktop: drop
local ssh pane -> watcher emitted -> respawnPane reattached the SAME remote
session (remote pane_pid unchanged 3939->3939); test session cleaned up, the
real host sessions left untouched.

Checks: tsc, eslint, check:frontend-syntax, check:public-assets, prettier
--check, build all green; tmux-manager/session-routes/session-manager/
sse-registry-parity suites pass.

(cherry picked from commit d13d58b1994eb6594fd2eadea208104d36204f9d)
2026-07-17 15:59:36 -04:00
Aamer AkhterandClaude Opus 4.8 897bfdff59 COD-109 terminate owned durable remote tmux sessions (propagate kill to remote)
Since COD-104 a remote session lives in a durable tmux server on the host and
outlives the local pane, so killing a tab only DETACHED — even for sessions we
own. Propagate `kill-session` to the remote for OWNED sessions in killSession's
owned path (after COD-105's non-owned detach-only early-return); non-owned
detach-only is untouched.

Reconciled onto upstream PR #145: #145 already upstreamed this exact owned-kill
propagation as `buildRemoteKillCommand({ remote, sessionId })` on the dedicated
`-L codeman-remote` socket (matching buildRemoteLaunchCommand) and wired it into
killSession (Strategy 3b, owned-only, fire-and-forget). The original COD-109
commit added a second `buildRemoteKillCommand(remote, name)` overload on the old
`-L codeman` socket plus a duplicate kill block — a compile error AND a wrong
socket post-#145 (owned sessions no longer live on `codeman`). This commit keeps
#145's socket-correct implementation and drops the duplicate; the required
test/remote-kill-command.test.ts is retargeted to #145's `{ remote, sessionId }`
signature and the `codeman-remote` socket.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 15:54:32 -04:00
Aamer Akhter fb013e9de0 COD-105 discover + attach existing remote tmux sessions (detach-not-kill)
Phase 2 of the remote-tmux arc. Discover codeman-* tmux sessions already
running on a remote host (created by the remote's own Codeman or another
instance) and attach to one this Codeman didn't launch, with detach-not-kill
ownership for non-owned sessions.

- remote-hosts.ts: listRemoteCodemanSessions (ssh, VITEST-guarded, never throws)
  + pure parseRemoteSessionList + buildRemoteListSessionsCommand. Parser splits
  on the LITERAL \t the remote tmux emits (next-3.7 does not expand \t) AND a
  real tab. toAttachedSessionRemote builds a non-owned SessionRemote; toSessionRemote
  now marks the COD-104 launch path owned:true.
- tmux-manager.ts: buildRemoteAttachCommand (sibling of buildRemoteLaunchCommand);
  buildRemoteSessionCommand selects attach vs launch by ownership. killSession gains
  a detach-not-kill early return for non-owned remote sessions: tears down only the
  LOCAL pane (kills local ssh -> remote attach detaches), NEVER issues a remote
  kill-session.
- types/session.ts: RemoteSessionInfo; SessionRemote.owned + remoteSessionName.
- schemas.ts: CreateSessionSchema.attachRemoteSession {hostId, remoteSessionName};
  fixed a pre-existing no-useless-escape lint error in the jumpHost regex.
- case-routes.ts: GET /api/remote-hosts/:hostId/sessions (explicit discovery).
- session-routes.ts: attachRemoteSession create path -> non-owned session.
- UI (index.html/session-ui.js/styles.css): explicit "Discover existing sessions"
  button + Attach action (owned:false). No auto-discover.

Verified on aa-desktop: discovered codeman-disco1, attached (attached=1, shared
view), killed local probe pane -> remote SURVIVED_DETACH (attached=0). Tests:
parse/attach-cmd/ownership unit + discovery route, session-routes + case-routes green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 55f5ada9db6d01518a4adf6b752e460b5df39524)
2026-07-17 15:49:24 -04:00
Aamer Akhter 7f24a132d0 COD-104 fix: skip remote tmux prereq check under VITEST (test-mode)
COD-104 wired checkRemoteTmuxAvailable into the remote-session create path,
but it does a real `ssh` via exec — so 2 remote-create tests in
session-routes.test.ts hit a ~10s ssh timeout and failed (422). Mirror
TmuxManager's IS_TEST_MODE no-op-shell-under-VITEST: short-circuit the live
probe to {ok:true} under vitest. Command construction stays covered by
buildRemoteTmuxCheckCommand unit tests. session-routes.test.ts now 61/61.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
(cherry picked from commit 6ae2c0b8160090a1f0f6b32a3fe8496d402ac2c6)
2026-07-17 15:43:06 -04:00
267 changed files with 48113 additions and 2698 deletions
+2 -2
View File
@@ -4,7 +4,7 @@ Codeman launches AI coding sessions with `--dangerously-skip-permissions`, so th
web UI is **by design a remote-code-execution surface for whoever can reach it**.
The entire security model exists to control *who* that is. Please read this before
exposing an instance beyond `localhost`. The full model lives in
[`docs/security-architecture.md`](docs/security-architecture.md).
[`docs/security-architecture.md`](../docs/security-architecture.md).
## Supported versions
@@ -75,4 +75,4 @@ subscribe and send time), and tmux session names discovered on the shared socket
are validated against the safe-name pattern before reaching any shell call site.
For the detailed rationale, defenses, and recommended secure setups, see
[`docs/security-architecture.md`](docs/security-architecture.md).
[`docs/security-architecture.md`](../docs/security-architecture.md).
+10 -2
View File
@@ -52,12 +52,20 @@ jobs:
OLD_TAG="aicodeman@${VERSION}"
NEW_TAG="codeman@${VERSION}"
# Update the GitHub release BEFORE deleting the old tag
# Update the GitHub release BEFORE deleting the old tag.
# make_latest pins the "Latest" badge to the Codeman release. This repo
# publishes TWO packages (aicodeman + xterm-zerolag-input), changesets
# creates a GitHub release for each, and GitHub awards "Latest" to
# whichever was published LAST. That is a race: 1.9.2 kept the badge,
# 1.9.4 lost it to xterm-zerolag-input@0.1.7 by two seconds. All package
# releases already exist by the time this step runs, so setting it here
# is deterministic.
RELEASE_ID=$(gh release view "$OLD_TAG" --json databaseId -q .databaseId 2>/dev/null || true)
if [ -n "$RELEASE_ID" ]; then
gh api -X PATCH "repos/${{ github.repository }}/releases/${RELEASE_ID}" \
-f tag_name="$NEW_TAG" \
-f name="$NEW_TAG"
-f name="$NEW_TAG" \
-f make_latest=true
fi
# Retag
+6
View File
@@ -2,6 +2,9 @@
.agents/
skills-lock.json
# Written by install.sh into end-user clones when setup finishes
.install-complete
# Dependencies
node_modules/
@@ -65,6 +68,9 @@ design-explorations/
# Artifacts that should not be tracked
test-results/
tmp/
# Machine-local working files (never meant for git). ANCHORED so only the root
# dir matches.
/pr/
# Root `public` (a symlink to scripts/remotion/public — local artifact). ANCHORED
# with a leading slash so it does NOT also match src/web/public (a bare `public`
# would swallow the whole web UI source dir and silently un-stage any new asset
+3
View File
@@ -26,3 +26,6 @@ src/web/public/terminal-ui.js
src/web/public/voice-input.js
src/web/public/upload.html
scripts/remotion/
# Hand-maintained; Prettier escapes underscores in glob paths and corrupts paragraphs.
CLAUDE.md
-8
View File
@@ -1,8 +0,0 @@
{
"singleQuote": true,
"semi": true,
"tabWidth": 2,
"printWidth": 120,
"trailingComma": "es5",
"endOfLine": "lf"
}
+732
View File
@@ -1,5 +1,737 @@
# aicodeman
## 1.14.2
### Patch Changes
- Four reported bugs fixed, and the Codeman agent skill from 1.14.1 gets its first published build with the fixes below alongside it.
## The Codeman agent skill
Introduced in 1.14.1 and the headline of this line. `skills/codeman` is a Claude Code skill that lets an agent running **inside** a Codeman session drive the HTTP API: start worker sessions, send them prompts, block until they finish, read their answers and clean up. It ships in the npm package and self-gates, so outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act and costs unrelated sessions nothing.
### Installing it
```bash
codeman skill install # ~/.claude/skills/codeman, every new Claude Code session sees it
codeman skill install --case myproject # just that case; linked cases resolve by name too
codeman skill uninstall # reverses either one
```
Or turn on **App Settings > Agent Skill** (`agentSkillEnabled`, synced, default off) and Codeman injects the skill into each case when a Claude session is created there.
Installs are marker-owned: a `skills/codeman` that Codeman did not write is never touched, a stale managed copy is refreshed in place, and a symlinked skill directory is refused rather than written through. Re-run `codeman skill install` after upgrading to refresh the copy. Turning `agentSkillEnabled` back off does **not** remove already-injected copies, because a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` directory; remove them per case with `codeman skill uninstall --case <name>`.
### Using it
Ask for orchestration in plain language ("spin up three workers, have them lint, typecheck and test in parallel, then report back") and the skill supplies the guard, the safety rules and the recipes. The flow it runs:
1. **Guard.** Re-runs a preamble on every shell call that refuses outside `CODEMAN_MUX=1`, reads `CODEMAN_API_URL` and `CODEMAN_SESSION_ID`, recovers a password from the data dir `.env` or the install's service definition if one is set, and defines a fail-closed `delete_session`. It re-runs it every call because shell state does not survive between an agent's tool calls.
2. **Start a worker** with `POST /api/v1/quick-start` (`mode` is any of `claude`, `shell`, `opencode`, `codex`, `gemini`, `antigravity`), checking `.success` before reading `.data.sessionId`.
3. **Wait until it is really ready.** A new session reports `idle` before its CLI has spawned, and a brand-new case shows a trust dialog first, so the skill waits for the composer's own status bar and treats the dialog as a bounded fallback.
4. **Send and wait in one call**: `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input`. It registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer. For `claude` workers it resolves on the `stop` hook, usually within seconds.
5. **Read the answer** from `GET /api/v1/sessions/:id/last-response`, which returns clean transcript text rather than a screen scrape.
6. **Clean up** with `delete_session`, for ids it created and nothing else.
Hook-less modes (`shell` and the external CLIs) have no `stop` signal and coarse lifecycle transitions, so the skill synchronizes those with a unique split marker and `wait-output ... from=buffer`. Worked fan-out flows, the per-mode signal table, error codes and the Docker/remote caveats live in the skill's `reference/` files, loaded on demand.
### The rules it encodes
Each of these silently wastes a run, which is why they are written down: every input must end with `\r` or Enter is never sent; input is single-line; a wait timeout is HTTP 200 with `wait.timedOut`, not an error; `stop` and `blocked` are `claude`-only; signals are edge-triggered with no history, so never fire-and-forget N prompts and then gather signal-waits one by one; a typed command echoes into the output stream, so markers must be split; a full-screen TUI stream is space-less, so match single tokens; and `pid != null` proves startup, not life, so `wait?until=exit` is the death check.
## Bug fixes
- **Web tabs: long-running proxied requests were aborted after 30 seconds with no server log (#237).** The proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which bounds the entire exchange rather than the wait for response headers, so a dashboard endpoint doing model inference and any actively streaming response both died at 30s as a generic unlogged 502 that read as an intermittent network error. The timeout now bounds time-to-headers only and is cleared the moment headers arrive, with the default raised to 300s (`CODEMAN_WEBVIEW_TIMEOUT_MS`). Header timeouts are logged with a sanitized identity (method plus origin plus path, never the query string, which can carry the dashboard's tokens). A browser that navigates away mid-request now aborts the upstream fetch, guarded by `writableFinished` so a completed response never triggers it. The WebSocket handshake keeps its own 30s budget via the new `CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`, since a handshake is connection establishment and waiting minutes on one only delays the browser's reconnect logic.
- **Web tabs: sandbox incompatibility with cookie-authenticated reverse proxies documented (#238).** `docs/web-tabs.md` now covers cookie auth in front of Codeman itself (Cloudflare Access and similar), where a sandboxed frame's asset and API requests carry no auth cookie, bounce to the login provider, and leave the embedded app apparently unstyled while trusted mode works. The Test button's result now states its own scope: it verifies server-to-upstream reachability, not how the page behaves in a sandboxed frame.
- **A described session tab now shows just the description (#232).** A session named `w2-foo-bar: some description` rendered both halves, so the generated id ate the width the chosen part needed. The tab shows the description alone, the `w<n>-<case>` id moves to the tooltip and stays in the session settings modal, and `aria-label` deliberately keeps the full name so screen readers still get the id. Undescribed tabs are unchanged. Right-click a tab to rename it inline. This also fixed a re-render loop: the incremental update compared against the full name, which a described tab never matched, so those tabs re-rendered on every pass.
- **`codeman status` now probes the running server (#230).** The command runs in its own fresh process and reported that process's always-stopped Ralph loop under a bare "Status:", which reads as "the server is down" while the service is running fine and agents are reachable. It now probes the real server (`CODEMAN_API_URL`, else https then http on the local port, overridable with `--url`) and reports reachability, version and live session state; any HTTP answer proves the server is up, including a 401 from a password-protected install. The Ralph loop keeps its own `codeman ralph status`. This complements `codeman web --status` from the daemon work: that answers "did I start a daemon", this answers "is a server running at all".
## 1.14.1
### Patch Changes
- The Codeman agent skill is now installable, so an agent running inside a Codeman session can drive the API without you pasting docs into its prompt. Plus six fixes to the packaged skill, each found by running it live against a real instance.
## What the skill is
`skills/codeman` is a Claude Code skill that teaches an agent inside a Codeman session how to start worker sessions, send them prompts, block until they finish, read their answers and clean up. It ships in the npm package. It self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so installing it globally costs unrelated sessions nothing.
## Installing it
Three ways, pick one:
```bash
codeman skill install # ~/.claude/skills/codeman, every new Claude Code session sees it
codeman skill install --case myproject # just that case; linked cases resolve by name too
codeman skill uninstall # reverses either one
```
Or turn on **App Settings > Agent Skill** (`agentSkillEnabled`, synced, default off) and Codeman injects the skill into each case when a Claude session is created there.
Installs are marker-owned: a `skills/codeman` that Codeman did not write is never touched, a stale managed copy is refreshed in place, and a symlinked skill directory is refused rather than written through. Re-run `codeman skill install` after upgrading Codeman to refresh the copy.
Note that turning `agentSkillEnabled` back off does **not** remove already-injected copies, because a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` directory. Remove them per case with `codeman skill uninstall --case <name>`.
## Using it
Once installed, just ask: "spin up three workers and have them lint, typecheck and test in parallel, then report back". The skill supplies the guard, the safety rules and the recipes. What it does under the hood:
**1. Guard.** Every Bash call re-runs a preamble that refuses outside `CODEMAN_MUX=1`, reads `CODEMAN_API_URL` and `CODEMAN_SESSION_ID`, recovers a password from the data dir `.env` or the install's service definition if one is set, and defines a fail-closed `delete_session`. It re-runs it every call because shell state does not survive between an agent's tool calls.
**2. Start a worker.**
```bash
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
```
`mode` is any of `claude`, `shell`, `opencode`, `codex`, `gemini`, `antigravity`.
**3. Wait until it is actually ready.** A new session reports `idle` before its CLI has spawned, and a brand-new case shows a trust dialog first, so the skill waits for the composer's own status bar and treats the dialog as a bounded fallback.
**4. Send a prompt and wait for the turn to end.**
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"codeman-agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
Send-and-wait registers the waiter before typing, which closes the race where a separate wait reports the previous turn's idle state as this turn's answer. For `claude` workers it resolves on the `stop` hook, typically within seconds.
**5. Read the answer.**
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text'
```
**6. Clean up.** `delete_session "$SID"`, for ids you created and nothing else.
Hook-less modes (`shell` and the external CLIs) have no `stop` signal and coarse lifecycle transitions, so the skill synchronizes those with a unique split marker and `wait-output ... from=buffer` instead. Worked fan-out flows, the per-mode signal table, error codes and the Docker/remote caveats live in the skill's `reference/` files, loaded on demand.
## The rules that bite
The skill documents these because each one silently wastes a run:
- **Every input must end with `\r`** or Enter is never sent and the text sits unsubmitted on the worker's prompt. `delivered:true` means "written to the pane", not "submitted".
- **Input is single-line.** Newlines are stripped.
- **A wait timeout is HTTP 200** with `wait.timedOut:true`, not an error. Loop over short waits; timeouts clamp to [1s, 600s] and the applied value comes back as `wait.timeoutMs`.
- **`stop` and `blocked` are `claude`-only.** Requesting them elsewhere is a 400.
- **Signals are edge-triggered with no history.** One that fires while no waiter is registered is unobservable afterwards, so never fire-and-forget N prompts and then gather signal-waits worker by worker.
- **Your typed command echoes into the output stream**, so a marker that appears verbatim in the input line matches before the command runs. Split it.
- **A full-screen TUI stream is space-less**, so match a single space-free token, never a phrase.
- **`pid != null` proves startup, not life.** A worker that dies inside its pane keeps `status:"idle"` and a pid. `wait?until=exit` is the death check.
## Fixes to the packaged skill
- **The self-delete guard failed open.** The old `is_self "$SID" || curl -X DELETE ...` shape meant an undefined `is_self` exited 127, the `||` branch fired, and the agent deleted its own session with the one guard bypassed. That is reachable because shell state does not survive between tool calls, so a partially re-pasted preamble was enough. The DELETE now lives inside a fail-closed `delete_session`, which also refuses an empty id and refuses when `$SELF` is unset or too short to prove the target is not the caller.
- **`clientId` was built from `$$`.** The pid changes between tool calls, so the documented "resend the identical request" loop stopped being recognized as a duplicate and retyped the prompt, submitting the turn twice. It is a fixed literal now.
- **`GET /api/v1/sessions/:id/last-response` was undocumented.** It returns the agent's final message as clean transcript text; the terminal scrape the skill previously recommended returns a wall of TUI repaint noise with the answer buried in it. It is now the documented read path for `claude` and `codex`, with the terminal buffer demoted to diagnosis and hook-less modes. Because the transcript flush lags the `stop` signal, the recipes poll it instead of reading once.
- **`quick-start` responses were never checked for `.success`.** On failure `.data.sessionId` is absent, `jq -r` prints the string `null`, and the flow burned its full readiness budget against `/api/v1/sessions/null` before reporting jq noise instead of the cause.
- **`codeman skill install --case <name>` could not resolve a linked case.** It hardcoded `~/codeman-cases/<name>` while the server resolves through `linked-cases.json` first, so it failed with "Case not found" for a case the web UI handled fine.
- **Documentation corrections**: `SESSION_BUSY` on `quick-start` is the 50-session cap rather than the waiter cap; `caseName` resolves linked cases, so a generic name can land a worker in a real repo; and the claim that a toggle-off sweep exists was wrong, so the per-case `skill uninstall` cleanup is now stated in both the README and the code.
## Also in this release
- **Terminal**: the wheel is no longer forwarded to codex, which ignores SGR mouse reports.
## 1.14.0
### Minor Changes
- Daemon mode and service install, plus subagent hook hardening and terminal/idle-checker fixes.
**New: run Codeman in the background without a terminal (#239, closes #231)**
- `codeman web -d` starts the server detached: it survives closing the shell, logs to `~/.codeman/web.log`, records a pidfile, and only reports success after the server actually answers `/api/status` (a port clash or missing dependency can never read as a clean start). `codeman web --status` and `codeman web --stop` manage it; `--stop` verifies the pid still looks like a Codeman server before signalling, so a recycled pid is never SIGTERMed.
- `codeman service install` / `status` / `uninstall`: installs a systemd user unit (Linux) or LaunchAgent (macOS) so the server comes back after reboots. The unit carries the installing shell's PATH (launchd's default PATH finds neither an nvm/Homebrew `node` nor `tmux`/`claude`), never contains `CODEMAN_PASSWORD`, and uses the same instance-scoped unit names as `install.sh` and the self-updater so no second copy can end up supervised.
- Both refuse to start a second server on one data dir (pidfile check plus a live probe): two servers on the shared tmux socket would attach to each other's sessions.
- Why `-d` exists at all: `nohup` does not protect a Node process, Node re-arms SIGHUP even when it inherits "ignore", so `nohup codeman web &` still dies on HUP. The detached relaunch (setsid) removes the controlling terminal instead.
**Subagent background-work hooks (#233, thanks @Lint111)**
- The background Bash rewake helper now also watches the top-level parent transcript when the hook fires inside a subagent: Claude records a subagent's Bash result in its own `subagents/agent-*.jsonl` but queues the completion in the lead session transcript, so subagents previously never woke. It can also inline a `CODEMAN_RESULT_BEGIN/END` marked report (up to 64 KiB) from the task output file into the wake feedback.
- New SubagentStop guard: a subagent that still owns live Monitor or background Bash processes is kept working instead of publishing an intermediate progress line as its final report. Ownership is verified against live process descriptors on `tasks/<id>.output`, so stale transcript text alone never blocks, and the guard fails open on systems without `/proc`.
- Existing cases self-heal to the new hooks on next launch.
**AI idle checker: stderr kept out of the verdict (#234, thanks @Lint111)**
The `claude -p` verdict command no longer merges stderr into the verdict file, where CLI warnings could turn a valid verdict into a parse error. On failures, the first 200 chars of stderr are attached to the diagnostic instead.
**Terminal: large final batches drain fully (#235, thanks @Lint111)**
A render-scheduling flag was cleared after the flush instead of before it, so when a large batch left a remainder behind, the remainder stayed unrendered until unrelated output arrived. This looked like truncated responses or shell commands that never finish. The flush now reschedules itself until the queue is empty.
**Docs and tests**
- README documents daemon mode and service install.
- Unique test port for the daemon-control suite.
## 1.13.0
### Minor Changes
- Agent wait primitives, the Codeman agent skill, a fix for hooks dying silently on HTTPS installs, and the tab-strip UX improvements from the previous batch.
**Agent wait primitives (new API surface, the reason this is a minor).** Three bounded long-polls let an agent driving Codeman from a shell block instead of poll:
- `GET /api/v1/sessions/:id/wait` blocks until a lifecycle signal fires (`until=stop,idle,working,blocked,exit`, `fresh=1` to require a new transition).
- `GET /api/v1/sessions/:id/wait-output` blocks until a literal substring appears in the session's output (`match=`, `nocase=`, `from=now|buffer`; never regex, by design).
- `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input` (send-and-wait) registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer.
Shared semantics: a timeout is HTTP 200 with `wait.timedOut: true` (callers loop over short waits; tunnels cut idle connections), timeouts are clamped to [1s, 600s] and echoed back as `wait.timeoutMs`, all three nest the result under `data.wait`, and `status`/`limitPaused` ride along. `stop`/`blocked` exist for `claude` mode only: requesting them explicitly elsewhere is a 400, the default set silently narrows and echoes what it waited on. Capacity caps (16 waiters per session, 128 process-wide) answer 409/429, waiter slots release on client hang-up, and shutdown resolves parked waiters instead of stranding them. Bounds are operator-tunable via `CODEMAN_WAIT_*` env vars.
Reliability details that came out of three verification rounds: a worker that dies inside its tmux pane is now detected at the mux layer (pane-death probe, ~750ms cache, a 3s watcher for waits already parked), so a corpse answers `exit` instead of `idle` and send-and-wait rolls back its dedup seq when the write went nowhere; output matching normalizes charset-designation escapes (a stock bash prompt's `ESC ( B` no longer breaks `match=tnode:`) and holds back partial escapes at chunk boundaries, so matches straddling PTY chunks are found.
**Codeman agent skill (`skills/codeman`).** A packaged skill that teaches an agent running inside a Codeman session to drive the API safely: guard preamble (refuses outside `CODEMAN_MUX=1`, resolves credentials from the data dir `.env` or the install's service definition), self-protection (`is_self` prefix check in both directions), readiness for claude workers (composer-first, trust dialog as bounded fallback), send-and-wait loops that cannot report a never-submitted prompt as success, marker-synchronized shell flows, fan-out patterns, and cleanup discipline. Ships in the npm package via the `files` entry.
**Hooks were dying silently on every HTTPS install (bug fix).** The generated hook curls lacked `-k`, so on `--https` installs (self-signed cert) every hook event (`stop`, `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `teammate_idle`, `task_completed`) failed TLS verification and the failure was swallowed, taking respawn's definitive idle signals with it. Hooks are now generated with `curl -sk`, and a staleness detector regenerates the on-disk hook config of already-created cases the next time a session starts in them. Relatedly, `CODEMAN_API_URL` is no longer exported with a guessed `http://localhost:3000` fallback (wrong scheme on HTTPS installs); it is omitted unless the server has stamped the real URL, so in-session guards fail closed.
**Tab strip (from the previous batch, reported by christianhaberl):** action icons (kill/pop-out) now appear on the active tab only, middle-click closes a tab, tab hover uses a fixed width with a sliding title instead of resizing the strip, and the pop-out button is opt-in (default off).
**Docs.** `docs/api-reference.md` gained the full long-polling contract (signals by mode, readiness, what the matcher sees, response discriminators); `docs/extending-codeman.md` and the README carry verified copy-paste orchestration recipes; `docs/architecture-invariants.md` records the load-bearing ordering, liveness, and edge-triggered-signal invariants. Net +163 tests (4300 passing in the CI sweep).
## 1.12.2
### Patch Changes
- Codex input fixes: all four bugs reported by @DodgyBadger traced to one root cause (the zero-lag local-echo overlay buffering keystrokes until Enter, which starves codex's per-keystroke composer) and fixed in terminal-ui.js:
- Slash command picker never appeared in codex sessions (#222): the "/" sat in the overlay until Enter, so codex never saw it. Codex-mode sessions now use plain PTY echo (same branch as shell), so the picker pops and live-filters as you type.
- Arrow keys dead while typing, backspace dead after Ctrl+Backspace (#218): arrows were forwarded to a still-empty composer while typed text sat pending, and after a control-char flush the overlay swallowed every backspace. Codex bypasses the overlay entirely now; the shared overlay branch (claude/gemini/opencode) additionally flushes pending text on composer nav keys, then hands the session to pass-through until Enter/Ctrl+C, and forwards backspace instead of swallowing it when the overlay has no state.
- Pasting displaced the typed prompt (#219): bracketed pastes (xterm terminal.paste with DECSET 2004 active) were forwarded without flushing pending typed text, so the paste landed first. The shared branch now flushes typed text first and delays the paste sequence by 80ms, because codex's paste-burst handling drops keystrokes that arrive in the same PTY read as a bracketed paste (verified against codex 0.147.0 at the byte level).
- Long prompts overflowed the bottom of the screen (#220): long typed prompts existed only in the overlay DOM so codex never grew its composer; with plain PTY echo the composer grows and rewraps normally.
Verified end to end against a real codex 0.147.0 TUI driven by a headless browser: the pre-fix build reproduces all four bugs, the fixed build passes 17/17 assertions. New CI test file test/local-echo-codex-gating.test.ts (41 tests) pins the nav-key classifier, per-mode overlay gating, the flush helper, and pass-through routing. Known upstream limitation: Ctrl+Backspace deletes one character, not a word (xterm.js sends 0x08; word-delete needs kitty CSI-u encoding that xterm.js 6.0.0 cannot emit).
Mobile keyboard viewport settling fixes by @Lint111 (#229): coalesce keyboard viewport settling so rapid visualViewport resize events during keyboard show/hide no longer thrash the terminal fit, and only arm the settle logic on a real keyboard transition instead of every viewport resize.
## 1.12.1
### Patch Changes
- Terminal scrollback fixes, round 2 of issue #205. A Claude pane's local buffer is hollow (tmux keeps no history for a repaint-mode pane), and both retest reports traced back to that fact. The scroll-to-top full-history re-pull now refuses to rewrite the terminal when the capture holds less than the browser already does, so it can no longer delete history mid-scroll on iPhone (a refused session also re-fetches far less often). When wheel-forwarding is unavailable on a Claude session (version probe failed, CLI older than 2.1.187, or the "Wheel Scrolls Local History" opt-out) and there is no local scrollback to scroll, wheel and touch now page the CLI's own transcript via coalesced PageUp/PageDown instead of doing nothing. The `claude --version` probe no longer caches a failed run for the server's lifetime (one timed-out probe used to silently disable wheel-forwarding on every device until restart); failures retry with backoff. Every scroll gesture now logs a one-line `[scroll]` routing decision to the browser console for direct diagnosis, and the opt-out setting's tooltip explains that the paging fallback is Claude-only (Codex has none).
- 2e69e28: Bound the process-tree walk that could take a machine down.
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup`
under WSL, none returned while the walk kept spawning more — ~13,000 `pgrep`
processes stuck in D-state out of ~39,000 total, load average above 13,000,
recoverable only by restarting WSL.
Now: one `ps` snapshot, breadth-first with a visited set, a depth cap and a node
cap, in a pure module (`proc-tree.ts`) that the regression tests exercise
directly. The snapshot is refreshed asynchronously, and the kill path forces a
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.
- ebfcac6: An input whose delivery fails can be retried instead of being lost for good.
Both input paths recorded the `(clientId, seq)` pair as applied and acknowledged
the frame _before_ knowing whether the write had landed — the POST route because
its mux write is fire-and-forget, the WebSocket handler because it ACKed
unconditionally. When the write then failed, the client dropped the frame from its
durable queue and the server rejected the retry as a duplicate: the reliable
delivery layer was guaranteeing exactly-once delivery of something that had never
been delivered.
The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so
the client redelivers. `Session.write()` reports whether it reached a PTY at all
instead of silently swallowing the data.
Response codes are unchanged: a session can legitimately have no PTY yet (created
but not started), so turning that into a failure status would be a contract change
of its own.
Note this does not remove the root cause: the POST still answers 200 before the
mux write is attempted, so a client that treats any 2xx as final still cannot
learn about that failure. Closing that would mean awaiting the tmux child in the
request path.
- 1a32e63: Routes that answer with `reply.raw.writeHead()` no longer drop the headers the
security hook set.
`writeHead` writes straight to the Node response and bypasses Fastify's header
store, so everything the `onRequest` hook granted was silently lost — including the
`Access-Control-Allow-Origin` it emits for localhost origins, and the
`X-Content-Type-Options` / `X-Frame-Options` / CSP headers. A localhost page could
therefore call every other `/api` endpoint cross-origin while its EventSource
failed CORS.
Affects `GET /api/events` and the three raw-writing routes in `file-routes.ts`
(`file-raw`, `tail-file`, `download`).
## 1.12.0
### Minor Changes
- Terminal scrollback overhaul (issue #205), fixing every reported scroll failure across shell and CLI sessions, desktop and mobile:
- Shell, OpenCode and Antigravity sessions finally have working scrollback: tmux's own client-side alternate-screen switch is stripped for tmux-backed sessions (narrow strip: alt-screen toggles only, keeping `clear`'s 3J and mouse DECSETs), so xterm stays in the normal buffer instead of a scrollback-less alt buffer where the wheel turned into shell history cycling and touch scrolling did nothing. Direct-PTY fallback sessions are untouched so fullscreen apps (vim/less/htop) keep the alt screen there.
- The wheel listener now runs in capture phase and owns the scroll: xterm's internal vscode-style viewport scroller consumed wheel events whenever local scrollback existed (and goes deaf entirely after a tab switch or replay resets the terminal), which silently killed wheel forwarding, made scrolling break after reload/tab switches, and let the CLI's input box scroll away. Local scrolling goes through buffer-level scrollLines and keeps working after resets; mouse-tracking apps and alternate-buffer sessions are passed through untouched.
- Wheel AND touch scrolling now forward to the CLI's own transcript for Codex and Claude 2.1.187+, at any scroll position (the viewport snaps home first), so the input box stays pinned on desktop and phones alike. Shift+wheel and the "Wheel scrolls local history" setting still pin local scrollback.
- Smooth scrolling: local wheel scrolling glides with an ease-out animation (fractional line accumulation, so slow trackpad drags track the finger instead of running ahead).
- Full tmux history on demand: the full-scrollback replay is now per session instead of once per page load, and scrolling up at the top of the buffer re-pulls the complete tmux history, recovering everything tmux's repaint bursts or tab switches removed from the browser's copy.
- Firefox wheel speed: wheel deltas are normalized by deltaMode (Firefox reports line units, previously read as pixels and slowed ~4x).
- Remote SSH Claude sessions now probe the CLI version over ssh (same connection options and login-shell wrapper as the real launch), so wheel forwarding works for them too instead of silently staying off.
Docs: scrollback analysis and fix plan recorded in docs/, architecture invariants updated (strip flavors, capture-phase wheel ownership, per-session full-history replay); docker agent-image rebuild warning and integration-guide link fixes from the preceding docs commits.
## 1.11.2
### Patch Changes
- Make Antigravity (`agy`) a first-class CLI everywhere, and stop presenting Gemini CLI as a consumer product now that it is enterprise-only.
Antigravity was already wired into the session layer, schemas, run-mode menu and remote/Docker command maps, but the surfaces around it were never updated. Gemini keeps full support; Antigravity now sits beside it.
Fixes:
- **Docker cases with `mode: 'antigravity'` were broken.** `docker/agent.Dockerfile` installs its CLIs from npm, and `agy` is not an npm package, so the binary was never in the image and the container died on command-not-found. It now gets its own installer step. The `--dir /usr/local/bin` flag is load-bearing: the installer's default `$HOME/.local/bin` resolves to root's home at build time and would be unreachable by the `agent` user the container runs as. Note the binary is roughly 190MB, making it the largest layer in the image, so rebuild with `node scripts/build-agent-image.mjs` when convenient.
- **Welcome screen** gained a "Run Antigravity" action, gated on `agy` being present like the other CLI buttons, styled with the same cyan identity as the toolbar run button and run-mode dot.
- **`install.sh`** now detects `agy` (search paths mirroring `antigravity-cli-resolver.ts`), counts it as a satisfying AI CLI so an Antigravity-only box is not told it has none, and recommends it instead of Gemini in the install hints.
Documentation corrections where it had become factually wrong: `architecture-invariants.md` described `isExternalCliMode()` as opencode/codex/gemini when the code has included antigravity for some time, said "all three modes", and omitted `ANTIGRAVITY_*` from the env-prefix allowlist row; the `agentType` enum in `cron-guide.md`, `SessionMode` in `cron-discovery.md`, and `RemoteCommandMode` in `remote-sessions.md` were all stale.
Also updated both READMEs (five CLIs, Gemini marked enterprise-only), the `antigravity` npm keyword, and comment drift in eight places. Test coverage added for the new welcome button.
Antigravity stores its state under `~/.gemini/antigravity-cli/` rather than a `~/.antigravity` directory, so the existing `.gemini` Docker credential seed already covers it. That is now recorded in a code comment so no dead configuration gets added later.
- b982c5d: Keep the brief Response Viewer output inside the same message card and Markdown wrapper used by the full conversation view, so opening the viewer without clicking More preserves the same readable formatting.
## 1.11.1
### Patch Changes
- fix(history): Past Sessions data quality, and gate the phone run picker on CLI availability
**Past Sessions data quality (#215).** Three bugs in the transcript scanner behind
the Cmd+K Session Manager and the phone overview's PAST SESSIONS list:
- Automated/SDK-driven transcripts (CI review bots and other tooling, which Claude
Code stamps with a non-`cli` `entrypoint`) were listed alongside real interactive
sessions even though they were never resumable. They are now excluded. Detection
scans every entrypoint-bearing message rather than stopping at the first, so a
transcript that began under an older Claude Code build and only later picked up a
non-`cli` entrypoint is no longer wrongly hidden.
- A resumed session could show a same-directory sibling's preview text as its own.
The `workingDir` backfill in `mergeUnifiedSessions()` now only ever applies to rows
that have no history entry of their own, so it can no longer overwrite a row's real
content with another conversation's.
- Sessions restarted many times accumulated enough bookkeeping lines to push the real
first prompt past the scanner's 16KB head-read window, leaving a blank row. The read
is now two-tier: 16KB first, escalating to 128KB only when that was not enough, which
is both correct and cheaper than reading 128KB unconditionally (measured on a real
transcript tree: 36% fewer bytes read, roughly 17.5% faster than the unconditional
version). Also restores the tail-read fallback for a file whose head read failed
outright (for example `EMFILE` while scanning hundreds of files), which had been
silently dropping the session from history.
Follow-up hardening on top of the above: the automated-transcript exclusion now
blocklists the SDK entrypoint shape (`sdk`, `sdk-cli`, `sdk-py`) instead of allowlisting
the exact value `cli`. Because the check hides rows, an allowlist failed closed on any
value Claude Code has not shipped yet: a future rename of the interactive entrypoint,
or a second interactive host, would have blanked the entire Past Sessions list with
nothing in the UI to explain it. An unrecognized automated entrypoint now costs a few
noisy rows instead, which is the annoyance this filter set out to fix rather than a
broken feature.
**Phone overview run picker (#214).** The "C" logo home screen's Run picker listed all
six backends regardless of what was installed, so tapping an uninstalled one produced a
failed launch instead of the entry simply not being offered. It is now gated on
`isCliAvailable()` exactly like the desktop toolbar's run-mode dropdown (shell exempt,
since it has no external CLI dependency and keeps the menu from ever being empty). The
picker is a hardcoded duplicate of the toolbar menu rather than a shared render, which
is why it never picked up the earlier gating work; a test now asserts that every mode
the picker offers is gated, so a newly added backend cannot silently drift again.
- 73315bc: fix(web): stop the Claude response viewer from following another session's conversation
The viewer re-derived a pane's live conversation by taking the newest
`~/.claude/history.jsonl` entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` run in the user's own terminal, so the eye followed whichever of those
was typed into last — and the adoption was written back to the session, so the
mispin persisted. Entries are now credited to a pane only when they land within
10s of that pane's own Enter and no other pane on the cwd submitted closer, the
same last-submit correlation the Codex locator already uses.
That correlation also has to survive a restart. `start()` resets
`claudeSessionId` to the launch id even when re-attaching to a mux session whose
CLI has since moved on via `/clear`, so a recovered pane pointed the viewer at
its pre-`/clear` transcript — and with the anchor itself living only in memory,
nothing corrected it until the user happened to type again. `lastSubmitAt` is
now persisted in `SessionState` and restored on boot recovery, so the viewer
re-derives the live conversation on its first poll.
## 1.11.0
### Minor Changes
- Two user-facing features since 1.10.0.
**Terminal: Ctrl+C copies the selection, interrupts when nothing is selected** (#211). Copying from the terminal previously worked only through the browser context menu: xterm turns Ctrl+C into 0x03 and cancels the keydown, so the muscle-memory copy failed silently and read as "no copy-paste at all". With a selection, Ctrl+C now copies it, shows the "Copied to clipboard" toast, clears the selection and sends nothing to the PTY; with no selection it falls through unchanged, so the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never interrupts. The shortcut is a normal registry entry (`copy-selection`), so it can be rebound or disabled in App Settings, and disabling it restores plain always-interrupt Ctrl+C. Copy goes through the Clipboard API with a hidden-textarea fallback, so it also works on plain-HTTP LAN installs.
**File Viewer: edit mode for text files** (#212). The file-preview overlay can now edit workspace text files in place, phone-first: `GET /api/sessions/:id/file-content?edit=1` reads for edit without the 500-line preview truncation (saving a truncated buffer would silently delete the rest) and returns a sha256 hash plus the detected EOL; `PUT /api/sessions/:id/file-content` saves. Edit-in-place only: there is no O_CREAT anywhere in the handler, so "never create, never delete" is structural. Confinement inherits the read path (realpath plus workspace boundary, ownership scoping) and adds sensitive-path and attachment-guard blocklists, a `.git/` subtree deny, and an extension allowlist (`svg` and `env` deliberately excluded). Optimistic concurrency is by content hash, so a file changed on disk mid-edit returns 409 with an overwrite option rather than clobbering. Writes are atomic (`wx` temp, fchmod, fsync, rename) which closes the validate-then-write TOCTOU window and cannot follow a pre-existing symlink. Binary and latin-1 content are refused via a NUL sniff plus a UTF-8 round-trip compare, and EOL is re-applied server-side so a textarea's LF normalization cannot turn a two-line edit of a CRLF file into a whole-file diff.
## 1.10.0
### Minor Changes
- Codeman 1.10.0.
**Every surface that offers a CLI now checks the CLI is actually there** (#200, #201). The welcome-screen run buttons, the run-mode dropdown and the App Settings "Codex CLI" tab used to be shown unconditionally, so picking one on a box without the binary spawned a session that errored out immediately. All of them now gate on a single server-injected availability object covering Claude, OpenCode, Codex, Gemini, Antigravity and cloudflared, so nothing flickers in after paint and the dropdown costs no round trips to open. Shell is never gated, which is what keeps the menu non-empty on a box with nothing installed, and unknown availability reads as available so a stale page can never leave a working install with nothing to click. Adds `isClaudeAvailable()` and `GET /api/claude/status`, the one CLI that had no availability check despite being the default. The Cloudflare Tunnel welcome button and its scan-to-connect QR are gated on `cloudflared` rather than shown regardless.
**Shell and remote-SSH sessions now launch a real login shell** (#209, #210). Local shell tabs match what tmux itself does for a pane with no `default-command`, picking up the `/etc/profile` and `/etc/profile.d/*` entries a systemd `--user` service never sourced. On remote SSH, `claude`/`opencode`/`codex`/`gemini`/`agy` are routed through the remote user's interactive login shell, fixing agent CLIs that silently failed with "command not found" because ssh's remote-command execution sees only sshd's minimal default PATH and not the `~/.local/bin` or `~/.opencode/bin` entries where those CLIs actually live. Shell mode uses the remote user's real shell instead of hardcoded bash. The login flags are applied only to shells verified to accept them, so an exotic passwd entry (nushell, elvish, xonsh) cannot produce a dead pane on arrival.
**A crashed remote pane is kept for diagnosis** (#210), which is how the PATH failure above was found: it previously destroyed the pane, the window and the whole remote session on exit, tearing the local ssh attach down with it and leaving a flap loop with no evidence. Scoped to `remain-on-exit failed`, so a clean `exit` still tears the session down and only a non-zero exit strands anything, and applied last in the tmux command chain so a remote tmux older than 3.2 cannot drop the other session options with it.
**Resumed sessions under a hidden directory get the right working directory** (#202). Claude Code's project-key encoder maps both `/` and `.` to `-`, and the decoder could not reconstruct a dot-prefixed component, so every session under `~/.codeman` (or any project nested beneath any dotdir) silently resolved to bare `$HOME`. The wrong `workingDir` then propagated into `state.json` and everything trusting it: CLAUDE.md lookup, paste-image directory, subagent and image watchers. A same-named non-dot sibling could also produce a doubled-slash path that failed every later string comparison.
**Launching a session no longer wipes the terminal you are looking at** (#180). All six run modes route through the shared ownership helpers instead of clearing and writing into whatever session happened to be active, Antigravity included.
**Codex terminal animations are configurable** (#181), and the App Settings "Codex CLI" tab appears only where the `codex` binary resolves, since both settings on it are handed to `codex` at launch.
## 1.9.9
### Patch Changes
- Two bug fixes.
**Plain shell sessions could not start when the server process had no `SHELL` (#208).** The tmux pane command for `mode: 'shell'` was the literal string `$SHELL`. That string is embedded in the `bash -c "..."` argument of the `respawn-pane` line, which is run through `/bin/sh -c`, so it was expanded by the _server_ process's shell against the _server_ process's environment rather than inside the pane. Containers and system-level systemd units do not set `SHELL`, so it expanded to nothing and the pane command ended in a dangling `&&`, giving `bash: -c: line 1: syntax error: unexpected end of file` and a pane that died instantly (status 2) while tmux session creation still reported success. The shell is now resolved in Node (`$SHELL`, then the passwd entry, then `/bin/bash`, `/bin/zsh`, `/bin/sh`), requiring an absolute path to an executable and skipping `nologin`-style stubs, then shell-quoted. Only local shell sessions were affected: agent CLI modes emit a real command, and Docker/remote-SSH cases already used a literal `exec bash -l`.
**A session name typed into the tab options could be silently dropped.** Two independent paths. In the Session Options modal, the Session Name input saves on blur while every autosave handler bails on a null `editingSessionId`, and `closeSessionOptions()` cleared that id before hiding the modal (hiding is what blurs the input), so the save always ran too late; Escape and backdrop-click lost the name with no PUT at all, and only the X button worked because mousedown blurs first. The focused modal field is now blurred before the id is cleared, which also covers the auto-compact prompt. Separately, the right-click inline rename could be destroyed mid-keystroke: the `_inlineRenameActive` guard was missing from `_renderSessionTabsImmediate()`, so a render queued just before the rename opened still rewrote the tab name's innerHTML, committing a truncated name or closing the rename outright. The debounced executor is now guarded too.
## 1.9.8
### Patch Changes
- **Fixed: sessions failed to start on macOS with `Error: posix_spawnp failed.`** (issues #6 and #204)
`node-pty@1.1.0` publishes its macOS prebuilt helper as `prebuilds/darwin-<arch>/spawn-helper` with mode 0644, i.e. no execute bit. macOS launches every PTY through that helper, so a stock install failed on every session start. The bug is macOS-only: `spawn-helper` is a mac-only gyp target and node-pty ships no Linux prebuild, so Linux always compiles a correctly-permissioned helper from source.
The previous fix chmodded only `build/Release/spawn-helper`, which on macOS does not exist (the prebuild is used, so node-gyp never runs), and it derived that path from `require.resolve('node-pty')`, landing on `<pkg>/lib/build/Release/...`. It was a no-op on every platform.
- New `scripts/fix-node-pty.mjs` (also `npm run fix:node-pty`) chmods every `spawn-helper` it finds, in `build/Release`, `build/Debug` and each `prebuilds/*/`, then verifies the result by actually opening a PTY. A `require()` alone passes on a broken install, because the helper is only touched at spawn time.
- `postinstall` no longer force-rebuilds node-pty from source on Node 22+. That step needed Xcode command line tools, cost 30-120s on every install, and deleted the `prebuilds/` tree before compiling, so a Mac without a compiler was left with no working binary at all. A rebuild now happens only when the chmod plus spawn probe still fails, and the prebuilds tree is backed up and restored around it.
- New `spawnPtyWithHelperRepair()` (`src/utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts`, so an install that is already broken repairs itself on the first failed spawn and retries in-process instead of showing a dead session. Unrelated spawn errors are rethrown untouched; a second failure carries the `npm run fix:node-pty` hint.
- `scripts/fix-node-pty.mjs` is now in the published `files` list, so global npm installs get the repair too.
- Direct-PTY Claude spawns use the resolved absolute binary path (new `getClaudeBinaryPath()`) instead of the bare name `claude`, so a CLI installed outside the server's PATH still launches.
Verified end to end on macOS 26.4 arm64: a stock `npm i` reproduces `posix_spawnp failed.`, and after the fix the same install spawns a PTY successfully with the prebuilds preserved.
**Added: phone home screen (session overview)**
Under 430px the "C" logo now opens a session overview (current sessions, past sessions, spaces) instead of the welcome overlay: on a small screen "which session needs me" beats "how do I start one". Rows resume a session in place, and "New session here" goes through the normal quick-start path so remote and Docker cases keep their routing. Per-device setting `mobileOverviewEnabled` (phones only, default ON) in App Settings. Tablet and desktop are unchanged.
**Added: guided Tailscale setup in `install.sh`**
The network-access prompt is now 3-way: Tailscale, LAN, or local-only. The Tailscale path binds loopback and walks through installing Tailscale, logging in, the operator grant, the tailnet HTTPS-certificates toggle, and `tailscale serve --bg <port>`, then verifies the result end to end with curl. That gives HTTPS on a real certificate with no app password and no `0.0.0.0` bind, which is also what PWA install and web push need. `install.sh tailscale` retrofits it onto an existing install, and `CODEMAN_TAILSCALE=1` presets the choice. Serve state is detected from `tailscale serve status --json`; the installer never runs `tailscale serve reset` and never touches serve mappings other than 443 to Codeman's port. README and `docs/security-architecture.md` updated to match.
**Docs**: replaced a real tailnet hostname with placeholders in `docs/web-tabs-fixes-plan.md`.
**xterm-zerolag-input**: npm description and keywords only, no code change.
## 1.9.7
### Patch Changes
- Antigravity run mode, plus opt-in entrance animations.
**Antigravity CLI backend (#207).** Antigravity (`agy`) joins Claude Code, shell, OpenCode, Codex and Gemini as a sixth session backend, following the same pluggable-resolver pattern: `utils/antigravity-cli-resolver.ts` resolves the CLI and `GET /api/antigravity/status` reports availability and path. `ANTIGRAVITY_*` is added to the `ALLOWED_ENV_PREFIXES` allowlist so env overrides stay CLI-scoped rather than blanket-forwarded. Like the other external CLIs it requires tmux with no direct PTY fallback, because secrets are injected through socket-scoped `tmux setenv` and never on the spawn command line. The UI gains a Run-dropdown entry, an agent-type option, an `ag` tab badge and toolbar colours; `runAntigravity()` routes remote and docker cases through `POST /api/quick-start` and skips the local status probe for them.
**Entrance animations (opt-in, OFF by default).** Optional animations for the four things that appear when work starts: session tabs, the terminal pane a session's CLI runs in, floating agent windows, and the connection lines tying a window back to its parent tab. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. Choose a look in App Settings > Appearance > Entrance Animations (per-device, stored in localStorage rather than the settings payload); `?animlab=1` opens a per-surface picker with a live preview that fakes tabs, a pane, a window and a line so styles can be compared without spawning sessions.
Three implementation notes worth knowing if you touch this: tabs and connection lines are destroyed mid-animation on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML, `_updateConnectionLinesImmediate()` clears the SVG), so both are tracked by id and re-applied to the fresh element with a negative `animation-delay` that resumes rather than restarts them; terminal-pane styles animate transform, opacity and clip-path only, because xterm's FitAddon derives rows and columns from the untransformed layout box and animating width or height there would resize the PTY; and window styles that transform also move the rect their connection line aims at, which is why the `beam` style animates opacity and filter only.
Also fixes an agent window spawning hidden (its agent belongs to a background tab): being `display:none` it never ran its animation, so `animationend` never fired and the entrance class plus its inline custom property stuck to the window permanently. Hidden windows now skip the entrance entirely.
## 1.9.6
### Patch Changes
- Two fixes from community PRs (thanks @Lint111):
- fix(transcripts): complete tools from user-entry results (#177). Claude transcripts record tool requests in assistant entries but commonly carry their results in user-role entries; the transcript watcher only completed tools from the older assistant-entry path, so Codeman could keep showing a tool as running after it had finished. The watcher now recognizes `tool_result` blocks in user entries, ends the active tool state, and emits `transcript:tool_end` with the correct tool name and error status. Watcher tests also moved from fixed sleeps to condition-based `vi.waitFor` assertions.
- fix(notifications): quiet lifecycle hook noise (#178). Notification preferences move to schema version 5: the drawer-only "Response complete" (stop) default is now off, and the migration disables only the legacy drawer-only shape, preserving any explicit browser, audio, or push delivery the user opted into. Teammate-idle and task-completed hooks now map to the existing opt-in subagent categories instead of the broadly enabled idle/stop alerts, so normal agent activity no longer floods the drawer. Local and server-hydrated preferences are normalized through the same migration path (server hydration used to revive the retired default on fresh browsers), and the notification storage key now uses the stable handheld identity so an unfolded foldable keeps its mobile defaults and storage key (tablets and desktops unaffected).
## 1.9.5
### Patch Changes
- Background-Bash rewake hook, hooks self-heal that preserves user hooks, and test-harness isolation.
- New `PostToolUse(Bash)` hook (PR #176): a self-contained `node -e` helper watches the session transcript for a background command's completion notification and uses Claude Code's `asyncRewake` to wake an idle agent (exit code 2), without injecting terminal input that could submit a user's draft. Works on Claude Code 2.1.207+; older CLIs strip the fields harmlessly.
- Hooks self-heal (`refreshStaleHookSecret` renamed to `refreshStaleCodemanHooks`) now replaces only Codeman-owned handlers, preserving user events, matchers, and sibling handlers in mixed configurations; `writeHooksConfig` merges instead of clobbering the hooks key at case creation (PR #176).
- Rewake helper hardening: self-terminates on its own 6h deadline and when orphaned; the marker is versioned (V2) with a version-agnostic ownership prefix so future script updates replace older handlers instead of duplicating them.
- Hook timeout units fixed: the hook `timeout` field is seconds (the CLI multiplies by 1000), so `HOOK_TIMEOUT_MS = 10000` gave curl hooks a ~2.8-hour effective timeout; now `HOOK_TIMEOUT_SECONDS = 10`.
- Test-harness isolation (PR #175): every test file gets a temporary `HOME`/`USERPROFILE` so tests cannot touch real Codeman state or delete real case directories, and `Session` attaches a raw-mode echo PTY instead of a real tmux client under Vitest. Fixes the quick-start suite deleting the real `~/codeman-cases/testcase`.
- CI stability: drain console-log rpc forwards before worker teardown (fixes a run-failing `EnvironmentTeardownError` with all tests passing); `test/webview-proxy.test.ts` no longer accidentally runs under the jsdom environment via a directive named in a comment.
- Release workflow pins the GitHub "Latest" badge to the Codeman release.
## 1.9.4
### Patch Changes
- Fix a latent bug where a partial settings PUT silently reset live service state, and trim the `xterm-zerolag-input` README callout.
- **`PUT /api/settings` no longer resets watchers on a partial body.** The three `toggleService` calls (subagent watcher, workflow-run watcher, image watcher) read the raw request body with `??` defaults, so every key a caller omitted was treated as "apply the default". A body of just `{statusLineTelemetry:true}` would START the subagent watcher and STOP the workflow and image watchers, undoing the persisted config. They now resolve from `merged` (persisted settings + incoming), the same convention the `tmuxHistoryLimit` branch in that handler already used, so any PUT reconciles services to the effective stored state. Nothing triggered this in practice because every shipped client sends a full settings payload rebuilt from the DOM, but it was a trap for the next partial-update caller.
- **Regression test**: `test/routes/system-routes-settings-partial-put.test.ts` (4 cases) pins both directions, omitted keys preserve state and explicit keys still take effect. Verified to fail against the pre-fix handler.
- **CLAUDE.md** records the rule under "Adding Features → App setting": anything acting on a setting in that handler must resolve from `merged`, never the request body.
- **`xterm-zerolag-input` README**: removed the links line (getcodeman.com / install one-liner / star link) from the Codeman callout above the demo GIF. The callout keeps its links in the heading and body.
## 1.9.3
### Patch Changes
- Plan-usage chip now defaults ON on desktop, plus the reworked `xterm-zerolag-input` README.
- **Plan-usage chip defaults ON (desktop).** The `showPlanUsageLimits` chip (live 5-hour and weekly plan usage from the Claude statusline) used to be opt-in and default OFF, so most users never saw it. Desktop now defaults ON; handhelds still default OFF so the phone header stays minimal and the `mobile-header-buttons-policy` guard keeps passing. Devices with an explicitly stored preference keep whatever they chose, so nobody's OFF gets overridden.
- **One resolver behind the chip.** Added `planUsageChipEnabled()` in settings-ui.js and routed all three call sites through it: the App Settings checkbox, the chip's visibility, and the create-time `statusLineTelemetry` flag in session-ui.js. Those three had independent `?? false` / `=== true` defaults, and a chip revealed without the telemetry flag renders `—` forever, so a default flip on one site alone would have shipped a permanently empty chip.
- **Cron button comment corrected.** The App Settings comment claimed "Cron button defaults ON" while the code, the template (`btn-cron--hidden`) and the CSS all default it OFF. Verified against a fresh browser profile: the button is hidden and its checkbox unchecked out of the box. Comment now matches, and states why the two halves stay consistent.
- **Docs.** CLAUDE.md, `docs/architecture-invariants.md` and `docs/usage-limits-display-plan.md` updated for the new default and the single-resolver rule; the stale `styles.css` comment claiming the server strips the chip's hidden class at render was corrected (display is per-device, so the client reveals it).
- **`xterm-zerolag-input` README rework** (0.1.5 shipped the content; this republishes with the graphic and promo changes): replaced the misaligned 8-line keystroke-flow diagram with a two-line stock-vs-zerolag contrast, added a Codeman callout above the demo GIF with links to getcodeman.com and the repo, and rewrote the Origin section so it argues the extraction story instead of repeating the promo.
## 1.9.2
### Patch Changes
- Rewrite the `xterm-zerolag-input` package README as a value-first document and correct the drift that had accumulated against the source.
- Added the side-by-side phone demo GIF (`docs/images/zerolag-demo-20260728.gif`) as the hero image, referenced by absolute raw URL so it renders on npmjs.com as well as GitHub. The two-phone comparison shows 0ms local echo next to a 600ms-2.7s server echo on the same session.
- New "Why this one" comparison table, an explicit list of target use cases (SSH web clients, cloud IDEs, mobile terminals, container consoles), and a bundle-size badge (6.1 kB gzipped, measured from the ESM build).
- Corrected the test-count badge from 78 to the actual 175 tests across 5 files, in both the package README and the Published Packages section of the root README.
- Removed the stale "Unicode/emoji rendered at single-cell width" limitation. CJK, fullwidth forms and emoji have had double-width rendering and visual-column positioning since the wide-character fix; the honest remaining caveat (per-code-point width summing over-counts ZWJ grapheme clusters) replaces it.
- Documented the previously undocumented public `setPrompt()` method for switching prompt strategies at runtime, and the new "Wide characters (CJK, emoji)" integration section covering the optional `Unicode11Addon` path and the built-in range-table fallback.
- Documented `backgroundColor: 'transparent'`, corrected the `foregroundColor` default, and updated the grid-alignment math to reflect visual-column positioning rather than character index.
No source changes, docs only.
## 1.9.1
### Patch Changes
- Narrow the Run dropdown, and close the last two gaps in web-tab asset rewriting.
**The Run dropdown was pinned at its full width.** It capped at 300px, and the recent-session rows wanted 326px, so it always rendered at the cap and reached further across the terminal than it needed to. Now 250px, chosen as the width at which a `~/<dir>/<repo>` + timestamp row still fits whole, since identifying a session to resume is what that list is for. Three fixes were needed to make the narrower menu degrade instead of clip: the saved-URL label now has its own element, because `text-overflow` on the row button did nothing (a bare text node inside a flex container becomes an anonymous flex item that ellipsis cannot reach); `.hist-dir` got `min-width: 0`, without which a flex item refuses to shrink below its own text and pushes the date out of the box; and history rows are held to the container width, because the list's `overflow-y: auto` implicitly makes `overflow-x: auto` and let each row size to its own content and scroll sideways. Phone and tablet widths are unchanged, being set separately in `mobile.css`.
**A dashboard's own `/api/...` assets are relayed again.** The `Referer`-keyed 404 fallback, which rescues a root-absolute asset that no rewrite layer could reach, refused everything under `/api` outright. Dashboards commonly serve their assets from exactly that namespace, so those requests had no rescue at all. The refusal is now precise: the relay runs before the API-shaped 404, and the auth exemption refuses only paths that resolve to a REAL Codeman route, with `/ws/` and `/q/` still refused by prefix.
Two findings shaped that fence, both from probing Fastify rather than reading it. `hasRoute()` matches the registered PATTERN literally, so `/api/sessions/abc` reports no match against a registered `/api/sessions/:id` and would have granted an unauthenticated exemption on a live session-scoped route; `findRoute()` performs the real lookup and is what the fence uses. And `@fastify/static` is mounted at `/`, so it registers a root catch-all matching every path, which has to count as "no real route" or the fence would refuse every referer-form request and break the rescue that already worked. A root catch-all is distinguishable because it is the only route whose wildcard param comes back equal to the whole request path. The fence fails closed, and both edges are pinned in `test/webview-auth-exemption.test.ts`.
**`url()` inside runtime CSS is rewritten.** Measuring the fallback against a purpose-built dashboard showed one sink no relay can reach: a `<style>` element built by page script has no URL of its own, so the browser sends an EMPTY `Referer` with the image request it triggers. The injected URL shim now rewrites root-absolute `url()` in `<style>` blocks, both as markup and when a `<style>` node is inserted. Verified in Chromium: a stylesheet-only `/api/hero.png` and a runtime `<style>` `/api/late.png` both load, where both previously failed. The remaining known gap is self-navigation via `location.href`, which cannot be patched because `Location.href` is unforgeable.
## 1.9.0
### Minor Changes
- 2667150: feat(mobile): browse and insert local file and folder paths
Add a root-confined filesystem picker to Link Existing and the extended mobile
keyboard bar. Selected paths remain editable at the active prompt, supported
images/documents/text files open in a safe inline preview, and a new one-tap
action clears only the current unsent input without invoking `/clear`.
### Patch Changes
- 3cff98f: Fix two multi-user scoping holes in the new filesystem path picker. `GET /api/filesystem/browse` and `GET /api/filesystem/preview` accept an optional `sessionId` that contributes the session's working directory as a browse root, but they resolved it straight off the session map without an ownership check, unlike the nine other session-scoped handlers in the same route file. A non-admin could therefore pin another user's working directory as a root simply by passing their session id, then list and preview files under it. Both endpoints now run `canAccessOwned` and report 404, which also avoids confirming that a session id exists.
Separately, `Home` and `CASES_DIR` were unconditional browse roots for every caller. Per-user spaces live at `<USER_SPACES_DIR>/<username>`, which is inside `homedir()`, so the `Home` root alone exposed every other user's workspace to any authenticated user. In multi-user mode a non-admin now gets only their own space plus anything explicitly listed in `CODEMAN_FILE_PICKER_ROOTS`; `/mnt/d` is no longer offered by default, since a broad host mount should be an explicit operator decision in a multi-user deployment. Admins keep the host-wide roots, and single-user mode is unchanged.
Both holes are regression-guarded in `test/routes/file-routes.test.ts`, verified to fail against the previous code. Multi-user mode is opt-in and off by default, so single-user installs were never affected.
- Web tabs: delete saved URLs from the Run dropdown, and fix images in proxied dashboards.
**Saved URLs are now manageable from the dropdown.** Each row under "Web / URL" gains a gear and an `x`, so a URL can be edited or deleted without first opening it as a tab. Previously the only delete path ran through the gear on an open tab, which was a dead end for a URL you no longer wanted open at all. Both controls stay permanently visible rather than hover-revealed, because the same menu is used on touch, and they get a larger hit box there. Deleting leaves the dropdown open on the remaining rows, and deleting the dashboard that is currently open also closes its tab and unmounts its frame.
**Runtime-injected images no longer 404.** A dashboard that renders its own markup from script (`card.innerHTML = '<img src="/api/hero?slug=x">'`, `img.src = '/api/slide'`) escaped every rewrite layer at once: `<base href>` never applies to a root-absolute URL, the server-side attribute rewrite only ever sees the initial document, and `runtimeUrlShim()` patched only `fetch`, `XMLHttpRequest`, `WebSocket` and `EventSource`. Those requests landed on Codeman's own root and 404'd, with a symptom that reads as an upstream fault: the dashboard's data loaded while every image stayed broken.
The shim now also covers the DOM URL sinks, so the request is never emitted in the first place and neither the `/api` fence in the 404 fallback nor the one in the auth middleware had to move. It wraps `innerHTML`, `outerHTML`, `insertAdjacentHTML` (including on `ShadowRoot`), `setAttribute`/`setAttributeNS`, and the `src`/`srcset`/`href`/`poster`/`data`/`action` property setters on img, source, media, video poster, script, iframe, embed, track, link, anchor, area, object and form, with a `MutationObserver` as a last net for sinks not patched above. Every rewrite routes through the same idempotent helper, which matters because unlike the server-side rewrite this one sees markup that may already be proxied, and a page re-injecting its own `outerHTML` would otherwise double-prefix. Everything is defensively guarded and marked so a double injection cannot wrap an already-wrapped setter.
Measured against a real dashboard: 693 image elements, 0 of them under the proxy prefix and 0 of 23 in-viewport images decoded before, 693 and 23 of 23 after. Covered by a new jsdom suite over the shim's DOM half and a new frontend suite over the dropdown rows. Known remaining gaps are documented in `docs/web-tabs.md`: a root-absolute `url()` inside a stylesheet injected at runtime, and self-navigation via `location.href`, which cannot be patched because `Location.href` is unforgeable.
Also in this release: a value-first README overhaul pointing at getcodeman.com, and the QR-auth distribution test now uses a chi-square check instead of a max-deviation threshold that failed on random variance.
- bca56b4: Normalize Claude conversations in the response viewer. A Claude transcript is an append-only event log, so one logical exchange spans many JSONL rows: tool-result rows, meta/image/skill rows, compact summaries, task and team notifications, sidechains, replayed assistant snapshots, and multi-block assistant output. The viewer rendered a card per row, which produced duplicate and truncated cards that read as lost responses. Cards are now built at real human-turn boundaries, replayed assistant snapshots are deduplicated, and sidechain rows (which belong to subagents, not the main conversation) no longer leak in. An identical prompt that legitimately recurs after an assistant reply is still kept as its own turn.
Measured over 40 real transcripts: 3108 cards became 621, duplicate cards dropped from 74 to 8 (all of them genuinely repeated turns), no assistant text was lost, and the non-`context=full` last-response text was byte-identical on every file.
Also rebinds recovered sessions to their transcript. `reconcileSessions()` can recover a lost mux session as a `restored-<uuid8>` placeholder with a stale working directory, which made transcript lookup by cwd find nothing. The placeholder still carries the first eight characters of the conversation UUID, so the viewer now rebinds to the matching top-level transcript when exactly one candidate matches.
## 1.8.3
### Patch Changes
- 8c089a4: Add four light UI and terminal skins: Paper Gray, Solarized Light, Catppuccin Latte, and Rosé Pine Dawn. The Skin picker now groups Light and Dark options, and each light skin ships a matching xterm ANSI palette plus `color-scheme: light` so native selects, date pickers and scrollbars stop rendering as dark OS widgets on a light page. Terminals set `minimumContrastRatio: 4.5` under a light skin (main terminal and teammate terminals both), which keeps CLI output that assumes a dark background readable, and `applyTerminalSkin()` now refreshes the zero-lag input overlay so typed-but-unflushed text does not keep the previous theme's colors.
Elevated surfaces (modals, command palette, dropdowns, subagent and ultracode windows, file preview, attachment tray, mobile sheets) now resolve through shared `--floating-bg` / `--control-*` / `--banner-bg-*` / `--modal-backdrop` / `--elevated-shadow` tokens instead of hardcoded near-black rgba, so they follow whichever skin is active. On the Daylight skins this lifts modals slightly off the page background; OG Codeman pins its own near-black value to keep that palette neutral.
Also defines twelve CSS compatibility aliases (`--bg-primary`, `--bg-secondary`, `--bg-tertiary`, `--text-primary`, `--text-secondary`, `--border-color`, `--accent-color`, `--success`, `--error`, `--danger`, `--font-mono`, `--shadow-lg`) that panels and overlays already referenced in about 79 places but which were never actually declared, so those rules silently resolved to nothing. Status badges and accent-tinted pills (search filter chips and result badges, session tab mode pills, respawn state, Ralph priority and circuit-breaker badges, tunnel and voice status, mobile case picker) no longer keep their pale light-on-dark ink under a light skin, where it measured 1.0 to 1.9:1 and made the search filter chips invisible.
New static regression `test/skin-themes.test.ts` guards the four-way parity between the CSS token block, the xterm palette, the pre-paint allowlist and the Settings picker.
## 1.8.2
### Patch Changes
- Web tabs: open dashboard URLs as tabs beside agent sessions, plus terminal link fixes.
**Web tabs.** The Run dropdown gains a "Web / URL" section. A saved URL renders as a tab in the same strip as Claude/Codex/Gemini sessions, with the same Alt+1-9 numbering, an icon picker, and per-device tab order. Frames stay mounted while hidden (LRU-bounded), so switching tabs never reloads a dashboard.
Dashboards are proxied through Codeman's own origin, because a direct iframe fails three ways at once: an HTTPS Codeman cannot embed a plain-HTTP target (mixed content, with no override at all on iOS Safari), many dashboards send `X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks cross-origin frames. Proxying dissolves all three and leaves the production CSP unchanged. The fetch happens server-side, so a tailnet-only or localhost-only dashboard is reachable from any device that can reach Codeman.
The proxy is not an API surface: it authenticates on a 192-bit capability in the path (memory-only, rolling TTL, bound to the minting user, revoked on edit or delete) and is exempt from the cookie and Origin checks, because a sandboxed iframe is opaque-origin and sends neither. The Host allowlist is never bypassed. Iframes omit `allow-same-origin` unless a URL is explicitly marked trusted, and `Authorization` plus the session cookie are stripped upstream in both modes so `CODEMAN_PASSWORD` cannot leak into a dashboard. Includes an HTTP and WebSocket proxy, redirect/cookie/`<base>` rewriting, a runtime URL shim for requests built by dashboard JavaScript, and CORS handling for the opaque-origin frame. New endpoints under `/api/webviews`, storage in `~/.codeman/webviews.json`, user guide in `docs/web-tabs.md`.
**Terminal links no longer truncate.** Three separate cuts, each producing a link that opened the wrong target or none at all:
- A single `&` ended the match, so every query string was cut. A WordPress edit link resolved to `?post=1479` and Claude Code's own `/login` URL was unusable. `&` is now part of a URL while `&&` remains a boundary.
- Links wider than the terminal were cut at the row boundary. The link provider now stitches continuation rows into one logical line and maps offsets back across rows. Handles both soft wraps (emulator, `isWrapped`) and hard wraps (a program wrapping its own output and emitting a newline, as Ink does), the latter being why the `/login` URL grew longer as the window was widened.
- Image and PDF paths were not matched at all, so pasted-screenshot paths rendered as plain text. They now link and open the file preview, which renders images inline.
**Also fixes** a pre-existing bug where `.toolbar`'s `backdrop-filter` created a stacking context that trapped the Run menu's z-index, letting the welcome overlay cover it: with no session open, every item in that menu (Claude Code included) was unclickable.
## 1.8.1
### Patch Changes
- Mobile toolbar: a dedicated Enter button, and Shell moves into the Run dropdown.
Submitting is a constant need on a touch keyboard, so on phones (≤430px) the toolbar slot that held "Shell" now holds a dark blue **Enter** button. Starting a shell, the far rarer action, moves into the expandable Run dropdown as `Terminal / Shell` (the Run button then reads "Run SH"). Desktop and tablet are unchanged: the green Run Shell button stays exactly where it was.
Enter is replayed through the terminal's own input path rather than posted to the input API. This matters because local echo is on by default on touch devices: the characters you type are buffered client-side and have not yet reached the PTY, so sending a bare carriage return would submit an empty line and leave your text stranded on screen. Replaying the keypress flushes the buffered text first, then submits.
Installer: re-runs and updates now preserve the existing network binding instead of silently reverting it, so upgrading no longer changes how the dashboard is reachable.
Default desktop header is cleaner: the file viewer is shown by default and the plan-usage chip is unchanged, while the token-count chip and lifecycle-log button now default off. Stored preferences are still honored.
Docs and repo housekeeping: fresh phone screenshots and a new hero GIF in both READMEs, contributor and total-commit badges, and a much shorter repo root. `SECURITY.md` moved to `.github/` (GitHub resolves it there, so the Security policy tab is unaffected), `SPEEDRUN.md` to `docs/`, the knip config to `config/`, and Prettier's config into the `"prettier"` key of `package.json`. `CLAUDE.md` was split so the always-loaded guidance is roughly half its former size, with the deep implementation detail preserved verbatim in `docs/architecture-invariants.md`.
## 1.8.0
### Minor Changes
- Installer: choose your network binding, with LAN access as the new guided default.
The install script now asks at the end of setup how the dashboard should be reachable:
1. Any device on your network (0.0.0.0), the default. The installer prompts for a dashboard password (hidden input, confirmed twice); declining a password requires an explicit confirmation and the install ends with a prominent warning explaining the exposure.
2. This machine only (127.0.0.1), the safer option for tunnel/Tailscale setups.
The choice is wired into the generated systemd unit and launchd plist (values escaped for each format), the run-now launch path, and the printed URLs, which now include the detected LAN IP for instant phone access. Non-interactive installs keep the safe loopback default unless CODEMAN_HOST is preset, and the server binary's own default binding (127.0.0.1) is unchanged, so npm and manual installs behave exactly as before. New installer env presets: CODEMAN_HOST and CODEMAN_PASSWORD skip the prompts for automation.
## 1.7.1
### Patch Changes
- Mobile and UI polish plus docs refresh.
- Mobile: the header brand collapses to a single "C" home button on phones (<430px), freeing header space for session tabs while keeping the same tap target. The compact letter lives in its own span so i18n custom branding keeps rewriting only the full wordmark.
- UI fix: the absolutely-centered toolbar voice button no longer overlaps the case picker's chevron and "+" button. Below ~1500px (or with long case names widening the left toolbar group) it now falls back into normal flex flow where overlap is impossible; wide viewports keep the centered layout.
- Docs: README gains a hero pitch block with deep links, npm version + GitHub stars badges, and a star CTA; CLAUDE.md core-files table synced (Infra docker modules, app.js line count); blog article images added under docs/images/blog/.
## 1.7.0
### Minor Changes
- Community release (thanks @shenlvkang-collab for all four PRs) plus documentation fixes.
- fix(mobile): per-device settings now key off a stable handheld classification (`MobileDetection.isHandheldDevice()`: touch plus UA form-factor tokens, with User-Agent Client Hints fallback) instead of the instantaneous viewport width, so an Android foldable that unfolds past the desktop breakpoint keeps `codeman-app-settings-mobile` and opt-ins such as the Response Viewer and Extended Keyboard Bar. Responsive layout stays width-driven. Adds an OPPO Find N5 (unfolded) device profile and a fold/unfold/reload Playwright regression test (mobile suite now 136 devices). (#162)
- fix(paths): `SAFE_PATH_PATTERN` now accepts Unicode letters and numbers (`\p{L}\p{N}` with the `u` flag), so working directories like `/mnt/d/AI/中文项目` validate in Create Session, Quick Run, and Scheduled Run. All shell-metacharacter, traversal, and absolute-path protections are unchanged. (#163)
- fix(ui): newly created run sessions render their tab immediately instead of waiting for the `session:created` SSE event (idempotent upsert from the POST response, with a `GET /api/sessions/:id` fallback for quick-start modes), and the Run button holds an in-flight lock (min 500 ms) so a double click cannot create duplicate sessions. (#164)
- feat(ui): the synced custom display name and per-device English/Simplified Chinese UI language are described in their own entry (#165); on top of that PR, `renderIndexHtml` no longer recomputes `windowTitle` on solo-session renders, so a detached window cannot reset the push-notification `hostTitle` prefix to the default name.
- docs: corrected the `sse-events.ts` fileoverview breakdown (148 event constants, was stale at 120; per-category counts refreshed, including Cron, Docker, Remote auto-reconnect, and Multi-user) and the CLAUDE.md SSE registry count; READMEs synced with the 1.6.2 installer behavior.
### Patch Changes
- 8d9fc41: Add a synced custom display name and a per-device English/Simplified Chinese browser UI language picker under App Settings → Display.
## 1.6.2
### Patch Changes
- Installer (install.sh) reliability and safety overhaul, prompted by a review of the Linux flow:
- Install-completion marker (`.install-complete`): a bare re-run only takes the quiet update path when a previous install actually finished. Previously, a first install that failed during npm install/build (or was interrupted) left `.git` behind, so the retry silently became an "update" and the user never got the launch menu, the `codeman`/`tmux-chooser` symlinks, the PATH entry, or the `sc` alias. The marker is refreshed by updates and cleared by uninstall when the app dir is kept; added to .gitignore for end-user clones.
- `update` no longer runs an unconditional `git reset --hard` over local changes: interactive runs are asked to stash (declining keeps everything and skips the update), headless runs auto-stash with a dated message (same policy as scripts/self-update.sh).
- Service setup is verified instead of asserted: after starting codeman-web, the installer polls `systemctl --user is-active` (up to 6s) and only then prints "Codeman is running now!"; failures print an honest warning plus status/journalctl hints. Uses `restart` instead of `start` so re-running the installer over an already-running service actually loads the new build. A missing user D-Bus session (e.g. bare `ssh host 'curl | bash'`) is detected up front with copy-paste recovery commands instead of dying mid-setup via `set -e`. macOS gets the equivalent `launchctl list` verification, and the update path verifies its service restart too. The Cloudflare tunnel-service offer is skipped when service setup failed.
- Headless consent guard: with no interactive terminal AND no explicit `CODEMAN_NONINTERACTIVE=1`, the installer now refuses (with instructions) to run sudo package installs (git/node/tmux) or third-party `curl | bash` AI CLI installers, instead of silently taking the default-yes prompts. Explicit `CODEMAN_NONINTERACTIVE=1` keeps the previous full-auto behavior for CI/automation.
- AI CLI gate now recognizes Codex and Gemini (search paths mirrored from the CLI resolvers), so a box with only Codex or Gemini installed is no longer forced to install Claude Code/OpenCode. The install menu gains a "Skip" option (with npm install hints for Codex/Gemini), and the final reminder lists all four CLIs.
Docs: CLAUDE.md documents `src/remote-reconnect.ts` (pure COD-108 auto-reconnect backoff/eligibility logic) in the Infra table and the remote-sessions pattern.
## 1.6.1
### Patch Changes
- **Admin Panel for multi-user mode.** Admins in multi-user mode now get a prominent Admin Panel button at the top of the page (header, admin-only; the template ships it hidden and `admin-ui.js` reveals it after identity boot; hidden on phones per the mobile header policy, where user management stays reachable via App Settings > Users). It opens a full Admin Panel modal: a users table with role, enabled/disabled status, bypass-permissions grant, live sessions, active logins, case count, and last login; per-user actions for Promote/Demote, Enable/Disable, Grant/Revoke bypass, Reset password (copyable one-time password), Force logout, and Delete (with an optional "also delete their files" step); and a proper add-user form (role, optional password, bypass checkbox) replacing the old prompt() flow. Each user's cases open in a drawer listing their case folders (modified date, live-session badge) with per-folder delete. Two new admin endpoints back this: `GET /api/admin/users/:username/cases` and `DELETE /api/admin/users/:username/cases/:caseName`, guarded like `deleteUserSpace` (symlinks refused, realpath confined to the user's space, folders in use by a live session refused with 409, audit-logged). The panel and the App Settings Users tab live-refresh on the SSE `admin:usersChanged` event (now wired in app.js). New coverage in `test/admin-routes.test.ts` (list/delete, traversal + symlink refusal, non-admin 403) and `test/admin-ui.test.ts` (button reveal gating, panel render, case drawer); verified end to end against a live multi-user instance with curl and Playwright.
**Also in this release:** README/docs synced with 1.6.0 (remote SSH cases, session manager, permissions) and fixed installer prompts when run via `curl | bash`.
**Recap of the recent feature line, for readers catching up:**
- **Multi-user mode (shipped 1.5.0, opt-in `--multiuser` / `CODEMAN_MULTIUSER=1`).** Named users with scrypt-hashed passwords, per-user case spaces under `~/codeman-users/<name>/cases`, and full ownership scoping of sessions, cases, cron jobs, scheduled runs, search, file previews, and SSE/WS streams. Non-admin users default to Claude's classifier-guarded `--permission-mode auto`; shell mode, cron `launchCommand`, and skip-permissions bypass switches require the per-user `canBypassPermissions` grant (now toggleable from the Admin Panel). Admin API with one-time passwords, last-admin invariants, and an append-only audit log; self-service `/api/me` password change; `codeman users add|passwd|list|rm` CLI. Off by default is byte-identical to single-user. Note: multi-user separates workspaces for a trusted team; it is not a security boundary (all sessions share the host OS account), so pair it with Docker cases for real isolation.
- **Docker cases (shipped 1.4.0/1.4.1).** A case can run inside an isolated per-case container (any of the five CLI backends), with one-click "Run in Docker" quick-create, durable in-container tmux that survives Codeman restarts and resumes conversations after container stops, hardened container creation (cap-drop ALL, no-new-privileges, non-root, memory/pid limits, never privileged, never the docker socket), commit-safe seeded credentials, config-drift detection, GPU passthrough, and portable export/import bundles to move a whole case between machines.
- **1.6.0 highlights.** Remote SSH cases with durable remote tmux (survives SSH drops, auto-reconnect, shared multi-client attach, discover + attach with detach-not-kill); the Cmd+K session palette and unified Session Manager with pinning, cross-device tab order, and first/last prompt search; full-scrollback replay; and the multi-user permission downgrade now threading through to remote launch/attach.
## 1.6.0
### Minor Changes
- Remote tmux durability, Session Manager polish, and an opt-in Cron button.
**Remote sessions: durability, discovery, and auto-reconnect** (PR #156 by @aakhter, COD-104 to COD-109)
- Durable remote launches survive an SSH drop: the agent runs inside `tmux -L codeman-remote new-session -A` on the remote host, and reconnecting lands back in the same session.
- Discover + attach: a "Discover existing sessions" action per remote host lists `codeman-*` tmux sessions on the host's canonical socket (started by the remote's own Codeman or another instance) and attaches to one. Attached (non-owned) sessions detach on tab close, never kill; a structural early-return in `killSession()` guarantees no remote `kill-session` can ever be issued for a session Codeman doesn't own (COD-105).
- Shared/collaborative sessions: per-session `window-size latest` so concurrent clients at different viewports don't clamp each other, plus a "shared - N clients" badge in discovery results (COD-106).
- Auto-reconnect watcher: a bounded-backoff (5s to 5m, ~6 attempts) watcher detects a dead remote pane and reattaches the still-running remote tmux session; intentional kills/detaches are guarded and never revived. Kill-switch setting `remoteAutoReconnect` (default on). SSE `remote:sessionDropped`/`sessionReconnected`/`reconnectExhausted`, with a manual Reconnect toast after exhaustion (COD-108).
- Owned durable sessions propagate `kill-session` to the remote on close (COD-109); the remote tmux prereq probe is skipped under the test runner (COD-104).
- All ssh command lines continue to flow through the single shell-safe `buildSshConnectionArgs()` (COD-107). New design doc: `docs/remote-sessions.md`.
- Maintainer additions: the discovery endpoint is admin-gated in multi-user mode, and the remote launch/attach chooser threads the multi-user permission downgrade (`claudeMode`/`allowedTools`) through to the remote agent.
**Session Manager: pinning, cross-device ordering, name/prompt retention** (PR #157 by @aakhter, COD-131/139/140/142/143/145)
- Session pinning: pin a session to the top of the Session Manager list (`POST /api/sessions/:id/pin`, `session:pinned` SSE, amber highlight + pin glyph). Pinned group orders most-recently-pinned first (COD-139).
- Pinned sessions survive kill: killing a pinned session demotes its record to a lightweight stopped entry instead of removing it, so it stays visible and resumable; cleanup skips pinned records (COD-142). The pin route also works on these persisted-only records, so a pinned-then-killed session can always be unpinned.
- Cross-device tab order: tab order syncs via server state (`PUT /api/session-order`, `session:orderChanged` SSE, persisted in `state.json`); the pushing device wins and server-only ids fall to the end, never dropped (COD-131).
- Resuming from the Session Manager keeps the session's original name instead of always synthesizing a fresh `w<N>-<dir>` one (COD-143).
- firstPrompt backfill for sessions whose Codeman id is not the transcript UUID (claudeSessionId join, then newest transcript in the same workingDir), and the most recent prompt is shown alongside the first and included in search (COD-140/145).
**Cron button now opt-in** (hidden by default)
- The Cron footer-toolbar button follows the same opt-in pattern as the Session Manager / Away Digest / File Viewer buttons: hidden by default, enable per device under App Settings -> Display -> Header Displays. Cron jobs themselves are unchanged.
Also: `docs/remote-sessions.md` synced with the shipped `-L codeman-remote` / `codeman-ssh-<id8>` naming.
## 1.5.1
### Patch Changes
- Docker session-mode deep-review fixes — the work intended for the skipped **1.4.2**, now merged onto the 1.5.x line — plus a recap of the multi-user mode shipped in 1.5.0.
**Docker resume actually works now.** `DockerCase.lastClaudeSessionId` was read at quick-start but never written, so the documented resume-after-container-stop never fired. Claude-mode docker panes now pin a deterministic conversation id (`claudeDockerPaneCommand()`): a fresh launch runs `claude --session-id <id> || claude --resume <id>` (a duplicate `--session-id` exits 1 "already in use", so the fallback resumes after a container stop/reboot — verified CLI behavior), an explicit resume runs `--resume <rid> || --session-id <sid>` so a stale id never dead-panes. The id is persisted at launch and again on hook / last-response conversation-id adoption. Verified end-to-end across a `docker stop` + relaunch and a full container recreate.
**Config-drift detection + recreate (was documented but entirely missing).** The `codeman.confighash` label was stamped but never read, so docker-host config edits silently never applied. Quick-start now compares via `checkDockerConfigDrift()` and refuses a drifted launch with `CONFLICT`; the UI confirms and calls the new `POST /api/docker-cases/:name/recreate` (refused while the case has live sessions), then relaunches with the new config. New SSE event `docker:containerRecreated`.
**Model picker now applies to docker sessions.** `modelOverride` was absent from `QuickStartSchema`, so the App Settings Claude Model choice was silently inert for docker runs. It is now accepted and applied via `updateCaseModel` for local and docker quick-starts (still rejected for remote, where the settings file would land on the wrong machine).
**Import hardening.** `importDockerBundle` validates the untrusted cross-machine manifest before trusting any field (`validateImportManifest`: engine/image/containerWorkdir/network/caseName/schemaVersion — a hostile `engine` could previously select the probe binary); the outer bundle tar gets the same member-traversal guard as the inner workspace tar; the quarantine image tag derives from the schema-validated case name.
**Remote-daemon correctness.** All docker probes and the base-image auto-build now honor a host's `context`/`daemonHost` (`dockerEngineArgv`) instead of always probing the local daemon.
**Smaller fixes:** commas are rejected in docker workspace/workdir/destination paths (a comma corrupts the `--mount type=bind,src=…` CSV spec, which shell escaping cannot protect); a dead `this.escapeHtml` reference in the exports refresh is fixed; `docker:importComplete` / `docker:containerRecreated` get frontend SSE listeners so other open tabs refresh; the File Viewer header button is hidden on phone headers like its siblings.
**Docs.** CLAUDE.md + READMEs synced with the current feature set, including a full zh-CN README re-translation.
**Multi-user mode (recap — shipped in 1.5.0).** Opt-in named users (`--multiuser` / `CODEMAN_MULTIUSER=1`, off by default) with per-user case spaces and full ownership scoping of sessions, cases, cron jobs, scheduled runs, search, file previews, and real-time SSE/WS streams. Non-admin users default to Claude's classifier-guarded `--permission-mode auto`; raw shell mode, cron `launchCommand`, skip-permissions, and the Codex/Gemini bypass switches require an explicit per-user `canBypassPermissions` grant. Machine-level resources are admin-only. Admin API (`/api/admin/users*`) with one-time passwords, last-admin invariants, and an append-only audit log; self-service `/api/me` + password change; and a `codeman users add|passwd|list|rm` CLI. Off by default is byte-identical to single-user. Note: multi-user separates workspaces for a trusted team; it is not a security boundary between mutually-distrusting users (all sessions share the host OS account) — pair with Docker cases for real isolation.
## 1.5.0
### Minor Changes
- 0ab2416: Opt-in multi-user mode (`--multiuser` / `CODEMAN_MULTIUSER=1`, off by default).
Named users with individually scrypt-hashed passwords in `~/.codeman/users.json`, per-user case spaces under `~/codeman-users/<name>/cases`, and ownership scoping of sessions (create/list/delete/mutate, incl. bulk delete), cases, cron jobs + run history, scheduled runs, search, file previews, session history, away digest, subagent/workflow monitors, and real-time SSE/WS streams (including the debounced session/task update path, clipboard, and push notifications). A non-admin's `workingDir` is realpath-confined to their own space at every spawn/link path (session create, quick-start, cron create/fire, scheduled runs, case link/docker-link, docker import). Non-admin users default to Claude's classifier-guarded `--permission-mode auto`; raw shell mode, cron `launchCommand`, skip-permissions, and the Codex/Gemini bypass switches require an explicit per-user `canBypassPermissions` grant (enforced at every spawn site incl. one-shots, plan generation, scheduled runs, and remote launches). Machine-level resources (remote/Docker hosts + host reads, mux sessions, orchestrator, tunnel, self-update, settings) are admin-only. Admin API (`/api/admin/users*`) with one-time passwords, last-admin invariants (validated before any teardown), and an append-only audit log; self-service `/api/me` + password change; a frontend admin Users tab + change-password modal; and `codeman users add|passwd|list|rm` CLI. Also adds a global `auto` Claude startup permission mode. When off, behavior is byte-identical to single-user.
Auth hardening: the login throttle verifies the password before consulting the per-account failure bucket (a correct password can never be locked out); the `mustChangePassword` lockbox covers the WebSocket terminal; the cookie fast-path re-validates identity against the store each request (so a CLI/admin delete/disable/demote takes effect promptly); a role/grant change revokes the target's sessions. (Known limitation: a bare CLI `codeman users passwd` reset — no delete — does not by itself revoke an already-active cookie until it expires; use `codeman users rm`, the admin API, or a restart to force-revoke.) Data-integrity hardening: the store distinguishes a missing users file from a corrupt/unreadable one (so a transient read error can't overwrite all accounts) and writes via a unique per-process temp file; the earlier fire-and-forget `touchLastLogin` corruption race is serialized.
Note: multi-user mode separates workspaces for a trusted team; it is not a security boundary between users (all sessions share the host OS account). Pair with Docker cases for real isolation.
## 1.4.1
### Patch Changes
- **Docker session mode** hardening + fixes, plus a File Viewer header button.
**What Docker session mode is** (recap): a case can run inside an isolated, hardened Docker container instead of on the host, and any of the CLI backends (Claude, Codex, Gemini, OpenCode, or a plain shell) runs inside it. It is a location overlay on cases — not a new session mode — and the container analog of remote-SSH cases: a local tmux pane `docker exec`s into a durable in-container tmux, with exactly one long-lived container per case that multiple sessions share. The workspace, credentials, and conversation transcripts are bind-mounted so the agent is authenticated and resumable; containers are hardened by default (`--cap-drop ALL`, `--security-opt no-new-privileges`, non-root, pids/memory caps, `--init`, never `--privileged` or the docker socket) and export-safe. Start one with the one-click "Run in Docker" checkbox on Create Case, or the Docker tab for full control.
This release fixes the rough edges found running it for real:
Docker cases:
- **Seamless Claude auth in containers**: `~/.claude.json` is no longer bind-mounted as a single file (a mount point that broke Claude's atomic-rename config writes — forcing re-auth and, via failed in-place writes, corrupting the host `~/.claude.json`). It is now seeded as a writable, onboarding-complete copy, so a docker session boots straight to the prompt (no theme picker, login, or folder-trust prompt).
- **Claude-state isolation**: containers no longer bind-mount the whole `~/.claude` directory (which wrote backups/tasks/teams/settings back into the host). Only `~/.claude/projects` transcripts are shared (host watchers + `--resume`); credentials, settings, and stats-cache are seeded as writable copies; everything else stays container-local.
- **Codex/Gemini/gcloud/opencode isolation**: same treatment — codex shares `sessions/` + `history.jsonl` (response-viewer + resume) and seeds `auth.json`/`config.toml`; gemini/gcloud/opencode are whole seed-copies. Containers never write their credential state back into the host dirs.
- **Base image auto-builds on first use**: a missing `codeman/agent:base` no longer blocks case creation or launch; it builds locally on first use (concurrency-safe, with SSE progress toasts).
- **UTF-8 locale**: containers set `LANG`/`LC_ALL=C.UTF-8` so tmux renders Claude's box-drawing correctly (fixes `qqqq` line artifacts).
- **Create Case UI**: larger, collapsed-by-default "Run in Docker" settings with a shorter hint; dockerized cases show a short `(docker)` tag (or the custom host id) in the case menus.
- **Tab naming**: docker/remote (and codex/gemini/opencode) sessions now follow the `w<n>-<case>` convention instead of `codeman-<id>`.
Other:
- **File Viewer header button** (opt-in via App Settings, Header Displays): toggle the file browser panel from the header.
- Fixed a timezone-boundary flaky test in the away-digest route suite.
## 1.4.0
### Minor Changes
+171 -109
View File
File diff suppressed because one or more lines are too long
+338 -153
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -13,7 +13,10 @@
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-22%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 22+"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
<img src="https://img.shields.io/badge/Tests-2861%20total-22c55e?style=flat-square" alt="Tests">
<a href="https://www.npmjs.com/package/aicodeman"><img src="https://img.shields.io/npm/v/aicodeman?style=flat-square&label=npm&color=22c55e" alt="npm version"></a>
<a href="https://github.com/Ark0N/Codeman/stargazers"><img src="https://img.shields.io/github/stars/Ark0N/Codeman?style=flat-square&color=eab308" alt="GitHub stars"></a>
<a href="https://github.com/Ark0N/Codeman/graphs/contributors"><img src="https://img.shields.io/github/contributors/Ark0N/Codeman?style=flat-square&color=3b82f6" alt="Contributors"></a>
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p>
<p align="center">
@@ -21,7 +24,33 @@
</p>
<p align="center">
<img src="docs/images/subagent-demo.gif" alt="Codeman — parallel subagent visualization" width="900">
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
```bash
curl -fsSL https://getcodeman.com/install | bash
```
```bash
codeman web
# Open http://localhost:3000 and start your first session
```
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, five CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
- **Nothing gets lost** - tmux persistence across restarts and network drops, exactly-once input delivery, full-scrollback replay
- **Self-hosted and private** - loopback-only by default, MIT licensed, no telemetry, runs entirely on your machine
<p align="center">
<img src="docs/images/codeman-tour-20260724.png" alt="Codeman dashboard tour: session tabs per case, one-click Run for new agents, live plan usage in the header" width="900">
</p>
---
@@ -29,20 +58,56 @@
## Quick Start - Installation
```bash
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it.
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [Codex](https://developers.openai.com/codex/cli) (any combination works). After install:
- **It asks first.** Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
- **Network or local-only, your choice.** The installer asks whether the dashboard should be reachable from other devices on your network (`0.0.0.0`, the default, with a strongly recommended password prompt) or from this machine only (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. A bare `codeman web` started by hand still defaults to loopback.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the five is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
# Open http://localhost:3000 and start your first session
```
**Sharing with a small team?** Start it in multi-user mode instead: each person gets their own login and workspace.
```bash
codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
```
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
<details>
<summary><strong>Run as a background service</strong></summary>
<summary><strong>Keep it running in the background</strong></summary>
To outlive the shell you started it in, without setting anything up:
```bash
codeman web -d # detach; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful SIGTERM; agents keep running in tmux
```
`-d` waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).
To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); `codeman service` is the equivalent for an `npm i -g aicodeman` install:
```bash
codeman service install # systemd user unit (Linux) or LaunchAgent (macOS)
codeman service status
codeman service uninstall
```
`service install` writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin`, so a Homebrew or nvm `node`, `tmux` or `claude` is invisible to a hand-written plist. It never copies `CODEMAN_PASSWORD` into the unit file; add that yourself if the service needs auth.
To write the unit by hand instead:
**Linux (systemd):**
@@ -103,96 +168,27 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
<summary><strong>Windows (WSL)</strong></summary>
```powershell
wsl bash -c "curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash"
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [Codex](https://developers.openai.com/codex/cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
---
## Using Codeman — A Human's Guide
A start-to-finish walkthrough for driving Codeman from the browser. If you just installed, this is where to begin.
### 1. Launch the server
```bash
codeman web # localhost:3000 (loopback only — safe default)
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
```
Open the printed URL. The page is a single dashboard; everything below happens there.
### 2. Create your first session
Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in its own tmux-backed terminal. You choose:
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.
### 3. Read the dashboard
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder.
- **Terminal (center)** — a real `xterm.js` terminal; full TUIs render correctly. Type directly and press **Enter** to send. `Shift+Enter` inserts a newline.
- **Side panels** — Respawn, Ralph, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).
### 4. Talk to the agent
- **Type prompts** straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
- **Paste or drag-and-drop images** directly into the session.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, with auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline.
### 5. Make it autonomous
| Mode | Use it for | Where |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| **Respawn** | Long unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: `solo-work`, `overnight-autonomous`, … | Respawn tab |
| **Ralph / Todo** | A self-driving loop that tracks a todo list and keeps working until done. | Ralph tab |
| **Orchestrator** | Turn one goal into a phased plan and drive it to completion across agents. | Orchestrator panel |
| **Cron** | Saved, named jobs on a schedule (`once`/`interval`/`daily`/`weekly`) that spawn a session and send a prompt when due. | ⏰ Cron button |
| **Auto-resume** | Automatically continue after a subscription rate-limit resets. | Respawn tab (top) |
### 6. Reach it from anywhere
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
### 7. Operate & maintain
- **App Settings** — model, effort, theme/skin, notifications, display toggles, per-CLI options.
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
- **Deploy your own changes** — see [Development](#development).
> ⚠️ **Safety:** if you're working _inside_ a Codeman-managed session (`echo $CODEMAN_MUX` → `1`), never run `tmux kill-session` / `pkill claude` directly — use the web UI or `./scripts/tmux-manager.sh`.
---
## Mobile-Optimized Web UI
The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work — not a desktop UI crammed onto a small screen.
<table>
<tr>
<td align="center" width="33%"><img src="docs/screenshots/mobile-landing-qr.png" alt="Mobile — landing page with QR auth" width="260"></td>
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-idle.png" alt="Mobile — idle session with keyboard accessory" width="260"></td>
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-active.png" alt="Mobile — active agent session" width="260"></td>
<td align="center" width="40%"><img src="docs/screenshots/mobile-session-keyboard-20260727.png" alt="Mobile — answering an agent's plan prompt with the keyboard accessory bar and Enter button" width="300"></td>
<td align="center" width="60%"><img src="docs/screenshots/mobile-toolbar-enter-20260727.png" alt="Mobile toolbar: accessory bar with /init, /clear, clipboard and Esc above the Run, case, stop, Enter, voice and settings controls" width="440"></td>
</tr>
<tr>
<td align="center"><em>Landing page with QR auth</em></td>
<td align="center"><em>Keyboard accessory bar</em></td>
<td align="center"><em>Agent working in real-time</em></td>
<td align="center"><em>Answering prompts by touch</em></td>
<td align="center"><em>Accessory bar + dedicated Enter button</em></td>
</tr>
</table>
@@ -211,6 +207,18 @@ The most responsive AI coding agent experience on any phone. Full xterm.js termi
<tr><td>Password typing on phone</td><td><b>QR code scan — instant auth</b></td></tr>
</table>
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard; destructive commands require a double-press to confirm, so you never fire one by accident
- **Dedicated Enter button** — replays the keypress through the terminal, so text buffered by local echo is flushed first rather than stranded
- **Swipe navigation & smart keyboard handling** — swipe left/right to switch sessions; toolbar and terminal shift up when the keyboard opens (`visualViewport` API)
- **Built for phones** — safe-area insets for notch and home indicator, 44px touch targets, bottom-sheet case picker, native momentum scrolling
```bash
codeman web --https
# Open on your phone: https://<your-ip>:3000
```
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended): the installer can set it up for you (choose **Tailscale** at the network-access prompt, or run `bash ~/.codeman/app/install.sh tailscale` on an existing install). That gives you `https://<your-machine>.<tailnet>.ts.net` with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone.
### Secure QR Code Authentication
Typing passwords on a phone keyboard is miserable. Codeman replaces it with **cryptographically secure single-use QR tokens** — scan the code displayed on your desktop and your phone is authenticated instantly.
@@ -219,47 +227,81 @@ Each QR encodes a URL containing a 6-character short code that maps to a 256-bit
The security design addresses all 6 critical QR auth flaws identified in ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025, which found 47 of the top-100 websites vulnerable): single-use enforcement, short TTL, cryptographic randomness, server-side generation, real-time desktop notification on scan (QRLjacking detection), and IP + User-Agent session binding with manual revocation. Dual-layer rate limiting (per-IP + global) makes brute force infeasible across 62^6 = 56.8 billion possible codes. Full security analysis: [`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
### Touch-Optimized Interface
- **Keyboard accessory bar** — `/init`, `/clear`, `/compact` quick-action buttons above the virtual keyboard. Destructive commands (`/clear`, `/compact`) require a double-press to confirm — first tap arms the button, second tap executes — so you never fire one by accident on a bumpy commute
- **Swipe navigation** — left/right on the terminal to switch sessions (80px threshold, 300ms)
- **Smart keyboard handling** — toolbar and terminal shift up when keyboard opens (uses `visualViewport` API with 100px threshold for iOS address bar drift)
- **Safe area support** — respects iPhone notch and home indicator via `env(safe-area-inset-*)`
- **44px touch targets** — all buttons meet iOS Human Interface Guidelines minimum sizes
- **Bottom sheet case picker** — slide-up modal replaces the desktop dropdown
- **Native momentum scrolling** — `-webkit-overflow-scrolling: touch` for buttery scroll
```bash
codeman web --https
# Open on your phone: https://<your-ip>:3000
```
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended) — it provides a private network so you can access `http://<tailscale-ip>:3000` from your phone without TLS certificates.
---
## Live Agent Visualization
## Using Codeman — A Human's Guide
Watch background agents work in real-time. Codeman monitors agent activity and displays each agent in a draggable floating window with animated Matrix-style connection lines back to the parent session.
A start-to-finish walkthrough for driving Codeman from the browser. If you just installed, this is where to begin.
<p align="center">
<img src="docs/images/subagent-spawn.png" alt="Subagent Visualization" width="900">
</p>
### 1. Launch the server
- **Floating terminal windows** — draggable, resizable panels for each agent with a live activity log showing every tool call, file read, and progress update as it happens
- **Connection lines** — animated green lines linking parent sessions to their child agents, updating in real-time as agents spawn and complete
- **Status & model badges** — green (active), yellow (idle), blue (completed) indicators with Haiku/Sonnet/Opus model color coding
- **Auto-behavior** — windows auto-open on spawn, auto-minimize on completion, tab badge shows "AGENT" or "AGENTS (n)" count
- **Nested agents** — supports 3-level hierarchies (lead session -> teammate agents -> sub-subagents)
```bash
codeman web # localhost:3000 (loopback only — safe default)
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
codeman web -d # detach: survives closing the shell (--status, --stop)
codeman service install # systemd/launchd service: comes back after reboots
```
**Agent Teams** — first-class support for Claude Code's native multi-agent teams (`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`). `TeamWatcher` polls `~/.claude/teams/`, matches teammates to their lead session, and surfaces them as live subagent windows with **team-aware idle detection** — so the Respawn Controller won't fire while teammates are still working. See [`docs/agent-teams/`](docs/agent-teams/).
Open the printed URL. The page is a single dashboard; everything below happens there.
### 2. Create your first session
Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in its own tmux-backed terminal. You choose:
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.
### 3. Read the dashboard
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder (tab order syncs across your devices).
- **Terminal (center)** — a real `xterm.js` terminal; full TUIs render correctly. Type directly and press **Enter** to send. `Shift+Enter` inserts a newline.
- **Side panels** — Respawn, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).
### 4. Talk to the agent
- **Type prompts** straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
- **Paste or drag-and-drop images** directly into the session.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, with auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline.
### 5. Make it autonomous
| Mode | Use it for | Where |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------- |
| **Respawn** | Long unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: `solo-work`, `overnight-autonomous`, … | Respawn tab |
| **Orchestrator** | Turn one goal into a phased plan and drive it to completion across agents. | Orchestrator panel |
| **Cron** | Saved, named jobs on a schedule (`once`/`interval`/`daily`/`weekly`) that spawn a session and send a prompt when due. | ⏰ Cron button _(opt-in: App Settings → Display → Header Displays)_ |
| **Auto-resume** | Automatically continue after a subscription rate-limit resets. | Respawn tab (top) |
### 6. Reach it from anywhere
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
### 7. Operate & maintain
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
- **Deploy your own changes** — see [Development](#development).
> ⚠️ **Safety:** if you're working _inside_ a Codeman-managed session (`echo $CODEMAN_MUX` → `1`), never run `tmux kill-session` / `pkill claude` directly — use the web UI or `./scripts/tmux-manager.sh`.
---
## Zero-Lag Input Overlay
<p align="center">
<img src="docs/images/zerolag-demo.gif" alt="Zerolag Demo — local echo vs server echo side-by-side" width="900">
<img src="docs/images/zerolag-demo-20260728.gif" alt="Zerolag demo: instant local echo next to 600ms-2.7s server echo, side by side on two phones" width="900">
</p>
When accessing your coding agent remotely (VPN, Tailscale, SSH tunnel), every keystroke normally takes 200-300ms to round-trip. Codeman implements a **Mosh-inspired local echo system** that makes typing feel instant regardless of latency.
@@ -276,6 +318,30 @@ A pixel-perfect DOM overlay inside xterm.js renders keystrokes at 0ms. Backgroun
---
## Live Agent Visualization
Watch background agents work in real-time. Codeman monitors agent activity and displays each agent in a draggable floating window with animated Matrix-style connection lines back to the parent session.
<p align="center">
<img src="docs/images/subagent-windows-20260724.png" alt="Subagent Visualization: three parallel Explore agents as floating windows with live tool-call feeds" width="900">
</p>
- **Floating terminal windows** — draggable, resizable panels for each agent with a live activity log showing every tool call, file read, and progress update as it happens
- **Connection lines** — animated green lines linking parent sessions to their child agents, updating in real-time as agents spawn and complete
- **Status & model badges** — green (active), yellow (idle), blue (completed) indicators with Haiku/Sonnet/Opus model color coding
- **Auto-behavior** — windows auto-open on spawn, auto-minimize on completion, tab badge shows "AGENT" or "AGENTS (n)" count
- **Nested agents** — supports 3-level hierarchies (lead session -> teammate agents -> sub-subagents)
Multi-agent Workflow runs ("ultracode") get the same treatment: a floating run window tracks the whole workflow live, with phases, per-agent token counts, and the current tool of every agent:
<p align="center">
<img src="docs/images/ultracode-window-20260724.png" alt="Ultracode workflow visualization: a live run window with per-agent tokens and phases" width="900">
</p>
**Agent Teams** — first-class support for Claude Code's native multi-agent teams (`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`). `TeamWatcher` polls `~/.claude/teams/`, matches teammates to their lead session, and surfaces them as live subagent windows with **team-aware idle detection** — so the Respawn Controller won't fire while teammates are still working. See [`docs/agent-teams/`](docs/agent-teams/).
---
## Respawn Controller
The core of autonomous work. When the agent goes idle, the Respawn Controller detects it, sends a continue prompt, cycles context management commands for fresh context, and resumes — running **24+ hours** completely unattended.
@@ -302,7 +368,7 @@ Beyond single-session respawn, the **Orchestrator** turns a high-level goal into
- **Crash-safe** — full state persists under the `orchestrator` key in `state.json`, so it survives restarts
- **Driven from the UI or API** — the Orchestrator panel, or `POST /api/orchestrator/start` → `/approve` → `/status` (10 endpoints)
> Distinct from Ralph (a single-session autonomous loop): the orchestrator coordinates multi-phase, multi-agent execution. Full design: [`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md).
> Full design: [`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md).
---
@@ -310,14 +376,18 @@ Beyond single-session respawn, the **Orchestrator** turns a high-level goal into
Run **20 parallel sessions** with full visibility — real-time xterm.js terminals at 60fps, per-session token and cost tracking, tab-based navigation, and one-click management.
<p align="center">
<img src="docs/screenshots/multi-session-dashboard.png" alt="Multi-Session Dashboard" width="800">
</p>
### Persistent Sessions
Every session runs inside **tmux** — sessions survive server restarts, network drops, and machine sleep. Auto-recovery on startup with dual redundancy. Ghost session discovery finds orphaned tmux sessions. Managed sessions are environment-tagged so the agent won't kill its own session.
### Session Manager & Command Palette
`Ctrl/Cmd/Alt+K` opens a fuzzy session palette; **Browse all sessions** opens the Session Manager: one deduped list of everything Codeman knows about (live sessions, past sessions from state and lifecycle history, and Claude transcripts), each row showing its first and most recent prompt.
- **Pinning**: pin a session to float it to the top of the list. Pinned sessions even survive kill (they demote to a lightweight stopped entry that stays visible and resumable).
- **Name retention**: resuming a past session keeps its original name instead of minting a new one.
- **Cross-device tab order**: drag-reordered tabs persist server-side, so your ordering follows you from desktop to phone.
### Hostname-Aware Window Title
Running Codeman on multiple hosts (laptop, dev box, NAS)? The browser tab title is `codeman:<hostname>` so you can tell which backend each tab points at without clicking in:
@@ -340,14 +410,6 @@ The title is templated into the served HTML on first byte, so it's correct from
Real-time desktop alerts when sessions need attention — `permission_prompt` and `elicitation_dialog` trigger critical red tab blinks, `idle_prompt` triggers yellow blinks. Click any notification to jump directly to the affected session. Hooks auto-configured per case directory.
### Ralph / Todo Tracking
Auto-detects Ralph Loops, `<promise>` tags, TodoWrite progress (`4/9 complete`), and iteration counters (`[5/50]`) with real-time progress rings and elapsed time tracking.
<p align="center">
<img src="docs/images/ralph-tracker-8tasks-44percent.png" alt="Ralph Loop Tracking" width="800">
</p>
### Run Summary
Click the chart icon on any session tab to see a timeline of everything that happened — respawn cycles, token milestones, auto-compact triggers, idle/working transitions, hook events, errors, and more.
@@ -364,18 +426,72 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, or **Codex** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
- **Voice input** — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Image input** — paste or drag-and-drop images straight into a session
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Display
- **Multi-monitor span** _(macOS)_ — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
- **File Viewer button** _(opt-in)_ — a header button that toggles the built-in file browser panel with one tap; enable under App Settings → Display → Header Displays
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean
- **OS notifications & hostname-aware titles** — desktop alerts and tab titles are prefixed `codeman:<host>` so multi-host setups stay unambiguous
---
## Isolated Docker Sessions
Run a case inside its own hardened Docker container instead of directly on your host — for security isolation, reproducible toolchains, and one-click portability.
- **One click** — on **New Case → Create New**, tick **🐳 Run in an isolated Docker container**. Codeman creates the case folder, spins up a container with default settings, and starts the agent inside it. No host/image/network fields to fill in.
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with `node scripts/build-agent-image.mjs`). Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
---
## Remote SSH Sessions
Point a case at another machine and run the agent **there**, over SSH, with the same dashboard, mobile UI, and autonomy features. Your laptop is just a window onto a session that lives on the remote host.
- **Durable by design**: the agent runs inside a dedicated tmux session on the remote host, so a dropped SSH connection, network change, or laptop sleep never kills the run. Reconnecting lands back in the same live conversation.
- **Auto-reconnect**: a bounded-backoff watcher notices a dead SSH pane and silently reattaches to the still-running remote session (kill-switch in settings; intentional kills are never revived).
- **Discover & attach**: list the `codeman-*` sessions already running on a host (started by that machine's own Codeman, or by another operator) and attach to one. Attached sessions you don't own **detach on tab close, never kill**.
- **Shared sessions**: several clients can attach the same remote session at different window sizes without clamping each other; discovery shows a "shared" badge with the client count.
- **Injection-safe**: every ssh command line flows through a single shell-escaping builder, and host/path/identity fields are schema-guarded.
Set it up under **New Case → Remote** (host, user, identity file, optional jump host). Full design: [`docs/remote-sessions.md`](docs/remote-sessions.md).
---
## Multi-User Mode (opt-in)
Share one Codeman with a small trusted team, each person getting their own login and workspace. **Off by default** — without the flag, nothing changes.
Enable with `codeman web --multiuser` (or `CODEMAN_MULTIUSER=1`). Create the first admin, then manage users from the CLI or the **Users** tab in App Settings:
```bash
codeman users add alice --admin # prompts for a password (or --password-stdin)
codeman users add bob # a regular user
codeman users list
```
- **Per-user spaces** — each user's cases live under `~/codeman-users/<name>/cases`; sessions, cases, search, and real-time events are scoped to their owner. Admins see everything.
- **Individually revocable logins** — named users with scrypt-hashed passwords in `~/.codeman/users.json`; disable, reset (one-time password), or delete an account at any time. Admin actions are audited to `~/.codeman/admin-audit.jsonl`.
- **Safer defaults for regular users** — non-admins run Claude in `--permission-mode auto` (Anthropic's classifier-guarded mode); raw shell sessions, cron `launchCommand`, and skip-permissions require an explicit per-user grant.
> ⚠️ **This separates workspaces; it does not sandbox users from each other.** Every session runs as the same OS account, so a determined user's agent can still reach another user's files. For real isolation, pair users with **Docker cases** or run separate instances under separate OS accounts. See [`docs/multi-user-plan.md`](docs/multi-user-plan.md) and the multi-user section of [`docs/security-architecture.md`](docs/security-architecture.md).
---
## Remote Access — Cloudflare Tunnel
Access Codeman from your phone or any device outside your local network using a free [Cloudflare quick tunnel](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/) — no port forwarding, no DNS, no static IP required.
@@ -499,13 +615,14 @@ When someone authenticates via QR, the desktop shows a notification toast with t
## Security
Codeman launches sessions with `--dangerously-skip-permissions`, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control _who_ that is. Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: [`docs/security-architecture.md`](docs/security-architecture.md). **Found a vulnerability?** See [`SECURITY.md`](SECURITY.md) for private disclosure and the list of known limitations.
By default Codeman launches sessions with `--dangerously-skip-permissions`, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control _who_ that is. (The startup permission mode is configurable; see below.) Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: [`docs/security-architecture.md`](docs/security-architecture.md). **Found a vulnerability?** See [`SECURITY.md`](.github/SECURITY.md) for private disclosure and the list of known limitations.
### Network & access
- **Loopback by default** — binds `127.0.0.1`, reachable only from the same machine, so the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` _starts but prints a loud warning_ with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with `--allow-unauthenticated-network`)
- **Loopback by default** — the server binary binds `127.0.0.1`, reachable only from the same machine, so the no-password default is safe out of the box (the guided installer asks about network access and configures the binding + password for you). Binding a non-loopback host without `CODEMAN_PASSWORD` _starts but prints a loud warning_ with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with `--allow-unauthenticated-network`)
- **Optional auth, real sessions** — HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Success issues an opaque 256-bit `codeman_session` cookie (`randomBytes(32)`) — validated server-side, not client-signed, so it can't be forged offline (24h TTL, auto-extend, device-context audit log)
- **Per-IP rate limiting** — 10 failed attempts → `429` with `Retry-After` (15-min decay). A valid cookie or correct password recovers _immediately_ even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter
- **Configurable permission mode** - `--dangerously-skip-permissions` is only the default. **App Settings → Claude CLI → Startup Mode** can switch new sessions to Anthropic's classifier-guarded `auto` mode (low-prompt, needs Claude Code 2.1.207+), `normal` prompting, or an explicit allowed-tools list. In multi-user mode, non-granted users are forced to `auto`, and shell sessions / skip-permissions require an explicit per-user grant
### Always-on browser hardening (v0.9.5)
@@ -519,7 +636,7 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` env-prefix allowlist gates which settings each CLI can receive
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
@@ -558,6 +675,8 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Alt/Option+[` / `Alt/Option+]` | Previous / next session |
| `Alt/Option+1`-`Alt/Option+9` | Switch to tab N (physical keys, so macOS Option layouts work) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+C` | Copy selection, or interrupt when nothing is selected |
| `Ctrl+Shift+C` | Copy selection (never interrupts) |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
@@ -572,6 +691,18 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
> **Shortcut: install the packaged agent skill.** Everything below (plus worked multi-worker recipes) ships as a Claude Code skill in [`skills/codeman`](skills/codeman/SKILL.md), so an agent inside a session can drive Codeman without you pasting docs into the prompt. Three ways to get it:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`: global, works for any skills-aware agent
> - `codeman skill install` (global) or `codeman skill install --case <name>`: for npm installs that never cloned the repo; `codeman skill uninstall` reverses it
> - **App Settings → Agent Skill** (`agentSkillEnabled`, default off): Codeman then injects the skill into each case on Claude session create; a user-authored `skills/codeman` in the case is never overwritten
>
> A global install (`codeman skill install`, or `npx skills add`) is picked up by **every new Claude Code session on the machine**, inside Codeman or not. The skill self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so a global install costs an idle session nothing.
>
> ⚠️ Turning `agentSkillEnabled` back off **does not remove already-injected copies** (a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` dir). Remove them per case with `codeman skill uninstall --case <name>`.
### Detect that you're inside Codeman
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
@@ -585,15 +716,21 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
### Rules of the road (read before you POST)
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
1. **Single-line input, ending in `\r`.** Programmatic input is sent as literal text, and Enter fires **only when the input contains a carriage return**: `{"input":"run tests\r"}`. Without the `\r` the text sits on the session's prompt unsubmitted (and a combined `wait` runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs the joined command `echo Aecho B`: send one line per call.
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard). ⚠️ A `401` replies with the bare string `Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error instead of showing the failure: check the status before parsing.
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
```bash
# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
# The fallback below fits a stock install; on a --https install set the https:// URL
# yourself and add -k to each curl (self-signed cert).
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
@@ -605,18 +742,63 @@ curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
# first-run trust dialog only as the fallback. (Probing trust first and sending
# a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
# so the probe matches stale text and the Enter lands in a ready composer.)
# Match single tokens: TUI text can reach the matcher without its spaces.
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # accept the first-run trust dialog
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. Read the terminal back
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
# writing, so it can't answer with the previous turn's idle state)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
# spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)
# 5. Stream live events (session output, agent activity, status)
# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. Or wait for a marker in the output (works for shell sessions too).
# ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
# typed line never contains it: your own keystrokes echo into the output
# stream, so an unsplit marker matches before the command has run. from=buffer
# catches a marker that printed before the wait landed.
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
# tail counts BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events" # Server-Sent Events
# 6. Schedule recurring work (cron-style job)
# 7. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
@@ -624,11 +806,11 @@ curl -s -X POST "$API/api/cron/jobs" \
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. Inspect background sub-agents and their transcripts
# 8. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
# 9. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq
```
@@ -641,7 +823,6 @@ codeman session start -d /path/to/repo # (s) start a session
codeman session list # list sessions
codeman session logs <id> # tail output
codeman task add "fix the failing test" # (t) queue a task
codeman ralph start --min-hours 8 # (r) launch the autonomous loop
codeman attach <path> # attach a Claude hook context
```
@@ -655,7 +836,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
## API
REST over Fastify — **~160 handlers across 18 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
@@ -663,8 +844,14 @@ REST over Fastify — **~160 handlers across 18 route modules**, plus an SSE str
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | List all |
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
| `GET` | `/api/sessions/:id/output` | Read terminal output |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`: `clientId`+`seq` = exactly-once; `wait` blocks until the turn ends) |
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
| `DELETE` | `/api/sessions/:id` | Delete session |
### Respawn
@@ -675,13 +862,6 @@ REST over Fastify — **~160 handlers across 18 route modules**, plus an SSE str
| `POST` | `/api/sessions/:id/respawn/stop` | Stop controller |
| `PUT` | `/api/sessions/:id/respawn/config` | Update config |
### Ralph / Todo
| Method | Endpoint | Description |
| ------ | -------------------------------- | ---------------------- |
| `GET` | `/api/sessions/:id/ralph-state` | Get loop state + todos |
| `POST` | `/api/sessions/:id/ralph-config` | Configure tracking |
### Orchestrator
| Method | Endpoint | Description |
@@ -722,6 +902,8 @@ REST over Fastify — **~160 handlers across 18 route modules**, plus an SSE str
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
---
## Architecture
@@ -744,7 +926,6 @@ flowchart TB
end
subgraph Detection["Detection Layer"]
RT["Ralph Tracker"]
SW["Subagent Watcher<br/><small>~/.claude/projects/*/subagents</small>"]
TW["Team Watcher<br/><small>~/.claude/teams/*</small>"]
end
@@ -755,7 +936,7 @@ flowchart TB
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
@@ -768,7 +949,6 @@ flowchart TB
SM --> RC
SM --> ORC
SM --> SS
S1 --> RT
S1 --> SCR
S2 --> SCR
RC --> SCR
@@ -787,7 +967,7 @@ flowchart TB
npm install
npx tsx src/index.ts web # Dev mode
npm run build # Production build
npm test # Run tests
npm run test:ci # Run tests (the CI suite; browser suites need extra setup)
```
See [CLAUDE.md](./CLAUDE.md) for full documentation.
@@ -817,7 +997,7 @@ Full details: [`docs/archive/code-structure-findings.md`](docs/archive/code-stru
[![npm](https://img.shields.io/npm/v/xterm-zerolag-input?style=flat-square&color=22c55e)](https://www.npmjs.com/package/xterm-zerolag-input)
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, configurable prompt detection, full state machine with 78 tests.
Instant keystroke feedback overlay for xterm.js. Eliminates perceived input latency over high-RTT connections by rendering typed characters immediately as a pixel-perfect DOM overlay. Zero dependencies, 6.1 kB gzipped, configurable prompt detection, CJK/emoji wide-character support, full state machine with 175 tests.
```bash
npm install xterm-zerolag-input
@@ -844,3 +1024,8 @@ MIT — see [LICENSE](LICENSE)
<p align="center">
<strong>Track sessions. Visualize agents. Control respawn. Let it run while you sleep.</strong>
</p>
<p align="center">
If Codeman saves you time, <a href="https://github.com/Ark0N/Codeman/stargazers">a star</a> helps other people find it.<br>
Bug reports and feature ideas are welcome in <a href="https://github.com/Ark0N/Codeman/issues">Issues</a>.
</p>
+412 -155
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -14,39 +14,73 @@
<p align="center">
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="License: MIT"></a>
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-18%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 18+"></a>
<a href="https://nodejs.org/"><img src="https://img.shields.io/badge/Node.js-22%2B-22c55e?style=flat-square&logo=node.js&logoColor=white" alt="Node.js 22+"></a>
<a href="https://www.typescriptlang.org/"><img src="https://img.shields.io/badge/TypeScript-5.9-3b82f6?style=flat-square&logo=typescript&logoColor=white" alt="TypeScript 5.9"></a>
<a href="https://fastify.dev/"><img src="https://img.shields.io/badge/Fastify-5.x-1e3a5f?style=flat-square&logo=fastify&logoColor=white" alt="Fastify"></a>
<img src="https://img.shields.io/badge/Tests-2861%20total-22c55e?style=flat-square" alt="Tests">
<a href="https://github.com/Ark0N/Codeman/graphs/contributors"><img src="https://img.shields.io/github/contributors/Ark0N/Codeman?style=flat-square&color=3b82f6" alt="Contributors"></a>
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p>
<p align="center">
<img src="docs/images/subagent-demo.gif" alt="Codeman — 并行子智能体可视化" width="900">
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — 并行子智能体可视化" width="900">
</p>
<p align="center">
<img src="docs/images/codeman-tour-20260724.png" alt="Codeman 仪表盘导览:按项目分组的会话标签页、一键 Run 启动新智能体、页头实时用量" width="900">
</p>
> 本文档由英文版 [`README.md`](README.md) 翻译而来。如有出入,以英文版为准。
---
## 快速开始 — 安装
一行命令即可安装(macOS 和 Linux,Windows 通过 WSL):
```bash
curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash
curl -fsSL https://getcodeman.com/install | bash
```
该脚本会在缺失时自动安装 Node.js 和 tmux,把 Codeman 克隆到 `~/.codeman/app` 并完成构建。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai) 或 [Codex](https://developers.openai.com/codex/cli)(任意组合均可)。安装完成后:
```bash
codeman web
# 打开 http://localhost:3000,开启你的第一个会话
```
安装器在每次系统改动前都会先询问;重跑同一条命令即可原地更新。详见[快速开始 — 安装](#快速开始--安装)。
---
## 快速开始 — 安装
```bash
curl -fsSL https://getcodeman.com/install | bash
```
该脚本会在缺失时自动安装 Node.js 和 tmux,把 Codeman 克隆到 `~/.codeman/app` 并完成构建。几点须知:
- **先询问,后改动。** 所有系统级改动(安装软件包、下载 AI CLI)都会先征求确认;结束时的菜单可选择:直接在本终端运行、安装为后台服务(systemd/launchd,开机自启),或暂不启动。不选就不会有任何后台进程。
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这五个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
# 打开 http://localhost:3000,开启你的第一个会话
```
**想和小团队共用一台?** 改用多用户模式启动:每人拥有自己的登录与工作空间。
```bash
codeman users add alice --admin # 创建第一个管理员账号
codeman web --multiuser # 命名登录 + 按用户隔离的案例空间
```
详见下文[多用户模式](#多用户模式可选启用)。
<details>
<summary><strong>作为后台服务运行</strong></summary>
安装器结尾的菜单(选项 2)可以帮你完成这一步,并在宣告成功前校验服务确实已启动。如需手动配置:
**Linux(systemd):**
```bash
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
@@ -69,6 +103,7 @@ loginctl enable-linger $USER
```
**macOS(launchd):**
```bash
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
@@ -96,16 +131,18 @@ cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
</details>
<details>
<summary><strong>Windows(WSL)</strong></summary>
```powershell
wsl bash -c "curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash"
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai) 或 [Codex](https://developers.openai.com/codex/cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
---
@@ -116,14 +153,12 @@ Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.
<table>
<tr>
<td align="center" width="33%"><img src="docs/screenshots/mobile-landing-qr.png" alt="移动端 — 带二维码认证的登录页" width="260"></td>
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-idle.png" alt="移动端 — 带键盘配件栏的空闲会话" width="260"></td>
<td align="center" width="33%"><img src="docs/screenshots/mobile-session-active.png" alt="移动端 — 活动中的智能体会话" width="260"></td>
<td align="center" width="40%"><img src="docs/screenshots/mobile-session-keyboard-20260727.png" alt="移动端 — 通过键盘配件栏与 Enter 按钮回答智能体的方案提示" width="300"></td>
<td align="center" width="60%"><img src="docs/screenshots/mobile-toolbar-enter-20260727.png" alt="移动端工具栏:配件栏的 /init、/clear、剪贴板与 Esc,下方是 Run、案例、停止、Enter、语音与设置控件" width="440"></td>
</tr>
<tr>
<td align="center"><em>带二维码认证的登录页</em></td>
<td align="center"><em>键盘配件栏</em></td>
<td align="center"><em>智能体实时工作中</em></td>
<td align="center"><em>触控回答提示</em></td>
<td align="center"><em>配件栏 + 独立 Enter 按钮</em></td>
</tr>
</table>
@@ -142,23 +177,10 @@ Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.
<tr><td>在手机上手打密码</td><td><b>扫二维码 —— 即时认证</b></td></tr>
</table>
### 安全的二维码认证
在手机键盘上输密码太痛苦了。Codeman 用**密码学安全的一次性二维码令牌**取而代之 —— 扫描桌面上显示的二维码,手机即刻完成认证。
每个二维码编码的是一个包含 6 字符短码的 URL,该短码在服务端映射到一个 256 位密钥(`crypto.randomBytes(32)`)。令牌每 **60 秒**自动轮换,**首次扫描即原子性消费**(重放永远失败),并采用**基于哈希的 `Map.get()` 查找**,不会通过响应时延泄露任何信息。短码只是一个不透明指针 —— 真正的密钥永远不会出现在浏览器历史、`Referer` 头或 Cloudflare 边缘日志中。
该安全设计覆盖了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025,该研究发现 Top-100 网站中有 47 个存在漏洞)所指出的全部 6 个关键二维码认证缺陷:强制一次性使用、短 TTL、密码学随机性、服务端生成、扫描时桌面实时通知(QRLjacking 检测),以及 IP + User-Agent 会话绑定与手动吊销。双层速率限制(按 IP + 全局)使得在 62^6 = 568 亿种可能短码空间内进行暴力破解变得不可行。完整安全分析见:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
### 触控优化界面
- **键盘配件栏** —— 在虚拟键盘上方提供 `/init`、`/clear`、`/compact` 快捷按钮。破坏性命令(`/clear`、`/compact`)需双击确认 —— 第一次点击「上膛」,第二次点击执行 —— 这样在颠簸的通勤路上也不会误触
- **滑动导航** —— 在终端上左右滑动切换会话(阈值 80px,300ms)
- **智能键盘处理** —— 键盘弹出时工具栏与终端整体上移(使用 `visualViewport` API,并对 iOS 地址栏漂移设置 100px 阈值)
- **安全区适配** —— 通过 `env(safe-area-inset-*)` 适配 iPhone 刘海与底部 Home 指示条
- **44px 触控目标** —— 所有按钮均满足 iOS 人机界面指南的最小尺寸
- **底部抽屉式 case 选择器** —— 用上滑模态框替代桌面端下拉菜单
- **原生惯性滚动** —— `-webkit-overflow-scrolling: touch`,丝滑流畅
- **键盘配件栏** —— 在虚拟键盘上方提供 `/init`、`/clear`、`/compact` 快捷按钮;破坏性命令需双击确认,绝不误触
- **独立的 Enter 按钮** —— 以按键方式回放,先冲刷本地回显缓冲的文本,不会让内容滞留在屏幕上
- **滑动导航与智能键盘处理** —— 左右滑动切换会话;键盘弹出时工具栏与终端整体上移(`visualViewport` API)
- **为手机而生** —— 刘海与 Home 指示条的安全区适配、44px 触控目标、底部抽屉式 case 选择器、原生惯性滚动
```bash
codeman web --https
@@ -167,30 +189,86 @@ codeman web --https
> `localhost` 走纯 HTTP 即可。从其他设备访问时请使用 `--https`,或使用 [Tailscale](https://tailscale.com/)(推荐)—— 它提供私有网络,让你无需 TLS 证书即可从手机访问 `http://<tailscale-ip>:3000`。
### 安全的二维码认证
在手机键盘上输密码太痛苦了。Codeman 用**密码学安全的一次性二维码令牌**取而代之 —— 扫描桌面上显示的二维码,手机即刻完成认证。
每个二维码编码的是一个包含 6 字符短码的 URL,该短码在服务端映射到一个 256 位密钥(`crypto.randomBytes(32)`)。令牌每 **60 秒**自动轮换,**首次扫描即原子性消费**(重放永远失败),并采用**基于哈希的 `Map.get()` 查找**,不会通过响应时延泄露任何信息。短码只是一个不透明指针 —— 真正的密钥永远不会出现在浏览器历史、`Referer` 头或 Cloudflare 边缘日志中。
该安全设计覆盖了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025,该研究发现 Top-100 网站中有 47 个存在漏洞)所指出的全部 6 个关键二维码认证缺陷:强制一次性使用、短 TTL、密码学随机性、服务端生成、扫描时桌面实时通知(QRLjacking 检测),以及 IP + User-Agent 会话绑定与手动吊销。双层速率限制(按 IP + 全局)使得在 62^6 = 568 亿种可能短码空间内进行暴力破解变得不可行。完整安全分析见:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
---
## 实时智能体可视化
## 使用 Codeman —— 人类操作指南
实时观看后台智能体工作。Codeman 监控智能体活动,将每个智能体显示在一个可拖拽的浮动窗口中,并用「黑客帝国」风格的动态连接线连回父会话。
从头到尾走一遍如何在浏览器里驾驭 Codeman。如果你刚装好,就从这里开始。
<p align="center">
<img src="docs/images/subagent-spawn.png" alt="子智能体可视化" width="900">
</p>
### 1. 启动服务器
- **浮动终端窗口** —— 每个智能体一个可拖拽、可调整大小的面板,带实时活动日志,逐条展示每一次工具调用、文件读取与进度更新
- **连接线** —— 用动态绿色线条连接父会话与其子智能体,随智能体的产生与完成实时更新
- **状态与模型徽标** —— 绿色(活动)、黄色(空闲)、蓝色(已完成)指示,并以 Haiku/Sonnet/Opus 的颜色编码区分模型
- **自动行为** —— 窗口在产生时自动打开、完成时自动最小化,标签徽标显示「AGENT」或「AGENTS (n)」计数
- **嵌套智能体** —— 支持 3 层层级(主会话 → 团队成员智能体 → 子-子智能体)
```bash
codeman web # localhost:3000(仅环回 —— 安全默认值)
codeman web --port 8080 # 自定义端口(或设置 CODEMAN_PORT)
codeman web --https # 自签名 TLS(仅远程访问时需要)
codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_PASSWORD(见「安全」)
```
**智能体团队(Agent Teams)** —— 一等公民式支持 Claude Code 原生的多智能体团队(`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`)。`TeamWatcher` 轮询 `~/.claude/teams/`,将团队成员匹配到其主会话,并以实时子智能体窗口呈现,且具备**团队感知的空闲检测** —— 因此当团队成员仍在工作时,重生控制器不会被触发。详见 [`docs/agent-teams/`](docs/agent-teams/)。
打开打印出的 URL。整个页面是一个单一仪表盘;下面的一切都在这里完成。
### 2. 创建你的第一个会话
点击 **+ New Session**(或 **Quick Start**)。一个会话就是一个运行在自己 tmux 终端里的 AI CLI。你可以选择:
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
点击启动 —— Codeman 通过真实 PTY 拉起 CLI,并经 SSE 流式传输到你的浏览器。
### 3. 读懂仪表盘
- **标签(顶部)** —— 每个会话一个。`Alt+1`–`9` 跳转,`Ctrl+Tab` 下一个,拖拽排序(标签顺序会跨设备同步)。
- **终端(中央)** —— 真实的 `xterm.js` 终端;完整 TUI 正常渲染。直接输入并按 **Enter** 发送。`Shift+Enter` 插入换行。
- **侧边面板** —— Respawn、Orchestrator、Cron、Subagents、Settings(从工具栏切换)。
### 4. 与智能体对话
- **直接在终端输入提示** —— 即使跨越重连,输入也是精确一次送达(连接中断绝不会丢失或重复发送提示)。
- **粘贴或拖放图片**,直接进入会话。
- **语音输入** —— `Ctrl+Shift+V`(Deepgram Nova-3,自动静音停止)。
- **附件** —— 注册外部文件/文档,并内联预览 Office/PDF。
### 5. 让它自主运行
| 模式 | 用途 | 位置 |
| ---------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ |
| **Respawn** | 长时间无人值守运行 —— 空闲/限额时自动重启 CLI,带自适应时序。预设:`solo-work`、`overnight-autonomous` 等 | Respawn 标签页 |
| **Orchestrator** | 把一个目标变成分阶段计划,并跨多个智能体推动完成。 | 编排器面板 |
| **Cron** | 已保存的、命名的定时任务(`once`/`interval`/`daily`/`weekly`),到期时拉起会话并发送提示。 | ⏰ Cron 按钮(可选启用:App Settings → Display → Header Displays) |
| **Auto-resume** | 订阅限额重置后自动继续。 | Respawn 标签页(顶部) |
### 6. 随时随地访问
- **手机/平板** —— UI 完全触控优化;扫描桌面上的**二维码**即可免密码登录。
- **网络之外** —— `./scripts/tunnel.sh start` 打开一条 Cloudflare 隧道(先设置 `CODEMAN_PASSWORD`)。
- **SSH** —— `sc` 选择器可从终端附着任意会话(`sc` 交互式,`sc 2` 快速附着,`sc -l` 列表)。
### 7. 运维与维护
- **App Settings** —— 模型、effort、权限启动模式、主题/皮肤、通知、显示开关、各 CLI 的专属选项,以及跨设备同步的自定义显示名称和按设备保存的英文/简体中文界面语言。
- **自更新** —— git-clone 安装可在 **Settings → Updates** 中原地更新。
- **部署你自己的改动** —— 见[开发](#开发)。
> ⚠️ **安全提示:** 如果你正在 Codeman 受管会话*内部*工作(`echo $CODEMAN_MUX` → `1`),绝不要直接运行 `tmux kill-session` / `pkill claude` —— 请使用 Web UI 或 `./scripts/tmux-manager.sh`。
---
## 零延迟输入叠加层
<p align="center">
<img src="docs/images/zerolag-demo.gif" alt="Zerolag 演示 —— 本地回显与服务端回显并排对比" width="900">
<img src="docs/images/zerolag-demo-20260728.gif" alt="Zerolag 演示:两台手机并排对比,即时本地回显与 600ms-2.7s 服务端回显" width="900">
</p>
远程访问你的编程智能体时(VPN、Tailscale、SSH 隧道),每次按键通常需要 200–300 毫秒往返。Codeman 实现了一套**受 Mosh 启发的本地回显系统**,无论延迟多高,打字都感觉即时。
@@ -207,6 +285,30 @@ xterm.js 内部一个像素级精准的 DOM 叠加层以 0ms 渲染按键。后
---
## 实时智能体可视化
实时观看后台智能体工作。Codeman 监控智能体活动,将每个智能体显示在一个可拖拽的浮动窗口中,并用「黑客帝国」风格的动态连接线连回父会话。
<p align="center">
<img src="docs/images/subagent-windows-20260724.png" alt="子智能体可视化 —— 三个并行 Explore 智能体的浮动窗口与实时工具调用日志" width="900">
</p>
- **浮动终端窗口** —— 每个智能体一个可拖拽、可调整大小的面板,带实时活动日志,逐条展示每一次工具调用、文件读取与进度更新
- **连接线** —— 用动态绿色线条连接父会话与其子智能体,随智能体的产生与完成实时更新
- **状态与模型徽标** —— 绿色(活动)、黄色(空闲)、蓝色(已完成)指示,并以 Haiku/Sonnet/Opus 的颜色编码区分模型
- **自动行为** —— 窗口在产生时自动打开、完成时自动最小化,标签徽标显示「AGENT」或「AGENTS (n)」计数
- **嵌套智能体** —— 支持 3 层层级(主会话 → 团队成员智能体 → 子-子智能体)
多智能体 Workflow 运行(「ultracode」)同样可视化:一个浮动运行窗口实时跟踪整个工作流,展示阶段、各智能体的 token 用量与当前工具:
<p align="center">
<img src="docs/images/ultracode-window-20260724.png" alt="Ultracode 工作流可视化 —— 实时运行窗口,含各智能体 token 与阶段" width="900">
</p>
**智能体团队(Agent Teams)** —— 一等公民式支持 Claude Code 原生的多智能体团队(`CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`)。`TeamWatcher` 轮询 `~/.claude/teams/`,将团队成员匹配到其主会话,并以实时子智能体窗口呈现,且具备**团队感知的空闲检测** —— 因此当团队成员仍在工作时,重生控制器不会被触发。详见 [`docs/agent-teams/`](docs/agent-teams/)。
---
## 重生控制器(Respawn Controller)
自主工作的核心。当智能体进入空闲,重生控制器会检测到,发送继续提示,循环执行上下文管理命令以获得全新上下文,然后恢复工作 —— 可完全无人值守运行 **24 小时以上**。
@@ -216,7 +318,7 @@ WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE →
```
- **多层空闲检测** —— 完成消息、AI 驱动的空闲检查、输出静默、token 稳定性
- **用量限额自动恢复**(*可选,默认关闭*)—— 当 Claude 因订阅用量限额而停止("You've hit your limit · resets 3pm")时,Codeman 会解析重置时间,等到限额刷新(外加 2 分钟安全缓冲)后自动关闭限额对话框并发送 `continue`,让通宵任务平稳跨过 5 小时窗口而不是停摆到早晨。可识别 Claude Code 各版本的全部限额消息格式;若仍受限会自动重试;计划在 Codeman 重启后依然生效;暂停期间会阻止重生循环,避免 `/clear` 清掉等待中的对话。在会话 Respawn 标签页顶部按会话启用
- **用量限额自动恢复**(_可选,默认关闭_)—— 当 Claude 因订阅用量限额而停止("You've hit your limit · resets 3pm")时,Codeman 会解析重置时间,等到限额刷新(外加 2 分钟安全缓冲)后自动关闭限额对话框并发送 `continue`,让通宵任务平稳跨过 5 小时窗口而不是停摆到早晨。可识别 Claude Code 各版本的全部限额消息格式;若仍受限会自动重试;计划在 Codeman 重启后依然生效;暂停期间会阻止重生循环,避免 `/clear` 清掉等待中的对话。在会话 Respawn 标签页顶部按会话启用
- **熔断器** —— 当 Claude 卡住时防止重生抖动(CLOSED → HALF_OPEN → OPEN 状态,跟踪连续无进展与重复错误)
- **健康评分** —— 0–100 健康分,分项涵盖循环成功率、熔断器状态、迭代进展与卡死恢复
- **内置预设** —— `solo-work`(3s 空闲,60min)、`subagent-workflow`(45s,240min)、`team-lead`(90s,480min)、`ralph-todo`(8s,480min)、`overnight-autonomous`(10s,480min)
@@ -233,7 +335,7 @@ WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE →
- **崩溃安全** —— 完整状态持久化在 `state.json` 的 `orchestrator` 键下,可在重启后存续
- **可从 UI 或 API 驱动** —— 编排器面板,或 `POST /api/orchestrator/start` → `/approve` → `/status`(共 10 个端点)
> 与 Ralph(单会话自主循环)不同:编排器协调多阶段、多智能体执行。完整设计:[`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md)。
> 完整设计:[`docs/orchestrator-loop-architecture.md`](docs/orchestrator-loop-architecture.md)。
---
@@ -241,14 +343,18 @@ WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE →
运行 **20 个并行会话**且全程可见 —— 60fps 的实时 xterm.js 终端、按会话的 token 与成本跟踪、基于标签的导航,以及一键管理。
<p align="center">
<img src="docs/screenshots/multi-session-dashboard.png" alt="多会话仪表盘" width="800">
</p>
### 持久化会话
每个会话都运行在 **tmux** 内 —— 会话可在服务器重启、网络中断与机器休眠后存续。启动时自动恢复,具备双重冗余。幽灵会话发现机制能找到孤立的 tmux 会话。受管会话带有环境标签,因此智能体不会杀掉自己的会话。
### 会话管理器与命令面板
`Ctrl/Cmd/Alt+K` 打开模糊搜索的会话面板;**Browse all sessions** 打开会话管理器:一份去重后的完整清单,涵盖 Codeman 所知的一切(活动会话、来自状态与生命周期历史的既往会话,以及 Claude 转录),每一行都显示其第一条与最近一条提示。
- **置顶(Pin)**:把会话固定到列表顶部。被置顶的会话甚至能挺过被杀掉(降级为一条轻量的已停止记录,依然可见、可恢复)。
- **名称保留**:从会话管理器恢复既往会话时保留其原有名称,而不是生成一个新名称。
- **跨设备标签顺序**:拖拽排序的标签顺序保存在服务端,你的排列会从桌面跟随到手机。
### 主机名感知的窗口标题
在多台主机上运行 Codeman(笔记本、开发机、NAS)?浏览器标签标题是 `codeman:<主机名>`,让你无需点进去就能分辨每个标签对应哪个后端:
@@ -262,23 +368,15 @@ codeman web --title-hostname dev-box # codeman:dev-box(用于覆盖嘈
### 智能 Token 管理
| 阈值 | 动作 | 结果 |
|-----------|--------|--------|
| 阈值 | 动作 | 结果 |
| --------------- | --------------- | ---------------------- |
| **110k tokens** | 自动 `/compact` | 上下文被摘要,工作继续 |
| **140k tokens** | 自动 `/clear` | 以 `/init` 全新开始 |
| **140k tokens** | 自动 `/clear` | 以 `/init` 全新开始 |
### 通知
当会话需要关注时实时桌面提醒 —— `permission_prompt` 与 `elicitation_dialog` 触发关键的红色标签闪烁,`idle_prompt` 触发黄色闪烁。点击任意通知即可直接跳转到相关会话。Hook 按 case 目录自动配置。
### Ralph / Todo 跟踪
自动检测 Ralph 循环、`<promise>` 标签、TodoWrite 进度(`4/9 complete`)以及迭代计数器(`[5/50]`),并提供实时进度环与已用时间跟踪。
<p align="center">
<img src="docs/images/ralph-tracker-8tasks-44percent.png" alt="Ralph 循环跟踪" width="800">
</p>
### 运行摘要(Run Summary)
点击任意会话标签上的图表图标,即可看到所发生一切的时间线 —— 重生周期、token 里程碑、自动 compact 触发、空闲/工作切换、hook 事件、错误等等。
@@ -296,17 +394,70 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode** 或 **Codex**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*` 与 `CODEX_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
- **语音输入** —— 用 Deepgram Nova-3 口述提示(带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
- **图像输入** —— 直接把图片粘贴或拖放进会话
- **手势控制** *(可选)* —— 一个 MediaPipe 手部追踪叠加层,可徒手抓取/拖动会话窗口并捏合按钮。用 `CODEMAN_GESTURE=1` + App Settings → Display 启用
- **多显示器横跨** *(macOS)* —— 一键打开一个横跨所有显示器最大化的浏览器窗口,让浮动的智能体/手势面板可以跨越物理拼接缝
- **手势控制** _(可选)_ —— 一个 MediaPipe 手部追踪叠加层,可徒手抓取/拖动会话窗口并捏合按钮。用 `CODEMAN_GESTURE=1` + App Settings → Display 启用
- **多显示器横跨** _(macOS)_ —— 一键打开一个横跨所有显示器最大化的浏览器窗口,让浮动的智能体/手势面板可以跨越物理拼接缝
- **文件查看器按钮** _(可选)_ —— 头部新增一个按钮,一键切换内置文件浏览器面板;在 App Settings → Display → Header Displays 中启用
- **CJK / 输入法支持** —— 完整支持中文 / 日文 / 韩文的组合输入
- **操作系统通知与主机名感知标题** —— 桌面提醒与标签标题以 `codeman:<host>` 为前缀,使多主机配置不再含糊
---
## 隔离的 Docker 会话
让案例(case)运行在专属的加固 Docker 容器里,而不是直接跑在主机上:获得安全隔离、可复现的工具链和一键可移植性。
- **一键启动** —— 在 **New Case → Create New** 中勾选 **🐳 Run in an isolated Docker container**。Codeman 会创建案例文件夹、用默认设置启动容器,并在容器内启动智能体。无需填写任何主机/镜像/网络字段。
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
前置条件:只需 Docker(或 Podman)。智能体基础镜像会在首次使用时自动构建,构建进度实时显示在 UI 中(也可用 `node scripts/build-agent-image.mjs` 预构建)。完整指南:[`docs/docker-cases.md`](docs/docker-cases.md)。
---
## 远程 SSH 会话
把案例(case)指向另一台机器,通过 SSH 让智能体**在那台机器上**运行,同时保留同样的仪表盘、移动端 UI 与自主运行特性。你的笔记本只是一扇窗口,会话本体活在远程主机上。
- **天生持久**:智能体运行在远程主机上一个专用的 tmux 会话里,SSH 断连、网络切换或笔记本休眠都不会中断任务。重新连接后回到同一个活跃对话。
- **自动重连**:一个带上限退避的监视器发现 SSH 面板断开后,会静默重新附着到仍在运行的远程会话(设置中有总开关;主动杀掉的会话绝不会被复活)。
- **发现与附着**:列出主机上已在运行的 `codeman-*` 会话(由那台机器自己的 Codeman 或其他操作者启动)并附着其一。非你所有的已附着会话在关闭标签时**只分离,绝不杀掉**。
- **共享会话**:多个客户端可以以不同窗口尺寸同时附着同一个远程会话而互不挤压;发现列表会显示带客户端计数的「shared」徽标。
- **注入安全**:所有 ssh 命令行都经由单一的 shell 转义构建器生成,主机/路径/身份文件字段均有模式校验。
在 **New Case → Remote** 中配置(主机、用户、身份文件、可选跳板机)。完整设计:[`docs/remote-sessions.md`](docs/remote-sessions.md)。
---
## 多用户模式(可选启用)
与一个小型互信团队共享同一个 Codeman,每人拥有自己的登录与工作空间。**默认关闭**:不加该开关时,行为与单用户完全一致。
用 `codeman web --multiuser`(或 `CODEMAN_MULTIUSER=1`)启用。创建第一个管理员后,可通过 CLI 或 App Settings 中的 **Users** 标签页管理用户:
```bash
codeman users add alice --admin # 提示输入密码(或 --password-stdin)
codeman users add bob # 普通用户
codeman users list
```
- **按用户的空间**:每个用户的案例位于 `~/codeman-users/<name>/cases`;会话、案例、搜索与实时事件都按属主隔离。管理员可以看到全部。
- **可单独吊销的登录**:命名用户的密码以 scrypt 哈希保存在 `~/.codeman/users.json`;可随时禁用、重置(一次性密码)或删除账号。管理员操作审计记录在 `~/.codeman/admin-audit.jsonl`。
- **普通用户的更安全默认值**:非管理员以 `--permission-mode auto` 运行 Claude(Anthropic 的分类器护栏模式);raw shell 会话、cron `launchCommand` 与跳过权限模式需要按用户显式授权。
> ⚠️ **这只是工作空间的划分,不是用户之间的沙箱。** 所有会话都以同一个操作系统账户运行,因此有心用户的智能体依然能触及他人的文件。若需要真正的隔离,请结合 **Docker 案例**,或在不同的操作系统账户下运行独立实例。参见 [`docs/multi-user-plan.md`](docs/multi-user-plan.md) 与 [`docs/security-architecture.md`](docs/security-architecture.md) 的多用户章节。
---
## 远程访问 —— Cloudflare 隧道
使用免费的 [Cloudflare 快速隧道](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/),从手机或本地网络外的任意设备访问 Codeman —— 无需端口转发、无需 DNS、无需静态 IP。
@@ -374,14 +525,14 @@ loginctl enable-linger $USER
该设计参考了 ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin)(USENIX Security 2025),该研究发现 Top-100 网站中有 47 个因横跨 42 个 CVE 的 6 个关键设计缺陷而易受二维码认证攻击。Codeman 全部六个都做了应对:
| USENIX 缺陷 | 缓解措施 |
|-------------|------------|
| **缺陷 1**:缺少一次性强制 | 令牌首次扫描即原子性消费 —— 重放永远失败 |
| **缺陷 2**:长生命周期令牌 | 60s TTL + 90s 宽限,由定时器自动轮换 |
| **缺陷 3**:可预测的令牌生成 | `crypto.randomBytes(32)` —— 256 位熵。短码采用拒绝采样以消除取模偏差 |
| **缺陷 4**:客户端令牌生成 | 仅服务端 —— 令牌在嵌入二维码前绝不离开服务器 |
| **缺陷 5**:缺少状态通知 | 桌面提示:*「设备 [IP] 已通过二维码认证(Safari)。不是你?[吊销]」* —— 实时 QRLjacking 检测 |
| **缺陷 6**:会话绑定不足 | 存储 IP + User-Agent 以供审计。通过 API 手动吊销会话。HttpOnly + Secure + SameSite=lax cookie |
| USENIX 缺陷 | 缓解措施 |
| ---------------------------- | --------------------------------------------------------------------------------------------- |
| **缺陷 1**:缺少一次性强制 | 令牌首次扫描即原子性消费 —— 重放永远失败 |
| **缺陷 2**:长生命周期令牌 | 60s TTL + 90s 宽限,由定时器自动轮换 |
| **缺陷 3**:可预测的令牌生成 | `crypto.randomBytes(32)` —— 256 位熵。短码采用拒绝采样以消除取模偏差 |
| **缺陷 4**:客户端令牌生成 | 仅服务端 —— 令牌在嵌入二维码前绝不离开服务器 |
| **缺陷 5**:缺少状态通知 | 桌面提示:_「设备 [IP] 已通过二维码认证(Safari)。不是你?[吊销]」_ —— 实时 QRLjacking 检测 |
| **缺陷 6**:会话绑定不足 | 存储 IP + User-Agent 以供审计。通过 API 手动吊销会话。HttpOnly + Secure + SameSite=lax cookie |
#### 时序安全的查找
@@ -406,23 +557,23 @@ URL 被刻意保持精简(`/q/` 路径 + 6 字符码 ≈ 53–56 个字符)
#### 威胁覆盖
| 威胁 | 为何无效 |
|--------|-------------------|
| **二维码截图被分享** | 一次性:首次扫描即消费。60s TTL:攻击者动手前已过期。桌面通知会立即提醒你。 |
| **重放攻击** | 原子性一次性消费 + 60s TTL。旧 URL 始终返回 401。 |
| 威胁 | 为何无效 |
| ----------------------- | ------------------------------------------------------------------------------------ |
| **二维码截图被分享** | 一次性:首次扫描即消费。60s TTL:攻击者动手前已过期。桌面通知会立即提醒你。 |
| **重放攻击** | 原子性一次性消费 + 60s TTL。旧 URL 始终返回 401。 |
| **Cloudflare 边缘日志** | 短码是不透明的 6 字符查找键,而非真正的 256 位令牌。一次性意味着从日志重放永远失败。 |
| **暴力破解** | 568 亿种组合、任意时刻约 2 个有效、双层速率限制,早在统计可行性之前就已拦截。 |
| **QRLjacking** | 60s 轮换迫使实时转发。桌面提示提供即时检测。自托管单用户场景使钓鱼难以成立。 |
| **时序攻击** | 基于哈希的 Map 查找 —— 无字符串比较时序泄露。 |
| **会话 cookie 窃取** | HttpOnly + Secure + SameSite=lax + 24h TTL。可在 `POST /api/auth/revoke` 手动吊销。 |
| **暴力破解** | 568 亿种组合、任意时刻约 2 个有效、双层速率限制,早在统计可行性之前就已拦截。 |
| **QRLjacking** | 60s 轮换迫使实时转发。桌面提示提供即时检测。自托管单用户场景使钓鱼难以成立。 |
| **时序攻击** | 基于哈希的 Map 查找 —— 无字符串比较时序泄露。 |
| **会话 cookie 窃取** | HttpOnly + Secure + SameSite=lax + 24h TTL。可在 `POST /api/auth/revoke` 手动吊销。 |
#### 横向对比
| 平台 | 模型 | 对比 |
|----------|-------|------------|
| **Discord** | 长生命周期令牌、无确认、[屡被利用](https://owasp.org/www-community/attacks/Qrljacking) | Codeman:一次性 + TTL + 通知 |
| **WhatsApp Web** | 手机确认「关联设备?」,约 60s 轮换 | 轮换相当;WhatsApp 额外加了显式确认(对单用户而言是可接受的取舍) |
| **Signal** | 临时公钥、端到端加密信道 | 加密更强,但 [2025 年仍被俄罗斯国家级行为者](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger)通过社会工程攻破 |
| 平台 | 模型 | 对比 |
| ---------------- | -------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Discord** | 长生命周期令牌、无确认、[屡被利用](https://owasp.org/www-community/attacks/Qrljacking) | Codeman:一次性 + TTL + 通知 |
| **WhatsApp Web** | 手机确认「关联设备?」,约 60s 轮换 | 轮换相当;WhatsApp 额外加了显式确认(对单用户而言是可接受的取舍) |
| **Signal** | 临时公钥、端到端加密信道 | 加密更强,但 [2025 年仍被俄罗斯国家级行为者](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger)通过社会工程攻破 |
> 完整设计理由、安全分析与实现细节:[`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
@@ -430,13 +581,14 @@ URL 被刻意保持精简(`/q/` 路径 + 6 字符码 ≈ 53–56 个字符)
## 安全
Codeman 用 `--dangerously-skip-permissions` 启动会话,因此 Web UI 在设计上对任何能访问到它的人都是一个远程代码执行面 —— 整套安全模型的存在就是为了控制*谁*能访问。近期加固(v0.9.0 + v0.9.5)封堵了那些常困扰自托管开发工具的浏览器驱动攻击路径。完整模型:[`docs/security-architecture.md`](docs/security-architecture.md)。
Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI 在设计上对任何能访问到它的人都是一个远程代码执行面 —— 整套安全模型的存在就是为了控制*谁*能访问。(启动权限模式可配置,见下文。)近期加固(v0.9.0 + v0.9.5)封堵了那些常困扰自托管开发工具的浏览器驱动攻击路径。完整模型:[`docs/security-architecture.md`](docs/security-architecture.md)。**发现了漏洞?** 私下披露方式与已知限制清单见 [`SECURITY.md`](.github/SECURITY.md)。
### 网络与访问
- **默认仅环回** —— 绑定 `127.0.0.1`,仅可从本机访问,因此「无密码」默认配置开箱即安全。在未设置 `CODEMAN_PASSWORD` 的情况下绑定非环回主机会*启动但打印一条醒目警告*,并给出三个具体修复方案(设置密码、环回 + 一个带认证的隧道,或用 `--allow-unauthenticated-network` 显式确认)
- **可选认证,真实会话** —— 通过 `CODEMAN_USERNAME`(默认 `admin`)/ `CODEMAN_PASSWORD` 的 HTTP Basic 认证。成功后签发一个不透明的 256 位 `codeman_session` cookie(`randomBytes(32)`)—— 服务端校验,而非客户端签名,因此无法离线伪造(24h TTL、自动延长、设备上下文审计日志)
- **按 IP 速率限制** —— 失败 10 次 → `429` 并带 `Retry-After`(15 分钟衰减)。即便攻击者在同一 IP 上猛攻,有效 cookie 或正确密码也能*立即*恢复 —— 这很重要,因为所有隧道流量共享同一个环回 IP。二维码认证有自己独立的限制器
- **可配置的权限模式**:`--dangerously-skip-permissions` 只是默认值。**App Settings → Claude CLI → Startup Mode** 可以把新会话切换为 Anthropic 的分类器护栏 `auto` 模式(低打扰,需要 Claude Code 2.1.207+)、`normal` 提示模式,或一份显式的允许工具列表。多用户模式下,未获授权的用户会被强制为 `auto`,shell 会话与跳过权限需要按用户显式授权
### 始终开启的浏览器加固(v0.9.5)
@@ -450,7 +602,7 @@ Codeman 用 `--dangerously-skip-permissions` 启动会话,因此 Web UI 在设
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
@@ -481,73 +633,176 @@ sc -l # 列出会话
> Ctrl 绑定在 macOS 上也接受 Cmd。
| 快捷键 | 动作 |
|----------|--------|
| `Ctrl/Cmd+W` | 杀掉当前会话 |
| `Ctrl/Cmd+Tab` | 下一个会话 |
| `Alt+1`–`Alt+9` | 切换到第 N 个标签 |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
| `Ctrl/Cmd +` / `-` | 字体大小 |
| `Ctrl/Cmd+?` | 键盘帮助 |
| `Shift+Enter` | 插入换行(发送到终端) |
| `Escape` | 关闭面板与模态框 |
| 快捷键 | 动作 |
| ------------------------------- | -------------------------------------------------------- |
| `Ctrl/Cmd+W` | 杀掉当前会话 |
| `Ctrl/Cmd/Option+K` | 查找已打开的会话或新建一个 |
| `Ctrl/Cmd+Tab` | 下一个会话 |
| `Alt/Option+[` / `Alt/Option+]` | 上一个 / 下一个会话 |
| `Alt/Option+1`–`Alt/Option+9` | 切换到第 N 个标签(按物理键位,macOS Option 布局也适用) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+C` | 复制选中内容;未选中时中断代理 |
| `Ctrl+Shift+C` | 复制选中内容(永不中断) |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
| `Ctrl/Cmd +` / `-` | 字体大小 |
| `Ctrl/Cmd+?` | 键盘帮助 |
| `Shift+Enter` | 插入换行(发送到终端) |
| `Escape` | 关闭面板与模态框 |
---
## 从智能体驱动 Codeman —— 编程指南
面向不经浏览器控制 Codeman 的 AI 智能体与自动化:一个拉起工作会话的智能体、一个 CI 机器人,或是**运行在 Codeman 会话*内部*、编排其他会话的 Claude Code**。UI 能做的一切都是 HTTP + CLI,因此智能体也能做。
### 检测自己身处 Codeman 内部
当 CLI 运行在 Codeman 受管会话中时,以下环境变量会被设置 —— 读取它们,别硬编码任何东西:
| 变量 | 含义 |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| `CODEMAN_MUX=1` | 你在一个受管 tmux 会话里。**绝不要** `tmux kill-session` / `pkill claude` / `pkill tmux` —— 你会杀掉自己或兄弟会话。 |
| `CODEMAN_API_URL` | API 的基础 URL(例如 `https://127.0.0.1:3000`)。下面每个调用都用它。 |
| `CODEMAN_SESSION_ID` | *你自己的*会话 id。用它避免对自己下手。 |
| `CODEMAN_HOOK_SECRET_FILE` | hook 密钥文件的路径(受管隧道开启时调用 `/api/hook-event` 必需)。 |
### 行路规则(POST 之前先读)
1. **只发单行输入。** 编程输入会作为字面文本 **+ Enter** 一次性发送。多行字符串会破坏智能体 TUI(Ink)—— 发送一行,或拆成多次调用。
2. **让输入幂等。** 在 `POST …/input` 上带上稳定的 `clientId` 和按会话单调递增的 `seq`。服务端会去重,因此连接中断后的重试不会重复投递提示。
3. **认证。** 若设置了 `CODEMAN_PASSWORD`,发送 HTTP Basic 认证(用户 `admin` 或 `CODEMAN_USERNAME`)或 `codeman_session` cookie。默认的环回安装无密码。缺失的 `Origin` 头被允许,因此普通 `curl` 可用;跨站的浏览器 origin 会被拒绝(CSRF 防护)。
4. **响应信封。** 多数端点返回 `{ "success": true, "data": … }`(错误:`{ "success": false, "error", "errorCode" }`)。少数遗留 GET 返回裸响应体 —— **两种都要处理**(`body.data ?? body`)。
5. **`/api/v1/*`** 是 `/api/*` 的稳定别名。
### 常用配方
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (若设置了密码,给每个调用加上 -u admin:"$CODEMAN_PASSWORD")
# 1. 看看有什么在运行
curl -s "$API/api/sessions" | jq '.data // .'
# 2. 拉起一个工作会话(「case」= 命名工作目录)
curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 3. 向会话发送提示(精确一次:clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
# 4. 读回终端内容
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 5. 流式接收实时事件(会话输出、智能体活动、状态)
curl -sN "$API/api/events" # Server-Sent Events
# 6. 调度周期性工作(cron 风格任务)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
"promptMode":"inline_text","promptText":"Update dependencies and open a PR",
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. 查看后台子智能体及其活动记录
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. 全系统快照(会话、设置、重生、统计)
curl -s "$API/api/status" | jq
```
### 或使用内置 CLI
同样的操作也有命令形式(`codeman <cmd>`,括号内为别名)—— 在会话内的 shell 工具里很顺手:
```bash
codeman session start -d /path/to/repo # (s) 启动会话
codeman session list # 列出会话
codeman session logs <id> # 查看输出
codeman task add "fix the failing test" # (t) 排入任务
codeman attach <path> # 附着 Claude hook 上下文
```
### Hook(事件*回流*到 Codeman)
Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission_prompt`、`idle_prompt`、`stop`、`task_completed` 等),让仪表盘实时响应。该端点在环回上免认证,但在受管隧道下需要 `X-Codeman-Hook-Secret` 头(从 `$CODEMAN_HOOK_SECRET_FILE` 读取)。通常你不需要手动调用它 —— Codeman 会自动接好 —— 但自主层正是靠它「看见」智能体在做什么。
> 完整端点列表与请求/响应形状见下文。
---
## API
基于 Fastify 的 REST —— **15 个路由模块中约 140 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。以下是一个有代表性的子集:
基于 Fastify 的 REST —— **20 个路由模块中约 190 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。所有响应都使用 `ApiResponse<T>` 信封(`{success, data}` / `{success, error, errorCode}`);`/api/v1/*` 是稳定别名。以下是一个有代表性的子集:
### 会话(Sessions)
| 方法 | 端点 | 说明 |
|--------|----------|-------------|
| `GET` | `/api/sessions` | 列出全部 |
| `POST` | `/api/quick-start` | 创建 case 并启动会话 |
| `DELETE` | `/api/sessions/:id` | 删除会话 |
| `POST` | `/api/sessions/:id/input` | 发送输入 |
| 方法 | 端点 | 说明 |
| -------- | -------------------------- | ------------------------------------------------------------------------------ |
| `GET` | `/api/sessions` | 列出全部 |
| `POST` | `/api/quick-start` | 创建 case + 启动会话(`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | 发送输入(`{input, useMux?, clientId?, seq?}` —— `clientId`+`seq` = 精确一次) |
| `GET` | `/api/sessions/:id/output` | 读取终端输出 |
| `GET` | `/api/sessions/unified` | 统一的活动 + 历史清单(会话管理器):`?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | 在会话管理器中置顶 / 取消置顶(`{pinned}`) |
| `PUT` | `/api/session-order` | 跨设备同步标签顺序(`{order: [ids]}`) |
| `DELETE` | `/api/sessions/:id` | 删除会话 |
### 重生(Respawn)
| 方法 | 端点 | 说明 |
|--------|----------|-------------|
| `POST` | `/api/sessions/:id/respawn/enable` | 启用,带配置与定时器 |
| `POST` | `/api/sessions/:id/respawn/stop` | 停止控制器 |
| `PUT` | `/api/sessions/:id/respawn/config` | 更新配置 |
### Ralph / Todo
| 方法 | 端点 | 说明 |
|--------|----------|-------------|
| `GET` | `/api/sessions/:id/ralph-state` | 获取循环状态 + todos |
| `POST` | `/api/sessions/:id/ralph-config` | 配置跟踪 |
| 方法 | 端点 | 说明 |
| ------ | ---------------------------------- | -------------------- |
| `POST` | `/api/sessions/:id/respawn/enable` | 启用,带配置与定时器 |
| `POST` | `/api/sessions/:id/respawn/stop` | 停止控制器 |
| `PUT` | `/api/sessions/:id/respawn/config` | 更新配置 |
### 编排器(Orchestrator)
| 方法 | 端点 | 说明 |
|--------|----------|-------------|
| `POST` | `/api/orchestrator/start` | 从目标启动编排 |
| `POST` | `/api/orchestrator/approve` | 批准生成的计划 |
| `GET` | `/api/orchestrator/status` | 当前阶段 + 进度 |
| `POST` | `/api/orchestrator/stop` | 停止并清理 |
| 方法 | 端点 | 说明 |
| ------ | --------------------------- | --------------- |
| `POST` | `/api/orchestrator/start` | 从目标启动编排 |
| `POST` | `/api/orchestrator/approve` | 批准生成的计划 |
| `GET` | `/api/orchestrator/status` | 当前阶段 + 进度 |
| `POST` | `/api/orchestrator/stop` | 停止并清理 |
### Cron(定时任务)
| 方法 | 端点 | 说明 |
| ---------------- | ---------------------------- | --------------------- |
| `GET` / `POST` | `/api/cron/jobs` | 列出 / 创建 cron 任务 |
| `PUT` / `DELETE` | `/api/cron/jobs/:id` | 更新 / 删除任务 |
| `PUT` | `/api/cron/jobs/:id/enabled` | 启用 / 禁用 |
| `POST` | `/api/cron/jobs/:id/run` | 立即运行 |
| `GET` | `/api/cron/jobs/:id/runs` | 运行历史 |
### 子智能体(Subagents)
| 方法 | 端点 | 说明 |
|--------|----------|-------------|
| `GET` | `/api/subagents` | 列出所有后台智能体 |
| `GET` | `/api/subagents/:id` | 智能体信息与状态 |
| `GET` | `/api/subagents/:id/transcript` | 完整活动记录 |
| `DELETE` | `/api/subagents/:id` | 杀掉智能体进程 |
| 方法 | 端点 | 说明 |
| -------- | ------------------------------- | ------------------ |
| `GET` | `/api/subagents` | 列出所有后台智能体 |
| `GET` | `/api/subagents/:id` | 智能体信息与状态 |
| `GET` | `/api/subagents/:id/transcript` | 完整活动记录 |
| `DELETE` | `/api/subagents/:id` | 杀掉智能体进程 |
### 系统(System)
| 方法 | 端点 | 说明 |
|--------|----------|-------------|
| `GET` | `/api/events` | SSE 流 |
| `GET` | `/api/status` | 完整应用状态 |
| `POST` | `/api/hook-event` | Hook 回调 |
| `GET` | `/api/system/update/check` | 检查新发行版 |
| `POST` | `/api/system/update` | 自更新(git-clone 安装) |
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
| 方法 | 端点 | 说明 |
| ------ | ------------------------------- | ---------------------------------------- |
| `GET` | `/api/events` | SSE 流 |
| `GET` | `/api/status` | 完整应用状态 |
| `POST` | `/api/hook-event` | Hook 回调 |
| `GET` | `/api/system/update/check` | 检查新发行版 |
| `POST` | `/api/system/update` | 自更新(git-clone 安装) |
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
---
@@ -571,7 +826,6 @@ flowchart TB
end
subgraph Detection["检测层"]
RT["Ralph 跟踪器"]
SW["子智能体监视器<br/><small>~/.claude/projects/*/subagents</small>"]
TW["团队监视器<br/><small>~/.claude/teams/*</small>"]
end
@@ -582,7 +836,7 @@ flowchart TB
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
@@ -595,7 +849,6 @@ flowchart TB
SM --> RC
SM --> ORC
SM --> SS
S1 --> RT
S1 --> SCR
S2 --> SCR
RC --> SCR
@@ -614,7 +867,7 @@ flowchart TB
npm install
npx tsx src/index.ts web # 开发模式
npm run build # 生产构建
npm test # 运行测试
npm run test:ci # 运行测试(CI 套件;浏览器套件需要额外环境)
```
完整文档见 [CLAUDE.md](./CLAUDE.md)。
@@ -625,14 +878,14 @@ npm test # 运行测试
本代码库经历了一次全面的 7 阶段重构,消除了上帝对象、集中了配置,并建立了模块化架构:
| 阶段 | 改了什么 | 影响 |
|-------|-------------|--------|
| **性能** | 缓存端点、SSE 自适应批处理、缓冲区分块 | 终端延迟低于 16ms |
| **路由抽取** | `server.ts` 拆分为 15 个领域路由模块 + 认证中间件 + 端口接口 | server.ts 代码量 **−67%**(6,736 → 2,254) |
| **领域拆分** | `types.ts` → 16 个领域文件、`ralph-tracker` → 7 个文件、`respawn-controller` → 5 个文件、`session` → 6 个文件 | 不再有上帝文件 |
| **前端模块** | `app.js` → 18 个抽取模块,横跨基础设施、领域与特性层 | app.js 核心降至 **约 3.4K 行** |
| **配置合并** | 约 70 个散落的魔法数字 → 10 个领域聚焦的配置文件 | 零跨文件重复 |
| **测试基础设施** | 共享 mock 库、12 个路由测试文件、统一的 MockSession | 路由处理器可通过 `app.inject()` 测试 |
| 阶段 | 改了什么 | 影响 |
| ---------------- | ------------------------------------------------------------------------------------------------------------- | ------------------------------------------ |
| **性能** | 缓存端点、SSE 自适应批处理、缓冲区分块 | 终端延迟低于 16ms |
| **路由抽取** | `server.ts` 拆分为 15 个领域路由模块 + 认证中间件 + 端口接口 | server.ts 代码量 **−67%**(6,736 → 2,254) |
| **领域拆分** | `types.ts` → 16 个领域文件、`ralph-tracker` → 7 个文件、`respawn-controller` → 5 个文件、`session` → 6 个文件 | 不再有上帝文件 |
| **前端模块** | `app.js` → 18 个抽取模块,横跨基础设施、领域与特性层 | app.js 核心降至 **约 3.4K 行** |
| **配置合并** | 约 70 个散落的魔法数字 → 10 个领域聚焦的配置文件 | 零跨文件重复 |
| **测试基础设施** | 共享 mock 库、12 个路由测试文件、统一的 MockSession | 路由处理器可通过 `app.inject()` 测试 |
完整细节:[`docs/archive/code-structure-findings.md`](docs/archive/code-structure-findings.md)
@@ -654,6 +907,10 @@ npm install xterm-zerolag-input
---
## 版本策略
Codeman 遵循 [SemVer](https://semver.org/)。版本号真正承诺的内容,以及哪些算内部实现(HTTP/SSE API、磁盘上的状态、实验性特性),都写在 [`docs/versioning-policy.md`](docs/versioning-policy.md) 中。如果你的脚本依赖 HTTP API,请锁定到确切版本。
## 许可证
MIT —— 见 [LICENSE](LICENSE)
View File
+1
View File
@@ -24,6 +24,7 @@ export default defineConfig({
'test/inline-rename.test.ts', // browser (Playwright)
'test/opencode-resize.test.ts', // browser (Playwright)
'test/webgl-fallback.test.ts', // browser (Playwright)
'test/terminal-copy-shortcut.test.ts', // browser (Playwright)
],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
+23 -2
View File
@@ -26,8 +26,8 @@ RUN apt-get update \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The agent CLIs (all four backends Codeman supports). Pinning is left to the
# rebuild cadence (see docs/docker-cases-plan.md, user-decision 2).
# The npm-published agent CLIs. Pinning is left to the rebuild cadence (see
# docs/docker-cases-plan.md, user-decision 2).
RUN npm install -g \
@anthropic-ai/claude-code \
@openai/codex \
@@ -35,14 +35,35 @@ RUN npm install -g \
opencode-ai \
&& npm cache clean --force
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
# the installer's default target is `$HOME/.local/bin`, which at build time is
# root's home and would be unreachable by the `agent` user the container runs as.
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
&& chmod 755 /usr/local/bin/agy \
&& agy --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
# only matters for a hand-run / Docker Desktop container. gid 0 + group-writable
# HOME (OpenShift arbitrary-uid convention) keeps $HOME writable for any uid.
# UTF-8 locale so tmux/Ink render Unicode box-drawing instead of VT100 ACS `q`
# glyphs (C.UTF-8 is built into glibc; no locales package needed). Codeman also
# sets these at run time so containers built before this line still get UTF-8.
ENV LANG=C.UTF-8 LC_ALL=C.UTF-8
ENV HOME=/home/agent
# `.claude` (+ `.claude/projects` mount point) and `.codex` (+ `.codex/sessions`) are
# pre-created gid-0 group-writable so the container owns its OWN credential config
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
View File
+703
View File
@@ -0,0 +1,703 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 6 IMPLEMENTED, uncommitted as of 2026-08-09. Steps 1 to 5 were
multi-round verified on 2026-08-08; step 6 (CLI install command + per-case injection +
`agentSkillEnabled`) was built 2026-08-09; see the step-6 entry at the end of
[§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and what is still open.
**Date**: 2026-08-08
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
---
## 0. Where this came from: what herdr does
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
multiplexer built around AI coding agents. Relevant findings from the research pass:
| Capability | How herdr does it |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
The commands the skill teaches the agent:
| Group | Commands |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| workspace | `workspace list`, `workspace create` |
| tab | `tab list --workspace <id>`, `tab create` |
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
### The honest comparison
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
What herdr genuinely does better is being **callable by the agent running inside it**. For
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
---
## 1. Gap analysis
| herdr capability | Codeman equivalent today | Gap |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
| `pane wait-output --match` | nothing | **missing** |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
the two real gaps.
---
## 2. Part 1: the Codeman agent skill
### 2.1 Goal
An agent running inside a Codeman session can discover and correctly drive Codeman without the
user pasting API docs into the prompt, and without inventing dangerous calls.
### 2.2 Layout and distribution
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
```
skills/
codeman/
SKILL.md <- single source of truth
reference/
endpoints.md <- full endpoint tables, loaded on demand
recipes.md <- worked multi-session orchestration examples
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
```
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
point of shipping a skill.
Install paths, in order of how a user gets it:
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
file. This is the path for users who installed via npm and never cloned the repo.
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
into context on every turn, so an always-on skill has a small permanent token cost, and we
should measure that we are buying something with it first.
### 2.3 SKILL.md content
Frontmatter, per the skills convention (`name` + `description` required):
```yaml
---
name: codeman
description: >-
Control Codeman, the session manager this agent is running inside: list sessions,
start worker sessions, send prompts, read terminal output, and wait for other agents
to finish. Only usable when CODEMAN_MUX=1.
---
```
Body sections, in order:
**1. Guard (first thing, non-negotiable).**
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
SELF="${CODEMAN_SESSION_ID:-}"
```
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
cannot identify is not one it should be driving.
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
- Single-line input only. Multi-line breaks the agent TUI (Ink).
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
- Prefer `/api/v1/*`, the stable alias.
**3. Safety rules (the section that does not exist anywhere today).**
- Never act on `$CODEMAN_SESSION_ID`. That is you.
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
**4. Recipes**, each one a single copy-pasteable curl:
| Task | Call |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| list sessions | `GET /api/v1/sessions` |
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
| read output | `GET /api/v1/sessions/:id/output` |
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
| watch sub-agents | `GET /api/v1/subagents` |
| schedule work | `POST /api/v1/cron/jobs` |
| clean up | `DELETE /api/v1/sessions/:id` |
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
part of the skill stays small.
### 2.4 An ergonomics guard worth adding server-side
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
equals the target id, with a clear error.
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
### 2.5 Verification
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
to "start a worker session that runs the test suite and tell me when it finishes".
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
lead agent waited rather than polling in a busy loop.
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
Never run this against `w1`/`w2`/`w3`.
### 2.6 Files touched
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
- `.claude/skills/codeman` symlink (new)
- `src/cli.ts` (new `skill install` subcommand)
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
from the raw body, per the partial-PUT invariant)
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
- `package.json` `files` array, so `skills/` ships to npm
- README pointer, `docs/extending-codeman.md` seam 3 pointer
---
## 3. Part 2: wait primitives
### 3.1 Goal
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
which a curl-driven agent cannot practically consume: it would have to hold a streaming
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
solve it with bounded long-poll endpoints.
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
new optional fields are non-breaking).
### 3.2 The signal model
A waiter resolves on the first of a set of signals. Sources that already exist:
| Signal | Source today |
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
waits forever on `stop` in a codex session.
### 3.3 Endpoint specs
#### A. `GET /api/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
Response (always 200 unless the session is missing or a cap is hit):
```json
{
"success": true,
"data": {
"signal": "stop",
"timedOut": false,
"immediate": false,
"ended": false,
"waitedMs": 8421,
"status": "idle",
"sessionId": "...",
"until": ["stop", "idle", "exit"],
"limitPaused": false
}
}
```
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
timeout was expected rather than a stall worth retrying hard.
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
can loop without treating every poll boundary as a failure. Errors are reserved for
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
`immediate: true` means the session was already in the requested state and `fresh` was not set.
#### B. `GET /api/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
generic one like `BUILD OK`. The skill's recipes must show that.
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
readability only; matching runs on the raw stripped text. Without it, a real pane's
`\r\n` padding between the prompt and the match fills the whole context window with
nothing, which was the first thing the live test showed.
#### C. `wait` on the existing input endpoint
`POST /api/sessions/:id/input` gains two optional fields:
```json
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
```
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
when the input contains a carriage return.)
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
"input delivered" and "session flips to working" there is a window where a naive
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
ships `agent prompt --wait` as its own thing.
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
Two behaviors to preserve carefully:
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
purpose (a tmux child process must not block the HTTP response). With `wait` present the
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
observable for the first time. The non-wait path must keep its current fire-and-forget shape
byte for byte.
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
original turn may be long over, and requiring a new transition would block a redelivery until
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
current state.
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
the waiter is registered. If registration then fails on a full pool, the handler must call
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
duplicate and the input is lost by the very mechanism reliable delivery exists for.
### 3.4 Module design
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
(same split as `self-update.ts`):
```ts
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
notifySignal(sessionId, signal: WaitSignal): void
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
notifyOutput(sessionId, chunk: string): void
cancelAll(sessionId, reason): void
```
Wiring points, all existing:
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
callbacks so a listener could be added lazily; that was deleted once it was clear no
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
the no-waiter check comes before the ANSI strip.
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
live-testing the delete path, not by the unit tests.
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
| Constant | Default | Why |
| ------------------------- | ------- | --------------------------------------- |
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
| `MAX_WAITERS_PER_SESSION` | 16 | |
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
### 3.5 Transport concerns
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
The skill's recipes must show the loop.
### 3.6 Edge cases to get right
| Case | Behavior |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
### 3.7 Tests
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
chunk-straddling output match, case-insensitive match.
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
path is unchanged (still returns before `writeViaMux` settles).
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
### 3.8 Files touched
- `src/config/agent-wait.ts` (new)
- `src/web/session-wait-registry.ts` (new)
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
generated client must send `undefined`, never `null`)
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
---
## 4. Deferred: parts 3 to 5
Not in scope now, kept here so they are not lost.
### Part 3: promote `blocked` to a first-class state
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
re-deriving it. herdr makes `blocked` a real state that rolls up.
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
mobile overview, the wait endpoints, and any external agent read one field.
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
the HTTP contract.
### Part 4: `GET /api/schema`
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
### Part 5: detection manifests instead of hardcoded patterns
CLI-specific readiness, blocked and usage-limit patterns live in code across
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
Bundled manifests plus local override only, no network.
### Explicit non-goals
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
runtime means third-party code inside a process that spawns agents with your credentials, on a
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
delegates to tmux, so PTYs already survive a self-update restart.
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
would double the surface for no capability gain.
---
## 5. Sequencing
| Step | Work | Gate |
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 | Docs: api-reference, extending-codeman, README | |
| 8 | COM (minor bump: new endpoints, new setting, new optional fields) | both CI and Release workflows green |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary? (Still open, not built with step 6.)
5. Regex support in `wait-output`: confirm literal-only for v1.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
### What shipped
| Piece | Files |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
### Bugs found in ADJACENT code, not in the new feature
These are the highest-value output of the exercise and none were on the plan:
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
without the flag, success with it, and the failure swallowed by the hook's own
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
have left every one of them broken).
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
only if the payload has a carriage return; without it the text sits in the composer
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
docs' own examples.
### Design decisions worth not re-litigating
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
waits because tunnels cut idle connections, and every poll boundary would otherwise be
indistinguishable from failure.
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
turn as this one. The waiter is registered before the write.
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
`--regex` because Rust's regex crate is linear-time.
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
`close`).
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
### Verification rounds
Six agents across three rounds, each verifying the previous round's work rather than its
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
success without running its task. Two traps recurred often enough to name:
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
in `afterEach` silently killed the registry for every later test in a file; three test
files sharing one session id against the process-wide registry let one file's leftover
waiter fail another's assertion. Any new wait test needs care on all three.
- **HTTP-only test instances.** Every isolated instance used during the build was plain
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
user actually runs.
### Resolved at wrap-up (2026-08-08, conclusion pass)
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
the skill** rather than patched. Signals are edge-triggered with no history, so a
`stop` that fires before its waiter registers is unobservable afterwards; a
`fresh=0` gather was rejected because the only `until` set that current state can
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
shell flows reliable; the limitation is recorded in
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
dialog emits no further `idle`; the false success is the startup transition).
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
documented; every send-and-wait retry loop now treats `duplicate:true` +
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
floor named; the auth fallback now also reads the supervisor definition
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
lines; `pid != null` is documented as startup-only, never liveness.
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
### Still open
- **Release checklist**: `package.json` `files` includes `skills`, which is still
untracked. `git add skills/` must be part of the release commit, or npm publishes
a tarball without the skill (a `files` entry that does not exist is silently
ignored, so nothing fails). `test/agent-skill.test.ts` reads the packaged source,
so CI at least fails loudly if the directory goes missing from a checkout.
- The 1.13.0 changeset is written under `.changeset/`; consuming it (COM flow),
the release commit, and the deploy remain.
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`) and converting
timeout-shaped test detections into fast assertions.
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
### Step 6 (2026-08-09): install command, per-case injection, the setting
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
| Piece | Where |
| ----- | ----- |
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
+341
View File
@@ -46,6 +46,20 @@ payload return `{ "success": true, "data": {} }`.
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
@@ -66,6 +80,333 @@ the HTTP status.
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
```
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
#### What the matcher actually sees
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
picture; the matcher sees the stream that painted it. For linear output the two
agree once escapes are stripped, but a full-screen TUI composes its picture with
cursor positioning, so what the pane shows and what the stream carries can differ.
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
with cursor moves rather than printing spaces, so screen text can reach the
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
this folder` matched, `Quick safety check` did not), so a multi-word `match`
against a TUI pane is unreliable rather than impossible. Match a **single
space-free token**, ideally one you printed yourself. Plain command output (a
shell worker, an `echo`) keeps its spaces.
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
it.** It is cut from the same normalized stream the match ran against, then
cleaned for display: remaining raw control bytes are removed (an agent pipes the
snippet into its own terminal, so a worker's bytes must not be able to reset that
display) and blank runs are collapsed. A printable needle that matched will appear
in it; a needle containing control bytes or a blank run may not survive verbatim.
```bash
MARK="DONE_$RANDOM"
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
```
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
any of them:
```json
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
2. `wait.timedOut`: a poll boundary. Loop again.
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
File diff suppressed because one or more lines are too long
+106 -26
View File
@@ -2,14 +2,18 @@
> Official documentation for Claude Code hooks system, extracted from [code.claude.com](https://code.claude.com/docs/en/hooks).
**Last Updated**: 2026-01-24
**Last Updated**: 2026-07-25
**Source**: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)
> This is a maintained summary, not an exhaustive copy of the upstream reference.
> Check the source link for event-specific schemas before adding a new hook.
---
## Overview
Hooks are automated scripts that execute at specific events during your Claude Code session. They allow you to:
- Validate, modify, or block tool usage
- Add context to prompts
- Implement custom workflows
@@ -21,12 +25,12 @@ Hooks are automated scripts that execute at specific events during your Claude C
Hooks are configured in settings files:
| File | Scope |
|------|-------|
| `~/.claude/settings.json` | User (global) |
| `.claude/settings.json` | Project |
| File | Scope |
| ----------------------------- | -------------------------- |
| `~/.claude/settings.json` | User (global) |
| `.claude/settings.json` | Project |
| `.claude/settings.local.json` | Local project (gitignored) |
| Plugin hook files | Plugin-specific |
| Plugin hook files | Plugin-specific |
### Basic Structure
@@ -49,8 +53,9 @@ Hooks are configured in settings files:
```
**Key Fields**:
- `matcher`: Pattern to match tool names (case-sensitive, supports regex like `Edit|Write` or `*` for all)
- `type`: `"command"` for bash or `"prompt"` for LLM-based evaluation
- `type`: `"command"`, `"http"`, `"mcp_tool"`, `"prompt"`, or `"agent"` where the event supports it
- `command`: Bash command to execute
- `prompt`: LLM prompt for evaluation (prompt-based hooks only)
- `timeout`: Optional timeout in seconds (default: 60)
@@ -59,6 +64,10 @@ Hooks are configured in settings files:
## Hook Events
Claude Code's current event surface is broader than the detailed subset below. In
particular, `TeammateIdle` and `TaskCompleted` are supported lifecycle events used
by Codeman; they are not stale or plugin-defined event names.
### PreToolUse
**When**: After Claude creates tool parameters, before processing the tool call.
@@ -66,15 +75,17 @@ Hooks are configured in settings files:
**Use Cases**: Approval, denial, or modification of tool calls.
**Common Matchers**:
- `Bash` - Shell commands
- `Write` - File writing
- `Edit` - File editing
- `Read` - File reading
- `Task` - Subagent tasks
- `Agent` - Subagent tasks
- `WebFetch`, `WebSearch` - Web operations
- `mcp__<server>__<tool>` - MCP tools
**Output Control**:
```json
{
"hookSpecificOutput": {
@@ -96,13 +107,14 @@ Hooks are configured in settings files:
**Use Cases**: Auto-approve or deny permissions.
**Output Control**:
```json
{
"hookSpecificOutput": {
"hookEventName": "PermissionRequest",
"decision": {
"behavior": "allow|deny",
"updatedInput": { },
"updatedInput": {},
"message": "deny reason",
"interrupt": false
}
@@ -117,6 +129,7 @@ Hooks are configured in settings files:
**Use Cases**: Provide feedback, run formatters/linters, log operations.
**Output Control**:
```json
{
"decision": "block",
@@ -128,15 +141,38 @@ Hooks are configured in settings files:
}
```
#### Asynchronous Rewake
Command hooks can set `"asyncRewake": true` to run asynchronously and wake an
idle Claude turn when the hook exits with code 2. The hook's stderr is delivered
to Claude as a system reminder. This implies `"async": true`; ordinary async
hooks do not wake an idle turn, and their output waits for the next interaction.
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
the background task ID from the Bash result, watches the originating transcript
and, for subagents, the top-level parent transcript for the matching completion
notification, and exits 2. Claude records a subagent's Bash result in its
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
The task ID keeps each wake targeted. The helper does not send terminal input,
so it cannot submit a user's partially written prompt.
For script-dispatched Codex work, `codex-run.sh` writes the final response
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
subagent discovery and dispatcher result delivery are separate contracts.
### Notification
**When**: When Claude Code sends notifications.
**Matchers**:
- `permission_prompt`
- `idle_prompt`
- `auth_success`
- `elicitation_dialog`
- `elicitation_complete`
- `elicitation_response`
### UserPromptSubmit
@@ -145,6 +181,7 @@ Hooks are configured in settings files:
**Use Cases**: Add context, validate, or block prompts.
**Output Control**:
```json
{
"decision": "block",
@@ -165,6 +202,7 @@ Hooks are configured in settings files:
**Use Cases**: **Ralph Wiggum loops** - block exit and refeed prompt.
**Output Control**:
```json
{
"decision": "block",
@@ -173,6 +211,7 @@ Hooks are configured in settings files:
```
Or to allow exit:
```json
{
"continue": true,
@@ -184,15 +223,42 @@ Or to allow exit:
### SubagentStop
**When**: When a subagent (Task tool call) finishes responding.
**When**: When a subagent (Agent tool call) finishes responding.
**Use Cases**: Control nested loops, verify subagent output.
The hook input includes `agent_id`, `agent_transcript_path`, and
`last_assistant_message`. Like `Stop`, a command hook can return
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
the reason back to it.
Codeman uses this to prevent premature reports from workers that still own live
Monitor or background-Bash processes. It derives candidate task IDs from the
subagent transcript, but requires a matching live Linux process descriptor for
`tasks/<id>.output`; historical task text by itself is not treated as active.
### TeammateIdle
**When**: When an agent-team teammate is about to go idle.
**Use Cases**: Reassign work, continue a teammate loop, or notify an orchestrator.
**Matcher Support**: None. The hook fires for every occurrence.
### TaskCompleted
**When**: When a task is about to be marked completed.
**Use Cases**: Validate completion or forward team progress to an external UI.
**Matcher Support**: None. The hook fires for every occurrence.
### PreCompact
**When**: Before a compact operation.
**Matchers**:
- `manual` - Invoked from `/compact`
- `auto` - Invoked from auto-compact
@@ -201,6 +267,7 @@ Or to allow exit:
**When**: When Claude Code starts or resumes a session.
**Matchers**:
- `startup` - Fresh start
- `resume` - From `--resume`, `--continue`, or `/resume`
- `clear` - From `/clear`
@@ -209,6 +276,7 @@ Or to allow exit:
**Use Cases**: Load development context, set environment variables.
**Persisting Environment Variables**:
```bash
#!/bin/bash
if [ -n "$CLAUDE_ENV_FILE" ]; then
@@ -219,6 +287,7 @@ exit 0
```
**Output Control**:
```json
{
"hookSpecificOutput": {
@@ -233,6 +302,7 @@ exit 0
**When**: When a session ends.
**Reason Values**:
- `clear`
- `logout`
- `prompt_input_exit`
@@ -254,7 +324,7 @@ Hooks receive JSON via stdin with common fields:
"permission_mode": "default",
"hook_event_name": "PreToolUse",
"tool_name": "Bash",
"tool_input": { },
"tool_input": {},
"tool_use_id": "toolu_01ABC123..."
}
```
@@ -262,6 +332,7 @@ Hooks receive JSON via stdin with common fields:
### Tool-Specific Input
**Bash**:
```json
{
"tool_name": "Bash",
@@ -274,6 +345,7 @@ Hooks receive JSON via stdin with common fields:
```
**Write**:
```json
{
"tool_name": "Write",
@@ -285,6 +357,7 @@ Hooks receive JSON via stdin with common fields:
```
**Edit**:
```json
{
"tool_name": "Edit",
@@ -302,11 +375,11 @@ Hooks receive JSON via stdin with common fields:
### Exit Codes
| Code | Behavior |
|------|----------|
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
| Code | Behavior |
| ----- | --------------------------------------------------------------------- |
| 0 | Success. `stdout` processed (shown in verbose or added as context) |
| 2 | Blocking error. Only `stderr` used. Blocks tool/prompt based on event |
| Other | Non-blocking error. `stderr` shown in verbose, execution continues |
### JSON Output (Exit Code 0)
@@ -323,7 +396,12 @@ Hooks receive JSON via stdin with common fields:
## Prompt-Based Hooks
For Stop and SubagentStop events, you can use LLM-based evaluation:
Prompt and agent handlers are supported by decision-oriented events including
`PreToolUse`, `PermissionRequest`, `PostToolUse`, `PostToolUseFailure`,
`PostToolBatch`, `UserPromptSubmit`, `Stop`, `SubagentStop`, `TaskCreated`, and
`TaskCompleted`. Check the upstream reference before choosing a handler type.
For example, a Stop event can use LLM-based evaluation:
```json
{
@@ -344,6 +422,7 @@ For Stop and SubagentStop events, you can use LLM-based evaluation:
```
**LLM Response Format**:
```json
{
"ok": true,
@@ -362,17 +441,18 @@ Hooks can be defined in Skills, Agents, and Slash Commands using frontmatter:
name: secure-operations
hooks:
PreToolUse:
- matcher: "Bash"
- matcher: 'Bash'
hooks:
- type: command
command: "./scripts/security-check.sh"
command: './scripts/security-check.sh'
---
```
These hooks:
- Are scoped to the component's lifecycle
- Only run when that component is active
- Support: PreToolUse, PostToolUse, Stop
- Support all hook events; a subagent-scoped `Stop` is converted to `SubagentStop`
---
@@ -550,11 +630,11 @@ exit 0
## Environment Variables
| Variable | Description |
|----------|-------------|
| `CLAUDE_PROJECT_DIR` | Project root directory |
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
| Variable | Description |
| -------------------- | ------------------------------------------------ |
| `CLAUDE_PROJECT_DIR` | Project root directory |
| `CLAUDE_CODE_REMOTE` | `"true"` for web, empty for CLI |
| `CLAUDE_ENV_FILE` | Path to write persistent env vars (SessionStart) |
---
@@ -593,4 +673,4 @@ Use `/hooks` command to view registered hooks and make changes.
---
*Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)*
_Source: [Claude Code Hooks Documentation](https://code.claude.com/docs/en/hooks)_
+2 -2
View File
@@ -44,9 +44,9 @@ records), kept distinct from the existing `ScheduledRun`.
## 2. Where agent/session types are defined
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity'`
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode}-cli-resolver.ts`.
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
## 3. Where input is sent into a session
+2 -2
View File
@@ -1,7 +1,7 @@
# Cron Jobs — User & Operator Guide
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini) session on a schedule and
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
+18 -2
View File
@@ -2,7 +2,7 @@
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` all work inside the container.
## One-time setup: build the base image
@@ -15,6 +15,21 @@ node scripts/build-agent-image.mjs # builds codeman/agent:base
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
```bash
node scripts/build-agent-image.mjs --no-cache
```
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
@@ -52,8 +67,9 @@ curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"cl
## Lifecycle
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
- **Container stop / host reboot** recreates the container and, when a resume id was captured, **resumes** the last conversation from the bind-mounted transcript.
- **Container stop / host reboot** restarts the container and **resumes** the last conversation from the bind-mounted transcript. Claude sessions launch with a pinned conversation id (`--session-id <sessionId>`, with a `--resume` fallback when the transcript already exists), and the case remembers its last conversation (`lastClaudeSessionId`), so a relaunch after the container was stopped, rebooted, or recreated continues where it left off.
- **Killing one session** only kills that session's in-container tmux session; the shared container stays up for sibling sessions.
- **Editing the docker host config** (image, memory, network, ...) is detected on the next launch: the desired config hash is compared against the container's `codeman.confighash` label, and a mismatch refuses the launch with a "config changed, recreate?" confirm. Confirming calls `POST /api/docker-cases/:name/recreate` (refused while sessions of the case are live), which removes the container so the next launch recreates it with the new config; the workspace and the conversation survive.
- **Deleting the case** `docker rm -f`s the container (the bind-mounted workspace on the host survives). An instance-scoped boot reaper removes containers whose case is gone.
## Isolation & security
+417
View File
@@ -0,0 +1,417 @@
# Extending Codeman
Codeman has no plugin runtime, and that is a deliberate choice rather than a
missing feature. A plugin runtime means running third-party code inside a process
that spawns agents with your credentials, on a server people routinely expose
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
exist, so it does not hand that away for an extension mechanism.
Instead there are four seams that already work, from any language, with nothing
installed:
| You want to | Use | Runs where |
| --- | --- | --- |
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
Everything below is covered by the stability promise in
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
envelope, `errorCode` values, and SSE event names are stable. Additive changes
(new endpoints, new optional fields, new events) are non-breaking. Breaking
changes ship under a new prefix (`/api/v2`).
## Before you start
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
authenticate once and keep the `codeman_session` cookie. With no password set,
Codeman is loopback-only and unauthenticated.
```bash
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
```
**Envelope.** Every response is `{"success": true, "data": ...}` or
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
is in [`api-reference.md`](api-reference.md).
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
the payload at the top level rather than under `data`. Read defensively with
`body.data ?? body`.
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
the status code before you parse, or a missing password looks like a broken endpoint.
**Already driving Codeman from an agent?** The README's
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
running inside Codeman find the API and avoid acting on itself. This page is for
code running *outside* a session.
## Seam 1: Web tabs
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
your agent sessions. You write a normal web page; Codeman handles embedding it.
```bash
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
-H 'Content-Type: application/json' \
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
```
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
framing check), `POST /api/v1/webviews/:id/open`.
### Why it is proxied
By default your page is served through Codeman's own origin at `/webview/:cap/*`
rather than framed directly. A direct iframe fails three ways at once: production
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
cross-origin frames. Proxying solves all three without weakening the CSP.
### The two things that will confuse you
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
`trusted: true`. Two consequences look like bugs in your own app:
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
patches the common DOM sinks, but if you construct URLs in an unusual way,
prefer relative paths.
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
itself renders fine, this is the area to look at.
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
the API that spawns agents. Only mark your own trusted code.
## Seam 2: SSE events
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
`event: <name>` plus `data: <json>`. There are 149 event names following a
`domain:action` convention, registered in `src/web/sse-events.ts`.
The ones most integrations want:
| Event | Meaning |
| --- | --- |
| `session:created`, `session:deleted` | A session appeared or went away |
| `session:idle` | The agent stopped working |
| `session:completion` | A completion message was detected |
| `session:exit`, `session:error` | The session ended or failed |
| `hook:permission_prompt` | The agent is asking for permission |
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
| `hook:task_completed`, `task:completed` | Work finished |
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
| `mux:died` | A multiplexer session died unexpectedly |
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
### Filtering
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
sessions you did not list. Lifecycle and metadata events are always delivered, so
you cannot accidentally filter away the thing you are listening for.
Pass `?clientId=<uuid>` to enable live filter updates through
`POST /api/v1/events/subscribe` without reconnecting the stream.
### Example: notify when any agent needs you
```js
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buf = '';
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buf += decoder.decode(value, { stream: true });
const frames = buf.split('\n\n');
buf = frames.pop() ?? '';
for (const frame of frames) {
const name = frame.match(/^event: (.+)$/m)?.[1];
const data = frame.match(/^data: (.+)$/m)?.[1];
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
}
}
```
## Seam 3: HTTP API and CLI
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
If the caller is an agent running _inside_ a Codeman session, install the packaged
agent skill instead of teaching it these calls by hand: `skills/codeman` in the repo
(`npx skills add Ark0N/Codeman --skill codeman -g`, or `codeman skill install
[--case <name>]`, or the synced `agentSkillEnabled` App Setting for automatic
per-case injection on Claude session create). The skill carries the guard, the
safety rules, and verified wait/orchestration recipes.
The common ones:
```bash
# List sessions (live + persisted + transcript history, deduped)
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
# Create a session
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
-H 'Content-Type: application/json' \
-d '{"workingDir":"/home/me/project","mode":"claude"}'
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
# when the input contains a carriage return; without it the text sits on the
# session's prompt unsubmitted)
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true}'
```
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
`seq` (monotonic per session). Send both and the server applies each pair
at-most-once, so retrying after a dropped connection cannot type the prompt
twice. Omit them entirely rather than sending `null`.
It also accepts `wait` and `waitTimeout`, which hold the response open until the
session finishes the turn you just started. `wait` is `true` (the default signal
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
under `data.wait`. Sending them changes nothing for callers that do not: without
`wait` the response is still `{"success": true, "data": {}}` and the write is still
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
a **tagged duplicate** (a pair the server already applied) skips the write but still
waits, answering from the session's current state rather than blocking for a
transition that already happened. It reports `"delivered": false, "duplicate": true`.
### Waiting instead of polling
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
output), and the `wait` field above. Full parameter and response tables are in
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
whether your integration works, and the last one is what actually bites:
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
waits rather than issuing one long one, because `tailscale serve` and cloudflared
both cut idle connections and a single 10-minute call is the pattern most likely
to die in the field.
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
default). Read it rather than assuming you got what you asked for.
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
even `idle` fires only once at startup, so send-and-wait there can only time out.
See the Gotchas below.
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
`\r` does not get past it, and the session's startup `idle` lands inside the wait
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
# 1. Start a worker session (creates the case if it does not exist yet).
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
# later step "succeeds" against nothing.
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
# the previous turn's idle state. Single line only, ending in \r (otherwise
# Enter is never sent and this wait times out on a turn that never started).
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
| jq -c '.data.wait')
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
for _ in $(seq 1 30); do
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
W=$("${CURL[@]}" \
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
done
jq -r 'if .ended or .aborted then "worker is not running"
elif .timedOut then "still working after 30 waits"
else "signal: \(.signal)" end' <<<"$W"
# 5. Read what it produced, then delete the session YOU created, by exact id.
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Waiting on a marker instead of a signal is the form that works in **every** mode,
and the only one that works on a `shell` session:
```bash
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
# into the output stream, so an unsplit marker matches before the command has run.
# `from=buffer` also catches a marker that printed before the wait registered.
N=$RANDOM
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
```
For shell scripting, the `codeman` CLI is the same surface without the HTTP
plumbing:
```
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
codeman ralph start|stop|status|reset codeman users add|passwd|list
codeman status | list | attach <path> codeman doctor
```
## Seam 4: Hooks
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
Codeman installs its own hooks automatically, but the endpoint is open to yours.
```json
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
```
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
event.
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
the loopback bypass requires the `X-Codeman-Hook-Secret` header
(`~/.codeman/hook-secret`) unconditionally.
## Gotchas
Every one of these has cost somebody real time.
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
call the API. Integrate server-side.
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
work. Cross-site origins are blocked by the CSRF guard.
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
shipped bugs more than once.
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
simple-request CSRF, so it is deliberate. Send `application/json`.
- **Prompts are single-line and must end with `\r`.** The server splits your text
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
Enter **only when the input contains a carriage return**. Without it your text
sits on the prompt unsubmitted, which is the single most common "the wait
endpoints don't work" report: the wait runs its full timeout on a turn that never
started. Newlines inside the string are stripped rather than rejected, so
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
per call.
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
unique to each call, and build it so the typed line never contains it (your own
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
rejected with a `400` rather than ignored.
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
line included), a partial escape at a chunk boundary is held back until its tail
arrives, and a match may straddle PTY chunks, so text you printed yourself
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
words with cursor moves, so its text can reach the matcher **without spaces** and
a multi-word match is unreliable there. Match one short space-free token, ideally
one you printed yourself, and keep it out of the typed line (your own keystrokes
echo into the stream).
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini` or
`antigravity` sessions.** They come from Claude Code hooks, which no other mode
installs, so only `idle`, `working` and `exit` exist there. Asking for them
explicitly is a `400`; omitting `until` is safe, since the server drops them from
the default set and echoes what it actually waited on as `wait.until`. Even in
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
written by Codeman < 1.13.0 against an `--https` install carries hook curls
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
time a session starts in that case.
- **Unwrap the envelope** before reading fields. `data` is not the response body.
## Publishing your integration
There is no registry and no review queue. Add the GitHub topic
**`codeman-integration`** to your public repository so others can find it, and
link back to Codeman in your README.
If a real ecosystem of these appears, a manifest format and an install command
become worth building. Until then, these four seams are the contract, and they
require nothing of you but HTTP.
## What Codeman deliberately does not have
- **No in-process plugin runtime.** See the reasoning at the top of this page.
- **No build or startup hooks** for third-party code. Run your own process.
- **No per-plugin config or state directories.** Manage your own files.
- **No sandbox for integration code**, because Codeman never launches it. Your
integration is your own process, started by you, with your permissions,
talking HTTP.
That last point is about integration code specifically, not about Codeman.
Sandboxing lives on a different axis here: the thing worth isolating is the
**agent**, and you isolate it per case with
[Docker cases](docker-cases.md), which run the agent in a hardened container with
a bind-mounted workspace and seeded (not shared) credentials. An integration that
creates or drives a Docker-backed session inherits that isolation for free, since
it is a property of the session rather than of the caller.
+432
View File
@@ -0,0 +1,432 @@
# File Viewer edit mode (issue #212)
Plan only. No implementation yet.
Goal: close the loop "agent writes a file, you review it in the viewer, tweak two lines, save, tell the
agent to continue" without hopping into the terminal, with the phone as the primary target.
Scope from the issue: an Edit toggle on text previews, a write endpoint that inherits the read path's
confinement, text-only, edit-in-place (no create, no delete, no rename), no editing through the
Docker/remote overlays.
---
## 1. What exists today
**Read path (backend), all in `src/web/routes/file-routes.ts`:**
| Route | Line | Notes |
| ------------------------------------ | ------ | ------------------------------------------------------------------ |
| `GET /api/sessions/:id/files` | `741` | Tree scan of `session.workingDir`, hidden files off by default |
| `GET /api/sessions/:id/file-content` | `865` | The text/preview classifier. `findSessionOrFail` + `validateSessionFilePath` |
| `GET /api/sessions/:id/file-raw` | `1018` | Bytes, 50MB cap |
| `GET /api/sessions/:id/file-preview` | `1254` | DOCX/PPTX to PDF, everything else redirects to `file-raw` |
| `GET /api/download` | `1384` | The only read route that also runs `isSensitivePath()` |
`file-content` classification order (`file-routes.ts:881-1011`): extension buckets (image / video / audio /
known-binary) return metadata only; otherwise the bytes are read, sniffed for a NUL in the first 8KB, and
either reported as `type:'binary'` or decoded as UTF-8 and **truncated to `lines` (default 500, hard cap
10000)**. Caps: `MAX_TEXT_FILE_SIZE` 10MB.
Confinement is `validateSessionFilePath()` (`src/web/route-helpers.ts:67`): `resolve()` then `realpathSync()`
then reject if the result is not under `workingDir`. Because it realpaths the *full* path, a symlink whose
target escapes the workspace is already rejected. Ownership is `findSessionOrFail()` which runs
`canAccessOwned()` (`route-helpers.ts:102`), a no-op outside multi-user mode.
**Read path (frontend), `src/web/public/panels-ui.js`:**
- `loadFileBrowser()` `2947`, `renderFileBrowserTree()` `2978`, click to `openFilePreview()` `3056`.
- `openFilePreview(filePath, sessionId, attachmentId)` `3193`: attachment-id branch, then docx/pptx, pdf,
svg branches, then the generic `file-content` fetch at `3274` with **`&lines=500` hardcoded**, rendering
text as `<pre><code>${escapeHtml(...)}</code></pre>` at `3298` and stashing `this.filePreviewContent`.
- `closeFilePreview()` `3308`, `copyFilePreviewContent()` `3751`.
- Markup: `src/web/public/index.html:420-432` (`filePreviewOverlay` / `-Title` / `-Body` / `-Footer`, two
header buttons: copy and close).
- CSS: `src/web/public/styles.css:9320-9430`. Overlay `z-index: 2000`, window `80vw/80vh`, capped
`900x700`. There are **no `.file-preview-*` rules in `mobile.css` at all**.
**Reachability on phones.** The header File Viewer button is hidden below 430px
(`mobile.css:482`, locked by `KNOWN_PHONE_HIDDEN` in `test/mobile-header-buttons-policy.test.ts`), so on a
phone the preview overlay is reached through:
1. an attachment card's **Preview** button (`panels-ui.js:3451`), which is exactly the "agent just wrote a
file" path the issue describes,
2. the attachment-history drawer (`panels-ui.js:3709`),
3. App Settings to Panels to **File Browser** (`showFileBrowser`, applied in `settings-ui.js:2202`; the
panel is mobile-styled at `mobile.css:1868`).
So edit mode is reachable on a phone today via (1) and (2) without touching the header policy. Improving
the entry point is listed as an open decision in section 10, not assumed.
---
## 2. Threat model, stated honestly
Anyone who can call this API can already reach `POST /api/sessions/:id/input` and type an arbitrary prompt
into an agent running with `--dangerously-skip-permissions`. A workspace-confined write endpoint therefore
does not create a new privilege tier for an authenticated caller.
What it *would* create if built carelessly is a **new host-write primitive reachable by path**, so the
things this plan actually defends against are:
1. **Path traversal / symlink escape** writing outside the workspace.
2. **TOCTOU**: a path component that becomes a symlink between validation and write.
3. **Cross-user writes** in multi-user mode (`canAccessOwned`).
4. **Silent data loss**, which is the highest-probability real-world failure here and gets its own section.
CSRF is already covered: `registerHostGuard()` (`src/web/middleware/auth.ts:555-578`) rejects any
non-safe-method request whose `Origin` is cross-site. The webview-capability exemption at that gate is
fenced to `GET`/`HEAD` for the Referer form (`auth.ts:161`) and to `/webview/:cap/*` paths for the path
form, so a proxied dashboard cannot reach a new `PUT /api/...`. Using `PUT` + `application/json` also
forces a preflight for any cross-origin attempt.
---
## 3. Backend design
### 3.1 New policy module: `src/config/file-editing.ts`
Pure, unit-testable, no IO (config lives in `src/config/`, no barrel, import the file directly).
```ts
export const MAX_EDITABLE_BYTES = 512 * 1024; // content cap, both directions
export const EDITABLE_EXTENSIONS: ReadonlySet<string>; // ts,tsx,js,jsx,mjs,cjs,json,jsonc,md,mdx,txt,
// css,scss,less,html,htm,xml,svg?,yml,yaml,toml,
// ini,cfg,conf,env?,sh,bash,zsh,fish,py,rb,go,rs,
// java,kt,swift,c,h,cpp,hpp,cs,php,sql,graphql,
// proto,lua,pl,r,jl,tf,gradle,csv,tsv,log,diff,patch
export const EDITABLE_BASENAMES: ReadonlySet<string>; // Dockerfile, Makefile, LICENSE, .gitignore,
// .prettierignore, .editorconfig, .nvmrc, ...
export function isEditableFileName(fileName: string): boolean;
export function isDeniedEditRelativePath(rel: string): boolean; // `.git/` subtree
export function detectEol(text: string): 'lf' | 'crlf';
export function applyEol(text: string, eol: 'lf' | 'crlf'): string;
```
Decisions baked in:
- **Allowlist, not blocklist**, per the issue and per the existing attachment-guard precedent.
- `svg` and `env` are deliberately marked with `?` above: `svg` is served as an untrusted octet-stream on
the read side (`file-routes.ts:118`) so allowing an edit is defensible, but I recommend **excluding
both** in v1. `.env` files are matched by `isSensitivePath()` anyway and would be rejected downstream;
excluding them at the allowlist keeps a single obvious refusal.
- `isDeniedEditRelativePath` blocks the `.git/` subtree: `.git/hooks/*` is code execution and a corrupt
index is unrecoverable-looking to a user who only wanted to fix a typo. Other dotfiles stay allowed but
are not reachable from the tree UI anyway (`showHidden=false`).
### 3.2 Read-for-edit: extend the existing GET
`GET /api/sessions/:id/file-content?path=<rel>&edit=1`
When `edit=1`:
- skip line truncation entirely (a truncated buffer must never become an edit buffer, see section 4.1),
- enforce `MAX_EDITABLE_BYTES` instead of `MAX_TEXT_FILE_SIZE` and answer 413 over it (as a structured
throw with `statusCode: 413`, the `throwFilesystemPickerError` pattern, since the central errorCode-to-
status map has no 413 entry; see the error-mechanics note in 3.3),
- run the editability gate (`isEditableFileName`, `isDeniedEditRelativePath`, `isSensitivePath`,
`isBlockedAttachmentPath`) and the content gate (NUL sniff plus UTF-8 round-trip, see 4.3),
- return `{ content, size, mtimeMs, totalLines, truncated: false, extension, editable: true, hash, eol }`.
`hash` is `sha256` hex of the exact on-disk bytes.
Non-`edit` responses gain **only** `editable: boolean` (additive, no shape change for existing consumers),
which is all the UI needs to decide whether to show the Edit button. No `hash` on plain reads: the Edit
action re-fetches with `edit=1` anyway (section 4.1), which is where the hash comes from, and hashing every
casual 10MB preview would be pure waste.
### 3.3 Write: `PUT /api/sessions/:id/file-content`
Body (new `FileWriteSchema` in `src/web/schemas.ts`, Zod v4):
```ts
{ path: string, content: string, baseHash: string, eol?: 'lf'|'crlf', force?: boolean }
```
Registered with an explicit route option `{ bodyLimit: 4 * 1024 * 1024 }`. **Fastify's default `bodyLimit`
is 1MB and this repo configures none**, and JSON escaping expands content: 2x for a file full of quotes or
backslashes, up to 6x for control characters (each serialized as a `\uXXXX` escape), so 512KB of content
can legitimately exceed 1MB on the wire; blowing the limit produces a raw `FST_ERR_CTP_BODY_TOO_LARGE`, not an `ApiResponse` envelope. Two
related sizing notes: `z.string().max()` counts **UTF-16 code units, not bytes**, so the schema's `.max()`
is only a coarse pre-filter and the real cap is an explicit `Buffer.byteLength(content, 'utf8')` check in
the handler (step 7a below); and 4MB comfortably bounds the worst-case expansion of a 512KB file without
inviting multi-MB bodies elsewhere.
**Error mechanics** (matters for both prod behavior and testability): a handler that *returns* a
`{success:false, errorCode}` envelope gets its HTTP status assigned centrally by the preSerialization hook
in `server.ts` (`httpStatusForErrorCode()`, `src/types/api.ts`), but the route-test harness
(`test/routes/_route-test-utils.ts`) installs only `installRouteErrorHandler`, **not** that hook, so
returned envelopes surface as HTTP 200 in tests. The PUT handler should therefore use the same
structured-**throw** pattern as the filesystem picker (`throwFilesystemPickerError`, `file-routes.ts:411`):
thrown `{statusCode, body}` errors are rendered identically in prod and in the harness, and they allow the
one status the code map cannot express (413). The error envelope itself is strictly
`{success:false, error, errorCode}`, **it has no data arm**, so no error response may carry extra payload.
Handler order (each step is a test case):
1. `findSessionOrFail(ctx, id, req)` (live sessions only, matching the read route, and it carries the
multi-user ownership check).
2. `parseBody(FileWriteSchema, req.body)`, then `Buffer.byteLength(content, 'utf8') <= MAX_EDITABLE_BYTES`
or 413 (the schema `.max()` alone cannot enforce a byte cap, see the sizing note above).
3. `validateSessionFilePath(session.workingDir, path)` or 404 (do not distinguish "outside workspace" from
"missing", matching the read route).
4. `isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)` or 403.
5. `isDeniedEditRelativePath(relativePath)` or 403.
6. `isEditableFileName(basename(resolvedPath))` or 400.
7. `stat`: must be `isFile()`, size within `MAX_EDITABLE_BYTES`, else 400/413. **No `O_CREAT` anywhere in
this handler**, which is what enforces edit-in-place.
8. Read current bytes, compute `hash`, run the NUL sniff and the UTF-8 round-trip check, else 400.
9. `hash !== baseHash && !force` gives **409 CONFLICT** (`ApiErrorCode.CONFLICT`, plain envelope; the error
arm carries no data, see the error-mechanics note). The client's conflict dialog gets fresh state by
re-fetching `edit=1`, which it needs for its Reload action anyway.
10. Build the output buffer: `applyEol(content, eol ?? detected-from-original)`; re-check
`Buffer.byteLength` against the cap.
11. Write atomically in the resolved parent directory:
`fs.open(<dir>/.<name>.codeman-tmp-<rand>, 'wx', stat.mode & 0o777)`, then `fchmod(stat.mode & 0o777)`
(open's mode argument is masked by the process umask, so the chmod is what actually preserves an
unusual mode), write, `fsync`, close, `fs.rename(tmp, resolvedPath)`, unlink the temp on any failure.
12. Re-stat, return `{ success: true, data: { path, size, mtimeMs, hash, totalLines } }`.
Why `O_EXCL` temp plus rename rather than truncate-in-place:
- `wx` cannot follow a pre-existing symlink, which closes the TOCTOU window from step 3 to step 11 without
needing `O_NOFOLLOW` gymnastics.
- `rename()` does not follow a symlink in the final component, so even if `resolvedPath` were swapped for a
symlink after validation, the symlink itself is replaced and the swap target is untouched.
- A crash mid-write leaves the original intact.
Caveat to document in the code comment: rename replaces the inode, so hardlinks to the file keep the old
content. That is the same trade-off vim makes by default and is preferable to a truncate window here.
No SSE event in v1. Nothing else in the app needs to know: `image-watcher.ts` only reacts to
`.png/.jpg/.jpeg/.gif/.webp/.bmp/.svg/.pdf/.docx/.pptx` adds (`image-watcher.ts:23-25`), none of which are
editable text, and the temp filename does not match either.
---
## 4. The five traps
These are the parts that turn a "small write endpoint" into a bug report.
### 4.1 Truncation (the data-loss trap)
The frontend fetches `&lines=500` (`panels-ui.js:3274`). Saving that buffer back would **delete every line
past 500**. Worse, the content hash of the full file would still match, so an optimistic-concurrency check
cannot catch it.
Mitigations, all three:
- The Edit affordance is only offered when the loaded payload came from `edit=1` (which never truncates).
Tapping Edit on an already-rendered preview **re-fetches** with `edit=1` before swapping in the editor.
- The read-for-edit path 413s above `MAX_EDITABLE_BYTES` rather than truncating, so "too big to edit here"
is an explicit refusal with a message, never a silent partial buffer.
- A test asserts `edit=1` never returns `truncated: true`.
### 4.2 Line endings
A `<textarea>`'s `.value` normalizes to LF. Saving a CRLF file naively rewrites every line, producing a
whole-file diff for a two-line change. So: the read returns the detected `eol`, the client echoes it back
unchanged, and the server re-applies it. Mixed-EOL files use the dominant style, which is lossy for the
minority lines; call that out in the response and accept it in v1.
### 4.3 Encoding
`buf.toString('utf-8')` on a latin-1 or otherwise non-UTF-8 file yields U+FFFD replacement characters, and
writing that back **corrupts the file**. The check is a round-trip:
`Buffer.from(decoded, 'utf8').equals(buf)`. If it fails, `editable: false` and the write is refused. This
also catches binary content that the NUL sniff misses. A UTF-8 BOM survives because it round-trips as a
leading U+FEFF; do not strip it.
### 4.4 Concurrency with the agent
The whole use case is editing a file the agent just wrote and may write again. `baseHash` plus 409 is the
guard. Do not use mtime alone: agents rewrite files within a single filesystem timestamp tick, and an
identical rewrite should not be reported as a conflict.
### 4.5 Symlinks and TOCTOU
Covered by `validateSessionFilePath` (escape) plus `wx` temp and `rename` (post-validation swap). One
intentional allowance: a symlink whose target is *inside* the workspace is edited through to its target,
because `validateSessionFilePath` returns the realpath. That matches what a user tapping the file expects.
---
## 5. Frontend design
All in `panels-ui.js` (prettier-exempt, hand-formatted; match the surrounding style), `index.html`,
`styles.css`, `mobile.css`.
### 5.1 State
```js
filePreviewEdit = { active, sessionId, path, baseHash, eol, original, dirty }
```
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
### 5.2 Markup (`index.html:420-432`)
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
untouched.
### 5.3 Behavior
- `openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
keyboard before the user has picked a line.
- `input` sets `dirty` and enables Save.
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
response, leave edit mode, re-render the read view from the local editor value (the response carries
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
itself carries no state (section 3.3, step 9).
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
guard.
- `copyFilePreviewContent()` copies the live editor value while editing.
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
real bug twice.
### 5.4 Mobile
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
the editor tracks the keyboard for free.
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
smaller `rem` value.
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
- No header-policy change is needed (section 1), so
`test/mobile-header-buttons-policy.test.ts` stays untouched.
### 5.5 i18n
`i18n.js` already skips `textarea`, `pre`, `code` and `.file-preview-content` in its `SKIP_SELECTOR`
(`i18n.js:20-38`), so file content is never translated. Add zh-CN entries for the new chrome: Edit, Save,
Cancel, Unsaved changes, Discard unsaved changes?, File changed on disk, Reload, Overwrite, Saved,
Too large to edit here.
---
## 6. Docker and remote cases
Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
---
## 7. Tests
| File | Kind | Covers |
| ------------------------------------------- | ----------- | ---------------------------------------------------------------------- |
| `test/file-editing-policy.test.ts` | pure unit | `isEditableFileName` (allow + deny + basenames), `isDeniedEditRelativePath`, `detectEol`/`applyEol` round-trip incl. mixed EOL, BOM preservation |
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
Route cases to assert explicitly:
1. happy path writes the bytes and returns a new hash
2. `../` and absolute paths give 404
3. symlink pointing outside the workspace gives 404
4. symlink pointing inside is written through to the target
5. non-allowlisted extension gives 400
6. `.git/config` gives 403
7. a `.env` in the workspace gives 403 (sensitive-path)
8. a file with a NUL byte gives 400
9. a latin-1 file that fails the UTF-8 round-trip gives 400
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
11. over `MAX_EDITABLE_BYTES` gives 413
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
14. CRLF file edited and saved stays CRLF
15. file mode is preserved across the temp-plus-rename
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
- `curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
`w1`/`w2`/`w3`; delete it by exact id afterwards.
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
assert the bytes on disk changed. Assert real state, not HTTP 200.
---
## 8. Docs and release
- This plan lives at `docs/file-viewer-edit-plan.md`.
- `docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
chain, the truncation invariant, and why temp-plus-rename.
- `CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
Remember `CLAUDE.md` is prettier-ignored on purpose.
- `docs/api-reference.md`: the new `PUT` and the `edit=1` query.
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
additive API surface, so **COM minor** when it ships.
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
and must pass `npm run format:check`.
---
## 9. Implementation order
Each phase is independently reviewable and leaves the tree working.
1. **Policy module + tests.** `src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
route wiring. (Small.)
2. **Read-for-edit.** `edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
tests. Nothing consumes it yet. (Small.)
3. **Write endpoint.** `FileWriteSchema`, `PUT` handler, `test/routes/file-write-routes.test.ts`. Fully
testable by curl before any UI exists. (Medium, the security-relevant part.)
4. **Desktop UI.** Edit button, textarea swap, Save/Cancel, dirty guard, 409 flow. (Medium.)
5. **Mobile pass.** `mobile.css` sizing against `--app-height`, font size, accessory-bar interaction,
real-device check. (Small but the part that decides whether the feature is actually usable.)
6. **Docs, i18n strings, changeset.**
---
## 10. Open decisions
1. **Editor widget.** Recommend a plain `<textarea>` for v1: zero dependencies, no CSP question, no bundle
growth, and it is the only thing guaranteed to behave with the iOS keyboard. CodeMirror-light with
syntax highlighting is a clean follow-up once the write path is proven. The issue allows either.
2. **Phone entry point.** Edit mode is reachable on a phone through attachment cards and the history
drawer without changing anything. A dedicated toolbar or overview affordance for "browse this session's
files" would make it discoverable, but it is a separate UX change and would need a decision against the
deliberately minimal phone header policy. Recommend deferring it and revisiting after the feature ships.
3. **`svg` editability.** Recommend excluded in v1 (it is deliberately treated as untrusted on the read
side). Easy to add later.
4. **Create / delete / rename.** Explicitly out of scope per the issue. Note that keeping `O_CREAT` out of
the handler is what makes that a structural property rather than a convention.
Binary file not shown.

After

Width:  |  Height:  |  Size: 357 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 941 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.0 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 357 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 82 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.0 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 28 MiB

Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 537 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 332 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 808 KiB

Binary file not shown.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 806 KiB

+282
View File
@@ -0,0 +1,282 @@
# Multi-User Mode: Design Plan
Status: **IMPLEMENTED on `feat/multiuser-mode`** (phases 1-5; opt-in, off by default). Target: opt-in multi-user support behind a `--multiuser` flag, with per-user case spaces and an admin panel for user management.
Shipped by phase:
- **Phase 1** (user store + mode plumbing + CLI): `src/user-store.ts` (scrypt, atomic 0600 writes, last-admin invariants, serialized read-modify-write), `src/config/multiuser.ts`, `codeman users add|passwd|list|rm`, `--multiuser` flag, bootstrap-on-first-boot. Tests: `test/user-store.test.ts`.
- **Phase 2** (multi-user auth): parallel async auth branch (`src/web/middleware/auth.ts`), `req.authUser`, per-username rate bucket, `mustChangePassword` lockbox, `GET /api/me` + `POST /api/me/password`, QR identity-bound minting, network-bind + tunnel exemptions, new error codes. Tests: `test/multiuser-auth.test.ts`.
- **Phase 3** (ownership threading): `Session.owner` at every create path + recovery mirror; `findSessionOrFail` owner check + list filtering; §6.3 permission policy (`resolveClaudeModeForUser` at all spawn sites incl. one-shots via `buildPromptArgs`; shell/launchCommand grant); per-user case spaces (`resolveCasesDir`) + owner-scoped case list + admin-only host CRUD; `workingDir` confinement; `sessionCapacityState` per-user cap. Tests: `test/ownership-scoping.test.ts`.
- **Phase 4** (event fan-out): WS owner gate; SSE per-client identity + `broadcast`/terminal-batch routing (`deriveSseHint`, fail-closed); `getLightState` per-identity filtering; file-route preview/thumbnail/history + `GET /api/search` scoping.
- **Phase 5** (admin API + frontend): `src/web/routes/admin-routes.ts` (user CRUD, one-time passwords, last-admin guards, session revoke/kill) + `src/web/admin-audit.ts`; `public/admin-ui.js` (identity boot, change-password modal + interceptor, admin Users tab). Tests: `test/admin-routes.test.ts`, `test/admin-ui.test.ts`.
Deferred follow-ups (documented, non-blocking): away-digest + subagent/workflow REST-list scoping, push-subscription identity/routing, per-user screenshot subdirs, `linked-cases.json` v2 owner field, `ScheduledRun.owner`, plan-orchestrator internal one-shot mode resolution, and a Playwright browser pass. Phase 6 (login form replacing Basic) remains out of scope.
## 1. Summary
Today Codeman is strictly single-user: one optional credential pair (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`), one shared `~/codeman-cases` folder, one global session list, and a global SSE/WS fan-out. This plan adds an opt-in **multi-user mode**:
- **Off by default.** Without the flag, behavior stays byte-identical to today (same auth path, same paths, same payloads). All new code is gated behind `isMultiUserMode()`.
- **`codeman web --multiuser`** (or `CODEMAN_MULTIUSER=1`) enables named users with individually hashed passwords stored in `~/.codeman/users.json`.
- **Each user gets their own space**: `~/codeman-users/<username>/cases/<case>` replaces the shared `~/codeman-cases` for that user. Sessions, cases, attachments, search, digests, and SSE events are scoped to their owner.
- **Admin panel** (App Settings, admin-only "Users" tab): create/delete users, change/reset passwords, enable/disable accounts, delete a user's space, see per-user live sessions and disk usage, force logout.
## 2. Threat Model (read first, be honest about this)
Multi-user mode is **workspace separation for a trusted team, NOT security isolation between mutually distrusting users**:
- Every session still runs as the **same OS account** with `claude --dangerously-skip-permissions`. Any user can ask their agent to `cat /home/<host>/codeman-users/otheruser/...`. The web layer enforces scoping; the agent layer cannot.
- **Shell sessions and custom launch commands are the bluntest holes**: `SessionMode = 'shell'` hands out a raw shell as the host account, and a cron job's `launchCommand` runs an arbitrary command; no Claude permission classifier is involved in either. These must be gated behind the same grant as bypass (section 6.3), otherwise the `auto`-mode mitigation below is theater.
- All sessions share one tmux socket (`-L codeman`), one `~/.claude` (transcripts, credentials, plan usage), one Claude subscription.
- Mitigation for stronger isolation: pair a user's cases with **Docker cases** (container per case, `docs/docker-cases.md`), or run separate Codeman instances per user (`CODEMAN_INSTANCE`, separate OS accounts). True per-user OS isolation is explicitly **out of scope** for this feature.
- Partial mitigation at the agent layer: non-admin users default to Claude's `auto` permission mode (section 6.3), whose safety classifier blocks destructive actions and credential exfiltration. That reduces, but does not eliminate, cross-user snooping; the `canBypassPermissions` grant reopens it and should be given deliberately.
This must be stated loudly in `docs/security-architecture.md`, the README section, and the admin panel UI ("Users share the host account; this separates workspaces, it does not sandbox users from each other").
Also note the flip side: multi-user mode strictly _improves_ today's network posture, because it removes the single shared password and gives every person their own revocable credential.
## 3. Activation and Mode Rules
| Condition | Behavior |
| ------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| No flag (default) | Exactly today's behavior. `users.json` is never read. Single-user auth via `CODEMAN_PASSWORD` if set. |
| `--multiuser` / `CODEMAN_MULTIUSER=1`, `users.json` has users | Multi-user auth active. `CODEMAN_PASSWORD` is ignored for login (warn if set). |
| `--multiuser`, no `users.json` (first boot) | Bootstrap: if `CODEMAN_USERNAME`/`CODEMAN_PASSWORD` are set, create that user as the initial admin and continue. Otherwise refuse to start with instructions to run `codeman users add <name> --admin`. Never start multi-user with zero users (there would be no way in). |
| `--multiuser` on a non-loopback bind | Allowed without `CODEMAN_PASSWORD`: `server.ts start()` treats "multi-user with >= 1 enabled user" as satisfying the auth requirement in the loud-warning check (wire into the existing `isLoopbackBindHost()` branch). |
| Flag later removed | Single-user mode again. Sessions/state that carry `owner` fields keep working (owner is simply ignored); user spaces remain on disk untouched. |
Plumbing: flag in `src/cli.ts` (web command), env in a new `src/config/multiuser.ts` exporting `isMultiUserMode()`. Per-instance like everything else: a beta instance (`CODEMAN_INSTANCE=beta`) has its own `users.json` via `dataPath()`.
## 4. Data Model and Disk Layout
### 4.1 `~/.codeman/users.json` (via `dataPath('users.json')`, mode 0600, atomic write: tmp + rename)
```jsonc
{
"version": 1,
"users": [
{
"username": "alice", // canonical lowercase slug
"role": "admin", // "admin" | "user"
"password": {
"algo": "scrypt", // node:crypto scrypt, no new deps
"N": 16384,
"r": 8,
"p": 1,
"salt": "<hex 32B>",
"hash": "<hex 64B>",
},
"disabled": false,
"mustChangePassword": false, // set by admin reset; gates all API access until changed
"canBypassPermissions": false, // permission-mode grant, see section 6.3; false for new users
"createdAt": 1752900000000,
"lastLoginAt": 1752900000000,
},
],
}
```
- **Username rules**: `^[a-z0-9][a-z0-9_-]{1,31}$` (it becomes a folder name), stored lowercase, unique case-insensitively. Reserve `admin`? No: any name can be admin; role is a field, not a name.
- **Hashing**: `scrypt` from `node:crypto` with per-user salt, compared via `timingSafeEqual`. Params stored per record so they can be raised later; verify tolerates old params and rehashes on next successful login.
- New module `src/user-store.ts` (mirrors the `remote-hosts.ts` / `docker-hosts.ts` pattern): `readUsers()`, `writeUsers()`, `verifyPassword()`, `createUser()`, `setPassword()`, `deleteUser()`, plus pure helpers (`isValidUsername`, `hashPassword`) that are unit-testable without IO. In-process cache with short TTL like `readSettings`, invalidated on every write; the short TTL also covers the CLI (section 10) editing `users.json` while the server runs (cross-process changes picked up within the TTL).
### 4.2 User spaces
```
~/codeman-users/
alice/
cases/
my-project/ <- same layout as today's ~/codeman-cases/<case>
bob/
cases/
```
- New helper in `route-helpers.ts`:
`resolveCasesDir(user?: AuthUser): string`
single-user mode: returns `CASES_DIR` (today's `~/codeman-cases`); multi-user: returns `join(USER_SPACES_DIR, user.username, 'cases')`, creating it lazily on first use.
- `CASES_DIR` stays exported for single-user code paths, but every route usage (see 6) switches to the resolver.
- The **user folder** (`~/codeman-users/<username>/`) is the deletion unit for "delete user + space" and leaves room for future per-user extras (uploads, exports) beside `cases/`.
- Legacy `~/codeman-cases` in multi-user mode: surfaces to admins only, as a read-only "Unassigned (legacy)" group in the case list, with an admin action `POST /api/admin/cases/assign { case, username }` that `fs.rename`s the folder into a user's space (same-filesystem move, cheap). No automatic migration.
## 5. Auth Pipeline Changes (`src/web/middleware/auth.ts`)
Keep the existing single-user branch untouched. Add a parallel multi-user branch selected once at registration time:
1. **Credential check**: Basic header parsed into `username:password`, verified against the user store (scrypt + `timingSafeEqual`). Disabled users fail closed.
2. **Cookie sessions**: same `codeman_session` cookie and `StaleExpirationMap`, but `AuthSessionRecord` gains `username` and `role`. All existing TTL/sliding/eviction logic reused. Eviction cap becomes per-user aware (evict oldest _of that user_ first) so one user cannot flush everyone's sessions by logging in 100 times.
3. **Request identity**: decorate `req.authUser = { username, role }` (Fastify decorateRequest). In single-user mode `req.authUser` is `{ username: 'admin', role: 'admin' }` when auth is on, and a synthetic admin when auth is off, so downstream code has ONE code path.
4. **Rate limiting**: keep the per-IP bucket; add a per-username failure bucket (same `StaleExpirationMap` pattern) so a botnet cannot brute-force one account across IPs, and one flaky user behind a NAT cannot lock out the rest.
5. **`mustChangePassword` gate**: when set, every API request except `GET /api/me`, `POST /api/me/password`, and static assets returns 403 with `errorCode: 'PASSWORD_CHANGE_REQUIRED'`; the frontend intercepts that code and shows the change-password modal.
6. **Password change vs Basic-auth caching**: browsers cache Basic credentials. After a password change we revoke all of that user's cookie sessions; the next request falls to Basic with stale creds, gets 401, and the browser re-prompts. Acceptable for v1; a proper login form is Phase 6 (see 15).
7. **Unchanged**: hook-secret loopback bypass (hooks authenticate the _instance_, not a user; the event maps to a session which has an owner), host guard, Origin/CSRF guard, security headers.
8. **WS upgrade identity** (`ws-routes.ts`): the global auth `onRequest` hook does run on the upgrade request (`@fastify/websocket` v11 runs hooks before the handshake; browsers send the session cookie), but the route handler itself only checks Host/Origin and never learns WHO authenticated. Multi-user: the handler reads the decorated `req.authUser` and closes 4003 unless owner or admin (section 6.4; identity plumbing lands in Phase 2, the owner check in Phase 4 once sessions have owners). Add a regression test that an upgrade with no credentials is rejected while auth is active: the handler-level Host/Origin gate alone must never be mistaken for auth.
9. **QR auth** (`/q/:code` redemption in `system-routes.ts`, minting in `tunnel-manager.ts`): today there is ONE global token, auto-rotated every 60s with a 90s grace window. A globally-rotating token cannot carry an identity (every logged-in user sees the same code), so multi-user mode replaces rotation with **on-demand minting**: an authenticated `POST /api/tunnel/qr` mints a single-use, short-TTL token bound to `req.authUser.username` (field on `QrTokenRecord`); redemption creates a cookie session for that user. Existing rate-limit buckets (`qrAuthFailures`, global `QR_RATE_LIMIT_MAX`) apply unchanged. Single-user mode keeps the rotating token.
New error codes in `src/types/api.ts`: `FORBIDDEN`, `PASSWORD_CHANGE_REQUIRED`, `USER_EXISTS`, `USER_NOT_FOUND`, `LAST_ADMIN`.
Role guard helper in `route-helpers.ts`: `requireAdmin(req, reply): boolean` used as the first line of every admin handler (403 `FORBIDDEN`), plus `requireOwnerOrAdmin(req, session)`.
## 6. Ownership Threading (the big refactor)
### 6.1 Sessions
- `Session` gains `owner?: string` (constructor option), persisted in `SessionState.owner`, included in `toState()`, round-tripped through recovery (`mux-sessions.json` entries carry it, `restoreMuxSessions` passes it back, exactly like `remote`/`docker`).
- Every session-creating path stamps the owner from `req.authUser`. Verified inventory of `new Session(...)` call sites: `POST /api/sessions` (session-routes.ts:444), `POST /api/quick-start` (:1956), `POST /api/run` one-shot (:1652), Ralph start (ralph-routes.ts:327), **cron** (cron-service.ts:352; `CronJob` gains `owner`, stamped at job create, launched as the job's owner), legacy `ScheduledRun` loop (server.ts:1603), plan generation + plan-orchestrator agents (plan-routes.ts:128, plan-orchestrator.ts:422/578; owner = requesting user), and recovery (server.ts:2225, next bullet). Two non-paths, also verified: **respawn never constructs a new Session** (it re-spawns the PTY on the same object, so `owner` survives automatically; no inheritance logic needed), and **orchestrator-loop creates no sessions** (it schedules work onto existing idle sessions via the task queue; its scoping requirement is different: it must only pick idle sessions owned by the goal's creator).
- Recovery: `owner` must ALSO be mirrored on `MuxSession` (mux-sessions.json) and read back mux-first like `remote`/`docker` (`muxSession.owner ?? savedState?.owner`, the server.ts:2246-2250 pattern), or a reboot erases ownership on the next persist.
- Every session-reading/mutating route filters: non-admin users only see and act on `session.owner === req.authUser.username`. Centralize in `findSessionOrFail` (route-helpers.ts:87; the owner check there covers the 6 route files that use it: system/session/respawn/ralph/file/plan-routes) and in the list endpoints (`GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status`). The Phase 3 audit must grep for BOTH `sessionManager.getSession` AND direct map access (`ctx.sessions.get(` / `.has(`): ws-routes and hook-event-routes reach sessions that way and bypass `findSessionOrFail`.
- Admins see everything; every session row carries `owner` so the UI can badge it.
### 6.2 Cases
- All `CASES_DIR` call sites switch to `resolveCasesDir(req.authUser)`: `case-routes.ts` (list/create/delete/CLAUDE.md scaffolding, name-collision checks, docker quickcreate), `session-routes.ts` (quick-start case resolution, the workingDir-inside-cases env-strip check), `ralph-routes.ts` (case path resolution), and `plan-routes.ts:231` (easy to miss). Case-name-to-path resolution is currently DUPLICATED (`resolveCasePath` in case-routes.ts:82 and an inline copy in quick-start, session-routes.ts:1846-1863); consolidate into one owner-aware resolver as part of this refactor instead of patching both copies.
- Registries that map case names to metadata become owner-scoped. `remote-cases.json`/`docker-cases.json` are arrays of objects, so entries simply gain `owner?: string` (absent = legacy: admin-only). `linked-cases.json` is a flat `Record<caseName, path>` with no room for a field: it needs a v2 shape (`{ "version": 2, "cases": { "<name>": { "path": "...", "owner": "..." } } }`) with read-time migration of the v1 form; it is read in two places (case-routes AND inline in quick-start), both must move to the new reader. Case names only need to be unique per user.
- **Remote hosts and Docker hosts are machine-level resources**: CRUD on `/api/docker-hosts` and remote-host endpoints becomes admin-only in multi-user mode; regular users can _use_ hosts on their own cases but not define them. (Docker containers exec as the host account; letting any user define arbitrary `docker run` args is admin-equivalent.)
- Case deletion, exports (`docker-exports/`), and imports check ownership; export filenames get an owner prefix to avoid collisions (fits the existing `^[a-zA-Z0-9._-]+\.tgz$` download guard).
- **Workspace confinement for non-admins (the linchpin, do not skip)**: today `POST /api/sessions` accepts ANY host directory as `workingDir` (the only check is `statSync().isDirectory()`, session-routes.ts:305-318), and file-routes/attachments confine reads to `session.workingDir`. Without a new rule the whole scoping story is circular: a user points a session at `~/codeman-users/bob` (or `/home`) and the web layer itself serves that subtree, no agent needed. Rule: in multi-user mode a non-admin's `workingDir` must realpath-resolve inside their own space, enforced at `POST /api/sessions`, `POST /api/run`, cron job create AND fire time (the dir can change owners between the two), and Ralph auto-configure. Admins are unrestricted. This one rule is what makes the section 6.4 file-route line ("own space or own sessions' workingDirs") meaningful.
### 6.3 Per-user Claude permission-mode policy
Codeman now ships a global **Startup Mode** picker (App Settings, Claude CLI tab: `settings.claudeMode`, values `dangerously-skip-permissions` (default) | `auto` | `normal` | `allowedTools`; `auto` emits `--permission-mode auto`, Anthropic's classifier-guarded low-prompt mode). Multi-user mode layers a per-user policy on top of it:
- **Default for regular users: `auto` only.** A non-admin's Claude sessions are forced to `--permission-mode auto` regardless of the global `claudeMode` setting. `normal` and `allowedTools` are also permitted (they are strictly more restrictive than auto), but `dangerously-skip-permissions` is NOT.
- **Bypass is an explicit admin grant**: `canBypassPermissions: true` on the user record (default `false`, section 4.1). Only with that grant does the global skip-permissions default (or a future per-user choice) apply to their sessions.
- **Admins** are unrestricted; the global setting applies to them as-is.
- **Single enforcement point**: a pure `resolveClaudeModeForUser(globalMode, user)` in `user-store.ts`, applied server-side at option-resolution time, BEFORE the Session constructor, so both downstream arg builders inherit it for free (`buildPermissionArgs` in session-cli-builder.ts for the direct-PTY path AND `buildClaudePermissionFlags` in tmux-manager.ts for tmux panes; there are two builders, not one). Call sites where `getClaudeModeConfig()` feeds a spawn: session-routes.ts:452/1964, ralph-routes.ts:334, cron-service.ts:360, and recovery (server.ts:2214/2233). Recovery re-reads the GLOBAL setting on reboot, so the resolver must run there with the RECOVERED owner, or a restart silently un-downgrades every restored session. Never resolved in the frontend, so it cannot be bypassed via payload.
- **Downgrade, don't error**: a non-granted user whose effective mode would be bypass gets `auto` silently (logged + surfaced as a badge on the session), so shared presets keep working.
- **Other CLIs' bypass equivalents** follow the same grant: Codex `--dangerously-bypass-approvals-and-sandbox` (`codexDangerouslyBypassApprovals`) and Gemini `--approval-mode yolo` are refused for non-granted users (Gemini falls back to `auto_edit`, Codex to its default sandbox). Whether this stays one grant or splits per-CLI is an open question (section 15).
- **Shell mode and custom launch commands follow the grant too**: `mode: 'shell'` sessions and cron `launchCommand` are arbitrary command execution as the host account, strictly stronger than any bypass flag, and no permission-mode downgrade applies to them. Non-granted users get 403 `FORBIDDEN` on shell session/quick-start creation and on cron jobs carrying `launchCommand` (checked at create AND at fire time). Folding them under `canBypassPermissions` keeps the model one-bit; section 15 asks whether it should split.
- **Admin UI**: a "Can skip permissions" toggle per user in the Users tab (PATCH field, section 8), with a warning echoing the section 2 threat model.
- Revoking the grant takes effect on the user's NEXT session start; live sessions are listed so the admin can restart them.
### 6.4 Everything else that lists or streams
| Surface | Scoping rule |
| -------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| SSE `/api/events` | Per-connection filter (see 7) |
| WS terminal (`ws-routes.ts`) | Handler reads `req.authUser` (section 5.8) and closes 4003 unless owner or admin; today it checks Host/Origin only and has no identity |
| `GET /api/search` | `harvestSources()` only over owned sessions |
| `GET /api/away-digest` | Aggregate only owned sessions/events |
| `GET /api/subagents`, workflow runs | Filter by owning session (`claudeSessionId -> session -> owner`); agents not attributable to any session: admin-only |
| Push (`push-routes.ts`) | Subscription records currently carry NO identity (keyed by endpoint only): `subscribe` stamps `username`. All 8 `PUSH_EVENT_MAP` events are session-scoped, so routing = resolve owner from `data.sessionId`, deliver to that owner's (plus admins') subscriptions. Legacy identity-less subscriptions: admin-only delivery |
| Screenshots `/api/screenshots` | Per-user subdir `~/.codeman/screenshots/<username>/` in multi-user mode. Note: `GET /:name` deliberately rejects `/` in names as traversal, so derive the subdir server-side from `req.authUser` and keep client-visible names flat |
| Attachments | Already session-scoped; inherits the session owner check. `attachmentConfineToWorkspace` is a global, default-OFF setting today: in multi-user mode it is FORCED ON for non-admins regardless of the setting (their attachments must resolve inside their own space); the setting keeps meaning what it means for admins |
| File routes (browse/preview) | Path allowlist adds: non-admin paths must resolve (realpath) inside their own space or their own sessions' workingDirs |
| Settings (`settings.json`) | Global, admin-only writes in multi-user mode; reads allowed (per-device display keys stay in localStorage as today). Per-user server settings: out of scope v1 |
| System ops (self-update, tunnel toggle, span-displays, docker image build) | Admin-only |
| `getLightState` init snapshot | Filtered per connection. Actual contents to filter (verified): `sessions`, `scheduledRuns`, `respawnStatus`, `subagents`, `workflowRuns`, `planUsage` (host-plan telemetry: admin-only); `globalStats` stays coarse-global. Cron jobs are NOT in the snapshot (they have their own REST route; filter there). The snapshot is cached process-wide (`LIGHT_STATE_CACHE_TTL_MS`): either key the cache per role/user or filter AFTER the cache on each send |
## 7. SSE Event Filtering
`/api/events` currently broadcasts everything to everyone. Ground truth first (verified): `broadcast()` lives in `SseStreamManager` (`sse-stream-manager.ts`), not server.ts; clients are keyed by the raw Fastify reply (`sseClients: Map<FastifyReply, Set<string> | null>`, plus `sseClientsById` for live filter updates); the existing `?sessions=` filter is a bandwidth optimization applied ONLY to `session:terminal` batches in `flushSessionTerminalBatch()`, while `broadcast()` itself loops ALL clients unconditionally. The single-client delivery primitive already exists (`sendSSE`, used for the per-connection init snapshot). Plan:
- At connection time, resolve `req.authUser` and store `{ username, role }` with the client. Concretely: extend `addClient(reply, sessionFilter, isRemote, clientId)` to take the identity and change the `sseClients` map value to `{ filter, identity }` (or add a parallel `Map<reply, identity>`); there is no per-client record object today to hang it on.
- `broadcast()` gains an optional routing hint: `broadcast(event, data, { sessionId?, adminOnly?, username? })`. Resolution order per client: admin sees all; `username` targets one user; `sessionId` resolves owner via SessionManager; `adminOnly` for machine-level events (docker image builds, tunnel, self-update); no hint = broadcast to all (connection status etc.).
- **Enforce the identity check in BOTH `broadcast()` AND `flushSessionTerminalBatch()`**: the terminal batch path does not go through `broadcast()`, and it carries the highest-value payload (raw terminal bytes).
- Sweep of the ~120 backend event constants in `sse-events.ts`: mechanically, everything `session:*`, `ralph:*`, `respawn:*`, `subagent:*`, `workflow:*`, `attachment:*`, `cron:*` (job owner) carries or can resolve a sessionId/owner; `docker:*`, `system:*`, tunnel and update events are adminOnly; a short tail needs case-by-case decisions during implementation.
- The existing `?sessions=` filter and `/api/events/subscribe` compose with (never override) the ownership filter: the subscription filter can only narrow within what the identity allows.
## 8. Admin API (`src/web/routes/admin-routes.ts`, new module + `AdminPort`)
All handlers: multi-user mode only (404 otherwise), `requireAdmin`, Zod schemas in `schemas.ts`, `ApiResponse` envelope, audit-logged.
| Endpoint | Behavior |
| ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `GET /api/admin/users` | List users + stats: role, disabled, createdAt, lastLoginAt, live session count, case count, space disk usage (best-effort async walk, cached 60s), active cookie-session count |
| `POST /api/admin/users` | Create: `{ username, role, password? }`. No password given: generate a one-time password, return it ONCE in the response, set `mustChangePassword` |
| `PATCH /api/admin/users/:username` | `{ role?, disabled?, canBypassPermissions? }`. Demoting/disabling the last enabled admin: 409 `LAST_ADMIN`. Disable also revokes cookie sessions. `canBypassPermissions` is the section 6.3 grant (default false) |
| `POST /api/admin/users/:username/reset-password` | Generates one-time password (returned once), sets `mustChangePassword`, revokes cookie sessions |
| `POST /api/admin/users/:username/logout` | Revoke all cookie sessions for that user. Honest limit under Basic auth: the browser silently re-sends cached credentials and gets a fresh cookie on the next request, so logout only truly ends QR-issued sessions; to actually lock someone out, disable the account or reset the password. Say so in the panel tooltip until Phase 6 |
| `DELETE /api/admin/users/:username` | `{ deleteSpace?: boolean }` (default false). Refuses last admin. Kills the user's live sessions first (normal kill flow, incl. docker/remote teardown per case), revokes cookies, removes from store. With `deleteSpace`: guarded recursive delete of `~/codeman-users/<username>` (realpath must be inside `USER_SPACES_DIR`, top-level dir must not be a symlink), plus their registry entries and push subscriptions |
| `POST /api/admin/cases/assign` | Move a legacy `~/codeman-cases/<case>` into a user's space (`fs.rename`) |
| Self-service `GET /api/me` | `{ username, role, mustChangePassword }` (works in single-user mode too: synthetic admin; the frontend uses it to decide whether to render admin UI) |
| Self-service `POST /api/me/password` | `{ currentPassword, newPassword }`, verifies current, min length 8, revokes other sessions, clears `mustChangePassword` |
**Audit log**: append-only `~/.codeman/admin-audit.jsonl` (same idiom as `session-lifecycle.jsonl`): timestamp, acting admin, action, target, request IP. User management without an audit trail is not acceptable even for a homelab tool.
SSE additions (both `sse-events.ts` and `constants.js`): `admin:usersChanged` (adminOnly; the panel re-fetches) and `auth:passwordChangeRequired` (targeted to the user).
## 9. Frontend
- **`GET /api/me` on boot** (app.js init): stores `window.__codemanUser`; everything below keys off it. Single-user mode returns the synthetic admin, so the UI needs no mode awareness beyond "am I admin".
- **Admin panel**: new tab "Users" in the App Settings modal (settings-ui.js), rendered only for admins in multi-user mode. Table of users with actions (create, reset password showing the one-time password in a copy-to-clipboard reveal, enable/disable, role toggle, logout, delete with a typed-username confirm for the delete-space variant). No new header button (mobile header policy test stays green; the settings modal is already reachable everywhere).
- **Change-password modal**: shown on `PASSWORD_CHANGE_REQUIRED` (fetch interceptor in api-client.js) and reachable from settings for self-service.
- **Owner badges**: admin's session tabs and the session palette/manager show `owner` on foreign sessions; regular users see no change.
- New module `admin-ui.js` if the settings-ui.js addition gets large (load order after settings-ui, before session-ui), else keep inside settings-ui.js. Follow the `@fileoverview` + `@loadorder` convention either way.
## 10. CLI Additions (`src/cli.ts`)
Headless bootstrap and recovery must not require the web UI:
```
codeman users add <name> [--admin] # prompts for password (hidden input), or --password-stdin
codeman users passwd <name> # reset password
codeman users list
codeman users rm <name> [--delete-space]
```
These operate directly on `users.json` via `user-store.ts` (no server needed), honoring `CODEMAN_INSTANCE`. This is also the answer to "locked out: last admin forgot password".
## 11. Limits and Config
- New `src/config/multiuser.ts`: `isMultiUserMode()`, `USER_SPACES_DIR` (`~/codeman-users`, overridable via `CODEMAN_USER_SPACES_DIR` for tests), `MAX_USERS` (default 25), per-user session cap (default: global cap / 2, env `CODEMAN_MAX_SESSIONS_PER_USER`).
- Cap enforcement is currently COPY-PASTED: the global `MAX_CONCURRENT_SESSIONS` (50, `config/map-limits.ts:25`) check appears at 6 independent sites (session-routes.ts:298/1622/1683, ralph-routes.ts:275, cron-service.ts:340, server.ts:1595). Do not add a 7th copy per site: extract one `assertSessionCapacity(ctx, owner?)` helper doing the global + per-user checks and use it everywhere, or the per-user cap WILL miss a path.
- Global limits (50 sessions, SSE clients 100, terminal buffers) are unchanged and shared; the per-user session cap is the fairness lever.
## 12. Compatibility Matrix
| Concern | Guarantee |
| ------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Default (no flag) | No behavior change. No new file reads on the hot path. All new fields optional in state |
| State round-trip | `SessionState.owner`, `MuxSession.owner`, `CronJob.owner`, registry `owner` fields are optional; old state loads clean; new state loaded by an old build ignores unknown fields (existing tolerant parsing) |
| Instance isolation | `users.json`, audit log, screenshots subdirs all via `dataPath()`; user spaces dir is shared across instances like `~/codeman-cases` is today (documented) |
| API versioning | HTTP API is internal per `docs/versioning-policy.md`; still, all changes are additive. Ship as a **minor** version |
| Hooks | Unchanged (instance-level hook secret; owner resolved from the session) |
## 13. Implementation Phases
Each phase is independently shippable behind the flag and ends with its tests green.
**Phase 1: user store + mode plumbing** (no behavior change yet)
`src/user-store.ts`, `src/config/multiuser.ts`, CLI `users` subcommands, bootstrap-on-first-boot logic, `users.json` schema + atomic writes.
Tests: `test/user-store.test.ts` (hashing, verify, params upgrade, username validation, atomic write, last-admin invariants; pure, no server).
**Phase 2: multi-user auth**
Auth middleware branch, `req.authUser` decoration, cookie records with username/role, per-username rate bucket, `mustChangePassword` gate, WS upgrade identity plumbing + unauthenticated-upgrade regression test (section 5.8), QR on-demand minting + identity binding (section 5.9), `GET /api/me`, `POST /api/me/password`, error codes, network-bind check integration.
Tests: `test/multiuser-auth.test.ts` (live server, unique port 3170+; wrong password, disabled user, cookie carries identity, per-user rate limit isolation, mustChangePassword lockbox, QR redemption identity). Reuse the `delete process.env.CODEMAN_PASSWORD` idiom from `test/setup.ts`.
**Phase 3: ownership threading**
Session `owner` + persistence + `MuxSession` mirror + recovery; `resolveCasesDir()` refactor across case/session/ralph/plan routes (consolidating the duplicated case-path resolution); registry owner fields incl. the linked-cases v2 shape; `findSessionOrFail` owner check + the direct-`sessions.get` audit; list filtering; owner stamping across ALL create paths from 6.1; **non-admin workingDir confinement** (6.2); permission-mode/shell/launchCommand policy (6.3); `assertSessionCapacity` helper + per-user cap.
Tests: `test/routes/ownership-scoping.test.ts` (inject-based: user A cannot read/kill/input user B's session, case lists are disjoint, admin sees both), extend `test/cron-service.test.ts` for owner stamping, recovery round-trip in the existing mux-recovery tests.
**Phase 4: event fan-out + remaining surfaces**
SSE routing hints + client identity (enforced in BOTH `broadcast()` and the terminal-batch flush), WS owner gate (identity landed in Phase 2), search/digest/subagent/workflow scoping, push subscription identity + owner routing, screenshot subdirs, file-route scoping, `getLightState` filtering + per-identity caching, admin-only system ops.
Tests: `test/sse-ownership.test.ts` (two SSE clients, event for A's session reaches only A + admin), WS upgrade rejection test, search/digest scoping tests.
**Phase 5: admin API + frontend**
`admin-routes.ts` + `AdminPort` + schemas + audit log + `admin:usersChanged`; settings-ui Users tab, change-password modal, owner badges, api-client interceptor.
Tests: `test/routes/admin-routes.test.ts` (CRUD, last-admin 409, one-time password flow, delete-space guard rails incl. symlink refusal), frontend vm-sandbox test following `test/run-mode-ui.test.ts` pattern, Playwright pass per the always-end-to-end rule before calling it done.
**Phase 6 (optional, later): login page**
Replace Basic with a form + `POST /api/login` in multi-user mode only (fixes browser credential caching UX, enables logout button). Explicitly deferred; Basic works for v1.
**Docs**: update `docs/security-architecture.md` (new section: multi-user model + threat model from section 2), `README.md` (short opt-in section), `CLAUDE.md` (Key Patterns entry + State Files + route/SSE counts), this file gets a "shipped" status stamp per phase.
## 14. Key Risks / Decisions Made
1. **Not a security boundary at the agent layer** (section 2). Decided: ship with loud documentation; Docker cases are the isolation story.
2. **`findSessionOrFail` as the single enforcement point** for ~30 session routes: any route that fetches sessions another way must be audited in Phase 3 (grep for `sessionManager.getSession` outside route-helpers).
3. **SSE sweep is the riskiest surface**: a missed event leaks metadata (not terminal content, which is session-scoped, but names/paths). Phase 4 includes a checklist pass over all ~138 events with the default flipped to "owner-scoped unless explicitly global": fail closed.
4. **Basic-auth password-change UX** is mediocre (browser re-prompt). Accepted for v1; Phase 6 fixes it properly.
5. **Legacy case migration** is manual (admin assigns). No silent moves of user data.
6. **Case-name uniqueness becomes per-user**; tmux session names already include the session id so no collision, but the `w<n>-<case>` tab naming and lifecycle-log rows should include the owner for disambiguation in admin views.
7. **`workingDir` confinement (6.2) is the single most load-bearing rule**: every file-serving and agent-spawning surface downstream trusts `session.workingDir`. Review and test it as carefully as the auth branch (foreign-space path, symlink into a foreign space, `..` traversal, cron fire-time re-check).
8. **The WS handler never sees identity today** (auth happens only in the global hook): the 5.8 wiring is new code on a security-sensitive path; cover unauthenticated, foreign-user, and admin upgrades with tests.
## 15. Open Questions (answer before Phase 3)
1. Should admins' own cases live in `~/codeman-users/<admin>/cases` (symmetric, proposed) or keep using legacy `~/codeman-cases`? Proposed: symmetric; legacy dir is a migration source only.
2. Per-user settings (respawn presets, notification prefs): global-only in v1. Worth a `users/<name>/settings.json` overlay later?
3. Should regular users be allowed to create Docker cases on admin-defined hosts (proposed: yes) or is Docker entirely admin-only?
4. Session handoff: does an admin need "reassign session/case to another user"? (Cheap to add next to `cases/assign`; not in v1 scope.)
5. Permission-mode grants (section 6.3): one `canBypassPermissions` flag covering Claude/Codex/Gemini bypass equivalents PLUS shell mode and cron `launchCommand` (proposed: one flag, keep it one-bit), or split into `canBypassPermissions` + `canRunArbitraryCommands`? And should admins be able to set a per-user DEFAULT mode (for example force `normal` for an intern) rather than just gating bypass?
6. OpenCode has no single bypass flag (its permission config rides `OPENCODE_CONFIG_CONTENT`): decide what the grant means there before Phase 3, or exclude OpenCode mode for non-granted users in v1.
+246
View File
@@ -0,0 +1,246 @@
# Remote Sessions (SSH)
Codeman can run a session's agent on a **remote host over SSH** instead of the
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, or a plain shell)
runs inside a `tmux` server **on the remote host**, so it survives the SSH
connection dropping; Codeman attaches to it the same way it attaches to a local
managed session.
This document covers the data model, the shell-safe SSH command construction
(COD-107), the durable-launch design (COD-104), and the operational caveats.
For the local session/mux machinery this builds on, see the **Mux** and
**Session** entries in `CLAUDE.md` → Architecture.
## Why it exists
A developer box (`AA-DESKTOP`) often needs to drive an agent on another machine —
a NAS, a build server, a host reachable only through a jump box or a
cloudflared SOCKS5 proxy. Rather than wrap `ssh` by hand per host, Codeman
stores reusable **remote hosts** + **remote cases** and reproduces the exact
connection the operator already uses (`ssh-aa-desktop`-style configs:
custom port, identity file, `-J` jump host, `-o ProxyCommand`).
## Data model
Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| Type | Role |
|------|------|
| `RemoteSshOptions` | The **HOW-to-reach** fields, shared by host + session: `identityFile`, `socksProxy` (`host:port`), `jumpHost` (`[user@]host[:port]`), `extraSshOptions` (`KEY=VALUE[]`). Every field optional — all-absent reproduces port-22, default-identity, directly-SSH-able behavior. |
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
- `~/.codeman/remote-hosts.json` — `readRemoteHosts()` / `writeRemoteHosts()`
- `~/.codeman/remote-cases.json` — `readRemoteCases()` / `writeRemoteCases()`
(Paths via `remoteHostsPath()` / `remoteCasesPath()`; both honor `CODEMAN_INSTANCE`
because the config dir is the instance data dir.)
On the live `Session`, the remote rides as `_remote?: SessionRemote`. When
attaching, `resolveMuxAttachCwd()` forces the cwd to `/tmp` for remote sessions —
the local working directory is meaningless on the remote box.
## SSH command construction (COD-107 — the injection surface)
**All** SSH command lines flow through one function so user-controlled fields are
escaped once and the launch + prereq probe can never drift apart:
```ts
// src/remote-hosts.ts
buildSshConnectionArgs(remote: RemoteSshOptions & Pick<RemoteHost, 'port'>): string[]
```
It returns the **ordered leading tokens** of an ssh command line (no `-t`, no
target, no remote command):
```
ssh -o BatchMode=yes
[-p <port>]
[-i <abs-identity>] # ~ / $HOME expanded, then shellescaped
[-J <jumpHost>] # shellescaped, single token
[-o ProxyCommand=nc -X 5 -x <socks> %h %p] # ONE shellescaped -o token
[-o <KEY=VALUE>] … # each extra option, shellescaped
```
Rules that keep this safe — **do not bypass them by hand-building an ssh line elsewhere:**
- **Every** user-controlled value (`-i`, `-J`, `-o`, ProxyCommand) is POSIX
single-quote `shellescape`d (`'…'` with embedded `'\''`). The helper mirrors
the one in `tmux-manager.ts`.
- **`~`/`$HOME` in `identityFile` is expanded at build time** (`expandIdentityPath`),
*before* escaping — ssh does not expand `~` inside `-i`, and the escaped value
never reaches a shell that would.
- **The ProxyCommand is one shellescaped `-o KEY=VALUE` token**, so its spaces and
the `%h`/`%p` placeholders reach ssh as a single argument. `%h %p` survive
verbatim — **ssh** expands them to the real host/port, not the shell.
- **Empty options ⇒ `['ssh', '-o BatchMode=yes']`** (+ `-p` only when set) —
byte-identical to the historical behavior.
Token construction is unit-tested independently of any live connection (see
`test/` for `buildSshConnectionArgs` / `buildRemoteTmuxCheckCommand` cases).
## Durable launch (COD-104)
`buildRemoteLaunchCommand({ mode, remote, sessionId })` in `tmux-manager.ts`
builds the command that launches (or **reattaches** to) the remote session:
```
ssh -o BatchMode=yes -t <connection-args> user@host \
'tmux -L codeman-remote new-session -A -s codeman-ssh-<id8> -c <remotePath> "cd <remotePath> && exec <cli>" \; \
set -t codeman-ssh-<id8> status off \; set -t codeman-ssh-<id8> mouse off \; \
set -t codeman-ssh-<id8> prefix C-q \; set -s escape-time 0 \; \
set -t codeman-ssh-<id8> window-size latest'
```
Key points:
- **`new-session -A -s codeman-ssh-<id8>`** = attach-if-exists-else-create, so a
reconnect (same deterministic `remoteTmuxSessionName(sessionId)` — `codeman-ssh-` +
the first 8 chars of the session id) lands back in
the **same** remote session rather than spawning a duplicate. This is what makes
the remote agent survive an SSH drop. The name deliberately fails
`SAFE_MUX_NAME_PATTERN` so a Codeman running ON the remote host never adopts it.
- **`-L codeman-remote`** = a DEDICATED socket for sessions launched by remote
Codemans, NOT the canonical `-L codeman` socket the remote host's own Codeman
uses. Options are set per-session (`set -t`), never `-g`, so a shared remote
tmux server's other sessions are untouched (#145 hardening). Note the
asymmetry: **discovery/attach (COD-105) target the canonical `-L codeman`
socket** — they join sessions the remote's own Codeman manages, while owned
durable launches live on `-L codeman-remote`.
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID *is*
the agent. The per-mode command comes from `remote.commands?.[mode]` or
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
pane command is independently quoted, so a `remotePath` with spaces is safe.
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
the prereq probe; `-t` is inserted right after `ssh -o BatchMode=yes`,
preserving historical token order.
### tmux prerequisite probe
Because durable remote sessions require tmux on the remote host,
`checkRemoteTmuxAvailable(host)` runs `command -v tmux` over SSH **before**
creating a remote case/session and returns a structured, never-throwing result:
- empty stdout / non-zero exit → *"remote host `<host>` needs tmux installed for
durable remote sessions"*
- stderr present → *"could not verify tmux on remote host `<host>`: `<stderr>`"*
(a real connection failure, surfaced to the operator)
- success → `{ ok: true, tmuxPath }`
It connects with the **identical** options as the launch
(`buildRemoteTmuxCheckCommand` reuses `buildSshConnectionArgs` and inserts
`-o ConnectTimeout=10`), so a proxied/custom-port/identity host that the launch
can reach also passes the probe (and vice-versa).
**Test-mode short-circuit:** under `VITEST` the probe returns
`{ ok: true, tmuxPath: '(test-mode)' }` without opening a socket — mirroring
`TmuxManager`'s no-op-shell-under-VITEST (`IS_TEST_MODE`). Without it, remote-case
create-path tests would hit a real ~10s ssh timeout. Only the live probe is
skipped; command construction is still asserted by unit tests.
## Ownership: launched vs. discovered-and-attached (COD-105)
COD-104 (above) was Phase 1 — Codeman *launches* a remote session and owns it.
COD-105 is Phase 2 — Codeman can also **discover** `codeman-*` tmux sessions
already running on a remote host (created by the remote's own Codeman or another
instance) and **attach** to one it didn't launch. Ownership decides what happens
when the tab closes.
`SessionRemote.owned` carries this:
- **`owned: true`** (or absent — legacy/COD-104 sessions persisted before this
field) — we launched it via `buildRemoteLaunchCommand` and may explicitly kill it.
- **`owned: false`** — discovered + attached; another Codeman owns the remote
session. `remoteSessionName` holds its existing tmux name. Closing the tab
**detaches**, never kills.
### Discovery
`listRemoteCodemanSessions(host)` lists the remote's `codeman-*` sessions:
- `buildRemoteListSessionsCommand()` runs `tmux -L codeman list-sessions -F "…"`
over SSH (connection args from the shared `buildSshConnectionArgs`, so discovery
connects identically to launch/probe). `2>/dev/null` swallows tmux's "no server
running" stderr.
- `parseRemoteSessionList()` is a **pure, unit-tested** parser. ⚠️ Quirk: the
remote tmux's `-F "…\t…"` format emits the **literal two-character `\t`**, not a
real tab (verified on tmux next-3.7), so the parser splits on `/\\t|\t/` (literal
backslash-t **or** a real tab, for builds that do expand it). It keeps only
`codeman-*` names, coerces types, and skips malformed lines.
- `listRemoteCodemanSessions()` **never throws** — unreachable host / no tmux / no
sessions all map to `[]`. Like the prereq probe, it **no-ops to `[]` under
`VITEST`** so a request path never opens a real ssh connection.
Discovery is **explicit** — the UI has a "Discover existing sessions" button per
host; Codeman never auto-discovers on host select.
### Attach vs. launch selection
`buildRemoteSessionCommand(mode, remote, sessionId)` in `tmux-manager.ts` picks the
remote command line by ownership:
- **`owned === false`** → `buildRemoteAttachCommand(remote, name)` — emits
`ssh … -t … 'tmux -L codeman attach -t <remoteSessionName>'`. It uses **`attach`,
NOT `new-session -A`**, so it only *joins* an existing session and never creates
one.
- **owned (default)** → `buildRemoteLaunchCommand` (the COD-104 path above).
### Detach-not-kill
`TmuxManager.killSession()` has an **early return for non-owned remote sessions**:
it tears down **only the LOCAL pane** holding the ssh client (`tmux -L codeman
kill-session` on *this* host's socket). Killing the local ssh sends SIGHUP to the
remote `tmux attach`, which **detaches** — the durable remote session survives.
The early return is a structural guarantee that **no code path can ever issue a
remote `kill-session` for a session we don't own** — the only `kill-session` run is
on the local socket, which never reaches the remote socket.
## API
Routes are registered in `src/web/routes/case-routes.ts`:
| Method | Path | Purpose |
|--------|------|---------|
| `GET` | `/api/remote-hosts` | List saved hosts |
| `POST` | `/api/remote-hosts` | Create a host |
| `PUT` | `/api/remote-hosts/:id` | Update a host |
| `DELETE` | `/api/remote-hosts/:id` | Delete a host |
| `GET` | `/api/remote-hosts/:hostId/sessions` | Discover `codeman-*` sessions on the host (COD-105; `listRemoteCodemanSessions`, never errors) |
| `POST` | `/api/cases/remote-link` | Link a case to a remote host (creates the `RemoteCase`) |
Attaching to a discovered session is a **session-create** path, not a host route:
`POST /api/sessions` accepts `attachRemoteSession: { hostId, remoteSessionName }`
(schema in `schemas.ts`; `remoteSessionName` must match `^codeman-[a-zA-Z0-9._-]+$`),
which `session-routes.ts` turns into a non-owned (`owned: false`) session.
Frontend touchpoints: the remote-host management UI is in `session-ui.js` /
`panels-ui.js`; a remote session is created by picking a remote host/case in the
session-create flow, or via the per-host **"Discover existing sessions"** button →
**Attach** action (creates an `owned: false` session).
## Security notes
- **`identityFile` is a path only — never key bytes.** Codeman stores the path and
passes it to `ssh -i`; the key never enters Codeman's state or the wire.
- The injection surface is the SSH option fields. The single-source
`buildSshConnectionArgs` + `shellescape` discipline (COD-107) is the control —
audit any new code path that constructs an ssh command to route through it
rather than concatenating options inline.
- `BatchMode=yes` means **no interactive password/passphrase prompts** — remote
hosts must be reachable with key-based or agent auth (or an unencrypted key the
agent has loaded). A host needing a passphrase will fail the probe with an ssh
diagnostic rather than hang.
## Related
- `CLAUDE.md` → Architecture → **Remote** row, and the **Remote sessions (SSH)**
Key Pattern.
- `docs/security-architecture.md` — overall network/auth model.
- COD-104 (tmux prereq + durable launch), COD-105 (discover + attach, detach-not-kill ownership), COD-107 (shell-safe connection args).
Binary file not shown.

Before

Width:  |  Height:  |  Size: 894 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 576 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 661 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 452 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 99 KiB

+301
View File
@@ -0,0 +1,301 @@
# Scrollback fix plan (issue #205)
Status: IMPLEMENTED on `fix/scrollback-shell-alt-screen` (2026-08-07), with one deliberate
divergence from the recommendation below. Kept for the diagnosis record; the measured evidence
behind it is `docs/scrollback-issues-analysis.md`, and the mechanisms as shipped are documented
in `docs/architecture-invariants.md` (§ Full-scrollback replay, § Terminal scrollback: strip
flavors and wheel/touch forwarding).
What shipped vs. what this doc proposed:
- **Bug A (deltaMode)**: implemented as specified (`_wheelScrollLines()` normalizes
line/page/pixel units, Shift-axis trap kept).
- **Bug B (shell scrollback)**: implemented via the NARROW alt-screen strip for tmux-backed
shell/opencode/antigravity plus the scroll-to-top `full=1` re-pull, NOT the recommended
approach (a) `tmux mouse on`. The measurements in the analysis doc showed the alt buffer
comes from tmux's own client-side `smcup` at attach (tmux never forwards a pane program's
alt-screen toggles), so stripping that one sequence fixes both symptoms with no selection
tradeoff, keeps vim/less/htop untouched, and the re-pull also covers the repaint-burst
history loss that `mouse on` would not have addressed.
- **Invariant change**: the "viewport-at-bottom gate stays" invariant below was deliberately
DROPPED for forwarding modes: a repaint-mode CLI keeps no real terminal scrollback, so the
gate pinned users to a buffer of stale frames whenever the viewport parked off-bottom.
Forwarding now snaps to bottom first; Shift+wheel and the opt-out setting keep local
scrollback reachable. Touch forwards through the same gate (the mobile half of the fix).
- **Finding 5 (remote probe)**: implemented (`probeRemoteCliVersion` over ssh, deferred at
session start, same login-shell wrapper as the launch).
## RETEST FAILED (2026-08-07, after v1.12.0 shipped) — analysis round 2
mtiller retested on 1.12.0 and reports it is NOT fixed (issue #205 comment, 2026-08-07 12:12 UTC;
issue reopened same day with clarifying questions: mouse vs trackpad, Shift+scroll behavior,
Claude vs shell session on the phone, and an iOS full-tab-kill to rule out stale JS). Two
failure signatures, now analyzed against the SHIPPED 1.12.0 code (not the pre-fix code):
1. **iPhone Safari (Claude session assumed)**: touch scrollback goes back only a limited
amount and sometimes REPEATS blocks of text; unreliable.
2. **Firefox on macOS (mouse)**: wheel does NOTHING at all, while Fn+Up (= PageUp) pages back
through INTACT text.
### Ruled out by code reading
- deltaMode mishandling: `_wheelScrollLinesFloat` normalizes line/page/pixel units correctly;
a Firefox line-mode notch yields ±3 lines. Not the bug.
- Ephemeral transport: `_sendInputEphemeral` (app.js) has a POST fallback when WS is down.
- Service worker: sw.js is network-first with cache fallback; it serves stale JS only when the
fetch FAILS (flaky mobile connection can do this — relevant to "unreliable" on the phone,
and the fixed `CACHE_NAME = 'codeman-v1'` never invalidates that offline copy).
### The load-bearing observation: PageUp works, the wheel does not
Fn+Up is a KEYBOARD event: xterm encodes PageUp and Claude pages its own transcript (intact
text proves Claude-side history is fine and the PTY input path is fine). The wheel path is the
capture-phase handler, and for a Claude session it has exactly two branches:
- **Forwarding branch** (`_shouldForwardWheelToApp` true): snap-to-bottom + SGR reports. If
this branch ran, the user would see the same paging motion Fn+Up produces. They see nothing.
- **Local branch** (gate false): `_smoothScrollBy` over xterm's local buffer. For a Claude
pane in repaint mode, tmux keeps `history_size≈0`, so `?full=1` returns roughly one frame:
the local buffer is structurally HOLLOW, the top-of-buffer re-pull recovers nothing, and the
wheel looks completely dead. **This matches every observed detail on Firefox.**
So the working hypothesis is that mtiller's sessions evaluate the gate FALSE. The gate
(`_shouldForwardWheelToApp`) has exactly four false-paths worth checking, in likelihood order:
1. **`terminalWheelLocalScrollback` opt-out is ON.** Plausible: a user whose scrolling was
broken on 1.11.x may well have toggled "Wheel scrolls local history" while trying to fix
it. On 1.12.0 that setting now routes the wheel to a hollow local buffer = dead wheel on
desktop AND the stale-repaint-frames experience on the phone (see below). Ask, or check
what the setting does on their export.
2. **`cliVersion` missing — CONFIRMED BUG, independent of whether it is mtiller's**:
`getClaudeCliVersion()` (utils/claude-cli-resolver.ts:124-148) caches its result
process-wide including FAILURE: on any exception it sets `_claudeVersion = null`, and the
guard is `!== undefined`, so a single failed/timed-out probe (5s `EXEC_TIMEOUT_MS`; PATH
under systemd/launchd; transient fs hiccup) at the FIRST Claude session start disables
wheel forwarding for every Claude session until the server restarts. Fix: cache success
permanently, but let failure retry (retry on next call, or a short negative-cache TTL).
Note that mtiller sees identical breakage on phone + iPad + laptop, which points at a
SERVER-side/session-side cause exactly like this (cliVersion is shared by all devices)
rather than anything browser-specific.
3. **Claude Code genuinely < 2.1.187** on their machine: gate false BY DESIGN, but the
resulting UX is a dead-end (no local history to fall back on).
4. mouseTrackingMode non-none (a DECSET leaked past the strip, e.g. emitted before attach or
split across chunks in a way the carry missed): would also kill the container handler via
the early return. Least likely, checkable via `terminal.modes.mouseTrackingMode` in console.
### The iPhone symptoms fit the same gate-false story
Touch with gate false = local `scrollLines()` over whatever repaint frames accumulated:
"repeats blocks of text" is literally what a buffer of successive overlapping repaint frames
looks like; "limited amount" is its thinness; "unreliable" is burst-dependence (finding 2)
PLUS the new re-pull being actively DESTRUCTIVE for repaint panes: `_maybeRefetchFullHistory`
does `_resetTerminalForReplay()` then writes the fetched capture, and when that capture is
one frame (Claude pane, `history_size≈0`) it REPLACES a multi-frame buffer with less than the
user had, mid-scroll. Stale pre-1.12 JS on the phone (suspended Safari tab) remains possible
until they confirm the tab kill.
### Fix directions, ranked
1. **Make the re-pull refuse downgrades** (`_maybeRefetchFullHistory`, app.js): if the fetched
capture would yield FEWER buffer rows than currently present, skip the reset+rewrite and
keep the richer buffer (optionally cache-mark the session "re-pull useless"). Small, safe,
kills the "got worse after scrolling to top" class. Consider skipping the re-pull entirely
for forwarding-capable modes where tmux keeps no history.
2. **Rescue the gate-false Claude dead-end with PageUp forwarding**: when mode is `claude`,
the gate is false, AND the local buffer has no scrollback (`baseY === 0`), translate wheel
lines into coalesced PageUp/PageDown key sends (mtiller just proved Claude pages correctly
on PageUp even on their version). Zero regression risk under that triple guard: sessions
with real local history keep local scrolling; only the currently-dead path changes.
Caveat: older Claude menus may react to PageUp; acceptable against "completely dead".
3. **Audit `getClaudeCliVersion()` failure caching** (utils/claude-cli-resolver.ts): a cached
empty probe must retry (with backoff), not poison the process.
4. **Guard the opt-out setting's footgun**: if `terminalWheelLocalScrollback` is ON for a
repaint-mode CLI session, local history is hollow; either scope the setting's effect to
modes with real local scrollback, or pair it with fix 2's PageUp fallback so it still
scrolls SOMETHING.
5. **Add a one-line gate diagnostic**: log (once per session, console) WHY the wheel chose
local vs forward: `{mode, cliVersion, optOut, trackingMode}`. The #205 thread is now two
rounds deep on guesswork a single console line would have answered.
### What shipped for round 2 (branch `fix/scrollback-205-round2`)
All five directions above, implemented as ranked:
1. **Downgrade guard** — `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the rows a
capture will occupy (ANSI stripped, `capture-pane -J` re-wrapping accounted for) and
`_maybeRefetchFullHistory` (app.js) skips the reset+rewrite when that is more than one
screen short of what xterm already holds. A refused session goes on
`_fullHistoryRepullUseless`, which raises its re-pull cooldown from 4s to 60s so a hollow
pane stops re-fetching. Measured A/B on a live Claude pane, same gesture, same buffer:
guard off → 341 rows collapse to 42 and every seeded row is gone; guard on → 341 rows
preserved. The tab-switch recovery it must not break still runs (shell buffer 401 → 44 on
a tab switch → 401 again after scrolling to the top).
2. **PageUp/PageDown fallback** — `_maybePageCliTranscript()` translates wheel/touch travel
into coalesced `\x1b[5~` / `\x1b[6~` under the triple guard (claude mode, forwarding gate
false, `baseY === 0`), through the same 40ms queue as the SGR reports. Half a screen of
travel per page: the page key always jumps a whole screen, and a 1:1 mapping was
unusably slow with a discrete wheel. Shift is excluded — it keeps meaning "local
scrollback". Verified live: opt-out ON on a Claude session sends real PageUp/PageDown to
the PTY where the wheel previously did nothing.
3. **Probe caching** — `getClaudeCliVersion()` no longer caches failure. Success is kept for
the process lifetime; a failed probe retries with a 1/2/4…15min backoff. The cache policy
is a pure function (`resolveClaudeCliVersion`) so the retry semantics are unit-testable
without spawning `claude`. The VITEST short-circuit now records nothing, where before it
wrote a permanent null.
4. **Opt-out footgun** — handled by pairing rather than by scoping: the setting keeps meaning
exactly what it says (the wheel goes local), and fix 2 catches the case where "local" is
empty. Scoping the setting away from repaint-mode CLIs would have silently overridden an
explicit user choice. The App Settings tooltip now says to leave it off for Claude/Codex.
5. **Diagnostic** — `_logScrollRouting()` prints one line per session per distinct decision:
`[scroll] <id> → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…,
cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. That
single line answers every open question in the list below.
Still unanswered by code alone: whether mtiller's Claude Code is genuinely older than
2.1.187 (false-path 3), and whether the iPhone was running stale JS. The diagnostic makes
both self-reporting, so the retest ask is now "open the console and paste the `[scroll]` line".
### What to get from mtiller (some already asked)
- Shift+scroll behavior on Firefox (distinguishes hollow-local from handler-not-firing).
- `claude --version` on the Mac (decides false-paths 2 vs 3).
- App Settings → Input → "Wheel scrolls local history" state (false-path 1).
- iPhone: Claude or shell session, and whether a full tab kill changes anything.
- Browser console: `app.terminalUi?.terminal?.modes?.mouseTrackingMode` (false-path 4).
## ROUND 3 (2026-08-09): Codex wheel dead — CONFIRMED AND FIXED
DodgyBadger (Codex latest, Chrome, Windows 11): mouse wheel does nothing in a CODEX session
while working fine in shell and web tabs; DRAGGING THE SCROLLBAR WORKS, so xterm's local
buffer demonstrably has content for their codex pane. Analysis against the shipped code:
- `_shouldForwardWheelToApp` returns true UNCONDITIONALLY for `codex` (no version gate, unlike
claude's `>= 2.1.187`), so every plain wheel tick is sent as SGR reports to Codex.
- The "verified to scroll its transcript on SGR wheel reports" claim for codex predates
current Codex builds; if Codex latest ignores SGR wheel, forwarding eats the gesture while
the healthy local scrollback (proven by the working scrollbar) sits unused.
- The #227 PageUp fallback cannot rescue this: it is gated to `claude` mode AND `baseY === 0`,
and codex here has real local scrollback. The `[scroll]` diagnostic will still say
`forward-sgr (mode=codex, ...)`, confirming the branch, worth asking the reporter to paste.
**CONFIRMED by the reporter's `[scroll]` line (2026-08-09, PR #227 comment)**:
`forward-sgr (mode=codex, cliVersion=unknown, localScrollbackOptOut=false, mouseTracking=none,
localScrollbackRows=967)`. Forwarding branch active, 967 rows of healthy local scrollback
unused, Codex ignoring the SGR reports. Environment: Codex latest, Chrome, Windows 11.
**Measured against codex-cli 0.147.0** (isolated `tmux -L codexwheel`, fake `CODEX_HOME/auth.json`,
history built with 401ing prompts), which settles it without needing a version gate at all:
| Probe | Result |
| ---------------------------------------------- | ----------------------------------------------- |
| `#{mouse_any_flag}` once the TUI is up | `0`: codex never enables mouse tracking |
| `#{alternate_on}` | `0`: inline viewport, not an alt-screen pager |
| `#{history_size}` while prompting | grows 3 → 32: the transcript goes to scrollback |
| 6 × `\x1b[<64;10;10M` written to the pane | pane capture byte-identical, nothing happens |
| control: literal `zz` | pane changes, so the probe can see changes |
| `\x1b[<0;12;5M` + release (the click-tap path) | no change either: taps are no-ops, not garbage |
Codex has no in-app pager to drive: its history lives in the terminal's own scrollback, which is
exactly what forwarding was stealing the gesture from. A version gate would be the wrong fix (and
`cliVersion=unknown` means there is no codex probe to gate on anyway).
**Fix (shipped):** `_shouldForwardWheelToApp` now returns true for `claude >= 2.1.187` and nothing
else. Codex falls to the normal local-scrollback path like shell/gemini/opencode, so wheel and touch
scroll the same history the scrollbar drag was already scrolling. The claude-only PageUp fallback is
untouched: codex never needs it, its local buffer is real. Taps stay hand-encoded for codex
(`_sessionUsesServerMouseStrip`), measured harmless, so click-to-position is merely unavailable
there rather than damaging. Lesson for the next mode added to the forward list: "it is a strip mode"
proves nothing, write a real SGR report into a live pane and diff the capture first.
Verified end-to-end in Chromium against a live codex session on an isolated instance
(`CODEMAN_INSTANCE=codexwheel`, port 5055, `envOverrides.CODEX_HOME` pointing at the fake auth
dir): trusted `page.mouse.wheel` up now logs
`[scroll] … → local-scrollback (mode=codex, …, localScrollbackRows=43)`, moves the viewport
39 → 4 (back to the Codex banner), and sends ZERO bytes to the PTY. Unit coverage:
`test/terminal-touch-tap.test.ts` ("only claude forwards — codex and gemini keep the local wheel").
Original plan follows.
## Reports
- **Issue #205** (https://github.com/Ark0N/Codeman/issues/205), OPEN:
- **jonocodes** (author, 2026-08-03): SHELL session. Host Mac M4, brew tmux. On Android, touch-scrolling the terminal does nothing. On desktop, the mouse wheel cycles shell command history (acts like Up/Down arrows) instead of scrolling the screen.
- **mtiller** (comment, 2026-08-06): "similar issue just with scrolling backward to see agent output. This is with Firefox on MacOS." (Claude session implied.)
- **Reddit r/selfhosted** comment `p21x6ts` by mmtiller (= mtiller on GitHub): scrolling broken enough across phone/iPad/laptop that they fall back to Claude's own remote-control feature. Churn-risk user who otherwise loves the product; fixing this has promo value beyond the bug itself.
## How scrolling works today (read this before touching anything)
Three independent paths, all in `src/web/public/terminal-ui.js` unless noted:
1. **Desktop wheel** (container `wheel` listener, ~line 421): ALWAYS `preventDefault()`s, then either
- forwards synthetic SGR wheel reports to the app (`_sendSyntheticSgrWheel`, coalesced every 40ms, fire-and-forget) when `_shouldForwardWheelToApp(ev)` (~line 2823) passes: no Shift held, opt-out setting `terminalWheelLocalScrollback` off, xterm `mouseTrackingMode === 'none'`, session mode is `claude` with `cliVersion >= 2.1.187` or `codex`, and viewport is at bottom;
- otherwise scrolls xterm's LOCAL scrollback via `terminal.scrollLines(lines)`.
- `lines` comes from `_wheelScrollLines(ev)` (~line 2818): `delta / 25`, i.e. it assumes PIXEL deltas.
- NOTE: xterm.js's own internal wheel handler sits on an element INSIDE the container, so it runs FIRST (bubble order) and is not suppressed by the container's `preventDefault`.
2. **Touch** (touchstart/move/end, ~lines 441-585): converts touch deltas to `terminal.scrollLines()` with momentum. Touch is ALWAYS local-scrollback, never forwarded to the app. Tap-to-position (touchend, ~line 533) is separate and already handles both mouse-tracking-on and server-strip cases.
3. **Server-side strip** (`_handleTerminalOutput`, `src/session.ts:1384`): for modes in `isAltScreenStripMode()` (`src/session.ts:179` = `codex | claude | gemini`), strips alt-screen switches (`?47/?1047/?1049`), scrollback erase (`3J`), and mouse-tracking DECSETs (`?1000-?1007` except `?1004` focus) so content stays in xterm's normal buffer with scrollback intact. Includes a chunk-boundary carry so split sequences can't leak. `shell` and `opencode` (and `antigravity`) are deliberately EXCLUDED: arbitrary shell programs (vim/less/htop) legitimately need the alt screen. There is a parity copy of this strip on the replay path (`src/web/routes/session-routes.ts`, ~line 1697) and a frontend parity check `_sessionUsesServerMouseStrip()` (terminal-ui.js ~line 2751). All three must stay in sync.
4. Related: full-scrollback replay (`GET .../terminal?full=1` on first buffer load) fills xterm local scrollback; client scrollback is hardcoded 50k (`DEFAULT_SCROLLBACK`, constants.js) vs tmux 100k.
## Diagnosis
### Bug A: Firefox wheel deltas (mtiller's desktop case)
`_wheelScrollLines()` divides by 25 assuming `WheelEvent.deltaY` is pixels (`deltaMode === 0`, Chrome/Safari behavior). Firefox commonly fires `deltaMode === 1` (LINE units, deltaY around 1-3 per notch), so `Math.round(3/25) = 0` and the `|| ±1` fallback yields 1 line per event. With a discrete mouse wheel that is 1 line per notch: scrolling feels dead/broken. This hits BOTH the local-scroll path and the forwarded path, since both use the same function.
**Fix**: normalize by `ev.deltaMode` in `_wheelScrollLines()`:
- `deltaMode 0` (pixels): current behavior, `delta / 25`.
- `deltaMode 1` (lines): use the delta directly (round, keep sign fallback).
- `deltaMode 2` (pages): `delta * terminal.rows` (or a sane page size).
Keep the existing Shift-axis trap intact: on macOS trackpads Shift+two-finger scroll arrives as a HORIZONTAL wheel (deltaX carries the magnitude, deltaY ~0); that's why the function reads deltaX when Shift is held (issue #154). Don't lose it.
**Verify**: don't trust this diagnosis blindly. First reproduce in real Firefox on macOS and log `deltaMode`/`deltaY` (Firefox trackpad input can arrive as pixels; external mouse as lines). Also confirm the session's `cliVersion` probe succeeded (a failed probe disables forwarding entirely, which would point elsewhere). Unit-test by dispatching synthetic `WheelEvent`s with explicit `deltaMode` values; a Playwright `firefox` project pass is the end-to-end check.
### Bug B: shell mode has NO working scrollback at all (jonocodes)
Chain: shell mode is excluded from the alt-screen strip (correctly) → tmux attaches on the alternate screen → xterm's alt buffer has zero scrollback. Consequences:
- **Wheel**: xterm's own internal wheel handler runs first and, in the alt buffer, converts wheel ticks into Up/Down arrow keys (alternateScroll behavior). The shell receives arrows → command history cycles. That is jonocodes' exact desktop symptom. The container handler's `scrollLines()` afterwards is a no-op (no scrollback in alt buffer).
- **Touch**: the touch handler's `scrollLines()` is equally a no-op → "scrolling does nothing" on Android. Exact symptom two.
- The real history exists the whole time in tmux's 100k-line buffer; nothing exposes it.
**Fix, recommended approach (a): enable tmux `mouse on` for shell sessions.**
- Server-side, set `mouse on` scoped to shell sessions' tmux sessions (`tmux set-option -t <session> mouse on` at create + on attach of recovered sessions). Do NOT set it globally on the socket: claude/codex/gemini sessions rely on the DECSET strip and must not change.
- What this buys, all natively: tmux enables mouse tracking on the outer terminal → xterm `mouseTrackingMode` goes non-none → the container handler stands down (line ~2830 check) and xterm's own encoder forwards wheel as SGR reports → tmux scrolls its OWN copy-mode history on wheel-up, auto-exits at bottom. The alt-scroll arrow conversion disappears too (tracking mode takes precedence). Desktop is fully fixed with no new endpoints.
- **Touch**: still needs one small client change: in the touchmove path, when the active session is `shell` AND `mouseTrackingMode !== 'none'`, convert accumulated lines to `_sendSyntheticSgrWheel(x, y, lines)` instead of `scrollLines()`. The 40ms coalescing already prevents the tmux process storm (each send is a tmux send-keys server-side; unbatched flicks would spawn dozens of processes: this constraint is documented at `_sendSyntheticSgrWheel`, do not bypass it).
- **Selection tradeoff to verify**: with tracking on, xterm hands drag events to tmux instead of doing local browser selection. Shift+drag still does local selection (xterm shift-override). Verify this UX on desktop before shipping; if it's unacceptable, fall back to approach (b).
- **Also verify**: vim/less/htop inside the shell still behave (they'll now receive real mouse events via tmux, generally an improvement); remote shell sessions run tmux on the REMOTE host (`tmux -L codeman-remote`) and need the same option set there if remote shells are in scope (fine to defer, note it in the changeset if skipped).
**Fallback approach (b), only if (a)'s selection tradeoff fails testing**: keep mouse off; when a shell session is in the alt buffer, have the client send scroll intents to a small server endpoint that drives `tmux copy-mode -e -t <pane>` + `send-keys -X -N <n> scroll-up/down`. Preserves selection semantics exactly, but needs a new endpoint, server-side batching, AND suppression of xterm's native alt-scroll arrow conversion (capture-phase wheel listener with `stopPropagation`, or `attachCustomWheelEventHandler` if the vendored xterm version has it). More moving parts; (a) should be tried first.
**Not acceptable**: adding `shell` to `isAltScreenStripMode()`. vim/less/htop need the alt screen; that exclusion is deliberate and documented.
### Bug C: mtiller's phone/iPad case — UNREPRODUCED, do not guess
Touch is always-local by design, and Claude sessions keep content in the normal buffer (strip), so touch scrollback "should" work there. Before coding anything: build a repro matrix (iPhone Safari / iPad Safari / Android Chrome × claude / shell) on the current release. Plausible candidates if it does reproduce: auto-scroll-to-bottom fighting user scrolls (`_noteTerminalUserScroll`, ~line 2004), or they were in shell sessions on mobile too (then Bug B covers it). Ask mtiller on #205 for session mode + Codeman version if the matrix comes up clean.
## Invariants the implementation MUST respect
- Shift+wheel always scrolls local scrollback; the trackpad Shift-axis handling from #154 stays.
- The `terminalWheelLocalScrollback` opt-out setting keeps working (pins plain wheel to local).
- The viewport-at-bottom gate stays: once the user scrolled up locally, wheel stays local until they return to bottom.
- 40ms SGR coalescing: never send per-event writes to the server.
- Strip parity triangle: `session.ts` live strip ↔ `session-routes.ts` replay strip ↔ `_sessionUsesServerMouseStrip()` in the frontend. If you touch mode lists, update all three.
- Don't add `opencode`/`antigravity` to any strip/forward list; their TUI wheel behavior is unverified (documented at `_shouldForwardWheelToApp`).
- The chunk-boundary sequence carry in `_handleTerminalOutput` must not be weakened.
## Testing (per repo rules)
- `npm test -- test/<file>.test.ts` only; never bare `npm test`. New test ports 3150+, never 3000.
- Browser-test traps (documented in CLAUDE.md Testing): drive input/scroll through real events (`page.mouse.wheel`, real touch), not app internals; headless Chromium reports `isTouchDevice()` false even with `hasTouch: true`; assert on real state (xterm viewport position, `tmux -L codeman capture-pane`), not HTTP 200.
- Shell-mode E2E: create a throwaway shell session, `seq 1 500`, then (1) wheel up on desktop shows earlier lines, not history cycling; (2) touch-scroll on a phone shows earlier lines; (3) `vim` + `less` still enter/leave the alt screen cleanly; (4) Shift+drag still selects text.
- Firefox E2E: Playwright `firefox` project, wheel over a Claude session's finished output, assert viewport moved more than 1 line per notch.
- End-to-end against the REAL environment before claiming done (standing user rule). w1/w2/w3 tmux sessions are the user's live sessions: never send input to them; create your own throwaway session and DELETE it by exact id when done.
## Related observation (not a reported bug, worth a look while in there)
The `claude --version` probe that feeds the forwarding gate runs only for local and docker sessions (`src/session.ts:1490` gates `!this._remote`; docker handled at :1507). Remote Claude sessions therefore never get `cliVersion` and silently keep local-only wheel. Harmless (local scrollback works) but inconsistent; cheap to fix by probing over ssh, or document as intended.
## Rollout
1. Bug A (deltaMode) is small and independent: can ship alone as a patch.
2. Bug B (shell scrollback) is the headline fix for #205: patch or minor per COM flow.
3. After deploy + verification: comment on #205 (what was fixed, what needs their retest), then reply to the Reddit comment `p21x6ts` with the release version. Both reporters gave environment details; address them specifically.
+255
View File
@@ -0,0 +1,255 @@
# Scrollback issues: analysis and test evidence
Covers GitHub issue **#205** ("Scrollback in terminal not working", jonocodes, shell mode,
Android + macOS desktop) and the follow-up comment on it from **mtiller** (Firefox on macOS,
"scrolling backward to see agent output"). Related closed issue: **#154** (fixed in 1.3.3).
Status: **analysis only, nothing implemented.** Measured against the live 1.11.2 instance on
2026-08-06 with throwaway `zz-*` shell sessions (all deleted afterwards; the user's `w*`
sessions were never touched).
---
## TL;DR
Five distinct problems, not one. #205 is fully explained by finding 1; findings 2 and 3 are
independent and hit **every** mode including Claude, and are the likely substance of the
"similar issue" follow-up.
| # | Problem | Modes affected | Severity | Confirmed |
| - | ------- | -------------- | -------- | --------- |
| 1 | xterm parked in the **alternate buffer** for the whole session, so there is no scrollback at all and the wheel is translated into Up/Down arrow keys | `shell`, `opencode`, `antigravity` | High | Reproduced end to end |
| 2 | **Bursty output silently destroys a screenful** of the browser's scrollback and adds ~1 row | all | High | Measured |
| 3 | **Tab switch collapses scrollback** to roughly one screen (`full=1` fires once per page load) | all | Medium | Measured |
| 4 | `deltaMode` is never read, so Firefox scrolls ~4x slower per notch | all, Firefox | Low | Static, needs reporter data |
| 5 | **Remote SSH Claude cases get no `claude --version` probe**, so wheel forwarding silently stays off (residual #154) | `claude` + remote | Medium | Static |
---
## Finding 1: shell / opencode / antigravity are stuck in xterm's alternate buffer
### Root cause
The local tmux **client** (the `tmux attach` that node-pty spawns) emits `smcup` as its very
first bytes on attach. Captured from a real PTY:
```
b'\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2J\x1b[?12l\x1b[?25h\x1b[?1000l...'
^^^^^^^^^^ enter alternate screen ^^^^^ application cursor keys ON
```
`Session._handleTerminalOutput()` strips `\x1b[?1049h` from the live stream, but only when
`isAltScreenStripMode(mode)` is true, and that is `claude | codex | gemini` only
(`src/session.ts:179`). For `shell`, `opencode` and `antigravity` the sequence reaches the
browser verbatim and xterm switches to the alternate buffer, where:
1. `buffer.active.type === 'alternate'` and `baseY` is pinned at 0, so there is **no
scrollback to reach**. `terminal.scrollLines()` is a no-op, which is why touch scrolling
on Android "does nothing".
2. xterm's own wheel listener takes over. From the vendored bundle
(`src/web/public/vendor/xterm.min.js`):
```js
if (!this.buffer.hasScrollback) {
if (ev.deltaY === 0) return false;
if (coreMouseService.consumeWheelEvent(...) === 0) return this.cancel(ev, true);
const seq = ESC + (decPrivateModes.applicationCursorKeys ? 'O' : '[') + (ev.deltaY < 0 ? 'A' : 'B');
coreService.triggerDataEvent(seq, true);
return this.cancel(ev, true);
}
```
tmux also set `\x1b[?1h`, so the emitted sequence is `\x1bOA`, i.e. **Up arrow**, straight
into the shell's readline. That is exactly the reported "the mouse wheel scrolls back
through previous commands, like pressing up".
3. `cancel(ev, true)` calls `preventDefault()` **and `stopPropagation()`**, and xterm's
listener sits on `terminal.element` (a child of Codeman's container). So Codeman's own
container wheel handler, `_shouldForwardWheelToApp` and `_wheelScrollLines` included, is
**never reached** for these modes. That whole path is dead code for shell.
### Reproduction (live instance, real browser)
Create a shell session with the page already open, print 150 lines, then dispatch 8 wheel-up
events over `.xterm-screen`:
```
t+1500 after shell start {"type":"alternate","length":35,"baseY":0}
t+3000 after shell start {"type":"alternate","length":35,"baseY":0}
after 150 live lines {"type":"alternate","length":35,"baseY":0}
WHEEL on live shell: {"ptyBytes":["OA","OA","OA","OA",
"OA","OA","OA","OA"],
"before":0,"after":0,"type":"alternate"}
```
Both reported symptoms, one root cause.
### Why it looks intermittent
The alternate-screen sequence only ever reaches the browser through the **live stream at
attach**. Neither replay path carries it:
- `?full=1` returns `capture-pane` output (`source: mux-full-history`), verified 0 hits for
`\x1b[?1049h`.
- `?tail=` returns the visible pane frame (`source: mux-visible`), also 0 hits; the shell byte
buffer was empty in every probe.
- `_resetTerminalForReplay()` calls `terminal.reset()`, which returns xterm to the normal
buffer.
So: watching a shell from creation leaves you in the alternate buffer until you reload or
switch tabs, at which point it silently starts working again. Then the next PTY attach (a
restart, or the auto-reattach in `selectSession()`) puts you back.
### Is stripping safe for shell? Probably yes when tmux-backed, and the current code comment is wrong about why
`src/session.ts:1404` says *"shell must keep the alt screen for vim/less/htop"*. For a
**tmux-backed** shell that reasoning does not hold: tmux is a full terminal emulator and never
forwards a pane's alternate-screen toggles to its client, it repaints instead. Measured per
phase on a real attach:
| phase | bytes | `?1049h` | `?1049l` | `?47/1047` |
| ----- | ----: | -------: | -------: | ---------: |
| attach | 772 | **1** | 0 | 0 |
| `seq 1 60` echo | 1402 | 0 | 0 | 0 |
| `less` open / end / quit | 284 / 230 / 321 | 0 | 0 | 0 |
| `vim` open / quit | 2200 / 646 | 0 | 0 | 0 |
`vim` and `less` inside tmux emit **zero** alternate-screen sequences to the client.
The caveat that does matter: `startShell()` falls back to a **direct PTY with no tmux** when
mux creation fails (`src/session.ts:1961`, `this._useMux = false`). In that path the inner
app's own `?1049h` does reach xterm, and a blanket strip would break vim/less/htop for real.
Any fix has to be conditional on `_useMux`, which is known server-side.
Second caveat: stripping alone buys less than it looks like, because of finding 2. It fixes
the wheel (no more phantom Up arrows) and it makes the `full=1` replay reachable, but live
output still will not accumulate.
---
## Finding 2: bursty output silently overwrites a screenful of browser scrollback
Independent of the alternate buffer, and it hits Claude sessions too.
tmux decides per flush whether to emit real linefeeds (which push rows into the outer
terminal's scrollback) or to repaint the pane rectangle with cursor addressing (which
overwrites the visible rows in place). When output outpaces its flush interval it coalesces
into a repaint, and one screenful of the browser's history is **destroyed**.
Measured on one session, same page, `rows = 36`:
| step | `baseY` | rows containing SEED | BURST | SLOW |
| ---- | ------: | -------------------: | ----: | ---: |
| after `?full=1` replay (120 seeded lines) | 86 | 120 | 0 | 0 |
| after 60 lines emitted as fast as possible | **87** (+1) | **86** (-34) | 35 | 0 |
| after 60 lines at ~16/s (`sleep 0.06`) | **148** (+61) | 86 | 35 | 60 |
The burst added **one** row of scrollback and ate **34** rows of existing history. The slow
run behaved correctly. So "I printed a bunch of lines and now I cannot scroll back" reproduces
without the alternate buffer being involved at all, and it is rate dependent, which is exactly
the kind of thing that reads as random flakiness.
Consequence: the browser's scrollback is effectively frozen at whatever the last `?full=1`
replay produced, minus a screen per burst. tmux's own history is fine throughout
(`history_size` kept growing, `history-limit` 2000), so the data is never actually lost
server-side, it just never reaches the browser again until a reload.
---
## Finding 3: switching tabs collapses a session's scrollback
`_initialFullBufferLoad` is true for the **first buffer load after a page load only**
(`app.js:4374`). Everything after that uses `?tail=`, which returns byte history plus the
visible pane frame. Worse, the snapshot restore path deliberately throws away the restored
xterm snapshot (which does carry scrollback) and replaces it with that frame
(`app.js:4316-4328` plus `needsRewrite`).
Measured, switching away from session A and back:
```
A: initial full=1 load {"len":152,"baseY":116,"AAA":150}
A: after switch away and back {"len": 87,"baseY": 51,"AAA": 59}
```
150 lines of history down to 59. Note also that the page's single `full=1` is consumed by
whichever session auto-selects at load, so **every other tab starts life with one frame of
history**.
---
## Finding 4: `deltaMode` is never read (Firefox)
`grep -rn "deltaMode" src/web/public packages` returns nothing. `_wheelScrollLines()`
(`terminal-ui.js:2818`) treats `deltaY` as pixels unconditionally:
```js
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
```
Chrome/WebKit report `deltaMode: 0` with `deltaY` around 100 to 120 px per notch, so about 4
to 5 lines. Firefox reports `deltaMode: 1` (`DOM_DELTA_LINE`) with `deltaY` around 3, so
`Math.round(3/25) === 0` and the `|| ±1` fallback yields **1 line per notch**, roughly 4x
slower. In Claude mode the same value caps the forwarded SGR report at 1 tick per event
instead of 4, so the transcript crawls too.
This is sluggishness, not breakage, so it is a plausible but unproven contributor to the
mtiller report. No Firefox build is installed under `~/.cache/ms-playwright` (chromium and
webkit only), so this was not measured. Worth asking the reporter for `deltaMode` / `deltaY`
from a live wheel event before acting on it.
---
## Finding 5: remote SSH Claude cases still have no version probe
`src/session.ts:1490` deliberately skips the deterministic `claude --version` probe for
remote sessions and defers to the startup-banner scrape, which the same comment block
describes as unreliable ("newer Claude Code builds don't print the banner and resumed sessions
never show it"). That is precisely the condition #154 was filed for: `cliVersion` empty means
`_shouldForwardWheelToApp()` returns false, wheel forwarding is off, and the user is left with
local scrollback that (per finding 2) does not accumulate.
Local and Docker Claude sessions are fine; verified all 7 live sessions report
`cliVersion=2.1.223`, so the 1.3.3 fix is still working there.
---
## Candidate directions (not decided)
Roughly in order of value per unit of risk.
1. **Extend the alternate-screen strip to tmux-backed `shell` / `opencode` / `antigravity`.**
Gate on `_useMux` so the direct-PTY fallback keeps vim/less/htop working. Kills the phantom
Up arrows and makes replayed history reachable. `isAltScreenStripMode()` currently takes
only `mode`, so it would need the mux flag threaded in, and
`test/claude-scrollback-strip.test.ts:16-17` plus `test/antigravity-mode.test.ts:116` pin
the current answers and would need updating.
2. **Re-pull `?full=1` when the user scrolls to the top of the buffer.** Directly addresses
findings 2 and 3 with machinery that already exists and is already proven to return
complete history (200/200 lines in the probe). Needs a guard against refetch storms.
3. **Stop discarding the xterm snapshot on tab switch**, or request `full=1` on the first load
per session rather than per page. Cheaper partial fix for finding 3 alone.
4. **Read `ev.deltaMode`** in `_wheelScrollLines()` and normalise line/page deltas to lines.
Small, self-contained, worth doing regardless of whether it is mtiller's actual bug.
5. **Probe the CLI version over SSH for remote Claude cases**, mirroring the deferred
in-container probe that Docker cases already use.
Option 1 alone does not fix #205's "print a bunch of lines then scroll" complaint; that needs
2 as well.
## Reproduction assets
Scripts used, in the session scratchpad
(`/tmp/claude-1000/-home-arkon-default-claudeman/597ffc9f-.../scratchpad/`):
- `ptycap.py` / `ptycap2.py`: PTY-level capture of the tmux client stream, per phase counts of
alternate-screen and mouse-tracking sequences.
- `sim.mjs`: replays a captured stream through `@xterm/headless` with and without the strip.
- `browser-test*.mjs`: Playwright against the live instance, reports `buffer.active.type`,
`baseY`, row content and the exact bytes xterm sends to the PTY on a wheel event.
`@xterm/headless` was installed with `npm i --no-save`, so `package.json` and the lockfile are
untouched.
+61 -8
View File
@@ -30,7 +30,8 @@ an explicit, guided opt‑in.
7. [Supply‑chain & build‑asset hardening](#7-supplychain--buildasset-hardening-cod28)
8. [Multi‑instance isolation](#8-multiinstance-isolation)
9. [Transport security headers](#9-transport-security-headers)
10. [Quick reference](#10-quick-reference)
10. [Docker container isolation](#10-docker-container-isolation)
11. [Quick reference](#11-quick-reference)
---
@@ -248,16 +249,28 @@ Ordered most‑to‑least recommended:
### A. Tailscale serve (recommended)
Bind loopback, let Tailscale front it on your tailnet with a real cert:
Bind loopback, let Tailscale front it on your tailnet with a real cert. **The
installer sets this up for you**: choose **Tailscale** at the network-access
prompt, or retrofit an existing install with:
```bash
codeman web --https # binds 127.0.0.1:3000
tailscale serve --bg https / http://127.0.0.1:3000
bash ~/.codeman/app/install.sh tailscale
```
Only devices on your tailnet can reach it; Tailscale handles identity. No app
password and no `0.0.0.0` bind required. (This is the maintainer's production
setup.)
The guided flow installs Tailscale if needed, walks through login and the
tailnet HTTPS-certificates toggle, and configures the equivalent of:
```bash
codeman web # binds 127.0.0.1:3000 (plain HTTP is fine here)
tailscale serve --bg 3000 # HTTPS at https://<node>.<tailnet>.ts.net
```
Only devices on your tailnet can reach it; Tailscale handles identity and
terminates TLS with a real Let's Encrypt certificate (so PWA install and web
push work). No app password and no `0.0.0.0` bind required. (This is the
maintainer's production setup.) `CODEMAN_TAILSCALE=1` presets the choice for
automation; the installer never runs `tailscale serve reset` and never touches
serve mappings other than `443 -> Codeman's port`.
### B. Authenticated cloudflared tunnel + password
@@ -471,7 +484,45 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
---
## 10. Quick reference
## 10. Docker container isolation
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — and `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
- **Instance isolation** — every managed container is labeled `codeman.instance=<CODEMAN_INSTANCE>`; the boot reaper reaps orphans of its OWN instance only, so a beta never removes a prod container. The in‑container tmux socket (`-L codeman-docker`) + session name (`codeman-dkr-*`) deliberately fail a nested Codeman's discovery pattern.
Full feature guide: [`docker-cases.md`](docker-cases.md).
---
## 10a. Multi‑user mode (opt‑in)
`codeman web --multiuser` (or `CODEMAN_MULTIUSER=1`) turns on named users with individually scrypt‑hashed passwords in `~/.codeman/users.json` (mode 0600). OFF by default; when off, nothing here applies and behavior is byte‑identical to single‑user. Design + phase status: [`multi-user-plan.md`](multi-user-plan.md).
- **It is workspace separation, NOT a security boundary between users.** Every session still runs as the SAME OS account with agent code that can read the whole host. Any user can ask their agent to `cat` another user's files; the WEB layer enforces scoping, the AGENT layer cannot. Mitigations: give non‑admins the default `auto` permission mode (classifier‑guarded), pair users with **Docker cases** (container per case) for real isolation, or run separate Codeman instances under separate OS accounts. Stated loudly in the admin panel and the plan's threat model (section 2).
- **It strictly improves network posture.** It removes the single shared `CODEMAN_PASSWORD` and gives each person a revocable credential; a non‑loopback bind and the tunnel‑enable guard are satisfied by "multi‑user with ≥1 enabled user" without a shared password.
- **Auth is a parallel branch** (`middleware/auth.ts`) that leaves the single‑user path untouched: per‑user scrypt verify (`timingSafeEqual`, timing‑equalized against user enumeration), identity‑carrying cookies, a per‑username failure bucket (a botnet can't brute one account across IPs; one NATed user can't lock out the rest), and a `mustChangePassword` lockbox. The hook‑secret loopback bypass, host guard, and Origin/CSRF guard are unchanged (hooks authenticate the INSTANCE, not a user).
- **Ownership is enforced server‑side only** and fails closed: `req.authUser` (a synthetic admin in single‑user), `findSessionOrFail` returns NOT_FOUND (never 403) for a foreign session, list/SSE/WS/file‑preview/search all filter by `session.owner`, and SSE routing defaults session‑scoped events to their owner (unresolved owner → withheld). The load‑bearing rule is **non‑admin `workingDir` confinement**: a non‑admin's session/one‑shot working dir must realpath‑resolve inside `~/codeman-users/<name>/cases`, checked BEFORE any disk write.
- **Privileged actions are a one‑bit grant** (`canBypassPermissions`, default off): only granted users (and admins) get `--dangerously-skip-permissions` (others are silently downgraded to `--permission-mode auto`), shell‑mode sessions, cron `launchCommand`, and other CLIs' bypass flags. Machine‑level resources (remote/Docker host definitions, tunnel, self‑update, settings writes) are admin‑only.
- **Admin actions are audited** append‑only to `~/.codeman/admin-audit.jsonl` (acting admin, action, target, IP). Passwords set by an admin create/reset are one‑time (returned once, force change). Under Basic auth, `logout` only truly ends QR‑issued sessions — to lock someone out, disable the account or reset the password (a proper login form is a deferred Phase 6).
---
## 10b. Web tabs (dashboard proxy)
A saved dashboard URL renders as a tab, served through Codeman's own origin at `/webview/<capability>/`. User guide: [`web-tabs.md`](web-tabs.md). Three properties carry the security weight:
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from).
---
## 11. Quick reference
| Env / flag | Effect |
|------------|--------|
@@ -482,6 +533,8 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
| `--https` | Enable TLS (adds HSTS) |
| `CODEMAN_INSTANCE` | Scope tmux socket + data dir for isolation |
| `CODEMAN_GESTURE=1` | Make the gesture overlay available (widens CSP) |
| `CODEMAN_DOCKER_BRIDGE_HOOKS=1` | Serve the hook endpoints on the docker bridge gateway (host‑internal, hooks‑only, `403` elsewhere) so in‑container hooks reach a loopback‑bound server — see §10 |
| `CODEMAN_DOCKER_BRIDGE_HOST` | Override the bridge gateway IP the hooks listener binds (default: auto‑detect) |
**Audit log:** session lifecycle and server start are recorded in
`~/.codeman/session-lifecycle.jsonl`.
+236
View File
@@ -0,0 +1,236 @@
# Tailscale Setup in the Installer (Plan)
Goal: make "Codeman over Tailscale, with real HTTPS" a first-class, guided path in
`install.sh`, instead of a one-line hint pointing at the docs. Today the safest
recommended deployment (loopback bind + `tailscale serve`) is exactly what the
maintainer's own prod runs, but a new user has to discover and wire it by hand.
The installer should do it for them.
Status: IMPLEMENTED (2026-08-04). `install.sh` carries the 3-way network
prompt, the guided Tailscale flow, and the `tailscale` subcommand; README,
`docs/security-architecture.md` section A, and CLAUDE.md are updated. Verified
live on the maintainer's prod host: `install.sh tailscale` took the idempotent
kept-as-is path against the existing serve mapping (recognizing the legacy
`https+insecure://` target), verified `https://<node>.ts.net/api/status`
end-to-end, and left `tailscale serve status` byte-identical. Items 1-4, 7,
and 10-12 of the manual matrix below still need a fresh machine to exercise.
## Why this is low-hanging fruit
Everything on the app side already works; this is almost purely installer UX:
- `.ts.net` is already in `DEFAULT_TRUSTED_HOST_SUFFIXES`
(`src/web/network-auth-policy.ts`), so the always-on Host/Origin guard accepts
`tailscale serve` traffic with zero configuration. No `CODEMAN_ALLOWED_HOSTS`
needed.
- The loopback bind is the server default and prints no warning; nothing to
acknowledge, no `CODEMAN_PASSWORD` strictly required (the tailnet is the auth
boundary; Tailscale authenticates the device before a packet ever reaches us).
- `tailscale serve` terminates TLS with a real Let's Encrypt certificate for
`<node>.<tailnet>.ts.net`. That gives users valid HTTPS with no self-signed
cert warnings, and (because it is a proper secure context) working service
worker, PWA install, and web push on phones. This is strictly better than
`codeman web --https` for remote access.
- SSE and WebSockets work through serve (proven by prod:
`https://tnode.tailf80371.ts.net` fronting `127.0.0.1:3000` daily).
- `docs/security-architecture.md` section "A. Tailscale serve (recommended)"
already documents this as the preferred setup; the installer just does not
implement it.
## UX design
### 1. The network-access prompt grows a Tailscale option
`choose_network_binding()` (install.sh:1051) currently offers two choices. New
menu, with Tailscale first when it can be recommended:
```
Network access
How should the Codeman dashboard be reachable?
1) Tailscale (recommended)
Private VPN access from your phone/laptop, real HTTPS,
no password needed. Works from anywhere, not just your Wi-Fi.
2) Any device on your network (0.0.0.0)
Open it straight from your phone or laptop on the same Wi-Fi.
Less safe: set a password so only you control your agents.
3) This machine only (127.0.0.1)
Safest. Reach it remotely via Tailscale or a tunnel later.
```
Choice mapping:
- Option 1 = bind `127.0.0.1` (unchanged server posture) + configure
`tailscale serve`. Internally it is option 3 plus the serve setup, so all
existing binding plumbing (`BIND_HOST`, service files, `read_existing_binding`)
is untouched.
- Options 2 and 3 behave exactly as today (renumbered).
- Default choice: 1 when tailscale is installed and logged in, or when an
existing serve mapping for our port is detected; otherwise keep today's
defaults (1 -> 2, 2 -> 3 renumbering, preserving the "existing setup wins"
rule). If tailscale is not installed, option 1 is still shown (the installer
offers to install it), but the default stays on the current behavior so a
bare Enter never pulls in new software.
- Password: after choosing Tailscale, offer the password prompt as optional
defense in depth with default skip ("the tailnet already authenticates your
devices; add one anyway?"). No `BIND_ACK` needed since the bind is loopback.
### 2. The Tailscale flow (state machine)
New `setup_tailscale_access()` runs after the binding choice, before service
setup, handling each state in order:
1. **Not installed.**
- Linux: offer to run the official installer
(`curl -fsSL https://tailscale.com/install.sh | sh`), which handles all
distros and enables `tailscaled` at boot. This mirrors our own
curl-pipe-bash story and avoids maintaining per-distro logic like the six
`install_cloudflared_*` functions.
- macOS: do not auto-install (the GUI app needs an interactive login).
Offer `brew install --cask tailscale` when brew exists, else print the
download link, then wait-and-retry or let the user skip.
- Declined install => fall back to plain loopback (option 3 behavior) and
print how to redo this later (`install.sh tailscale`, see below).
2. **Installed but logged out** (`tailscale status --json` ->
`.BackendState == "NeedsLogin"` or `"Stopped"`).
- Run `tailscale up` (via `run_as_root` if needed). It prints an auth URL
that works headless (user opens it on any device). Poll
`.BackendState == "Running"` with a friendly spinner + timeout; on
timeout, skip gracefully with re-run instructions.
3. **Running: grant operator (Linux).** `sudo tailscale set --operator=$USER`
so serve configuration (now and in the future) does not need root. Skip
silently if we are already operator (probe: `tailscale serve status`
exits 0) or sudo is declined; fall back to `run_as_root tailscale serve ...`.
4. **HTTPS availability check.** `.CertDomains` empty or
`.CurrentTailnet.MagicDNSEnabled == false` means the tailnet has not enabled
MagicDNS / HTTPS certificates. Print the exact two toggles with the admin
URL (https://login.tailscale.com/admin/dns: enable MagicDNS, then enable
HTTPS Certificates), then offer "I enabled it, re-check" / "skip for now".
No silent HTTP fallback: the pitch is real HTTPS, and a plain-HTTP serve
would break the PWA/push story. Skipping falls back to loopback + re-run
instructions.
5. **Existing serve config check** (`tailscale serve status --json`).
- Already proxying to our port (443 -> `127.0.0.1:$PORT`): keep it, report
it, done. Re-running the installer must be idempotent.
- Port 443 occupied by a DIFFERENT target: never clobber it. Ask whether to
replace it or skip. (Prod itself has a second serve on :5000; blind
`tailscale serve reset` would destroy user config. NEVER use `reset`.)
6. **Configure.** `tailscale serve --bg $PORT` where `$PORT` is the install's
Codeman port (default 3000; honor a preset `CODEMAN_PORT`). Serve targets
plain HTTP on loopback; TLS terminates at tailscaled with the real cert.
The `--bg` config persists in tailscaled state across reboots, so no extra
service unit is needed.
(Note: do NOT combine this with `codeman web --https`; that is what forces
the awkward `https+insecure://` proxy target prod historically used. New
installs should keep Codeman on plain HTTP behind serve.)
7. **Verify end-to-end.** Derive the URL from `.Self.DNSName` (strip the
trailing dot) and curl `https://<dnsname>/api/status` after the service is
up, retrying for ~30s: the first request can be slow while the Let's
Encrypt cert is issued. Print success with the URL, or the observed error
with `tailscale serve status` output on failure. This follows the "always
test before claiming it works" rule; a blind "done!" is not acceptable.
### 3. Closing summary and security notice
- The final summary gains a "Remote Access (Tailscale)" block, printed above
the cloudflared block, showing the actual URL:
```
Remote Access (Tailscale):
https://tnode.tailf80371.ts.net (any device on your tailnet, HTTPS)
tailscale serve status # inspect
```
- `print_security_notice()` third branch (loopback) gets a variant: when a
serve mapping for our port is detected, lead with "reachable on your tailnet
at https://... (HTTPS, tailnet-only)" instead of the generic "do ONE of"
list. Detection is dynamic (query `tailscale serve status --json` at print
time), no marker persisted anywhere: tailscaled's own state is the single
source of truth, so external changes never drift against a stale flag.
### 4. Standalone entry point: `install.sh tailscale`
Add a `tailscale` subcommand next to `update` / `uninstall` in the existing
dispatch. It runs `setup_tailscale_access()` against the already-installed
service (reads the port from the service file, requires an existing install).
This serves:
- existing installs that predate the feature,
- users who picked "this machine only" and changed their mind,
- every "skip for now" branch above, all of which print this exact command.
One implementation, two entry points. No separate `scripts/tailscale-setup.sh`
(unlike cloudflared, there is no long-running process for a `tunnel.sh`-style
start/stop wrapper to manage; tailscaled owns the lifecycle).
### 5. Non-interactive / automation
- `CODEMAN_TAILSCALE=1` presets choice 1 (analogous to presetting
`CODEMAN_HOST`). In non-interactive runs it only proceeds through states
that need no human (already installed + logged in + HTTPS-enabled tailnet);
anything requiring interaction (login URL, admin-console toggle, replacing a
foreign serve mapping) warns and falls back to loopback. It never installs
tailscale non-interactively.
- `CODEMAN_NONINTERACTIVE=1` with an existing serve mapping: preserve it, same
"never silently loosen/change" policy as `read_existing_binding`.
- Document both in the header comment block of install.sh (the env-var
reference at the top) and in the README.
## Edge cases and decisions
| Case | Decision |
| ---- | -------- |
| macOS GUI app without `tailscale` on PATH | `get_tailscale_path()` helper mirroring `get_cloudflared_path()`: check PATH, then `/Applications/Tailscale.app/Contents/MacOS/Tailscale`. All calls go through it. |
| Tailnet HTTPS certs disabled | Guided admin-console instructions + re-check loop; skip falls back to loopback. Never configure plain-HTTP serve. |
| Port 443 serve exists for another app | Prompt replace/skip; never `tailscale serve reset` (destroys unrelated mappings). |
| First cert issuance latency | Verify step retries ~30s and says why the first load may be slow. |
| `tailscale up` needs auth | Print the auth URL prominently, poll with timeout, skip gracefully. Works headless. |
| Custom `CODEMAN_PORT` | Serve target uses the actual port; `install.sh tailscale` re-reads it from the service file. |
| Funnel (public internet) | OUT OF SCOPE for v1. If ever added it must mirror the tunnel guard: refuse without `CODEMAN_PASSWORD` (`isUnauthenticatedNetworkAcknowledged`). Funnel exposes to the whole internet and is a different risk class than tailnet-only serve. Mention `tailscale funnel` in docs only, with the password warning. |
| Uninstall | Best effort: if `serve status --json` shows 443 proxying to our port, run the targeted `tailscale serve --https=443 off` (still accepted by current CLIs); if the CLI rejects it, print manual instructions. Never touch other mappings, never uninstall tailscale itself. |
| User already fronting Codeman some other way (reverse proxy etc.) | The serve check only looks at tailscale state; other proxies are invisible and unaffected (same stance as the loopback-exemption note in security-architecture). |
## What does NOT change
- Server code: no changes required. Host guard already trusts `.ts.net`,
loopback bind is already the default, SSE/WS already work through serve.
- The two existing binding options and their semantics, `read_existing_binding`
preservation, and the LAN+password flow.
- `scripts/tunnel.sh` / cloudflared support (stays as the "no Tailscale
account" alternative).
- The security model: this feature only ever narrows exposure (loopback +
authenticated overlay), never widens it.
## Files touched (implementation inventory)
| File | Change |
| ---- | ------ |
| `install.sh` | New: `check_tailscale`, `get_tailscale_path`, `tailscale_status_field` (jq-free JSON field extraction; the installer cannot assume jq: use `sed`/`grep` like existing helpers or `tailscale status --json` piped to `node -e` since node is guaranteed post-install), `offer_install_tailscale`, `ensure_tailscale_login`, `ensure_tailscale_operator`, `ensure_tailnet_https`, `setup_tailscale_serve`, `verify_tailscale_access`, `setup_tailscale_access` (orchestrator). Modified: `choose_network_binding` (3-way menu), summary block, `print_security_notice`, subcommand dispatch (`tailscale`), `uninstall` (targeted serve removal), header env-var docs (`CODEMAN_TAILSCALE`). |
| `README.md` | Remote-access section: promote the Tailscale path with the one-liner and `install.sh tailscale`; keep the tailscale-IP HTTP note for non-serve users but recommend serve + HTTPS. |
| `docs/security-architecture.md` | Section A gains "the installer can set this up for you" + `install.sh tailscale` pointer. |
| `CLAUDE.md` | One line in Scripts & Tunnel: installer offers Tailscale setup (`install.sh tailscale` to redo). |
| `test/` | No unit tests possible for interactive bash + a live tailnet; guard with `shellcheck install.sh` (already the norm) and the manual matrix below. |
## Manual test matrix (before release)
1. Linux + tailscale absent: install offered, declined => loopback fallback + hint.
2. Linux + tailscale absent: install accepted => full flow => URL verified.
3. Logged out => auth URL flow => Running => serve configured.
4. Tailnet with HTTPS certs disabled => guided instructions => re-check => success; and the skip branch.
5. Re-run installer with serve already configured => idempotent, preserved, reported.
6. Second serve mapping on another port present => untouched (prod-like state).
7. Port 443 already proxying another target => replace/skip prompt honored.
8. `install.sh tailscale` on an existing loopback install (the retrofit path).
9. `CODEMAN_NONINTERACTIVE=1` re-run => preserves everything, no prompts.
10. macOS (Mac mini `arbbot` box): GUI-app CLI path detection + full flow.
11. Uninstall removes only our 443 mapping, leaves others.
12. Phone check: PWA install + push from the `https://*.ts.net` origin.
## Release
Changeset: `minor` (new documented installer capability + new `CODEMAN_TAILSCALE`
env var). The feature is installer-only, so it ships with zero risk to running
servers; `install.sh update` does not invoke the new flow (updates never rewrite
access config), only fresh installs and the explicit `install.sh tailscale`
subcommand do.
+10 -28
View File
@@ -43,38 +43,22 @@ const syncData = DEC_SYNC_START + data + DEC_SYNC_END;
this.broadcast('session:terminal', { id: sessionId, data: syncData });
```
## Client-Side Implementation (`app.js`)
## Client-Side Implementation (`terminal-ui.js`)
### `batchTerminalWrite(data)`
1. Checks if flicker filter is enabled (optional, per-session)
2. If flicker filter active: buffers screen-clear patterns (`ESC[2J`, `ESC[H ESC[J`, `ESC[nA`)
3. Accumulates data in `pendingWrites`
4. Schedules `requestAnimationFrame` if not already scheduled
5. On rAF callback: checks for incomplete sync blocks (start without end)
6. If incomplete: waits up to 50ms via `syncWaitTimeout`
7. Calls `flushPendingWrites()` when complete
### `extractSyncSegments(data)`
- Parses DEC 2026 markers, returns array of content segments
- Content before sync blocks returned as-is
- Content inside sync blocks returned without markers
- Incomplete blocks (start without end) returned with marker for next chunk
4. Calls `_scheduleTerminalWriteFlush()` if no flush is pending
5. The yielded callback clears its scheduled flag before calling `flushPendingWrites()`
6. Large batches schedule their own next chunk until the queue is empty
### `flushPendingWrites()`
```javascript
const segments = extractSyncSegments(this.pendingWrites);
this.pendingWrites = ''; // Clear before writing
for (const segment of segments) {
if (segment && !segment.startsWith(DEC_SYNC_START)) {
terminal.write(segment); // Skip incomplete blocks (start with marker)
}
}
```
Note: Segments starting with `DEC_SYNC_START` are incomplete blocks awaiting more data. These are skipped (discarded if timeout forces flush).
- Joins the queued terminal data and passes DEC 2026 markers through to xterm.js 6, which handles synchronized output natively.
- Writes at most 32KB per yield for Codex and 64KB for other modes.
- Requeues the remainder and immediately schedules another safe yield. A final large response therefore drains without waiting for another SSE event.
### `chunkedTerminalWrite(buffer, chunkSize=128KB)`
@@ -116,17 +100,15 @@ When detected, buffers 50ms of subsequent output before flushing atomically.
## Edge Cases
- **Incomplete sync blocks**: 50ms timeout forces flush (content discarded to prevent freeze)
- **Incomplete sync blocks**: xterm.js retains synchronized output until its closing marker
- **Large buffers**: Chunked writing prevents UI freeze
- **Server shutdown**: Skips batching via `_isStopping` flag
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
- **SSE reconnect**: `handleInit()` clears all pending write state
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
## DEC Mode 2026 Compatibility
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
Terminals that natively support DEC 2026 buffer and render atomically. Codeman uses xterm.js 6, so the client passes the markers through instead of parsing or discarding partial blocks.
**Supporting terminals:** WezTerm, Kitty, Ghostty, iTerm2 3.5+, Windows Terminal, VSCode terminal
@@ -135,4 +117,4 @@ Terminals that natively support DEC 2026 will buffer and render atomically. Term
| File | Key Functions |
|------|---------------|
| `src/web/server.ts` | `batchTerminalData()`, `flushTerminalBatches()`, `broadcast()` |
| `src/web/public/app.js` | `batchTerminalWrite()`, `extractSyncSegments()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
| `src/web/public/terminal-ui.js` | `batchTerminalWrite()`, `_scheduleTerminalWriteFlush()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
+303
View File
@@ -0,0 +1,303 @@
# Terminal smart copy (Ctrl+C) plan
Issue: [#211](https://github.com/Ark0N/Codeman/issues/211) "Terminal: Ctrl+C should copy when text is selected (interrupt otherwise)".
Origin: r/selfhosted feedback, "Biggest stumbling block is apparent lack of copy-paste in the terminal."
Status: **implemented and shipped** on 2026-08-05 (this document is kept as the rationale record). It was first served as an isolated beta over Tailscale for manual sign-off, then landed. Section 2 is the research that shaped the design, sections 4 to 6 describe what was built.
---
## 1. What the issue asks for
- Text selected in the terminal + `Ctrl+C` -> copy the selection, toast, clear the selection, do NOT send the byte to the PTY.
- No selection + `Ctrl+C` -> unchanged, the interrupt (`0x03`) reaches the PTY.
- `Ctrl+Shift+C` as an explicit copy chord.
- The selection check must run before the shortcut registry dispatch so a rebind cannot cost the user their interrupt key.
- Paste is out of scope (it already works via `Ctrl+V`, which terminal-ui.js routes to the image/text paste trap).
## 2. Verified current behavior
### 2.1 xterm cancels the Ctrl+C keydown, so no copy can happen
`src/web/public/vendor/xterm.min.js` (xterm 6.x), `_keyDown`:
```js
_keyDown(x){ if(this._keyDownHandled=!1, this._keyDownSeen=!0,
this._customKeyEventHandler && this._customKeyEventHandler(x)===!1) return !1;
... evaluateKeyboardEvent(...) ... this.cancel(x) ... }
```
Two consequences that shape the design:
1. The custom handler runs **first**, before xterm evaluates the key. Returning `false` exits before `cancel(x)`, so returning `false` does **not** call `preventDefault()` for us.
2. When the handler returns `true`, xterm turns Ctrl+C into `0x03` and cancels the event, which is why the browser's own copy command never runs.
Probe (headless chromium against an isolated server on port 3174, selection active, real focus on `.xterm-helper-textarea`, synthetic Ctrl+C keydown):
```json
{ "hasSelection": true, "defaultPrevented": true, "dataSeen": ["\"\\u0003\""],
"clipboardAfter": "SENTINEL-BEFORE", "stillHasSelection": false }
```
So today: interrupt byte sent, clipboard untouched, and xterm drops the selection anyway. The last point matters, "copy then clear the selection" is not a behavior change in how the selection feels, it is what already happens on any keypress.
### 2.2 Why right-click Copy works today
xterm registers a `copy` listener on its root element that substitutes the selection text:
```js
this._register(addDisposableListener(this.element,"copy",(k=>{ this.hasSelection() && copyHandler(k,this._selectionService) })))
```
Second probe (port 3175, real `page.keyboard.press('Control+c')`, custom handler patched to return `false` for Ctrl+C without `preventDefault`):
```json
{ "dataSeen": [], "copyEvents": ["xterm-element"],
"clipboardAfter": "native-copy-probe-line\n...", "stillHasSelection": true }
```
So a "return false and let the browser copy" implementation would also work in Chromium. It is rejected below (section 3.3) because it gives no toast, does not clear the selection, and leans on per-browser behavior of the copy command when the focused element is xterm's empty helper textarea.
### 2.3 The document-level capture handler will not interfere
`setupEventListeners()` in `src/web/public/app.js:989` runs on document capture, before xterm's textarea listener. Its registry loop skips any entry whose action is not in the local `SHORTCUT_ACTIONS` map:
```js
if (shortcut.disabled || !shortcut.action) continue;
const action = SHORTCUT_ACTIONS[shortcut.action];
if (!action) continue;
```
This is exactly how `command-palette` already behaves: it is a full registry entry (rebindable and disableable in App Settings) whose dispatch happens in a dedicated, focus-aware gate rather than the generic loop. The new copy entry follows that pattern, so the capture handler falls through untouched and the terminal handler owns the decision.
### 2.4 Registry matching rules that constrain the bindings
`matchesShortcutEvent()` (`app.js:4890`):
- Ctrl and Cmd are interchangeable as the primary modifier, so a `['ctrl']` binding also matches Cmd+C on macOS. That is fine here: with a selection it copies (same result the native macOS path gives today), without one it falls through.
- Every other modifier must be declared exactly: `if (mods.includes('shift') !== !!e.shiftKey) return false`. So `Ctrl+Shift+C` needs its own binding, a plain `ctrl+c` binding will never swallow it.
- `binding.code` wins when present, otherwise `binding.key` is compared case-insensitively.
### 2.5 Where selection is actually possible
- The server strips mouse-tracking DECSETs for `claude`, `codex`, and `gemini` (`isAltScreenStripMode`, `src/session.ts:179`), which is why plain drag-select works in those tabs even though the TUI has mouse tracking on.
- `shell`, `opencode`, and `antigravity` keep mouse reporting, so xterm requires `Shift`+drag to force a selection there. Worth one line in the docs, it is not a code change.
- Touch devices deliberately disable selection entirely (`body.touch-device .terminal-container .xterm{user-select:none !important}`, `styles.css:3196`), and phones have no Ctrl key. This feature is desktop and hardware-keyboard only, with no mobile regression surface.
### 2.6 Helpers that already exist and should be reused
| Need | Existing code |
| --- | --- |
| Clipboard write with an HTTP-safe fallback | `_copyText(text)` in `app.js:1887` (Clipboard API, then hidden textarea + `execCommand`) |
| Toast | `showToast(message, type)` in `panels-ui.js:4385` |
| Translated string | `'Copied to clipboard'` already in `i18n.js:453` |
| Focus-aware chord gate to copy the shape of | `shouldOpenCommandPaletteFromShortcut(e)` in `panels-ui.js:285` |
| Buffer-wide copy (currently unreferenced) | `copyTerminal()` in `terminal-ui.js:2615` |
`_copyText` matters more than it looks: `install.sh`'s LAN option serves plain HTTP, where `navigator.clipboard` is undefined. The issue's suggested `navigator.clipboard.writeText` alone would silently do nothing for those users, the `execCommand` fallback covers them.
## 3. Design
### 3.1 Behavior
| Chord | Selection present | No selection |
| --- | --- | --- |
| `Ctrl+C` (and Cmd+C, per registry equivalence) | copy, toast, clear selection, swallow the key | fall through, xterm sends `0x03` (interrupt) |
| `Ctrl+Shift+C` | copy, toast, clear selection, swallow the key | swallow, no-op (see 3.2) |
| Shortcut disabled in App Settings | never copies, `Ctrl+C` is always the interrupt | unchanged |
| Rebound to another chord | that chord copies when a selection exists | plain `Ctrl+C` is always the interrupt |
### 3.2 Why `Ctrl+Shift+C` with no selection is swallowed rather than forwarded
Today `Ctrl+Shift+C` produces `0x03` as well (the shift is irrelevant to the control byte), so forwarding would be "no regression". But once the chord is advertised as *the explicit copy key*, letting it interrupt a running agent when the selection happens to be empty is a footgun with no upside. Swallowing costs nothing: a user who wants to interrupt has `Ctrl+C` right there.
The rule in code is "no selection and the matched chord had Shift -> swallow", not a hardcoded key check, so it stays correct under rebinds.
### 3.3 Why an explicit clipboard write rather than falling through to the native copy
Probe 2 showed the native path works in Chromium, but the explicit write is chosen because it:
- gives the "Copied to clipboard" toast, which is the discoverability half of the issue,
- clears the selection so a second `Ctrl+C` interrupts (the smart-copy contract),
- works on plain-HTTP LAN installs through `_copyText`'s `execCommand` fallback,
- does not depend on how each browser treats a copy command issued while an empty textarea has focus.
### 3.4 Why no new app setting
Per-shortcut enable/disable and rebinding already exist in App Settings -> Shortcuts and are driven by the registry. A user who wants "Ctrl+C is always interrupt" unchecks one box. Adding a `terminalSmartCopy` setting would duplicate that and would drag in the per-device vs synced decision (`displayKeys` + `.strict()` `SettingsUpdateSchema`) for no gain.
## 4. Code changes, file by file
### 4.1 `src/web/public/app.js`, registry entry
Add to `DEFAULT_SHORTCUTS` (after the `clear-terminal` entry, ~line 351) so the Terminal group stays together:
```js
{
id: 'copy-selection',
group: 'Terminal',
label: 'Copy Selection',
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js,
// deliberately NOT in SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults, which would cost the user the interrupt key.
action: 'copyTerminalSelection',
},
```
Match on `key`, not `code`. xterm decides what byte to emit from the produced character, so intercepting the physical `KeyC` on a layout where it does not produce "c" would diverge from what xterm would have sent.
The `action` string is required for App Settings to render the row as configurable (`configurable = !!shortcut.action && Array.isArray(shortcut.bindings)`, `settings-ui.js:2624`). Do **not** add `copyTerminalSelection` to `SHORTCUT_ACTIONS`.
### 4.2 `src/web/public/terminal-ui.js`, the gate
New prototype method, modeled on `shouldOpenCommandPaletteFromShortcut`:
```js
shouldCopyTerminalSelectionFromShortcut(ev) {
if (!ev || ev.type !== 'keydown') return false; // the handler also runs for keypress/keyup
if (!ev.ctrlKey && !ev.metaKey && !ev.altKey) return false; // hot path: plain typing exits here
const registryAvailable =
typeof this.getShortcutRegistry === 'function' && typeof this.matchesShortcutEvent === 'function';
const entry = registryAvailable
? this.getShortcutRegistry().find((s) => s.id === 'copy-selection')
: null;
if (entry) return !entry.disabled && this.matchesShortcutEvent(ev, entry);
return (ev.key || '').toLowerCase() === 'c' && !ev.altKey; // fallback for isolated harnesses
}
```
### 4.3 `src/web/public/terminal-ui.js`, the branch
Inside `attachCustomKeyEventHandler` (`terminal-ui.js:133`), after the command-palette gate and before the `Ctrl+V` branch:
```js
// Smart copy (#211): with a selection, Ctrl+C copies instead of sending ^C.
// With no selection it MUST fall through (return true, no preventDefault) or
// the interrupt key is lost. Ctrl+Shift+C is the explicit chord and never
// falls through: an "explicit copy" that interrupts the agent is a footgun.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
}
return true;
}
```
`preventDefault()` is explicit because returning `false` alone does not cancel the event (section 2.1), and without it the browser would run its own copy on top of ours.
### 4.4 `src/web/public/terminal-ui.js`, the copy action
```js
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const ok = await this._copyText(selection);
if (ok) {
this.terminal.clearSelection?.();
this.showToast('Copied to clipboard', 'success');
} else {
this.showToast('Failed to copy', 'error');
}
// _copyText's execCommand fallback focuses a temp textarea; restore the
// terminal (this.terminal.focus is the CJK-aware router, not xterm's raw focus).
this.terminal.focus();
return ok;
}
```
The selection text is captured **before** the first `await`, and `navigator.clipboard.writeText` is reached in the same task as the keydown, so user activation still holds.
### 4.5 `src/web/public/i18n.js`
`'Copied to clipboard'` exists. Add `'Failed to copy': '复制失败'` (the error path is new to this surface).
### 4.6 Documentation
| File | Change |
| --- | --- |
| `README.md` shortcut table (~line 648) | `\| `Ctrl/Cmd+C` \| Copy selection (interrupts when nothing is selected) \|` and a `Ctrl+Shift+C` row |
| `src/web/public/index.html` help modal, Terminal section (~line 641) | `<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection / Interrupt</div>` plus the Ctrl+Shift+C row. Keep the existing negative assertion in `help-modal-shortcuts.test.ts` in mind (it forbids `Ctrl+K`, `C` is fine) |
| `CLAUDE.md` "Keyboard shortcuts" line | add `Ctrl+C` (copy selection, else interrupt) and `Ctrl+Shift+C` |
| `docs/architecture-invariants.md` -> "Command palette and shortcut registry" | append the invariant: the no-selection path must return `true` without `preventDefault`, the branch is keydown-only, and `copyTerminalSelection` must stay out of `SHORTCUT_ACTIONS` |
The shortcut overlay (`Ctrl+?`) and App Settings -> Shortcuts are registry-driven and pick the entry up with no edit.
## 5. Edge cases and risks
| Case | Handling |
| --- | --- |
| Handler also fires for `keypress`/`keyup` | gated on `ev.type === 'keydown'`. xterm's `_keyPress` bails on ctrl combos anyway, so no stray byte |
| CJK IME composing | the existing `isComposing || keyCode === 229` guard is the first line of the handler and stays first |
| Local echo overlay has unsent `pendingText` | the copy branch returns before `onData`, so `pendingText`, flushed offsets and the durable input queue are untouched. The no-selection path is byte-identical to today, including the "control char flushes buffered text then sends `0x03`" logic at `terminal-ui.js:895` |
| Plain HTTP (LAN install) | `_copyText` falls back to `execCommand`, then focus is restored |
| Clipboard write rejected (permissions policy, no gesture) | error toast, right-click Copy still available |
| Whitespace-only or empty selection | `getSelection()` empty string is treated as "no selection", so Ctrl+C still interrupts |
| macOS Cmd+C | registry treats ctrl/meta as interchangeable, so with a selection it takes our path (same visible result as today's native copy), without one it falls through |
| Chrome/Firefox `Ctrl+Shift+C` is the devtools inspect chord | browser-level and may still toggle devtools, our copy runs regardless. Document as a caveat, `Ctrl+C` is the primary path |
| Selection in a tab whose TUI owns the mouse (`shell`/`opencode`/`antigravity`) | unchanged, `Shift`+drag selects, then Ctrl+C copies |
| Web tab (iframe dashboard) focused | xterm handler never runs, browser-native copy inside the iframe |
| Teammate/subagent terminals (`panels-ui.js:2268`, `onData` wired) | same limitation exists there, out of scope for this PR (section 8) |
## 6. Test plan
New file `test/terminal-copy-selection.test.ts` (node env, `vm` harness in the style of `test/command-palette-ui.test.ts`), covering `shouldCopyTerminalSelectionFromShortcut` in isolation:
1. Ctrl+C keydown -> true, keyup/keypress of the same chord -> false.
2. Ctrl+Shift+C -> true, plain `c` -> false, Ctrl+K -> false.
3. Registry entry `disabled: true` -> false for every chord.
4. Rebound entry (for example Alt+Y) -> true for the rebind, false for Ctrl+C.
5. Missing registry (harness without `getShortcutRegistry`) -> falls back to the `c` check.
Static assertions appended to `test/keyboard-shortcuts.test.ts` (this suite already pins the xterm-handler chokepoint):
6. `DEFAULT_SHORTCUTS` contains `id: 'copy-selection'` and `SHORTCUT_ACTIONS` does **not** contain `copyTerminalSelection` (the interrupt-safety invariant).
7. `terminal-ui.js` contains the `shouldCopyTerminalSelectionFromShortcut` branch and a `return true` no-selection fall-through.
8. README + help modal rows exist (mirrors the existing palette/Alt-nav doc assertions).
`test/help-modal-shortcuts.test.ts`: add `expectShortcut(helpModal, ['Ctrl', 'C'], 'Copy Selection')`.
New browser test `test/terminal-copy-shortcut.test.ts` (Playwright, port **3174**, free per a scan of `test/`), following `test/webgl-fallback.test.ts`: boot `WebServer`, grant `clipboard-read`/`clipboard-write`, `terminal.write()` a known line, `selectLines()`, real `page.keyboard.press('Control+c')`, then assert clipboard content, empty `onData` capture, cleared selection and the toast. Second case: no selection, assert `onData` saw `\u0003` and the clipboard is unchanged.
Per repo convention, browser suites are excluded from CI, so add the filename to the exclude list in `config/vitest.ci.config.ts` and run it locally.
Regression runs: `npm test -- test/keyboard-shortcuts.test.ts`, `test/help-modal-shortcuts.test.ts`, `test/command-palette-ui.test.ts`, `test/input-send-order.test.ts`, then `npm run test:ci`.
## 7. Manual verification before COM (CLAUDE.md rule)
Against a throwaway session on the live instance (`curl -sk https://localhost:3000/...`, never w1/w2/w3):
1. Select output with the mouse, press Ctrl+C, confirm the toast, paste elsewhere, confirm the agent did not stop.
2. Press Ctrl+C again with nothing selected, confirm the agent interrupts.
3. Type a few characters with local echo on (phone or `localEchoEnabled` forced), press Ctrl+C with no selection, confirm buffered text plus interrupt behave as before.
4. Uncheck the shortcut in App Settings -> Shortcuts, confirm Ctrl+C always interrupts even with a selection.
5. Rebind it, confirm the new chord copies and Ctrl+C reverts to pure interrupt.
6. Repeat 1 and 2 in an `opencode` or `shell` tab using Shift+drag to select.
7. Load over plain HTTP (`--host` LAN or `http://127.0.0.1:<port>`) and confirm the `execCommand` fallback copies and focus returns to the terminal.
8. Mobile smoke: confirm nothing changed (selection is CSS-disabled, no Ctrl key).
## 8. Out of scope, follow-ups worth filing separately
- **Teammate/subagent terminals** (`panels-ui.js:2268`) have the same blocked-copy problem. One `attachCustomKeyEventHandler` reusing `copyTerminalSelection` would fix them, but it touches a different surface and deserves its own change.
- **A mobile copy affordance.** Selection is disabled on touch, so phones still cannot copy terminal text. The unreferenced `copyTerminal()` (whole buffer) plus a keyboard-accessory "Copy" button would be the cheapest answer.
- **Right-click context menu** with Copy/Paste, better discoverability than any chord, but a bigger UI surface.
- **`copyTerminal()` cleanup**: it uses raw `navigator.clipboard` rather than `_copyText`, so it would fail on plain HTTP if ever wired up.
## 9. PR mechanics
- Branch off `master` (verify with `git branch --show-current`, the tree is shared), stage explicit paths only.
- Files touched: `src/web/public/app.js`, `src/web/public/terminal-ui.js`, `src/web/public/i18n.js`, `src/web/public/index.html`, `README.md`, `CLAUDE.md`, `docs/architecture-invariants.md`, `docs/terminal-copy-shortcut-plan.md`, three test files, `config/vitest.ci.config.ts`.
- `index.html`, `app.js` and `terminal-ui.js` are `.prettierignore`d hand-formatted assets, match the surrounding style by hand. `npm run check:public-assets` and `npm run check:frontend-syntax` are the guards.
- No changeset in this PR: a merged, unconsumed changeset turns the Release workflow red until the next COM, and the COM flow writes release notes covering everything since the last tag (current version is 1.10.0).
- Close #211 from the PR body.
Rough size: about 60 lines of product code, most of the work is the tests and the four documentation surfaces.
+1 -1
View File
@@ -1,6 +1,6 @@
# Plan Usage Limits Display — Design & As-Built
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** Opt-in via App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`, default OFF). Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`). **Default changed in 1.9.3: desktop now defaults ON, handhelds stay OFF, resolved via `planUsageChipEnabled()`.** The per-device notes further down describing it as opt-in/synced record the original 2026-06-14 shape, not current behavior. Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
>
> Two surfaces from one `statusLine` callback:
> - **Header chip** (top-right) — account-wide **plan limits**: `5h 35% · 7d 38%`, per-window green/yellow/red.
+1 -1
View File
@@ -75,5 +75,5 @@ allowance. The commitments above take effect at `1.0.0`.
## See also
- `CLAUDE.md` — the COM release workflow (changesets, version bump, deploy)
- `SECURITY.md` — security reporting and the supported-version policy
- `.github/SECURITY.md` — security reporting and the supported-version policy
- `docs/security-architecture.md` — the full trust model
+201
View File
@@ -0,0 +1,201 @@
<!-- Design doc drafted 2026-07-28 from WWDC26 session 224 research. STATUS: PLANNED, NOT IMPLEMENTED. Blocked on macOS 27 "Golden Gate" (beta now, GA expected fall 2026). -->
# VM Cases (macOS Virtualization framework), Implementation Plan
## Status
PLANNED, nothing implemented. This is the design + phased execution plan for a native-macOS VM isolation tier for cases ("the VM subsystem"), modeled on Docker cases (`docs/docker-cases-plan.md`). Testbed prerequisite: a macOS 27 host (see Section 8).
**⚠ DESIGN DIRECTION (owner, 2026-07-29): the subsystem is GUI-first.** Users want real macOS desktops, not headless SSH machines. Guests may be macOS (GUI-only in practice) or Linux (GUI or headless). Key decision 3 below carries the full consequences; anything in this doc that reads as "Linux-first / headless-first" predates this and has been revised.
**2026-07-29: Phase 0 substantially validated on the beta testbed; full Apple-stack reference now lives in [`docs/vm-subsystem-apple-stack.md`](vm-subsystem-apple-stack.md)** (API surfaces, beta bugs, our empirical results, and design implications). Plan-relevant corrections from that work: vmnet's topology/port-forwarding APIs are macOS 26 (only the loopback fix is 27); guest provisioning is macOS-guests-only (Linux stays cloud-init, proven working); DiskImageKit has NO flatten/merge, so the `export` subcommand ships the layer chain (or flattens in-guest) instead of flattening; seed ISOs are base-build-time only, never attached at case runtime; per-case EFI variable stores are mandatory; guest health checks read DHCP leases, never serial/ping.
## 1. Context and motivation
WWDC 2026 session 224 ("Expand the Capabilities of your Virtualization App", https://developer.apple.com/videos/play/wwdc2026/224/) shipped the missing pieces for programmatic, fleet-style VM management on macOS:
- **`VZMacGuestProvisioningOptions`**: automated first-boot setup of a macOS guest (user account, auto-login, SSH enabled) with zero interactive setup.
- **DiskImageKit**: stacked disk images on the Apple Sparse Image Format (ASIF): a read-only base layer plus cheap per-VM cache/overlay layers. Direct analog of Docker image layers + writable container layer.
- **vmnet framework**: custom network topologies and port forwarding from the host process.
- **`VZCustomVirtioDevice`**: custom low-latency host<->guest channels (Linux guests).
- **AccessoryAccess**: USB passthrough (not relevant to Codeman, out of scope).
Codeman's isolation story today is Docker cases. On macOS, Docker means Docker Desktop / a Linux VM anyway, with weaker fidelity and a heavyweight dependency. The Virtualization framework gives hardware-virtualized per-case sandboxes natively, with a layered-image story that mirrors what `scripts/build-agent-image.mjs` does for Docker. This is the premium native-macOS tier ON TOP of Docker cases, never a replacement (Docker remains the cross-platform story; the Linux prod box cannot use any of this).
## 2. Platform reality (hard constraints)
| Constraint | Detail |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Host OS | macOS 27 "Golden Gate" required for the new APIs (dev beta since 2026-06-08, public beta since 2026-07-13, GA expected fall 2026) |
| Host hardware | Apple Silicon only (macOS 27 dropped Intel). Testbed: the owner's dedicated MacBook (Section 8); the M4 Mac mini (macOS 26.4, runs the second Codeman install) stays on stable + untouched |
| Guest provisioning | `VZMacGuestProvisioningOptions` needs macOS 27 on BOTH host and guest. Linux guests provision via cloud-init instead |
| macOS guest concurrency | **Hard kernel cap: 2 concurrent macOS VMs per host. MEASURED on 27 beta 4 (2026-07-29), not inferred**: the 3rd VM is refused instantly with `VZErrorDomain` code 6 while 39% of RAM is free, so more hardware does NOT raise it. Since macOS GUI guests are the headline use case, this is a real product capacity limit to schedule around and surface in the UI. Linux guests are uncapped (resource-bound only) |
| Language | Virtualization framework is Swift/ObjC only; Node cannot call it. Requires a Swift helper binary (Key decision 2) |
| Entitlement | Host process needs `com.apple.security.virtualization`. Fine for a locally built dev binary; distribution needs signing thought (Section 9) |
| Nested virtualization | Linux-guest-only on M3+. A macOS 27 VM cannot dependably host its own guests, so the host-side APIs must be tested on bare-metal 27 (dual-boot) |
| CI | Cannot run in CI (needs beta macOS on Apple Silicon). Same answer as tmux/docker: no-op all VM IO under `VITEST`, unit-test the pure parts |
## 3. Goal and user stories
Add "VM cases" to Codeman: a case can point at a per-case virtual machine on a macOS host, and any CLI backend runs inside it over the existing remote-SSH session machinery. A LOCATION OVERLAY on cases, exactly like remote-SSH and Docker cases, NEVER a sixth `SessionMode`.
- As a Mac user, I link a case to a VM so an autonomous run executes behind a hardware virtualization boundary (stronger than Docker's shared kernel) while file viewing, transcripts, and hooks keep working.
- Per-case VMs are instant and cheap: a shared provisioned base image plus a per-case overlay, not a full image copy per case.
- Killing a session kills only its in-guest tmux; the VM stays up while sibling sessions remain; case delete tears the VM down.
- I export a case's VM overlay as a portable artifact (mirror of `docker-exports/`), secrets excluded.
- On a non-mac host, or a Mac without the helper, the feature is invisible: zero UI, zero probes, zero errors.
Non-goals for the MVP: USB passthrough, custom Virtio channels (Phase 3 candidate), macOS-guest fleets (capped at 2 anyway), Kubernetes-style orchestration, Intel Macs.
## 4. Architecture
```
Codeman (Node, unchanged session layer)
| JSON over stdout (same pattern as shelling out to docker/tmux)
v
codeman-vm (Swift package: CLI + per-VM GUI runner app in the console session)
| Virtualization / DiskImageKit / vmnet
v
per-case VM (macOS or Linux)
|-- GUI mode: VZVirtualMachineView in a window --> guest screen sharing --> browser (noVNC)
|-- shell: SSH on vmnet IP --> existing remote-SSH tmux machinery
^ VirtioFS: host case dir mounted at the SAME absolute path
```
Note the runner is a **GUI app in the console user's session**, not a detached daemon: a daemon-launched VM cannot render, which is fatal for macOS guests and for Linux desktop cases.
### Key decision 1: location overlay, not a mode
Identical reasoning to Docker/remote-SSH (see CLAUDE.md): the session layer, respawn, Ralph, recovery, and quick-start plumbing all stay untouched. `SessionMode` stays five-valued. State mirrors the Docker pair: `~/.codeman/vm-hosts.json` + `vm-cases.json`, new `src/vm-hosts.ts` with the storage + pure helpers split.
### Key decision 2: Swift helper CLI (`codeman-vm`)
The framework is Swift-only, so all VM work lives in a SwiftPM package (`packages/codeman-vm/`), a CLI with a stable JSON contract:
- `create-base --guest linux|macos`: build the shared base image. Linux: boot an arm64 cloud image with EFI + cloud-init, install Node 22 + tmux + the four CLIs (same inventory as `docker/agent.Dockerfile`), seal as base ASIF. macOS: IPSW restore + `VZMacGuestProvisioningOptions` (agent user, SSH on), then **desktop-readiness baking**, which is mandatory for GUI guests: suppress the per-user first-login assistant (`com.apple.SetupAssistant` keys + the User Template), enable auto-login (`autoLoginUser` + `/etc/kcpassword`), disable screensaver/lock/display-sleep, and set a static wallpaper (animated "aerials" wallpaper is unusable over remote display). ⚠ Use RAW (not ASIF) for macOS guest disks until the beta's macOS-guest space-reclamation bug is fixed.
- `create <case>`: DiskImageKit stacked image: shared read-only base + fresh per-case overlay. Near-instant, space-efficient.
- `start <case>` / `stop` / `status` / `ip`: lifecycle + vmnet NAT; `ip` reports the guest SSH endpoint.
- `export <case>` / `import`: flatten overlay + workspace tar + manifest, credentials excluded (mirror of docker-export).
A VM dies with its owning process, so `start` spawns a DETACHED per-VM runner process (analog of the detached `scripts/self-update.sh` trick) rather than a monolithic daemon; `status` talks to it over a unix socket in the instance data dir (`dataPath()`, never a hardcoded `~/.codeman` path).
### Key decision 3: multi-guest, and GUI is a first-class mode (REVISED 2026-07-29 by the repo owner)
The subsystem supports both macOS and Linux guests, and a guest runs in one of two **display modes**:
| | macOS guest | Linux guest |
| --- | --- | --- |
| **GUI mode** | **the point of the feature**; a real macOS desktop. Mandatory: nothing renders without an attached `VZVirtualMachineView` in an unlocked host session | supported (EFI + virtio-gpu framebuffer) for desktop Linux cases |
| **Headless mode** | not offered: a macOS guest with no view renders nothing, so a "headless macOS desktop" is a contradiction. SSH-only macOS is possible but is not what this feature is for | supported and cheap; the natural mode for agent/CI work, driven over SSH |
Consequences that flow from GUI being first-class:
- VM processes are **GUI apps in the console user's session** (LaunchAgent / `launchctl asuser`), never daemons. A daemon-launched VM cannot render.
- **The host is part of the product surface**: it must auto-login, never lock, never sleep, and keep a live WindowServer. Host lock == every VM's screen goes black, so the screen lock is effectively a global kill switch for every VM display on the machine. The product must own these host settings rather than treat them as user preference.
- **FileVault conflicts with unattended GUI hosting** and the trade-off must be a deliberate choice: FileVault disables auto-login, so a full-disk-encrypted host needs a human at a keyboard (or a remote screen-sharing session) after every reboot before any VM can render. Options are (a) FileVault on, accept manual login per boot, (b) FileVault off on a dedicated VM host so it boots straight into a rendering session, or (c) FileVault on plus a remote-unlock runbook. Codeman should detect the state and tell the user which one they are in instead of silently serving black screens.
- **Guests must be desktop-ready, not just booted**: auto-login, no screensaver/lock, and the per-user first-login assistant pre-suppressed at base-image time (`com.apple.SetupAssistant` keys, plus the User Template so later accounts inherit it). Otherwise the user connects to a login prompt or a setup wizard, which is exactly what happened during the first hands-on run.
- **Capacity is capped for macOS**: at most 2 concurrent macOS VMs per host, confirmed by our own test on 27 beta 4 (3rd refused with `VZErrorDomain` 6 at 39% free RAM; it is a kernel quota, so bigger hardware does not help). Scheduling must queue or evict beyond 2, the UI must explain why, and the scheduler should tolerate the acknowledged slot-leak bug (a slot occupied with nothing running, host-reboot to clear). Linux guests are uncapped and bounded only by host resources, which is the lever for scaling case counts on one machine.
- **Access is via the guest's own screen**, viewable in a browser through the noVNC chain (see `docs/vm-subsystem-apple-stack.md` §8), so no client-version or client-install requirements land on the user.
Provisioning per guest type: `VZMacGuestProvisioningOptions` for macOS (needs 27-on-27, first-boot-only, and does NOT skip the per-user wizard), cloud-init NoCloud seed ISO for Linux (proven working).
### Key decision 3b: the GUI VM host profile, and supervision that catches black screens
GUI hosting only works if the host is configured for it and supervised. This profile was derived the hard way on the testbed (prototyped there 2026-07-30) and should be what `codeman-vm` installs and verifies:
**Host profile** (the product should own these, not leave them to preference):
1. **No login barrier.** Either FileVault off + auto-login (a dedicated VM host boots straight into a rendering session, fully unattended), or FileVault on and remote reboots done with `sudo fdesetup authrestart`, where the pre-boot unlock *is* the login so the machine returns already logged in with encryption intact. **`authrestart` is VERIFIED on the testbed (2026-07-30): the host rebooted remotely and came back with a live logged-in console session, FileVault still enabled, no password prompt** — this is the recommended pattern for an encrypted GUI VM host. Plain reboots on a FileVault host always need a human, so Codeman should detect that combination and warn instead of serving black screens.
2. **Never lock**: lock policy off (needs the account password, so it is a setup step, not a scriptable one) plus `caffeinate -d -i -m -u` re-armed per session.
3. **Never sleep**: `pmset -a sleep 0 displaysleep 0 disablesleep 1`; a physical display is NOT required (a lid-closed laptop renders fine, only an unlocked session matters). Note OS updates reset these.
4. **Session-independent control plane**: run VPN/remote access as a system service, never a session app, and keep the access chain (forwards, VNC proxies, web endpoints) in LaunchDaemons so a session restart cannot sever operator access.
**Supervision** must be a **root LaunchDaemon**, not a user LaunchAgent. This is the load-bearing detail: a user agent cannot launch a GUI app into the Aqua session, so its restart attempts fail *silently* (the child dies instantly, leaving an empty log while the supervisor cheerfully reports success). A root daemon can, via `launchctl asuser <uid> sudo -u <user> …`, and those launches persist. Prototyped and verified on the testbed 2026-07-30; a working supervisor runs on a short interval and:
- Restarts the runner when the process is gone **or when its log shows `WindowServer event port death`**, which means it is permanently blind while still looking alive.
- Defers restarts while the console is at the login window, and launches into whichever session actually exists (resolve the console user with `stat -f %Su /dev/console`, never a hardcoded one).
- Re-points the guest port-forward whenever the guest's NAT lease changes, which happens on **every guest boot** under plain NAT. A vmnet DHCP reservation for a stable per-case IP is the better long-term answer.
- **Re-applies host power settings**, because `pmset -a disablesleep 1` does NOT survive a reboot (caught on the supervisor's first run after a real reboot) and OS updates reset it too.
- Re-arms the keep-awake helper, which dies with its session.
- Ideally also samples the guest framebuffer for non-black content, since a black screen is the one symptom common to every failure mode here.
`pgrep` alone is worthless for health: every failure mode in this session presented as a healthy process.
### Key decision 4: sessions ride the existing remote-SSH machinery
A provisioned guest is literally an SSH host on a vmnet IP. Session launch = the remote-SSH flow with the host swapped in: durable remote `tmux -L codeman-remote`, session names failing `SAFE_MUX_NAME_PATTERN` on purpose, EVERY ssh command line through `buildSshConnectionArgs()` (command-injection invariant), run flows through `POST /api/quick-start` (never `POST /api/sessions`, which stat-validates `workingDir` locally). What is genuinely new is only lifecycle (create/start/stop/export) and the vm-hosts/vm-cases overlay state.
### Key decision 5: workspace via VirtioFS at the same absolute path
Mirror the Docker bind-mount invariant: the case workspace is a real host directory shared into the guest via VirtioFS and mounted at the SAME absolute path. That keeps file-routes/watchers on real host bytes and makes the in-guest transcript projHash match the host. Without this, transcripts/attachments/file viewer all silently degrade.
### Key decision 6: credentials seeded, hooks bridged
- Credentials are SEEDED (read-only share, copied into the guest once at create), never shared read-write, and excluded from exports: byte-for-byte the Docker cases rule and rationale.
- Hooks: on the loopback-only prod bind a guest cannot reach `127.0.0.1:3000`. Mirror `CODEMAN_DOCKER_BRIDGE_HOOKS` with a `CODEMAN_VM_BRIDGE_HOOKS` opt-in listener on the vmnet gateway IP; otherwise idle detection falls back to output-based, same as Docker.
### Key decision 7: drift and teardown copy Docker semantics verbatim
Config hash label on the VM (guest type, cpu/mem, share list); a drifted launch is REFUSED, never silently launched stale. One VM per case shared by all sessions; session kill = in-guest tmux kill only; case delete = stop + remove overlay; instance-scoped boot reaper for orphaned runner processes.
## 5. Implementation phases
**Phase 0, testbed (no repo code):** dedicated MacBook on the macOS 27 beta, remotely accessible over the tailnet (setup protocol in Section 8), Xcode 27 beta, then a throwaway Swift script proving the loop: create base -> overlay -> boot -> ssh in. This validates 80% of the design before any Codeman code.
**Phase 1, `codeman-vm` helper:** SwiftPM package, the six subcommands above, JSON contract doc, detached runner + unix-socket status, Linux base image build. Deliverable is testable entirely without Codeman.
**Phase 2, Codeman integration:** types (`VmHost`/`VmCase`/`SessionVm`), `src/vm-hosts.ts` (+ pure helpers: config hash, arg building, endpoint parsing), Zod schemas, `case-routes` link/unlink + listing, `quick-start` vm branch reusing the remote-SSH launch path, `Session` threading + recovery round-trip, `VITEST` no-op layer, unit tests. Feature-detect: darwin + arm64 + helper binary present, else invisible.
**Phase 3, polish:** export/import UI, frontend Create Case "VM" tab + case-picker labels, SSE `vm:*` events, macOS-guest opt-in with cap surfaced, custom-Virtio input channel exploration, CLAUDE.md Key Pattern + `docs/vm-cases.md` + COM.
## 6. Testing
- Pure helpers unit-tested (ports pattern from `docker-hosts.ts`: 26 tests there, aim similar).
- All helper-invoking IO no-ops under `VITEST` (the `IS_TEST_MODE` pattern in `tmux-manager.ts`).
- End-to-end verification happens ON the beta MacBook, per the always-end-to-end rule: real base build, real per-case overlay boot, real quick-start into the guest, workspace round-trip through VirtioFS, session-delete keeps VM up, case-delete removes it.
- CI never runs the real path; the static guards are type-level + unit-level only.
## 7. Risks
1. **Beta API churn**: everything here targets beta SDKs; symbol/behavior changes are likely before fall GA. Mitigation: Phase 0/1 are throwaway-tolerant; no Codeman-side commitment until the helper contract survives a beta cycle.
2. **New artifact class**: Codeman ships pure TypeScript today; a Swift binary changes build/distribution (build-on-install via `xcrun swift build` on macs with Xcode CLT? prebuilt signed binary per release?). Needs an owner decision; local dev build is fine for the whole beta period.
3. **Entitlement/signing**: `com.apple.security.virtualization` is trivial for local dev, real for distribution.
4. **Adoption gating**: users need macOS 27 + Apple Silicon for months after GA. Docker cases remain the default recommendation; VM cases ship dark (feature-detected) with zero cost to everyone else.
## 8. Beta testbed plan: dedicated MacBook (actionable now)
Testbed is a dedicated MacBook the owner sacrifices to the beta (after a full backup). This supersedes the earlier dual-boot-the-Mini idea (git history has it): a dedicated machine means no OS-switching, no downtime for the Mini's live Codeman, and no FileVault pre-boot headaches.
**Sequencing rule that makes it headless: configure ALL remote access on the CURRENT macOS first, THEN upgrade in place.** An in-place beta upgrade preserves Remote Login, Tailscale, user accounts, and auto-login, so there is no Setup Assistant and no post-install physical step. (A fresh install would boot into GUI-only Setup Assistant with no SSH, which on a headless box is a dead end.)
Confirmed hardware (2026-07-28): MacBook, M3, 16 GB RAM, 256 GB disk with ~100 GB free. Verdict: green. M3 = eligible + nested-virt capable; 16 GB = host + 2-3 concurrent Linux guests (macOS guest = one at a time); 100 GB = fits with discipline: install Xcode 27 beta with the macOS platform only (skipping iOS/watchOS/tvOS simulators saves 15-20 GB), and defer any macOS guest base (~30 GB) to an external SSD or until actually needed. Linux guests + sparse ASIF overlays are the comfortable path.
### Pre-upgrade checklist (owner, physical, once)
1. Full backup (Time Machine or clone); the machine should be considered beta-only afterwards.
2. Tailscale: install, sign into the tailnet, confirm it appears in `tailscale status` from another node.
3. System Settings -> General -> Sharing: **Remote Login ON** (SSH) and **Screen Sharing ON** (for the rare GUI-only moments: Xcode license, Apple Account dialogs).
4. **FileVault stays ON** (owner decision 2026-07-28, security over convenience). Consequences: auto-login is unavailable, but FileVault's pre-boot unlock doubles as login, so an unlocked boot still lands in a live GUI session; planned remote reboots go through `sudo fdesetup authrestart` (unlocks for exactly one restart); an UNPLANNED reboot (beta kernel panic, battery drain) parks the machine at the pre-boot screen, no SSH/Tailscale, until the password is typed physically. If the testbed goes silent, suspect this first. Keep it on AC so the battery absorbs power blips.
5. Beta enrollment (manual): sign into the Apple Account in System Settings; System Settings -> General -> Software Update -> **Beta Updates** -> select the **macOS 27 Developer Beta** (preferred: framework fixes land weeks earlier than public beta; free since 2023 after accepting the agreement once at developer.apple.com; public-beta alternative: enroll at beta.apple.com). Then run the offered upgrade: plugged in, lid open, trusted network.
6. Send over: tailnet name/IP, username, and a first-login password (key install + lockdown happens remotely right after).
### Post-upgrade setup (remote, over the tailnet)
1. Verify: `sw_vers` reports 27.x, SSH reachable.
2. Server-ize the laptop: `sudo pmset -a sleep 0 disksleep 0 disablesleep 1` (lid-closed operation without an external display), `womp 1` (wake on network), `sudo systemsetup -setrestartpowerfailure on`. Keep on AC power.
3. Install the controlling host's SSH key, then disable password auth.
4. Xcode 27 beta install (the one step needing the owner's Apple Account sign-in once, doable via Screen Sharing from anywhere); `xcode-select`, license accept, verify `swift --version` + the 27 SDK (`xcrun --show-sdk-version`).
5. Phase 0 prototype loop, all remote from here: Linux guest base image (no 27-on-27 provisioning dependency), DiskImageKit overlay, boot, vmnet NAT, ssh into the guest, run `claude --version` inside.
6. Only after that loop works: start Phase 1 in `packages/codeman-vm/`.
## 9. Open decisions (owner)
1. Linux base distro/image for the default guest (proposal: Ubuntu 24.04 arm64 cloud image, matching the docker agent image's userland).
2. Helper distribution for GA: build-on-install vs prebuilt signed binary vs "bring your own Xcode".
3. Ship dark behind `CODEMAN_VM_CASES=1` for the first release, or feature-detect only?
4. Export format parity with docker-exports (one manifest schema for both?).
## References
- Session 224: https://developer.apple.com/videos/play/wwdc2026/224/
- Fleet-angle writeup: https://bitrise.io/blog/post/wwdc26-the-virtualization-framework-updates-that-matter-for-large-mac-fleets
- Beta timeline: https://www.macworld.com/article/3189014/apple-july-2026-ios-ipados-macos-27-public-betas-tv-arcade-releases.html
- Internal analogs: `docs/docker-cases-plan.md` (architecture template), `docs/remote-sessions.md` (session transport), `docs/architecture-invariants.md#docker-cases`
+281
View File
@@ -0,0 +1,281 @@
<!-- Reference doc for the VM subsystem (Codeman VM cases). Compiled 2026-07-29 from: Apple DocC JSON backend, macOS 27 beta 4 SDK on the testbed, a multi-source web research sweep, and hands-on prototyping on a MacBook Air M3 running macOS 27.0 beta (26A5388g). Companion to vm-cases-plan.md (the Codeman integration plan). -->
# The VM Subsystem: Apple Virtualization Stack Reference (macOS 27 "Golden Gate")
"VM subsystem" is the working name for Codeman's native-macOS VM isolation tier and everything under it. This document is the single place for what the Apple stack actually provides, what we have verified ourselves on the beta, and what is known-broken. The Codeman-side design lives in `docs/vm-cases-plan.md`.
**Research method note:** Apple's HTML doc pages are JS-rendered and come back empty to fetchers. The working route is the DocC JSON backend: `https://developer.apple.com/tutorials/data/documentation/<path>.json` (page content) and `https://developer.apple.com/tutorials/data/index/<framework>` (full symbol tree with per-symbol `beta` flags). Everything below marked "Apple docs" was parsed from that backend directly.
## 1. Component map and minimum OS versions
| Component | What it is | Min host OS | Notes |
| --- | --- | --- | --- |
| Virtualization.framework core | VMs, EFI/Linux boot, virtio devices, VirtioFS | macOS 11-13 era | Unchanged basics; our prototype uses nothing newer than macOS 13 APIs except the DiskImageKit bridge |
| **DiskImageKit** | ASIF + raw disk images, layered stacks | **macOS 27** | Swift-only, no ObjC headers. Section 2 |
| **Guest provisioning** | First-boot account/SSH setup for macOS guests | **macOS 27 host AND guest** | Mac guests only as of beta 4. Section 3 |
| vmnet topology/port-forward/DHCP APIs | Custom networks, port forwarding | **macOS 26** (NOT 27) | 27 adds exactly one fix: loopback port forwarding. Section 4 |
| `VZVmnetNetworkDeviceAttachment` | In-process vmnet attach | macOS 26 | |
| **`VZCustomVirtioDevice`** family | Custom paravirt devices | **macOS 27** | Linux guests only, custom guest driver required. Section 5 |
| AccessoryAccess (USB passthrough) | USB claim + attach to VMs | macOS 27 | Requires paid-team provisioning profile, Dock app. Out of scope for Codeman. Section 6 |
Corrections to the WWDC-session framing we started with: vmnet's topology family is a macOS 26 story (129 symbols, zero beta-flagged in 27); provisioning does NOT currently extend beyond macOS guests despite the generic-looking `VZGuestProvisioningOptions` base class; DiskImageKit has no attach/mount API at all (it is a file-format library that hands `DiskImage` objects to Virtualization, no `/dev/diskN`, no root needed, no entitlement documented).
## 2. DiskImageKit (macOS 27, Swift-only)
Public framework, `/System/Library/Frameworks/DiskImageKit.framework`. No ObjC headers; the API surface lives in the `.swiftinterface`. Verified present in the CLT 27 beta 4 SDK, and our prototype compiled against it with plain `swiftc` on the first attempt.
### API surface (complete as of beta 4)
```swift
class DiskImage {
convenience init(creating: some DiskImage.CreationConfiguration) throws
convenience init(opening: some OpenConfigurationProtocol) throws
func appending(any DiskImage.CreationConfiguration & DiskImage.StackableLayer) throws -> any StackedImage
func appending(consuming DiskImage) throws -> any StackedImage // reattach an existing layer; validates parentUUID
func truncate(blockCount: Int) throws // stacked: affects top layer; does NOT resize guest fs
var blockCount, blockSize, format, layerType, layerUUID, parentUUID, openMode, size, url
}
protocol StackedImage: DiskImage { var layers: [DiskImage] }
struct OpenConfiguration { init(url:mode:); Mode = automatic | readOnly | readWrite }
// CreationConfiguration statics: .asif(url:blockCount:blockSize:), .asifLayer(url:type:), .raw(url:blockCount:)
// DiskImage.LayerType: .cache | .overlay | .overlay(blockCount:)
// DiskImage.BlockSize: .bytes512 | .bytes4096
// Errors: CorruptedImageError, IncompatibleStackingError(reason), InvalidBlockCountError, UnsupportedFormatError
```
Bridge into Virtualization is a new beta convenience init on the existing attachment class. Note there is no `readOnly:` parameter; read-only-ness comes from each layer's own `openMode`:
```swift
VZDiskImageStorageDeviceAttachment(diskImage: stack, cachingMode: .automatic, synchronizationMode: .full)
```
### Stacking rules (Apple docs, verbatim where quoted)
- ASIF works standalone or stacked. "You can only use RAW images as standalone images or as **base** images in stacked configurations." Upper layers are always ASIF.
- **One cache layer per stack**, any number of overlays conceptually, "shallow stacks perform better" (WWDC 224). No published max-depth guidance.
- "Layers are processed from bottom (base) to top. The **topmost layer determines the stack's size and receives all writes**." `.overlay(blockCount:)` therefore also grows the virtual disk.
- UUID chaining: appending sets the child's `parentUUID` to the parent's `layerUUID`. Raw bases have no UUID. "The layer UUID **changes if the layer is written to**", and reattaching a mismatched layer throws `IncompatibleStackingError`. This is the mechanism that makes a shared read-only base safe.
- Base sharing across multiple VMs is the stated design intent ("can be shared across multiple VMs"), with the WWDC caveat that per-VM auxiliary files (EFI variable store, macOS auxiliary storage) must be duplicated per VM, never shared.
- **There is no flatten/merge.** An overlay cannot be merged back into its base (confirmed by Howard Oakley's coverage plus an independent hands-on report). Export/move flows must ship the layer chain, or flatten inside a guest (dd to a fresh attached image).
### Known issues and adoption
- **ASIF space reclamation is broken for macOS guests on the beta** (deleted files never return space, survives reboots). Linux guests reclaim correctly on both raw and ASIF via `fstrim -av`. Single detailed field report, unrefuted. Since the VM subsystem targets macOS guests, the practical rule until this is fixed is: back macOS guest disks with RAW, and revisit ASIF stacking for macOS guests each beta (stacking still works, the disks just never shrink).
- **Zero shipping adopters anywhere.** tart has a design issue with no activity; nobody has published working DiskImageKit code. Everything must be treated as field-untested (and our own testing bears that out, Section 8).
- Framework binary grew every beta (588 → 598 across betas 1-4); expect churn until GA.
- Release notes list no DiskImageKit known issues in any beta, which given the above says more about the notes than the framework.
## 3. Guest provisioning (macOS guests only)
```swift
class VZGuestProvisioningOptions: NSObject { func validate() throws } // "use one of its subclasses"
class VZMacGuestProvisioningOptions: VZGuestProvisioningOptions {
var fullName, username, password: String
var logsInAutomatically: Bool
var enablesRemoteLogin: Bool // SSH
}
// Wiring: VZMacOSVirtualMachineStartOptions.guestProvisioningOptions (Mac-typed)
// .setGuestProvisioning(_:) throws (validating setter)
```
- **Requires macOS 27 on host AND guest.** Older guests **silently ignore** the options (no error).
- **First boot after restore only.** Cannot reconfigure an already-provisioned VM; property changes after start are no-ops.
- The base class is forward-looking scaffolding; its only subclass is Mac. A Linux/cloud-init analogue may come later; do not assume it lands in 27.0. For Linux guests, cloud-init NoCloud seed ISOs remain the provisioning path (proven working, Section 8).
- Field-verified behavior (third-party hands-on, beta 3): provisioned account gets full admin + sudo; Setup Assistant fully skipped; SSH reachable ~48 s after first boot. **Race**: the account is created late in first boot (~T+54 s), after LaunchDaemons start (~T+33 s), so anything at daemon-level must wait for the account to exist.
- Open Apple-acknowledged bug: provisioned users are invisible to `CSIdentityQueryExecute()` (FB23716201).
- IPSW acquisition gotcha for automation: `VZMacOSRestoreImage.latestSupported` tracks the latest *release* (returned 26.5.2), not the installed beta; beta IPSWs must be fetched from the seed CDN explicitly.
## 4. vmnet: a macOS 26 feature set, one macOS 27 fix
Everything interesting shipped in macOS 26: `vmnet_network_create`, `vmnet_network_configuration_create`, `..._add_port_forwarding_rule`, `..._add_dhcp_reservation`, subnet/prefix/MTU/external-interface setters, NAT44/NAT66/DHCP/DNS-proxy/RA disables, plus serialization (`vmnet_network_copy_serialization` / `_create_with_serialization`) for handing networks across processes. `VZVmnetNetworkDeviceAttachment` is macOS 26.
macOS 27's only change (beta 4 release notes, verbatim): "The vmnet port forwarding APIs now support port forwarding when communicating over loopback." That closes the old gap where the host could not reach its own forwarded ports via 127.0.0.1 (confirmed working by the original bug reporter). Directly relevant to Codeman's loopback-bound production server talking to per-case guests.
Gotchas:
- vmnet networks are **not persisted**; they die with the owning process. Persist settings yourself and recreate (or serialize across processes).
- The `com.apple.vm.networking` entitlement is still restricted ("contact your Apple representative", though DTS says most requests are approved). The plain `VZNATNetworkDeviceAttachment` needs no special entitlement and is what our prototype uses.
- Ecosystem signal: tart's maintainer is not adopting in-process vmnet (prefers their separate-process softnet), so field testing of these APIs is thin.
## 5. VZCustomVirtioDevice (macOS 27, Linux guests only)
14 new types (`VZCustomVirtioDevice(+Configuration/Delegate/Provider)`, `VZVirtioQueue(+Element)`, `VZVirtioFeatureSet`, shared-memory-region types, `VZGuestMemoryMapping`), wired via `VZVirtualMachineConfiguration.customVirtioDevices`. Mandatory for guest discovery: `deviceID`, `pciClassID`, `pciSubclassID`, `virtioQueueCount`. You must write the Linux guest driver (Virtio spec 1.3/1.4). Threading contract: the framework calls the device/delegate on a serial queue (`deviceQueue`, defaulting to the VM's queue). Zero public adopters. For the VM subsystem this is a Phase 3+ option for a low-latency host-guest channel; SSH over NAT is proven and sufficient for now.
## 6. Signing and entitlements
- **Core loop (VZ + DiskImageKit + provisioning): ad-hoc signing with only `com.apple.security.virtualization` suffices.** Verified by us on beta 4 (plain `codesign --entitlements ... -s -` on a `swiftc` binary) and independently by third parties on beta 3. DiskImageKit documents no entitlement at all.
- **Over-entitling is the actual trap.** Adding `com.apple.application-identifier`/team-identifier keys without an embedded provisioning profile hangs the process before `main` (watchdog kill); shipping `com.apple.vm.networking` unauthorized gets AMFI SIGKILL at exec (exit 137, no crash report, even for `--version`). Keep the entitlements plist to exactly the one key.
- **USB passthrough breaks the ad-hoc story**: `com.apple.developer.accessory-access.usb` is profile-restricted (any paid team, no ad-hoc), additionally requires `com.apple.security.device.usb`, and `AAUSBAccessoryManager` presents UI, so it wants a Dock app, not a headless CLI. Out of scope for Codeman.
- No Xcode required for any of the above: the CLT beta (~500 MB via `softwareupdate`) carries the full macOS 27 SDK including DiskImageKit and compiles/signs everything.
## 7. Ecosystem state (July 2026)
- **tart is now `openai/tart`** (moved from cirruslabs, mid-2026) and **relicensed to FSL-1.1-ALv2** (no longer permissive). Provisioning support shipped in 2.33.0. Old cirruslabs URLs and license assumptions are stale.
- VirtualBuddy shipped provisioning ("Skip Setup Assistant") in 2.2 betas; had to add account-detail validation and a workaround installer for the cross-version bug below.
- lima is deliberately waiting for GA before touching macOS 27 APIs.
- **Code-Hex/vz (Go bindings) is dormant** (no commits since Feb 2026, no macOS 27 APIs), so the entire Go ecosystem (podman-machine, colima) currently has no path to these APIs. Swift is the only realistic binding today, which validates the VM subsystem's Swift-helper design.
- Useful pattern if ever supporting older SDKs: resolve new classes via `NSClassFromString` at runtime (no link-time dependency), fail gracefully when absent.
- **Cross-version restore bug**: installing a macOS 27 guest from IPSW on a macOS 26 host fails at 77-78% (`VZErrorDomain 10007`); fixed in 26.6b3 + Xcode 27b4 era, with a nasty MobileDevice.pkg trap (installing it from Xcode 27 beta on a 26 host requires a full macOS reinstall to undo). Not relevant to our 27-host testbed, very relevant to anyone on a 26 host.
## 8. Our empirical results (beta 4, 26A5388g, MacBook Air M3, 2026-07-29)
Prototype tooling, all in `~/vm-lab/` on the testbed, compiled with CLT-only `swiftc` and ad-hoc signed with the single virtualization entitlement:
| Tool | Purpose |
| --- | --- |
| `vzboot.swift` | Linux guest: EFI boot + virtio disk/net/entropy + NAT + optional cloud-init seed ISO + serial on stdio |
| `vzstack.swift` | Same, but boots a DiskImageKit stack (read-only raw base + ASIF overlay) |
| `vzmac.swift` | macOS guest: `install` (IPSW restore into a bundle) and `run` (boot, `--provision` for first-boot account/SSH) |
| `vzmacgui.swift` | macOS guest in a real window via `VZVirtualMachineView` (required for the guest to render at all) |
| `setup-seed.sh` | Builds a cloud-init NoCloud seed ISO with `hdiutil makehybrid` (volume label `cidata`) |
| `vncproxy.py` | RFB proxy that advertises only security type 2, so version-skewed/browser clients can authenticate |
| noVNC + `websockify` | Browser access; `websockify --web noVNC-<ver> 0.0.0.0:<port> 127.0.0.1:<proxy>` |
| `vmwatchdog.sh` + `vmaccess.sh` | Supervision: root LaunchDaemon that restarts a blind/dead runner, re-points the forward, re-applies `pmset`, re-arms keep-awake; plus a keeper for the proxy/web endpoints |
Host-side diagnostics written during this work (in the session scratchpad, not on the testbed): `vnclogin.py` (Apple DH auth + session open, distinguishes "credentials rejected" from "authorized but session refused"), `vncshot.py` (decodes the raw framebuffer to PNG and reports non-black pixel counts, plus optional synthetic wake input), `relay.py` (plain TCP relay used to bridge a tailnet peer to a LAN-only host), `sshpw.py` (pty-driven password SSH for the one-time key bootstrap into a freshly provisioned guest).
### Proven working
1. **Boot**: Debian 12 arm64 cloud images (nocloud and genericcloud variants) boot under `VZEFIBootLoader` + `VZGenericPlatformConfiguration`.
2. **Networking**: `VZNATNetworkDeviceAttachment` gives the guest a `192.168.64.x` DHCP lease from the host's bootpd (leases visible in `/var/db/dhcpd_leases`, bridge is `bridge100`).
3. **cloud-init provisioning**: NoCloud seed ISO (built with `hdiutil makehybrid -iso -joliet -default-volume-name cidata`) created a `codeman` user with SSH key + passwordless sudo on first boot; `ssh codeman@<lease-ip>` from the host works with key auth.
4. **DiskImageKit stack mechanics**: opening a raw base `.readOnly`, appending an ASIF overlay (`ASIFCreationConfiguration.layer(url:type:.overlay)`), attaching via `init(diskImage:)`, and booting it. The overlay received ~44 MB of boot-time writes while the **base file's SHA-256 stayed bit-identical**, which is the write-isolation property the whole per-case design rests on.
5. **Reattach**: reopening an existing overlay and `appending(consuming:)` onto the same base passes UUID validation.
6. **macOS guest install (added later the same day)**: `VZMacOSInstaller` restore of the 27.0 IPSW (26A5388g, fetched from the seed CDN via appledb; same build as host) into a sparse 64 GiB raw disk + auxiliary storage: INSTALL-OK on the first attempt, ~25 minutes.
7. **Headless guest provisioning WORKS**: `VZMacGuestProvisioningOptions` via `setGuestProvisioning` (username, password, `enablesRemoteLogin`, `logsInAutomatically=false`) produced, with zero GUI interaction: an account with full admin (groups include `80(admin)`, `com.apple.access_ssh`), Remote Login on from first boot, port 22 reachable ~140 s after first-boot start, hostname auto-derived from the account ("Codemans-Virtual-Machine"). SSH password auth is on by default, so the bootstrap path is: pty-driven password login once to install `authorized_keys`, key auth thereafter. Note the provisioned account's sudo is NOT passwordless (`echo <pass> | sudo -S ...`), and provisioning is first-boot-only (later boots take no options and just boot).
8. **Slot-leak bug NOT reproduced on 26A5388g**: a guest-initiated `shutdown -h now` fired `guestDidStop` cleanly and an immediate relaunch started fine (SSH-ready again in ~75 s), so FB22967193 (VM slot leaked on guest-initiated shutdown, host reboot to recover) did not manifest after one cycle. Either fixed in beta 4 or needs more cycles to trigger.
### Unstable / under investigation (beta-quality territory)
Boot reliability degraded over a ~15-VM session on one host boot, ending with reproducible silent hangs (VM process alive, 0% CPU, no DHCP, no ARP, nothing on serial):
- A genericcloud base that had been booted read-write once (cloud-init first boot) subsequently hung on every boot **with the seed ISO still attached**, while booting **without** the seed succeeded, then later runs failed in both configurations. The seed correlation is strong but was observed while host state was already suspect, so it needs a retest from a clean baseline.
- The first stack-boot "success" that later wedged turned out (via DHCP lease timestamp arithmetic) never to have reached the network at all; its overlay growth was pre-network boot writes.
- Working hypothesis, matching a class of acknowledged beta bugs (e.g. the VM-slot counter that leaks on guest-initiated shutdown, FB22967193, where only a host reboot recovers): accumulated hypervisor/vmnet state on the host degrades boots. Requires a host reboot + a disciplined retest matrix to confirm.
### Display rendering: the single most important operational finding
**A VZ macOS guest renders nothing unless a `VZVirtualMachineView` is attached AND the host session is actually drawing.** Verified byte-for-byte: the guest's own screen sharing serves an all-zero framebuffer (0 non-black bytes across 400 KB samples, with a sane pixel format: `rmax/gmax/bmax = 255`, shifts 16/8/0), in-guest `screencapture` fails with "could not create image from display", and no `IODisplayWrangler` shows up in the guest's `ioreg`. Three distinct states all produce black:
1. **Headless** (VM run with no view attached).
2. **View attached, host session locked.** The lock screen suspends drawing and the guest's virtual GPU produces no frames.
3. **View attached, but the app lost its WindowServer connection** (see the incident below): black permanently until the app is restarted.
**Consequence for the VM subsystem: rendering is a first-class requirement, not an optional extra (owner decision 2026-07-29).** The product serves GUI desktops: mandatory for macOS guests, optional-but-supported for Linux guests (which can also run headless over SSH). Any VM in GUI mode must be launched by an app that attaches a `VZVirtualMachineView`, from inside a host GUI session that is logged in and unlocked. That makes the following non-negotiable parts of the design, not workarounds:
- VMs run as **GUI apps in the console user's session** (launched via a LaunchAgent or `launchctl asuser`), never as daemons.
- The **host must auto-login and never lock or sleep**; a locked host is equivalent to a powered-off display for every VM on it.
- The **guest must auto-login, never lock, and have its first-login assistant pre-suppressed**, or the "desktop" a user connects to is a password prompt or a setup wizard.
- A VM app that loses its WindowServer connection is **permanently blind** and must be restarted; supervision has to detect that, not just check that the process is alive.
- The **2-concurrent-macOS-VM cap** becomes a real capacity limit for the product, so it must be surfaced in the UI and tested (still untested worldwide as of this writing).
### Incident 2026-07-29: `killall -HUP loginwindow` (never do this on a remote Mac)
Applying a wallpaper change on the testbed with `killall -HUP loginwindow` restarted the host's login session. Three consequences:
1. **The Mac dropped off the tailnet entirely.** Tailscale's App Store build is a GUI app living in the user session, so killing the session killed the VPN; remote access was gone until someone logged in. Recovery came from a second machine on the same LAN: it could still SSH in, and then relay ports back over the tailnet (a plain TCP relay on a tailnet-connected LAN peer is a good out-of-band path worth keeping ready).
2. **The VM app lost its WindowServer connection** (`HIToolbox: received notification of WindowServer event port death`) while surviving as a process. Every later black screen traced to this, and nothing guest-side could fix it; only restarting the app restored rendering.
3. The session's `caffeinate` died, so the host resumed auto-locking.
Rule: on a remote Mac, never run session-level commands (`killall -HUP loginwindow`, `pkill -u <user>`, logout, fast user switching). `killall WallpaperAgent` alone is session-safe. Before any such command, enumerate what depends on that session: VPN, VM processes, port forwards, keep-awake helpers.
### Keeping host and guest usable unattended
- **Host**: `caffeinate -d -i -m -u` prevents display sleep but does NOT override the lock policy. "Require password after screen saver begins or display is turned off → Never" must be set in System Settings; it needs the account password, so a passwordless-sudo shell cannot script it, and turning it off does NOT dismiss a lock that is already engaged (one more unlock is always needed). `pmset -a disablesleep 1` keeps a lid-closed laptop awake but **does not survive a reboot**, and OS updates reset it too, so a supervisor should re-apply it rather than assume it sticks.
- **Rebooting an encrypted host**: use `sudo fdesetup authrestart`. FileVault's pre-boot unlock doubles as the login, so the machine returns with a **live logged-in console session** and encryption intact, no password prompt, and supervision can then bring the VMs back by itself. Verified 2026-07-30. A plain `reboot` parks at the lock screen and blacks out every VM until a human logs in.
- **Guest**: set `autoLoginUser` plus a valid `/etc/kcpassword` (XOR-obfuscated password file, key `7D 89 52 23 D2 BC DE A3`, payload zero-padded to a multiple of 12). `sysadminctl -autologin` fails with `SACSetAutoLoginPassword error:22` on provisioned accounts, and a fresh guest has no Python, so generate the bytes on the controlling host and copy them in. Then `pmset -a displaysleep 0 sleep 0 disablesleep 1`, `defaults -currentHost write com.apple.screensaver idleTime 0`, `defaults write com.apple.screensaver askForPassword 0`, and `caffeinate` inside the guest. ⚠ `autoLoginUser` was observed being wiped by failed `sysadminctl -autologin` attempts; verify it after each boot until stable.
- **Wallpaper**: animated "aerials" wallpaper is brutal over VNC. The provider lives in `~/Library/Application Support/com.apple.wallpaper/Store/Index.plist` under several keys (`AllSpacesAndDisplays:Desktop`, `:Idle`, and `SystemDefault:*` which is what the login/lock screen uses). Switch each `Provider` to `com.apple.wallpaper.choice.solid-color` with PlistBuddy and restart `WallpaperAgent`. The login-window copy is cached and only refreshes on a later login cycle.
### Remote GUI/SSH access to a guest (recipe, verified 2026-07-29)
The guest lives on the host-private NAT bridge, so remote access is guest-service + host-forward:
1. **In the macOS guest** (over ssh), use ONE mechanism, fully activated. The reliable form is Remote Management in a single kickstart call:
```
sudo .../RemoteManagement/ARDAgent.app/Contents/Resources/kickstart \
-activate -configure -access -on \
-clientopts -setvnclegacy -vnclegacy yes -setvncpw -vncpw <8-char-pw> \
-allowAccessFor -allUsers -privs -all -restart -agent -menu
```
⚠ **Half-configured states authenticate but refuse the session.** Loading `com.apple.screensharing` while Remote Management is deactivated (or vice versa) produces an Apple-client error that names the wrong culprit: *"Screen Sharing is not permitted on <host>. Disable and re-enable Screen Sharing or Remote Management in System Settings"*. A raw-protocol client can still authenticate AND open a framebuffer in that state, so protocol-level tests pass while every Apple client fails. The remedy is exactly what the dialog says, done over ssh: `launchctl unload -w …screensharing.plist`, `kickstart -deactivate -configure -access -off`, `pkill screensharingd`, then the single activate call above.
Notes: `launchctl enable system/com.apple.screensharing` fails with "Could not find service" on this build; `load -w` is the plain-Screen-Sharing path if you deliberately want it instead of Remote Management. Apple clients negotiate `RSA-SRP` (auth type 33) and the guest logs `Authentication: SUCCEEDED :: User Name: … :: Type: RSA-SRP` on success, which is the definitive server-side confirmation.
2. **On the host**: a gateway port-forward makes the guest's 5900 reachable from the whole tailnet without per-client tunnels: self-authorize the host's own key, then `ssh -N -g -L 0.0.0.0:5901:<guest-ip>:5900 <user>@localhost` (nohup'd).
⚠⚠ **NEVER forward on host port 5900.** If the host has Screen Sharing enabled (our testbed does, from the pre-upgrade checklist), launchd already owns 5900 socket-activated. The `ssh -L` bind then fails with "Address already in use" **while the tunnel process keeps running**, so every symptom of success is present (process alive, port answers, real RFB banner) yet **every connection reaches the HOST's login window, not the guest**. This cost us an hour: guest credentials failed against the host's screensharingd, which reads exactly like broken guest auth, and we chased the (real, but irrelevant) provisioned-account identity bug. Diagnostics that would have caught it instantly: `sudo lsof -nP -iTCP:5900 -sTCP:LISTEN` showing `launchd` rather than `ssh`, or the guest's own logs showing NO auth attempts during a failed login. Always use a distinct host port and verify with `lsof` that the forward owns it.
⚠ `-g` binds all interfaces, so the forward is also visible on the host's LAN; the VNC layer still requires the account or VNC password. ⚠ The forward pins the guest IP, which changes per boot under plain NAT; re-point it after a guest reboot (the proper fix is a vmnet DHCP reservation, macOS 26 API, once we move off plain `VZNATNetworkDeviceAttachment`).
Verified working: with the forward on 5901, both a provisioned account and a `sysadminctl`-created one authenticate successfully (RFB `SecurityResult` = 0) against the guest. The guest offers security types `[30, 33, 36, 2, 35]`, i.e. Apple DH/SRP **plus classic type 2**, so non-Apple VNC clients work with the legacy password once ARD's `-setvnclegacy` is set. (The host's screensharingd, by contrast, offered no type 2, which is itself a tell that you are talking to the wrong machine.)
3. **SSH from any tailnet device**: `ssh -J <host-user>@<host> codeman@<guest-ip>` (jump through the host), after adding the connecting machine's key to the guest's `authorized_keys`.
**Client-version incompatibility (macOS 27 servers vs older Screen Sharing clients)**: an older Mac's Screen Sharing client fails Apple's `RSA-SRP` handshake against macOS 27 servers, logging `Authentication: FAILED :: User Name: <user> :: Type: RSA-SRP` server-side, while a macOS 27 client authenticates against the same servers without issue. This was verified against BOTH a macOS 27 guest and a macOS 27 host with the operator's own account, so it is a client-side version skew, not configuration, and no server-side change fixes it. Same family as the documented "macOS 26 host cannot install a 27 guest" bug. Practical workaround: bypass Apple auth entirely with classic VNC auth (security type 2), which macOS offers only when Remote Management legacy VNC is enabled. Two ways to consume it: any third-party VNC client, or a browser via noVNC.
**Browser-based access chain (zero client install, version-proof)**, all hosted on the Mac:
```
browser --HTTP/WS--> websockify (+ noVNC static files)
--> type-2-only proxy # rewrites the server's security-type list to [2]
--> ssh -L forward # loopback hop; see the Local Network note below
--> guest:5900
```
Notes learned the hard way: (a) **never bind the forward on host port 5900** (see the launchd warning above); (b) a Python proxy cannot reach the guest subnet directly because macOS **Local Network privacy** denies headless CLI binaries, surfacing as `No route to host`, so point the proxy at a loopback `ssh -L` forward instead (Apple-signed `ssh` is unaffected); (c) noVNC needs `?resize=scale` or Scaling Mode → Local Scaling, otherwise a Retina host screen (2940x1912) is unusable in a browser window; (d) noVNC speaks security type 2 only, which is exactly why the proxy rewrite is needed.
**Debugging technique that settled all of this**: a ~80-line Python RFB client (scratchpad `vnclogin.py`) that implements Apple DH auth (security type 30) and continues through `ClientInit`/`ServerInit`. It reports the server's `SecurityResult` plus the framebuffer size and desktop name, which separates "credentials rejected" from "authorized but session refused" without any GUI client. Pair it with `log stream --predicate 'process == "screensharingd"'` inside the guest, and drive a REAL Apple client headlessly from the host with `sudo launchctl asuser <uid> sudo -u <user> osascript -e 'tell application "Screen Sharing" to open location "vnc://user:pass@host:port"'`, verifying the result via `lsof -nP -iTCP -a -p <pid>` (an ESTABLISHED socket to the target) since `screencapture` fails on a lid-closed laptop ("could not create image from display"). Tailscale was never implicated: both the raw client and Apple's client work over the tailnet address once the guest service is fully activated.
### Hard-won operational lessons (write these into any tooling)
- **Silent serial is normal, not failure.** Debian's GRUB/kernel log to the graphics console; nothing attaches a getty to hvc0 by default. The reliable boot signal is the DHCP lease (or passive `tcpdump -i bridge100`), never the serial port and never a quick ping (BSD ping's first packet often dies to ARP latency; passive capture showed "dead" guests alive).
- **DHCP lease entries carry truth**: `name=` shows the guest hostname, and the lease timestamps order events; stale entries linger, so compare timestamps before attributing a lease to a boot.
- **Never boot a base image read-write.** Every RW boot mutates it (dhclient lease cache, journal, cloud-init state) and destroys experiment reproducibility, exactly why the production design only ever boots bases under overlays. Provision INTO the base once at base-build time, or provision per-case overlays with the seed, then detach the seed.
- **A killed SSH client does not kill a remote `nohup`'d VM**, and the survivor holds the EFI variable store lock: "The EFI variable store is already in use" (`VZErrorDomain 50002`) means a zombie VM process, `pkill` it.
- **EFI variable stores are per-VM state.** Fresh stores boot reliably; reuse across different VM instances is at minimum suspect on this beta (Apple's own guidance for cloned VMs is one store per VM). Cheap policy: one store per case, created with the overlay, deleted with it.
- **Downloads from cloud.debian.org mirrors truncate silently**; always verify byte count against origin `Content-Length` and resume with `curl -C -`.
- The remote host's default shell is zsh: `=` -prefixed words (`echo ===`) explode via zsh's `=cmd` expansion; keep separators zsh-safe in automation.
### The 2-concurrent-macOS-VM cap: TESTED AND CONFIRMED on macOS 27 beta 4 (2026-07-29)
We measured it, which as far as we can tell nobody had published for macOS 27. Method: `cp -c -R` the guest bundle (APFS clonefile, instant and **zero additional disk**), regenerate the machine identifier per clone (`VZMacMachineIdentifier()` written to `machine.id`; the hardware model is reused), then launch VMs until one is refused.
Result: VM #1 (8 GB, GUI) and VM #2 (4 GB, headless) ran concurrently without complaint. VM #3 was refused **instantly** at `vm.start`:
```
VZErrorDomain Code=6 "The maximum supported number of active virtual machines has been reached."
NSLocalizedFailure = "The number of virtual machines exceeds the limit."
```
**This is a licensing/kernel quota, not a resource limit**: the refusal came with **39% of system memory free** on a 16 GB host, and adding RAM or CPU cannot raise it. It matches the pre-27 behavior (`hv_apple_isa_vm_quota`), so nothing changed in 27 despite the framework's other additions. Linux guests are unaffected and are bounded only by host resources.
Design consequences: macOS-guest capacity per host is **hard-capped at 2**, so a GUI-macOS-per-case product must schedule around it (queue, evict idle VMs, or scale across hosts) and surface it in the UI. Also relevant: the acknowledged slot-leak bug (a guest-initiated shutdown failing to release a slot, recoverable only by host reboot) is far more damaging under a cap of 2 than it sounds; we did not reproduce it on beta 4, but any scheduler should treat "slot appears used but nothing is running" as a real state.
### Not yet tested
- Cache layers (`LayerType.cache`), `.overlay(blockCount:)` disk growth, stack depth performance, VirtioFS + stack combination, `truncate`, ASIF disks for macOS guests (raw used so far; ASIF has the reclamation bug).
- One more scripting lesson from this session: inner `ssh` calls inside a piped `sh -s` script MUST use `-n`, or they consume the remainder of the script from stdin and it silently never runs.
### Session timeline (what was actually established, 2026-07-29)
Linux path: base image download (with resume, mirrors truncate) → `vzboot` compiles against the beta SDK first try → EFI boot → NAT DHCP lease → cloud-init seed provisions a user with the host's SSH key → `ssh` into the guest works → DiskImageKit stack boots with an ASIF overlay taking all writes while the base stays SHA-identical. Later Linux boots became unreliable on an un-rebooted host (silent hangs, 0% CPU, no DHCP); a clean-baseline retest is still pending.
macOS path: seed-CDN IPSW (matched to the host build) → `VZMacOSInstaller` restore, ~25 min, first try → first boot with `VZMacGuestProvisioningOptions` creates an admin account with Remote Login on, no interaction needed, SSH reachable ~140 s later → key bootstrap over a one-time password login → guest shutdown/relaunch clean (the slot-leak bug did not reproduce) → GUI access fought through a port collision, a client-version incompatibility, the rendering dependency, and a self-inflicted session kill, ending with a browser-based path plus a guest hardened to auto-login and never lock.
**Lifecycle verified (stop → start), 2026-07-30**: an in-guest `shutdown -h now` fires `guestDidStop` and the runner app exits on its own; relaunching from the same bundle boots the guest in ~2 minutes straight into an auto-logged-in desktop, and the VM slot is released cleanly (an immediate restart works, so the slot-leak bug did not bite). Two operational notes: the guest takes a **new NAT lease on every boot**, so any port-forward must be re-pointed (or use a vmnet DHCP reservation), and a host reboot resets `pmset -a disablesleep`.
⚠ **Provisioning does NOT skip the per-user first-login assistant.** `VZMacGuestProvisioningOptions` skips the initial Setup Assistant (account creation, region, Apple Account) so the machine is immediately reachable, but the first time anyone actually logs into a desktop, macOS still presents its per-user wizard (Apple Intelligence, Siri, privacy, appearance, Touch ID). The operator hit exactly this. For a GUI-first product this MUST be pre-suppressed during base-image creation by writing `com.apple.SetupAssistant` keys for every account that will log in, and into `/System/Library/User Template/English.lproj/Library/Preferences/` so accounts created later inherit it.
⚠ **A partial key list is worse than none**, because the wizard simply shows the panes you missed and the operator has to click through them again after every fresh login (we hit this twice). The set that finally silenced macOS 27 beta 4: `DidSeeCloudSetup`, `DidSeeSiriSetup`, `DidSeePrivacy`, `DidSeeAppearanceSetup`, `DidSeeTouchIDSetup`, `DidSeeAvatarSetup`, `DidSeeScreenTime`, `DidSeeApplePaySetup`, `DidSeeSafariImport`, `DidSeeAccessibility`, **`DidSeeActivationLock`, `DidSeeAppStore`, `DidSeeLockdownMode`** (the three easy to miss), plus the Express-Settings flags **`SkipExpressSettingsUpdating`** and **`SkipFirstLoginOptimization`**, and the version markers `LastSeenCloudProductVersion` / `LastSeenBuddyBuildVersion` / `PreviousSystemVersion` / `PreviousBuildVersion` matching the guest build. Verify afterwards by reading the domain back and checking that no `DidSee*` key is still `0`. Note these keys change between macOS releases, so base-image creation should re-verify per OS version rather than trust a hardcoded list.
## 9. Design implications for Codeman's VM subsystem
0. **GUI is a first-class mode, and for macOS guests it is the whole point (owner decision, 2026-07-29).** The subsystem serves real desktops, not only headless SSH boxes. macOS guests are GUI-only in practice (nothing renders without an attached view). Linux guests are supported in BOTH modes: GUI when the case wants a desktop, headless-over-SSH when it wants a cheap agent sandbox. The costs of the GUI path are in §8 "Display rendering": VMs as GUI apps in a live session, a host that never locks, guests that auto-login with their first-login wizard pre-suppressed, and the macOS concurrency cap as a real capacity limit.
1. **The macOS-specific liabilities are accepted costs, not reasons to avoid macOS guests**: provisioning is macOS-only and first-boot-only, ASIF space reclamation is broken for macOS guests on the beta (use RAW disks for macOS guests until fixed), and the 2-VM cap applies. Plan around each: RAW-backed macOS disks, provisioning baked into base-image creation, and capacity limits surfaced in the UI.
2. **Base immutability is not just hygiene, it is load-bearing**: DiskImageKit's UUID invalidation plus our sha-stability proof make a read-only shared base per image-generation the core artifact. Bases are built once (seed attached), then only ever opened `.readOnly` under per-case overlays.
3. **Seed ISOs are a base-build-time tool only.** Never attach a seed to a routine case boot (correlated with boot hangs on the beta, and semantically wrong anyway since cloud-init already ran).
4. **Per-case files**: overlay ASIF + EFI variable store live and die together with the case.
5. **Export = ship the layer chain** (base ref + overlay + manifest), not flatten; there is no flatten API. In-guest `dd` to a fresh image is the fallback for a true single-file export.
6. **Health checking must be lease/API based**, not serial/ping based, and Codeman's `codeman-vm status` should read `/var/db/dhcpd_leases` (or use vmnet DHCP reservations for deterministic per-case IPs, a macOS 26 API).
7. **Run `fstrim` periodically in Linux guests** (or mount with discard) so overlays stay sparse.
8. **Entitlements plist stays minimal** (exactly `com.apple.security.virtualization`) to dodge the AMFI/watchdog traps.
9. **Expect beta churn**: pin findings to build numbers (this doc: 26A5388g) and retest each beta; the framework binaries changed every beta so far.
10. **A macOS guest is only "ready" when its desktop is ready**, which is a stricter bar than "the VM booted". Readiness means: VM app running with a live WindowServer connection, guest auto-logged-in (not at a login or lock screen), first-login assistant suppressed, and the guest's screen sharing serving a non-black framebuffer. Health checks should sample the framebuffer for non-black content, because every failure mode in this session (headless run, locked host, dead WindowServer, locked guest, setup wizard) presents as a perfectly healthy-looking process with a black or useless screen.
10b. **Supervision must run as a root LaunchDaemon.** A user LaunchAgent cannot launch a GUI app into the Aqua session; its restarts fail silently (child dies instantly, empty log, supervisor reports success). Root + `launchctl asuser <uid> sudo -u <user> …` works and the launched process persists. This bit us on the first supervisor implementation and is easy to repeat.
11. **Remote-access plumbing belongs in the helper CLI, not in ad-hoc shell**: a `codeman-vm` implementation should own port selection (never 5900), forward lifecycle across guest IP changes (or better, vmnet DHCP reservations for stable per-case IPs), and a documented browser path, because every failure in this session came from hand-rolled plumbing rather than from the Virtualization APIs themselves.
12. **Never let control-plane connectivity depend on a GUI session** on a remote Mac host: prefer a Tailscale system service over the App Store app, and keep a LAN-adjacent peer able to relay as an out-of-band recovery path.
## Sources
Apple DocC JSON backend (diskimagekit, virtualization, vmnet trees; macOS 27 release notes) | WWDC26 session 224 https://developer.apple.com/videos/play/wwdc2026/224/ | eclecticlight.co ASIF/virtualization coverage | developer.apple.com/forums threads 839343 (CSIdentity bug), 830118 (cross-version restore), 830119 (VM-slot leak), 830383 (VM cap), 834822 + 831902 (USB entitlements), 822658 (vmnet loopback) | openai/tart issues 1261/1263/1268/1269/1285 | Spooky-Labs provisioning design doc | VirtualBuddy 2.2 release notes | lima-vm discussions | our own test transcripts on the testbed (`~/vm-lab/*.log`, this repo's session)
+190
View File
@@ -0,0 +1,190 @@
# Web tabs: two fixes (planned + implemented 2026-07-28)
Both found against the saved dashboard
`https://<your-host>.<your-tailnet>.ts.net:4000` (Bio-Hacking-Dashboard).
Kept because the root-cause analysis of the second one is not obvious from the
resulting diff.
Status: **both implemented and verified end-to-end.** The one deliberate
non-change is recorded at the bottom.
---
## Bug 1: saved URLs could not be deleted from the Run dropdown
### What happened
The "Web / URL" section of the Run dropdown listed every saved dashboard as a
single clickable row whose only action was "open". Deleting required opening the
dashboard as a tab, clicking the tab's gear, then Delete in the modal, so a URL
you no longer wanted open at all could not be removed without first opening it.
### What shipped
- `renderWebviewMenuItems()` (`src/web/public/webview-tabs.js`) now renders each
saved URL as a `.run-mode-row--web` flex row: the open button, a gear
(`showWebviewModal`), and an `x` (`deleteWebviewById`). Nested buttons are
invalid HTML, hence the wrapper rather than a button inside a button.
- `deleteWebview()` split into the modal entry point, the new row entry point
`deleteWebviewById(id)`, and the shared `_confirmAndDeleteWebview(id)`.
- Both side buttons call `event.stopPropagation()` so the click does not also
open the dashboard.
- The dropdown's outside-click handler (`session-ui.js`) closes when the click
target is not inside `#runModeMenu`, and the row is gone by the time the delete
resolves, so `deleteWebviewById` re-asserts `.active` on the menu. Verified in a
browser: deleting one of several URLs leaves you looking at the rest of the list.
- CSS in `styles.css` (`.run-mode-row--web`, `.run-mode-row-btn`) plus a larger
touch target in `mobile.css`. The side buttons are permanently visible rather
than hover-revealed, because this menu is used on touch.
No server change: `DELETE /api/webviews/:id` already existed, owner-scoped, and
already revoked the capability and broadcast `WebviewChanged`.
---
## Bug 2: images did not load in a proxied dashboard
### Reproduction (before the fix)
```
CAP=<from POST /api/webviews/<id>/open>
# A) upstream direct -> 200 image/jpeg 118150
curl -sk "https://<your-host>.<your-tailnet>.ts.net:4000/api/hero?slug=120-minutes-in-nature"
# B) through the proxy prefix -> 200 image/jpeg 118150
curl -sk "https://localhost:3000/webview/$CAP/api/hero?slug=120-minutes-in-nature"
# C) what the browser ACTUALLY requested -> 404 {"errorCode":"NOT_FOUND"}
curl -sk -H "Referer: https://localhost:3000/webview/$CAP/" \
"https://localhost:3000/api/hero?slug=120-minutes-in-nature"
# D) same shape but NOT under /api -> 200 (referer fallback rescues it)
curl -sk -H "Referer: https://localhost:3000/webview/$CAP/" "https://localhost:3000/styles.css"
```
The proxy itself was fine (B). The failure was entirely about which URL the
browser ended up requesting (C).
### Root cause
The dashboard builds its image markup at runtime with root-absolute URLs:
`c.innerHTML = '<img class="thumb" src="/api/hero?slug=...">'`, `img.src =
slideSrc(...)` returning `/api/slide?owner=...`, `/api/story`, `/api/video`, and a
nested `<iframe src="/api/preview?slug=...">`.
All three rewrite layers missed that shape:
1. `<base href="/webview/<cap>/">` only affects **relative** URLs. A root-absolute
`/api/hero` ignores the base path and resolves against Codeman's origin.
2. `rewriteHtml()` only runs over the **initial HTML document**. This markup is
created later by page script. (The static header `<img src="/api/logo">` DID
work, having been rewritten at proxy time, which is why only the
runtime-injected images were broken.)
3. `runtimeUrlShim()` patched only `fetch`, `XMLHttpRequest.open`, `WebSocket` and
`EventSource`, so the dashboard's **data** loaded while its **pictures** did
not.
The safety net was fenced off from `/api` in two places, both deliberate:
`server.ts`'s not-found handler returns the API-envelope 404 before reaching
`tryWebviewRefererFallback`, and `middleware/auth.ts` refuses the Referer-form
auth exemption for `/api/`, `/ws/`, `/q/`.
### What shipped
`runtimeUrlShim()` in `src/web/webview-proxy.ts` now also covers the DOM sinks, so
a root-absolute `/api/...` request is never emitted in the first place and neither
security fence had to move:
- `innerHTML` / `outerHTML` / `insertAdjacentHTML` (and `ShadowRoot.innerHTML`),
- `setAttribute` / `setAttributeNS`,
- the `src`/`srcset`/`href`/`poster`/`data`/`action` property setters on img,
source, media, video poster, script, iframe, embed, track, link, anchor, area,
object and form,
- a `MutationObserver` as a last net for any sink not patched above (it costs one
wasted 404 per node, since the browser starts fetching on insert, so it is a net
and not the mechanism).
Two details that mattered:
- Every rewrite routes through the existing idempotent `rw()` rather than a blind
prefix concat. The first draft used the server-side regex shape and
double-prefixed markup that was already proxied (a page re-injecting its own
`outerHTML`); the jsdom test caught it.
- Everything stays inside `try`/`catch` and is marked `__cmrw`, so a double
injection cannot wrap an already-wrapped setter, and nothing can throw into a
page we do not control.
### Verification
- `test/webview-proxy.test.ts` gained a jsdom `runtimeUrlShim DOM sinks` block:
innerHTML, insertAdjacentHTML, property setters, setAttribute, srcset candidate
lists, the MutationObserver net via an unpatched sink
(`createContextualFragment`), idempotence, re-injected markup, empty `src`, and
the pass-throughs (relative, cross-origin, `#hash`, `data:`). 73 tests pass.
- End-to-end in a real browser against an isolated instance
(`CODEMAN_INSTANCE=wvtest`, port 3151), with prod's old build as the negative
control:
| | before (prod, old build) | after (fixed) |
| --- | --- | --- |
| images found | 693 | 693 |
| src under the proxy prefix | 0 | 693 |
| in-viewport images decoded | 0 / 23 | 23 / 23 |
| sample src | `/api/hero?slug=...` | `/webview/<cap>/api/hero?slug=...` |
(The dashboard marks thumbs `loading="lazy"`, so only in-viewport images are
ever fetched. All 27 proxied image responses returned 200.)
---
## Follow-up (same day): the `/api` referer fallback, done safely
Originally deferred, then implemented on request. Both gates had to move, and the
auth one is the security-sensitive half: auth runs in `onRequest`, before routing,
so it cannot tell a real Codeman API route from a 404, and simply dropping the
`/api` fence would let a page holding a capability forge a `Referer` and reach
Codeman's **real** API unauthenticated.
What shipped:
- `server.ts`: `tryWebviewRefererFallback` is tried **before** the API-shaped 404.
Reaching that handler already proves no route matched, and the relay declines
unless the `Referer` carries a live capability, so unknown `/api` paths still
get the envelope.
- `middleware/auth.ts`: the `/api/` prefix refusal is replaced by
`matchesRegisteredRoute()`, which refuses the exemption for any path that
resolves to a real route. `/ws/` and `/q/` stay refused by prefix.
Two findings that decided the implementation, both established by probing Fastify
rather than by reading its docs:
- **`hasRoute()` is the wrong tool and would have been a hole.** It matches the
registered PATTERN literally, so `hasRoute({url: '/api/sessions/abc'})` returns
false against a registered `/api/sessions/:id` and would have handed out an
exemption on a live, session-scoped API route. `findRoute()` performs the real
radix-tree lookup and is what the fence uses.
- **`@fastify/static` is mounted at `/`, so it registers a root catch-all that
matches every path.** A match on it means "heading for the 404 handler", not
"real route", and it is distinguishable because a root catch-all is the only
route whose `*` param comes back equal to the whole request path. Without that
carve-out the fence would have refused every referer-form request and broken the
rescue that already worked.
The fence fails closed, and `test/webview-auth-exemption.test.ts` pins both edges
(a concrete URL onto a parametric API route stays 401; the dashboard's own
`/api/...` namespace is served).
### And the CSS gap, which the fallback could NOT close
Testing the fallback against a purpose-built upstream showed the runtime-injected
stylesheet case is unreachable by any relay: a `<style>` element has no URL of its
own, so Chromium sends an **empty `Referer`** with the image request it triggers
and there is nothing to key on. Measured directly:
| sink | Referer the browser sends | fixed by |
| --- | --- | --- |
| `url()` in a proxied `.css` | the stylesheet's proxied URL | the referer relay |
| `url()` in a runtime `<style>` | *empty* | `rwCss()` in the shim |
So the shim also rewrites `url()` inside `<style>` blocks, both when they arrive as
markup and when a `<style>` node is inserted (via the existing MutationObserver).
The only gap left is self-navigation via `location.href = '/x'`, which cannot be
patched because `Location.href` is unforgeable.
+179
View File
@@ -0,0 +1,179 @@
# Web Tabs (dashboards as Codeman tabs)
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port
4000, as a tab beside your Claude/Codex/Antigravity sessions. Codeman becomes one mission
control instead of Codeman plus a pile of browser tabs.
## Using it
1. Click the chevron next to **Run** to expand the dropdown.
2. Under **Web / URL**, pick **Add dashboard...**
3. Give it a name and a URL, optionally hit **Test**, then **Save**.
The dashboard opens as a tab immediately, and appears in the Run dropdown from then
on. Web tabs sit in the same strip as session tabs, continue the same `Alt+1..9`
numbering, and carry a globe icon so they never read as a running agent.
Closing a tab (the `x`) only closes it. The saved dashboard stays in the dropdown.
To delete it for good, use the `x` on its **dropdown row** (the tab's own `x` is
close, not delete). Each dropdown row also has a gear for editing, so a saved URL
can be changed or removed without opening it first.
Switching tabs does **not** reload a dashboard. Frames stay alive in the background,
so a dashboard that took a while to authenticate is still there when you come back.
Past six live frames the least-recently-viewed one is dropped to bound memory
(`CODEMAN_MAX_LIVE_WEBVIEW_FRAMES`).
## Why dashboards are proxied
A plain `<iframe src="http://your-box:4000">` does not work in the setup Codeman
actually ships in, for three separate reasons:
| Blocker | What happens |
| ------------------- | ---------------------------------------------------------------------------------------------- |
| **Mixed content** | Production serves HTTPS (behind `tailscale serve`). Browsers hard-block `http://` iframes on an HTTPS page, with no override, and none at all on iOS Safari. |
| **Framing refusal** | Grafana, Portainer, Home Assistant and many others send `X-Frame-Options: DENY` or `frame-ancestors 'none'`. |
| **Codeman's CSP** | `default-src 'self'` means `frame-src` falls back to `'self'`, so a cross-origin iframe is blocked before it starts. |
Serving the dashboard **through Codeman's own origin** dissolves all three. So by
default a web tab loads `/webview/<capability>/` on Codeman, and Codeman relays to
the dashboard: stripping the framing refusal, rewriting redirects, cookies and
root-absolute URLs, and relaying WebSockets so live panels actually update.
A useful consequence: the dashboard is fetched **from the Codeman server**, so a
tailnet-only or `localhost`-only dashboard works from any device that can reach
Codeman, including a phone that is not on the tailnet.
`direct` mode (a plain cross-origin iframe) still exists and is cheaper, but it only
works for an HTTPS dashboard that permits framing. The **Test** button probes from
the server and tells you which mode applies. Note what Test actually verifies:
**server-to-upstream reachability, nothing else**. It does not exercise the browser
sandbox, cookies, CORS, CSP, or any reverse proxy sitting in front of Codeman, so a
passing Test does not guarantee the embedded page will render (see the
cookie-authenticated reverse proxy caveat below).
## The sandbox, and when to turn it off
Because a proxied dashboard is served from Codeman's own address, it is
*same-origin with Codeman* as far as the browser is concerned. Left unchecked, its
JavaScript could read the Codeman page and call the API that spawns agents.
So the iframe is sandboxed **without** `allow-same-origin` by default. The page runs
in an opaque origin: it cannot touch Codeman, and it gets no cookies or
`localStorage` of its own.
Unchecking **Open sandboxed** grants `allow-same-origin`. Do that only for a
dashboard you fully trust, and only if you need it, which in practice means a
dashboard with its own login that stores a session in a cookie or `localStorage`.
Even in trusted mode, Codeman never forwards its own credentials upstream: the
`Authorization` header and the `codeman_session` cookie are stripped on the way out,
so `CODEMAN_PASSWORD` cannot leak into a dashboard.
⚠️ **Sandboxed tabs may not work when Codeman itself is behind a
cookie-authenticated reverse proxy** (Cloudflare Access, Authelia, oauth2-proxy and
similar). The sandboxed frame is opaque-origin, so its stylesheet, script, and API
requests do not carry the proxy's authentication cookie; the proxy redirects them to
the login provider, where CORS/CSP kills them, and the embedded app renders
unstyled or broken while the Codeman page around it works fine. Trusted mode
(**Open sandboxed** off) keeps a real origin and the cookie, so it works. The
**Test** button cannot catch this: it checks that the Codeman *server* can reach the
upstream, not that a sandboxed *browser* frame can load assets through the public
authentication layer.
## How the proxy authenticates
A sandboxed iframe is opaque-origin, so every request it makes is cross-site: the
`SameSite=lax` session cookie is not sent, and writes and WebSocket upgrades arrive
with `Origin: null`. Cookie auth cannot work.
Instead, opening a dashboard mints a **capability**: 192 bits of entropy in the URL
path, held in memory only, with a rolling 12-hour TTL, bound to the user who minted
it, and granting exactly one thing, relaying bytes to that one saved URL. Editing or
deleting a dashboard revokes it, and a server restart invalidates every outstanding
capability (tabs re-mint transparently on next click).
## Limits and env vars
| Variable | Default | Meaning |
| ------------------------------------ | ------- | ------------------------------------------ |
| `CODEMAN_MAX_WEBVIEWS` | 50 | Saved dashboards per owner |
| `CODEMAN_MAX_LIVE_WEBVIEW_FRAMES` | 6 | Iframes kept mounted at once |
| `CODEMAN_WEBVIEW_CAPABILITY_TTL_MS` | 12h | Rolling capability lifetime |
| `CODEMAN_WEBVIEW_TIMEOUT_MS` | 30000 | Upstream request timeout |
| `CODEMAN_WEBVIEW_PROBE_TIMEOUT_MS` | 8000 | Timeout for the Test button |
| `CODEMAN_MAX_WEBVIEW_HTML_BYTES` | 8MB | Largest HTML document rewritten |
| `CODEMAN_MAX_WEBVIEW_SOCKETS` | 8 | Concurrent proxied WebSockets per dashboard |
Saved dashboards live in `~/.codeman/webviews.json`. Which tabs you have open is
per-device (`localStorage`), since that is workspace layout rather than config.
## How a dashboard's own API calls keep working
Worth knowing, because it is where this feature does its least obvious work. Three
layers cooperate so a dashboard talking to its own backend just works:
1. `<base href>` handles relative URLs in the markup.
2. Attribute rewriting handles root-absolute `src`/`href`/`action` in the page the
proxy serves.
3. A small injected script rebases URLs built at **runtime**, which the first two
cannot see: `fetch('/api/data')` and `new WebSocket('/live')`, but equally
`card.innerHTML = '<img src="/api/hero">'`, `img.src = '/api/slide'`, and
`url(/img.png)` inside a `<style>` the page injects. That second group is why
images are covered too. A dashboard that renders its thumbnails from script
would otherwise show all its data and none of its pictures, because `<base>`
does not apply to root-absolute URLs and the attribute rewriting only ever saw
the initial document.
4. As a last resort, a request that still lands on Codeman's own root is relayed
using its `Referer` to identify the dashboard. This only fires for a request
that already missed every Codeman route, and never for one that resolves to a
real route, which is what keeps it from being an authentication bypass.
On top of that, the proxy answers those requests with CORS headers. That sounds
wrong for same-host requests, but a sandboxed iframe has an *opaque* origin, so the
browser treats every one of its `fetch`/XHR calls as cross-origin even though the
URL is on Codeman itself. Without those headers, a dashboard renders perfectly and
then every API call fails, which looks like the dashboard being broken.
## Known limits
- **Exotic loaders.** The layers above cover normal `fetch`/XHR/WebSocket/
EventSource, normal markup, the DOM sinks a page uses to build markup at runtime,
and `url()` inside stylesheets. Something that constructs requests by an unusual
route can still slip through. Symptom: the page renders but a panel stays empty.
- **Root-absolute `location` navigation.** A dashboard that navigates itself with
`location.href = '/login'` escapes the prefix, because `Location.href` is
unforgeable and cannot be patched the way the other sinks are. A relative
`location.href = 'login'` is fine (`<base>` covers it).
- **Cross-origin redirects are not followed.** If a dashboard bounces to a different
host (an external SSO provider, say), the proxy hands the redirect back unchanged
rather than relaying it, because relaying would make this an open proxy. Use
**Open in new tab** for those.
- **Login-protected dashboards need trusted mode**, since a sandboxed frame has no
cookie jar. A server-side per-dashboard cookie jar would lift this and is the
natural next step if it becomes annoying.
- **Cookie-authenticated reverse proxies in front of Codeman break sandboxed tabs**
(#238). The sandboxed frame's requests carry no auth cookie, so the proxy bounces
them to its login provider and the app loads broken while Test reports reachable.
Use trusted mode behind Cloudflare Access and friends; see the warning above.
- **Slow endpoints and the upstream timeout** (#237). The proxy waits
`CODEMAN_WEBVIEW_TIMEOUT_MS` (default 300s) for the upstream's response *headers*,
then streams the body without any time bound; a header timeout is logged
server-side and answered as a 502 that names the limit. WebSocket handshakes use
the separate `CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS` (default 30s).
- **Not a security boundary.** The proxy reaches whatever the Codeman server can
reach. That is not an escalation for someone who already commands
`--dangerously-skip-permissions` agents, but in multi-user mode it does mean a
non-admin user's dashboard is fetched from the server's network position.
## Where the code lives
| Concern | File |
| ------------------------ | --------------------------------------- |
| Pure rewrite helpers | `src/web/webview-proxy.ts` |
| Routes + proxy + sockets | `src/web/routes/webview-routes.ts` |
| Capability tokens | `src/webview-capabilities.ts` |
| Persistence | `src/webview-store.ts` |
| Limits | `src/config/webview-limits.ts` |
| Frontend | `src/web/public/webview-tabs.js` |
| Auth exemption | `src/web/middleware/auth.ts` |
+1150 -56
View File
File diff suppressed because it is too large Load Diff
+4 -3
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.4.0",
"version": "1.14.2",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.4.0",
"version": "1.14.2",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -34,6 +34,7 @@
"qrcode": "^1.5.4",
"uuid": "^14.0.0",
"web-push": "^3.6.7",
"ws": "^8.21.0",
"zod": "^4.3.6"
},
"bin": {
@@ -12332,7 +12333,7 @@
}
},
"packages/xterm-zerolag-input": {
"version": "0.1.4",
"version": "0.1.8",
"license": "MIT",
"devDependencies": {
"jsdom": "^24.1.3",
+31 -7
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.4.0",
"version": "1.14.2",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -22,6 +22,7 @@
"test:coverage": "vitest run --config config/vitest.config.ts --coverage",
"test:ci": "vitest run --config config/vitest.ci.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
@@ -32,25 +33,45 @@
"changeset": "changeset",
"version-packages": "changeset version && npm install --package-lock-only && node scripts/check-lockfile-sync.mjs",
"check:lockfile": "node scripts/check-lockfile-sync.mjs",
"knip": "npx --yes knip@latest",
"knip": "npx --yes knip@latest --config config/knip.json",
"release": "changeset publish"
},
"prettier": {
"singleQuote": true,
"semi": true,
"tabWidth": 2,
"printWidth": 120,
"trailingComma": "es5",
"endOfLine": "lf"
},
"workspaces": [
".",
"packages/*"
],
"keywords": [
"claude",
"claude-code",
"claude-ai",
"claude",
"anthropic",
"ai-agent",
"automation",
"opencode",
"codex",
"antigravity",
"gemini-cli",
"ai-agents",
"agent",
"session-manager",
"self-hosted",
"developer-tools",
"tmux",
"terminal",
"xterm",
"docker",
"mosh",
"local-echo",
"web-dashboard",
"cli",
"llm",
"autonomous-agent",
"ralph-loop"
"automation"
],
"author": "arkon",
"license": "MIT",
@@ -75,6 +96,7 @@
"qrcode": "^1.5.4",
"uuid": "^14.0.0",
"web-push": "^3.6.7",
"ws": "^8.21.0",
"zod": "^4.3.6"
},
"devDependencies": {
@@ -136,6 +158,8 @@
"files": [
"dist",
"scripts/postinstall.js",
"scripts/fix-node-pty.mjs",
"skills",
"LICENSE",
"README.md"
]
+64
View File
@@ -1,5 +1,69 @@
# xterm-zerolag-input
## 0.1.8
### Patch Changes
- **Fixed: sessions failed to start on macOS with `Error: posix_spawnp failed.`** (issues #6 and #204)
`node-pty@1.1.0` publishes its macOS prebuilt helper as `prebuilds/darwin-<arch>/spawn-helper` with mode 0644, i.e. no execute bit. macOS launches every PTY through that helper, so a stock install failed on every session start. The bug is macOS-only: `spawn-helper` is a mac-only gyp target and node-pty ships no Linux prebuild, so Linux always compiles a correctly-permissioned helper from source.
The previous fix chmodded only `build/Release/spawn-helper`, which on macOS does not exist (the prebuild is used, so node-gyp never runs), and it derived that path from `require.resolve('node-pty')`, landing on `<pkg>/lib/build/Release/...`. It was a no-op on every platform.
- New `scripts/fix-node-pty.mjs` (also `npm run fix:node-pty`) chmods every `spawn-helper` it finds, in `build/Release`, `build/Debug` and each `prebuilds/*/`, then verifies the result by actually opening a PTY. A `require()` alone passes on a broken install, because the helper is only touched at spawn time.
- `postinstall` no longer force-rebuilds node-pty from source on Node 22+. That step needed Xcode command line tools, cost 30-120s on every install, and deleted the `prebuilds/` tree before compiling, so a Mac without a compiler was left with no working binary at all. A rebuild now happens only when the chmod plus spawn probe still fails, and the prebuilds tree is backed up and restored around it.
- New `spawnPtyWithHelperRepair()` (`src/utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts`, so an install that is already broken repairs itself on the first failed spawn and retries in-process instead of showing a dead session. Unrelated spawn errors are rethrown untouched; a second failure carries the `npm run fix:node-pty` hint.
- `scripts/fix-node-pty.mjs` is now in the published `files` list, so global npm installs get the repair too.
- Direct-PTY Claude spawns use the resolved absolute binary path (new `getClaudeBinaryPath()`) instead of the bare name `claude`, so a CLI installed outside the server's PATH still launches.
Verified end to end on macOS 26.4 arm64: a stock `npm i` reproduces `posix_spawnp failed.`, and after the fix the same install spawns a PTY successfully with the prebuilds preserved.
**Added: phone home screen (session overview)**
Under 430px the "C" logo now opens a session overview (current sessions, past sessions, spaces) instead of the welcome overlay: on a small screen "which session needs me" beats "how do I start one". Rows resume a session in place, and "New session here" goes through the normal quick-start path so remote and Docker cases keep their routing. Per-device setting `mobileOverviewEnabled` (phones only, default ON) in App Settings. Tablet and desktop are unchanged.
**Added: guided Tailscale setup in `install.sh`**
The network-access prompt is now 3-way: Tailscale, LAN, or local-only. The Tailscale path binds loopback and walks through installing Tailscale, logging in, the operator grant, the tailnet HTTPS-certificates toggle, and `tailscale serve --bg <port>`, then verifies the result end to end with curl. That gives HTTPS on a real certificate with no app password and no `0.0.0.0` bind, which is also what PWA install and web push need. `install.sh tailscale` retrofits it onto an existing install, and `CODEMAN_TAILSCALE=1` presets the choice. Serve state is detected from `tailscale serve status --json`; the installer never runs `tailscale serve reset` and never touches serve mappings other than 443 to Codeman's port. README and `docs/security-architecture.md` updated to match.
**Docs**: replaced a real tailnet hostname with placeholders in `docs/web-tabs-fixes-plan.md`.
**xterm-zerolag-input**: npm description and keywords only, no code change.
## 0.1.7
### Patch Changes
- Fix a latent bug where a partial settings PUT silently reset live service state, and trim the `xterm-zerolag-input` README callout.
- **`PUT /api/settings` no longer resets watchers on a partial body.** The three `toggleService` calls (subagent watcher, workflow-run watcher, image watcher) read the raw request body with `??` defaults, so every key a caller omitted was treated as "apply the default". A body of just `{statusLineTelemetry:true}` would START the subagent watcher and STOP the workflow and image watchers, undoing the persisted config. They now resolve from `merged` (persisted settings + incoming), the same convention the `tmuxHistoryLimit` branch in that handler already used, so any PUT reconciles services to the effective stored state. Nothing triggered this in practice because every shipped client sends a full settings payload rebuilt from the DOM, but it was a trap for the next partial-update caller.
- **Regression test**: `test/routes/system-routes-settings-partial-put.test.ts` (4 cases) pins both directions, omitted keys preserve state and explicit keys still take effect. Verified to fail against the pre-fix handler.
- **CLAUDE.md** records the rule under "Adding Features → App setting": anything acting on a setting in that handler must resolve from `merged`, never the request body.
- **`xterm-zerolag-input` README**: removed the links line (getcodeman.com / install one-liner / star link) from the Codeman callout above the demo GIF. The callout keeps its links in the heading and body.
## 0.1.6
### Patch Changes
- Plan-usage chip now defaults ON on desktop, plus the reworked `xterm-zerolag-input` README.
- **Plan-usage chip defaults ON (desktop).** The `showPlanUsageLimits` chip (live 5-hour and weekly plan usage from the Claude statusline) used to be opt-in and default OFF, so most users never saw it. Desktop now defaults ON; handhelds still default OFF so the phone header stays minimal and the `mobile-header-buttons-policy` guard keeps passing. Devices with an explicitly stored preference keep whatever they chose, so nobody's OFF gets overridden.
- **One resolver behind the chip.** Added `planUsageChipEnabled()` in settings-ui.js and routed all three call sites through it: the App Settings checkbox, the chip's visibility, and the create-time `statusLineTelemetry` flag in session-ui.js. Those three had independent `?? false` / `=== true` defaults, and a chip revealed without the telemetry flag renders `—` forever, so a default flip on one site alone would have shipped a permanently empty chip.
- **Cron button comment corrected.** The App Settings comment claimed "Cron button defaults ON" while the code, the template (`btn-cron--hidden`) and the CSS all default it OFF. Verified against a fresh browser profile: the button is hidden and its checkbox unchecked out of the box. Comment now matches, and states why the two halves stay consistent.
- **Docs.** CLAUDE.md, `docs/architecture-invariants.md` and `docs/usage-limits-display-plan.md` updated for the new default and the single-resolver rule; the stale `styles.css` comment claiming the server strips the chip's hidden class at render was corrected (display is per-device, so the client reveals it).
- **`xterm-zerolag-input` README rework** (0.1.5 shipped the content; this republishes with the graphic and promo changes): replaced the misaligned 8-line keystroke-flow diagram with a two-line stock-vs-zerolag contrast, added a Codeman callout above the demo GIF with links to getcodeman.com and the repo, and rewrote the Origin section so it argues the extraction story instead of repeating the promo.
## 0.1.5
### Patch Changes
- Rewrite the `xterm-zerolag-input` package README as a value-first document and correct the drift that had accumulated against the source.
- Added the side-by-side phone demo GIF (`docs/images/zerolag-demo-20260728.gif`) as the hero image, referenced by absolute raw URL so it renders on npmjs.com as well as GitHub. The two-phone comparison shows 0ms local echo next to a 600ms-2.7s server echo on the same session.
- New "Why this one" comparison table, an explicit list of target use cases (SSH web clients, cloud IDEs, mobile terminals, container consoles), and a bundle-size badge (6.1 kB gzipped, measured from the ESM build).
- Corrected the test-count badge from 78 to the actual 175 tests across 5 files, in both the package README and the Published Packages section of the root README.
- Removed the stale "Unicode/emoji rendered at single-cell width" limitation. CJK, fullwidth forms and emoji have had double-width rendering and visual-column positioning since the wide-character fix; the honest remaining caveat (per-code-point width summing over-counts ZWJ grapheme clusters) replaces it.
- Documented the previously undocumented public `setPrompt()` method for switching prompt strategies at runtime, and the new "Wide characters (CJK, emoji)" integration section covering the optional `Unicode11Addon` path and the built-in range-table fallback.
- Documented `backgroundColor: 'transparent'`, corrected the `foregroundColor` default, and updated the grid-alignment math to reflect visual-column positioning rather than character index.
No source changes, docs only.
## 0.1.4
### Patch Changes
+162 -105
View File
@@ -1,45 +1,64 @@
<p align="center">
<h1 align="center">xterm-zerolag-input</h1>
<p align="center">
Instant keystroke feedback overlay for <a href="https://xtermjs.org/">xterm.js</a><br>
<em>Eliminates perceived input latency over high-RTT connections</em>
<strong>Make typing feel instant in <a href="https://xtermjs.org/">xterm.js</a>, no matter how far away the server is.</strong><br>
<em>A pixel-perfect local echo overlay. Client-side only. Zero dependencies.</em>
</p>
<p align="center">
<a href="https://www.npmjs.com/package/xterm-zerolag-input"><img src="https://img.shields.io/npm/v/xterm-zerolag-input?style=flat-square&color=22c55e" alt="npm"></a>
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="MIT"></a>
<img src="https://img.shields.io/badge/Dependencies-0-22c55e?style=flat-square" alt="Zero deps">
<img src="https://img.shields.io/badge/Tests-78-22c55e?style=flat-square" alt="78 tests">
<img src="https://img.shields.io/badge/xterm.js-v5%20%7C%20v7+-3b82f6?style=flat-square" alt="xterm.js">
<img src="https://img.shields.io/badge/Dependencies-0-22c55e?style=flat-square" alt="Zero dependencies">
<img src="https://img.shields.io/badge/Size-6.1%20kB%20gzip-22c55e?style=flat-square" alt="6.1 kB gzipped">
<img src="https://img.shields.io/badge/Tests-175-22c55e?style=flat-square" alt="175 tests">
<img src="https://img.shields.io/badge/xterm.js-v5%20%7C%20v7+-3b82f6?style=flat-square" alt="xterm.js v5 and v7+">
</p>
</p>
> ### Made for [**Codeman**](https://getcodeman.com)
>
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Antigravity sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
>
> That last part is why this library exists. The demo below is a real Codeman session on two phones.
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/zerolag-demo-20260728.gif" alt="Side-by-side phones typing into the same remote session: with zerolag the text appears at 0ms, without it every keystroke waits 600ms to 2.7s for the server echo" width="900">
</p>
<p align="center">
<em>Two phones, the same remote session, the same slow link.<br>
Left: the zerolag overlay paints every keystroke at <strong>0ms</strong>. Right: stock xterm.js waits <strong>600ms to 2.7s</strong> for the server to echo it back.</em>
</p>
---
## The Problem
## The 30-second version
When using xterm.js over a remote connection (SSH web clients, cloud IDEs, mobile terminals), every keystroke takes a full round-trip to the server before appearing on screen. At 100-500ms RTT, typing feels sluggish and unresponsive. Users type blind, make mistakes they can't see, and the experience feels broken.
## The Solution
`xterm-zerolag-input` renders typed characters **immediately** as a pixel-perfect DOM overlay positioned on the terminal's character grid. The overlay covers the terminal canvas at the prompt location, showing characters instantly while the server echo travels back. Once the server responds, the overlay seamlessly disappears and the real terminal text takes over.
Over a remote connection, xterm.js shows you a character only after it has flown to the server and back. At 100-500ms RTT that reads as broken: you type ahead of the screen, you cannot see your typos, and you start pecking one key at a time to stay in sync.
```
Keystroke Flow:
┌─── DOM overlay (instant, 0ms)
User types 'h' ─── onData('h') ───┤
└─── Your app sends to PTY ──→ Server
│
Server echoes 'h' ←──────────────────────────────────────────────────┘
│ (200-500ms RTT)
└──→ terminal.write('h') ──→ overlay.clear()
(server output replaces overlay — seamless transition)
stock xterm.js keypress ─────── 300 ms ───────→ character appears
with zerolag keypress → character appears · echo lands later, unseen
```
**No changes to your backend needed.** The addon is purely client-side.
Same keystroke, same link. The only difference is who you wait for: the server, or nobody.
## Origin
`xterm-zerolag-input` paints your keystrokes **immediately**, as an absolutely-positioned DOM overlay locked to the terminal's character grid. The byte still goes to the PTY exactly as before, so nothing about your shell changes. When the server echo lands 300ms later, the overlay clears and the real terminal text takes over on the same pixels. The handoff is invisible.
This library was extracted from [Codeman](https://github.com/Ark0N/Codeman), mission control for AI coding agents — multi-session management, real-time agent visualization, autonomous respawn loops, and a mobile-first web UI for Claude Code, OpenCode, and Codex. The local echo system was built to make mobile and remote access feel instant, then battle-tested across thousands of hours of real usage. After 3 deep code audits, it was extracted into this standalone library with 78 tests covering every state transition.
**No backend changes. No protocol. No server support.** It is a client-side addon that never touches the wire.
## Why this one
| | |
|---|---|
| **Survives full-screen TUIs** | Ink, blessed, and friends repaint the whole screen constantly. The overlay is a separate DOM layer they cannot reach, so it does not get clobbered. |
| **Pixel-matched to the canvas** | Each character is its own absolutely-positioned `<span>` at exact cell coordinates, so it does not drift out of the grid like normal DOM text flow. |
| **Wide characters included** | CJK, fullwidth forms and emoji render double-width and position by visual column, using the terminal's Unicode addon when one is loaded. |
| **Backspace that actually works** | A three-layer cascade (unsent, in-flight, already on screen) tells you exactly what to forward to the PTY, so editing works through any mix of typed, flushed and tab-completed text. |
| **You keep control of input** | The addon never hooks `onData` for you. You decide what gets echoed and what gets forwarded, which is what makes char-at-a-time, buffered, and multi-session tab switching all possible. |
| **Small and self-contained** | 6.1 kB gzipped, zero runtime dependencies, dual CJS/ESM with full type declarations. |
| **Proven under load** | Extracted from [Codeman](https://getcodeman.com), hardened over thousands of hours of real remote and mobile usage, 175 tests over every state transition. |
Built for anything that puts a terminal behind a network hop: SSH web clients, cloud IDEs, mobile terminals, Kubernetes and container consoles, remote agent dashboards, browser-based dev environments.
## Install
@@ -47,12 +66,9 @@ This library was extracted from [Codeman](https://github.com/Ark0N/Codeman), mis
npm install xterm-zerolag-input
```
- **Zero runtime dependencies**
- Compatible with both `xterm` (pre-5.4) and `@xterm/xterm` (5.4+)
- Dual CJS/ESM build with full TypeScript declarations
- Works with canvas, WebGL, and DOM renderers
Works with both `xterm` (pre-5.4) and `@xterm/xterm` (5.4+), and with the canvas, WebGL and DOM renderers.
## Quick Start
## Quick start
```typescript
import { Terminal } from '@xterm/xterm';
@@ -61,7 +77,7 @@ import { ZerolagInputAddon } from 'xterm-zerolag-input';
const terminal = new Terminal();
terminal.open(document.getElementById('terminal')!);
// 1. Create addon with your prompt character
// 1. Create the addon with your prompt character
const zerolag = new ZerolagInputAddon({
prompt: { type: 'character', char: '$', offset: 2 },
});
@@ -75,7 +91,7 @@ terminal.onData((data) => {
ws.send(text + '\r');
} else if (data === '\x7f') {
const source = zerolag.removeChar();
if (source === 'flushed') ws.send(data); // only backspace text already in PTY
if (source === 'flushed') ws.send(data); // only backspace text already in the PTY
} else if (data.length === 1 && data.charCodeAt(0) >= 32) {
zerolag.addChar(data);
}
@@ -87,26 +103,29 @@ terminal.onWriteParsed(() => {
});
```
## Why This Is Hard
That is the whole integration. Everything below is for tuning it.
Most terminal UIs can't do local echo because:
## Why this is hard
1. **Buffer writes corrupt**: Frameworks like [Ink](https://github.com/vadimdemedes/ink) (React for terminals) redraw the entire screen on every state change. Writing directly to the terminal buffer gets immediately overwritten.
Most terminal UIs cannot do local echo, for three reasons:
2. **Cursor position lies**: In Ink, `buffer.cursorY` reflects internal state (near the status bar), not the visible prompt. You can't trust it.
1. **Buffer writes get corrupted.** Frameworks like [Ink](https://github.com/vadimdemedes/ink) (React for terminals) redraw the entire screen on every state change. Anything written straight into the terminal buffer is overwritten immediately.
3. **Font matching**: Canvas/WebGL renderers use their own text shaping. A DOM overlay must pixel-match the canvas grid — normal DOM text flow drifts due to sub-pixel glyph width differences.
2. **Cursor position lies.** In Ink, `buffer.cursorY` reflects internal render state (often near a status bar), not the visible prompt. You cannot trust it.
This library solves all three by:
- Using a **DOM overlay** that Ink can't touch (separate z-index layer)
- **Scanning the buffer** bottom-up for the prompt character instead of trusting cursor position
- Rendering each character as an **absolutely-positioned `<span>`** at exact cell-grid coordinates
3. **Fonts do not line up.** Canvas and WebGL renderers do their own text shaping. A DOM overlay has to pixel-match that grid, and normal DOM text flow drifts as sub-pixel glyph widths accumulate.
This library answers all three:
- a **DOM overlay** on its own z-index layer, which Ink cannot touch
- **bottom-up buffer scanning** for the prompt instead of trusting the cursor
- **one absolutely-positioned `<span>` per character** at exact cell-grid coordinates
---
## Prompt Detection
## Prompt detection
The addon needs to know where user input starts. It scans the terminal buffer bottom-up for the prompt. Three strategies:
The addon needs to know where user input starts. It scans the terminal buffer bottom-up. Three strategies:
### Character (default)
@@ -118,17 +137,17 @@ The addon needs to know where user input starts. It scans the terminal buffer bo
{ type: 'character', char: '%', offset: 2 }
// Fish / Starship: ❯
{ type: 'character', char: '\u276f', offset: 2 }
{ type: 'character', char: '❯', offset: 2 }
// Simple arrow: >
{ type: 'character', char: '>', offset: 2 }
```
`offset` = characters between the prompt marker and where user input begins (e.g., `"$ "` = 2).
`offset` = characters between the prompt marker and where user input begins (`"$ "` = 2).
### Regex
For complex prompts. The `g` flag is safely stripped to prevent `lastIndex` mutation.
For complex prompts. The `g` flag is stripped safely, so there is no `lastIndex` mutation.
```typescript
{ type: 'regex', pattern: /\$\s*$/, offset: 2 }
@@ -150,77 +169,88 @@ Full control:
}
```
### Switching prompts at runtime
If one terminal hosts several CLIs with different prompts, swap the strategy in place:
```typescript
zerolag.setPrompt({ type: 'character', char: '❯', offset: 2 });
```
`setPrompt()` clears the cached prompt position and re-renders if anything is pending, so a mode switch cannot leave the overlay pinned to the old column.
---
## API Reference
## API reference
### `ZerolagInputAddon`
Implements xterm.js `ITerminalAddon`. The addon does **not** hook `terminal.onData()` — you wire your own input handler and call these methods. This gives you full control over which keystrokes are echoed vs forwarded.
Implements the xterm.js `ITerminalAddon` interface. It deliberately does **not** hook `terminal.onData()`: you wire your own handler and call these methods, which is what gives you control over which keystrokes are echoed and which are forwarded.
### Input
| Method | Returns | Description |
|--------|---------|-------------|
| `addChar(char)` | `void` | Add a single printable character. Auto-detects existing buffer text on first keystroke. |
| `addChar(char)` | `void` | Add a single printable character. Auto-detects existing buffer text on the first keystroke. |
| `appendText(text)` | `void` | Append multiple characters (paste). |
| `removeChar()` | `'pending'` \| `'flushed'` \| `false` | Remove last char. See [backspace handling](#backspace-handling). |
| `clear()` | `void` | Clear all state, hide overlay. Call on Enter/Ctrl+C/Escape. |
| `removeChar()` | `'pending'` \| `'flushed'` \| `false` | Remove the last character. See [backspace handling](#backspace-handling). |
| `clear()` | `void` | Clear all state and hide the overlay. Call on Enter, Ctrl+C, Escape. |
### Backspace Handling
### Backspace handling
`removeChar()` cascades through three layers and tells you what it removed:
| Return | Source | Your action |
|--------|--------|-------------|
| `'pending'` | Unsent text (never transmitted to PTY) | Do nothing |
| `'flushed'` | Text already sent to PTY | Send `\x7f` backspace to PTY |
| `'pending'` | Unsent text (never transmitted to the PTY) | Do nothing |
| `'flushed'` | Text already sent to the PTY | Send `\x7f` to the PTY |
| `false` | Nothing to remove | Do nothing |
The cascade: pending text first, then flushed text, then auto-detect buffer text (handles tab completion). This means backspace "just works" through any combination of typed, flushed, and tab-completed text.
The cascade order is pending text, then flushed text, then auto-detected buffer text (which is what makes backspace work after tab completion). Backspace "just works" across any combination of typed, in-flight and completed text.
### Flushed Text
### Flushed text
"Flushed" = sent to PTY but echo hasn't arrived yet. Happens during tab switches and tab completion.
"Flushed" means sent to the PTY but the echo has not arrived yet. This happens during tab switches and tab completion.
| Method | Description |
|--------|-------------|
| `setFlushed(count, text, render?)` | Mark text as flushed. Pass `render=false` during tab-switch restore (buffer not loaded yet). |
| `setFlushed(count, text, render?)` | Mark text as flushed. Pass `render=false` during tab-switch restore, when the buffer is not loaded yet. |
| `getFlushed()` | Returns `{ count, text }`. |
| `clearFlushed()` | Clear flushed state when server echo arrives. |
| `clearFlushed()` | Clear flushed state once the server echo arrives. |
### Buffer Detection
### Buffer detection
Scan the terminal for text that exists after the prompt but wasn't typed through the overlay.
Finds text that exists after the prompt but was never typed through the overlay.
| Method | Description |
|--------|-------------|
| `detectBufferText()` | Scan and return detected text (or `null`). Sets it as flushed. Guarded: runs once per `clear()` cycle. |
| `detectBufferText()` | Scan and return the detected text (or `null`), marking it flushed. Guarded: runs once per `clear()` cycle. |
| `resetBufferDetection()` | Re-enable detection. |
| `suppressBufferDetection()` | Block detection until next `clear()`. Use for sessions with UI framework text after the prompt. |
| `undoDetection()` | Undo last detection — clears flushed state, re-enables detection. For tab completion retry. |
| `suppressBufferDetection()` | Block detection until the next `clear()`. Use for sessions that render UI framework text after the prompt. |
| `undoDetection()` | Undo the last detection: clears flushed state and re-enables detection. For tab-completion retries. |
### Rendering
| Method | Description |
|--------|-------------|
| `rerender()` | Force re-render. Call after buffer reloads, screen redraws, resizes, reconnects. |
| `refreshFont()` | Re-cache font properties from terminal. Call after font size or theme changes. |
| `rerender()` | Force a re-render. Call after buffer reloads, screen redraws, resizes and reconnects. |
| `refreshFont()` | Re-cache font and color properties from the terminal. Call after a font size or theme change. |
### Prompt Utilities
### Prompt
| Method | Description |
|--------|-------------|
| `findPrompt()` | Find prompt position. Returns `{ row, col }` or `null`. |
| `readPromptText()` | Read text after prompt marker. Returns string or `null`. |
| `setPrompt(finder)` | Replace the prompt detection strategy at runtime. |
| `findPrompt()` | Find the prompt position. Returns `{ row, col }` or `null`. |
| `readPromptText()` | Read the text after the prompt marker. Returns a string or `null`. |
### State
| Property | Type | Description |
|----------|------|-------------|
| `pendingText` | `string` | Unacknowledged text (read-only) |
| `hasPending` | `boolean` | `true` if overlay has any content |
| `state` | `ZerolagInputState` | Full snapshot: pendingText, flushedLength, flushedText, visible, promptPosition |
| `hasPending` | `boolean` | `true` if the overlay has any content |
| `state` | `ZerolagInputState` | Full snapshot: `pendingText`, `flushedLength`, `flushedText`, `visible`, `promptPosition` |
### Options
@@ -228,23 +258,23 @@ Scan the terminal for text that exists after the prompt but wasn't typed through
{
prompt?: PromptFinder, // Default: { type: 'character', char: '>', offset: 2 }
zIndex?: number, // Default: 7
backgroundColor?: string, // Default: from terminal theme
foregroundColor?: string, // Default: from computed .xterm-rows style
backgroundColor?: string, // Default: terminal theme background ('transparent' to disable)
foregroundColor?: string, // Default: terminal theme / computed .xterm-rows style
showCursor?: boolean, // Default: true
cursorColor?: string, // Default: from terminal theme
cursorColor?: string, // Default: terminal theme cursor
scrollDebounceMs?: number, // Default: 50
}
```
---
## Integration Patterns
## Integration patterns
### Buffered Input (hold until Enter)
### Buffered input (hold until Enter)
The quick start example above. Characters accumulate in the overlay and are sent on Enter. Best for remote shells where you want to batch input.
The quick start above. Characters accumulate in the overlay and go out on Enter. Best for remote shells where you want to batch input.
### Char-at-a-Time (send immediately)
### Char-at-a-time (send immediately)
```typescript
terminal.onData((data) => {
@@ -256,12 +286,14 @@ terminal.onData((data) => {
ws.send(data);
} else if (data.length === 1 && data.charCodeAt(0) >= 32) {
zerolag.addChar(data);
ws.send(data); // send immediately — overlay shows while echo travels back
ws.send(data); // overlay shows the char while the echo travels back
}
});
```
### Tab Switching (multi-session)
This is the mode that keeps shell features intact: tab completion, `Ctrl+R` history search, and readline bindings all still work, because every byte still reaches the PTY.
### Tab switching (multi-session)
```typescript
function switchToSession(newId: string) {
@@ -280,19 +312,19 @@ function switchToSession(newId: string) {
const saved = savedState.get(newId);
if (saved) zerolag.setFlushed(saved.count, saved.text, false); // silent
// Render after buffer loads
// Render after the buffer loads
terminal.write('', () => zerolag.rerender());
}
```
### Tab Completion
### Tab completion
```typescript
const baseline = zerolag.readPromptText();
zerolag.clear();
sendToPty('\t');
// After response:
// After the response:
zerolag.resetBufferDetection();
const detected = zerolag.detectBufferText();
if (detected && detected !== baseline) {
@@ -302,7 +334,7 @@ if (detected && detected !== baseline) {
}
```
### Resize / Font / Reconnect
### Resize, font, reconnect
```typescript
fitAddon.fit();
@@ -311,14 +343,31 @@ zerolag.rerender();
terminal.options.fontSize = 18;
zerolag.refreshFont();
function onReconnect() { zerolag.rerender(); }
function onReconnect() {
zerolag.rerender();
}
```
### Wide characters (CJK, emoji)
Wide characters work out of the box: the overlay measures each character's cell width, renders double-width spans for wide ones, and positions later characters by visual column instead of character index. Line wrapping is computed in columns too, so a wrapped Japanese or Chinese line lands on the same cells the server will use.
For exact Unicode 11+ widths, load xterm's Unicode addon and the overlay will defer to it:
```typescript
import { Unicode11Addon } from '@xterm/addon-unicode11';
terminal.loadAddon(new Unicode11Addon());
terminal.unicode.activeVersion = '11';
```
Without it, a built-in range table covers Hangul, Kana, CJK Unified (including Ext A through G), fullwidth forms and the emoji planes.
---
## How It Works
## How it works
### DOM Structure
### DOM structure
```
div.xterm-screen (position: relative)
@@ -326,53 +375,61 @@ div.xterm-screen (position: relative)
├── div.xterm-selection (z-index: 1)
├── div.xterm-helpers (z-index: 5)
├── div.xterm-decoration-container (z-index: 6-7)
└── div[zerolag overlay] (z-index: 7) ← our overlay (invisible to Ink)
└── div[zerolag overlay] (z-index: 7) ← our overlay, invisible to Ink
```
### Per-Character Grid Alignment
### Per-character grid alignment
Each character is an absolutely-positioned `<span>`:
```
left = charIndex * cellWidth (CSS pixels)
top = lineIndex * cellHeight (CSS pixels)
width = cellWidth (exact cell width)
left = visualColumn * cellWidth (CSS pixels)
top = lineIndex * cellHeight (CSS pixels)
width = cellWidth * charCellWidth (1 cell, or 2 for wide characters)
```
This avoids sub-pixel drift from normal DOM text flow.
Positioning by visual column instead of letting the browser lay out text is what removes sub-pixel drift.
### Font Matching
### Font matching
1. `fontFamily`, `fontSize`, `fontWeight` from `terminal.options`
2. `letterSpacing` from computed style of `.xterm-rows`
3. `-webkit-font-smoothing: antialiased` (matches canvas grayscale)
2. `letterSpacing` from the computed style of `.xterm-rows`
3. `-webkit-font-smoothing: antialiased` (matches canvas grayscale AA)
4. `font-feature-settings: 'liga' 0, 'calt' 0` (no ligatures)
5. `text-rendering: geometricPrecision`
### Cell Dimensions
### Cell dimensions
- **xterm.js v5.x**: `terminal._core._renderService.dimensions.css.cell` (private API)
- **xterm.js v7+**: `terminal.dimensions.css.cell` (public API, auto-detected)
### Prompt Column Locking
### Prompt column locking
When flushed text exists, the prompt column is locked to prevent jitter from full-screen redraws. Row changes are allowed (output can scroll the prompt).
While flushed text exists the prompt column is locked, so a full-screen redraw cannot make the overlay jitter sideways. Row changes are still allowed, because output legitimately scrolls the prompt.
### Scroll Awareness
### Scroll awareness
Overlay hides when scrolled up (`viewportY !== baseY`). Debounced re-render when scrolling back to bottom.
The overlay hides when the viewport is scrolled up (`viewportY !== baseY`) and re-renders, debounced, when you scroll back to the bottom.
---
## Known Limitations
## Known limitations
- **Canvas/WebGL font mismatch**: Minor sub-pixel differences possible. Per-character absolute positioning minimizes this.
- **Unicode/emoji**: Multi-byte characters occupy variable cell widths — rendered at single-cell width, causing misalignment.
- **Password prompts**: Overlay shows characters that aren't echoed. Call `clear()` when you detect no-echo mode.
- **Prompt in output**: If `$` appears in command output, prompt detection may find the wrong position. Use regex or custom finder.
- **Canvas and WebGL font mismatch**: minor sub-pixel differences are still possible. Per-character absolute positioning keeps them small.
- **Grapheme clusters**: widths are summed per code point, so ZWJ emoji sequences (for example 👨‍👩‍👧) and combining marks can be over-counted. Single-code-point emoji and CJK are correct.
- **Password prompts**: the overlay will happily show characters the server is not echoing. Call `clear()` when you detect a no-echo prompt.
- **Prompt characters in output**: if your prompt marker also appears in command output, detection can latch onto the wrong line. Use a regex or a custom finder.
---
## Origin
[Codeman](https://getcodeman.com) needed this before anyone else did. A coding agent you drive from your phone over a tunnel is unusable if every keystroke costs a round trip.
So the overlay was built there, ran in production for thousands of hours, and survived three deep code audits before being pulled out into this standalone library with its tests intact. Nothing was reimplemented for the extraction: the engine here is the one Codeman ships.
Want the whole thing? [**getcodeman.com**](https://getcodeman.com) · [github.com/Ark0N/Codeman](https://github.com/Ark0N/Codeman)
## License
MIT — [Codeman](https://github.com/Ark0N/Codeman) Contributors
MIT, [Codeman](https://github.com/Ark0N/Codeman) Contributors
+10 -2
View File
@@ -1,7 +1,7 @@
{
"name": "xterm-zerolag-input",
"version": "0.1.4",
"description": "Instant keystroke feedback overlay for xterm.js — eliminates perceived input latency over high-RTT connections",
"version": "0.1.8",
"description": "Instant keystroke feedback overlay for xterm.js: Mosh-inspired local echo that removes perceived input latency over SSH, tunnels and other high-RTT connections",
"type": "module",
"main": "dist/index.cjs",
"module": "dist/index.js",
@@ -26,8 +26,16 @@
"xterm",
"xterm.js",
"terminal",
"web-terminal",
"local-echo",
"local echo",
"mosh",
"input-latency",
"latency",
"zero-lag",
"keystroke",
"ssh",
"remote-terminal",
"overlay",
"addon"
],
+2
View File
@@ -67,6 +67,7 @@ appendFileSync(
// 4. Minify frontend assets
run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --outfile=dist/web/public/input-cjk.js --allow-overwrite');
run('minify i18n.js', 'npx esbuild dist/web/public/i18n.js --minify --outfile=dist/web/public/i18n.js --allow-overwrite');
run('minify sanitize-html.js', 'npx esbuild dist/web/public/sanitize-html.js --minify --outfile=dist/web/public/sanitize-html.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify terminal-ui.js', 'npx esbuild dist/web/public/terminal-ui.js --minify --outfile=dist/web/public/terminal-ui.js --allow-overwrite');
@@ -86,6 +87,7 @@ console.log('\n[build] content-hash cache busting');
'styles.css',
'mobile.css',
'constants.js',
'i18n.js',
'mobile-handlers.js',
'voice-input.js',
'notification-manager.js',
+482
View File
@@ -0,0 +1,482 @@
#!/usr/bin/env node
/**
* capture-readme-gifs.mjs
*
* Deterministic README GIFs — no real server, Claude CLI, or tmux. Reuses the
* mock-injection pipeline from capture-readme-screenshots.mjs (static file
* server + page.route mocks), drives a scripted timeline in the page, records
* it with Playwright video, and converts to GIF via ffmpeg palette encoding.
*
* Scenes:
* 1. subagent-demo.gif — terminal spawns 3 parallel agents; floating agent
* windows open one by one and stream tool-call activity live (driven
* through the real _onSubagentDiscovered/_onSubagentToolCall handlers).
* 2. zerolag-demo.gif — side-by-side typing: instant local echo (zerolag)
* vs bursty ~350 ms server echo, rendered with the vendored xterm.
*
* Usage: node scripts/capture-readme-gifs.mjs
* SCREENSHOT_OUT_DIR=/path/to/review node scripts/capture-readme-gifs.mjs
* Output: docs/images/ (or flat into SCREENSHOT_OUT_DIR)
* Requires: ffmpeg
*/
import { chromium } from 'playwright';
import { execSync } from 'child_process';
import { mkdtempSync, rmSync } from 'fs';
import { tmpdir } from 'os';
import { join } from 'path';
import {
PORT,
SESSION_IDS,
STANDARD_SESSIONS,
buildInitPayload,
startStaticServer,
setupRoutes,
injectState,
outPath,
RST, GRN, YEL, MAG, CYN, GRY, BOLD,
} from './capture-readme-screenshots.mjs';
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const GIF_COLORS = 192;
// ─── ffmpeg conversion (palette recipe from capture-subagent-gif.mjs) ────────
function webmToGif(videoPath, gifPath, { ss, duration, width, fps }) {
// One GLOBAL palette (default stats_mode=full) + ordered dither: per-frame
// palettes (stats_mode=single:new=1) make dirty rectangles visibly mismatch
// on flat dark UI, and error-diffusion dither shimmers between frames.
const filters = `fps=${fps},scale=${width}:-1:flags=lanczos`;
execSync(
`ffmpeg -y -loglevel error -ss ${ss.toFixed(2)} -t ${duration} -i "${videoPath}" ` +
`-vf "${filters},split[s0][s1];[s0]palettegen=max_colors=${GIF_COLORS}:reserve_transparent=0[p];` +
`[s1][p]paletteuse=dither=bayer:bayer_scale=5:diff_mode=rectangle" "${gifPath}"`,
{ stdio: 'inherit' }
);
}
// ─── Scene 1: subagent demo ──────────────────────────────────────────────────
const SUBAGENT_VIEWPORT = { width: 1440, height: 810 };
// Terminal content visible before the agents spawn
const TERMINAL_PRESPAWN = [
'',
`${GRN}●${RST} Working on ${CYN}/home/arkon/codeman-cases/testcase${RST} - I'll use the ${BOLD}Task tool${RST} to spawn parallel agents.`,
'',
`${GRN}●${RST} ${BOLD}Read${RST}(/home/arkon/codeman-cases/testcase/CLAUDE.md)`,
` ${GRY}░${RST} Read ${BOLD}127${RST} lines ${GRY}│${RST} ${CYN}1.2KB${RST}`,
'',
`${GRN}●${RST} ${BOLD}Bash${RST}(find . -name "*.ts" -not -path "*/node_modules/*" | head -20)`,
` ${GRY}░${RST} ./src/index.ts`,
` ${GRY}░${RST} ./src/session.ts`,
` ${GRY}░${RST} ./src/web/server.ts`,
` ${GRY}░${RST} ${GRY}... (17 more)${RST}`,
'',
`${GRN}●${RST} I'll spawn 3 parallel research agents to analyze different parts of the codebase simultaneously.`,
'',
].join('\r\n');
function makeAgent(agentId, description, startedOffsetMs) {
return {
agentId,
sessionId: 'claude-sess-w1-0001',
projectHash: 'abc123',
filePath: `/tmp/${agentId}.jsonl`,
startedAt: new Date(Date.now() - startedOffsetMs).toISOString(),
lastActivityAt: Date.now(),
status: 'active',
toolCallCount: 0,
entryCount: 0,
fileSize: 4000,
description,
model: 'claude-haiku-4-5-20251001',
modelShort: 'haiku',
totalInputTokens: 0,
totalOutputTokens: 0,
parentSessionId: SESSION_IDS.w1,
};
}
// Timeline events: t (ms from scene start) + kind
// term — write raw data to the session terminal
// discover — register subagent + open + position its floating window
// tool — stream a tool call into an agent window
// msg — stream an assistant message into an agent window
// complete — flip an agent to completed
function buildSubagentTimeline() {
const T = (lines) => lines.join('\r\n') + '\r\n';
const tool = (t, agentId, name, input) => ({ t, kind: 'tool', agentId, tool: name, input });
const msg = (t, agentId, text) => ({ t, kind: 'msg', agentId, text });
return [
{
t: 600,
kind: 'term',
data: T([
`${GRN}●${RST} ${BOLD}Task${RST}(Find and document all API endpoints in src/)`,
` ${GRY}░${RST} Spawned ${CYN}agent-001${RST} ${GRY}(haiku)${RST}`,
'',
]),
},
{
t: 1000,
kind: 'discover',
agent: makeAgent('agent-001', 'Find and document all API endpoints in src/', 2000),
x: 440, y: 45,
},
tool(1500, 'agent-001', 'Glob', { pattern: 'src/**/*.ts' }),
{
t: 2000,
kind: 'term',
data: T([
`${GRN}●${RST} ${BOLD}Task${RST}(Explore and understand test structure in test/)`,
` ${GRY}░${RST} Spawned ${CYN}agent-002${RST} ${GRY}(haiku)${RST}`,
'',
]),
},
tool(2200, 'agent-001', 'Read', { file_path: '/home/arkon/codeman/src/web/server.ts' }),
{
t: 2500,
kind: 'discover',
agent: makeAgent('agent-002', 'Explore and understand test structure in test/', 1200),
x: 880, y: 45,
},
tool(3000, 'agent-002', 'Glob', { pattern: 'test/**/*.test.ts' }),
{
t: 3300,
kind: 'term',
data: T([
`${GRN}●${RST} ${BOLD}Task${RST}(Analyze TypeScript type definitions in src/types.ts)`,
` ${GRY}░${RST} Spawned ${CYN}agent-003${RST} ${GRY}(haiku)${RST}`,
'',
]),
},
tool(3500, 'agent-001', 'Grep', { pattern: 'app\\.get|app\\.post|app\\.delete', path: 'src/' }),
{
t: 3800,
kind: 'discover',
agent: makeAgent('agent-003', 'Analyze TypeScript type definitions in src/types.ts', 400),
x: 660, y: 400,
},
tool(4100, 'agent-002', 'Read', { file_path: '/home/arkon/codeman/test/respawn-test-utils.ts' }),
{
t: 4500,
kind: 'term',
data: T([
`${MAG}✻${RST} ${YEL}Waiting for agents...${RST} ${GRY}(${BOLD}esc${RST}${GRY} to interrupt · 32s · ↓ 1.7k tokens · thinking)${RST}`,
'',
]),
},
tool(4700, 'agent-003', 'Read', { file_path: '/home/arkon/codeman/src/types.ts' }),
tool(5200, 'agent-001', 'Read', { file_path: '/home/arkon/codeman/src/web/schemas.ts' }),
tool(5600, 'agent-002', 'Read', { file_path: '/home/arkon/codeman/config/vitest.config.ts' }),
tool(6100, 'agent-003', 'Grep', { pattern: 'export (interface|type)', path: 'src/types/' }),
msg(6700, 'agent-001', 'Found 47 API endpoints across server.ts. Documenting REST paths...'),
tool(7100, 'agent-002', 'Grep', { pattern: 'const PORT =', path: 'test/' }),
msg(7700, 'agent-002', 'Analyzing test patterns: MockSession, unique ports, fileParallelism: false...'),
tool(8100, 'agent-003', 'Read', { file_path: '/home/arkon/codeman/src/types/index.ts' }),
msg(8700, 'agent-003', 'Mapped 38 exported interfaces across 15 domain files. Building summary...'),
{
t: 9300,
kind: 'term',
data: T([
`${GRN}●${RST} ${CYN}agent-001${RST}: ${GRY}12 tool calls — Glob, Read(server.ts), Grep(endpoints)...${RST}`,
`${GRN}●${RST} ${CYN}agent-002${RST}: ${GRY}8 tool calls — Glob, Read(test-utils), Read(vitest.config)...${RST}`,
`${GRN}●${RST} ${CYN}agent-003${RST}: ${GRY}7 tool calls — Read(types.ts), Grep(interface)...${RST}`,
'',
]),
},
tool(10100, 'agent-001', 'Glob', { pattern: 'src/web/routes/*.ts' }),
tool(10600, 'agent-002', 'Read', { file_path: '/home/arkon/codeman/test/setup.ts' }),
tool(11100, 'agent-003', 'Grep', { pattern: 'assertNever', path: 'src/' }),
{
t: 11600,
kind: 'term',
data: T([`${GRN}●${RST} ${GRY}171.8k, 13s${RST} ${GRY}│${RST} ${GRY}1.7k tokens${RST} ${GRY}│${RST} ${GRY}thinking${RST}`, '']),
},
];
}
const SUBAGENT_TAIL_HOLD = 2500; // hold the final frame
async function recordSubagentScene(browser, videoDir) {
console.log('\n1/2 Recording subagent-demo...');
const context = await browser.newContext({
viewport: SUBAGENT_VIEWPORT,
deviceScaleFactor: 1,
recordVideo: { dir: videoDir, size: SUBAGENT_VIEWPORT },
});
const recStart = Date.now();
const page = await context.newPage();
page.setDefaultTimeout(30000);
// Start with NO subagents — they appear during the recording
const initPayload = buildInitPayload(STANDARD_SESSIONS);
await setupRoutes(page, initPayload, TERMINAL_PRESPAWN);
await page.goto(`http://localhost:${PORT}`, { waitUntil: 'domcontentloaded' });
await injectState(page, initPayload, TERMINAL_PRESPAWN, SESSION_IDS.w1);
await page.evaluate(() => {
try { window.app?.fitAddon?.fit(); } catch {}
window.app?.terminal?.scrollToBottom();
});
await sleep(500);
const timeline = buildSubagentTimeline();
const totalMs = Math.max(...timeline.map((e) => e.t)) + SUBAGENT_TAIL_HOLD;
const sceneStart = Date.now();
// Run the whole timeline inside the page so events interleave naturally
await page.evaluate((events) => {
const app = window.app;
for (const ev of events) {
setTimeout(() => {
try {
if (ev.kind === 'term') {
app.terminal.write(ev.data);
app.terminal.scrollToBottom();
} else if (ev.kind === 'discover') {
app._onSubagentDiscovered(ev.agent);
app.openSubagentWindow(ev.agent.agentId);
// The spawn animation (400ms) lands on the auto-grid; glide to our tile after it
setTimeout(() => {
const win = app.subagentWindows.get(ev.agent.agentId);
if (win?.element) {
win.element.style.transition = 'left 0.25s ease, top 0.25s ease';
win.element.style.left = `${ev.x}px`;
win.element.style.top = `${ev.y}px`;
}
}, 520);
setTimeout(() => {
const win = app.subagentWindows.get(ev.agent.agentId);
if (win?.element) win.element.style.transition = '';
app.updateConnectionLines();
}, 850);
} else if (ev.kind === 'tool') {
app._onSubagentToolCall({
agentId: ev.agentId,
tool: ev.tool,
input: ev.input,
timestamp: new Date().toISOString(),
});
} else if (ev.kind === 'msg') {
app._onSubagentMessage({
agentId: ev.agentId,
role: 'assistant',
text: ev.text,
timestamp: new Date().toISOString(),
});
} else if (ev.kind === 'complete') {
app._onSubagentCompleted({ agentId: ev.agentId, timestamp: new Date().toISOString() });
}
} catch (err) {
console.error('timeline event failed', ev, err);
}
}, ev.t);
}
}, timeline);
await sleep(totalMs + 500);
await page.close();
const videoPath = await page.video().path();
await context.close();
return {
videoPath,
ss: (sceneStart - recStart) / 1000 - 0.4,
duration: (totalMs + 400) / 1000,
};
}
// ─── Scene 2: zerolag typing comparison ──────────────────────────────────────
const ZEROLAG_VIEWPORT = { width: 1280, height: 470 };
const TYPED_TEXT = 'echo "zero lag typing from anywhere"';
const TYPE_INTERVAL_MS = 110;
const REMOTE_FLUSH_MS = 350; // server-echo pane flushes queued chars in bursts
const ZEROLAG_TAIL_HOLD = 1800;
const ZEROLAG_HTML = `<!DOCTYPE html>
<html>
<head>
<link rel="stylesheet" href="http://localhost:${PORT}/vendor/xterm.css">
<script src="http://localhost:${PORT}/vendor/xterm.min.js"></script>
<style>
* { margin: 0; box-sizing: border-box; }
body {
width: 1280px; height: 470px; background: #0a0a0c;
display: flex; align-items: center; justify-content: center; gap: 48px;
font-family: -apple-system, 'Segoe UI', Roboto, sans-serif;
}
.pane { width: 560px; }
.card {
background: #131316; border: 1px solid rgba(255,255,255,0.08);
border-radius: 10px; overflow: hidden;
box-shadow: 0 8px 32px rgba(0,0,0,0.45);
}
.card-head {
display: flex; align-items: baseline; gap: 10px;
padding: 12px 16px; border-bottom: 1px solid rgba(255,255,255,0.06);
}
.dot { width: 9px; height: 9px; border-radius: 50%; align-self: center; }
.title { font-size: 15px; font-weight: 600; color: #e8e8ea; }
.sub { font-size: 12.5px; color: #8b8b92; }
.term { padding: 16px 8px 12px 16px; height: 165px; }
.good .dot { background: #22c55e; box-shadow: 0 0 8px rgba(34,197,94,0.7); }
.bad .dot { background: #ef4444; box-shadow: 0 0 8px rgba(239,68,68,0.7); }
.tag {
margin-top: 14px; text-align: center; font-size: 14.5px; color: #7e7e86;
}
.tag b { color: #22c55e; font-weight: 600; }
.bad-tag b { color: #ef4444; }
</style>
</head>
<body>
<div class="pane">
<div class="card good">
<div class="card-head">
<span class="dot"></span>
<span class="title">With zerolag-input</span>
<span class="sub">instant local echo</span>
</div>
<div class="term" id="termLeft"></div>
</div>
<div class="tag">keystrokes echo in <b>0 ms</b></div>
</div>
<div class="pane">
<div class="card bad">
<div class="card-head">
<span class="dot"></span>
<span class="title">Without</span>
<span class="sub">server round-trip echo</span>
</div>
<div class="term" id="termRight"></div>
</div>
<div class="tag bad-tag">keystrokes echo after <b>~350 ms</b></div>
</div>
</body>
</html>`;
async function recordZerolagScene(browser, videoDir) {
console.log('\n2/2 Recording zerolag-demo...');
const context = await browser.newContext({
viewport: ZEROLAG_VIEWPORT,
deviceScaleFactor: 1,
recordVideo: { dir: videoDir, size: ZEROLAG_VIEWPORT },
});
const recStart = Date.now();
const page = await context.newPage();
page.setDefaultTimeout(30000);
await page.setContent(ZEROLAG_HTML, { waitUntil: 'load' });
await page.waitForFunction(() => typeof Terminal !== 'undefined');
await page.evaluate(() => {
const theme = {
background: '#131316',
foreground: '#e8e8ea',
cursor: '#22c55e',
cursorAccent: '#131316',
};
const mk = (id) => {
const term = new Terminal({
cols: 44,
rows: 5,
fontSize: 20,
fontFamily: "'SF Mono', 'Cascadia Code', Menlo, monospace",
cursorBlink: true,
cursorStyle: 'block',
theme,
});
term.open(document.getElementById(id));
term.write('\x1b[32m❯\x1b[0m ');
return term;
};
window.termLeft = mk('termLeft');
window.termRight = mk('termRight');
});
await sleep(600);
const sceneStart = Date.now();
const typingMs = TYPED_TEXT.length * TYPE_INTERVAL_MS;
const totalMs = typingMs + REMOTE_FLUSH_MS + ZEROLAG_TAIL_HOLD;
await page.evaluate(
({ text, interval, flushEvery }) => {
let i = 0;
const remoteQueue = [];
const typer = setInterval(() => {
if (i >= text.length) { clearInterval(typer); return; }
const ch = text[i++];
window.termLeft.write(ch); // local echo: instant
remoteQueue.push(ch); // server echo: waits for the round-trip
}, interval);
const flusher = setInterval(() => {
if (remoteQueue.length) window.termRight.write(remoteQueue.splice(0).join(''));
if (i >= text.length && remoteQueue.length === 0) clearInterval(flusher);
}, flushEvery);
},
{ text: TYPED_TEXT, interval: TYPE_INTERVAL_MS, flushEvery: REMOTE_FLUSH_MS }
);
await sleep(totalMs + 400);
await page.close();
const videoPath = await page.video().path();
await context.close();
return {
videoPath,
ss: (sceneStart - recStart) / 1000 - 0.6, // small lead-in with idle cursors
duration: (totalMs + 600) / 1000,
};
}
// ─── Main ────────────────────────────────────────────────────────────────────
async function main() {
console.log('='.repeat(60));
console.log('Codeman README GIF Capture');
console.log('='.repeat(60));
const server = await startStaticServer();
const videoDir = mkdtempSync(join(tmpdir(), 'codeman-gifs-'));
let browser;
try {
browser = await chromium.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', '--disable-gpu'],
});
const sub = await recordSubagentScene(browser, videoDir);
const subGif = outPath('images', 'subagent-demo.gif');
webmToGif(sub.videoPath, subGif, { ss: Math.max(0, sub.ss), duration: sub.duration, width: 960, fps: 8 });
console.log(` Saved: ${subGif}`);
const zl = await recordZerolagScene(browser, videoDir);
const zlGif = outPath('images', 'zerolag-demo.gif');
webmToGif(zl.videoPath, zlGif, { ss: Math.max(0, zl.ss), duration: zl.duration, width: 900, fps: 10 });
console.log(` Saved: ${zlGif}`);
console.log('\nDone.');
} catch (err) {
console.error('\nFatal error:', err.message);
console.error(err.stack);
process.exitCode = 1;
} finally {
if (browser) await browser.close().catch(() => {});
server.close();
rmSync(videoDir, { recursive: true, force: true });
}
}
process.on('SIGINT', () => process.exit(1));
main();
+65 -13
View File
@@ -50,7 +50,10 @@ async function newCtx(browser) {
try {
localStorage.setItem('codeman:skin', skin);
localStorage.setItem('codeman-font-size', String(font));
const blob = { skin, showFileBrowser: false, showProjectInsights: false };
const blob = { skin, showFileBrowser: false, showProjectInsights: false, showTokenCount: false };
// Don't auto-hide subagent windows that belong to a non-active tab — the
// subagent scene re-homes agents and needs both windows visible at once.
blob.subagentActiveTabOnly = false;
if (planUsage) blob.showPlanUsageLimits = true;
localStorage.setItem('codeman-app-settings', JSON.stringify(blob));
} catch {
@@ -136,9 +139,9 @@ async function sceneSubagent(browser) {
const sessions = await listSessions(page);
const targetId = process.env.SUBAGENT_SID || (sessions.find((s) => s.mode === 'claude') || sessions[0])?.id;
if (targetId) await page.evaluate((id) => window.app.selectSession(id), targetId);
// Wait (up to ~25s) for live subagents to arrive via SSE into app.subagents.
// Wait (up to ~45s) for live subagents to arrive via SSE into app.subagents.
let agents = [];
for (let i = 0; i < 25; i++) {
for (let i = 0; i < 45; i++) {
agents = await page.evaluate(() =>
Array.from(window.app.subagents?.entries?.() || []).map(([id, a]) => ({ id, name: a.name ?? a.agentType ?? '' }))
);
@@ -151,6 +154,44 @@ async function sceneSubagent(browser) {
await context.close();
return;
}
// The window body renders from app.subagentActivity, which fills ONLY from live
// SSE tool-call/progress events — a fresh client never gets past activity replayed.
// So sit connected and wait for live activity to accumulate, then open the two
// agents that actually have content (otherwise the windows read "No activity yet").
let active = [];
for (let i = 0; i < 100; i++) {
active = await page.evaluate(() =>
Array.from(window.app.subagentActivity?.entries?.() || [])
.filter(([, arr]) => Array.isArray(arr) && arr.length >= 1)
.map(([id, arr]) => ({ id, n: arr.length }))
.sort((a, b) => b.n - a.n)
);
if (active.length >= 2) break;
// xhigh-effort agents churn in bursts between long thinking pauses, so be
// patient (~150s); accept a single populated window after ~45s if that's all.
if (i >= 30 && active.length >= 1) break;
await sleep(1500);
}
console.log(' agents with live activity:', JSON.stringify(active));
const openIds = (active.length ? active : agents).map((a) => a.id);
// Capture-only DOM nudge: on fresh dev sessions, a tab's claudeSessionId stays the
// Codeman id and never becomes the real Claude conversation UUID, so the window
// open-gate (claudeSessionId === agent.sessionId) + the activeTabOnly hide rule both
// fail. Re-home the chosen agents onto the active tab and align its claudeSessionId
// to the agents' (shared) sessionId so the windows open AND show their live activity.
await page.evaluate(
(ids) => {
const activeId = window.app.activeSessionId;
const tab = window.app.sessions.get(activeId);
ids.slice(0, 2).forEach((id) => {
const a = window.app.subagents.get(id);
if (!a) return;
a.parentSessionId = activeId;
if (tab && a.sessionId) tab.claudeSessionId = a.sessionId;
});
},
openIds
);
await page.evaluate(
(ids) => {
ids.slice(0, 2).forEach((id) => {
@@ -159,22 +200,33 @@ async function sceneSubagent(browser) {
} catch {}
});
},
agents.map((a) => a.id)
openIds
);
await sleep(2000);
await page.evaluate(() => {
// Viewport-relative tiling: center two subagent windows over the terminal so
// the layout adapts to whatever VW/VH the capture uses (e.g. the HQ 1100×650
// recipe) instead of overflowing at narrower widths.
const wins = Array.from(window.app.subagentWindows.values());
const place = [
{ left: 360, top: 60, w: 430, h: 330 },
{ left: 810, top: 60, w: 430, h: 330 },
];
const W = window.innerWidth;
const H = window.innerHeight;
const winW = Math.min(440, Math.floor((W - 60) / 2 - 10));
const winH = Math.min(360, Math.floor(H * 0.56));
const top = Math.floor(H * 0.16);
const gap = 16;
const totalW = winW * 2 + gap;
const startLeft = Math.max(16, Math.floor((W - totalW) / 2));
wins.slice(0, 2).forEach((win, i) => {
const el = win.element;
const p = place[i];
el.style.left = p.left + 'px';
el.style.top = p.top + 'px';
el.style.width = p.w + 'px';
el.style.height = p.h + 'px';
// Force visible: a freshly opened window may be hidden by the activeTabOnly
// rule before we override it (we also seed subagentActiveTabOnly:false).
win.hidden = false;
win.minimized = false;
el.style.display = 'flex';
el.style.left = startLeft + i * (winW + gap) + 'px';
el.style.top = top + 'px';
el.style.width = winW + 'px';
el.style.height = winH + 'px';
});
});
await sleep(1500);
File diff suppressed because it is too large Load Diff
+1
View File
@@ -79,6 +79,7 @@ const main = async () => {
showMonitor: false,
showSubagents: false,
showProjectInsights: false,
showTokenCount: false,
};
if (planUsage) blob.showPlanUsageLimits = true;
localStorage.setItem('codeman-app-settings', JSON.stringify(blob));
+2
View File
@@ -126,6 +126,7 @@ async function capture() {
showMonitor: false,
showProjectInsights: false,
showFileBrowser: false,
showTokenCount: false,
});
localStorage.setItem('codeman-app-settings', JSON.stringify(existing));
});
@@ -230,6 +231,7 @@ async function capture() {
showMonitor: false,
showProjectInsights: false,
showFileBrowser: false,
showTokenCount: false,
});
localStorage.setItem('codeman-app-settings', JSON.stringify(existing));
});
+1
View File
@@ -221,6 +221,7 @@ async function configureSettings(page) {
subagentTrackingEnabled: true,
subagentActiveTabOnly: false, // Show all subagents regardless of active tab
showMonitor: true,
showTokenCount: false,
};
localStorage.setItem('codeman-app-settings', JSON.stringify(settings));
});
+261
View File
@@ -0,0 +1,261 @@
#!/usr/bin/env node
/**
* @fileoverview Repairs node-pty's macOS `spawn-helper` and verifies that a PTY
* can really be spawned. Called by `scripts/postinstall.js` on every install and
* exposed as `npm run fix:node-pty` for repairing an install after the fact.
*
* Why this exists (issues #6 and #204):
*
* node-pty@1.1.0 publishes its macOS prebuilt helper as
* `prebuilds/darwin-<arch>/spawn-helper` with mode 0644, i.e. no execute bit.
* On macOS node-pty launches every PTY through that helper with posix_spawnp,
* which then fails EACCES and surfaces as `Error: posix_spawnp failed.` on every
* session start.
*
* It is macOS-exclusive twice over: `spawn-helper` is an `OS=="mac"` gyp target,
* and pty.cc only spawns it under `#if defined(__APPLE__)`. node-pty ships
* prebuilds for darwin and win32 only, so Linux always compiles from source
* (which produces an executable helper) and never sees the bug.
*
* The repair is a chmod, NOT a rebuild: the prebuilt binary itself is fine, and
* requiring a from-source rebuild would make every macOS install depend on Xcode
* command line tools. A rebuild is attempted only when a chmod plus a real spawn
* probe still can't get a working PTY, and the prebuilds tree is backed up first
* so a failed rebuild can never leave the install worse than it started.
*/
import { chmodSync, cpSync, existsSync, readdirSync, rmSync, statSync } from 'node:fs';
import { execSync } from 'node:child_process';
import { createRequire } from 'node:module';
import { tmpdir } from 'node:os';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
const require = createRequire(import.meta.url);
/** Errors that mean "the native module or its helper is unusable", i.e. worth a rebuild. */
const NATIVE_FAILURE_PATTERN = /posix_spawnp|spawn-helper|Failed to load native module|Cannot find module/i;
/**
* Locates the installed node-pty package directory.
*
* @returns {string|null} Absolute path to the package root, or null if not installed.
*/
export function findNodePtyDir() {
// package.json first: node-pty declares no "exports" map, so the subpath resolves,
// and it lands on the package root directly. require.resolve('node-pty') would give
// <pkg>/lib/index.js, which is one directory deeper than callers expect.
try {
return dirname(require.resolve('node-pty/package.json'));
} catch {
/* fall through */
}
try {
return join(dirname(require.resolve('node-pty')), '..');
} catch {
return null;
}
}
/**
* Lists every `spawn-helper` shipped in a node-pty install.
*
* node-pty's own loader (lib/utils.js) checks `build/Release`, `build/Debug` and
* then `prebuilds/<platform>-<arch>`, and takes the helper from whichever
* directory the native module loaded out of, so all of them must be executable,
* not just the one this machine happens to use today.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @returns {string[]} Absolute paths of the helpers that exist on disk.
*/
export function listSpawnHelpers(ptyDir) {
const dirs = [join(ptyDir, 'build', 'Release'), join(ptyDir, 'build', 'Debug')];
const prebuilds = join(ptyDir, 'prebuilds');
if (existsSync(prebuilds)) {
try {
for (const entry of readdirSync(prebuilds, { withFileTypes: true })) {
if (entry.isDirectory()) dirs.push(join(prebuilds, entry.name));
}
} catch {
/* unreadable prebuilds dir: nothing to repair there */
}
}
return dirs.map((d) => join(d, 'spawn-helper')).filter((p) => existsSync(p));
}
/**
* Adds the execute bit to every `spawn-helper` that is missing it.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @returns {{ repaired: string[], failed: Array<{ path: string, error: string }> }}
*/
export function repairSpawnHelpers(ptyDir) {
const repaired = [];
const failed = [];
for (const helper of listSpawnHelpers(ptyDir)) {
try {
const mode = statSync(helper).mode & 0o777;
if ((mode & 0o111) === 0o111) continue; // already executable by all
chmodSync(helper, mode | 0o755);
repaired.push(helper);
} catch (err) {
failed.push({ path: helper, error: err instanceof Error ? err.message : String(err) });
}
}
return { repaired, failed };
}
/**
* Proves node-pty works by actually opening a PTY, which is the only check that
* exercises the spawn-helper path that breaks. A `require` alone would pass on a
* broken install, because the helper is only touched at spawn time.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @returns {{ ok: boolean, error?: string, nativeFailure?: boolean }}
*/
export function verifyPtySpawn(ptyDir) {
let child;
try {
const pty = require(ptyDir); // directory require → node-pty's "main" (lib/index.js)
const file = process.platform === 'win32' ? process.env.COMSPEC || 'cmd.exe' : '/bin/echo';
const args = process.platform === 'win32' ? ['/c', 'exit'] : ['codeman-node-pty-check'];
child = pty.spawn(file, args, {
name: 'xterm-color',
cols: 80,
rows: 24,
cwd: tmpdir(),
env: process.env,
});
return { ok: true };
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
return { ok: false, error: message, nativeFailure: NATIVE_FAILURE_PATTERN.test(message) };
} finally {
try {
child?.kill();
} catch {
/* the probe child exits on its own anyway */
}
}
}
/**
* Rebuilds node-pty from source, preserving the prebuilds tree across a failure.
*
* node-pty's install script deletes `prebuilds/` as soon as
* `npm_config_build_from_source` is set and only then shells out to node-gyp, so
* a machine without a compiler toolchain would otherwise be left with neither a
* prebuilt nor a compiled binary.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @param {string} cwd Directory to run npm from (the package root that owns node_modules).
* @returns {{ ok: boolean, error?: string }}
*/
function rebuildFromSource(ptyDir, cwd) {
const prebuilds = join(ptyDir, 'prebuilds');
const backup = join(ptyDir, '.prebuilds-codeman-backup');
let backedUp = false;
if (existsSync(prebuilds)) {
try {
rmSync(backup, { recursive: true, force: true });
cpSync(prebuilds, backup, { recursive: true });
backedUp = true;
} catch {
/* best effort: proceed without a safety net rather than skip the repair */
}
}
try {
execSync('npm rebuild node-pty --build-from-source', { cwd, stdio: 'pipe', timeout: 300000 });
return { ok: true };
} catch (err) {
if (backedUp && !existsSync(prebuilds)) {
try {
cpSync(backup, prebuilds, { recursive: true });
} catch {
/* nothing further we can do */
}
}
return { ok: false, error: err instanceof Error ? err.message : String(err) };
} finally {
rmSync(backup, { recursive: true, force: true });
}
}
/**
* Full repair flow: chmod, verify, and only rebuild if a working PTY still can't
* be opened.
*
* @param {object} [options]
* @param {(line: string) => void} [options.log] Progress sink (default: silent).
* @param {(line: string) => void} [options.warn] Warning sink (default: same as log).
* @param {boolean} [options.allowRebuild] Permit a from-source rebuild (default: true).
* @returns {Promise<{ ok: boolean, repaired: string[], rebuilt: boolean, reason?: string }>}
*/
export async function fixNodePty(options = {}) {
const log = options.log ?? (() => {});
const warn = options.warn ?? log;
const allowRebuild = options.allowRebuild ?? true;
const ptyDir = findNodePtyDir();
if (!ptyDir) {
return { ok: false, repaired: [], rebuilt: false, reason: 'node-pty is not installed' };
}
const { repaired, failed } = repairSpawnHelpers(ptyDir);
for (const f of failed) warn(`could not chmod ${f.path}: ${f.error}`);
if (repaired.length > 0) {
log(`made node-pty spawn-helper executable (${repaired.length} file${repaired.length === 1 ? '' : 's'})`);
}
const first = verifyPtySpawn(ptyDir);
if (first.ok) return { ok: true, repaired, rebuilt: false };
if (!allowRebuild || !first.nativeFailure) {
return { ok: false, repaired, rebuilt: false, reason: first.error };
}
warn(`node-pty could not open a PTY (${first.error}), rebuilding from source...`);
const projectRoot = join(dirname(fileURLToPath(import.meta.url)), '..');
const rebuild = rebuildFromSource(ptyDir, projectRoot);
if (!rebuild.ok) {
return { ok: false, repaired, rebuilt: false, reason: `rebuild failed: ${rebuild.error}` };
}
const after = repairSpawnHelpers(ptyDir);
repaired.push(...after.repaired);
const second = verifyPtySpawn(ptyDir);
return second.ok
? { ok: true, repaired, rebuilt: true }
: { ok: false, repaired, rebuilt: true, reason: second.error };
}
// ---------------------------------------------------------------------------
// CLI: node scripts/fix-node-pty.mjs [--quiet]
// ---------------------------------------------------------------------------
const isDirectRun = process.argv[1] && fileURLToPath(import.meta.url) === process.argv[1];
if (isDirectRun) {
const quiet = process.argv.includes('--quiet');
const say = (line) => {
if (!quiet) console.log(line);
};
const result = await fixNodePty({ log: say, warn: (line) => console.warn(line) });
if (result.ok) {
say(result.repaired.length > 0 || result.rebuilt ? 'node-pty repaired, PTY spawning works' : 'node-pty is healthy');
process.exit(0);
}
console.error(`node-pty is not usable: ${result.reason}`);
console.error('Try: cd node_modules/node-pty && npx node-gyp rebuild');
process.exit(1);
}
+21 -24
View File
@@ -6,7 +6,7 @@
*/
import { execSync, spawn } from 'child_process';
import { chmodSync, existsSync } from 'fs';
import { existsSync } from 'fs';
import { homedir, platform } from 'os';
import { join } from 'path';
import { createRequire } from 'module';
@@ -148,35 +148,32 @@ if (majorVersion < MIN_NODE_VERSION) {
}
// ----------------------------------------------------------------------------
// 1b. Fix node-pty spawn-helper permissions (macOS posix_spawnp fix)
// 1b. Repair + verify node-pty (macOS posix_spawnp fix, issues #6 and #204)
//
// node-pty ships its macOS spawn-helper without the execute bit, which breaks
// every session start on macOS. fixNodePty() chmods it, then proves a PTY can
// actually be opened, and only falls back to a from-source rebuild if that
// still fails. See scripts/fix-node-pty.mjs for the full story.
// ----------------------------------------------------------------------------
try {
const require = createRequire(import.meta.url);
const ptyPath = join(require.resolve('node-pty'), '..');
const spawnHelper = join(ptyPath, 'build', 'Release', 'spawn-helper');
if (existsSync(spawnHelper)) {
chmodSync(spawnHelper, 0o755);
console.log(colors.green('✓ node-pty spawn-helper permissions fixed'));
}
} catch {
// Non-critical — only affects macOS with prebuilt binaries
}
const { fixNodePty } = await import('./fix-node-pty.mjs');
const result = await fixNodePty({
log: (line) => console.log(colors.dim(` ${line}`)),
warn: (line) => console.log(colors.yellow(`⚠ ${line}`)),
});
// ----------------------------------------------------------------------------
// 1c. Rebuild node-pty from source for Node.js 22+ compatibility
// ----------------------------------------------------------------------------
if (majorVersion >= 22) {
try {
console.log(colors.dim(' Rebuilding node-pty from source for Node.js 22+...'));
execSync('npm rebuild node-pty --build-from-source', { stdio: 'pipe', timeout: 120000 });
console.log(colors.green('✓ node-pty rebuilt from source'));
} catch {
if (result.ok) {
console.log(colors.green('✓ node-pty verified') + colors.dim(' (PTY spawn works)'));
} else {
hasWarnings = true;
console.log(colors.yellow('⚠ Failed to rebuild node-pty from source'));
console.log(colors.dim(' You may need to run: npm rebuild node-pty --build-from-source'));
console.log(colors.yellow(`⚠ node-pty is not usable: ${result.reason}`));
console.log(colors.dim(' Sessions will fail to start. Try: ') + colors.cyan('npm run fix:node-pty'));
}
} catch (err) {
hasWarnings = true;
console.log(colors.yellow(`⚠ Could not verify node-pty: ${err.message}`));
console.log(colors.dim(' If sessions fail to start, run: ') + colors.cyan('npm run fix:node-pty'));
}
// ----------------------------------------------------------------------------
+344
View File
@@ -0,0 +1,344 @@
---
name: codeman
description: >-
Drive Codeman, the session manager this agent is running inside, over its HTTP API:
list sessions, start worker sessions, send them prompts, block until they finish
(wait / wait-output / send-and-wait), read their output, and clean up. Use when asked
to orchestrate or parallelize work across Codeman sessions, watch another session, or
start and manage workers. Only usable inside a Codeman-managed session
(CODEMAN_MUX=1); refuse to act otherwise.
---
# Driving Codeman from inside a session
You are an agent running inside a Codeman-managed terminal session. Codeman is the
server that spawned you; its HTTP API can start, prompt, watch, and delete other
sessions. Every recipe below was verified live. Full endpoint tables and
troubleshooting: [reference/endpoints.md](reference/endpoints.md). Worked multi-worker
flows: [reference/recipes.md](reference/recipes.md).
## 0. Guard, and the one thing that breaks every recipe below
⚠️ **Your shell state does not survive between tool calls.** Each Bash call starts a
fresh shell, so `$API`, `$SELF`, the `CURL` array and `delete_session` are all gone by
the next call, and `$$` is a different pid. Three consequences, all of which have
teeth:
- **Re-run this entire preamble at the top of every Bash call that touches the API.**
Running it once and assuming it stuck is the single most likely way to break a run.
- **Never re-paste only half of it.** The delete guard below is written so that a
missing definition deletes nothing, but that only holds if you never hand-roll a
`DELETE` of your own.
- **Never put `$$` in a `clientId`.** It changes per call, so the "resend the identical
request" loop in §3 would stop being a duplicate and would **retype the prompt**,
submitting the turn twice. Use a fixed literal (`codeman-agent-1` below).
Only real environment variables (`CODEMAN_*`) survive, which is why this preamble
rebuilds everything else from them.
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Codeman does NOT hand a session the server password. If one is set, the two
# in-reach copies are the data dir's .env (the same fallback `codeman attach`
# uses — hand-authored; nothing ever writes it) and the supervisor definition
# that install.sh wrote the password into, which is where a stock
# password-protected install actually keeps it. The data dir is wherever the
# hook-secret file lives. Values may be quoted or `export`-prefixed.
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
fi
if [ -z "${CODEMAN_PASSWORD:-}" ]; then # stock installs: install.sh puts it in the service definition
UNIT="$HOME/.config/systemd/user/codeman-web.service"
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if [ -f "$UNIT" ]; then
# install.sh backslash-escapes " and \ in the unit value; undo it or a password
# containing either recovers wrong and auth fails.
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1 | sed 's/\\\(["\\]\)/\1/g')
elif [ -f "$PLIST" ]; then
# install.sh XML-escapes the plist value; undo it (&amp; LAST, mirroring escape order).
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p' \
| sed -e 's/&lt;/</g' -e 's/&gt;/>/g' -e 's/&amp;/\&/g')
fi
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-signed cert)
# Fail-CLOSED session delete. The DELETE lives INSIDE the guard on purpose: the older
# `is_self "$SID" || curl -X DELETE ...` shape failed OPEN, because an undefined
# is_self exits 127 and the `||` branch then ran the delete completely unguarded.
# Undefined delete_session is "command not found", which deletes nothing.
delete_session() {
local id="${1:-}"
[ -n "$id" ] || { echo "refusing: empty session id"; return 1; }
[ "${#SELF}" -ge 8 ] || { echo "refusing: \$SELF unset or too short to prove this is not me"; return 1; }
# ids appear in full AND 8-char form (Docker exports a truncated $SELF; mux names and
# UI surfaces carry 8-char ids), so compare by prefix in BOTH directions. Equality or
# a one-directional check each miss a real combination, and the miss deletes you.
case "$id" in "$SELF"*) echo "refusing: $id is me"; return 1 ;; esac
case "$SELF" in "$id"*) echo "refusing: $id is me"; return 1 ;; esac
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
}
CID=codeman-agent-1 # FIXED literal, never "agent-$$" (see §0)
```
- If `CODEMAN_MUX` is not `1`, **stop and say so**. Do not guess an API URL; a server
you are not part of is not yours to drive.
- **A 401 is plain text, not the JSON envelope**, so on a password-protected server
every `jq` in these recipes dies with `jq: parse error` instead of showing
`UNAUTHORIZED`. If that happens, check the status with `-w '%{http_code}'`; if it
is 401 and neither fallback above found a credential, **stop and tell the user
you need credentials**. The hook-secret bypass covers only `/api/hook-event` and
`/api/status-telemetry`, never session control.
- These endpoints first ship in Codeman **1.13.0**, but do not gate on the version
number: a dev build can serve them while reporting an older version. Probe
instead: `GET .../wait` on a real session id answering 404 with an `.error`
starting `Route ` means the server predates the wait endpoints (fall back to
polling `GET .../terminal?tail=` and say so); `Session ... not found` means your
session id is wrong, not the server.
## 1. Safety rules — read before any mutating call
You are yourself a session on this server, and the API has **no undo**.
- **Never act on your own session, and know that `delete_session` is the ONLY guard.**
The server has no self-protection: a session that DELETEs its own id succeeds and
dies silently (verified live). **Always delete through `delete_session "$SID"` from
§0; never write a bare `curl -X DELETE` and never reintroduce the
`is_self … || curl -X DELETE …` shape.** That older form failed open: with the
function undefined (a half-re-pasted preamble, see §0) bash returns 127, the `||`
branch fires, and the delete runs with no self-check at all. Wrapping the request
inside the guard is what makes a lost preamble delete nothing instead of deleting
you. Apply the same prefix-both-directions reasoning before any kill, respawn, or
input call you write by hand.
- **Mutating calls you may make unprompted** (this is an allowlist):
`POST /api/v1/quick-start`, `POST /api/v1/sessions/:id/input`, and
`DELETE /api/v1/sessions/:id` **only** for a session you created in this
conversation, by exact id. Keep a list of the ids you create. Everything else
mutating needs the user to have asked for it.
- **Never call these** unless the user explicitly asked, naming the target:
- `DELETE /api/cases/:name` — recursively **deletes a real directory of the user's
code** from disk. One wrong case name destroys work that was never yours.
- `DELETE /api/sessions` (no id) and `DELETE /api/subagents` (no id) — bulk kills.
- respawn / ralph / orchestrator / cron mutations — respawn runs `/clear` (wipes a
conversation), orchestrator state is a single global slot, cron jobs outlive you.
- `PUT /api/settings`, `POST /api/system/update` — global UI settings; server restart.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. The API is the only interface.
- Sessions count against a 50-session cap and case creation is uncapped: clean up every
session you start, and don't retry `quick-start` in a loop.
## 2. Rules of the road
- **End every input with `\r`** — literally the two characters `\r` inside the JSON
string. Codeman types the text and sends Enter **only when the input contains a
carriage return**; without it your command sits unsubmitted on the worker's prompt
and everything downstream times out. `{"input":"run the tests\r",...}`. No response
field catches this: `delivered:true` means "written to the pane", **not**
"submitted" — a `\r`-less send still reports `delivered:true` and then every wait
times out, which is why the loops below are bounded and check the terminal.
- **Single-line input only.** Newlines are stripped; one line per call.
- **Build request bodies with `jq -n` for any prompt you did not author as a
literal.** The inline `-d '{"input":"'"$P"'\r"}'` pattern breaks on the first
double quote, backslash, or `$` in a real prompt:
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
- **Exactly-once delivery**: always send a stable `clientId` and a monotonic
per-session `seq` on `POST .../input`. A retry after a dropped connection then
cannot double-type the prompt. Increment `seq` for each NEW input; reuse the same
pair only to re-ask about the same delivery.
- **Envelope**: success is `{"success":true,"data":…}`, errors are
`{"success":false,"error","errorCode"}`. Read `.data`. Use `/api/v1/*` paths.
- **A wait timeout is HTTP 200**, `{wait:{timedOut:true,signal:null}}` — not an error.
Loop over short waits (60 s); proxies cut long-idle connections. Timeouts are
**clamped** (ceiling 600 s): read back `wait.timeoutMs` for what was applied.
- **`stop` and `blocked` fire for `claude` sessions only** (Claude Code hooks). On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`, requesting them explicitly is a
400 — and lifecycle transitions there are coarse (a short shell command may emit
**no** `idle` transition at all, verified live), so synchronize those modes with
output markers, not signals.
- **Your typed command echoes into the output stream**, so a marker that appears
verbatim in the input line matches **before the command runs**. Always split the
marker (recipe below), keep it unique per call, and use `from=buffer` so a marker
that printed before your wait landed is still found. Matching is literal — no regex.
- **Match single space-free tokens against TUI output.** A full-screen TUI (claude,
codex, …) positions text with cursor movements, not literal spaces, so the stripped
stream can read `Yes,Itrustthisfolder` and a multi-word match is unreliable there —
whether a phrase keeps its spaces depends on how the TUI happened to draw it
(observed live: some match, some never fire). Plain command output (shell workers,
`echo` lines) keeps real spaces.
## 3. Recipes (each verified live)
**List sessions / find yourself** — metadata only, safe to poll:
```bash
"${CURL[@]}" "$API/api/v1/sessions" | jq '.data[] | {id, name, mode, status}'
"${CURL[@]}" "$API/api/v1/sessions" | jq --arg s "$SELF" '.data[] | select(.id | startswith($s))'
```
**Start a claude worker and wait until it is actually ready.** A new session reports
`idle` before its CLI has spawned, and a brand-new case shows a **trust dialog**
first, so neither "wait for idle" nor "wait for ❯" means ready (the trust dialog
contains `❯` too — observed live). Codeman *can* auto-accept that dialog itself, but
the accept rides a stream match that misses on some runs (both outcomes seen live),
so wait for the composer first and handle the dialog only as the bounded fallback —
never send a blind Enter up front (if auto-accept already fired, it lands in the
composer). Stage 1 is short on purpose: an already-trusted case matches `bypass` in
under a second, while a **virgin case can never pass stage 1** (the dialog is up, so
the composer is not) and always pays it in full before the fallback runs — the long
budget belongs to stage 3, after the dialog is answered:
```bash
# ALWAYS check .success: on failure `.data.sessionId` is null, jq -r prints the string
# "null", and the flow below then burns its full readiness budget against
# /api/v1/sessions/null before reporting jq noise instead of the actual cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
if [ -z "$SID" ]; then
# SESSION_BUSY here is the 50-session cap, not the waiter cap; FORBIDDEN/CONFLICT/
# OPERATION_FAILED/INVALID_INPUT are the others. None are retryable in a loop.
jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed; stopping."
exit 1
fi
for _ in $(seq 1 30); do # bounded: a bad SID would otherwise poll forever
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# ⚠️ pid != null proves STARTUP only, never life: a worker that later dies inside
# its pane keeps status "idle" and a pid (the local tmux attach client, not the
# worker). The death check is wait?until=exit, below.
SEQ=1 # $CID came from the §0 preamble; do NOT rebuild it from $$
# the composer's status bar ("bypass permissions on") is the ready marker — Codeman
# spawns claude in bypass mode. Single-token matches only: TUI text is space-less.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# composer never appeared → the trust dialog is probably still up; accept it once
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null || \
{ echo "worker $SID never became ready; inspect terminal?tail="; }
fi
```
**Send a prompt and wait for the turn to finish** (claude workers — the call to
prefer). It registers the waiter *before* typing, closing the race where a separate
wait sees the previous turn's idle state. Loop by resending the **identical** request:
the repeat is a tagged duplicate (same `clientId`+`seq`) that does not retype but
answers from the session's current state. Verified: the stop hook resolves this in
seconds; a duplicate resend answers in ~20 ms without retyping.
```bash
for TRY in $(seq 1 10); do # BOUNDED: a \r-less send never produces a signal and resends are no-op duplicates
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"run the tests, then summarize in one line\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ',"wait":true,"waitTimeout":60000}')
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # two straight timeouts: prompt sitting unsubmitted?
continue
fi
# Resolved — but a duplicate answering immediately reports the session's CURRENT
# state ("it is idle now"), NOT that a new turn ran. A \r-less send lands exactly
# here on try 2 (verified live), so check the terminal before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' | tail -5
# your prompt still on the ❯ composer line = never submitted (missing \r);
# submit it with {"input":"\r"} (the only recovery), then loop again
fi
break
done
SEQ=$((SEQ+1)); jq '.data.wait.signal, .data.status' <<<"$R"
```
Read the outcome in this order: `wait.signal != null` → done (`stop` is definitive;
`idle` is heuristic) — **unless** it arrived as `duplicate:true` + `immediate:true`,
which only says the session is idle *now* and must be confirmed from the terminal
(above); `wait.timedOut` → loop again (bounded); `wait.ended` → session gone, stop.
If the loop exhausts its cap, do not keep looping: read the terminal, report what
you see, and remember that a still-typed-but-unsubmitted prompt (missing `\r`) can
only be recovered by submitting it with `{"input":"\r"}`.
**Shell worker + completion marker** — the pattern for `shell` mode (no hooks there).
The typed line must not contain the marker verbatim (the input echo would match
instantly — observed live), so build it with a variable the worker's shell expands:
```bash
N="${RANDOM}_$$"; MARK="DONE_$N" # unique per call: tmux repaints replay old text
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}'
SEQ=$((SEQ+1))
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=120000' \
| jq -r '.data.wait | {matched, snippet}'
```
The typed line shows `${M}_…`, the real output shows `DONE_… rc=<exit code>`, and the
snippet carries the exit code back to you.
**Read a worker's answer.** For `claude` and `codex` workers this is the read path:
`last-response` returns the agent's final message as clean text, taken from the
transcript rather than the screen, so it carries none of the TUI's box-drawing or
repaint noise.
```bash
for _ in $(seq 1 10); do # the transcript write LAGS the stop signal
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
```
`.data` is `{text, timestamp}`. ⚠️ **Poll it, do not read it once.** `text` is written
from the transcript file, which is flushed slightly *after* the `stop` hook fires, so a
single read taken the instant send-and-wait returns comes back `""` even though the
turn finished (verified live: empty on the first call, full text seconds later). `text`
is also `""` before the worker's first completed turn, and always `""` for modes with
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, verified live), which is
why the loop above is bounded rather than open-ended. Fall back to the terminal buffer there, tail in **bytes**
(`textOutput` in `GET .../output` stays empty for interactive sessions; don't use it):
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=3000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
```
⚠️ Do not use that pipeline to read a **claude/codex** answer. A full-screen TUI draws
with cursor moves, so the stripped buffer is largely one long line: `tail -30` has
almost nothing to split on and you get a wall of repaint noise with the answer buried
in it (verified live, side by side with `last-response` returning the exact prose).
The terminal buffer is for *diagnosis* (is my prompt sitting unsubmitted?), not for
reading answers. Avoid `?full=1` (entire tmux scrollback, a context bomb) unless doing
a post-mortem.
**Detect a dead worker cheaply**: `GET .../wait?until=exit&timeout=60000` answers
immediately (`signal:"exit"`, `immediate:true`) if the PTY is gone — including a
worker that exited *inside* its pane, which `GET .../sessions/:id` keeps reporting
as `status:"idle"` with a pid (that pid is the local tmux attach client, not the
worker). The wait routes are the only liveness check; a worker dying while a wait
is parked resolves it within ~3 s. A session deleted mid-wait resolves in ~1 s.
**Clean up** — only ids you created, one at a time, always through the §0 helper:
```bash
delete_session "$SID"
```
Everything else (endpoint tables, per-mode signal table, error codes, capacity
limits, Docker/remote caveats): [reference/endpoints.md](reference/endpoints.md).
Fan-out orchestration and blocked-worker handling:
[reference/recipes.md](reference/recipes.md).
+229
View File
@@ -0,0 +1,229 @@
# Codeman API reference for agents
Loaded on demand from the `codeman` skill. Assumes the guard variables from SKILL.md
(`$API`, `$SELF`, `"${CURL[@]}"`). Canonical contract: `docs/api-reference.md` in the
Codeman repo; this file is the agent-relevant subset, verified live.
## Envelope and errors
Every JSON response: `{"success":true,"data":…}` or
`{"success":false,"error":"…","errorCode":"…"}`. Branch on `errorCode`:
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| `INVALID_INPUT` | 400 | malformed request; the message names the bad field |
| `UNAUTHORIZED` | 401 | auth required or failed (send `-u user:password`). ⚠️ The 401 body is plain text, NOT this envelope — `jq` dies with a parse error, see the guard in SKILL.md |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | on a **wait**: this session's waiter cap (16, combined signal+output) is full. On **quick-start**: the 50-session cap is full, so clean up before starting more |
| `CONFLICT` / `ALREADY_EXISTS` | 409 | conflicts with current state |
| `OPERATION_FAILED` | 422 | well-formed but could not be completed |
| `RATE_LIMITED` | 429 | per-owner or process-wide waiter pool is full — back off; switching sessions will not help |
| `INTERNAL_ERROR` | 500 | server bug |
`SESSION_BUSY` vs `RATE_LIMITED` on the wait endpoints is deliberate: the first means
"too many waiters on *this* session", the second means the *pool* is full.
## Sessions
| Task | Call |
|------|------|
| list sessions (metadata only, ~1.5 KB each, safe to poll) | `GET /api/v1/sessions` |
| one session (has `.data.pid`, `null` until the PTY spawns) | `GET /api/v1/sessions/:id` — ⚠️ **not a liveness check**: a worker that dies inside its pane keeps `status:"idle"` and a pid (the tmux attach client); `wait?until=exit` is the death check |
| unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check |
| start case + session in one call | `POST /api/v1/quick-start` |
| send input | `POST /api/v1/sessions/:id/input` |
| **read a worker's answer** (claude/codex) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}` — clean transcript text, no TUI noise. ⚠️ **Poll it**: the transcript flush lags the `stop` signal, so a read taken the instant send-and-wait returns is `""` (verified live). Also `""` before the first completed turn, and always `""` for `shell`/`opencode`/`gemini`/`antigravity` (no transcript) |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer` — for *diagnosis* (unsubmitted prompt?), not for reading answers |
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
| background agents, one session | `GET /api/v1/sessions/:id/subagents` |
| background agents, global list | `GET /api/v1/subagents` (admin-only in multi-user mode) |
| server status / version | `GET /api/v1/status` → `.data.version` |
| delete one session (yours only, via `delete_session`) | `DELETE /api/v1/sessions/:id` — never call it bare; the fail-closed helper in SKILL.md §0 is the only self-protection that exists |
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Use
`last-response` for claude/codex answers; only fall back to `terminal?tail=` for
hook-less modes, or to diagnose a prompt that was never submitted, and strip ANSI:
```bash
… | jq -r '.data.terminalBuffer' | sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g'
```
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
— `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing — do not retry it in a loop, and remember the name.
⚠️ **Branch on `.success` before reading `.data.sessionId`.** On any failure the field
is absent, `jq -r` prints the literal string `null`, and every later call then targets
`/api/v1/sessions/null`, burning the full readiness budget and reporting jq noise
instead of the real cause. Failure modes here are `SESSION_BUSY` (the **50-session
cap**, not the waiter cap), `FORBIDDEN`, `CONFLICT`, `OPERATION_FAILED` and
`INVALID_INPUT`; none of them are retryable in a loop.
⚠️ `caseName` resolves through the linked-cases registry first, so a name that happens
to match a case the user linked in lands in that **real repo**, not a fresh scratch
directory. Pick distinctive scratch names, and use a linked name deliberately when you
do want a worker in an existing checkout.
`POST /api/v1/sessions/:id/input` body:
`{"input":"one line\r","useMux":true,"clientId":"agent-1","seq":1}` plus optionally
`"wait"` / `"waitTimeout"` (below).
- ⚠️ **The input must contain `\r`** (the JSON escape, i.e. a real carriage return)
**or Enter is never sent**: the text is typed onto the worker's prompt and sits
there unsubmitted. Verified live — this is the number-one silent failure, and no
response field catches it: `delivered:true` means "written to the pane", not
"submitted". A `\r`-less send with `wait` reports `delivered:true` and then every
wait on that turn times out. Without `wait`, fire-and-forget returns an **empty**
`{"success":true,"data":{}}` — no `delivered`, no `duplicate`; those fields exist
only on the `wait` variant, so a fire-and-forget flow gets no delivery
confirmation at all.
- `input` must be single-line (newlines are stripped). To send a bare Enter (confirm
a dialog), send `{"input":"\r"}`.
- `clientId`+`seq` give exactly-once delivery: the server applies each pair at most
once. Increment `seq` per new input.
## The wait primitives
Three bounded long-polls. Shared semantics:
- **Timeout = HTTP 200** with `wait.timedOut:true`. Loop over short waits (60 s);
`tailscale serve` / cloudflared cut idle connections.
- Timeouts are **clamped** to `[1000, 600000]` ms (operator-tunable); the applied
value is echoed as `wait.timeoutMs` — read it back, never assume.
- All three nest the result under `.data.wait`, same shape, so one helper parses all.
- `.data.status` (post-wait `SessionStatus`) and `.data.limitPaused` ride along.
`limitPaused:true` means the session is paused on a usage limit and will emit
nothing until reset — a timeout is then *expected*; do not retry hard, and do not
kill the worker.
### Signals by mode
| Signal | Meaning | Available for |
|--------|---------|---------------|
| `idle` | output stabilized + prompt detected — heuristic, can flap mid-turn | every mode |
| `working` | session started producing output | every mode |
| `stop` | Claude Code `stop` hook — the definitive end-of-turn | `claude` only |
| `blocked` | `permission_prompt` / `elicitation_dialog` hook — the worker needs an answer | `claude` only |
| `exit` | PTY exited or session deleted | every mode |
Default `until` set: `stop,idle,exit`. On non-claude modes the server silently drops
`stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
`["idle","exit"]` on shell); requesting them *explicitly* there is a 400 naming the
mode. ⚠️ On hook-less modes the lifecycle signals are also **coarse in practice**: a
short shell command produced **no** `idle` transition within 60 s (verified live), so
a `fresh=1` / fresh-delivery wait can burn its whole timeout while the work finished
long ago. Synchronize hook-less modes with `wait-output` markers instead.
Two more places hooks go missing even in claude mode: **Docker cases** need
`CODEMAN_DOCKER_BRIDGE_HOOKS=1` on the server (without it only `idle`/`working`/
`exit` arrive), and **remote-SSH cases** run the agent on another host whose hooks may
never reach this server. When unsure, ask for `stop,idle,exit`.
⚠️ **Signals are edge-triggered with no history.** A signal that fires while no
waiter is registered is gone; no later wait can observe it (`until=stop` on a worker
whose turn already ended just times out, with or without `fresh` — verified live).
Register the waiter before the event can happen: send-and-wait does exactly that,
and `wait-output` markers with `from=buffer` are latched by construction. Never
fire-and-forget N prompts and then gather signal-waits worker by worker; every
worker that finishes before its gather is unobservable (see recipes.md Flow 3b).
### `GET /api/v1/sessions/:id/wait`
| Param | Default | Notes |
|-------|---------|-------|
| `until` | `stop,idle,exit` | comma list; unknown token → 400 naming it |
| `timeout` | 60000 | ms, clamped; applied value echoed as `wait.timeoutMs` |
| `fresh` | `0` | `1` requires an actual *transition*, ignoring the state at call time |
⚠️ A session whose PTY has not spawned (`pid:null`) or has exited counts as `exit`
**right now**: with the default set the call answers immediately
(`signal:"exit", immediate:true`). That is how you detect a dead worker cheaply — but
it also means "wait for my just-created session" needs the readiness recipe in
SKILL.md, not this endpoint.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Default | Notes |
|-------|---------|-------|
| `match` | required | literal substring, 1–200 chars, ANSI-stripped; chunk-straddling matches found; **no regex** — a `regex=` param is a 400 |
| `nocase` | `0` | case-insensitive compare; snippet keeps original casing |
| `from` | `now` | `buffer` scans the tail (~256 KB) of existing output first |
| `timeout` | 60000 | same clamp |
Four traps, all observed live:
1. **The echo of your own typed command is output.** A marker appearing verbatim in
the input line matches the moment the text is typed, before the command runs.
Split the marker with a shell variable: send `M=DONE; …; echo ${M}_1234\r`, wait
on `DONE_1234`.
2. **`from=now` misses text printed before the wait landed** — a marker echoed just
before the request registered timed out at full length. After sending a command,
always wait with `from=buffer`.
3. **`from=now` can also match too much**: tmux repaints old screen content as
ordinary output on attach/resize/redraw, so a *generic* marker (`BUILD OK`)
matches stale text. Unique-per-call markers (`DONE_$RANDOM`) make both `from`
modes safe.
4. **TUI output can be space-less in the stream.** Full-screen TUIs (claude, codex,
…) position words with cursor-movement escapes rather than literal spaces, so
the stripped stream can read `Yes,Itrustthisfolder` while the pane shows the
spaced phrase. Whether a given phrase keeps its spaces depends on how the TUI
drew it (observed live: some multi-word matches fire, some never do), so treat
multi-word matches against TUI screens as unreliable and match a **single
space-free token** (`trust`, `bypass`). Plain command output (shell workers,
`echo` lines) keeps real spaces and multi-word matches work there.
Build the query with `-G --data-urlencode` (a `+` in a hand-built query decodes to a
space). Result extras: `wait.matched`, `wait.match`, `wait.snippet` (bounded window
around the match, blank runs collapsed — the snippet is often all you need to read).
### `POST /api/v1/sessions/:id/input` with `wait`
| Field | Notes |
|-------|-------|
| `wait` | `true` (default signal set) or the same comma grammar as `until`; absent = historical fire-and-forget |
| `waitTimeout` | ms, same clamp |
Registers the waiter **before** typing, which closes the race where send-then-wait
sees the previous turn's idle state and returns instantly. Response adds `delivered`
and `duplicate` beside the standard `wait` object.
A **tagged duplicate** (same `clientId`+`seq` already applied) does not retype but
still honors `wait`, answering from the session's *current* state instead of
requiring a new transition (`delivered:false, duplicate:true` — verified: ~20 ms,
command ran exactly once). That is what makes the resend-identical-request loop in
SKILL.md correct: iteration 1 delivers and needs a transition; later iterations
resolve immediately if the turn ended in between. ⚠️ The flip side: a duplicate's
`immediate:true` answer is the current state and nothing more — an idle worker
whose prompt was never submitted (missing `\r`) produces the same
`signal:"idle", immediate:true` as one that finished the turn. Confirm from
`terminal?tail=` before reporting success; SKILL.md's loop shows where.
### Outcome parsing, in order
1. `wait.signal != null` (or `wait.matched == true`) — the thing happened.
`wait.immediate:true` rides along and means the condition already held at call
time; if that is not what you meant, you wanted `fresh=1` or send-and-wait.
2. `wait.timedOut` — poll boundary; loop again.
3. `wait.ended` — session deleted/torn down mid-wait; stop looping.
## Troubleshooting
| Symptom | Cause / fix |
|---------|-------------|
| every curl fails with a certificate error | you dropped `-k`; `CODEMAN_API_URL` is HTTPS with a self-signed cert |
| `jq: parse error` on every call | plain-text 401s: the server has a password. Check with `-w '%{http_code}'`, use the guard's `.env` fallback, and if no `.env` exists, stop and ask the user for credentials |
| input arrives but nothing happens; later waits all time out | the input had no `\r`, so Enter was never sent; the text is sitting on the worker's prompt. **Submitting it with `{"input":"\r"}` is the ONLY recovery** — Ctrl+U (0x15) and Esc do NOT clear the composer (verified live) — and the flush costs one turn in which the worker reasons about the junk; open the next real prompt with "ignore the garbled line above:" |
| `GET .../sessions/$CODEMAN_SESSION_ID` 404s | Docker case: the env id is truncated to 8 chars; find yourself with `startswith($SELF)`, and always self-compare by prefix, in both directions |
| `CODEMAN_MUX` unset but you seem to be in a session | remote-SSH case: the env vars are not exported there. Fail closed — refuse to act |
| connection refused from inside a container | a loopback-bound server is unreachable from a container, and `CODEMAN_DOCKER_BRIDGE_HOOKS=1` does **not** fix that: it opens a hooks-only listener, so hook events start flowing but `/api/v1/*` stays refused. Driving the API from inside a Docker case needs a reachable bind (an operator decision); report it, don't retry |
| wait routes 404 on a valid session id | read the `.error` text: a `Route ...` prefix means the server predates the wait endpoints (< 1.13.0; a dev build can serve them while reporting an older version, so probe, never version-compare) — poll `terminal?tail=` and say so. `Session ... not found` means your id is wrong, not the server |
| wait on `stop` never resolves | non-claude mode, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept missed; use the readiness recipe in SKILL.md (wait for `bypass` first, accept the dialog only as the bounded fallback) |
| `wait-output` times out although the pane shows the text | multi-word match against a TUI screen; the stream has no spaces there — match one token |
| `wait-output` matched instantly with stale text | generic marker + tmux repaint; use `DONE_$RANDOM` |
| 409 `SESSION_BUSY` on a wait | too many concurrent waiters on that session (cap 16 combined); reuse one wait per worker |
| 429 `RATE_LIMITED` on a wait | global/owner waiter pool full; back off, do not switch sessions |
+279
View File
@@ -0,0 +1,279 @@
# Worked orchestration flows
Loaded on demand from the `codeman` skill. Every flow assumes the SKILL.md §0 preamble
is in scope (`$API`, `$SELF`, `$CID`, `"${CURL[@]}"`, `delete_session`).
⚠️ **That preamble does not survive between tool calls**, so re-run it at the top of
every Bash call that uses these flows, in full. Re-pasting only part of it is the
failure mode the fail-closed `delete_session` exists to contain, and a `clientId` you
rebuild from `$$` changes per call, which turns the duplicate-resend loop in Flow 1
into a second typed prompt.
Track every session id you create; delete them (and only them) when done. The two
silent killers: **every input ends with `\r`**, and **markers must be split** so the
typed-line echo does not match them.
## Flow 1: claude worker, end to end
Start a worker, get it truly ready (trust dialog included), give it a task, wait for
the turn to finish, read the answer, clean up. Verified live: the stop hook resolves
the send-and-wait within seconds of the turn ending.
```bash
# 1. start (returns before the CLI inside is ready). ALWAYS check .success: on failure
# .data.sessionId is null, jq -r yields the string "null", and every step below
# then runs against /api/v1/sessions/null and reports jq noise, not the cause.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-tests","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed"; exit 1; }
CREATED+=("$SID") # the cleanup list
SEQ=1 # $CID is the fixed literal from §0; never rebuild it from $$
# 2. readiness. "wait for idle" or "wait for ❯" is NOT readiness: a fresh session
# reports idle before anything spawned, and the first-run trust dialog contains ❯.
# Codeman CAN auto-accept that dialog, but the accept misses on some runs (both
# outcomes seen live), so: composer marker first, dialog only as the bounded
# fallback (a blind Enter up front would land in an already-ready composer).
# Stage 1 is SHORT on purpose: an already-trusted case matches in <1 s, while a
# virgin case can never pass it (the dialog is up) and always pays it in full —
# the long budget belongs to stage 3, after the dialog is answered.
# Single-token matches only: TUI text is space-less in the stream.
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# (pid != null proves startup only — a worker that later dies inside its pane keeps
# status "idle" and a pid. The death check is wait?until=exit.)
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null || echo "worker $SID not ready; inspect terminal?tail="
fi
# 3. send-and-wait, looping on the IDENTICAL request (tagged duplicate: no retype).
# BOUNDED (a \r-less send would otherwise loop forever), body built with jq -n so
# quotes/backslashes/$ in a real prompt survive; note the appended \r.
PROMPT='run the unit tests and summarize failures in one line'
BODY=$(jq -n --arg p "$PROMPT" --arg c "$CID" --argjson s "$SEQ" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:60000}')
for TRY in $(seq 1 10); do
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' --data-binary "$BODY")
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
jq -e '.data.limitPaused' <<<"$R" >/dev/null && sleep 60 # usage-limit pause: silence is expected
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # is the prompt sitting unsubmitted?
continue
fi
# Resolved — but duplicate + immediate is only "the session is idle NOW", which a
# never-submitted (\r-less) prompt also produces. Check before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5
# prompt still on the ❯ composer line = never submitted; {"input":"\r"} is the
# only recovery, then loop again
fi
break
done
SEQ=$((SEQ+1))
# 4. interpret
case "$(jq -r '.data.wait.signal' <<<"$R")" in
stop) : ;; # definitive end of turn
idle) : ;; # heuristic — and if it rode a duplicate with
# immediate:true, it proves nothing ran (step 3)
exit) echo "worker died" ;;
null) jq -e '.data.wait.ended' <<<"$R" >/dev/null && echo "worker deleted mid-wait" ;;
esac
# 5. read the answer. For a claude worker this is last-response: clean transcript text,
# no TUI repaint noise. Do NOT scrape the terminal for this — a full-screen TUI
# draws with cursor moves, so the stripped buffer is nearly one long line and the
# answer arrives buried in redraw garbage.
# POLL it: the transcript flush lags the stop signal, so a single read taken the
# instant step 3 returned comes back "" even though the turn finished (verified live).
for _ in $(seq 1 10); do
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# (.data is {text,timestamp}; text is also "" before the first completed turn and
# always "" for shell/opencode/gemini/antigravity, which have no transcript — use
# the terminal tail there, and here only to diagnose an unsubmitted prompt.)
# 6. clean up — exact id, own list only, through the fail-closed §0 helper
delete_session "$SID"
```
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
re-ask about the same delivery (the duplicate-wait loop above).
## Flow 2: shell worker running a build, marker-synchronized
`shell` sessions have no hooks (`stop`/`blocked` are a 400 there), and their lifecycle
signals are coarse — a short command may emit no `idle` transition at all (verified
live), so send-and-wait can burn its whole timeout. The reliable pattern is a split,
unique marker plus `wait-output from=buffer`:
```bash
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"builder","mode":"shell"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "quick-start failed"; exit 1; }
CREATED+=("$SID")
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# Split marker: the typed line carries ${M}_N, only the OUTPUT carries DONE_N.
# An unsplit marker matches the echo of your own keystrokes before the build runs.
N="${RANDOM}_$$"; MARK="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"codeman-build-1","seq":1}'
for TRY in $(seq 1 30); do # BOUNDED (30 min): a \r-less send makes an uncapped loop infinite
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null && break
jq -e '.data.wait.ended' <<<"$R" >/dev/null && { echo "worker gone"; break; }
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # command still sitting unsubmitted?
done
jq -r '.data.wait.snippet' <<<"$R" # e.g. "DONE_123_456 rc=0" — the exit code rides the marker line
```
## Flow 3: fan out N workers, gather as each finishes
Start everything first, then gather. One in-flight wait per worker — the per-session
waiter cap is 16 and abandoned concurrent waits pile up against it.
```bash
declare -A WORKER MARKS
for task in lint typecheck unit; do
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"fan-'"$task"'","mode":"shell"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; echo "$task: spawn failed"; continue; }
WORKER[$task]=$SID; CREATED+=("$SID")
done
for task in "${!WORKER[@]}"; do
SID=${WORKER[$task]}
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
N="${task}_${RANDOM}"; MARKS[$task]="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run '"$task"'; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"codeman-fan-'"$task"'","seq":1}'
done
for task in "${!WORKER[@]}"; do # sequential gather; each wait blocks until that worker is done
for TRY in $(seq 1 30); do # BOUNDED per worker, same reasoning as Flow 2
R=$("${CURL[@]}" -G "$API/api/v1/sessions/${WORKER[$task]}/wait-output" \
--data-urlencode "match=${MARKS[$task]}" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
jq -e '.data.wait.matched or .data.wait.ended' <<<"$R" >/dev/null && break
done
echo "$task: $(jq -r '.data.wait.snippet // "worker gone"' <<<"$R" | tail -1)"
done
```
## Flow 3b: fan out N CLAUDE workers
Send-and-wait is synchronous, so the shell-flow shape ("send everything, then
gather") does not translate directly: the send *is* the wait, and worker 2's prompt
would not go out until worker 1's turn ended. Two working patterns, both verified
live (and one anti-pattern, measured failing, replaced by B):
**A. Background the send-and-waits** (simplest; each resolved on `stop` while the
other was still running):
```bash
sendwait() { # $1=sid $2=prompt $3=seq — assumes the worker passed Flow 1's readiness
local body; body=$(jq -n --arg p "$2" --argjson s "$3" --arg c "codeman-fan-$1" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:600000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$1/input" \
-H 'Content-Type: application/json' --data-binary "$body" > "/tmp/fan-$1.json"
}
( sendwait "$SID1" 'refactor module A and reply DONE' 2 & \
sendwait "$SID2" 'write tests for module B and reply DONE' 2 & wait )
jq -c '.data.wait | {signal, waitedMs}' /tmp/fan-"$SID1".json /tmp/fan-"$SID2".json
```
One in-flight wait per worker keeps you far from the 16-per-session waiter cap.
**B. Fire-and-forget, then gather with output markers.** If you must send every
prompt before waiting on anything, do **not** gather with signal waits: signals
are edge-triggered with no history, so a `stop` that fires before the gather
reaches that worker is gone and unobservable afterwards — `fresh=1` cannot help,
and neither can omitting it (measured: worker 2's turn ended at +2 s, its
sequential `until=stop,exit&fresh=1` gather burned its full bounded 300 s and
reported nothing). Gather instead on a marker each worker prints itself, which
`from=buffer` re-finds no matter when it appeared:
```bash
# SIDS[1], SIDS[2] = worker ids that already passed Flow 1's readiness.
# The typed prompt must NOT contain the finished marker verbatim (your keystrokes
# echo into the output stream and would match instantly), so ask for it in halves:
declare -A TOK
for i in 1 2; do
TOK[$i]="${RANDOM}_$i"
BODY=$(jq -n --arg p "do task $i; when completely done print the word WORKDONE immediately followed by _${TOK[$i]}" \
--arg c "codeman-fan-$i" --argjson s 2 '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/${SIDS[$i]}/input" \
-H 'Content-Type: application/json' --data-binary "$BODY"
done
for i in 1 2; do # order no longer matters: the marker is latched in the buffer
"${CURL[@]}" -G "$API/api/v1/sessions/${SIDS[$i]}/wait-output" \
--data-urlencode "match=WORKDONE_${TOK[$i]}" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=600000' | jq -c '.data.wait | {matched, snippet}'
done
```
Use A unless you genuinely need to send everything before waiting on anything: A
needs no marker discipline, and resolves on the definitive `stop` instead of on
the worker remembering to print a token.
## Flow 4: watch for a worker stuck on a permission prompt
Claude workers can block on a permission dialog. `blocked` is a wait signal
(claude-mode only), so watch for it and surface the question to the user instead of
guessing an answer:
```bash
R=$("${CURL[@]}" "$API/api/v1/sessions/$SID/wait?until=stop,blocked,exit&timeout=60000")
if [ "$(jq -r '.data.wait.signal' <<<"$R")" = blocked ]; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' | grep -v '^[[:space:]]*$' | tail -15
# show this to the user and ask how to answer; do NOT auto-confirm another
# session's permission prompt
fi
```
## Cleanup discipline
At the end of the conversation (or on abort), delete exactly what you created:
```bash
for id in "${CREATED[@]}"; do
delete_session "$id"
done
```
- Only ids from your own `CREATED` list. Never enumerate `/api/v1/sessions` and
delete by pattern; other sessions belong to the user.
- Always go through `delete_session`. It refuses an empty id, refuses when `$SELF` is
unset or too short to prove the target is not you, and prefix-checks in both
directions. A hand-written `curl -X DELETE`, or the old
`is_self "$id" || curl -X DELETE …`, has none of that: an undefined `is_self` exits
127 and the `||` branch deletes unguarded.
- If you created a *case* purely as scratch and the user confirmed it is disposable,
`DELETE /api/v1/cases/:name` removes it — but that recursively deletes the
directory from disk, so never do it without the user's explicit go-ahead for that
exact name.
+31 -3
View File
@@ -135,6 +135,7 @@ export abstract class AiCheckerBase<
// Active check state
protected checkMuxName: string | null = null;
protected checkTempFile: string | null = null;
protected checkStderrFile: string | null = null;
protected checkPromptFile: string | null = null;
protected checkPollTimer: NodeJS.Timeout | null = null;
protected checkTimeoutTimer: NodeJS.Timeout | null = null;
@@ -376,6 +377,7 @@ export abstract class AiCheckerBase<
const shortId = this.sessionId.slice(0, 8);
const timestamp = Date.now();
this.checkTempFile = join(tmpdir(), `${this.tempFilePrefix}-${shortId}-${timestamp}.txt`);
this.checkStderrFile = join(tmpdir(), `${this.tempFilePrefix}-stderr-${shortId}-${timestamp}.txt`);
this.checkPromptFile = join(tmpdir(), `${this.tempFilePrefix}-prompt-${shortId}-${timestamp}.txt`);
this.checkMuxName = `${this.muxNamePrefix}${shortId}`;
@@ -386,6 +388,7 @@ export abstract class AiCheckerBase<
// Ensure output temp file exists (empty) so we can poll it
writeFileSync(this.checkTempFile, '');
writeFileSync(this.checkStderrFile, '');
// Write prompt to file to avoid E2BIG error (argument list too long)
// The prompt can be 16KB+ which exceeds shell argument limits
@@ -396,7 +399,7 @@ export abstract class AiCheckerBase<
const modelArg = `--model "${this.config.model.replace(/"/g, '\\"')}"`;
const augmentedPath = getAugmentedPath();
const claudeCmd = `cat "${this.checkPromptFile}" | claude -p ${modelArg} --output-format text`;
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2>&1; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
const fullCmd = `export PATH="${augmentedPath}"; ${claudeCmd} > "${this.checkTempFile}" 2> "${this.checkStderrFile}"; echo "${this.doneMarker}" >> "${this.checkTempFile}"; rm -f "${this.checkPromptFile}"`;
// Spawn tmux session
try {
@@ -461,18 +464,32 @@ export abstract class AiCheckerBase<
const output = content.replace(this.doneMarker, '').trim();
if (!output) {
return this.createErrorResult(`Empty output from ${this.checkDescription}`, durationMs);
const stderr = this.readStderrDiagnostic();
const detail = stderr ? `: ${stderr}` : '';
return this.createErrorResult(`Empty output from ${this.checkDescription}${detail}`, durationMs);
}
// Delegate to subclass for verdict parsing
const parsed = this.parseVerdict(output);
if (!parsed) {
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"`, durationMs);
const stderr = this.readStderrDiagnostic();
const detail = stderr ? `; stderr: "${stderr}"` : '';
return this.createErrorResult(`Could not parse verdict from: "${output.substring(0, 100)}"${detail}`, durationMs);
}
return this.createResult(parsed.verdict, parsed.reasoning, durationMs);
}
private readStderrDiagnostic(): string {
if (!this.checkStderrFile || !existsSync(this.checkStderrFile)) return '';
try {
return readFileSync(this.checkStderrFile, 'utf-8').trim().substring(0, 200);
} catch {
return '';
}
}
private cleanupCheck(): void {
// Clear poll timer
if (this.checkPollTimer) {
@@ -509,6 +526,17 @@ export abstract class AiCheckerBase<
this.checkTempFile = null;
}
if (this.checkStderrFile) {
try {
if (existsSync(this.checkStderrFile)) {
unlinkSync(this.checkStderrFile);
}
} catch {
// Best effort cleanup
}
this.checkStderrFile = null;
}
if (this.checkPromptFile) {
try {
if (existsSync(this.checkPromptFile)) {
+615 -73
View File
@@ -12,15 +12,20 @@ import chalk from 'chalk';
import { createRequire } from 'module';
import http from 'node:http';
import https from 'node:https';
import { readFileSync } from 'node:fs';
import { isAbsolute } from 'node:path';
import { existsSync, readFileSync } from 'node:fs';
import { isAbsolute, join } from 'node:path';
import { homedir } from 'node:os';
import { dataPath } from './config/instance.js';
import { installAgentSkillInto, removeAgentSkillFrom, type AgentSkillApplyResult } from './hooks-config.js';
import { getSessionManager } from './session-manager.js';
import { getTaskQueue } from './task-queue.js';
import { getRalphLoop } from './ralph-loop.js';
import { getStore } from './state-store.js';
import { getErrorMessage } from './types.js';
import { isSupportedAttachmentExtension } from './attachment-registry.js';
import { daemonStatus, startDaemon, stopDaemon, type WebLaunchOptions } from './daemon-control.js';
import { installService, serviceStatus, uninstallService } from './service-installer.js';
import { isLoopbackBindHost, isUnauthenticatedNetworkAcknowledged } from './web/network-auth-policy.js';
const require = createRequire(import.meta.url);
const pkg = require('../package.json') as { version: string };
@@ -116,6 +121,112 @@ program
console.log(makeAttachmentMagicLink(filePath));
});
// ============ Skill Commands ============
/** Same registry the server resolves case names through (mirrors `case-routes.ts`). */
const LINKED_CASES_FILE = dataPath('linked-cases.json');
/**
* Case name to directory, checking `linked-cases.json` FIRST and falling back to the
* shared single-user cases dir. Mirrors `resolveCasePath()` in `case-routes.ts`, which
* is what the web UI and `quick-start` use. Without the linked-cases lookup this
* command rejected every case linked in from outside `~/codeman-cases` with
* "Case not found", even though the server resolved the same name fine.
*
* Sync and tolerant on purpose: a missing or malformed registry means "no linked
* cases", never a crash.
*/
function resolveCliCasePath(name: string): string {
try {
const linked = JSON.parse(readFileSync(LINKED_CASES_FILE, 'utf-8')) as Record<string, string>;
const target = linked?.[name];
if (typeof target === 'string' && target) return target;
} catch {
// no registry yet, or unreadable/invalid JSON: fall through to the cases dir
}
return join(homedir(), 'codeman-cases', name);
}
/**
* Resolve where `skill install` / `skill uninstall` operate. Global is
* `~/.claude/skills/codeman` (Claude Code's user-scope skill dir, read by every new
* session); `--case <name>` targets `<case>/.claude/skills/codeman`, resolved through
* `resolveCliCasePath()` above. The web server's automatic per-case injection
* (`agentSkillEnabled`) covers multi-user spaces; this CLI is a local operator tool
* and stays single-user.
*/
function resolveSkillTarget(options: { case?: string }): string {
if (options.case) {
const casePath = resolveCliCasePath(options.case);
if (!existsSync(casePath)) {
console.error(chalk.red(`✗ Case not found: ${casePath}`));
process.exit(1);
}
return join(casePath, '.claude', 'skills', 'codeman');
}
return join(homedir(), '.claude', 'skills', 'codeman');
}
/** Print an AgentSkillApplyResult for humans; exit non-zero when nothing was done. */
function reportSkillResult(result: AgentSkillApplyResult, target: string): void {
const messages: Record<AgentSkillApplyResult, { ok: boolean; text: string }> = {
installed: { ok: true, text: `Agent skill installed: ${target}` },
refreshed: { ok: true, text: `Agent skill refreshed (was stale): ${target}` },
unchanged: { ok: true, text: `Agent skill already up to date: ${target}` },
removed: { ok: true, text: `Agent skill removed: ${target}` },
absent: { ok: true, text: `Nothing to remove at ${target}` },
foreign: {
ok: false,
text: `${target} exists but is not Codeman-managed (no marker), refusing to touch it. Remove it yourself if you want the packaged skill there.`,
},
symlink: {
ok: false,
text: `${target} (or its parent) is a symlink, refusing to write through it.`,
},
};
const message = messages[result];
if (message.ok) {
console.log(chalk.green(`✓ ${message.text}`));
} else {
console.error(chalk.red(`✗ ${message.text}`));
process.exit(1);
}
}
const skillCmd = program
.command('skill')
.description('Manage the Codeman agent skill (lets an agent inside a session drive the API)');
skillCmd
.command('install')
.description('Install the agent skill globally (~/.claude/skills/codeman) or into one case')
.option('-g, --global', 'Install into ~/.claude/skills/codeman, picked up by every new session (the default)')
.option('-c, --case <name>', 'Install into <case>/.claude/skills/codeman instead (linked cases resolve too)')
.action(async (options: { global?: boolean; case?: string }) => {
try {
const target = resolveSkillTarget(options);
reportSkillResult(await installAgentSkillInto(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
skillCmd
.command('uninstall')
.description('Remove a Codeman-managed agent skill copy (never touches a user-authored one)')
.option('-g, --global', 'Remove from ~/.claude/skills/codeman (the default)')
.option('-c, --case <name>', 'Remove from <case>/.claude/skills/codeman instead (linked cases resolve too)')
.action(async (options: { global?: boolean; case?: string }) => {
try {
const target = resolveSkillTarget(options);
reportSkillResult(await removeAgentSkillFrom(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
// ============ Session Commands ============
const sessionCmd = program.command('session').alias('s').description('Manage Claude sessions');
@@ -466,47 +577,152 @@ function printStats(stats: ReturnType<ReturnType<typeof getRalphLoop>['getStats'
// ============ Utility Commands ============
/** What probing the web server found. */
interface WebServerProbe {
reachable: boolean;
/** The URL that answered, or the first candidate when nothing did. */
url: string;
statusCode?: number;
version?: string;
authRequired?: boolean;
/** Live session states from `/api/status`, when the probe could read them. */
sessions?: Array<{ status?: string }>;
}
/**
* GET `<base>/api/status` with a short timeout, tolerating the self-signed cert an
* `--https` install uses. ANY HTTP answer proves the server is up: a 401 just
* means it wants credentials (sent when available, same env → data-dir `.env`
* fallback as `codeman attach`).
*/
function probeWebServerAt(base: string): Promise<WebServerProbe | null> {
let url: URL;
try {
url = new URL('/api/status', base);
} catch {
return Promise.resolve(null);
}
const envFile = readCodemanEnv();
const username = process.env.CODEMAN_USERNAME || envFile.CODEMAN_USERNAME || 'admin';
const password = process.env.CODEMAN_PASSWORD || envFile.CODEMAN_PASSWORD;
const transport = url.protocol === 'https:' ? https : http;
const headers: Record<string, string> = { Accept: 'application/json' };
if (password) {
headers.Authorization = `Basic ${Buffer.from(`${username}:${password}`).toString('base64')}`;
}
return new Promise((resolve) => {
const req = transport.request(
{
protocol: url.protocol,
hostname: url.hostname,
port: url.port,
method: 'GET',
path: url.pathname,
rejectUnauthorized: false,
headers,
timeout: 3000,
},
(res) => {
const chunks: Buffer[] = [];
let received = 0;
res.on('data', (chunk: Buffer) => {
received += chunk.length;
if (received <= 1024 * 1024) chunks.push(chunk);
});
res.on('end', () => {
const statusCode = res.statusCode ?? 0;
if (statusCode === 401) {
resolve({ reachable: true, url: base, statusCode, authRequired: true });
return;
}
let version: string | undefined;
let sessions: Array<{ status?: string }> | undefined;
try {
const parsed = JSON.parse(Buffer.concat(chunks).toString('utf-8')) as {
data?: { version?: unknown; sessions?: unknown };
};
const data = parsed?.data ?? (parsed as { version?: unknown; sessions?: unknown });
if (typeof data?.version === 'string') version = data.version;
if (Array.isArray(data?.sessions)) sessions = data.sessions as Array<{ status?: string }>;
} catch {
// Not JSON, but still an answer, so still running.
}
resolve({ reachable: true, url: base, statusCode, version, sessions });
});
}
);
req.on('timeout', () => req.destroy(new Error('timeout')));
req.on('error', () => resolve(null));
req.end();
});
}
program
.command('status')
.description('Show overall status')
.action(() => {
const manager = getSessionManager();
const queue = getTaskQueue();
const loop = getRalphLoop();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
const storedValues = Object.values(stored);
const taskCounts = queue.getCount();
const loopStatus = loop.status;
// Use live sessions if available, otherwise fall back to stored state
const activeCount = sessions.length || storedValues.filter((s) => s.status !== 'stopped').length;
const idleCount = sessions.length
? sessions.filter((s) => s.isIdle()).length
: storedValues.filter((s) => s.status === 'idle').length;
const busyCount = sessions.length
? sessions.filter((s) => s.isBusy()).length
: storedValues.filter((s) => s.status === 'busy').length;
.description('Show whether the Codeman web server is running, plus session/task state')
.option('--url <url>', 'Server URL to probe (defaults to CODEMAN_API_URL, then local port)')
.action(async (options: { url?: string }) => {
// Issue #230: this command runs in its own fresh process, and the old output
// reported THAT process's (always-stopped) Ralph loop under a bare "Status:",
// reading as "the server is down" while the web service ran fine. Probe the
// real server first; the Ralph loop has its own `codeman ralph status`.
const port = process.env.CODEMAN_PORT || '3000';
const candidates = options.url
? [options.url]
: process.env.CODEMAN_API_URL
? [process.env.CODEMAN_API_URL]
: [`https://127.0.0.1:${port}`, `http://127.0.0.1:${port}`];
let probe: WebServerProbe = { reachable: false, url: candidates[0] };
for (const candidate of candidates) {
const answer = await probeWebServerAt(candidate);
if (answer) {
probe = answer;
break;
}
}
console.log(chalk.bold('\nCodeman Status'));
console.log('─'.repeat(40));
console.log(chalk.bold('\nSessions:'));
console.log(` Active: ${activeCount}`);
console.log(` Idle: ${idleCount}`);
console.log(` Busy: ${busyCount}`);
console.log(chalk.bold('\nWeb Server:'));
if (probe.reachable) {
const version = probe.version ? ` (v${probe.version})` : '';
console.log(` Status: ${chalk.green('running')}${version} at ${probe.url}`);
if (probe.authRequired) {
console.log(chalk.gray(' (answers 401: set CODEMAN_PASSWORD/CODEMAN_USERNAME to see session details)'));
}
} else {
console.log(` Status: ${chalk.red('not reachable')} at ${candidates.join(' or ')}`);
console.log(
chalk.gray(' (start it with `codeman web`, or check your service: systemctl --user status codeman-web)')
);
}
// Prefer the server's live view; fall back to the shared saved state, labeled
// as such, so the numbers are never silently a different thing.
if (probe.sessions) {
const live = probe.sessions;
console.log(chalk.bold('\nSessions (live, from the server):'));
console.log(` Total: ${live.length}`);
console.log(` Idle: ${live.filter((s) => s.status === 'idle').length}`);
console.log(` Busy: ${live.filter((s) => s.status === 'busy').length}`);
} else {
const manager = getSessionManager();
const storedValues = Object.values(manager.getStoredSessions());
console.log(chalk.bold('\nSessions (from saved state):'));
console.log(` Active: ${storedValues.filter((s) => s.status !== 'stopped').length}`);
console.log(` Idle: ${storedValues.filter((s) => s.status === 'idle').length}`);
console.log(` Busy: ${storedValues.filter((s) => s.status === 'busy').length}`);
}
const taskCounts = getTaskQueue().getCount();
console.log(chalk.bold('\nTasks:'));
console.log(` Total: ${taskCounts.total}`);
console.log(` Pending: ${taskCounts.pending}`);
console.log(` Running: ${taskCounts.running}`);
console.log(` Completed: ${taskCounts.completed}`);
console.log(` Failed: ${taskCounts.failed}`);
const statusColor = loopStatus === 'running' ? chalk.green : loopStatus === 'paused' ? chalk.yellow : chalk.gray;
console.log(chalk.bold('\nRalph Loop:'));
console.log(` Status: ${statusColor(loopStatus)}`);
console.log('');
});
@@ -572,56 +788,382 @@ program
console.log('');
});
// ============ Web / daemon / service Commands ============
/** Shared option set for the commands that can launch a web server. */
function addWebLaunchOptions(cmd: Command): Command {
return cmd
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
)
.option(
'--multiuser',
'Enable opt-in multi-user mode (named users in ~/.codeman/users.json; env: CODEMAN_MULTIUSER)'
);
}
/** Normalize commander's strings into the shape daemon-control/service-installer take. */
function toWebLaunchOptions(options: {
host: string;
port: string;
https?: boolean;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}): WebLaunchOptions {
const port = parseInt(options.port, 10);
if (!Number.isInteger(port) || port <= 0 || port > 65535) {
console.error(chalk.red(`✗ Invalid port: ${options.port}`));
process.exit(1);
}
return {
host: options.host,
port,
https: !!options.https,
titleHostname: options.titleHostname,
allowUnauthenticatedNetwork: !!options.allowUnauthenticatedNetwork,
multiuser: !!options.multiuser,
};
}
/**
* The server prints this itself, but into a log file nobody reads when it is
* detached or supervised. Repeat it where the operator is actually looking.
*/
function warnIfUnauthenticatedNetwork(launch: WebLaunchOptions): void {
if (isLoopbackBindHost(launch.host)) return;
if (isUnauthenticatedNetworkAcknowledged(launch.allowUnauthenticatedNetwork)) return;
console.log(
chalk.yellow(
`⚠ Binding ${launch.host} without CODEMAN_PASSWORD: anyone who can reach this port gets terminal control.`
)
);
console.log(chalk.yellow(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
}
// Web interface command
program
.command('web')
.description('Start the web interface')
.option('-H, --host <host>', 'Host to bind to', process.env.CODEMAN_HOST || '127.0.0.1')
.option('-p, --port <port>', 'Port to listen on (env: CODEMAN_PORT)', process.env.CODEMAN_PORT || '3000')
.option('--https', 'Enable HTTPS with self-signed certificate (only needed for remote access, not localhost)')
.option('--title-hostname <hostname>', 'Override the hostname shown in the browser title')
.option(
'--allow-unauthenticated-network',
'Allow non-loopback web access without CODEMAN_PASSWORD (dangerous; terminal control is exposed)'
)
.action(async (options) => {
const { startWebServer } = await import('./web/server.js');
const host = options.host;
const port = parseInt(options.port, 10);
const https = !!options.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = !!options.allowUnauthenticatedNetwork;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
const webCmd = addWebLaunchOptions(program.command('web').description('Start the web interface'))
.option('-d, --daemon', 'Run detached in the background; survives the shell, logs to <data dir>/web.log')
.option('--stop', 'Stop a server started with --daemon')
.option('--status', 'Report whether a detached server is running');
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
webCmd.action(async (options) => {
// The flag is surfaced to the rest of the process via the env var so
// isMultiUserMode() has a single source of truth (see config/multiuser.ts).
if (options.multiuser) process.env.CODEMAN_MULTIUSER = '1';
const launch = toWebLaunchOptions(options);
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
if (options.stop) {
const result = await stopDaemon(launch);
if (result.ok && result.reason === 'not-running') {
console.log(chalk.gray(`○ ${result.message}`));
return;
}
if (result.ok) {
console.log(chalk.green(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(chalk.gray(' Your agents keep running in tmux.'));
return;
}
console.error(chalk.red(`✗ ${result.message ?? 'Could not stop the server'}`));
process.exit(1);
}
if (options.status) {
const status = await daemonStatus(launch);
if (status.responding) {
const version = status.version ? ` (v${status.version})` : '';
console.log(chalk.green(`✓ Responding at ${status.url}${version}`));
} else {
console.log(chalk.yellow(`○ Nothing answering at ${status.url}`));
}
console.log(` Daemon pid: ${status.running ? chalk.green(String(status.pid)) : chalk.gray('not running')}`);
console.log(chalk.gray(` Pidfile: ${status.pidFile}`));
console.log(chalk.gray(` Log: ${status.logPath}`));
if (!status.running && status.responding) {
console.log(chalk.gray(' (running, but not started with --daemon: probably a service or a foreground run)'));
}
return;
}
if (options.daemon) {
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Starting Codeman in the background...'));
const result = await startDaemon(launch);
if (result.ok) {
console.log(chalk.green(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(chalk.gray(` Logs: ${result.logPath}`));
console.log(chalk.gray(' Stop it with: codeman web --stop'));
console.log(chalk.gray(' Want it back after a reboot? codeman service install'));
return;
}
console.error(chalk.red(`\n✗ ${result.message ?? 'Failed to start'}`));
process.exit(1);
}
const { startWebServer } = await import('./web/server.js');
const host = launch.host;
const port = launch.port;
const https = launch.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = launch.allowUnauthenticatedNetwork ?? false;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
try {
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
process.exit(1);
}
});
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
// Supervised service: the "still there after a reboot" answer, where `web -d` is
// the "still there after I close this shell" one (issue #231).
const serviceCmd = program
.command('service')
.description('Manage the background service (systemd user unit on Linux, LaunchAgent on macOS)');
addWebLaunchOptions(
serviceCmd.command('install').description('Install and start the service, then verify it answers')
).action(async (options) => {
const launch = toWebLaunchOptions(options);
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Installing the Codeman service...'));
const result = await installService(launch);
for (const warning of result.warnings ?? []) console.log(chalk.yellow(`⚠ ${warning}`));
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(chalk.gray(` Unit: ${result.unitPath}`));
if (process.env.CODEMAN_PASSWORD) {
console.log(
chalk.yellow(
' Note: CODEMAN_PASSWORD was NOT copied into the unit file. Add it there yourself if the service needs auth.'
)
);
}
});
serviceCmd
.command('uninstall')
.description('Stop the service and remove its unit file')
.action(() => {
const result = uninstallService();
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
});
addWebLaunchOptions(
serviceCmd.command('status').description('Show whether the service is installed and running')
).action(async (options) => {
const status = await serviceStatus(toWebLaunchOptions(options));
if (!status.kind) {
console.log(chalk.yellow(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
return;
}
console.log(` Supervisor: ${status.kind} (${status.name})`);
console.log(` Unit file: ${status.installed ? chalk.green(status.unitPath) : chalk.gray('not installed')}`);
console.log(` Loaded: ${status.loaded ? chalk.green('yes') : chalk.gray('no')}`);
const version = status.version ? ` (v${status.version})` : '';
console.log(
` Responding: ${status.responding ? chalk.green(`yes at ${status.url}${version}`) : chalk.gray(`no at ${status.url}`)}`
);
});
// ============ Multi-user Commands ============
//
// Operate directly on ~/.codeman/users.json (via user-store) with NO running
// server, honoring CODEMAN_INSTANCE. This is the headless bootstrap path and the
// recovery answer to "locked out: last admin forgot password".
/** Read a password from stdin without echoing. Falls back to plain read on non-TTY. */
function promptHiddenPassword(question: string): Promise<string> {
const stdin = process.stdin;
if (!stdin.isTTY || typeof stdin.setRawMode !== 'function') {
// Non-interactive: read a single line from stdin.
return new Promise((resolve) => {
let buf = '';
stdin.setEncoding('utf8');
stdin.on('data', (d) => (buf += d));
stdin.on('end', () => resolve(buf.replace(/\r?\n$/, '')));
});
}
return new Promise((resolve) => {
process.stdout.write(question);
let input = '';
stdin.setRawMode(true);
stdin.resume();
stdin.setEncoding('utf8');
const onData = (chunk: string) => {
for (const c of chunk) {
if (c === '\n' || c === '\r' || c === '\u0004') {
stdin.setRawMode!(false);
stdin.pause();
stdin.removeListener('data', onData);
process.stdout.write('\n');
resolve(input);
return;
} else if (c === '\u0003') {
process.stdout.write('\n');
process.exit(1);
} else if (c === '\u007f' || c === '\b') {
input = input.slice(0, -1);
} else {
input += c;
}
process.exit(0);
};
process.on('SIGTERM', () => shutdown('SIGTERM'));
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
}
};
stdin.on('data', onData);
});
}
function readAllStdin(): Promise<string> {
return new Promise((resolve) => {
let buf = '';
process.stdin.setEncoding('utf8');
process.stdin.on('data', (d) => (buf += d));
process.stdin.on('end', () => resolve(buf.replace(/\r?\n$/, '')));
});
}
const usersCmd = program.command('users').description('Manage multi-user accounts (~/.codeman/users.json)');
usersCmd
.command('add <name>')
.description('Create a user (prompts for password; use --password-stdin for scripts)')
.option('--admin', 'Create as an admin')
.option('--password-stdin', 'Read the password from stdin instead of prompting')
.action(async (name, options) => {
const { createUser, isValidUsername } = await import('./user-store.js');
if (!isValidUsername(name)) {
console.error(chalk.red('✗ Username must be lowercase, start alphanumeric, 2-32 chars ([a-z0-9_-])'));
process.exit(1);
}
try {
let password: string;
if (options.passwordStdin) {
password = await readAllStdin();
} else {
password = await promptHiddenPassword('New password: ');
const confirm = await promptHiddenPassword('Confirm password: ');
if (password !== confirm) {
console.error(chalk.red('✗ Passwords do not match'));
process.exit(1);
}
}
if (!password || password.length < 8) {
console.error(chalk.red('✗ Password must be at least 8 characters'));
process.exit(1);
}
const user = await createUser({ username: name, role: options.admin ? 'admin' : 'user', password });
console.log(chalk.green(`✓ Created ${user.role} "${user.username}"`));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
usersCmd
.command('passwd <name>')
.description('Reset a user password')
.option('--password-stdin', 'Read the new password from stdin instead of prompting')
.action(async (name, options) => {
const { setPassword } = await import('./user-store.js');
try {
let password: string;
if (options.passwordStdin) {
password = await readAllStdin();
} else {
password = await promptHiddenPassword('New password: ');
const confirm = await promptHiddenPassword('Confirm password: ');
if (password !== confirm) {
console.error(chalk.red('✗ Passwords do not match'));
process.exit(1);
}
}
await setPassword(name, password, { mustChangePassword: false });
console.log(chalk.green(`✓ Password updated for "${name}"`));
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
usersCmd
.command('list')
.alias('ls')
.description('List all users')
.action(async () => {
const { readUsers } = await import('./user-store.js');
const users = await readUsers(true);
if (users.length === 0) {
console.log(chalk.yellow('No users defined (run: codeman users add <name> --admin)'));
return;
}
console.log(chalk.bold('\nUsers:'));
for (const u of users) {
const role = u.role === 'admin' ? chalk.magenta('admin') : chalk.cyan('user ');
const state = u.disabled ? chalk.red('disabled') : chalk.green('enabled ');
const flags = [u.mustChangePassword ? 'must-change-pw' : '', u.canBypassPermissions ? 'can-bypass' : '']
.filter(Boolean)
.join(' ');
console.log(` ${role} ${state} ${u.username}${flags ? chalk.gray(` [${flags}]`) : ''}`);
}
console.log('');
});
usersCmd
.command('rm <name>')
.description('Delete a user')
.option('--delete-space', "Also delete the user's ~/codeman-users/<name> space")
.action(async (name, options) => {
const { deleteUser, deleteUserSpace } = await import('./user-store.js');
try {
await deleteUser(name);
if (options.deleteSpace) {
await deleteUserSpace(name);
console.log(chalk.green(`✓ Deleted user "${name}" and their space`));
} else {
console.log(chalk.green(`✓ Deleted user "${name}" (space left on disk)`));
}
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
+112
View File
@@ -0,0 +1,112 @@
/**
* @fileoverview Bounds for the agent wait primitives.
*
* These back the blocking endpoints an agent uses to orchestrate other sessions
* (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the
* `wait` field on `POST /api/sessions/:id/input`). Plan: `docs/agent-control-plan.md`.
*
* Why every value is bounded:
* - An unbounded long-poll is a socket leak. A caller that asks for a 12-hour wait
* and walks away holds a connection (and a waiter, and a timer) until the process
* restarts, so `MAX_WAIT_MS` is a hard ceiling applied server-side.
* - `DEFAULT_WAIT_MS` is deliberately short (60s). Production is reached through
* `tailscale serve` and users also run cloudflared tunnels; both can cut an idle
* connection, so the documented pattern is a client-side loop over short waits
* rather than one very long call. Fastify itself is happy to hold the request
* (`requestTimeout` defaults to 0, and `keepAliveTimeout` applies between
* requests, not to an in-flight one), the intermediaries are the constraint.
* - The waiter caps mirror `MAX_SSE_CLIENTS` in `map-limits.ts`: each pending
* waiter costs an open HTTP response plus a timer, so the pool is capped rather
* than queued. Exceeding a cap is an explicit error, never a silent wait.
* - There are THREE caps, not two, because a process-wide pool with no per-user
* dimension lets one user deny the primitive to everyone else. `middleware/auth.ts`
* already treats that shape as a bug (its `userFailures` bucket exists so "one user
* behind a NAT can't lock out everyone else"); `MAX_WAITERS_PER_OWNER` is the same
* idea for waiters. It applies only when the caller has an owner, so single-user
* mode is byte-identical to having no owner cap at all.
*
* All values are env-overridable and clamped to sane hard bounds, so a typo in an
* env var degrades to the default instead of disabling the protection.
*
* @module config/agent-wait
*/
/** Absolute floor for any wait, in ms. Sub-second waits are polling, not waiting. */
export const MIN_WAIT_MS = 1_000;
/** Ceiling the operator-configurable maximum is itself clamped to. */
const HARD_MAX_WAIT_MS = 3_600_000;
function envInt(name: string, fallback: number, min: number, max: number): number {
const raw = parseInt(process.env[name] || '', 10);
if (!Number.isFinite(raw) || raw <= 0) return fallback;
return Math.max(min, Math.min(max, raw));
}
/** Longest a single wait may block. Requests above this are clamped down, not rejected. */
export const MAX_WAIT_MS = envInt('CODEMAN_WAIT_MAX_MS', 600_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS);
/** Used when the caller omits `timeout`. Never exceeds MAX_WAIT_MS. */
export const DEFAULT_WAIT_MS = Math.min(
envInt('CODEMAN_WAIT_DEFAULT_MS', 60_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS),
MAX_WAIT_MS
);
/** Concurrent waiters (signal + output) allowed against one session. */
export const MAX_WAITERS_PER_SESSION = envInt('CODEMAN_WAIT_MAX_PER_SESSION', 16, 1, 256);
/**
* Ceiling the operator-configurable total is itself clamped to.
*
* 512 rather than the 4096 this started at. Every other knob in this file degrades
* safely on a bad value; a 4096 ceiling instead lets a well-meaning operator turn the
* protection into the problem, since 4096 concurrent held responses (each an open
* socket, a timer and a pending promise) exceeds the 1024 soft `RLIMIT_NOFILE` that is
* still the default on most Linux distros, before counting PTYs, SSE clients and
* WebSockets. 512 is ~5x `MAX_SSE_CLIENTS` (100, the pool this one is modelled on), so
* the knob stays useful for a busy orchestration host while the whole server still fits
* inside a default fd budget with room to spare.
*/
const HARD_MAX_WAITERS_TOTAL = 512;
/** Concurrent waiters allowed across every session in the process. */
export const MAX_WAITERS_TOTAL = envInt('CODEMAN_WAIT_MAX_TOTAL', 128, 1, HARD_MAX_WAITERS_TOTAL);
/**
* Concurrent waiters allowed for one owner (multi-user mode's `Session.owner`).
*
* Sits between the per-session cap (16) and the process-wide one (128): high enough
* that one user orchestrating several workers at once never trips it, low enough that
* a single user cannot occupy the whole pool and deny the primitive to everyone else,
* admin included. Ignored entirely when the caller has no owner, which is every
* request in single-user mode.
*/
export const MAX_WAITERS_PER_OWNER = envInt('CODEMAN_WAIT_MAX_PER_OWNER', 48, 1, HARD_MAX_WAITERS_TOTAL);
/** Bounds on the literal `match` string accepted by wait-output. */
export const MIN_MATCH_LENGTH = 1;
export const MAX_MATCH_LENGTH = 200;
/**
* Tail of the terminal buffer scanned by `wait-output?from=buffer`.
*
* The buffer itself runs to 32MB. Scanning all of it would be an ANSI strip over
* 32MB (a full second copy) on a request an agent may issue in a loop, and the
* question `from=buffer` answers is "did this appear recently", not "ever". The
* tail is continuous with the live stream, since `_terminalBuffer.append(data)`
* and `emit('terminal', data)` receive the same bytes.
*/
export const MAX_BUFFER_SCAN_BYTES = envInt('CODEMAN_WAIT_BUFFER_SCAN_BYTES', 256 * 1024, 4 * 1024, 8 * 1024 * 1024);
/** Characters of surrounding output returned either side of a wait-output match. */
export const MAX_SNIPPET_CONTEXT = 80;
/**
* Clamp a caller-supplied timeout into [MIN_WAIT_MS, MAX_WAIT_MS].
* Absent / non-numeric / non-finite input falls back to DEFAULT_WAIT_MS.
*/
export function clampWaitMs(value: unknown): number {
const n = typeof value === 'string' ? Number(value) : value;
if (typeof n !== 'number' || !Number.isFinite(n)) return DEFAULT_WAIT_MS;
return Math.max(MIN_WAIT_MS, Math.min(MAX_WAIT_MS, Math.trunc(n)));
}
+7 -2
View File
@@ -31,5 +31,10 @@ export const AUTH_FAILURE_WINDOW_MS = 15 * 60 * 1000;
// Hooks
// ============================================================================
/** Timeout for Claude Code hook curl commands (ms) */
export const HOOK_TIMEOUT_MS = 10000;
/**
* Timeout for Claude Code hook curl commands, in SECONDS: the hook `timeout`
* field is seconds (the CLI multiplies by 1000). The predecessor constant
* `HOOK_TIMEOUT_MS = 10000` fed the same field, so those hooks effectively had a
* ~2.8-hour timeout; 10 seconds is the originally intended budget.
*/
export const HOOK_TIMEOUT_SECONDS = 10;
+8
View File
@@ -98,6 +98,14 @@ export const DEPENDENCY_REGISTRY: ToolDependency[] = [
usedBy: ['Gemini sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['gemini'], versionArg: '--version' } }],
},
{
id: 'antigravity',
label: 'Antigravity CLI',
category: 'core',
required: false,
usedBy: ['Antigravity sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['agy'], versionArg: '--version' } }],
},
{
id: 'libreoffice',
label: 'LibreOffice',
+152
View File
@@ -0,0 +1,152 @@
/**
* @fileoverview File Viewer edit-mode policy (issue #212).
*
* Pure, IO-free policy for which workspace files the in-viewer editor may read
* for editing and write back. Consumed by the `edit=1` branch of
* `GET /api/sessions/:id/file-content` and by `PUT /api/sessions/:id/file-content`
* in `src/web/routes/file-routes.ts`.
*
* Design (docs/file-viewer-edit-plan.md):
* - ALLOWLIST of text extensions/basenames, not a blocklist — matching the
* attachment-guard precedent. `svg` and `env` are deliberately absent: svg is
* treated as untrusted on the read side, and `.env` is sensitive-path blocked
* anyway; excluding them here keeps a single obvious refusal.
* - The `.git/` subtree is denied outright: `.git/hooks/*` is code execution and
* a corrupted index looks unrecoverable to a user who wanted to fix a typo.
* - EOL helpers exist because a browser <textarea> normalizes to LF; the server
* re-applies the file's original ending so a two-line edit of a CRLF file does
* not become a whole-file diff. Mixed-EOL files normalize to the dominant
* style (documented lossy edge).
*/
/** Hard cap for edit-mode reads AND writes (bytes of file content). */
export const MAX_EDITABLE_BYTES = 512 * 1024;
/** Lowercase extensions (no dot) the editor will open and save. */
export const EDITABLE_EXTENSIONS: ReadonlySet<string> = new Set([
// JS/TS ecosystem
'ts',
'tsx',
'js',
'jsx',
'mjs',
'cjs',
'json',
'jsonc',
// Docs / plain text
'md',
'mdx',
'txt',
'rst',
'adoc',
// Web
'css',
'scss',
'less',
'html',
'htm',
'xml',
// Config
'yml',
'yaml',
'toml',
'ini',
'cfg',
'conf',
'properties',
// Shell
'sh',
'bash',
'zsh',
'fish',
// Languages
'py',
'rb',
'go',
'rs',
'java',
'kt',
'swift',
'c',
'h',
'cpp',
'hpp',
'cc',
'cs',
'php',
'sql',
'graphql',
'proto',
'lua',
'pl',
'r',
'jl',
'tf',
'gradle',
// Data / misc text
'csv',
'tsv',
'log',
'diff',
'patch',
]);
/** Extensionless (or dot-led) file names that are still editable text. */
export const EDITABLE_BASENAMES: ReadonlySet<string> = new Set([
'dockerfile',
'makefile',
'license',
'readme',
'changelog',
'authors',
'codeowners',
'procfile',
'.gitignore',
'.gitattributes',
'.dockerignore',
'.prettierignore',
'.prettierrc',
'.editorconfig',
'.nvmrc',
'.npmrc',
'.eslintignore',
]);
/** Whether a file name (basename only) is eligible for in-viewer editing. */
export function isEditableFileName(fileName: string): boolean {
const lower = fileName.toLowerCase();
if (EDITABLE_BASENAMES.has(lower)) return true;
const dot = lower.lastIndexOf('.');
// No extension (or a bare dotfile like `.bashrc`): only the basename list applies.
if (dot <= 0) return false;
return EDITABLE_EXTENSIONS.has(lower.slice(dot + 1));
}
/**
* Whether a workspace-relative path is denied for editing regardless of its
* extension. Currently: anything inside a `.git` directory at any depth.
*/
export function isDeniedEditRelativePath(relativePath: string): boolean {
return relativePath.split('/').some((segment) => segment === '.git');
}
export type FileEol = 'lf' | 'crlf';
/** Dominant line-ending style of a text buffer (LF when tied or single-line). */
export function detectEol(text: string): FileEol {
let crlf = 0;
let lf = 0;
for (let i = 0; i < text.length; i++) {
if (text.charCodeAt(i) === 10) {
if (i > 0 && text.charCodeAt(i - 1) === 13) crlf++;
else lf++;
}
}
return crlf > lf ? 'crlf' : 'lf';
}
/** Normalize every line ending in `text` to the requested style. */
export function applyEol(text: string, eol: FileEol): string {
const normalized = text.replace(/\r\n/g, '\n');
return eol === 'crlf' ? normalized.replace(/\n/g, '\r\n') : normalized;
}
+63
View File
@@ -0,0 +1,63 @@
/**
* @fileoverview Multi-user mode gating + limits (opt-in, off by default).
*
* Multi-user mode is enabled by `codeman web --multiuser` (which sets
* `CODEMAN_MULTIUSER=1`) or the env var directly. When OFF, behavior is
* byte-identical to today: `users.json` is never read and all ownership scoping
* is bypassed. Everything here is per-instance like the rest of Codeman: a beta
* instance (`CODEMAN_INSTANCE=beta`) has its own `users.json` via `dataPath()`,
* and its user spaces live under the same shared `~/codeman-users` as prod (like
* `~/codeman-cases`), unless `CODEMAN_USER_SPACES_DIR` overrides it.
*
* See `docs/multi-user-plan.md` sections 3, 4.2, and 11.
*/
import { homedir } from 'node:os';
import { join } from 'node:path';
import { MAX_CONCURRENT_SESSIONS } from './map-limits.js';
/**
* Whether multi-user mode is active. Read from the environment each call so it is
* stable for the process lifetime (env does not change after boot) and trivially
* overridable in tests. Accepts `1` or `true`.
*/
export function isMultiUserMode(): boolean {
const v = process.env.CODEMAN_MULTIUSER;
return v === '1' || v === 'true';
}
/**
* Root of per-user spaces: `~/codeman-users` (sibling of `~/codeman-cases`).
* Overridable via `CODEMAN_USER_SPACES_DIR` (used by tests). Resolved lazily so a
* test can point it at a temp dir before the first call.
*/
export function getUserSpacesDir(): string {
return process.env.CODEMAN_USER_SPACES_DIR || join(homedir(), 'codeman-users');
}
/** Absolute path to a user's top-level space: `<USER_SPACES_DIR>/<username>[/segments]`. */
export function userSpacePath(username: string, ...segments: string[]): string {
return join(getUserSpacesDir(), username, ...segments);
}
/** Absolute path to a user's cases dir: `<USER_SPACES_DIR>/<username>/cases`. */
export function userCasesDir(username: string): string {
return join(getUserSpacesDir(), username, 'cases');
}
/** Maximum number of user accounts (default 25, env `CODEMAN_MAX_USERS`). */
export function maxUsers(): number {
const n = Number(process.env.CODEMAN_MAX_USERS);
return Number.isInteger(n) && n > 0 ? n : 25;
}
/**
* Per-user concurrent-session cap (the fairness lever). Defaults to half the
* global cap; overridable via `CODEMAN_MAX_SESSIONS_PER_USER`. The global cap
* (MAX_CONCURRENT_SESSIONS) still applies on top and is shared across users.
*/
export function maxSessionsPerUser(): number {
const n = Number(process.env.CODEMAN_MAX_SESSIONS_PER_USER);
if (Number.isInteger(n) && n > 0) return n;
return Math.max(1, Math.floor(MAX_CONCURRENT_SESSIONS / 2));
}
+32
View File
@@ -0,0 +1,32 @@
/**
* @fileoverview Supervisor identity (systemd unit name / launchd job label).
*
* Three things now write or look for the same supervisor job: `install.sh`, the
* in-app self-updater (`web/self-update.ts` detects it to decide how to restart),
* and `codeman service install`. The names live here so they cannot drift apart,
* because a mismatch is silent in the worst way: `service install` would happily
* create a SECOND job alongside the installer's, and two servers sharing one data
* dir and one tmux socket attach PTYs to each other's live sessions
* (see config/instance.ts).
*
* The names are instance-scoped for exactly that reason: a `CODEMAN_INSTANCE=beta`
* build writing `com.codeman.web` would overwrite the production LaunchAgent. The
* DEFAULT instance keeps the historical names byte-identical, so existing installs
* and every unit install.sh has already written are unaffected.
*
* @module config/service-names
*/
import { CODEMAN_INSTANCE } from './instance.js';
/**
* Instance name reduced to characters that are safe in a filename and in a
* launchd label. `CODEMAN_INSTANCE` is arbitrary operator input.
*/
const SAFE_INSTANCE = CODEMAN_INSTANCE.replace(/[^A-Za-z0-9_-]/g, '').slice(0, 32);
/** systemd user unit: `codeman-web.service`, or `codeman-web-beta.service` for a beta. */
export const SYSTEMD_UNIT = `codeman-web${SAFE_INSTANCE ? `-${SAFE_INSTANCE}` : ''}.service`;
/** launchd job label: `com.codeman.web`, or `com.codeman.beta.web` for a beta. */
export const LAUNCHD_LABEL = SAFE_INSTANCE ? `com.codeman.${SAFE_INSTANCE}.web` : 'com.codeman.web';
+66
View File
@@ -0,0 +1,66 @@
/**
* Limits and timeouts for web tabs (dashboards embedded as Codeman tabs).
*
* Every value here bounds something an untrusted-ish upstream controls: how many
* dashboards can be saved, how long the server will wait on one, how much of a
* response it will buffer before rewriting HTML, and how many sockets a single
* dashboard may hold open. Env-overridable in the same style as the other config
* modules.
*/
function envInt(name: string, fallback: number): number {
const parsed = parseInt(process.env[name] || '', 10);
return Number.isFinite(parsed) && parsed > 0 ? parsed : fallback;
}
/** Max saved webviews (per owner in multi-user mode). */
export const MAX_WEBVIEWS = envInt('CODEMAN_MAX_WEBVIEWS', 50);
/**
* Max iframes kept mounted at once. Switching tabs must not reload a dashboard,
* so frames stay alive while hidden; past this many, the least-recently-viewed
* frame is evicted. Consumed by the frontend via `GET /api/webviews`.
*/
export const MAX_LIVE_WEBVIEW_FRAMES = envInt('CODEMAN_MAX_LIVE_WEBVIEW_FRAMES', 6);
/** How long a minted proxy capability stays valid (rolling, refreshed on use). */
export const WEBVIEW_CAPABILITY_TTL_MS = envInt('CODEMAN_WEBVIEW_CAPABILITY_TTL_MS', 12 * 60 * 60 * 1000);
/** Max concurrent capabilities held in memory before the oldest are dropped. */
export const MAX_WEBVIEW_CAPABILITIES = 200;
/**
* How long a proxied HTTP request waits for the upstream's RESPONSE HEADERS.
*
* This bounds time-to-headers only, never an actively streaming body: the proxy
* clears the timer the moment headers arrive (issue #237: the old 30s
* `AbortSignal.timeout` bounded the whole fetch and killed slow AI/model endpoints
* and long streams alike, as a silent 502). 300s because "the app is thinking" is
* normal for the dashboards people proxy; abandoned upstreams are reclaimed by the
* client-hangup abort, not by this value, so a generous default costs nothing.
*/
export const WEBVIEW_UPSTREAM_TIMEOUT_MS = envInt('CODEMAN_WEBVIEW_TIMEOUT_MS', 300_000);
/** Shorter timeout for the editor's "Test" probe, which a human is waiting on. */
export const WEBVIEW_PROBE_TIMEOUT_MS = envInt('CODEMAN_WEBVIEW_PROBE_TIMEOUT_MS', 8_000);
/**
* WebSocket upgrade handshake timeout. Deliberately decoupled from
* WEBVIEW_UPSTREAM_TIMEOUT_MS: a handshake is connection establishment, and waiting
* minutes on one only delays the browser's reconnect logic. Matches the pre-#237
* behavior (the handshake used to ride the 30s upstream timeout).
*/
export const WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS = envInt('CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS', 30_000);
/**
* Max bytes of an HTML response buffered for `<base>` injection and link
* rewriting. Larger HTML documents stream through untouched: the rewrite is a
* convenience, and buffering an unbounded upstream body is a memory hazard.
*/
export const MAX_WEBVIEW_HTML_REWRITE_BYTES = envInt('CODEMAN_MAX_WEBVIEW_HTML_BYTES', 8 * 1024 * 1024);
/** Max concurrent proxied WebSockets per webview (mirrors MAX_WS_PER_SESSION). */
export const MAX_WEBVIEW_SOCKETS = envInt('CODEMAN_MAX_WEBVIEW_SOCKETS', 8);
/** URL path prefix the proxy is mounted at. Single source of truth. */
export const WEBVIEW_PROXY_PREFIX = '/webview';
+35 -4
View File
@@ -15,6 +15,8 @@ import { SseEvent } from '../web/sse-events.js';
import { CronJobSchema } from '../web/schemas.js';
import { getErrorMessage, createErrorResponse, ApiErrorCode } from '../types/api.js';
import { MAX_CONCURRENT_SESSIONS, MAX_CRON_JOBS, MAX_CRON_RUN_HISTORY } from '../config/map-limits.js';
import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../user-store.js';
import { sessionCapacityState, isWorkingDirAllowedForUsername } from '../web/route-helpers.js';
import { CRON_READY_MAX_ATTEMPTS, CRON_READY_SETTLE_MS } from '../config/server-timing.js';
import {
DEFAULT_BLOCKED_TREES,
@@ -25,6 +27,7 @@ import { validateSessionFilePath } from '../web/route-helpers.js';
import { computeNextRunAt, dueKeyFor } from './cron-time.js';
import type { SessionPort, EventPort, ConfigPort, InfraPort } from '../web/ports/index.js';
import type { CronJob, CronJobRun, CronJobRunStatus, TriggerType } from '../types/cron.js';
import type { GeminiConfig } from '../types/session.js';
import type { CronJobInput } from './cron-input.js';
/** The subset of the route context the cron depends on. */
@@ -108,7 +111,7 @@ export class CronService {
// ──────────────────────────── Mutations ───────────────────────────
createJob(input: CronJobInput): CronJob {
createJob(input: CronJobInput, owner?: string): CronJob {
if (Object.keys(this.store.getCronJobs()).length >= MAX_CRON_JOBS) {
throw this.badRequest(`Maximum number of cron jobs (${MAX_CRON_JOBS}) reached`);
}
@@ -117,6 +120,7 @@ export class CronService {
const job: CronJob = {
id: uuidv4(),
name: input.name,
owner,
agentType: input.agentType,
workingDir: input.workingDir,
launchCommand: input.launchCommand,
@@ -328,6 +332,12 @@ export class CronService {
return this.failRun(job, run, 'workingDir does not exist');
}
// Section 6.3: defense-in-depth workingDir confinement re-check at FIRE time against the
// owner's CURRENT space (complements the create/update gate). No-op in single-user / unset owner.
if (!(await isWorkingDirAllowedForUsername(job.owner, job.workingDir))) {
return this.failRun(job, run, 'workingDir is outside the owner workspace');
}
// Recurring jobs: close the still-open session created by this job's
// previous run before launching the next (default ON, opt-out via
// autoClosePreviousSession:false) — otherwise an unattended interval/daily
@@ -336,10 +346,21 @@ export class CronService {
await this.closePreviousRunSessions(job, run.id);
}
// Respect the global session cap.
if (this.deps.sessions.size >= MAX_CONCURRENT_SESSIONS) {
// Respect the global cap AND the owner's per-user cap (multi-user).
const cap = sessionCapacityState(this.deps.sessions, job.owner);
if (cap.atGlobalCap) {
return this.failRun(job, run, `Maximum concurrent sessions (${MAX_CONCURRENT_SESSIONS}) reached`);
}
if (cap.atUserCap) {
return this.failRun(job, run, `Owner's per-user session limit reached`);
}
// Section 6.3: re-resolve the owner's grant at FIRE time (it may have been revoked
// since create). Gates shell/launchCommand AND clamps the external-CLI bypass below.
const ownerGranted = await canUsernameRunPrivilegedCommands(job.owner);
if ((job.agentType === 'shell' || job.launchCommand) && !ownerGranted) {
return this.failRun(job, run, 'Owner lacks the can-bypass-permissions grant for shell/launchCommand jobs');
}
// Create + start the session (mirrors the quick-start route flow).
let session: Session;
@@ -348,7 +369,15 @@ export class CronService {
const globalNice = await this.deps.getGlobalNiceConfig();
const modelConfig = await this.deps.getModelConfig();
const claudeModeConfig = await this.deps.getClaudeModeConfig();
const effectiveClaudeMode = await resolveClaudeModeForUsername(claudeModeConfig.claudeMode, job.owner);
const model = mode !== 'shell' ? modelConfig?.defaultModel || undefined : undefined;
// Section 6.3: cron carries no per-CLI config, so buildGeminiCommand(undefined)
// would default a non-granted owner to `--approval-mode yolo` (classifier-free) —
// materialize auto_edit for a non-granted gemini owner, mirroring the route clamp
// (#15). Granted/admin/single-user leave it undefined → yolo parity. Codex's absent
// config already defaults to the safe sandbox, so no clamp is needed there.
const geminiConfig: GeminiConfig | undefined =
mode === 'gemini' && !ownerGranted ? { approvalMode: 'auto_edit' } : undefined;
session = new Session({
workingDir: job.workingDir,
mode,
@@ -357,8 +386,10 @@ export class CronService {
useMux: true,
niceConfig: globalNice,
model,
claudeMode: claudeModeConfig.claudeMode,
claudeMode: effectiveClaudeMode,
allowedTools: claudeModeConfig.allowedTools,
geminiConfig,
owner: job.owner,
});
this.deps.addSession(session);
this.store.incrementSessionsCreated();
+496
View File
@@ -0,0 +1,496 @@
/**
* @fileoverview Detached `codeman web` control: start (-d), stop, status.
*
* Backs `codeman web -d`, `codeman web --stop` and `codeman web --status`. The
* server itself is unchanged; this module re-launches the SAME entry script in a
* new session (`detached: true` calls setsid), so the child has no controlling
* terminal and no shell job entry. That is what actually makes it outlive the
* shell: `nohup` does not, because Node re-arms SIGHUP to its default disposition
* even when it inherits "ignore", and `cli.ts` installs a SIGHUP handler that
* shuts the server down gracefully (issue #231).
*
* Two rules shape the rest of the module:
*
* 1. **Never start a second server on one data dir.** `~/.codeman` and the
* `tmux -L codeman` socket are process-wide (config/instance.ts), so a second
* instance discovers and attaches PTYs to the first one's live sessions and
* starts resizing them. A double `-d` therefore has to be a hard error, which
* means checking both the pidfile AND the port before spawning.
* 2. **Never report success we have not seen.** The parent polls `/api/status`
* until the child answers (or dies) before printing a URL. A port clash or a
* missing dependency otherwise looks exactly like a clean start.
*
* Pure helpers (arg building, URL building, pidfile parsing, the process-identity
* check) are exported separately so they can be unit-tested without spawning.
*
* @module daemon-control
*/
import { spawn, execFileSync } from 'node:child_process';
import { appendFileSync, closeSync, existsSync, openSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import http from 'node:http';
import https from 'node:https';
import { dataPath } from './config/instance.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** How long to wait for a freshly spawned server to answer `/api/status`. */
const START_TIMEOUT_MS = 30_000;
/** How long to wait for a SIGTERM'd server to actually exit before giving up. */
const STOP_TIMEOUT_MS = 15_000;
/** Poll interval while waiting for either of the above. */
const POLL_INTERVAL_MS = 250;
/** The `web` command's options, as far as a detached relaunch cares about them. */
export interface WebLaunchOptions {
host: string;
port: number;
https: boolean;
titleHostname?: string;
allowUnauthenticatedNetwork?: boolean;
multiuser?: boolean;
}
export interface StartResult {
ok: boolean;
pid?: number;
url?: string;
/** Machine-readable failure cause; `undefined` on success. */
reason?: 'already-running' | 'exited' | 'timeout';
message?: string;
logPath: string;
}
export interface StopResult {
ok: boolean;
pid?: number;
reason?: 'not-running' | 'foreign-pid' | 'timeout' | 'no-pidfile-but-responding';
message?: string;
}
export interface DaemonStatus {
pid: number | null;
/** The pid in the pidfile is alive AND still looks like a Codeman web process. */
running: boolean;
/** Something answered `/api/status` at the expected address. */
responding: boolean;
version?: string;
url: string;
pidFile: string;
logPath: string;
}
// ─────────────────────────────────────────────────────────────────────────────
// Pure helpers
// ─────────────────────────────────────────────────────────────────────────────
/** Rebuild the `web` argv for the child, dropping the daemon flags themselves. */
export function buildWebArgs(options: WebLaunchOptions): string[] {
const args = ['web', '--host', options.host, '--port', String(options.port)];
if (options.https) args.push('--https');
if (options.titleHostname) args.push('--title-hostname', options.titleHostname);
if (options.allowUnauthenticatedNetwork) args.push('--allow-unauthenticated-network');
if (options.multiuser) args.push('--multiuser');
return args;
}
/**
* Connectable address for this bind. A wildcard bind is not itself connectable,
* so `0.0.0.0` / `::` become loopback; a bare IPv6 literal gets bracketed.
*/
export function buildBaseUrl(options: WebLaunchOptions): string {
const protocol = options.https ? 'https' : 'http';
let host = options.host.trim();
if (host === '0.0.0.0' || host === '::' || host === '') host = '127.0.0.1';
if (host.includes(':') && !host.startsWith('[')) host = `[${host}]`;
return `${protocol}://${host}:${options.port}`;
}
/** The endpoint polled for readiness. */
export function buildStatusUrl(options: WebLaunchOptions): string {
return `${buildBaseUrl(options)}/api/status`;
}
/** Parse a pidfile body. Rejects garbage, and pid 1 (init is never ours). */
export function parsePidFileContents(text: string): number | null {
const trimmed = text.trim();
if (!/^\d+$/.test(trimmed)) return null;
const pid = Number.parseInt(trimmed, 10);
if (!Number.isSafeInteger(pid) || pid <= 1) return null;
return pid;
}
/**
* Does this command line look like a Codeman web server?
*
* Pids are recycled, and a stale pidfile pointing at whatever inherited the
* number is a live footgun: `codeman web --stop` must not SIGTERM an unrelated
* process. Both the npm bin (`codeman`/`aicodeman`) and the direct entry
* (`node dist/index.js web`, `tsx src/index.ts web`) have to match.
*/
export function looksLikeCodemanWeb(command: string | null | undefined): boolean {
if (!command) return false;
if (!/(^|\s)web(\s|$)/.test(command)) return false;
return /(^|[/\s])(ai)?codeman(\s|$)/.test(command) || /index\.(js|ts)(\s|$)/.test(command);
}
// ─────────────────────────────────────────────────────────────────────────────
// Paths
// ─────────────────────────────────────────────────────────────────────────────
/**
* Resolved at call time, not module load: tests swap `HOME` per file, and the
* data dir is derived from it (see test/setup.ts).
*/
export function pidFilePath(): string {
return dataPath('web.pid');
}
/** Where a detached server's stdout/stderr is appended. */
export function logFilePath(): string {
return dataPath('web.log');
}
// ─────────────────────────────────────────────────────────────────────────────
// Process probing
// ─────────────────────────────────────────────────────────────────────────────
/** Signal 0 liveness check. EPERM means the pid exists but is not ours. */
export function isProcessAlive(pid: number): boolean {
try {
process.kill(pid, 0);
return true;
} catch (err) {
return (err as NodeJS.ErrnoException).code === 'EPERM';
}
}
/** Full command line of a pid, or null. `-o command=` is portable to macOS. */
export function readProcessCommand(pid: number): string | null {
try {
const out = execFileSync('ps', ['-o', 'command=', '-p', String(pid)], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
});
return out.trim() || null;
} catch {
return null;
}
}
/** Read the pidfile, returning null when it is missing, empty or malformed. */
export function readPidFile(): number | null {
const file = pidFilePath();
if (!existsSync(file)) return null;
try {
return parsePidFileContents(readFileSync(file, 'utf-8'));
} catch {
return null;
}
}
function removePidFile(): void {
try {
unlinkSync(pidFilePath());
} catch {
/* already gone */
}
}
/** Pid of a live Codeman web server recorded in the pidfile, or null. */
export function readLivePid(): number | null {
const pid = readPidFile();
if (pid === null) return null;
if (!isProcessAlive(pid)) return null;
// A recycled pid is not ours. `ps` can also legitimately fail (containers with
// no procps); treat "cannot tell" as ours rather than orphaning the pidfile.
const command = readProcessCommand(pid);
if (command !== null && !looksLikeCodemanWeb(command)) return null;
return pid;
}
// ─────────────────────────────────────────────────────────────────────────────
// HTTP readiness probe
// ─────────────────────────────────────────────────────────────────────────────
export interface ProbeResult {
/** A Codeman server answered. A 401 counts: auth is active, the server is up. */
up: boolean;
version?: string;
}
/**
* Probe `/api/status`. Self-signed certs are accepted (`--https` generates one),
* and 401 counts as up because `CODEMAN_PASSWORD` gates that route. The body is
* checked so an unrelated service squatting on the port is not read as success.
*/
export function probeServer(url: string, timeoutMs = 2000): Promise<ProbeResult> {
return new Promise((resolve) => {
let settled = false;
const done = (result: ProbeResult) => {
if (settled) return;
settled = true;
resolve(result);
};
let target: URL;
try {
target = new URL(url);
} catch {
done({ up: false });
return;
}
const transport = target.protocol === 'https:' ? https : http;
const req = transport.request(
{
protocol: target.protocol,
hostname: target.hostname,
port: target.port,
path: target.pathname,
method: 'GET',
rejectUnauthorized: false,
timeout: timeoutMs,
headers: { Accept: 'application/json' },
},
(res) => {
if (res.statusCode === 401) {
res.resume();
done({ up: true });
return;
}
let body = '';
res.setEncoding('utf-8');
res.on('data', (chunk: string) => {
if (body.length < 4096) body += chunk;
});
res.on('end', () => {
if (!body.includes('"success"')) {
done({ up: false });
return;
}
let version: string | undefined;
try {
version = (JSON.parse(body) as { data?: { version?: string } }).data?.version;
} catch {
/* body was truncated at 4KB; up is still true */
}
done({ up: true, version });
});
res.on('error', () => done({ up: false }));
}
);
req.on('timeout', () => {
req.destroy();
done({ up: false });
});
req.on('error', () => done({ up: false }));
req.end();
});
}
function sleep(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
// ─────────────────────────────────────────────────────────────────────────────
// Start / stop / status
// ─────────────────────────────────────────────────────────────────────────────
/**
* The script to relaunch. `process.execArgv` is carried over with it so a dev
* run under tsx (whose execArgv holds the tsx loader flags) re-launches through
* tsx instead of handing a `.ts` file to bare node.
*/
function entryScript(): string {
const script = process.argv[1];
if (!script) throw new Error('cannot determine the codeman entry script to relaunch');
return script;
}
/** Marks one launch in the append-only log so a tail cannot mix two runs. */
const LOG_SEPARATOR = '=== codeman web start';
/**
* Last few lines of the daemon log, for reporting a failed start. The log is
* append-only across launches, so the tail starts at the last separator when
* there is one: otherwise a crash report is padded with the previous run's
* cheerful startup banner.
*/
export function tailLog(maxLines = 15): string {
try {
const lines = readFileSync(logFilePath(), 'utf-8').trimEnd().split('\n');
const start = lines.map((line) => line.startsWith(LOG_SEPARATOR)).lastIndexOf(true);
const current = start === -1 ? lines : lines.slice(start + 1);
return current.slice(-maxLines).join('\n');
} catch {
return '';
}
}
/**
* Spawn a detached `codeman web` and wait until it answers before returning.
* Refuses when a server is already up on this data dir (see rule 1 in the module
* docblock).
*/
export async function startDaemon(options: WebLaunchOptions): Promise<StartResult> {
const logPath = logFilePath();
const url = buildBaseUrl(options);
const statusUrl = buildStatusUrl(options);
const existingPid = readLivePid();
if (existingPid !== null) {
return {
ok: false,
reason: 'already-running',
pid: existingPid,
logPath,
message: `a Codeman server is already running (pid ${existingPid}). Stop it with \`codeman web --stop\` first.`,
};
}
const alreadyServing = await probeServer(statusUrl, 1500);
if (alreadyServing.up) {
return {
ok: false,
reason: 'already-running',
logPath,
url,
message: `something is already serving ${url}. Two servers on one data dir attach to each other's tmux sessions, so refusing to start.`,
};
}
// A pidfile that survived a crash: the process is gone, so it is just litter.
if (readPidFile() !== null) removePidFile();
const args = buildWebArgs(options);
try {
appendFileSync(logPath, `\n${LOG_SEPARATOR} ${new Date().toISOString()} ===\n`, 'utf-8');
} catch {
/* the spawn below reports a genuinely unwritable log */
}
const logFd = openSync(logPath, 'a');
let child;
try {
child = spawn(process.execPath, [...process.execArgv, entryScript(), ...args], {
detached: true,
stdio: ['ignore', logFd, logFd],
env: process.env,
});
} finally {
closeSync(logFd);
}
let exited = false;
child.on('exit', () => {
exited = true;
});
child.on('error', () => {
exited = true;
});
const pid = child.pid;
if (pid === undefined) {
return { ok: false, reason: 'exited', logPath, message: 'failed to spawn the server process' };
}
writeFileSync(pidFilePath(), `${pid}\n`, 'utf-8');
const deadline = Date.now() + START_TIMEOUT_MS;
while (Date.now() < deadline) {
if (exited) {
removePidFile();
child.unref();
return {
ok: false,
reason: 'exited',
logPath,
message: `the server exited during startup. Last lines of ${logPath}:\n${tailLog()}`,
};
}
const probe = await probeServer(statusUrl, 1000);
if (probe.up) {
child.unref();
return { ok: true, pid, url, logPath };
}
await sleep(POLL_INTERVAL_MS);
}
child.unref();
return {
ok: false,
reason: 'timeout',
pid,
url,
logPath,
message: `the server did not answer ${url} within ${START_TIMEOUT_MS / 1000}s. It may still be starting; check ${logPath}.`,
};
}
/** SIGTERM the recorded server and wait for it to actually exit. */
export async function stopDaemon(options: WebLaunchOptions): Promise<StopResult> {
const pid = readPidFile();
if (pid === null) {
const probe = await probeServer(buildStatusUrl(options), 1500);
if (probe.up) {
return {
ok: false,
reason: 'no-pidfile-but-responding',
message:
'a server is responding but there is no pidfile, so it was not started with `-d`. If it is a service use `codeman service uninstall` (or stop the unit); otherwise `pkill -f "index.js web"`.',
};
}
return { ok: true, reason: 'not-running', message: 'no daemon is running; nothing to stop' };
}
if (!isProcessAlive(pid)) {
removePidFile();
return { ok: true, pid, message: `stale pidfile removed (pid ${pid} was not running)` };
}
const command = readProcessCommand(pid);
if (command !== null && !looksLikeCodemanWeb(command)) {
return {
ok: false,
reason: 'foreign-pid',
pid,
message: `pid ${pid} is not a Codeman server (${command}). Refusing to signal it; delete ${pidFilePath()} if it is stale.`,
};
}
// SIGTERM, never SIGKILL: cli.ts flushes state on the way out.
try {
process.kill(pid, 'SIGTERM');
} catch (err) {
return { ok: false, reason: 'foreign-pid', pid, message: `could not signal pid ${pid}: ${String(err)}` };
}
const deadline = Date.now() + STOP_TIMEOUT_MS;
while (Date.now() < deadline) {
if (!isProcessAlive(pid)) {
removePidFile();
return { ok: true, pid };
}
await sleep(POLL_INTERVAL_MS);
}
return {
ok: false,
reason: 'timeout',
pid,
message: `pid ${pid} did not exit within ${STOP_TIMEOUT_MS / 1000}s. Force it with \`kill -9 ${pid}\` if you are sure.`,
};
}
/** Report on both halves: the recorded process, and whether the port answers. */
export async function daemonStatus(options: WebLaunchOptions): Promise<DaemonStatus> {
const url = buildBaseUrl(options);
const pid = readPidFile();
const probe = await probeServer(buildStatusUrl(options), 2000);
return {
pid,
running: readLivePid() !== null,
responding: probe.up,
version: probe.version,
url,
pidFile: pidFilePath(),
logPath: logFilePath(),
};
}
+57 -8
View File
@@ -105,6 +105,45 @@ export function parseLoadedImageRef(loadOutput: string): string | null {
return null;
}
/**
* Validate an imported bundle's manifest BEFORE any of its fields are trusted.
* A bundle is cross-machine input (potentially authored by someone else), and its
* fields flow into stored host/case config that the schema layer never sees:
* `engine` becomes the probe/launch binary selector, `image`/`containerWorkdir`
* reach the shellescaped launch string, `network` is a create arg. Mirror the
* DockerHostSchema/DockerCaseLinkSchema constraints here (throwing, since this is
* not a web-layer module). Exported for unit tests.
*/
export function validateImportManifest(manifest: DockerExportManifest): void {
const fail = (msg: string): never => {
throw new Error(`invalid bundle manifest: ${msg}`);
};
if (manifest.schemaVersion !== DOCKER_EXPORT_SCHEMA) {
fail(`unsupported export schema version ${manifest.schemaVersion} (expected ${DOCKER_EXPORT_SCHEMA})`);
}
if (manifest.mode !== 'full' && manifest.mode !== 'workspace') fail(`unknown mode ${String(manifest.mode)}`);
if (manifest.engine !== 'docker' && manifest.engine !== 'podman') fail(`unknown engine ${String(manifest.engine)}`);
if (typeof manifest.caseName !== 'string' || !/^[a-zA-Z0-9_-]+$/.test(manifest.caseName)) fail('bad caseName');
if (
typeof manifest.image !== 'string' ||
manifest.image.length > 512 ||
!/^[a-zA-Z0-9][\w./:@-]*$/.test(manifest.image)
) {
fail('bad image reference');
}
if (
typeof manifest.containerWorkdir !== 'string' ||
manifest.containerWorkdir.length > 2000 ||
!manifest.containerWorkdir.startsWith('/') ||
// comma: --mount specs are comma-delimited CSV; shell escaping cannot protect it
/[`$\\"'\n\r;&|<>,]/.test(manifest.containerWorkdir)
) {
fail('bad containerWorkdir');
}
if (!['bridge', 'none', 'custom'].includes(manifest.network)) fail(`unknown network ${String(manifest.network)}`);
if (typeof manifest.checksums !== 'object' || manifest.checksums === null) fail('missing checksums');
}
// ========== IO helpers ==========
function run(
@@ -338,25 +377,34 @@ export async function importDockerBundle(params: {
destWorkspace: string;
engine: DockerEngine;
timestamp: number;
/** Schema-validated destination case name; the quarantine tag derives from THIS,
* never from the (attacker-authored) manifest.caseName. */
newCaseName: string;
}): Promise<ImportResult> {
const { bundlePath, destWorkspace, engine, timestamp } = params;
const { bundlePath, destWorkspace, engine, timestamp, newCaseName } = params;
const argv: string[] = [engine === 'podman' ? 'podman' : 'docker'];
if (IS_TEST_MODE) {
const raw = await fs.readFile(bundlePath, 'utf-8').catch(() => '{}');
return { manifest: JSON.parse(raw) as DockerExportManifest, workspacePath: destWorkspace };
const manifest = JSON.parse(raw) as DockerExportManifest;
validateImportManifest(manifest);
return { manifest, workspacePath: destWorkspace };
}
const stageDir = `${destWorkspace}.import-stage-${timestamp}`;
mkdirSync(stageDir, { recursive: true });
try {
await run('tar', ['-xzf', bundlePath, '-C', stageDir], { timeout: 300_000 });
// Outer-bundle traversal guard (defense in depth: GNU/bsd tar already refuse
// `..`/absolute members by default, but the bundle is cross-machine input).
const { stdout: bundleMembers } = await run('tar', ['-tzf', bundlePath], { timeout: 60_000 });
for (const member of bundleMembers.split('\n').filter(Boolean)) {
if (!isSafeTarMember(member)) throw new Error(`unsafe path in bundle archive: ${member}`);
}
await run('tar', ['--no-same-owner', '-xzf', bundlePath, '-C', stageDir], { timeout: 300_000 });
const manifestRaw = await fs.readFile(join(stageDir, 'manifest.json'), 'utf-8');
const manifest = JSON.parse(manifestRaw) as DockerExportManifest;
if (manifest.schemaVersion !== DOCKER_EXPORT_SCHEMA) {
throw new Error(`unsupported export schema version ${manifest.schemaVersion} (expected ${DOCKER_EXPORT_SCHEMA})`);
}
validateImportManifest(manifest);
// Integrity: verify checksums before trusting any member.
const workspaceTar = join(stageDir, 'workspace.tar');
@@ -385,8 +433,9 @@ export async function importDockerBundle(params: {
const { stdout } = await run(argv[0], [...argv.slice(1), 'load', '-i', imageTar], { timeout: 300_000 });
const loadedRef = parseLoadedImageRef(stdout);
if (!loadedRef) throw new Error('could not determine loaded image ref');
// Quarantine: re-tag by the loaded ref/id, never trusting the bundle's original tag.
importedImage = importedImageTag(manifest.caseName, timestamp);
// Quarantine: re-tag by the loaded ref/id, never trusting the bundle's original
// tag; the tag name derives from the caller's schema-validated newCaseName.
importedImage = importedImageTag(newCaseName, timestamp);
await run(argv[0], [...argv.slice(1), 'tag', loadedRef, importedImage], { timeout: 60_000 });
}
+440 -34
View File
@@ -21,13 +21,15 @@
* @module docker-hosts
*/
import { existsSync, mkdirSync } from 'node:fs';
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join } from 'node:path';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import { homedir } from 'node:os';
import { createHash } from 'node:crypto';
import { execFile } from 'node:child_process';
import { execFile, spawn } from 'node:child_process';
import { promisify } from 'node:util';
import { dataPath } from './config/instance.js';
import type {
DockerCase,
DockerCommandMode,
@@ -107,6 +109,24 @@ export async function writeDockerCases(configDir: string, cases: DockerCase[]):
await writeJsonArray(configDir, dockerCasesPath(configDir), cases);
}
/**
* Persist the case's last Claude conversation id (the `--resume` seed for the
* container-recreated relaunch, docs/docker-cases-plan.md two-layer durability).
* Keyed by container name so callers that only hold a SessionDocker can update it.
* No-op when the id is unchanged or the case is gone.
*/
export async function persistDockerCaseClaudeSessionId(
configDir: string,
containerName: string,
claudeSessionId: string
): Promise<void> {
const cases = await readDockerCases(configDir);
const idx = cases.findIndex((c) => (c.container ?? dockerContainerName(c.name)) === containerName);
if (idx === -1 || cases[idx].lastClaudeSessionId === claudeSessionId) return;
cases[idx] = { ...cases[idx], lastClaudeSessionId: claudeSessionId };
await writeDockerCases(configDir, cases);
}
// ========== Naming / display / defaults ==========
/** Per-case container name. Mirrors how remote derives a stable name from the case. */
@@ -123,6 +143,7 @@ export function defaultDockerCommandForMode(mode: SessionMode): string {
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
antigravity: 'exec agy',
};
return commands[mode as DockerCommandMode] || commands.shell;
}
@@ -388,37 +409,310 @@ export function buildDockerCreateArgs(ctx: DockerCreateContext): string[] {
return args;
}
/**
* PURE argv for building the agent base image locally (the programmatic mirror of
* scripts/build-agent-image.mjs): `build -f <dockerfile> -t <image> [--no-cache]
* <contextDir>`. Kept pure + unit-testable; the caller prepends the engine binary.
*/
export function agentImageBuildArgs(dockerfile: string, image: string, contextDir: string, noCache = false): string[] {
return ['build', '-f', dockerfile, '-t', image, ...(noCache ? ['--no-cache'] : []), contextDir];
}
// ========== Credential mount resolution (IO) ==========
/** Host cred paths mapped to their in-container HOME location. */
const CREDENTIAL_PATHS: Array<{ rel: string }> = [
{ rel: '.claude' },
{ rel: '.claude.json' },
{ rel: '.codex' },
{ rel: '.gemini' },
{ rel: '.config/gcloud' },
{ rel: '.config/opencode' },
/** Container Claude config dir (created gid-0 writable in the image). */
export const CONTAINER_CLAUDE_DIR = `${CONTAINER_HOME}/.claude`;
/** In-container path of the seeded (writable) `~/.claude.json`. */
export const CLAUDE_JSON_HOME = `${CONTAINER_HOME}/.claude.json`;
/** In-container path of the read-only host-seeded `~/.claude.json` (copied into HOME at launch). */
export const CLAUDE_JSON_SEED = `${CONTAINER_HOME}/.codeman/claude.seed.json`;
/** Read-only seed paths for the files copied into the container's `.claude`. */
const CLAUDE_CREDS_SEED = `${CONTAINER_HOME}/.codeman/claude-creds.seed.json`;
const CLAUDE_SETTINGS_SEED = `${CONTAINER_HOME}/.codeman/claude-settings.seed.json`;
const CLAUDE_STATS_SEED = `${CONTAINER_HOME}/.codeman/claude-stats.seed.json`;
/** Staging root for read-only host-cred seed mounts (codex/gemini/gcloud/opencode). */
const CRED_SEED_DIR = `${CONTAINER_HOME}/.codeman/cred-seeds`;
/**
* PURE: merge the host `~/.claude.json` into a config that makes an
* already-authenticated Claude skip its INTERACTIVE onboarding inside the container
* (the host file itself lacks these flags — the host install is grandfathered, so a
* verbatim copy still triggers the theme picker + login wizard + folder-trust
* prompt). Forces `hasCompletedOnboarding`, a `theme` (so the theme picker is
* skipped), and marks the workspace project trusted + onboarded. Auth still comes
* from the copied `oauthAccount` + the dir-mounted `~/.claude/.credentials.json`.
*/
export function buildSeamlessClaudeConfig(
hostConfig: Record<string, unknown>,
workspacePath: string,
theme = 'dark'
): Record<string, unknown> {
const merged: Record<string, unknown> = { ...hostConfig };
merged.hasCompletedOnboarding = true;
if (typeof merged.theme !== 'string') merged.theme = theme;
const projects = { ...((merged.projects as Record<string, Record<string, unknown>> | undefined) ?? {}) };
const existing = (projects[workspacePath] as Record<string, unknown> | undefined) ?? {};
const seenCount = existing.projectOnboardingSeenCount;
projects[workspacePath] = {
...existing,
hasTrustDialogAccepted: true,
hasCompletedProjectOnboarding: true,
projectOnboardingSeenCount: typeof seenCount === 'number' && seenCount > 0 ? seenCount : 1,
};
merged.projects = projects;
return merged;
}
/** Best-effort read of the host `~/.claude/settings.json` theme (drives the seed's theme). */
function readHostClaudeTheme(home: string): string | undefined {
try {
const parsed = JSON.parse(readFileSync(join(home, '.claude', 'settings.json'), 'utf-8')) as { theme?: unknown };
return typeof parsed.theme === 'string' ? parsed.theme : undefined;
} catch {
return undefined;
}
}
/**
* Resolve the read-only seed mount for `~/.claude.json`. Reads the host file, merges
* in the seamless-onboarding flags + workspace trust (buildSeamlessClaudeConfig),
* writes the result to a per-container seed file under `~/.codeman/docker-seeds/`,
* and returns its mount. The launch chain copies it to `~/.claude.json` inside HOME
* once — giving Claude a NORMAL writable, already-onboarded config (no atomic-rename
* EBUSY, no re-auth, no theme/trust prompts). Falls back to the RAW host file when
* parse/write fails (auth still works; the wizard may show). Returns null when the
* host has no `~/.claude.json`. IO; under VITEST returns the raw mount (no write).
*/
export function resolveClaudeJsonSeedMount(
home: string = homedir(),
containerName?: string,
workspacePath?: string
): DockerMount | null {
const src = join(home, '.claude.json');
if (!existsSync(src)) return null;
const rawMount: DockerMount = { src, dst: CLAUDE_JSON_SEED, readonly: true };
if (IS_TEST_MODE || !containerName || !workspacePath) return rawMount;
try {
const hostConfig = JSON.parse(readFileSync(src, 'utf-8')) as Record<string, unknown>;
const merged = buildSeamlessClaudeConfig(hostConfig, workspacePath, readHostClaudeTheme(home) ?? 'dark');
const seedsDir = dataPath('docker-seeds');
if (!existsSync(seedsDir)) mkdirSync(seedsDir, { recursive: true });
const seedFile = join(seedsDir, `${containerName}.json`);
writeFileSync(seedFile, JSON.stringify(merged), { mode: 0o600 });
return { src: seedFile, dst: CLAUDE_JSON_SEED, readonly: true };
} catch {
return rawMount; // partial host write / unreadable — auth still carries, wizard may show
}
}
/** A file (or dir, when `recursive`) copied into the container HOME once at launch
* (`[ -e to ] || cp [-a] from to`). */
export interface DockerSeedCopy {
from: string;
to: string;
/** `cp -a` for whole-directory credential seeds (gemini/gcloud/opencode). */
recursive?: boolean;
}
export interface DockerClaudeArtifacts {
/** Bind mounts to add: the shared `projects/` transcripts (RW) + read-only seed files. */
mounts: DockerMount[];
/** Files copied into the container's writable HOME/.claude (+ HOME/.claude.json) at launch. */
seedCopies: DockerSeedCopy[];
}
/**
* Resolve the ISOLATED Claude artifacts for a docker session (replaces the old
* whole-`~/.claude` RW mount that polluted the host). Shares ONLY what must cross
* the boundary and seeds the rest as writable copies:
* - `~/.claude/projects` → RW dir mount (transcripts: host watchers + `--resume`).
* - `~/.claude.json` → merged onboarding seed, copied to HOME (no re-auth/wizard).
* - `~/.claude/.credentials.json` + `~/.claude/settings.json` → read-only seeds
* copied into the container's own `~/.claude` (token + global prefs carry in;
* the container refreshes its own copy and never writes back to the host).
* Everything else Claude writes (backups, tasks, teams, session-env, history) stays
* container-local. IO (reads host files, writes the merged `.claude.json` seed).
*/
export function resolveDockerClaudeArtifacts(
home: string,
containerName: string,
workspacePath: string
): DockerClaudeArtifacts {
const mounts: DockerMount[] = [];
const seedCopies: DockerSeedCopy[] = [];
// The ONE genuinely-shared part: conversation transcripts (dir mount → renames work).
const projectsSrc = join(home, '.claude', 'projects');
if (existsSync(projectsSrc)) {
mounts.push({ src: projectsSrc, dst: `${CONTAINER_CLAUDE_DIR}/projects` });
}
// ~/.claude.json → merged, onboarding-complete seed at HOME root.
const jsonSeed = resolveClaudeJsonSeedMount(home, containerName, workspacePath);
if (jsonSeed) {
mounts.push(jsonSeed);
seedCopies.push({ from: CLAUDE_JSON_SEED, to: CLAUDE_JSON_HOME });
}
// credentials (token) + settings (theme/model/effort/permissions) + stats-cache
// (drives the model/effort status indicator) → writable copies inside the
// container's own ~/.claude (never a wholesale mount → no host pollution).
const files: Array<[rel: string, seed: string, dest: string]> = [
['.credentials.json', CLAUDE_CREDS_SEED, `${CONTAINER_CLAUDE_DIR}/.credentials.json`],
['settings.json', CLAUDE_SETTINGS_SEED, `${CONTAINER_CLAUDE_DIR}/settings.json`],
['stats-cache.json', CLAUDE_STATS_SEED, `${CONTAINER_CLAUDE_DIR}/stats-cache.json`],
];
for (const [rel, seed, dest] of files) {
const src = join(home, '.claude', rel);
if (existsSync(src)) {
mounts.push({ src, dst: seed, readonly: true });
seedCopies.push({ from: seed, to: dest });
}
}
return { mounts, seedCopies };
}
/**
* Per-CLI credential-store isolation policy (the codex/gemini/gcloud/opencode analog
* of resolveDockerClaudeArtifacts). Codex is the direct Claude-analog: its
* `sessions/` rollouts + `history.jsonl` are read HOST-SIDE (response-viewer +
* `codex resume`), so they are SHARED (RW), while `auth.json`/`config.toml` are
* seeded. The other three have no host-read/resume dependency and are fully
* seed-copied (writable copy in the container, no write-back to the host).
*/
interface CredStorePolicy {
/** Path relative to HOME (host + container), e.g. '.codex' or '.config/gcloud'. */
rel: string;
/** Subdirs bind-mounted RW (shared: resume + host reads). */
shareDirs?: string[];
/** Files bind-mounted RW (append-only, e.g. codex history.jsonl — never renamed). */
shareFiles?: string[];
/** Files seeded (RO mount → cp) into the container's own copy. */
seedFiles?: string[];
/** Seed the WHOLE dir (RO mount → cp -a) — for stores with no shared/host-read state. */
seedWhole?: boolean;
}
const CRED_STORES: CredStorePolicy[] = [
{ rel: '.codex', shareDirs: ['sessions'], shareFiles: ['history.jsonl'], seedFiles: ['auth.json', 'config.toml'] },
// Also covers Antigravity: `agy` nests its whole state (auth `jetski_state.pbtxt`,
// `conversations/`, `knowledge/`) under `~/.gemini/antigravity-cli/`, so it needs no
// entry of its own. There is no `~/.antigravity` credential dir to add.
{ rel: '.gemini', seedWhole: true },
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
];
/**
* Resolve which host credential dirs/files EXIST and map them to their container
* HOME location. Only-existing avoids docker auto-creating root-owned empty dirs
* in the user's home. `~/.claude` also carries the transcripts (bind-mounted so
* host watchers + `--resume` see them) and is therefore mounted read-WRITE.
* Resolve the ISOLATED codex/gemini/gcloud/opencode artifacts (replaces the old
* whole-dir RW mounts that let each in-container CLI write its refreshed tokens +
* session state back into the host). Every path is existsSync-gated (on most hosts
* only a subset exists). Pure-ish IO (no writes; just existence checks + mount specs).
*/
export function resolveCredentialMounts(home: string = homedir()): DockerMount[] {
export function resolveDockerCredentialArtifacts(home: string = homedir()): DockerClaudeArtifacts {
const mounts: DockerMount[] = [];
for (const { rel } of CREDENTIAL_PATHS) {
const src = join(home, rel);
if (existsSync(src)) {
mounts.push({ src, dst: `${CONTAINER_HOME}/${rel}` });
const seedCopies: DockerSeedCopy[] = [];
for (const store of CRED_STORES) {
const hostBase = join(home, store.rel);
if (!existsSync(hostBase)) continue;
const containerBase = `${CONTAINER_HOME}/${store.rel}`;
const seedName = store.rel.replace(/\//g, '-'); // '.config/gcloud' → '.config-gcloud'
if (store.seedWhole) {
const seed = `${CRED_SEED_DIR}/${seedName}`;
mounts.push({ src: hostBase, dst: seed, readonly: true });
seedCopies.push({ from: seed, to: containerBase, recursive: true });
continue;
}
for (const sub of store.shareDirs ?? []) {
const src = join(hostBase, sub);
if (existsSync(src)) mounts.push({ src, dst: `${containerBase}/${sub}` });
}
for (const file of store.shareFiles ?? []) {
const src = join(hostBase, file);
if (existsSync(src)) mounts.push({ src, dst: `${containerBase}/${file}` });
}
for (const file of store.seedFiles ?? []) {
const src = join(hostBase, file);
if (existsSync(src)) {
const seed = `${CRED_SEED_DIR}/${seedName}-${file}`;
mounts.push({ src, dst: seed, readonly: true });
seedCopies.push({ from: seed, to: `${containerBase}/${file}` });
}
}
}
return mounts;
return { mounts, seedCopies };
}
// ========== Daemon probes (IO; no-op under VITEST) ==========
/**
* UNESCAPED argv prefix for execFile-based probes. The shellescaped
* buildDockerBaseArgs variant is for interpolation into the `bash -c` launch
* string; argv arrays must NOT carry literal quotes (mirror of docker-export's
* dockerArgv).
*/
function dockerEngineArgv(docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>): string[] {
const argv: string[] = [docker.engine === 'podman' ? 'podman' : 'docker'];
if (docker.context) argv.push('--context', docker.context);
if (docker.daemonHost) argv.push('-H', docker.daemonHost);
return argv;
}
export interface DockerDriftStatus {
/** Container exists (daemon reachable AND a container with this name is present). */
exists: boolean;
running: boolean;
/** The desired configHash no longer matches the container's codeman.confighash label. */
drifted: boolean;
currentHash?: string;
}
/**
* Drift check (docs/docker-cases-plan.md §4): compare the DESIRED configHash
* against the existing container's `codeman.confighash` label so docker-host
* config edits actually take effect instead of being silently ignored by the
* idempotent inspect-or-create launch chain. `exists:false` (no container /
* daemon down) means there is nothing to drift. No-op under VITEST.
*/
export async function checkDockerConfigDrift(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName' | 'configHash'>
): Promise<DockerDriftStatus> {
if (IS_TEST_MODE) return { exists: false, running: false, drifted: false };
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(
argv[0],
[
...argv.slice(1),
'inspect',
'-f',
'{{.State.Running}}\t{{index .Config.Labels "codeman.confighash"}}',
docker.containerName,
],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const [running = '', hash = ''] = stdout.trim().split('\t');
return { exists: true, running: running === 'true', drifted: hash !== docker.configHash, currentHash: hash };
} catch {
return { exists: false, running: false, drifted: false };
}
}
/**
* `docker rm -f` the case container (the recreate-on-drift confirm action; the
* launch chain recreates it with the new config on next start). Workspace +
* transcripts ride bind mounts and survive; the conversation resumes via the
* case's lastClaudeSessionId. No-op under VITEST.
*/
export async function removeDockerContainer(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost' | 'containerName'>
): Promise<void> {
if (IS_TEST_MODE) return;
const argv = dockerEngineArgv(docker);
await execFileAsync(argv[0], [...argv.slice(1), 'rm', '-f', docker.containerName], { timeout: 30_000 });
}
export interface DockerAvailability {
ok: boolean;
engine: DockerEngine;
@@ -497,11 +791,16 @@ export async function checkDockerAvailable(engine?: DockerEngine): Promise<Docke
};
}
/** Is the base image present locally? (never triggers an auto-pull). */
export async function checkDockerImagePresent(engine: DockerEngine, image: string): Promise<boolean> {
/** Is the base image present on the host's daemon? (never triggers an auto-pull).
* Honors context/daemonHost so a remote-daemon host is probed on the RIGHT daemon. */
export async function checkDockerImagePresent(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>,
image: string
): Promise<boolean> {
if (IS_TEST_MODE) return true;
const argv = dockerEngineArgv(docker);
try {
await execFileAsync(engine, ['image', 'inspect', '--format', '{{.Id}}', image], {
await execFileAsync(argv[0], [...argv.slice(1), 'image', 'inspect', '--format', '{{.Id}}', image], {
timeout: DOCKER_PROBE_TIMEOUT_MS,
});
return true;
@@ -510,6 +809,108 @@ export async function checkDockerImagePresent(engine: DockerEngine, image: strin
}
}
export interface EnsureImageResult {
ok: boolean;
/** true when this call actually ran a build (vs. the image already existing). */
built: boolean;
alreadyPresent: boolean;
error?: string;
}
/** In-flight builds keyed by `engine:image`, so concurrent callers share ONE build. */
const inFlightImageBuilds = new Map<string, Promise<EnsureImageResult>>();
/**
* Resolve the repo's Dockerfile + build context. Works from BOTH src (dev/tsx) and
* dist/index.js (esbuild prod: dist sits at repo root), since both are one level
* under the repo root. Returns null when the Dockerfile is absent (npm-global
* installs don't ship docker/ — Docker cases are a git-clone feature).
*/
function resolveAgentDockerfile(): { dockerfile: string; contextDir: string } | null {
const repoRoot = join(dirname(fileURLToPath(import.meta.url)), '..');
const dockerfile = join(repoRoot, 'docker', 'agent.Dockerfile');
return existsSync(dockerfile) ? { dockerfile, contextDir: repoRoot } : null;
}
/**
* Ensure the agent base image exists, BUILDING it locally on first use so a missing
* image is never a hard blocker (decision: "build locally on first use",
* docs/docker-cases-plan.md). Idempotent, concurrency-safe (one build per
* engine:image shared by concurrent callers), and a no-op under VITEST. Only the
* DEFAULT image is auto-built — we can never build a user's custom ref, and the
* `--pull=never` invariant forbids pulling. `onProgress` receives build output
* lines for SSE surfacing.
*/
export async function ensureAgentBaseImage(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>,
image: string,
opts: { onProgress?: (line: string) => void; noCache?: boolean } = {}
): Promise<EnsureImageResult> {
if (IS_TEST_MODE) return { ok: true, built: false, alreadyPresent: true };
if (await checkDockerImagePresent(docker, image)) {
return { ok: true, built: false, alreadyPresent: true };
}
if (image !== DEFAULT_AGENT_IMAGE) {
return {
ok: false,
built: false,
alreadyPresent: false,
error: `image ${image} is not present and only ${DEFAULT_AGENT_IMAGE} is auto-built. Build or pull ${image} yourself.`,
};
}
const key = `${docker.engine}:${image}`;
const existing = inFlightImageBuilds.get(key);
if (existing) return existing;
const build = buildAgentImage(docker, image, opts).finally(() => inFlightImageBuilds.delete(key));
inFlightImageBuilds.set(key, build);
return build;
}
function buildAgentImage(
docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>,
image: string,
opts: { onProgress?: (line: string) => void; noCache?: boolean }
): Promise<EnsureImageResult> {
const resolved = resolveAgentDockerfile();
if (!resolved) {
return Promise.resolve({
ok: false,
built: false,
alreadyPresent: false,
error: `docker/agent.Dockerfile not found in this install; clone the repo or build ${image} manually`,
});
}
const argv = dockerEngineArgv(docker);
const args = [
...argv.slice(1),
...agentImageBuildArgs(resolved.dockerfile, image, resolved.contextDir, opts.noCache),
];
return new Promise<EnsureImageResult>((resolve) => {
// async spawn (NEVER spawnSync) so a multi-minute build never wedges the event loop.
const child = spawn(argv[0], args, { stdio: ['ignore', 'pipe', 'pipe'] });
const forward = (buf: Buffer) => {
for (const line of buf.toString('utf-8').split('\n')) {
const trimmed = line.trimEnd();
if (trimmed) opts.onProgress?.(trimmed);
}
};
child.stdout?.on('data', forward);
child.stderr?.on('data', forward);
child.on('error', (err) => {
resolve({
ok: false,
built: false,
alreadyPresent: false,
error: `could not spawn ${argv[0]} build: ${err.message}`,
});
});
child.on('exit', (code) => {
if (code === 0) resolve({ ok: true, built: true, alreadyPresent: false });
else resolve({ ok: false, built: false, alreadyPresent: false, error: `${argv[0]} build failed (exit ${code})` });
});
});
}
export interface DockerTmuxCheckResult {
ok: boolean;
tmuxPath?: string;
@@ -524,21 +925,21 @@ export interface DockerTmuxCheckResult {
* (`--pull=never`). No-op under VITEST. Mirror of checkRemoteTmuxAvailable.
*/
export async function checkDockerTmuxAvailable(
docker: Pick<SessionDocker, 'engine' | 'image'>
docker: Pick<SessionDocker, 'engine' | 'image' | 'context' | 'daemonHost'>
): Promise<DockerTmuxCheckResult> {
if (IS_TEST_MODE) return { ok: true, tmuxPath: '/usr/bin/tmux' };
const engine = docker.engine;
if (!(await checkDockerImagePresent(engine, docker.image))) {
if (!(await checkDockerImagePresent(docker, docker.image))) {
return {
ok: false,
imageMissing: true,
error: `base image ${docker.image} not present: build it with 'node scripts/build-agent-image.mjs' (or pull it)`,
error: `image ${docker.image} not present (the default image is auto-built on first use; a custom image must be built or pulled first)`,
};
}
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(
engine,
['run', '--rm', '--pull=never', docker.image, 'sh', '-lc', 'command -v tmux'],
argv[0],
[...argv.slice(1), 'run', '--rm', '--pull=never', docker.image, 'sh', '-lc', 'command -v tmux'],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const tmuxPath = stdout.trim();
@@ -635,16 +1036,21 @@ export async function reapOrphanedDockerContainers(
* Returns undefined on any failure. No-op under VITEST.
*/
export async function probeDockerCliVersion(
docker: Pick<SessionDocker, 'engine' | 'containerName'>,
docker: Pick<SessionDocker, 'engine' | 'containerName' | 'context' | 'daemonHost'>,
mode: SessionMode
): Promise<string | undefined> {
if (IS_TEST_MODE) return undefined;
const bin = mode === 'shell' ? null : mode;
if (!bin) return undefined;
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(docker.engine, ['exec', docker.containerName, bin, '--version'], {
timeout: DOCKER_PROBE_TIMEOUT_MS,
});
const { stdout } = await execFileAsync(
argv[0],
[...argv.slice(1), 'exec', docker.containerName, bin, '--version'],
{
timeout: DOCKER_PROBE_TIMEOUT_MS,
}
);
const match = stdout.trim().match(/\d+\.\d+\.\d+/);
return match ? match[0] : stdout.trim() || undefined;
} catch {
+570 -25
View File
@@ -10,26 +10,29 @@
* Key exports:
* - `generateHooksConfig()` — returns hooks object for settings.local.json
* - `writeHooksConfig(casePath)` — writes hooks + env config to disk
* - `ensureCodemanHooks(casePath)` — safely installs/updates hooks for a managed case
* - `updateCaseEnvVars(casePath, envVars)` — merges env vars into settings
*
* Hook events generated: `idle_prompt`, `permission_prompt`, `elicitation_dialog`,
* `stop`, `teammate_idle`, `task_completed`
*
* Hook categories: `Notification` (3 matchers), `Stop` (1), `TeammateIdle` (1),
* `TaskCompleted` (1)
* Hook categories: `Notification` (3 matchers), `Stop` (1), `SubagentStop` (1),
* `TeammateIdle` (1), `TaskCompleted` (1), `PostToolUse` (1 self-contained
* background Bash rewake)
*
* @dependencies types (HookEventType), config/auth-config (HOOK_TIMEOUT_MS)
* @dependencies types (HookEventType), config/auth-config (HOOK_TIMEOUT_SECONDS)
* @consumedby web/server (session creation), session-cli-builder (env setup)
*
* @module hooks-config
*/
import { existsSync } from 'node:fs';
import { readFile, writeFile, mkdir } from 'node:fs/promises';
import { join } from 'node:path';
import { readFile, writeFile, mkdir, lstat, readdir, unlink, rmdir } from 'node:fs/promises';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import type { HookEventType } from './types.js';
import { HOOK_TIMEOUT_MS } from './config/auth-config.js';
import { HOOK_TIMEOUT_SECONDS } from './config/auth-config.js';
/**
* Serializes read-modify-write access to a `settings.local.json` path. Every
@@ -40,6 +43,230 @@ import { HOOK_TIMEOUT_MS } from './config/auth-config.js';
* are independent; the map self-prunes when a path's chain goes idle.
*/
const settingsWriteLocks = new Map<string, Promise<unknown>>();
/**
* Version-agnostic ownership prefix: every rewake script version embeds a marker
* starting with this, and `isCodemanHookHandler` matches on the prefix. That way a
* version bump replaces the old handler instead of duplicating it (matching on the
* full versioned marker would disown every older script).
*/
const BACKGROUND_WAKE_MARKER_PREFIX = 'CODEMAN_BACKGROUND_REWAKE_V';
/**
* Current script version. Bump the suffix whenever `generateBackgroundWakeScript`
* changes: `refreshStaleCodemanHooks` treats the absence of the CURRENT marker as
* stale, so healed cases pick up the new script on next launch.
*/
const BACKGROUND_WAKE_MARKER = `${BACKGROUND_WAKE_MARKER_PREFIX}3`;
const SUBAGENT_STOP_GUARD_MARKER_PREFIX = 'CODEMAN_SUBAGENT_STOP_GUARD_V';
const SUBAGENT_STOP_GUARD_MARKER = `${SUBAGENT_STOP_GUARD_MARKER_PREFIX}1`;
const BACKGROUND_WAKE_TIMEOUT_SECONDS = 6 * 60 * 60;
/**
* Inline Node helper for Claude Code's `asyncRewake` hook.
*
* A background Bash tool returns immediately with a task ID, then Claude writes
* its completion as a queue-operation in the top-level transcript. Subagent hooks
* receive their own transcript path even though their completion is parent-owned,
* so the helper watches both paths. Watching durable records avoids injecting
* terminal input (which could submit a user's draft).
* The helper is embedded in settings via `node -e`, so it has no script path
* that can go stale after an install or plugin-cache cleanup.
*
* Self-terminating: Claude Code enforces the hook timeout, but the helper does not
* rely on it. It exits on its own deadline (same budget) and when orphaned
* (`ppid === 1`), so a dead session cannot leave a poller stat-ing the transcript
* forever. The ppid check misses subreaper setups; the deadline is the backstop.
*/
export function generateBackgroundWakeScript(): string {
return [
"const fs = require('node:fs');",
"const path = require('node:path');",
`const ${BACKGROUND_WAKE_MARKER} = true;`,
`const deadline = Date.now() + ${BACKGROUND_WAKE_TIMEOUT_SECONDS} * 1000;`,
"const RESULT_BEGIN = '=== CODEMAN_RESULT_BEGIN ===';",
"const RESULT_END = '=== CODEMAN_RESULT_END ===';",
'const MAX_RESULT_CHARS = 65536;',
'let input = {};',
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
'function findTaskId(value) {',
" const idKeys = new Set(['taskId', 'task_id', 'shellId', 'shell_id', 'backgroundTaskId', 'background_task_id']);",
' const stack = [value];',
' const seen = new Set();',
' while (stack.length > 0) {',
' const current = stack.pop();',
" if (!current || typeof current !== 'object' || seen.has(current)) continue;",
' seen.add(current);',
' for (const [key, nested] of Object.entries(current)) {',
" if (idKeys.has(key) && typeof nested === 'string' && /^[A-Za-z0-9_-]+$/.test(nested)) return nested;",
" if (nested && typeof nested === 'object') stack.push(nested);",
' }',
' }',
" const serialized = JSON.stringify(value ?? '');",
' const messageMatch = serialized.match(/Command running in background with ID:\\s*([A-Za-z0-9_-]+)/i);',
' if (messageMatch) return messageMatch[1];',
' const pathMatch = serialized.match(/[\\\\/]tasks[\\\\/]([A-Za-z0-9_-]+)\\.output/i);',
' return pathMatch ? pathMatch[1] : null;',
'}',
'const taskId = findTaskId(input.tool_response);',
"const transcriptPath = typeof input.transcript_path === 'string' ? input.transcript_path : '';",
'if (!taskId || !transcriptPath) process.exit(0);',
'const transcriptPaths = [transcriptPath];',
'const sessionDir = path.dirname(path.dirname(transcriptPath));',
"if (typeof input.agent_id === 'string' && path.basename(path.dirname(transcriptPath)) === 'subagents' &&",
" typeof input.session_id === 'string' && path.basename(sessionDir) === input.session_id) {",
" transcriptPaths.push(sessionDir + '.jsonl');",
'}',
'const transcripts = [...new Set(transcriptPaths)].map((transcript) => {',
' let position = 0;',
' try { position = Math.max(0, fs.statSync(transcript).size - 262144); } catch {}',
" return { path: transcript, position, carry: '' };",
'});',
'if (!transcripts.some((transcript) => fs.existsSync(transcript.path))) process.exit(0);',
'function readMarkedResult(outputPath) {',
" if (!outputPath || !path.isAbsolute(outputPath) || path.basename(outputPath) !== taskId + '.output') return '';",
" if (path.basename(path.dirname(outputPath)) !== 'tasks') return '';",
' try {',
' const size = fs.statSync(outputPath).size;',
' const length = Math.min(size, MAX_RESULT_CHARS * 2);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(outputPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
' fs.closeSync(fd);',
" const text = buffer.subarray(0, bytes).toString('utf8');",
' const begin = text.lastIndexOf(RESULT_BEGIN);',
' const end = text.indexOf(RESULT_END, begin + RESULT_BEGIN.length);',
" if (begin < 0 || end < 0) return '';",
' let result = text.slice(begin + RESULT_BEGIN.length, end).trim();',
" if (!result) return '';",
' if (result.length > MAX_RESULT_CHARS) {',
' const half = Math.floor(MAX_RESULT_CHARS / 2);',
" result = result.slice(0, half) + '\\n\\n[report truncated by Codeman]\\n\\n' + result.slice(-half);",
' }',
" return '\\n\\nCompleted task report:\\n<codeman-background-result>\\n' + result + '\\n</codeman-background-result>';",
" } catch { return ''; }",
'}',
'function inspect(text) {',
' for (const line of text.split(/\\r?\\n/)) {',
' if (!line.includes(taskId)) continue;',
' let entry;',
' try { entry = JSON.parse(line); } catch { continue; }',
" if (entry.type !== 'queue-operation' || entry.operation !== 'enqueue' || typeof entry.content !== 'string') continue;",
" if (!entry.content.includes('<task-id>' + taskId + '</task-id>')) continue;",
' const status = entry.content.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
' if (!status) continue;',
' const output = entry.content.match(/<output-file>([^<]+)<\\/output-file>/i);',
" const outputPath = output ? output[1].trim() : '';",
" const location = outputPath ? ' Read ' + outputPath + ' and' : '';",
' const result = readMarkedResult(outputPath);',
" console.error('Background command ' + taskId + ' ' + status[1].toLowerCase() + '.' + location + ' continue the task.' + result);",
' process.exit(2);',
' }',
'}',
'function pollTranscript(transcript) {',
' try {',
' const size = fs.statSync(transcript.path).size;',
" if (size < transcript.position) { transcript.position = 0; transcript.carry = ''; }",
' if (size > transcript.position) {',
' const length = Math.min(size - transcript.position, 1048576);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(transcript.path, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, transcript.position);',
' fs.closeSync(fd);',
' transcript.position += bytes;',
" transcript.carry = (transcript.carry + buffer.subarray(0, bytes).toString('utf8')).slice(-262144);",
' inspect(transcript.carry);',
' }',
' } catch {}',
'}',
'function poll() {',
' if (Date.now() > deadline || process.ppid === 1) process.exit(0);',
' for (const transcript of transcripts) pollTranscript(transcript);',
' setTimeout(poll, 1000);',
'}',
'poll();',
].join('\n');
}
/**
* Keep a Claude subagent alive while its Monitor or background Bash work is live.
* Claude otherwise can publish the worker's last progress sentence as an Agent
* result when one watcher ends, even if other tracked tasks are still running.
*/
export function generateSubagentStopGuardScript(): string {
return [
"const fs = require('node:fs');",
`const ${SUBAGENT_STOP_GUARD_MARKER} = true;`,
'let input = {};',
"try { input = JSON.parse(fs.readFileSync(0, 'utf8') || '{}'); } catch { process.exit(0); }",
"const transcriptPath = typeof input.agent_transcript_path === 'string' ? input.agent_transcript_path : '';",
'if (!transcriptPath) process.exit(0);',
'let text;',
'try {',
' const size = fs.statSync(transcriptPath).size;',
' const length = Math.min(size, 16 * 1024 * 1024);',
' const buffer = Buffer.allocUnsafe(length);',
" const fd = fs.openSync(transcriptPath, 'r');",
' const bytes = fs.readSync(fd, buffer, 0, length, size - length);',
' fs.closeSync(fd);',
" text = buffer.subarray(0, bytes).toString('utf8');",
'} catch { process.exit(0); }',
'const launched = new Set();',
'const finished = new Set();',
'function inspectToolResult(value) {',
" const serialized = typeof value === 'string' ? value : JSON.stringify(value ?? '');",
' for (const match of serialized.matchAll(/Command running in background with ID:\\s*([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
' for (const match of serialized.matchAll(/Monitor started \\(task ([A-Za-z0-9_-]+)/gi)) launched.add(match[1]);',
'}',
'function inspectNotifications(value) {',
" if (typeof value !== 'string' || !value.includes('<task-notification>')) return;",
' for (const match of value.matchAll(/<task-notification>([\\s\\S]*?)<\\/task-notification>/gi)) {',
' const body = match[1];',
' const id = body.match(/<task-id>([^<]+)<\\/task-id>/i);',
' const status = body.match(/<status>(completed|failed|killed|error)<\\/status>/i);',
' if (id && status) finished.add(id[1].trim());',
' }',
'}',
'for (const line of text.split(/\\r?\\n/)) {',
' let entry;',
' try { entry = JSON.parse(line); } catch { continue; }',
' const content = entry && entry.message ? entry.message.content : undefined;',
' if (Array.isArray(content)) {',
' for (const block of content) {',
" if (block && block.type === 'tool_result') inspectToolResult(block.content);",
" if (block && block.type === 'text') inspectNotifications(block.text);",
' }',
' } else {',
' inspectNotifications(content);',
' }',
' inspectNotifications(entry && entry.content);',
'}',
'function findLiveTasks(candidates) {',
' const live = new Set();',
" if (candidates.size === 0 || !fs.existsSync('/proc')) return live;",
' let processIds;',
" try { processIds = fs.readdirSync('/proc').filter((name) => /^\\d+$/.test(name)); } catch { return live; }",
' for (const processId of processIds) {',
" for (const descriptor of ['0', '1', '2']) {",
' let target;',
" try { target = fs.readlinkSync('/proc/' + processId + '/fd/' + descriptor); } catch { continue; }",
' const match = target.match(/[\\/]tasks[\\/]([A-Za-z0-9_-]+)\\.output(?: \\(deleted\\))?$/);',
' if (match && candidates.has(match[1])) live.add(match[1]);',
' }',
' if (live.size === candidates.size) break;',
' }',
' return live;',
'}',
'const unfinished = new Set([...launched].filter((taskId) => !finished.has(taskId)));',
'const active = [...findLiveTasks(unfinished)];',
'if (active.length === 0) process.exit(0);',
'const shown = active.slice(0, 8);',
"const suffix = active.length > shown.length ? ' and ' + (active.length - shown.length) + ' more' : '';",
'process.stdout.write(JSON.stringify({',
" decision: 'block',",
" reason: 'You still own active background work (' + shown.join(', ') + suffix + '). Do not return an intermediate progress message as your final report. Process the task notifications or keep actively polling until every task completes, then return one complete summary.',",
'}));',
].join('\n');
}
function withSettingsLock<T>(path: string, fn: () => Promise<T>): Promise<T> {
const prev = settingsWriteLocks.get(path) ?? Promise.resolve();
const run = prev.then(fn, fn); // run after the prior writer, regardless of its outcome
@@ -75,7 +302,11 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
const curlCmd = (event: HookEventType) =>
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"event":"${event}","sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$HOOK_DATA" | ` +
`curl -s -X POST "$CODEMAN_API_URL/api/hook-event" ` +
// `-k`, same as the statusline exporter: CODEMAN_API_URL is loopback HTTPS with
// a self-signed cert on --https/tailscale installs. Without it curl exits 60,
// the `|| true` swallows it, and ALL SIX hook events die silently: respawn loses
// its definitive idle signals and the wait endpoints lose stop/blocked.
`curl -sk -X POST "$CODEMAN_API_URL/api/hook-event" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- ` +
@@ -86,36 +317,134 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
Notification: [
{
matcher: 'idle_prompt',
hooks: [{ type: 'command', command: curlCmd('idle_prompt'), timeout: HOOK_TIMEOUT_MS }],
hooks: [{ type: 'command', command: curlCmd('idle_prompt'), timeout: HOOK_TIMEOUT_SECONDS }],
},
{
matcher: 'permission_prompt',
hooks: [{ type: 'command', command: curlCmd('permission_prompt'), timeout: HOOK_TIMEOUT_MS }],
hooks: [{ type: 'command', command: curlCmd('permission_prompt'), timeout: HOOK_TIMEOUT_SECONDS }],
},
{
matcher: 'elicitation_dialog',
hooks: [{ type: 'command', command: curlCmd('elicitation_dialog'), timeout: HOOK_TIMEOUT_MS }],
hooks: [{ type: 'command', command: curlCmd('elicitation_dialog'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
Stop: [
{
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_MS }],
hooks: [{ type: 'command', command: curlCmd('stop'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
SubagentStop: [
{
hooks: [
{
type: 'command',
command: 'node',
args: ['-e', generateSubagentStopGuardScript()],
timeout: HOOK_TIMEOUT_SECONDS,
},
],
},
],
TeammateIdle: [
{
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: HOOK_TIMEOUT_MS }],
hooks: [{ type: 'command', command: curlCmd('teammate_idle'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
TaskCompleted: [
{
hooks: [{ type: 'command', command: curlCmd('task_completed'), timeout: HOOK_TIMEOUT_MS }],
hooks: [{ type: 'command', command: curlCmd('task_completed'), timeout: HOOK_TIMEOUT_SECONDS }],
},
],
PostToolUse: [
{
matcher: 'Bash',
hooks: [
{
type: 'command',
command: 'node',
args: ['-e', generateBackgroundWakeScript()],
asyncRewake: true,
timeout: BACKGROUND_WAKE_TIMEOUT_SECONDS,
},
],
},
],
},
};
}
function isCodemanHookHandler(value: unknown): boolean {
try {
const serialized = JSON.stringify(value);
// Prefixes, not versioned markers: older script versions must still be ours.
return (
serialized.includes('/api/hook-event') ||
serialized.includes(BACKGROUND_WAKE_MARKER_PREFIX) ||
serialized.includes(SUBAGENT_STOP_GUARD_MARKER_PREFIX)
);
} catch {
return false;
}
}
/**
* Replace only Codeman-owned command handlers while preserving user events,
* matcher entries, and sibling handlers in mixed entries.
*/
function mergeCodemanHooks(existingValue: unknown, generated: Record<string, unknown[]>): Record<string, unknown[]> {
const existing =
existingValue && typeof existingValue === 'object' && !Array.isArray(existingValue)
? (existingValue as Record<string, unknown>)
: {};
const merged: Record<string, unknown[]> = {};
for (const eventName of new Set([...Object.keys(existing), ...Object.keys(generated)])) {
const existingEntries = Array.isArray(existing[eventName]) ? (existing[eventName] as unknown[]) : [];
const generatedEntries = generated[eventName];
if (!generatedEntries) {
merged[eventName] = existingEntries;
continue;
}
const entries: unknown[] = [];
let insertedGenerated = false;
for (const entry of existingEntries) {
if (!entry || typeof entry !== 'object' || Array.isArray(entry)) {
if (!isCodemanHookHandler(entry)) entries.push(entry);
continue;
}
const record = entry as Record<string, unknown>;
if (!Array.isArray(record.hooks)) {
if (isCodemanHookHandler(record)) {
if (!insertedGenerated) {
entries.push(...generatedEntries);
insertedGenerated = true;
}
} else {
entries.push(entry);
}
continue;
}
const retainedHandlers = record.hooks.filter((handler) => !isCodemanHookHandler(handler));
const removedCodemanHandler = retainedHandlers.length !== record.hooks.length;
if (removedCodemanHandler && !insertedGenerated) {
entries.push(...generatedEntries);
insertedGenerated = true;
}
if (retainedHandlers.length > 0 || !removedCodemanHandler) {
entries.push(retainedHandlers.length === record.hooks.length ? entry : { ...record, hooks: retainedHandlers });
}
}
if (!insertedGenerated) entries.push(...generatedEntries);
merged[eventName] = entries;
}
return merged;
}
/**
* Remove a subset of env keys from .claude/settings.local.json.env if present.
* Used during the disk→tmux-setenv migration: when the caller is actively setting
@@ -237,29 +566,67 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
}
const hooksConfig = generateHooksConfig();
const merged = { ...existing, ...hooksConfig };
const merged = {
...existing,
hooks: mergeCodemanHooks(existing.hooks, hooksConfig.hooks),
};
await writeFile(settingsPath, JSON.stringify(merged, null, 2) + '\n');
});
}
/**
* Self-heal a case's hooks block so the COD-91 unconditional hook-secret gate keeps
* accepting its hook events.
* Ensures an explicitly managed case has the current Codeman hooks.
*
* Unlike `refreshStaleCodemanHooks`, this may add Codeman handlers to a valid
* user-owned settings file. It is therefore reserved for case quick-starts,
* where the user has explicitly asked Codeman to manage that workspace. A
* malformed existing file is left untouched rather than replaced.
*/
export async function ensureCodemanHooks(casePath: string): Promise<void> {
const claudeDir = join(casePath, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
await withSettingsLock(settingsPath, async () => {
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
let existing: Record<string, unknown> = {};
try {
const parsed: unknown = JSON.parse(await readFile(settingsPath, 'utf-8'));
if (!parsed || typeof parsed !== 'object' || Array.isArray(parsed)) return;
existing = parsed as Record<string, unknown>;
} catch (err) {
if ((err as NodeJS.ErrnoException).code !== 'ENOENT') return;
}
const generated = generateHooksConfig();
const hooks = mergeCodemanHooks(existing.hooks, generated.hooks);
if (JSON.stringify(existing.hooks ?? {}) === JSON.stringify(hooks)) return;
await writeFile(settingsPath, JSON.stringify({ ...existing, hooks }, null, 2) + '\n');
});
}
/**
* Self-heal a case's Codeman-owned hooks block.
*
* `writeHooksConfig` only runs when a case is first CREATED. Cases created before the
* X-Codeman-Hook-Secret header was added (COD-54, 2026-06-10) keep hook curls in their
* settings.local.json that POST to /api/hook-event WITHOUT the secret — which, once the
* gate requires it unconditionally (COD-91), silently 401 on a password-protected
* install. This refreshes the hooks block so those stale curls regain the header.
* gate requires it unconditionally (COD-91), silently 401 on a password-protected install.
* Older Codeman blocks also lack the current background Bash async-rewake hook or the
* SubagentStop guard. A further stale shape: hook curls without `-k`, which exit 60 on
* every --https/tailscale install (the cert is self-signed), swallowed by the hooks'
* own `|| true` — all six hook events die silently. Refresh any of these stale shapes
* on launch so existing cases heal.
*
* Deliberately surgical: regenerates ONLY when settings.local.json already contains
* Codeman's own hook curls (they target `/api/hook-event`) that lack the secret header.
* No-op when the file/hooks are absent (we never impose hooks on a user who removed
* them), when the hooks aren't ours, or when the secret is already present — so it never
* clobbers a user's customizations and is cheap enough to call on every Claude spawn.
* Codeman's own hook curls (they target `/api/hook-event`) and they are stale. No-op
* when the file/hooks are absent (we never impose hooks on a user who removed them) or
* when the hooks aren't ours, so it is cheap enough to call on every Claude spawn.
*/
export async function refreshStaleHookSecret(casePath: string): Promise<void> {
export async function refreshStaleCodemanHooks(casePath: string): Promise<void> {
const settingsPath = join(casePath, '.claude', 'settings.local.json');
if (!existsSync(settingsPath)) return;
await withSettingsLock(settingsPath, async () => {
@@ -274,8 +641,18 @@ export async function refreshStaleHookSecret(casePath: string): Promise<void> {
// The generated curl carries this header literal (see generateHooksConfig); its
// absence on our own hooks means they predate COD-54 and need regenerating.
const hasSecret = hooksJson.includes('X-Codeman-Hook-Secret');
if (!isOurs || hasSecret) return;
const merged = { ...existing, ...generateHooksConfig() };
const hasBackgroundWake = hooksJson.includes(BACKGROUND_WAKE_MARKER);
// The pre--k curl shape: `curl -sk -X POST` does not contain `curl -s -X POST`
// as a substring, so this cleanly identifies hook curls that die with exit 60
// on a self-signed HTTPS install.
const hasTlsFlaglessCurl = hooksJson.includes('curl -s -X POST');
const hasSubagentStopGuard = hooksJson.includes(SUBAGENT_STOP_GUARD_MARKER);
if (!isOurs || (hasSecret && hasBackgroundWake && hasSubagentStopGuard && !hasTlsFlaglessCurl)) return;
const generated = generateHooksConfig();
const merged = {
...existing,
hooks: mergeCodemanHooks(existing.hooks, generated.hooks),
};
await writeFile(settingsPath, JSON.stringify(merged, null, 2) + '\n');
});
}
@@ -344,3 +721,171 @@ export async function applyStatusLineConfig(casePath: string, enabled: boolean):
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
// ─── Agent skill injection ───────────────────────────────────────────────────
/**
* Version-agnostic ownership prefix for the injected agent skill, same pattern as
* `BACKGROUND_WAKE_MARKER_PREFIX`: ownership is decided on the prefix so a wording
* change in the full marker cannot disown every previously injected copy.
*/
const AGENT_SKILL_MARKER_PREFIX = '<!-- codeman-managed-agent-skill';
/**
* Marker appended to the injected SKILL.md. Its presence is what makes a copy OURS:
* install/refresh/remove all refuse to touch a `skills/codeman` whose SKILL.md lacks
* it, so a user's hand-authored or hand-edited-and-de-marked skill is never clobbered.
*/
const AGENT_SKILL_MARKER = `${AGENT_SKILL_MARKER_PREFIX}: installed by Codeman; edits are overwritten while the agent-skill setting is on -->`;
/**
* Packaged source of the skill: `skills/codeman/` at the package root. Resolved
* relative to this module so it works from `src/` (tsx dev), `dist/` (tsc build),
* and an npm install (`files` includes `skills`), all of which sit one level below
* the package root.
*/
function agentSkillSourceDir(): string {
return join(dirname(fileURLToPath(import.meta.url)), '..', 'skills', 'codeman');
}
interface AgentSkillFile {
/** Path relative to the target skill dir (e.g. `reference/endpoints.md`). */
relPath: string;
content: string;
}
/**
* Read the packaged skill: SKILL.md (marker appended) plus every markdown file
* under `reference/`. Enumerated from disk rather than a hardcoded manifest so a
* new reference file ships without touching this module.
*/
async function readAgentSkillSource(): Promise<AgentSkillFile[]> {
const src = agentSkillSourceDir();
const skill = await readFile(join(src, 'SKILL.md'), 'utf-8');
const files: AgentSkillFile[] = [{ relPath: 'SKILL.md', content: `${skill.trimEnd()}\n\n${AGENT_SKILL_MARKER}\n` }];
let referenceNames: string[] = [];
try {
referenceNames = (await readdir(join(src, 'reference'))).filter((name) => name.endsWith('.md')).sort();
} catch {
// no reference dir in the source; SKILL.md alone is still a valid skill
}
for (const name of referenceNames) {
files.push({ relPath: join('reference', name), content: await readFile(join(src, 'reference', name), 'utf-8') });
}
return files;
}
async function isSymlink(path: string): Promise<boolean> {
try {
return (await lstat(path)).isSymbolicLink();
} catch {
return false;
}
}
/** What an install/remove actually did, so callers (CLI, logs) can say so. */
export type AgentSkillApplyResult =
| 'installed' // fresh copy written
| 'refreshed' // our copy was stale and got rewritten
| 'unchanged' // our copy already matches the packaged source
| 'removed' // our copy deleted
| 'absent' // nothing there to remove
| 'foreign' // a copy exists but is not ours; left untouched
| 'symlink'; // the skill dir (or its parent) is a symlink; left untouched
/**
* Install or refresh the Codeman agent skill into `skillDir` (a `.../codeman`
* directory, e.g. `<case>/.claude/skills/codeman` or `~/.claude/skills/codeman`).
*
* Refuses two shapes rather than writing through them:
* - a SYMLINK at the skill dir or its `skills/` parent: this repo's own dogfooding
* layout (`.claude/skills/codeman -> ../../skills/codeman`) would otherwise have
* the injector overwrite the repo source through the link;
* - a FOREIGN copy (SKILL.md present without our marker): that is the user's own
* skill, and per the statusLine rule we never clobber what we did not write.
*
* Idempotent and cheap: unchanged files are not rewritten, so calling on every
* session create causes no mtime churn.
*/
export async function installAgentSkillInto(skillDir: string): Promise<AgentSkillApplyResult> {
if ((await isSymlink(dirname(skillDir))) || (await isSymlink(skillDir))) return 'symlink';
let existing: string | null = null;
try {
existing = await readFile(join(skillDir, 'SKILL.md'), 'utf-8');
} catch {
// absent: fresh install
}
if (existing !== null && !existing.includes(AGENT_SKILL_MARKER_PREFIX)) return 'foreign';
const files = await readAgentSkillSource();
let changed = false;
for (const file of files) {
const target = join(skillDir, file.relPath);
let current: string | null = null;
try {
current = await readFile(target, 'utf-8');
} catch {
// missing: will be written
}
if (current === file.content) continue;
await mkdir(dirname(target), { recursive: true });
await writeFile(target, file.content);
changed = true;
}
if (!changed) return 'unchanged';
return existing === null ? 'installed' : 'refreshed';
}
/**
* Remove a Codeman-managed skill copy from `skillDir`. Same ownership and symlink
* refusals as the install path. Deletes only files the packaged source would have
* written (never `rm -rf`, so a user's extra files in the directory survive), then
* prunes the directories bottom-up if they emptied.
*/
export async function removeAgentSkillFrom(skillDir: string): Promise<AgentSkillApplyResult> {
if ((await isSymlink(dirname(skillDir))) || (await isSymlink(skillDir))) return 'symlink';
let existing: string | null = null;
try {
existing = await readFile(join(skillDir, 'SKILL.md'), 'utf-8');
} catch {
return 'absent';
}
if (!existing.includes(AGENT_SKILL_MARKER_PREFIX)) return 'foreign';
// Manifest-based, with SKILL.md as the fallback when the packaged source is
// unreadable: removal must still work on an install whose skills/ dir went missing.
const files = await readAgentSkillSource().catch((): AgentSkillFile[] => [{ relPath: 'SKILL.md', content: '' }]);
for (const file of files) {
await unlink(join(skillDir, file.relPath)).catch(() => {});
}
await rmdir(join(skillDir, 'reference')).catch(() => {}); // fails when non-empty, fine
await rmdir(skillDir).catch(() => {});
await rmdir(dirname(skillDir)).catch(() => {}); // prune `.claude/skills` if now empty
return 'removed';
}
/**
* Add or remove the Codeman agent skill in `<case>/.claude/skills/codeman`,
* mirroring `applyStatusLineConfig`'s shape. Gated by the synced `agentSkillEnabled`
* app setting (default OFF); callers gate on Claude mode, since the skill is discovered
* via `.claude/skills/`, which only Claude Code reads.
*
* Call-site policy is ADD-ONLY on session create (callers pass `enabled: true` or
* skip the call), for the statusLine reason: sessions in a repo share one `.claude/`
* dir, so a single create while the setting is off must not yank the skill out from
* under other live sessions.
*
* ⚠️ Consequence: turning `agentSkillEnabled` OFF sweeps nothing. There is deliberately
* no server-side toggle-off sweep (it would have to walk every case, including ones
* with live sessions, and would hit exactly the shared-`.claude/` hazard above), so
* already-injected copies stay on disk until removed per case with
* `codeman skill uninstall --case <name>`. The `enabled: false` branch here backs that
* CLI and the tests; it has no server call site. Keep the README's Agent Skill note in
* sync if this ever changes.
*/
export async function applyAgentSkill(casePath: string, enabled: boolean): Promise<AgentSkillApplyResult> {
const skillDir = join(casePath, '.claude', 'skills', 'codeman');
return enabled ? installAgentSkillInto(skillDir) : removeAgentSkillFrom(skillDir);
}
+9
View File
@@ -17,6 +17,7 @@ import type {
CodexConfig,
EffortLevel,
GeminiConfig,
AntigravityConfig,
SessionRemote,
SessionDocker,
} from './types.js';
@@ -39,6 +40,8 @@ export interface MuxSession {
remote?: SessionRemote;
/** Docker execution metadata for local tmux sessions wrapping `docker exec` */
docker?: SessionDocker;
/** Owning username in multi-user mode (round-tripped through recovery like remote/docker) */
owner?: string;
/** Session mode */
mode: SessionMode;
/** Whether webserver is attached to this session */
@@ -72,6 +75,7 @@ export interface CreateSessionOptions {
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Ephemeral — not written to disk. */
@@ -84,6 +88,8 @@ export interface CreateSessionOptions {
remote?: SessionRemote;
/** Docker execution metadata for local tmux sessions wrapping `docker exec` */
docker?: SessionDocker;
/** Owning username in multi-user mode; persisted for recovery. */
owner?: string;
}
/** Options for respawning a dead pane. */
@@ -98,6 +104,7 @@ export interface RespawnPaneOptions {
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
@@ -110,6 +117,8 @@ export interface RespawnPaneOptions {
remote?: SessionRemote;
/** Docker execution metadata for local tmux sessions wrapping `docker exec` */
docker?: SessionDocker;
/** Owning username (multi-user); redundant on respawn since the Session object survives, kept for shape parity. */
owner?: string;
}
/** Options for pane buffer capture (COD-47 full-history mode). */
+22 -2
View File
@@ -20,7 +20,7 @@ import type { TerminalMultiplexer } from './mux-interface.js';
import { existsSync, mkdirSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { RESEARCH_AGENT_PROMPT, PLANNER_PROMPT } from './prompts/index.js';
import { getErrorMessage, type PlanItem } from './types.js';
import { getErrorMessage, type PlanItem, type ClaudeMode } from './types.js';
// Re-export for backward compatibility
export type { PlanItem };
@@ -130,18 +130,28 @@ export class PlanOrchestrator {
private taskDescription = '';
private researchModel: string;
private plannerModel: string;
// Multi-user permission threading: the resolved claudeMode/owner/allowedTools for the
// internal research/planner one-shots. Left undefined = today's single-user behavior
// (the caller threads the resolved global mode, byte-identical when !isMultiUserMode()).
private claudeMode?: ClaudeMode;
private owner?: string;
private allowedTools?: string;
constructor(
mux: TerminalMultiplexer,
workingDir: string = process.cwd(),
outputDir?: string,
modelConfig?: { defaultModel?: string; agentTypeOverrides?: Record<string, string> }
modelConfig?: { defaultModel?: string; agentTypeOverrides?: Record<string, string> },
security?: { claudeMode?: ClaudeMode; owner?: string; allowedTools?: string }
) {
this.mux = mux;
this.workingDir = workingDir;
this.outputDir = outputDir;
this.researchModel = modelConfig?.agentTypeOverrides?.explore || modelConfig?.defaultModel || DEFAULT_MODEL;
this.plannerModel = modelConfig?.agentTypeOverrides?.review || modelConfig?.defaultModel || DEFAULT_MODEL;
this.claudeMode = security?.claudeMode;
this.owner = security?.owner;
this.allowedTools = security?.allowedTools;
}
private saveAgentOutput(agentType: string, prompt: string, result: unknown, durationMs: number): void {
@@ -424,6 +434,12 @@ export class PlanOrchestrator {
mux: this.mux,
useMux: false,
mode: 'claude',
// Section 6.3: run this one-shot under the caller-resolved permission mode/owner so a
// non-granted multi-user user cannot regain --dangerously-skip-permissions. Undefined
// (single-user, not threaded) is byte-identical to today (Session keeps its default).
claudeMode: this.claudeMode,
allowedTools: this.allowedTools,
owner: this.owner,
});
this.runningSessions.add(session);
@@ -580,6 +596,10 @@ export class PlanOrchestrator {
mux: this.mux,
useMux: false,
mode: 'claude',
// Section 6.3: same permission-mode/owner threading as the research one-shot above.
claudeMode: this.claudeMode,
allowedTools: this.allowedTools,
owner: this.owner,
});
this.runningSessions.add(session);
+87
View File
@@ -0,0 +1,87 @@
/**
* @fileoverview Bounded descendant walk over a process-tree snapshot.
*
* Split out of `tmux-manager.ts` so the traversal can be unit-tested directly. It
* previously lived as a private method, which meant the regression test had to keep
* its own copy of the algorithm — a test that passes while the shipped code rots.
*
* ## The incident this guards against
*
* On 2026-07-30 an unbounded version of this walk took a machine down. It ran
* `pgrep -P <pid>` once per node and recursed with no visited set, no depth limit and
* no node cap. Across ~28 adopted tmux trees the fan-out exploded, and because each
* `pgrep` blocks in the WSL kernel while reading `/proc/<pid>/cgroup`, none of them
* returned while the walk kept spawning more. Result: ~13,000 `pgrep` processes stuck
* in D-state out of ~39,000 total, load average above 13,000, and a machine only
* recoverable by restarting WSL — which cost every running session.
*
* Three properties make that impossible, and each has a test:
* 1. a cycle terminates instead of looping (stale snapshots can contain one),
* 2. depth is capped,
* 3. node count is capped.
*
* The fourth property — spawning nothing per node — is structural: this function
* takes a snapshot and cannot spawn anything at all.
*
* @module proc-tree
*/
/** Maximum generations to descend. Deeper than any real agent process tree. */
export const PROC_WALK_MAX_DEPTH = 10;
/** Hard ceiling on collected descendants. A backstop, not an expected limit. */
export const PROC_WALK_MAX_NODES = 500;
export interface WalkOptions {
maxDepth?: number;
maxNodes?: number;
/**
* Called once when a cap truncated the result, with which cap it was. Both are
* reported: a silent depth cap would hide a deep tree just as effectively as a
* silent node cap hides a wide one, and the whole point of this module is that
* truncation is visible rather than mysterious.
*/
onTruncated?: (pid: number, cap: number, reason: 'nodes' | 'depth') => void;
}
/**
* All descendants of `pid`, breadth-first and bounded.
*
* @param pid root of the walk; never included in the result
* @param byParent parent pid → child pids, from ONE `ps` snapshot
*/
export function collectDescendants(
pid: number,
byParent: ReadonlyMap<number, readonly number[]>,
opts: WalkOptions = {}
): number[] {
const maxDepth = opts.maxDepth ?? PROC_WALK_MAX_DEPTH;
const maxNodes = opts.maxNodes ?? PROC_WALK_MAX_NODES;
const out: number[] = [];
const visited = new Set<number>([pid]);
let frontier = [pid];
for (let depth = 0; depth < maxDepth && frontier.length; depth += 1) {
const next: number[] = [];
for (const parent of frontier) {
for (const child of byParent.get(parent) ?? []) {
if (visited.has(child)) continue; // a real tree has no cycles, a stale
visited.add(child); // snapshot can still produce one
out.push(child);
next.push(child);
if (out.length >= maxNodes) {
opts.onTruncated?.(pid, maxNodes, 'nodes');
return out;
}
}
}
frontier = next;
// Ran out of generations while descendants were still queued: the tree is
// deeper than the cap and the result is incomplete.
if (depth === maxDepth - 1 && frontier.length > 0) {
opts.onTruncated?.(pid, maxDepth, 'depth');
}
}
return out;
}
+25 -10
View File
@@ -9,10 +9,23 @@
import { existsSync, readFileSync, writeFileSync, mkdirSync } from 'node:fs';
import { join } from 'node:path';
import webpush from 'web-push';
import type { VapidKeys, PushSubscriptionRecord } from './types.js';
import type { VapidKeys, PushSubscriptionRecord, UserRole } from './types.js';
import { Debouncer } from './utils/index.js';
import { getDataDir } from './config/instance.js';
/**
* A push subscription plus the multi-user owner identity stamped at subscribe time.
* `username`/`role` are undefined in single-user mode (and for legacy records saved
* before this field existed). sendPushNotifications uses them to scope a
* session-notification to its owner's devices (+ admins) instead of fanning out to
* every user. Kept as a store-local widening of PushSubscriptionRecord so the shared
* type stays untouched; the extra keys serialize/persist transparently.
*/
export type OwnedPushSubscriptionRecord = PushSubscriptionRecord & {
username?: string;
role?: UserRole;
};
const DATA_DIR = getDataDir();
const KEYS_FILE = join(DATA_DIR, 'push-keys.json');
const SUBS_FILE = join(DATA_DIR, 'push-subscriptions.json');
@@ -20,7 +33,7 @@ const SAVE_DEBOUNCE_MS = 500;
export class PushSubscriptionStore {
private vapidKeys: VapidKeys | null = null;
private subscriptions: Map<string, PushSubscriptionRecord> = new Map();
private subscriptions: Map<string, OwnedPushSubscriptionRecord> = new Map();
private saveDeb = new Debouncer(SAVE_DEBOUNCE_MS);
private _disposed = false;
@@ -67,17 +80,19 @@ export class PushSubscriptionStore {
}
/** Register or update a push subscription (deduplicates by endpoint) */
addSubscription(sub: Omit<PushSubscriptionRecord, 'lastUsedAt'>): PushSubscriptionRecord {
addSubscription(sub: Omit<OwnedPushSubscriptionRecord, 'lastUsedAt'>): OwnedPushSubscriptionRecord {
// Check for existing subscription with same endpoint
for (const [existingId, existing] of this.subscriptions) {
if (existing.endpoint === sub.endpoint) {
// Update existing
const updated: PushSubscriptionRecord = {
// Update existing (re-stamp owner identity so it tracks the current caller)
const updated: OwnedPushSubscriptionRecord = {
...existing,
keys: sub.keys,
userAgent: sub.userAgent,
lastUsedAt: Date.now(),
pushPreferences: sub.pushPreferences,
username: sub.username,
role: sub.role,
};
this.subscriptions.set(existingId, updated);
this.scheduleSave();
@@ -86,7 +101,7 @@ export class PushSubscriptionStore {
}
// New subscription
const record: PushSubscriptionRecord = {
const record: OwnedPushSubscriptionRecord = {
...sub,
lastUsedAt: Date.now(),
};
@@ -96,7 +111,7 @@ export class PushSubscriptionStore {
}
/** Update push preferences for a subscription */
updatePreferences(id: string, preferences: Record<string, boolean>): PushSubscriptionRecord | null {
updatePreferences(id: string, preferences: Record<string, boolean>): OwnedPushSubscriptionRecord | null {
const sub = this.subscriptions.get(id);
if (!sub) return null;
sub.pushPreferences = preferences;
@@ -124,12 +139,12 @@ export class PushSubscriptionStore {
}
/** Get all subscriptions */
getAll(): PushSubscriptionRecord[] {
getAll(): OwnedPushSubscriptionRecord[] {
return Array.from(this.subscriptions.values());
}
/** Get a single subscription by ID */
get(id: string): PushSubscriptionRecord | null {
get(id: string): OwnedPushSubscriptionRecord | null {
return this.subscriptions.get(id) ?? null;
}
@@ -138,7 +153,7 @@ export class PushSubscriptionStore {
if (!existsSync(SUBS_FILE)) return;
try {
const raw = readFileSync(SUBS_FILE, 'utf-8');
const arr = JSON.parse(raw) as PushSubscriptionRecord[];
const arr = JSON.parse(raw) as OwnedPushSubscriptionRecord[];
for (const sub of arr) {
this.subscriptions.set(sub.id, sub);
}
+264 -5
View File
@@ -8,6 +8,7 @@ import type {
RemoteCase,
RemoteCommandMode,
RemoteHost,
RemoteSessionInfo,
RemoteSshOptions,
SessionMode,
SessionRemote,
@@ -57,16 +58,61 @@ export async function writeRemoteCases(configDir: string, cases: RemoteCase[]):
await writeJsonArray(configDir, remoteCasesPath(configDir), cases);
}
/**
* The remote user's login shell, defaulted and quoted.
*
* The default is belt-and-braces, not a live bug: an empty `$SHELL` would expand
* to `exec -i -l`, which the shell reads as `exec -i` — "not found", pane dead on
* arrival, the #208 failure all over again (verified: `sh -c 'exec $SHELL -i -l'`
* with SHELL unset prints `exec: -i: not found`). In practice tmux always exports
* SHELL into a pane from its own `default-shell` option, so the command as USED
* here is safe either way (also verified). The default matters because these
* strings are the seed values a per-host `commands.*` override is edited from, and
* nothing constrains where an edited one ends up running. Quoted for a shell path
* containing spaces. `/bin/sh` exists on every POSIX host.
*/
const REMOTE_LOGIN_SHELL = '"${SHELL:-/bin/sh}"';
/**
* Run `command` through the remote user's interactive login shell, so per-user
* PATH entries (~/.local/bin, ~/.opencode/bin, …) are resolved before the CLI name
* is looked up. ssh's remote-command execution is neither interactive nor login,
* so a bare `exec claude` sees only sshd's minimal default PATH and dies with
* "command not found" (exit 127).
*
* Shells that take neither flag (nushell, elvish, …) cannot be detected from here
* the way `loginShellArgs()` detects them locally, since the shell is whatever the
* REMOTE passwd says. A host like that is what the per-host `commands.*` override
* is for.
*/
export function remoteLoginShellCommand(command: string): string {
return `exec ${REMOTE_LOGIN_SHELL} -i -l -c ${shellescape(command)}`;
}
export function defaultRemoteCommandForMode(mode: SessionMode): string {
// Agent CLIs (claude/opencode/codex/gemini/antigravity) are typically installed
// under per-user paths like ~/.local/bin or ~/.opencode/bin, added to PATH only by
// the remote user's interactive-login shell startup files (~/.zshrc etc.). ssh's
// remote-command execution is neither interactive nor login, so a bare `exec
// claude` sees only sshd's minimal default PATH and fails with "command not
// found" (exit 127) — confirmed via `tmux capture-pane` on the
// remain-on-exit-preserved dead pane. Route through `$SHELL -i -l -c`, the same
// fix already used for shell mode below, so PATH is fully resolved before the
// CLI name is looked up.
const commands: Record<RemoteCommandMode, string> = {
shell: 'exec bash -l',
// $SHELL, not a hardcoded bash: sshd sets it from the remote user's
// /etc/passwd entry, so this launches their actual login shell (zsh,
// fish, etc.). -i -l so it sources rc files (~/.zshrc etc.), matching
// the local shell-mode launch.
shell: `exec ${REMOTE_LOGIN_SHELL} -i -l`,
// Mirror the LOCAL claude default so the remote agent runs non-interactively
// (no trust-folder/permission prompt that nothing on the remote answers). The
// per-host `commands.claude` override stays the escape hatch.
claude: 'exec claude --dangerously-skip-permissions',
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
claude: remoteLoginShellCommand('claude --dangerously-skip-permissions'),
opencode: remoteLoginShellCommand('opencode'),
codex: remoteLoginShellCommand('codex'),
gemini: remoteLoginShellCommand('gemini'),
antigravity: remoteLoginShellCommand('agy'),
};
return commands[mode as RemoteCommandMode] || commands.shell;
}
@@ -173,6 +219,14 @@ export interface RemoteTmuxCheckResult {
export async function checkRemoteTmuxAvailable(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions
): Promise<RemoteTmuxCheckResult> {
// Under vitest, never open a real ssh connection — mirrors TmuxManager's
// no-op-shell-under-VITEST (IS_TEST_MODE). Without this, remote-case
// create-path tests hit a real ~10s ssh timeout. The command construction is
// covered by buildRemoteTmuxCheckCommand unit tests; only the live probe is
// short-circuited here.
if (process.env.VITEST) {
return { ok: true, tmuxPath: '(test-mode)' };
}
const command = buildRemoteTmuxCheckCommand(host);
try {
const { stdout } = await execAsync(command, { timeout: 15_000 });
@@ -202,6 +256,172 @@ export async function checkRemoteTmuxAvailable(
}
}
/**
* The CLI binary each session mode runs on the remote host. Antigravity's
* binary is `agy` (the mode name is not the command); shell has no CLI to
* probe, so it is absent.
*/
const REMOTE_CLI_BIN: Partial<Record<SessionMode, string>> = {
claude: 'claude',
opencode: 'opencode',
codex: 'codex',
gemini: 'gemini',
antigravity: 'agy',
};
/**
* Build the SSH command that reads the remote CLI's version (`claude --version`
* on the remote host). The version query is routed through
* `remoteLoginShellCommand` (the SAME `$SHELL -i -l -c` wrapper the real
* launch uses), because agent CLIs live on PATH only after the remote user's
* interactive-login startup files run (see defaultRemoteCommandForMode); a bare
* `claude --version` over ssh exits 127. Connection options come from the
* shared `buildSshConnectionArgs`, so the probe reaches exactly the hosts the
* launch can reach. Returns null for modes with no CLI (shell).
*/
export function buildRemoteCliVersionProbeCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): string | null {
const bin = REMOTE_CLI_BIN[mode];
if (!bin) return null;
return [
...buildSshConnectionArgs(host),
remoteSshTarget(host),
shellescape(remoteLoginShellCommand(`${bin} --version`)),
].join(' ');
}
/**
* Read the CLI version installed ON THE REMOTE HOST. Feeds Session.cliVersion
* for remote sessions: the deterministic local probe deliberately skips them
* (it would report the LOCAL host's claude), and the startup-banner scrape is
* unreliable (newer Claude Code builds print no banner; resumed sessions never
* do), which left cliVersion undefined and silently disabled wheel-forwarding
* to the CLI transcript (residual #154, noted in the #205 analysis). The
* version is parsed as the first semver in stdout, never raw output: an
* interactive-login shell may echo rc-file noise around it. Returns undefined
* on any failure. No-op under VITEST (mirrors checkRemoteTmuxAvailable).
*/
export async function probeRemoteCliVersion(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): Promise<string | undefined> {
if (process.env.VITEST) return undefined;
const command = buildRemoteCliVersionProbeCommand(host, mode);
if (!command) return undefined;
try {
const { stdout } = await execAsync(command, { timeout: 15_000 });
const match = stdout.match(/\d+\.\d+\.\d+/);
return match ? match[0] : undefined;
} catch {
return undefined;
}
}
/**
* COD-105 — build the SSH command that lists `codeman-*` tmux sessions on a
* remote host's canonical `-L codeman` socket.
*
* `list-sessions` exits NON-ZERO with empty output when no sessions exist (and
* the server isn't running), so `2>/dev/null` swallows tmux's "no server
* running" stderr; the caller treats a non-zero exit / empty output as "no
* sessions" rather than an error.
*
* COD-107 — connection options come from the shared `buildSshConnectionArgs`, so
* discovery connects with the SAME port/identity/proxy/jump-host as the launch
* and the tmux prereq probe.
*/
export function buildRemoteListSessionsCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions
): string {
const [ssh, ...connectionArgs] = buildSshConnectionArgs(host);
const parts = [ssh, connectionArgs[0], '-o ConnectTimeout=10', ...connectionArgs.slice(1)];
// The tmux list-sessions invocation is passed as ONE shell-quoted argument so
// the remote login shell runs it verbatim. The `-F` format uses literal `\t`
// separators (tmux expands them); `2>/dev/null` is inside the quoted command.
const remoteCmd =
'tmux -L codeman list-sessions -F "#{session_name}\\t#{session_attached}\\t#{session_created}\\t#{session_windows}" 2>/dev/null';
parts.push(remoteSshTarget(host), shellescape(remoteCmd));
return parts.join(' ');
}
/**
* COD-105 — pure parser for the `tmux list-sessions -F` output emitted by
* `buildRemoteListSessionsCommand`. Factored out so the parse is unit-testable
* without opening a real ssh connection.
*
* - Splits each non-empty line into [name, attached, created, windows] on the
* field separator. IMPORTANT: the remote tmux's `-F "…\t…"` format does NOT
* expand `\t` to a real tab — it emits the LITERAL two-character sequence
* `\t` (verified on aa-desktop / tmux next-3.7). So we split on the literal
* backslash-t sequence; we also tolerate a real tab in case a tmux build
* does expand it. (A real TAB is the regex `\t`; a literal backslash-t is the
* regex `\\t`.)
* - Keeps ONLY sessions whose name starts with `codeman-` (ignores foreign tmux
* sessions that happen to share the socket).
* - Coerces: `attached` → boolean (`'1'`), `created`/`windows` → finite ints.
* - Skips malformed lines (wrong column count or non-numeric created/windows)
* rather than emitting garbage.
*/
export function parseRemoteSessionList(stdout: string): RemoteSessionInfo[] {
const out: RemoteSessionInfo[] = [];
for (const rawLine of stdout.split('\n')) {
const line = rawLine.trim();
if (!line) continue;
// Split on a literal `\t` (backslash + t, what the remote tmux emits) OR a
// real tab character. `/\\t|\t/` = the two-char sequence, or a TAB.
const cols = line.split(/\\t|\t/);
if (cols.length !== 4) continue;
const [name, attachedStr, createdStr, windowsStr] = cols;
if (!name.startsWith('codeman-')) continue;
const created = Number(createdStr);
const windows = Number(windowsStr);
if (!Number.isFinite(created) || !Number.isFinite(windows)) continue;
// COD-106 — `session_attached` is the CLIENT COUNT (not a 0/1 flag); >1 = shared.
const attachedNum = Number(attachedStr.trim());
const attachedClients = Number.isFinite(attachedNum) ? Math.max(0, Math.trunc(attachedNum)) : 0;
out.push({
name,
attached: attachedClients > 0,
attachedClients,
created: Math.trunc(created),
windows: Math.trunc(windows),
});
}
return out;
}
/**
* COD-105 — discover `codeman-*` tmux sessions already running on a remote host
* (created by the remote's own Codeman, another instance, or this one), so the
* operator can attach to one this Codeman didn't launch.
*
* NEVER throws: returns `[]` on unreachable host / no tmux / no sessions
* (`list-sessions` exits non-zero with empty output when there are none).
*
* VITEST guard — like `checkRemoteTmuxAvailable`, returns `[]` under test so a
* real ssh never runs in a request path (which would make route tests hit a
* ~10s timeout). The command construction is covered by
* `buildRemoteListSessionsCommand` and the parse by `parseRemoteSessionList`.
*/
export async function listRemoteCodemanSessions(
remote: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions
): Promise<RemoteSessionInfo[]> {
if (process.env.VITEST) {
return [];
}
const command = buildRemoteListSessionsCommand(remote);
try {
const { stdout } = await execAsync(command, { timeout: 15_000 });
return parseRemoteSessionList(stdout);
} catch {
// Unreachable host, no tmux server, or no sessions (non-zero exit). All map
// to "nothing to attach to" — never surface as an error to the caller.
return [];
}
}
export function remoteDisplayPath(
remote: Pick<SessionRemote, 'username' | 'host' | 'remotePath'> | { username: string; host: string; path: string }
): string {
@@ -218,6 +438,10 @@ export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): Sessi
port: host.port,
remotePath: remoteCase.remotePath,
commands: host.commands,
// COD-105 — the COD-104 launch path creates the remote session, so we own it
// (an explicit kill may propagate a remote kill-session). Discovered+attached
// sessions go through `toAttachedSessionRemote` with `owned: false`.
owned: true,
// COD-107 — carry the advanced SSH options from host config into the session
// so the launch/prereq commands connect the same way the operator configured.
identityFile: host.identityFile,
@@ -226,3 +450,38 @@ export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): Sessi
extraSshOptions: host.extraSshOptions,
};
}
/**
* COD-105 — build a NON-owned `SessionRemote` for ATTACHING to a `codeman-*`
* session already running on a remote host (discovered via
* `listRemoteCodemanSessions`). The resulting session's pane runs
* `tmux -L codeman attach -t <remoteSessionName>` (see
* `buildRemoteAttachCommand`), and because we did NOT create the remote session,
* `owned: false` means closing the tab DETACHES rather than killing it.
*
* `remotePath` is informational here (the attached remote session keeps its own
* cwd); we record the host's nominal path so display helpers still show
* `user@host:path`.
*/
export function toAttachedSessionRemote(
host: RemoteHost,
remoteSessionName: string,
remotePath: string
): SessionRemote {
return {
hostId: host.id,
label: host.label,
host: host.host,
username: host.username,
port: host.port,
remotePath,
commands: host.commands,
// Discovered + attached — another Codeman created it. Detach-not-kill.
owned: false,
remoteSessionName,
identityFile: host.identityFile,
socksProxy: host.socksProxy,
jumpHost: host.jumpHost,
extraSshOptions: host.extraSshOptions,
};
}
+184
View File
@@ -0,0 +1,184 @@
/**
* @fileoverview Pure logic for the remote-session auto-reconnect watcher (COD-108).
*
* COD-104 made remote tmux sessions durable + idempotently reattachable, but a
* reconnect only fired at explicit trigger points. COD-108 adds a continuous
* watcher (in `TmuxManager`) that detects a dead remote pane and emits
* `remoteSessionDropped`; `SessionManager`/server then reassembles the respawn
* options and reattaches (re-running the idempotent remote command).
*
* This module holds the SIDE-EFFECT-FREE pieces so they can be unit-tested
* without real tmux:
* - the bounded exponential **backoff schedule** (attempt → delay, capped),
* - the per-session **reconnect state** shape,
* - the **eligibility decision** (`decideReconnect`) given a session + its
* reconnect state + the current time + the guard set.
*
* The watcher in `tmux-manager.ts` owns the live `isPaneDead` probe and the
* timers; everything here is pure and deterministic (time is injected).
*
* @module remote-reconnect
*/
/**
* Bounded exponential backoff delays (ms) between reconnect attempts.
* Attempt N (1-based) waits `BACKOFF_SCHEDULE_MS[N-1]` from the previous emit
* before the next emit is eligible. After the last entry the session is
* considered `reconnect-exhausted` and the watcher stops emitting for it.
*
* 5s, 15s, 45s, 2m, 5m, 5m → ~6 attempts spanning ~13 minutes.
*/
export const BACKOFF_SCHEDULE_MS: readonly number[] = [5_000, 15_000, 45_000, 120_000, 300_000, 300_000];
/** Maximum number of reconnect attempts before exhaustion. */
export const MAX_RECONNECT_ATTEMPTS = BACKOFF_SCHEDULE_MS.length;
/**
* Delay (ms) to wait AFTER emitting attempt `attempt` (1-based) before the next
* attempt is eligible. `attempt <= 0` returns the first delay; an attempt at or
* beyond the cap returns the last delay (callers should check exhaustion via
* {@link isExhausted} rather than relying on this for the stop decision).
*
* Pure — no clock, no I/O.
*/
export function reconnectDelayForAttempt(attempt: number): number {
if (!Number.isFinite(attempt) || attempt <= 1) return BACKOFF_SCHEDULE_MS[0];
const idx = Math.min(Math.floor(attempt) - 1, BACKOFF_SCHEDULE_MS.length - 1);
return BACKOFF_SCHEDULE_MS[idx];
}
/** Whether `attempts` reconnect emits have reached/exceeded the cap. Pure. */
export function isExhausted(attempts: number): boolean {
return attempts >= MAX_RECONNECT_ATTEMPTS;
}
/**
* Per-session reconnect bookkeeping held by the watcher. All time values are
* epoch ms. `inFlight` guards against stacking respawns when a tick fires while
* a previous reattach is still running. `exhaustedEmitted` ensures the
* `remoteReconnectExhausted` event fires at most once per session.
*/
export interface RemoteReconnectState {
/** Number of `remoteSessionDropped` emits so far (advances per emit). */
attempts: number;
/** Earliest time (epoch ms) the next emit is eligible. 0 = eligible now. */
nextEligibleAt: number;
/** A reattach triggered by a prior emit is currently running. */
inFlight: boolean;
/** Cap reached — stop auto-retrying for this session. */
exhausted: boolean;
/** The `remoteReconnectExhausted` SSE event has already been emitted. */
exhaustedEmitted: boolean;
}
/** A fresh reconnect state (no attempts, immediately eligible). Pure. */
export function freshReconnectState(): RemoteReconnectState {
return { attempts: 0, nextEligibleAt: 0, inFlight: false, exhausted: false, exhaustedEmitted: false };
}
/**
* Advance the backoff after an emit at time `now`. Increments `attempts` and
* schedules `nextEligibleAt = now + delay`. Returns a NEW state object (does
* not mutate the input). Pure.
*
* NOTE: this does NOT set `exhausted`. Exhaustion is a decision the watcher
* makes on the FOLLOWING tick (via {@link decideReconnect} → `exhaust`), so the
* `remoteReconnectExhausted` event fires exactly once after the final attempt's
* backoff window elapses — not pre-emptively on the last emit.
*/
export function advanceBackoff(state: RemoteReconnectState, now: number): RemoteReconnectState {
const attempts = state.attempts + 1;
const delay = reconnectDelayForAttempt(attempts);
return {
...state,
attempts,
nextEligibleAt: now + delay,
};
}
/** Reset after a successful reattach — back to a fresh, eligible state. Pure. */
export function resetReconnectState(): RemoteReconnectState {
return freshReconnectState();
}
/** Minimal session view the decision needs (avoids importing MuxSession here). */
export interface ReconnectSessionView {
sessionId: string;
/** Truthy when this is a remote (SSH-wrapped) session. */
isRemote: boolean;
/** Result of `isPaneDead(muxName)` for this session. */
paneDead: boolean;
}
/**
* Decision outcomes for a single watcher tick on one session.
* - `emit` → emit `remoteSessionDropped { sessionId, attempt }`, then
* advance backoff (attempt = the returned `attempt`).
* - `exhaust` → cap reached this tick; emit `remoteReconnectExhausted` once.
* - `skip` → do nothing (not remote / pane alive / guarded / in-flight /
* not yet due / already exhausted).
*/
export type ReconnectAction =
| { kind: 'emit'; attempt: number }
| { kind: 'exhaust' }
| { kind: 'skip'; reason: ReconnectSkipReason };
export type ReconnectSkipReason =
| 'not-remote'
| 'pane-alive'
| 'guarded'
| 'in-flight'
| 'not-due'
| 'exhausted'
| 'disabled';
export interface DecideReconnectInput {
session: ReconnectSessionView;
state: RemoteReconnectState | undefined;
/** Session is in the intentional-teardown guard set (killed/detached/stopping). */
guarded: boolean;
/** Kill-switch: `remoteAutoReconnect` setting. When false, never reconnect. */
enabled: boolean;
now: number;
}
/**
* PURE eligibility decision for one session on one tick. No clock, no I/O — all
* inputs are passed in. The watcher translates the result into emits + state
* transitions.
*
* Order of guards (most-decisive first):
* 1. kill-switch off → skip:disabled
* 2. not a remote session → skip:not-remote
* 3. pane is alive → skip:pane-alive
* 4. intentional teardown guard → skip:guarded (NEVER revive a killed tab)
* 5. a reattach already running → skip:in-flight (no stacked respawns)
* 6. already exhausted → skip:exhausted (one exhaust emit, then quiet)
* 7. cap reached this tick → exhaust
* 8. not yet due (backoff) → skip:not-due
* 9. otherwise → emit (attempt = attempts + 1)
*/
export function decideReconnect(input: DecideReconnectInput): ReconnectAction {
const { session, state, guarded, enabled, now } = input;
if (!enabled) return { kind: 'skip', reason: 'disabled' };
if (!session.isRemote) return { kind: 'skip', reason: 'not-remote' };
if (!session.paneDead) return { kind: 'skip', reason: 'pane-alive' };
// Intentional kill / detach must NEVER be auto-revived.
if (guarded) return { kind: 'skip', reason: 'guarded' };
const s = state ?? freshReconnectState();
// Only one reconnect in flight per session — don't stack respawns.
if (s.inFlight) return { kind: 'skip', reason: 'in-flight' };
if (s.exhausted) return { kind: 'skip', reason: 'exhausted' };
// Cap reached: surface exhaustion once, then go quiet.
if (isExhausted(s.attempts)) return { kind: 'exhaust' };
// Backoff gate — only emit when due.
if (now < s.nextEligibleAt) return { kind: 'skip', reason: 'not-due' };
return { kind: 'emit', attempt: s.attempts + 1 };
}
+401
View File
@@ -0,0 +1,401 @@
/**
* @fileoverview `codeman service install|uninstall|status`: write and load the
* systemd user unit (Linux) or LaunchAgent (macOS) that supervises `codeman web`.
*
* This is the "always running" half of issue #231, next to the "detached right
* now" half in daemon-control.ts. `install.sh` already does this for people who
* install with the one-liner; this exists for `npm i -g aicodeman` users, who
* otherwise have to hand-write a plist.
*
* Two details are load-bearing and easy to get wrong by hand:
*
* - **PATH.** launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` and systemd's
* user manager is nearly as bare, so a Homebrew or nvm `node`, `tmux` or
* `claude` is simply not found and sessions fail in a way that reads as a
* Codeman bug. The unit therefore carries the PATH of the shell that ran the
* install, with the running node's own directory in front.
* - **The job name.** It is the one `install.sh` and the self-updater already use
* (config/service-names.ts), so re-running install.sh later updates this unit
* instead of supervising a second copy of the server.
*
* Secrets are deliberately NOT written here. `CODEMAN_PASSWORD` in the installing
* shell is not copied into the unit; the caller is told where to add it instead,
* because a unit file is long-lived, world-readable by default, and gets copied
* into bug reports.
*
* The file writers are pure string builders so they can be unit-tested without
* touching launchctl/systemctl.
*
* @module service-installer
*/
import { execFileSync } from 'node:child_process';
import { existsSync, mkdirSync, unlinkSync, writeFileSync } from 'node:fs';
import { homedir, userInfo } from 'node:os';
import { dirname, join } from 'node:path';
import { LAUNCHD_LABEL, SYSTEMD_UNIT } from './config/service-names.js';
import { CODEMAN_INSTANCE } from './config/instance.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import {
buildBaseUrl,
buildStatusUrl,
buildWebArgs,
logFilePath,
probeServer,
type WebLaunchOptions,
} from './daemon-control.js';
export type ServiceKind = 'launchd' | 'systemd';
/** Everything a unit file needs, resolved from the environment by the caller. */
export interface ServicePlan {
kind: ServiceKind;
/** systemd unit filename or launchd label. */
name: string;
nodePath: string;
/** Runner flags carried over from the current process (tsx loader in dev). */
execArgv: string[];
scriptPath: string;
args: string[];
env: Record<string, string>;
logPath: string;
workingDir: string;
}
export interface ServiceActionResult {
ok: boolean;
message: string;
/** Path of the unit/plist that was written or removed. */
unitPath?: string;
warnings?: string[];
}
export interface ServiceStatusResult {
kind: ServiceKind | null;
name: string;
unitPath: string;
installed: boolean;
loaded: boolean;
responding: boolean;
version?: string;
url: string;
}
/** Directories worth having on PATH even when the installing shell lacked them. */
const FALLBACK_PATH_DIRS = ['/opt/homebrew/bin', '/usr/local/bin', '/usr/bin', '/bin', '/usr/sbin', '/sbin'];
// ─────────────────────────────────────────────────────────────────────────────
// Pure builders
// ─────────────────────────────────────────────────────────────────────────────
/** XML text escaping for plist `<string>` values. */
export function xmlEscape(value: string): string {
return value
.replace(/&/g, '&amp;')
.replace(/</g, '&lt;')
.replace(/>/g, '&gt;')
.replace(/"/g, '&quot;')
.replace(/'/g, '&apos;');
}
/**
* PATH for the supervised process: the running node's directory first (so an nvm
* or Homebrew node is used rather than whatever the supervisor finds), then the
* installing shell's PATH, then the fallbacks that are still missing.
*
* `node_modules/.bin` entries are dropped. npm and npx inject those for the
* lifetime of one command, and baking a project's local bin dir into a unit file
* that outlives the checkout is how a service ends up running a binary the
* operator deleted months ago.
*/
export function buildServicePath(nodeDir: string, currentPath: string, home: string): string {
const seen = new Set<string>();
const ordered: string[] = [];
const push = (dir: string) => {
const trimmed = dir.trim();
if (!trimmed || seen.has(trimmed)) return;
if (/(^|\/)node_modules\/\.bin\/?$/.test(trimmed)) return;
seen.add(trimmed);
ordered.push(trimmed);
};
push(nodeDir);
for (const dir of currentPath.split(':')) push(dir);
push(join(home, '.local', 'bin'));
for (const dir of FALLBACK_PATH_DIRS) push(dir);
return ordered.join(':');
}
/** Environment written into the unit. Never includes secrets (see module docs). */
export function buildServiceEnv(
nodeDir: string,
currentPath: string,
home: string,
lang?: string
): Record<string, string> {
const env: Record<string, string> = {
PATH: buildServicePath(nodeDir, currentPath, home),
HOME: home,
LANG: lang || 'en_US.UTF-8',
};
if (CODEMAN_INSTANCE) env.CODEMAN_INSTANCE = CODEMAN_INSTANCE;
return env;
}
/** systemd accepts double-quoted values; escape the two characters that matter. */
export function systemdQuote(value: string): string {
return `"${value.replace(/\\/g, '\\\\').replace(/"/g, '\\"')}"`;
}
export function buildLaunchAgentPlist(plan: ServicePlan): string {
const programArguments = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
.map((arg) => ` <string>${xmlEscape(arg)}</string>`)
.join('\n');
const environment = Object.entries(plan.env)
.map(([key, value]) => ` <key>${xmlEscape(key)}</key>\n <string>${xmlEscape(value)}</string>`)
.join('\n');
return `<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>${xmlEscape(plan.name)}</string>
<key>ProgramArguments</key>
<array>
${programArguments}
</array>
<key>EnvironmentVariables</key>
<dict>
${environment}
</dict>
<key>WorkingDirectory</key>
<string>${xmlEscape(plan.workingDir)}</string>
<key>RunAtLoad</key>
<true/>
<key>KeepAlive</key>
<true/>
<key>ThrottleInterval</key>
<integer>10</integer>
<key>StandardOutPath</key>
<string>${xmlEscape(plan.logPath)}</string>
<key>StandardErrorPath</key>
<string>${xmlEscape(plan.logPath)}</string>
</dict>
</plist>
`;
}
export function buildSystemdUnit(plan: ServicePlan): string {
const execStart = [plan.nodePath, ...plan.execArgv, plan.scriptPath, ...plan.args]
.map((arg) => (/[\s"'\\]/.test(arg) ? systemdQuote(arg) : arg))
.join(' ');
const environment = Object.entries(plan.env)
.map(([key, value]) => `Environment=${systemdQuote(`${key}=${value}`)}`)
.join('\n');
return `[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
WorkingDirectory=${plan.workingDir}
ExecStart=${execStart}
Restart=always
RestartSec=10
# Agents keep running in tmux when the server restarts, so only signal the
# server itself.
KillMode=process
${environment}
StandardOutput=journal
StandardError=journal
SyslogIdentifier=codeman
LimitNOFILE=65536
[Install]
WantedBy=default.target
`;
}
// ─────────────────────────────────────────────────────────────────────────────
// Environment resolution
// ─────────────────────────────────────────────────────────────────────────────
export function detectServiceKind(): ServiceKind | null {
if (process.platform === 'darwin') return 'launchd';
if (process.platform === 'linux') return 'systemd';
return null;
}
export function unitPathFor(kind: ServiceKind): string {
return kind === 'launchd'
? join(homedir(), 'Library', 'LaunchAgents', `${LAUNCHD_LABEL}.plist`)
: join(homedir(), '.config', 'systemd', 'user', SYSTEMD_UNIT);
}
function entryScript(): string {
const script = process.argv[1];
if (!script) throw new Error('cannot determine the codeman entry script to supervise');
return script;
}
/** Resolve a full plan from the current process and the requested web options. */
export function resolveServicePlan(kind: ServiceKind, options: WebLaunchOptions): ServicePlan {
const home = homedir();
return {
kind,
name: kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT,
nodePath: process.execPath,
execArgv: [...process.execArgv],
scriptPath: entryScript(),
args: buildWebArgs(options),
env: buildServiceEnv(dirname(process.execPath), process.env.PATH || '', home, process.env.LANG),
logPath: logFilePath(),
workingDir: home,
};
}
function run(command: string, args: string[]): { ok: boolean; output: string } {
try {
const output = execFileSync(command, args, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'pipe'],
});
return { ok: true, output: output.trim() };
} catch (err) {
const e = err as { stderr?: Buffer | string; message?: string };
const stderr = typeof e.stderr === 'string' ? e.stderr : e.stderr?.toString('utf-8');
return { ok: false, output: (stderr || e.message || '').trim() };
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Install / uninstall / status
// ─────────────────────────────────────────────────────────────────────────────
/**
* Write the unit, load it, and confirm the server actually answers before
* reporting success. `launchctl load` and `systemctl enable` are both quiet about
* a job that starts and immediately dies, which is the whole reason install.sh
* verifies too.
*/
export async function installService(options: WebLaunchOptions): Promise<ServiceActionResult> {
const kind = detectServiceKind();
if (!kind) {
return { ok: false, message: `no supported supervisor on ${process.platform}; use \`codeman web -d\` instead` };
}
const plan = resolveServicePlan(kind, options);
const unitPath = unitPathFor(kind);
const warnings: string[] = [];
mkdirSync(dirname(unitPath), { recursive: true });
if (kind === 'launchd') {
const uid = process.getuid?.() ?? 0;
// Unload any previous copy first, otherwise bootstrap fails with "service
// already loaded" and leaves the OLD job running against the NEW file.
run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
writeFileSync(unitPath, buildLaunchAgentPlist(plan), { encoding: 'utf-8', mode: 0o600 });
const bootstrap = run('launchctl', ['bootstrap', `gui/${uid}`, unitPath]);
if (!bootstrap.ok) {
const legacy = run('launchctl', ['load', unitPath]);
if (!legacy.ok) {
return {
ok: false,
unitPath,
message: `wrote ${unitPath} but launchctl refused to load it: ${bootstrap.output}`,
};
}
}
} else {
writeFileSync(unitPath, buildSystemdUnit(plan), { encoding: 'utf-8', mode: 0o600 });
const reload = run('systemctl', ['--user', 'daemon-reload']);
if (!reload.ok) {
return {
ok: false,
unitPath,
message: `wrote ${unitPath} but \`systemctl --user daemon-reload\` failed: ${reload.output}`,
};
}
const enable = run('systemctl', ['--user', 'enable', '--now', SYSTEMD_UNIT]);
if (!enable.ok) {
return { ok: false, unitPath, message: `wrote ${unitPath} but enabling it failed: ${enable.output}` };
}
// Without lingering the unit stops at logout, which is exactly what someone
// installing a service does not want. Best effort: it needs polkit rights.
const linger = run('loginctl', ['enable-linger', userInfo().username]);
if (!linger.ok) {
warnings.push(
`could not enable lingering, so the service will stop when you log out. Run: sudo loginctl enable-linger ${userInfo().username}`
);
}
}
const url = buildBaseUrl(options);
const statusUrl = buildStatusUrl(options);
const deadline = Date.now() + 30_000;
while (Date.now() < deadline) {
const probe = await probeServer(statusUrl, 1000);
if (probe.up) {
return { ok: true, unitPath, warnings, message: `service installed and responding at ${url}` };
}
await new Promise((resolve) => setTimeout(resolve, 500));
}
const hint =
kind === 'launchd' ? `tail -20 ${plan.logPath}` : `journalctl --user -u ${SYSTEMD_UNIT} -n 20 --no-pager`;
return {
ok: false,
unitPath,
warnings,
message: `wrote and loaded ${unitPath}, but nothing answered ${url} within 30s. Check: ${hint}`,
};
}
export function uninstallService(): ServiceActionResult {
const kind = detectServiceKind();
if (!kind) return { ok: false, message: `no supported supervisor on ${process.platform}` };
const unitPath = unitPathFor(kind);
if (!existsSync(unitPath)) {
return { ok: false, unitPath, message: `no service installed at ${unitPath}` };
}
if (kind === 'launchd') {
const uid = process.getuid?.() ?? 0;
const bootout = run('launchctl', ['bootout', `gui/${uid}/${LAUNCHD_LABEL}`]);
if (!bootout.ok) run('launchctl', ['unload', unitPath]);
} else {
run('systemctl', ['--user', 'disable', '--now', SYSTEMD_UNIT]);
}
try {
unlinkSync(unitPath);
} catch (err) {
return { ok: false, unitPath, message: `stopped the service but could not remove ${unitPath}: ${String(err)}` };
}
if (kind === 'systemd') run('systemctl', ['--user', 'daemon-reload']);
return { ok: true, unitPath, message: `service stopped and ${unitPath} removed. Your tmux sessions are untouched.` };
}
export async function serviceStatus(options: WebLaunchOptions): Promise<ServiceStatusResult> {
const kind = detectServiceKind();
const url = buildBaseUrl(options);
if (!kind) {
return { kind: null, name: '', unitPath: '', installed: false, loaded: false, responding: false, url };
}
const unitPath = unitPathFor(kind);
const name = kind === 'launchd' ? LAUNCHD_LABEL : SYSTEMD_UNIT;
const installed = existsSync(unitPath);
const loaded =
kind === 'launchd'
? run('launchctl', ['list', LAUNCHD_LABEL]).ok
: run('systemctl', ['--user', 'is-active', SYSTEMD_UNIT]).output === 'active';
const probe = await probeServer(buildStatusUrl(options), 2000);
return { kind, name, unitPath, installed, loaded, responding: probe.up, version: probe.version, url };
}
+105 -3
View File
@@ -28,9 +28,15 @@ export type UnifiedSessionItem = {
lastActivityAt?: number;
claudeSessionId?: string;
firstPrompt?: string;
/** Most recent user prompt from the transcript (COD-145), parallel to firstPrompt. */
lastPrompt?: string;
sizeBytes?: number;
projectKey?: string;
remote?: boolean;
/** Pinned to the top of the session manager list (COD-139). */
pinned?: boolean;
/** When the session was pinned (epoch ms) — orders the pinned group desc. */
pinnedAt?: number;
sources: string[];
stats?: { memoryMB: number; cpuPercent: number };
};
@@ -46,6 +52,8 @@ export type LiveSessionInput = {
createdAt?: number;
lastActivityAt?: number;
claudeSessionId?: string;
pinned?: boolean;
pinnedAt?: number;
};
/** Persisted session view (subset of `SessionState`). */
@@ -59,6 +67,8 @@ export type PersistedSessionInput = {
lastActivityAt?: number;
/** Claude conversation ID this session resumes (`SessionState.resumeSessionId`). */
claudeSessionId?: string;
pinned?: boolean;
pinnedAt?: number;
};
/** Lifecycle audit-log view. Entries are expected NEWEST-first (the order `SessionLifecycleLog.query()` returns). */
@@ -77,6 +87,8 @@ export type HistoryInput = {
sizeBytes: number;
lastModified: string;
firstPrompt?: string;
/** Most recent user prompt from the transcript (COD-145). */
lastPrompt?: string;
projectKey?: string;
};
@@ -149,6 +161,7 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
overwrite(item, 'workingDir', h.workingDir);
overwrite(item, 'sizeBytes', h.sizeBytes);
overwrite(item, 'firstPrompt', h.firstPrompt);
overwrite(item, 'lastPrompt', h.lastPrompt);
overwrite(item, 'projectKey', h.projectKey);
const ms = Date.parse(h.lastModified);
if (!Number.isNaN(ms) && item.lastActivityAt === undefined) item.lastActivityAt = ms;
@@ -175,6 +188,8 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
overwrite(item, 'workingDir', p.workingDir);
overwrite(item, 'createdAt', p.createdAt);
overwrite(item, 'lastActivityAt', p.lastActivityAt);
overwrite(item, 'pinned', p.pinned);
overwrite(item, 'pinnedAt', p.pinnedAt);
}
// 4) live (highest precedence)
@@ -189,6 +204,8 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
overwrite(item, 'createdAt', v.createdAt);
overwrite(item, 'lastActivityAt', v.lastActivityAt);
overwrite(item, 'claudeSessionId', v.claudeSessionId);
overwrite(item, 'pinned', v.pinned);
overwrite(item, 'pinnedAt', v.pinnedAt);
}
// 5) mux stats + remote flag (create item if mux-only)
@@ -200,6 +217,75 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
if (m.remote !== undefined) item.remote = m.remote;
}
// firstPrompt backfill (COD-140): the only source that sets firstPrompt is the
// transcript-history view, keyed by the Claude transcript file's UUID. A live/persisted
// row keyed by its Codeman id only inherits firstPrompt when that id happens to equal an
// on-disk transcript UUID. When it doesn't (stale/wrong claudeSessionId, post-/clear new
// uuid, resumed/attached/worktree session, transcript not yet flushed), the row shows
// "(no prompt captured)" even though a real transcript for that working dir exists under a
// different UUID. Backfill from the already-passed history: first try the claudeSessionId
// join, then the newest transcript in the same workingDir. Never overwrite a non-empty
// firstPrompt (so rows keyed to their own transcript are untouched).
//
// The workingDir guess is a last resort and MUST be skipped for any item that
// already has its own 'history' entry (step 1 above already gave it a real,
// direct scan of its own transcript). Without this guard, a history row whose
// OWN extraction genuinely failed (oversized first message, etc.) silently
// inherited the newest OTHER session's opening line from the same directory —
// not a blank, but actively wrong: old sessions displayed today's conversation
// as if it were their own. A row with no 'history' source at all (its
// transcript hasn't been linked/scanned under its own id yet) has no such
// direct attempt to prefer, so the guess remains a reasonable stand-in there.
const firstPromptByUuid = new Map<string, string>();
const firstPromptByWorkingDir = new Map<string, { prompt: string; ms: number }>();
// COD-145: lastPrompt rides the same backfill (build parallel indexes; never overwrite).
const lastPromptByUuid = new Map<string, string>();
const lastPromptByWorkingDir = new Map<string, { prompt: string; ms: number }>();
for (const h of sources.history ?? []) {
const ms = Date.parse(h.lastModified);
const ts = Number.isNaN(ms) ? -Infinity : ms;
if (h.firstPrompt) {
firstPromptByUuid.set(h.sessionId, h.firstPrompt);
if (h.workingDir) {
const existing = firstPromptByWorkingDir.get(h.workingDir);
if (!existing || ts > existing.ms) {
firstPromptByWorkingDir.set(h.workingDir, { prompt: h.firstPrompt, ms: ts });
}
}
}
if (h.lastPrompt) {
lastPromptByUuid.set(h.sessionId, h.lastPrompt);
if (h.workingDir) {
const existing = lastPromptByWorkingDir.get(h.workingDir);
if (!existing || ts > existing.ms) {
lastPromptByWorkingDir.set(h.workingDir, { prompt: h.lastPrompt, ms: ts });
}
}
}
}
for (const item of map.values()) {
const hasOwnHistoryEntry = item.sources.includes('history');
if (!item.firstPrompt) {
// never overwrite an existing non-empty prompt
const byUuid = item.claudeSessionId ? firstPromptByUuid.get(item.claudeSessionId) : undefined;
if (byUuid) {
item.firstPrompt = byUuid;
} else if (item.workingDir && !hasOwnHistoryEntry) {
const byDir = firstPromptByWorkingDir.get(item.workingDir);
if (byDir) item.firstPrompt = byDir.prompt;
}
}
if (!item.lastPrompt) {
const byUuid = item.claudeSessionId ? lastPromptByUuid.get(item.claudeSessionId) : undefined;
if (byUuid) {
item.lastPrompt = byUuid;
} else if (item.workingDir && !hasOwnHistoryEntry) {
const byDir = lastPromptByWorkingDir.get(item.workingDir);
if (byDir) item.lastPrompt = byDir.prompt;
}
}
}
// Meaningfulness floor: keep real rows, drop bare lifecycle/mux-only noise.
const kept: UnifiedSessionItem[] = [];
for (const item of map.values()) {
@@ -211,8 +297,24 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
if (isReal) kept.push(item);
}
// Stable sort: lastActivityAt desc (undefined last), createdAt desc, sessionId asc.
// Stable sort (COD-139): pinned group first (pinnedAt desc, most-recently-pinned
// first), then unpinned by lastActivityAt desc (undefined last), createdAt desc,
// sessionId asc.
kept.sort((a, b) => {
const pa = a.pinned === true;
const pb = b.pinned === true;
if (pa !== pb) return pa ? -1 : 1; // pinned floats above unpinned
if (pa && pb) {
// Both pinned: most-recently-pinned first (undefined pinnedAt sorts last).
const ta = a.pinnedAt;
const tb = b.pinnedAt;
if (ta !== tb) {
if (ta === undefined) return 1;
if (tb === undefined) return -1;
return tb - ta;
}
// tie-break falls through to the activity/createdAt/id rules below.
}
const la = a.lastActivityAt;
const lb = b.lastActivityAt;
if (la !== lb) {
@@ -234,7 +336,7 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
}
/**
* Case-insensitive substring filter (name + firstPrompt + workingDir + sessionId)
* Case-insensitive substring filter (name + firstPrompt + lastPrompt + workingDir + sessionId)
* with offset/limit paging. `total` is the filtered count BEFORE paging.
*/
export function filterAndPaginate(
@@ -244,7 +346,7 @@ export function filterAndPaginate(
const q = (opts.q ?? '').trim().toLowerCase();
const filtered = q
? items.filter((it) => {
const hay = [it.name, it.firstPrompt, it.workingDir, it.sessionId]
const hay = [it.name, it.firstPrompt, it.lastPrompt, it.workingDir, it.sessionId]
.filter((v): v is string => typeof v === 'string')
.join(' ')
.toLowerCase();
+18 -4
View File
@@ -21,6 +21,8 @@ function buildPermissionArgs(claudeMode: ClaudeMode, allowedTools?: string): str
switch (claudeMode) {
case 'dangerously-skip-permissions':
return ['--dangerously-skip-permissions'];
case 'auto':
return ['--permission-mode', 'auto'];
case 'allowedTools':
if (allowedTools) {
return ['--allowedTools', allowedTools];
@@ -80,8 +82,16 @@ export function buildInteractiveArgs(
* @param model - Optional model override
* @returns Array of CLI arguments
*/
export function buildPromptArgs(prompt: string, model?: string): string[] {
const args = ['-p', '--verbose', '--dangerously-skip-permissions', '--output-format', 'stream-json'];
export function buildPromptArgs(
prompt: string,
model?: string,
claudeMode: ClaudeMode = 'dangerously-skip-permissions',
allowedTools?: string
): string[] {
// Respect the session's permission mode instead of always skipping, so a
// multi-user non-granted user's one-shot runs classifier-guarded (auto) rather
// than with full bypass. Defaults to skip-permissions (unchanged single-user).
const args = ['-p', '--verbose', ...buildPermissionArgs(claudeMode, allowedTools), '--output-format', 'stream-json'];
if (model) {
args.push('--model', model);
}
@@ -111,7 +121,10 @@ export function buildClaudeEnv(sessionId: string): Record<string, string | undef
// Inform Claude it's running within Codeman (helps prevent self-termination)
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// CODEMAN_API_URL rides in via the process.env spread when the server has
// stamped it (WebServer.start()); no fallback: a hardcoded one was the wrong
// scheme on HTTPS installs, and a present-with-undefined key would serialize
// as the literal "CODEMAN_API_URL=undefined" (COD-115).
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
@@ -169,7 +182,8 @@ export function buildShellEnv(sessionId: string): Record<string, string | undefi
TERM: 'xterm-256color',
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// CODEMAN_API_URL rides in via the process.env spread when set; no fallback
// (same reasoning as buildClaudeEnv above).
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
+68
View File
@@ -0,0 +1,68 @@
/**
* @fileoverview Pure helpers for the global session tab-order (COD-131).
*
* Tab order (drag-and-drop reorder + Ctrl+Shift+{/}) is persisted server-side
* so it follows the user across devices. The server is authoritative; the
* browser's localStorage (`codeman-session-order`) is the offline fallback.
*
* These helpers are pure (no IO) so they can be unit-tested in isolation and
* reused by both the PUT /api/session-order route and the StateStore accessor.
*
* - `normalizeSessionOrder` coerces arbitrary input into a clean string[]
* (non-empty strings only, deduped with first occurrence winning).
* - `mergeSessionOrder` lets the pushing device's order win, while preserving
* any server-only ids the pushing device didn't know about — they fall to the
* END in their existing relative order, never dropped.
*/
/**
* Coerce arbitrary input into a clean ordered list of session ids:
* keep only non-empty strings and dedup (first occurrence wins).
*
* @param order - unknown input (expected to be a string[], but defensive)
* @returns a normalized string[] (empty array for non-array / all-junk input)
*/
export function normalizeSessionOrder(order: unknown): string[] {
if (!Array.isArray(order)) {
return [];
}
const seen = new Set<string>();
const result: string[] = [];
for (const entry of order) {
if (typeof entry !== 'string' || entry.length === 0) {
continue;
}
if (seen.has(entry)) {
continue;
}
seen.add(entry);
result.push(entry);
}
return result;
}
/**
* Merge an incoming order from a pushing device with the existing server order.
*
* The incoming order wins; any ids present in `existing` but NOT in `incoming`
* are appended at the END, preserving their relative order. This is the
* "server-only ids the pushing device didn't know about fall to the end, never
* dropped" rule.
*
* Both arguments are normalized first, so callers may pass raw input safely.
*
* @param incoming - the order the pushing device wants
* @param existing - the current server-side order
* @returns the merged, normalized order
*/
export function mergeSessionOrder(incoming: string[], existing: string[]): string[] {
const normalizedIncoming = normalizeSessionOrder(incoming);
const incomingSet = new Set(normalizedIncoming);
const merged = [...normalizedIncoming];
for (const id of normalizeSessionOrder(existing)) {
if (!incomingSet.has(id)) {
merged.push(id);
}
}
return merged;
}
+342 -92
View File
@@ -49,10 +49,12 @@ import {
type CodexConfig,
type EffortLevel,
type GeminiConfig,
type AntigravityConfig,
type SessionRemote,
type SessionDocker,
} from './types.js';
import { probeDockerCliVersion } from './docker-hosts.js';
import { probeRemoteCliVersion } from './remote-hosts.js';
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
@@ -65,6 +67,9 @@ import {
MAX_SESSION_TOKENS,
execPattern,
getClaudeCliVersion,
getClaudeBinaryPath,
spawnPtyWithHelperRepair,
resolveLocalShell,
} from './utils/index.js';
import {
MAX_TERMINAL_BUFFER_SIZE,
@@ -140,7 +145,7 @@ const NEWLINE_SPLIT_PATTERN = /\r?\n/;
/** True for external-CLI run modes (non-Claude) that use their own TUI and output format. */
export function isExternalCliMode(mode: SessionMode): boolean {
return mode === 'opencode' || mode === 'codex' || mode === 'gemini';
return mode === 'opencode' || mode === 'codex' || mode === 'gemini' || mode === 'antigravity';
}
function getModeLabel(mode: SessionMode): string {
@@ -151,6 +156,8 @@ function getModeLabel(mode: SessionMode): string {
return 'Codex';
case 'gemini':
return 'Gemini';
case 'antigravity':
return 'Antigravity';
case 'shell':
return 'Shell';
case 'claude':
@@ -174,14 +181,54 @@ export function isAltScreenStripMode(mode: SessionMode): boolean {
return mode === 'codex' || mode === 'claude' || mode === 'gemini';
}
/**
* Modes that need the NARROW strip: alt-screen toggles only, leaving `\x1b[3J`
* and the mouse-tracking DECSETs alone. Applies to every mode `isAltScreenStripMode`
* excludes, but ONLY when the session is tmux-backed (`useMux`).
*
* The bug (issue #205): the tmux CLIENT emits `smcup` (`\x1b[?1049h`) as its first
* bytes on attach, before any program has run. Unstripped, xterm.js parks in the
* alternate buffer for the whole session, where `baseY` is pinned at 0 (no
* scrollback to reach, so touch scrolling is a no-op) and xterm's own wheel handler
* translates the wheel into `\x1bOA`/`\x1bOB` cursor keys — which readline receives
* as shell history navigation. Both reported symptoms, one sequence.
*
* Why this is safe under tmux, despite the old "shell must keep the alt screen for
* vim/less/htop" reasoning: tmux is a full terminal emulator and NEVER forwards a
* pane's alt-screen toggles to its client, it repaints instead. Captured from a real
* attach, `\x1b[?1049h` appears exactly once (at attach) and vim/less/htop sessions
* inside the pane emit zero. So the only thing stripped here is tmux's own smcup.
*
* Why it is gated on `useMux`: `startShell()`/`startInteractive()` fall back to a
* DIRECT PTY when mux creation fails. There the inner program's `\x1b[?1049h` really
* does reach xterm, and stripping it would break vim/less/htop for real.
*
* Why it is narrower than the full strip: with tmux `mouse off`, a mouse-aware
* program in the pane (htop, vim with `set mouse=a`) still gets its DECSETs passed
* through to the client, so stripping those would break its mouse support. And
* `\x1b[3J` from a user's own `clear` is a deliberate "wipe my scrollback".
*/
export function isMuxAltScreenOnlyStripMode(mode: SessionMode, useMux: boolean): boolean {
return useMux && !isAltScreenStripMode(mode);
}
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
/** PTY fallback geometry when tmux can't be queried (matches pre-#80 hardcoded values). */
const DEFAULT_PTY_COLS = 120;
const DEFAULT_PTY_ROWS = 40;
const TMUX_DISPLAY_TIMEOUT_MS = 2000;
const IS_TEST_MODE = !!process.env.VITEST;
/**
* Echo transport for the test-mode PTY attach. Raw mode disables the tty line
* discipline, so each input byte flows back exactly once and immediately; without
* it, tty echo doubles every line and canonical buffering holds bytes until Enter.
*/
const TEST_PTY_SCRIPT = 'if (process.stdin.isTTY) process.stdin.setRawMode(true); process.stdin.pipe(process.stdout);';
/** Delay before the in-container Claude CLI version probe (lets the container start). */
const DOCKER_CLI_VERSION_PROBE_DELAY_MS = 3000;
/** Delay before the over-ssh Claude CLI version probe (keeps session start off the ssh round-trip). */
const REMOTE_CLI_VERSION_PROBE_DELAY_MS = 3000;
/**
* Ask tmux for the current window geometry of `muxName` so a re-attaching PTY
@@ -342,6 +389,11 @@ export class Session extends EventEmitter {
// Image watcher setting (per-session toggle)
private _imageWatcherEnabled: boolean = false;
// Pin state (COD-139) — pinned sessions float to the top of the session
// manager list, ordered by pinnedAt descending (most-recently-pinned first).
private _pinned: boolean = false;
private _pinnedAt: number | null = null;
// Flicker filter setting (per-session toggle, applied on frontend)
private _flickerFilterEnabled: boolean = false;
@@ -391,6 +443,8 @@ export class Session extends EventEmitter {
private _codexConfig: CodexConfig | undefined;
// Gemini configuration (only for mode === 'gemini')
private _geminiConfig: GeminiConfig | undefined;
// Antigravity configuration (only for mode === 'antigravity')
private _antigravityConfig: AntigravityConfig | undefined;
private _resumeSessionId: string | undefined;
// Ephemeral env overrides (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Exported by tmux
@@ -412,6 +466,10 @@ export class Session extends EventEmitter {
// local tmux + `docker exec`. The container is per-CASE (shared by sibling sessions).
private readonly _docker?: SessionDocker;
// Owning username in multi-user mode (undefined in single-user). Stamped at create
// from req.authUser and round-tripped through recovery like _remote/_docker.
private _owner?: string;
// Session color for visual differentiation
private _color: import('./types.js').SessionColor = 'default';
@@ -473,6 +531,8 @@ export class Session extends EventEmitter {
codexConfig?: CodexConfig;
/** Gemini configuration (only for mode === 'gemini') */
geminiConfig?: GeminiConfig;
/** Antigravity configuration (only for mode === 'antigravity') */
antigravityConfig?: AntigravityConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
/** Extra env vars exported to the CLI at spawn time (no disk persistence) */
@@ -483,10 +543,14 @@ export class Session extends EventEmitter {
tmuxHistoryLimit?: number;
/** Restored per-session attachment history. May include server-private external paths. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/** Restored wall-clock ms of the pane's last Enter (see `lastSubmitAt`). */
lastSubmitAt?: number;
/** Remote execution metadata for sessions launched through SSH inside local tmux. */
remote?: SessionRemote;
/** Docker execution metadata for sessions launched inside a container via local tmux. */
docker?: SessionDocker;
/** Owning username (multi-user mode); undefined in single-user. */
owner?: string;
}
) {
super();
@@ -507,6 +571,12 @@ export class Session extends EventEmitter {
this._lastActivityAt = this.createdAt;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = config.resumeSessionId || this.id;
// Restored from state.json on boot recovery. start() resets _claudeSessionId
// to the launch id even when re-attaching to a mux session whose CLI has
// moved on (a `/clear` before the restart), so this anchor is what lets the
// response viewer re-derive the live conversation without waiting for the
// user to type again.
this._lastSubmitAt = config.lastSubmitAt ?? 0;
this._mux = config.mux || null;
this._useMux = config.useMux ?? (this._mux !== null && this._mux.isAvailable());
this._muxSession = config.muxSession || null;
@@ -544,6 +614,11 @@ export class Session extends EventEmitter {
this._geminiConfig = config.geminiConfig;
}
// Apply Antigravity configuration
if (config.antigravityConfig) {
this._antigravityConfig = config.antigravityConfig;
}
// Apply env overrides (exported at spawn, not persisted to disk).
// Legacy migration: pre-0.7.2 carried effort as the CLAUDE_CODE_EFFORT_LEVEL env var,
// which hard-locks /effort switching. Extract it into _effort (--settings soft default)
@@ -561,6 +636,7 @@ export class Session extends EventEmitter {
this._tmuxHistoryLimit = config.tmuxHistoryLimit ?? DEFAULT_TMUX_HISTORY_LIMIT;
this._remote = config.remote;
this._docker = config.docker;
this._owner = config.owner;
if (config.attachmentHistory && config.attachmentHistory.length > 0) {
this.restoreAttachmentHistory(config.attachmentHistory);
}
@@ -667,6 +743,16 @@ export class Session extends EventEmitter {
return this._docker;
}
/** Owning username in multi-user mode, else undefined. */
get owner(): string | undefined {
return this._owner;
}
/** Set the owning username (used by recovery to restore ownership). */
set owner(username: string | undefined) {
this._owner = username;
}
// Adopt a Claude conversation ID observed from an external source (e.g. hook
// payload). In interactive PTY mode Claude CLI emits no JSON to stdout, so
// `_handleJsonMessage` never sees `session_id`; hooks are the only signal
@@ -681,6 +767,15 @@ export class Session extends EventEmitter {
return this._muxSession?.muxName ?? null;
}
/**
* True when this session's PTY is a tmux client rather than the program itself.
* Read by the replay-side alt-screen strip, which must apply the same
* `useMux` gate as the live strip (isMuxAltScreenOnlyStripMode).
*/
get usesMux(): boolean {
return this._useMux;
}
get totalCost(): number {
return this._totalCost;
}
@@ -975,6 +1070,26 @@ export class Session extends EventEmitter {
this._imageWatcherEnabled = enabled;
}
/** Whether this session is pinned to the top of the session manager (COD-139). */
get pinned(): boolean {
return this._pinned;
}
/** When the session was pinned (epoch ms), or null when unpinned. */
get pinnedAt(): number | null {
return this._pinnedAt;
}
/**
* Set pin state (COD-139). Pinning stamps pinnedAt with now so the pinned
* group orders most-recently-pinned first; unpinning clears it. Idempotent:
* re-pinning an already-pinned session refreshes its pinnedAt.
*/
setPinned(pinned: boolean): void {
this._pinned = pinned;
this._pinnedAt = pinned ? Date.now() : null;
}
get flickerFilterEnabled(): boolean {
return this._flickerFilterEnabled;
}
@@ -1027,6 +1142,7 @@ export class Session extends EventEmitter {
workingDir: this.workingDir,
remote: this._remote,
docker: this._docker,
owner: this._owner,
currentTaskId: this._currentTaskId,
createdAt: this.createdAt,
lastActivityAt: this._lastActivityAt,
@@ -1040,6 +1156,8 @@ export class Session extends EventEmitter {
autoResumeEnabled: this._autoOps.autoResumeEnabled,
autoResumeAt: this._autoOps.autoResumeAt ?? undefined,
imageWatcherEnabled: this._imageWatcherEnabled,
pinned: this._pinned || undefined,
pinnedAt: this._pinned ? (this._pinnedAt ?? undefined) : undefined,
totalCost: this._totalCost,
inputTokens: this._totalInputTokens,
outputTokens: this._totalOutputTokens,
@@ -1059,6 +1177,7 @@ export class Session extends EventEmitter {
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
resumeSessionId: this._resumeSessionId,
effort: this._effort,
// COD-118: runtime-only — surfaced so the frontend can require explicit user
@@ -1067,6 +1186,7 @@ export class Session extends EventEmitter {
// recovery can re-attach.
respawnBlocked: this._respawnBlocked || undefined,
attachmentHistory: this.attachmentHistory.length > 0 ? this.attachmentHistory : undefined,
lastSubmitAt: this._lastSubmitAt || undefined,
// envOverrides intentionally NOT on the public SessionState type — they must not
// leak into SSE / GET /api/sessions broadcasts (schema allows OPENCODE_*, which
// can carry secrets). For disk persistence, session-manager calls
@@ -1203,24 +1323,33 @@ export class Session extends EventEmitter {
// No extra sleep — createSession() already waits for tmux readiness
}
// Attach to the mux session via PTY
// Prevent tmux from letting the newest browser attach dictate global window
// size; accepted Codeman resize events update it explicitly below.
mux.setManualWindowSize?.(this._muxSession!.muxName);
// Integration tests need a live input/output transport without attaching to
// the host's tmux server or agent CLI. Production still uses the real mux.
if (!IS_TEST_MODE) {
// Prevent tmux from letting the newest browser attach dictate global window
// size; accepted Codeman resize events update it explicitly below.
mux.setManualWindowSize?.(this._muxSession!.muxName);
}
// Query existing tmux window size so re-attach matches (avoids flicker from 120x40 default).
// MUST go through the dedicated socket (mux.muxSocket); a bare `tmux display` hits the
// default server, always fails for our socketed sessions, and silently falls back to 120x40.
const { cols: ptyCols, rows: ptyRows } = queryTmuxWindowSize(this._muxSession!.muxName, mux.muxSocket);
const { cols: ptyCols, rows: ptyRows } = IS_TEST_MODE
? { cols: DEFAULT_PTY_COLS, rows: DEFAULT_PTY_ROWS }
: queryTmuxWindowSize(this._muxSession!.muxName, mux.muxSocket);
const attachCommand = IS_TEST_MODE ? process.execPath : mux.getAttachCommand();
const attachArgs = IS_TEST_MODE ? ['-e', TEST_PTY_SCRIPT] : mux.getAttachArgs(this._muxSession!.muxName);
try {
this.ptyProcess = pty.spawn(mux.getAttachCommand(), mux.getAttachArgs(this._muxSession!.muxName), {
name: 'xterm-256color',
cols: ptyCols,
rows: ptyRows,
cwd: resolveMuxAttachCwd(this.workingDir, this._remote, this._docker),
// COD-75: codex/gemini get COLORTERM=truecolor — mirrors buildEnvExports()
// in tmux-manager.ts so the attach client and the tmux session agree.
env: buildMuxAttachEnv(this.mode === 'codex' || this.mode === 'gemini'),
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(attachCommand, attachArgs, {
name: 'xterm-256color',
cols: ptyCols,
rows: ptyRows,
cwd: resolveMuxAttachCwd(this.workingDir, this._remote, this._docker),
// COD-75: codex/gemini/antigravity get COLORTERM=truecolor — mirrors buildEnvExports()
// in tmux-manager.ts so the attach client and the tmux session agree.
env: buildMuxAttachEnv(this.mode === 'codex' || this.mode === 'gemini' || this.mode === 'antigravity'),
})
);
} catch (spawnErr) {
console.error(`[Session] Failed to spawn PTY for ${options.spawnErrLabel}:`, spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
@@ -1230,6 +1359,71 @@ export class Session extends EventEmitter {
return { isRestored };
}
/**
* COD-108 — re-establish a dropped REMOTE session. Triggered by the
* `TmuxManager` remote-reconnect watcher (via `remoteSessionDropped`): the
* watcher detects a dead remote pane, the session owner reassembles the SAME
* `RespawnPaneOptions` used for Claude-idle respawns and calls
* `respawnPane()` directly. For a remote session that re-runs
* `buildRemoteSessionCommand` (owned → `new-session -A`, non-owned →
* `attach`), which idempotently REATTACHES the still-running durable remote
* tmux session — scrollback + agent intact (proven COD-104/105).
*
* Deliberately does NOT route through the Claude-idle respawn-controller —
* this is a transport re-establish, not a `/clear`/`/compact` cycle.
*
* @returns true if the pane was respawned (reattach issued), false otherwise.
*/
async reattachRemote(): Promise<boolean> {
if (!this._remote) return false; // not a remote session
if (!this._useMux || !this._mux || !this._muxSession) return false;
const mux = this._mux;
// If tmux lost the whole session (not just a dead pane), there is nothing to
// respawn into — a genuine death, leave it for normal recovery/reconcile.
if (!mux.muxSessionExists(this._muxSession.muxName)) {
console.log('[Session] reattachRemote: mux session gone, skipping:', this._muxSession.muxName);
return false;
}
const newPid = await mux.respawnPane(this._buildRespawnPaneOptions());
if (!newPid) {
console.error('[Session] reattachRemote: respawnPane failed for', this._muxSession.muxName);
return false;
}
console.log('[Session] reattachRemote: reattached remote session', this._muxSession.muxName, 'pid', newPid);
return true;
}
/**
* Assemble the {@link RespawnPaneOptions} for this session. Single source of
* truth shared by interactive start, shell start (via their inline copies),
* and {@link reattachRemote} so the remote reattach path can never drift from
* the spawn path.
*/
private _buildRespawnPaneOptions(): import('./mux-interface.js').RespawnPaneOptions {
return {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
niceConfig: this._niceConfig,
model: this._model,
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
owner: this._owner,
};
}
private _handleTerminalOutput(data: string): void {
// Codex AND Claude Code emit sequences that wipe xterm.js scrollback, plus
// mouse-tracking enables that hijack the scroll wheel so the user can't reach
@@ -1250,9 +1444,16 @@ export class Session extends EventEmitter {
// SSE/WS stream carries them, keeping everything in the main buffer with
// scrollback intact. These are controlled TUIs whose cursor-positioned
// redraws overwrite only the cells they target, so non-erased rows keep
// their content. Gated to Codex/Claude (isAltScreenStripMode) — shell must
// keep the alt screen for vim/less/htop.
if (isAltScreenStripMode(this.mode)) {
// their content. Gated to Codex/Claude/Gemini (isAltScreenStripMode).
//
// Every OTHER mode (shell/opencode/antigravity) gets the NARROW strip when it
// is tmux-backed: alt-screen toggles only, because the sequence that breaks
// scrollback there is tmux's own client-side smcup at attach, not anything the
// program in the pane emitted (issue #205, see isMuxAltScreenOnlyStripMode).
// 3J and the mouse DECSETs stay, so `clear` and mouse-aware TUIs keep working.
const fullStrip = isAltScreenStripMode(this.mode);
const altOnlyStrip = !fullStrip && isMuxAltScreenOnlyStripMode(this.mode, this._useMux);
if (fullStrip || altOnlyStrip) {
// Reassemble sequences split across PTY chunk boundaries first: a chunk
// ending mid-sequence ('\x1b[?104' now, '9h' next) would slip past the
// strip below and leave xterm stuck in the scrollback-less alt buffer
@@ -1268,13 +1469,15 @@ export class Session extends EventEmitter {
data = data.slice(0, -splitTail[0].length);
if (!data) return;
}
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
// eslint-disable-next-line no-control-regex
data = data.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '');
if (fullStrip) {
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
}
}
// Scan terminal output for attachment requests. `codeman://attach?...` is an
@@ -1334,8 +1537,8 @@ export class Session extends EventEmitter {
// never show it — which left cliVersion undefined and silently disabled
// wheel-forwarding to Claude's own transcript (the only route to history in
// repaint/alt-screen mode; issue #154). Remote sessions run claude on
// another host, so a local probe wouldn't reflect their version — skip them
// and let the banner scrape handle those. Cached process-wide, best-effort.
// another host, so a local probe wouldn't reflect their version; they get
// their own over-ssh probe below. Cached process-wide, best-effort.
if (this.mode === 'claude' && !this._remote && !this._docker && !this._cliVersion) {
const probedVersion = getClaudeCliVersion();
if (probedVersion) {
@@ -1374,28 +1577,38 @@ export class Session extends EventEmitter {
}, DOCKER_CLI_VERSION_PROBE_DELAY_MS);
}
// Remote sessions run claude on ANOTHER HOST, so neither the local nor the
// docker probe applies, and the banner-scrape fallback they were left with
// is the unreliable path #154 was filed for, so remote Claude cases silently
// never got wheel-forwarding (noted in the #205 analysis). Probe over ssh,
// deferred so session start never waits on the ssh round-trip.
if (this.mode === 'claude' && this._remote && !this._cliVersion) {
const remoteMeta = this._remote;
setTimeout(() => {
if (this._isStopped || this._cliVersion) return;
void probeRemoteCliVersion(remoteMeta, this.mode)
.then((version) => {
if (!version || this._isStopped || this._cliVersion) return;
this._cliVersion = version;
this.emit('cliInfoUpdated', {
version: this._cliVersion,
model: this._cliModel,
accountType: this._cliAccountType,
latestVersion: this._cliLatestVersion,
});
})
.catch(() => {
/* best-effort */
});
}, REMOTE_CLI_VERSION_PROBE_DELAY_MS);
}
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
const { isRestored } = await this._setupOrAttachMuxSession({
respawnPaneOptions: {
sessionId: this.id,
workingDir: this.workingDir,
mode: this.mode,
niceConfig: this._niceConfig,
model: this._model,
claudeMode: this._claudeMode,
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
},
// Single source of truth shared with reattachRemote() (COD-108).
respawnPaneOptions: this._buildRespawnPaneOptions(),
createSessionOptions: {
sessionId: this.id,
workingDir: this.workingDir,
@@ -1408,12 +1621,14 @@ export class Session extends EventEmitter {
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
owner: this._owner,
},
spawnErrLabel: 'mux attachment',
});
@@ -1488,18 +1703,24 @@ export class Session extends EventEmitter {
if (this.mode === 'gemini') {
throw new Error('Gemini sessions require tmux. Direct PTY fallback is not supported.');
}
// Antigravity sessions require tmux for env override injection via setenv
if (this.mode === 'antigravity') {
throw new Error('Antigravity sessions require tmux. Direct PTY fallback is not supported.');
}
try {
// Pass --session-id to use the SAME ID as the Codeman session
// This ensures subagents can be directly matched to the correct tab
const args = buildInteractiveArgs(this.id, this._claudeMode, this._model, this._allowedTools, this._effort);
this.ptyProcess = pty.spawn('claude', args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(getClaudeBinaryPath(), args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
})
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY:', spawnErr);
this._status = 'stopped';
@@ -1523,7 +1744,7 @@ export class Session extends EventEmitter {
// === Auto-accept workspace trust dialog ===
// Claude CLI 2.x shows "Yes, I trust this folder" prompt on first launch per directory.
// Codeman sessions always use --dangerously-skip-permissions, so auto-accept.
// Codeman sessions run permission-skipping or classifier-guarded (auto) modes, so auto-accept.
if (!this._trustDialogAccepted && data.includes('trust this folder')) {
this._trustDialogAccepted = true;
console.log(`[Session] Auto-accepting workspace trust dialog for: ${this.id}`);
@@ -1765,8 +1986,9 @@ export class Session extends EventEmitter {
this._resetBuffers();
// Use user's default shell or bash
const shell = process.env.SHELL || '/bin/bash';
// Use user's default shell, falling back to a shell that actually exists.
// Shared with the tmux pane command so both paths launch the same binary.
const shell = resolveLocalShell();
console.log(
'[Session] Starting shell session with:',
shell + (this._useMux ? ` (with ${this._mux!.backend})` : '')
@@ -1785,6 +2007,7 @@ export class Session extends EventEmitter {
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
owner: this._owner,
},
createSessionOptions: {
sessionId: this.id,
@@ -1796,6 +2019,7 @@ export class Session extends EventEmitter {
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
owner: this._owner,
},
spawnErrLabel: 'shell mux attachment',
});
@@ -1820,13 +2044,15 @@ export class Session extends EventEmitter {
// Fallback to direct PTY if mux is not used
if (!this.ptyProcess) {
try {
this.ptyProcess = pty.spawn(shell, [], {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildShellEnv(this.id),
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(shell, [], {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildShellEnv(this.id),
})
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn shell PTY:', spawnErr);
this._status = 'stopped';
@@ -1923,17 +2149,19 @@ export class Session extends EventEmitter {
model ? `(model: ${model})` : ''
);
const args = buildPromptArgs(prompt, model);
const args = buildPromptArgs(prompt, model, this._claudeMode, this._allowedTools);
try {
this.ptyProcess = pty.spawn('claude', args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(getClaudeBinaryPath(), args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
})
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY for runPrompt:', spawnErr);
this.emit(
@@ -2392,36 +2620,41 @@ export class Session extends EventEmitter {
* For interactive sessions, this is how you send user input to Claude.
* Remember to include `\r` (carriage return) to simulate pressing Enter.
*
* @param data - The input data to send (text, escape sequences, etc.)
*
* @example
* ```typescript
* session.write('hello world'); // Text only, no Enter
* session.write('\r'); // Enter key
* session.write('ls -la\r'); // Command with Enter
* ```
*
* @param data - The input data to send (text, escape sequences, etc.)
* @returns true if the data reached a PTY. A session whose PTY is gone still
* discards the data, but it used to do so with no signal at all — which is how
* input could disappear while the caller believed it had been delivered.
*/
write(data: string): void {
this._trackCodexSubmit(data);
if (this.ptyProcess) {
this.ptyProcess.write(data);
}
write(data: string): boolean {
this._trackSubmit(data);
if (!this.ptyProcess) return false;
this.ptyProcess.write(data);
return true;
}
// ── Codex thread tracking ─────────────────────────────────────────────
// When a codex pane last submitted a message (Enter). The response-viewer
// correlates this against ~/.codex/history.jsonl entry timestamps to find
// the thread the pane is ACTUALLY on — the only signal that survives
// /resume, /new and /fork typed inside the codex TUI itself.
private _codexLastSubmitAt = 0;
// ── Conversation tracking ─────────────────────────────────────────────
// When this pane last submitted a message (Enter). The response-viewer
// correlates this against the CLI's own history.jsonl entry timestamps to
// find the conversation the pane is ACTUALLY on — the only signal that
// survives /clear, /resume, /new and /fork typed inside the TUI itself,
// none of which announce themselves on the PTY's stdout.
private _lastSubmitAt = 0;
get codexLastSubmitAt(): number {
return this._codexLastSubmitAt;
/** Wall-clock ms of this pane's last Enter; 0 if it has never submitted. */
get lastSubmitAt(): number {
return this._lastSubmitAt;
}
private _trackCodexSubmit(data: string): void {
if (this.mode === 'codex' && (data.includes('\r') || data.includes('\n'))) {
this._codexLastSubmitAt = Date.now();
private _trackSubmit(data: string): void {
if (data.includes('\r') || data.includes('\n')) {
this._lastSubmitAt = Date.now();
}
}
@@ -2461,6 +2694,23 @@ export class Session extends EventEmitter {
return true;
}
/**
* Undo the bookkeeping of {@link shouldApplyInput} for a delivery that failed.
*
* Without this, the reliable-delivery layer guarantees exactly-once delivery of
* something that may never have been delivered: the seq is recorded as applied
* BEFORE the write is attempted, so a client retry — the very mechanism the seq
* exists for — is rejected as a duplicate and the input is lost for good.
*
* Only rolls back if `seq` is still the newest recorded one; a later input has
* already superseded it and must not be re-opened.
*/
forgetInputSeq(clientId: string, seq: number): void {
if (this._appliedInputSeq.get(clientId) === seq) {
this._appliedInputSeq.set(clientId, seq - 1);
}
}
/**
* Sends input via the terminal multiplexer's direct input mechanism.
*
@@ -2478,7 +2728,7 @@ export class Session extends EventEmitter {
* ```
*/
async writeViaMux(data: string): Promise<boolean> {
this._trackCodexSubmit(data);
this._trackSubmit(data);
if (this._mux && this._muxSession) {
return this._mux.sendInput(this.id, data);
}
@@ -2587,7 +2837,7 @@ export class Session extends EventEmitter {
if (this.ptyProcess && (dimsChanged || options.force)) {
this._ptyCols = cols;
this._ptyRows = rows;
if (this._mux && this._muxSession) {
if (!IS_TEST_MODE && this._mux && this._muxSession) {
this._mux.resizeWindow?.(this._muxSession.muxName, cols, rows);
}
this.ptyProcess.resize(cols, rows);
+34
View File
@@ -278,6 +278,9 @@ export class StateStore {
if (this.state.cronJobRuns) {
parts.push(`"cronJobRuns":${JSON.stringify(this.state.cronJobRuns)}`);
}
if (this.state.sessionOrder) {
parts.push(`"sessionOrder":${JSON.stringify(this.state.sessionOrder)}`);
}
return `{${parts.join(',')}}`;
}
@@ -485,6 +488,25 @@ export class StateStore {
this.save();
}
/**
* COD-142: Remove a session's persisted record on kill UNLESS it is pinned.
* A pinned session is demoted to a lightweight `stopped` record (pin retained)
* so it stays visible in the session-manager pinned group and survives restart.
* Unpinned sessions are fully removed (unchanged behavior).
* @returns 'preserved' if demoted to stopped+pinned, 'removed' if deleted, 'absent' if no record existed.
*/
demoteOrRemoveSession(id: string): 'preserved' | 'removed' | 'absent' {
const existing = this.state.sessions[id];
if (!existing) return 'absent';
if (existing.pinned === true) {
// Demote in place: keep identity/resume fields + pin, mark stopped, clear live runtime.
this.setSession(id, { ...existing, status: 'stopped', pid: null });
return 'preserved';
}
this.removeSession(id);
return 'removed';
}
/**
* Cleans up stale sessions from state that don't have corresponding active sessions.
* @param activeSessionIds - Set of currently active session IDs
@@ -499,6 +521,7 @@ export class StateStore {
for (const sessionId of allSessionIds) {
if (!activeSessionIds.has(sessionId)) {
if (this.state.sessions[sessionId]?.pinned === true) continue; // COD-142: pinned records persist even with no live session
const name = this.state.sessions[sessionId]?.name;
cleaned.push({ id: sessionId, name });
delete this.state.sessions[sessionId];
@@ -630,6 +653,17 @@ export class StateStore {
this.save();
}
/** Returns the global tab order (ordered sessionIds), [] if unset. COD-131. */
getSessionOrder(): string[] {
return this.state.sessionOrder ?? [];
}
/** Persists the global tab order (ordered sessionIds) and triggers a debounced save. COD-131. */
setSessionOrder(order: string[]): void {
this.state.sessionOrder = order;
this.save();
}
/** Resets all state to initial values and saves immediately. */
reset(): void {
this.state = createInitialState();
+571 -62
View File
@@ -22,7 +22,8 @@
*/
import { EventEmitter } from 'node:events';
import { execSync, exec } from 'node:child_process';
import { collectDescendants } from './proc-tree.js';
import { execSync, exec, execFile } from 'node:child_process';
import { promisify } from 'node:util';
const execAsync = promisify(exec);
@@ -43,12 +44,18 @@ import {
type CodexConfig,
type EffortLevel,
type GeminiConfig,
type AntigravityConfig,
type SessionRemote,
type SessionDocker,
type DockerCommandMode,
} from './types.js';
import { buildEffortCliArgs } from './session-cli-builder.js';
import { buildSshConnectionArgs, defaultRemoteCommandForMode, remoteSshTarget } from './remote-hosts.js';
import {
buildSshConnectionArgs,
defaultRemoteCommandForMode,
remoteLoginShellCommand,
remoteSshTarget,
} from './remote-hosts.js';
import {
buildDockerBaseArgs,
buildDockerCreateArgs,
@@ -56,9 +63,11 @@ import {
CONTAINER_HOME,
defaultDockerCommandForMode,
hostGatewayAlias,
resolveCredentialMounts,
resolveDockerClaudeArtifacts,
resolveDockerCredentialArtifacts,
type DockerCreateContext,
type DockerMount,
type DockerSeedCopy,
} from './docker-hosts.js';
import {
wrapWithNice,
@@ -67,6 +76,9 @@ import {
resolveOpenCodeDir,
resolveCodexDir,
resolveGeminiDir,
resolveAntigravityDir,
resolveLocalShell,
loginShellArgs,
} from './utils/index.js';
import type {
TerminalMultiplexer,
@@ -76,12 +88,29 @@ import type {
RespawnPaneOptions,
PaneCaptureOptions,
} from './mux-interface.js';
import {
decideReconnect,
advanceBackoff,
freshReconnectState,
resetReconnectState,
type RemoteReconnectState,
} from './remote-reconnect.js';
// ============================================================================
// Timing Constants
// ============================================================================
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** How long a cached process snapshot stays usable. */
const PROC_SNAPSHOT_TTL_MS = 2000;
/**
* How long the kill path waits for a fresh snapshot before giving up on it.
* Shorter than EXEC_TIMEOUT_MS on purpose: killSession has two further strategies
* (process group, tmux kill-session) and must reach them even when `ps` is wedged.
*/
const PROC_SNAPSHOT_WAIT_MS = 1500;
import { DEFAULT_TMUX_HISTORY_LIMIT, DEFAULT_TERMINAL_BUFFER_MAX_BYTES } from './config/terminal-history.js';
/**
@@ -108,6 +137,9 @@ const GRACEFUL_SHUTDOWN_WAIT_MS = 100;
/** Default stats collection interval (2 seconds) */
const DEFAULT_STATS_INTERVAL_MS = 2000;
/** Default remote-reconnect watcher poll interval (5 seconds) — COD-108 */
const DEFAULT_REMOTE_RECONNECT_INTERVAL_MS = 5000;
/** Stable cwd for tmux server/pane launch; actual session cwd is reached inside the pane. */
const TMUX_LAUNCH_CWD = '/tmp';
@@ -132,6 +164,20 @@ const IS_TEST_MODE = !!process.env.VITEST;
/** Path to persisted mux session metadata */
const MUX_SESSIONS_FILE = dataPath('mux-sessions.json');
/**
* COD-108 kill-switch: `remoteAutoReconnect` app setting (default ON). Read at
* call time (like headroom routing) so a settings change takes effect without a
* restart. Absent/non-boolean ⇒ true (feature on).
*/
function isRemoteAutoReconnectEnabled(): boolean {
try {
const s = JSON.parse(readFileSync(dataPath('settings.json'), 'utf8')) as Record<string, unknown>;
return typeof s.remoteAutoReconnect === 'boolean' ? s.remoteAutoReconnect : true;
} catch {
return true;
}
}
/** Regex to validate tmux session names (only allow safe characters) */
const SAFE_MUX_NAME_PATTERN = /^codeman-[a-f0-9-]+$/;
@@ -561,6 +607,8 @@ function buildClaudePermissionFlags(claudeMode?: ClaudeMode, allowedTools?: stri
switch (mode) {
case 'dangerously-skip-permissions':
return ' --dangerously-skip-permissions';
case 'auto':
return ' --permission-mode auto';
case 'allowedTools':
if (allowedTools) {
// Sanitize: allow tool names with patterns like Bash(git:*), space/comma-separated
@@ -612,6 +660,10 @@ export function buildCodexCommand(config?: CodexConfig): string {
parts.push('--dangerously-bypass-approvals-and-sandbox');
}
if (config?.animations !== undefined) {
parts.push('--config', `tui.animations=${config.animations ? 'true' : 'false'}`);
}
if (config?.model) {
const safeModel = /^[a-zA-Z0-9._\-/]+$/.test(config.model) ? config.model : undefined;
if (safeModel) parts.push('--model', safeModel);
@@ -654,6 +706,34 @@ function buildGeminiCommand(config?: GeminiConfig): string {
return parts.join(' ');
}
/**
* Build the Antigravity CLI (agy) command with appropriate flags.
*
* Unlike gemini's yolo default, `--dangerously-skip-permissions` is only added
* when the config explicitly asks for it (the frontend sends it for parity with
* Codeman's Claude default; the multi-user clamp strips it for non-granted owners,
* and an ABSENT config stays at agy's own prompting default — safe like Codex).
*/
function buildAntigravityCommand(config?: AntigravityConfig): string {
const parts = ['agy'];
if (config?.dangerouslySkipPermissions) {
parts.push('--dangerously-skip-permissions');
}
if (config?.model) {
const safeModel = /^[a-zA-Z0-9._\-/]+$/.test(config.model) ? config.model : undefined;
if (safeModel) parts.push('--model', safeModel);
}
if (config?.resumeConversationId) {
const safeId = /^[a-zA-Z0-9._-]+$/.test(config.resumeConversationId) ? config.resumeConversationId : undefined;
if (safeId) parts.push('--conversation', safeId);
}
return parts.join(' ');
}
/**
* Build the spawn command for any session mode.
* Shared by createSession() and respawnPane() to avoid duplication.
@@ -672,7 +752,7 @@ function buildEffortSettingsFlag(effort?: EffortLevel): string {
return flag && value ? ` ${flag} '${value}'` : '';
}
function buildSpawnCommand(options: {
export function buildSpawnCommand(options: {
mode: SessionMode;
sessionId: string;
model?: string;
@@ -681,6 +761,7 @@ function buildSpawnCommand(options: {
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
resumeSessionId?: string;
effort?: EffortLevel;
}): string {
@@ -711,7 +792,24 @@ function buildSpawnCommand(options: {
if (options.mode === 'gemini') {
return buildGeminiCommand(options.geminiConfig);
}
return '$SHELL';
if (options.mode === 'antigravity') {
return buildAntigravityCommand(options.antigravityConfig);
}
// #208: NOT the literal '$SHELL'. This string is embedded in the `bash -c "…"`
// argument of the respawn-pane line, which execSync runs through `/bin/sh -c`,
// so a `$SHELL` here is expanded by the SERVER process's shell against the
// SERVER process's env — empty in containers and system systemd units, leaving
// the pane command ending in a dangling `&&` ("syntax error: unexpected end of
// file", pane dead on arrival). Resolve it in Node and quote the result.
// #209: launch it as a LOGIN shell, which is what tmux itself does for a pane
// with no `default-command`, so a Codeman shell tab matches a hand-started tmux
// one. That is what picks up /etc/profile and /etc/profile.d/* — a systemd
// --user service never sourced them, so its PATH is what every pane inherited.
// The flags come from loginShellArgs() rather than being hardcoded: they are
// appended to a path that ultimately comes from the passwd entry, and a shell
// that rejects an unknown flag exits on the spot, which is #208 all over again.
const shell = resolveLocalShell();
return `${shellescape(shell)}${loginShellArgs(shell)}`;
}
/**
@@ -775,9 +873,23 @@ export function buildRemoteLaunchCommand(options: {
mode: SessionMode;
remote: SessionRemote;
sessionId: string;
claudeMode?: ClaudeMode;
allowedTools?: string;
}): string {
const { mode, remote, sessionId } = options;
const modeCommand = remote.commands?.[mode] || defaultRemoteCommandForMode(mode);
const { mode, remote, sessionId, claudeMode, allowedTools } = options;
// §6.3: honor the session's EFFECTIVE claude permission mode on remote instead of
// hardcoding --dangerously-skip-permissions, so a non-granted multi-user user's
// downgraded 'auto' actually reaches the remote agent (the default command otherwise
// ignored claudeMode). A per-host `commands.claude` override stays authoritative
// (admin's explicit choice). Wrapped in `$SHELL -i -l -c` for the same reason as
// `defaultRemoteCommandForMode`: `claude` lives under a per-user PATH entry that
// only an interactive login shell resolves (see that function's comment).
const override = remote.commands?.[mode];
const modeCommand = override
? override
: mode === 'claude'
? remoteLoginShellCommand(`claude${buildClaudePermissionFlags(claudeMode, allowedTools)}`)
: defaultRemoteCommandForMode(mode);
const remoteName = remoteTmuxSessionName(sessionId);
// Innermost: the command tmux runs in the new pane. Run via `/bin/sh -c` by
@@ -795,6 +907,31 @@ export function buildRemoteLaunchCommand(options: {
`set -t ${remoteName} mouse off`,
`set -t ${remoteName} prefix C-q`,
'set -s escape-time 0',
// COD-106 — shared/collaborative sessions: tmux defaults to sizing a window
// to the SMALLEST attached client, so two Codemans at different viewports
// would fight (clamp to the smaller). `window-size latest` sizes to the
// most-recently-active client instead, so concurrent clients coexist.
// Per-session scoped (`set -t <name>`, matching #145's hardening) so a shared
// remote tmux server's other sessions keep their own sizing behavior.
`set -t ${remoteName} window-size latest`,
// #210: keep a CRASHED pane so the failure is still on screen. Without this,
// tmux destroys the pane -> window -> session (and, being the only session,
// the whole remote server) the instant the pane command exits, which tears the
// local `ssh -t` attach down with it; reconnect's `-A` then builds a fresh
// session and the cycle can repeat as a flap loop with no evidence surviving.
// That is how the exit-127 PATH bug fixed above stayed invisible.
//
// `failed`, NOT `on`: `on` keeps the pane on a CLEAN exit too, so typing
// `exit` in a remote shell leaves a dead pane behind, the session outlives it,
// and the next launch's `-A` reattaches to that corpse ("Pane is dead (status
// 0)") instead of starting a shell — verified against a real tmux. `failed`
// keeps the pane only on a non-zero exit, which is exactly the diagnostic case.
//
// LAST in the chain on purpose: tmux aborts the remaining commands of a `\;`
// sequence once one errors (also verified), and `failed` needs tmux >= 3.2 on
// the REMOTE host. Trailing, a rejection costs only this option; leading, it
// would silently drop status/mouse/prefix/escape-time/window-size with it.
`set -t ${remoteName} remain-on-exit failed`,
].join(' \\; ');
// ssh runs its trailing args through the remote login shell, so the entire
@@ -856,24 +993,50 @@ export function dockerTmuxSessionName(sessionId: string): string {
const RESUME_ID_SAFE = /^[A-Za-z0-9._-]+$/;
/**
* Append the CLI-specific resume flag to a pane command. Only fires when the
* in-container tmux is RE-CREATED (`new-session -A` makes the flag inert on a
* live reattach), i.e. exactly when the previous live agent was lost and we want
* to resume the conversation from the bind-mounted transcript.
* Append the CLI-specific resume flag to a pane command (codex/gemini/antigravity). Only fires
* when the in-container tmux is RE-CREATED (`new-session -A` makes the flag inert
* on a live reattach), i.e. exactly when the previous live agent was lost and we
* want to resume the conversation from the bind-mounted transcript. Claude mode
* uses claudeDockerPaneCommand instead.
*/
function appendResumeFlag(modeCommand: string, mode: SessionMode, resumeId: string): string {
if (!RESUME_ID_SAFE.test(resumeId)) return modeCommand;
switch (mode) {
case 'claude':
case 'gemini':
return `${modeCommand} --resume ${resumeId}`;
case 'codex':
return `${modeCommand} resume ${resumeId}`;
case 'antigravity':
return `${modeCommand} --conversation ${resumeId}`;
default:
return modeCommand; // shell / opencode: no resume
}
}
/**
* Claude-mode pane command with a DETERMINISTIC conversation id (the docker analog
* of buildSpawnCommand's --resume/--session-id logic). A fresh launch passes
* `--session-id <sessionId>`, so the in-container conversation id is knowable
* host-side (resume-id capture + subagent/workflow correlation) WITHOUT relying on
* hook reachability. When the in-container tmux was re-created after a container
* stop/reboot, the same command re-runs against the surviving transcript:
* `--session-id` exits 1 ("already in use") and the `||` fallback RESUMES that
* conversation (verified CLI behavior). An explicit resumeId gets the local
* builder's shape — resume first, session-id fallback — so a stale id never
* dead-panes. The leading `exec ` is stripped: an exec'd first branch could never
* fall back.
*/
function claudeDockerPaneCommand(modeCommand: string, sessionId: string, resumeId?: string): string {
if (!RESUME_ID_SAFE.test(sessionId)) return modeCommand; // defensive — ids are server-minted uuids
const cmd = modeCommand.replace(/^exec\s+/, '');
const rid = resumeId && RESUME_ID_SAFE.test(resumeId) ? resumeId : undefined;
if (rid && rid !== sessionId) {
return `${cmd} --resume ${rid} || ${cmd} --session-id ${sessionId}`;
}
const cid = rid ?? sessionId;
return `${cmd} --session-id ${cid} || ${cmd} --resume ${cid}`;
}
/** Fully-resolved inputs for buildDockerLaunchCommand (pure). */
export interface DockerLaunchOptions {
mode: SessionMode;
@@ -885,6 +1048,14 @@ export interface DockerLaunchOptions {
execEnv: Record<string, string>;
/** exec-time NAME-ONLY env forwarded from Codeman's process env (codex/gemini keys) */
execEnvNames: string[];
/**
* Files to copy from read-only seed mounts into the container's writable HOME once
* before launch (guarded so reconnects never clobber). Isolates Claude state: the
* merged `~/.claude.json`, plus `~/.claude/.credentials.json` + `settings.json`,
* are writable copies (not host mounts), so the container never re-auths and never
* writes its runtime state back into the host `~/.claude`.
*/
seedCopies?: DockerSeedCopy[];
}
/**
@@ -895,7 +1066,7 @@ export interface DockerLaunchOptions {
* command -> `docker exec … sh -lc '<tmux>'` -> tmux `'<paneCommand>'`.
*/
export function buildDockerLaunchCommand(opts: DockerLaunchOptions): string {
const { mode, docker, sessionId, resumeSessionId, createContext, execEnv, execEnvNames } = opts;
const { mode, docker, sessionId, resumeSessionId, createContext, execEnv, execEnvNames, seedCopies } = opts;
const base = buildDockerBaseArgs(docker).join(' ');
const createArgs = buildDockerCreateArgs(createContext).join(' ');
const name = shellescape(docker.containerName);
@@ -905,7 +1076,11 @@ export function buildDockerLaunchCommand(opts: DockerLaunchOptions): string {
const sid = sessionId.slice(0, 8);
let modeCommand = docker.commands?.[mode as DockerCommandMode] || defaultDockerCommandForMode(mode);
if (resumeSessionId) modeCommand = appendResumeFlag(modeCommand, mode, resumeSessionId);
if (mode === 'claude') {
modeCommand = claudeDockerPaneCommand(modeCommand, sessionId, resumeSessionId);
} else if (resumeSessionId) {
modeCommand = appendResumeFlag(modeCommand, mode, resumeSessionId);
}
// Run by tmux via /bin/sh -c, so the path is shell-quoted here. `exec` makes the
// pane PID the agent itself.
const paneCommand = `cd ${workdir} && ${modeCommand}`;
@@ -932,7 +1107,7 @@ export function buildDockerLaunchCommand(opts: DockerLaunchOptions): string {
for (const extra of docker.extraExecArgs ?? []) execEnvFlags.push(shellescape(extra));
const imageMissingMsg = shellescape(
`Codeman: base image ${docker.image} not present (build: node scripts/build-agent-image.mjs)`
`Codeman: base image ${docker.image} not present (it is normally auto-built on first use)`
);
const startFailMsg = shellescape(`Codeman: container ${docker.containerName} failed to start (docker daemon down?)`);
@@ -940,7 +1115,18 @@ export function buildDockerLaunchCommand(opts: DockerLaunchOptions): string {
// create-if-missing (idempotent): reconnect / boot recovery re-runs this exact chain.
const ensure = `${base} inspect ${name} >/dev/null 2>&1 || ${base} ${createArgs}`;
const start = `${base} start ${name} >/dev/null 2>&1 || { echo ${startFailMsg}; exit 1; }`;
const execCmd = `exec ${base} exec -it --workdir ${workdir} ${execEnvFlags.join(' ')} ${name} sh -lc ${shellescape(tmuxInvocation)}`;
// Seed writable credential config from read-only host mounts ONCE per container
// (guarded by [ -e ] so reconnects never clobber in-container config; `cp -a` for
// whole-dir credential seeds). mkdir -p the parent so a file seed works even when
// no sibling share-mount pre-created the dir. Paths are fixed CONTAINER_HOME
// constants (no shell metachars), so the whole inner command is shell-quoted once.
const seedSteps = (seedCopies ?? []).map((s) => {
const cp = s.recursive ? 'cp -a' : 'cp';
const parent = s.to.slice(0, s.to.lastIndexOf('/'));
return `mkdir -p ${parent} 2>/dev/null; [ -e ${s.to} ] || ${cp} ${s.from} ${s.to} 2>/dev/null || true`;
});
const innerCmd = seedSteps.length ? `${seedSteps.join(' ; ')} ; ${tmuxInvocation}` : tmuxInvocation;
const execCmd = `exec ${base} exec -it --workdir ${workdir} ${execEnvFlags.join(' ')} ${name} sh -lc ${shellescape(innerCmd)}`;
return [imageCheck, ensure, start, execCmd].join(' ; ');
}
@@ -991,12 +1177,29 @@ export function resolveDockerLaunchOptions(
: ['--user', `${uid}:0`]; // Linux: host uid + GID 0 (OpenShift arbitrary-uid writable HOME)
const gatewayAlias = hostGatewayAlias(docker.engine);
const credentialMounts: DockerMount[] = docker.mountCredentials ? resolveCredentialMounts(home) : [];
const credentialMounts: DockerMount[] = [];
const extraMounts: DockerMount[] = [];
// Isolated credential state (Claude + codex/gemini/gcloud/opencode): each store
// shares ONLY what a host feature / --resume needs (Claude projects/, codex
// sessions/+history) and seeds everything else (tokens, settings, configs) as
// writable copies, so the container is authed WITHOUT re-auth and WITHOUT writing
// its runtime state back into the host dirs. Only when credentials are mounted.
let seedCopies: DockerSeedCopy[] = [];
if (docker.mountCredentials) {
const claudeArtifacts = resolveDockerClaudeArtifacts(home, docker.containerName, docker.containerWorkdir);
const credArtifacts = resolveDockerCredentialArtifacts(home);
extraMounts.push(...claudeArtifacts.mounts, ...credArtifacts.mounts);
seedCopies = [...claudeArtifacts.seedCopies, ...credArtifacts.seedCopies];
}
const envCreate: Record<string, string> = {
HOME: CONTAINER_HOME,
TERM: 'xterm-256color',
COLORTERM: 'truecolor',
// Force a UTF-8 locale (the base image defaults to POSIX/C). Without this, tmux
// runs in non-UTF-8 mode and renders Claude's Unicode box-drawing (─│┌┐) as raw
// VT100 ACS glyphs (`qqqq…`). `C.UTF-8` is built into glibc (no locale-gen).
LANG: 'C.UTF-8',
LC_ALL: 'C.UTF-8',
// Give claude a temp dir it will own inside HOME. Its default `/tmp/claude-<uid>`
// is refused when that path pre-exists root-owned — which happens when the
// workspace bind-mount path traverses it (e.g. a workspace under /tmp/claude-<uid>).
@@ -1031,6 +1234,11 @@ export function resolveDockerLaunchOptions(
const execEnv: Record<string, string> = {
TERM: 'xterm-256color',
COLORTERM: 'truecolor',
// UTF-8 at exec time too, so the tmux CLIENT this exec launches is UTF-8 and
// renders box-drawing correctly even when reattaching to a container created
// before this fix (client_utf8 is per-client, resolved from the exec's locale).
LANG: 'C.UTF-8',
LC_ALL: 'C.UTF-8',
CODEMAN_SESSION_ID: sessionId.slice(0, 8),
CODEMAN_MUX: '1',
};
@@ -1043,7 +1251,57 @@ export function resolveDockerLaunchOptions(
? ['GEMINI_API_KEY', 'GOOGLE_API_KEY']
: [];
return { mode, docker, sessionId, resumeSessionId, createContext, execEnv, execEnvNames };
return { mode, docker, sessionId, resumeSessionId, createContext, execEnv, execEnvNames, seedCopies };
}
/**
* COD-105 — build the SSH command that ATTACHES to an EXISTING `codeman-*` tmux
* session on the remote host (one this Codeman didn't create — discovered via
* `listRemoteCodemanSessions`). Sibling of `buildRemoteLaunchCommand`.
*
* Emits:
* ssh -o BatchMode=yes -t [<COD-107 connection opts>] user@host \
* 'tmux -L codeman attach -t <session>'
*
* - `attach` (NOT `new-session -A`) so we only join an existing session; the
* remote session keeps running independent of us, which is exactly why the
* resulting Codeman session is NON-OWNED (see `SessionRemote.owned`): closing
* the local tab must detach, never `kill-session` the remote.
* - The remote session name is shell-escaped so a value with metachars stays a
* single token inside the quoted tmux invocation.
* - COD-107 — connection options (`-p`, `-i`, `-J`, SOCKS `-o ProxyCommand`,
* arbitrary `-o`) come from the shared `buildSshConnectionArgs`, so attach
* connects identically to launch / discovery / the prereq probe. `-t` sits
* right after `ssh -o BatchMode=yes` (a PTY is required for interactive tmux).
*/
export function buildRemoteAttachCommand(remote: SessionRemote, remoteSessionName: string): string {
const tmuxInvocation = `tmux -L codeman attach -t ${shellescape(remoteSessionName)}`;
const [ssh, batchMode, ...connectionArgs] = buildSshConnectionArgs(remote);
const sshParts = [ssh, batchMode, '-t', ...connectionArgs, remoteSshTarget(remote), shellescape(tmuxInvocation)];
return sshParts.join(' ');
}
/**
* COD-105 — choose the right remote ssh command for a session's ownership:
* - NON-owned (`remote.owned === false`): ATTACH to a discovered remote tmux
* session by its EXISTING name (`remote.remoteSessionName`, falling back to
* this session's deterministic name). We only join — never create.
* - owned (default): LAUNCH/attach-or-create via `buildRemoteLaunchCommand`
* (COD-104), which we then own and may explicitly kill.
*/
function buildRemoteSessionCommand(options: {
mode: SessionMode;
remote: SessionRemote;
sessionId: string;
claudeMode?: ClaudeMode;
allowedTools?: string;
}): string {
const { remote, sessionId } = options;
if (remote.owned === false) {
const target = remote.remoteSessionName || remoteTmuxSessionName(sessionId);
return buildRemoteAttachCommand(remote, target);
}
return buildRemoteLaunchCommand(options);
}
/**
@@ -1199,6 +1457,17 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
/** Track last-known pane count per session to avoid unnecessary tmux set-option calls */
private lastPaneCount: Map<string, number> = new Map();
// ── COD-108 remote-reconnect watcher state ────────────────────────────────
/** Periodic watcher that re-establishes dropped remote sessions. */
private remoteReconnectInterval: NodeJS.Timeout | null = null;
/** Per-session backoff/attempt bookkeeping (sessionId → state). */
private reconnectState: Map<string, RemoteReconnectState> = new Map();
/**
* Sessions excluded from auto-reconnect because they are being intentionally
* torn down (killed/detached/stopping). A guarded session is NEVER revived.
*/
private reconnectGuard: Set<string> = new Set();
private trueColorConfigured = false;
constructor() {
@@ -1307,8 +1576,10 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const exports = [
'export LANG=en_US.UTF-8',
'export LC_ALL=en_US.UTF-8',
mode === 'codex' || mode === 'gemini' ? 'export COLORTERM=truecolor' : 'unset COLORTERM',
...(mode === 'codex' || mode === 'gemini' ? ['unset NO_COLOR'] : []),
mode === 'codex' || mode === 'gemini' || mode === 'antigravity'
? 'export COLORTERM=truecolor'
: 'unset COLORTERM',
...(mode === 'codex' || mode === 'gemini' || mode === 'antigravity' ? ['unset NO_COLOR'] : []),
// Stamp each Codex pane with a unique originator so the response-viewer
// can locate THIS pane's rollout exactly — codex writes the value into
// session_meta.originator of every rollout it creates. Without it,
@@ -1318,7 +1589,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
// Only exported when the server has stamped the real URL (scheme+host+port,
// set in WebServer.start()). A hardcoded fallback here exported the wrong
// scheme on HTTPS installs; leaving the variable unset makes in-session
// guards fail closed instead of curling a URL that was never right.
...(process.env.CODEMAN_API_URL ? [`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL}`] : []),
// Path only (not the secret value): hook curl commands cat the file at
// execution time, so the COD-54 hook secret stays off the command line.
`export CODEMAN_HOOK_SECRET_FILE="${dataPath('hook-secret')}"`,
@@ -1391,6 +1666,10 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const dir = resolveGeminiDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
if (mode === 'antigravity') {
const dir = resolveAntigravityDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
return { pathExport: '', dir: null };
}
@@ -1438,12 +1717,14 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
envOverrides,
effort,
historyLimit = DEFAULT_TMUX_HISTORY_LIMIT,
remote,
docker,
owner,
} = options;
const muxName = `codeman-${sessionId.slice(0, 8)}`;
@@ -1464,6 +1745,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
workingDir,
remote,
docker,
owner,
mode,
attached: false,
name,
@@ -1487,6 +1769,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
if (mode === 'gemini' && !cliDir) {
throw new Error('Gemini CLI not found. Install with: npm install -g @google/gemini-cli');
}
if (mode === 'antigravity' && !cliDir) {
throw new Error(
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash'
);
}
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
@@ -1499,6 +1786,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
effort,
});
@@ -1512,7 +1800,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const fullCmd = docker
? buildDockerLaunchCommand(resolveDockerLaunchOptions(mode, docker, sessionId, resumeSessionId))
: remote
? buildRemoteLaunchCommand({ mode, remote, sessionId })
? buildRemoteSessionCommand({ mode, remote, sessionId, claudeMode, allowedTools })
: localFullCmd;
// Create tmux session in three steps to handle cold-start (no server running)
@@ -1640,6 +1928,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
workingDir,
remote,
docker,
owner,
mode,
attached: false,
name,
@@ -1720,6 +2009,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
envOverrides,
effort,
@@ -1757,6 +2047,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
effort,
});
@@ -1766,7 +2057,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const fullCmd = docker
? buildDockerLaunchCommand(resolveDockerLaunchOptions(mode, docker, sessionId, resumeSessionId))
: remote
? buildRemoteLaunchCommand({ mode, remote, sessionId })
? buildRemoteSessionCommand({ mode, remote, sessionId, claudeMode, allowedTools })
: localFullCmd;
try {
@@ -1818,27 +2109,102 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
// Get all child process PIDs recursively
private getChildPids(pid: number): number[] {
const pids: number[] = [];
try {
const output = execSync(`pgrep -P ${pid}`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (output) {
for (const childPid of output
.split('\n')
.map((p) => parseInt(p, 10))
.filter((p) => !Number.isNaN(p))) {
pids.push(childPid);
pids.push(...this.getChildPids(childPid));
/** One `ps` snapshot of the whole process table, cached briefly. */
private static procSnapshot: { at: number; byParent: Map<number, number[]> } | null = null;
/** Single-flight guard so a hung `ps` cannot pile up parallel refreshes. */
private static procRefresh: { started: number; promise: Promise<Map<number, number[]>> } | null = null;
/**
* Fork ONE `ps` asynchronously and cache the parent -> children map.
*
* Async on purpose: a synchronous fork here would block the event loop on every
* stats tick, and under the procfs pathology this module exists to survive,
* `execSync`'s timeout cannot return at all (spawnSync waits for the unkillable
* child) — freezing the whole server where a hung async poll only costs staleness.
*/
private static refreshProcSnapshot(): Promise<Map<number, number[]>> {
const inFlight = TmuxManager.procRefresh;
// Reuse an in-flight refresh — unless it is old enough to be presumed stuck.
if (inFlight && Date.now() - inFlight.started < EXEC_TIMEOUT_MS * 2) return inFlight.promise;
const started = Date.now();
const promise = new Promise<Map<number, number[]>>((resolve) => {
execFile('ps', ['-eo', 'pid=,ppid='], { timeout: EXEC_TIMEOUT_MS, maxBuffer: 8 * 1024 * 1024 }, (err, out) => {
if (TmuxManager.procRefresh?.started === started) TmuxManager.procRefresh = null;
if (err) {
// ANY error, not just an empty one: a timed-out or truncated `ps` yields
// partial output, and caching that as fresh would make whole subtrees
// invisible — including to the kill path. Stale beats wrong.
console.error('[TmuxManager] process snapshot failed:', err);
resolve(TmuxManager.procSnapshot?.byParent ?? new Map());
return;
}
}
} catch {
// No children or command failed
const byParent = new Map<number, number[]>();
for (const line of String(out).split('\n')) {
const parts = line.trim().split(/\s+/);
if (parts.length < 2) continue;
const pid = parseInt(parts[0], 10);
const ppid = parseInt(parts[1], 10);
if (Number.isNaN(pid) || Number.isNaN(ppid)) continue;
const list = byParent.get(ppid);
if (list) list.push(pid);
else byParent.set(ppid, [pid]);
}
TmuxManager.procSnapshot = { at: Date.now(), byParent };
resolve(byParent);
});
});
TmuxManager.procRefresh = { started, promise };
return promise;
}
/**
* Best snapshot WITHOUT forking: returns the cache, kicking off a background
* refresh when it has gone stale, and never blocks. Stats and window-title
* consumers tolerate data one interval old; nothing that KILLS may use this.
*/
private childrenByParent(): Map<number, number[]> {
const cached = TmuxManager.procSnapshot;
if (!cached || Date.now() - cached.at >= PROC_SNAPSHOT_TTL_MS) {
void TmuxManager.refreshProcSnapshot();
}
return pids;
return cached?.byParent ?? new Map();
}
/**
* Descendants from a snapshot that is not the cached one — the kill path's variant.
*
* killSession re-scans for survivors between SIGTERM and SIGKILL, and the wait in
* between (200ms) sits far inside the cache TTL (2000ms): reading the cache there
* returns the pre-SIGTERM state verbatim, so children spawned since are invisible
* and SIGKILL aims at stale PIDs, guarded only by kill(pid, 0) — which cannot
* detect PID reuse.
*
* It forces a refresh rather than guaranteeing recency: an already-running refresh
* is reused, so the snapshot can predate this call by up to one `ps` runtime. A
* strict postdate guarantee would mean chaining a second `ps` behind every
* in-flight one, which is the fork storm this code exists to avoid.
*
* Bounded by design: waiting forever would freeze killSession before it reaches
* its process-group and tmux fallbacks.
*/
private async getChildPidsFresh(pid: number): Promise<number[]> {
let byParent: ReadonlyMap<number, readonly number[]>;
try {
byParent = await Promise.race([
TmuxManager.refreshProcSnapshot(),
new Promise<never>((_, reject) =>
setTimeout(() => reject(new Error('proc snapshot timeout')), PROC_SNAPSHOT_WAIT_MS)
),
]);
} catch {
console.warn('[TmuxManager] process snapshot did not return in time; using the cached one');
byParent = TmuxManager.procSnapshot?.byParent ?? new Map<number, number[]>();
}
return collectDescendants(pid, byParent, {
onTruncated: (root, cap, reason) =>
console.warn(`[TmuxManager] descendant walk for ${root} hit the ${cap}-${reason} cap; truncating`),
});
}
// Check if a process is still alive
@@ -1882,9 +2248,16 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
return false;
}
// COD-108: an intentional kill/detach must NEVER be auto-revived by the
// remote-reconnect watcher. Guard BEFORE any teardown so a tick that fires
// mid-kill (especially the non-owned DETACH early-return below, where the
// dead local pane would otherwise look reconnectable) sees the guard.
this.guardRemoteReconnect(sessionId);
// TEST MODE: Remove from memory only — NEVER touch real tmux sessions
if (IS_TEST_MODE) {
this.sessions.delete(sessionId);
this.clearRemoteReconnectState(sessionId);
this.emit('sessionKilled', { sessionId });
return true;
}
@@ -1896,6 +2269,40 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
return false;
}
// COD-105 — DETACH-NOT-KILL for NON-owned remote sessions.
//
// When this session was created by ATTACHING a remote tmux session another
// Codeman owns (`remote.owned === false`), closing the tab must NOT propagate
// a remote `tmux kill-session` — that would nuke work the remote's own
// Codeman (or another instance) still relies on. We tear down ONLY the LOCAL
// pane that holds the ssh client: killing the local ssh sends SIGHUP to its
// remote `tmux attach`, which DETACHES (the durable remote session survives).
//
// This early return is the structural guarantee: no code below this point
// (now or in future for owned sessions) can ever issue a remote kill-session
// for a non-owned session. The only `kill-session` we run is on OUR LOCAL
// socket (`this.tmux()` = `tmux -L codeman` on THIS host), which kills the
// local pane — it does NOT reach the REMOTE socket.
if (session.remote && session.remote.owned === false) {
console.log(`[TmuxManager] DETACH (non-owned remote): tearing down local pane only for ${session.muxName}`);
if (isValidMuxName(session.muxName)) {
try {
// Local socket only — detaches the remote session by killing the local ssh pane.
execSync(`${this.tmux()} kill-session -t "${session.muxName}" 2>/dev/null`, {
timeout: EXEC_TIMEOUT_MS,
});
} catch {
// Local pane may already be gone.
}
}
this.lastPaneCount.delete(session.muxName);
this.sessions.delete(sessionId);
this.clearRemoteReconnectState(sessionId);
this.saveSessions();
this.emit('sessionKilled', { sessionId });
return true;
}
// Get current PID (may have changed)
const currentPid = this.getPanePid(session.muxName) || session.pid;
@@ -1904,7 +2311,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const allPids: number[] = [currentPid];
// Strategy 1: Kill all child processes recursively
let childPids = this.getChildPids(currentPid);
let childPids = await this.getChildPidsFresh(currentPid);
if (childPids.length > 0) {
console.log(`[TmuxManager] Found ${childPids.length} child processes to kill`);
allPids.push(...childPids);
@@ -1921,7 +2328,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
await new Promise((resolve) => setTimeout(resolve, TMUX_KILL_WAIT_MS));
childPids = this.getChildPids(currentPid);
childPids = await this.getChildPidsFresh(currentPid);
for (const childPid of childPids) {
if (this.isProcessAlive(childPid)) {
try {
@@ -2000,6 +2407,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
this.lastPaneCount.delete(session.muxName);
this.sessions.delete(sessionId);
this.clearRemoteReconnectState(sessionId);
this.saveSessions();
this.emit('sessionKilled', { sessionId });
@@ -2066,6 +2474,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
} else {
dead.push(sessionId);
this.sessions.delete(sessionId);
this.clearRemoteReconnectState(sessionId);
this.emit('sessionDied', { sessionId });
}
}
@@ -2133,17 +2542,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const [rss, cpu] = psOutput.split(/\s+/).map((x) => parseFloat(x) || 0);
// From the shared snapshot: this runs per session on every stats tick, and a
// pgrep per session was a fork per session per interval.
let childCount = 0;
try {
const childOutput = (
await execAsync(`pgrep -P ${session.pid} | wc -l`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
})
).stdout.trim();
childCount = parseInt(childOutput, 10) || 0;
childCount = (this.childrenByParent().get(session.pid) ?? []).length;
} catch {
// No children or command failed
// No children or snapshot unavailable
}
return {
@@ -2177,17 +2582,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// Step 1: Get descendant PIDs
const descendantMap = new Map<number, number[]>();
const pgrepOutput = (
await execAsync(
`for p in ${sessionPids.join(' ')}; do children=$(pgrep -P $p 2>/dev/null | tr '\\n' ','); echo "$p:$children"; done`,
{
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}
)
).stdout.trim();
// Derived from the ONE snapshot instead of a shell loop that forks a pgrep
// per session — the shape that turned into a fork storm under load.
const byParent = this.childrenByParent();
const childLines = sessionPids.map((p) => `${p}:${(byParent.get(p) ?? []).join(',')}`).join('\n');
for (const line of pgrepOutput.split('\n')) {
for (const line of childLines.split('\n')) {
const [pidStr, childrenStr] = line.split(':');
const sessionPid = parseInt(pidStr, 10);
if (!Number.isNaN(sessionPid)) {
@@ -2337,9 +2737,118 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
this.lastPaneCount.clear();
}
// ── COD-108 remote-session auto-reconnect watcher ─────────────────────────
/**
* Start the remote-reconnect watcher (COD-108). Each tick, for every tracked
* session with `session.remote` whose local pane is DEAD, not intentionally
* guarded, and within its backoff budget, emit `remoteSessionDropped` so the
* session owner reattaches (re-running the idempotent remote command rejoins
* the durable remote tmux session). After the attempt cap, emit
* `remoteReconnectExhausted` once and go quiet.
*
* No-op tick body under `IS_TEST_MODE` (mirrors `startMouseModeSync`): tests
* drive the logic deterministically via {@link runRemoteReconnectTick}.
*/
startRemoteReconnectWatcher(intervalMs: number = DEFAULT_REMOTE_RECONNECT_INTERVAL_MS): void {
if (this.remoteReconnectInterval) {
clearInterval(this.remoteReconnectInterval);
}
this.remoteReconnectInterval = setInterval(() => {
if (IS_TEST_MODE) return;
try {
this.runRemoteReconnectTick(Date.now(), isRemoteAutoReconnectEnabled());
} catch (err) {
console.error('[TmuxManager] Remote reconnect watcher error:', err);
}
}, intervalMs);
}
stopRemoteReconnectWatcher(): void {
if (this.remoteReconnectInterval) {
clearInterval(this.remoteReconnectInterval);
this.remoteReconnectInterval = null;
}
}
/**
* Run ONE watcher tick. Extracted (and given an injected `now`/`enabled`) so
* the reconnect logic is deterministically testable even though the live
* `setInterval` body no-ops under test mode. For each remote session it
* applies the pure {@link decideReconnect} decision and translates the result
* into events + backoff/state transitions. Public for tests + the watcher.
*/
runRemoteReconnectTick(now: number, enabled: boolean): void {
for (const session of this.sessions.values()) {
if (!session.remote) continue;
const sessionId = session.sessionId;
const state = this.reconnectState.get(sessionId);
const action = decideReconnect({
session: {
sessionId,
isRemote: true,
paneDead: this.isPaneDead(session.muxName),
},
state,
guarded: this.reconnectGuard.has(sessionId),
enabled,
now,
});
if (action.kind === 'emit') {
const base = state ?? freshReconnectState();
// Mark in-flight + advance backoff BEFORE emitting so a re-entrant tick
// (or a synchronous listener) can never stack a second reconnect.
this.reconnectState.set(sessionId, { ...advanceBackoff(base, now), inFlight: true });
this.emit('remoteSessionDropped', { sessionId, attempt: action.attempt });
} else if (action.kind === 'exhaust') {
const base = state ?? freshReconnectState();
if (!base.exhaustedEmitted) {
this.reconnectState.set(sessionId, { ...base, exhausted: true, exhaustedEmitted: true });
this.emit('remoteReconnectExhausted', { sessionId });
}
}
// 'skip' → nothing to do.
}
}
/**
* Tell the watcher a reattach attempt for `sessionId` finished. On success,
* reset the backoff so the session is healthy again; on failure, just clear
* the in-flight flag so the next due tick can retry under the existing
* backoff schedule. Called by the session owner after `respawnPane`.
*/
noteRemoteReconnect(sessionId: string, success: boolean): void {
if (success) {
this.reconnectState.set(sessionId, resetReconnectState());
return;
}
const state = this.reconnectState.get(sessionId);
if (state) this.reconnectState.set(sessionId, { ...state, inFlight: false });
}
/**
* Exclude a session from auto-reconnect (intentional teardown). Adds it to the
* guard set and drops any backoff state so a closed/killed tab — especially a
* non-owned remote DETACH — is never auto-revived. Idempotent.
*/
guardRemoteReconnect(sessionId: string): void {
this.reconnectGuard.add(sessionId);
this.reconnectState.delete(sessionId);
}
/** Clear all per-session reconnect + guard state (e.g. when a session is removed). */
clearRemoteReconnectState(sessionId: string): void {
this.reconnectState.delete(sessionId);
this.reconnectGuard.delete(sessionId);
}
destroy(): void {
this.stopStatsCollection();
this.stopMouseModeSync();
this.stopRemoteReconnectWatcher();
this.reconnectState.clear();
this.reconnectGuard.clear();
}
registerSession(session: MuxSession): void {
+33 -16
View File
@@ -40,6 +40,7 @@ interface TranscriptContentBlock {
text?: string;
name?: string;
input?: Record<string, unknown>;
tool_use_id?: string;
content?: string;
is_error?: boolean;
}
@@ -328,10 +329,7 @@ export class TranscriptWatcher extends EventEmitter {
this.handleResultEntry(entry);
break;
case 'user':
// User message means new turn, reset some state
this.state.isComplete = false;
this.state.hasError = false;
this.state.errorMessage = null;
this.handleUserEntry(entry);
break;
case 'system':
// System messages are informational
@@ -360,23 +358,42 @@ export class TranscriptWatcher extends EventEmitter {
this.state.currentTool = block.name;
this.emit('transcript:tool_start', block.name);
} else if (block.type === 'tool_result') {
// Tool completed
const wasError = block.is_error === true;
const toolName = this.state.currentTool;
this.state.toolExecuting = false;
this.state.currentTool = null;
if (toolName) {
this.emit('transcript:tool_end', toolName, wasError);
}
if (wasError && block.content) {
this.state.hasError = true;
this.state.errorMessage = String(block.content).slice(0, 200);
}
this.handleToolResult(block);
}
}
}
}
private handleUserEntry(entry: TranscriptEntry): void {
// A user-authored prompt starts a turn, while Claude tool results also use
// user entries. Reset turn state first, then close any completed tool.
this.state.isComplete = false;
this.state.hasError = false;
this.state.errorMessage = null;
const content = entry.message?.content;
if (!Array.isArray(content)) return;
for (const block of content) {
if (block.type === 'tool_result') {
this.handleToolResult(block);
}
}
}
private handleToolResult(block: TranscriptContentBlock): void {
const wasError = block.is_error === true;
const toolName = this.state.currentTool;
this.state.toolExecuting = false;
this.state.currentTool = null;
if (toolName) {
this.emit('transcript:tool_end', toolName, wasError);
}
if (wasError && block.content) {
this.state.hasError = true;
this.state.errorMessage = String(block.content).slice(0, 200);
}
}
private handleResultEntry(entry: TranscriptEntry): void {
// Result entry indicates completion
this.state.isComplete = true;

Some files were not shown because too many files have changed in this diff Show More