Compare commits

...
Author SHA1 Message Date
Codeman maintainer d26f26fe34 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:18:03 +02:00
Codeman maintainer fa18eeef35 feat: tab action icons on the active tab only, middle-click closes tabs
Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.

Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.

The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer a9f26bd03a feat: fixed-width tab hover with sliding title, pop-out button now opt-in
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).

The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer 8dc8b164a7 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:48:12 +02:00
Ark0N 2524759655 Merge pull request #229 from Lint111/feat/keyboard-viewport-settle
fix(mobile): coalesce keyboard viewport settling
2026-08-08 12:51:04 +02:00
Codeman maintainer 1f164bc8d2 fix(mobile): only arm the viewport settle on a real keyboard transition
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 12:06:08 +02:00
lior 0a1439b1e9 test(mobile): make the coalescing test actually exercise the settle path
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.

Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.

Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.

Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).
2026-08-08 08:41:04 +03:00
lior 66abe6c70a fix(mobile): coalesce keyboard viewport settling 2026-08-08 08:13:58 +03:00
Codeman maintainer fa1700da5b chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:38:35 +02:00
Ark0N aed1e59ee3 Merge pull request #227 from Ark0N/fix/scrollback-205-round2
fix(terminal): scrollback round 2 for #205 (re-pull downgrade guard, PageUp fallback, CLI version probe retry)
2026-08-08 01:36:03 +02:00
Ark0N 7f6d18b398 Merge pull request #226 from christianhaberl/fix/input-loss-on-failed-delivery
fix(api,ws): an input whose delivery fails can be retried instead of being lost
2026-08-08 01:31:10 +02:00
Ark0N 52571c7fd4 Merge pull request #225 from christianhaberl/fix/bound-the-process-tree-walk
fix(mux): bound the process-tree walk — unbounded pgrep recursion can take a machine down
2026-08-08 01:31:00 +02:00
Ark0N cb95a8562c Merge pull request #224 from christianhaberl/fix/raw-writehead-drops-security-headers
fix(http): raw writeHead routes drop every header the security hook set
2026-08-08 01:30:47 +02:00
Codeman maintainer 3cb7e30636 fix(ui): scope the wheel-opt-out tooltip's paging fallback to Claude
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:29:48 +02:00
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00
Claudia 9d27cc0bab docs: merge the stacked doc comments the previous commits left behind
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.

- write() had two: the original description with @param and @example, then a
  @returns-only block added on top, which dropped the params and examples from
  hover. Merged into one. The @returns wording is also honest now — write() still
  discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
  and its declaration, leaving that function undocumented on hover and the doc
  attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:29:46 +02:00
Claudia 84132d3025 fix(ws): pin the withheld ACK with a test, and correct the changeset
Two blockers from the pre-submission gate, both reproduced before fixing.

1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
   code says the opposite in as many words ("NOT an error response,
   deliberately"), the commit message says response codes are unchanged, and the
   test asserts the 200. It was a leftover sentence from an earlier iteration that
   would have shipped into the CHANGELOG announcing an API contract change that
   does not exist — and errorCode values are SemVer-relevant per
   docs/versioning-policy.md.

2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
   master left all 9 tests green, while the commit message sells "plus the whole
   WebSocket path" as part of the fix. Three tests added against the real WS
   route — ACK on delivery, ACK withheld and seq re-opened when the write did not
   land, and a deduplicated frame still ACKed so the client can drop it. Verified
   the other way round: with ws-routes.ts reverted, the middle one fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:52:44 +02:00
Codeman maintainer cc163792e5 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:43:47 +02:00
Ark0N 6f1ff17ccc Merge pull request #223 from Ark0N/fix/scrollback-shell-alt-screen
fix: terminal scrollback overhaul for shell and CLI sessions (#205)
2026-08-07 13:42:47 +02:00
Codeman maintainer f262b8cb69 feat(terminal): gentler glide start and fractional wheel accumulation
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:36:27 +02:00
Codeman maintainer 5f2b491d99 feat(terminal): ease-out smooth scrolling for the local wheel path
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:30:34 +02:00
Codeman maintainer c067167dbc fix(terminal): take the wheel in capture phase; xterm's scroller is deaf after reset
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.

The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.

Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:03:17 +02:00
Codeman maintainer ad2ca9b575 docs: record the #205 scrollback mechanisms and the shipped fix plan
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:42 +02:00
Codeman maintainer a7a1cef3d6 fix(session): probe the Claude CLI version over ssh for remote sessions
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:41 +02:00
Codeman maintainer a1d7ec02e9 fix(terminal): forward touch scrolls to the CLI transcript on mobile
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:25 +02:00
Codeman maintainer dfa43928af docs: record the scrollback analysis and its measurements for #205 2026-08-07 04:33:15 +02:00
Codeman maintainer adbb74cd5a fix(terminal): keep the CLI's input box pinned when scrolling with the wheel
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.

_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:

- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
  history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
  REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
  furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
  non-empty row, so any session with trailing blank rows was left off-bottom
  and every later wheel event went local without the user ever scrolling.

Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.

Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
2026-08-07 04:27:16 +02:00
Codeman maintainer eb8d11ffc3 fix(terminal): restore shell scrollback, recover history lost to tmux repaints
Four fixes for the scrollback reports in #205 (plus its follow-up comment).

1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
   ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
   (\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
   to claude/codex/gemini, so it reached the browser verbatim. In the alternate
   buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
   and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
   which readline receives as shell history navigation. Both reported symptoms,
   one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
   modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
   own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
   the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
   never forwards a pane's alt-screen toggles to its client, it repaints;
   captured from a real attach, vim/less/htop emit zero.

2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
   onto tmux's history, and tmux repaints the pane rectangle instead of emitting
   linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
   scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
   same 60 lines emitted slowly added all 60. Scrolling up at the top now
   re-pulls the full tmux scrollback and holds the user's place. Verified
   end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.

3. The full-scrollback replay was gated on a single "first load after page load"
   flag, which whichever session auto-selected consumed, so every other tab
   started with one visible frame. Now tracked per session.

4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
   per notch) scrolled one line where Chrome scrolls four or five, and capped
   the forwarded SGR report at one tick. Line and page deltas are now converted,
   and a pure horizontal swipe no longer falls through to a phantom -1.

Analysis and measurements: docs/scrollback-issues-analysis.md
2026-08-07 04:06:54 +02:00
Claudia ebfcac6ad1 fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.

When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.

- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
  is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
  client redelivers.
- `Session.write()` reports whether it reached a PTY.

Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.

What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.

9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:36:33 +02:00
Claudia 2e69e28e71 fix(mux): bound the process-tree walk — it can take a machine down
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.

Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.

- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
  with a single-flight guard. Async matters: under the same procfs pathology,
  `execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
  which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
  visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
  both reporting when they truncate. Pure so the regression tests can exercise the
  shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
  re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
  would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
  bounded, so a wedged `ps` cannot stop killSession from reaching its
  process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
  fresh; a truncated table would make whole subtrees invisible to the kill path.

13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:34:47 +02:00
Claudia 1a32e63765 fix(http): raw writeHead routes lost every header the security hook set
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.

The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.

Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.

Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:31:50 +02:00
Codeman maintainer d41f28bc14 docs(docker): warn that a plain agent-image rebuild keeps stale CLIs
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.

That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.

Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.

No changeset: docs-only, rides the next release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 09:02:28 +02:00
Codeman maintainer 322f21ef9f docs(extending): scope the "no sandbox" claim, point at Docker cases
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:

- Integration code cannot be sandboxed by Codeman because Codeman never
  launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
  documented isolation story.

Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:46:22 +02:00
Codeman maintainer c2d973cb2d docs: link the integration guide from both READMEs, fix three inaccuracies
Adds a pointer to docs/extending-codeman.md at the end of the API section in
README.md and README.zh-CN.md, so the guide is reachable from where people
read about endpoints rather than only from CLAUDE.md.

Reading the README's programmatic guide alongside the new page surfaced three
errors in it, all now fixed:

- POST /api/sessions/:id/input takes `useMux`, not `useScreen`. The latter is
  a legacy name that no longer appears in the schema.
- The page told integrators to send `\r` to submit. With `useMux: true` the
  server delivers text and Enter as two separate writes (writeViaMux does
  send-keys -l then send-keys Enter), so appending `\r` is wrong.
- "Unwrap the envelope" was incomplete: a few legacy GETs put the payload at
  the top level, so the advice is now `body.data ?? body`.

Also cross-references the README's programmatic guide, which covers the
in-session case (CODEMAN_MUX, CODEMAN_API_URL, CODEMAN_SESSION_ID,
CODEMAN_HOOK_SECRET_FILE) that the new page deliberately does not duplicate,
and documents the optional clientId/seq exactly-once fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:34:58 +02:00
Codeman maintainer 84e31c0ee1 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:31:02 +02:00
Codeman maintainer 0d0b772619 feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.

Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
  mode 'antigravity' died on command-not-found. agy is not on npm, so it
  gets its own installer step. --dir /usr/local/bin is load-bearing: the
  default $HOME/.local/bin resolves to root's home at build time and is
  unreachable by the `agent` user the container runs as. Verified inside
  codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
  ~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
  present like the other CLI buttons, with a cyan identity matching the
  toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
  it as a satisfying AI CLI, and recommends it over Gemini in the install
  hints. Detection only, no new auto-install path.

Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
  opencode/codex/gemini when the code has included antigravity for a
  while, said "all three modes", and omitted ANTIGRAVITY_ from the env
  prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
  remote-sessions' RemoteCommandMode were all stale.

Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.

test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.

Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.

isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:22:29 +02:00
Ark0N bfff20a093 Merge pull request #216 from shenlvkang-collab/fix/response-viewer-brief-format
fix(web): align brief Response Viewer formatting
2026-08-06 07:22:13 +02:00
codeman-local b982c5d0e0 fix(web): align brief response viewer formatting 2026-08-06 10:16:15 +08:00
Codeman maintainer f50c922240 docs: add extending-codeman.md, the third-party integration guide
Codeman has no plugin runtime by design: running third-party code inside
the process that spawns agents, on a server people expose over a tunnel,
would trade away the security posture that is a reason to use it. But it
already has four extension seams that work from any language with nothing
installed, and they were undocumented.

Documents web tabs (render your own UI as a tab), the SSE event channel
(react when an agent needs you), the HTTP API plus the codeman CLI (drive
it from a script), and hook events. Every endpoint, schema field, event
name and header in the page was read from source and then verified against
a running instance, including the localhost-only CORS behavior and the SSE
framing the example depends on.

Also corrects a stale line in CLAUDE.md: it claimed the HTTP/SSE API was
internal/unstable, which contradicts docs/versioning-policy.md, where the
API under /api/v1 was finalized as part of the stable surface for the 1.0
cut. No new stability commitment is made here; the page makes an existing
one discoverable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 02:03:39 +02:00
Codeman maintainer de5b048c3f docs(vm): VM cases plan + Apple virtualization stack reference
Two design/reference docs for the planned native-macOS VM isolation tier
("VM cases"), a location overlay on cases in the same shape as Docker and
remote-SSH cases, never a sixth SessionMode. Nothing is implemented; both
docs are marked PLANNED and are blocked on macOS 27 GA.

- vm-cases-plan.md: the Codeman-side design and phased plan. Swift helper
  CLI, DiskImageKit base + per-case overlay, sessions riding the existing
  remote-SSH machinery, VirtioFS workspace at the same absolute path, and
  seeded credentials, each mirroring an established Docker-cases rule.

- vm-subsystem-apple-stack.md: what the Apple stack actually provides,
  measured on the 27 beta rather than inferred from the WWDC session. Of
  note: the 2-concurrent-macOS-VM cap is a kernel quota (refused at 39%
  free RAM, so more hardware does not help), DiskImageKit has no flatten
  API so exports must ship the layer chain, and a macOS guest renders
  nothing without an attached view in an unlocked host session.

No credentials, hostnames, tailnet addresses or account names in either
file; every host/guest reference is a placeholder.

Also joins a table row that a stray blank line had split off into its own
malformed table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 01:28:17 +02:00
Codeman maintainer 12a5f5919e chore: version packages 2026-08-05 22:36:51 +02:00
Codeman maintainer ecd3f3f32a harden(history): exclude automated transcripts by SDK shape, not by "not cli"
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.

Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.

Test fails against the pre-fix line and passes after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 21:47:06 +02:00
Ark0N c19d884a51 Merge pull request #215 from timkjr/fix/past-sessions-history-quality
fix(history): three Past Sessions data-quality bugs (automated-session noise, cross-contaminated previews, blank restart-heavy rows)
2026-08-05 21:44:47 +02:00
Ark0N 22e77a1827 Merge pull request #214 from timkjr/fix/mobile-overview-run-gating
fix(mobile): gate the phone overview's run picker on CLI availability
2026-08-05 21:44:42 +02:00
Ark0N b641560040 Merge pull request #203 from shenlvkang-collab/contrib/claude-viewer-session-pin
fix(web): pin the Claude response viewer to the pane's own conversation
2026-08-05 21:44:37 +02:00
timkjr 8300c15cbd fix(history): entrypoint detection was first-field-wins, plus a two-tier head read
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).

Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 09f5f28017 docs(test): correct an overclaiming comment in the tail-fallback regression test
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 251706be3b harden: scope entrypoint detection to message lines, add fallback coverage
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:

1. extractTranscriptEntrypoint() scanned any line containing the
   substring "entrypoint", not specifically the first "type":"user"/
   "type":"assistant" message line (unlike its sibling
   extractFirstUserPrompt, which does scope to type). A transcript
   that started under an older Claude Code version (no entrypoint
   field) and got resumed under a newer one mid-conversation could
   pick up the field from a much later message than the true first
   one, misattributing the session's origin. Scoped it to match.

2. Added a regression test proving the tail-read fallback still
   engages correctly when bookkeeping accumulation exceeds even the
   new 128KB head window, not just the 16KB it previously blanked at.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 18b473f0e4 fix(history): raise the transcript head-read window to fit restart bookkeeping
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).

Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 a2aed38073 fix(unified-sessions): stop the firstPrompt workingDir backfill from cross-contaminating history rows
COD-140's backfill was meant to cover live/persisted rows whose Codeman
id doesn't match an on-disk transcript UUID, guessing from the newest
transcript in the same workingDir as a last resort. It was also firing
for pure history rows whose OWN transcript scan already ran (and
genuinely found nothing, e.g. an oversized first message) -- those got
silently backfilled with the newest OTHER session's opening line from
the same directory. Not a blank row, but actively wrong: old sessions
displayed a completely unrelated (often today's live) conversation's
first prompt as if it were their own.

Skip the workingDir guess for any item that already has its own
'history' source -- it already had a real, direct attempt. Rows with
no history source at all (their transcript isn't linked/scanned under
their own id yet) still get the guess, matching the original intent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 e888c65c52 fix(history): exclude non-interactive (SDK-driven) transcripts from Past Sessions
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.

Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 1ea39de650 fix(mobile): gate the phone overview's run picker on CLI availability
MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu is a separate,
hardcoded duplicate of the toolbar's #runModeMenu (mobile-overview.js
is a newer feature that mirrors the toolbar menu's look/behavior
rather than reusing its render), so it never picked up #201's
isCliAvailable() gating and offered every backend regardless of what
the server actually has installed.

Gate it the same way: skip an entry unless isCliAvailable(mode),
shell always exempt. Added functional + static regression tests
mirroring the toolbar menu's own test pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:15 -05:00
Codeman maintainer e2a644997e chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 09:01:45 +02:00
Ark0N cd5a101626 Merge pull request #213 from Ark0N/feat/file-viewer-edit-mode
File Viewer: edit mode for text files (edit + save in the viewer)
2026-08-05 09:00:22 +02:00
Codeman maintainer 4ea781c80f feat(file-viewer): edit mode for text files (edit + save in the viewer)
Closes #212. The file-preview overlay can now edit workspace text files in
place, phone-first: agent writes a file, you review it in the viewer, tweak
two lines, save, tell the agent to continue.

Backend (file-routes.ts, policy in src/config/file-editing.ts):
- GET file-content?edit=1: read-for-edit that never truncates (a truncated
  buffer must never become an edit buffer), 512KB cap (413 over it), and
  returns the sha256 hash + detected EOL the client echoes back on save.
- PUT /api/sessions/:id/file-content: edit-in-place only, with no O_CREAT
  anywhere in the handler. Confinement matches the read path (realpath +
  workspace boundary + ownership via findSessionOrFail), plus sensitive-path
  and attachment-guard blocklists, a .git subtree deny, and an extension
  allowlist (svg and env deliberately excluded). Optimistic concurrency via
  baseHash: mismatch is a 409 unless force. Writes are wx-temp + fchmod +
  fsync + rename, closing the validate-then-write TOCTOU window.
- Corruption guards: NUL sniff + UTF-8 round-trip compare (refuses binary
  and latin-1), and server-side EOL re-application so a textarea's LF
  normalization cannot rewrite every line of a CRLF file.
- Plain reads gain an additive editable flag the UI keys the button off.

Frontend (panels-ui.js + overlay markup/styles):
- Edit button on editable text previews; textarea editor with Save/Cancel,
  dirty indicator, discard-confirm on cancel/close, and a conflict dialog
  that offers overwrite (force) when the file changed on disk mid-edit.
- Phone: full-bleed window sized by --app-height so the editor and Save bar
  track the OS keyboard; 16px editor font (iOS zoom guard); no autofocus.
- zh-CN strings for the new chrome.

Tests: pure policy unit tests plus a route suite that deliberately does NOT
mock node:fs. It runs against a real temp workspace so symlink escapes,
write-through of in-workspace symlinks, mode preservation, CRLF round-trip,
409/force, and the no-create property are exercised for real. Also verified
end to end on an isolated beta instance: 39-check curl matrix, Playwright
desktop flow (real clicks and typing, bytes asserted on disk, live conflict
with an external rewrite), and a 393px phone profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 08:44:47 +02:00
Codeman maintainer d9123de9eb feat(terminal): Ctrl+C copies the selection, interrupts when nothing is selected
Closes #211. Copying from the terminal only worked through the browser
context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown,
so the muscle-memory copy failed silently and read as "no copy-paste at all".

With a selection, Ctrl+C now copies it, toasts, clears the selection and
sends nothing to the PTY. With no selection it falls through unchanged, so
the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never
falls through: an explicit copy that interrupts a running agent because the
selection happened to be empty would be a footgun.

Three details that keep the interrupt safe:

- The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the
  no-selection path returns true WITHOUT preventDefault. xterm calls the
  custom handler before its own cancel(), so returning false alone does not
  cancel the event; the copy path therefore calls preventDefault explicitly,
  or the browser would run its native copy on top of ours.
- copy-selection is a full registry entry (rebindable and disableable in App
  Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the
  same trick command-palette uses: the generic capture loop preventDefaults
  every match it dispatches, which would cost the user the interrupt key.
- The gate is keydown-only, since the custom handler also runs for keypress
  and keyup.

Copy goes through _copyText (Clipboard API, then hidden-textarea +
execCommand) rather than raw navigator.clipboard, because install.sh's LAN
option serves plain HTTP where navigator.clipboard is undefined; the
fallback steals focus, so the terminal is refocused afterwards.

Tests: test/terminal-copy-selection.test.ts pins the gate and the
SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real
key presses in chromium and asserts on the clipboard plus the bytes xterm
emitted (browser-driven, so excluded from test:ci like the other Playwright
suites). Verified manually on an isolated beta instance before landing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 02:38:57 +02:00
Codeman maintainer 1e5f6c8ee1 chore: version packages 2026-08-05 02:10:55 +02:00
Codeman maintainer 5d2899907e fix(cli-gating): gate the tunnel button instead of deleting it, and cover antigravity
Follow-up to #200 and #201, which gate the welcome buttons and the run-mode
dropdown on whether the CLI is actually installed. Four corrections:

1. #200 also DELETED the Cloudflare Tunnel welcome button and the QR widget
   outright. Its rationale is right (offering a tunnel where cloudflared is not
   installed is a bad default) but the conclusion overshoots: the welcome QR is
   the whole scan-to-connect-from-your-phone flow, and deleting it left a large
   block of live tunnel code in settings-ui.js driving elements that no longer
   existed. Both are restored and the button is gated on cloudflared, which is
   what the stated rationale actually asks for. New cloudflared-resolver.ts
   mirrors the CLI resolvers, and TunnelManager now shares its search path so
   the button and the spawn can never disagree about where cloudflared lives.

2. Antigravity was missing from the run-mode gating, the one run mode LEAST
   likely to be installed. It slipped past because #201 predates it. Covered
   now, plus a static test that fails if a sixth mode reaches the dropdown
   without being gated, so the next one cannot slip the same way.

3. The per-surface fetches are replaced by the injected availability object
   already used for the Codex settings tab, so the codebase has one mechanism
   rather than two. The status routes buy nothing as a gating source: every
   resolver memoizes its PATH probe server-side, so a fetch is exactly as stale
   as an injected value while costing a round trip every time the dropdown opens
   and leaving the welcome buttons to flicker in after paint. The routes
   themselves stay, including the /api/claude/status that #200 adds.

4. Unknown availability now reads as AVAILABLE for run buttons. Both PRs hid the
   button on a failed fetch, so a blip left a working install with nothing to
   click; a genuinely missing CLI only ever produced an error toast. The Codex
   settings TAB keeps the opposite default, since hiding it costs nothing.

The dropdown query is also scoped to the menu: `.run-mode-option` is the class
the saved-dashboard and history rows use too, and a document-wide querySelector
would have found whichever came first in the DOM.

Fixes a latent environment-sensitivity in 816d900 while here: the index-title
test asserted the template was untouched apart from the title, which held only
on a machine with no codex installed.

Verified end-to-end against a real server on an isolated instance+socket, with
Playwright: gemini/codex hidden and claude/opencode/antigravity/shell shown,
matching this host, tunnel button back, Codex settings tab still hidden, no
console errors. Full test:ci sweep green (3902 tests).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:53:43 +02:00
Codeman maintainer 8facd5e7e7 Merge pull request #201 from timkjr/pr/gate-run-mode-dropdown
fix(run-mode): gate dropdown entries on CLI availability
2026-08-05 01:40:47 +02:00
Codeman maintainer b54094a4c8 Merge pull request #200 from timkjr/pr/gate-gemini-drop-tunnel-button
fix(welcome): gate CLI welcome buttons on actual availability
2026-08-05 01:40:44 +02:00
Codeman maintainer 2b89f35599 fix(shell,remote-ssh): allowlist the login flags, and keep only CRASHED remote panes
Follow-up to #209 and #210. Both land a real fix (a pane that is a login shell
picks up /etc/profile and the per-user PATH entries an ssh remote command never
sees, which is what was failing agent CLIs with exit 127). Three corrections:

1. `-i -l` is no longer hardcoded onto the resolved shell. That path ultimately
   comes from the passwd entry, which is user data and can name anything, and a
   shell that rejects an unknown flag exits on the spot: nushell, elvish and xonsh
   take neither flag, so a user with one of those in passwd would have gotten a
   dead pane on arrival, which is exactly the #208 failure #209 builds on top of.
   loginShellArgs() applies them only to the POSIX-family shells verified to
   accept both, and a test really launches every allowlisted shell present on the
   machine rather than trusting the set. csh/tcsh are excluded deliberately: tcsh
   honors -l only when it is the ONLY flag.

2. `remain-on-exit on` -> `failed`, moved LAST in the tmux command chain. `on`
   keeps the pane after a CLEAN exit too, so typing `exit` in a remote shell
   stranded a dead pane, the session outlived it, and the next launch's `-A`
   reattached to that corpse: "Pane is dead (status 0)" instead of a shell,
   permanently, on the DEFAULT path. Verified against a real tmux, as was the
   fix: `failed` tears the session down on status 0 and keeps the pane on 127
   with the "command not found" still on screen, which is the case #210 wanted.
   It is last because tmux aborts the remaining commands of a `\;` sequence once
   one errors (also verified) and `failed` needs tmux >= 3.2 on the REMOTE host;
   leading, a rejection there would have silently dropped status/mouse/prefix/
   escape-time/window-size along with it.

3. `$SHELL` -> `"${SHELL:-/bin/sh}"`, via one shared remoteLoginShellCommand()
   helper instead of the string being rebuilt in tmux-manager as well.

Also corrects the rationale both PRs carried: a tmux pane already hands the shell
a tty, so it was interactive all along ($- contains i for a bare /bin/bash in a
pane) and ~/.bashrc was always being sourced. `-l` is the flag doing the work.

End-to-end verified, not just unit-tested: the emitted remote pane command was
run through all three quoting layers under a minimal sshd-style PATH with the
CLI installed only on a login-shell PATH entry, and it resolved and launched the
CLI with its arguments intact and a space-containing remote path preserved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:40:37 +02:00
Codeman maintainer ee670c38f6 Merge pull request #210 from timkjr/fix/remote-ssh-login-shell
fix(remote-ssh): route shell + agent CLIs through a real interactive login shell
2026-08-05 01:34:23 +02:00
Codeman maintainer ad57109dcf Merge pull request #209 from timkjr/fix/shell-login-shell
fix(shell): launch shell tabs as an interactive login shell
2026-08-05 01:34:22 +02:00
Codeman maintainer c15b8345b5 fix(history): never treat the empty split segment as a directory name
Follow-up to #202. The dotdir decode landed there was reachable only when
nothing else matched first, and in the greedy half it was not reachable at all.

decodeProjectKey() splits the project key on '-', so the '/.' that the encoder
collapses leaves an EMPTY segment behind. Both loops offered that empty string
as a candidate directory name, and isDir(current + '/' + '') stats current + '/',
which always succeeds. So the empty segment matched unconditionally:

  - backtracking half: ~/.sib resolved to "/home/x//sib" whenever a non-dot
    sibling ~/sib existed (wrong directory, and a doubled slash that then fails
    every string comparison against session.workingDir). Without a sibling it
    only backtracked out by luck.
  - greedy half: that loop is shortest-match-first, so the empty candidate
    matched on the FIRST iteration and set matched=true, leaving #202's dotdir
    branch permanently dead there.

An empty string is never a real path component, so skip it in both loops. The
unmatched tail then has to handle the empty segment too, or it would append a
bare '/' and re-introduce the '//' path it just stopped producing; it now emits
the dotdir guess instead, which is what the encoder implies.

Regression test asserts both halves: the dotdir wins over the non-dot sibling,
and the result never contains '//'. Verified it fails on #202 as merged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:34:16 +02:00
Codeman maintainer 45ae9f4064 Merge pull request #202 from timkjr/fix/dotdir-workingdir-decode
fix: decode dotdir working directories in history session scanning
2026-08-05 01:32:47 +02:00
Codeman maintainer 816d900857 feat(settings): show the Codex CLI tab only where codex is installed
Both settings on the App Settings "Codex CLI" tab (bypass approvals, animated
status effects) are handed to `codex` at launch, so on an instance where the
binary does not resolve the tab offers choices nothing can act on. Gate it on
availability instead.

renderIndexHtml injects window.__codemanCodexAvailable, mirroring the existing
gesture-availability flag, and settings-ui.js hides the tab button when it is
absent. Injected rather than fetched on modal open so the tab cannot flicker in
and back out; isCodexAvailable() memoizes its PATH probe, so the per-render cost
is nil. Installing codex later needs a restart, exactly like the
/api/codex/status route that already backs the Run menu. Solo popups skip the
probe since they have no settings modal.

Only the tab BUTTON is toggled. The panel already carries
.modal-tab-content.hidden unless it is the selected tab and openAppSettings()
always reopens on Display, so an unreachable button keeps the panel unreachable.
The inputs stay in the DOM and are still populated and read back on save, so a
user without codex cannot silently wipe the codex preferences of an instance
that has it. Animations stay off by default for new local Codex sessions.

Verified in a browser on this host, which has no codex: the flag is absent, the
Codex tab is hidden while the other tabs are unaffected, and saving App Settings
with the tab hidden leaves codexAnimationsEnabled/codexDangerouslyBypassApprovals
untouched. With the flag forced on, the tab appears, its panel opens, and
toggling the visible slider persists. The openAppSettings coupling test was
checked to fail when the call is removed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:14:53 +02:00
Ark0N ddc267c6ff Merge pull request #181 from Lint111/agent/split-codex-animations
feat(codex): make terminal animations configurable
2026-08-05 00:20:38 +02:00
timkjrandClaude Sonnet 5 d66007053b fix(shell): launch shell tabs as an interactive login shell
Shell-mode sessions resolve to an absolute shell path (issue #208's
fix) but launch it bare, with no -i/-l flags. Without those, the
spawned shell runs as a non-interactive child of the non-interactive
`bash -c` that launches the pane, so it never sources ~/.zshrc or
~/.bashrc — silently dropping aliases, PATH additions, and tool init
(zoxide, nvm, etc.).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:12:28 -05:00
timkjrandClaude Sonnet 5 f470f3a4e7 fix: decode dotdir working directories in history session scanning
decodeProjectKey() couldn't recover a dotdir path (e.g. ~/.codeman) from
Claude Code's encoded project-key names: the encoder maps both '/' and
'.' to '-', so the decoder's candidate joins never matched a hidden
directory on disk. It silently fell through to bare $HOME instead,
which corrupted workingDir for any resumed session under a dotdir case
(observed on ~/.codeman itself: history rows and state.json recorded
"/home/timkjr" instead of "/home/timkjr/.codeman").

Add a dot-prefixed candidate to both the backtracking decoder and its
greedy fallback so a leading empty split segment (the signature of a
literal '.' in the original path) is retried as a hidden directory.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 17:11:51 -05:00
Codeman maintainer cb3eecad9b Merge branch 'master' into pr181 2026-08-05 00:04:43 +02:00
Ark0N db24fc6d7e Merge pull request #180 from Lint111/agent/split-preserve-active-launch
fix(sessions): preserve active terminal during launches
2026-08-04 23:53:14 +02:00
Codeman maintainer 292ba2c775 fix(sessions): route antigravity launches through the ownership helpers
runAntigravity() landed on master after this branch was cut, so it kept the
exact pattern the rest of this PR removes: terminal.clear() plus direct
writeln into whatever session happened to be active. Merging master in
surfaced it, leaving one of six run modes still wiping the active session's
xterm on launch.

Also adds regression coverage that can actually see the bug. The existing
test drives the three helpers directly, so it stays green even when a run*()
function is reverted to writing at the terminal itself: reverting
runClaude()'s call site keeps all 16 tests passing. The new static guard
scans session-ui.js and fails if any run*() body touches
this.terminal.clear/writeln, which catches a regressed call site and would
have caught runAntigravity on its own. A second unit test covers the
home-screen path that nothing exercised: with no active session, launch
progress must still clear and render in the terminal.

Verified in a browser against a live instance. With a session active,
runShell() and runAntigravity() leave its terminal untouched (clear() calls:
0, writes: 0) and emit one info toast; on master the same run wipes the
session's marker text. The session-less home screen still clears and writes
exactly as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 23:47:17 +02:00
timkjrandClaude Sonnet 5 e803186dfe fix(remote-ssh): route claude/opencode/codex/gemini/antigravity through login shell
remain-on-exit (previous commit) preserved dead remote panes instead of
destroying them, which revealed the real failure: `exec claude`/`exec
opencode` ran under ssh's non-interactive, non-login remote-command
shell, which only sees sshd's minimal default PATH — not the ~/.zshrc
PATH entries where these CLIs actually live (e.g. ~/.local/bin,
~/.opencode/bin). Wrap them in `$SHELL -i -l -c '<cmd>'`, mirroring the
fix shell mode already had.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 16:41:27 -05:00
timkjrandClaude Sonnet 5 474efd9023 fix(remote-ssh): use remote user's real shell, keep dead panes alive
Remote shell-mode sessions hardcoded 'exec bash -l', ignoring the
remote user's actual login shell. sshd sets $SHELL from the remote
user's /etc/passwd entry, so 'exec $SHELL -i -l' launches their real
shell (zsh, fish, etc.) with rc files sourced, same fix as the local
shell-mode launch.

Also set remain-on-exit on the remote tmux session. It was only ever
set on the local socket, so if the remote command exited for any
reason -- even something transient -- tmux destroyed the pane, window,
and (being the only session) the whole remote server, tearing down the
local ssh attach along with it and leaving no trace to diagnose. The
local pane saw this as an instant clean exit, and reconnect's -A then
created a fresh session, which could repeat as a flap loop with no
evidence surviving between attempts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-04 16:39:29 -05:00
Codeman maintainer 03bb40c78a Merge branch 'master' into pr180 2026-08-04 23:37:59 +02:00
timkjr 3ea1ea28f0 fix(welcome): gate Claude and Opencode buttons on CLI availability too
Extends the Gemini gating from bb7fb9e to the other welcome-screen
buttons that had the same problem: shown unconditionally even when the
underlying CLI isn't installed.

- Add isClaudeAvailable() (claude-cli-resolver.ts) and GET
  /api/claude/status, mirroring the existing opencode/codex/gemini
  resolvers and status endpoints.
- Opencode already had a working /api/opencode/status the welcome
  screen just wasn't checking; wire it up the same way.
- Refactor loadGeminiAvailability() into a shared
  _loadCliAvailability(buttonId, statusUrl) helper instead of
  duplicating the fetch/try-catch three times.

Run-mode dropdown entries (Opencode/Codex) are intentionally left
unconditional here — follow-up PR.
2026-08-04 16:22:21 -05:00
timkjr 62008fb408 fix(welcome): gate Gemini button on availability, drop unconditional tunnel button
- Remove the always-visible Cloudflare Tunnel welcome button and QR
  widget; offering it regardless of whether cloudflared is installed
  is a bad default.
- Hide the "Run Gemini" welcome button by default and only show it
  when /api/gemini/status reports available:true, via new
  loadGeminiAvailability() called from showWelcome().
2026-08-04 16:21:55 -05:00
timkjr 660b320a67 fix(run-mode): gate dropdown entries on CLI availability
Follow-up to the welcome-screen gating (#200): the run-mode dropdown
(gear menu next to Run) had the same problem — Claude/Opencode/Codex/
Gemini entries were always shown regardless of whether the CLI is
actually installed, so picking one could spawn a session that
immediately errors out.

- Add _refreshRunModeAvailability() (session-ui.js), called each time
  the dropdown opens; hides entries whose /api/<cli>/status reports
  unavailable.
- Shell is intentionally never gated (no external CLI dependency).

Depends on isClaudeAvailable()/GET /api/claude/status, which don't
exist on upstream/master yet — duplicated here from #200 so this PR
is self-contained and independently mergeable. Once #200 lands this
branch should be rebased onto master, which will collapse the
duplicate cleanly.
2026-08-04 16:21:17 -05:00
Codeman maintainer 529d8fa8ea chore: version packages 2026-08-04 23:09:00 +02:00
Codeman maintainer 19af37977a fix(ui): stop dropping the session name typed in the options modal
Two independent ways a tab description could be typed in and silently lost.

1. Session Options modal (deterministic). The Session Name input saves on
   blur, and every autosave handler in the modal bails on a null
   editingSessionId. closeSessionOptions() cleared that id BEFORE hiding the
   modal, and hiding it is what blurs the input, so the save always ran too
   late and returned early. Escape and backdrop-click lost the name with no
   PUT at all; only the X button worked, because mousedown blurs the input
   before the click handler runs. Fix: blur the focused modal field first,
   then clear the id. That also covers the auto-compact prompt, which saves
   on change and had the same fate.

2. Right-click inline rename (racy). The _inlineRenameActive guard from #81
   sits in renderSessionTabs() (the scheduler) and _fullRenderSessionTabs(),
   but not in _renderSessionTabsImmediate() (the debounced executor). A
   render queued in the ~100ms before the rename opened still fires and the
   incremental branch rewrites .tab-name's innerHTML, destroying the input
   mid-keystroke: it commits a truncated name, or, if it lands before the
   first keystroke, closes the rename so everything typed after goes
   nowhere. Fix: guard the executor too. finishRename() re-renders on both
   commit and cancel, so a render dropped there is picked back up.

Verified end-to-end against a live server on an isolated instance: all three
modal close paths now persist the name, and the rename input survives a
render mid-typing. Both regression tests were checked to fail with their fix
reverted; the render one was vacuous at first because the synthetic tab sat
on <body> instead of inside #sessionTabs, so it now builds the tab in the
real container.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 17:41:02 +02:00
Codeman maintainer 23f258a85d chore: version packages
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).

Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.

Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 15:02:46 +02:00
Codeman maintainer aa4f423d8a chore(gitignore): ignore the root pr/ working dir
pr/ holds machine-local promo drafts that are never meant for git. Anchored with
a leading slash so it matches only the root dir, matching the /public entry below
it, rather than swallowing any nested pr/ elsewhere in the tree.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 13:56:52 +02:00
Codeman maintainer 1b1057d9e0 chore: version packages 2026-08-04 13:01:28 +02:00
Codeman maintainer 26cbbe0dcb feat(cli): Antigravity run mode
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.

- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
  resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
  env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
  injected via socket-scoped `tmux setenv`, never on the spawn command line, so
  the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
  colours. `runAntigravity()` routes remote/docker cases through
  `POST /api/quick-start` and skips the local status probe for them.

Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 12:59:40 +02:00
Codeman maintainer 1113d34ca8 feat(ui): opt-in entrance animations for tabs, terminal pane, agent windows and connection lines
All OFF by default (the `legacy` theme), so an untouched install behaves exactly
as before and every mark/apply hook short-circuits on its first line. Opt in via
App Settings > Appearance > Entrance Animations; per-surface control and a live
preview lab at ?animlab=1.

Surfaces and styles:
- Tabs: slide, pop, crt, unroll, boot, flip. A batch launched together cascades
  by a configurable stagger.
- Terminal pane: crt, boot, wipe, slide, fade.
- Agent windows: fly (the pre-existing tab-to-window flight, still the default),
  crt, materialize, unfold, beam, pop.
- Connection lines: draw, packet, fade.

Three constraints drove the design:

1. Tabs and connection lines are DESTROYED mid-animation on every re-render:
   _fullRenderSessionTabs() replaces the strip's innerHTML and
   _updateConnectionLinesImmediate() does `svg.innerHTML = ''`, both of which run
   constantly while sessions and agents spawn. Each is tracked by id and
   re-applied to the fresh element with a NEGATIVE animation-delay so it resumes
   at the same offset instead of restarting or snapping. Verified on the real
   path: a forced rebuild mid-draw resumed at -0.243s.

2. Terminal-pane styles animate transform/opacity/clip-path ONLY. xterm's
   FitAddon derives rows+cols from getComputedStyle(parent).width/height, the
   untransformed layout box, so transforms are invisible to it; animating
   width/height/padding would have resized the PTY. Verified by forcing
   fitAddon.fit() eight times mid-animation: dimensions held at 178x38.

3. A window entrance that transforms also moves the rect its connection line
   aims at (crt drifts it 109px, pop 81px). `beam` animates opacity/filter only
   (0px drift) so its line can draw toward a stable target; the others refresh
   the lines on animationend.

Also fixes: an agent window spawning hidden (its agent belongs to a background
tab) is display:none, so its animation never runs and animationend never fires,
which left the entrance class and its inline custom property stuck on the window
permanently. Hidden windows now skip the entrance entirely.

Styles persist to their own codeman:*Anim localStorage keys, keeping them
per-device without touching the .strict() SettingsUpdateSchema.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 12:31:52 +02:00
shenlvkang-collab ab7a703e90 fix(web): keep the viewer's conversation anchor across a Codeman restart
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.

Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.

Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
2026-08-03 21:22:33 +08:00
Codeman maintainer 8a31f10b7d chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-03 14:09:17 +02:00
Ark0N 2891ae0d6d Merge pull request #178 from Lint111/agent/split-notification-noise
fix(notifications): quiet lifecycle hook noise
2026-08-03 14:07:54 +02:00
Ark0N 17b86b1007 Merge pull request #177 from Lint111/agent/split-transcript-tool-results
fix(transcripts): complete tools from user results
2026-08-03 14:05:38 +02:00
shenlvkang-collabandClaude Opus 5 73315bc351 fix(web): pin the Claude response viewer to the pane's own conversation
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.

Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:53:42 +08:00
lior 94e3aae57d feat(codex): make terminal animations configurable 2026-07-29 03:39:07 +03:00
lior 0a039239e4 fix(sessions): preserve active terminal during launches 2026-07-29 03:30:38 +03:00
lior 67eb5b43eb fix(notifications): quiet lifecycle hook noise 2026-07-28 23:20:01 +03:00
lior 4a4720cb62 fix(transcripts): complete tools from user results 2026-07-28 23:18:10 +03:00
129 changed files with 21488 additions and 728 deletions
+3
View File
@@ -68,6 +68,9 @@ design-explorations/
# Artifacts that should not be tracked
test-results/
tmp/
# Machine-local working files (never meant for git). ANCHORED so only the root
# dir matches.
/pr/
# Root `public` (a symlink to scripts/remotion/public — local artifact). ANCHORED
# with a leading slash so it does NOT also match src/web/public (a bare `public`
# would swallow the whole web UI source dir and silently un-stage any new asset
+281
View File
@@ -1,5 +1,286 @@
# aicodeman
## 1.13.0
### Minor Changes
- Agent wait primitives, the Codeman agent skill, a fix for hooks dying silently on HTTPS installs, and the tab-strip UX improvements from the previous batch.
**Agent wait primitives (new API surface, the reason this is a minor).** Three bounded long-polls let an agent driving Codeman from a shell block instead of poll:
- `GET /api/v1/sessions/:id/wait` blocks until a lifecycle signal fires (`until=stop,idle,working,blocked,exit`, `fresh=1` to require a new transition).
- `GET /api/v1/sessions/:id/wait-output` blocks until a literal substring appears in the session's output (`match=`, `nocase=`, `from=now|buffer`; never regex, by design).
- `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input` (send-and-wait) registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer.
Shared semantics: a timeout is HTTP 200 with `wait.timedOut: true` (callers loop over short waits; tunnels cut idle connections), timeouts are clamped to [1s, 600s] and echoed back as `wait.timeoutMs`, all three nest the result under `data.wait`, and `status`/`limitPaused` ride along. `stop`/`blocked` exist for `claude` mode only: requesting them explicitly elsewhere is a 400, the default set silently narrows and echoes what it waited on. Capacity caps (16 waiters per session, 128 process-wide) answer 409/429, waiter slots release on client hang-up, and shutdown resolves parked waiters instead of stranding them. Bounds are operator-tunable via `CODEMAN_WAIT_*` env vars.
Reliability details that came out of three verification rounds: a worker that dies inside its tmux pane is now detected at the mux layer (pane-death probe, ~750ms cache, a 3s watcher for waits already parked), so a corpse answers `exit` instead of `idle` and send-and-wait rolls back its dedup seq when the write went nowhere; output matching normalizes charset-designation escapes (a stock bash prompt's `ESC ( B` no longer breaks `match=tnode:`) and holds back partial escapes at chunk boundaries, so matches straddling PTY chunks are found.
**Codeman agent skill (`skills/codeman`).** A packaged skill that teaches an agent running inside a Codeman session to drive the API safely: guard preamble (refuses outside `CODEMAN_MUX=1`, resolves credentials from the data dir `.env` or the install's service definition), self-protection (`is_self` prefix check in both directions), readiness for claude workers (composer-first, trust dialog as bounded fallback), send-and-wait loops that cannot report a never-submitted prompt as success, marker-synchronized shell flows, fan-out patterns, and cleanup discipline. Ships in the npm package via the `files` entry.
**Hooks were dying silently on every HTTPS install (bug fix).** The generated hook curls lacked `-k`, so on `--https` installs (self-signed cert) every hook event (`stop`, `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `teammate_idle`, `task_completed`) failed TLS verification and the failure was swallowed, taking respawn's definitive idle signals with it. Hooks are now generated with `curl -sk`, and a staleness detector regenerates the on-disk hook config of already-created cases the next time a session starts in them. Relatedly, `CODEMAN_API_URL` is no longer exported with a guessed `http://localhost:3000` fallback (wrong scheme on HTTPS installs); it is omitted unless the server has stamped the real URL, so in-session guards fail closed.
**Tab strip (from the previous batch, reported by christianhaberl):** action icons (kill/pop-out) now appear on the active tab only, middle-click closes a tab, tab hover uses a fixed width with a sliding title instead of resizing the strip, and the pop-out button is opt-in (default off).
**Docs.** `docs/api-reference.md` gained the full long-polling contract (signals by mode, readiness, what the matcher sees, response discriminators); `docs/extending-codeman.md` and the README carry verified copy-paste orchestration recipes; `docs/architecture-invariants.md` records the load-bearing ordering, liveness, and edge-triggered-signal invariants. Net +163 tests (4300 passing in the CI sweep).
## 1.12.2
### Patch Changes
- Codex input fixes: all four bugs reported by @DodgyBadger traced to one root cause (the zero-lag local-echo overlay buffering keystrokes until Enter, which starves codex's per-keystroke composer) and fixed in terminal-ui.js:
- Slash command picker never appeared in codex sessions (#222): the "/" sat in the overlay until Enter, so codex never saw it. Codex-mode sessions now use plain PTY echo (same branch as shell), so the picker pops and live-filters as you type.
- Arrow keys dead while typing, backspace dead after Ctrl+Backspace (#218): arrows were forwarded to a still-empty composer while typed text sat pending, and after a control-char flush the overlay swallowed every backspace. Codex bypasses the overlay entirely now; the shared overlay branch (claude/gemini/opencode) additionally flushes pending text on composer nav keys, then hands the session to pass-through until Enter/Ctrl+C, and forwards backspace instead of swallowing it when the overlay has no state.
- Pasting displaced the typed prompt (#219): bracketed pastes (xterm terminal.paste with DECSET 2004 active) were forwarded without flushing pending typed text, so the paste landed first. The shared branch now flushes typed text first and delays the paste sequence by 80ms, because codex's paste-burst handling drops keystrokes that arrive in the same PTY read as a bracketed paste (verified against codex 0.147.0 at the byte level).
- Long prompts overflowed the bottom of the screen (#220): long typed prompts existed only in the overlay DOM so codex never grew its composer; with plain PTY echo the composer grows and rewraps normally.
Verified end to end against a real codex 0.147.0 TUI driven by a headless browser: the pre-fix build reproduces all four bugs, the fixed build passes 17/17 assertions. New CI test file test/local-echo-codex-gating.test.ts (41 tests) pins the nav-key classifier, per-mode overlay gating, the flush helper, and pass-through routing. Known upstream limitation: Ctrl+Backspace deletes one character, not a word (xterm.js sends 0x08; word-delete needs kitty CSI-u encoding that xterm.js 6.0.0 cannot emit).
Mobile keyboard viewport settling fixes by @Lint111 (#229): coalesce keyboard viewport settling so rapid visualViewport resize events during keyboard show/hide no longer thrash the terminal fit, and only arm the settle logic on a real keyboard transition instead of every viewport resize.
## 1.12.1
### Patch Changes
- Terminal scrollback fixes, round 2 of issue #205. A Claude pane's local buffer is hollow (tmux keeps no history for a repaint-mode pane), and both retest reports traced back to that fact. The scroll-to-top full-history re-pull now refuses to rewrite the terminal when the capture holds less than the browser already does, so it can no longer delete history mid-scroll on iPhone (a refused session also re-fetches far less often). When wheel-forwarding is unavailable on a Claude session (version probe failed, CLI older than 2.1.187, or the "Wheel Scrolls Local History" opt-out) and there is no local scrollback to scroll, wheel and touch now page the CLI's own transcript via coalesced PageUp/PageDown instead of doing nothing. The `claude --version` probe no longer caches a failed run for the server's lifetime (one timed-out probe used to silently disable wheel-forwarding on every device until restart); failures retry with backoff. Every scroll gesture now logs a one-line `[scroll]` routing decision to the browser console for direct diagnosis, and the opt-out setting's tooltip explains that the paging fallback is Claude-only (Codex has none).
- 2e69e28: Bound the process-tree walk that could take a machine down.
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup`
under WSL, none returned while the walk kept spawning more — ~13,000 `pgrep`
processes stuck in D-state out of ~39,000 total, load average above 13,000,
recoverable only by restarting WSL.
Now: one `ps` snapshot, breadth-first with a visited set, a depth cap and a node
cap, in a pure module (`proc-tree.ts`) that the regression tests exercise
directly. The snapshot is refreshed asynchronously, and the kill path forces a
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.
- ebfcac6: An input whose delivery fails can be retried instead of being lost for good.
Both input paths recorded the `(clientId, seq)` pair as applied and acknowledged
the frame _before_ knowing whether the write had landed — the POST route because
its mux write is fire-and-forget, the WebSocket handler because it ACKed
unconditionally. When the write then failed, the client dropped the frame from its
durable queue and the server rejected the retry as a duplicate: the reliable
delivery layer was guaranteeing exactly-once delivery of something that had never
been delivered.
The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so
the client redelivers. `Session.write()` reports whether it reached a PTY at all
instead of silently swallowing the data.
Response codes are unchanged: a session can legitimately have no PTY yet (created
but not started), so turning that into a failure status would be a contract change
of its own.
Note this does not remove the root cause: the POST still answers 200 before the
mux write is attempted, so a client that treats any 2xx as final still cannot
learn about that failure. Closing that would mean awaiting the tmux child in the
request path.
- 1a32e63: Routes that answer with `reply.raw.writeHead()` no longer drop the headers the
security hook set.
`writeHead` writes straight to the Node response and bypasses Fastify's header
store, so everything the `onRequest` hook granted was silently lost — including the
`Access-Control-Allow-Origin` it emits for localhost origins, and the
`X-Content-Type-Options` / `X-Frame-Options` / CSP headers. A localhost page could
therefore call every other `/api` endpoint cross-origin while its EventSource
failed CORS.
Affects `GET /api/events` and the three raw-writing routes in `file-routes.ts`
(`file-raw`, `tail-file`, `download`).
## 1.12.0
### Minor Changes
- Terminal scrollback overhaul (issue #205), fixing every reported scroll failure across shell and CLI sessions, desktop and mobile:
- Shell, OpenCode and Antigravity sessions finally have working scrollback: tmux's own client-side alternate-screen switch is stripped for tmux-backed sessions (narrow strip: alt-screen toggles only, keeping `clear`'s 3J and mouse DECSETs), so xterm stays in the normal buffer instead of a scrollback-less alt buffer where the wheel turned into shell history cycling and touch scrolling did nothing. Direct-PTY fallback sessions are untouched so fullscreen apps (vim/less/htop) keep the alt screen there.
- The wheel listener now runs in capture phase and owns the scroll: xterm's internal vscode-style viewport scroller consumed wheel events whenever local scrollback existed (and goes deaf entirely after a tab switch or replay resets the terminal), which silently killed wheel forwarding, made scrolling break after reload/tab switches, and let the CLI's input box scroll away. Local scrolling goes through buffer-level scrollLines and keeps working after resets; mouse-tracking apps and alternate-buffer sessions are passed through untouched.
- Wheel AND touch scrolling now forward to the CLI's own transcript for Codex and Claude 2.1.187+, at any scroll position (the viewport snaps home first), so the input box stays pinned on desktop and phones alike. Shift+wheel and the "Wheel scrolls local history" setting still pin local scrollback.
- Smooth scrolling: local wheel scrolling glides with an ease-out animation (fractional line accumulation, so slow trackpad drags track the finger instead of running ahead).
- Full tmux history on demand: the full-scrollback replay is now per session instead of once per page load, and scrolling up at the top of the buffer re-pulls the complete tmux history, recovering everything tmux's repaint bursts or tab switches removed from the browser's copy.
- Firefox wheel speed: wheel deltas are normalized by deltaMode (Firefox reports line units, previously read as pixels and slowed ~4x).
- Remote SSH Claude sessions now probe the CLI version over ssh (same connection options and login-shell wrapper as the real launch), so wheel forwarding works for them too instead of silently staying off.
Docs: scrollback analysis and fix plan recorded in docs/, architecture invariants updated (strip flavors, capture-phase wheel ownership, per-session full-history replay); docker agent-image rebuild warning and integration-guide link fixes from the preceding docs commits.
## 1.11.2
### Patch Changes
- Make Antigravity (`agy`) a first-class CLI everywhere, and stop presenting Gemini CLI as a consumer product now that it is enterprise-only.
Antigravity was already wired into the session layer, schemas, run-mode menu and remote/Docker command maps, but the surfaces around it were never updated. Gemini keeps full support; Antigravity now sits beside it.
Fixes:
- **Docker cases with `mode: 'antigravity'` were broken.** `docker/agent.Dockerfile` installs its CLIs from npm, and `agy` is not an npm package, so the binary was never in the image and the container died on command-not-found. It now gets its own installer step. The `--dir /usr/local/bin` flag is load-bearing: the installer's default `$HOME/.local/bin` resolves to root's home at build time and would be unreachable by the `agent` user the container runs as. Note the binary is roughly 190MB, making it the largest layer in the image, so rebuild with `node scripts/build-agent-image.mjs` when convenient.
- **Welcome screen** gained a "Run Antigravity" action, gated on `agy` being present like the other CLI buttons, styled with the same cyan identity as the toolbar run button and run-mode dot.
- **`install.sh`** now detects `agy` (search paths mirroring `antigravity-cli-resolver.ts`), counts it as a satisfying AI CLI so an Antigravity-only box is not told it has none, and recommends it instead of Gemini in the install hints.
Documentation corrections where it had become factually wrong: `architecture-invariants.md` described `isExternalCliMode()` as opencode/codex/gemini when the code has included antigravity for some time, said "all three modes", and omitted `ANTIGRAVITY_*` from the env-prefix allowlist row; the `agentType` enum in `cron-guide.md`, `SessionMode` in `cron-discovery.md`, and `RemoteCommandMode` in `remote-sessions.md` were all stale.
Also updated both READMEs (five CLIs, Gemini marked enterprise-only), the `antigravity` npm keyword, and comment drift in eight places. Test coverage added for the new welcome button.
Antigravity stores its state under `~/.gemini/antigravity-cli/` rather than a `~/.antigravity` directory, so the existing `.gemini` Docker credential seed already covers it. That is now recorded in a code comment so no dead configuration gets added later.
- b982c5d: Keep the brief Response Viewer output inside the same message card and Markdown wrapper used by the full conversation view, so opening the viewer without clicking More preserves the same readable formatting.
## 1.11.1
### Patch Changes
- fix(history): Past Sessions data quality, and gate the phone run picker on CLI availability
**Past Sessions data quality (#215).** Three bugs in the transcript scanner behind
the Cmd+K Session Manager and the phone overview's PAST SESSIONS list:
- Automated/SDK-driven transcripts (CI review bots and other tooling, which Claude
Code stamps with a non-`cli` `entrypoint`) were listed alongside real interactive
sessions even though they were never resumable. They are now excluded. Detection
scans every entrypoint-bearing message rather than stopping at the first, so a
transcript that began under an older Claude Code build and only later picked up a
non-`cli` entrypoint is no longer wrongly hidden.
- A resumed session could show a same-directory sibling's preview text as its own.
The `workingDir` backfill in `mergeUnifiedSessions()` now only ever applies to rows
that have no history entry of their own, so it can no longer overwrite a row's real
content with another conversation's.
- Sessions restarted many times accumulated enough bookkeeping lines to push the real
first prompt past the scanner's 16KB head-read window, leaving a blank row. The read
is now two-tier: 16KB first, escalating to 128KB only when that was not enough, which
is both correct and cheaper than reading 128KB unconditionally (measured on a real
transcript tree: 36% fewer bytes read, roughly 17.5% faster than the unconditional
version). Also restores the tail-read fallback for a file whose head read failed
outright (for example `EMFILE` while scanning hundreds of files), which had been
silently dropping the session from history.
Follow-up hardening on top of the above: the automated-transcript exclusion now
blocklists the SDK entrypoint shape (`sdk`, `sdk-cli`, `sdk-py`) instead of allowlisting
the exact value `cli`. Because the check hides rows, an allowlist failed closed on any
value Claude Code has not shipped yet: a future rename of the interactive entrypoint,
or a second interactive host, would have blanked the entire Past Sessions list with
nothing in the UI to explain it. An unrecognized automated entrypoint now costs a few
noisy rows instead, which is the annoyance this filter set out to fix rather than a
broken feature.
**Phone overview run picker (#214).** The "C" logo home screen's Run picker listed all
six backends regardless of what was installed, so tapping an uninstalled one produced a
failed launch instead of the entry simply not being offered. It is now gated on
`isCliAvailable()` exactly like the desktop toolbar's run-mode dropdown (shell exempt,
since it has no external CLI dependency and keeps the menu from ever being empty). The
picker is a hardcoded duplicate of the toolbar menu rather than a shared render, which
is why it never picked up the earlier gating work; a test now asserts that every mode
the picker offers is gated, so a newly added backend cannot silently drift again.
- 73315bc: fix(web): stop the Claude response viewer from following another session's conversation
The viewer re-derived a pane's live conversation by taking the newest
`~/.claude/history.jsonl` entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` run in the user's own terminal, so the eye followed whichever of those
was typed into last — and the adoption was written back to the session, so the
mispin persisted. Entries are now credited to a pane only when they land within
10s of that pane's own Enter and no other pane on the cwd submitted closer, the
same last-submit correlation the Codex locator already uses.
That correlation also has to survive a restart. `start()` resets
`claudeSessionId` to the launch id even when re-attaching to a mux session whose
CLI has since moved on via `/clear`, so a recovered pane pointed the viewer at
its pre-`/clear` transcript — and with the anchor itself living only in memory,
nothing corrected it until the user happened to type again. `lastSubmitAt` is
now persisted in `SessionState` and restored on boot recovery, so the viewer
re-derives the live conversation on its first poll.
## 1.11.0
### Minor Changes
- Two user-facing features since 1.10.0.
**Terminal: Ctrl+C copies the selection, interrupts when nothing is selected** (#211). Copying from the terminal previously worked only through the browser context menu: xterm turns Ctrl+C into 0x03 and cancels the keydown, so the muscle-memory copy failed silently and read as "no copy-paste at all". With a selection, Ctrl+C now copies it, shows the "Copied to clipboard" toast, clears the selection and sends nothing to the PTY; with no selection it falls through unchanged, so the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never interrupts. The shortcut is a normal registry entry (`copy-selection`), so it can be rebound or disabled in App Settings, and disabling it restores plain always-interrupt Ctrl+C. Copy goes through the Clipboard API with a hidden-textarea fallback, so it also works on plain-HTTP LAN installs.
**File Viewer: edit mode for text files** (#212). The file-preview overlay can now edit workspace text files in place, phone-first: `GET /api/sessions/:id/file-content?edit=1` reads for edit without the 500-line preview truncation (saving a truncated buffer would silently delete the rest) and returns a sha256 hash plus the detected EOL; `PUT /api/sessions/:id/file-content` saves. Edit-in-place only: there is no O_CREAT anywhere in the handler, so "never create, never delete" is structural. Confinement inherits the read path (realpath plus workspace boundary, ownership scoping) and adds sensitive-path and attachment-guard blocklists, a `.git/` subtree deny, and an extension allowlist (`svg` and `env` deliberately excluded). Optimistic concurrency is by content hash, so a file changed on disk mid-edit returns 409 with an overwrite option rather than clobbering. Writes are atomic (`wx` temp, fchmod, fsync, rename) which closes the validate-then-write TOCTOU window and cannot follow a pre-existing symlink. Binary and latin-1 content are refused via a NUL sniff plus a UTF-8 round-trip compare, and EOL is re-applied server-side so a textarea's LF normalization cannot turn a two-line edit of a CRLF file into a whole-file diff.
## 1.10.0
### Minor Changes
- Codeman 1.10.0.
**Every surface that offers a CLI now checks the CLI is actually there** (#200, #201). The welcome-screen run buttons, the run-mode dropdown and the App Settings "Codex CLI" tab used to be shown unconditionally, so picking one on a box without the binary spawned a session that errored out immediately. All of them now gate on a single server-injected availability object covering Claude, OpenCode, Codex, Gemini, Antigravity and cloudflared, so nothing flickers in after paint and the dropdown costs no round trips to open. Shell is never gated, which is what keeps the menu non-empty on a box with nothing installed, and unknown availability reads as available so a stale page can never leave a working install with nothing to click. Adds `isClaudeAvailable()` and `GET /api/claude/status`, the one CLI that had no availability check despite being the default. The Cloudflare Tunnel welcome button and its scan-to-connect QR are gated on `cloudflared` rather than shown regardless.
**Shell and remote-SSH sessions now launch a real login shell** (#209, #210). Local shell tabs match what tmux itself does for a pane with no `default-command`, picking up the `/etc/profile` and `/etc/profile.d/*` entries a systemd `--user` service never sourced. On remote SSH, `claude`/`opencode`/`codex`/`gemini`/`agy` are routed through the remote user's interactive login shell, fixing agent CLIs that silently failed with "command not found" because ssh's remote-command execution sees only sshd's minimal default PATH and not the `~/.local/bin` or `~/.opencode/bin` entries where those CLIs actually live. Shell mode uses the remote user's real shell instead of hardcoded bash. The login flags are applied only to shells verified to accept them, so an exotic passwd entry (nushell, elvish, xonsh) cannot produce a dead pane on arrival.
**A crashed remote pane is kept for diagnosis** (#210), which is how the PATH failure above was found: it previously destroyed the pane, the window and the whole remote session on exit, tearing the local ssh attach down with it and leaving a flap loop with no evidence. Scoped to `remain-on-exit failed`, so a clean `exit` still tears the session down and only a non-zero exit strands anything, and applied last in the tmux command chain so a remote tmux older than 3.2 cannot drop the other session options with it.
**Resumed sessions under a hidden directory get the right working directory** (#202). Claude Code's project-key encoder maps both `/` and `.` to `-`, and the decoder could not reconstruct a dot-prefixed component, so every session under `~/.codeman` (or any project nested beneath any dotdir) silently resolved to bare `$HOME`. The wrong `workingDir` then propagated into `state.json` and everything trusting it: CLAUDE.md lookup, paste-image directory, subagent and image watchers. A same-named non-dot sibling could also produce a doubled-slash path that failed every later string comparison.
**Launching a session no longer wipes the terminal you are looking at** (#180). All six run modes route through the shared ownership helpers instead of clearing and writing into whatever session happened to be active, Antigravity included.
**Codex terminal animations are configurable** (#181), and the App Settings "Codex CLI" tab appears only where the `codex` binary resolves, since both settings on it are handed to `codex` at launch.
## 1.9.9
### Patch Changes
- Two bug fixes.
**Plain shell sessions could not start when the server process had no `SHELL` (#208).** The tmux pane command for `mode: 'shell'` was the literal string `$SHELL`. That string is embedded in the `bash -c "..."` argument of the `respawn-pane` line, which is run through `/bin/sh -c`, so it was expanded by the _server_ process's shell against the _server_ process's environment rather than inside the pane. Containers and system-level systemd units do not set `SHELL`, so it expanded to nothing and the pane command ended in a dangling `&&`, giving `bash: -c: line 1: syntax error: unexpected end of file` and a pane that died instantly (status 2) while tmux session creation still reported success. The shell is now resolved in Node (`$SHELL`, then the passwd entry, then `/bin/bash`, `/bin/zsh`, `/bin/sh`), requiring an absolute path to an executable and skipping `nologin`-style stubs, then shell-quoted. Only local shell sessions were affected: agent CLI modes emit a real command, and Docker/remote-SSH cases already used a literal `exec bash -l`.
**A session name typed into the tab options could be silently dropped.** Two independent paths. In the Session Options modal, the Session Name input saves on blur while every autosave handler bails on a null `editingSessionId`, and `closeSessionOptions()` cleared that id before hiding the modal (hiding is what blurs the input), so the save always ran too late; Escape and backdrop-click lost the name with no PUT at all, and only the X button worked because mousedown blurs first. The focused modal field is now blurred before the id is cleared, which also covers the auto-compact prompt. Separately, the right-click inline rename could be destroyed mid-keystroke: the `_inlineRenameActive` guard was missing from `_renderSessionTabsImmediate()`, so a render queued just before the rename opened still rewrote the tab name's innerHTML, committing a truncated name or closing the rename outright. The debounced executor is now guarded too.
## 1.9.8
### Patch Changes
- **Fixed: sessions failed to start on macOS with `Error: posix_spawnp failed.`** (issues #6 and #204)
`node-pty@1.1.0` publishes its macOS prebuilt helper as `prebuilds/darwin-<arch>/spawn-helper` with mode 0644, i.e. no execute bit. macOS launches every PTY through that helper, so a stock install failed on every session start. The bug is macOS-only: `spawn-helper` is a mac-only gyp target and node-pty ships no Linux prebuild, so Linux always compiles a correctly-permissioned helper from source.
The previous fix chmodded only `build/Release/spawn-helper`, which on macOS does not exist (the prebuild is used, so node-gyp never runs), and it derived that path from `require.resolve('node-pty')`, landing on `<pkg>/lib/build/Release/...`. It was a no-op on every platform.
- New `scripts/fix-node-pty.mjs` (also `npm run fix:node-pty`) chmods every `spawn-helper` it finds, in `build/Release`, `build/Debug` and each `prebuilds/*/`, then verifies the result by actually opening a PTY. A `require()` alone passes on a broken install, because the helper is only touched at spawn time.
- `postinstall` no longer force-rebuilds node-pty from source on Node 22+. That step needed Xcode command line tools, cost 30-120s on every install, and deleted the `prebuilds/` tree before compiling, so a Mac without a compiler was left with no working binary at all. A rebuild now happens only when the chmod plus spawn probe still fails, and the prebuilds tree is backed up and restored around it.
- New `spawnPtyWithHelperRepair()` (`src/utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts`, so an install that is already broken repairs itself on the first failed spawn and retries in-process instead of showing a dead session. Unrelated spawn errors are rethrown untouched; a second failure carries the `npm run fix:node-pty` hint.
- `scripts/fix-node-pty.mjs` is now in the published `files` list, so global npm installs get the repair too.
- Direct-PTY Claude spawns use the resolved absolute binary path (new `getClaudeBinaryPath()`) instead of the bare name `claude`, so a CLI installed outside the server's PATH still launches.
Verified end to end on macOS 26.4 arm64: a stock `npm i` reproduces `posix_spawnp failed.`, and after the fix the same install spawns a PTY successfully with the prebuilds preserved.
**Added: phone home screen (session overview)**
Under 430px the "C" logo now opens a session overview (current sessions, past sessions, spaces) instead of the welcome overlay: on a small screen "which session needs me" beats "how do I start one". Rows resume a session in place, and "New session here" goes through the normal quick-start path so remote and Docker cases keep their routing. Per-device setting `mobileOverviewEnabled` (phones only, default ON) in App Settings. Tablet and desktop are unchanged.
**Added: guided Tailscale setup in `install.sh`**
The network-access prompt is now 3-way: Tailscale, LAN, or local-only. The Tailscale path binds loopback and walks through installing Tailscale, logging in, the operator grant, the tailnet HTTPS-certificates toggle, and `tailscale serve --bg <port>`, then verifies the result end to end with curl. That gives HTTPS on a real certificate with no app password and no `0.0.0.0` bind, which is also what PWA install and web push need. `install.sh tailscale` retrofits it onto an existing install, and `CODEMAN_TAILSCALE=1` presets the choice. Serve state is detected from `tailscale serve status --json`; the installer never runs `tailscale serve reset` and never touches serve mappings other than 443 to Codeman's port. README and `docs/security-architecture.md` updated to match.
**Docs**: replaced a real tailnet hostname with placeholders in `docs/web-tabs-fixes-plan.md`.
**xterm-zerolag-input**: npm description and keywords only, no code change.
## 1.9.7
### Patch Changes
- Antigravity run mode, plus opt-in entrance animations.
**Antigravity CLI backend (#207).** Antigravity (`agy`) joins Claude Code, shell, OpenCode, Codex and Gemini as a sixth session backend, following the same pluggable-resolver pattern: `utils/antigravity-cli-resolver.ts` resolves the CLI and `GET /api/antigravity/status` reports availability and path. `ANTIGRAVITY_*` is added to the `ALLOWED_ENV_PREFIXES` allowlist so env overrides stay CLI-scoped rather than blanket-forwarded. Like the other external CLIs it requires tmux with no direct PTY fallback, because secrets are injected through socket-scoped `tmux setenv` and never on the spawn command line. The UI gains a Run-dropdown entry, an agent-type option, an `ag` tab badge and toolbar colours; `runAntigravity()` routes remote and docker cases through `POST /api/quick-start` and skips the local status probe for them.
**Entrance animations (opt-in, OFF by default).** Optional animations for the four things that appear when work starts: session tabs, the terminal pane a session's CLI runs in, floating agent windows, and the connection lines tying a window back to its parent tab. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. Choose a look in App Settings > Appearance > Entrance Animations (per-device, stored in localStorage rather than the settings payload); `?animlab=1` opens a per-surface picker with a live preview that fakes tabs, a pane, a window and a line so styles can be compared without spawning sessions.
Three implementation notes worth knowing if you touch this: tabs and connection lines are destroyed mid-animation on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML, `_updateConnectionLinesImmediate()` clears the SVG), so both are tracked by id and re-applied to the fresh element with a negative `animation-delay` that resumes rather than restarts them; terminal-pane styles animate transform, opacity and clip-path only, because xterm's FitAddon derives rows and columns from the untransformed layout box and animating width or height there would resize the PTY; and window styles that transform also move the rect their connection line aims at, which is why the `beam` style animates opacity and filter only.
Also fixes an agent window spawning hidden (its agent belongs to a background tab): being `display:none` it never ran its animation, so `animationend` never fired and the entrance class plus its inline custom property stuck to the window permanently. Hidden windows now skip the entrance entirely.
## 1.9.6
### Patch Changes
- Two fixes from community PRs (thanks @Lint111):
- fix(transcripts): complete tools from user-entry results (#177). Claude transcripts record tool requests in assistant entries but commonly carry their results in user-role entries; the transcript watcher only completed tools from the older assistant-entry path, so Codeman could keep showing a tool as running after it had finished. The watcher now recognizes `tool_result` blocks in user entries, ends the active tool state, and emits `transcript:tool_end` with the correct tool name and error status. Watcher tests also moved from fixed sleeps to condition-based `vi.waitFor` assertions.
- fix(notifications): quiet lifecycle hook noise (#178). Notification preferences move to schema version 5: the drawer-only "Response complete" (stop) default is now off, and the migration disables only the legacy drawer-only shape, preserving any explicit browser, audio, or push delivery the user opted into. Teammate-idle and task-completed hooks now map to the existing opt-in subagent categories instead of the broadly enabled idle/stop alerts, so normal agent activity no longer floods the drawer. Local and server-hydrated preferences are normalized through the same migration path (server hydration used to revive the retired default on fresh browsers), and the notification storage key now uses the stable handheld identity so an unfolded foldable keeps its mobile defaults and storage key (tablets and desktops unaffected).
## 1.9.5
### Patch Changes
+26 -15
View File
@@ -47,7 +47,7 @@ The production server caches static files for 1 year, `immutable` (`maxAge: '1y'
## COM Shorthand (Deployment)
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI + documented env vars are public; the HTTP/SSE API, on-disk state, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI, documented env vars, **and the HTTP/SSE API under `/api/v1`**: endpoint paths, response envelope, `errorCode` values and SSE event names are public/stable; on-disk state, internal TS modules, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Third-party integration surfaces are documented in `docs/extending-codeman.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
When user says "COM":
@@ -74,13 +74,13 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.9.5 (must match `package.json`)
**Version**: 1.13.0 (must match `package.json`)
## Project Overview
Codeman is a Claude Code session manager with web interface and autonomous Ralph Loop. Spawns Claude CLI via PTY, streams via SSE, supports respawn cycling for 24+ hour autonomous runs.
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, node-pty, xterm.js. Supports Claude Code, OpenCode, Codex (OpenAI), and Gemini (Google) CLIs via pluggable CLI resolvers (`SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`).
**Tech Stack**: TypeScript (ES2022/NodeNext, strict mode), Node.js, Fastify, node-pty, xterm.js. Supports Claude Code, OpenCode, Codex (OpenAI), Gemini (Google, enterprise-only since Google's June 2026 consumer cutover), and Antigravity (`agy`, Google) CLIs via pluggable CLI resolvers (`SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity'`).
**TypeScript Strictness** (see `tsconfig.json`): `noUnusedLocals`, `noUnusedParameters`, `noImplicitReturns`, `noImplicitOverride`, `noFallthroughCasesInSwitch`, `allowUnreachableCode: false`, `allowUnusedLabels: false`.
@@ -102,7 +102,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Test coverage | `npm run test:coverage` |
| Dead-code sweep | `npm run knip` (config in `config/knip.json`, passed via `--config`) |
| Rebuild gesture overlay | `npm run build:gesture` (esbuild `packages/gesture-control/src/codeman/entry.ts` → `src/web/public/gesture/gesture-codeman.js`; commit the result) |
| Build the docker agent image | `node scripts/build-agent-image.mjs` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`/`--no-cache`) |
| Build the docker agent image | `node scripts/build-agent-image.mjs --no-cache` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`). ⚠ **Always `--no-cache`** — a plain rebuild re-uses the cached `npm install -g` layer and silently keeps the CLIs frozen at their original versions, which once shipped a BROKEN codex while reporting success. See `docs/docker-cases.md` |
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
@@ -122,14 +122,15 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`)
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` env vars** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.)
- **`envOverrides` flow `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` / `ANTIGRAVITY_*` env vars** — Set via `POST /api/sessions { envOverrides }`, stored on `Session._envOverrides`, exported by `tmux-manager.buildEnvExports()` at spawn time, persisted in `SessionState.envOverrides`. **Do NOT** write these to `<case>/.claude/settings.local.json` — that's the old path and creates UI/disk drift. (`GOOGLE_*` is the deliberately-broad Vertex-AI namespace for Gemini — see Multi-CLI prefix discipline.)
- **Effort is NOT an env var** — never carry effort as `CLAUDE_CODE_EFFORT_LEVEL`: the env var hard-locks effort and blocks in-session `/effort` switching (incl. ultracode). It flows as the dedicated `effort` payload field → `Session._effort` → `claude --effort <level>` for regular levels incl. `max` (the settings `effortLevel` key is `enum(["low","medium","high","xhigh"]).catch(undefined)` — `max` gets SILENTLY dropped there), or `claude --settings '{"ultracode":true}'` for ultracode (rejected by `--effort`). Both are soft defaults the user can override anytime. Legacy env-var entries are auto-migrated by the Session constructor and unset from tmux sessions in `applyEnvOverrides()`. See `buildEffortCliArgs()` in `session-cli-builder.ts`, tests in `test/effort-injection.test.ts`
- **Model choice flows via `settings.local.json`, NOT `--model` or env** — the App Settings **Claude Model** picker (`claudeModel` in `settings.json`) is read by `session-ui.js` at session create (wins over the legacy 1M-Opus toggles `opusContext1m`/`opusContext1mEnabled`), sent as the `modelOverride` payload field, and `updateCaseModel()` (`hooks-config.ts`) writes/deletes the `model` key in `<case>/.claude/settings.local.json`. This is the intended exception to the envOverrides rule above: model legitimately lives in `settings.local.json` (a soft default — in-session `/model` still works); env vars do not
- **Multi-CLI prefix discipline** — env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*`) and the `ALLOWED_ENV_PREFIXES` allowlist in `schemas.ts` enforces this. Gemini additionally allowlists the **broad `GOOGLE_*`** namespace (intentional: Vertex AI auth needs `GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI`; it is the loosest allowlist entry, affecting only the user's own spawned CLI). When adding a setting, decide which CLI(s) it applies to and gate the env export accordingly. Never blanket-forward all prefixes. Resolver design pattern: `docs/opencode-integration.md`
- **Multi-CLI prefix discipline** — env-var prefix is CLI-specific (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*` vs `ANTIGRAVITY_*`) and the `ALLOWED_ENV_PREFIXES` allowlist in `schemas.ts` enforces this. Gemini additionally allowlists the **broad `GOOGLE_*`** namespace (intentional: Vertex AI auth needs `GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI`; it is the loosest allowlist entry, affecting only the user's own spawned CLI). When adding a setting, decide which CLI(s) it applies to and gate the env export accordingly. Never blanket-forward all prefixes. Resolver design pattern: `docs/opencode-integration.md`
- **Zod `.optional()` rejects `null`** — accepts `undefined` only. When the frontend builds a request body with `JSON.stringify`, an explicit `null` field is preserved on the wire and fails validation with `INVALID_INPUT`. Convert `null` → `undefined` before stringifying (e.g. `field: value ?? undefined`), or declare the schema `.nullish()`. This has caused real shipped bugs twice
- **`xterm-zerolag-input` is single-source** — the local-echo overlay source lives ONLY in `packages/xterm-zerolag-input/src/`, and is bundled into the **gitignored** `src/web/public/vendor/xterm-zerolag-input.js` (dev, by `scripts/postinstall.js`) and `dist/.../vendor/` (prod, by `scripts/build.mjs`). `app.js` only **consumes** it via `new LocalEchoOverlay(terminal)`; there is no inline copy. So: change the package source, then rerun the bundle step (`npm install` for dev, `npm run build` for prod). **Never hand-edit `app.js` for overlay behavior, and never commit the gitignored vendor bundle.** Always test on mobile after touching it. → [architecture-invariants#xterm-zerolag-input-is-single-source](docs/architecture-invariants.md#xterm-zerolag-input-is-single-source), `docs/local-echo-overlay-plan.md`
- **Default bind is loopback-only; non-loopback without a password starts but warns** — the server defaults to `--host 127.0.0.1`. Binding non-loopback (`--host`/`-H`/`CODEMAN_HOST`) without `CODEMAN_PASSWORD` starts anyway but prints a loud warning; `--allow-unauthenticated-network` / `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` acknowledges it. ⚠️ The production systemd unit passes no `--host`, so prod binds **localhost only**: reach it via `tailscale serve`/tunnel to `127.0.0.1`. A loopback bind is reachable through a same-host tunnel but NOT by a browser hitting the box's LAN IP. `install.sh` is separate and prompts for the binding (defaulting to LAN + a password), and preserves the existing binding on re-runs. → [architecture-invariants#default-bind-and-the-non-loopback-warning-path](docs/architecture-invariants.md#default-bind-and-the-non-loopback-warning-path), `docs/security-architecture.md`
- **Instance isolation / multi-instance attach danger** — the data dir (`~/.codeman`) and tmux socket (`tmux -L codeman`) are PROCESS-WIDE and shared by every Codeman on the machine, derived from `CODEMAN_INSTANCE` via `src/config/instance.ts`. ⚠️ A 2nd instance on the SAME socket **discovers and attaches PTYs to the first instance's live sessions**, resizing and mutating them. `$HOME` isolation is NOT enough because tmux is system-global. To run two instances, give each a distinct `CODEMAN_INSTANCE` (scopes dir + socket together), or set `CODEMAN_TMUX_SOCKET` + `CODEMAN_DATA_DIR` individually; `scripts/run-beta.sh` does this for a beta alongside prod. **Any new `~/.codeman/...` path MUST go through `dataPath()`**, never `join(homedir(), '.codeman', …)`. → [architecture-invariants#instance-isolation-and-the-multi-instance-attach-danger](docs/architecture-invariants.md#instance-isolation-and-the-multi-instance-attach-danger)
- **node-pty's macOS `spawn-helper` ships without `+x`** (issues #6, #204): `node-pty@1.1.0` publishes `prebuilds/darwin-<arch>/spawn-helper` as mode 0644, and macOS launches every PTY through it, so a stock macOS install fails every session start with `Error: posix_spawnp failed.` **Linux can never reproduce it**: `spawn-helper` is an `OS=="mac"` gyp target and node-pty ships no Linux prebuild, so node-gyp always emits an executable helper there. ⚠️ Look in **`prebuilds/<platform>-<arch>/`**, not just `build/Release/`, which does not exist on macOS. Repair is a chmod, never a mandatory rebuild (that would require Xcode CLI tools and deletes `prebuilds/` before compiling): `npm run fix:node-pty` chmods every helper then proves it by really opening a PTY. `spawnPtyWithHelperRepair()` (`utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts` and self-heals a broken install on the first failure. → [architecture-invariants#node-ptys-macos-spawn-helper-must-be-executable](docs/architecture-invariants.md#node-ptys-macos-spawn-helper-must-be-executable)
- **Headless screenshots: `deviceScaleFactor` MUST be 1, and write unique filenames** — under DSF=2 xterm's WebGL renderer draws glyphs at ~2× nominal size while still *reporting* nominal cell dims, so only the pixels reveal it and only the terminal font looks wrong. And overwriting a fixed output path leaves OS image viewers showing the old render, which reads as "the fix didn't work"; `scripts/capture-real-overview.mjs` mints a timestamped filename per run. Seed the per-device `localStorage` keys (`codeman:skin`, `codeman-font-size`, `codeman-app-settings`) so the capture matches a real device. → [architecture-invariants#headless-screenshot-capture](docs/architecture-invariants.md#headless-screenshot-capture)
**Import conventions**: Utils from `./utils`, types from `./types` (barrel), config from specific `./config/*` files.
@@ -157,7 +158,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| **Attachments** | `src/attachment-registry.ts`, `attachment-magic`, `generated-artifact-attachments`, `session-attachment-history`, `document-preview-cache`, `document-thumbnailer`, `document-conversion-limiter`, `config/attachment-guard` | See Key Patterns |
| **Plan** | `src/plan-orchestrator.ts`, `src/prompts/*.ts`, `src/templates/` (`claude-md.ts` + `case-template.md`) | `templates/` holds the CLAUDE.md scaffold generated into new cases |
| **Web** | `src/web/server.ts` ★, `sse-events.ts`, `routes/*.ts` (20 modules + barrel; `session-routes.ts` ★), `route-helpers.ts`, `ports/*.ts`, `middleware/auth.ts`, `schemas.ts`, `self-update.ts`, `plan-usage-latest.ts`, `ws-connection-registry.ts`, `heic-jpeg-converter.ts` + `heic-jpeg-worker.ts` | |
| **Frontend** | `src/web/public/app.js` (~5K lines, core) + 23 modules + `sw.js` | See Frontend section for the load order, which is authoritative |
| **Frontend** | `src/web/public/app.js` (~5K lines, core) + 25 modules + `sw.js` | See Frontend section for the load order, which is authoritative |
| **Types** | `src/types/index.ts` (barrel) → 20 domain files; also `src/types.ts` root re-export | See `@fileoverview` in index.ts |
★ = Large, central file (>50KB) — read its `@fileoverview` first. All files have `@fileoverview` JSDoc — read that before diving in. Discovery aid: `grep -l '@fileoverview' src/web/routes/*.ts` lists all route modules; same grep works for `src/types/`, `src/web/public/*.js`.
@@ -179,6 +180,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Input**: `session.writeViaMux()` for programmatic/curl input via tmux `send-keys -l` + `send-keys Enter`, single-line only. Interactive **browser** input goes through a durable **exactly-once** layer: a stable `clientId` + monotonic per-session `seq` persisted to localStorage until the server ACKs, so a dropped link cannot lose or double-deliver a prompt. `ws-connection-registry.ts` supersedes only same-TAB reconnects, so two tabs on one session coexist. → [architecture-invariants#input-delivery-and-ws-resilience](docs/architecture-invariants.md#input-delivery-and-ws-resilience)
**Agent wait primitives**: bounded long-polls so an agent driving Codeman from a shell can block instead of poll: `GET /api/sessions/:id/wait` (lifecycle signal), `GET /api/sessions/:id/wait-output` (literal substring, **never** regex) and `wait`/`waitTimeout` on `POST /api/sessions/:id/input`. Registry in `session-wait-registry.ts` (pure, no `Session` reference), bounds in `config/agent-wait.ts`. ⚠️ **A timeout is a 200** (`wait.timedOut`), never an error, so callers loop over short waits. ⚠️ `stop`/`blocked` come from Claude Code hooks and therefore fire for **`claude` mode ONLY** (`shell` installs none either); asking for one explicitly on another mode is a 400, the default set silently drops them. ⚠️ Send-and-wait registers the waiter BEFORE the write (a separate POST-then-wait races and reports the PREVIOUS turn), and both teardown paths must `notifySignal('exit')` BEFORE `cancelAll()`. ⚠️ Client-hangup abort listens on **`reply.raw`** guarded by `writableFinished`: on `req.raw`, `close` fires when the request BODY ends, which on a POST killed every send-and-wait instantly and no `app.inject()` test could see it. ⚠️ Worker liveness cannot come from `session.pid` — for a tmux session that is the local attach client, which outlives a worker dying inside its pane — so it is probed at the mux layer (`isPaneDead`, ~750 ms cache) on blocking waits only, never on the input hot path. ⚠️ Signals are edge-triggered with no history: one that fires with no waiter registered is unobservable afterwards, so gather fan-outs with send-and-wait or latched `wait-output` markers, never fire-and-forget-then-sequential-signal-waits. → [architecture-invariants#agent-wait-primitives](docs/architecture-invariants.md#agent-wait-primitives), `docs/api-reference.md`
**Idle detection**: Multi-layer (completion message → AI check → output silence → token stability). See `docs/respawn-state-machine.md`.
**Auto-resume on usage limit** (opt-in per session, top of the Respawn tab): when Claude halts on a subscription limit, `usage-limit-patterns.ts` (pure, unit-tested) parses the reset time and `SessionAutoOps` arms a timer for reset+2min, then sends Esc + `continue`. ⚠️ Respawn cycles are blocked while paused (`isLimitPaused` guard in `onIdleDetected`), which is what prevents `/clear` from wiping the paused conversation. Claude-mode only. → [architecture-invariants#auto-resume-on-usage-limit](docs/architecture-invariants.md#auto-resume-on-usage-limit)
@@ -193,7 +196,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Docker cases**: a case can point at a **container**, with any of the five CLI backends running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a sixth `SessionMode`**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design)
**External CLI modes (OpenCode, Codex, Gemini)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All three **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. → [architecture-invariants#external-cli-modes-opencode-codex-gemini](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini)
**External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token/CLI-info parsing, ❯-prompt readiness); these CLIs render their own TUIs, so readiness is output stabilization instead. All four **require tmux with no direct PTY fallback**, because secrets are injected via socket-scoped `tmux setenv` and never on the spawn command line. ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope; reading the raw shape silently breaks the run. ⚠️ **The local-echo overlay is DISABLED for codex sessions** (`_updateLocalEchoState` in terminal-ui.js, same branch as shell): codex's composer reacts per keystroke ("/" pops a live-filtering picker, arrows edit server-side state, the composer grows as it wraps), so buffer-until-Enter starved it into issues #218/#219/#220/#222. Codex also **drops keystrokes that share a PTY read with a bracketed paste**, so flushed text and the paste sequence must go out as separate delayed writes (mirroring the Enter branch's delayed `\r`). Tests: `test/local-echo-codex-gating.test.ts`. → [architecture-invariants#external-cli-modes-opencode-codex-gemini](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini)
**Run launch synchronization**: the Run entrypoint holds an in-flight lock and disables `#runBtn` for the whole launch (≥500ms), so a double click cannot create duplicate sessions with the same `w<n>-<case>` name. `_ensureCreatedSessionVisible()` runs before `selectSession()`, and `_onSessionCreated()` stays an idempotent upsert, so POST-first and SSE-first ordering both produce exactly one rendered tab. → [architecture-invariants#run-launch-synchronization](docs/architecture-invariants.md#run-launch-synchronization)
@@ -205,7 +208,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. Only the FIRST buffer load after a page load requests `full=1`; tab switches keep the cheap `?tail=` path. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). ⚠️ That re-pull must never DOWNGRADE the buffer: a repaint-mode CLI pane keeps no tmux history, so its capture is one frame and the reset+rewrite would delete history mid-scroll — `_replayWouldShrinkBuffer()` refuses it and slows that session's cooldown to 60s. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for codex/claude ≥ 2.1.187 at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). ⚠️ When that gate is FALSE on a claude session whose local buffer is hollow (`baseY === 0`), the gesture becomes coalesced PageUp/PageDown key sends (`_maybePageCliTranscript`) instead of a no-op; ⚠️ and `getClaudeCliVersion()` must never cache a FAILED probe (one timeout used to disable forwarding process-wide until restart). `_logScrollRouting()` prints the routing decision and its inputs once per session — read it before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Self-update** (App Settings → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. → [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
@@ -213,6 +218,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Filesystem path picker** (Link Existing "Browse" + the mobile keyboard's `📁 Path` key): lazy one-directory browsing via `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` for the tapped file. Inserts the path **without** Enter, so the prompt is never submitted; the sibling `⌫ All` key clears only the unsent prompt and must never send the agent's `/clear`. ⚠️ This is a **second file-serving surface and inherits neither the attachment confinement nor its ownership scoping** — it allowlists Home, `CASES_DIR`, `/mnt/d` and `CODEMAN_FILE_PICKER_ROOTS`, blocks sensitive trees, and rejects symlink escapes **after** `realpath`. ⚠️ The optional `sessionId` is an ownership boundary that must be `canAccessOwned`-checked by hand (it does not go through `findSessionOrFail`), and in multi-user mode a non-admin gets only their own `userSpacePath` as a root: per-user spaces live INSIDE `homedir()`, so a `Home` root exposes every other user's workspace. Previews go through the same global conversion limiter, and Markdown/TXT/JSON are served as inert `text/plain`. → [architecture-invariants#filesystem-path-picker](docs/architecture-invariants.md#filesystem-path-picker)
**File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md`
**Ultracode / workflow-run visualization** (opt-in, default OFF): the Workflow tool writes a completion artifact only at run *end*, so live in-flight runs exist solely as transcript dirs. `workflow-run-watcher.ts` therefore synthesizes ACTIVE runs from transcripts until the completion artifact appears and supersedes them. It is **STANDALONE** and deliberately never imports or touches `subagent-watcher.ts`, despite reading the same tree. Two independent toggles: `showUltracodeAgents` (docked panel) and `ultracodeFloatingWindows` (floating windows); the watcher starts if **either** is on. → [architecture-invariants#ultracode--workflow-run-visualization](docs/architecture-invariants.md#ultracode-and-workflow-run-visualization)
**Cross-session search**: `GET /api/search` federates an in-memory search over session metadata, run-summary events, and attachment-history entries. The pure core `searchSources()` does substring matching with hard per-type caps: **no regex (so no ReDoS) and no filesystem reads (so no traversal)**. The server-private `externalPath` is never read. → [architecture-invariants#cross-session-search](docs/architecture-invariants.md#cross-session-search)
@@ -229,9 +236,13 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
### Frontend
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `sanitize-html.js`(5.6) → `app.js`(6) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `ultracode-panel.js`(11.5) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `webview-tabs.js`(12.5) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData).
Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. Load order: `constants.js`(1) → `i18n.js`(1.5) → `mobile-handlers.js`(2) → `voice-input.js`(3) → `notification-manager.js`(4) → `keyboard-accessory.js`(5) → `input-cjk.js`(5.5) → `sanitize-html.js`(5.6) → `app.js`(6) → `terminal-ui.js`(7) → `respawn-ui.js`(8) → `ralph-panel.js`(9) → `orchestrator-panel.js`(9.5) → `cron-ui.js`(9.7) → `settings-ui.js`(10) → `panels-ui.js`(11) → `ultracode-panel.js`(11.5) → `admin-ui.js`(11.7) → `session-ui.js`(12) → `webview-tabs.js`(12.5) → `mobile-overview.js`(12.55) → `entrance-animations.js`(12.6) → `ralph-wizard.js`(13) → `api-client.js`(14) → `subagent-windows.js`(15) → `ultracode-windows.js`(15.5) → `image-input.js`(16). `i18n.js` translates static + newly inserted application DOM while skipping terminal/response/file/user-name surfaces; `input-cjk.js` handles CJK IME composition via an always-visible textarea below the terminal (`window.cjkActive` blocks xterm's onData).
**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry)
**Entrance animations** (`entrance-animations.js`, all OFF by default): opt-in animations for the four things that appear when work starts, chosen per surface via `data-tab-anim` / `data-term-anim` / `data-win-anim` / `data-line-anim` on `<html>`. Defaults are the `legacy` theme, so an untouched install behaves exactly as before and every hook short-circuits on its first line. ⚠️ Tabs and connection lines are **destroyed mid-animation** on every re-render (`_fullRenderSessionTabs()` replaces the strip's innerHTML; `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''`), so both are tracked by id and re-applied to the fresh element with a **negative `animation-delay`** to resume rather than restart. ⚠️ The terminal-pane styles may animate **transform / opacity / clip-path only**, xterm's FitAddon derives rows+cols from `getComputedStyle(parent).width/height`, so animating width/height/padding there would resize the PTY. ⚠️ Window styles other than `beam` transform the window, which moves the rect its connection line is aimed at; `beam` deliberately animates opacity/filter only so its line can draw toward a stable target. Persisted to its own `codeman:*Anim` localStorage keys (per-device, deliberately NOT in the `.strict()` `SettingsUpdateSchema`); picker in App Settings → Appearance, full per-surface lab at `?animlab=1`.
**Phone overview home screen** (`mobile-overview.js`, phones only, per-device `mobileOverviewEnabled`, default ON): under 430px the "C" logo shows a session overview (NEEDS YOU / CURRENT SESSIONS / PAST SESSIONS) instead of the welcome overlay; tablet and desktop are unchanged. The branch lives in `showWelcome()`/`hideWelcome()` (terminal-ui.js) behind `shouldUseMobileOverview()`, which is **width-driven** (`getDeviceType() === 'mobile'`) because this is a layout decision, unlike the settings namespace which stays handheld-based. ⚠️ The container ships with the `hidden` attribute and only this module removes it: never give `.mobile-overview` a bare `display` rule, since desktop does not load `mobile.css` (`media="(max-width: 1023px)"`) and would then render it unstyled. Live re-renders ride on the tail of `_renderSessionTabsImmediate()` (every state change it needs already funnels there); PAST rows come from one `_fetchUnifiedSessions(60)` per home-screen visit and resume through the shared `resumeHistorySession()`, so they behave exactly like the welcome screen's Resume list. ⚠️ Two things must stay in lockstep with surfaces outside this module, because divergence reads as a bug rather than a style: the split Run button carries the **toolbar's own classes** (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) so the per-backend gradient and the light-skin overrides apply unchanged (mobile.css must therefore set no `background`/`color` on it), and row status uses the **session-tab language** (green dot when fine, `pulse` while working, yellow blinking row when waiting for input, red blinking row when a question is pending, mirroring `tab-alert-idle`/`tab-alert-action`). The picker mirrors the toolbar run-mode menu (`setRunMode()` + `run()`, `openWebviewFromMenu()` for saved dashboards) and deliberately omits its Recent-Sessions block, since PAST SESSIONS is that. Status pills carry `data-i18n-skip` (generic words like "idle" collide with state strings elsewhere).
**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry)
**Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema.
@@ -255,7 +266,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active). Rebindable via the registry.
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+C (copy selection, else interrupt) / Ctrl+Shift+C (copy, never interrupts), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active). Rebindable via the registry.
### Security
@@ -272,7 +283,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
| **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR and hook-secret have separate buckets, so neither can lock out login |
| **Hook bypass** | `/api/hook-event` + `/api/status-telemetry` skip Basic auth (localhost-only, schema-validated), but when auth is active the loopback bypass requires `X-Codeman-Hook-Secret` **unconditionally** (Codeman cannot detect a user's own loopback reverse proxy) |
| **Tunnel** | Enabling a tunnel **refuses** without `CODEMAN_PASSWORD` unless exposure is acknowledged via `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` or the per-request `acknowledgeUnauthTunnel:true` action field (never persisted) |
| **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`) |
| **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`/`ANTIGRAVITY_*`) |
| **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS |
**Security-relevant env vars**: `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks), `CODEMAN_ALLOWED_HOSTS` (extra Host/Origin allowlist entries for reverse proxies; bare `.suffix` matches subdomains), `CODEMAN_DOCKER_BRIDGE_HOOKS=1` (opt-in hooks-only listener on the docker bridge gateway).
@@ -283,7 +294,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
### API Routes
~199 handlers across 21 route files in `src/web/routes/`: system (45), sessions (32), cases (27), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
~200 handlers across 21 route files in `src/web/routes/`: system (45), sessions (34), cases (27), files (16), orchestrator (10), ralph (9), cron (9), admin (8), plan (8), respawn (7), webviews (6 + the `/webview/:cap/*` proxy), mux (5), push (4), scheduled (4, legacy `ScheduledRun`), me (2), teams (2), search (1), hooks (1), clipboard (1), status-telemetry (1), ws (1 WebSocket). Each file has `@fileoverview` with endpoint details.
**HTTP contract** (stable since 0.9.x, see `docs/versioning-policy.md`; full envelope/status/error-code/SSE spec in `docs/api-reference.md`): responses use the `ApiResponse<T>` envelope — `{ success: true, data? }` or `{ success: false, error, errorCode }` (`src/types/api.ts`). `/api/v1/*` is a versioned alias of `/api/*` (URL rewrite in `server.ts`).
@@ -352,6 +363,6 @@ Two constraints worth knowing before you touch them: the env-derived PTY buffer
## Scripts & Tunnel
**`install.sh`** (repo root, 69KB) is the public entry point: `curl -fsSL <raw url> | bash` installs Node/tmux if missing, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service. It prompts for the network binding (LAN default + password prompt) and preserves the existing binding on re-runs via `read_existing_binding()`. `install.sh update` and `install.sh uninstall` also exist; `CODEMAN_NONINTERACTIVE=1` approves system changes for automation.
**`install.sh`** (repo root, 69KB) is the public entry point: `curl -fsSL <raw url> | bash` installs Node/tmux if missing, clones to `~/.codeman/app`, builds, and offers a systemd/launchd service. The network-access prompt is 3-way: **Tailscale** (loopback bind + guided `tailscale serve --bg <port>` HTTPS setup: install/login/operator/tailnet-HTTPS-toggle, then curl-verified end-to-end), **LAN** (0.0.0.0 + password prompt), or **local-only**; it preserves the existing binding on re-runs via `read_existing_binding()`. Tailscale state is detected dynamically from `tailscale serve status --json` (no marker files); the installer must NEVER `tailscale serve reset` or touch serve mappings other than 443→Codeman's port (users have unrelated serve config). `install.sh update`, `install.sh uninstall`, and `install.sh tailscale` (retrofit Tailscale access onto an existing install) also exist; `CODEMAN_NONINTERACTIVE=1` approves system changes for automation, `CODEMAN_TAILSCALE=1` presets the Tailscale choice (never installs Tailscale non-interactively).
Other key scripts: `scripts/tmux-manager.sh` (safe tmux mgmt), `scripts/tunnel.sh [quick|named] start|stop|status|url` (quick = random trycloudflare URL, default; `named setup|enable` = fixed-hostname tunnel via `scripts/codeman-tunnel-named.service`; bare `start|stop|url` still means quick), `scripts/run-beta.sh` (isolated beta instance), `scripts/build-agent-image.mjs` (docker base image), `scripts/self-update.sh` (detached updater). Production services: `scripts/codeman-web.service`, `scripts/codeman-tunnel.service`. **Always set `CODEMAN_PASSWORD`** before exposing via tunnel.
+81 -23
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Gemini &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -27,7 +27,7 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, four CLIs** - run [Claude Code, OpenCode, Codex, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, five CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -68,7 +68,7 @@ This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, a
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works). The installer detects whichever of the four is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the five is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -151,7 +151,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
@@ -197,7 +197,7 @@ codeman web --https
# Open on your phone: https://<your-ip>:3000
```
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended) — it provides a private network so you can access `http://<tailscale-ip>:3000` from your phone without TLS certificates.
> `localhost` works over plain HTTP. Use `--https` when accessing from another device, or use [Tailscale](https://tailscale.com/) (recommended): the installer can set it up for you (choose **Tailscale** at the network-access prompt, or run `bash ~/.codeman/app/install.sh tailscale` on an existing install). That gives you `https://<your-machine>.<tailnet>.ts.net` with a real certificate: private to your tailnet, no password required, and PWA install + push notifications work on your phone.
### Secure QR Code Authentication
@@ -231,7 +231,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -404,7 +404,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
@@ -426,7 +426,7 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
@@ -612,7 +612,7 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
@@ -651,6 +651,8 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Alt/Option+[` / `Alt/Option+]` | Previous / next session |
| `Alt/Option+1`-`Alt/Option+9` | Switch to tab N (physical keys, so macOS Option layouts work) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+C` | Copy selection, or interrupt when nothing is selected |
| `Ctrl+Shift+C` | Copy selection (never interrupts) |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
@@ -678,15 +680,21 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
### Rules of the road (read before you POST)
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
1. **Single-line input, ending in `\r`.** Programmatic input is sent as literal text, and Enter fires **only when the input contains a carriage return**: `{"input":"run tests\r"}`. Without the `\r` the text sits on the session's prompt unsubmitted (and a combined `wait` runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs the joined command `echo Aecho B`: send one line per call.
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard). ⚠️ A `401` replies with the bare string `Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error instead of showing the failure: check the status before parsing.
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
```bash
# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
# The fallback below fits a stock install; on a --https install set the https:// URL
# yourself and add -k to each curl (self-signed cert).
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
@@ -698,18 +706,63 @@ curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
# first-run trust dialog only as the fallback. (Probing trust first and sending
# a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
# so the probe matches stale text and the Enter lands in a ready composer.)
# Match single tokens: TUI text can reach the matcher without its spaces.
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # accept the first-run trust dialog
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. Read the terminal back
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
# writing, so it can't answer with the previous turn's idle state)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
# spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)
# 5. Stream live events (session output, agent activity, status)
# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. Or wait for a marker in the output (works for shell sessions too).
# ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
# typed line never contains it: your own keystrokes echo into the output
# stream, so an unsplit marker matches before the command has run. from=buffer
# catches a marker that printed before the wait landed.
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
# tail counts BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events" # Server-Sent Events
# 6. Schedule recurring work (cron-style job)
# 7. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
@@ -717,11 +770,11 @@ curl -s -X POST "$API/api/cron/jobs" \
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. Inspect background sub-agents and their transcripts
# 8. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
# 9. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq
```
@@ -747,7 +800,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
## API
REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
@@ -755,8 +808,11 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | List all |
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
| `GET` | `/api/sessions/:id/output` | Read terminal output |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`: `clientId`+`seq` = exactly-once; `wait` blocks until the turn ends) |
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
@@ -810,6 +866,8 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
---
## Architecture
@@ -842,7 +900,7 @@ flowchart TB
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
+12 -8
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Gemini &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -58,7 +58,7 @@ curl -fsSL https://getcodeman.com/install | bash
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可)。安装器会自动检测这四个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这五个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
@@ -141,7 +141,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
@@ -221,7 +221,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Gemini` 或 `Terminal`(普通 shell)。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
@@ -394,7 +394,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
@@ -416,7 +416,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
@@ -602,7 +602,7 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
@@ -641,6 +641,8 @@ sc -l # 列出会话
| `Alt/Option+[` / `Alt/Option+]` | 上一个 / 下一个会话 |
| `Alt/Option+1`–`Alt/Option+9` | 切换到第 N 个标签(按物理键位,macOS Option 布局也适用) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+C` | 复制选中内容;未选中时中断代理 |
| `Ctrl+Shift+C` | 复制选中内容(永不中断) |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
@@ -800,6 +802,8 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
---
## 架构
@@ -832,7 +836,7 @@ flowchart TB
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
+1
View File
@@ -24,6 +24,7 @@ export default defineConfig({
'test/inline-rename.test.ts', // browser (Playwright)
'test/opencode-resize.test.ts', // browser (Playwright)
'test/webgl-fallback.test.ts', // browser (Playwright)
'test/terminal-copy-shortcut.test.ts', // browser (Playwright)
],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
+13 -3
View File
@@ -26,8 +26,8 @@ RUN apt-get update \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The agent CLIs (all four backends Codeman supports). Pinning is left to the
# rebuild cadence (see docs/docker-cases-plan.md, user-decision 2).
# The npm-published agent CLIs. Pinning is left to the rebuild cadence (see
# docs/docker-cases-plan.md, user-decision 2).
RUN npm install -g \
@anthropic-ai/claude-code \
@openai/codex \
@@ -35,6 +35,15 @@ RUN npm install -g \
opencode-ai \
&& npm cache clean --force
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
# the installer's default target is `$HOME/.local/bin`, which at build time is
# root's home and would be unreachable by the `agent` user the container runs as.
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
&& chmod 755 /usr/local/bin/agy \
&& agy --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
@@ -50,7 +59,8 @@ ENV HOME=/home/agent
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir.)
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions \
+662
View File
@@ -0,0 +1,662 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 5 IMPLEMENTED and multi-round verified, uncommitted as of 2026-08-08.
Step 6 (CLI install command + per-case injection + `agentSkillEnabled`) is not built.
See [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and what is still open.
**Date**: 2026-08-08
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
---
## 0. Where this came from: what herdr does
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
multiplexer built around AI coding agents. Relevant findings from the research pass:
| Capability | How herdr does it |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
The commands the skill teaches the agent:
| Group | Commands |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| workspace | `workspace list`, `workspace create` |
| tab | `tab list --workspace <id>`, `tab create` |
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
### The honest comparison
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
What herdr genuinely does better is being **callable by the agent running inside it**. For
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
---
## 1. Gap analysis
| herdr capability | Codeman equivalent today | Gap |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
| `pane wait-output --match` | nothing | **missing** |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
the two real gaps.
---
## 2. Part 1: the Codeman agent skill
### 2.1 Goal
An agent running inside a Codeman session can discover and correctly drive Codeman without the
user pasting API docs into the prompt, and without inventing dangerous calls.
### 2.2 Layout and distribution
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
```
skills/
codeman/
SKILL.md <- single source of truth
reference/
endpoints.md <- full endpoint tables, loaded on demand
recipes.md <- worked multi-session orchestration examples
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
```
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
point of shipping a skill.
Install paths, in order of how a user gets it:
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
file. This is the path for users who installed via npm and never cloned the repo.
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
into context on every turn, so an always-on skill has a small permanent token cost, and we
should measure that we are buying something with it first.
### 2.3 SKILL.md content
Frontmatter, per the skills convention (`name` + `description` required):
```yaml
---
name: codeman
description: >-
Control Codeman, the session manager this agent is running inside: list sessions,
start worker sessions, send prompts, read terminal output, and wait for other agents
to finish. Only usable when CODEMAN_MUX=1.
---
```
Body sections, in order:
**1. Guard (first thing, non-negotiable).**
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
SELF="${CODEMAN_SESSION_ID:-}"
```
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
cannot identify is not one it should be driving.
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
- Single-line input only. Multi-line breaks the agent TUI (Ink).
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
- Prefer `/api/v1/*`, the stable alias.
**3. Safety rules (the section that does not exist anywhere today).**
- Never act on `$CODEMAN_SESSION_ID`. That is you.
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
**4. Recipes**, each one a single copy-pasteable curl:
| Task | Call |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| list sessions | `GET /api/v1/sessions` |
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
| read output | `GET /api/v1/sessions/:id/output` |
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
| watch sub-agents | `GET /api/v1/subagents` |
| schedule work | `POST /api/v1/cron/jobs` |
| clean up | `DELETE /api/v1/sessions/:id` |
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
part of the skill stays small.
### 2.4 An ergonomics guard worth adding server-side
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
equals the target id, with a clear error.
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
### 2.5 Verification
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
to "start a worker session that runs the test suite and tell me when it finishes".
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
lead agent waited rather than polling in a busy loop.
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
Never run this against `w1`/`w2`/`w3`.
### 2.6 Files touched
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
- `.claude/skills/codeman` symlink (new)
- `src/cli.ts` (new `skill install` subcommand)
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
from the raw body, per the partial-PUT invariant)
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
- `package.json` `files` array, so `skills/` ships to npm
- README pointer, `docs/extending-codeman.md` seam 3 pointer
---
## 3. Part 2: wait primitives
### 3.1 Goal
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
which a curl-driven agent cannot practically consume: it would have to hold a streaming
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
solve it with bounded long-poll endpoints.
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
new optional fields are non-breaking).
### 3.2 The signal model
A waiter resolves on the first of a set of signals. Sources that already exist:
| Signal | Source today |
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
waits forever on `stop` in a codex session.
### 3.3 Endpoint specs
#### A. `GET /api/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
Response (always 200 unless the session is missing or a cap is hit):
```json
{
"success": true,
"data": {
"signal": "stop",
"timedOut": false,
"immediate": false,
"ended": false,
"waitedMs": 8421,
"status": "idle",
"sessionId": "...",
"until": ["stop", "idle", "exit"],
"limitPaused": false
}
}
```
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
timeout was expected rather than a stall worth retrying hard.
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
can loop without treating every poll boundary as a failure. Errors are reserved for
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
`immediate: true` means the session was already in the requested state and `fresh` was not set.
#### B. `GET /api/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
generic one like `BUILD OK`. The skill's recipes must show that.
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
readability only; matching runs on the raw stripped text. Without it, a real pane's
`\r\n` padding between the prompt and the match fills the whole context window with
nothing, which was the first thing the live test showed.
#### C. `wait` on the existing input endpoint
`POST /api/sessions/:id/input` gains two optional fields:
```json
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
```
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
when the input contains a carriage return.)
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
"input delivered" and "session flips to working" there is a window where a naive
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
ships `agent prompt --wait` as its own thing.
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
Two behaviors to preserve carefully:
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
purpose (a tmux child process must not block the HTTP response). With `wait` present the
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
observable for the first time. The non-wait path must keep its current fire-and-forget shape
byte for byte.
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
original turn may be long over, and requiring a new transition would block a redelivery until
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
current state.
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
the waiter is registered. If registration then fails on a full pool, the handler must call
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
duplicate and the input is lost by the very mechanism reliable delivery exists for.
### 3.4 Module design
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
(same split as `self-update.ts`):
```ts
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
notifySignal(sessionId, signal: WaitSignal): void
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
notifyOutput(sessionId, chunk: string): void
cancelAll(sessionId, reason): void
```
Wiring points, all existing:
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
callbacks so a listener could be added lazily; that was deleted once it was clear no
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
the no-waiter check comes before the ANSI strip.
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
live-testing the delete path, not by the unit tests.
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
| Constant | Default | Why |
| ------------------------- | ------- | --------------------------------------- |
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
| `MAX_WAITERS_PER_SESSION` | 16 | |
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
### 3.5 Transport concerns
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
The skill's recipes must show the loop.
### 3.6 Edge cases to get right
| Case | Behavior |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
### 3.7 Tests
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
chunk-straddling output match, case-insensitive match.
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
path is unchanged (still returns before `writeViaMux` settles).
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
### 3.8 Files touched
- `src/config/agent-wait.ts` (new)
- `src/web/session-wait-registry.ts` (new)
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
generated client must send `undefined`, never `null`)
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
---
## 4. Deferred: parts 3 to 5
Not in scope now, kept here so they are not lost.
### Part 3: promote `blocked` to a first-class state
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
re-deriving it. herdr makes `blocked` a real state that rolls up.
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
mobile overview, the wait endpoints, and any external agent read one field.
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
the HTTP contract.
### Part 4: `GET /api/schema`
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
### Part 5: detection manifests instead of hardcoded patterns
CLI-specific readiness, blocked and usage-limit patterns live in code across
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
Bundled manifests plus local override only, no network.
### Explicit non-goals
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
runtime means third-party code inside a process that spawns agents with your credentials, on a
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
delegates to tmux, so PTYs already survive a self-update restart.
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
would double the surface for no capability gain.
---
## 5. Sequencing
| Step | Work | Gate |
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | settings partial-PUT test, case-creation test |
| 7 | Docs: api-reference, extending-codeman, README | |
| 8 | COM (minor bump: new endpoints, new setting, new optional fields) | both CI and Release workflows green |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. `skills/` at the repo root, accepted despite the short-root rule? (Recommended yes, the
install one-liner depends on it.)
2. `agentSkillEnabled` default: OFF for the first release then flip, or ON immediately?
3. Auto-inject the skill into every case's `.claude/skills/`, or global install only?
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary?
5. Regex support in `wait-output`: confirm literal-only for v1.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
### What shipped
| Piece | Files |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
### Bugs found in ADJACENT code, not in the new feature
These are the highest-value output of the exercise and none were on the plan:
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
without the flag, success with it, and the failure swallowed by the hook's own
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
have left every one of them broken).
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
only if the payload has a carriage return; without it the text sits in the composer
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
docs' own examples.
### Design decisions worth not re-litigating
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
waits because tunnels cut idle connections, and every poll boundary would otherwise be
indistinguishable from failure.
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
turn as this one. The waiter is registered before the write.
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
`--regex` because Rust's regex crate is linear-time.
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
`close`).
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
### Verification rounds
Six agents across three rounds, each verifying the previous round's work rather than its
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
success without running its task. Two traps recurred often enough to name:
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
in `afterEach` silently killed the registry for every later test in a file; three test
files sharing one session id against the process-wide registry let one file's leftover
waiter fail another's assertion. Any new wait test needs care on all three.
- **HTTP-only test instances.** Every isolated instance used during the build was plain
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
user actually runs.
### Resolved at wrap-up (2026-08-08, conclusion pass)
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
the skill** rather than patched. Signals are edge-triggered with no history, so a
`stop` that fires before its waiter registers is unobservable afterwards; a
`fresh=0` gather was rejected because the only `until` set that current state can
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
shell flows reliable; the limitation is recorded in
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
dialog emits no further `idle`; the false success is the startup transition).
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
documented; every send-and-wait retry loop now treats `duplicate:true` +
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
floor named; the auth fallback now also reads the supervisor definition
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
lines; `pid != null` is documented as startup-only, never liveness.
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
### Still open
- **Release checklist**: `package.json` `files` includes `skills`, which is still
untracked. `git add skills/` must be part of the release commit, or npm publishes
a tarball without the skill (a `files` entry that does not exist is silently
ignored, so nothing fails).
- The 1.13.0 changeset is written under `.changeset/`; consuming it (COM flow),
the release commit, and the deploy remain.
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`) and converting
timeout-shaped test detections into fast assertions.
+341
View File
@@ -46,6 +46,20 @@ payload return `{ "success": true, "data": {} }`.
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
@@ -66,6 +80,333 @@ the HTTP status.
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
```
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
#### What the matcher actually sees
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
picture; the matcher sees the stream that painted it. For linear output the two
agree once escapes are stripped, but a full-screen TUI composes its picture with
cursor positioning, so what the pane shows and what the stream carries can differ.
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
with cursor moves rather than printing spaces, so screen text can reach the
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
this folder` matched, `Quick safety check` did not), so a multi-word `match`
against a TUI pane is unreliable rather than impossible. Match a **single
space-free token**, ideally one you printed yourself. Plain command output (a
shell worker, an `echo`) keeps its spaces.
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
it.** It is cut from the same normalized stream the match ran against, then
cleaned for display: remaining raw control bytes are removed (an agent pipes the
snippet into its own terminal, so a worker's bytes must not be able to reset that
display) and blank runs are collapsed. A printable needle that matched will appear
in it; a needle containing control bytes or a blank run may not survive verbatim.
```bash
MARK="DONE_$RANDOM"
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
```
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
any of them:
```json
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
2. `wait.timedOut`: a poll boundary. Loop again.
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
File diff suppressed because one or more lines are too long
+2 -2
View File
@@ -44,9 +44,9 @@ records), kept distinct from the existing `ScheduledRun`.
## 2. Where agent/session types are defined
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity'`
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode}-cli-resolver.ts`.
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
## 3. Where input is sent into a session
+2 -2
View File
@@ -1,7 +1,7 @@
# Cron Jobs — User & Operator Guide
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini) session on a schedule and
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
+16 -1
View File
@@ -2,7 +2,7 @@
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` all work inside the container.
## One-time setup: build the base image
@@ -15,6 +15,21 @@ node scripts/build-agent-image.mjs # builds codeman/agent:base
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
```bash
node scripts/build-agent-image.mjs --no-cache
```
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
+410
View File
@@ -0,0 +1,410 @@
# Extending Codeman
Codeman has no plugin runtime, and that is a deliberate choice rather than a
missing feature. A plugin runtime means running third-party code inside a process
that spawns agents with your credentials, on a server people routinely expose
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
exist, so it does not hand that away for an extension mechanism.
Instead there are four seams that already work, from any language, with nothing
installed:
| You want to | Use | Runs where |
| --- | --- | --- |
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
Everything below is covered by the stability promise in
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
envelope, `errorCode` values, and SSE event names are stable. Additive changes
(new endpoints, new optional fields, new events) are non-breaking. Breaking
changes ship under a new prefix (`/api/v2`).
## Before you start
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
authenticate once and keep the `codeman_session` cookie. With no password set,
Codeman is loopback-only and unauthenticated.
```bash
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
```
**Envelope.** Every response is `{"success": true, "data": ...}` or
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
is in [`api-reference.md`](api-reference.md).
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
the payload at the top level rather than under `data`. Read defensively with
`body.data ?? body`.
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
the status code before you parse, or a missing password looks like a broken endpoint.
**Already driving Codeman from an agent?** The README's
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
running inside Codeman find the API and avoid acting on itself. This page is for
code running *outside* a session.
## Seam 1: Web tabs
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
your agent sessions. You write a normal web page; Codeman handles embedding it.
```bash
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
-H 'Content-Type: application/json' \
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
```
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
framing check), `POST /api/v1/webviews/:id/open`.
### Why it is proxied
By default your page is served through Codeman's own origin at `/webview/:cap/*`
rather than framed directly. A direct iframe fails three ways at once: production
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
cross-origin frames. Proxying solves all three without weakening the CSP.
### The two things that will confuse you
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
`trusted: true`. Two consequences look like bugs in your own app:
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
patches the common DOM sinks, but if you construct URLs in an unusual way,
prefer relative paths.
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
itself renders fine, this is the area to look at.
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
the API that spawns agents. Only mark your own trusted code.
## Seam 2: SSE events
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
`event: <name>` plus `data: <json>`. There are 149 event names following a
`domain:action` convention, registered in `src/web/sse-events.ts`.
The ones most integrations want:
| Event | Meaning |
| --- | --- |
| `session:created`, `session:deleted` | A session appeared or went away |
| `session:idle` | The agent stopped working |
| `session:completion` | A completion message was detected |
| `session:exit`, `session:error` | The session ended or failed |
| `hook:permission_prompt` | The agent is asking for permission |
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
| `hook:task_completed`, `task:completed` | Work finished |
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
| `mux:died` | A multiplexer session died unexpectedly |
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
### Filtering
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
sessions you did not list. Lifecycle and metadata events are always delivered, so
you cannot accidentally filter away the thing you are listening for.
Pass `?clientId=<uuid>` to enable live filter updates through
`POST /api/v1/events/subscribe` without reconnecting the stream.
### Example: notify when any agent needs you
```js
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buf = '';
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buf += decoder.decode(value, { stream: true });
const frames = buf.split('\n\n');
buf = frames.pop() ?? '';
for (const frame of frames) {
const name = frame.match(/^event: (.+)$/m)?.[1];
const data = frame.match(/^data: (.+)$/m)?.[1];
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
}
}
```
## Seam 3: HTTP API and CLI
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
The common ones:
```bash
# List sessions (live + persisted + transcript history, deduped)
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
# Create a session
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
-H 'Content-Type: application/json' \
-d '{"workingDir":"/home/me/project","mode":"claude"}'
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
# when the input contains a carriage return; without it the text sits on the
# session's prompt unsubmitted)
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true}'
```
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
`seq` (monotonic per session). Send both and the server applies each pair
at-most-once, so retrying after a dropped connection cannot type the prompt
twice. Omit them entirely rather than sending `null`.
It also accepts `wait` and `waitTimeout`, which hold the response open until the
session finishes the turn you just started. `wait` is `true` (the default signal
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
under `data.wait`. Sending them changes nothing for callers that do not: without
`wait` the response is still `{"success": true, "data": {}}` and the write is still
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
a **tagged duplicate** (a pair the server already applied) skips the write but still
waits, answering from the session's current state rather than blocking for a
transition that already happened. It reports `"delivered": false, "duplicate": true`.
### Waiting instead of polling
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
output), and the `wait` field above. Full parameter and response tables are in
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
whether your integration works, and the last one is what actually bites:
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
waits rather than issuing one long one, because `tailscale serve` and cloudflared
both cut idle connections and a single 10-minute call is the pattern most likely
to die in the field.
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
default). Read it rather than assuming you got what you asked for.
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
even `idle` fires only once at startup, so send-and-wait there can only time out.
See the Gotchas below.
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
`\r` does not get past it, and the session's startup `idle` lands inside the wait
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
# 1. Start a worker session (creates the case if it does not exist yet).
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
# later step "succeeds" against nothing.
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
# the previous turn's idle state. Single line only, ending in \r (otherwise
# Enter is never sent and this wait times out on a turn that never started).
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
| jq -c '.data.wait')
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
for _ in $(seq 1 30); do
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
W=$("${CURL[@]}" \
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
done
jq -r 'if .ended or .aborted then "worker is not running"
elif .timedOut then "still working after 30 waits"
else "signal: \(.signal)" end' <<<"$W"
# 5. Read what it produced, then delete the session YOU created, by exact id.
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Waiting on a marker instead of a signal is the form that works in **every** mode,
and the only one that works on a `shell` session:
```bash
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
# into the output stream, so an unsplit marker matches before the command has run.
# `from=buffer` also catches a marker that printed before the wait registered.
N=$RANDOM
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
```
For shell scripting, the `codeman` CLI is the same surface without the HTTP
plumbing:
```
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
codeman ralph start|stop|status|reset codeman users add|passwd|list
codeman status | list | attach <path> codeman doctor
```
## Seam 4: Hooks
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
Codeman installs its own hooks automatically, but the endpoint is open to yours.
```json
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
```
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
event.
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
the loopback bypass requires the `X-Codeman-Hook-Secret` header
(`~/.codeman/hook-secret`) unconditionally.
## Gotchas
Every one of these has cost somebody real time.
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
call the API. Integrate server-side.
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
work. Cross-site origins are blocked by the CSRF guard.
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
shipped bugs more than once.
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
simple-request CSRF, so it is deliberate. Send `application/json`.
- **Prompts are single-line and must end with `\r`.** The server splits your text
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
Enter **only when the input contains a carriage return**. Without it your text
sits on the prompt unsubmitted, which is the single most common "the wait
endpoints don't work" report: the wait runs its full timeout on a turn that never
started. Newlines inside the string are stripped rather than rejected, so
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
per call.
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
unique to each call, and build it so the typed line never contains it (your own
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
rejected with a `400` rather than ignored.
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
line included), a partial escape at a chunk boundary is held back until its tail
arrives, and a match may straddle PTY chunks, so text you printed yourself
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
words with cursor moves, so its text can reach the matcher **without spaces** and
a multi-word match is unreliable there. Match one short space-free token, ideally
one you printed yourself, and keep it out of the typed line (your own keystrokes
echo into the stream).
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini` or
`antigravity` sessions.** They come from Claude Code hooks, which no other mode
installs, so only `idle`, `working` and `exit` exist there. Asking for them
explicitly is a `400`; omitting `until` is safe, since the server drops them from
the default set and echoes what it actually waited on as `wait.until`. Even in
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
written by Codeman < 1.13.0 against an `--https` install carries hook curls
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
time a session starts in that case.
- **Unwrap the envelope** before reading fields. `data` is not the response body.
## Publishing your integration
There is no registry and no review queue. Add the GitHub topic
**`codeman-integration`** to your public repository so others can find it, and
link back to Codeman in your README.
If a real ecosystem of these appears, a manifest format and an install command
become worth building. Until then, these four seams are the contract, and they
require nothing of you but HTTP.
## What Codeman deliberately does not have
- **No in-process plugin runtime.** See the reasoning at the top of this page.
- **No build or startup hooks** for third-party code. Run your own process.
- **No per-plugin config or state directories.** Manage your own files.
- **No sandbox for integration code**, because Codeman never launches it. Your
integration is your own process, started by you, with your permissions,
talking HTTP.
That last point is about integration code specifically, not about Codeman.
Sandboxing lives on a different axis here: the thing worth isolating is the
**agent**, and you isolate it per case with
[Docker cases](docker-cases.md), which run the agent in a hardened container with
a bind-mounted workspace and seeded (not shared) credentials. An integration that
creates or drives a Docker-backed session inherits that isolation for free, since
it is a property of the session rather than of the caller.
+432
View File
@@ -0,0 +1,432 @@
# File Viewer edit mode (issue #212)
Plan only. No implementation yet.
Goal: close the loop "agent writes a file, you review it in the viewer, tweak two lines, save, tell the
agent to continue" without hopping into the terminal, with the phone as the primary target.
Scope from the issue: an Edit toggle on text previews, a write endpoint that inherits the read path's
confinement, text-only, edit-in-place (no create, no delete, no rename), no editing through the
Docker/remote overlays.
---
## 1. What exists today
**Read path (backend), all in `src/web/routes/file-routes.ts`:**
| Route | Line | Notes |
| ------------------------------------ | ------ | ------------------------------------------------------------------ |
| `GET /api/sessions/:id/files` | `741` | Tree scan of `session.workingDir`, hidden files off by default |
| `GET /api/sessions/:id/file-content` | `865` | The text/preview classifier. `findSessionOrFail` + `validateSessionFilePath` |
| `GET /api/sessions/:id/file-raw` | `1018` | Bytes, 50MB cap |
| `GET /api/sessions/:id/file-preview` | `1254` | DOCX/PPTX to PDF, everything else redirects to `file-raw` |
| `GET /api/download` | `1384` | The only read route that also runs `isSensitivePath()` |
`file-content` classification order (`file-routes.ts:881-1011`): extension buckets (image / video / audio /
known-binary) return metadata only; otherwise the bytes are read, sniffed for a NUL in the first 8KB, and
either reported as `type:'binary'` or decoded as UTF-8 and **truncated to `lines` (default 500, hard cap
10000)**. Caps: `MAX_TEXT_FILE_SIZE` 10MB.
Confinement is `validateSessionFilePath()` (`src/web/route-helpers.ts:67`): `resolve()` then `realpathSync()`
then reject if the result is not under `workingDir`. Because it realpaths the *full* path, a symlink whose
target escapes the workspace is already rejected. Ownership is `findSessionOrFail()` which runs
`canAccessOwned()` (`route-helpers.ts:102`), a no-op outside multi-user mode.
**Read path (frontend), `src/web/public/panels-ui.js`:**
- `loadFileBrowser()` `2947`, `renderFileBrowserTree()` `2978`, click to `openFilePreview()` `3056`.
- `openFilePreview(filePath, sessionId, attachmentId)` `3193`: attachment-id branch, then docx/pptx, pdf,
svg branches, then the generic `file-content` fetch at `3274` with **`&lines=500` hardcoded**, rendering
text as `<pre><code>${escapeHtml(...)}</code></pre>` at `3298` and stashing `this.filePreviewContent`.
- `closeFilePreview()` `3308`, `copyFilePreviewContent()` `3751`.
- Markup: `src/web/public/index.html:420-432` (`filePreviewOverlay` / `-Title` / `-Body` / `-Footer`, two
header buttons: copy and close).
- CSS: `src/web/public/styles.css:9320-9430`. Overlay `z-index: 2000`, window `80vw/80vh`, capped
`900x700`. There are **no `.file-preview-*` rules in `mobile.css` at all**.
**Reachability on phones.** The header File Viewer button is hidden below 430px
(`mobile.css:482`, locked by `KNOWN_PHONE_HIDDEN` in `test/mobile-header-buttons-policy.test.ts`), so on a
phone the preview overlay is reached through:
1. an attachment card's **Preview** button (`panels-ui.js:3451`), which is exactly the "agent just wrote a
file" path the issue describes,
2. the attachment-history drawer (`panels-ui.js:3709`),
3. App Settings to Panels to **File Browser** (`showFileBrowser`, applied in `settings-ui.js:2202`; the
panel is mobile-styled at `mobile.css:1868`).
So edit mode is reachable on a phone today via (1) and (2) without touching the header policy. Improving
the entry point is listed as an open decision in section 10, not assumed.
---
## 2. Threat model, stated honestly
Anyone who can call this API can already reach `POST /api/sessions/:id/input` and type an arbitrary prompt
into an agent running with `--dangerously-skip-permissions`. A workspace-confined write endpoint therefore
does not create a new privilege tier for an authenticated caller.
What it *would* create if built carelessly is a **new host-write primitive reachable by path**, so the
things this plan actually defends against are:
1. **Path traversal / symlink escape** writing outside the workspace.
2. **TOCTOU**: a path component that becomes a symlink between validation and write.
3. **Cross-user writes** in multi-user mode (`canAccessOwned`).
4. **Silent data loss**, which is the highest-probability real-world failure here and gets its own section.
CSRF is already covered: `registerHostGuard()` (`src/web/middleware/auth.ts:555-578`) rejects any
non-safe-method request whose `Origin` is cross-site. The webview-capability exemption at that gate is
fenced to `GET`/`HEAD` for the Referer form (`auth.ts:161`) and to `/webview/:cap/*` paths for the path
form, so a proxied dashboard cannot reach a new `PUT /api/...`. Using `PUT` + `application/json` also
forces a preflight for any cross-origin attempt.
---
## 3. Backend design
### 3.1 New policy module: `src/config/file-editing.ts`
Pure, unit-testable, no IO (config lives in `src/config/`, no barrel, import the file directly).
```ts
export const MAX_EDITABLE_BYTES = 512 * 1024; // content cap, both directions
export const EDITABLE_EXTENSIONS: ReadonlySet<string>; // ts,tsx,js,jsx,mjs,cjs,json,jsonc,md,mdx,txt,
// css,scss,less,html,htm,xml,svg?,yml,yaml,toml,
// ini,cfg,conf,env?,sh,bash,zsh,fish,py,rb,go,rs,
// java,kt,swift,c,h,cpp,hpp,cs,php,sql,graphql,
// proto,lua,pl,r,jl,tf,gradle,csv,tsv,log,diff,patch
export const EDITABLE_BASENAMES: ReadonlySet<string>; // Dockerfile, Makefile, LICENSE, .gitignore,
// .prettierignore, .editorconfig, .nvmrc, ...
export function isEditableFileName(fileName: string): boolean;
export function isDeniedEditRelativePath(rel: string): boolean; // `.git/` subtree
export function detectEol(text: string): 'lf' | 'crlf';
export function applyEol(text: string, eol: 'lf' | 'crlf'): string;
```
Decisions baked in:
- **Allowlist, not blocklist**, per the issue and per the existing attachment-guard precedent.
- `svg` and `env` are deliberately marked with `?` above: `svg` is served as an untrusted octet-stream on
the read side (`file-routes.ts:118`) so allowing an edit is defensible, but I recommend **excluding
both** in v1. `.env` files are matched by `isSensitivePath()` anyway and would be rejected downstream;
excluding them at the allowlist keeps a single obvious refusal.
- `isDeniedEditRelativePath` blocks the `.git/` subtree: `.git/hooks/*` is code execution and a corrupt
index is unrecoverable-looking to a user who only wanted to fix a typo. Other dotfiles stay allowed but
are not reachable from the tree UI anyway (`showHidden=false`).
### 3.2 Read-for-edit: extend the existing GET
`GET /api/sessions/:id/file-content?path=<rel>&edit=1`
When `edit=1`:
- skip line truncation entirely (a truncated buffer must never become an edit buffer, see section 4.1),
- enforce `MAX_EDITABLE_BYTES` instead of `MAX_TEXT_FILE_SIZE` and answer 413 over it (as a structured
throw with `statusCode: 413`, the `throwFilesystemPickerError` pattern, since the central errorCode-to-
status map has no 413 entry; see the error-mechanics note in 3.3),
- run the editability gate (`isEditableFileName`, `isDeniedEditRelativePath`, `isSensitivePath`,
`isBlockedAttachmentPath`) and the content gate (NUL sniff plus UTF-8 round-trip, see 4.3),
- return `{ content, size, mtimeMs, totalLines, truncated: false, extension, editable: true, hash, eol }`.
`hash` is `sha256` hex of the exact on-disk bytes.
Non-`edit` responses gain **only** `editable: boolean` (additive, no shape change for existing consumers),
which is all the UI needs to decide whether to show the Edit button. No `hash` on plain reads: the Edit
action re-fetches with `edit=1` anyway (section 4.1), which is where the hash comes from, and hashing every
casual 10MB preview would be pure waste.
### 3.3 Write: `PUT /api/sessions/:id/file-content`
Body (new `FileWriteSchema` in `src/web/schemas.ts`, Zod v4):
```ts
{ path: string, content: string, baseHash: string, eol?: 'lf'|'crlf', force?: boolean }
```
Registered with an explicit route option `{ bodyLimit: 4 * 1024 * 1024 }`. **Fastify's default `bodyLimit`
is 1MB and this repo configures none**, and JSON escaping expands content: 2x for a file full of quotes or
backslashes, up to 6x for control characters (each serialized as a `\uXXXX` escape), so 512KB of content
can legitimately exceed 1MB on the wire; blowing the limit produces a raw `FST_ERR_CTP_BODY_TOO_LARGE`, not an `ApiResponse` envelope. Two
related sizing notes: `z.string().max()` counts **UTF-16 code units, not bytes**, so the schema's `.max()`
is only a coarse pre-filter and the real cap is an explicit `Buffer.byteLength(content, 'utf8')` check in
the handler (step 7a below); and 4MB comfortably bounds the worst-case expansion of a 512KB file without
inviting multi-MB bodies elsewhere.
**Error mechanics** (matters for both prod behavior and testability): a handler that *returns* a
`{success:false, errorCode}` envelope gets its HTTP status assigned centrally by the preSerialization hook
in `server.ts` (`httpStatusForErrorCode()`, `src/types/api.ts`), but the route-test harness
(`test/routes/_route-test-utils.ts`) installs only `installRouteErrorHandler`, **not** that hook, so
returned envelopes surface as HTTP 200 in tests. The PUT handler should therefore use the same
structured-**throw** pattern as the filesystem picker (`throwFilesystemPickerError`, `file-routes.ts:411`):
thrown `{statusCode, body}` errors are rendered identically in prod and in the harness, and they allow the
one status the code map cannot express (413). The error envelope itself is strictly
`{success:false, error, errorCode}`, **it has no data arm**, so no error response may carry extra payload.
Handler order (each step is a test case):
1. `findSessionOrFail(ctx, id, req)` (live sessions only, matching the read route, and it carries the
multi-user ownership check).
2. `parseBody(FileWriteSchema, req.body)`, then `Buffer.byteLength(content, 'utf8') <= MAX_EDITABLE_BYTES`
or 413 (the schema `.max()` alone cannot enforce a byte cap, see the sizing note above).
3. `validateSessionFilePath(session.workingDir, path)` or 404 (do not distinguish "outside workspace" from
"missing", matching the read route).
4. `isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)` or 403.
5. `isDeniedEditRelativePath(relativePath)` or 403.
6. `isEditableFileName(basename(resolvedPath))` or 400.
7. `stat`: must be `isFile()`, size within `MAX_EDITABLE_BYTES`, else 400/413. **No `O_CREAT` anywhere in
this handler**, which is what enforces edit-in-place.
8. Read current bytes, compute `hash`, run the NUL sniff and the UTF-8 round-trip check, else 400.
9. `hash !== baseHash && !force` gives **409 CONFLICT** (`ApiErrorCode.CONFLICT`, plain envelope; the error
arm carries no data, see the error-mechanics note). The client's conflict dialog gets fresh state by
re-fetching `edit=1`, which it needs for its Reload action anyway.
10. Build the output buffer: `applyEol(content, eol ?? detected-from-original)`; re-check
`Buffer.byteLength` against the cap.
11. Write atomically in the resolved parent directory:
`fs.open(<dir>/.<name>.codeman-tmp-<rand>, 'wx', stat.mode & 0o777)`, then `fchmod(stat.mode & 0o777)`
(open's mode argument is masked by the process umask, so the chmod is what actually preserves an
unusual mode), write, `fsync`, close, `fs.rename(tmp, resolvedPath)`, unlink the temp on any failure.
12. Re-stat, return `{ success: true, data: { path, size, mtimeMs, hash, totalLines } }`.
Why `O_EXCL` temp plus rename rather than truncate-in-place:
- `wx` cannot follow a pre-existing symlink, which closes the TOCTOU window from step 3 to step 11 without
needing `O_NOFOLLOW` gymnastics.
- `rename()` does not follow a symlink in the final component, so even if `resolvedPath` were swapped for a
symlink after validation, the symlink itself is replaced and the swap target is untouched.
- A crash mid-write leaves the original intact.
Caveat to document in the code comment: rename replaces the inode, so hardlinks to the file keep the old
content. That is the same trade-off vim makes by default and is preferable to a truncate window here.
No SSE event in v1. Nothing else in the app needs to know: `image-watcher.ts` only reacts to
`.png/.jpg/.jpeg/.gif/.webp/.bmp/.svg/.pdf/.docx/.pptx` adds (`image-watcher.ts:23-25`), none of which are
editable text, and the temp filename does not match either.
---
## 4. The five traps
These are the parts that turn a "small write endpoint" into a bug report.
### 4.1 Truncation (the data-loss trap)
The frontend fetches `&lines=500` (`panels-ui.js:3274`). Saving that buffer back would **delete every line
past 500**. Worse, the content hash of the full file would still match, so an optimistic-concurrency check
cannot catch it.
Mitigations, all three:
- The Edit affordance is only offered when the loaded payload came from `edit=1` (which never truncates).
Tapping Edit on an already-rendered preview **re-fetches** with `edit=1` before swapping in the editor.
- The read-for-edit path 413s above `MAX_EDITABLE_BYTES` rather than truncating, so "too big to edit here"
is an explicit refusal with a message, never a silent partial buffer.
- A test asserts `edit=1` never returns `truncated: true`.
### 4.2 Line endings
A `<textarea>`'s `.value` normalizes to LF. Saving a CRLF file naively rewrites every line, producing a
whole-file diff for a two-line change. So: the read returns the detected `eol`, the client echoes it back
unchanged, and the server re-applies it. Mixed-EOL files use the dominant style, which is lossy for the
minority lines; call that out in the response and accept it in v1.
### 4.3 Encoding
`buf.toString('utf-8')` on a latin-1 or otherwise non-UTF-8 file yields U+FFFD replacement characters, and
writing that back **corrupts the file**. The check is a round-trip:
`Buffer.from(decoded, 'utf8').equals(buf)`. If it fails, `editable: false` and the write is refused. This
also catches binary content that the NUL sniff misses. A UTF-8 BOM survives because it round-trips as a
leading U+FEFF; do not strip it.
### 4.4 Concurrency with the agent
The whole use case is editing a file the agent just wrote and may write again. `baseHash` plus 409 is the
guard. Do not use mtime alone: agents rewrite files within a single filesystem timestamp tick, and an
identical rewrite should not be reported as a conflict.
### 4.5 Symlinks and TOCTOU
Covered by `validateSessionFilePath` (escape) plus `wx` temp and `rename` (post-validation swap). One
intentional allowance: a symlink whose target is *inside* the workspace is edited through to its target,
because `validateSessionFilePath` returns the realpath. That matches what a user tapping the file expects.
---
## 5. Frontend design
All in `panels-ui.js` (prettier-exempt, hand-formatted; match the surrounding style), `index.html`,
`styles.css`, `mobile.css`.
### 5.1 State
```js
filePreviewEdit = { active, sessionId, path, baseHash, eol, original, dirty }
```
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
### 5.2 Markup (`index.html:420-432`)
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
untouched.
### 5.3 Behavior
- `openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
keyboard before the user has picked a line.
- `input` sets `dirty` and enables Save.
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
response, leave edit mode, re-render the read view from the local editor value (the response carries
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
itself carries no state (section 3.3, step 9).
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
guard.
- `copyFilePreviewContent()` copies the live editor value while editing.
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
real bug twice.
### 5.4 Mobile
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
the editor tracks the keyboard for free.
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
smaller `rem` value.
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
- No header-policy change is needed (section 1), so
`test/mobile-header-buttons-policy.test.ts` stays untouched.
### 5.5 i18n
`i18n.js` already skips `textarea`, `pre`, `code` and `.file-preview-content` in its `SKIP_SELECTOR`
(`i18n.js:20-38`), so file content is never translated. Add zh-CN entries for the new chrome: Edit, Save,
Cancel, Unsaved changes, Discard unsaved changes?, File changed on disk, Reload, Overwrite, Saved,
Too large to edit here.
---
## 6. Docker and remote cases
Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
---
## 7. Tests
| File | Kind | Covers |
| ------------------------------------------- | ----------- | ---------------------------------------------------------------------- |
| `test/file-editing-policy.test.ts` | pure unit | `isEditableFileName` (allow + deny + basenames), `isDeniedEditRelativePath`, `detectEol`/`applyEol` round-trip incl. mixed EOL, BOM preservation |
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
Route cases to assert explicitly:
1. happy path writes the bytes and returns a new hash
2. `../` and absolute paths give 404
3. symlink pointing outside the workspace gives 404
4. symlink pointing inside is written through to the target
5. non-allowlisted extension gives 400
6. `.git/config` gives 403
7. a `.env` in the workspace gives 403 (sensitive-path)
8. a file with a NUL byte gives 400
9. a latin-1 file that fails the UTF-8 round-trip gives 400
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
11. over `MAX_EDITABLE_BYTES` gives 413
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
14. CRLF file edited and saved stays CRLF
15. file mode is preserved across the temp-plus-rename
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
- `curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
`w1`/`w2`/`w3`; delete it by exact id afterwards.
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
assert the bytes on disk changed. Assert real state, not HTTP 200.
---
## 8. Docs and release
- This plan lives at `docs/file-viewer-edit-plan.md`.
- `docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
chain, the truncation invariant, and why temp-plus-rename.
- `CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
Remember `CLAUDE.md` is prettier-ignored on purpose.
- `docs/api-reference.md`: the new `PUT` and the `edit=1` query.
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
additive API surface, so **COM minor** when it ships.
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
and must pass `npm run format:check`.
---
## 9. Implementation order
Each phase is independently reviewable and leaves the tree working.
1. **Policy module + tests.** `src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
route wiring. (Small.)
2. **Read-for-edit.** `edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
tests. Nothing consumes it yet. (Small.)
3. **Write endpoint.** `FileWriteSchema`, `PUT` handler, `test/routes/file-write-routes.test.ts`. Fully
testable by curl before any UI exists. (Medium, the security-relevant part.)
4. **Desktop UI.** Edit button, textarea swap, Save/Cancel, dirty guard, 409 flow. (Medium.)
5. **Mobile pass.** `mobile.css` sizing against `--app-height`, font size, accessory-bar interaction,
real-device check. (Small but the part that decides whether the feature is actually usable.)
6. **Docs, i18n strings, changeset.**
---
## 10. Open decisions
1. **Editor widget.** Recommend a plain `<textarea>` for v1: zero dependencies, no CSP question, no bundle
growth, and it is the only thing guaranteed to behave with the iOS keyboard. CodeMirror-light with
syntax highlighting is a clean follow-up once the write path is proven. The issue allows either.
2. **Phone entry point.** Edit mode is reachable on a phone through attachment cards and the history
drawer without changing anything. A dedicated toolbar or overview affordance for "browse this session's
files" would make it discoverable, but it is a separate UX change and would need a decision against the
deliberately minimal phone header policy. Recommend deferring it and revisiting after the feature ships.
3. **`svg` editability.** Recommend excluded in v1 (it is deliberately treated as untrusted on the read
side). Easy to add later.
4. **Create / delete / rename.** Explicitly out of scope per the issue. Note that keeping `O_CREAT` out of
the handler is what makes that a structural property rather than a convention.
+3 -3
View File
@@ -1,7 +1,7 @@
# Remote Sessions (SSH)
Codeman can run a session's agent on a **remote host over SSH** instead of the
local machine. The agent (Claude, OpenCode, Codex, Gemini, or a plain shell)
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, or a plain shell)
runs inside a `tmux` server **on the remote host**, so it survives the SSH
connection dropping; Codeman attaches to it the same way it attaches to a local
managed session.
@@ -30,7 +30,7 @@ Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini'>` — the modes that can run remotely. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
@@ -115,7 +115,7 @@ Key points:
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID *is*
the agent. The per-mode command comes from `remote.commands?.[mode]` or
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
`exec codex` / `exec gemini` / `exec bash -l`).
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
pane command is independently quoted, so a `remotePath` with spaces is safe.
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
+250
View File
@@ -0,0 +1,250 @@
# Scrollback fix plan (issue #205)
Status: IMPLEMENTED on `fix/scrollback-shell-alt-screen` (2026-08-07), with one deliberate
divergence from the recommendation below. Kept for the diagnosis record; the measured evidence
behind it is `docs/scrollback-issues-analysis.md`, and the mechanisms as shipped are documented
in `docs/architecture-invariants.md` (§ Full-scrollback replay, § Terminal scrollback: strip
flavors and wheel/touch forwarding).
What shipped vs. what this doc proposed:
- **Bug A (deltaMode)**: implemented as specified (`_wheelScrollLines()` normalizes
line/page/pixel units, Shift-axis trap kept).
- **Bug B (shell scrollback)**: implemented via the NARROW alt-screen strip for tmux-backed
shell/opencode/antigravity plus the scroll-to-top `full=1` re-pull, NOT the recommended
approach (a) `tmux mouse on`. The measurements in the analysis doc showed the alt buffer
comes from tmux's own client-side `smcup` at attach (tmux never forwards a pane program's
alt-screen toggles), so stripping that one sequence fixes both symptoms with no selection
tradeoff, keeps vim/less/htop untouched, and the re-pull also covers the repaint-burst
history loss that `mouse on` would not have addressed.
- **Invariant change**: the "viewport-at-bottom gate stays" invariant below was deliberately
DROPPED for forwarding modes: a repaint-mode CLI keeps no real terminal scrollback, so the
gate pinned users to a buffer of stale frames whenever the viewport parked off-bottom.
Forwarding now snaps to bottom first; Shift+wheel and the opt-out setting keep local
scrollback reachable. Touch forwards through the same gate (the mobile half of the fix).
- **Finding 5 (remote probe)**: implemented (`probeRemoteCliVersion` over ssh, deferred at
session start, same login-shell wrapper as the launch).
## RETEST FAILED (2026-08-07, after v1.12.0 shipped) — analysis round 2
mtiller retested on 1.12.0 and reports it is NOT fixed (issue #205 comment, 2026-08-07 12:12 UTC;
issue reopened same day with clarifying questions: mouse vs trackpad, Shift+scroll behavior,
Claude vs shell session on the phone, and an iOS full-tab-kill to rule out stale JS). Two
failure signatures, now analyzed against the SHIPPED 1.12.0 code (not the pre-fix code):
1. **iPhone Safari (Claude session assumed)**: touch scrollback goes back only a limited
amount and sometimes REPEATS blocks of text; unreliable.
2. **Firefox on macOS (mouse)**: wheel does NOTHING at all, while Fn+Up (= PageUp) pages back
through INTACT text.
### Ruled out by code reading
- deltaMode mishandling: `_wheelScrollLinesFloat` normalizes line/page/pixel units correctly;
a Firefox line-mode notch yields ±3 lines. Not the bug.
- Ephemeral transport: `_sendInputEphemeral` (app.js) has a POST fallback when WS is down.
- Service worker: sw.js is network-first with cache fallback; it serves stale JS only when the
fetch FAILS (flaky mobile connection can do this — relevant to "unreliable" on the phone,
and the fixed `CACHE_NAME = 'codeman-v1'` never invalidates that offline copy).
### The load-bearing observation: PageUp works, the wheel does not
Fn+Up is a KEYBOARD event: xterm encodes PageUp and Claude pages its own transcript (intact
text proves Claude-side history is fine and the PTY input path is fine). The wheel path is the
capture-phase handler, and for a Claude session it has exactly two branches:
- **Forwarding branch** (`_shouldForwardWheelToApp` true): snap-to-bottom + SGR reports. If
this branch ran, the user would see the same paging motion Fn+Up produces. They see nothing.
- **Local branch** (gate false): `_smoothScrollBy` over xterm's local buffer. For a Claude
pane in repaint mode, tmux keeps `history_size≈0`, so `?full=1` returns roughly one frame:
the local buffer is structurally HOLLOW, the top-of-buffer re-pull recovers nothing, and the
wheel looks completely dead. **This matches every observed detail on Firefox.**
So the working hypothesis is that mtiller's sessions evaluate the gate FALSE. The gate
(`_shouldForwardWheelToApp`) has exactly four false-paths worth checking, in likelihood order:
1. **`terminalWheelLocalScrollback` opt-out is ON.** Plausible: a user whose scrolling was
broken on 1.11.x may well have toggled "Wheel scrolls local history" while trying to fix
it. On 1.12.0 that setting now routes the wheel to a hollow local buffer = dead wheel on
desktop AND the stale-repaint-frames experience on the phone (see below). Ask, or check
what the setting does on their export.
2. **`cliVersion` missing — CONFIRMED BUG, independent of whether it is mtiller's**:
`getClaudeCliVersion()` (utils/claude-cli-resolver.ts:124-148) caches its result
process-wide including FAILURE: on any exception it sets `_claudeVersion = null`, and the
guard is `!== undefined`, so a single failed/timed-out probe (5s `EXEC_TIMEOUT_MS`; PATH
under systemd/launchd; transient fs hiccup) at the FIRST Claude session start disables
wheel forwarding for every Claude session until the server restarts. Fix: cache success
permanently, but let failure retry (retry on next call, or a short negative-cache TTL).
Note that mtiller sees identical breakage on phone + iPad + laptop, which points at a
SERVER-side/session-side cause exactly like this (cliVersion is shared by all devices)
rather than anything browser-specific.
3. **Claude Code genuinely < 2.1.187** on their machine: gate false BY DESIGN, but the
resulting UX is a dead-end (no local history to fall back on).
4. mouseTrackingMode non-none (a DECSET leaked past the strip, e.g. emitted before attach or
split across chunks in a way the carry missed): would also kill the container handler via
the early return. Least likely, checkable via `terminal.modes.mouseTrackingMode` in console.
### The iPhone symptoms fit the same gate-false story
Touch with gate false = local `scrollLines()` over whatever repaint frames accumulated:
"repeats blocks of text" is literally what a buffer of successive overlapping repaint frames
looks like; "limited amount" is its thinness; "unreliable" is burst-dependence (finding 2)
PLUS the new re-pull being actively DESTRUCTIVE for repaint panes: `_maybeRefetchFullHistory`
does `_resetTerminalForReplay()` then writes the fetched capture, and when that capture is
one frame (Claude pane, `history_size≈0`) it REPLACES a multi-frame buffer with less than the
user had, mid-scroll. Stale pre-1.12 JS on the phone (suspended Safari tab) remains possible
until they confirm the tab kill.
### Fix directions, ranked
1. **Make the re-pull refuse downgrades** (`_maybeRefetchFullHistory`, app.js): if the fetched
capture would yield FEWER buffer rows than currently present, skip the reset+rewrite and
keep the richer buffer (optionally cache-mark the session "re-pull useless"). Small, safe,
kills the "got worse after scrolling to top" class. Consider skipping the re-pull entirely
for forwarding-capable modes where tmux keeps no history.
2. **Rescue the gate-false Claude dead-end with PageUp forwarding**: when mode is `claude`,
the gate is false, AND the local buffer has no scrollback (`baseY === 0`), translate wheel
lines into coalesced PageUp/PageDown key sends (mtiller just proved Claude pages correctly
on PageUp even on their version). Zero regression risk under that triple guard: sessions
with real local history keep local scrolling; only the currently-dead path changes.
Caveat: older Claude menus may react to PageUp; acceptable against "completely dead".
3. **Audit `getClaudeCliVersion()` failure caching** (utils/claude-cli-resolver.ts): a cached
empty probe must retry (with backoff), not poison the process.
4. **Guard the opt-out setting's footgun**: if `terminalWheelLocalScrollback` is ON for a
repaint-mode CLI session, local history is hollow; either scope the setting's effect to
modes with real local scrollback, or pair it with fix 2's PageUp fallback so it still
scrolls SOMETHING.
5. **Add a one-line gate diagnostic**: log (once per session, console) WHY the wheel chose
local vs forward: `{mode, cliVersion, optOut, trackingMode}`. The #205 thread is now two
rounds deep on guesswork a single console line would have answered.
### What shipped for round 2 (branch `fix/scrollback-205-round2`)
All five directions above, implemented as ranked:
1. **Downgrade guard** — `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the rows a
capture will occupy (ANSI stripped, `capture-pane -J` re-wrapping accounted for) and
`_maybeRefetchFullHistory` (app.js) skips the reset+rewrite when that is more than one
screen short of what xterm already holds. A refused session goes on
`_fullHistoryRepullUseless`, which raises its re-pull cooldown from 4s to 60s so a hollow
pane stops re-fetching. Measured A/B on a live Claude pane, same gesture, same buffer:
guard off → 341 rows collapse to 42 and every seeded row is gone; guard on → 341 rows
preserved. The tab-switch recovery it must not break still runs (shell buffer 401 → 44 on
a tab switch → 401 again after scrolling to the top).
2. **PageUp/PageDown fallback** — `_maybePageCliTranscript()` translates wheel/touch travel
into coalesced `\x1b[5~` / `\x1b[6~` under the triple guard (claude mode, forwarding gate
false, `baseY === 0`), through the same 40ms queue as the SGR reports. Half a screen of
travel per page: the page key always jumps a whole screen, and a 1:1 mapping was
unusably slow with a discrete wheel. Shift is excluded — it keeps meaning "local
scrollback". Verified live: opt-out ON on a Claude session sends real PageUp/PageDown to
the PTY where the wheel previously did nothing.
3. **Probe caching** — `getClaudeCliVersion()` no longer caches failure. Success is kept for
the process lifetime; a failed probe retries with a 1/2/4…15min backoff. The cache policy
is a pure function (`resolveClaudeCliVersion`) so the retry semantics are unit-testable
without spawning `claude`. The VITEST short-circuit now records nothing, where before it
wrote a permanent null.
4. **Opt-out footgun** — handled by pairing rather than by scoping: the setting keeps meaning
exactly what it says (the wheel goes local), and fix 2 catches the case where "local" is
empty. Scoping the setting away from repaint-mode CLIs would have silently overridden an
explicit user choice. The App Settings tooltip now says to leave it off for Claude/Codex.
5. **Diagnostic** — `_logScrollRouting()` prints one line per session per distinct decision:
`[scroll] <id> → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…,
cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. That
single line answers every open question in the list below.
Still unanswered by code alone: whether mtiller's Claude Code is genuinely older than
2.1.187 (false-path 3), and whether the iPhone was running stale JS. The diagnostic makes
both self-reporting, so the retest ask is now "open the console and paste the `[scroll]` line".
### What to get from mtiller (some already asked)
- Shift+scroll behavior on Firefox (distinguishes hollow-local from handler-not-firing).
- `claude --version` on the Mac (decides false-paths 2 vs 3).
- App Settings → Input → "Wheel scrolls local history" state (false-path 1).
- iPhone: Claude or shell session, and whether a full tab kill changes anything.
- Browser console: `app.terminalUi?.terminal?.modes?.mouseTrackingMode` (false-path 4).
Original plan follows.
## Reports
- **Issue #205** (https://github.com/Ark0N/Codeman/issues/205), OPEN:
- **jonocodes** (author, 2026-08-03): SHELL session. Host Mac M4, brew tmux. On Android, touch-scrolling the terminal does nothing. On desktop, the mouse wheel cycles shell command history (acts like Up/Down arrows) instead of scrolling the screen.
- **mtiller** (comment, 2026-08-06): "similar issue just with scrolling backward to see agent output. This is with Firefox on MacOS." (Claude session implied.)
- **Reddit r/selfhosted** comment `p21x6ts` by mmtiller (= mtiller on GitHub): scrolling broken enough across phone/iPad/laptop that they fall back to Claude's own remote-control feature. Churn-risk user who otherwise loves the product; fixing this has promo value beyond the bug itself.
## How scrolling works today (read this before touching anything)
Three independent paths, all in `src/web/public/terminal-ui.js` unless noted:
1. **Desktop wheel** (container `wheel` listener, ~line 421): ALWAYS `preventDefault()`s, then either
- forwards synthetic SGR wheel reports to the app (`_sendSyntheticSgrWheel`, coalesced every 40ms, fire-and-forget) when `_shouldForwardWheelToApp(ev)` (~line 2823) passes: no Shift held, opt-out setting `terminalWheelLocalScrollback` off, xterm `mouseTrackingMode === 'none'`, session mode is `claude` with `cliVersion >= 2.1.187` or `codex`, and viewport is at bottom;
- otherwise scrolls xterm's LOCAL scrollback via `terminal.scrollLines(lines)`.
- `lines` comes from `_wheelScrollLines(ev)` (~line 2818): `delta / 25`, i.e. it assumes PIXEL deltas.
- NOTE: xterm.js's own internal wheel handler sits on an element INSIDE the container, so it runs FIRST (bubble order) and is not suppressed by the container's `preventDefault`.
2. **Touch** (touchstart/move/end, ~lines 441-585): converts touch deltas to `terminal.scrollLines()` with momentum. Touch is ALWAYS local-scrollback, never forwarded to the app. Tap-to-position (touchend, ~line 533) is separate and already handles both mouse-tracking-on and server-strip cases.
3. **Server-side strip** (`_handleTerminalOutput`, `src/session.ts:1384`): for modes in `isAltScreenStripMode()` (`src/session.ts:179` = `codex | claude | gemini`), strips alt-screen switches (`?47/?1047/?1049`), scrollback erase (`3J`), and mouse-tracking DECSETs (`?1000-?1007` except `?1004` focus) so content stays in xterm's normal buffer with scrollback intact. Includes a chunk-boundary carry so split sequences can't leak. `shell` and `opencode` (and `antigravity`) are deliberately EXCLUDED: arbitrary shell programs (vim/less/htop) legitimately need the alt screen. There is a parity copy of this strip on the replay path (`src/web/routes/session-routes.ts`, ~line 1697) and a frontend parity check `_sessionUsesServerMouseStrip()` (terminal-ui.js ~line 2751). All three must stay in sync.
4. Related: full-scrollback replay (`GET .../terminal?full=1` on first buffer load) fills xterm local scrollback; client scrollback is hardcoded 50k (`DEFAULT_SCROLLBACK`, constants.js) vs tmux 100k.
## Diagnosis
### Bug A: Firefox wheel deltas (mtiller's desktop case)
`_wheelScrollLines()` divides by 25 assuming `WheelEvent.deltaY` is pixels (`deltaMode === 0`, Chrome/Safari behavior). Firefox commonly fires `deltaMode === 1` (LINE units, deltaY around 1-3 per notch), so `Math.round(3/25) = 0` and the `|| ±1` fallback yields 1 line per event. With a discrete mouse wheel that is 1 line per notch: scrolling feels dead/broken. This hits BOTH the local-scroll path and the forwarded path, since both use the same function.
**Fix**: normalize by `ev.deltaMode` in `_wheelScrollLines()`:
- `deltaMode 0` (pixels): current behavior, `delta / 25`.
- `deltaMode 1` (lines): use the delta directly (round, keep sign fallback).
- `deltaMode 2` (pages): `delta * terminal.rows` (or a sane page size).
Keep the existing Shift-axis trap intact: on macOS trackpads Shift+two-finger scroll arrives as a HORIZONTAL wheel (deltaX carries the magnitude, deltaY ~0); that's why the function reads deltaX when Shift is held (issue #154). Don't lose it.
**Verify**: don't trust this diagnosis blindly. First reproduce in real Firefox on macOS and log `deltaMode`/`deltaY` (Firefox trackpad input can arrive as pixels; external mouse as lines). Also confirm the session's `cliVersion` probe succeeded (a failed probe disables forwarding entirely, which would point elsewhere). Unit-test by dispatching synthetic `WheelEvent`s with explicit `deltaMode` values; a Playwright `firefox` project pass is the end-to-end check.
### Bug B: shell mode has NO working scrollback at all (jonocodes)
Chain: shell mode is excluded from the alt-screen strip (correctly) → tmux attaches on the alternate screen → xterm's alt buffer has zero scrollback. Consequences:
- **Wheel**: xterm's own internal wheel handler runs first and, in the alt buffer, converts wheel ticks into Up/Down arrow keys (alternateScroll behavior). The shell receives arrows → command history cycles. That is jonocodes' exact desktop symptom. The container handler's `scrollLines()` afterwards is a no-op (no scrollback in alt buffer).
- **Touch**: the touch handler's `scrollLines()` is equally a no-op → "scrolling does nothing" on Android. Exact symptom two.
- The real history exists the whole time in tmux's 100k-line buffer; nothing exposes it.
**Fix, recommended approach (a): enable tmux `mouse on` for shell sessions.**
- Server-side, set `mouse on` scoped to shell sessions' tmux sessions (`tmux set-option -t <session> mouse on` at create + on attach of recovered sessions). Do NOT set it globally on the socket: claude/codex/gemini sessions rely on the DECSET strip and must not change.
- What this buys, all natively: tmux enables mouse tracking on the outer terminal → xterm `mouseTrackingMode` goes non-none → the container handler stands down (line ~2830 check) and xterm's own encoder forwards wheel as SGR reports → tmux scrolls its OWN copy-mode history on wheel-up, auto-exits at bottom. The alt-scroll arrow conversion disappears too (tracking mode takes precedence). Desktop is fully fixed with no new endpoints.
- **Touch**: still needs one small client change: in the touchmove path, when the active session is `shell` AND `mouseTrackingMode !== 'none'`, convert accumulated lines to `_sendSyntheticSgrWheel(x, y, lines)` instead of `scrollLines()`. The 40ms coalescing already prevents the tmux process storm (each send is a tmux send-keys server-side; unbatched flicks would spawn dozens of processes: this constraint is documented at `_sendSyntheticSgrWheel`, do not bypass it).
- **Selection tradeoff to verify**: with tracking on, xterm hands drag events to tmux instead of doing local browser selection. Shift+drag still does local selection (xterm shift-override). Verify this UX on desktop before shipping; if it's unacceptable, fall back to approach (b).
- **Also verify**: vim/less/htop inside the shell still behave (they'll now receive real mouse events via tmux, generally an improvement); remote shell sessions run tmux on the REMOTE host (`tmux -L codeman-remote`) and need the same option set there if remote shells are in scope (fine to defer, note it in the changeset if skipped).
**Fallback approach (b), only if (a)'s selection tradeoff fails testing**: keep mouse off; when a shell session is in the alt buffer, have the client send scroll intents to a small server endpoint that drives `tmux copy-mode -e -t <pane>` + `send-keys -X -N <n> scroll-up/down`. Preserves selection semantics exactly, but needs a new endpoint, server-side batching, AND suppression of xterm's native alt-scroll arrow conversion (capture-phase wheel listener with `stopPropagation`, or `attachCustomWheelEventHandler` if the vendored xterm version has it). More moving parts; (a) should be tried first.
**Not acceptable**: adding `shell` to `isAltScreenStripMode()`. vim/less/htop need the alt screen; that exclusion is deliberate and documented.
### Bug C: mtiller's phone/iPad case — UNREPRODUCED, do not guess
Touch is always-local by design, and Claude sessions keep content in the normal buffer (strip), so touch scrollback "should" work there. Before coding anything: build a repro matrix (iPhone Safari / iPad Safari / Android Chrome × claude / shell) on the current release. Plausible candidates if it does reproduce: auto-scroll-to-bottom fighting user scrolls (`_noteTerminalUserScroll`, ~line 2004), or they were in shell sessions on mobile too (then Bug B covers it). Ask mtiller on #205 for session mode + Codeman version if the matrix comes up clean.
## Invariants the implementation MUST respect
- Shift+wheel always scrolls local scrollback; the trackpad Shift-axis handling from #154 stays.
- The `terminalWheelLocalScrollback` opt-out setting keeps working (pins plain wheel to local).
- The viewport-at-bottom gate stays: once the user scrolled up locally, wheel stays local until they return to bottom.
- 40ms SGR coalescing: never send per-event writes to the server.
- Strip parity triangle: `session.ts` live strip ↔ `session-routes.ts` replay strip ↔ `_sessionUsesServerMouseStrip()` in the frontend. If you touch mode lists, update all three.
- Don't add `opencode`/`antigravity` to any strip/forward list; their TUI wheel behavior is unverified (documented at `_shouldForwardWheelToApp`).
- The chunk-boundary sequence carry in `_handleTerminalOutput` must not be weakened.
## Testing (per repo rules)
- `npm test -- test/<file>.test.ts` only; never bare `npm test`. New test ports 3150+, never 3000.
- Browser-test traps (documented in CLAUDE.md Testing): drive input/scroll through real events (`page.mouse.wheel`, real touch), not app internals; headless Chromium reports `isTouchDevice()` false even with `hasTouch: true`; assert on real state (xterm viewport position, `tmux -L codeman capture-pane`), not HTTP 200.
- Shell-mode E2E: create a throwaway shell session, `seq 1 500`, then (1) wheel up on desktop shows earlier lines, not history cycling; (2) touch-scroll on a phone shows earlier lines; (3) `vim` + `less` still enter/leave the alt screen cleanly; (4) Shift+drag still selects text.
- Firefox E2E: Playwright `firefox` project, wheel over a Claude session's finished output, assert viewport moved more than 1 line per notch.
- End-to-end against the REAL environment before claiming done (standing user rule). w1/w2/w3 tmux sessions are the user's live sessions: never send input to them; create your own throwaway session and DELETE it by exact id when done.
## Related observation (not a reported bug, worth a look while in there)
The `claude --version` probe that feeds the forwarding gate runs only for local and docker sessions (`src/session.ts:1490` gates `!this._remote`; docker handled at :1507). Remote Claude sessions therefore never get `cliVersion` and silently keep local-only wheel. Harmless (local scrollback works) but inconsistent; cheap to fix by probing over ssh, or document as intended.
## Rollout
1. Bug A (deltaMode) is small and independent: can ship alone as a patch.
2. Bug B (shell scrollback) is the headline fix for #205: patch or minor per COM flow.
3. After deploy + verification: comment on #205 (what was fixed, what needs their retest), then reply to the Reddit comment `p21x6ts` with the release version. Both reporters gave environment details; address them specifically.
+255
View File
@@ -0,0 +1,255 @@
# Scrollback issues: analysis and test evidence
Covers GitHub issue **#205** ("Scrollback in terminal not working", jonocodes, shell mode,
Android + macOS desktop) and the follow-up comment on it from **mtiller** (Firefox on macOS,
"scrolling backward to see agent output"). Related closed issue: **#154** (fixed in 1.3.3).
Status: **analysis only, nothing implemented.** Measured against the live 1.11.2 instance on
2026-08-06 with throwaway `zz-*` shell sessions (all deleted afterwards; the user's `w*`
sessions were never touched).
---
## TL;DR
Five distinct problems, not one. #205 is fully explained by finding 1; findings 2 and 3 are
independent and hit **every** mode including Claude, and are the likely substance of the
"similar issue" follow-up.
| # | Problem | Modes affected | Severity | Confirmed |
| - | ------- | -------------- | -------- | --------- |
| 1 | xterm parked in the **alternate buffer** for the whole session, so there is no scrollback at all and the wheel is translated into Up/Down arrow keys | `shell`, `opencode`, `antigravity` | High | Reproduced end to end |
| 2 | **Bursty output silently destroys a screenful** of the browser's scrollback and adds ~1 row | all | High | Measured |
| 3 | **Tab switch collapses scrollback** to roughly one screen (`full=1` fires once per page load) | all | Medium | Measured |
| 4 | `deltaMode` is never read, so Firefox scrolls ~4x slower per notch | all, Firefox | Low | Static, needs reporter data |
| 5 | **Remote SSH Claude cases get no `claude --version` probe**, so wheel forwarding silently stays off (residual #154) | `claude` + remote | Medium | Static |
---
## Finding 1: shell / opencode / antigravity are stuck in xterm's alternate buffer
### Root cause
The local tmux **client** (the `tmux attach` that node-pty spawns) emits `smcup` as its very
first bytes on attach. Captured from a real PTY:
```
b'\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2J\x1b[?12l\x1b[?25h\x1b[?1000l...'
^^^^^^^^^^ enter alternate screen ^^^^^ application cursor keys ON
```
`Session._handleTerminalOutput()` strips `\x1b[?1049h` from the live stream, but only when
`isAltScreenStripMode(mode)` is true, and that is `claude | codex | gemini` only
(`src/session.ts:179`). For `shell`, `opencode` and `antigravity` the sequence reaches the
browser verbatim and xterm switches to the alternate buffer, where:
1. `buffer.active.type === 'alternate'` and `baseY` is pinned at 0, so there is **no
scrollback to reach**. `terminal.scrollLines()` is a no-op, which is why touch scrolling
on Android "does nothing".
2. xterm's own wheel listener takes over. From the vendored bundle
(`src/web/public/vendor/xterm.min.js`):
```js
if (!this.buffer.hasScrollback) {
if (ev.deltaY === 0) return false;
if (coreMouseService.consumeWheelEvent(...) === 0) return this.cancel(ev, true);
const seq = ESC + (decPrivateModes.applicationCursorKeys ? 'O' : '[') + (ev.deltaY < 0 ? 'A' : 'B');
coreService.triggerDataEvent(seq, true);
return this.cancel(ev, true);
}
```
tmux also set `\x1b[?1h`, so the emitted sequence is `\x1bOA`, i.e. **Up arrow**, straight
into the shell's readline. That is exactly the reported "the mouse wheel scrolls back
through previous commands, like pressing up".
3. `cancel(ev, true)` calls `preventDefault()` **and `stopPropagation()`**, and xterm's
listener sits on `terminal.element` (a child of Codeman's container). So Codeman's own
container wheel handler, `_shouldForwardWheelToApp` and `_wheelScrollLines` included, is
**never reached** for these modes. That whole path is dead code for shell.
### Reproduction (live instance, real browser)
Create a shell session with the page already open, print 150 lines, then dispatch 8 wheel-up
events over `.xterm-screen`:
```
t+1500 after shell start {"type":"alternate","length":35,"baseY":0}
t+3000 after shell start {"type":"alternate","length":35,"baseY":0}
after 150 live lines {"type":"alternate","length":35,"baseY":0}
WHEEL on live shell: {"ptyBytes":["OA","OA","OA","OA",
"OA","OA","OA","OA"],
"before":0,"after":0,"type":"alternate"}
```
Both reported symptoms, one root cause.
### Why it looks intermittent
The alternate-screen sequence only ever reaches the browser through the **live stream at
attach**. Neither replay path carries it:
- `?full=1` returns `capture-pane` output (`source: mux-full-history`), verified 0 hits for
`\x1b[?1049h`.
- `?tail=` returns the visible pane frame (`source: mux-visible`), also 0 hits; the shell byte
buffer was empty in every probe.
- `_resetTerminalForReplay()` calls `terminal.reset()`, which returns xterm to the normal
buffer.
So: watching a shell from creation leaves you in the alternate buffer until you reload or
switch tabs, at which point it silently starts working again. Then the next PTY attach (a
restart, or the auto-reattach in `selectSession()`) puts you back.
### Is stripping safe for shell? Probably yes when tmux-backed, and the current code comment is wrong about why
`src/session.ts:1404` says *"shell must keep the alt screen for vim/less/htop"*. For a
**tmux-backed** shell that reasoning does not hold: tmux is a full terminal emulator and never
forwards a pane's alternate-screen toggles to its client, it repaints instead. Measured per
phase on a real attach:
| phase | bytes | `?1049h` | `?1049l` | `?47/1047` |
| ----- | ----: | -------: | -------: | ---------: |
| attach | 772 | **1** | 0 | 0 |
| `seq 1 60` echo | 1402 | 0 | 0 | 0 |
| `less` open / end / quit | 284 / 230 / 321 | 0 | 0 | 0 |
| `vim` open / quit | 2200 / 646 | 0 | 0 | 0 |
`vim` and `less` inside tmux emit **zero** alternate-screen sequences to the client.
The caveat that does matter: `startShell()` falls back to a **direct PTY with no tmux** when
mux creation fails (`src/session.ts:1961`, `this._useMux = false`). In that path the inner
app's own `?1049h` does reach xterm, and a blanket strip would break vim/less/htop for real.
Any fix has to be conditional on `_useMux`, which is known server-side.
Second caveat: stripping alone buys less than it looks like, because of finding 2. It fixes
the wheel (no more phantom Up arrows) and it makes the `full=1` replay reachable, but live
output still will not accumulate.
---
## Finding 2: bursty output silently overwrites a screenful of browser scrollback
Independent of the alternate buffer, and it hits Claude sessions too.
tmux decides per flush whether to emit real linefeeds (which push rows into the outer
terminal's scrollback) or to repaint the pane rectangle with cursor addressing (which
overwrites the visible rows in place). When output outpaces its flush interval it coalesces
into a repaint, and one screenful of the browser's history is **destroyed**.
Measured on one session, same page, `rows = 36`:
| step | `baseY` | rows containing SEED | BURST | SLOW |
| ---- | ------: | -------------------: | ----: | ---: |
| after `?full=1` replay (120 seeded lines) | 86 | 120 | 0 | 0 |
| after 60 lines emitted as fast as possible | **87** (+1) | **86** (-34) | 35 | 0 |
| after 60 lines at ~16/s (`sleep 0.06`) | **148** (+61) | 86 | 35 | 60 |
The burst added **one** row of scrollback and ate **34** rows of existing history. The slow
run behaved correctly. So "I printed a bunch of lines and now I cannot scroll back" reproduces
without the alternate buffer being involved at all, and it is rate dependent, which is exactly
the kind of thing that reads as random flakiness.
Consequence: the browser's scrollback is effectively frozen at whatever the last `?full=1`
replay produced, minus a screen per burst. tmux's own history is fine throughout
(`history_size` kept growing, `history-limit` 2000), so the data is never actually lost
server-side, it just never reaches the browser again until a reload.
---
## Finding 3: switching tabs collapses a session's scrollback
`_initialFullBufferLoad` is true for the **first buffer load after a page load only**
(`app.js:4374`). Everything after that uses `?tail=`, which returns byte history plus the
visible pane frame. Worse, the snapshot restore path deliberately throws away the restored
xterm snapshot (which does carry scrollback) and replaces it with that frame
(`app.js:4316-4328` plus `needsRewrite`).
Measured, switching away from session A and back:
```
A: initial full=1 load {"len":152,"baseY":116,"AAA":150}
A: after switch away and back {"len": 87,"baseY": 51,"AAA": 59}
```
150 lines of history down to 59. Note also that the page's single `full=1` is consumed by
whichever session auto-selects at load, so **every other tab starts life with one frame of
history**.
---
## Finding 4: `deltaMode` is never read (Firefox)
`grep -rn "deltaMode" src/web/public packages` returns nothing. `_wheelScrollLines()`
(`terminal-ui.js:2818`) treats `deltaY` as pixels unconditionally:
```js
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
```
Chrome/WebKit report `deltaMode: 0` with `deltaY` around 100 to 120 px per notch, so about 4
to 5 lines. Firefox reports `deltaMode: 1` (`DOM_DELTA_LINE`) with `deltaY` around 3, so
`Math.round(3/25) === 0` and the `|| ±1` fallback yields **1 line per notch**, roughly 4x
slower. In Claude mode the same value caps the forwarded SGR report at 1 tick per event
instead of 4, so the transcript crawls too.
This is sluggishness, not breakage, so it is a plausible but unproven contributor to the
mtiller report. No Firefox build is installed under `~/.cache/ms-playwright` (chromium and
webkit only), so this was not measured. Worth asking the reporter for `deltaMode` / `deltaY`
from a live wheel event before acting on it.
---
## Finding 5: remote SSH Claude cases still have no version probe
`src/session.ts:1490` deliberately skips the deterministic `claude --version` probe for
remote sessions and defers to the startup-banner scrape, which the same comment block
describes as unreliable ("newer Claude Code builds don't print the banner and resumed sessions
never show it"). That is precisely the condition #154 was filed for: `cliVersion` empty means
`_shouldForwardWheelToApp()` returns false, wheel forwarding is off, and the user is left with
local scrollback that (per finding 2) does not accumulate.
Local and Docker Claude sessions are fine; verified all 7 live sessions report
`cliVersion=2.1.223`, so the 1.3.3 fix is still working there.
---
## Candidate directions (not decided)
Roughly in order of value per unit of risk.
1. **Extend the alternate-screen strip to tmux-backed `shell` / `opencode` / `antigravity`.**
Gate on `_useMux` so the direct-PTY fallback keeps vim/less/htop working. Kills the phantom
Up arrows and makes replayed history reachable. `isAltScreenStripMode()` currently takes
only `mode`, so it would need the mux flag threaded in, and
`test/claude-scrollback-strip.test.ts:16-17` plus `test/antigravity-mode.test.ts:116` pin
the current answers and would need updating.
2. **Re-pull `?full=1` when the user scrolls to the top of the buffer.** Directly addresses
findings 2 and 3 with machinery that already exists and is already proven to return
complete history (200/200 lines in the probe). Needs a guard against refetch storms.
3. **Stop discarding the xterm snapshot on tab switch**, or request `full=1` on the first load
per session rather than per page. Cheaper partial fix for finding 3 alone.
4. **Read `ev.deltaMode`** in `_wheelScrollLines()` and normalise line/page deltas to lines.
Small, self-contained, worth doing regardless of whether it is mtiller's actual bug.
5. **Probe the CLI version over SSH for remote Claude cases**, mirroring the deferred
in-container probe that Docker cases already use.
Option 1 alone does not fix #205's "print a bunch of lines then scroll" complaint; that needs
2 as well.
## Reproduction assets
Scripts used, in the session scratchpad
(`/tmp/claude-1000/-home-arkon-default-claudeman/597ffc9f-.../scratchpad/`):
- `ptycap.py` / `ptycap2.py`: PTY-level capture of the tmux client stream, per phase counts of
alternate-screen and mouse-tracking sequences.
- `sim.mjs`: replays a captured stream through `@xterm/headless` with and without the strip.
- `browser-test*.mjs`: Playwright against the live instance, reports `buffer.active.type`,
`baseY`, row content and the exact bytes xterm sends to the PTY on a wheel event.
`@xterm/headless` was installed with `npm i --no-save`, so `package.json` and the lockfile are
untouched.
+19 -7
View File
@@ -249,16 +249,28 @@ Ordered most‑to‑least recommended:
### A. Tailscale serve (recommended)
Bind loopback, let Tailscale front it on your tailnet with a real cert:
Bind loopback, let Tailscale front it on your tailnet with a real cert. **The
installer sets this up for you**: choose **Tailscale** at the network-access
prompt, or retrofit an existing install with:
```bash
codeman web --https # binds 127.0.0.1:3000
tailscale serve --bg https / http://127.0.0.1:3000
bash ~/.codeman/app/install.sh tailscale
```
Only devices on your tailnet can reach it; Tailscale handles identity. No app
password and no `0.0.0.0` bind required. (This is the maintainer's production
setup.)
The guided flow installs Tailscale if needed, walks through login and the
tailnet HTTPS-certificates toggle, and configures the equivalent of:
```bash
codeman web # binds 127.0.0.1:3000 (plain HTTP is fine here)
tailscale serve --bg 3000 # HTTPS at https://<node>.<tailnet>.ts.net
```
Only devices on your tailnet can reach it; Tailscale handles identity and
terminates TLS with a real Let's Encrypt certificate (so PWA install and web
push work). No app password and no `0.0.0.0` bind required. (This is the
maintainer's production setup.) `CODEMAN_TAILSCALE=1` presets the choice for
automation; the installer never runs `tailscale serve reset` and never touches
serve mappings other than `443 -> Codeman's port`.
### B. Authenticated cloudflared tunnel + password
@@ -477,7 +489,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini`, `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — and `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
+236
View File
@@ -0,0 +1,236 @@
# Tailscale Setup in the Installer (Plan)
Goal: make "Codeman over Tailscale, with real HTTPS" a first-class, guided path in
`install.sh`, instead of a one-line hint pointing at the docs. Today the safest
recommended deployment (loopback bind + `tailscale serve`) is exactly what the
maintainer's own prod runs, but a new user has to discover and wire it by hand.
The installer should do it for them.
Status: IMPLEMENTED (2026-08-04). `install.sh` carries the 3-way network
prompt, the guided Tailscale flow, and the `tailscale` subcommand; README,
`docs/security-architecture.md` section A, and CLAUDE.md are updated. Verified
live on the maintainer's prod host: `install.sh tailscale` took the idempotent
kept-as-is path against the existing serve mapping (recognizing the legacy
`https+insecure://` target), verified `https://<node>.ts.net/api/status`
end-to-end, and left `tailscale serve status` byte-identical. Items 1-4, 7,
and 10-12 of the manual matrix below still need a fresh machine to exercise.
## Why this is low-hanging fruit
Everything on the app side already works; this is almost purely installer UX:
- `.ts.net` is already in `DEFAULT_TRUSTED_HOST_SUFFIXES`
(`src/web/network-auth-policy.ts`), so the always-on Host/Origin guard accepts
`tailscale serve` traffic with zero configuration. No `CODEMAN_ALLOWED_HOSTS`
needed.
- The loopback bind is the server default and prints no warning; nothing to
acknowledge, no `CODEMAN_PASSWORD` strictly required (the tailnet is the auth
boundary; Tailscale authenticates the device before a packet ever reaches us).
- `tailscale serve` terminates TLS with a real Let's Encrypt certificate for
`<node>.<tailnet>.ts.net`. That gives users valid HTTPS with no self-signed
cert warnings, and (because it is a proper secure context) working service
worker, PWA install, and web push on phones. This is strictly better than
`codeman web --https` for remote access.
- SSE and WebSockets work through serve (proven by prod:
`https://tnode.tailf80371.ts.net` fronting `127.0.0.1:3000` daily).
- `docs/security-architecture.md` section "A. Tailscale serve (recommended)"
already documents this as the preferred setup; the installer just does not
implement it.
## UX design
### 1. The network-access prompt grows a Tailscale option
`choose_network_binding()` (install.sh:1051) currently offers two choices. New
menu, with Tailscale first when it can be recommended:
```
Network access
How should the Codeman dashboard be reachable?
1) Tailscale (recommended)
Private VPN access from your phone/laptop, real HTTPS,
no password needed. Works from anywhere, not just your Wi-Fi.
2) Any device on your network (0.0.0.0)
Open it straight from your phone or laptop on the same Wi-Fi.
Less safe: set a password so only you control your agents.
3) This machine only (127.0.0.1)
Safest. Reach it remotely via Tailscale or a tunnel later.
```
Choice mapping:
- Option 1 = bind `127.0.0.1` (unchanged server posture) + configure
`tailscale serve`. Internally it is option 3 plus the serve setup, so all
existing binding plumbing (`BIND_HOST`, service files, `read_existing_binding`)
is untouched.
- Options 2 and 3 behave exactly as today (renumbered).
- Default choice: 1 when tailscale is installed and logged in, or when an
existing serve mapping for our port is detected; otherwise keep today's
defaults (1 -> 2, 2 -> 3 renumbering, preserving the "existing setup wins"
rule). If tailscale is not installed, option 1 is still shown (the installer
offers to install it), but the default stays on the current behavior so a
bare Enter never pulls in new software.
- Password: after choosing Tailscale, offer the password prompt as optional
defense in depth with default skip ("the tailnet already authenticates your
devices; add one anyway?"). No `BIND_ACK` needed since the bind is loopback.
### 2. The Tailscale flow (state machine)
New `setup_tailscale_access()` runs after the binding choice, before service
setup, handling each state in order:
1. **Not installed.**
- Linux: offer to run the official installer
(`curl -fsSL https://tailscale.com/install.sh | sh`), which handles all
distros and enables `tailscaled` at boot. This mirrors our own
curl-pipe-bash story and avoids maintaining per-distro logic like the six
`install_cloudflared_*` functions.
- macOS: do not auto-install (the GUI app needs an interactive login).
Offer `brew install --cask tailscale` when brew exists, else print the
download link, then wait-and-retry or let the user skip.
- Declined install => fall back to plain loopback (option 3 behavior) and
print how to redo this later (`install.sh tailscale`, see below).
2. **Installed but logged out** (`tailscale status --json` ->
`.BackendState == "NeedsLogin"` or `"Stopped"`).
- Run `tailscale up` (via `run_as_root` if needed). It prints an auth URL
that works headless (user opens it on any device). Poll
`.BackendState == "Running"` with a friendly spinner + timeout; on
timeout, skip gracefully with re-run instructions.
3. **Running: grant operator (Linux).** `sudo tailscale set --operator=$USER`
so serve configuration (now and in the future) does not need root. Skip
silently if we are already operator (probe: `tailscale serve status`
exits 0) or sudo is declined; fall back to `run_as_root tailscale serve ...`.
4. **HTTPS availability check.** `.CertDomains` empty or
`.CurrentTailnet.MagicDNSEnabled == false` means the tailnet has not enabled
MagicDNS / HTTPS certificates. Print the exact two toggles with the admin
URL (https://login.tailscale.com/admin/dns: enable MagicDNS, then enable
HTTPS Certificates), then offer "I enabled it, re-check" / "skip for now".
No silent HTTP fallback: the pitch is real HTTPS, and a plain-HTTP serve
would break the PWA/push story. Skipping falls back to loopback + re-run
instructions.
5. **Existing serve config check** (`tailscale serve status --json`).
- Already proxying to our port (443 -> `127.0.0.1:$PORT`): keep it, report
it, done. Re-running the installer must be idempotent.
- Port 443 occupied by a DIFFERENT target: never clobber it. Ask whether to
replace it or skip. (Prod itself has a second serve on :5000; blind
`tailscale serve reset` would destroy user config. NEVER use `reset`.)
6. **Configure.** `tailscale serve --bg $PORT` where `$PORT` is the install's
Codeman port (default 3000; honor a preset `CODEMAN_PORT`). Serve targets
plain HTTP on loopback; TLS terminates at tailscaled with the real cert.
The `--bg` config persists in tailscaled state across reboots, so no extra
service unit is needed.
(Note: do NOT combine this with `codeman web --https`; that is what forces
the awkward `https+insecure://` proxy target prod historically used. New
installs should keep Codeman on plain HTTP behind serve.)
7. **Verify end-to-end.** Derive the URL from `.Self.DNSName` (strip the
trailing dot) and curl `https://<dnsname>/api/status` after the service is
up, retrying for ~30s: the first request can be slow while the Let's
Encrypt cert is issued. Print success with the URL, or the observed error
with `tailscale serve status` output on failure. This follows the "always
test before claiming it works" rule; a blind "done!" is not acceptable.
### 3. Closing summary and security notice
- The final summary gains a "Remote Access (Tailscale)" block, printed above
the cloudflared block, showing the actual URL:
```
Remote Access (Tailscale):
https://tnode.tailf80371.ts.net (any device on your tailnet, HTTPS)
tailscale serve status # inspect
```
- `print_security_notice()` third branch (loopback) gets a variant: when a
serve mapping for our port is detected, lead with "reachable on your tailnet
at https://... (HTTPS, tailnet-only)" instead of the generic "do ONE of"
list. Detection is dynamic (query `tailscale serve status --json` at print
time), no marker persisted anywhere: tailscaled's own state is the single
source of truth, so external changes never drift against a stale flag.
### 4. Standalone entry point: `install.sh tailscale`
Add a `tailscale` subcommand next to `update` / `uninstall` in the existing
dispatch. It runs `setup_tailscale_access()` against the already-installed
service (reads the port from the service file, requires an existing install).
This serves:
- existing installs that predate the feature,
- users who picked "this machine only" and changed their mind,
- every "skip for now" branch above, all of which print this exact command.
One implementation, two entry points. No separate `scripts/tailscale-setup.sh`
(unlike cloudflared, there is no long-running process for a `tunnel.sh`-style
start/stop wrapper to manage; tailscaled owns the lifecycle).
### 5. Non-interactive / automation
- `CODEMAN_TAILSCALE=1` presets choice 1 (analogous to presetting
`CODEMAN_HOST`). In non-interactive runs it only proceeds through states
that need no human (already installed + logged in + HTTPS-enabled tailnet);
anything requiring interaction (login URL, admin-console toggle, replacing a
foreign serve mapping) warns and falls back to loopback. It never installs
tailscale non-interactively.
- `CODEMAN_NONINTERACTIVE=1` with an existing serve mapping: preserve it, same
"never silently loosen/change" policy as `read_existing_binding`.
- Document both in the header comment block of install.sh (the env-var
reference at the top) and in the README.
## Edge cases and decisions
| Case | Decision |
| ---- | -------- |
| macOS GUI app without `tailscale` on PATH | `get_tailscale_path()` helper mirroring `get_cloudflared_path()`: check PATH, then `/Applications/Tailscale.app/Contents/MacOS/Tailscale`. All calls go through it. |
| Tailnet HTTPS certs disabled | Guided admin-console instructions + re-check loop; skip falls back to loopback. Never configure plain-HTTP serve. |
| Port 443 serve exists for another app | Prompt replace/skip; never `tailscale serve reset` (destroys unrelated mappings). |
| First cert issuance latency | Verify step retries ~30s and says why the first load may be slow. |
| `tailscale up` needs auth | Print the auth URL prominently, poll with timeout, skip gracefully. Works headless. |
| Custom `CODEMAN_PORT` | Serve target uses the actual port; `install.sh tailscale` re-reads it from the service file. |
| Funnel (public internet) | OUT OF SCOPE for v1. If ever added it must mirror the tunnel guard: refuse without `CODEMAN_PASSWORD` (`isUnauthenticatedNetworkAcknowledged`). Funnel exposes to the whole internet and is a different risk class than tailnet-only serve. Mention `tailscale funnel` in docs only, with the password warning. |
| Uninstall | Best effort: if `serve status --json` shows 443 proxying to our port, run the targeted `tailscale serve --https=443 off` (still accepted by current CLIs); if the CLI rejects it, print manual instructions. Never touch other mappings, never uninstall tailscale itself. |
| User already fronting Codeman some other way (reverse proxy etc.) | The serve check only looks at tailscale state; other proxies are invisible and unaffected (same stance as the loopback-exemption note in security-architecture). |
## What does NOT change
- Server code: no changes required. Host guard already trusts `.ts.net`,
loopback bind is already the default, SSE/WS already work through serve.
- The two existing binding options and their semantics, `read_existing_binding`
preservation, and the LAN+password flow.
- `scripts/tunnel.sh` / cloudflared support (stays as the "no Tailscale
account" alternative).
- The security model: this feature only ever narrows exposure (loopback +
authenticated overlay), never widens it.
## Files touched (implementation inventory)
| File | Change |
| ---- | ------ |
| `install.sh` | New: `check_tailscale`, `get_tailscale_path`, `tailscale_status_field` (jq-free JSON field extraction; the installer cannot assume jq: use `sed`/`grep` like existing helpers or `tailscale status --json` piped to `node -e` since node is guaranteed post-install), `offer_install_tailscale`, `ensure_tailscale_login`, `ensure_tailscale_operator`, `ensure_tailnet_https`, `setup_tailscale_serve`, `verify_tailscale_access`, `setup_tailscale_access` (orchestrator). Modified: `choose_network_binding` (3-way menu), summary block, `print_security_notice`, subcommand dispatch (`tailscale`), `uninstall` (targeted serve removal), header env-var docs (`CODEMAN_TAILSCALE`). |
| `README.md` | Remote-access section: promote the Tailscale path with the one-liner and `install.sh tailscale`; keep the tailscale-IP HTTP note for non-serve users but recommend serve + HTTPS. |
| `docs/security-architecture.md` | Section A gains "the installer can set this up for you" + `install.sh tailscale` pointer. |
| `CLAUDE.md` | One line in Scripts & Tunnel: installer offers Tailscale setup (`install.sh tailscale` to redo). |
| `test/` | No unit tests possible for interactive bash + a live tailnet; guard with `shellcheck install.sh` (already the norm) and the manual matrix below. |
## Manual test matrix (before release)
1. Linux + tailscale absent: install offered, declined => loopback fallback + hint.
2. Linux + tailscale absent: install accepted => full flow => URL verified.
3. Logged out => auth URL flow => Running => serve configured.
4. Tailnet with HTTPS certs disabled => guided instructions => re-check => success; and the skip branch.
5. Re-run installer with serve already configured => idempotent, preserved, reported.
6. Second serve mapping on another port present => untouched (prod-like state).
7. Port 443 already proxying another target => replace/skip prompt honored.
8. `install.sh tailscale` on an existing loopback install (the retrofit path).
9. `CODEMAN_NONINTERACTIVE=1` re-run => preserves everything, no prompts.
10. macOS (Mac mini `arbbot` box): GUI-app CLI path detection + full flow.
11. Uninstall removes only our 443 mapping, leaves others.
12. Phone check: PWA install + push from the `https://*.ts.net` origin.
## Release
Changeset: `minor` (new documented installer capability + new `CODEMAN_TAILSCALE`
env var). The feature is installer-only, so it ships with zero risk to running
servers; `install.sh update` does not invoke the new flow (updates never rewrite
access config), only fresh installs and the explicit `install.sh tailscale`
subcommand do.
+303
View File
@@ -0,0 +1,303 @@
# Terminal smart copy (Ctrl+C) plan
Issue: [#211](https://github.com/Ark0N/Codeman/issues/211) "Terminal: Ctrl+C should copy when text is selected (interrupt otherwise)".
Origin: r/selfhosted feedback, "Biggest stumbling block is apparent lack of copy-paste in the terminal."
Status: **implemented and shipped** on 2026-08-05 (this document is kept as the rationale record). It was first served as an isolated beta over Tailscale for manual sign-off, then landed. Section 2 is the research that shaped the design, sections 4 to 6 describe what was built.
---
## 1. What the issue asks for
- Text selected in the terminal + `Ctrl+C` -> copy the selection, toast, clear the selection, do NOT send the byte to the PTY.
- No selection + `Ctrl+C` -> unchanged, the interrupt (`0x03`) reaches the PTY.
- `Ctrl+Shift+C` as an explicit copy chord.
- The selection check must run before the shortcut registry dispatch so a rebind cannot cost the user their interrupt key.
- Paste is out of scope (it already works via `Ctrl+V`, which terminal-ui.js routes to the image/text paste trap).
## 2. Verified current behavior
### 2.1 xterm cancels the Ctrl+C keydown, so no copy can happen
`src/web/public/vendor/xterm.min.js` (xterm 6.x), `_keyDown`:
```js
_keyDown(x){ if(this._keyDownHandled=!1, this._keyDownSeen=!0,
this._customKeyEventHandler && this._customKeyEventHandler(x)===!1) return !1;
... evaluateKeyboardEvent(...) ... this.cancel(x) ... }
```
Two consequences that shape the design:
1. The custom handler runs **first**, before xterm evaluates the key. Returning `false` exits before `cancel(x)`, so returning `false` does **not** call `preventDefault()` for us.
2. When the handler returns `true`, xterm turns Ctrl+C into `0x03` and cancels the event, which is why the browser's own copy command never runs.
Probe (headless chromium against an isolated server on port 3174, selection active, real focus on `.xterm-helper-textarea`, synthetic Ctrl+C keydown):
```json
{ "hasSelection": true, "defaultPrevented": true, "dataSeen": ["\"\\u0003\""],
"clipboardAfter": "SENTINEL-BEFORE", "stillHasSelection": false }
```
So today: interrupt byte sent, clipboard untouched, and xterm drops the selection anyway. The last point matters, "copy then clear the selection" is not a behavior change in how the selection feels, it is what already happens on any keypress.
### 2.2 Why right-click Copy works today
xterm registers a `copy` listener on its root element that substitutes the selection text:
```js
this._register(addDisposableListener(this.element,"copy",(k=>{ this.hasSelection() && copyHandler(k,this._selectionService) })))
```
Second probe (port 3175, real `page.keyboard.press('Control+c')`, custom handler patched to return `false` for Ctrl+C without `preventDefault`):
```json
{ "dataSeen": [], "copyEvents": ["xterm-element"],
"clipboardAfter": "native-copy-probe-line\n...", "stillHasSelection": true }
```
So a "return false and let the browser copy" implementation would also work in Chromium. It is rejected below (section 3.3) because it gives no toast, does not clear the selection, and leans on per-browser behavior of the copy command when the focused element is xterm's empty helper textarea.
### 2.3 The document-level capture handler will not interfere
`setupEventListeners()` in `src/web/public/app.js:989` runs on document capture, before xterm's textarea listener. Its registry loop skips any entry whose action is not in the local `SHORTCUT_ACTIONS` map:
```js
if (shortcut.disabled || !shortcut.action) continue;
const action = SHORTCUT_ACTIONS[shortcut.action];
if (!action) continue;
```
This is exactly how `command-palette` already behaves: it is a full registry entry (rebindable and disableable in App Settings) whose dispatch happens in a dedicated, focus-aware gate rather than the generic loop. The new copy entry follows that pattern, so the capture handler falls through untouched and the terminal handler owns the decision.
### 2.4 Registry matching rules that constrain the bindings
`matchesShortcutEvent()` (`app.js:4890`):
- Ctrl and Cmd are interchangeable as the primary modifier, so a `['ctrl']` binding also matches Cmd+C on macOS. That is fine here: with a selection it copies (same result the native macOS path gives today), without one it falls through.
- Every other modifier must be declared exactly: `if (mods.includes('shift') !== !!e.shiftKey) return false`. So `Ctrl+Shift+C` needs its own binding, a plain `ctrl+c` binding will never swallow it.
- `binding.code` wins when present, otherwise `binding.key` is compared case-insensitively.
### 2.5 Where selection is actually possible
- The server strips mouse-tracking DECSETs for `claude`, `codex`, and `gemini` (`isAltScreenStripMode`, `src/session.ts:179`), which is why plain drag-select works in those tabs even though the TUI has mouse tracking on.
- `shell`, `opencode`, and `antigravity` keep mouse reporting, so xterm requires `Shift`+drag to force a selection there. Worth one line in the docs, it is not a code change.
- Touch devices deliberately disable selection entirely (`body.touch-device .terminal-container .xterm{user-select:none !important}`, `styles.css:3196`), and phones have no Ctrl key. This feature is desktop and hardware-keyboard only, with no mobile regression surface.
### 2.6 Helpers that already exist and should be reused
| Need | Existing code |
| --- | --- |
| Clipboard write with an HTTP-safe fallback | `_copyText(text)` in `app.js:1887` (Clipboard API, then hidden textarea + `execCommand`) |
| Toast | `showToast(message, type)` in `panels-ui.js:4385` |
| Translated string | `'Copied to clipboard'` already in `i18n.js:453` |
| Focus-aware chord gate to copy the shape of | `shouldOpenCommandPaletteFromShortcut(e)` in `panels-ui.js:285` |
| Buffer-wide copy (currently unreferenced) | `copyTerminal()` in `terminal-ui.js:2615` |
`_copyText` matters more than it looks: `install.sh`'s LAN option serves plain HTTP, where `navigator.clipboard` is undefined. The issue's suggested `navigator.clipboard.writeText` alone would silently do nothing for those users, the `execCommand` fallback covers them.
## 3. Design
### 3.1 Behavior
| Chord | Selection present | No selection |
| --- | --- | --- |
| `Ctrl+C` (and Cmd+C, per registry equivalence) | copy, toast, clear selection, swallow the key | fall through, xterm sends `0x03` (interrupt) |
| `Ctrl+Shift+C` | copy, toast, clear selection, swallow the key | swallow, no-op (see 3.2) |
| Shortcut disabled in App Settings | never copies, `Ctrl+C` is always the interrupt | unchanged |
| Rebound to another chord | that chord copies when a selection exists | plain `Ctrl+C` is always the interrupt |
### 3.2 Why `Ctrl+Shift+C` with no selection is swallowed rather than forwarded
Today `Ctrl+Shift+C` produces `0x03` as well (the shift is irrelevant to the control byte), so forwarding would be "no regression". But once the chord is advertised as *the explicit copy key*, letting it interrupt a running agent when the selection happens to be empty is a footgun with no upside. Swallowing costs nothing: a user who wants to interrupt has `Ctrl+C` right there.
The rule in code is "no selection and the matched chord had Shift -> swallow", not a hardcoded key check, so it stays correct under rebinds.
### 3.3 Why an explicit clipboard write rather than falling through to the native copy
Probe 2 showed the native path works in Chromium, but the explicit write is chosen because it:
- gives the "Copied to clipboard" toast, which is the discoverability half of the issue,
- clears the selection so a second `Ctrl+C` interrupts (the smart-copy contract),
- works on plain-HTTP LAN installs through `_copyText`'s `execCommand` fallback,
- does not depend on how each browser treats a copy command issued while an empty textarea has focus.
### 3.4 Why no new app setting
Per-shortcut enable/disable and rebinding already exist in App Settings -> Shortcuts and are driven by the registry. A user who wants "Ctrl+C is always interrupt" unchecks one box. Adding a `terminalSmartCopy` setting would duplicate that and would drag in the per-device vs synced decision (`displayKeys` + `.strict()` `SettingsUpdateSchema`) for no gain.
## 4. Code changes, file by file
### 4.1 `src/web/public/app.js`, registry entry
Add to `DEFAULT_SHORTCUTS` (after the `clear-terminal` entry, ~line 351) so the Terminal group stays together:
```js
{
id: 'copy-selection',
group: 'Terminal',
label: 'Copy Selection',
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js,
// deliberately NOT in SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults, which would cost the user the interrupt key.
action: 'copyTerminalSelection',
},
```
Match on `key`, not `code`. xterm decides what byte to emit from the produced character, so intercepting the physical `KeyC` on a layout where it does not produce "c" would diverge from what xterm would have sent.
The `action` string is required for App Settings to render the row as configurable (`configurable = !!shortcut.action && Array.isArray(shortcut.bindings)`, `settings-ui.js:2624`). Do **not** add `copyTerminalSelection` to `SHORTCUT_ACTIONS`.
### 4.2 `src/web/public/terminal-ui.js`, the gate
New prototype method, modeled on `shouldOpenCommandPaletteFromShortcut`:
```js
shouldCopyTerminalSelectionFromShortcut(ev) {
if (!ev || ev.type !== 'keydown') return false; // the handler also runs for keypress/keyup
if (!ev.ctrlKey && !ev.metaKey && !ev.altKey) return false; // hot path: plain typing exits here
const registryAvailable =
typeof this.getShortcutRegistry === 'function' && typeof this.matchesShortcutEvent === 'function';
const entry = registryAvailable
? this.getShortcutRegistry().find((s) => s.id === 'copy-selection')
: null;
if (entry) return !entry.disabled && this.matchesShortcutEvent(ev, entry);
return (ev.key || '').toLowerCase() === 'c' && !ev.altKey; // fallback for isolated harnesses
}
```
### 4.3 `src/web/public/terminal-ui.js`, the branch
Inside `attachCustomKeyEventHandler` (`terminal-ui.js:133`), after the command-palette gate and before the `Ctrl+V` branch:
```js
// Smart copy (#211): with a selection, Ctrl+C copies instead of sending ^C.
// With no selection it MUST fall through (return true, no preventDefault) or
// the interrupt key is lost. Ctrl+Shift+C is the explicit chord and never
// falls through: an "explicit copy" that interrupts the agent is a footgun.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
}
return true;
}
```
`preventDefault()` is explicit because returning `false` alone does not cancel the event (section 2.1), and without it the browser would run its own copy on top of ours.
### 4.4 `src/web/public/terminal-ui.js`, the copy action
```js
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const ok = await this._copyText(selection);
if (ok) {
this.terminal.clearSelection?.();
this.showToast('Copied to clipboard', 'success');
} else {
this.showToast('Failed to copy', 'error');
}
// _copyText's execCommand fallback focuses a temp textarea; restore the
// terminal (this.terminal.focus is the CJK-aware router, not xterm's raw focus).
this.terminal.focus();
return ok;
}
```
The selection text is captured **before** the first `await`, and `navigator.clipboard.writeText` is reached in the same task as the keydown, so user activation still holds.
### 4.5 `src/web/public/i18n.js`
`'Copied to clipboard'` exists. Add `'Failed to copy': '复制失败'` (the error path is new to this surface).
### 4.6 Documentation
| File | Change |
| --- | --- |
| `README.md` shortcut table (~line 648) | `\| `Ctrl/Cmd+C` \| Copy selection (interrupts when nothing is selected) \|` and a `Ctrl+Shift+C` row |
| `src/web/public/index.html` help modal, Terminal section (~line 641) | `<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection / Interrupt</div>` plus the Ctrl+Shift+C row. Keep the existing negative assertion in `help-modal-shortcuts.test.ts` in mind (it forbids `Ctrl+K`, `C` is fine) |
| `CLAUDE.md` "Keyboard shortcuts" line | add `Ctrl+C` (copy selection, else interrupt) and `Ctrl+Shift+C` |
| `docs/architecture-invariants.md` -> "Command palette and shortcut registry" | append the invariant: the no-selection path must return `true` without `preventDefault`, the branch is keydown-only, and `copyTerminalSelection` must stay out of `SHORTCUT_ACTIONS` |
The shortcut overlay (`Ctrl+?`) and App Settings -> Shortcuts are registry-driven and pick the entry up with no edit.
## 5. Edge cases and risks
| Case | Handling |
| --- | --- |
| Handler also fires for `keypress`/`keyup` | gated on `ev.type === 'keydown'`. xterm's `_keyPress` bails on ctrl combos anyway, so no stray byte |
| CJK IME composing | the existing `isComposing || keyCode === 229` guard is the first line of the handler and stays first |
| Local echo overlay has unsent `pendingText` | the copy branch returns before `onData`, so `pendingText`, flushed offsets and the durable input queue are untouched. The no-selection path is byte-identical to today, including the "control char flushes buffered text then sends `0x03`" logic at `terminal-ui.js:895` |
| Plain HTTP (LAN install) | `_copyText` falls back to `execCommand`, then focus is restored |
| Clipboard write rejected (permissions policy, no gesture) | error toast, right-click Copy still available |
| Whitespace-only or empty selection | `getSelection()` empty string is treated as "no selection", so Ctrl+C still interrupts |
| macOS Cmd+C | registry treats ctrl/meta as interchangeable, so with a selection it takes our path (same visible result as today's native copy), without one it falls through |
| Chrome/Firefox `Ctrl+Shift+C` is the devtools inspect chord | browser-level and may still toggle devtools, our copy runs regardless. Document as a caveat, `Ctrl+C` is the primary path |
| Selection in a tab whose TUI owns the mouse (`shell`/`opencode`/`antigravity`) | unchanged, `Shift`+drag selects, then Ctrl+C copies |
| Web tab (iframe dashboard) focused | xterm handler never runs, browser-native copy inside the iframe |
| Teammate/subagent terminals (`panels-ui.js:2268`, `onData` wired) | same limitation exists there, out of scope for this PR (section 8) |
## 6. Test plan
New file `test/terminal-copy-selection.test.ts` (node env, `vm` harness in the style of `test/command-palette-ui.test.ts`), covering `shouldCopyTerminalSelectionFromShortcut` in isolation:
1. Ctrl+C keydown -> true, keyup/keypress of the same chord -> false.
2. Ctrl+Shift+C -> true, plain `c` -> false, Ctrl+K -> false.
3. Registry entry `disabled: true` -> false for every chord.
4. Rebound entry (for example Alt+Y) -> true for the rebind, false for Ctrl+C.
5. Missing registry (harness without `getShortcutRegistry`) -> falls back to the `c` check.
Static assertions appended to `test/keyboard-shortcuts.test.ts` (this suite already pins the xterm-handler chokepoint):
6. `DEFAULT_SHORTCUTS` contains `id: 'copy-selection'` and `SHORTCUT_ACTIONS` does **not** contain `copyTerminalSelection` (the interrupt-safety invariant).
7. `terminal-ui.js` contains the `shouldCopyTerminalSelectionFromShortcut` branch and a `return true` no-selection fall-through.
8. README + help modal rows exist (mirrors the existing palette/Alt-nav doc assertions).
`test/help-modal-shortcuts.test.ts`: add `expectShortcut(helpModal, ['Ctrl', 'C'], 'Copy Selection')`.
New browser test `test/terminal-copy-shortcut.test.ts` (Playwright, port **3174**, free per a scan of `test/`), following `test/webgl-fallback.test.ts`: boot `WebServer`, grant `clipboard-read`/`clipboard-write`, `terminal.write()` a known line, `selectLines()`, real `page.keyboard.press('Control+c')`, then assert clipboard content, empty `onData` capture, cleared selection and the toast. Second case: no selection, assert `onData` saw `\u0003` and the clipboard is unchanged.
Per repo convention, browser suites are excluded from CI, so add the filename to the exclude list in `config/vitest.ci.config.ts` and run it locally.
Regression runs: `npm test -- test/keyboard-shortcuts.test.ts`, `test/help-modal-shortcuts.test.ts`, `test/command-palette-ui.test.ts`, `test/input-send-order.test.ts`, then `npm run test:ci`.
## 7. Manual verification before COM (CLAUDE.md rule)
Against a throwaway session on the live instance (`curl -sk https://localhost:3000/...`, never w1/w2/w3):
1. Select output with the mouse, press Ctrl+C, confirm the toast, paste elsewhere, confirm the agent did not stop.
2. Press Ctrl+C again with nothing selected, confirm the agent interrupts.
3. Type a few characters with local echo on (phone or `localEchoEnabled` forced), press Ctrl+C with no selection, confirm buffered text plus interrupt behave as before.
4. Uncheck the shortcut in App Settings -> Shortcuts, confirm Ctrl+C always interrupts even with a selection.
5. Rebind it, confirm the new chord copies and Ctrl+C reverts to pure interrupt.
6. Repeat 1 and 2 in an `opencode` or `shell` tab using Shift+drag to select.
7. Load over plain HTTP (`--host` LAN or `http://127.0.0.1:<port>`) and confirm the `execCommand` fallback copies and focus returns to the terminal.
8. Mobile smoke: confirm nothing changed (selection is CSS-disabled, no Ctrl key).
## 8. Out of scope, follow-ups worth filing separately
- **Teammate/subagent terminals** (`panels-ui.js:2268`) have the same blocked-copy problem. One `attachCustomKeyEventHandler` reusing `copyTerminalSelection` would fix them, but it touches a different surface and deserves its own change.
- **A mobile copy affordance.** Selection is disabled on touch, so phones still cannot copy terminal text. The unreferenced `copyTerminal()` (whole buffer) plus a keyboard-accessory "Copy" button would be the cheapest answer.
- **Right-click context menu** with Copy/Paste, better discoverability than any chord, but a bigger UI surface.
- **`copyTerminal()` cleanup**: it uses raw `navigator.clipboard` rather than `_copyText`, so it would fail on plain HTTP if ever wired up.
## 9. PR mechanics
- Branch off `master` (verify with `git branch --show-current`, the tree is shared), stage explicit paths only.
- Files touched: `src/web/public/app.js`, `src/web/public/terminal-ui.js`, `src/web/public/i18n.js`, `src/web/public/index.html`, `README.md`, `CLAUDE.md`, `docs/architecture-invariants.md`, `docs/terminal-copy-shortcut-plan.md`, three test files, `config/vitest.ci.config.ts`.
- `index.html`, `app.js` and `terminal-ui.js` are `.prettierignore`d hand-formatted assets, match the surrounding style by hand. `npm run check:public-assets` and `npm run check:frontend-syntax` are the guards.
- No changeset in this PR: a merged, unconsumed changeset turns the Release workflow red until the next COM, and the COM flow writes release notes covering everything since the last tag (current version is 1.10.0).
- Close #211 from the PR body.
Rough size: about 60 lines of product code, most of the work is the tests and the four documentation surfaces.
+201
View File
@@ -0,0 +1,201 @@
<!-- Design doc drafted 2026-07-28 from WWDC26 session 224 research. STATUS: PLANNED, NOT IMPLEMENTED. Blocked on macOS 27 "Golden Gate" (beta now, GA expected fall 2026). -->
# VM Cases (macOS Virtualization framework), Implementation Plan
## Status
PLANNED, nothing implemented. This is the design + phased execution plan for a native-macOS VM isolation tier for cases ("the VM subsystem"), modeled on Docker cases (`docs/docker-cases-plan.md`). Testbed prerequisite: a macOS 27 host (see Section 8).
**⚠ DESIGN DIRECTION (owner, 2026-07-29): the subsystem is GUI-first.** Users want real macOS desktops, not headless SSH machines. Guests may be macOS (GUI-only in practice) or Linux (GUI or headless). Key decision 3 below carries the full consequences; anything in this doc that reads as "Linux-first / headless-first" predates this and has been revised.
**2026-07-29: Phase 0 substantially validated on the beta testbed; full Apple-stack reference now lives in [`docs/vm-subsystem-apple-stack.md`](vm-subsystem-apple-stack.md)** (API surfaces, beta bugs, our empirical results, and design implications). Plan-relevant corrections from that work: vmnet's topology/port-forwarding APIs are macOS 26 (only the loopback fix is 27); guest provisioning is macOS-guests-only (Linux stays cloud-init, proven working); DiskImageKit has NO flatten/merge, so the `export` subcommand ships the layer chain (or flattens in-guest) instead of flattening; seed ISOs are base-build-time only, never attached at case runtime; per-case EFI variable stores are mandatory; guest health checks read DHCP leases, never serial/ping.
## 1. Context and motivation
WWDC 2026 session 224 ("Expand the Capabilities of your Virtualization App", https://developer.apple.com/videos/play/wwdc2026/224/) shipped the missing pieces for programmatic, fleet-style VM management on macOS:
- **`VZMacGuestProvisioningOptions`**: automated first-boot setup of a macOS guest (user account, auto-login, SSH enabled) with zero interactive setup.
- **DiskImageKit**: stacked disk images on the Apple Sparse Image Format (ASIF): a read-only base layer plus cheap per-VM cache/overlay layers. Direct analog of Docker image layers + writable container layer.
- **vmnet framework**: custom network topologies and port forwarding from the host process.
- **`VZCustomVirtioDevice`**: custom low-latency host<->guest channels (Linux guests).
- **AccessoryAccess**: USB passthrough (not relevant to Codeman, out of scope).
Codeman's isolation story today is Docker cases. On macOS, Docker means Docker Desktop / a Linux VM anyway, with weaker fidelity and a heavyweight dependency. The Virtualization framework gives hardware-virtualized per-case sandboxes natively, with a layered-image story that mirrors what `scripts/build-agent-image.mjs` does for Docker. This is the premium native-macOS tier ON TOP of Docker cases, never a replacement (Docker remains the cross-platform story; the Linux prod box cannot use any of this).
## 2. Platform reality (hard constraints)
| Constraint | Detail |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Host OS | macOS 27 "Golden Gate" required for the new APIs (dev beta since 2026-06-08, public beta since 2026-07-13, GA expected fall 2026) |
| Host hardware | Apple Silicon only (macOS 27 dropped Intel). Testbed: the owner's dedicated MacBook (Section 8); the M4 Mac mini (macOS 26.4, runs the second Codeman install) stays on stable + untouched |
| Guest provisioning | `VZMacGuestProvisioningOptions` needs macOS 27 on BOTH host and guest. Linux guests provision via cloud-init instead |
| macOS guest concurrency | **Hard kernel cap: 2 concurrent macOS VMs per host. MEASURED on 27 beta 4 (2026-07-29), not inferred**: the 3rd VM is refused instantly with `VZErrorDomain` code 6 while 39% of RAM is free, so more hardware does NOT raise it. Since macOS GUI guests are the headline use case, this is a real product capacity limit to schedule around and surface in the UI. Linux guests are uncapped (resource-bound only) |
| Language | Virtualization framework is Swift/ObjC only; Node cannot call it. Requires a Swift helper binary (Key decision 2) |
| Entitlement | Host process needs `com.apple.security.virtualization`. Fine for a locally built dev binary; distribution needs signing thought (Section 9) |
| Nested virtualization | Linux-guest-only on M3+. A macOS 27 VM cannot dependably host its own guests, so the host-side APIs must be tested on bare-metal 27 (dual-boot) |
| CI | Cannot run in CI (needs beta macOS on Apple Silicon). Same answer as tmux/docker: no-op all VM IO under `VITEST`, unit-test the pure parts |
## 3. Goal and user stories
Add "VM cases" to Codeman: a case can point at a per-case virtual machine on a macOS host, and any CLI backend runs inside it over the existing remote-SSH session machinery. A LOCATION OVERLAY on cases, exactly like remote-SSH and Docker cases, NEVER a sixth `SessionMode`.
- As a Mac user, I link a case to a VM so an autonomous run executes behind a hardware virtualization boundary (stronger than Docker's shared kernel) while file viewing, transcripts, and hooks keep working.
- Per-case VMs are instant and cheap: a shared provisioned base image plus a per-case overlay, not a full image copy per case.
- Killing a session kills only its in-guest tmux; the VM stays up while sibling sessions remain; case delete tears the VM down.
- I export a case's VM overlay as a portable artifact (mirror of `docker-exports/`), secrets excluded.
- On a non-mac host, or a Mac without the helper, the feature is invisible: zero UI, zero probes, zero errors.
Non-goals for the MVP: USB passthrough, custom Virtio channels (Phase 3 candidate), macOS-guest fleets (capped at 2 anyway), Kubernetes-style orchestration, Intel Macs.
## 4. Architecture
```
Codeman (Node, unchanged session layer)
| JSON over stdout (same pattern as shelling out to docker/tmux)
v
codeman-vm (Swift package: CLI + per-VM GUI runner app in the console session)
| Virtualization / DiskImageKit / vmnet
v
per-case VM (macOS or Linux)
|-- GUI mode: VZVirtualMachineView in a window --> guest screen sharing --> browser (noVNC)
|-- shell: SSH on vmnet IP --> existing remote-SSH tmux machinery
^ VirtioFS: host case dir mounted at the SAME absolute path
```
Note the runner is a **GUI app in the console user's session**, not a detached daemon: a daemon-launched VM cannot render, which is fatal for macOS guests and for Linux desktop cases.
### Key decision 1: location overlay, not a mode
Identical reasoning to Docker/remote-SSH (see CLAUDE.md): the session layer, respawn, Ralph, recovery, and quick-start plumbing all stay untouched. `SessionMode` stays five-valued. State mirrors the Docker pair: `~/.codeman/vm-hosts.json` + `vm-cases.json`, new `src/vm-hosts.ts` with the storage + pure helpers split.
### Key decision 2: Swift helper CLI (`codeman-vm`)
The framework is Swift-only, so all VM work lives in a SwiftPM package (`packages/codeman-vm/`), a CLI with a stable JSON contract:
- `create-base --guest linux|macos`: build the shared base image. Linux: boot an arm64 cloud image with EFI + cloud-init, install Node 22 + tmux + the four CLIs (same inventory as `docker/agent.Dockerfile`), seal as base ASIF. macOS: IPSW restore + `VZMacGuestProvisioningOptions` (agent user, SSH on), then **desktop-readiness baking**, which is mandatory for GUI guests: suppress the per-user first-login assistant (`com.apple.SetupAssistant` keys + the User Template), enable auto-login (`autoLoginUser` + `/etc/kcpassword`), disable screensaver/lock/display-sleep, and set a static wallpaper (animated "aerials" wallpaper is unusable over remote display). ⚠ Use RAW (not ASIF) for macOS guest disks until the beta's macOS-guest space-reclamation bug is fixed.
- `create <case>`: DiskImageKit stacked image: shared read-only base + fresh per-case overlay. Near-instant, space-efficient.
- `start <case>` / `stop` / `status` / `ip`: lifecycle + vmnet NAT; `ip` reports the guest SSH endpoint.
- `export <case>` / `import`: flatten overlay + workspace tar + manifest, credentials excluded (mirror of docker-export).
A VM dies with its owning process, so `start` spawns a DETACHED per-VM runner process (analog of the detached `scripts/self-update.sh` trick) rather than a monolithic daemon; `status` talks to it over a unix socket in the instance data dir (`dataPath()`, never a hardcoded `~/.codeman` path).
### Key decision 3: multi-guest, and GUI is a first-class mode (REVISED 2026-07-29 by the repo owner)
The subsystem supports both macOS and Linux guests, and a guest runs in one of two **display modes**:
| | macOS guest | Linux guest |
| --- | --- | --- |
| **GUI mode** | **the point of the feature**; a real macOS desktop. Mandatory: nothing renders without an attached `VZVirtualMachineView` in an unlocked host session | supported (EFI + virtio-gpu framebuffer) for desktop Linux cases |
| **Headless mode** | not offered: a macOS guest with no view renders nothing, so a "headless macOS desktop" is a contradiction. SSH-only macOS is possible but is not what this feature is for | supported and cheap; the natural mode for agent/CI work, driven over SSH |
Consequences that flow from GUI being first-class:
- VM processes are **GUI apps in the console user's session** (LaunchAgent / `launchctl asuser`), never daemons. A daemon-launched VM cannot render.
- **The host is part of the product surface**: it must auto-login, never lock, never sleep, and keep a live WindowServer. Host lock == every VM's screen goes black, so the screen lock is effectively a global kill switch for every VM display on the machine. The product must own these host settings rather than treat them as user preference.
- **FileVault conflicts with unattended GUI hosting** and the trade-off must be a deliberate choice: FileVault disables auto-login, so a full-disk-encrypted host needs a human at a keyboard (or a remote screen-sharing session) after every reboot before any VM can render. Options are (a) FileVault on, accept manual login per boot, (b) FileVault off on a dedicated VM host so it boots straight into a rendering session, or (c) FileVault on plus a remote-unlock runbook. Codeman should detect the state and tell the user which one they are in instead of silently serving black screens.
- **Guests must be desktop-ready, not just booted**: auto-login, no screensaver/lock, and the per-user first-login assistant pre-suppressed at base-image time (`com.apple.SetupAssistant` keys, plus the User Template so later accounts inherit it). Otherwise the user connects to a login prompt or a setup wizard, which is exactly what happened during the first hands-on run.
- **Capacity is capped for macOS**: at most 2 concurrent macOS VMs per host, confirmed by our own test on 27 beta 4 (3rd refused with `VZErrorDomain` 6 at 39% free RAM; it is a kernel quota, so bigger hardware does not help). Scheduling must queue or evict beyond 2, the UI must explain why, and the scheduler should tolerate the acknowledged slot-leak bug (a slot occupied with nothing running, host-reboot to clear). Linux guests are uncapped and bounded only by host resources, which is the lever for scaling case counts on one machine.
- **Access is via the guest's own screen**, viewable in a browser through the noVNC chain (see `docs/vm-subsystem-apple-stack.md` §8), so no client-version or client-install requirements land on the user.
Provisioning per guest type: `VZMacGuestProvisioningOptions` for macOS (needs 27-on-27, first-boot-only, and does NOT skip the per-user wizard), cloud-init NoCloud seed ISO for Linux (proven working).
### Key decision 3b: the GUI VM host profile, and supervision that catches black screens
GUI hosting only works if the host is configured for it and supervised. This profile was derived the hard way on the testbed (prototyped there 2026-07-30) and should be what `codeman-vm` installs and verifies:
**Host profile** (the product should own these, not leave them to preference):
1. **No login barrier.** Either FileVault off + auto-login (a dedicated VM host boots straight into a rendering session, fully unattended), or FileVault on and remote reboots done with `sudo fdesetup authrestart`, where the pre-boot unlock *is* the login so the machine returns already logged in with encryption intact. **`authrestart` is VERIFIED on the testbed (2026-07-30): the host rebooted remotely and came back with a live logged-in console session, FileVault still enabled, no password prompt** — this is the recommended pattern for an encrypted GUI VM host. Plain reboots on a FileVault host always need a human, so Codeman should detect that combination and warn instead of serving black screens.
2. **Never lock**: lock policy off (needs the account password, so it is a setup step, not a scriptable one) plus `caffeinate -d -i -m -u` re-armed per session.
3. **Never sleep**: `pmset -a sleep 0 displaysleep 0 disablesleep 1`; a physical display is NOT required (a lid-closed laptop renders fine, only an unlocked session matters). Note OS updates reset these.
4. **Session-independent control plane**: run VPN/remote access as a system service, never a session app, and keep the access chain (forwards, VNC proxies, web endpoints) in LaunchDaemons so a session restart cannot sever operator access.
**Supervision** must be a **root LaunchDaemon**, not a user LaunchAgent. This is the load-bearing detail: a user agent cannot launch a GUI app into the Aqua session, so its restart attempts fail *silently* (the child dies instantly, leaving an empty log while the supervisor cheerfully reports success). A root daemon can, via `launchctl asuser <uid> sudo -u <user> …`, and those launches persist. Prototyped and verified on the testbed 2026-07-30; a working supervisor runs on a short interval and:
- Restarts the runner when the process is gone **or when its log shows `WindowServer event port death`**, which means it is permanently blind while still looking alive.
- Defers restarts while the console is at the login window, and launches into whichever session actually exists (resolve the console user with `stat -f %Su /dev/console`, never a hardcoded one).
- Re-points the guest port-forward whenever the guest's NAT lease changes, which happens on **every guest boot** under plain NAT. A vmnet DHCP reservation for a stable per-case IP is the better long-term answer.
- **Re-applies host power settings**, because `pmset -a disablesleep 1` does NOT survive a reboot (caught on the supervisor's first run after a real reboot) and OS updates reset it too.
- Re-arms the keep-awake helper, which dies with its session.
- Ideally also samples the guest framebuffer for non-black content, since a black screen is the one symptom common to every failure mode here.
`pgrep` alone is worthless for health: every failure mode in this session presented as a healthy process.
### Key decision 4: sessions ride the existing remote-SSH machinery
A provisioned guest is literally an SSH host on a vmnet IP. Session launch = the remote-SSH flow with the host swapped in: durable remote `tmux -L codeman-remote`, session names failing `SAFE_MUX_NAME_PATTERN` on purpose, EVERY ssh command line through `buildSshConnectionArgs()` (command-injection invariant), run flows through `POST /api/quick-start` (never `POST /api/sessions`, which stat-validates `workingDir` locally). What is genuinely new is only lifecycle (create/start/stop/export) and the vm-hosts/vm-cases overlay state.
### Key decision 5: workspace via VirtioFS at the same absolute path
Mirror the Docker bind-mount invariant: the case workspace is a real host directory shared into the guest via VirtioFS and mounted at the SAME absolute path. That keeps file-routes/watchers on real host bytes and makes the in-guest transcript projHash match the host. Without this, transcripts/attachments/file viewer all silently degrade.
### Key decision 6: credentials seeded, hooks bridged
- Credentials are SEEDED (read-only share, copied into the guest once at create), never shared read-write, and excluded from exports: byte-for-byte the Docker cases rule and rationale.
- Hooks: on the loopback-only prod bind a guest cannot reach `127.0.0.1:3000`. Mirror `CODEMAN_DOCKER_BRIDGE_HOOKS` with a `CODEMAN_VM_BRIDGE_HOOKS` opt-in listener on the vmnet gateway IP; otherwise idle detection falls back to output-based, same as Docker.
### Key decision 7: drift and teardown copy Docker semantics verbatim
Config hash label on the VM (guest type, cpu/mem, share list); a drifted launch is REFUSED, never silently launched stale. One VM per case shared by all sessions; session kill = in-guest tmux kill only; case delete = stop + remove overlay; instance-scoped boot reaper for orphaned runner processes.
## 5. Implementation phases
**Phase 0, testbed (no repo code):** dedicated MacBook on the macOS 27 beta, remotely accessible over the tailnet (setup protocol in Section 8), Xcode 27 beta, then a throwaway Swift script proving the loop: create base -> overlay -> boot -> ssh in. This validates 80% of the design before any Codeman code.
**Phase 1, `codeman-vm` helper:** SwiftPM package, the six subcommands above, JSON contract doc, detached runner + unix-socket status, Linux base image build. Deliverable is testable entirely without Codeman.
**Phase 2, Codeman integration:** types (`VmHost`/`VmCase`/`SessionVm`), `src/vm-hosts.ts` (+ pure helpers: config hash, arg building, endpoint parsing), Zod schemas, `case-routes` link/unlink + listing, `quick-start` vm branch reusing the remote-SSH launch path, `Session` threading + recovery round-trip, `VITEST` no-op layer, unit tests. Feature-detect: darwin + arm64 + helper binary present, else invisible.
**Phase 3, polish:** export/import UI, frontend Create Case "VM" tab + case-picker labels, SSE `vm:*` events, macOS-guest opt-in with cap surfaced, custom-Virtio input channel exploration, CLAUDE.md Key Pattern + `docs/vm-cases.md` + COM.
## 6. Testing
- Pure helpers unit-tested (ports pattern from `docker-hosts.ts`: 26 tests there, aim similar).
- All helper-invoking IO no-ops under `VITEST` (the `IS_TEST_MODE` pattern in `tmux-manager.ts`).
- End-to-end verification happens ON the beta MacBook, per the always-end-to-end rule: real base build, real per-case overlay boot, real quick-start into the guest, workspace round-trip through VirtioFS, session-delete keeps VM up, case-delete removes it.
- CI never runs the real path; the static guards are type-level + unit-level only.
## 7. Risks
1. **Beta API churn**: everything here targets beta SDKs; symbol/behavior changes are likely before fall GA. Mitigation: Phase 0/1 are throwaway-tolerant; no Codeman-side commitment until the helper contract survives a beta cycle.
2. **New artifact class**: Codeman ships pure TypeScript today; a Swift binary changes build/distribution (build-on-install via `xcrun swift build` on macs with Xcode CLT? prebuilt signed binary per release?). Needs an owner decision; local dev build is fine for the whole beta period.
3. **Entitlement/signing**: `com.apple.security.virtualization` is trivial for local dev, real for distribution.
4. **Adoption gating**: users need macOS 27 + Apple Silicon for months after GA. Docker cases remain the default recommendation; VM cases ship dark (feature-detected) with zero cost to everyone else.
## 8. Beta testbed plan: dedicated MacBook (actionable now)
Testbed is a dedicated MacBook the owner sacrifices to the beta (after a full backup). This supersedes the earlier dual-boot-the-Mini idea (git history has it): a dedicated machine means no OS-switching, no downtime for the Mini's live Codeman, and no FileVault pre-boot headaches.
**Sequencing rule that makes it headless: configure ALL remote access on the CURRENT macOS first, THEN upgrade in place.** An in-place beta upgrade preserves Remote Login, Tailscale, user accounts, and auto-login, so there is no Setup Assistant and no post-install physical step. (A fresh install would boot into GUI-only Setup Assistant with no SSH, which on a headless box is a dead end.)
Confirmed hardware (2026-07-28): MacBook, M3, 16 GB RAM, 256 GB disk with ~100 GB free. Verdict: green. M3 = eligible + nested-virt capable; 16 GB = host + 2-3 concurrent Linux guests (macOS guest = one at a time); 100 GB = fits with discipline: install Xcode 27 beta with the macOS platform only (skipping iOS/watchOS/tvOS simulators saves 15-20 GB), and defer any macOS guest base (~30 GB) to an external SSD or until actually needed. Linux guests + sparse ASIF overlays are the comfortable path.
### Pre-upgrade checklist (owner, physical, once)
1. Full backup (Time Machine or clone); the machine should be considered beta-only afterwards.
2. Tailscale: install, sign into the tailnet, confirm it appears in `tailscale status` from another node.
3. System Settings -> General -> Sharing: **Remote Login ON** (SSH) and **Screen Sharing ON** (for the rare GUI-only moments: Xcode license, Apple Account dialogs).
4. **FileVault stays ON** (owner decision 2026-07-28, security over convenience). Consequences: auto-login is unavailable, but FileVault's pre-boot unlock doubles as login, so an unlocked boot still lands in a live GUI session; planned remote reboots go through `sudo fdesetup authrestart` (unlocks for exactly one restart); an UNPLANNED reboot (beta kernel panic, battery drain) parks the machine at the pre-boot screen, no SSH/Tailscale, until the password is typed physically. If the testbed goes silent, suspect this first. Keep it on AC so the battery absorbs power blips.
5. Beta enrollment (manual): sign into the Apple Account in System Settings; System Settings -> General -> Software Update -> **Beta Updates** -> select the **macOS 27 Developer Beta** (preferred: framework fixes land weeks earlier than public beta; free since 2023 after accepting the agreement once at developer.apple.com; public-beta alternative: enroll at beta.apple.com). Then run the offered upgrade: plugged in, lid open, trusted network.
6. Send over: tailnet name/IP, username, and a first-login password (key install + lockdown happens remotely right after).
### Post-upgrade setup (remote, over the tailnet)
1. Verify: `sw_vers` reports 27.x, SSH reachable.
2. Server-ize the laptop: `sudo pmset -a sleep 0 disksleep 0 disablesleep 1` (lid-closed operation without an external display), `womp 1` (wake on network), `sudo systemsetup -setrestartpowerfailure on`. Keep on AC power.
3. Install the controlling host's SSH key, then disable password auth.
4. Xcode 27 beta install (the one step needing the owner's Apple Account sign-in once, doable via Screen Sharing from anywhere); `xcode-select`, license accept, verify `swift --version` + the 27 SDK (`xcrun --show-sdk-version`).
5. Phase 0 prototype loop, all remote from here: Linux guest base image (no 27-on-27 provisioning dependency), DiskImageKit overlay, boot, vmnet NAT, ssh into the guest, run `claude --version` inside.
6. Only after that loop works: start Phase 1 in `packages/codeman-vm/`.
## 9. Open decisions (owner)
1. Linux base distro/image for the default guest (proposal: Ubuntu 24.04 arm64 cloud image, matching the docker agent image's userland).
2. Helper distribution for GA: build-on-install vs prebuilt signed binary vs "bring your own Xcode".
3. Ship dark behind `CODEMAN_VM_CASES=1` for the first release, or feature-detect only?
4. Export format parity with docker-exports (one manifest schema for both?).
## References
- Session 224: https://developer.apple.com/videos/play/wwdc2026/224/
- Fleet-angle writeup: https://bitrise.io/blog/post/wwdc26-the-virtualization-framework-updates-that-matter-for-large-mac-fleets
- Beta timeline: https://www.macworld.com/article/3189014/apple-july-2026-ios-ipados-macos-27-public-betas-tv-arcade-releases.html
- Internal analogs: `docs/docker-cases-plan.md` (architecture template), `docs/remote-sessions.md` (session transport), `docs/architecture-invariants.md#docker-cases`
+281
View File
@@ -0,0 +1,281 @@
<!-- Reference doc for the VM subsystem (Codeman VM cases). Compiled 2026-07-29 from: Apple DocC JSON backend, macOS 27 beta 4 SDK on the testbed, a multi-source web research sweep, and hands-on prototyping on a MacBook Air M3 running macOS 27.0 beta (26A5388g). Companion to vm-cases-plan.md (the Codeman integration plan). -->
# The VM Subsystem: Apple Virtualization Stack Reference (macOS 27 "Golden Gate")
"VM subsystem" is the working name for Codeman's native-macOS VM isolation tier and everything under it. This document is the single place for what the Apple stack actually provides, what we have verified ourselves on the beta, and what is known-broken. The Codeman-side design lives in `docs/vm-cases-plan.md`.
**Research method note:** Apple's HTML doc pages are JS-rendered and come back empty to fetchers. The working route is the DocC JSON backend: `https://developer.apple.com/tutorials/data/documentation/<path>.json` (page content) and `https://developer.apple.com/tutorials/data/index/<framework>` (full symbol tree with per-symbol `beta` flags). Everything below marked "Apple docs" was parsed from that backend directly.
## 1. Component map and minimum OS versions
| Component | What it is | Min host OS | Notes |
| --- | --- | --- | --- |
| Virtualization.framework core | VMs, EFI/Linux boot, virtio devices, VirtioFS | macOS 11-13 era | Unchanged basics; our prototype uses nothing newer than macOS 13 APIs except the DiskImageKit bridge |
| **DiskImageKit** | ASIF + raw disk images, layered stacks | **macOS 27** | Swift-only, no ObjC headers. Section 2 |
| **Guest provisioning** | First-boot account/SSH setup for macOS guests | **macOS 27 host AND guest** | Mac guests only as of beta 4. Section 3 |
| vmnet topology/port-forward/DHCP APIs | Custom networks, port forwarding | **macOS 26** (NOT 27) | 27 adds exactly one fix: loopback port forwarding. Section 4 |
| `VZVmnetNetworkDeviceAttachment` | In-process vmnet attach | macOS 26 | |
| **`VZCustomVirtioDevice`** family | Custom paravirt devices | **macOS 27** | Linux guests only, custom guest driver required. Section 5 |
| AccessoryAccess (USB passthrough) | USB claim + attach to VMs | macOS 27 | Requires paid-team provisioning profile, Dock app. Out of scope for Codeman. Section 6 |
Corrections to the WWDC-session framing we started with: vmnet's topology family is a macOS 26 story (129 symbols, zero beta-flagged in 27); provisioning does NOT currently extend beyond macOS guests despite the generic-looking `VZGuestProvisioningOptions` base class; DiskImageKit has no attach/mount API at all (it is a file-format library that hands `DiskImage` objects to Virtualization, no `/dev/diskN`, no root needed, no entitlement documented).
## 2. DiskImageKit (macOS 27, Swift-only)
Public framework, `/System/Library/Frameworks/DiskImageKit.framework`. No ObjC headers; the API surface lives in the `.swiftinterface`. Verified present in the CLT 27 beta 4 SDK, and our prototype compiled against it with plain `swiftc` on the first attempt.
### API surface (complete as of beta 4)
```swift
class DiskImage {
convenience init(creating: some DiskImage.CreationConfiguration) throws
convenience init(opening: some OpenConfigurationProtocol) throws
func appending(any DiskImage.CreationConfiguration & DiskImage.StackableLayer) throws -> any StackedImage
func appending(consuming DiskImage) throws -> any StackedImage // reattach an existing layer; validates parentUUID
func truncate(blockCount: Int) throws // stacked: affects top layer; does NOT resize guest fs
var blockCount, blockSize, format, layerType, layerUUID, parentUUID, openMode, size, url
}
protocol StackedImage: DiskImage { var layers: [DiskImage] }
struct OpenConfiguration { init(url:mode:); Mode = automatic | readOnly | readWrite }
// CreationConfiguration statics: .asif(url:blockCount:blockSize:), .asifLayer(url:type:), .raw(url:blockCount:)
// DiskImage.LayerType: .cache | .overlay | .overlay(blockCount:)
// DiskImage.BlockSize: .bytes512 | .bytes4096
// Errors: CorruptedImageError, IncompatibleStackingError(reason), InvalidBlockCountError, UnsupportedFormatError
```
Bridge into Virtualization is a new beta convenience init on the existing attachment class. Note there is no `readOnly:` parameter; read-only-ness comes from each layer's own `openMode`:
```swift
VZDiskImageStorageDeviceAttachment(diskImage: stack, cachingMode: .automatic, synchronizationMode: .full)
```
### Stacking rules (Apple docs, verbatim where quoted)
- ASIF works standalone or stacked. "You can only use RAW images as standalone images or as **base** images in stacked configurations." Upper layers are always ASIF.
- **One cache layer per stack**, any number of overlays conceptually, "shallow stacks perform better" (WWDC 224). No published max-depth guidance.
- "Layers are processed from bottom (base) to top. The **topmost layer determines the stack's size and receives all writes**." `.overlay(blockCount:)` therefore also grows the virtual disk.
- UUID chaining: appending sets the child's `parentUUID` to the parent's `layerUUID`. Raw bases have no UUID. "The layer UUID **changes if the layer is written to**", and reattaching a mismatched layer throws `IncompatibleStackingError`. This is the mechanism that makes a shared read-only base safe.
- Base sharing across multiple VMs is the stated design intent ("can be shared across multiple VMs"), with the WWDC caveat that per-VM auxiliary files (EFI variable store, macOS auxiliary storage) must be duplicated per VM, never shared.
- **There is no flatten/merge.** An overlay cannot be merged back into its base (confirmed by Howard Oakley's coverage plus an independent hands-on report). Export/move flows must ship the layer chain, or flatten inside a guest (dd to a fresh attached image).
### Known issues and adoption
- **ASIF space reclamation is broken for macOS guests on the beta** (deleted files never return space, survives reboots). Linux guests reclaim correctly on both raw and ASIF via `fstrim -av`. Single detailed field report, unrefuted. Since the VM subsystem targets macOS guests, the practical rule until this is fixed is: back macOS guest disks with RAW, and revisit ASIF stacking for macOS guests each beta (stacking still works, the disks just never shrink).
- **Zero shipping adopters anywhere.** tart has a design issue with no activity; nobody has published working DiskImageKit code. Everything must be treated as field-untested (and our own testing bears that out, Section 8).
- Framework binary grew every beta (588 → 598 across betas 1-4); expect churn until GA.
- Release notes list no DiskImageKit known issues in any beta, which given the above says more about the notes than the framework.
## 3. Guest provisioning (macOS guests only)
```swift
class VZGuestProvisioningOptions: NSObject { func validate() throws } // "use one of its subclasses"
class VZMacGuestProvisioningOptions: VZGuestProvisioningOptions {
var fullName, username, password: String
var logsInAutomatically: Bool
var enablesRemoteLogin: Bool // SSH
}
// Wiring: VZMacOSVirtualMachineStartOptions.guestProvisioningOptions (Mac-typed)
// .setGuestProvisioning(_:) throws (validating setter)
```
- **Requires macOS 27 on host AND guest.** Older guests **silently ignore** the options (no error).
- **First boot after restore only.** Cannot reconfigure an already-provisioned VM; property changes after start are no-ops.
- The base class is forward-looking scaffolding; its only subclass is Mac. A Linux/cloud-init analogue may come later; do not assume it lands in 27.0. For Linux guests, cloud-init NoCloud seed ISOs remain the provisioning path (proven working, Section 8).
- Field-verified behavior (third-party hands-on, beta 3): provisioned account gets full admin + sudo; Setup Assistant fully skipped; SSH reachable ~48 s after first boot. **Race**: the account is created late in first boot (~T+54 s), after LaunchDaemons start (~T+33 s), so anything at daemon-level must wait for the account to exist.
- Open Apple-acknowledged bug: provisioned users are invisible to `CSIdentityQueryExecute()` (FB23716201).
- IPSW acquisition gotcha for automation: `VZMacOSRestoreImage.latestSupported` tracks the latest *release* (returned 26.5.2), not the installed beta; beta IPSWs must be fetched from the seed CDN explicitly.
## 4. vmnet: a macOS 26 feature set, one macOS 27 fix
Everything interesting shipped in macOS 26: `vmnet_network_create`, `vmnet_network_configuration_create`, `..._add_port_forwarding_rule`, `..._add_dhcp_reservation`, subnet/prefix/MTU/external-interface setters, NAT44/NAT66/DHCP/DNS-proxy/RA disables, plus serialization (`vmnet_network_copy_serialization` / `_create_with_serialization`) for handing networks across processes. `VZVmnetNetworkDeviceAttachment` is macOS 26.
macOS 27's only change (beta 4 release notes, verbatim): "The vmnet port forwarding APIs now support port forwarding when communicating over loopback." That closes the old gap where the host could not reach its own forwarded ports via 127.0.0.1 (confirmed working by the original bug reporter). Directly relevant to Codeman's loopback-bound production server talking to per-case guests.
Gotchas:
- vmnet networks are **not persisted**; they die with the owning process. Persist settings yourself and recreate (or serialize across processes).
- The `com.apple.vm.networking` entitlement is still restricted ("contact your Apple representative", though DTS says most requests are approved). The plain `VZNATNetworkDeviceAttachment` needs no special entitlement and is what our prototype uses.
- Ecosystem signal: tart's maintainer is not adopting in-process vmnet (prefers their separate-process softnet), so field testing of these APIs is thin.
## 5. VZCustomVirtioDevice (macOS 27, Linux guests only)
14 new types (`VZCustomVirtioDevice(+Configuration/Delegate/Provider)`, `VZVirtioQueue(+Element)`, `VZVirtioFeatureSet`, shared-memory-region types, `VZGuestMemoryMapping`), wired via `VZVirtualMachineConfiguration.customVirtioDevices`. Mandatory for guest discovery: `deviceID`, `pciClassID`, `pciSubclassID`, `virtioQueueCount`. You must write the Linux guest driver (Virtio spec 1.3/1.4). Threading contract: the framework calls the device/delegate on a serial queue (`deviceQueue`, defaulting to the VM's queue). Zero public adopters. For the VM subsystem this is a Phase 3+ option for a low-latency host-guest channel; SSH over NAT is proven and sufficient for now.
## 6. Signing and entitlements
- **Core loop (VZ + DiskImageKit + provisioning): ad-hoc signing with only `com.apple.security.virtualization` suffices.** Verified by us on beta 4 (plain `codesign --entitlements ... -s -` on a `swiftc` binary) and independently by third parties on beta 3. DiskImageKit documents no entitlement at all.
- **Over-entitling is the actual trap.** Adding `com.apple.application-identifier`/team-identifier keys without an embedded provisioning profile hangs the process before `main` (watchdog kill); shipping `com.apple.vm.networking` unauthorized gets AMFI SIGKILL at exec (exit 137, no crash report, even for `--version`). Keep the entitlements plist to exactly the one key.
- **USB passthrough breaks the ad-hoc story**: `com.apple.developer.accessory-access.usb` is profile-restricted (any paid team, no ad-hoc), additionally requires `com.apple.security.device.usb`, and `AAUSBAccessoryManager` presents UI, so it wants a Dock app, not a headless CLI. Out of scope for Codeman.
- No Xcode required for any of the above: the CLT beta (~500 MB via `softwareupdate`) carries the full macOS 27 SDK including DiskImageKit and compiles/signs everything.
## 7. Ecosystem state (July 2026)
- **tart is now `openai/tart`** (moved from cirruslabs, mid-2026) and **relicensed to FSL-1.1-ALv2** (no longer permissive). Provisioning support shipped in 2.33.0. Old cirruslabs URLs and license assumptions are stale.
- VirtualBuddy shipped provisioning ("Skip Setup Assistant") in 2.2 betas; had to add account-detail validation and a workaround installer for the cross-version bug below.
- lima is deliberately waiting for GA before touching macOS 27 APIs.
- **Code-Hex/vz (Go bindings) is dormant** (no commits since Feb 2026, no macOS 27 APIs), so the entire Go ecosystem (podman-machine, colima) currently has no path to these APIs. Swift is the only realistic binding today, which validates the VM subsystem's Swift-helper design.
- Useful pattern if ever supporting older SDKs: resolve new classes via `NSClassFromString` at runtime (no link-time dependency), fail gracefully when absent.
- **Cross-version restore bug**: installing a macOS 27 guest from IPSW on a macOS 26 host fails at 77-78% (`VZErrorDomain 10007`); fixed in 26.6b3 + Xcode 27b4 era, with a nasty MobileDevice.pkg trap (installing it from Xcode 27 beta on a 26 host requires a full macOS reinstall to undo). Not relevant to our 27-host testbed, very relevant to anyone on a 26 host.
## 8. Our empirical results (beta 4, 26A5388g, MacBook Air M3, 2026-07-29)
Prototype tooling, all in `~/vm-lab/` on the testbed, compiled with CLT-only `swiftc` and ad-hoc signed with the single virtualization entitlement:
| Tool | Purpose |
| --- | --- |
| `vzboot.swift` | Linux guest: EFI boot + virtio disk/net/entropy + NAT + optional cloud-init seed ISO + serial on stdio |
| `vzstack.swift` | Same, but boots a DiskImageKit stack (read-only raw base + ASIF overlay) |
| `vzmac.swift` | macOS guest: `install` (IPSW restore into a bundle) and `run` (boot, `--provision` for first-boot account/SSH) |
| `vzmacgui.swift` | macOS guest in a real window via `VZVirtualMachineView` (required for the guest to render at all) |
| `setup-seed.sh` | Builds a cloud-init NoCloud seed ISO with `hdiutil makehybrid` (volume label `cidata`) |
| `vncproxy.py` | RFB proxy that advertises only security type 2, so version-skewed/browser clients can authenticate |
| noVNC + `websockify` | Browser access; `websockify --web noVNC-<ver> 0.0.0.0:<port> 127.0.0.1:<proxy>` |
| `vmwatchdog.sh` + `vmaccess.sh` | Supervision: root LaunchDaemon that restarts a blind/dead runner, re-points the forward, re-applies `pmset`, re-arms keep-awake; plus a keeper for the proxy/web endpoints |
Host-side diagnostics written during this work (in the session scratchpad, not on the testbed): `vnclogin.py` (Apple DH auth + session open, distinguishes "credentials rejected" from "authorized but session refused"), `vncshot.py` (decodes the raw framebuffer to PNG and reports non-black pixel counts, plus optional synthetic wake input), `relay.py` (plain TCP relay used to bridge a tailnet peer to a LAN-only host), `sshpw.py` (pty-driven password SSH for the one-time key bootstrap into a freshly provisioned guest).
### Proven working
1. **Boot**: Debian 12 arm64 cloud images (nocloud and genericcloud variants) boot under `VZEFIBootLoader` + `VZGenericPlatformConfiguration`.
2. **Networking**: `VZNATNetworkDeviceAttachment` gives the guest a `192.168.64.x` DHCP lease from the host's bootpd (leases visible in `/var/db/dhcpd_leases`, bridge is `bridge100`).
3. **cloud-init provisioning**: NoCloud seed ISO (built with `hdiutil makehybrid -iso -joliet -default-volume-name cidata`) created a `codeman` user with SSH key + passwordless sudo on first boot; `ssh codeman@<lease-ip>` from the host works with key auth.
4. **DiskImageKit stack mechanics**: opening a raw base `.readOnly`, appending an ASIF overlay (`ASIFCreationConfiguration.layer(url:type:.overlay)`), attaching via `init(diskImage:)`, and booting it. The overlay received ~44 MB of boot-time writes while the **base file's SHA-256 stayed bit-identical**, which is the write-isolation property the whole per-case design rests on.
5. **Reattach**: reopening an existing overlay and `appending(consuming:)` onto the same base passes UUID validation.
6. **macOS guest install (added later the same day)**: `VZMacOSInstaller` restore of the 27.0 IPSW (26A5388g, fetched from the seed CDN via appledb; same build as host) into a sparse 64 GiB raw disk + auxiliary storage: INSTALL-OK on the first attempt, ~25 minutes.
7. **Headless guest provisioning WORKS**: `VZMacGuestProvisioningOptions` via `setGuestProvisioning` (username, password, `enablesRemoteLogin`, `logsInAutomatically=false`) produced, with zero GUI interaction: an account with full admin (groups include `80(admin)`, `com.apple.access_ssh`), Remote Login on from first boot, port 22 reachable ~140 s after first-boot start, hostname auto-derived from the account ("Codemans-Virtual-Machine"). SSH password auth is on by default, so the bootstrap path is: pty-driven password login once to install `authorized_keys`, key auth thereafter. Note the provisioned account's sudo is NOT passwordless (`echo <pass> | sudo -S ...`), and provisioning is first-boot-only (later boots take no options and just boot).
8. **Slot-leak bug NOT reproduced on 26A5388g**: a guest-initiated `shutdown -h now` fired `guestDidStop` cleanly and an immediate relaunch started fine (SSH-ready again in ~75 s), so FB22967193 (VM slot leaked on guest-initiated shutdown, host reboot to recover) did not manifest after one cycle. Either fixed in beta 4 or needs more cycles to trigger.
### Unstable / under investigation (beta-quality territory)
Boot reliability degraded over a ~15-VM session on one host boot, ending with reproducible silent hangs (VM process alive, 0% CPU, no DHCP, no ARP, nothing on serial):
- A genericcloud base that had been booted read-write once (cloud-init first boot) subsequently hung on every boot **with the seed ISO still attached**, while booting **without** the seed succeeded, then later runs failed in both configurations. The seed correlation is strong but was observed while host state was already suspect, so it needs a retest from a clean baseline.
- The first stack-boot "success" that later wedged turned out (via DHCP lease timestamp arithmetic) never to have reached the network at all; its overlay growth was pre-network boot writes.
- Working hypothesis, matching a class of acknowledged beta bugs (e.g. the VM-slot counter that leaks on guest-initiated shutdown, FB22967193, where only a host reboot recovers): accumulated hypervisor/vmnet state on the host degrades boots. Requires a host reboot + a disciplined retest matrix to confirm.
### Display rendering: the single most important operational finding
**A VZ macOS guest renders nothing unless a `VZVirtualMachineView` is attached AND the host session is actually drawing.** Verified byte-for-byte: the guest's own screen sharing serves an all-zero framebuffer (0 non-black bytes across 400 KB samples, with a sane pixel format: `rmax/gmax/bmax = 255`, shifts 16/8/0), in-guest `screencapture` fails with "could not create image from display", and no `IODisplayWrangler` shows up in the guest's `ioreg`. Three distinct states all produce black:
1. **Headless** (VM run with no view attached).
2. **View attached, host session locked.** The lock screen suspends drawing and the guest's virtual GPU produces no frames.
3. **View attached, but the app lost its WindowServer connection** (see the incident below): black permanently until the app is restarted.
**Consequence for the VM subsystem: rendering is a first-class requirement, not an optional extra (owner decision 2026-07-29).** The product serves GUI desktops: mandatory for macOS guests, optional-but-supported for Linux guests (which can also run headless over SSH). Any VM in GUI mode must be launched by an app that attaches a `VZVirtualMachineView`, from inside a host GUI session that is logged in and unlocked. That makes the following non-negotiable parts of the design, not workarounds:
- VMs run as **GUI apps in the console user's session** (launched via a LaunchAgent or `launchctl asuser`), never as daemons.
- The **host must auto-login and never lock or sleep**; a locked host is equivalent to a powered-off display for every VM on it.
- The **guest must auto-login, never lock, and have its first-login assistant pre-suppressed**, or the "desktop" a user connects to is a password prompt or a setup wizard.
- A VM app that loses its WindowServer connection is **permanently blind** and must be restarted; supervision has to detect that, not just check that the process is alive.
- The **2-concurrent-macOS-VM cap** becomes a real capacity limit for the product, so it must be surfaced in the UI and tested (still untested worldwide as of this writing).
### Incident 2026-07-29: `killall -HUP loginwindow` (never do this on a remote Mac)
Applying a wallpaper change on the testbed with `killall -HUP loginwindow` restarted the host's login session. Three consequences:
1. **The Mac dropped off the tailnet entirely.** Tailscale's App Store build is a GUI app living in the user session, so killing the session killed the VPN; remote access was gone until someone logged in. Recovery came from a second machine on the same LAN: it could still SSH in, and then relay ports back over the tailnet (a plain TCP relay on a tailnet-connected LAN peer is a good out-of-band path worth keeping ready).
2. **The VM app lost its WindowServer connection** (`HIToolbox: received notification of WindowServer event port death`) while surviving as a process. Every later black screen traced to this, and nothing guest-side could fix it; only restarting the app restored rendering.
3. The session's `caffeinate` died, so the host resumed auto-locking.
Rule: on a remote Mac, never run session-level commands (`killall -HUP loginwindow`, `pkill -u <user>`, logout, fast user switching). `killall WallpaperAgent` alone is session-safe. Before any such command, enumerate what depends on that session: VPN, VM processes, port forwards, keep-awake helpers.
### Keeping host and guest usable unattended
- **Host**: `caffeinate -d -i -m -u` prevents display sleep but does NOT override the lock policy. "Require password after screen saver begins or display is turned off → Never" must be set in System Settings; it needs the account password, so a passwordless-sudo shell cannot script it, and turning it off does NOT dismiss a lock that is already engaged (one more unlock is always needed). `pmset -a disablesleep 1` keeps a lid-closed laptop awake but **does not survive a reboot**, and OS updates reset it too, so a supervisor should re-apply it rather than assume it sticks.
- **Rebooting an encrypted host**: use `sudo fdesetup authrestart`. FileVault's pre-boot unlock doubles as the login, so the machine returns with a **live logged-in console session** and encryption intact, no password prompt, and supervision can then bring the VMs back by itself. Verified 2026-07-30. A plain `reboot` parks at the lock screen and blacks out every VM until a human logs in.
- **Guest**: set `autoLoginUser` plus a valid `/etc/kcpassword` (XOR-obfuscated password file, key `7D 89 52 23 D2 BC DE A3`, payload zero-padded to a multiple of 12). `sysadminctl -autologin` fails with `SACSetAutoLoginPassword error:22` on provisioned accounts, and a fresh guest has no Python, so generate the bytes on the controlling host and copy them in. Then `pmset -a displaysleep 0 sleep 0 disablesleep 1`, `defaults -currentHost write com.apple.screensaver idleTime 0`, `defaults write com.apple.screensaver askForPassword 0`, and `caffeinate` inside the guest. ⚠ `autoLoginUser` was observed being wiped by failed `sysadminctl -autologin` attempts; verify it after each boot until stable.
- **Wallpaper**: animated "aerials" wallpaper is brutal over VNC. The provider lives in `~/Library/Application Support/com.apple.wallpaper/Store/Index.plist` under several keys (`AllSpacesAndDisplays:Desktop`, `:Idle`, and `SystemDefault:*` which is what the login/lock screen uses). Switch each `Provider` to `com.apple.wallpaper.choice.solid-color` with PlistBuddy and restart `WallpaperAgent`. The login-window copy is cached and only refreshes on a later login cycle.
### Remote GUI/SSH access to a guest (recipe, verified 2026-07-29)
The guest lives on the host-private NAT bridge, so remote access is guest-service + host-forward:
1. **In the macOS guest** (over ssh), use ONE mechanism, fully activated. The reliable form is Remote Management in a single kickstart call:
```
sudo .../RemoteManagement/ARDAgent.app/Contents/Resources/kickstart \
-activate -configure -access -on \
-clientopts -setvnclegacy -vnclegacy yes -setvncpw -vncpw <8-char-pw> \
-allowAccessFor -allUsers -privs -all -restart -agent -menu
```
⚠ **Half-configured states authenticate but refuse the session.** Loading `com.apple.screensharing` while Remote Management is deactivated (or vice versa) produces an Apple-client error that names the wrong culprit: *"Screen Sharing is not permitted on <host>. Disable and re-enable Screen Sharing or Remote Management in System Settings"*. A raw-protocol client can still authenticate AND open a framebuffer in that state, so protocol-level tests pass while every Apple client fails. The remedy is exactly what the dialog says, done over ssh: `launchctl unload -w …screensharing.plist`, `kickstart -deactivate -configure -access -off`, `pkill screensharingd`, then the single activate call above.
Notes: `launchctl enable system/com.apple.screensharing` fails with "Could not find service" on this build; `load -w` is the plain-Screen-Sharing path if you deliberately want it instead of Remote Management. Apple clients negotiate `RSA-SRP` (auth type 33) and the guest logs `Authentication: SUCCEEDED :: User Name: … :: Type: RSA-SRP` on success, which is the definitive server-side confirmation.
2. **On the host**: a gateway port-forward makes the guest's 5900 reachable from the whole tailnet without per-client tunnels: self-authorize the host's own key, then `ssh -N -g -L 0.0.0.0:5901:<guest-ip>:5900 <user>@localhost` (nohup'd).
⚠⚠ **NEVER forward on host port 5900.** If the host has Screen Sharing enabled (our testbed does, from the pre-upgrade checklist), launchd already owns 5900 socket-activated. The `ssh -L` bind then fails with "Address already in use" **while the tunnel process keeps running**, so every symptom of success is present (process alive, port answers, real RFB banner) yet **every connection reaches the HOST's login window, not the guest**. This cost us an hour: guest credentials failed against the host's screensharingd, which reads exactly like broken guest auth, and we chased the (real, but irrelevant) provisioned-account identity bug. Diagnostics that would have caught it instantly: `sudo lsof -nP -iTCP:5900 -sTCP:LISTEN` showing `launchd` rather than `ssh`, or the guest's own logs showing NO auth attempts during a failed login. Always use a distinct host port and verify with `lsof` that the forward owns it.
⚠ `-g` binds all interfaces, so the forward is also visible on the host's LAN; the VNC layer still requires the account or VNC password. ⚠ The forward pins the guest IP, which changes per boot under plain NAT; re-point it after a guest reboot (the proper fix is a vmnet DHCP reservation, macOS 26 API, once we move off plain `VZNATNetworkDeviceAttachment`).
Verified working: with the forward on 5901, both a provisioned account and a `sysadminctl`-created one authenticate successfully (RFB `SecurityResult` = 0) against the guest. The guest offers security types `[30, 33, 36, 2, 35]`, i.e. Apple DH/SRP **plus classic type 2**, so non-Apple VNC clients work with the legacy password once ARD's `-setvnclegacy` is set. (The host's screensharingd, by contrast, offered no type 2, which is itself a tell that you are talking to the wrong machine.)
3. **SSH from any tailnet device**: `ssh -J <host-user>@<host> codeman@<guest-ip>` (jump through the host), after adding the connecting machine's key to the guest's `authorized_keys`.
**Client-version incompatibility (macOS 27 servers vs older Screen Sharing clients)**: an older Mac's Screen Sharing client fails Apple's `RSA-SRP` handshake against macOS 27 servers, logging `Authentication: FAILED :: User Name: <user> :: Type: RSA-SRP` server-side, while a macOS 27 client authenticates against the same servers without issue. This was verified against BOTH a macOS 27 guest and a macOS 27 host with the operator's own account, so it is a client-side version skew, not configuration, and no server-side change fixes it. Same family as the documented "macOS 26 host cannot install a 27 guest" bug. Practical workaround: bypass Apple auth entirely with classic VNC auth (security type 2), which macOS offers only when Remote Management legacy VNC is enabled. Two ways to consume it: any third-party VNC client, or a browser via noVNC.
**Browser-based access chain (zero client install, version-proof)**, all hosted on the Mac:
```
browser --HTTP/WS--> websockify (+ noVNC static files)
--> type-2-only proxy # rewrites the server's security-type list to [2]
--> ssh -L forward # loopback hop; see the Local Network note below
--> guest:5900
```
Notes learned the hard way: (a) **never bind the forward on host port 5900** (see the launchd warning above); (b) a Python proxy cannot reach the guest subnet directly because macOS **Local Network privacy** denies headless CLI binaries, surfacing as `No route to host`, so point the proxy at a loopback `ssh -L` forward instead (Apple-signed `ssh` is unaffected); (c) noVNC needs `?resize=scale` or Scaling Mode → Local Scaling, otherwise a Retina host screen (2940x1912) is unusable in a browser window; (d) noVNC speaks security type 2 only, which is exactly why the proxy rewrite is needed.
**Debugging technique that settled all of this**: a ~80-line Python RFB client (scratchpad `vnclogin.py`) that implements Apple DH auth (security type 30) and continues through `ClientInit`/`ServerInit`. It reports the server's `SecurityResult` plus the framebuffer size and desktop name, which separates "credentials rejected" from "authorized but session refused" without any GUI client. Pair it with `log stream --predicate 'process == "screensharingd"'` inside the guest, and drive a REAL Apple client headlessly from the host with `sudo launchctl asuser <uid> sudo -u <user> osascript -e 'tell application "Screen Sharing" to open location "vnc://user:pass@host:port"'`, verifying the result via `lsof -nP -iTCP -a -p <pid>` (an ESTABLISHED socket to the target) since `screencapture` fails on a lid-closed laptop ("could not create image from display"). Tailscale was never implicated: both the raw client and Apple's client work over the tailnet address once the guest service is fully activated.
### Hard-won operational lessons (write these into any tooling)
- **Silent serial is normal, not failure.** Debian's GRUB/kernel log to the graphics console; nothing attaches a getty to hvc0 by default. The reliable boot signal is the DHCP lease (or passive `tcpdump -i bridge100`), never the serial port and never a quick ping (BSD ping's first packet often dies to ARP latency; passive capture showed "dead" guests alive).
- **DHCP lease entries carry truth**: `name=` shows the guest hostname, and the lease timestamps order events; stale entries linger, so compare timestamps before attributing a lease to a boot.
- **Never boot a base image read-write.** Every RW boot mutates it (dhclient lease cache, journal, cloud-init state) and destroys experiment reproducibility, exactly why the production design only ever boots bases under overlays. Provision INTO the base once at base-build time, or provision per-case overlays with the seed, then detach the seed.
- **A killed SSH client does not kill a remote `nohup`'d VM**, and the survivor holds the EFI variable store lock: "The EFI variable store is already in use" (`VZErrorDomain 50002`) means a zombie VM process, `pkill` it.
- **EFI variable stores are per-VM state.** Fresh stores boot reliably; reuse across different VM instances is at minimum suspect on this beta (Apple's own guidance for cloned VMs is one store per VM). Cheap policy: one store per case, created with the overlay, deleted with it.
- **Downloads from cloud.debian.org mirrors truncate silently**; always verify byte count against origin `Content-Length` and resume with `curl -C -`.
- The remote host's default shell is zsh: `=` -prefixed words (`echo ===`) explode via zsh's `=cmd` expansion; keep separators zsh-safe in automation.
### The 2-concurrent-macOS-VM cap: TESTED AND CONFIRMED on macOS 27 beta 4 (2026-07-29)
We measured it, which as far as we can tell nobody had published for macOS 27. Method: `cp -c -R` the guest bundle (APFS clonefile, instant and **zero additional disk**), regenerate the machine identifier per clone (`VZMacMachineIdentifier()` written to `machine.id`; the hardware model is reused), then launch VMs until one is refused.
Result: VM #1 (8 GB, GUI) and VM #2 (4 GB, headless) ran concurrently without complaint. VM #3 was refused **instantly** at `vm.start`:
```
VZErrorDomain Code=6 "The maximum supported number of active virtual machines has been reached."
NSLocalizedFailure = "The number of virtual machines exceeds the limit."
```
**This is a licensing/kernel quota, not a resource limit**: the refusal came with **39% of system memory free** on a 16 GB host, and adding RAM or CPU cannot raise it. It matches the pre-27 behavior (`hv_apple_isa_vm_quota`), so nothing changed in 27 despite the framework's other additions. Linux guests are unaffected and are bounded only by host resources.
Design consequences: macOS-guest capacity per host is **hard-capped at 2**, so a GUI-macOS-per-case product must schedule around it (queue, evict idle VMs, or scale across hosts) and surface it in the UI. Also relevant: the acknowledged slot-leak bug (a guest-initiated shutdown failing to release a slot, recoverable only by host reboot) is far more damaging under a cap of 2 than it sounds; we did not reproduce it on beta 4, but any scheduler should treat "slot appears used but nothing is running" as a real state.
### Not yet tested
- Cache layers (`LayerType.cache`), `.overlay(blockCount:)` disk growth, stack depth performance, VirtioFS + stack combination, `truncate`, ASIF disks for macOS guests (raw used so far; ASIF has the reclamation bug).
- One more scripting lesson from this session: inner `ssh` calls inside a piped `sh -s` script MUST use `-n`, or they consume the remainder of the script from stdin and it silently never runs.
### Session timeline (what was actually established, 2026-07-29)
Linux path: base image download (with resume, mirrors truncate) → `vzboot` compiles against the beta SDK first try → EFI boot → NAT DHCP lease → cloud-init seed provisions a user with the host's SSH key → `ssh` into the guest works → DiskImageKit stack boots with an ASIF overlay taking all writes while the base stays SHA-identical. Later Linux boots became unreliable on an un-rebooted host (silent hangs, 0% CPU, no DHCP); a clean-baseline retest is still pending.
macOS path: seed-CDN IPSW (matched to the host build) → `VZMacOSInstaller` restore, ~25 min, first try → first boot with `VZMacGuestProvisioningOptions` creates an admin account with Remote Login on, no interaction needed, SSH reachable ~140 s later → key bootstrap over a one-time password login → guest shutdown/relaunch clean (the slot-leak bug did not reproduce) → GUI access fought through a port collision, a client-version incompatibility, the rendering dependency, and a self-inflicted session kill, ending with a browser-based path plus a guest hardened to auto-login and never lock.
**Lifecycle verified (stop → start), 2026-07-30**: an in-guest `shutdown -h now` fires `guestDidStop` and the runner app exits on its own; relaunching from the same bundle boots the guest in ~2 minutes straight into an auto-logged-in desktop, and the VM slot is released cleanly (an immediate restart works, so the slot-leak bug did not bite). Two operational notes: the guest takes a **new NAT lease on every boot**, so any port-forward must be re-pointed (or use a vmnet DHCP reservation), and a host reboot resets `pmset -a disablesleep`.
⚠ **Provisioning does NOT skip the per-user first-login assistant.** `VZMacGuestProvisioningOptions` skips the initial Setup Assistant (account creation, region, Apple Account) so the machine is immediately reachable, but the first time anyone actually logs into a desktop, macOS still presents its per-user wizard (Apple Intelligence, Siri, privacy, appearance, Touch ID). The operator hit exactly this. For a GUI-first product this MUST be pre-suppressed during base-image creation by writing `com.apple.SetupAssistant` keys for every account that will log in, and into `/System/Library/User Template/English.lproj/Library/Preferences/` so accounts created later inherit it.
⚠ **A partial key list is worse than none**, because the wizard simply shows the panes you missed and the operator has to click through them again after every fresh login (we hit this twice). The set that finally silenced macOS 27 beta 4: `DidSeeCloudSetup`, `DidSeeSiriSetup`, `DidSeePrivacy`, `DidSeeAppearanceSetup`, `DidSeeTouchIDSetup`, `DidSeeAvatarSetup`, `DidSeeScreenTime`, `DidSeeApplePaySetup`, `DidSeeSafariImport`, `DidSeeAccessibility`, **`DidSeeActivationLock`, `DidSeeAppStore`, `DidSeeLockdownMode`** (the three easy to miss), plus the Express-Settings flags **`SkipExpressSettingsUpdating`** and **`SkipFirstLoginOptimization`**, and the version markers `LastSeenCloudProductVersion` / `LastSeenBuddyBuildVersion` / `PreviousSystemVersion` / `PreviousBuildVersion` matching the guest build. Verify afterwards by reading the domain back and checking that no `DidSee*` key is still `0`. Note these keys change between macOS releases, so base-image creation should re-verify per OS version rather than trust a hardcoded list.
## 9. Design implications for Codeman's VM subsystem
0. **GUI is a first-class mode, and for macOS guests it is the whole point (owner decision, 2026-07-29).** The subsystem serves real desktops, not only headless SSH boxes. macOS guests are GUI-only in practice (nothing renders without an attached view). Linux guests are supported in BOTH modes: GUI when the case wants a desktop, headless-over-SSH when it wants a cheap agent sandbox. The costs of the GUI path are in §8 "Display rendering": VMs as GUI apps in a live session, a host that never locks, guests that auto-login with their first-login wizard pre-suppressed, and the macOS concurrency cap as a real capacity limit.
1. **The macOS-specific liabilities are accepted costs, not reasons to avoid macOS guests**: provisioning is macOS-only and first-boot-only, ASIF space reclamation is broken for macOS guests on the beta (use RAW disks for macOS guests until fixed), and the 2-VM cap applies. Plan around each: RAW-backed macOS disks, provisioning baked into base-image creation, and capacity limits surfaced in the UI.
2. **Base immutability is not just hygiene, it is load-bearing**: DiskImageKit's UUID invalidation plus our sha-stability proof make a read-only shared base per image-generation the core artifact. Bases are built once (seed attached), then only ever opened `.readOnly` under per-case overlays.
3. **Seed ISOs are a base-build-time tool only.** Never attach a seed to a routine case boot (correlated with boot hangs on the beta, and semantically wrong anyway since cloud-init already ran).
4. **Per-case files**: overlay ASIF + EFI variable store live and die together with the case.
5. **Export = ship the layer chain** (base ref + overlay + manifest), not flatten; there is no flatten API. In-guest `dd` to a fresh image is the fallback for a true single-file export.
6. **Health checking must be lease/API based**, not serial/ping based, and Codeman's `codeman-vm status` should read `/var/db/dhcpd_leases` (or use vmnet DHCP reservations for deterministic per-case IPs, a macOS 26 API).
7. **Run `fstrim` periodically in Linux guests** (or mount with discard) so overlays stay sparse.
8. **Entitlements plist stays minimal** (exactly `com.apple.security.virtualization`) to dodge the AMFI/watchdog traps.
9. **Expect beta churn**: pin findings to build numbers (this doc: 26A5388g) and retest each beta; the framework binaries changed every beta so far.
10. **A macOS guest is only "ready" when its desktop is ready**, which is a stricter bar than "the VM booted". Readiness means: VM app running with a live WindowServer connection, guest auto-logged-in (not at a login or lock screen), first-login assistant suppressed, and the guest's screen sharing serving a non-black framebuffer. Health checks should sample the framebuffer for non-black content, because every failure mode in this session (headless run, locked host, dead WindowServer, locked guest, setup wizard) presents as a perfectly healthy-looking process with a black or useless screen.
10b. **Supervision must run as a root LaunchDaemon.** A user LaunchAgent cannot launch a GUI app into the Aqua session; its restarts fail silently (child dies instantly, empty log, supervisor reports success). Root + `launchctl asuser <uid> sudo -u <user> …` works and the launched process persists. This bit us on the first supervisor implementation and is easy to repeat.
11. **Remote-access plumbing belongs in the helper CLI, not in ad-hoc shell**: a `codeman-vm` implementation should own port selection (never 5900), forward lifecycle across guest IP changes (or better, vmnet DHCP reservations for stable per-case IPs), and a documented browser path, because every failure in this session came from hand-rolled plumbing rather than from the Virtualization APIs themselves.
12. **Never let control-plane connectivity depend on a GUI session** on a remote Mac host: prefer a Tailscale system service over the App Store app, and keep a LAN-adjacent peer able to relay as an out-of-band recovery path.
## Sources
Apple DocC JSON backend (diskimagekit, virtualization, vmnet trees; macOS 27 release notes) | WWDC26 session 224 https://developer.apple.com/videos/play/wwdc2026/224/ | eclecticlight.co ASIF/virtualization coverage | developer.apple.com/forums threads 839343 (CSIdentity bug), 830118 (cross-version restore), 830119 (VM-slot leak), 830383 (VM cap), 834822 + 831902 (USB entitlements), 822658 (vmnet loopback) | openai/tart issues 1261/1263/1268/1269/1285 | Spooky-Labs provisioning design doc | VirtualBuddy 2.2 release notes | lima-vm discussions | our own test transcripts on the testbed (`~/vm-lab/*.log`, this repo's session)
+2 -2
View File
@@ -1,7 +1,7 @@
# Web tabs: two fixes (planned + implemented 2026-07-28)
Both found against the saved dashboard
`https://macminis-mac-mini.tailf80371.ts.net:4000` (Bio-Hacking-Dashboard).
`https://<your-host>.<your-tailnet>.ts.net:4000` (Bio-Hacking-Dashboard).
Kept because the root-cause analysis of the second one is not obvious from the
resulting diff.
@@ -49,7 +49,7 @@ already revoked the capability and broadcast `WebviewChanged`.
```
CAP=<from POST /api/webviews/<id>/open>
# A) upstream direct -> 200 image/jpeg 118150
curl -sk "https://macminis-mac-mini.tailf80371.ts.net:4000/api/hero?slug=120-minutes-in-nature"
curl -sk "https://<your-host>.<your-tailnet>.ts.net:4000/api/hero?slug=120-minutes-in-nature"
# B) through the proxy prefix -> 200 image/jpeg 118150
curl -sk "https://localhost:3000/webview/$CAP/api/hero?slug=120-minutes-in-nature"
# C) what the browser ACTUALLY requested -> 404 {"errorCode":"NOT_FOUND"}
+1 -1
View File
@@ -1,7 +1,7 @@
# Web Tabs (dashboards as Codeman tabs)
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port
4000, as a tab beside your Claude/Codex/Gemini sessions. Codeman becomes one mission
4000, as a tab beside your Claude/Codex/Antigravity sessions. Codeman becomes one mission
control instead of Codeman plus a pile of browser tabs.
## Using it
+612 -30
View File
@@ -21,6 +21,16 @@
# non-interactive default is 127.0.0.1)
# CODEMAN_PASSWORD - Preset the dashboard password (skips the
# password prompt when binding to the network)
# CODEMAN_TAILSCALE=1 - Preset the Tailscale choice: bind loopback and
# front it with `tailscale serve` HTTPS (skips
# the network prompt; never installs Tailscale
# in non-interactive runs)
#
# Subcommands:
# install.sh update - Update an existing install
# install.sh uninstall - Remove services, symlinks and (optionally) data
# install.sh tailscale - Set up (or repair) Tailscale serve HTTPS access
# for an existing install
set -euo pipefail
@@ -51,6 +61,14 @@ EXISTING_HOST=""
EXISTING_PASSWORD=""
EXISTING_ACK="0"
# Tailscale serve URL configured or detected during this run
# (setup_tailscale_access / detect_tailscale_serve_url). Empty when the
# Tailscale path was not taken or not completed.
TAILSCALE_SERVE_URL=""
# Set to 1 when serve commands must go through sudo because granting the user
# tailscale "operator" rights failed (ensure_tailscale_operator).
TS_NEED_ROOT="0"
# puppeteer is a devDependency used only by scripts/browser-comparison.mjs — its
# ~150MB chrome-headless-shell download is never needed to build or run Codeman.
# Skipping it avoids a slow download and a fatal install failure when a prior
@@ -98,6 +116,14 @@ GEMINI_SEARCH_PATHS=(
"$HOME/bin/gemini"
)
# Antigravity CLI search paths (from src/utils/antigravity-cli-resolver.ts)
ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/.local/bin/agy"
"$HOME/.antigravity/bin/agy"
"/usr/local/bin/agy"
"$HOME/bin/agy"
)
# ============================================================================
# Color Output
# ============================================================================
@@ -177,13 +203,28 @@ print_security_notice() {
echo -e " For access from OUTSIDE your network, prefer Tailscale or a tunnel."
echo -e " ${DIM}Details: docs/security-architecture.md${NC}"
else
echo -e " ${YELLOW}${BOLD}Security:${NC}"
echo -e " Codeman binds ${BOLD}127.0.0.1${NC} (this machine only) — no password needed by default."
echo -e " To reach it from another device, do ONE of:"
echo -e " ${CYAN}•${NC} tailscale serve / cloudflared tunnel ${DIM}(recommended)${NC}, or"
echo -e " ${CYAN}•${NC} ${CYAN}codeman web --host 0.0.0.0${NC} AND set ${CYAN}CODEMAN_PASSWORD${NC}"
echo -e " A non-loopback bind without a password still starts, but warns loudly."
echo -e " ${DIM}Details: docs/security-architecture.md${NC}"
# Loopback bind: when a tailscale serve mapping fronts it, lead with
# the actual URL instead of the generic "do ONE of" list. Detection is
# dynamic (tailscaled state is the single source of truth).
local notice_ts_url="$TAILSCALE_SERVE_URL"
if [[ -z "$notice_ts_url" ]]; then
notice_ts_url=$(detect_tailscale_serve_url 2>/dev/null) || notice_ts_url=""
fi
if [[ -n "$notice_ts_url" ]]; then
echo -e " ${YELLOW}${BOLD}Security:${NC}"
echo -e " Codeman binds ${BOLD}127.0.0.1${NC}, fronted by Tailscale serve:"
echo -e " reachable at ${BOLD}$notice_ts_url${NC} (HTTPS, your tailnet only)."
echo -e " Tailscale authenticates every device before traffic reaches Codeman."
echo -e " ${DIM}Details: docs/security-architecture.md${NC}"
else
echo -e " ${YELLOW}${BOLD}Security:${NC}"
echo -e " Codeman binds ${BOLD}127.0.0.1${NC} (this machine only) — no password needed by default."
echo -e " To reach it from another device, do ONE of:"
echo -e " ${CYAN}•${NC} tailscale serve / cloudflared tunnel ${DIM}(recommended)${NC}, or"
echo -e " ${CYAN}•${NC} ${CYAN}codeman web --host 0.0.0.0${NC} AND set ${CYAN}CODEMAN_PASSWORD${NC}"
echo -e " A non-loopback bind without a password still starts, but warns loudly."
echo -e " ${DIM}Details: docs/security-architecture.md${NC}"
fi
fi
echo ""
}
@@ -460,6 +501,34 @@ get_gemini_path() {
done
}
check_antigravity() {
if command -v agy &>/dev/null; then
return 0
fi
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
get_antigravity_path() {
if command -v agy &>/dev/null; then
command -v agy
return
fi
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
check_cloudflared() {
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
@@ -1050,6 +1119,7 @@ read_existing_binding() {
# preset. The server binary itself still defaults to 127.0.0.1 either way.
choose_network_binding() {
# Preset via environment: honor it and skip the prompt entirely.
# CODEMAN_TAILSCALE=1 composes with a loopback (or absent) CODEMAN_HOST.
if [[ -n "${CODEMAN_HOST:-}" ]]; then
BIND_HOST="$CODEMAN_HOST"
BIND_PASSWORD="${CODEMAN_PASSWORD:-}"
@@ -1057,6 +1127,20 @@ choose_network_binding() {
BIND_ACK="1"
fi
info "Network binding preset via CODEMAN_HOST: $BIND_HOST"
if [[ "${CODEMAN_TAILSCALE:-0}" == "1" ]]; then
if [[ "$BIND_HOST" == "127.0.0.1" ]]; then
setup_tailscale_access || true
else
warn "CODEMAN_TAILSCALE=1 ignored: CODEMAN_HOST=$BIND_HOST is not loopback."
fi
fi
return 0
fi
if [[ "${CODEMAN_TAILSCALE:-0}" == "1" ]]; then
BIND_HOST="127.0.0.1"
BIND_PASSWORD="${CODEMAN_PASSWORD:-}"
info "Tailscale access preset via CODEMAN_TAILSCALE=1"
setup_tailscale_access || true
return 0
fi
@@ -1077,21 +1161,48 @@ choose_network_binding() {
return 0
fi
# Default follows the existing setup when there is one, else network.
local default_choice="1"
# Tailscale state, for the menu hint and the default choice. Detection
# only; never installs, logs in, or prompts for sudo here.
local ts_hint="will be installed for you" ts_ready="0" ts_detected_url=""
if check_tailscale; then
ts_hint="installed, needs login"
if command -v node &>/dev/null && [[ "$(ts_status_field 's.BackendState')" == "Running" ]]; then
ts_ready="1"
ts_hint="already connected"
ts_detected_url=$(detect_tailscale_serve_url) || ts_detected_url=""
if [[ -n "$ts_detected_url" ]]; then
ts_hint="already serving Codeman"
fi
fi
fi
# Defaults: an existing setup wins (existing loopback installs default to
# Tailscale only when its serve mapping is already present); fresh installs
# default to Tailscale when it is already connected, else network access.
# A bare Enter never pulls in new software.
local default_choice="2"
if [[ "$EXISTING_FOUND" == "1" && "$EXISTING_HOST" == "127.0.0.1" ]]; then
default_choice="2"
if [[ -n "$ts_detected_url" ]]; then
default_choice="1"
else
default_choice="3"
fi
elif [[ "$EXISTING_FOUND" != "1" && "$ts_ready" == "1" ]]; then
default_choice="1"
fi
echo -e " ${BOLD}Network access${NC}"
echo ""
echo -e " How should the Codeman dashboard be reachable?"
echo ""
echo -e " ${CYAN}1)${NC} ${BOLD}Any device on your network${NC} ${DIM}(0.0.0.0)${NC}"
echo -e " Open it straight from your phone or laptop."
echo -e " ${CYAN}1)${NC} ${BOLD}Tailscale${NC} ${DIM}($ts_hint)${NC}"
echo -e " Private VPN access from your phone or laptop, anywhere."
echo -e " Real HTTPS, no password needed: your tailnet is the login."
echo -e " ${CYAN}2)${NC} ${BOLD}Any device on your network${NC} ${DIM}(0.0.0.0)${NC}"
echo -e " Open it straight from your phone or laptop on the same Wi-Fi."
echo -e " ${YELLOW}Less safe: set a password so only you control your agents.${NC}"
echo -e " ${CYAN}2)${NC} ${BOLD}This machine only${NC} ${DIM}(127.0.0.1)${NC}"
echo -e " Safest. Reach it remotely via Tailscale or a tunnel."
echo -e " ${CYAN}3)${NC} ${BOLD}This machine only${NC} ${DIM}(127.0.0.1)${NC}"
echo -e " Safest. Reach it remotely via Tailscale or a tunnel later."
echo ""
if [[ "$EXISTING_FOUND" == "1" ]]; then
echo -e " ${DIM}Current setup: $EXISTING_HOST$([[ -n "$EXISTING_PASSWORD" ]] && echo ", password set"). Enter keeps it.${NC}"
@@ -1100,21 +1211,55 @@ choose_network_binding() {
local bind_choice=""
while true; do
echo -en "${CYAN}Choose [1/2] (default $default_choice):${NC} " >&2
echo -en "${CYAN}Choose [1/2/3] (default $default_choice):${NC} " >&2
read_reply bind_choice || bind_choice="$default_choice"
bind_choice="${bind_choice:-$default_choice}"
case "$bind_choice" in
1|2) break ;;
*) echo "Please enter 1 or 2." >&2 ;;
1|2|3) break ;;
*) echo "Please enter 1, 2, or 3." >&2 ;;
esac
done
if [[ "$bind_choice" == "2" ]]; then
if [[ "$bind_choice" == "3" ]]; then
BIND_HOST="127.0.0.1"
success "Binding 127.0.0.1 (this machine only)"
return 0
fi
if [[ "$bind_choice" == "1" ]]; then
BIND_HOST="127.0.0.1"
setup_tailscale_access || true
# Password is optional here: the tailnet already authenticates devices.
# An existing password is always kept (never silently loosen).
if [[ -n "$EXISTING_PASSWORD" ]]; then
BIND_PASSWORD="$EXISTING_PASSWORD"
info "Keeping the existing dashboard password"
elif [[ -n "${CODEMAN_PASSWORD:-}" ]]; then
BIND_PASSWORD="$CODEMAN_PASSWORD"
info "Using CODEMAN_PASSWORD from the environment"
elif prompt_yes_no "Add a dashboard password too? (optional; your tailnet already authenticates your devices)" "n"; then
local ts_pw="" ts_pw2=""
while true; do
echo -en "${CYAN}Dashboard password:${NC} " >&2
read_secret ts_pw || ts_pw=""
if [[ -z "$ts_pw" ]]; then
info "No password set"
break
fi
echo -en "${CYAN}Confirm password:${NC} " >&2
read_secret ts_pw2 || ts_pw2=""
if [[ "$ts_pw" == "$ts_pw2" ]]; then
BIND_PASSWORD="$ts_pw"
success "Password set (login user: admin)"
break
fi
echo "Passwords do not match, try again." >&2
done
fi
return 0
fi
# Keep a custom non-loopback host from a previous install (e.g. a specific
# interface IP); otherwise bind all interfaces.
if [[ "$EXISTING_FOUND" == "1" && -n "$EXISTING_HOST" && "$EXISTING_HOST" != "127.0.0.1" ]]; then
@@ -1162,6 +1307,403 @@ choose_network_binding() {
return 0
}
# ============================================================================
# Tailscale Access (loopback bind fronted by `tailscale serve` HTTPS)
# ============================================================================
# The recommended remote-access setup: Codeman stays on 127.0.0.1 and
# tailscaled fronts it with a real Let's Encrypt certificate for
# https://<node>.<tailnet>.ts.net, reachable from the user's tailnet only.
# The app side needs zero configuration (.ts.net is in the server's trusted
# host suffixes). All state lives in tailscaled: no marker files, `tailscale
# serve status` is the single source of truth, and `--bg` config persists
# across reboots on its own.
#
# Safety rule for every function here: NEVER `tailscale serve reset` and never
# touch mappings other than 443 -> Codeman's port. Users may have unrelated
# serve config (other ports, other apps) that a reset would destroy.
get_tailscale_path() {
if command -v tailscale &>/dev/null; then
command -v tailscale
return 0
fi
# macOS GUI app (App Store or brew cask) ships the CLI inside the bundle
# and does not put it on PATH.
if [[ -x "/Applications/Tailscale.app/Contents/MacOS/Tailscale" ]]; then
echo "/Applications/Tailscale.app/Contents/MacOS/Tailscale"
return 0
fi
return 1
}
check_tailscale() {
get_tailscale_path >/dev/null 2>&1
}
ts_cmd() {
local ts_bin
ts_bin=$(get_tailscale_path) || return 127
"$ts_bin" "$@"
}
# Serve mutations need root or "operator" rights on Linux; TS_NEED_ROOT is set
# by ensure_tailscale_operator when the operator grant failed. Detection paths
# run with TS_NEED_ROOT=0 and must never trigger a sudo prompt.
ts_cmd_serve() {
local ts_bin
ts_bin=$(get_tailscale_path) || return 127
if [[ "$TS_NEED_ROOT" == "1" ]]; then
run_as_root "$ts_bin" "$@"
else
"$ts_bin" "$@"
fi
}
# ts_status_field <js-expr>: evaluate an expression against the parsed
# `tailscale status --json` object bound to `s`, printing the result (empty on
# any error). node is guaranteed at every call site (the installer installs it
# before the binding prompt; the subcommand requires a completed install).
ts_status_field() {
ts_cmd status --json 2>/dev/null | node -e '
let d = "";
process.stdin.on("data", (c) => (d += c));
process.stdin.on("end", () => {
try {
const s = JSON.parse(d);
const v = eval(process.argv[1]);
if (v !== undefined && v !== null && v !== false) process.stdout.write(String(v));
} catch {}
});
' "$1" 2>/dev/null
}
# Print the local port that the :443 web handler proxies to, empty when 443 is
# unconfigured. Any scheme counts (http://, and https+insecure:// from setups
# where Codeman itself runs --https), so legacy configs are recognized as ours.
ts_serve_443_target_port() {
ts_cmd_serve serve status --json 2>/dev/null | node -e '
let d = "";
process.stdin.on("data", (c) => (d += c));
process.stdin.on("end", () => {
try {
const s = JSON.parse(d);
for (const [hostport, cfg] of Object.entries(s.Web || {})) {
if (!hostport.endsWith(":443")) continue;
const proxy = cfg && cfg.Handlers && cfg.Handlers["/"] && cfg.Handlers["/"].Proxy;
if (!proxy) continue;
const m = String(proxy).match(/:(\d+)\/?$/);
if (m) process.stdout.write(m[1]);
return;
}
} catch {}
});
' 2>/dev/null
}
# Print https://<node>.<tailnet>.ts.net when tailscale is running AND serve
# already forwards 443 to Codeman's port; print nothing otherwise. Safe to call
# anywhere (no sudo, no side effects); used by the security notice, uninstall,
# and the re-run default.
detect_tailscale_serve_url() {
check_tailscale || return 0
command -v node &>/dev/null || return 0
[[ "$(ts_status_field 's.BackendState')" == "Running" ]] || return 0
local port="${CODEMAN_PORT:-3000}"
[[ "$(ts_serve_443_target_port)" == "$port" ]] || return 0
local dns
dns=$(ts_status_field 's.Self && s.Self.DNSName')
[[ -n "$dns" ]] || return 0
echo "https://${dns%.}"
}
tailscale_retrofit_hint() {
warn "$1: falling back to local-only access (127.0.0.1)."
echo -e " ${DIM}Set up Tailscale access any time later with:${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC}" >&2
}
offer_install_tailscale() {
if [[ "$NONINTERACTIVE" == "1" ]]; then
info "Tailscale is not installed; skipping (non-interactive runs never install it)."
return 1
fi
headless_guard "install Tailscale (curl | sh from tailscale.com)"
if [[ "$(uname -s)" == "Darwin" ]]; then
if command -v brew &>/dev/null; then
if ! prompt_yes_no "Tailscale is not installed. Install it now with Homebrew?" "y"; then
return 1
fi
if ! brew install --cask tailscale; then
warn "Homebrew install failed."
return 1
fi
open -a Tailscale 2>/dev/null || true
info "Log in via the Tailscale menu-bar app if it asks."
else
info "Install the Tailscale app first: https://tailscale.com/download/macos"
if ! prompt_yes_no "Continue once Tailscale is installed?" "n"; then
return 1
fi
fi
else
if ! prompt_yes_no "Tailscale is not installed. Install it now (official installer from tailscale.com)?" "y"; then
return 1
fi
info "Running the official Tailscale installer (it may ask for sudo)..."
# When piped (curl | bash), stdin is our pipe: give the child installer
# the real terminal so its own sudo prompt works.
if [[ -e /dev/tty ]]; then
if ! sh -c "$(download_to_stdout https://tailscale.com/install.sh)" < /dev/tty; then
warn "Tailscale installation failed."
return 1
fi
else
if ! sh -c "$(download_to_stdout https://tailscale.com/install.sh)"; then
warn "Tailscale installation failed."
return 1
fi
fi
fi
if ! check_tailscale; then
warn "tailscale was not found after the install."
return 1
fi
success "Tailscale installed"
return 0
}
ensure_tailscale_login() {
local state
state=$(ts_status_field 's.BackendState')
if [[ "$state" == "Running" ]]; then
return 0
fi
if [[ "$NONINTERACTIVE" == "1" ]] || ! has_tty; then
warn "Tailscale is installed but not connected (state: ${state:-unknown})."
return 1
fi
info "Tailscale needs to log in to your tailnet."
echo -e " ${DIM}A login URL will be printed: open it on any device. Waiting up to 5 minutes.${NC}"
local ts_bin up_ok="0"
ts_bin=$(get_tailscale_path) || return 1
if [[ "$(uname -s)" == "Darwin" ]]; then
# The GUI app's CLI runs as the user; no root needed.
if "$ts_bin" up --timeout=300s; then up_ok="1"; fi
else
if [[ -e /dev/tty ]]; then
if run_as_root "$ts_bin" up --timeout=300s < /dev/tty; then up_ok="1"; fi
else
if run_as_root "$ts_bin" up --timeout=300s; then up_ok="1"; fi
fi
fi
if [[ "$up_ok" != "1" ]]; then
if [[ "$(uname -s)" == "Darwin" ]]; then
info "If the CLI cannot log in, open the Tailscale app, log in there, then run:"
info " bash $INSTALL_DIR/install.sh tailscale"
fi
return 1
fi
[[ "$(ts_status_field 's.BackendState')" == "Running" ]]
}
# Linux: `tailscale serve` needs root or operator rights. Grant operator once
# (with the user's consent via sudo) so serve config never needs sudo again;
# fall back to sudo-per-command when the grant fails.
ensure_tailscale_operator() {
if [[ "$(uname -s)" == "Darwin" ]] || [[ $EUID -eq 0 ]]; then
return 0
fi
if ts_cmd serve status &>/dev/null; then
return 0
fi
if ! command -v sudo &>/dev/null; then
warn "No sudo available; tailscale serve configuration may fail without root."
TS_NEED_ROOT="1"
return 0
fi
info "Granting your user Tailscale 'operator' rights (one-time sudo; lets serve run without root)..."
local ts_bin
ts_bin=$(get_tailscale_path) || return 0
if run_as_root "$ts_bin" set --operator="$USER" 2>/dev/null && ts_cmd serve status &>/dev/null; then
success "Operator rights granted"
return 0
fi
warn "Could not grant operator rights; serve commands will use sudo."
TS_NEED_ROOT="1"
return 0
}
# HTTPS certificates are a per-tailnet admin toggle. Serve without them cannot
# terminate TLS, and a plain-HTTP fallback would silently break the "real
# HTTPS" promise (PWA install, web push), so guide the user through enabling
# them instead of degrading.
ensure_tailnet_https() {
while true; do
local magic cert
magic=$(ts_status_field 's.CurrentTailnet && s.CurrentTailnet.MagicDNSEnabled ? "1" : ""')
cert=$(ts_status_field 'Array.isArray(s.CertDomains) && s.CertDomains.length > 0 ? "1" : ""')
if [[ "$magic" == "1" && "$cert" == "1" ]]; then
return 0
fi
warn "Your tailnet has not enabled HTTPS certificates yet (a one-time admin toggle)."
echo -e " Open ${CYAN}https://login.tailscale.com/admin/dns${NC} and enable:" >&2
if [[ "$magic" == "1" ]]; then
echo -e " ${CYAN}1.${NC} MagicDNS ${GREEN}(already on)${NC}" >&2
else
echo -e " ${CYAN}1.${NC} MagicDNS" >&2
fi
if [[ "$cert" == "1" ]]; then
echo -e " ${CYAN}2.${NC} HTTPS Certificates ${GREEN}(already on)${NC}" >&2
else
echo -e " ${CYAN}2.${NC} HTTPS Certificates" >&2
fi
if [[ "$NONINTERACTIVE" == "1" ]] || ! has_tty; then
return 1
fi
if ! prompt_yes_no "Re-check now? (answering no skips Tailscale setup)" "y"; then
return 1
fi
done
}
setup_tailscale_serve() {
local port="${CODEMAN_PORT:-3000}"
local dns url existing
dns=$(ts_status_field 's.Self && s.Self.DNSName')
if [[ -z "$dns" ]]; then
warn "Could not determine this machine's tailnet DNS name."
return 1
fi
url="https://${dns%.}"
existing=$(ts_serve_443_target_port)
if [[ "$existing" == "$port" ]]; then
TAILSCALE_SERVE_URL="$url"
success "Tailscale serve already forwards $url to port $port (kept as-is)"
return 0
fi
if [[ -n "$existing" ]]; then
warn "tailscale serve already forwards $url (port 443) to local port $existing."
if ! prompt_yes_no "Replace that mapping with Codeman (port $port)?" "n"; then
info "Keeping the existing mapping."
return 1
fi
fi
info "Configuring: tailscale serve --bg $port"
local serve_out
if serve_out=$(ts_cmd_serve serve --bg "$port" 2>&1); then
TAILSCALE_SERVE_URL="$url"
success "Tailscale HTTPS enabled: $url"
echo -e " ${DIM}(persists across reboots; inspect with: tailscale serve status)${NC}"
return 0
fi
warn "tailscale serve failed:"
printf '%s\n' "$serve_out" | sed 's/^/ /' >&2
return 1
}
# Curl the ts.net URL until it answers. 200 = reachable; 401 = reachable behind
# the dashboard password. The first request can be slow while tailscaled
# obtains the Let's Encrypt certificate.
verify_tailscale_access() {
if [[ -z "$TAILSCALE_SERVE_URL" ]]; then
return 0
fi
if ! command -v curl &>/dev/null; then
info "curl not available; open $TAILSCALE_SERVE_URL to verify."
return 0
fi
info "Verifying $TAILSCALE_SERVE_URL (first load can take ~30s while the HTTPS certificate is issued)..."
local i http_code
for ((i = 1; i <= 10; i++)); do
http_code=$(curl -skm 10 -o /dev/null -w '%{http_code}' "$TAILSCALE_SERVE_URL/api/status" 2>/dev/null) || http_code=""
if [[ "$http_code" == "200" || "$http_code" == "401" ]]; then
success "Reachable: $TAILSCALE_SERVE_URL"
return 0
fi
sleep 3
done
warn "Could not reach $TAILSCALE_SERVE_URL/api/status yet."
warn "It may need another minute (certificate issuance). Inspect: tailscale serve status"
warn "If Codeman itself runs with --https, the serve target must be:"
warn " tailscale serve --bg https+insecure://localhost:${CODEMAN_PORT:-3000}"
return 1
}
# Orchestrator: walk every state (not installed -> logged out -> operator ->
# tailnet HTTPS -> serve) and end with TAILSCALE_SERVE_URL set, or fall back
# gracefully (the caller keeps the loopback bind either way).
setup_tailscale_access() {
TAILSCALE_SERVE_URL=""
if ! check_tailscale; then
if ! offer_install_tailscale; then
tailscale_retrofit_hint "Tailscale is not installed"
return 1
fi
fi
if ! command -v node &>/dev/null; then
tailscale_retrofit_hint "node is not on PATH yet"
return 1
fi
if ! ensure_tailscale_login; then
tailscale_retrofit_hint "Tailscale is not connected"
return 1
fi
ensure_tailscale_operator
if ! ensure_tailnet_https; then
tailscale_retrofit_hint "HTTPS certificates are not enabled for your tailnet"
return 1
fi
if ! setup_tailscale_serve; then
tailscale_retrofit_hint "tailscale serve could not be configured"
return 1
fi
return 0
}
# `install.sh tailscale`: retrofit Tailscale access onto an existing install
# (also the target of every "set it up later" hint above).
setup_tailscale_subcommand() {
print_banner
if ! command -v node &>/dev/null; then
die "node is required. Install Codeman first (run the installer without arguments)."
fi
read_existing_binding
if [[ "$EXISTING_FOUND" == "1" && -n "$EXISTING_HOST" && "$EXISTING_HOST" != "127.0.0.1" ]]; then
warn "Your service binds $EXISTING_HOST (network-wide). Tailscale serve will work, but the"
warn "dashboard stays reachable on your LAN too. Re-run the installer and choose Tailscale"
warn "to switch to the tighter loopback-only bind."
echo ""
fi
if ! setup_tailscale_access; then
exit 1
fi
# Verify end-to-end only when Codeman is actually answering locally.
local port="${CODEMAN_PORT:-3000}" server_up="0"
if command -v curl &>/dev/null; then
if curl -skm 5 -o /dev/null "http://127.0.0.1:$port/api/status" 2>/dev/null ||
curl -skm 5 -o /dev/null "https://127.0.0.1:$port/api/status" 2>/dev/null; then
server_up="1"
fi
fi
if [[ "$server_up" == "1" ]]; then
verify_tailscale_access || true
else
info "Codeman does not appear to be running on port $port right now."
info "Once it is, open: $TAILSCALE_SERVE_URL"
fi
BIND_HOST="${EXISTING_HOST:-127.0.0.1}"
BIND_PASSWORD="$EXISTING_PASSWORD"
print_security_notice
}
# ============================================================================
# Service Setup (Linux systemd / macOS launchd)
# ============================================================================
@@ -1487,11 +2029,12 @@ main() {
fi
fi
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini)
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini, Antigravity)
local has_claude=false
local has_opencode=false
local has_codex=false
local has_gemini=false
local has_antigravity=false
info "Checking AI CLI tools..."
if check_claude; then
@@ -1510,17 +2053,21 @@ main() {
has_gemini=true
success "Gemini CLI found at $(get_gemini_path)"
fi
if check_antigravity; then
has_antigravity=true
success "Antigravity CLI found at $(get_antigravity_path)"
fi
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" ]]; then
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" ]]; then
echo ""
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, or Gemini."
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, or Gemini."
headless_guard "install an AI CLI (curl | bash from its vendor)"
echo ""
echo -e " ${BOLD}Which AI CLI would you like to install?${NC}"
echo -e " ${CYAN}1)${NC} Claude Code (Anthropic)"
echo -e " ${CYAN}2)${NC} OpenCode (open-source)"
echo -e " ${CYAN}3)${NC} Both"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Gemini)"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Antigravity)"
echo ""
local cli_choice=""
@@ -1565,8 +2112,8 @@ main() {
if [[ "$cli_choice" == "4" ]]; then
warn "Skipping AI CLI install. Codeman will run, but sessions need a CLI to drive."
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: npm install -g @google/gemini-cli (Gemini)"
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity)"
elif [[ "$has_claude" == "false" ]] && [[ "$has_opencode" == "false" ]]; then
die "The selected AI CLI failed to install. Install one manually and re-run the installer."
fi
@@ -1773,10 +2320,19 @@ main() {
echo ""
if [[ "$service_ok" == "true" ]]; then
# With Tailscale configured, prove the URL actually answers now
# that the server is up (never claim success blindly).
if [[ -n "$TAILSCALE_SERVE_URL" ]]; then
verify_tailscale_access || true
echo ""
fi
echo -e " ${GREEN}${BOLD}Codeman is running now!${NC}"
echo ""
echo -e " ${CYAN}# Open in browser${NC}"
if [[ "$BIND_HOST" == "0.0.0.0" ]]; then
if [[ -n "$TAILSCALE_SERVE_URL" ]]; then
echo -e " $TAILSCALE_SERVE_URL ${DIM}(any device on your tailnet, HTTPS)${NC}"
echo -e " http://localhost:3000 ${DIM}(this machine)${NC}"
elif [[ "$BIND_HOST" == "0.0.0.0" ]]; then
echo -e " http://$(detect_lan_ip):3000 ${DIM}(any device on your network)${NC}"
echo -e " http://localhost:3000 ${DIM}(this machine)${NC}"
else
@@ -1822,10 +2378,21 @@ main() {
echo ""
echo -e " ${CYAN}# Open in browser${NC}"
echo -e " http://localhost:3000"
if [[ -n "$TAILSCALE_SERVE_URL" ]]; then
echo -e " $TAILSCALE_SERVE_URL ${DIM}(any device on your tailnet, once running)${NC}"
fi
fi
echo ""
fi
if [[ -n "$TAILSCALE_SERVE_URL" ]]; then
echo -e " ${BOLD}Remote Access (Tailscale):${NC}"
echo ""
echo -e " $TAILSCALE_SERVE_URL ${DIM}(HTTPS, any device on your tailnet)${NC}"
echo -e " ${CYAN}tailscale serve status${NC} # Inspect the mapping"
echo ""
fi
if check_cloudflared; then
echo -e " ${BOLD}Remote Access (Cloudflare Tunnel):${NC}"
echo ""
@@ -1846,12 +2413,12 @@ main() {
echo -e " https://github.com/Ark0N/Codeman"
echo ""
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini; then
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity; then
echo -e " ${YELLOW}${BOLD}Reminder:${NC} Install at least one AI CLI to start using Codeman:"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}npm install -g @google/gemini-cli${NC} # Gemini"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}curl -fsSL https://antigravity.google/cli/install.sh | bash${NC} # Antigravity"
echo ""
fi
@@ -1982,6 +2549,20 @@ uninstall() {
success "Removed LaunchDaemon"
fi
# Remove OUR tailscale serve mapping (443 -> Codeman's port) only. Other
# serve config stays untouched, and never `tailscale serve reset`.
local ts_url=""
ts_url=$(detect_tailscale_serve_url 2>/dev/null) || ts_url=""
if [[ -n "$ts_url" ]]; then
if prompt_yes_no "Remove the Tailscale serve mapping for Codeman ($ts_url)?" "y"; then
if ts_cmd_serve serve --https=443 off 2>/dev/null; then
success "Removed tailscale serve mapping"
else
warn "Could not remove it automatically. Run: tailscale serve --https=443 off"
fi
fi
fi
# Remove symlinks
local symlink_dir="$HOME/.local/bin"
if [[ -L "$symlink_dir/codeman" ]]; then
@@ -2030,6 +2611,7 @@ uninstall() {
case "${1:-}" in
update) update ;;
uninstall) uninstall ;;
tailscale) setup_tailscale_subcommand ;;
*)
# Only a COMPLETED install re-runs as a quiet update. A partial one
# (clone succeeded but build/menu never finished) lacks the marker and
+3 -3
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.9.5",
"version": "1.13.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.9.5",
"version": "1.13.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -12333,7 +12333,7 @@
}
},
"packages/xterm-zerolag-input": {
"version": "0.1.7",
"version": "0.1.8",
"license": "MIT",
"devDependencies": {
"jsdom": "^24.1.3",
+5 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.9.5",
"version": "1.13.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -22,6 +22,7 @@
"test:coverage": "vitest run --config config/vitest.config.ts --coverage",
"test:ci": "vitest run --config config/vitest.ci.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
"lint:fix": "eslint --config config/eslint.config.js 'src/**/*.ts' --fix",
@@ -54,6 +55,7 @@
"anthropic",
"opencode",
"codex",
"antigravity",
"gemini-cli",
"ai-agents",
"agent",
@@ -156,6 +158,8 @@
"files": [
"dist",
"scripts/postinstall.js",
"scripts/fix-node-pty.mjs",
"skills",
"LICENSE",
"README.md"
]
+29
View File
@@ -1,5 +1,34 @@
# xterm-zerolag-input
## 0.1.8
### Patch Changes
- **Fixed: sessions failed to start on macOS with `Error: posix_spawnp failed.`** (issues #6 and #204)
`node-pty@1.1.0` publishes its macOS prebuilt helper as `prebuilds/darwin-<arch>/spawn-helper` with mode 0644, i.e. no execute bit. macOS launches every PTY through that helper, so a stock install failed on every session start. The bug is macOS-only: `spawn-helper` is a mac-only gyp target and node-pty ships no Linux prebuild, so Linux always compiles a correctly-permissioned helper from source.
The previous fix chmodded only `build/Release/spawn-helper`, which on macOS does not exist (the prebuild is used, so node-gyp never runs), and it derived that path from `require.resolve('node-pty')`, landing on `<pkg>/lib/build/Release/...`. It was a no-op on every platform.
- New `scripts/fix-node-pty.mjs` (also `npm run fix:node-pty`) chmods every `spawn-helper` it finds, in `build/Release`, `build/Debug` and each `prebuilds/*/`, then verifies the result by actually opening a PTY. A `require()` alone passes on a broken install, because the helper is only touched at spawn time.
- `postinstall` no longer force-rebuilds node-pty from source on Node 22+. That step needed Xcode command line tools, cost 30-120s on every install, and deleted the `prebuilds/` tree before compiling, so a Mac without a compiler was left with no working binary at all. A rebuild now happens only when the chmod plus spawn probe still fails, and the prebuilds tree is backed up and restored around it.
- New `spawnPtyWithHelperRepair()` (`src/utils/node-pty-repair.ts`) wraps every `pty.spawn()` in `session.ts`, so an install that is already broken repairs itself on the first failed spawn and retries in-process instead of showing a dead session. Unrelated spawn errors are rethrown untouched; a second failure carries the `npm run fix:node-pty` hint.
- `scripts/fix-node-pty.mjs` is now in the published `files` list, so global npm installs get the repair too.
- Direct-PTY Claude spawns use the resolved absolute binary path (new `getClaudeBinaryPath()`) instead of the bare name `claude`, so a CLI installed outside the server's PATH still launches.
Verified end to end on macOS 26.4 arm64: a stock `npm i` reproduces `posix_spawnp failed.`, and after the fix the same install spawns a PTY successfully with the prebuilds preserved.
**Added: phone home screen (session overview)**
Under 430px the "C" logo now opens a session overview (current sessions, past sessions, spaces) instead of the welcome overlay: on a small screen "which session needs me" beats "how do I start one". Rows resume a session in place, and "New session here" goes through the normal quick-start path so remote and Docker cases keep their routing. Per-device setting `mobileOverviewEnabled` (phones only, default ON) in App Settings. Tablet and desktop are unchanged.
**Added: guided Tailscale setup in `install.sh`**
The network-access prompt is now 3-way: Tailscale, LAN, or local-only. The Tailscale path binds loopback and walks through installing Tailscale, logging in, the operator grant, the tailnet HTTPS-certificates toggle, and `tailscale serve --bg <port>`, then verifies the result end to end with curl. That gives HTTPS on a real certificate with no app password and no `0.0.0.0` bind, which is also what PWA install and web push need. `install.sh tailscale` retrofits it onto an existing install, and `CODEMAN_TAILSCALE=1` presets the choice. Serve state is detected from `tailscale serve status --json`; the installer never runs `tailscale serve reset` and never touches serve mappings other than 443 to Codeman's port. README and `docs/security-architecture.md` updated to match.
**Docs**: replaced a real tailnet hostname with placeholders in `docs/web-tabs-fixes-plan.md`.
**xterm-zerolag-input**: npm description and keywords only, no code change.
## 0.1.7
### Patch Changes
+1 -1
View File
@@ -16,7 +16,7 @@
> ### Made for [**Codeman**](https://getcodeman.com)
>
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Gemini sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Antigravity sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
>
> That last part is why this library exists. The demo below is a real Codeman session on two phones.
+10 -2
View File
@@ -1,7 +1,7 @@
{
"name": "xterm-zerolag-input",
"version": "0.1.7",
"description": "Instant keystroke feedback overlay for xterm.js — eliminates perceived input latency over high-RTT connections",
"version": "0.1.8",
"description": "Instant keystroke feedback overlay for xterm.js: Mosh-inspired local echo that removes perceived input latency over SSH, tunnels and other high-RTT connections",
"type": "module",
"main": "dist/index.cjs",
"module": "dist/index.js",
@@ -26,8 +26,16 @@
"xterm",
"xterm.js",
"terminal",
"web-terminal",
"local-echo",
"local echo",
"mosh",
"input-latency",
"latency",
"zero-lag",
"keystroke",
"ssh",
"remote-terminal",
"overlay",
"addon"
],
+261
View File
@@ -0,0 +1,261 @@
#!/usr/bin/env node
/**
* @fileoverview Repairs node-pty's macOS `spawn-helper` and verifies that a PTY
* can really be spawned. Called by `scripts/postinstall.js` on every install and
* exposed as `npm run fix:node-pty` for repairing an install after the fact.
*
* Why this exists (issues #6 and #204):
*
* node-pty@1.1.0 publishes its macOS prebuilt helper as
* `prebuilds/darwin-<arch>/spawn-helper` with mode 0644, i.e. no execute bit.
* On macOS node-pty launches every PTY through that helper with posix_spawnp,
* which then fails EACCES and surfaces as `Error: posix_spawnp failed.` on every
* session start.
*
* It is macOS-exclusive twice over: `spawn-helper` is an `OS=="mac"` gyp target,
* and pty.cc only spawns it under `#if defined(__APPLE__)`. node-pty ships
* prebuilds for darwin and win32 only, so Linux always compiles from source
* (which produces an executable helper) and never sees the bug.
*
* The repair is a chmod, NOT a rebuild: the prebuilt binary itself is fine, and
* requiring a from-source rebuild would make every macOS install depend on Xcode
* command line tools. A rebuild is attempted only when a chmod plus a real spawn
* probe still can't get a working PTY, and the prebuilds tree is backed up first
* so a failed rebuild can never leave the install worse than it started.
*/
import { chmodSync, cpSync, existsSync, readdirSync, rmSync, statSync } from 'node:fs';
import { execSync } from 'node:child_process';
import { createRequire } from 'node:module';
import { tmpdir } from 'node:os';
import { dirname, join } from 'node:path';
import { fileURLToPath } from 'node:url';
const require = createRequire(import.meta.url);
/** Errors that mean "the native module or its helper is unusable", i.e. worth a rebuild. */
const NATIVE_FAILURE_PATTERN = /posix_spawnp|spawn-helper|Failed to load native module|Cannot find module/i;
/**
* Locates the installed node-pty package directory.
*
* @returns {string|null} Absolute path to the package root, or null if not installed.
*/
export function findNodePtyDir() {
// package.json first: node-pty declares no "exports" map, so the subpath resolves,
// and it lands on the package root directly. require.resolve('node-pty') would give
// <pkg>/lib/index.js, which is one directory deeper than callers expect.
try {
return dirname(require.resolve('node-pty/package.json'));
} catch {
/* fall through */
}
try {
return join(dirname(require.resolve('node-pty')), '..');
} catch {
return null;
}
}
/**
* Lists every `spawn-helper` shipped in a node-pty install.
*
* node-pty's own loader (lib/utils.js) checks `build/Release`, `build/Debug` and
* then `prebuilds/<platform>-<arch>`, and takes the helper from whichever
* directory the native module loaded out of, so all of them must be executable,
* not just the one this machine happens to use today.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @returns {string[]} Absolute paths of the helpers that exist on disk.
*/
export function listSpawnHelpers(ptyDir) {
const dirs = [join(ptyDir, 'build', 'Release'), join(ptyDir, 'build', 'Debug')];
const prebuilds = join(ptyDir, 'prebuilds');
if (existsSync(prebuilds)) {
try {
for (const entry of readdirSync(prebuilds, { withFileTypes: true })) {
if (entry.isDirectory()) dirs.push(join(prebuilds, entry.name));
}
} catch {
/* unreadable prebuilds dir: nothing to repair there */
}
}
return dirs.map((d) => join(d, 'spawn-helper')).filter((p) => existsSync(p));
}
/**
* Adds the execute bit to every `spawn-helper` that is missing it.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @returns {{ repaired: string[], failed: Array<{ path: string, error: string }> }}
*/
export function repairSpawnHelpers(ptyDir) {
const repaired = [];
const failed = [];
for (const helper of listSpawnHelpers(ptyDir)) {
try {
const mode = statSync(helper).mode & 0o777;
if ((mode & 0o111) === 0o111) continue; // already executable by all
chmodSync(helper, mode | 0o755);
repaired.push(helper);
} catch (err) {
failed.push({ path: helper, error: err instanceof Error ? err.message : String(err) });
}
}
return { repaired, failed };
}
/**
* Proves node-pty works by actually opening a PTY, which is the only check that
* exercises the spawn-helper path that breaks. A `require` alone would pass on a
* broken install, because the helper is only touched at spawn time.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @returns {{ ok: boolean, error?: string, nativeFailure?: boolean }}
*/
export function verifyPtySpawn(ptyDir) {
let child;
try {
const pty = require(ptyDir); // directory require → node-pty's "main" (lib/index.js)
const file = process.platform === 'win32' ? process.env.COMSPEC || 'cmd.exe' : '/bin/echo';
const args = process.platform === 'win32' ? ['/c', 'exit'] : ['codeman-node-pty-check'];
child = pty.spawn(file, args, {
name: 'xterm-color',
cols: 80,
rows: 24,
cwd: tmpdir(),
env: process.env,
});
return { ok: true };
} catch (err) {
const message = err instanceof Error ? err.message : String(err);
return { ok: false, error: message, nativeFailure: NATIVE_FAILURE_PATTERN.test(message) };
} finally {
try {
child?.kill();
} catch {
/* the probe child exits on its own anyway */
}
}
}
/**
* Rebuilds node-pty from source, preserving the prebuilds tree across a failure.
*
* node-pty's install script deletes `prebuilds/` as soon as
* `npm_config_build_from_source` is set and only then shells out to node-gyp, so
* a machine without a compiler toolchain would otherwise be left with neither a
* prebuilt nor a compiled binary.
*
* @param {string} ptyDir Absolute path to the node-pty package root.
* @param {string} cwd Directory to run npm from (the package root that owns node_modules).
* @returns {{ ok: boolean, error?: string }}
*/
function rebuildFromSource(ptyDir, cwd) {
const prebuilds = join(ptyDir, 'prebuilds');
const backup = join(ptyDir, '.prebuilds-codeman-backup');
let backedUp = false;
if (existsSync(prebuilds)) {
try {
rmSync(backup, { recursive: true, force: true });
cpSync(prebuilds, backup, { recursive: true });
backedUp = true;
} catch {
/* best effort: proceed without a safety net rather than skip the repair */
}
}
try {
execSync('npm rebuild node-pty --build-from-source', { cwd, stdio: 'pipe', timeout: 300000 });
return { ok: true };
} catch (err) {
if (backedUp && !existsSync(prebuilds)) {
try {
cpSync(backup, prebuilds, { recursive: true });
} catch {
/* nothing further we can do */
}
}
return { ok: false, error: err instanceof Error ? err.message : String(err) };
} finally {
rmSync(backup, { recursive: true, force: true });
}
}
/**
* Full repair flow: chmod, verify, and only rebuild if a working PTY still can't
* be opened.
*
* @param {object} [options]
* @param {(line: string) => void} [options.log] Progress sink (default: silent).
* @param {(line: string) => void} [options.warn] Warning sink (default: same as log).
* @param {boolean} [options.allowRebuild] Permit a from-source rebuild (default: true).
* @returns {Promise<{ ok: boolean, repaired: string[], rebuilt: boolean, reason?: string }>}
*/
export async function fixNodePty(options = {}) {
const log = options.log ?? (() => {});
const warn = options.warn ?? log;
const allowRebuild = options.allowRebuild ?? true;
const ptyDir = findNodePtyDir();
if (!ptyDir) {
return { ok: false, repaired: [], rebuilt: false, reason: 'node-pty is not installed' };
}
const { repaired, failed } = repairSpawnHelpers(ptyDir);
for (const f of failed) warn(`could not chmod ${f.path}: ${f.error}`);
if (repaired.length > 0) {
log(`made node-pty spawn-helper executable (${repaired.length} file${repaired.length === 1 ? '' : 's'})`);
}
const first = verifyPtySpawn(ptyDir);
if (first.ok) return { ok: true, repaired, rebuilt: false };
if (!allowRebuild || !first.nativeFailure) {
return { ok: false, repaired, rebuilt: false, reason: first.error };
}
warn(`node-pty could not open a PTY (${first.error}), rebuilding from source...`);
const projectRoot = join(dirname(fileURLToPath(import.meta.url)), '..');
const rebuild = rebuildFromSource(ptyDir, projectRoot);
if (!rebuild.ok) {
return { ok: false, repaired, rebuilt: false, reason: `rebuild failed: ${rebuild.error}` };
}
const after = repairSpawnHelpers(ptyDir);
repaired.push(...after.repaired);
const second = verifyPtySpawn(ptyDir);
return second.ok
? { ok: true, repaired, rebuilt: true }
: { ok: false, repaired, rebuilt: true, reason: second.error };
}
// ---------------------------------------------------------------------------
// CLI: node scripts/fix-node-pty.mjs [--quiet]
// ---------------------------------------------------------------------------
const isDirectRun = process.argv[1] && fileURLToPath(import.meta.url) === process.argv[1];
if (isDirectRun) {
const quiet = process.argv.includes('--quiet');
const say = (line) => {
if (!quiet) console.log(line);
};
const result = await fixNodePty({ log: say, warn: (line) => console.warn(line) });
if (result.ok) {
say(result.repaired.length > 0 || result.rebuilt ? 'node-pty repaired, PTY spawning works' : 'node-pty is healthy');
process.exit(0);
}
console.error(`node-pty is not usable: ${result.reason}`);
console.error('Try: cd node_modules/node-pty && npx node-gyp rebuild');
process.exit(1);
}
+21 -24
View File
@@ -6,7 +6,7 @@
*/
import { execSync, spawn } from 'child_process';
import { chmodSync, existsSync } from 'fs';
import { existsSync } from 'fs';
import { homedir, platform } from 'os';
import { join } from 'path';
import { createRequire } from 'module';
@@ -148,35 +148,32 @@ if (majorVersion < MIN_NODE_VERSION) {
}
// ----------------------------------------------------------------------------
// 1b. Fix node-pty spawn-helper permissions (macOS posix_spawnp fix)
// 1b. Repair + verify node-pty (macOS posix_spawnp fix, issues #6 and #204)
//
// node-pty ships its macOS spawn-helper without the execute bit, which breaks
// every session start on macOS. fixNodePty() chmods it, then proves a PTY can
// actually be opened, and only falls back to a from-source rebuild if that
// still fails. See scripts/fix-node-pty.mjs for the full story.
// ----------------------------------------------------------------------------
try {
const require = createRequire(import.meta.url);
const ptyPath = join(require.resolve('node-pty'), '..');
const spawnHelper = join(ptyPath, 'build', 'Release', 'spawn-helper');
if (existsSync(spawnHelper)) {
chmodSync(spawnHelper, 0o755);
console.log(colors.green('✓ node-pty spawn-helper permissions fixed'));
}
} catch {
// Non-critical — only affects macOS with prebuilt binaries
}
const { fixNodePty } = await import('./fix-node-pty.mjs');
const result = await fixNodePty({
log: (line) => console.log(colors.dim(` ${line}`)),
warn: (line) => console.log(colors.yellow(`⚠ ${line}`)),
});
// ----------------------------------------------------------------------------
// 1c. Rebuild node-pty from source for Node.js 22+ compatibility
// ----------------------------------------------------------------------------
if (majorVersion >= 22) {
try {
console.log(colors.dim(' Rebuilding node-pty from source for Node.js 22+...'));
execSync('npm rebuild node-pty --build-from-source', { stdio: 'pipe', timeout: 120000 });
console.log(colors.green('✓ node-pty rebuilt from source'));
} catch {
if (result.ok) {
console.log(colors.green('✓ node-pty verified') + colors.dim(' (PTY spawn works)'));
} else {
hasWarnings = true;
console.log(colors.yellow('⚠ Failed to rebuild node-pty from source'));
console.log(colors.dim(' You may need to run: npm rebuild node-pty --build-from-source'));
console.log(colors.yellow(`⚠ node-pty is not usable: ${result.reason}`));
console.log(colors.dim(' Sessions will fail to start. Try: ') + colors.cyan('npm run fix:node-pty'));
}
} catch (err) {
hasWarnings = true;
console.log(colors.yellow(`⚠ Could not verify node-pty: ${err.message}`));
console.log(colors.dim(' If sessions fail to start, run: ') + colors.cyan('npm run fix:node-pty'));
}
// ----------------------------------------------------------------------------
+274
View File
@@ -0,0 +1,274 @@
---
name: codeman
description: >-
Drive Codeman, the session manager this agent is running inside, over its HTTP API:
list sessions, start worker sessions, send them prompts, block until they finish
(wait / wait-output / send-and-wait), read their output, and clean up. Use when asked
to orchestrate or parallelize work across Codeman sessions, watch another session, or
start and manage workers. Only usable inside a Codeman-managed session
(CODEMAN_MUX=1); refuse to act otherwise.
---
# Driving Codeman from inside a session
You are an agent running inside a Codeman-managed terminal session. Codeman is the
server that spawned you; its HTTP API can start, prompt, watch, and delete other
sessions. Every recipe below was verified live. Full endpoint tables and
troubleshooting: [reference/endpoints.md](reference/endpoints.md). Worked multi-worker
flows: [reference/recipes.md](reference/recipes.md).
## 0. Guard — run this before anything else
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "Not inside a Codeman-managed session; refusing to act."; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Codeman does NOT hand a session the server password. If one is set, the two
# in-reach copies are the data dir's .env (the same fallback `codeman attach`
# uses — hand-authored; nothing ever writes it) and the supervisor definition
# that install.sh wrote the password into, which is where a stock
# password-protected install actually keeps it. The data dir is wherever the
# hook-secret file lives. Values may be quoted or `export`-prefixed.
ENV_FILE="${CODEMAN_HOOK_SECRET_FILE:+${CODEMAN_HOOK_SECRET_FILE%hook-secret}.env}"
envval() { sed -n "s/^\(export \)\{0,1\}$1=//p" "$ENV_FILE" | tail -1 | sed 's/^"\(.*\)"$/\1/; s/^'\''\(.*\)'\''$/\1/'; }
if [ -z "${CODEMAN_PASSWORD:-}" ] && [ -n "$ENV_FILE" ] && [ -f "$ENV_FILE" ]; then
CODEMAN_USERNAME=$(envval CODEMAN_USERNAME)
CODEMAN_PASSWORD=$(envval CODEMAN_PASSWORD)
fi
if [ -z "${CODEMAN_PASSWORD:-}" ]; then # stock installs: install.sh puts it in the service definition
UNIT="$HOME/.config/systemd/user/codeman-web.service"
PLIST="$HOME/Library/LaunchAgents/com.codeman.web.plist"
if [ -f "$UNIT" ]; then
CODEMAN_PASSWORD=$(sed -n 's/^Environment="CODEMAN_PASSWORD=\(.*\)"$/\1/p' "$UNIT" | head -1)
elif [ -f "$PLIST" ]; then
CODEMAN_PASSWORD=$(awk '/<key>CODEMAN_PASSWORD<\/key>/{getline; print}' "$PLIST" | sed -n 's/.*<string>\(.*\)<\/string>.*/\1/p')
fi
fi
AUTH=(); [ -n "${CODEMAN_PASSWORD:-}" ] && AUTH=(-u "${CODEMAN_USERNAME:-admin}:$CODEMAN_PASSWORD")
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on https (self-signed cert)
```
- If `CODEMAN_MUX` is not `1`, **stop and say so**. Do not guess an API URL; a server
you are not part of is not yours to drive.
- **A 401 is plain text, not the JSON envelope**, so on a password-protected server
every `jq` in these recipes dies with `jq: parse error` instead of showing
`UNAUTHORIZED`. If that happens, check the status with `-w '%{http_code}'`; if it
is 401 and neither fallback above found a credential, **stop and tell the user
you need credentials**. The hook-secret bypass covers only `/api/hook-event` and
`/api/status-telemetry`, never session control.
- These endpoints first ship in Codeman **1.13.0**, but do not gate on the version
number: a dev build can serve them while reporting an older version. Probe
instead: `GET .../wait` on a real session id answering 404 with an `.error`
starting `Route ` means the server predates the wait endpoints (fall back to
polling `GET .../terminal?tail=` and say so); `Session ... not found` means your
session id is wrong, not the server.
## 1. Safety rules — read before any mutating call
You are yourself a session on this server, and the API has **no undo**.
- **Never act on your own session — and know that this check is the ONLY guard.**
The server has no self-protection: a session that DELETEs its own id succeeds and
dies silently (verified live). Session ids appear in both full and 8-character
forms (Docker cases export a truncated `$SELF`; mux names and UI surfaces carry
8-char ids), so compare by prefix **in both directions**, never by equality:
```bash
is_self() { case "$1" in "$SELF"*) return 0 ;; esac; case "$SELF" in "$1"*) return 0 ;; esac; return 1; }
```
One-directional or equality checks each miss a real combination (full `$SELF` vs
a target you transcribed in 8-char form, or truncated `$SELF` vs a full target)
and the miss deletes you. Check `is_self` before every `DELETE`, kill, respawn,
or input call.
- **Mutating calls you may make unprompted** (this is an allowlist):
`POST /api/v1/quick-start`, `POST /api/v1/sessions/:id/input`, and
`DELETE /api/v1/sessions/:id` **only** for a session you created in this
conversation, by exact id. Keep a list of the ids you create. Everything else
mutating needs the user to have asked for it.
- **Never call these** unless the user explicitly asked, naming the target:
- `DELETE /api/cases/:name` — recursively **deletes a real directory of the user's
code** from disk. One wrong case name destroys work that was never yours.
- `DELETE /api/sessions` (no id) and `DELETE /api/subagents` (no id) — bulk kills.
- respawn / ralph / orchestrator / cron mutations — respawn runs `/clear` (wipes a
conversation), orchestrator state is a single global slot, cron jobs outlive you.
- `PUT /api/settings`, `POST /api/system/update` — global UI settings; server restart.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. The API is the only interface.
- Sessions count against a 50-session cap and case creation is uncapped: clean up every
session you start, and don't retry `quick-start` in a loop.
## 2. Rules of the road
- **End every input with `\r`** — literally the two characters `\r` inside the JSON
string. Codeman types the text and sends Enter **only when the input contains a
carriage return**; without it your command sits unsubmitted on the worker's prompt
and everything downstream times out. `{"input":"run the tests\r",...}`. No response
field catches this: `delivered:true` means "written to the pane", **not**
"submitted" — a `\r`-less send still reports `delivered:true` and then every wait
times out, which is why the loops below are bounded and check the terminal.
- **Single-line input only.** Newlines are stripped; one line per call.
- **Build request bodies with `jq -n` for any prompt you did not author as a
literal.** The inline `-d '{"input":"'"$P"'\r"}'` pattern breaks on the first
double quote, backslash, or `$` in a real prompt:
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
- **Exactly-once delivery**: always send a stable `clientId` and a monotonic
per-session `seq` on `POST .../input`. A retry after a dropped connection then
cannot double-type the prompt. Increment `seq` for each NEW input; reuse the same
pair only to re-ask about the same delivery.
- **Envelope**: success is `{"success":true,"data":…}`, errors are
`{"success":false,"error","errorCode"}`. Read `.data`. Use `/api/v1/*` paths.
- **A wait timeout is HTTP 200**, `{wait:{timedOut:true,signal:null}}` — not an error.
Loop over short waits (60 s); proxies cut long-idle connections. Timeouts are
**clamped** (ceiling 600 s): read back `wait.timeoutMs` for what was applied.
- **`stop` and `blocked` fire for `claude` sessions only** (Claude Code hooks). On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`, requesting them explicitly is a
400 — and lifecycle transitions there are coarse (a short shell command may emit
**no** `idle` transition at all, verified live), so synchronize those modes with
output markers, not signals.
- **Your typed command echoes into the output stream**, so a marker that appears
verbatim in the input line matches **before the command runs**. Always split the
marker (recipe below), keep it unique per call, and use `from=buffer` so a marker
that printed before your wait landed is still found. Matching is literal — no regex.
- **Match single space-free tokens against TUI output.** A full-screen TUI (claude,
codex, …) positions text with cursor movements, not literal spaces, so the stripped
stream can read `Yes,Itrustthisfolder` and a multi-word match is unreliable there —
whether a phrase keeps its spaces depends on how the TUI happened to draw it
(observed live: some match, some never fire). Plain command output (shell workers,
`echo` lines) keeps real spaces.
## 3. Recipes (each verified live)
**List sessions / find yourself** — metadata only, safe to poll:
```bash
"${CURL[@]}" "$API/api/v1/sessions" | jq '.data[] | {id, name, mode, status}'
"${CURL[@]}" "$API/api/v1/sessions" | jq --arg s "$SELF" '.data[] | select(.id | startswith($s))'
```
**Start a claude worker and wait until it is actually ready.** A new session reports
`idle` before its CLI has spawned, and a brand-new case shows a **trust dialog**
first, so neither "wait for idle" nor "wait for ❯" means ready (the trust dialog
contains `❯` too — observed live). Codeman *can* auto-accept that dialog itself, but
the accept rides a stream match that misses on some runs (both outcomes seen live),
so wait for the composer first and handle the dialog only as the bounded fallback —
never send a blind Enter up front (if auto-accept already fired, it lands in the
composer). Stage 1 is short on purpose: an already-trusted case matches `bypass` in
under a second, while a **virgin case can never pass stage 1** (the dialog is up, so
the composer is not) and always pays it in full before the fallback runs — the long
budget belongs to stage 3, after the dialog is answered:
```bash
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
for _ in $(seq 1 30); do # bounded: a bad SID would otherwise poll forever
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# ⚠️ pid != null proves STARTUP only, never life: a worker that later dies inside
# its pane keeps status "idle" and a pid (the local tmux attach client, not the
# worker). The death check is wait?until=exit, below.
CID="agent-$$"; SEQ=1
# the composer's status bar ("bypass permissions on") is the ready marker — Codeman
# spawns claude in bypass mode. Single-token matches only: TUI text is space-less.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# composer never appeared → the trust dialog is probably still up; accept it once
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null || \
{ echo "worker $SID never became ready; inspect terminal?tail="; }
fi
```
**Send a prompt and wait for the turn to finish** (claude workers — the call to
prefer). It registers the waiter *before* typing, closing the race where a separate
wait sees the previous turn's idle state. Loop by resending the **identical** request:
the repeat is a tagged duplicate (same `clientId`+`seq`) that does not retype but
answers from the session's current state. Verified: the stop hook resolves this in
seconds; a duplicate resend answers in ~20 ms without retyping.
```bash
for TRY in $(seq 1 10); do # BOUNDED: a \r-less send never produces a signal and resends are no-op duplicates
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"run the tests, then summarize in one line\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ',"wait":true,"waitTimeout":60000}')
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # two straight timeouts: prompt sitting unsubmitted?
continue
fi
# Resolved — but a duplicate answering immediately reports the session's CURRENT
# state ("it is idle now"), NOT that a new turn ran. A \r-less send lands exactly
# here on try 2 (verified live), so check the terminal before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' | tail -5
# your prompt still on the ❯ composer line = never submitted (missing \r);
# submit it with {"input":"\r"} (the only recovery), then loop again
fi
break
done
SEQ=$((SEQ+1)); jq '.data.wait.signal, .data.status' <<<"$R"
```
Read the outcome in this order: `wait.signal != null` → done (`stop` is definitive;
`idle` is heuristic) — **unless** it arrived as `duplicate:true` + `immediate:true`,
which only says the session is idle *now* and must be confirmed from the terminal
(above); `wait.timedOut` → loop again (bounded); `wait.ended` → session gone, stop.
If the loop exhausts its cap, do not keep looping: read the terminal, report what
you see, and remember that a still-typed-but-unsubmitted prompt (missing `\r`) can
only be recovered by submitting it with `{"input":"\r"}`.
**Shell worker + completion marker** — the pattern for `shell` mode (no hooks there).
The typed line must not contain the marker verbatim (the input echo would match
instantly — observed live), so build it with a variable the worker's shell expands:
```bash
N="${RANDOM}_$$"; MARK="DONE_$N" # unique per call: tmux repaints replay old text
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}'
SEQ=$((SEQ+1))
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=120000' \
| jq -r '.data.wait | {matched, snippet}'
```
The typed line shows `${M}_…`, the real output shows `DONE_… rc=<exit code>`, and the
snippet carries the exit code back to you.
**Read a worker's output** — the terminal buffer, tail in **bytes** (`textOutput` in
`GET .../output` stays empty for interactive sessions; don't use it):
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=3000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
```
Avoid `?full=1` (entire tmux scrollback, a context bomb) unless doing a post-mortem.
**Detect a dead worker cheaply**: `GET .../wait?until=exit&timeout=60000` answers
immediately (`signal:"exit"`, `immediate:true`) if the PTY is gone — including a
worker that exited *inside* its pane, which `GET .../sessions/:id` keeps reporting
as `status:"idle"` with a pid (that pid is the local tmux attach client, not the
worker). The wait routes are the only liveness check; a worker dying while a wait
is parked resolves it within ~3 s. A session deleted mid-wait resolves in ~1 s.
**Clean up** — only ids you created, `is_self`-checked, one at a time:
```bash
is_self "$SID" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Everything else (endpoint tables, per-mode signal table, error codes, capacity
limits, Docker/remote caveats): [reference/endpoints.md](reference/endpoints.md).
Fan-out orchestration and blocked-worker handling:
[reference/recipes.md](reference/recipes.md).
+214
View File
@@ -0,0 +1,214 @@
# Codeman API reference for agents
Loaded on demand from the `codeman` skill. Assumes the guard variables from SKILL.md
(`$API`, `$SELF`, `"${CURL[@]}"`). Canonical contract: `docs/api-reference.md` in the
Codeman repo; this file is the agent-relevant subset, verified live.
## Envelope and errors
Every JSON response: `{"success":true,"data":…}` or
`{"success":false,"error":"…","errorCode":"…"}`. Branch on `errorCode`:
| `errorCode` | HTTP | Meaning |
|-------------|------|---------|
| `INVALID_INPUT` | 400 | malformed request; the message names the bad field |
| `UNAUTHORIZED` | 401 | auth required or failed (send `-u user:password`). ⚠️ The 401 body is plain text, NOT this envelope — `jq` dies with a parse error, see the guard in SKILL.md |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap (16, combined signal+output) is full |
| `CONFLICT` / `ALREADY_EXISTS` | 409 | conflicts with current state |
| `OPERATION_FAILED` | 422 | well-formed but could not be completed |
| `RATE_LIMITED` | 429 | per-owner or process-wide waiter pool is full — back off; switching sessions will not help |
| `INTERNAL_ERROR` | 500 | server bug |
`SESSION_BUSY` vs `RATE_LIMITED` on the wait endpoints is deliberate: the first means
"too many waiters on *this* session", the second means the *pool* is full.
## Sessions
| Task | Call |
|------|------|
| list sessions (metadata only, ~1.5 KB each, safe to poll) | `GET /api/v1/sessions` |
| one session (has `.data.pid`, `null` until the PTY spawns) | `GET /api/v1/sessions/:id` — ⚠️ **not a liveness check**: a worker that dies inside its pane keeps `status:"idle"` and a pid (the tmux attach client); `wait?until=exit` is the death check |
| unified list incl. history | `GET /api/v1/sessions/unified` → `.data.sessions[]` (NOT `.data[]`), and it folds in transcript history from the whole machine — never use it to verify cleanup; `GET /api/v1/sessions` is the cleanup check |
| start case + session in one call | `POST /api/v1/quick-start` |
| send input | `POST /api/v1/sessions/:id/input` |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer` |
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
| background agents of a session | `GET /api/v1/subagents` |
| server status / version | `GET /api/v1/status` → `.data.version` |
| delete one session (yours, `is_self`-checked) | `DELETE /api/v1/sessions/:id` |
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Read
`terminal?tail=` instead and strip ANSI:
```bash
… | jq -r '.data.terminalBuffer' | sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g'
```
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
— `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing — do not retry it in a loop, and remember the name.
`POST /api/v1/sessions/:id/input` body:
`{"input":"one line\r","useMux":true,"clientId":"agent-1","seq":1}` plus optionally
`"wait"` / `"waitTimeout"` (below).
- ⚠️ **The input must contain `\r`** (the JSON escape, i.e. a real carriage return)
**or Enter is never sent**: the text is typed onto the worker's prompt and sits
there unsubmitted. Verified live — this is the number-one silent failure, and no
response field catches it: `delivered:true` means "written to the pane", not
"submitted". A `\r`-less send with `wait` reports `delivered:true` and then every
wait on that turn times out. Without `wait`, fire-and-forget returns an **empty**
`{"success":true,"data":{}}` — no `delivered`, no `duplicate`; those fields exist
only on the `wait` variant, so a fire-and-forget flow gets no delivery
confirmation at all.
- `input` must be single-line (newlines are stripped). To send a bare Enter (confirm
a dialog), send `{"input":"\r"}`.
- `clientId`+`seq` give exactly-once delivery: the server applies each pair at most
once. Increment `seq` per new input.
## The wait primitives
Three bounded long-polls. Shared semantics:
- **Timeout = HTTP 200** with `wait.timedOut:true`. Loop over short waits (60 s);
`tailscale serve` / cloudflared cut idle connections.
- Timeouts are **clamped** to `[1000, 600000]` ms (operator-tunable); the applied
value is echoed as `wait.timeoutMs` — read it back, never assume.
- All three nest the result under `.data.wait`, same shape, so one helper parses all.
- `.data.status` (post-wait `SessionStatus`) and `.data.limitPaused` ride along.
`limitPaused:true` means the session is paused on a usage limit and will emit
nothing until reset — a timeout is then *expected*; do not retry hard, and do not
kill the worker.
### Signals by mode
| Signal | Meaning | Available for |
|--------|---------|---------------|
| `idle` | output stabilized + prompt detected — heuristic, can flap mid-turn | every mode |
| `working` | session started producing output | every mode |
| `stop` | Claude Code `stop` hook — the definitive end-of-turn | `claude` only |
| `blocked` | `permission_prompt` / `elicitation_dialog` hook — the worker needs an answer | `claude` only |
| `exit` | PTY exited or session deleted | every mode |
Default `until` set: `stop,idle,exit`. On non-claude modes the server silently drops
`stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
`["idle","exit"]` on shell); requesting them *explicitly* there is a 400 naming the
mode. ⚠️ On hook-less modes the lifecycle signals are also **coarse in practice**: a
short shell command produced **no** `idle` transition within 60 s (verified live), so
a `fresh=1` / fresh-delivery wait can burn its whole timeout while the work finished
long ago. Synchronize hook-less modes with `wait-output` markers instead.
Two more places hooks go missing even in claude mode: **Docker cases** need
`CODEMAN_DOCKER_BRIDGE_HOOKS=1` on the server (without it only `idle`/`working`/
`exit` arrive), and **remote-SSH cases** run the agent on another host whose hooks may
never reach this server. When unsure, ask for `stop,idle,exit`.
⚠️ **Signals are edge-triggered with no history.** A signal that fires while no
waiter is registered is gone; no later wait can observe it (`until=stop` on a worker
whose turn already ended just times out, with or without `fresh` — verified live).
Register the waiter before the event can happen: send-and-wait does exactly that,
and `wait-output` markers with `from=buffer` are latched by construction. Never
fire-and-forget N prompts and then gather signal-waits worker by worker; every
worker that finishes before its gather is unobservable (see recipes.md Flow 3b).
### `GET /api/v1/sessions/:id/wait`
| Param | Default | Notes |
|-------|---------|-------|
| `until` | `stop,idle,exit` | comma list; unknown token → 400 naming it |
| `timeout` | 60000 | ms, clamped; applied value echoed as `wait.timeoutMs` |
| `fresh` | `0` | `1` requires an actual *transition*, ignoring the state at call time |
⚠️ A session whose PTY has not spawned (`pid:null`) or has exited counts as `exit`
**right now**: with the default set the call answers immediately
(`signal:"exit", immediate:true`). That is how you detect a dead worker cheaply — but
it also means "wait for my just-created session" needs the readiness recipe in
SKILL.md, not this endpoint.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Default | Notes |
|-------|---------|-------|
| `match` | required | literal substring, 1–200 chars, ANSI-stripped; chunk-straddling matches found; **no regex** — a `regex=` param is a 400 |
| `nocase` | `0` | case-insensitive compare; snippet keeps original casing |
| `from` | `now` | `buffer` scans the tail (~256 KB) of existing output first |
| `timeout` | 60000 | same clamp |
Four traps, all observed live:
1. **The echo of your own typed command is output.** A marker appearing verbatim in
the input line matches the moment the text is typed, before the command runs.
Split the marker with a shell variable: send `M=DONE; …; echo ${M}_1234\r`, wait
on `DONE_1234`.
2. **`from=now` misses text printed before the wait landed** — a marker echoed just
before the request registered timed out at full length. After sending a command,
always wait with `from=buffer`.
3. **`from=now` can also match too much**: tmux repaints old screen content as
ordinary output on attach/resize/redraw, so a *generic* marker (`BUILD OK`)
matches stale text. Unique-per-call markers (`DONE_$RANDOM`) make both `from`
modes safe.
4. **TUI output can be space-less in the stream.** Full-screen TUIs (claude, codex,
…) position words with cursor-movement escapes rather than literal spaces, so
the stripped stream can read `Yes,Itrustthisfolder` while the pane shows the
spaced phrase. Whether a given phrase keeps its spaces depends on how the TUI
drew it (observed live: some multi-word matches fire, some never do), so treat
multi-word matches against TUI screens as unreliable and match a **single
space-free token** (`trust`, `bypass`). Plain command output (shell workers,
`echo` lines) keeps real spaces and multi-word matches work there.
Build the query with `-G --data-urlencode` (a `+` in a hand-built query decodes to a
space). Result extras: `wait.matched`, `wait.match`, `wait.snippet` (bounded window
around the match, blank runs collapsed — the snippet is often all you need to read).
### `POST /api/v1/sessions/:id/input` with `wait`
| Field | Notes |
|-------|-------|
| `wait` | `true` (default signal set) or the same comma grammar as `until`; absent = historical fire-and-forget |
| `waitTimeout` | ms, same clamp |
Registers the waiter **before** typing, which closes the race where send-then-wait
sees the previous turn's idle state and returns instantly. Response adds `delivered`
and `duplicate` beside the standard `wait` object.
A **tagged duplicate** (same `clientId`+`seq` already applied) does not retype but
still honors `wait`, answering from the session's *current* state instead of
requiring a new transition (`delivered:false, duplicate:true` — verified: ~20 ms,
command ran exactly once). That is what makes the resend-identical-request loop in
SKILL.md correct: iteration 1 delivers and needs a transition; later iterations
resolve immediately if the turn ended in between. ⚠️ The flip side: a duplicate's
`immediate:true` answer is the current state and nothing more — an idle worker
whose prompt was never submitted (missing `\r`) produces the same
`signal:"idle", immediate:true` as one that finished the turn. Confirm from
`terminal?tail=` before reporting success; SKILL.md's loop shows where.
### Outcome parsing, in order
1. `wait.signal != null` (or `wait.matched == true`) — the thing happened.
`wait.immediate:true` rides along and means the condition already held at call
time; if that is not what you meant, you wanted `fresh=1` or send-and-wait.
2. `wait.timedOut` — poll boundary; loop again.
3. `wait.ended` — session deleted/torn down mid-wait; stop looping.
## Troubleshooting
| Symptom | Cause / fix |
|---------|-------------|
| every curl fails with a certificate error | you dropped `-k`; `CODEMAN_API_URL` is HTTPS with a self-signed cert |
| `jq: parse error` on every call | plain-text 401s: the server has a password. Check with `-w '%{http_code}'`, use the guard's `.env` fallback, and if no `.env` exists, stop and ask the user for credentials |
| input arrives but nothing happens; later waits all time out | the input had no `\r`, so Enter was never sent; the text is sitting on the worker's prompt. **Submitting it with `{"input":"\r"}` is the ONLY recovery** — Ctrl+U (0x15) and Esc do NOT clear the composer (verified live) — and the flush costs one turn in which the worker reasons about the junk; open the next real prompt with "ignore the garbled line above:" |
| `GET .../sessions/$CODEMAN_SESSION_ID` 404s | Docker case: the env id is truncated to 8 chars; find yourself with `startswith($SELF)`, and always self-compare by prefix, in both directions |
| `CODEMAN_MUX` unset but you seem to be in a session | remote-SSH case: the env vars are not exported there. Fail closed — refuse to act |
| connection refused from inside a container | a loopback-bound server is unreachable from a container, and `CODEMAN_DOCKER_BRIDGE_HOOKS=1` does **not** fix that: it opens a hooks-only listener, so hook events start flowing but `/api/v1/*` stays refused. Driving the API from inside a Docker case needs a reachable bind (an operator decision); report it, don't retry |
| wait routes 404 on a valid session id | read the `.error` text: a `Route ...` prefix means the server predates the wait endpoints (< 1.13.0; a dev build can serve them while reporting an older version, so probe, never version-compare) — poll `terminal?tail=` and say so. `Session ... not found` means your id is wrong, not the server |
| wait on `stop` never resolves | non-claude mode, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept missed; use the readiness recipe in SKILL.md (wait for `bypass` first, accept the dialog only as the bounded fallback) |
| `wait-output` times out although the pane shows the text | multi-word match against a TUI screen; the stream has no spaces there — match one token |
| `wait-output` matched instantly with stale text | generic marker + tmux repaint; use `DONE_$RANDOM` |
| 409 `SESSION_BUSY` on a wait | too many concurrent waiters on that session (cap 16 combined); reuse one wait per worker |
| 429 `RATE_LIMITED` on a wait | global/owner waiter pool full; back off, do not switch sessions |
+249
View File
@@ -0,0 +1,249 @@
# Worked orchestration flows
Loaded on demand from the `codeman` skill. Every flow assumes the guard preamble from
SKILL.md ran (`$API`, `$SELF`, `"${CURL[@]}"`, `is_self`). Track every session id you
create; delete them (and only them) when done. Remember the two silent killers:
**every input ends with `\r`**, and **markers must be split** so the typed-line echo
does not match them.
## Flow 1: claude worker, end to end
Start a worker, get it truly ready (trust dialog included), give it a task, wait for
the turn to finish, read the answer, clean up. Verified live: the stop hook resolves
the send-and-wait within seconds of the turn ending.
```bash
# 1. start (returns before the CLI inside is ready)
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-tests","mode":"claude"}' | jq -r '.data.sessionId')
CREATED+=("$SID") # the cleanup list
CID="agent-$$"; SEQ=1
# 2. readiness. "wait for idle" or "wait for ❯" is NOT readiness: a fresh session
# reports idle before anything spawned, and the first-run trust dialog contains ❯.
# Codeman CAN auto-accept that dialog, but the accept misses on some runs (both
# outcomes seen live), so: composer marker first, dialog only as the bounded
# fallback (a blind Enter up front would land in an already-ready composer).
# Stage 1 is SHORT on purpose: an already-trusted case matches in <1 s, while a
# virgin case can never pass it (the dialog is up) and always pays it in full —
# the long budget belongs to stage 3, after the dialog is answered.
# Single-token matches only: TUI text is space-less in the stream.
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# (pid != null proves startup only — a worker that later dies inside its pane keeps
# status "idle" and a pid. The death check is wait?until=exit.)
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null || echo "worker $SID not ready; inspect terminal?tail="
fi
# 3. send-and-wait, looping on the IDENTICAL request (tagged duplicate: no retype).
# BOUNDED (a \r-less send would otherwise loop forever), body built with jq -n so
# quotes/backslashes/$ in a real prompt survive; note the appended \r.
PROMPT='run the unit tests and summarize failures in one line'
BODY=$(jq -n --arg p "$PROMPT" --arg c "$CID" --argjson s "$SEQ" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:60000}')
for TRY in $(seq 1 10); do
R=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' --data-binary "$BODY")
if jq -e '.data.wait.timedOut' <<<"$R" >/dev/null; then
jq -e '.data.limitPaused' <<<"$R" >/dev/null && sleep 60 # usage-limit pause: silence is expected
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # is the prompt sitting unsubmitted?
continue
fi
# Resolved — but duplicate + immediate is only "the session is idle NOW", which a
# never-submitted (\r-less) prompt also produces. Check before believing it:
if jq -e '.data.duplicate and .data.wait.immediate' <<<"$R" >/dev/null; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5
# prompt still on the ❯ composer line = never submitted; {"input":"\r"} is the
# only recovery, then loop again
fi
break
done
SEQ=$((SEQ+1))
# 4. interpret
case "$(jq -r '.data.wait.signal' <<<"$R")" in
stop) : ;; # definitive end of turn
idle) : ;; # heuristic — and if it rode a duplicate with
# immediate:true, it proves nothing ran (step 3)
exit) echo "worker died" ;;
null) jq -e '.data.wait.ended' <<<"$R" >/dev/null && echo "worker deleted mid-wait" ;;
esac
# 5. read the answer: terminal tail (BYTES), ANSI-stripped. textOutput stays empty
# for interactive sessions; terminal?full=1 is a context bomb.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=4000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' -e 's/\x1b([B0]//g' | grep -v '^[[:space:]]*$' | tail -30
# 6. clean up — exact id, own list only, self-check
is_self "$SID" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
re-ask about the same delivery (the duplicate-wait loop above).
## Flow 2: shell worker running a build, marker-synchronized
`shell` sessions have no hooks (`stop`/`blocked` are a 400 there), and their lifecycle
signals are coarse — a short command may emit no `idle` transition at all (verified
live), so send-and-wait can burn its whole timeout. The reliable pattern is a split,
unique marker plus `wait-output from=buffer`:
```bash
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"builder","mode":"shell"}' | jq -r '.data.sessionId')
CREATED+=("$SID")
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
# Split marker: the typed line carries ${M}_N, only the OUTPUT carries DONE_N.
# An unsplit marker matches the echo of your own keystrokes before the build runs.
N="${RANDOM}_$$"; MARK="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run build; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"build-'$$'","seq":1}'
for TRY in $(seq 1 30); do # BOUNDED (30 min): a \r-less send makes an uncapped loop infinite
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
jq -e '.data.wait.matched' <<<"$R" >/dev/null && break
jq -e '.data.wait.ended' <<<"$R" >/dev/null && { echo "worker gone"; break; }
[ "$TRY" = 2 ] && "${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" \
| jq -r '.data.terminalBuffer' | tail -5 # command still sitting unsubmitted?
done
jq -r '.data.wait.snippet' <<<"$R" # e.g. "DONE_123_456 rc=0" — the exit code rides the marker line
```
## Flow 3: fan out N workers, gather as each finishes
Start everything first, then gather. One in-flight wait per worker — the per-session
waiter cap is 16 and abandoned concurrent waits pile up against it.
```bash
declare -A WORKER MARKS
for task in lint typecheck unit; do
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"fan-'"$task"'","mode":"shell"}' | jq -r '.data.sessionId')
WORKER[$task]=$SID; CREATED+=("$SID")
done
for task in "${!WORKER[@]}"; do
SID=${WORKER[$task]}
for _ in $(seq 1 30); do
[ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ] && break; sleep 1
done
N="${task}_${RANDOM}"; MARKS[$task]="DONE_$N"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"M=DONE; npm run '"$task"'; echo ${M}_'"$N"' rc=$?\r","useMux":true,"clientId":"fan-'$$'","seq":1}'
done
for task in "${!WORKER[@]}"; do # sequential gather; each wait blocks until that worker is done
for TRY in $(seq 1 30); do # BOUNDED per worker, same reasoning as Flow 2
R=$("${CURL[@]}" -G "$API/api/v1/sessions/${WORKER[$task]}/wait-output" \
--data-urlencode "match=${MARKS[$task]}" --data-urlencode 'from=buffer' --data-urlencode 'timeout=60000')
jq -e '.data.wait.matched or .data.wait.ended' <<<"$R" >/dev/null && break
done
echo "$task: $(jq -r '.data.wait.snippet // "worker gone"' <<<"$R" | tail -1)"
done
```
## Flow 3b: fan out N CLAUDE workers
Send-and-wait is synchronous, so the shell-flow shape ("send everything, then
gather") does not translate directly: the send *is* the wait, and worker 2's prompt
would not go out until worker 1's turn ended. Two working patterns, both verified
live (and one anti-pattern, measured failing, replaced by B):
**A. Background the send-and-waits** (simplest; each resolved on `stop` while the
other was still running):
```bash
sendwait() { # $1=sid $2=prompt $3=seq — assumes the worker passed Flow 1's readiness
local body; body=$(jq -n --arg p "$2" --argjson s "$3" \
'{input:($p+"\r"),useMux:true,clientId:"fan-'$$'",seq:$s,wait:true,waitTimeout:600000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$1/input" \
-H 'Content-Type: application/json' --data-binary "$body" > "/tmp/fan-$1.json"
}
( sendwait "$SID1" 'refactor module A and reply DONE' 2 & \
sendwait "$SID2" 'write tests for module B and reply DONE' 2 & wait )
jq -c '.data.wait | {signal, waitedMs}' /tmp/fan-"$SID1".json /tmp/fan-"$SID2".json
```
One in-flight wait per worker keeps you far from the 16-per-session waiter cap.
**B. Fire-and-forget, then gather with output markers.** If you must send every
prompt before waiting on anything, do **not** gather with signal waits: signals
are edge-triggered with no history, so a `stop` that fires before the gather
reaches that worker is gone and unobservable afterwards — `fresh=1` cannot help,
and neither can omitting it (measured: worker 2's turn ended at +2 s, its
sequential `until=stop,exit&fresh=1` gather burned its full bounded 300 s and
reported nothing). Gather instead on a marker each worker prints itself, which
`from=buffer` re-finds no matter when it appeared:
```bash
# SIDS[1], SIDS[2] = worker ids that already passed Flow 1's readiness.
# The typed prompt must NOT contain the finished marker verbatim (your keystrokes
# echo into the output stream and would match instantly), so ask for it in halves:
declare -A TOK
for i in 1 2; do
TOK[$i]="${RANDOM}_$i"
BODY=$(jq -n --arg p "do task $i; when completely done print the word WORKDONE immediately followed by _${TOK[$i]}" \
--arg c "fan-$$" --argjson s 2 '{input:($p+"\r"),useMux:true,clientId:$c,seq:$s}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/${SIDS[$i]}/input" \
-H 'Content-Type: application/json' --data-binary "$BODY"
done
for i in 1 2; do # order no longer matters: the marker is latched in the buffer
"${CURL[@]}" -G "$API/api/v1/sessions/${SIDS[$i]}/wait-output" \
--data-urlencode "match=WORKDONE_${TOK[$i]}" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=600000' | jq -c '.data.wait | {matched, snippet}'
done
```
Use A unless you genuinely need to send everything before waiting on anything: A
needs no marker discipline, and resolves on the definitive `stop` instead of on
the worker remembering to print a token.
## Flow 4: watch for a worker stuck on a permission prompt
Claude workers can block on a permission dialog. `blocked` is a wait signal
(claude-mode only), so watch for it and surface the question to the user instead of
guessing an answer:
```bash
R=$("${CURL[@]}" "$API/api/v1/sessions/$SID/wait?until=stop,blocked,exit&timeout=60000")
if [ "$(jq -r '.data.wait.signal' <<<"$R")" = blocked ]; then
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=2000" | jq -r '.data.terminalBuffer' \
| sed -e 's/\x1b\[[0-9;?]*[a-zA-Z]//g' | grep -v '^[[:space:]]*$' | tail -15
# show this to the user and ask how to answer; do NOT auto-confirm another
# session's permission prompt
fi
```
## Cleanup discipline
At the end of the conversation (or on abort), delete exactly what you created:
```bash
for id in "${CREATED[@]}"; do
is_self "$id" || "${CURL[@]}" -X DELETE "$API/api/v1/sessions/$id"
done
```
- Only ids from your own `CREATED` list. Never enumerate `/api/v1/sessions` and
delete by pattern; other sessions belong to the user.
- If you created a *case* purely as scratch and the user confirmed it is disposable,
`DELETE /api/v1/cases/:name` removes it — but that recursively deletes the
directory from disk, so never do it without the user's explicit go-ahead for that
exact name.
+112
View File
@@ -0,0 +1,112 @@
/**
* @fileoverview Bounds for the agent wait primitives.
*
* These back the blocking endpoints an agent uses to orchestrate other sessions
* (`GET /api/sessions/:id/wait`, `GET /api/sessions/:id/wait-output`, and the
* `wait` field on `POST /api/sessions/:id/input`). Plan: `docs/agent-control-plan.md`.
*
* Why every value is bounded:
* - An unbounded long-poll is a socket leak. A caller that asks for a 12-hour wait
* and walks away holds a connection (and a waiter, and a timer) until the process
* restarts, so `MAX_WAIT_MS` is a hard ceiling applied server-side.
* - `DEFAULT_WAIT_MS` is deliberately short (60s). Production is reached through
* `tailscale serve` and users also run cloudflared tunnels; both can cut an idle
* connection, so the documented pattern is a client-side loop over short waits
* rather than one very long call. Fastify itself is happy to hold the request
* (`requestTimeout` defaults to 0, and `keepAliveTimeout` applies between
* requests, not to an in-flight one), the intermediaries are the constraint.
* - The waiter caps mirror `MAX_SSE_CLIENTS` in `map-limits.ts`: each pending
* waiter costs an open HTTP response plus a timer, so the pool is capped rather
* than queued. Exceeding a cap is an explicit error, never a silent wait.
* - There are THREE caps, not two, because a process-wide pool with no per-user
* dimension lets one user deny the primitive to everyone else. `middleware/auth.ts`
* already treats that shape as a bug (its `userFailures` bucket exists so "one user
* behind a NAT can't lock out everyone else"); `MAX_WAITERS_PER_OWNER` is the same
* idea for waiters. It applies only when the caller has an owner, so single-user
* mode is byte-identical to having no owner cap at all.
*
* All values are env-overridable and clamped to sane hard bounds, so a typo in an
* env var degrades to the default instead of disabling the protection.
*
* @module config/agent-wait
*/
/** Absolute floor for any wait, in ms. Sub-second waits are polling, not waiting. */
export const MIN_WAIT_MS = 1_000;
/** Ceiling the operator-configurable maximum is itself clamped to. */
const HARD_MAX_WAIT_MS = 3_600_000;
function envInt(name: string, fallback: number, min: number, max: number): number {
const raw = parseInt(process.env[name] || '', 10);
if (!Number.isFinite(raw) || raw <= 0) return fallback;
return Math.max(min, Math.min(max, raw));
}
/** Longest a single wait may block. Requests above this are clamped down, not rejected. */
export const MAX_WAIT_MS = envInt('CODEMAN_WAIT_MAX_MS', 600_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS);
/** Used when the caller omits `timeout`. Never exceeds MAX_WAIT_MS. */
export const DEFAULT_WAIT_MS = Math.min(
envInt('CODEMAN_WAIT_DEFAULT_MS', 60_000, MIN_WAIT_MS, HARD_MAX_WAIT_MS),
MAX_WAIT_MS
);
/** Concurrent waiters (signal + output) allowed against one session. */
export const MAX_WAITERS_PER_SESSION = envInt('CODEMAN_WAIT_MAX_PER_SESSION', 16, 1, 256);
/**
* Ceiling the operator-configurable total is itself clamped to.
*
* 512 rather than the 4096 this started at. Every other knob in this file degrades
* safely on a bad value; a 4096 ceiling instead lets a well-meaning operator turn the
* protection into the problem, since 4096 concurrent held responses (each an open
* socket, a timer and a pending promise) exceeds the 1024 soft `RLIMIT_NOFILE` that is
* still the default on most Linux distros, before counting PTYs, SSE clients and
* WebSockets. 512 is ~5x `MAX_SSE_CLIENTS` (100, the pool this one is modelled on), so
* the knob stays useful for a busy orchestration host while the whole server still fits
* inside a default fd budget with room to spare.
*/
const HARD_MAX_WAITERS_TOTAL = 512;
/** Concurrent waiters allowed across every session in the process. */
export const MAX_WAITERS_TOTAL = envInt('CODEMAN_WAIT_MAX_TOTAL', 128, 1, HARD_MAX_WAITERS_TOTAL);
/**
* Concurrent waiters allowed for one owner (multi-user mode's `Session.owner`).
*
* Sits between the per-session cap (16) and the process-wide one (128): high enough
* that one user orchestrating several workers at once never trips it, low enough that
* a single user cannot occupy the whole pool and deny the primitive to everyone else,
* admin included. Ignored entirely when the caller has no owner, which is every
* request in single-user mode.
*/
export const MAX_WAITERS_PER_OWNER = envInt('CODEMAN_WAIT_MAX_PER_OWNER', 48, 1, HARD_MAX_WAITERS_TOTAL);
/** Bounds on the literal `match` string accepted by wait-output. */
export const MIN_MATCH_LENGTH = 1;
export const MAX_MATCH_LENGTH = 200;
/**
* Tail of the terminal buffer scanned by `wait-output?from=buffer`.
*
* The buffer itself runs to 32MB. Scanning all of it would be an ANSI strip over
* 32MB (a full second copy) on a request an agent may issue in a loop, and the
* question `from=buffer` answers is "did this appear recently", not "ever". The
* tail is continuous with the live stream, since `_terminalBuffer.append(data)`
* and `emit('terminal', data)` receive the same bytes.
*/
export const MAX_BUFFER_SCAN_BYTES = envInt('CODEMAN_WAIT_BUFFER_SCAN_BYTES', 256 * 1024, 4 * 1024, 8 * 1024 * 1024);
/** Characters of surrounding output returned either side of a wait-output match. */
export const MAX_SNIPPET_CONTEXT = 80;
/**
* Clamp a caller-supplied timeout into [MIN_WAIT_MS, MAX_WAIT_MS].
* Absent / non-numeric / non-finite input falls back to DEFAULT_WAIT_MS.
*/
export function clampWaitMs(value: unknown): number {
const n = typeof value === 'string' ? Number(value) : value;
if (typeof n !== 'number' || !Number.isFinite(n)) return DEFAULT_WAIT_MS;
return Math.max(MIN_WAIT_MS, Math.min(MAX_WAIT_MS, Math.trunc(n)));
}
+8
View File
@@ -98,6 +98,14 @@ export const DEPENDENCY_REGISTRY: ToolDependency[] = [
usedBy: ['Gemini sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['gemini'], versionArg: '--version' } }],
},
{
id: 'antigravity',
label: 'Antigravity CLI',
category: 'core',
required: false,
usedBy: ['Antigravity sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['agy'], versionArg: '--version' } }],
},
{
id: 'libreoffice',
label: 'LibreOffice',
+152
View File
@@ -0,0 +1,152 @@
/**
* @fileoverview File Viewer edit-mode policy (issue #212).
*
* Pure, IO-free policy for which workspace files the in-viewer editor may read
* for editing and write back. Consumed by the `edit=1` branch of
* `GET /api/sessions/:id/file-content` and by `PUT /api/sessions/:id/file-content`
* in `src/web/routes/file-routes.ts`.
*
* Design (docs/file-viewer-edit-plan.md):
* - ALLOWLIST of text extensions/basenames, not a blocklist — matching the
* attachment-guard precedent. `svg` and `env` are deliberately absent: svg is
* treated as untrusted on the read side, and `.env` is sensitive-path blocked
* anyway; excluding them here keeps a single obvious refusal.
* - The `.git/` subtree is denied outright: `.git/hooks/*` is code execution and
* a corrupted index looks unrecoverable to a user who wanted to fix a typo.
* - EOL helpers exist because a browser <textarea> normalizes to LF; the server
* re-applies the file's original ending so a two-line edit of a CRLF file does
* not become a whole-file diff. Mixed-EOL files normalize to the dominant
* style (documented lossy edge).
*/
/** Hard cap for edit-mode reads AND writes (bytes of file content). */
export const MAX_EDITABLE_BYTES = 512 * 1024;
/** Lowercase extensions (no dot) the editor will open and save. */
export const EDITABLE_EXTENSIONS: ReadonlySet<string> = new Set([
// JS/TS ecosystem
'ts',
'tsx',
'js',
'jsx',
'mjs',
'cjs',
'json',
'jsonc',
// Docs / plain text
'md',
'mdx',
'txt',
'rst',
'adoc',
// Web
'css',
'scss',
'less',
'html',
'htm',
'xml',
// Config
'yml',
'yaml',
'toml',
'ini',
'cfg',
'conf',
'properties',
// Shell
'sh',
'bash',
'zsh',
'fish',
// Languages
'py',
'rb',
'go',
'rs',
'java',
'kt',
'swift',
'c',
'h',
'cpp',
'hpp',
'cc',
'cs',
'php',
'sql',
'graphql',
'proto',
'lua',
'pl',
'r',
'jl',
'tf',
'gradle',
// Data / misc text
'csv',
'tsv',
'log',
'diff',
'patch',
]);
/** Extensionless (or dot-led) file names that are still editable text. */
export const EDITABLE_BASENAMES: ReadonlySet<string> = new Set([
'dockerfile',
'makefile',
'license',
'readme',
'changelog',
'authors',
'codeowners',
'procfile',
'.gitignore',
'.gitattributes',
'.dockerignore',
'.prettierignore',
'.prettierrc',
'.editorconfig',
'.nvmrc',
'.npmrc',
'.eslintignore',
]);
/** Whether a file name (basename only) is eligible for in-viewer editing. */
export function isEditableFileName(fileName: string): boolean {
const lower = fileName.toLowerCase();
if (EDITABLE_BASENAMES.has(lower)) return true;
const dot = lower.lastIndexOf('.');
// No extension (or a bare dotfile like `.bashrc`): only the basename list applies.
if (dot <= 0) return false;
return EDITABLE_EXTENSIONS.has(lower.slice(dot + 1));
}
/**
* Whether a workspace-relative path is denied for editing regardless of its
* extension. Currently: anything inside a `.git` directory at any depth.
*/
export function isDeniedEditRelativePath(relativePath: string): boolean {
return relativePath.split('/').some((segment) => segment === '.git');
}
export type FileEol = 'lf' | 'crlf';
/** Dominant line-ending style of a text buffer (LF when tied or single-line). */
export function detectEol(text: string): FileEol {
let crlf = 0;
let lf = 0;
for (let i = 0; i < text.length; i++) {
if (text.charCodeAt(i) === 10) {
if (i > 0 && text.charCodeAt(i - 1) === 13) crlf++;
else lf++;
}
}
return crlf > lf ? 'crlf' : 'lf';
}
/** Normalize every line ending in `text` to the requested style. */
export function applyEol(text: string, eol: FileEol): string {
const normalized = text.replace(/\r\n/g, '\n');
return eol === 'crlf' ? normalized.replace(/\n/g, '\r\n') : normalized;
}
+4
View File
@@ -143,6 +143,7 @@ export function defaultDockerCommandForMode(mode: SessionMode): string {
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
antigravity: 'exec agy',
};
return commands[mode as DockerCommandMode] || commands.shell;
}
@@ -595,6 +596,9 @@ interface CredStorePolicy {
const CRED_STORES: CredStorePolicy[] = [
{ rel: '.codex', shareDirs: ['sessions'], shareFiles: ['history.jsonl'], seedFiles: ['auth.json', 'config.toml'] },
// Also covers Antigravity: `agy` nests its whole state (auth `jetski_state.pbtxt`,
// `conversations/`, `knowledge/`) under `~/.gemini/antigravity-cli/`, so it needs no
// entry of its own. There is no `~/.antigravity` credential dir to add.
{ rel: '.gemini', seedWhole: true },
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
+14 -4
View File
@@ -173,7 +173,11 @@ export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
const curlCmd = (event: HookEventType) =>
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"event":"${event}","sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$HOOK_DATA" | ` +
`curl -s -X POST "$CODEMAN_API_URL/api/hook-event" ` +
// `-k`, same as the statusline exporter: CODEMAN_API_URL is loopback HTTPS with
// a self-signed cert on --https/tailscale installs. Without it curl exits 60,
// the `|| true` swallows it, and ALL SIX hook events die silently: respawn loses
// its definitive idle signals and the wait endpoints lose stop/blocked.
`curl -sk -X POST "$CODEMAN_API_URL/api/hook-event" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- ` +
@@ -433,8 +437,10 @@ export async function writeHooksConfig(casePath: string): Promise<void> {
* X-Codeman-Hook-Secret header was added (COD-54, 2026-06-10) keep hook curls in their
* settings.local.json that POST to /api/hook-event WITHOUT the secret — which, once the
* gate requires it unconditionally (COD-91), silently 401 on a password-protected install.
* Older Codeman blocks also lack the background Bash async-rewake hook. Refresh either
* stale shape on launch so existing cases gain both current behaviors.
* Older Codeman blocks also lack the background Bash async-rewake hook. A third stale
* shape: hook curls without `-k`, which exit 60 on every --https/tailscale install (the
* cert is self-signed), swallowed by the hooks' own `|| true` — all six hook events die
* silently. Refresh any of these stale shapes on launch so existing cases heal.
*
* Deliberately surgical: regenerates ONLY when settings.local.json already contains
* Codeman's own hook curls (they target `/api/hook-event`) and they are stale. No-op
@@ -457,7 +463,11 @@ export async function refreshStaleCodemanHooks(casePath: string): Promise<void>
// absence on our own hooks means they predate COD-54 and need regenerating.
const hasSecret = hooksJson.includes('X-Codeman-Hook-Secret');
const hasBackgroundWake = hooksJson.includes(BACKGROUND_WAKE_MARKER);
if (!isOurs || (hasSecret && hasBackgroundWake)) return;
// The pre--k curl shape: `curl -sk -X POST` does not contain `curl -s -X POST`
// as a substring, so this cleanly identifies hook curls that die with exit 60
// on a self-signed HTTPS install.
const hasTlsFlaglessCurl = hooksJson.includes('curl -s -X POST');
if (!isOurs || (hasSecret && hasBackgroundWake && !hasTlsFlaglessCurl)) return;
const generated = generateHooksConfig();
const merged = {
...existing,
+3
View File
@@ -17,6 +17,7 @@ import type {
CodexConfig,
EffortLevel,
GeminiConfig,
AntigravityConfig,
SessionRemote,
SessionDocker,
} from './types.js';
@@ -74,6 +75,7 @@ export interface CreateSessionOptions {
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Ephemeral — not written to disk. */
@@ -102,6 +104,7 @@ export interface RespawnPaneOptions {
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
+87
View File
@@ -0,0 +1,87 @@
/**
* @fileoverview Bounded descendant walk over a process-tree snapshot.
*
* Split out of `tmux-manager.ts` so the traversal can be unit-tested directly. It
* previously lived as a private method, which meant the regression test had to keep
* its own copy of the algorithm — a test that passes while the shipped code rots.
*
* ## The incident this guards against
*
* On 2026-07-30 an unbounded version of this walk took a machine down. It ran
* `pgrep -P <pid>` once per node and recursed with no visited set, no depth limit and
* no node cap. Across ~28 adopted tmux trees the fan-out exploded, and because each
* `pgrep` blocks in the WSL kernel while reading `/proc/<pid>/cgroup`, none of them
* returned while the walk kept spawning more. Result: ~13,000 `pgrep` processes stuck
* in D-state out of ~39,000 total, load average above 13,000, and a machine only
* recoverable by restarting WSL — which cost every running session.
*
* Three properties make that impossible, and each has a test:
* 1. a cycle terminates instead of looping (stale snapshots can contain one),
* 2. depth is capped,
* 3. node count is capped.
*
* The fourth property — spawning nothing per node — is structural: this function
* takes a snapshot and cannot spawn anything at all.
*
* @module proc-tree
*/
/** Maximum generations to descend. Deeper than any real agent process tree. */
export const PROC_WALK_MAX_DEPTH = 10;
/** Hard ceiling on collected descendants. A backstop, not an expected limit. */
export const PROC_WALK_MAX_NODES = 500;
export interface WalkOptions {
maxDepth?: number;
maxNodes?: number;
/**
* Called once when a cap truncated the result, with which cap it was. Both are
* reported: a silent depth cap would hide a deep tree just as effectively as a
* silent node cap hides a wide one, and the whole point of this module is that
* truncation is visible rather than mysterious.
*/
onTruncated?: (pid: number, cap: number, reason: 'nodes' | 'depth') => void;
}
/**
* All descendants of `pid`, breadth-first and bounded.
*
* @param pid root of the walk; never included in the result
* @param byParent parent pid → child pids, from ONE `ps` snapshot
*/
export function collectDescendants(
pid: number,
byParent: ReadonlyMap<number, readonly number[]>,
opts: WalkOptions = {}
): number[] {
const maxDepth = opts.maxDepth ?? PROC_WALK_MAX_DEPTH;
const maxNodes = opts.maxNodes ?? PROC_WALK_MAX_NODES;
const out: number[] = [];
const visited = new Set<number>([pid]);
let frontier = [pid];
for (let depth = 0; depth < maxDepth && frontier.length; depth += 1) {
const next: number[] = [];
for (const parent of frontier) {
for (const child of byParent.get(parent) ?? []) {
if (visited.has(child)) continue; // a real tree has no cycles, a stale
visited.add(child); // snapshot can still produce one
out.push(child);
next.push(child);
if (out.length >= maxNodes) {
opts.onTruncated?.(pid, maxNodes, 'nodes');
return out;
}
}
}
frontier = next;
// Ran out of generations while descendants were still queued: the tree is
// deeper than the cap and the result is incomplete.
if (depth === maxDepth - 1 && frontier.length > 0) {
opts.onTruncated?.(pid, maxDepth, 'depth');
}
}
return out;
}
+113 -5
View File
@@ -58,16 +58,61 @@ export async function writeRemoteCases(configDir: string, cases: RemoteCase[]):
await writeJsonArray(configDir, remoteCasesPath(configDir), cases);
}
/**
* The remote user's login shell, defaulted and quoted.
*
* The default is belt-and-braces, not a live bug: an empty `$SHELL` would expand
* to `exec -i -l`, which the shell reads as `exec -i` — "not found", pane dead on
* arrival, the #208 failure all over again (verified: `sh -c 'exec $SHELL -i -l'`
* with SHELL unset prints `exec: -i: not found`). In practice tmux always exports
* SHELL into a pane from its own `default-shell` option, so the command as USED
* here is safe either way (also verified). The default matters because these
* strings are the seed values a per-host `commands.*` override is edited from, and
* nothing constrains where an edited one ends up running. Quoted for a shell path
* containing spaces. `/bin/sh` exists on every POSIX host.
*/
const REMOTE_LOGIN_SHELL = '"${SHELL:-/bin/sh}"';
/**
* Run `command` through the remote user's interactive login shell, so per-user
* PATH entries (~/.local/bin, ~/.opencode/bin, …) are resolved before the CLI name
* is looked up. ssh's remote-command execution is neither interactive nor login,
* so a bare `exec claude` sees only sshd's minimal default PATH and dies with
* "command not found" (exit 127).
*
* Shells that take neither flag (nushell, elvish, …) cannot be detected from here
* the way `loginShellArgs()` detects them locally, since the shell is whatever the
* REMOTE passwd says. A host like that is what the per-host `commands.*` override
* is for.
*/
export function remoteLoginShellCommand(command: string): string {
return `exec ${REMOTE_LOGIN_SHELL} -i -l -c ${shellescape(command)}`;
}
export function defaultRemoteCommandForMode(mode: SessionMode): string {
// Agent CLIs (claude/opencode/codex/gemini/antigravity) are typically installed
// under per-user paths like ~/.local/bin or ~/.opencode/bin, added to PATH only by
// the remote user's interactive-login shell startup files (~/.zshrc etc.). ssh's
// remote-command execution is neither interactive nor login, so a bare `exec
// claude` sees only sshd's minimal default PATH and fails with "command not
// found" (exit 127) — confirmed via `tmux capture-pane` on the
// remain-on-exit-preserved dead pane. Route through `$SHELL -i -l -c`, the same
// fix already used for shell mode below, so PATH is fully resolved before the
// CLI name is looked up.
const commands: Record<RemoteCommandMode, string> = {
shell: 'exec bash -l',
// $SHELL, not a hardcoded bash: sshd sets it from the remote user's
// /etc/passwd entry, so this launches their actual login shell (zsh,
// fish, etc.). -i -l so it sources rc files (~/.zshrc etc.), matching
// the local shell-mode launch.
shell: `exec ${REMOTE_LOGIN_SHELL} -i -l`,
// Mirror the LOCAL claude default so the remote agent runs non-interactively
// (no trust-folder/permission prompt that nothing on the remote answers). The
// per-host `commands.claude` override stays the escape hatch.
claude: 'exec claude --dangerously-skip-permissions',
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
claude: remoteLoginShellCommand('claude --dangerously-skip-permissions'),
opencode: remoteLoginShellCommand('opencode'),
codex: remoteLoginShellCommand('codex'),
gemini: remoteLoginShellCommand('gemini'),
antigravity: remoteLoginShellCommand('agy'),
};
return commands[mode as RemoteCommandMode] || commands.shell;
}
@@ -211,6 +256,69 @@ export async function checkRemoteTmuxAvailable(
}
}
/**
* The CLI binary each session mode runs on the remote host. Antigravity's
* binary is `agy` (the mode name is not the command); shell has no CLI to
* probe, so it is absent.
*/
const REMOTE_CLI_BIN: Partial<Record<SessionMode, string>> = {
claude: 'claude',
opencode: 'opencode',
codex: 'codex',
gemini: 'gemini',
antigravity: 'agy',
};
/**
* Build the SSH command that reads the remote CLI's version (`claude --version`
* on the remote host). The version query is routed through
* `remoteLoginShellCommand` (the SAME `$SHELL -i -l -c` wrapper the real
* launch uses), because agent CLIs live on PATH only after the remote user's
* interactive-login startup files run (see defaultRemoteCommandForMode); a bare
* `claude --version` over ssh exits 127. Connection options come from the
* shared `buildSshConnectionArgs`, so the probe reaches exactly the hosts the
* launch can reach. Returns null for modes with no CLI (shell).
*/
export function buildRemoteCliVersionProbeCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): string | null {
const bin = REMOTE_CLI_BIN[mode];
if (!bin) return null;
return [
...buildSshConnectionArgs(host),
remoteSshTarget(host),
shellescape(remoteLoginShellCommand(`${bin} --version`)),
].join(' ');
}
/**
* Read the CLI version installed ON THE REMOTE HOST. Feeds Session.cliVersion
* for remote sessions: the deterministic local probe deliberately skips them
* (it would report the LOCAL host's claude), and the startup-banner scrape is
* unreliable (newer Claude Code builds print no banner; resumed sessions never
* do), which left cliVersion undefined and silently disabled wheel-forwarding
* to the CLI transcript (residual #154, noted in the #205 analysis). The
* version is parsed as the first semver in stdout, never raw output: an
* interactive-login shell may echo rc-file noise around it. Returns undefined
* on any failure. No-op under VITEST (mirrors checkRemoteTmuxAvailable).
*/
export async function probeRemoteCliVersion(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): Promise<string | undefined> {
if (process.env.VITEST) return undefined;
const command = buildRemoteCliVersionProbeCommand(host, mode);
if (!command) return undefined;
try {
const { stdout } = await execAsync(command, { timeout: 15_000 });
const match = stdout.match(/\d+\.\d+\.\d+/);
return match ? match[0] : undefined;
} catch {
return undefined;
}
}
/**
* COD-105 — build the SSH command that lists `codeman-*` tmux sessions on a
* remote host's canonical `-L codeman` socket.
+13 -2
View File
@@ -226,6 +226,16 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
// different UUID. Backfill from the already-passed history: first try the claudeSessionId
// join, then the newest transcript in the same workingDir. Never overwrite a non-empty
// firstPrompt (so rows keyed to their own transcript are untouched).
//
// The workingDir guess is a last resort and MUST be skipped for any item that
// already has its own 'history' entry (step 1 above already gave it a real,
// direct scan of its own transcript). Without this guard, a history row whose
// OWN extraction genuinely failed (oversized first message, etc.) silently
// inherited the newest OTHER session's opening line from the same directory —
// not a blank, but actively wrong: old sessions displayed today's conversation
// as if it were their own. A row with no 'history' source at all (its
// transcript hasn't been linked/scanned under its own id yet) has no such
// direct attempt to prefer, so the guess remains a reasonable stand-in there.
const firstPromptByUuid = new Map<string, string>();
const firstPromptByWorkingDir = new Map<string, { prompt: string; ms: number }>();
// COD-145: lastPrompt rides the same backfill (build parallel indexes; never overwrite).
@@ -254,12 +264,13 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
}
}
for (const item of map.values()) {
const hasOwnHistoryEntry = item.sources.includes('history');
if (!item.firstPrompt) {
// never overwrite an existing non-empty prompt
const byUuid = item.claudeSessionId ? firstPromptByUuid.get(item.claudeSessionId) : undefined;
if (byUuid) {
item.firstPrompt = byUuid;
} else if (item.workingDir) {
} else if (item.workingDir && !hasOwnHistoryEntry) {
const byDir = firstPromptByWorkingDir.get(item.workingDir);
if (byDir) item.firstPrompt = byDir.prompt;
}
@@ -268,7 +279,7 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
const byUuid = item.claudeSessionId ? lastPromptByUuid.get(item.claudeSessionId) : undefined;
if (byUuid) {
item.lastPrompt = byUuid;
} else if (item.workingDir) {
} else if (item.workingDir && !hasOwnHistoryEntry) {
const byDir = lastPromptByWorkingDir.get(item.workingDir);
if (byDir) item.lastPrompt = byDir.prompt;
}
+6 -2
View File
@@ -121,7 +121,10 @@ export function buildClaudeEnv(sessionId: string): Record<string, string | undef
// Inform Claude it's running within Codeman (helps prevent self-termination)
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// CODEMAN_API_URL rides in via the process.env spread when the server has
// stamped it (WebServer.start()); no fallback: a hardcoded one was the wrong
// scheme on HTTPS installs, and a present-with-undefined key would serialize
// as the literal "CODEMAN_API_URL=undefined" (COD-115).
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
@@ -179,7 +182,8 @@ export function buildShellEnv(sessionId: string): Record<string, string | undefi
TERM: 'xterm-256color',
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// CODEMAN_API_URL rides in via the process.env spread when set; no fallback
// (same reasoning as buildClaudeEnv above).
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
+206 -66
View File
@@ -49,10 +49,12 @@ import {
type CodexConfig,
type EffortLevel,
type GeminiConfig,
type AntigravityConfig,
type SessionRemote,
type SessionDocker,
} from './types.js';
import { probeDockerCliVersion } from './docker-hosts.js';
import { probeRemoteCliVersion } from './remote-hosts.js';
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
@@ -65,6 +67,9 @@ import {
MAX_SESSION_TOKENS,
execPattern,
getClaudeCliVersion,
getClaudeBinaryPath,
spawnPtyWithHelperRepair,
resolveLocalShell,
} from './utils/index.js';
import {
MAX_TERMINAL_BUFFER_SIZE,
@@ -140,7 +145,7 @@ const NEWLINE_SPLIT_PATTERN = /\r?\n/;
/** True for external-CLI run modes (non-Claude) that use their own TUI and output format. */
export function isExternalCliMode(mode: SessionMode): boolean {
return mode === 'opencode' || mode === 'codex' || mode === 'gemini';
return mode === 'opencode' || mode === 'codex' || mode === 'gemini' || mode === 'antigravity';
}
function getModeLabel(mode: SessionMode): string {
@@ -151,6 +156,8 @@ function getModeLabel(mode: SessionMode): string {
return 'Codex';
case 'gemini':
return 'Gemini';
case 'antigravity':
return 'Antigravity';
case 'shell':
return 'Shell';
case 'claude':
@@ -174,6 +181,37 @@ export function isAltScreenStripMode(mode: SessionMode): boolean {
return mode === 'codex' || mode === 'claude' || mode === 'gemini';
}
/**
* Modes that need the NARROW strip: alt-screen toggles only, leaving `\x1b[3J`
* and the mouse-tracking DECSETs alone. Applies to every mode `isAltScreenStripMode`
* excludes, but ONLY when the session is tmux-backed (`useMux`).
*
* The bug (issue #205): the tmux CLIENT emits `smcup` (`\x1b[?1049h`) as its first
* bytes on attach, before any program has run. Unstripped, xterm.js parks in the
* alternate buffer for the whole session, where `baseY` is pinned at 0 (no
* scrollback to reach, so touch scrolling is a no-op) and xterm's own wheel handler
* translates the wheel into `\x1bOA`/`\x1bOB` cursor keys — which readline receives
* as shell history navigation. Both reported symptoms, one sequence.
*
* Why this is safe under tmux, despite the old "shell must keep the alt screen for
* vim/less/htop" reasoning: tmux is a full terminal emulator and NEVER forwards a
* pane's alt-screen toggles to its client, it repaints instead. Captured from a real
* attach, `\x1b[?1049h` appears exactly once (at attach) and vim/less/htop sessions
* inside the pane emit zero. So the only thing stripped here is tmux's own smcup.
*
* Why it is gated on `useMux`: `startShell()`/`startInteractive()` fall back to a
* DIRECT PTY when mux creation fails. There the inner program's `\x1b[?1049h` really
* does reach xterm, and stripping it would break vim/less/htop for real.
*
* Why it is narrower than the full strip: with tmux `mouse off`, a mouse-aware
* program in the pane (htop, vim with `set mouse=a`) still gets its DECSETs passed
* through to the client, so stripping those would break its mouse support. And
* `\x1b[3J` from a user's own `clear` is a deliberate "wipe my scrollback".
*/
export function isMuxAltScreenOnlyStripMode(mode: SessionMode, useMux: boolean): boolean {
return useMux && !isAltScreenStripMode(mode);
}
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
/** PTY fallback geometry when tmux can't be queried (matches pre-#80 hardcoded values). */
@@ -189,6 +227,8 @@ const IS_TEST_MODE = !!process.env.VITEST;
const TEST_PTY_SCRIPT = 'if (process.stdin.isTTY) process.stdin.setRawMode(true); process.stdin.pipe(process.stdout);';
/** Delay before the in-container Claude CLI version probe (lets the container start). */
const DOCKER_CLI_VERSION_PROBE_DELAY_MS = 3000;
/** Delay before the over-ssh Claude CLI version probe (keeps session start off the ssh round-trip). */
const REMOTE_CLI_VERSION_PROBE_DELAY_MS = 3000;
/**
* Ask tmux for the current window geometry of `muxName` so a re-attaching PTY
@@ -403,6 +443,8 @@ export class Session extends EventEmitter {
private _codexConfig: CodexConfig | undefined;
// Gemini configuration (only for mode === 'gemini')
private _geminiConfig: GeminiConfig | undefined;
// Antigravity configuration (only for mode === 'antigravity')
private _antigravityConfig: AntigravityConfig | undefined;
private _resumeSessionId: string | undefined;
// Ephemeral env overrides (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Exported by tmux
@@ -489,6 +531,8 @@ export class Session extends EventEmitter {
codexConfig?: CodexConfig;
/** Gemini configuration (only for mode === 'gemini') */
geminiConfig?: GeminiConfig;
/** Antigravity configuration (only for mode === 'antigravity') */
antigravityConfig?: AntigravityConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
/** Extra env vars exported to the CLI at spawn time (no disk persistence) */
@@ -499,6 +543,8 @@ export class Session extends EventEmitter {
tmuxHistoryLimit?: number;
/** Restored per-session attachment history. May include server-private external paths. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/** Restored wall-clock ms of the pane's last Enter (see `lastSubmitAt`). */
lastSubmitAt?: number;
/** Remote execution metadata for sessions launched through SSH inside local tmux. */
remote?: SessionRemote;
/** Docker execution metadata for sessions launched inside a container via local tmux. */
@@ -525,6 +571,12 @@ export class Session extends EventEmitter {
this._lastActivityAt = this.createdAt;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = config.resumeSessionId || this.id;
// Restored from state.json on boot recovery. start() resets _claudeSessionId
// to the launch id even when re-attaching to a mux session whose CLI has
// moved on (a `/clear` before the restart), so this anchor is what lets the
// response viewer re-derive the live conversation without waiting for the
// user to type again.
this._lastSubmitAt = config.lastSubmitAt ?? 0;
this._mux = config.mux || null;
this._useMux = config.useMux ?? (this._mux !== null && this._mux.isAvailable());
this._muxSession = config.muxSession || null;
@@ -562,6 +614,11 @@ export class Session extends EventEmitter {
this._geminiConfig = config.geminiConfig;
}
// Apply Antigravity configuration
if (config.antigravityConfig) {
this._antigravityConfig = config.antigravityConfig;
}
// Apply env overrides (exported at spawn, not persisted to disk).
// Legacy migration: pre-0.7.2 carried effort as the CLAUDE_CODE_EFFORT_LEVEL env var,
// which hard-locks /effort switching. Extract it into _effort (--settings soft default)
@@ -710,6 +767,15 @@ export class Session extends EventEmitter {
return this._muxSession?.muxName ?? null;
}
/**
* True when this session's PTY is a tmux client rather than the program itself.
* Read by the replay-side alt-screen strip, which must apply the same
* `useMux` gate as the live strip (isMuxAltScreenOnlyStripMode).
*/
get usesMux(): boolean {
return this._useMux;
}
get totalCost(): number {
return this._totalCost;
}
@@ -1111,6 +1177,7 @@ export class Session extends EventEmitter {
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
resumeSessionId: this._resumeSessionId,
effort: this._effort,
// COD-118: runtime-only — surfaced so the frontend can require explicit user
@@ -1119,6 +1186,7 @@ export class Session extends EventEmitter {
// recovery can re-attach.
respawnBlocked: this._respawnBlocked || undefined,
attachmentHistory: this.attachmentHistory.length > 0 ? this.attachmentHistory : undefined,
lastSubmitAt: this._lastSubmitAt || undefined,
// envOverrides intentionally NOT on the public SessionState type — they must not
// leak into SSE / GET /api/sessions broadcasts (schema allows OPENCODE_*, which
// can carry secrets). For disk persistence, session-manager calls
@@ -1271,15 +1339,17 @@ export class Session extends EventEmitter {
const attachCommand = IS_TEST_MODE ? process.execPath : mux.getAttachCommand();
const attachArgs = IS_TEST_MODE ? ['-e', TEST_PTY_SCRIPT] : mux.getAttachArgs(this._muxSession!.muxName);
try {
this.ptyProcess = pty.spawn(attachCommand, attachArgs, {
name: 'xterm-256color',
cols: ptyCols,
rows: ptyRows,
cwd: resolveMuxAttachCwd(this.workingDir, this._remote, this._docker),
// COD-75: codex/gemini get COLORTERM=truecolor — mirrors buildEnvExports()
// in tmux-manager.ts so the attach client and the tmux session agree.
env: buildMuxAttachEnv(this.mode === 'codex' || this.mode === 'gemini'),
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(attachCommand, attachArgs, {
name: 'xterm-256color',
cols: ptyCols,
rows: ptyRows,
cwd: resolveMuxAttachCwd(this.workingDir, this._remote, this._docker),
// COD-75: codex/gemini/antigravity get COLORTERM=truecolor — mirrors buildEnvExports()
// in tmux-manager.ts so the attach client and the tmux session agree.
env: buildMuxAttachEnv(this.mode === 'codex' || this.mode === 'gemini' || this.mode === 'antigravity'),
})
);
} catch (spawnErr) {
console.error(`[Session] Failed to spawn PTY for ${options.spawnErrLabel}:`, spawnErr);
this.emit('error', `Failed to attach to mux session: ${spawnErr}`);
@@ -1343,6 +1413,7 @@ export class Session extends EventEmitter {
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
@@ -1373,9 +1444,16 @@ export class Session extends EventEmitter {
// SSE/WS stream carries them, keeping everything in the main buffer with
// scrollback intact. These are controlled TUIs whose cursor-positioned
// redraws overwrite only the cells they target, so non-erased rows keep
// their content. Gated to Codex/Claude (isAltScreenStripMode) — shell must
// keep the alt screen for vim/less/htop.
if (isAltScreenStripMode(this.mode)) {
// their content. Gated to Codex/Claude/Gemini (isAltScreenStripMode).
//
// Every OTHER mode (shell/opencode/antigravity) gets the NARROW strip when it
// is tmux-backed: alt-screen toggles only, because the sequence that breaks
// scrollback there is tmux's own client-side smcup at attach, not anything the
// program in the pane emitted (issue #205, see isMuxAltScreenOnlyStripMode).
// 3J and the mouse DECSETs stay, so `clear` and mouse-aware TUIs keep working.
const fullStrip = isAltScreenStripMode(this.mode);
const altOnlyStrip = !fullStrip && isMuxAltScreenOnlyStripMode(this.mode, this._useMux);
if (fullStrip || altOnlyStrip) {
// Reassemble sequences split across PTY chunk boundaries first: a chunk
// ending mid-sequence ('\x1b[?104' now, '9h' next) would slip past the
// strip below and leave xterm stuck in the scrollback-less alt buffer
@@ -1391,13 +1469,15 @@ export class Session extends EventEmitter {
data = data.slice(0, -splitTail[0].length);
if (!data) return;
}
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
// eslint-disable-next-line no-control-regex
data = data.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '');
if (fullStrip) {
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
}
}
// Scan terminal output for attachment requests. `codeman://attach?...` is an
@@ -1457,8 +1537,8 @@ export class Session extends EventEmitter {
// never show it — which left cliVersion undefined and silently disabled
// wheel-forwarding to Claude's own transcript (the only route to history in
// repaint/alt-screen mode; issue #154). Remote sessions run claude on
// another host, so a local probe wouldn't reflect their version — skip them
// and let the banner scrape handle those. Cached process-wide, best-effort.
// another host, so a local probe wouldn't reflect their version; they get
// their own over-ssh probe below. Cached process-wide, best-effort.
if (this.mode === 'claude' && !this._remote && !this._docker && !this._cliVersion) {
const probedVersion = getClaudeCliVersion();
if (probedVersion) {
@@ -1497,6 +1577,32 @@ export class Session extends EventEmitter {
}, DOCKER_CLI_VERSION_PROBE_DELAY_MS);
}
// Remote sessions run claude on ANOTHER HOST, so neither the local nor the
// docker probe applies, and the banner-scrape fallback they were left with
// is the unreliable path #154 was filed for, so remote Claude cases silently
// never got wheel-forwarding (noted in the #205 analysis). Probe over ssh,
// deferred so session start never waits on the ssh round-trip.
if (this.mode === 'claude' && this._remote && !this._cliVersion) {
const remoteMeta = this._remote;
setTimeout(() => {
if (this._isStopped || this._cliVersion) return;
void probeRemoteCliVersion(remoteMeta, this.mode)
.then((version) => {
if (!version || this._isStopped || this._cliVersion) return;
this._cliVersion = version;
this.emit('cliInfoUpdated', {
version: this._cliVersion,
model: this._cliModel,
accountType: this._cliAccountType,
latestVersion: this._cliLatestVersion,
});
})
.catch(() => {
/* best-effort */
});
}, REMOTE_CLI_VERSION_PROBE_DELAY_MS);
}
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
@@ -1515,6 +1621,7 @@ export class Session extends EventEmitter {
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
@@ -1596,18 +1703,24 @@ export class Session extends EventEmitter {
if (this.mode === 'gemini') {
throw new Error('Gemini sessions require tmux. Direct PTY fallback is not supported.');
}
// Antigravity sessions require tmux for env override injection via setenv
if (this.mode === 'antigravity') {
throw new Error('Antigravity sessions require tmux. Direct PTY fallback is not supported.');
}
try {
// Pass --session-id to use the SAME ID as the Codeman session
// This ensures subagents can be directly matched to the correct tab
const args = buildInteractiveArgs(this.id, this._claudeMode, this._model, this._allowedTools, this._effort);
this.ptyProcess = pty.spawn('claude', args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(getClaudeBinaryPath(), args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
})
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY:', spawnErr);
this._status = 'stopped';
@@ -1873,8 +1986,9 @@ export class Session extends EventEmitter {
this._resetBuffers();
// Use user's default shell or bash
const shell = process.env.SHELL || '/bin/bash';
// Use user's default shell, falling back to a shell that actually exists.
// Shared with the tmux pane command so both paths launch the same binary.
const shell = resolveLocalShell();
console.log(
'[Session] Starting shell session with:',
shell + (this._useMux ? ` (with ${this._mux!.backend})` : '')
@@ -1930,13 +2044,15 @@ export class Session extends EventEmitter {
// Fallback to direct PTY if mux is not used
if (!this.ptyProcess) {
try {
this.ptyProcess = pty.spawn(shell, [], {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildShellEnv(this.id),
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(shell, [], {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
env: buildShellEnv(this.id),
})
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn shell PTY:', spawnErr);
this._status = 'stopped';
@@ -2036,14 +2152,16 @@ export class Session extends EventEmitter {
const args = buildPromptArgs(prompt, model, this._claudeMode, this._allowedTools);
try {
this.ptyProcess = pty.spawn('claude', args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
});
this.ptyProcess = spawnPtyWithHelperRepair(() =>
pty.spawn(getClaudeBinaryPath(), args, {
name: 'xterm-256color',
cols: 120,
rows: 40,
cwd: this.workingDir,
// Merge envOverrides after buildClaudeEnv so user settings shadow defaults.
env: { ...buildClaudeEnv(this.id), ...(this._envOverrides ?? {}) },
})
);
} catch (spawnErr) {
console.error('[Session] Failed to spawn Claude PTY for runPrompt:', spawnErr);
this.emit(
@@ -2502,36 +2620,41 @@ export class Session extends EventEmitter {
* For interactive sessions, this is how you send user input to Claude.
* Remember to include `\r` (carriage return) to simulate pressing Enter.
*
* @param data - The input data to send (text, escape sequences, etc.)
*
* @example
* ```typescript
* session.write('hello world'); // Text only, no Enter
* session.write('\r'); // Enter key
* session.write('ls -la\r'); // Command with Enter
* ```
*
* @param data - The input data to send (text, escape sequences, etc.)
* @returns true if the data reached a PTY. A session whose PTY is gone still
* discards the data, but it used to do so with no signal at all — which is how
* input could disappear while the caller believed it had been delivered.
*/
write(data: string): void {
this._trackCodexSubmit(data);
if (this.ptyProcess) {
this.ptyProcess.write(data);
}
write(data: string): boolean {
this._trackSubmit(data);
if (!this.ptyProcess) return false;
this.ptyProcess.write(data);
return true;
}
// ── Codex thread tracking ─────────────────────────────────────────────
// When a codex pane last submitted a message (Enter). The response-viewer
// correlates this against ~/.codex/history.jsonl entry timestamps to find
// the thread the pane is ACTUALLY on — the only signal that survives
// /resume, /new and /fork typed inside the codex TUI itself.
private _codexLastSubmitAt = 0;
// ── Conversation tracking ─────────────────────────────────────────────
// When this pane last submitted a message (Enter). The response-viewer
// correlates this against the CLI's own history.jsonl entry timestamps to
// find the conversation the pane is ACTUALLY on — the only signal that
// survives /clear, /resume, /new and /fork typed inside the TUI itself,
// none of which announce themselves on the PTY's stdout.
private _lastSubmitAt = 0;
get codexLastSubmitAt(): number {
return this._codexLastSubmitAt;
/** Wall-clock ms of this pane's last Enter; 0 if it has never submitted. */
get lastSubmitAt(): number {
return this._lastSubmitAt;
}
private _trackCodexSubmit(data: string): void {
if (this.mode === 'codex' && (data.includes('\r') || data.includes('\n'))) {
this._codexLastSubmitAt = Date.now();
private _trackSubmit(data: string): void {
if (data.includes('\r') || data.includes('\n')) {
this._lastSubmitAt = Date.now();
}
}
@@ -2571,6 +2694,23 @@ export class Session extends EventEmitter {
return true;
}
/**
* Undo the bookkeeping of {@link shouldApplyInput} for a delivery that failed.
*
* Without this, the reliable-delivery layer guarantees exactly-once delivery of
* something that may never have been delivered: the seq is recorded as applied
* BEFORE the write is attempted, so a client retry — the very mechanism the seq
* exists for — is rejected as a duplicate and the input is lost for good.
*
* Only rolls back if `seq` is still the newest recorded one; a later input has
* already superseded it and must not be re-opened.
*/
forgetInputSeq(clientId: string, seq: number): void {
if (this._appliedInputSeq.get(clientId) === seq) {
this._appliedInputSeq.set(clientId, seq - 1);
}
}
/**
* Sends input via the terminal multiplexer's direct input mechanism.
*
@@ -2588,7 +2728,7 @@ export class Session extends EventEmitter {
* ```
*/
async writeViaMux(data: string): Promise<boolean> {
this._trackCodexSubmit(data);
this._trackSubmit(data);
if (this._mux && this._muxSession) {
return this._mux.sendInput(this.id, data);
}
+225 -49
View File
@@ -22,7 +22,8 @@
*/
import { EventEmitter } from 'node:events';
import { execSync, exec } from 'node:child_process';
import { collectDescendants } from './proc-tree.js';
import { execSync, exec, execFile } from 'node:child_process';
import { promisify } from 'node:util';
const execAsync = promisify(exec);
@@ -43,12 +44,18 @@ import {
type CodexConfig,
type EffortLevel,
type GeminiConfig,
type AntigravityConfig,
type SessionRemote,
type SessionDocker,
type DockerCommandMode,
} from './types.js';
import { buildEffortCliArgs } from './session-cli-builder.js';
import { buildSshConnectionArgs, defaultRemoteCommandForMode, remoteSshTarget } from './remote-hosts.js';
import {
buildSshConnectionArgs,
defaultRemoteCommandForMode,
remoteLoginShellCommand,
remoteSshTarget,
} from './remote-hosts.js';
import {
buildDockerBaseArgs,
buildDockerCreateArgs,
@@ -69,6 +76,9 @@ import {
resolveOpenCodeDir,
resolveCodexDir,
resolveGeminiDir,
resolveAntigravityDir,
resolveLocalShell,
loginShellArgs,
} from './utils/index.js';
import type {
TerminalMultiplexer,
@@ -91,6 +101,16 @@ import {
// ============================================================================
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
/** How long a cached process snapshot stays usable. */
const PROC_SNAPSHOT_TTL_MS = 2000;
/**
* How long the kill path waits for a fresh snapshot before giving up on it.
* Shorter than EXEC_TIMEOUT_MS on purpose: killSession has two further strategies
* (process group, tmux kill-session) and must reach them even when `ps` is wedged.
*/
const PROC_SNAPSHOT_WAIT_MS = 1500;
import { DEFAULT_TMUX_HISTORY_LIMIT, DEFAULT_TERMINAL_BUFFER_MAX_BYTES } from './config/terminal-history.js';
/**
@@ -640,6 +660,10 @@ export function buildCodexCommand(config?: CodexConfig): string {
parts.push('--dangerously-bypass-approvals-and-sandbox');
}
if (config?.animations !== undefined) {
parts.push('--config', `tui.animations=${config.animations ? 'true' : 'false'}`);
}
if (config?.model) {
const safeModel = /^[a-zA-Z0-9._\-/]+$/.test(config.model) ? config.model : undefined;
if (safeModel) parts.push('--model', safeModel);
@@ -682,6 +706,34 @@ function buildGeminiCommand(config?: GeminiConfig): string {
return parts.join(' ');
}
/**
* Build the Antigravity CLI (agy) command with appropriate flags.
*
* Unlike gemini's yolo default, `--dangerously-skip-permissions` is only added
* when the config explicitly asks for it (the frontend sends it for parity with
* Codeman's Claude default; the multi-user clamp strips it for non-granted owners,
* and an ABSENT config stays at agy's own prompting default — safe like Codex).
*/
function buildAntigravityCommand(config?: AntigravityConfig): string {
const parts = ['agy'];
if (config?.dangerouslySkipPermissions) {
parts.push('--dangerously-skip-permissions');
}
if (config?.model) {
const safeModel = /^[a-zA-Z0-9._\-/]+$/.test(config.model) ? config.model : undefined;
if (safeModel) parts.push('--model', safeModel);
}
if (config?.resumeConversationId) {
const safeId = /^[a-zA-Z0-9._-]+$/.test(config.resumeConversationId) ? config.resumeConversationId : undefined;
if (safeId) parts.push('--conversation', safeId);
}
return parts.join(' ');
}
/**
* Build the spawn command for any session mode.
* Shared by createSession() and respawnPane() to avoid duplication.
@@ -709,6 +761,7 @@ export function buildSpawnCommand(options: {
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
resumeSessionId?: string;
effort?: EffortLevel;
}): string {
@@ -739,7 +792,24 @@ export function buildSpawnCommand(options: {
if (options.mode === 'gemini') {
return buildGeminiCommand(options.geminiConfig);
}
return '$SHELL';
if (options.mode === 'antigravity') {
return buildAntigravityCommand(options.antigravityConfig);
}
// #208: NOT the literal '$SHELL'. This string is embedded in the `bash -c "…"`
// argument of the respawn-pane line, which execSync runs through `/bin/sh -c`,
// so a `$SHELL` here is expanded by the SERVER process's shell against the
// SERVER process's env — empty in containers and system systemd units, leaving
// the pane command ending in a dangling `&&` ("syntax error: unexpected end of
// file", pane dead on arrival). Resolve it in Node and quote the result.
// #209: launch it as a LOGIN shell, which is what tmux itself does for a pane
// with no `default-command`, so a Codeman shell tab matches a hand-started tmux
// one. That is what picks up /etc/profile and /etc/profile.d/* — a systemd
// --user service never sourced them, so its PATH is what every pane inherited.
// The flags come from loginShellArgs() rather than being hardcoded: they are
// appended to a path that ultimately comes from the passwd entry, and a shell
// that rejects an unknown flag exits on the spot, which is #208 all over again.
const shell = resolveLocalShell();
return `${shellescape(shell)}${loginShellArgs(shell)}`;
}
/**
@@ -811,13 +881,14 @@ export function buildRemoteLaunchCommand(options: {
// hardcoding --dangerously-skip-permissions, so a non-granted multi-user user's
// downgraded 'auto' actually reaches the remote agent (the default command otherwise
// ignored claudeMode). A per-host `commands.claude` override stays authoritative
// (admin's explicit choice). For the DEFAULT single-user config (skip), the emitted
// command is byte-identical to before. Non-claude modes are unchanged.
// (admin's explicit choice). Wrapped in `$SHELL -i -l -c` for the same reason as
// `defaultRemoteCommandForMode`: `claude` lives under a per-user PATH entry that
// only an interactive login shell resolves (see that function's comment).
const override = remote.commands?.[mode];
const modeCommand = override
? override
: mode === 'claude'
? `exec claude${buildClaudePermissionFlags(claudeMode, allowedTools)}`
? remoteLoginShellCommand(`claude${buildClaudePermissionFlags(claudeMode, allowedTools)}`)
: defaultRemoteCommandForMode(mode);
const remoteName = remoteTmuxSessionName(sessionId);
@@ -843,6 +914,24 @@ export function buildRemoteLaunchCommand(options: {
// Per-session scoped (`set -t <name>`, matching #145's hardening) so a shared
// remote tmux server's other sessions keep their own sizing behavior.
`set -t ${remoteName} window-size latest`,
// #210: keep a CRASHED pane so the failure is still on screen. Without this,
// tmux destroys the pane -> window -> session (and, being the only session,
// the whole remote server) the instant the pane command exits, which tears the
// local `ssh -t` attach down with it; reconnect's `-A` then builds a fresh
// session and the cycle can repeat as a flap loop with no evidence surviving.
// That is how the exit-127 PATH bug fixed above stayed invisible.
//
// `failed`, NOT `on`: `on` keeps the pane on a CLEAN exit too, so typing
// `exit` in a remote shell leaves a dead pane behind, the session outlives it,
// and the next launch's `-A` reattaches to that corpse ("Pane is dead (status
// 0)") instead of starting a shell — verified against a real tmux. `failed`
// keeps the pane only on a non-zero exit, which is exactly the diagnostic case.
//
// LAST in the chain on purpose: tmux aborts the remaining commands of a `\;`
// sequence once one errors (also verified), and `failed` needs tmux >= 3.2 on
// the REMOTE host. Trailing, a rejection costs only this option; leading, it
// would silently drop status/mouse/prefix/escape-time/window-size with it.
`set -t ${remoteName} remain-on-exit failed`,
].join(' \\; ');
// ssh runs its trailing args through the remote login shell, so the entire
@@ -904,7 +993,7 @@ export function dockerTmuxSessionName(sessionId: string): string {
const RESUME_ID_SAFE = /^[A-Za-z0-9._-]+$/;
/**
* Append the CLI-specific resume flag to a pane command (codex/gemini). Only fires
* Append the CLI-specific resume flag to a pane command (codex/gemini/antigravity). Only fires
* when the in-container tmux is RE-CREATED (`new-session -A` makes the flag inert
* on a live reattach), i.e. exactly when the previous live agent was lost and we
* want to resume the conversation from the bind-mounted transcript. Claude mode
@@ -917,6 +1006,8 @@ function appendResumeFlag(modeCommand: string, mode: SessionMode, resumeId: stri
return `${modeCommand} --resume ${resumeId}`;
case 'codex':
return `${modeCommand} resume ${resumeId}`;
case 'antigravity':
return `${modeCommand} --conversation ${resumeId}`;
default:
return modeCommand; // shell / opencode: no resume
}
@@ -1485,8 +1576,10 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const exports = [
'export LANG=en_US.UTF-8',
'export LC_ALL=en_US.UTF-8',
mode === 'codex' || mode === 'gemini' ? 'export COLORTERM=truecolor' : 'unset COLORTERM',
...(mode === 'codex' || mode === 'gemini' ? ['unset NO_COLOR'] : []),
mode === 'codex' || mode === 'gemini' || mode === 'antigravity'
? 'export COLORTERM=truecolor'
: 'unset COLORTERM',
...(mode === 'codex' || mode === 'gemini' || mode === 'antigravity' ? ['unset NO_COLOR'] : []),
// Stamp each Codex pane with a unique originator so the response-viewer
// can locate THIS pane's rollout exactly — codex writes the value into
// session_meta.originator of every rollout it creates. Without it,
@@ -1496,7 +1589,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
// Only exported when the server has stamped the real URL (scheme+host+port,
// set in WebServer.start()). A hardcoded fallback here exported the wrong
// scheme on HTTPS installs; leaving the variable unset makes in-session
// guards fail closed instead of curling a URL that was never right.
...(process.env.CODEMAN_API_URL ? [`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL}`] : []),
// Path only (not the secret value): hook curl commands cat the file at
// execution time, so the COD-54 hook secret stays off the command line.
`export CODEMAN_HOOK_SECRET_FILE="${dataPath('hook-secret')}"`,
@@ -1569,6 +1666,10 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const dir = resolveGeminiDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
if (mode === 'antigravity') {
const dir = resolveAntigravityDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
return { pathExport: '', dir: null };
}
@@ -1616,6 +1717,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
envOverrides,
effort,
@@ -1667,6 +1769,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
if (mode === 'gemini' && !cliDir) {
throw new Error('Gemini CLI not found. Install with: npm install -g @google/gemini-cli');
}
if (mode === 'antigravity' && !cliDir) {
throw new Error(
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash'
);
}
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
@@ -1679,6 +1786,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
effort,
});
@@ -1901,6 +2009,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
envOverrides,
effort,
@@ -1938,6 +2047,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
openCodeConfig,
codexConfig,
geminiConfig,
antigravityConfig,
resumeSessionId,
effort,
});
@@ -1999,27 +2109,102 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
// Get all child process PIDs recursively
private getChildPids(pid: number): number[] {
const pids: number[] = [];
try {
const output = execSync(`pgrep -P ${pid}`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (output) {
for (const childPid of output
.split('\n')
.map((p) => parseInt(p, 10))
.filter((p) => !Number.isNaN(p))) {
pids.push(childPid);
pids.push(...this.getChildPids(childPid));
/** One `ps` snapshot of the whole process table, cached briefly. */
private static procSnapshot: { at: number; byParent: Map<number, number[]> } | null = null;
/** Single-flight guard so a hung `ps` cannot pile up parallel refreshes. */
private static procRefresh: { started: number; promise: Promise<Map<number, number[]>> } | null = null;
/**
* Fork ONE `ps` asynchronously and cache the parent -> children map.
*
* Async on purpose: a synchronous fork here would block the event loop on every
* stats tick, and under the procfs pathology this module exists to survive,
* `execSync`'s timeout cannot return at all (spawnSync waits for the unkillable
* child) — freezing the whole server where a hung async poll only costs staleness.
*/
private static refreshProcSnapshot(): Promise<Map<number, number[]>> {
const inFlight = TmuxManager.procRefresh;
// Reuse an in-flight refresh — unless it is old enough to be presumed stuck.
if (inFlight && Date.now() - inFlight.started < EXEC_TIMEOUT_MS * 2) return inFlight.promise;
const started = Date.now();
const promise = new Promise<Map<number, number[]>>((resolve) => {
execFile('ps', ['-eo', 'pid=,ppid='], { timeout: EXEC_TIMEOUT_MS, maxBuffer: 8 * 1024 * 1024 }, (err, out) => {
if (TmuxManager.procRefresh?.started === started) TmuxManager.procRefresh = null;
if (err) {
// ANY error, not just an empty one: a timed-out or truncated `ps` yields
// partial output, and caching that as fresh would make whole subtrees
// invisible — including to the kill path. Stale beats wrong.
console.error('[TmuxManager] process snapshot failed:', err);
resolve(TmuxManager.procSnapshot?.byParent ?? new Map());
return;
}
}
} catch {
// No children or command failed
const byParent = new Map<number, number[]>();
for (const line of String(out).split('\n')) {
const parts = line.trim().split(/\s+/);
if (parts.length < 2) continue;
const pid = parseInt(parts[0], 10);
const ppid = parseInt(parts[1], 10);
if (Number.isNaN(pid) || Number.isNaN(ppid)) continue;
const list = byParent.get(ppid);
if (list) list.push(pid);
else byParent.set(ppid, [pid]);
}
TmuxManager.procSnapshot = { at: Date.now(), byParent };
resolve(byParent);
});
});
TmuxManager.procRefresh = { started, promise };
return promise;
}
/**
* Best snapshot WITHOUT forking: returns the cache, kicking off a background
* refresh when it has gone stale, and never blocks. Stats and window-title
* consumers tolerate data one interval old; nothing that KILLS may use this.
*/
private childrenByParent(): Map<number, number[]> {
const cached = TmuxManager.procSnapshot;
if (!cached || Date.now() - cached.at >= PROC_SNAPSHOT_TTL_MS) {
void TmuxManager.refreshProcSnapshot();
}
return pids;
return cached?.byParent ?? new Map();
}
/**
* Descendants from a snapshot that is not the cached one — the kill path's variant.
*
* killSession re-scans for survivors between SIGTERM and SIGKILL, and the wait in
* between (200ms) sits far inside the cache TTL (2000ms): reading the cache there
* returns the pre-SIGTERM state verbatim, so children spawned since are invisible
* and SIGKILL aims at stale PIDs, guarded only by kill(pid, 0) — which cannot
* detect PID reuse.
*
* It forces a refresh rather than guaranteeing recency: an already-running refresh
* is reused, so the snapshot can predate this call by up to one `ps` runtime. A
* strict postdate guarantee would mean chaining a second `ps` behind every
* in-flight one, which is the fork storm this code exists to avoid.
*
* Bounded by design: waiting forever would freeze killSession before it reaches
* its process-group and tmux fallbacks.
*/
private async getChildPidsFresh(pid: number): Promise<number[]> {
let byParent: ReadonlyMap<number, readonly number[]>;
try {
byParent = await Promise.race([
TmuxManager.refreshProcSnapshot(),
new Promise<never>((_, reject) =>
setTimeout(() => reject(new Error('proc snapshot timeout')), PROC_SNAPSHOT_WAIT_MS)
),
]);
} catch {
console.warn('[TmuxManager] process snapshot did not return in time; using the cached one');
byParent = TmuxManager.procSnapshot?.byParent ?? new Map<number, number[]>();
}
return collectDescendants(pid, byParent, {
onTruncated: (root, cap, reason) =>
console.warn(`[TmuxManager] descendant walk for ${root} hit the ${cap}-${reason} cap; truncating`),
});
}
// Check if a process is still alive
@@ -2126,7 +2311,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const allPids: number[] = [currentPid];
// Strategy 1: Kill all child processes recursively
let childPids = this.getChildPids(currentPid);
let childPids = await this.getChildPidsFresh(currentPid);
if (childPids.length > 0) {
console.log(`[TmuxManager] Found ${childPids.length} child processes to kill`);
allPids.push(...childPids);
@@ -2143,7 +2328,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
await new Promise((resolve) => setTimeout(resolve, TMUX_KILL_WAIT_MS));
childPids = this.getChildPids(currentPid);
childPids = await this.getChildPidsFresh(currentPid);
for (const childPid of childPids) {
if (this.isProcessAlive(childPid)) {
try {
@@ -2357,17 +2542,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const [rss, cpu] = psOutput.split(/\s+/).map((x) => parseFloat(x) || 0);
// From the shared snapshot: this runs per session on every stats tick, and a
// pgrep per session was a fork per session per interval.
let childCount = 0;
try {
const childOutput = (
await execAsync(`pgrep -P ${session.pid} | wc -l`, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
})
).stdout.trim();
childCount = parseInt(childOutput, 10) || 0;
childCount = (this.childrenByParent().get(session.pid) ?? []).length;
} catch {
// No children or command failed
// No children or snapshot unavailable
}
return {
@@ -2401,17 +2582,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// Step 1: Get descendant PIDs
const descendantMap = new Map<number, number[]>();
const pgrepOutput = (
await execAsync(
`for p in ${sessionPids.join(' ')}; do children=$(pgrep -P $p 2>/dev/null | tr '\\n' ','); echo "$p:$children"; done`,
{
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}
)
).stdout.trim();
// Derived from the ONE snapshot instead of a shell loop that forks a pgrep
// per session — the shape that turned into a fork storm under load.
const byParent = this.childrenByParent();
const childLines = sessionPids.map((p) => `${p}:${(byParent.get(p) ?? []).join(',')}`).join('\n');
for (const line of pgrepOutput.split('\n')) {
for (const line of childLines.split('\n')) {
const [pidStr, childrenStr] = line.split(':');
const sessionPid = parseInt(pidStr, 10);
if (!Number.isNaN(sessionPid)) {
+33 -16
View File
@@ -40,6 +40,7 @@ interface TranscriptContentBlock {
text?: string;
name?: string;
input?: Record<string, unknown>;
tool_use_id?: string;
content?: string;
is_error?: boolean;
}
@@ -328,10 +329,7 @@ export class TranscriptWatcher extends EventEmitter {
this.handleResultEntry(entry);
break;
case 'user':
// User message means new turn, reset some state
this.state.isComplete = false;
this.state.hasError = false;
this.state.errorMessage = null;
this.handleUserEntry(entry);
break;
case 'system':
// System messages are informational
@@ -360,23 +358,42 @@ export class TranscriptWatcher extends EventEmitter {
this.state.currentTool = block.name;
this.emit('transcript:tool_start', block.name);
} else if (block.type === 'tool_result') {
// Tool completed
const wasError = block.is_error === true;
const toolName = this.state.currentTool;
this.state.toolExecuting = false;
this.state.currentTool = null;
if (toolName) {
this.emit('transcript:tool_end', toolName, wasError);
}
if (wasError && block.content) {
this.state.hasError = true;
this.state.errorMessage = String(block.content).slice(0, 200);
}
this.handleToolResult(block);
}
}
}
}
private handleUserEntry(entry: TranscriptEntry): void {
// A user-authored prompt starts a turn, while Claude tool results also use
// user entries. Reset turn state first, then close any completed tool.
this.state.isComplete = false;
this.state.hasError = false;
this.state.errorMessage = null;
const content = entry.message?.content;
if (!Array.isArray(content)) return;
for (const block of content) {
if (block.type === 'tool_result') {
this.handleToolResult(block);
}
}
}
private handleToolResult(block: TranscriptContentBlock): void {
const wasError = block.is_error === true;
const toolName = this.state.currentTool;
this.state.toolExecuting = false;
this.state.currentTool = null;
if (toolName) {
this.emit('transcript:tool_end', toolName, wasError);
}
if (wasError && block.content) {
this.state.hasError = true;
this.state.errorMessage = String(block.content).slice(0, 200);
}
}
private handleResultEntry(entry: TranscriptEntry): void {
// Result entry indicates completion
this.state.isComplete = true;
+5 -20
View File
@@ -15,10 +15,8 @@
import { EventEmitter } from 'node:events';
import { spawn, type ChildProcess } from 'node:child_process';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { randomBytes } from 'node:crypto';
import { resolveCloudflaredPath } from './utils/cloudflared-resolver.js';
import {
QR_TOKEN_TTL_MS,
QR_TOKEN_GRACE_MS,
@@ -95,23 +93,10 @@ export class TunnelManager extends EventEmitter {
private resolveCloudflared(): string | null {
if (this.cloudflaredPath) return this.cloudflaredPath;
// Check ~/.local/bin first (common user install location)
const localBin = join(homedir(), '.local', 'bin', 'cloudflared');
if (existsSync(localBin)) {
this.cloudflaredPath = localBin;
return localBin;
}
// Check /usr/local/bin
const usrLocalBin = '/usr/local/bin/cloudflared';
if (existsSync(usrLocalBin)) {
this.cloudflaredPath = usrLocalBin;
return usrLocalBin;
}
// Fall back to PATH
this.cloudflaredPath = 'cloudflared';
return 'cloudflared';
// Shared with the welcome-screen availability check, so the button and the
// spawn can never disagree about where cloudflared lives.
this.cloudflaredPath = resolveCloudflaredPath() ?? 'cloudflared';
return this.cloudflaredPath;
}
/** Clear all pending timers */
+16
View File
@@ -97,6 +97,22 @@ export interface FilesystemBrowseData {
truncated: boolean;
}
/** Response payload for `PUT /api/sessions/:id/file-content` (File Viewer edit mode). */
export interface FileWriteData {
/** Workspace-relative path as submitted */
path: string;
/** Size of the written content in bytes */
size: number;
/** mtime of the file after the write */
mtimeMs: number;
/** sha256 hex of the written bytes — the client's next baseHash */
hash: string;
/** Line count of the written content */
totalLines: number;
/** Line-ending style that was applied */
eol: 'lf' | 'crlf';
}
export type CleanupResourceType = 'timer' | 'interval' | 'watcher' | 'listener' | 'stream';
/**
+34 -4
View File
@@ -8,12 +8,13 @@
* - SessionConfig — creation-time config (id, workingDir, createdAt)
* - SessionOutput — captured stdout/stderr/exitCode
* - SessionStatus — 'idle' | 'busy' | 'stopped' | 'error'
* - SessionMode — 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' (which CLI backend)
* - SessionMode — 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity' (which CLI backend)
* - ClaudeMode — CLI permission mode ('dangerously-skip-permissions' | 'auto' | 'normal' | 'allowedTools')
* - SessionColor — visual differentiation color
* - OpenCodeConfig — OpenCode-specific settings (model, autoAllowTools, continueSession)
* - CodexConfig — Codex (OpenAI CLI)-specific settings (model, resumeSessionId)
* - GeminiConfig — Gemini CLI-specific settings (model, approvalMode, resumeSession)
* - AntigravityConfig — Antigravity CLI (agy) settings (model, dangerouslySkipPermissions, resumeConversationId)
*
* Cross-domain relationships:
* - SessionState.respawnConfig embeds RespawnConfig (respawn domain)
@@ -42,9 +43,12 @@ export type SessionStatus = 'idle' | 'busy' | 'stopped' | 'error';
export type ClaudeMode = 'dangerously-skip-permissions' | 'auto' | 'normal' | 'allowedTools';
/** Session mode: which CLI backend a session runs */
export type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini';
export type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity';
export type RemoteCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
export type RemoteCommandMode = Extract<
SessionMode,
'shell' | 'claude' | 'opencode' | 'codex' | 'gemini' | 'antigravity'
>;
/**
* Advanced SSH connection options shared by RemoteHost and SessionRemote.
@@ -150,7 +154,10 @@ export interface RemoteSessionInfo {
// into the same long-lived container. See `docs/docker-cases-plan.md`.
/** Which CLI backends a Docker case can run (same set as remote). */
export type DockerCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
export type DockerCommandMode = Extract<
SessionMode,
'shell' | 'claude' | 'opencode' | 'codex' | 'gemini' | 'antigravity'
>;
/** Container engine. Docker and Podman differ in the uid/userns + host-gateway alias. */
export type DockerEngine = 'docker' | 'podman';
@@ -298,6 +305,8 @@ export interface CodexConfig {
resumeSessionId?: string;
/** Bypass approval prompts (passes --dangerously-bypass-approvals-and-sandbox) */
dangerouslyBypassApprovals?: boolean;
/** Enable Codex's decorative TUI animations. Disable to reduce remote terminal redraws. */
animations?: boolean;
/** Browser rendering strategy for Codex sessions. Hybrid TUI is the only supported mode. */
renderMode?: CodexRenderMode;
}
@@ -312,6 +321,16 @@ export interface GeminiConfig {
resumeSession?: string;
}
/** Antigravity CLI (agy) session configuration */
export interface AntigravityConfig {
/** Model identifier. Passed via --model. */
model?: string;
/** Auto-approve all tool permission requests (passes --dangerously-skip-permissions). Absent = agy's default prompting. */
dangerouslySkipPermissions?: boolean;
/** Resume a previous conversation by ID (passed via --conversation). */
resumeConversationId?: string;
}
/**
* Configuration for creating a new session
*/
@@ -453,12 +472,23 @@ export interface SessionState {
codexConfig?: CodexConfig;
/** Gemini-specific configuration (only for mode === 'gemini') */
geminiConfig?: GeminiConfig;
/** Antigravity-specific configuration (only for mode === 'antigravity') */
antigravityConfig?: AntigravityConfig;
/** Claude conversation session ID to resume after reboot (set by restore script) */
resumeSessionId?: string;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/** Sanitized per-session attachment history. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/**
* Wall-clock ms of this pane's last Enter (Session.lastSubmitAt). Persisted
* because it is the response-viewer's only anchor for re-deriving the pane's
* live conversation after a Codeman restart: `start()` resets
* `claudeSessionId` to the launch id even when re-attaching to a mux session
* whose CLI has since moved on via `/clear`, and the correlation cannot run
* again until the pane's own Enter is known.
*/
lastSubmitAt?: number;
/**
* PTY-exit circuit breaker tripped — respawn blocked until an explicit restart
* (COD-118). Runtime-only: never restored on boot (fresh server = fresh breaker).
+65
View File
@@ -0,0 +1,65 @@
/**
* @fileoverview Resolve the Antigravity CLI (`agy`) binary across common install paths.
*
* Mirrors gemini-cli-resolver.ts. Google's installer (antigravity.google/cli/install.sh)
* places the binary at ~/.local/bin/agy; the other locations cover manual installs.
*
* @module utils/antigravity-cli-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
/** Common directories where the Antigravity CLI binary may be installed */
const ANTIGRAVITY_SEARCH_DIRS = [
join(homedir(), '.local', 'bin'),
join(homedir(), '.antigravity', 'bin'),
'/usr/local/bin',
join(homedir(), 'bin'),
];
/** Cached directory containing the agy binary (empty string = searched but not found) */
let _antigravityDir: string | null = null;
/**
* Finds the directory containing the `agy` binary.
* Checks `which agy` first, then falls back to common install locations.
*
* @returns Directory path, or null if not found
*/
export function resolveAntigravityDir(): string | null {
if (_antigravityDir !== null) return _antigravityDir || null;
try {
const result = execSync('which agy', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
_antigravityDir = dirname(result);
return _antigravityDir;
}
} catch {
// agy not in PATH, will check common locations
}
for (const dir of ANTIGRAVITY_SEARCH_DIRS) {
if (existsSync(join(dir, 'agy'))) {
_antigravityDir = dir;
return _antigravityDir;
}
}
_antigravityDir = '';
return null;
}
/**
* Check if the Antigravity CLI is available on the system.
*/
export function isAntigravityAvailable(): boolean {
return resolveAntigravityDir() !== null;
}
+119 -25
View File
@@ -26,6 +26,15 @@ const CLAUDE_SEARCH_DIRS = [
/** Cached directory containing the claude binary (empty string = searched but not found) */
let _claudeDir: string | null = null;
/**
* Returns true if the Claude CLI binary can be located (via `which` or one of
* the common install directories). Mirrors `isGeminiAvailable`/`isAntigravityAvailable`/`isOpenCodeAvailable`/
* `isCodexAvailable` in the sibling resolvers.
*/
export function isClaudeAvailable(): boolean {
return findClaudeDir() !== null;
}
/**
* Finds the directory containing the `claude` binary.
* Checks `which claude` first, then falls back to common install locations.
@@ -59,6 +68,21 @@ export function findClaudeDir(): string | null {
return null;
}
/**
* Returns an absolute path to the `claude` binary, falling back to the bare
* name `'claude'` when it cannot be located (so PATH resolution still gets a
* chance).
*
* Preferred over passing `'claude'` to `pty.spawn()`: a PTY child resolves the
* command against the environment it is handed, and an install that lives in
* `~/.local/bin` or `~/.claude/local` is frequently absent from the PATH the
* server process inherited (issue #6).
*/
export function getClaudeBinaryPath(): string {
const dir = findClaudeDir();
return dir ? join(dir, 'claude') : 'claude';
}
/** Cached augmented PATH string */
let _augmentedPath: string | null = null;
@@ -84,12 +108,100 @@ export function getAugmentedPath(): string {
return _augmentedPath;
}
/** Cached `claude --version` result: string = version, null = probed but unavailable, undefined = not probed */
let _claudeVersion: string | null | undefined = undefined;
/**
* Cache state for the `claude --version` probe.
*
* `version` is only ever set from a SUCCESSFUL probe and then kept for the
* process lifetime (the binary can't change under a running server without a
* restart). Failures are tracked separately so they expire.
*/
export interface ClaudeVersionProbeState {
/** Successful probe result; `undefined` until one succeeds. */
version?: string;
/** Consecutive failed probes (drives the retry backoff). */
failures: number;
/** Timestamp of the most recent failed probe. */
lastFailureAt: number;
}
/** First retry window after a failed probe. */
const VERSION_PROBE_BASE_RETRY_MS = 60_000;
/** Ceiling for the doubling backoff, so a permanently missing binary settles down. */
const VERSION_PROBE_MAX_RETRY_MS = 15 * 60_000;
/**
* How long to wait before re-probing after `failures` consecutive failures:
* 1min, 2min, 4min… capped at 15min. Exported for tests.
*/
export function claudeVersionRetryDelayMs(failures: number): number {
if (failures <= 0) return 0;
return Math.min(VERSION_PROBE_BASE_RETRY_MS * 2 ** (failures - 1), VERSION_PROBE_MAX_RETRY_MS);
}
/**
* Cache policy for the version probe, pure apart from the `state` it mutates
* and the injected `probe` (exported so tests can drive it with a fake clock).
*
* Success is cached forever; FAILURE is not. That asymmetry is the fix for a
* real shipped bug: the old cache stored `null` on any exception and guarded on
* `!== undefined`, so a single failed probe — a 5s `EXEC_TIMEOUT_MS` timeout, a
* PATH-starved systemd/launchd environment, a transient fs hiccup — at the FIRST
* Claude session start left `cliVersion` undefined for EVERY Claude session
* until the server restarted. An undefined `cliVersion` silently disables
* wheel-forwarding to Claude's own transcript (`_shouldForwardWheelToApp`),
* which is the only route to history in repaint mode: a dead wheel on every
* device at once, matching the issue #205 retest reports.
*
* Retries back off so a genuinely absent binary still can't spawn a probe per
* session start.
*/
export function resolveClaudeCliVersion(
state: ClaudeVersionProbeState,
now: number,
probe: () => string | null
): string | null {
if (state.version !== undefined) return state.version;
if (state.failures > 0 && now - state.lastFailureAt < claudeVersionRetryDelayMs(state.failures)) return null;
let version: string | null = null;
try {
version = probe();
} catch {
version = null;
}
if (version) {
state.version = version;
state.failures = 0;
state.lastFailureAt = 0;
return version;
}
state.failures += 1;
state.lastFailureAt = now;
return null;
}
const _claudeVersionState: ClaudeVersionProbeState = { failures: 0, lastFailureAt: 0 };
/** One `claude --version` run. Throws on spawn/timeout failure. */
function probeClaudeCliVersion(): string | null {
const dir = findClaudeDir();
const bin = dir ? join(dir, 'claude') : 'claude';
// execFileSync (no shell) — the resolved path may contain spaces, and there
// is no untrusted input, but avoid a shell either way.
const out = execFileSync(bin, ['--version'], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
env: { ...process.env, PATH: getAugmentedPath() },
});
const match = out.match(/(\d+\.\d+\.\d+)/);
return match ? match[1] : null;
}
/**
* Returns the installed Claude CLI version (e.g. `"2.1.210"`), or null if it
* can't be determined. Runs `claude --version` once and caches the result.
* can't be determined. Runs `claude --version` at most once per successful
* resolution; failed probes retry with backoff (see `resolveClaudeCliVersion`).
*
* This is a deterministic alternative to scraping the interactive startup
* banner (`parseClaudeCodeInfo` in session.ts): newer Claude Code builds don't
@@ -98,28 +210,10 @@ let _claudeVersion: string | null | undefined = undefined;
* gated on it (e.g. wheel-forwarding to Claude's transcript — issue #154).
*/
export function getClaudeCliVersion(): string | null {
if (_claudeVersion !== undefined) return _claudeVersion;
// Keep the test suite hermetic — never spawn a real `claude` subprocess under
// vitest (matches IS_TEST_MODE in tmux-manager). Tests that need a version set
// it on the session directly.
if (process.env.VITEST) {
_claudeVersion = null;
return _claudeVersion;
}
try {
const dir = findClaudeDir();
const bin = dir ? join(dir, 'claude') : 'claude';
// execFileSync (no shell) — the resolved path may contain spaces, and there
// is no untrusted input, but avoid a shell either way.
const out = execFileSync(bin, ['--version'], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
env: { ...process.env, PATH: getAugmentedPath() },
});
const match = out.match(/(\d+\.\d+\.\d+)/);
_claudeVersion = match ? match[1] : null;
} catch {
_claudeVersion = null;
}
return _claudeVersion;
// it on the session directly. Deliberately does NOT touch the cache state:
// recording a phantom failure here would be the very poisoning this fixes.
if (process.env.VITEST) return null;
return resolveClaudeCliVersion(_claudeVersionState, Date.now(), probeClaudeCliVersion);
}
+65
View File
@@ -0,0 +1,65 @@
/**
* @fileoverview Resolve the `cloudflared` binary across common install paths.
*
* Mirrors the CLI resolvers (antigravity-cli-resolver.ts et al), for the same reason
* they exist: the welcome screen should not offer a button whose only possible
* outcome is an error toast.
*
* The search list is deliberately the SAME one `TunnelManager.resolveCloudflared()`
* has always used, and that method now delegates here so the two can never drift.
* The difference is the fallback: this module answers "is it installed?" honestly
* with null, while the tunnel manager keeps falling back to the bare name so a
* cloudflared that only exists somewhere on the tunnel process's PATH still
* starts. A stricter answer there would turn a working tunnel into a refusal.
*
* @module utils/cloudflared-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
/** Common directories where the cloudflared binary may be installed */
const CLOUDFLARED_SEARCH_DIRS = [join(homedir(), '.local', 'bin'), '/usr/local/bin'];
/** Cached path to the cloudflared binary (empty string = searched but not found) */
let _cloudflaredPath: string | null = null;
/**
* Finds the `cloudflared` binary.
*
* @returns Absolute path, or null if not found
*/
export function resolveCloudflaredPath(): string | null {
if (_cloudflaredPath !== null) return _cloudflaredPath || null;
for (const dir of CLOUDFLARED_SEARCH_DIRS) {
const candidate = join(dir, 'cloudflared');
if (existsSync(candidate)) {
_cloudflaredPath = candidate;
return candidate;
}
}
try {
const result = execSync('which cloudflared', { encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS }).trim();
if (result && existsSync(result)) {
_cloudflaredPath = result;
return result;
}
} catch {
// Not on PATH either.
}
_cloudflaredPath = ''; // mark as searched, not found
return null;
}
/**
* Check if cloudflared is available on the system.
*/
export function isCloudflaredAvailable(): boolean {
return resolveCloudflaredPath() !== null;
}
+4 -1
View File
@@ -26,7 +26,10 @@ export { isSafePushEndpoint } from './push-endpoint-validation.js';
export { stringSimilarity, fuzzyPhraseMatch, todoContentHash } from './string-similarity.js';
export { assertNever } from './type-safety.js';
export { wrapWithNice } from './nice-wrapper.js';
export { findClaudeDir, getAugmentedPath, getClaudeCliVersion } from './claude-cli-resolver.js';
export { resolveLocalShell, loginShellArgs } from './shell-resolver.js';
export { findClaudeDir, getAugmentedPath, getClaudeCliVersion, getClaudeBinaryPath } from './claude-cli-resolver.js';
export { spawnPtyWithHelperRepair } from './node-pty-repair.js';
export { resolveOpenCodeDir } from './opencode-cli-resolver.js';
export { resolveCodexDir, isCodexAvailable } from './codex-cli-resolver.js';
export { resolveGeminiDir, isGeminiAvailable } from './gemini-cli-resolver.js';
export { resolveAntigravityDir, isAntigravityAvailable } from './antigravity-cli-resolver.js';
+153
View File
@@ -0,0 +1,153 @@
/**
* @fileoverview Runtime self-heal for node-pty's macOS `spawn-helper`.
*
* node-pty@1.1.0 ships its macOS prebuilt helper as
* `prebuilds/darwin-<arch>/spawn-helper` with mode 0644 (no execute bit). On
* macOS every PTY is launched through that helper via posix_spawnp, so a
* non-executable helper turns every session start into
* `Error: posix_spawnp failed.` (issues #6 and #204). The bug is macOS-only:
* `spawn-helper` is an `OS=="mac"` gyp target and pty.cc only spawns it under
* `#if defined(__APPLE__)`, and node-pty ships no Linux prebuild, so Linux always
* compiles a correctly-permissioned helper from source.
*
* `scripts/fix-node-pty.mjs` fixes this at install time. This module is the
* safety net for installs that are already broken: the first PTY spawn that
* fails this way is repaired and retried in-process, so the user never sees a
* dead session. If the retry still fails, the thrown error carries the manual
* repair command instead of a bare "posix_spawnp failed".
*
* @module utils/node-pty-repair
*/
import { chmodSync, existsSync, readdirSync, statSync } from 'node:fs';
import { createRequire } from 'node:module';
import { dirname, join } from 'node:path';
const require = createRequire(import.meta.url);
/** The one-line fix appended to errors we could not repair automatically. */
export const SPAWN_HELPER_FIX_HINT =
'node-pty cannot execute its spawn-helper. Repair it with: npm run fix:node-pty ' +
'(or: chmod +x node_modules/node-pty/prebuilds/*/spawn-helper)';
/** Set once a repair has been attempted, so a genuinely broken install cannot chmod-storm. */
let repairAttempted = false;
/**
* True when an error is node-pty failing to launch its spawn-helper.
*
* The native throw site is `throw Napi::Error::New(napiEnv, "posix_spawnp failed.")`
* in pty.cc, reached only on Apple platforms.
*/
export function isSpawnHelperFailure(err: unknown): boolean {
const message = err instanceof Error ? err.message : String(err ?? '');
return /posix_spawnp|spawn-helper/i.test(message);
}
/** Locates the installed node-pty package root, or null when it can't be resolved. */
export function findNodePtyDir(): string | null {
// node-pty declares no "exports" map, so the package.json subpath resolves and
// lands on the package root. require.resolve('node-pty') would return
// <pkg>/lib/index.js, one level deeper than callers need.
try {
return dirname(require.resolve('node-pty/package.json'));
} catch {
/* fall through */
}
try {
return join(dirname(require.resolve('node-pty')), '..');
} catch {
return null;
}
}
/**
* Lists every `spawn-helper` present in a node-pty install.
*
* node-pty's loader checks `build/Release`, `build/Debug`, then
* `prebuilds/<platform>-<arch>`, and takes the helper from whichever directory
* the native module loaded out of, so every copy has to be executable, not just
* the one this machine happens to use.
*/
export function listSpawnHelpers(ptyDir: string): string[] {
const dirs = [join(ptyDir, 'build', 'Release'), join(ptyDir, 'build', 'Debug')];
const prebuilds = join(ptyDir, 'prebuilds');
if (existsSync(prebuilds)) {
try {
for (const entry of readdirSync(prebuilds, { withFileTypes: true })) {
if (entry.isDirectory()) dirs.push(join(prebuilds, entry.name));
}
} catch {
/* unreadable prebuilds dir: nothing to repair there */
}
}
return dirs.map((d) => join(d, 'spawn-helper')).filter((p) => existsSync(p));
}
/**
* Adds the execute bit to every `spawn-helper` missing it.
*
* @param ptyDir - node-pty package root; resolved automatically when omitted.
* @returns Paths actually changed (empty when nothing needed repair, or node-pty
* is missing, or the files are not writable).
*/
export function repairSpawnHelperPermissions(ptyDir?: string): string[] {
const dir = ptyDir ?? findNodePtyDir();
if (!dir) return [];
const repaired: string[] = [];
for (const helper of listSpawnHelpers(dir)) {
try {
const mode = statSync(helper).mode & 0o777;
if ((mode & 0o111) === 0o111) continue;
chmodSync(helper, mode | 0o755);
repaired.push(helper);
} catch {
// Read-only install (or not ours to chmod): fall through to the hint.
}
}
return repaired;
}
/** Wraps an error so the message carries the actionable repair command. */
function withFixHint(err: unknown): Error {
const message = err instanceof Error ? err.message : String(err);
return new Error(`${message}. ${SPAWN_HELPER_FIX_HINT}`, { cause: err });
}
/**
* Runs a `pty.spawn()` call, repairing a non-executable spawn-helper and
* retrying once if that is why it failed.
*
* Any unrelated spawn error is rethrown untouched, so this stays invisible on
* every platform but a broken macOS install.
*
* @param spawn - The `pty.spawn(...)` call to run.
* @param ptyDir - node-pty package root; resolved automatically when omitted.
*/
export function spawnPtyWithHelperRepair<T>(spawn: () => T, ptyDir?: string): T {
try {
return spawn();
} catch (err) {
if (!isSpawnHelperFailure(err)) throw err;
if (repairAttempted) throw withFixHint(err);
repairAttempted = true;
const repaired = repairSpawnHelperPermissions(ptyDir);
if (repaired.length === 0) throw withFixHint(err);
console.warn(`[node-pty] spawn-helper was not executable, repaired ${repaired.join(', ')} and retrying`);
try {
return spawn();
} catch (retryErr) {
throw withFixHint(retryErr);
}
}
}
/** Test seam: forget that a repair was already attempted in this process. */
export function resetSpawnHelperRepairState(): void {
repairAttempted = false;
}
+111
View File
@@ -0,0 +1,111 @@
/**
* @fileoverview Resolve a real, launchable login shell for `mode: 'shell'` sessions.
*
* The tmux pane command for a local shell session used to be the literal string
* `$SHELL`. That string is embedded in the `bash -c "…"` argument of the
* `respawn-pane` line, which `execSync` hands to `/bin/sh -c` — so `$SHELL` was
* expanded by the SERVER process's shell (not the pane's), against the SERVER
* process's env. Containers and system-level systemd units do not set `SHELL`,
* so the expansion produced an empty string and the pane command ended in a
* dangling `&&`:
*
* bash -c "cd \"/case\" && ulimit … && export … && "
* -> bash: -c: line 1: syntax error: unexpected end of file
*
* The pane then died instantly (status 2) while tmux creation itself reported
* success, which is exactly what issue #208 saw. Resolving the shell HERE, in
* Node, removes the shell-expansion layer entirely and guarantees a non-empty
* absolute path.
*
* @module utils/shell-resolver
*/
import { accessSync, constants } from 'node:fs';
import { userInfo } from 'node:os';
/** Last-resort shells, in preference order. `/bin/sh` exists on every POSIX host. */
const FALLBACK_SHELLS = ['/bin/bash', '/bin/zsh', '/bin/sh'];
/**
* Shells that exist and are executable but immediately exit — a service account's
* passwd entry commonly points at one, which would look identical to the crash
* this module exists to prevent.
*/
const NON_INTERACTIVE_SHELLS = new Set(['nologin', 'false', 'true', 'sync']);
/**
* Shells verified to accept BOTH `-i` and `-l`. Deliberately an allowlist, not a
* blocklist: a shell that rejects an unknown flag exits immediately, which is the
* dead-pane-on-arrival failure this module exists to prevent (#208). The passwd
* entry is user data and can name anything — nushell, elvish, and xonsh all take
* neither flag in this form, so they get a bare launch instead of a dead tab.
*
* csh/tcsh are excluded on purpose: tcsh honors `-l` only when it is the ONLY
* flag, so `-i -l` would silently not be a login shell there anyway.
*/
const LOGIN_FLAG_SHELLS = new Set(['sh', 'bash', 'dash', 'ash', 'zsh', 'ksh', 'ksh93', 'mksh', 'pdksh', 'fish']);
function isUsableShell(candidate: string): boolean {
if (!candidate.startsWith('/')) return false;
const base = candidate.slice(candidate.lastIndexOf('/') + 1);
if (NON_INTERACTIVE_SHELLS.has(base)) return false;
try {
accessSync(candidate, constants.X_OK);
return true;
} catch {
return false;
}
}
/**
* Resolve an absolute path to an interactive shell, preferring the user's own.
*
* Order: `$SHELL` -> the passwd entry -> `/bin/bash` -> `/bin/zsh` -> `/bin/sh`.
* Every candidate must be an absolute path to an executable that is not a
* nologin-style stub. Always returns a non-empty string.
*/
export function resolveLocalShell(): string {
const candidates: string[] = [];
const envShell = process.env.SHELL?.trim();
if (envShell) candidates.push(envShell);
try {
// Throws when the uid has no /etc/passwd entry (common for `--user` containers).
const passwdShell = userInfo().shell?.trim();
if (passwdShell) candidates.push(passwdShell);
} catch {
/* no passwd entry — fall through to the static fallbacks */
}
candidates.push(...FALLBACK_SHELLS);
for (const candidate of candidates) {
if (isUsableShell(candidate)) return candidate;
}
// Nothing was verifiable (exotic/read-restricted image). /bin/sh is still the
// best guess and is far better than emitting an empty command.
return '/bin/sh';
}
/**
* Flags that make `shellPath` a login shell, or `''` when it takes none we trust.
*
* A tmux pane already hands the shell a tty, so it is interactive with or without
* `-i` (verified: `$-` contains `i` for a bare `/bin/bash` in a pane, which is why
* `~/.bashrc` has always been sourced). The flag that actually changes anything is
* `-l`: it makes the pane a LOGIN shell, matching what tmux itself does when it
* spawns a pane with no `default-command`, and picking up the `/etc/profile` and
* `/etc/profile.d/*` PATH entries that a systemd-spawned server never sourced.
*
* `-i` is kept alongside it because for bash the two select different files —
* login reads `~/.bash_profile`, interactive-non-login reads `~/.bashrc` — and
* asking for both is the closest thing to "the shell the user actually gets".
*
* Returns a string ready to append to an already-escaped shell path.
*/
export function loginShellArgs(shellPath: string): string {
const base = shellPath.slice(shellPath.lastIndexOf('/') + 1);
return LOGIN_FLAG_SHELLS.has(base) ? ' -i -l' : '';
}
+203 -36
View File
@@ -349,6 +349,22 @@ const DEFAULT_SHORTCUTS = [
bindings: [{ modifiers: ['ctrl'], key: 'l' }],
action: 'clearTerminal',
},
{
id: 'copy-selection',
group: 'Terminal',
label: 'Copy Selection',
// Bindings match on `key`, not `code`: xterm decides which byte to emit from the
// PRODUCED character, so intercepting a physical KeyC that doesn't produce "c"
// would diverge from the chord that actually sends ^C.
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js and
// deliberately absent from SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults on a match, which would cost the user the interrupt key.
action: 'copyTerminalSelection',
},
{
id: 'increase-font',
group: 'Terminal',
@@ -495,7 +511,18 @@ class CodemanApp {
this._initGeneration = 0; // dedup concurrent handleInit calls
this._initFallbackTimer = null; // fallback timer if SSE init doesn't arrive
this._selectGeneration = 0; // cancel stale selectSession loads
this._initialFullBufferLoad = true; // first buffer load after a page load fetches full tmux scrollback (COD-47)
// Sessions whose full tmux scrollback has already been replayed this page load
// (COD-47). Tracked PER SESSION rather than as a single "first load" flag: the
// flag was consumed by whichever session auto-selected at page load, so every
// OTHER tab started life with one visible frame of history (issue #205).
this._fullHistoryLoaded = new Set();
// Cooldown per session for the scroll-to-top "load more history" re-pull.
this._fullHistoryRepullAt = new Map(); // Map<sessionId, timestamp>
this._fullHistoryRepullInFlight = false;
// Sessions whose last re-pull came back THINNER than the live buffer (a
// repaint-mode CLI pane, where tmux keeps no history of its own). The pull is
// refused for those and retried far more slowly — see _maybeRefetchFullHistory.
this._fullHistoryRepullUseless = new Set();
this.terminalLoadStates = new Map(); // Map<sessionId, { generation, phase }>
this.respawnStatus = {};
this.respawnTimers = {}; // Track timed respawn timers
@@ -800,6 +827,11 @@ class CodemanApp {
this.applyLocalization();
this.applyTabWrapSettings();
this.applyMonitorVisibility();
this._setupTabMiddleClickClose();
// Must run before the first session:created can arrive: markSessionTabEntering()
// ignores ids until this sets up its state, which is what keeps the tabs
// restored on page load from animating.
this.initEntranceAnimations?.();
// Remove mobile-init class now that JS has applied visibility settings.
// The inline <script> in <head> added this to prevent flash-of-content on mobile.
document.documentElement.classList.remove('mobile-init');
@@ -1563,6 +1595,12 @@ class CodemanApp {
this.sessionOrder.push(data.id);
this.saveSessionOrder();
}
// Idempotent per id: the POST response and the session:created event both
// land here, and a batch launched together cascades in creation order.
this.markSessionTabEntering?.(data.id);
// The pane is one shared element, so it is only marked here and played when
// this session is actually selected (see selectSession).
this.markTerminalEntering?.(data.id);
this.renderSessionTabs();
this.updateCost();
// Start stats polling when first session appears
@@ -1897,6 +1935,37 @@ class CodemanApp {
}
}
/** Build one response-viewer message so the brief and full views share markup and CSS. */
_buildResponseViewerMessage(text, role, agentLabel) {
const div = document.createElement('div');
const isUser = role === 'user';
div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');
const roleBadge = document.createElement('div');
roleBadge.className = 'rv-role ' + (isUser ? 'rv-role-user' : 'rv-role-assistant');
roleBadge.textContent = isUser ? 'You' : agentLabel;
div.appendChild(roleBadge);
const renderedText = document.createElement('div');
renderedText.className = 'rv-text';
renderedText.innerHTML = this._renderMarkdown(text);
div.appendChild(renderedText);
return div;
}
_getResponseViewerAgentLabel() {
const mode = this.sessions.get(this.activeSessionId)?.mode;
return mode === 'codex'
? 'Codex'
: mode === 'gemini'
? 'Gemini'
: mode === 'antigravity'
? 'Antigravity'
: mode === 'opencode'
? 'OpenCode'
: 'Claude';
}
async toggleResponseViewer() {
const viewer = document.getElementById('responseViewer');
const backdrop = document.getElementById('responseViewerBackdrop');
@@ -1919,7 +1988,7 @@ class CodemanApp {
// Source 2: Terminal buffer fallback — strip ANSI, drop Claude CLI chrome.
// Claude + shell only: _cleanTerminalBuffer knows Claude CLI's output, and
// shell sessions have no transcript source at all; for TUI modes
// (codex/opencode/gemini) it yields repaint garbage, so a clear
// (codex/opencode/gemini/antigravity) it yields repaint garbage, so a clear
// placeholder beats a messy screen dump there.
const sessionMode = this.sessions.get(this.activeSessionId)?.mode || 'claude';
if (!lastResponse && (sessionMode === 'claude' || sessionMode === 'shell')) {
@@ -1932,7 +2001,11 @@ class CodemanApp {
const body = document.getElementById('responseViewerBody');
if (lastResponse) {
body.innerHTML = this._renderMarkdown(lastResponse);
// Keep the brief view inside the same message wrapper as the full
// conversation view. The wrapper supplies the card, role badge and
// descendant markdown styles that direct body children do not get.
body.innerHTML = '';
body.appendChild(this._buildResponseViewerMessage(lastResponse, 'assistant', this._getResponseViewerAgentLabel()));
this._bindResponseViewerInteractions(body);
} else {
body.textContent =
@@ -1972,26 +2045,10 @@ class CodemanApp {
}
// Render conversation thread
const mode = this.sessions.get(this.activeSessionId)?.mode;
const agentLabel =
mode === 'codex' ? 'Codex' : mode === 'gemini' ? 'Gemini' : mode === 'opencode' ? 'OpenCode' : 'Claude';
const agentLabel = this._getResponseViewerAgentLabel();
body.innerHTML = '';
for (const msg of messages) {
const div = document.createElement('div');
const isUser = msg.role === 'user';
div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');
const role = document.createElement('div');
role.className = 'rv-role ' + (isUser ? 'rv-role-user' : 'rv-role-assistant');
role.textContent = isUser ? 'You' : agentLabel;
div.appendChild(role);
const text = document.createElement('div');
text.className = 'rv-text';
text.innerHTML = this._renderMarkdown(msg.text);
div.appendChild(text);
body.appendChild(div);
body.appendChild(this._buildResponseViewerMessage(msg.text, msg.role, agentLabel));
}
this._bindResponseViewerInteractions(body);
@@ -3280,6 +3337,13 @@ class CodemanApp {
}
_renderSessionTabsImmediate() {
// Same guard as renderSessionTabs()/_fullRenderSessionTabs(): the incremental
// branch below rewrites .tab-name's innerHTML, which destroys the inline rename
// <input> mid-keystroke. Guarding only the scheduler is not enough: a render
// debounced just BEFORE the rename opened still fires ~100ms later and lands
// here directly. finishRename() re-renders on both commit and cancel, so a
// render dropped here is picked back up when the rename settles.
if (this._inlineRenameActive) return;
const container = this.$('sessionTabs');
const existingTabs = container.querySelectorAll('.session-tab[data-id]');
const existingIds = new Set([...existingTabs].map(t => t.dataset.id));
@@ -3426,11 +3490,12 @@ class CodemanApp {
}
}
} else if (minimizedCount > 0 && !subagentBadgeEl) {
// Need to add badge - insert before gear icon
// Need to add badge - insert before the action-icon overlay so the
// badge stays a direct child of the tab (outside .tab-actions)
const badgeHtml = this.renderSubagentTabBadge(id, minimizedAgents);
const gearEl = tab.querySelector('.tab-gear');
if (gearEl) {
gearEl.insertAdjacentHTML('beforebegin', badgeHtml);
const actionsEl = tab.querySelector('.tab-actions');
if (actionsEl) {
actionsEl.insertAdjacentHTML('beforebegin', badgeHtml);
}
} else if (minimizedCount === 0 && subagentBadgeEl) {
// Count went to 0 - remove badge
@@ -3443,6 +3508,13 @@ class CodemanApp {
}
this.updateTabOverflowMode();
// After the wrap measurement: the `unroll` style starts tabs at max-width 0,
// so measuring mid-animation would decide the wrap on collapsed widths.
this._applyTabEntrances?.();
// Phone overview rides on this one call: every state change it cares about
// (create, delete, idle, working, exit, hook alerts via updateTabAlertFromHooks)
// already funnels through here. No-ops unless that surface is showing.
this._refreshMobileOverviewIfVisible?.();
}
// Auto-wrap desktop session tabs to a second row when they overflow one row,
@@ -3477,6 +3549,25 @@ class CodemanApp {
container.classList.toggle('tabs-auto-wrap', shouldWrap);
}
// Middle-click closes a tab, mirroring browser tab strips. Session tabs go
// through requestCloseSession (the same confirm modal as the x button), web
// tabs through closeWebviewTab (same as theirs). Delegated on the container:
// tabs are re-rendered wholesale, the container is stable.
_setupTabMiddleClickClose() {
const container = this.$('sessionTabs');
if (!container || this._tabAuxClickBound) return;
this._tabAuxClickBound = true;
container.addEventListener('auxclick', (e) => {
if (e.button !== 1) return;
const tab = e.target.closest?.('.session-tab');
if (!tab) return;
e.preventDefault();
e.stopPropagation();
if (tab.dataset.id) this.requestCloseSession(tab.dataset.id);
else if (tab.dataset.webviewId) this.closeWebviewTab?.(tab.dataset.webviewId);
});
}
_fullRenderSessionTabs() {
if (this._inlineRenameActive) return;
const container = this.$('sessionTabs');
@@ -3532,7 +3623,7 @@ class CodemanApp {
<span class="tab-status ${status}" aria-hidden="true"></span>
<span class="tab-info">
<span class="tab-name-row">
${mode === 'shell' ? '<span class="tab-mode shell" aria-hidden="true">sh</span>' : mode === 'opencode' ? '<span class="tab-mode opencode" aria-hidden="true">oc</span>' : mode === 'codex' ? '<span class="tab-mode codex" aria-hidden="true">cx</span>' : mode === 'gemini' ? '<span class="tab-mode gemini" aria-hidden="true">gm</span>' : ''}
${mode === 'shell' ? '<span class="tab-mode shell" aria-hidden="true">sh</span>' : mode === 'opencode' ? '<span class="tab-mode opencode" aria-hidden="true">oc</span>' : mode === 'codex' ? '<span class="tab-mode codex" aria-hidden="true">cx</span>' : mode === 'gemini' ? '<span class="tab-mode gemini" aria-hidden="true">gm</span>' : mode === 'antigravity' ? '<span class="tab-mode antigravity" aria-hidden="true">ag</span>' : ''}
<span class="tab-name" data-session-id="${id}">${(() => { const p = parseSessionPrefix(name); return p && p.suffix ? '<span class="tab-prefix">' + escapeHtml(p.prefix) + '</span><span class="tab-suffix">: ' + escapeHtml(p.suffix) + '</span>' : escapeHtml(name); })()}</span>
<span class="tab-detached-badge" aria-hidden="true">detached</span>
</span>
@@ -3541,9 +3632,7 @@ class CodemanApp {
${hasRunningTasks ? `<span class="tab-badge" onclick="event.stopPropagation(); app.toggleTaskPanel()" aria-label="${taskStats.running} running tasks">${taskStats.running}</span>` : ''}
${subagentBadge}
${ultracodeBadge}
<span class="tab-gear" onclick="event.stopPropagation(); app.openSessionOptions(${escapeHtml(JSON.stringify(id))})" title="Session options" aria-label="Session options" tabindex="0">&#x2699;</span>
<span class="tab-detach" onclick="event.stopPropagation(); app.detachSession(${escapeHtml(JSON.stringify(id))})" title="Open in a new window" aria-label="Open session in a new window" tabindex="0">&#x29C9;</span>
<span class="tab-close" onclick="event.stopPropagation(); app.requestCloseSession(${escapeHtml(JSON.stringify(id))})" title="Close session" aria-label="Close session" tabindex="0">&times;</span>
<span class="tab-actions"><span class="tab-gear" onclick="event.stopPropagation(); app.openSessionOptions(${escapeHtml(JSON.stringify(id))})" title="Session options" aria-label="Session options" tabindex="0">&#x2699;</span><span class="tab-detach" onclick="event.stopPropagation(); app.detachSession(${escapeHtml(JSON.stringify(id))})" title="Open in a new window" aria-label="Open session in a new window" tabindex="0">&#x29C9;</span><span class="tab-close" onclick="event.stopPropagation(); app.requestCloseSession(${escapeHtml(JSON.stringify(id))})" title="Close session" aria-label="Close session" tabindex="0">&times;</span></span>
</div>`);
_tabIdx++;
}
@@ -3569,6 +3658,9 @@ class CodemanApp {
// toggle (applyTabWrapSettings calls this) which would otherwise leave a stale
// tabs-auto-wrap class until the next content render.
this.updateTabOverflowMode();
// Newly created tabs animate in; a re-render mid-cascade resumes them rather
// than restarting, since this rebuild just destroyed the animating elements.
this._applyTabEntrances?.();
}
// Set up arrow key navigation for session tabs (accessibility)
@@ -4036,6 +4128,72 @@ class CodemanApp {
this.terminal.write('\x1b[3J\x1b[H\x1b[2J');
}
/**
* "Load more history": re-pull the whole tmux scrollback when the user scrolls up
* while already at the top of what the browser has.
*
* xterm's buffer is only ever a WINDOW onto tmux's real history, and two things
* shrink it. tmux repaints the pane rectangle instead of emitting linefeeds
* whenever output outpaces its flush interval, which OVERWRITES already-rendered
* scrollback rather than pushing rows into it (measured: a 60-line burst added 1
* row and destroyed 34, while the same 60 lines emitted slowly added all 60). And
* a tab switch replays only the visible frame. Either way tmux still holds
* everything (history-limit 100k by default), so the fix is to go ask for it with
* the same `?full=1` capture a page reload uses (issue #205).
*
* On demand rather than automatic because that capture is unbounded-ish work: at
* the default history limit it can be megabytes, which is fine to pay when the
* user is explicitly reaching for history and not fine on every tab switch.
*
* NEVER a downgrade: for a repaint-mode CLI pane tmux keeps no history of its
* own, so the capture can be THINNER than what xterm already holds and the
* reset+rewrite below would delete history mid-scroll. `_replayWouldShrinkBuffer`
* (terminal-ui.js) is the guard, and a session that produced one useless re-pull
* gets a much longer cooldown so a hollow pane stops re-fetching megabytes on
* every scroll-up (issue #205, round 2).
*/
async _maybeRefetchFullHistory() {
const sessionId = this.activeSessionId;
if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return;
if (this.detachedSessions?.has(sessionId)) return;
const now = Date.now();
// Momentum scrolling fires this dozens of times per flick, and a burst of new
// output is the normal reason to want a re-pull, so cooldown rather than latch.
const cooldown = this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000;
if (now - (this._fullHistoryRepullAt.get(sessionId) || 0) < cooldown) return;
this._fullHistoryRepullAt.set(sessionId, now);
this._fullHistoryRepullInFlight = true;
try {
const res = await fetch(`/api/sessions/${sessionId}/terminal?full=1`);
const buffer = (await res.json())?.data?.terminalBuffer;
// Bail on a tab switch mid-fetch: writing here would paint another session's
// history into the terminal the user is now looking at.
if (!buffer || this.activeSessionId !== sessionId) return;
if (this._replayWouldShrinkBuffer(buffer)) {
(this._fullHistoryRepullUseless ||= new Set()).add(sessionId);
this._logScrollRouting?.('repull-refused-downgrade');
return;
}
this._fullHistoryRepullUseless?.delete(sessionId);
const rowsBefore = this.terminal.buffer.active.length;
this._resetTerminalForReplay();
await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
if (this.activeSessionId !== sessionId) return;
this.terminalBufferCache.set(sessionId, buffer);
// Hold the user's place. The replay is a superset that grew the buffer
// UPWARD, so what used to be row 0 (what they were looking at) is now `delta`
// rows down; scrolling there reveals the recovered history above it instead
// of teleporting them to the bottom the way a normal buffer load does.
const delta = this.terminal.buffer.active.length - rowsBefore;
if (delta > 0) this.terminal.scrollToLine(delta);
else this.terminal.scrollToTop();
} catch {
// Transient (offline, 5xx) — the next scroll-up past the cooldown retries.
} finally {
this._fullHistoryRepullInFlight = false;
}
}
_shouldFocusTerminalForTabSwitch() {
if (typeof MobileDetection === 'undefined' || !MobileDetection.isTouchDevice()) {
return true;
@@ -4100,6 +4258,10 @@ class CodemanApp {
// selectSession or reconnect catches up.
this._updateSseSubscription(sessionId);
this.hideWelcome();
// Terminal-pane entrance: plays for a freshly created session, and on every
// switch when that option is on. Transform/opacity/clip-path only, xterm's
// FitAddon reads the untransformed layout box, so this cannot reach the PTY.
this.playTerminalEntrance?.(sessionId);
// Clear idle hooks on view, but keep action hooks until user interacts
this.clearPendingHooks(sessionId, 'idle_prompt');
// Instant active-class toggle (no 100ms debounce), then schedule full render for badges/status
@@ -4215,7 +4377,7 @@ class CodemanApp {
// (viewport + scrollback + colors) for an instant first paint. For codex
// this is also a correctness fix — its byte-stream replay shows only the
// latest TUI frame (the idle welcome banner) because codex doesn't include
// earlier conversation in its current redraw. For claude/opencode/gemini
// earlier conversation in its current redraw. For claude/opencode/gemini/antigravity
// the replay is already complete, so the snapshot is purely a faster,
// scroll-preserving first paint before the canonical fetch reconciles.
//
@@ -4302,11 +4464,14 @@ class CodemanApp {
this._setTerminalLoadState(sessionId, selectGen, 'fetching');
_crashDiag.log('FETCH_START');
// The FIRST buffer load after a page load requests the full tmux scrollback
// (?full=1, COD-47) so history that scrolled off the server's byte buffer
// comes back after a reload. Tab switches keep the fast ?tail= frame path.
const useFullHistory = this._initialFullBufferLoad === true;
this._initialFullBufferLoad = false;
// The first load OF EACH SESSION this page load requests the full tmux
// scrollback (?full=1, COD-47) so history that scrolled off the server's byte
// buffer comes back. Later switches to an already-replayed session keep the
// fast ?tail= frame path, which is why this is a Set and not a flag: the flag
// version gave the full replay to the auto-selected tab and one frame of
// history to every other one (issue #205).
const useFullHistory = !this._fullHistoryLoaded.has(sessionId);
if (useFullHistory) this._fullHistoryLoaded.add(sessionId);
const res = await fetch(
useFullHistory
? `/api/sessions/${sessionId}/terminal?full=1`
@@ -4610,7 +4775,9 @@ class CodemanApp {
? 'Kill Tmux & Codex'
: session.mode === 'gemini'
? 'Kill Tmux & Gemini'
: 'Kill Tmux & Claude Code';
: session.mode === 'antigravity'
? 'Kill Tmux & Antigravity'
: 'Kill Tmux & Claude Code';
}
document.getElementById('closeConfirmModal').classList.add('active');
+754
View File
@@ -0,0 +1,754 @@
/**
* @fileoverview Entrance animations for the four things that appear when work
* starts: session TABS, the main TERMINAL pane a session's CLI runs in, floating
* agent WINDOWS, and the CONNECTION LINES tying a window back to its parent tab.
* One picker per surface, plus themes that set all four to a matching look.
*
* Everything is OFF by default (the `legacy` theme), so an untouched install
* behaves exactly as it did before this module existed. Opt in via App Settings
* → Appearance → Entrance Animations.
*
* Four constraints shape the design:
*
* 1. `_fullRenderSessionTabs()` replaces the tab strip's entire innerHTML, and
* `_updateConnectionLinesImmediate()` does `svg.innerHTML = ''` and rebuilds
* every path. Both run constantly while sessions and agents are spawning, so
* an animating tab or line element is DESTROYED mid-flight. Those two are
* therefore tracked by id in `_tabEnterActive` / `_lineEnterActive` and
* re-applied to the fresh element with a NEGATIVE animation-delay, resuming at
* the same offset instead of restarting or snapping to the end. Windows are
* stable DOM and need none of this.
* 2. Connection-line geometry comes from `getBoundingClientRect()` on the window.
* A window entrance that starts with a transform would move that rect, so the
* `beam` style (which must hold still while its line draws toward it) animates
* opacity and filter only. Every other style refreshes the lines when it ends.
* 3. The terminal pane is ONE shared element, so its entrance is marked at
* session creation but played at selection: a session created in the
* background must not animate the pane the user is currently looking at. Its
* styles are also restricted to transform/opacity/clip-path (see below).
* 4. Nothing may animate on page load or reconnect replay. Only ids that pass
* through `markSessionTabEntering()` animate, and `_tabEnterSeen` makes that
* once-per-id even though the POST response and the SSE event both call
* `_onSessionCreated`.
*
* Styles are selected by `data-tab-anim` / `data-term-anim` / `data-win-anim` /
* `data-line-anim` on <html>; the keyframes live in styles.css. `?animlab=1`
* opens a floating picker that fakes tabs, a pane replay, a window and a line,
* so styles can be compared without spawning real sessions or agents.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (tab render pipeline), subagent-windows.js (window + line hooks)
* @dependency constants.js (escapeHtml)
* @loadorder 12.6 of 16, after webview-tabs.js, before ralph-wizard.js
*/
/** Tab entrance styles. `key` doubles as the `data-tab-anim` value. */
const TAB_ANIM_STYLES = [
{ key: 'slide', label: 'Slide', blurb: 'Drifts in from the right.', duration: 380 },
{ key: 'pop', label: 'Pop', blurb: 'Springs past full size, then settles.', duration: 460 },
{ key: 'crt', label: 'CRT', blurb: 'Snaps open as a hot line, then unfolds.', duration: 520 },
{ key: 'unroll', label: 'Unroll', blurb: 'The strip makes room and the tab widens in.', duration: 480 },
{ key: 'boot', label: 'Boot', blurb: 'Flickers on under a green scan sweep.', duration: 720 },
{ key: 'flip', label: 'Flip', blurb: 'Drops in as a card hinged on its top edge.', duration: 520 },
{ key: 'off', label: 'Off', blurb: 'Tabs just appear.', duration: 0 },
];
/**
* Window entrance styles. `fly` is the pre-existing behaviour (the window flies
* out of its parent tab via a JS transition in subagent-windows.js); every other
* style positions the window at its resting spot and runs a CSS animation there.
*/
const WIN_ANIM_STYLES = [
{ key: 'fly', label: 'Fly from tab', blurb: 'Current behaviour: flies out of the tab, scaling up.', duration: 400 },
{ key: 'crt', label: 'CRT', blurb: 'Bursts open as a hot line, then unfolds vertically.', duration: 560 },
{ key: 'materialize', label: 'Materialize', blurb: 'Resolves out of a blur with a short glitch.', duration: 620 },
{ key: 'unfold', label: 'Unfold', blurb: 'Hinges down from its top edge in 3D.', duration: 560 },
{ key: 'beam', label: 'Beam down', blurb: 'Waits for its line to reach it, then materializes.', duration: 620 },
{ key: 'pop', label: 'Pop', blurb: 'Springs open from its centre.', duration: 460 },
{ key: 'off', label: 'Off', blurb: 'Windows just appear.', duration: 0 },
];
/** Connection-line entrance styles. `key` doubles as the `data-line-anim` value. */
const LINE_ANIM_STYLES = [
{ key: 'draw', label: 'Draw', blurb: 'Draws itself from the tab down to the window.', duration: 420 },
{ key: 'packet', label: 'Packet', blurb: 'Line fades in, then a bright packet runs down it.', duration: 700 },
{ key: 'fade', label: 'Fade', blurb: 'Simply fades in.', duration: 300 },
{ key: 'off', label: 'Off', blurb: 'Lines just appear.', duration: 0 },
];
/**
* Main-terminal entrance styles: the pane a session's CLI actually runs in.
*
* ⚠ These may only animate transform, opacity and clip-path. xterm's FitAddon
* derives rows/cols from `getComputedStyle(parent).width/height`, which reports
* the untransformed layout box, so transforms are safe, but animating width,
* height or padding would feed wrong dimensions into `resize()` and through to
* the PTY. Colour washes go on `.terminal-container::before`, never a `filter`
* on the container: that would blur a full-screen WebGL canvas every frame.
*/
const TERM_ANIM_STYLES = [
{ key: 'crt', label: 'CRT', blurb: 'Power-on: a hot line that expands to full height.', duration: 560 },
{ key: 'boot', label: 'Boot', blurb: 'Flickers on under a green scan sweep.', duration: 760 },
{ key: 'wipe', label: 'Wipe', blurb: 'Reveals top-to-bottom behind a bright edge.', duration: 520 },
{ key: 'slide', label: 'Slide up', blurb: 'Rises into place from below.', duration: 420 },
{ key: 'fade', label: 'Fade', blurb: 'Quiet fade with a touch of scale.', duration: 340 },
{ key: 'off', label: 'Off', blurb: 'Current behaviour: the pane just appears.', duration: 0 },
];
/** How long a `beam` window waits before materializing. Just under the line draw. */
const BEAM_HOLD_MS = 360;
/** One-click combinations that read as a single look. */
const ANIM_THEMES = [
{ key: 'terminal', label: 'Terminal', tab: 'crt', win: 'crt', line: 'draw', term: 'crt' },
{ key: 'beamdown', label: 'Beam down', tab: 'crt', win: 'beam', line: 'draw', term: 'wipe' },
{ key: 'quiet', label: 'Quiet', tab: 'slide', win: 'materialize', line: 'fade', term: 'fade' },
{ key: 'playful', label: 'Playful', tab: 'pop', win: 'pop', line: 'packet', term: 'slide' },
{ key: 'legacy', label: 'Legacy', tab: 'off', win: 'fly', line: 'off', term: 'off' },
];
/**
* Defaults are the `legacy` theme: every entrance OFF, and agent windows on the
* `fly` behaviour Codeman already had before this module existed. So a user who
* never opens the picker sees exactly the pre-existing UI, and each mark/apply
* hook short-circuits on its first line. Opt in via App Settings → Appearance →
* Entrance Animations, which persists to the localStorage keys below.
*/
const TAB_ANIM_DEFAULT = 'off';
const WIN_ANIM_DEFAULT = 'fly';
const LINE_ANIM_DEFAULT = 'off';
const TERM_ANIM_DEFAULT = 'off';
const TAB_ANIM_STAGGER_DEFAULT = 90;
/** A new id joins the current cascade if it arrives within this of the last one. */
const TAB_ANIM_BATCH_WINDOW_MS = 600;
const ANIM_KEYS = {
tab: 'codeman:tabAnim',
win: 'codeman:winAnim',
line: 'codeman:lineAnim',
term: 'codeman:termAnim',
termSwitch: 'codeman:termAnimOnSwitch',
stagger: 'codeman:tabAnimStagger',
speed: 'codeman:tabAnimSpeed',
};
Object.assign(CodemanApp.prototype, {
// ── Setup ─────────────────────────────────────────────────────────────────
/** Resolve every style/timing and stamp them on <html>. Called once at startup. */
initEntranceAnimations() {
this._tabEnterSeen = new Set();
this._tabEnterActive = new Map();
this._lineEnterActive = new Map();
this._tabEnterBatchIndex = 0;
this._tabEnterLastMarkTs = 0;
const params = new URLSearchParams(location.search);
// ?tabanim= / ?winanim= / ?lineanim= win for a single load, so a style can be
// tried without persisting over whatever is saved.
const pick = (param, styles, storeKey, fallback) => {
const fromUrl = params.get(param);
if (styles.some((s) => s.key === fromUrl)) return fromUrl;
return this._animRead(storeKey, fallback);
};
this.setTabAnimStyle(pick('tabanim', TAB_ANIM_STYLES, ANIM_KEYS.tab, TAB_ANIM_DEFAULT), { persist: false });
this.setWinAnimStyle(pick('winanim', WIN_ANIM_STYLES, ANIM_KEYS.win, WIN_ANIM_DEFAULT), { persist: false });
this.setLineAnimStyle(pick('lineanim', LINE_ANIM_STYLES, ANIM_KEYS.line, LINE_ANIM_DEFAULT), { persist: false });
this.setTermAnimStyle(pick('termanim', TERM_ANIM_STYLES, ANIM_KEYS.term, TERM_ANIM_DEFAULT), { persist: false });
this.setTermAnimOnSwitch(this._animRead(ANIM_KEYS.termSwitch, '0') === '1', { persist: false });
this.setTabAnimStagger(Number(this._animRead(ANIM_KEYS.stagger, TAB_ANIM_STAGGER_DEFAULT)), { persist: false });
this.setAnimSpeed(Number(this._animRead(ANIM_KEYS.speed, 1)), { persist: false });
if (params.get('animlab') === '1' || params.get('tabanimlab') === '1') this.openAnimLab();
},
_animRead(key, fallback) {
try {
const raw = localStorage.getItem(key);
return raw === null || raw === '' ? fallback : raw;
} catch {
return fallback;
}
},
_animWrite(key, value) {
try {
localStorage.setItem(key, String(value));
} catch {
/* private mode / quota, the in-memory value still applies for this load */
}
},
_setAnimStyle(prop, key, styles, fallback, attr, storeKey, persist) {
const style = styles.some((s) => s.key === key) ? key : fallback;
this[prop] = style;
document.documentElement.setAttribute(attr, style);
if (persist) this._animWrite(storeKey, style);
// Keep the lab's highlight honest when a style is set from anywhere other
// than the lab's own buttons (a theme, a URL param, the console).
this._syncAnimLab?.();
},
setTabAnimStyle(key, { persist = true } = {}) {
// prettier-ignore
this._setAnimStyle('_tabAnimStyle', key, TAB_ANIM_STYLES, TAB_ANIM_DEFAULT, 'data-tab-anim', ANIM_KEYS.tab, persist);
},
setWinAnimStyle(key, { persist = true } = {}) {
// prettier-ignore
this._setAnimStyle('_winAnimStyle', key, WIN_ANIM_STYLES, WIN_ANIM_DEFAULT, 'data-win-anim', ANIM_KEYS.win, persist);
},
setLineAnimStyle(key, { persist = true } = {}) {
// prettier-ignore
this._setAnimStyle('_lineAnimStyle', key, LINE_ANIM_STYLES, LINE_ANIM_DEFAULT, 'data-line-anim', ANIM_KEYS.line, persist);
},
setTermAnimStyle(key, { persist = true } = {}) {
// prettier-ignore
this._setAnimStyle('_termAnimStyle', key, TERM_ANIM_STYLES, TERM_ANIM_DEFAULT, 'data-term-anim', ANIM_KEYS.term, persist);
},
/** Replay the terminal entrance on every tab switch, not just on a new session. */
setTermAnimOnSwitch(on, { persist = true } = {}) {
this._termAnimOnSwitch = !!on;
if (persist) this._animWrite(ANIM_KEYS.termSwitch, on ? '1' : '0');
this._syncAnimLab?.();
},
/** Apply a theme: one look across tabs, windows, lines and the terminal. */
setAnimTheme(themeKey) {
const theme = ANIM_THEMES.find((t) => t.key === themeKey);
if (!theme) return;
this.setTabAnimStyle(theme.tab);
this.setWinAnimStyle(theme.win);
this.setLineAnimStyle(theme.line);
this.setTermAnimStyle(theme.term);
this._syncEntranceAnimSetting?.();
},
/** The theme matching the four current styles, or 'custom' for a lab mix. */
currentAnimTheme() {
const match = ANIM_THEMES.find(
(t) =>
t.tab === this._tabAnimStyle &&
t.win === this._winAnimStyle &&
t.line === this._lineAnimStyle &&
t.term === this._termAnimStyle
);
return match ? match.key : 'custom';
},
// ── App Settings picker ───────────────────────────────────────────────────
//
// Wired straight to setAnimTheme() rather than through saveAppSettings(): the
// styles live in their own localStorage keys, so they stay per-device and never
// reach `PUT /api/settings`, whose schema is .strict() and would reject them.
_syncEntranceAnimSetting() {
const sel = document.getElementById('appSettingsEntranceAnim');
if (!sel) return;
sel.value = this.currentAnimTheme();
if (!sel.dataset.bound) {
sel.dataset.bound = '1';
sel.addEventListener('change', () => {
// 'custom' is a readout of a lab mix, not something you can select into.
if (sel.value === 'custom') sel.value = this.currentAnimTheme();
else this.setAnimTheme(sel.value);
});
}
},
setTabAnimStagger(ms, { persist = true } = {}) {
const value = Number.isFinite(ms) ? Math.max(0, Math.min(400, Math.round(ms))) : TAB_ANIM_STAGGER_DEFAULT;
this._tabAnimStagger = value;
if (persist) this._animWrite(ANIM_KEYS.stagger, value);
},
setAnimSpeed(multiplier, { persist = true } = {}) {
const value = Number.isFinite(multiplier) ? Math.max(0.25, Math.min(3, multiplier)) : 1;
this._animSpeed = value;
// Every keyframe block reads this, so one variable retimes all of them.
document.documentElement.style.setProperty('--anim-enter-scale', String(1 / value));
if (persist) this._animWrite(ANIM_KEYS.speed, value);
},
_styleDuration(styles, key) {
return (styles.find((s) => s.key === key)?.duration || 0) / (this._animSpeed || 1);
},
_tabAnimDuration() {
return this._styleDuration(TAB_ANIM_STYLES, this._tabAnimStyle);
},
_winAnimDuration() {
return this._styleDuration(WIN_ANIM_STYLES, this._winAnimStyle);
},
_lineAnimDuration() {
return this._styleDuration(LINE_ANIM_STYLES, this._lineAnimStyle);
},
_termAnimDuration() {
return this._styleDuration(TERM_ANIM_STYLES, this._termAnimStyle);
},
// ── Tabs ──────────────────────────────────────────────────────────────────
/** Queue a session id to animate on its next render. Idempotent per id. */
markSessionTabEntering(id) {
if (!id || this._tabAnimStyle === 'off') return;
if (!this._tabEnterSeen) return; // initEntranceAnimations() has not run yet
if (this._tabEnterSeen.has(id)) return;
this._tabEnterSeen.add(id);
const now = performance.now();
// A launch landing well after the previous one starts its own cascade rather
// than inheriting a large stale offset.
if (now - this._tabEnterLastMarkTs > TAB_ANIM_BATCH_WINDOW_MS) this._tabEnterBatchIndex = 0;
this._tabEnterLastMarkTs = now;
this._tabEnterActive.set(id, {
startTs: now,
staggerMs: this._tabEnterBatchIndex * this._tabAnimStagger,
});
this._tabEnterBatchIndex += 1;
},
/** Attach (or resume) the entrance animation on freshly rendered tabs. */
_applyTabEntrances() {
const active = this._tabEnterActive;
if (!active || active.size === 0) return;
if (this._tabAnimStyle === 'off') {
active.clear();
return;
}
const container = this.$('sessionTabs');
if (!container) return;
const now = performance.now();
const duration = this._tabAnimDuration();
for (const [id, state] of active) {
const elapsed = now - state.startTs;
// Ran to completion while the element was detached, nothing left to show.
if (elapsed > state.staggerMs + duration + 50) {
active.delete(id);
continue;
}
const tab = container.querySelector(`.session-tab[data-id="${CSS.escape(id)}"]`);
if (!tab) continue; // not rendered yet; a later render picks it up
// Already running on this element. Re-stamping the delay would jump it, and
// the incremental render path can reach here for the same element.
if (tab.classList.contains('tab-enter')) continue;
// Negative delay resumes mid-animation, so an unrelated re-render mid-cascade
// does not restart the tab or make it snap.
tab.style.setProperty('--tab-enter-delay', `${state.staggerMs - elapsed}ms`);
tab.classList.add('tab-enter');
const done = () => {
tab.classList.remove('tab-enter');
tab.style.removeProperty('--tab-enter-delay');
this._tabEnterActive?.delete(id);
};
tab.addEventListener('animationend', done, { once: true });
tab.addEventListener('animationcancel', done, { once: true });
}
},
// ── Main terminal ─────────────────────────────────────────────────────────
/**
* Queue the terminal pane to animate the first time this session is shown.
* Marked at creation but PLAYED at selection, because the pane is one shared
* element: a session created in the background must not animate the pane the
* user is currently looking at.
*/
markTerminalEntering(id) {
if (!id || this._termAnimStyle === 'off') return;
if (!this._termEnterPending) this._termEnterPending = new Set();
this._termEnterPending.add(id);
},
/** Play the terminal entrance for `sessionId`, if it is owed one. */
playTerminalEntrance(sessionId) {
const style = this._termAnimStyle || TERM_ANIM_DEFAULT;
if (style === 'off') return;
const owed = sessionId && this._termEnterPending?.delete(sessionId);
if (!owed && !this._termAnimOnSwitch) return;
const el = this.$('terminalContainer');
if (!el) return;
// Restart cleanly when switching tabs faster than the animation runs.
el.classList.remove('term-enter');
void el.offsetWidth;
el.classList.add('term-enter');
clearTimeout(this._termEnterTimer);
const done = () => {
el.classList.remove('term-enter');
clearTimeout(this._termEnterTimer);
};
el.addEventListener('animationend', done, { once: true });
el.addEventListener('animationcancel', done, { once: true });
// Backstop: a backgrounded tab never fires animationend, which would leave
// the pane stuck at its 0% keyframe (invisible) when you come back to it.
this._termEnterTimer = setTimeout(done, this._termAnimDuration() + 900);
},
// ── Windows ───────────────────────────────────────────────────────────────
/**
* True when the window should be parked on its parent tab and flown to its
* resting spot by subagent-windows.js. False for every CSS-animated style,
* which needs the window to start at its final position.
*/
windowEntranceFliesFromTab() {
return (this._winAnimStyle || WIN_ANIM_DEFAULT) === 'fly';
},
/**
* Run the window entrance on an already-positioned window.
* @param {HTMLElement} win a .subagent-window or .ultracode-window
*/
applyWindowEntrance(win) {
if (!win) return;
const style = this._winAnimStyle || WIN_ANIM_DEFAULT;
if (style === 'off' || style === 'fly') return;
// `beam` holds the window still (opacity/filter only) while its line draws
// toward it, then materializes. Everything else starts immediately.
if (style === 'beam') win.style.setProperty('--win-enter-delay', `${BEAM_HOLD_MS / (this._animSpeed || 1)}ms`);
win.classList.add('win-enter');
const done = () => {
win.classList.remove('win-enter');
win.style.removeProperty('--win-enter-delay');
// A transformed window reports a transformed rect, so lines drawn while it
// was animating are slightly off. Redraw once it has settled.
this.updateConnectionLines?.();
};
win.addEventListener('animationend', done, { once: true });
win.addEventListener('animationcancel', done, { once: true });
},
// ── Connection lines ──────────────────────────────────────────────────────
/** Queue an agent's connection line to draw itself on the next line rebuild. */
markConnectionLineEntering(agentId) {
if (!agentId || this._lineAnimStyle === 'off') return;
if (!this._lineEnterActive) return;
if (this._lineEnterActive.has(agentId)) return;
this._lineEnterActive.set(agentId, { startTs: performance.now() });
},
/**
* Attach (or resume) the draw-in animation on freshly rebuilt paths. Called at
* the end of `_updateConnectionLinesImmediate()`, which has just thrown away
* and recreated every path element.
*/
_applyLineEntrances(svg) {
const active = this._lineEnterActive;
if (!active || active.size === 0 || !svg) return;
const style = this._lineAnimStyle || LINE_ANIM_DEFAULT;
if (style === 'off') {
active.clear();
return;
}
const now = performance.now();
const duration = this._lineAnimDuration();
for (const [agentId, state] of active) {
const elapsed = now - state.startTs;
if (elapsed > duration + 50) {
active.delete(agentId);
continue;
}
const path = svg.querySelector(`path[data-agent-id="${CSS.escape(agentId)}"]`);
if (!path) continue;
const len = Math.max(1, Math.round(path.getTotalLength()));
path.style.setProperty('--line-len', `${len}px`);
path.style.setProperty('--line-enter-delay', `${-elapsed}ms`);
path.classList.add('line-enter');
// `packet` rides a bright dash ON TOP of the normal dashed line, so the base
// line keeps its look instead of being taken over by the animation.
if (style === 'packet') {
const packet = path.cloneNode(false);
packet.removeAttribute('data-agent-id');
packet.setAttribute('class', 'connection-line-packet');
packet.style.setProperty('--line-len', `${len}px`);
packet.style.setProperty('--line-enter-delay', `${-elapsed}ms`);
svg.appendChild(packet);
}
}
},
// ── Lab (compare styles without spawning sessions or agents) ───────────────
/** Floating picker: switch styles per surface and replay fake entrances. */
openAnimLab() {
if (document.getElementById('animLab')) return;
const group = (title, styles, attr) => `
<div class="anim-lab-group">
<div class="anim-lab-group-title">${title}</div>
${styles
.map(
(s) => `<button type="button" class="anim-lab-style" data-attr="${attr}" data-style="${s.key}">
<strong>${escapeHtml(s.label)}</strong><em>${escapeHtml(s.blurb)}</em>
</button>`
)
.join('')}
</div>`;
const panel = document.createElement('div');
panel.id = 'animLab';
panel.className = 'anim-lab';
panel.innerHTML = `
<div class="anim-lab-head">
<span>Entrance lab</span>
<button type="button" class="anim-lab-close" aria-label="Close">&times;</button>
</div>
<div class="anim-lab-themes">
${ANIM_THEMES.map((t) => `<button type="button" data-theme="${t.key}">${escapeHtml(t.label)}</button>`).join('')}
</div>
<div class="anim-lab-scroll">
${group('Tabs', TAB_ANIM_STYLES, 'tab')}
${group('Terminal pane', TERM_ANIM_STYLES, 'term')}
<label class="anim-lab-check">
<input type="checkbox" data-check="termSwitch"> Also on every tab switch
</label>
${group('Agent windows', WIN_ANIM_STYLES, 'win')}
${group('Connection lines', LINE_ANIM_STYLES, 'line')}
</div>
<label class="anim-lab-range">Tab stagger <output data-out="stagger"></output>
<input type="range" data-range="stagger" min="0" max="260" step="10">
</label>
<label class="anim-lab-range">Speed <output data-out="speed"></output>
<input type="range" data-range="speed" min="0.5" max="2" step="0.1">
</label>
<div class="anim-lab-demo">
<span>Replay</span>
<button type="button" data-demo="tabs">Tabs</button>
<button type="button" data-demo="term">Pane</button>
<button type="button" data-demo="window">Window</button>
<button type="button" data-demo="all">All</button>
</div>
<p class="anim-lab-hint">Fake tabs, window and line, removed after the run. Real launches use the same timing.</p>
`;
document.body.appendChild(panel);
panel.querySelector('.anim-lab-close').addEventListener('click', () => this.closeAnimLab());
panel.querySelectorAll('button[data-theme]').forEach((btn) => {
btn.addEventListener('click', () => {
this.setAnimTheme(btn.dataset.theme);
this._syncAnimLab();
this.demoEntrance('all');
});
});
panel.querySelectorAll('.anim-lab-style').forEach((btn) => {
btn.addEventListener('click', () => {
const { attr, style } = btn.dataset;
if (attr === 'tab') this.setTabAnimStyle(style);
else if (attr === 'win') this.setWinAnimStyle(style);
else if (attr === 'term') this.setTermAnimStyle(style);
else this.setLineAnimStyle(style);
this._syncAnimLab();
this.demoEntrance({ tab: 'tabs', term: 'term' }[attr] || 'all');
});
});
panel.querySelector('input[data-check="termSwitch"]').addEventListener('change', (e) => {
this.setTermAnimOnSwitch(e.target.checked);
});
panel.querySelectorAll('input[data-range]').forEach((input) => {
input.addEventListener('input', () => {
if (input.dataset.range === 'stagger') this.setTabAnimStagger(Number(input.value));
else this.setAnimSpeed(Number(input.value));
this._syncAnimLab();
});
input.addEventListener('change', () => this.demoEntrance('all'));
});
panel.querySelectorAll('button[data-demo]').forEach((btn) => {
btn.addEventListener('click', () => this.demoEntrance(btn.dataset.demo));
});
this._syncAnimLab();
},
closeAnimLab() {
document.getElementById('animLab')?.remove();
this._clearEntranceDemo();
},
_syncAnimLab() {
const panel = document.getElementById('animLab');
if (!panel) return;
const current = {
tab: this._tabAnimStyle,
win: this._winAnimStyle,
line: this._lineAnimStyle,
term: this._termAnimStyle,
};
panel.querySelectorAll('.anim-lab-style').forEach((btn) => {
btn.classList.toggle('selected', current[btn.dataset.attr] === btn.dataset.style);
});
panel.querySelectorAll('button[data-theme]').forEach((btn) => {
const t = ANIM_THEMES.find((x) => x.key === btn.dataset.theme);
btn.classList.toggle('selected', !!t && ['tab', 'win', 'line', 'term'].every((k) => t[k] === current[k]));
});
const check = panel.querySelector('input[data-check="termSwitch"]');
if (check) check.checked = !!this._termAnimOnSwitch;
panel.querySelector('input[data-range="stagger"]').value = String(this._tabAnimStagger);
panel.querySelector('output[data-out="stagger"]').textContent = `${this._tabAnimStagger}ms`;
panel.querySelector('input[data-range="speed"]').value = String(this._animSpeed);
panel.querySelector('output[data-out="speed"]').textContent = `${this._animSpeed.toFixed(1)}x`;
},
_clearEntranceDemo() {
clearTimeout(this._animDemoTimer);
clearTimeout(this._animDemoWindowTimer);
document.querySelectorAll('.session-tab[data-demo]').forEach((el) => el.remove());
document.querySelectorAll('.subagent-window[data-demo]').forEach((el) => el.remove());
document.getElementById('animLabLines')?.remove();
},
/**
* The demo line gets its OWN svg overlay rather than sharing #connectionLines.
* That overlay is rebuilt from real windows via `svg.innerHTML = ''`, so a fake
* path dropped into it is erased the moment anything triggers a redraw -
* including the demo window's own entrance finishing.
*/
_demoLineSvg() {
let svg = document.getElementById('animLabLines');
if (!svg) {
svg = document.createElementNS('http://www.w3.org/2000/svg', 'svg');
svg.id = 'animLabLines';
svg.setAttribute('class', 'connection-lines-svg');
document.body.appendChild(svg);
}
return svg;
},
/** @param {'tabs'|'term'|'window'|'all'} what */
demoEntrance(what = 'all') {
this._clearEntranceDemo();
// The pane is a real, shared element rather than a throwaway, so replay it
// through the same entry point a real launch uses (bypassing the owed-id
// check, which only exists to keep background sessions from hijacking it).
if (what === 'term' || what === 'all') {
const wasOnSwitch = this._termAnimOnSwitch;
this._termAnimOnSwitch = true;
this.playTerminalEntrance(null);
this._termAnimOnSwitch = wasOnSwitch;
}
if (what === 'term') return;
const tabCount = what === 'window' ? 1 : 4;
const tabs = this._demoTabs(tabCount);
let total = Math.max(this._tabAnimDuration() + tabCount * this._tabAnimStagger, this._termAnimDuration());
if (what !== 'tabs') {
// Give the tab cascade a beat, so the window reads as coming out of it.
const lead = what === 'all' ? Math.min(320, this._tabAnimStagger * 2) : 0;
this._animDemoWindowTimer = setTimeout(() => this._demoWindow(tabs[0]), lead);
total = Math.max(total, lead + this._winAnimDuration() + this._lineAnimDuration() + BEAM_HOLD_MS);
}
this._animDemoTimer = setTimeout(() => this._clearEntranceDemo(), total + 1800);
},
/** Append `count` throwaway tabs and run the tab entrance on them. */
_demoTabs(count) {
const container = this.$('sessionTabs');
if (!container) return [];
const made = [];
const base = this.sessions?.size || 0;
for (let i = 0; i < count; i++) {
const tab = document.createElement('div');
tab.className = 'session-tab';
tab.dataset.demo = '1';
tab.dataset.color = ['green', 'blue', 'purple', 'orange', 'pink', 'yellow', 'red'][i % 7];
tab.innerHTML = `
<span class="tab-number">${base + i + 1}</span>
<span class="tab-status idle" aria-hidden="true"></span>
<span class="tab-info"><span class="tab-name-row">
<span class="tab-name">w${base + i + 1}-demo</span>
</span></span>`;
container.appendChild(tab);
made.push(tab);
if (this._tabAnimStyle !== 'off') {
tab.style.setProperty('--tab-enter-delay', `${i * this._tabAnimStagger}ms`);
void tab.offsetWidth; // force layout so the class add starts a fresh run
tab.classList.add('tab-enter');
}
}
return made;
},
/** Append a throwaway agent window under `originTab`, plus its connection line. */
_demoWindow(originTab) {
const tabRect = originTab?.getBoundingClientRect();
const win = document.createElement('div');
win.className = 'subagent-window';
win.dataset.demo = '1';
win.style.width = '360px';
win.style.height = '220px';
win.style.left = `${Math.max(24, (tabRect?.left ?? 120) - 40)}px`;
win.style.top = `${(tabRect?.bottom ?? 60) + 160}px`;
win.style.zIndex = '1001';
win.innerHTML = `
<div class="subagent-window-header">
<div class="subagent-window-title"><span class="icon">&#129302;</span><span class="id">demo-agent</span>
<span class="status running">running</span></div>
</div>
<div class="subagent-window-body"><div class="subagent-empty">Preview window</div></div>`;
document.body.appendChild(win);
this.applyWindowEntrance(win);
this._demoLine(tabRect, win);
},
/** Draw a fake tab→window line into the lab's own overlay and animate it. */
_demoLine(tabRect, win) {
if (!tabRect) return;
const svg = this._demoLineSvg();
const winRect = win.getBoundingClientRect();
const x1 = tabRect.left + tabRect.width / 2;
const y1 = tabRect.bottom;
const x2 = winRect.left + winRect.width / 2;
const y2 = winRect.top;
const midY = (y1 + y2) / 2;
const path = document.createElementNS('http://www.w3.org/2000/svg', 'path');
path.setAttribute('d', `M ${x1} ${y1} C ${x1} ${midY}, ${x2} ${midY}, ${x2} ${y2}`);
path.setAttribute('class', 'connection-line');
path.dataset.demo = '1';
svg.appendChild(path);
if ((this._lineAnimStyle || LINE_ANIM_DEFAULT) === 'off') return;
const len = Math.max(1, Math.round(path.getTotalLength()));
path.style.setProperty('--line-len', `${len}px`);
path.style.setProperty('--line-enter-delay', '0ms');
path.classList.add('line-enter');
if (this._lineAnimStyle === 'packet') {
const packet = path.cloneNode(false);
packet.setAttribute('class', 'connection-line-packet');
packet.dataset.demo = '1';
packet.style.setProperty('--line-len', `${len}px`);
packet.style.setProperty('--line-enter-delay', '0ms');
svg.appendChild(packet);
}
},
});
+26
View File
@@ -103,6 +103,7 @@
'Run Claude Code': '运行 Claude Code',
'Run OpenCode': '运行 OpenCode',
'Run Gemini': '运行 Gemini',
'Run Antigravity': '运行 Antigravity',
'Run Shell': '运行 Shell',
'Select AI backend': '选择 AI 后端',
'Create New Case': '新建案例',
@@ -226,6 +227,7 @@
'Redraw Terminal Button': '重绘终端按钮',
'Tab Bar': '标签栏',
'Tall Tabs (Name + Folder)': '双行标签(名称 + 文件夹)',
'Pop-out Button on Tabs': '标签页弹出窗口按钮',
Panels: '面板',
Monitor: '监视器',
'Project Insights': '项目洞察',
@@ -349,6 +351,26 @@
'Show Shortcuts': '显示快捷键',
'Full shortcut reference': '完整快捷键参考',
// Mobile overview (phone home screen)
'Needs you': '需要你',
'Current sessions': '当前会话',
'Past sessions': '历史会话',
'Show all past sessions': '显示全部历史会话',
'Show fewer': '收起',
'Choose what to run': '选择运行方式',
'Web / URL': '网页 / 链接',
'Add URL…': '添加链接…',
'Nothing running. Hit Run to start something.': '当前没有运行中的会话。点击“运行”开始。',
'No past conversations yet': '尚无历史对话',
'Loading…': '加载中…',
// Status pills are deliberately NOT listed: they are single generic words
// ("idle", "done", "error") that also appear as state strings elsewhere, so
// they carry data-i18n-skip in the DOM instead of a translation entry here.
'Overview Home Screen': '概览主页',
'On phones, the C logo opens a session overview (needs you / spaces / idle) instead of the welcome screen':
'在手机上,点击 C 图标打开会话概览(需要你 / 空间 / 空闲),而不是欢迎页',
Phone: '手机',
// Session/case dialogs
'Session Options': '会话选项',
'Session Name': '会话名称',
@@ -430,6 +452,7 @@
'Respawn Blocked': '重生已阻止',
'Task Complete': '任务完成',
'Copied to clipboard': '已复制到剪贴板',
'Failed to copy': '复制失败',
'Checking…': '正在检查…',
'Starting…': '正在启动…',
'Starting update…': '正在开始更新…',
@@ -579,6 +602,9 @@
'Select a run to view its agents': '选择一次运行以查看其智能体',
'Source type filter': '来源类型筛选',
'Copy content': '复制内容',
'Edit file': '编辑文件',
'Unsaved changes': '未保存的更改',
Saved: '已保存',
'Export as JSON': '导出为 JSON',
'Export as Markdown': '导出为 Markdown',
'Mark all read': '全部标为已读',
+71 -7
View File
@@ -311,19 +311,23 @@
<h1 class="welcome-title">Codeman</h1>
<p class="welcome-desc">Manage AI Coding tools in persistent tmux sessions.</p>
<div class="welcome-actions">
<button class="welcome-btn welcome-btn-claude" onclick="app.setRunMode('claude'); app.runClaude()">
<button class="welcome-btn welcome-btn-claude" id="welcomeClaudeBtn" style="display: none;" onclick="app.setRunMode('claude'); app.runClaude()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Claude Code
</button>
<button class="welcome-btn welcome-btn-tunnel" id="welcomeTunnelBtn" onclick="app.toggleTunnelFromWelcome()">
<button class="welcome-btn welcome-btn-tunnel" id="welcomeTunnelBtn" style="display: none;" onclick="app.toggleTunnelFromWelcome()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><path d="M12 2L2 7l10 5 10-5-10-5z"/><path d="M2 17l10 5 10-5"/><path d="M2 12l10 5 10-5"/></svg>
Cloudflare Tunnel
</button>
<button class="welcome-btn welcome-btn-opencode" onclick="app.setRunMode('opencode'); app.runOpenCode()">
<button class="welcome-btn welcome-btn-opencode" id="welcomeOpencodeBtn" style="display: none;" onclick="app.setRunMode('opencode'); app.runOpenCode()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run OpenCode
</button>
<button class="welcome-btn welcome-btn-gemini" onclick="app.setRunMode('gemini'); app.runGemini()">
<button class="welcome-btn welcome-btn-antigravity" id="welcomeAntigravityBtn" style="display: none;" onclick="app.setRunMode('antigravity'); app.runAntigravity()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Antigravity
</button>
<button class="welcome-btn welcome-btn-gemini" id="welcomeGeminiBtn" style="display: none;" onclick="app.setRunMode('gemini'); app.runGemini()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Gemini
</button>
@@ -379,6 +383,12 @@
<button class="welcome-ralph-link" onclick="app.showRalphWizard()">Start Ralph Loop &rarr;</button>
</div>
</div>
<!-- Mobile Overview (phone home screen, shown in place of the welcome
overlay when the "C" logo is tapped). Ships `hidden`: only
mobile-overview.js removes it, and only when the phone gate passes, so
desktop (which never loads mobile.css) can never render it. -->
<div class="mobile-overview" id="mobileOverview" hidden></div>
</main>
<!-- Project Insights Panel (shows file-viewing Bash commands) -->
@@ -416,11 +426,18 @@
<div class="file-preview-header">
<span class="file-preview-title" id="filePreviewTitle">file.ts</span>
<div class="file-preview-actions">
<button class="btn-icon-sm file-preview-edit-btn" id="filePreviewEditBtn" onclick="app.enterFilePreviewEdit()" title="Edit file" aria-label="Edit file" hidden><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M17 3a2.85 2.83 0 1 1 4 4L7.5 20.5 2 22l1.5-5.5z"/></svg></button>
<button class="btn-icon-sm" onclick="app.copyFilePreviewContent()" title="Copy content">&#x2398;</button>
<button class="btn-icon-sm" onclick="app.closeFilePreview()" title="Close">&times;</button>
</div>
</div>
<div class="file-preview-body" id="filePreviewBody"></div>
<div class="file-preview-editbar" id="filePreviewEditBar" hidden>
<span class="file-preview-dirty" id="filePreviewDirty" hidden>Unsaved changes</span>
<span class="file-preview-editbar-spacer"></span>
<button class="file-preview-editbar-btn" onclick="app.cancelFilePreviewEdit()">Cancel</button>
<button class="file-preview-editbar-btn file-preview-editbar-btn--save" id="filePreviewSaveBtn" onclick="app.saveFilePreviewEdit()" disabled>Save</button>
</div>
<div class="file-preview-footer" id="filePreviewFooter"></div>
</div>
</div>
@@ -464,6 +481,9 @@
<button class="run-mode-option" data-mode="gemini" onclick="app.setRunMode('gemini')">
<span class="run-mode-dot gemini"></span>Gemini
</button>
<button class="run-mode-option" data-mode="antigravity" onclick="app.setRunMode('antigravity')">
<span class="run-mode-dot antigravity"></span>Antigravity
</button>
<div class="run-mode-sep"></div>
<button class="run-mode-option" data-mode="shell" onclick="app.setRunMode('shell')">
<span class="run-mode-dot shell"></span>Terminal / Shell
@@ -630,6 +650,8 @@
<section class="shortcut-section">
<h4>Terminal</h4>
<div class="shortcuts-grid">
<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection (interrupts when nothing is selected)</div>
<div><kbd>Ctrl</kbd>+<kbd>Shift</kbd>+<kbd>C</kbd></div><div>Copy Selection</div>
<div><kbd>Ctrl</kbd>+<kbd>L</kbd></div><div>Clear Terminal</div>
<div><kbd>Ctrl</kbd>+<kbd>+</kbd></div><div>Increase Font</div>
<div><kbd>Ctrl</kbd>+<kbd>-</kbd></div><div>Decrease Font</div>
@@ -733,6 +755,7 @@
<option value="opencode">OpenCode</option>
<option value="codex">Codex</option>
<option value="gemini">Gemini</option>
<option value="antigravity">Antigravity</option>
</select>
</div>
<div class="form-row"><label>Working Directory</label><input type="text" id="schWorkingDir" placeholder="/absolute/path"></div>
@@ -1276,6 +1299,20 @@
</optgroup>
</select>
</div>
<div class="settings-item settings-item-multiline" title="How new session tabs, the terminal pane, agent windows and their connection lines appear. This device only. For per-surface control and a live preview, add ?animlab=1 to the URL.">
<div class="settings-item-text">
<span class="settings-item-label">Entrance Animations</span>
<span class="settings-item-desc">How new tabs, panes and agent windows arrive</span>
</div>
<select id="appSettingsEntranceAnim" class="form-select settings-inline-select">
<option value="legacy">Off (default)</option>
<option value="terminal">Terminal (CRT)</option>
<option value="beamdown">Beam down</option>
<option value="quiet">Quiet</option>
<option value="playful">Playful</option>
<option value="custom">Custom (set in the lab)</option>
</select>
</div>
<div class="settings-item" id="appSettingsWebglRendererItem" title="Use the GPU-accelerated WebGL terminal renderer (desktop only). Turn off to force the DOM renderer if you hit GPU glitches. Codeman also auto-falls-back to the DOM renderer after repeated GPU stalls.">
<span class="settings-item-label">WebGL Renderer</span>
<label class="switch switch-sm">
@@ -1285,7 +1322,7 @@
</div>
<!-- Input Section -->
<div class="settings-section-header">Input</div>
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Turn on if scrolling back through history doesn't work (e.g. macOS trackpad in Claude sessions). Shift+wheel always reaches local scrollback regardless.">
<div class="settings-item settings-item-multiline" title="Scroll the terminal's own local scrollback with a plain mouse wheel / two-finger swipe, instead of forwarding the wheel to the CLI's transcript. Leave OFF for Claude/Codex sessions: those CLIs redraw the screen in place and keep no local scrollback, so the wheel would have almost nothing to scroll. In Claude sessions Codeman then falls back to paging the CLI's own transcript; Codex sessions have no such fallback, so the wheel goes dead. Shift+wheel always reaches local scrollback regardless.">
<div class="settings-item-text">
<span class="settings-item-label">Wheel Scrolls Local History</span>
<span class="settings-item-desc">Plain wheel/trackpad pages the terminal scrollback</span>
@@ -1424,6 +1461,16 @@
</label>
</div>
<!-- Phone Section (only meaningful under 430px, hidden elsewhere by settings-ui.js) -->
<div class="settings-section-header" id="appSettingsPhoneSection">Phone</div>
<div class="settings-item" id="appSettingsMobileOverviewItem" title="On phones, the C logo opens a session overview (needs you / spaces / idle) instead of the welcome screen">
<span class="settings-item-label">Overview Home Screen</span>
<label class="switch switch-sm">
<input type="checkbox" id="appSettingsMobileOverview" checked>
<span class="slider"></span>
</label>
</div>
<!-- Tab Bar Section -->
<div class="settings-section-header">Tab Bar</div>
<div class="settings-item" title="Show folder path below tab name and allow tab bar to wrap into multiple rows">
@@ -1433,6 +1480,13 @@
<span class="slider"></span>
</label>
</div>
<div class="settings-item" title="Show the open-in-new-window (pop-out) button when hovering a session tab">
<span class="settings-item-label">Pop-out Button on Tabs</span>
<label class="switch switch-sm">
<input type="checkbox" id="appSettingsShowTabDetachButton">
<span class="slider"></span>
</label>
</div>
<!-- Panels Section -->
<div class="settings-section-header">Panels</div>
@@ -1654,6 +1708,14 @@
</label>
<span class="form-hint">Start new Codex sessions with --dangerously-bypass-approvals-and-sandbox</span>
</div>
<div class="form-row form-row-switch">
<label>Animated Status Effects</label>
<label class="switch">
<input type="checkbox" id="appSettingsCodexAnimations">
<span class="slider"></span>
</label>
<span class="form-hint">Decorative Codex TUI motion for new local sessions. Leave off to reduce remote and mobile redraws.</span>
</div>
</div>
<!-- Models Tab -->
<div class="modal-tab-content hidden" id="settings-models">
@@ -1848,7 +1910,7 @@
<input type="checkbox" id="eventIdleAudio">
<div class="event-label">Response complete</div>
<input type="checkbox" id="eventStopEnabled" checked>
<input type="checkbox" id="eventStopEnabled">
<input type="checkbox" id="eventStopBrowser">
<input type="checkbox" id="eventStopPush">
<input type="checkbox" id="eventStopAudio">
@@ -2141,7 +2203,7 @@
<div class="form-row">
<label>Image</label>
<input type="text" id="dockerImage" placeholder="codeman/agent:base" autocomplete="off" autocapitalize="off" spellcheck="false">
<span class="form-hint">Build it once with <code>node scripts/build-agent-image.mjs</code>. Contains node + claude/codex/gemini + tmux.</span>
<span class="form-hint">Build it once with <code>node scripts/build-agent-image.mjs</code>. Contains node + claude/codex/gemini/opencode/agy + tmux.</span>
</div>
<div class="form-row">
<label>Network</label>
@@ -2626,6 +2688,8 @@
<script defer src="admin-ui.js"></script>
<script defer src="session-ui.js"></script>
<script defer src="webview-tabs.js"></script>
<script defer src="mobile-overview.js"></script>
<script defer src="entrance-animations.js"></script>
<script defer src="ralph-wizard.js"></script>
<script defer src="api-client.js"></script>
<script defer src="subagent-windows.js"></script>
+62 -37
View File
@@ -209,9 +209,13 @@ const MobileDetection = {
* Also handles terminal scrolling and toolbar repositioning via visualViewport API.
*/
const KeyboardHandler = {
VIEWPORT_SETTLE_MS: 80,
lastViewportHeight: 0,
keyboardVisible: false,
initialViewportHeight: 0,
_viewportSettleTimer: null,
_settleScrollToBottom: false,
_settlePending: false,
/** Initialize keyboard handling */
init() {
@@ -276,6 +280,12 @@ const KeyboardHandler = {
window.removeEventListener('scroll', this._windowScrollHandler);
this._windowScrollHandler = null;
}
if (this._viewportSettleTimer) {
clearTimeout(this._viewportSettleTimer);
this._viewportSettleTimer = null;
}
this._settleScrollToBottom = false;
this._settlePending = false;
},
/** Handle viewport resize (keyboard show/hide) */
@@ -313,6 +323,7 @@ const KeyboardHandler = {
}
this.updateLayoutForKeyboard();
this._deferViewportSettle();
this.lastViewportHeight = currentHeight;
},
@@ -414,32 +425,9 @@ const KeyboardHandler = {
// iOS Safari may scroll the document to reveal xterm's hidden textarea.
window.scrollTo(0, 0);
// Refit terminal locally AND send resize to server so Claude Code (Ink)
// knows the actual terminal dimensions. Without this, Ink redraws at the
// old (larger) row count when the user types, causing content to scroll
// off the visible area with each keystroke.
// Note: the throttledResize handler still suppresses ongoing resize events
// while keyboard is up — this one-shot resize on open/close is sufficient.
setTimeout(() => {
if (typeof app !== 'undefined' && app.terminal) {
if (app.fitAddon)
try {
app.fitAddon.fit();
} catch {}
// Eliminate terminal row quantization gap: xterm can only show whole
// rows, so leftover pixels create dead space below the last row.
// Shrink .main's paddingBottom by the gap so the terminal fills flush
// to the accessory bar.
this._shrinkPaddingToFit();
app.terminal.scrollToBottom();
app._syncMobileHelperTextareaToCursor?.();
app._localEchoOverlay?.rerender?.();
// Send resize to server so PTY dimensions match xterm
this._sendTerminalResize();
}
// Reset again after fit/resize in case layout changes triggered scroll
window.scrollTo(0, 0);
}, 150);
// visualViewport emits multiple heights throughout the OS animation.
// Re-schedule on every event and fit only after the final height settles.
this._scheduleViewportSettle({ scrollToBottom: true });
// Reposition subagent windows to stack from bottom (above keyboard)
if (typeof app !== 'undefined') app.relayoutMobileSubagentWindows();
@@ -454,22 +442,59 @@ const KeyboardHandler = {
this.resetLayout();
// Refit terminal, scroll to bottom, and send resize to restore original dimensions
setTimeout(() => {
if (typeof app !== 'undefined' && app.fitAddon) {
try {
app.fitAddon.fit();
} catch {}
if (app.terminal) app.terminal.scrollToBottom();
// Send resize to server to restore full terminal size
this._sendTerminalResize();
}
}, 100);
this._scheduleViewportSettle({ scrollToBottom: true });
// Reposition subagent windows to stack from top (below header)
if (typeof app !== 'undefined') app.relayoutMobileSubagentWindows();
},
/**
* Coalesce the keyboard animation into one final xterm reflow and PTY resize.
* Only a real show/hide transition arms the settle work; ongoing viewport
* resize events merely push a pending settle back (_deferViewportSettle).
* A viewport change that never crosses the show/hide thresholds must not
* refit: keyboard detection can miss a fine-grained OS animation entirely
* (each step under 150px, with the baseline chasing the animation), and the
* container is then mid-animation with no keyboard CSS compensation, so a
* fit against it resizes the PTY to transient dims and the SIGWINCH thrash
* garbles the transcript.
*/
_scheduleViewportSettle({ scrollToBottom = false } = {}) {
this._settleScrollToBottom = this._settleScrollToBottom || scrollToBottom;
this._settlePending = true;
this._armViewportSettleTimer();
},
/** Push a pending settle back while the viewport is still animating; no-op otherwise. */
_deferViewportSettle() {
if (!this._settlePending) return;
this._armViewportSettleTimer();
},
_armViewportSettleTimer() {
if (this._viewportSettleTimer) clearTimeout(this._viewportSettleTimer);
this._viewportSettleTimer = setTimeout(() => {
this._viewportSettleTimer = null;
this._settlePending = false;
const shouldScrollToBottom = this._settleScrollToBottom;
this._settleScrollToBottom = false;
if (typeof app !== 'undefined' && app.terminal) {
if (app.fitAddon) {
try {
app.fitAddon.fit();
} catch {}
}
if (this.keyboardVisible) this._shrinkPaddingToFit();
if (shouldScrollToBottom) app.terminal.scrollToBottom();
app._syncMobileHelperTextareaToCursor?.();
app._localEchoOverlay?.rerender?.();
this._sendTerminalResize();
}
window.scrollTo(0, 0);
}, this.VIEWPORT_SETTLE_MS);
},
/** Send current terminal dimensions to the server (one-shot, for keyboard open/close) */
_sendTerminalResize() {
if (typeof app === 'undefined' || !app.activeSessionId || !app.fitAddon) return;
+690
View File
@@ -0,0 +1,690 @@
/**
* @fileoverview Phone home screen: a scrolling overview of what every session is
* doing, shown instead of the welcome overlay when the "C" logo is tapped.
*
* The welcome screen answers "how do I start something"; on a phone the more
* urgent question is "which of my sessions is blocked on me". This surface
* answers that first: NEEDS YOU (pending permission/question/idle hooks and
* errored sessions), then SPACES (cases, expandable to their sessions), then
* WORKING and IDLE / DONE.
*
* PHONE ONLY. The gate is `shouldUseMobileOverview()` (viewport < 430px, not a
* popped-out solo window, per-device setting on). Tablet and desktop keep the
* welcome overlay untouched. The container ships with the `hidden` attribute and
* only this module removes it, so desktop (which never loads mobile.css) cannot
* render an unstyled overview even if a class rule leaked.
*
* Everything renders from state the page already holds (`this.sessions`,
* `this.cases`, `this.pendingHooks`) — no endpoint, no SSE event, no schema.
* `buildMobileOverviewModel()` is pure and unit-tested (test/mobile-overview.test.ts).
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js (this.sessions, this.cases, this.pendingHooks, selectSession, run)
* @dependency mobile-handlers.js (MobileDetection)
* @dependency session-ui.js (selectQuickStartCase for "New session here")
* @loadorder 12.55 of 16, after webview-tabs.js, before entrance-animations.js
*/
/** Viewport width that counts as a phone. Matches the mobile.css phone block. */
const MOBILE_OVERVIEW_PHONE_QUERY = '(max-width: 430px)';
/** Sort rank per state: the most demanding thing sorts first inside a section. */
const MOBILE_OVERVIEW_STATE_RANK = {
needs: 0,
error: 1,
waiting: 2,
working: 3,
idle: 4,
done: 5,
};
/** How many past conversations show before the "Show all" toggle. */
const MOBILE_OVERVIEW_PAST_LIMIT = 8;
/**
* Backends offered by the Run picker, mirroring the toolbar's run-mode menu
* (`#runModeMenu` in index.html). `short` is the badge on the Run button itself.
*/
const MOBILE_OVERVIEW_RUN_MODES = [
{ mode: 'claude', label: 'Claude Code', short: 'Claude' },
{ mode: 'opencode', label: 'OpenCode', short: 'OpenCode' },
{ mode: 'codex', label: 'Codex', short: 'Codex' },
{ mode: 'gemini', label: 'Gemini', short: 'Gemini' },
{ mode: 'antigravity', label: 'Antigravity', short: 'Antigravity' },
{ mode: 'shell', label: 'Terminal / Shell', short: 'Shell' },
];
/** Pill copy per state. Kept short: a phone row has ~90px for it. */
const MOBILE_OVERVIEW_PILL_LABEL = {
needs: 'needs you',
error: 'error',
waiting: 'waiting',
working: 'working',
idle: 'idle',
done: 'done',
};
Object.assign(CodemanApp.prototype, {
// ═══════════════════════════════════════════════════════════════
// Model (pure)
// ═══════════════════════════════════════════════════════════════
/**
* Classify one session.
* Order matters: an action hook outranks everything (it is literally blocking
* the agent), and a pending idle_prompt outranks a stale 'busy' status because
* the hook is the newer signal.
* @param {object} session session state from this.sessions
* @param {Set<string>|undefined} hooks pending hook types for that session
* @returns {'needs'|'error'|'waiting'|'working'|'idle'|'done'}
*/
_mobileOverviewState(session, hooks) {
if (hooks && (hooks.has('permission_prompt') || hooks.has('elicitation_dialog'))) return 'needs';
if (session.status === 'error') return 'error';
if (hooks && hooks.has('idle_prompt')) return 'waiting';
if (session.status === 'busy') return 'working';
if (session.status === 'stopped') return 'done';
return 'idle';
},
/**
* Longest-prefix match of a workingDir against the case list, so a session
* started in a subdirectory still belongs to its case. Mirrors the matching in
* `_resolveCaseLabel()` (terminal-ui.js) but returns the case itself.
* @returns {object|null} the matching case, or null when the dir is outside every case
*/
_mobileOverviewCaseFor(workingDir, cases) {
if (!workingDir) return null;
let best = null;
for (const c of cases || []) {
if (!c || !c.path) continue;
if (workingDir === c.path) return c;
if (workingDir.startsWith(c.path + '/') && (!best || c.path.length > best.path.length)) {
best = c;
}
}
return best;
},
/**
* Build the whole overview model. PURE: reads only its argument, touches no DOM
* and no `this` state, so it can be unit-tested against plain objects.
*
* @param {object} input
* @param {Map<string, object>|Array} input.sessions live sessions (this.sessions)
* @param {Array} input.cases case list (this.cases)
* @param {Array<string>} [input.sessionOrder] the user's tab order, used as the tiebreak
* @param {Map<string, Set<string>>} [input.pendingHooks] this.pendingHooks
* @param {Array} [input.history] unified session items (GET /api/sessions/unified)
* @returns {{needsYou: Array, current: Array, past: Array, sessionCount: number}}
*/
buildMobileOverviewModel(input) {
const cases = Array.isArray(input && input.cases) ? input.cases : [];
const order = Array.isArray(input && input.sessionOrder) ? input.sessionOrder : [];
const pendingHooks = (input && input.pendingHooks) || new Map();
const raw = (input && input.sessions) || [];
const sessions = typeof raw.values === 'function' ? Array.from(raw.values()) : Array.from(raw);
const rows = sessions.map((session) => {
const matched = this._mobileOverviewCaseFor(session.workingDir, cases);
const state = this._mobileOverviewState(session, pendingHooks.get && pendingHooks.get(session.id));
const orderIndex = order.indexOf(session.id);
return {
id: session.id,
name: this.getSessionName ? this.getSessionName(session) : session.name || session.id.slice(0, 8),
mode: session.mode || 'claude',
caseName: matched ? matched.name : '',
dir: this._shortenHomePath ? this._shortenHomePath(session.workingDir) : session.workingDir || '',
state,
pill: MOBILE_OVERVIEW_PILL_LABEL[state] || state,
orderIndex: orderIndex === -1 ? Number.MAX_SAFE_INTEGER : orderIndex,
};
});
const bySeverityThenOrder = (a, b) => {
const rank = MOBILE_OVERVIEW_STATE_RANK[a.state] - MOBILE_OVERVIEW_STATE_RANK[b.state];
return rank !== 0 ? rank : a.orderIndex - b.orderIndex;
};
const inSection = (states) => rows.filter((r) => states.includes(r.state)).sort(bySeverityThenOrder);
// Past = conversations from the unified list that are not currently live.
// The endpoint already folds a transcript into its owning session (via the
// claudeSessionId alias map), so a plain id check is enough to avoid listing
// a running session twice.
const liveIds = new Set(rows.map((r) => r.id));
const past = (Array.isArray(input && input.history) ? input.history : [])
.filter((item) => item && item.sessionId && !liveIds.has(item.sessionId))
.map((item) => {
const matched = this._mobileOverviewCaseFor(item.workingDir, cases);
const dir = item.workingDir || '';
// The transcript reader emits the literal "(no content)" for a
// conversation it could not pull a prompt from; that is not a title.
const prompt = (item.firstPrompt || '').trim();
const title = prompt && prompt !== '(no content)' ? prompt : '';
return {
id: item.sessionId,
claudeSessionId: item.claudeSessionId || '',
workingDir: dir,
name: item.name || '',
title: title || item.name || dir.split('/').pop() || item.sessionId.slice(0, 8),
mode: item.mode || 'claude',
caseName: matched ? matched.name : '',
dir: this._shortenHomePath ? this._shortenHomePath(dir) : dir,
at: item.lastActivityAt || item.createdAt || 0,
};
})
.sort((a, b) => b.at - a.at);
return {
needsYou: inSection(['needs', 'error', 'waiting']),
current: inSection(['working', 'idle', 'done']),
past,
sessionCount: rows.length,
};
},
// ═══════════════════════════════════════════════════════════════
// Gate + visibility
// ═══════════════════════════════════════════════════════════════
/**
* Phone-only gate. Width-driven (not `isHandheldDevice()`): this is a LAYOUT
* decision, and an unfolded foldable with a tablet-width viewport should get
* the tablet welcome screen. Per-device settings identity is a separate
* question and deliberately stays handheld-based.
*/
shouldUseMobileOverview() {
if (this.isSoloWindow) return false;
const settings = this.loadAppSettingsFromStorage ? this.loadAppSettingsFromStorage() : {};
if (settings.mobileOverviewEnabled === false) return false;
if (typeof MobileDetection !== 'undefined' && MobileDetection.getDeviceType) {
return MobileDetection.getDeviceType() === 'mobile';
}
return !!(window.matchMedia && window.matchMedia(MOBILE_OVERVIEW_PHONE_QUERY).matches);
},
/** True while the overview is the visible home surface. */
isMobileOverviewVisible() {
const el = document.getElementById('mobileOverview');
return !!el && el.classList.contains('visible');
},
showMobileOverview() {
const el = document.getElementById('mobileOverview');
if (!el) return;
el.hidden = false;
el.classList.add('visible');
this._wireMobileOverview(el);
this.renderMobileOverview();
void this.loadMobileOverviewHistory();
},
hideMobileOverview() {
const el = document.getElementById('mobileOverview');
if (!el) return;
this._closeMobileOverviewRunMenu();
el.classList.remove('visible');
el.hidden = true;
},
/**
* Past conversations, fetched once per home-screen visit. The unified list is
* the same source the welcome screen resumes from, so a row resumed here and a
* row resumed there behave identically. Failures leave the section out rather
* than showing an error: the live sessions above it are the important part.
*/
async loadMobileOverviewHistory() {
if (this._mobileOverviewHistoryLoading) return;
this._mobileOverviewHistoryLoading = true;
try {
this._mobileOverviewHistory = await this._fetchUnifiedSessions(60);
} catch (err) {
console.warn('[mobile-overview] history load failed:', err);
this._mobileOverviewHistory = this._mobileOverviewHistory || [];
} finally {
this._mobileOverviewHistoryLoading = false;
if (this.isMobileOverviewVisible()) this.renderMobileOverview();
}
},
/** Re-render only when the surface is actually showing (called from the tab renderer). */
_refreshMobileOverviewIfVisible() {
if (!this.isMobileOverviewVisible()) return;
this._debouncedCall('mobileOverview', () => this.renderMobileOverview(), 150);
},
/**
* One delegated click listener for every row, plus a breakpoint listener so
* rotating or unfolding while on the home screen swaps to the right surface
* instead of stranding a phone layout on a tablet-width viewport.
*/
_wireMobileOverview(el) {
if (this._mobileOverviewWired) return;
this._mobileOverviewWired = true;
el.addEventListener('click', (event) => {
const target = event.target && event.target.closest && event.target.closest('[data-mo-action]');
if (!target) return;
const action = target.dataset.moAction;
if (action === 'session') {
this._closeMobileOverviewRunMenu();
void this.selectSession(target.dataset.moSession);
} else if (action === 'resume') {
this._closeMobileOverviewRunMenu();
void this.resumeMobileOverviewSession(target.dataset.moSession);
} else if (action === 'more-past') {
this._mobileOverviewShowAllPast = !this._mobileOverviewShowAllPast;
this.renderMobileOverview();
} else if (action === 'run') {
this._closeMobileOverviewRunMenu();
void this.run();
} else if (action === 'run-menu') {
this._toggleMobileOverviewRunMenu();
} else if (action === 'run-mode') {
// Picking a backend both selects it (so the Run button keeps meaning what
// you last chose, exactly like the toolbar) and launches it: on a phone
// the pick IS the intent to start.
this._closeMobileOverviewRunMenu();
this.setRunMode(target.dataset.moMode);
void this.run();
} else if (action === 'run-webview') {
this._closeMobileOverviewRunMenu();
void this.openWebviewFromMenu(target.dataset.moWebview);
} else if (action === 'run-add-url') {
this._closeMobileOverviewRunMenu();
this.showWebviewModal();
}
});
if (window.matchMedia) {
const mq = window.matchMedia(MOBILE_OVERVIEW_PHONE_QUERY);
const onChange = () => {
// Only relevant while a home surface is up; entering a session re-decides
// through hideWelcome()/showWelcome() anyway.
if (this.activeSessionId) return;
if (typeof this.showWelcome === 'function') this.showWelcome();
};
if (mq.addEventListener) mq.addEventListener('change', onChange);
else if (mq.addListener) mq.addListener(onChange);
}
},
_toggleMobileOverviewRunMenu() {
this._mobileOverviewRunMenuOpen = !this._mobileOverviewRunMenuOpen;
this.renderMobileOverview();
},
_closeMobileOverviewRunMenu() {
if (!this._mobileOverviewRunMenuOpen) return;
this._mobileOverviewRunMenuOpen = false;
if (this.isMobileOverviewVisible()) this.renderMobileOverview();
},
/**
* Resume a past conversation. Delegates to the same resumeHistorySession() the
* welcome screen's Resume list uses, so name synthesis, envOverrides and the
* resumeSessionId wiring stay in one place.
*/
async resumeMobileOverviewSession(sessionId) {
const row = (this._mobileOverviewPastRows || []).find((r) => r.id === sessionId);
if (!row || !row.workingDir) return;
await this.resumeHistorySession(row.claudeSessionId || row.id, row.workingDir, row.name || undefined);
},
// ═══════════════════════════════════════════════════════════════
// Render
// ═══════════════════════════════════════════════════════════════
renderMobileOverview() {
const el = document.getElementById('mobileOverview');
if (!el) return;
const model = this.buildMobileOverviewModel({
sessions: this.sessions,
cases: this.cases,
sessionOrder: this.sessionOrder,
pendingHooks: this.pendingHooks,
history: this._mobileOverviewHistory,
});
// Resume needs the workingDir/claudeSessionId off the row the user tapped.
this._mobileOverviewPastRows = model.past;
el.replaceChildren();
el.appendChild(this._buildMobileOverviewTop());
if (model.needsYou.length) {
el.appendChild(
this._buildMobileOverviewSection(
'Needs you',
model.needsYou.length,
model.needsYou.map((r) => this._buildMobileOverviewRow(r))
)
);
}
el.appendChild(
this._buildMobileOverviewSection(
'Current sessions',
model.current.length,
model.current.map((r) => this._buildMobileOverviewRow(r)),
'Nothing running. Hit Run to start something.'
)
);
el.appendChild(
this._buildMobileOverviewSection(
'Past sessions',
model.past.length,
this._buildMobileOverviewPast(model),
this._mobileOverviewHistory ? 'No past conversations yet' : 'Loading…'
)
);
},
/**
* Past conversations, newest first and capped: the unified list can run to
* dozens, and this section sits below the live ones on purpose.
*/
_buildMobileOverviewPast(model) {
const showAll = !!this._mobileOverviewShowAllPast;
const visible = showAll ? model.past : model.past.slice(0, MOBILE_OVERVIEW_PAST_LIMIT);
const children = visible.map((r) => this._buildMobileOverviewPastRow(r));
const hiddenCount = model.past.length - visible.length;
if (hiddenCount > 0 || showAll) {
const toggle = document.createElement('button');
toggle.type = 'button';
toggle.className = 'mobile-overview-more';
toggle.dataset.moAction = 'more-past';
const label = document.createElement('span');
label.textContent = showAll ? 'Show fewer' : 'Show all past sessions';
toggle.appendChild(label);
if (!showAll) {
const count = document.createElement('span');
count.className = 'mobile-overview-more-count';
count.setAttribute('data-i18n-skip', '');
count.textContent = String(hiddenCount);
toggle.appendChild(count);
}
children.push(toggle);
}
return children;
},
_buildMobileOverviewTop() {
const wrap = document.createElement('div');
wrap.className = 'mobile-overview-header';
const top = document.createElement('div');
top.className = 'mobile-overview-top';
const brand = document.createElement('span');
brand.className = 'mobile-overview-brand';
brand.textContent = (window.CodemanI18n && window.CodemanI18n.displayName) || 'Codeman';
brand.setAttribute('data-i18n-skip', '');
top.appendChild(brand);
// Split button carrying the TOOLBAR's own classes (`btn-toolbar btn-run
// mode-<mode>` / `btn-run-gear`), so the per-backend gradient, border and
// text color come from the same rules as the Run button in the toolbar and
// stay in sync with it for free. mobile.css only sizes it.
const group = document.createElement('div');
group.className = 'mobile-overview-run-group';
const mode = this.runMode || 'claude';
const run = document.createElement('button');
run.className = `btn-toolbar btn-run mode-${mode} mobile-overview-run`;
run.type = 'button';
run.dataset.moAction = 'run';
const runLabel = document.createElement('span');
runLabel.textContent = 'Run';
run.appendChild(runLabel);
const runMode = document.createElement('span');
runMode.className = 'mobile-overview-run-mode';
runMode.setAttribute('data-i18n-skip', '');
runMode.textContent = MOBILE_OVERVIEW_RUN_MODES.find((m) => m.mode === mode)?.short || mode;
run.appendChild(runMode);
group.appendChild(run);
const caret = document.createElement('button');
caret.className = `btn-toolbar btn-run-gear mode-${mode} mobile-overview-run-caret`;
caret.type = 'button';
caret.dataset.moAction = 'run-menu';
caret.setAttribute('aria-label', 'Choose what to run');
caret.setAttribute('aria-expanded', String(!!this._mobileOverviewRunMenuOpen));
// An SVG chevron, not a "⌄" glyph: the character carries its own baseline
// offset, so it sits visibly low in a flex-centered box no matter what the
// line-height says. A path is centered by geometry. Same shape the toolbar's
// run-mode gear uses.
caret.appendChild(this._buildMobileOverviewChevron());
group.appendChild(caret);
top.appendChild(group);
wrap.appendChild(top);
if (this._mobileOverviewRunMenuOpen) wrap.appendChild(this._buildMobileOverviewRunMenu());
return wrap;
},
/** Down chevron as SVG (see the note at its call site). */
_buildMobileOverviewChevron() {
const svg = document.createElementNS('http://www.w3.org/2000/svg', 'svg');
svg.setAttribute('viewBox', '0 0 24 24');
svg.setAttribute('fill', 'none');
svg.setAttribute('stroke', 'currentColor');
svg.setAttribute('stroke-width', '2.5');
svg.setAttribute('stroke-linecap', 'round');
svg.setAttribute('stroke-linejoin', 'round');
svg.setAttribute('aria-hidden', 'true');
const path = document.createElementNS('http://www.w3.org/2000/svg', 'path');
path.setAttribute('d', 'M6 9l6 6 6-6');
svg.appendChild(path);
return svg;
},
/**
* The Run picker: the same backends as the toolbar's run-mode menu, plus saved
* web tabs. Deliberately no "Recent Sessions" block, unlike the toolbar menu:
* past conversations have their own section further down this screen.
*
* Gated the same way as the toolbar's #runModeMenu (isCliAvailable(), shell
* exempt) — this list is a separate, hardcoded duplicate of the toolbar's menu
* rather than a shared render, so it never picked up #201's gating and offered
* every backend regardless of what's actually installed.
*/
_buildMobileOverviewRunMenu() {
const menu = document.createElement('div');
menu.className = 'mobile-overview-run-menu';
const current = this.runMode || 'claude';
for (const entry of MOBILE_OVERVIEW_RUN_MODES) {
if (entry.mode !== 'shell' && !this.isCliAvailable(entry.mode)) continue;
const option = document.createElement('button');
option.type = 'button';
option.className = 'mobile-overview-run-option' + (entry.mode === current ? ' selected' : '');
option.dataset.moAction = 'run-mode';
option.dataset.moMode = entry.mode;
const dot = document.createElement('span');
dot.className = 'run-mode-dot ' + entry.mode;
dot.setAttribute('aria-hidden', 'true');
option.appendChild(dot);
const label = document.createElement('span');
label.textContent = entry.label;
option.appendChild(label);
menu.appendChild(option);
}
const header = document.createElement('div');
header.className = 'mobile-overview-run-header';
header.textContent = 'Web / URL';
menu.appendChild(header);
for (const webview of this.webviews ? this.webviews.values() : []) {
const option = document.createElement('button');
option.type = 'button';
option.className = 'mobile-overview-run-option';
option.dataset.moAction = 'run-webview';
option.dataset.moWebview = webview.id;
const dot = document.createElement('span');
dot.className = 'run-mode-dot web';
dot.setAttribute('aria-hidden', 'true');
option.appendChild(dot);
const label = document.createElement('span');
// A dashboard name is user content.
label.className = 'case-name';
label.textContent = webview.name;
option.appendChild(label);
menu.appendChild(option);
}
const add = document.createElement('button');
add.type = 'button';
add.className = 'mobile-overview-run-option mobile-overview-run-option--add';
add.dataset.moAction = 'run-add-url';
const addDot = document.createElement('span');
addDot.className = 'run-mode-dot web';
addDot.setAttribute('aria-hidden', 'true');
add.appendChild(addDot);
const addLabel = document.createElement('span');
addLabel.textContent = 'Add URL…';
add.appendChild(addLabel);
menu.appendChild(add);
return menu;
},
_buildMobileOverviewSection(title, count, children, emptyText) {
const section = document.createElement('section');
section.className = 'mobile-overview-section';
const heading = document.createElement('h2');
heading.className = 'mobile-overview-heading';
const label = document.createElement('span');
label.textContent = title;
heading.appendChild(label);
const badge = document.createElement('span');
badge.className = 'mobile-overview-heading-count';
badge.textContent = String(count);
badge.setAttribute('data-i18n-skip', '');
heading.appendChild(badge);
section.appendChild(heading);
if (!children.length && emptyText) {
const empty = document.createElement('p');
empty.className = 'mobile-overview-empty';
empty.textContent = emptyText;
section.appendChild(empty);
return section;
}
for (const child of children) section.appendChild(child);
return section;
},
/**
* A session row. The state class drives the same visual language as the
* session tabs: green dot when it is fine (pulsing while working), a yellow
* blinking row when it wants input, a red blinking row when it asked a
* question. Anything else here would mean two different meanings for the same
* colors on one screen.
*/
_buildMobileOverviewRow(row) {
const item = document.createElement('button');
item.type = 'button';
item.className = 'mobile-overview-row mobile-overview-row--' + row.state;
item.dataset.moAction = 'session';
item.dataset.moSession = row.id;
const dot = document.createElement('span');
dot.className = 'mobile-overview-dot mobile-overview-dot--' + row.state;
dot.setAttribute('aria-hidden', 'true');
item.appendChild(dot);
const body = document.createElement('span');
body.className = 'mobile-overview-row-body';
const line1 = document.createElement('span');
line1.className = 'mobile-overview-row-title';
const name = document.createElement('span');
// .session-name is in the i18n skip list: a session name is user content.
name.className = 'session-name';
name.textContent = row.name;
line1.appendChild(name);
if (row.caseName) {
const meta = document.createElement('span');
meta.className = 'mobile-overview-row-case case-name';
meta.textContent = ' · ' + row.caseName;
line1.appendChild(meta);
}
body.appendChild(line1);
const line2 = document.createElement('span');
line2.className = 'mobile-overview-row-sub';
line2.setAttribute('data-i18n-skip', '');
line2.textContent = row.mode + (row.dir ? ' · ' + row.dir : '');
body.appendChild(line2);
item.appendChild(body);
const pill = document.createElement('span');
pill.className = 'mobile-overview-pill mobile-overview-pill--' + row.state;
// Skipped by i18n on purpose: the labels are generic single words ("idle",
// "done", "error") that collide with state strings on other surfaces.
pill.setAttribute('data-i18n-skip', '');
pill.textContent = row.pill;
item.appendChild(pill);
const chevron = document.createElement('span');
chevron.className = 'mobile-overview-chevron';
chevron.setAttribute('aria-hidden', 'true');
chevron.textContent = '›';
item.appendChild(chevron);
return item;
},
/** A past conversation. Tapping it resumes, which creates a fresh session. */
_buildMobileOverviewPastRow(row) {
const item = document.createElement('button');
item.type = 'button';
item.className = 'mobile-overview-row mobile-overview-row--past';
item.dataset.moAction = 'resume';
item.dataset.moSession = row.id;
const dot = document.createElement('span');
dot.className = 'mobile-overview-dot mobile-overview-dot--past';
dot.setAttribute('aria-hidden', 'true');
item.appendChild(dot);
const body = document.createElement('span');
body.className = 'mobile-overview-row-body';
const title = document.createElement('span');
// A first prompt is user content, never app copy.
title.className = 'mobile-overview-row-title session-name';
title.textContent = row.title;
body.appendChild(title);
const sub = document.createElement('span');
sub.className = 'mobile-overview-row-sub';
sub.setAttribute('data-i18n-skip', '');
const when = row.at && this._formatTimeAgo ? this._formatTimeAgo(row.at) : '';
sub.textContent = [row.caseName || row.dir, when].filter(Boolean).join(' · ');
body.appendChild(sub);
item.appendChild(body);
const pill = document.createElement('span');
pill.className = 'mobile-overview-pill mobile-overview-pill--past';
pill.textContent = 'resume';
pill.setAttribute('data-i18n-skip', '');
item.appendChild(pill);
const chevron = document.createElement('span');
chevron.className = 'mobile-overview-chevron';
chevron.setAttribute('aria-hidden', 'true');
chevron.textContent = '›';
item.appendChild(chevron);
return item;
},
});
+474
View File
@@ -839,6 +839,20 @@ html.mobile-init .file-browser-panel {
border-color: rgba(96, 165, 250, 0.5);
}
/* Antigravity mode colors on mobile */
.btn-toolbar.btn-run.mode-antigravity,
.btn-toolbar.btn-run-gear.mode-antigravity {
background: #0b2b33;
border-color: rgba(34, 211, 238, 0.3);
color: #cffafe;
}
.btn-toolbar.btn-run.mode-antigravity:active,
.btn-toolbar.btn-run-gear.mode-antigravity:active {
background: #0e7490;
border-color: rgba(34, 211, 238, 0.5);
}
/* Run mode dropdown menu — positioned above toolbar on mobile */
.run-mode-menu {
bottom: 100%;
@@ -1857,6 +1871,33 @@ html.mobile-init .file-browser-panel {
bottom: calc(44px + 2rem + var(--safe-area-bottom));
}
/* File preview window: full screen on phones. --app-height tracks the visual
viewport (KeyboardHandler), so the edit textarea + Save bar stay above the
OS keyboard instead of hiding behind it. Footer/edit bar pad for the home
indicator. */
.file-preview-window {
width: 100vw;
max-width: 100vw;
height: var(--app-height, 100vh);
max-height: var(--app-height, 100vh);
border-radius: 0;
border-left: none;
border-right: none;
}
.file-preview-editbar {
padding-bottom: calc(0.4rem + var(--safe-area-bottom));
}
.file-preview-overlay .file-preview-footer {
padding-bottom: calc(0.35rem + var(--safe-area-bottom));
}
/* >=16px or iOS Safari auto-zooms the page on focus */
.file-preview-body textarea.file-preview-editor {
font-size: 16px;
}
/* Notification drawer - full width on mobile, includes safe area padding */
.notification-drawer {
width: 100%;
@@ -2245,6 +2286,433 @@ html.mobile-init .file-browser-panel {
border-radius: 8px;
transform: none !important;
}
/* ---- Mobile Overview (phone home screen) ----
Shown in place of the welcome overlay when the C logo is tapped. Styled only
with :root tokens so every skin, including the four light ones, works with
no override block. Never give .mobile-overview itself a display value: the
element ships with [hidden] and only .visible may turn it on. */
.mobile-overview.visible {
position: absolute;
top: 0;
left: 0;
right: 0;
bottom: 0;
z-index: 10;
display: flex;
flex-direction: column;
gap: 0.75rem;
padding: 0.75rem 0.6rem calc(1rem + var(--safe-area-bottom));
background: var(--bg-dark);
overflow-y: auto;
-webkit-overflow-scrolling: touch;
overscroll-behavior: contain;
}
/* No side padding: the wordmark's left edge lines up with the session cards
below it, and the Run button's right edge with theirs. */
.mobile-overview-top {
display: flex;
align-items: center;
justify-content: space-between;
gap: 0.5rem;
padding: 0;
}
/* Same accent as the "C" in the header (`.logo` uses --accent-hover), and the
same 48px box as the Run button opposite it, so the two are centered on one
line by construction rather than by eye. */
.mobile-overview-brand {
display: inline-flex;
align-items: center;
min-height: 48px;
font-size: 1.5rem;
font-weight: 800;
line-height: 1;
color: var(--accent-hover);
letter-spacing: -0.02em;
}
/* Split Run button. Colors come from the toolbar's own
`.btn-toolbar.btn-run.mode-<backend>` rules in styles.css (the element
carries those classes), so this button and the one in the toolbar are the
same control in two places. Only size/shape is set here, and NOTHING that
would override the mode gradient. */
.mobile-overview-header {
position: relative;
display: flex;
flex-direction: column;
}
.mobile-overview-run-group {
display: flex;
align-items: stretch;
flex-shrink: 0;
}
/* Sized well past the 44px touch minimum: this is the primary action on the
home screen, and the caret is a separate target right next to it.
⚠️ The !important is required, not decorative: the phone block clamps every
.btn-toolbar to `height/min-height/max-height: 26px !important`, and this
button deliberately carries .btn-toolbar to inherit the run-mode gradient.
Without matching !important (max-height included) it renders 26px tall. */
.mobile-overview-run-group .mobile-overview-run {
display: inline-flex;
align-items: center;
gap: 0.4rem;
min-height: 48px !important;
height: 48px !important;
max-height: 48px !important;
padding: 0 1.15rem !important;
font-size: 0.95rem !important;
line-height: 1;
font-weight: 700;
font-family: inherit;
border-radius: 10px 0 0 10px !important;
}
.mobile-overview-run-mode {
font-weight: 600;
font-size: 0.85rem;
opacity: 0.8;
}
.mobile-overview-run-group .mobile-overview-run-caret {
display: inline-flex;
align-items: center;
justify-content: center;
min-height: 48px !important;
height: 48px !important;
max-height: 48px !important;
min-width: 48px !important;
width: 48px !important;
padding: 0 !important;
line-height: 1;
font-family: inherit;
border-radius: 0 10px 10px 0 !important;
}
/* The phone block shrinks every run-gear glyph to 10px; this one is a 48px
tap target, so it gets a proportionate chevron. `display:block` drops the
inline-baseline gap that would otherwise push it a pixel low. */
.mobile-overview-run-group .mobile-overview-run-caret svg {
display: block;
width: 20px !important;
height: 20px !important;
margin: 0 !important;
}
.mobile-overview-run-menu {
display: flex;
flex-direction: column;
gap: 0.15rem;
margin: 0.5rem 0 0;
padding: 0.35rem;
border: 1px solid var(--border);
border-radius: 12px;
background: var(--floating-bg);
box-shadow: var(--elevated-shadow);
}
.mobile-overview-run-option {
display: flex;
align-items: center;
gap: 0.5rem;
min-height: 44px;
padding: 0 0.6rem;
border: none;
border-radius: 8px;
background: transparent;
color: var(--text);
font-family: inherit;
font-size: 0.82rem;
text-align: left;
}
.mobile-overview-run-option.selected {
background: var(--control-bg-hover);
}
.mobile-overview-run-option:active {
background: var(--bg-hover);
}
.mobile-overview-run-option--add {
color: var(--accent);
}
.mobile-overview-run-header {
padding: 0.4rem 0.6rem 0.2rem;
font-size: 0.6rem;
font-weight: 700;
letter-spacing: 0.09em;
text-transform: uppercase;
color: var(--text-muted);
}
.mobile-overview-section {
display: flex;
flex-direction: column;
gap: 0.4rem;
}
/* Flush left with the wordmark and the cards: one left edge down the screen. */
.mobile-overview-heading {
display: flex;
align-items: center;
gap: 0.35rem;
margin: 0.35rem 0 0;
padding: 0;
font-size: 0.65rem;
font-weight: 700;
letter-spacing: 0.09em;
text-transform: uppercase;
color: var(--text-muted);
}
.mobile-overview-heading-count {
color: var(--text-dim);
font-weight: 600;
}
.mobile-overview-empty {
margin: 0;
padding: 0.5rem 0;
font-size: 0.75rem;
color: var(--text-muted);
}
/* Row: 56px tap target, dot + two-line body + pill + chevron */
.mobile-overview-row {
display: flex;
align-items: center;
gap: 0.6rem;
width: 100%;
min-height: 56px;
padding: 0.5rem 0.7rem;
border: 1px solid var(--border);
border-radius: 12px;
background: var(--bg-card);
color: var(--text);
font-family: inherit;
text-align: left;
}
.mobile-overview-row:active {
background: var(--bg-hover);
}
/* Attention states mirror the session tabs exactly: red blink when the agent
asked something (permission / question), yellow blink when it is waiting for
a prompt. Same hues and same cadence as tab-blink-red / tab-blink-yellow in
styles.css; the keyframes are re-declared here only because a tab's resting
background is transparent while a row's is the card color. */
.mobile-overview-row--needs {
border-color: var(--red);
animation: mobile-overview-blink-red 2.5s ease-in-out infinite;
}
.mobile-overview-row--waiting {
border-color: var(--yellow);
animation: mobile-overview-blink-yellow 3.5s ease-in-out infinite;
}
.mobile-overview-row--error {
border-color: var(--red);
}
@keyframes mobile-overview-blink-red {
0%,
100% {
background: var(--bg-card);
border-color: var(--border);
}
50% {
background: rgba(239, 68, 68, 0.14);
border-color: var(--red);
}
}
@keyframes mobile-overview-blink-yellow {
0%,
100% {
background: var(--bg-card);
border-color: var(--border);
}
50% {
background: rgba(234, 179, 8, 0.12);
border-color: var(--yellow);
}
}
.mobile-overview-row-body {
display: flex;
flex-direction: column;
gap: 0.15rem;
flex: 1;
min-width: 0;
}
.mobile-overview-row-title {
display: block;
font-size: 0.9rem;
font-weight: 600;
color: var(--text);
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
}
.mobile-overview-row-case {
color: var(--text-dim);
font-weight: 500;
}
.mobile-overview-row-sub {
display: block;
font-size: 0.68rem;
color: var(--text-muted);
white-space: nowrap;
overflow: hidden;
text-overflow: ellipsis;
}
.mobile-overview-dot {
flex-shrink: 0;
width: 9px;
height: 9px;
border-radius: 50%;
background: var(--text-muted);
}
.mobile-overview-dot--needs,
.mobile-overview-dot--error {
background: var(--red);
}
.mobile-overview-dot--waiting {
background: var(--yellow);
}
/* Same as .session-tab .tab-status: green when the session is fine, and the
shared `pulse` keyframes while it is working. */
.mobile-overview-dot--working {
background: var(--green);
animation: pulse 1.5s infinite;
will-change: opacity;
}
.mobile-overview-dot--idle {
background: var(--green);
}
.mobile-overview-dot--done {
background: var(--text-muted);
opacity: 0.5;
}
.mobile-overview-pill {
flex-shrink: 0;
display: inline-flex;
align-items: center;
gap: 0.25rem;
padding: 0.2rem 0.45rem;
border: 1px solid var(--border-light);
border-radius: 999px;
font-size: 0.62rem;
font-weight: 600;
color: var(--text-dim);
white-space: nowrap;
}
.mobile-overview-pill--needs,
.mobile-overview-pill--error {
border-color: var(--red);
color: var(--red);
}
.mobile-overview-pill--waiting {
border-color: var(--yellow);
color: var(--yellow);
}
.mobile-overview-pill--idle,
.mobile-overview-pill--working {
border-color: var(--green);
color: var(--green);
}
.mobile-overview-chevron {
flex-shrink: 0;
color: var(--text-muted);
font-size: 1rem;
line-height: 1;
}
/* Past conversations: quieter than a live row, since tapping one starts work */
.mobile-overview-row--past {
background: transparent;
border-style: dashed;
}
.mobile-overview-row--past .mobile-overview-row-title {
font-weight: 500;
color: var(--text-dim);
}
.mobile-overview-dot--past {
background: transparent;
border: 1px solid var(--border-light);
}
.mobile-overview-pill--past {
border-color: var(--border-light);
color: var(--text-muted);
}
.mobile-overview-more {
display: flex;
align-items: center;
justify-content: center;
gap: 0.35rem;
min-height: 40px;
border: none;
border-radius: 10px;
background: var(--control-bg);
color: var(--text-dim);
font-family: inherit;
font-size: 0.72rem;
font-weight: 600;
}
.mobile-overview-more-count {
padding: 0.05rem 0.35rem;
border-radius: 999px;
background: var(--control-bg-hover);
color: var(--text-muted);
}
}
/* The overview's attention blink is an alert, so it stays visible without
motion: hold the alert color instead of animating to it. */
@media (prefers-reduced-motion: reduce) {
.mobile-overview-row--needs,
.mobile-overview-row--waiting {
animation: none;
}
.mobile-overview-row--needs {
background: rgba(239, 68, 68, 0.14);
}
.mobile-overview-row--waiting {
background: rgba(234, 179, 8, 0.12);
}
.mobile-overview-dot--working {
animation: none;
}
}
/* Light-skin compatibility for mobile-only chrome. These components predate
@@ -2286,6 +2754,12 @@ html:is([data-skin="paper-gray"], [data-skin="solarized-light"], [data-skin="cat
color: #ffffff;
}
html:is([data-skin="paper-gray"], [data-skin="solarized-light"], [data-skin="catppuccin-latte"], [data-skin="rose-pine-dawn"]) :is(.btn-toolbar.btn-run.mode-antigravity, .btn-toolbar.btn-run-gear.mode-antigravity) {
background: linear-gradient(135deg, #0e7490, #0891b2);
border-color: #155e75;
color: #ffffff;
}
html:is([data-skin="paper-gray"], [data-skin="solarized-light"], [data-skin="catppuccin-latte"], [data-skin="rose-pine-dawn"]) .btn-toolbar.btn-run-gear {
border-left-color: var(--control-border-hover) !important;
}
+97 -42
View File
@@ -9,7 +9,7 @@
* 5. Audio alerts (Web Audio API beep, user-opt-in)
*
* Features:
* - Per-event-type preferences (enabled, browser, audio, push) with v1→v4 migration
* - Per-event-type preferences (enabled, browser, audio, push) with v1→v5 migration
* - Device-specific defaults (notifications disabled on mobile by default)
* - 5s notification grouping window to batch rapid-fire events
* - 100-notification cap with oldest eviction
@@ -21,7 +21,7 @@
* @param {CodemanApp} app - Reference to the main app instance
*
* @dependency constants.js (STUCK_THRESHOLD_DEFAULT_MS, timing constants)
* @dependency mobile-handlers.js (MobileDetection.getDeviceType for device-specific defaults)
* @dependency mobile-handlers.js (MobileDetection stable handheld identity/device type)
* @loadorder 4 of 15 — loaded after voice-input.js, before keyboard-accessory.js
*/
@@ -65,12 +65,19 @@ class NotificationManager {
});
}
loadPreferences() {
_usesMobilePreferences() {
return (
MobileDetection.isHandheldDevice?.() ??
MobileDetection.getDeviceType() === 'mobile'
);
}
getDefaultPreferences() {
const defaultEventTypes = {
permission_prompt: { enabled: true, browser: true, audio: true, push: false },
elicitation_dialog: { enabled: true, browser: true, audio: true, push: false },
idle_prompt: { enabled: true, browser: true, audio: false, push: false },
stop: { enabled: true, browser: false, audio: false, push: false },
stop: { enabled: false, browser: false, audio: false, push: false },
session_error: { enabled: true, browser: true, audio: false, push: false },
respawn_cycle: { enabled: true, browser: false, audio: false, push: false },
token_milestone: { enabled: true, browser: false, audio: false, push: false },
@@ -80,8 +87,8 @@ class NotificationManager {
};
// Device-specific defaults: mobile has notifications disabled by default
const isMobile = MobileDetection.getDeviceType() === 'mobile';
const defaults = {
const isMobile = this._usesMobilePreferences();
return {
enabled: !isMobile, // Disabled on mobile by default
browserNotifications: !isMobile,
audioAlerts: false,
@@ -92,51 +99,97 @@ class NotificationManager {
muteInfo: false,
// Per-event-type preferences
eventTypes: defaultEventTypes,
_version: 4,
_version: 5,
};
}
/**
* Apply the complete v1→v5 migration to either local or server-hydrated
* preferences. Keeping one normalization path prevents fresh browsers from
* reviving retired drawer-only hook defaults.
*/
normalizePreferences(rawPreferences) {
const defaults = this.getDefaultPreferences();
if (
!rawPreferences ||
typeof rawPreferences !== 'object' ||
Array.isArray(rawPreferences)
) {
return defaults;
}
const prefs = {
...rawPreferences,
eventTypes:
rawPreferences.eventTypes &&
typeof rawPreferences.eventTypes === 'object' &&
!Array.isArray(rawPreferences.eventTypes)
? Object.fromEntries(
Object.entries(rawPreferences.eventTypes).map(([key, value]) => [
key,
value && typeof value === 'object' ? { ...value } : value,
])
)
: undefined,
};
const version = Number.isInteger(prefs._version) ? prefs._version : 0;
// Migrate: v1 had browserNotifications defaulting to false
if (version < 2) {
prefs.browserNotifications = true;
}
// Migrate: v2 -> v3 adds eventTypes
if (version < 3) {
prefs.eventTypes = { ...defaults.eventTypes };
}
// Migrate: v3 -> v4 adds push field to all eventTypes
if (version < 4 && prefs.eventTypes) {
for (const key of Object.keys(prefs.eventTypes)) {
if (prefs.eventTypes[key] && prefs.eventTypes[key].push === undefined) {
prefs.eventTypes[key].push = false;
}
}
}
// Migrate: v4 -> v5 removes the drawer-only Response Complete default.
// Preserve users who opted into any external delivery channel.
if (version < 5) {
const stopPref = prefs.eventTypes?.stop;
if (
stopPref?.enabled === true &&
!stopPref.browser &&
!stopPref.audio &&
!stopPref.push
) {
stopPref.enabled = false;
}
}
return {
...defaults,
...prefs,
eventTypes: { ...defaults.eventTypes, ...prefs.eventTypes },
_version: 5,
};
}
loadPreferences() {
try {
const storageKey = this.getStorageKey();
const saved = localStorage.getItem(storageKey);
if (saved) {
const prefs = JSON.parse(saved);
// Migrate: v1 had browserNotifications defaulting to false
if (!prefs._version || prefs._version < 2) {
prefs.browserNotifications = true;
prefs._version = 2;
}
// Migrate: v2 -> v3 adds eventTypes
if (prefs._version < 3) {
prefs.eventTypes = defaultEventTypes;
prefs._version = 3;
localStorage.setItem(storageKey, JSON.stringify(prefs));
}
// Migrate: v3 -> v4 adds push field to all eventTypes
if (prefs._version < 4) {
if (prefs.eventTypes) {
for (const key of Object.keys(prefs.eventTypes)) {
if (prefs.eventTypes[key] && prefs.eventTypes[key].push === undefined) {
prefs.eventTypes[key].push = false;
}
}
}
prefs._version = 4;
localStorage.setItem(storageKey, JSON.stringify(prefs));
}
// Merge with defaults to ensure all eventTypes exist
return {
...defaults,
...prefs,
eventTypes: { ...defaultEventTypes, ...prefs.eventTypes },
};
const normalized = this.normalizePreferences(JSON.parse(saved));
localStorage.setItem(storageKey, JSON.stringify(normalized));
return normalized;
}
} catch (_e) { /* ignore */ }
return defaults;
return this.getDefaultPreferences();
}
// Get storage key for notification prefs (device-specific)
getStorageKey() {
const isMobile = MobileDetection.getDeviceType() === 'mobile';
return isMobile ? 'codeman-notification-prefs-mobile' : 'codeman-notification-prefs';
return this._usesMobilePreferences()
? 'codeman-notification-prefs-mobile'
: 'codeman-notification-prefs';
}
savePreferences() {
@@ -163,8 +216,10 @@ class NotificationManager {
'exit-gate': 'ralph_complete',
'subagent-spawn': 'subagent_spawn',
'subagent-complete': 'subagent_complete',
'hook-teammate-idle': 'idle_prompt',
'hook-task-completed': 'stop',
// Team lifecycle hooks are agent activity, not session-idle/stop alerts.
// Reuse the existing opt-in agent categories instead of making them noisy.
'hook-teammate-idle': 'subagent_spawn',
'hook-task-completed': 'subagent_complete',
};
const eventTypeKey = categoryToEventType[category] || category;
+186 -3
View File
@@ -426,7 +426,7 @@ Object.assign(CodemanApp.prototype, {
_buildCommandPaletteNewSessionItem(query = '') {
const mode = this.runMode || this._runMode || 'claude';
const labels = { claude: 'Claude', opencode: 'OpenCode', codex: 'Codex', gemini: 'Gemini' };
const labels = { claude: 'Claude', opencode: 'OpenCode', codex: 'Codex', gemini: 'Gemini', antigravity: 'Antigravity' };
const caseName = this._findCommandPaletteCaseMatch(query) || document.getElementById('quickStartCase')?.value || 'testcase';
return {
id: 'new-session',
@@ -3200,6 +3200,9 @@ Object.assign(CodemanApp.prototype, {
if (!overlay || !bodyEl) return;
// Edit mode: reset any prior editor state whenever a preview (re)loads.
this._resetFilePreviewEdit();
// Show overlay with loading state
overlay.classList.add('visible');
titleEl.textContent = filePath;
@@ -3298,6 +3301,13 @@ Object.assign(CodemanApp.prototype, {
bodyEl.innerHTML = `<pre><code>${escapeHtml(data.content)}</code></pre>`;
const truncNote = data.truncated ? ` (showing 500/${data.totalLines} lines)` : '';
footerEl.textContent = `${data.totalLines} lines \u2022 ${this.formatFileSize(data.size)}${truncNote}`;
// Edit affordance only when the server says an edit=1 re-fetch would
// succeed (workspace text file inside the allowlist and size cap).
if (data.editable) {
this.filePreviewEditTarget = { sessionId, filePath };
const editBtn = this.$('filePreviewEditBtn');
if (editBtn) editBtn.hidden = false;
}
}
} catch (err) {
console.error('Failed to preview file:', err);
@@ -3306,6 +3316,8 @@ Object.assign(CodemanApp.prototype, {
},
closeFilePreview() {
if (this.filePreviewEdit?.dirty && !confirm('Discard unsaved changes?')) return;
this._resetFilePreviewEdit();
const overlay = this.$('filePreviewOverlay');
if (overlay) {
overlay.classList.remove('visible');
@@ -3313,6 +3325,172 @@ Object.assign(CodemanApp.prototype, {
this.filePreviewContent = '';
},
// ═══════════════════════════════════════════════════════════════
// File Viewer edit mode (issue #212 — docs/file-viewer-edit-plan.md)
// ═══════════════════════════════════════════════════════════════
_resetFilePreviewEdit() {
this.filePreviewEdit = null;
this.filePreviewEditTarget = null;
const editBtn = this.$('filePreviewEditBtn');
if (editBtn) editBtn.hidden = true;
const editBar = this.$('filePreviewEditBar');
if (editBar) editBar.hidden = true;
const dirtyEl = this.$('filePreviewDirty');
if (dirtyEl) dirtyEl.hidden = true;
const saveBtn = this.$('filePreviewSaveBtn');
if (saveBtn) {
saveBtn.disabled = true;
saveBtn.textContent = 'Save';
}
},
async enterFilePreviewEdit() {
const target = this.filePreviewEditTarget;
if (!target || this.filePreviewEdit) return;
const bodyEl = this.$('filePreviewBody');
const footerEl = this.$('filePreviewFooter');
if (!bodyEl) return;
// Always re-fetch with edit=1: the preview buffer may be line-truncated and
// a truncated buffer must never become an edit buffer. Parse the envelope
// even on non-ok responses so the specific refusal ("too large to edit
// here") reaches the toast instead of a generic failure.
let data;
try {
const res = await fetch(
`/api/sessions/${target.sessionId}/file-content?path=${encodeURIComponent(target.filePath)}&edit=1`
);
const result = await res.json().catch(() => null);
if (!result || result.success !== true) {
throw new Error(result?.error || `Failed to load file for editing (HTTP ${res.status})`);
}
data = result.data;
} catch (err) {
this.showToast(err.message, 'error');
return;
}
this.filePreviewEdit = {
sessionId: target.sessionId,
filePath: target.filePath,
baseHash: data.hash,
eol: data.eol,
original: data.content,
dirty: false,
saving: false,
};
const textarea = document.createElement('textarea');
textarea.className = 'file-preview-editor';
textarea.spellcheck = false;
textarea.setAttribute('autocapitalize', 'off');
textarea.setAttribute('autocorrect', 'off');
textarea.setAttribute('autocomplete', 'off');
textarea.wrap = 'off';
textarea.value = data.content;
textarea.addEventListener('input', () => this._onFilePreviewEditInput());
bodyEl.innerHTML = '';
bodyEl.appendChild(textarea);
// Deliberately no autofocus: on phones that would pop the OS keyboard
// before the user has scrolled to the line they want to change.
const editBtn = this.$('filePreviewEditBtn');
if (editBtn) editBtn.hidden = true;
const editBar = this.$('filePreviewEditBar');
if (editBar) editBar.hidden = false;
if (footerEl) {
const eolNote = data.eol === 'crlf' ? ' • CRLF' : '';
footerEl.textContent = `Editing • ${data.totalLines} lines • ${this.formatFileSize(data.size)}${eolNote}`;
}
},
_onFilePreviewEditInput() {
const edit = this.filePreviewEdit;
if (!edit) return;
const textarea = this.$('filePreviewBody')?.querySelector('textarea.file-preview-editor');
if (!textarea) return;
edit.dirty = textarea.value !== edit.original;
const dirtyEl = this.$('filePreviewDirty');
if (dirtyEl) dirtyEl.hidden = !edit.dirty;
const saveBtn = this.$('filePreviewSaveBtn');
if (saveBtn) saveBtn.disabled = !edit.dirty || edit.saving;
},
cancelFilePreviewEdit() {
const edit = this.filePreviewEdit;
if (!edit) return;
if (edit.dirty && !confirm('Discard unsaved changes?')) return;
const { sessionId, filePath } = edit;
this._resetFilePreviewEdit();
this.openFilePreview(filePath, sessionId);
},
async saveFilePreviewEdit(force = false) {
const edit = this.filePreviewEdit;
if (!edit || edit.saving) return;
const textarea = this.$('filePreviewBody')?.querySelector('textarea.file-preview-editor');
if (!textarea) return;
edit.saving = true;
const saveBtn = this.$('filePreviewSaveBtn');
if (saveBtn) {
saveBtn.disabled = true;
saveBtn.textContent = 'Saving…';
}
const restoreSaveState = () => {
edit.saving = false;
if (saveBtn) saveBtn.textContent = 'Save';
this._onFilePreviewEditInput();
};
let result = null;
let status = 0;
try {
const res = await fetch(`/api/sessions/${edit.sessionId}/file-content`, {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
path: edit.filePath,
content: textarea.value,
baseHash: edit.baseHash,
eol: edit.eol ?? undefined, // Zod .optional() rejects null
force: force || undefined,
}),
});
status = res.status;
result = await res.json().catch(() => null);
} catch (err) {
restoreSaveState();
this.showToast(`Save failed: ${err.message}`, 'error');
return;
}
if (status === 409 || result?.errorCode === 'CONFLICT') {
restoreSaveState();
if (
confirm(
'File changed on disk since you loaded it.\nOK overwrites it with your version; Cancel keeps your draft open.'
)
) {
this.saveFilePreviewEdit(true);
}
return;
}
if (!result || result.success !== true) {
restoreSaveState();
this.showToast(`Save failed: ${result?.error || `HTTP ${status}`}`, 'error');
return;
}
const { sessionId, filePath } = edit;
this._resetFilePreviewEdit();
this.showToast('Saved', 'success');
// Re-open in read mode — re-fetching shows the truth on disk (including the
// server-side EOL normalization) rather than trusting the local buffer.
this.openFilePreview(filePath, sessionId);
},
// ═══════════════════════════════════════════════════════════════
// Attachment Cards (detected documents/images)
// ═══════════════════════════════════════════════════════════════
@@ -3749,8 +3927,13 @@ Object.assign(CodemanApp.prototype, {
},
copyFilePreviewContent() {
if (this.filePreviewContent) {
navigator.clipboard.writeText(this.filePreviewContent).then(() => {
// While editing, copy the live editor buffer (not the stale preview text).
const editTextarea = this.filePreviewEdit
? this.$('filePreviewBody')?.querySelector('textarea.file-preview-editor')
: null;
const content = editTextarea ? editTextarea.value : this.filePreviewContent;
if (content) {
navigator.clipboard.writeText(content).then(() => {
this.showToast('Copied to clipboard', 'success');
}).catch(() => {
this.showToast('Failed to copy', 'error');
+12
View File
@@ -848,6 +848,18 @@ Object.assign(CodemanApp.prototype, {
},
closeSessionOptions() {
// Commit the field the user was still editing BEFORE editingSessionId is
// cleared. The Session Name input saves on blur (and the auto-compact prompt
// on change), and every autosave handler bails out on `!this.editingSessionId`.
// Hiding the modal blurs the focused input on its own, but that happens after
// the id is gone, so Escape / backdrop-click silently dropped what was typed.
// (Clicking the X worked only because mousedown blurs the input first.)
const modal = document.getElementById('sessionOptionsModal');
const focused = document.activeElement;
if (focused && modal && modal.contains(focused) && typeof focused.blur === 'function') {
focused.blur();
}
this.editingSessionId = null;
// Stop run summary auto-refresh if it was running
this.stopRunSummaryAutoRefresh();
+147 -36
View File
@@ -1,5 +1,5 @@
/**
* @fileoverview Quick start (case loading, session spawning for Claude/Shell/OpenCode/Codex/Gemini),
* @fileoverview Quick start (case loading, session spawning for Claude/Shell/OpenCode/Codex/Gemini/Antigravity),
* session options modal (per-session settings, color picker, rename),
* session options tabs (Ralph config tab), case settings (CRUD, links),
* create case modal, and mobile case picker.
@@ -316,6 +316,9 @@ Object.assign(CodemanApp.prototype, {
select.dataset.listenerAdded = 'true';
}
this.setupQuickStartCasePicker();
// The phone overview labels rows with their case name, and a case rename or
// link does not go through the session-tab renderer.
this._refreshMobileOverviewIfVisible?.();
} catch (err) {
console.error('Failed to load cases:', err);
}
@@ -370,7 +373,7 @@ Object.assign(CodemanApp.prototype, {
this._renderSessionTabsImmediate?.();
},
/** Run using the selected mode (Claude Code, OpenCode, Codex, or Gemini) */
/** Run using the selected mode (Claude Code, OpenCode, Codex, Gemini, or Antigravity) */
async run() {
if (this._runInFlight) return;
@@ -394,6 +397,9 @@ Object.assign(CodemanApp.prototype, {
if (mode === 'gemini') {
return await this.runGemini();
}
if (mode === 'antigravity') {
return await this.runAntigravity();
}
if (mode === 'shell') {
return await this.runShell();
}
@@ -435,6 +441,7 @@ Object.assign(CodemanApp.prototype, {
// Load history sessions when menu opens
if (menu.classList.contains('active')) {
this._loadRunModeHistory();
this._refreshRunModeAvailability(menu);
const close = (ev) => {
if (!menu.contains(ev.target)) {
menu.classList.remove('active');
@@ -445,6 +452,25 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* #201: hides run-mode dropdown entries for CLIs that aren't installed, so
* picking one doesn't spawn a session that immediately errors out.
*
* Shell has no external CLI dependency and is never gated, which is also what
* guarantees the menu is never empty. Scoped to `menu` rather than the document:
* `.run-mode-option` is also the class the saved-dashboard rows and the history
* rows use, and a bare querySelector would find whichever came first in the DOM.
*
* Antigravity is in this list even though #201 predates it — it is a run mode
* like the rest, and `agy` is the LEAST likely of the five to be installed.
*/
_refreshRunModeAvailability(menu) {
for (const mode of ['claude', 'opencode', 'codex', 'gemini', 'antigravity']) {
const btn = menu.querySelector(`.run-mode-option[data-mode="${mode}"]`);
if (btn) btn.style.display = this.isCliAvailable(mode) ? 'flex' : 'none';
}
},
async _loadRunModeHistory() {
const container = document.getElementById('runModeHistory');
if (!container) return;
@@ -503,7 +529,7 @@ Object.assign(CodemanApp.prototype, {
gearBtn.className = `btn-toolbar btn-run-gear mode-${mode}`;
}
if (label) {
label.textContent = mode === 'opencode' ? 'Run OC' : mode === 'codex' ? 'Run CX' : mode === 'gemini' ? 'Run GM' : mode === 'shell' ? 'Run SH' : 'Run';
label.textContent = mode === 'opencode' ? 'Run OC' : mode === 'codex' ? 'Run CX' : mode === 'gemini' ? 'Run GM' : mode === 'antigravity' ? 'Run AG' : mode === 'shell' ? 'Run SH' : 'Run';
}
},
@@ -576,13 +602,43 @@ Object.assign(CodemanApp.prototype, {
return startNumber;
},
/**
* Launch progress may use the terminal only on the session-less home screen.
* When another session is active, mutating the shared xterm would serialize
* launch chrome into that session's snapshot during the subsequent switch.
*/
_beginSessionLaunchStatus(message, ansiColor = '1;32') {
const ownsTerminal = !this.activeSessionId;
if (ownsTerminal) {
this.terminal.clear();
this.terminal.writeln(`\x1b[${ansiColor}m ${message}\x1b[0m`);
this.terminal.writeln('');
} else {
this.showToast?.(message, 'info');
}
return ownsTerminal;
},
_appendSessionLaunchStatus(ownsTerminal, message, ansiColor = '90') {
if (!ownsTerminal || this.activeSessionId) return;
this.terminal.writeln(`\x1b[${ansiColor}m ${message}\x1b[0m`);
},
_reportSessionLaunchError(ownsTerminal, message) {
if (ownsTerminal && !this.activeSessionId) {
this.terminal.writeln(`\x1b[1;31m Error: ${message}\x1b[0m`);
} else {
this.showToast?.(message, 'error');
}
},
async runClaude() {
const caseName = document.getElementById('quickStartCase').value || 'testcase';
const tabCount = Math.min(20, Math.max(1, parseInt(document.getElementById('tabCount').value) || 1));
this.terminal.clear();
this.terminal.writeln(`\x1b[1;32m Starting ${tabCount} Claude session(s) in ${caseName}...\x1b[0m`);
this.terminal.writeln('');
const ownsLaunchTerminal = this._beginSessionLaunchStatus(
`Starting ${tabCount} Claude session(s) in ${caseName}...`
);
// Focus terminal NOW, in the synchronous user-gesture context (button click).
// iOS Safari ignores programmatic focus() after any await, so this must happen
// before the first async call. The keyboard opens here and stays open through
@@ -665,7 +721,7 @@ Object.assign(CodemanApp.prototype, {
await this._ensureCreatedSessionVisible(data.data.sessionId, data.data.session);
remoteIds.push(data.data.sessionId);
}
this.terminal.writeln(`\x1b[90m All ${tabCount} remote session(s) ready\x1b[0m`);
this._appendSessionLaunchStatus(ownsLaunchTerminal, `All ${tabCount} remote session(s) ready`);
if (remoteIds[0]) {
await this.selectSession(remoteIds[0]);
this.loadQuickStartCases();
@@ -700,7 +756,7 @@ Object.assign(CodemanApp.prototype, {
const modelOverride = globalSettings.claudeModel || (useOpus1m ? 'opus[1m]' : '');
// Step 1: Create all sessions in parallel
this.terminal.writeln(`\x1b[90m Creating ${tabCount} session(s)...\x1b[0m`);
this._appendSessionLaunchStatus(ownsLaunchTerminal, `Creating ${tabCount} session(s)...`);
const createPromises = sessionNames.map(name =>
fetch('/api/sessions', {
method: 'POST',
@@ -741,12 +797,12 @@ Object.assign(CodemanApp.prototype, {
));
// Step 3: Start all sessions in parallel (biggest speedup)
this.terminal.writeln(`\x1b[90m Starting ${tabCount} session(s) in parallel...\x1b[0m`);
this._appendSessionLaunchStatus(ownsLaunchTerminal, `Starting ${tabCount} session(s) in parallel...`);
await Promise.all(sessionIds.map(id =>
fetch(`/api/sessions/${id}/interactive`, { method: 'POST' })
));
this.terminal.writeln(`\x1b[90m All ${tabCount} sessions ready\x1b[0m`);
this._appendSessionLaunchStatus(ownsLaunchTerminal, `All ${tabCount} sessions ready`);
// Auto-switch to the new session using selectSession (does proper refresh)
if (firstSessionId) {
@@ -756,7 +812,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.focus();
} catch (err) {
this.terminal.writeln(`\x1b[1;31m Error: ${err.message}\x1b[0m`);
this._reportSessionLaunchError(ownsLaunchTerminal, err.message);
}
},
@@ -799,9 +855,10 @@ Object.assign(CodemanApp.prototype, {
const caseName = document.getElementById('quickStartCase').value || 'testcase';
const shellCount = Math.min(20, Math.max(1, parseInt(document.getElementById('shellCount').value) || 1));
this.terminal.clear();
this.terminal.writeln(`\x1b[1;33m Starting ${shellCount} Shell session(s) in ${caseName}...\x1b[0m`);
this.terminal.writeln('');
const ownsLaunchTerminal = this._beginSessionLaunchStatus(
`Starting ${shellCount} Shell session(s) in ${caseName}...`,
'1;33'
);
try {
// Get the case path
@@ -907,7 +964,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.focus();
} catch (err) {
this.terminal.writeln(`\x1b[1;31m Error: ${err.message}\x1b[0m`);
this._reportSessionLaunchError(ownsLaunchTerminal, err.message);
}
},
@@ -918,9 +975,7 @@ Object.assign(CodemanApp.prototype, {
const _runLoc = (this.cases || []).find(c => c.name === caseName)?.location;
const isRemote = _runLoc === 'remote' || _runLoc === 'docker';
this.terminal.clear();
this.terminal.writeln(`\x1b[1;32m Starting OpenCode session in ${caseName}...\x1b[0m`);
this.terminal.writeln('');
const ownsLaunchTerminal = this._beginSessionLaunchStatus(`Starting OpenCode session in ${caseName}...`);
// Focus in sync gesture context (see runClaude comment)
this.terminal.focus();
@@ -930,8 +985,10 @@ Object.assign(CodemanApp.prototype, {
const statusRes = await fetch('/api/opencode/status');
const status = (await statusRes.json()).data;
if (!status.available) {
this.terminal.writeln('\x1b[1;31m OpenCode CLI not found.\x1b[0m');
this.terminal.writeln('\x1b[90m Install with: curl -fsSL https://opencode.ai/install | bash\x1b[0m');
this._reportSessionLaunchError(
ownsLaunchTerminal,
'OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash'
);
return;
}
}
@@ -964,7 +1021,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.focus();
} catch (err) {
this.terminal.writeln(`\x1b[1;31m Error: ${err.message}\x1b[0m`);
this._reportSessionLaunchError(ownsLaunchTerminal, err.message);
}
},
@@ -975,9 +1032,7 @@ Object.assign(CodemanApp.prototype, {
const _runLoc = (this.cases || []).find(c => c.name === caseName)?.location;
const isRemote = _runLoc === 'remote' || _runLoc === 'docker';
this.terminal.clear();
this.terminal.writeln(`\x1b[1;32m Starting Codex session in ${caseName}...\x1b[0m`);
this.terminal.writeln('');
const ownsLaunchTerminal = this._beginSessionLaunchStatus(`Starting Codex session in ${caseName}...`);
this.terminal.focus();
try {
@@ -985,8 +1040,10 @@ Object.assign(CodemanApp.prototype, {
const statusRes = await fetch('/api/codex/status');
const status = (await statusRes.json()).data;
if (!status.available) {
this.terminal.writeln('\x1b[1;31m Codex CLI not found.\x1b[0m');
this.terminal.writeln('\x1b[90m Install with: npm install -g @openai/codex\x1b[0m');
this._reportSessionLaunchError(
ownsLaunchTerminal,
'Codex CLI not found. Install with: npm install -g @openai/codex'
);
return;
}
}
@@ -1003,6 +1060,7 @@ Object.assign(CodemanApp.prototype, {
...(isRemote ? {} : {
codexConfig: {
dangerouslyBypassApprovals: globalSettings.codexDangerouslyBypassApprovals ?? false,
animations: globalSettings.codexAnimationsEnabled ?? false,
renderMode: 'hybrid',
},
...(Object.keys(envOverrides).length > 0 ? { envOverrides } : {}),
@@ -1021,7 +1079,7 @@ Object.assign(CodemanApp.prototype, {
this.terminal.focus();
} catch (err) {
this.terminal.writeln(`\x1b[1;31m Error: ${err.message}\x1b[0m`);
this._reportSessionLaunchError(ownsLaunchTerminal, err.message);
}
},
@@ -1032,9 +1090,7 @@ Object.assign(CodemanApp.prototype, {
const _runLoc = (this.cases || []).find(c => c.name === caseName)?.location;
const isRemote = _runLoc === 'remote' || _runLoc === 'docker';
this.terminal.clear();
this.terminal.writeln(`\x1b[1;32m Starting Gemini session in ${caseName}...\x1b[0m`);
this.terminal.writeln('');
const ownsLaunchTerminal = this._beginSessionLaunchStatus(`Starting Gemini session in ${caseName}...`);
this.terminal.focus();
try {
@@ -1042,8 +1098,10 @@ Object.assign(CodemanApp.prototype, {
const statusRes = await fetch('/api/gemini/status');
const status = (await statusRes.json()).data;
if (!status.available) {
this.terminal.writeln('\x1b[1;31m Gemini CLI not found.\x1b[0m');
this.terminal.writeln('\x1b[90m Install with: npm install -g @google/gemini-cli\x1b[0m');
this._reportSessionLaunchError(
ownsLaunchTerminal,
'Gemini CLI not found. Install with: npm install -g @google/gemini-cli'
);
return;
}
}
@@ -1072,7 +1130,58 @@ Object.assign(CodemanApp.prototype, {
this.terminal.focus();
} catch (err) {
this.terminal.writeln(`\x1b[1;31m Error: ${err.message}\x1b[0m`);
this._reportSessionLaunchError(ownsLaunchTerminal, err.message);
}
},
async runAntigravity() {
const caseName = document.getElementById('quickStartCase').value || 'testcase';
// Remote/docker cases run agy on the OTHER side — skip the local status probe and the
// local-only config/env below (quick-start rejects them for remote cases).
const _runLoc = (this.cases || []).find(c => c.name === caseName)?.location;
const isRemote = _runLoc === 'remote' || _runLoc === 'docker';
const ownsLaunchTerminal = this._beginSessionLaunchStatus(`Starting Antigravity session in ${caseName}...`);
this.terminal.focus();
try {
if (!isRemote) {
const statusRes = await fetch('/api/antigravity/status');
const status = (await statusRes.json()).data;
if (!status.available) {
this._reportSessionLaunchError(
ownsLaunchTerminal,
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash'
);
return;
}
}
const envOverrides = this.buildEnvOverrides(this.getCaseSettings(caseName), this.loadAppSettingsFromStorage());
const res = await fetch('/api/quick-start', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
caseName,
mode: 'antigravity',
sessionName: `w${this._nextCaseSessionStartNumber(caseName)}-${caseName}`,
...(isRemote ? {} : {
antigravityConfig: { dangerouslySkipPermissions: true },
...(Object.keys(envOverrides).length > 0 ? { envOverrides } : {}),
}),
})
});
const data = await res.json();
if (!data.success) throw new Error(data.error || 'Failed to start Antigravity');
await this._ensureCreatedSessionVisible(data.data.sessionId, data.data.session);
if (data.data.sessionId) {
await this.selectSession(data.data.sessionId);
}
this.terminal.focus();
} catch (err) {
this._reportSessionLaunchError(ownsLaunchTerminal, err.message);
}
},
@@ -1088,7 +1197,7 @@ Object.assign(CodemanApp.prototype, {
this.editingSessionId = sessionId;
// Reset to an appropriate tab — Summary for external CLIs (Respawn/Ralph are Claude-only)
const isAltMode = session.mode === 'opencode' || session.mode === 'codex' || session.mode === 'gemini';
const isAltMode = session.mode === 'opencode' || session.mode === 'codex' || session.mode === 'gemini' || session.mode === 'antigravity';
this.switchOptionsTab(isAltMode ? 'summary' : 'respawn');
// Update respawn status display and buttons
@@ -1118,7 +1227,7 @@ Object.assign(CodemanApp.prototype, {
}
// Hide Claude-specific options for external CLI sessions
const isExternalCli = session.mode === 'opencode' || session.mode === 'codex' || session.mode === 'gemini';
const isExternalCli = session.mode === 'opencode' || session.mode === 'codex' || session.mode === 'gemini' || session.mode === 'antigravity';
const claudeOnlyEls = document.querySelectorAll('[data-claude-only]');
claudeOnlyEls.forEach(el => { el.style.display = isExternalCli ? 'none' : ''; });
@@ -2538,6 +2647,8 @@ Object.defineProperty(CodemanApp.prototype, 'runMode', {
},
set(mode) {
this._runMode =
mode === 'opencode' || mode === 'codex' || mode === 'gemini' || mode === 'claude' ? mode : 'claude';
mode === 'opencode' || mode === 'codex' || mode === 'gemini' || mode === 'antigravity' || mode === 'claude'
? mode
: 'claude';
},
});
+107 -4
View File
@@ -312,6 +312,10 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsShowFileViewerButton').checked = settings.showFileViewerButton ?? defaults.showFileViewerButton ?? true;
document.getElementById('appSettingsShowAttachmentsButton').checked = settings.showAttachmentsButton ?? defaults.showAttachmentsButton ?? false;
document.getElementById('appSettingsSkin').value = settings.skin ?? defaults.skin ?? 'daylight-blue';
// Entrance animations. Deliberately NOT part of the settings payload: the
// styles persist to their own localStorage keys via setAnimTheme(), which
// keeps them per-device without touching the .strict() SettingsUpdateSchema.
this._syncEntranceAnimSetting?.();
// WebGL renderer (desktop only — mobile always uses the DOM renderer, so hide
// the toggle there so it can't promise something that won't apply).
document.getElementById('appSettingsWebglRenderer').checked = settings.webglRendererEnabled ?? defaults.webglRendererEnabled ?? true;
@@ -327,6 +331,14 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsShowMultiMonitorButton').checked = settings.showMultiMonitorButton ?? defaults.showMultiMonitorButton ?? false;
document.getElementById('appSettingsShowPlanUsageLimits').checked = this.planUsageChipEnabled(settings);
document.getElementById('appSettingsShowRedrawButton').checked = settings.showRedrawButton ?? defaults.showRedrawButton ?? false;
// Phone overview home screen: only meaningful under 430px, so the row is
// hidden elsewhere rather than offering a toggle that changes nothing.
document.getElementById('appSettingsMobileOverview').checked = settings.mobileOverviewEnabled ?? defaults.mobileOverviewEnabled ?? false;
const phoneOnly = MobileDetection.getDeviceType() === 'mobile' ? '' : 'none';
const mobileOverviewItem = document.getElementById('appSettingsMobileOverviewItem');
if (mobileOverviewItem) mobileOverviewItem.style.display = phoneOnly;
const phoneSection = document.getElementById('appSettingsPhoneSection');
if (phoneSection) phoneSection.style.display = phoneOnly;
// Session Manager, Away Digest and Cron buttons all default OFF (opt-in under
// Display → Header Displays; the Cron button also ships with btn-cron--hidden
// in the template, so an unchecked box and a hidden button stay consistent).
@@ -351,6 +363,7 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsCjkInput').checked = settings.cjkInputEnabled ?? defaults.cjkInputEnabled ?? false;
document.getElementById('appSettingsExtendedKeyboardBar').checked = settings.extendedKeyboardBar ?? false;
document.getElementById('appSettingsTabTwoRows').checked = settings.tabTwoRows ?? defaults.tabTwoRows ?? false;
document.getElementById('appSettingsShowTabDetachButton').checked = settings.showTabDetachButton ?? defaults.showTabDetachButton ?? false;
// Claude CLI settings
const claudeModeSelect = document.getElementById('appSettingsClaudeMode');
const allowedToolsRow = document.getElementById('allowedToolsRow');
@@ -361,9 +374,15 @@ Object.assign(CodemanApp.prototype, {
claudeModeSelect.onchange = () => {
allowedToolsRow.style.display = claudeModeSelect.value === 'allowedTools' ? '' : 'none';
};
// Codex CLI settings
// Codex CLI settings. The inputs are always populated (and always read back
// by saveAppSettings), even when the tab is hidden below, so a user without
// codex installed can never silently wipe the codex prefs of an instance
// that does have it.
document.getElementById('appSettingsCodexDangerouslyBypassApprovals').checked =
settings.codexDangerouslyBypassApprovals ?? false;
document.getElementById('appSettingsCodexAnimations').checked =
settings.codexAnimationsEnabled ?? false;
this._applyCodexSettingsVisibility();
// Claude Permissions settings
document.getElementById('appSettingsAgentTeams').checked = settings.agentTeamsEnabled ?? false;
document.getElementById('appSettingsClaudeModel').value = settings.claudeModel ?? '';
@@ -410,7 +429,7 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('eventIdleAudio').checked = idlePref.audio ?? false;
// Response complete (stop)
const stopPref = eventTypes.stop || {};
document.getElementById('eventStopEnabled').checked = stopPref.enabled ?? true;
document.getElementById('eventStopEnabled').checked = stopPref.enabled ?? false;
document.getElementById('eventStopBrowser').checked = stopPref.browser ?? false;
document.getElementById('eventStopPush').checked = stopPref.push ?? false;
document.getElementById('eventStopAudio').checked = stopPref.audio ?? false;
@@ -473,6 +492,29 @@ Object.assign(CodemanApp.prototype, {
this.activeFocusTrap.activate();
},
/**
* Show the App Settings "Codex CLI" tab only on instances where the codex
* binary actually resolves. Both settings on it (approval bypass, animated
* status effects) are passed to `codex` at launch, so on a box without codex
* the tab is a promise nothing can keep.
*
* Availability comes from the injected `window.__codemanCliAvailable`, shared
* with the welcome buttons and the run-mode dropdown, so the tab never flickers
* in and back out. Only the tab BUTTON is toggled: the panel already carries
* `.modal-tab-content.hidden` unless it is the selected tab, and
* openAppSettings() always reopens on Display, so an unreachable button is
* enough to keep the panel unreachable.
*
* Note the inverted default versus the run buttons: an UNKNOWN flag hides this
* tab. Hiding a settings tab costs a user nothing (the values stay in the DOM
* and are still saved), whereas hiding a run button would leave a working
* install with nothing to click.
*/
_applyCodexSettingsVisibility() {
const btn = document.querySelector('#appSettingsModal .modal-tab-btn[data-tab="settings-codex"]');
if (btn) btn.style.display = window.__codemanCliAvailable?.codex === true ? '' : 'none';
},
switchSettingsTab(tabName) {
const modal = document.getElementById('appSettingsModal');
// Toggle active class on tab buttons
@@ -682,6 +724,43 @@ Object.assign(CodemanApp.prototype, {
this._updatePollTimer = setInterval(poll, 1500);
},
/**
* Is `tool` installed on the server? Reads `window.__codemanCliAvailable`,
* injected by renderIndexHtml (see the comment there for why this is injected
* rather than fetched per surface).
*
* Unknown reads as AVAILABLE. A missing flag means the page was rendered by a
* build that predates the injection, or by a solo popup: hiding every run
* button on a doubt would leave nothing to click, and the pre-existing failure
* mode for a genuinely missing CLI is just an error toast.
*/
isCliAvailable(tool) {
const flags = window.__codemanCliAvailable;
if (!flags || typeof flags !== 'object') return true;
return flags[tool] !== false;
},
/**
* #200: show a welcome-screen button only where the thing it launches exists.
* The markup ships them hidden, so an old cached page can never flash a button
* for a tool this server does not have.
*/
applyWelcomeCliVisibility() {
const buttons = [
['welcomeClaudeBtn', 'claude'],
['welcomeOpencodeBtn', 'opencode'],
['welcomeAntigravityBtn', 'antigravity'],
['welcomeGeminiBtn', 'gemini'],
// Not a run mode, same reasoning: offering a Cloudflare Tunnel on a box
// without cloudflared can only ever produce "cloudflared not found".
['welcomeTunnelBtn', 'cloudflared'],
];
for (const [id, tool] of buttons) {
const btn = document.getElementById(id);
if (btn) btn.style.display = this.isCliAvailable(tool) ? 'flex' : 'none';
}
},
async loadTunnelStatus() {
try {
const res = await fetch('/api/tunnel/status');
@@ -1449,6 +1528,7 @@ Object.assign(CodemanApp.prototype, {
showMultiMonitorButton: document.getElementById('appSettingsShowMultiMonitorButton').checked,
showPlanUsageLimits: document.getElementById('appSettingsShowPlanUsageLimits').checked,
showRedrawButton: document.getElementById('appSettingsShowRedrawButton').checked,
mobileOverviewEnabled: document.getElementById('appSettingsMobileOverview').checked,
showSessionButton: document.getElementById('appSettingsShowSessionButton').checked,
showAwayDigestButton: document.getElementById('appSettingsShowAwayDigestButton').checked,
showCronButton: document.getElementById('appSettingsShowCronButton').checked,
@@ -1463,12 +1543,14 @@ Object.assign(CodemanApp.prototype, {
webglRendererEnabled: document.getElementById('appSettingsWebglRenderer').checked,
extendedKeyboardBar: document.getElementById('appSettingsExtendedKeyboardBar').checked,
tabTwoRows: document.getElementById('appSettingsTabTwoRows').checked,
showTabDetachButton: document.getElementById('appSettingsShowTabDetachButton').checked,
skin: document.getElementById('appSettingsSkin').value,
// Claude CLI settings
claudeMode: document.getElementById('appSettingsClaudeMode').value,
allowedTools: document.getElementById('appSettingsAllowedTools').value.trim(),
// Codex CLI settings
codexDangerouslyBypassApprovals: document.getElementById('appSettingsCodexDangerouslyBypassApprovals').checked,
codexAnimationsEnabled: document.getElementById('appSettingsCodexAnimations').checked,
// Claude Permissions settings
agentTeamsEnabled: document.getElementById('appSettingsAgentTeams').checked,
claudeModel: document.getElementById('appSettingsClaudeModel').value,
@@ -1589,7 +1671,7 @@ Object.assign(CodemanApp.prototype, {
audio: document.getElementById('eventSubagentAudio').checked,
},
},
_version: 4,
_version: 5,
};
if (this.notificationManager) {
this.notificationManager.preferences = notifPrefsToSave;
@@ -1612,6 +1694,10 @@ Object.assign(CodemanApp.prototype, {
// Apply CJK input visibility immediately
this._updateCjkInputState();
// The phone home surface (overview vs welcome) may have just been toggled.
// Only re-decide while a home screen is actually up.
if (!this.activeSessionId) this.showWelcome();
// Apply keyboard bar mode
KeyboardAccessoryBar.setMode(settings.extendedKeyboardBar ? 'extended' : 'simple');
@@ -1642,6 +1728,9 @@ Object.assign(CodemanApp.prototype, {
showSessionButton: _ssb,
showAwayDigestButton: _adb,
showCronButton: _crb,
showTabDetachButton: _tdb,
// Phone-only home surface, and absent from SettingsUpdateSchema (.strict()).
mobileOverviewEnabled: _mov,
...serverSettings
} = settings;
try {
@@ -1810,6 +1899,10 @@ Object.assign(CodemanApp.prototype, {
showSessionButton: false,
showAwayDigestButton: false,
showCronButton: false,
// Phone home screen: the C logo opens the session overview instead of the
// welcome screen. ON by default here, and the escape hatch if it ever
// misbehaves on a device (the gate treats only an explicit false as off).
mobileOverviewEnabled: true,
// Remote auto-reconnect (COD-108) — on by default
remoteAutoReconnect: true,
// Input
@@ -1909,6 +2002,13 @@ Object.assign(CodemanApp.prototype, {
applyHeaderVisibilitySettings() {
const settings = this.loadAppSettingsFromStorage();
const defaults = this.getDefaultSettings();
// Tab pop-out (open-in-new-window) button: opt-in (App Settings → Tab Bar,
// default OFF, per-device). Mirrored as a class on <html>: styles.css hides
// .tab-detach without it (a tab that is already detached keeps its icon as
// the re-focus affordance for the popped-out window).
const showTabDetach = settings.showTabDetachButton ?? defaults.showTabDetachButton ?? false;
document.documentElement.classList.toggle('tabs-show-detach', showTabDetach);
const compactHeader = MobileDetection.getDeviceType() !== 'desktop';
const showFontControls = compactHeader ? false : (settings.showFontControls ?? defaults.showFontControls ?? false);
const showSystemStats = compactHeader ? false : (settings.showSystemStats ?? defaults.showSystemStats ?? true);
@@ -2265,6 +2365,8 @@ Object.assign(CodemanApp.prototype, {
'language',
'terminalWheelLocalScrollback',
'showSessionButton', 'showAwayDigestButton', 'showCronButton',
'showTabDetachButton',
'mobileOverviewEnabled',
]);
// The plan-usage chip is a PER-DEVICE display setting (desktop default ON,
// handheld default OFF): desktop can show it while mobile stays hidden. It
@@ -2295,7 +2397,8 @@ Object.assign(CodemanApp.prototype, {
if (notificationPreferences && this.notificationManager) {
const localNotifPrefs = localStorage.getItem(this.notificationManager.getStorageKey());
if (!localNotifPrefs) {
this.notificationManager.preferences = notificationPreferences;
this.notificationManager.preferences =
this.notificationManager.normalizePreferences(notificationPreferences);
this.notificationManager.savePreferences();
}
}
+817 -4
View File
@@ -328,7 +328,8 @@ html:is([data-skin="paper-gray"], [data-skin="solarized-light"], [data-skin="cat
.search-filter-chip.active,
.search-badge-session,
.history-view-all-btn,
.session-tab .tab-mode.gemini
.session-tab .tab-mode.gemini,
.session-tab .tab-mode.antigravity
) {
color: var(--accent-d);
}
@@ -638,6 +639,674 @@ body {
100% { box-shadow: 0 0 0px 0px rgba(34, 197, 94, 0); background: transparent; }
}
/* ── Tab entrance animations (tab-entrance.js) ─────────────────────────────
Style is picked by html[data-tab-anim]; the stagger arrives as an inline
--tab-enter-delay per tab (negative when an entrance is resuming after a
re-render). Duration is scaled by --anim-enter-scale so one slider retimes
every style.
Colour/glow effects live on ::before, never on the tab itself: .session-tab.active
sets background/border/box-shadow with !important (and a newly created session is
normally the active one), and !important beats an animation. Transform and opacity
are unaffected, so the tab keeps those. */
.session-tab.tab-enter {
animation-delay: var(--tab-enter-delay, 0ms);
animation-fill-mode: both;
will-change: transform, opacity;
}
.session-tab.tab-enter::before {
content: "";
position: absolute;
inset: -1px;
border-radius: inherit;
pointer-events: none;
opacity: 0;
animation-delay: var(--tab-enter-delay, 0ms);
animation-fill-mode: both;
}
/* Slide, drifts in from the right and settles. */
html[data-tab-anim="slide"] .session-tab.tab-enter {
animation-name: tab-enter-slide;
animation-duration: calc(380ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.22, 1, 0.36, 1);
}
@keyframes tab-enter-slide {
from { opacity: 0; transform: translate3d(24px, 3px, 0) scale(0.96); }
to { opacity: 1; transform: none; }
}
/* Pop, springs past full size, then settles back. */
html[data-tab-anim="pop"] .session-tab.tab-enter {
animation-name: tab-enter-pop;
animation-duration: calc(460ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes tab-enter-pop {
0% { opacity: 0; transform: scale(0.4); }
55% { opacity: 1; transform: scale(1.13); }
76% { opacity: 1; transform: scale(0.965); }
100% { opacity: 1; transform: scale(1); }
}
/* CRT, snaps open as a hot line, then unfolds vertically. */
html[data-tab-anim="crt"] .session-tab.tab-enter {
animation-name: tab-enter-crt;
animation-duration: calc(520ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.3, 0.9, 0.3, 1);
}
html[data-tab-anim="crt"] .session-tab.tab-enter::before {
animation-name: tab-enter-crt-flash;
animation-duration: calc(520ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes tab-enter-crt {
0% { opacity: 0; transform: scale3d(0.04, 0.09, 1); }
20% { opacity: 1; transform: scale3d(1, 0.09, 1); }
56% { opacity: 1; transform: scale3d(1, 1.14, 1); }
78% { opacity: 1; transform: scale3d(1, 0.95, 1); }
100% { opacity: 1; transform: none; }
}
@keyframes tab-enter-crt-flash {
0% { opacity: 1; background: rgba(205, 255, 225, 0.95); box-shadow: 0 0 26px 8px rgba(0, 255, 102, 0.6); }
20% { opacity: 0.9; background: rgba(150, 255, 195, 0.7); box-shadow: 0 0 20px 6px rgba(0, 255, 102, 0.45); }
60% { opacity: 0.32; background: rgba(0, 255, 102, 0.14); box-shadow: 0 0 12px 2px rgba(0, 255, 102, 0.22); }
100% { opacity: 0; background: transparent; box-shadow: none; }
}
/* Unroll, the strip makes room and the tab widens into it. */
html[data-tab-anim="unroll"] .session-tab.tab-enter {
animation-name: tab-enter-unroll;
animation-duration: calc(480ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.22, 1, 0.36, 1);
overflow: hidden;
}
@keyframes tab-enter-unroll {
0% { max-width: 0; opacity: 0; padding-left: 0; padding-right: 0; transform: translateY(3px); }
30% { opacity: 1; }
100% { max-width: 900px; opacity: 1; padding-left: 0.6rem; padding-right: 0.6rem; transform: none; }
}
/* Boot, flickers on under a green scan sweep, like a display warming up. */
html[data-tab-anim="boot"] .session-tab.tab-enter {
animation-name: tab-enter-boot;
animation-duration: calc(720ms * var(--anim-enter-scale, 1));
animation-timing-function: linear;
}
html[data-tab-anim="boot"] .session-tab.tab-enter::before {
background: linear-gradient(
100deg,
transparent 0%,
rgba(0, 255, 102, 0.06) 34%,
rgba(200, 255, 220, 0.5) 50%,
rgba(0, 255, 102, 0.06) 66%,
transparent 100%
);
background-size: 280% 100%;
animation-name: tab-enter-boot-scan;
animation-duration: calc(720ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-in-out;
}
@keyframes tab-enter-boot {
0% { opacity: 0; transform: translate3d(0, 5px, 0); }
8% { opacity: 0.85; }
15% { opacity: 0.12; }
23% { opacity: 1; }
31% { opacity: 0.35; }
42% { opacity: 1; transform: none; }
55% { opacity: 0.72; }
66% { opacity: 1; }
100% { opacity: 1; transform: none; }
}
@keyframes tab-enter-boot-scan {
0% { opacity: 0; background-position: 150% 0; }
12% { opacity: 1; }
86% { opacity: 1; }
100% { opacity: 0; background-position: -50% 0; }
}
/* Flip, drops in as a card hinged on its top edge. */
html[data-tab-anim="flip"] .session-tab.tab-enter {
animation-name: tab-enter-flip;
animation-duration: calc(520ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.3, 0.9, 0.3, 1);
transform-origin: 50% 0%;
}
@keyframes tab-enter-flip {
0% { opacity: 0; transform: perspective(700px) rotateX(-92deg); }
55% { opacity: 1; transform: perspective(700px) rotateX(13deg); }
78% { opacity: 1; transform: perspective(700px) rotateX(-5deg); }
100% { opacity: 1; transform: perspective(700px) rotateX(0deg); }
}
/* ── Entrance lab (?animlab=1) ───────────────────────────────────────────── */
.anim-lab {
position: fixed;
right: 14px;
bottom: 14px;
z-index: 3100;
width: 300px;
max-height: 88vh;
display: flex;
flex-direction: column;
padding: 10px 12px 12px;
border: 1px solid var(--border-light);
border-radius: 10px;
background: rgba(12, 16, 14, 0.96);
box-shadow: 0 12px 40px rgba(0, 0, 0, 0.55);
color: var(--text);
font-size: 0.72rem;
backdrop-filter: blur(6px);
}
.anim-lab-head {
display: flex;
align-items: center;
justify-content: space-between;
margin-bottom: 8px;
font-weight: 600;
letter-spacing: 0.02em;
}
.anim-lab-close {
background: none;
border: none;
color: var(--text-dim);
font-size: 1.1rem;
line-height: 1;
cursor: pointer;
}
.anim-lab-close:hover { color: var(--text); }
.anim-lab-themes {
display: flex;
flex-wrap: wrap;
gap: 4px;
margin-bottom: 10px;
}
.anim-lab-themes button {
padding: 3px 8px;
border: 1px solid var(--border-light);
border-radius: 999px;
background: rgba(255, 255, 255, 0.04);
color: var(--text-dim);
font-size: 0.66rem;
cursor: pointer;
}
.anim-lab-themes button:hover { background: rgba(34, 197, 94, 0.16); color: var(--text); }
.anim-lab-themes button.selected {
border-color: rgba(0, 255, 102, 0.7);
background: rgba(34, 197, 94, 0.2);
color: #fff;
}
.anim-lab-scroll {
flex: 1;
min-height: 0;
overflow-y: auto;
margin-bottom: 8px;
}
.anim-lab-scroll::-webkit-scrollbar { width: 5px; }
.anim-lab-scroll::-webkit-scrollbar-thumb { background: var(--border); border-radius: 3px; }
.anim-lab-group { margin-bottom: 10px; }
.anim-lab-group-title {
margin-bottom: 4px;
font-size: 0.62rem;
font-weight: 700;
letter-spacing: 0.08em;
text-transform: uppercase;
color: var(--text-muted);
}
.anim-lab-style {
display: flex;
flex-direction: column;
gap: 1px;
width: 100%;
margin-bottom: 3px;
padding: 5px 8px;
border: 1px solid transparent;
border-radius: 6px;
background: rgba(255, 255, 255, 0.03);
color: var(--text-dim);
text-align: left;
cursor: pointer;
}
.anim-lab-style:hover { background: rgba(34, 197, 94, 0.09); color: var(--text); }
.anim-lab-style.selected {
border-color: rgba(0, 255, 102, 0.6);
background: rgba(34, 197, 94, 0.14);
color: #fff;
}
.anim-lab-style strong { font-size: 0.73rem; font-weight: 600; }
.anim-lab-style em { font-style: normal; font-size: 0.64rem; opacity: 0.7; }
.anim-lab-range {
display: grid;
grid-template-columns: auto 1fr;
align-items: center;
gap: 4px 8px;
margin-bottom: 6px;
color: var(--text-dim);
}
.anim-lab-range output { justify-self: end; color: var(--text); font-variant-numeric: tabular-nums; }
.anim-lab-range input { grid-column: 1 / -1; width: 100%; accent-color: #00ff66; }
.anim-lab-demo {
display: flex;
align-items: center;
gap: 6px;
margin-top: 8px;
color: var(--text-dim);
}
.anim-lab-demo button {
flex: 1;
padding: 4px 0;
border: 1px solid var(--border-light);
border-radius: 6px;
background: rgba(255, 255, 255, 0.04);
color: var(--text);
cursor: pointer;
}
.anim-lab-demo button:hover { background: rgba(34, 197, 94, 0.16); border-color: rgba(0, 255, 102, 0.5); }
.anim-lab-hint { margin: 8px 0 0; font-size: 0.64rem; opacity: 0.55; }
/* ── Window entrance animations ────────────────────────────────────────────
Applied to a window already sitting at its resting position. The `fly` style
is NOT here: it is the pre-existing JS transition in subagent-windows.js that
flies the window out of its parent tab, and it skips this class entirely. */
.subagent-window.win-enter,
.ultracode-window.win-enter {
animation-delay: var(--win-enter-delay, 0ms);
animation-fill-mode: both;
will-change: transform, opacity, filter;
}
.subagent-window.win-enter::before,
.ultracode-window.win-enter::before {
content: "";
position: absolute;
inset: 0;
z-index: 5;
border-radius: inherit;
pointer-events: none;
opacity: 0;
animation-delay: var(--win-enter-delay, 0ms);
animation-fill-mode: both;
}
/* CRT, bursts open as a hot line, then unfolds vertically. */
html[data-win-anim="crt"] .subagent-window.win-enter,
html[data-win-anim="crt"] .ultracode-window.win-enter {
animation-name: win-enter-crt;
animation-duration: calc(560ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.3, 0.9, 0.3, 1);
}
html[data-win-anim="crt"] .subagent-window.win-enter::before,
html[data-win-anim="crt"] .ultracode-window.win-enter::before {
animation-name: win-enter-crt-flash;
animation-duration: calc(560ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes win-enter-crt {
0% { opacity: 0; transform: scale3d(0.55, 0.008, 1); }
22% { opacity: 1; transform: scale3d(1, 0.01, 1); }
58% { opacity: 1; transform: scale3d(1, 1.05, 1); }
80% { opacity: 1; transform: scale3d(1, 0.98, 1); }
100% { opacity: 1; transform: none; }
}
@keyframes win-enter-crt-flash {
0% { opacity: 1; background: rgba(205, 255, 225, 0.9); box-shadow: 0 0 40px 14px rgba(0, 255, 102, 0.55); }
22% { opacity: 0.8; background: rgba(120, 255, 180, 0.5); box-shadow: 0 0 30px 10px rgba(0, 255, 102, 0.4); }
60% { opacity: 0.25; background: rgba(0, 255, 102, 0.1); box-shadow: 0 0 14px 3px rgba(0, 255, 102, 0.18); }
100% { opacity: 0; background: transparent; box-shadow: none; }
}
/* Materialize, resolves out of a blur with a short glitch. */
html[data-win-anim="materialize"] .subagent-window.win-enter,
html[data-win-anim="materialize"] .ultracode-window.win-enter {
animation-name: win-enter-materialize;
animation-duration: calc(620ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.22, 1, 0.36, 1);
}
@keyframes win-enter-materialize {
0% { opacity: 0; transform: scale(0.94); filter: blur(16px) brightness(1.7); }
35% { opacity: 0.85; filter: blur(6px) brightness(1.3); }
45% { opacity: 0.4; }
55% { opacity: 0.95; filter: blur(3px) brightness(1.15); }
70% { opacity: 0.75; }
100% { opacity: 1; transform: none; filter: none; }
}
/* Unfold, hinges down from the top edge. */
html[data-win-anim="unfold"] .subagent-window.win-enter,
html[data-win-anim="unfold"] .ultracode-window.win-enter {
animation-name: win-enter-unfold;
animation-duration: calc(560ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.3, 0.9, 0.3, 1);
transform-origin: 50% 0%;
}
@keyframes win-enter-unfold {
0% { opacity: 0; transform: perspective(1600px) rotateX(-88deg); }
55% { opacity: 1; transform: perspective(1600px) rotateX(9deg); }
80% { opacity: 1; transform: perspective(1600px) rotateX(-3deg); }
100% { opacity: 1; transform: perspective(1600px) rotateX(0deg); }
}
/* Beam down, holds still (so its connection line can draw toward a stable rect),
then materializes out of a green wash. NO transform: a transformed window
reports a transformed rect and the line would aim at the wrong place. */
html[data-win-anim="beam"] .subagent-window.win-enter,
html[data-win-anim="beam"] .ultracode-window.win-enter {
animation-name: win-enter-beam;
animation-duration: calc(620ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
html[data-win-anim="beam"] .subagent-window.win-enter::before,
html[data-win-anim="beam"] .ultracode-window.win-enter::before {
animation-name: win-enter-beam-wash;
animation-duration: calc(620ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes win-enter-beam {
0% { opacity: 0; filter: brightness(2.6) saturate(0.4); }
20% { opacity: 0.55; }
32% { opacity: 0.2; }
48% { opacity: 0.9; filter: brightness(1.5) saturate(0.8); }
100% { opacity: 1; filter: none; }
}
@keyframes win-enter-beam-wash {
0% { opacity: 0.9; background: linear-gradient(180deg, rgba(190, 255, 215, 0.85), rgba(0, 255, 102, 0)); }
45% { opacity: 0.5; background: linear-gradient(180deg, rgba(0, 255, 102, 0.35), rgba(0, 255, 102, 0)); }
100% { opacity: 0; background: transparent; }
}
/* Pop, springs open from the centre. */
html[data-win-anim="pop"] .subagent-window.win-enter,
html[data-win-anim="pop"] .ultracode-window.win-enter {
animation-name: win-enter-pop;
animation-duration: calc(460ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes win-enter-pop {
0% { opacity: 0; transform: scale(0.55); }
58% { opacity: 1; transform: scale(1.045); }
80% { opacity: 1; transform: scale(0.99); }
100% { opacity: 1; transform: scale(1); }
}
/* ── Main terminal pane entrance animations ────────────────────────────────
⚠ transform / opacity / clip-path ONLY. xterm's FitAddon derives rows+cols
from getComputedStyle(parent).width/height, the untransformed layout box -
so these are invisible to it, but animating width/height/padding here would
resize the PTY mid-animation. Colour washes go on ::before, never as a
`filter` on the container: that would blur a full-screen WebGL canvas every
frame. `will-change` is deliberately not set, the base rule needs its
`will-change: contents` for terminal compositing. */
.terminal-container.term-enter {
animation-fill-mode: both;
}
.terminal-container.term-enter::before {
content: "";
position: absolute;
inset: 0;
z-index: 4;
pointer-events: none;
opacity: 0;
animation-fill-mode: both;
}
/* CRT, power-on: a hot line that expands to full height. */
html[data-term-anim="crt"] .terminal-container.term-enter {
animation-name: term-enter-crt;
animation-duration: calc(560ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.3, 0.9, 0.3, 1);
}
html[data-term-anim="crt"] .terminal-container.term-enter::before {
animation-name: term-enter-crt-flash;
animation-duration: calc(560ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes term-enter-crt {
0% { opacity: 0; transform: scale3d(0.6, 0.004, 1); }
20% { opacity: 1; transform: scale3d(1, 0.006, 1); }
58% { opacity: 1; transform: scale3d(1, 1.03, 1); }
80% { opacity: 1; transform: scale3d(1, 0.99, 1); }
100% { opacity: 1; transform: none; }
}
@keyframes term-enter-crt-flash {
0% { opacity: 1; background: rgba(210, 255, 230, 0.85); }
20% { opacity: 0.7; background: rgba(120, 255, 180, 0.4); }
60% { opacity: 0.2; background: rgba(0, 255, 102, 0.08); }
100% { opacity: 0; background: transparent; }
}
/* Boot, flickers on under a green scan sweep. */
html[data-term-anim="boot"] .terminal-container.term-enter {
animation-name: term-enter-boot;
animation-duration: calc(760ms * var(--anim-enter-scale, 1));
animation-timing-function: linear;
}
html[data-term-anim="boot"] .terminal-container.term-enter::before {
background: linear-gradient(
180deg,
transparent 0%,
rgba(0, 255, 102, 0.05) 40%,
rgba(200, 255, 220, 0.28) 50%,
rgba(0, 255, 102, 0.05) 60%,
transparent 100%
);
background-size: 100% 300%;
animation-name: term-enter-boot-scan;
animation-duration: calc(760ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-in-out;
}
@keyframes term-enter-boot {
0% { opacity: 0; transform: scale(0.995); }
8% { opacity: 0.85; }
15% { opacity: 0.1; }
24% { opacity: 1; }
33% { opacity: 0.35; }
45% { opacity: 1; transform: none; }
58% { opacity: 0.7; }
70% { opacity: 1; }
100% { opacity: 1; transform: none; }
}
@keyframes term-enter-boot-scan {
0% { opacity: 0; background-position: 0 -150%; }
12% { opacity: 1; }
86% { opacity: 1; }
100% { opacity: 0; background-position: 0 150%; }
}
/* Wipe, reveals top-to-bottom behind a bright edge. */
html[data-term-anim="wipe"] .terminal-container.term-enter {
animation-name: term-enter-wipe;
animation-duration: calc(520ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.35, 0.85, 0.3, 1);
}
html[data-term-anim="wipe"] .terminal-container.term-enter::before {
background: linear-gradient(180deg, rgba(0, 255, 102, 0.16) 0%, rgba(190, 255, 215, 0.5) 82%, transparent 100%);
animation-name: term-enter-wipe-edge;
animation-duration: calc(520ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.35, 0.85, 0.3, 1);
}
@keyframes term-enter-wipe {
0% { opacity: 1; clip-path: inset(0 0 100% 0); }
100% { opacity: 1; clip-path: inset(0 0 0 0); }
}
@keyframes term-enter-wipe-edge {
0% { opacity: 0.9; clip-path: inset(0 0 100% 0); }
80% { opacity: 0.45; }
100% { opacity: 0; clip-path: inset(0 0 0 0); }
}
/* Slide up, rises into place from below. */
html[data-term-anim="slide"] .terminal-container.term-enter {
animation-name: term-enter-slide;
animation-duration: calc(420ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.22, 1, 0.36, 1);
}
@keyframes term-enter-slide {
0% { opacity: 0; transform: translate3d(0, 34px, 0); }
100% { opacity: 1; transform: none; }
}
/* Fade, quiet, with a touch of scale. */
html[data-term-anim="fade"] .terminal-container.term-enter {
animation-name: term-enter-fade;
animation-duration: calc(340ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes term-enter-fade {
0% { opacity: 0; transform: scale(0.985); }
100% { opacity: 1; transform: none; }
}
/* ── Connection-line entrance animations ───────────────────────────────────
`--line-len` is the measured path length, stamped inline by
_applyLineEntrances(); `--line-enter-delay` is negative when an entrance is
resuming after the SVG was rebuilt underneath it. */
.connection-line.line-enter {
animation-delay: var(--line-enter-delay, 0ms);
animation-fill-mode: both;
}
/* Draw, the line paints itself from the tab down to the window. Overrides the
resting 5 3 dash to a single full-length dash for the duration. */
html[data-line-anim="draw"] .connection-line.line-enter {
stroke-dasharray: var(--line-len);
animation-name: line-enter-draw;
animation-duration: calc(420ms * var(--anim-enter-scale, 1));
animation-timing-function: cubic-bezier(0.32, 0.8, 0.3, 1);
}
@keyframes line-enter-draw {
from { stroke-dashoffset: var(--line-len); opacity: 1; }
to { stroke-dashoffset: 0; opacity: 0.9; }
}
/* Fade, plain opacity ramp, dashes intact. */
html[data-line-anim="fade"] .connection-line.line-enter {
animation-name: line-enter-fade;
animation-duration: calc(300ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
@keyframes line-enter-fade {
from { opacity: 0; }
to { opacity: 0.9; }
}
/* Packet, the base line fades in and a separate bright dash rides down it. The
packet is its own cloned path so the base line keeps its dashed look. */
html[data-line-anim="packet"] .connection-line.line-enter {
animation-name: line-enter-fade;
animation-duration: calc(300ms * var(--anim-enter-scale, 1));
animation-timing-function: ease-out;
}
.connection-line-packet {
fill: none;
stroke: #cdfbe3;
stroke-width: 4;
stroke-linecap: round;
pointer-events: none;
stroke-dasharray: 20px 100000px;
filter: drop-shadow(0 0 4px rgba(0, 255, 102, 0.95)) drop-shadow(0 0 10px rgba(59, 130, 246, 0.7));
animation-name: line-enter-packet;
animation-duration: calc(700ms * var(--anim-enter-scale, 1));
animation-delay: var(--line-enter-delay, 0ms);
animation-timing-function: cubic-bezier(0.4, 0, 0.3, 1);
animation-fill-mode: both;
}
@keyframes line-enter-packet {
0% { stroke-dashoffset: 20px; opacity: 0; }
12% { opacity: 1; }
85% { opacity: 1; }
100% { stroke-dashoffset: calc(-1 * var(--line-len)); opacity: 0; }
}
.anim-lab-check {
display: flex;
align-items: center;
gap: 6px;
margin: -4px 0 12px;
padding: 0 2px;
color: var(--text-dim);
font-size: 0.66rem;
cursor: pointer;
}
.anim-lab-check input { accent-color: #00ff66; }
@media (prefers-reduced-motion: reduce) {
.session-tab.tab-enter,
.session-tab.tab-enter::before,
.subagent-window.win-enter,
.subagent-window.win-enter::before,
.ultracode-window.win-enter,
.ultracode-window.win-enter::before,
.terminal-container.term-enter,
.terminal-container.term-enter::before,
.connection-line.line-enter,
.connection-line-packet {
animation: none !important;
}
}
.session-tab .tab-status {
width: 6px;
height: 6px;
@@ -752,7 +1421,11 @@ body {
transition: opacity 0.05s ease-out, width 0.05s ease-out, padding 0.05s ease-out;
}
.session-tab:hover .tab-close {
/* Icons expand on the ACTIVE tab only (in flow): selection is a deliberate
click, so the width change never happens while aiming at a tab, background
tabs keep their full title on hover, and a stray click can only switch.
Middle-click closes any tab (_setupTabMiddleClickClose in app.js). */
.session-tab.active .tab-close {
opacity: 1;
width: auto;
padding: 0.15rem 0.35rem;
@@ -1314,7 +1987,7 @@ body {
transition: opacity 0.15s, width 0.15s, padding 0.15s, transform 0.2s;
}
.session-tab:hover .tab-gear {
.session-tab.active .tab-gear {
opacity: 1;
width: auto;
padding: 0 0.3rem;
@@ -1339,7 +2012,7 @@ body {
overflow: hidden;
transition: opacity 0.15s, width 0.15s, padding 0.15s;
}
.session-tab:hover .tab-detach {
.session-tab.active .tab-detach {
opacity: 1;
width: auto;
padding: 0 0.3rem;
@@ -1378,6 +2051,29 @@ body {
display: inline-flex;
}
/* ===== Tab action icons: active tab only =================================
All three per-tab icons live in a .tab-actions wrapper (in flow; it adds
no width of its own while the children keep the width:0 collapse above).
They expand only on the ACTIVE tab: selection is a deliberate click, so
the tab-strip geometry never shifts while the pointer is aiming, hovering
a background tab changes nothing (full title stays readable), and a stray
click can only switch sessions. Middle-click closes any tab. The phone
layout in mobile.css follows the same active-only pattern with its own
sizing; the tablet touch fallback there keeps icons always visible. */
.session-tab .tab-actions {
display: flex;
align-items: center;
}
/* Pop-out button is opt-in (App Settings → Tab Bar, default off; per-device).
settings-ui.js mirrors the setting as the tabs-show-detach class on <html>.
A tab that is ALREADY detached keeps its icon regardless: it is the
re-focus affordance for the popped-out window. */
html:not(.tabs-show-detach) .session-tab:not(.detached) .tab-detach {
display: none;
}
/* ===== Solo (detached single-session) window chrome ===================== */
body.solo-mode .session-tabs,
body.solo-mode .header-system-stats,
@@ -1447,6 +2143,11 @@ body.solo-mode .btn-lifecycle-log {
color: #8ab4f8;
}
.session-tab .tab-mode.antigravity {
background: rgba(34, 211, 238, 0.2);
color: #22d3ee;
}
/* Timer Banner - Compact */
.timer-banner {
display: flex;
@@ -2631,6 +3332,23 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
transform: translateY(-1px);
}
/* Antigravity: cyan identity, matching .btn-toolbar.btn-run.mode-antigravity and
.run-mode-dot.antigravity so the welcome action reads as the same backend. */
.welcome-btn-antigravity {
background: linear-gradient(135deg, #0b2b33 0%, #0e7490 55%, #0891b2 100%);
border-color: rgba(34, 211, 238, 0.4);
color: #cffafe;
box-shadow: 0 2px 8px rgba(34, 211, 238, 0.16), inset 0 1px 0 rgba(255, 255, 255, 0.06);
}
.welcome-btn-antigravity:hover {
background: linear-gradient(135deg, #124450 0%, #0891b2 55%, #06b6d4 100%);
box-shadow: 0 4px 20px rgba(34, 211, 238, 0.3), 0 0 40px rgba(8, 145, 178, 0.12), inset 0 1px 0 rgba(255, 255, 255, 0.08);
border-color: rgba(103, 232, 249, 0.5);
color: #ecfeff;
transform: translateY(-1px);
}
.welcome-btn-gemini {
background: linear-gradient(135deg, #10243f 0%, #174ea6 55%, #4f46e5 100%);
border-color: rgba(96, 165, 250, 0.4);
@@ -3531,6 +4249,22 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
color: #eff6ff;
}
/* Antigravity mode colors */
.btn-toolbar.btn-run.mode-antigravity,
.btn-toolbar.btn-run-gear.mode-antigravity {
background: linear-gradient(135deg, #0b2b33 0%, #0e7490 55%, #0891b2 100%);
border-color: rgba(34, 211, 238, 0.5);
color: #cffafe;
box-shadow: 0 1px 2px rgba(0, 0, 0, 0.2), inset 0 1px 0 rgba(255, 255, 255, 0.06);
}
.btn-toolbar.btn-run.mode-antigravity:hover,
.btn-toolbar.btn-run-gear.mode-antigravity:hover {
background: linear-gradient(135deg, #124450 0%, #0891b2 55%, #06b6d4 100%);
box-shadow: 0 0 12px rgba(34, 211, 238, 0.35), 0 2px 8px rgba(8, 145, 178, 0.2), inset 0 1px 0 rgba(255, 255, 255, 0.08);
border-color: rgba(103, 232, 249, 0.6);
color: #ecfeff;
}
/* Dropdown menu */
.run-mode-menu {
display: none;
@@ -3590,6 +4324,7 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
.run-mode-dot.opencode { background: #10b981; }
.run-mode-dot.codex { background: #a855f7; }
.run-mode-dot.gemini { background: #8ab4f8; }
.run-mode-dot.antigravity { background: #22d3ee; }
.run-mode-dot.shell { background: #94a3b8; }
/* Phone-only Enter button (see index.html). Hidden by default at every width;
@@ -8739,6 +9474,84 @@ kbd {
flex-shrink: 0;
}
/* ---- File Viewer edit mode (issue #212) ---- */
.file-preview-body textarea.file-preview-editor {
display: block;
width: 100%;
height: 100%;
margin: 0;
padding: 0.75rem;
border: none;
outline: none;
resize: none;
background: var(--bg-dark);
color: var(--text);
font-family: var(--font-mono);
font-size: 0.8rem;
line-height: 1.5;
white-space: pre;
overflow-wrap: normal;
overflow: auto;
tab-size: 4;
}
.file-preview-editbar {
display: flex;
align-items: center;
gap: 0.5rem;
padding: 0.4rem 0.75rem;
border-top: 1px solid var(--border);
background: var(--bg-input);
flex-shrink: 0;
}
.file-preview-editbar[hidden] {
display: none;
}
.file-preview-editbar-spacer {
flex: 1;
}
.file-preview-dirty {
font-size: 0.7rem;
color: var(--warning, #e5c07b);
}
.file-preview-dirty::before {
content: '\25CF ';
}
.file-preview-editbar-btn {
padding: 0.3rem 0.9rem;
font-size: 0.75rem;
border-radius: 6px;
border: 1px solid var(--control-border);
background: var(--bg-input);
color: var(--text);
cursor: pointer;
}
.file-preview-editbar-btn:hover {
background: var(--bg-hover, rgba(255, 255, 255, 0.08));
}
.file-preview-editbar-btn--save {
background: var(--accent);
border-color: var(--accent);
color: #fff;
}
.file-preview-editbar-btn--save:hover {
background: var(--accent-hover);
}
.file-preview-editbar-btn--save:disabled {
opacity: 0.45;
cursor: default;
}
/* ========== Log Viewer Windows (Floating) ========== */
.log-viewer-window {
+27 -3
View File
@@ -465,6 +465,10 @@ Object.assign(CodemanApp.prototype, {
if (typeof this._appendUltracodeAgentConnectionLines === 'function') {
this._appendUltracodeAgentConnectionLines(svg, rects);
}
// Every path above was just created from scratch, so any line entrance in
// flight has to be re-attached here (resumed via a negative animation-delay).
this._applyLineEntrances?.(svg);
},
// ═══════════════════════════════════════════════════════════════
@@ -729,12 +733,17 @@ Object.assign(CodemanApp.prototype, {
</div>
`;
// Only the `fly` entrance style parks the window on its parent tab; every
// CSS-animated style (entrance-animations.js) needs it to start at its
// resting position, or the animation would play at the wrong place.
const flyFromTab = !!parentTab && !isMobile && this.windowEntranceFliesFromTab?.() !== false;
// If we have a parent tab, start window at tab position for spawn animation
if (isMobile) {
// Mobile: position using top (keyboard-aware positioning calculated above)
win.style.top = `${finalY}px`;
win.style.bottom = 'auto';
} else if (parentTab) {
} else if (flyFromTab) {
const tabRect = parentTab.getBoundingClientRect();
win.style.left = `${tabRect.left}px`;
win.style.top = `${tabRect.bottom}px`;
@@ -812,7 +821,7 @@ Object.assign(CodemanApp.prototype, {
this.subagentWindows.get(agentId).resizeObserver = resizeObserver;
// Animate to final position if spawning from tab (desktop only)
if (parentTab && !isMobile) {
if (flyFromTab) {
requestAnimationFrame(() => {
win.style.transition = 'all 0.4s cubic-bezier(0.34, 1.56, 0.64, 1)';
win.style.left = `${finalX}px`;
@@ -824,11 +833,26 @@ Object.assign(CodemanApp.prototype, {
setTimeout(() => {
win.style.transition = '';
win.classList.remove('spawning');
// The line only becomes meaningful once the window has landed, so its
// draw-in starts here rather than at spawn.
this.markConnectionLineEntering?.(agentId);
this.updateConnectionLines();
}, 400);
});
} else {
// No animation (mobile uses CSS positioning), just update connection lines
// CSS-animated entrance styles (and mobile) run in place. The window is
// already at its resting position, so the connection line can be drawn
// against a correct rect right away and animate alongside it.
//
// Skipped when the window spawns hidden (its agent belongs to a background
// tab): a display:none element never runs its animation, so `animationend`
// would never fire and the entrance class plus its inline custom property
// would stick to the window forever. Nothing is visible to animate anyway,
// and revealing it later is a tab switch, not a spawn.
if (!shouldHide) {
this.applyWindowEntrance?.(win);
this.markConnectionLineEntering?.(agentId);
}
this.updateConnectionLines();
}
+543 -19
View File
@@ -23,6 +23,44 @@
// short window, only the app's synthetic tap-to-position mouse event should
// reach xterm.
const TOUCH_COMPAT_MOUSE_SUPPRESS_MS = 450;
// Escape sequences occupy no terminal cells, so they must come out before a
// captured line's WIDTH can be measured (_estimateReplayRows). Covers OSC,
// CSI, charset designators and the short escapes tmux emits; deliberately
// approximate — this feeds a size comparison, not a renderer.
// eslint-disable-next-line no-control-regex
const REPLAY_ESCAPE_RE =
/\x1b\][^\x07\x1b]*(?:\x07|\x1b\\)|\x1b\[[0-9;?<>=!]*[ -/]*[@-~]|\x1b[()#][0-9A-Za-z]|\x1b[=>78M]/g;
// PageUp / PageDown as xterm.js encodes them. Used as the LAST-RESORT scroll
// gesture for a repaint-mode CLI whose local buffer holds no scrollback
// (_maybePageCliTranscript).
const KEY_PAGE_UP = '\x1b[5~';
const KEY_PAGE_DOWN = '\x1b[6~';
// Wheel/touch travel (in lines) that adds up to one PageUp/PageDown. Half a
// screen rather than a full one: the page key always jumps a whole screen, so
// a 1:1 mapping made the fallback feel unreachably slow with a discrete mouse
// wheel (Firefox reports 3 lines a notch → 12 notches per page). Overshooting
// the finger is the right trade against a gesture that otherwise does nothing.
const PAGE_KEY_SCREEN_FRACTION = 0.5;
// Bound on page keys emitted from one gesture batch, mirroring the SGR tick
// cap: a fling must not build a backlog that keeps paging after it stops.
const PAGE_KEY_MAX_PER_BATCH = 3;
// Composer navigation keys as xterm.js encodes user keystrokes: plain and
// modified arrows (CSI A-D, CSI 1;mA-D, SS3 A-D), Home/End (CSI H/F, SS3
// H/F, CSI 1~/4~), Insert/Delete/PgUp/PgDn (CSI 2~/3~/5~/6~, optional
// modifier). Deliberately EXCLUDES terminal query responses that also
// arrive via onData (DA `\x1b[?1;2c`, CPR `\x1b[12;34R`) and function
// keys, so only genuine cursor/editing keys trigger the local-echo flush.
// eslint-disable-next-line no-control-regex
const COMPOSER_NAV_KEY_PATTERN = /^\x1b(?:\[(?:[ABCDHF]|1;[2-8][ABCDHF]|[1-8](?:;[2-8])?~)|O[ABCDHF])$/;
// Prefix xterm.js puts on terminal.paste() payloads while the application
// has bracketed-paste mode (DECSET 2004) enabled. Codex, Claude Code and
// tmux all enable it, so browser pastes arrive as one onData chunk of
// `\x1b[200~<text>\x1b[201~`.
const BRACKETED_PASTE_START = '\x1b[200~';
function isComposerNavKey(data) {
return COMPOSER_NAV_KEY_PATTERN.test(data);
}
function isTerminalQueryResponse(data) {
return TERMINAL_QUERY_RESPONSE_PATTERN.test(data) || TERMINAL_OSC_RESPONSE_PATTERN.test(data);
@@ -60,8 +98,15 @@
global.CodemanTerminalInput = {
isTerminalQueryResponse,
shouldSuppressTerminalQueryResponse,
isComposerNavKey,
BRACKETED_PASTE_START,
USER_SCROLL_STICKY_SUPPRESS_MS,
TOUCH_COMPAT_MOUSE_SUPPRESS_MS,
REPLAY_ESCAPE_RE,
KEY_PAGE_UP,
KEY_PAGE_DOWN,
PAGE_KEY_SCREEN_FRACTION,
PAGE_KEY_MAX_PER_BATCH,
};
global.CODEMAN_XTERM_THEMES = CODEMAN_XTERM_THEMES;
global.codemanCurrentXtermTheme = currentXtermTheme;
@@ -158,6 +203,30 @@ Object.assign(CodemanApp.prototype, {
return false;
}
// Smart copy (#211): with a selection, Ctrl+C copies it instead of sending
// ^C. With NO selection the branch must fall through (return true, and no
// preventDefault) or the interrupt key is lost, which is the whole reason
// the selection check runs before any registry dispatch. Ctrl+Shift+C is
// the explicit copy chord and never falls through: an "explicit copy" that
// interrupts a running agent because the selection happened to be empty is
// a footgun with no upside.
// NOTE: returning false does NOT cancel the event (xterm's _keyDown calls
// this handler before its own cancel()), so preventDefault is explicit:
// without it the browser runs its native copy on top of ours.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
}
return true;
}
// Ctrl+V / Cmd+V: intercept before xterm sends ^V to PTY.
// Route through our paste trap which handles both images and text.
if ((ev.ctrlKey || ev.metaKey) && ev.key === 'v' && ev.type === 'keydown') {
@@ -391,31 +460,79 @@ Object.assign(CodemanApp.prototype, {
// ignores wheel reports); older versions DO capture wheel as option
// navigation, so they keep the local wheel.
// Shift+wheel always scrolls xterm's local scrollback (Codeman's restored
// history lives there), and once the viewport left the bottom the wheel
// stays local until the user scrolls back down — so both scrollbacks stay
// reachable without a mode switch.
// history lives there); the plain wheel stays on the CLI's transcript for
// those modes regardless of scroll position, so the CLI's input box never
// slides off the screen (see _shouldForwardWheelToApp).
//
// CAPTURE phase, deliberately, and Codeman owns the scroll. xterm's
// viewport is a vscode-style ScrollableElement that consumes wheel events
// itself (preventDefault + stopPropagation) whenever it believes a
// scrollbar exists, does NOT consult attachCustomWheelEventHandler, and —
// measured on the live instance — goes DEAF after terminal.reset(): a tab
// switch or full-history replay leaves its scroll dimensions stale, after
// which wheel events neither scroll nor propagate reliably. A bubble-phase
// listener here therefore never fired once local scrollback existed
// (measured: _shouldForwardWheelToApp call count stayed 0 while xterm
// scrolled), and after a tab switch NOTHING scrolled at all — the "input
// box scrolls up then it fights", "works at first, breaks after a tab
// switch" reports on #205.
//
// So: capture runs ancestors-first; this handler sees every wheel first
// and stops propagation, keeping xterm's scroller out of it entirely.
// Local scrolling goes through terminal.scrollLines() — buffer-level, so
// it keeps working after resets — with our own deltaMode normalization
// (_wheelScrollLines) covering Firefox's line-unit wheels. Two cases still
// belong to xterm and are passed through untouched:
// - mouseTrackingMode active: xterm's own encoder forwards the wheel to
// the PTY (htop/vim with mouse on in a shell pane);
// - alternate buffer (direct-PTY fallback running vim/less): xterm's
// alt-scroll handling converts the wheel to cursor keys, which is what
// those apps expect.
container.addEventListener(
'wheel',
(ev) => {
const trackingMode = this.terminal?.modes?.mouseTrackingMode;
if (trackingMode && trackingMode !== 'none') return;
if (this.terminal?.buffer?.active?.type === 'alternate') return;
ev.preventDefault();
const lines = this._wheelScrollLines(ev);
ev.stopPropagation();
if (this._shouldForwardWheelToApp(ev)) {
this._sendSyntheticSgrWheel(ev.clientX, ev.clientY, lines);
this._logScrollRouting('forward-sgr');
this._forwardScrollToApp(ev.clientX, ev.clientY, this._wheelScrollLines(ev));
return;
}
// Local scrolling accumulates FRACTIONAL lines: a macOS trackpad emits
// a stream of tiny pixel deltas, and rounding each one to a whole line
// (the ±1 fallback) made slow drags scroll faster than the finger.
const lines = this._wheelScrollLinesFloat(ev);
// …unless there is no local scrollback to scroll, in which case page the
// CLI's own transcript instead of doing nothing (_maybePageCliTranscript).
if (this._maybePageCliTranscript(ev, lines)) return;
this._logScrollRouting('local-scrollback');
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
this._smoothScrollBy(lines);
},
{ passive: false }
{ passive: false, capture: true }
);
// Touch scrolling — use terminal.scrollLines() for all devices.
// xterm.js DOM renderer doesn't populate xterm-viewport's scroll area,
// so native CSS scrolling (overflow-y: scroll + touch-action: pan-y)
// has nothing to scroll. Instead, convert touch deltas into scrollLines()
// calls, matching the wheel handler above.
// calls, matching the wheel handler above, including the forwarding
// branch: for the sessions whose wheel goes to the CLI's own transcript
// (_shouldForwardWheelToApp), a touch drag must go there too, or every
// phone/tablet swipe scrolls the local buffer of stale repaint frames and
// drags the CLI's pinned input box off the screen (issue #205's mobile
// half). Same gate, so Shift has no touch analog but the local-scrollback
// opt-out setting and the CLI-version gate apply to touch exactly as they
// do to the wheel — including the PageUp/PageDown fallback the wheel uses
// when that gate is false and there is no local scrollback to scroll
// (_maybePageCliTranscript), which is what keeps a swipe from being a
// complete no-op on a phone.
{
const cellHeight = () => this.terminal._core?._renderService?.dimensions?.css?.cell?.height || 13;
let touchLastX = 0;
let touchLastY = 0;
let velocity = 0;
let lastTime = 0;
@@ -429,7 +546,16 @@ Object.assign(CodemanApp.prototype, {
if (!isTouching && Math.abs(velocity) > 0.3) {
// Momentum phase — convert pixel velocity to lines
const lines = Math.round(velocity / cellHeight());
if (lines !== 0) this.terminal.scrollLines(lines);
if (lines !== 0) {
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
// Flick momentum keeps feeding the CLI's transcript from the last
// touch point; the 40ms coalescer batches the per-frame reports.
this._forwardScrollToApp(touchLastX, touchLastY, lines);
} else if (!this._maybePageCliTranscript({ shiftKey: false }, lines)) {
this.terminal.scrollLines(lines);
this._maybeLoadMoreHistoryOnScroll(lines);
}
}
velocity *= 0.92;
scrollFrame = requestAnimationFrame(scrollLoop);
} else if (!isTouching) {
@@ -450,6 +576,7 @@ Object.assign(CodemanApp.prototype, {
'touchstart',
(ev) => {
if (ev.touches.length === 1) {
touchLastX = ev.touches[0].clientX;
touchLastY = ev.touches[0].clientY;
touchStartY = touchLastY;
velocity = 0;
@@ -485,13 +612,21 @@ Object.assign(CodemanApp.prototype, {
const delta = touchLastY - touchY; // positive = scroll down
pixelAccum += delta;
velocity = delta * 1.2;
touchLastX = ev.touches[0].clientX;
touchLastY = touchY;
// Convert accumulated pixels to whole lines
const ch = cellHeight();
const lines = Math.trunc(pixelAccum / ch);
if (lines !== 0) {
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
this._logScrollRouting('forward-sgr');
this._forwardScrollToApp(touchLastX, touchLastY, lines);
} else if (!this._maybePageCliTranscript({ shiftKey: false }, lines)) {
this._logScrollRouting('local-scrollback');
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
this._maybeLoadMoreHistoryOnScroll(lines);
}
pixelAccum -= lines * ch;
}
}
@@ -752,11 +887,23 @@ Object.assign(CodemanApp.prototype, {
}
this._lastTerminalData = { data, time: performance.now() };
// ── Local Echo Pass-through ──
// After a composer nav key (arrow/Home/End/Delete) the real cursor may
// sit mid-text, where the overlay's append-only buffering would corrupt
// both the preview and the submitted text. Such sessions are handed
// back to plain PTY echo until Enter or Ctrl+C submits/cancels the
// composer line (see the nav-key branch below).
const echoPassthrough =
this._localEchoEnabled && this._echoPassthroughSessions?.has(this.activeSessionId);
if (echoPassthrough && (data === '\r' || data === '\x03')) {
this._echoPassthroughSessions.delete(this.activeSessionId);
}
// ── Local Echo Mode ──
// When enabled, keystrokes are buffered locally in the overlay for
// instant visual feedback. Nothing is sent to the PTY until Enter
// (or a control char) is pressed — avoids out-of-order char delivery.
if (this._localEchoEnabled) {
if (this._localEchoEnabled && !echoPassthrough) {
if (data === '\x7f') {
const source = this._localEchoOverlay?.removeChar();
if (source === 'flushed') {
@@ -773,9 +920,16 @@ Object.assign(CodemanApp.prototype, {
}
this._pendingInput += data;
flushInput();
} else if (source === false) {
// Nothing pending, nothing flushed, nothing detected. The
// composer may still hold text the overlay cannot see (buffer
// detection is suppressed after a control-char flush), so
// forward the backspace instead of swallowing it (issue #218);
// an empty composer ignores it.
this._pendingInput += data;
flushInput();
}
// 'pending' = removed unsent text (no PTY backspace needed)
// false = nothing to remove (swallow the backspace)
return;
}
if (/^[\r\n]+$/.test(data)) {
@@ -817,6 +971,41 @@ Object.assign(CodemanApp.prototype, {
// Single-byte ESC (user pressing Escape) still falls through to
// the control char handler below.
if (data.length > 1 && data.charCodeAt(0) === 27) {
// Bracketed paste (terminal.paste() while DECSET 2004 is on):
// flush typed-but-unsent overlay text FIRST so the pasted block
// lands after it in the composer, not before it (issue #219).
// The paste sequence gets its own delayed write: Codex's
// paste-burst handling drops keystrokes that arrive in the SAME
// PTY read as a bracketed paste (verified against codex 0.147),
// mirroring the delayed \r in the Enter branch above.
if (data.startsWith(window.CodemanTerminalInput.BRACKETED_PASTE_START)) {
const hadPending = !!this._localEchoOverlay?.pendingText;
this._flushLocalEchoPending();
if (hadPending) {
flushInput();
setTimeout(() => {
this._pendingInput += data;
flushInput();
}, 80);
} else {
this._pendingInput += data;
flushInput();
}
return;
}
// Composer nav keys (arrows, Home/End, Delete, PgUp/PgDn):
// flush unsent text so the key edits the real composer state,
// then hand the session to plain PTY echo until Enter/Ctrl+C.
// The cursor may now sit mid-text, where append-only buffering
// cannot track edits (issue #218).
if (window.CodemanTerminalInput.isComposerNavKey(data)) {
this._flushLocalEchoPending();
if (!this._echoPassthroughSessions) this._echoPassthroughSessions = new Set();
this._echoPassthroughSessions.add(this.activeSessionId);
this._pendingInput += data;
flushInput();
return;
}
// Multi-byte escape sequence — forward to PTY without clearing
// overlay/flushed state (terminal response, not user input)
this._pendingInput += data;
@@ -1178,10 +1367,24 @@ Object.assign(CodemanApp.prototype, {
},
showWelcome() {
// Phones get the session overview instead of the welcome screen: on a small
// screen "which session is blocked on me" beats "how do I start one". The
// gate lives in mobile-overview.js; every other device falls through
// unchanged. Both surfaces are toggled here so a breakpoint change (rotate,
// unfold) swaps cleanly instead of showing both.
if (this.shouldUseMobileOverview?.()) {
const overlay = document.getElementById('welcomeOverlay');
if (overlay) overlay.classList.remove('visible');
this.showMobileOverview();
this._updateCjkInputState?.();
return;
}
this.hideMobileOverview?.();
const overlay = document.getElementById('welcomeOverlay');
if (overlay) {
overlay.classList.add('visible');
this.loadTunnelStatus();
this.applyWelcomeCliVisibility();
this.loadHistorySessions();
this.initSearchPanel();
}
@@ -1191,6 +1394,7 @@ Object.assign(CodemanApp.prototype, {
},
hideWelcome() {
this.hideMobileOverview?.();
const overlay = document.getElementById('welcomeOverlay');
if (overlay) {
overlay.classList.remove('visible');
@@ -1369,7 +1573,7 @@ Object.assign(CodemanApp.prototype, {
}
titleSpan.appendChild(document.createTextNode(s.name || s.firstPrompt || shortDir));
// Badge row: mode (claude/codex/opencode/gemini/shell) + a LIVE pill.
// Badge row: mode (claude/codex/opencode/gemini/antigravity/shell) + a LIVE pill.
const badgeRow = document.createElement('div');
badgeRow.className = 'history-item-badges';
if (s.mode) {
@@ -1970,6 +2174,121 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the
* buffer while scrolling up is the user reaching for history the browser does
* not have, so pull the rest of tmux's scrollback (issue #205, see
* _maybeRefetchFullHistory). Must be called AFTER scrollLines(), since the
* check is on the resulting position, and it is deliberately not folded into
* _noteTerminalUserScroll for exactly that reason. Cheap: one integer compare
* per scroll event, and the pull itself is cooldown-guarded.
*/
_maybeLoadMoreHistoryOnScroll(lines) {
if (lines >= 0) return;
if (this.terminal?.buffer?.active?.viewportY === 0) this._maybeRefetchFullHistory?.();
},
/**
* Rows a `?full=1` capture will occupy once written into xterm.
*
* tmux joins wrapped rows in that capture (`capture-pane -J`), so a long
* logical line re-wraps into several xterm rows on write and a bare newline
* count would undershoot; escape sequences occupy no cells and come out
* first. Approximate by construction (it ignores double-width glyphs), which
* is fine: the only consumer is a coarse size comparison
* (_replayWouldShrinkBuffer), and it runs once per cooldown-guarded re-pull.
*/
_estimateReplayRows(text, cols) {
if (typeof text !== 'string' || !text) return 0;
const width = cols > 0 ? cols : 80;
const plain = text.replace(window.CodemanTerminalInput.REPLAY_ESCAPE_RE, '');
let rows = 0;
for (const line of plain.split('\n')) {
const cells = line.endsWith('\r') ? line.length - 1 : line.length;
rows += cells > width ? Math.ceil(cells / width) : 1;
}
return rows;
},
/**
* DOWNGRADE GUARD for the scroll-to-top re-pull (issue #205, round 2).
*
* `_maybeRefetchFullHistory` resets the terminal and rewrites it from the
* capture, which is a straight win when tmux holds more than the browser —
* the burst-repaint and tab-switch losses it was built for. But a repaint-mode
* CLI pane keeps NO tmux history of its own (`history_size≈0` measured for a
* Claude pane), so there the capture is roughly ONE frame while xterm may hold
* hundreds of rows of replayed frames. Rewriting then DESTROYS history
* mid-scroll: exactly the "goes back a limited amount, repeats blocks, gets
* worse when I reach the top" report from the 1.12.0 retest.
*
* So refuse when the capture is smaller, with a one-screen tolerance because
* both sides are estimates: `buffer.active.length` includes the blank rows
* below the last line, and _estimateReplayRows can only approximate wrapping.
* Only a capture that is worse by more than a full screen counts as a
* downgrade, which leaves every genuine recovery case untouched.
*/
_replayWouldShrinkBuffer(capture) {
const term = this.terminal;
const rowsNow = term?.buffer?.active?.length || 0;
if (!rowsNow) return false;
const screen = term?.rows || 24;
return this._estimateReplayRows(capture, term?.cols) + screen < rowsNow;
},
/**
* Ease-out smooth scrolling for the local wheel path. The capture-phase
* wheel handler owns local scrolling (xterm's own smooth scroller is
* bypassed, see the listener comment), so without this every notch was an
* instant multi-line jump. Wheel deltas accumulate into a pending line
* count (fractional — see _wheelScrollLinesFloat) and drain ~22% per
* animation frame with a one-line floor, so a single notch starts with a
* gentle step and glides to an exact landing; more notches mid-glide deepen
* the pending count, which reads as natural acceleration. A sub-line
* residual stays pending until further input pushes it past a whole line
* (that is what makes slow trackpad drags track the finger). Direction
* reversals cancel arithmetically. The pending amount is dropped when the
* active session changes mid-glide — leftover momentum must never scroll
* the tab the user just switched to.
*/
_smoothScrollBy(lines) {
if (!lines) return;
this._smoothScrollPending = (this._smoothScrollPending || 0) + lines;
this._smoothScrollSession = this.activeSessionId;
if (this._smoothScrollFrame) return;
const step = () => {
this._smoothScrollFrame = null;
const pending = this._smoothScrollPending || 0;
if (!pending) return;
if (this.activeSessionId !== this._smoothScrollSession) {
this._smoothScrollPending = 0;
return;
}
if (Math.abs(pending) < 1) return; // sub-line residual: wait for more input
const eased = pending * 0.22;
const move = pending > 0 ? Math.max(1, Math.floor(eased)) : Math.min(-1, Math.ceil(eased));
this._smoothScrollPending = pending - move;
this.terminal.scrollLines(move);
this._maybeLoadMoreHistoryOnScroll(move);
if (Math.abs(this._smoothScrollPending) >= 1) this._smoothScrollFrame = requestAnimationFrame(step);
};
this._smoothScrollFrame = requestAnimationFrame(step);
},
/**
* Hand a scroll gesture (wheel tick or touch drag, already converted to
* lines) to the CLI as synthetic SGR wheel reports. SGR coordinates address
* the LIVE screen (the bottom `rows` of the buffer), so a report computed
* from a scrolled-up viewport would hit-test a different row entirely, and
* forwarding while the user stares at stale scrollback looks like the
* gesture is dead. Snap back first: the gesture then always acts on what the
* CLI is drawing now.
*/
_forwardScrollToApp(clientX, clientY, lines) {
if (!this._terminalViewportAtBottom()) this.terminal.scrollToBottom();
this._sendSyntheticSgrWheel(clientX, clientY, lines);
},
_hasRecentUserScrollUp() {
if (typeof this._lastUserScrollUpAt !== 'number') return false;
return performance.now() - this._lastUserScrollUpAt < window.CodemanTerminalInput.USER_SCROLL_STICKY_SUPPRESS_MS;
@@ -2065,6 +2384,23 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* Flush the local-echo overlay's unsent text into `_pendingInput` (no
* trailing Enter) and reset overlay + flushed-state tracking. Used before
* forwarding sequences that must arrive AFTER the typed text (bracketed
* paste, composer nav keys). The caller forwards its own sequence: nav keys
* ride the same write, pastes get a delayed second write because codex
* drops keys that share a PTY read with a bracketed paste.
*/
_flushLocalEchoPending() {
const text = this._localEchoOverlay?.pendingText || '';
this._localEchoOverlay?.clear();
this._localEchoOverlay?.suppressBufferDetection();
this._flushedOffsets?.delete(this.activeSessionId);
this._flushedTexts?.delete(this.activeSessionId);
if (text) this._pendingInput += text;
},
/**
* Update local echo overlay state based on settings.
* Enabled whenever the setting is on — works during idle AND busy.
@@ -2104,8 +2440,14 @@ Object.assign(CodemanApp.prototype, {
}
},
});
} else if (session.mode === 'shell') {
} else if (session.mode === 'shell' || session.mode === 'codex') {
// Shell mode: the shell provides its own PTY echo so the overlay isn't needed.
// Codex mode: the composer is fully interactive per keystroke. Typing
// "/" pops a live-filtering command picker (issue #222), the composer
// grows and rewraps as it fills (#220), pastes are bracketed (#219)
// and arrows/history edit server-side state (#218). Buffering
// keystrokes until Enter starves all of that, so codex sessions use
// plain PTY echo like shell.
// Disable it by clearing any pending text.
this._localEchoOverlay.clear();
this._localEchoEnabled = false;
@@ -2509,7 +2851,11 @@ Object.assign(CodemanApp.prototype, {
/** Insert editable text at the active prompt without pressing Enter. */
insertTerminalText(text) {
if (!this.activeSessionId || !text) return;
if (this._localEchoEnabled && this._localEchoOverlay) {
if (
this._localEchoEnabled &&
this._localEchoOverlay &&
!this._echoPassthroughSessions?.has(this.activeSessionId)
) {
this._localEchoOverlay.appendText(text);
} else {
this.sendInput(text).catch(() => {});
@@ -2597,6 +2943,46 @@ Object.assign(CodemanApp.prototype, {
// intentionally empty
},
// Registry-aware gate for the smart-copy chord (#211). Mirrors
// shouldOpenCommandPaletteFromShortcut(): honors a rebound or disabled
// 'copy-selection' entry, and falls back to the default chord when the
// registry isn't available (isolated test harnesses).
// Returning true only means "this chord asked to copy", the CALLER decides
// what happens when there is no selection, so the interrupt stays intact.
shouldCopyTerminalSelectionFromShortcut(ev) {
// The custom key handler also runs for keypress/keyup; only keydown decides.
if (!ev || ev.type !== 'keydown') return false;
// Hot path: every dispatchable chord needs Ctrl/Cmd/Alt, so plain typing
// exits before any registry work.
if (!ev.ctrlKey && !ev.metaKey && !ev.altKey) return false;
const registryAvailable =
typeof this.getShortcutRegistry === 'function' && typeof this.matchesShortcutEvent === 'function';
const entry = registryAvailable ? this.getShortcutRegistry().find((s) => s.id === 'copy-selection') : null;
if (entry) return !entry.disabled && this.matchesShortcutEvent(ev, entry);
return !ev.altKey && (ev.key || '').toLowerCase() === 'c';
},
// Copy the current terminal selection. Goes through _copyText (Clipboard API,
// then a hidden-textarea + execCommand fallback) because install.sh's LAN
// option serves plain HTTP, where navigator.clipboard is undefined.
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const ok = await this._copyText(selection);
if (ok) {
// Clearing is what makes a second Ctrl+C an interrupt (and xterm already
// drops the selection on any keypress, so this matches existing feel).
this.terminal.clearSelection?.();
this.showToast('Copied to clipboard', 'success');
} else {
this.showToast('Failed to copy', 'error');
}
// The execCommand fallback focuses a temp textarea, so hand focus back. This
// is the CJK-aware focus router, not xterm's raw focus().
this.terminal.focus();
return ok;
},
async copyTerminal() {
try {
const buffer = this.terminal.buffer.active;
@@ -2736,9 +3122,29 @@ Object.assign(CodemanApp.prototype, {
// deltaY≈0 collapses to a fixed ±1 line/tick and the gesture can't page through
// history on a trackpad (issue #154). Non-Shift and mouse-wheel paths are
// unchanged (they carry deltaY). The `|| ±1` keeps sub-25px deltas moving.
//
// `deltaMode` says what UNIT the delta is in, and ignoring it made every
// non-pixel browser scroll ~4x too slowly: Firefox reports DOM_DELTA_LINE (1)
// with deltaY≈3 per notch, so the pixel math rounded to 0 and fell through to
// the ±1 fallback — one line per notch, versus 4-5 for Chrome's ~110px. In
// Claude mode the same value also capped the forwarded SGR report at one tick.
_wheelScrollLines(ev) {
const lines = this._wheelScrollLinesFloat(ev);
if (!lines) return 0; // pure horizontal swipe: don't fall through to -1
return Math.round(lines) || (lines > 0 ? 1 : -1);
},
/** Unrounded variant for the smooth local-scroll path, which accumulates
* sub-line fractions across events instead of forcing every tiny trackpad
* delta to a whole ±1 line. Same unit handling and Shift-axis trap. */
_wheelScrollLinesFloat(ev) {
const delta = ev.shiftKey && Math.abs(ev.deltaX) > Math.abs(ev.deltaY) ? ev.deltaX : ev.deltaY;
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
if (!delta) return 0;
return ev.deltaMode === 1 // DOM_DELTA_LINE (Firefox mouse wheel)
? delta
: ev.deltaMode === 2 // DOM_DELTA_PAGE
? delta * (this.terminal?.rows || 24)
: delta / 25; // DOM_DELTA_PIXEL (Chrome/WebKit, and every trackpad)
},
_shouldForwardWheelToApp(ev) {
@@ -2747,6 +3153,16 @@ Object.assign(CodemanApp.prototype, {
// plain wheel to xterm's own scrollback like pre-#144, for users who prefer
// it over forwarding the wheel to the CLI's transcript (issue #154). Cheap —
// loadAppSettingsFromStorage() is cache-backed.
//
// FOOTGUN, and why it is handled downstream rather than here: for a
// repaint-mode CLI that local scrollback is EMPTY (tmux keeps no history for
// the pane), so this setting can silently convert a working wheel into a
// dead one — a plausible reading of the #205 retest, where a user whose
// scrolling was broken on 1.11.x may well have flipped it while hunting for
// a fix. Scoping the setting away from those modes would be the other
// option, but it would override an explicit user choice; instead the caller
// falls through to _maybePageCliTranscript, so the gesture still pages the
// CLI's transcript and the setting keeps meaning exactly what it says.
if (this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback) return false;
const mode = this.terminal?.modes?.mouseTrackingMode;
if (mode && mode !== 'none') return false;
@@ -2757,7 +3173,23 @@ Object.assign(CodemanApp.prototype, {
} else if (sessionMode !== 'codex') {
return false;
}
return this._terminalViewportAtBottom();
// Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so
// that leaving the bottom handed the wheel back to local scrollback and both
// histories stayed reachable without a mode switch. In practice that inverted
// the behavior users actually want: a repaint-mode CLI keeps NO terminal
// scrollback of its own (tmux reports history_size=0 for a Claude pane), so
// xterm's buffer holds only Codeman's REPLAYED repaint frames. Scrolling that
// locally drags the CLI's own pinned furniture (the prompt box, the status
// line) up the screen and shows stale frames underneath, which reads as "the
// window scrolled away" rather than "I am reading history".
//
// And it was easy to fall into: scrollToLastNonEmptyLine() parks the viewport
// `rows - 2` above the last non-empty row, so any tab switch onto a session
// with trailing blank rows left the viewport off-bottom and every later wheel
// went local. Forwarding unconditionally keeps the CLI's transcript as the
// plain wheel's target and its input box fixed in place; local scrollback is
// still on Shift+wheel and on the "Wheel scrolls local history" opt-out above.
return true;
},
// Encode wheel ticks as SGR reports (button 64 = up, 65 = down) at the pointer
@@ -2774,13 +3206,105 @@ Object.assign(CodemanApp.prototype, {
if (!pos) return;
const btn = lines < 0 ? 64 : 65;
const ticks = Math.min(Math.abs(lines), 5);
this._queueScrollBytes(`\x1b[<${btn};${pos.col};${pos.row}M`.repeat(ticks));
},
/**
* Shared 40ms coalescer for every byte a scroll gesture sends to the PTY (SGR
* wheel reports and the PageUp/PageDown fallback alike). Each flush becomes a
* tmux send-keys server-side, so per-event writes would spawn a process storm
* on a single flick; the queue is bounded so a wild scroll can't build a
* backlog that keeps scrolling after the finger stops.
*/
_queueScrollBytes(data) {
if (!data || !this.activeSessionId) return;
const queued = this._wheelSgrQueue || '';
if (queued.length > 512) return;
this._wheelSgrQueue = queued + `\x1b[<${btn};${pos.col};${pos.row}M`.repeat(ticks);
this._wheelSgrQueue = queued + data;
if (this._wheelSgrFlushTimer) return;
this._wheelSgrFlushTimer = setTimeout(() => this._flushWheelSgrQueue(), 40);
},
/**
* True when this session's LOCAL scrollback is structurally empty: a Claude
* pane in repaint mode, where tmux reports `history_size≈0` and every frame
* overwrites the last, so xterm's normal buffer never grows past one screen
* (`baseY === 0`). Scrolling that buffer is a no-op no matter how the gesture
* is routed — the "wheel does nothing at all" half of the #205 retest.
*/
_localScrollbackIsHollow() {
const mode = this.sessions?.get(this.activeSessionId)?.mode || 'claude';
if (mode !== 'claude') return false;
const buf = this.terminal?.buffer?.active;
if (!buf || buf.type === 'alternate') return false;
return (buf.baseY || 0) === 0;
},
/**
* LAST-RESORT scroll for a hollow local buffer: translate gesture lines into
* coalesced PageUp/PageDown key sends so the CLI pages its OWN transcript.
*
* The rescue path for every way `_shouldForwardWheelToApp` can come back false
* on a Claude session that has no local history to fall back on: the CLI
* version probe failed or is genuinely older than 2.1.187, or the user turned
* on "Wheel scrolls local history" (which pins the wheel to a buffer that,
* for a repaint-mode CLI, is empty — the setting's footgun). Before this, all
* of those produced a completely dead gesture; the #205 reporter proved the
* keyboard route works by paging back through intact text with Fn+Up.
*
* Triple-guarded (claude mode + gate false + `baseY === 0`), so a session with
* real local scrollback is never touched. Shift is excluded on purpose: it is
* the explicit "give me local scrollback" gesture and must keep that meaning.
*
* @returns true when the gesture was consumed here (the caller must not also
* scroll locally).
*/
_maybePageCliTranscript(ev, lines) {
if (!lines || ev?.shiftKey || !this.activeSessionId) return false;
if (!this._localScrollbackIsHollow()) return false;
// Leftover travel belongs to the tab it was made on.
if (this._pageKeySession !== this.activeSessionId) {
this._pageKeySession = this.activeSessionId;
this._pageKeyPending = 0;
}
const tuning = window.CodemanTerminalInput;
const perPage = Math.max(2, Math.round((this.terminal?.rows || 24) * tuning.PAGE_KEY_SCREEN_FRACTION));
const pending = (this._pageKeyPending || 0) + lines;
const pages = Math.trunc(pending / perPage);
this._pageKeyPending = pending - pages * perPage;
if (pages) {
const key = pages < 0 ? tuning.KEY_PAGE_UP : tuning.KEY_PAGE_DOWN;
this._queueScrollBytes(key.repeat(Math.min(Math.abs(pages), tuning.PAGE_KEY_MAX_PER_BATCH)));
}
this._logScrollRouting('page-keys');
return true;
},
/**
* One line in the console saying WHY a scroll gesture went where it went.
*
* Issue #205 ran two rounds of remote guesswork — is the CLI version probe
* empty, is the opt-out setting on, did a mouse DECSET leak past the strip? —
* that this single log answers directly. Logged once per session per distinct
* decision, so a steady gesture stays silent and a CHANGE (e.g. the version
* arriving late and flipping the route) still prints.
*/
_logScrollRouting(decision) {
const sessionId = this.activeSessionId || '(none)';
const session = this.sessions?.get(sessionId);
const optOut = !!this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback;
const tracking = this.terminal?.modes?.mouseTrackingMode || 'none';
const baseY = this.terminal?.buffer?.active?.baseY ?? -1;
const signature = `${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${baseY > 0}`;
if (!this._scrollRoutingLogged) this._scrollRoutingLogged = new Map();
if (this._scrollRoutingLogged.get(sessionId) === signature) return;
this._scrollRoutingLogged.set(sessionId, signature);
console.log(
`[scroll] ${sessionId} → ${decision} (mode=${session?.mode || '?'}, cliVersion=${session?.cliVersion || 'unknown'}, ` +
`localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, localScrollbackRows=${baseY})`
);
},
_flushWheelSgrQueue() {
this._wheelSgrFlushTimer = null;
const data = this._wheelSgrQueue;
+4
View File
@@ -210,6 +210,9 @@ Object.assign(CodemanApp.prototype, {
document.body.appendChild(win);
// Drop the spawn class on the next frame so the transition runs.
requestAnimationFrame(() => win.classList.remove('spawning'));
// Then hand over to the chosen entrance style (no-op for `fly`/`off`, which
// leave the small scale-in transition above as the whole animation).
this.applyWindowEntrance?.(win);
const header = win.querySelector('.ultracode-window-header');
const dragListeners = this.makeWindowDraggable(win, header);
@@ -586,6 +589,7 @@ Object.assign(CodemanApp.prototype, {
document.body.appendChild(win);
requestAnimationFrame(() => win.classList.remove('spawning'));
this.applyWindowEntrance?.(win);
const header = win.querySelector('.ultracode-window-header');
const dragListeners = this.makeWindowDraggable(win, header);
+1 -2
View File
@@ -97,8 +97,7 @@ Object.assign(CodemanApp.prototype, {
<span class="tab-name">${escapeHtml(webview.name)}</span>
</span>
</span>
<span class="tab-gear" onclick="event.stopPropagation(); app.showWebviewModal(${jsonId})" title="URL settings" aria-label="URL settings" tabindex="0">&#x2699;</span>
<span class="tab-close" onclick="event.stopPropagation(); app.closeWebviewTab(${jsonId})" title="Close tab" aria-label="Close web tab" tabindex="0">&times;</span>
<span class="tab-actions"><span class="tab-gear" onclick="event.stopPropagation(); app.showWebviewModal(${jsonId})" title="URL settings" aria-label="URL settings" tabindex="0">&#x2699;</span><span class="tab-close" onclick="event.stopPropagation(); app.closeWebviewTab(${jsonId})" title="Close tab" aria-label="Close web tab" tabindex="0">&times;</span></span>
</div>`);
idx++;
}
+283 -5
View File
@@ -1,12 +1,16 @@
/**
* @fileoverview File browser and streaming routes.
* Provides directory listing, file content preview, raw file serving, and tail streaming.
* Provides directory listing, file content preview, raw file serving, tail
* streaming, and the File Viewer edit-mode write path (edit=1 read +
* PUT /api/sessions/:id/file-content; policy in src/config/file-editing.ts,
* design in docs/file-viewer-edit-plan.md).
*/
import { FastifyInstance, type FastifyReply } from 'fastify';
import { basename as pathBasename, extname, isAbsolute, join, relative, resolve, sep } from 'node:path';
import { basename as pathBasename, dirname, extname, isAbsolute, join, relative, resolve, sep } from 'node:path';
import { createReadStream, realpathSync, type ReadStream } from 'node:fs';
import fs from 'node:fs/promises';
import { createHash, randomBytes } from 'node:crypto';
import { homedir } from 'node:os';
import type {
ApiResponse,
@@ -14,6 +18,7 @@ import type {
FilesystemBrowseEntry,
FilesystemBrowseRoot,
FilesystemPreviewKind,
FileWriteData,
} from '../../types.js';
import { ApiErrorCode, createErrorResponse, getErrorMessage } from '../../types.js';
import { fileStreamManager } from '../../file-stream-manager.js';
@@ -44,7 +49,14 @@ import type { SessionAttachmentHistoryItem, SessionState } from '../../types/ses
import { isSensitivePath } from '../sensitive-path.js';
import { SseEvent } from '../sse-events.js';
import type { ConfigPort, EventPort, SessionPort } from '../ports/index.js';
import { FilesystemBrowseQuerySchema, FilesystemPreviewQuerySchema } from '../schemas.js';
import { FilesystemBrowseQuerySchema, FilesystemPreviewQuerySchema, FileWriteSchema } from '../schemas.js';
import {
MAX_EDITABLE_BYTES,
applyEol,
detectEol,
isDeniedEditRelativePath,
isEditableFileName,
} from '../../config/file-editing.js';
const MIME_TYPES: Record<string, string> = {
png: 'image/png',
@@ -453,6 +465,71 @@ function appendDownloadFlag(url: string): string {
return `${url}${url.includes('?') ? '&' : '?'}download=true`;
}
// ===== File Viewer edit mode (issue #212) =====
// Policy lives in src/config/file-editing.ts; design in docs/file-viewer-edit-plan.md.
function sha256Hex(buf: Buffer): string {
return createHash('sha256').update(buf).digest('hex');
}
/** NUL byte in the first 8KB — same binary signal the plain read path uses. */
function sniffsBinary(buf: Buffer): boolean {
const sniffLength = Math.min(buf.length, 8192);
for (let i = 0; i < sniffLength; i++) {
if (buf[i] === 0) return true;
}
return false;
}
/**
* Structured-throw variant for the edit read/write paths. Identical mechanics to
* throwFilesystemPickerError (rendered by the central route error handler both
* in prod and in the app.inject() test harness); a separate name only so edit
* failures grep distinctly.
*/
function throwFileEditError(statusCode: number, code: ApiErrorCode, message: string): never {
throw Object.assign(new Error(message), {
statusCode,
body: createErrorResponse(code, message),
});
}
/**
* Gate a resolved workspace file for edit-mode read/write. Throws a structured
* error when the file may not be edited; returns void when it may. Order
* matters for the message a user sees: confinement (the caller's 404) →
* sensitive/blocked (403) → .git (403) → extension allowlist (400).
*/
function assertEditableTarget(resolvedPath: string, relativePath: string, blockedTrees: readonly string[]): void {
if (isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, blockedTrees)) {
throwFileEditError(403, ApiErrorCode.FORBIDDEN, 'Editing this file is blocked');
}
if (isDeniedEditRelativePath(relativePath)) {
throwFileEditError(403, ApiErrorCode.FORBIDDEN, 'Files under .git cannot be edited');
}
if (!isEditableFileName(pathBasename(resolvedPath))) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'This file type is not editable');
}
}
/**
* Decode a candidate edit buffer, refusing binary and non-UTF-8 content. The
* round-trip compare is what protects against silent corruption: decoding
* latin-1 (or any non-UTF-8) bytes yields U+FFFD replacements, and writing
* those back would destroy the original bytes. A UTF-8 BOM round-trips and is
* deliberately preserved.
*/
function decodeEditableText(buf: Buffer): string {
if (sniffsBinary(buf)) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Binary files cannot be edited');
}
const text = buf.toString('utf8');
if (!Buffer.from(text, 'utf8').equals(buf)) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Only UTF-8 text files can be edited');
}
return text;
}
function getSessionAttachmentHistory(
ctx: SessionPort & ConfigPort,
sessionId: string,
@@ -563,6 +640,25 @@ async function buildExternalAttachmentRouteItem(
}
}
/**
* Headers Fastify already put on the reply, in a shape `writeHead` accepts.
*
* `reply.raw.writeHead()` writes straight to the Node response and bypasses
* Fastify's header store, so anything the security `onRequest` hook granted — CORS
* for localhost origins, nosniff, frame-options, CSP — is silently dropped on every
* route that answers this way. Spread this first and let the route's own headers
* win over it.
*/
function inheritedHeaders(reply: {
getHeaders(): NodeJS.Dict<number | string | string[]>;
}): Record<string, number | string | string[]> {
const out: Record<string, number | string | string[]> = {};
for (const [name, value] of Object.entries(reply.getHeaders())) {
if (value !== undefined) out[name] = value;
}
return out;
}
export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & EventPort & ConfigPort): void {
// Lazy filesystem listing for the Link Existing and mobile input path pickers.
app.get('/api/filesystem/browse', async (req, reply): Promise<ApiResponse<FilesystemBrowseData>> => {
@@ -864,7 +960,12 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
// Get file content for preview (File Browser)
app.get('/api/sessions/:id/file-content', async (req) => {
const { id } = req.params as { id: string };
const { path: filePath, lines, raw } = req.query as { path?: string; lines?: string; raw?: string };
const {
path: filePath,
lines,
raw,
edit,
} = req.query as { path?: string; lines?: string; raw?: string; edit?: string };
const session = findSessionOrFail(ctx, id, req);
if (!filePath) {
@@ -876,7 +977,52 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
if (!validated) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'File not found');
}
const { resolvedPath } = validated;
const { resolvedPath, relativePath } = validated;
// Read-for-edit: never truncated (a truncated buffer must never become an
// edit buffer), tighter size cap, full editability gate, and the hash/eol
// the client must echo back on PUT. Outside the shared try/catch below so
// its structured errors keep their status codes instead of collapsing into
// OPERATION_FAILED.
if (edit === '1' || edit === 'true') {
const guard = await loadAttachmentGuardConfig();
assertEditableTarget(resolvedPath, relativePath, guard.blockedTrees);
let editStat;
try {
editStat = await fs.stat(resolvedPath);
} catch {
throwFileEditError(404, ApiErrorCode.NOT_FOUND, 'File not found');
}
if (!editStat.isFile()) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Only regular files can be edited');
}
if (editStat.size > MAX_EDITABLE_BYTES) {
throwFileEditError(
413,
ApiErrorCode.INVALID_INPUT,
`File too large to edit here (${Math.ceil(editStat.size / 1024)}KB > ${MAX_EDITABLE_BYTES / 1024}KB limit)`
);
}
const editBuf = await fs.readFile(resolvedPath);
const editText = decodeEditableText(editBuf);
return {
success: true,
data: {
path: filePath,
content: editText,
size: editBuf.length,
mtimeMs: editStat.mtimeMs,
totalLines: editText.split('\n').length,
truncated: false,
extension: filePath.split('.').pop()?.toLowerCase() || '',
editable: true,
hash: sha256Hex(editBuf),
eol: detectEol(editText),
},
};
}
try {
const stat = await fs.stat(resolvedPath);
@@ -998,6 +1144,19 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const truncatedContent = allLines.length > maxLines;
const displayContent = truncatedContent ? allLines.slice(0, maxLines).join('\n') : content;
// Additive edit-mode advertisement: whether an edit=1 re-fetch would
// succeed. The UTF-8 round-trip compare is a cheap memcmp and mirrors
// decodeEditableText; no hash here — the Edit action re-fetches with
// edit=1, which is where the baseHash comes from.
const guard = await loadAttachmentGuardConfig();
const editable =
isEditableFileName(pathBasename(resolvedPath)) &&
!isDeniedEditRelativePath(relativePath) &&
!isSensitivePath(resolvedPath) &&
!isBlockedAttachmentPath(resolvedPath, guard.blockedTrees) &&
stat.size <= MAX_EDITABLE_BYTES &&
Buffer.from(content, 'utf8').equals(buf);
return {
success: true,
data: {
@@ -1007,6 +1166,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
totalLines: allLines.length,
truncated: truncatedContent,
extension: ext,
editable,
},
};
} catch (err) {
@@ -1014,6 +1174,121 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
}
});
// File Viewer edit mode: save a text file back into the session workspace.
// Edit-in-place ONLY — there is deliberately no O_CREAT path in this handler,
// so it can never create, and it never deletes. Confinement is identical to
// the read path (realpath + workspace boundary + ownership via
// findSessionOrFail), plus the sensitive-path/attachment-guard blocklists and
// the extension allowlist. Concurrency is optimistic: the client echoes the
// sha256 it loaded (baseHash) and a mismatch is a 409 unless force is set.
// bodyLimit: JSON escaping can expand content up to ~6x (each control char
// becomes \uXXXX), so the 512KB content cap needs headroom over Fastify's
// 1MB default.
app.put(
'/api/sessions/:id/file-content',
{ bodyLimit: 4 * 1024 * 1024 },
async (req): Promise<ApiResponse<FileWriteData>> => {
const { id } = req.params as { id: string };
const session = findSessionOrFail(ctx, id, req);
const body = parseBody(FileWriteSchema, req.body);
// Exact byte cap — the schema's .max() counts UTF-16 code units and is
// only a coarse pre-filter.
if (Buffer.byteLength(body.content, 'utf8') > MAX_EDITABLE_BYTES) {
throwFileEditError(413, ApiErrorCode.INVALID_INPUT, `Content too large (${MAX_EDITABLE_BYTES / 1024}KB limit)`);
}
const validated = validateSessionFilePath(session.workingDir, body.path);
if (!validated) {
// Covers missing files, traversal, and symlink escapes alike — a write
// target that fails confinement is reported identically to a missing
// one, matching the read route.
throwFileEditError(404, ApiErrorCode.NOT_FOUND, 'File not found');
}
const { resolvedPath, relativePath } = validated;
const guard = await loadAttachmentGuardConfig();
assertEditableTarget(resolvedPath, relativePath, guard.blockedTrees);
let stat;
try {
stat = await fs.stat(resolvedPath);
} catch {
throwFileEditError(404, ApiErrorCode.NOT_FOUND, 'File not found');
}
if (!stat.isFile()) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Only regular files can be edited');
}
if (stat.size > MAX_EDITABLE_BYTES) {
throwFileEditError(
413,
ApiErrorCode.INVALID_INPUT,
`File too large to edit here (${MAX_EDITABLE_BYTES / 1024}KB limit)`
);
}
const currentBuf = await fs.readFile(resolvedPath);
const currentText = decodeEditableText(currentBuf);
const currentHash = sha256Hex(currentBuf);
if (currentHash !== body.baseHash && !body.force) {
throwFileEditError(
409,
ApiErrorCode.CONFLICT,
'File changed on disk since it was loaded — reload it or overwrite'
);
}
// Re-apply the file's original line endings (a <textarea> normalizes to
// LF; without this a two-line edit of a CRLF file rewrites every line).
const eol = body.eol ?? detectEol(currentText);
const outText = applyEol(body.content, eol);
const outBuf = Buffer.from(outText, 'utf8');
if (outBuf.length > MAX_EDITABLE_BYTES) {
throwFileEditError(413, ApiErrorCode.INVALID_INPUT, `Content too large (${MAX_EDITABLE_BYTES / 1024}KB limit)`);
}
// Atomic replace: O_EXCL temp in the same directory, then rename.
// 'wx' cannot follow a pre-existing symlink and rename() replaces (not
// follows) a symlink in the final component, which closes the
// validate-then-write TOCTOU window. fchmod because open()'s mode is
// masked by the process umask; fsync so the rename never publishes a
// partially-durable file. Trade-off (same as vim's default): the inode
// changes, so hardlinks keep the old content.
const fileMode = stat.mode & 0o777;
const tmpPath = join(
dirname(resolvedPath),
`.${pathBasename(resolvedPath)}.codeman-tmp-${randomBytes(6).toString('hex')}`
);
let handle;
try {
handle = await fs.open(tmpPath, 'wx', fileMode);
await handle.chmod(fileMode);
await handle.writeFile(outBuf);
await handle.sync();
await handle.close();
handle = undefined;
await fs.rename(tmpPath, resolvedPath);
} catch (err) {
if (handle) await handle.close().catch(() => {});
await fs.unlink(tmpPath).catch(() => {});
throwFileEditError(500, ApiErrorCode.OPERATION_FAILED, `Failed to save file: ${getErrorMessage(err)}`);
}
const newStat = await fs.stat(resolvedPath).catch(() => undefined);
return {
success: true,
data: {
path: body.path,
size: outBuf.length,
mtimeMs: newStat?.mtimeMs ?? Date.now(),
hash: sha256Hex(outBuf),
totalLines: outText.split('\n').length,
eol,
},
};
}
);
// Serve raw file content (for images/binary files)
app.get('/api/sessions/:id/file-raw', async (req, reply) => {
const { id } = req.params as { id: string };
@@ -1081,6 +1356,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const basename = rawBasename.replace(/["\\\r\n]/g, '_');
if (download === 'true' || ext === 'svg') {
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': ext === 'svg' ? 'application/octet-stream' : mimeTypes[ext] || 'application/octet-stream',
'Content-Disposition': `attachment; filename="${basename}"`,
'Content-Length': content.length,
@@ -1320,6 +1596,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
// Set up SSE headers
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
@@ -1453,6 +1730,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const content = await fs.readFile(resolvedPath);
// Bypass Fastify compression — write directly to raw response
reply.raw.writeHead(200, {
...inheritedHeaders(reply),
'Content-Type': mimeTypes[ext] || 'application/octet-stream',
'Content-Disposition': `attachment; filename="${filename}"`,
'Content-Length': content.length,
+22
View File
@@ -10,6 +10,7 @@ import { HookEventSchema, isValidWorkingDir } from '../schemas.js';
import { sanitizeHookData, parseBody } from '../route-helpers.js';
import { persistDockerCaseClaudeSessionId } from '../../docker-hosts.js';
import { getDataDir } from '../../config/instance.js';
import { sessionWaits, hooksAvailableForMode } from '../session-wait-registry.js';
import type { SessionPort, EventPort, RespawnPort, ConfigPort, InfraPort } from '../ports/index.js';
export function registerHookEventRoutes(
@@ -22,6 +23,27 @@ export function registerHookEventRoutes(
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Session not found');
}
// Wake anything blocked on `GET /api/sessions/:id/wait`. Hooks are the only
// DEFINITIVE signals Codeman gets (`idle` is inferred from output stabilization
// and can flap mid-turn), so these two are what an orchestrating agent should
// wait on.
//
// Gated on the session's MODE, matching `resolveWaitSignals` on the read side.
// Without it the guard is one-sided: a caller cannot ASK for `stop` on a shell or
// codex session, but this endpoint would happily deliver one for it. Hook events
// carry no identity beyond a per-instance secret shared by every case, so this is
// also the cheap half of the forgery surface — a `stop` claimed for a session that
// could never legitimately emit one is now dropped instead of steering another
// agent's control flow.
const waitSession = ctx.sessions.get(sessionId);
if (waitSession && hooksAvailableForMode(waitSession.mode)) {
if (event === 'stop') {
sessionWaits.notifySignal(sessionId, 'stop');
} else if (event === 'permission_prompt' || event === 'elicitation_dialog') {
sessionWaits.notifySignal(sessionId, 'blocked');
}
}
// Signal the respawn controller based on hook event type
const controller = ctx.respawnControllers.get(sessionId);
if (controller) {
File diff suppressed because it is too large Load Diff
+21 -1
View File
@@ -374,9 +374,19 @@ export function registerSystemRoutes(
});
// ═══════════════════════════════════════════════════════════════
// CLI Integrations (OpenCode, Codex, Gemini)
// CLI Integrations (Claude, OpenCode, Codex, Gemini, Antigravity)
// ═══════════════════════════════════════════════════════════════
// ========== Claude ==========
app.get('/api/claude/status', async () => {
const { isClaudeAvailable, findClaudeDir } = await import('../../utils/claude-cli-resolver.js');
return {
available: isClaudeAvailable(),
path: findClaudeDir(),
};
});
// ========== OpenCode ==========
app.get('/api/opencode/status', async () => {
@@ -405,6 +415,16 @@ export function registerSystemRoutes(
};
});
// ========== Antigravity ==========
app.get('/api/antigravity/status', async () => {
const { isAntigravityAvailable, resolveAntigravityDir } = await import('../../utils/antigravity-cli-resolver.js');
return {
available: isAntigravityAvailable(),
path: resolveAntigravityDir(),
};
});
// ═══════════════════════════════════════════════════════════════
// State & Lifecycle (cleanup, lifecycle log, stats)
// ═══════════════════════════════════════════════════════════════
+8 -2
View File
@@ -180,13 +180,19 @@ export function registerWsRoutes(app: FastifyInstance, ctx: SessionPort, getHost
const cid = typeof msg.cid === 'string' ? msg.cid : null;
const seq = Number.isInteger(msg.seq) ? (msg.seq as number) : null;
const apply = cid && seq !== null ? session.shouldApplyInput(cid, seq) : true;
let delivered = true;
if (apply) {
// Typed input from a claim-holding desktop keeps the claim "hot"
// and re-asserts the desktop layout after a mobile override.
if (holdsDesktopClaim) session.noteDesktopActivity();
session.write(msg.d);
delivered = session.write(msg.d);
// A session whose PTY is gone swallows the write. ACKing anyway told
// the client to drop the frame from its durable queue and left the seq
// burnt, so the retry that reliable delivery exists for was rejected as
// a duplicate: the input was lost for good.
if (!delivered && cid && seq !== null) session.forgetInputSeq(cid, seq);
}
if (seq !== null && socket.readyState === 1) {
if (delivered && seq !== null && socket.readyState === 1) {
socket.send(`{"t":"ia","seq":${seq}}`);
}
} else if (
+114 -6
View File
@@ -16,6 +16,8 @@ import {
MIN_TERMINAL_BUFFER_BYTES,
MIN_TERMINAL_SCROLLBACK_LINES,
} from '../config/terminal-history.js';
import { MAX_EDITABLE_BYTES } from '../config/file-editing.js';
import { MIN_MATCH_LENGTH, MAX_MATCH_LENGTH } from '../config/agent-wait.js';
// ========== Path Validation ==========
@@ -83,10 +85,34 @@ export const FilesystemPreviewQuerySchema = z.object({
.optional(),
});
/**
* Body validation for `PUT /api/sessions/:id/file-content` (File Viewer edit
* mode). `content.max()` counts UTF-16 code units, which for UTF-8 output is
* always <= the byte length, so it is a coarse pre-filter that never rejects
* valid content; the handler enforces the exact MAX_EDITABLE_BYTES byte cap.
* Workspace containment and symlink resolution are enforced by the route via
* validateSessionFilePath after parsing.
*/
export const FileWriteSchema = z
.object({
path: z
.string()
.min(1)
.max(4096)
.refine((p) => !p.includes('\0') && !p.includes('\n') && !p.includes('\r'), {
message: 'Invalid path',
}),
content: z.string().max(MAX_EDITABLE_BYTES),
baseHash: z.string().regex(/^[a-f0-9]{64}$/, 'baseHash must be a sha256 hex digest'),
eol: z.enum(['lf', 'crlf']).optional(),
force: z.boolean().optional(),
})
.strict();
// ========== Env Var Allowlist ==========
/** Allowlisted env var key prefixes */
const ALLOWED_ENV_PREFIXES = ['CLAUDE_CODE_', 'OPENCODE_', 'CODEX_', 'GEMINI_', 'GOOGLE_'];
const ALLOWED_ENV_PREFIXES = ['CLAUDE_CODE_', 'OPENCODE_', 'CODEX_', 'GEMINI_', 'GOOGLE_', 'ANTIGRAVITY_'];
/** Env var keys that are always blocked (security-sensitive) */
const BLOCKED_ENV_KEYS = new Set([
@@ -116,7 +142,7 @@ const safeEnvOverridesSchema = z
},
{
message:
'envOverrides contains blocked or disallowed env var keys. Only CLAUDE_CODE_*, OPENCODE_*, CODEX_*, GEMINI_*, and GOOGLE_* keys are allowed.',
'envOverrides contains blocked or disallowed env var keys. Only CLAUDE_CODE_*, OPENCODE_*, CODEX_*, GEMINI_*, GOOGLE_*, and ANTIGRAVITY_* keys are allowed.',
}
);
@@ -182,6 +208,7 @@ const CodexConfigSchema = z
.regex(/^[a-zA-Z0-9_-]+$/)
.optional(),
dangerouslyBypassApprovals: z.boolean().optional(),
animations: z.boolean().optional(),
renderMode: z
.enum(['scrollback', 'hybrid'])
.optional()
@@ -206,9 +233,26 @@ const GeminiConfigSchema = z
})
.optional();
/** Schema for Antigravity CLI (agy)-specific configuration */
const AntigravityConfigSchema = z
.object({
model: z
.string()
.max(100)
.regex(/^[a-zA-Z0-9._\-/]+$/)
.optional(),
dangerouslySkipPermissions: z.boolean().optional(),
resumeConversationId: z
.string()
.max(100)
.regex(/^[a-zA-Z0-9._-]+$/)
.optional(),
})
.optional();
export const CreateSessionSchema = z.object({
workingDir: safePathSchema.optional(),
mode: z.enum(['claude', 'shell', 'opencode', 'codex', 'gemini']).optional(),
mode: z.enum(['claude', 'shell', 'opencode', 'codex', 'gemini', 'antigravity']).optional(),
name: z.string().max(100).optional(),
envOverrides: safeEnvOverridesSchema,
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
@@ -220,6 +264,7 @@ export const CreateSessionSchema = z.object({
openCodeConfig: OpenCodeConfigSchema,
codexConfig: CodexConfigSchema,
geminiConfig: GeminiConfigSchema,
antigravityConfig: AntigravityConfigSchema,
/** Resume a previous Claude conversation by its session ID (used for reboot recovery) */
resumeSessionId: z
.string()
@@ -328,6 +373,7 @@ const RemoteCommandOverridesSchema = z
opencode: z.string().min(1).max(300).optional(),
codex: z.string().min(1).max(300).optional(),
gemini: z.string().min(1).max(300).optional(),
antigravity: z.string().min(1).max(300).optional(),
})
.strict()
.optional();
@@ -471,7 +517,7 @@ export const DockerHostSchema = z.object({
mountCredentials: z.boolean().optional(),
hooksEnabled: z.boolean().optional(),
resumeOnStart: z.boolean().optional(),
commands: RemoteCommandOverridesSchema, // same shell/claude/opencode/codex/gemini shape
commands: RemoteCommandOverridesSchema, // same shell/claude/opencode/codex/gemini/antigravity shape
extraCreateArgs: z
.array(
z
@@ -600,10 +646,11 @@ export const QuickStartSchema = z.object({
* a real host dir, so the settings file crosses the bind mount); rejected for
* remote cases (the file would be written on the WRONG machine). */
modelOverride: z.string().max(50).optional(),
mode: z.enum(['claude', 'shell', 'opencode', 'codex', 'gemini']).optional(),
mode: z.enum(['claude', 'shell', 'opencode', 'codex', 'gemini', 'antigravity']).optional(),
openCodeConfig: OpenCodeConfigSchema,
codexConfig: CodexConfigSchema,
geminiConfig: GeminiConfigSchema,
antigravityConfig: AntigravityConfigSchema,
envOverrides: safeEnvOverridesSchema,
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort: effortLevelSchema,
@@ -756,6 +803,7 @@ export const SettingsUpdateSchema = z
allowedTools: z.string().max(2000).optional(),
// Codex CLI settings
codexDangerouslyBypassApprovals: z.boolean().optional(),
codexAnimationsEnabled: z.boolean().optional(),
// Terminal history and retention
terminalScrollbackLines: z
.number()
@@ -864,6 +912,66 @@ export const SessionInputWithLimitSchema = z.object({
// unset rather than sending null. See docs/reliable-input-delivery.md.
seq: z.number().int().nonnegative().optional(),
clientId: z.string().max(128).optional(),
// Send-and-wait (agent orchestration): `true` for the default signal set, or the
// same grammar as `GET .../wait` — a comma string or an array of signals. Absent
// means the historical fire-and-forget behavior, byte for byte.
//
// `.nullish()`, not `.optional()`: a third-party caller building the body with
// JSON.stringify keeps an explicit null on the wire, and `.optional()` rejects it
// with INVALID_INPUT. That gotcha has shipped as a real bug twice.
wait: z.union([z.boolean(), z.string().max(120), z.array(z.string().max(120)).max(8)]).nullish(),
// Unbounded above: the effective value is clamped to MAX_WAIT_MS server-side and
// returned as `data.wait.timeoutMs`, so a caller that asks for 24h sees what it
// actually got. A `.max()` here would turn the same documented clamp into a 400 for
// large-enough guesses, which is the one behaviour an agent cannot predict.
waitTimeout: z.number().int().positive().nullish(),
});
/**
* Query validation for `GET /api/sessions/:id/wait` (agent wait primitives).
*
* Everything arrives as a string. `timeout` is coerced and bounded here, then
* clamped again to the operator's ceiling by `clampWaitMs()` — the schema bound
* only keeps an absurd number out of the arithmetic. A non-numeric `timeout` is a
* 400 rather than a silent fallback, so an agent never believes it asked for a
* longer wait than it got; the value actually applied comes back as
* `data.wait.timeoutMs`, which is what makes the clamp observable. `until` is
* parsed by `parseWaitSignals()`, which reports unknown tokens instead of
* dropping them.
*
* `until` accepts an ARRAY as well as the comma string: `?until=stop&until=exit`
* is how most HTTP clients express a list, Fastify's query parser delivers a
* repeated parameter as an array, and `parseWaitSignals()` has always handled
* both. Rejecting the repeated form left that branch unreachable and 400'd the
* more natural spelling.
*/
export const SessionWaitQuerySchema = z.object({
until: z.union([z.string().max(120), z.array(z.string().max(120)).max(8)]).optional(),
// No upper bound on purpose. The contract is "clamped to [MIN_WAIT_MS, MAX_WAIT_MS]",
// and a `.max()` here contradicted it: `timeout=99999999` was a 400 mid-fan-out while
// `timeout=600001` was silently clamped, so the same documented rule produced two
// different outcomes depending on how big the caller's guess was. `clampWaitMs()`
// bounds every finite value, and `.int()` still rejects `Infinity`/`1e999` and junk.
timeout: z.coerce.number().int().positive().optional(),
fresh: z.enum(['0', '1', 'true', 'false']).optional(),
});
/**
* Query validation for `GET /api/sessions/:id/wait-output`.
*
* `match` is a LITERAL substring, never a pattern: `search-service.ts` avoids regex
* so there is no ReDoS surface, and this endpoint is more exposed still (the pattern
* would be caller-supplied and the input is a live stream). The length bound is a
* second reason the carry buffer stays small. The route separately rejects a `regex`
* parameter outright rather than ignoring it.
*/
export const SessionWaitOutputQuerySchema = z.object({
match: z.string().min(MIN_MATCH_LENGTH).max(MAX_MATCH_LENGTH),
nocase: z.enum(['0', '1', 'true', 'false']).optional(),
from: z.enum(['now', 'buffer']).optional(),
// Unbounded above for the same reason as SessionWaitQuerySchema.timeout: clamping is
// the documented contract, so a large value must clamp rather than 400.
timeout: z.coerce.number().int().positive().optional(),
});
// ========== Session Mutation Routes ==========
@@ -957,7 +1065,7 @@ const noNewlines = (v: string) => !/[\r\n]/.test(v);
/** Shared field shape for creating/updating a scheduled job. */
const CronJobBaseSchema = z.object({
name: z.string().min(1).max(200),
agentType: z.enum(['claude', 'shell', 'opencode', 'codex', 'gemini']),
agentType: z.enum(['claude', 'shell', 'opencode', 'codex', 'gemini', 'antigravity']),
workingDir: safePathSchema,
launchCommand: z.string().max(2000).refine(noNewlines, 'launchCommand must be a single line').optional(),
promptMode: z.enum(['inline_text', 'prompt_file_path']),
+81
View File
@@ -85,6 +85,7 @@ import {
attachSessionListeners,
detachSessionListeners,
} from './session-listener-wiring.js';
import { sessionWaits } from './session-wait-registry.js';
import {
wireRespawnListeners,
setupTimedRespawn,
@@ -805,7 +806,24 @@ export class WebServer extends EventEmitter {
const clientId =
typeof query.clientId === 'string' && SSE_CLIENT_ID_RE.test(query.clientId) ? query.clientId : undefined;
// Carry over the headers the security hook already set on this reply.
//
// writeHead goes straight to the Node response and bypasses Fastify's header
// store, so everything the onRequest hook granted is silently dropped —
// including the Access-Control-Allow-Origin it emits for localhost origins.
// The result is an internal contradiction: a localhost page may call every
// /api endpoint cross-origin, but its EventSource fails CORS. The security
// headers (nosniff, frame-options, CSP) were lost the same way.
//
// The other raw-writeHead routes live in file-routes.ts and share a helper;
// this one keeps its own copy so the server does not import from a route
// module it registers.
const inherited: Record<string, number | string | string[]> = {};
for (const [name, value] of Object.entries(reply.getHeaders())) {
if (value !== undefined) inherited[name] = value;
}
reply.raw.writeHead(200, {
...inherited,
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
Connection: 'keep-alive',
@@ -1230,6 +1248,16 @@ export class WebServer extends EventEmitter {
}
}
// Release anything blocked on this session, in the documented order: 'exit'
// first so an until=exit caller gets its signal, then cancelAll so everyone
// else resolves with ended:true instead of timing out.
//
// The 'exit' here is NOT redundant with the PTY-exit listener: listeners are
// detached a few lines above, before `session.stop()`, so on a delete the
// session's own exit event never reaches the registry.
sessionWaits.notifySignal(sessionId, 'exit');
sessionWaits.cancelAll(sessionId);
this.broadcast(SseEvent.SessionDeleted, { id: sessionId });
}
@@ -1288,6 +1316,49 @@ export class WebServer extends EventEmitter {
// actual on/off. We expose `__codemanGestureAvailable` so the settings UI can
// show the toggle only when the feature is available, and inject the bundle
// (served same-origin from /gesture/, so 'self' covers it) only when enabled.
// Tool availability (#200/#201): the welcome-screen run buttons, the run-mode
// dropdown entries and the App Settings "Codex CLI" tab are all offers that a
// box without the binary cannot keep — picking one spawns a session that
// errors out immediately. One object answers all three.
//
// INJECTED, not fetched per surface. The `/api/<cli>/status` routes exist and
// stay (they mirror each other and are a fine API surface), but as the source
// for UI gating they buy nothing: every resolver memoizes its PATH probe on
// the server, so a fetch is exactly as stale as an injected value, while
// costing a round trip each time the dropdown opens and leaving the welcome
// buttons to flicker in after paint. Installing a CLI later needs a server
// restart either way. Memoized probes also make this cheap per render.
//
// Solo popups skip it: no settings modal, no welcome screen, no run menu.
if (!soloSessionId) {
const [
{ isClaudeAvailable },
{ isOpenCodeAvailable },
{ isCodexAvailable },
{ isGeminiAvailable },
{ isAntigravityAvailable },
{ isCloudflaredAvailable },
] = await Promise.all([
import('../utils/claude-cli-resolver.js'),
import('../utils/opencode-cli-resolver.js'),
import('../utils/codex-cli-resolver.js'),
import('../utils/gemini-cli-resolver.js'),
import('../utils/antigravity-cli-resolver.js'),
import('../utils/cloudflared-resolver.js'),
]);
const available = {
claude: isClaudeAvailable(),
opencode: isOpenCodeAvailable(),
codex: isCodexAvailable(),
gemini: isGeminiAvailable(),
antigravity: isAntigravityAvailable(),
cloudflared: isCloudflaredAvailable(),
};
html = html.replace(
'</head>',
`<script>window.__codemanCliAvailable=${JSON.stringify(available)};</script>\n</head>`
);
}
if (!soloSessionId && process.env.CODEMAN_GESTURE === '1') {
html = html.replace('</head>', `<script>window.__codemanGestureAvailable=true;</script>\n</head>`);
if (settings.gestureControlEnabled === true) {
@@ -2461,9 +2532,14 @@ export class WebServer extends EventEmitter {
openCodeConfig: muxSession.mode === 'opencode' ? savedState?.openCodeConfig : undefined,
codexConfig: muxSession.mode === 'codex' ? savedState?.codexConfig : undefined,
geminiConfig: muxSession.mode === 'gemini' ? savedState?.geminiConfig : undefined,
antigravityConfig: muxSession.mode === 'antigravity' ? savedState?.antigravityConfig : undefined,
envOverrides: savedEnvOverrides,
effort: savedState?.effort,
attachmentHistory: savedAttachmentHistory,
// The pane's last Enter. Without it the response viewer would show
// the launch conversation until the user types again, even though
// the re-attached CLI is on a post-`/clear` one.
lastSubmitAt: savedState?.lastSubmitAt,
// Remote SSH metadata must round-trip on recovery: without it the
// attach cwd falls back to the (nonexistent-locally) remote path and
// respawn rebuilds a LOCAL command, breaking the pane and silently
@@ -2780,6 +2856,11 @@ export class WebServer extends EventEmitter {
// Gracefully close all SSE connections and clear batching state
this.sse.stop();
// Release every pending long-poll waiter. Their timers are deliberately not
// unref'd (an unref'd timer can let the process exit mid-wait and strand the
// response), so without this a 10-minute wait holds shutdown open.
sessionWaits.cancelEverything();
this.lastRecordedTokens.clear();
// Stop multiplexer and flush pending saves
+28
View File
@@ -27,6 +27,7 @@ import type { RalphStatusBlock, CircuitBreakerStatus } from '../types.js';
import { SseEvent } from './sse-events.js';
import { getLifecycleLog } from '../session-lifecycle-log.js';
import { fileStreamManager } from '../file-stream-manager.js';
import { sessionWaits } from './session-wait-registry.js';
/** Stored listener references for session cleanup (prevents memory leaks) */
export interface SessionListenerRefs {
@@ -92,6 +93,9 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Batches PTY output → broadcasts `session:terminal` at 16-50ms intervals */
terminal: (data) => {
// Feeds `GET /api/sessions/:id/wait-output`. No-ops with a single Map lookup
// when nothing is waiting, which is the case on virtually every chunk.
sessionWaits.notifyOutput(session.id, data);
deps.batchTerminalData(session.id, data);
},
@@ -137,6 +141,28 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Broadcasts `session:exit` + `session:updated` — PTY process exited; cleans up respawn, timers, listeners */
exit: (code) => {
// Before anything that can throw: a caller blocked on this session must learn
// the process died rather than sit until its timeout.
//
// Both halves are required, in this order — the same pair `_doCleanupSession`
// uses on the delete path, for the same reason. `notifySignal` resolves ONLY
// waiters that asked for `exit`; everyone else (`until=working`, `until=stop`,
// every wait-output) would keep a slot in the process-wide pool until their
// timeout, on a session whose feeds this very handler is about to tear down:
// `removeSessionListenerRefs` below detaches the `terminal` listener that is
// the only input to `notifyOutput`, and the `idle`/`working` listeners with it.
// Nothing can reach those waiters afterwards, so holding them is a guaranteed
// ten-minute lie. `cancelAll` answers them `ended: true`, which the plan's §3.6
// specifies for exactly this case ("Never hang").
//
// Safe against the respawn cycle: a respawn writes `/clear` + a kickstart
// prompt through the mux and never restarts the PTY, so it emits no `exit` and
// cannot cancel an orchestrating agent's wait. And for an agent driving a
// worker this is the right trade even when the PTY exit was only a tmux
// DETACH: `ended` means "re-check and re-issue", one extra round trip, versus
// burning the caller's entire timeout learning nothing.
sessionWaits.notifySignal(session.id, 'exit');
sessionWaits.cancelAll(session.id);
getLifecycleLog().log({
event: 'exit',
sessionId: session.id,
@@ -187,6 +213,7 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Broadcasts `session:working` — Claude started processing */
working: () => {
sessionWaits.notifySignal(session.id, 'working');
deps.broadcast(SseEvent.SessionWorking, { id: session.id });
const tracker = deps.getRunSummaryTracker(session.id);
if (tracker) {
@@ -197,6 +224,7 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
/** Broadcasts `session:idle` — Claude finished processing, waiting for input */
idle: () => {
sessionWaits.notifySignal(session.id, 'idle');
deps.broadcast(SseEvent.SessionIdle, { id: session.id });
deps.broadcastSessionStateDebounced(session.id);
const tracker = deps.getRunSummaryTracker(session.id);
File diff suppressed because it is too large Load Diff
+125
View File
@@ -0,0 +1,125 @@
import { describe, expect, it } from 'vitest';
import { CreateSessionSchema, QuickStartSchema } from '../src/web/schemas.js';
import { buildSpawnCommand } from '../src/tmux-manager.js';
import { defaultDockerCommandForMode } from '../src/docker-hosts.js';
import { defaultRemoteCommandForMode } from '../src/remote-hosts.js';
import { isExternalCliMode, isAltScreenStripMode } from '../src/session.js';
describe('Antigravity mode schemas', () => {
it('accepts Antigravity session creation config', () => {
const parsed = CreateSessionSchema.parse({
workingDir: '/tmp',
mode: 'antigravity',
antigravityConfig: {
model: 'gemini-3-pro',
dangerouslySkipPermissions: true,
},
});
expect(parsed.mode).toBe('antigravity');
expect(parsed.antigravityConfig).toEqual({
model: 'gemini-3-pro',
dangerouslySkipPermissions: true,
});
});
it('accepts Antigravity quick-start config', () => {
const parsed = QuickStartSchema.parse({
caseName: 'antigravity-case',
mode: 'antigravity',
antigravityConfig: {
resumeConversationId: 'conv-1234abcd',
},
});
expect(parsed.mode).toBe('antigravity');
expect(parsed.antigravityConfig?.resumeConversationId).toBe('conv-1234abcd');
});
it('rejects unsafe Antigravity model strings', () => {
expect(() =>
CreateSessionSchema.parse({
workingDir: '/tmp',
mode: 'antigravity',
antigravityConfig: { model: 'agy; rm -rf /' },
})
).toThrow();
});
it('allows ANTIGRAVITY_* env overrides and still rejects unknown prefixes', () => {
const parsed = CreateSessionSchema.parse({
workingDir: '/tmp',
mode: 'antigravity',
envOverrides: { ANTIGRAVITY_LOG_LEVEL: 'debug' },
});
expect(parsed.envOverrides).toEqual({ ANTIGRAVITY_LOG_LEVEL: 'debug' });
expect(() =>
CreateSessionSchema.parse({
workingDir: '/tmp',
envOverrides: { RANDOM_PREFIX_KEY: 'x' },
})
).toThrow();
});
});
describe('Antigravity spawn command', () => {
it('builds a bare agy command when no config is sent (safe default, no bypass)', () => {
const cmd = buildSpawnCommand({ mode: 'antigravity', sessionId: 'abc12345' });
expect(cmd).toBe('agy');
});
it('adds --dangerously-skip-permissions only when explicitly requested', () => {
const cmd = buildSpawnCommand({
mode: 'antigravity',
sessionId: 'abc12345',
antigravityConfig: { dangerouslySkipPermissions: true, model: 'gemini-3-pro' },
});
expect(cmd).toBe('agy --dangerously-skip-permissions --model gemini-3-pro');
});
it('passes --conversation for resume and drops unsafe ids', () => {
expect(
buildSpawnCommand({
mode: 'antigravity',
sessionId: 'abc12345',
antigravityConfig: { resumeConversationId: 'conv-99' },
})
).toBe('agy --conversation conv-99');
expect(
buildSpawnCommand({
mode: 'antigravity',
sessionId: 'abc12345',
antigravityConfig: { resumeConversationId: 'x; rm -rf /' },
})
).toBe('agy');
});
it('drops unsafe model strings from the spawn command', () => {
expect(
buildSpawnCommand({
mode: 'antigravity',
sessionId: 'abc12345',
antigravityConfig: { model: 'a`b' },
})
).toBe('agy');
});
});
describe('Antigravity mode gates', () => {
it('is an external CLI mode (readiness/ralph/respawn gating)', () => {
expect(isExternalCliMode('antigravity')).toBe(true);
});
it('is NOT an alt-screen strip mode (unverified Go TUI, like opencode)', () => {
expect(isAltScreenStripMode('antigravity')).toBe(false);
});
it('has docker/remote default commands', () => {
expect(defaultDockerCommandForMode('antigravity')).toBe('exec agy');
// Routed through an interactive login shell so per-user PATH entries resolve —
// same fix as the other remote agent CLIs (see defaultRemoteCommandForMode).
expect(defaultRemoteCommandForMode('antigravity')).toBe('exec "${SHELL:-/bin/sh}" -i -l -c \'agy\'');
});
});
+109
View File
@@ -0,0 +1,109 @@
/**
* Issue #205, round 2: `getClaudeCliVersion()` used to cache FAILURE forever.
*
* It stored `null` on any exception and guarded on `!== undefined`, so a single
* failed probe — the 5s exec timeout, a PATH-starved systemd/launchd
* environment, a transient fs hiccup — at the first Claude session start left
* `cliVersion` undefined for every Claude session until the server restarted.
* An undefined `cliVersion` silently disables wheel-forwarding to Claude's own
* transcript (`_shouldForwardWheelToApp`), which is the only route to history
* for a repaint-mode pane: a dead wheel on every device at once, which is what
* the reporter described (phone + iPad + laptop all broken together points at a
* SERVER-side cause, not a browser one).
*
* The probe itself can't run under vitest (it would spawn a real `claude`), so
* these drive the cache policy directly with an injected probe and clock.
*/
import { describe, expect, it, vi } from 'vitest';
import {
claudeVersionRetryDelayMs,
getClaudeCliVersion,
resolveClaudeCliVersion,
type ClaudeVersionProbeState,
} from '../src/utils/claude-cli-resolver.js';
const freshState = (): ClaudeVersionProbeState => ({ failures: 0, lastFailureAt: 0 });
describe('claude --version probe caching', () => {
it('probes once on success and never spawns again', () => {
const state = freshState();
const probe = vi.fn(() => '2.1.223');
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBe('2.1.223');
expect(resolveClaudeCliVersion(state, 2_000, probe)).toBe('2.1.223');
expect(resolveClaudeCliVersion(state, 9_999_999, probe)).toBe('2.1.223');
expect(probe).toHaveBeenCalledTimes(1);
});
it('RETRIES after a failed probe instead of poisoning the process', () => {
const state = freshState();
const probe = vi
.fn<() => string | null>()
.mockImplementationOnce(() => {
throw new Error('spawn claude ETIMEDOUT'); // the shipped failure mode
})
.mockImplementationOnce(() => '2.1.223');
// First session start: probe blows up, no version.
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
// Immediately after, the negative cache holds — no probe storm.
expect(resolveClaudeCliVersion(state, 30_000, probe)).toBeNull();
expect(probe).toHaveBeenCalledTimes(1);
// Once the retry window elapses, the next session start probes again and
// wheel-forwarding comes back without a server restart.
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBe('2.1.223');
expect(probe).toHaveBeenCalledTimes(2);
});
it('treats an unparseable version like a failure (retryable, not cached)', () => {
const state = freshState();
const probe = vi.fn<() => string | null>(() => null); // e.g. output without a x.y.z
expect(resolveClaudeCliVersion(state, 1_000, probe)).toBeNull();
expect(resolveClaudeCliVersion(state, 61_000, probe)).toBeNull();
expect(probe).toHaveBeenCalledTimes(2);
expect(state.version).toBeUndefined(); // nothing cached as "known bad"
});
it('clears the failure streak once a probe succeeds', () => {
const state = freshState();
const probe = vi
.fn<() => string | null>()
.mockImplementationOnce(() => null)
.mockImplementationOnce(() => '2.1.223');
resolveClaudeCliVersion(state, 1_000, probe);
expect(state.failures).toBe(1);
resolveClaudeCliVersion(state, 61_000, probe);
expect(state.failures).toBe(0);
expect(state.lastFailureAt).toBe(0);
});
it('backs off so a genuinely missing binary cannot probe on every session start', () => {
expect(claudeVersionRetryDelayMs(0)).toBe(0);
expect(claudeVersionRetryDelayMs(1)).toBe(60_000);
expect(claudeVersionRetryDelayMs(2)).toBe(120_000);
expect(claudeVersionRetryDelayMs(3)).toBe(240_000);
// Capped, so it keeps retrying forever without ever spinning.
expect(claudeVersionRetryDelayMs(50)).toBe(15 * 60_000);
const state = freshState();
const probe = vi.fn<() => string | null>(() => null);
resolveClaudeCliVersion(state, 0, probe); // failure 1 → retry at 60s
resolveClaudeCliVersion(state, 30_000, probe); // still inside the window
expect(probe).toHaveBeenCalledTimes(1);
resolveClaudeCliVersion(state, 60_000, probe); // failure 2 → retry at 120s
resolveClaudeCliVersion(state, 119_000, probe);
expect(probe).toHaveBeenCalledTimes(2);
resolveClaudeCliVersion(state, 180_001, probe);
expect(probe).toHaveBeenCalledTimes(3);
});
it('stays hermetic under vitest without recording a phantom failure', () => {
// The guard returns before the probe, and — unlike the old code, which wrote
// null into the cache here — leaves the cache untouched.
expect(getClaudeCliVersion()).toBeNull();
expect(getClaudeCliVersion()).toBeNull();
});
});
+58 -2
View File
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { Session, isAltScreenStripMode } from '../src/session.js';
import { Session, isAltScreenStripMode, isMuxAltScreenOnlyStripMode } from '../src/session.js';
type SessionInternals = {
_handleTerminalOutput(data: string): void;
@@ -81,7 +81,7 @@ describe('Claude terminal scrollback strip', () => {
});
});
describe('Shell terminal output is NOT stripped (vim/less/htop need the alt screen)', () => {
describe('Shell terminal output on a DIRECT PTY is NOT stripped (vim/less/htop need the alt screen)', () => {
it('leaves alt-screen toggles, scrollback-erase, and mouse-tracking intact for shell', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell' });
@@ -91,3 +91,59 @@ describe('Shell terminal output is NOT stripped (vim/less/htop need the alt scre
expect(session.terminalBuffer).toBe(vimLike);
});
});
describe('isMuxAltScreenOnlyStripMode', () => {
it('covers exactly the modes the full strip does not, and only under tmux', () => {
for (const mode of ['shell', 'opencode', 'antigravity'] as const) {
expect(isMuxAltScreenOnlyStripMode(mode, true)).toBe(true);
// Direct-PTY fallback: the program's own alt screen really does reach xterm.
expect(isMuxAltScreenOnlyStripMode(mode, false)).toBe(false);
}
// The full strip already owns these; never double-gate them here.
for (const mode of ['claude', 'codex', 'gemini'] as const) {
expect(isMuxAltScreenOnlyStripMode(mode, true)).toBe(false);
}
});
});
describe('tmux-backed shell: strip tmux’s own client smcup, keep everything else (#205)', () => {
it('drops alt-screen toggles so xterm keeps a scrollback buffer', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
// What a real `tmux attach` emits as its first bytes.
handleOutput(session, '\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2Jprompt$ ');
expect(session.terminalBuffer).toBe('\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2Jprompt$ ');
expect(session.terminalBuffer).not.toContain('\x1b[?1049h');
});
it('KEEPS 3J and mouse-tracking, unlike the full strip', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
// `clear` legitimately wipes scrollback; htop/vim mouse modes are passed
// through by tmux even with `mouse off` and must keep working.
handleOutput(session, '\x1b[3J\x1b[?1002h\x1b[?1006hhtop\x1b[?1006l\x1b[?1002l');
expect(session.terminalBuffer).toBe('\x1b[3J\x1b[?1002h\x1b[?1006hhtop\x1b[?1006l\x1b[?1002l');
});
it('reassembles alt-screen sequences split across PTY chunk boundaries', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
const emitted: string[] = [];
session.on('terminal', (data) => emitted.push(data));
handleOutput(session, 'before\x1b[?104');
handleOutput(session, '9h after');
expect(session.terminalBuffer).toBe('before after');
expect(emitted).toEqual(['before', ' after']);
});
it('applies to opencode and antigravity too', () => {
for (const mode of ['opencode', 'antigravity'] as const) {
const session = new Session({ workingDir: '/tmp', mode, useMux: true });
handleOutput(session, '\x1b[?1049hTUI\x1b[3J');
expect(session.terminalBuffer).toBe('TUI\x1b[3J');
}
});
});
+98
View File
@@ -0,0 +1,98 @@
/**
* @fileoverview Unit tests for the File Viewer edit-mode policy module.
*
* Pure functions only — no IO, no server.
* Port: N/A (no server)
*/
import { describe, it, expect } from 'vitest';
import {
MAX_EDITABLE_BYTES,
applyEol,
detectEol,
isDeniedEditRelativePath,
isEditableFileName,
} from '../src/config/file-editing.js';
describe('file-editing policy', () => {
describe('isEditableFileName', () => {
it('allows common text extensions', () => {
for (const name of ['a.ts', 'b.md', 'c.json', 'd.py', 'style.css', 'notes.txt', 'x.yml', 'Q.SQL']) {
expect(isEditableFileName(name), name).toBe(true);
}
});
it('allows well-known basenames regardless of case', () => {
for (const name of ['Dockerfile', 'Makefile', 'LICENSE', '.gitignore', '.editorconfig', '.nvmrc']) {
expect(isEditableFileName(name), name).toBe(true);
}
});
it('rejects binary/media/document extensions', () => {
for (const name of ['a.png', 'b.pdf', 'c.docx', 'd.zip', 'e.woff2', 'f.mp4', 'g.exe']) {
expect(isEditableFileName(name), name).toBe(false);
}
});
it('rejects svg and env (deliberate v1 exclusions)', () => {
expect(isEditableFileName('image.svg')).toBe(false);
expect(isEditableFileName('config.env')).toBe(false);
});
it('rejects extensionless and unknown-dotfile names not on the basename list', () => {
expect(isEditableFileName('somebinary')).toBe(false);
expect(isEditableFileName('.bashrc')).toBe(false);
expect(isEditableFileName('archive.xyz')).toBe(false);
});
});
describe('isDeniedEditRelativePath', () => {
it('denies anything inside a .git directory at any depth', () => {
expect(isDeniedEditRelativePath('.git/config')).toBe(true);
expect(isDeniedEditRelativePath('.git/hooks/pre-commit')).toBe(true);
expect(isDeniedEditRelativePath('sub/module/.git/HEAD')).toBe(true);
});
it('allows non-.git paths, including names merely containing "git"', () => {
expect(isDeniedEditRelativePath('src/index.ts')).toBe(false);
expect(isDeniedEditRelativePath('.github/workflows/ci.yml')).toBe(false);
expect(isDeniedEditRelativePath('digits/file.md')).toBe(false);
expect(isDeniedEditRelativePath('.gitignore')).toBe(false);
});
});
describe('detectEol / applyEol', () => {
it('detects LF, CRLF, and defaults to LF for single-line text', () => {
expect(detectEol('a\nb\nc')).toBe('lf');
expect(detectEol('a\r\nb\r\nc')).toBe('crlf');
expect(detectEol('no newline at all')).toBe('lf');
expect(detectEol('')).toBe('lf');
});
it('picks the dominant style for mixed-EOL text', () => {
expect(detectEol('a\r\nb\r\nc\nd')).toBe('crlf');
expect(detectEol('a\nb\nc\r\nd')).toBe('lf');
});
it('applyEol round-trips a textarea-normalized (LF) buffer back to CRLF', () => {
const original = 'line1\r\nline2\r\nline3';
const textareaValue = original.replace(/\r\n/g, '\n');
expect(applyEol(textareaValue, detectEol(original))).toBe(original);
});
it('applyEol is idempotent and never doubles CR', () => {
expect(applyEol('a\r\nb', 'crlf')).toBe('a\r\nb');
expect(applyEol('a\r\nb', 'lf')).toBe('a\nb');
expect(applyEol('a\nb', 'lf')).toBe('a\nb');
});
it('preserves a UTF-8 BOM through the EOL rewrite', () => {
const withBom = 'hello\nworld';
expect(applyEol(withBom, 'crlf')).toBe('hello\r\nworld');
});
});
it('exposes a sane editable-bytes cap', () => {
expect(MAX_EDITABLE_BYTES).toBe(512 * 1024);
});
});
+9
View File
@@ -20,6 +20,15 @@ describe('frontend public asset tooling', () => {
expect(appJs.includes(0)).toBe(false);
});
it('uses the same message wrapper for brief and full response views', () => {
const appJs = readFileSync(resolve(repoRoot, 'src/web/public/app.js'), 'utf8');
expect(appJs).toContain("body.appendChild(this._buildResponseViewerMessage(lastResponse, 'assistant'");
expect(appJs).toContain('body.appendChild(this._buildResponseViewerMessage(msg.text, msg.role, agentLabel));');
expect(appJs).toContain("div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');");
expect(appJs).toContain("renderedText.className = 'rv-text';");
});
it('runs the public asset check script', () => {
expect(() => {
execFileSync('npm', ['run', 'check:public-assets', '--silent'], {
+2
View File
@@ -52,6 +52,8 @@ describe('help modal shortcuts', () => {
});
it('documents terminal input shortcuts without advertising stale run shortcuts', () => {
expectShortcut(helpModal, ['Ctrl', 'C'], 'Copy Selection');
expectShortcut(helpModal, ['Ctrl', 'Shift', 'C'], 'Copy Selection');
expectShortcut(helpModal, ['Ctrl', 'L'], 'Clear Terminal');
expectShortcut(helpModal, ['Ctrl', '+'], 'Increase Font');
expectShortcut(helpModal, ['Ctrl', '-'], 'Decrease Font');
+19
View File
@@ -133,6 +133,25 @@ describe('refreshStaleCodemanHooks', () => {
expect(after.hooks.CustomEvent).toEqual(customEvent);
});
// A case can be current on the secret AND the background-wake hook and still carry
// the `-k`-less curl shape, which exits 60 against a self-signed HTTPS API and is
// swallowed by `|| true` — every hook event dead, silently. The refresh must treat
// that as a third stale shape.
it('heals a current-looking block whose hook curls lack -k (HTTPS self-signed installs)', async () => {
const { generateHooksConfig } = await import('../src/hooks-config.js');
const flagless = JSON.parse(JSON.stringify(generateHooksConfig()).replaceAll('curl -sk ', 'curl -s '));
writeFileSync(settingsPath, JSON.stringify({ hooks: flagless.hooks }, null, 2));
await refreshStaleCodemanHooks(dir);
const after = readFileSync(settingsPath, 'utf-8');
expect(after).toContain('curl -sk -X POST');
expect(after).not.toContain('curl -s -X POST');
// and the pass is convergent: a second refresh must not rewrite
await refreshStaleCodemanHooks(dir);
expect(readFileSync(settingsPath, 'utf-8')).toBe(after);
});
it('is a no-op when settings.local.json is absent (does not create one)', async () => {
await refreshStaleCodemanHooks(dir);
expect(existsSync(settingsPath)).toBe(false);
+11
View File
@@ -95,6 +95,17 @@ describe('generateHooksConfig', () => {
expect(notifHooks[0].hooks[0].command).toContain('|| true');
});
// On --https/tailscale installs CODEMAN_API_URL is HTTPS with a self-signed cert.
// A `-k`-less hook curl exits 60 there, the `|| true` swallows it, and every hook
// event (stop, permission_prompt, elicitation_dialog, idle_prompt, teammate_idle,
// task_completed) dies silently — killing respawn's idle signals and the wait
// endpoints' stop/blocked. The statusline exporter always carried -k; the hooks must too.
it('every hook curl tolerates a self-signed HTTPS API (curl -sk)', () => {
const serialized = JSON.stringify(generateHooksConfig());
expect(serialized).toContain('curl -sk -X POST');
expect(serialized).not.toContain('curl -s -X POST');
});
it('should set timeout to 10 seconds (hook timeout fields are seconds)', () => {
const config = generateHooksConfig();
const notifHooks = config.hooks.Notification as Array<{ hooks: Array<{ timeout: number }> }>;
+93
View File
@@ -89,4 +89,97 @@ describe('Stable HTTP contract (live server)', () => {
expect(body.success).toBe(false);
expect(body.errorCode).toBe('INVALID_INPUT');
});
/**
* The agent wait primitives, through the REAL pipeline.
*
* Their own route tests hand-roll a partial copy of the preSerialization hook that
* maps errorCode to status but does NOT wrap bare payloads — so nothing there
* proves these routes emit a correct envelope, a correct status, or work through
* the /api/v1 alias, and one assertion in them pins `{}` for a response no client
* will ever receive. This is the file whose docstring already claims that scope.
*/
describe('agent wait primitives', () => {
let sessionId: string;
beforeAll(async () => {
const res = await fetch(`${base}/api/sessions`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({}),
});
sessionId = (await res.json()).data.session.id;
expect(sessionId).toBeDefined();
});
afterAll(async () => {
await fetch(`${base}/api/sessions/${sessionId}`, { method: 'DELETE' });
});
it('answers a wait timeout as a 200 inside the envelope, on the /api/v1 alias', async () => {
// A timeout is the long-poll SUCCEEDING at "did this happen within N ms?"; a
// 4xx/5xx here would make every poll boundary indistinguishable from a failure.
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?until=working&timeout=1000`);
expect(res.status).toBe(200);
expect(res.headers.get('cache-control')).toBe('no-store');
const body = await res.json();
expect(body.success).toBe(true);
expect(body.data.sessionId).toBe(sessionId);
// The one shape all three wait endpoints share.
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.signal).toBeNull();
expect(body.data.wait.timeoutMs).toBe(1000);
expect(body.data.wait.until).toEqual(['working']);
});
it('returns a contract-shaped 400 for an unknown until token', async () => {
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?until=stpo`);
expect(res.status).toBe(400);
const body = await res.json();
expect(body.success).toBe(false);
expect(body.errorCode).toBe('INVALID_INPUT');
expect(body.error).toContain('stpo');
});
it('returns a contract-shaped 400 naming the bad query parameter', async () => {
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait?timeout=30s`);
expect(res.status).toBe(400);
const body = await res.json();
expect(body.errorCode).toBe('INVALID_INPUT');
expect(body.error).toContain('timeout');
});
it('wraps the non-wait input response as { success: true, data: {} }', async () => {
// What a client actually receives on the fire-and-forget path — NOT the bare
// `{}` the handler returns and the route tests assert.
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/input`, {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ input: 'hello' }),
});
expect(res.status).toBe(200);
expect(await res.json()).toEqual({ success: true, data: {} });
});
it('serves wait-output through the same envelope', async () => {
const res = await fetch(`${base}/api/v1/sessions/${sessionId}/wait-output?match=NEVER_APPEARS&timeout=1000`);
expect(res.status).toBe(200);
const body = await res.json();
expect(body.success).toBe(true);
expect(body.data.wait.matched).toBe(false);
expect(body.data.wait.timedOut).toBe(true);
expect(body.data.wait.match).toBe('NEVER_APPEARS');
});
it('404s an unknown session on both new routes, with the error envelope', async () => {
for (const path of ['wait?until=idle', 'wait-output?match=x']) {
const res = await fetch(`${base}/api/v1/sessions/nonexistent/${path}`);
expect(res.status).toBe(404);
const body = await res.json();
expect(body.success).toBe(false);
expect(body.errorCode).toBe('NOT_FOUND');
}
});
});
});
+113
View File
@@ -245,6 +245,119 @@ describe('Inline rename input', () => {
expect(result.threw).toBe(false);
});
it('Render guard: _renderSessionTabsImmediate() does not destroy an open rename input', async () => {
await resetState();
// The debounced tab render is scheduled by renderSessionTabs() but EXECUTED by
// _renderSessionTabsImmediate(). A render queued just before the rename opened
// still fires ~100ms later and lands in the executor directly, so the guard has
// to live there too, otherwise the incremental branch rewrites .tab-name's
// innerHTML and the user's half-typed description is lost.
//
// The tab MUST live inside the real #sessionTabs container and be the only
// session in app.sessions: the renderer walks that container, so a synthetic
// node parked on <body> would make this test pass with the guard removed.
const result = await page.evaluate(() => {
const app = (
window as unknown as {
app: {
sessions: Map<string, { id: string; name: string; status: string }>;
sessionOrder: string[];
startInlineRename: (id: string) => void;
_renderSessionTabsImmediate: () => void;
_activeRename: unknown;
};
}
).app;
const id = 'render-race';
app.sessions.set(id, { id, name: 'w9-case', status: 'idle' });
app.sessionOrder = [id];
const container = document.getElementById('sessionTabs') as HTMLElement;
const tab = document.createElement('div');
tab.setAttribute('data-test-tab', '1');
tab.className = 'session-tab';
tab.dataset.id = id;
tab.innerHTML =
'<span class="tab-status idle"></span><span class="tab-info"><span class="tab-name-row">' +
`<span class="tab-name" data-session-id="${id}">w9-case</span>` +
'</span></span>';
container.appendChild(tab);
app.startInlineRename(id);
const input = document.querySelector('input.tab-rename-input') as HTMLInputElement | null;
if (!input) return { opened: false };
input.value = 'half-typed';
// Exactly what a debounce timer queued before the rename would do.
app._renderSessionTabsImmediate();
const after = document.querySelector('input.tab-rename-input') as HTMLInputElement | null;
return {
opened: true,
stillInDom: !!after && document.body.contains(after),
value: after?.value ?? null,
renameStillActive: !!app._activeRename,
};
});
expect(result.opened).toBe(true);
expect(result.stillInDom).toBe(true);
expect(result.value).toBe('half-typed');
expect(result.renameStillActive).toBe(true);
});
it('Modal: closeSessionOptions() commits the Session Name field before clearing the id', async () => {
await resetState();
// Every autosave handler in the session-options modal bails on a null
// editingSessionId, and hiding the modal blurs the focused input. If the id is
// cleared first, the blur-driven save is dropped and the typed name vanishes,
// which is what Escape and backdrop-click used to do.
const result = await page.evaluate(async () => {
const app = (
window as unknown as {
app: {
editingSessionId: string | null;
sessions: Map<string, { id: string; name: string }>;
closeSessionOptions: () => void;
};
}
).app;
app.sessions.set('modal-id', { id: 'modal-id', name: 'w9-case' });
app.editingSessionId = 'modal-id';
const nameInput = document.getElementById('modalSessionName') as HTMLInputElement;
const modal = document.getElementById('sessionOptionsModal') as HTMLElement;
modal.classList.add('active');
// The Session Name field lives on the modal's Context tab, which is hidden
// until selected: a hidden input cannot take focus.
document.getElementById('context-tab')?.classList.remove('hidden');
nameInput.value = 'mydesc';
nameInput.focus();
const wasFocused = document.activeElement === nameInput;
let putBody: string | null = null;
const origFetch = window.fetch;
window.fetch = (async (input: RequestInfo | URL, init?: RequestInit) => {
if (String(input).includes('/api/sessions/modal-id/name')) putBody = String(init?.body ?? '');
return new Response('{"success":true}', { status: 200 });
}) as typeof window.fetch;
app.closeSessionOptions();
await new Promise((r) => setTimeout(r, 30));
window.fetch = origFetch;
modal.classList.remove('active');
return { wasFocused, putBody, editingAfter: app.editingSessionId };
});
expect(result.wasFocused).toBe(true);
// Prefixed session: the suffix the user typed is appended to the w9-case prefix.
expect(result.putBody).toContain('w9-case: mydesc');
expect(result.editingAfter).toBe(null);
});
it('Re-entry: starting rename while one is active aborts the previous one', async () => {
await resetState();
expect(await startRename('first-id', 'First')).toBe(true);
+18
View File
@@ -58,4 +58,22 @@ describe('keyboard shortcuts', () => {
expect(appSource).toContain('if (this.matchesShortcutEvent(e, shortcut))');
expect(appSource).toContain('if (shortcut.disabled || !shortcut.action) continue;');
});
it('keeps the interrupt when Ctrl+C copies a selection (#211)', () => {
// The xterm handler owns this decision, and the no-selection path must fall
// through with NO preventDefault so xterm still evaluates Ctrl+C into 0x03.
expect(terminalUiSource).toContain('this.shouldCopyTerminalSelectionFromShortcut?.(ev)');
expect(terminalUiSource).toMatch(
/const selection = this\.terminal\.hasSelection\?\.\(\) \? this\.terminal\.getSelection\(\) : '';/
);
expect(terminalUiSource).toContain('void this.copyTerminalSelection(selection);');
expect(appSource).toContain("id: 'copy-selection'");
});
it('documents the terminal copy shortcut in help and README', () => {
expect(helpHtml).toContain('<kbd>Ctrl</kbd>+<kbd>C</kbd>');
expect(helpHtml).toContain('<kbd>Ctrl</kbd>+<kbd>Shift</kbd>+<kbd>C</kbd>');
expect(readme).toContain('`Ctrl/Cmd+C`');
expect(readme).toContain('`Ctrl+Shift+C`');
});
});
+240
View File
@@ -0,0 +1,240 @@
/**
* @fileoverview Local-echo gating and input-ordering helpers for codex
* sessions (issues #218/#219/#220/#222).
*
* Codex's composer is interactive per keystroke: typing "/" pops a
* live-filtering command picker (#222), the composer grows as it wraps
* (#220), pastes arrive bracketed (#219) and arrows edit server-side state
* (#218). The buffer-until-Enter local echo overlay starves all of that, so
* codex-mode sessions must use plain PTY echo like shell. The shared overlay
* branch (claude/gemini/opencode) additionally flushes typed-but-unsent text
* before forwarding bracketed pastes and composer nav keys, and hands the
* session to pass-through after a nav key.
*
* Loaded via `vm` with a stubbed context (no jsdom), mirroring
* test/input-send-order.test.ts. End-to-end behavior was verified against a
* real codex 0.147.0 TUI in tmux through a headless browser.
*/
import { readFileSync } from 'node:fs';
import { performance } from 'node:perf_hooks';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
type OverlayStub = {
pendingText: string;
cleared: number;
suppressed: number;
prompts: unknown[];
clear(): void;
suppressBufferDetection(): void;
setPrompt(p: unknown): void;
appendText: ReturnType<typeof vi.fn>;
};
type AppInstance = {
activeSessionId: string | null;
sessions: Map<string, { mode: string }>;
terminal?: { focus: () => void };
_localEchoEnabled?: boolean;
_localEchoOverlay?: OverlayStub;
_pendingInput: string;
_flushedOffsets?: Map<string, number>;
_flushedTexts?: Map<string, string>;
_echoPassthroughSessions?: Set<string>;
loadAppSettingsFromStorage: () => Record<string, unknown>;
sendInput: ReturnType<typeof vi.fn>;
_updateLocalEchoState(): void;
_flushLocalEchoPending(): void;
insertTerminalText(text: string): void;
};
function loadContext() {
const read = (f: string) => readFileSync(resolve(import.meta.dirname, `../src/web/public/${f}`), 'utf8');
const windowStub: Record<string, unknown> = {
addEventListener: vi.fn(),
removeEventListener: vi.fn(),
};
const context = vm.createContext({
console,
performance,
setInterval: vi.fn(),
clearInterval: vi.fn(),
setTimeout,
clearTimeout,
requestAnimationFrame: vi.fn(),
HTMLCanvasElement: class HTMLCanvasElement {},
WebSocket: { OPEN: 1 },
fetch: vi.fn(),
document: { addEventListener: vi.fn(), documentElement: { dataset: {} } },
localStorage: {
length: 0,
key: vi.fn(),
getItem: vi.fn(),
setItem: vi.fn(),
removeItem: vi.fn(),
},
window: windowStub,
MobileDetection: {
isTouchDevice: () => true,
isHandheldDevice: () => false,
getDeviceType: () => 'desktop',
},
});
vm.runInContext(
`${read('constants.js')}\n${read('app.js')}\n${read('terminal-ui.js')}\nglobalThis.__CodemanApp = CodemanApp;`,
context
);
const CodemanApp = (context as unknown as { __CodemanApp: { prototype: object } }).__CodemanApp;
return {
CodemanApp,
terminalInput: (windowStub as { CodemanTerminalInput?: Record<string, unknown> }).CodemanTerminalInput!,
};
}
const { CodemanApp, terminalInput } = loadContext();
const isComposerNavKey = terminalInput.isComposerNavKey as (data: string) => boolean;
function makeOverlay(pending = ''): OverlayStub {
return {
pendingText: pending,
cleared: 0,
suppressed: 0,
prompts: [],
clear() {
this.cleared++;
this.pendingText = '';
},
suppressBufferDetection() {
this.suppressed++;
},
setPrompt(p: unknown) {
this.prompts.push(p);
},
appendText: vi.fn(),
};
}
function makeApp(mode: string, overlay = makeOverlay()): AppInstance {
const app = Object.create(CodemanApp.prototype) as AppInstance;
app.activeSessionId = 's1';
app.sessions = new Map([['s1', { mode }]]);
app._localEchoOverlay = overlay;
app._pendingInput = '';
app._flushedOffsets = new Map([['s1', 3]]);
app._flushedTexts = new Map([['s1', 'abc']]);
app.loadAppSettingsFromStorage = () => ({ localEchoEnabled: true });
app.sendInput = vi.fn().mockResolvedValue(undefined);
return app;
}
describe('CodemanTerminalInput.isComposerNavKey', () => {
it.each([
'\x1b[A',
'\x1b[B',
'\x1b[C',
'\x1b[D',
'\x1b[H',
'\x1b[F',
'\x1bOA',
'\x1bOD',
'\x1bOH',
'\x1bOF',
'\x1b[1;5C', // Ctrl+Right
'\x1b[1;2A', // Shift+Up
'\x1b[3~', // Delete
'\x1b[3;5~', // Ctrl+Delete
'\x1b[5~', // PgUp
'\x1b[6~', // PgDn
'\x1b[1~', // Home variant
'\x1b[4~', // End variant
])('classifies %j as a composer nav key', (seq) => {
expect(isComposerNavKey(seq)).toBe(true);
});
it.each([
'\x1b[?1;2c', // DA1 response
'\x1b[>0;276;0c', // DA2 response
'\x1b[12;34R', // CPR response
'\x1b[1;3R', // CPR response (small coords)
'\x1b[0n', // DSR response
'\x1b[15~', // F5 (function keys stay out)
'\x1b[200~hi\x1b[201~', // bracketed paste
'\x1b[?u', // kitty keyboard query response
'\x1bOP', // F1
'\x1b',
'a',
'abc',
'\r',
])('does NOT classify %j as a composer nav key', (seq) => {
expect(isComposerNavKey(seq)).toBe(false);
});
it('exports the bracketed paste prefix xterm puts on terminal.paste()', () => {
expect(terminalInput.BRACKETED_PASTE_START).toBe('\x1b[200~');
});
});
describe('_updateLocalEchoState mode gating', () => {
it('disables the overlay for codex sessions even with the setting ON (issues #218/#219/#220/#222)', () => {
const overlay = makeOverlay('pending');
const app = makeApp('codex', overlay);
app._updateLocalEchoState();
expect(app._localEchoEnabled).toBe(false);
expect(overlay.cleared).toBeGreaterThan(0);
});
it('disables the overlay for shell sessions (PTY provides its own echo)', () => {
const app = makeApp('shell');
app._updateLocalEchoState();
expect(app._localEchoEnabled).toBe(false);
});
it.each(['claude', 'gemini', 'opencode'])('keeps the overlay enabled for %s sessions', (mode) => {
const overlay = makeOverlay();
const app = makeApp(mode, overlay);
app._updateLocalEchoState();
expect(app._localEchoEnabled).toBe(true);
expect(overlay.prompts.length).toBeGreaterThan(0);
});
});
describe('_flushLocalEchoPending', () => {
it('moves pending text into _pendingInput and resets overlay + flushed tracking', () => {
const overlay = makeOverlay('hello');
const app = makeApp('claude', overlay);
app._flushLocalEchoPending();
expect(app._pendingInput).toBe('hello');
expect(overlay.cleared).toBe(1);
expect(overlay.suppressed).toBe(1);
expect(app._flushedOffsets!.has('s1')).toBe(false);
expect(app._flushedTexts!.has('s1')).toBe(false);
});
it('appends nothing when the overlay is empty', () => {
const app = makeApp('claude', makeOverlay(''));
app._flushLocalEchoPending();
expect(app._pendingInput).toBe('');
});
});
describe('insertTerminalText pass-through routing', () => {
it('appends to the overlay while local echo is buffering', () => {
const overlay = makeOverlay();
const app = makeApp('claude', overlay);
app._localEchoEnabled = true;
app.insertTerminalText('path.txt');
expect(overlay.appendText).toHaveBeenCalledWith('path.txt');
expect(app.sendInput).not.toHaveBeenCalled();
});
it('sends directly while the session is in nav-key pass-through', () => {
const overlay = makeOverlay();
const app = makeApp('claude', overlay);
app._localEchoEnabled = true;
app._echoPassthroughSessions = new Set(['s1']);
app.insertTerminalText('path.txt');
expect(app.sendInput).toHaveBeenCalledWith('path.txt');
expect(overlay.appendText).not.toHaveBeenCalled();
});
});
+339
View File
@@ -0,0 +1,339 @@
// Port: none (pure model + static markup assertions — no browser, no server).
//
// The phone home screen (src/web/public/mobile-overview.js) replaces the welcome
// overlay under 430px. Its grouping logic is the part that can silently go wrong:
// a session blocked on a permission prompt landing in "idle" is exactly the bug
// this surface exists to prevent. buildMobileOverviewModel() is pure for that
// reason, so it can be exercised here against plain objects.
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it } from 'vitest';
const PUBLIC = resolve(import.meta.dirname, '../src/web/public');
/** Minimal fake DOM node — enough surface for mobile-overview.js's programmatic builders. */
function fakeElement(): any {
const el: any = {
className: '',
type: '',
dataset: {},
style: {},
children: [] as any[],
setAttribute() {},
appendChild(child: any) {
el.children.push(child);
return child;
},
};
return el;
}
function loadOverviewApp(overrides: Record<string, any> = {}) {
const CodemanApp = function CodemanApp(this: any) {};
const context = vm.createContext({
CodemanApp,
console,
window: {},
document: {
getElementById: () => null,
createElement: () => fakeElement(),
createElementNS: () => fakeElement(),
},
MobileDetection: { getDeviceType: () => 'mobile' },
});
vm.runInContext(readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8'), context, {
filename: 'mobile-overview.js',
});
const app = new (CodemanApp as any)();
app.getSessionName = (session: any) => session.name || session.workingDir?.split('/').pop() || session.id.slice(0, 8);
app._shortenHomePath = (p: string) => (p || '').replace(/^\/home\/[^/]+\//, '~/');
app.loadAppSettingsFromStorage = () => ({});
Object.assign(app, overrides);
return app;
}
const CASES = [
{ name: 'claudeman', path: '/home/arkon/default/claudeman', location: 'local' },
{ name: 'beta', path: '/home/arkon/codeman-cases/beta', location: 'local' },
{ name: 'boxed', path: '/srv/boxed', location: 'docker' },
];
function session(over: Record<string, any>) {
return { id: 'x', status: 'idle', mode: 'claude', workingDir: '/home/arkon/default/claudeman', ...over };
}
describe('mobile overview model', () => {
it('routes a session with a pending permission prompt into NEEDS YOU, not idle', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'a', status: 'idle' })],
cases: CASES,
pendingHooks: new Map([['a', new Set(['permission_prompt'])]]),
});
expect(model.needsYou.map((r: any) => r.id)).toEqual(['a']);
expect(model.current).toHaveLength(0);
expect(model.needsYou[0].state).toBe('needs');
expect(model.needsYou[0].pill).toBe('needs you');
});
it('ranks an action hook above an idle hook above a stale busy status', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
// An idle_prompt hook on a session the server still calls 'busy': the hook
// is the newer signal, so it must win.
sessions: [
session({ id: 'busy-with-idle-hook', status: 'busy' }),
session({ id: 'elicit', status: 'busy' }),
session({ id: 'plain-busy', status: 'busy' }),
],
cases: CASES,
pendingHooks: new Map([
['busy-with-idle-hook', new Set(['idle_prompt'])],
['elicit', new Set(['elicitation_dialog'])],
]),
});
expect(model.needsYou.map((r: any) => r.id)).toEqual(['elicit', 'busy-with-idle-hook']);
expect(model.current.map((r: any) => r.id)).toEqual(['plain-busy']);
});
it('buckets busy / idle / stopped / error and labels each pill', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [
session({ id: 'w', status: 'busy' }),
session({ id: 'i', status: 'idle' }),
session({ id: 'd', status: 'stopped' }),
session({ id: 'e', status: 'error' }),
],
cases: CASES,
});
// Everything that is not blocked on you shares one "current" section,
// most demanding first.
expect(model.current.map((r: any) => [r.id, r.pill])).toEqual([
['w', 'working'],
['i', 'idle'],
['d', 'done'],
]);
expect(model.needsYou.map((r: any) => r.pill)).toEqual(['error']);
expect(model.sessionCount).toBe(4);
});
it('keeps the user tab order as the tiebreak inside a section', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'first' }), session({ id: 'second' }), session({ id: 'third' })],
cases: CASES,
sessionOrder: ['third', 'first', 'second'],
});
expect(model.current.map((r: any) => r.id)).toEqual(['third', 'first', 'second']);
});
it('matches a session started in a subdirectory to its case (longest prefix)', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [
session({ id: 'sub', workingDir: '/home/arkon/default/claudeman/src/web' }),
session({ id: 'outside', workingDir: '/tmp/scratch' }),
],
cases: [...CASES, { name: 'claudeman-web', path: '/home/arkon/default/claudeman/src/web' }],
});
const rows = Object.fromEntries(model.current.map((r: any) => [r.id, r.caseName]));
expect(rows.sub).toBe('claudeman-web');
expect(rows.outside).toBe('');
});
it('lists past conversations newest first and never repeats a live session', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [session({ id: 'live-1' })],
cases: CASES,
history: [
// Same id as the running session: the unified list includes live rows,
// and showing one in both sections would be a duplicate.
{ sessionId: 'live-1', workingDir: '/home/arkon/default/claudeman', lastActivityAt: 500 },
{
sessionId: 'old-a',
workingDir: '/home/arkon/codeman-cases/beta',
firstPrompt: 'fix the mobile header',
claudeSessionId: 'claude-uuid-a',
lastActivityAt: 100,
},
{
sessionId: 'old-b',
workingDir: '/home/arkon/default/claudeman',
name: 'w4-claudeman',
lastActivityAt: 400,
},
],
});
expect(model.past.map((r: any) => r.id)).toEqual(['old-b', 'old-a']);
expect(model.past[1]).toMatchObject({
title: 'fix the mobile header',
caseName: 'beta',
claudeSessionId: 'claude-uuid-a',
workingDir: '/home/arkon/codeman-cases/beta',
});
// A row with no prompt falls back to its name, so it is never a bare UUID.
expect(model.past[0].title).toBe('w4-claudeman');
});
it('does not title a past row with the transcript reader placeholder', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({
sessions: [],
cases: CASES,
history: [
{ sessionId: 'blank', workingDir: '/home/arkon/default/claudeman', firstPrompt: '(no content)' },
{ sessionId: 'spaces', workingDir: '/home/arkon/codeman-cases/beta', firstPrompt: ' ' },
],
});
expect(model.past.map((r: any) => r.title)).toEqual(['claudeman', 'beta']);
});
it('accepts the live Map as-is and survives an empty state', () => {
const app = loadOverviewApp();
const fromMap = app.buildMobileOverviewModel({
sessions: new Map([['a', session({ id: 'a' })]]),
cases: CASES,
});
expect(fromMap.current.map((r: any) => r.id)).toEqual(['a']);
const empty = app.buildMobileOverviewModel({});
expect(empty).toMatchObject({ needsYou: [], current: [], past: [], sessionCount: 0 });
});
it('no longer builds a spaces section', () => {
const app = loadOverviewApp();
const model = app.buildMobileOverviewModel({ sessions: [session({ id: 'a' })], cases: CASES });
expect(model.spaces).toBeUndefined();
});
});
describe('mobile overview gate', () => {
it('is phone-width only, off in solo windows, and off when explicitly disabled', () => {
expect(loadOverviewApp().shouldUseMobileOverview()).toBe(true);
expect(loadOverviewApp({ isSoloWindow: true }).shouldUseMobileOverview()).toBe(false);
expect(
loadOverviewApp({
loadAppSettingsFromStorage: () => ({ mobileOverviewEnabled: false }),
}).shouldUseMobileOverview()
).toBe(false);
// An unset value must read as ON: phones that already have saved settings
// from before this feature existed have no key for it.
expect(loadOverviewApp({ loadAppSettingsFromStorage: () => ({ skin: 'og' }) }).shouldUseMobileOverview()).toBe(
true
);
});
});
describe('mobile overview wiring', () => {
const html = readFileSync(resolve(PUBLIC, 'index.html'), 'utf8');
const mobileCss = readFileSync(resolve(PUBLIC, 'mobile.css'), 'utf8');
const moduleSrc = readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8');
it('speaks the same status language as the session tabs', () => {
// A session that is fine reads green on the tabs; anything else here would
// mean two meanings for one color on the same screen.
expect(mobileCss).toMatch(/\.mobile-overview-dot--idle\s*\{\s*background:\s*var\(--green\)/);
expect(mobileCss).toMatch(/\.mobile-overview-dot--working\s*\{[^}]*var\(--green\)[^}]*animation:\s*pulse/);
// Waiting-for-input blinks yellow, asked-a-question blinks red, same as
// tab-alert-idle / tab-alert-action.
expect(mobileCss).toMatch(/\.mobile-overview-row--waiting\s*\{[^}]*animation:\s*mobile-overview-blink-yellow/);
expect(mobileCss).toMatch(/\.mobile-overview-row--needs\s*\{[^}]*animation:\s*mobile-overview-blink-red/);
expect(mobileCss).toContain('@keyframes mobile-overview-blink-red');
expect(mobileCss).toContain('@keyframes mobile-overview-blink-yellow');
// The alert must survive reduced-motion as a held color, not vanish.
expect(mobileCss).toMatch(/prefers-reduced-motion[^}]*\}[\s\S]*?\.mobile-overview-row--needs/);
});
it('reuses the toolbar Run button classes instead of its own palette', () => {
// The per-backend gradient lives in styles.css keyed on
// `.btn-toolbar.btn-run.mode-<backend>` (and light skins override exactly
// those); carrying the same classes keeps both Run buttons identical.
expect(moduleSrc).toContain('btn-toolbar btn-run mode-');
expect(moduleSrc).toContain('btn-toolbar btn-run-gear mode-');
// The two button rules (not the dropdown below them) must set no color at
// all, or they would win over the mode gradient.
const buttonRules = mobileCss.match(/\.mobile-overview-run(-caret)?\s*\{[^}]*\}/g) || [];
expect(buttonRules.length).toBe(2);
for (const rule of buttonRules) {
expect(rule).not.toMatch(/\b(background|color)\s*:/);
}
});
it('ships the container hidden and loads the module', () => {
expect(html).toMatch(/<div class="mobile-overview" id="mobileOverview" hidden><\/div>/);
expect(html).toContain('<script defer src="mobile-overview.js"></script>');
});
it('never gives .mobile-overview a bare display rule', () => {
// Desktop does not load mobile.css at all, so the [hidden] attribute is the
// only thing keeping the overview off desktop. A bare
// `.mobile-overview { display: … }` rule would beat the UA [hidden] rule.
const bareDisplay = /\.mobile-overview\s*\{[^}]*display\s*:/;
expect(bareDisplay.test(mobileCss)).toBe(false);
expect(mobileCss).toContain('.mobile-overview.visible {');
});
it('styles the overview from skin tokens rather than hardcoded colors', () => {
// Skins re-point the :root tokens, so a hex literal here is a rule that
// silently stays dark on the four light skins.
const rules = mobileCss.match(/\.mobile-overview[^{}]*\{[^}]*\}/g) || [];
expect(rules.length).toBeGreaterThan(10);
const hardcoded = rules.flatMap((rule) => rule.match(/:\s*#[0-9a-f]{3,8}\b/gi) || []);
expect(hardcoded).toEqual([]);
});
});
describe('mobile overview run picker (CLI availability gating)', () => {
function modeButtons(menu: any): string[] {
return menu.children.filter((c: any) => c.dataset.moAction === 'run-mode').map((c: any) => c.dataset.moMode);
}
// #201 gated the toolbar's #runModeMenu on isCliAvailable(); this phone-only
// picker (MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu) is a
// separate, hardcoded duplicate of that menu rather than a shared render, so
// it silently offered every backend regardless of what the server reported.
it('hides run modes the server reports as unavailable, keeps shell always', () => {
const app = loadOverviewApp({
runMode: 'claude',
isCliAvailable: (tool: string) => tool === 'claude',
});
const menu = app._buildMobileOverviewRunMenu();
expect(modeButtons(menu)).toEqual(['claude', 'shell']);
});
it('shows every mode when every CLI is available', () => {
const app = loadOverviewApp({
runMode: 'claude',
isCliAvailable: () => true,
});
const menu = app._buildMobileOverviewRunMenu();
expect(modeButtons(menu)).toEqual(['claude', 'opencode', 'codex', 'gemini', 'antigravity', 'shell']);
});
it('gates every mode the picker actually offers', () => {
// Catches a new backend being added to MOBILE_OVERVIEW_RUN_MODES without
// being gated — the same class of bug that let this list drift from the
// toolbar menu's gating in the first place.
const src = readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8');
const modesBlock = src.slice(
src.indexOf('const MOBILE_OVERVIEW_RUN_MODES'),
src.indexOf('];', src.indexOf('const MOBILE_OVERVIEW_RUN_MODES')) + 2
);
const offered = [...modesBlock.matchAll(/mode: '([^']+)'/g)].map((m) => m[1]);
expect(offered).toContain('antigravity');
const fn = src.slice(src.indexOf('_buildMobileOverviewRunMenu() {'));
const gate = fn.slice(0, fn.indexOf('const header'));
expect(gate).toContain('isCliAvailable');
});
});
+140
View File
@@ -323,6 +323,146 @@ describe('Virtual Keyboard', () => {
expect(mainPadding).toBe('');
});
it('coalesces keyboard animation frames into one final terminal fit', async () => {
const result = await page.evaluate(async () => {
// `app.terminal` and `app.fitAddon` are only assigned by initTerminal(),
// which needs a selected session this harness never creates. Both are
// null at rest, and the settle callback returns early on a falsy
// terminal — so without stand-ins this test cannot reach the behavior
// it asserts. Install the minimum surface the callback touches.
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
const originalScrollToBottom = app.terminal.scrollToBottom.bind(app.terminal);
let fits = 0;
let resizes = 0;
let bottomRestores = 0;
app.fitAddon.fit = () => {
fits++;
};
KeyboardHandler._sendTerminalResize = () => {
resizes++;
};
app.terminal.scrollToBottom = () => {
bottomRestores++;
};
KeyboardHandler._scheduleViewportSettle({ scrollToBottom: true });
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler._scheduleViewportSettle();
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler._scheduleViewportSettle();
await new Promise((resolve) => setTimeout(resolve, 50));
const beforeFinalSettle = { fits, resizes, bottomRestores };
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS));
const afterFinalSettle = { fits, resizes, bottomRestores };
app.fitAddon.fit = originalFit;
KeyboardHandler._sendTerminalResize = originalSendResize;
app.terminal.scrollToBottom = originalScrollToBottom;
if (!hadFitAddon) app.fitAddon = null;
if (!hadTerminal) app.terminal = null;
return { beforeFinalSettle, afterFinalSettle };
});
expect(result.beforeFinalSettle).toEqual({ fits: 0, resizes: 0, bottomRestores: 0 });
expect(result.afterFinalSettle).toEqual({ fits: 1, resizes: 1, bottomRestores: 1 });
});
// Behavioral counterpart to the test above, driven through the PUBLIC entry
// point rather than the internal scheduler. Before this change each
// onKeyboardShow armed its own uncoalesced 150ms setTimeout, so a keyboard
// animation that reports several viewport steps refit the terminal once per
// step — the visible symptom being repeated reflow while the keyboard slides
// up. This asserts the observable outcome (one refit for a burst) and so
// fails on master by COUNT, not by a missing method.
it('refits once for a burst of keyboard viewport steps', async () => {
const counts = await page.evaluate(async () => {
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
let fits = 0;
app.fitAddon.fit = () => {
fits++;
};
KeyboardHandler._sendTerminalResize = () => {};
// Three viewport steps in quick succession, as a keyboard animation
// produces on a real device.
KeyboardHandler.onKeyboardShow();
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler.onKeyboardShow();
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler.onKeyboardShow();
// Well past both the coalescing window and master's fixed 150ms timer.
await new Promise((resolve) => setTimeout(resolve, 400));
app.fitAddon.fit = originalFit;
KeyboardHandler._sendTerminalResize = originalSendResize;
if (!hadFitAddon) app.fitAddon = null;
if (!hadTerminal) app.terminal = null;
return fits;
});
// Coalesced: one refit for the whole burst. Master fires one per step.
expect(counts).toBe(1);
});
// A viewport resize with NO pending show/hide transition must not arm settle
// work of its own: keyboard detection can miss a fine-grained OS animation
// entirely (sub-150px steps with the baseline chasing the animation), and a
// fit against that mid-animation, uncompensated layout resizes the PTY to
// transient dims. The resulting SIGWINCH thrash duplicates prompts and
// garbles the transcript. Wiggles may only push a pending settle back.
it('does not refit on viewport wiggles without a keyboard transition', async () => {
const result = await page.evaluate(async () => {
const hadTerminal = app.terminal !== null && app.terminal !== undefined;
const hadFitAddon = app.fitAddon !== null && app.fitAddon !== undefined;
if (!hadTerminal) app.terminal = { scrollToBottom() {} };
if (!hadFitAddon) app.fitAddon = { fit() {}, proposeDimensions: () => null };
const originalFit = app.fitAddon.fit.bind(app.fitAddon);
const originalSendResize = KeyboardHandler._sendTerminalResize.bind(KeyboardHandler);
let fits = 0;
app.fitAddon.fit = () => {
fits++;
};
KeyboardHandler._sendTerminalResize = () => {};
// Wiggle only: nothing pending, so nothing may fire.
KeyboardHandler._deferViewportSettle();
KeyboardHandler._deferViewportSettle();
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS + 80));
const wiggleOnly = fits;
// A real transition arms the work; a following wiggle defers it but the
// settle still fires exactly once.
KeyboardHandler._scheduleViewportSettle({ scrollToBottom: true });
await new Promise((resolve) => setTimeout(resolve, 30));
KeyboardHandler._deferViewportSettle();
await new Promise((resolve) => setTimeout(resolve, KeyboardHandler.VIEWPORT_SETTLE_MS + 80));
const afterTransition = fits;
app.fitAddon.fit = originalFit;
KeyboardHandler._sendTerminalResize = originalSendResize;
if (!hadFitAddon) app.fitAddon = null;
if (!hadTerminal) app.terminal = null;
return { wiggleOnly, afterTransition };
});
expect(result.wiggleOnly).toBe(0);
expect(result.afterTransition).toBe(1);
});
it('accessory bar has the simple-mode action buttons', async () => {
const actions = await page.evaluate(() => {
return Array.from(document.querySelectorAll('.keyboard-accessory-bar [data-action]')).map(
+38 -5
View File
@@ -4,6 +4,7 @@
*/
import { EventEmitter } from 'node:events';
import { vi } from 'vitest';
import type { SessionStatus } from '../../src/types.js';
/**
* Enhanced mock session for testing RespawnController.
@@ -12,13 +13,22 @@ import { vi } from 'vitest';
export class MockSession extends EventEmitter {
id: string;
workingDir: string = '/tmp/test-workdir';
status: 'idle' | 'working' = 'idle';
pid: number = 12345;
/**
* The REAL union, deliberately. This used to be `'idle' | 'working'`, and
* `'working'` is not a `SessionStatus` at all — so `signalForStatus()` fell to its
* `default: null` branch in every route test and the busy / stopped / error halves
* of the immediate-resolve mapping had zero coverage while appearing to be tested.
*/
status: SessionStatus = 'idle';
/** `null` once the PTY is gone (or before it has ever started) — see `pid` in Session. */
pid: number | null = 12345;
isWorking: boolean = false;
private _activeChildProcesses: { pid: number; command: string }[] = [];
ralphTracker: null = null;
writeBuffer: string[] = [];
terminalBuffer: string = '';
/** Mirrors Session.lastSubmitAt — the response viewer credits history entries by it. */
lastSubmitAt: number = 0;
private _muxName: string | null = null;
@@ -28,13 +38,22 @@ export class MockSession extends EventEmitter {
this._muxName = `codeman-test-${id.slice(0, 8)}`;
}
/** Direct PTY write (used by session.write()) */
write(data: string): void {
/**
* Set to simulate a session whose PTY is gone: both write paths report failure,
* which is the state in which input used to disappear silently.
*/
failWrites = false;
/** Direct PTY write (used by session.write()). Mirrors the real boolean return. */
write(data: string): boolean {
if (this.failWrites) return false;
this.writeBuffer.push(data);
return true;
}
/** Write via mux (used by respawn controller) */
async writeViaMux(data: string): Promise<boolean> {
if (this.failWrites) return false;
this.writeBuffer.push(data);
return true;
}
@@ -42,6 +61,10 @@ export class MockSession extends EventEmitter {
/** Exactly-once input dedup — mirrors Session.shouldApplyInput so route tests
* exercising the reliable-delivery path behave like production. */
private _appliedInputSeq = new Map<string, number>();
forgetInputSeq(clientId: string, seq: number): void {
if (this._appliedInputSeq.get(clientId) === seq) this._appliedInputSeq.set(clientId, seq - 1);
}
shouldApplyInput(clientId: string, seq: number): boolean {
const last = this._appliedInputSeq.get(clientId);
if (last !== undefined && seq <= last) return false;
@@ -98,7 +121,9 @@ export class MockSession extends EventEmitter {
/** Simulate working state with spinner */
simulateWorking(text: string = 'Thinking'): void {
this.simulateTerminalOutput(`${text}... \u280b`);
this.status = 'working';
// 'busy' is what the real Session sets while a turn is in flight; the old
// 'working' here was the event name, not a status value.
this.status = 'busy';
this.emit('working');
}
@@ -167,6 +192,14 @@ export class MockSession extends EventEmitter {
return this._muxName;
}
/**
* Mirrors `Session.usesMux`. True by default because that is the normal
* configuration, and it is what makes a route's pane-liveness probe reachable:
* `session.pid` is the tmux ATTACH CLIENT, so a mux-backed session's worker can be
* dead while `pid` is still a live number.
*/
usesMux: boolean = true;
/** Check for active child processes (mock returns configurable list) */
getActiveChildProcesses(): { pid: number; command: string }[] {
return this._activeChildProcesses;
+215
View File
@@ -0,0 +1,215 @@
/**
* Regression tests for the node-pty spawn-helper repair (issues #6, #204).
*
* node-pty@1.1.0 publishes `prebuilds/darwin-<arch>/spawn-helper` with mode 0644,
* so on macOS every PTY spawn dies with `posix_spawnp failed.`. These tests pin
* the two things the old fix got wrong: it looked ONLY in `build/Release` (which
* does not exist on macOS, where the prebuilt binary is used), and it never
* checked whether the repair actually worked.
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import { chmodSync, mkdirSync, mkdtempSync, rmSync, statSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import {
isSpawnHelperFailure,
listSpawnHelpers,
repairSpawnHelperPermissions,
spawnPtyWithHelperRepair,
resetSpawnHelperRepairState,
SPAWN_HELPER_FIX_HINT,
} from '../src/utils/node-pty-repair.js';
/** Builds a fake node-pty tree; each entry is a directory that gets a spawn-helper. */
function makeFakePtyDir(helpers: Array<{ dir: string; mode: number }>): string {
const root = mkdtempSync(join(tmpdir(), 'codeman-node-pty-'));
for (const { dir, mode } of helpers) {
const full = join(root, dir);
mkdirSync(full, { recursive: true });
const helper = join(full, 'spawn-helper');
writeFileSync(helper, '#!/bin/sh\nexit 0\n');
chmodSync(helper, mode);
}
return root;
}
function modeOf(path: string): number {
return statSync(path).mode & 0o777;
}
describe('node-pty spawn-helper repair', () => {
const created: string[] = [];
beforeEach(() => resetSpawnHelperRepairState());
afterEach(() => {
for (const dir of created.splice(0)) rmSync(dir, { recursive: true, force: true });
});
function fixture(helpers: Array<{ dir: string; mode: number }>): string {
const root = makeFakePtyDir(helpers);
created.push(root);
return root;
}
describe('isSpawnHelperFailure', () => {
it('matches the native error node-pty throws on macOS', () => {
expect(isSpawnHelperFailure(new Error('posix_spawnp failed.'))).toBe(true);
});
it('matches errors that name the helper directly', () => {
expect(isSpawnHelperFailure(new Error('ENOENT: no such file, spawn-helper'))).toBe(true);
});
it('ignores unrelated spawn failures', () => {
expect(isSpawnHelperFailure(new Error('cwd does not exist'))).toBe(false);
expect(isSpawnHelperFailure(undefined)).toBe(false);
});
});
describe('listSpawnHelpers', () => {
it('finds the prebuilt helper, which is the ONLY one that exists on macOS', () => {
// A stock macOS install has no build/ directory at all: node-pty ships a
// darwin prebuild, so node-gyp never runs.
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
expect(listSpawnHelpers(root)).toEqual([join(root, 'prebuilds', 'darwin-arm64', 'spawn-helper')]);
});
it('finds helpers across build/Release, build/Debug and every prebuilds arch', () => {
const root = fixture([
{ dir: 'build/Release', mode: 0o755 },
{ dir: 'build/Debug', mode: 0o644 },
{ dir: 'prebuilds/darwin-arm64', mode: 0o644 },
{ dir: 'prebuilds/darwin-x64', mode: 0o644 },
]);
expect(listSpawnHelpers(root).sort()).toEqual(
[
join(root, 'build', 'Release', 'spawn-helper'),
join(root, 'build', 'Debug', 'spawn-helper'),
join(root, 'prebuilds', 'darwin-arm64', 'spawn-helper'),
join(root, 'prebuilds', 'darwin-x64', 'spawn-helper'),
].sort()
);
});
it('returns nothing for a Linux install, which has no spawn-helper at all', () => {
const root = fixture([]);
mkdirSync(join(root, 'build', 'Release'), { recursive: true });
writeFileSync(join(root, 'build', 'Release', 'pty.node'), 'stub');
expect(listSpawnHelpers(root)).toEqual([]);
});
});
describe('repairSpawnHelperPermissions', () => {
it('adds the execute bit to the 0644 prebuilt helper', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
const helper = join(root, 'prebuilds', 'darwin-arm64', 'spawn-helper');
const repaired = repairSpawnHelperPermissions(root);
expect(repaired).toEqual([helper]);
expect(modeOf(helper) & 0o111).toBe(0o111);
});
it('is a no-op on an already-executable helper', () => {
const root = fixture([{ dir: 'build/Release', mode: 0o755 }]);
expect(repairSpawnHelperPermissions(root)).toEqual([]);
expect(modeOf(join(root, 'build', 'Release', 'spawn-helper'))).toBe(0o755);
});
it('repairs every copy, not just the first one found', () => {
const root = fixture([
{ dir: 'prebuilds/darwin-arm64', mode: 0o644 },
{ dir: 'prebuilds/darwin-x64', mode: 0o644 },
]);
expect(repairSpawnHelperPermissions(root)).toHaveLength(2);
for (const arch of ['darwin-arm64', 'darwin-x64']) {
expect(modeOf(join(root, 'prebuilds', arch, 'spawn-helper')) & 0o111).toBe(0o111);
}
});
it('preserves the non-execute permission bits it was given', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o640 }]);
const helper = join(root, 'prebuilds', 'darwin-arm64', 'spawn-helper');
repairSpawnHelperPermissions(root);
expect(modeOf(helper)).toBe(0o755 | 0o640);
});
it('returns nothing when node-pty has no helper to repair', () => {
expect(repairSpawnHelperPermissions(fixture([]))).toEqual([]);
});
});
describe('spawnPtyWithHelperRepair', () => {
it('passes the spawn result straight through when nothing is wrong', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o755 }]);
expect(spawnPtyWithHelperRepair(() => 'pty', root)).toBe('pty');
});
it('repairs and retries once after a posix_spawnp failure', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
const helper = join(root, 'prebuilds', 'darwin-arm64', 'spawn-helper');
let attempts = 0;
const result = spawnPtyWithHelperRepair(() => {
attempts++;
// Mirror the real failure: node-pty only throws while the helper is 0644.
if ((modeOf(helper) & 0o111) !== 0o111) throw new Error('posix_spawnp failed.');
return 'pty';
}, root);
expect(result).toBe('pty');
expect(attempts).toBe(2);
expect(modeOf(helper) & 0o111).toBe(0o111);
});
it('rethrows unrelated errors untouched, without chmodding anything', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
const helper = join(root, 'prebuilds', 'darwin-arm64', 'spawn-helper');
expect(() =>
spawnPtyWithHelperRepair(() => {
throw new Error('cwd does not exist');
}, root)
).toThrow('cwd does not exist');
expect(modeOf(helper)).toBe(0o644);
});
it('surfaces the manual fix command when the retry still fails', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
expect(() =>
spawnPtyWithHelperRepair(() => {
throw new Error('posix_spawnp failed.');
}, root)
).toThrow(SPAWN_HELPER_FIX_HINT);
});
it('surfaces the fix command when there is no helper to repair', () => {
const root = fixture([]);
expect(() =>
spawnPtyWithHelperRepair(() => {
throw new Error('posix_spawnp failed.');
}, root)
).toThrow(SPAWN_HELPER_FIX_HINT);
});
it('does not chmod-storm: only the first failure triggers a repair attempt', () => {
const root = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
const boom = () => {
throw new Error('posix_spawnp failed.');
};
expect(() => spawnPtyWithHelperRepair(boom, root)).toThrow(SPAWN_HELPER_FIX_HINT);
// Second call: repair already attempted, so it fails fast with the hint and
// never re-walks the tree.
const untouched = fixture([{ dir: 'prebuilds/darwin-arm64', mode: 0o644 }]);
expect(() => spawnPtyWithHelperRepair(boom, untouched)).toThrow(SPAWN_HELPER_FIX_HINT);
expect(modeOf(join(untouched, 'prebuilds', 'darwin-arm64', 'spawn-helper'))).toBe(0o644);
});
});
});

Some files were not shown because too many files have changed in this diff Show More