Compare commits

...
Author SHA1 Message Date
Codeman maintainer cc163792e5 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:43:47 +02:00
Ark0N 6f1ff17ccc Merge pull request #223 from Ark0N/fix/scrollback-shell-alt-screen
fix: terminal scrollback overhaul for shell and CLI sessions (#205)
2026-08-07 13:42:47 +02:00
Codeman maintainer f262b8cb69 feat(terminal): gentler glide start and fractional wheel accumulation
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:36:27 +02:00
Codeman maintainer 5f2b491d99 feat(terminal): ease-out smooth scrolling for the local wheel path
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:30:34 +02:00
Codeman maintainer c067167dbc fix(terminal): take the wheel in capture phase; xterm's scroller is deaf after reset
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.

The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.

Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:03:17 +02:00
Codeman maintainer ad2ca9b575 docs: record the #205 scrollback mechanisms and the shipped fix plan
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:42 +02:00
Codeman maintainer a7a1cef3d6 fix(session): probe the Claude CLI version over ssh for remote sessions
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:41 +02:00
Codeman maintainer a1d7ec02e9 fix(terminal): forward touch scrolls to the CLI transcript on mobile
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:25 +02:00
Codeman maintainer dfa43928af docs: record the scrollback analysis and its measurements for #205 2026-08-07 04:33:15 +02:00
Codeman maintainer adbb74cd5a fix(terminal): keep the CLI's input box pinned when scrolling with the wheel
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.

_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:

- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
  history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
  REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
  furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
  non-empty row, so any session with trailing blank rows was left off-bottom
  and every later wheel event went local without the user ever scrolling.

Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.

Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
2026-08-07 04:27:16 +02:00
Codeman maintainer eb8d11ffc3 fix(terminal): restore shell scrollback, recover history lost to tmux repaints
Four fixes for the scrollback reports in #205 (plus its follow-up comment).

1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
   ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
   (\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
   to claude/codex/gemini, so it reached the browser verbatim. In the alternate
   buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
   and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
   which readline receives as shell history navigation. Both reported symptoms,
   one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
   modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
   own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
   the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
   never forwards a pane's alt-screen toggles to its client, it repaints;
   captured from a real attach, vim/less/htop emit zero.

2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
   onto tmux's history, and tmux repaints the pane rectangle instead of emitting
   linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
   scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
   same 60 lines emitted slowly added all 60. Scrolling up at the top now
   re-pulls the full tmux scrollback and holds the user's place. Verified
   end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.

3. The full-scrollback replay was gated on a single "first load after page load"
   flag, which whichever session auto-selected consumed, so every other tab
   started with one visible frame. Now tracked per session.

4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
   per notch) scrolled one line where Chrome scrolls four or five, and capped
   the forwarded SGR report at one tick. Line and page deltas are now converted,
   and a pure horizontal swipe no longer falls through to a phantom -1.

Analysis and measurements: docs/scrollback-issues-analysis.md
2026-08-07 04:06:54 +02:00
Codeman maintainer d41f28bc14 docs(docker): warn that a plain agent-image rebuild keeps stale CLIs
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.

That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.

Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.

No changeset: docs-only, rides the next release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 09:02:28 +02:00
Codeman maintainer 322f21ef9f docs(extending): scope the "no sandbox" claim, point at Docker cases
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:

- Integration code cannot be sandboxed by Codeman because Codeman never
  launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
  documented isolation story.

Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:46:22 +02:00
Codeman maintainer c2d973cb2d docs: link the integration guide from both READMEs, fix three inaccuracies
Adds a pointer to docs/extending-codeman.md at the end of the API section in
README.md and README.zh-CN.md, so the guide is reachable from where people
read about endpoints rather than only from CLAUDE.md.

Reading the README's programmatic guide alongside the new page surfaced three
errors in it, all now fixed:

- POST /api/sessions/:id/input takes `useMux`, not `useScreen`. The latter is
  a legacy name that no longer appears in the schema.
- The page told integrators to send `\r` to submit. With `useMux: true` the
  server delivers text and Enter as two separate writes (writeViaMux does
  send-keys -l then send-keys Enter), so appending `\r` is wrong.
- "Unwrap the envelope" was incomplete: a few legacy GETs put the payload at
  the top level, so the advice is now `body.data ?? body`.

Also cross-references the README's programmatic guide, which covers the
in-session case (CODEMAN_MUX, CODEMAN_API_URL, CODEMAN_SESSION_ID,
CODEMAN_HOOK_SECRET_FILE) that the new page deliberately does not duplicate,
and documents the optional clientId/seq exactly-once fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:34:58 +02:00
Codeman maintainer 84e31c0ee1 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:31:02 +02:00
Codeman maintainer 0d0b772619 feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.

Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
  mode 'antigravity' died on command-not-found. agy is not on npm, so it
  gets its own installer step. --dir /usr/local/bin is load-bearing: the
  default $HOME/.local/bin resolves to root's home at build time and is
  unreachable by the `agent` user the container runs as. Verified inside
  codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
  ~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
  present like the other CLI buttons, with a cyan identity matching the
  toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
  it as a satisfying AI CLI, and recommends it over Gemini in the install
  hints. Detection only, no new auto-install path.

Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
  opencode/codex/gemini when the code has included antigravity for a
  while, said "all three modes", and omitted ANTIGRAVITY_ from the env
  prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
  remote-sessions' RemoteCommandMode were all stale.

Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.

test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.

Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.

isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:22:29 +02:00
Ark0N bfff20a093 Merge pull request #216 from shenlvkang-collab/fix/response-viewer-brief-format
fix(web): align brief Response Viewer formatting
2026-08-06 07:22:13 +02:00
codeman-local b982c5d0e0 fix(web): align brief response viewer formatting 2026-08-06 10:16:15 +08:00
Codeman maintainer f50c922240 docs: add extending-codeman.md, the third-party integration guide
Codeman has no plugin runtime by design: running third-party code inside
the process that spawns agents, on a server people expose over a tunnel,
would trade away the security posture that is a reason to use it. But it
already has four extension seams that work from any language with nothing
installed, and they were undocumented.

Documents web tabs (render your own UI as a tab), the SSE event channel
(react when an agent needs you), the HTTP API plus the codeman CLI (drive
it from a script), and hook events. Every endpoint, schema field, event
name and header in the page was read from source and then verified against
a running instance, including the localhost-only CORS behavior and the SSE
framing the example depends on.

Also corrects a stale line in CLAUDE.md: it claimed the HTTP/SSE API was
internal/unstable, which contradicts docs/versioning-policy.md, where the
API under /api/v1 was finalized as part of the stable surface for the 1.0
cut. No new stability commitment is made here; the page makes an existing
one discoverable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 02:03:39 +02:00
Codeman maintainer de5b048c3f docs(vm): VM cases plan + Apple virtualization stack reference
Two design/reference docs for the planned native-macOS VM isolation tier
("VM cases"), a location overlay on cases in the same shape as Docker and
remote-SSH cases, never a sixth SessionMode. Nothing is implemented; both
docs are marked PLANNED and are blocked on macOS 27 GA.

- vm-cases-plan.md: the Codeman-side design and phased plan. Swift helper
  CLI, DiskImageKit base + per-case overlay, sessions riding the existing
  remote-SSH machinery, VirtioFS workspace at the same absolute path, and
  seeded credentials, each mirroring an established Docker-cases rule.

- vm-subsystem-apple-stack.md: what the Apple stack actually provides,
  measured on the 27 beta rather than inferred from the WWDC session. Of
  note: the 2-concurrent-macOS-VM cap is a kernel quota (refused at 39%
  free RAM, so more hardware does not help), DiskImageKit has no flatten
  API so exports must ship the layer chain, and a macOS guest renders
  nothing without an attached view in an unlocked host session.

No credentials, hostnames, tailnet addresses or account names in either
file; every host/guest reference is a placeholder.

Also joins a table row that a stray blank line had split off into its own
malformed table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 01:28:17 +02:00
Codeman maintainer 12a5f5919e chore: version packages 2026-08-05 22:36:51 +02:00
Codeman maintainer ecd3f3f32a harden(history): exclude automated transcripts by SDK shape, not by "not cli"
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.

Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.

Test fails against the pre-fix line and passes after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 21:47:06 +02:00
Ark0N c19d884a51 Merge pull request #215 from timkjr/fix/past-sessions-history-quality
fix(history): three Past Sessions data-quality bugs (automated-session noise, cross-contaminated previews, blank restart-heavy rows)
2026-08-05 21:44:47 +02:00
Ark0N 22e77a1827 Merge pull request #214 from timkjr/fix/mobile-overview-run-gating
fix(mobile): gate the phone overview's run picker on CLI availability
2026-08-05 21:44:42 +02:00
Ark0N b641560040 Merge pull request #203 from shenlvkang-collab/contrib/claude-viewer-session-pin
fix(web): pin the Claude response viewer to the pane's own conversation
2026-08-05 21:44:37 +02:00
timkjr 8300c15cbd fix(history): entrypoint detection was first-field-wins, plus a two-tier head read
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).

Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 09f5f28017 docs(test): correct an overclaiming comment in the tail-fallback regression test
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 251706be3b harden: scope entrypoint detection to message lines, add fallback coverage
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:

1. extractTranscriptEntrypoint() scanned any line containing the
   substring "entrypoint", not specifically the first "type":"user"/
   "type":"assistant" message line (unlike its sibling
   extractFirstUserPrompt, which does scope to type). A transcript
   that started under an older Claude Code version (no entrypoint
   field) and got resumed under a newer one mid-conversation could
   pick up the field from a much later message than the true first
   one, misattributing the session's origin. Scoped it to match.

2. Added a regression test proving the tail-read fallback still
   engages correctly when bookkeeping accumulation exceeds even the
   new 128KB head window, not just the 16KB it previously blanked at.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 18b473f0e4 fix(history): raise the transcript head-read window to fit restart bookkeeping
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).

Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 a2aed38073 fix(unified-sessions): stop the firstPrompt workingDir backfill from cross-contaminating history rows
COD-140's backfill was meant to cover live/persisted rows whose Codeman
id doesn't match an on-disk transcript UUID, guessing from the newest
transcript in the same workingDir as a last resort. It was also firing
for pure history rows whose OWN transcript scan already ran (and
genuinely found nothing, e.g. an oversized first message) -- those got
silently backfilled with the newest OTHER session's opening line from
the same directory. Not a blank row, but actively wrong: old sessions
displayed a completely unrelated (often today's live) conversation's
first prompt as if it were their own.

Skip the workingDir guess for any item that already has its own
'history' source -- it already had a real, direct attempt. Rows with
no history source at all (their transcript isn't linked/scanned under
their own id yet) still get the guess, matching the original intent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 e888c65c52 fix(history): exclude non-interactive (SDK-driven) transcripts from Past Sessions
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.

Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 1ea39de650 fix(mobile): gate the phone overview's run picker on CLI availability
MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu is a separate,
hardcoded duplicate of the toolbar's #runModeMenu (mobile-overview.js
is a newer feature that mirrors the toolbar menu's look/behavior
rather than reusing its render), so it never picked up #201's
isCliAvailable() gating and offered every backend regardless of what
the server actually has installed.

Gate it the same way: skip an entry unless isCliAvailable(mode),
shell always exempt. Added functional + static regression tests
mirroring the toolbar menu's own test pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:15 -05:00
Codeman maintainer e2a644997e chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 09:01:45 +02:00
Ark0N cd5a101626 Merge pull request #213 from Ark0N/feat/file-viewer-edit-mode
File Viewer: edit mode for text files (edit + save in the viewer)
2026-08-05 09:00:22 +02:00
Codeman maintainer 4ea781c80f feat(file-viewer): edit mode for text files (edit + save in the viewer)
Closes #212. The file-preview overlay can now edit workspace text files in
place, phone-first: agent writes a file, you review it in the viewer, tweak
two lines, save, tell the agent to continue.

Backend (file-routes.ts, policy in src/config/file-editing.ts):
- GET file-content?edit=1: read-for-edit that never truncates (a truncated
  buffer must never become an edit buffer), 512KB cap (413 over it), and
  returns the sha256 hash + detected EOL the client echoes back on save.
- PUT /api/sessions/:id/file-content: edit-in-place only, with no O_CREAT
  anywhere in the handler. Confinement matches the read path (realpath +
  workspace boundary + ownership via findSessionOrFail), plus sensitive-path
  and attachment-guard blocklists, a .git subtree deny, and an extension
  allowlist (svg and env deliberately excluded). Optimistic concurrency via
  baseHash: mismatch is a 409 unless force. Writes are wx-temp + fchmod +
  fsync + rename, closing the validate-then-write TOCTOU window.
- Corruption guards: NUL sniff + UTF-8 round-trip compare (refuses binary
  and latin-1), and server-side EOL re-application so a textarea's LF
  normalization cannot rewrite every line of a CRLF file.
- Plain reads gain an additive editable flag the UI keys the button off.

Frontend (panels-ui.js + overlay markup/styles):
- Edit button on editable text previews; textarea editor with Save/Cancel,
  dirty indicator, discard-confirm on cancel/close, and a conflict dialog
  that offers overwrite (force) when the file changed on disk mid-edit.
- Phone: full-bleed window sized by --app-height so the editor and Save bar
  track the OS keyboard; 16px editor font (iOS zoom guard); no autofocus.
- zh-CN strings for the new chrome.

Tests: pure policy unit tests plus a route suite that deliberately does NOT
mock node:fs. It runs against a real temp workspace so symlink escapes,
write-through of in-workspace symlinks, mode preservation, CRLF round-trip,
409/force, and the no-create property are exercised for real. Also verified
end to end on an isolated beta instance: 39-check curl matrix, Playwright
desktop flow (real clicks and typing, bytes asserted on disk, live conflict
with an external rewrite), and a 393px phone profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 08:44:47 +02:00
Codeman maintainer d9123de9eb feat(terminal): Ctrl+C copies the selection, interrupts when nothing is selected
Closes #211. Copying from the terminal only worked through the browser
context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown,
so the muscle-memory copy failed silently and read as "no copy-paste at all".

With a selection, Ctrl+C now copies it, toasts, clears the selection and
sends nothing to the PTY. With no selection it falls through unchanged, so
the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never
falls through: an explicit copy that interrupts a running agent because the
selection happened to be empty would be a footgun.

Three details that keep the interrupt safe:

- The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the
  no-selection path returns true WITHOUT preventDefault. xterm calls the
  custom handler before its own cancel(), so returning false alone does not
  cancel the event; the copy path therefore calls preventDefault explicitly,
  or the browser would run its native copy on top of ours.
- copy-selection is a full registry entry (rebindable and disableable in App
  Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the
  same trick command-palette uses: the generic capture loop preventDefaults
  every match it dispatches, which would cost the user the interrupt key.
- The gate is keydown-only, since the custom handler also runs for keypress
  and keyup.

Copy goes through _copyText (Clipboard API, then hidden-textarea +
execCommand) rather than raw navigator.clipboard, because install.sh's LAN
option serves plain HTTP where navigator.clipboard is undefined; the
fallback steals focus, so the terminal is refocused afterwards.

Tests: test/terminal-copy-selection.test.ts pins the gate and the
SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real
key presses in chromium and asserts on the clipboard plus the bytes xterm
emitted (browser-driven, so excluded from test:ci like the other Playwright
suites). Verified manually on an isolated beta instance before landing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 02:38:57 +02:00
shenlvkang-collab ab7a703e90 fix(web): keep the viewer's conversation anchor across a Codeman restart
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.

Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.

Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
2026-08-03 21:22:33 +08:00
shenlvkang-collabandClaude Opus 5 73315bc351 fix(web): pin the Claude response viewer to the pane's own conversation
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.

Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:53:42 +08:00
66 changed files with 5232 additions and 233 deletions
+110
View File
@@ -1,5 +1,115 @@
# aicodeman
## 1.12.0
### Minor Changes
- Terminal scrollback overhaul (issue #205), fixing every reported scroll failure across shell and CLI sessions, desktop and mobile:
- Shell, OpenCode and Antigravity sessions finally have working scrollback: tmux's own client-side alternate-screen switch is stripped for tmux-backed sessions (narrow strip: alt-screen toggles only, keeping `clear`'s 3J and mouse DECSETs), so xterm stays in the normal buffer instead of a scrollback-less alt buffer where the wheel turned into shell history cycling and touch scrolling did nothing. Direct-PTY fallback sessions are untouched so fullscreen apps (vim/less/htop) keep the alt screen there.
- The wheel listener now runs in capture phase and owns the scroll: xterm's internal vscode-style viewport scroller consumed wheel events whenever local scrollback existed (and goes deaf entirely after a tab switch or replay resets the terminal), which silently killed wheel forwarding, made scrolling break after reload/tab switches, and let the CLI's input box scroll away. Local scrolling goes through buffer-level scrollLines and keeps working after resets; mouse-tracking apps and alternate-buffer sessions are passed through untouched.
- Wheel AND touch scrolling now forward to the CLI's own transcript for Codex and Claude 2.1.187+, at any scroll position (the viewport snaps home first), so the input box stays pinned on desktop and phones alike. Shift+wheel and the "Wheel scrolls local history" setting still pin local scrollback.
- Smooth scrolling: local wheel scrolling glides with an ease-out animation (fractional line accumulation, so slow trackpad drags track the finger instead of running ahead).
- Full tmux history on demand: the full-scrollback replay is now per session instead of once per page load, and scrolling up at the top of the buffer re-pulls the complete tmux history, recovering everything tmux's repaint bursts or tab switches removed from the browser's copy.
- Firefox wheel speed: wheel deltas are normalized by deltaMode (Firefox reports line units, previously read as pixels and slowed ~4x).
- Remote SSH Claude sessions now probe the CLI version over ssh (same connection options and login-shell wrapper as the real launch), so wheel forwarding works for them too instead of silently staying off.
Docs: scrollback analysis and fix plan recorded in docs/, architecture invariants updated (strip flavors, capture-phase wheel ownership, per-session full-history replay); docker agent-image rebuild warning and integration-guide link fixes from the preceding docs commits.
## 1.11.2
### Patch Changes
- Make Antigravity (`agy`) a first-class CLI everywhere, and stop presenting Gemini CLI as a consumer product now that it is enterprise-only.
Antigravity was already wired into the session layer, schemas, run-mode menu and remote/Docker command maps, but the surfaces around it were never updated. Gemini keeps full support; Antigravity now sits beside it.
Fixes:
- **Docker cases with `mode: 'antigravity'` were broken.** `docker/agent.Dockerfile` installs its CLIs from npm, and `agy` is not an npm package, so the binary was never in the image and the container died on command-not-found. It now gets its own installer step. The `--dir /usr/local/bin` flag is load-bearing: the installer's default `$HOME/.local/bin` resolves to root's home at build time and would be unreachable by the `agent` user the container runs as. Note the binary is roughly 190MB, making it the largest layer in the image, so rebuild with `node scripts/build-agent-image.mjs` when convenient.
- **Welcome screen** gained a "Run Antigravity" action, gated on `agy` being present like the other CLI buttons, styled with the same cyan identity as the toolbar run button and run-mode dot.
- **`install.sh`** now detects `agy` (search paths mirroring `antigravity-cli-resolver.ts`), counts it as a satisfying AI CLI so an Antigravity-only box is not told it has none, and recommends it instead of Gemini in the install hints.
Documentation corrections where it had become factually wrong: `architecture-invariants.md` described `isExternalCliMode()` as opencode/codex/gemini when the code has included antigravity for some time, said "all three modes", and omitted `ANTIGRAVITY_*` from the env-prefix allowlist row; the `agentType` enum in `cron-guide.md`, `SessionMode` in `cron-discovery.md`, and `RemoteCommandMode` in `remote-sessions.md` were all stale.
Also updated both READMEs (five CLIs, Gemini marked enterprise-only), the `antigravity` npm keyword, and comment drift in eight places. Test coverage added for the new welcome button.
Antigravity stores its state under `~/.gemini/antigravity-cli/` rather than a `~/.antigravity` directory, so the existing `.gemini` Docker credential seed already covers it. That is now recorded in a code comment so no dead configuration gets added later.
- b982c5d: Keep the brief Response Viewer output inside the same message card and Markdown wrapper used by the full conversation view, so opening the viewer without clicking More preserves the same readable formatting.
## 1.11.1
### Patch Changes
- fix(history): Past Sessions data quality, and gate the phone run picker on CLI availability
**Past Sessions data quality (#215).** Three bugs in the transcript scanner behind
the Cmd+K Session Manager and the phone overview's PAST SESSIONS list:
- Automated/SDK-driven transcripts (CI review bots and other tooling, which Claude
Code stamps with a non-`cli` `entrypoint`) were listed alongside real interactive
sessions even though they were never resumable. They are now excluded. Detection
scans every entrypoint-bearing message rather than stopping at the first, so a
transcript that began under an older Claude Code build and only later picked up a
non-`cli` entrypoint is no longer wrongly hidden.
- A resumed session could show a same-directory sibling's preview text as its own.
The `workingDir` backfill in `mergeUnifiedSessions()` now only ever applies to rows
that have no history entry of their own, so it can no longer overwrite a row's real
content with another conversation's.
- Sessions restarted many times accumulated enough bookkeeping lines to push the real
first prompt past the scanner's 16KB head-read window, leaving a blank row. The read
is now two-tier: 16KB first, escalating to 128KB only when that was not enough, which
is both correct and cheaper than reading 128KB unconditionally (measured on a real
transcript tree: 36% fewer bytes read, roughly 17.5% faster than the unconditional
version). Also restores the tail-read fallback for a file whose head read failed
outright (for example `EMFILE` while scanning hundreds of files), which had been
silently dropping the session from history.
Follow-up hardening on top of the above: the automated-transcript exclusion now
blocklists the SDK entrypoint shape (`sdk`, `sdk-cli`, `sdk-py`) instead of allowlisting
the exact value `cli`. Because the check hides rows, an allowlist failed closed on any
value Claude Code has not shipped yet: a future rename of the interactive entrypoint,
or a second interactive host, would have blanked the entire Past Sessions list with
nothing in the UI to explain it. An unrecognized automated entrypoint now costs a few
noisy rows instead, which is the annoyance this filter set out to fix rather than a
broken feature.
**Phone overview run picker (#214).** The "C" logo home screen's Run picker listed all
six backends regardless of what was installed, so tapping an uninstalled one produced a
failed launch instead of the entry simply not being offered. It is now gated on
`isCliAvailable()` exactly like the desktop toolbar's run-mode dropdown (shell exempt,
since it has no external CLI dependency and keeps the menu from ever being empty). The
picker is a hardcoded duplicate of the toolbar menu rather than a shared render, which
is why it never picked up the earlier gating work; a test now asserts that every mode
the picker offers is gated, so a newly added backend cannot silently drift again.
- 73315bc: fix(web): stop the Claude response viewer from following another session's conversation
The viewer re-derived a pane's live conversation by taking the newest
`~/.claude/history.jsonl` entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` run in the user's own terminal, so the eye followed whichever of those
was typed into last — and the adoption was written back to the session, so the
mispin persisted. Entries are now credited to a pane only when they land within
10s of that pane's own Enter and no other pane on the cwd submitted closer, the
same last-submit correlation the Codex locator already uses.
That correlation also has to survive a restart. `start()` resets
`claudeSessionId` to the launch id even when re-attaching to a mux session whose
CLI has since moved on via `/clear`, so a recovered pane pointed the viewer at
its pre-`/clear` transcript — and with the anchor itself living only in memory,
nothing corrected it until the user happened to type again. `lastSubmitAt` is
now persisted in `SessionState` and restored on boot recovery, so the viewer
re-derives the live conversation on its first poll.
## 1.11.0
### Minor Changes
- Two user-facing features since 1.10.0.
**Terminal: Ctrl+C copies the selection, interrupts when nothing is selected** (#211). Copying from the terminal previously worked only through the browser context menu: xterm turns Ctrl+C into 0x03 and cancels the keydown, so the muscle-memory copy failed silently and read as "no copy-paste at all". With a selection, Ctrl+C now copies it, shows the "Copied to clipboard" toast, clears the selection and sends nothing to the PTY; with no selection it falls through unchanged, so the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never interrupts. The shortcut is a normal registry entry (`copy-selection`), so it can be rebound or disabled in App Settings, and disabling it restores plain always-interrupt Ctrl+C. Copy goes through the Clipboard API with a hidden-textarea fallback, so it also works on plain-HTTP LAN installs.
**File Viewer: edit mode for text files** (#212). The file-preview overlay can now edit workspace text files in place, phone-first: `GET /api/sessions/:id/file-content?edit=1` reads for edit without the 500-line preview truncation (saving a truncated buffer would silently delete the rest) and returns a sha256 hash plus the detected EOL; `PUT /api/sessions/:id/file-content` saves. Edit-in-place only: there is no O_CREAT anywhere in the handler, so "never create, never delete" is structural. Confinement inherits the read path (realpath plus workspace boundary, ownership scoping) and adds sensitive-path and attachment-guard blocklists, a `.git/` subtree deny, and an extension allowlist (`svg` and `env` deliberately excluded). Optimistic concurrency is by content hash, so a file changed on disk mid-edit returns 409 with an overwrite option rather than clobbering. Writes are atomic (`wx` temp, fchmod, fsync, rename) which closes the validate-then-write TOCTOU window and cannot follow a pre-existing symlink. Binary and latin-1 content are refused via a NUL sniff plus a UTF-8 round-trip compare, and EOL is re-applied server-side so a textarea's LF normalization cannot turn a two-line edit of a CRLF file into a whole-file diff.
## 1.10.0
### Minor Changes
+10 -6
View File
@@ -47,7 +47,7 @@ The production server caches static files for 1 year, `immutable` (`maxAge: '1y'
## COM Shorthand (Deployment)
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI + documented env vars are public; the HTTP/SSE API, on-disk state, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
Uses [Semantic Versioning](https://semver.org/) (`MAJOR.MINOR.PATCH`) via `@changesets/cli`. What SemVer actually covers (the CLI, documented env vars, **and the HTTP/SSE API under `/api/v1`**: endpoint paths, response envelope, `errorCode` values and SSE event names are public/stable; on-disk state, internal TS modules, and experimental features are internal/unstable) is defined in `docs/versioning-policy.md`. Third-party integration surfaces are documented in `docs/extending-codeman.md`. Security reporting + known limitations live in `.github/SECURITY.md`.
When user says "COM":
@@ -74,7 +74,7 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.10.0 (must match `package.json`)
**Version**: 1.12.0 (must match `package.json`)
## Project Overview
@@ -102,7 +102,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Test coverage | `npm run test:coverage` |
| Dead-code sweep | `npm run knip` (config in `config/knip.json`, passed via `--config`) |
| Rebuild gesture overlay | `npm run build:gesture` (esbuild `packages/gesture-control/src/codeman/entry.ts` → `src/web/public/gesture/gesture-codeman.js`; commit the result) |
| Build the docker agent image | `node scripts/build-agent-image.mjs` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`/`--no-cache`) |
| Build the docker agent image | `node scripts/build-agent-image.mjs --no-cache` (builds `codeman/agent:base` from `docker/agent.Dockerfile`; prerequisite for Docker cases; `--engine`/`--image`). ⚠ **Always `--no-cache`** — a plain rebuild re-uses the cached `npm install -g` layer and silently keeps the CLIs frozen at their original versions, which once shipped a BROKEN codex while reporting success. See `docs/docker-cases.md` |
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
@@ -206,7 +206,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. Only the FIRST buffer load after a page load requests `full=1`; tab switches keep the cheap `?tail=` path. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the entire tmux scrollback, bounded by the configured history limit. On success the capture is returned ALONE (`source='mux-full-history'`), superseding the byte buffer so nothing duplicates. The first load of EACH session per page load requests `full=1` (`_fullHistoryLoaded` Set); tab switches keep the cheap `?tail=` path, and scrolling up at the TOP of the buffer re-pulls `full=1` on demand (cooldown-guarded — tmux repaints bursty output in place, so browser scrollback shrinks while tmux's history stays complete). → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Terminal scrollback strip + wheel/touch forwarding** (#205): codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity get a NARROW strip (alt-screen toggles only — it removes tmux's own attach-time `smcup`, which otherwise parks xterm in the scrollback-less alt buffer and turns the wheel into arrow keys). ⚠️ Gated on `useMux`: direct-PTY fallback sessions must keep the alt screen for vim/less/htop. Wheel AND touch forward to the CLI transcript for codex/claude ≥ 2.1.187 at ANY scroll position (snap-to-bottom first); Shift+wheel and the `terminalWheelLocalScrollback` setting stay local. `_wheelScrollLines()` reads `ev.deltaMode` (Firefox = LINE units). → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Self-update** (App Settings → Updates): in-app updater for git-clone installs supervised by systemd/launchd (`systemd`, `launchd`, `launchd-daemon`, else `none` → "restart manually"). The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` that outlives the restart and writes progress to `update-status.json`, which the browser polls across the connection drop. `src/web/self-update.ts` splits pure helpers (unit-tested) from IO wrappers. npm installs report as non-updatable. → [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
@@ -214,6 +216,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Filesystem path picker** (Link Existing "Browse" + the mobile keyboard's `📁 Path` key): lazy one-directory browsing via `GET /api/filesystem/browse`, with `GET /api/filesystem/preview` for the tapped file. Inserts the path **without** Enter, so the prompt is never submitted; the sibling `⌫ All` key clears only the unsent prompt and must never send the agent's `/clear`. ⚠️ This is a **second file-serving surface and inherits neither the attachment confinement nor its ownership scoping** — it allowlists Home, `CASES_DIR`, `/mnt/d` and `CODEMAN_FILE_PICKER_ROOTS`, blocks sensitive trees, and rejects symlink escapes **after** `realpath`. ⚠️ The optional `sessionId` is an ownership boundary that must be `canAccessOwned`-checked by hand (it does not go through `findSessionOrFail`), and in multi-user mode a non-admin gets only their own `userSpacePath` as a root: per-user spaces live INSIDE `homedir()`, so a `Home` root exposes every other user's workspace. Previews go through the same global conversion limiter, and Markdown/TXT/JSON are served as inert `text/plain`. → [architecture-invariants#filesystem-path-picker](docs/architecture-invariants.md#filesystem-path-picker)
**File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md`
**Ultracode / workflow-run visualization** (opt-in, default OFF): the Workflow tool writes a completion artifact only at run *end*, so live in-flight runs exist solely as transcript dirs. `workflow-run-watcher.ts` therefore synthesizes ACTIVE runs from transcripts until the completion artifact appears and supersedes them. It is **STANDALONE** and deliberately never imports or touches `subagent-watcher.ts`, despite reading the same tree. Two independent toggles: `showUltracodeAgents` (docked panel) and `ultracodeFloatingWindows` (floating windows); the watcher starts if **either** is on. → [architecture-invariants#ultracode--workflow-run-visualization](docs/architecture-invariants.md#ultracode-and-workflow-run-visualization)
**Cross-session search**: `GET /api/search` federates an in-memory search over session metadata, run-summary events, and attachment-history entries. The pure core `searchSources()` does substring matching with hard per-type caps: **no regex (so no ReDoS) and no filesystem reads (so no traversal)**. The server-private `externalPath` is never read. → [architecture-invariants#cross-session-search](docs/architecture-invariants.md#cross-session-search)
@@ -236,7 +240,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
**Phone overview home screen** (`mobile-overview.js`, phones only, per-device `mobileOverviewEnabled`, default ON): under 430px the "C" logo shows a session overview (NEEDS YOU / CURRENT SESSIONS / PAST SESSIONS) instead of the welcome overlay; tablet and desktop are unchanged. The branch lives in `showWelcome()`/`hideWelcome()` (terminal-ui.js) behind `shouldUseMobileOverview()`, which is **width-driven** (`getDeviceType() === 'mobile'`) because this is a layout decision, unlike the settings namespace which stays handheld-based. ⚠️ The container ships with the `hidden` attribute and only this module removes it: never give `.mobile-overview` a bare `display` rule, since desktop does not load `mobile.css` (`media="(max-width: 1023px)"`) and would then render it unstyled. Live re-renders ride on the tail of `_renderSessionTabsImmediate()` (every state change it needs already funnels there); PAST rows come from one `_fetchUnifiedSessions(60)` per home-screen visit and resume through the shared `resumeHistorySession()`, so they behave exactly like the welcome screen's Resume list. ⚠️ Two things must stay in lockstep with surfaces outside this module, because divergence reads as a bug rather than a style: the split Run button carries the **toolbar's own classes** (`btn-toolbar btn-run mode-<backend>` / `btn-run-gear`) so the per-backend gradient and the light-skin overrides apply unchanged (mobile.css must therefore set no `background`/`color` on it), and row status uses the **session-tab language** (green dot when fine, `pulse` while working, yellow blinking row when waiting for input, red blinking row when a question is pending, mirroring `tab-alert-idle`/`tab-alert-action`). The picker mirrors the toolbar run-mode menu (`setRunMode()` + `run()`, `openWebviewFromMenu()` for saved dashboards) and deliberately omits its Recent-Sessions block, since PAST SESSIONS is that. Status pills carry `data-i18n-skip` (generic words like "idle" collide with state strings elsewhere).
**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry)
**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)** lives in that same handler: with a selection it copies, with none it must `return true` **without** `preventDefault()` or the interrupt is lost. `copyTerminalSelection` is deliberately absent from `SHORTCUT_ACTIONS` because the generic capture loop preventDefaults every match it dispatches. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry)
**Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema.
@@ -260,7 +264,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L
**Respawn presets**: `solo-work` (3s/60min), `subagent-workflow` (45s/240min), `team-lead` (90s/480min), `ralph-todo` (8s/480min), `overnight-autonomous` (10s/480min).
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active). Rebindable via the registry.
**Keyboard shortcuts**: Escape (close), Ctrl+? (shortcut overlay), Ctrl/Cmd/Alt+K (session palette), Ctrl+W (kill), Ctrl+Tab (next), Alt+[/] (prev/next tab), Alt+1-9 (switch tab), Ctrl+Shift+{/} (move tab left/right), Shift+Enter or Ctrl+Enter (newline), Ctrl+C (copy selection, else interrupt) / Ctrl+Shift+C (copy, never interrupts), Ctrl+L (clear), Ctrl+Shift+R (restore size), Ctrl+Shift+V (voice input), Ctrl/Cmd +/- (font), Shift+Wheel (local scrollback when mouse passthrough is active). Rebindable via the registry.
### Security
+14 -10
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Gemini &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -27,7 +27,7 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, four CLIs** - run [Claude Code, OpenCode, Codex, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, five CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -68,7 +68,7 @@ This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, a
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works). The installer detects whichever of the four is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the five is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -151,7 +151,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
@@ -231,7 +231,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -404,7 +404,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
@@ -426,7 +426,7 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
@@ -612,7 +612,7 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
@@ -651,6 +651,8 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Alt/Option+[` / `Alt/Option+]` | Previous / next session |
| `Alt/Option+1`-`Alt/Option+9` | Switch to tab N (physical keys, so macOS Option layouts work) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+C` | Copy selection, or interrupt when nothing is selected |
| `Ctrl+Shift+C` | Copy selection (never interrupts) |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
@@ -810,6 +812,8 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
---
## Architecture
@@ -842,7 +846,7 @@ flowchart TB
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
+12 -8
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Gemini &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -58,7 +58,7 @@ curl -fsSL https://getcodeman.com/install | bash
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可)。安装器会自动检测这四个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这五个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
@@ -141,7 +141,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
@@ -221,7 +221,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Gemini` 或 `Terminal`(普通 shell)。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
@@ -394,7 +394,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
@@ -416,7 +416,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
@@ -602,7 +602,7 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
@@ -641,6 +641,8 @@ sc -l # 列出会话
| `Alt/Option+[` / `Alt/Option+]` | 上一个 / 下一个会话 |
| `Alt/Option+1`–`Alt/Option+9` | 切换到第 N 个标签(按物理键位,macOS Option 布局也适用) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+C` | 复制选中内容;未选中时中断代理 |
| `Ctrl+Shift+C` | 复制选中内容(永不中断) |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
@@ -800,6 +802,8 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
---
## 架构
@@ -832,7 +836,7 @@ flowchart TB
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
+1
View File
@@ -24,6 +24,7 @@ export default defineConfig({
'test/inline-rename.test.ts', // browser (Playwright)
'test/opencode-resize.test.ts', // browser (Playwright)
'test/webgl-fallback.test.ts', // browser (Playwright)
'test/terminal-copy-shortcut.test.ts', // browser (Playwright)
],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
+13 -3
View File
@@ -26,8 +26,8 @@ RUN apt-get update \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The agent CLIs (all four backends Codeman supports). Pinning is left to the
# rebuild cadence (see docs/docker-cases-plan.md, user-decision 2).
# The npm-published agent CLIs. Pinning is left to the rebuild cadence (see
# docs/docker-cases-plan.md, user-decision 2).
RUN npm install -g \
@anthropic-ai/claude-code \
@openai/codex \
@@ -35,6 +35,15 @@ RUN npm install -g \
opencode-ai \
&& npm cache clean --force
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
# the installer's default target is `$HOME/.local/bin`, which at build time is
# root's home and would be unreachable by the `agent` user the container runs as.
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
&& chmod 755 /usr/local/bin/agy \
&& agy --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
@@ -50,7 +59,8 @@ ENV HOME=/home/agent
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir.)
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions \
+37 -4
View File
@@ -18,9 +18,9 @@ Implementation detail extracted from `CLAUDE.md` so that file stays small enough
## Session launch modes
### External CLI modes (OpenCode, Codex, Gemini)
### External CLI modes (OpenCode, Codex, Gemini, Antigravity)
**External CLI modes (OpenCode, Codex, Gemini)**: `isExternalCliMode()` in `session.ts` (`mode === 'opencode' || 'codex' || 'gemini'`) gates Claude-specific behavior — Ralph tracker, BashToolParser, token/CLI-info parsing, and ❯-prompt readiness detection are all skipped (these CLIs render their own TUIs; readiness = output stabilization instead). All three modes **require tmux — no direct PTY fallback** — because secrets are injected via `tmux setenv` (socket-scoped `${this.tmux()} setenv`, never on the spawn command line): OpenCode gets `OPENCODE_CONFIG_CONTENT` etc., Codex gets `OPENAI_API_KEY`/`CODEX_API_KEY`/`CODEX_HOME` (`setCodexEnvVars`), Gemini gets `GEMINI_API_KEY`/`GOOGLE_API_KEY`/`GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI` etc. (`setGeminiEnvVars`, all in `tmux-manager.ts`). Codex specifics: command built by `buildCodexCommand()` (`--model`, `resume <id>`, `--dangerously-bypass-approvals-and-sandbox` from the `codexConfig` payload / `codexDangerouslyBypassApprovals` app setting; `renderMode` is schema-coerced to `'hybrid'`, the only supported mode). Gemini specifics: command built by `buildGeminiCommand()` (`--skip-trust` always, `--approval-mode <default|auto_edit|yolo|plan>` defaulting to `yolo` for parity with Claude's `--dangerously-skip-permissions`, `--model`, `--resume` from the `geminiConfig` payload); availability via `GET /api/gemini/status` — session/quick-start routes fail with `OPERATION_FAILED` + install hint (`npm install -g @google/gemini-cli`) when missing. Codex AND Gemini export `COLORTERM=truecolor` + unset `NO_COLOR` (other modes unset `COLORTERM`); Gemini joins `isAltScreenStripMode()` (Codex/Claude/Gemini are Ink TUIs that repaint inline → strip alt-screen/`3J` so scrollback survives). Codex availability via `GET /api/codex/status`. Frontend: run-mode dropdown → `runCodex()`/`runGemini()` in `session-ui.js` ("Run CX"/"Run GM" labels), App Settings → Codex CLI tab; Respawn/Ralph options are Claude-only, so session options open on the Summary tab for external CLI sessions. ⚠️ `run*()` MUST unwrap the `{success,data}` envelope (`(await res.json()).data.available` / `data.data.sessionId`) — reading the raw shape silently breaks the run. Tests: `test/run-mode-ui.test.ts` + `test/gemini-mode.test.ts` (vm-sandbox harness, no real DOM).
**External CLI modes (OpenCode, Codex, Gemini, Antigravity)**: `isExternalCliMode()` in `session.ts` (`mode === 'opencode' || 'codex' || 'gemini' || 'antigravity'`) gates Claude-specific behavior — Ralph tracker, BashToolParser, token/CLI-info parsing, and ❯-prompt readiness detection are all skipped (these CLIs render their own TUIs; readiness = output stabilization instead). All four modes **require tmux — no direct PTY fallback** — because secrets are injected via `tmux setenv` (socket-scoped `${this.tmux()} setenv`, never on the spawn command line): OpenCode gets `OPENCODE_CONFIG_CONTENT` etc., Codex gets `OPENAI_API_KEY`/`CODEX_API_KEY`/`CODEX_HOME` (`setCodexEnvVars`), Gemini gets `GEMINI_API_KEY`/`GOOGLE_API_KEY`/`GOOGLE_CLOUD_PROJECT`/`GOOGLE_APPLICATION_CREDENTIALS`/`GOOGLE_GENAI_USE_VERTEXAI` etc. (`setGeminiEnvVars`, all in `tmux-manager.ts`). Codex specifics: command built by `buildCodexCommand()` (`--model`, `resume <id>`, `--dangerously-bypass-approvals-and-sandbox` from the `codexConfig` payload / `codexDangerouslyBypassApprovals` app setting; `renderMode` is schema-coerced to `'hybrid'`, the only supported mode). Gemini specifics: command built by `buildGeminiCommand()` (`--skip-trust` always, `--approval-mode <default|auto_edit|yolo|plan>` defaulting to `yolo` for parity with Claude's `--dangerously-skip-permissions`, `--model`, `--resume` from the `geminiConfig` payload); availability via `GET /api/gemini/status` — session/quick-start routes fail with `OPERATION_FAILED` + install hint (`npm install -g @google/gemini-cli`) when missing. Codex AND Gemini export `COLORTERM=truecolor` + unset `NO_COLOR` (other modes unset `COLORTERM`); Gemini joins `isAltScreenStripMode()` (Codex/Claude/Gemini are Ink TUIs that repaint inline → strip alt-screen/`3J` so scrollback survives). Codex availability via `GET /api/codex/status`. Antigravity specifics: command built by `buildAntigravityCommand()` (`--model`, `--conversation <id>` resume, `--dangerously-skip-permissions` from the `antigravityConfig` payload); availability via `GET /api/antigravity/status` — routes fail with `OPERATION_FAILED` + install hint (`curl -fsSL https://antigravity.google/cli/install.sh | bash`) when missing. Unlike the other three it is NOT an npm package (standalone binary, `~/.local/bin/agy`), which is why `docker/agent.Dockerfile` installs it with its own `--dir /usr/local/bin` step rather than in the `npm install -g` line, and why it does NOT join `isAltScreenStripMode()`. Frontend: run-mode dropdown → `runCodex()`/`runGemini()` in `session-ui.js` ("Run CX"/"Run GM" labels), App Settings → Codex CLI tab; Respawn/Ralph options are Claude-only, so session options open on the Summary tab for external CLI sessions. ⚠️ `run*()` MUST unwrap the `{success,data}` envelope (`(await res.json()).data.available` / `data.data.sessionId`) — reading the raw shape silently breaks the run. Tests: `test/run-mode-ui.test.ts` + `test/gemini-mode.test.ts` (vm-sandbox harness, no real DOM).
### Remote sessions over SSH
@@ -58,7 +58,15 @@ Implementation detail extracted from `CLAUDE.md` so that file stays small enough
### Full-scrollback replay
**Full-scrollback replay** (COD-164/#148): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). Only the FIRST buffer load after a page load requests `full=1` (one-shot `_initialFullBufferLoad` flag in app.js); tab switches keep the cheap `?tail=` visible-frame path. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`.
**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load OF EACH SESSION per page load requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other tab one frame of history); later switches keep the cheap `?tail=` visible-frame path. On top of that, scrolling up while already at the TOP of the buffer re-pulls `full=1` on demand (`_maybeRefetchFullHistory`, 4s per-session cooldown, in-flight + tab-switch guards, viewport position held across the replay). The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`.
### Terminal scrollback: strip flavors and wheel/touch forwarding
**Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend `_sessionUsesServerMouseStrip()` mirror stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`.
**Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (codex, claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`.
**The wheel listener is CAPTURE-phase and Codeman owns the scroll** (#205 follow-up, measured on the live instance): xterm's viewport is a vscode-style ScrollableElement that consumes wheel events itself (preventDefault + stopPropagation) whenever it believes a scrollbar exists, ignores `attachCustomWheelEventHandler`, and goes DEAF after `terminal.reset()` — a tab switch or full-history replay leaves its scroll dimensions stale, after which wheel events neither scroll nor propagate reliably. A bubble-phase container listener therefore never fired once local scrollback existed (forwarding, deltaMode and the top-of-buffer re-pull all silently dead exactly on sessions WITH history), and after a tab switch nothing scrolled at all ("works at first, breaks after a tab switch"). The container wheel listener is `{capture: true}`, stops propagation, and scrolls locally via buffer-level `terminal.scrollLines()` (immune to the stale scroller). ⚠️ Two cases are deliberately passed through untouched, in this order BEFORE preventDefault: `mouseTrackingMode !== 'none'` (xterm's encoder forwards the wheel to the PTY — htop/vim with mouse on) and `buffer.active.type === 'alternate'` (direct-PTY vim/less: xterm's alt-scroll converts the wheel to cursor keys). Do not "simplify" this back to a bubble listener or re-delegate local scrolling to xterm's viewport. E2E guard: the reload → tab-switch → wheel matrix in the #205 verification scripts.
### Run launch synchronization
@@ -87,6 +95,19 @@ Implementation detail extracted from `CLAUDE.md` so that file stays small enough
The general rule: **any new endpoint that turns a caller-supplied `sessionId` into a filesystem path is an ownership boundary**, whether or not it goes through `findSessionOrFail`.
### File Viewer edit mode
**File Viewer edit mode** (issue #212, design in `docs/file-viewer-edit-plan.md`): the file-preview overlay can edit workspace text files in place — `GET /api/sessions/:id/file-content?edit=1` (read-for-edit) + `PUT /api/sessions/:id/file-content` (save), policy in `src/config/file-editing.ts`, UI in `panels-ui.js`. This is the **only file surface that writes**, so it carries every rule the read surfaces have plus its own:
- **Confinement is the read path's, plus write-only gates.** `findSessionOrFail` (ownership) → `validateSessionFilePath` (realpath + workspace boundary; escapes report as 404, same as reads) → sensitive-path + attachment-guard blocklists (403) → `.git/` subtree deny (403 — `.git/hooks/*` is code execution) → extension **allowlist** (400; `svg` and `env` deliberately excluded). ⚠️ **There is no `O_CREAT` anywhere in the handler** — that absence is what makes "edit-in-place only, never create" a structural property instead of a convention. Do not add a create path without treating it as a new security surface.
- **A truncated buffer must never become an edit buffer.** The plain preview truncates to `lines` (default 500); saving such a buffer would silently delete everything past the cut, and the hash check cannot catch it (the loaded prefix hashes differently from the full file, which reads as an ordinary conflict at best). `edit=1` therefore never truncates — it 413s over `MAX_EDITABLE_BYTES` (512KB) instead — and the frontend always re-fetches with `edit=1` before swapping in the textarea, even though the preview already holds content.
- **Concurrency is optimistic by content hash, not mtime.** The client echoes the sha256 it loaded (`baseHash`); mismatch → 409 CONFLICT (plain envelope — the error arm carries no data; the client re-fetches `edit=1` for fresh state) unless `force:true`. mtime alone is wrong: agents rewrite files within one timestamp tick.
- **Writes are `wx` temp + `fchmod` + `fsync` + `rename` in the target's directory.** `wx` cannot follow a pre-existing symlink and `rename()` replaces (not follows) a symlink final component, which closes the validate-then-write TOCTOU window; `fchmod` because `open()`'s mode argument is masked by the umask; a symlink whose target is *inside* the workspace is deliberately written through (validation returns the realpath). Trade-off (same as vim): the inode changes, so hardlinks keep old content.
- **Corruption guards**: NUL-sniff + UTF-8 **round-trip compare** (`Buffer.from(buf.toString('utf8'), 'utf8').equals(buf)`) refuse binary and non-UTF-8 files — decoding latin-1 yields U+FFFD replacements and writing those back destroys the original bytes. EOL is detected server-side and re-applied on save because a `<textarea>` normalizes to LF (a two-line edit of a CRLF file must not become a whole-file diff).
- **Two size caps on the wire**: the Zod `.max()` counts UTF-16 code units (coarse pre-filter, 400) while the handler's `Buffer.byteLength` check enforces the real byte cap (413); the route sets `bodyLimit: 4MB` because JSON escaping can expand 512KB of content past Fastify's 1MB default. Error paths **throw** structured `{statusCode, body}` errors (`throwFileEditError`) rather than returning envelopes — the central preSerialization status-mapping hook is absent from the route-test harness, and 413 has no errorCode mapping at all.
Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write-routes.test.ts` (deliberately **unmocked fs** against a real temp workspace — symlink/TOCTOU/mode behavior must be exercised for real).
### Ultracode and workflow-run visualization
**Ultracode / Workflow-run visualization** (opt-in `showUltracodeAgents`, default OFF; released 1.1.2): the Workflow tool ("ultracode") writes a COMPLETION artifact per run at `~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_*.json` (written only at run end); LIVE in-flight runs exist only as transcript dirs at `…/subagents/workflows/wf_<id>/` (journal.jsonl + `agent-*.jsonl`). `workflow-run-watcher.ts` (STANDALONE — deliberately never imports/touches `subagent-watcher.ts`; separate singleton, though it independently reads the same `subagents/workflows/` tree) scans BOTH sources via periodic poll + per-directory chokidar watchers with per-source mtime skip (LRU agentStatCache + journalCache), synthesizing ACTIVE runs (live per-agent tokens/tools/state from transcripts, title/phases from the workflow script) until the completion `wf_*.json` appears and supersedes, and broadcasts SSE `workflow:run_discovered`/`run_updated`/`run_removed`. The watcher is started when **either** `showUltracodeAgents` **or** `ultracodeFloatingWindows` is on (`server.ts` `isWorkflowAgentTrackingEnabled()` returns `(showUltracodeAgents ?? false) || (ultracodeFloatingWindows ?? false)`). Served via `GET /api/workflows` (optional `?minutes=` filter) and `GET /api/workflows/:runId`. Frontend `ultracode-panel.js` renders a docked master-detail view (LEFT: runs + phases; RIGHT: per-agent tokens + tool-calls; click an agent card → its live transcript via client-side `agentId` join). **Additionally**, `ultracode-windows.js` auto-pops a draggable **floating window per active run** (gated on a **DEDICATED** `ultracodeFloatingWindows` toggle, default OFF — independent of the dock panel's `showUltracodeAgents`; see `_ultracodeFloatingEnabled()`), connected by a glowing line to the originating session tab (resolved by `session.claudeSessionId === run.sessionUuid`) — same line idiom as subagent windows, drawn into the shared `#connectionLines` SVG from the tail of `_updateConnectionLinesImmediate`. The window auto-closes ~8s after its run finishes; explicit dismissals are remembered. Clicking an agent card opens an **in-page** connected transcript window (not a browser popup); both run and transcript windows minimize **into** the originating session tab as a merged `ULTRA` badge (🧬 runs / 📄 transcripts) with a restore/dismiss dropdown — minimized runs are skipped by auto-pop. Gesture beta: floating subagent/ultracode windows are pinch-draggable (a `window` grab kind in `entry.ts`). Types: `src/types/workflow-run.ts`. Config: `src/config/workflow-config.ts`.
@@ -138,6 +159,14 @@ The general rule: **any new endpoint that turns a caller-supplied `sessionId` in
**Command palette + shortcut registry** (COD-151/153/157/192, #146): `Ctrl/Cmd/Alt+K` opens the session palette (fuzzy search over live sessions; "Browse all sessions" → the Session Manager modal backed by `GET /api/sessions/unified`); the quick-start case `<select>` is fronted by a searchable picker (`buildCasePickerOptions`/`formatCasePickerLabel` — remote cases render `name @ hostId`). Shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js; overrides persist under `settings.shortcutOverrides` via `saveAppSettingsToStorage`); App Settings → Shortcuts renders capture/disable rows; `Ctrl+?` opens the registry-driven overlay (footer links to the full `#helpModal` reference). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM — keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over.
**Terminal smart copy** (#211): `Ctrl+C` copies the selection when there is one and stays the interrupt when there isn't. Three rules keep that split honest, and breaking any of them silently costs the user their interrupt key:
1. The branch lives in `attachCustomKeyEventHandler` (terminal-ui.js) and the **no-selection path returns `true` with no `preventDefault()`**, so xterm still evaluates `Ctrl+C` into `0x03`. Returning `false` alone does not cancel the event either way: xterm's `_keyDown` calls the custom handler *before* its own `cancel()`, which is exactly why the copy path calls `preventDefault()` explicitly (otherwise the browser also runs its native copy on top).
2. `copyTerminalSelection` is a registry `action` **deliberately missing from `SHORTCUT_ACTIONS`** (same trick as `command-palette`): the entry stays rebindable and disableable in App Settings, while the generic document-capture loop, which `preventDefault()`s every match it dispatches, skips it and lets the terminal handler decide.
3. The gate is keydown-only (the custom handler also runs for `keypress`/`keyup`), and `Ctrl+Shift+C` never falls through to the PTY: an "explicit copy" chord that interrupts a running agent because the selection happened to be empty is a footgun with no upside.
Copy goes through `_copyText()` (Clipboard API, then hidden-textarea + `execCommand`), not raw `navigator.clipboard`, because `install.sh`'s LAN option serves plain HTTP where `navigator.clipboard` is undefined; the fallback steals focus, so the terminal is refocused afterwards. Related: xterm registers its own `copy` listener on the terminal element gated on `hasSelection()`, which is why right-click → Copy has always worked. Selection itself is unavailable on touch devices by design (`user-select: none` on the terminal subtree), and in `shell`/`opencode`/`antigravity` tabs the TUI owns the mouse, so selecting there needs Shift+drag. Tests: `test/terminal-copy-selection.test.ts` (gate + wiring invariants), `test/terminal-copy-shortcut.test.ts` (browser, real key presses).
### WebGL renderer toggle
**WebGL renderer toggle** (#140, `webglRendererEnabled`): per-device (`displayKeys` set, stripped from the server payload — NOT in `SettingsUpdateSchema`, which is `.strict()`). The GPU-stall watchdog's sticky `codeman-webgl-disabled` marker survives page loads; it's cleared only by an explicit OFF→ON save transition or `?webgl=force` (`shouldSkipWebGL` in constants.js). `?nowebgl` still forces the DOM renderer per-load.
@@ -147,6 +176,10 @@ The general rule: **any new endpoint that turns a caller-supplied `sessionId` in
**Multi-monitor button** (header, top-right; the notification bell it sits beside stays hidden — notifications live in Settings → Notifications). `app.launchMultiMonitor()` (in `panels-ui.js`) POSTs `/api/system/span-displays`, which spawns `scripts/span-codeman.sh` — a fresh, maximized browser `--app` window sized to the union of all displays (macOS; needs "Displays have separate Spaces" OFF). Supports the gesture layer's in-page floating session panels dragging across the physical monitor seam. **Opt-in:** hidden by default; enable under App Settings → Display → **Header Displays** ("Multi-monitor Button", `showMultiMonitorButton`). The button carries a `btn-multimonitor--hidden` class in the template; `renderIndexHtml` strips that class at render when the setting is on (a unique class token, not a brittle match on the aria-label/style copy), and `applyHeaderVisibilitySettings()` toggles the same class live on save. Solo (detached) windows hide it via `body.solo-mode`.
**Response-viewer (eye) button** (header) is likewise **hidden by default** — enable under App Settings → Display → **Response Viewer** (`showResponseViewer`). Works for Claude AND Codex sessions (#152): Codex last-responses are located via a 4-layer rollout resolution under `CODEX_HOME` (history pin → originator match → resume-UUID → cwd fallback with other-pane exclusion), with injected-context filtering and event/legacy dedup — tests in `test/routes/session-routes-codex-last-response.test.ts`.
⚠️ **A Claude pane's conversation is identified by the pane's own Enter, never by "newest entry for this cwd".** `~/.claude/history.jsonl` records every submitted prompt as `{project, sessionId, timestamp}`, and `/clear` moves the pane to a fresh `<uuid>.jsonl` that nothing on the PTY announces — so the viewer has to re-derive the live conversation. Keying that off `project` alone was the bug: a cwd is shared with every other Codeman tab on it, with tabs long since closed, and with any plain `claude` the user runs in their own terminal, so the eye followed whichever of those conversations was typed into last and showed a stranger's transcript. `resolveActiveClaudeSessionIdFromHistory()` instead credits an entry to a pane only when it lands within `CLAUDE_SUBMIT_MATCH_MS` of that pane's `Session.lastSubmitAt` **and** no other pane on the same cwd submitted closer — the same last-submit correlation the Codex locator uses. With no correlated entry the pane keeps the id it has: a viewer one turn behind beats a viewer showing someone else's conversation.
⚠️ **`Session.lastSubmitAt` is persisted state, not a runtime counter.** `start()` reassigns `_claudeSessionId = resumeSessionId || id` on every launch — including the re-attach path for a mux session that survived the restart — so a recovered pane always points the viewer at its *launch* conversation, even when the CLI moved on via `/clear` hours earlier. The submit anchor is the only thing that can correct that without user input, so it round-trips through `SessionState.lastSubmitAt` and is restored in `restoreMuxSessions()`. Drop it from `toState()` and recovered panes silently show the pre-`/clear` transcript until the user types again. Restoring a *stale* anchor is safe: the resolver's staleness guard rejects any candidate transcript older than the one the pane is currently on, which is exactly the shape of a respawn into a fresh conversation.
⚠️ **Claude transcripts are grouped at real human-turn boundaries, not per JSONL row.** A Claude transcript is an append-only event log, so one logical exchange spans many rows: tool-result rows, meta/image/skill rows, compact summaries, task/team notifications, sidechains, replayed assistant snapshots, and multi-block assistant output. Rendering a card per row was the bug: it produced duplicate and truncated cards that looked like the viewer had lost the response. The grouping walks to the next genuine user turn and dedups replayed assistant snapshots while preserving the tool/task/skill/compact/team metadata filtering. Related: a recovered `restored-<uuid8>` tmux placeholder carries a **stale cwd**, so transcript lookup by working directory finds nothing; it rebinds to the matching top-level Claude transcript UUID instead when that match is unambiguous. Tests: `test/routes/session-routes-claude-last-response.test.ts`. Purely client-side (no `renderIndexHtml` step): the template ships with `btn-response-viewer-header--hidden` and `applyHeaderVisibilitySettings()` (settings-ui.js) toggles it after settings load. Hiding must go through that marker class — the base rule is `display:inline-flex !important`, so an inline style can't override it. `showResponseViewer` is in the `displayKeys` per-device set (settings-ui.js), so it does NOT sync across devices.
**File Viewer button** (header, 1.4.1) is **shown by default on desktop** since `211f3c0` (post-1.8.0): toggle under App Settings → Display → **Header Displays** → File Viewer (`showFileViewerButton`, in the per-device `displayKeys` set, fallback default `true`). Purely client-side like the response viewer: the template now ships the button VISIBLE (no `--hidden` class) and `applyHeaderVisibilitySettings()` toggles the `btn-file-viewer--hidden` marker class after settings load; phones still hide it via mobile.css. The button toggles the file-browser panel open/closed without opening the settings modal (`panels-ui.js`). The same commit set the **default desktop header** to WS/CPU/MEM + File Viewer + gear: the token-count chip (`showTokenCount`, no settings-UI toggle) and the lifecycle-log button (`showLifecycleLog`) both default **OFF** now (templates ship them hidden; stored prefs still honored). The plan-usage chip default is unchanged (opt-in, see Plan-usage chip). The **Cron toolbar button** joined the same opt-in pattern in 1.6.0: template ships `btn-cron--hidden`, `applyHeaderVisibilitySettings()` toggles it via the per-device `showCronButton` setting (default OFF, App Settings → Display → Header Displays); cron jobs themselves are unaffected.
@@ -191,7 +224,7 @@ The general rule: **any new endpoint that turns a caller-supplied `sessionId` in
| **Rate limit** | 10 failed auth/IP → 429 (15min decay). QR has separate limiter |
| **Hook bypass** | `/api/hook-event` (and `/api/status-telemetry`, the statusLine exporter) skip Basic auth (localhost-only, schema-validated). When auth is active (`CODEMAN_PASSWORD` set), the loopback bypass requires the per-instance `X-Codeman-Hook-Secret` header **unconditionally** — COD-54 introduced it tunnel-gated; COD-91 (PR #127) made it always-on because Codeman can't detect a user's own loopback reverse proxy (own cloudflared/`tailscale serve`/nginx → 127.0.0.1), closing that residual plain-bypass gap. Hook curls cat the secret file at exec time via `$CODEMAN_HOOK_SECRET_FILE` (session env, `config/hook-secret.ts`); a missing/wrong secret gets 401 and rate-limits in a dedicated bucket (never locks out login). Tunnel enable **refuses** without `CODEMAN_PASSWORD` unless exposure is acknowledged — via `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` (env, COD-55) **or** the per-request `acknowledgeUnauthTunnel:true` action field (1.1.9): the welcome/settings tunnel toggle pops a security confirm dialog and, on confirm, resends with that flag (server logs a loud warning on every passwordless tunnel start; curl/API stay refused without password/env/flag). The flag is an action field, never persisted |
| **Env vars** | `CODEMAN_MUX` (managed session), `CODEMAN_API_URL` (auto-set for hooks), `CODEMAN_ALLOWED_HOSTS` (extra Host/Origin allowlist entries for reverse proxies, comma-separated; bare `.suffix` matches subdomains), `CODEMAN_DOCKER_BRIDGE_HOOKS`=1 (opt-in hooks-only listener on the docker bridge gateway so in-container hooks reach a loopback-bound server; bind IP from `CODEMAN_DOCKER_BRIDGE_HOST` or auto-detect) |
| **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`) |
| **Validation** | Zod schemas, Unicode-aware path allowlist regex, env prefix allowlist (`CLAUDE_CODE_*`/`OPENCODE_*`/`CODEX_*`/`GEMINI_*`/`GOOGLE_*`/`ANTIGRAVITY_*`) |
| **Headers** | CORS localhost-only, CSP, X-Frame-Options, HSTS if HTTPS |
## Performance and limits
+2 -2
View File
@@ -44,9 +44,9 @@ records), kept distinct from the existing `ScheduledRun`.
## 2. Where agent/session types are defined
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity'`
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode}-cli-resolver.ts`.
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
## 3. Where input is sent into a session
+2 -2
View File
@@ -1,7 +1,7 @@
# Cron Jobs — User & Operator Guide
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini) session on a schedule and
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
+16 -1
View File
@@ -2,7 +2,7 @@
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` all work inside the container.
## One-time setup: build the base image
@@ -15,6 +15,21 @@ node scripts/build-agent-image.mjs # builds codeman/agent:base
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
```bash
node scripts/build-agent-image.mjs --no-cache
```
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
+255
View File
@@ -0,0 +1,255 @@
# Extending Codeman
Codeman has no plugin runtime, and that is a deliberate choice rather than a
missing feature. A plugin runtime means running third-party code inside a process
that spawns agents with your credentials, on a server people routinely expose
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
exist, so it does not hand that away for an extension mechanism.
Instead there are four seams that already work, from any language, with nothing
installed:
| You want to | Use | Runs where |
| --- | --- | --- |
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
Everything below is covered by the stability promise in
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
envelope, `errorCode` values, and SSE event names are stable. Additive changes
(new endpoints, new optional fields, new events) are non-breaking. Breaking
changes ship under a new prefix (`/api/v2`).
## Before you start
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
authenticate once and keep the `codeman_session` cookie. With no password set,
Codeman is loopback-only and unauthenticated.
```bash
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
```
**Envelope.** Every response is `{"success": true, "data": ...}` or
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
is in [`api-reference.md`](api-reference.md).
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
the payload at the top level rather than under `data`. Read defensively with
`body.data ?? body`.
**Already driving Codeman from an agent?** The README's
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
running inside Codeman find the API and avoid acting on itself. This page is for
code running *outside* a session.
## Seam 1: Web tabs
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
your agent sessions. You write a normal web page; Codeman handles embedding it.
```bash
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
-H 'Content-Type: application/json' \
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
```
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
framing check), `POST /api/v1/webviews/:id/open`.
### Why it is proxied
By default your page is served through Codeman's own origin at `/webview/:cap/*`
rather than framed directly. A direct iframe fails three ways at once: production
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
cross-origin frames. Proxying solves all three without weakening the CSP.
### The two things that will confuse you
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
`trusted: true`. Two consequences look like bugs in your own app:
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
patches the common DOM sinks, but if you construct URLs in an unusual way,
prefer relative paths.
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
itself renders fine, this is the area to look at.
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
the API that spawns agents. Only mark your own trusted code.
## Seam 2: SSE events
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
`event: <name>` plus `data: <json>`. There are 149 event names following a
`domain:action` convention, registered in `src/web/sse-events.ts`.
The ones most integrations want:
| Event | Meaning |
| --- | --- |
| `session:created`, `session:deleted` | A session appeared or went away |
| `session:idle` | The agent stopped working |
| `session:completion` | A completion message was detected |
| `session:exit`, `session:error` | The session ended or failed |
| `hook:permission_prompt` | The agent is asking for permission |
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
| `hook:task_completed`, `task:completed` | Work finished |
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
| `mux:died` | A multiplexer session died unexpectedly |
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
### Filtering
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
sessions you did not list. Lifecycle and metadata events are always delivered, so
you cannot accidentally filter away the thing you are listening for.
Pass `?clientId=<uuid>` to enable live filter updates through
`POST /api/v1/events/subscribe` without reconnecting the stream.
### Example: notify when any agent needs you
```js
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buf = '';
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buf += decoder.decode(value, { stream: true });
const frames = buf.split('\n\n');
buf = frames.pop() ?? '';
for (const frame of frames) {
const name = frame.match(/^event: (.+)$/m)?.[1];
const data = frame.match(/^data: (.+)$/m)?.[1];
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
}
}
```
## Seam 3: HTTP API and CLI
Around 199 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
The common ones:
```bash
# List sessions (live + persisted + transcript history, deduped)
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
# Create a session
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
-H 'Content-Type: application/json' \
-d '{"workingDir":"/home/me/project","mode":"claude"}'
# Send a prompt (single-line only)
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests","useMux":true}'
```
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
`seq` (monotonic per session). Send both and the server applies each pair
at-most-once, so retrying after a dropped connection cannot type the prompt
twice. Omit them entirely rather than sending `null`.
For shell scripting, the `codeman` CLI is the same surface without the HTTP
plumbing:
```
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
codeman ralph start|stop|status|reset codeman users add|passwd|list
codeman status | list | attach <path> codeman doctor
```
## Seam 4: Hooks
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
Codeman installs its own hooks automatically, but the endpoint is open to yours.
```json
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
```
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
event.
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
the loopback bypass requires the `X-Codeman-Hook-Secret` header
(`~/.codeman/hook-secret`) unconditionally.
## Gotchas
Every one of these has cost somebody real time.
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
call the API. Integrate server-side.
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
work. Cross-site origins are blocked by the CSRF guard.
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
shipped bugs more than once.
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
simple-request CSRF, so it is deliberate. Send `application/json`.
- **Prompts are single-line.** With `useMux: true` the server delivers your text
and then Enter as two separate writes, so you do not append `\r` yourself. A
multi-line string breaks the agent's Ink-based input handling: send one line,
or split it across calls.
- **Unwrap the envelope** before reading fields. `data` is not the response body.
## Publishing your integration
There is no registry and no review queue. Add the GitHub topic
**`codeman-integration`** to your public repository so others can find it, and
link back to Codeman in your README.
If a real ecosystem of these appears, a manifest format and an install command
become worth building. Until then, these four seams are the contract, and they
require nothing of you but HTTP.
## What Codeman deliberately does not have
- **No in-process plugin runtime.** See the reasoning at the top of this page.
- **No build or startup hooks** for third-party code. Run your own process.
- **No per-plugin config or state directories.** Manage your own files.
- **No sandbox for integration code**, because Codeman never launches it. Your
integration is your own process, started by you, with your permissions,
talking HTTP.
That last point is about integration code specifically, not about Codeman.
Sandboxing lives on a different axis here: the thing worth isolating is the
**agent**, and you isolate it per case with
[Docker cases](docker-cases.md), which run the agent in a hardened container with
a bind-mounted workspace and seeded (not shared) credentials. An integration that
creates or drives a Docker-backed session inherits that isolation for free, since
it is a property of the session rather than of the caller.
+432
View File
@@ -0,0 +1,432 @@
# File Viewer edit mode (issue #212)
Plan only. No implementation yet.
Goal: close the loop "agent writes a file, you review it in the viewer, tweak two lines, save, tell the
agent to continue" without hopping into the terminal, with the phone as the primary target.
Scope from the issue: an Edit toggle on text previews, a write endpoint that inherits the read path's
confinement, text-only, edit-in-place (no create, no delete, no rename), no editing through the
Docker/remote overlays.
---
## 1. What exists today
**Read path (backend), all in `src/web/routes/file-routes.ts`:**
| Route | Line | Notes |
| ------------------------------------ | ------ | ------------------------------------------------------------------ |
| `GET /api/sessions/:id/files` | `741` | Tree scan of `session.workingDir`, hidden files off by default |
| `GET /api/sessions/:id/file-content` | `865` | The text/preview classifier. `findSessionOrFail` + `validateSessionFilePath` |
| `GET /api/sessions/:id/file-raw` | `1018` | Bytes, 50MB cap |
| `GET /api/sessions/:id/file-preview` | `1254` | DOCX/PPTX to PDF, everything else redirects to `file-raw` |
| `GET /api/download` | `1384` | The only read route that also runs `isSensitivePath()` |
`file-content` classification order (`file-routes.ts:881-1011`): extension buckets (image / video / audio /
known-binary) return metadata only; otherwise the bytes are read, sniffed for a NUL in the first 8KB, and
either reported as `type:'binary'` or decoded as UTF-8 and **truncated to `lines` (default 500, hard cap
10000)**. Caps: `MAX_TEXT_FILE_SIZE` 10MB.
Confinement is `validateSessionFilePath()` (`src/web/route-helpers.ts:67`): `resolve()` then `realpathSync()`
then reject if the result is not under `workingDir`. Because it realpaths the *full* path, a symlink whose
target escapes the workspace is already rejected. Ownership is `findSessionOrFail()` which runs
`canAccessOwned()` (`route-helpers.ts:102`), a no-op outside multi-user mode.
**Read path (frontend), `src/web/public/panels-ui.js`:**
- `loadFileBrowser()` `2947`, `renderFileBrowserTree()` `2978`, click to `openFilePreview()` `3056`.
- `openFilePreview(filePath, sessionId, attachmentId)` `3193`: attachment-id branch, then docx/pptx, pdf,
svg branches, then the generic `file-content` fetch at `3274` with **`&lines=500` hardcoded**, rendering
text as `<pre><code>${escapeHtml(...)}</code></pre>` at `3298` and stashing `this.filePreviewContent`.
- `closeFilePreview()` `3308`, `copyFilePreviewContent()` `3751`.
- Markup: `src/web/public/index.html:420-432` (`filePreviewOverlay` / `-Title` / `-Body` / `-Footer`, two
header buttons: copy and close).
- CSS: `src/web/public/styles.css:9320-9430`. Overlay `z-index: 2000`, window `80vw/80vh`, capped
`900x700`. There are **no `.file-preview-*` rules in `mobile.css` at all**.
**Reachability on phones.** The header File Viewer button is hidden below 430px
(`mobile.css:482`, locked by `KNOWN_PHONE_HIDDEN` in `test/mobile-header-buttons-policy.test.ts`), so on a
phone the preview overlay is reached through:
1. an attachment card's **Preview** button (`panels-ui.js:3451`), which is exactly the "agent just wrote a
file" path the issue describes,
2. the attachment-history drawer (`panels-ui.js:3709`),
3. App Settings to Panels to **File Browser** (`showFileBrowser`, applied in `settings-ui.js:2202`; the
panel is mobile-styled at `mobile.css:1868`).
So edit mode is reachable on a phone today via (1) and (2) without touching the header policy. Improving
the entry point is listed as an open decision in section 10, not assumed.
---
## 2. Threat model, stated honestly
Anyone who can call this API can already reach `POST /api/sessions/:id/input` and type an arbitrary prompt
into an agent running with `--dangerously-skip-permissions`. A workspace-confined write endpoint therefore
does not create a new privilege tier for an authenticated caller.
What it *would* create if built carelessly is a **new host-write primitive reachable by path**, so the
things this plan actually defends against are:
1. **Path traversal / symlink escape** writing outside the workspace.
2. **TOCTOU**: a path component that becomes a symlink between validation and write.
3. **Cross-user writes** in multi-user mode (`canAccessOwned`).
4. **Silent data loss**, which is the highest-probability real-world failure here and gets its own section.
CSRF is already covered: `registerHostGuard()` (`src/web/middleware/auth.ts:555-578`) rejects any
non-safe-method request whose `Origin` is cross-site. The webview-capability exemption at that gate is
fenced to `GET`/`HEAD` for the Referer form (`auth.ts:161`) and to `/webview/:cap/*` paths for the path
form, so a proxied dashboard cannot reach a new `PUT /api/...`. Using `PUT` + `application/json` also
forces a preflight for any cross-origin attempt.
---
## 3. Backend design
### 3.1 New policy module: `src/config/file-editing.ts`
Pure, unit-testable, no IO (config lives in `src/config/`, no barrel, import the file directly).
```ts
export const MAX_EDITABLE_BYTES = 512 * 1024; // content cap, both directions
export const EDITABLE_EXTENSIONS: ReadonlySet<string>; // ts,tsx,js,jsx,mjs,cjs,json,jsonc,md,mdx,txt,
// css,scss,less,html,htm,xml,svg?,yml,yaml,toml,
// ini,cfg,conf,env?,sh,bash,zsh,fish,py,rb,go,rs,
// java,kt,swift,c,h,cpp,hpp,cs,php,sql,graphql,
// proto,lua,pl,r,jl,tf,gradle,csv,tsv,log,diff,patch
export const EDITABLE_BASENAMES: ReadonlySet<string>; // Dockerfile, Makefile, LICENSE, .gitignore,
// .prettierignore, .editorconfig, .nvmrc, ...
export function isEditableFileName(fileName: string): boolean;
export function isDeniedEditRelativePath(rel: string): boolean; // `.git/` subtree
export function detectEol(text: string): 'lf' | 'crlf';
export function applyEol(text: string, eol: 'lf' | 'crlf'): string;
```
Decisions baked in:
- **Allowlist, not blocklist**, per the issue and per the existing attachment-guard precedent.
- `svg` and `env` are deliberately marked with `?` above: `svg` is served as an untrusted octet-stream on
the read side (`file-routes.ts:118`) so allowing an edit is defensible, but I recommend **excluding
both** in v1. `.env` files are matched by `isSensitivePath()` anyway and would be rejected downstream;
excluding them at the allowlist keeps a single obvious refusal.
- `isDeniedEditRelativePath` blocks the `.git/` subtree: `.git/hooks/*` is code execution and a corrupt
index is unrecoverable-looking to a user who only wanted to fix a typo. Other dotfiles stay allowed but
are not reachable from the tree UI anyway (`showHidden=false`).
### 3.2 Read-for-edit: extend the existing GET
`GET /api/sessions/:id/file-content?path=<rel>&edit=1`
When `edit=1`:
- skip line truncation entirely (a truncated buffer must never become an edit buffer, see section 4.1),
- enforce `MAX_EDITABLE_BYTES` instead of `MAX_TEXT_FILE_SIZE` and answer 413 over it (as a structured
throw with `statusCode: 413`, the `throwFilesystemPickerError` pattern, since the central errorCode-to-
status map has no 413 entry; see the error-mechanics note in 3.3),
- run the editability gate (`isEditableFileName`, `isDeniedEditRelativePath`, `isSensitivePath`,
`isBlockedAttachmentPath`) and the content gate (NUL sniff plus UTF-8 round-trip, see 4.3),
- return `{ content, size, mtimeMs, totalLines, truncated: false, extension, editable: true, hash, eol }`.
`hash` is `sha256` hex of the exact on-disk bytes.
Non-`edit` responses gain **only** `editable: boolean` (additive, no shape change for existing consumers),
which is all the UI needs to decide whether to show the Edit button. No `hash` on plain reads: the Edit
action re-fetches with `edit=1` anyway (section 4.1), which is where the hash comes from, and hashing every
casual 10MB preview would be pure waste.
### 3.3 Write: `PUT /api/sessions/:id/file-content`
Body (new `FileWriteSchema` in `src/web/schemas.ts`, Zod v4):
```ts
{ path: string, content: string, baseHash: string, eol?: 'lf'|'crlf', force?: boolean }
```
Registered with an explicit route option `{ bodyLimit: 4 * 1024 * 1024 }`. **Fastify's default `bodyLimit`
is 1MB and this repo configures none**, and JSON escaping expands content: 2x for a file full of quotes or
backslashes, up to 6x for control characters (each serialized as a `\uXXXX` escape), so 512KB of content
can legitimately exceed 1MB on the wire; blowing the limit produces a raw `FST_ERR_CTP_BODY_TOO_LARGE`, not an `ApiResponse` envelope. Two
related sizing notes: `z.string().max()` counts **UTF-16 code units, not bytes**, so the schema's `.max()`
is only a coarse pre-filter and the real cap is an explicit `Buffer.byteLength(content, 'utf8')` check in
the handler (step 7a below); and 4MB comfortably bounds the worst-case expansion of a 512KB file without
inviting multi-MB bodies elsewhere.
**Error mechanics** (matters for both prod behavior and testability): a handler that *returns* a
`{success:false, errorCode}` envelope gets its HTTP status assigned centrally by the preSerialization hook
in `server.ts` (`httpStatusForErrorCode()`, `src/types/api.ts`), but the route-test harness
(`test/routes/_route-test-utils.ts`) installs only `installRouteErrorHandler`, **not** that hook, so
returned envelopes surface as HTTP 200 in tests. The PUT handler should therefore use the same
structured-**throw** pattern as the filesystem picker (`throwFilesystemPickerError`, `file-routes.ts:411`):
thrown `{statusCode, body}` errors are rendered identically in prod and in the harness, and they allow the
one status the code map cannot express (413). The error envelope itself is strictly
`{success:false, error, errorCode}`, **it has no data arm**, so no error response may carry extra payload.
Handler order (each step is a test case):
1. `findSessionOrFail(ctx, id, req)` (live sessions only, matching the read route, and it carries the
multi-user ownership check).
2. `parseBody(FileWriteSchema, req.body)`, then `Buffer.byteLength(content, 'utf8') <= MAX_EDITABLE_BYTES`
or 413 (the schema `.max()` alone cannot enforce a byte cap, see the sizing note above).
3. `validateSessionFilePath(session.workingDir, path)` or 404 (do not distinguish "outside workspace" from
"missing", matching the read route).
4. `isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)` or 403.
5. `isDeniedEditRelativePath(relativePath)` or 403.
6. `isEditableFileName(basename(resolvedPath))` or 400.
7. `stat`: must be `isFile()`, size within `MAX_EDITABLE_BYTES`, else 400/413. **No `O_CREAT` anywhere in
this handler**, which is what enforces edit-in-place.
8. Read current bytes, compute `hash`, run the NUL sniff and the UTF-8 round-trip check, else 400.
9. `hash !== baseHash && !force` gives **409 CONFLICT** (`ApiErrorCode.CONFLICT`, plain envelope; the error
arm carries no data, see the error-mechanics note). The client's conflict dialog gets fresh state by
re-fetching `edit=1`, which it needs for its Reload action anyway.
10. Build the output buffer: `applyEol(content, eol ?? detected-from-original)`; re-check
`Buffer.byteLength` against the cap.
11. Write atomically in the resolved parent directory:
`fs.open(<dir>/.<name>.codeman-tmp-<rand>, 'wx', stat.mode & 0o777)`, then `fchmod(stat.mode & 0o777)`
(open's mode argument is masked by the process umask, so the chmod is what actually preserves an
unusual mode), write, `fsync`, close, `fs.rename(tmp, resolvedPath)`, unlink the temp on any failure.
12. Re-stat, return `{ success: true, data: { path, size, mtimeMs, hash, totalLines } }`.
Why `O_EXCL` temp plus rename rather than truncate-in-place:
- `wx` cannot follow a pre-existing symlink, which closes the TOCTOU window from step 3 to step 11 without
needing `O_NOFOLLOW` gymnastics.
- `rename()` does not follow a symlink in the final component, so even if `resolvedPath` were swapped for a
symlink after validation, the symlink itself is replaced and the swap target is untouched.
- A crash mid-write leaves the original intact.
Caveat to document in the code comment: rename replaces the inode, so hardlinks to the file keep the old
content. That is the same trade-off vim makes by default and is preferable to a truncate window here.
No SSE event in v1. Nothing else in the app needs to know: `image-watcher.ts` only reacts to
`.png/.jpg/.jpeg/.gif/.webp/.bmp/.svg/.pdf/.docx/.pptx` adds (`image-watcher.ts:23-25`), none of which are
editable text, and the temp filename does not match either.
---
## 4. The five traps
These are the parts that turn a "small write endpoint" into a bug report.
### 4.1 Truncation (the data-loss trap)
The frontend fetches `&lines=500` (`panels-ui.js:3274`). Saving that buffer back would **delete every line
past 500**. Worse, the content hash of the full file would still match, so an optimistic-concurrency check
cannot catch it.
Mitigations, all three:
- The Edit affordance is only offered when the loaded payload came from `edit=1` (which never truncates).
Tapping Edit on an already-rendered preview **re-fetches** with `edit=1` before swapping in the editor.
- The read-for-edit path 413s above `MAX_EDITABLE_BYTES` rather than truncating, so "too big to edit here"
is an explicit refusal with a message, never a silent partial buffer.
- A test asserts `edit=1` never returns `truncated: true`.
### 4.2 Line endings
A `<textarea>`'s `.value` normalizes to LF. Saving a CRLF file naively rewrites every line, producing a
whole-file diff for a two-line change. So: the read returns the detected `eol`, the client echoes it back
unchanged, and the server re-applies it. Mixed-EOL files use the dominant style, which is lossy for the
minority lines; call that out in the response and accept it in v1.
### 4.3 Encoding
`buf.toString('utf-8')` on a latin-1 or otherwise non-UTF-8 file yields U+FFFD replacement characters, and
writing that back **corrupts the file**. The check is a round-trip:
`Buffer.from(decoded, 'utf8').equals(buf)`. If it fails, `editable: false` and the write is refused. This
also catches binary content that the NUL sniff misses. A UTF-8 BOM survives because it round-trips as a
leading U+FEFF; do not strip it.
### 4.4 Concurrency with the agent
The whole use case is editing a file the agent just wrote and may write again. `baseHash` plus 409 is the
guard. Do not use mtime alone: agents rewrite files within a single filesystem timestamp tick, and an
identical rewrite should not be reported as a conflict.
### 4.5 Symlinks and TOCTOU
Covered by `validateSessionFilePath` (escape) plus `wx` temp and `rename` (post-validation swap). One
intentional allowance: a symlink whose target is *inside* the workspace is edited through to its target,
because `validateSessionFilePath` returns the realpath. That matches what a user tapping the file expects.
---
## 5. Frontend design
All in `panels-ui.js` (prettier-exempt, hand-formatted; match the surrounding style), `index.html`,
`styles.css`, `mobile.css`.
### 5.1 State
```js
filePreviewEdit = { active, sessionId, path, baseHash, eol, original, dirty }
```
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
### 5.2 Markup (`index.html:420-432`)
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
untouched.
### 5.3 Behavior
- `openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
keyboard before the user has picked a line.
- `input` sets `dirty` and enables Save.
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
response, leave edit mode, re-render the read view from the local editor value (the response carries
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
itself carries no state (section 3.3, step 9).
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
guard.
- `copyFilePreviewContent()` copies the live editor value while editing.
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
real bug twice.
### 5.4 Mobile
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
the editor tracks the keyboard for free.
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
smaller `rem` value.
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
- No header-policy change is needed (section 1), so
`test/mobile-header-buttons-policy.test.ts` stays untouched.
### 5.5 i18n
`i18n.js` already skips `textarea`, `pre`, `code` and `.file-preview-content` in its `SKIP_SELECTOR`
(`i18n.js:20-38`), so file content is never translated. Add zh-CN entries for the new chrome: Edit, Save,
Cancel, Unsaved changes, Discard unsaved changes?, File changed on disk, Reload, Overwrite, Saved,
Too large to edit here.
---
## 6. Docker and remote cases
Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
---
## 7. Tests
| File | Kind | Covers |
| ------------------------------------------- | ----------- | ---------------------------------------------------------------------- |
| `test/file-editing-policy.test.ts` | pure unit | `isEditableFileName` (allow + deny + basenames), `isDeniedEditRelativePath`, `detectEol`/`applyEol` round-trip incl. mixed EOL, BOM preservation |
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
Route cases to assert explicitly:
1. happy path writes the bytes and returns a new hash
2. `../` and absolute paths give 404
3. symlink pointing outside the workspace gives 404
4. symlink pointing inside is written through to the target
5. non-allowlisted extension gives 400
6. `.git/config` gives 403
7. a `.env` in the workspace gives 403 (sensitive-path)
8. a file with a NUL byte gives 400
9. a latin-1 file that fails the UTF-8 round-trip gives 400
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
11. over `MAX_EDITABLE_BYTES` gives 413
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
14. CRLF file edited and saved stays CRLF
15. file mode is preserved across the temp-plus-rename
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
- `curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
`w1`/`w2`/`w3`; delete it by exact id afterwards.
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
assert the bytes on disk changed. Assert real state, not HTTP 200.
---
## 8. Docs and release
- This plan lives at `docs/file-viewer-edit-plan.md`.
- `docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
chain, the truncation invariant, and why temp-plus-rename.
- `CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
Remember `CLAUDE.md` is prettier-ignored on purpose.
- `docs/api-reference.md`: the new `PUT` and the `edit=1` query.
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
additive API surface, so **COM minor** when it ships.
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
and must pass `npm run format:check`.
---
## 9. Implementation order
Each phase is independently reviewable and leaves the tree working.
1. **Policy module + tests.** `src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
route wiring. (Small.)
2. **Read-for-edit.** `edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
tests. Nothing consumes it yet. (Small.)
3. **Write endpoint.** `FileWriteSchema`, `PUT` handler, `test/routes/file-write-routes.test.ts`. Fully
testable by curl before any UI exists. (Medium, the security-relevant part.)
4. **Desktop UI.** Edit button, textarea swap, Save/Cancel, dirty guard, 409 flow. (Medium.)
5. **Mobile pass.** `mobile.css` sizing against `--app-height`, font size, accessory-bar interaction,
real-device check. (Small but the part that decides whether the feature is actually usable.)
6. **Docs, i18n strings, changeset.**
---
## 10. Open decisions
1. **Editor widget.** Recommend a plain `<textarea>` for v1: zero dependencies, no CSP question, no bundle
growth, and it is the only thing guaranteed to behave with the iOS keyboard. CodeMirror-light with
syntax highlighting is a clean follow-up once the write path is proven. The issue allows either.
2. **Phone entry point.** Edit mode is reachable on a phone through attachment cards and the history
drawer without changing anything. A dedicated toolbar or overview affordance for "browse this session's
files" would make it discoverable, but it is a separate UX change and would need a decision against the
deliberately minimal phone header policy. Recommend deferring it and revisiting after the feature ships.
3. **`svg` editability.** Recommend excluded in v1 (it is deliberately treated as untrusted on the read
side). Easy to add later.
4. **Create / delete / rename.** Explicitly out of scope per the issue. Note that keeping `O_CREAT` out of
the handler is what makes that a structural property rather than a convention.
+3 -3
View File
@@ -1,7 +1,7 @@
# Remote Sessions (SSH)
Codeman can run a session's agent on a **remote host over SSH** instead of the
local machine. The agent (Claude, OpenCode, Codex, Gemini, or a plain shell)
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, or a plain shell)
runs inside a `tmux` server **on the remote host**, so it survives the SSH
connection dropping; Codeman attaches to it the same way it attaches to a local
managed session.
@@ -30,7 +30,7 @@ Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini'>` — the modes that can run remotely. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
@@ -115,7 +115,7 @@ Key points:
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID *is*
the agent. The per-mode command comes from `remote.commands?.[mode]` or
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
`exec codex` / `exec gemini` / `exec bash -l`).
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
pane command is independently quoted, so a `remotePath` with spaces is safe.
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
+112
View File
@@ -0,0 +1,112 @@
# Scrollback fix plan (issue #205)
Status: IMPLEMENTED on `fix/scrollback-shell-alt-screen` (2026-08-07), with one deliberate
divergence from the recommendation below. Kept for the diagnosis record; the measured evidence
behind it is `docs/scrollback-issues-analysis.md`, and the mechanisms as shipped are documented
in `docs/architecture-invariants.md` (§ Full-scrollback replay, § Terminal scrollback: strip
flavors and wheel/touch forwarding).
What shipped vs. what this doc proposed:
- **Bug A (deltaMode)**: implemented as specified (`_wheelScrollLines()` normalizes
line/page/pixel units, Shift-axis trap kept).
- **Bug B (shell scrollback)**: implemented via the NARROW alt-screen strip for tmux-backed
shell/opencode/antigravity plus the scroll-to-top `full=1` re-pull, NOT the recommended
approach (a) `tmux mouse on`. The measurements in the analysis doc showed the alt buffer
comes from tmux's own client-side `smcup` at attach (tmux never forwards a pane program's
alt-screen toggles), so stripping that one sequence fixes both symptoms with no selection
tradeoff, keeps vim/less/htop untouched, and the re-pull also covers the repaint-burst
history loss that `mouse on` would not have addressed.
- **Invariant change**: the "viewport-at-bottom gate stays" invariant below was deliberately
DROPPED for forwarding modes: a repaint-mode CLI keeps no real terminal scrollback, so the
gate pinned users to a buffer of stale frames whenever the viewport parked off-bottom.
Forwarding now snaps to bottom first; Shift+wheel and the opt-out setting keep local
scrollback reachable. Touch forwards through the same gate (the mobile half of the fix).
- **Finding 5 (remote probe)**: implemented (`probeRemoteCliVersion` over ssh, deferred at
session start, same login-shell wrapper as the launch).
Original plan follows.
## Reports
- **Issue #205** (https://github.com/Ark0N/Codeman/issues/205), OPEN:
- **jonocodes** (author, 2026-08-03): SHELL session. Host Mac M4, brew tmux. On Android, touch-scrolling the terminal does nothing. On desktop, the mouse wheel cycles shell command history (acts like Up/Down arrows) instead of scrolling the screen.
- **mtiller** (comment, 2026-08-06): "similar issue just with scrolling backward to see agent output. This is with Firefox on MacOS." (Claude session implied.)
- **Reddit r/selfhosted** comment `p21x6ts` by mmtiller (= mtiller on GitHub): scrolling broken enough across phone/iPad/laptop that they fall back to Claude's own remote-control feature. Churn-risk user who otherwise loves the product; fixing this has promo value beyond the bug itself.
## How scrolling works today (read this before touching anything)
Three independent paths, all in `src/web/public/terminal-ui.js` unless noted:
1. **Desktop wheel** (container `wheel` listener, ~line 421): ALWAYS `preventDefault()`s, then either
- forwards synthetic SGR wheel reports to the app (`_sendSyntheticSgrWheel`, coalesced every 40ms, fire-and-forget) when `_shouldForwardWheelToApp(ev)` (~line 2823) passes: no Shift held, opt-out setting `terminalWheelLocalScrollback` off, xterm `mouseTrackingMode === 'none'`, session mode is `claude` with `cliVersion >= 2.1.187` or `codex`, and viewport is at bottom;
- otherwise scrolls xterm's LOCAL scrollback via `terminal.scrollLines(lines)`.
- `lines` comes from `_wheelScrollLines(ev)` (~line 2818): `delta / 25`, i.e. it assumes PIXEL deltas.
- NOTE: xterm.js's own internal wheel handler sits on an element INSIDE the container, so it runs FIRST (bubble order) and is not suppressed by the container's `preventDefault`.
2. **Touch** (touchstart/move/end, ~lines 441-585): converts touch deltas to `terminal.scrollLines()` with momentum. Touch is ALWAYS local-scrollback, never forwarded to the app. Tap-to-position (touchend, ~line 533) is separate and already handles both mouse-tracking-on and server-strip cases.
3. **Server-side strip** (`_handleTerminalOutput`, `src/session.ts:1384`): for modes in `isAltScreenStripMode()` (`src/session.ts:179` = `codex | claude | gemini`), strips alt-screen switches (`?47/?1047/?1049`), scrollback erase (`3J`), and mouse-tracking DECSETs (`?1000-?1007` except `?1004` focus) so content stays in xterm's normal buffer with scrollback intact. Includes a chunk-boundary carry so split sequences can't leak. `shell` and `opencode` (and `antigravity`) are deliberately EXCLUDED: arbitrary shell programs (vim/less/htop) legitimately need the alt screen. There is a parity copy of this strip on the replay path (`src/web/routes/session-routes.ts`, ~line 1697) and a frontend parity check `_sessionUsesServerMouseStrip()` (terminal-ui.js ~line 2751). All three must stay in sync.
4. Related: full-scrollback replay (`GET .../terminal?full=1` on first buffer load) fills xterm local scrollback; client scrollback is hardcoded 50k (`DEFAULT_SCROLLBACK`, constants.js) vs tmux 100k.
## Diagnosis
### Bug A: Firefox wheel deltas (mtiller's desktop case)
`_wheelScrollLines()` divides by 25 assuming `WheelEvent.deltaY` is pixels (`deltaMode === 0`, Chrome/Safari behavior). Firefox commonly fires `deltaMode === 1` (LINE units, deltaY around 1-3 per notch), so `Math.round(3/25) = 0` and the `|| ±1` fallback yields 1 line per event. With a discrete mouse wheel that is 1 line per notch: scrolling feels dead/broken. This hits BOTH the local-scroll path and the forwarded path, since both use the same function.
**Fix**: normalize by `ev.deltaMode` in `_wheelScrollLines()`:
- `deltaMode 0` (pixels): current behavior, `delta / 25`.
- `deltaMode 1` (lines): use the delta directly (round, keep sign fallback).
- `deltaMode 2` (pages): `delta * terminal.rows` (or a sane page size).
Keep the existing Shift-axis trap intact: on macOS trackpads Shift+two-finger scroll arrives as a HORIZONTAL wheel (deltaX carries the magnitude, deltaY ~0); that's why the function reads deltaX when Shift is held (issue #154). Don't lose it.
**Verify**: don't trust this diagnosis blindly. First reproduce in real Firefox on macOS and log `deltaMode`/`deltaY` (Firefox trackpad input can arrive as pixels; external mouse as lines). Also confirm the session's `cliVersion` probe succeeded (a failed probe disables forwarding entirely, which would point elsewhere). Unit-test by dispatching synthetic `WheelEvent`s with explicit `deltaMode` values; a Playwright `firefox` project pass is the end-to-end check.
### Bug B: shell mode has NO working scrollback at all (jonocodes)
Chain: shell mode is excluded from the alt-screen strip (correctly) → tmux attaches on the alternate screen → xterm's alt buffer has zero scrollback. Consequences:
- **Wheel**: xterm's own internal wheel handler runs first and, in the alt buffer, converts wheel ticks into Up/Down arrow keys (alternateScroll behavior). The shell receives arrows → command history cycles. That is jonocodes' exact desktop symptom. The container handler's `scrollLines()` afterwards is a no-op (no scrollback in alt buffer).
- **Touch**: the touch handler's `scrollLines()` is equally a no-op → "scrolling does nothing" on Android. Exact symptom two.
- The real history exists the whole time in tmux's 100k-line buffer; nothing exposes it.
**Fix, recommended approach (a): enable tmux `mouse on` for shell sessions.**
- Server-side, set `mouse on` scoped to shell sessions' tmux sessions (`tmux set-option -t <session> mouse on` at create + on attach of recovered sessions). Do NOT set it globally on the socket: claude/codex/gemini sessions rely on the DECSET strip and must not change.
- What this buys, all natively: tmux enables mouse tracking on the outer terminal → xterm `mouseTrackingMode` goes non-none → the container handler stands down (line ~2830 check) and xterm's own encoder forwards wheel as SGR reports → tmux scrolls its OWN copy-mode history on wheel-up, auto-exits at bottom. The alt-scroll arrow conversion disappears too (tracking mode takes precedence). Desktop is fully fixed with no new endpoints.
- **Touch**: still needs one small client change: in the touchmove path, when the active session is `shell` AND `mouseTrackingMode !== 'none'`, convert accumulated lines to `_sendSyntheticSgrWheel(x, y, lines)` instead of `scrollLines()`. The 40ms coalescing already prevents the tmux process storm (each send is a tmux send-keys server-side; unbatched flicks would spawn dozens of processes: this constraint is documented at `_sendSyntheticSgrWheel`, do not bypass it).
- **Selection tradeoff to verify**: with tracking on, xterm hands drag events to tmux instead of doing local browser selection. Shift+drag still does local selection (xterm shift-override). Verify this UX on desktop before shipping; if it's unacceptable, fall back to approach (b).
- **Also verify**: vim/less/htop inside the shell still behave (they'll now receive real mouse events via tmux, generally an improvement); remote shell sessions run tmux on the REMOTE host (`tmux -L codeman-remote`) and need the same option set there if remote shells are in scope (fine to defer, note it in the changeset if skipped).
**Fallback approach (b), only if (a)'s selection tradeoff fails testing**: keep mouse off; when a shell session is in the alt buffer, have the client send scroll intents to a small server endpoint that drives `tmux copy-mode -e -t <pane>` + `send-keys -X -N <n> scroll-up/down`. Preserves selection semantics exactly, but needs a new endpoint, server-side batching, AND suppression of xterm's native alt-scroll arrow conversion (capture-phase wheel listener with `stopPropagation`, or `attachCustomWheelEventHandler` if the vendored xterm version has it). More moving parts; (a) should be tried first.
**Not acceptable**: adding `shell` to `isAltScreenStripMode()`. vim/less/htop need the alt screen; that exclusion is deliberate and documented.
### Bug C: mtiller's phone/iPad case — UNREPRODUCED, do not guess
Touch is always-local by design, and Claude sessions keep content in the normal buffer (strip), so touch scrollback "should" work there. Before coding anything: build a repro matrix (iPhone Safari / iPad Safari / Android Chrome × claude / shell) on the current release. Plausible candidates if it does reproduce: auto-scroll-to-bottom fighting user scrolls (`_noteTerminalUserScroll`, ~line 2004), or they were in shell sessions on mobile too (then Bug B covers it). Ask mtiller on #205 for session mode + Codeman version if the matrix comes up clean.
## Invariants the implementation MUST respect
- Shift+wheel always scrolls local scrollback; the trackpad Shift-axis handling from #154 stays.
- The `terminalWheelLocalScrollback` opt-out setting keeps working (pins plain wheel to local).
- The viewport-at-bottom gate stays: once the user scrolled up locally, wheel stays local until they return to bottom.
- 40ms SGR coalescing: never send per-event writes to the server.
- Strip parity triangle: `session.ts` live strip ↔ `session-routes.ts` replay strip ↔ `_sessionUsesServerMouseStrip()` in the frontend. If you touch mode lists, update all three.
- Don't add `opencode`/`antigravity` to any strip/forward list; their TUI wheel behavior is unverified (documented at `_shouldForwardWheelToApp`).
- The chunk-boundary sequence carry in `_handleTerminalOutput` must not be weakened.
## Testing (per repo rules)
- `npm test -- test/<file>.test.ts` only; never bare `npm test`. New test ports 3150+, never 3000.
- Browser-test traps (documented in CLAUDE.md Testing): drive input/scroll through real events (`page.mouse.wheel`, real touch), not app internals; headless Chromium reports `isTouchDevice()` false even with `hasTouch: true`; assert on real state (xterm viewport position, `tmux -L codeman capture-pane`), not HTTP 200.
- Shell-mode E2E: create a throwaway shell session, `seq 1 500`, then (1) wheel up on desktop shows earlier lines, not history cycling; (2) touch-scroll on a phone shows earlier lines; (3) `vim` + `less` still enter/leave the alt screen cleanly; (4) Shift+drag still selects text.
- Firefox E2E: Playwright `firefox` project, wheel over a Claude session's finished output, assert viewport moved more than 1 line per notch.
- End-to-end against the REAL environment before claiming done (standing user rule). w1/w2/w3 tmux sessions are the user's live sessions: never send input to them; create your own throwaway session and DELETE it by exact id when done.
## Related observation (not a reported bug, worth a look while in there)
The `claude --version` probe that feeds the forwarding gate runs only for local and docker sessions (`src/session.ts:1490` gates `!this._remote`; docker handled at :1507). Remote Claude sessions therefore never get `cliVersion` and silently keep local-only wheel. Harmless (local scrollback works) but inconsistent; cheap to fix by probing over ssh, or document as intended.
## Rollout
1. Bug A (deltaMode) is small and independent: can ship alone as a patch.
2. Bug B (shell scrollback) is the headline fix for #205: patch or minor per COM flow.
3. After deploy + verification: comment on #205 (what was fixed, what needs their retest), then reply to the Reddit comment `p21x6ts` with the release version. Both reporters gave environment details; address them specifically.
+255
View File
@@ -0,0 +1,255 @@
# Scrollback issues: analysis and test evidence
Covers GitHub issue **#205** ("Scrollback in terminal not working", jonocodes, shell mode,
Android + macOS desktop) and the follow-up comment on it from **mtiller** (Firefox on macOS,
"scrolling backward to see agent output"). Related closed issue: **#154** (fixed in 1.3.3).
Status: **analysis only, nothing implemented.** Measured against the live 1.11.2 instance on
2026-08-06 with throwaway `zz-*` shell sessions (all deleted afterwards; the user's `w*`
sessions were never touched).
---
## TL;DR
Five distinct problems, not one. #205 is fully explained by finding 1; findings 2 and 3 are
independent and hit **every** mode including Claude, and are the likely substance of the
"similar issue" follow-up.
| # | Problem | Modes affected | Severity | Confirmed |
| - | ------- | -------------- | -------- | --------- |
| 1 | xterm parked in the **alternate buffer** for the whole session, so there is no scrollback at all and the wheel is translated into Up/Down arrow keys | `shell`, `opencode`, `antigravity` | High | Reproduced end to end |
| 2 | **Bursty output silently destroys a screenful** of the browser's scrollback and adds ~1 row | all | High | Measured |
| 3 | **Tab switch collapses scrollback** to roughly one screen (`full=1` fires once per page load) | all | Medium | Measured |
| 4 | `deltaMode` is never read, so Firefox scrolls ~4x slower per notch | all, Firefox | Low | Static, needs reporter data |
| 5 | **Remote SSH Claude cases get no `claude --version` probe**, so wheel forwarding silently stays off (residual #154) | `claude` + remote | Medium | Static |
---
## Finding 1: shell / opencode / antigravity are stuck in xterm's alternate buffer
### Root cause
The local tmux **client** (the `tmux attach` that node-pty spawns) emits `smcup` as its very
first bytes on attach. Captured from a real PTY:
```
b'\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2J\x1b[?12l\x1b[?25h\x1b[?1000l...'
^^^^^^^^^^ enter alternate screen ^^^^^ application cursor keys ON
```
`Session._handleTerminalOutput()` strips `\x1b[?1049h` from the live stream, but only when
`isAltScreenStripMode(mode)` is true, and that is `claude | codex | gemini` only
(`src/session.ts:179`). For `shell`, `opencode` and `antigravity` the sequence reaches the
browser verbatim and xterm switches to the alternate buffer, where:
1. `buffer.active.type === 'alternate'` and `baseY` is pinned at 0, so there is **no
scrollback to reach**. `terminal.scrollLines()` is a no-op, which is why touch scrolling
on Android "does nothing".
2. xterm's own wheel listener takes over. From the vendored bundle
(`src/web/public/vendor/xterm.min.js`):
```js
if (!this.buffer.hasScrollback) {
if (ev.deltaY === 0) return false;
if (coreMouseService.consumeWheelEvent(...) === 0) return this.cancel(ev, true);
const seq = ESC + (decPrivateModes.applicationCursorKeys ? 'O' : '[') + (ev.deltaY < 0 ? 'A' : 'B');
coreService.triggerDataEvent(seq, true);
return this.cancel(ev, true);
}
```
tmux also set `\x1b[?1h`, so the emitted sequence is `\x1bOA`, i.e. **Up arrow**, straight
into the shell's readline. That is exactly the reported "the mouse wheel scrolls back
through previous commands, like pressing up".
3. `cancel(ev, true)` calls `preventDefault()` **and `stopPropagation()`**, and xterm's
listener sits on `terminal.element` (a child of Codeman's container). So Codeman's own
container wheel handler, `_shouldForwardWheelToApp` and `_wheelScrollLines` included, is
**never reached** for these modes. That whole path is dead code for shell.
### Reproduction (live instance, real browser)
Create a shell session with the page already open, print 150 lines, then dispatch 8 wheel-up
events over `.xterm-screen`:
```
t+1500 after shell start {"type":"alternate","length":35,"baseY":0}
t+3000 after shell start {"type":"alternate","length":35,"baseY":0}
after 150 live lines {"type":"alternate","length":35,"baseY":0}
WHEEL on live shell: {"ptyBytes":["OA","OA","OA","OA",
"OA","OA","OA","OA"],
"before":0,"after":0,"type":"alternate"}
```
Both reported symptoms, one root cause.
### Why it looks intermittent
The alternate-screen sequence only ever reaches the browser through the **live stream at
attach**. Neither replay path carries it:
- `?full=1` returns `capture-pane` output (`source: mux-full-history`), verified 0 hits for
`\x1b[?1049h`.
- `?tail=` returns the visible pane frame (`source: mux-visible`), also 0 hits; the shell byte
buffer was empty in every probe.
- `_resetTerminalForReplay()` calls `terminal.reset()`, which returns xterm to the normal
buffer.
So: watching a shell from creation leaves you in the alternate buffer until you reload or
switch tabs, at which point it silently starts working again. Then the next PTY attach (a
restart, or the auto-reattach in `selectSession()`) puts you back.
### Is stripping safe for shell? Probably yes when tmux-backed, and the current code comment is wrong about why
`src/session.ts:1404` says *"shell must keep the alt screen for vim/less/htop"*. For a
**tmux-backed** shell that reasoning does not hold: tmux is a full terminal emulator and never
forwards a pane's alternate-screen toggles to its client, it repaints instead. Measured per
phase on a real attach:
| phase | bytes | `?1049h` | `?1049l` | `?47/1047` |
| ----- | ----: | -------: | -------: | ---------: |
| attach | 772 | **1** | 0 | 0 |
| `seq 1 60` echo | 1402 | 0 | 0 | 0 |
| `less` open / end / quit | 284 / 230 / 321 | 0 | 0 | 0 |
| `vim` open / quit | 2200 / 646 | 0 | 0 | 0 |
`vim` and `less` inside tmux emit **zero** alternate-screen sequences to the client.
The caveat that does matter: `startShell()` falls back to a **direct PTY with no tmux** when
mux creation fails (`src/session.ts:1961`, `this._useMux = false`). In that path the inner
app's own `?1049h` does reach xterm, and a blanket strip would break vim/less/htop for real.
Any fix has to be conditional on `_useMux`, which is known server-side.
Second caveat: stripping alone buys less than it looks like, because of finding 2. It fixes
the wheel (no more phantom Up arrows) and it makes the `full=1` replay reachable, but live
output still will not accumulate.
---
## Finding 2: bursty output silently overwrites a screenful of browser scrollback
Independent of the alternate buffer, and it hits Claude sessions too.
tmux decides per flush whether to emit real linefeeds (which push rows into the outer
terminal's scrollback) or to repaint the pane rectangle with cursor addressing (which
overwrites the visible rows in place). When output outpaces its flush interval it coalesces
into a repaint, and one screenful of the browser's history is **destroyed**.
Measured on one session, same page, `rows = 36`:
| step | `baseY` | rows containing SEED | BURST | SLOW |
| ---- | ------: | -------------------: | ----: | ---: |
| after `?full=1` replay (120 seeded lines) | 86 | 120 | 0 | 0 |
| after 60 lines emitted as fast as possible | **87** (+1) | **86** (-34) | 35 | 0 |
| after 60 lines at ~16/s (`sleep 0.06`) | **148** (+61) | 86 | 35 | 60 |
The burst added **one** row of scrollback and ate **34** rows of existing history. The slow
run behaved correctly. So "I printed a bunch of lines and now I cannot scroll back" reproduces
without the alternate buffer being involved at all, and it is rate dependent, which is exactly
the kind of thing that reads as random flakiness.
Consequence: the browser's scrollback is effectively frozen at whatever the last `?full=1`
replay produced, minus a screen per burst. tmux's own history is fine throughout
(`history_size` kept growing, `history-limit` 2000), so the data is never actually lost
server-side, it just never reaches the browser again until a reload.
---
## Finding 3: switching tabs collapses a session's scrollback
`_initialFullBufferLoad` is true for the **first buffer load after a page load only**
(`app.js:4374`). Everything after that uses `?tail=`, which returns byte history plus the
visible pane frame. Worse, the snapshot restore path deliberately throws away the restored
xterm snapshot (which does carry scrollback) and replaces it with that frame
(`app.js:4316-4328` plus `needsRewrite`).
Measured, switching away from session A and back:
```
A: initial full=1 load {"len":152,"baseY":116,"AAA":150}
A: after switch away and back {"len": 87,"baseY": 51,"AAA": 59}
```
150 lines of history down to 59. Note also that the page's single `full=1` is consumed by
whichever session auto-selects at load, so **every other tab starts life with one frame of
history**.
---
## Finding 4: `deltaMode` is never read (Firefox)
`grep -rn "deltaMode" src/web/public packages` returns nothing. `_wheelScrollLines()`
(`terminal-ui.js:2818`) treats `deltaY` as pixels unconditionally:
```js
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
```
Chrome/WebKit report `deltaMode: 0` with `deltaY` around 100 to 120 px per notch, so about 4
to 5 lines. Firefox reports `deltaMode: 1` (`DOM_DELTA_LINE`) with `deltaY` around 3, so
`Math.round(3/25) === 0` and the `|| ±1` fallback yields **1 line per notch**, roughly 4x
slower. In Claude mode the same value caps the forwarded SGR report at 1 tick per event
instead of 4, so the transcript crawls too.
This is sluggishness, not breakage, so it is a plausible but unproven contributor to the
mtiller report. No Firefox build is installed under `~/.cache/ms-playwright` (chromium and
webkit only), so this was not measured. Worth asking the reporter for `deltaMode` / `deltaY`
from a live wheel event before acting on it.
---
## Finding 5: remote SSH Claude cases still have no version probe
`src/session.ts:1490` deliberately skips the deterministic `claude --version` probe for
remote sessions and defers to the startup-banner scrape, which the same comment block
describes as unreliable ("newer Claude Code builds don't print the banner and resumed sessions
never show it"). That is precisely the condition #154 was filed for: `cliVersion` empty means
`_shouldForwardWheelToApp()` returns false, wheel forwarding is off, and the user is left with
local scrollback that (per finding 2) does not accumulate.
Local and Docker Claude sessions are fine; verified all 7 live sessions report
`cliVersion=2.1.223`, so the 1.3.3 fix is still working there.
---
## Candidate directions (not decided)
Roughly in order of value per unit of risk.
1. **Extend the alternate-screen strip to tmux-backed `shell` / `opencode` / `antigravity`.**
Gate on `_useMux` so the direct-PTY fallback keeps vim/less/htop working. Kills the phantom
Up arrows and makes replayed history reachable. `isAltScreenStripMode()` currently takes
only `mode`, so it would need the mux flag threaded in, and
`test/claude-scrollback-strip.test.ts:16-17` plus `test/antigravity-mode.test.ts:116` pin
the current answers and would need updating.
2. **Re-pull `?full=1` when the user scrolls to the top of the buffer.** Directly addresses
findings 2 and 3 with machinery that already exists and is already proven to return
complete history (200/200 lines in the probe). Needs a guard against refetch storms.
3. **Stop discarding the xterm snapshot on tab switch**, or request `full=1` on the first load
per session rather than per page. Cheaper partial fix for finding 3 alone.
4. **Read `ev.deltaMode`** in `_wheelScrollLines()` and normalise line/page deltas to lines.
Small, self-contained, worth doing regardless of whether it is mtiller's actual bug.
5. **Probe the CLI version over SSH for remote Claude cases**, mirroring the deferred
in-container probe that Docker cases already use.
Option 1 alone does not fix #205's "print a bunch of lines then scroll" complaint; that needs
2 as well.
## Reproduction assets
Scripts used, in the session scratchpad
(`/tmp/claude-1000/-home-arkon-default-claudeman/597ffc9f-.../scratchpad/`):
- `ptycap.py` / `ptycap2.py`: PTY-level capture of the tmux client stream, per phase counts of
alternate-screen and mouse-tracking sequences.
- `sim.mjs`: replays a captured stream through `@xterm/headless` with and without the strip.
- `browser-test*.mjs`: Playwright against the live instance, reports `buffer.active.type`,
`baseY`, row content and the exact bytes xterm sends to the PTY on a wheel event.
`@xterm/headless` was installed with `npm i --no-save`, so `package.json` and the lockfile are
untouched.
+1 -1
View File
@@ -489,7 +489,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini`, `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — and `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
+303
View File
@@ -0,0 +1,303 @@
# Terminal smart copy (Ctrl+C) plan
Issue: [#211](https://github.com/Ark0N/Codeman/issues/211) "Terminal: Ctrl+C should copy when text is selected (interrupt otherwise)".
Origin: r/selfhosted feedback, "Biggest stumbling block is apparent lack of copy-paste in the terminal."
Status: **implemented and shipped** on 2026-08-05 (this document is kept as the rationale record). It was first served as an isolated beta over Tailscale for manual sign-off, then landed. Section 2 is the research that shaped the design, sections 4 to 6 describe what was built.
---
## 1. What the issue asks for
- Text selected in the terminal + `Ctrl+C` -> copy the selection, toast, clear the selection, do NOT send the byte to the PTY.
- No selection + `Ctrl+C` -> unchanged, the interrupt (`0x03`) reaches the PTY.
- `Ctrl+Shift+C` as an explicit copy chord.
- The selection check must run before the shortcut registry dispatch so a rebind cannot cost the user their interrupt key.
- Paste is out of scope (it already works via `Ctrl+V`, which terminal-ui.js routes to the image/text paste trap).
## 2. Verified current behavior
### 2.1 xterm cancels the Ctrl+C keydown, so no copy can happen
`src/web/public/vendor/xterm.min.js` (xterm 6.x), `_keyDown`:
```js
_keyDown(x){ if(this._keyDownHandled=!1, this._keyDownSeen=!0,
this._customKeyEventHandler && this._customKeyEventHandler(x)===!1) return !1;
... evaluateKeyboardEvent(...) ... this.cancel(x) ... }
```
Two consequences that shape the design:
1. The custom handler runs **first**, before xterm evaluates the key. Returning `false` exits before `cancel(x)`, so returning `false` does **not** call `preventDefault()` for us.
2. When the handler returns `true`, xterm turns Ctrl+C into `0x03` and cancels the event, which is why the browser's own copy command never runs.
Probe (headless chromium against an isolated server on port 3174, selection active, real focus on `.xterm-helper-textarea`, synthetic Ctrl+C keydown):
```json
{ "hasSelection": true, "defaultPrevented": true, "dataSeen": ["\"\\u0003\""],
"clipboardAfter": "SENTINEL-BEFORE", "stillHasSelection": false }
```
So today: interrupt byte sent, clipboard untouched, and xterm drops the selection anyway. The last point matters, "copy then clear the selection" is not a behavior change in how the selection feels, it is what already happens on any keypress.
### 2.2 Why right-click Copy works today
xterm registers a `copy` listener on its root element that substitutes the selection text:
```js
this._register(addDisposableListener(this.element,"copy",(k=>{ this.hasSelection() && copyHandler(k,this._selectionService) })))
```
Second probe (port 3175, real `page.keyboard.press('Control+c')`, custom handler patched to return `false` for Ctrl+C without `preventDefault`):
```json
{ "dataSeen": [], "copyEvents": ["xterm-element"],
"clipboardAfter": "native-copy-probe-line\n...", "stillHasSelection": true }
```
So a "return false and let the browser copy" implementation would also work in Chromium. It is rejected below (section 3.3) because it gives no toast, does not clear the selection, and leans on per-browser behavior of the copy command when the focused element is xterm's empty helper textarea.
### 2.3 The document-level capture handler will not interfere
`setupEventListeners()` in `src/web/public/app.js:989` runs on document capture, before xterm's textarea listener. Its registry loop skips any entry whose action is not in the local `SHORTCUT_ACTIONS` map:
```js
if (shortcut.disabled || !shortcut.action) continue;
const action = SHORTCUT_ACTIONS[shortcut.action];
if (!action) continue;
```
This is exactly how `command-palette` already behaves: it is a full registry entry (rebindable and disableable in App Settings) whose dispatch happens in a dedicated, focus-aware gate rather than the generic loop. The new copy entry follows that pattern, so the capture handler falls through untouched and the terminal handler owns the decision.
### 2.4 Registry matching rules that constrain the bindings
`matchesShortcutEvent()` (`app.js:4890`):
- Ctrl and Cmd are interchangeable as the primary modifier, so a `['ctrl']` binding also matches Cmd+C on macOS. That is fine here: with a selection it copies (same result the native macOS path gives today), without one it falls through.
- Every other modifier must be declared exactly: `if (mods.includes('shift') !== !!e.shiftKey) return false`. So `Ctrl+Shift+C` needs its own binding, a plain `ctrl+c` binding will never swallow it.
- `binding.code` wins when present, otherwise `binding.key` is compared case-insensitively.
### 2.5 Where selection is actually possible
- The server strips mouse-tracking DECSETs for `claude`, `codex`, and `gemini` (`isAltScreenStripMode`, `src/session.ts:179`), which is why plain drag-select works in those tabs even though the TUI has mouse tracking on.
- `shell`, `opencode`, and `antigravity` keep mouse reporting, so xterm requires `Shift`+drag to force a selection there. Worth one line in the docs, it is not a code change.
- Touch devices deliberately disable selection entirely (`body.touch-device .terminal-container .xterm{user-select:none !important}`, `styles.css:3196`), and phones have no Ctrl key. This feature is desktop and hardware-keyboard only, with no mobile regression surface.
### 2.6 Helpers that already exist and should be reused
| Need | Existing code |
| --- | --- |
| Clipboard write with an HTTP-safe fallback | `_copyText(text)` in `app.js:1887` (Clipboard API, then hidden textarea + `execCommand`) |
| Toast | `showToast(message, type)` in `panels-ui.js:4385` |
| Translated string | `'Copied to clipboard'` already in `i18n.js:453` |
| Focus-aware chord gate to copy the shape of | `shouldOpenCommandPaletteFromShortcut(e)` in `panels-ui.js:285` |
| Buffer-wide copy (currently unreferenced) | `copyTerminal()` in `terminal-ui.js:2615` |
`_copyText` matters more than it looks: `install.sh`'s LAN option serves plain HTTP, where `navigator.clipboard` is undefined. The issue's suggested `navigator.clipboard.writeText` alone would silently do nothing for those users, the `execCommand` fallback covers them.
## 3. Design
### 3.1 Behavior
| Chord | Selection present | No selection |
| --- | --- | --- |
| `Ctrl+C` (and Cmd+C, per registry equivalence) | copy, toast, clear selection, swallow the key | fall through, xterm sends `0x03` (interrupt) |
| `Ctrl+Shift+C` | copy, toast, clear selection, swallow the key | swallow, no-op (see 3.2) |
| Shortcut disabled in App Settings | never copies, `Ctrl+C` is always the interrupt | unchanged |
| Rebound to another chord | that chord copies when a selection exists | plain `Ctrl+C` is always the interrupt |
### 3.2 Why `Ctrl+Shift+C` with no selection is swallowed rather than forwarded
Today `Ctrl+Shift+C` produces `0x03` as well (the shift is irrelevant to the control byte), so forwarding would be "no regression". But once the chord is advertised as *the explicit copy key*, letting it interrupt a running agent when the selection happens to be empty is a footgun with no upside. Swallowing costs nothing: a user who wants to interrupt has `Ctrl+C` right there.
The rule in code is "no selection and the matched chord had Shift -> swallow", not a hardcoded key check, so it stays correct under rebinds.
### 3.3 Why an explicit clipboard write rather than falling through to the native copy
Probe 2 showed the native path works in Chromium, but the explicit write is chosen because it:
- gives the "Copied to clipboard" toast, which is the discoverability half of the issue,
- clears the selection so a second `Ctrl+C` interrupts (the smart-copy contract),
- works on plain-HTTP LAN installs through `_copyText`'s `execCommand` fallback,
- does not depend on how each browser treats a copy command issued while an empty textarea has focus.
### 3.4 Why no new app setting
Per-shortcut enable/disable and rebinding already exist in App Settings -> Shortcuts and are driven by the registry. A user who wants "Ctrl+C is always interrupt" unchecks one box. Adding a `terminalSmartCopy` setting would duplicate that and would drag in the per-device vs synced decision (`displayKeys` + `.strict()` `SettingsUpdateSchema`) for no gain.
## 4. Code changes, file by file
### 4.1 `src/web/public/app.js`, registry entry
Add to `DEFAULT_SHORTCUTS` (after the `clear-terminal` entry, ~line 351) so the Terminal group stays together:
```js
{
id: 'copy-selection',
group: 'Terminal',
label: 'Copy Selection',
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js,
// deliberately NOT in SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults, which would cost the user the interrupt key.
action: 'copyTerminalSelection',
},
```
Match on `key`, not `code`. xterm decides what byte to emit from the produced character, so intercepting the physical `KeyC` on a layout where it does not produce "c" would diverge from what xterm would have sent.
The `action` string is required for App Settings to render the row as configurable (`configurable = !!shortcut.action && Array.isArray(shortcut.bindings)`, `settings-ui.js:2624`). Do **not** add `copyTerminalSelection` to `SHORTCUT_ACTIONS`.
### 4.2 `src/web/public/terminal-ui.js`, the gate
New prototype method, modeled on `shouldOpenCommandPaletteFromShortcut`:
```js
shouldCopyTerminalSelectionFromShortcut(ev) {
if (!ev || ev.type !== 'keydown') return false; // the handler also runs for keypress/keyup
if (!ev.ctrlKey && !ev.metaKey && !ev.altKey) return false; // hot path: plain typing exits here
const registryAvailable =
typeof this.getShortcutRegistry === 'function' && typeof this.matchesShortcutEvent === 'function';
const entry = registryAvailable
? this.getShortcutRegistry().find((s) => s.id === 'copy-selection')
: null;
if (entry) return !entry.disabled && this.matchesShortcutEvent(ev, entry);
return (ev.key || '').toLowerCase() === 'c' && !ev.altKey; // fallback for isolated harnesses
}
```
### 4.3 `src/web/public/terminal-ui.js`, the branch
Inside `attachCustomKeyEventHandler` (`terminal-ui.js:133`), after the command-palette gate and before the `Ctrl+V` branch:
```js
// Smart copy (#211): with a selection, Ctrl+C copies instead of sending ^C.
// With no selection it MUST fall through (return true, no preventDefault) or
// the interrupt key is lost. Ctrl+Shift+C is the explicit chord and never
// falls through: an "explicit copy" that interrupts the agent is a footgun.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
}
return true;
}
```
`preventDefault()` is explicit because returning `false` alone does not cancel the event (section 2.1), and without it the browser would run its own copy on top of ours.
### 4.4 `src/web/public/terminal-ui.js`, the copy action
```js
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const ok = await this._copyText(selection);
if (ok) {
this.terminal.clearSelection?.();
this.showToast('Copied to clipboard', 'success');
} else {
this.showToast('Failed to copy', 'error');
}
// _copyText's execCommand fallback focuses a temp textarea; restore the
// terminal (this.terminal.focus is the CJK-aware router, not xterm's raw focus).
this.terminal.focus();
return ok;
}
```
The selection text is captured **before** the first `await`, and `navigator.clipboard.writeText` is reached in the same task as the keydown, so user activation still holds.
### 4.5 `src/web/public/i18n.js`
`'Copied to clipboard'` exists. Add `'Failed to copy': '复制失败'` (the error path is new to this surface).
### 4.6 Documentation
| File | Change |
| --- | --- |
| `README.md` shortcut table (~line 648) | `\| `Ctrl/Cmd+C` \| Copy selection (interrupts when nothing is selected) \|` and a `Ctrl+Shift+C` row |
| `src/web/public/index.html` help modal, Terminal section (~line 641) | `<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection / Interrupt</div>` plus the Ctrl+Shift+C row. Keep the existing negative assertion in `help-modal-shortcuts.test.ts` in mind (it forbids `Ctrl+K`, `C` is fine) |
| `CLAUDE.md` "Keyboard shortcuts" line | add `Ctrl+C` (copy selection, else interrupt) and `Ctrl+Shift+C` |
| `docs/architecture-invariants.md` -> "Command palette and shortcut registry" | append the invariant: the no-selection path must return `true` without `preventDefault`, the branch is keydown-only, and `copyTerminalSelection` must stay out of `SHORTCUT_ACTIONS` |
The shortcut overlay (`Ctrl+?`) and App Settings -> Shortcuts are registry-driven and pick the entry up with no edit.
## 5. Edge cases and risks
| Case | Handling |
| --- | --- |
| Handler also fires for `keypress`/`keyup` | gated on `ev.type === 'keydown'`. xterm's `_keyPress` bails on ctrl combos anyway, so no stray byte |
| CJK IME composing | the existing `isComposing || keyCode === 229` guard is the first line of the handler and stays first |
| Local echo overlay has unsent `pendingText` | the copy branch returns before `onData`, so `pendingText`, flushed offsets and the durable input queue are untouched. The no-selection path is byte-identical to today, including the "control char flushes buffered text then sends `0x03`" logic at `terminal-ui.js:895` |
| Plain HTTP (LAN install) | `_copyText` falls back to `execCommand`, then focus is restored |
| Clipboard write rejected (permissions policy, no gesture) | error toast, right-click Copy still available |
| Whitespace-only or empty selection | `getSelection()` empty string is treated as "no selection", so Ctrl+C still interrupts |
| macOS Cmd+C | registry treats ctrl/meta as interchangeable, so with a selection it takes our path (same visible result as today's native copy), without one it falls through |
| Chrome/Firefox `Ctrl+Shift+C` is the devtools inspect chord | browser-level and may still toggle devtools, our copy runs regardless. Document as a caveat, `Ctrl+C` is the primary path |
| Selection in a tab whose TUI owns the mouse (`shell`/`opencode`/`antigravity`) | unchanged, `Shift`+drag selects, then Ctrl+C copies |
| Web tab (iframe dashboard) focused | xterm handler never runs, browser-native copy inside the iframe |
| Teammate/subagent terminals (`panels-ui.js:2268`, `onData` wired) | same limitation exists there, out of scope for this PR (section 8) |
## 6. Test plan
New file `test/terminal-copy-selection.test.ts` (node env, `vm` harness in the style of `test/command-palette-ui.test.ts`), covering `shouldCopyTerminalSelectionFromShortcut` in isolation:
1. Ctrl+C keydown -> true, keyup/keypress of the same chord -> false.
2. Ctrl+Shift+C -> true, plain `c` -> false, Ctrl+K -> false.
3. Registry entry `disabled: true` -> false for every chord.
4. Rebound entry (for example Alt+Y) -> true for the rebind, false for Ctrl+C.
5. Missing registry (harness without `getShortcutRegistry`) -> falls back to the `c` check.
Static assertions appended to `test/keyboard-shortcuts.test.ts` (this suite already pins the xterm-handler chokepoint):
6. `DEFAULT_SHORTCUTS` contains `id: 'copy-selection'` and `SHORTCUT_ACTIONS` does **not** contain `copyTerminalSelection` (the interrupt-safety invariant).
7. `terminal-ui.js` contains the `shouldCopyTerminalSelectionFromShortcut` branch and a `return true` no-selection fall-through.
8. README + help modal rows exist (mirrors the existing palette/Alt-nav doc assertions).
`test/help-modal-shortcuts.test.ts`: add `expectShortcut(helpModal, ['Ctrl', 'C'], 'Copy Selection')`.
New browser test `test/terminal-copy-shortcut.test.ts` (Playwright, port **3174**, free per a scan of `test/`), following `test/webgl-fallback.test.ts`: boot `WebServer`, grant `clipboard-read`/`clipboard-write`, `terminal.write()` a known line, `selectLines()`, real `page.keyboard.press('Control+c')`, then assert clipboard content, empty `onData` capture, cleared selection and the toast. Second case: no selection, assert `onData` saw `\u0003` and the clipboard is unchanged.
Per repo convention, browser suites are excluded from CI, so add the filename to the exclude list in `config/vitest.ci.config.ts` and run it locally.
Regression runs: `npm test -- test/keyboard-shortcuts.test.ts`, `test/help-modal-shortcuts.test.ts`, `test/command-palette-ui.test.ts`, `test/input-send-order.test.ts`, then `npm run test:ci`.
## 7. Manual verification before COM (CLAUDE.md rule)
Against a throwaway session on the live instance (`curl -sk https://localhost:3000/...`, never w1/w2/w3):
1. Select output with the mouse, press Ctrl+C, confirm the toast, paste elsewhere, confirm the agent did not stop.
2. Press Ctrl+C again with nothing selected, confirm the agent interrupts.
3. Type a few characters with local echo on (phone or `localEchoEnabled` forced), press Ctrl+C with no selection, confirm buffered text plus interrupt behave as before.
4. Uncheck the shortcut in App Settings -> Shortcuts, confirm Ctrl+C always interrupts even with a selection.
5. Rebind it, confirm the new chord copies and Ctrl+C reverts to pure interrupt.
6. Repeat 1 and 2 in an `opencode` or `shell` tab using Shift+drag to select.
7. Load over plain HTTP (`--host` LAN or `http://127.0.0.1:<port>`) and confirm the `execCommand` fallback copies and focus returns to the terminal.
8. Mobile smoke: confirm nothing changed (selection is CSS-disabled, no Ctrl key).
## 8. Out of scope, follow-ups worth filing separately
- **Teammate/subagent terminals** (`panels-ui.js:2268`) have the same blocked-copy problem. One `attachCustomKeyEventHandler` reusing `copyTerminalSelection` would fix them, but it touches a different surface and deserves its own change.
- **A mobile copy affordance.** Selection is disabled on touch, so phones still cannot copy terminal text. The unreferenced `copyTerminal()` (whole buffer) plus a keyboard-accessory "Copy" button would be the cheapest answer.
- **Right-click context menu** with Copy/Paste, better discoverability than any chord, but a bigger UI surface.
- **`copyTerminal()` cleanup**: it uses raw `navigator.clipboard` rather than `_copyText`, so it would fail on plain HTTP if ever wired up.
## 9. PR mechanics
- Branch off `master` (verify with `git branch --show-current`, the tree is shared), stage explicit paths only.
- Files touched: `src/web/public/app.js`, `src/web/public/terminal-ui.js`, `src/web/public/i18n.js`, `src/web/public/index.html`, `README.md`, `CLAUDE.md`, `docs/architecture-invariants.md`, `docs/terminal-copy-shortcut-plan.md`, three test files, `config/vitest.ci.config.ts`.
- `index.html`, `app.js` and `terminal-ui.js` are `.prettierignore`d hand-formatted assets, match the surrounding style by hand. `npm run check:public-assets` and `npm run check:frontend-syntax` are the guards.
- No changeset in this PR: a merged, unconsumed changeset turns the Release workflow red until the next COM, and the COM flow writes release notes covering everything since the last tag (current version is 1.10.0).
- Close #211 from the PR body.
Rough size: about 60 lines of product code, most of the work is the tests and the four documentation surfaces.
+201
View File
@@ -0,0 +1,201 @@
<!-- Design doc drafted 2026-07-28 from WWDC26 session 224 research. STATUS: PLANNED, NOT IMPLEMENTED. Blocked on macOS 27 "Golden Gate" (beta now, GA expected fall 2026). -->
# VM Cases (macOS Virtualization framework), Implementation Plan
## Status
PLANNED, nothing implemented. This is the design + phased execution plan for a native-macOS VM isolation tier for cases ("the VM subsystem"), modeled on Docker cases (`docs/docker-cases-plan.md`). Testbed prerequisite: a macOS 27 host (see Section 8).
**⚠ DESIGN DIRECTION (owner, 2026-07-29): the subsystem is GUI-first.** Users want real macOS desktops, not headless SSH machines. Guests may be macOS (GUI-only in practice) or Linux (GUI or headless). Key decision 3 below carries the full consequences; anything in this doc that reads as "Linux-first / headless-first" predates this and has been revised.
**2026-07-29: Phase 0 substantially validated on the beta testbed; full Apple-stack reference now lives in [`docs/vm-subsystem-apple-stack.md`](vm-subsystem-apple-stack.md)** (API surfaces, beta bugs, our empirical results, and design implications). Plan-relevant corrections from that work: vmnet's topology/port-forwarding APIs are macOS 26 (only the loopback fix is 27); guest provisioning is macOS-guests-only (Linux stays cloud-init, proven working); DiskImageKit has NO flatten/merge, so the `export` subcommand ships the layer chain (or flattens in-guest) instead of flattening; seed ISOs are base-build-time only, never attached at case runtime; per-case EFI variable stores are mandatory; guest health checks read DHCP leases, never serial/ping.
## 1. Context and motivation
WWDC 2026 session 224 ("Expand the Capabilities of your Virtualization App", https://developer.apple.com/videos/play/wwdc2026/224/) shipped the missing pieces for programmatic, fleet-style VM management on macOS:
- **`VZMacGuestProvisioningOptions`**: automated first-boot setup of a macOS guest (user account, auto-login, SSH enabled) with zero interactive setup.
- **DiskImageKit**: stacked disk images on the Apple Sparse Image Format (ASIF): a read-only base layer plus cheap per-VM cache/overlay layers. Direct analog of Docker image layers + writable container layer.
- **vmnet framework**: custom network topologies and port forwarding from the host process.
- **`VZCustomVirtioDevice`**: custom low-latency host<->guest channels (Linux guests).
- **AccessoryAccess**: USB passthrough (not relevant to Codeman, out of scope).
Codeman's isolation story today is Docker cases. On macOS, Docker means Docker Desktop / a Linux VM anyway, with weaker fidelity and a heavyweight dependency. The Virtualization framework gives hardware-virtualized per-case sandboxes natively, with a layered-image story that mirrors what `scripts/build-agent-image.mjs` does for Docker. This is the premium native-macOS tier ON TOP of Docker cases, never a replacement (Docker remains the cross-platform story; the Linux prod box cannot use any of this).
## 2. Platform reality (hard constraints)
| Constraint | Detail |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Host OS | macOS 27 "Golden Gate" required for the new APIs (dev beta since 2026-06-08, public beta since 2026-07-13, GA expected fall 2026) |
| Host hardware | Apple Silicon only (macOS 27 dropped Intel). Testbed: the owner's dedicated MacBook (Section 8); the M4 Mac mini (macOS 26.4, runs the second Codeman install) stays on stable + untouched |
| Guest provisioning | `VZMacGuestProvisioningOptions` needs macOS 27 on BOTH host and guest. Linux guests provision via cloud-init instead |
| macOS guest concurrency | **Hard kernel cap: 2 concurrent macOS VMs per host. MEASURED on 27 beta 4 (2026-07-29), not inferred**: the 3rd VM is refused instantly with `VZErrorDomain` code 6 while 39% of RAM is free, so more hardware does NOT raise it. Since macOS GUI guests are the headline use case, this is a real product capacity limit to schedule around and surface in the UI. Linux guests are uncapped (resource-bound only) |
| Language | Virtualization framework is Swift/ObjC only; Node cannot call it. Requires a Swift helper binary (Key decision 2) |
| Entitlement | Host process needs `com.apple.security.virtualization`. Fine for a locally built dev binary; distribution needs signing thought (Section 9) |
| Nested virtualization | Linux-guest-only on M3+. A macOS 27 VM cannot dependably host its own guests, so the host-side APIs must be tested on bare-metal 27 (dual-boot) |
| CI | Cannot run in CI (needs beta macOS on Apple Silicon). Same answer as tmux/docker: no-op all VM IO under `VITEST`, unit-test the pure parts |
## 3. Goal and user stories
Add "VM cases" to Codeman: a case can point at a per-case virtual machine on a macOS host, and any CLI backend runs inside it over the existing remote-SSH session machinery. A LOCATION OVERLAY on cases, exactly like remote-SSH and Docker cases, NEVER a sixth `SessionMode`.
- As a Mac user, I link a case to a VM so an autonomous run executes behind a hardware virtualization boundary (stronger than Docker's shared kernel) while file viewing, transcripts, and hooks keep working.
- Per-case VMs are instant and cheap: a shared provisioned base image plus a per-case overlay, not a full image copy per case.
- Killing a session kills only its in-guest tmux; the VM stays up while sibling sessions remain; case delete tears the VM down.
- I export a case's VM overlay as a portable artifact (mirror of `docker-exports/`), secrets excluded.
- On a non-mac host, or a Mac without the helper, the feature is invisible: zero UI, zero probes, zero errors.
Non-goals for the MVP: USB passthrough, custom Virtio channels (Phase 3 candidate), macOS-guest fleets (capped at 2 anyway), Kubernetes-style orchestration, Intel Macs.
## 4. Architecture
```
Codeman (Node, unchanged session layer)
| JSON over stdout (same pattern as shelling out to docker/tmux)
v
codeman-vm (Swift package: CLI + per-VM GUI runner app in the console session)
| Virtualization / DiskImageKit / vmnet
v
per-case VM (macOS or Linux)
|-- GUI mode: VZVirtualMachineView in a window --> guest screen sharing --> browser (noVNC)
|-- shell: SSH on vmnet IP --> existing remote-SSH tmux machinery
^ VirtioFS: host case dir mounted at the SAME absolute path
```
Note the runner is a **GUI app in the console user's session**, not a detached daemon: a daemon-launched VM cannot render, which is fatal for macOS guests and for Linux desktop cases.
### Key decision 1: location overlay, not a mode
Identical reasoning to Docker/remote-SSH (see CLAUDE.md): the session layer, respawn, Ralph, recovery, and quick-start plumbing all stay untouched. `SessionMode` stays five-valued. State mirrors the Docker pair: `~/.codeman/vm-hosts.json` + `vm-cases.json`, new `src/vm-hosts.ts` with the storage + pure helpers split.
### Key decision 2: Swift helper CLI (`codeman-vm`)
The framework is Swift-only, so all VM work lives in a SwiftPM package (`packages/codeman-vm/`), a CLI with a stable JSON contract:
- `create-base --guest linux|macos`: build the shared base image. Linux: boot an arm64 cloud image with EFI + cloud-init, install Node 22 + tmux + the four CLIs (same inventory as `docker/agent.Dockerfile`), seal as base ASIF. macOS: IPSW restore + `VZMacGuestProvisioningOptions` (agent user, SSH on), then **desktop-readiness baking**, which is mandatory for GUI guests: suppress the per-user first-login assistant (`com.apple.SetupAssistant` keys + the User Template), enable auto-login (`autoLoginUser` + `/etc/kcpassword`), disable screensaver/lock/display-sleep, and set a static wallpaper (animated "aerials" wallpaper is unusable over remote display). ⚠ Use RAW (not ASIF) for macOS guest disks until the beta's macOS-guest space-reclamation bug is fixed.
- `create <case>`: DiskImageKit stacked image: shared read-only base + fresh per-case overlay. Near-instant, space-efficient.
- `start <case>` / `stop` / `status` / `ip`: lifecycle + vmnet NAT; `ip` reports the guest SSH endpoint.
- `export <case>` / `import`: flatten overlay + workspace tar + manifest, credentials excluded (mirror of docker-export).
A VM dies with its owning process, so `start` spawns a DETACHED per-VM runner process (analog of the detached `scripts/self-update.sh` trick) rather than a monolithic daemon; `status` talks to it over a unix socket in the instance data dir (`dataPath()`, never a hardcoded `~/.codeman` path).
### Key decision 3: multi-guest, and GUI is a first-class mode (REVISED 2026-07-29 by the repo owner)
The subsystem supports both macOS and Linux guests, and a guest runs in one of two **display modes**:
| | macOS guest | Linux guest |
| --- | --- | --- |
| **GUI mode** | **the point of the feature**; a real macOS desktop. Mandatory: nothing renders without an attached `VZVirtualMachineView` in an unlocked host session | supported (EFI + virtio-gpu framebuffer) for desktop Linux cases |
| **Headless mode** | not offered: a macOS guest with no view renders nothing, so a "headless macOS desktop" is a contradiction. SSH-only macOS is possible but is not what this feature is for | supported and cheap; the natural mode for agent/CI work, driven over SSH |
Consequences that flow from GUI being first-class:
- VM processes are **GUI apps in the console user's session** (LaunchAgent / `launchctl asuser`), never daemons. A daemon-launched VM cannot render.
- **The host is part of the product surface**: it must auto-login, never lock, never sleep, and keep a live WindowServer. Host lock == every VM's screen goes black, so the screen lock is effectively a global kill switch for every VM display on the machine. The product must own these host settings rather than treat them as user preference.
- **FileVault conflicts with unattended GUI hosting** and the trade-off must be a deliberate choice: FileVault disables auto-login, so a full-disk-encrypted host needs a human at a keyboard (or a remote screen-sharing session) after every reboot before any VM can render. Options are (a) FileVault on, accept manual login per boot, (b) FileVault off on a dedicated VM host so it boots straight into a rendering session, or (c) FileVault on plus a remote-unlock runbook. Codeman should detect the state and tell the user which one they are in instead of silently serving black screens.
- **Guests must be desktop-ready, not just booted**: auto-login, no screensaver/lock, and the per-user first-login assistant pre-suppressed at base-image time (`com.apple.SetupAssistant` keys, plus the User Template so later accounts inherit it). Otherwise the user connects to a login prompt or a setup wizard, which is exactly what happened during the first hands-on run.
- **Capacity is capped for macOS**: at most 2 concurrent macOS VMs per host, confirmed by our own test on 27 beta 4 (3rd refused with `VZErrorDomain` 6 at 39% free RAM; it is a kernel quota, so bigger hardware does not help). Scheduling must queue or evict beyond 2, the UI must explain why, and the scheduler should tolerate the acknowledged slot-leak bug (a slot occupied with nothing running, host-reboot to clear). Linux guests are uncapped and bounded only by host resources, which is the lever for scaling case counts on one machine.
- **Access is via the guest's own screen**, viewable in a browser through the noVNC chain (see `docs/vm-subsystem-apple-stack.md` §8), so no client-version or client-install requirements land on the user.
Provisioning per guest type: `VZMacGuestProvisioningOptions` for macOS (needs 27-on-27, first-boot-only, and does NOT skip the per-user wizard), cloud-init NoCloud seed ISO for Linux (proven working).
### Key decision 3b: the GUI VM host profile, and supervision that catches black screens
GUI hosting only works if the host is configured for it and supervised. This profile was derived the hard way on the testbed (prototyped there 2026-07-30) and should be what `codeman-vm` installs and verifies:
**Host profile** (the product should own these, not leave them to preference):
1. **No login barrier.** Either FileVault off + auto-login (a dedicated VM host boots straight into a rendering session, fully unattended), or FileVault on and remote reboots done with `sudo fdesetup authrestart`, where the pre-boot unlock *is* the login so the machine returns already logged in with encryption intact. **`authrestart` is VERIFIED on the testbed (2026-07-30): the host rebooted remotely and came back with a live logged-in console session, FileVault still enabled, no password prompt** — this is the recommended pattern for an encrypted GUI VM host. Plain reboots on a FileVault host always need a human, so Codeman should detect that combination and warn instead of serving black screens.
2. **Never lock**: lock policy off (needs the account password, so it is a setup step, not a scriptable one) plus `caffeinate -d -i -m -u` re-armed per session.
3. **Never sleep**: `pmset -a sleep 0 displaysleep 0 disablesleep 1`; a physical display is NOT required (a lid-closed laptop renders fine, only an unlocked session matters). Note OS updates reset these.
4. **Session-independent control plane**: run VPN/remote access as a system service, never a session app, and keep the access chain (forwards, VNC proxies, web endpoints) in LaunchDaemons so a session restart cannot sever operator access.
**Supervision** must be a **root LaunchDaemon**, not a user LaunchAgent. This is the load-bearing detail: a user agent cannot launch a GUI app into the Aqua session, so its restart attempts fail *silently* (the child dies instantly, leaving an empty log while the supervisor cheerfully reports success). A root daemon can, via `launchctl asuser <uid> sudo -u <user> …`, and those launches persist. Prototyped and verified on the testbed 2026-07-30; a working supervisor runs on a short interval and:
- Restarts the runner when the process is gone **or when its log shows `WindowServer event port death`**, which means it is permanently blind while still looking alive.
- Defers restarts while the console is at the login window, and launches into whichever session actually exists (resolve the console user with `stat -f %Su /dev/console`, never a hardcoded one).
- Re-points the guest port-forward whenever the guest's NAT lease changes, which happens on **every guest boot** under plain NAT. A vmnet DHCP reservation for a stable per-case IP is the better long-term answer.
- **Re-applies host power settings**, because `pmset -a disablesleep 1` does NOT survive a reboot (caught on the supervisor's first run after a real reboot) and OS updates reset it too.
- Re-arms the keep-awake helper, which dies with its session.
- Ideally also samples the guest framebuffer for non-black content, since a black screen is the one symptom common to every failure mode here.
`pgrep` alone is worthless for health: every failure mode in this session presented as a healthy process.
### Key decision 4: sessions ride the existing remote-SSH machinery
A provisioned guest is literally an SSH host on a vmnet IP. Session launch = the remote-SSH flow with the host swapped in: durable remote `tmux -L codeman-remote`, session names failing `SAFE_MUX_NAME_PATTERN` on purpose, EVERY ssh command line through `buildSshConnectionArgs()` (command-injection invariant), run flows through `POST /api/quick-start` (never `POST /api/sessions`, which stat-validates `workingDir` locally). What is genuinely new is only lifecycle (create/start/stop/export) and the vm-hosts/vm-cases overlay state.
### Key decision 5: workspace via VirtioFS at the same absolute path
Mirror the Docker bind-mount invariant: the case workspace is a real host directory shared into the guest via VirtioFS and mounted at the SAME absolute path. That keeps file-routes/watchers on real host bytes and makes the in-guest transcript projHash match the host. Without this, transcripts/attachments/file viewer all silently degrade.
### Key decision 6: credentials seeded, hooks bridged
- Credentials are SEEDED (read-only share, copied into the guest once at create), never shared read-write, and excluded from exports: byte-for-byte the Docker cases rule and rationale.
- Hooks: on the loopback-only prod bind a guest cannot reach `127.0.0.1:3000`. Mirror `CODEMAN_DOCKER_BRIDGE_HOOKS` with a `CODEMAN_VM_BRIDGE_HOOKS` opt-in listener on the vmnet gateway IP; otherwise idle detection falls back to output-based, same as Docker.
### Key decision 7: drift and teardown copy Docker semantics verbatim
Config hash label on the VM (guest type, cpu/mem, share list); a drifted launch is REFUSED, never silently launched stale. One VM per case shared by all sessions; session kill = in-guest tmux kill only; case delete = stop + remove overlay; instance-scoped boot reaper for orphaned runner processes.
## 5. Implementation phases
**Phase 0, testbed (no repo code):** dedicated MacBook on the macOS 27 beta, remotely accessible over the tailnet (setup protocol in Section 8), Xcode 27 beta, then a throwaway Swift script proving the loop: create base -> overlay -> boot -> ssh in. This validates 80% of the design before any Codeman code.
**Phase 1, `codeman-vm` helper:** SwiftPM package, the six subcommands above, JSON contract doc, detached runner + unix-socket status, Linux base image build. Deliverable is testable entirely without Codeman.
**Phase 2, Codeman integration:** types (`VmHost`/`VmCase`/`SessionVm`), `src/vm-hosts.ts` (+ pure helpers: config hash, arg building, endpoint parsing), Zod schemas, `case-routes` link/unlink + listing, `quick-start` vm branch reusing the remote-SSH launch path, `Session` threading + recovery round-trip, `VITEST` no-op layer, unit tests. Feature-detect: darwin + arm64 + helper binary present, else invisible.
**Phase 3, polish:** export/import UI, frontend Create Case "VM" tab + case-picker labels, SSE `vm:*` events, macOS-guest opt-in with cap surfaced, custom-Virtio input channel exploration, CLAUDE.md Key Pattern + `docs/vm-cases.md` + COM.
## 6. Testing
- Pure helpers unit-tested (ports pattern from `docker-hosts.ts`: 26 tests there, aim similar).
- All helper-invoking IO no-ops under `VITEST` (the `IS_TEST_MODE` pattern in `tmux-manager.ts`).
- End-to-end verification happens ON the beta MacBook, per the always-end-to-end rule: real base build, real per-case overlay boot, real quick-start into the guest, workspace round-trip through VirtioFS, session-delete keeps VM up, case-delete removes it.
- CI never runs the real path; the static guards are type-level + unit-level only.
## 7. Risks
1. **Beta API churn**: everything here targets beta SDKs; symbol/behavior changes are likely before fall GA. Mitigation: Phase 0/1 are throwaway-tolerant; no Codeman-side commitment until the helper contract survives a beta cycle.
2. **New artifact class**: Codeman ships pure TypeScript today; a Swift binary changes build/distribution (build-on-install via `xcrun swift build` on macs with Xcode CLT? prebuilt signed binary per release?). Needs an owner decision; local dev build is fine for the whole beta period.
3. **Entitlement/signing**: `com.apple.security.virtualization` is trivial for local dev, real for distribution.
4. **Adoption gating**: users need macOS 27 + Apple Silicon for months after GA. Docker cases remain the default recommendation; VM cases ship dark (feature-detected) with zero cost to everyone else.
## 8. Beta testbed plan: dedicated MacBook (actionable now)
Testbed is a dedicated MacBook the owner sacrifices to the beta (after a full backup). This supersedes the earlier dual-boot-the-Mini idea (git history has it): a dedicated machine means no OS-switching, no downtime for the Mini's live Codeman, and no FileVault pre-boot headaches.
**Sequencing rule that makes it headless: configure ALL remote access on the CURRENT macOS first, THEN upgrade in place.** An in-place beta upgrade preserves Remote Login, Tailscale, user accounts, and auto-login, so there is no Setup Assistant and no post-install physical step. (A fresh install would boot into GUI-only Setup Assistant with no SSH, which on a headless box is a dead end.)
Confirmed hardware (2026-07-28): MacBook, M3, 16 GB RAM, 256 GB disk with ~100 GB free. Verdict: green. M3 = eligible + nested-virt capable; 16 GB = host + 2-3 concurrent Linux guests (macOS guest = one at a time); 100 GB = fits with discipline: install Xcode 27 beta with the macOS platform only (skipping iOS/watchOS/tvOS simulators saves 15-20 GB), and defer any macOS guest base (~30 GB) to an external SSD or until actually needed. Linux guests + sparse ASIF overlays are the comfortable path.
### Pre-upgrade checklist (owner, physical, once)
1. Full backup (Time Machine or clone); the machine should be considered beta-only afterwards.
2. Tailscale: install, sign into the tailnet, confirm it appears in `tailscale status` from another node.
3. System Settings -> General -> Sharing: **Remote Login ON** (SSH) and **Screen Sharing ON** (for the rare GUI-only moments: Xcode license, Apple Account dialogs).
4. **FileVault stays ON** (owner decision 2026-07-28, security over convenience). Consequences: auto-login is unavailable, but FileVault's pre-boot unlock doubles as login, so an unlocked boot still lands in a live GUI session; planned remote reboots go through `sudo fdesetup authrestart` (unlocks for exactly one restart); an UNPLANNED reboot (beta kernel panic, battery drain) parks the machine at the pre-boot screen, no SSH/Tailscale, until the password is typed physically. If the testbed goes silent, suspect this first. Keep it on AC so the battery absorbs power blips.
5. Beta enrollment (manual): sign into the Apple Account in System Settings; System Settings -> General -> Software Update -> **Beta Updates** -> select the **macOS 27 Developer Beta** (preferred: framework fixes land weeks earlier than public beta; free since 2023 after accepting the agreement once at developer.apple.com; public-beta alternative: enroll at beta.apple.com). Then run the offered upgrade: plugged in, lid open, trusted network.
6. Send over: tailnet name/IP, username, and a first-login password (key install + lockdown happens remotely right after).
### Post-upgrade setup (remote, over the tailnet)
1. Verify: `sw_vers` reports 27.x, SSH reachable.
2. Server-ize the laptop: `sudo pmset -a sleep 0 disksleep 0 disablesleep 1` (lid-closed operation without an external display), `womp 1` (wake on network), `sudo systemsetup -setrestartpowerfailure on`. Keep on AC power.
3. Install the controlling host's SSH key, then disable password auth.
4. Xcode 27 beta install (the one step needing the owner's Apple Account sign-in once, doable via Screen Sharing from anywhere); `xcode-select`, license accept, verify `swift --version` + the 27 SDK (`xcrun --show-sdk-version`).
5. Phase 0 prototype loop, all remote from here: Linux guest base image (no 27-on-27 provisioning dependency), DiskImageKit overlay, boot, vmnet NAT, ssh into the guest, run `claude --version` inside.
6. Only after that loop works: start Phase 1 in `packages/codeman-vm/`.
## 9. Open decisions (owner)
1. Linux base distro/image for the default guest (proposal: Ubuntu 24.04 arm64 cloud image, matching the docker agent image's userland).
2. Helper distribution for GA: build-on-install vs prebuilt signed binary vs "bring your own Xcode".
3. Ship dark behind `CODEMAN_VM_CASES=1` for the first release, or feature-detect only?
4. Export format parity with docker-exports (one manifest schema for both?).
## References
- Session 224: https://developer.apple.com/videos/play/wwdc2026/224/
- Fleet-angle writeup: https://bitrise.io/blog/post/wwdc26-the-virtualization-framework-updates-that-matter-for-large-mac-fleets
- Beta timeline: https://www.macworld.com/article/3189014/apple-july-2026-ios-ipados-macos-27-public-betas-tv-arcade-releases.html
- Internal analogs: `docs/docker-cases-plan.md` (architecture template), `docs/remote-sessions.md` (session transport), `docs/architecture-invariants.md#docker-cases`
+281
View File
@@ -0,0 +1,281 @@
<!-- Reference doc for the VM subsystem (Codeman VM cases). Compiled 2026-07-29 from: Apple DocC JSON backend, macOS 27 beta 4 SDK on the testbed, a multi-source web research sweep, and hands-on prototyping on a MacBook Air M3 running macOS 27.0 beta (26A5388g). Companion to vm-cases-plan.md (the Codeman integration plan). -->
# The VM Subsystem: Apple Virtualization Stack Reference (macOS 27 "Golden Gate")
"VM subsystem" is the working name for Codeman's native-macOS VM isolation tier and everything under it. This document is the single place for what the Apple stack actually provides, what we have verified ourselves on the beta, and what is known-broken. The Codeman-side design lives in `docs/vm-cases-plan.md`.
**Research method note:** Apple's HTML doc pages are JS-rendered and come back empty to fetchers. The working route is the DocC JSON backend: `https://developer.apple.com/tutorials/data/documentation/<path>.json` (page content) and `https://developer.apple.com/tutorials/data/index/<framework>` (full symbol tree with per-symbol `beta` flags). Everything below marked "Apple docs" was parsed from that backend directly.
## 1. Component map and minimum OS versions
| Component | What it is | Min host OS | Notes |
| --- | --- | --- | --- |
| Virtualization.framework core | VMs, EFI/Linux boot, virtio devices, VirtioFS | macOS 11-13 era | Unchanged basics; our prototype uses nothing newer than macOS 13 APIs except the DiskImageKit bridge |
| **DiskImageKit** | ASIF + raw disk images, layered stacks | **macOS 27** | Swift-only, no ObjC headers. Section 2 |
| **Guest provisioning** | First-boot account/SSH setup for macOS guests | **macOS 27 host AND guest** | Mac guests only as of beta 4. Section 3 |
| vmnet topology/port-forward/DHCP APIs | Custom networks, port forwarding | **macOS 26** (NOT 27) | 27 adds exactly one fix: loopback port forwarding. Section 4 |
| `VZVmnetNetworkDeviceAttachment` | In-process vmnet attach | macOS 26 | |
| **`VZCustomVirtioDevice`** family | Custom paravirt devices | **macOS 27** | Linux guests only, custom guest driver required. Section 5 |
| AccessoryAccess (USB passthrough) | USB claim + attach to VMs | macOS 27 | Requires paid-team provisioning profile, Dock app. Out of scope for Codeman. Section 6 |
Corrections to the WWDC-session framing we started with: vmnet's topology family is a macOS 26 story (129 symbols, zero beta-flagged in 27); provisioning does NOT currently extend beyond macOS guests despite the generic-looking `VZGuestProvisioningOptions` base class; DiskImageKit has no attach/mount API at all (it is a file-format library that hands `DiskImage` objects to Virtualization, no `/dev/diskN`, no root needed, no entitlement documented).
## 2. DiskImageKit (macOS 27, Swift-only)
Public framework, `/System/Library/Frameworks/DiskImageKit.framework`. No ObjC headers; the API surface lives in the `.swiftinterface`. Verified present in the CLT 27 beta 4 SDK, and our prototype compiled against it with plain `swiftc` on the first attempt.
### API surface (complete as of beta 4)
```swift
class DiskImage {
convenience init(creating: some DiskImage.CreationConfiguration) throws
convenience init(opening: some OpenConfigurationProtocol) throws
func appending(any DiskImage.CreationConfiguration & DiskImage.StackableLayer) throws -> any StackedImage
func appending(consuming DiskImage) throws -> any StackedImage // reattach an existing layer; validates parentUUID
func truncate(blockCount: Int) throws // stacked: affects top layer; does NOT resize guest fs
var blockCount, blockSize, format, layerType, layerUUID, parentUUID, openMode, size, url
}
protocol StackedImage: DiskImage { var layers: [DiskImage] }
struct OpenConfiguration { init(url:mode:); Mode = automatic | readOnly | readWrite }
// CreationConfiguration statics: .asif(url:blockCount:blockSize:), .asifLayer(url:type:), .raw(url:blockCount:)
// DiskImage.LayerType: .cache | .overlay | .overlay(blockCount:)
// DiskImage.BlockSize: .bytes512 | .bytes4096
// Errors: CorruptedImageError, IncompatibleStackingError(reason), InvalidBlockCountError, UnsupportedFormatError
```
Bridge into Virtualization is a new beta convenience init on the existing attachment class. Note there is no `readOnly:` parameter; read-only-ness comes from each layer's own `openMode`:
```swift
VZDiskImageStorageDeviceAttachment(diskImage: stack, cachingMode: .automatic, synchronizationMode: .full)
```
### Stacking rules (Apple docs, verbatim where quoted)
- ASIF works standalone or stacked. "You can only use RAW images as standalone images or as **base** images in stacked configurations." Upper layers are always ASIF.
- **One cache layer per stack**, any number of overlays conceptually, "shallow stacks perform better" (WWDC 224). No published max-depth guidance.
- "Layers are processed from bottom (base) to top. The **topmost layer determines the stack's size and receives all writes**." `.overlay(blockCount:)` therefore also grows the virtual disk.
- UUID chaining: appending sets the child's `parentUUID` to the parent's `layerUUID`. Raw bases have no UUID. "The layer UUID **changes if the layer is written to**", and reattaching a mismatched layer throws `IncompatibleStackingError`. This is the mechanism that makes a shared read-only base safe.
- Base sharing across multiple VMs is the stated design intent ("can be shared across multiple VMs"), with the WWDC caveat that per-VM auxiliary files (EFI variable store, macOS auxiliary storage) must be duplicated per VM, never shared.
- **There is no flatten/merge.** An overlay cannot be merged back into its base (confirmed by Howard Oakley's coverage plus an independent hands-on report). Export/move flows must ship the layer chain, or flatten inside a guest (dd to a fresh attached image).
### Known issues and adoption
- **ASIF space reclamation is broken for macOS guests on the beta** (deleted files never return space, survives reboots). Linux guests reclaim correctly on both raw and ASIF via `fstrim -av`. Single detailed field report, unrefuted. Since the VM subsystem targets macOS guests, the practical rule until this is fixed is: back macOS guest disks with RAW, and revisit ASIF stacking for macOS guests each beta (stacking still works, the disks just never shrink).
- **Zero shipping adopters anywhere.** tart has a design issue with no activity; nobody has published working DiskImageKit code. Everything must be treated as field-untested (and our own testing bears that out, Section 8).
- Framework binary grew every beta (588 → 598 across betas 1-4); expect churn until GA.
- Release notes list no DiskImageKit known issues in any beta, which given the above says more about the notes than the framework.
## 3. Guest provisioning (macOS guests only)
```swift
class VZGuestProvisioningOptions: NSObject { func validate() throws } // "use one of its subclasses"
class VZMacGuestProvisioningOptions: VZGuestProvisioningOptions {
var fullName, username, password: String
var logsInAutomatically: Bool
var enablesRemoteLogin: Bool // SSH
}
// Wiring: VZMacOSVirtualMachineStartOptions.guestProvisioningOptions (Mac-typed)
// .setGuestProvisioning(_:) throws (validating setter)
```
- **Requires macOS 27 on host AND guest.** Older guests **silently ignore** the options (no error).
- **First boot after restore only.** Cannot reconfigure an already-provisioned VM; property changes after start are no-ops.
- The base class is forward-looking scaffolding; its only subclass is Mac. A Linux/cloud-init analogue may come later; do not assume it lands in 27.0. For Linux guests, cloud-init NoCloud seed ISOs remain the provisioning path (proven working, Section 8).
- Field-verified behavior (third-party hands-on, beta 3): provisioned account gets full admin + sudo; Setup Assistant fully skipped; SSH reachable ~48 s after first boot. **Race**: the account is created late in first boot (~T+54 s), after LaunchDaemons start (~T+33 s), so anything at daemon-level must wait for the account to exist.
- Open Apple-acknowledged bug: provisioned users are invisible to `CSIdentityQueryExecute()` (FB23716201).
- IPSW acquisition gotcha for automation: `VZMacOSRestoreImage.latestSupported` tracks the latest *release* (returned 26.5.2), not the installed beta; beta IPSWs must be fetched from the seed CDN explicitly.
## 4. vmnet: a macOS 26 feature set, one macOS 27 fix
Everything interesting shipped in macOS 26: `vmnet_network_create`, `vmnet_network_configuration_create`, `..._add_port_forwarding_rule`, `..._add_dhcp_reservation`, subnet/prefix/MTU/external-interface setters, NAT44/NAT66/DHCP/DNS-proxy/RA disables, plus serialization (`vmnet_network_copy_serialization` / `_create_with_serialization`) for handing networks across processes. `VZVmnetNetworkDeviceAttachment` is macOS 26.
macOS 27's only change (beta 4 release notes, verbatim): "The vmnet port forwarding APIs now support port forwarding when communicating over loopback." That closes the old gap where the host could not reach its own forwarded ports via 127.0.0.1 (confirmed working by the original bug reporter). Directly relevant to Codeman's loopback-bound production server talking to per-case guests.
Gotchas:
- vmnet networks are **not persisted**; they die with the owning process. Persist settings yourself and recreate (or serialize across processes).
- The `com.apple.vm.networking` entitlement is still restricted ("contact your Apple representative", though DTS says most requests are approved). The plain `VZNATNetworkDeviceAttachment` needs no special entitlement and is what our prototype uses.
- Ecosystem signal: tart's maintainer is not adopting in-process vmnet (prefers their separate-process softnet), so field testing of these APIs is thin.
## 5. VZCustomVirtioDevice (macOS 27, Linux guests only)
14 new types (`VZCustomVirtioDevice(+Configuration/Delegate/Provider)`, `VZVirtioQueue(+Element)`, `VZVirtioFeatureSet`, shared-memory-region types, `VZGuestMemoryMapping`), wired via `VZVirtualMachineConfiguration.customVirtioDevices`. Mandatory for guest discovery: `deviceID`, `pciClassID`, `pciSubclassID`, `virtioQueueCount`. You must write the Linux guest driver (Virtio spec 1.3/1.4). Threading contract: the framework calls the device/delegate on a serial queue (`deviceQueue`, defaulting to the VM's queue). Zero public adopters. For the VM subsystem this is a Phase 3+ option for a low-latency host-guest channel; SSH over NAT is proven and sufficient for now.
## 6. Signing and entitlements
- **Core loop (VZ + DiskImageKit + provisioning): ad-hoc signing with only `com.apple.security.virtualization` suffices.** Verified by us on beta 4 (plain `codesign --entitlements ... -s -` on a `swiftc` binary) and independently by third parties on beta 3. DiskImageKit documents no entitlement at all.
- **Over-entitling is the actual trap.** Adding `com.apple.application-identifier`/team-identifier keys without an embedded provisioning profile hangs the process before `main` (watchdog kill); shipping `com.apple.vm.networking` unauthorized gets AMFI SIGKILL at exec (exit 137, no crash report, even for `--version`). Keep the entitlements plist to exactly the one key.
- **USB passthrough breaks the ad-hoc story**: `com.apple.developer.accessory-access.usb` is profile-restricted (any paid team, no ad-hoc), additionally requires `com.apple.security.device.usb`, and `AAUSBAccessoryManager` presents UI, so it wants a Dock app, not a headless CLI. Out of scope for Codeman.
- No Xcode required for any of the above: the CLT beta (~500 MB via `softwareupdate`) carries the full macOS 27 SDK including DiskImageKit and compiles/signs everything.
## 7. Ecosystem state (July 2026)
- **tart is now `openai/tart`** (moved from cirruslabs, mid-2026) and **relicensed to FSL-1.1-ALv2** (no longer permissive). Provisioning support shipped in 2.33.0. Old cirruslabs URLs and license assumptions are stale.
- VirtualBuddy shipped provisioning ("Skip Setup Assistant") in 2.2 betas; had to add account-detail validation and a workaround installer for the cross-version bug below.
- lima is deliberately waiting for GA before touching macOS 27 APIs.
- **Code-Hex/vz (Go bindings) is dormant** (no commits since Feb 2026, no macOS 27 APIs), so the entire Go ecosystem (podman-machine, colima) currently has no path to these APIs. Swift is the only realistic binding today, which validates the VM subsystem's Swift-helper design.
- Useful pattern if ever supporting older SDKs: resolve new classes via `NSClassFromString` at runtime (no link-time dependency), fail gracefully when absent.
- **Cross-version restore bug**: installing a macOS 27 guest from IPSW on a macOS 26 host fails at 77-78% (`VZErrorDomain 10007`); fixed in 26.6b3 + Xcode 27b4 era, with a nasty MobileDevice.pkg trap (installing it from Xcode 27 beta on a 26 host requires a full macOS reinstall to undo). Not relevant to our 27-host testbed, very relevant to anyone on a 26 host.
## 8. Our empirical results (beta 4, 26A5388g, MacBook Air M3, 2026-07-29)
Prototype tooling, all in `~/vm-lab/` on the testbed, compiled with CLT-only `swiftc` and ad-hoc signed with the single virtualization entitlement:
| Tool | Purpose |
| --- | --- |
| `vzboot.swift` | Linux guest: EFI boot + virtio disk/net/entropy + NAT + optional cloud-init seed ISO + serial on stdio |
| `vzstack.swift` | Same, but boots a DiskImageKit stack (read-only raw base + ASIF overlay) |
| `vzmac.swift` | macOS guest: `install` (IPSW restore into a bundle) and `run` (boot, `--provision` for first-boot account/SSH) |
| `vzmacgui.swift` | macOS guest in a real window via `VZVirtualMachineView` (required for the guest to render at all) |
| `setup-seed.sh` | Builds a cloud-init NoCloud seed ISO with `hdiutil makehybrid` (volume label `cidata`) |
| `vncproxy.py` | RFB proxy that advertises only security type 2, so version-skewed/browser clients can authenticate |
| noVNC + `websockify` | Browser access; `websockify --web noVNC-<ver> 0.0.0.0:<port> 127.0.0.1:<proxy>` |
| `vmwatchdog.sh` + `vmaccess.sh` | Supervision: root LaunchDaemon that restarts a blind/dead runner, re-points the forward, re-applies `pmset`, re-arms keep-awake; plus a keeper for the proxy/web endpoints |
Host-side diagnostics written during this work (in the session scratchpad, not on the testbed): `vnclogin.py` (Apple DH auth + session open, distinguishes "credentials rejected" from "authorized but session refused"), `vncshot.py` (decodes the raw framebuffer to PNG and reports non-black pixel counts, plus optional synthetic wake input), `relay.py` (plain TCP relay used to bridge a tailnet peer to a LAN-only host), `sshpw.py` (pty-driven password SSH for the one-time key bootstrap into a freshly provisioned guest).
### Proven working
1. **Boot**: Debian 12 arm64 cloud images (nocloud and genericcloud variants) boot under `VZEFIBootLoader` + `VZGenericPlatformConfiguration`.
2. **Networking**: `VZNATNetworkDeviceAttachment` gives the guest a `192.168.64.x` DHCP lease from the host's bootpd (leases visible in `/var/db/dhcpd_leases`, bridge is `bridge100`).
3. **cloud-init provisioning**: NoCloud seed ISO (built with `hdiutil makehybrid -iso -joliet -default-volume-name cidata`) created a `codeman` user with SSH key + passwordless sudo on first boot; `ssh codeman@<lease-ip>` from the host works with key auth.
4. **DiskImageKit stack mechanics**: opening a raw base `.readOnly`, appending an ASIF overlay (`ASIFCreationConfiguration.layer(url:type:.overlay)`), attaching via `init(diskImage:)`, and booting it. The overlay received ~44 MB of boot-time writes while the **base file's SHA-256 stayed bit-identical**, which is the write-isolation property the whole per-case design rests on.
5. **Reattach**: reopening an existing overlay and `appending(consuming:)` onto the same base passes UUID validation.
6. **macOS guest install (added later the same day)**: `VZMacOSInstaller` restore of the 27.0 IPSW (26A5388g, fetched from the seed CDN via appledb; same build as host) into a sparse 64 GiB raw disk + auxiliary storage: INSTALL-OK on the first attempt, ~25 minutes.
7. **Headless guest provisioning WORKS**: `VZMacGuestProvisioningOptions` via `setGuestProvisioning` (username, password, `enablesRemoteLogin`, `logsInAutomatically=false`) produced, with zero GUI interaction: an account with full admin (groups include `80(admin)`, `com.apple.access_ssh`), Remote Login on from first boot, port 22 reachable ~140 s after first-boot start, hostname auto-derived from the account ("Codemans-Virtual-Machine"). SSH password auth is on by default, so the bootstrap path is: pty-driven password login once to install `authorized_keys`, key auth thereafter. Note the provisioned account's sudo is NOT passwordless (`echo <pass> | sudo -S ...`), and provisioning is first-boot-only (later boots take no options and just boot).
8. **Slot-leak bug NOT reproduced on 26A5388g**: a guest-initiated `shutdown -h now` fired `guestDidStop` cleanly and an immediate relaunch started fine (SSH-ready again in ~75 s), so FB22967193 (VM slot leaked on guest-initiated shutdown, host reboot to recover) did not manifest after one cycle. Either fixed in beta 4 or needs more cycles to trigger.
### Unstable / under investigation (beta-quality territory)
Boot reliability degraded over a ~15-VM session on one host boot, ending with reproducible silent hangs (VM process alive, 0% CPU, no DHCP, no ARP, nothing on serial):
- A genericcloud base that had been booted read-write once (cloud-init first boot) subsequently hung on every boot **with the seed ISO still attached**, while booting **without** the seed succeeded, then later runs failed in both configurations. The seed correlation is strong but was observed while host state was already suspect, so it needs a retest from a clean baseline.
- The first stack-boot "success" that later wedged turned out (via DHCP lease timestamp arithmetic) never to have reached the network at all; its overlay growth was pre-network boot writes.
- Working hypothesis, matching a class of acknowledged beta bugs (e.g. the VM-slot counter that leaks on guest-initiated shutdown, FB22967193, where only a host reboot recovers): accumulated hypervisor/vmnet state on the host degrades boots. Requires a host reboot + a disciplined retest matrix to confirm.
### Display rendering: the single most important operational finding
**A VZ macOS guest renders nothing unless a `VZVirtualMachineView` is attached AND the host session is actually drawing.** Verified byte-for-byte: the guest's own screen sharing serves an all-zero framebuffer (0 non-black bytes across 400 KB samples, with a sane pixel format: `rmax/gmax/bmax = 255`, shifts 16/8/0), in-guest `screencapture` fails with "could not create image from display", and no `IODisplayWrangler` shows up in the guest's `ioreg`. Three distinct states all produce black:
1. **Headless** (VM run with no view attached).
2. **View attached, host session locked.** The lock screen suspends drawing and the guest's virtual GPU produces no frames.
3. **View attached, but the app lost its WindowServer connection** (see the incident below): black permanently until the app is restarted.
**Consequence for the VM subsystem: rendering is a first-class requirement, not an optional extra (owner decision 2026-07-29).** The product serves GUI desktops: mandatory for macOS guests, optional-but-supported for Linux guests (which can also run headless over SSH). Any VM in GUI mode must be launched by an app that attaches a `VZVirtualMachineView`, from inside a host GUI session that is logged in and unlocked. That makes the following non-negotiable parts of the design, not workarounds:
- VMs run as **GUI apps in the console user's session** (launched via a LaunchAgent or `launchctl asuser`), never as daemons.
- The **host must auto-login and never lock or sleep**; a locked host is equivalent to a powered-off display for every VM on it.
- The **guest must auto-login, never lock, and have its first-login assistant pre-suppressed**, or the "desktop" a user connects to is a password prompt or a setup wizard.
- A VM app that loses its WindowServer connection is **permanently blind** and must be restarted; supervision has to detect that, not just check that the process is alive.
- The **2-concurrent-macOS-VM cap** becomes a real capacity limit for the product, so it must be surfaced in the UI and tested (still untested worldwide as of this writing).
### Incident 2026-07-29: `killall -HUP loginwindow` (never do this on a remote Mac)
Applying a wallpaper change on the testbed with `killall -HUP loginwindow` restarted the host's login session. Three consequences:
1. **The Mac dropped off the tailnet entirely.** Tailscale's App Store build is a GUI app living in the user session, so killing the session killed the VPN; remote access was gone until someone logged in. Recovery came from a second machine on the same LAN: it could still SSH in, and then relay ports back over the tailnet (a plain TCP relay on a tailnet-connected LAN peer is a good out-of-band path worth keeping ready).
2. **The VM app lost its WindowServer connection** (`HIToolbox: received notification of WindowServer event port death`) while surviving as a process. Every later black screen traced to this, and nothing guest-side could fix it; only restarting the app restored rendering.
3. The session's `caffeinate` died, so the host resumed auto-locking.
Rule: on a remote Mac, never run session-level commands (`killall -HUP loginwindow`, `pkill -u <user>`, logout, fast user switching). `killall WallpaperAgent` alone is session-safe. Before any such command, enumerate what depends on that session: VPN, VM processes, port forwards, keep-awake helpers.
### Keeping host and guest usable unattended
- **Host**: `caffeinate -d -i -m -u` prevents display sleep but does NOT override the lock policy. "Require password after screen saver begins or display is turned off → Never" must be set in System Settings; it needs the account password, so a passwordless-sudo shell cannot script it, and turning it off does NOT dismiss a lock that is already engaged (one more unlock is always needed). `pmset -a disablesleep 1` keeps a lid-closed laptop awake but **does not survive a reboot**, and OS updates reset it too, so a supervisor should re-apply it rather than assume it sticks.
- **Rebooting an encrypted host**: use `sudo fdesetup authrestart`. FileVault's pre-boot unlock doubles as the login, so the machine returns with a **live logged-in console session** and encryption intact, no password prompt, and supervision can then bring the VMs back by itself. Verified 2026-07-30. A plain `reboot` parks at the lock screen and blacks out every VM until a human logs in.
- **Guest**: set `autoLoginUser` plus a valid `/etc/kcpassword` (XOR-obfuscated password file, key `7D 89 52 23 D2 BC DE A3`, payload zero-padded to a multiple of 12). `sysadminctl -autologin` fails with `SACSetAutoLoginPassword error:22` on provisioned accounts, and a fresh guest has no Python, so generate the bytes on the controlling host and copy them in. Then `pmset -a displaysleep 0 sleep 0 disablesleep 1`, `defaults -currentHost write com.apple.screensaver idleTime 0`, `defaults write com.apple.screensaver askForPassword 0`, and `caffeinate` inside the guest. ⚠ `autoLoginUser` was observed being wiped by failed `sysadminctl -autologin` attempts; verify it after each boot until stable.
- **Wallpaper**: animated "aerials" wallpaper is brutal over VNC. The provider lives in `~/Library/Application Support/com.apple.wallpaper/Store/Index.plist` under several keys (`AllSpacesAndDisplays:Desktop`, `:Idle`, and `SystemDefault:*` which is what the login/lock screen uses). Switch each `Provider` to `com.apple.wallpaper.choice.solid-color` with PlistBuddy and restart `WallpaperAgent`. The login-window copy is cached and only refreshes on a later login cycle.
### Remote GUI/SSH access to a guest (recipe, verified 2026-07-29)
The guest lives on the host-private NAT bridge, so remote access is guest-service + host-forward:
1. **In the macOS guest** (over ssh), use ONE mechanism, fully activated. The reliable form is Remote Management in a single kickstart call:
```
sudo .../RemoteManagement/ARDAgent.app/Contents/Resources/kickstart \
-activate -configure -access -on \
-clientopts -setvnclegacy -vnclegacy yes -setvncpw -vncpw <8-char-pw> \
-allowAccessFor -allUsers -privs -all -restart -agent -menu
```
⚠ **Half-configured states authenticate but refuse the session.** Loading `com.apple.screensharing` while Remote Management is deactivated (or vice versa) produces an Apple-client error that names the wrong culprit: *"Screen Sharing is not permitted on <host>. Disable and re-enable Screen Sharing or Remote Management in System Settings"*. A raw-protocol client can still authenticate AND open a framebuffer in that state, so protocol-level tests pass while every Apple client fails. The remedy is exactly what the dialog says, done over ssh: `launchctl unload -w …screensharing.plist`, `kickstart -deactivate -configure -access -off`, `pkill screensharingd`, then the single activate call above.
Notes: `launchctl enable system/com.apple.screensharing` fails with "Could not find service" on this build; `load -w` is the plain-Screen-Sharing path if you deliberately want it instead of Remote Management. Apple clients negotiate `RSA-SRP` (auth type 33) and the guest logs `Authentication: SUCCEEDED :: User Name: … :: Type: RSA-SRP` on success, which is the definitive server-side confirmation.
2. **On the host**: a gateway port-forward makes the guest's 5900 reachable from the whole tailnet without per-client tunnels: self-authorize the host's own key, then `ssh -N -g -L 0.0.0.0:5901:<guest-ip>:5900 <user>@localhost` (nohup'd).
⚠⚠ **NEVER forward on host port 5900.** If the host has Screen Sharing enabled (our testbed does, from the pre-upgrade checklist), launchd already owns 5900 socket-activated. The `ssh -L` bind then fails with "Address already in use" **while the tunnel process keeps running**, so every symptom of success is present (process alive, port answers, real RFB banner) yet **every connection reaches the HOST's login window, not the guest**. This cost us an hour: guest credentials failed against the host's screensharingd, which reads exactly like broken guest auth, and we chased the (real, but irrelevant) provisioned-account identity bug. Diagnostics that would have caught it instantly: `sudo lsof -nP -iTCP:5900 -sTCP:LISTEN` showing `launchd` rather than `ssh`, or the guest's own logs showing NO auth attempts during a failed login. Always use a distinct host port and verify with `lsof` that the forward owns it.
⚠ `-g` binds all interfaces, so the forward is also visible on the host's LAN; the VNC layer still requires the account or VNC password. ⚠ The forward pins the guest IP, which changes per boot under plain NAT; re-point it after a guest reboot (the proper fix is a vmnet DHCP reservation, macOS 26 API, once we move off plain `VZNATNetworkDeviceAttachment`).
Verified working: with the forward on 5901, both a provisioned account and a `sysadminctl`-created one authenticate successfully (RFB `SecurityResult` = 0) against the guest. The guest offers security types `[30, 33, 36, 2, 35]`, i.e. Apple DH/SRP **plus classic type 2**, so non-Apple VNC clients work with the legacy password once ARD's `-setvnclegacy` is set. (The host's screensharingd, by contrast, offered no type 2, which is itself a tell that you are talking to the wrong machine.)
3. **SSH from any tailnet device**: `ssh -J <host-user>@<host> codeman@<guest-ip>` (jump through the host), after adding the connecting machine's key to the guest's `authorized_keys`.
**Client-version incompatibility (macOS 27 servers vs older Screen Sharing clients)**: an older Mac's Screen Sharing client fails Apple's `RSA-SRP` handshake against macOS 27 servers, logging `Authentication: FAILED :: User Name: <user> :: Type: RSA-SRP` server-side, while a macOS 27 client authenticates against the same servers without issue. This was verified against BOTH a macOS 27 guest and a macOS 27 host with the operator's own account, so it is a client-side version skew, not configuration, and no server-side change fixes it. Same family as the documented "macOS 26 host cannot install a 27 guest" bug. Practical workaround: bypass Apple auth entirely with classic VNC auth (security type 2), which macOS offers only when Remote Management legacy VNC is enabled. Two ways to consume it: any third-party VNC client, or a browser via noVNC.
**Browser-based access chain (zero client install, version-proof)**, all hosted on the Mac:
```
browser --HTTP/WS--> websockify (+ noVNC static files)
--> type-2-only proxy # rewrites the server's security-type list to [2]
--> ssh -L forward # loopback hop; see the Local Network note below
--> guest:5900
```
Notes learned the hard way: (a) **never bind the forward on host port 5900** (see the launchd warning above); (b) a Python proxy cannot reach the guest subnet directly because macOS **Local Network privacy** denies headless CLI binaries, surfacing as `No route to host`, so point the proxy at a loopback `ssh -L` forward instead (Apple-signed `ssh` is unaffected); (c) noVNC needs `?resize=scale` or Scaling Mode → Local Scaling, otherwise a Retina host screen (2940x1912) is unusable in a browser window; (d) noVNC speaks security type 2 only, which is exactly why the proxy rewrite is needed.
**Debugging technique that settled all of this**: a ~80-line Python RFB client (scratchpad `vnclogin.py`) that implements Apple DH auth (security type 30) and continues through `ClientInit`/`ServerInit`. It reports the server's `SecurityResult` plus the framebuffer size and desktop name, which separates "credentials rejected" from "authorized but session refused" without any GUI client. Pair it with `log stream --predicate 'process == "screensharingd"'` inside the guest, and drive a REAL Apple client headlessly from the host with `sudo launchctl asuser <uid> sudo -u <user> osascript -e 'tell application "Screen Sharing" to open location "vnc://user:pass@host:port"'`, verifying the result via `lsof -nP -iTCP -a -p <pid>` (an ESTABLISHED socket to the target) since `screencapture` fails on a lid-closed laptop ("could not create image from display"). Tailscale was never implicated: both the raw client and Apple's client work over the tailnet address once the guest service is fully activated.
### Hard-won operational lessons (write these into any tooling)
- **Silent serial is normal, not failure.** Debian's GRUB/kernel log to the graphics console; nothing attaches a getty to hvc0 by default. The reliable boot signal is the DHCP lease (or passive `tcpdump -i bridge100`), never the serial port and never a quick ping (BSD ping's first packet often dies to ARP latency; passive capture showed "dead" guests alive).
- **DHCP lease entries carry truth**: `name=` shows the guest hostname, and the lease timestamps order events; stale entries linger, so compare timestamps before attributing a lease to a boot.
- **Never boot a base image read-write.** Every RW boot mutates it (dhclient lease cache, journal, cloud-init state) and destroys experiment reproducibility, exactly why the production design only ever boots bases under overlays. Provision INTO the base once at base-build time, or provision per-case overlays with the seed, then detach the seed.
- **A killed SSH client does not kill a remote `nohup`'d VM**, and the survivor holds the EFI variable store lock: "The EFI variable store is already in use" (`VZErrorDomain 50002`) means a zombie VM process, `pkill` it.
- **EFI variable stores are per-VM state.** Fresh stores boot reliably; reuse across different VM instances is at minimum suspect on this beta (Apple's own guidance for cloned VMs is one store per VM). Cheap policy: one store per case, created with the overlay, deleted with it.
- **Downloads from cloud.debian.org mirrors truncate silently**; always verify byte count against origin `Content-Length` and resume with `curl -C -`.
- The remote host's default shell is zsh: `=` -prefixed words (`echo ===`) explode via zsh's `=cmd` expansion; keep separators zsh-safe in automation.
### The 2-concurrent-macOS-VM cap: TESTED AND CONFIRMED on macOS 27 beta 4 (2026-07-29)
We measured it, which as far as we can tell nobody had published for macOS 27. Method: `cp -c -R` the guest bundle (APFS clonefile, instant and **zero additional disk**), regenerate the machine identifier per clone (`VZMacMachineIdentifier()` written to `machine.id`; the hardware model is reused), then launch VMs until one is refused.
Result: VM #1 (8 GB, GUI) and VM #2 (4 GB, headless) ran concurrently without complaint. VM #3 was refused **instantly** at `vm.start`:
```
VZErrorDomain Code=6 "The maximum supported number of active virtual machines has been reached."
NSLocalizedFailure = "The number of virtual machines exceeds the limit."
```
**This is a licensing/kernel quota, not a resource limit**: the refusal came with **39% of system memory free** on a 16 GB host, and adding RAM or CPU cannot raise it. It matches the pre-27 behavior (`hv_apple_isa_vm_quota`), so nothing changed in 27 despite the framework's other additions. Linux guests are unaffected and are bounded only by host resources.
Design consequences: macOS-guest capacity per host is **hard-capped at 2**, so a GUI-macOS-per-case product must schedule around it (queue, evict idle VMs, or scale across hosts) and surface it in the UI. Also relevant: the acknowledged slot-leak bug (a guest-initiated shutdown failing to release a slot, recoverable only by host reboot) is far more damaging under a cap of 2 than it sounds; we did not reproduce it on beta 4, but any scheduler should treat "slot appears used but nothing is running" as a real state.
### Not yet tested
- Cache layers (`LayerType.cache`), `.overlay(blockCount:)` disk growth, stack depth performance, VirtioFS + stack combination, `truncate`, ASIF disks for macOS guests (raw used so far; ASIF has the reclamation bug).
- One more scripting lesson from this session: inner `ssh` calls inside a piped `sh -s` script MUST use `-n`, or they consume the remainder of the script from stdin and it silently never runs.
### Session timeline (what was actually established, 2026-07-29)
Linux path: base image download (with resume, mirrors truncate) → `vzboot` compiles against the beta SDK first try → EFI boot → NAT DHCP lease → cloud-init seed provisions a user with the host's SSH key → `ssh` into the guest works → DiskImageKit stack boots with an ASIF overlay taking all writes while the base stays SHA-identical. Later Linux boots became unreliable on an un-rebooted host (silent hangs, 0% CPU, no DHCP); a clean-baseline retest is still pending.
macOS path: seed-CDN IPSW (matched to the host build) → `VZMacOSInstaller` restore, ~25 min, first try → first boot with `VZMacGuestProvisioningOptions` creates an admin account with Remote Login on, no interaction needed, SSH reachable ~140 s later → key bootstrap over a one-time password login → guest shutdown/relaunch clean (the slot-leak bug did not reproduce) → GUI access fought through a port collision, a client-version incompatibility, the rendering dependency, and a self-inflicted session kill, ending with a browser-based path plus a guest hardened to auto-login and never lock.
**Lifecycle verified (stop → start), 2026-07-30**: an in-guest `shutdown -h now` fires `guestDidStop` and the runner app exits on its own; relaunching from the same bundle boots the guest in ~2 minutes straight into an auto-logged-in desktop, and the VM slot is released cleanly (an immediate restart works, so the slot-leak bug did not bite). Two operational notes: the guest takes a **new NAT lease on every boot**, so any port-forward must be re-pointed (or use a vmnet DHCP reservation), and a host reboot resets `pmset -a disablesleep`.
⚠ **Provisioning does NOT skip the per-user first-login assistant.** `VZMacGuestProvisioningOptions` skips the initial Setup Assistant (account creation, region, Apple Account) so the machine is immediately reachable, but the first time anyone actually logs into a desktop, macOS still presents its per-user wizard (Apple Intelligence, Siri, privacy, appearance, Touch ID). The operator hit exactly this. For a GUI-first product this MUST be pre-suppressed during base-image creation by writing `com.apple.SetupAssistant` keys for every account that will log in, and into `/System/Library/User Template/English.lproj/Library/Preferences/` so accounts created later inherit it.
⚠ **A partial key list is worse than none**, because the wizard simply shows the panes you missed and the operator has to click through them again after every fresh login (we hit this twice). The set that finally silenced macOS 27 beta 4: `DidSeeCloudSetup`, `DidSeeSiriSetup`, `DidSeePrivacy`, `DidSeeAppearanceSetup`, `DidSeeTouchIDSetup`, `DidSeeAvatarSetup`, `DidSeeScreenTime`, `DidSeeApplePaySetup`, `DidSeeSafariImport`, `DidSeeAccessibility`, **`DidSeeActivationLock`, `DidSeeAppStore`, `DidSeeLockdownMode`** (the three easy to miss), plus the Express-Settings flags **`SkipExpressSettingsUpdating`** and **`SkipFirstLoginOptimization`**, and the version markers `LastSeenCloudProductVersion` / `LastSeenBuddyBuildVersion` / `PreviousSystemVersion` / `PreviousBuildVersion` matching the guest build. Verify afterwards by reading the domain back and checking that no `DidSee*` key is still `0`. Note these keys change between macOS releases, so base-image creation should re-verify per OS version rather than trust a hardcoded list.
## 9. Design implications for Codeman's VM subsystem
0. **GUI is a first-class mode, and for macOS guests it is the whole point (owner decision, 2026-07-29).** The subsystem serves real desktops, not only headless SSH boxes. macOS guests are GUI-only in practice (nothing renders without an attached view). Linux guests are supported in BOTH modes: GUI when the case wants a desktop, headless-over-SSH when it wants a cheap agent sandbox. The costs of the GUI path are in §8 "Display rendering": VMs as GUI apps in a live session, a host that never locks, guests that auto-login with their first-login wizard pre-suppressed, and the macOS concurrency cap as a real capacity limit.
1. **The macOS-specific liabilities are accepted costs, not reasons to avoid macOS guests**: provisioning is macOS-only and first-boot-only, ASIF space reclamation is broken for macOS guests on the beta (use RAW disks for macOS guests until fixed), and the 2-VM cap applies. Plan around each: RAW-backed macOS disks, provisioning baked into base-image creation, and capacity limits surfaced in the UI.
2. **Base immutability is not just hygiene, it is load-bearing**: DiskImageKit's UUID invalidation plus our sha-stability proof make a read-only shared base per image-generation the core artifact. Bases are built once (seed attached), then only ever opened `.readOnly` under per-case overlays.
3. **Seed ISOs are a base-build-time tool only.** Never attach a seed to a routine case boot (correlated with boot hangs on the beta, and semantically wrong anyway since cloud-init already ran).
4. **Per-case files**: overlay ASIF + EFI variable store live and die together with the case.
5. **Export = ship the layer chain** (base ref + overlay + manifest), not flatten; there is no flatten API. In-guest `dd` to a fresh image is the fallback for a true single-file export.
6. **Health checking must be lease/API based**, not serial/ping based, and Codeman's `codeman-vm status` should read `/var/db/dhcpd_leases` (or use vmnet DHCP reservations for deterministic per-case IPs, a macOS 26 API).
7. **Run `fstrim` periodically in Linux guests** (or mount with discard) so overlays stay sparse.
8. **Entitlements plist stays minimal** (exactly `com.apple.security.virtualization`) to dodge the AMFI/watchdog traps.
9. **Expect beta churn**: pin findings to build numbers (this doc: 26A5388g) and retest each beta; the framework binaries changed every beta so far.
10. **A macOS guest is only "ready" when its desktop is ready**, which is a stricter bar than "the VM booted". Readiness means: VM app running with a live WindowServer connection, guest auto-logged-in (not at a login or lock screen), first-login assistant suppressed, and the guest's screen sharing serving a non-black framebuffer. Health checks should sample the framebuffer for non-black content, because every failure mode in this session (headless run, locked host, dead WindowServer, locked guest, setup wizard) presents as a perfectly healthy-looking process with a black or useless screen.
10b. **Supervision must run as a root LaunchDaemon.** A user LaunchAgent cannot launch a GUI app into the Aqua session; its restarts fail silently (child dies instantly, empty log, supervisor reports success). Root + `launchctl asuser <uid> sudo -u <user> …` works and the launched process persists. This bit us on the first supervisor implementation and is easy to repeat.
11. **Remote-access plumbing belongs in the helper CLI, not in ad-hoc shell**: a `codeman-vm` implementation should own port selection (never 5900), forward lifecycle across guest IP changes (or better, vmnet DHCP reservations for stable per-case IPs), and a documented browser path, because every failure in this session came from hand-rolled plumbing rather than from the Virtualization APIs themselves.
12. **Never let control-plane connectivity depend on a GUI session** on a remote Mac host: prefer a Tailscale system service over the App Store app, and keep a LAN-adjacent peer able to relay as an out-of-band recovery path.
## Sources
Apple DocC JSON backend (diskimagekit, virtualization, vmnet trees; macOS 27 release notes) | WWDC26 session 224 https://developer.apple.com/videos/play/wwdc2026/224/ | eclecticlight.co ASIF/virtualization coverage | developer.apple.com/forums threads 839343 (CSIdentity bug), 830118 (cross-version restore), 830119 (VM-slot leak), 830383 (VM cap), 834822 + 831902 (USB entitlements), 822658 (vmnet loopback) | openai/tart issues 1261/1263/1268/1269/1285 | Spooky-Labs provisioning design doc | VirtualBuddy 2.2 release notes | lima-vm discussions | our own test transcripts on the testbed (`~/vm-lab/*.log`, this repo's session)
+1 -1
View File
@@ -1,7 +1,7 @@
# Web Tabs (dashboards as Codeman tabs)
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port
4000, as a tab beside your Claude/Codex/Gemini sessions. Codeman becomes one mission
4000, as a tab beside your Claude/Codex/Antigravity sessions. Codeman becomes one mission
control instead of Codeman plus a pile of browser tabs.
## Using it
+52 -11
View File
@@ -116,6 +116,14 @@ GEMINI_SEARCH_PATHS=(
"$HOME/bin/gemini"
)
# Antigravity CLI search paths (from src/utils/antigravity-cli-resolver.ts)
ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/.local/bin/agy"
"$HOME/.antigravity/bin/agy"
"/usr/local/bin/agy"
"$HOME/bin/agy"
)
# ============================================================================
# Color Output
# ============================================================================
@@ -493,6 +501,34 @@ get_gemini_path() {
done
}
check_antigravity() {
if command -v agy &>/dev/null; then
return 0
fi
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
get_antigravity_path() {
if command -v agy &>/dev/null; then
command -v agy
return
fi
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
check_cloudflared() {
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
@@ -1993,11 +2029,12 @@ main() {
fi
fi
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini)
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini, Antigravity)
local has_claude=false
local has_opencode=false
local has_codex=false
local has_gemini=false
local has_antigravity=false
info "Checking AI CLI tools..."
if check_claude; then
@@ -2016,17 +2053,21 @@ main() {
has_gemini=true
success "Gemini CLI found at $(get_gemini_path)"
fi
if check_antigravity; then
has_antigravity=true
success "Antigravity CLI found at $(get_antigravity_path)"
fi
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" ]]; then
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" ]]; then
echo ""
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, or Gemini."
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, or Gemini."
headless_guard "install an AI CLI (curl | bash from its vendor)"
echo ""
echo -e " ${BOLD}Which AI CLI would you like to install?${NC}"
echo -e " ${CYAN}1)${NC} Claude Code (Anthropic)"
echo -e " ${CYAN}2)${NC} OpenCode (open-source)"
echo -e " ${CYAN}3)${NC} Both"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Gemini)"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Antigravity)"
echo ""
local cli_choice=""
@@ -2071,8 +2112,8 @@ main() {
if [[ "$cli_choice" == "4" ]]; then
warn "Skipping AI CLI install. Codeman will run, but sessions need a CLI to drive."
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: npm install -g @google/gemini-cli (Gemini)"
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity)"
elif [[ "$has_claude" == "false" ]] && [[ "$has_opencode" == "false" ]]; then
die "The selected AI CLI failed to install. Install one manually and re-run the installer."
fi
@@ -2372,12 +2413,12 @@ main() {
echo -e " https://github.com/Ark0N/Codeman"
echo ""
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini; then
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity; then
echo -e " ${YELLOW}${BOLD}Reminder:${NC} Install at least one AI CLI to start using Codeman:"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}npm install -g @google/gemini-cli${NC} # Gemini"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}curl -fsSL https://antigravity.google/cli/install.sh | bash${NC} # Antigravity"
echo ""
fi
+2 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.10.0",
"version": "1.12.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.10.0",
"version": "1.12.0",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
+2 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.10.0",
"version": "1.12.0",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -55,6 +55,7 @@
"anthropic",
"opencode",
"codex",
"antigravity",
"gemini-cli",
"ai-agents",
"agent",
+1 -1
View File
@@ -16,7 +16,7 @@
> ### Made for [**Codeman**](https://getcodeman.com)
>
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Gemini sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Antigravity sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
>
> That last part is why this library exists. The demo below is a real Codeman session on two phones.
+152
View File
@@ -0,0 +1,152 @@
/**
* @fileoverview File Viewer edit-mode policy (issue #212).
*
* Pure, IO-free policy for which workspace files the in-viewer editor may read
* for editing and write back. Consumed by the `edit=1` branch of
* `GET /api/sessions/:id/file-content` and by `PUT /api/sessions/:id/file-content`
* in `src/web/routes/file-routes.ts`.
*
* Design (docs/file-viewer-edit-plan.md):
* - ALLOWLIST of text extensions/basenames, not a blocklist — matching the
* attachment-guard precedent. `svg` and `env` are deliberately absent: svg is
* treated as untrusted on the read side, and `.env` is sensitive-path blocked
* anyway; excluding them here keeps a single obvious refusal.
* - The `.git/` subtree is denied outright: `.git/hooks/*` is code execution and
* a corrupted index looks unrecoverable to a user who wanted to fix a typo.
* - EOL helpers exist because a browser <textarea> normalizes to LF; the server
* re-applies the file's original ending so a two-line edit of a CRLF file does
* not become a whole-file diff. Mixed-EOL files normalize to the dominant
* style (documented lossy edge).
*/
/** Hard cap for edit-mode reads AND writes (bytes of file content). */
export const MAX_EDITABLE_BYTES = 512 * 1024;
/** Lowercase extensions (no dot) the editor will open and save. */
export const EDITABLE_EXTENSIONS: ReadonlySet<string> = new Set([
// JS/TS ecosystem
'ts',
'tsx',
'js',
'jsx',
'mjs',
'cjs',
'json',
'jsonc',
// Docs / plain text
'md',
'mdx',
'txt',
'rst',
'adoc',
// Web
'css',
'scss',
'less',
'html',
'htm',
'xml',
// Config
'yml',
'yaml',
'toml',
'ini',
'cfg',
'conf',
'properties',
// Shell
'sh',
'bash',
'zsh',
'fish',
// Languages
'py',
'rb',
'go',
'rs',
'java',
'kt',
'swift',
'c',
'h',
'cpp',
'hpp',
'cc',
'cs',
'php',
'sql',
'graphql',
'proto',
'lua',
'pl',
'r',
'jl',
'tf',
'gradle',
// Data / misc text
'csv',
'tsv',
'log',
'diff',
'patch',
]);
/** Extensionless (or dot-led) file names that are still editable text. */
export const EDITABLE_BASENAMES: ReadonlySet<string> = new Set([
'dockerfile',
'makefile',
'license',
'readme',
'changelog',
'authors',
'codeowners',
'procfile',
'.gitignore',
'.gitattributes',
'.dockerignore',
'.prettierignore',
'.prettierrc',
'.editorconfig',
'.nvmrc',
'.npmrc',
'.eslintignore',
]);
/** Whether a file name (basename only) is eligible for in-viewer editing. */
export function isEditableFileName(fileName: string): boolean {
const lower = fileName.toLowerCase();
if (EDITABLE_BASENAMES.has(lower)) return true;
const dot = lower.lastIndexOf('.');
// No extension (or a bare dotfile like `.bashrc`): only the basename list applies.
if (dot <= 0) return false;
return EDITABLE_EXTENSIONS.has(lower.slice(dot + 1));
}
/**
* Whether a workspace-relative path is denied for editing regardless of its
* extension. Currently: anything inside a `.git` directory at any depth.
*/
export function isDeniedEditRelativePath(relativePath: string): boolean {
return relativePath.split('/').some((segment) => segment === '.git');
}
export type FileEol = 'lf' | 'crlf';
/** Dominant line-ending style of a text buffer (LF when tied or single-line). */
export function detectEol(text: string): FileEol {
let crlf = 0;
let lf = 0;
for (let i = 0; i < text.length; i++) {
if (text.charCodeAt(i) === 10) {
if (i > 0 && text.charCodeAt(i - 1) === 13) crlf++;
else lf++;
}
}
return crlf > lf ? 'crlf' : 'lf';
}
/** Normalize every line ending in `text` to the requested style. */
export function applyEol(text: string, eol: FileEol): string {
const normalized = text.replace(/\r\n/g, '\n');
return eol === 'crlf' ? normalized.replace(/\n/g, '\r\n') : normalized;
}
+3
View File
@@ -596,6 +596,9 @@ interface CredStorePolicy {
const CRED_STORES: CredStorePolicy[] = [
{ rel: '.codex', shareDirs: ['sessions'], shareFiles: ['history.jsonl'], seedFiles: ['auth.json', 'config.toml'] },
// Also covers Antigravity: `agy` nests its whole state (auth `jetski_state.pbtxt`,
// `conversations/`, `knowledge/`) under `~/.gemini/antigravity-cli/`, so it needs no
// entry of its own. There is no `~/.antigravity` credential dir to add.
{ rel: '.gemini', seedWhole: true },
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
+63
View File
@@ -256,6 +256,69 @@ export async function checkRemoteTmuxAvailable(
}
}
/**
* The CLI binary each session mode runs on the remote host. Antigravity's
* binary is `agy` (the mode name is not the command); shell has no CLI to
* probe, so it is absent.
*/
const REMOTE_CLI_BIN: Partial<Record<SessionMode, string>> = {
claude: 'claude',
opencode: 'opencode',
codex: 'codex',
gemini: 'gemini',
antigravity: 'agy',
};
/**
* Build the SSH command that reads the remote CLI's version (`claude --version`
* on the remote host). The version query is routed through
* `remoteLoginShellCommand` (the SAME `$SHELL -i -l -c` wrapper the real
* launch uses), because agent CLIs live on PATH only after the remote user's
* interactive-login startup files run (see defaultRemoteCommandForMode); a bare
* `claude --version` over ssh exits 127. Connection options come from the
* shared `buildSshConnectionArgs`, so the probe reaches exactly the hosts the
* launch can reach. Returns null for modes with no CLI (shell).
*/
export function buildRemoteCliVersionProbeCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): string | null {
const bin = REMOTE_CLI_BIN[mode];
if (!bin) return null;
return [
...buildSshConnectionArgs(host),
remoteSshTarget(host),
shellescape(remoteLoginShellCommand(`${bin} --version`)),
].join(' ');
}
/**
* Read the CLI version installed ON THE REMOTE HOST. Feeds Session.cliVersion
* for remote sessions: the deterministic local probe deliberately skips them
* (it would report the LOCAL host's claude), and the startup-banner scrape is
* unreliable (newer Claude Code builds print no banner; resumed sessions never
* do), which left cliVersion undefined and silently disabled wheel-forwarding
* to the CLI transcript (residual #154, noted in the #205 analysis). The
* version is parsed as the first semver in stdout, never raw output: an
* interactive-login shell may echo rc-file noise around it. Returns undefined
* on any failure. No-op under VITEST (mirrors checkRemoteTmuxAvailable).
*/
export async function probeRemoteCliVersion(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): Promise<string | undefined> {
if (process.env.VITEST) return undefined;
const command = buildRemoteCliVersionProbeCommand(host, mode);
if (!command) return undefined;
try {
const { stdout } = await execAsync(command, { timeout: 15_000 });
const match = stdout.match(/\d+\.\d+\.\d+/);
return match ? match[0] : undefined;
} catch {
return undefined;
}
}
/**
* COD-105 — build the SSH command that lists `codeman-*` tmux sessions on a
* remote host's canonical `-L codeman` socket.
+13 -2
View File
@@ -226,6 +226,16 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
// different UUID. Backfill from the already-passed history: first try the claudeSessionId
// join, then the newest transcript in the same workingDir. Never overwrite a non-empty
// firstPrompt (so rows keyed to their own transcript are untouched).
//
// The workingDir guess is a last resort and MUST be skipped for any item that
// already has its own 'history' entry (step 1 above already gave it a real,
// direct scan of its own transcript). Without this guard, a history row whose
// OWN extraction genuinely failed (oversized first message, etc.) silently
// inherited the newest OTHER session's opening line from the same directory —
// not a blank, but actively wrong: old sessions displayed today's conversation
// as if it were their own. A row with no 'history' source at all (its
// transcript hasn't been linked/scanned under its own id yet) has no such
// direct attempt to prefer, so the guess remains a reasonable stand-in there.
const firstPromptByUuid = new Map<string, string>();
const firstPromptByWorkingDir = new Map<string, { prompt: string; ms: number }>();
// COD-145: lastPrompt rides the same backfill (build parallel indexes; never overwrite).
@@ -254,12 +264,13 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
}
}
for (const item of map.values()) {
const hasOwnHistoryEntry = item.sources.includes('history');
if (!item.firstPrompt) {
// never overwrite an existing non-empty prompt
const byUuid = item.claudeSessionId ? firstPromptByUuid.get(item.claudeSessionId) : undefined;
if (byUuid) {
item.firstPrompt = byUuid;
} else if (item.workingDir) {
} else if (item.workingDir && !hasOwnHistoryEntry) {
const byDir = firstPromptByWorkingDir.get(item.workingDir);
if (byDir) item.firstPrompt = byDir.prompt;
}
@@ -268,7 +279,7 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
const byUuid = item.claudeSessionId ? lastPromptByUuid.get(item.claudeSessionId) : undefined;
if (byUuid) {
item.lastPrompt = byUuid;
} else if (item.workingDir) {
} else if (item.workingDir && !hasOwnHistoryEntry) {
const byDir = lastPromptByWorkingDir.get(item.workingDir);
if (byDir) item.lastPrompt = byDir.prompt;
}
+114 -25
View File
@@ -54,6 +54,7 @@ import {
type SessionDocker,
} from './types.js';
import { probeDockerCliVersion } from './docker-hosts.js';
import { probeRemoteCliVersion } from './remote-hosts.js';
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
@@ -180,6 +181,37 @@ export function isAltScreenStripMode(mode: SessionMode): boolean {
return mode === 'codex' || mode === 'claude' || mode === 'gemini';
}
/**
* Modes that need the NARROW strip: alt-screen toggles only, leaving `\x1b[3J`
* and the mouse-tracking DECSETs alone. Applies to every mode `isAltScreenStripMode`
* excludes, but ONLY when the session is tmux-backed (`useMux`).
*
* The bug (issue #205): the tmux CLIENT emits `smcup` (`\x1b[?1049h`) as its first
* bytes on attach, before any program has run. Unstripped, xterm.js parks in the
* alternate buffer for the whole session, where `baseY` is pinned at 0 (no
* scrollback to reach, so touch scrolling is a no-op) and xterm's own wheel handler
* translates the wheel into `\x1bOA`/`\x1bOB` cursor keys — which readline receives
* as shell history navigation. Both reported symptoms, one sequence.
*
* Why this is safe under tmux, despite the old "shell must keep the alt screen for
* vim/less/htop" reasoning: tmux is a full terminal emulator and NEVER forwards a
* pane's alt-screen toggles to its client, it repaints instead. Captured from a real
* attach, `\x1b[?1049h` appears exactly once (at attach) and vim/less/htop sessions
* inside the pane emit zero. So the only thing stripped here is tmux's own smcup.
*
* Why it is gated on `useMux`: `startShell()`/`startInteractive()` fall back to a
* DIRECT PTY when mux creation fails. There the inner program's `\x1b[?1049h` really
* does reach xterm, and stripping it would break vim/less/htop for real.
*
* Why it is narrower than the full strip: with tmux `mouse off`, a mouse-aware
* program in the pane (htop, vim with `set mouse=a`) still gets its DECSETs passed
* through to the client, so stripping those would break its mouse support. And
* `\x1b[3J` from a user's own `clear` is a deliberate "wipe my scrollback".
*/
export function isMuxAltScreenOnlyStripMode(mode: SessionMode, useMux: boolean): boolean {
return useMux && !isAltScreenStripMode(mode);
}
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
/** PTY fallback geometry when tmux can't be queried (matches pre-#80 hardcoded values). */
@@ -195,6 +227,8 @@ const IS_TEST_MODE = !!process.env.VITEST;
const TEST_PTY_SCRIPT = 'if (process.stdin.isTTY) process.stdin.setRawMode(true); process.stdin.pipe(process.stdout);';
/** Delay before the in-container Claude CLI version probe (lets the container start). */
const DOCKER_CLI_VERSION_PROBE_DELAY_MS = 3000;
/** Delay before the over-ssh Claude CLI version probe (keeps session start off the ssh round-trip). */
const REMOTE_CLI_VERSION_PROBE_DELAY_MS = 3000;
/**
* Ask tmux for the current window geometry of `muxName` so a re-attaching PTY
@@ -509,6 +543,8 @@ export class Session extends EventEmitter {
tmuxHistoryLimit?: number;
/** Restored per-session attachment history. May include server-private external paths. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/** Restored wall-clock ms of the pane's last Enter (see `lastSubmitAt`). */
lastSubmitAt?: number;
/** Remote execution metadata for sessions launched through SSH inside local tmux. */
remote?: SessionRemote;
/** Docker execution metadata for sessions launched inside a container via local tmux. */
@@ -535,6 +571,12 @@ export class Session extends EventEmitter {
this._lastActivityAt = this.createdAt;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = config.resumeSessionId || this.id;
// Restored from state.json on boot recovery. start() resets _claudeSessionId
// to the launch id even when re-attaching to a mux session whose CLI has
// moved on (a `/clear` before the restart), so this anchor is what lets the
// response viewer re-derive the live conversation without waiting for the
// user to type again.
this._lastSubmitAt = config.lastSubmitAt ?? 0;
this._mux = config.mux || null;
this._useMux = config.useMux ?? (this._mux !== null && this._mux.isAvailable());
this._muxSession = config.muxSession || null;
@@ -725,6 +767,15 @@ export class Session extends EventEmitter {
return this._muxSession?.muxName ?? null;
}
/**
* True when this session's PTY is a tmux client rather than the program itself.
* Read by the replay-side alt-screen strip, which must apply the same
* `useMux` gate as the live strip (isMuxAltScreenOnlyStripMode).
*/
get usesMux(): boolean {
return this._useMux;
}
get totalCost(): number {
return this._totalCost;
}
@@ -1135,6 +1186,7 @@ export class Session extends EventEmitter {
// recovery can re-attach.
respawnBlocked: this._respawnBlocked || undefined,
attachmentHistory: this.attachmentHistory.length > 0 ? this.attachmentHistory : undefined,
lastSubmitAt: this._lastSubmitAt || undefined,
// envOverrides intentionally NOT on the public SessionState type — they must not
// leak into SSE / GET /api/sessions broadcasts (schema allows OPENCODE_*, which
// can carry secrets). For disk persistence, session-manager calls
@@ -1392,9 +1444,16 @@ export class Session extends EventEmitter {
// SSE/WS stream carries them, keeping everything in the main buffer with
// scrollback intact. These are controlled TUIs whose cursor-positioned
// redraws overwrite only the cells they target, so non-erased rows keep
// their content. Gated to Codex/Claude (isAltScreenStripMode) — shell must
// keep the alt screen for vim/less/htop.
if (isAltScreenStripMode(this.mode)) {
// their content. Gated to Codex/Claude/Gemini (isAltScreenStripMode).
//
// Every OTHER mode (shell/opencode/antigravity) gets the NARROW strip when it
// is tmux-backed: alt-screen toggles only, because the sequence that breaks
// scrollback there is tmux's own client-side smcup at attach, not anything the
// program in the pane emitted (issue #205, see isMuxAltScreenOnlyStripMode).
// 3J and the mouse DECSETs stay, so `clear` and mouse-aware TUIs keep working.
const fullStrip = isAltScreenStripMode(this.mode);
const altOnlyStrip = !fullStrip && isMuxAltScreenOnlyStripMode(this.mode, this._useMux);
if (fullStrip || altOnlyStrip) {
// Reassemble sequences split across PTY chunk boundaries first: a chunk
// ending mid-sequence ('\x1b[?104' now, '9h' next) would slip past the
// strip below and leave xterm stuck in the scrollback-less alt buffer
@@ -1410,13 +1469,15 @@ export class Session extends EventEmitter {
data = data.slice(0, -splitTail[0].length);
if (!data) return;
}
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
// eslint-disable-next-line no-control-regex
data = data.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '');
if (fullStrip) {
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
}
}
// Scan terminal output for attachment requests. `codeman://attach?...` is an
@@ -1476,8 +1537,8 @@ export class Session extends EventEmitter {
// never show it — which left cliVersion undefined and silently disabled
// wheel-forwarding to Claude's own transcript (the only route to history in
// repaint/alt-screen mode; issue #154). Remote sessions run claude on
// another host, so a local probe wouldn't reflect their version — skip them
// and let the banner scrape handle those. Cached process-wide, best-effort.
// another host, so a local probe wouldn't reflect their version; they get
// their own over-ssh probe below. Cached process-wide, best-effort.
if (this.mode === 'claude' && !this._remote && !this._docker && !this._cliVersion) {
const probedVersion = getClaudeCliVersion();
if (probedVersion) {
@@ -1516,6 +1577,32 @@ export class Session extends EventEmitter {
}, DOCKER_CLI_VERSION_PROBE_DELAY_MS);
}
// Remote sessions run claude on ANOTHER HOST, so neither the local nor the
// docker probe applies, and the banner-scrape fallback they were left with
// is the unreliable path #154 was filed for, so remote Claude cases silently
// never got wheel-forwarding (noted in the #205 analysis). Probe over ssh,
// deferred so session start never waits on the ssh round-trip.
if (this.mode === 'claude' && this._remote && !this._cliVersion) {
const remoteMeta = this._remote;
setTimeout(() => {
if (this._isStopped || this._cliVersion) return;
void probeRemoteCliVersion(remoteMeta, this.mode)
.then((version) => {
if (!version || this._isStopped || this._cliVersion) return;
this._cliVersion = version;
this.emit('cliInfoUpdated', {
version: this._cliVersion,
model: this._cliModel,
accountType: this._cliAccountType,
latestVersion: this._cliLatestVersion,
});
})
.catch(() => {
/* best-effort */
});
}, REMOTE_CLI_VERSION_PROBE_DELAY_MS);
}
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
@@ -2543,26 +2630,28 @@ export class Session extends EventEmitter {
* ```
*/
write(data: string): void {
this._trackCodexSubmit(data);
this._trackSubmit(data);
if (this.ptyProcess) {
this.ptyProcess.write(data);
}
}
// ── Codex thread tracking ─────────────────────────────────────────────
// When a codex pane last submitted a message (Enter). The response-viewer
// correlates this against ~/.codex/history.jsonl entry timestamps to find
// the thread the pane is ACTUALLY on — the only signal that survives
// /resume, /new and /fork typed inside the codex TUI itself.
private _codexLastSubmitAt = 0;
// ── Conversation tracking ─────────────────────────────────────────────
// When this pane last submitted a message (Enter). The response-viewer
// correlates this against the CLI's own history.jsonl entry timestamps to
// find the conversation the pane is ACTUALLY on — the only signal that
// survives /clear, /resume, /new and /fork typed inside the TUI itself,
// none of which announce themselves on the PTY's stdout.
private _lastSubmitAt = 0;
get codexLastSubmitAt(): number {
return this._codexLastSubmitAt;
/** Wall-clock ms of this pane's last Enter; 0 if it has never submitted. */
get lastSubmitAt(): number {
return this._lastSubmitAt;
}
private _trackCodexSubmit(data: string): void {
if (this.mode === 'codex' && (data.includes('\r') || data.includes('\n'))) {
this._codexLastSubmitAt = Date.now();
private _trackSubmit(data: string): void {
if (data.includes('\r') || data.includes('\n')) {
this._lastSubmitAt = Date.now();
}
}
@@ -2619,7 +2708,7 @@ export class Session extends EventEmitter {
* ```
*/
async writeViaMux(data: string): Promise<boolean> {
this._trackCodexSubmit(data);
this._trackSubmit(data);
if (this._mux && this._muxSession) {
return this._mux.sendInput(this.id, data);
}
+1 -1
View File
@@ -982,7 +982,7 @@ export function dockerTmuxSessionName(sessionId: string): string {
const RESUME_ID_SAFE = /^[A-Za-z0-9._-]+$/;
/**
* Append the CLI-specific resume flag to a pane command (codex/gemini). Only fires
* Append the CLI-specific resume flag to a pane command (codex/gemini/antigravity). Only fires
* when the in-container tmux is RE-CREATED (`new-session -A` makes the flag inert
* on a live reattach), i.e. exactly when the previous live agent was lost and we
* want to resume the conversation from the bind-mounted transcript. Claude mode
+16
View File
@@ -97,6 +97,22 @@ export interface FilesystemBrowseData {
truncated: boolean;
}
/** Response payload for `PUT /api/sessions/:id/file-content` (File Viewer edit mode). */
export interface FileWriteData {
/** Workspace-relative path as submitted */
path: string;
/** Size of the written content in bytes */
size: number;
/** mtime of the file after the write */
mtimeMs: number;
/** sha256 hex of the written bytes — the client's next baseHash */
hash: string;
/** Line count of the written content */
totalLines: number;
/** Line-ending style that was applied */
eol: 'lf' | 'crlf';
}
export type CleanupResourceType = 'timer' | 'interval' | 'watcher' | 'listener' | 'stream';
/**
+9
View File
@@ -480,6 +480,15 @@ export interface SessionState {
effort?: EffortLevel;
/** Sanitized per-session attachment history. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/**
* Wall-clock ms of this pane's last Enter (Session.lastSubmitAt). Persisted
* because it is the response-viewer's only anchor for re-deriving the pane's
* live conversation after a Codeman restart: `start()` resets
* `claudeSessionId` to the launch id even when re-attaching to a mux session
* whose CLI has since moved on via `/clear`, and the correlation cannot run
* again until the pane's own Enter is known.
*/
lastSubmitAt?: number;
/**
* PTY-exit circuit breaker tripped — respawn blocked until an explicit restart
* (COD-118). Runtime-only: never restored on boot (fresh server = fresh breaker).
+1 -1
View File
@@ -28,7 +28,7 @@ let _claudeDir: string | null = null;
/**
* Returns true if the Claude CLI binary can be located (via `which` or one of
* the common install directories). Mirrors `isGeminiAvailable`/`isOpenCodeAvailable`/
* the common install directories). Mirrors `isGeminiAvailable`/`isAntigravityAvailable`/`isOpenCodeAvailable`/
* `isCodexAvailable` in the sibling resolvers.
*/
export function isClaudeAvailable(): boolean {
+1 -1
View File
@@ -1,7 +1,7 @@
/**
* @fileoverview Resolve the `cloudflared` binary across common install paths.
*
* Mirrors the CLI resolvers (gemini-cli-resolver.ts et al), for the same reason
* Mirrors the CLI resolvers (antigravity-cli-resolver.ts et al), for the same reason
* they exist: the welcome screen should not offer a button whose only possible
* outcome is an error toast.
*
+124 -27
View File
@@ -349,6 +349,22 @@ const DEFAULT_SHORTCUTS = [
bindings: [{ modifiers: ['ctrl'], key: 'l' }],
action: 'clearTerminal',
},
{
id: 'copy-selection',
group: 'Terminal',
label: 'Copy Selection',
// Bindings match on `key`, not `code`: xterm decides which byte to emit from the
// PRODUCED character, so intercepting a physical KeyC that doesn't produce "c"
// would diverge from the chord that actually sends ^C.
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js and
// deliberately absent from SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults on a match, which would cost the user the interrupt key.
action: 'copyTerminalSelection',
},
{
id: 'increase-font',
group: 'Terminal',
@@ -495,7 +511,14 @@ class CodemanApp {
this._initGeneration = 0; // dedup concurrent handleInit calls
this._initFallbackTimer = null; // fallback timer if SSE init doesn't arrive
this._selectGeneration = 0; // cancel stale selectSession loads
this._initialFullBufferLoad = true; // first buffer load after a page load fetches full tmux scrollback (COD-47)
// Sessions whose full tmux scrollback has already been replayed this page load
// (COD-47). Tracked PER SESSION rather than as a single "first load" flag: the
// flag was consumed by whichever session auto-selected at page load, so every
// OTHER tab started life with one visible frame of history (issue #205).
this._fullHistoryLoaded = new Set();
// Cooldown per session for the scroll-to-top "load more history" re-pull.
this._fullHistoryRepullAt = new Map(); // Map<sessionId, timestamp>
this._fullHistoryRepullInFlight = false;
this.terminalLoadStates = new Map(); // Map<sessionId, { generation, phase }>
this.respawnStatus = {};
this.respawnTimers = {}; // Track timed respawn timers
@@ -1907,6 +1930,37 @@ class CodemanApp {
}
}
/** Build one response-viewer message so the brief and full views share markup and CSS. */
_buildResponseViewerMessage(text, role, agentLabel) {
const div = document.createElement('div');
const isUser = role === 'user';
div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');
const roleBadge = document.createElement('div');
roleBadge.className = 'rv-role ' + (isUser ? 'rv-role-user' : 'rv-role-assistant');
roleBadge.textContent = isUser ? 'You' : agentLabel;
div.appendChild(roleBadge);
const renderedText = document.createElement('div');
renderedText.className = 'rv-text';
renderedText.innerHTML = this._renderMarkdown(text);
div.appendChild(renderedText);
return div;
}
_getResponseViewerAgentLabel() {
const mode = this.sessions.get(this.activeSessionId)?.mode;
return mode === 'codex'
? 'Codex'
: mode === 'gemini'
? 'Gemini'
: mode === 'antigravity'
? 'Antigravity'
: mode === 'opencode'
? 'OpenCode'
: 'Claude';
}
async toggleResponseViewer() {
const viewer = document.getElementById('responseViewer');
const backdrop = document.getElementById('responseViewerBackdrop');
@@ -1929,7 +1983,7 @@ class CodemanApp {
// Source 2: Terminal buffer fallback — strip ANSI, drop Claude CLI chrome.
// Claude + shell only: _cleanTerminalBuffer knows Claude CLI's output, and
// shell sessions have no transcript source at all; for TUI modes
// (codex/opencode/gemini) it yields repaint garbage, so a clear
// (codex/opencode/gemini/antigravity) it yields repaint garbage, so a clear
// placeholder beats a messy screen dump there.
const sessionMode = this.sessions.get(this.activeSessionId)?.mode || 'claude';
if (!lastResponse && (sessionMode === 'claude' || sessionMode === 'shell')) {
@@ -1942,7 +1996,11 @@ class CodemanApp {
const body = document.getElementById('responseViewerBody');
if (lastResponse) {
body.innerHTML = this._renderMarkdown(lastResponse);
// Keep the brief view inside the same message wrapper as the full
// conversation view. The wrapper supplies the card, role badge and
// descendant markdown styles that direct body children do not get.
body.innerHTML = '';
body.appendChild(this._buildResponseViewerMessage(lastResponse, 'assistant', this._getResponseViewerAgentLabel()));
this._bindResponseViewerInteractions(body);
} else {
body.textContent =
@@ -1982,26 +2040,10 @@ class CodemanApp {
}
// Render conversation thread
const mode = this.sessions.get(this.activeSessionId)?.mode;
const agentLabel =
mode === 'codex' ? 'Codex' : mode === 'gemini' ? 'Gemini' : mode === 'antigravity' ? 'Antigravity' : mode === 'opencode' ? 'OpenCode' : 'Claude';
const agentLabel = this._getResponseViewerAgentLabel();
body.innerHTML = '';
for (const msg of messages) {
const div = document.createElement('div');
const isUser = msg.role === 'user';
div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');
const role = document.createElement('div');
role.className = 'rv-role ' + (isUser ? 'rv-role-user' : 'rv-role-assistant');
role.textContent = isUser ? 'You' : agentLabel;
div.appendChild(role);
const text = document.createElement('div');
text.className = 'rv-text';
text.innerHTML = this._renderMarkdown(msg.text);
div.appendChild(text);
body.appendChild(div);
body.appendChild(this._buildResponseViewerMessage(msg.text, msg.role, agentLabel));
}
this._bindResponseViewerInteractions(body);
@@ -4063,6 +4105,58 @@ class CodemanApp {
this.terminal.write('\x1b[3J\x1b[H\x1b[2J');
}
/**
* "Load more history": re-pull the whole tmux scrollback when the user scrolls up
* while already at the top of what the browser has.
*
* xterm's buffer is only ever a WINDOW onto tmux's real history, and two things
* shrink it. tmux repaints the pane rectangle instead of emitting linefeeds
* whenever output outpaces its flush interval, which OVERWRITES already-rendered
* scrollback rather than pushing rows into it (measured: a 60-line burst added 1
* row and destroyed 34, while the same 60 lines emitted slowly added all 60). And
* a tab switch replays only the visible frame. Either way tmux still holds
* everything (history-limit 100k by default), so the fix is to go ask for it with
* the same `?full=1` capture a page reload uses (issue #205).
*
* On demand rather than automatic because that capture is unbounded-ish work: at
* the default history limit it can be megabytes, which is fine to pay when the
* user is explicitly reaching for history and not fine on every tab switch.
*/
async _maybeRefetchFullHistory() {
const sessionId = this.activeSessionId;
if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return;
if (this.detachedSessions?.has(sessionId)) return;
const now = Date.now();
// Momentum scrolling fires this dozens of times per flick, and a burst of new
// output is the normal reason to want a re-pull, so cooldown rather than latch.
if (now - (this._fullHistoryRepullAt.get(sessionId) || 0) < 4000) return;
this._fullHistoryRepullAt.set(sessionId, now);
this._fullHistoryRepullInFlight = true;
try {
const res = await fetch(`/api/sessions/${sessionId}/terminal?full=1`);
const buffer = (await res.json())?.data?.terminalBuffer;
// Bail on a tab switch mid-fetch: writing here would paint another session's
// history into the terminal the user is now looking at.
if (!buffer || this.activeSessionId !== sessionId) return;
const rowsBefore = this.terminal.buffer.active.length;
this._resetTerminalForReplay();
await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
if (this.activeSessionId !== sessionId) return;
this.terminalBufferCache.set(sessionId, buffer);
// Hold the user's place. The replay is a superset that grew the buffer
// UPWARD, so what used to be row 0 (what they were looking at) is now `delta`
// rows down; scrolling there reveals the recovered history above it instead
// of teleporting them to the bottom the way a normal buffer load does.
const delta = this.terminal.buffer.active.length - rowsBefore;
if (delta > 0) this.terminal.scrollToLine(delta);
else this.terminal.scrollToTop();
} catch {
// Transient (offline, 5xx) — the next scroll-up past the cooldown retries.
} finally {
this._fullHistoryRepullInFlight = false;
}
}
_shouldFocusTerminalForTabSwitch() {
if (typeof MobileDetection === 'undefined' || !MobileDetection.isTouchDevice()) {
return true;
@@ -4246,7 +4340,7 @@ class CodemanApp {
// (viewport + scrollback + colors) for an instant first paint. For codex
// this is also a correctness fix — its byte-stream replay shows only the
// latest TUI frame (the idle welcome banner) because codex doesn't include
// earlier conversation in its current redraw. For claude/opencode/gemini
// earlier conversation in its current redraw. For claude/opencode/gemini/antigravity
// the replay is already complete, so the snapshot is purely a faster,
// scroll-preserving first paint before the canonical fetch reconciles.
//
@@ -4333,11 +4427,14 @@ class CodemanApp {
this._setTerminalLoadState(sessionId, selectGen, 'fetching');
_crashDiag.log('FETCH_START');
// The FIRST buffer load after a page load requests the full tmux scrollback
// (?full=1, COD-47) so history that scrolled off the server's byte buffer
// comes back after a reload. Tab switches keep the fast ?tail= frame path.
const useFullHistory = this._initialFullBufferLoad === true;
this._initialFullBufferLoad = false;
// The first load OF EACH SESSION this page load requests the full tmux
// scrollback (?full=1, COD-47) so history that scrolled off the server's byte
// buffer comes back. Later switches to an already-replayed session keep the
// fast ?tail= frame path, which is why this is a Set and not a flag: the flag
// version gave the full replay to the auto-selected tab and one frame of
// history to every other one (issue #205).
const useFullHistory = !this._fullHistoryLoaded.has(sessionId);
if (useFullHistory) this._fullHistoryLoaded.add(sessionId);
const res = await fetch(
useFullHistory
? `/api/sessions/${sessionId}/terminal?full=1`
+4
View File
@@ -451,6 +451,7 @@
'Respawn Blocked': '重生已阻止',
'Task Complete': '任务完成',
'Copied to clipboard': '已复制到剪贴板',
'Failed to copy': '复制失败',
'Checking…': '正在检查…',
'Starting…': '正在启动…',
'Starting update…': '正在开始更新…',
@@ -600,6 +601,9 @@
'Select a run to view its agents': '选择一次运行以查看其智能体',
'Source type filter': '来源类型筛选',
'Copy content': '复制内容',
'Edit file': '编辑文件',
'Unsaved changes': '未保存的更改',
Saved: '已保存',
'Export as JSON': '导出为 JSON',
'Export as Markdown': '导出为 Markdown',
'Mark all read': '全部标为已读',
+14 -1
View File
@@ -323,6 +323,10 @@
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run OpenCode
</button>
<button class="welcome-btn welcome-btn-antigravity" id="welcomeAntigravityBtn" style="display: none;" onclick="app.setRunMode('antigravity'); app.runAntigravity()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Antigravity
</button>
<button class="welcome-btn welcome-btn-gemini" id="welcomeGeminiBtn" style="display: none;" onclick="app.setRunMode('gemini'); app.runGemini()">
<svg width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2"><polygon points="5 3 19 12 5 21 5 3"/></svg>
Run Gemini
@@ -422,11 +426,18 @@
<div class="file-preview-header">
<span class="file-preview-title" id="filePreviewTitle">file.ts</span>
<div class="file-preview-actions">
<button class="btn-icon-sm file-preview-edit-btn" id="filePreviewEditBtn" onclick="app.enterFilePreviewEdit()" title="Edit file" aria-label="Edit file" hidden><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M17 3a2.85 2.83 0 1 1 4 4L7.5 20.5 2 22l1.5-5.5z"/></svg></button>
<button class="btn-icon-sm" onclick="app.copyFilePreviewContent()" title="Copy content">&#x2398;</button>
<button class="btn-icon-sm" onclick="app.closeFilePreview()" title="Close">&times;</button>
</div>
</div>
<div class="file-preview-body" id="filePreviewBody"></div>
<div class="file-preview-editbar" id="filePreviewEditBar" hidden>
<span class="file-preview-dirty" id="filePreviewDirty" hidden>Unsaved changes</span>
<span class="file-preview-editbar-spacer"></span>
<button class="file-preview-editbar-btn" onclick="app.cancelFilePreviewEdit()">Cancel</button>
<button class="file-preview-editbar-btn file-preview-editbar-btn--save" id="filePreviewSaveBtn" onclick="app.saveFilePreviewEdit()" disabled>Save</button>
</div>
<div class="file-preview-footer" id="filePreviewFooter"></div>
</div>
</div>
@@ -639,6 +650,8 @@
<section class="shortcut-section">
<h4>Terminal</h4>
<div class="shortcuts-grid">
<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection (interrupts when nothing is selected)</div>
<div><kbd>Ctrl</kbd>+<kbd>Shift</kbd>+<kbd>C</kbd></div><div>Copy Selection</div>
<div><kbd>Ctrl</kbd>+<kbd>L</kbd></div><div>Clear Terminal</div>
<div><kbd>Ctrl</kbd>+<kbd>+</kbd></div><div>Increase Font</div>
<div><kbd>Ctrl</kbd>+<kbd>-</kbd></div><div>Decrease Font</div>
@@ -2183,7 +2196,7 @@
<div class="form-row">
<label>Image</label>
<input type="text" id="dockerImage" placeholder="codeman/agent:base" autocomplete="off" autocapitalize="off" spellcheck="false">
<span class="form-hint">Build it once with <code>node scripts/build-agent-image.mjs</code>. Contains node + claude/codex/gemini + tmux.</span>
<span class="form-hint">Build it once with <code>node scripts/build-agent-image.mjs</code>. Contains node + claude/codex/gemini/opencode/agy + tmux.</span>
</div>
<div class="form-row">
<label>Network</label>
+6
View File
@@ -486,6 +486,11 @@ Object.assign(CodemanApp.prototype, {
* The Run picker: the same backends as the toolbar's run-mode menu, plus saved
* web tabs. Deliberately no "Recent Sessions" block, unlike the toolbar menu:
* past conversations have their own section further down this screen.
*
* Gated the same way as the toolbar's #runModeMenu (isCliAvailable(), shell
* exempt) — this list is a separate, hardcoded duplicate of the toolbar's menu
* rather than a shared render, so it never picked up #201's gating and offered
* every backend regardless of what's actually installed.
*/
_buildMobileOverviewRunMenu() {
const menu = document.createElement('div');
@@ -493,6 +498,7 @@ Object.assign(CodemanApp.prototype, {
const current = this.runMode || 'claude';
for (const entry of MOBILE_OVERVIEW_RUN_MODES) {
if (entry.mode !== 'shell' && !this.isCliAvailable(entry.mode)) continue;
const option = document.createElement('button');
option.type = 'button';
option.className = 'mobile-overview-run-option' + (entry.mode === current ? ' selected' : '');
+27
View File
@@ -1871,6 +1871,33 @@ html.mobile-init .file-browser-panel {
bottom: calc(44px + 2rem + var(--safe-area-bottom));
}
/* File preview window: full screen on phones. --app-height tracks the visual
viewport (KeyboardHandler), so the edit textarea + Save bar stay above the
OS keyboard instead of hiding behind it. Footer/edit bar pad for the home
indicator. */
.file-preview-window {
width: 100vw;
max-width: 100vw;
height: var(--app-height, 100vh);
max-height: var(--app-height, 100vh);
border-radius: 0;
border-left: none;
border-right: none;
}
.file-preview-editbar {
padding-bottom: calc(0.4rem + var(--safe-area-bottom));
}
.file-preview-overlay .file-preview-footer {
padding-bottom: calc(0.35rem + var(--safe-area-bottom));
}
/* >=16px or iOS Safari auto-zooms the page on focus */
.file-preview-body textarea.file-preview-editor {
font-size: 16px;
}
/* Notification drawer - full width on mobile, includes safe area padding */
.notification-drawer {
width: 100%;
+185 -2
View File
@@ -3200,6 +3200,9 @@ Object.assign(CodemanApp.prototype, {
if (!overlay || !bodyEl) return;
// Edit mode: reset any prior editor state whenever a preview (re)loads.
this._resetFilePreviewEdit();
// Show overlay with loading state
overlay.classList.add('visible');
titleEl.textContent = filePath;
@@ -3298,6 +3301,13 @@ Object.assign(CodemanApp.prototype, {
bodyEl.innerHTML = `<pre><code>${escapeHtml(data.content)}</code></pre>`;
const truncNote = data.truncated ? ` (showing 500/${data.totalLines} lines)` : '';
footerEl.textContent = `${data.totalLines} lines \u2022 ${this.formatFileSize(data.size)}${truncNote}`;
// Edit affordance only when the server says an edit=1 re-fetch would
// succeed (workspace text file inside the allowlist and size cap).
if (data.editable) {
this.filePreviewEditTarget = { sessionId, filePath };
const editBtn = this.$('filePreviewEditBtn');
if (editBtn) editBtn.hidden = false;
}
}
} catch (err) {
console.error('Failed to preview file:', err);
@@ -3306,6 +3316,8 @@ Object.assign(CodemanApp.prototype, {
},
closeFilePreview() {
if (this.filePreviewEdit?.dirty && !confirm('Discard unsaved changes?')) return;
this._resetFilePreviewEdit();
const overlay = this.$('filePreviewOverlay');
if (overlay) {
overlay.classList.remove('visible');
@@ -3313,6 +3325,172 @@ Object.assign(CodemanApp.prototype, {
this.filePreviewContent = '';
},
// ═══════════════════════════════════════════════════════════════
// File Viewer edit mode (issue #212 — docs/file-viewer-edit-plan.md)
// ═══════════════════════════════════════════════════════════════
_resetFilePreviewEdit() {
this.filePreviewEdit = null;
this.filePreviewEditTarget = null;
const editBtn = this.$('filePreviewEditBtn');
if (editBtn) editBtn.hidden = true;
const editBar = this.$('filePreviewEditBar');
if (editBar) editBar.hidden = true;
const dirtyEl = this.$('filePreviewDirty');
if (dirtyEl) dirtyEl.hidden = true;
const saveBtn = this.$('filePreviewSaveBtn');
if (saveBtn) {
saveBtn.disabled = true;
saveBtn.textContent = 'Save';
}
},
async enterFilePreviewEdit() {
const target = this.filePreviewEditTarget;
if (!target || this.filePreviewEdit) return;
const bodyEl = this.$('filePreviewBody');
const footerEl = this.$('filePreviewFooter');
if (!bodyEl) return;
// Always re-fetch with edit=1: the preview buffer may be line-truncated and
// a truncated buffer must never become an edit buffer. Parse the envelope
// even on non-ok responses so the specific refusal ("too large to edit
// here") reaches the toast instead of a generic failure.
let data;
try {
const res = await fetch(
`/api/sessions/${target.sessionId}/file-content?path=${encodeURIComponent(target.filePath)}&edit=1`
);
const result = await res.json().catch(() => null);
if (!result || result.success !== true) {
throw new Error(result?.error || `Failed to load file for editing (HTTP ${res.status})`);
}
data = result.data;
} catch (err) {
this.showToast(err.message, 'error');
return;
}
this.filePreviewEdit = {
sessionId: target.sessionId,
filePath: target.filePath,
baseHash: data.hash,
eol: data.eol,
original: data.content,
dirty: false,
saving: false,
};
const textarea = document.createElement('textarea');
textarea.className = 'file-preview-editor';
textarea.spellcheck = false;
textarea.setAttribute('autocapitalize', 'off');
textarea.setAttribute('autocorrect', 'off');
textarea.setAttribute('autocomplete', 'off');
textarea.wrap = 'off';
textarea.value = data.content;
textarea.addEventListener('input', () => this._onFilePreviewEditInput());
bodyEl.innerHTML = '';
bodyEl.appendChild(textarea);
// Deliberately no autofocus: on phones that would pop the OS keyboard
// before the user has scrolled to the line they want to change.
const editBtn = this.$('filePreviewEditBtn');
if (editBtn) editBtn.hidden = true;
const editBar = this.$('filePreviewEditBar');
if (editBar) editBar.hidden = false;
if (footerEl) {
const eolNote = data.eol === 'crlf' ? ' • CRLF' : '';
footerEl.textContent = `Editing • ${data.totalLines} lines • ${this.formatFileSize(data.size)}${eolNote}`;
}
},
_onFilePreviewEditInput() {
const edit = this.filePreviewEdit;
if (!edit) return;
const textarea = this.$('filePreviewBody')?.querySelector('textarea.file-preview-editor');
if (!textarea) return;
edit.dirty = textarea.value !== edit.original;
const dirtyEl = this.$('filePreviewDirty');
if (dirtyEl) dirtyEl.hidden = !edit.dirty;
const saveBtn = this.$('filePreviewSaveBtn');
if (saveBtn) saveBtn.disabled = !edit.dirty || edit.saving;
},
cancelFilePreviewEdit() {
const edit = this.filePreviewEdit;
if (!edit) return;
if (edit.dirty && !confirm('Discard unsaved changes?')) return;
const { sessionId, filePath } = edit;
this._resetFilePreviewEdit();
this.openFilePreview(filePath, sessionId);
},
async saveFilePreviewEdit(force = false) {
const edit = this.filePreviewEdit;
if (!edit || edit.saving) return;
const textarea = this.$('filePreviewBody')?.querySelector('textarea.file-preview-editor');
if (!textarea) return;
edit.saving = true;
const saveBtn = this.$('filePreviewSaveBtn');
if (saveBtn) {
saveBtn.disabled = true;
saveBtn.textContent = 'Saving…';
}
const restoreSaveState = () => {
edit.saving = false;
if (saveBtn) saveBtn.textContent = 'Save';
this._onFilePreviewEditInput();
};
let result = null;
let status = 0;
try {
const res = await fetch(`/api/sessions/${edit.sessionId}/file-content`, {
method: 'PUT',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({
path: edit.filePath,
content: textarea.value,
baseHash: edit.baseHash,
eol: edit.eol ?? undefined, // Zod .optional() rejects null
force: force || undefined,
}),
});
status = res.status;
result = await res.json().catch(() => null);
} catch (err) {
restoreSaveState();
this.showToast(`Save failed: ${err.message}`, 'error');
return;
}
if (status === 409 || result?.errorCode === 'CONFLICT') {
restoreSaveState();
if (
confirm(
'File changed on disk since you loaded it.\nOK overwrites it with your version; Cancel keeps your draft open.'
)
) {
this.saveFilePreviewEdit(true);
}
return;
}
if (!result || result.success !== true) {
restoreSaveState();
this.showToast(`Save failed: ${result?.error || `HTTP ${status}`}`, 'error');
return;
}
const { sessionId, filePath } = edit;
this._resetFilePreviewEdit();
this.showToast('Saved', 'success');
// Re-open in read mode — re-fetching shows the truth on disk (including the
// server-side EOL normalization) rather than trusting the local buffer.
this.openFilePreview(filePath, sessionId);
},
// ═══════════════════════════════════════════════════════════════
// Attachment Cards (detected documents/images)
// ═══════════════════════════════════════════════════════════════
@@ -3749,8 +3927,13 @@ Object.assign(CodemanApp.prototype, {
},
copyFilePreviewContent() {
if (this.filePreviewContent) {
navigator.clipboard.writeText(this.filePreviewContent).then(() => {
// While editing, copy the live editor buffer (not the stale preview text).
const editTextarea = this.filePreviewEdit
? this.$('filePreviewBody')?.querySelector('textarea.file-preview-editor')
: null;
const content = editTextarea ? editTextarea.value : this.filePreviewContent;
if (content) {
navigator.clipboard.writeText(content).then(() => {
this.showToast('Copied to clipboard', 'success');
}).catch(() => {
this.showToast('Failed to copy', 'error');
+1 -1
View File
@@ -1,5 +1,5 @@
/**
* @fileoverview Quick start (case loading, session spawning for Claude/Shell/OpenCode/Codex/Gemini),
* @fileoverview Quick start (case loading, session spawning for Claude/Shell/OpenCode/Codex/Gemini/Antigravity),
* session options modal (per-session settings, color picker, rename),
* session options tabs (Ralph config tab), case settings (CRUD, links),
* create case modal, and mobile case picker.
+1
View File
@@ -748,6 +748,7 @@ Object.assign(CodemanApp.prototype, {
const buttons = [
['welcomeClaudeBtn', 'claude'],
['welcomeOpencodeBtn', 'opencode'],
['welcomeAntigravityBtn', 'antigravity'],
['welcomeGeminiBtn', 'gemini'],
// Not a run mode, same reasoning: offering a Cloudflare Tunnel on a box
// without cloudflared can only ever produce "cloudflared not found".
+95
View File
@@ -3305,6 +3305,23 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
transform: translateY(-1px);
}
/* Antigravity: cyan identity, matching .btn-toolbar.btn-run.mode-antigravity and
.run-mode-dot.antigravity so the welcome action reads as the same backend. */
.welcome-btn-antigravity {
background: linear-gradient(135deg, #0b2b33 0%, #0e7490 55%, #0891b2 100%);
border-color: rgba(34, 211, 238, 0.4);
color: #cffafe;
box-shadow: 0 2px 8px rgba(34, 211, 238, 0.16), inset 0 1px 0 rgba(255, 255, 255, 0.06);
}
.welcome-btn-antigravity:hover {
background: linear-gradient(135deg, #124450 0%, #0891b2 55%, #06b6d4 100%);
box-shadow: 0 4px 20px rgba(34, 211, 238, 0.3), 0 0 40px rgba(8, 145, 178, 0.12), inset 0 1px 0 rgba(255, 255, 255, 0.08);
border-color: rgba(103, 232, 249, 0.5);
color: #ecfeff;
transform: translateY(-1px);
}
.welcome-btn-gemini {
background: linear-gradient(135deg, #10243f 0%, #174ea6 55%, #4f46e5 100%);
border-color: rgba(96, 165, 250, 0.4);
@@ -9430,6 +9447,84 @@ kbd {
flex-shrink: 0;
}
/* ---- File Viewer edit mode (issue #212) ---- */
.file-preview-body textarea.file-preview-editor {
display: block;
width: 100%;
height: 100%;
margin: 0;
padding: 0.75rem;
border: none;
outline: none;
resize: none;
background: var(--bg-dark);
color: var(--text);
font-family: var(--font-mono);
font-size: 0.8rem;
line-height: 1.5;
white-space: pre;
overflow-wrap: normal;
overflow: auto;
tab-size: 4;
}
.file-preview-editbar {
display: flex;
align-items: center;
gap: 0.5rem;
padding: 0.4rem 0.75rem;
border-top: 1px solid var(--border);
background: var(--bg-input);
flex-shrink: 0;
}
.file-preview-editbar[hidden] {
display: none;
}
.file-preview-editbar-spacer {
flex: 1;
}
.file-preview-dirty {
font-size: 0.7rem;
color: var(--warning, #e5c07b);
}
.file-preview-dirty::before {
content: '\25CF ';
}
.file-preview-editbar-btn {
padding: 0.3rem 0.9rem;
font-size: 0.75rem;
border-radius: 6px;
border: 1px solid var(--control-border);
background: var(--bg-input);
color: var(--text);
cursor: pointer;
}
.file-preview-editbar-btn:hover {
background: var(--bg-hover, rgba(255, 255, 255, 0.08));
}
.file-preview-editbar-btn--save {
background: var(--accent);
border-color: var(--accent);
color: #fff;
}
.file-preview-editbar-btn--save:hover {
background: var(--accent-hover);
}
.file-preview-editbar-btn--save:disabled {
opacity: 0.45;
cursor: default;
}
/* ========== Log Viewer Windows (Floating) ========== */
.log-viewer-window {
+237 -14
View File
@@ -158,6 +158,30 @@ Object.assign(CodemanApp.prototype, {
return false;
}
// Smart copy (#211): with a selection, Ctrl+C copies it instead of sending
// ^C. With NO selection the branch must fall through (return true, and no
// preventDefault) or the interrupt key is lost, which is the whole reason
// the selection check runs before any registry dispatch. Ctrl+Shift+C is
// the explicit copy chord and never falls through: an "explicit copy" that
// interrupts a running agent because the selection happened to be empty is
// a footgun with no upside.
// NOTE: returning false does NOT cancel the event (xterm's _keyDown calls
// this handler before its own cancel()), so preventDefault is explicit:
// without it the browser runs its native copy on top of ours.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
}
return true;
}
// Ctrl+V / Cmd+V: intercept before xterm sends ^V to PTY.
// Route through our paste trap which handles both images and text.
if ((ev.ctrlKey || ev.metaKey) && ev.key === 'v' && ev.type === 'keydown') {
@@ -391,31 +415,71 @@ Object.assign(CodemanApp.prototype, {
// ignores wheel reports); older versions DO capture wheel as option
// navigation, so they keep the local wheel.
// Shift+wheel always scrolls xterm's local scrollback (Codeman's restored
// history lives there), and once the viewport left the bottom the wheel
// stays local until the user scrolls back down — so both scrollbacks stay
// reachable without a mode switch.
// history lives there); the plain wheel stays on the CLI's transcript for
// those modes regardless of scroll position, so the CLI's input box never
// slides off the screen (see _shouldForwardWheelToApp).
//
// CAPTURE phase, deliberately, and Codeman owns the scroll. xterm's
// viewport is a vscode-style ScrollableElement that consumes wheel events
// itself (preventDefault + stopPropagation) whenever it believes a
// scrollbar exists, does NOT consult attachCustomWheelEventHandler, and —
// measured on the live instance — goes DEAF after terminal.reset(): a tab
// switch or full-history replay leaves its scroll dimensions stale, after
// which wheel events neither scroll nor propagate reliably. A bubble-phase
// listener here therefore never fired once local scrollback existed
// (measured: _shouldForwardWheelToApp call count stayed 0 while xterm
// scrolled), and after a tab switch NOTHING scrolled at all — the "input
// box scrolls up then it fights", "works at first, breaks after a tab
// switch" reports on #205.
//
// So: capture runs ancestors-first; this handler sees every wheel first
// and stops propagation, keeping xterm's scroller out of it entirely.
// Local scrolling goes through terminal.scrollLines() — buffer-level, so
// it keeps working after resets — with our own deltaMode normalization
// (_wheelScrollLines) covering Firefox's line-unit wheels. Two cases still
// belong to xterm and are passed through untouched:
// - mouseTrackingMode active: xterm's own encoder forwards the wheel to
// the PTY (htop/vim with mouse on in a shell pane);
// - alternate buffer (direct-PTY fallback running vim/less): xterm's
// alt-scroll handling converts the wheel to cursor keys, which is what
// those apps expect.
container.addEventListener(
'wheel',
(ev) => {
const trackingMode = this.terminal?.modes?.mouseTrackingMode;
if (trackingMode && trackingMode !== 'none') return;
if (this.terminal?.buffer?.active?.type === 'alternate') return;
ev.preventDefault();
const lines = this._wheelScrollLines(ev);
ev.stopPropagation();
if (this._shouldForwardWheelToApp(ev)) {
this._sendSyntheticSgrWheel(ev.clientX, ev.clientY, lines);
this._forwardScrollToApp(ev.clientX, ev.clientY, this._wheelScrollLines(ev));
return;
}
// Local scrolling accumulates FRACTIONAL lines: a macOS trackpad emits
// a stream of tiny pixel deltas, and rounding each one to a whole line
// (the ±1 fallback) made slow drags scroll faster than the finger.
const lines = this._wheelScrollLinesFloat(ev);
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
this._smoothScrollBy(lines);
},
{ passive: false }
{ passive: false, capture: true }
);
// Touch scrolling — use terminal.scrollLines() for all devices.
// xterm.js DOM renderer doesn't populate xterm-viewport's scroll area,
// so native CSS scrolling (overflow-y: scroll + touch-action: pan-y)
// has nothing to scroll. Instead, convert touch deltas into scrollLines()
// calls, matching the wheel handler above.
// calls, matching the wheel handler above, including the forwarding
// branch: for the sessions whose wheel goes to the CLI's own transcript
// (_shouldForwardWheelToApp), a touch drag must go there too, or every
// phone/tablet swipe scrolls the local buffer of stale repaint frames and
// drags the CLI's pinned input box off the screen (issue #205's mobile
// half). Same gate, so Shift has no touch analog but the local-scrollback
// opt-out setting and the CLI-version gate apply to touch exactly as they
// do to the wheel.
{
const cellHeight = () => this.terminal._core?._renderService?.dimensions?.css?.cell?.height || 13;
let touchLastX = 0;
let touchLastY = 0;
let velocity = 0;
let lastTime = 0;
@@ -429,7 +493,16 @@ Object.assign(CodemanApp.prototype, {
if (!isTouching && Math.abs(velocity) > 0.3) {
// Momentum phase — convert pixel velocity to lines
const lines = Math.round(velocity / cellHeight());
if (lines !== 0) this.terminal.scrollLines(lines);
if (lines !== 0) {
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
// Flick momentum keeps feeding the CLI's transcript from the last
// touch point; the 40ms coalescer batches the per-frame reports.
this._forwardScrollToApp(touchLastX, touchLastY, lines);
} else {
this.terminal.scrollLines(lines);
this._maybeLoadMoreHistoryOnScroll(lines);
}
}
velocity *= 0.92;
scrollFrame = requestAnimationFrame(scrollLoop);
} else if (!isTouching) {
@@ -450,6 +523,7 @@ Object.assign(CodemanApp.prototype, {
'touchstart',
(ev) => {
if (ev.touches.length === 1) {
touchLastX = ev.touches[0].clientX;
touchLastY = ev.touches[0].clientY;
touchStartY = touchLastY;
velocity = 0;
@@ -485,13 +559,19 @@ Object.assign(CodemanApp.prototype, {
const delta = touchLastY - touchY; // positive = scroll down
pixelAccum += delta;
velocity = delta * 1.2;
touchLastX = ev.touches[0].clientX;
touchLastY = touchY;
// Convert accumulated pixels to whole lines
const ch = cellHeight();
const lines = Math.trunc(pixelAccum / ch);
if (lines !== 0) {
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
if (this._shouldForwardWheelToApp({ shiftKey: false })) {
this._forwardScrollToApp(touchLastX, touchLastY, lines);
} else {
this._noteTerminalUserScroll(lines);
this.terminal.scrollLines(lines);
this._maybeLoadMoreHistoryOnScroll(lines);
}
pixelAccum -= lines * ch;
}
}
@@ -1384,7 +1464,7 @@ Object.assign(CodemanApp.prototype, {
}
titleSpan.appendChild(document.createTextNode(s.name || s.firstPrompt || shortDir));
// Badge row: mode (claude/codex/opencode/gemini/shell) + a LIVE pill.
// Badge row: mode (claude/codex/opencode/gemini/antigravity/shell) + a LIVE pill.
const badgeRow = document.createElement('div');
badgeRow.className = 'history-item-badges';
if (s.mode) {
@@ -1985,6 +2065,73 @@ Object.assign(CodemanApp.prototype, {
}
},
/**
* Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the
* buffer while scrolling up is the user reaching for history the browser does
* not have, so pull the rest of tmux's scrollback (issue #205, see
* _maybeRefetchFullHistory). Must be called AFTER scrollLines(), since the
* check is on the resulting position, and it is deliberately not folded into
* _noteTerminalUserScroll for exactly that reason. Cheap: one integer compare
* per scroll event, and the pull itself is cooldown-guarded.
*/
_maybeLoadMoreHistoryOnScroll(lines) {
if (lines >= 0) return;
if (this.terminal?.buffer?.active?.viewportY === 0) this._maybeRefetchFullHistory?.();
},
/**
* Ease-out smooth scrolling for the local wheel path. The capture-phase
* wheel handler owns local scrolling (xterm's own smooth scroller is
* bypassed, see the listener comment), so without this every notch was an
* instant multi-line jump. Wheel deltas accumulate into a pending line
* count (fractional — see _wheelScrollLinesFloat) and drain ~22% per
* animation frame with a one-line floor, so a single notch starts with a
* gentle step and glides to an exact landing; more notches mid-glide deepen
* the pending count, which reads as natural acceleration. A sub-line
* residual stays pending until further input pushes it past a whole line
* (that is what makes slow trackpad drags track the finger). Direction
* reversals cancel arithmetically. The pending amount is dropped when the
* active session changes mid-glide — leftover momentum must never scroll
* the tab the user just switched to.
*/
_smoothScrollBy(lines) {
if (!lines) return;
this._smoothScrollPending = (this._smoothScrollPending || 0) + lines;
this._smoothScrollSession = this.activeSessionId;
if (this._smoothScrollFrame) return;
const step = () => {
this._smoothScrollFrame = null;
const pending = this._smoothScrollPending || 0;
if (!pending) return;
if (this.activeSessionId !== this._smoothScrollSession) {
this._smoothScrollPending = 0;
return;
}
if (Math.abs(pending) < 1) return; // sub-line residual: wait for more input
const eased = pending * 0.22;
const move = pending > 0 ? Math.max(1, Math.floor(eased)) : Math.min(-1, Math.ceil(eased));
this._smoothScrollPending = pending - move;
this.terminal.scrollLines(move);
this._maybeLoadMoreHistoryOnScroll(move);
if (Math.abs(this._smoothScrollPending) >= 1) this._smoothScrollFrame = requestAnimationFrame(step);
};
this._smoothScrollFrame = requestAnimationFrame(step);
},
/**
* Hand a scroll gesture (wheel tick or touch drag, already converted to
* lines) to the CLI as synthetic SGR wheel reports. SGR coordinates address
* the LIVE screen (the bottom `rows` of the buffer), so a report computed
* from a scrolled-up viewport would hit-test a different row entirely, and
* forwarding while the user stares at stale scrollback looks like the
* gesture is dead. Snap back first: the gesture then always acts on what the
* CLI is drawing now.
*/
_forwardScrollToApp(clientX, clientY, lines) {
if (!this._terminalViewportAtBottom()) this.terminal.scrollToBottom();
this._sendSyntheticSgrWheel(clientX, clientY, lines);
},
_hasRecentUserScrollUp() {
if (typeof this._lastUserScrollUpAt !== 'number') return false;
return performance.now() - this._lastUserScrollUpAt < window.CodemanTerminalInput.USER_SCROLL_STICKY_SUPPRESS_MS;
@@ -2612,6 +2759,46 @@ Object.assign(CodemanApp.prototype, {
// intentionally empty
},
// Registry-aware gate for the smart-copy chord (#211). Mirrors
// shouldOpenCommandPaletteFromShortcut(): honors a rebound or disabled
// 'copy-selection' entry, and falls back to the default chord when the
// registry isn't available (isolated test harnesses).
// Returning true only means "this chord asked to copy", the CALLER decides
// what happens when there is no selection, so the interrupt stays intact.
shouldCopyTerminalSelectionFromShortcut(ev) {
// The custom key handler also runs for keypress/keyup; only keydown decides.
if (!ev || ev.type !== 'keydown') return false;
// Hot path: every dispatchable chord needs Ctrl/Cmd/Alt, so plain typing
// exits before any registry work.
if (!ev.ctrlKey && !ev.metaKey && !ev.altKey) return false;
const registryAvailable =
typeof this.getShortcutRegistry === 'function' && typeof this.matchesShortcutEvent === 'function';
const entry = registryAvailable ? this.getShortcutRegistry().find((s) => s.id === 'copy-selection') : null;
if (entry) return !entry.disabled && this.matchesShortcutEvent(ev, entry);
return !ev.altKey && (ev.key || '').toLowerCase() === 'c';
},
// Copy the current terminal selection. Goes through _copyText (Clipboard API,
// then a hidden-textarea + execCommand fallback) because install.sh's LAN
// option serves plain HTTP, where navigator.clipboard is undefined.
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const ok = await this._copyText(selection);
if (ok) {
// Clearing is what makes a second Ctrl+C an interrupt (and xterm already
// drops the selection on any keypress, so this matches existing feel).
this.terminal.clearSelection?.();
this.showToast('Copied to clipboard', 'success');
} else {
this.showToast('Failed to copy', 'error');
}
// The execCommand fallback focuses a temp textarea, so hand focus back. This
// is the CJK-aware focus router, not xterm's raw focus().
this.terminal.focus();
return ok;
},
async copyTerminal() {
try {
const buffer = this.terminal.buffer.active;
@@ -2751,9 +2938,29 @@ Object.assign(CodemanApp.prototype, {
// deltaY≈0 collapses to a fixed ±1 line/tick and the gesture can't page through
// history on a trackpad (issue #154). Non-Shift and mouse-wheel paths are
// unchanged (they carry deltaY). The `|| ±1` keeps sub-25px deltas moving.
//
// `deltaMode` says what UNIT the delta is in, and ignoring it made every
// non-pixel browser scroll ~4x too slowly: Firefox reports DOM_DELTA_LINE (1)
// with deltaY≈3 per notch, so the pixel math rounded to 0 and fell through to
// the ±1 fallback — one line per notch, versus 4-5 for Chrome's ~110px. In
// Claude mode the same value also capped the forwarded SGR report at one tick.
_wheelScrollLines(ev) {
const lines = this._wheelScrollLinesFloat(ev);
if (!lines) return 0; // pure horizontal swipe: don't fall through to -1
return Math.round(lines) || (lines > 0 ? 1 : -1);
},
/** Unrounded variant for the smooth local-scroll path, which accumulates
* sub-line fractions across events instead of forcing every tiny trackpad
* delta to a whole ±1 line. Same unit handling and Shift-axis trap. */
_wheelScrollLinesFloat(ev) {
const delta = ev.shiftKey && Math.abs(ev.deltaX) > Math.abs(ev.deltaY) ? ev.deltaX : ev.deltaY;
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
if (!delta) return 0;
return ev.deltaMode === 1 // DOM_DELTA_LINE (Firefox mouse wheel)
? delta
: ev.deltaMode === 2 // DOM_DELTA_PAGE
? delta * (this.terminal?.rows || 24)
: delta / 25; // DOM_DELTA_PIXEL (Chrome/WebKit, and every trackpad)
},
_shouldForwardWheelToApp(ev) {
@@ -2772,7 +2979,23 @@ Object.assign(CodemanApp.prototype, {
} else if (sessionMode !== 'codex') {
return false;
}
return this._terminalViewportAtBottom();
// Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so
// that leaving the bottom handed the wheel back to local scrollback and both
// histories stayed reachable without a mode switch. In practice that inverted
// the behavior users actually want: a repaint-mode CLI keeps NO terminal
// scrollback of its own (tmux reports history_size=0 for a Claude pane), so
// xterm's buffer holds only Codeman's REPLAYED repaint frames. Scrolling that
// locally drags the CLI's own pinned furniture (the prompt box, the status
// line) up the screen and shows stale frames underneath, which reads as "the
// window scrolled away" rather than "I am reading history".
//
// And it was easy to fall into: scrollToLastNonEmptyLine() parks the viewport
// `rows - 2` above the last non-empty row, so any tab switch onto a session
// with trailing blank rows left the viewport off-bottom and every later wheel
// went local. Forwarding unconditionally keeps the CLI's transcript as the
// plain wheel's target and its input box fixed in place; local scrollback is
// still on Shift+wheel and on the "Wheel scrolls local history" opt-out above.
return true;
},
// Encode wheel ticks as SGR reports (button 64 = up, 65 = down) at the pointer
+261 -5
View File
@@ -1,12 +1,16 @@
/**
* @fileoverview File browser and streaming routes.
* Provides directory listing, file content preview, raw file serving, and tail streaming.
* Provides directory listing, file content preview, raw file serving, tail
* streaming, and the File Viewer edit-mode write path (edit=1 read +
* PUT /api/sessions/:id/file-content; policy in src/config/file-editing.ts,
* design in docs/file-viewer-edit-plan.md).
*/
import { FastifyInstance, type FastifyReply } from 'fastify';
import { basename as pathBasename, extname, isAbsolute, join, relative, resolve, sep } from 'node:path';
import { basename as pathBasename, dirname, extname, isAbsolute, join, relative, resolve, sep } from 'node:path';
import { createReadStream, realpathSync, type ReadStream } from 'node:fs';
import fs from 'node:fs/promises';
import { createHash, randomBytes } from 'node:crypto';
import { homedir } from 'node:os';
import type {
ApiResponse,
@@ -14,6 +18,7 @@ import type {
FilesystemBrowseEntry,
FilesystemBrowseRoot,
FilesystemPreviewKind,
FileWriteData,
} from '../../types.js';
import { ApiErrorCode, createErrorResponse, getErrorMessage } from '../../types.js';
import { fileStreamManager } from '../../file-stream-manager.js';
@@ -44,7 +49,14 @@ import type { SessionAttachmentHistoryItem, SessionState } from '../../types/ses
import { isSensitivePath } from '../sensitive-path.js';
import { SseEvent } from '../sse-events.js';
import type { ConfigPort, EventPort, SessionPort } from '../ports/index.js';
import { FilesystemBrowseQuerySchema, FilesystemPreviewQuerySchema } from '../schemas.js';
import { FilesystemBrowseQuerySchema, FilesystemPreviewQuerySchema, FileWriteSchema } from '../schemas.js';
import {
MAX_EDITABLE_BYTES,
applyEol,
detectEol,
isDeniedEditRelativePath,
isEditableFileName,
} from '../../config/file-editing.js';
const MIME_TYPES: Record<string, string> = {
png: 'image/png',
@@ -453,6 +465,71 @@ function appendDownloadFlag(url: string): string {
return `${url}${url.includes('?') ? '&' : '?'}download=true`;
}
// ===== File Viewer edit mode (issue #212) =====
// Policy lives in src/config/file-editing.ts; design in docs/file-viewer-edit-plan.md.
function sha256Hex(buf: Buffer): string {
return createHash('sha256').update(buf).digest('hex');
}
/** NUL byte in the first 8KB — same binary signal the plain read path uses. */
function sniffsBinary(buf: Buffer): boolean {
const sniffLength = Math.min(buf.length, 8192);
for (let i = 0; i < sniffLength; i++) {
if (buf[i] === 0) return true;
}
return false;
}
/**
* Structured-throw variant for the edit read/write paths. Identical mechanics to
* throwFilesystemPickerError (rendered by the central route error handler both
* in prod and in the app.inject() test harness); a separate name only so edit
* failures grep distinctly.
*/
function throwFileEditError(statusCode: number, code: ApiErrorCode, message: string): never {
throw Object.assign(new Error(message), {
statusCode,
body: createErrorResponse(code, message),
});
}
/**
* Gate a resolved workspace file for edit-mode read/write. Throws a structured
* error when the file may not be edited; returns void when it may. Order
* matters for the message a user sees: confinement (the caller's 404) →
* sensitive/blocked (403) → .git (403) → extension allowlist (400).
*/
function assertEditableTarget(resolvedPath: string, relativePath: string, blockedTrees: readonly string[]): void {
if (isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, blockedTrees)) {
throwFileEditError(403, ApiErrorCode.FORBIDDEN, 'Editing this file is blocked');
}
if (isDeniedEditRelativePath(relativePath)) {
throwFileEditError(403, ApiErrorCode.FORBIDDEN, 'Files under .git cannot be edited');
}
if (!isEditableFileName(pathBasename(resolvedPath))) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'This file type is not editable');
}
}
/**
* Decode a candidate edit buffer, refusing binary and non-UTF-8 content. The
* round-trip compare is what protects against silent corruption: decoding
* latin-1 (or any non-UTF-8) bytes yields U+FFFD replacements, and writing
* those back would destroy the original bytes. A UTF-8 BOM round-trips and is
* deliberately preserved.
*/
function decodeEditableText(buf: Buffer): string {
if (sniffsBinary(buf)) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Binary files cannot be edited');
}
const text = buf.toString('utf8');
if (!Buffer.from(text, 'utf8').equals(buf)) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Only UTF-8 text files can be edited');
}
return text;
}
function getSessionAttachmentHistory(
ctx: SessionPort & ConfigPort,
sessionId: string,
@@ -864,7 +941,12 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
// Get file content for preview (File Browser)
app.get('/api/sessions/:id/file-content', async (req) => {
const { id } = req.params as { id: string };
const { path: filePath, lines, raw } = req.query as { path?: string; lines?: string; raw?: string };
const {
path: filePath,
lines,
raw,
edit,
} = req.query as { path?: string; lines?: string; raw?: string; edit?: string };
const session = findSessionOrFail(ctx, id, req);
if (!filePath) {
@@ -876,7 +958,52 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
if (!validated) {
return createErrorResponse(ApiErrorCode.NOT_FOUND, 'File not found');
}
const { resolvedPath } = validated;
const { resolvedPath, relativePath } = validated;
// Read-for-edit: never truncated (a truncated buffer must never become an
// edit buffer), tighter size cap, full editability gate, and the hash/eol
// the client must echo back on PUT. Outside the shared try/catch below so
// its structured errors keep their status codes instead of collapsing into
// OPERATION_FAILED.
if (edit === '1' || edit === 'true') {
const guard = await loadAttachmentGuardConfig();
assertEditableTarget(resolvedPath, relativePath, guard.blockedTrees);
let editStat;
try {
editStat = await fs.stat(resolvedPath);
} catch {
throwFileEditError(404, ApiErrorCode.NOT_FOUND, 'File not found');
}
if (!editStat.isFile()) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Only regular files can be edited');
}
if (editStat.size > MAX_EDITABLE_BYTES) {
throwFileEditError(
413,
ApiErrorCode.INVALID_INPUT,
`File too large to edit here (${Math.ceil(editStat.size / 1024)}KB > ${MAX_EDITABLE_BYTES / 1024}KB limit)`
);
}
const editBuf = await fs.readFile(resolvedPath);
const editText = decodeEditableText(editBuf);
return {
success: true,
data: {
path: filePath,
content: editText,
size: editBuf.length,
mtimeMs: editStat.mtimeMs,
totalLines: editText.split('\n').length,
truncated: false,
extension: filePath.split('.').pop()?.toLowerCase() || '',
editable: true,
hash: sha256Hex(editBuf),
eol: detectEol(editText),
},
};
}
try {
const stat = await fs.stat(resolvedPath);
@@ -998,6 +1125,19 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
const truncatedContent = allLines.length > maxLines;
const displayContent = truncatedContent ? allLines.slice(0, maxLines).join('\n') : content;
// Additive edit-mode advertisement: whether an edit=1 re-fetch would
// succeed. The UTF-8 round-trip compare is a cheap memcmp and mirrors
// decodeEditableText; no hash here — the Edit action re-fetches with
// edit=1, which is where the baseHash comes from.
const guard = await loadAttachmentGuardConfig();
const editable =
isEditableFileName(pathBasename(resolvedPath)) &&
!isDeniedEditRelativePath(relativePath) &&
!isSensitivePath(resolvedPath) &&
!isBlockedAttachmentPath(resolvedPath, guard.blockedTrees) &&
stat.size <= MAX_EDITABLE_BYTES &&
Buffer.from(content, 'utf8').equals(buf);
return {
success: true,
data: {
@@ -1007,6 +1147,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
totalLines: allLines.length,
truncated: truncatedContent,
extension: ext,
editable,
},
};
} catch (err) {
@@ -1014,6 +1155,121 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
}
});
// File Viewer edit mode: save a text file back into the session workspace.
// Edit-in-place ONLY — there is deliberately no O_CREAT path in this handler,
// so it can never create, and it never deletes. Confinement is identical to
// the read path (realpath + workspace boundary + ownership via
// findSessionOrFail), plus the sensitive-path/attachment-guard blocklists and
// the extension allowlist. Concurrency is optimistic: the client echoes the
// sha256 it loaded (baseHash) and a mismatch is a 409 unless force is set.
// bodyLimit: JSON escaping can expand content up to ~6x (each control char
// becomes \uXXXX), so the 512KB content cap needs headroom over Fastify's
// 1MB default.
app.put(
'/api/sessions/:id/file-content',
{ bodyLimit: 4 * 1024 * 1024 },
async (req): Promise<ApiResponse<FileWriteData>> => {
const { id } = req.params as { id: string };
const session = findSessionOrFail(ctx, id, req);
const body = parseBody(FileWriteSchema, req.body);
// Exact byte cap — the schema's .max() counts UTF-16 code units and is
// only a coarse pre-filter.
if (Buffer.byteLength(body.content, 'utf8') > MAX_EDITABLE_BYTES) {
throwFileEditError(413, ApiErrorCode.INVALID_INPUT, `Content too large (${MAX_EDITABLE_BYTES / 1024}KB limit)`);
}
const validated = validateSessionFilePath(session.workingDir, body.path);
if (!validated) {
// Covers missing files, traversal, and symlink escapes alike — a write
// target that fails confinement is reported identically to a missing
// one, matching the read route.
throwFileEditError(404, ApiErrorCode.NOT_FOUND, 'File not found');
}
const { resolvedPath, relativePath } = validated;
const guard = await loadAttachmentGuardConfig();
assertEditableTarget(resolvedPath, relativePath, guard.blockedTrees);
let stat;
try {
stat = await fs.stat(resolvedPath);
} catch {
throwFileEditError(404, ApiErrorCode.NOT_FOUND, 'File not found');
}
if (!stat.isFile()) {
throwFileEditError(400, ApiErrorCode.INVALID_INPUT, 'Only regular files can be edited');
}
if (stat.size > MAX_EDITABLE_BYTES) {
throwFileEditError(
413,
ApiErrorCode.INVALID_INPUT,
`File too large to edit here (${MAX_EDITABLE_BYTES / 1024}KB limit)`
);
}
const currentBuf = await fs.readFile(resolvedPath);
const currentText = decodeEditableText(currentBuf);
const currentHash = sha256Hex(currentBuf);
if (currentHash !== body.baseHash && !body.force) {
throwFileEditError(
409,
ApiErrorCode.CONFLICT,
'File changed on disk since it was loaded — reload it or overwrite'
);
}
// Re-apply the file's original line endings (a <textarea> normalizes to
// LF; without this a two-line edit of a CRLF file rewrites every line).
const eol = body.eol ?? detectEol(currentText);
const outText = applyEol(body.content, eol);
const outBuf = Buffer.from(outText, 'utf8');
if (outBuf.length > MAX_EDITABLE_BYTES) {
throwFileEditError(413, ApiErrorCode.INVALID_INPUT, `Content too large (${MAX_EDITABLE_BYTES / 1024}KB limit)`);
}
// Atomic replace: O_EXCL temp in the same directory, then rename.
// 'wx' cannot follow a pre-existing symlink and rename() replaces (not
// follows) a symlink in the final component, which closes the
// validate-then-write TOCTOU window. fchmod because open()'s mode is
// masked by the process umask; fsync so the rename never publishes a
// partially-durable file. Trade-off (same as vim's default): the inode
// changes, so hardlinks keep the old content.
const fileMode = stat.mode & 0o777;
const tmpPath = join(
dirname(resolvedPath),
`.${pathBasename(resolvedPath)}.codeman-tmp-${randomBytes(6).toString('hex')}`
);
let handle;
try {
handle = await fs.open(tmpPath, 'wx', fileMode);
await handle.chmod(fileMode);
await handle.writeFile(outBuf);
await handle.sync();
await handle.close();
handle = undefined;
await fs.rename(tmpPath, resolvedPath);
} catch (err) {
if (handle) await handle.close().catch(() => {});
await fs.unlink(tmpPath).catch(() => {});
throwFileEditError(500, ApiErrorCode.OPERATION_FAILED, `Failed to save file: ${getErrorMessage(err)}`);
}
const newStat = await fs.stat(resolvedPath).catch(() => undefined);
return {
success: true,
data: {
path: body.path,
size: outBuf.length,
mtimeMs: newStat?.mtimeMs ?? Date.now(),
hash: sha256Hex(outBuf),
totalLines: outText.split('\n').length,
eol,
},
};
}
);
// Serve raw file content (for images/binary files)
app.get('/api/sessions/:id/file-raw', async (req, reply) => {
const { id } = req.params as { id: string };
+228 -85
View File
@@ -21,7 +21,7 @@ import {
type GeminiConfig,
type AntigravityConfig,
} from '../../types.js';
import { Session, isAltScreenStripMode } from '../../session.js';
import { Session, isAltScreenStripMode, isMuxAltScreenOnlyStripMode } from '../../session.js';
import { SseEvent } from '../sse-events.js';
import {
CreateSessionSchema,
@@ -415,7 +415,7 @@ export function registerSessionRoutes(
//
// For keys the caller is actively setting, strip any stale disk entry a prior
// Codeman version may have written. Scope limited to:
// - Claude mode (OpenCode/Codex/Gemini don't read .claude/settings.local.json)
// - Claude mode (OpenCode/Codex/Gemini/Antigravity don't read .claude/settings.local.json)
// - workingDir inside CASES_DIR / the per-user case space (Codeman's managed
// territory — we never mutate .claude/settings.local.json in arbitrary user
// repos that POST /api/sessions can target, as those may have hand-authored
@@ -965,80 +965,97 @@ export function registerSessionRoutes(
// ========== Get Last Response (from transcript JSONL) ==========
// Resolves the most recent Claude conversation id for a session's cwd by
// tailing ~/.claude/history.jsonl. After `/clear`, Claude Code keeps writing
// to a new <uuid>.jsonl; history.jsonl is the only source-of-truth update
// that does not rely on project-local hooks (we intentionally don't install
// hooks in arbitrary user repos, see the POST /api/sessions comment).
// How far apart a ~/.claude/history.jsonl entry and a pane's Enter may be and
// still be the same submission. Claude appends to history as it accepts the
// prompt, so the true gap is milliseconds — this is slack for a loaded box,
// not a search radius.
const CLAUDE_SUBMIT_MATCH_MS = 10_000;
// history.jsonl grows forever; only the tail can hold entries near a submit.
const CLAUDE_HISTORY_TAIL_BYTES = 256 * 1024;
// Resolves the Claude conversation id THIS pane is currently on by matching
// ~/.claude/history.jsonl (which logs every submitted prompt as
// {project, sessionId, timestamp}) against the pane's last Enter. After
// `/clear` Claude keeps writing to a new <uuid>.jsonl, and history.jsonl is
// the only source-of-truth update that does not rely on project-local hooks
// (we intentionally don't install hooks in arbitrary user repos, see the
// POST /api/sessions comment).
//
// Entries from OTHER Codeman sessions in the same cwd are filtered out by
// their known claudeSessionIds so concurrent tabs don't shadow each other,
// as long as each has had its id resolved at least once.
// The pane's own Enter is what makes an entry OURS. `project` alone is not:
// a cwd is shared by every other Codeman tab on it, by tabs long since
// closed, and by any plain `claude` the user runs in their own terminal —
// adopting the newest entry for the cwd pinned the viewer to whichever of
// those conversations was typed in last, so the eye showed a stranger's
// transcript. With no correlated entry we keep the id we have; a viewer one
// turn behind beats a viewer showing someone else's conversation.
const claudeHistoryPinCache = new LRUMap<string, { submitAt: number; claudeSessionId: string }>({ maxSize: 1024 });
async function resolveActiveClaudeSessionIdFromHistory(
session: Session,
projectsDir: string
): Promise<string | null> {
const historyPath = join(homedir(), '.claude', 'history.jsonl');
const submitAt = session.lastSubmitAt;
if (!submitAt) return null; // never typed through Codeman — nothing to credit
const cached = claudeHistoryPinCache.get(session.id);
if (cached && cached.submitAt === submitAt) return cached.claudeSessionId;
// Ids another live pane is already pinned to can never be ours, and every
// pane sharing this cwd competes for the entry we are about to claim —
// including non-Claude panes, since a shell pane can run `claude` too.
const otherClaudeIds = new Set<string>();
const otherSubmits: number[] = [];
for (const s of ctx.sessions.values()) {
if (s.id !== session.id && s.workingDir === session.workingDir && s.claudeSessionId) {
otherClaudeIds.add(s.claudeSessionId);
}
if (s.id === session.id || s.workingDir !== session.workingDir) continue;
if (s.claudeSessionId) otherClaudeIds.add(s.claudeSessionId);
if (s.lastSubmitAt) otherSubmits.push(s.lastSubmitAt);
}
let candidateSid: string | null = null;
try {
const content = await fs.readFile(historyPath, 'utf8');
const lines = content.split('\n');
for (let i = lines.length - 1; i >= 0; i--) {
const line = lines[i];
if (!line) continue;
try {
const entry = JSON.parse(line) as { project?: string; sessionId?: string };
if (
entry.project === session.workingDir &&
typeof entry.sessionId === 'string' &&
!otherClaudeIds.has(entry.sessionId)
) {
candidateSid = entry.sessionId;
break;
}
} catch {
// Skip unparseable lines
}
}
} catch {
return null;
}
if (!candidateSid || candidateSid === session.id) return candidateSid;
const historyPath = join(homedir(), '.claude', 'history.jsonl');
const stat = await fs.stat(historyPath).catch(() => null);
if (!stat || stat.size === 0) return null;
const tail = await readFileTail(historyPath, Buffer.alloc(CLAUDE_HISTORY_TAIL_BYTES), stat.size);
if (!tail) return null;
// Safety: only adopt if the candidate's jsonl is more recently written
// than our initial conversation's jsonl. Blocks stale ids inherited from
// a prior Codeman session that happened to share this cwd.
try {
const projectDirs = await fs.readdir(projectsDir);
let best: { sessionId: string; dist: number } | undefined;
for (const line of tail.split('\n')) {
if (!line) continue;
let entry: { project?: string; sessionId?: string; timestamp?: number };
try {
entry = JSON.parse(line) as typeof entry;
} catch {
continue; // the first tail line is usually cut mid-JSON
}
const { sessionId, timestamp } = entry;
if (entry.project !== session.workingDir) continue;
if (typeof sessionId !== 'string' || !sessionId) continue;
if (typeof timestamp !== 'number') continue;
if (otherClaudeIds.has(sessionId)) continue;
const dist = Math.abs(timestamp - submitAt);
if (dist > CLAUDE_SUBMIT_MATCH_MS) continue;
if (otherSubmits.some((other) => Math.abs(timestamp - other) < dist)) continue; // another pane is closer
if (!best || dist <= best.dist) best = { sessionId, dist }; // ties: the newer entry wins
}
if (!best) return null;
// Sanity: the conversation we switch to must exist on disk and must not be
// staler than the one we are leaving. A `/clear` successor never is.
const currentSessionId = session.claudeSessionId || session.id;
if (best.sessionId !== currentSessionId) {
const projectDirs = await fs.readdir(projectsDir).catch(() => null);
if (!projectDirs) return null;
let candidateMtime = 0;
let initialMtime = 0;
let currentMtime = 0;
for (const projDir of projectDirs) {
try {
const cs = await fs.stat(join(projectsDir, projDir, `${candidateSid}.jsonl`));
if (cs.mtimeMs > candidateMtime) candidateMtime = cs.mtimeMs;
} catch {
/* not in this dir */
}
try {
const is = await fs.stat(join(projectsDir, projDir, `${session.id}.jsonl`));
if (is.mtimeMs > initialMtime) initialMtime = is.mtimeMs;
} catch {
/* not in this dir */
}
const candidateStat = await fs.stat(join(projectsDir, projDir, `${best.sessionId}.jsonl`)).catch(() => null);
if (candidateStat && candidateStat.mtimeMs > candidateMtime) candidateMtime = candidateStat.mtimeMs;
const currentStat = await fs.stat(join(projectsDir, projDir, `${currentSessionId}.jsonl`)).catch(() => null);
if (currentStat && currentStat.mtimeMs > currentMtime) currentMtime = currentStat.mtimeMs;
}
if (candidateMtime === 0) return null;
if (initialMtime > 0 && candidateMtime <= initialMtime) return null;
} catch {
return null;
if (candidateMtime === 0) return null; // transcript not written yet — retry next poll
if (currentMtime > 0 && candidateMtime < currentMtime) return null;
}
return candidateSid;
claudeHistoryPinCache.set(session.id, { submitAt, claudeSessionId: best.sessionId });
return best.sessionId;
}
interface ClaudeResponseMessage {
@@ -1236,6 +1253,11 @@ export function registerSessionRoutes(
const activeId = await resolveActiveClaudeSessionIdFromHistory(session, projectsDir);
if (activeId && activeId !== session.claudeSessionId) {
session.adoptClaudeSessionId(activeId);
// Flush the Enter that vouched for this adoption to state.json. A `/clear`
// emits no completion event, so without this the anchor could still be
// unpersisted when the server restarts — and recovery would fall back to
// the launch conversation.
ctx.persistSessionState(session);
// Docker sessions: keep the case's resume seed following the live conversation.
if (session.docker) {
void persistDockerCaseClaudeSessionId(CODEMAN_CONFIG_DIR, session.docker.containerName, activeId).catch(
@@ -1309,7 +1331,7 @@ export function registerSessionRoutes(
return { cwd, originator };
}
// The pane's last Enter (Session.codexLastSubmitAt) correlated against
// The pane's last Enter (Session.lastSubmitAt) correlated against
// ~/.codex/history.jsonl, which logs every submitted user message as
// {session_id, ts}. This identifies the thread the pane is ACTUALLY on and
// is the only signal that survives /resume, /new and /fork typed inside the
@@ -1318,10 +1340,10 @@ export function registerSessionRoutes(
// can't steal the attribution.
const codexHistoryPinCache = new LRUMap<string, { submitAt: number; threadId: string }>({ maxSize: 1024 });
async function resolveCodexThreadFromHistory(
session: { id: string; codexLastSubmitAt?: number },
session: { id: string; lastSubmitAt?: number },
codexHome: string
): Promise<string | null> {
const submitAt = session.codexLastSubmitAt || 0;
const submitAt = session.lastSubmitAt || 0;
if (!submitAt) return null;
const cached = codexHistoryPinCache.get(session.id);
if (cached && cached.submitAt === submitAt) return cached.threadId;
@@ -1335,8 +1357,8 @@ export function registerSessionRoutes(
const WINDOW_MS = 15_000;
const otherSubmits: number[] = [];
for (const s of ctx.sessions.values()) {
if (s.id !== session.id && s.mode === 'codex' && s.codexLastSubmitAt) {
otherSubmits.push(s.codexLastSubmitAt);
if (s.id !== session.id && s.mode === 'codex' && s.lastSubmitAt) {
otherSubmits.push(s.lastSubmitAt);
}
}
@@ -1379,7 +1401,7 @@ export function registerSessionRoutes(
async function findActiveCodexFile(session: {
id: string;
workingDir: string;
codexLastSubmitAt?: number;
lastSubmitAt?: number;
codexConfig?: { resumeSessionId?: string };
}): Promise<string | null> {
const codexHome = process.env.CODEX_HOME || join(process.env.HOME || '/tmp', '.codex');
@@ -1681,6 +1703,11 @@ export function registerSessionRoutes(
.replace(ALT_SCREEN_TOGGLE_PATTERN, '')
.replace(ERASE_SCROLLBACK_PATTERN, '')
.replace(MOUSE_TRACKING_PATTERN, '');
} else if (isMuxAltScreenOnlyStripMode(session.mode, session.usesMux)) {
// tmux-backed shell/opencode/antigravity: drop tmux's own client smcup only.
// A byte buffer recorded before the live-side strip existed can still carry
// it, and one replayed `\x1b[?1049h` re-parks xterm in the alt buffer (#205).
strippedBuffer = strippedBuffer.replace(ALT_SCREEN_TOGGLE_PATTERN, '');
}
if (tailBytes > 0 && strippedBuffer.length > tailBytes) {
@@ -2005,7 +2032,7 @@ export function registerSessionRoutes(
if (!host) return createErrorResponse(ApiErrorCode.NOT_FOUND, 'Remote host not found');
// Per-session config that is applied to the LOCAL tmux/CLI wrapper (env vars via
// tmux setenv, effort/model CLI args, codex/gemini/opencode config) does NOT
// tmux setenv, effort/model CLI args, codex/gemini/antigravity/opencode config) does NOT
// cross ssh, so it would silently no-op. Reject rather than pretend it worked —
// remote command/env customization goes through the per-host command override.
if (
@@ -2342,7 +2369,7 @@ export function registerSessionRoutes(
});
ctx.broadcast(SseEvent.SessionInteractive, { id: session.id, mode: 'shell' });
} else {
// 'claude', 'opencode', 'codex', and 'gemini' modes use startInteractive()
// 'claude', 'opencode', 'codex', 'gemini', and 'antigravity' modes use startInteractive()
await session.startInteractive();
getLifecycleLog().log({
event: 'started',
@@ -2456,6 +2483,64 @@ export function registerSessionRoutes(
return undefined;
}
/**
* Is this `entrypoint` value an automated/SDK-driven invocation?
*
* ⚠️ Deliberately a BLOCKLIST on the SDK shape, not an allowlist on `'cli'`.
* The exclusion below hides rows, so an allowlist fails CLOSED on any value
* Claude Code has not shipped yet: the day it stamps a new interactive
* entrypoint (a rename, or a second interactive host), every transcript stops
* matching `'cli'` and the whole Past Sessions list goes blank with nothing in
* the UI to explain it. A blocklist fails OPEN instead — an automated
* entrypoint we do not recognize yet costs a few noisy rows, which is the
* annoyance this filter set out to fix rather than a broken feature.
*
* Observed values: `cli` (interactive), `sdk-cli` / `sdk-py` (automated).
*/
function isAutomatedEntrypoint(entrypoint: string): boolean {
return /^sdk(-|$)/.test(entrypoint);
}
/**
* The `entrypoint` field Claude Code stamps on its own message records:
* 'cli' for a real interactive session, something else (e.g. 'sdk-py') for
* an SDK/automated invocation. Used to exclude non-interactive transcripts
* (CI review bots, etc.) from the resumable history list — they were never
* something a user can resume into.
*
* Scans every `"type":"user"`/`"type":"assistant"` line with an entrypoint
* field — not just the first one — and returns 'cli' the moment ANY of them
* carries it. A transcript is excluded only when every entrypoint-bearing
* message says something else; "first field wins" would misattribute a
* transcript that started under an older Claude Code version (no entrypoint
* on its true first message) and later picked up a non-'cli' entrypoint on
* some later message, wrongly hiding a genuinely interactive session. This
* deliberately errs toward keeping a session visible: one real interactive
* message anywhere is enough. Returns undefined ("unknown", fail-open) only
* when nothing scanned carries the field at all.
*/
function extractTranscriptEntrypoint(text: string): string | undefined {
let start = 0;
let sawNonCli: string | undefined;
while (start < text.length) {
const end = text.indexOf('\n', start);
const line = end === -1 ? text.slice(start) : text.slice(start, end);
start = end === -1 ? text.length : end + 1;
if (!line.includes('"type":"user"') && !line.includes('"type":"assistant"')) continue;
if (!line.includes('"entrypoint"')) continue;
try {
const entry = JSON.parse(line);
if ((entry.type === 'user' || entry.type === 'assistant') && typeof entry.entrypoint === 'string') {
if (entry.entrypoint === 'cli') return 'cli';
sawNonCli ??= entry.entrypoint;
}
} catch {
// Malformed/truncated line — skip
}
}
return sawNonCli;
}
/**
* Extract the text of the LAST user message from a JSONL transcript chunk
* (COD-145). Mirrors `extractFirstUserPrompt` exactly — same user-message
@@ -2668,7 +2753,7 @@ export function registerSessionRoutes(
return finalExists ? current : process.env.HOME || '/tmp';
}
/** Read the first 16KB of a file for content sniffing. */
/** Read the first `buf.length` bytes of a file for content sniffing. */
async function readFileHead(path: string, buf: Buffer): Promise<string | null> {
try {
const fd = await fs.open(path, 'r');
@@ -2711,7 +2796,12 @@ export function registerSessionRoutes(
// Scan a single project directory and return all valid history sessions in it.
// Reused by both the global overview and the single-folder drill-down.
async function scanProjectDir(projPath: string, projDir: string, headBuf: Buffer): Promise<HistorySession[]> {
async function scanProjectDir(
projPath: string,
projDir: string,
smallHeadBuf: Buffer,
headBuf: Buffer
): Promise<HistorySession[]> {
const out: HistorySession[] = [];
const stat = await fs.stat(projPath).catch(() => null);
if (!stat?.isDirectory()) return out;
@@ -2729,22 +2819,45 @@ export function registerSessionRoutes(
if (!fileStat) continue;
if (fileStat.size < 4000) continue;
let firstPrompt: string | undefined;
const head = await readFileHead(filePath, headBuf);
const hasConversation = (text: string) =>
text.includes('"type":"user"') || text.includes('"type":"assistant"') || text.includes('"type":"summary"');
// Two-tier head read: try the cheap smallHeadBuf (16KB) size first -- enough
// for the vast majority of transcripts -- and only escalate to the full
// headBuf (128KB) when that wasn't enough. Reading 128KB unconditionally for
// EVERY file in the directory roughly quadrupled the cost of a full scan
// (measured against a real ~/.claude/projects tree: ~4x both bytes read and
// wall time) to fix a problem only ~28% of files actually have. Escalating
// resolves the restart-bookkeeping case (the reason 128KB exists at all)
// without ever touching the tail-read fallback below for most of that 28%.
let head = await readFileHead(filePath, smallHeadBuf);
let foundContent = head ? hasConversation(head) : false;
let firstPrompt = head ? extractFirstUserPrompt(head) : undefined;
if ((!foundContent || !firstPrompt) && head !== null && fileStat.size > smallHeadBuf.length) {
const biggerHead = await readFileHead(filePath, headBuf);
if (biggerHead) {
head = biggerHead;
if (!foundContent) foundContent = hasConversation(head);
if (!firstPrompt) firstPrompt = extractFirstUserPrompt(head);
}
}
let tail: string | null = null;
if (!foundContent && fileStat.size > 16384) {
// `head === null` (a failed read -- e.g. EMFILE while scanning hundreds of
// files) must also get a shot at the tail, not just "file bigger than the
// head buffer". Losing this dropped the session from history entirely
// instead of giving it a second chance, for any file at or under the head
// buffer size whose head read happened to fail.
if (!foundContent && (head === null || fileStat.size > headBuf.length)) {
const tailBuf = Buffer.alloc(32768);
tail = await readFileTail(filePath, tailBuf, fileStat.size);
if (tail) foundContent = hasConversation(tail);
}
if (!foundContent) continue;
if (head) firstPrompt = extractFirstUserPrompt(head);
if (!firstPrompt && fileStat.size > 65536) {
// firstPrompt was already attempted from head (both tiers) above; this is
// purely the tail fallback for whatever's left unresolved.
if (!firstPrompt && (head === null || fileStat.size > headBuf.length)) {
if (!tail) {
const tailBuf = Buffer.alloc(32768);
tail = await readFileTail(filePath, tailBuf, fileStat.size);
@@ -2754,15 +2867,37 @@ export function registerSessionRoutes(
// COD-145: last (most recent) user prompt lives near the END of the file, so
// prefer the tail. For large files where no tail was read yet, read one
// (mirrors the firstPrompt > 65536 block). Small files fit in `head`, which
// then contains the whole transcript — scan it for the last match instead.
if (!tail && fileStat.size > 65536) {
// (mirrors the firstPrompt > headBuf.length block). Small files fit in `head`,
// which then contains the whole transcript — scan it for the last match instead.
if (!tail && fileStat.size > headBuf.length) {
const tailBuf = Buffer.alloc(32768);
tail = await readFileTail(filePath, tailBuf, fileStat.size);
}
const lastPrompt =
(tail ? extractLastUserPrompt(tail) : undefined) ?? (head ? extractLastUserPrompt(head) : undefined);
// Automated/SDK-driven invocations (CI review bots, etc.) write transcripts
// into the same ~/.claude/projects tree as interactive sessions but were
// never something a user can resume into — no PTY, no running process, and
// their "conversation" is typically a single one-shot prompt (often with a
// full diff embedded, which is exactly why it dwarfs this scanner's read
// windows and shows up above as blank or as an identical boilerplate
// sentence across many rows). Checked last, so it reuses whatever `head`/
// `tail` the prompt extraction above already read rather than triggering
// an extra file read. Missing entrypoint (older transcripts) reads as
// interactive — fail open, matching every other gating check in this
// codebase.
//
// head and tail are checked independently and merged with "cli wins" (not
// a first-truthy-value `??` chain): a large file's head might land on a
// non-'cli' message while a real interactive message sits in the tail (or
// vice versa), and either one being 'cli' is enough to keep the session.
const headEntrypoint = head ? extractTranscriptEntrypoint(head) : undefined;
const tailEntrypoint = tail ? extractTranscriptEntrypoint(tail) : undefined;
const entrypoint =
headEntrypoint === 'cli' || tailEntrypoint === 'cli' ? 'cli' : (headEntrypoint ?? tailEntrypoint);
if (entrypoint && isAutomatedEntrypoint(entrypoint)) continue;
out.push({
sessionId,
workingDir,
@@ -2779,7 +2914,13 @@ export function registerSessionRoutes(
app.get('/api/history/sessions', async (req) => {
const query = req.query as { projectKey?: string; offset?: string; limit?: string };
const projectsDir = join(process.env.HOME || '/tmp', '.claude', 'projects');
const headBuf = Buffer.alloc(16384);
// scanProjectDir tries smallHeadBuf (16KB, the original size) first for every
// file and only escalates to headBuf (128KB) when that wasn't enough — see the
// comment at the escalation site in scanProjectDir for why unconditional 128KB
// reads were too expensive to keep. 128KB matches the existing precedent
// elsewhere in this file (line ~1431).
const smallHeadBuf = Buffer.alloc(16384);
const headBuf = Buffer.alloc(131072);
// Multi-user: this scans the host-wide ~/.claude/projects tree, so a non-admin
// must only see history whose decoded workingDir is inside their own case space.
// Do NOT trust the caller-supplied projectKey — confine on the decoded path.
@@ -2797,7 +2938,7 @@ export function registerSessionRoutes(
const offset = Math.max(0, parseInt(query.offset || '0', 10) || 0);
const limit = Math.min(100, Math.max(1, parseInt(query.limit || '20', 10) || 20));
const projPath = join(projectsDir, query.projectKey);
let all = await scanProjectDir(projPath, query.projectKey, headBuf);
let all = await scanProjectDir(projPath, query.projectKey, smallHeadBuf, headBuf);
// Confine to the caller's workspace (a projectKey maps to a single foreign cwd).
if (scopeHistory) all = all.filter((r) => isWorkingDirAllowed(user, r.workingDir));
all.sort((a, b) => new Date(b.lastModified).getTime() - new Date(a.lastModified).getTime());
@@ -2810,7 +2951,7 @@ export function registerSessionRoutes(
const projectDirs = await fs.readdir(projectsDir);
for (const projDir of projectDirs) {
const projPath = join(projectsDir, projDir);
const list = await scanProjectDir(projPath, projDir, headBuf);
const list = await scanProjectDir(projPath, projDir, smallHeadBuf, headBuf);
results.push(...list);
}
} catch {
@@ -2886,11 +3027,13 @@ export function registerSessionRoutes(
const history: HistoryInput[] = [];
try {
const projectsDir = join(process.env.HOME || '/tmp', '.claude', 'projects');
const headBuf = Buffer.alloc(16384);
// See the sibling allocation above for why there are two sizes.
const smallHeadBuf = Buffer.alloc(16384);
const headBuf = Buffer.alloc(131072);
const projectDirs = await fs.readdir(projectsDir);
for (const projDir of projectDirs) {
const projPath = join(projectsDir, projDir);
const list = await scanProjectDir(projPath, projDir, headBuf);
const list = await scanProjectDir(projPath, projDir, smallHeadBuf, headBuf);
for (const h of list) {
history.push({
sessionId: h.sessionId,
+26 -1
View File
@@ -16,6 +16,7 @@ import {
MIN_TERMINAL_BUFFER_BYTES,
MIN_TERMINAL_SCROLLBACK_LINES,
} from '../config/terminal-history.js';
import { MAX_EDITABLE_BYTES } from '../config/file-editing.js';
// ========== Path Validation ==========
@@ -83,6 +84,30 @@ export const FilesystemPreviewQuerySchema = z.object({
.optional(),
});
/**
* Body validation for `PUT /api/sessions/:id/file-content` (File Viewer edit
* mode). `content.max()` counts UTF-16 code units, which for UTF-8 output is
* always <= the byte length, so it is a coarse pre-filter that never rejects
* valid content; the handler enforces the exact MAX_EDITABLE_BYTES byte cap.
* Workspace containment and symlink resolution are enforced by the route via
* validateSessionFilePath after parsing.
*/
export const FileWriteSchema = z
.object({
path: z
.string()
.min(1)
.max(4096)
.refine((p) => !p.includes('\0') && !p.includes('\n') && !p.includes('\r'), {
message: 'Invalid path',
}),
content: z.string().max(MAX_EDITABLE_BYTES),
baseHash: z.string().regex(/^[a-f0-9]{64}$/, 'baseHash must be a sha256 hex digest'),
eol: z.enum(['lf', 'crlf']).optional(),
force: z.boolean().optional(),
})
.strict();
// ========== Env Var Allowlist ==========
/** Allowlisted env var key prefixes */
@@ -491,7 +516,7 @@ export const DockerHostSchema = z.object({
mountCredentials: z.boolean().optional(),
hooksEnabled: z.boolean().optional(),
resumeOnStart: z.boolean().optional(),
commands: RemoteCommandOverridesSchema, // same shell/claude/opencode/codex/gemini shape
commands: RemoteCommandOverridesSchema, // same shell/claude/opencode/codex/gemini/antigravity shape
extraCreateArgs: z
.array(
z
+4
View File
@@ -2508,6 +2508,10 @@ export class WebServer extends EventEmitter {
envOverrides: savedEnvOverrides,
effort: savedState?.effort,
attachmentHistory: savedAttachmentHistory,
// The pane's last Enter. Without it the response viewer would show
// the launch conversation until the user types again, even though
// the re-attached CLI is on a post-`/clear` one.
lastSubmitAt: savedState?.lastSubmitAt,
// Remote SSH metadata must round-trip on recovery: without it the
// attach cwd falls back to the (nonexistent-locally) remote path and
// respawn rebuilds a LOCAL command, breaking the pane and silently
+58 -2
View File
@@ -1,5 +1,5 @@
import { describe, expect, it } from 'vitest';
import { Session, isAltScreenStripMode } from '../src/session.js';
import { Session, isAltScreenStripMode, isMuxAltScreenOnlyStripMode } from '../src/session.js';
type SessionInternals = {
_handleTerminalOutput(data: string): void;
@@ -81,7 +81,7 @@ describe('Claude terminal scrollback strip', () => {
});
});
describe('Shell terminal output is NOT stripped (vim/less/htop need the alt screen)', () => {
describe('Shell terminal output on a DIRECT PTY is NOT stripped (vim/less/htop need the alt screen)', () => {
it('leaves alt-screen toggles, scrollback-erase, and mouse-tracking intact for shell', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell' });
@@ -91,3 +91,59 @@ describe('Shell terminal output is NOT stripped (vim/less/htop need the alt scre
expect(session.terminalBuffer).toBe(vimLike);
});
});
describe('isMuxAltScreenOnlyStripMode', () => {
it('covers exactly the modes the full strip does not, and only under tmux', () => {
for (const mode of ['shell', 'opencode', 'antigravity'] as const) {
expect(isMuxAltScreenOnlyStripMode(mode, true)).toBe(true);
// Direct-PTY fallback: the program's own alt screen really does reach xterm.
expect(isMuxAltScreenOnlyStripMode(mode, false)).toBe(false);
}
// The full strip already owns these; never double-gate them here.
for (const mode of ['claude', 'codex', 'gemini'] as const) {
expect(isMuxAltScreenOnlyStripMode(mode, true)).toBe(false);
}
});
});
describe('tmux-backed shell: strip tmux’s own client smcup, keep everything else (#205)', () => {
it('drops alt-screen toggles so xterm keeps a scrollback buffer', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
// What a real `tmux attach` emits as its first bytes.
handleOutput(session, '\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2Jprompt$ ');
expect(session.terminalBuffer).toBe('\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2Jprompt$ ');
expect(session.terminalBuffer).not.toContain('\x1b[?1049h');
});
it('KEEPS 3J and mouse-tracking, unlike the full strip', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
// `clear` legitimately wipes scrollback; htop/vim mouse modes are passed
// through by tmux even with `mouse off` and must keep working.
handleOutput(session, '\x1b[3J\x1b[?1002h\x1b[?1006hhtop\x1b[?1006l\x1b[?1002l');
expect(session.terminalBuffer).toBe('\x1b[3J\x1b[?1002h\x1b[?1006hhtop\x1b[?1006l\x1b[?1002l');
});
it('reassembles alt-screen sequences split across PTY chunk boundaries', () => {
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
const emitted: string[] = [];
session.on('terminal', (data) => emitted.push(data));
handleOutput(session, 'before\x1b[?104');
handleOutput(session, '9h after');
expect(session.terminalBuffer).toBe('before after');
expect(emitted).toEqual(['before', ' after']);
});
it('applies to opencode and antigravity too', () => {
for (const mode of ['opencode', 'antigravity'] as const) {
const session = new Session({ workingDir: '/tmp', mode, useMux: true });
handleOutput(session, '\x1b[?1049hTUI\x1b[3J');
expect(session.terminalBuffer).toBe('TUI\x1b[3J');
}
});
});
+98
View File
@@ -0,0 +1,98 @@
/**
* @fileoverview Unit tests for the File Viewer edit-mode policy module.
*
* Pure functions only — no IO, no server.
* Port: N/A (no server)
*/
import { describe, it, expect } from 'vitest';
import {
MAX_EDITABLE_BYTES,
applyEol,
detectEol,
isDeniedEditRelativePath,
isEditableFileName,
} from '../src/config/file-editing.js';
describe('file-editing policy', () => {
describe('isEditableFileName', () => {
it('allows common text extensions', () => {
for (const name of ['a.ts', 'b.md', 'c.json', 'd.py', 'style.css', 'notes.txt', 'x.yml', 'Q.SQL']) {
expect(isEditableFileName(name), name).toBe(true);
}
});
it('allows well-known basenames regardless of case', () => {
for (const name of ['Dockerfile', 'Makefile', 'LICENSE', '.gitignore', '.editorconfig', '.nvmrc']) {
expect(isEditableFileName(name), name).toBe(true);
}
});
it('rejects binary/media/document extensions', () => {
for (const name of ['a.png', 'b.pdf', 'c.docx', 'd.zip', 'e.woff2', 'f.mp4', 'g.exe']) {
expect(isEditableFileName(name), name).toBe(false);
}
});
it('rejects svg and env (deliberate v1 exclusions)', () => {
expect(isEditableFileName('image.svg')).toBe(false);
expect(isEditableFileName('config.env')).toBe(false);
});
it('rejects extensionless and unknown-dotfile names not on the basename list', () => {
expect(isEditableFileName('somebinary')).toBe(false);
expect(isEditableFileName('.bashrc')).toBe(false);
expect(isEditableFileName('archive.xyz')).toBe(false);
});
});
describe('isDeniedEditRelativePath', () => {
it('denies anything inside a .git directory at any depth', () => {
expect(isDeniedEditRelativePath('.git/config')).toBe(true);
expect(isDeniedEditRelativePath('.git/hooks/pre-commit')).toBe(true);
expect(isDeniedEditRelativePath('sub/module/.git/HEAD')).toBe(true);
});
it('allows non-.git paths, including names merely containing "git"', () => {
expect(isDeniedEditRelativePath('src/index.ts')).toBe(false);
expect(isDeniedEditRelativePath('.github/workflows/ci.yml')).toBe(false);
expect(isDeniedEditRelativePath('digits/file.md')).toBe(false);
expect(isDeniedEditRelativePath('.gitignore')).toBe(false);
});
});
describe('detectEol / applyEol', () => {
it('detects LF, CRLF, and defaults to LF for single-line text', () => {
expect(detectEol('a\nb\nc')).toBe('lf');
expect(detectEol('a\r\nb\r\nc')).toBe('crlf');
expect(detectEol('no newline at all')).toBe('lf');
expect(detectEol('')).toBe('lf');
});
it('picks the dominant style for mixed-EOL text', () => {
expect(detectEol('a\r\nb\r\nc\nd')).toBe('crlf');
expect(detectEol('a\nb\nc\r\nd')).toBe('lf');
});
it('applyEol round-trips a textarea-normalized (LF) buffer back to CRLF', () => {
const original = 'line1\r\nline2\r\nline3';
const textareaValue = original.replace(/\r\n/g, '\n');
expect(applyEol(textareaValue, detectEol(original))).toBe(original);
});
it('applyEol is idempotent and never doubles CR', () => {
expect(applyEol('a\r\nb', 'crlf')).toBe('a\r\nb');
expect(applyEol('a\r\nb', 'lf')).toBe('a\nb');
expect(applyEol('a\nb', 'lf')).toBe('a\nb');
});
it('preserves a UTF-8 BOM through the EOL rewrite', () => {
const withBom = 'hello\nworld';
expect(applyEol(withBom, 'crlf')).toBe('hello\r\nworld');
});
});
it('exposes a sane editable-bytes cap', () => {
expect(MAX_EDITABLE_BYTES).toBe(512 * 1024);
});
});
+9
View File
@@ -20,6 +20,15 @@ describe('frontend public asset tooling', () => {
expect(appJs.includes(0)).toBe(false);
});
it('uses the same message wrapper for brief and full response views', () => {
const appJs = readFileSync(resolve(repoRoot, 'src/web/public/app.js'), 'utf8');
expect(appJs).toContain("body.appendChild(this._buildResponseViewerMessage(lastResponse, 'assistant'");
expect(appJs).toContain('body.appendChild(this._buildResponseViewerMessage(msg.text, msg.role, agentLabel));');
expect(appJs).toContain("div.className = 'rv-message ' + (isUser ? 'rv-msg-user' : 'rv-msg-assistant');");
expect(appJs).toContain("renderedText.className = 'rv-text';");
});
it('runs the public asset check script', () => {
expect(() => {
execFileSync('npm', ['run', 'check:public-assets', '--silent'], {
+2
View File
@@ -52,6 +52,8 @@ describe('help modal shortcuts', () => {
});
it('documents terminal input shortcuts without advertising stale run shortcuts', () => {
expectShortcut(helpModal, ['Ctrl', 'C'], 'Copy Selection');
expectShortcut(helpModal, ['Ctrl', 'Shift', 'C'], 'Copy Selection');
expectShortcut(helpModal, ['Ctrl', 'L'], 'Clear Terminal');
expectShortcut(helpModal, ['Ctrl', '+'], 'Increase Font');
expectShortcut(helpModal, ['Ctrl', '-'], 'Decrease Font');
+18
View File
@@ -58,4 +58,22 @@ describe('keyboard shortcuts', () => {
expect(appSource).toContain('if (this.matchesShortcutEvent(e, shortcut))');
expect(appSource).toContain('if (shortcut.disabled || !shortcut.action) continue;');
});
it('keeps the interrupt when Ctrl+C copies a selection (#211)', () => {
// The xterm handler owns this decision, and the no-selection path must fall
// through with NO preventDefault so xterm still evaluates Ctrl+C into 0x03.
expect(terminalUiSource).toContain('this.shouldCopyTerminalSelectionFromShortcut?.(ev)');
expect(terminalUiSource).toMatch(
/const selection = this\.terminal\.hasSelection\?\.\(\) \? this\.terminal\.getSelection\(\) : '';/
);
expect(terminalUiSource).toContain('void this.copyTerminalSelection(selection);');
expect(appSource).toContain("id: 'copy-selection'");
});
it('documents the terminal copy shortcut in help and README', () => {
expect(helpHtml).toContain('<kbd>Ctrl</kbd>+<kbd>C</kbd>');
expect(helpHtml).toContain('<kbd>Ctrl</kbd>+<kbd>Shift</kbd>+<kbd>C</kbd>');
expect(readme).toContain('`Ctrl/Cmd+C`');
expect(readme).toContain('`Ctrl+Shift+C`');
});
});
+66 -1
View File
@@ -12,13 +12,34 @@ import { describe, expect, it } from 'vitest';
const PUBLIC = resolve(import.meta.dirname, '../src/web/public');
/** Minimal fake DOM node — enough surface for mobile-overview.js's programmatic builders. */
function fakeElement(): any {
const el: any = {
className: '',
type: '',
dataset: {},
style: {},
children: [] as any[],
setAttribute() {},
appendChild(child: any) {
el.children.push(child);
return child;
},
};
return el;
}
function loadOverviewApp(overrides: Record<string, any> = {}) {
const CodemanApp = function CodemanApp(this: any) {};
const context = vm.createContext({
CodemanApp,
console,
window: {},
document: { getElementById: () => null },
document: {
getElementById: () => null,
createElement: () => fakeElement(),
createElementNS: () => fakeElement(),
},
MobileDetection: { getDeviceType: () => 'mobile' },
});
vm.runInContext(readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8'), context, {
@@ -272,3 +293,47 @@ describe('mobile overview wiring', () => {
expect(hardcoded).toEqual([]);
});
});
describe('mobile overview run picker (CLI availability gating)', () => {
function modeButtons(menu: any): string[] {
return menu.children.filter((c: any) => c.dataset.moAction === 'run-mode').map((c: any) => c.dataset.moMode);
}
// #201 gated the toolbar's #runModeMenu on isCliAvailable(); this phone-only
// picker (MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu) is a
// separate, hardcoded duplicate of that menu rather than a shared render, so
// it silently offered every backend regardless of what the server reported.
it('hides run modes the server reports as unavailable, keeps shell always', () => {
const app = loadOverviewApp({
runMode: 'claude',
isCliAvailable: (tool: string) => tool === 'claude',
});
const menu = app._buildMobileOverviewRunMenu();
expect(modeButtons(menu)).toEqual(['claude', 'shell']);
});
it('shows every mode when every CLI is available', () => {
const app = loadOverviewApp({
runMode: 'claude',
isCliAvailable: () => true,
});
const menu = app._buildMobileOverviewRunMenu();
expect(modeButtons(menu)).toEqual(['claude', 'opencode', 'codex', 'gemini', 'antigravity', 'shell']);
});
it('gates every mode the picker actually offers', () => {
// Catches a new backend being added to MOBILE_OVERVIEW_RUN_MODES without
// being gated — the same class of bug that let this list drift from the
// toolbar menu's gating in the first place.
const src = readFileSync(resolve(PUBLIC, 'mobile-overview.js'), 'utf8');
const modesBlock = src.slice(
src.indexOf('const MOBILE_OVERVIEW_RUN_MODES'),
src.indexOf('];', src.indexOf('const MOBILE_OVERVIEW_RUN_MODES')) + 2
);
const offered = [...modesBlock.matchAll(/mode: '([^']+)'/g)].map((m) => m[1]);
expect(offered).toContain('antigravity');
const fn = src.slice(src.indexOf('_buildMobileOverviewRunMenu() {'));
const gate = fn.slice(0, fn.indexOf('const header'));
expect(gate).toContain('isCliAvailable');
});
});
+2
View File
@@ -19,6 +19,8 @@ export class MockSession extends EventEmitter {
ralphTracker: null = null;
writeBuffer: string[] = [];
terminalBuffer: string = '';
/** Mirrors Session.lastSubmitAt — the response viewer credits history entries by it. */
lastSubmitAt: number = 0;
private _muxName: string | null = null;
+32 -1
View File
@@ -20,7 +20,12 @@
import { homedir } from 'node:os';
import { describe, it, expect } from 'vitest';
import { buildSshConnectionArgs, buildRemoteTmuxCheckCommand, remoteSshTarget } from '../src/remote-hosts.js';
import {
buildSshConnectionArgs,
buildRemoteTmuxCheckCommand,
buildRemoteCliVersionProbeCommand,
remoteSshTarget,
} from '../src/remote-hosts.js';
import { buildRemoteLaunchCommand } from '../src/tmux-manager.js';
import type { SessionRemote } from '../src/types.js';
@@ -198,3 +203,29 @@ describe('COD-107 buildRemoteTmuxCheckCommand — same connection options as the
expect(buildRemoteTmuxCheckCommand({ username: 'ubuntu', host: '10.0.0.42', port: 2222 })).toContain('-p 2222');
});
});
describe('buildRemoteCliVersionProbeCommand: remote CLI version over the same connection (#205)', () => {
it('routes the version query through the interactive-login shell wrapper, like the launch', () => {
const cmd = buildRemoteCliVersionProbeCommand(baseRemote, 'claude');
// Same PATH-resolution wrapper as defaultRemoteCommandForMode: a bare
// `claude --version` over ssh sees only sshd's minimal PATH (exit 127).
expect(cmd).toBe(
'ssh -o BatchMode=yes -o ConnectTimeout=10 ubuntu@10.0.0.42 ' +
`'exec "\${SHELL:-/bin/sh}" -i -l -c '\\''claude --version'\\'''`
);
});
it('uses the shared connection args (proxy/identity/port), so it reaches what the launch reaches', () => {
const cmd = buildRemoteCliVersionProbeCommand(aaDesktop, 'claude');
expect(cmd).toContain('-o BatchMode=yes');
expect(cmd).toContain('-p 2222');
expect(cmd).toContain(`-i '${HOME}/.ssh/remote_ed25519'`);
expect(cmd).toContain("-o 'ProxyCommand=nc -X 5 -x 127.0.0.1:1080 %h %p'");
expect(cmd).toContain('aakht@192.168.55.170');
});
it('maps antigravity to its real binary name and shell to no probe at all', () => {
expect(buildRemoteCliVersionProbeCommand(baseRemote, 'antigravity')).toContain('agy --version');
expect(buildRemoteCliVersionProbeCommand(baseRemote, 'shell')).toBeNull();
});
});
+354
View File
@@ -0,0 +1,354 @@
/**
* @fileoverview File Viewer edit mode — read-for-edit (`edit=1`) and
* `PUT /api/sessions/:id/file-content` (issue #212).
*
* Deliberately does NOT mock node:fs — every case runs against a real temp
* workspace so the confinement (realpath + workspace boundary), the symlink
* behavior, the atomic temp+rename write, and mode preservation are exercised
* for real, not against a mock's assumptions.
*
* Uses app.inject() — no real HTTP ports needed.
* Port: N/A (app.inject doesn't open ports)
*/
import { describe, it, expect, beforeEach, afterEach } from 'vitest';
import {
mkdtempSync,
mkdirSync,
writeFileSync,
readFileSync,
symlinkSync,
chmodSync,
statSync,
realpathSync,
readdirSync,
rmSync,
existsSync,
} from 'node:fs';
import { createHash } from 'node:crypto';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { createRouteTestHarness, type RouteTestHarness } from './_route-test-utils.js';
import { registerFileRoutes } from '../../src/web/routes/file-routes.js';
import { MAX_EDITABLE_BYTES } from '../../src/config/file-editing.js';
function sha256(data: string | Buffer): string {
return createHash('sha256').update(data).digest('hex');
}
describe('file viewer edit mode (real fs)', () => {
let harness: RouteTestHarness;
let workDir: string;
let outsideDir: string;
const sessionId = 'test-session-1';
const putFile = (path: string, body: Record<string, unknown>) =>
harness.app.inject({
method: 'PUT',
url: `/api/sessions/${sessionId}/file-content`,
payload: { path, ...body },
});
const getEdit = (path: string) =>
harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=${encodeURIComponent(path)}&edit=1`,
});
beforeEach(async () => {
harness = await createRouteTestHarness(registerFileRoutes);
// realpath: on some hosts tmpdir() contains a symlinked component, which
// would make validateSessionFilePath's relative() check misfire.
workDir = realpathSync(mkdtempSync(join(tmpdir(), 'codeman-edit-ws-')));
outsideDir = realpathSync(mkdtempSync(join(tmpdir(), 'codeman-edit-out-')));
harness.ctx._session.workingDir = workDir;
});
afterEach(async () => {
await harness.app.close();
rmSync(workDir, { recursive: true, force: true });
rmSync(outsideDir, { recursive: true, force: true });
});
// ========== GET ?edit=1 ==========
describe('GET /api/sessions/:id/file-content?edit=1', () => {
it('returns the FULL content (never truncated) with hash and eol', async () => {
const content = Array.from({ length: 800 }, (_, i) => `line ${i + 1}`).join('\n');
writeFileSync(join(workDir, 'long.md'), content);
const res = await getEdit('long.md');
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.success).toBe(true);
expect(body.data.content).toBe(content);
expect(body.data.truncated).toBe(false);
expect(body.data.totalLines).toBe(800);
expect(body.data.editable).toBe(true);
expect(body.data.hash).toBe(sha256(content));
expect(body.data.eol).toBe('lf');
});
it('reports crlf for a CRLF file', async () => {
writeFileSync(join(workDir, 'dos.txt'), 'a\r\nb\r\nc');
const res = await getEdit('dos.txt');
expect(res.json().data.eol).toBe('crlf');
});
it('413s above MAX_EDITABLE_BYTES instead of truncating', async () => {
writeFileSync(join(workDir, 'big.log'), 'x'.repeat(MAX_EDITABLE_BYTES + 1));
const res = await getEdit('big.log');
expect(res.statusCode).toBe(413);
expect(res.json().success).toBe(false);
});
it('400s for a non-allowlisted extension', async () => {
writeFileSync(join(workDir, 'data.xyz'), 'text');
const res = await getEdit('data.xyz');
expect(res.statusCode).toBe(400);
});
it('400s for binary content even with a text extension', async () => {
writeFileSync(join(workDir, 'fake.txt'), Buffer.from([0x68, 0x00, 0x69]));
const res = await getEdit('fake.txt');
expect(res.statusCode).toBe(400);
});
});
describe('plain read editable flag', () => {
it('advertises editable:true for an editable text file', async () => {
writeFileSync(join(workDir, 'notes.md'), 'hello');
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=notes.md`,
});
expect(res.json().data.editable).toBe(true);
});
it('advertises editable:false for a non-allowlisted extension', async () => {
writeFileSync(join(workDir, 'schema.xsd'), '<xml/>');
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${sessionId}/file-content?path=schema.xsd`,
});
const data = res.json().data;
expect(data.content).toBeDefined();
expect(data.editable).toBe(false);
});
});
// ========== PUT ==========
describe('PUT /api/sessions/:id/file-content', () => {
it('happy path: writes the bytes, returns new hash, leaves no temp files', async () => {
const original = 'line one\nline two\n';
writeFileSync(join(workDir, 'notes.md'), original);
const updated = 'line one EDITED\nline two\n';
const res = await putFile('notes.md', { content: updated, baseHash: sha256(original) });
expect(res.statusCode).toBe(200);
const body = res.json();
expect(body.success).toBe(true);
expect(body.data.hash).toBe(sha256(updated));
expect(body.data.eol).toBe('lf');
expect(body.data.size).toBe(Buffer.byteLength(updated));
expect(readFileSync(join(workDir, 'notes.md'), 'utf8')).toBe(updated);
const leftovers = readdirSync(workDir).filter((n) => n.includes('codeman-tmp'));
expect(leftovers).toEqual([]);
});
it('404s on ../ traversal without touching the outside file', async () => {
const target = join(outsideDir, 'victim.md');
writeFileSync(target, 'safe');
// Build a relative path that resolves outside the workspace.
const traversal = `..${target.startsWith('/') ? target : `/${target}`}`;
const res = await putFile(traversal, { content: 'pwned', baseHash: sha256('safe') });
expect(res.statusCode).toBe(404);
expect(readFileSync(target, 'utf8')).toBe('safe');
});
it('404s on an absolute path outside the workspace', async () => {
const target = join(outsideDir, 'victim2.md');
writeFileSync(target, 'safe');
const res = await putFile(target, { content: 'pwned', baseHash: sha256('safe') });
expect(res.statusCode).toBe(404);
expect(readFileSync(target, 'utf8')).toBe('safe');
});
it('404s a symlink pointing outside the workspace and never follows it', async () => {
const target = join(outsideDir, 'secret.md');
writeFileSync(target, 'outside');
symlinkSync(target, join(workDir, 'sneaky.md'));
const res = await putFile('sneaky.md', { content: 'pwned', baseHash: sha256('outside') });
expect(res.statusCode).toBe(404);
expect(readFileSync(target, 'utf8')).toBe('outside');
});
it('writes THROUGH a symlink whose target is inside the workspace', async () => {
writeFileSync(join(workDir, 'real.md'), 'original');
symlinkSync(join(workDir, 'real.md'), join(workDir, 'alias.md'));
const res = await putFile('alias.md', { content: 'via alias', baseHash: sha256('original') });
expect(res.statusCode).toBe(200);
expect(readFileSync(join(workDir, 'real.md'), 'utf8')).toBe('via alias');
});
it('400s a non-allowlisted extension', async () => {
writeFileSync(join(workDir, 'blob.xyz'), 'text');
const res = await putFile('blob.xyz', { content: 'nope', baseHash: sha256('text') });
expect(res.statusCode).toBe(400);
expect(readFileSync(join(workDir, 'blob.xyz'), 'utf8')).toBe('text');
});
it('403s inside .git even for an allowlisted-looking name', async () => {
mkdirSync(join(workDir, '.git'));
writeFileSync(join(workDir, '.git', 'config.ini'), '[core]');
const res = await putFile('.git/config.ini', { content: 'x', baseHash: sha256('[core]') });
expect(res.statusCode).toBe(403);
});
it('rejects a .env file (allowlist first, sensitive-path as backstop)', async () => {
writeFileSync(join(workDir, '.env'), 'SECRET=1');
const res = await putFile('.env', { content: 'SECRET=2', baseHash: sha256('SECRET=1') });
expect([400, 403]).toContain(res.statusCode);
expect(readFileSync(join(workDir, '.env'), 'utf8')).toBe('SECRET=1');
});
it('400s when the current file contains a NUL byte', async () => {
writeFileSync(join(workDir, 'weird.txt'), Buffer.from([0x61, 0x00, 0x62]));
const res = await putFile('weird.txt', { content: 'ab', baseHash: sha256(Buffer.from([0x61, 0x00, 0x62])) });
expect(res.statusCode).toBe(400);
});
it('400s when the current file is not valid UTF-8 (latin-1)', async () => {
const latin1 = Buffer.from('caf\xe9 au lait', 'latin1');
writeFileSync(join(workDir, 'legacy.txt'), latin1);
const res = await putFile('legacy.txt', { content: 'cafe au lait', baseHash: sha256(latin1) });
expect(res.statusCode).toBe(400);
expect(readFileSync(join(workDir, 'legacy.txt'))).toEqual(latin1);
});
it('409s on a stale baseHash and succeeds with force:true', async () => {
writeFileSync(join(workDir, 'contested.md'), 'agent version');
const res = await putFile('contested.md', { content: 'my version', baseHash: sha256('older version') });
expect(res.statusCode).toBe(409);
expect(res.json().errorCode).toBe('CONFLICT');
expect(readFileSync(join(workDir, 'contested.md'), 'utf8')).toBe('agent version');
const forced = await putFile('contested.md', {
content: 'my version',
baseHash: sha256('older version'),
force: true,
});
expect(forced.statusCode).toBe(200);
expect(readFileSync(join(workDir, 'contested.md'), 'utf8')).toBe('my version');
});
it('rejects oversized ASCII content at the schema pre-filter (400)', async () => {
writeFileSync(join(workDir, 'small.md'), 'ok');
const res = await putFile('small.md', {
content: 'x'.repeat(MAX_EDITABLE_BYTES + 1),
baseHash: sha256('ok'),
});
expect(res.statusCode).toBe(400);
expect(readFileSync(join(workDir, 'small.md'), 'utf8')).toBe('ok');
});
it('413s multibyte content that passes the code-unit pre-filter but exceeds the byte cap', async () => {
writeFileSync(join(workDir, 'small.md'), 'ok');
// '€' is 1 UTF-16 code unit but 3 UTF-8 bytes: 200k units (< 512Ki cap)
// becomes ~586KB on disk, so only the handler's byteLength check catches it.
const res = await putFile('small.md', {
content: '€'.repeat(200_000),
baseHash: sha256('ok'),
});
expect(res.statusCode).toBe(413);
expect(readFileSync(join(workDir, 'small.md'), 'utf8')).toBe('ok');
});
it('404s a missing file and creates nothing (edit-in-place only)', async () => {
const res = await putFile('brand-new.md', { content: 'hello', baseHash: sha256('hello') });
expect(res.statusCode).toBe(404);
expect(existsSync(join(workDir, 'brand-new.md'))).toBe(false);
});
it('400s a malformed baseHash at the schema layer', async () => {
writeFileSync(join(workDir, 'a.md'), 'x');
const res = await putFile('a.md', { content: 'y', baseHash: 'not-a-hash' });
expect(res.statusCode).toBe(400);
expect(res.json().errorCode).toBe('INVALID_INPUT');
});
it('preserves CRLF line endings across a textarea-normalized save', async () => {
const original = 'first\r\nsecond\r\nthird';
writeFileSync(join(workDir, 'dos.txt'), original);
// Client sends LF-normalized content + the eol it was told at load time.
const res = await putFile('dos.txt', {
content: 'first\nsecond EDITED\nthird',
baseHash: sha256(original),
eol: 'crlf',
});
expect(res.statusCode).toBe(200);
expect(readFileSync(join(workDir, 'dos.txt'), 'utf8')).toBe('first\r\nsecond EDITED\r\nthird');
});
it('re-applies the original EOL even when the client omits eol', async () => {
const original = 'a\r\nb';
writeFileSync(join(workDir, 'implicit.txt'), original);
const res = await putFile('implicit.txt', { content: 'a\nb\nc', baseHash: sha256(original) });
expect(res.statusCode).toBe(200);
expect(readFileSync(join(workDir, 'implicit.txt'), 'utf8')).toBe('a\r\nb\r\nc');
});
it('preserves the file mode across the temp+rename', async () => {
const p = join(workDir, 'script.sh');
writeFileSync(p, '#!/bin/sh\necho hi\n');
chmodSync(p, 0o750);
const res = await putFile('script.sh', {
content: '#!/bin/sh\necho bye\n',
baseHash: sha256('#!/bin/sh\necho hi\n'),
});
expect(res.statusCode).toBe(200);
expect(statSync(p).mode & 0o777).toBe(0o750);
});
it('rejects unknown body keys (.strict() schema)', async () => {
writeFileSync(join(workDir, 'a.md'), 'x');
const res = await putFile('a.md', { content: 'y', baseHash: sha256('x'), evil: true });
expect(res.statusCode).toBe(400);
});
});
// ========== multi-user scoping ==========
describe('multi-user ownership', () => {
it("404s a non-admin writing to a session they don't own", async () => {
const prev = process.env.CODEMAN_MULTIUSER;
process.env.CODEMAN_MULTIUSER = '1';
try {
const scoped = await createRouteTestHarness(registerFileRoutes, {
authUser: { username: 'mallory', role: 'user' },
});
scoped.ctx._session.workingDir = workDir;
writeFileSync(join(workDir, 'owned.md'), 'admin file');
const res = await scoped.app.inject({
method: 'PUT',
url: `/api/sessions/${sessionId}/file-content`,
payload: { path: 'owned.md', content: 'stolen', baseHash: sha256('admin file') },
});
expect(res.statusCode).toBe(404);
expect(readFileSync(join(workDir, 'owned.md'), 'utf8')).toBe('admin file');
await scoped.app.close();
} finally {
if (prev === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = prev;
}
});
});
});
@@ -8,7 +8,7 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import Fastify, { type FastifyInstance } from 'fastify';
import fastifyCookie from '@fastify/cookie';
import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs';
import { mkdirSync, mkdtempSync, rmSync, utimesSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
@@ -176,3 +176,131 @@ describe('GET /api/sessions/:id/last-response (claude)', () => {
]);
});
});
/**
* Which conversation a pane is on is decided by the pane's own Enter, not by
* "newest entry for this cwd" — a cwd is shared with every other tab on it,
* with tabs long since closed, and with any plain `claude` the user runs in
* their own terminal.
*/
describe('GET /api/sessions/:id/last-response (claude conversation pinning)', () => {
let harness: LocalHarness;
let testHome: string;
let previousHome: string | undefined;
const WORKDIR = '/workspace';
const NOW = 1_770_000_000_000;
beforeEach(async () => {
testHome = mkdtempSync(join(tmpdir(), 'codeman-claude-pin-'));
previousHome = process.env.HOME;
process.env.HOME = testHome;
harness = await createEnvelopeHarness();
});
afterEach(async () => {
if (previousHome === undefined) delete process.env.HOME;
else process.env.HOME = previousHome;
rmSync(testHome, { recursive: true, force: true });
await harness.app.close();
});
/** A transcript whose only assistant turn is `text`, stamped at `mtimeMs`. */
function writeTranscript(conversationId: string, text: string, mtimeMs: number): void {
const projectDir = join(testHome, '.claude', 'projects', '-workspace');
mkdirSync(projectDir, { recursive: true });
const path = join(projectDir, `${conversationId}.jsonl`);
writeFileSync(
path,
JSON.stringify({
type: 'assistant',
timestamp: new Date(mtimeMs).toISOString(),
message: { content: [{ type: 'text', text }] },
})
);
utimesSync(path, mtimeMs / 1000, mtimeMs / 1000);
}
function writeHistory(entries: Array<{ sessionId: string; timestamp: number; project?: string }>): void {
const claudeDir = join(testHome, '.claude');
mkdirSync(claudeDir, { recursive: true });
writeFileSync(
join(claudeDir, 'history.jsonl'),
entries
.map((entry) => JSON.stringify({ display: 'prompt', project: entry.project ?? WORKDIR, ...entry }))
.join('\n')
);
}
/** Replaces the pre-seeded mock session with a Claude pane in WORKDIR. */
function addPane(id: string, conversationId: string, lastSubmitAt: number) {
const base = harness.ctx._session;
const pane = Object.create(Object.getPrototypeOf(base)) as typeof base & {
claudeSessionId: string;
lastSubmitAt: number;
adoptClaudeSessionId: ReturnType<typeof vi.fn>;
};
Object.assign(pane, base, { id, mode: 'claude', workingDir: WORKDIR, docker: undefined });
pane.claudeSessionId = conversationId;
pane.lastSubmitAt = lastSubmitAt;
pane.adoptClaudeSessionId = vi.fn((newId: string) => {
pane.claudeSessionId = newId;
});
harness.ctx.sessions.set(id, pane);
return pane;
}
async function getLastResponse(sessionId: string) {
const response = await harness.app.inject({ method: 'GET', url: `/api/sessions/${sessionId}/last-response` });
return JSON.parse(response.body).data as { text: string };
}
it('does not adopt a conversation from another claude process sharing the cwd', async () => {
// The pane typed hours ago; a `claude` running in the user's own terminal
// is the newest thing in this cwd. Before this fix the viewer followed it.
const pane = addPane('pane-1', 'pane-conversation', NOW - 6 * 3600_000);
writeTranscript('pane-conversation', 'my own answer', NOW - 6 * 3600_000);
writeTranscript('someone-elses-conversation', 'a stranger answer', NOW);
writeHistory([{ sessionId: 'someone-elses-conversation', timestamp: NOW }]);
expect(await getLastResponse('pane-1')).toEqual({ text: 'my own answer', timestamp: expect.any(String) });
expect(pane.adoptClaudeSessionId).not.toHaveBeenCalled();
});
it('follows /clear onto the new conversation the pane submitted into', async () => {
const pane = addPane('pane-1', 'before-clear', NOW);
writeTranscript('before-clear', 'answer before clear', NOW - 60_000);
writeTranscript('after-clear', 'answer after clear', NOW + 500);
writeHistory([{ sessionId: 'after-clear', timestamp: NOW + 120 }]);
expect(await getLastResponse('pane-1')).toEqual({ text: 'answer after clear', timestamp: expect.any(String) });
expect(pane.adoptClaudeSessionId).toHaveBeenCalledWith('after-clear');
});
it('stays put when the pane has never submitted through Codeman', async () => {
const pane = addPane('pane-1', 'pane-conversation', 0);
writeTranscript('pane-conversation', 'my own answer', NOW - 60_000);
writeTranscript('unrelated-conversation', 'a stranger answer', NOW);
writeHistory([{ sessionId: 'unrelated-conversation', timestamp: NOW }]);
expect(await getLastResponse('pane-1')).toEqual({ text: 'my own answer', timestamp: expect.any(String) });
expect(pane.adoptClaudeSessionId).not.toHaveBeenCalled();
});
it('credits a shared-cwd entry to the pane whose Enter is closest to it', async () => {
const near = addPane('pane-near', 'near-conversation', NOW);
const far = addPane('pane-far', 'far-conversation', NOW - 4_000);
writeTranscript('near-conversation', 'near answer', NOW - 60_000);
writeTranscript('far-conversation', 'far answer', NOW - 60_000);
writeTranscript('fresh-conversation', 'the freshly cleared answer', NOW + 500);
writeHistory([{ sessionId: 'fresh-conversation', timestamp: NOW + 100 }]);
// Both panes are inside the match window; only the closest may claim it.
expect(await getLastResponse('pane-far')).toEqual({ text: 'far answer', timestamp: expect.any(String) });
expect(far.adoptClaudeSessionId).not.toHaveBeenCalled();
expect(await getLastResponse('pane-near')).toEqual({
text: 'the freshly cleared answer',
timestamp: expect.any(String),
});
expect(near.adoptClaudeSessionId).toHaveBeenCalledWith('fresh-conversation');
});
});
@@ -104,7 +104,7 @@ describe('GET /api/sessions/:id/last-response (codex)', () => {
let codexHome: string;
let prevCodexHome: string | undefined;
// eslint-disable-next-line @typescript-eslint/no-explicit-any
let session: any; // MockSession, loosened for codex-only fields (codexConfig, codexLastSubmitAt)
let session: any; // MockSession, loosened for codex-only fields (codexConfig, lastSubmitAt)
let workdir: string;
/** Write a rollout under CODEX_HOME/sessions/<date>/ with a controlled mtime. */
@@ -230,7 +230,7 @@ describe('GET /api/sessions/:id/last-response (codex)', () => {
it('history.jsonl pin (pane last-submit correlation) outranks the originator match', async () => {
const submitAtSec = BASE_MTIME + 500;
session.codexLastSubmitAt = submitAtSec * 1000;
session.lastSubmitAt = submitAtSec * 1000;
writeHistory([{ session_id: UUID_B, ts: submitAtSec }]);
// Originator-stamped rollout exists and is NEWER, but the pane /resume'd onto
+247
View File
@@ -1350,6 +1350,253 @@ describe('session-routes', () => {
expect(row.workingDir).toBe(dotDir);
expect(row.workingDir).not.toContain('//');
});
it('excludes non-interactive (SDK-driven) transcripts from the history list', async () => {
// CI review bots and other automated tools write transcripts into the same
// ~/.claude/projects tree as interactive sessions (entrypoint "sdk-py" etc.)
// but were never something a user can resume into — no PTY, no running
// process. They cluttered Past Sessions as blank rows or identical
// boilerplate ("Review this change for security vulnerabilities...").
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-entrypoint-test');
await mkdir(projPath, { recursive: true });
const cliId = '33333333-3333-3333-3333-333333333333';
const sdkId = '44444444-4444-4444-4444-444444444444';
const noEntrypointId = '55555555-5555-5555-5555-555555555555';
const cliLine =
JSON.stringify({ type: 'user', entrypoint: 'cli', message: { role: 'user', content: 'a real question' } }) +
'\n';
const sdkLine =
JSON.stringify({
type: 'user',
entrypoint: 'sdk-py',
message: { role: 'user', content: 'Review this change for security vulnerabilities.' },
}) + '\n';
// Older transcripts predate the entrypoint field entirely — must still show.
const noEntrypointLine =
JSON.stringify({ type: 'user', message: { role: 'user', content: 'a pre-entrypoint session' } }) + '\n';
await writeFile(join(projPath, `${cliId}.jsonl`), cliLine + '#'.repeat(4200 - cliLine.length));
await writeFile(join(projPath, `${sdkId}.jsonl`), sdkLine + '#'.repeat(4200 - sdkLine.length));
await writeFile(
join(projPath, `${noEntrypointId}.jsonl`),
noEntrypointLine + '#'.repeat(4200 - noEntrypointLine.length)
);
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-entrypoint-test',
});
expect(res.statusCode).toBe(200);
const ids = JSON.parse(res.body).data.sessions.map((s: { sessionId: string }) => s.sessionId);
expect(ids).toContain(cliId);
expect(ids).toContain(noEntrypointId);
expect(ids).not.toContain(sdkId);
});
it('ignores a bookkeeping line that happens to mention "entrypoint" outside a real message record', async () => {
// Scanning must anchor on "type":"user"/"assistant" lines specifically,
// not any line that happens to contain the substring "entrypoint".
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-entrypoint-bookkeeping-test');
await mkdir(projPath, { recursive: true });
const sessionId = '77777777-7777-7777-7777-777777777777';
const bookkeepingLine = JSON.stringify({ type: 'mode', mode: 'normal', entrypoint: 'sdk-py' }) + '\n';
const realLine =
JSON.stringify({ type: 'user', entrypoint: 'cli', message: { role: 'user', content: 'a real message' } }) +
'\n';
// scanProjectDir skips files under 4000 bytes.
const body = bookkeepingLine + realLine;
await writeFile(join(projPath, `${sessionId}.jsonl`), body + '#'.repeat(4200 - body.length));
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-entrypoint-bookkeeping-test',
});
expect(res.statusCode).toBe(200);
const ids = JSON.parse(res.body).data.sessions.map((s: { sessionId: string }) => s.sessionId);
expect(ids).toContain(sessionId);
});
it('shows a session with ANY interactive (cli) message, even if an earlier message was automated', async () => {
// "First field wins" would have misattributed this: an old transcript
// whose true first message predates the entrypoint field, later resumed
// under something automated (entrypoint: 'sdk-py' on message 2), then
// continued interactively by a real person (entrypoint: 'cli' on message
// 3). Stopping at the first entrypoint-bearing line found ('sdk-py')
// would wrongly exclude a session a human genuinely used. One real
// interactive message anywhere is enough to keep it visible.
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-entrypoint-any-cli-test');
await mkdir(projPath, { recursive: true });
const sessionId = '99999999-9999-9999-9999-999999999999';
const firstLine =
JSON.stringify({ type: 'user', message: { role: 'user', content: 'pre-entrypoint-field message' } }) + '\n';
const automatedLine =
JSON.stringify({
type: 'user',
entrypoint: 'sdk-py',
message: { role: 'user', content: 'an automated follow-up' },
}) + '\n';
const interactiveLine =
JSON.stringify({
type: 'user',
entrypoint: 'cli',
message: { role: 'user', content: 'a real person continued this' },
}) + '\n';
const body = firstLine + automatedLine + interactiveLine;
await writeFile(join(projPath, `${sessionId}.jsonl`), body + '#'.repeat(4200 - body.length));
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-entrypoint-any-cli-test',
});
expect(res.statusCode).toBe(200);
const ids = JSON.parse(res.body).data.sessions.map((s: { sessionId: string }) => s.sessionId);
expect(ids).toContain(sessionId);
});
it('still excludes a session where every entrypoint-bearing message is automated', async () => {
// Mirror of the previous test with no 'cli' message anywhere — proves the
// "any cli wins" fix isn't just failing open unconditionally.
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-entrypoint-all-automated-test');
await mkdir(projPath, { recursive: true });
const sessionId = 'aaaaaaaa-aaaa-aaaa-aaaa-aaaaaaaaaaaa';
const firstLine =
JSON.stringify({
type: 'user',
entrypoint: 'sdk-py',
message: { role: 'user', content: 'Review this change for security vulnerabilities.' },
}) + '\n';
const secondLine =
JSON.stringify({
type: 'assistant',
entrypoint: 'sdk-py',
message: { role: 'assistant', content: [{ type: 'text', text: 'Looking at the diff...' }] },
}) + '\n';
const body = firstLine + secondLine;
await writeFile(join(projPath, `${sessionId}.jsonl`), body + '#'.repeat(4200 - body.length));
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-entrypoint-all-automated-test',
});
expect(res.statusCode).toBe(200);
const ids = JSON.parse(res.body).data.sessions.map((s: { sessionId: string }) => s.sessionId);
expect(ids).not.toContain(sessionId);
});
it('keeps a session whose entrypoint is an unrecognized non-SDK value (fail open)', async () => {
// The exclusion is a blocklist on the SDK shape, NOT an allowlist on 'cli'.
// An allowlist fails CLOSED on any value Claude Code has not shipped yet:
// the day it stamps a new interactive entrypoint, nothing matches 'cli' and
// the entire Past Sessions list silently goes blank. Excluding only what we
// positively recognize as automated fails open instead — a few noisy rows,
// not a dead feature.
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-entrypoint-unknown-test');
await mkdir(projPath, { recursive: true });
const sessionId = 'bbbbbbbb-bbbb-bbbb-bbbb-bbbbbbbbbbbb';
const line =
JSON.stringify({
type: 'user',
entrypoint: 'cli-next',
message: { role: 'user', content: 'a question from a future interactive host' },
}) + '\n';
await writeFile(join(projPath, `${sessionId}.jsonl`), line + '#'.repeat(4200 - line.length));
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-entrypoint-unknown-test',
});
expect(res.statusCode).toBe(200);
const ids = JSON.parse(res.body).data.sessions.map((s: { sessionId: string }) => s.sessionId);
expect(ids).toContain(sessionId);
});
it('finds the real first prompt past a large run of pre-message bookkeeping lines', async () => {
// A session restarted many times over a long conversation accumulates a batch
// of small bookkeeping lines (mode/permission-mode/last-prompt/queue-operation)
// per restart, ahead of the real first message. With enough restarts these can
// push the genuine first prompt past a 16KB head-read window even though the
// message itself is tiny — the row showed up blank despite having real content.
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-bookkeeping-test');
await mkdir(projPath, { recursive: true });
const sessionId = '66666666-6666-6666-6666-666666666666';
const bookkeepingLine = JSON.stringify({ type: 'mode', mode: 'normal', sessionId }) + '\n';
// > 16KB (the old head-read size) but well under 128KB (the new one).
const prefix = bookkeepingLine.repeat(Math.ceil(20000 / bookkeepingLine.length));
const realLine =
JSON.stringify({
type: 'user',
entrypoint: 'cli',
message: { role: 'user', content: 'the real first message' },
}) + '\n';
expect(prefix.length).toBeGreaterThan(16384);
await writeFile(join(projPath, `${sessionId}.jsonl`), prefix + realLine);
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-bookkeeping-test',
});
expect(res.statusCode).toBe(200);
const row = JSON.parse(res.body).data.sessions.find((s: { sessionId: string }) => s.sessionId === sessionId);
expect(row).toBeDefined();
expect(row.firstPrompt).toBe('the real first message');
});
it('still falls back to the tail read when bookkeeping alone exceeds the new 128KB head window', async () => {
// Raising the head buffer to 128KB helps most restart-heavy sessions, but an
// even more extreme case (many more restarts) can still exceed it. This
// proves the tail-read fallback itself is intact after the threshold
// rewrite (`fileStat.size > headBuf.length` replacing the old hardcoded
// 16384/65536) — the fallback's own logic, not the exact threshold value,
// is what could have silently broken (e.g. a copy-paste slip that dropped
// the `> headBuf.length` check entirely). The real message sits near the
// end of the file, well inside the 32KB tail window, so a working fallback
// finds it; a broken one leaves the row blank exactly like the bug this
// whole fix addresses.
const home = process.env.HOME as string;
const projPath = join(home, '.claude', 'projects', 'proj-tail-fallback-test');
await mkdir(projPath, { recursive: true });
const sessionId = '88888888-8888-8888-8888-888888888888';
const bookkeepingLine = JSON.stringify({ type: 'mode', mode: 'normal', sessionId }) + '\n';
// Comfortably past the new 128KB head window (was 16KB), so the head read
// never reaches a single "type":"user"/"assistant"/"summary" line.
const prefix = bookkeepingLine.repeat(Math.ceil(140000 / bookkeepingLine.length));
const realLine =
JSON.stringify({
type: 'user',
entrypoint: 'cli',
message: { role: 'user', content: 'found via tail fallback' },
}) + '\n';
expect(prefix.length).toBeGreaterThan(131072);
await writeFile(join(projPath, `${sessionId}.jsonl`), prefix + realLine);
const res = await harness.app.inject({
method: 'GET',
url: '/api/history/sessions?projectKey=proj-tail-fallback-test',
});
expect(res.statusCode).toBe(200);
const row = JSON.parse(res.body).data.sessions.find((s: { sessionId: string }) => s.sessionId === sessionId);
expect(row).toBeDefined();
expect(row.firstPrompt).toBe('found via tail fallback');
});
});
// ========== POST /api/sessions (with resumeSessionId) ==========
+14 -1
View File
@@ -351,7 +351,13 @@ describe('Codex quick start settings', () => {
function loadUi(flags: Record<string, boolean> | undefined) {
const CodemanApp = function CodemanApp(this: any) {};
const welcomeBtns: Record<string, { style: { display: string } }> = {};
for (const id of ['welcomeClaudeBtn', 'welcomeOpencodeBtn', 'welcomeGeminiBtn', 'welcomeTunnelBtn']) {
for (const id of [
'welcomeClaudeBtn',
'welcomeOpencodeBtn',
'welcomeAntigravityBtn',
'welcomeGeminiBtn',
'welcomeTunnelBtn',
]) {
welcomeBtns[id] = { style: { display: 'PRISTINE' } };
}
const modeBtns: Record<string, { style: { display: string } }> = {};
@@ -394,6 +400,7 @@ describe('Codex quick start settings', () => {
app.applyWelcomeCliVisibility();
expect(welcomeBtns.welcomeClaudeBtn.style.display).toBe('flex');
expect(welcomeBtns.welcomeOpencodeBtn.style.display).toBe('none');
expect(welcomeBtns.welcomeAntigravityBtn.style.display).toBe('none');
expect(welcomeBtns.welcomeGeminiBtn.style.display).toBe('none');
// #200 originally DELETED the tunnel button and its QR outright; it is gated
// on cloudflared instead, so a box that has cloudflared keeps the feature.
@@ -402,6 +409,12 @@ describe('Codex quick start settings', () => {
const withTunnel = loadUi({ ...ALL_OFF, cloudflared: true });
withTunnel.app.applyWelcomeCliVisibility();
expect(withTunnel.welcomeBtns.welcomeTunnelBtn.style.display).toBe('flex');
// Antigravity is a first-class welcome action, gated on `agy` like the rest.
const withAgy = loadUi({ ...ALL_OFF, antigravity: true });
withAgy.app.applyWelcomeCliVisibility();
expect(withAgy.welcomeBtns.welcomeAntigravityBtn.style.display).toBe('flex');
expect(withAgy.welcomeBtns.welcomeClaudeBtn.style.display).toBe('none');
});
it('gates every run mode in the dropdown, antigravity included, and never shell', () => {
@@ -270,6 +270,75 @@ describe('mergeUnifiedSessions', () => {
expect(live!.firstPrompt).toBeUndefined();
});
it('does NOT borrow a sibling transcript for a history-only row whose own extraction failed (no cross-contamination)', () => {
// A pure history row already got its own real scan (step 1 keys it under its
// OWN sessionId) — if that extraction genuinely failed (oversized first
// message, noise-filtered, etc.), the workingDir guess must not paper over
// it with an unrelated session's opening line. Regression: an old session
// in a shared workingDir was displaying TODAY's live session's firstPrompt
// as its own, because the guess didn't check whether this row already had
// its own (failed) attempt.
const merged = mergeUnifiedSessions({
history: [
// This session's own transcript scan found no usable prompt.
{
sessionId: 'old-uuid',
workingDir: '/shared',
sizeBytes: 5000,
lastModified: '2026-01-01T00:00:00.000Z',
firstPrompt: undefined,
},
// A much newer, unrelated session in the same directory.
{
sessionId: 'newer-uuid',
workingDir: '/shared',
sizeBytes: 6000,
lastModified: '2026-06-01T00:00:00.000Z',
firstPrompt: "today's real prompt",
},
],
});
const old = merged.find((m) => m.sessionId === 'old-uuid');
expect(old).toBeDefined();
expect(old!.firstPrompt).toBeUndefined();
});
it('leaves a RESUMED session blank rather than borrowing a sibling, once its own transcript is aliased in', () => {
// The exact scenario COD-140's own comment lists first: a live/persisted row
// whose claudeSessionId aliases to an on-disk transcript. Once that alias
// successfully folds the transcript's own (failed) extraction into this row
// (sources includes 'history'), it must NOT then fall through to the
// workingDir guess and borrow an unrelated sibling's prompt -- same bug as
// the plain history-only case above, but for the resumed-session path the
// backfill mechanism was actually built for.
const merged = mergeUnifiedSessions({
live: [{ id: 'codeman-resumed', status: 'working', claudeSessionId: 'resumed-uuid', workingDir: '/shared' }],
history: [
// The resumed session's OWN transcript -- aliased in via claudeSessionId,
// but its own extraction found nothing.
{
sessionId: 'resumed-uuid',
workingDir: '/shared',
sizeBytes: 5000,
lastModified: '2026-01-01T00:00:00.000Z',
firstPrompt: undefined,
},
// An unrelated, newer sibling in the same directory.
{
sessionId: 'sibling-uuid',
workingDir: '/shared',
sizeBytes: 6000,
lastModified: '2026-06-01T00:00:00.000Z',
firstPrompt: "unrelated sibling's prompt",
},
],
});
const resumed = merged.find((m) => m.sessionId === 'codeman-resumed');
expect(resumed).toBeDefined();
expect([...resumed!.sources].sort()).toEqual(['history', 'live']);
expect(resumed!.firstPrompt).toBeUndefined();
});
// COD-145: lastPrompt backfill — mirrors the COD-140 firstPrompt path so the
// most-recent user prompt also reaches live rows whose id ≠ transcript UUID.
it('backfills lastPrompt onto a live session by claudeSessionId join (uuid-join)', () => {
+55
View File
@@ -0,0 +1,55 @@
/**
* @fileoverview The pane's last-Enter timestamp must survive a Codeman restart.
*
* `start()` resets `claudeSessionId` to the launch id even when re-attaching to
* a mux session whose CLI has since moved on (a `/clear` before the restart), so
* `lastSubmitAt` is the response viewer's only anchor for re-deriving the live
* conversation. If it is not persisted, a recovered pane shows the pre-`/clear`
* transcript until the user happens to type again — hours, in practice.
*
* Port: N/A (no server needed)
*/
import { describe, it, expect } from 'vitest';
import { Session } from '../src/session.js';
describe('session submit anchor', () => {
it('records the pane Enter and carries it into persisted state', () => {
const session = new Session({ workingDir: '/tmp' });
expect(session.lastSubmitAt).toBe(0);
expect(session.toState().lastSubmitAt).toBeUndefined();
const before = Date.now();
session.write('hello\r');
const after = Date.now();
expect(session.lastSubmitAt).toBeGreaterThanOrEqual(before);
expect(session.lastSubmitAt).toBeLessThanOrEqual(after);
expect(session.toState().lastSubmitAt).toBe(session.lastSubmitAt);
});
it('leaves the anchor unset for keystrokes that never submit', () => {
const session = new Session({ workingDir: '/tmp' });
session.write('hello');
session.write('\x1b[A'); // arrow-up: history recall, not a submit
expect(session.lastSubmitAt).toBe(0);
expect(session.toState().lastSubmitAt).toBeUndefined();
});
it('restores the anchor from persisted state on boot recovery', () => {
const submitted = new Session({ workingDir: '/tmp' });
submitted.write('prompt\r');
const persisted = submitted.toState();
const recovered = new Session({ workingDir: '/tmp', lastSubmitAt: persisted.lastSubmitAt });
expect(recovered.lastSubmitAt).toBe(submitted.lastSubmitAt);
expect(recovered.toState().lastSubmitAt).toBe(submitted.lastSubmitAt);
});
it('starts a pane with no persisted anchor at zero rather than NaN', () => {
const recovered = new Session({ workingDir: '/tmp', lastSubmitAt: undefined });
expect(recovered.lastSubmitAt).toBe(0);
});
});
+147
View File
@@ -0,0 +1,147 @@
/**
* Smart-copy chord gate (#211).
*
* `Ctrl+C` has to keep meaning "interrupt" whenever nothing is selected, so the
* decision is split in two: `shouldCopyTerminalSelectionFromShortcut()` only
* answers "did this chord ask to copy", and the caller in the xterm custom key
* handler decides what to do when there is no selection. These tests pin the
* gate itself (registry-aware, keydown-only) plus the static invariants that
* keep the generic capture loop from ever swallowing the interrupt.
*
* Strategy: run terminal-ui.js in a vm with a stub CodemanApp, the same harness
* shape test/command-palette-ui.test.ts uses for panels-ui.js. No DOM, no xterm.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it } from 'vitest';
const APP_SOURCE = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8');
type Shortcut = {
id: string;
disabled?: boolean;
bindings?: Array<{ modifiers?: string[]; key?: string; code?: string }>;
};
function loadTerminalHarness(registry?: Shortcut[]) {
const CodemanApp = function CodemanApp(this: unknown) {};
const context = vm.createContext({
CodemanApp,
window: {},
document: { getElementById: () => null, querySelector: () => null },
console,
MobileDetection: { isTouchDevice: () => false, getDeviceType: () => 'desktop' },
Object,
});
const terminalUi = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-ui.js'), 'utf8');
vm.runInContext(terminalUi, context, { filename: 'terminal-ui.js' });
const app = new (CodemanApp as unknown as new () => Record<string, any>)();
if (registry) {
app.getShortcutRegistry = () => registry;
// Real implementation, copied by reference from app.js semantics: ctrl/meta are
// interchangeable, every other modifier must be declared by the binding.
app.matchesShortcutEvent = (e: any, shortcut: Shortcut) => {
if (!shortcut || !Array.isArray(shortcut.bindings)) return false;
return shortcut.bindings.some((binding) => {
const mods = binding.modifiers || [];
const wantsPrimary = mods.includes('ctrl') || mods.includes('meta');
if (wantsPrimary !== !!(e.ctrlKey || e.metaKey)) return false;
if (mods.includes('shift') !== !!e.shiftKey) return false;
if (mods.includes('alt') !== !!e.altKey) return false;
if (binding.code && e.code === binding.code) return true;
if (binding.key && typeof e.key === 'string' && e.key.toLowerCase() === binding.key.toLowerCase()) return true;
return false;
});
};
}
return app;
}
const DEFAULT_REGISTRY: Shortcut[] = [
{
id: 'copy-selection',
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
},
];
function keydown(over: Record<string, unknown> = {}) {
return {
type: 'keydown',
key: 'c',
code: 'KeyC',
ctrlKey: true,
shiftKey: false,
altKey: false,
metaKey: false,
...over,
};
}
describe('terminal smart-copy gate', () => {
it('matches the default Ctrl+C and Ctrl+Shift+C chords', () => {
const app = loadTerminalHarness(DEFAULT_REGISTRY);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown())).toBe(true);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ key: 'C', shiftKey: true }))).toBe(true);
// Cmd+C on macOS: the registry treats ctrl/meta as interchangeable.
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ ctrlKey: false, metaKey: true }))).toBe(true);
});
it('ignores plain typing and unrelated chords', () => {
const app = loadTerminalHarness(DEFAULT_REGISTRY);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ ctrlKey: false }))).toBe(false);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ key: 'k', code: 'KeyK' }))).toBe(false);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ key: 'v', code: 'KeyV' }))).toBe(false);
});
it('only decides on keydown (the handler also runs for keypress and keyup)', () => {
const app = loadTerminalHarness(DEFAULT_REGISTRY);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ type: 'keypress' }))).toBe(false);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ type: 'keyup' }))).toBe(false);
expect(app.shouldCopyTerminalSelectionFromShortcut(null)).toBe(false);
});
it('honors a disabled shortcut so Ctrl+C goes back to being the interrupt', () => {
const app = loadTerminalHarness([{ ...DEFAULT_REGISTRY[0], disabled: true }]);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown())).toBe(false);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ key: 'C', shiftKey: true }))).toBe(false);
});
it('honors a rebound chord and stops claiming the old one', () => {
const app = loadTerminalHarness([{ id: 'copy-selection', bindings: [{ modifiers: ['alt'], key: 'y' }] }]);
expect(
app.shouldCopyTerminalSelectionFromShortcut(keydown({ ctrlKey: false, altKey: true, key: 'y', code: 'KeyY' }))
).toBe(true);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown())).toBe(false);
});
it('falls back to the default chord when no registry is available', () => {
const app = loadTerminalHarness(); // no getShortcutRegistry / matchesShortcutEvent
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown())).toBe(true);
expect(app.shouldCopyTerminalSelectionFromShortcut(keydown({ key: 'x', code: 'KeyX' }))).toBe(false);
});
});
describe('smart-copy wiring invariants', () => {
it('registers copy-selection in the shortcut registry', () => {
expect(APP_SOURCE).toContain("id: 'copy-selection'");
expect(APP_SOURCE).toContain("action: 'copyTerminalSelection'");
});
it('keeps copyTerminalSelection OUT of SHORTCUT_ACTIONS', () => {
// The generic capture loop preventDefaults on every match it dispatches. If
// the copy action were reachable from there, Ctrl+C would be swallowed with
// no selection and the user would lose the interrupt key.
const actionsBlock = APP_SOURCE.slice(
APP_SOURCE.indexOf('const SHORTCUT_ACTIONS = {'),
APP_SOURCE.indexOf('// Use capture to handle before terminal')
);
expect(actionsBlock.length).toBeGreaterThan(0);
expect(actionsBlock).not.toContain('copyTerminalSelection');
});
});
+166
View File
@@ -0,0 +1,166 @@
/**
* Smart copy in a real browser (#211).
*
* The gate itself is unit-tested in test/terminal-copy-selection.test.ts. What
* can only be proven in a browser is the half that decides whether the PTY sees
* an interrupt: xterm calls the custom key handler BEFORE its own cancel(), so
* returning false does not preventDefault, and a synthetic KeyboardEvent never
* triggers a browser default action. Both facts mean the copy/interrupt split
* has to be driven with real key presses.
*
* Assertions are on real state: what landed on the clipboard, and what xterm
* emitted through onData (the bytes that would reach the PTY).
*
* Browser-driven, so it is excluded from `npm run test:ci` like the other
* Playwright suites. Run locally: npm test -- test/terminal-copy-shortcut.test.ts
*
* Port: 3174 (per MEMORY.md, ports 3150+ for tests)
*/
import { describe, it, expect, beforeAll, afterAll } from 'vitest';
import { chromium, type Browser, type Page } from 'playwright';
import { WebServer } from '../src/web/server.js';
const PORT = 3174;
const BASE_URL = `http://localhost:${PORT}`;
describe('terminal Ctrl+C smart copy', () => {
let server: WebServer;
let browser: Browser;
let page: Page;
beforeAll(async () => {
server = new WebServer(PORT, false, true);
await server.start();
browser = await chromium.launch({ headless: true });
const context = await browser.newContext({ permissions: ['clipboard-read', 'clipboard-write'] });
page = await context.newPage();
await page.goto(BASE_URL, { waitUntil: 'domcontentloaded' });
await page.waitForFunction(() => (window as any).app?.terminal, null, { timeout: 30000 });
// The first write after load can be dropped while the app finishes wiring
// its render pipeline, so poll until one really lands in the buffer.
await page.waitForFunction(
async () => {
const term = (window as any).app.terminal;
await new Promise((r) => term.write('\r\nWARMUP\r\n', r));
const buf = term.buffer.active;
for (let i = 0; i < buf.length; i++) {
if (buf.getLine(i)?.translateToString(true).includes('WARMUP')) return true;
}
return false;
},
null,
{ timeout: 20000, polling: 500 }
);
}, 90000);
afterAll(async () => {
if (browser) await browser.close();
if (server) await server.stop();
}, 60000);
/** Write a marker line, optionally select it, and reset the capture state. */
async function setup(line: string, select: boolean, overrides: Record<string, unknown> = {}) {
await page.evaluate(
async ({ line, select, overrides }) => {
const app = (window as any).app;
const term = app.terminal;
const settings = app.loadAppSettingsFromStorage();
settings.shortcutOverrides = overrides;
app.saveAppSettingsToStorage(settings);
(window as any).__data = [];
if (!(window as any).__dataHooked) {
term.onData((d: string) => (window as any).__data.push(d));
(window as any).__dataHooked = true;
}
await new Promise((r) => term.write('\r\n' + line + '\r\n', r));
term.clearSelection();
if (select) {
const buf = term.buffer.active;
let row = -1;
for (let i = 0; i < buf.length; i++) {
if (buf.getLine(i)?.translateToString(true).includes(line)) row = i;
}
if (row === -1) throw new Error('marker line not found in buffer');
term.select(0, row, line.length);
if (!(term.getSelection() || '').trim()) throw new Error('selection is empty');
}
document.querySelector('.xterm-helper-textarea')!.dispatchEvent(new Event('focus'));
(document.querySelector('.xterm-helper-textarea') as HTMLElement).focus();
await navigator.clipboard.writeText('SENTINEL');
},
{ line, select, overrides }
);
}
async function outcome() {
await page.waitForTimeout(350);
return page.evaluate(async () => ({
data: (window as any).__data as string[],
clipboard: (await navigator.clipboard.readText()).trim(),
hasSelection: (window as any).app.terminal.hasSelection(),
}));
}
it('copies the selection and sends nothing to the PTY', async () => {
await setup('COPY-CASE-SELECTED', true);
await page.keyboard.press('Control+c');
const res = await outcome();
expect(res.clipboard).toBe('COPY-CASE-SELECTED');
expect(res.data).toEqual([]);
expect(res.hasSelection).toBe(false); // cleared, so a second Ctrl+C interrupts
});
it('still interrupts when nothing is selected', async () => {
await setup('COPY-CASE-UNSELECTED', false);
await page.keyboard.press('Control+c');
const res = await outcome();
expect(res.data).toEqual(['\x03']);
expect(res.clipboard).toBe('SENTINEL');
});
it('copies on the explicit Ctrl+Shift+C chord', async () => {
await setup('COPY-CASE-EXPLICIT', true);
await page.keyboard.press('Control+Shift+C');
const res = await outcome();
expect(res.clipboard).toBe('COPY-CASE-EXPLICIT');
expect(res.data).toEqual([]);
});
it('never interrupts on Ctrl+Shift+C with an empty selection', async () => {
await setup('COPY-CASE-EXPLICIT-EMPTY', false);
await page.keyboard.press('Control+Shift+C');
const res = await outcome();
expect(res.data).toEqual([]);
expect(res.clipboard).toBe('SENTINEL');
});
it('restores the plain interrupt when the shortcut is disabled', async () => {
await setup('COPY-CASE-DISABLED', true, { 'copy-selection': { disabled: true } });
await page.keyboard.press('Control+c');
const res = await outcome();
expect(res.data).toEqual(['\x03']);
expect(res.clipboard).toBe('SENTINEL');
});
it('follows a rebind, and Ctrl+C goes back to pure interrupt', async () => {
await setup('COPY-CASE-REBOUND', true, { 'copy-selection': { bindings: [{ modifiers: ['alt'], key: 'y' }] } });
await page.keyboard.press('Alt+y');
const rebound = await outcome();
expect(rebound.clipboard).toBe('COPY-CASE-REBOUND');
expect(rebound.data).toEqual([]);
await setup('COPY-CASE-REBOUND-2', true, { 'copy-selection': { bindings: [{ modifiers: ['alt'], key: 'y' }] } });
await page.keyboard.press('Control+c');
const res = await outcome();
expect(res.data).toEqual(['\x03']);
});
it('leaves Ctrl+V on the paste trap', async () => {
await setup('COPY-CASE-PASTE', false);
await page.keyboard.press('Control+v');
const res = await outcome();
expect(res.data.join('')).toContain('SENTINEL'); // pasted text, not ^V
expect(res.data.join('')).not.toContain('\x16');
});
});
+60 -3
View File
@@ -311,7 +311,7 @@ describe('terminal touch tap mouse guard', () => {
expect(sent).toEqual(['\x1b[<0;7;4M\x1b[<0;7;4m']);
});
it('wheel: forwards to the app only for verified sessions at the buffer bottom without Shift', () => {
it('wheel: forwards to the app for verified sessions without Shift, at ANY scroll position', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187' }]]);
@@ -323,8 +323,14 @@ describe('terminal touch tap mouse guard', () => {
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
expect(app._shouldForwardWheelToApp({ shiftKey: true })).toBe(false); // Shift = local scrollback
app.terminal.buffer.active.viewportY = 10; // browsing local scrollback
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
// Scrolled up into local scrollback still forwards. Gating this on the
// viewport being at the bottom is what let a repaint-mode CLI's own prompt
// box scroll off the screen: scrollToLastNonEmptyLine() parks the viewport
// above the bottom, so a tab switch silently pinned the wheel to local
// scrollback full of stale replayed frames. The wheel handler snaps the
// viewport back to the bottom before encoding the report instead.
app.terminal.buffer.active.viewportY = 10;
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
app.terminal.buffer.active.viewportY = 50;
app.terminal.modes.mouseTrackingMode = 'vt200'; // xterm's own encoder live
@@ -335,6 +341,24 @@ describe('terminal touch tap mouse guard', () => {
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
});
it('wheel: converts deltaMode line/page units instead of assuming pixels', () => {
const { app } = loadTerminalUiHarness();
app.terminal = { rows: 40 };
// DOM_DELTA_PIXEL (Chrome/WebKit, and every trackpad): ~110px per notch.
expect(app._wheelScrollLines({ deltaY: 110, deltaX: 0, deltaMode: 0, shiftKey: false })).toBe(4);
// DOM_DELTA_LINE (Firefox mouse wheel): deltaY is already lines. Read as
// pixels this rounded to 0 and fell through to the ±1 fallback.
expect(app._wheelScrollLines({ deltaY: 3, deltaX: 0, deltaMode: 1, shiftKey: false })).toBe(3);
expect(app._wheelScrollLines({ deltaY: -3, deltaX: 0, deltaMode: 1, shiftKey: false })).toBe(-3);
// DOM_DELTA_PAGE: one page is one screenful.
expect(app._wheelScrollLines({ deltaY: 1, deltaX: 0, deltaMode: 2, shiftKey: false })).toBe(40);
// A pure horizontal swipe must not fall through to a phantom -1.
expect(app._wheelScrollLines({ deltaY: 0, deltaX: 90, deltaMode: 0, shiftKey: false })).toBe(0);
// Shift + macOS trackpad reports the magnitude on deltaX (issue #154).
expect(app._wheelScrollLines({ deltaY: 0, deltaX: -100, deltaMode: 0, shiftKey: true })).toBe(-4);
});
it('wheel: gates claude forwarding on CLI version 2.1.187+ (unknown or older stays local)', () => {
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
@@ -438,6 +462,39 @@ describe('terminal touch tap mouse guard', () => {
expect(sent).toHaveLength(1);
});
it('forwarded scrolls (wheel AND touch) snap the viewport home first, then encode SGR ticks', () => {
const { app } = loadTerminalUiHarness();
const sent: Array<{ id: string; data: string }> = [];
app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude' }]]);
app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data });
const scrolledToBottom: boolean[] = [];
app.terminal = {
cols: 80,
rows: 24,
// Scrolled up into local scrollback: SGR coordinates address the LIVE
// screen, so the report would hit-test the wrong row without the snap.
buffer: { active: { viewportY: 10, baseY: 50 } },
scrollToBottom: () => scrolledToBottom.push(true),
element: {
querySelector: () => ({ getBoundingClientRect: () => ({ left: 0, top: 0 }) }),
},
_core: { _renderService: { dimensions: { css: { cell: { width: 8, height: 16 } } } } },
};
app._forwardScrollToApp(50, 50, -3);
expect(scrolledToBottom).toEqual([true]);
app._flushWheelSgrQueue();
expect(sent).toEqual([{ id: 'sess-1', data: '\x1b[<64;7;4M'.repeat(3) }]);
// Already at the bottom: no snap, just the report.
app.terminal.buffer.active.viewportY = 50;
app._forwardScrollToApp(50, 50, 2);
expect(scrolledToBottom).toHaveLength(1);
app._flushWheelSgrQueue();
expect(sent).toHaveLength(2);
});
it('allows trusted mouse events after the tap window expires', () => {
const { app, setNow } = loadTerminalUiHarness();
const { element, dispatch } = createElementHarness();