Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.
Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.
The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).
The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.
1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
terminal and rewrites it from the capture, which is a win when tmux holds
more than the browser, but for a repaint-mode pane that capture is roughly
ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
(escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
rewrite when it is more than one screen short; a refused session's cooldown
goes from 4s to 60s so a hollow pane stops re-fetching megabytes.
2. A false forwarding gate on a Claude session no longer means a dead gesture.
Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
SGR reports, at half a screen of travel per page key. Shift is excluded: it
keeps meaning "local scrollback".
3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
exception and guarded on !== undefined, so one timed-out or PATH-starved
probe at the first Claude session start disabled wheel-forwarding for every
Claude session until the server restarted, which fits a report of breakage on
phone, tablet and laptop at once. Success is still cached for the process
lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
a pure function so the semantics are testable without spawning claude.
4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
scoping: the setting keeps meaning exactly what it says, and fix 2 catches
the case where "local" is empty. The App Settings tooltip now says to leave
it off for Claude/Codex sessions.
5. _logScrollRouting() prints one line per session per distinct decision:
forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
mode, cliVersion, the opt-out state, mouse tracking and local scrollback
depth. #205 ran two rounds of remote guesswork over questions that line
answers directly.
Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.
- write() had two: the original description with @param and @example, then a
@returns-only block added on top, which dropped the params and examples from
hover. Merged into one. The @returns wording is also honest now — write() still
discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
and its declaration, leaving that function undocumented on hover and the doc
attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two blockers from the pre-submission gate, both reproduced before fixing.
1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
code says the opposite in as many words ("NOT an error response,
deliberately"), the commit message says response codes are unchanged, and the
test asserts the 200. It was a leftover sentence from an earlier iteration that
would have shipped into the CHANGELOG announcing an API contract change that
does not exist — and errorCode values are SemVer-relevant per
docs/versioning-policy.md.
2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
master left all 9 tests green, while the commit message sells "plus the whole
WebSocket path" as part of the fix. Three tests added against the real WS
route — ACK on delivery, ACK withheld and seq re-opened when the write did not
land, and a deduplicated frame still ACKed so the client can drop it. Verified
the other way round: with ws-routes.ts reverted, the middle one fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.
The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.
Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.
_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:
- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
non-empty row, so any session with trailing blank rows was left off-bottom
and every later wheel event went local without the user ever scrolling.
Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.
Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
Four fixes for the scrollback reports in #205 (plus its follow-up comment).
1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
(\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
to claude/codex/gemini, so it reached the browser verbatim. In the alternate
buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
which readline receives as shell history navigation. Both reported symptoms,
one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
never forwards a pane's alt-screen toggles to its client, it repaints;
captured from a real attach, vim/less/htop emit zero.
2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
onto tmux's history, and tmux repaints the pane rectangle instead of emitting
linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
same 60 lines emitted slowly added all 60. Scrolling up at the top now
re-pulls the full tmux scrollback and holds the user's place. Verified
end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.
3. The full-scrollback replay was gated on a single "first load after page load"
flag, which whichever session auto-selected consumed, so every other tab
started with one visible frame. Now tracked per session.
4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
per notch) scrolled one line where Chrome scrolls four or five, and capped
the forwarded SGR report at one tick. Line and page deltas are now converted,
and a pure horizontal swipe no longer falls through to a phantom -1.
Analysis and measurements: docs/scrollback-issues-analysis.md
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.
Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.
- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
with a single-flight guard. Async matters: under the same procfs pathology,
`execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
both reporting when they truncate. Pure so the regression tests can exercise the
shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
bounded, so a wedged `ps` cannot stop killSession from reaching its
process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
fresh; a truncated table would make whole subtrees invisible to the kill path.
13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.
The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.
Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.
Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.
That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.
Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.
No changeset: docs-only, rides the next release.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:
- Integration code cannot be sandboxed by Codeman because Codeman never
launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
documented isolation story.
Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a pointer to docs/extending-codeman.md at the end of the API section in
README.md and README.zh-CN.md, so the guide is reachable from where people
read about endpoints rather than only from CLAUDE.md.
Reading the README's programmatic guide alongside the new page surfaced three
errors in it, all now fixed:
- POST /api/sessions/:id/input takes `useMux`, not `useScreen`. The latter is
a legacy name that no longer appears in the schema.
- The page told integrators to send `\r` to submit. With `useMux: true` the
server delivers text and Enter as two separate writes (writeViaMux does
send-keys -l then send-keys Enter), so appending `\r` is wrong.
- "Unwrap the envelope" was incomplete: a few legacy GETs put the payload at
the top level, so the advice is now `body.data ?? body`.
Also cross-references the README's programmatic guide, which covers the
in-session case (CODEMAN_MUX, CODEMAN_API_URL, CODEMAN_SESSION_ID,
CODEMAN_HOOK_SECRET_FILE) that the new page deliberately does not duplicate,
and documents the optional clientId/seq exactly-once fields.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.
Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
mode 'antigravity' died on command-not-found. agy is not on npm, so it
gets its own installer step. --dir /usr/local/bin is load-bearing: the
default $HOME/.local/bin resolves to root's home at build time and is
unreachable by the `agent` user the container runs as. Verified inside
codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
present like the other CLI buttons, with a cyan identity matching the
toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
it as a satisfying AI CLI, and recommends it over Gemini in the install
hints. Detection only, no new auto-install path.
Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
opencode/codex/gemini when the code has included antigravity for a
while, said "all three modes", and omitted ANTIGRAVITY_ from the env
prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
remote-sessions' RemoteCommandMode were all stale.
Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.
test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.
Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.
isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Codeman has no plugin runtime by design: running third-party code inside
the process that spawns agents, on a server people expose over a tunnel,
would trade away the security posture that is a reason to use it. But it
already has four extension seams that work from any language with nothing
installed, and they were undocumented.
Documents web tabs (render your own UI as a tab), the SSE event channel
(react when an agent needs you), the HTTP API plus the codeman CLI (drive
it from a script), and hook events. Every endpoint, schema field, event
name and header in the page was read from source and then verified against
a running instance, including the localhost-only CORS behavior and the SSE
framing the example depends on.
Also corrects a stale line in CLAUDE.md: it claimed the HTTP/SSE API was
internal/unstable, which contradicts docs/versioning-policy.md, where the
API under /api/v1 was finalized as part of the stable surface for the 1.0
cut. No new stability commitment is made here; the page makes an existing
one discoverable.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two design/reference docs for the planned native-macOS VM isolation tier
("VM cases"), a location overlay on cases in the same shape as Docker and
remote-SSH cases, never a sixth SessionMode. Nothing is implemented; both
docs are marked PLANNED and are blocked on macOS 27 GA.
- vm-cases-plan.md: the Codeman-side design and phased plan. Swift helper
CLI, DiskImageKit base + per-case overlay, sessions riding the existing
remote-SSH machinery, VirtioFS workspace at the same absolute path, and
seeded credentials, each mirroring an established Docker-cases rule.
- vm-subsystem-apple-stack.md: what the Apple stack actually provides,
measured on the 27 beta rather than inferred from the WWDC session. Of
note: the 2-concurrent-macOS-VM cap is a kernel quota (refused at 39%
free RAM, so more hardware does not help), DiskImageKit has no flatten
API so exports must ship the layer chain, and a macOS guest renders
nothing without an attached view in an unlocked host session.
No credentials, hostnames, tailnet addresses or account names in either
file; every host/guest reference is a placeholder.
Also joins a table row that a stray blank line had split off into its own
malformed table.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.
Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.
Test fails against the pre-fix line and passes after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes#212. The file-preview overlay can now edit workspace text files in
place, phone-first: agent writes a file, you review it in the viewer, tweak
two lines, save, tell the agent to continue.
Backend (file-routes.ts, policy in src/config/file-editing.ts):
- GET file-content?edit=1: read-for-edit that never truncates (a truncated
buffer must never become an edit buffer), 512KB cap (413 over it), and
returns the sha256 hash + detected EOL the client echoes back on save.
- PUT /api/sessions/:id/file-content: edit-in-place only, with no O_CREAT
anywhere in the handler. Confinement matches the read path (realpath +
workspace boundary + ownership via findSessionOrFail), plus sensitive-path
and attachment-guard blocklists, a .git subtree deny, and an extension
allowlist (svg and env deliberately excluded). Optimistic concurrency via
baseHash: mismatch is a 409 unless force. Writes are wx-temp + fchmod +
fsync + rename, closing the validate-then-write TOCTOU window.
- Corruption guards: NUL sniff + UTF-8 round-trip compare (refuses binary
and latin-1), and server-side EOL re-application so a textarea's LF
normalization cannot rewrite every line of a CRLF file.
- Plain reads gain an additive editable flag the UI keys the button off.
Frontend (panels-ui.js + overlay markup/styles):
- Edit button on editable text previews; textarea editor with Save/Cancel,
dirty indicator, discard-confirm on cancel/close, and a conflict dialog
that offers overwrite (force) when the file changed on disk mid-edit.
- Phone: full-bleed window sized by --app-height so the editor and Save bar
track the OS keyboard; 16px editor font (iOS zoom guard); no autofocus.
- zh-CN strings for the new chrome.
Tests: pure policy unit tests plus a route suite that deliberately does NOT
mock node:fs. It runs against a real temp workspace so symlink escapes,
write-through of in-workspace symlinks, mode preservation, CRLF round-trip,
409/force, and the no-create property are exercised for real. Also verified
end to end on an isolated beta instance: 39-check curl matrix, Playwright
desktop flow (real clicks and typing, bytes asserted on disk, live conflict
with an external rewrite), and a 393px phone profile.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes#211. Copying from the terminal only worked through the browser
context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown,
so the muscle-memory copy failed silently and read as "no copy-paste at all".
With a selection, Ctrl+C now copies it, toasts, clears the selection and
sends nothing to the PTY. With no selection it falls through unchanged, so
the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never
falls through: an explicit copy that interrupts a running agent because the
selection happened to be empty would be a footgun.
Three details that keep the interrupt safe:
- The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the
no-selection path returns true WITHOUT preventDefault. xterm calls the
custom handler before its own cancel(), so returning false alone does not
cancel the event; the copy path therefore calls preventDefault explicitly,
or the browser would run its native copy on top of ours.
- copy-selection is a full registry entry (rebindable and disableable in App
Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the
same trick command-palette uses: the generic capture loop preventDefaults
every match it dispatches, which would cost the user the interrupt key.
- The gate is keydown-only, since the custom handler also runs for keypress
and keyup.
Copy goes through _copyText (Clipboard API, then hidden-textarea +
execCommand) rather than raw navigator.clipboard, because install.sh's LAN
option serves plain HTTP where navigator.clipboard is undefined; the
fallback steals focus, so the terminal is refocused afterwards.
Tests: test/terminal-copy-selection.test.ts pins the gate and the
SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real
key presses in chromium and asserts on the clipboard plus the bytes xterm
emitted (browser-driven, so excluded from test:ci like the other Playwright
suites). Verified manually on an isolated beta instance before landing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #200 and #201, which gate the welcome buttons and the run-mode
dropdown on whether the CLI is actually installed. Four corrections:
1. #200 also DELETED the Cloudflare Tunnel welcome button and the QR widget
outright. Its rationale is right (offering a tunnel where cloudflared is not
installed is a bad default) but the conclusion overshoots: the welcome QR is
the whole scan-to-connect-from-your-phone flow, and deleting it left a large
block of live tunnel code in settings-ui.js driving elements that no longer
existed. Both are restored and the button is gated on cloudflared, which is
what the stated rationale actually asks for. New cloudflared-resolver.ts
mirrors the CLI resolvers, and TunnelManager now shares its search path so
the button and the spawn can never disagree about where cloudflared lives.
2. Antigravity was missing from the run-mode gating, the one run mode LEAST
likely to be installed. It slipped past because #201 predates it. Covered
now, plus a static test that fails if a sixth mode reaches the dropdown
without being gated, so the next one cannot slip the same way.
3. The per-surface fetches are replaced by the injected availability object
already used for the Codex settings tab, so the codebase has one mechanism
rather than two. The status routes buy nothing as a gating source: every
resolver memoizes its PATH probe server-side, so a fetch is exactly as stale
as an injected value while costing a round trip every time the dropdown opens
and leaving the welcome buttons to flicker in after paint. The routes
themselves stay, including the /api/claude/status that #200 adds.
4. Unknown availability now reads as AVAILABLE for run buttons. Both PRs hid the
button on a failed fetch, so a blip left a working install with nothing to
click; a genuinely missing CLI only ever produced an error toast. The Codex
settings TAB keeps the opposite default, since hiding it costs nothing.
The dropdown query is also scoped to the menu: `.run-mode-option` is the class
the saved-dashboard and history rows use too, and a document-wide querySelector
would have found whichever came first in the DOM.
Fixes a latent environment-sensitivity in 816d900 while here: the index-title
test asserted the template was untouched apart from the title, which held only
on a machine with no codex installed.
Verified end-to-end against a real server on an isolated instance+socket, with
Playwright: gemini/codex hidden and claude/opencode/antigravity/shell shown,
matching this host, tunnel button back, Codex settings tab still hidden, no
console errors. Full test:ci sweep green (3902 tests).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #209 and #210. Both land a real fix (a pane that is a login shell
picks up /etc/profile and the per-user PATH entries an ssh remote command never
sees, which is what was failing agent CLIs with exit 127). Three corrections:
1. `-i -l` is no longer hardcoded onto the resolved shell. That path ultimately
comes from the passwd entry, which is user data and can name anything, and a
shell that rejects an unknown flag exits on the spot: nushell, elvish and xonsh
take neither flag, so a user with one of those in passwd would have gotten a
dead pane on arrival, which is exactly the #208 failure #209 builds on top of.
loginShellArgs() applies them only to the POSIX-family shells verified to
accept both, and a test really launches every allowlisted shell present on the
machine rather than trusting the set. csh/tcsh are excluded deliberately: tcsh
honors -l only when it is the ONLY flag.
2. `remain-on-exit on` -> `failed`, moved LAST in the tmux command chain. `on`
keeps the pane after a CLEAN exit too, so typing `exit` in a remote shell
stranded a dead pane, the session outlived it, and the next launch's `-A`
reattached to that corpse: "Pane is dead (status 0)" instead of a shell,
permanently, on the DEFAULT path. Verified against a real tmux, as was the
fix: `failed` tears the session down on status 0 and keeps the pane on 127
with the "command not found" still on screen, which is the case #210 wanted.
It is last because tmux aborts the remaining commands of a `\;` sequence once
one errors (also verified) and `failed` needs tmux >= 3.2 on the REMOTE host;
leading, a rejection there would have silently dropped status/mouse/prefix/
escape-time/window-size along with it.
3. `$SHELL` -> `"${SHELL:-/bin/sh}"`, via one shared remoteLoginShellCommand()
helper instead of the string being rebuilt in tmux-manager as well.
Also corrects the rationale both PRs carried: a tmux pane already hands the shell
a tty, so it was interactive all along ($- contains i for a bare /bin/bash in a
pane) and ~/.bashrc was always being sourced. `-l` is the flag doing the work.
End-to-end verified, not just unit-tested: the emitted remote pane command was
run through all three quoting layers under a minimal sshd-style PATH with the
CLI installed only on a login-shell PATH entry, and it resolved and launched the
CLI with its arguments intact and a space-containing remote path preserved.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-up to #202. The dotdir decode landed there was reachable only when
nothing else matched first, and in the greedy half it was not reachable at all.
decodeProjectKey() splits the project key on '-', so the '/.' that the encoder
collapses leaves an EMPTY segment behind. Both loops offered that empty string
as a candidate directory name, and isDir(current + '/' + '') stats current + '/',
which always succeeds. So the empty segment matched unconditionally:
- backtracking half: ~/.sib resolved to "/home/x//sib" whenever a non-dot
sibling ~/sib existed (wrong directory, and a doubled slash that then fails
every string comparison against session.workingDir). Without a sibling it
only backtracked out by luck.
- greedy half: that loop is shortest-match-first, so the empty candidate
matched on the FIRST iteration and set matched=true, leaving #202's dotdir
branch permanently dead there.
An empty string is never a real path component, so skip it in both loops. The
unmatched tail then has to handle the empty segment too, or it would append a
bare '/' and re-introduce the '//' path it just stopped producing; it now emits
the dotdir guess instead, which is what the encoder implies.
Regression test asserts both halves: the dotdir wins over the non-dot sibling,
and the result never contains '//'. Verified it fails on #202 as merged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both settings on the App Settings "Codex CLI" tab (bypass approvals, animated
status effects) are handed to `codex` at launch, so on an instance where the
binary does not resolve the tab offers choices nothing can act on. Gate it on
availability instead.
renderIndexHtml injects window.__codemanCodexAvailable, mirroring the existing
gesture-availability flag, and settings-ui.js hides the tab button when it is
absent. Injected rather than fetched on modal open so the tab cannot flicker in
and back out; isCodexAvailable() memoizes its PATH probe, so the per-render cost
is nil. Installing codex later needs a restart, exactly like the
/api/codex/status route that already backs the Run menu. Solo popups skip the
probe since they have no settings modal.
Only the tab BUTTON is toggled. The panel already carries
.modal-tab-content.hidden unless it is the selected tab and openAppSettings()
always reopens on Display, so an unreachable button keeps the panel unreachable.
The inputs stay in the DOM and are still populated and read back on save, so a
user without codex cannot silently wipe the codex preferences of an instance
that has it. Animations stay off by default for new local Codex sessions.
Verified in a browser on this host, which has no codex: the flag is absent, the
Codex tab is hidden while the other tabs are unaffected, and saving App Settings
with the tab hidden leaves codexAnimationsEnabled/codexDangerouslyBypassApprovals
untouched. With the flag forced on, the tab appears, its panel opens, and
toggling the visible slider persists. The openAppSettings coupling test was
checked to fail when the call is removed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
runAntigravity() landed on master after this branch was cut, so it kept the
exact pattern the rest of this PR removes: terminal.clear() plus direct
writeln into whatever session happened to be active. Merging master in
surfaced it, leaving one of six run modes still wiping the active session's
xterm on launch.
Also adds regression coverage that can actually see the bug. The existing
test drives the three helpers directly, so it stays green even when a run*()
function is reverted to writing at the terminal itself: reverting
runClaude()'s call site keeps all 16 tests passing. The new static guard
scans session-ui.js and fails if any run*() body touches
this.terminal.clear/writeln, which catches a regressed call site and would
have caught runAntigravity on its own. A second unit test covers the
home-screen path that nothing exercised: with no active session, launch
progress must still clear and render in the terminal.
Verified in a browser against a live instance. With a session active,
runShell() and runAntigravity() leave its terminal untouched (clear() calls:
0, writes: 0) and emit one info toast; on master the same run wipes the
session's marker text. The session-less home screen still clears and writes
exactly as before.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two independent ways a tab description could be typed in and silently lost.
1. Session Options modal (deterministic). The Session Name input saves on
blur, and every autosave handler in the modal bails on a null
editingSessionId. closeSessionOptions() cleared that id BEFORE hiding the
modal, and hiding it is what blurs the input, so the save always ran too
late and returned early. Escape and backdrop-click lost the name with no
PUT at all; only the X button worked, because mousedown blurs the input
before the click handler runs. Fix: blur the focused modal field first,
then clear the id. That also covers the auto-compact prompt, which saves
on change and had the same fate.
2. Right-click inline rename (racy). The _inlineRenameActive guard from #81
sits in renderSessionTabs() (the scheduler) and _fullRenderSessionTabs(),
but not in _renderSessionTabsImmediate() (the debounced executor). A
render queued in the ~100ms before the rename opened still fires and the
incremental branch rewrites .tab-name's innerHTML, destroying the input
mid-keystroke: it commits a truncated name, or, if it lands before the
first keystroke, closes the rename so everything typed after goes
nowhere. Fix: guard the executor too. finishRename() re-renders on both
commit and cancel, so a render dropped there is picked back up.
Verified end-to-end against a live server on an isolated instance: all three
modal close paths now persist the name, and the rename input survives a
render mid-typing. Both regression tests were checked to fail with their fix
reverted; the render one was vacuous at first because the synthetic tab sat
on <body> instead of inside #sessionTabs, so it now builds the tab in the
real container.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).
Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.
Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pr/ holds machine-local promo drafts that are never meant for git. Anchored with
a leading slash so it matches only the root dir, matching the /public entry below
it, rather than swallowing any nested pr/ elsewhere in the tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.
- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
injected via socket-scoped `tmux setenv`, never on the spawn command line, so
the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
colours. `runAntigravity()` routes remote/docker cases through
`POST /api/quick-start` and skips the local status probe for them.
Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All OFF by default (the `legacy` theme), so an untouched install behaves exactly
as before and every mark/apply hook short-circuits on its first line. Opt in via
App Settings > Appearance > Entrance Animations; per-surface control and a live
preview lab at ?animlab=1.
Surfaces and styles:
- Tabs: slide, pop, crt, unroll, boot, flip. A batch launched together cascades
by a configurable stagger.
- Terminal pane: crt, boot, wipe, slide, fade.
- Agent windows: fly (the pre-existing tab-to-window flight, still the default),
crt, materialize, unfold, beam, pop.
- Connection lines: draw, packet, fade.
Three constraints drove the design:
1. Tabs and connection lines are DESTROYED mid-animation on every re-render:
_fullRenderSessionTabs() replaces the strip's innerHTML and
_updateConnectionLinesImmediate() does `svg.innerHTML = ''`, both of which run
constantly while sessions and agents spawn. Each is tracked by id and
re-applied to the fresh element with a NEGATIVE animation-delay so it resumes
at the same offset instead of restarting or snapping. Verified on the real
path: a forced rebuild mid-draw resumed at -0.243s.
2. Terminal-pane styles animate transform/opacity/clip-path ONLY. xterm's
FitAddon derives rows+cols from getComputedStyle(parent).width/height, the
untransformed layout box, so transforms are invisible to it; animating
width/height/padding would have resized the PTY. Verified by forcing
fitAddon.fit() eight times mid-animation: dimensions held at 178x38.
3. A window entrance that transforms also moves the rect its connection line
aims at (crt drifts it 109px, pop 81px). `beam` animates opacity/filter only
(0px drift) so its line can draw toward a stable target; the others refresh
the lines on animationend.
Also fixes: an agent window spawning hidden (its agent belongs to a background
tab) is display:none, so its animation never runs and animationend never fires,
which left the entrance class and its inline custom property stuck on the window
permanently. Hidden windows now skip the entrance entirely.
Styles persist to their own codeman:*Anim localStorage keys, keeping them
per-device without touching the .strict() SettingsUpdateSchema.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-ups from the PR #175/#176 reviews:
- Rewake helper self-terminates on its own 6h deadline and when orphaned,
instead of relying on Claude Code to reap the poller
- Rewake marker versioned (V2) with a version-agnostic ownership prefix, so
future script updates replace older handlers instead of duplicating them;
regression test covers the V1 to V2 swap
- HOOK_TIMEOUT_MS renamed to HOOK_TIMEOUT_SECONDS = 10: the hook timeout
field is seconds (the CLI multiplies by 1000), so the curl hooks have
effectively had a ~2.8h timeout since COD-54
- Test echo PTY switches to raw mode: each input byte echoes exactly once
(tty line discipline doubled every line and buffered until Enter)
- test/setup.ts: drain in-flight console-log rpc forwards before environment
teardown (fixes the EnvironmentTeardownError that failed CI twice on the
merge commit with all 3820 tests passing), clean the temp home on process
exit (fully-skipped files leaked it), fix the Windows Playwright cache
fallback path
- test/webview-proxy.test.ts: stop naming the vitest environment directive in
prose; vitest matches it inside comments and silently ran the whole file
under the jsdom environment while the comment claimed node
- CLAUDE.md: document the temp-HOME and echo-PTY test isolation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The workspace publishes two packages, changesets creates a GitHub release
for each, and GitHub awards "Latest" to whichever was published last. That
is a race: 1.9.2 kept the badge, 1.9.4 lost it to xterm-zerolag-input@0.1.7
by two seconds. Set make_latest in the rename PATCH, which runs after every
package release already exists.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PUT /api/settings service toggles now resolve from `merged` (persisted +
incoming) instead of the raw request body, so a partial PUT no longer
starts the subagent watcher and stops the workflow + image watchers by
treating every omitted key as "apply the default". Pinned by a 4-case
regression test verified to fail against the old handler.
Also trims the links line from the Codeman callout in the
xterm-zerolag-input README.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Plan-usage chip defaults ON on desktop (handhelds stay OFF), resolved
through a single planUsageChipEnabled() helper so the checkbox, the chip
and the create-time statusLineTelemetry flag cannot disagree. Correct the
stale "Cron button defaults ON" comment (it is OFF in code, template and
CSS) and the styles.css comment claiming the server strips the chip's
hidden class.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace the misaligned 8-line keystroke-flow diagram (its branch sat two
columns off the junction it attached to) with a two-lane contrast that
makes the same point in two lines: stock xterm.js waiting 300ms vs the
overlay painting immediately. The mechanism detail it was annotating
moved into the following paragraph.
Add a Codeman callout between the badges and the demo GIF, with links to
getcodeman.com, the install one-liner and the repo, and rewrite the
bottom Origin section so it argues credibility instead of repeating the
promo.
Not released: the npm page updates only on publish, so the next COM
needs an "xterm-zerolag-input": patch changeset for this to ship.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Rewrite the xterm-zerolag-input README (hero demo GIF, value-first
structure) and fix its drift against the source: 175 tests not 78,
CJK/emoji wide-char support documented instead of listed as a
limitation, setPrompt() documented.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
The response-viewer detail now lives in architecture-invariants, so the
Claude turn-grouping and restored-placeholder rebind notes moved there.
Changeset rewritten to record the measured effect on real transcripts.
- One compact Mobile-Optimized Web UI section: the two current screenshots
(mobile-session-keyboard, mobile-toolbar-enter) side by side, comparison
table, condensed feature bullets, then QR auth
- Drop the outdated black-background phone screenshots
(mobile-landing-qr.png, mobile-session-active.png)
- Same restructure in README.zh-CN.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Both picker endpoints are a second file-serving surface, and they
inherited neither the attachment guard's confinement nor its ownership
scoping. Two separate holes:
1. `sessionId` contributes that session's workingDir as a browse root,
but it was resolved straight off ctx.sessions/ctx.store with no owner
check, unlike the nine other session-scoped handlers in this file. A
non-admin could pin ANOTHER user's working directory as a root just
by passing their session id, then list and preview underneath it. Now
runs canAccessOwned and reports 404, which also avoids confirming
that a session id exists.
2. `Home` and `CASES_DIR` were unconditional roots for every caller.
Per-user spaces live at <USER_SPACES_DIR>/<username>, which is INSIDE
homedir(), so the Home root alone exposed every other user's
workspace. A multi-user non-admin now gets only their own
userSpacePath plus anything explicitly listed in
CODEMAN_FILE_PICKER_ROOTS. /mnt/d is dropped as well: a broad host
mount should be an explicit operator decision in a multi-user
deployment, and operators who want it can name it in that env var.
Admins and single-user mode keep the host-wide roots, so behavior is
unchanged unless CODEMAN_MULTIUSER is on (opt-in, off by default).
All three discriminating tests were verified to fail against the
previous code: browse and preview both returned 200 instead of 404, and
the roots came back as [Home, Codeman Cases, ...] instead of [My Space].
Full suite green, 3784 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The zerolag composition renders on a pure black page background, which
read as an outdated screenshot when placed right under the hero. Top of
the README now shows only the current-skin visuals (subagent gif + tour).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Move the install one-liner (curl getcodeman.com/install | bash) and value bullets to the top so the pitch and quick start fit in the first scrolls
- Promote Zero-Lag Input Overlay to right after the hero, with a new side-by-side phone demo gif generated from the current zerolag master
- Switch all install commands (incl. WSL) to the getcodeman.com short URL
- Remove outdated screenshots (multi-session-dashboard.png, ralph-tracker-8tasks-44percent.png) and the old zerolag-demo.gif
- Remove Ralph tracker content: tracking section, API table, CLI example, autonomy-table row, architecture-diagram node
- Mirror all changes in README.zh-CN.md
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The short-code distribution test asserted that no base62 character
deviated more than 15% from its expected count. That statistic is the
maximum over 62 correlated near-normal cells, so its tail is fat: at
n=36000 the per-cell relative SD is ~4.1%, which puts the 15% bound at
|z| ~ 3.65, and taken as a max over 62 cells it fires on a perfectly
uniform generator about 1.6% of the time. Measured directly: 48 spurious
failures in 3000 simulated runs. It had been rerun-to-green repeatedly
and most recently red-herringed a PR review.
Chi-square is the correct test for "is this multinomial uniform", and
unlike 0.15 its threshold is derivable. df=61, Wilson-Hilferty puts the
p=1e-6 critical value at ~129, so the bound is 130.
Power is unchanged. Removing rejection sampling from generateShortCode
reintroduces modulo bias (256 % 62 = 8, so eight characters draw five
chances per 256 instead of four) and was verified against the real code
in an isolated worktree: chi-square 243.06 against the 130 limit. The
threshold sits in a wide empty gap, 3000 clean runs peaked at 104 while
200 biased runs bottomed out at 174.5.
Also iterate the alphabet explicitly rather than the observed keys, so a
character that never appears counts as zero instead of being skipped.
Verified: 30/30 consecutive runs of the real test pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
Route counts reconciled against master's numbering (files 14 -> 16 for
the two new filesystem endpoints, total ~197 -> ~199) and the path
picker's detail moved into architecture-invariants under its own
section.
The light skins themed the app chrome, but a class of status badges and
accent-tinted pills still hardcode pale light-on-dark ink (#cdddff,
#9dc0ff, #ffc107, #fff) over a low-alpha tint. Measured on a rendered
page, that lands at 1.0 to 1.9:1 under all four light skins: the search
filter chips (Sessions / Events / Files) render as empty blue pills.
Re-point the ink at each skin's own dark tokens and keep the tint as the
category signal, which moves the same components to 3.2 to 14:1.
Also pin --floating-bg on the OG skin. The new :root default is slate
rgba(31,38,48,.96), which suits the Daylight palettes (their glass
header is already rgba(31,38,48,.85)) but repaints OG's modals, command
palette and floating windows away from the neutral near-black that skin
is built on.
Verified against a live instance across all seven skins, plus a real
shell session for terminal ANSI output. Full suite green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both were unreferenced by either README and are now kept in Ark0N/gittrend
under assets/codeman-demos, alongside the source recordings they were cut from.
As with the GIF removal, this does not shrink the repository: the blobs remain
in history and only new checkouts stop carrying them.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Neither README references it: both switched to the dated
subagent-demo-20260724.gif in 8e9f254, which kept this file only so external
hotlinks would keep resolving. Removing it now at the maintainer's request.
Note this does NOT shrink the repository. The blob stays in history, so clone
size is unchanged; only new checkouts stop carrying the 29MB file. Actually
reclaiming the space needs a history rewrite, which would invalidate every
existing clone and is a separate decision.
The three capture scripts that write docs/images/subagent-demo.gif are
unaffected: they create the file, they do not read it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three separate truncations, each cutting a clickable link short so it opened the
wrong target (or nothing at all).
1. A single `&` ended the match. It is a query-parameter separator, so every real
query string was cut: a WordPress edit link resolved to `?post=1479` and opened
the post list instead of the editor, and Claude Code's own `/login` URL was not
usable at all. `&` is now part of a URL; `&&` stays a boundary, since that is
the shell operator and never appears inside one. A lone trailing `&` is still
trimmed as punctuation.
2. Links longer than the terminal is wide were cut at the row boundary. xterm
calls the link provider once per visible ROW and translateToString returns only
that row, despite a comment here claiming it handled wrapping. The provider now
stitches the continuation rows back into one logical line and maps match offsets
back to (x, y), so a link can span rows.
Two kinds of continuation exist and handling only the first is not enough. A
SOFT wrap is the emulator running out of columns, which flags the next row
`isWrapped`. A HARD wrap is the program wrapping its own output and emitting a
real newline, which flags nothing: Ink does this, which is why the /login URL
was cut at the window edge and why the clickable part grew when the window was
widened. A row that fills the full width is now treated as continuing into the
next, that being the only trace a hard wrap leaves behind. Bounded to 12 rows so
a screenful of wide output cannot make every hover re-scan the viewport.
3. Image and PDF paths were not matched at all. `.claude-images/paste-*.png`, what
Codeman writes for a pasted screenshot, rendered as plain text. Those extensions
are now linked and open the file preview, which renders images inline, rather
than the log viewer, which would show binary noise.
Verified in a real terminal: a 450-char /login URL hard-wrapped across 5 rows with
zero isWrapped flags in the buffer (so a genuine hard wrap, not the soft case)
links intact, as do soft-wrapped URLs and a wrapped attachment path. Regression
cases added to link-provider-regex.test.ts, which extracts the patterns from the
shipped source so they cannot drift. Its existing ReDoS guard still passes, which
matters because this changes a pattern that once froze the tab on hover.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a "Web / URL" section to the Run dropdown. A saved URL renders as a tab in
the same strip as Claude/Codex/Gemini sessions, with the same Alt+1..9 numbering,
so Codeman is one mission control instead of Codeman plus a pile of browser tabs.
A webview is NOT a sixth SessionMode: no PTY, no tmux, no respawn, no idle
detection. It is a separate resource sharing only the tab strip and the main
content area, the same call that keeps Docker and remote-SSH as case overlays.
Dashboards are proxied through Codeman's own origin, because a direct iframe
fails three ways at once in the shipped deployment: prod serves HTTPS behind
tailscale serve, so http:// targets are hard-blocked as mixed content (with no
override at all on iOS Safari); Grafana/Portainer-class dashboards send
X-Frame-Options: DENY; and our own default-src 'self' CSP blocks cross-origin
frames. Proxying dissolves all three and leaves the production CSP byte-for-byte
unchanged, since /webview/... is already covered by 'self'. A useful side effect:
the fetch happens server-side, so a tailnet-only dashboard is reachable from a
phone that is not on the tailnet.
The proxy is not an API surface. It authenticates on a 192-bit capability in the
path (memory-only, rolling TTL, bound to the minting user, revoked on edit or
delete) and is correspondingly exempt from the cookie and Origin checks, because
a sandboxed iframe is opaque-origin: it sends no SameSite=lax cookie and its
writes arrive with Origin: null. The Host allowlist is never bypassed. A second
Referer-keyed form of the exemption exists for root-absolute assets and is fenced
to safe methods on non-/api, non-/ws, non-/q paths.
Iframes omit allow-same-origin unless a URL is explicitly marked trusted, since a
proxied page is served from Codeman's own origin and could otherwise read this
document and drive the agent-spawning API. Authorization and codeman_session are
stripped upstream in BOTH modes, so CODEMAN_PASSWORD cannot leak into a dashboard.
Two things only a real browser reveals, both presenting as the dashboard's own
"Failed to fetch" while the page itself renders fine:
- Runtime-built root-absolute URLs (fetch('/api/data')) escape <base href> and
land on Codeman's root. Widening the Referer fallback into /api would trade
security for it, so an injected shim patches fetch/XHR/WebSocket/EventSource
inside the frame instead, removing the class rather than the guard.
- An opaque-origin document CORS-checks every request, including to the host it
was served from. Script/css/img loads are not CORS-checked, which is why the
page renders while its API calls die. The proxy now emits CORS headers and
answers preflights itself. registerSecurityHeaders answered every OPTIONS with
a bare 204 before routing, carrying no ACAO for Origin: null, so that
short-circuit now exempts a valid capability.
Neither is reproducible with curl, which does not enforce CORS.
Also fixes a pre-existing bug found on the way: .toolbar has backdrop-filter,
making it a stacking context that trapped .run-mode-menu's z-index:1000, so
.welcome-overlay painted over the whole Run menu. With no session open, every
item in it (Claude Code included) was unclickable.
Verified end to end against a real tailnet dashboard: live data, WebSocket push,
no failed requests, and switching tabs does not reload the frame. 98 new tests
cover the pure rewrite helpers, the CORS helper, the shim's rewrite logic, route
CRUD, and every edge of the auth exemption.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Feature + layout changes:
- Phone toolbar: Enter replaced Shell below 430px; shell launching moved into
the Run dropdown. Documents the ordering (setRunMode -> run -> runShell) and
that runMode is a loose string server-side so new modes need no schema edit.
- Repo root layout: config/ holds knip.json, Prettier config is the package.json
"prettier" key, SECURITY.md is under .github/, and the list of files that must
stay at the root with the reason each one is pinned there.
- Pointer to docs/SPEEDRUN.md, which nothing linked to after the move.
Traps worth not rediscovering:
- sendEnterKey MUST use triggerDataEvent, not sendInput or a raw POST. Local
echo is on by default on touch devices, so typed text is buffered client-side
and a bare CR submits an empty line while the text stays stranded. Cost me two
wrong fixes before the real cause surfaced.
- styles.css nests skin overrides under html:not([data-skin="og"]), giving a
bare .btn-toolbar rule (0,2,1) which outranks .btn-toolbar.btn-x (0,2,0) in
mobile.css whatever the load order. Explains why mobile.css needs !important.
- Browser tests pass vacuously on mobile input: sendInput() bypasses the overlay,
and headless Chromium reports isTouchDevice() false even with hasTouch, so the
local-echo branch never runs. Assert on overlay state and the tmux pane.
- The working tree is shared with other agent sessions: check the branch before
every commit (a commit silently landed on feat/web-tabs today and the push to
master reported "Everything up-to-date"), push with HEAD:master rather than
checking master out, and never git add -A.
- COM step 5 no longer tells you to git add -A, which has swept another
session's WIP into a release before.
Verified: 33 relative links and 30 invariants anchors all resolve, and every
factual claim re-checked against the tree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Continues trimming the repo root so the README is reached with less scrolling.
Root files: 19 -> 15 across both passes.
- knip.json -> config/knip.json, joining eslint.config.js and the vitest
configs. `npm run knip` now passes --config explicitly. Verified by A/B: the
run from the new location produces byte-identical findings and the same five
configuration hints as from the root, so knip resolves its globs relative to
cwd rather than the config file. Those hints are pre-existing, not caused by
the move.
- .prettierrc -> the "prettier" key in package.json, a config source Prettier
reads natively, so editor format-on-save keeps working with no --config flag
anywhere. Verified live: `npm run format:check` still passes across src/**,
which it could not if the config had been lost (Prettier's defaults are
double quotes at 80 columns and would flag nearly every file).
.prettierignore deliberately stays at the root: Prettier resolves it relative
to cwd, so moving it would require threading --ignore-path through every
script and would break editor integration.
Everything else in the root is load-bearing: .editorconfig (walks up from the
edited file), .nvmrc/.npmrc (read from the project root), tsconfig.json (bare
`tsc` discovers it), LICENSE (GitHub license detection), install.sh (its raw
URL is the published one-liner in the README and cannot move without breaking
every copy in the wild), plus the five documented .md files.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Trims the repo root listing so the README is reached with less scrolling.
Only these two were movable; the other five root .md files are load-bearing
and stay put:
- README.md / README.zh-CN.md — the landing page and the language-switcher
entry point
- CLAUDE.md — Claude Code loads project instructions from the ROOT path;
moving it silently breaks every future session in this repo
- AGENTS.md — the agent-convention file Codex reads from the root and injects
as context (see the comments in session-routes.ts)
- CHANGELOG.md — the changesets default writer emits it next to package.json,
so moving it breaks `npm run version-packages`
GitHub officially resolves .github/SECURITY.md, so the Security policy tab
keeps working. Inbound links updated in both READMEs, CLAUDE.md and
docs/versioning-policy.md. CHANGELOG.md also names SECURITY.md but is left
alone: it is a historical record, not a live reference.
The move broke a link the other direction too: SECURITY.md pointed at
docs/security-architecture.md, which from .github/ resolved to
.github/docs/... — repointed to ../docs/. All relative links in the six
touched files verified resolving (62 links, 0 broken).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two captures from an actual phone, replacing the placeholder-ish shots:
- Mobile table, middle cell: the full-height capture with the keyboard open,
showing the accessory bar and the new Enter button while answering a plan
prompt. Supersedes mobile-session-question-20260727.png from this morning,
which showed the pre-Enter toolbar.
- Touch-Optimized Interface: the cropped toolbar capture as a standalone
560px figure, where a near-square crop reads better than it would squeezed
into a 260px table cell.
Picking the tall capture for the table also fixes a row the earlier square
shot had left lopsided: the three cells now render 473 / 482 / 469px tall
instead of 473 / 263 / 469.
Adds a "Dedicated Enter button" bullet documenting the behaviour, including
why it replays the keypress (local-echo flush) rather than sending a bare
carriage return, and that shell launching moved into the Run dropdown.
Both READMEs updated so EN and zh-CN stay in sync.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
On phones the toolbar slot held "Shell", which starts a rarely-needed session
type. Sending Enter is a constant need on a touch keyboard, so the slot now
holds a dark blue Enter button and shell launching moves into the expandable
Run dropdown (Terminal / Shell, label "Run SH"). Desktop and tablet are
unchanged: the green Run Shell button stays exactly where it was.
Enter goes through xterm's own input path:
coreService.triggerDataEvent('\r', true)
NOT through sendInput() or a direct POST to /input. localEchoEnabled defaults
to MobileDetection.isTouchDevice(), so on a phone the characters you type are
buffered client-side in the LocalEchoOverlay and have never reached the PTY.
The onData Enter branch in terminal-ui.js is what flushes that buffer before
sending \r. A bare \r submits an empty line and leaves the typed text stranded
on screen, which presents as "the Enter button does nothing". Replaying the
keypress reuses the overlay flush, the flushed-offset cleanup and the 80ms
text-before-CR ordering instead of reimplementing them.
Verified with local echo forced on: before the fix the overlay still held
"echo OLD_WAY" after Enter; after it, pendingText is empty and the command
executes in the pane.
The !important on the Enter button's colors is required, not habit: styles.css
nests its skin overrides inside `html:not([data-skin="og"]) { … }`, so a plain
.btn-toolbar there resolves to (0,2,1) and outranks .btn-toolbar.btn-enter at
(0,2,0). Without it the button renders in generic toolbar grey.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replaces the middle cell of the Mobile-Optimized Web UI table in both
READMEs. The new shot shows an agent's multiple-choice prompt being answered
on a phone, with the touch accessory bar and bottom toolbar visible, which
demonstrates more of the mobile UI than the old idle-session capture.
Uses a dated filename per the convention the other 2026-07 images follow.
That also avoids GitHub's image cache serving the old picture, which an
in-place overwrite of mobile-session-idle.png would have risked. The old
file is left on disk so any existing external link to it keeps working.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two live shields.io badges in the header block of both READMEs, linking to
the contributors graph and the commit history. Colors reuse the existing
palette (3b82f6, 1e3a5f) and keep the flat-square style.
Verified both endpoints render real data matching the GitHub API
(contributors: 13, commits: 1.5k against 1,460 on master) and that master
is the default branch, so the /commits/master link target is correct.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
CLAUDE.md was 110KB (~27.5k tokens) loaded into every session, with 30 lines
carrying 49% of the bytes as single-paragraph walls (the Docker cases entry
alone was 9,388 chars). Extract the implementation detail verbatim into
docs/architecture-invariants.md (41 sections) and leave the rule plus a
pointer inline. Result: 59.5KB, ~14.9k tokens, 46% smaller.
Also:
- Add CLAUDE.md to .prettierignore. Prettier's markdown printer escapes
underscores in the glob-heavy paths used throughout, which had already
corrupted the Ultracode paragraph (agent-*.jsonl became agent-\_.jsonl,
collapsing backtick spans). npm run format:check is unaffected; its globs
are src/** only.
- Move version archaeology (PR numbers, ticket ids, commit shas, "was X now
Y" lineage) into the invariants doc, keeping the rules and their reasoning
inline.
- De-duplicate the Core Files table against Key Patterns.
- Document install.sh in Scripts, and why Prettier's scope is deliberately
narrow (14 hand-formatted public JS modules are guarded by
check:public-assets and check:frontend-syntax instead).
Two factual fixes found while verifying: displayKeys is a client-side merge
policy, not a wire filter, and showResponseViewer / showPlanUsageLimits /
language are declared in SettingsUpdateSchema and do persist server-side; and
the respawn route count is 7, not 18.
Verified: 30/30 cross-doc pointers resolve, 1,184 of 1,190 backticked
identifiers from the original survive (the 6 others are dropped archaeology
or the prettier-corrupted spellings), 59 table rows well-formed,
format:check and check:frontend-syntax clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replaces the dated subagent-spawn.png with the recaptured floating-windows
still (clean header, three haiku Explore agents, connector lines) and adds
the live ultracode workflow-run window below the feature bullets, in both
READMEs. Dated filenames so caches never serve a stale render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Replaces the 29MB subagent-demo.gif reference with a recaptured 6s loop:
three haiku Explore agents pop as floating windows (25fps through the pop,
bayer dither, 1080px) on the new clean header. Adds the annotated dashboard
tour screenshot below the feature bullets in both READMEs. Old GIF file kept
on disk so external hotlinks stay alive; new files use dated names so caches
can never serve a stale render.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The default desktop header right cluster is now: WS, CPU, MEM, File Viewer
folder button, 5H/7D plan-usage chips, gear. The token-count chip and the
lifecycle-log document button default OFF (both still honor stored prefs),
and the File Viewer button defaults ON (phones keep hiding it via mobile.css).
Templates ship the hidden/shown state so nothing flashes before settings load.
Capture scripts seed showTokenCount:false so screenshots match regardless of
server defaults.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Updating must never silently loosen security. The update path already never
rewrites service files; this covers the remaining gap, re-running the full
installer over an existing setup:
- read_existing_binding() parses the current systemd unit or launchd plist
(a pre-1.8 service without our env lines counts as loopback).
- The network-access prompt defaults to the CURRENT setup instead of the
network default, shows what that setup is, and Enter keeps it, including
a custom non-loopback host and the existing password.
- Non-interactive re-installs adopt the existing binding wholesale.
- The update path's closing security notice now reflects the service's
actual binding instead of the generic loopback text.
Round-trip escaping tested for both formats (quotes, backslashes, XML
specials) plus the legacy-unit, preserve, and Enter-keeps flows.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The loopback-only default was safe but left most installs unreachable from
the devices people actually use. The installer now asks at the end of setup:
1) Any device on your network (0.0.0.0), the default. Prompts for a
dashboard password (confirmed twice); skipping it requires an explicit
confirmation and prints a big red warning as the final output.
2) This machine only (127.0.0.1), the safer option, for tunnel/Tailscale
setups.
The choice flows into the systemd unit, the launchd plist (values escaped
for both formats), the run-now exec path, and the printed URLs (LAN IP
detection included). Non-interactive installs keep the safe loopback
default unless CODEMAN_HOST is preset; the server binary's own default
binding is unchanged.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>