Follow-up to #285. Violet sits close to the terminal's own dim foreground,
so the arcs lost contrast exactly where they cross text, which is most of
their length. Blue reads at a glance on the dark skins and on the light
ones.
Colour still comes from a token every skin block already defines and tunes
for its own background (--session-blue instead of --session-purple), so it
stays one rule for all seven skins with no per-skin override, and the two
blues are not even the same: --session-blue is per palette while the
subagent rule hardcodes #3b82f6.
Hue no longer separates this layer from the subagent lines, so the
separation now rests entirely on shape (a lineage arc hangs under the strip
and never reaches a window), weight and dash pattern. Noted in the rule.
CSS only: no geometry, no markup, no settings.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The lines that join a tab to the workers its codeman skill spawned were
drawn with numbers tuned against two tabs sitting side by side, and they
degraded in exactly the two situations the feature is actually used in.
1. A spawned worker is appended to the END of the strip, so the real span
between a lead and its worker is 800-1500px. With the dip clamped at
44px that is a 33px sag: the arc reads as a straight line drawn across
the terminal instead of a bracket hanging under the strip. The dip now
grows at 0.085/px and clamps at 104.
2. When the desktop strip wraps (tabs-two-rows / tabs-auto-wrap), a parent
on row 1 and its child on row 2 are ~14px apart, and the cross-row
branch drew parent-bottom to child-TOP: a flat line hidden inside the
row gap, with siblings overprinting each other. Both ends now anchor on
the tab BOTTOM with the control points below the LOWER row, so a wrapped
pair gets the same bracket a flat strip gets. That deletes the branch:
one shape covers both.
Visibility, at 1:1 rather than in a zoomed mockup: 2 -> 2.5px stroke,
4 4 -> 5 5 dashes (lineage-flow moves with them, -16 -> -20), opacity
.55 -> .72, and a second wider glow so the contrast comes from the halo
rather than from more weight, keeping the line under the subagent lines'
3px. A working child is bright (.95) outside the reduced-motion block, so
turning motion off no longer also dims every worker's arc. Sibling nesting
6 -> 8px and the direction dot 3 -> 3.5px to match the heavier stroke.
Verified at 1:1 in a harness driving the real styles.css and the real
computeLineagePath over three layouts (adjacent workers, workers at the
far end of a full strip, wrapped two-row strip) on a dark and a light
skin. test/session-lineage-lines.test.ts pins both regressions.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs in the File Viewer's media player, both reproduced in a real
browser against an 18MB mp4 before and after the fix.
1. Closing the preview left the video playing. closeFilePreview() only
dropped the overlay's `visible` class, which is display:none and
nothing else, so the audio kept going with no visible player to pause.
Detaching the element is not a fix either: a detached HTMLMediaElement
plays on until it is garbage collected. _stopFilePreviewMedia() now
pauses, drops src and load()s every media element (also on re-open,
where overwriting innerHTML had the same effect), which additionally
aborts the in-flight download.
2. The scrub bar was inert. file-raw read the whole file and answered
200 with no Accept-Ranges, so Chrome reported video.seekable as
[0, 0] and silently reverted `currentTime = x`; Safari refuses to
start such media at all. Raw bodies are now streamed and range-aware:
Accept-Ranges: bytes on every response, 206 + Content-Range for a
Range request, 416 for one past EOF, and a malformed spec ignored
(200) per RFC 9110. Parsing is pure in src/web/http-range.ts.
Measured on tmp/codeman-crt-v5-66s.mp4 (18MB, 66.6s):
before seekable [0, 0] seek to 56.6s reverted to 3.9s close: still playing
after seekable [0, 66.56] seek to 56.6s landed at 60.2s close: paused, NETWORK_EMPTY
Range slices are byte-identical to `dd`, the full-file path is
byte-identical to the file, and the SVG octet-stream/attachment
hardening and the 50MB cap are unchanged (the cap is still checked
before the range).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
PR #282 added pi across the prominent surfaces but left the enumerations
that read as exhaustive: the env-prefix allowlist (missing PI_*), the
external-CLI list for stop/blocked, cron's agent types (also missing
antigravity), the narrow-strip mode list, and the claude-only caveats in the
cron and Read My Mind guides. Both READMEs and the four affected docs now agree
with the schema.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Heal a stalled SSE stream: the server's :keepalive comment becomes a named
sse:heartbeat event (comments are invisible to EventSource by spec), and the
client gains a staleness watchdog that forces a reconnect after three missed
beats. Also applies a confirmed rename locally instead of waiting on SSE.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.
The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.
Server:
- `sse:heartbeat` under a new Transport category in the event registry
(155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
The write stays per-client rather than going through `broadcast()`: the frame
carries no session data, so it needs no multi-user owner routing.
Client:
- `computeSseStale()` in constants.js, a pure policy beside
`computeConnectionLossUi`. Stale only when the transport believes it is
`connected`, the device is online, and no frame has arrived for 45s (three
missed heartbeats). The `connected`-only guard is also the loop breaker: a
forced reconnect leaves that state immediately, so the watchdog cannot re-fire
while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
`_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
from one place instead of three that can drift. The heartbeat's own listener
is a no-op that exists only to be registered, since `EventSource` drops named
events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
at the top of `connectSSE()` and nowhere else (its only teardown path).
Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
already rebuilds from the server. `visibilitychange` -> visible checks too,
riding the existing listener, since a background tab's timers are throttled
and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
delays heartbeats, the failure mode is "silently reconnects every 45s", which
is undebuggable from a field report without it.
Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).
Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.
Event names are part of the stable API contract, so this is a MINOR bump.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.
Pi is a different shape of CLI from the other four, and three decisions
follow from that:
- It has NO permission prompts and no sandbox, so there is no
--dangerously-skip-permissions analog and none was invented. The
privilege-shaped knob is the tri-state approveProjectTrust, which makes
pi load and EXECUTE repo-local .pi/extensions TypeScript and install
missing project packages. clampExternalCliBypassForOwner() therefore
puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
--no-approve even when no config was sent, because pi's own default is
a prompt the session user could answer themselves. That helper had zero
test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
context, so admitting them would widen the allowlist for every mode at
once. Auth goes through pi's /login or the server's own environment.
--api-key is deliberately never wired: it would put a provider secret on
the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
mode is runtime-switchable via /settings; that flip was measured to put
the pane into the alt screen, which the strip would have corrupted.
pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.
Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).
Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
App Settings opened on a System section that mixed the two things worth seeing
immediately (what this install runs, whether a newer release is waiting) with
three groups nobody sets twice (CLAUDE.md template path, default working
directory, image watcher, Cloudflare tunnel).
Split in two. **Updates** is now the first section and carries only the current
version and the update action, so the modal opens on it and the second thing in
reach is Terminal & Input, where Local Echo lives. **System** keeps Paths,
Automation and Remote access and tails the document, last in the rail.
Also fixes the admin-ui load-order test, which broke on this branch: it located
the modules with a bare `indexOf('session-ui.js')`, and the modal markup now
cites those modules in comments well above the script tags, so it was comparing
a comment against a `<script src>`. It matches the script tag itself now.
A docs pass landed in this worktree while the preview was up (a respawn loop on
the throwaway session it was serving), and it is the documentation this work
needed, so it is reviewed and kept rather than thrown away.
- docs/architecture-invariants.md gains a "Settings surface" section: the one
`:is()` scope and why the id-only list preserves specificity, the anatomy,
the two meanings of the rail, the deliberate two sizes, the phone strip, the
Add Case adapter, the flex-summary chevron trap, the Respawn ordering, the
retired tab chrome, and the live preview's clone-the-chip-icon rule.
- Settings paths are repointed everywhere they moved: Display -> Header &
Panels (header buttons, cron, multi-monitor, response viewer, file viewer),
Settings -> App Settings -> System -> Updates, Panels -> Header & Panels ->
Cross-session features (Read My Mind), Display -> Terminal & Input (gesture
control), Claude Model -> Models -> New Claude sessions.
- Stale counts refreshed (route modules, frontend modules, type files, config
files) and the typecheck script named.
- browser-testing-guide gains the three modal ids and the `set-*` selectors.
- The styles.css block comment covers all three modals.
Two claims it got wrong are corrected here: an external-CLI session opens
Session Options on the Session tab (`switchOptionsTab('context')`), not
Summary - measured in the browser - and the Cron toggle lives under Header &
Panels -> Scheduling, with no "Header Displays" step under it any more.
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.
Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.
Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.
- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
guard as the terminal socket, plus caps on concurrency, stream length and
frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
Claude, then a configured Deepgram key, then the browser
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The footer buttons shipped with class="btn btn-secondary/primary", but no
.btn or .btn-secondary rule exists in this codebase, so all four rendered
as unstyled UA buttons. Moved them to the btn-toolbar convention every
other modal footer uses, with a scoped flex-row footer rule (btn-toolbar
is display:flex, block-level) mirroring the runSummaryModal footer.
Send's accent needs a (0,4,0) re-assert: the skin block's bare
.btn-toolbar rule is (0,2,1) under html:not([data-skin="og"]) and beats
.btn-toolbar.btn-primary (0,2,0), the same specificity trap CLAUDE.md
documents for mobile.css. Scoped to this modal; the repo-wide greying of
btn-primary on non-OG skins is pre-existing and left as a design call.
The empty-result copy now points at the steer note sitting right below
it ("Add a steer note and Rethink to try again"), zh-CN updated.
Verified with the steer E2E (still green) plus desktop, phone (390px),
and error-phase screenshots; static guards extended to pin the footer
convention and the accent re-assert.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds the optional free-text steer note to the Read My Mind modal: a
dashed input under the suggestions ("no, I meant the mobile bug") that
rides along as `steer` on every Rethink. The API already accepted it;
this wires the frontend end of the contract.
- Shown whenever Rethink is live (ready AND empty-result phases),
hidden only while a prediction runs; typed text survives re-runs.
- Enter in the field triggers Rethink, mirroring the prompt field's
Enter-to-send; a fresh open clears it with the rethink memory.
- Trimmed and capped to the schema's 2000 chars on the way out; a
plain open still sends an empty body (neither steer nor rejected).
- zh-CN strings for the placeholder and aria-label, phone-sized
touch target in mobile.css, static guards in the phase-3 test.
Verified with a browser E2E against a live dev server (stubbed predict
endpoint): payload contents, phase visibility, Enter wiring, and
reset-on-reopen all asserted with real keystrokes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The Read My Mind modal grows up and reaches phones:
- Alternate suggestions (the predictor's verify/redirect kinds) now render
as tappable rows below the main field. Tapping one swaps it into the
editable field; the edit you were making folds back into the row you
leave, so toggling between alternates never loses typing. Rethink now
records the WHOLE shown set (main + alternates) as rejected.
- Phones get a 🧠 key on the keyboard accessory bar (both simple and
extended layouts), gated on the same synced readMyMindEnabled setting
via an rmm-enabled marker class on the BAR element: setMode() rebuilds
the buttons' innerHTML, so per-key state would be wiped. Synced at init
and re-synced by applyHeaderVisibilitySettings() on every settings
apply, so a live toggle needs no reload. The header button stays off
phones.
- On phones the modal renders as a small dialog (mirrors modal-sm) instead
of the full-screen default, with wrap-friendly finger-sized footer
buttons. Not modal-sm itself: that caps desktop width at 340px and this
modal wants 560px there.
- On touch devices the ready/swap paths no longer focus the field, so the
OS keyboard does not pop over the alternates that just rendered.
- New static guard test/readmymind-phone-key.test.ts pins the dual-template
key, the marker-class gating, the phone-hidden header button, the
small-dialog phone modal, and the no-innerHTML discipline.
Part 2 of phase 3 (rethink steering, the free-text steer note) is next;
the API already accepts steer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two home-screen reports from @jordan8037310, both about history that is
present but unreachable.
#260 — "Resume Conversation" rendered 4 rows, then a button that appended
every remaining row into a `max-height: 240px` box, so 35 conversations
landed in a four-row scroll well with no ordering or filtering. Rendering
now goes through `_renderHistoryList()` over a cached corpus: 10 rows to
start, Show more/Show less that grows and shrinks the box (the height cap
is class-driven, `.history-list.expanded`), plus a filter box (name,
folder, #case label, prompts), a sort control (recent / name / folder,
pinned rows still first) and a shown-of-total count. A filter implies
expansion, so every match is visible, and the whole header hides as one
unit while a federated search is active. The A-Z sort keys off the same
string the row renders, since most rows are transcript-backed and carry
no session name at all.
#261 — the search box could not match a past project by folder name:
`harvestSources()` built its session corpus from the live in-memory map,
while past sessions come from `/api/sessions/unified` (lifecycle log +
transcript scan). Folding that scan into the request path would have cost
the search its no-filesystem-reads property, so the corpus arrives via a
bounded snapshot instead: `session-history-index.ts` is published as a
side effect of `/api/sessions/unified` (the home screen fetches it on
open, which is the same screen the search box lives on) and rebuilt
fire-and-forget, single-flight and TTL-guarded when a search finds it
stale. A result for a closed session now resumes the conversation rather
than selecting a tab that no longer exists, and is badged RESUME.
The snapshot is stored unscoped with a per-row owner and re-filtered
through canAccessOwned() on read, so multi-user sees exactly what
/api/sessions/unified exposes: own sessions only, host-wide transcript
history admin-only. Live rows are harvested first and win the dedupe.
Verified end-to-end against a real instance with 60 past sessions: cold
process answers its first search without history and its second with it;
folder-name queries return resume targets; clicking one posts the right
resumeSessionId + workingDir.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an exact-key tier (ALLOWED_ENV_KEYS) beside ALLOWED_ENV_PREFIXES in
schemas.ts, admitting CLAUDE_CONFIG_DIR so a case can run on a separate
Claude subscription (client-billed accounts). Exact match only: other
CLAUDE_* keys and near-misses like CLAUDE_CONFIG_DIR_EXTRA stay rejected,
blocked keys stay blocked. The key also survives getEnvOverridesForPersist()
(a path, not a secret; dropping it would silently switch a rebuilt session
back to the default account after a reboot).
Docs cover the transcript caveat: a relocated config dir writes transcripts
outside ~/.claude/projects, so response viewer / subagent windows /
ultracode / Read My Mind go blind for that session unless projects is
symlinked back into the shared tree.
Design and spec contributed by @jordan8037310 in #255. Closes#255.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
docs/readmymind.md covers phase 1 as a user guide: how to enable the synced
readMyMindEnabled setting via the API (no UI checkbox until phase 2), exactly
what is and is not captured, the hooks dependency (Docker bridge / remote-SSH
caveats), storage and wipe paths, curl examples for the three endpoints, the
agent-skill ground rules, and a troubleshooting table. Cross-linked from the
CLAUDE.md Key Patterns entry and the api-reference section.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).
- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
(tool_result-only entries stay silent); capture wiring in server.ts is
claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing
@DodgyBadger's proposal in #236: clone a public repository straight into
codeman-cases/<name> and register it as a normal local case.
POST /api/cases/clone is synchronous by design (request held open, bounded
by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation
surface. Success broadcasts the usual case:created event, so the case
still appears when a proxy idle-timeout kills the request mid-clone.
POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can
say, while the user is still typing, whether the URL is cloneable without
credentials, what its default branch is, and which branches/tags exist.
Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env,
ls-remote parse, stderr classification) and a thin IO half, so every
security decision is unit-testable without spawning anything:
- `<name>::<payload>` transports are refused as a family, not by name:
ext:: is the famous one, but any of them dispatches to git-remote-<name>
and turns a clone into arbitrary command execution.
- A leading `-` is refused AND every spawn puts `--` before the operands.
Either alone is one edit away from being a hole.
- argv arrays, never a shell. URLs carrying user:password@ are refused.
- gitNonInteractiveEnv() closes all four ways git can block on a prompt
with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM).
HOME/PATH stay inherited, so a user's own credential helper or ssh agent
keeps working; Codeman itself collects and stores nothing.
- The timeout signals the process GROUP, since clone fans out into
git-remote-https/index-pack children that outlive a signal to the parent.
- Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs
each) and a global 2-op pool, so N large clones cannot exhaust the host.
Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks
are merged into whatever .claude/settings.local.json the repo shipped, and
a repo that ships its own Claude settings is reported back as a warning
(those hooks run locally as soon as a session starts there). A failed clone
removes only the directory the attempt created, and refuses a pre-existing
destination outright, so it can never squat on a case name.
Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only
inside the caller's own case space. Local-path/file:// sources are the
exception and stay admin-only there.
UI: live verdict under the URL field, case name filled from the parsed repo
until the user types their own, branch/tag as a datalist of the remote's
real refs, optional shallow clone, and a Brain picker (installed CLIs only)
that points the Run button at the chosen agent. Starting a session stays
opt-in. The tab hides itself when the server reports no git.
Tests: the pure half exhaustively (every refusal has a case), plus real git
against a real local bare repo for clone/ref/timeout/cleanup, and a
route-level suite with unmocked fs that clones through the endpoint.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One switch now governs the whole feature: with approvalsInboxEnabled off
(the default), sendPushNotifications strips the actions and approvalId
from permission push payloads, so the buttons no longer render at all
(pre-inbox they rendered and did nothing). The page-side action relay is
gated the same way for stale notifications sent before the toggle
flipped. Only the store and answer endpoints keep running, so enabling
the toggle surfaces anything already pending immediately.
sendPushNotifications is async now (cached settings read); all call
sites were already fire-and-forget. Covered by three new payload tests
alongside the existing hostTitle suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Owner decision: every Approvals Inbox UI surface (header bell, drawer,
phone overview answer strips, reload seeding) now requires enabling
approvalsInboxEnabled in App Settings -> Panels; only an explicit true
turns it on. The store, endpoints, and push Approve/Deny actions keep
running regardless (the push buttons are already opt-in per subscription).
Also replaces em-dashes with plain punctuation across the newly authored
comments, docs, and strings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Permission dialogs, AskUserQuestion questions and idle prompts from every
session now land in a server-side inbox (web/approval-inbox.ts, one item per
session, claude-mode only) and are answerable in place: a header bell + drawer
on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and
working push Approve/Deny buttons (previously dead ends, now answered straight
from sw.js with no tab open). Pending alerts survive reloads because the
frontend seeds from GET /api/approvals on init.
Answering sends the digit / Esc / prompt text through the existing tmux input
path; option digits are accepted only when they match options parsed from the
captured pane frame, and the answer path re-captures the pane first so a
dialog that already left the screen refuses with 409 instead of typing into
the composer. New elicitation_complete / elicitation_response hook matchers
resolve question items the moment they are answered in the terminal;
refreshStaleCodemanHooks heals existing cases.
Verified end-to-end against a live claude session: a real AskUserQuestion
dialog parsed into 5 option buttons and was answered from the drawer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude Code v2.1.224+ gives sessions ListAgents/SendMessage and a per-session
inbox socket. Codeman's claude workers are ordinary local Claude Code sessions,
so the agent skill now teaches task delivery and result collection over
messaging where available (multi-line exactly-once messages, mid-turn steering,
latched replies), with the HTTP primitives keeping spawn, readiness,
synchronization, liveness and delete, and a bounded fallback to the HTTP
recipes whenever the feature is absent. All mechanics verified live against
claude-cli 2.1.226.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
README.zh-CN.md taught a recipe that cannot work: its programmatic-input example had no
trailing `\r`, so Enter was never sent and the prompt sat unsubmitted forever, and its
read step used `/output`, whose `textOutput` is always empty for interactive tmux-backed
sessions. A reader following the Chinese README walked into both of the silent failures
the English one warns about. Its agent/automation section is now brought in line with
README.md: the `\r` rule and every example that needs it, and the correct read path.
CLAUDE.md's "Single-line prompts only" gotcha described the newline restriction but
never mentioned that input must end with `\r` or Enter is never sent, which is the most
common silent failure when driving the API.
docs/agent-control-plan.md asserted as still-open several things that shipped in 1.14.1
and 1.14.2 (the wait endpoints, the packaged skill, the install CLI, agentSkillEnabled).
The status header and the stale bullets now match reality; the historical design content
is untouched, since the document is a record.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Independent post-build review found three gaps, all one family: input that
changes the composer without a prediction leaves the DISPLAYED cursor stale
for one RTT, and anchoring a new run on it painted ghosts one cell off
(blank-neutral, so they lived out the full TTL: "tehh" on
backspace-then-retype, exactly on the links the feature targets).
Fix: the addon now HOLDS new predictions after any such edit (backspace with
nothing outstanding = deleting echoed text, clearPredictions, and now also
IME/plain-paste 'text' commits, which the hook clears like 'clear') until
the next PARSED write releases the hold. The inline predictChar reconcile
deliberately does not count: only the emitter pass or the public
reconcile() is the display-caught-up contract. Worst case is exactly one
unpredicted keystroke, whose own echo releases the hold. Also patched the
one bypass path the PR had missed: _handleCjkInput now clears predictions
like insertTerminalText and the other bypass sends.
Package suite 230, vm gating 85, E2E 10/10 all green after the change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ci.yml runs the xterm-zerolag-input suite (Layers 1-3) after the root
npm ci (workspaces hoisting; no separate install). CLAUDE.md and
architecture-invariants.md rewrite the codex echo story: predictive
write-through with the wire-neutrality, separate-bundle, composer-gate,
baseY and blank-neutral invariants spelled out; the single-source section
now covers both vendor bundles and why their entry points differ.
Changeset: minor for aicodeman + xterm-zerolag-input.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Recorder (scripts/dev/record-codex-frames.mjs) captures real codex 0.147
TUI output through the production pipeline (tmux status-off + the codex-mode
full strip from session.ts) into JSONL fixtures with keystroke injection
points; analyzer replays them through @xterm/headless for the measurements
in docs/predictive-echo-plan.md. Composer signature /^> /-style (U+203A),
modal and wrapped rows correctly rejected, echo is unstyled default-fg,
tmux delivers echo as minimal in-place deltas.
Package: types.ts gains optional cursorX/cursorY, getCell, onWriteParsed,
onResize (all additive); prediction-renderer.ts renders per-glyph spans
keyed by prediction seq; @xterm/headless@^6.0.0 devDep for replay tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reported by @DodgyBadger.
#237: the proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which
bounded the ENTIRE exchange rather than the wait for response headers. A dashboard
endpoint doing model inference, and any actively streaming response, both died at
30s as a generic 502 that Codeman never logged, so it read as an intermittent
network error. The timeout now bounds time-to-headers only and is cleared the
moment headers arrive, so a slow endpoint and a long stream both survive. The
default moves to 300s because "the app is thinking" is normal for the dashboards
people proxy; abandoned upstreams are reclaimed by the client-hangup abort rather
than by this value.
A browser that navigates away mid-request now aborts the upstream fetch, guarded
by `writableFinished` for the same reason as `abortOnClientHangUp` in
session-routes: `close` also fires after a completed response and must not abort
anything. Header timeouts are logged as a warning with a sanitized identity
(method plus origin plus path, never the query string, which can carry the
dashboard's tokens), and a client hangup is deliberately not warned since nobody
is listening and it would read as the dashboard being broken.
The WebSocket handshake keeps its own 30s budget
(`CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`), decoupled from the request timeout:
a handshake is connection establishment, and waiting minutes on one only delays
the browser's reconnect logic.
#238: the web-tab guide covered sandboxed dashboards having no cookies, but not
cookie authentication in front of Codeman itself (Cloudflare Access and similar),
where a sandboxed frame's asset and API requests carry no auth cookie, bounce to
the login provider, and leave the embedded app looking unstyled or broken while
trusted mode works. Documented, and the Test button's result now says it probes
server-to-upstream reachability only, not how the page behaves in a sandboxed
frame.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.
Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
Case names resolve through linked-cases.json first, mirroring the
server's resolveCasePath(), so a case linked in from outside
~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
skill is never touched, and a symlinked skill dir is refused (this
repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
ports/config-port.ts, server.ts, session-routes.ts (add-only injection
on Claude session create and quick-start), plus the App Settings toggle.
Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
Shell state does not survive between agent tool calls, and an undefined
is_self exited 127, firing the `||` branch and deleting the caller's own
session with the one guard bypassed. The request now lives inside the
guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
tool call, so the documented resend-identical-request loop stopped being
a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
workers. It returns clean transcript text; the terminal scrape it
replaces returns a wall of TUI repaint noise. Its transcript flush lags
the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
yielded the literal session id "null" and burned the whole readiness
budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
corrected the hooks-config comment that claimed a toggle-off sweep
exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
and that caseName resolves linked cases, so a generic name can land a
worker in a real repo.
Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.
Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.
_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.
Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two ways to keep the server running, split by how long it should last.
`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.
`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.
Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.
The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two defects in the background-task hook scripts.
SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.
The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.
The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.
12 tests fail on unmodified master, e.g.
expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.
1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
terminal and rewrites it from the capture, which is a win when tmux holds
more than the browser, but for a repaint-mode pane that capture is roughly
ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
(escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
rewrite when it is more than one screen short; a refused session's cooldown
goes from 4s to 60s so a hollow pane stops re-fetching megabytes.
2. A false forwarding gate on a Claude session no longer means a dead gesture.
Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
SGR reports, at half a screen of travel per page key. Shift is excluded: it
keeps meaning "local scrollback".
3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
exception and guarded on !== undefined, so one timed-out or PATH-starved
probe at the first Claude session start disabled wheel-forwarding for every
Claude session until the server restarted, which fits a report of breakage on
phone, tablet and laptop at once. Success is still cached for the process
lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
a pure function so the semantics are testable without spawning claude.
4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
scoping: the setting keeps meaning exactly what it says, and fix 2 catches
the case where "local" is empty. The App Settings tooltip now says to leave
it off for Claude/Codex sessions.
5. _logScrollRouting() prints one line per session per distinct decision:
forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
mode, cliVersion, the opt-out state, mouse tracking and local scrollback
depth. #205 ran two rounds of remote guesswork over questions that line
answers directly.
Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.
The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.
Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>