The Read My Mind modal grows up and reaches phones:
- Alternate suggestions (the predictor's verify/redirect kinds) now render
as tappable rows below the main field. Tapping one swaps it into the
editable field; the edit you were making folds back into the row you
leave, so toggling between alternates never loses typing. Rethink now
records the WHOLE shown set (main + alternates) as rejected.
- Phones get a 🧠 key on the keyboard accessory bar (both simple and
extended layouts), gated on the same synced readMyMindEnabled setting
via an rmm-enabled marker class on the BAR element: setMode() rebuilds
the buttons' innerHTML, so per-key state would be wiped. Synced at init
and re-synced by applyHeaderVisibilitySettings() on every settings
apply, so a live toggle needs no reload. The header button stays off
phones.
- On phones the modal renders as a small dialog (mirrors modal-sm) instead
of the full-screen default, with wrap-friendly finger-sized footer
buttons. Not modal-sm itself: that caps desktop width at 340px and this
modal wants 560px there.
- On touch devices the ready/swap paths no longer focus the field, so the
OS keyboard does not pop over the alternates that just rendered.
- New static guard test/readmymind-phone-key.test.ts pins the dual-template
key, the marker-class gating, the phone-hidden header button, the
small-dialog phone modal, and the no-innerHTML discipline.
Part 2 of phase 3 (rethink steering, the free-text steer note) is next;
the API already accepts steer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Round 2 of the #251 review: settingsWriteBlocker covered only
writeHooksConfig and updateCaseModel, while applyStatusLineConfig,
stripCaseEnvKeys, updateCaseEnvVars, refreshStaleCodemanHooks and
ensureCodemanHooks still wrote the same repository-controlled path
unguarded (applyStatusLineConfig was demonstrated writing through a
symlinked settings.local.json).
All seven writers now go through withSafeSettingsWrite(), which runs
the blocker check INSIDE the per-path settings lock and then hands the
writer its claudeDir/settingsPath; none of them touch the settings path
directly anymore. Test pins all seven against a symlinked
settings.local.json at once (link target must stay byte-identical).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Addresses all four findings from the #251 review:
- Scaffolding no longer writes through repository-controlled symlinks.
The guard lives in hooks-config.ts (settingsWriteBlocker) so it also
covers quick-start/docker/ralph writers, not just the clone route:
refuses a symlinked .claude or settings.local.json, a .claude that is
a file, or one resolving outside the case. The clone route surfaces
the refusal as a user-visible warning, and the CLAUDE.md write checks
presence via lstat so a BROKEN repo-shipped symlink counts as present
(existsSync follows links and would have created the outside target).
- Failed-clone cleanup can no longer delete a concurrent winner's tree:
git clones into an attempt-owned temp sibling (.<name>.cloning-<rand>)
which is atomically renamed into place; the loser reports
DESTINATION_EXISTS and only ever removes its own temp dir.
- decodeURIComponent(url.pathname) is guarded: malformed percent-escapes
now come back as BAD_SYNTAX instead of an uncaught URIError 500.
- The git pool's waiter queue is bounded (CODEMAN_MAX_GIT_QUEUE, default
16): overflow answers BUSY immediately (HTTP 429 via RATE_LIMITED),
and queue time counts against the operation's own deadline.
Tests: hostile symlink fixture repo (route level), settingsWriteBlocker
units, concurrent same-destination race, temp-dir leak assertions,
percent-escape rejection, and a fake-git pool-bounds suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an exact-key tier (ALLOWED_ENV_KEYS) beside ALLOWED_ENV_PREFIXES in
schemas.ts, admitting CLAUDE_CONFIG_DIR so a case can run on a separate
Claude subscription (client-billed accounts). Exact match only: other
CLAUDE_* keys and near-misses like CLAUDE_CONFIG_DIR_EXTRA stay rejected,
blocked keys stay blocked. The key also survives getEnvOverridesForPersist()
(a path, not a secret; dropping it would silently switch a rebuilt session
back to the default account after a reboot).
Docs cover the transcript caveat: a relocated config dir writes transcripts
outside ~/.claude/projects, so response viewer / subagent windows /
ultracode / Read My Mind go blind for that session unless projects is
symlinked back into the shared tree.
Design and spec contributed by @jordan8037310 in #255. Closes#255.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The brand "C" got a 44px-wide hit box in the previous commit but was capped at
36px tall by the bar it sits in. The phone header is now 44px, so the one
control that gets you back to the home screen is square at the platform
minimum, and every other header control gains the same 8px.
Redefined as --header-height inside the phone media query rather than as a
literal, so the panels positioned off that token (file browser, project
insights, plan overlays) follow the bar instead of drifting 8px underneath it;
.app's top offset is derived from it for the same reason. The header also stops
top-aligning its children on phones: that read as centred in a 36px bar whose
contents were ~31px, and leaves a visible gap under everything at 44px.
Costs 8px of terminal height on a phone.
Verified on a real isolated instance at 390px: header 44px, button 44x44
spanning the bar, a touch tap at (4,41) - inside the new area, outside the old
one - reaches the home screen, tabs centred, and content still clears the fixed
header. Tablet (48px) and desktop are untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The welcome overlay centers ~560px of content in a ~1400px window, so both
gutters are dead space. The left one now carries the open tabs as a vertical
list (home-sessions.js): one row per live session plus saved web tabs, in TAB
order rather than by urgency, because the row badges are the Alt+1..9 indices.
Clicking a row enters that session.
Working state is deliberately the phone's, exactly: a pulsing green dot ringed
by the same tab-load-spin the tab strip uses while a tab loads, now with a green
halo added on both surfaces so "working" reads identically wherever you see it.
The column is position:absolute so the centered content never moves, which is
why it needs a width gate in two places (HOME_SESSIONS_MIN_WIDTH = 1180 in JS,
a max-width: 1179px media query as the backstop for a resize that outruns the
matchMedia listener). A test pins the two equal. State classification is reused
from mobile-overview.js rather than re-derived, so the two home screens cannot
disagree about what counts as needing you.
Phones keep the mobile overview, and their brand "C" was a 0.85rem inline span,
roughly a 12x13px target on the one control that gets you back to that screen.
It is now a 44px-wide button filling the full header height, with the glyph
scaled to match. 44 is horizontal only: the phone header is pinned to 36px and
clips overflow, so a true 44x44 would mean taking height off the terminal.
Verified end to end against a real isolated instance (own tmux socket + data
dir): 18 browser checks covering render, live update through the tab renderer,
the working dot's animation/glow/ring, row click, the narrow-window gate, the
phone fallback, and a real touch tap on the far corner of the new hit box.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).
- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
(tool_result-only entries stay silent); capture wiring in server.ts is
claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
selectSession() ends with scrollToLastNonEmptyLine(), which parks the viewport
one row ABOVE the bottom for any session whose buffer is taller than the screen
and ends in blank rows, so that is the normal state after a tab switch. Nothing
pinned that a tap there still leaves the keyboard reachable.
The blocker reduced in #173 came back through exactly that gap in #244: a tap
classifier that treats "viewport is scrolled up" as a reason to blur, paired
with touchstart preventDefault cancelling the compatibility click, closes both
routes to focus on the same gesture and strands document.activeElement on
<body> with no way to type. The prompt row is no exception.
Measured on a 390x844 viewport, claude-mode session, dispatched touch gesture:
master leaves focus on textarea.xterm-helper-textarea, PR #244's terminal-ui.js
leaves it on body. Green here, red against that branch.
The test also pins the half that IS correct: SGR coordinates are meaningless
off-bottom, so the tap must send no mouse report.
It has to be a dispatched gesture. Calling the touchend handler directly
bypasses touchstart's preventDefault, which is half of what closes the focus
path, so a direct call reports the right intent and still misses the bug.
test/mobile/keyboard.test.ts: 4 failed | 32 passed (36), against 4 failed |
31 passed (35) without it. Same four pre-existing failures either way.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The working dot is the one glance-state a phone needs: busy tabs now get a
9px pulsing dot with a green glow (idle stays 4px). The glow needs !important
because the skin block's no-halo rule outranks mobile.css.
The simple keyboard accessory bar swaps /clear for Tab (/clear and /compact
stay in the extended bar with their double-tap confirm). The tab action now
flushes locally-buffered prompt text to the PTY before sending \t, so
completion applies to what was just typed instead of an empty composer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing
@DodgyBadger's proposal in #236: clone a public repository straight into
codeman-cases/<name> and register it as a normal local case.
POST /api/cases/clone is synchronous by design (request held open, bounded
by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation
surface. Success broadcasts the usual case:created event, so the case
still appears when a proxy idle-timeout kills the request mid-clone.
POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can
say, while the user is still typing, whether the URL is cloneable without
credentials, what its default branch is, and which branches/tags exist.
Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env,
ls-remote parse, stderr classification) and a thin IO half, so every
security decision is unit-testable without spawning anything:
- `<name>::<payload>` transports are refused as a family, not by name:
ext:: is the famous one, but any of them dispatches to git-remote-<name>
and turns a clone into arbitrary command execution.
- A leading `-` is refused AND every spawn puts `--` before the operands.
Either alone is one edit away from being a hole.
- argv arrays, never a shell. URLs carrying user:password@ are refused.
- gitNonInteractiveEnv() closes all four ways git can block on a prompt
with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM).
HOME/PATH stay inherited, so a user's own credential helper or ssh agent
keeps working; Codeman itself collects and stores nothing.
- The timeout signals the process GROUP, since clone fans out into
git-remote-https/index-pack children that outlive a signal to the parent.
- Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs
each) and a global 2-op pool, so N large clones cannot exhaust the host.
Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks
are merged into whatever .claude/settings.local.json the repo shipped, and
a repo that ships its own Claude settings is reported back as a warning
(those hooks run locally as soon as a session starts there). A failed clone
removes only the directory the attempt created, and refuses a pre-existing
destination outright, so it can never squat on a case name.
Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only
inside the caller's own case space. Local-path/file:// sources are the
exception and stay admin-only there.
UI: live verdict under the URL field, case name filled from the parsed repo
until the user types their own, branch/tag as a datalist of the remote's
real refs, optional shallow clone, and a Brain picker (installed CLIs only)
that points the Run button at the chosen agent. Starting a session stays
opt-in. The tab hides itself when the server reports no git.
Tests: the pure half exhaustively (every refusal has a case), plus real git
against a real local bare repo for clone/ref/timeout/cleanup, and a
route-level suite with unmocked fs that clones through the endpoint.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The picker behind Link Existing's "Browse" and the mobile keyboard's Path key
refused every path with a dot-prefixed segment, so `.github/workflows/ci.yml`
could not be selected and a hidden folder could not be opened at all. It gains
the same `.*` toggle as the File Viewer: default OFF, per-device, and applied to
both the listing and the preview endpoint, which re-resolves the path
independently.
That dotfile filter was quietly doing security work. The picker's roots include
Home, so with every hidden path unreachable the shared blocklist never had to
name the credentials that live in dot-directories. Lifting the filter removes
that accident, so `isSensitivePath` now covers them explicitly: SSH keys at any
depth rather than only under $HOME, GPG keyrings, AWS/GCloud/Azure/Docker/
Kubernetes credentials, npm, Yarn, git, gh, netrc, PyPI, RubyGems, Cargo and
Terraform tokens, .pgpass and .my.cnf, and the Claude and Codeman agent
credentials. `~/.codeman/` and `~/.claude/` stay attachable as trees, since the
publish skill and the review-card loop read from them; only their secret-bearing
members are named.
Everything else still applies with the toggle on: blocked trees, sensitive
files, root confinement, ownership scoping and symlink-escape checks. A hidden
entry whose realpath is a secret is dropped from the listing, and opening it is
refused.
Follows #221
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
One switch now governs the whole feature: with approvalsInboxEnabled off
(the default), sendPushNotifications strips the actions and approvalId
from permission push payloads, so the buttons no longer render at all
(pre-inbox they rendered and did nothing). The page-side action relay is
gated the same way for stale notifications sent before the toggle
flipped. Only the store and answer endpoints keep running, so enabling
the toggle surfaces anything already pending immediately.
sendPushNotifications is async now (cached settings read); all call
sites were already fire-and-forget. Covered by three new payload tests
alongside the existing hostTitle suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A session on a fresh directory sat on Claude's "Quick safety check: Is
this a project you created or one you trust?" dialog until a human
pressed Enter. Reproduced on a new case, then read off the wire:
1.\x1b[C Yes,\x1b[C I\x1b[C trust\x1b[C this\x1b[C folder
tmux repaints a row by writing each word followed by a cursor-forward
escape instead of a space, and Ink colours each word separately, so
`data.includes('trust this folder')` could never match a chunk. The
spaces are not there to strip: they were never sent. The auto-accept has
been dead for every session that hit the dialog.
Match on whitespace-free, ANSI-free, lowercased text instead
(`compactScreenText`), which survives both that repaint style and the
spaced full-screen redraw.
Answering means pressing Enter into a session, so three guards bound it:
- Read the RENDERED SCREEN (capturePaneText), not the chunk. The terminal
buffer is append-only and keeps the dialog in its tail long after it
has been answered, so a retry driven off the buffer would type into a
live session. Direct-PTY sessions, which have no pane, fall back to a
short buffer tail.
- Require a trust phrase AND the dialog's own confirm affordance. One
phrase is not enough, since an agent's transcript can quote it.
- Only look during the first 90s of the pane's life, and cap it at three
attempts. Ink can drop a keystroke while it is still mounting the
widget, which is the other half of why sessions got stuck, but a
dialog that will not clear must not become an Enter loop.
Verified end to end on a fresh case: dialog answered on attempt 1, one
Enter sent in total, session went straight to the composer and answered a
prompt. Before the fix the same flow parked on the dialog indefinitely.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Opening Codeman with nothing reachable (phone off the tailnet, VPN down,
server stopped) rendered a normal-looking UI: the service worker serves the
cached app shell, every /api call fails, and the only tell was an 8px red dot
in the header corner. On a phone that reads as "there are no sessions".
Two surfaces, chosen by whether there is anything worth looking at:
- Full-screen overlay while no server state has loaded this page load. It
names the host, lists the three things to check (network, VPN/Tailscale,
server), counts down to the next retry, and offers "Retry now" plus
"Show cached view" to demote itself to the banner.
- Non-blocking banner once state HAS loaded, so a mid-session drop leaves the
terminal scrollback readable.
A 2.5s grace keeps a COM deploy (SSE is back in ~200ms) from flashing the
banner every release; navigator.onLine === false skips the grace, since the
device saying "no network" is never a blip. Retry re-arms the terminal
WebSocket as well as SSE: planWsReconnect can give up outright, and the SSE
backoff caps at 30s, so waiting it out is not always an option.
The decision is pure (computeConnectionLossUi in constants.js, unit-tested in
a node VM like the WS reconnect policy); app.js only writes the DOM.
Owner decision: every Approvals Inbox UI surface (header bell, drawer,
phone overview answer strips, reload seeding) now requires enabling
approvalsInboxEnabled in App Settings -> Panels; only an explicit true
turns it on. The store, endpoints, and push Approve/Deny actions keep
running regardless (the push buttons are already opt-in per subscription).
Also replaces em-dashes with plain punctuation across the newly authored
comments, docs, and strings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The tree endpoint has accepted `showHidden=true` since it was written; the
panel hardcoded `showHidden=false`, so dot-prefixed entries were unreachable
from the File Viewer and opening one meant guessing its path.
Adds a `.*` toggle to the panel header. It re-fetches instead of re-rendering
the cached tree (the filtering is server-side), preserves the expanded
directories so toggling does not collapse the tree, and persists per-device to
its own `codeman:fileBrowserShowHidden` key. That key is deliberately not part
of the app-settings object, which `saveAppSettings()` rebuilds from the
settings-modal DOM and would drop it on the next save.
Default is OFF, so an untouched install behaves exactly as before.
Closes#221
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every working Claude session reported `status: "idle"` about two seconds
into its turn. Measured on live workers: two sessions mid-tool-call at 13
and 17 minutes both read `idle` while their panes showed
`✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`.
Two things had drifted apart:
1. The working indicator changed. Claude animates the glyph through
`· ✢ ✳ ∗ ✻ ✽` and randomizes the gerund per turn, so neither
SPINNER_PATTERN (braille, no longer drawn) nor the keyword list
(Thinking/Writing/Reading/Running) matches a turn anymore.
2. A `❯` sighting is not the end of a turn. Claude redraws the composer
roughly once a second all the way through one, and that redraw armed
the "2s later, call it idle" timer.
Matching the new status line in the STREAM does not fix it either: tmux
ships partial repaints, so the complete line reached the PTY about once
every 20 seconds while the `❯` arrived every second.
So the decision moves off the stream:
- An unbroken run of repaints marks a turn as started. Sampled once a
second for 12s over six live sessions, the two working ones produced
output in 12/12 windows and the four idle ones in 0/12. Pure helpers in
session-activity.ts carry the thresholds.
- Idle now needs the pane to go quiet AND the screen to agree.
`_confirmIdle()` asks tmux what is rendered (new `capturePaneText()`,
one plain `capture-pane`, floored at 1.5s per session and only ever at
a transition) and re-checks every 5s while the screen still shows work.
A turn can sit silent for tens of seconds inside one tool call, so
silence alone proves nothing.
- The same screen check vetoes keystroke echo, which is a steady stream
of repaints too but is not work.
CLAUDE_WORKING_LINE_PATTERN matches the `… (elapsed)` shape rather than
the glyph, because the FINISHED line (`✻ Cooked for 2m 49s`) carries the
same glyph and would otherwise pin a session at working forever.
Claude mode only. An external CLI has no `❯`, so nothing would arm the
confirmation and such a session would latch busy.
respawn-patterns.hasWorkingPattern() had the same blind spot (its gerund
list cannot see "Actualizing"), so it takes the pattern as an extra
signal. That can only make respawn less eager, never more.
Idle now lands about 3 to 5 seconds after a turn ends instead of 2
seconds into one. Verified end to end against a live worker, sampled
against the CLI's own "esc to interrupt" footer as independent ground
truth: busy for all 25s of a turn, idle 3s after it ended.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Permission dialogs, AskUserQuestion questions and idle prompts from every
session now land in a server-side inbox (web/approval-inbox.ts, one item per
session, claude-mode only) and are answerable in place: a header bell + drawer
on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and
working push Approve/Deny buttons (previously dead ends, now answered straight
from sw.js with no tab open). Pending alerts survive reloads because the
frontend seeds from GET /api/approvals on init.
Answering sends the digit / Esc / prompt text through the existing tmux input
path; option digits are accepted only when they match options parsed from the
captured pane frame, and the answer path re-captures the pane first so a
dialog that already left the screen refuses with 409 instead of typing into
the composer. New elicitation_complete / elicitation_response hook matchers
resolve question items the moment they are answered in the terminal;
refreshStaleCodemanHooks heals existing cases.
Verified end-to-end against a live claude session: a real AskUserQuestion
dialog parsed into 5 option buttons and was answered from the drawer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Follow-ups to #241 (thanks @Lint111), from an independent review of that PR. The
script is a real fix for a real gap; these are the four defects the review found,
each reproduced before and after.
1. A wrong-but-fresh output was never repaired. The zerolag bundle is finished by a
SECOND step (the alias append), so anything landing between esbuild and the
append is permanent: the file looks complete, carries a current mtime, and the
mtime-only cache reports "up to date" forever while the suite dies on
`LocalEchoOverlay is not defined`. Reproduced by replaying #241's own two
commits: running the first and then pulling the second kept the broken bundle.
Fixed twice over, because the two halves address different cases. Builds now go
to a temp file and `renameSync` into place, so this script can never publish a
half-written output (that also covers an interrupted esbuild or copy, and two
concurrent runs). And `isFresh` verifies the bundle actually contains its alias
tail, which is what repairs a file an EARLIER version already poisoned; a rename
alone cannot fix what is already on disk.
2. Freshness compared against the entry file only, but esbuild bundles its four
siblings too, so editing overlay-renderer.ts left the suite testing a stale
overlay while reporting "up to date". Editing those siblings is exactly the
single-source workflow CLAUDE.md mandates. It now stats every `.ts` in the
package source dir. A full rebuild is ~2s, so the cache was not buying much.
3. `execFileSync('npx', ...)` passed no cwd, unlike scripts/build.mjs, so a run from
another directory missed the repo's pinned esbuild and would fetch an unpinned
one from the registry. Both calls now pass `cwd: ROOT`.
4. Every invocation in test/mobile/README.md was a bare `npx vitest`, which skips
the `pretest:mobile` hook npm only fires for `npm run test:mobile`, so the
documented commands all bypassed the fix. Rewritten, with a note on why.
Also: an esbuild failure printed a raw stack; it now names the asset and its input,
matching the missing-input message. And the header comment no longer implies the
vendor dir is always empty: scripts/postinstall.js already writes these same seven
outputs, so what this script adds is freshness and independence from install time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three gaps found while auditing the agent skill.
`codeman skill install` / `uninstall` had no tests at all, including the linked-case
resolution that shipped in 1.14.2 with nothing guarding it. Covered now: global target
resolution, `--case` resolving through linked-cases.json, `--case` falling back to the
cases dir for an unlinked name, a missing or malformed registry degrading to the
fallback instead of throwing, and a nonexistent case being rejected. `resolveSkillTarget`
called `process.exit(1)` for a missing case, which would have killed the test runner, so
the pure resolution is split out and exported; CLI behavior is unchanged.
The `POST /api/sessions` injection call site was never exercised, because the shared
route mock hardcoded the gate off. The mock's gate is overridable per test now (default
still off, since other tests rely on that), and there is coverage that the path injects
when the setting is on, does not when it is off, and is claude-mode gated.
Nothing guarded skills/codeman/reference/endpoints.md against drifting from the routes
it documents, which is how it drifted in the first place. A static guard parses the
endpoints out of the markdown and asserts each is really registered, tolerating the
/api/v1 alias and path params.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Independent post-build review found three gaps, all one family: input that
changes the composer without a prediction leaves the DISPLAYED cursor stale
for one RTT, and anchoring a new run on it painted ghosts one cell off
(blank-neutral, so they lived out the full TTL: "tehh" on
backspace-then-retype, exactly on the links the feature targets).
Fix: the addon now HOLDS new predictions after any such edit (backspace with
nothing outstanding = deleting echoed text, clearPredictions, and now also
IME/plain-paste 'text' commits, which the hook clears like 'clear') until
the next PARSED write releases the hold. The inline predictChar reconcile
deliberately does not count: only the emitter pass or the public
reconcile() is the display-caught-up contract. Worst case is exactly one
unpredicted keystroke, whose own echo releases the hold. Also patched the
one bypass path the PR had missed: _handleCjkInput now clears predictions
like insertTerminalText and the other bypass sends.
Package suite 230, vm gating 85, E2E 10/10 all green after the change.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Out-of-process lab server (VITEST markers stripped so tmux/codex are real),
CODEMAN_INSTANCE=codexlab on port 3222, throwaway CODEX_HOME with a fake
key. Ten scenarios: bundle smoke, predict+converge typing, the #218 arrow
retest (submitted text exact), the #222 live picker, the #219 paste order,
the #220 wrap, the trust-modal ghost eliminator, the localEchoEnabled kill
switch, the end-to-end byte-identity trace (predictor active vs null), and
a display-delayed 300ms-RTT run pinning instant spans with exact pixel
geometry plus arrow-edit correctness under lag.
Live-TUI hardening learned the hard way: codex Ctrl+U kills only to line
start (End first), a fake-key submit leaves a Reconnecting loop that can
kill codex seconds later (retry-cancel + composer stability probe; the
submitting scenario runs after all composer-state ones), and typing must
wait for the predictWhen gate itself, not merely a rendered composer.
CI-excluded like the other Playwright suites; skips cleanly when codex is
not installed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
terminal-ui.js: _localEchoPolicy ('buffer'|'predict'|'off') computed at the
end of _updateLocalEchoState with _localEchoEnabled keeping its exact 1.12.2
values; _predictHookOnData called as a plain statement between the buffer
block and Normal Mode (visual-only, try/catch, never returns, never touches
_pendingInput); classifyPredictInput + isCodexComposerRow (baseY-based,
measured /^> /-signature gate) on CodemanTerminalInput; construction beside
the LocalEchoOverlay from the separate bundle with graceful absence;
insertTerminalText/clearTerminalInput/setFontSize/applyTerminalSkin clear or
refresh predictions. app.js: fields + tab-switch and SSE-reconnect clears.
voice-input '\r' branch and keyboard-accessory sendKey clear predictions
(both bypass onData). sendEnterKey needs no change: codex falls through to
the immediate-flush branch.
Layer 4 vm tests: classify truth table (20 cases), composer-row gate incl.
the baseY pin, policy matrix with the 1.12.2 invariants untouched, wire
neutrality + throwing-predictor pins. Stale mobile keyboard codex-buffering
tests repointed at claude; new codex twin asserts write-through streaming,
prediction spans and TTL self-heal.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Reported by @DodgyBadger.
#237: the proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which
bounded the ENTIRE exchange rather than the wait for response headers. A dashboard
endpoint doing model inference, and any actively streaming response, both died at
30s as a generic 502 that Codeman never logged, so it read as an intermittent
network error. The timeout now bounds time-to-headers only and is cleared the
moment headers arrive, so a slow endpoint and a long stream both survive. The
default moves to 300s because "the app is thinking" is normal for the dashboards
people proxy; abandoned upstreams are reclaimed by the client-hangup abort rather
than by this value.
A browser that navigates away mid-request now aborts the upstream fetch, guarded
by `writableFinished` for the same reason as `abortOnClientHangUp` in
session-routes: `close` also fires after a completed response and must not abort
anything. Header timeouts are logged as a warning with a sanitized identity
(method plus origin plus path, never the query string, which can carry the
dashboard's tokens), and a client hangup is deliberately not warned since nobody
is listening and it would read as the dashboard being broken.
The WebSocket handshake keeps its own 30s budget
(`CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`), decoupled from the request timeout:
a handshake is connection establishment, and waiting minutes on one only delays
the browser's reconnect logic.
#238: the web-tab guide covered sandboxed dashboards having no cookies, but not
cookie authentication in front of Codeman itself (Cloudflare Access and similar),
where a sandboxed frame's asset and API requests carry no auth cookie, bounce to
the login provider, and leave the embedded app looking unstyled or broken while
trusted mode works. Documented, and the Test button's result now says it probes
server-to-upstream reachability only, not how the page behaves in a sandboxed
frame.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.
Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
Case names resolve through linked-cases.json first, mirroring the
server's resolveCasePath(), so a case linked in from outside
~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
skill is never touched, and a symlinked skill dir is refused (this
repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
ports/config-port.ts, server.ts, session-routes.ts (add-only injection
on Claude session create and quick-start), plus the App Settings toggle.
Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
Shell state does not survive between agent tool calls, and an undefined
is_self exited 127, firing the `||` branch and deleting the caller's own
session with the one guard bypassed. The request now lives inside the
guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
tool call, so the documented resend-identical-request loop stopped being
a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
workers. It returns clean transcript text; the terminal scrape it
replaces returns a wall of TUI repaint noise. Its transcript flush lags
the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
yielded the literal session id "null" and burned the whole readiness
budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
corrected the hooks-config comment that claimed a toggle-off sweep
exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
and that caseName resolves linked cases, so a generic name can land a
worker in a real repo.
Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.
Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.
_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.
Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two ways to keep the server running, split by how long it should last.
`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.
`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.
Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.
The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two defects in the background-task hook scripts.
SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.
The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.
The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.
12 tests fail on unmodified master, e.g.
expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".
stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.
Two tests, both failing on master:
expected 'export PATH="…' to contain ' 2> "'
expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.
Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.
Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.
Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).