SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three post-merge fixes for the external-CLI response viewer:
- ?context=full blocks now carry role ('user' for prompts, 'assistant'
for response/status/tool). The frontend's loadFullContext() renders
via msg.role, so the roleless blocks lost the "You" badge and every
turn rendered as the agent. kind/label/text are unchanged and the
frontend needs no change.
- normalizeDividerStatusLine() dropped its backtracking regex
(/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
minutes at 10,000), and pane text is agent-controlled with buffers up
to 32MB. Replaced by a linear counter walk with the identical accept
set and captured content, pinned char-for-char against the old regex
by a brute-force corpus test plus a hostile-input regression test
that fails by timeout with the RegExp version (same approach as the
glob-matcher hardening in 68ae9a8).
- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
empty-viewer symptom the transcript branch exists to fix. The list
stays a local duplicate of isExternalCliMode() (importing session.ts
would drag node-pty into the pure module); a new exhaustive parity
test asserts the two mode sets can no longer drift.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.
For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.
Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.
The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.
Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.
compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.
The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.
Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
validateSessionFilePath realpath-resolves the candidate path but compared it
against the raw sessionWorkingDir. When the workspace is itself reached through
a symlink the two sides live in different namespaces, so relative() reports a
spurious `../` and every file in that workspace is judged an escape — reads and
writes in the session are refused wholesale.
That is not an exotic setup: os.tmpdir() hands back a symlinked path on macOS
(/tmp -> /private/tmp), and symlinked project directories and bind-mounted case
paths hit it too.
Resolve both sides and compare canonical to canonical. This only makes the
comparison honest — it does not widen it. The candidate keeps its own realpath,
so a symlink pointing out of the workspace and a ../ traversal are still
refused, and a workspace that cannot be resolved now fails closed.
Three stubs in file-routes.test.ts used a blanket
realpathSync.mockReturnValue(escapeTarget), which answers the same path for the
workspace and the candidate; with both sides resolved that makes an escape look
contained. They now use the input-aware mockImplementation idiom the rest of
that file already uses, so the workspace resolves to itself and only the
candidate escapes. Verified they still bite: removing the confinement check
turns all of them red.
Adds test/route-helpers-symlink-confinement.test.ts, which exercises the
function against a real symlinked workspace on disk and pins the negative cases
(../ escape, symlink-out, missing file) alongside the fix.
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Persist the 'I checked it' state of yellow idle tab alerts across reloads and devices.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled,
default ON) moved from a session-routes-local helper into hooks-config.ts as
applyWorkspaceHooks(workspace, install?), and the claude session-create sites
that bypassed it now go through it: cron job fires (cron-service), legacy
scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research
and planner one-shots. A cron or scheduled run firing in a linked case that
never had an interactive session ran hook-blind (no stop for completion
detection, no tab alert on a blocking dialog).
The shared core also carries the two guards every caller needs: a workspace
that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot
recovery sweep used to resurrect a deleted repo as an empty tree holding only
.claude/settings.local.json), and all errors are swallowed since a create must
never fail on hooks. Route handlers keep resolving the setting through their
ConfigPort and pass it in; non-route callers omit it and the core reads
settings.json itself (absent key or unreadable file = ON).
Two adjacent gates tightened in session-routes:
- the docker quick-start hooks branch excluded the five external CLIs but let
`shell` through, contradicting its own rule that only claude reads .claude
hooks; it is now gated on mode === 'claude'
- the statusLine exporter call in POST /api/sessions got the same
!remote && body.workingDir guard the hooks call got in 499d355 (it mkdirs the
same way, so a remote attach created a junk user@host:session dir locally and
a cwd-fallback create wrote into $HOME)
plan-routes' one-shot deliberately stays out: its workingDir is process.cwd(),
exactly the target 499d355 forbids writing into. restoreMuxSessions stays out
too: the boot sweep already covers recovered workspaces.
Tests: quick-start existing-case install, docker claude-installs/shell-does-not,
and the core directly (default-ON install, OFF add-nothing, OFF still heals a
stale block, malformed file untouched, vanished workspace skipped); the remote
and cwd-fallback regressions now also send statusLineTelemetry:true to pin the
statusLine guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A claude-mode attachRemoteSession create overwrites workingDir with the
user@host:session pseudo-path, which is a RELATIVE path locally — the old
refresh-only call no-op'd on it, but ensureCodemanHooks mkdirs, so it
created a junk local directory. And with workingDir omitted the cwd
fallback reaches the hooks write unvalidated; under installer-created
services cwd is $HOME, so hooks materialized in ~/.claude/settings.local.json.
Both guarded at the applyWorkspaceHooks call site; regression tests prove
the remote attach leaves no junk dir and the no-workingDir create leaves
the server cwd untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A .json/.log/.yaml/code path outside the session workspace was refused as an
unsupported type, and clicking one in the terminal made it worse: text goes to
the log viewer, which spawns `tail -f` and allows only the workspace, /var/log
and ~/logs, so it answered "Path must be within working directory or allowed
log directories" while the same path clicked in the response viewer previewed
fine. Two surfaces, two answers, for a file the session can already cat.
- TEXT_ATTACHMENT_EXTENSIONS IS EDITABLE_EXTENSIONS (config/file-editing.ts),
not a second curated list that would drift from it. The rule reads: if the
viewer would open a file for editing inside the workspace, the same file
outside it can be read. The suffix was never the confidentiality gate here,
the path guard is (sensitive-file blocklist, /root and /etc trees, realpath
before the check), and it still runs on every registration.
- Widening what can be READ must not widen what can RUN. html/htm join svg in
serveRawFile's download-only branch, so markup is never served with a
renderable type on our own origin; other text goes out as inert
text/plain; charset=utf-8 with nosniff, matching what the path picker does.
The preview reads through fetch(), which ignores the disposition, so a
clicked .html still shows its source.
- ~/.codeman*/state.json joins isSensitivePath. It persists
SessionState.envOverrides and the env allowlist admits key-shaped names
(GEMINI_API_KEY, CLAUDE_CODE_*), so it can hold a live credential. Same
treatment as hook-secret and users.json, and the rest of the tree stays
attachable.
- The terminal sends an out-of-workspace path to the preview instead of the log
viewer. In-workspace text keeps the tail viewer, which is the point of it, and
file-stream-manager's allowlist is untouched: no `tail -f` on arbitrary host
paths.
- The by-id text preview is bounded like the workspace one: a Range request for
the first 512KB (a real partial read, not a discarded 50MB download) plus a
500-line cap, with the footer saying so.
Verified on an isolated instance: a 1.1MB external log opens in ~1.8s showing
500 lines with "showing first 500 lines" in the footer; json, yaml and code
preview; an .html carrying a script tag renders as source and does not execute;
.svg is still refused; a terminal click on an external .yaml opens the preview
with no log viewer and no attachment card; an in-workspace .log still opens the
streaming tail viewer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A clip an agent wrote inside the workspace played with a working scrub bar,
while the same file in /tmp was refused as an unsupported type. The workspace
preview classified media with its own inline extension sets and the attachment
allowlist had no media at all, so the two paths disagreed about what a video is.
- VIDEO_ATTACHMENT_EXTENSIONS and AUDIO_ATTACHMENT_EXTENSIONS now live in
attachment-registry.ts and are imported by file-content's classification, so
both paths answer the same. mp4/webm/mov/m4v/ogv and
mp3/wav/ogg/oga/m4a/aac/flac/opus join the attachment allowlist.
- Real MIME types for those extensions. Without one the raw route falls back to
application/octet-stream, which a <video> refuses to decode: the player
renders and then does nothing.
- getAttachmentType() gained the video and audio members of
AttachmentDetectedType. Attachment cards have no per-type CSS and their
thumbnail falls back to the type label, since the thumbnailer has no media
branch and answers 204 rather than spawning a converter.
- The preview overlay's by-id branch renders <video>/<audio> with the same
markup as the workspace branch, playsinline included. Serving was already
range-aware, so seeking works.
The image-watcher keeps its own narrow detection list (png/pdf/docx/pptx), so
this does not start popping cards for every video an agent writes. Text types
that are not md or txt (.json, .log, code files) remain out of the allowlist by
choice and still report what is previewable instead.
Verified on an isolated instance: an external mp4 and mp3 both play, seek, and
report the right duration, matching the in-workspace clip exactly, and a click
on an external mp4 in the terminal opens the player with no attachment card.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A path an agent prints was already underlined in the terminal, but clicking
one opened the preview overlay on "File not found": file-content/file-raw
resolve against the session workingDir and refuse anything outside it, and the
paths agents print most (a /tmp capture, Claude's own scratchpad, another
checkout) are outside it by definition. In the response viewer those paths were
not links at all.
- openFilePreview() detects an out-of-workspace path and registers it through
POST /api/sessions/:id/attachments first, rendering by attachment id. That is
the surface built for live external files, so the server-side guard is
unchanged: secret trees blocked, symlinks resolved, extension allowlist. The
workspace routes keep refusing escapes exactly as before.
- New optional `notify` field on that route. `notify: false` suppresses only the
attachment:detected broadcast, so a click does not also pop a card announcing
the file already filling the screen. Default stays true for the CLI and
publish callers.
- _linkifyFilePaths() links paths in rendered response-viewer markdown. It walks
text nodes and builds anchors with DOM APIs (the source is model output; never
a string rebuild of sanitized markup), skips subtrees already inside an <a>,
and keeps the message text byte-identical so copy-code is unaffected.
- One path pattern in constants.js now feeds both the xterm link provider and
the chat linkifier, a fresh instance per call since lastIndex is per-object
state. It picks up /Users and /mnt roots (nothing was clickable on macOS or
WSL), plus docx/pptx and video/audio extensions.
- .file-preview-overlay moves to z-index 5100, above the response viewer at
5000. At its old 2000 a path clicked in the chat opened the overlay behind the
panel it was launched from.
Verified end to end on an isolated instance, desktop and phone viewport: real
clicks in the terminal and the chat both render the image, external md and pdf
render, /etc/hosts is still refused, workspace previews unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Installing hooks into any workspace a Claude session runs in is the right
default, but it takes a decision away from a user who deliberately removed
them: nothing on disk distinguishes "removed on purpose" from "never had any",
so they would come back on the next session create.
Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs ->
Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks
block that is already present is still refreshed when stale (COD-91), but one
is never added.
Every create path routes through one applyWorkspaceHooks() helper so the gate
cannot apply to some paths only, and the boot-time recovery sweep honours it too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A session in a linked case (or any pre-existing repo) ran with no hooks block
at all: writeHooksConfig only fires when Codeman CREATES the case directory,
and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven
surface was therefore dead in exactly the place most sessions run: no tab
alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no
Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn,
and no stop/blocked for the agent wait endpoints.
Both session-create paths and restoreMuxSessions() now call
ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and
leaves a malformed settings file untouched. Claude Code re-reads
settings.local.json, so a session already running in the workspace starts
firing hooks without a restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs in the File Viewer's media player, both reproduced in a real
browser against an 18MB mp4 before and after the fix.
1. Closing the preview left the video playing. closeFilePreview() only
dropped the overlay's `visible` class, which is display:none and
nothing else, so the audio kept going with no visible player to pause.
Detaching the element is not a fix either: a detached HTMLMediaElement
plays on until it is garbage collected. _stopFilePreviewMedia() now
pauses, drops src and load()s every media element (also on re-open,
where overwriting innerHTML had the same effect), which additionally
aborts the in-flight download.
2. The scrub bar was inert. file-raw read the whole file and answered
200 with no Accept-Ranges, so Chrome reported video.seekable as
[0, 0] and silently reverted `currentTime = x`; Safari refuses to
start such media at all. Raw bodies are now streamed and range-aware:
Accept-Ranges: bytes on every response, 206 + Content-Range for a
Range request, 416 for one past EOF, and a malformed spec ignored
(200) per RFC 9110. Parsing is pure in src/web/http-range.ts.
Measured on tmp/codeman-crt-v5-66s.mp4 (18MB, 66.6s):
before seekable [0, 0] seek to 56.6s reverted to 3.9s close: still playing
after seekable [0, 66.56] seek to 56.6s landed at 60.2s close: paused, NETWORK_EMPTY
Range slices are byte-identical to `dd`, the full-file path is
byte-identical to the file, and the SVG octet-stream/attachment
hardening and the 50MB cap are unchanged (the cap is still checked
before the range).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Closes#259, closes#258. Both bottom out in the same gap: nothing tracked
whether the user was following live output or reading history.
#259 — the keyboard path forced the terminal to the bottom unconditionally
(onKeyboardShow/onKeyboardHide passed scrollToBottom:true, applied with no
check), so opening the keyboard while scrolled up yanked the user down. The
settle cycle now captures intent on its FIRST event, before any fit() has
reflowed the buffer, and returns to that anchor when the user was reading.
A later capture would read an already-moved viewportY, which is why the
capture point matters. The param is renamed restoreScroll to match.
Separately, flushPendingWrites gated viewport preservation on
_hasRecentUserScrollUp(), a 1500ms decay window, so a user who scrolled up and
then actually READ for longer lost protection mid-read. Being scrolled up IS
the intent however long ago it was expressed, so it now keys off position.
The recency window stays as a race guard on the sticky scroll-to-bottom.
The full-history repull already held the user's place and is unchanged.
#258 — truncation was reported by a grey line written INTO the terminal
("earlier output truncated"), which scrolls away with the output it describes,
cannot be acted on, and said the same thing whether the rest was one click away
or gone forever. The server set one `truncated` boolean at two sites meaning
opposite things, and the client discarded fullSize and source entirely.
The route now reports truncationReason ('tail' = intentional partial replay,
the rest is retained; 'capped' = the byte ceiling dropped it) plus
retainedBytes, and 'capped' is not downgraded by a later tail cut. The client
renders a dismissible banner outside terminal output with three honest states:
recoverable (offers Load full history), at-ceiling, and exhausted. The Load
button forces past the scroll cooldown but NOT past _replayWouldShrinkBuffer,
which still refuses a downgrade for repaint-mode panes.
The banner is an overlay, not a flex child: FitAddon derives rows/cols from the
terminal parent's computed height, so occupying real layout space would SIGWINCH
the CLI on every truncation-state change.
Verified in a real browser on the 7 skins: banner text and button clear 4.5:1
contrast on all of them, and terminal height is byte-identical with the banner
shown. The first cut used --bg-elevated and --accent-muted, which do not exist,
so light skins rendered a hardcoded dark bar under dark text; it now uses only
tokens every skin redefines.
test/terminal-scroll-intent.test.ts lives outside test/mobile/ deliberately —
that suite is excluded from test:ci, so a guard placed there is invisible to CI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.
Pi is a different shape of CLI from the other four, and three decisions
follow from that:
- It has NO permission prompts and no sandbox, so there is no
--dangerously-skip-permissions analog and none was invented. The
privilege-shaped knob is the tri-state approveProjectTrust, which makes
pi load and EXECUTE repo-local .pi/extensions TypeScript and install
missing project packages. clampExternalCliBypassForOwner() therefore
puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
--no-approve even when no config was sent, because pi's own default is
a prompt the session user could answer themselves. That helper had zero
test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
context, so admitting them would widen the allowlist for every mode at
once. Auth goes through pi's /login or the server's own environment.
--api-key is deliberately never wired: it would put a provider secret on
the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
mode is runtime-switchable via /settings; that flip was measured to put
the pane into the alt screen, which the strip would have corrupted.
pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.
Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).
Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.
Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.
Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.
- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
guard as the terminal socket, plus caps on concurrency, stream length and
frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
Claude, then a configured Deepgram key, then the browser
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two home-screen reports from @jordan8037310, both about history that is
present but unreachable.
#260 — "Resume Conversation" rendered 4 rows, then a button that appended
every remaining row into a `max-height: 240px` box, so 35 conversations
landed in a four-row scroll well with no ordering or filtering. Rendering
now goes through `_renderHistoryList()` over a cached corpus: 10 rows to
start, Show more/Show less that grows and shrinks the box (the height cap
is class-driven, `.history-list.expanded`), plus a filter box (name,
folder, #case label, prompts), a sort control (recent / name / folder,
pinned rows still first) and a shown-of-total count. A filter implies
expansion, so every match is visible, and the whole header hides as one
unit while a federated search is active. The A-Z sort keys off the same
string the row renders, since most rows are transcript-backed and carry
no session name at all.
#261 — the search box could not match a past project by folder name:
`harvestSources()` built its session corpus from the live in-memory map,
while past sessions come from `/api/sessions/unified` (lifecycle log +
transcript scan). Folding that scan into the request path would have cost
the search its no-filesystem-reads property, so the corpus arrives via a
bounded snapshot instead: `session-history-index.ts` is published as a
side effect of `/api/sessions/unified` (the home screen fetches it on
open, which is the same screen the search box lives on) and rebuilt
fire-and-forget, single-flight and TTL-guarded when a search finds it
stale. A result for a closed session now resumes the conversation rather
than selecting a tab that no longer exists, and is badged RESUME.
The snapshot is stored unscoped with a per-row owner and re-filtered
through canAccessOwned() on read, so multi-user sees exactly what
/api/sessions/unified exposes: own sessions only, host-wide transcript
history admin-only. Live rows are harvested first and win the dedupe.
Verified end-to-end against a real instance with 60 past sessions: cold
process answers its first search without history and its second with it;
folder-name queries return resume targets; clicking one posts the right
resumeSessionId + workingDir.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses all four findings from the #251 review:
- Scaffolding no longer writes through repository-controlled symlinks.
The guard lives in hooks-config.ts (settingsWriteBlocker) so it also
covers quick-start/docker/ralph writers, not just the clone route:
refuses a symlinked .claude or settings.local.json, a .claude that is
a file, or one resolving outside the case. The clone route surfaces
the refusal as a user-visible warning, and the CLAUDE.md write checks
presence via lstat so a BROKEN repo-shipped symlink counts as present
(existsSync follows links and would have created the outside target).
- Failed-clone cleanup can no longer delete a concurrent winner's tree:
git clones into an attempt-owned temp sibling (.<name>.cloning-<rand>)
which is atomically renamed into place; the loser reports
DESTINATION_EXISTS and only ever removes its own temp dir.
- decodeURIComponent(url.pathname) is guarded: malformed percent-escapes
now come back as BAD_SYNTAX instead of an uncaught URIError 500.
- The git pool's waiter queue is bounded (CODEMAN_MAX_GIT_QUEUE, default
16): overflow answers BUSY immediately (HTTP 429 via RATE_LIMITED),
and queue time counts against the operation's own deadline.
Tests: hostile symlink fixture repo (route level), settingsWriteBlocker
units, concurrent same-destination race, temp-dir leak assertions,
percent-escape rejection, and a fake-git pool-bounds suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).
- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
(tool_result-only entries stay silent); capture wiring in server.ts is
claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing
@DodgyBadger's proposal in #236: clone a public repository straight into
codeman-cases/<name> and register it as a normal local case.
POST /api/cases/clone is synchronous by design (request held open, bounded
by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation
surface. Success broadcasts the usual case:created event, so the case
still appears when a proxy idle-timeout kills the request mid-clone.
POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can
say, while the user is still typing, whether the URL is cloneable without
credentials, what its default branch is, and which branches/tags exist.
Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env,
ls-remote parse, stderr classification) and a thin IO half, so every
security decision is unit-testable without spawning anything:
- `<name>::<payload>` transports are refused as a family, not by name:
ext:: is the famous one, but any of them dispatches to git-remote-<name>
and turns a clone into arbitrary command execution.
- A leading `-` is refused AND every spawn puts `--` before the operands.
Either alone is one edit away from being a hole.
- argv arrays, never a shell. URLs carrying user:password@ are refused.
- gitNonInteractiveEnv() closes all four ways git can block on a prompt
with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM).
HOME/PATH stay inherited, so a user's own credential helper or ssh agent
keeps working; Codeman itself collects and stores nothing.
- The timeout signals the process GROUP, since clone fans out into
git-remote-https/index-pack children that outlive a signal to the parent.
- Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs
each) and a global 2-op pool, so N large clones cannot exhaust the host.
Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks
are merged into whatever .claude/settings.local.json the repo shipped, and
a repo that ships its own Claude settings is reported back as a warning
(those hooks run locally as soon as a session starts there). A failed clone
removes only the directory the attempt created, and refuses a pre-existing
destination outright, so it can never squat on a case name.
Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only
inside the caller's own case space. Local-path/file:// sources are the
exception and stay admin-only there.
UI: live verdict under the URL field, case name filled from the parsed repo
until the user types their own, branch/tag as a datalist of the remote's
real refs, optional shallow clone, and a Brain picker (installed CLIs only)
that points the Run button at the chosen agent. Starting a session stays
opt-in. The tab hides itself when the server reports no git.
Tests: the pure half exhaustively (every refusal has a case), plus real git
against a real local bare repo for clone/ref/timeout/cleanup, and a
route-level suite with unmocked fs that clones through the endpoint.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The picker behind Link Existing's "Browse" and the mobile keyboard's Path key
refused every path with a dot-prefixed segment, so `.github/workflows/ci.yml`
could not be selected and a hidden folder could not be opened at all. It gains
the same `.*` toggle as the File Viewer: default OFF, per-device, and applied to
both the listing and the preview endpoint, which re-resolves the path
independently.
That dotfile filter was quietly doing security work. The picker's roots include
Home, so with every hidden path unreachable the shared blocklist never had to
name the credentials that live in dot-directories. Lifting the filter removes
that accident, so `isSensitivePath` now covers them explicitly: SSH keys at any
depth rather than only under $HOME, GPG keyrings, AWS/GCloud/Azure/Docker/
Kubernetes credentials, npm, Yarn, git, gh, netrc, PyPI, RubyGems, Cargo and
Terraform tokens, .pgpass and .my.cnf, and the Claude and Codeman agent
credentials. `~/.codeman/` and `~/.claude/` stay attachable as trees, since the
publish skill and the review-card loop read from them; only their secret-bearing
members are named.
Everything else still applies with the toggle on: blocked trees, sensitive
files, root confinement, ownership scoping and symlink-escape checks. A hidden
entry whose realpath is a secret is dropped from the listing, and opening it is
refused.
Follows #221
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Owner decision: every Approvals Inbox UI surface (header bell, drawer,
phone overview answer strips, reload seeding) now requires enabling
approvalsInboxEnabled in App Settings -> Panels; only an explicit true
turns it on. The store, endpoints, and push Approve/Deny actions keep
running regardless (the push buttons are already opt-in per subscription).
Also replaces em-dashes with plain punctuation across the newly authored
comments, docs, and strings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Permission dialogs, AskUserQuestion questions and idle prompts from every
session now land in a server-side inbox (web/approval-inbox.ts, one item per
session, claude-mode only) and are answerable in place: a header bell + drawer
on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and
working push Approve/Deny buttons (previously dead ends, now answered straight
from sw.js with no tab open). Pending alerts survive reloads because the
frontend seeds from GET /api/approvals on init.
Answering sends the digit / Esc / prompt text through the existing tmux input
path; option digits are accepted only when they match options parsed from the
captured pane frame, and the answer path re-captures the pane first so a
dialog that already left the screen refuses with 409 instead of typing into
the composer. New elicitation_complete / elicitation_response hook matchers
resolve question items the moment they are answered in the terminal;
refreshStaleCodemanHooks heals existing cases.
Verified end-to-end against a live claude session: a real AskUserQuestion
dialog parsed into 5 option buttons and was answered from the drawer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three gaps found while auditing the agent skill.
`codeman skill install` / `uninstall` had no tests at all, including the linked-case
resolution that shipped in 1.14.2 with nothing guarding it. Covered now: global target
resolution, `--case` resolving through linked-cases.json, `--case` falling back to the
cases dir for an unlinked name, a missing or malformed registry degrading to the
fallback instead of throwing, and a nonexistent case being rejected. `resolveSkillTarget`
called `process.exit(1)` for a missing case, which would have killed the test runner, so
the pure resolution is split out and exported; CLI behavior is unchanged.
The `POST /api/sessions` injection call site was never exercised, because the shared
route mock hardcoded the gate off. The mock's gate is overridable per test now (default
still off, since other tests rely on that), and there is coverage that the path injects
when the setting is on, does not when it is off, and is claude-mode gated.
Nothing guarded skills/codeman/reference/endpoints.md against drifting from the routes
it documents, which is how it drifted in the first place. A static guard parses the
endpoints out of the markdown and asserts each is really registered, tolerating the
/api/v1 alias and path params.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two blockers from the pre-submission gate, both reproduced before fixing.
1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
code says the opposite in as many words ("NOT an error response,
deliberately"), the commit message says response codes are unchanged, and the
test asserts the 200. It was a leftover sentence from an earlier iteration that
would have shipped into the CHANGELOG announcing an API contract change that
does not exist — and errorCode values are SemVer-relevant per
docs/versioning-policy.md.
2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
master left all 9 tests green, while the commit message sells "plus the whole
WebSocket path" as part of the fix. Three tests added against the real WS
route — ACK on delivery, ACK withheld and seq re-opened when the write did not
land, and a deduplicated frame still ACKed so the client can drop it. Verified
the other way round: with ws-routes.ts reverted, the middle one fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.
Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.
Test fails against the pre-fix line and passes after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).
Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:
1. extractTranscriptEntrypoint() scanned any line containing the
substring "entrypoint", not specifically the first "type":"user"/
"type":"assistant" message line (unlike its sibling
extractFirstUserPrompt, which does scope to type). A transcript
that started under an older Claude Code version (no entrypoint
field) and got resumed under a newer one mid-conversation could
pick up the field from a much later message than the true first
one, misattributing the session's origin. Scoped it to match.
2. Added a regression test proving the tail-read fallback still
engages correctly when bookkeeping accumulation exceeds even the
new 128KB head window, not just the 16KB it previously blanked at.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).
Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.
Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>