The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.
Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:
- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
judges the RESOLVED addresses of a name and refuses when any is blocked. This
is what closes DNS rebinding, which a hostname-string check cannot.
Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.
Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.
Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.
Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.
OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.
Guard rails:
- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
branching reappears outside `stock.ts`, in any of its four shapes (`===`,
`!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
negated forms, which is how 36 of them survived an earlier pass. Every
allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
`privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
wire field is separate, bridged only by `legacyConfigAliases`. Getting
`privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
clamp's only handle on a CLI's privilege switch, and a wrong name clamps
nothing with no error and no failing test — so `schema.ts` rejects an entry
naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
`allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
thunks). A module-level const freezes at first import, so a CLI enabled while
the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
(`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
rather than measured. A test pins the list so it cannot quietly grow.
Three user-visible changes, all deliberate and named:
- `probeDockerCliVersion()` derives the in-container binary from the registry
rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
hardcoded map it replaces omitted while its own comment said the rule was
"every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
install hint is the install command rather than a docs URL, five CLIs gain
hints they never had, and the row order follows the catalog.
Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).
Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.
Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:
- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
the whole context-relative path, so `docker/.env` — which the deployment's own
README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
was picked up by `COPY . .` and baked into the image at
/opt/codeman/docker/.env. Verified in both directions against a real build
context: with a canary secret in docker/.env, the unfixed ignore file lets
/ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
reported "Case not found" on exactly the deployment the override exists for.
Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
literal on purpose: that one migrates the historical ~/claudeman-cases
directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
joins the documented list of files that genuinely belong in the repo root.
The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.
Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.
That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.
Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
Small cleanup items from upstream review (Ark0N/Codeman#353):
- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
installer target (~/.omp/bin was an earlier unverified guess, confirmed
wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
SessionMode incl. shell -- not eighth), matched the install-path guidance
to the resolver fix, updated the version example to the actually-tested
18.0.8, and added a Docker-section caveat: --resume pinning does not
currently reach an in-container omp process, since Docker panes never see
ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
already linked to the -omp anchor, so the link was dead) and added an OMP
specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
and split two CSS lines that had two declarations jammed onto one line.
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.
DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.
The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:
- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
exported to make it testable: the exact "resume this OMP row from
history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
continuation silently degrades to omp's own ambiguous --continue,
in both call sites (session create and respawn pinning) - previously
silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
directory in omp-transcript.ts's parser, so a corrupted/malformed
session file can't point a downstream resume at a relative or empty
path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
case in mangleOmpWorkingDir(): the review's suggested realpath() fix
assumes omp itself resolves symlinks before mangling, which is
unconfirmed - guessing wrong there would trade one silent mismatch
for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
session-routes.ts (antigravity/opencode dynamic import line-wrapping)
that was blocking the pre-commit formatting gate on this file.
Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.
Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.
Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.
Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.
Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.
Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.
That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).
Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
Fifteen review findings on the dsh mode, the serious ones first:
- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
_configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
every dsh pane and applyEnvOverrides() lands after it, so a non-granted
owner who could redirect the base URL would have the operator's key sent
as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
The HERDR triple is set via LOCAL tmux setenv, which crosses neither
docker exec nor ssh, so such a session can never post a hook event and
the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
option parser cannot read a third-party TUI's frames, so an answer was a
blind keystroke into a foreign composer), and the push notification
carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
reports inside a 60s window (the TUI retries with backoff, so a retried
'working' could land after 'blocked' and resolve an approval whose
dialog was still on screen); 4xx responses exit 0 instead of retrying,
so one misconfigured session cannot feed the auth rate-limit bucket
until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
racing POSTs used to pick the same port and orphan the winner), and the
readiness poll / timeout paths only clear or stop the singleton while it
is still theirs. First click actually opens the tab now
(refreshWebviews, not the nonexistent loadWebviews). DELETE
/api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
create paths (impl moved into the resolver so all three share it) and no
longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
the remote branch; the Ralph auto-enable list gained deepseek;
HookEventType gained agent_working; the phone overview run menu filters
managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
child that reads stdin eats the rest of the script), bounds the exec
with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
identity (it rendered as an unstyled UA-grey button); stale markup
comment about the web shortcut rewritten; clamp docs updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four review findings on the worker-transcript feature:
- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
transcripts live in the container's / remote host's own ~/.dsh, which
the local reader can never see, so the transcript path returned
'nothing said yet' forever and an agent polling such a worker starved
on an answer that existed. Gated on !session.docker && !session.remote
(statically pinned) and documented in the integration guide.
- last-response reads are memoized on (path, mtime, size, blocks): the
skill's last_text polls once per second, and each poll decompressed and
reparsed the whole file on the event loop even when nothing had been
appended. An unchanged poll now costs one stat.
- The pairing ladder's comment claimed /new is served by step 2; in truth
the boot-window transcript wins for as long as it exists (deliberately:
preferring newest-eligible would hand a worker its busier sibling's
reply). The comment now states the real tradeoff instead of the
aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
corrupt frame truncates the decode there, which is the safe behavior.
- stripReasoningPrefix no longer runs on user prompt text, so a prompt
containing a literal </think> renders whole in blocks view.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.
dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:
1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
(one-shot and streaming alike) stops at the first frame end: a real
56-line transcript decoded as 1 line / 158 bytes -- the session header
alone, i.e. a silent truncation that reads as "nothing said yet"
forever. `zstdFrameRanges()` walks frame and block headers to find
exact boundaries; splitting on the 4-byte magic would corrupt
everything after a magic sequence occurring inside compressed data.
zstd is resolved at RUNTIME because it landed in Node 22.15 while the
project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
surfaced as `Turn error: …` (and a non-error early stop as
`Turn ended: …`) rather than as an empty string, which an agent reads
as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
the streamed deltas fill in only for a step that never finalized, so a
partial answer is readable mid-turn and never doubled. "Finalized" is
tracked as a set of steps rather than as non-empty text, because a
step whose whole reply was reasoning strips to '' at the `</think>`
boundary and would otherwise resurrect the raw deltas in its place.
Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.
An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.
The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:
- Exactly one server. A second click reuses the running one instead of racing
it for a port; the session flow could not do this at all, because two clicks
were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
own /api against the browser authority, and a Codeman reachable at both
loopback and a tailnet name has two. Reusing a server fenced for the other
origin renders a page whose every call 403s, which reads as a broken
dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
signalled at once, which also means it would outlive Codeman and hold its
port against the next start - the exact EADDRINUSE this feature already got
wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
for a stack trace to land, so a failed spawn reports its own tail.
The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.
`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.
Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.
1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
precisely the port a DeepSeek user is most likely to be serving on already,
so the launch died with EADDRINUSE against the user's own server. The port
now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
free loopback port by BINDING it (a connect probe cannot tell "free" from
"listening but not answering yet").
2. It opened the tab unconditionally. The crashed server left a saved dashboard
pointing at nothing, with the failure only visible in a shell tab nobody had
a reason to look at. The launch now polls the existing webview probe until
the URL answers, and on timeout reports the error naming the shell tab
instead of persisting a dead dashboard.
3. The saved tab was untrusted, so the frame was sandboxed without
`allow-same-origin` and the dashboard was broken twice over: the dsh
client-runtime reads `localStorage` while loading its plugins and died there
("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
every `/api` call no matter which authority `--trusted-host` named. Passing
`location.host` only means anything once the frame actually carries that
origin, so `--trusted-host` had never once done its job. The managed tab is
now created `trusted: true`.
That trade is real and deliberate: a trusted proxied frame is same-origin
with Codeman and can reach Codeman's API. It is defensible only because this
dashboard is an agent harness Codeman just started itself, on loopback, which
can already run code as the user. It is not a precedent for trusting
third-party dashboards, which is why it is set at this one call site rather
than defaulted.
Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.
`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.
Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
Three review findings on the DeepSeek Harness mode, plus one the third exposed.
1. The multi-user clamp was bypassable by a sibling field on the same request.
clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
_configureDeepSeek(), so a non-granted owner sending
envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
instance: a session created with permissionMode "read-only" and that override
ran with DSH_PERMISSION_MODE=danger-full-access in its pane.
Every other CLI's bypass is a command-line flag reachable only through the
per-CLI config, which is why the config clamp alone is the whole gate for
them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
that list because it aims the launcher at a profile tree whose plugin code
runs at boot, before any approval row can apply. Verified end to end in real
multi-user mode: a non-granted user sending both now gets workspace-write and
no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.
2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
signals only the direct child, and a plugin install fans out into
package-manager children that keep the inherited stdio pipes open, so `close`
never fires and the held-open request leaks with no route-level deadline.
Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
6s and both fan-out children were alive. Now detached: true plus negative-pid
SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
last-resort reap for a grandchild that escaped the group. Same probe after the
change: close fires, direct child and both grandchildren dead.
3. hooksAvailableForMode() promised more than a dsh session can deliver.
deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
triple is the only reason a dsh session posts hook events, so `until=stop` was
accepted and then blocked for the caller's whole timeout: the exact
infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
takes HookCapabilityOptions and every call site passes sessionHookOptions(),
with the deepseek arm reading `!== false` so a forgotten one degrades to the
old behaviour. The refusal names the setting rather than saying "no Claude
Code hooks", which would send the caller hunting a bug that is really a
setting they chose. Profile conformance stays unknowable at request time and
is documented as such. The stale "True for `claude` and nothing else" docblock
is corrected.
4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
capture read Claude's own transcript, and adding deepseek silently widened
both to a mode that has none. They compare mode === 'claude' directly now, and
a static check pins them there.
Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Post-merge follow-ups for PR #329 (shared CLI executable resolution):
- Negative-cache resolution misses with a doubling backoff (1min -> 5min
cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
shared resolver cached success only, so a missing CLI re-ran the whole
chain - ending in a synchronous interactive login-shell spawn bounded by
the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
attempt, stalling the event loop each time, forever. Success still caches
for the process lifetime, so an installed CLI is picked up within minutes
without a restart. Tests drive the backoff via an injectable clock
(createCliExecutableResolver `now` option, threaded through the
createPiResolverForTest / createAntigravityResolverForTest wrappers).
- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
pi/claude --version probes: execFileSync's timeout only SENDS the kill
signal and then keeps waiting for the child to exit, and interactive bash
ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
survived the timeout and blocked the server permanently.
- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
test pinned the deletion): under vitest the production resolver host now
replaces un-injected IO primitives with inert stubs - no real PATH
scanning, no login-shell spawns - and probePiVersion never executes a
`pi` candidate again (`pi` is a generic binary name, so route tests
hitting /api/pi/status executed whatever binary the machine carried).
Tests opt in through the runCommand/isExecutableFile injection hooks or
allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
test is replaced by behavioral pins, including a real-executable fixture
in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
is ever removed again.
- Wire the six get*NotFoundMessage() exports (previously dead) into their
intended call sites: the createSession throws in tmux-manager and the
availability gates on POST /api/sessions and POST /api/quick-start in
session-routes, replacing a third hardcoded copy of the text. A not-found
error now names where resolution looked (server PATH, login shell,
checked directories). npm run knip no longer reports any unused export
from the resolver modules.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three post-merge fixes for the external-CLI response viewer:
- ?context=full blocks now carry role ('user' for prompts, 'assistant'
for response/status/tool). The frontend's loadFullContext() renders
via msg.role, so the roleless blocks lost the "You" badge and every
turn rendered as the agent. kind/label/text are unchanged and the
frontend needs no change.
- normalizeDividerStatusLine() dropped its backtracking regex
(/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
minutes at 10,000), and pane text is agent-controlled with buffers up
to 32MB. Replaced by a linear counter walk with the identical accept
set and captured content, pinned char-for-char against the old regex
by a brute-force corpus test plus a hostile-input regression test
that fails by timeout with the RegExp version (same approach as the
glob-matcher hardening in 68ae9a8).
- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
empty-viewer symptom the transcript branch exists to fix. The list
stays a local duplicate of isExternalCliMode() (importing session.ts
would drag node-pty into the pure module); a new exhaustive parity
test asserts the two mode sets can no longer drift.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.
Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.
Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.
Tests: 24 cases in test/repo-status.test.ts.
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.
For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.
Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.
The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.
Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.
compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.
The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.
Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Persist the 'I checked it' state of yellow idle tab alerts across reloads and devices.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled,
default ON) moved from a session-routes-local helper into hooks-config.ts as
applyWorkspaceHooks(workspace, install?), and the claude session-create sites
that bypassed it now go through it: cron job fires (cron-service), legacy
scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research
and planner one-shots. A cron or scheduled run firing in a linked case that
never had an interactive session ran hook-blind (no stop for completion
detection, no tab alert on a blocking dialog).
The shared core also carries the two guards every caller needs: a workspace
that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot
recovery sweep used to resurrect a deleted repo as an empty tree holding only
.claude/settings.local.json), and all errors are swallowed since a create must
never fail on hooks. Route handlers keep resolving the setting through their
ConfigPort and pass it in; non-route callers omit it and the core reads
settings.json itself (absent key or unreadable file = ON).
Two adjacent gates tightened in session-routes:
- the docker quick-start hooks branch excluded the five external CLIs but let
`shell` through, contradicting its own rule that only claude reads .claude
hooks; it is now gated on mode === 'claude'
- the statusLine exporter call in POST /api/sessions got the same
!remote && body.workingDir guard the hooks call got in 499d355 (it mkdirs the
same way, so a remote attach created a junk user@host:session dir locally and
a cwd-fallback create wrote into $HOME)
plan-routes' one-shot deliberately stays out: its workingDir is process.cwd(),
exactly the target 499d355 forbids writing into. restoreMuxSessions stays out
too: the boot sweep already covers recovered workspaces.
Tests: quick-start existing-case install, docker claude-installs/shell-does-not,
and the core directly (default-ON install, OFF add-nothing, OFF still heals a
stale block, malformed file untouched, vanished workspace skipped); the remote
and cwd-fallback regressions now also send statusLineTelemetry:true to pin the
statusLine guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three follow-ups from the 1.19.0 review of the activity-ordered home screens:
- Hook events now ride the same debounced session state broadcast the
working/idle handlers use. The blocked group ranks on lastActivityAt, and
without this a permission prompt raised after page load kept ranking by
whatever stamp the browser loaded with.
- A working row with no submit stamp now shows the lastActivityAt fallback
its sort anchor already uses: a row must never be ranked by a number it
does not display.
- Alt+digit resolves through the live-session projection the render paints
(sessionOrder minus dead ids), so a stale id cannot shift every painted
number off its target, web tabs included. New tests pin both surfaces to
one shared order and the numbering to the live projection.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A claude-mode attachRemoteSession create overwrites workingDir with the
user@host:session pseudo-path, which is a RELATIVE path locally — the old
refresh-only call no-op'd on it, but ensureCodemanHooks mkdirs, so it
created a junk local directory. And with workingDir omitted the cwd
fallback reaches the hooks write unvalidated; under installer-created
services cwd is $HOME, so hooks materialized in ~/.claude/settings.local.json.
Both guarded at the applyWorkspaceHooks call site; regression tests prove
the remote attach leaves no junk dir and the no-workingDir create leaves
the server cwd untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A .json/.log/.yaml/code path outside the session workspace was refused as an
unsupported type, and clicking one in the terminal made it worse: text goes to
the log viewer, which spawns `tail -f` and allows only the workspace, /var/log
and ~/logs, so it answered "Path must be within working directory or allowed
log directories" while the same path clicked in the response viewer previewed
fine. Two surfaces, two answers, for a file the session can already cat.
- TEXT_ATTACHMENT_EXTENSIONS IS EDITABLE_EXTENSIONS (config/file-editing.ts),
not a second curated list that would drift from it. The rule reads: if the
viewer would open a file for editing inside the workspace, the same file
outside it can be read. The suffix was never the confidentiality gate here,
the path guard is (sensitive-file blocklist, /root and /etc trees, realpath
before the check), and it still runs on every registration.
- Widening what can be READ must not widen what can RUN. html/htm join svg in
serveRawFile's download-only branch, so markup is never served with a
renderable type on our own origin; other text goes out as inert
text/plain; charset=utf-8 with nosniff, matching what the path picker does.
The preview reads through fetch(), which ignores the disposition, so a
clicked .html still shows its source.
- ~/.codeman*/state.json joins isSensitivePath. It persists
SessionState.envOverrides and the env allowlist admits key-shaped names
(GEMINI_API_KEY, CLAUDE_CODE_*), so it can hold a live credential. Same
treatment as hook-secret and users.json, and the rest of the tree stays
attachable.
- The terminal sends an out-of-workspace path to the preview instead of the log
viewer. In-workspace text keeps the tail viewer, which is the point of it, and
file-stream-manager's allowlist is untouched: no `tail -f` on arbitrary host
paths.
- The by-id text preview is bounded like the workspace one: a Range request for
the first 512KB (a real partial read, not a discarded 50MB download) plus a
500-line cap, with the footer saying so.
Verified on an isolated instance: a 1.1MB external log opens in ~1.8s showing
500 lines with "showing first 500 lines" in the footer; json, yaml and code
preview; an .html carrying a script tag renders as source and does not execute;
.svg is still refused; a terminal click on an external .yaml opens the preview
with no log viewer and no attachment card; an in-workspace .log still opens the
streaming tail viewer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A clip an agent wrote inside the workspace played with a working scrub bar,
while the same file in /tmp was refused as an unsupported type. The workspace
preview classified media with its own inline extension sets and the attachment
allowlist had no media at all, so the two paths disagreed about what a video is.
- VIDEO_ATTACHMENT_EXTENSIONS and AUDIO_ATTACHMENT_EXTENSIONS now live in
attachment-registry.ts and are imported by file-content's classification, so
both paths answer the same. mp4/webm/mov/m4v/ogv and
mp3/wav/ogg/oga/m4a/aac/flac/opus join the attachment allowlist.
- Real MIME types for those extensions. Without one the raw route falls back to
application/octet-stream, which a <video> refuses to decode: the player
renders and then does nothing.
- getAttachmentType() gained the video and audio members of
AttachmentDetectedType. Attachment cards have no per-type CSS and their
thumbnail falls back to the type label, since the thumbnailer has no media
branch and answers 204 rather than spawning a converter.
- The preview overlay's by-id branch renders <video>/<audio> with the same
markup as the workspace branch, playsinline included. Serving was already
range-aware, so seeking works.
The image-watcher keeps its own narrow detection list (png/pdf/docx/pptx), so
this does not start popping cards for every video an agent writes. Text types
that are not md or txt (.json, .log, code files) remain out of the allowlist by
choice and still report what is previewable instead.
Verified on an isolated instance: an external mp4 and mp3 both play, seek, and
report the right duration, matching the in-workspace clip exactly, and a click
on an external mp4 in the terminal opens the player with no attachment card.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A path an agent prints was already underlined in the terminal, but clicking
one opened the preview overlay on "File not found": file-content/file-raw
resolve against the session workingDir and refuse anything outside it, and the
paths agents print most (a /tmp capture, Claude's own scratchpad, another
checkout) are outside it by definition. In the response viewer those paths were
not links at all.
- openFilePreview() detects an out-of-workspace path and registers it through
POST /api/sessions/:id/attachments first, rendering by attachment id. That is
the surface built for live external files, so the server-side guard is
unchanged: secret trees blocked, symlinks resolved, extension allowlist. The
workspace routes keep refusing escapes exactly as before.
- New optional `notify` field on that route. `notify: false` suppresses only the
attachment:detected broadcast, so a click does not also pop a card announcing
the file already filling the screen. Default stays true for the CLI and
publish callers.
- _linkifyFilePaths() links paths in rendered response-viewer markdown. It walks
text nodes and builds anchors with DOM APIs (the source is model output; never
a string rebuild of sanitized markup), skips subtrees already inside an <a>,
and keeps the message text byte-identical so copy-code is unaffected.
- One path pattern in constants.js now feeds both the xterm link provider and
the chat linkifier, a fresh instance per call since lastIndex is per-object
state. It picks up /Users and /mnt roots (nothing was clickable on macOS or
WSL), plus docx/pptx and video/audio extensions.
- .file-preview-overlay moves to z-index 5100, above the response viewer at
5000. At its old 2000 a path clicked in the chat opened the overlay behind the
panel it was launched from.
Verified end to end on an isolated instance, desktop and phone viewport: real
clicks in the terminal and the chat both render the image, external md and pdf
render, /etc/hosts is still refused, workspace previews unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Installing hooks into any workspace a Claude session runs in is the right
default, but it takes a decision away from a user who deliberately removed
them: nothing on disk distinguishes "removed on purpose" from "never had any",
so they would come back on the next session create.
Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs ->
Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks
block that is already present is still refreshed when stale (COD-91), but one
is never added.
Every create path routes through one applyWorkspaceHooks() helper so the gate
cannot apply to some paths only, and the boot-time recovery sweep honours it too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A session in a linked case (or any pre-existing repo) ran with no hooks block
at all: writeHooksConfig only fires when Codeman CREATES the case directory,
and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven
surface was therefore dead in exactly the place most sessions run: no tab
alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no
Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn,
and no stop/blocked for the agent wait endpoints.
Both session-create paths and restoreMuxSessions() now call
ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and
leaves a malformed settings file untouched. Claude Code re-reads
settings.local.json, so a session already running in the workspace starts
firing hooks without a restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two live failures from one root cause: Claude Code loads a same-named
user-level skill (~/.claude/skills/codeman, written once by `codeman skill
install`) over the fresh per-case copy, and nothing ever refreshed it. A
stale Aug-9 copy (pre fast-path, pre lineage header) made every agent-driven
spawn run the old recipes: workers spawned serially with pid polls and
without X-Codeman-Parent-Session, so the web UI drew no lineage arcs.
- refreshUserAgentSkill(): session create now refreshes a marker-owned
user-level copy (refresh-only: absent copies are not installed,
foreign/symlink copies stay untouched).
- seedAgentSessionPreamble(): local claude session create pre-seeds the
skill's preamble into ${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh,
single-sourced from the new skills/codeman/preamble.sh, so the skill's §0
bootstrap collapses to a two-line loader instead of a ~150-line paste the
model has to type out (measured ~47s of generation per run).
- SKILL.md: §0 now leads with the loader and keeps the full block as the
stale/missing fallback; explicit verbatim-paste warning (a hand-assembled
preamble is how the header and the fast-path functions got lost);
spawn_worker also sends parentSessionId in the body as defense in depth;
preamble stamp bumped to 1.18.3 so pre-fix cached preambles self-heal.
- test/agent-skill.test.ts pins preamble.sh byte-identical to the SKILL.md
heredoc and covers seeding (XDG + HOME fallback, 0600) and the user-level
refresh (absent/stale/foreign).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two bugs in the File Viewer's media player, both reproduced in a real
browser against an 18MB mp4 before and after the fix.
1. Closing the preview left the video playing. closeFilePreview() only
dropped the overlay's `visible` class, which is display:none and
nothing else, so the audio kept going with no visible player to pause.
Detaching the element is not a fix either: a detached HTMLMediaElement
plays on until it is garbage collected. _stopFilePreviewMedia() now
pauses, drops src and load()s every media element (also on re-open,
where overwriting innerHTML had the same effect), which additionally
aborts the in-flight download.
2. The scrub bar was inert. file-raw read the whole file and answered
200 with no Accept-Ranges, so Chrome reported video.seekable as
[0, 0] and silently reverted `currentTime = x`; Safari refuses to
start such media at all. Raw bodies are now streamed and range-aware:
Accept-Ranges: bytes on every response, 206 + Content-Range for a
Range request, 416 for one past EOF, and a malformed spec ignored
(200) per RFC 9110. Parsing is pure in src/web/http-range.ts.
Measured on tmp/codeman-crt-v5-66s.mp4 (18MB, 66.6s):
before seekable [0, 0] seek to 56.6s reverted to 3.9s close: still playing
after seekable [0, 66.56] seek to 56.6s landed at 60.2s close: paused, NETWORK_EMPTY
Range slices are byte-identical to `dd`, the full-file path is
byte-identical to the file, and the SVG octet-stream/attachment
hardening and the 50MB cap are unchanged (the cap is still checked
before the range).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>