#356 and #371 fixed the same leak two ways. #356 pointed CODEMAN_DATA_DIR at a
second throwaway directory and cleaned it up in afterAll and on exit; #371
deletes the variable along with CODEMAN_INSTANCE and CODEMAN_TMUX_SOCKET, so
`getDataDir()` falls back to `homedir()`, which the temp HOME already redirects.
Merged as they were, setup.ts set the variable and deleted it a few lines
later, and the second directory was created for nothing.
The strip wins: same protection, one tree to clean up, and the isolation test
#371 adds pins the list statically. The extra directory, its restore and its
two rmSync calls go, the vitest config `env` entries that set the same variable
go (they were documented as inert and would now be contradicted by the setup
file either way), the two test comments that described the old mechanism are
reworded, and CLAUDE.md's testing paragraph names the three stripped variables
and why CODEMAN_INSTANCE has to be stripped in the setup file rather than a hook.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 stopped a bare suite run from overwriting the production
`remote-hosts.json` by pointing `CODEMAN_DATA_DIR` at a throwaway dir, and it
gated every case-tree delete on the temp HOME. Both changes are right; the
explanation written next to them is not. It says `os.homedir()` reads
/etc/passwd rather than `$HOME` on Linux, which would mean the temp HOME in
test/setup.ts never worked. It does: libuv checks the env var before the passwd
entry (measured: `HOME=/tmp/x node -e 'console.log(os.homedir())'` prints
/tmp/x), and CLAUDE.md's testing section relies on exactly that.
What bypasses the temp HOME is `CODEMAN_DATA_DIR` itself. `getDataDir()` reads
it as an absolute override before it looks at `homedir()`, so one inherited from
the shell (a second instance, a beta run) sends the whole suite at the real data
dir. That is the case setup.ts now closes, and #371 names the same variable from
the other direction.
The comments in setup.ts, the `safeRmHomeTree` helper, the voice-routes and
case-clone tests now say that, and the containment gate is described as what it
is: defense in depth. CLAUDE.md's testing paragraph gets the same note so the
next reader does not chase a homedir() bug that does not exist.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.
The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).
Two smaller things in the same area:
- The cached answer was never invalidated, so after one successful reattach a
stale `true` would have revived the NEXT clean exit (the original bug back
after the first transport drop), and a cached `false` from a clean exit would
have left a manually restarted session with auto-reconnect permanently off.
The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
unreachable host stacked up to three ssh processes per dead session. An
in-flight set caps it at one.
The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.
Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.
Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.
OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.
Guard rails:
- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
branching reappears outside `stock.ts`, in any of its four shapes (`===`,
`!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
negated forms, which is how 36 of them survived an earlier pass. Every
allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
`privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
wire field is separate, bridged only by `legacyConfigAliases`. Getting
`privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
clamp's only handle on a CLI's privilege switch, and a wrong name clamps
nothing with no error and no failing test — so `schema.ts` rejects an entry
naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
`allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
thunks). A module-level const freezes at first import, so a CLI enabled while
the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
(`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
rather than measured. A test pins the list so it cannot quietly grow.
Three user-visible changes, all deliberate and named:
- `probeDockerCliVersion()` derives the in-container binary from the registry
rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
hardcoded map it replaces omitted while its own comment said the rule was
"every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
install hint is the install command rather than a docs URL, five CLIs gain
hints they never had, and the row order follows the catalog.
Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).
Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.
Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:
- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
the whole context-relative path, so `docker/.env` — which the deployment's own
README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
was picked up by `COPY . .` and baked into the image at
/opt/codeman/docker/.env. Verified in both directions against a real build
context: with a canary secret in docker/.env, the unfixed ignore file lets
/ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
reported "Case not found" on exactly the deployment the override exists for.
Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
literal on purpose: that one migrates the historical ~/claudeman-cases
directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
joins the documented list of files that genuinely belong in the repo root.
The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.
Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
1.24.2. Also corrects two numbers in the CLAUDE.md trust-dialog paragraph that
was written while the fix was still uncommitted: the keystroke cap is 6, not 3,
and the scan now schedules its own follow-up read rather than waiting on PTY
output that a static dialog never produces.
Audit of CLAUDE.md against the tree. Verified still accurate: the 31-module
frontend load order (matches index.html exactly), SSE registry parity at
157 = 157 (confirmed by running the parity test), config/ 21 files, types/ 22
domain files, 136 mobile device profiles, the version line, and every Quick
Reference command.
Drift corrected: 24 route modules to 25, ~220 handlers to ~227, system-routes
51 to 56, app.js ~5K lines to ~6.7K and 30 modules to 31, install.sh 92KB to
104KB. Completed the CLI resolver inventory, which was missing
deepseek-cli-resolver and omp-cli-resolver even though both modes are
documented, and named the shared cli-executable-resolver lookup chain.
Filled the gaps found by sweeping every src module against the file:
- Owner tab layouts (COD-359) had 6 source modules, 7 test files, 2 routes, an
SSE event and a state.json key, with zero mentions anywhere in CLAUDE.md or
docs/. The paragraph records the four things a reader would otherwise get
wrong: it is backend-only as of 1.24.1 with no frontend consumer, the service
is the sole mutation boundary, it projects onto PUT /api/session-order rather
than replacing it, and reconciliation is gated on a successful restore.
- codeman doctor and codeman users were undocumented top-level CLI commands.
- Four subsystems whose invariants lived only in their @fileoverview:
the workspace-trust dialog recognizer, proc-tree's bounded walk (the
2026-07-30 incident that took a machine down), deepseek-web-server (one
child process, deliberately not a shell session), and the Files panel
search matcher (globs are never compiled to a RegExp).
Also fixes a stale "156 event types" comment in constants.js (actual: 157) and
a contradiction in AGENTS.md, which still carried the retired "never run the
full suite inside a managed tmux session" rule against CLAUDE.md's current
"npm test is the gate and is safe to run bare".
Note: the trust-dialog paragraph documents trustDialogNextKey(), which is part
of a sibling session's in-flight fix for the Claude Code 2.1.252 layout change
(unnumbered, reversed options with "No, exit" highlighted, so a blind carriage
return picks exit and kills the pane). That fix was uncommitted in the shared
tree when this landed, so the doc leads the code until it is committed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NMN8UuvdBim3iM87reuQ9Z
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.
The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
Codex plan usage in the header chip (#346), a visible inline rename in
the session sidebar (#345), and the install.sh Tailscale re-run fix plus
the README network-access prompt description.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fifteen review findings on the dsh mode, the serious ones first:
- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
_configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
every dsh pane and applyEnvOverrides() lands after it, so a non-granted
owner who could redirect the base URL would have the operator's key sent
as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
The HERDR triple is set via LOCAL tmux setenv, which crosses neither
docker exec nor ssh, so such a session can never post a hook event and
the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
option parser cannot read a third-party TUI's frames, so an answer was a
blind keystroke into a foreign composer), and the push notification
carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
reports inside a 60s window (the TUI retries with backoff, so a retried
'working' could land after 'blocked' and resolve an approval whose
dialog was still on screen); 4xx responses exit 0 instead of retrying,
so one misconfigured session cannot feed the auth rate-limit bucket
until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
racing POSTs used to pick the same port and orphan the winner), and the
readiness poll / timeout paths only clear or stop the singleton while it
is still theirs. First click actually opens the tab now
(refreshWebviews, not the nonexistent loadWebviews). DELETE
/api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
create paths (impl moved into the resolver so all three share it) and no
longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
the remote branch; the Ralph auto-enable list gained deepseek;
HookEventType gained agent_working; the phone overview run menu filters
managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
child that reads stdin eats the rest of the script), bounds the exec
with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
identity (it rendered as an unstyled UA-grey button); stale markup
comment about the web shortcut rewritten; clamp docs updated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.
dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.
Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):
- `spawn_worker` grows a deepseek branch that gates on the harness
composer. ⚠️ Readiness there is NOT the stop signal: the harness
reports idle at BOOT ~300 ms before its composer paints (measured
2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
quick-start resolves on the boot edge, reports a turn that never ran,
and strands the prompt in a pane not yet taking input. Waiting for the
composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
concurrent call. Case names still have to be unique -- the mode never
disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
default set. That set also carries `idle`, which for an external CLI is
inferred from output stabilization: on a dsh worker whose TUI repaints
rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
with three minutes left to run. It also makes a wrong mode loud -- the
modes that cannot deliver `stop` answer 400 before writing anything,
instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
tagged duplicate, so the server truthfully reports `delivered:false`
about a write it skipped, and §1's cleanup then read a completed turn
as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
because the harness default still asks and a worker parked on an
approval row cannot finish a fan-out. The multi-user clamp still
applies.
Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).
The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
Three review findings on the DeepSeek Harness mode, plus one the third exposed.
1. The multi-user clamp was bypassable by a sibling field on the same request.
clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
_configureDeepSeek(), so a non-granted owner sending
envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
instance: a session created with permissionMode "read-only" and that override
ran with DSH_PERMISSION_MODE=danger-full-access in its pane.
Every other CLI's bypass is a command-line flag reachable only through the
per-CLI config, which is why the config clamp alone is the whole gate for
them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
that list because it aims the launcher at a profile tree whose plugin code
runs at boot, before any approval row can apply. Verified end to end in real
multi-user mode: a non-granted user sending both now gets workspace-write and
no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.
2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
signals only the direct child, and a plugin install fans out into
package-manager children that keep the inherited stdio pipes open, so `close`
never fires and the held-open request leaks with no route-level deadline.
Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
6s and both fan-out children were alive. Now detached: true plus negative-pid
SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
last-resort reap for a grandchild that escaped the group. Same probe after the
change: close fires, direct child and both grandchildren dead.
3. hooksAvailableForMode() promised more than a dsh session can deliver.
deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
triple is the only reason a dsh session posts hook events, so `until=stop` was
accepted and then blocked for the caller's whole timeout: the exact
infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
takes HookCapabilityOptions and every call site passes sessionHookOptions(),
with the deepseek arm reading `!== false` so a forgotten one degrades to the
old behaviour. The refusal names the setting rather than saying "no Claude
Code hooks", which would send the caller hunting a bug that is really a
setting they chose. Profile conformance stays unknowable at request time and
is documented as such. The stale "True for `claude` and nothing else" docblock
is corrected.
4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
capture read Claude's own transcript, and adding deepseek silently widened
both to a mode that has none. They compare mode === 'claude' directly now, and
a static check pins them there.
Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.
- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
20s in-place clock are the existing rich-sidebar ones, classified by
_mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
desktop home rail and the phone overview cannot disagree about what
"working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
existing [data-tab-orientation='vertical'] rule keeps matching both variants
untouched. A flip of detail ALONE still forces a full render (the stamps line
is emitted by the row template, not toggled by CSS) and re-runs
applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
never :is() - an :is() list takes its most specific argument, which would
lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
mid-word, the same reason the rich sidebar is 300px, so a rail that has never
been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
which also keeps the settings select on a named choice). A width the user has
chosen is never overridden: below 288px the created stamp is dropped rather
than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
applySessionListLayout(); a leaked interval would rewrite stamps in a list
that no longer has any.
Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.
Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Post-merge follow-ups from the deep review of #334 and #335, so they ship in
the same release as the features.
Tab-layout foundation (#335):
- PUT /api/session-order drops unknown/foreign ids again instead of 400ing
the whole write, in both the owner and the admin path (single-user requests
are the synthetic admin, so that path is the one the browser hits). The
frontend debounces its reorder push and swallows errors, so a session
deleted inside the debounce window silently cost the user the entire
reorder - and the endpoint sits on the stable /api/v1 surface, where the
pre-layout server merged leniently.
- A failed mux restore no longer locks explicit deletions into 500s for the
process lifetime: runSessionDeletion and webviewDeleted degrade to
best-effort without layout coordination, while the automated stale sweep
(runStaleSessionCleanup) stays fail-closed.
- sse-events doc comment: no 'suppressed' hook event exists; hooks stay 8.
- registerSessionWithLayout resolves its owner through ownerLayoutKey()
instead of a hardcoded '@single'.
Vertical rail (#334) - all rail-awareness gaps in sidebar-only predicates,
unified behind the new _isVerticalTabList() (sidebar OR rail):
- Drag-reorder read the insertion side from clientX in the rail, so
before/after was effectively arbitrary on vertical rows; the drag-over
indicators now draw as top/bottom edges there like the sidebar's.
- The active tab is scrolled into view in the rail (Alt+N/palette selection
used to leave the row below the fold).
- Floating subagent/ultracode windows anchor to the RIGHT of rail tabs, and
the connector redraw gates (render tail + strip scroll) cover the rail.
- Server-seeded tabOrientation is applied when the async settings load
resolves, not only at boot, so a fresh device shows the rail immediately.
- The pre-paint script stamps data-tab-orientation and --tab-rail-width
(sidebar-wins and solo carve-outs included), removing the flash of the
header strip on every vertical-mode load.
- The session name font defaults to 12px, the sidebar's historical 0.75rem
size, so installs that never touch the new slider are not restyled.
Also documents the rail in CLAUDE.md (second #sessionTabs host, mover
ordering, the axis-predicate rule) and gives tab-rail-resize.js its
@dependency/@loadorder header. Full gate green (6093 tests); the excluded
browser suite was run by hand - only the known environmental failures
(opencode/codex binaries) remain, identical to pristine master.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The agent skill's endpoints.md is what other agents read as ground truth, and
three of its facts went stale when grok landed:
- `/api/v1/grok/status` was added to the probe list, but the sentence after it
still said only Pi's response carries `.data.version`. Grok's carries it for
the same reason (a squatted binary name), and an agent that trusts the old
wording has no way to tell a misresolved grok from an absent one.
- the `active-tools` bullet listed grok among the modes it stays empty for, then
claimed in the same breath that `isExternalCliMode` "lists only those five".
- its three source line refs had all drifted: `isExternalCliMode` is now
session.ts:174-183 (it was already wrong before this branch), the external-CLI
early return is session.ts:2261, and TEXT_COMMAND_PATTERN is
bash-tool-parser.ts:89.
CLAUDE.md and architecture-invariants.md counted modes in their Docker-cases and
Web-tabs paragraphs ("any of the five CLI backends", "never a sixth
SessionMode"). Both numbers were already stale before grok (antigravity and pi
had made it seven) and grok is now in the agent image, so the counts are gone
rather than incremented: the invariant those sentences carry is that Docker and
web tabs are not modes at all, which no number has ever helped state. The two
plan docs keep their original wording, being historical design records.
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codeman tui supersedes the sc bash chooser: it reaches sessions 10+,
carries the server's real states instead of a static list, and leaves an
attach with one key. Every place that told a user to run sc now names the
tui equivalent, including the two wiki pages and install.sh's next-steps
banner. Both wiki pages also carried the wrong detach chord (Ctrl+A D;
the socket's prefix is C-b), which the tui makes moot.
Source comments that explained themselves as "the sc -l replacement" now
just say what they do. docs/tui-plan.md and CHANGELOG.md are historical
records and keep their references.
The data-dir half was already spelled out; the socket half only lived in
a function docstring, and the TUI is the first code that shells out to
`tmux -L` from a process that is not the server.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The user guide covers what the dashboard is (and is not), the two
non-interactive fast paths, the four groups and their ordering, the full
keymap, what answering an approval does server-side, and the SSH/narrow
and degraded cases. The example frame is a real 100x30 capture against
the E2E fake server, not a drawing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Confirming an AskUserQuestion left its tab flowing red for the rest of
the turn (owner report: ~8 minutes on a running session, with no dialog
anywhere on screen). Two separate bugs, both live-verified.
The re-capture erased the evidence the staleness check runs on. Claude
Code fires the Notification behind the dialog (measured 6-7s on v2.1.237,
documented up to ~30s), so the 600ms re-capture routinely lands on a
frame the user has ALREADY answered, parses nothing, and applyCapture
overwrote item.options with undefined. A MISSING options is how "we never
could read this dialog" is expressed, and those items stay answerable by
design, so a cleared field was indistinguishable from a never-parsed one
and the item became permanently unsweepable: it survived every
GET /api/approvals and every page reload, cleared only on `stop`, and
still accepted an answer, sending a bare `1` into a composer with no
dialog under it. applyCapture is now ADD-ONLY for options.
Nothing ran the staleness check while a page was open. It lived only in
GET /api/approvals, which seedApprovals() calls on init and reconnect, so
`stop` was the first thing that ever cleared an answered dialog. The
`working` signal now runs the pane-VERIFIED variant (resolveIfDialogGone
-> verifyStillAnswerable): the heuristic only decides when to look, the
screen decides the outcome, so the existing "working can flap" rule is
respected.
A frame that parses no options is now conclusive in two cases, and only
those, so an unreadable capture still keeps the alert: the item once
parsed options, or the frame shows Claude actively running a turn. A
modal dialog BLOCKS the turn, so the two cannot coexist - measured, a
live-dialog frame carries neither the elapsed-timer spinner nor the
"esc to interrupt" footer, which the dialog replaces with "Enter to
select". That second signal is reached by a delayed staleness pass (3s)
scheduled alongside the re-capture, which closes the late-hook case where
the prompt is answered before the hook lands: nothing ever parses, `stop`
may have gone by already, and the alert outlived reloads until the 12h
TTL. The pass is deliberately later than RECAPTURE_DELAY_MS, whose whole
reason for existing is that the hook can beat Ink to the screen.
Frontend: _onHookElicitationComplete cleared only the elicitation entry,
but an AskUserQuestion arrives as permission_prompt, so it was clearing
the wrong alert; it now clears both, matching the server's kind-agnostic
APPROVAL_RESOLVING_EVENTS.
Verified end to end on an isolated beta instance, not just in unit tests:
before, resolution could only come from the stop route (approval:resolved
always immediately preceding hook:stop); after, it arrives from the new
paths, and a simulated late hook resolves at +3.12s with no stop, no
working signal and no GET, while the pane is still working. Tests use
frames captured off a live pane and each new one was confirmed to fail
against the old behaviour.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:
$ [<0;88;20Mecho hello
bash: 0: No such file or directory
The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.
What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.
Details that are easy to get wrong:
* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
without turning tracking on is not asking about clicks, and counting
those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
broadcastSessionStateDebounced: the flag flips when a dialog opens, and
the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
the CLI re-emits its DECSET, which tmux does at client attach.
Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>