#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.
The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.
Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.
Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 and #371 fixed the same leak two ways. #356 pointed CODEMAN_DATA_DIR at a
second throwaway directory and cleaned it up in afterAll and on exit; #371
deletes the variable along with CODEMAN_INSTANCE and CODEMAN_TMUX_SOCKET, so
`getDataDir()` falls back to `homedir()`, which the temp HOME already redirects.
Merged as they were, setup.ts set the variable and deleted it a few lines
later, and the second directory was created for nothing.
The strip wins: same protection, one tree to clean up, and the isolation test
#371 adds pins the list statically. The extra directory, its restore and its
two rmSync calls go, the vitest config `env` entries that set the same variable
go (they were documented as inert and would now be contradicted by the setup
file either way), the two test comments that described the old mechanism are
reworded, and CLAUDE.md's testing paragraph names the three stripped variables
and why CODEMAN_INSTANCE has to be stripped in the setup file rather than a hook.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
The loader refuses a `clis.json` with any group/world permission bit, read bits
included, so a file created with a normal umask (0644) is ignored. That is a
defensible posture for a file that chooses the binaries Codeman spawns, but two
things around it made the override feature look dead: the warning said
"group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was
returned to a caller nobody wired up, so nothing anywhere printed it. A user
following the docs got silence.
The message now names the rule and the command that satisfies it, the loader
logs every warning once on first load (the result is memoized, so once per
process), the module header stops claiming that nothing ever writes (the
quarantine rename of a malformed file is a write, on first use) and the
registry doc gains a short section on the override file with the 0600
requirement in it. Whether the check should relax to writable bits only is a
separate decision; this keeps the shipped behaviour and makes it visible.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 stopped a bare suite run from overwriting the production
`remote-hosts.json` by pointing `CODEMAN_DATA_DIR` at a throwaway dir, and it
gated every case-tree delete on the temp HOME. Both changes are right; the
explanation written next to them is not. It says `os.homedir()` reads
/etc/passwd rather than `$HOME` on Linux, which would mean the temp HOME in
test/setup.ts never worked. It does: libuv checks the env var before the passwd
entry (measured: `HOME=/tmp/x node -e 'console.log(os.homedir())'` prints
/tmp/x), and CLAUDE.md's testing section relies on exactly that.
What bypasses the temp HOME is `CODEMAN_DATA_DIR` itself. `getDataDir()` reads
it as an absolute override before it looks at `homedir()`, so one inherited from
the shell (a second instance, a beta run) sends the whole suite at the real data
dir. That is the case setup.ts now closes, and #371 names the same variable from
the other direction.
The comments in setup.ts, the `safeRmHomeTree` helper, the voice-routes and
case-clone tests now say that, and the containment gate is described as what it
is: defense in depth. CLAUDE.md's testing paragraph gets the same note so the
next reader does not chase a homedir() bug that does not exist.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.
The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).
Two smaller things in the same area:
- The cached answer was never invalidated, so after one successful reattach a
stale `true` would have revived the NEXT clean exit (the original bug back
after the first transport drop), and a cached `false` from a clean exit would
have left a manually restarted session with auto-reconnect permanently off.
The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
unreachable host stacked up to three ssh processes per dead session. An
in-flight set caps it at one.
The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
#356 introduced safeRmHomeTree/isUnderTestHome to stop tests from deleting
the PRODUCTION ~/codeman-cases tree on platforms where os.homedir() ignores
the $HOME override -- but only applied it to the one file caught doing it
live. CASES_DIR has no CODEMAN_DATA_DIR-style env override at all, so every
other test file's raw rmSync(join(CASES_DIR, ...)) was the same unguarded
pattern, just not yet triggered.
Routes every CASES_DIR delete in these 10 files through safeRmHomeTree:
cli-skill-target, edge-cases, integration-flows, operation-lightspeed,
ralph-integration, routes/case-clone-routes, routes/voice-routes,
session-cleanup, sse-events, sse-subscription-filter.
Also fixes one instance in case-clone-routes.test.ts that mkdirSync'd then
rmSync'd a CASES_DIR path directly with no guard at all -- the exact
clobbering pattern #356 exists to prevent, found by extending the sweep.
Held as a separate commit (and intended as a separate PR once #356 merges)
rather than folding into #356 -- keeps the already-checked skinny fix
reviewable on its own; this is the same bug class applied broadly, not new
functionality.
Verified: all 10 files pass (180 tests), npm run typecheck clean.
The "never writes hooks for a remote attach" test stubbed CODEMAN_DATA_DIR
to a separate throwaway dir just for this write, but session-routes.ts's
CODEMAN_CONFIG_DIR is a module-load-time constant frozen at test/setup.ts's
sandboxed dir before this test ever runs. The fixture landed somewhere the
route handler could never read, so the remote-host lookup silently failed
(NOT_FOUND) and the test passed for the wrong reason -- createErrorResponse
never sets reply.code(), so Fastify's default 200 made the NOT_FOUND branch
and the intended success branch indistinguishable by status code alone.
Write straight to getDataDir() instead, matching the docker-hosts fixture
convention already used elsewhere in this file. Verified the fix actually
exercises the success path (host resolves, 200 with a real session), not
just an accidental 200 from the error branch.
PR #356 stopped the remote-hosts.json fixture write from clobbering prod.
Two holes in the same file remain:
1. The quick-start afterEach still ran rmSync(CASES_DIR, recursive).
CASES_DIR is join(homedir(), 'codeman-cases'), and on Linux builds
where os.homedir() reads /etc/passwd instead of $HOME it resolves to
the PROD case tree - so a full-suite run deleted the real
~/codeman-cases. Add a shared safeRmHomeTree() containment gate that
only deletes a path under the redirected test HOME.
2. setup.ts teardown did rmSync(process.env.CODEMAN_DATA_DIR ?? '') AFTER
restoring the env - if a pre-existing prod CODEMAN_DATA_DIR was set,
that deleted prod. Capture the throwaway dir in a const and clean that.
A broader test-isolation sweep (10 files: cli-skill-target, edge-cases,
integration-flows, operation-lightspeed, ralph-integration,
case-clone-routes, voice-routes, session-cleanup, sse-events,
sse-subscription-filter) also applies the same containment gates to every
per-case delete. It is intentionally NOT included here to keep this PR
skinny; it is identified and available on request.
session-routes-workspace-hooks.test.ts wrote its h1/box/10.0.0.5 host
fixture into getDataDir()/remote-hosts.json. getDataDir() resolves via
homedir() → ~/.codeman (INSTANCE_SUFFIX='' by default), and overriding
HOME in test/setup.ts does NOT change os.homedir() on Linux — so every
full-suite run silently overwrote the PRODUCTION remote-hosts.json,
wiping user-defined remote hosts, emptying the launch-case dropdown and
breaking remote session creation (found live 2026-08-29).
The vitest v4 test.env config key is ignored (probe confirmed the
worker still saw CODEMAN_DATA_DIR=undefined), so the reliable fix is
stubbing the env inside the test: the fixture write now goes to a
throwaway /tmp dir via vi.stubEnv + finally unstub. Verified: prod
remote-hosts.json hash is identical before and after the suite run.
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
Adds `claude-fable-5-1` to the App Settings model picker and the five task-routing selects, mirroring how Fable 5 is already offered: a base option with data-ctx="1" plus its [1m] companion row. No settings-ui.js logic change, since the cards and the 1M switch are built from those options.
`test/setup.ts` gives every test file a temp HOME so the suite cannot touch the
real Codeman tree, and strips the env vars that would leak past it — but the
list only covered auth and the gesture flag. The three vars
`src/config/instance.ts` derives the data dir and tmux socket from were missing,
and they reach past the temp HOME:
- **`CODEMAN_DATA_DIR` is the one that matters.** It is an ABSOLUTE override
read in `getDataDir()`, so it bypasses HOME entirely: a developer who exports
it — or a shell left over from `codeman web -d` — has the suite reading and
WRITING their real `state.json`, `users.json`, `intents.json` and
`hook-secret`.
- **`CODEMAN_INSTANCE`** moves the data dir to `~/.codeman-<name>` and the
socket to `codeman-<name>`. Inside the temp HOME that is not data loss, but it
silently changes the paths tests assert on — and `scripts/run-beta.sh` exports
it, so any shell that has run a beta carries it.
- **`CODEMAN_TMUX_SOCKET`** renames the socket `resolveTmuxSocketName()`
returns. `TmuxManager` no-ops its shell commands under vitest, so this is
assertion drift rather than a stray `tmux -L` against prod — same class of
leak, same one-line fix.
They are deleted in the setup file rather than in a hook because
`CODEMAN_INSTANCE` is captured into a module-level const the first time
`config/instance.ts` is imported; a `beforeEach` would already be too late.
`test/test-env-isolation.test.ts` pins the whole list in two halves, because the
obvious half is not enough: asserting the vars are unset passes trivially on a
machine that never set them, so a removed `delete` line would sail through on
almost every box and on CI. The static half reads `setup.ts` and asserts each
name is deleted there, which fails everywhere. An anti-drift check catches the
other direction — a var stripped in `setup.ts` but never given a reason in the
list — and is scoped to the strip section so the teardown's restores are not
mistaken for strips.
Verified by demonstrating the leak: with the `CODEMAN_DATA_DIR` line removed and
the var exported, the runtime assertion fails; with the line restored it passes.
Full suite: no new failures against an upstream/master baseline.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Review item 4 named THREE live tables duplicating registry data. Two are now
read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`);
the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither
wired nor annotated, which is the state that item explicitly rules out.
It is not wired because the shape cannot express the live table: `credStore` is
ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the
CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares
none at all even though `.dsh` is seeded. Wiring it means making the field an
array and correcting those two entries — a change to credential seeding, which
is at once the worst thing in that file to get wrong and the least covered by
tests, since every docker IO path is no-op'd under vitest. It belongs in its own
change, measured against a real container.
So it is annotated instead, at the field, in the type's declared-for-later
header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the
last of which means wiring it later makes a test fail rather than leaving a
stale comment behind.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
CI caught three cron-service failures. Both are mine, from converting cron's
per-mode ladders to capability reads without checking what each ladder's scope
actually was.
**The pre-flight.** cron only ever pre-flighted `deepseek` — dsh is a profile
LAUNCHER, so "installed" is not "runnable" and a bare `dsh` can boot a profile
that cannot drive a pane. I replaced that with an unscoped
`resolveCliLaunchError(mode)`, which pre-flights EVERY mode, so a claude cron
job on a box with no claude binary now failed with "Claude CLI not found"
instead of reaching tmux-manager's own throw. Three tests assert the latter.
It is now gated on `discovery.launcherProfile !== undefined`, which is
byte-identical to the `mode === 'deepseek'` check it replaces and generalises to
the next launcher. The equivalent HTTP-route conversion was already scoped (to
`capabilities.external`, matching what that route has always pre-flighted); I
simply failed to carry the same reasoning across.
**The model.** cron's ladder was `mode !== 'shell' && mode !== 'deepseek'`, and
I read it as `capabilities.model.source === 'claude-settings-file'` — which is
the HTTP route's question, not cron's. There, every external CLI reads its model
from its own config object earlier in the chain, so only claude reaches the
global default; cron has no such config, so the same expression silently
narrowed the default model from eight modes to one. Now `!== 'none'`, which is
exactly the two entries the ladder excluded. Not caught by a test — found by
re-deriving each ladder's scope after the first failure.
Also names a fourth deliberate behaviour change in the changeset, found while
tracing these: `session.ts` carried a hand-written list of modes with no
direct-PTY fallback and OMP was missing from it, though CLAUDE.md's own text
says "all eight require tmux". `requiresMux` comes off the entry now, so an omp
session whose mux creation fails refuses instead of silently starting outside
tmux.
Verified by diffing failing tests BY NAME against an upstream/master baseline,
rather than by file as before — which is how the regression slipped through: the
three new failures landed inside a file already failing for unrelated
Windows-path reasons, and the aggregate count happened to collide.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.
Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.
Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.
OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.
Guard rails:
- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
branching reappears outside `stock.ts`, in any of its four shapes (`===`,
`!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
negated forms, which is how 36 of them survived an earlier pass. Every
allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
`privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
wire field is separate, bridged only by `legacyConfigAliases`. Getting
`privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
clamp's only handle on a CLI's privilege switch, and a wrong name clamps
nothing with no error and no failing test — so `schema.ts` rejects an entry
naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
`allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
thunks). A module-level const freezes at first import, so a CLI enabled while
the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
(`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
rather than measured. A test pins the list so it cannot quietly grow.
Three user-visible changes, all deliberate and named:
- `probeDockerCliVersion()` derives the in-container binary from the registry
rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
hardcoded map it replaces omitted while its own comment said the rule was
"every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
install hint is the install command rather than a docs URL, five CLIs gain
hints they never had, and the row order follows the catalog.
Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
docker/server.Dockerfile installed Debian's `docker.io` to get a client for the
socket mounted by Compose. That package is the full ENGINE: even with
--no-install-recommends it pulls 15 packages including containerd, runc, dmsetup
and iptables, none of which a container that only talks to a mounted socket can
use, and it ships Docker 20.10.24 (2023).
Copy the CLI and the buildx plugin from the official docker:29-cli image
instead. Measured on the same node:22-bookworm-slim base: 266 MB -> 108 MB, so
158 MB smaller with a current CLI (29.7.2) in place of a two-year-old one.
Three things verified rather than assumed, by building the real image and
running it:
- docker:cli is an ALPINE image, so copying a binary into this Debian one is
only safe because the binaries are static Go builds (ldd: "Not a valid dynamic
program"). In the built image, `docker --version`, `docker ps` and
`docker build` all work against a mounted host socket as the unprivileged
runtime user.
- buildx is copied on purpose. scripts/build-agent-image.mjs shells out to
`docker build` and Codeman auto-builds the agent image on the first Docker
case. Without the plugin that still works today — CLI 29 falls back to the
classic builder, tested — but that builder is deprecated and will be dropped,
so the plugin keeps the path supported.
- docker-compose is NOT copied: Codeman never shells out to it.
Pinned to the 29 major, matching how the base images here are pinned.
The header chip silently shrank from "5h 4% · 7d 52%" to a lone "7d 52%", which
reads as half the feature breaking rather than as an idle window.
Nothing was broken. Claude Code documents `rate_limits.five_hour` as "present
only while the API reports it and its resets_at has not passed", so between
5-hour session windows the key simply leaves the statusline payload. Codeman's
snapshot replaces the Claude half wholesale on every sample, so the segment
disappeared until usage opened a new window. Confirmed against a live 2.1.252
session by capturing real statusline payloads on an isolated tmux socket: the
boot render carries no `rate_limits` at all, and the first post-response render
carries both windows.
The slot now stays, with a dimmed em dash. Claude only: a missing CODEX bucket
means that plan has no such limit rather than an idle window, so those stay
omitted (pinned by the existing test). The placeholder can never stand alone
either — hasWindows() still gates the row, so a provider reporting nothing
renders nothing rather than a row of dashes. The tooltip says "5-hour limit: no
active session window" instead of dropping the line.
Verified in a real browser against a dev instance: the idle chip renders
"5h — · 7d 52%" with the dash at opacity 0.55 in --text-dim while the live
value keeps its green, and the chip holds its shape (100px idle vs 107px with
both windows).
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).
Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.
Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:
- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
the whole context-relative path, so `docker/.env` — which the deployment's own
README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
was picked up by `COPY . .` and baked into the image at
/opt/codeman/docker/.env. Verified in both directions against a real build
context: with a canary secret in docker/.env, the unfixed ignore file lets
/ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
reported "Case not found" on exactly the deployment the override exists for.
Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
literal on purpose: that one migrates the historical ~/claudeman-cases
directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
joins the documented list of files that genuinely belong in the repo root.
The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.
Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
Which conversation a pane is on was re-derived by correlating
~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is
bumped only by input that flows through Codeman's own write path
(Session.write / writeViaMux). A user who attaches to the pane's tmux session
directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at
its first line for that pane's whole life and the response viewer stayed pinned
to the launch conversation, showing a pre-/clear transcript indefinitely.
A UserPromptSubmit hook reports the live conversation id from inside the CLI
process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a
fact rather than a correlation: it never consults workingDir, so it cannot be
claimed by a sibling pane on the same folder, a closed tab, or a bare `claude`
in the user's terminal. A pane holding such an id skips the correlation
entirely, so the number of prompts eligible for cwd-based guessing goes DOWN,
never up — the naive alternative (relax the guard, or synthesize an anchor from
PTY activity) is the reverted bug the resolver's own comment describes.
The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted"
rather than "typed into Codeman's web terminal". Conversations vouched for
first-hand — and only those — extend a persisted claudeSessionChain, whose tail
re-pins the conversation when a surviving tmux session is re-attached after a
restart. ⚠️ start() resets the id at THREE points and the last one runs
unconditionally after the mux branch, so the tail is applied there too; patching
only the mux branch looks right and silently does nothing.
⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code
injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0
- stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd
already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the
shell binds the redirection to `true`, which never runs on the success path. The
discard is opt-in so the five SSE-fed events keep byte-identical command text
and no workspace's settings file is rewritten for them. The staleness marker is
quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted
needle never matches and the gate would rewrite every workspace on every spawn.
Existing workspaces heal on their next Claude spawn through the staleness sweep.
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.
One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.
A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.
This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
mobile.css hides #createCaseModal's .set-foot below 860px, and that modal's
header — unlike Settings' — carries no set-head-save. So on a phone the
Create/Link button existed nowhere and the modal could not be submitted at all.
Adds the header button and drives both together through switchCaseModalTab()
and submitCaseModal(), so whichever one is pressed the other shows the same
pending state and is equally unclickable. Following the Settings pattern also
means Add Case picks up the existing .set-head-actions:has(.set-head-save) tray
and .set-head-save sizing with no new CSS; the mobile.css comment that still
listed Add Case as a lone-× sheet is corrected to match.
1.24.2. Also corrects two numbers in the CLAUDE.md trust-dialog paragraph that
was written while the fix was still uncommitted: the keystroke cap is 6, not 3,
and the scan now schedules its own follow-up read rather than waiting on PTY
output that a static dialog never produces.
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.
Claude Code 2.1.252 rewrote the dialog. It used to be
❯ 1. Yes, I trust this folder
2. No, exit
and is now unnumbered, reversed, and highlights the option that quits:
❯ No, exit
Yes, I trust this folder
Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.
trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.
Two things only a live pane showed:
- The scan ran solely from the PTY onData handler. The arrow that moves the
cursor is the last output the pane produces, so the first fix parked every
session with the cursor sitting on the right option and no Enter ever sent.
It now schedules its own follow-up read (_trustDialogTimer, cleared in
_clearAllTimers()), offset past the scan throttle so the chain cannot break
on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.
The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.
Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
Audit of CLAUDE.md against the tree. Verified still accurate: the 31-module
frontend load order (matches index.html exactly), SSE registry parity at
157 = 157 (confirmed by running the parity test), config/ 21 files, types/ 22
domain files, 136 mobile device profiles, the version line, and every Quick
Reference command.
Drift corrected: 24 route modules to 25, ~220 handlers to ~227, system-routes
51 to 56, app.js ~5K lines to ~6.7K and 30 modules to 31, install.sh 92KB to
104KB. Completed the CLI resolver inventory, which was missing
deepseek-cli-resolver and omp-cli-resolver even though both modes are
documented, and named the shared cli-executable-resolver lookup chain.
Filled the gaps found by sweeping every src module against the file:
- Owner tab layouts (COD-359) had 6 source modules, 7 test files, 2 routes, an
SSE event and a state.json key, with zero mentions anywhere in CLAUDE.md or
docs/. The paragraph records the four things a reader would otherwise get
wrong: it is backend-only as of 1.24.1 with no frontend consumer, the service
is the sole mutation boundary, it projects onto PUT /api/session-order rather
than replacing it, and reconciliation is gated on a successful restore.
- codeman doctor and codeman users were undocumented top-level CLI commands.
- Four subsystems whose invariants lived only in their @fileoverview:
the workspace-trust dialog recognizer, proc-tree's bounded walk (the
2026-07-30 incident that took a machine down), deepseek-web-server (one
child process, deliberately not a shell session), and the Files panel
search matcher (globs are never compiled to a RegExp).
Also fixes a stale "156 event types" comment in constants.js (actual: 157) and
a contradiction in AGENTS.md, which still carried the retired "never run the
full suite inside a managed tmux session" rule against CLAUDE.md's current
"npm test is the gate and is safe to run bare".
Note: the trust-dialog paragraph documents trustDialogNextKey(), which is part
of a sibling session's in-flight fix for the Claude Code 2.1.252 layout change
(unnumbered, reversed options with "No, exit" highlighted, so a blind carriage
return picks exit and kills the pane). That fix was uncommitted in the shared
tree when this landed, so the doc leads the code until it is committed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NMN8UuvdBim3iM87reuQ9Z
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.
That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.
Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
Format check failed twice, on different files each time, because three prettier
versions were in play: package.json says ^3.4.0, package-lock pins 3.8.3 (CI runs
npm ci, so that is the one CI uses), and the local node_modules had 3.9.6. Files
formatted with 3.9.6 were then "fixed" with 3.4.2, pushing session-routes and
system-routes onto a third style — every version change moved the failure to a
different set of files.
Line-break placement in `await import` and a union type only; no logic changes.
The run menu still offered every mode for an attached container. The browser's
actual request showed why:
POST /api/docker-cases/adopt-preflight -> 400
{"error":"Invalid input: expected object, received string"}
_api serializes `body` and sets Content-Type itself, and three call sites each
passed an already-stringified body, so it was encoded twice and the server saw a
JSON string where it expects an object. curl was fine throughout, so nothing in
the server logs pointed at it.
Also fixes the design defect underneath: a failed probe fell through to "do not
gate", which silently offered every mode. When the container has been recreated,
is stopped, or the engine is unreachable, the user sees claude, clicks it, and
it can only fail — with the reason visible nowhere. A failed probe now hides
every agent mode (Shell needs no CLI and stays) and shows the server's own
reason at the top of the menu.
Two static guards switched from a character window to brace matching. They
sliced between two call sites, and _loadRunModeHistory's call appears above its
definition, so the slice came out empty and the assertion verified nothing —
the same trap twice in one file.
Both paths in the adoption form had to be typed. Each gets a Browse button
using the same path-input-group markup Link Existing uses, so the two look and
behave alike.
What they can browse differs, and that is the point. The host workspace path
reuses the existing host picker. The container workdir cannot: an adopted
container has nothing mounted at a matching host path, so a host listing would
be a different filesystem — and getting this field wrong is the source of the
opaque OCI chdir error at launch, which makes it the field that most needs to
be clickable.
Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no
writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks
directories with a trailing slash and keeps names with spaces intact.
PathPicker takes an optional fetchListing source rather than being forked: the
container variant only swaps where the rows come from, and reuses the rendering,
navigation, Up and Choose/Select unchanged.
Link Existing's Browse did nothing: GET /api/filesystem/browse answered 403
"No filesystem browse roots are available".
Two rules were fighting. /root is a default blocked tree in the attachment
guard, and Codeman running as root — containers, plenty of servers — makes
homedir() exactly /root, so the picker's own allowlisted Home root was blocked;
the other candidates live under it or do not exist. The root list came out
empty and there was nothing the user could open.
The blocked trees exist to keep ~/.ssh and friends out of reach, not to seal off
the user's own home. Only trees that would swallow a configured root whole are
dropped now: /root goes when Home is it (or sits inside it), /etc holds no
configured root and is untouched. Secrets stay protected — isSensitivePath
independently matches .ssh/, .env and credentials* at any depth, and it is what
the directory probe asks about.
⚠️ Navigation must reuse the same narrowed list the roots were chosen with.
Handing the raw trees downstream admits a root and then refuses every path
inside it, which reads as a picker that opens and does nothing.
Attaching a container, picking claude and hitting Run gave one line —
`execvp(3) failed.: No such file or directory` — and the run-mode menu offered
every mode. Three separate defects, found on a real deployment.
TmuxManager.createSession resolved the CLI directory without distinguishing a
docker session, so a host with no claude threw, the catch fell back to a direct
PTY, and that PTY exec'd the CLI on the HOST. The failure surfaced as a bare
execvp error naming nothing. A docker session runs its CLI inside the container;
the host does not need it. All eight modes now sit behind a cliRunsInContainer
guard, and whether the container has the CLI is settled by the adoption
preflight or the image gate before launch.
The running check used a bare double quote and command substitution. The whole
chain is embedded in an outer `bash -c "…"`, so the unescaped quote closed that
string early and the remainder was re-tokenized. It is now a `grep -qx` pipeline
using only the single-quote form every other line in the builder already uses.
Claude Code refuses --dangerously-skip-permissions as root. Our base image runs
a non-root user, so an owned container never hit this; an adopted container's
user belongs to its owner and is frequently root, and keeping the flag killed
the pane with a message visible only inside the container. The preflight now
reports runsAsRoot and the launch chain drops the flag for it.
The menu also showed every mode because the container CLI probe only started
when the menu opened. It is warmed when the case is selected instead.
Storing the container's CLIs on the case at attach time left two gaps: a case
linked before that field existed has none at all, and a container's CLIs can be
installed or removed long after it was linked. A real deployment hit the first
one — the host had only codex, the container only claude, and with no stored
list the menu still gated on the host and hid the mode that actually worked.
The probe now runs when a container case is selected, reusing the existing
adopt-preflight endpoint, so there is no new backend surface. Results are cached
per case for the page's lifetime, since the menu opens often and the probe is a
`docker exec` round trip; a concurrent probe for the same case is deduplicated
with an in-flight marker.
A failed probe leaves the cache empty, which the caller reads as "unknown" and
therefore does not gate. Hiding every mode because one probe failed is worse
than offering one that turns out to be missing, which the launch path already
refuses with a specific message.
The repaint only happens while the menu is still open, so a late answer cannot
make the list jump under a user who already closed it.
The adoption preflight used the mode name as the binary name. claude, codex,
opencode, gemini and pi happen to match, so it never showed — but antigravity
ships as `agy` and deepseek as `dsh`, so a container that has either was
reported as not having it, and the mode was silently dropped from the case.
Adds a MODE_BINARIES map, single-sourced with defaultDockerCommandForMode, which
launches those same binaries. Probing and result filtering share one `binaryFor`
so the two cannot drift apart.
The run-mode dropdown hides CLIs that are not installed on the HOST (#201). That
is right for local sessions and wrong for a container case, whose agents run
inside the container: a host with no claude installed hides the mode while the
container ships one, which is exactly what happened on a real deployment.
The adoption preflight already probes what the container has, so that result is
persisted on the case and surfaced through CaseInfo. Docker cases gate on it;
every other case keeps the host probe unchanged.
An absent list reads as "do not gate" rather than "nothing available": an owned
container runs our base image, which ships every CLI, and treating unknown as
empty would leave the menu with Shell alone.
Typing a container name from memory is error-prone. The field becomes a native
datalist: pick from the engine's containers, type to filter, or type a name that
is not listed (the engine may be remote, or the container may not exist yet).
A datalist gives all three natively, so no dropdown state machine is introduced.
Adds listDockerContainers and GET /api/docker-hosts/:hostId/containers, following
the listRemoteCodemanSessions discovery precedent: read-only and never throwing,
so an unreachable daemon returns an empty list and the field degrades to plain
text instead of erroring.
Stopped containers stay in the list, sorted after running ones and labelled.
Attaching does require a running container, but hiding stopped ones turns "my
container is not in the list" into a dead end, while showing
`Exited (137) 8 days ago` says exactly what to fix.
The new strings were English only. Adding entries surfaced a deeper problem: the
translator matches whole text nodes and skips `code`/`pre`, so an inline `<code>`
mid-sentence splits a hint into fragments that can never match an entry — which is
why the panel's existing "Build it once with <code>...</code>" hint was never
translated either.
Drops the inline markup from the new hints so each is a single text node, then
adds the zh-CN entries. The brand name goes through the existing {name}
placeholder.
Server-side error bodies are deliberately not added: the client receives them
already interpolated with a concrete container name, so a template key could
never match.
Attaching lived only on the Docker tab, but the place users look for anything
container-shaped is the "Run in an isolated Docker container" checkbox on Create
New. A feature nobody can find is a feature nobody has.
Adds a one-click link there that switches to the Docker tab, turns the toggle on
and focuses the container field. Reuses switchCaseModalTab and the existing sync
helper; no new CSS.
Two defects that only a real container exposes.
The probe chained `command -v X && echo X` with semicolons, and a script's exit
status is its last command's. A container without the last probed CLI made the
whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was
reported as "could not exec into the container". A missing CLI is data here, not
failure, so the script now ends with `exit 0`.
containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned
container only because the create-time bind mount puts the host directory at that
exact path; attaching mounts nothing, so the two are independent facts. A host
path absent inside the container makes `docker exec --workdir` fail with an OCI
chdir error that surfaces in the pane as a bare "execvp failed". The preflight now
proves the directory exists inside the container and refuses at link time.