diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 0715c2ff..f42b844d 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -58,7 +58,7 @@ Model is NOT a session field: it is a composition entry in the profile's config ### Docker cases -**Docker cases** (shipped 1.4.0; user guide `docs/docker-cases.md`, design `docs/docker-cases-plan.md`): a case can point at a **container** instead of a local/remote path, and any of the CLI run modes runs INSIDE it. Like remote-SSH, it is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own** (`SessionMode` is unchanged). Storage `~/.codeman/docker-hosts.json` + `docker-cases.json` via `src/docker-hosts.ts` (direct mirror of `remote-hosts.ts`: `readDockerHosts`/`readDockerCases`, `toSessionDocker`, `dockerDisplayPath`, and the PURE builders `buildDockerBaseArgs`/`buildDockerCreateArgs`/`containerApiUrl`/`hostGatewayAlias`/`dockerConfigHash`). CRUD `/api/docker-hosts` + `/api/cases/docker-link`, plus **one-click** `/api/cases/docker-quickcreate` (Create New "Run in Docker" checkbox → case folder in `CASES_DIR` + auto-provisioned shared `default` host + auto-start a session inside; an expandable Template picker Small/Medium/Large/GPU or any override creates a per-case `q-` host), and export/import (`/api/docker-cases/:name/export`, `/api/docker-cases/import`, `GET/DELETE /api/docker-exports`) — all in `case-routes.ts`. Run flows route through `POST /api/quick-start` like remote (session-routes.ts docker branch, skips LOCAL CLI-availability gates). **Launch model**: exactly one long-lived container **per case** (`codeman-case-`, PID1 `sleep infinity` under `--init`); a LOCAL tmux pane runs `docker exec -it` into a **durable in-container tmux** on dedicated socket `-L codeman-docker`, session `codeman-dkr-` (deliberately fails `SAFE_MUX_NAME_PATTERN` so a Codeman running INSIDE the container never adopts it, exactly like remote's `codeman-ssh-`). Builders `buildDockerLaunchCommand`/`buildDockerKillCommand` in `tmux-manager.ts` (image-check → `docker inspect||create` → start → exec, all idempotent). The container is **shared by all sessions of the case**: `buildDockerKillCommand` kills ONLY that session's in-container tmux session, NEVER `docker stop` while siblings remain; `docker rm -f` happens only on case-delete (plus an instance-scoped boot reaper keyed on the `codeman.instance` label). **Two-layer durability/resume** (the central design point): (1) Codeman-PROCESS restart with the container still up → `tmux new-session -A` reattaches the SAME live agent (paneCommand ignored); (2) container stop/reboot/OOM → inner tmux is gone, so the re-run pane command resumes the conversation from the bind-mounted transcript: claude mode pins a DETERMINISTIC conversation id via `claudeDockerPaneCommand()` (`tmux-manager.ts`) — fresh launch `claude --session-id || claude --resume ` (a duplicate `--session-id` exits 1 "already in use", so the fallback RESUMES after a container stop; verified CLI behavior), explicit resume `--resume || --session-id ` so a stale id never dead-panes (leading `exec ` is stripped — an exec'd first branch could never fall back); codex `resume ` / gemini `--resume` keep `appendResumeFlag`. The resume id rides `resumeSessionId` through create/respawn options and persists on `DockerCase.lastClaudeSessionId` via `persistDockerCaseClaudeSessionId()` (written at quick-start launch, and again on hook/last-response conversation-id adoption so post-`/clear` switches track; seeded back when `resumeOnStart`, default true); `-A` makes the pane command self-selecting (inert on reattach, active only when tmux was re-created). **Config drift** (`dockerConfigHash` → `codeman.confighash` label): quick-start compares via `checkDockerConfigDrift()` and REFUSES a drifted launch with `CONFLICT`; the UI confirm calls `POST /api/docker-cases/:name/recreate` (refused while case sessions are live) which `docker rm -f`s so the next launch recreates with the new config — host config edits actually take effect. **Workspace** is a REAL host dir bind-mounted at the SAME absolute path (mirror, `dst==src`), so `Session.workingDir = hostWorkspacePath` keeps file-routes/attachments/watchers on real host bytes AND the in-container transcript projHash matches the host so subagent/workflow correlation (and thus resume-id capture) works; `resolveMuxAttachCwd` returns `/tmp` for docker (the local pane only runs `docker exec`). **Creds** arrive commit-safe and ISOLATED (1.4.1; replaced the whole-dir RW mounts that let in-container CLIs write refreshed tokens/state back to the host): shared RW across the boundary is ONLY what host-side reads/resume need (`~/.claude/projects` transcripts; codex `sessions/` + `history.jsonl` for response-viewer/`codex resume`); everything else is SEEDED (RO mount, copied into container HOME once at launch via `[ -e ] || cp`; the container refreshes its own copy and never writes back): `~/.claude.json` is merged through `buildSeamlessClaudeConfig()` (forces `hasCompletedOnboarding` + theme + workspace trust, so no login wizard/theme picker/trust prompt inside the container), plus `.claude/{.credentials.json,settings.json,stats-cache.json}`, plus whole-dir seeds for `~/.gemini`/`~/.config/{gcloud,opencode}` (`resolveDockerClaudeArtifacts`/`resolveDockerCredentialArtifacts` in `docker-hosts.ts`). Bind mounts are physically excluded from `docker commit`, so exports stay secret-free; API-key CLIs get exec-time NAME-ONLY `--env OPENAI_API_KEY` (no `=value`); the SEALED profile is `mountCredentials:false` + `network:none`. NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket. **Hardening** on every create: `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit`, `--memory`==`--memory-swap`, non-root via `--user :0` (Linux, GID 0 for writable HOME) / `--userns=keep-id` (podman rootless) / baked uid (Docker Desktop), `--pull=never`, `--init`. Base image `codeman/agent:base` is BUILT LOCALLY from `docker/agent.Dockerfile` (node22 + tmux + claude/codex/gemini/opencode, OpenShift arbitrary-uid HOME, `C.UTF-8` locale so tmux/Ink render real box-drawing glyphs; Codeman also sets `LANG`/`LC_ALL` at run time for containers built before that line) via `scripts/build-agent-image.mjs` OR **auto-built on first use** (1.4.1: `ensureAgentBaseImage()` in `docker-hosts.ts`; idempotent + concurrency-safe, only the DEFAULT image ref is ever auto-built, `--pull=never` stays absolute; build output streams over SSE `docker:imageBuildStarted`/`imageBuildProgress`/`imageBuildComplete`/`imageBuildFailed`, and quick-create returns `imageBuilding:true` while the first launch awaits the gate); tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent bare-exec fallback. **Hooks + model**: the workspace-scaffolding block DOES run for docker (writes `.claude/settings.local.json` + the CLAUDE.md scaffold into the real host dir), so `modelOverride` works via `settings.local.json` — it is a `QuickStartSchema` field applied for local AND docker quick-starts (`updateCaseModel`), sent by the frontend docker run path (the one deliberate difference from remote, which rejects it); `effort`/`envOverrides`/`codexConfig`/`geminiConfig`/`openCodeConfig` stay rejected. In-container hook curls hit `containerApiUrl(process.env.CODEMAN_API_URL, engine)` (swaps ONLY the hostname to the gateway alias, preserving scheme+port so prod HTTPS still works); the host guard allowlists both `host.docker.internal`/`host.containers.internal` (`DOCKER_HOST_GATEWAY_ALIASES` in `network-auth-policy.ts`). ⚠️ On a **loopback-only** bind (the prod default) a container cannot reach 127.0.0.1, so in-container hooks fire ONLY when `CODEMAN_DOCKER_BRIDGE_HOOKS=1` — an opt-in SECOND listener on the docker bridge gateway (`_startDockerBridgeHooksListener` in `server.ts`; gateway auto-detected via `detectDockerBridgeGateway`, or set `CODEMAN_DOCKER_BRIDGE_HOST`) that serves ONLY the hook endpoints (403 for any other path) into the same secret-gated pipeline; otherwise idle detection falls back to output-based through the docker-exec PTY. Container-set `CLAUDE_CODE_TMPDIR` keeps claude launching regardless of workspace path. `SessionState.docker`/`MuxSession.docker` round-trip through recovery. Every docker IO path is `IS_TEST_MODE` (VITEST) no-op'd; the pure builders are unit-tested. **Export/import** (`src/docker-export.ts`): full-image (`docker commit` + `save | gzip` + workspace tar + manifest) or workspace-only → one portable `~/.codeman/docker-exports/-.codeman-container.tgz`; import validates per-member sha256, traversal-guards the workspace tar, `docker load`s + quarantine-retags the image (`codeman/imported-:`, never overwriting a local tag); a `saveImageToTar` stream `pipeline` avoids truncation. **GPU** passthrough (`gpus` → `--gpus`, needs the NVIDIA container toolkit) and **elastic disk** (no `--storage-opt` cap, so container storage grows with data). SSE `docker:exportComplete`/`exportFailed`/`importComplete` (both registries). **UI** in `session-ui.js`: Create Case **Docker** tab (collapsed/compact form since 1.4.1), the one-click checkbox + Template picker, short `(docker)` case-menu tags, and a Manage-tab Export button; docker AND remote sessions name their tabs `w-` via the shared `_nextCaseSessionStartNumber()` so all tabs follow one naming convention. **Adopting an already-running container** (`DockerCase.owned === false`, `POST /api/cases/docker-adopt` + the read-only `POST /api/docker-cases/adopt-preflight`, `GET /api/docker-hosts/:hostId/containers`, `POST /api/docker-cases/browse`): the mirror of remote-SSH's `owned:false` attach. For an adopted container the launch chain only LOOKS and then execs — no image gate (the image is theirs), no create, and above all no `start`, since starting a container we do not own is precisely the mutation adoption promises never to perform; a missing or stopped container fails closed with an actionable message. Credential seeding is skipped too (those copies read from create-time read-only mounts that do not exist here, and writing host credentials into someone's container is not ours to do), so its CLIs must already be authenticated inside it. Absent `owned` = owned, so every pre-existing case is byte-identical. ⚠️ The guarantee is NEGATIVE, so it cannot be observed by using the feature — only by asserting the mutating verbs are absent — and it is therefore enforced at four deliberately independent layers: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION (no shape of caller bug can produce a `docker stop`/`rm` for a container we do not own), `removeDockerContainer` refuses again at the lowest layer, `checkDockerConfigDrift` reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would always report drift and the launch gate would 409 forever, offering a recreate we may not perform), and the orphan reaper skips it through a check independent of the two conditions that already cover it. ⚠️ **The export path is the one place that still touched the container** and both halves had to be closed: a full export `docker commit`s it (refused for an adopted case — it packages someone else's container, with their logins, into a bundle Codeman hands out) and even a workspace-only export `docker pause`d it first for snapshot consistency (skipped: the freeze stops the owner's processes for as long as the tar takes). ⚠️ `owned` is applied AFTER the config hash; `dockerConfigHash` takes an explicit field list, so ownership can never shift an existing case's hash and mass-trip the drift gate, whose only remedy is "recreate the container". ⚠️ The container workdir is verified INSIDE the container: it defaults to `hostWorkspacePath` for an OWNED case only because the create-time bind mount puts the host directory at that exact path, and adoption mounts nothing, so the two are independent facts — without the check `docker exec --workdir ` fails with an OCI chdir error the pane surfaces as a bare `execvp failed`. ⚠️ Run-mode availability comes from the CONTAINER (`availableModes`), live-probed rather than trusted from attach time: gating the dropdown on host CLI availability (#201) is right for local sessions and wrong here. The probe modes and the BINARY each mode looks for both come from the CLI registry (`enabledCliIds()` / `discovery.binaries[0]`), never a local table — a hand-written list silently froze once already, missing `omp` and hiding that mode on every docker case; `antigravity` ships as `agy` and `deepseek` as `dsh`, so a mode-name probe reports both missing on a container that has them, and a mode with no binary (`shell`) is reported available without a lookup. ⚠️ **A FAILED probe means opposite things per ownership.** For an adopted case it is a real fault (only the user can start that container). For an owned case it is the NORMAL state before the first session — the launch chain creates the container on demand — so recording it as an error hid every agent mode on every freshly linked Docker case behind "start it yourself first", for a container Codeman was about to create itself; `CaseInfo.docker.owned` exists on the wire so the frontend can tell the two apart. ⚠️ Claude launches WITHOUT `--dangerously-skip-permissions` when the container's exec user is root: Claude Code refuses the flag as root ("cannot be used with root/sudo privileges", still true in 2.1.261) and the refusal is visible only inside the container, so the pane just dies. Our base image runs a non-root user and never hits it; an adopted container's user belongs to its owner and is frequently root. Which flag to drop is a per-CLI fact, so it is `overlays.docker.rootCommand` in the registry rather than an id branch. ⚠️ **Admin-only in multi-user mode**, unlike `docker-link` right next to it: linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has already confined to the caller's space, while an adopted container's mounts are whatever its owner gave it — one mounting `/` hands the adopter a shell over the whole host, defeating exactly the workspace scoping that mode exists to enforce. The container listing and the in-container directory browser are gated with it (both are machine-level reads over containers belonging to anyone); the preflight is NOT, because the run menu probes it for every docker case, so it admits a non-admin only for a container already linked to a case they can access. Tests: `test/docker-adopted-container.test.ts`. +**Docker cases** (shipped 1.4.0; user guide `docs/docker-cases.md`, design `docs/docker-cases-plan.md`): a case can point at a **container** instead of a local/remote path, and any of the CLI run modes runs INSIDE it. Like remote-SSH, it is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own** (`SessionMode` is unchanged). Storage `~/.codeman/docker-hosts.json` + `docker-cases.json` via `src/docker-hosts.ts` (direct mirror of `remote-hosts.ts`: `readDockerHosts`/`readDockerCases`, `toSessionDocker`, `dockerDisplayPath`, and the PURE builders `buildDockerBaseArgs`/`buildDockerCreateArgs`/`containerApiUrl`/`hostGatewayAlias`/`dockerConfigHash`). CRUD `/api/docker-hosts` + `/api/cases/docker-link`, plus **one-click** `/api/cases/docker-quickcreate` (Create New "Run in Docker" checkbox → case folder in `CASES_DIR` + auto-provisioned shared `default` host + auto-start a session inside; an expandable Template picker Small/Medium/Large/GPU or any override creates a per-case `q-` host), and export/import (`/api/docker-cases/:name/export`, `/api/docker-cases/import`, `GET/DELETE /api/docker-exports`) — all in `case-routes.ts`. Run flows route through `POST /api/quick-start` like remote (session-routes.ts docker branch, skips LOCAL CLI-availability gates). **Launch model**: exactly one long-lived container **per case** (`codeman-case-`, PID1 `sleep infinity` under `--init`); a LOCAL tmux pane runs `docker exec -it` into a **durable in-container tmux** on dedicated socket `-L codeman-docker`, session `codeman-dkr-` (deliberately fails `SAFE_MUX_NAME_PATTERN` so a Codeman running INSIDE the container never adopts it, exactly like remote's `codeman-ssh-`). Builders `buildDockerLaunchCommand`/`buildDockerKillCommand` in `tmux-manager.ts` (image-check → `docker inspect||create` → start → exec, all idempotent). The container is **shared by all sessions of the case**: `buildDockerKillCommand` kills ONLY that session's in-container tmux session, NEVER `docker stop` while siblings remain; `docker rm -f` happens only on case-delete (plus an instance-scoped boot reaper keyed on the `codeman.instance` label). **Two-layer durability/resume** (the central design point): (1) Codeman-PROCESS restart with the container still up → `tmux new-session -A` reattaches the SAME live agent (paneCommand ignored); (2) container stop/reboot/OOM → inner tmux is gone, so the re-run pane command resumes the conversation from the bind-mounted transcript: claude mode pins a DETERMINISTIC conversation id via `claudeDockerPaneCommand()` (`tmux-manager.ts`) — fresh launch `claude --session-id || claude --resume ` (a duplicate `--session-id` exits 1 "already in use", so the fallback RESUMES after a container stop; verified CLI behavior), explicit resume `--resume || --session-id ` so a stale id never dead-panes (leading `exec ` is stripped — an exec'd first branch could never fall back); codex `resume ` / gemini `--resume` keep `appendResumeFlag`. The resume id rides `resumeSessionId` through create/respawn options and persists on `DockerCase.lastClaudeSessionId` via `persistDockerCaseClaudeSessionId()` (written at quick-start launch, and again on hook/last-response conversation-id adoption so post-`/clear` switches track; seeded back when `resumeOnStart`, default true); `-A` makes the pane command self-selecting (inert on reattach, active only when tmux was re-created). **Config drift** (`dockerConfigHash` → `codeman.confighash` label): quick-start compares via `checkDockerConfigDrift()` and REFUSES a drifted launch with `CONFLICT`; the UI confirm calls `POST /api/docker-cases/:name/recreate` (refused while case sessions are live) which `docker rm -f`s so the next launch recreates with the new config — host config edits actually take effect. **Workspace** is a REAL host dir bind-mounted at the SAME absolute path (mirror, `dst==src`), so `Session.workingDir = hostWorkspacePath` keeps file-routes/attachments/watchers on real host bytes AND the in-container transcript projHash matches the host so subagent/workflow correlation (and thus resume-id capture) works; `resolveMuxAttachCwd` returns `/tmp` for docker (the local pane only runs `docker exec`). **Creds** arrive commit-safe and ISOLATED (1.4.1; replaced the whole-dir RW mounts that let in-container CLIs write refreshed tokens/state back to the host): shared RW across the boundary is ONLY what host-side reads/resume need (`~/.claude/projects` transcripts; codex `sessions/` + `history.jsonl` for response-viewer/`codex resume`); everything else is SEEDED (RO mount, copied into container HOME once at launch via `[ -e ] || cp`; the container refreshes its own copy and never writes back): `~/.claude.json` is merged through `buildSeamlessClaudeConfig()` (forces `hasCompletedOnboarding` + theme + workspace trust, so no login wizard/theme picker/trust prompt inside the container), plus `.claude/{.credentials.json,settings.json,stats-cache.json}`, plus whole-dir seeds for `~/.gemini`/`~/.config/{gcloud,opencode}` (`resolveDockerClaudeArtifacts`/`resolveDockerCredentialArtifacts` in `docker-hosts.ts`). Bind mounts are physically excluded from `docker commit`, so exports stay secret-free; API-key CLIs get exec-time NAME-ONLY `--env OPENAI_API_KEY` (no `=value`); the SEALED profile is `mountCredentials:false` + `network:none`. NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket. **Hardening** on every create: `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit`, `--memory`==`--memory-swap`, non-root via `--user :0` (Linux, GID 0 for writable HOME) / `--userns=keep-id` (podman rootless) / baked uid (Docker Desktop), `--pull=never`, `--init`. Base image `codeman/agent:base` is BUILT LOCALLY from `docker/agent.Dockerfile` (node22 + tmux + every enabled npm CLI from the registry, since the `CLI_NPM_PACKAGES` build arg is generated from `stock.ts` and entries carrying `agentImageLayer` or no `npmPackage` get their own layers, see `docs/docker-cases.md`; OpenShift arbitrary-uid HOME, `C.UTF-8` locale so tmux/Ink render real box-drawing glyphs; Codeman also sets `LANG`/`LC_ALL` at run time for containers built before that line) via `scripts/build-agent-image.mjs` OR **auto-built on first use** (1.4.1: `ensureAgentBaseImage()` in `docker-hosts.ts`; idempotent + concurrency-safe, only the DEFAULT image ref is ever auto-built, `--pull=never` stays absolute; build output streams over SSE `docker:imageBuildStarted`/`imageBuildProgress`/`imageBuildComplete`/`imageBuildFailed`, and quick-create returns `imageBuilding:true` while the first launch awaits the gate); tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent bare-exec fallback. **Hooks + model**: the workspace-scaffolding block DOES run for docker (writes `.claude/settings.local.json` + the CLAUDE.md scaffold into the real host dir), so `modelOverride` works via `settings.local.json` — it is a `QuickStartSchema` field applied for local AND docker quick-starts (`updateCaseModel`), sent by the frontend docker run path (the one deliberate difference from remote, which rejects it); `effort`/`envOverrides`/`codexConfig`/`geminiConfig`/`openCodeConfig` stay rejected. In-container hook curls hit `containerApiUrl(process.env.CODEMAN_API_URL, engine)` (swaps ONLY the hostname to the gateway alias, preserving scheme+port so prod HTTPS still works); the host guard allowlists both `host.docker.internal`/`host.containers.internal` (`DOCKER_HOST_GATEWAY_ALIASES` in `network-auth-policy.ts`). ⚠️ On a **loopback-only** bind (the prod default) a container cannot reach 127.0.0.1, so in-container hooks fire ONLY when `CODEMAN_DOCKER_BRIDGE_HOOKS=1` — an opt-in SECOND listener on the docker bridge gateway (`_startDockerBridgeHooksListener` in `server.ts`; gateway auto-detected via `detectDockerBridgeGateway`, or set `CODEMAN_DOCKER_BRIDGE_HOST`) that serves ONLY the hook endpoints (403 for any other path) into the same secret-gated pipeline; otherwise idle detection falls back to output-based through the docker-exec PTY. Container-set `CLAUDE_CODE_TMPDIR` keeps claude launching regardless of workspace path. `SessionState.docker`/`MuxSession.docker` round-trip through recovery. Every docker IO path is `IS_TEST_MODE` (VITEST) no-op'd; the pure builders are unit-tested. **Export/import** (`src/docker-export.ts`): full-image (`docker commit` + `save | gzip` + workspace tar + manifest) or workspace-only → one portable `~/.codeman/docker-exports/-.codeman-container.tgz`; import validates per-member sha256, traversal-guards the workspace tar, `docker load`s + quarantine-retags the image (`codeman/imported-:`, never overwriting a local tag); a `saveImageToTar` stream `pipeline` avoids truncation. **GPU** passthrough (`gpus` → `--gpus`, needs the NVIDIA container toolkit) and **elastic disk** (no `--storage-opt` cap, so container storage grows with data). SSE `docker:exportComplete`/`exportFailed`/`importComplete` (both registries). **UI** in `session-ui.js`: Create Case **Docker** tab (collapsed/compact form since 1.4.1), the one-click checkbox + Template picker, short `(docker)` case-menu tags, and a Manage-tab Export button; docker AND remote sessions name their tabs `w-` via the shared `_nextCaseSessionStartNumber()` so all tabs follow one naming convention. **Adopting an already-running container** (`DockerCase.owned === false`, `POST /api/cases/docker-adopt` + the read-only `POST /api/docker-cases/adopt-preflight`, `GET /api/docker-hosts/:hostId/containers`, `POST /api/docker-cases/browse`): the mirror of remote-SSH's `owned:false` attach. For an adopted container the launch chain only LOOKS and then execs — no image gate (the image is theirs), no create, and above all no `start`, since starting a container we do not own is precisely the mutation adoption promises never to perform; a missing or stopped container fails closed with an actionable message. Credential seeding is skipped too (those copies read from create-time read-only mounts that do not exist here, and writing host credentials into someone's container is not ours to do), so its CLIs must already be authenticated inside it. Absent `owned` = owned, so every pre-existing case is byte-identical. ⚠️ The guarantee is NEGATIVE, so it cannot be observed by using the feature — only by asserting the mutating verbs are absent — and it is therefore enforced at four deliberately independent layers: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION (no shape of caller bug can produce a `docker stop`/`rm` for a container we do not own), `removeDockerContainer` refuses again at the lowest layer, `checkDockerConfigDrift` reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would always report drift and the launch gate would 409 forever, offering a recreate we may not perform), and the orphan reaper skips it through a check independent of the two conditions that already cover it. ⚠️ **The export path is the one place that still touched the container** and both halves had to be closed: a full export `docker commit`s it (refused for an adopted case — it packages someone else's container, with their logins, into a bundle Codeman hands out) and even a workspace-only export `docker pause`d it first for snapshot consistency (skipped: the freeze stops the owner's processes for as long as the tar takes). ⚠️ `owned` is applied AFTER the config hash; `dockerConfigHash` takes an explicit field list, so ownership can never shift an existing case's hash and mass-trip the drift gate, whose only remedy is "recreate the container". ⚠️ The container workdir is verified INSIDE the container: it defaults to `hostWorkspacePath` for an OWNED case only because the create-time bind mount puts the host directory at that exact path, and adoption mounts nothing, so the two are independent facts — without the check `docker exec --workdir ` fails with an OCI chdir error the pane surfaces as a bare `execvp failed`. ⚠️ Run-mode availability comes from the CONTAINER (`availableModes`), live-probed rather than trusted from attach time: gating the dropdown on host CLI availability (#201) is right for local sessions and wrong here. The probe modes and the BINARY each mode looks for both come from the CLI registry (`enabledCliIds()` / `discovery.binaries[0]`), never a local table — a hand-written list silently froze once already, missing `omp` and hiding that mode on every docker case; `antigravity` ships as `agy` and `deepseek` as `dsh`, so a mode-name probe reports both missing on a container that has them, and a mode with no binary (`shell`) is reported available without a lookup. ⚠️ **A FAILED probe means opposite things per ownership.** For an adopted case it is a real fault (only the user can start that container). For an owned case it is the NORMAL state before the first session — the launch chain creates the container on demand — so recording it as an error hid every agent mode on every freshly linked Docker case behind "start it yourself first", for a container Codeman was about to create itself; `CaseInfo.docker.owned` exists on the wire so the frontend can tell the two apart. ⚠️ Claude launches WITHOUT `--dangerously-skip-permissions` when the container's exec user is root: Claude Code refuses the flag as root ("cannot be used with root/sudo privileges", still true in 2.1.261) and the refusal is visible only inside the container, so the pane just dies. Our base image runs a non-root user and never hits it; an adopted container's user belongs to its owner and is frequently root. Which flag to drop is a per-CLI fact, so it is `overlays.docker.rootCommand` in the registry rather than an id branch. ⚠️ **Admin-only in multi-user mode**, unlike `docker-link` right next to it: linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has already confined to the caller's space, while an adopted container's mounts are whatever its owner gave it — one mounting `/` hands the adopter a shell over the whole host, defeating exactly the workspace scoping that mode exists to enforce. The container listing and the in-container directory browser are gated with it (both are machine-level reads over containers belonging to anyone); the preflight is NOT, because the run menu probes it for every docker case, so it admits a non-admin only for a container already linked to a case they can access. Tests: `test/docker-adopted-container.test.ts`. Tests: `test/docker-hosts.test.ts`, `test/docker-exec-options.test.ts`, `test/docker-export.test.ts`, `test/network-host-guard.test.ts`. diff --git a/docs/cli-registry.md b/docs/cli-registry.md index d60d7319..e684b611 100644 --- a/docs/cli-registry.md +++ b/docs/cli-registry.md @@ -156,9 +156,9 @@ Three rules, and the middle one is why the embed matters: That is mechanical rather than a promise. `CLI_INSTALL_CMD_TRUSTED` is written only from the generated block and is the only array the installer ever runs or displays — there is no second -array a refresh could rewrite, because there is no refresh. `test/install-sh-invariants.test.ts` -asserts as much: the embedded commands are exactly the registry's, and nothing in `install.sh` -`eval`s. +array a refresh could rewrite, because there is no refresh. `test/cli-catalog-sync.test.ts` +asserts that the embedded commands are exactly the registry's, and +`test/install-sh-invariants.test.ts` that nothing in `install.sh` `eval`s. ### bash 3.2 @@ -181,7 +181,7 @@ A module-level const freezes at first import, and the failure is asymmetric: a C 1. Add a `CliEntry` to `stock.ts`. 2. Run `npm run generate:cli-catalog` and commit **both** artifacts (`config/clis.stock.json` and `install.sh`). The installer's detection, its install menu, its reminder text and the Docker agent image all follow from that one step — this is what makes upstream `b6d0f1fa` ("wire OMP into install.sh's CLI detection, it had none") impossible rather than merely fixed. 3. Add a golden spawn-command pin to `test/cli-registry-spawn-golden.test.ts`, a row to `test/cli-capability-predicates.test.ts`, its remote/docker commands to `test/location-overlay-commands.test.ts`, and its search paths to `test/install-sh-detection-parity.test.ts`. -4. Only if it cannot install with a plain `npm install -g `: give it a layer in `docker/agent.Dockerfile` and a reason in `AGENT_IMAGE_SPECIAL_CASES` (`scripts/lib/cli-catalog.mjs`). The coverage test requires both, so an exclusion cannot quietly become an omission. +4. Only if it cannot install with a plain `npm install -g `: give it a layer in `docker/agent.Dockerfile` and set `discovery.install.agentImageLayer: { kind: 'dedicated', reason }` on its entry in `stock.ts`. `test/docker-agent-image-coverage.test.ts` requires both, so an exclusion cannot quietly become an omission. An entry with no `npmPackage` needs only the Dockerfile layer, since it never enters the shared npm layer in the first place. 5. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code. ## See also diff --git a/docs/docker-cases.md b/docs/docker-cases.md index f930c126..f4664a3e 100644 --- a/docs/docker-cases.md +++ b/docs/docker-cases.md @@ -34,9 +34,12 @@ must not change what is inside an image tagged `codeman/agent:base`, or two mach that tag hold different images and every cache decision downstream is a lie. Each entry's `enabled` flag IS honoured, so a CLI that ships disabled is never baked in. -Five CLIs keep hand-written layers, because the registry cannot express what makes them -special (as a REGISTRY field now — `discovery.install.agentImageLayer` in `stock.ts` — rather -than an id-keyed table duplicated between the two producers of the image's build args): +Five CLIs keep hand-written layers, for two different reasons that are easy to conflate. +`antigravity`, `grok` and `omp` declare no `npmPackage` at all, so they never enter the shared +npm layer and each gets a vendor-installer layer instead. `pi` and `deepseek` ARE on npm but +carry `discovery.install.agentImageLayer` in `stock.ts` (a REGISTRY field, rather than an +id-keyed table duplicated between the two producers of the image's build args), which pulls +them out of the shared layer because a plain `npm install -g` is not enough for them: | CLI | Why it is not in the shared npm layer | | ------------- | ------------------------------------------------------------------------------------- | diff --git a/install.sh b/install.sh index 18920df5..5be359c2 100755 --- a/install.sh +++ b/install.sh @@ -452,8 +452,9 @@ dsh_banner_probe() { # DeepSeek stays a hand-written special case ON PURPOSE: the registry expresses # its identity check as `discovery.identity.regex`, a JavaScript regex, and # translating that into a `grep` pattern at install time is a transformation -# nobody should be performing on a security-adjacent check. The parity test pins -# that the registry still demands "DeepSeek Harness", so an upstream banner +# nobody should be performing on a security-adjacent check. Instead +# test/install-sh-invariants.test.ts pins the grep below against the registry's +# `discovery.identity.regex`, so the two cannot drift apart: an upstream banner # change fails a test instead of silently mis-detecting here. _cli_candidate_ok() { case "$1" in diff --git a/test/install-sh-invariants.test.ts b/test/install-sh-invariants.test.ts index 93ad731d..126c406d 100644 --- a/test/install-sh-invariants.test.ts +++ b/test/install-sh-invariants.test.ts @@ -14,6 +14,7 @@ import { describe, expect, it } from 'vitest'; import { spawnSync } from 'node:child_process'; import { readFileSync } from 'node:fs'; import { fileURLToPath } from 'node:url'; +import { STOCK_CLIS } from '../src/config/cli-registry/stock.js'; const INSTALL_SH = fileURLToPath(new URL('../install.sh', import.meta.url)); const SOURCE = readFileSync(INSTALL_SH, 'utf-8'); @@ -167,6 +168,23 @@ describe('install.sh runtime safety', () => { }); }); +describe('install.sh DeepSeek identity probe', () => { + it('greps for the same banner the registry identity regex demands', () => { + // dsh_banner_probe is the ONE hand-written identity check left in the script (the + // registry's is a JavaScript regex, deliberately not translated into grep at install + // time). The two are pinned to each other here so an upstream banner change fails + // this test instead of mis-detecting on one side only. + const grepLine = CODE_LINES.find((line) => line.includes('grep -qi "DeepSeek Harness"')); + expect(grepLine, 'the dsh banner grep is gone or its literal changed').toBeDefined(); + + const deepseek = STOCK_CLIS.find((entry) => entry.id === 'deepseek'); + const identity = deepseek?.discovery.identity; + expect(identity, 'the deepseek entry no longer declares an identity probe').toBeDefined(); + expect(identity?.arg).toBe('--help'); + expect(new RegExp(identity!.regex, 'i').test('DeepSeek Harness')).toBe(true); + }); +}); + describe('install.sh AI CLI install menu', () => { // The menu is the one interactive path in the script, which is why it used to be the // only part nothing exercised: choosing "s" (Skip) once fell straight into the shared