diff --git a/CLAUDE.md b/CLAUDE.md index f0558d38..80e7bf5b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -217,7 +217,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Docker cases**: a case can point at a **container**, with any of the CLI run modes running inside it. Like remote-SSH this is a **LOCATION OVERLAY on cases, never a `SessionMode` of its own**. Exactly one long-lived container **per case**, shared by all its sessions, so killing a session kills only that session's in-container tmux and **never** `docker stop` while siblings remain. The workspace is a real host dir bind-mounted at the **same absolute path**, which is what keeps file-routes/watchers on real host bytes and makes the in-container transcript projHash match the host. Credentials are **seeded** (RO mount, copied into the container once) rather than shared RW, so in-container CLIs never write refreshed tokens back to the host, and bind mounts are excluded from `docker commit` so exports stay secret-free. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** Config drift is detected via a label hash and a drifted launch is REFUSED rather than silently launched with stale config. ⚠️ A case may instead **ADOPT** a container the user already runs (`DockerCase.owned === false`, mirror of remote-SSH's `owned:false`): Codeman only `exec`s into it and never creates, starts, stops, restarts or removes it, so a missing or stopped container FAILS CLOSED with an actionable message instead of being fixed. Absent = owned, so existing cases are byte-identical. The guarantee is enforced at four independent layers because it cannot be observed by using the feature: `buildDockerStopCommand`/`buildDockerRemoveCommand` throw during pure STRING CONSTRUCTION, `removeDockerContainer` refuses again, drift reports "none" (an adopted container carries no `codeman.confighash` label, so a real comparison would 409 the launch forever), and the boot reaper skips it. ⚠️ Two lifecycle touches the original design missed and that are easy to re-introduce: the full-image export `docker commit`s the container (refused for an adopted case) and the workspace export `docker pause`s it first (skipped — it freezes the owner's processes for the length of the tar). ⚠️ `owned` is applied AFTER `dockerConfigHash`, which takes an explicit field list, or every pre-existing case would trip the drift gate at once. ⚠️ Run modes for a container case come from the CONTAINER (`availableModes`, live-probed): gating the run menu on HOST CLIs (#201) is right for local sessions and wrong here, since a host with no `claude` may run a container that ships one. ⚠️ **A failed probe means opposite things per ownership** — for an ADOPTED case it is a fault worth reporting, for an OWNED one it is the NORMAL state before the first session (the launch chain creates the container), so treating it as a fault hid every agent mode on every freshly linked Docker case behind "start it yourself first". That is why `CaseInfo.docker.owned` is on the wire. ⚠️ Claude is launched WITHOUT `--dangerously-skip-permissions` when the container's exec user is root (Claude Code refuses the flag as root and the refusal is visible only inside the container); which flag to drop is a per-CLI fact, so it is the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is **admin-only in multi-user mode**, unlike `docker-link`: linking creates OUR container, whose one bind mount `isWorkingDirAllowed` has already confined, while an adopted container's mounts belong to its owner and one mounting `/` hands the adopter the host. The same reasoning admin-gates the container listing and the in-container directory browser; the preflight instead admits a non-admin for a container already linked to a case they own, because the run menu probes it for every docker case. ⚠️ On the loopback-only prod bind a container cannot reach 127.0.0.1, so in-container hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; otherwise idle detection falls back to output-based. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` (user guide), `docs/docker-cases-plan.md` (design) -**Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases above). ⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case ` reported "Case not found" on exactly the deployment the override exists for. ⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. ⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing, so any bind source the runtime user must write to has to be pre-created and chowned. `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. ⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update below and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. `docs/docker-compose.md` + `docker/README.md` (user guides) +**Docker Compose deployment** (`docker/`, contributed): Codeman itself runs in a container and spawns Docker cases as **SIBLING** containers through the mounted host socket (Docker-outside-of-Docker), never nested. That inverts one assumption the bare-host path takes for granted: the daemon no longer shares Codeman's filesystem, so a bind source valid *inside* Codeman means nothing to it. `resolveDockerDaemonMountSource()` translates sources under HOME into the daemon's namespace via `CODEMAN_DOCKER_HOST_HOME`, and `CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount so a workspace resolves to the SAME absolute path on both sides (which is what keeps the transcript projHash matching, per Docker cases above). ⚠️ **`CODEMAN_CASES_PATH` must move every consumer or none**: it is resolved once in `config/cases-dir.ts` because `src/cli.ts` resolves case paths too, and when only the server's `CASES_DIR` learned the override, `codeman skill install --case ` reported "Case not found" on exactly the deployment the override exists for. ⚠️ **`.dockerignore` patterns match the WHOLE context-relative path**, so a bare `.env` line excludes only the ROOT file: `docker/.env` (which holds `CODEMAN_PASSWORD` and any provider keys) rode `COPY . .` into the image until `**/.env` was added — verified in both directions with a real build context. ⚠️ A Compose LONG-form bind (`type: bind`) **creates a missing host source directory ROOT-OWNED** rather than refusing. `Start-Codeman.sh` pre-creates both `CODEMAN_APPDATA_PATH` and `CODEMAN_CASES_PATH` on the host before `up`, which is what keeps the daemon from ever having to materialise either as root in the first place; the container ALSO starts as root (`cap_add: [CHOWN, DAC_OVERRIDE, SETGID, SETUID]` against the base `cap_drop: ALL`) so `docker/entrypoint.sh` can correct a bind source that turns up root-owned anyway (a restored backup, a cleared directory, plain `docker compose up` run without the script) before dropping to `PUID:PGID` via `setpriv` — it refuses instead of chowning a directory owned by neither root nor `PUID:PGID`, since that ownership is not this container's to reassign. `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. ⚠️ The deployment ALSO self-updates in place (the repo bind mount at `/opt/codeman` + a restart-by-exiting supervisor) — see Self-update below and `docs/docker-self-update.md` before touching `server.Dockerfile`, the compose file or `.env.example`, since each is an input to the updater's environment gate. `docs/docker-compose.md` + `docker/README.md` (user guides) **CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. **No code outside `stock.ts` may branch on a CLI id**; behaviour that genuinely differs is either a capability field or a NAMED PROFILE selected by one (`profiles.ts`), and `test/cli-registry-no-id-branching.test.ts` fails the build if an id check reappears — it matches `===`, `!==`, `case '':` and `[...].includes(mode)`, because an earlier `===`-only version let 36 negated branches survive the conversion (including a seven-mode Ralph chain whose own comment asked the next person to keep it in step with `isExternalCliMode()` by hand). ⚠️ Config contains no shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. ⚠️ Two capability fields carry a REGEX from config (`discovery.version.regex` and `capabilities.workDetect.workingLine`) and both must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. ⚠️ Six fields are DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/`keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured, so re-measure before wiring one up; the list is pinned so it cannot quietly grow. Spawn commands are pinned as literal strings in `test/cli-registry-spawn-golden.test.ts`, remote/docker pane commands in `test/location-overlay-commands.test.ts`. ⚠️ That second golden no longer covers **remote claude or remote omp**: both now have their own arm in `buildRemoteLaunchCommand` (a `--session-id || --resume` pair, and `--continue`, so a respawn continues the same conversation) and never reach `defaultRemoteCommandForMode`, which is what that test asserts. Their real pins are `toContain` substrings in `test/tmux-manager.test.ts` and `test/remote-shared-sessions.test.ts`; changing either arm will NOT fail the golden. ⚠️ Anything reading the registry resolves it AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks) — a module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. `~/.codeman/clis.json` overrides any entry (read-only in this release; nothing writes it, so importing the registry has no filesystem side effects). → `docs/cli-registry.md` @@ -388,7 +388,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L ## State Files -All in `~/.codeman/`: `state.json` (sessions, settings, respawn, orchestrator, cron jobs/runs, owner tab layouts), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` + `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log), `update-status.json` (self-updater progress, polled across the service restart), `docker-env-applied.json` (Compose deployment only: sha256 of the Dockerfile + compose file the running container was built from, written by `Start-Codeman.sh`, read by the self-updater's environment gate), `linked-cases.json`, `webviews.json` (saved web-tab dashboard URLs), `remote-hosts.json` + `remote-cases.json`, `docker-hosts.json` + `docker-cases.json` + `docker-exports/`, `subagent-window-states.json` + `subagent-parents.json` (subagent window layout, GET/PUT `/api/subagent-window-states`/`-parents`), `hook-secret` (per-instance), `users.json` (multi-user, mode 0600) + `admin-audit.jsonl`, `intents.json` (Read My Mind intent profiles, mode 0600), `certs/` (self-signed TLS for `--https`), `.env` (CODEMAN_USERNAME/PASSWORD fallback for the `codeman attach` CLI). Transient: `self-update-runner.sh`. Multi-user case spaces live OUTSIDE the data dir at `~/codeman-users//cases` (shared across instances like `~/codeman-cases`, override `CODEMAN_USER_SPACES_DIR`). +All in `~/.codeman/`: `state.json` (sessions, settings, respawn, orchestrator, cron jobs/runs, owner tab layouts), `mux-sessions.json` (tmux recovery), `settings.json` (user prefs), `push-keys.json` + `push-subscriptions.json`, `session-lifecycle.jsonl` (audit log), `update-status.json` (self-updater progress, polled across the service restart), `docker-env-applied.json` (Compose deployment only: sha256 of the Dockerfile + compose file the running container was built from, written by `Start-Codeman.sh`, read by the self-updater's environment gate), `docker-build-source.json` (Compose deployment only: the checkout's HEAD commit and `package-lock.json` hash the `codeman-node-modules`/`codeman-dist` volumes currently reflect, written by both `Start-Codeman.sh` and a successful in-place self-update, compared to detect and refresh a volume left stale by an externally-triggered rebuild), `linked-cases.json`, `webviews.json` (saved web-tab dashboard URLs), `remote-hosts.json` + `remote-cases.json`, `docker-hosts.json` + `docker-cases.json` + `docker-exports/`, `subagent-window-states.json` + `subagent-parents.json` (subagent window layout, GET/PUT `/api/subagent-window-states`/`-parents`), `hook-secret` (per-instance), `users.json` (multi-user, mode 0600) + `admin-audit.jsonl`, `intents.json` (Read My Mind intent profiles, mode 0600), `certs/` (self-signed TLS for `--https`), `.env` (CODEMAN_USERNAME/PASSWORD fallback for the `codeman attach` CLI). Transient: `self-update-runner.sh`. Multi-user case spaces live OUTSIDE the data dir at `~/codeman-users//cases` (shared across instances like `~/codeman-cases`, override `CODEMAN_USER_SPACES_DIR`). **Generated top-level dirs** (all gitignored — don't edit or commit): `dist/` (esbuild output), `out/`, `coverage/`, `test-results/`, `tmp/`, `screenshots-echo-diag/`. The committed gesture bundle (`src/web/public/gesture/gesture-codeman.js`) IS tracked, but its runtime wasm/model assets (`src/web/public/gesture/wasm/`, `*.task`) are fetched and gitignored. diff --git a/docker/.env.example b/docker/.env.example index 12770034..1027b955 100644 --- a/docker/.env.example +++ b/docker/.env.example @@ -25,7 +25,7 @@ CODEMAN_APPDATA_PATH=/mnt/user/appdata/codeman # start script detects it from the compose file's own location, so it only needs # setting for direct `docker compose` use or a checkout kept elsewhere. Point it # at a directory that is not a git checkout and in-app updates are unavailable. -# CODEMAN_REPO_PATH=/mnt/user/appdata/Coding/codeman/app +# CODEMAN_REPO_PATH=/mnt/user/appdata/codeman/app # Required for Docker cases. This must be an absolute path on the Docker host. # Codeman and each isolated case use this same path, so it cannot be a diff --git a/docker/Start-Codeman.sh b/docker/Start-Codeman.sh index 92a37424..84c6b211 100644 --- a/docker/Start-Codeman.sh +++ b/docker/Start-Codeman.sh @@ -15,11 +15,16 @@ fi # Naming a Compose file explicitly disables Compose's automatic discovery of # the override file, so it has to be added back by hand. Without this, local # customisation in docker-compose.override.yml is silently ignored. The -# candidates are checked in Compose's own precedence order. +# candidates are checked in Compose's own precedence order - measured on +# Compose v5.5.0 with both present: it uses `.yml` and ignores `.yaml`. +override_yml="$script_dir/docker-compose.override.yml" +override_yaml="$script_dir/docker-compose.override.yaml" +if [[ -f "$override_yml" && -f "$override_yaml" ]]; then + printf 'Warning: both %s and %s exist; Compose uses .yml and ignores .yaml.\n' \ + "$override_yml" "$override_yaml" >&2 +fi compose_files=(-f "$compose_file") -for override_file in \ - "$script_dir/docker-compose.override.yaml" \ - "$script_dir/docker-compose.override.yml"; do +for override_file in "$override_yml" "$override_yaml"; do if [[ -f "$override_file" ]]; then compose_files+=(-f "$override_file") printf 'Using Compose override file: %s\n' "$override_file" @@ -31,6 +36,10 @@ appdata_path=$( "${compose_command[@]}" config --environment | awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }' ) +cases_path=$( + "${compose_command[@]}" config --environment | + awk -F= '$1 == "CODEMAN_CASES_PATH" { sub(/^[^=]*=/, ""); print; exit }' +) docker_socket=$( "${compose_command[@]}" config --environment | awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }' @@ -50,6 +59,26 @@ if [[ ! -d "$appdata_path" ]]; then mkdir -p -- "$appdata_path" fi +if [[ -z "$cases_path" ]]; then + printf 'Error: CODEMAN_CASES_PATH is not set in %s\n' "$env_file" >&2 + exit 1 +fi + +# Pre-creating this here, exactly like CODEMAN_APPDATA_PATH above, means Compose +# never has to materialise a missing bind source itself - which it does as +# root:root - so the in-container entrypoint's chown never has to run for this +# path at all. Unlike appdata, an EXISTING cases directory is left exactly as +# it is: the README explicitly allows pointing this at a normal projects +# directory the host account already owns, so no ownership check happens here. +if [[ ! -d "$cases_path" ]]; then + if [[ "$EUID" == '0' ]]; then + printf 'Error: Refusing to create CODEMAN_CASES_PATH as root: %s\n' "$cases_path" >&2 + printf 'Create it as the unprivileged account that should run Codeman, then retry.\n' >&2 + exit 1 + fi + mkdir -p -- "$cases_path" +fi + if owner_ids=$(stat -c '%u:%g' -- "$appdata_path" 2>/dev/null); then : elif owner_ids=$(stat -f '%u:%g' "$appdata_path" 2>/dev/null); then @@ -109,6 +138,11 @@ fi # Reads HEAD without requiring a `git` binary on the host — this script # otherwise checks the checkout only by testing for `.git` as a directory, and # resolving refs by hand keeps that the same "no host git needed" guarantee. +# ⚠️ A worktree checkout has `.git` as a FILE (`gitdir: `), not a +# directory, so this returns nothing there and the volume-refresh check below +# silently no-ops — consistent with the `-d .git` test used everywhere else in +# this script, not a special case, but worth knowing if a worktree checkout +# stops picking up a stale-volume refresh it should have caught. git_head_commit() { local git_dir="$1/.git" head_ref ref_path [[ -d "$git_dir" ]] || return 1 @@ -192,8 +226,23 @@ if [[ -n "$dockerfile_sha" ]]; then # no-op, so there is no fresh-install case this needs to avoid. printf 'Source changed since the last start; refreshing: %s\n' "${volumes_to_refresh[*]}" "${compose_command[@]}" down + # `com.docker.compose.volume` is the volume KEY, not a project-qualified + # name - a second stack on the same host (a beta instance started with a + # different COMPOSE_PROJECT_NAME, say) that also declares a volume keyed + # `codeman-dist` shares that label, and `head -n1` would pick whichever + # the daemon happens to list first. Scope the lookup to THIS stack's own + # resolved project name so it can only ever match this stack's volume. + project_name=$( + "${compose_command[@]}" config --format json 2>/dev/null | + sed -n 's/^ "name": "\(.*\)",\{0,1\}$/\1/p' | head -n1 + ) for key in "${volumes_to_refresh[@]}"; do - volume_name=$(docker volume ls -q --filter "label=com.docker.compose.volume=$key" | head -n1) + volume_name=$( + docker volume ls -q \ + --filter "label=com.docker.compose.volume=$key" \ + --filter "label=com.docker.compose.project=$project_name" | + head -n1 + ) [[ -n "$volume_name" ]] && docker volume rm -- "$volume_name" done fi diff --git a/docker/entrypoint.sh b/docker/entrypoint.sh index 45cf8d5a..a8f2d00c 100755 --- a/docker/entrypoint.sh +++ b/docker/entrypoint.sh @@ -25,12 +25,30 @@ fi for target in "${HOME:-}" "${CODEMAN_CASES_PATH:-}"; do [ -n "$target" ] && [ -d "$target" ] || continue - [ "$(stat -c '%u:%g' "$target")" = "${PUID}:${PGID}" ] && continue + owner=$(stat -c '%u:%g' "$target") + [ "$owner" = "${PUID}:${PGID}" ] && continue - # Deliberately not fatal. A bind mount backed by NFS, CIFS or a rootless - # daemon can refuse chown while still being perfectly writable, and those - # deployments must keep working. A warning is more useful than a container - # that will not start. + # Only ever correct a directory the DAEMON created (root-owned, because + # neither PUID nor PGID existed yet when it materialised the missing bind + # source). Anything else - a host tree that legitimately belongs to some + # OTHER account, such as an existing CODEMAN_CASES_PATH the README already + # allows pointing at a normal project directory - is not this container's + # to reassign; recursively chowning it on every mismatch silently rewrote + # a credentials tree or a projects directory to PUID:PGID with one log + # line to explain it. Refuse instead, the same way Start-Codeman.sh already + # refuses to touch a root-owned appdata directory it did not expect. + if [ "${owner%%:*}" != '0' ]; then + printf 'entrypoint: %s is owned by %s, which is neither root nor PUID:PGID (%s:%s).\n' \ + "$target" "$owner" "$PUID" "$PGID" >&2 + printf 'entrypoint: refusing to change ownership of a directory this container did not create.\n' >&2 + printf 'entrypoint: either chown it on the host, or set PUID/PGID to match its current owner.\n' >&2 + exit 1 + fi + + # Deliberately not fatal for a root-owned directory. A bind mount backed by + # NFS, CIFS or a rootless daemon can refuse chown while still being + # perfectly writable, and those deployments must keep working. A warning is + # more useful than a container that will not start. if chown -R "${PUID}:${PGID}" "$target" 2>/dev/null; then printf 'entrypoint: corrected ownership of %s to %s:%s\n' "$target" "$PUID" "$PGID" else @@ -45,4 +63,10 @@ done supplementary=$(id -G | tr ' ' '\n' | grep -vx 0 | paste -sd, -) [ -n "$supplementary" ] || supplementary="$PGID" -exec setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" "$@" +# --bounding-set -all: with the reuid/regid drop above, CapPrm/CapEff are +# already empty, but the bounding set otherwise still lists everything +# cap_add granted (visible as a nonzero CapBnd even post-drop). no-new-privileges +# already makes that moot - nothing can regain a capability outside the +# bounding set - but clearing it too is free and matches what "drops to +# PUID:PGID" actually promises. +exec setpriv --reuid "$PUID" --regid "$PGID" --groups "$supplementary" --bounding-set -all "$@" diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index c8a37475..27ec4a43 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -71,6 +71,24 @@ COPY --from=docker:29-cli \ # Keep credentials out of the image. Users authenticate these CLIs at runtime # through Codeman sessions, and the configured host bind mount retains state. # +# Installed into a DEDICATED prefix, /opt/codeman-cli, not the base image's +# default /usr/local. A session needs write access to wherever these CLIs live +# so it can self-update one in place (observed via Codex's own +# `npm install -g @openai/codex`, which renames the old package directory +# aside before installing the new one — a rename needs write access to the +# PARENT directory, not just the target, so the runtime account needs that +# access at the directory level). Chowning /usr/local/bin and +# /usr/local/lib/node_modules directly to get it would ALSO hand away +# entrypoint.sh (COPY'd to /usr/local/bin below, root-owned, executed as root +# on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node +# binary: owning the DIRECTORY is enough to rename it aside and drop a +# replacement, even though the file itself stays root-owned, which would let a +# compromised session arrange for its own script to run as root at the next +# restart — undoing the "the server itself never runs privileged" guarantee +# the entrypoint exists to provide. /opt/codeman-cli holds nothing else to +# escalate through, so owning it is exactly the CLI-update access it needs and +# no more. +# # ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are # a function of WHEN their image was built, not of any commit — so a Codeman # release that depends on newer CLI behaviour (the trust-dialog handling is @@ -82,6 +100,8 @@ COPY --from=docker:29-cli \ # # Bump these deliberately, in a release. `--no-cache` is still needed to rebuild # this layer when only the pins change upstream. +ENV NPM_CONFIG_PREFIX=/opt/codeman-cli +ENV PATH=/opt/codeman-cli/bin:$PATH RUN npm install --global \ @anthropic-ai/claude-code@2.1.258 \ @google/gemini-cli@0.58.0 \ @@ -94,14 +114,10 @@ RUN npm install --global \ # requested GID may not exist in the base image, and a host UID such as 1000 may # already belong to the baked `node` account, so handle both cases explicitly. # -# The trailing chown hands the globally-installed CLIs to that same account. -# They were `npm install --global`-ed above while still root, so -# /usr/local/lib/node_modules (and the /usr/local/bin symlinks pointing into it) -# start out root-owned; a session running as the unprivileged runtime user then -# hits EACCES the moment it tries to self-update one in place (observed via -# Codex's own `npm install -g @openai/codex`, which renames the old package dir -# aside before installing the new one — a rename needs write access to the -# PARENT directory, not just the target, so this must chown the whole tree). +# The trailing chown hands the CLI prefix (/opt/codeman-cli, populated above) +# to that same account, so a session can self-update one of the CLIs in place. +# /usr/local stays root-owned throughout — see the comment on the npm install +# above for why that boundary matters. RUN set -eux; \ case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \ case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \ @@ -130,7 +146,7 @@ RUN set -eux; \ --shell /bin/bash \ "${CODEMAN_RUNTIME_USER}"; \ fi; \ - chown -R "${PUID}:${PGID}" /usr/local/lib/node_modules /usr/local/bin + chown -R "${PUID}:${PGID}" /opt/codeman-cli WORKDIR /opt/codeman diff --git a/docs/docker-compose.md b/docs/docker-compose.md index f2add195..da85e203 100644 --- a/docs/docker-compose.md +++ b/docs/docker-compose.md @@ -15,7 +15,7 @@ The application container mounts the Docker daemon socket so Codeman can create ## Start -Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/Coding/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes. +Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes. ```sh cp docker/.env.example docker/.env