mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
bb8ada7e5fc5a3991db4392f972e99b4c8d53c12
31
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2bda191471 |
docs(docker): describe the root start and drop, and keep the override file out of the image
docker/README.md and docs/docker-compose.md now say that the container starts as root, corrects a daemon-created bind source and drops to PUID:PGID with setpriv, which capabilities that needs, and that a compose file written elsewhere must carry them. The README's PowerShell example runs Compose from inside docker/ so the override file is discovered, instead of the `-f docker/docker-compose.yaml` form its own Local customisation section warns silently drops it, and the reverse-proxy section no longer asks for an override file now that docker-compose.yaml forwards CODEMAN_ALLOWED_HOSTS itself. .dockerignore excludes docker-compose.override.* everywhere: it is the documented home for host-specific settings and rode `COPY . .` into the image, the same shape as the docker/.env exclusion above it (verified with a scratch build context: the override files and docker/.env are absent, .env.example and the compose file present). CLAUDE.md's Compose paragraph carries the corrected cap list, the writability probe, and the two traps behind it (KILL is for tini, the CLI prefix is appended to PATH). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
f92883704e |
fix(docker): create the cases dir with the runtime owner and record a refresh only when it happened
Start-Codeman.sh created CODEMAN_CASES_PATH with a plain `mkdir -p` BEFORE it derived PUID/PGID from the appdata directory, so the new directory landed as the invoking user's uid and primary gid. On a host set up the way the README suggests (`chown -R 99:100 <appdata>`) that gid is not PGID, and the container refused to start on a directory the script had just made. PUID/PGID are now derived first and the directory is chowned to them right after creation, with a clear host-side error when that is not possible. As root this always works, which also retires the old "refusing to create as root" branch for this path. The build-artefact volume refresh had three holes. The docker-build-source.json marker was written whether or not a volume had actually been removed, and the project name came from a sed over `docker compose config --format json` keyed on two-space indentation: an empty name made the label filter match nothing, nothing was removed, and the marker recorded the new HEAD, so the check never fired again while the stale volume kept serving old code. The name is now parsed indentation-agnostically, an empty result falls back to `down --volumes` (the documented reset; both volumes re-seed from the image by a plain copy), the marker is written only after a successful refresh, and a failed `docker volume rm` warns and leaves the marker alone instead of aborting under set -e with the stack down. The image is also built BEFORE `down`, so the deployment is offline only for the recreate rather than for the whole rebuild. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1851d80f3a |
fix(docker): keep SIGTERM reaching the server, pin root's PATH, probe writability
Three changes to how the Compose container starts as root and drops to PUID:PGID, each reproduced on Docker 29.1.3 / Compose v5.5.0 with a minimal image of the same shape as server.Dockerfile. - cap_add gains KILL. `init: true` makes tini PID 1, and tini stays root while the entrypoint drops the server to PUID. Signalling a process of a different uid needs CAP_KILL, and `cap_drop: ALL` had removed it, so every `docker compose down`/`restart` ended in `[FATAL tini (1)] Unexpected error when forwarding signal: 'Operation not permitted'` and the server being SIGKILLed instead of running `server.stop()`. Measured: without KILL the trap never fires, with it the child logs `GOT SIGTERM`. - /opt/codeman-cli/bin is appended to PATH, never prepended, and entrypoint.sh pins its own PATH to the system directories before its first command. The prefix is chowned to the runtime account so sessions can update the agent CLIs in place, and the root entrypoint resolved stat/chown/setpriv by bare name through it: a `setpriv` planted there by the unprivileged uid ran as uid 0 at the next start. The image's full PATH is handed back to the server at the exec (`env PATH=...`), since Codeman resolves the CLIs through it. - The ownership gate becomes a writability probe. A directory owned by neither root nor PUID:PGID is no longer refused on ownership alone; it is tested with `setpriv --reuid PUID --regid PGID --groups <same groups> test -w`, the exact identity the server gets, so a group-writable tree, an ACL or a CIFS/NFS mount reporting some unrelated uid all pass, and the refusal names path, owner and PUID:PGID. Root-owned directories are still chowned first. Also: a pre-flight runs the drop before touching anything and, when it fails, prints the cap_add list the compose file needs, so an out-of-tree compose file (Unraid's Compose Manager) gets a one-line diagnosis instead of a restart loop; `--bounding-set -all` is gone, since it is a silent no-op without CAP_SETPCAP; a root:root Docker socket now produces a warning that Docker cases will not work rather than silently losing group 0 at the drop; and CODEMAN_ALLOWED_HOSTS is forwarded from .env with an empty default (documented as a commented entry in .env.example so the parity test and the updater's env gate both stay quiet). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
a29e1f61ef |
Merge pull request #377 from opticon454/bugfix-docker-user-perms
fix(docker): bind-mount ownership, Compose override discovery, and the default runtime account |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
7af4dbc0f8 |
feat(docker): derive the agent image's npm CLI list from the catalogue
docker/agent.Dockerfile hardcoded the four npm-published CLIs it installs, one of the several lists that had to be kept in step with the registry by hand. It now takes them as `ARG CLI_NPM_PACKAGES`, supplied by scripts/build-agent-image.mjs from config/clis.stock.json, with the default set to today's list so a bare `docker build` still produces the same image. The arg is expanded unquoted because word splitting is what turns the list into several arguments, which is exactly why every token is validated against ^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side; a package name carrying a space or a metacharacter is refused rather than reaching the RUN line. Verified by building the layer: four packages in, four arguments out, and the default still applies with no arg. The list is filtered on each entry's `enabled` flag — the field whose absence was the maintainer's §3 finding, where a CLI shipping disabled still got baked into every image. No stock entry is disabled today, so that assertion would pass vacuously; a unit test feeds the pure helper a fabricated disabled entry so the fix is covered now rather than the first time someone ships one. ⚠️ It reads the STOCK catalogue, never the merged registry. A user's ~/.codeman/clis.json must not change what is inside an image tagged codeman/agent:base, or two machines holding that tag hold different images. Four CLIs keep hand-written layers because the registry cannot describe what makes them special: pi's --ignore-scripts, deepseek's pnpm companion and dsh-tui profile, and the three standalone installers. Rather than extend the schema for a Docker-only benefit, the coverage test requires each to carry a written reason AND still be present, so an exclusion cannot quietly become an omission. There are two producers of this command line and there have to be — the .mjs cannot import TypeScript, and src/docker-hosts.ts builds the same argv for the in-app auto-build — so a parity test pins them together, package list, arg pairs and rendered argv. Their order is pinned too: a different order is a different RUN string and so a needless cache miss between the two build paths. docker/server.Dockerfile is deliberately NOT edited (PRs #373 and #377 both modify it); its narrower list is asserted as a declared omission list instead, so the divergence is reviewable without touching the file. Also fixes the in-app hint at index.html, which the new coverage test caught still omitting omp. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
c179daf869 |
fix(docker): re-assert /opt/codeman-cli ownership every start, not just at build
/opt/codeman-cli is chowned to PUID:PGID once, at image build time, from the PUID/PGID build args. That bake only happens when the image is actually rebuilt (`docker compose up --build`, which Start-Codeman.sh always does) — a deployment that runs the compose file directly instead (Unraid's Compose Manager, a native systemd unit, any plain `docker compose up`/`restart`) can change PUID/PGID in .env and restart without ever rebuilding. The container then runs as the NEW uid via entrypoint's setpriv (Linux needs no /etc/passwd entry to setuid to an arbitrary number) while the CLI directory is still owned by the OLD one baked into the image layer — silently breaking the self-update-a-CLI- in-place fix that directory exists for. Unlike HOME/CODEMAN_CASES_PATH, this one is pure image content Codeman itself populated, never host data that might legitimately belong to someone else, so there is no ownership to be careful about — it is always correct for it to be owned by whoever the container is about to run as. Re-assert it unconditionally on every start. Verified live: built an image with PUID=99/PGID=100, ran it with PUID=1234/PGID=4321 (no rebuild, simulating a changed .env restarted directly), confirmed /opt/codeman-cli ends up 1234:4321-owned and is genuinely writable by the running process. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
ae32daf135 |
fix(docker): address maintainer review on #377
Two real bugs the review caught, both verified live against a real build on the Unraid host: 1. entrypoint.sh's chown fired on ANY ownership mismatch, not just a directory the daemon itself created root-owned. A host tree legitimately owned by some other account - an existing CODEMAN_CASES_PATH the README already allows pointing at a normal projects directory, or appdata under a different PUID/PGID convention than the one in use - got silently recursively re-owned with one log line to explain it. Now gated on the target actually being root-owned; anything else is a clean refusal naming the directory, its owner, and PUID/PGID. Start-Codeman.sh also now pre-creates CODEMAN_CASES_PATH the same way it already did CODEMAN_APPDATA_PATH, so Compose never has to materialise a missing bind source as root in the first place - the in-container chown becomes a safety net, not the primary mechanism. 2. The CLI-update chown (chown -R .../node_modules /usr/local/bin) handed the runtime account write access to entrypoint.sh itself (root-owned, executed as root on every container start with CHOWN/DAC_OVERRIDE/SETUID/SETGID) and the node binary - owning the DIRECTORY is enough to rename it aside and drop a replacement, which would let a compromised session arrange for its own script to run as root at the next restart. The four CLIs now install into a dedicated /opt/codeman-cli prefix (NPM_CONFIG_PREFIX); only that directory is chowned, /usr/local stays root-owned throughout. Smaller fixes from the same review: - Start-Codeman.sh's volume-refresh label filter wasn't project-scoped: a second Compose stack on the same host sharing the `codeman-dist` volume KEY could have had ITS volume deleted. Added a com.docker.compose.project filter, resolved from this stack's own `compose config --format json`. - Override-file precedence was backwards (checked .yaml before .yml; Compose actually prefers .yml) - swapped, plus a warning when both exist. - entrypoint.sh's setpriv now also passes --bounding-set -all, so CapBnd actually clears post-drop rather than just CapPrm/CapEff. - A comment on git_head_commit() noting it returns nothing for a worktree checkout (.git as a file), consistent with the script's existing -d .git convention elsewhere. - Doc drift: CLAUDE.md's Docker Compose section still described the old pre-created-and-chowned-by-hand model and didn't mention the root-then-drop entrypoint; the state-files list was missing docker-build-source.json; docs/docker-compose.md and docker/.env.example still had the pre-rename `Coding/codeman` path in one place each. Verified end to end against a real build on the Unraid host: a root-owned bind source is corrected as before; a directory owned by neither root nor PUID:PGID is refused rather than silently rewritten; a correctly-owned directory is left alone entirely; the four CLIs resolve via PATH from /opt/codeman-cli while /usr/local/bin, /usr/local/lib/node_modules and entrypoint.sh itself stay root-owned; CapBnd is fully cleared post-drop. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
8fe3f34fc5 |
fix(docker): detect and refresh stale build-artefact volumes
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from the image only while empty, so a rebuilt image's fresh dist/ node_modules sat unused behind old volume content until something cleared it. The in-app self-updater never hit this (it rebuilds INSIDE the running container, into the very volume already in use), but a `docker compose build` triggered from outside it — Start-Codeman.sh, after a manual `git pull` — did: the container came back up looking unchanged, serving stale compiled routes against current source. Start-Codeman.sh now compares the checkout's HEAD commit and package-lock.json hash against a recorded marker (docker-build-source.json) and clears just the affected volume(s) before its own --build when either moved. The in-place self-update path writes that same marker after a successful build, so the two mechanisms agree on what the volumes currently reflect — without it, the next plain Start-Codeman.sh run would see the HEAD self-update just checked out, not recognise it as already accounted for, and wipe the volumes self-update just correctly rebuilt right back to the older baked image. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
89e2cb5814 |
fix(docker): let the runtime account update its own global CLIs
The four CLIs (claude, gemini, codex, opencode) are npm-installed globally as root during the image build, before the unprivileged runtime account exists. A session running as that account (e.g. a codex-mode terminal) then hits EACCES the moment it tries to update one in place, because npm renames the old package directory aside before installing the new one, which needs write access to the parent (/usr/local/lib/node_modules), not just the target package. Chown that tree plus /usr/local/bin's CLI symlinks to PUID:PGID in the same step that creates/renames the runtime account. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
d38bf33a69 |
docs(docker): document the reverse-proxy host allowlist
CODEMAN_ALLOWED_HOSTS is a real, documented application setting (the Host- header allowlist in network-auth-policy.ts), but docker-compose.yaml does not forward it from .env into the container - Compose only passes through variables explicitly listed under environment:, and this is not one of them. Set without that passthrough, any request through a reverse proxy is rejected with 403 Forbidden: host not allowed before it reaches any handler, and nothing in the Docker deployment docs said why. Document the variable and the override needed to forward it, using the Local customisation mechanism already described above it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
9702126046 |
chore(docker): name the default runtime account codeman
CODEMAN_RUNTIME_USER defaulted to `opencode`, which no longer matches the project and is confusing in a deployment whose every other identifier is codeman. Rename the default in .env.example and in the Dockerfile ARG that mirrors it, and correct the example comment that referred to /home/opencode/codeman-cases. Also drop the `Coding/` component from the example application-data path. CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH now suggest /mnt/user/appdata/codeman and its codeman-cases child, matching the account name and removing a directory level that meant nothing outside the original author's host. README.md is updated to match, including the chown example. The npm package `opencode-ai` and the references to the OpenCode CLI are deliberately left alone: those name a different tool, not this account. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
748bbf5423 |
fix(docker): honour docker-compose.override.yml in Start-Codeman.sh
Naming a Compose file with -f disables Compose's automatic discovery of the override file, so Start-Codeman.sh silently ignored docker-compose.override.yml. Any local customisation placed in the conventional override file was dropped without warning, and the only way to notice was to inspect the running container. Collect the -f arguments into an array, append the override file when one is present, and reuse that array for the final launch so the two cannot drift apart again. Both .yml and .yaml are checked, in Compose's own precedence order, and the chosen file is reported on startup. Document the override file in docker/README.md, including the two things that are easy to get wrong: it is ignored when -f is passed without naming it, and it cannot remove a key such as ports, which Compose concatenates. Add the override file to .gitignore. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
10876aa440 |
fix(docker): correct bind-mount ownership before dropping privileges
Compose binds CODEMAN_APPDATA_PATH and CODEMAN_CASES_PATH from the host. When either path does not exist yet - a first run, a cleared application-data directory, a restored backup - the Docker daemon creates it owned by root. The server runs unprivileged as CODEMAN_RUNTIME_USER, so it cannot create its own state directory, and the container restarts forever on: Failed to start web server: EACCES: permission denied, mkdir '/home/<user>/.codeman' Start-Codeman.sh already worked around this by preparing the directory on the host, so the failure only appears when Compose is run directly, which the README documents as a supported path. Add docker/entrypoint.sh, which starts as root, corrects the ownership of both bind mounts, then drops to PUID:PGID with setpriv. The Dockerfile's USER instruction is replaced by that entrypoint and CMD is unchanged. docker-compose.yaml adds back only the four capabilities the chown and the privilege drop require, so cap_drop: ALL continues to remove everything else. Two guards keep existing deployments working: - A container started with an explicit `user:` is left alone. The entrypoint execs straight through, with no elevation and no chown. - A chown that fails is a warning, not an error. Bind mounts backed by NFS, CIFS or a rootless daemon can refuse chown while remaining perfectly writable, and those deployments must keep starting. PUID and PGID are also exported as runtime environment defaults so the image behaves correctly when run without Compose, rather than depending on build args alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
99ad9cb236 |
fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for the shipped deployment: `restart: unless-stopped` relaunches it. The updater verified that policy through the Docker socket and, when it could not (no socket mounted), failed open and exited anyway. Failing open is the correct choice for the GATE, where refusing would block every install without a socket, but not for the kill: a container the daemon does not restart goes down for good, with no UI left to recover it from. That is exactly the case a plain `docker run` of this image without `--restart` produces, and the image sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path. The decision now happens server-side, where both the socket and the Compose env are reachable, and rides down to the script as `--restart-by-exit 0|1`. It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there and only there, since that file is what sets the restart policy; the image ENV deliberately does not) or when the daemon confirmed an auto-restart policy. Otherwise the build still lands, the status becomes `completed-needs-manual-restart` with the `docker restart` hint, and the server keeps running. The shipped deployment is unchanged in effect: with the socket it was already confirmed, and without it the declaration now covers it. Also: a root-run `Start-Codeman.sh` (common on Unraid) created the fingerprint baseline's `.codeman` directory before the container's first start and left it root-owned, which the unprivileged server could then never write its own state into. It is chowned to PUID:PGID when running as root. Verified with a real image build of the merged tree (classic builder; this box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain present, the four CLIs at their pins, docker/.env absent, and `docker inspect $HOSTNAME` returns the restart policy through the mounted socket as that user. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu |
||
|
|
66eb01ba8f |
feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
|
||
|
|
826ddaa9aa |
build(docker): ship the Docker CLI in the Compose image, not the whole engine
docker/server.Dockerfile installed Debian's `docker.io` to get a client for the socket mounted by Compose. That package is the full ENGINE: even with --no-install-recommends it pulls 15 packages including containerd, runc, dmsetup and iptables, none of which a container that only talks to a mounted socket can use, and it ships Docker 20.10.24 (2023). Copy the CLI and the buildx plugin from the official docker:29-cli image instead. Measured on the same node:22-bookworm-slim base: 266 MB -> 108 MB, so 158 MB smaller with a current CLI (29.7.2) in place of a two-year-old one. Three things verified rather than assumed, by building the real image and running it: - docker:cli is an ALPINE image, so copying a binary into this Debian one is only safe because the binaries are static Go builds (ldd: "Not a valid dynamic program"). In the built image, `docker --version`, `docker ps` and `docker build` all work against a mounted host socket as the unprivileged runtime user. - buildx is copied on purpose. scripts/build-agent-image.mjs shells out to `docker build` and Codeman auto-builds the agent image on the first Docker case. Without the plugin that still works today — CLI 29 falls back to the classic builder, tested — but that builder is deprecated and will be dropped, so the plugin keeps the path supported. - docker-compose is NOT copied: Codeman never shells out to it. Pinned to the 29 major, matching how the base images here are pinned. |
||
|
|
2a32b5064a |
Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases as SIBLING containers through the mounted host socket (Docker-outside-of-Docker). Resolved the README conflict (master had grown to eight CLIs since the branch was cut) and moved the Compose blurb out of the feature bullets into Quick Start, next to the other ways of starting Codeman. Three review findings from the PR discussion are fixed here rather than left for a follow-up, because two of them are shipped-image problems: - `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against the whole context-relative path, so `docker/.env` — which the deployment's own README tells the user to fill with CODEMAN_PASSWORD and provider API keys — was picked up by `COPY . .` and baked into the image at /opt/codeman/docker/.env. Verified in both directions against a real build context: with a canary secret in docker/.env, the unfixed ignore file lets /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for the checked-in .env.example) leaves only the example behind. - `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which still hardcoded ~/codeman-cases, so `codeman skill install --case <name>` reported "Case not found" on exactly the deployment the override exists for. Both now resolve through config/cases-dir.ts. state-store.ts keeps its own literal on purpose: that one migrates the historical ~/claudeman-cases directory by name and is about the old default, not the active location. - CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore joins the documented list of files that genuinely belong in the repo root. The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/ deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's lastClaudeSessionId was handed to every non-claude CLI. Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax, public assets, lockfile. |
||
|
|
d5b5f8f618 |
fix(docker): make the dsh profile install survive pnpm's build-script gate
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an image without pnpm dies at exit 127 and takes the whole build with it. That PR also pinned an allowlist of the two packages whose lifecycle scripts pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm, unlike npm, refuses dependency build scripts by default and FAILS the install over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two names would have let the next tree break the build the same way. Allowing them wholesale is also the exposure this image already accepts three layers up, where `npm install -g` runs the install scripts of every transitive dep of the five CLIs above with no gate at all. Also correct a comment in the `/api/deepseek/install-profile` route that asserted the opposite of what #352 proved ("dsh bundles its own package manager, so no system pnpm is required"). The route's behavior is already right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the fix. Documented the prerequisite in docs/deepseek-integration.md, and taught the docker-cases image smoke test about `dsh`/`omp` plus the profile check that `dsh --version` does NOT cover. |
||
|
|
7762809202 |
Merge pull request #350 from opticon454/bugfix/dsh-pnpm
fix(docker): install pnpm for DeepSeek profile |
||
|
|
d74cde759b |
feat(omp): install omp in the docker agent image, isolate its credentials
OMP had full routing at the Docker layer (default pane command, schema) but
was never actually installed in docker/agent.Dockerfile, and had no
credential-isolation entry in docker-hosts.ts's CRED_STORES - a Docker-mode
OMP session would have failed with "omp: command not found", and even with
the binary present would have had no config/auth seeded, despite the README
already claiming OMP has "seamless auth, isolated credentials" in Docker.
- docker/agent.Dockerfile: install omp via its own installer (standalone
binary, same shape as grok/antigravity - not on npm). Verified against a
real --no-cache build: the installer actually targets ~/.local/bin, not
~/.omp/bin as the resolver's OMP_SEARCH_DIRS ordering would suggest -
confirmed omp/18.0.8 installs and runs correctly inside the image.
- src/docker-hosts.ts: add a .omp/agent CRED_STORES entry. Unlike every
sibling CLI in this family, sessions/ is SHARED (RW), not seeded: Codeman
reads ~/.omp/agent/sessions/**/*.jsonl host-side for history recovery and
--resume pinning (omp-transcript.ts, omp-session-resolver.ts), the same
reason codex's sessions/ is shared rather than seeded. Seeding it instead
would silently break the kill-survival feature for Docker cases. Only the
small config files (config.yml/mcp.json/models.yml/settings.yml) are
seeded; the SQLite caches and terminal-sessions/ stay container-local.
- test/docker-hosts.test.ts: pin the new CRED_STORES entry's behavior.
Found in passing (NOT fixed here, unrelated and pre-existing on master): the
agent image's DeepSeek (dsh) plugin-install step currently fails on a fresh
build ("pnpm not found on PATH"), confirmed via git diff against
origin/master that this line is untouched by this branch. Worth a separate
issue/PR.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
26b4ffbb0f | fix(docker): install pnpm for DeepSeek profile | ||
|
|
b85f7659b7 | feat(docker): add Compose deployment support | ||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9cfd8e8989 |
fix(docker): survive xAI installer's own /usr/local/bin/grok symlink
The agent-image grok step copied /root/.grok/bin/grok onto /usr/local/bin/grok with cp -L. Newer versions of xAI's install.sh already create /usr/local/bin/grok as a symlink to that same binary, so the copy failed with 'same file' and the --no-cache rebuild died at the grok layer (2026-08-24). Stage the copy under a temp name, drop whatever the installer left at the destination, then move into place - correct against both old and new installers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
c5b59633d8 |
feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code, OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose tab identity, welcome button, run-mode entry, cron agentType, Docker and remote-SSH command defaults, and clone-repo Brain option. Pi is a different shape of CLI from the other four, and three decisions follow from that: - It has NO permission prompts and no sandbox, so there is no --dangerously-skip-permissions analog and none was invented. The privilege-shaped knob is the tri-state approveProjectTrust, which makes pi load and EXECUTE repo-local .pi/extensions TypeScript and install missing project packages. clampExternalCliBypassForOwner() therefore puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets --no-approve even when no config was sent, because pi's own default is a prompt the session user could answer themselves. That helper had zero test coverage; it now has coverage for all four CLIs. - Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode context, so admitting them would widen the allowlist for every mode at once. Auth goes through pi's /login or the server's own environment. --api-key is deliberately never wired: it would put a provider secret on the spawn command line. - pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the main screen with terminal-owned scrollback, and its 0.84.0 fullscreen mode is runtime-switchable via /settings; that flip was measured to put the pane into the alt screen, which the strip would have corrupted. pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires semver-shaped output, because `pi` is a short generic name a stray binary can shadow; GET /api/pi/status surfaces path and version so a misresolution is diagnosable rather than presenting as a broken mode. Docker installs pi in its own --ignore-scripts step so that flag cannot affect the other four CLIs, and seeds its credentials per-file rather than whole-dir (~/.pi/agent also holds sessions, extensions and package trees). Verified end to end against pi 0.84.1 on an isolated instance: resolver search-dir fallback, flag construction, piConfig persistence across a full server restart, the trust prompt and its --no-approve suppression, the rose Run button on the default daylight-blue skin (the nested skin block eats per-mode gradients unless the rule lives inside it), and the buffer local-echo policy, which pi tolerates where codex did not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d0b772619 |
feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to the surfaces around it, while Gemini CLI stayed documented as a consumer product despite being enterprise-only since Google's cutover. Gemini keeps full support; Antigravity now sits beside it everywhere. Functional fixes: - docker/agent.Dockerfile never installed agy, so a docker case with mode 'antigravity' died on command-not-found. agy is not on npm, so it gets its own installer step. --dir /usr/local/bin is load-bearing: the default $HOME/.local/bin resolves to root's home at build time and is unreachable by the `agent` user the container runs as. Verified inside codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is ~190MB, the largest layer in the image. - Welcome screen gained a Run Antigravity action, gated on agy being present like the other CLI buttons, with a cyan identity matching the toolbar run button and run-mode dot. - install.sh now detects agy (search paths mirroring the resolver), counts it as a satisfying AI CLI, and recommends it over Gemini in the install hints. Detection only, no new auto-install path. Docs corrected where they were factually wrong: - architecture-invariants documented isExternalCliMode() as opencode/codex/gemini when the code has included antigravity for a while, said "all three modes", and omitted ANTIGRAVITY_ from the env prefix allowlist row. - cron-guide's agentType enum, cron-discovery's SessionMode, and remote-sessions' RemoteCommandMode were all stale. Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only), package.json keyword, and comment drift in 8 places. test/run-mode-ui.test.ts now covers the new welcome button; verified it fails without the settings-ui wiring. Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not ~/.antigravity, so the existing .gemini docker credential seed already covers it. Recorded as a comment so nobody adds dead config later. isAltScreenStripMode() deliberately still excludes antigravity: whether its TUI needs the alt-screen strip is a behavioural question that needs a real agy session, not a guess. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ca731c67b3 |
feat(docker): harden session mode + File Viewer button (v1.4.1)
Docker cases: seamless Claude auth (seed ~/.claude.json instead of the corruption-prone single-file mount), full credential-store isolation for claude + codex/gemini/gcloud/opencode (share only transcripts/rollouts, seed the rest), auto-build the base image on first use, C.UTF-8 locale (fixes box-drawing), collapsed/shortened Create-Case UI + short "(docker)" case-menu tags, and w<n>-<case> tab naming for docker/remote sessions. Also: opt-in File Viewer header button; fix a TZ-boundary flaky test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5f4c89b990 |
fix(docker): auto-assign agent uid (node:22 already occupies uid 1000)
node:22-bookworm-slim ships a `node` user at uid 1000, so `useradd -u 1000` failed. Auto-assign the uid and rely on gid-0 + group-writable HOME so any runtime `--user <hostUid>:0` can write $HOME. Verified: image builds; toolchain (node/tmux/claude/codex/gemini/opencode) present; `--user 1000:0` writes /home/agent and `claude --version` runs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
54615e2371 |
feat(docker): agent base image + local build script
docker/agent.Dockerfile: node:22 + claude/codex/gemini/opencode CLIs + git/ tmux/ripgrep/curl, secret-free, OpenShift arbitrary-uid-writable HOME (gid 0). scripts/build-agent-image.mjs: local build (decision "build locally on first use"), docker/podman auto-detect, --engine/--image/--no-cache flags. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |