mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
fix/self-update-stalled-status
6
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
697b05b118 |
fix(docker): merge-time fixes for Update-Codeman.sh (#465)
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose label after a plain `down`, instead of `down --volumes` (which also takes any volume an override file declares while the message named two). `down --volumes` remains only as a warned fallback when the project name cannot be resolved. - Report a failing first `docker compose config --format json` call with a clear error instead of exiting silently under `set -e`. - Filter empty label lines in the collision guard so an unlabelled container cannot hide a real collision; name the moved-checkout exit in its error. - Comments no longer cite a guard or incident in Start-Codeman.sh that does not exist; the README states the real gap (a Node base-image bump leaves codeman-node-modules stale because the lockfile did not move). - docs: Update-Codeman.sh in the docker-self-update.md short-version table and a mention in docker-compose.md; "Major updates" moved under "Updating" in docker/README.md. - test: smoke test covers the new sequence, the config failure and the empty-line case; quiet stdio; @fileoverview names the fourth concern. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
5b4878df3b |
fix(docker): address Ark0N's PR review on Update-Codeman.sh — fix the handoff, the build/down ordering, and default-clear the build volumes
Three blockers, all fixed and verified by actually running the script (not just string-matching it): 1. `exec "$script_dir/Start-Codeman.sh"` failed EACCES/exit 126 on every checkout, since Start-Codeman.sh is committed non-executable (100644) — the same fact my own second commit on this branch established. Fixed to `exec bash "$script_dir/Start-Codeman.sh"`. 2. `down` ran before `build --no-cache`, so Codeman and every session it was running were offline for the entire rebuild, and a build failure left the stack down with nothing to bring it back — the exact ordering mistake Start-Codeman.sh's own "Build BEFORE taking the stack down" comment exists to prevent. Reordered to build, then down, then hand off. 3. The default path could throw the rebuild away: codeman-node-modules/ codeman-dist only re-seed from the image while EMPTY, Start-Codeman.sh only clears them when it detects the checkout's HEAD or package-lock.json moved, and neither condition is true for the Dockerfile-only change this script exists for — so a plain `bash docker/Update-Codeman.sh` rebuilt an image whose fresh node_modules/dist then sat unused behind the old volumes. Made clearing them the default; `--keep-volumes` opts out (replaces the old `--volumes`/`-v` flag, which is no longer needed since clearing is now the default). Smaller items from the same review, also fixed: - The --no-cache build now derives PUID/PGID from CODEMAN_APPDATA_PATH's owner first, via the identical owner_of() helper Start-Codeman.sh uses (parity-tested) — without it, the build used Compose's default 1000:1000 regardless of the real appdata owner (99:100 on the Unraid layout docker/README.md documents), and Start-Codeman.sh's own correctly-PUID'd build during the handoff would then rebuild those layers anyway, so the --no-cache image never actually shipped. - docker/README.md's "rebuilds ... only when it detects ... moved" wrongly described BOTH the rebuild and the volume-clearing as conditional; Start-Codeman.sh rebuilds on every start, only the volume-clearing is conditional. Corrected, and reworded around the new default. - --help/-h now prints usage and exits 0 instead of falling into the unrecognised-argument branch. - "the ONLY named volumes this stack declares" now says docker-compose.yaml specifically, since a docker-compose.override.yml could add more. New tests: PUID/PGID derivation parity with Start-Codeman.sh's owner_of(), --help handling, and — the one that actually catches blocker #1, which five source-string-matching tests did not — a real end-to-end smoke test: a synthetic deployment, a stub `docker` on PATH logging every invocation, the real script executed via a real subprocess. Confirms the real command sequence (build --no-cache, then down --volumes or plain down, then evidence the handoff genuinely ran Start-Codeman.sh) and that a working handoff fails honestly at Start-Codeman.sh's own later check rather than with EACCES. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
9ba90a674a |
chore(docker): add Update-Codeman.sh for scripted major-update rebuilds
docker/README.md and docs/docker-self-update.md both already point operators at "stop the stack, rebuild, restart" for anything the in-app updater refuses to apply (a changed server.Dockerfile, a changed docker-compose.yaml, or a new required .env key) — but that was a manual, hand-typed procedure with no script of its own, unlike every other start/update path this deployment has. docker/Update-Codeman.sh scripts it: `docker compose down`, then an unconditional `docker compose build --no-cache` (a major update should be certain of what actually ships, not reuse whatever layers happened to be cached), then hands off to the existing Start-Codeman.sh for the same careful PUID/PGID, override-file and fingerprint handling every other start already goes through — rather than reimplementing any of that by hand and risking it drifting out of step. An optional --volumes/-v flag also removes the codeman-node-modules/ codeman-dist named volumes, the scripted form of the "Resetting the build artefacts" procedure docs/docker-self-update.md already documents by hand. Safe: those two are the only named volumes this stack declares; application data and case workspaces are host bind mounts, never touched by `docker compose down` either way. Docs updated: a "Major updates" section in docker/README.md, and a pointer from docs/docker-self-update.md's existing "Resetting the build artefacts" troubleshooting entry. Tests: extended test/docker-entrypoint.test.ts (the existing home for Start-Codeman.sh's own static checks) with a bash -n parse check, the down-before-build-before-handoff ordering, the --volumes flag's effect, unrecognised-argument handling, and byte-for-byte agreement with Start-Codeman.sh's own override-file resolution logic (so `down` here and `up` there can never target different Compose files). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD |
||
|
|
8fe3f34fc5 |
fix(docker): detect and refresh stale build-artefact volumes
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded from the image only while empty, so a rebuilt image's fresh dist/ node_modules sat unused behind old volume content until something cleared it. The in-app self-updater never hit this (it rebuilds INSIDE the running container, into the very volume already in use), but a `docker compose build` triggered from outside it — Start-Codeman.sh, after a manual `git pull` — did: the container came back up looking unchanged, serving stale compiled routes against current source. Start-Codeman.sh now compares the checkout's HEAD commit and package-lock.json hash against a recorded marker (docker-build-source.json) and clears just the affected volume(s) before its own --build when either moved. The in-place self-update path writes that same marker after a successful build, so the two mechanisms agree on what the volumes currently reflect — without it, the next plain Start-Codeman.sh run would see the HEAD self-update just checked out, not recognise it as already accounted for, and wipe the volumes self-update just correctly rebuilt right back to the older baked image. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru |
||
|
|
99ad9cb236 |
fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for the shipped deployment: `restart: unless-stopped` relaunches it. The updater verified that policy through the Docker socket and, when it could not (no socket mounted), failed open and exited anyway. Failing open is the correct choice for the GATE, where refusing would block every install without a socket, but not for the kill: a container the daemon does not restart goes down for good, with no UI left to recover it from. That is exactly the case a plain `docker run` of this image without `--restart` produces, and the image sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path. The decision now happens server-side, where both the socket and the Compose env are reachable, and rides down to the script as `--restart-by-exit 0|1`. It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there and only there, since that file is what sets the restart policy; the image ENV deliberately does not) or when the daemon confirmed an auto-restart policy. Otherwise the build still lands, the status becomes `completed-needs-manual-restart` with the `docker restart` hint, and the server keeps running. The shipped deployment is unchanged in effect: with the socket it was already confirmed, and without it the declaration now covers it. Also: a root-run `Start-Codeman.sh` (common on Unraid) created the fingerprint baseline's `.codeman` directory before the container's first start and left it root-owned, which the unprivileged server could then never write its own state into. It is chowned to PUID:PGID when running as root. Verified with a real image build of the merged tree (classic builder; this box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain present, the four CLIs at their pins, docker/.env absent, and `docker inspect $HOSTNAME` returns the restart policy through the mounted socket as that user. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu |
||
|
|
66eb01ba8f |
feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
|