* fix(self-update): stop a stalled status from blocking every later update
A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.
- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
is older than the stale window; applied on every read (start + status
poll) and persisted. The live updater heartbeats every few seconds, so a
running update never trips it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down
On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.
- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
then SIGKILL it. tmux sessions live outside the server and survive.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded
from the image only while empty, so a rebuilt image's fresh dist/
node_modules sat unused behind old volume content until something
cleared it. The in-app self-updater never hit this (it rebuilds INSIDE
the running container, into the very volume already in use), but a
`docker compose build` triggered from outside it — Start-Codeman.sh,
after a manual `git pull` — did: the container came back up looking
unchanged, serving stale compiled routes against current source.
Start-Codeman.sh now compares the checkout's HEAD commit and
package-lock.json hash against a recorded marker
(docker-build-source.json) and clears just the affected volume(s)
before its own --build when either moved.
The in-place self-update path writes that same marker after a
successful build, so the two mechanisms agree on what the volumes
currently reflect — without it, the next plain Start-Codeman.sh run
would see the HEAD self-update just checked out, not recognise it as
already accounted for, and wipe the volumes self-update just correctly
rebuilt right back to the older baked image.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.
The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.
Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.
Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
A KeepAlive system-level LaunchDaemon (the right setup for headless Macs,
where no GUI login means LaunchAgents never start) is now detected as
supervisor 'launchd-daemon': the updater kills the server PID (passed via
--server-pid) and launchd respawns it on the new dist/ — no root needed.
Detection requires the daemon plist to be bootstrapped AND KeepAlive=true.
Also: on boot, a 'completed-needs-manual-restart' status auto-completes
when the running version matches the staged target, so the stale
'restart Codeman to apply' instruction no longer lingers in the UI.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The updater wrote the status once per phase, so the minute-plus npm install and
build steps left the UI frozen on a single label. Add:
- a heartbeat in scripts/self-update.sh (run_step wrapper) that refreshes
update-status.json every ~3s during the install/build steps with the latest
output line; full output is still mirrored to the update log.
- a frontend (settings-ui.js) that, during non-terminal phases, shows the live
status message plus a ticking total-elapsed counter instead of only the static
phase label.
Takes effect when updating FROM a build that contains it — the detached runner
script (staged from scripts/self-update.sh) and the polling frontend are both
the from-version's copies.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Update Codeman from the web UI: a "Check for updates" button queries GitHub
for the latest tagged release (git ls-remote fallback) and shows release
notes; "Update now" runs git checkout <tag> → npm install → npm run build →
restart, streaming live progress that survives the service restart.
- Release-tag channel; dirty trees auto-stashed (left for manual git stash pop)
- Cross-platform restart: systemd / launchd / manual, detected at runtime
- Updater runs detached (systemd-run --scope on Linux, setsid on macOS) so the
restart it triggers can't kill the build mid-flight
- Build-failure rollback to the pre-update commit; boot reconcile with an
update-id/freshness guard; 409 concurrency lock; runner staged outside the
repo; strict tag validation; CODEMAN_DISABLE_SELF_UPDATE kill-switch
- Endpoints: GET /api/system/update/check, POST /api/system/update,
GET /api/system/update/status
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>