fix(self-update): stalled status and hung shutdown on launchd-daemon installs (#478)

* fix(self-update): stop a stalled status from blocking every later update

A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.

- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
  is older than the stale window; applied on every read (start + status
  poll) and persisted. The live updater heartbeats every few seconds, so a
  running update never trips it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down

On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.

- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
  server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
  then SIGKILL it. tmux sessions live outside the server and survive.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Codeman maintainer <noreply@anthropic.com>
This commit is contained in:
Ark0N
2026-09-24 01:35:27 +02:00
committed by GitHub
co-authored by Claude Opus 5.5
parent dd230b0b6e
commit c46e87fd7a
4 changed files with 106 additions and 3 deletions
+20 -1
View File
@@ -74,6 +74,15 @@ echo "[self-update] $(date) start tag=$TAG supervisor=$SUPERVISOR repo=$REPO"
export PATH="$(dirname "$NODE"):$HOME/.local/bin:$HOME/.npm-global/bin:/usr/local/bin:/opt/homebrew/bin:$PATH"
export GIT_TERMINAL_PROMPT=0
# --node is the server's process.execPath, a VERSIONED path (Homebrew resolves it
# into Cellar/node/<ver>/). A `brew upgrade node` under a long-running server
# deletes it, and every status write then failed, so the status stayed "queued"
# forever. Fall back to whatever node is on PATH.
if [ ! -x "$NODE" ]; then
echo "[self-update] WARN: $NODE is not executable, falling back to node on PATH"
NODE="$(command -v node || echo node)"
fi
TO_VERSION="${TAG##*@}" # codeman@0.9.4 → 0.9.4 (tag is validated upstream)
STASH_REF=""
MANUAL_CMD=""
@@ -276,7 +285,17 @@ case "$SUPERVISOR" in
# domain needs root, but we don't need it — kill the server and launchd
# respawns it on the new dist/ within ThrottleInterval seconds.
if [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null; then
: # respawn is launchd's job from here
# Respawn is launchd's job, but only once the old process EXITS. A graceful
# shutdown that hangs leaves the port closed and the service down, so
# escalate to SIGKILL (tmux sessions live outside the server and survive).
for _ in $(seq 1 30); do
kill -0 "$SERVER_PID" 2>/dev/null || break
sleep 1
done
if kill -0 "$SERVER_PID" 2>/dev/null; then
echo "[self-update] server pid $SERVER_PID still alive 30s after SIGTERM, sending SIGKILL"
kill -9 "$SERVER_PID" 2>/dev/null || true
fi
else
MANUAL_CMD="sudo launchctl kickstart -k system/com.codeman.web"
write_status "completed-needs-manual-restart" "Update staged — restart Codeman to apply v$TO_VERSION."