mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
feat(cli): codeman web -d and codeman service install (#231)
Two ways to keep the server running, split by how long it should last. `codeman web -d` relaunches the same entry script detached (setsid), with `--stop` and `--status` alongside it. A pidfile and log live in the data dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its default disposition even when it inherits "ignore", and cli.ts handles SIGHUP with a graceful shutdown, so a delivered HUP still stops the server. Removing the shell's ability to send one is the fix. `codeman service install|uninstall|status` writes and loads the systemd user unit or the LaunchAgent, with the installing shell's PATH baked in (launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner installs; this is for npm globals. Both refuse to start when a server is already up on the data dir, since a second instance on the shared tmux socket attaches PTYs to the first one's live sessions. Both poll /api/status until the child answers or dies rather than reporting a success they have not seen. `--stop` checks the pid still looks like a Codeman server before signalling it. The systemd unit name and launchd label move to config/service-names.ts so install.sh, detectSupervisor() and service install cannot drift into supervising two copies. Instance-scoped, unchanged for the default instance. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -140,6 +140,22 @@ Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write
|
||||
|
||||
**Away digest** (COD-41/#136): `GET /api/away-digest?range=&since=&until=&lastViewed=` aggregates "what happened while you were away" from the lifecycle log + run-summary events + live sessions + daily token stats + recently-completed subagents into needs-attention/completed/still-running/idle/informational sections. Pure aggregator in `web/away-digest.ts` (`resolveAwayDigestRange()` validates the window — `since-last-visit`/`1h`/`today`/`24h`/`custom`, server-local TZ; `buildAwayDigest()` classifies). Header-button modal in `panels-ui.js` (button hidden on phones — regression-guarded). ⚠️ Returns `{success:true,digest}` (a legacy raw-ish shape, consistent with the other raw GET handlers in `system-routes.ts` — `{entries}`/`{config}`/`{files}`/`getSystemStats()`); frontend + tests read `.digest`. Subagent lookback is a fixed 60-min window regardless of range.
|
||||
|
||||
### Detached start and service install
|
||||
|
||||
**`codeman web -d` / `codeman service install`** (issue #231). Two answers to "keep it running", split by how long: `-d` survives the shell, the service survives a reboot. `src/daemon-control.ts` and `src/service-installer.ts`, both splitting pure builders (argv, URLs, pidfile parsing, unit-file text) from the IO.
|
||||
|
||||
Why a flag at all, when `nohup codeman web &` looks like it should work: it does not reliably. **Node re-arms SIGHUP to its default disposition even when it inherits "ignore" from `nohup`** (verified: `nohup node script.js &` then `kill -HUP` prints "Hangup" and dies; `/proc/<pid>/status` shows SIGHUP absent from `SigIgn`, where `nohup sleep` has it set). `cli.ts` then adds a SIGHUP handler that shuts the server down gracefully, so a delivered HUP always stops it. What actually works is removing the shell's ability to send one: `disown` in the user's shell, or `detached: true` (setsid) here. zsh HUPs running jobs on exit by default, bash does not on a clean `exit` but does when it receives SIGHUP itself, which is why "does `&` survive?" gets opposite answers on macOS and Linux.
|
||||
|
||||
Invariants:
|
||||
|
||||
- **Never a second server on one data dir.** `-d` and `service install` both check the pidfile AND probe `/api/status` first, and refuse. This is the instance-isolation hazard, not politeness: a second instance on the shared `tmux -L codeman` socket discovers the first one's live sessions, attaches PTYs to them and resizes them (see [Instance isolation](#instance-isolation-and-the-multi-instance-attach-danger)).
|
||||
- **Never report success that was not observed.** The parent polls `/api/status` until the child answers or exits, then prints the URL or the tail of `web.log`. A `launchctl load`, a `systemctl enable --now` and a plain spawn are all silent about a server that starts and dies half a second later, which is why `install.sh` verifies too. The log is append-only across launches, so each start writes a separator line and the failure tail begins there.
|
||||
- **A 401 counts as up.** `CODEMAN_PASSWORD` gates `/api/status`, so requiring a 200 would make readiness detection fail on exactly the installs that took security advice. The body is still checked for `"success"` so an unrelated service squatting on the port is not mistaken for Codeman.
|
||||
- **`--stop` verifies identity before signalling.** Pids are recycled; a stale pidfile plus a blind `process.kill` is how a tool SIGTERMs someone's database. `ps -o command=` (portable to macOS) must still look like a Codeman web process. When `ps` itself fails, the pid is treated as ours rather than orphaning the pidfile.
|
||||
- **One source of truth for the job name.** `src/config/service-names.ts` holds the systemd unit name and launchd label used by `install.sh`, `detectSupervisor()` in self-update, and `service install`. Drift here is silent and bad: `service install` would supervise a SECOND copy alongside the installer's. The names are instance-scoped (`CODEMAN_INSTANCE=beta` → `codeman-web-beta.service` / `com.codeman.beta.web`) so a beta cannot overwrite the production unit, and are byte-identical to the historical names for the default instance.
|
||||
- **PATH is the reason hand-written units fail.** launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` and systemd's user manager is nearly as bare, so a Homebrew or nvm `node`, `tmux` or `claude` is simply absent. The unit therefore carries the installing shell's PATH with the running node's directory in front; `node_modules/.bin` entries are dropped, since npx injects those for one command and they would outlive the checkout.
|
||||
- **No secrets in unit files.** `CODEMAN_PASSWORD` present in the installing shell is NOT copied into the plist/unit; the operator is told to add it. `CODEMAN_INSTANCE` IS copied, because without it a supervised beta would silently run against the production data dir and tmux socket.
|
||||
|
||||
### Self-update
|
||||
|
||||
**Self-update** (App Settings → Updates): in-app updater for **git-clone installs** supervised by systemd/launchd. Supervisors: `systemd` (user unit), `launchd` (GUI LaunchAgent, gui-domain kickstart), `launchd-daemon` (KeepAlive system LaunchDaemon on headless Macs — restarts rootlessly by killing the server PID and letting launchd respawn it; detected only when the daemon is bootstrapped AND KeepAlive), else `none` → "restart manually" message; on next boot a manual-restart status auto-completes when the running version matches the target. The update restarts the very process running it, so the real work runs in a DETACHED `scripts/self-update.sh` (`git checkout <release tag> && npm install && npm run build && restart`) that outlives the restart; it writes progress to `dataPath('update-status.json')`, which the browser polls across the connection drop. Channel = latest `codeman@X.Y.Z` release tag; dirty trees are auto-stashed. `src/web/self-update.ts` splits PURE helpers (semver/tag parsing, reconcile decision — unit-tested) from IO wrappers (`getInstallInfo`/`checkForUpdate`/`startUpdate`/`reconcileUpdateOnBoot`). Routes: `GET /api/system/update/check`, `POST /api/system/update`, `GET /api/system/update/status`. Types: `src/types/update.ts`. npm installs report as non-updatable.
|
||||
|
||||
Reference in New Issue
Block a user