`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no depth limit and no node cap. Two further sites forked a `pgrep` per session on every stats tick. Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in D-state out of ~39,000 total, load average above 13,000, and a machine only recoverable by restarting WSL — which cost every running session. Every diagnostic command timed out too, because they read /proc as well. - ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously with a single-flight guard. Async matters: under the same procfs pathology, `execSync`'s timeout cannot return (spawnSync waits for the unkillable child), which would freeze the server where a hung async poll only costs staleness. - The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a visited set (a stale snapshot can contain a cycle), a depth cap and a node cap, both reporting when they truncate. Pure so the regression tests can exercise the shipped code rather than a copy of it. - The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is bounded, so a wedged `ps` cannot stop killSession from reaching its process-group and tmux fallbacks. - Any `ps` error keeps the previous snapshot instead of caching partial output as fresh; a truncated table would make whole subtrees invisible to the kill path. 13 tests, including one that drives TmuxManager itself — with the caps bypassed at the call site, 3 of them fail. The snapshot refresh is stubbed there, because otherwise the manager runs a real `ps`, replaces the fixture, and the test silently measures the machine's own process tree instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
829 B
aicodeman
| aicodeman |
|---|
| patch |
Bound the process-tree walk that could take a machine down.
getChildPids ran pgrep -P <pid> per node and recursed with no visited set, no
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
and because each pgrep blocks in the kernel while reading /proc/<pid>/cgroup
under WSL, none returned while the walk kept spawning more — ~13,000 pgrep
processes stuck in D-state out of ~39,000 total, load average above 13,000,
recoverable only by restarting WSL.
Now: one ps snapshot, breadth-first with a visited set, a depth cap and a node
cap, in a pure module (proc-tree.ts) that the regression tests exercise
directly. The snapshot is refreshed asynchronously, and the kill path forces a
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.