mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-07 16:09:43 +02:00
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no depth limit and no node cap. Two further sites forked a `pgrep` per session on every stats tick. Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in D-state out of ~39,000 total, load average above 13,000, and a machine only recoverable by restarting WSL — which cost every running session. Every diagnostic command timed out too, because they read /proc as well. - ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously with a single-flight guard. Async matters: under the same procfs pathology, `execSync`'s timeout cannot return (spawnSync waits for the unkillable child), which would freeze the server where a hung async poll only costs staleness. - The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a visited set (a stale snapshot can contain a cycle), a depth cap and a node cap, both reporting when they truncate. Pure so the regression tests can exercise the shipped code rather than a copy of it. - The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is bounded, so a wedged `ps` cannot stop killSession from reaching its process-group and tmux fallbacks. - Any `ps` error keeps the previous snapshot instead of caching partial output as fresh; a truncated table would make whole subtrees invisible to the kill path. 13 tests, including one that drives TmuxManager itself — with the caps bypassed at the call site, 3 of them fail. The snapshot refresh is stubbed there, because otherwise the manager runs a real `ps`, replaces the fixture, and the test silently measures the machine's own process tree instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
18 lines
829 B
Markdown
18 lines
829 B
Markdown
---
|
|
'aicodeman': patch
|
|
---
|
|
|
|
Bound the process-tree walk that could take a machine down.
|
|
|
|
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
|
|
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
|
|
and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup`
|
|
under WSL, none returned while the walk kept spawning more — ~13,000 `pgrep`
|
|
processes stuck in D-state out of ~39,000 total, load average above 13,000,
|
|
recoverable only by restarting WSL.
|
|
|
|
Now: one `ps` snapshot, breadth-first with a visited set, a depth cap and a node
|
|
cap, in a pure module (`proc-tree.ts`) that the regression tests exercise
|
|
directly. The snapshot is refreshed asynchronously, and the kill path forces a
|
|
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.
|