Two blockers from the pre-submission gate, both reproduced before fixing.
1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
code says the opposite in as many words ("NOT an error response,
deliberately"), the commit message says response codes are unchanged, and the
test asserts the 200. It was a leftover sentence from an earlier iteration that
would have shipped into the CHANGELOG announcing an API contract change that
does not exist — and errorCode values are SemVer-relevant per
docs/versioning-policy.md.
2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
master left all 9 tests green, while the commit message sells "plus the whole
WebSocket path" as part of the fix. Three tests added against the real WS
route — ACK on delivery, ACK withheld and seq re-opened when the write did not
land, and a deduplicated frame still ACKed so the client can drop it. Verified
the other way round: with ws-routes.ts reverted, the middle one fails.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.
Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.
- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
with a single-flight guard. Async matters: under the same procfs pathology,
`execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
both reporting when they truncate. Pure so the regression tests can exercise the
shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
bounded, so a wedged `ps` cannot stop killSession from reaching its
process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
fresh; a truncated table would make whole subtrees invisible to the kill path.
13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.
The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.
Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.
Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.
Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.
Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.
Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only CLAUDE.md conflicted: master restructured it into the short-rule +
docs/architecture-invariants.md pointer layout while this PR was open.
The response-viewer detail now lives in architecture-invariants, so the
Claude turn-grouping and restored-placeholder rebind notes moved there.
Changeset rewritten to record the measured effect on real transcripts.
Both picker endpoints are a second file-serving surface, and they
inherited neither the attachment guard's confinement nor its ownership
scoping. Two separate holes:
1. `sessionId` contributes that session's workingDir as a browse root,
but it was resolved straight off ctx.sessions/ctx.store with no owner
check, unlike the nine other session-scoped handlers in this file. A
non-admin could pin ANOTHER user's working directory as a root just
by passing their session id, then list and preview underneath it. Now
runs canAccessOwned and reports 404, which also avoids confirming
that a session id exists.
2. `Home` and `CASES_DIR` were unconditional roots for every caller.
Per-user spaces live at <USER_SPACES_DIR>/<username>, which is INSIDE
homedir(), so the Home root alone exposed every other user's
workspace. A multi-user non-admin now gets only their own
userSpacePath plus anything explicitly listed in
CODEMAN_FILE_PICKER_ROOTS. /mnt/d is dropped as well: a broad host
mount should be an explicit operator decision in a multi-user
deployment, and operators who want it can name it in that env var.
Admins and single-user mode keep the host-wide roots, so behavior is
unchanged unless CODEMAN_MULTIUSER is on (opt-in, off by default).
All three discriminating tests were verified to fail against the
previous code: browse and preview both returned 200 instead of 404, and
the roots came back as [Home, Codeman Cases, ...] instead of [My Space].
Full suite green, 3784 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The opt-in multi-user feature's only enforcement is web-layer scoping
(all sessions share one OS account). An adversarial review found 8 critical
+ 7 high cross-user holes that defeated it, plus mediums; all fixed here.
Single-user (flag-off) behavior stays byte-identical apart from documented
consistency deltas.
Ownership / confinement:
- DELETE /api/sessions (bulk) + /:id now owner-scope / findSessionOrFail
- quick-start, cron (create+fire), scheduled runs confine workingDir to the
owner's space; case link/docker-link/docker-import confine the host path
- resolveCasePath no longer resolves linked cases for non-admins; foreign
remote/docker cases are skipped (fall through to the caller's own local case)
- history, subagents/workflows, mux-sessions, orchestrator, cron run-history,
away-digest, and remote/docker host reads are owner- or admin-scoped
Permission policy (section 6.3):
- non-granted users are downgraded at every spawn site incl. legacy
/api/scheduled, PlanOrchestrator one-shots, remote launch, and the cron-fire
gemini/codex bypass switches; resolveClaudeModeForUsername now fails closed
Auth / store:
- verify-first login throttle (a correct password is never locked out),
/ws terminal subject to the change-password lockbox, cookie fast-path
re-validates identity live, role/grant changes revoke sessions, admin delete
runs the last-admin guard before any teardown
- users.json: distinguish missing (ENOENT) from corrupt/unreadable so a bad
read can't overwrite all accounts; unique per-process temp write path
Event streams:
- debounced session:updated + batched task:updated, clipboard, and push
notifications route by owner (fail closed); getLightState hides machine-wide
globalStats from non-admins
Tests: two suites updated to assert the fixed (secure) behavior. tsc, eslint,
and test:ci all green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stamp the plan doc with shipped-by-phase status; add the multi-user Key Patterns
entry + State Files + case-spaces note to CLAUDE.md; add a multi-user section to
the security architecture (threat model: workspace separation, not a security
boundary) and a README opt-in section; add a minor changeset.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Release 1.3.5. Consumes the changeset from PR #155: re-issue the
codeman_session cookie on every authenticated request so the browser cookie
lifetime tracks the server-side sliding TTL, fixing the recurring native Basic
Auth dialog during active use.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The codeman_session cookie was only set on the Basic Auth path with a fixed
lifetime from login and never refreshed, while the server-side session store
slides its TTL (refreshOnGet). So the browser cookie expired mid-use, the next
request arrived cookie-less and fell through to Basic Auth, popping the native
username/password dialog — perceived as a random logout while actively working.
Re-issue the cookie on every authenticated (valid-cookie) request so the browser
lifetime tracks the server-side sliding TTL.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>