Both input paths recorded the (clientId, seq) pair as applied and acknowledged the frame BEFORE knowing whether the write had landed: the POST route because its mux write is fire-and-forget so the response never waits on a tmux child, the WebSocket handler because it ACKed unconditionally. When the write then failed, the client dropped the frame from its durable queue and the server rejected the retry as a duplicate. The reliable-delivery layer was guaranteeing exactly-once delivery of something that had never been delivered — and `Session.write()` returned void, so a session whose PTY was gone swallowed the data with no signal at all. - `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq is still the newest one; a later input has superseded it and must not re-open. - The WebSocket handler withholds its ACK when the write did not land, so the client redelivers. - `Session.write()` reports whether it reached a PTY. Response codes are unchanged, deliberately: a session can legitimately have no PTY yet, and turning that into a failure status would be a contract change of its own. What this does NOT do: remove the root cause. The POST still answers 200 before the mux write is attempted, so a client that treats any 2xx as final cannot learn about that failure. What closes is the narrower window — the write failed AND the ACK never reached the client — plus the whole WebSocket path. Closing the rest would mean awaiting the tmux child inside the request. 9 tests. They drive the HTTP route, not only the Session primitives: with the rollback removed from the route, 2 of them fail. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
1.2 KiB
aicodeman
| aicodeman |
|---|
| patch |
An input whose delivery fails can be retried instead of being lost for good.
Both input paths recorded the (clientId, seq) pair as applied and acknowledged
the frame before knowing whether the write had landed — the POST route because
its mux write is fire-and-forget, the WebSocket handler because it ACKed
unconditionally. When the write then failed, the client dropped the frame from its
durable queue and the server rejected the retry as a duplicate: the reliable
delivery layer was guaranteeing exactly-once delivery of something that had never
been delivered.
The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so
the client redelivers. Session.write() reports whether it reached a PTY at all
instead of silently swallowing the data, and the non-mux POST branch — whose
response has not gone out yet — answers OPERATION_FAILED rather than a cheerful 200.
Note this does not remove the root cause: the POST still answers 200 before the mux write is attempted, so a client that treats any 2xx as final still cannot learn about that failure. Closing that would mean awaiting the tmux child in the request path.