Files
Codeman/.changeset/input-delivery-retryable.md
T
Claudia 84132d3025 fix(ws): pin the withheld ACK with a test, and correct the changeset
Two blockers from the pre-submission gate, both reproduced before fixing.

1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
   code says the opposite in as many words ("NOT an error response,
   deliberately"), the commit message says response codes are unchanged, and the
   test asserts the 200. It was a leftover sentence from an earlier iteration that
   would have shipped into the CHANGELOG announcing an API contract change that
   does not exist — and errorCode values are SemVer-relevant per
   docs/versioning-policy.md.

2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
   master left all 9 tests green, while the commit message sells "plus the whole
   WebSocket path" as part of the fix. Three tests added against the real WS
   route — ACK on delivery, ACK withheld and seq re-opened when the write did not
   land, and a deduplicated frame still ACKed so the client can drop it. Verified
   the other way round: with ws-routes.ts reverted, the middle one fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:52:44 +02:00

1.2 KiB

aicodeman
aicodeman
patch

An input whose delivery fails can be retried instead of being lost for good.

Both input paths recorded the (clientId, seq) pair as applied and acknowledged the frame before knowing whether the write had landed — the POST route because its mux write is fire-and-forget, the WebSocket handler because it ACKed unconditionally. When the write then failed, the client dropped the frame from its durable queue and the server rejected the retry as a duplicate: the reliable delivery layer was guaranteeing exactly-once delivery of something that had never been delivered.

The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so the client redelivers. Session.write() reports whether it reached a PTY at all instead of silently swallowing the data.

Response codes are unchanged: a session can legitimately have no PTY yet (created but not started), so turning that into a failure status would be a contract change of its own.

Note this does not remove the root cause: the POST still answers 200 before the mux write is attempted, so a client that treats any 2xx as final still cannot learn about that failure. Closing that would mean awaiting the tmux child in the request path.