Files
Codeman/.changeset/input-delivery-retryable.md
T
Claudia ebfcac6ad1 fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.

When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.

- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
  is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
  client redelivers.
- `Session.write()` reports whether it reached a PTY.

Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.

What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.

9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:36:33 +02:00

1.2 KiB

aicodeman
aicodeman
patch

An input whose delivery fails can be retried instead of being lost for good.

Both input paths recorded the (clientId, seq) pair as applied and acknowledged the frame before knowing whether the write had landed — the POST route because its mux write is fire-and-forget, the WebSocket handler because it ACKed unconditionally. When the write then failed, the client dropped the frame from its durable queue and the server rejected the retry as a duplicate: the reliable delivery layer was guaranteeing exactly-once delivery of something that had never been delivered.

The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so the client redelivers. Session.write() reports whether it reached a PTY at all instead of silently swallowing the data, and the non-mux POST branch — whose response has not gone out yet — answers OPERATION_FAILED rather than a cheerful 200.

Note this does not remove the root cause: the POST still answers 200 before the mux write is attempted, so a client that treats any 2xx as final still cannot learn about that failure. Closing that would mean awaiting the tmux child in the request path.