Browser input is delivered exactly once by (clientId, seq). The server records a
watermark per clientId and discards anything not above it as a duplicate — but
acknowledged it with an ACK indistinguishable from "applied". The client then
dropped the record from its queue, the UI looked perfectly normal, and the
terminal received nothing at all.
The counter is persisted to localStorage through a debounced write. Kill the page
between "sent" and "persisted" and the restored counter is below the server's
watermark, after which every keystroke lands under it, is discarded, and is
ACKed. Reloading does not help: the clientId is restored from localStorage
alongside that stale counter. Measured on a real session — typing into the same
session from a fresh browser (new clientId, no watermark on the server) worked
perfectly, which is what localised the fault to client state.
Three changes:
- on rejection the server replies {"t":"ia",seq,"dup":true,"last":<watermark>}.
It still ACKs, so the client can drop the record from its queue, but it now
says the input was not applied and supplies the number needed to climb out.
- on `dup` the client lifts its counter above the watermark and re-queues.
⚠️ Only records whose FIRST delivery is being retried are re-sent: a retry
judged duplicate means the mechanism is working (the original did arrive), and
re-sending would type the same text twice.
- the counter is now persisted synchronously. The queue payload can stay
debounced, but the counter is the thing that has to survive a crash, and
leaving it on the lossiest path cancels the only guarantee there is.
⚠️ Reading the watermark is defensive: the session arrives through a structured
port, and a port missing that method must not take the whole input path down —
a throw inside the handler means the ACK is never sent and the record is stuck in
the client queue forever, which is worse than the ambiguity being fixed. A mock
port's test timeout is what exposed this.
(cherry picked from commit 05bb7081cc)
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Scopes real-time streams and the init snapshot so a multi-user client only
receives what it owns. No-op in single-user mode (identity-less clients).
- WS terminal (ws-routes): owner gate after the session lookup. A non-admin may
only attach to their own session (close 4003); the global auth hook already
ran on the upgrade and decorated req.authUser, so an unauthenticated upgrade
never reaches the handler.
- SSE (sse-stream-manager): per-client identity stored at addClient; broadcast()
and the terminal-batch flush both enforce a routing hint via canDeliver().
WebServer.broadcast auto-derives the hint (deriveSseHint): session-scoped event
families resolve the owner from the payload's session id (fail closed when the
owner can't be resolved), machine-level families (docker/tunnel/update/system/
cron) + host-plan telemetry are admin-only, everything else stays global. Raw
terminal bytes resolve the owner once and are withheld from non-owners.
- getLightState is filtered per connection AFTER the shared cache (sessions,
respawnStatus, subagents, workflowRuns by owner; scheduledRuns + planUsage
admin-only); applied to both the SSE init snapshot and GET /api/status.
- file-routes: getKnownSessionWorkingDir + getSessionAttachmentHistory (the
preview/thumbnail/history helpers that bypass findSessionOrFail) now owner-check
the session, closing a cross-user file-read path.
- GET /api/search: harvestSources is owner-scoped.
Deferred to a follow-up (documented in docs/multi-user-plan.md): away-digest +
subagent/workflow REST list scoping, push-subscription identity + routing,
per-user screenshot subdirs. The live-event versions of these are already routed
by the SSE hint; only the on-demand REST aggregates remain global for admins-only
follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MAX_WS_PER_SESSION was gated by a bare Map<sessionId,number> counter,
incremented on upgrade and decremented only on the old socket's async
close. A client that dropped and immediately reconnected could land its
new upgrade before the old socket's close fired, briefly over-counting and
tripping a spurious 4008 (-> HTTP fallback). The limit also counted raw
sockets, so a reconnecting client consumed a new slot instead of its own.
Replace the counter with WsConnectionRegistry (new pure, unit-tested module)
that tracks live sockets per session keyed by clientId. A same-cid upgrade
SUPERSEDES its own socket (evicts the stale one with close 4010, reuses the
slot, no net count change) -> a reconnect can never be rejected by the cap.
The reliable-input protocol (shouldApplyInput(cid,seq)) already assumes one
logical client per cid per session, so same-cid eviction is principled, not
a regression of multi-tab (which already collides on seq). Slots are freed
EAGERLY on error/terminate, not just async close; close is identity-matched
so a superseded socket's late close is a no-op. cid-less upgrades are
admitted anonymously up to the cap and never evict (backward-compat).
Client sends cid on the WS upgrade URL (?cid=, encoded, omitted if absent).
Tests: ws-connection-registry.test.ts (reconnect-reclaim at cap, rejects
N+1th distinct, eager-terminate frees slot, cid-less up-to-limit + no-evict,
late-close-no-evict, per-session isolation) + route integration in
ws-routes.test.ts (real upgrade through the cap). 45/45 across registry +
ws-routes + input-send-order + ws-reconnect-plan; tsc 0, build, prettier,
frontend-syntax clean.
Root cause of the WS->HTTP->WS flapping: the v1.1.15 input-delivery merge left a
call to the now-undefined _flushHttpFallbackQueuesViaWs() in ws.onopen, so every
(re)connect threw a TypeError BEFORE _onWsReady() ran -- durable input was never
re-flushed over the fresh socket, the 2s redeliver sweep then saw stale unacked
frames and force-closed the socket, reconnect, throw again: a self-sustaining
flap loop. Remove the dead call (_onWsReady, 10 lines below, is its replacement).
Resilience + observability:
- Pure CodemanWsReconnect.plan(code, attempt) (constants.js, TDD, 6 tests):
<4004 -> fast reconnect (immediate jittered first retry, faster backoff);
4008/unknown->=4004 -> bounded retry-fallback (HTTP no longer sticks until a
tab switch); 4004/4009 -> give up (session gone). Wired into onclose.
- Redeliver sweep force-closes only a SILENT socket (no recent recv), not one
actively delivering output/ACKs -- stops self-inflicted flaps while typing.
- Client logs WS close code/reason to crash-diag; server logs [ws]
open/close/terminate/4008 (console -> journald; Fastify runs logger:false).
Verified: 6/6 unit, tsc 0, frontend-syntax + prettier clean, build; beta WS
reaches connected with zero console errors (onopen TypeError gone),
_wsLastRecvAt tracked, server [ws] lines emit.
A "sent" prompt could vanish with no trace on a flaky connection (e.g. a train):
with local echo on, Enter cleared the overlay then sent over the WebSocket
fire-and-forget. On a half-open socket (readyState===OPEN, dead TCP) ws.send()
doesn't throw, so the frame was silently discarded, nothing was enqueued, and
navigator.onLine stayed true — the prompt was lost and never resent.
Replace the best-effort offline queue with a durable, acknowledged delivery layer:
- Client (app.js): every input frame is recorded with a stable clientId +
monotonic per-session seq and persisted to localStorage BEFORE delivery, and
only dropped on a server ACK. Delivered over WS (acked via {t:'ia',seq}) or,
when the socket is down, POST in seq order (HTTP 2xx = ACK). A 2s sweep
force-reconnects a WS whose oldest frame is unacked past 4s (half-open sockets
never recover on their own); on reconnect/reload all pending frames re-deliver.
Survives reconnects AND page reloads. Connection indicator shows pending count.
- Server: Session.shouldApplyInput(clientId, seq) applies each frame exactly once
(bounded MRU map); ws-routes + POST /input dedup a redelivered seq but still ACK
it (200 / {t:'ia'}), so an at-least-once resend can never type the prompt twice.
Untagged input (curl/legacy) applies unconditionally — no behavior change.
- terminal-ui.js sendInput() (voice / keyboard-accessory / paste) now routes
through the same durable layer.
Tests: test/reliable-input-dedup.test.ts (exactly-once semantics on the real
Session) + POST /input dedup route tests. Design: docs/reliable-input-delivery.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Auto-resume on usage limit ("token pause" control, opt-in checkbox at the
top of the Respawn tab, off by default):
- usage-limit-patterns.ts (new, pure): detects all Claude Code limit
messages (1.0.x-2.1.x eras incl. "5-hour limit reached - resets 8pm",
"You've hit your limit - resets 1:40pm (TZ)", weekly date forms, raw
"usage limit reached|<epoch>") and parses the reset time. Conservative:
no parseable future reset time, no action.
- SessionAutoOps: arms a timer at reset+2min, sends Esc (dismisses the
rate-limit dialog) + "continue"; dedups footer redraws, retries every
5min on stale times, cancels when Claude starts working, persists and
re-arms across Codeman restarts (SessionState.autoResumeEnabled/At).
- Respawn guard: cycles are blocked while limit-paused so /clear cannot
wipe the paused conversation (respawnBlocked reason 'usage_limit').
- POST /api/sessions/:id/auto-resume; SSE session:limitPauseScheduled/
limitResume/limitResumeCancelled; toasts + status line in the modal.
- Respawn tab tidied: single-row prompt fields, merged behavior row.
Mobile fixes (0.9.8 regressions, user-reported):
- Resize arbitration is now activity-based: a desktop sizing claim only
blocks phone resizes while the desktop typed within 90s
(Session.DESKTOP_CLAIM_IDLE_MS). Idle desktop -> phone takes the pane;
next desktop keystroke re-asserts the desktop layout server-side
(noteDesktopActivity via ws-routes input). Phones re-send dims every
30s (visible tab only, skipped while the keyboard is open) so attaching
under a hot claim self-corrects. Fixes the desktop-width-stream-in-
narrow-xterm soup (mid-word wraps, tmux dot fill, Ink overdraw).
- Cross-device reflows (takeover/re-assert) emit a debounced needsRefresh
so all clients reload the buffer instead of stacking ghost Ink frames.
- Keyboard accessory/toolbar lift restored: measure keyboardOffset
against window.innerHeight (layout viewport), not the shrunken .app -
on iOS the offset computed to 0, leaving both bars hidden behind the
OS keyboard with a dead gap above.
- Removed the mobile header utility ("three dots") toggle entirely;
the headerRight tray stays collapsed on small viewports.
Tests: usage-limit-patterns (36), session-auto-resume (21), resize
arbitration (+6), session routes (+4), respawn guard (+2); MockSession
auto-resume/sizing stubs; mobile tabs test updated for toggle removal.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review follow-ups on PR #111 (rebased onto master post-#112/#113):
Resize arbitration redesigned (review blocker 2): the previous
'cols < _ptyCols' guard froze a mobile-only session's PTY at the spawn
default — narrow phones rendered clipped and could never re-fit. The
guard now uses connection-scoped desktop sizing claims instead:
ws-routes registers a claim on a desktop-typed resize and releases it
on socket close (or when the same connection later reports a small
viewport), and Session.resize() ignores mobile/tablet resizes only
while at least one desktop connection holds a claim. A phone alone
fully controls its size (shrink, rows-only shrink, re-grow); a phone
glancing at a desktop-driven session can no longer reflow it.
mobile-handlers' keyboard open/close resize now declares its viewport
type so it participates in arbitration. Tests rewritten to cover
mobile-only shrink/rows-only/re-grow, claim/release lifecycle, multi-
claim behavior, and untyped legacy resizes; ws-routes test covers the
claim lifecycle over a real socket.
Solo/detached header restored (review blocker 3): index.html had
removed #soloSessionTitle and #soloRedockBtn, which _applySoloMode
still references — every detached window hit a null deref. Both are
back alongside the new mobile utility toggle.
Desktop leak fixed (review should-fix): .mobile-header-utility-toggle
had no rule outside the <=768px media queries, so the raw button
rendered on desktop. styles.css now hides it by default; the mobile/
tablet queries re-enable it.
Visual-regression baselines reverted to master (review should-fix):
the 18 contributor-machine PNGs are environment-specific (8 of the
behavioral tests already report environment-sensitive failures across
machines); re-baseline deliberately on the canonical machine instead.
The 24 behavioral keyboard/layout/tabs tests are kept as-is.
AGENTS.md trimmed to a pointer at CLAUDE.md (review should-fix) to
avoid drift between duplicated guidance.
Also dropped a dead getAttachmentHistoryForPersist stub (codex-branch
residue — no such method exists in src/).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Mobile-focused fixes for the web UI: keyboard-accessory layout and
overlap, native input visibility above the keyboard, CJK input handling,
terminal touch scrolling, tab-menu tap targets, mic-recording glow
containment, and mobile resize/keyboard-state handling on tab switch,
plus mobile visual-regression test coverage and snapshots.
Co-Authored-By: Saqeb Akhter <saqeb.akhter@gmail.com>
Adds an always-on Host-header allowlist and a cross-site Origin/CSRF guard,
hardens the text/plain body parser, validates the WebSocket upgrade origin,
and escapes AI-derived fields in the subagent panel. Closes the two
CRITICALs and 5 HIGHs from the 2026-06-09 adversarial security review.
- C1: no Host allowlist -> DNS rebinding drove the full API (RCE) on the
default no-auth loopback install. New registerHostGuard rejects rebound
custom domains; allows loopback, any IP literal, the bind host,
*.ts.net / *.trycloudflare.com / *.cfargotunnel.com, the active managed
tunnel, and CODEMAN_ALLOWED_HOSTS.
- C2: a global text/plain parser JSON-parsed every body, enabling cross-site
simple-request CSRF. Parser now keeps the raw string; /api/crash-diag
self-parses; the global Origin guard rejects cross-site state changes.
- H1/H3/H6: self-update, session create/input, and settings/tunnel toggles
were CSRF-triggerable -> now covered by the Origin guard.
- H4: the subagent activity panel injected raw AI tool names/inputs into
innerHTML (executed under CSP 'unsafe-inline'). All sinks now escapeHtml'd.
- H5: the WebSocket upgrade had no Origin/Host check (CSWSH) -> now validated.
A missing Origin is allowed so curl/CLI and Claude Code hooks keep working;
custom reverse-proxy domains need CODEMAN_ALLOWED_HOSTS=host,.suffix.
Deferred: H2 (self-update tag signing, needs signing infra) and CSP
'unsafe-inline' removal (needs a nonce migration).
Tests: test/network-host-guard.test.ts (19), test/routes/ws-routes.test.ts
updated. Report: docs/reports/security-review-2026-06-09.md
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Close WebSocket when session exits (exit event listener) to prevent
orphaned listeners and stale writes to dead PTY
- Add readyState guard in onTerminal to stop buffering after socket closes
- Simplify heartbeat: remove redundant alive flag, use pongTimeout only
- Add exponential backoff reconnection on unexpected WS close (skip for
server rejections 4004/4008/4009)
- Clear CJK textarea on session switch to prevent wrong-session input
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
WebSocket route: add socket error handler to prevent process crashes, enforce
per-session connection limit (max 5), track/decrement counts on close.
CJK input: add destroy() method with proper listener cleanup, guard against
double-init, add maxlength/aria-label to textarea, use language-neutral
placeholder, explicitly clear cjkActive on hide.
install.sh: fix update() to use $BRANCH and $REPO_URL instead of hardcoded
origin/master — fork users were silently switched back to master on update.
README: fix broken markdown table (paragraph concatenated into last cell),
add CODEMAN_NODE_VERSION to env var table.
Tests: add 8 new test cases for batch coalescing, flush threshold, unknown
message types, connection limit, heartbeat, readyState guards. Import
MAX_INPUT_LENGTH from config, add connectWs timeout, replace setTimeout
with vi.waitFor in cleanup test.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Detect stale connections that TCP keepalive won't catch for minutes,
especially through tunnels and proxies. Pings every 30s with a 10s
pong timeout — if the client doesn't respond, the socket is terminated
and all timers cleaned up.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The HTTP resize route validates via ResizeSchema (cols: 1-500, rows:
1-200, integers only). The WS handler only checked typeof === 'number',
allowing floats, negatives, and extreme values through to ptyProcess.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace per-keystroke HTTP POST + SSE terminal output with a single
bidirectional WebSocket connection for dramatically lower input latency.
The existing SSE+POST paths remain fully functional as fallback.
Server-side: ws-routes.ts provides /ws/sessions/:id/terminal with 8ms
micro-batching and 16KB flush threshold. Each batch is wrapped in
DEC 2026 synchronized update markers so xterm.js renders atomically —
Ink's DA capability negotiation fails through the PTY→server→WS proxy
chain, so without server-injected markers, cursor-up redraws flicker.
Frontend: _connectWs/_disconnectWs manage per-session WS lifecycle.
Input and resize use WS fast path with HTTP POST fallback. SSE terminal
events are suppressed when WS is active to prevent double rendering.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>