Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Scopes real-time streams and the init snapshot so a multi-user client only
receives what it owns. No-op in single-user mode (identity-less clients).
- WS terminal (ws-routes): owner gate after the session lookup. A non-admin may
only attach to their own session (close 4003); the global auth hook already
ran on the upgrade and decorated req.authUser, so an unauthenticated upgrade
never reaches the handler.
- SSE (sse-stream-manager): per-client identity stored at addClient; broadcast()
and the terminal-batch flush both enforce a routing hint via canDeliver().
WebServer.broadcast auto-derives the hint (deriveSseHint): session-scoped event
families resolve the owner from the payload's session id (fail closed when the
owner can't be resolved), machine-level families (docker/tunnel/update/system/
cron) + host-plan telemetry are admin-only, everything else stays global. Raw
terminal bytes resolve the owner once and are withheld from non-owners.
- getLightState is filtered per connection AFTER the shared cache (sessions,
respawnStatus, subagents, workflowRuns by owner; scheduledRuns + planUsage
admin-only); applied to both the SSE init snapshot and GET /api/status.
- file-routes: getKnownSessionWorkingDir + getSessionAttachmentHistory (the
preview/thumbnail/history helpers that bypass findSessionOrFail) now owner-check
the session, closing a cross-user file-read path.
- GET /api/search: harvestSources is owner-scoped.
Deferred to a follow-up (documented in docs/multi-user-plan.md): away-digest +
subagent/workflow REST list scoping, push-subscription identity + routing,
per-user screenshot subdirs. The live-event versions of these are already routed
by the SSE hint; only the on-demand REST aggregates remain global for admins-only
follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MAX_WS_PER_SESSION was gated by a bare Map<sessionId,number> counter,
incremented on upgrade and decremented only on the old socket's async
close. A client that dropped and immediately reconnected could land its
new upgrade before the old socket's close fired, briefly over-counting and
tripping a spurious 4008 (-> HTTP fallback). The limit also counted raw
sockets, so a reconnecting client consumed a new slot instead of its own.
Replace the counter with WsConnectionRegistry (new pure, unit-tested module)
that tracks live sockets per session keyed by clientId. A same-cid upgrade
SUPERSEDES its own socket (evicts the stale one with close 4010, reuses the
slot, no net count change) -> a reconnect can never be rejected by the cap.
The reliable-input protocol (shouldApplyInput(cid,seq)) already assumes one
logical client per cid per session, so same-cid eviction is principled, not
a regression of multi-tab (which already collides on seq). Slots are freed
EAGERLY on error/terminate, not just async close; close is identity-matched
so a superseded socket's late close is a no-op. cid-less upgrades are
admitted anonymously up to the cap and never evict (backward-compat).
Client sends cid on the WS upgrade URL (?cid=, encoded, omitted if absent).
Tests: ws-connection-registry.test.ts (reconnect-reclaim at cap, rejects
N+1th distinct, eager-terminate frees slot, cid-less up-to-limit + no-evict,
late-close-no-evict, per-session isolation) + route integration in
ws-routes.test.ts (real upgrade through the cap). 45/45 across registry +
ws-routes + input-send-order + ws-reconnect-plan; tsc 0, build, prettier,
frontend-syntax clean.
Root cause of the WS->HTTP->WS flapping: the v1.1.15 input-delivery merge left a
call to the now-undefined _flushHttpFallbackQueuesViaWs() in ws.onopen, so every
(re)connect threw a TypeError BEFORE _onWsReady() ran -- durable input was never
re-flushed over the fresh socket, the 2s redeliver sweep then saw stale unacked
frames and force-closed the socket, reconnect, throw again: a self-sustaining
flap loop. Remove the dead call (_onWsReady, 10 lines below, is its replacement).
Resilience + observability:
- Pure CodemanWsReconnect.plan(code, attempt) (constants.js, TDD, 6 tests):
<4004 -> fast reconnect (immediate jittered first retry, faster backoff);
4008/unknown->=4004 -> bounded retry-fallback (HTTP no longer sticks until a
tab switch); 4004/4009 -> give up (session gone). Wired into onclose.
- Redeliver sweep force-closes only a SILENT socket (no recent recv), not one
actively delivering output/ACKs -- stops self-inflicted flaps while typing.
- Client logs WS close code/reason to crash-diag; server logs [ws]
open/close/terminate/4008 (console -> journald; Fastify runs logger:false).
Verified: 6/6 unit, tsc 0, frontend-syntax + prettier clean, build; beta WS
reaches connected with zero console errors (onopen TypeError gone),
_wsLastRecvAt tracked, server [ws] lines emit.
A "sent" prompt could vanish with no trace on a flaky connection (e.g. a train):
with local echo on, Enter cleared the overlay then sent over the WebSocket
fire-and-forget. On a half-open socket (readyState===OPEN, dead TCP) ws.send()
doesn't throw, so the frame was silently discarded, nothing was enqueued, and
navigator.onLine stayed true — the prompt was lost and never resent.
Replace the best-effort offline queue with a durable, acknowledged delivery layer:
- Client (app.js): every input frame is recorded with a stable clientId +
monotonic per-session seq and persisted to localStorage BEFORE delivery, and
only dropped on a server ACK. Delivered over WS (acked via {t:'ia',seq}) or,
when the socket is down, POST in seq order (HTTP 2xx = ACK). A 2s sweep
force-reconnects a WS whose oldest frame is unacked past 4s (half-open sockets
never recover on their own); on reconnect/reload all pending frames re-deliver.
Survives reconnects AND page reloads. Connection indicator shows pending count.
- Server: Session.shouldApplyInput(clientId, seq) applies each frame exactly once
(bounded MRU map); ws-routes + POST /input dedup a redelivered seq but still ACK
it (200 / {t:'ia'}), so an at-least-once resend can never type the prompt twice.
Untagged input (curl/legacy) applies unconditionally — no behavior change.
- terminal-ui.js sendInput() (voice / keyboard-accessory / paste) now routes
through the same durable layer.
Tests: test/reliable-input-dedup.test.ts (exactly-once semantics on the real
Session) + POST /input dedup route tests. Design: docs/reliable-input-delivery.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Auto-resume on usage limit ("token pause" control, opt-in checkbox at the
top of the Respawn tab, off by default):
- usage-limit-patterns.ts (new, pure): detects all Claude Code limit
messages (1.0.x-2.1.x eras incl. "5-hour limit reached - resets 8pm",
"You've hit your limit - resets 1:40pm (TZ)", weekly date forms, raw
"usage limit reached|<epoch>") and parses the reset time. Conservative:
no parseable future reset time, no action.
- SessionAutoOps: arms a timer at reset+2min, sends Esc (dismisses the
rate-limit dialog) + "continue"; dedups footer redraws, retries every
5min on stale times, cancels when Claude starts working, persists and
re-arms across Codeman restarts (SessionState.autoResumeEnabled/At).
- Respawn guard: cycles are blocked while limit-paused so /clear cannot
wipe the paused conversation (respawnBlocked reason 'usage_limit').
- POST /api/sessions/:id/auto-resume; SSE session:limitPauseScheduled/
limitResume/limitResumeCancelled; toasts + status line in the modal.
- Respawn tab tidied: single-row prompt fields, merged behavior row.
Mobile fixes (0.9.8 regressions, user-reported):
- Resize arbitration is now activity-based: a desktop sizing claim only
blocks phone resizes while the desktop typed within 90s
(Session.DESKTOP_CLAIM_IDLE_MS). Idle desktop -> phone takes the pane;
next desktop keystroke re-asserts the desktop layout server-side
(noteDesktopActivity via ws-routes input). Phones re-send dims every
30s (visible tab only, skipped while the keyboard is open) so attaching
under a hot claim self-corrects. Fixes the desktop-width-stream-in-
narrow-xterm soup (mid-word wraps, tmux dot fill, Ink overdraw).
- Cross-device reflows (takeover/re-assert) emit a debounced needsRefresh
so all clients reload the buffer instead of stacking ghost Ink frames.
- Keyboard accessory/toolbar lift restored: measure keyboardOffset
against window.innerHeight (layout viewport), not the shrunken .app -
on iOS the offset computed to 0, leaving both bars hidden behind the
OS keyboard with a dead gap above.
- Removed the mobile header utility ("three dots") toggle entirely;
the headerRight tray stays collapsed on small viewports.
Tests: usage-limit-patterns (36), session-auto-resume (21), resize
arbitration (+6), session routes (+4), respawn guard (+2); MockSession
auto-resume/sizing stubs; mobile tabs test updated for toggle removal.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review follow-ups on PR #111 (rebased onto master post-#112/#113):
Resize arbitration redesigned (review blocker 2): the previous
'cols < _ptyCols' guard froze a mobile-only session's PTY at the spawn
default — narrow phones rendered clipped and could never re-fit. The
guard now uses connection-scoped desktop sizing claims instead:
ws-routes registers a claim on a desktop-typed resize and releases it
on socket close (or when the same connection later reports a small
viewport), and Session.resize() ignores mobile/tablet resizes only
while at least one desktop connection holds a claim. A phone alone
fully controls its size (shrink, rows-only shrink, re-grow); a phone
glancing at a desktop-driven session can no longer reflow it.
mobile-handlers' keyboard open/close resize now declares its viewport
type so it participates in arbitration. Tests rewritten to cover
mobile-only shrink/rows-only/re-grow, claim/release lifecycle, multi-
claim behavior, and untyped legacy resizes; ws-routes test covers the
claim lifecycle over a real socket.
Solo/detached header restored (review blocker 3): index.html had
removed #soloSessionTitle and #soloRedockBtn, which _applySoloMode
still references — every detached window hit a null deref. Both are
back alongside the new mobile utility toggle.
Desktop leak fixed (review should-fix): .mobile-header-utility-toggle
had no rule outside the <=768px media queries, so the raw button
rendered on desktop. styles.css now hides it by default; the mobile/
tablet queries re-enable it.
Visual-regression baselines reverted to master (review should-fix):
the 18 contributor-machine PNGs are environment-specific (8 of the
behavioral tests already report environment-sensitive failures across
machines); re-baseline deliberately on the canonical machine instead.
The 24 behavioral keyboard/layout/tabs tests are kept as-is.
AGENTS.md trimmed to a pointer at CLAUDE.md (review should-fix) to
avoid drift between duplicated guidance.
Also dropped a dead getAttachmentHistoryForPersist stub (codex-branch
residue — no such method exists in src/).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Mobile-focused fixes for the web UI: keyboard-accessory layout and
overlap, native input visibility above the keyboard, CJK input handling,
terminal touch scrolling, tab-menu tap targets, mic-recording glow
containment, and mobile resize/keyboard-state handling on tab switch,
plus mobile visual-regression test coverage and snapshots.
Co-Authored-By: Saqeb Akhter <saqeb.akhter@gmail.com>
Adds an always-on Host-header allowlist and a cross-site Origin/CSRF guard,
hardens the text/plain body parser, validates the WebSocket upgrade origin,
and escapes AI-derived fields in the subagent panel. Closes the two
CRITICALs and 5 HIGHs from the 2026-06-09 adversarial security review.
- C1: no Host allowlist -> DNS rebinding drove the full API (RCE) on the
default no-auth loopback install. New registerHostGuard rejects rebound
custom domains; allows loopback, any IP literal, the bind host,
*.ts.net / *.trycloudflare.com / *.cfargotunnel.com, the active managed
tunnel, and CODEMAN_ALLOWED_HOSTS.
- C2: a global text/plain parser JSON-parsed every body, enabling cross-site
simple-request CSRF. Parser now keeps the raw string; /api/crash-diag
self-parses; the global Origin guard rejects cross-site state changes.
- H1/H3/H6: self-update, session create/input, and settings/tunnel toggles
were CSRF-triggerable -> now covered by the Origin guard.
- H4: the subagent activity panel injected raw AI tool names/inputs into
innerHTML (executed under CSP 'unsafe-inline'). All sinks now escapeHtml'd.
- H5: the WebSocket upgrade had no Origin/Host check (CSWSH) -> now validated.
A missing Origin is allowed so curl/CLI and Claude Code hooks keep working;
custom reverse-proxy domains need CODEMAN_ALLOWED_HOSTS=host,.suffix.
Deferred: H2 (self-update tag signing, needs signing infra) and CSP
'unsafe-inline' removal (needs a nonce migration).
Tests: test/network-host-guard.test.ts (19), test/routes/ws-routes.test.ts
updated. Report: docs/reports/security-review-2026-06-09.md
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Close WebSocket when session exits (exit event listener) to prevent
orphaned listeners and stale writes to dead PTY
- Add readyState guard in onTerminal to stop buffering after socket closes
- Simplify heartbeat: remove redundant alive flag, use pongTimeout only
- Add exponential backoff reconnection on unexpected WS close (skip for
server rejections 4004/4008/4009)
- Clear CJK textarea on session switch to prevent wrong-session input
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
WebSocket route: add socket error handler to prevent process crashes, enforce
per-session connection limit (max 5), track/decrement counts on close.
CJK input: add destroy() method with proper listener cleanup, guard against
double-init, add maxlength/aria-label to textarea, use language-neutral
placeholder, explicitly clear cjkActive on hide.
install.sh: fix update() to use $BRANCH and $REPO_URL instead of hardcoded
origin/master — fork users were silently switched back to master on update.
README: fix broken markdown table (paragraph concatenated into last cell),
add CODEMAN_NODE_VERSION to env var table.
Tests: add 8 new test cases for batch coalescing, flush threshold, unknown
message types, connection limit, heartbeat, readyState guards. Import
MAX_INPUT_LENGTH from config, add connectWs timeout, replace setTimeout
with vi.waitFor in cleanup test.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Detect stale connections that TCP keepalive won't catch for minutes,
especially through tunnels and proxies. Pings every 30s with a 10s
pong timeout — if the client doesn't respond, the socket is terminated
and all timers cleaned up.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The HTTP resize route validates via ResizeSchema (cols: 1-500, rows:
1-200, integers only). The WS handler only checked typeof === 'number',
allowing floats, negatives, and extreme values through to ptyProcess.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Replace per-keystroke HTTP POST + SSE terminal output with a single
bidirectional WebSocket connection for dramatically lower input latency.
The existing SSE+POST paths remain fully functional as fallback.
Server-side: ws-routes.ts provides /ws/sessions/:id/terminal with 8ms
micro-batching and 16KB flush threshold. Each batch is wrapped in
DEC 2026 synchronized update markers so xterm.js renders atomically —
Ink's DA capability negotiation fails through the PTY→server→WS proxy
chain, so without server-injected markers, cursor-up redraws flicker.
Frontend: _connectWs/_disconnectWs manage per-session WS lifecycle.
Input and resize use WS fast path with HTTP POST fallback. SSE terminal
events are suppressed when WS is active to prevent double rendering.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>