Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.
A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.
The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.
The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.
Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three follow-ups to the source gate, each one measured rather than reasoned.
A pane already drawing at the size the client just requested is left alone. The
replay runs at `dimsAfterLoad`, so it can only change what is on screen if the
pane was drawing at some other size; when the reported geometry already IS that
size, the second pass captures the identical frame and pays a full reload to do
it, including a visible re-flash, a dropped and reopened WebSocket and a deleted
xterm snapshot. That equality is the signature of a clamp rather than a race:
`getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` does not, so a
terminal narrower than 40 columns or shorter than 10 rows reports a pane
permanently bigger than itself and replayed on every tab switch without ever
converging. A race never produces the equality, since its premise is that the
pane was still at the size it was asked to leave. The declined-resize case does
not produce it either, so that one still costs the single capped attempt and
needs the pane-ownership question this does not touch.
The full-history re-arm is unreachable and now says so. A pass that consumed the
flag sent `full=1`, and the route answers `full=1` with `mux-full-history` or
`history`, never `mux-visible`, so the source gate already rules out every such
pass. The line stays for the invariant, but its comment no longer reads as if a
page load retries, and the suite pins that it does not.
The response no longer reports geometry for a body that carries no capture. The
full-history path writes `capturedGeometry` from the cursor query and then
returns '' for a pane holding nothing visible, which drops the source to
`history` with the geometry already recorded: a `full=1` request whose capture
reported 100x50 and returned nothing answered `source: "history"` with both
fields set. Nothing acted on it, because the client ignores geometry on any
other source, but the field said a frame had been drawn at a size when none had.
The browser stub now derives `source` from the request the way the route does,
rather than answering `full=1` with `mux-visible`, which the route cannot
produce. Each case reaches a visible-frame response the way production does, by
not being the first select of the page. Three cases pin the new behaviour and
each fails without its guard: the clamp case sees two fetches instead of one,
the scope case and the full-history case both see a replay the gate forbids, and
the width case sees one fetch instead of two.
The changeset now describes the change from 1.29.x rather than the difference
between the two commits on this branch.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Only a visible-frame capture positions its rows absolutely, so only that frame
can be damaged by a terminal of the wrong size. A `full=1` body is linear
scrollback closed by a relative cursor move, which is relative precisely so the
browser's row count need not match the pane's, and a `history` body is the byte
stream, which carries no row alignment to protect. The geometry comparison ran
on all three, so it fired most often on the one response it cannot help:
`_fullHistoryLoaded` is empty on the first select of every non-shell session per
page, and a session whose pane a desktop tab holds too tall to ever fit then
paid a second whole-scrollback capture, reset and replay on every page load and
every first tab switch.
`framePositionsRowsAbsolutely` gates both the captured-geometry comparison and
`sizeMovedUnderLoad`. A size that moved under a byte-stream or scrollback replay
is healed by xterm's own reflow plus the SIGWINCH the trailing `sendResize`
already sends.
A pane WIDER than the terminal damages the same frame a second way, so
`captureCols` is now compared rather than only logged. `formatPaneSnapshot`
paints each row out to the pane's own width, so a narrower browser wraps every
painted row, and the wrap on the last one scrolls the whole frame up by a row.
The terminal response no longer falls back to `session.ptyCols`/`ptyRows` when
the capture reported no geometry. The cursor query is what produces the absolute
addressing in the first place, so a capture that lost it returned a raw frame
that was never positioned, and a byte-history response was never positioned
either. Naming the session's own PTY size there described a frame that does not
exist and invited a repair for damage that is not present. `_ptyCols` is also
written only by `resize()` while the PTY is spawned at the size queried from
tmux, so it can be wrong on its own terms. Both fields are now absent instead,
and the `Session` getters added for that fallback go with it.
Two browser cases cover the new behaviour and each fails without its fix: a
`mux-full-history` response with both dimensions mismatched asserts one fetch
(two without the gate), and a `mux-visible` response wider than the terminal
but short enough to fit asserts two (one without the width comparison).
Corrects a claim in the comment above `capturedGeometry` in tmux-manager.ts.
Both replay paths do not address rows absolutely; the full-history one ends in a
relative move, which is the whole reason the gate is right.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A visible-frame capture repaints each row at an absolute position, counting up
to the pane's height. A terminal shorter than that clamps every address past its
own height onto its last line. The overflow rows then overwrite one another, and
the rows underneath are lost. Replaying a real 50-row capture into a 30-row
terminal rendered 28 lines of a 45-line command and drew the frame twice.
Nothing in the response said what height the frame was built for, so the client
could not detect this. A capture now reports the geometry it was really taken at
through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response
carries it as `captureCols` and `captureRows`. When the captured pane is taller
than the terminal, or the size that produced the capture did not survive the
load, `selectSession` replays once at the size that stuck. `resizeRetry` caps
that at one attempt, so two competing fits cannot trade replays forever.
The retry re-arms the full-history flag only when the pass that ran had consumed
it. A tab switch takes the bounded tail, so its retry takes the tail too:
clearing the flag unconditionally would upgrade that switch into a fresh
scrollback capture the user never asked for, which the route's own comments put
at tens of megabytes.
What this repairs is a capture that won a race against the resize meant to
precede it. It does not repair a capture whose pane was too tall because
`Session.resize` declined the resize outright, which it does for a small
viewport while a desktop viewport's size claim is live. The retry re-sends the
same declined resize and captures the same pane, and `resizeRetry` then stops
it. Repairing that means changing who owns the pane size, which is a policy
question this does not touch. The reported geometry still helps there, because
the client can see the mismatch at all rather than being blind to it.
Follows #395, #396 and #397, which fixed the other ways the replayed frame and
the terminal could disagree.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review follow-up on the wake-on-LAN PR (five findings, all of them about the
state the feature keeps and the budgets it inherits):
- Wake state is dropped by `WebServer.cleanupSession` instead of the two delete
routes, so it now goes with the session on EVERY cleanup path (cron, admin,
scheduled-run teardown, error paths) instead of surviving with up to 4 KB of
the user's buffered keystrokes. `registerSessionRoutes` returns the registry
so the server can own its lifetime without the wake-capable code living in
`server.ts`; the wiring guard is updated to allow that and gains a second
assertion that `server.ts` calls nothing but `drop`/`stop` on it.
- `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets
a wake-state entry — the input gate runs on every keystroke, so that entry
used to be allocated for every session the user types in.
- An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being
head-trimmed and then written as a fragment: one paste is one `input` value
and was never typed character by character, so its tail is a partial command
the user never sent. The drop is logged.
- The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s)
like the create/attach paths, instead of inheriting the 90 s session default
that the dashboard's reverse proxy cuts off at 60 s.
- `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable
sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart
during a wake no longer waits the poll out.
- The banner/toast wording keys off a new `queuedInput` flag on the two SSE
events, which is true only when the server actually holds bytes: browser
keystrokes travel over the WebSocket, which never passes through the
registry, so the wake BUTTON must not promise queued input. The failed-wake
path also stops pattern-matching the error message (it re-asks the
reachability route) and the WoL dialog says "admin-only" instead of "host not
found" for a non-admin in multi-user mode.
Pressing Run on a remote case whose host was asleep failed with
`could not verify tmux on remote host 192.168.50.137: …` — an ssh error that
blames tmux for a machine that is merely suspended. The only wake paths were
typed input on an established session and the banner's Wake button, so OPENING a
session (the moment the user actually decides to use that host) had none.
`RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness
machinery for a host that has no session yet, and is wired into the two
user-initiated create paths: `POST /api/quick-start` for a remote case (before
the tmux prereq probe, which is what surfaced the misleading error) and
`POST /api/sessions` with `attachRemoteSession`. A host without a wake target is
not even probed, so its behavior and latency are byte-identical. The wake is
blocking — the caller gets the session or an error — but bounded by
REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default,
because the dashboard sits behind a reverse proxy whose default
`proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while
the session was still being created. The budget has to cover the whole request
(40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which
is why it is 40 s and not 45. A timeout now says the host did not come back, and
an unreachable host without a wake target says so instead of pointing at tmux.
The wiring is deliberately in the HTTP ROUTE, never in the shared session
service: `cron-service.ts` builds sessions there with nobody waiting on the
answer, and a wake on that path would power the host on for every schedule —
the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted
(importers of `remote-wake`, and `ensureHostAwake` having exactly one caller
file), so a future caller has to come through the guard test. A rejection from
the wake IO is caught too: a broken target must fail the wake, not the route.
`remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the
session-less case, where "input is queued" would be untrue; the toast then reads
"the session starts when it is back".
Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the
route behavior is covered by new tests in session-routes.test.ts with an injected
registry, so no test opens a real socket or ssh.
Review of the previous commit found the guard inverted: the three skips keyed
on `?full=1`, which is only what the client asked for. When the capture comes
back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at
all — the reply falls back to the byte history, which IS a stream of
successive frames and still needs stripping. Gating on the request returned it
whole: measured at 82KB against 4KB for the same buffer without `full=1`. A
direct-PTY session takes that path on every first selection, not only during
an outage. The skips now key on `isFullCapture`, meaning a capture arrived.
Three further defects the same review surfaced, all on this path:
Keeping the trailing rows is only sound when a cursor move follows to count
back up from them. On the two branches where the cursor query fails there is
no move, so the caret was left at the bottom of the pane — worse than before.
The cursor is now read first and settles both decisions together.
The move is relative rather than absolute. `CUP` numbers rows from the top of
the browser's screen, so it is only right while the browser's row count equals
the pane's, and `resizeWindow` does not wait for tmux, so a capture can be
taken before a requested resize applies. Measured against real tmux with a
browser four rows shorter than the pane: the absolute move lands on a blank
row, the relative one lands on the caret's row.
An all-blank pane no longer reads as content. Retaining trailing rows and
appending a move made it non-empty, and the caller treats non-empty as "replay
this", so a blank screen would have replaced real history — the downgrade
`_replayWouldShrinkBuffer` refuses, arriving from the server side where that
guard cannot see it.
The documentation claimed one line per screen row. `-J` joins a hard-wrapped
row into its logical line, so that is false whenever any row wrapped: measured
at 10 lines for a 12-row pane. Both entries now say what actually holds, and
the stale "NOT repositioned" contract in the mux interface is updated too.
Tests: the byte-history fallback is stripped, an empty capture leaves history
intact, and the extracted helpers are unit-tested directly rather than through
source-text matching. The slice window in the capture test is bounded at the
next method, having overrun into its neighbours.
Switching to a session left the caret one row below the composer's input
line, on the box border, and every cursor-relative update the CLI sent
afterwards was measured from the wrong row. Any fresh output repaired it,
because the CLI then repainted the whole frame.
Two things were wrong with the full-history replay, and they compound.
The capture never restored the cursor. The visible-frame path ends with an
absolute cursor move back to the pane's position; the linear path returned
its text and left the caret wherever the last character landed, which for an
agent CLI is the bottom-most row carrying text — the status line.
The rows it addressed did not line up with the pane's rows either. Four
transforms ran over the capture and each can delete a line: the trailing
blank rows were stripped, redraw-bloat stripping ran, the trim that cuts
everything above the Claude banner ran, and leading whitespace was removed.
All four are right for a byte stream of successive frames. A capture is the
rendered pane, one line per screen row, so each deletion shifted the frame
out from under the restored cursor.
The full-history path now appends the pane's own cursor position and keeps
every row, so row N of the reply is row N of the pane. The visible-frame and
tail paths are untouched.
Restoring the cursor is what makes row alignment load-bearing here, and
neither CLAUDE.md nor the architecture invariants said so — which is how
four line-deleting transforms accumulated on the path. Both now record it.
Verified against a live 315x59 pane: the reply carries 59 rows, its row 55
is the composer's input line matching tmux, and it ends with the cursor move
that lands there.
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.
DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.
That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).
Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
Closes#259, closes#258. Both bottom out in the same gap: nothing tracked
whether the user was following live output or reading history.
#259 — the keyboard path forced the terminal to the bottom unconditionally
(onKeyboardShow/onKeyboardHide passed scrollToBottom:true, applied with no
check), so opening the keyboard while scrolled up yanked the user down. The
settle cycle now captures intent on its FIRST event, before any fit() has
reflowed the buffer, and returns to that anchor when the user was reading.
A later capture would read an already-moved viewportY, which is why the
capture point matters. The param is renamed restoreScroll to match.
Separately, flushPendingWrites gated viewport preservation on
_hasRecentUserScrollUp(), a 1500ms decay window, so a user who scrolled up and
then actually READ for longer lost protection mid-read. Being scrolled up IS
the intent however long ago it was expressed, so it now keys off position.
The recency window stays as a race guard on the sticky scroll-to-bottom.
The full-history repull already held the user's place and is unchanged.
#258 — truncation was reported by a grey line written INTO the terminal
("earlier output truncated"), which scrolls away with the output it describes,
cannot be acted on, and said the same thing whether the rest was one click away
or gone forever. The server set one `truncated` boolean at two sites meaning
opposite things, and the client discarded fullSize and source entirely.
The route now reports truncationReason ('tail' = intentional partial replay,
the rest is retained; 'capped' = the byte ceiling dropped it) plus
retainedBytes, and 'capped' is not downgraded by a later tail cut. The client
renders a dismissible banner outside terminal output with three honest states:
recoverable (offers Load full history), at-ceiling, and exhausted. The Load
button forces past the scroll cooldown but NOT past _replayWouldShrinkBuffer,
which still refuses a downgrade for repaint-mode panes.
The banner is an overlay, not a flex child: FitAddon derives rows/cols from the
terminal parent's computed height, so occupying real layout space would SIGWINCH
the CLI on every truncation-state change.
Verified in a real browser on the 7 skins: banner text and button clear 4.5:1
contrast on all of them, and terminal height is byte-identical with the banner
shown. The first cut used --bg-elevated and --accent-muted, which do not exist,
so light skins rendered a hardcoded dark bar under dark text; it now uses only
tokens every skin redefines.
test/terminal-scroll-intent.test.ts lives outside test/mobile/ deliberately —
that suite is excluded from test:ci, so a guard placed there is invisible to CI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.
Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.
Test fails against the pre-fix line and passes after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).
Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:
1. extractTranscriptEntrypoint() scanned any line containing the
substring "entrypoint", not specifically the first "type":"user"/
"type":"assistant" message line (unlike its sibling
extractFirstUserPrompt, which does scope to type). A transcript
that started under an older Claude Code version (no entrypoint
field) and got resumed under a newer one mid-conversation could
pick up the field from a much later message than the true first
one, misattributing the session's origin. Scoped it to match.
2. Added a regression test proving the tail-read fallback still
engages correctly when bookkeeping accumulation exceeds even the
new 128KB head window, not just the 16KB it previously blanked at.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).
Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.
Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up to #202. The dotdir decode landed there was reachable only when
nothing else matched first, and in the greedy half it was not reachable at all.
decodeProjectKey() splits the project key on '-', so the '/.' that the encoder
collapses leaves an EMPTY segment behind. Both loops offered that empty string
as a candidate directory name, and isDir(current + '/' + '') stats current + '/',
which always succeeds. So the empty segment matched unconditionally:
- backtracking half: ~/.sib resolved to "/home/x//sib" whenever a non-dot
sibling ~/sib existed (wrong directory, and a doubled slash that then fails
every string comparison against session.workingDir). Without a sibling it
only backtracked out by luck.
- greedy half: that loop is shortest-match-first, so the empty candidate
matched on the FIRST iteration and set matched=true, leaving #202's dotdir
branch permanently dead there.
An empty string is never a real path component, so skip it in both loops. The
unmatched tail then has to handle the empty segment too, or it would append a
bare '/' and re-introduce the '//' path it just stopped producing; it now emits
the dotdir guess instead, which is what the encoder implies.
Regression test asserts both halves: the dotdir wins over the non-dot sibling,
and the result never contains '//'. Verified it fails on #202 as merged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
decodeProjectKey() couldn't recover a dotdir path (e.g. ~/.codeman) from
Claude Code's encoded project-key names: the encoder maps both '/' and
'.' to '-', so the decoder's candidate joins never matched a hidden
directory on disk. It silently fell through to bare $HOME instead,
which corrupted workingDir for any resumed session under a dotdir case
(observed on ~/.codeman itself: history rows and state.json recorded
"/home/timkjr" instead of "/home/timkjr/.codeman").
Add a dot-prefixed candidate to both the backtracking decoder and its
greedy fallback so a leading empty split segment (the signature of a
literal '.' in the original path) is retried as a hidden directory.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Includes review fixes: explicit ?full=1 trigger wired from initial page load, capture maxBuffer sized from config with -S line bound, capture returned alone (no byte-buffer duplication), early byte-cap before normalization.
- UI: add the missing data-tab="case-remote" tab button; dispatch it through
submitCaseModal()/switchCaseModalTab() to linkRemoteCase() (was dead code).
- Restore: restoreMuxSessions() now passes remote (muxSession.remote ??
savedState.remote) into the Session constructor, so remote metadata round-trips
on restart instead of reattaching from a local cwd / respawning LOCAL / being
erased from state.json. Recovery tests added.
- Run flows: runClaude()/runShell() route remote cases through /api/quick-start
(POST /api/sessions stat-validates workingDir locally); run*() skip the
/api/*/status pre-check and omit inert config/env for remote cases.
- Quick-start: resolve the remote case BEFORE the local CLI availability gates and
skip isCodex/Gemini/OpenCodeAvailable() when remote; REJECT
envOverrides/effort/codex/gemini/openCode config for remote (they don't cross
ssh) instead of silently dropping them.
- Injection: reject $, backtick, $( in remotePath + identityFile at the schema
layer (they survive shellescape into the bash -c launch double-quote layer).
Regression tests for $(...) and backtick payloads added.
- Remote socket/name: launch on a DEDICATED -L codeman-remote socket under a
codeman-ssh-<id> name that fails a remote Codeman's SAFE_MUX_NAME_PATTERN, so a
remote instance can't adopt the session; scope tmux set-options per-session
(never -g) so they don't mutate other sessions.
- Kill: best-effort ssh 'tmux -L codeman-remote kill-session' on remote session
kill (fire-and-forget, never blocks/throws the local kill) so the remote agent
isn't orphaned forever.
- Probe: wire checkRemoteTmuxAvailable() into POST /api/quick-start (structured
OPERATION_FAILED) and as courtesy validation in remote-link; add a default
-o ConnectTimeout=10 to buildSshConnectionArgs (overridable via extraSshOptions).
- Command default: remote claude default is now
'exec claude --dangerously-skip-permissions' (per-host override stays the escape
hatch), mirroring local non-interactive semantics.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Replace the 'missing ?tail means reload' overload with an explicit ?full=1
query param: the frontend's first buffer load after a page load (selectSession)
now requests full=1, tab switches keep ?tail=, and the legacy no-param callers
(response-viewer fallback, clearTerminal refresh) keep the cheap visible-frame
path — the COD-47 feature was previously unreachable from a real reload.
- When the full-history capture succeeds, return it ALONE instead of prepending
the byte buffer + \x1b[H\x1b[2J: the capture is the rendered superset of the
byte history, and ED2 clears only the viewport so the concat replayed the whole
conversation twice in xterm scrollback. The history+clear+frame concat stays
for the visible-frame/tab-switch path.
- Pass an explicit execSync maxBuffer for the full-history capture (configured
terminalBufferMaxBytes + slack) — the 1MB Node default ENOBUFS-killed exactly
the multi-MB captures the feature exists for; log ENOBUFS concisely instead of
dumping the truncated stdout.
- Bound the capture itself via -S -<N> derived from the configured tmux
history limit (was unbounded -S -), and add -J so lines hard-wrapped at the
capture-time pane width reflow in the browser xterm.
- Cap the concatenated buffer to terminalBufferMaxBytes EARLY (before the
regex normalization passes) so multi-MB captures don't stall the event loop
normalizing bytes that get sliced away.
- Tests: route tests updated for ?full=1 semantics (capture-alone response,
config-forwarded capture bounds, byte-history fallback, no-param requests
stay on the visible-frame path); source-scan tests cover the bounded -J -S -<N>
flags and explicit maxBuffer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Breaker reset is now explicit-only: POST /api/sessions/:id/interactive no
longer unconditionally resets the PTY-exit breaker (that endpoint IS the
frontend's automatic re-attach path, so the breaker could never trip on the
COD-115 crash loop and any tab click silently re-armed it). The route accepts
a schema-validated optional body flag {clearBreaker:true}
(InteractiveStartSchema) and resets only when it is sent.
- Frontend restart control: app.js selectSession keeps the bare auto-attach
(no body, never clears); when the selected session has respawnBlocked it asks
for explicit user confirmation and only then re-POSTs with clearBreaker:true.
respawnBlocked is surfaced via SessionState/toState() (runtime-only, not
restored on boot so recovery can re-attach).
- Trip observability: WebServer.setupSessionListeners() is now idempotent
(skips while refs are attached) and the re-attach routes (/interactive,
/interactive-respawn, /shell) re-run it, restoring the wiring that the exit
handler detaches on every PTY exit — without this the 5th-exit trip had
guaranteed zero listeners (no SSE, no push, no persist, no run-summary).
- Push notification: added SessionRespawnBreakerTripped to PUSH_EVENT_MAP
('Session crash loop stopped', urgency critical) with an exit-count body
branch; previously sendPushNotifications silently no-oped.
- Minor: buildMuxAttachEnv() truecolor param is now actually passed
(codex/gemini, mirrors buildEnvExports); buildClaudeEnv() uses delete for
COLORTERM/CLAUDECODE (same node-pty "KEY=undefined" quirk as COD-115).
- Tests: route tests assert auto-reattach does NOT reset, clearBreaker resets,
invalid flag rejected, and listener re-wiring on /interactive + /shell;
real-wiring lifecycle tests (createSessionListeners/attach/detach) prove the
exit-detach gap and that re-setup keeps the 5th-exit trip observable;
PUSH_EVENT_MAP regression guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Event-loop blockage: HEIC decode/encode (CPU-synchronous libheif WASM +
jpeg-js) now runs in a per-conversion worker_threads Worker
(src/web/heic-jpeg-worker.ts, spawned by heic-jpeg-converter.ts) with
resourceLimits and a 30s hard timeout that terminates the worker —
verified end-to-end under tsx and against compiled dist/ output with a
real iPhone HEIC (event-loop max stall 52ms during conversion).
- No server-side concurrency cap: conversions now acquire a slot from the
existing global runWithConversionLimit() pool (document-conversion-limiter),
bounding peak decode memory/CPU across simultaneous uploads.
- Decompression bomb: header-declared dimensions are read via heic-decode's
allocation-free `.all` path and rejected above 64MP BEFORE decode() can
allocate width*height*4 bytes (a <300-byte crafted file can declare
30000x30000 = 3.6GB). Regression-tested with a crafted ISOBMFF fixture
against the real heic-decode WASM (test/heic-jpeg-core.test.ts).
- Mislabeled HEIC (documented Android/MIUI case): conversion now routes on
ftyp magic-byte sniff of the raw buffer regardless of declared
ext/Content-Type, so a HEIF uploaded as image/jpeg converts instead of
415ing; the magic-mismatch 415 only fires for genuinely unrecognized bytes.
- Brand allowlist narrowed to what heic-decode's isHeic() accepts
(heim/heis/hevm/hevs dropped — they could only ever fail conversion).
- Converted-output size: the JPEG result is checked against
MAX_PASTE_IMAGE_BYTES (jpeg-js can inflate a within-limit HEIC past the cap).
- Deps: heic-convert replaced with its underlying heic-decode + jpeg-js
(the wrapper could not expose the pre-decode dimension check); lockfile
synced, drops pngjs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
When a browser pastes an HEIC file without normalising it first, the
paste-image route now converts it to JPEG server-side via heic-convert
before writing to .claude-images/. Magic-byte validation confirms the
output is valid JPEG. Adds type declarations for the heic-convert package.
Co-authored-by: Saqeb Akhter <saqeb.akhter@gmail.com>
A full page reload (GET /api/sessions/:id/terminal with no ?tail=) now captures
the ENTIRE tmux scrollback via capture-pane -p -e -S -, so users get back history
that scrolled off Codeman's byte buffer. Tab switches (?tail=N) keep the fast
visible-frame capture.
- tmux-manager capturePaneBuffer/captureActivePaneBuffer take { fullHistory }:
full-history returns raw linear scrollback (skips the single-screen
formatPaneSnapshot repaint, which would clip multi-screen history).
- /terminal selects full-history on full reload, visible on tail; caps the
payload at the configured terminalBufferMaxBytes (keeps most-recent bytes,
line-aligned) and returns source/fullSize/truncated metadata.
Verified: tsc 0, tmux-capture-full-history 5/5, session-routes 68/68.
Caveat: lines tmux already evicted past its history-limit can't be recovered.
- src/remote-hosts.ts: add missing execAsync = promisify(exec) that was
implied by intermediate commits not in the cherry-pick set
- src/web/routes/session-routes.ts: add getDataDir import and
readRemoteCases/readRemoteHosts/toSessionRemote for remote case support
in quick-start; narrow casePath string|null via resolvedCasePath cast
- test/routes/session-routes.test.ts: add vi.hoisted remoteStore mock for
remote-hosts.js; fix 'creates session from remote case' test to use
/api/quick-start (remote cases are not supported on /api/sessions)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
A "sent" prompt could vanish with no trace on a flaky connection (e.g. a train):
with local echo on, Enter cleared the overlay then sent over the WebSocket
fire-and-forget. On a half-open socket (readyState===OPEN, dead TCP) ws.send()
doesn't throw, so the frame was silently discarded, nothing was enqueued, and
navigator.onLine stayed true — the prompt was lost and never resent.
Replace the best-effort offline queue with a durable, acknowledged delivery layer:
- Client (app.js): every input frame is recorded with a stable clientId +
monotonic per-session seq and persisted to localStorage BEFORE delivery, and
only dropped on a server ACK. Delivered over WS (acked via {t:'ia',seq}) or,
when the socket is down, POST in seq order (HTTP 2xx = ACK). A 2s sweep
force-reconnects a WS whose oldest frame is unacked past 4s (half-open sockets
never recover on their own); on reconnect/reload all pending frames re-deliver.
Survives reconnects AND page reloads. Connection indicator shows pending count.
- Server: Session.shouldApplyInput(clientId, seq) applies each frame exactly once
(bounded MRU map); ws-routes + POST /input dedup a redelivered seq but still ACK
it (200 / {t:'ia'}), so an at-least-once resend can never type the prompt twice.
Untagged input (curl/legacy) applies unconditionally — no behavior change.
- terminal-ui.js sendInput() (voice / keyboard-accessory / paste) now routes
through the same durable layer.
Tests: test/reliable-input-dedup.test.ts (exactly-once semantics on the real
Session) + POST /input dedup route tests. Design: docs/reliable-input-delivery.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Switching away from a session and back replayed only the server's byte
history. For TUI modes (codex especially) that shows just the latest
repaint — the idle banner — because the TUI drops earlier conversation
from its current frame. This restores the actual on-screen view.
Two complementary mechanisms:
- Client: load xterm's SerializeAddon and snapshot the rendered state
(viewport + scrollback + colors) per session on switch-away, restoring
it for an instant first paint on switch-back. The snapshot is only the
first paint — the canonical /terminal frame is still fetched and
reconciled (restoredSnapshot/clearedForBusy force the replay). Snapshots
are LRU-bounded in memory (<=20) and persisted to localStorage
(<=256KB each, <=10 sessions, stale-pruned) so they survive tab discard.
- Server: GET /api/sessions/:id/terminal prepends the live tmux pane
buffer (via the existing captureActivePaneBuffer) ahead of the byte
history, cleared between, so replay reflects the current frame.
Also fix formatPaneSnapshot dropping the rightmost column of every
captured row: it painted to cols - 1 out of caution about last-column
autowrap, but every row is followed by an absolute cursor-position CSI
that cancels xterm's pending-wrap, so painting the full width is safe.
The SerializeAddon is built from @xterm/addon-serialize (new dependency)
into the vendor bundle by postinstall.js (dev) and build.mjs (prod),
matching how the other xterm addons are vendored.
Auto-resume on usage limit ("token pause" control, opt-in checkbox at the
top of the Respawn tab, off by default):
- usage-limit-patterns.ts (new, pure): detects all Claude Code limit
messages (1.0.x-2.1.x eras incl. "5-hour limit reached - resets 8pm",
"You've hit your limit - resets 1:40pm (TZ)", weekly date forms, raw
"usage limit reached|<epoch>") and parses the reset time. Conservative:
no parseable future reset time, no action.
- SessionAutoOps: arms a timer at reset+2min, sends Esc (dismisses the
rate-limit dialog) + "continue"; dedups footer redraws, retries every
5min on stale times, cancels when Claude starts working, persists and
re-arms across Codeman restarts (SessionState.autoResumeEnabled/At).
- Respawn guard: cycles are blocked while limit-paused so /clear cannot
wipe the paused conversation (respawnBlocked reason 'usage_limit').
- POST /api/sessions/:id/auto-resume; SSE session:limitPauseScheduled/
limitResume/limitResumeCancelled; toasts + status line in the modal.
- Respawn tab tidied: single-row prompt fields, merged behavior row.
Mobile fixes (0.9.8 regressions, user-reported):
- Resize arbitration is now activity-based: a desktop sizing claim only
blocks phone resizes while the desktop typed within 90s
(Session.DESKTOP_CLAIM_IDLE_MS). Idle desktop -> phone takes the pane;
next desktop keystroke re-asserts the desktop layout server-side
(noteDesktopActivity via ws-routes input). Phones re-send dims every
30s (visible tab only, skipped while the keyboard is open) so attaching
under a hot claim self-corrects. Fixes the desktop-width-stream-in-
narrow-xterm soup (mid-word wraps, tmux dot fill, Ink overdraw).
- Cross-device reflows (takeover/re-assert) emit a debounced needsRefresh
so all clients reload the buffer instead of stacking ghost Ink frames.
- Keyboard accessory/toolbar lift restored: measure keyboardOffset
against window.innerHeight (layout viewport), not the shrunken .app -
on iOS the offset computed to 0, leaving both bars hidden behind the
OS keyboard with a dead gap above.
- Removed the mobile header utility ("three dots") toggle entirely;
the headerRight tray stays collapsed on small viewports.
Tests: usage-limit-patterns (36), session-auto-resume (21), resize
arbitration (+6), session routes (+4), respawn guard (+2); MockSession
auto-resume/sizing stubs; mobile tabs test updated for toggle removal.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Mobile-focused fixes for the web UI: keyboard-accessory layout and
overlap, native input visibility above the keyboard, CJK input handling,
terminal touch scrolling, tab-menu tap targets, mic-recording glow
containment, and mobile resize/keyboard-state handling on tab switch,
plus mobile visual-regression test coverage and snapshots.
Co-Authored-By: Saqeb Akhter <saqeb.akhter@gmail.com>
Point 1 of the v1.0 lock-in: commit to a stable HTTP API (the cleanest, fullest form).
Core (centralized):
- Every JSON /api response now uses ONE envelope via a Fastify preSerialization hook (src/web/server.ts): success -> { success:true, data:<payload> }; error -> { success:false, error, errorCode } with a conventional HTTP status. Non-JSON routes (file-raw, tail-file SSE, download, screenshots, /q redirect, WS) are skipped.
- Error-code -> HTTP status is a single source of truth (httpStatusForErrorCode in src/types/api.ts): 400/401/404/409/422/429/500. Expanded ApiErrorCode (added UNAUTHORIZED, CONFLICT, RATE_LIMITED). Errors are no longer HTTP 200.
- Versioned alias: /api/v1/* rewrites to /api/* (rewriteApiV1Url), so external clients pin to a stable surface while the bundled UI keeps using /api/*.
- Handlers stripped of manual 'success:true' (50 across 14 route files) so they return bare payloads the hook wraps uniformly; fixed the mux DELETE {success:<bool>} envelope collision (-> {killed}).
Frontend (48 call sites across 10 files):
- _apiJson() auto-unwraps { success:true, data } -> data (null on error), so most bare-shape readers are transparent. Raw-fetch sites relocate payload reads under .data; success/res.ok/error checks unchanged.
Docs: new docs/api-reference.md (envelope, status table, error codes, /api/v1, SSE); versioning-policy.md flipped — the HTTP/SSE API is now part of the stable, SemVer-covered surface.
Verification: full unit/route suite green (2680 passed) incl. ~166 updated assertions across 24 test files; typecheck/lint/format/frontend-syntax clean; a headless-chromium smoke loaded the migrated UI and drove the panels with 0 console/page errors; /api/status and /api/v1/status confirmed returning the uniform envelope live.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(test): share route error handler with test harness + fix stale assertions
The route test harness built a bare Fastify instance without the production
global error handler (server.ts), so structured errors thrown by route helpers
(findSessionOrFail → 404, parseBody → 400) fell through to Fastify's default
handler — yielding a `{statusCode,error,message}` body instead of the
`{success:false,...}` shape, and the tests asserted the old implicit-200
behavior. 51 route tests across 7 files were red.
- Extract the handler into src/web/route-error-handler.ts; server.ts and the
test harness now install the identical handler (single source of truth).
- Correct stale assertions across route test files: throw-based error paths
now assert 404 (unknown session) / 400 (invalid body); genuine in-handler
`return createErrorResponse(...)` paths (200 + success:false) left untouched.
- Reformat a few test files prettier flagged (pre-existing non-compliance).
Route suite: 307/307 passing (was 256/307). No production behavior change.
* test(respawn): mock child_process so AI checker never spawns real processes
respawn-controller.test.ts drives the AI idle checker (ai-checker-base), whose
runCheck() spawns a real `tmux new-session` running `claude -p`. The AI-enabled
tests only assert the ai_checking state transition (then cancel/stop), so the
spawn produced stray real tmux sessions and claude processes on every run — the
reason `npm test` (full suite) was unsafe to run inside a managed session.
Mock node:child_process here (mirroring ai-idle-checker.test.ts), spreading the
real module so `exec` stays intact for transitively-imported modules
(tmux-manager calls promisify(exec) at load). With this, the full non-mobile
suite runs without spawning any real tmux/claude.
---------
Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
* fix: isolate codeman tmux sessions
* fix(tmux): unify all sessions onto a single dedicated socket
Remove the per-session `tmuxSocket` field that recorded which tmux server
each session lived on (default vs the `codeman` socket). That field was a
persisted cache of physical reality and could drift — causing live sessions
to be wrongly marked dead ("tab shows no session found") and spawning
duplicate "Restored:" tabs.
All Codeman sessions now live on one process-wide socket (`tmux -L codeman`,
overridable via CODEMAN_TMUX_SOCKET), exposed via TmuxManager.muxSocket on
the TerminalMultiplexer interface. reconcileSessions() collapses from a
multi-socket scan (locate / re-pin / cross-socket dedup) to a single
`list-panes` query. loadSessions() strips the obsolete field from on-disk
records so it stops being written back.
Also fix two sibling bare-`tmux` call sites the unification would otherwise
leave broken (same #80 regression class — bare tmux hits the user's default
server and never finds a session on the codeman socket):
- session.ts queryTmuxWindowSize(): add `-L <socket>` (was silently falling
back to 120x40 on re-attach, losing scrollback)
- session-routes.ts send-key (Shift+Enter / Ctrl+Enter newline): route
through ctx.mux.muxSocket
SSH chooser scripts (tmux-manager.sh, tmux-chooser.sh) route every tmux call
through `tmux -L $CODEMAN_TMUX_SOCKET`, matching the TS default.
---------
Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
- Filter empty sessions from history API (check for conversation content)
- Add --resume fallback to new session if resume fails (prevents dead panes)
- Pass resumeSessionId through respawnPane for dead pane recovery
- Persist respawn presets and runMode to server settings (cross-device sync)
- Fix mobile touch handling for Recent Sessions dropdown (DOM API + touch CSS)
Consolidate duplicated MockSession/MockStateStore into test/mocks/,
migrate respawn tests to shared mocks, and add 58 route tests for
session, system, and respawn endpoints using Fastify app.inject().
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>