`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.
dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:
1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
(one-shot and streaming alike) stops at the first frame end: a real
56-line transcript decoded as 1 line / 158 bytes -- the session header
alone, i.e. a silent truncation that reads as "nothing said yet"
forever. `zstdFrameRanges()` walks frame and block headers to find
exact boundaries; splitting on the 4-byte magic would corrupt
everything after a magic sequence occurring inside compressed data.
zstd is resolved at RUNTIME because it landed in Node 22.15 while the
project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
surfaced as `Turn error: …` (and a non-error early stop as
`Turn ended: …`) rather than as an empty string, which an agent reads
as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
the streamed deltas fill in only for a step that never finalized, so a
partial answer is readable mid-turn and never doubled. "Finalized" is
tracked as a set of steps rather than as non-empty text, because a
step whose whole reply was reasoning strips to '' at the `</think>`
boundary and would otherwise resurrect the raw deltas in its place.
Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.
An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
Three review findings on the DeepSeek Harness mode, plus one the third exposed.
1. The multi-user clamp was bypassable by a sibling field on the same request.
clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
_configureDeepSeek(), so a non-granted owner sending
envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
instance: a session created with permissionMode "read-only" and that override
ran with DSH_PERMISSION_MODE=danger-full-access in its pane.
Every other CLI's bypass is a command-line flag reachable only through the
per-CLI config, which is why the config clamp alone is the whole gate for
them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
that list because it aims the launcher at a profile tree whose plugin code
runs at boot, before any approval row can apply. Verified end to end in real
multi-user mode: a non-granted user sending both now gets workspace-write and
no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.
2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
signals only the direct child, and a plugin install fans out into
package-manager children that keep the inherited stdio pipes open, so `close`
never fires and the held-open request leaks with no route-level deadline.
Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
6s and both fan-out children were alive. Now detached: true plus negative-pid
SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
last-resort reap for a grandchild that escaped the group. Same probe after the
change: close fires, direct child and both grandchildren dead.
3. hooksAvailableForMode() promised more than a dsh session can deliver.
deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
triple is the only reason a dsh session posts hook events, so `until=stop` was
accepted and then blocked for the caller's whole timeout: the exact
infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
takes HookCapabilityOptions and every call site passes sessionHookOptions(),
with the deepseek arm reading `!== false` so a forgotten one degrades to the
old behaviour. The refusal names the setting rather than saying "no Claude
Code hooks", which would send the caller hunting a bug that is really a
setting they chose. Profile conformance stays unknowable at request time and
is documented as such. The stale "True for `claude` and nothing else" docblock
is corrected.
4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
capture read Claude's own transcript, and adding deepseek silently widened
both to a mode that has none. They compare mode === 'claude' directly now, and
a static check pins them there.
Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.
DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:
1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
$DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
`base` -- the interactive terminal front door is always a third-party
plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
button gates on the latter, because reporting only the binary would spawn a
pane that dies on arrival. When the binary is present but no profile is,
the run menu offers to install one (POST /api/deepseek/install-profile).
2. The permission switch is an env var, not a flag. The harness has no
command-line permission option; its sandbox/approval rows read
DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
own workspace-write, which still asks, so the multi-user clamp is the
only-if-sent branch and clamps to workspace-write, never read-only.
3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
earned that. The terminal front door reports idle/working/blocked to a
supervising process over a generic env-gated contract; a generated shim
(deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
report to /api/hook-event as stop / agent_working / permission_prompt. So a
dsh session gets definitive respawn triggers, real wait-endpoint signals and
real Approvals Inbox items instead of output-stabilization guesswork.
`agent_working` is new (157th SSE constant) and joins
APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
alert at once.
The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.
Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.
Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.
Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:
$ [<0;88;20Mecho hello
bash: 0: No such file or directory
The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.
What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.
Details that are easy to get wrong:
* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
without turning tracking on is not asking about clicks, and counting
those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
broadcastSessionStateDebounced: the flag flips when a dialog opens, and
the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
the CLI re-emits its DECSET, which tmux does at client attach.
Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Root cause of the reviewer's mass-bump measurement (17 of 17 sessions with an
identical lastActivityAt): every restart restamps all sessions in the
constructor loop, and the boot auto-attach's repaint re-bumps the rest within
the same second. A 12-minute steady-state sample shows NO ambient mass bump,
so restarts are the whole story, and Codeman restarts on every deploy.
The stamp now has a display twin: recovery threads the previous run's
lastActivityAt from state.json into the wire-visible stamp (getter + toState),
and a 15s settle window keeps the attach repaint from overwriting it. Real
actions (input, task assignment, respawn) always write through. The private
stamp keeps its boot-anchored semantics untouched, because the idle
confirmation reads it as how long the pane has been quiet.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Second review pass on #282, the three items left open after f4dcfbe.
1. `codeman doctor` and the run mode disagreed about pi. The registry entry
accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
`--version` output, so the Dependencies panel could report an installed Pi CLI
on a box where Run Pi stays hidden, which reads as a broken mode rather than a
missing install. Both sides now share one exported PI_VERSION_REGEX, and
PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
shape check is reported MISSING instead of installed-with-unknown-version.
Only pi sets it; every other tool keeps its current behaviour.
2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
strips the alt-screen toggles anyway. What exclusion actually preserves is
`\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
main screen and is mouse-aware). Comment and changeset now say that, and state
the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
shell session.
3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
agents a backend does not exist and understating class-wide caveats by one
mode. All updated, plus stale session.ts line references refreshed.
Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.
Pi is a different shape of CLI from the other four, and three decisions
follow from that:
- It has NO permission prompts and no sandbox, so there is no
--dangerously-skip-permissions analog and none was invented. The
privilege-shaped knob is the tri-state approveProjectTrust, which makes
pi load and EXECUTE repo-local .pi/extensions TypeScript and install
missing project packages. clampExternalCliBypassForOwner() therefore
puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
--no-approve even when no config was sent, because pi's own default is
a prompt the session user could answer themselves. That helper had zero
test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
context, so admitting them would widen the allowlist for every mode at
once. Auth goes through pi's /login or the server's own environment.
--api-key is deliberately never wired: it would put a provider secret on
the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
mode is runtime-switchable via /settings; that flip was measured to put
the pane into the alt screen, which the strip would have corrupted.
pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.
Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).
Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.16.6: phone overview started/idle stamps, plus fixes for the
selection-dialog keyboard lockout, the accessory bar arrows bypassing the
local-echo overlay, and recovered sessions being restamped as newly created
on every server restart.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.
Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
sources: pending approval dialog, user goals, last assistant turn tail,
recent prompts, tool activity, git workspace signals, away context,
sibling sessions, rethink state; 30 KB budget, whole-section drop from
the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
keeps only a 500-char snippet) and git signal collection (execFile,
2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
session, opus by default (readMyMindModel setting), strict JSON
contract with 1-3 suggestions (continue / verify / redirect), newline
stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
prediction in flight per session (409 CONFLICT), rethink body
{ steer, rejected }; ownership via findSessionOrFail
Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
until readMyMindEnabled is ON, desktop only (phone key is phase 3);
modal with editable suggestion + rationale and Send / Insert /
Rethink / Dismiss; suggestion text rendered via value/textContent only
and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
strings
Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds an exact-key tier (ALLOWED_ENV_KEYS) beside ALLOWED_ENV_PREFIXES in
schemas.ts, admitting CLAUDE_CONFIG_DIR so a case can run on a separate
Claude subscription (client-billed accounts). Exact match only: other
CLAUDE_* keys and near-misses like CLAUDE_CONFIG_DIR_EXTRA stay rejected,
blocked keys stay blocked. The key also survives getEnvOverridesForPersist()
(a path, not a secret; dropping it would silently switch a rebuilt session
back to the default account after a reboot).
Docs cover the transcript caveat: a relocated config dir writes transcripts
outside ~/.claude/projects, so response viewer / subagent windows /
ultracode / Read My Mind go blind for that session unless projects is
symlinked back into the shared tree.
Design and spec contributed by @jordan8037310 in #255. Closes#255.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A session on a fresh directory sat on Claude's "Quick safety check: Is
this a project you created or one you trust?" dialog until a human
pressed Enter. Reproduced on a new case, then read off the wire:
1.\x1b[C Yes,\x1b[C I\x1b[C trust\x1b[C this\x1b[C folder
tmux repaints a row by writing each word followed by a cursor-forward
escape instead of a space, and Ink colours each word separately, so
`data.includes('trust this folder')` could never match a chunk. The
spaces are not there to strip: they were never sent. The auto-accept has
been dead for every session that hit the dialog.
Match on whitespace-free, ANSI-free, lowercased text instead
(`compactScreenText`), which survives both that repaint style and the
spaced full-screen redraw.
Answering means pressing Enter into a session, so three guards bound it:
- Read the RENDERED SCREEN (capturePaneText), not the chunk. The terminal
buffer is append-only and keeps the dialog in its tail long after it
has been answered, so a retry driven off the buffer would type into a
live session. Direct-PTY sessions, which have no pane, fall back to a
short buffer tail.
- Require a trust phrase AND the dialog's own confirm affordance. One
phrase is not enough, since an agent's transcript can quote it.
- Only look during the first 90s of the pane's life, and cap it at three
attempts. Ink can drop a keystroke while it is still mounting the
widget, which is the other half of why sessions got stuck, but a
dialog that will not clear must not become an Enter loop.
Verified end to end on a fresh case: dialog answered on attempt 1, one
Enter sent in total, session went straight to the composer and answered a
prompt. Before the fix the same flow parked on the dialog indefinitely.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Every working Claude session reported `status: "idle"` about two seconds
into its turn. Measured on live workers: two sessions mid-tool-call at 13
and 17 minutes both read `idle` while their panes showed
`✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`.
Two things had drifted apart:
1. The working indicator changed. Claude animates the glyph through
`· ✢ ✳ ∗ ✻ ✽` and randomizes the gerund per turn, so neither
SPINNER_PATTERN (braille, no longer drawn) nor the keyword list
(Thinking/Writing/Reading/Running) matches a turn anymore.
2. A `❯` sighting is not the end of a turn. Claude redraws the composer
roughly once a second all the way through one, and that redraw armed
the "2s later, call it idle" timer.
Matching the new status line in the STREAM does not fix it either: tmux
ships partial repaints, so the complete line reached the PTY about once
every 20 seconds while the `❯` arrived every second.
So the decision moves off the stream:
- An unbroken run of repaints marks a turn as started. Sampled once a
second for 12s over six live sessions, the two working ones produced
output in 12/12 windows and the four idle ones in 0/12. Pure helpers in
session-activity.ts carry the thresholds.
- Idle now needs the pane to go quiet AND the screen to agree.
`_confirmIdle()` asks tmux what is rendered (new `capturePaneText()`,
one plain `capture-pane`, floored at 1.5s per session and only ever at
a transition) and re-checks every 5s while the screen still shows work.
A turn can sit silent for tens of seconds inside one tool call, so
silence alone proves nothing.
- The same screen check vetoes keystroke echo, which is a steady stream
of repaints too but is not work.
CLAUDE_WORKING_LINE_PATTERN matches the `… (elapsed)` shape rather than
the glyph, because the FINISHED line (`✻ Cooked for 2m 49s`) carries the
same glyph and would otherwise pin a session at working forever.
Claude mode only. An external CLI has no `❯`, so nothing would arm the
confirmation and such a session would latch busy.
respawn-patterns.hasWorkingPattern() had the same blind spot (its gerund
list cannot see "Actualizing"), so it takes the pattern as an extra
signal. That can only make respawn less eager, never more.
Idle now lands about 3 to 5 seconds after a turn ends instead of 2
seconds into one. Verified end to end against a live worker, sampled
against the CLI's own "esc to interrupt" footer as independent ground
truth: busy for all 25s of a turn, idle 3s after it ended.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.
- write() had two: the original description with @param and @example, then a
@returns-only block added on top, which dropped the params and examples from
hover. Merged into one. The @returns wording is also honest now — write() still
discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
and its declaration, leaving that function undocumented on hover and the doc
attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Four fixes for the scrollback reports in #205 (plus its follow-up comment).
1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
(\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
to claude/codex/gemini, so it reached the browser verbatim. In the alternate
buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
which readline receives as shell history navigation. Both reported symptoms,
one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
never forwards a pane's alt-screen toggles to its client, it repaints;
captured from a real attach, vim/less/htop emit zero.
2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
onto tmux's history, and tmux repaints the pane rectangle instead of emitting
linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
same 60 lines emitted slowly added all 60. Scrolling up at the top now
re-pulls the full tmux scrollback and holds the user's place. Verified
end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.
3. The full-scrollback replay was gated on a single "first load after page load"
flag, which whichever session auto-selected consumed, so every other tab
started with one visible frame. Now tracked per session.
4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
per notch) scrolled one line where Chrome scrolls four or five, and capped
the forwarded SGR report at one tick. Line and page deltas are now converted,
and a pure horizontal swipe no longer falls through to a phantom -1.
Analysis and measurements: docs/scrollback-issues-analysis.md
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.
When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.
- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
client redelivers.
- `Session.write()` reports whether it reached a PTY.
Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.
What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.
9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).
Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.
Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.
- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
injected via socket-scoped `tmux setenv`, never on the spawn command line, so
the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
colours. `runAntigravity()` routes remote/docker cases through
`POST /api/quick-start` and skips the local status probe for them.
Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.
Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.
Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.
Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-ups from the PR #175/#176 reviews:
- Rewake helper self-terminates on its own 6h deadline and when orphaned,
instead of relying on Claude Code to reap the poller
- Rewake marker versioned (V2) with a version-agnostic ownership prefix, so
future script updates replace older handlers instead of duplicating them;
regression test covers the V1 to V2 swap
- HOOK_TIMEOUT_MS renamed to HOOK_TIMEOUT_SECONDS = 10: the hook timeout
field is seconds (the CLI multiplies by 1000), so the curl hooks have
effectively had a ~2.8h timeout since COD-54
- Test echo PTY switches to raw mode: each input byte echoes exactly once
(tty line discipline doubled every line and buffered until Enter)
- test/setup.ts: drain in-flight console-log rpc forwards before environment
teardown (fixes the EnvironmentTeardownError that failed CI twice on the
merge commit with all 3820 tests passing), clean the temp home on process
exit (fully-skipped files leaked it), fix the Windows Playwright cache
fallback path
- test/webview-proxy.test.ts: stop naming the vitest environment directive in
prose; vitest matches it inside comments and silently ran the whole file
under the jsdom environment while the comment claimed node
- CLAUDE.md: document the temp-HOME and echo-PTY test isolation
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolutions (sse-events.ts / constants.js / app.js): unions of the docker/
multi-user event registrations from master with the session-order/pin events
from this branch.
Additions on top of the merge:
- POST /api/sessions/:id/pin now falls back to the persisted store record when
no live session exists: COD-142 deliberately preserves pinned records after
kill (and cleanupStaleSessions skips them), so without this a pinned-then-
killed session could never be unpinned. Owner-scoped in multi-user mode.
- SessionOrderUpdateSchema bounds (id <= 100 chars, <= 500 entries) so a buggy
client can't persist megabytes into state.json; empty strings still flow to
normalizeSessionOrder which drops them.
- Route tests for the persisted-record pin fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Resolutions:
- session.ts: keep the extracted _buildRespawnPaneOptions() helper (COD-108)
and add master's docker/owner fields to it
- tmux-manager.ts: docker branch first, then remote via buildRemoteSessionCommand
(now an options object threading claudeMode/allowedTools into
buildRemoteLaunchCommand, preserving the 6.3 multi-user permission downgrade)
- case-routes.ts: keep master's adminOnly helper; gate the new COD-105 discovery
endpoint admin-only in multi-user mode (hosts are machine-level infra)
- settings-ui.js: union of remoteAutoReconnect + master's header-button defaults
- session-routes.ts: union of imports; session gets remote + owner
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Threads per-user ownership through sessions, cases, cron, and the permission
policy. All scoping is a no-op in single-user mode (isMultiUserMode() guards).
Sessions
- Session.owner stamped at every create path from req.authUser / job.owner:
POST /api/sessions, /api/run, /api/quick-start, ralph start, cron launch,
plan generation. Round-trips through recovery (MuxSession.owner mirror, read
muxSession.owner ?? savedState?.owner) and the mux layer.
- findSessionOrFail(ctx, id, req) now does a NOT_FOUND owner check (never 403, so
other users' session existence is not leaked); wired at ~50 call sites.
- List endpoints filtered by owner: GET /api/sessions, /api/sessions/unified
(live+persisted+lifecycle scoped, host-wide transcripts admin-only), cron jobs.
Permission policy (section 6.3)
- resolveClaudeModeForUsername wraps getClaudeModeConfig at every spawn site so a
non-granted user is forced to --permission-mode auto (bypass -> auto), including
recovery (or a reboot would un-downgrade). buildPromptArgs now respects the
session's claudeMode, closing the one-shot (runPrompt) bypass hole.
- Shell mode and cron launchCommand require canBypassPermissions: 403 at
POST /api/sessions, /api/quick-start create, cron job create, AND cron fire time
(re-checked against the owner's current grant).
Cases
- resolveCasesDir(user): per-user ~/codeman-users/<name>/cases in multi-user, the
shared ~/codeman-cases otherwise. All case CRUD + ralph + plan + quick-start
resolve through it. resolveCasePath is owner-aware.
- GET /api/cases scoped per user (own folders; legacy linked cases admin-only;
remote/docker cases owner-filtered). RemoteCase/DockerCase gain owner, stamped
at link/quickcreate/import.
- Remote + Docker host CRUD is admin-only.
- Non-admin workingDir confinement (the linchpin): realpath must resolve inside the
user's space, enforced at POST /api/sessions and /api/run BEFORE any disk write.
Limits
- sessionCapacityState / sessionCapacityMessage centralize the global + per-user
cap (CODEMAN_MAX_SESSIONS_PER_USER, default global/2), replacing the 6 copy-pasted
MAX_CONCURRENT_SESSIONS checks.
Tests: test/ownership-scoping.test.ts (case isolation, host-CRUD gate, workingDir +
shell gates, and the scoping helpers). Deferred to phase 4: WS owner gate, SSE
fan-out filtering, file-route preview/thumbnail helper scoping, push routing.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a parallel multi-user auth branch (the single-user Basic-auth path is
left byte-identical). Off unless CODEMAN_MULTIUSER/--multiuser.
- middleware/auth.ts: mode-selecting registerAuthMiddleware. New async
multi-user hook verifies username:password against the user store (scrypt),
mints identity-carrying cookies, decorates req.authUser, enforces a per-IP
AND per-username failure bucket, and the mustChangePassword lockbox. The
hook-secret loopback bypass is now a single shared helper used by both
branches. FastifyRequest.authUser module augmentation.
- ports/auth-port.ts: AuthSessionRecord gains username/role/mustChangePassword.
- user-store.ts: verifyPassword (timing-equalized against user enumeration).
- route-helpers.ts: getAuthUser (synthetic admin fallback), canAccessOwned,
requireAdmin, revokeUserSessions; findSessionOrFail gains an optional req for
a NOT_FOUND owner check (dormant until phase 3 wires callers).
- routes/me-routes.ts: GET /api/me (synthetic admin in single-user) and
POST /api/me/password (verify current, min 8, clear mustChangePassword,
revoke other sessions).
- QR: QrTokenRecord + AuthSessionRecord carry a username; tunnel-manager
mintUserToken / consumeTokenWithIdentity / getQrSvgForCode; /q/:code binds
the cookie to the token's user (rejects identity-less tokens in multi-user);
GET /api/tunnel/qr mints a per-user token. Single-user keeps the rotating token.
- server.ts: bootstrap the initial admin from CODEMAN_USERNAME/PASSWORD on first
boot (refuse to start with no users); multi-user with >= 1 user satisfies the
non-loopback auth requirement and the tunnel-enable guard; userFailures bucket
disposal.
- types/api.ts: FORBIDDEN, PASSWORD_CHANGE_REQUIRED, USER_EXISTS, USER_NOT_FOUND,
LAST_ADMIN error codes (message + status wired).
- Session.owner field + getter/setter, SessionState.owner, MuxSession.owner,
CreateSessionOptions.owner (foundation for phase 3 ownership threading).
Tests: test/multiuser-auth.test.ts (10, live server on 3170/3171). Existing auth
suite (auth-security, qr-auth, cod54-hook-event, network-auth-policy) unchanged
and green; full test:ci sweep passes (3519 tests).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds Anthropic's classifier-guarded low-prompt mode (--permission-mode
auto) as a fourth ClaudeMode alongside skip-permissions/normal/allowedTools.
Wired through both spawn paths (buildPermissionArgs for direct PTY,
buildClaudePermissionFlags for tmux), the getClaudeModeConfig validator,
and the App Settings Startup Mode picker. Exports buildSpawnCommand for
test coverage.
This is the prerequisite for multi-user mode section 6.3, which downgrades
non-granted users' sessions to 'auto'.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Session: _docker field, constructor config, toState, createSessionOptions/
respawnPaneOptions (both interactive + shell paths), docker getter
- resolveMuxAttachCwd returns /tmp for docker sessions (local wrapper only execs)
- skip the LOCAL claude version probe for docker; probe the IN-CONTAINER version
instead (deferred) so wheel-forwarding stays enabled (#154)
- server restoreMuxSessions round-trips MuxSession.docker / SessionState.docker
- full CI suite green (3444 passed)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pin/unpin a session via POST /api/sessions/:id/pin {pinned}; pinned sessions
sort above unpinned in the unified session manager list (COD-121), ordered by
pinnedAt descending. Pin state lives on SessionState, persists to state.json,
and survives reload/reconnect/restart (persisted-input carries pinned; the
merge skips undefined so a recovered live session can't clobber it). New SSE
event session:pinned re-sorts the open list live across clients. Pin/Unpin
affordance in the session-row kebab menu with a 📌 glyph + amber highlight.
(cherry picked from commit 82749747039afcd4a3104f6a97ce7d3c2ddd048d)
Continuous remote-only reconnect watcher closing the COD-104 durability
arc: when a remote session's local ssh pane dies mid-run, re-establish it
automatically instead of leaving a dead pane until the user pokes it.
Design decisions (per cod108 design doc):
- D1 event->owner: TmuxManager watcher DETECTS a dead remote pane and emits
`remoteSessionDropped`; the session owner (server) reassembles the same
RespawnPaneOptions and calls Session.reattachRemote() -> respawnPane, which
re-runs the idempotent remote command (owned new-session -A / non-owned
attach) and REJOINS the still-running durable remote tmux session. The
watcher never reassembles options itself, and never routes through the
Claude-idle respawn-controller.
- D2 bounded backoff: per-session exponential backoff [5s,15s,45s,2m,5m,5m],
reset on a successful reattach, `remoteReconnectExhausted` emitted once after
the cap. Pure, unit-tested schedule + eligibility decision.
- D3 always-on + kill-switch: `remoteAutoReconnect` app setting (default ON),
read each tick; when false the watcher does nothing.
Guards: killSession() (incl. the non-owned DETACH early-return) and shutdown
add the session to an intentional-teardown guard set + clear its backoff
BEFORE teardown, so a closed/killed tab is never auto-revived. Exactly one
reconnect in flight per session (inFlight guard prevents stacked respawns).
Per-session reconnect/guard state cleared on session removal.
New: src/remote-reconnect.ts (pure backoff + decideReconnect), TmuxManager
startRemoteReconnectWatcher/stop + runRemoteReconnectTick + noteRemoteReconnect
+ guardRemoteReconnect + clearRemoteReconnectState; Session.reattachRemote()
(+ extracted _buildRespawnPaneOptions, shared with interactive start); server
wiring + watcher start; 3 SSE events (sse-events.ts + constants.js in sync,
broadcast + app.js exhausted "Reconnect" affordance); remoteAutoReconnect
schema + settings-ui toggle.
Tests: test/remote-auto-reconnect.test.ts (21) - pure schedule, eligibility
(guarded never reconnects, non-remote/pane-alive/not-due skip, over-cap
exhaust), and manager-level integration (dead remote pane -> dropped ->
backoff -> exhausted; guarded emits nothing; reset-on-success; kill-switch
off; state-cleared-on-remove). Verified real-remote against aa-desktop: drop
local ssh pane -> watcher emitted -> respawnPane reattached the SAME remote
session (remote pane_pid unchanged 3939->3939); test session cleaned up, the
real host sessions left untouched.
Checks: tsc, eslint, check:frontend-syntax, check:public-assets, prettier
--check, build all green; tmux-manager/session-routes/session-manager/
sse-registry-parity suites pass.
(cherry picked from commit d13d58b1994eb6594fd2eadea208104d36204f9d)
Deterministic claude --version probe seeds cliVersion so wheel-forwarding
to Claude's transcript engages (banner scrape was unreliable on 2.1.187+
and resumed sessions). Shift+wheel reads the dominant axis so a trackpad's
horizontal Shift-scroll reaches local scrollback. New per-device
"Wheel Scrolls Local History" opt-out. Wheel reports use a fire-and-forget
send path so they no longer flicker the pending-bytes indicator.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Includes review fixes: full route-test coverage for the rollout locator/parser (originator/uuid/pin resolution, dedup, injected-context filtering), LRU caches, multi-block text joins.
# Conflicts:
# src/web/routes/session-routes.ts
- Breaker reset is now explicit-only: POST /api/sessions/:id/interactive no
longer unconditionally resets the PTY-exit breaker (that endpoint IS the
frontend's automatic re-attach path, so the breaker could never trip on the
COD-115 crash loop and any tab click silently re-armed it). The route accepts
a schema-validated optional body flag {clearBreaker:true}
(InteractiveStartSchema) and resets only when it is sent.
- Frontend restart control: app.js selectSession keeps the bare auto-attach
(no body, never clears); when the selected session has respawnBlocked it asks
for explicit user confirmation and only then re-POSTs with clearBreaker:true.
respawnBlocked is surfaced via SessionState/toState() (runtime-only, not
restored on boot so recovery can re-attach).
- Trip observability: WebServer.setupSessionListeners() is now idempotent
(skips while refs are attached) and the re-attach routes (/interactive,
/interactive-respawn, /shell) re-run it, restoring the wiring that the exit
handler detaches on every PTY exit — without this the 5th-exit trip had
guaranteed zero listeners (no SSE, no push, no persist, no run-summary).
- Push notification: added SessionRespawnBreakerTripped to PUSH_EVENT_MAP
('Session crash loop stopped', urgency critical) with an exit-count body
branch; previously sendPushNotifications silently no-oped.
- Minor: buildMuxAttachEnv() truecolor param is now actually passed
(codex/gemini, mirrors buildEnvExports); buildClaudeEnv() uses delete for
COLORTERM/CLAUDECODE (same node-pty "KEY=undefined" quirk as COD-115).
- Tests: route tests assert auto-reattach does NOT reset, clearBreaker resets,
invalid flag rejected, and listener re-wiring on /interactive + /shell;
real-wiring lifecycle tests (createSessionListeners/attach/detach) prove the
exit-detach gap and that re-setup keeps the 5th-exit trip observable;
PUSH_EVENT_MAP regression guard.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Pass the attachment request `source` through the server deps lambda and make
it a required param on SessionListenerDeps.registerAttachment + the wiring
event type (the 2-arg lambda silently dropped `source`, force-confining every
codex-generated artifact — the feature never worked outside the workspace);
new test/session-listener-wiring.test.ts asserts the pass-through
- Gate the Codex `Saved to: file://` scanner on mode === 'codex' via a
codexArtifacts option threaded from the session call site; magic links stay
mode-agnostic; tests assert claude/shell sessions never emit codex-generated
requests
- Decide the generated-artifact trust policy on the realpath-RESOLVED path
(unresolvable → force-confined) and anchor the ~/.codex marker dirs to
os.homedir() prefixes with startsWith instead of substring matching; symlink
escape + unanchored-marker regression tests added
- Run the Codex scanner on stripAnsi'd data so trailing SGR sequences don't
ride into the captured URL; styled 'Saved to:' test added
- Extend generateFirstPageThumbnail with jpg/jpeg/gif/webp passthrough and
per-extension content types (mirrors the png passthrough) so the PR's new
image formats render real thumbnails instead of 204 letter-tiles
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The response-viewer (eye) currently reads only ~/.claude/projects — for
Codex panes it falls back to a raw terminal-buffer dump. This adds a
Codex-aware reader with exact per-pane rollout attribution.
Locating THIS pane's rollout (~/.codex/sessions/**), in confidence order:
1. history match — Session tracks the pane's last Enter
(codexLastSubmitAt); correlating it against ~/.codex/history.jsonl
{session_id, ts} entries identifies the thread the pane is ACTUALLY
on, surviving /resume, /new and /fork typed inside the codex TUI.
An entry is credited to the pane whose Enter is closest, so menu
keystrokes in other panes can't steal attribution.
2. originator match — codex panes are spawned with
CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_<sessionId>, which codex
(verified on 0.144.1) writes into session_meta.originator of every
rollout it creates.
3. resume-id match — resumed rollouts keep their original session_meta
(codex appends without rewriting), but the uuid is in the filename.
4. cwd+mtime heuristic — case-blind compare (codex records launch-time
path case) and rollouts claimed by other panes are excluded.
Reader details: user turns come from event_msg/user_message (real input
only — AGENTS.md / environment_context injections never appear there),
deduped against legacy response_item rows per-text so mixed-version
rollouts keep full history; image inputs render an [image xN]
placeholder; session_meta identity is cached per path (write-once).
Frontend: thread role label follows session mode (Codex/Gemini/
OpenCode); the terminal-buffer fallback is Claude-only — TUI modes show
a clear placeholder instead of a repaint dump.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Defense-in-depth after COD-115. If the interactive PTY exits non-zero
repeatedly within a short window, recovery/reconnect paths recreate it
indefinitely (COD-115 saw 114 'exited with code: 1' events + orphans).
- New pure InteractivePtyExitBreaker (session-pty-exit-breaker.ts):
injectable time, sliding window, clean-exit resets counter, stays
tripped until reset(). Defaults: threshold 5, window 10s.
- Session records each interactive PTY exit in the breaker; on trip it
flips _status to 'error', sets _respawnBlocked, emits
respawnBreakerTripped. startInteractive() refuses to respawn while
blocked, so all recovery/reconnect callers stop looping uniformly.
- Explicit user restart (POST /api/sessions/:id/interactive) calls
resetRespawnBreaker() so intentional restarts are never blocked.
- New SSE event session:respawnBreakerTripped wired in sse-events.ts +
constants.js (registries in sync) + session-listener-wiring.ts;
minimal diagnostic toast in app.js.
- Tests: test/respawn-pty-breaker.test.ts (pure trip/reset/window +
MockSession session-level trip/reset).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>