- _wsState now transitions through the full lifecycle: _connectWs() sets
'connecting', ws.onopen (inside the this._ws === ws guard) sets 'connected',
_disconnectWs() resets to 'disconnected' — the connection chip's "WS" state
was previously unreachable (stuck on "WS…"/"HTTP" forever).
- WS registry supersede is now keyed per TAB: the upgrade URL sends
cid = clientId + ':' + per-page nonce (reusing the constructor's page UUID),
while input frames keep the bare browser clientId for seq dedup — two
tabs/windows on one session coexist instead of 4010-evicting each other in a
perpetual 5s ping-pong; a genuine same-tab reconnect still supersedes.
- Exponential backoff engages: _disconnectWs() no longer zeroes
_wsReconnectAttempts (it's called at the top of _connectWs, so every retry
replanned at attempt 0 → ~0ms tight reconnect loop during outages); onopen
resets the counter on success.
- styles.css: add .connection-dot.connected (green) and .connection-dot.fallback
(yellow) — both states rendered an invisible dot (no rule existed).
- Remove smuggled dead code: resolveMonitorRowLabels/CodemanMonitorLabels
(COD-122, no consumer, referenced test doesn't exist) and the never-written
_wsLastClose/_wsInputSendCount/_httpFallbackSendCount diagnostics.
- Tests: new test/ws-state-lifecycle.test.ts drives the REAL
_connectWs/onopen/onclose/timer cycle (state transitions, escalating backoff
delays, composite cid on the upgrade URL); registry two-tab coexistence test;
static check that every emitted connection-dot class has a styles.css rule.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A freshly created shell session rendered blank until a tab-switch. selectSession()
fetches the terminal buffer, but for a just-started shell that fetch resolves before
the PTY emits its prompt, so the buffer is empty; the prompt then arrives as a live
SSE event queued during the load and _finishBufferLoad() discarded it. The discard is
correct for an established session (its fetched buffer already contains that output),
but harmful when the load painted nothing.
_finishBufferLoad(owner, { flushQueued }) now REPLAYS the queued events through
batchTerminalWrite (after _isLoadingBuffer is cleared, so they write through, not
re-queue) instead of discarding. selectSession passes flushQueued only in the empty
branch (no fresh buffer + no cache), so the established-session de-dup path is
unchanged. TDD: test/terminal-buffer-flush.test.ts exercises the real begin/finish
mixin (vm-harness, no jsdom).
_updateConnectionIndicator() ran on every keystroke (_reliableSend) and
every ACK (_ackDelivery), unconditionally writing display/className/
textContent/title. During fast typing the rendered output is usually
identical between calls, so those were wasted main-thread DOM writes.
Extracted the branch logic into a pure DOM-free _computeConnectionDescriptor()
returning { display, dotClass, text, title } (every branch/string preserved
verbatim; hidden state normalizes the three non-display fields to '' so the
compare is well-defined). _updateConnectionIndicator() now computes the
descriptor, compares all four fields against a cached _lastIndicatorDescriptor,
and early-returns when unchanged — otherwise caches and writes the DOM exactly
as before (display always; dotClass/text/title only when shown). First call
renders (cache starts null). Perf only, no behavior change.
Tests: test/connection-indicator.test.ts — 9 descriptor cases pinning the
exact strings per state + 4 skip cases (first call writes; two identical calls
write DOM once via counting setters; state change and hidden->shown re-render).
31/31 with input-send-order regression; build, frontend-syntax, prettier clean.
MAX_WS_PER_SESSION was gated by a bare Map<sessionId,number> counter,
incremented on upgrade and decremented only on the old socket's async
close. A client that dropped and immediately reconnected could land its
new upgrade before the old socket's close fired, briefly over-counting and
tripping a spurious 4008 (-> HTTP fallback). The limit also counted raw
sockets, so a reconnecting client consumed a new slot instead of its own.
Replace the counter with WsConnectionRegistry (new pure, unit-tested module)
that tracks live sockets per session keyed by clientId. A same-cid upgrade
SUPERSEDES its own socket (evicts the stale one with close 4010, reuses the
slot, no net count change) -> a reconnect can never be rejected by the cap.
The reliable-input protocol (shouldApplyInput(cid,seq)) already assumes one
logical client per cid per session, so same-cid eviction is principled, not
a regression of multi-tab (which already collides on seq). Slots are freed
EAGERLY on error/terminate, not just async close; close is identity-matched
so a superseded socket's late close is a no-op. cid-less upgrades are
admitted anonymously up to the cap and never evict (backward-compat).
Client sends cid on the WS upgrade URL (?cid=, encoded, omitted if absent).
Tests: ws-connection-registry.test.ts (reconnect-reclaim at cap, rejects
N+1th distinct, eager-terminate frees slot, cid-less up-to-limit + no-evict,
late-close-no-evict, per-session isolation) + route integration in
ws-routes.test.ts (real upgrade through the cap). 45/45 across registry +
ws-routes + input-send-order + ws-reconnect-plan; tsc 0, build, prettier,
frontend-syntax clean.
A reliable-input frame could be stranded forever if its server ACK
({t:'ia',seq}) was lost while the WebSocket kept delivering other output.
_drainSession's WS fast path skips records with sentAt!==0, and after
COD-134 the sweep only force-closes a *silent* socket -- so a lost ACK on
an otherwise-live socket (stale && !silent) was never re-sent.
_redeliverSweep now, for an active-WS session whose oldest unacked frame
is stale but the socket is NOT silent, resets sentAt=0 on every stale
unacked frame and lets the existing _drainSession re-drive them over the
live socket (server dedups by seq). The stale && silent force-close
remains the fallback for a genuinely half-open socket. Restores the
exactly-once recovery guarantee without reintroducing the flap.
Tests: new failing-first COD-135 cases in test/input-send-order.test.ts
(re-drive on live socket; leave not-yet-stale alone; keep stale+silent
force-close). 18/18 across input-send-order + reliable-input-dedup +
ws-reconnect-plan; tsc 0, frontend-syntax, build all clean.
Root cause of the WS->HTTP->WS flapping: the v1.1.15 input-delivery merge left a
call to the now-undefined _flushHttpFallbackQueuesViaWs() in ws.onopen, so every
(re)connect threw a TypeError BEFORE _onWsReady() ran -- durable input was never
re-flushed over the fresh socket, the 2s redeliver sweep then saw stale unacked
frames and force-closed the socket, reconnect, throw again: a self-sustaining
flap loop. Remove the dead call (_onWsReady, 10 lines below, is its replacement).
Resilience + observability:
- Pure CodemanWsReconnect.plan(code, attempt) (constants.js, TDD, 6 tests):
<4004 -> fast reconnect (immediate jittered first retry, faster backoff);
4008/unknown->=4004 -> bounded retry-fallback (HTTP no longer sticks until a
tab switch); 4004/4009 -> give up (session gone). Wired into onclose.
- Redeliver sweep force-closes only a SILENT socket (no recent recv), not one
actively delivering output/ACKs -- stops self-inflicted flaps while typing.
- Client logs WS close code/reason to crash-diag; server logs [ws]
open/close/terminate/4008 (console -> journald; Fastify runs logger:false).
Verified: 6/6 unit, tsc 0, frontend-syntax + prettier clean, build; beta WS
reaches connected with zero console errors (onopen TypeError gone),
_wsLastRecvAt tracked, server [ws] lines emit.
The upstream v1.1.15 merge spliced upstream's transport-object indicator
body onto local's _connectionStatus-based _updateConnectionIndicator()
without defining `transport`, so every transport.* reference threw
ReferenceError on any queued state. That hid the "WS" status and, because
_reliableSend() updates the indicator before _drainSession(), made every
keystroke skip immediate delivery (input flushed only on the 2s sweep =
typing lag).
- Rewrite _updateConnectionIndicator() to show the terminal WebSocket
transport from _wsState (WS / HTTP / WS… / Offline), falling back to the
SSE _connectionStatus only on the idle dashboard.
- Only annotate a backlog (· N queued) above 4 bytes so normal typing no
longer flickers "sending 1B" on each key press.
- test/connection-indicator.test.ts (new): transport labels, the >4B
threshold, an exhaustive never-throws guard for the ReferenceError, and
the _reliableSend -> _drainSession invariant (typing-lag guard).
- test/input-send-order.test.ts: reconcile to local's durable input layer
(the prior coalescing-fallback tests had been failing since 1255e28).
Introduce src/config/terminal-history.ts: one place for terminal scrollback,
tmux history-limit, and PTY buffer byte caps, each overridable via env var or
the settings object and bounds-clamped via resolveTerminalHistoryConfig().
Defaults match the prior hardcoded values, so this is behavior-neutral. Wires
the resolver through buffer-limits, tmux-manager (incl. a setHistoryLimit so a
settings change applies live), session, server, system-routes, session-routes,
schemas, and the config port. Adds 4 optional settings keys (terminalScrollback
Lines, tmuxHistoryLimit, terminalBufferMaxBytes, terminalBufferTrimBytes) with
bounds + a trim<=max cross-check.
Gemini (PR #134) blockers:
- runGemini() now unwraps the {success,data} envelope: status check reads
.data.available, quick-start reads data.data.sessionId (was reading the raw
shape, so the Run-Gemini button could never start a session).
- setGeminiEnvVars() now uses the socket-scoped ${this.tmux()} setenv instead of
bare tmux — Gemini/Google auth env vars were targeting the wrong tmux server
and silently failing on every install.
Gemini parity polish:
- gemini tab-mode badge ('gm') + .tab-mode.gemini CSS; kill-dialog label
'Kill Tmux & Gemini'; codeman doctor dependency-registry entry; export
isGeminiAvailable from utils barrel; COLORTERM=truecolor + unset NO_COLOR;
add gemini to isAltScreenStripMode (Ink TUI, repaints inline like Codex/Claude).
- Revert 4 system-routes.test.ts envelope assertions weakened to
(body.message ?? body.error) back to (body.success === false).
- Add a runGemini() vm-sandbox test that drives the envelope path end-to-end.
Ralph todo-config (PR #135): maxTodos/todoExpirationMinutes are now persisted
and read back — surfaced via the loopState getter (RalphTrackerState) into
toState()/SSE broadcast and restored in restoreState(), mirroring maxIterations.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Ralph settings modal sent maxTodos/todoExpirationMinutes but RalphConfigSchema
(zod) stripped them and the ralph-config route never applied them, so the inputs
were silent no-ops.
Fix: add both as optional positive-int fields to RalphConfigSchema; destructure
and apply them in the ralph-config route (matching the maxIterations pattern).
RalphTracker had no setters (the values were module constants) — added per-instance
_maxTodos/_todoExpiryMs (defaulting to the same constants, behavior unchanged),
switched the eviction + expiry sites to read them, and added
setMaxTodos/setTodoExpirationMinutes (minutes→ms) + getters.
Test: route test POSTs the two fields and asserts the route applies them to the
tracker. Verified RED (setters not called — fields stripped) → GREEN. 34/34
ralph-routes tests pass; tsc + eslint(src) + prettier + build clean. Frontend
already sent the fields (no change).
A "sent" prompt could vanish with no trace on a flaky connection (e.g. a train):
with local echo on, Enter cleared the overlay then sent over the WebSocket
fire-and-forget. On a half-open socket (readyState===OPEN, dead TCP) ws.send()
doesn't throw, so the frame was silently discarded, nothing was enqueued, and
navigator.onLine stayed true — the prompt was lost and never resent.
Replace the best-effort offline queue with a durable, acknowledged delivery layer:
- Client (app.js): every input frame is recorded with a stable clientId +
monotonic per-session seq and persisted to localStorage BEFORE delivery, and
only dropped on a server ACK. Delivered over WS (acked via {t:'ia',seq}) or,
when the socket is down, POST in seq order (HTTP 2xx = ACK). A 2s sweep
force-reconnects a WS whose oldest frame is unacked past 4s (half-open sockets
never recover on their own); on reconnect/reload all pending frames re-deliver.
Survives reconnects AND page reloads. Connection indicator shows pending count.
- Server: Session.shouldApplyInput(clientId, seq) applies each frame exactly once
(bounded MRU map); ws-routes + POST /input dedup a redelivered seq but still ACK
it (200 / {t:'ia'}), so an at-least-once resend can never type the prompt twice.
Untagged input (curl/legacy) applies unconditionally — no behavior change.
- terminal-ui.js sendInput() (voice / keyboard-accessory / paste) now routes
through the same durable layer.
Tests: test/reliable-input-dedup.test.ts (exactly-once semantics on the real
Session) + POST /input dedup route tests. Design: docs/reliable-input-delivery.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- Daylight Blue: Cloudflare Tunnel welcome button is now purple (was orange),
keeping Claude blue / Tunnel purple / OpenCode green distinct.
- Allow enabling the Cloudflare tunnel with no CODEMAN_PASSWORD via the UI: the
toggle now pops a security confirm dialog and, on confirm, sends an explicit
per-request acknowledgeUnauthTunnel:true (new action field, never persisted).
Server logs a loud warning whenever a passwordless public tunnel starts.
curl/API/CLI stay refused unless password/env/flag — no accidental exposure.
Tests: extend test/routes/system-routes-tunnel-guard.test.ts (ack allows + not
persisted; ack:false still refuses). Verified e2e on an isolated instance
(purple button, confirm dialog, retry carries the flag, no real tunnel opened).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Terminal scroll-up intermittently broke for Claude sessions (most visible on
iPhone). Claude Code periodically emits alt-screen switches (?1049h/?47h/?1047h),
scrollback-erase (3J), and mouse-tracking enables for full-screen UIs, which move
xterm.js to the scrollback-less alt buffer / wipe saved lines / hijack the wheel.
Codeman stripped these but only for codex mode.
Share the strip via isAltScreenStripMode(mode) = codex || claude, applied at both
sites that were codex-only: the live PTY stream (Session._handleTerminalOutput,
incl. the chunk-boundary carry) and the /terminal buffer replay. shell stays
excluded (vim/less/htop need the alt screen); opencode unchanged.
Tests: test/claude-scrollback-strip.test.ts (8 new); codex strip tests unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Workflow runtime writes workflows/wf_<id>.json only at completion (always
terminal), so workflow-run-watcher never saw a run until it was already done and
the ACTIVE-gated floating window never popped. The watcher now also scans
subagents/workflows/wf_<id>/ and synthesizes a minimal running record (agentId
slots preserved for the transcript-click join, lastActivityAt from mtimes,
done/running from the journal), superseded by the real wf_<id>.json at
completion. Standalone (no subagent-watcher import). Verified e2e on a real
in-flight run; +6 unit tests. Bumps 1.1.3 -> 1.1.4.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Opt-in (showUltracodeAgents, default OFF) panel that visualizes ultracode /
Workflow-tool runs like Claude Code's "working agents" TUI: LEFT = runs + phases
(selectable tasks), RIGHT = each run's agents with model, live state, tokens
burned, and tool calls.
Standalone — ZERO edits to subagent-watcher.ts. A new workflow-run-watcher.ts
singleton globs the run-state tree (~/.claude/projects/*/*/workflows/wf_*.json,
disjoint from the transcript tree), strips the heavy script/scriptPath/result/logs
fields (174KB -> ~25KB/run), and emits workflow:run_* SSE events. The LEFT list
ships lightweight summaries (getLightState replay + SSE); the RIGHT pane fetches
the full run (with agents[]) via GET /api/workflows/:runId on selection.
Backend: workflow-run-watcher.ts, types/workflow-run.ts, config/workflow-config.ts,
3 SSE events, getLightState workflowRuns replay, GET /api/workflows[/:runId],
showUltracodeAgents schema key + boot-gate (default OFF) + live toggleService.
Frontend: ultracode-panel.js (debounced master-detail render, run/phase select),
header launcher (btn-ultracode-agents--hidden marker -> mobile-guard-exempt),
App Settings toggle (SYNCED, deliberately not in displayKeys).
Agent states on disk are start|progress|done (start=queued; done has
durationMs/resultPreview). Tests: workflow-run-watcher (9), workflow-routes (3).
Verified: tsc/lint/prettier/frontend-syntax/public-assets/mobile-header-guard
clean; full test:ci green (2986 passed); live server + Playwright e2e against 25
real runs (28-agent grid, phase filter, OFF hides launcher).
Design: docs/ultracode-agent-viz-plan.md (rev. 3).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two follow-ups to db93491 (the 2026-06 CC meta.json format change), after
reverse-engineering the new on-disk layout with a live current-CC subagent +
1Hz fs poller:
(1) Workflow recursion — the Workflow tool nests its agents at
subagents/workflows/{wf}/agent-{id}.jsonl, one level below the flat
subagents/ scan, so they were never tracked. Add watchWorkflowDirs()
(driven from scanForSubagents) to descend and watch each workflow dir
(idempotent; fs.watch recursive is unsupported on Linux, so the ~5s
periodic scan re-drives it — same latency as new-session discovery).
Require the `agent-` prefix in the flat readdir + watch callback so a
workflow dir's sibling journal.jsonl can't register a bogus "journal" agent.
E2E verified against real ~/.claude/projects: 32 workflow-nested agents
discovered (wf_fa35c1d8-4a9), 0 bogus journal agents.
(2) Transcript timing — empirically the per-agent .jsonl IS written at the
standard subagents/ path and grows incrementally (tailable); the
/tmp/.../tasks/<id>.output the prior probe found is just a symlink back to
it. meta.json lands at spawn, the .jsonl a beat later. Add a meta→transcript
upgrade in registerAgentFile: when an agent registered meta-only gets its
sibling .jsonl, re-point filePath, drop the stale sidecar context, start
tailing, and emit subagent:updated (not a duplicate discovered). Corrects the
now-inaccurate "no transcript to tail" doc comment on registerAgentMeta.
Tests: 2 new cases (workflow-nested discovery; journal.jsonl not registered).
All 56 pass; tsc/lint/format clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude Code changed its subagent on-disk format (~2026-06-14): TUI Task
subagents now write `agent-{id}.meta.json` ({agentType,description,toolUseId})
into the session's `subagents/` dir and no longer reliably write a per-agent
`agent-{id}.jsonl` transcript there. The watcher discovered agents ONLY by
`.jsonl`, so it tracked zero — subagent windows and the monitor's "N TRACKED"
showed nothing.
- Add `registerAgentMeta()`: discover from the meta sidecar (description from
meta.description/agentType), prefer a sibling `.jsonl` transcript when present
(richer), never tail a meta file.
- Initial scan + directory watcher now handle `.meta.json` alongside `.jsonl`.
- Tests: 2 new cases (meta-only discovery; prefer-.jsonl-when-present).
Verified e2e against a real ~/.claude/projects fixture.
Known follow-ups (not in scope): meta-only agents have no per-agent transcript
to tail (no live tool-call feed, status stays 'active'); workflow agents under
`subagents/workflows/{wf}/agent-*.jsonl` are still missed by the flat scan.
Also adds the README screenshot tooling used to surface this:
- capture-real-overview.mjs: DSF=2 + ?nowebgl crisp path (DOM renderer avoids
the WebGL glyph-doubling at deviceScaleFactor>1).
- capture-readme-real.mjs: real-instance desktop-scene capture (dashboard/
monitor/subagent) for an isolated beta seeded from prod settings.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The File Browser preview and Attachments preview share openFilePreview(),
but the workspace branch (via /file-content) misclassified several types the
attachments viewer handled fine:
- SVG was reported as type:image, but file-raw serves SVG as octet-stream +
attachment (XSS hardening), so the <img> broke. Now fetched and rendered via
a same-origin image/svg+xml blob <img> (safe; <img> never runs SVG scripts).
file-raw's SVG hardening is unchanged.
- Audio (mp3/wav/ogg/m4a/aac/flac/opus) was type:binary -> "Cannot preview".
Now classified as audio and rendered with <audio controls>; file-raw gained
the matching audio/video MIME types so playback works.
- Binary formats not in the hardcoded list (xlsx/doc/zip/...) were decoded as
UTF-8 and dumped as mojibake. Replaced the static list with a NUL-byte
content sniff that flags arbitrary binaries; the binary fallback now offers a
Download link instead of dead-ending.
Adds route tests for audio, known-binary (xlsx), and NUL-sniff classification.
Verified end-to-end on an isolated instance + headless browser.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Review polish on the desktop tab auto-wrap:
- Auto-wrap is purely width-driven, but updateTabOverflowMode() was only called at the
tail of _renderSessionTabsImmediate (SSE content renders). Window resize — the primary
trigger for tabs crossing the one-row overflow threshold — never re-evaluated it, so
narrowing/widening the window left the wrap state stale until an unrelated status event
fired a render. Call it from the debounced window-resize handler (no-op on
mobile/tablet, where the method bails).
- Move the re-evaluation into _fullRenderSessionTabs() as well, so the incremental
branch's two early `_fullRenderSessionTabs(); return;` paths (badge add/remove, which
change tab width) and the manual two-rows toggle (applyTabWrapSettings → _fullRender…)
re-evaluate too. The latter also fixes a transient where enabling manual two-rows while
auto-wrap was on left both classes set (clipping folder tabs to 96px) until the next
render.
- Add boundary cases to the policy test: exact fit and the +1 sub-pixel tolerance (no
wrap), 2px over (wrap), and a single overflowing tab (no wrap).
Verified: tab-overflow test passes; tsc, check:frontend-syntax, check:public-assets,
prettier all clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Making the hook-event secret unconditionally required closes the own-loopback-proxy gap,
but it would also silently 401 the hook curls baked into cases created BEFORE the secret
header existed (COD-54, 2026-06-10): writeHooksConfig only runs at case CREATION, so an
existing/linked case on a password-protected install keeps secret-less curls that the new
gate rejects (degrading idle/stop/teammate/task signalling with no error surfaced).
No-password installs are unaffected — the gate isn't registered without CODEMAN_PASSWORD.
Add `refreshStaleHookSecret(casePath)` and call it on Claude-mode spawns in
POST /api/sessions and POST /api/quick-start (existing-case branch). It regenerates the
hooks block ONLY when settings.local.json already holds Codeman's own hook curls (they
target /api/hook-event) that lack the X-Codeman-Hook-Secret header — a no-op when the
hooks are absent, not ours, or already current, so it never clobbers user customizations
and is cheap on every spawn. Fresh cases are unaffected (writeHooksConfig already wrote
the secret). withSettingsLock serializes it with the model/statusLine writers.
Verified: new test/hook-secret-selfheal.test.ts 5/5 (heal + key-preservation + no-op on
current/foreign/absent/malformed); the PR's cod54 + auth-security suites still pass
(36); tsc, lint, format:check, and npm run build all clean (symbol present in dist).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The PR gates CJK textarea visibility on an active session
(`showCjk = cjkUserEnabled && !!activeSessionId`) so the fixed-position textarea no
longer floats over the welcome overlay. That intentionally changes the behavior the
existing `shows the CJK textarea on mobile only for server override` test asserted —
it set `_serverCjkOverride = true` on a fresh page (no active session) and expected the
textarea visible, which now (correctly) resolves to hidden. The test lives in
test/mobile/** (excluded from CI), so it wasn't caught by the PR's green CI.
Update the test to verify the new, intended behavior: with the server override on it
stays hidden on the welcome screen (no active session) and is revealed once a session
is active. This is a co-authored review fix; the original change is TeigenZhang's.
Verified: tsc, check:frontend-syntax, check:public-assets, prettier all clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The PR migrated the app.js tab-nav handler to physical e.code but left xterm's
pass-through gate (terminal-ui.js) matching ev.key digits. Consequences:
- Alt+[ / Alt+] (the new bindings) were never in the gate, so xterm sent ESC[ / ESC]
to the PTY on every platform AS WELL AS switching the session.
- Alt+digit on a remapped macOS Option layout (Option+1 -> "¡") didn't match the
ev.key '0'-'9' gate either, so xterm injected ESC<char> — on exactly the layouts
this PR exists to fix.
Update the xterm gate to mirror app.js exactly: suppress when
`ev.altKey && !ctrl && !shift && /^(Digit[1-9]|BracketLeft|BracketRight)$/.test(ev.code)`.
Returning false there tells xterm not to write to the PTY, so the shortcut switches
the tab with no stray escape sequence.
Also: relabel the docs Alt/Option (the mechanism is layout/OS-independent, so the
shortcut works for Linux/Windows Alt users too — "Option" alone was Mac-only wording),
and add a keyboard-shortcuts test asserting terminal-ui.js gates on the same physical
codes so this desync can't regress (a grep the original test missed).
Verified: keyboard-shortcuts test 4/4, check:frontend-syntax, check:public-assets,
format:check all clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Review fixes on top of the `codeman doctor` checker:
- Node minVersion 18.0.0 -> 22.0.0. package.json engines is ">=22.0.0" and the docs/CI
require Node 22+, so doctor was green-lighting Node 18-21 (a false pass).
- Remove the phantom `gemini` registry entry. Codeman has no Gemini backend
(SessionMode = 'claude' | 'shell' | 'opencode' | 'codex'); the entry advertised a
dependency that nothing uses.
- Add `pdftoppm` (poppler) to the office group. document-thumbnailer.ts calls pdftoppm
with no fallback as the sole PDF/Office first-page thumbnail renderer, yet it was
absent from the registry, so doctor never reported it missing.
- Fix the `--category` mismatch: the help advertised `documents|media` categories that
the ToolCategory type/registry never defined, and an unknown category silently
produced an empty "all healthy" table. Introduce TOOL_CATEGORIES as the single source
of truth (type + help + validation); an invalid `--category` now errors with the
valid list and exits 2.
Verified: tsc, lint, format:check all clean; both dependency tests pass (20);
`doctor` runs correctly (Node 22.22 ok, pdftoppm detected, no gemini), `--category media`
errors with exit 2, `--category office` lists libreoffice/pdftoppm/msoffice.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Review fixes on top of the DOMPurify mXSS hardening:
- Remove `USE_PROFILES: { html: true }` from the sanitize-html.js config. DOMPurify
treats USE_PROFILES and ALLOWED_TAGS/ALLOWED_ATTR as mutually exclusive — with a
profile set it resets the allow-lists to the full HTML profile and silently ignores
the curated lists, so the tight markdown-only allowlist was dead config (still
XSS-safe via FORBID + core, but far broader than intended: <button>/<input>/
<details>/<audio>/<select>/<label> all survived). Dropping USE_PROFILES puts the
curated ALLOWED_TAGS/ALLOWED_ATTR back in force; FORBID_TAGS/FORBID_ATTR stay as
defense-in-depth and DOMPurify keeps its default safe-URI handling.
- Rewrite test/markdown-sanitizer.test.ts to run in the default node environment with
an in-test jsdom window instead of a per-file jsdom environment. That environment
externalizes node:fs/node:path under vite, so the suite failed to load in isolation
("No such built-in module: node:") and only survived the full CI run because an
earlier node-env test happened to pre-cache node:fs — order-dependent and fragile.
The rewrite is order-robust and adds an "allowlist is actually enforced" block
(non-markdown tags must be dropped) that fails if USE_PROFILES is reintroduced.
Verified: 25/25 tests pass standalone under config/vitest.ci.config.ts; tsc, lint,
format:check, check:frontend-syntax, check:public-assets, and npm run build all clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Tab-switch shortcuts matched e.key, so on macOS Option+1 emits a special
character ('¡', not '1') and the shortcut silently failed. Switch to physical
e.code (Digit1-9), which is layout-independent. Also adds Option+[ / Option+]
for previous / next session. Help modal + README updated.
Test: test/keyboard-shortcuts.test.ts.
When desktop session tabs overflow one row, wrap them to a second row instead
of horizontal scroll — unless the user has pinned the manual two-row layout
(tabTwoRows). Mobile/tablet keep horizontal scroll. The wrap policy
(shouldAutoWrapTabs) lives in constants.js as a pure, unit-testable function;
updateTabOverflowMode() measures overflow after each tab render and toggles
.tabs-auto-wrap.
Test: test/tab-overflow.test.ts (vm-loads constants.js, asserts the policy).
The COD-39 attachments button was hard-visible in the header — first on
mobile, then (after the mobile-only hide) still on desktop. Make it a
proper opt-in App Settings → Display toggle ("Attachments Button"),
default OFF everywhere, mirroring the Response Viewer button:
- index.html: button ships with the `btn-attachments-history--hidden`
marker; new settings checkbox #appSettingsShowAttachmentsButton.
- styles.css: base `display:inline-flex !important` + a more-specific
`--hidden` rule (same pattern as the response viewer).
- settings-ui.js: load/save/getDefaultSettings(false) + a live toggle in
applyHeaderVisibilitySettings. Per-device and NON-leaking — added to
displayKeys AND stripped from the server payload, so enabling it on
desktop never makes it appear on mobile (or any other device). No
server-side render step (purely client display, like the eye button).
- mobile.css: dropped the now-redundant phone-only hide — the opt-in
marker hides it everywhere by default; the per-device toggle governs
both desktop and phone.
Tests updated: the CI static guard drops btn-attachments-history from the
phone-hidden lock (it's opt-in now, excluded from the default-visible
enumeration — the guard still gates any NEW default-visible button); the
real-browser E2E now asserts default-hidden on a desktop-class viewport
and visible after enabling the setting.
Verified on a real desktop browser: hidden by default, the settings
toggle exists, enabling it shows the button. tsc + frontend-syntax +
prettier + public-asset checks + both test suites green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The COD-39 attachment-history header button was visible on the cramped
phone header. Hide it on phones alongside the settings gear and lifecycle
log (the mobile header is intentionally minimal — those controls live in
the toolbar). One-line addition to the existing @media (max-width: 430px)
display:none block in mobile.css.
This is the second time a header control leaked onto mobile (the
plan-usage chip was the first), so add two regression guards:
- test/mobile-header-buttons-policy.test.ts — a pure static analysis of
index.html + mobile.css (no browser), so it runs in the normal CI sweep
(the test/mobile/** Playwright suite is EXCLUDED from CI and never gated
this). It enumerates every default-visible header button and fails when
one has no phone-visibility decision — either a mobile.css hide rule or
an explicit MOBILE_VISIBLE_ALLOWLIST entry. A new header button now
forces that decision. Verified it fails on the pre-fix state and passes
after.
- test/mobile/header-buttons.test.ts — real-browser E2E in the mobile
suite: asserts the attachments/settings/lifecycle buttons are hidden on
an emulated iPhone 14 Pro and the attachments button is visible on a
desktop-class tablet.
tsc + lint + prettier + both new tests green. Only CSS + tests changed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Remove the dead mobile-collapsed header tray that hid the entire header-right cluster (incl. the opt-in response-viewer eye) on phones/tablets, and update the mobile test to assert inline reachability. Eye stays hidden by default (showResponseViewer).
Removing the dead `mobile-collapsed` tray (this PR) means the test that
asserted the headerRight tray *stays collapsed* on mobile now contradicts
the code and would fail when run. Flip it: with the three-dot utility
toggle gone, the header-right utilities must flow inline and stay
reachable on small viewports. The response-viewer eye itself remains
hidden by default (showResponseViewer opt-in), so this only re-exposes
the already-default-visible utilities inline.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
COD-54 gated the /api/hook-event + /api/status-telemetry localhost bypass
behind the shared X-Codeman-Hook-Secret only WHILE a managed tunnel was
running, keeping a plain localhost bypass otherwise. But Codeman can't detect
a user's OWN loopback reverse proxy (their own `cloudflared --url`,
`tailscale serve`, nginx -> 127.0.0.1), which proxies internet traffic into
the loopback origin with req.ip === 127.0.0.1 — so that setup kept the unsafe
plain bypass.
Require the secret on the loopback bypass unconditionally. Managed-session
hooks already always present it (X-Codeman-Hook-Secret from
$CODEMAN_HOOK_SECRET_FILE, generated for every instance), so the legitimate
hook channel is unaffected; only the previously-unguarded own-proxy path is
now rejected. Drops the now-unused getTunnelRunning param from
registerAuthMiddleware.
Tests: cod54-hook-event-auth (tunnel-down now also requires the secret, plus
a good-secret positive case); auth-security (hook tests present the secret to
reach schema validation).
The previous _sanitizeHtml was a denylist over agent/transcript markdown
rendered via innerHTML; it missed style attributes and the svg/math mXSS
namespaces — e.g. <svg><style><img src=x onerror=alert(1)></style></svg>
re-serialized into a live <img onerror>.
Vendor DOMPurify 3.4.8 (allowlist) following the existing marked.min.js
vendor pattern (same-origin, CSP script-src 'self'; not in package.json so
no lockfile drift). New sanitize-html.js wires a hardened allowlist config
(FORBID style/svg/math/script/iframe/object/embed/form; no data attrs);
app.js _sanitizeHtml delegates to it with a fail-closed escape-all fallback.
index.html loads dompurify -> sanitize-html -> app.js (defer); build.mjs
minifies + content-hashes sanitize-html.js.
Test: test/markdown-sanitizer.test.ts (jsdom, real shipping artifacts) —
mXSS payloads neutralized + legit markdown preserved.
Environment-aware dependency probe (linux|darwin|win32|wsl) with a static
registry, an injectable ProbeHost seam for testing, grouped table + `--json`
output, and a non-zero exit when a required dependency is missing/outdated.
Node and tmux are the only hard-required tools; the agent CLIs and document
converters (LibreOffice / MS Office via WSL interop) are optional. CI-safe
unit tests (no tmux, injected host).
Follow-up fixes applied during review of PR #121 (all confirmed minor/nit;
no blockers). Security posture verified sound (externalPath never leaves
toState()/the list route; re-registration runs the guard).
- fix(recovery): restoreAttachmentHistory now skips malformed/legacy saved
items (null, non-object, missing source/fileName) instead of throwing inside
the Session constructor — a corrupt __attachmentHistory entry could otherwise
abort the entire mux-recovery loop. (P1)
- fix(routes): the attachment-list route degrades a single failing entry to
{missing:true} instead of failing the whole drawer. (INT-4)
- fix(ui): give the attachments header button a positioning context so the
unread badge anchors to the icon, not the header bar. (F1/CSS-1)
- fix(ui): cancel the debounced history refresh on drawer close and guard it
against a stale session/closed drawer. (F3)
- fix(ui): re-show ("Card") of a detected item now uses the item's own
timestamp so the cardId is stable — focuses the existing card instead of
stacking duplicates. (F4)
- fix(ui): Escape now closes the drawer, matching every other panel. (UX-1)
- fix(ui): badge shows "99+" past 99 (was an inconsistent 100/99 cap). (BADGE-1)
- style: drop the duplicate @keyframes notif-badge-pulse (dead CSS). (INT-1/CSS-3)
- style: empty-state used three undefined CSS custom properties
(--text-primary/--border-color/--bg-tertiary) → use the defined
--text/--border-light/--bg-input tokens. (CSS-2)
- test: add constructor restore round-trip + malformed-item resilience tests.
Deferred (noted for author): broadcasting the full 100-item history in every
session-state SSE event (payload bloat), "unread" badge semantics, making the
header button opt-in, and app.inject route tests for the two new endpoints.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Stacks on COD-38: accumulates a per-session attachment history and exposes it
through a slide-in drawer with an unread badge, so attachments stay reachable
after their cards are dismissed.
Backend:
- session-attachment-history: history state — dedupe by source path / relative
path, newest-first, 100-item cap, and externalPath sanitization (the absolute
host path is server-private and never leaves toState()).
- session.ts: _attachmentHistory + getter (sanitized) / upsert / restore /
getAttachmentHistoryForPersist; restored from saved state in the constructor.
- file-routes: GET /attachments (list — resolves each entry to live metadata +
routes; external entries are re-registered) and GET /attachments/:id
(metadata poll). The by-id route guards via the registry's TOCTOU-safe
resolveServableAttachmentPath.
- server.ts: detected/registered attachments upsert into history and persist;
the private (externalPath-bearing) history rides on disk under
__attachmentHistory, separate from the sanitized public copy, and is restored
on mux-session recovery.
- types/session.ts: SessionAttachmentHistoryItem + SessionState.attachmentHistory.
Frontend:
- panels-ui: the drawer (lazy-built), unread badge, list render with per-item
preview/download/open/"Card" (reshow) actions, and live refresh of the open
drawer on new detections.
- app.js: history state + per-session badge/cleanup wiring.
- index.html / styles.css / mobile.css: header button + badge and the drawer.
Verified: tsc / eslint / prettier / frontend-syntax / public-assets clean; new
history-module unit tests pass; full test:ci green (2866 passed); badge, drawer
open/render/reshow/close verified in-browser.