Commit Graph
1621 Commits
Author SHA1 Message Date
Codeman maintainer 5aa59c70cc feat(predictive-echo): Phase 0 codex fixtures, measurements, package scaffolding
Recorder (scripts/dev/record-codex-frames.mjs) captures real codex 0.147
TUI output through the production pipeline (tmux status-off + the codex-mode
full strip from session.ts) into JSONL fixtures with keystroke injection
points; analyzer replays them through @xterm/headless for the measurements
in docs/predictive-echo-plan.md. Composer signature /^> /-style (U+203A),
modal and wrapped rows correctly rejected, echo is unstyled default-fg,
tmux delivers echo as minimal in-place deltas.

Package: types.ts gains optional cursorX/cursorY, getCell, onWriteParsed,
onResize (all additive); prediction-renderer.ts renders per-glyph spans
keyed by prediction seq; @xterm/headless@^6.0.0 devDep for replay tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 04:10:18 +02:00
Codeman maintainer ffccde4f7d fix(cli): codeman status probes the running server (#230)
Reported by @mtiller.

`codeman status` runs in its own fresh process, and reported THAT process's
always-stopped Ralph loop under a bare "Status:", which reads as "the web server
is down" while the service is running fine and agents are reachable. It now probes
the real server first (`CODEMAN_API_URL`, else https then http on the local port,
overridable with `--url`) and reports reachability, version and live session
state. Any HTTP answer proves the server is up, including a 401 from a
password-protected install. The Ralph loop keeps its own `codeman ralph status`.

This complements `codeman web --status` from the daemon work: that answers "did I
start a daemon", this answers "is a server running at all", which is what the bare
command was already being used for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:24 +02:00
Codeman maintainer bec3da3d31 fix(ui): a described session tab shows just the description (#232)
Reported by @mtiller.

A session named `w2-foo-bar: some description` rendered both halves on the tab, so
the generated id ate the width that the part the user actually chose needed. The
tab now shows the description alone and the `w<n>-<case>` id moves to the tooltip,
where it stays available without being read every time. It is still shown in the
session settings modal. Undescribed tabs are unchanged.

`aria-label` deliberately keeps the FULL name, so screen readers still get the id.

Also fixes a re-render loop this exposed: the incremental update compared
`nameEl.textContent` against the full name, which for a described tab never
matched, so those tabs re-rendered on every pass. The compare now targets the
display label.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:24 +02:00
Codeman maintainer 6e89eb9ec1 fix(web-tabs): bound time-to-headers, not the whole proxied exchange (#237, #238)
Reported by @DodgyBadger.

#237: the proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which
bounded the ENTIRE exchange rather than the wait for response headers. A dashboard
endpoint doing model inference, and any actively streaming response, both died at
30s as a generic 502 that Codeman never logged, so it read as an intermittent
network error. The timeout now bounds time-to-headers only and is cleared the
moment headers arrive, so a slow endpoint and a long stream both survive. The
default moves to 300s because "the app is thinking" is normal for the dashboards
people proxy; abandoned upstreams are reclaimed by the client-hangup abort rather
than by this value.

A browser that navigates away mid-request now aborts the upstream fetch, guarded
by `writableFinished` for the same reason as `abortOnClientHangUp` in
session-routes: `close` also fires after a completed response and must not abort
anything. Header timeouts are logged as a warning with a sanitized identity
(method plus origin plus path, never the query string, which can carry the
dashboard's tokens), and a client hangup is deliberately not warned since nobody
is listening and it would read as the dashboard being broken.

The WebSocket handshake keeps its own 30s budget
(`CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`), decoupled from the request timeout:
a handshake is connection establishment, and waiting minutes on one only delays
the browser's reconnect logic.

#238: the web-tab guide covered sandboxed dashboards having no cookies, but not
cookie authentication in front of Codeman itself (Cloudflare Access and similar),
where a sandboxed frame's asset and API requests carry no auth cookie, bounce to
the login provider, and leave the embedded app looking unstyled or broken while
trusted mode works. Documented, and the Test button's result now says it probes
server-to-upstream reachability only, not how the page behaves in a sandboxed
frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:23 +02:00
Codeman maintainer 94aa53c65b chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
codeman@1.14.1
2026-08-09 03:38:40 +02:00
Codeman maintainer e88b971bb7 feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.

Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
  Case names resolve through linked-cases.json first, mirroring the
  server's resolveCasePath(), so a case linked in from outside
  ~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
  hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
  skill is never touched, and a symlinked skill dir is refused (this
  repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
  ports/config-port.ts, server.ts, session-routes.ts (add-only injection
  on Claude session create and quick-start), plus the App Settings toggle.

Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
  Shell state does not survive between agent tool calls, and an undefined
  is_self exited 127, firing the `||` branch and deleting the caller's own
  session with the one guard bypassed. The request now lives inside the
  guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
  tool call, so the documented resend-identical-request loop stopped being
  a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
  workers. It returns clean transcript text; the terminal scrape it
  replaces returns a wall of TUI repaint noise. Its transcript flush lags
  the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
  yielded the literal session id "null" and burned the whole readiness
  budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
  corrected the hooks-config comment that claimed a toggle-off sweep
  exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
  and that caseName resolves linked cases, so a generic name can land a
  worker in a real repo.

Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:30:15 +02:00
Codeman maintainer 8406c497e2 fix(terminal): stop forwarding the wheel to codex, it ignores SGR reports
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.

Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.

_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.

Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:22:06 +02:00
Codeman maintainer 40b4aba043 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codeman@1.14.0
2026-08-09 02:35:59 +02:00
Codeman maintainer 4b44988bfc test: give daemon-control tests a unique port (3212 was already taken)
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:22:01 +02:00
Codeman maintainer 316d0a4c82 Merge pull request #233 from Lint111/feat/hooks-config
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:21:52 +02:00
Ark0N 1184720648 Merge pull request #239 from Ark0N/feat/daemon-mode
feat(cli): codeman web -d and codeman service install (#231)
2026-08-09 02:19:50 +02:00
Ark0N b067aad9b6 Merge pull request #235 from Lint111/feat/deferred-terminal-flush
fix(terminal): drain deferred output without a wake event
2026-08-09 02:19:32 +02:00
Ark0N 19a3d7c773 Merge pull request #234 from Lint111/feat/ai-checker-stderr
fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
2026-08-09 02:19:10 +02:00
Codeman maintainer 085f4acb60 feat(cli): codeman web -d and codeman service install (#231)
Two ways to keep the server running, split by how long it should last.

`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.

`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.

Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.

The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:34:55 +02:00
Codeman maintainer d26f26fe34 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
codeman@1.13.0
2026-08-09 01:18:03 +02:00
lior 091df2b6d8 fix(terminal): drain deferred output without a wake event 2026-08-08 23:00:36 +03:00
lior 5f775b1ab1 fix(hooks): guard subagent stops and rewake from the parent transcript
Two defects in the background-task hook scripts.

SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.

The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.

The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.

12 tests fail on unmodified master, e.g.
  expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
  expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
2026-08-08 22:31:38 +03:00
lior da51193264 fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".

stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.

Two tests, both failing on master:
  expected 'export PATH="…' to contain ' 2> "'
  expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
2026-08-08 22:30:39 +03:00
Codeman maintainer fa18eeef35 feat: tab action icons on the active tab only, middle-click closes tabs
Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.

Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.

The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer a9f26bd03a feat: fixed-width tab hover with sliding title, pop-out button now opt-in
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).

The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer 8dc8b164a7 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codeman@1.12.2
2026-08-08 13:48:12 +02:00
Ark0N 2524759655 Merge pull request #229 from Lint111/feat/keyboard-viewport-settle
fix(mobile): coalesce keyboard viewport settling
2026-08-08 12:51:04 +02:00
Codeman maintainer 1f164bc8d2 fix(mobile): only arm the viewport settle on a real keyboard transition
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 12:06:08 +02:00
lior 0a1439b1e9 test(mobile): make the coalescing test actually exercise the settle path
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.

Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.

Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.

Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).
2026-08-08 08:41:04 +03:00
lior 66abe6c70a fix(mobile): coalesce keyboard viewport settling 2026-08-08 08:13:58 +03:00
Codeman maintainer fa1700da5b chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codeman@1.12.1
2026-08-08 01:38:35 +02:00
Ark0N aed1e59ee3 Merge pull request #227 from Ark0N/fix/scrollback-205-round2
fix(terminal): scrollback round 2 for #205 (re-pull downgrade guard, PageUp fallback, CLI version probe retry)
2026-08-08 01:36:03 +02:00
Ark0N 7f6d18b398 Merge pull request #226 from christianhaberl/fix/input-loss-on-failed-delivery
fix(api,ws): an input whose delivery fails can be retried instead of being lost
2026-08-08 01:31:10 +02:00
Ark0N 52571c7fd4 Merge pull request #225 from christianhaberl/fix/bound-the-process-tree-walk
fix(mux): bound the process-tree walk — unbounded pgrep recursion can take a machine down
2026-08-08 01:31:00 +02:00
Ark0N cb95a8562c Merge pull request #224 from christianhaberl/fix/raw-writehead-drops-security-headers
fix(http): raw writeHead routes drop every header the security hook set
2026-08-08 01:30:47 +02:00
Codeman maintainer 3cb7e30636 fix(ui): scope the wheel-opt-out tooltip's paging fallback to Claude
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:29:48 +02:00
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00
Claudia 9d27cc0bab docs: merge the stacked doc comments the previous commits left behind
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.

- write() had two: the original description with @param and @example, then a
  @returns-only block added on top, which dropped the params and examples from
  hover. Merged into one. The @returns wording is also honest now — write() still
  discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
  and its declaration, leaving that function undocumented on hover and the doc
  attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:29:46 +02:00
Claudia 84132d3025 fix(ws): pin the withheld ACK with a test, and correct the changeset
Two blockers from the pre-submission gate, both reproduced before fixing.

1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
   code says the opposite in as many words ("NOT an error response,
   deliberately"), the commit message says response codes are unchanged, and the
   test asserts the 200. It was a leftover sentence from an earlier iteration that
   would have shipped into the CHANGELOG announcing an API contract change that
   does not exist — and errorCode values are SemVer-relevant per
   docs/versioning-policy.md.

2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
   master left all 9 tests green, while the commit message sells "plus the whole
   WebSocket path" as part of the fix. Three tests added against the real WS
   route — ACK on delivery, ACK withheld and seq re-opened when the write did not
   land, and a deduplicated frame still ACKed so the client can drop it. Verified
   the other way round: with ws-routes.ts reverted, the middle one fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:52:44 +02:00
Codeman maintainer cc163792e5 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
codeman@1.12.0
2026-08-07 13:43:47 +02:00
Ark0N 6f1ff17ccc Merge pull request #223 from Ark0N/fix/scrollback-shell-alt-screen
fix: terminal scrollback overhaul for shell and CLI sessions (#205)
2026-08-07 13:42:47 +02:00
Codeman maintainer f262b8cb69 feat(terminal): gentler glide start and fractional wheel accumulation
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:36:27 +02:00
Codeman maintainer 5f2b491d99 feat(terminal): ease-out smooth scrolling for the local wheel path
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:30:34 +02:00
Codeman maintainer c067167dbc fix(terminal): take the wheel in capture phase; xterm's scroller is deaf after reset
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.

The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.

Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:03:17 +02:00
Codeman maintainer ad2ca9b575 docs: record the #205 scrollback mechanisms and the shipped fix plan
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:42 +02:00
Codeman maintainer a7a1cef3d6 fix(session): probe the Claude CLI version over ssh for remote sessions
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:41 +02:00
Codeman maintainer a1d7ec02e9 fix(terminal): forward touch scrolls to the CLI transcript on mobile
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:25 +02:00
Codeman maintainer dfa43928af docs: record the scrollback analysis and its measurements for #205 2026-08-07 04:33:15 +02:00
Codeman maintainer adbb74cd5a fix(terminal): keep the CLI's input box pinned when scrolling with the wheel
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.

_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:

- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
  history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
  REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
  furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
  non-empty row, so any session with trailing blank rows was left off-bottom
  and every later wheel event went local without the user ever scrolling.

Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.

Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
2026-08-07 04:27:16 +02:00
Codeman maintainer eb8d11ffc3 fix(terminal): restore shell scrollback, recover history lost to tmux repaints
Four fixes for the scrollback reports in #205 (plus its follow-up comment).

1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
   ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
   (\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
   to claude/codex/gemini, so it reached the browser verbatim. In the alternate
   buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
   and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
   which readline receives as shell history navigation. Both reported symptoms,
   one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
   modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
   own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
   the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
   never forwards a pane's alt-screen toggles to its client, it repaints;
   captured from a real attach, vim/less/htop emit zero.

2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
   onto tmux's history, and tmux repaints the pane rectangle instead of emitting
   linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
   scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
   same 60 lines emitted slowly added all 60. Scrolling up at the top now
   re-pulls the full tmux scrollback and holds the user's place. Verified
   end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.

3. The full-scrollback replay was gated on a single "first load after page load"
   flag, which whichever session auto-selected consumed, so every other tab
   started with one visible frame. Now tracked per session.

4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
   per notch) scrolled one line where Chrome scrolls four or five, and capped
   the forwarded SGR report at one tick. Line and page deltas are now converted,
   and a pure horizontal swipe no longer falls through to a phantom -1.

Analysis and measurements: docs/scrollback-issues-analysis.md
2026-08-07 04:06:54 +02:00
Claudia ebfcac6ad1 fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.

When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.

- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
  is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
  client redelivers.
- `Session.write()` reports whether it reached a PTY.

Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.

What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.

9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:36:33 +02:00
Claudia 2e69e28e71 fix(mux): bound the process-tree walk — it can take a machine down
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.

Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.

- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
  with a single-flight guard. Async matters: under the same procfs pathology,
  `execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
  which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
  visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
  both reporting when they truncate. Pure so the regression tests can exercise the
  shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
  re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
  would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
  bounded, so a wedged `ps` cannot stop killSession from reaching its
  process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
  fresh; a truncated table would make whole subtrees invisible to the kill path.

13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:34:47 +02:00
Claudia 1a32e63765 fix(http): raw writeHead routes lost every header the security hook set
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.

The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.

Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.

Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:31:50 +02:00
Codeman maintainer d41f28bc14 docs(docker): warn that a plain agent-image rebuild keeps stale CLIs
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.

That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.

Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.

No changeset: docs-only, rides the next release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 09:02:28 +02:00
Codeman maintainer 322f21ef9f docs(extending): scope the "no sandbox" claim, point at Docker cases
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:

- Integration code cannot be sandboxed by Codeman because Codeman never
  launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
  documented isolation story.

Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:46:22 +02:00