Compare commits

...
46 Commits
Author SHA1 Message Date
Codeman maintainer 848ab48b0a chore: version packages (1.33.2)
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:35:45 +02:00
Codeman maintainer 0b106b03eb chore: add the cron paste-mode fix to the landing changeset
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:25:00 +02:00
Codeman maintainer 54c591c84d Merge origin/master (cron paste-mode Enter fix) into the landing branch 2026-09-28 17:24:52 +02:00
Codeman maintainer 2eece4f8f9 fix(cron): send a paste-mode prompt's Enter as its own write
A cron job in "Paste (direct)" input mode wrote `<text>\r` into the pane
in one piece. Claude Code (measured on 2.1.283) takes a burst of about a
hundred characters as a paste, so the `\r` landed as a newline and the
prompt sat unsent on the composer while the run reported `prompt_sent`.

Delivery now lives in `deliverCronPrompt()`. Paste mode writes the text
raw, waits CRON_PASTE_ENTER_DELAY_MS (300 ms), sends `\r` as a separate
write down the same PTY (so it cannot overtake the text), and arms the
session's composer check through the new public
`Session.verifySubmitted()`, which re-presses Enter while the prompt is
still visibly unsent. A session with nothing to write to now fails the
run instead of reporting the prompt as sent. Typed mode is unchanged.

Verified on an isolated instance: a paste-mode job with a 104-character
prompt submitted on the first Enter and Claude answered.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 17:10:12 +02:00
Codeman maintainer 4d165d3fb1 chore: add the input-delivery fix to the landing changeset
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:46:04 +02:00
Codeman maintainer d5ffc22f4a Merge the input-delivery fix from master
fix(input): deliver API prompts through tmux so their Enter is not lost
2026-09-28 16:45:27 +02:00
Codeman maintainer c2dfc775a3 fix(input): deliver API prompts through tmux so their Enter is not lost
A prompt posted to /api/sessions/:id/input without `useMux` was written
into the pane in one piece. Claude Code (measured on 2.1.283) takes a
`<text>\r` burst of about a hundred characters or more as a paste, so the
trailing `\r` landed as a newline in the composer and the prompt sat there
unsent while the route answered 200. A later raw `\r` did not recover it;
a tmux `send-keys Enter` did. Short prompts submitted, which is why it
looked random. The same stranding was seen with Codex and OpenCode.

A plain prompt (printable text plus exactly one trailing `\r`, detected by
`isPlainPromptInput()`) now goes through `writeViaMux` even without
`useMux`: the text is typed, Enter is pressed as its own key, and the
SubmitVerifier re-presses it while the prompt is still on the composer.
The write is awaited, since the browser's POST fallback sends frames one
at a time and a following keystroke must not overtake the Enter. Raw
frames (escape sequences, bracketed paste, a line feed, a bare `\r`) and
an explicit `useMux: false` keep the direct write.

Verified on an isolated instance: the 239- and 104-character prompts that
stranded (at +1 s, at +50 s on ultracode, and on a warm session) all
submitted on the first Enter with no `useMux`. The phone's local-echo
flow (a burst, then its `\r` as a separate write) was measured unaffected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:44:59 +02:00
Codeman maintainer e439cf0ef3 chore: changeset for the 2026-09-28 landing
Folds the #490 and #492 contributor changesets (the latter said minor) into one patch changeset with the Thanks section, one paragraph per change and the fixes applied while landing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:30:56 +02:00
Codeman maintainer 1f4c390e12 fix(terminal): merge-time fixes for #498
- _logScrollRouting() reports cliMouseTracking, the gate's new input, in both
  the de-dup signature and the console line (xterm's own mouseTracking stays
  'none' for Claude, so it gave no reason for a no).
- Restore two guard tests the new gate made vacuous: the local-scrollback
  opt-out footgun test and the codex/gemini "no version rescues it" fixtures
  now set cliMouseTracking: true, so removing the opt-out or re-adding codex to
  the gate fails again.
- Update the comments and architecture-invariants lines that still described
  the version-only rule (wheel handler header, gate doc, the false paths of
  _maybePageCliTranscript, "holds a tracking mode on continuously").
- Name both fullscreen switches (CLAUDE_CODE_NO_FLICKER=1 and "tui":
  "fullscreen" in ~/.claude/settings.json) in the code comment, the invariants
  and the two wiki pages.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:30:40 +02:00
Codeman maintainer 714050fe8a fix(terminal): merge-time fixes for #494
- Skip and latch a bounded Shell window once the browser is at xterm's
  scrollback cap (scrollback + rows): a 1 MiB window of short lines can carry
  more rows than the browser can ever hold, so it replayed and re-captured on
  every scroll-to-top with no 60 s back-off.
- Label a replayed bounded window 'tail' even when the capture was byte-capped,
  so the banner keeps offering Load full history instead of calling the rest
  unrecoverable.
- Pin GET /terminal?full=1&tail=<n> in the route tests: full-history source,
  truncationReason 'tail', and the closing relative cursor move survive the cut.
- Log the bounded skip via _logScrollRouting('repull-skipped-bounded').

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:30:40 +02:00
Codeman maintainer dfd3df8289 fix(build): merge-time fixes for #500
- pre-push hook: skip with a notice when npm is not on PATH (GUI git
  clients and IDEs often run hooks with a minimal PATH), instead of
  blocking every push on "npm: not found"; real-push test with a
  stripped PATH
- test/git-hooks.test.ts: pin GIT_CONFIG_NOSYSTEM=1 and
  GIT_CONFIG_GLOBAL=/dev/null around the resolveGitHooksDir tests, so
  an exported global or a system core.hooksPath no longer fails them
- watch tsconfig.json, .prettierignore and .editorconfig too:
  typecheck and format:check read them
- check:browser-excludes: fail loudly when the vitest list output and
  the walked test/**/*.test.ts tree share no path (format drift would
  otherwise pass vacuously)
- Reword the PRE_PUSH_MARKER comment: bumping its version would make every
  installed v1 hook read as foreign and never refresh again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:56 +02:00
Codeman maintainer a0fbd1d28d docs(registry): note the cliMouseTracking half of claude's wheel rule (#498)
- claude's declared-for-later wheelForward says the live rule in
  _shouldForwardWheelToApp is the version AND the server-published
  cliMouseTracking flag, so whoever wires the field up needs both

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:29 +02:00
Codeman maintainer 272b56d47b fix(session): merge-time fixes for #491
- claude watchingLine: the lookahead keys on "Artifact" alone, so a
  footer truncated mid-chip ("1 Artifact…", "1 Artifact comm…") is still
  refused instead of reporting the shell beside it; comment follows
- test: both truncations return no watching label
- invariants: a chip that waits on a human never counts as watching, and
  the ^ anchor is what stops the retry past the chip

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:29 +02:00
Codeman maintainer 1645ef5f5c fix(docker): merge-time fixes for #492
- test: the complete-identity case now checks the combined
  agentImageBuildArgPairs() argv on both producers, so the manual
  build-agent-image.mjs path cannot drop the identity unnoticed
- both producers: GIT_IDENTITY_BUILD_ARGS carries the mirror/parity
  warning its gh/az neighbour has
- the partial-identity error names CODEMAN_AGENT_IMAGE_GIT_USER_NAME and
  CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL; test regex follows
- wiki Docker-Cases: mention the identity variables next to the gh/az
  switches

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:29 +02:00
Codeman maintainer 627b76739c fix(docker): merge-time fixes for #490
- test: every ENV PATH= line in server.Dockerfile must start $PATH:, and
  the ~/.local/bin append is pinned alongside /opt/codeman-cli/bin
- invariants + CLAUDE.md: the append-only PATH rule names ~/.local/bin too
- docker-compose.md: Settings-installed CLIs live in ~/.local on the
  app-data mount; reinstall once after upgrading; hand-run npm installs
  need --prefix ~/.local
- installEnv() JSDoc describes the in-container npm prefix redirect

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:28:29 +02:00
Codeman maintainer 83e39c40a1 Merge pull request #504 from Ark0N/fix/phone-tab-strip
fix(mobile): make the phone header tab strip read as live tabs
2026-09-28 16:21:10 +02:00
Codeman maintainer fec0409315 Merge pull request #500 from aakhter/pr/prepush-browser-excludes
build: add a browser-test exclusion check and a pre-push static-check hook
2026-09-28 16:21:09 +02:00
Codeman maintainer b4954c14cd Merge pull request #494 from timkjr/fix/shell-scroll-history
fix(terminal): let a Shell pane's scroll-up reach tmux history

# Conflicts:
#	docs/wiki/The-Dashboard.md
2026-09-28 16:21:08 +02:00
Codeman maintainer c9f47b095a Merge pull request #498 from JDProfresh/fix/claude-inline-scroll
fix(terminal): only forward scroll to Claude while it tracks the mouse
2026-09-28 16:20:57 +02:00
Codeman maintainer 47ac16d6ab Merge pull request #492 from opticon454/feature/static-git-identity
feat(docker): configure static git identity
2026-09-28 16:20:56 +02:00
Codeman maintainer 6d147c1bf1 Merge pull request #490 from opticon454/feature/docker-uv-uvx
fix(docker): keep CLIs installed from Settings across container updates
2026-09-28 16:20:54 +02:00
Codeman maintainer 92921b9107 Merge pull request #491 from irisitymichaelgrundberg/fix/artifact-comment-monitor-needs-you
fix(session): alert for an agent waiting on artifact comments

# Conflicts:
#	src/config/cli-registry/stock.ts
2026-09-28 16:20:52 +02:00
Codeman maintainer 614c7e6cd5 Merge pull request #501 from aakhter/pr/webview-sse-owner
fix(webview): route webview:changed only to its owner in multi-user mode
2026-09-28 16:20:35 +02:00
Codeman maintainer 7659ca8b44 fix(session): keep a tab working while Claude waits for its own workers
When Claude hands work to an ultracode workflow or background agents, it
ends its own turn and closes it with `✻ Waiting for 1 dynamic workflow to
finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when the
workers report back. The pane sits quiet with the composer up, so the idle
probe called the session idle for the whole wait. At phone width the
workflow's progress row also drops its ticking timer, so nothing on screen
changes for minutes.

A new optional registry field, `capabilities.workDetect.awaitingLine`,
names that closing row, and `_probePaneWorking()` counts it as work.
Claude renders the row once from a snapshot and never redraws it, so the
same words stay on screen after the workers finish. `isAwaitingWorkers()`
therefore tests only the newest column-0 row directly above the composer,
never the whole pane and never the PTY stream; a follow-up turn always
puts rows of its own there. The column-0 anchor also keeps an agent from
holding its own tab busy by printing the sentence.

Verified against the live Mac mini pane that reported the bug (2.1.283),
and end to end on an isolated instance: an ultracode session running a
90 s workflow at 46 columns stayed busy through the wait and the
follow-up turn, then went idle 6 s after that turn closed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 16:18:27 +02:00
Codeman maintainer 61037082d1 fix(mobile): make the phone header tab strip read as live tabs
On a phone every inactive tab rendered transparent: grey 11px text
floating in unmarked gaps, a boxed Alt+N digit in each tab (a phone has
no Alt key), names capped at 50px so a shared `w1-` prefix was most of
what showed, and the tab that did not fit was chopped mid-word against
the connection dot. The strip looked like a row of disabled labels.

Phone block of mobile.css only:
- Every header tab is a chip, filled and bordered from the skin's
  --control-* tokens, name in --text at weight 500. Written
  `:where(.header) .session-tab` so it stays at (0,1,0): the per-colour
  left border still wins, and sidebar layout (where the list leaves the
  header) is untouched.
- The Alt+N digit is hidden in the header; inactive tabs drop their
  empty .tab-actions container, which padded the chip's right side.
- Name cap 50px -> 80px, status dot 4px -> 6px, strip gap 2px -> 6px.
- Scroll-driven edge fade: a mask on the strip whose widths follow its
  own inline scroll timeline (registered @property lengths), so the
  clipped tab dissolves into the edge. No JS; a strip that does not
  overflow gets no mask, and browsers without scroll timelines keep the
  old hard edge.

The tap-zone arithmetic comment is updated for the numberless phone
tabs and the bigger dot (the required reserve drops from 38px to 36px;
the 44px min-width stays). test/mobile-tab-strip-chips.test.ts pins the
(0,1,0) selector, the top-level @property registration and the
timeline-after-shorthand order, each of which fails silently otherwise.
test/mobile/tabs.test.ts follows the new name cap.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-28 15:23:22 +02:00
Aamer Akhter e71971cab4 fix(webview): route webview:changed only to its owner in multi-user mode
webview:changed carried only {action, id} and the SSE routing hint had no
webview: branch, so every connected client received it: in multi-user mode
any user saw the ids of other users' web-tab creates, edits and deletes.
The event now carries the web tab's owner (from the stored record) and is
routed to that owner plus admins. Single-user delivery is unchanged.
2026-09-26 22:45:35 -04:00
Aamer Akhter e60b5a8a2c build: address review on the pre-push hook and hooks-dir resolution
resolveGitHooksDir now returns a directory only when it is the repo's own
<git-common-dir>/hooks (compared on canonical paths), so a core.hooksPath
elsewhere, global or repo-local, is never written to by postinstall, while a
core.hooksPath pointing back at the repo's own .git/hooks still resolves.

The pre-push hook skips with a one-line notice when a pushed ref is not the
checked-out HEAD (tags peeled) or when git status shows uncommitted or
untracked changes under a path the checks read (src, config, scripts, test,
package.json, package-lock.json, install.sh), since the checks read the
working tree rather than the pushed commit.

Also: honest timing (~10-40s instead of ~15s), CLAUDE.md Session Safety note
on CODEMAN_SKIP_PREPUSH for another session's WIP, 14 (not 9) Playwright
tests, and a note that the browser-excludes check only sees direct imports.
2026-09-26 22:44:14 -04:00
Aamer Akhter 1d85909a06 build: add a browser-test exclusion check and a pre-push static-check hook
npm run check:browser-excludes finds tests that import a browser driver and
asks `vitest list` whether the CI config still collects them; wired into CI.
npm install now also installs a marker-owned pre-push hook that runs the
static CI checks (~15s). Skip with CODEMAN_SKIP_PREPUSH=1; hand-written
hooks are left alone.
2026-09-26 18:45:17 -04:00
JD 1da2fa2529 fix(terminal): only forward scroll to Claude while it tracks the mouse
Claude 2.1.280 renders inline by default: no alt screen, no mouse tracking, transcript in real scrollback. The version-only gate still sent every wheel tick and touch swipe as SGR reports, which Claude ignores, so scrolling a Claude session was dead while codex (routed locally) worked. Gate forwarding on the server-recorded cliMouseTracking flag, which fullscreen mode (CLAUDE_CODE_NO_FLICKER=1) sets.
2026-09-26 15:49:58 -04:00
Codeman maintainer 45ea2e1d32 docs(readme): ask readers to star the project
Adds a centered star call-to-action under the badge row in both the
English and Simplified Chinese READMEs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-26 04:39:13 +02:00
timkjrandClaude Sonnet 5 f6aa50239f fix(terminal): skip a bounded Shell window before the downgrade guard
A window cut at the tail size can be smaller than the browser's buffer
while tmux still holds more. The downgrade guard reads that as "tmux has
nothing more to give", which is true of an unbounded capture only, so a
bounded window reaching it marked the session exhausted and removed Load
full history from the banner.

The bounded skip now runs first, so such a window never reaches the
exhausted path, and it no longer writes banner state: relabelling it from
the bounded payload would call a terminal holding all of a Load full
history pull "the most recent 1 MiB".

A skipped window that came back truncated cannot reach anything older
than the browser shows, and every ask costs the server a synchronous
capture-pane of the whole history (tail is applied after the capture), so
it puts the session on the 60 s cooldown. An untruncated one keeps 4 s.

_replayWouldShrinkBuffer takes optional pre-estimated rows so a megabyte
capture is not scanned twice. CLAUDE.md's Full-scrollback replay entry no
longer says Shell never pulls on ordinary scroll.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-25 20:17:22 -05:00
timkjrandClaude Opus 5.5 9676e90133 fix(terminal): let a Shell pane's scroll-up reach tmux history
A burst of output leaves a Shell pane with about one screen of browser
scrollback, because tmux repaints the burst instead of scrolling it,
while tmux itself keeps every line. Shell declined the scroll-to-top
re-pull other modes use, and the Load full history button renders only
once a replay was truncated, so a Shell tab under 1 MiB could not
scroll back at all.

The scroll gesture now pulls ?full=1&tail=TERMINAL_TAIL_SIZE, the same
bound a tab switch loads; the route's existing tail cut marks longer
histories 'tail', so the banner still offers the unbounded pull. A
window no longer than the browser's buffer is not rewritten.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 19:05:16 -05:00
Devvyn bf73a84732 fix(docker): address git identity review 2026-09-25 22:27:26 +08:00
Devvyn 8d358aaa26 feat(docker): configure static git identity 2026-09-25 22:25:11 +08:00
Michael GrundbergandClaude Opus 5.5 a9b48320a3 fix(session): alert for an agent waiting on artifact comments
An agent that publishes an artifact arms a monitor for its comments and
ends its turn. Claude Code shows that on the footer as `1 Artifact
comment monitor`, and #473 put that chip on the list of background work,
so the session counted as watching and its idle prompt opened already
acknowledged. Unlike every other chip on the list, that monitor waits
on the user: the agent hears nothing until somebody comments.

Claude's `watchingLine` now refuses any footer that carries the chip,
through a lookahead over the whole row, so a shell running beside the
monitor cannot report the session as watching either.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-25 07:50:25 +02:00
DevvynandClaude Sonnet 5 95a3b87062 chore: drop changesets already released in 1.33.1
The pnpm and uv/uvx changesets describe work upstream shipped in 1.33.1
(#485, #487), so keeping them would repeat those notes in the next release.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-25 11:14:47 +08:00
DevvynandClaude Sonnet 5 8cef31086b fix(docker): persist CLIs installed from Settings across container updates
The image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is image content, so
Update-Codeman.sh discarded every npm-installed CLI (dsh, pi). In the Compose
container, POST /api/clis/:id/install now installs into ~/.local on the
persistent home mount, and ~/.local/bin is appended to the image PATH.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-25 08:56:20 +08:00
Devvyn 55790964b7 Merge remote-tracking branch 'upstream/master' into feature/docker-uv-uvx 2026-09-25 08:30:56 +08:00
Codeman maintainer 5ae574374f docs(changelog): add the Thanks section to 1.33.1
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-09-24 23:47:08 +02:00
Devvyn d6c3386102 Revert "feat(docker): add sudo to the agent image"
This reverts commit e98127a804.
2026-09-24 22:19:20 +08:00
Devvyn 10c263a5b8 Revert "feat(docker): install sudo with passwordless access for the agent user"
This reverts commit b070c9ee65.
2026-09-24 22:19:13 +08:00
DevvynandClaude Sonnet 5 b070c9ee65 feat(docker): install sudo with passwordless access for the agent user
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-24 22:17:38 +08:00
DevvynandClaude Sonnet 5 e98127a804 feat(docker): add sudo to the agent image
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-24 22:17:22 +08:00
DevvynandClaude Sonnet 5 3e3a4612e6 feat(docker): install libsecret-1-0 for the Azure DevOps MCP
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-24 22:03:24 +08:00
DevvynandClaude Sonnet 5 a5283c565d feat(docker): install uv and uvx in server and agent images
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-24 21:04:18 +08:00
DevvynandClaude Sonnet 5 b46588f247 fix(docker): install pnpm in the Compose server image
`dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run
menu's "DeepSeek - add a terminal profile" button failed with
`dsh: pnpm not found on PATH` (exit 127) on the server image. The agent
image already installs pnpm for the same reason (#352). Pin pnpm@12.6.0
in the runtime-writable CLI prefix and note it in the DeepSeek doc.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
2026-09-24 21:04:06 +08:00
65 changed files with 3178 additions and 119 deletions
+1 -1
View File
@@ -10,7 +10,7 @@
"name": "codeman", "name": "codeman",
"source": "./plugins/codeman", "source": "./plugins/codeman",
"description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", "description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.33.1", "version": "1.33.2",
"author": { "author": {
"name": "Ark0N", "name": "Ark0N",
"url": "https://github.com/Ark0N" "url": "https://github.com/Ark0N"
+5 -2
View File
@@ -28,12 +28,15 @@ The frontend is plain JS served from `src/web/public/` with no bundler in dev: e
CI runs all of these, so save yourself a round trip: CI runs all of these, so save yourself a round trip:
```bash ```bash
npm run typecheck # tsc --noEmit, strict mode npm run typecheck # tsc --noEmit, strict mode
npm run lint npm run lint
npm run format:check npm run format:check
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
npm run check:browser-excludes # every browser-driven test is kept out of `npm test`
``` ```
`npm install` also installs a `pre-push` git hook that runs these static checks (about 10-40s, machine-dependent) and blocks the push if one fails. It skips itself when you push something other than the checked-out HEAD, or when the tree has uncommitted changes the checks would read. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; it never replaces a `pre-push` hook of your own.
### Tests ### Tests
```bash ```bash
+7
View File
@@ -34,6 +34,13 @@ jobs:
- name: Frontend JS syntax check - name: Frontend JS syntax check
run: npm run check:frontend-syntax run: npm run check:frontend-syntax
# Asks `vitest list` what CI would actually collect, rather than matching
# filenames: a browser-driven test missing from BROWSER_TEST_GLOBS
# (config/test-suites.ts) passes locally and dies in the test job with
# "browserType.launch: Executable doesn't exist".
- name: Browser-test exclusion check
run: npm run check:browser-excludes
- name: Format check - name: Format check
run: npm run format:check run: npm run format:check
+38 -1
View File
@@ -1,10 +1,47 @@
# aicodeman # aicodeman
## 1.33.2
### Patch Changes
- e439cf0: ### Thanks
- @aakhter for keeping web-tab events private to their owner in multi-user mode (#501), with end-to-end isolation tests that fail without the fix, and for the browser-test exclusion check and pre-push hook (#500), including the hooks-dir resolution that never writes outside the repo's own `.git/hooks`.
- @opticon454 for keeping CLIs installed from Settings across Docker container updates (#490) and for the static Git identity for the Docker images (#492).
- @timkjr for letting a Shell pane's scroll-to-top reach tmux history (#494), with tests that fail on the commit before each fix.
- @JDProfresh for tracking down why wheel and touch scrolling did nothing in Claude's default inline view (#498), with the tmux measurements that proved it.
- @irisitymichaelgrundberg for the follow-up that makes an agent waiting on artifact comments raise its alert again (#491).
**A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends.
**Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. The same fix reaches cron jobs in "Paste (direct)" input mode, which reported `prompt_sent` for a prompt that never left the composer: the text is written raw, Enter follows as its own write 300 ms later, and the session presses it again while the prompt is still unsent. A cron run with no session to write to now fails instead of reporting the prompt as sent.
**Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback.
**Shell panes scroll back into tmux history (#494).** Scrolling to the top of a Shell pane now pulls the most recent 1 MiB of its tmux history, so output that arrived in a burst is reachable without pressing **Load full history**, which still loads the rest.
**An agent waiting on artifact comments alerts again (#491).** A session whose agent published an artifact and is waiting for somebody to comment on it now raises the normal idle alert and lands in NEEDS YOU, instead of being treated as busy with background work.
**Web-tab changes stay private in multi-user mode (#501).** The `webview:changed` event reached every connected user, exposing the ids of other users' web-tab creates, edits and deletes. It now carries the tab's owner and reaches that owner plus admins only. Single-user mode is unchanged apart from a new optional `owner` field on the event.
**Phone header tabs look like tabs (#504).** On phones every header tab is now a chip with a fill and a border, the Alt+N digit (a keyboard hint a phone cannot use) is hidden, names get 80px instead of 50px, and the strip fades at whichever edge still has tabs scrolled out of view.
**Docker: CLIs installed from Settings survive container updates (#490).** On the Compose deployment, CLIs installed from App Settings (DeepSeek, Pi and other npm-based CLIs) now go to `~/.local` on the persistent home mount, and `~/.local/bin` is on the image PATH, so recreating the container no longer discards them. Anything installed from Settings before this release has to be installed once more after the rebuild.
**Docker: a static Git identity for the server and agent images (#492).** Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` (or `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` / `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` on a bare-host install) and the identity is written to `/etc/gitconfig` in the server image and the Docker-case agent image. A half-set pair is refused on every build path. Both Docker changes edit `server.Dockerfile`, so the in-app updater asks Compose deployments to rebuild with `docker/Start-Codeman.sh` instead of updating in place.
**Contributor tooling (#500).** `npm run check:browser-excludes`, now a CI step, fails when a test that drives a real browser is still collected by `npm test`. `npm install` also installs a pre-push hook that runs the static CI checks before a push; it steps aside when the pushed ref is not HEAD or the tree has uncommitted changes the checks would read, and `CODEMAN_SKIP_PREPUSH=1 git push` skips it once.
**Fixes applied while landing.** A Shell pane's scroll-to-top (#494) no longer re-pulls the same window on every gesture once the browser's 50,000-row scrollback is full, and a pull that hit the byte cap no longer claims the older history is gone. The scroll-routing diagnostics (#498) now log whether Claude has mouse tracking on. The artifact-comment check (#491) also refuses a footer cut off in the middle of the chip. The pre-push hook (#500) steps aside when `npm` is not on PATH, as in some GUI git clients, instead of blocking every push. The Docker Git identity error (#492) names the two variables to set. New tests pin the image PATH order, the identity on both agent-image build paths, and the `?full=1&tail=` terminal route.
## 1.33.1 ## 1.33.1
### Patch Changes ### Patch Changes
- **Finished sessions close themselves (#486).** A session whose agent you ended with `/exit` is now closed the same way the X button closes it, so finished sessions stop piling up on the board; the conversation stays resumable from the Resume list and the lifecycle log records "agent exited cleanly (status 0)". Only an explicit exit status 0 with no signal, confirmed by two pane reads, qualifies: a crashed or OOM-killed agent keeps its row with the exit code on the tab. The phone overview and desktop home rail now say `exited` instead of `idle`, reboot restore no longer offers to rebuild a session whose agent had exited, and closing one session no longer deletes the `.claude-images` directory that a sibling session in the same case still uses. Thanks @irisitymichaelgrundberg. - ### Thanks
- @irisitymichaelgrundberg for closing sessions whose agent exited cleanly (#486), built carefully around every way a pane exit can lie (a SIGKILL with no status, a single misread), with the `.claude-images` guard split into its own commit as asked.
- @opticon454 for the live-refreshing case picker and Manage search (#483), and for the uv/uvx, libsecret and pnpm additions to the Docker images (#487, #485).
**Finished sessions close themselves (#486).** A session whose agent you ended with `/exit` is now closed the same way the X button closes it, so finished sessions stop piling up on the board; the conversation stays resumable from the Resume list and the lifecycle log records "agent exited cleanly (status 0)". Only an explicit exit status 0 with no signal, confirmed by two pane reads, qualifies: a crashed or OOM-killed agent keeps its row with the exit code on the tab. The phone overview and desktop home rail now say `exited` instead of `idle`, reboot restore no longer offers to rebuild a session whose agent had exited, and closing one session no longer deletes the `.claude-images` directory that a sibling session in the same case still uses. Thanks @irisitymichaelgrundberg.
**Search in the phone Select Case sheet (#488).** The bottom sheet gains a "Search cases" field that filters by name (every word must match, any order, ignoring case), Enter picks the case when exactly one row is left, and Escape clears then closes. Also fixes a dead band under Create New Case and a list shorter than the sheet could show. **Search in the phone Select Case sheet (#488).** The bottom sheet gains a "Search cases" field that filters by name (every word must match, any order, ignoring case), Enter picks the case when exactly one row is left, and Escape clears then closes. Also fixes a dead band under Create New Case and a list shorter than the sheet could show.
+12 -7
View File
@@ -34,6 +34,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
- To land a commit on master **without** switching branches (which would yank the tree out from under the other session): `git push origin HEAD:master` then `git branch -f master HEAD`. Never `git checkout master` to "fix" it. - To land a commit on master **without** switching branches (which would yank the tree out from under the other session): `git push origin HEAD:master` then `git branch -f master HEAD`. Never `git checkout master` to "fix" it.
- **Never `git add -A`/`git add .`** — stage explicit paths. A sweep will pick up another session's WIP. - **Never `git add -A`/`git add .`** — stage explicit paths. A sweep will pick up another session's WIP.
- Another session's broken WIP can block `npm run build`, since `tsc` is the first step and the build gates on it. That is not your bug to fix. ⚠️ `tsc` still EMITS on type errors, so a failed `npm run build` leaves a rebuilt `dist/index.js` compiled from their tree; check what it pulled in before restarting the service. To deploy frontend-only changes past a blocked `tsc`, run the asset stage of `scripts/build.mjs` (everything after the `tsc`/`chmod` lines is independent of it). - Another session's broken WIP can block `npm run build`, since `tsc` is the first step and the build gates on it. That is not your bug to fix. ⚠️ `tsc` still EMITS on type errors, so a failed `npm run build` leaves a rebuilt `dist/index.js` compiled from their tree; check what it pulled in before restarting the service. To deploy frontend-only changes past a blocked `tsc`, run the asset stage of `scripts/build.mjs` (everything after the `tsc`/`chmod` lines is independent of it).
- **A pre-push failure in a file you did not touch is another session's WIP.** Push with `CODEMAN_SKIP_PREPUSH=1 git push` and leave it alone. (The hook already skips itself when the tree has uncommitted changes in a path it checks, so this mostly happens once the other session has committed.)
## CRITICAL: Always Test Before Deploying ## CRITICAL: Always Test Before Deploying
@@ -77,7 +78,7 @@ When user says "COM":
CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed. CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed.
**Version**: 1.33.1 (must match `package.json`) **Version**: 1.33.2 (must match `package.json`)
## Project Overview ## Project Overview
@@ -112,6 +113,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) | | Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) |
| Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) | | Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) |
| Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) | | Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) |
| Browser-test exclusion check | `npm run check:browser-excludes` (`scripts/check-browser-test-excludes.mjs`; runs in CI, <1s). Fails if a test importing playwright/puppeteer is still collected by `config/vitest.ci.config.ts`; add it to `BROWSER_TEST_GLOBS` in `config/test-suites.ts` |
| Pre-push hook | Installed by `npm install` (`scripts/git-hooks.mjs`, via postinstall): runs the static CI checks (~10-40s) before `git push`. Skip once: `CODEMAN_SKIP_PREPUSH=1 git push`. Skips itself with a notice when HEAD is not the pushed commit or the tree has uncommitted changes the checks would read. Marker-owned, so a hand-written `pre-push` is never overwritten; installs ONLY into the repo's own `<git-common-dir>/hooks` (worktree-safe; a `core.hooksPath` elsewhere, e.g. a global one, is left alone) |
| Excluded-suite runners | `npm run test:browser` · `npm run test:mobile` · `npm run test:perf` · `npm run test:all` (everything, environmental failures included) — see Testing | | Excluded-suite runners | `npm run test:browser` · `npm run test:mobile` · `npm run test:perf` · `npm run test:all` (everything, environmental failures included) — see Testing |
| Production start | `npm run start` | | Production start | `npm run start` |
| Production logs | `journalctl --user -u codeman-web -f` | | Production logs | `journalctl --user -u codeman-web -f` |
@@ -120,7 +123,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
| Dependency doctor | `codeman doctor` (alias `check-deps`; `--json`, `--category core\|office\|other`). Probes Node/Claude CLI/tmux/LibreOffice/MS Office against `config/dependency-registry.ts`; engine is pure given an injectable `ProbeHost` | | Dependency doctor | `codeman doctor` (alias `check-deps`; `--json`, `--category core\|office\|other`). Probes Node/Claude CLI/tmux/LibreOffice/MS Office against `config/dependency-registry.ts`; engine is pure given an injectable `ProbeHost` |
| Multi-user accounts | `codeman users add <name>` / `passwd <name>` / `list` / `rm <name>` (writes `~/.codeman/users.json`, mode 0600; see Multi-user mode) | | Multi-user accounts | `codeman users add <name>` / `passwd <name>` / `list` / `rm <name>` (writes `~/.codeman/users.json`, mode 0600; see Multi-user mode) |
**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 9 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`). **CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `check:browser-excludes`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 14 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`).
**Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`) — config lives in the **`"prettier"` key of `package.json`**, not a `.prettierrc` (keeps the repo root short; editors read it natively). `.prettierignore` stays at the root because Prettier resolves it relative to cwd. ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`. **Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`) — config lives in the **`"prettier"` key of `package.json`**, not a `.prettierrc` (keeps the repo root short; editors read it natively). `.prettierignore` stays at the root because Prettier resolves it relative to cwd. ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`.
@@ -128,7 +131,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
## Common Gotchas ## Common Gotchas
- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`. ⚠️ **Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints** while still taking the typed text (measured 2026-09-19: an Enter at 28 s stranded the prompt, one at 51 s submitted it), so text+`\r` sent at readiness sits unsent with `0 tokens` and a `wait` burns its timeout. So the SERVER verifies every programmatic write that carried a `\r`: `SubmitVerifier` (`session-submit-verifier.ts`, armed from `writeViaMux`) reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the LAST composer line (the CLI's own `promptGlyph`) verifiably still holds the head of what was sent; an empty composer, other text, or no composer line at all (a shell, a direct-PTY session) ends it, and a newer write replaces the schedule. The skill's `sendwait` keeps its own copy of the loop (`_composer_text` in `skills/codeman/preamble.sh`) for servers that predate this. The `shift+tab` footer only means the composer painted, never that Enter is accepted - **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`. ⚠️ **Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints** while still taking the typed text (measured 2026-09-19: an Enter at 28 s stranded the prompt, one at 51 s submitted it), so text+`\r` sent at readiness sits unsent with `0 tokens` and a `wait` burns its timeout. So the SERVER verifies every programmatic write that carried a `\r`: `SubmitVerifier` (`session-submit-verifier.ts`, armed from `writeViaMux`) reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the LAST composer line (the CLI's own `promptGlyph`) verifiably still holds the head of what was sent; an empty composer, other text, or no composer line at all (a shell, a direct-PTY session) ends it, and a newer write replaces the schedule. The skill's `sendwait` keeps its own copy of the loop (`_composer_text` in `skills/codeman/preamble.sh`) for servers that predate this. The `shift+tab` footer only means the composer painted, never that Enter is accepted. ⚠️ **A prompt must never be written into the pane as ONE burst**: Claude Code 2.1.283 takes a `<text>\r` burst of ~100+ chars as a paste, its `\r` lands as a NEWLINE and the prompt strands (a later raw `\r` does not recover it, a tmux `send-keys Enter` does). So `POST .../input` routes a plain prompt (`isPlainPromptInput()`, route-helpers.ts: printable text + exactly one trailing `\r`) through `writeViaMux` even without `useMux`, AWAITED so the browser's serialized POST fallback keeps frame order; raw frames and an explicit `useMux:false` keep the direct write
- **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks - **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks
- **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`) - **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`)
- **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically) - **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically)
@@ -203,6 +206,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer all through a turn, and its working line (`✻ Actualizing… (13m 23s · …)`) is invisible to `SPINNER_PATTERN`, keyword lists and the raw stream. `_confirmIdle()` (session.ts) requires the pane to go quiet AND the SCREEN (`capturePaneText()` + the working-line pattern) to agree; a sustained run of repaints (`session-activity.ts`) marks a turn as started. ⚠️ The composer glyph and working line are per-CLI registry DATA (`capabilities.workDetect`), never Claude constants; a CLI declaring neither falls back to Claude's pair. ⚠️ `workingLine` is config-supplied and runs on the PTY hot path, so it must compile through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()` (ReDoS guard; null, not throw). → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line) ⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer all through a turn, and its working line (`✻ Actualizing… (13m 23s · …)`) is invisible to `SPINNER_PATTERN`, keyword lists and the raw stream. `_confirmIdle()` (session.ts) requires the pane to go quiet AND the SCREEN (`capturePaneText()` + the working-line pattern) to agree; a sustained run of repaints (`session-activity.ts`) marks a turn as started. ⚠️ The composer glyph and working line are per-CLI registry DATA (`capabilities.workDetect`), never Claude constants; a CLI declaring neither falls back to Claude's pair. ⚠️ `workingLine` is config-supplied and runs on the PTY hot path, so it must compile through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()` (ReDoS guard; null, not throw). → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line)
⚠️ **A turn that ENDED waiting for its own workers is working, not idle.** Claude closes such a turn with `✻ Waiting for 1 dynamic workflow to finish` (background agents / ultracode) and resumes by itself; `capabilities.workDetect.awaitingLine` makes the idle probe count it as work. ⚠️ Claude never redraws that row, so it stays on screen after the workers finish: test it ONLY as the newest column-0 row above the composer (`isAwaitingWorkers()`), never pane-wide and never on the stream. → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line)
⚠️ **A quiet pane is not always a pane that wants you.** A CLI can declare an optional `capabilities.workDetect.watchingLine` (a monitor, background shell or cloud hand-off it is still running); the idle probe reads it into `Session.watching` and `notePrompt()` opens that idle item ALREADY acknowledged, so no surface alerts. Only `idle` is eligible, and the label is pane-derived and prompt-injectable, so a pattern must anchor on chrome only that CLI draws. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`. ⚠️ **A quiet pane is not always a pane that wants you.** A CLI can declare an optional `capabilities.workDetect.watchingLine` (a monitor, background shell or cloud hand-off it is still running); the idle probe reads it into `Session.watching` and `notePrompt()` opens that idle item ALREADY acknowledged, so no surface alerts. Only `idle` is eligible, and the label is pane-derived and prompt-injectable, so a pattern must anchor on chrome only that CLI draws. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`.
**An exited agent in a live pane** (`paneExit`, #446): panes use `remain-on-exit on`, so `/exit` leaves a pane, session and pid that look alive; `TmuxManager.startPaneExitWatcher()` publishes `SessionState.paneExit` via `session:updated`. ⚠️ Never set `status: 'error'` or null the `pid` for it; the field is TRI-STATE (absent = UNKNOWN, never alive, scoped by `Session.paneExitApplies`); an absent `#{pane_dead_status}` is not 0; a path that starts a command in a pane must clear the record AND persist. A clean exit is CLOSED via `cleanupSession()` (`pane-exit-sweep.ts`): only an explicit numeric status 0 with no signal, confirmed by 2 reads, with no start/attach in flight (`paneLifecycleInFlight`) and not within 10 s of one (a startup error keeps its row); a crashed agent keeps its row. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit) **An exited agent in a live pane** (`paneExit`, #446): panes use `remain-on-exit on`, so `/exit` leaves a pane, session and pid that look alive; `TmuxManager.startPaneExitWatcher()` publishes `SessionState.paneExit` via `session:updated`. ⚠️ Never set `status: 'error'` or null the `pid` for it; the field is TRI-STATE (absent = UNKNOWN, never alive, scoped by `Session.paneExitApplies`); an absent `#{pane_dead_status}` is not 0; a path that starts a command in a pane must clear the record AND persist. A clean exit is CLOSED via `cleanupSession()` (`pane-exit-sweep.ts`): only an explicit numeric status 0 with no signal, confirmed by 2 reads, with no start/attach in flight (`paneLifecycleInFlight`) and not within 10 s of one (a startup error keeps its row); a crashed agent keeps its row. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit)
@@ -227,9 +232,9 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Docker cases**: a case can point at a **container** running any CLI mode inside it, a **LOCATION OVERLAY on cases, never a `SessionMode`**. One long-lived container **per case**, shared by its sessions: killing a session kills only its in-container tmux, **never** `docker stop` while siblings remain. The workspace is bind-mounted at the **same absolute path**. Credentials are **seeded**, never shared RW. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** A drifted config (label hash) REFUSES the launch. ⚠️ An **adopted** container (`DockerCase.owned === false`) is only `exec`ed into: never create, start, stop, restart, remove, `docker commit` or `docker pause` it; fail closed. Test `owned === false`, never truthiness. ⚠️ Apply `owned` AFTER `dockerConfigHash`. ⚠️ Run modes come from the CONTAINER (`availableModes`), and a failed probe is normal for an OWNED case. ⚠️ Root exec user drops the bypass flag via the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is admin-only in multi-user mode. ⚠️ Loopback prod needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for in-container hooks. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` **Docker cases**: a case can point at a **container** running any CLI mode inside it, a **LOCATION OVERLAY on cases, never a `SessionMode`**. One long-lived container **per case**, shared by its sessions: killing a session kills only its in-container tmux, **never** `docker stop` while siblings remain. The workspace is bind-mounted at the **same absolute path**. Credentials are **seeded**, never shared RW. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** A drifted config (label hash) REFUSES the launch. ⚠️ An **adopted** container (`DockerCase.owned === false`) is only `exec`ed into: never create, start, stop, restart, remove, `docker commit` or `docker pause` it; fail closed. Test `owned === false`, never truthiness. ⚠️ Apply `owned` AFTER `dockerConfigHash`. ⚠️ Run modes come from the CONTAINER (`availableModes`), and a failed probe is normal for an OWNED case. ⚠️ Root exec user drops the bypass flag via the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is admin-only in multi-user mode. ⚠️ Loopback prod needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for in-container hooks. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md`
**Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md` **Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` and `~/.local/bin` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md`
**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `<Mode>Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries. It is WRITTEN only by the opt-in CLI management routes (`cliManagementEnabled`, default OFF; `/api/clis`, `cli-registry-routes.ts`), and only through `mutateRegistryFile()` in `registry-writer.ts`, which serializes mutations and refuses (409) a file that does not parse or has group/world permission bits rather than overwriting it. Importing the registry still writes nothing. → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md` **CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`, `capabilities.workDetect.awaitingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `<Mode>Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries. It is WRITTEN only by the opt-in CLI management routes (`cliManagementEnabled`, default OFF; `/api/clis`, `cli-registry-routes.ts`), and only through `mutateRegistryFile()` in `registry-writer.ts`, which serializes mutations and refuses (409) a file that does not parse or has group/world permission bits rather than overwriting it. Importing the registry still writes nothing. → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md`
**External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token parsing, ❯ readiness; readiness is output stabilization); work detection is per-CLI `capabilities.workDetect` data, not this gate. All eight **require tmux, no direct PTY fallback** (secrets go via socket-scoped `tmux setenv`, never the command line). ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope. ⚠️ **Codex uses predictive write-through echo, never the buffer overlay**: `_predictHookOnData` must never `return` (wire path stays byte-identical), and flushed text and a bracketed paste must go out as separate delayed writes. ⚠️ **Pi**: no bypass flag, never invent one; `approveProjectTrust` executes repo code, so it is in the clamp's **materialize** branch; never wire `--api-key`. ⚠️ **Grok**: `alwaysApprove` is stripped for non-granted owners (only-if-sent). ⚠️ **DeepSeek**: the agent is a PROFILE (Run gates on `isDeepSeekRunnable()`); the permission switch is the `DSH_PERMISSION_MODE` env var, so `clampEnvOverridesForOwner()` must DROP `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for non-granted owners; `hooksAvailableForMode()` is per-SESSION for it (pass `sessionHookOptions(session)`) and is never a stand-in for `mode === 'claude'`; answers come from `deepseek-transcript.ts`, paired by header `cwd` + boot window, never newest-mtime. ⚠️ **OMP**: `OMP_AUTH_BROKER_URL`/`_TOKEN` are clamped the same way. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp) **External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token parsing, ❯ readiness; readiness is output stabilization); work detection is per-CLI `capabilities.workDetect` data, not this gate. All eight **require tmux, no direct PTY fallback** (secrets go via socket-scoped `tmux setenv`, never the command line). ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope. ⚠️ **Codex uses predictive write-through echo, never the buffer overlay**: `_predictHookOnData` must never `return` (wire path stays byte-identical), and flushed text and a bracketed paste must go out as separate delayed writes. ⚠️ **Pi**: no bypass flag, never invent one; `approveProjectTrust` executes repo code, so it is in the clamp's **materialize** branch; never wire `--api-key`. ⚠️ **Grok**: `alwaysApprove` is stripped for non-granted owners (only-if-sent). ⚠️ **DeepSeek**: the agent is a PROFILE (Run gates on `isDeepSeekRunnable()`); the permission switch is the `DSH_PERMISSION_MODE` env var, so `clampEnvOverridesForOwner()` must DROP `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for non-granted owners; `hooksAvailableForMode()` is per-SESSION for it (pass `sessionHookOptions(session)`) and is never a stand-in for `mode === 'claude'`; answers come from `deepseek-transcript.ts`, paired by header `cwd` + boot window, never newest-mtime. ⚠️ **OMP**: `OMP_AUTH_BROKER_URL`/`_TOKEN` are clamped the same way. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp)
@@ -265,7 +270,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit)
**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and Shell loads the rest only via **Load full history**, never on ordinary scroll. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay)
**Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions)
@@ -275,7 +280,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph
**Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which focuses a hidden `contenteditable` trap and reads the clipboard from the paste event landing there; images upload and their paths are typed in, text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event** (Firefox delivers two per keypress: the `execCommand('paste')` event and the keydown's default action); the one-shot flag lives on the trap, never on a browser check. ⚠️ Do not remove the `execCommand('paste')` call: on some mobile engines it is the only route into the trap, and the trap is the only place image blobs are read. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv) **Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which focuses a hidden `contenteditable` trap and reads the clipboard from the paste event landing there; images upload and their paths are typed in, text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event** (Firefox delivers two per keypress: the `execCommand('paste')` event and the keydown's default action); the one-shot flag lives on the trap, never on a browser check. ⚠️ Do not remove the `execCommand('paste')` call: on some mobile engines it is the only route into the trap, and the trap is the only place image blobs are read. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv)
**Terminal scrollback strip + wheel/touch forwarding**: codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only). ⚠️ Gated on `useMux`: direct-PTY sessions must keep the alt screen. Wheel and touch forward to the CLI for **claude ≥ 2.1.187 ONLY**; ⚠️ never re-add codex without a fresh measurement (it ignores SGR wheel reports). ⚠️ `getClaudeCliVersion()` must never cache a FAILED probe. ⚠️ Hand-report clicks only while the CLI has mouse tracking on: `_shouldReportMouseToCli()` gates all three report sites on `cliMouseTracking` (from `_recordStrippedMouseMode()`, session.ts), or a plain shell prints the reports as literal text. Read `_logScrollRouting()` before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding) **Terminal scrollback strip + wheel/touch forwarding**: codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only). ⚠️ Gated on `useMux`: direct-PTY sessions must keep the alt screen. Wheel and touch forward to the CLI for **claude ≥ 2.1.187 ONLY, and only while it has mouse tracking on** (`cliMouseTracking`: fullscreen claude sets it, its default inline renderer does not and scrolls locally like codex); ⚠️ never re-add codex without a fresh measurement (it ignores SGR wheel reports). ⚠️ `getClaudeCliVersion()` must never cache a FAILED probe. ⚠️ Hand-report clicks only while the CLI has mouse tracking on: `_shouldReportMouseToCli()` gates all three report sites on `cliMouseTracking` (from `_recordStrippedMouseMode()`, session.ts), or a plain shell prints the reports as literal text. Read `_logScrollRouting()` before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding)
**Detached start + service install**: `codeman web -d` relaunches the same entry script `detached:true` (setsid); `nohup` is not what makes it survive. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile + `/api/status` probe), or a second instance attaches to the first one's live sessions. ⚠️ Never report success not observed: poll `/api/status` until the child answers or dies. `--stop` must verify the pid still looks like Codeman (`ps -o command=`) before signalling. Unit/label names live only in `config/service-names.ts`. `service install` bakes the installing shell's PATH into the unit and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install) **Detached start + service install**: `codeman web -d` relaunches the same entry script `detached:true` (setsid); `nohup` is not what makes it survive. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile + `/api/status` probe), or a second instance attaches to the first one's live sessions. ⚠️ Never report success not observed: poll `/api/status` until the child answers or dies. `--stop` must verify the pid still looks like Codeman (`ps -o command=`) before signalling. Unit/label names live only in `config/service-names.ts`. `service install` bakes the installing shell's PATH into the unit and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install)
**Self-update** (App Settings → System → Updates): in-app updater for git-clone installs under a supervisor (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none`). The work runs in a DETACHED `scripts/self-update.sh` writing `update-status.json`, polled across the restart; pure helpers in `src/web/self-update.ts`. ⚠️ Compose: the restart kills the script, so nothing may be appended after the `restarting` marker; the repo must stay a host bind mount over `/opt/codeman` and the image must keep devDependencies + toolchain. ⚠️ `evaluateEnvironmentGate()` refuses releases that change `server.Dockerfile`/`docker-compose.yaml` or add `.env.example` keys, re-evaluated on `POST /api/system/update`; unknowns fail OPEN, but the exit-to-restart needs `--restart-by-exit 1` (`CODEMAN_RESTART_BY_EXIT=1` only in the Compose file). ⚠️ Keep the agent CLIs in `server.Dockerfile` pinned. → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update) **Self-update** (App Settings → System → Updates): in-app updater for git-clone installs under a supervisor (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none`). The work runs in a DETACHED `scripts/self-update.sh` writing `update-status.json`, polled across the restart; pure helpers in `src/web/self-update.ts`. ⚠️ Compose: the restart kills the script, so nothing may be appended after the `restarting` marker; the repo must stay a host bind mount over `/opt/codeman` and the image must keep devDependencies + toolchain. ⚠️ `evaluateEnvironmentGate()` refuses releases that change `server.Dockerfile`/`docker-compose.yaml` or add `.env.example` keys, re-evaluated on `POST /api/system/update`; unknowns fail OPEN, but the exit-to-restart needs `--restart-by-exit 1` (`CODEMAN_RESTART_BY_EXIT=1` only in the Compose file). ⚠️ Keep the agent CLIs in `server.Dockerfile` pinned. → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update)
+4
View File
@@ -19,6 +19,10 @@
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a> <a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p> </p>
<p align="center">
⭐ <strong>Like Codeman? <a href="https://github.com/Ark0N/Codeman">Give it a star on GitHub!</a></strong> It takes one click and helps more people find the project. ⭐
</p>
<p align="center"> <p align="center">
<strong>English</strong> &bull; <a href="README.zh-CN.md">简体中文</a> <strong>English</strong> &bull; <a href="README.zh-CN.md">简体中文</a>
</p> </p>
+4
View File
@@ -23,6 +23,10 @@
<a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a> <a href="https://github.com/Ark0N/Codeman/commits/master"><img src="https://img.shields.io/github/commit-activity/t/Ark0N/Codeman?style=flat-square&color=1e3a5f" alt="Total commits"></a>
</p> </p>
<p align="center">
⭐ <strong>喜欢 Codeman?<a href="https://github.com/Ark0N/Codeman">在 GitHub 上给它点个 Star 吧!</a></strong>只需轻点一下,就能帮助更多人发现这个项目。⭐
</p>
<p align="center"> <p align="center">
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — 并行子智能体可视化" width="900"> <img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — 并行子智能体可视化" width="900">
</p> </p>
+6
View File
@@ -15,6 +15,12 @@ TZ=Australia/Perth
# this value rebuilds the image with a matching account. # this value rebuilds the image with a matching account.
CODEMAN_RUNTIME_USER=codeman CODEMAN_RUNTIME_USER=codeman
# Optional Git identity for commits made by Codeman and Docker-case agents. These values
# are written to each image's system Git configuration when it is rebuilt, so
# deployments can configure a consistent default. Set both values together.
# GIT_USER_NAME=
# GIT_USER_EMAIL=
# Required. Persistent Codeman application data, CLI credentials, and session # Required. Persistent Codeman application data, CLI credentials, and session
# state are stored here on the host and mounted at the runtime account's home # state are stored here on the host and mounted at the runtime account's home
# directory in the container. # directory in the container.
+21
View File
@@ -67,6 +67,27 @@ two volumes are removed, by name within this Compose project; any volume a
`docker-compose.override.yml` adds is left alone, and application data and `docker-compose.override.yml` adds is left alone, and application data and
case workspaces are host bind mounts, never touched either way. case workspaces are host bind mounts, never touched either way.
## Git commit identity
Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` before rebuilding:
```sh
GIT_USER_NAME='Your Name'
GIT_USER_EMAIL='you@example.com'
```
Compose passes the values to the Codeman server build, and to the server process
when it builds Docker-case agent images. Both images write the pair to Git's
system configuration during their build, so commits retain the same identity
after a container or agent image is recreated. Set both values together; an
image build with only one value fails rather than using a partial identity. An
identity already present in `CODEMAN_APPDATA_PATH`'s `~/.gitconfig` overrides
the server image's system-level default.
Run `bash docker/Start-Codeman.sh` after changing the server values. Rebuild an
existing agent image with `node scripts/build-agent-image.mjs --no-cache` in the
server container, then recreate any Docker cases that should use it.
## Private repositories (GitHub and Azure DevOps) ## Private repositories (GitHub and Azure DevOps)
The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`. The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`.
+15
View File
@@ -254,6 +254,21 @@ RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& chgrp -R 0 /home/agent \ && chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent && chmod -R g=u /home/agent
# Docker cases have a fresh, container-owned home directory. Declare the
# optional identity here so changing it invalidates only this final layer, then
# configure Git's system defaults. A user-level config still takes precedence.
ARG GIT_USER_EMAIL=
ARG GIT_USER_NAME=
RUN set -eux; \
if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \
if [ -z "${GIT_USER_NAME}" ] || [ -z "${GIT_USER_EMAIL}" ]; then \
echo 'Git user name and email must both be set when configuring Git identity' >&2; \
exit 1; \
fi; \
git config --system user.name "${GIT_USER_NAME}"; \
git config --system user.email "${GIT_USER_EMAIL}"; \
fi
USER agent USER agent
WORKDIR /home/agent WORKDIR /home/agent
+6
View File
@@ -7,6 +7,8 @@ services:
dockerfile: docker/server.Dockerfile dockerfile: docker/server.Dockerfile
args: args:
CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER} CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER}
GIT_USER_EMAIL: ${GIT_USER_EMAIL:-}
GIT_USER_NAME: ${GIT_USER_NAME:-}
PGID: ${PGID:-1000} PGID: ${PGID:-1000}
PUID: ${PUID:-1000} PUID: ${PUID:-1000}
image: ${CODEMAN_IMAGE} image: ${CODEMAN_IMAGE}
@@ -32,6 +34,10 @@ services:
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH} CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT} CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH} CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
# Passed through only so Codeman can use the same identity when it builds
# the Docker-case agent image.
CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: ${GIT_USER_EMAIL:-}
CODEMAN_AGENT_IMAGE_GIT_USER_NAME: ${GIT_USER_NAME:-}
# Extra Host-header allowlist entries for a reverse-proxied deployment # Extra Host-header allowlist entries for a reverse-proxied deployment
# (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it # (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it
# defaults to empty rather than requiring a line in every .env. # defaults to empty rather than requiring a line in every .env.
+18
View File
@@ -219,6 +219,9 @@ RUN set -eux; \
COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/ COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/
ENV NPM_CONFIG_PREFIX=/opt/codeman-cli ENV NPM_CONFIG_PREFIX=/opt/codeman-cli
ENV PATH=$PATH:/opt/codeman-cli/bin ENV PATH=$PATH:/opt/codeman-cli/bin
# CLIs installed at runtime (Settings -> CLIs, npm redirected to ~/.local by installEnv()) live on the
# persistent home mount, so they survive a container recreate. Appended for the same reason as above.
ENV PATH=$PATH:/home/${CODEMAN_RUNTIME_USER}/.local/bin
# pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which # pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which
# this image leaves to be installed at runtime, see SERVER_INTENTIONAL_OMISSIONS in # this image leaves to be installed at runtime, see SERVER_INTENTIONAL_OMISSIONS in
# test/docker-agent-image-coverage.test.ts) spawns a literal `pnpm` with no npm # test/docker-agent-image-coverage.test.ts) spawns a literal `pnpm` with no npm
@@ -299,6 +302,21 @@ EXPOSE 3000
COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod 0755 /usr/local/bin/entrypoint.sh RUN chmod 0755 /usr/local/bin/entrypoint.sh
# Declare the optional identity immediately before configuring it so a change
# invalidates only this final layer. This is declarative setup: a persisted
# ~/.gitconfig in CODEMAN_APPDATA_PATH still overrides the system-level values.
ARG GIT_USER_EMAIL=
ARG GIT_USER_NAME=
RUN set -eux; \
if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \
if [ -z "${GIT_USER_NAME}" ] || [ -z "${GIT_USER_EMAIL}" ]; then \
echo 'Git user name and email must both be set when configuring Git identity' >&2; \
exit 1; \
fi; \
git config --system user.name "${GIT_USER_NAME}"; \
git config --system user.email "${GIT_USER_EMAIL}"; \
fi
ENTRYPOINT ["/usr/local/bin/entrypoint.sh"] ENTRYPOINT ["/usr/local/bin/entrypoint.sh"]
CMD ["node", "dist/index.js", "web"] CMD ["node", "dist/index.js", "web"]
+9
View File
@@ -311,6 +311,15 @@ worker's prompt but never submitted, and the wait then runs its full timeout on
turn that never started. Verified live; this is the most common silent failure on turn that never started. Verified live; this is the most common silent failure on
this endpoint. this endpoint.
A **plain prompt** (printable text followed by exactly one `\r`, nothing else) is
delivered through tmux even without `useMux`: the text is typed, Enter is pressed as
a separate key, and the server re-presses Enter while the prompt is still visibly
sitting on the composer. Written straight into the pane in one piece, a prompt of
about a hundred characters or more is taken as a paste by Claude Code, its `\r`
becomes a newline, and the prompt stays unsent (measured on 2.1.283). Any other
input (escape sequences, a bracketed-paste frame, a line feed, a bare `\r`) keeps
the raw write, and an explicit `"useMux": false` forces it.
```bash ```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \ curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \ -H 'Content-Type: application/json' \
File diff suppressed because one or more lines are too long
+19 -3
View File
@@ -46,8 +46,9 @@ interface CliEntry {
launch: CliLaunch; // the structured argv template launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id capabilities: CliCapabilities; // what every call site reads instead of the id
// .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how // .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines?, awaitingLine? }
// this CLI's pane shows work, and how it shows work it started in the background // — how this CLI's pane shows work, work it started in the background, and a turn
// that ended waiting for workers it will resume from
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
} }
``` ```
@@ -56,7 +57,7 @@ interface CliEntry {
### Regexes that come from config ### Regexes that come from config
Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. Four capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine` and `capabilities.workDetect.awaitingLine`. All four go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing.
`workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused. `workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused.
@@ -66,6 +67,9 @@ agents` while a monitor, a backgrounded shell or a cloud session is live. Codema
into `Session.watching`, and an idle prompt from such a session opens already acknowledged, into `Session.watching`, and an idle prompt from such a session opens already acknowledged,
so a pane waiting for its own background work never raises an alert a human cannot answer. so a pane waiting for its own background work never raises an alert a human cannot answer.
Group 1 is the label, and a CLI that declares no pattern reports no background work. Group 1 is the label, and a CLI that declares no pattern reports no background work.
Claude's Artifact comment monitor is the one chip that does not count. It waits for a human
to comment on a page the agent published, so Claude's pattern refuses any footer that
carries it, and the idle alert goes out as usual.
Two CLIs declare such a row today, and they put it in different places. Claude writes its Two CLIs declare such a row today, and they put it in different places. Claude writes its
chip on the last row of the screen, so it keeps the default one-row window and anchors on chip on the last row of the screen, so it keeps the default one-row window and anchors on
@@ -76,6 +80,18 @@ entry declares `watchingLines: 3` and matches that row end to end. Both were mea
against live panes rather than read out of a binary, which is the standard for adding a against live panes rather than read out of a binary, which is the standard for adding a
third. third.
`awaitingLine` covers the quiet pane that is neither idle nor watching: a turn that ENDED
to wait for workers the CLI will resume from by itself. When background agents or an
ultracode workflow are still running at turn end, Claude closes the turn with
`✻ Waiting for 1 dynamic workflow to finish` instead of `✻ Brewed for 1m 18s`, and a pane
showing that row counts as working. ⚠️ Claude renders the row once and never redraws it, so
the words are still on screen after the workers report back and the follow-up turn ends.
The pattern is therefore never run over the whole pane: `isAwaitingWorkers()`
(`session-activity.ts`) walks up from the composer past blank, framed and indented rows and
tests only the first row that starts in column 0, which is the newest transcript row. Claude
starts its own rows in column 0 and the agent's prose never does, so the anchor also keeps an
agent from holding its own session busy.
That label is the one value in the registry that an AGENT can influence, because it comes off That label is the one value in the registry that an AGENT can influence, because it comes off
the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern
for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few
+2
View File
@@ -83,6 +83,8 @@ Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.jso
The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated. The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated.
Set `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` and `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` together to configure the agent image's Git identity. Rebuild an existing `codeman/agent:base` with `node scripts/build-agent-image.mjs --no-cache`, then recreate Docker-case containers so they use the rebuilt image.
## Quickest path: one-click "Run in Docker" ## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in. On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
+2
View File
@@ -6,6 +6,8 @@ For the Compose configuration, environment settings, storage migration, and macv
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image. The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
CLIs installed from **App Settings → Agents & CLIs → CLI management** (DeepSeek Harness, Pi, and the other npm-based ones) go to `~/.local` on the `CODEMAN_APPDATA_PATH` mount, so they survive an image rebuild and a container recreate. Releases up to 1.33.1 installed them into the image instead, so a CLI installed from Settings on one of those has to be installed again once after the rebuild. The same applies to a hand-run `npm install -g` inside a session: it writes to the image prefix (`/opt/codeman-cli`) and is lost on the next rebuild, so use `npm install -g --prefix ~/.local <package>` instead.
It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings. It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings.
## Prerequisites ## Prerequisites
+7
View File
@@ -42,10 +42,17 @@ npm run typecheck
npm run lint npm run lint
npm run format:check npm run format:check
npm run check:frontend-syntax npm run check:frontend-syntax
npm run check:browser-excludes
npm test -- test/<file>.test.ts # one file, the normal way npm test -- test/<file>.test.ts # one file, the normal way
npm run test:ci # the full CI sweep npm run test:ci # the full CI sweep
``` ```
`npm install` installs a `pre-push` git hook that runs the static checks above (about 10-40s,
machine-dependent) and blocks a push that would fail them. It skips itself when you push
something other than the checked-out HEAD, or when the tree has uncommitted changes the
checks would read. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; a
`pre-push` hook of your own is never overwritten.
**Never run bare `npm test`.** The default configuration includes browser-driven Playwright **Never run bare `npm test`.** The default configuration includes browser-driven Playwright
suites that need a live server, Chromium, and environment-specific baselines; they hang or suites that need a live server, Chromium, and environment-specific baselines; they hang or
fail on a normal machine. `test:ci` is the honest "run everything". fail on a normal machine. `test:ci` is the honest "run everything".
+6
View File
@@ -147,6 +147,12 @@ in `docker-compose.override.yml`), then rebuild the image with `--no-cache`.
`docker/README.md` ("Private repositories") has the details and the matching switches for `docker/README.md` ("Private repositories") has the details and the matching switches for
the server image. the server image.
To give agents a fixed Git commit identity, set `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` and
`CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` together in that same environment (in the Docker
deployment, set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` instead, which feeds
both images). An existing `codeman/agent:base` only picks it up after a `--no-cache` rebuild
and recreated case containers; `docker/README.md` ("Git commit identity") has the details.
## Isolation ## Isolation
Every container runs hardened by default: Every container runs hardened by default:
+4
View File
@@ -127,6 +127,10 @@ plain prose is not a dialog, so an agent that starts a monitor and then writes "
should I target?" is quiet along with the rest — check a watching session yourself if it has should I target?" is quiet along with the rest — check a watching session yourself if it has
been quiet longer than the work it is waiting for should take. been quiet longer than the work it is waiting for should take.
An agent waiting for your comments on an artifact it published never counts as watching.
Claude shows that as "1 Artifact comment monitor", but the agent hears nothing until you
comment, so the session alerts you like any other quiet session.
## The phone overview ## The phone overview
On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then
+7 -4
View File
@@ -151,10 +151,13 @@ Worth knowing:
- **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open. - **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open.
Shell sessions open from a bounded recent tail so a large transcript cannot stall tab Shell sessions open from a bounded recent tail so a large transcript cannot stall tab
switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling switching. Scrolling to the top of a Shell pane pulls the most recent 1 MiB of its tmux
and automatic output recovery stay within the bounded browser buffer. history; press **Load full history** to pull the rest explicitly. Automatic output
- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude recovery stays within the bounded browser buffer.
versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is - **Wheel and touch scrolling** are forwarded into Claude's own transcript when a recent
Claude runs fullscreen (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in
`~/.claude/settings.json`), so the wheel scrolls the conversation rather than the terminal.
Claude's default inline view keeps its history in the terminal and scrolls locally. `Shift+Wheel` is
always local scrollback. Other CLIs scroll locally. always local scrollback. Other CLIs scroll locally.
- **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not. - **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not.
`Ctrl+Shift+C` always copies. `Ctrl+Shift+C` always copies.
+5 -2
View File
@@ -164,8 +164,11 @@ Scrollback behaviour depends on the CLI, and Codeman adjusts what it strips per
Things to try: Things to try:
- `Shift+Wheel` always scrolls the local buffer, whatever else is going on. - `Shift+Wheel` always scrolls the local buffer, whatever else is going on.
- On Claude sessions with a recent CLI, the wheel is forwarded into Claude's own transcript, - On Claude sessions running fullscreen (recent CLI with mouse tracking on), the wheel is
so it scrolls the conversation rather than the terminal buffer. That is intended. forwarded into Claude's own transcript, so it scrolls the conversation rather than the
terminal buffer. That is intended. Claude's default inline view scrolls locally; turn
fullscreen on with `CLAUDE_CODE_NO_FLICKER=1` or `"tui": "fullscreen"` in
`~/.claude/settings.json`.
- Scrolling to the very top pulls the full tmux scrollback again on demand. - Scrolling to the very top pulls the full tmux scrollback again on demand.
### The wheel does nothing in a Codex session ### The wheel does nothing in a Codex session
+2 -2
View File
@@ -1,12 +1,12 @@
{ {
"name": "aicodeman", "name": "aicodeman",
"version": "1.33.1", "version": "1.33.2",
"lockfileVersion": 3, "lockfileVersion": 3,
"requires": true, "requires": true,
"packages": { "packages": {
"": { "": {
"name": "aicodeman", "name": "aicodeman",
"version": "1.33.1", "version": "1.33.2",
"hasInstallScript": true, "hasInstallScript": true,
"license": "MIT", "license": "MIT",
"workspaces": [ "workspaces": [
+2 -1
View File
@@ -1,6 +1,6 @@
{ {
"name": "aicodeman", "name": "aicodeman",
"version": "1.33.1", "version": "1.33.2",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence", "description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module", "type": "module",
"main": "dist/index.js", "main": "dist/index.js",
@@ -28,6 +28,7 @@
"pretest:mobile": "node scripts/prepare-test-vendor.mjs", "pretest:mobile": "node scripts/prepare-test-vendor.mjs",
"test:mobile": "vitest run --config test/mobile/vitest.config.ts", "test:mobile": "vitest run --config test/mobile/vitest.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs", "check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"check:browser-excludes": "node scripts/check-browser-test-excludes.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs", "fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit && tsc -p config/tsconfig.scripts.json", "typecheck": "tsc --noEmit && tsc -p config/tsconfig.scripts.json",
"lint": "eslint --config config/eslint.config.js 'src/**/*.ts'", "lint": "eslint --config config/eslint.config.js 'src/**/*.ts'",
+1 -1
View File
@@ -1,7 +1,7 @@
{ {
"name": "codeman", "name": "codeman",
"description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", "description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.",
"version": "1.33.1", "version": "1.33.2",
"author": { "author": {
"name": "Ark0N", "name": "Ark0N",
"url": "https://github.com/Ark0N" "url": "https://github.com/Ark0N"
+185
View File
@@ -0,0 +1,185 @@
#!/usr/bin/env node
/**
* Browser-test exclusion check.
*
* `npm run test:ci` must never try to drive a real browser: CI runners (and any
* clean checkout) have no chromium, so such a file dies with
* `browserType.launch: Executable doesn't exist` and takes the whole suite with
* it. `config/vitest.ci.config.ts` therefore excludes every browser-driven test
* via `BROWSER_TEST_GLOBS` in `config/test-suites.ts`. That list is maintained
* BY HAND, and a new browser test simply does not appear in it unless someone
* remembers. The omission is invisible on a developer machine that has run
* `npx playwright install`, where the test passes, and only shows up on a clean
* runner.
*
* Two deliberate design choices:
*
* 1. **Detection is by CONTENT, not filename.** Matching `*.browser.test.ts`
* would miss the browser tests that predate that convention
* (`inline-rename`, `opencode-resize`, `webgl-fallback`,
* `terminal-copy-shortcut`, `codex-predictive-echo`). What actually makes a
* file dangerous is importing a browser driver, so that is what is tested.
* ⚠️ Only a DIRECT import is seen: a test that reaches playwright through a
* helper module (e.g. `test/mobile/helpers/browser.ts`) is not detected, so
* such a test still has to be added to `BROWSER_TEST_GLOBS` by hand.
*
* 2. **The exclusion side is answered by vitest itself**, via
* `vitest list --filesOnly`, rather than by re-implementing glob matching
* against the config's `exclude` array. Patterns there include `test/mobile/**`
* and `perf-*`; a hand-rolled matcher that disagreed with vitest by even one
* edge case would report a gap that does not exist, or miss one that does.
* Asking the real resolver cannot drift from the real behaviour.
*
* The pure pieces are exported for test/check-browser-test-excludes.test.ts; the
* check itself only runs when this file is executed directly.
*/
import { readdirSync, readFileSync } from 'node:fs';
import { join, dirname, relative, sep, resolve } from 'node:path';
import { fileURLToPath } from 'node:url';
import { execFileSync } from 'node:child_process';
const ROOT = join(dirname(fileURLToPath(import.meta.url)), '..');
const CI_CONFIG = join('config', 'vitest.ci.config.ts');
const SUITES_FILE = join('config', 'test-suites.ts');
/** Importing any one of these means the test needs a real browser binary. */
const BROWSER_DRIVER =
/\bfrom\s+['"](?:playwright|playwright-core|@playwright\/test|puppeteer|puppeteer-core)['"]|\b(?:require|import)\(\s*['"](?:playwright|playwright-core|@playwright\/test|puppeteer|puppeteer-core)['"]\s*\)/;
/** @param {string} source */
export function importsBrowserDriver(source) {
return BROWSER_DRIVER.test(source);
}
/** @param {string} dir @returns {string[]} */
function walk(dir) {
const out = [];
for (const entry of readdirSync(dir, { withFileTypes: true })) {
const path = join(dir, entry.name);
if (entry.isDirectory()) out.push(...walk(path));
else if (entry.isFile() && entry.name.endsWith('.test.ts')) out.push(path);
}
return out;
}
/**
* Every `*.test.ts` under `<root>/test`, as sorted repo-relative POSIX paths (the form
* `vitest list` prints).
*
* @param {string} root
* @returns {string[]}
*/
export function findTestFiles(root) {
return walk(join(root, 'test'))
.map((file) => relative(root, file).split(sep).join('/'))
.sort();
}
/**
* The subset of {@link findTestFiles} that imports a browser driver.
*
* @param {string} root
* @returns {string[]}
*/
export function findBrowserTests(root) {
return findTestFiles(root).filter((file) => importsBrowserDriver(readFileSync(join(root, file), 'utf8')));
}
/**
* Parse `vitest list --filesOnly` output into a set of repo-relative paths. Stray
* blank or decorative lines are ignored rather than assuming the format is pristine.
*
* @param {string} output
* @returns {Set<string>}
*/
export function parseVitestFileList(output) {
return new Set(
output
.split('\n')
.map((line) => line.trim())
.filter((line) => line.endsWith('.test.ts'))
.map((line) => line.replace(/^\.\//, ''))
);
}
/**
* Whether the `vitest list` paths and the walked tree name at least one file in common.
* False means the two sides are not speaking the same path format (absolute paths, backslashes
* or a new prefix after a vitest upgrade), and then {@link findLeaks} would find nothing
* against a perfectly non-empty listing.
*
* @param {Set<string>} ciFiles
* @param {string[]} testFiles
*/
export function listingMatchesTree(ciFiles, testFiles) {
return testFiles.some((file) => ciFiles.has(file));
}
/**
* @param {string[]} browserTests
* @param {Set<string>} ciFiles
* @returns {string[]} browser-driven files that the CI config would still collect
*/
export function findLeaks(browserTests, ciFiles) {
return browserTests.filter((file) => ciFiles.has(file));
}
function main() {
const testFiles = findTestFiles(ROOT);
const browserTests = findBrowserTests(ROOT);
let collected;
try {
collected = execFileSync('npx', ['vitest', 'list', '--config', CI_CONFIG, '--filesOnly'], {
cwd: ROOT,
encoding: 'utf8',
stdio: ['ignore', 'pipe', 'pipe'],
});
} catch (err) {
console.error('✗ could not enumerate the CI test set via `vitest list`.');
console.error(err.stderr ? err.stderr.toString() : String(err));
process.exit(1);
}
const ciFiles = parseVitestFileList(collected);
if (ciFiles.size === 0) {
// An empty list would make every browser test look excluded: fail rather than pass vacuously.
console.error('✗ `vitest list` reported no test files; refusing to pass on an empty CI set.');
process.exit(1);
}
// Same vacuous pass, one step removed: a listing whose paths never match the tree. This guard,
// not `vitest list --json`, is the answer to format drift: the JSON form prints absolute paths
// that would need canonicalizing against ROOT (symlinked checkouts), and its shape can drift too.
if (!listingMatchesTree(ciFiles, testFiles)) {
const sample = [...ciFiles].slice(0, 3).join(', ');
console.error(
`✗ none of the ${ciFiles.size} paths \`vitest list\` reported (e.g. ${sample}) is one of the ${testFiles.length} test/**/*.test.ts files; its output format has probably changed.`
);
process.exit(1);
}
const leaked = findLeaks(browserTests, ciFiles);
if (leaked.length > 0) {
console.error(`✗ ${leaked.length} browser-driven test file(s) are NOT excluded from ${CI_CONFIG}:\n`);
for (const file of leaked) console.error(` ${file}`);
console.error(`
These import a browser driver, so on a runner with no chromium they fail with
"browserType.launch: Executable doesn't exist" and take the suite down. Add each
to BROWSER_TEST_GLOBS in ${SUITES_FILE} (${CI_CONFIG} derives its excludes from
it, and \`npm run test:browser\` its includes).
They may well pass on this machine; that is the trap. To reproduce a clean
runner locally:
PLAYWRIGHT_BROWSERS_PATH=\$(mktemp -d) PUPPETEER_CACHE_DIR=\$(mktemp -d) npm run test:ci`);
process.exit(1);
}
console.log(
`✓ all ${browserTests.length} browser-driven test files are excluded from the CI suite (${ciFiles.size} files collected)`
);
}
if (process.argv[1] && resolve(process.argv[1]) === fileURLToPath(import.meta.url)) {
main();
}
+253
View File
@@ -0,0 +1,253 @@
/**
* @fileoverview Git hook bodies + install policy, shared by scripts/postinstall.js and
* pinned by test/git-hooks.test.ts.
*
* Why a pre-push hook: the static CI job (lockfile, typecheck, lint, format, frontend
* syntax, ...) fails often on things a contributor could have caught locally in seconds,
* and finding out after a push costs a full CI round-trip plus a fix-up commit. Running
* the same checks before the push surfaces those failures in ~10-40s instead (12s on a fast
* workstation, ~35s measured elsewhere; typecheck, format:check and lint dominate).
*
* Why pre-PUSH and not pre-commit: a commit is cheap and local, a push is what CI and
* reviewers pick up. And why the STATIC tier only: the unit/integration suite takes
* minutes, which nobody tolerates per push, so a hook that ran it would be bypassed
* within a day. The checks below mirror the static CI job.
*
* ⚠️ The checks read the WORKING TREE, not the commits being pushed. So the hook skips
* (with a one-line notice) whenever the two can differ: when HEAD is not the commit being
* pushed, and when `git status` shows uncommitted or untracked changes in a path a check
* reads ({@link PRE_PUSH_WATCHED_PATHS}). In a checkout shared by several agent sessions
* the second case is usually another session's WIP, which must not block this push.
*
* ⚠️ This installer is deliberately MARKER-OWNED, unlike the older pre-commit installer in
* postinstall.js which overwrites whatever it finds. A developer's own pre-push hook must
* survive `npm install`.
*/
import { execFileSync } from 'node:child_process';
import { chmodSync, existsSync, mkdirSync, readFileSync, realpathSync, writeFileSync } from 'node:fs';
import { basename, dirname, join, resolve } from 'node:path';
/**
* Ownership marker. ⚠️ Never bump the version suffix: ownership is matched on this exact
* string, so a `v2` would read every installed `v1` hook as foreign and never refresh it.
* A changed body still reaches installed hooks, because the refresh compares the whole file.
*/
export const PRE_PUSH_MARKER = '# codeman-managed-hook: pre-push v1';
/**
* Checks that make up the fast tier, cheapest first so failures surface sooner. Each entry
* is the argument list for `npm run`, and each is a step of the static job in
* .github/workflows/ci.yml (test/git-hooks.test.ts pins that every script exists).
*/
export const PRE_PUSH_CHECKS = [
['check:lockfile'],
['generate:cli-catalog', '--', '--check'],
['check:browser-excludes'],
['check:frontend-syntax'],
['format:check'],
['lint'],
['typecheck'],
];
/**
* Paths whose uncommitted state would leak into a check, so a dirty one makes the hook skip.
* Derived from what each check reads: src/ (format:check, lint, typecheck,
* check:frontend-syntax), config/ (eslint + vitest configs, test-suites.ts, the CLI
* catalogue), scripts/ (every check is a script there, and typecheck's second pass compiles
* one), test/ (check:browser-excludes scans it and runs `vitest list` over it),
* package.json + package-lock.json (check:lockfile), install.sh (generate:cli-catalog
* --check diffs its generated block), tsconfig.json (typecheck, and
* config/tsconfig.scripts.json extends it) and .prettierignore + .editorconfig
* (format:check; the Prettier CLI honours .editorconfig by default).
*/
export const PRE_PUSH_WATCHED_PATHS = [
'src',
'config',
'scripts',
'test',
'package.json',
'package-lock.json',
'install.sh',
'tsconfig.json',
'.prettierignore',
'.editorconfig',
];
/**
* Render the pre-push hook script.
*
* POSIX sh, not bash: this ships to whatever shell the contributor's git uses.
*/
export function renderPrePushHook() {
const runs = PRE_PUSH_CHECKS.map((args) => `run_check ${args.join(' ')}`).join('\n');
const watched = PRE_PUSH_WATCHED_PATHS.join(' ');
return `#!/bin/sh
${PRE_PUSH_MARKER}
# Installed by scripts/postinstall.js. Edit scripts/git-hooks.mjs, not this file:
# it is regenerated on npm install. Delete the marker line above to take ownership
# and the installer will leave your version alone.
#
# Skip once: CODEMAN_SKIP_PREPUSH=1 git push
# Skip always: remove this file.
[ "$CODEMAN_SKIP_PREPUSH" = "1" ] && exit 0
repo_root=$(git rev-parse --show-toplevel 2>/dev/null) || exit 0
cd "$repo_root" || exit 0
# Nothing to check without dependencies (fresh clone, or a worktree that never ran
# npm install). Warn rather than blocking the push on a setup detail.
if [ ! -d node_modules ]; then
echo "pre-push: node_modules missing, skipping checks (run 'npm install' to enable them)."
exit 0
fi
# GUI git clients and IDEs often run hooks with a minimal PATH that lacks an nvm or
# Homebrew Node. Every check would then fail with "npm: not found", so skip instead.
command -v npm >/dev/null 2>&1 || { echo "pre-push: npm not on PATH, skipping checks."; exit 0; }
# git feeds us "<localref> <localsha> <remoteref> <remotesha>" per ref. A deletion has an
# all-zero local sha and no tree worth checking; if every ref is a deletion, skip.
# The checks below read the working tree, so they only say something about a pushed commit
# that IS the checked-out HEAD (tags are peeled to their commit first).
head=$(git rev-parse -q --verify HEAD 2>/dev/null)
has_content=0
not_head=''
while read -r localref localsha _remoteref _remotesha; do
[ -z "$localsha" ] && continue
case "$localsha" in
0000000000000000000000000000000000000000) ;;
*)
has_content=1
commit=$(git rev-parse -q --verify "$localsha^{commit}" 2>/dev/null)
[ -n "$head" ] && [ "$commit" = "$head" ] || not_head="$localref"
;;
esac
done
[ "$has_content" = "0" ] && exit 0
if [ -n "$not_head" ]; then
echo "pre-push: skipping static checks: $not_head is not the checked-out HEAD, and the checks read the working tree."
exit 0
fi
# Uncommitted or untracked changes in a path a check reads would be judged instead of the
# pushed commit. In a checkout shared by several sessions that is usually someone else's WIP.
if [ -n "$(git --no-optional-locks status --porcelain -- ${watched} 2>/dev/null)" ]; then
echo "pre-push: skipping static checks: uncommitted changes under ${watched} would be checked instead of the pushed commit."
exit 0
fi
log=$(mktemp "\${TMPDIR:-/tmp}/codeman-prepush.XXXXXX") || exit 0
trap 'rm -f "$log"' EXIT
failed=''
run_check() {
if ! npm run --silent "$@" >"$log" 2>&1; then
echo ""
echo "pre-push: FAILED npm run $*"
tail -n 25 "$log"
failed="$failed $1"
fi
}
echo "pre-push: running static checks (~10-40s)..."
${runs}
if [ -n "$failed" ]; then
echo ""
echo "pre-push: blocked by:$failed"
echo "Fix, or push anyway with: CODEMAN_SKIP_PREPUSH=1 git push"
exit 1
fi
echo "pre-push: static checks passed."
exit 0
`;
}
/**
* Decide what to do with an existing hook file.
*
* @param {{ existing: string | null | undefined, next: string }} args
* @returns {'write' | 'up-to-date' | 'skip-foreign'}
*/
export function planHookInstall({ existing, next }) {
if (existing === null || existing === undefined || existing.trim() === '') return 'write';
if (!existing.includes(PRE_PUSH_MARKER)) return 'skip-foreign';
return existing === next ? 'up-to-date' : 'write';
}
/** @param {string} cwd @param {string[]} args */
function git(cwd, args) {
return execFileSync('git', args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'] }).trim();
}
/**
* realpath() that tolerates a missing leaf: a fresh `.git` may have no `hooks/` yet, so
* canonicalize the parent and re-append the name. Throws if the parent is missing too.
*
* @param {string} path
*/
function canonicalPath(path) {
return existsSync(path) ? realpathSync(path) : join(realpathSync(dirname(path)), basename(path));
}
/**
* Resolve the hooks directory for the checkout rooted at `repoRoot`, or null when there
* is nothing to install into.
*
* Asks git (`--git-path hooks`) rather than assuming `<root>/.git/hooks`: in a worktree
* `.git` is a FILE pointing at the parent repo, so the hooks live under
* `--git-common-dir`.
*
* ⚠️ Returns a directory ONLY when it is this repository's own `<git-common-dir>/hooks`.
* `--git-path hooks` also reports `core.hooksPath`, and that setting is often GLOBAL (a
* shared hooks directory used by every repo on the machine); installing there would
* overwrite the user's own hooks and run Codeman's checks on unrelated repos. A
* `core.hooksPath` that points back at the repo's own hooks dir still resolves, because
* the comparison is on canonical paths rather than on whether the setting exists.
*
* Also returns null unless `repoRoot` is itself the top of a work tree. Without that guard,
* a copy of this package sitting inside SOMEONE ELSE's repository (e.g. under their
* node_modules) would resolve to their hooks directory and install Codeman's hook there.
*
* @param {string} repoRoot
* @returns {string | null}
*/
export function resolveGitHooksDir(repoRoot) {
try {
const top = git(repoRoot, ['rev-parse', '--show-toplevel']);
if (!top || realpathSync(top) !== realpathSync(repoRoot)) return null;
// Both are printed relative to the cwd (repoRoot) unless already absolute.
const hooks = git(repoRoot, ['rev-parse', '--git-path', 'hooks']);
const common = git(repoRoot, ['rev-parse', '--git-common-dir']);
if (!hooks || !common) return null;
const own = join(realpathSync(resolve(repoRoot, common)), 'hooks');
return canonicalPath(resolve(repoRoot, hooks)) === own ? own : null;
} catch {
return null;
}
}
/**
* Install (or refresh) the managed pre-push hook in `hooksDir`, honouring
* {@link planHookInstall}: a hook without the marker is never touched.
*
* @param {string} hooksDir
* @returns {'write' | 'up-to-date' | 'skip-foreign'}
*/
export function installPrePushHook(hooksDir) {
const path = join(hooksDir, 'pre-push');
const next = renderPrePushHook();
const existing = existsSync(path) ? readFileSync(path, 'utf8') : null;
const action = planHookInstall({ existing, next });
if (action === 'write') {
mkdirSync(hooksDir, { recursive: true });
writeFileSync(path, next, { mode: 0o755 });
chmodSync(path, 0o755); // `mode` only applies when the file is created
}
return action;
}
+26 -1
View File
@@ -64,6 +64,15 @@ export const GIT_HOST_CLI_BUILD_ARGS = [
['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'], ['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'],
]; ];
/**
* Environment variable → Dockerfile ARG for the image's system Git identity.
* ⚠️ Mirrored by `GIT_IDENTITY_BUILD_ARGS` in `src/docker-hosts.ts`; the parity test pins them.
*/
export const GIT_IDENTITY_BUILD_ARGS = [
['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'],
['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'],
];
/** /**
* The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable * The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable
* contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same * contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same
@@ -82,9 +91,25 @@ export function gitHostCliBuildArgPairs(env) {
return pairs; return pairs;
} }
/** The `--build-arg` pairs for Git identity, requiring either both values or neither. */
export function gitIdentityBuildArgPairs(env) {
const pairs = GIT_IDENTITY_BUILD_ARGS.map(([envName, argName]) => [argName, env[envName] ?? '']);
const configured = pairs.filter(([, value]) => value !== '');
if (configured.length === 0) return [];
if (configured.length !== pairs.length) {
const names = GIT_IDENTITY_BUILD_ARGS.map(([envName]) => envName).join(' and ');
throw new Error(`${names} must both be set when configuring Git identity`);
}
return pairs;
}
/** The `--build-arg` pairs the agent image takes. PURE given `env`. */ /** The `--build-arg` pairs the agent image takes. PURE given `env`. */
export function agentImageBuildArgPairs(catalog, env = process.env) { export function agentImageBuildArgPairs(catalog, env = process.env) {
return [['CLI_NPM_PACKAGES', agentImageNpmPackages(catalog).join(' ')], ...gitHostCliBuildArgPairs(env)]; return [
['CLI_NPM_PACKAGES', agentImageNpmPackages(catalog).join(' ')],
...gitHostCliBuildArgPairs(env),
...gitIdentityBuildArgPairs(env),
];
} }
/** Read the committed catalogue. IO. */ /** Read the committed catalogue. IO. */
+16 -4
View File
@@ -356,14 +356,17 @@ if (!isGlobalInstall) {
} }
// ---------------------------------------------------------------------------- // ----------------------------------------------------------------------------
// 5. Install git pre-commit hook (format check) // 5. Install git hooks (pre-commit format check, pre-push static checks)
// ---------------------------------------------------------------------------- // ----------------------------------------------------------------------------
if (!isGlobalInstall) { if (!isGlobalInstall) {
try { try {
const { writeFileSync, mkdirSync } = await import('fs'); const { writeFileSync, mkdirSync } = await import('fs');
const gitHooksDir = join(import.meta.dirname, '..', '.git', 'hooks'); const { resolveGitHooksDir, installPrePushHook } = await import('./git-hooks.mjs');
if (existsSync(join(import.meta.dirname, '..', '.git'))) { // Resolved through git, not `../.git/hooks`: in a worktree `.git` is a file.
// null when this directory is not the top of a git checkout.
const gitHooksDir = resolveGitHooksDir(join(import.meta.dirname, '..'));
if (gitHooksDir) {
mkdirSync(gitHooksDir, { recursive: true }); mkdirSync(gitHooksDir, { recursive: true });
const hook = `#!/bin/bash const hook = `#!/bin/bash
# Auto-installed by postinstall — prevents CI format failures # Auto-installed by postinstall — prevents CI format failures
@@ -379,9 +382,18 @@ fi
const hookPath = join(gitHooksDir, 'pre-commit'); const hookPath = join(gitHooksDir, 'pre-commit');
writeFileSync(hookPath, hook, { mode: 0o755 }); writeFileSync(hookPath, hook, { mode: 0o755 });
console.log(colors.green('✓ Git pre-commit hook installed (prettier check)')); console.log(colors.green('✓ Git pre-commit hook installed (prettier check)'));
// Unlike the pre-commit hook above, this one is marker-owned: a pre-push
// hook the developer wrote themselves is left alone.
const action = installPrePushHook(gitHooksDir);
if (action === 'write') {
console.log(colors.green('✓ Git pre-push hook installed') + colors.dim(' (static CI checks, ~10-40s)'));
} else if (action === 'skip-foreign') {
console.log(colors.dim(' Existing pre-push hook left untouched (not Codeman-managed)'));
}
} }
} catch { } catch {
// Non-critical — git hook is a convenience // Non-critical — git hooks are a convenience
} }
} }
+9
View File
@@ -338,6 +338,15 @@ const capabilitiesSchema = z
// Bounded hard: this is how far up the screen a config file may push the search, // Bounded hard: this is how far up the screen a config file may push the search,
// and every row it adds is one more row the agent itself may be able to write. // and every row it adds is one more row the agent itself may be able to write.
watchingLines: z.number().int().min(1).max(8).optional(), watchingLines: z.number().int().min(1).max(8).optional(),
// Same guard again: tested against a pane row every time a session settles.
awaitingLine: z
.string()
.min(1)
.refine(
(src) => compileVersionRegex(src) !== null,
'awaitingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers'
)
.optional(),
}) })
.strict() .strict()
// A window with nothing to search is a typo, not a configuration. Refused at LOAD // A window with nothing to search is a typo, not a configuration. Refused at LOAD
+20 -1
View File
@@ -237,7 +237,24 @@ const CLAUDE: CliEntry = {
// carry a count. A footer that ever drew the chip as its only item would report no // carry a count. A footer that ever drew the chip as its only item would report no
// watching rather than open that door. See `watchingLabel()` in // watching rather than open that door. See `watchingLabel()` in
// `session-activity.ts`. // `session-activity.ts`.
watchingLine: String.raw`·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?|Artifact comment monitors?))`, // ⚠️ An Artifact comment monitor is the one chip that waits on the user. The agent
// has published a page and hears nothing until somebody comments on it, so the
// lookahead refuses the whole row while that chip is on it, whatever else is
// running beside it. The `^` is what makes the lookahead judge the row once:
// without it the engine retries from each later position, and a start past the
// chip reports the shell beside it. The lookahead keys on "Artifact" alone, so a
// footer cut off mid-chip (`· 1 Artifact…`, `· 1 Artifact comm…`) is still refused;
// no other chip on this row says "Artifact". Counting the chip as watching kept the
// idle alert quiet for a session that was waiting for a human.
watchingLine: String.raw`^(?!.*Artifact).*?·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?))`,
// When a turn ends while background agents or an ultracode workflow are still
// running, Claude swaps its `✻ Brewed for 1m 18s` closing row for
// `✻ Waiting for 2 background agents and 1 dynamic workflow to finish` and resumes
// by itself when they report back. Read from the 2.1.283 bundle (the turn-duration
// renderer) and a live pane on 2026-09-28. The row is a snapshot taken at turn end
// and never redrawn, which is why only the newest row above the composer counts.
// Anchored on column 0: Claude's own rows start there, the agent's prose never does.
awaitingLine: String.raw`^✻ Waiting for \d+ (?:background agents?|dynamic workflows?)\b`,
}, },
requiresMux: false, requiresMux: false,
// Claude installs Codeman's own hooks block into every workspace it runs in, so its // Claude installs Codeman's own hooks block into every workspace it runs in, so its
@@ -246,6 +263,8 @@ const CLAUDE: CliEntry = {
transcript: 'claude-jsonl', transcript: 'claude-jsonl',
altScreen: 'strip-full', altScreen: 'strip-full',
echo: { policy: 'buffer', anchor: { kind: 'glyph', glyph: '❯', offset: 2 } }, echo: { policy: 'buffer', anchor: { kind: 'glyph', glyph: '❯', offset: 2 } },
// Declared-for-later: the live rule (`_shouldForwardWheelToApp`, terminal-ui.js) is this version
// AND the server-published `cliMouseTracking` flag (#498), so wiring this field up needs both.
wheelForward: { mode: 'version-gated', minVersion: '2.1.187' }, wheelForward: { mode: 'version-gated', minVersion: '2.1.187' },
keyboardAccessory: 'agent', keyboardAccessory: 'agent',
privilegedCommandGate: false, privilegedCommandGate: false,
+13
View File
@@ -362,6 +362,19 @@ export interface CliCapabilities {
* alert. See `watchingLabel()` in `session-activity.ts`. * alert. See `watchingLabel()` in `session-activity.ts`.
*/ */
watchingLines?: number; watchingLines?: number;
/**
* Source of a regex matching the row this CLI closes a turn with when it ended that
* turn to WAIT for workers it started and will resume on its own once they finish,
* e.g. Claude's `✻ Waiting for 1 dynamic workflow to finish`. A pane showing it counts
* as working, not idle: nothing is being asked of the user, and the next turn starts
* without them.
*
* Unlike `workingLine` this is never searched across the pane. The CLI prints the row
* once and never updates it, so the copy from an earlier turn is still on screen after
* the workers are done. Only the newest transcript row directly above the composer is
* tested. See `isAwaitingWorkers()` in `session-activity.ts`.
*/
awaitingLine?: string;
}; };
/** /**
* How many columns this CLI indents its transcript body by, so a copy taken from its * How many columns this CLI indents its transcript body by, so a copy taken from its
+8
View File
@@ -99,3 +99,11 @@ export const STALE_DATA_MAX_AGE_MS = 60 * 60 * 1000;
/** Standard 5-minute inactivity timeout for streams and caches (ms) */ /** Standard 5-minute inactivity timeout for streams and caches (ms) */
export const INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000; export const INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000;
/**
* Gap between a paste-mode cron prompt's text and its Enter (ms). The two must be
* separate writes: Claude Code takes a raw `<text>\r` burst of about a hundred
* characters as a paste and turns its `\r` into a newline. A separate `\r` 80 ms
* after the text was measured to submit; this leaves room for a longer prompt.
*/
export const CRON_PASTE_ENTER_DELAY_MS = 300;
+38 -9
View File
@@ -20,7 +20,7 @@ import { getErrorMessage, createErrorResponse, ApiErrorCode } from '../types/api
import { MAX_CONCURRENT_SESSIONS, MAX_CRON_JOBS, MAX_CRON_RUN_HISTORY } from '../config/map-limits.js'; import { MAX_CONCURRENT_SESSIONS, MAX_CRON_JOBS, MAX_CRON_RUN_HISTORY } from '../config/map-limits.js';
import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../user-store.js'; import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../user-store.js';
import { sessionCapacityState, isWorkingDirAllowedForUsername } from '../web/route-helpers.js'; import { sessionCapacityState, isWorkingDirAllowedForUsername } from '../web/route-helpers.js';
import { CRON_READY_MAX_ATTEMPTS, CRON_READY_SETTLE_MS } from '../config/server-timing.js'; import { CRON_PASTE_ENTER_DELAY_MS, CRON_READY_MAX_ATTEMPTS, CRON_READY_SETTLE_MS } from '../config/server-timing.js';
import { import {
DEFAULT_BLOCKED_TREES, DEFAULT_BLOCKED_TREES,
isBlockedAttachmentPath, isBlockedAttachmentPath,
@@ -100,6 +100,41 @@ const CRON_WORKING_DIR_BLOCKED_TREES: readonly string[] = [...DEFAULT_BLOCKED_TR
/** Prompt delivery is single-line only (writeViaMux/Ink constraint). */ /** Prompt delivery is single-line only (writeViaMux/Ink constraint). */
const HAS_NEWLINE = /[\r\n]/; const HAS_NEWLINE = /[\r\n]/;
/** The three session calls prompt delivery needs, so it can be tested without a PTY. */
type CronPromptTarget = Pick<Session, 'write' | 'writeViaMux' | 'verifySubmitted'>;
/**
* Send a cron job's (single-line) prompt into its session and press Enter.
*
* `typed` goes through the mux: the text is typed, Enter is its own key, and the
* session re-presses it while the prompt is still on the composer.
*
* `paste` writes the text straight into the PTY, and must send its Enter as a
* SEPARATE write. It used to send `<text>\r` in one piece, and Claude Code (measured
* on 2.1.283) takes a burst of about a hundred characters as a paste, so the `\r`
* landed as a newline and the prompt sat unsent while the run reported
* `prompt_sent`. The Enter goes down the same PTY as the text, so it cannot overtake
* it, and the same composer check then covers a CLI that was not taking Enter yet.
*
* @returns false when the session had no PTY or mux to write to
*/
export async function deliverCronPrompt(
target: CronPromptTarget,
prompt: string,
inputMode: CronJob['inputMode'],
wait: (ms: number) => Promise<void> = delay
): Promise<boolean> {
if (inputMode !== 'paste') {
return target.writeViaMux(prompt.endsWith('\r') ? prompt : `${prompt}\r`);
}
const text = prompt.replace(/[\r\n]+$/, '');
if (!target.write(text)) return false;
await wait(CRON_PASTE_ENTER_DELAY_MS);
if (!target.write('\r')) return false;
target.verifySubmitted(text);
return true;
}
/** Order-insensitive equality for the weekly-days arrays. */ /** Order-insensitive equality for the weekly-days arrays. */
function sameDays(a: number[] | undefined, b: number[] | undefined): boolean { function sameDays(a: number[] | undefined, b: number[] | undefined): boolean {
const x = [...(a ?? [])].sort((p, q) => p - q); const x = [...(a ?? [])].sort((p, q) => p - q);
@@ -623,15 +658,9 @@ export class CronService {
const s = this.deps.sessions.get(sessionId); const s = this.deps.sessions.get(sessionId);
if (!s) return; if (!s) return;
try { try {
const payload = prompt.endsWith('\r') ? prompt : `${prompt}\r`; const delivered = await deliverCronPrompt(s, prompt, job.inputMode);
let delivered = true;
if (job.inputMode === 'paste') {
s.write(payload);
} else {
delivered = await s.writeViaMux(payload);
}
if (!delivered) { if (!delivered) {
this.failRun(job, run, 'Failed to send prompt: mux write failed'); this.failRun(job, run, 'Failed to send prompt: the session could not be written to');
return; return;
} }
run.status = 'prompt_sent'; run.status = 'prompt_sent';
+31 -2
View File
@@ -615,6 +615,15 @@ export const GIT_HOST_CLI_BUILD_ARGS: ReadonlyArray<readonly [string, string]> =
['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'], ['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'],
]; ];
/**
* Environment variable → Dockerfile ARG for the image's system Git identity.
* ⚠️ Mirrors `GIT_IDENTITY_BUILD_ARGS` in `scripts/lib/cli-catalog.mjs`; the parity test pins them.
*/
export const GIT_IDENTITY_BUILD_ARGS: ReadonlyArray<readonly [string, string]> = [
['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'],
['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'],
];
/** /**
* The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable * The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable
* contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same * contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same
@@ -633,9 +642,29 @@ export function gitHostCliBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string,
return pairs; return pairs;
} }
/**
* The `--build-arg` pairs for a configured Git identity. An absent pair leaves
* Git unconfigured, preserving existing deployments; a partial pair is refused.
* ⚠️ Mirrors `gitIdentityBuildArgPairs()` in `scripts/lib/cli-catalog.mjs`; the parity test pins them.
*/
export function gitIdentityBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, string]> {
const pairs = GIT_IDENTITY_BUILD_ARGS.map(([envName, argName]) => [argName, env[envName] ?? ''] as [string, string]);
const configured = pairs.filter(([, value]) => value !== '');
if (configured.length === 0) return [];
if (configured.length !== pairs.length) {
const names = GIT_IDENTITY_BUILD_ARGS.map(([envName]) => envName).join(' and ');
throw new Error(`${names} must both be set when configuring Git identity`);
}
return pairs;
}
/** The `--build-arg` pairs the agent image takes. PURE given `env`. */ /** The `--build-arg` pairs the agent image takes. PURE given `env`. */
export function agentImageBuildArgPairs(env: NodeJS.ProcessEnv = process.env): Array<[string, string]> { export function agentImageBuildArgPairs(env: NodeJS.ProcessEnv = process.env): Array<[string, string]> {
return [['CLI_NPM_PACKAGES', agentImageNpmPackages().join(' ')], ...gitHostCliBuildArgPairs(env)]; return [
['CLI_NPM_PACKAGES', agentImageNpmPackages().join(' ')],
...gitHostCliBuildArgPairs(env),
...gitIdentityBuildArgPairs(env),
];
} }
// ========== Credential mount resolution (IO) ========== // ========== Credential mount resolution (IO) ==========
@@ -1204,7 +1233,7 @@ function buildAgentImage(
try { try {
buildArgPairs = agentImageBuildArgPairs(); buildArgPairs = agentImageBuildArgPairs();
} catch (err) { } catch (err) {
// A malformed CODEMAN_AGENT_IMAGE_INSTALL_* value: report it like any other build failure. // A malformed CODEMAN_AGENT_IMAGE_* value: report it like any other build failure.
return Promise.resolve({ ok: false, built: false, alreadyPresent: false, error: String((err as Error).message) }); return Promise.resolve({ ok: false, built: false, alreadyPresent: false, error: String((err as Error).message) });
} }
const argv = dockerEngineArgv(docker); const argv = dockerEngineArgv(docker);
+56
View File
@@ -158,3 +158,59 @@ export function watchingLabel(
} }
return null; return null;
} }
/**
* How many rows above the composer the turn's closing row may sit. Between the two Claude
* draws only its composer border and, sometimes, a right-aligned hint
* (`new task? /clear to save 169.1k tokens`), so this leaves room for a blank row or two
* and no more. A bound, not a tuning knob: the walk must never reach far enough up the
* transcript to find an old turn's row.
*/
export const AWAITING_SEARCH_ROWS = 6;
/** A row that opens with a box-drawing character is the composer's frame, not transcript. */
const COMPOSER_FRAME_ROW = /^[─-╿]/;
/**
* Whether the pane's newest turn ended by handing off to workers the CLI will wait for,
* e.g. Claude's `✻ Waiting for 1 dynamic workflow to finish`.
*
* Such a pane is quiet and shows its composer, so every other signal calls it idle, yet
* nothing is being asked of the user: the CLI resumes by itself when the workers report
* back. That is why a session in this state counts as working.
*
* ⚠️ The row is a snapshot. Claude renders it once, at the end of the turn, and never
* updates it, so after the workers finish the same words are still on screen above the
* follow-up turn. Matching them anywhere on the pane would pin the session busy for as
* long as they stay visible. Only the newest transcript row counts: the walk starts at
* the composer (the LAST row carrying `promptGlyph`), steps up past blank rows, the
* composer's frame and anything indented (a right-aligned hint, a wrapped continuation),
* and tests the first row that starts in column 0. A follow-up turn always puts rows of
* its own there, so the stale copy is never the one tested.
*
* @param promptGlyph the CLI's composer glyph (`capabilities.workDetect.promptGlyph`)
* @returns false when the screen shows no composer, which is no evidence either way
*/
export function isAwaitingWorkers(paneText: string | null | undefined, pattern: RegExp, promptGlyph: string): boolean {
if (!paneText) return false;
const rows = stripAnsi(paneText)
.split('\n')
.map((row) => row.trimEnd());
let composer = -1;
for (let i = rows.length - 1; i >= 0; i--) {
// Claude has drawn its composer both bare (`❯ …` between rules) and boxed (`│ ❯ … │`).
if (rows[i].replace(/^[\s│]+/, '').startsWith(promptGlyph)) {
composer = i;
break;
}
}
if (composer < 0) return false;
for (let i = composer - 1; i >= Math.max(0, composer - AWAITING_SEARCH_ROWS); i--) {
const row = rows[i];
if (row === '' || /^\s/.test(row) || COMPOSER_FRAME_ROW.test(row)) continue;
// Same reasoning as watchingLabel(): a caller's `g` flag must not make this flap.
pattern.lastIndex = 0;
return pattern.test(row);
}
return false;
}
+32 -1
View File
@@ -85,6 +85,7 @@ import {
isSustainedActivity, isSustainedActivity,
isPaneQuiet, isPaneQuiet,
watchingLabel, watchingLabel,
isAwaitingWorkers,
WATCHING_TAIL_LINES, WATCHING_TAIL_LINES,
IDLE_RECHECK_MS, IDLE_RECHECK_MS,
PANE_PROBE_MIN_INTERVAL_MS, PANE_PROBE_MIN_INTERVAL_MS,
@@ -531,6 +532,8 @@ export class Session extends EventEmitter {
private _watchingLineRe: RegExp | null | undefined = undefined; private _watchingLineRe: RegExp | null | undefined = undefined;
/** Resolved with the pattern above: how many rows at the foot of the screen to search. */ /** Resolved with the pattern above: how many rows at the foot of the screen to search. */
private _watchingWindow = WATCHING_TAIL_LINES; private _watchingWindow = WATCHING_TAIL_LINES;
/** Lazily compiled `capabilities.workDetect.awaitingLine`. See _awaitingLinePattern(). */
private _awaitingLineRe: RegExp | null | undefined = undefined;
private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up) private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up)
private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog
private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read
@@ -3093,7 +3096,10 @@ export class Session extends EventEmitter {
if (now - this._lastPaneProbeAt < PANE_PROBE_MIN_INTERVAL_MS) return this._lastPaneProbeWorking; if (now - this._lastPaneProbeAt < PANE_PROBE_MIN_INTERVAL_MS) return this._lastPaneProbeWorking;
this._lastPaneProbeAt = now; this._lastPaneProbeAt = now;
const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null; const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null;
this._lastPaneProbeWorking = text === null ? null : this._workingLinePattern().test(text); // A turn that ended by handing off to workers the CLI waits for is work too: the
// composer is up and the pane is quiet, but the next turn starts without the user.
this._lastPaneProbeWorking =
text === null ? null : this._workingLinePattern().test(text) || this._paneAwaitsWorkers(text);
this._readWatching(text); this._readWatching(text);
return this._lastPaneProbeWorking; return this._lastPaneProbeWorking;
} }
@@ -3151,6 +3157,21 @@ export class Session extends EventEmitter {
return this._watchingLineRe; return this._watchingLineRe;
} }
/**
* Whether the newest turn on this screen ended waiting for workers the CLI started
* (Claude's `✻ Waiting for 1 dynamic workflow to finish`). False for a CLI whose
* registry entry declares no `awaitingLine`. See `isAwaitingWorkers()`.
*/
private _paneAwaitsWorkers(paneText: string): boolean {
if (this._awaitingLineRe === undefined) {
const src = getCli(this.mode)?.capabilities.workDetect?.awaitingLine;
this._awaitingLineRe = src ? compileVersionRegex(src) : null;
}
if (!this._awaitingLineRe) return false;
const glyph = getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '❯';
return isAwaitingWorkers(paneText, this._awaitingLineRe, glyph);
}
/** /**
* The regex matching this CLI's "a turn is running" status line. * The regex matching this CLI's "a turn is running" status line.
* *
@@ -4186,6 +4207,16 @@ export class Session extends EventEmitter {
return false; return false;
} }
/**
* Arm the composer check for a prompt that went out some other way than
* `writeViaMux`, e.g. cron's paste mode, which writes the body raw and its Enter
* separately. `text` is what the composer line starts with while the prompt is still
* unsent; the check re-presses Enter only while that holds.
*/
verifySubmitted(text: string): void {
this._verifySubmitted(`${text}\r`);
}
/** /**
* Arm the composer check for a write that carried Enter (session-submit-verifier.ts): * Arm the composer check for a write that carried Enter (session-submit-verifier.ts):
* Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints, * Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints,
+61 -7
View File
@@ -6320,10 +6320,15 @@ class CodemanApp {
if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return; if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return;
if (this.detachedSessions?.has(sessionId)) return; if (this.detachedSessions?.has(sessionId)) return;
const session = this.sessions.get(sessionId); const session = this.sessions.get(sessionId);
// A shell's full capture can be many megabytes. Replaying it from an // A shell's full capture can be many megabytes, and replaying all of it from
// ordinary scroll gesture blocks xterm's main thread, so keep that cost // an ordinary scroll gesture blocks xterm's main thread. So a shell scroll
// behind the explicit "Load full history" button. // pulls a BOUNDED window of tmux's full history (the same 1 MiB a tab switch
if (!force && session?.mode === 'shell') return; // loads, but of the scrollback rather than the visible frame) and the
// unbounded pull stays behind the "Load full history" button. Declining
// outright left a shell pane about one screen of browser scrollback after any
// burst, and the button only renders once a replay was truncated, so a young
// shell tab had no way back to output tmux was still holding.
const boundedShellPull = !force && session?.mode === 'shell';
const now = Date.now(); const now = Date.now();
// Momentum scrolling fires this dozens of times per flick, and a burst of new // Momentum scrolling fires this dozens of times per flick, and a burst of new
// output is the normal reason to want a re-pull, so cooldown rather than latch. // output is the normal reason to want a re-pull, so cooldown rather than latch.
@@ -6336,7 +6341,12 @@ class CodemanApp {
this._fullHistoryRepullInFlight = true; this._fullHistoryRepullInFlight = true;
try { try {
const requestStartedAt = performance.now(); const requestStartedAt = performance.now();
const capture = await this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true }); const capture = await this._fetchTerminalCapture(
boundedShellPull
? `/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`
: `/api/sessions/${sessionId}/terminal?full=1`,
{ full: true }
);
const headersReceivedAt = capture.headersAt; const headersReceivedAt = capture.headersAt;
const payload = capture.json?.data ?? {}; const payload = capture.json?.data ?? {};
const bodyParsedAt = performance.now(); const bodyParsedAt = performance.now();
@@ -6357,7 +6367,44 @@ class CodemanApp {
// Bail on a tab switch mid-fetch: writing here would paint another session's // Bail on a tab switch mid-fetch: writing here would paint another session's
// history into the terminal the user is now looking at. // history into the terminal the user is now looking at.
if (!buffer || this.activeSessionId !== sessionId) return; if (!buffer || this.activeSessionId !== sessionId) return;
if (this._replayWouldShrinkBuffer(buffer)) { const windowRows = this._estimateReplayRows(buffer, this.terminal.cols);
// A bounded window no longer than the browser's buffer buys nothing, and
// resetting to rewrite it would jump the viewport on every scroll that
// outlasts the cooldown at the top. This runs BEFORE the downgrade guard
// on purpose: that guard reads "smaller than the browser" as "tmux has
// nothing more to give", which is true of an unbounded capture but not of a
// window cut at the tail size, so a bounded window must never reach the
// exhausted path, which would take Load full history off the banner while
// tmux still holds the rest. Nothing was written here, so the banner state
// is left as the load that produced it set it: re-labelling it from this
// payload would call a terminal that holds ALL of a Load full history pull
// "the most recent 1 MiB".
//
// A browser already at xterm's cap buys nothing either. xterm keeps at most
// `scrollback + rows` rows (DEFAULT_SCROLLBACK 50k) while tmux keeps 100k
// lines by default, so a 1 MiB window of short lines can render to more rows
// than the browser can ever hold, and `windowRows <= rowsNow` then never
// comes true: without this every scroll-to-top would reset and re-parse it.
const rowsNow = this.terminal.buffer.active.length;
const scrollbackCap = this.terminal.options?.scrollback || 0;
const browserFull = scrollbackCap > 0 && rowsNow >= scrollbackCap + this.terminal.rows;
if (boundedShellPull && (windowRows <= rowsNow || browserFull)) {
// An untruncated window IS all of tmux's history, so nothing is missing,
// and the next burst of output can put more in tmux than the browser has:
// keep the normal 4 s cooldown. A truncated one is the opposite case, since
// the gesture can never reach anything older than what the browser already
// shows, and every ask costs the server a synchronous capture-pane of the
// whole history (`tail` is applied after the capture): back off to 60 s.
// A full browser backs off too, since no window can ever fit in it.
// Trade-off: only a successful replay clears that latch, so a tab switch or
// burst that shrinks the browser's buffer below the window can leave a
// scroll-to-top inert for up to a minute. Load full history (`force`)
// bypasses the cooldown, and the latch is bounded, never permanent.
if (payload.truncated || browserFull) (this._fullHistoryRepullUseless ||= new Set()).add(sessionId);
this._logScrollRouting?.('repull-skipped-bounded');
return;
}
if (this._replayWouldShrinkBuffer(buffer, windowRows)) {
timing.refused = true; timing.refused = true;
timing.totalMs = performance.now() - requestStartedAt; timing.totalMs = performance.now() - requestStartedAt;
this._recordTerminalLoadTiming(timing); this._recordTerminalLoadTiming(timing);
@@ -6368,7 +6415,14 @@ class CodemanApp {
this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); this._setHistoryTruncation(sessionId, { ...payload, exhausted: true });
return; return;
} }
this._setHistoryTruncation(sessionId, payload); // A bounded window that was cut is always recoverable: a capture over the
// byte cap keeps `truncationReason: 'capped'` through the tail cut, and that
// would tell the user the rest "cannot be recovered" and drop Load full
// history, whose unbounded pull returns up to the cap itself.
this._setHistoryTruncation(
sessionId,
boundedShellPull && payload.truncated ? { ...payload, truncationReason: 'tail' } : payload
);
this._fullHistoryRepullUseless?.delete(sessionId); this._fullHistoryRepullUseless?.delete(sessionId);
const rowsBefore = this.terminal.buffer.active.length; const rowsBefore = this.terminal.buffer.active.length;
const replayStartedAt = performance.now(); const replayStartedAt = performance.now();
+117 -16
View File
@@ -332,6 +332,39 @@ html.mobile-init .file-browser-panel {
} }
} }
/* Edge fade for the phone's header tab strip (used in the block below).
Registered so the keyframes can interpolate them as lengths; @property is
only valid at the top level, hence out here. */
@property --tab-strip-fade-start {
syntax: '<length>';
inherits: false;
initial-value: 0px;
}
@property --tab-strip-fade-end {
syntax: '<length>';
inherits: false;
initial-value: 0px;
}
/* Driven by the strip's own scroll position, not by time: at the start only the
end edge fades, at the end only the start edge, and in between both. */
@keyframes tab-strip-edge-fade {
0% {
--tab-strip-fade-start: 0px;
--tab-strip-fade-end: 28px;
}
10%,
90% {
--tab-strip-fade-start: 28px;
--tab-strip-fade-end: 28px;
}
100% {
--tab-strip-fade-start: 28px;
--tab-strip-fade-end: 0px;
}
}
/* ============================================================================ /* ============================================================================
Phone Breakpoint (<600px) Phone Breakpoint (<600px)
============================================================================ */ ============================================================================ */
@@ -652,7 +685,7 @@ html.mobile-init .file-browser-panel {
overscroll-behavior-x: contain; overscroll-behavior-x: contain;
scrollbar-width: none; scrollbar-width: none;
max-height: 36px; max-height: 36px;
gap: 2px; gap: 6px;
padding: 0; padding: 0;
} }
@@ -660,6 +693,36 @@ html.mobile-init .file-browser-panel {
display: none; display: none;
} }
/* Fade the strip's edges while there is more to scroll to, so the tab that
does not fit dissolves into the edge instead of being cut mid-word against
the connection dot. Scroll-driven, no JS: the timeline is the strip's own
inline scroll (keyframes + registered properties above this block). A strip
that does not overflow has an INACTIVE timeline, so the animation applies
nothing and both widths stay at their registered 0px, which is no mask at
all. Browsers without scroll timelines skip the block and keep the hard
edge. Header only: in sidebar layout the same list scrolls vertically. */
@supports (animation-timeline: scroll()) {
.header .session-tabs {
-webkit-mask-image: linear-gradient(
to right,
transparent,
#000 var(--tab-strip-fade-start),
#000 calc(100% - var(--tab-strip-fade-end)),
transparent
);
mask-image: linear-gradient(
to right,
transparent,
#000 var(--tab-strip-fade-start),
#000 calc(100% - var(--tab-strip-fade-end)),
transparent
);
/* The shorthand resets animation-timeline, so the timeline comes after. */
animation: tab-strip-edge-fade linear both;
animation-timeline: scroll(self inline);
}
}
/* Smaller tabs for mobile */ /* Smaller tabs for mobile */
.session-tab { .session-tab {
flex-shrink: 0; flex-shrink: 0;
@@ -671,10 +734,46 @@ html.mobile-init .file-browser-panel {
border-radius: 4px; border-radius: 4px;
} }
/* Smaller status indicator on mobile */ /* Every tab in the header strip is a chip, not only the active one. Left
transparent, the strip read as a row of disabled labels: grey 11px text
floating in unmarked gaps, with nothing saying "tap me". Fill and border
come from the skin's control tokens, so the four light skins (which repaint
the header with --glass-bg) get a matching chip with no override block, and
the active tab's !important fill and border in styles.css still win.
`:where(.header)` keeps this at (0,1,0): the per-colour left border
(`.session-tab[data-color="red"]`, (0,2,0)) must still outrank the
border-color here, and in sidebar layout the list leaves the header, so
its rows are untouched. */
:where(.header) .session-tab {
border-radius: 8px;
background: var(--control-bg-hover);
border-color: var(--control-border-hover);
color: var(--text);
}
:where(.header) .session-tab .tab-name {
font-weight: 500;
}
/* Only the active tab shows its action icons on a phone (see below), so on
every other tab the container is empty but still a flex item, and its gap
made the chip visibly wider on the right than on the left. */
:where(.header) .session-tab:not(.active) .tab-actions {
display: none;
}
/* The boxed digit is the Alt+1..9 shortcut hint. A phone has no Alt key, so
here it was only a second grey box inside every tab, and 20px of the name's
width. Every header tab is therefore numberless on a phone, which is the
case the active-tab reserve below is already sized for. */
:where(.header) .session-tab .tab-number {
display: none;
}
/* Status dot: 6px so an idle green reads at arm's length (4px was a speck). */
.session-tab .tab-status { .session-tab .tab-status {
width: 4px; width: 6px;
height: 4px; height: 6px;
} }
/* The working dot is the one glance-state a phone needs: keep idle tiny, but /* The working dot is the one glance-state a phone needs: keep idle tiny, but
@@ -710,9 +809,12 @@ html.mobile-init .file-browser-panel {
opacity: 0.5; opacity: 0.5;
} }
/* Truncate tab names more aggressively on mobile */ /* Truncate tab names on mobile. 80px, not the old 50px: session names share
a `w1-` style prefix, and at 50px "w1-ingest-pipeline" became "w1-inge…"
and a clipped tab just "w1-", which says nothing about which session it is.
The 20px the hidden tab number gave back pays for most of the difference. */
.session-tab .tab-name { .session-tab .tab-name {
max-width: 50px; max-width: 80px;
overflow: hidden; overflow: hidden;
text-overflow: ellipsis; text-overflow: ellipsis;
} }
@@ -726,20 +828,19 @@ html.mobile-init .file-browser-panel {
difference instead, which costs a little strip space on exactly one tab difference instead, which costs a little strip space on exactly one tab
and keeps tap-to-switch the majority of it. and keeps tap-to-switch the majority of it.
⚠️ The floor is set by the 10th tab onward, NOT by the numbered tabs you ⚠️ The floor is set by a NUMBERLESS tab. `.tab-number` is rendered only
are looking at. `.tab-number` is rendered only for `_tabIdx < 9` (app.js), for `_tabIdx < 9` (app.js), and the header hides it on phones altogether
so tab 10 loses 16px + a 4px gap off its left and its centre sits 10px (above), so every phone tab is that case now; a numbered one would sit 10px
further right. The centre clears the icons when further left and hide the problem. The centre clears the icons when
reserved > icons + rightEdge - leftRunUp - gap reserved > icons + rightEdge - leftRunUp - gap
= 50 + 9 - 17 - 4 = 38px = 50 + 9 - 19 - 4 = 36px
with icons = gear 32 + close 20 - close's -2px margin, leftRunUp = border 1 with icons = gear 32 + close 20 - close's -2px margin, leftRunUp = border 1
+ padding 8 + status dot 4 + gap 4, and rightEdge = padding 8 + border 1. + padding 8 + status dot 6 + gap 4, and rightEdge = padding 8 + border 1.
Hit testing snaps to whole pixels, so 39px still lands on the gear: the Hit testing snaps to whole pixels, so a centre half a pixel short still
practical floor is 40px and 44px keeps 4px of headroom. A NUMBERED tab lands on the gear: the practical floor was measured at 40px (with the
clears it at 20px, so reasoning from the tabs on screen is exactly what older 4px dot) and 44px keeps headroom. Pinned by
would put the centre back on the gear. Pinned by
test/mobile-tab-tap-zones.test.ts. */ test/mobile-tab-tap-zones.test.ts. */
.session-tab.active .tab-name { .session-tab.active .tab-name {
min-width: 44px; min-width: 44px;
+40 -18
View File
@@ -650,8 +650,9 @@ Object.assign(CodemanApp.prototype, {
} }
// Mouse wheel: forward to the TUI only for sessions verified to handle SGR // Mouse wheel: forward to the TUI only for sessions verified to handle SGR
// wheel reports (claude 2.1.187+ — see _shouldForwardWheelToApp), local // wheel reports (claude 2.1.187+ while it tracks the mouse, which only its
// scrollback otherwise. Claude Code 2.1.187+ scrolls its own // fullscreen renderer does; see _shouldForwardWheelToApp), local scrollback
// otherwise. Claude Code 2.1.187+ scrolls its own
// transcript on SGR wheel reports — scrolled-away tool blocks re-render // transcript on SGR wheel reports — scrolled-away tool blocks re-render
// live and stay clickable — and its select menus no longer capture wheel // live and stay clickable — and its select menus no longer capture wheel
// as option navigation (verified against 2.1.202: /model menu highlight // as option navigation (verified against 2.1.202: /model menu highlight
@@ -723,7 +724,7 @@ Object.assign(CodemanApp.prototype, {
// phone/tablet swipe scrolls the local buffer of stale repaint frames and // phone/tablet swipe scrolls the local buffer of stale repaint frames and
// drags the CLI's pinned input box off the screen (issue #205's mobile // drags the CLI's pinned input box off the screen (issue #205's mobile
// half). Same gate, so Shift has no touch analog but the local-scrollback // half). Same gate, so Shift has no touch analog but the local-scrollback
// opt-out setting and the CLI-version gate apply to touch exactly as they // opt-out setting and the version/tracking gate apply to touch exactly as they
// do to the wheel — including the PageUp/PageDown fallback the wheel uses // do to the wheel — including the PageUp/PageDown fallback the wheel uses
// when that gate is false and there is no local scrollback to scroll // when that gate is false and there is no local scrollback to scroll
// (_maybePageCliTranscript), which is what keeps a swipe from being a // (_maybePageCliTranscript), which is what keeps a swipe from being a
@@ -3304,9 +3305,9 @@ Object.assign(CodemanApp.prototype, {
/** /**
* Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the * Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the
* buffer while scrolling up gives the app a chance to pull the rest of tmux's * buffer while scrolling up gives the app a chance to pull the rest of tmux's
* scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions decline * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions pull a
* automatic pulls because their captures can be large; their banner button is * bounded window because their captures can be large; their banner button is
* the explicit path. Must be called AFTER scrollLines(), since the check is on * the unbounded path. Must be called AFTER scrollLines(), since the check is on
* the resulting position, and it is deliberately not folded into * the resulting position, and it is deliberately not folded into
* _noteTerminalUserScroll for exactly that reason. * _noteTerminalUserScroll for exactly that reason.
*/ */
@@ -3354,13 +3355,17 @@ Object.assign(CodemanApp.prototype, {
* below the last line, and _estimateReplayRows can only approximate wrapping. * below the last line, and _estimateReplayRows can only approximate wrapping.
* Only a capture that is worse by more than a full screen counts as a * Only a capture that is worse by more than a full screen counts as a
* downgrade, which leaves every genuine recovery case untouched. * downgrade, which leaves every genuine recovery case untouched.
*
* A caller that already estimated the capture's rows passes them as
* `estimatedRows`, so a megabyte capture is not scanned twice.
*/ */
_replayWouldShrinkBuffer(capture) { _replayWouldShrinkBuffer(capture, estimatedRows) {
const term = this.terminal; const term = this.terminal;
const rowsNow = term?.buffer?.active?.length || 0; const rowsNow = term?.buffer?.active?.length || 0;
if (!rowsNow) return false; if (!rowsNow) return false;
const screen = term?.rows || 24; const screen = term?.rows || 24;
return this._estimateReplayRows(capture, term?.cols) + screen < rowsNow; const rows = estimatedRows ?? this._estimateReplayRows(capture, term?.cols);
return rows + screen < rowsNow;
}, },
/** /**
@@ -5082,9 +5087,10 @@ Object.assign(CodemanApp.prototype, {
// Wheel forwarding gate for the container wheel handler: no Shift override, // Wheel forwarding gate for the container wheel handler: no Shift override,
// xterm's own encoder dormant, viewport at the bottom, and a TUI VERIFIED to // xterm's own encoder dormant, viewport at the bottom, and a TUI VERIFIED to
// scroll its transcript on SGR wheel reports — which today is claude 2.1.187+ // scroll its transcript on SGR wheel reports, which today is claude 2.1.187+
// and nothing else (older Claude Code captures wheel as select-menu option // with mouse tracking on (fullscreen) and nothing else (older Claude Code
// navigation; an unknown version is treated as older). Gemini and codex are // captures wheel as select-menu option navigation, an unknown version is
// treated as older, and inline Claude ignores it). Gemini and codex are
// strip modes too but keep the local wheel — taps/clicks are still forwarded // strip modes too but keep the local wheel — taps/clicks are still forwarded
// for them (harmless no-ops at worst). // for them (harmless no-ops at worst).
// //
@@ -5151,6 +5157,15 @@ Object.assign(CodemanApp.prototype, {
const sessionMode = session?.mode || 'claude'; const sessionMode = session?.mode || 'claude';
if (sessionMode !== 'claude') return false; if (sessionMode !== 'claude') return false;
if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false; if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false;
// Only while Claude is actually listening for the mouse. In its default
// inline renderer (2.1.280 measured: alternate_on=0, mouse_any_flag=0) the
// transcript lives in real scrollback, like codex, and SGR wheel reports are
// ignored, so forwarding made every swipe and wheel tick dead. Fullscreen
// (CLAUDE_CODE_NO_FLICKER=1, or "tui": "fullscreen" in ~/.claude/settings.json)
// turns on alt-screen + mode 1003/1006, which the server records as
// cliMouseTracking. A stale-false flag after a server restart falls through
// to _maybePageCliTranscript, so it never goes dead.
if (session?.cliMouseTracking !== true) return false;
// Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so // Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so
// that leaving the bottom handed the wheel back to local scrollback and both // that leaving the bottom handed the wheel back to local scrollback and both
// histories stayed reachable without a mode switch. In practice that inverted // histories stayed reachable without a mode switch. In practice that inverted
@@ -5224,11 +5239,13 @@ Object.assign(CodemanApp.prototype, {
* *
* The rescue path for every way `_shouldForwardWheelToApp` can come back false * The rescue path for every way `_shouldForwardWheelToApp` can come back false
* on a Claude session that has no local history to fall back on: the CLI * on a Claude session that has no local history to fall back on: the CLI
* version probe failed or is genuinely older than 2.1.187, or the user turned * version probe failed or is genuinely older than 2.1.187, the CLI's mouse
* on "Wheel scrolls local history" (which pins the wheel to a buffer that, * tracking flag is unset (the inline renderer, or fullscreen right after a
* for a repaint-mode CLI, is empty — the setting's footgun). Before this, all * server restart), or the user turned on "Wheel scrolls local history" (which
* of those produced a completely dead gesture; the #205 reporter proved the * pins the wheel to a buffer that, for a repaint-mode CLI, is empty: the
* keyboard route works by paging back through intact text with Fn+Up. * setting's footgun). Before this, all of those produced a completely dead
* gesture; the #205 reporter proved the keyboard route works by paging back
* through intact text with Fn+Up.
* *
* Triple-guarded (claude mode + gate false + `baseY === 0`), so a session with * Triple-guarded (claude mode + gate false + `baseY === 0`), so a session with
* real local scrollback is never touched. Shift is excluded on purpose: it is * real local scrollback is never touched. Shift is excluded on purpose: it is
@@ -5272,14 +5289,19 @@ Object.assign(CodemanApp.prototype, {
const session = this.sessions?.get(sessionId); const session = this.sessions?.get(sessionId);
const optOut = !!this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback; const optOut = !!this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback;
const tracking = this.terminal?.modes?.mouseTrackingMode || 'none'; const tracking = this.terminal?.modes?.mouseTrackingMode || 'none';
// xterm's own mode above stays 'none' for a strip mode (the server removes
// the DECSETs), so the CLI's real tracking state is reported separately.
const cliTracking = session?.cliMouseTracking === true;
const baseY = this.terminal?.buffer?.active?.baseY ?? -1; const baseY = this.terminal?.buffer?.active?.baseY ?? -1;
const signature = `${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${baseY > 0}`; const signature =
`${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${cliTracking}|${baseY > 0}`;
if (!this._scrollRoutingLogged) this._scrollRoutingLogged = new Map(); if (!this._scrollRoutingLogged) this._scrollRoutingLogged = new Map();
if (this._scrollRoutingLogged.get(sessionId) === signature) return; if (this._scrollRoutingLogged.get(sessionId) === signature) return;
this._scrollRoutingLogged.set(sessionId, signature); this._scrollRoutingLogged.set(sessionId, signature);
console.log( console.log(
`[scroll] ${sessionId} → ${decision} (mode=${session?.mode || '?'}, cliVersion=${session?.cliVersion || 'unknown'}, ` + `[scroll] ${sessionId} → ${decision} (mode=${session?.mode || '?'}, cliVersion=${session?.cliVersion || 'unknown'}, ` +
`localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, localScrollbackRows=${baseY})` `localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, cliMouseTracking=${cliTracking}, ` +
`localScrollbackRows=${baseY})`
); );
}, },
+20
View File
@@ -437,6 +437,26 @@ export function parseBody<T>(schema: z.ZodType<T>, body: unknown, errorMessage?:
return result.data; return result.data;
} }
/**
* Whether an input body is a plain prompt: printable text followed by exactly one
* carriage return, and nothing else.
*
* That is the shape a script, a bot or a curl call sends to submit a prompt, and the
* one that must NOT be written into the pane in one piece. Measured on Claude Code
* 2.1.283 (2026-09-28): a direct write of `<text>\r` arrives as a single burst, and a
* burst of about a hundred characters or more is taken as a paste, so its trailing
* `\r` lands as a NEWLINE in the composer and the prompt sits there unsent. A later
* bare `\r` written the same way does not recover it; a tmux `send-keys Enter` does.
* Short bursts (tens of characters) submit, which is why the failure looked random.
*
* Anything with another control character (escape sequences, a bracketed-paste frame,
* a line feed, a tab, C1 controls) is raw terminal input and keeps the direct write.
*/
export function isPlainPromptInput(input: string): boolean {
// eslint-disable-next-line no-control-regex -- matching control characters is the point
return /^[^\x00-\x1f\x7f-\x9f]+\r$/.test(input);
}
/** /**
* Persist session state and broadcast a SessionUpdated event. * Persist session state and broadcast a SessionUpdated event.
* Replaces the repeated two-line pattern across route handlers. * Replaces the repeated two-line pattern across route handlers.
+10 -1
View File
@@ -210,13 +210,22 @@ const installsInFlight = new Set<string>();
/** /**
* The server's environment minus every `CODEMAN_*` variable. An install script is third-party * The server's environment minus every `CODEMAN_*` variable. An install script is third-party
* code, and those variables carry Codeman's own secrets and wiring (`CODEMAN_PASSWORD`, the * code, and those variables carry Codeman's own secrets and wiring (`CODEMAN_PASSWORD`, the
* data dir, the tmux socket), none of which an installer needs. * data dir, the tmux socket), none of which an installer needs. Inside the Docker Compose
* container (`CODEMAN_IN_CONTAINER=1`) it also points `NPM_CONFIG_PREFIX` at `$HOME/.local`,
* so an `npm install -g` lands on the persistent home mount instead of the image.
*/ */
export function installEnv(source: NodeJS.ProcessEnv = process.env): NodeJS.ProcessEnv { export function installEnv(source: NodeJS.ProcessEnv = process.env): NodeJS.ProcessEnv {
const env: NodeJS.ProcessEnv = {}; const env: NodeJS.ProcessEnv = {};
for (const [key, value] of Object.entries(source)) { for (const [key, value] of Object.entries(source)) {
if (!key.startsWith('CODEMAN_')) env[key] = value; if (!key.startsWith('CODEMAN_')) env[key] = value;
} }
// ⚠️ In the Docker Compose deployment the image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is IMAGE
// content: `Update-Codeman.sh` recreates the container and every CLI installed there (dsh, pi, ...)
// vanishes. HOME is the persistent bind mount and `~/.local/bin` is already on every resolver's search
// list, so npm-based installs are redirected there. curl|bash installers already target HOME.
if (source.CODEMAN_IN_CONTAINER === '1' && source.HOME) {
env.NPM_CONFIG_PREFIX = `${source.HOME}/.local`;
}
return env; return env;
} }
+20 -1
View File
@@ -91,6 +91,7 @@ import {
findSessionOrFail, findSessionOrFail,
getAuthUser, getAuthUser,
isAdmin, isAdmin,
isPlainPromptInput,
isWorkingDirAllowed, isWorkingDirAllowed,
ownerFor, ownerFor,
parseBody, parseBody,
@@ -1835,10 +1836,18 @@ export function registerSessionRoutes(
// the wrong recovery — wait longer, when the truth is "restart the worker". // the wrong recovery — wait longer, when the truth is "restart the worker".
let delivered = false; let delivered = false;
// A plain prompt (`<text>\r`) goes through the mux even when the caller did not
// ask for it: written straight into the pane it arrives as one burst, and Claude
// Code takes a long burst as a paste whose `\r` becomes a newline, so the prompt
// sat unsent (see isPlainPromptInput). The mux path types the text, presses Enter
// separately and arms the SubmitVerifier. An explicit `useMux: false` keeps the
// raw write for a caller that really wants it.
const autoMux = useMux === undefined && isPlainPromptInput(inputStr);
if (duplicate) { if (duplicate) {
// Redelivery of an already-applied input: skip the write, but still honor the // Redelivery of an already-applied input: skip the write, but still honor the
// wait, since the caller's question ("tell me when this settles") is unanswered. // wait, since the caller's question ("tell me when this settles") is unanswered.
} else if (useMux && waitPromise) { } else if ((useMux || autoMux) && waitPromise) {
// The response is already staying open for the wait, so the tmux write can be // The response is already staying open for the wait, so the tmux write can be
// awaited here. This is the ONE path where a writeViaMux failure is observable. // awaited here. This is the ONE path where a writeViaMux failure is observable.
const ok = await session.writeViaMux(inputStr, { fromUser: true }).catch(() => false); const ok = await session.writeViaMux(inputStr, { fromUser: true }).catch(() => false);
@@ -1849,6 +1858,16 @@ export function registerSessionRoutes(
delivered = session.write(inputStr, { fromUser: true }); delivered = session.write(inputStr, { fromUser: true });
if (!delivered) undoOnFailure(); if (!delivered) undoOnFailure();
} }
} else if (autoMux) {
// Awaited, unlike the explicit useMux branch below. This shape also reaches here
// from the browser's POST fallback, which sends its frames one at a time and
// waits for each 2xx; answering only once Enter has gone out is what keeps the
// next keystroke from overtaking it.
const ok = await session.writeViaMux(inputStr, { fromUser: true }).catch(() => false);
if (!ok) {
console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`);
if (!session.write(inputStr, { fromUser: true })) undoOnFailure();
}
} else if (useMux) { } else if (useMux) {
// Fire-and-forget: don't block the HTTP response on a tmux child process. // Fire-and-forget: don't block the HTTP response on a tmux child process.
// Fallback to a direct write on failure. Unchanged from before send-and-wait. // Fallback to a direct write on failure. Unchanged from before send-and-wait.
+15 -3
View File
@@ -165,7 +165,11 @@ function registerCrudRoutes(app: FastifyInstance, ctx: EventPort & TabLayoutPort
}); });
throw error; throw error;
} }
ctx.broadcast(SseEvent.WebviewChanged, { action: 'created', id: created.id }); ctx.broadcast(SseEvent.WebviewChanged, {
action: 'created',
id: created.id,
owner: ownerLayoutKey(created.owner),
});
return { success: true, data: created }; return { success: true, data: created };
}); });
@@ -195,7 +199,11 @@ function registerCrudRoutes(app: FastifyInstance, ctx: EventPort & TabLayoutPort
// Any edit invalidates the outstanding capability. Otherwise a token minted // Any edit invalidates the outstanding capability. Otherwise a token minted
// against the OLD url keeps proxying to it after the user repointed the tab. // against the OLD url keeps proxying to it after the user repointed the tab.
webviewCapabilities.revokeWebview(id); webviewCapabilities.revokeWebview(id);
ctx.broadcast(SseEvent.WebviewChanged, { action: 'updated', id }); ctx.broadcast(SseEvent.WebviewChanged, {
action: 'updated',
id,
owner: ownerLayoutKey(updated.owner),
});
return { success: true, data: updated }; return { success: true, data: updated };
}); });
@@ -231,7 +239,11 @@ function registerCrudRoutes(app: FastifyInstance, ctx: EventPort & TabLayoutPort
webviewCapabilities.revokeWebview(id); webviewCapabilities.revokeWebview(id);
socketCounts.delete(id); socketCounts.delete(id);
ctx.broadcast(SseEvent.WebviewChanged, { action: 'deleted', id }); ctx.broadcast(SseEvent.WebviewChanged, {
action: 'deleted',
id,
owner: ownerLayoutKey(result.owner),
});
return { success: true, data: { id } }; return { success: true, data: { id } };
}); });
+7
View File
@@ -96,6 +96,7 @@ import { PushSubscriptionStore } from '../push-store.js';
import webpush from 'web-push'; import webpush from 'web-push';
import { SseStreamManager } from './sse-stream-manager.js'; import { SseStreamManager } from './sse-stream-manager.js';
import { deriveTabLayoutSseHint } from './tab-layout-sse.js'; import { deriveTabLayoutSseHint } from './tab-layout-sse.js';
import { deriveWebviewSseHint } from './webview-sse.js';
import { import {
type SessionListenerRefs, type SessionListenerRefs,
createSessionListeners, createSessionListeners,
@@ -2461,6 +2462,12 @@ export class WebServer extends EventEmitter {
if (event.startsWith('tab:')) { if (event.startsWith('tab:')) {
return deriveTabLayoutSseHint(data); return deriveTabLayoutSseHint(data);
} }
// Saved-webview invalidations carry the trusted resource owner. Route them to
// that owner (plus admins), so an admin editing a user's web tab notifies the
// user, and no other user learns the ids of someone else's web tabs.
if (event.startsWith('webview:')) {
return deriveWebviewSseHint(data);
}
// Session-scoped families: resolve the owner from the payload's session id. // Session-scoped families: resolve the owner from the payload's session id.
const SESSION_PREFIXES = [ const SESSION_PREFIXES = [
'session:', 'session:',
+4 -2
View File
@@ -481,8 +481,10 @@ export const AuthPasswordChangeRequired = 'auth:passwordChangeRequired' as const
export const SessionOrderChanged = 'session:orderChanged' as const; export const SessionOrderChanged = 'session:orderChanged' as const;
/** A saved web tab (dashboard URL) was created, updated or deleted. /** A saved web tab (dashboard URL) was created, updated or deleted.
* Payload: `{ action: 'created' | 'updated' | 'deleted', id }`. The client * Payload: `{ action: 'created' | 'updated' | 'deleted', id, owner }`. The client
* re-fetches the list rather than patching from the payload. */ * re-fetches the list rather than patching from the payload. `owner` is the web
* tab's owner (`'@single'` when multi-user mode is off); in multi-user mode the
* event is delivered only to that owner and admins (`deriveWebviewSseHint`). */
export const WebviewChanged = 'webview:changed' as const; export const WebviewChanged = 'webview:changed' as const;
/** Owner-scoped layout invalidation. Payload contains only `{ owner, version }`. */ /** Owner-scoped layout invalidation. Payload contains only `{ owner, version }`. */
export const TabLayoutChanged = 'tab:layoutChanged' as const; export const TabLayoutChanged = 'tab:layoutChanged' as const;
+6
View File
@@ -0,0 +1,6 @@
/** @fileoverview Trusted owner routing metadata for saved-webview invalidations. */
import type { SseRoutingHint } from './sse-stream-manager.js';
export function deriveWebviewSseHint(data: unknown): SseRoutingHint {
return { username: (data as { owner?: string }).owner, sessionScoped: true };
}
@@ -22,6 +22,8 @@ import {
agentImageNpmPackages as mjsPackages, agentImageNpmPackages as mjsPackages,
GIT_HOST_CLI_BUILD_ARGS as mjsGitHostArgs, GIT_HOST_CLI_BUILD_ARGS as mjsGitHostArgs,
gitHostCliBuildArgPairs as mjsGitHostPairs, gitHostCliBuildArgPairs as mjsGitHostPairs,
GIT_IDENTITY_BUILD_ARGS as mjsGitIdentityArgs,
gitIdentityBuildArgPairs as mjsGitIdentityPairs,
} from '../scripts/lib/cli-catalog.mjs'; } from '../scripts/lib/cli-catalog.mjs';
import { import {
agentImageBuildArgPairs as tsPairs, agentImageBuildArgPairs as tsPairs,
@@ -29,6 +31,8 @@ import {
agentImageNpmPackages as tsPackages, agentImageNpmPackages as tsPackages,
GIT_HOST_CLI_BUILD_ARGS as tsGitHostArgs, GIT_HOST_CLI_BUILD_ARGS as tsGitHostArgs,
gitHostCliBuildArgPairs as tsGitHostPairs, gitHostCliBuildArgPairs as tsGitHostPairs,
GIT_IDENTITY_BUILD_ARGS as tsGitIdentityArgs,
gitIdentityBuildArgPairs as tsGitIdentityPairs,
} from '../src/docker-hosts.js'; } from '../src/docker-hosts.js';
const CATALOG = JSON.parse(readFileSync(fileURLToPath(new URL('../config/clis.stock.json', import.meta.url)), 'utf-8')); const CATALOG = JSON.parse(readFileSync(fileURLToPath(new URL('../config/clis.stock.json', import.meta.url)), 'utf-8'));
@@ -145,3 +149,52 @@ describe('optional gh / az in the agent image: both producers pass the same swit
} }
}); });
}); });
describe('Git identity in the agent image: both producers pass the same settings', () => {
it('maps the Git environment variables to matching Dockerfile ARGs', () => {
expect(tsGitIdentityArgs).toEqual(mjsGitIdentityArgs);
expect(tsGitIdentityArgs).toEqual([
['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'],
['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'],
]);
});
it('passes a complete identity and omits an absent identity', () => {
const identity = {
CODEMAN_AGENT_IMAGE_GIT_USER_NAME: 'Ada Lovelace',
CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: 'ada@example.com',
};
const expected: Array<[string, string]> = [
['GIT_USER_NAME', 'Ada Lovelace'],
['GIT_USER_EMAIL', 'ada@example.com'],
];
expect(tsGitIdentityPairs(identity)).toEqual(expected);
expect(mjsGitIdentityPairs(identity)).toEqual(expected);
expect(tsGitIdentityPairs({})).toEqual([]);
expect(mjsGitIdentityPairs({})).toEqual([]);
// The combined argv, not just the helper: the manual build path could drop the identity otherwise.
expect(tsPairs(identity)).toEqual(mjsPairs(CATALOG, identity));
expect(tsPairs(identity)).toEqual(expect.arrayContaining(expected));
});
it('refuses a partial identity in both build paths', () => {
for (const identity of [
{ CODEMAN_AGENT_IMAGE_GIT_USER_NAME: 'Ada Lovelace' },
{ CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: 'ada@example.com' },
]) {
const named = /CODEMAN_AGENT_IMAGE_GIT_USER_NAME and CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL must both be set/;
expect(() => tsGitIdentityPairs(identity)).toThrow(named);
expect(() => mjsGitIdentityPairs(identity)).toThrow(named);
}
});
it('both Dockerfiles configure system Git identity from the build arguments', () => {
for (const file of ['../docker/agent.Dockerfile', '../docker/server.Dockerfile']) {
const dockerfile = readFileSync(fileURLToPath(new URL(file, import.meta.url)), 'utf-8');
expect(dockerfile, file).toMatch(/^ARG GIT_USER_NAME=$/m);
expect(dockerfile, file).toMatch(/^ARG GIT_USER_EMAIL=$/m);
expect(dockerfile, file).toContain('git config --system user.name "${GIT_USER_NAME}"');
expect(dockerfile, file).toContain('git config --system user.email "${GIT_USER_EMAIL}"');
}
});
});
+121
View File
@@ -0,0 +1,121 @@
/**
* @fileoverview scripts/check-browser-test-excludes.mjs: the detection side (which test
* files need a real browser) and the leak computation. The exclusion side is vitest's own
* `vitest list`, which `npm run check:browser-excludes` exercises for real in CI.
*/
import { afterAll, beforeAll, describe, expect, it } from 'vitest';
import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import {
findBrowserTests,
findLeaks,
findTestFiles,
importsBrowserDriver,
listingMatchesTree,
parseVitestFileList,
} from '../scripts/check-browser-test-excludes.mjs';
import { BROWSER_TEST_GLOBS } from '../config/test-suites';
const repoRoot = resolve(import.meta.dirname, '..');
// Fixture sources are assembled from the module name at runtime, so THIS file never contains
// a literal driver import and is not itself flagged by the checker it tests.
const fromImport = (mod: string) => `import { chromium, type Browser } from '${mod}';\n`;
describe('importsBrowserDriver', () => {
it.each([
fromImport('playwright'),
fromImport('playwright-core').replace(/'/g, '"'),
fromImport('@playwright/test'),
fromImport('puppeteer'),
`import type { Page } from '${'playwright'}';`,
`const { chromium } = require('${'playwright'}');`,
`const pw = await import('${'playwright'}');`,
])('flags %s', (src) => {
expect(importsBrowserDriver(src)).toBe(true);
});
it.each([
"import { describe } from 'vitest';",
"// needs ms-playwright's cache dir\nconst dir = '.cache/ms-playwright';",
fromImport('./playwright-helpers'),
fromImport('playwright-extra-thing'),
])('ignores %s', (src) => {
expect(importsBrowserDriver(src)).toBe(false);
});
});
describe('findBrowserTests (fixture tree)', () => {
let root: string;
beforeAll(() => {
root = mkdtempSync(join(tmpdir(), 'codeman-browser-excludes-'));
const put = (rel: string, src: string) => {
mkdirSync(join(root, rel, '..'), { recursive: true });
writeFileSync(join(root, rel), src);
};
put('test/unit.test.ts', "import { it } from 'vitest';\n");
put('test/legacy-name.test.ts', fromImport('playwright'));
put('test/new.browser.test.ts', fromImport('playwright'));
put('test/nested/deep.test.ts', fromImport('puppeteer'));
put('test/helpers/browser.ts', fromImport('playwright')); // not a test file
});
afterAll(() => rmSync(root, { recursive: true, force: true }));
it('finds driver imports by content, recursively, as sorted repo-relative paths', () => {
expect(findBrowserTests(root)).toEqual([
'test/legacy-name.test.ts',
'test/nested/deep.test.ts',
'test/new.browser.test.ts',
]);
});
it('lists every test file, browser-driven or not, in the same form', () => {
expect(findTestFiles(root)).toEqual([
'test/legacy-name.test.ts',
'test/nested/deep.test.ts',
'test/new.browser.test.ts',
'test/unit.test.ts',
]);
});
});
describe('parseVitestFileList + findLeaks', () => {
it('keeps only test paths and normalizes a leading ./', () => {
const out = '\n./test/a.test.ts\ntest/b.test.ts\nsome banner line\n test/c.test.ts \n';
expect([...parseVitestFileList(out)].sort()).toEqual(['test/a.test.ts', 'test/b.test.ts', 'test/c.test.ts']);
});
it('reports exactly the browser tests the CI set still collects', () => {
const ci = new Set(['test/unit.test.ts', 'test/legacy-name.test.ts']);
expect(findLeaks(['test/legacy-name.test.ts', 'test/new.browser.test.ts'], ci)).toEqual([
'test/legacy-name.test.ts',
]);
expect(findLeaks(['test/new.browser.test.ts'], ci)).toEqual([]);
});
it('flags a non-empty listing whose paths never match the tree instead of passing vacuously', () => {
const tree = ['test/legacy-name.test.ts', 'test/unit.test.ts'];
// e.g. a vitest upgrade that starts printing absolute paths: nothing leaks, but only
// because nothing matches, so the checker must refuse rather than report success.
const drifted = parseVitestFileList('/repo/test/legacy-name.test.ts\n/repo/test/unit.test.ts\n');
expect(drifted.size).toBe(2);
expect(findLeaks(['test/legacy-name.test.ts'], drifted)).toEqual([]);
expect(listingMatchesTree(drifted, tree)).toBe(false);
const healthy = parseVitestFileList('test/legacy-name.test.ts\ntest/unit.test.ts\n');
expect(listingMatchesTree(healthy, tree)).toBe(true);
});
});
describe('against this repository', () => {
it('detects every file already listed in BROWSER_TEST_GLOBS', () => {
// If detection stopped recognising a known browser test, the checker would go blind to
// exactly the class of file it exists for.
const detected = new Set(findBrowserTests(repoRoot));
const literals = BROWSER_TEST_GLOBS.filter((g) => !/[*?[{]/.test(g));
expect(literals.length).toBeGreaterThan(0);
for (const file of literals) expect(detected, file).toContain(file);
});
});
+18
View File
@@ -165,6 +165,24 @@ describe('workDetect.workingLine is guarded like every other config regex', () =
expect(compileVersionRegex(src), `${entry.id} declares a watchingLine the guard refuses`).not.toBeNull(); expect(compileVersionRegex(src), `${entry.id} declares a watchingLine the guard refuses`).not.toBeNull();
} }
}); });
it('holds the optional awaitingLine to the same guard', () => {
expectRejected((e) => {
(e.capabilities as Record<string, unknown>).workDetect = {
promptGlyph: '>',
workingLine: 'working',
awaitingLine: '(a+)+b',
};
}, 'it is tested against a pane row every time a session settles');
});
it('accepts every shipped awaitingLine', () => {
for (const entry of STOCK_CLIS) {
const src = entry.capabilities.workDetect?.awaitingLine;
if (!src) continue;
expect(compileVersionRegex(src), `${entry.id} declares an awaitingLine the guard refuses`).not.toBeNull();
}
});
}); });
describe('no shell text can reach the command line', () => { describe('no shell text can reach the command line', () => {
+81 -1
View File
@@ -16,7 +16,13 @@ import { describe, it, expect, beforeEach, vi } from 'vitest';
import { existsSync, mkdtempSync, mkdirSync, writeFileSync, symlinkSync } from 'node:fs'; import { existsSync, mkdtempSync, mkdirSync, writeFileSync, symlinkSync } from 'node:fs';
import { tmpdir } from 'node:os'; import { tmpdir } from 'node:os';
import { join } from 'node:path'; import { join } from 'node:path';
import { CronService, clampCronExternalCliConfigs, type CronDeps } from '../src/cron/cron-service.js'; import {
CronService,
clampCronExternalCliConfigs,
deliverCronPrompt,
type CronDeps,
} from '../src/cron/cron-service.js';
import { CRON_PASTE_ENTER_DELAY_MS } from '../src/config/server-timing.js';
import { CronJobSchema } from '../src/web/schemas.js'; import { CronJobSchema } from '../src/web/schemas.js';
import { MAX_CRON_JOBS } from '../src/config/map-limits.js'; import { MAX_CRON_JOBS } from '../src/config/map-limits.js';
import type { CronJob, CronJobRun } from '../src/types/cron.js'; import type { CronJob, CronJobRun } from '../src/types/cron.js';
@@ -705,3 +711,77 @@ describe('clampCronExternalCliConfigs', () => {
} }
}); });
}); });
/**
* Paste mode used to write `<text>\r` in one piece. Claude Code takes a raw burst of
* about a hundred characters as a paste and turns its `\r` into a newline, so the
* prompt sat unsent on the composer while the run said `prompt_sent`.
*/
describe('deliverCronPrompt', () => {
const PROMPT =
'Reply with only the word ok and nothing else, this sentence is padding to reach about one hundred chars.';
function fakeTarget(ok = true) {
const calls: string[] = [];
const target = {
write: vi.fn((d: string) => {
calls.push(`write:${JSON.stringify(d)}`);
return ok;
}),
writeViaMux: vi.fn(async (d: string) => {
calls.push(`mux:${JSON.stringify(d)}`);
return ok;
}),
verifySubmitted: vi.fn((t: string) => {
calls.push(`verify:${JSON.stringify(t)}`);
}),
};
return { target, calls };
}
const noWait = async (): Promise<void> => {};
it('paste mode writes the text and its Enter separately, then arms the composer check', async () => {
const { target, calls } = fakeTarget();
const waits: number[] = [];
const ok = await deliverCronPrompt(target, PROMPT, 'paste', async (ms) => {
waits.push(ms);
calls.push('wait');
});
expect(ok).toBe(true);
expect(calls).toEqual([
`write:${JSON.stringify(PROMPT)}`,
'wait',
'write:"\\r"',
`verify:${JSON.stringify(PROMPT)}`,
]);
expect(waits).toEqual([CRON_PASTE_ENTER_DELAY_MS]);
expect(target.writeViaMux).not.toHaveBeenCalled();
});
it('never puts the Enter in the same write as the text', async () => {
const { target } = fakeTarget();
await deliverCronPrompt(target, `${PROMPT}\r`, 'paste', noWait);
for (const [data] of target.write.mock.calls) {
expect(data === '\r' || !data.includes('\r')).toBe(true);
}
});
it('typed mode is unchanged: one mux write that carries the Enter', async () => {
const { target, calls } = fakeTarget();
await deliverCronPrompt(target, PROMPT, 'typed', noWait);
expect(calls).toEqual([`mux:${JSON.stringify(`${PROMPT}\r`)}`]);
});
it('reports a session it could not write to, instead of claiming the prompt went out', async () => {
const { target } = fakeTarget(false);
expect(await deliverCronPrompt(target, PROMPT, 'paste', noWait)).toBe(false);
expect(target.verifySubmitted).not.toHaveBeenCalled();
});
});
+3 -4
View File
@@ -110,15 +110,14 @@ describe('docker-compose.yaml cap_add covers what entrypoint.sh and init:true ne
}); });
describe('the runtime-owned CLI prefix never shadows root commands', () => { describe('the runtime-owned CLI prefix never shadows root commands', () => {
it('server.Dockerfile appends /opt/codeman-cli/bin to PATH rather than prepending it', () => { it('server.Dockerfile appends the runtime-writable CLI dirs to PATH rather than prepending them', () => {
const pathLines = dockerfile.split('\n').filter((l) => /^ENV PATH=/.test(l)); const pathLines = dockerfile.split('\n').filter((l) => /^ENV PATH=/.test(l));
expect(pathLines.length).toBeGreaterThan(0); expect(pathLines.length).toBeGreaterThan(0);
for (const line of pathLines) { for (const line of pathLines) {
expect(line, 'a writable prefix ahead of $PATH lets a planted setpriv run as root').not.toMatch( expect(line, 'a writable prefix ahead of $PATH lets a planted setpriv run as root').toMatch(/^ENV PATH=\$PATH:/);
/^ENV PATH=\/opt\/codeman-cli/
);
} }
expect(pathLines).toContain('ENV PATH=$PATH:/opt/codeman-cli/bin'); expect(pathLines).toContain('ENV PATH=$PATH:/opt/codeman-cli/bin');
expect(pathLines).toContain('ENV PATH=$PATH:/home/${CODEMAN_RUNTIME_USER}/.local/bin');
}); });
it('entrypoint.sh pins PATH to the system directories before its first command', () => { it('entrypoint.sh pins PATH to the system directories before its first command', () => {
+468
View File
@@ -0,0 +1,468 @@
/**
* @fileoverview The pre-push hook that scripts/postinstall.js installs (scripts/git-hooks.mjs).
*
* Two properties matter more than the hook's contents, because the older pre-commit
* installer gets both wrong and this one must not copy it:
* 1. It is MARKER-OWNED: a hook the developer wrote by hand is never overwritten.
* 2. The hooks directory is resolved through git, since in a worktree `.git` is a FILE
* and `<root>/.git/hooks` does not exist, and it is ONLY ever the repo's own
* `<git-common-dir>/hooks`: a `core.hooksPath` elsewhere (typically a global one) is
* never written to.
*
* ⚠️ Every filesystem/git test here runs against THROWAWAY repositories under a temp dir.
* Never point the installer at this checkout: its hooks directory is shared with every
* worktree of it, including whatever the developer is running right now.
*/
import { afterAll, beforeAll, describe, expect, it } from 'vitest';
import { execFileSync, spawnSync } from 'node:child_process';
import {
chmodSync,
mkdirSync,
mkdtempSync,
readFileSync,
realpathSync,
rmSync,
statSync,
symlinkSync,
writeFileSync,
} from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import {
PRE_PUSH_CHECKS,
PRE_PUSH_MARKER,
PRE_PUSH_WATCHED_PATHS,
installPrePushHook,
planHookInstall,
renderPrePushHook,
resolveGitHooksDir,
} from '../scripts/git-hooks.mjs';
const repoRoot = resolve(import.meta.dirname, '..');
const read = (rel: string) => readFileSync(resolve(repoRoot, rel), 'utf8');
/** git with no user/system config leaking in (a global core.hooksPath would redirect everything). */
const GIT_ENV = {
...process.env,
GIT_CONFIG_NOSYSTEM: '1',
GIT_CONFIG_GLOBAL: '/dev/null',
GIT_AUTHOR_NAME: 'test',
GIT_AUTHOR_EMAIL: 'test@example.invalid',
GIT_COMMITTER_NAME: 'test',
GIT_COMMITTER_EMAIL: 'test@example.invalid',
CODEMAN_SKIP_PREPUSH: '',
};
function git(cwd: string, args: string[], env: NodeJS.ProcessEnv = GIT_ENV): string {
return execFileSync('git', args, { cwd, env, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] }).trim();
}
let scratch: string;
beforeAll(() => {
scratch = realpathSync(mkdtempSync(join(tmpdir(), 'codeman-git-hooks-')));
});
afterAll(() => {
rmSync(scratch, { recursive: true, force: true });
});
let counter = 0;
function newRepo(): string {
const dir = join(scratch, `repo-${++counter}`);
mkdirSync(dir, { recursive: true });
git(dir, ['init', '-q', '-b', 'main']);
git(dir, ['commit', '-q', '--allow-empty', '-m', 'init']);
return dir;
}
describe('pre-push hook body', () => {
const hook = renderPrePushHook();
it('carries the ownership marker', () => {
expect(hook).toContain(PRE_PUSH_MARKER);
});
it('runs every configured check through npm, and nothing slow', () => {
for (const args of PRE_PUSH_CHECKS) {
expect(hook).toContain(`run_check ${args.join(' ')}`);
}
expect(hook).toContain('npm run --silent "$@"');
// The whole point of the tier: the minutes-long suites stay out of a per-push hook.
expect(hook).not.toMatch(/\btest:(ci|browser|mobile|perf|all)\b/);
});
it('is POSIX sh', () => {
expect(hook.startsWith('#!/bin/sh\n')).toBe(true);
const r = spawnSync('sh', ['-n'], { input: hook });
expect(r.status).toBe(0);
});
});
describe('pre-push checks match the static CI job', () => {
const scripts = JSON.parse(read('package.json')).scripts as Record<string, string>;
const ci = read('.github/workflows/ci.yml');
it.each(PRE_PUSH_CHECKS.map((args) => [args.join(' ')] as const))('%s is a real script that CI runs', (joined) => {
const [name] = joined.split(' ');
expect(scripts[name], `package.json has no "${name}" script`).toBeTypeOf('string');
expect(ci).toContain(`npm run ${joined}`);
});
});
describe('planHookInstall', () => {
const hook = renderPrePushHook();
it('writes when no hook exists', () => {
expect(planHookInstall({ existing: null, next: hook })).toBe('write');
});
it('refuses to clobber a hook it does not own', () => {
expect(planHookInstall({ existing: '#!/bin/sh\nmake lint\n', next: hook })).toBe('skip-foreign');
});
it('refreshes its own hook when the body changed', () => {
expect(planHookInstall({ existing: `#!/bin/sh\n${PRE_PUSH_MARKER}\necho old\n`, next: hook })).toBe('write');
});
it('is idempotent when already current', () => {
expect(planHookInstall({ existing: hook, next: hook })).toBe('up-to-date');
});
it('treats an empty file as absent rather than foreign', () => {
expect(planHookInstall({ existing: ' \n', next: hook })).toBe('write');
});
});
describe('resolveGitHooksDir (temp repos)', () => {
// resolveGitHooksDir runs git with process.env, so an exported GIT_CONFIG_GLOBAL or a system
// gitconfig carrying core.hooksPath would otherwise redirect every expectation below.
// test/setup.ts swaps HOME, which only covers ~/.gitconfig.
const ambient = {
GIT_CONFIG_GLOBAL: process.env.GIT_CONFIG_GLOBAL,
GIT_CONFIG_NOSYSTEM: process.env.GIT_CONFIG_NOSYSTEM,
};
beforeAll(() => {
process.env.GIT_CONFIG_NOSYSTEM = '1';
process.env.GIT_CONFIG_GLOBAL = '/dev/null';
});
afterAll(() => {
for (const [k, v] of Object.entries(ambient)) {
if (v === undefined) delete process.env[k];
else process.env[k] = v;
}
});
it('resolves <root>/.git/hooks in a plain checkout', () => {
const repo = newRepo();
expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks'));
});
it('resolves the SHARED hooks dir from a worktree, where .git is a file', () => {
const repo = newRepo();
const wt = join(scratch, `wt-${counter}`);
git(repo, ['worktree', 'add', '-q', wt, '-b', 'wt-branch']);
expect(statSync(join(wt, '.git')).isFile()).toBe(true);
expect(resolveGitHooksDir(wt)).toBe(join(repo, '.git', 'hooks'));
});
it('returns null outside any git checkout', () => {
const dir = join(scratch, `plain-${++counter}`);
mkdirSync(dir);
expect(resolveGitHooksDir(dir)).toBeNull();
});
it('returns null when a repo-local core.hooksPath points outside the repo', () => {
const repo = newRepo();
const outside = join(scratch, `shared-hooks-${counter}`);
mkdirSync(outside);
git(repo, ['config', 'core.hooksPath', outside]);
expect(resolveGitHooksDir(repo)).toBeNull();
});
it('returns null when core.hooksPath points at a directory that does not exist yet', () => {
const repo = newRepo();
git(repo, ['config', 'core.hooksPath', join(scratch, `missing-${counter}`, 'hooks')]);
expect(resolveGitHooksDir(repo)).toBeNull();
});
it("still resolves when core.hooksPath points at the repo's OWN .git/hooks", () => {
const repo = newRepo();
git(repo, ['config', 'core.hooksPath', join(repo, '.git', 'hooks')]);
expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks'));
});
it('resolves before .git/hooks exists (compares the would-be path)', () => {
const repo = newRepo();
rmSync(join(repo, '.git', 'hooks'), { recursive: true, force: true });
expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks'));
});
it('returns null under a GLOBAL core.hooksPath, from a checkout and from a worktree', () => {
const repo = newRepo();
const wt = join(scratch, `wt-global-${counter}`);
git(repo, ['worktree', 'add', '-q', wt, '-b', 'wt-global']);
const globalHooks = join(scratch, `global-hooks-${counter}`);
mkdirSync(globalHooks);
const globalConfig = join(scratch, `gitconfig-${counter}`);
writeFileSync(globalConfig, `[core]\n\thooksPath = ${globalHooks}\n`);
// resolveGitHooksDir runs git with the ambient environment, so scope the fake global
// config to this test through process.env (never the developer's real ~/.gitconfig).
const saved = {
GIT_CONFIG_GLOBAL: process.env.GIT_CONFIG_GLOBAL,
GIT_CONFIG_NOSYSTEM: process.env.GIT_CONFIG_NOSYSTEM,
};
process.env.GIT_CONFIG_GLOBAL = globalConfig;
process.env.GIT_CONFIG_NOSYSTEM = '1';
try {
expect(git(repo, ['rev-parse', '--git-path', 'hooks'], { ...GIT_ENV, GIT_CONFIG_GLOBAL: globalConfig })).toBe(
globalHooks
);
expect(resolveGitHooksDir(repo)).toBeNull();
expect(resolveGitHooksDir(wt)).toBeNull();
} finally {
for (const [k, v] of Object.entries(saved)) {
if (v === undefined) delete process.env[k];
else process.env[k] = v;
}
}
// Control: the same repo resolves again once the global setting is gone.
expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks'));
});
it("returns null for a copy nested inside someone else's repo (e.g. under node_modules)", () => {
const repo = newRepo();
const nested = join(repo, 'node_modules', 'aicodeman');
mkdirSync(nested, { recursive: true });
expect(resolveGitHooksDir(nested)).toBeNull();
});
});
describe('installPrePushHook (temp repos)', () => {
it('writes an executable hook into a fresh repo', () => {
const hooks = join(newRepo(), '.git', 'hooks');
expect(installPrePushHook(hooks)).toBe('write');
const path = join(hooks, 'pre-push');
expect(readFileSync(path, 'utf8')).toBe(renderPrePushHook());
expect(statSync(path).mode & 0o111).not.toBe(0);
expect(installPrePushHook(hooks)).toBe('up-to-date');
});
it('leaves a foreign pre-push hook byte-identical', () => {
const hooks = join(newRepo(), '.git', 'hooks');
const path = join(hooks, 'pre-push');
const mine = '#!/bin/sh\n# my own hook\nexit 0\n';
writeFileSync(path, mine, { mode: 0o755 });
expect(installPrePushHook(hooks)).toBe('skip-foreign');
expect(readFileSync(path, 'utf8')).toBe(mine);
});
it('refreshes a stale managed hook and keeps it executable', () => {
const hooks = join(newRepo(), '.git', 'hooks');
const path = join(hooks, 'pre-push');
writeFileSync(path, `#!/bin/sh\n${PRE_PUSH_MARKER}\necho old\n`, { mode: 0o644 });
expect(installPrePushHook(hooks)).toBe('write');
expect(readFileSync(path, 'utf8')).toBe(renderPrePushHook());
expect(statSync(path).mode & 0o111).not.toBe(0);
});
});
/**
* Drive the rendered hook through a real `git push` to a local bare remote. The repo gets a
* stub package.json whose check scripts only record that they ran, so this exercises the
* hook's control flow (ref parsing, skips, blocking) without running the real checks.
*/
describe('the installed hook on a real push (temp repos)', () => {
function setup(opts: { failing?: string; nodeModules?: boolean } = {}) {
const repo = newRepo();
const remote = join(scratch, `remote-${counter}.git`);
git(scratch, ['init', '-q', '--bare', remote]);
git(repo, ['remote', 'add', 'origin', remote]);
const log = join(repo, 'ran.log');
const scripts: Record<string, string> = {};
for (const [name] of PRE_PUSH_CHECKS) {
scripts[name] =
name === opts.failing ? `echo ${name} >> ran.log && echo boom-${name} && exit 1` : `echo ${name} >> ran.log`;
}
writeFileSync(join(repo, 'package.json'), JSON.stringify({ name: 'hook-fixture', private: true, scripts }));
writeFileSync(join(repo, '.gitignore'), 'node_modules/\nran.log\n');
git(repo, ['add', 'package.json', '.gitignore']);
git(repo, ['commit', '-q', '-m', 'fixture']);
if (opts.nodeModules !== false) mkdirSync(join(repo, 'node_modules'));
installPrePushHook(join(repo, '.git', 'hooks'));
chmodSync(join(repo, '.git', 'hooks', 'pre-push'), 0o755);
const ran = () => {
try {
return readFileSync(log, 'utf8').trim().split('\n').filter(Boolean);
} catch {
return [];
}
};
const push = (args: string[], env: NodeJS.ProcessEnv = {}) =>
spawnSync('git', ['push', ...args], { cwd: repo, env: { ...GIT_ENV, ...env }, encoding: 'utf8' });
return { repo, remote, ran, push };
}
/** What the stubs record: npm appends the args after `--` to the script, so they prove forwarding. */
const expectedRuns = PRE_PUSH_CHECKS.map((args) => args.filter((a) => a !== '--').join(' '));
it('runs every check before a push, in order', () => {
const { ran, push } = setup();
const r = push(['-q', 'origin', 'main']);
expect(r.status, r.stderr + r.stdout).toBe(0);
expect(ran()).toEqual(expectedRuns);
});
it('blocks the push when a check fails, but still runs the rest', () => {
const { ran, push, remote } = setup({ failing: 'lint' });
const r = push(['origin', 'main']);
expect(r.status).not.toBe(0);
expect(r.stdout + r.stderr).toContain('pre-push: FAILED npm run lint');
expect(r.stdout + r.stderr).toContain('boom-lint');
expect(ran()).toEqual(expectedRuns);
expect(spawnSync('git', ['rev-parse', '--verify', '-q', 'refs/heads/main'], { cwd: remote }).status).not.toBe(0);
});
it('CODEMAN_SKIP_PREPUSH=1 skips every check', () => {
const { ran, push } = setup({ failing: 'lint' });
const r = push(['-q', 'origin', 'main'], { CODEMAN_SKIP_PREPUSH: '1' });
expect(r.status, r.stderr).toBe(0);
expect(ran()).toEqual([]);
});
it('a delete-only push skips the checks', () => {
const { ran, push, repo } = setup({ failing: 'lint' });
expect(push(['-q', 'origin', 'main'], { CODEMAN_SKIP_PREPUSH: '1' }).status).toBe(0);
git(repo, ['branch', 'doomed']);
expect(push(['-q', 'origin', 'doomed'], { CODEMAN_SKIP_PREPUSH: '1' }).status).toBe(0);
const r = push(['-q', 'origin', '--delete', 'doomed']);
expect(r.status, r.stderr).toBe(0);
expect(ran()).toEqual([]);
});
it('skips when the pushed ref is not the checked-out HEAD', () => {
const { ran, push, repo } = setup({ failing: 'lint' });
git(repo, ['branch', 'other']);
git(repo, ['commit', '-q', '--allow-empty', '-m', 'only on main']);
git(repo, ['checkout', '-q', 'other']);
// HEAD is `other`; pushing `main` would check a working tree that is not main's.
const r = push(['origin', 'main']);
expect(r.status, r.stderr).toBe(0);
expect(r.stdout + r.stderr).toContain(
'pre-push: skipping static checks: refs/heads/main is not the checked-out HEAD'
);
expect(ran()).toEqual([]);
});
it('skips when any one of several pushed refs is not HEAD', () => {
const { ran, push, repo } = setup({ failing: 'lint' });
git(repo, ['branch', 'behind']);
git(repo, ['commit', '-q', '--allow-empty', '-m', 'ahead']);
const r = push(['origin', 'main', 'behind']);
expect(r.status, r.stderr).toBe(0);
expect(r.stdout + r.stderr).toContain('is not the checked-out HEAD');
expect(ran()).toEqual([]);
});
it('still checks an annotated tag that points at HEAD (the tag is peeled)', () => {
const { ran, push, repo } = setup();
git(repo, ['tag', '-a', 'v1', '-m', 'v1']);
const r = push(['-q', 'origin', 'v1']);
expect(r.status, r.stderr + r.stdout).toBe(0);
expect(ran()).toEqual(expectedRuns);
});
it.each([
'src/wip.ts',
'config/wip.json',
'scripts/wip.mjs',
'test/wip.test.ts',
'install.sh',
'tsconfig.json',
'.prettierignore',
'.editorconfig',
])('skips when %s is untracked (another session may own it)', (rel) => {
const { ran, push, repo } = setup({ failing: 'lint' });
mkdirSync(join(repo, rel, '..'), { recursive: true });
writeFileSync(join(repo, rel), 'wip\n');
const r = push(['origin', 'main']);
expect(r.status, r.stderr).toBe(0);
expect(r.stdout + r.stderr).toContain('pre-push: skipping static checks: uncommitted changes under');
expect(ran()).toEqual([]);
});
it('skips when a tracked package.json has an unstaged edit', () => {
const { ran, push, repo } = setup({ failing: 'lint' });
const pkg = join(repo, 'package.json');
writeFileSync(pkg, readFileSync(pkg, 'utf8') + '\n');
const r = push(['origin', 'main']);
expect(r.status, r.stderr).toBe(0);
expect(r.stdout + r.stderr).toContain('uncommitted changes under');
expect(ran()).toEqual([]);
});
it('still checks when the only uncommitted changes are outside the watched paths', () => {
const { ran, push, repo } = setup({ failing: 'lint' });
mkdirSync(join(repo, 'docs'));
writeFileSync(join(repo, 'docs', 'notes.md'), 'draft\n');
writeFileSync(join(repo, 'README.md'), 'draft\n');
const r = push(['origin', 'main']);
expect(r.status).not.toBe(0);
expect(r.stdout + r.stderr).toContain('pre-push: FAILED npm run lint');
expect(ran()).toEqual(expectedRuns);
});
it('watches exactly the paths the checks read', () => {
expect(PRE_PUSH_WATCHED_PATHS).toEqual([
'src',
'config',
'scripts',
'test',
'package.json',
'package-lock.json',
'install.sh',
'tsconfig.json',
'.prettierignore',
'.editorconfig',
]);
});
it('skips (never blocks) when node_modules is absent', () => {
const { ran, push } = setup({ failing: 'lint', nodeModules: false });
const r = push(['origin', 'main']);
expect(r.status, r.stderr).toBe(0);
expect(r.stdout + r.stderr).toContain('node_modules missing');
expect(ran()).toEqual([]);
});
it('skips (never blocks) when npm is not on PATH, as under a GUI git client', () => {
const { ran, push } = setup({ failing: 'lint' });
// A PATH holding only what git and the hook need, and no npm/node. Symlinks rather than
// the real directories, since /usr/bin usually holds npm right next to git.
const bin = join(scratch, `bin-${counter}`);
mkdirSync(bin);
for (const tool of ['git', 'sh', 'mktemp', 'tail', 'rm', 'cat']) {
const found = spawnSync('sh', ['-c', `command -v ${tool}`], { encoding: 'utf8' }).stdout.trim();
expect(found, `${tool} not found on the test PATH`).toMatch(/^\//);
symlinkSync(found, join(bin, tool));
}
expect(spawnSync('sh', ['-c', 'command -v npm'], { env: { PATH: bin } }).status).not.toBe(0);
const r = push(['origin', 'main'], { PATH: bin });
expect(r.status, r.stderr + r.stdout).toBe(0);
expect(r.stdout + r.stderr).toContain('pre-push: npm not on PATH, skipping checks.');
expect(ran()).toEqual([]);
});
});
describe('postinstall wiring', () => {
const postinstall = read('scripts/postinstall.js');
it('installs the pre-push hook through the shared module', () => {
expect(postinstall).toContain("import('./git-hooks.mjs')");
expect(postinstall).toContain('installPrePushHook(gitHooksDir)');
});
it('resolves the hooks dir through git, so worktrees work', () => {
expect(postinstall).toContain('resolveGitHooksDir(');
expect(postinstall).not.toContain("join(import.meta.dirname, '..', '.git', 'hooks')");
});
});
+6 -3
View File
@@ -122,7 +122,7 @@ describe('the in-terminal truncation line is gone (static guard)', () => {
expect(app).not.toContain('earlier output truncated for performance'); expect(app).not.toContain('earlier output truncated for performance');
}); });
it('loads a bounded shell tail first and keeps full history user-triggered', () => { it('loads a bounded shell tail first and keeps unbounded full history user-triggered', () => {
const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8');
expect(app).toContain("session?.mode !== 'shell' && !this._fullHistoryLoaded.has(sessionId)"); expect(app).toContain("session?.mode !== 'shell' && !this._fullHistoryLoaded.has(sessionId)");
expect(app).toContain("!restoredSnapshot && session?.mode !== 'shell'"); expect(app).toContain("!restoredSnapshot && session?.mode !== 'shell'");
@@ -131,10 +131,13 @@ describe('the in-terminal truncation line is gone (static guard)', () => {
// an abort deadline (a `?full=1` body can be megabytes and used to hang // an abort deadline (a `?full=1` body can be megabytes and used to hang
// indefinitely on a stalled mobile link). The URL and the full-vs-tail // indefinitely on a stalled mobile link). The URL and the full-vs-tail
// decision this guard exists to pin are unchanged. // decision this guard exists to pin are unchanged.
expect(app).toContain('this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true })'); expect(app).toContain(': `/api/sessions/${sessionId}/terminal?full=1`,\n { full: true }');
expect(app).toContain("if (this.sessions.get(sessionId)?.mode !== 'shell')"); expect(app).toContain("if (this.sessions.get(sessionId)?.mode !== 'shell')");
expect(app).toContain("if (session?.mode === 'shell')"); expect(app).toContain("if (session?.mode === 'shell')");
expect(app).toContain("if (!force && session?.mode === 'shell') return;"); // A shell scroll gesture pulls a BOUNDED window of full history; only the
// button pulls all of it (behaviour pinned in shell-scroll-history-pull.test.ts).
expect(app).toContain("const boundedShellPull = !force && session?.mode === 'shell';");
expect(app).toContain('`/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`');
expect(app).toContain("trigger: force ? 'full-history-button' : 'full-history-scroll'"); expect(app).toContain("trigger: force ? 'full-history-button' : 'full-history-scroll'");
}); });
+166
View File
@@ -0,0 +1,166 @@
/**
* @fileoverview The phone header's tab strip must read as live tabs.
*
* It used to render every inactive tab transparent: grey 11px text floating in
* unmarked gaps, a boxed Alt+N digit in each (a phone has no Alt key), names cut
* to 50px so a shared `w1-` prefix was most of what showed, and the tab that did
* not fit chopped mid-word against the connection dot. On a phone it looked like
* a row of disabled labels.
*
* The fix is four small rules in the phone block of mobile.css, and each has a
* way to be silently undone, which is what this file fences:
*
* - The chip rule is written `:where(.header) .session-tab` so it stays at
* (0,1,0). Written `.header .session-tab` it would be (0,2,0), tie with the
* per-colour `.session-tab[data-color="red"]` left border in styles.css, and
* win on source order (mobile.css loads later): every colour-tagged tab would
* lose its identity stripe.
* - The edge fade is scroll-DRIVEN (no JS). Its two widths must be registered
* with @property to interpolate, and @property is only valid at the top
* level: nested inside the phone @media it is dropped, the keyframes stop
* interpolating, and the fade snaps between states instead of following the
* scroll position.
* - `animation` is a shorthand that resets `animation-timeline`, so the
* timeline must be declared AFTER it or the fade silently becomes a 0s time
* animation.
*
* Parsed with postcss because the declarations live in nested at-rules. The
* rendered result (chips on dark and light skins, the fade at both scroll ends)
* was checked in a browser; this is the cheap regression fence. Port: N/A.
*/
import { readFileSync } from 'node:fs';
import { resolve } from 'node:path';
import postcss, { type AtRule, type Declaration, type Rule } from 'postcss';
import { describe, expect, it } from 'vitest';
const CSS = readFileSync(resolve(import.meta.dirname, '../src/web/public/mobile.css'), 'utf8');
const ROOT = postcss.parse(CSS);
const PHONE_QUERY = '(max-width: 599px)';
/** Declarations of the rule matching `selector` inside the phone block (later rules win). */
function phoneDeclarations(selector: string): Record<string, string> {
const found: Record<string, string> = {};
ROOT.walkAtRules('media', (atRule) => {
if (atRule.params !== PHONE_QUERY) return;
atRule.walkRules((rule: Rule) => {
if (!rule.selectors.map((s) => s.trim()).includes(selector)) return;
rule.walkDecls((decl: Declaration) => {
found[decl.prop] = decl.value.trim();
});
});
});
return found;
}
/** The `@property` rule for `name`, wherever it sits. */
function propertyRule(name: string): AtRule | undefined {
let hit: AtRule | undefined;
ROOT.walkAtRules('property', (atRule) => {
if (atRule.params.trim() === name) hit = atRule;
});
return hit;
}
function declsOf(node: AtRule | Rule): Record<string, string> {
const out: Record<string, string> = {};
node.each((child) => {
if (child.type === 'decl') out[child.prop] = child.value.trim();
});
return out;
}
describe('phone header tab strip', () => {
describe('chips', () => {
const chip = phoneDeclarations(':where(.header) .session-tab');
it('gives every header tab a fill and border from the skin control tokens', () => {
// Tokens, not literals: the four light skins repaint the header with
// --glass-bg and define their own --control-* values.
expect(chip.background).toMatch(/^var\(--control-bg/);
expect(chip['border-color']).toMatch(/^var\(--control-border/);
expect(chip.color).toBe('var(--text)');
});
it('keeps the chip selector at (0,1,0) so per-colour borders still win', () => {
// The lookup above only matches the exact `:where(.header)` spelling, so a
// rewrite to `.header .session-tab` leaves it empty and fails here.
expect(Object.keys(chip).length).toBeGreaterThan(0);
expect(phoneDeclarations('.header .session-tab')).toEqual({});
});
it('hides the Alt+N digit, which a phone has no key for', () => {
expect(phoneDeclarations(':where(.header) .session-tab .tab-number').display).toBe('none');
});
it('drops the empty action container on inactive tabs only', () => {
// The active tab's gear and close live in .tab-actions, so the rule must
// stay scoped to :not(.active).
expect(phoneDeclarations(':where(.header) .session-tab:not(.active) .tab-actions').display).toBe('none');
expect(phoneDeclarations(':where(.header) .session-tab .tab-actions')).toEqual({});
});
it('leaves enough name to get past a shared w1- prefix', () => {
const maxWidth = Number.parseFloat(phoneDeclarations('.session-tab .tab-name')['max-width'] ?? '');
expect(maxWidth).toBeGreaterThanOrEqual(72);
});
});
describe('scroll-driven edge fade', () => {
it('registers both fade widths at the top level, as lengths starting at 0px', () => {
for (const name of ['--tab-strip-fade-start', '--tab-strip-fade-end']) {
const rule = propertyRule(name);
expect(rule, `${name} is not registered`).toBeDefined();
// Nested in @media it is invalid and silently ignored.
expect(rule!.parent?.type, `${name} must be top level`).toBe('root');
const d = declsOf(rule!);
expect(d.syntax).toBe("'<length>'");
expect(d['initial-value']).toBe('0px');
}
});
it('fades only the far edge at the start and only the near edge at the end', () => {
let frames: Record<string, Record<string, string>> = {};
ROOT.walkAtRules('keyframes', (atRule) => {
if (atRule.params.trim() !== 'tab-strip-edge-fade') return;
frames = {};
atRule.each((node) => {
if (node.type !== 'rule') return;
for (const sel of node.selectors) frames[sel.trim()] = declsOf(node);
});
});
expect(frames['0%']?.['--tab-strip-fade-start']).toBe('0px');
expect(Number.parseFloat(frames['0%']?.['--tab-strip-fade-end'] ?? '0')).toBeGreaterThan(0);
expect(frames['100%']?.['--tab-strip-fade-end']).toBe('0px');
expect(Number.parseFloat(frames['100%']?.['--tab-strip-fade-start'] ?? '0')).toBeGreaterThan(0);
});
it('masks the header strip behind a scroll-timeline feature check, timeline after the shorthand', () => {
let strip: Rule | undefined;
ROOT.walkAtRules('media', (media) => {
if (media.params !== PHONE_QUERY) return;
media.walkAtRules('supports', (supports) => {
if (!/animation-timeline:\s*scroll\(\)/.test(supports.params)) return;
supports.walkRules((rule) => {
if (rule.selectors.map((s) => s.trim()).includes('.header .session-tabs')) strip = rule;
});
});
});
expect(strip, 'no @supports-gated .header .session-tabs rule in the phone block').toBeDefined();
const props: string[] = [];
const d: Record<string, string> = {};
strip!.each((node) => {
if (node.type !== 'decl') return;
props.push(node.prop);
d[node.prop] = node.value.replace(/\s+/g, ' ').trim();
});
for (const prop of ['mask-image', '-webkit-mask-image']) {
expect(d[prop]).toContain('var(--tab-strip-fade-start)');
expect(d[prop]).toContain('var(--tab-strip-fade-end)');
}
expect(d.animation).toContain('tab-strip-edge-fade');
expect(d['animation-timeline']).toBe('scroll(self inline)');
expect(props.indexOf('animation-timeline')).toBeGreaterThan(props.indexOf('animation'));
});
});
});
+4 -3
View File
@@ -88,9 +88,10 @@ describe('Tab Navigation', () => {
if (tabNameExists) { if (tabNameExists) {
const maxWidth = await getCSSProperty(page, SELECTORS.TAB_NAME, 'max-width'); const maxWidth = await getCSSProperty(page, SELECTORS.TAB_NAME, 'max-width');
const maxWidthPx = parseFloat(maxWidth); const maxWidthPx = parseFloat(maxWidth);
// Should be 50px on mobile // 80px on phones: wide enough to get past a shared `w1-` prefix,
expect(maxWidthPx).toBeLessThanOrEqual(60); // still short enough that several tabs fit the strip.
expect(maxWidthPx).toBeGreaterThan(0); expect(maxWidthPx).toBeLessThanOrEqual(96);
expect(maxWidthPx).toBeGreaterThanOrEqual(72);
} }
}); });
+6
View File
@@ -713,6 +713,12 @@ describe('registry writes are serialized and never clobber a file the reader wou
} }
expect(installEnv({ CODEMAN_PASSWORD: 'x', HOME: '/h' })).toEqual({ HOME: '/h' }); expect(installEnv({ CODEMAN_PASSWORD: 'x', HOME: '/h' })).toEqual({ HOME: '/h' });
}); });
it('redirects npm installs to the persistent HOME inside the Compose container', () => {
expect(
installEnv({ CODEMAN_IN_CONTAINER: '1', HOME: '/home/codeman', NPM_CONFIG_PREFIX: '/opt/codeman-cli' })
).toEqual({ HOME: '/home/codeman', NPM_CONFIG_PREFIX: '/home/codeman/.local' });
});
}); });
/** /**
+105 -1
View File
@@ -14,11 +14,12 @@
import fastifyCookie from '@fastify/cookie'; import fastifyCookie from '@fastify/cookie';
import Fastify, { type FastifyInstance } from 'fastify'; import Fastify, { type FastifyInstance } from 'fastify';
import { afterEach, beforeEach, describe, expect, it } from 'vitest'; import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest';
import { Session } from '../../src/session.js'; import { Session } from '../../src/session.js';
import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js'; import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js';
import { installRouteErrorHandler } from '../../src/web/route-error-handler.js'; import { installRouteErrorHandler } from '../../src/web/route-error-handler.js';
import { isPlainPromptInput } from '../../src/web/route-helpers.js';
import { registerSessionRoutes } from '../../src/web/routes/session-routes.js'; import { registerSessionRoutes } from '../../src/web/routes/session-routes.js';
import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js'; import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js';
@@ -155,3 +156,106 @@ describe('POST /api/sessions/:id/input rollback wiring', () => {
expect(session.shouldApplyInput('c2', 5)).toBe(true); expect(session.shouldApplyInput('c2', 5)).toBe(true);
}); });
}); });
/**
* A plain prompt goes through the mux even when the caller did not say `useMux`.
*
* Measured on Claude Code 2.1.283: a direct write of `<text>\r` arrives as one burst,
* a burst of about a hundred characters is taken as a paste, and its `\r` lands as a
* newline in the composer, so a script's prompt sat there unsent while the route
* answered 200. The mux path types the text, presses Enter on its own, and arms the
* submit verifier.
*/
describe('POST /api/sessions/:id/input plain-prompt routing', () => {
let harness: { app: FastifyInstance; ctx: MockRouteContext };
beforeEach(async () => {
harness = await createEnvelopeHarness();
});
afterEach(async () => {
await harness.app.close();
});
const post = (body: Record<string, unknown>) =>
harness.app.inject({ method: 'POST', url: '/api/sessions/test-session-1/input', payload: body });
const spies = () => {
const session = harness.ctx.sessions.get('test-session-1')!;
return { session, viaMux: vi.spyOn(session, 'writeViaMux'), direct: vi.spyOn(session, 'write') };
};
const LONG_PROMPT =
'Reply with only the word ok and nothing else, this sentence is padding to reach about one hundred chars.\r';
it('sends a prompt with no useMux through the mux, not as one burst', async () => {
const { viaMux, direct } = spies();
const res = await post({ input: LONG_PROMPT });
expect(res.statusCode).toBe(200);
expect(viaMux).toHaveBeenCalledWith(LONG_PROMPT, { fromUser: true });
expect(direct).not.toHaveBeenCalled();
});
it('answers only once the mux write is done, so the next frame cannot overtake it', async () => {
// The browser's POST fallback sends frames one at a time and waits for each 2xx;
// a fire-and-forget write here would let its next keystroke land before the Enter.
const { viaMux } = spies();
let finished = false;
viaMux.mockImplementation(async () => {
await new Promise((r) => setTimeout(r, 30));
finished = true;
return true;
});
await post({ input: 'ok\r', clientId: 'browser-1', seq: 1 });
expect(finished).toBe(true);
});
it('falls back to the direct write when the mux write fails', async () => {
const { viaMux, direct } = spies();
viaMux.mockResolvedValue(false);
await post({ input: 'hello\r' });
expect(direct).toHaveBeenCalledWith('hello\r', { fromUser: true });
});
it('keeps the raw write for an explicit useMux: false', async () => {
const { viaMux, direct } = spies();
await post({ input: LONG_PROMPT, useMux: false });
expect(direct).toHaveBeenCalledWith(LONG_PROMPT, { fromUser: true });
expect(viaMux).not.toHaveBeenCalled();
});
it.each([
['a bare Enter', '\r'],
['text with no Enter', 'hello'],
['an arrow key', '\x1b[A'],
['a bracketed paste frame', '\x1b[200~line one\nline two\x1b[201~'],
['a line feed inside', 'line one\nline two\r'],
['two Enters', 'hello\r\r'],
['a tab', 'a\tb\r'],
])('leaves %s on the direct write', async (_label, input) => {
const { viaMux, direct } = spies();
await post({ input });
expect(direct).toHaveBeenCalledWith(input, { fromUser: true });
expect(viaMux).not.toHaveBeenCalled();
});
});
describe('isPlainPromptInput', () => {
it('accepts printable text ending in exactly one carriage return', () => {
expect(isPlainPromptInput('run the tests\r')).toBe(true);
expect(isPlainPromptInput('ünïcødé and emoji 🚀\r')).toBe(true);
});
it('refuses anything carrying another control character', () => {
for (const input of ['\r', 'x', 'x\n', 'x\r\n', 'x\r\r', '\x1b[Ax\r', 'a\tb\r', 'x\x7f\r', 'x\u009b\r', '\rx']) {
expect(isPlainPromptInput(input), JSON.stringify(input)).toBe(false);
}
});
});
+47
View File
@@ -920,6 +920,53 @@ describe('session-routes', () => {
expect(res.headers['server-timing']).toMatch(/^capture;dur=\d+\.\d, prepare;dur=\d+\.\d, total;dur=\d+\.\d$/); expect(res.headers['server-timing']).toMatch(/^capture;dur=\d+\.\d, prepare;dur=\d+\.\d, total;dur=\d+\.\d$/);
}); });
it('full reload with a tail (?full=1&tail=) cuts the full capture to its newest bytes, cursor restore intact', async () => {
// A Shell scroll-to-top asks for exactly this (`_maybeRefetchFullHistory`):
// tmux's whole scrollback, bounded to the tab-switch tail size. The client
// relies on all three answers below, so a refactor that dropped the tail on
// a full capture (an unbounded pull from an ordinary scroll) or cut off the
// closing cursor move (a caret parked below the prompt) must fail here.
const tail = 1024 * 1024;
const oldestMarker = 'BOUNDED_OLDEST_LINE_00001';
const newestMarker = 'BOUNDED_NEWEST_LINE_40000';
const rows: string[] = [oldestMarker];
for (let i = 2; i < 40_000; i++) rows.push(`shell history line ${String(i).padStart(5, '0')} lorem ipsum`);
rows.push(newestMarker);
// What formatCursorRestore appends: up from the last row, then the column.
const cursorRestore = '\x1b[3A\r\x1b[2C';
const fullHistoryCapture = `${rows.join('\r\n')}${cursorRestore}`;
expect(fullHistoryCapture.length).toBeGreaterThan(tail);
harness.ctx._session.mode = 'shell';
harness.ctx._session.terminalBuffer = '';
const captureSpy = vi.fn((_name: string, opts?: { fullHistory?: boolean }) =>
opts?.fullHistory ? fullHistoryCapture : 'only the visible frame'
);
(harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = captureSpy;
const res = await harness.app.inject({
method: 'GET',
url: `/api/sessions/${harness.ctx._sessionId}/terminal?full=1&tail=${tail}`,
});
expect(res.statusCode).toBe(200);
const body = JSON.parse(res.body);
// Still the scrollback, not the visible frame a plain `?tail=` gets.
expect(captureSpy).toHaveBeenCalledWith(
harness.ctx._session.muxName,
expect.objectContaining({ fullHistory: true })
);
expect(body.data.source).toBe('mux-full-history');
// Recoverable, not 'capped': Load full history can still bring the rest back.
expect(body.data.truncated).toBe(true);
expect(body.data.truncationReason).toBe('tail');
expect(body.data.fullSize).toBe(fullHistoryCapture.length);
expect(body.data.terminalBuffer.length).toBeLessThanOrEqual(tail);
expect(body.data.terminalBuffer).toContain(newestMarker);
expect(body.data.terminalBuffer).not.toContain(oldestMarker);
expect(body.data.terminalBuffer.endsWith(`${newestMarker}${cursorRestore}`)).toBe(true);
});
it('full reload (?full=1) returns the tmux capture ALONE — byte history is not duplicated', async () => { it('full reload (?full=1) returns the tmux capture ALONE — byte history is not duplicated', async () => {
// The full-history capture is the rendered form of everything already in // The full-history capture is the rendered form of everything already in
// the byte buffer; prepending the byte history would replay the whole // the byte buffer; prepending the byte history would replay the whole
+233
View File
@@ -0,0 +1,233 @@
/**
* A session whose turn ended waiting for workers it started counts as working.
*
* The bug this pins: when Claude hands work to background agents or an ultracode
* workflow, it ends its turn and closes it with `✻ Waiting for 1 dynamic workflow to
* finish` instead of `✻ Brewed for 1m 18s`. The pane goes quiet with the composer up, so
* every other signal called the session idle while it was plainly busy, and it resumes
* on its own the moment the workers report back.
*
* The chrome rows below (the closing row, the right-aligned hint, the composer rules, the
* footer, the workflow progress row) are verbatim from a live Claude Code 2.1.283 pane on
* 2026-09-28 (`tmux -L codeman capture-pane -p`, 64 columns). The prose and the names are
* invented.
*/
import { describe, expect, it, vi, afterEach } from 'vitest';
import { Session } from '../src/session.js';
import { getCli } from '../src/config/cli-registry/index.js';
import { compileVersionRegex } from '../src/config/cli-registry/patterns.js';
import { isAwaitingWorkers, AWAITING_SEARCH_ROWS, IDLE_SILENCE_MS } from '../src/session-activity.js';
/** The registry's own pattern, which is what every consumer runs. */
const CLAUDE_AWAITING = compileVersionRegex(getCli('claude')!.capabilities.workDetect!.awaitingLine!)!;
const RULE = '────────────────────────────────────────────────────────────────';
const NAMED_RULE = '──────────────────────────────────────────────────── w1-demo ─';
/** Everything Claude draws from the composer down while a workflow runs. */
const COMPOSER_AND_FOOTER = [
NAMED_RULE,
'❯ sounds good, go ahead',
RULE,
' Opus 5.5 (1M context) in:285,618 out:581 ctx:29%',
' ⏵⏵ bypass permissions on (shift+tab to cycle) · ← 2 agents',
'',
' ◯ docs-research ▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱ ↓ 1.1m',
];
/** A pane whose newest turn closed with `closing`, then the optional hint row. */
function frame(closing: string, { hint = true, body = [] as string[] } = {}): string {
return [
'⏺ The research is running now: three tracks, each checked by',
' a second agent.',
'',
' When it is done I will rewrite the plan.',
'',
...body,
closing,
...(hint ? [' 286199 tokens'] : []),
...COMPOSER_AND_FOOTER,
'',
].join('\n');
}
const WAITING_WORKFLOW = frame('✻ Waiting for 1 dynamic workflow to finish');
const WAITING_BOTH = frame('✻ Waiting for 2 background agents and 1 dynamic workflow to finish');
const WAITING_AGENT = frame('✻ Waiting for 1 background agent to finish', { hint: false });
const DONE = frame('✻ Brewed for 1m 18s · done 3:04 PM');
/**
* The same screen after the workers reported back and the follow-up turn ended. The
* waiting row is a snapshot Claude never redraws, so it is STILL on screen, just no longer
* the newest row.
*/
const FOLLOW_UP_DONE = frame('✻ Cooked for 12s', {
body: [
'✻ Waiting for 1 dynamic workflow to finish',
'',
'⏺ All three tracks are back. The plan is rewritten and pushed.',
'',
],
});
describe('isAwaitingWorkers', () => {
it('reads the closing row of a turn that handed off to a workflow', () => {
expect(isAwaitingWorkers(WAITING_WORKFLOW, CLAUDE_AWAITING, '❯')).toBe(true);
});
it('reads every form Claude builds the row in', () => {
expect(isAwaitingWorkers(WAITING_BOTH, CLAUDE_AWAITING, '❯')).toBe(true);
expect(isAwaitingWorkers(WAITING_AGENT, CLAUDE_AWAITING, '❯')).toBe(true);
expect(isAwaitingWorkers(frame('✻ Waiting for 3 background agents to finish'), CLAUDE_AWAITING, '❯')).toBe(true);
expect(isAwaitingWorkers(frame('✻ Waiting for 2 dynamic workflows to finish'), CLAUDE_AWAITING, '❯')).toBe(true);
});
it('leaves an ordinary turn end alone', () => {
expect(isAwaitingWorkers(DONE, CLAUDE_AWAITING, '❯')).toBe(false);
});
it('ignores a stale waiting row once a newer turn has closed below it', () => {
// The trap the whole positional walk exists for: matching the words anywhere on the
// screen would pin the session busy until they scrolled away.
expect(FOLLOW_UP_DONE).toContain('Waiting for 1 dynamic workflow to finish');
expect(isAwaitingWorkers(FOLLOW_UP_DONE, CLAUDE_AWAITING, '❯')).toBe(false);
});
it('refuses the words when the agent wrote them', () => {
// Claude's own rows start in column 0; the agent's prose sits behind `⏺ ` or is
// indented, so an agent cannot keep itself busy by printing the sentence.
expect(isAwaitingWorkers(frame('⏺ ✻ Waiting for 1 background agent to finish'), CLAUDE_AWAITING, '❯')).toBe(false);
expect(isAwaitingWorkers(frame(' ✻ Waiting for 1 background agent to finish'), CLAUDE_AWAITING, '❯')).toBe(false);
});
it('says nothing about a screen with no composer on it', () => {
const noComposer = WAITING_WORKFLOW.replace('❯ sounds good, go ahead', ' 1. Yes 2. No');
expect(isAwaitingWorkers(noComposer, CLAUDE_AWAITING, '❯')).toBe(false);
expect(isAwaitingWorkers('', CLAUDE_AWAITING, '❯')).toBe(false);
expect(isAwaitingWorkers(null, CLAUDE_AWAITING, '❯')).toBe(false);
});
it('finds the composer in the boxed layout too', () => {
const boxed = [
'✻ Waiting for 1 dynamic workflow to finish',
'╭──────────────────────────────────────╮',
'│ ❯ │',
'╰──────────────────────────────────────╯',
' ⏵⏵ bypass permissions on · ← 1 agent',
].join('\n');
expect(isAwaitingWorkers(boxed, CLAUDE_AWAITING, '❯')).toBe(true);
});
it('reads a coloured capture', () => {
const coloured = WAITING_WORKFLOW.replace(
'✻ Waiting for 1 dynamic workflow to finish',
'\u001b[2m✻\u001b[0m \u001b[2mWaiting for \u001b[1m1\u001b[22m dynamic workflow to finish\u001b[0m'
);
expect(isAwaitingWorkers(coloured, CLAUDE_AWAITING, '❯')).toBe(true);
});
it('stops looking a few rows above the composer', () => {
const farAway = [
'✻ Waiting for 1 dynamic workflow to finish',
...Array.from({ length: AWAITING_SEARCH_ROWS }, () => ''),
...COMPOSER_AND_FOOTER,
].join('\n');
expect(isAwaitingWorkers(farAway, CLAUDE_AWAITING, '❯')).toBe(false);
});
it('survives a pattern handed to it with the global flag set', () => {
const global = new RegExp(CLAUDE_AWAITING.source, 'g');
expect(isAwaitingWorkers(WAITING_WORKFLOW, global, '❯')).toBe(true);
expect(isAwaitingWorkers(WAITING_WORKFLOW, global, '❯')).toBe(true);
});
});
/** A composer repaint: the frame Claude ships roughly once a second while working. */
const COMPOSER_REPAINT =
'\x1b[31;1H\x1b[38;5;246m❯\xa0\x1b[39m\x1b[0m\x1b[33;1H \x1b[38;5;246mOpus 5 in:143,699 out:669 ctx:14%\x1b[39m';
type SessionInternals = {
_handleTerminalOutput(data: string): void;
_detectInteractiveActivity(data: string): void;
};
function feed(session: Session, data: string): void {
const internals = session as unknown as SessionInternals;
internals._handleTerminalOutput(data);
internals._detectInteractiveActivity(data);
}
/** A session whose mux reports a scripted screen for the pane probe to read. */
function withFakePane(read: () => string, mode: 'claude' | 'codex' = 'claude'): Session {
const mux = {
isAvailable: () => true,
capturePaneText: () => read(),
} as unknown as NonNullable<ConstructorParameters<typeof Session>[0]>['mux'];
return new Session({
workingDir: '/tmp',
mode,
mux,
muxSession: { muxName: 'codeman-test', sessionId: 'test', createdAt: Date.now() },
} as ConstructorParameters<typeof Session>[0]);
}
/** Run one turn and let it end, which is when the probe reads the screen. */
function runAndSettle(session: Session, repaint: string = COMPOSER_REPAINT): void {
for (let i = 0; i < 3; i++) {
feed(session, repaint);
vi.advanceTimersByTime(1000);
}
vi.advanceTimersByTime(IDLE_SILENCE_MS + 2000);
}
describe('Session status while its workers run', () => {
afterEach(() => {
vi.useRealTimers();
});
it('stays working when the turn ends waiting for a workflow', () => {
vi.useFakeTimers();
const session = withFakePane(() => WAITING_WORKFLOW);
runAndSettle(session);
expect(session.status).toBe('busy');
expect(session.isWorking).toBe(true);
});
it('goes idle once the follow-up turn closes, although the old row is still on screen', () => {
vi.useFakeTimers();
const screen = { text: WAITING_WORKFLOW };
const session = withFakePane(() => screen.text);
runAndSettle(session);
expect(session.status).toBe('busy');
screen.text = FOLLOW_UP_DONE;
// The probe keeps re-reading the screen on its own slow cadence while it says busy,
// with no PTY output needed to trigger it.
vi.advanceTimersByTime(IDLE_SILENCE_MS + 10_000);
expect(session.status).toBe('idle');
expect(session.isWorking).toBe(false);
});
it('reports an ordinary turn end as idle, as before', () => {
vi.useFakeTimers();
const session = withFakePane(() => DONE);
runAndSettle(session);
expect(session.status).toBe('idle');
});
it('does not apply to a CLI whose registry entry declares no awaitingLine', () => {
vi.useFakeTimers();
expect(getCli('codex')?.capabilities.workDetect?.awaitingLine).toBeUndefined();
const codexScreen = ['✻ Waiting for 1 dynamic workflow to finish', '', '› Ask Codex to do anything', ''].join('\n');
const session = withFakePane(() => codexScreen, 'codex');
runAndSettle(session, '\x1b[31;1H\x1b[38;5;246m›\xa0\x1b[39m\x1b[0m');
expect(session.status).toBe('idle');
});
});
+39 -1
View File
@@ -133,7 +133,6 @@ describe('watchingLabel', () => {
'1 MCP task', '1 MCP task',
'1 background dynamic workflow', '1 background dynamic workflow',
'2 remote dynamic workflows', '2 remote dynamic workflows',
'1 Artifact comment monitor',
'2 teams', '2 teams',
]; ];
for (const label of labels) { for (const label of labels) {
@@ -141,6 +140,45 @@ describe('watchingLabel', () => {
} }
}); });
it('reports no watching while the agent waits for comments on an artifact', () => {
// An agent that publishes an artifact arms a monitor for its comments and ends its
// turn. That monitor waits on the user, so the idle alert has to reach them. The
// singular footer is a live capture from 2026-09-25; the plural is assumed.
expect(
watchingLabel(pane('⏵⏵ bypass permissions on · 1 Artifact comment monitor · ← for agents'), CLAUDE_WATCHING)
).toBeNull();
expect(
watchingLabel(pane('⏵⏵ bypass permissions on · 2 Artifact comment monitors · ← for agents'), CLAUDE_WATCHING)
).toBeNull();
});
it('lets a comment monitor outrank other background work on the same row', () => {
// A shell beside the monitor is still running, but the agent needs the user all the
// same, and the chip order on the footer must not decide that. The second row is
// the one that needs the `^` in front of the lookahead.
expect(
watchingLabel(
pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comment monitor · ← for agents'),
CLAUDE_WATCHING
)
).toBeNull();
expect(
watchingLabel(
pane('⏵⏵ bypass permissions on · 1 Artifact comment monitor · 1 shell · ← for agents'),
CLAUDE_WATCHING
)
).toBeNull();
});
it('still refuses a footer cut off in the middle of the comment monitor', () => {
expect(
watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comment moni…'), CLAUDE_WATCHING)
).toBeNull();
// Cut before "comment" is complete: the lookahead keys on "Artifact" alone for these.
expect(watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comm…'), CLAUDE_WATCHING)).toBeNull();
expect(watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact…'), CLAUDE_WATCHING)).toBeNull();
});
it('says nothing about a pane that is running nothing', () => { it('says nothing about a pane that is running nothing', () => {
expect(watchingLabel(NOTHING_RUNNING, CLAUDE_WATCHING)).toBeNull(); expect(watchingLabel(NOTHING_RUNNING, CLAUDE_WATCHING)).toBeNull();
expect(watchingLabel('', CLAUDE_WATCHING)).toBeNull(); expect(watchingLabel('', CLAUDE_WATCHING)).toBeNull();
+360
View File
@@ -0,0 +1,360 @@
/**
* @fileoverview A shell pane's scroll-up must reach the history tmux still holds.
*
* tmux repaints a burst of output instead of scrolling it, so after `cat` of a
* file longer than the screen the browser holds about one screen of scrollback
* while tmux holds all of it. Other modes recover it by re-pulling `?full=1`
* when the wheel reaches the top (`_maybeRefetchFullHistory`, issue #205). Shell
* declined that gesture outright to keep a multi-megabyte capture off xterm's
* main thread, leaving only the "Load full history" button, and that button
* renders only once a replay was truncated. A young shell tab therefore had no
* way to scroll back at all.
*
* The gesture now pulls a BOUNDED window (`?full=1&tail=TERMINAL_TAIL_SIZE`),
* the button stays the unbounded path, and a window the browser already holds
* in full is not rewritten. Neither is one a browser at xterm's scrollback cap
* could never hold, and a window cut from a byte-capped capture is still labelled
* recoverable, since Load full history can reach past it.
*
* ORDER MATTERS: that skip must run BEFORE the downgrade guard. The guard reads
* "smaller than the browser" as "tmux has nothing more to give", which is true of
* an unbounded capture and false of a window cut at the tail size, so a bounded
* window that reached it marked the session exhausted and took Load full history
* off the banner while tmux still held the rest. The second block below drives
* the real `_setHistoryTruncation` and the real `computeHistoryTruncationNotice`
* to pin what the user is actually told.
*
* The method is extracted from app.js and run in a `vm` against stubs (no jsdom
* on this box; see connection-indicator.test.ts), with the REAL row estimators
* from terminal-ui.js, which decide both the downgrade and the no-gain skip.
*/
import { readFileSync } from 'node:fs';
import { performance } from 'node:perf_hooks';
import { resolve } from 'node:path';
import vm from 'node:vm';
import { describe, expect, it, vi } from 'vitest';
const PUBLIC = resolve(import.meta.dirname, '../src/web/public');
const TERMINAL_TAIL_SIZE = 1024 * 1024;
function methodSource(source: string, method: string): string {
const start = source.search(new RegExp(`^ {2}(?:async )?${method}\\(`, 'm'));
expect(start, `${method} not found`).toBeGreaterThan(-1);
const next = /^ {2}(?:async )?[A-Za-z_$][\w$]*\(/m.exec(source.slice(start + 1));
return next ? source.slice(start, start + 1 + next.index) : source.slice(start);
}
/** Real terminal-ui.js mixin, for `_estimateReplayRows` / `_replayWouldShrinkBuffer`. */
function loadTerminalMixin(): Record<string, unknown> {
const source = readFileSync(resolve(PUBLIC, 'terminal-ui.js'), 'utf8');
const FakeCodemanApp = function () {} as unknown as { prototype: Record<string, unknown> };
const context = vm.createContext({
console,
performance,
setTimeout,
clearTimeout,
setInterval: vi.fn(),
clearInterval: vi.fn(),
requestAnimationFrame: vi.fn(),
CodemanApp: FakeCodemanApp,
window: { addEventListener: vi.fn(), removeEventListener: vi.fn() },
document: { addEventListener: vi.fn() },
});
vm.runInContext(source, context);
return FakeCodemanApp.prototype;
}
function loadRefetch(): (this: unknown, opts?: { force?: boolean }) => Promise<void> {
const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8');
const body = methodSource(app, '_maybeRefetchFullHistory');
const context = vm.createContext({ performance, TERMINAL_TAIL_SIZE, TERMINAL_CHUNK_SIZE: 32 * 1024 });
return vm.runInContext(`({ ${body} })._maybeRefetchFullHistory`, context);
}
/** The REAL `_setHistoryTruncation`, so the banner state a pull leaves behind is what production would hold. */
function loadSetHistoryTruncation(): (this: unknown, sessionId: string, payload?: Record<string, unknown>) => void {
const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8');
const body = methodSource(app, '_setHistoryTruncation');
return vm.runInContext(`({ ${body} })._setHistoryTruncation`, vm.createContext({}));
}
/** The REAL banner decision from constants.js: what the user is told, and whether Load full history is offered. */
function loadNotice() {
const context = vm.createContext({ console, window: {}, document: {}, navigator: { userAgent: 'test' } });
vm.runInContext(
`${readFileSync(resolve(PUBLIC, 'constants.js'), 'utf8')}\n;globalThis.__notice = computeHistoryTruncationNotice;`,
context,
{ filename: 'constants.js' }
);
return (context as { __notice: (s: Record<string, unknown>) => { visible: boolean; canLoadMore: boolean } }).__notice;
}
const mixin = loadTerminalMixin();
const refetch = loadRefetch();
const setHistoryTruncation = loadSetHistoryTruncation();
const computeNotice = loadNotice();
const lines = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\r\n');
/** A `?full=1&tail=` answer whose window was CUT at the tail size: tmux holds ~3 MiB, the window carries 1 MiB. */
const TAIL_CUT = {
truncated: true,
truncationReason: 'tail',
fullSize: 3 * 1024 * 1024,
retainedBytes: TERMINAL_TAIL_SIZE,
source: 'mux-full-history',
};
function makeApp(
mode: string,
{
bufferRows,
capture,
payload = {},
scrollback = 0,
}: { bufferRows: number; capture: string; payload?: Record<string, unknown>; scrollback?: number }
) {
const urls: string[] = [];
const app = {
activeSessionId: 's1',
sessions: new Map([['s1', { mode }]]),
detachedSessions: new Set<string>(),
_fullHistoryRepullInFlight: false,
_isLoadingBuffer: false,
_fullHistoryRepullAt: new Map<string, number>(),
_fullHistoryRepullUseless: new Set<string>(),
terminalBufferCache: new Map<string, string>(),
terminal: {
cols: 80,
rows: 30,
// xterm's scrollback option; 0 leaves the browser-cap check out of a test.
options: { scrollback },
buffer: { active: { length: bufferRows } },
scrollToLine: vi.fn(),
scrollToTop: vi.fn(),
},
_estimateReplayRows: mixin._estimateReplayRows,
_replayWouldShrinkBuffer: mixin._replayWouldShrinkBuffer,
_fetchTerminalCapture: vi.fn(async (url: string) => {
urls.push(url);
return {
headersAt: performance.now(),
headers: { get: () => '' },
json: { data: { terminalBuffer: capture, source: 'mux-full-history', ...payload } },
};
}),
_recordTerminalLoadTiming: vi.fn(),
_logScrollRouting: vi.fn(),
// The real method behind a spy, so a test sees both what it was called with
// and the banner state (`_historyTruncation`) it leaves behind.
_historyTruncation: new Map<string, unknown>(),
_renderHistoryTruncationBanner: vi.fn(),
_setHistoryTruncation: vi.fn((sessionId: string, p?: Record<string, unknown>): void => {
setHistoryTruncation.call(app, sessionId, p);
}),
_resetTerminalForReplay: vi.fn(),
_bufferLoadFinishOpts: vi.fn(() => ({})),
chunkedTerminalWrite: vi.fn(async () => ({ parsedAt: performance.now(), bufferLength: 400, completed: true })),
_syncStickyScrollBaseline: vi.fn(),
};
return { app, urls };
}
describe('shell scroll-up pulls a bounded window of tmux history', () => {
it('a shell scroll gesture requests full history bounded by the tail size', async () => {
const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) });
await refetch.call(app);
expect(urls).toEqual([`/api/sessions/s1/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`]);
// It then actually replays the recovered history.
expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1);
expect(app.chunkedTerminalWrite).toHaveBeenCalledTimes(1);
});
it('the Load full history button stays unbounded for a shell', async () => {
const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) });
await refetch.call(app, { force: true });
expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']);
});
it('other modes keep the unbounded scroll pull', async () => {
const { app, urls } = makeApp('claude', { bufferRows: 40, capture: lines(300) });
await refetch.call(app);
expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']);
});
it('a bounded window the browser already holds is not rewritten', async () => {
// Browser already has every row the window carries: resetting to rewrite
// it would jump the viewport on every scroll that outlasts the cooldown.
const { app } = makeApp('shell', { bufferRows: 320, capture: lines(300) });
await refetch.call(app);
expect(app._resetTerminalForReplay).not.toHaveBeenCalled();
expect(app.chunkedTerminalWrite).not.toHaveBeenCalled();
// Not latched as useless: more output can put more history in tmux.
expect(app._fullHistoryRepullUseless.has('s1')).toBe(false);
// Nothing was written, so the banner state is left exactly as it was.
expect(app._setHistoryTruncation).not.toHaveBeenCalled();
// …but the skip is visible to someone diagnosing "scroll-to-top does nothing".
expect(app._logScrollRouting).toHaveBeenCalledWith('repull-skipped-bounded');
});
it("a browser at xterm's scrollback cap stops replaying a window it can never hold, and backs off", async () => {
// xterm keeps at most `scrollback + rows` rows while tmux keeps 100k lines, so
// a 1 MiB window of short lines can carry more rows than the browser ever will.
// `windowRows <= rows held` then never comes true, and every scroll-to-top
// past the cooldown reset and re-parsed the window. Untruncated on purpose:
// the back-off has to come from the full browser, not from `truncated`.
const { app } = makeApp('shell', { bufferRows: 40, capture: lines(2000), scrollback: 1000 });
// The first pull has room to grow, so it replays.
await refetch.call(app);
expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1);
expect(app._fullHistoryRepullUseless.has('s1')).toBe(false);
// xterm kept only the last `scrollback + rows` of the 2000 rows written.
app.terminal.buffer.active.length = 1000 + 30;
// Past the 4 s cooldown: the browser is full, so nothing is replayed and the
// session backs off for a minute.
app._fullHistoryRepullAt.set('s1', Date.now() - 5000);
await refetch.call(app);
expect(app._fetchTerminalCapture).toHaveBeenCalledTimes(2);
expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1);
expect(app.chunkedTerminalWrite).toHaveBeenCalledTimes(1);
expect(app._fullHistoryRepullUseless.has('s1')).toBe(true);
// So a scroll 10 s later does not even ask the server for another capture.
app._fullHistoryRepullAt.set('s1', Date.now() - 10_000);
await refetch.call(app);
expect(app._fetchTerminalCapture).toHaveBeenCalledTimes(2);
expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1);
});
it('a replayed window cut from a byte-capped capture still offers Load full history', async () => {
// The route keeps `truncationReason: 'capped'` through the tail cut when the
// full capture exceeded the byte cap. On a bounded window that is not "gone for
// good": the unbounded pull behind the button returns up to the cap itself.
const capped = { ...TAIL_CUT, truncationReason: 'capped', fullSize: 40 * 1024 * 1024 };
const { app } = makeApp('shell', { bufferRows: 40, capture: lines(300), payload: capped });
await refetch.call(app);
expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1);
expect(app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ truncationReason: 'tail' }));
const notice = computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>);
expect(notice.visible).toBe(true);
expect(notice.canLoadMore).toBe(true);
// The button's own unbounded pull is the one place 'capped' is the truth.
const button = makeApp('shell', { bufferRows: 40, capture: lines(300), payload: capped });
await refetch.call(button.app, { force: true });
expect(computeNotice(button.app._historyTruncation.get('s1') as Record<string, unknown>).canLoadMore).toBe(false);
});
});
describe('a skipped bounded window never damages the Load full history banner', () => {
it('a tail-cut window smaller than the browser is not replayed and never marks the session exhausted', async () => {
// The browser holds far more rows than a 1 MiB window carries, and tmux holds
// ~3 MiB. The downgrade guard reads that as "tmux has nothing more to give",
// which is true of an unbounded capture and false of a window cut at the tail.
const { app } = makeApp('shell', { bufferRows: 5000, capture: lines(300), payload: TAIL_CUT });
// The tab load that put this session on screen left it truncated and recoverable.
app._setHistoryTruncation('s1', TAIL_CUT);
app._setHistoryTruncation.mockClear();
await refetch.call(app);
expect(app._resetTerminalForReplay).not.toHaveBeenCalled();
expect(app.chunkedTerminalWrite).not.toHaveBeenCalled();
// Not even a relabel: a skipped window writes nothing, banner state included.
expect(app._setHistoryTruncation).not.toHaveBeenCalled();
const notice = computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>);
expect(notice.visible).toBe(true);
// "Earlier output is no longer kept" would be a lie: tmux still holds ~2 MiB more.
expect(notice.canLoadMore).toBe(true);
});
it('a skip right after Load full history leaves the banner as that load set it', async () => {
// Load full history replayed everything, so nothing is truncated any more.
const afterLoadFullHistory = {
truncated: false,
fullSize: 3 * 1024 * 1024,
retainedBytes: 3 * 1024 * 1024,
source: 'mux-full-history',
};
const { app } = makeApp('shell', { bufferRows: 5000, capture: lines(300), payload: TAIL_CUT });
app._setHistoryTruncation('s1', afterLoadFullHistory);
const before = structuredClone(app._historyTruncation.get('s1'));
app._setHistoryTruncation.mockClear();
await refetch.call(app);
// Relabelling it from the bounded payload would call a terminal that holds ALL
// of the history "the most recent 1.0 MB".
expect(app._setHistoryTruncation).not.toHaveBeenCalled();
expect(app._historyTruncation.get('s1')).toEqual(before);
expect(computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>).visible).toBe(false);
});
it('backs off for a minute after a truncated skip, and keeps the 4 s cooldown after an untruncated one', async () => {
// Truncated: the gesture cannot reach anything older than the browser shows, and
// every ask costs the server a synchronous capture of the whole history.
// Only just larger than the window, so the downgrade guard does not fire here:
// the back-off has to come from the skip itself.
const cut = makeApp('shell', { bufferRows: 320, capture: lines(300), payload: TAIL_CUT });
await refetch.call(cut.app);
expect(cut.app._fullHistoryRepullUseless.has('s1')).toBe(true);
expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(1);
// Well past 4 s, still inside the minute: no second capture.
cut.app._fullHistoryRepullAt.set('s1', Date.now() - 10_000);
await refetch.call(cut.app);
expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(1);
cut.app._fullHistoryRepullAt.set('s1', Date.now() - 61_000);
await refetch.call(cut.app);
expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(2);
// Untruncated: it IS all of tmux's history, and the next burst can add to it.
const whole = makeApp('shell', { bufferRows: 320, capture: lines(300) });
await refetch.call(whole.app);
expect(whole.app._fullHistoryRepullUseless.has('s1')).toBe(false);
whole.app._fullHistoryRepullAt.set('s1', Date.now() - 5000);
await refetch.call(whole.app);
expect(whole.app._fetchTerminalCapture).toHaveBeenCalledTimes(2);
});
it('the downgrade guard still refuses an unbounded capture smaller than the browser, and still marks it exhausted', async () => {
// Reordering must not weaken the guard it moved above: a repaint-mode pane's
// capture really is one frame, and rewriting with it would destroy history.
const oneFrame = makeApp('claude', { bufferRows: 300, capture: lines(36) });
await refetch.call(oneFrame.app);
expect(oneFrame.app._resetTerminalForReplay).not.toHaveBeenCalled();
expect(oneFrame.app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ exhausted: true }));
expect(oneFrame.app._fullHistoryRepullUseless.has('s1')).toBe(true);
// The button is unbounded too, so the same guard governs it for a shell.
const button = makeApp('shell', { bufferRows: 5000, capture: lines(300) });
await refetch.call(button.app, { force: true });
expect(button.app._resetTerminalForReplay).not.toHaveBeenCalled();
expect(button.app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ exhausted: true }));
});
});
describe('_replayWouldShrinkBuffer takes rows the caller already estimated', () => {
const shrink = mixin._replayWouldShrinkBuffer as (this: unknown, capture: string, rows?: number) => boolean;
const make = () => ({
terminal: { cols: 80, rows: 30, buffer: { active: { length: 200 } } },
_estimateReplayRows: vi.fn(mixin._estimateReplayRows as (t: string, c: number) => number),
});
it('does not scan the capture again when handed the estimate', () => {
const ctx = make();
expect(shrink.call(ctx, lines(300), 300)).toBe(false);
expect(shrink.call(ctx, lines(300), 5)).toBe(true);
expect(ctx._estimateReplayRows).not.toHaveBeenCalled();
});
it('still estimates for itself when called the old way', () => {
const ctx = make();
expect(shrink.call(ctx, lines(300))).toBe(false);
expect(ctx._estimateReplayRows).toHaveBeenCalledTimes(1);
});
});
+24 -4
View File
@@ -47,11 +47,13 @@ function loadTerminalUiHarness() {
} }
/** A Claude session whose local buffer holds exactly one screen (baseY 0). */ /** A Claude session whose local buffer holds exactly one screen (baseY 0). */
function hollowClaudeApp(overrides: { cliVersion?: string; rows?: number } = {}) { function hollowClaudeApp(overrides: { cliVersion?: string; rows?: number; cliMouseTracking?: boolean } = {}) {
const { app, logs } = loadTerminalUiHarness(); const { app, logs } = loadTerminalUiHarness();
const sent: Array<{ id: string; data: string }> = []; const sent: Array<{ id: string; data: string }> = [];
app.activeSessionId = 'sess-1'; app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: overrides.cliVersion }]]); app.sessions = new Map([
['sess-1', { mode: 'claude', cliVersion: overrides.cliVersion, cliMouseTracking: overrides.cliMouseTracking }],
]);
app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data }); app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data });
app.terminal = { app.terminal = {
cols: 80, cols: 80,
@@ -111,12 +113,20 @@ describe('full-history re-pull downgrade guard (issue #205 round 2)', () => {
// Anchor on the open paren, not the full empty signature: the method takes // Anchor on the open paren, not the full empty signature: the method takes
// options since #258 ({ force }) and this guard is about ORDER, not arity. // options since #258 ({ force }) and this guard is about ORDER, not arity.
const start = source.indexOf('async _maybeRefetchFullHistory('); const start = source.indexOf('async _maybeRefetchFullHistory(');
const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer)', start); // Also anchored on the open paren: the guard is handed the rows the caller
// already estimated, and this test is about ORDER, not the argument list.
const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer', start);
const boundedSkip = source.indexOf('boundedShellPull && (windowRows <= rowsNow || browserFull)', start);
const reset = source.indexOf('this._resetTerminalForReplay()', start); const reset = source.indexOf('this._resetTerminalForReplay()', start);
expect(start).toBeGreaterThan(-1); expect(start).toBeGreaterThan(-1);
expect(guard).toBeGreaterThan(start); expect(guard).toBeGreaterThan(start);
expect(guard).toBeLessThan(reset); // refuse first, only then reset+rewrite expect(guard).toBeLessThan(reset); // refuse first, only then reset+rewrite
// A bounded shell window is skipped BEFORE the guard sees it: the guard reads
// "smaller than the browser" as "tmux has nothing more", which a window cut at
// the tail size does not mean (see shell-scroll-history-pull.test.ts).
expect(boundedSkip).toBeGreaterThan(start);
expect(boundedSkip).toBeLessThan(guard);
// A hollow pane must also stop re-fetching megabytes on every scroll-up. // A hollow pane must also stop re-fetching megabytes on every scroll-up.
expect(source).toContain('this._fullHistoryRepullUseless'); expect(source).toContain('this._fullHistoryRepullUseless');
expect(source).toContain('this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000'); expect(source).toContain('this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000');
@@ -186,7 +196,8 @@ describe('PageUp/PageDown fallback for a hollow local buffer (issue #205 round 2
// "Wheel scrolls local history" ON pins the wheel to a buffer that, for a // "Wheel scrolls local history" ON pins the wheel to a buffer that, for a
// repaint-mode CLI, is empty — a user who flipped it while hunting for a fix // repaint-mode CLI, is empty — a user who flipped it while hunting for a fix
// on 1.11.x would have ended up with a completely dead wheel on 1.12.0. // on 1.11.x would have ended up with a completely dead wheel on 1.12.0.
const { app, sent } = hollowClaudeApp({ cliVersion: '2.1.223' }); // gate would forward… // Version and tracking both qualify, so the opt-out is the only thing saying no.
const { app, sent } = hollowClaudeApp({ cliVersion: '2.1.223', cliMouseTracking: true }); // gate would forward…
app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: true }); app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: true });
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); // …but the opt-out wins expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); // …but the opt-out wins
@@ -228,10 +239,19 @@ describe('scroll routing diagnostic (issue #205 round 2)', () => {
expect(logs[0]).toContain('cliVersion=2.1.100'); expect(logs[0]).toContain('cliVersion=2.1.100');
expect(logs[0]).toContain('localScrollbackOptOut=false'); expect(logs[0]).toContain('localScrollbackOptOut=false');
expect(logs[0]).toContain('mouseTracking=none'); expect(logs[0]).toContain('mouseTracking=none');
// The gate's real tracking input: xterm's own mode above is always 'none'
// for Claude, since the server strips the DECSETs.
expect(logs[0]).toContain('cliMouseTracking=false');
app._logScrollRouting('page-keys'); // a changed route still prints app._logScrollRouting('page-keys'); // a changed route still prints
expect(logs).toHaveLength(2); expect(logs).toHaveLength(2);
expect(logs[1]).toContain('page-keys'); expect(logs[1]).toContain('page-keys');
// The CLI turning tracking on changes the gate, so it prints again.
app.sessions.get('sess-1').cliMouseTracking = true;
app._logScrollRouting('page-keys');
expect(logs).toHaveLength(3);
expect(logs[2]).toContain('cliMouseTracking=true');
}); });
it('reports an unknown CLI version, the false-path that disables forwarding', () => { it('reports an unknown CLI version, the false-path that disables forwarding', () => {
+26 -5
View File
@@ -585,7 +585,7 @@ describe('terminal touch tap mouse guard', () => {
it('wheel: forwards to the app for verified sessions without Shift, at ANY scroll position', () => { it('wheel: forwards to the app for verified sessions without Shift, at ANY scroll position', () => {
const { app } = loadTerminalUiHarness(); const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1'; app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187' }]]); app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187', cliMouseTracking: true }]]);
app.terminal = { app.terminal = {
modes: { mouseTrackingMode: 'none' }, modes: { mouseTrackingMode: 'none' },
buffer: { active: { viewportY: 50, baseY: 50 } }, buffer: { active: { viewportY: 50, baseY: 50 } },
@@ -638,7 +638,7 @@ describe('terminal touch tap mouse guard', () => {
buffer: { active: { viewportY: 50, baseY: 50 } }, buffer: { active: { viewportY: 50, baseY: 50 } },
}; };
const withVersion = (cliVersion?: string) => { const withVersion = (cliVersion?: string) => {
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion }]]); app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion, cliMouseTracking: true }]]);
return app._shouldForwardWheelToApp({ shiftKey: false }); return app._shouldForwardWheelToApp({ shiftKey: false });
}; };
@@ -650,6 +650,25 @@ describe('terminal touch tap mouse guard', () => {
expect(withVersion('garbage')).toBe(false); // unparseable → assume older expect(withVersion('garbage')).toBe(false); // unparseable → assume older
}); });
it('wheel: inline claude (no mouse tracking) keeps the local wheel', () => {
// Claude 2.1.280's default inline renderer never enables mouse tracking and
// keeps its transcript in real scrollback, so SGR wheel reports are ignored.
// Forwarding there made every swipe dead on iOS Safari while codex scrolled.
const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1';
app.terminal = {
modes: { mouseTrackingMode: 'none' },
buffer: { active: { viewportY: 50, baseY: 50 } },
};
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.280' }]]);
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.280', cliMouseTracking: false }]]);
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
// Fullscreen (CLAUDE_CODE_NO_FLICKER=1) turns tracking on → forwarding resumes.
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.280', cliMouseTracking: true }]]);
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true);
});
it('wheel: only claude forwards — codex and gemini keep the local wheel', () => { it('wheel: only claude forwards — codex and gemini keep the local wheel', () => {
const { app } = loadTerminalUiHarness(); const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1'; app.activeSessionId = 'sess-1';
@@ -664,17 +683,19 @@ describe('terminal touch tap mouse guard', () => {
// (the codex transcript lives there — inline viewport, no in-app pager) sat unused. // (the codex transcript lives there — inline viewport, no in-app pager) sat unused.
app.sessions = new Map([['sess-1', { mode: 'codex' }]]); app.sessions = new Map([['sess-1', { mode: 'codex' }]]);
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'codex', cliVersion: '9.9.9' }]]); // no version rescues it // Tracking on and a high version, so only the mode check can say no: without
// them the gate is false for claude too and this would pin nothing.
app.sessions = new Map([['sess-1', { mode: 'codex', cliVersion: '9.9.9', cliMouseTracking: true }]]); // no version rescues it
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
app.sessions = new Map([['sess-1', { mode: 'gemini', cliVersion: '9.9.9' }]]); // unverified TUI app.sessions = new Map([['sess-1', { mode: 'gemini', cliVersion: '9.9.9', cliMouseTracking: true }]]); // unverified TUI
expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false);
}); });
it('wheel: the local-scrollback opt-out pins the plain wheel to local scrollback (issue #154)', () => { it('wheel: the local-scrollback opt-out pins the plain wheel to local scrollback (issue #154)', () => {
const { app } = loadTerminalUiHarness(); const { app } = loadTerminalUiHarness();
app.activeSessionId = 'sess-1'; app.activeSessionId = 'sess-1';
app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187' }]]); app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187', cliMouseTracking: true }]]);
app.terminal = { app.terminal = {
modes: { mouseTrackingMode: 'none' }, modes: { mouseTrackingMode: 'none' },
buffer: { active: { viewportY: 50, baseY: 50 } }, buffer: { active: { viewportY: 50, baseY: 50 } },
+217
View File
@@ -0,0 +1,217 @@
/**
* @fileoverview Owner routing of the `webview:changed` SSE event.
*
* Saved web tabs are owner-scoped in multi-user mode (`canAccessOwned` on every
* CRUD route), but their invalidation event used to carry no owner and fell
* through `deriveSseHint`'s global branch, so every connected user learned the
* ids of every other user's web-tab creates, edits and deletes. The event now
* carries the resource owner and routes to that owner plus admins; single-user
* mode (no SSE identity) still delivers it to every client.
*/
import fs from 'node:fs/promises';
import os from 'node:os';
import path from 'node:path';
import Fastify, { type FastifyInstance, type FastifyReply } from 'fastify';
import fastifyCookie from '@fastify/cookie';
import fastifyWebsocket from '@fastify/websocket';
import { afterEach, beforeEach, describe, expect, it } from 'vitest';
import { CleanupManager } from '../src/utils/index.js';
import type { AuthUser } from '../src/types.js';
import { installRouteErrorHandler } from '../src/web/route-error-handler.js';
import { registerWebviewRoutes } from '../src/web/routes/webview-routes.js';
import { WebServer } from '../src/web/server.js';
import { SseEvent } from '../src/web/sse-events.js';
import { SseStreamManager, type SseRoutingHint } from '../src/web/sse-stream-manager.js';
import { deriveWebviewSseHint } from '../src/web/webview-sse.js';
function client() {
const writes: string[] = [];
return {
writes,
reply: { raw: { write: (chunk: string) => (writes.push(chunk), true) } } as unknown as FastifyReply,
};
}
/** The server's real event → routing-hint derivation, without starting the server. */
function serverHint(event: string, data: unknown): SseRoutingHint | undefined {
const server = new WebServer(3999, false, true) as unknown as {
deriveSseHint(event: string, data: unknown): SseRoutingHint | undefined;
};
return server.deriveSseHint(event, data);
}
describe('deriveWebviewSseHint', () => {
it('routes to the exact owner and fails closed when the owner is missing', () => {
expect(deriveWebviewSseHint({ action: 'updated', id: 'w1', owner: 'alice' })).toEqual({
username: 'alice',
sessionScoped: true,
});
expect(deriveWebviewSseHint({ action: 'updated', id: 'w1' })).toEqual({
username: undefined,
sessionScoped: true,
});
});
it('is what the server derives for the webview: family (never the global branch)', () => {
const payload = { action: 'created', id: 'w1', owner: 'alice' };
expect(serverHint(SseEvent.WebviewChanged, payload)).toEqual({ username: 'alice', sessionScoped: true });
expect(serverHint(SseEvent.WebviewChanged, { action: 'deleted', id: 'w1' })).not.toBeUndefined();
});
});
describe('webview:changed delivery', () => {
it('reaches the owner and admins, never another ordinary user', () => {
const cleanup = new CleanupManager();
const manager = new SseStreamManager({ getSessionStateWithRespawn: () => null }, cleanup);
const alice = client();
const bob = client();
const admin = client();
manager.addClient(alice.reply, null, false, undefined, { username: 'alice', role: 'user' });
manager.addClient(bob.reply, null, false, undefined, { username: 'bob', role: 'user' });
manager.addClient(admin.reply, null, false, undefined, { username: 'root', role: 'admin' });
const payload = { action: 'deleted', id: 'w1', owner: 'alice' };
manager.broadcast(SseEvent.WebviewChanged, payload, serverHint(SseEvent.WebviewChanged, payload));
expect(alice.writes).toEqual(['event: webview:changed\ndata: {"action":"deleted","id":"w1","owner":"alice"}\n\n']);
expect(admin.writes).toEqual(alice.writes);
expect(bob.writes).toEqual([]);
cleanup.dispose();
});
it('single-user mode: clients without an identity all still receive it', () => {
const cleanup = new CleanupManager();
const manager = new SseStreamManager({ getSessionStateWithRespawn: () => null }, cleanup);
const tabA = client();
const tabB = client();
manager.addClient(tabA.reply, null, false, undefined, undefined);
manager.addClient(tabB.reply, null, false, undefined, undefined);
const payload = { action: 'created', id: 'w1', owner: '@single' };
manager.broadcast(SseEvent.WebviewChanged, payload, serverHint(SseEvent.WebviewChanged, payload));
expect(tabA.writes).toHaveLength(1);
expect(tabB.writes).toEqual(tabA.writes);
cleanup.dispose();
});
});
describe('webview routes → SSE, end to end', () => {
let tmpDir: string;
let savedDataDir: string | undefined;
let savedMode: string | undefined;
let cleanup: CleanupManager;
let manager: SseStreamManager;
const apps: FastifyInstance[] = [];
beforeEach(async () => {
tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), 'codeman-webview-sse-'));
savedDataDir = process.env.CODEMAN_DATA_DIR;
savedMode = process.env.CODEMAN_MULTIUSER;
process.env.CODEMAN_DATA_DIR = tmpDir;
cleanup = new CleanupManager();
manager = new SseStreamManager({ getSessionStateWithRespawn: () => null }, cleanup);
});
afterEach(async () => {
for (const app of apps.splice(0)) await app.close();
cleanup.dispose();
if (savedDataDir === undefined) delete process.env.CODEMAN_DATA_DIR;
else process.env.CODEMAN_DATA_DIR = savedDataDir;
if (savedMode === undefined) delete process.env.CODEMAN_MULTIUSER;
else process.env.CODEMAN_MULTIUSER = savedMode;
await fs.rm(tmpDir, { recursive: true, force: true }).catch(() => {});
});
/** A route app acting as `authUser`, whose broadcasts go through the server's routing. */
async function appAs(authUser: AuthUser | undefined): Promise<FastifyInstance> {
const app = Fastify({ logger: false });
await app.register(fastifyCookie);
await app.register(fastifyWebsocket);
app.decorateRequest('authUser', undefined);
app.addHook('onRequest', async (req) => {
req.authUser = authUser;
});
registerWebviewRoutes(app, {
broadcast: (event: string, data: unknown) => manager.broadcast(event, data, serverHint(event, data)),
tabLayouts: { webviewCreated: async () => {}, webviewDeleted: async () => {} },
} as never);
installRouteErrorHandler(app);
await app.ready();
apps.push(app);
return app;
}
it("multi-user: another user's SSE stream never sees a web-tab create, edit or delete", async () => {
process.env.CODEMAN_MULTIUSER = '1';
const aliceApp = await appAs({ username: 'alice', role: 'user' });
const alice = client();
const bob = client();
const admin = client();
manager.addClient(alice.reply, null, false, undefined, { username: 'alice', role: 'user' });
manager.addClient(bob.reply, null, false, undefined, { username: 'bob', role: 'user' });
manager.addClient(admin.reply, null, false, undefined, { username: 'root', role: 'admin' });
const created = await aliceApp.inject({
method: 'POST',
url: '/api/webviews',
payload: { name: 'Grafana', url: 'http://127.0.0.1:4000/' },
});
expect(created.statusCode).toBe(200);
const id = created.json().data.id as string;
const patched = await aliceApp.inject({ method: 'PATCH', url: `/api/webviews/${id}`, payload: { name: 'G2' } });
expect(patched.statusCode).toBe(200);
expect((await aliceApp.inject({ method: 'DELETE', url: `/api/webviews/${id}` })).statusCode).toBe(200);
const expected = ['created', 'updated', 'deleted'].map(
(action) => `event: webview:changed\ndata: ${JSON.stringify({ action, id, owner: 'alice' })}\n\n`
);
expect(alice.writes).toEqual(expected);
expect(admin.writes).toEqual(expected);
expect(bob.writes).toEqual([]);
});
it("multi-user: an admin editing a user's web tab notifies that user, not a bystander", async () => {
process.env.CODEMAN_MULTIUSER = '1';
const aliceApp = await appAs({ username: 'alice', role: 'user' });
const adminApp = await appAs({ username: 'root', role: 'admin' });
const id = (
await aliceApp.inject({
method: 'POST',
url: '/api/webviews',
payload: { name: 'G', url: 'http://127.0.0.1:4000/' },
})
).json().data.id as string;
const alice = client();
const bob = client();
manager.addClient(alice.reply, null, false, undefined, { username: 'alice', role: 'user' });
manager.addClient(bob.reply, null, false, undefined, { username: 'bob', role: 'user' });
expect((await adminApp.inject({ method: 'DELETE', url: `/api/webviews/${id}` })).statusCode).toBe(200);
expect(alice.writes).toEqual([
`event: webview:changed\ndata: ${JSON.stringify({ action: 'deleted', id, owner: 'alice' })}\n\n`,
]);
expect(bob.writes).toEqual([]);
});
it('single-user: every client still receives the event', async () => {
const soloApp = await appAs(undefined);
const tabA = client();
const tabB = client();
manager.addClient(tabA.reply, null, false, undefined, undefined);
manager.addClient(tabB.reply, null, false, undefined, undefined);
const created = await soloApp.inject({
method: 'POST',
url: '/api/webviews',
payload: { name: 'G', url: 'http://127.0.0.1:4000/' },
});
expect(created.statusCode).toBe(200);
expect(tabA.writes).toEqual([
`event: webview:changed\ndata: ${JSON.stringify({ action: 'created', id: created.json().data.id, owner: '@single' })}\n\n`,
]);
expect(tabB.writes).toEqual(tabA.writes);
});
});