From b46588f247f5f6bb72145260370f4544490a3bf4 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 11:33:00 +0800 Subject: [PATCH 01/46] fix(docker): install pnpm in the Compose server image `dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` (exit 127) on the server image. The agent image already installs pnpm for the same reason (#352). Pin pnpm@12.6.0 in the runtime-writable CLI prefix and note it in the DeepSeek doc. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- .changeset/serverpnpm3.md | 5 +++++ docker/server.Dockerfile | 8 ++++++++ docs/deepseek-integration.md | 4 +++- 3 files changed, 16 insertions(+), 1 deletion(-) create mode 100644 .changeset/serverpnpm3.md diff --git a/.changeset/serverpnpm3.md b/.changeset/serverpnpm3.md new file mode 100644 index 00000000..86fe7bb7 --- /dev/null +++ b/.changeset/serverpnpm3.md @@ -0,0 +1,5 @@ +--- +"aicodeman": patch +--- + +Install pnpm in the Docker Compose server image. `dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` in that image. Because this changes `server.Dockerfile`, the in-app updater will ask Compose deployments to rebuild the image (`Update-Codeman.sh`) rather than apply this release in place. diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index 09b4607a..8e957d26 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -214,11 +214,19 @@ RUN set -eux; \ # system directories for the root part of the start. ENV NPM_CONFIG_PREFIX=/opt/codeman-cli ENV PATH=$PATH:/opt/codeman-cli/bin +# pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which +# this image leaves to be installed at runtime, see SERVER_INTENTIONAL_OMISSIONS in +# test/docker-agent-image-coverage.test.ts) spawns a literal `pnpm` with no npm +# fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed +# with `dsh: pnpm not found on PATH` (exit 127) on this image. The agent image +# already carries it for the same reason (#352). It lives in the same +# runtime-writable prefix as the CLIs, so a session can update it in place. RUN npm install --global \ @anthropic-ai/claude-code@2.1.258 \ @google/gemini-cli@0.58.0 \ @openai/codex@0.152.1 \ opencode-ai@1.18.26 \ + pnpm@12.6.0 \ && npm cache clean --force # Keep the web server and every local Codeman session unprivileged. PUID and diff --git a/docs/deepseek-integration.md b/docs/deepseek-integration.md index 8232af92..6b4fd61c 100644 --- a/docs/deepseek-integration.md +++ b/docs/deepseek-integration.md @@ -53,7 +53,9 @@ that spawns a literal `pnpm` with no npm fallback, so without one it exits 127 w surfaces that same line as the install error. `npm install -g pnpm` (or `corepack enable pnpm`) is the fix. This is what broke the Docker agent image in [#352](https://github.com/Ark0N/Codeman/issues/352); the image now installs pnpm -alongside `dsh`. +alongside `dsh`. The Compose server image (`docker/server.Dockerfile`) does not +ship `dsh`, since it is installed at runtime, but it does ship pnpm so the UI +button works there too. Codeman's default is `@deepseek-harness-tui/dsh-tui` because it is by a wide margin the most used community TUI, it is MIT, and it implements the status From a5283c565db6aadd176852822f5ba3fd23d07a31 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 21:04:18 +0800 Subject: [PATCH 02/46] feat(docker): install uv and uvx in server and agent images Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- .changeset/uvxdocker1.md | 5 +++++ docker/agent.Dockerfile | 4 ++++ docker/server.Dockerfile | 4 ++++ 3 files changed, 13 insertions(+) create mode 100644 .changeset/uvxdocker1.md diff --git a/.changeset/uvxdocker1.md b/.changeset/uvxdocker1.md new file mode 100644 index 00000000..32e2d9b8 --- /dev/null +++ b/.changeset/uvxdocker1.md @@ -0,0 +1,5 @@ +--- +"aicodeman": patch +--- + +Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. The server image also carries `pnpm` for `dsh plugin`. diff --git a/docker/agent.Dockerfile b/docker/agent.Dockerfile index f7b46058..1082ee71 100644 --- a/docker/agent.Dockerfile +++ b/docker/agent.Dockerfile @@ -126,6 +126,10 @@ RUN set -eux; \ # A different order is a different RUN string, which is a different layer hash and # so a needless cache miss between a bare `docker build` and a scripted one. ARG CLI_NPM_PACKAGES="@anthropic-ai/claude-code opencode-ai @openai/codex @google/gemini-cli" +# uv/uvx: MCP servers are commonly launched with `uvx ` (e.g. the Nginx +# Proxy Manager MCP), and Codex failed to enable them with "uvx not found". Copied +# from the pinned upstream image into root-owned /usr/local/bin, never pip-installed. +COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/ RUN npm install -g ${CLI_NPM_PACKAGES} \ && npm cache clean --force diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index 8e957d26..939cf3fa 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -212,6 +212,10 @@ RUN set -eux; \ # minimal image of this exact shape). The four CLIs live only in this prefix, # so they still resolve; entrypoint.sh additionally pins its own PATH to the # system directories for the root part of the start. +# uv/uvx: MCP servers are commonly launched with `uvx ` (e.g. the Nginx +# Proxy Manager MCP), and Codex failed to enable them with "uvx not found". Copied +# from the pinned upstream image into root-owned /usr/local/bin, never pip-installed. +COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/ ENV NPM_CONFIG_PREFIX=/opt/codeman-cli ENV PATH=$PATH:/opt/codeman-cli/bin # pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which From 3e3a4612e66a0dbd024e86d061947badbe5fbc52 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:03:24 +0800 Subject: [PATCH 03/46] feat(docker): install libsecret-1-0 for the Azure DevOps MCP Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- .changeset/uvxdocker1.md | 2 ++ docker/agent.Dockerfile | 1 + docker/server.Dockerfile | 1 + 3 files changed, 4 insertions(+) diff --git a/.changeset/uvxdocker1.md b/.changeset/uvxdocker1.md index 32e2d9b8..b4c74474 100644 --- a/.changeset/uvxdocker1.md +++ b/.changeset/uvxdocker1.md @@ -3,3 +3,5 @@ --- Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. The server image also carries `pnpm` for `dsh plugin`. + +Both images also install `libsecret-1-0`, the native library the `keytar` dependency of the Azure DevOps MCP (`@azure-devops/mcp`) needs; without it the server crashes before answering the MCP initialize handshake. diff --git a/docker/agent.Dockerfile b/docker/agent.Dockerfile index 1082ee71..fa7dfdd9 100644 --- a/docker/agent.Dockerfile +++ b/docker/agent.Dockerfile @@ -17,6 +17,7 @@ FROM node:22-bookworm-slim RUN apt-get update \ && apt-get install -y --no-install-recommends \ git \ + libsecret-1-0 \ tmux \ ripgrep \ curl \ diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index 939cf3fa..99142abc 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -39,6 +39,7 @@ RUN apt-get update \ curl \ g++ \ git \ + libsecret-1-0 \ make \ openssh-client \ procps \ From e98127a80414736c47f5c349038dacd23d2948c2 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:17:22 +0800 Subject: [PATCH 04/46] feat(docker): add sudo to the agent image Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- .changeset/uvxdocker1.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.changeset/uvxdocker1.md b/.changeset/uvxdocker1.md index b4c74474..0535c953 100644 --- a/.changeset/uvxdocker1.md +++ b/.changeset/uvxdocker1.md @@ -5,3 +5,5 @@ Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. The server image also carries `pnpm` for `dsh plugin`. Both images also install `libsecret-1-0`, the native library the `keytar` dependency of the Azure DevOps MCP (`@azure-devops/mcp`) needs; without it the server crashes before answering the MCP initialize handshake. + +The agent image also installs `sudo` with passwordless access for the `agent` user, so a session can install system packages itself. The Compose server image is unchanged here: it runs with `no-new-privileges` and `cap_drop: ALL`, where `sudo` cannot work. From b070c9ee65fb9ef47af105bea96396fd5894bf45 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:17:38 +0800 Subject: [PATCH 05/46] feat(docker): install sudo with passwordless access for the agent user Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- docker/agent.Dockerfile | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docker/agent.Dockerfile b/docker/agent.Dockerfile index fa7dfdd9..359fd8d9 100644 --- a/docker/agent.Dockerfile +++ b/docker/agent.Dockerfile @@ -18,6 +18,7 @@ RUN apt-get update \ && apt-get install -y --no-install-recommends \ git \ libsecret-1-0 \ + sudo \ tmux \ ripgrep \ curl \ @@ -244,6 +245,8 @@ ENV HOME=/home/agent # Codeman's own host-side history/resume reads), and neither kind of artifact # creates its own parent directory. RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \ + && echo 'agent ALL=(ALL) NOPASSWD:ALL' > /etc/sudoers.d/agent \ + && chmod 0440 /etc/sudoers.d/agent \ && mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \ /home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent /home/agent/.grok \ /home/agent/.dsh /home/agent/.omp/agent \ From 10c263a5b832e33bfc9b76e4a2af49fe2959dc08 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:19:13 +0800 Subject: [PATCH 06/46] Revert "feat(docker): install sudo with passwordless access for the agent user" This reverts commit b070c9ee65fb9ef47af105bea96396fd5894bf45. --- docker/agent.Dockerfile | 3 --- 1 file changed, 3 deletions(-) diff --git a/docker/agent.Dockerfile b/docker/agent.Dockerfile index 359fd8d9..fa7dfdd9 100644 --- a/docker/agent.Dockerfile +++ b/docker/agent.Dockerfile @@ -18,7 +18,6 @@ RUN apt-get update \ && apt-get install -y --no-install-recommends \ git \ libsecret-1-0 \ - sudo \ tmux \ ripgrep \ curl \ @@ -245,8 +244,6 @@ ENV HOME=/home/agent # Codeman's own host-side history/resume reads), and neither kind of artifact # creates its own parent directory. RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \ - && echo 'agent ALL=(ALL) NOPASSWD:ALL' > /etc/sudoers.d/agent \ - && chmod 0440 /etc/sudoers.d/agent \ && mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \ /home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent /home/agent/.grok \ /home/agent/.dsh /home/agent/.omp/agent \ From d6c3386102482c2a18f02a194aa059426aa64394 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Thu, 24 Sep 2026 22:19:20 +0800 Subject: [PATCH 07/46] Revert "feat(docker): add sudo to the agent image" This reverts commit e98127a80414736c47f5c349038dacd23d2948c2. --- .changeset/uvxdocker1.md | 2 -- 1 file changed, 2 deletions(-) diff --git a/.changeset/uvxdocker1.md b/.changeset/uvxdocker1.md index 0535c953..b4c74474 100644 --- a/.changeset/uvxdocker1.md +++ b/.changeset/uvxdocker1.md @@ -5,5 +5,3 @@ Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. The server image also carries `pnpm` for `dsh plugin`. Both images also install `libsecret-1-0`, the native library the `keytar` dependency of the Azure DevOps MCP (`@azure-devops/mcp`) needs; without it the server crashes before answering the MCP initialize handshake. - -The agent image also installs `sudo` with passwordless access for the `agent` user, so a session can install system packages itself. The Compose server image is unchanged here: it runs with `no-new-privileges` and `cap_drop: ALL`, where `sudo` cannot work. From 8cef31086b487701686edbc1e5d505fe0558e726 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Fri, 25 Sep 2026 08:56:20 +0800 Subject: [PATCH 08/46] fix(docker): persist CLIs installed from Settings across container updates The image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is image content, so Update-Codeman.sh discarded every npm-installed CLI (dsh, pi). In the Compose container, POST /api/clis/:id/install now installs into ~/.local on the persistent home mount, and ~/.local/bin is appended to the image PATH. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- .changeset/cli9a1d.md | 5 +++++ docker/server.Dockerfile | 3 +++ src/web/routes/cli-registry-routes.ts | 7 +++++++ test/routes/cli-registry-routes.test.ts | 6 ++++++ 4 files changed, 21 insertions(+) create mode 100644 .changeset/cli9a1d.md diff --git a/.changeset/cli9a1d.md b/.changeset/cli9a1d.md new file mode 100644 index 00000000..0ecbf8df --- /dev/null +++ b/.changeset/cli9a1d.md @@ -0,0 +1,5 @@ +--- +"aicodeman": patch +--- + +Docker Compose: CLIs installed from Settings (DeepSeek, Pi and any other npm-based CLI) survive `Update-Codeman.sh`. The image's `NPM_CONFIG_PREFIX` (`/opt/codeman-cli`) is image content and was discarded when the container was recreated; `POST /api/clis/:id/install` now installs into `~/.local` on the persistent home mount when running in the container, and `~/.local/bin` is on the image PATH. diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index 99142abc..305fc392 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -219,6 +219,9 @@ RUN set -eux; \ COPY --from=ghcr.io/astral-sh/uv:0.9 /uv /uvx /usr/local/bin/ ENV NPM_CONFIG_PREFIX=/opt/codeman-cli ENV PATH=$PATH:/opt/codeman-cli/bin +# CLIs installed at runtime (Settings -> CLIs, npm redirected to ~/.local by installEnv()) live on the +# persistent home mount, so they survive a container recreate. Appended for the same reason as above. +ENV PATH=$PATH:/home/${CODEMAN_RUNTIME_USER}/.local/bin # pnpm is not an agent CLI: it is here because `dsh plugin` (DeepSeek Harness, which # this image leaves to be installed at runtime, see SERVER_INTENTIONAL_OMISSIONS in # test/docker-agent-image-coverage.test.ts) spawns a literal `pnpm` with no npm diff --git a/src/web/routes/cli-registry-routes.ts b/src/web/routes/cli-registry-routes.ts index 5b79b428..62a54e4f 100644 --- a/src/web/routes/cli-registry-routes.ts +++ b/src/web/routes/cli-registry-routes.ts @@ -217,6 +217,13 @@ export function installEnv(source: NodeJS.ProcessEnv = process.env): NodeJS.Proc for (const [key, value] of Object.entries(source)) { if (!key.startsWith('CODEMAN_')) env[key] = value; } + // ⚠️ In the Docker Compose deployment the image sets NPM_CONFIG_PREFIX=/opt/codeman-cli, which is IMAGE + // content: `Update-Codeman.sh` recreates the container and every CLI installed there (dsh, pi, ...) + // vanishes. HOME is the persistent bind mount and `~/.local/bin` is already on every resolver's search + // list, so npm-based installs are redirected there. curl|bash installers already target HOME. + if (source.CODEMAN_IN_CONTAINER === '1' && source.HOME) { + env.NPM_CONFIG_PREFIX = `${source.HOME}/.local`; + } return env; } diff --git a/test/routes/cli-registry-routes.test.ts b/test/routes/cli-registry-routes.test.ts index 34d7cda0..bc33b1f8 100644 --- a/test/routes/cli-registry-routes.test.ts +++ b/test/routes/cli-registry-routes.test.ts @@ -713,6 +713,12 @@ describe('registry writes are serialized and never clobber a file the reader wou } expect(installEnv({ CODEMAN_PASSWORD: 'x', HOME: '/h' })).toEqual({ HOME: '/h' }); }); + + it('redirects npm installs to the persistent HOME inside the Compose container', () => { + expect( + installEnv({ CODEMAN_IN_CONTAINER: '1', HOME: '/home/codeman', NPM_CONFIG_PREFIX: '/opt/codeman-cli' }) + ).toEqual({ HOME: '/home/codeman', NPM_CONFIG_PREFIX: '/home/codeman/.local' }); + }); }); /** From 95a3b87062a29f7ff57fc7372b94b0bba80c69cd Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Fri, 25 Sep 2026 11:14:47 +0800 Subject: [PATCH 09/46] chore: drop changesets already released in 1.33.1 The pnpm and uv/uvx changesets describe work upstream shipped in 1.33.1 (#485, #487), so keeping them would repeat those notes in the next release. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n --- .changeset/serverpnpm3.md | 5 ----- .changeset/uvxdocker1.md | 7 ------- 2 files changed, 12 deletions(-) delete mode 100644 .changeset/serverpnpm3.md delete mode 100644 .changeset/uvxdocker1.md diff --git a/.changeset/serverpnpm3.md b/.changeset/serverpnpm3.md deleted file mode 100644 index 86fe7bb7..00000000 --- a/.changeset/serverpnpm3.md +++ /dev/null @@ -1,5 +0,0 @@ ---- -"aicodeman": patch ---- - -Install pnpm in the Docker Compose server image. `dsh plugin` spawns a literal `pnpm` with no npm fallback, so the Run menu's "DeepSeek - add a terminal profile" button failed with `dsh: pnpm not found on PATH` in that image. Because this changes `server.Dockerfile`, the in-app updater will ask Compose deployments to rebuild the image (`Update-Codeman.sh`) rather than apply this release in place. diff --git a/.changeset/uvxdocker1.md b/.changeset/uvxdocker1.md deleted file mode 100644 index b4c74474..00000000 --- a/.changeset/uvxdocker1.md +++ /dev/null @@ -1,7 +0,0 @@ ---- -"aicodeman": patch ---- - -Install `uv` and `uvx` in the Compose server image and the agent image, so MCP servers launched with `uvx` (such as the Nginx Proxy Manager MCP) can be enabled by Codex instead of failing with `uvx` not found. The server image also carries `pnpm` for `dsh plugin`. - -Both images also install `libsecret-1-0`, the native library the `keytar` dependency of the Azure DevOps MCP (`@azure-devops/mcp`) needs; without it the server crashes before answering the MCP initialize handshake. From a9b48320a37ea41d882f6664f310ca4d54be97e3 Mon Sep 17 00:00:00 2001 From: Michael Grundberg Date: Fri, 25 Sep 2026 07:47:03 +0200 Subject: [PATCH 10/46] fix(session): alert for an agent waiting on artifact comments An agent that publishes an artifact arms a monitor for its comments and ends its turn. Claude Code shows that on the footer as `1 Artifact comment monitor`, and #473 put that chip on the list of background work, so the session counted as watching and its idle prompt opened already acknowledged. Unlike every other chip on the list, that monitor waits on the user: the agent hears nothing until somebody comments. Claude's `watchingLine` now refuses any footer that carries the chip, through a lookahead over the whole row, so a shell running beside the monitor cannot report the session as watching either. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/cli-registry.md | 3 ++ docs/wiki/Notifications-And-Approvals.md | 4 +++ src/config/cli-registry/stock.ts | 10 ++++++- test/session-watching.test.ts | 37 +++++++++++++++++++++++- 4 files changed, 52 insertions(+), 2 deletions(-) diff --git a/docs/cli-registry.md b/docs/cli-registry.md index 48cfee37..7298f385 100644 --- a/docs/cli-registry.md +++ b/docs/cli-registry.md @@ -66,6 +66,9 @@ agents` while a monitor, a backgrounded shell or a cloud session is live. Codema into `Session.watching`, and an idle prompt from such a session opens already acknowledged, so a pane waiting for its own background work never raises an alert a human cannot answer. Group 1 is the label, and a CLI that declares no pattern reports no background work. +Claude's Artifact comment monitor is the one chip that does not count. It waits for a human +to comment on a page the agent published, so Claude's pattern refuses any footer that +carries it, and the idle alert goes out as usual. Two CLIs declare such a row today, and they put it in different places. Claude writes its chip on the last row of the screen, so it keeps the default one-row window and anchors on diff --git a/docs/wiki/Notifications-And-Approvals.md b/docs/wiki/Notifications-And-Approvals.md index 41472d1a..e7964f6b 100644 --- a/docs/wiki/Notifications-And-Approvals.md +++ b/docs/wiki/Notifications-And-Approvals.md @@ -127,6 +127,10 @@ plain prose is not a dialog, so an agent that starts a monitor and then writes " should I target?" is quiet along with the rest — check a watching session yourself if it has been quiet longer than the work it is waiting for should take. +An agent waiting for your comments on an artifact it published never counts as watching. +Claude shows that as "1 Artifact comment monitor", but the agent hears nothing until you +comment, so the session alerts you like any other quiet session. + ## The phone overview On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index 94d2e3b7..f491f5ee 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -237,7 +237,15 @@ const CLAUDE: CliEntry = { // carry a count. A footer that ever drew the chip as its only item would report no // watching rather than open that door. See `watchingLabel()` in // `session-activity.ts`. - watchingLine: String.raw`·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?|Artifact comment monitors?))`, + // ⚠️ An Artifact comment monitor is the one chip that waits on the user. The agent + // has published a page and hears nothing until somebody comments on it, so the + // lookahead refuses the whole row while that chip is on it, whatever else is + // running beside it. The `^` is what makes the lookahead judge the row once: + // without it the engine retries from each later position, and a start past the + // chip reports the shell beside it. The lookahead stops short of "monitor", so a + // footer cut off mid-chip still counts. Counting the chip as watching kept the + // idle alert quiet for a session that was waiting for a human. + watchingLine: String.raw`^(?!.*Artifact comment).*?·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?))`, }, requiresMux: false, // Claude installs Codeman's own hooks block into every workspace it runs in, so its diff --git a/test/session-watching.test.ts b/test/session-watching.test.ts index d520ed32..44b940c4 100644 --- a/test/session-watching.test.ts +++ b/test/session-watching.test.ts @@ -133,7 +133,6 @@ describe('watchingLabel', () => { '1 MCP task', '1 background dynamic workflow', '2 remote dynamic workflows', - '1 Artifact comment monitor', '2 teams', ]; for (const label of labels) { @@ -141,6 +140,42 @@ describe('watchingLabel', () => { } }); + it('reports no watching while the agent waits for comments on an artifact', () => { + // An agent that publishes an artifact arms a monitor for its comments and ends its + // turn. That monitor waits on the user, so the idle alert has to reach them. The + // singular footer is a live capture from 2026-09-25; the plural is assumed. + expect( + watchingLabel(pane('⏵⏵ bypass permissions on · 1 Artifact comment monitor · ← for agents'), CLAUDE_WATCHING) + ).toBeNull(); + expect( + watchingLabel(pane('⏵⏵ bypass permissions on · 2 Artifact comment monitors · ← for agents'), CLAUDE_WATCHING) + ).toBeNull(); + }); + + it('lets a comment monitor outrank other background work on the same row', () => { + // A shell beside the monitor is still running, but the agent needs the user all the + // same, and the chip order on the footer must not decide that. The second row is + // the one that needs the `^` in front of the lookahead. + expect( + watchingLabel( + pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comment monitor · ← for agents'), + CLAUDE_WATCHING + ) + ).toBeNull(); + expect( + watchingLabel( + pane('⏵⏵ bypass permissions on · 1 Artifact comment monitor · 1 shell · ← for agents'), + CLAUDE_WATCHING + ) + ).toBeNull(); + }); + + it('still refuses a footer cut off in the middle of the comment monitor', () => { + expect( + watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comment moni…'), CLAUDE_WATCHING) + ).toBeNull(); + }); + it('says nothing about a pane that is running nothing', () => { expect(watchingLabel(NOTHING_RUNNING, CLAUDE_WATCHING)).toBeNull(); expect(watchingLabel('', CLAUDE_WATCHING)).toBeNull(); From 8d358aaa263f7f532d4dd9bd0ddcaa7d9965e55d Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Fri, 25 Sep 2026 21:09:50 +0800 Subject: [PATCH 11/46] feat(docker): configure static git identity --- docker/.env.example | 6 +++ docker/README.md | 19 ++++++++++ docker/agent.Dockerfile | 14 +++++++ docker/docker-compose.yaml | 6 +++ docker/server.Dockerfile | 14 +++++++ scripts/lib/cli-catalog.mjs | 23 +++++++++++- src/docker-hosts.ts | 26 ++++++++++++- test/agent-image-build-args-parity.test.ts | 43 ++++++++++++++++++++++ 8 files changed, 149 insertions(+), 2 deletions(-) diff --git a/docker/.env.example b/docker/.env.example index 95531008..8e929697 100644 --- a/docker/.env.example +++ b/docker/.env.example @@ -15,6 +15,12 @@ TZ=Australia/Perth # this value rebuilds the image with a matching account. CODEMAN_RUNTIME_USER=codeman +# Required for Git commits made by Codeman and Docker-case agents. These values +# are written to each image's system Git configuration when it is rebuilt, so +# they remain available even when the runtime home directory is a fresh mount. +GIT_USER_NAME= +GIT_USER_EMAIL= + # Required. Persistent Codeman application data, CLI credentials, and session # state are stored here on the host and mounted at the runtime account's home # directory in the container. diff --git a/docker/README.md b/docker/README.md index 87aa2f2d..0358add4 100644 --- a/docker/README.md +++ b/docker/README.md @@ -67,6 +67,25 @@ two volumes are removed, by name within this Compose project; any volume a `docker-compose.override.yml` adds is left alone, and application data and case workspaces are host bind mounts, never touched either way. +## Git commit identity + +Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` before rebuilding: + +```sh +GIT_USER_NAME='Your Name' +GIT_USER_EMAIL='you@example.com' +``` + +Compose passes the values to the Codeman server build, and to the server process +when it builds Docker-case agent images. Both images write the pair to Git's +system configuration during their build, so commits retain the same identity +after a container or agent image is recreated. Set both values together; an +image build with only one value fails rather than using a partial identity. + +Run `bash docker/Start-Codeman.sh` after changing the server values. Rebuild an +existing agent image with `node scripts/build-agent-image.mjs --no-cache` in the +server container, then recreate any Docker cases that should use it. + ## Private repositories (GitHub and Azure DevOps) The images can include the GitHub CLI (`gh`) and the Azure CLI (`az`, with the `azure-devops` extension), wired into the system Git configuration as credential helpers, so Codeman can clone private repositories. Both are **opt-in and off by default**, and are turned on per host in `docker-compose.override.yml`. diff --git a/docker/agent.Dockerfile b/docker/agent.Dockerfile index fa7dfdd9..93cb1c3b 100644 --- a/docker/agent.Dockerfile +++ b/docker/agent.Dockerfile @@ -12,6 +12,9 @@ # writable even though the uid is not the baked 1000. FROM node:22-bookworm-slim +ARG GIT_USER_EMAIL= +ARG GIT_USER_NAME= + # Base toolchain. `curl` is needed for the hook callbacks (`curl -sk $CODEMAN_API_URL`), # `procps` for `ps`, `tmux` for the durable in-container session. RUN apt-get update \ @@ -27,6 +30,17 @@ RUN apt-get update \ openssh-client \ && rm -rf /var/lib/apt/lists/* +# Docker cases run with a fresh, container-owned home directory. Configure Git +# at the system level during the build so the identity supplied in docker/.env +# remains stable after an agent image rebuild. Refuse an incomplete identity. +RUN set -eux; \ + if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \ + test -n "${GIT_USER_NAME}"; \ + test -n "${GIT_USER_EMAIL}"; \ + git config --system user.name "${GIT_USER_NAME}"; \ + git config --system user.email "${GIT_USER_EMAIL}"; \ + fi + # GitHub CLI and Azure CLI (+ the azure-devops extension) with the same system # git credential helpers as docker/server.Dockerfile, so an agent in a Docker # case can clone and push to private GitHub / Azure DevOps repositories. The diff --git a/docker/docker-compose.yaml b/docker/docker-compose.yaml index d26b2a97..9b50e8f1 100644 --- a/docker/docker-compose.yaml +++ b/docker/docker-compose.yaml @@ -7,6 +7,8 @@ services: dockerfile: docker/server.Dockerfile args: CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER} + GIT_USER_EMAIL: ${GIT_USER_EMAIL:-} + GIT_USER_NAME: ${GIT_USER_NAME:-} PGID: ${PGID:-1000} PUID: ${PUID:-1000} image: ${CODEMAN_IMAGE} @@ -32,6 +34,10 @@ services: CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH} CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT} CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH} + # Passed through only so Codeman can use the same identity when it builds + # the Docker-case agent image. + GIT_USER_EMAIL: ${GIT_USER_EMAIL:-} + GIT_USER_NAME: ${GIT_USER_NAME:-} # Extra Host-header allowlist entries for a reverse-proxied deployment # (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it # defaults to empty rather than requiring a line in every .env. diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index 99142abc..10f58ec0 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -25,6 +25,8 @@ RUN npm ci \ FROM node:22-bookworm-slim ARG CODEMAN_RUNTIME_USER=codeman +ARG GIT_USER_EMAIL= +ARG GIT_USER_NAME= ARG PUID=1000 ARG PGID=1000 @@ -48,6 +50,18 @@ RUN apt-get update \ tmux \ && rm -rf /var/lib/apt/lists/* +# A runtime home is normally a bind mount, so user-level Git configuration is +# not durable across a fresh deployment. Keep the operator-supplied identity in +# the image's system config instead. Both values are required together to avoid +# producing commits with a misleading partial identity. +RUN set -eux; \ + if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \ + test -n "${GIT_USER_NAME}"; \ + test -n "${GIT_USER_EMAIL}"; \ + git config --system user.name "${GIT_USER_NAME}"; \ + git config --system user.email "${GIT_USER_EMAIL}"; \ + fi + # The Docker CLI, taken from the official image rather than Debian's `docker.io`. # That package is the full ENGINE: with --no-install-recommends it still pulls 15 # packages including containerd, runc, dmsetup and iptables, none of which a diff --git a/scripts/lib/cli-catalog.mjs b/scripts/lib/cli-catalog.mjs index 6380b3fe..23edb136 100644 --- a/scripts/lib/cli-catalog.mjs +++ b/scripts/lib/cli-catalog.mjs @@ -64,6 +64,12 @@ export const GIT_HOST_CLI_BUILD_ARGS = [ ['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'], ]; +/** Environment variables passed through to the agent image's system Git configuration. */ +export const GIT_IDENTITY_BUILD_ARGS = [ + ['GIT_USER_NAME', 'GIT_USER_NAME'], + ['GIT_USER_EMAIL', 'GIT_USER_EMAIL'], +]; + /** * The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable * contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same @@ -82,9 +88,24 @@ export function gitHostCliBuildArgPairs(env) { return pairs; } +/** The `--build-arg` pairs for Git identity, requiring either both values or neither. */ +export function gitIdentityBuildArgPairs(env) { + const pairs = GIT_IDENTITY_BUILD_ARGS.map(([envName, argName]) => [argName, env[envName] ?? '']); + const configured = pairs.filter(([, value]) => value !== ''); + if (configured.length === 0) return []; + if (configured.length !== pairs.length) { + throw new Error('GIT_USER_NAME and GIT_USER_EMAIL must both be set when configuring Git identity'); + } + return pairs; +} + /** The `--build-arg` pairs the agent image takes. PURE given `env`. */ export function agentImageBuildArgPairs(catalog, env = process.env) { - return [['CLI_NPM_PACKAGES', agentImageNpmPackages(catalog).join(' ')], ...gitHostCliBuildArgPairs(env)]; + return [ + ['CLI_NPM_PACKAGES', agentImageNpmPackages(catalog).join(' ')], + ...gitHostCliBuildArgPairs(env), + ...gitIdentityBuildArgPairs(env), + ]; } /** Read the committed catalogue. IO. */ diff --git a/src/docker-hosts.ts b/src/docker-hosts.ts index 5509e89b..269dd540 100644 --- a/src/docker-hosts.ts +++ b/src/docker-hosts.ts @@ -615,6 +615,12 @@ export const GIT_HOST_CLI_BUILD_ARGS: ReadonlyArray = ['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'], ]; +/** Environment variables passed through to the agent image's system Git configuration. */ +export const GIT_IDENTITY_BUILD_ARGS: ReadonlyArray = [ + ['GIT_USER_NAME', 'GIT_USER_NAME'], + ['GIT_USER_EMAIL', 'GIT_USER_EMAIL'], +]; + /** * The `--build-arg` pairs for the optional git-host CLIs. PURE. An unset or empty variable * contributes NOTHING, so the Dockerfile's own default (off) applies and the argv is the same @@ -633,9 +639,27 @@ export function gitHostCliBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, return pairs; } +/** + * The `--build-arg` pairs for a configured Git identity. An absent pair leaves + * Git unconfigured, preserving existing deployments; a partial pair is refused. + */ +export function gitIdentityBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, string]> { + const pairs = GIT_IDENTITY_BUILD_ARGS.map(([envName, argName]) => [argName, env[envName] ?? ''] as [string, string]); + const configured = pairs.filter(([, value]) => value !== ''); + if (configured.length === 0) return []; + if (configured.length !== pairs.length) { + throw new Error('GIT_USER_NAME and GIT_USER_EMAIL must both be set when configuring Git identity'); + } + return pairs; +} + /** The `--build-arg` pairs the agent image takes. PURE given `env`. */ export function agentImageBuildArgPairs(env: NodeJS.ProcessEnv = process.env): Array<[string, string]> { - return [['CLI_NPM_PACKAGES', agentImageNpmPackages().join(' ')], ...gitHostCliBuildArgPairs(env)]; + return [ + ['CLI_NPM_PACKAGES', agentImageNpmPackages().join(' ')], + ...gitHostCliBuildArgPairs(env), + ...gitIdentityBuildArgPairs(env), + ]; } // ========== Credential mount resolution (IO) ========== diff --git a/test/agent-image-build-args-parity.test.ts b/test/agent-image-build-args-parity.test.ts index 030235a8..1f29a4a9 100644 --- a/test/agent-image-build-args-parity.test.ts +++ b/test/agent-image-build-args-parity.test.ts @@ -22,6 +22,8 @@ import { agentImageNpmPackages as mjsPackages, GIT_HOST_CLI_BUILD_ARGS as mjsGitHostArgs, gitHostCliBuildArgPairs as mjsGitHostPairs, + GIT_IDENTITY_BUILD_ARGS as mjsGitIdentityArgs, + gitIdentityBuildArgPairs as mjsGitIdentityPairs, } from '../scripts/lib/cli-catalog.mjs'; import { agentImageBuildArgPairs as tsPairs, @@ -29,6 +31,8 @@ import { agentImageNpmPackages as tsPackages, GIT_HOST_CLI_BUILD_ARGS as tsGitHostArgs, gitHostCliBuildArgPairs as tsGitHostPairs, + GIT_IDENTITY_BUILD_ARGS as tsGitIdentityArgs, + gitIdentityBuildArgPairs as tsGitIdentityPairs, } from '../src/docker-hosts.js'; const CATALOG = JSON.parse(readFileSync(fileURLToPath(new URL('../config/clis.stock.json', import.meta.url)), 'utf-8')); @@ -145,3 +149,42 @@ describe('optional gh / az in the agent image: both producers pass the same swit } }); }); + +describe('Git identity in the agent image: both producers pass the same settings', () => { + it('maps the Git environment variables to matching Dockerfile ARGs', () => { + expect(tsGitIdentityArgs).toEqual(mjsGitIdentityArgs); + expect(tsGitIdentityArgs).toEqual([ + ['GIT_USER_NAME', 'GIT_USER_NAME'], + ['GIT_USER_EMAIL', 'GIT_USER_EMAIL'], + ]); + }); + + it('passes a complete identity and omits an absent identity', () => { + const identity = { GIT_USER_NAME: 'Ada Lovelace', GIT_USER_EMAIL: 'ada@example.com' }; + const expected: Array<[string, string]> = [ + ['GIT_USER_NAME', 'Ada Lovelace'], + ['GIT_USER_EMAIL', 'ada@example.com'], + ]; + expect(tsGitIdentityPairs(identity)).toEqual(expected); + expect(mjsGitIdentityPairs(identity)).toEqual(expected); + expect(tsGitIdentityPairs({})).toEqual([]); + expect(mjsGitIdentityPairs({})).toEqual([]); + }); + + it('refuses a partial identity in both build paths', () => { + for (const identity of [{ GIT_USER_NAME: 'Ada Lovelace' }, { GIT_USER_EMAIL: 'ada@example.com' }]) { + expect(() => tsGitIdentityPairs(identity)).toThrow(/GIT_USER_NAME and GIT_USER_EMAIL/); + expect(() => mjsGitIdentityPairs(identity)).toThrow(/GIT_USER_NAME and GIT_USER_EMAIL/); + } + }); + + it('both Dockerfiles configure system Git identity from the build arguments', () => { + for (const file of ['../docker/agent.Dockerfile', '../docker/server.Dockerfile']) { + const dockerfile = readFileSync(fileURLToPath(new URL(file, import.meta.url)), 'utf-8'); + expect(dockerfile, file).toMatch(/^ARG GIT_USER_NAME=$/m); + expect(dockerfile, file).toMatch(/^ARG GIT_USER_EMAIL=$/m); + expect(dockerfile, file).toContain('git config --system user.name "${GIT_USER_NAME}"'); + expect(dockerfile, file).toContain('git config --system user.email "${GIT_USER_EMAIL}"'); + } + }); +}); From bf73a8473219c13f1755813bb76546d89a0d3399 Mon Sep 17 00:00:00 2001 From: Devvyn <22340871+opticon454@users.noreply.github.com> Date: Fri, 25 Sep 2026 22:27:26 +0800 Subject: [PATCH 12/46] fix(docker): address git identity review --- .changeset/static-git-identity.md | 5 ++++ docker/.env.example | 8 +++--- docker/README.md | 4 ++- docker/agent.Dockerfile | 29 +++++++++++----------- docker/docker-compose.yaml | 4 +-- docker/server.Dockerfile | 29 +++++++++++----------- docs/docker-cases.md | 2 ++ scripts/lib/cli-catalog.mjs | 6 ++--- src/docker-hosts.ts | 8 +++--- test/agent-image-build-args-parity.test.ts | 18 +++++++++----- 10 files changed, 65 insertions(+), 48 deletions(-) create mode 100644 .changeset/static-git-identity.md diff --git a/.changeset/static-git-identity.md b/.changeset/static-git-identity.md new file mode 100644 index 00000000..345e124a --- /dev/null +++ b/.changeset/static-git-identity.md @@ -0,0 +1,5 @@ +--- +"aicodeman": minor +--- + +Docker deployments can configure a static Git commit identity for server and Docker-case agent images. diff --git a/docker/.env.example b/docker/.env.example index 8e929697..fbdb1229 100644 --- a/docker/.env.example +++ b/docker/.env.example @@ -15,11 +15,11 @@ TZ=Australia/Perth # this value rebuilds the image with a matching account. CODEMAN_RUNTIME_USER=codeman -# Required for Git commits made by Codeman and Docker-case agents. These values +# Optional Git identity for commits made by Codeman and Docker-case agents. These values # are written to each image's system Git configuration when it is rebuilt, so -# they remain available even when the runtime home directory is a fresh mount. -GIT_USER_NAME= -GIT_USER_EMAIL= +# deployments can configure a consistent default. Set both values together. +# GIT_USER_NAME= +# GIT_USER_EMAIL= # Required. Persistent Codeman application data, CLI credentials, and session # state are stored here on the host and mounted at the runtime account's home diff --git a/docker/README.md b/docker/README.md index 0358add4..5328fb65 100644 --- a/docker/README.md +++ b/docker/README.md @@ -80,7 +80,9 @@ Compose passes the values to the Codeman server build, and to the server process when it builds Docker-case agent images. Both images write the pair to Git's system configuration during their build, so commits retain the same identity after a container or agent image is recreated. Set both values together; an -image build with only one value fails rather than using a partial identity. +image build with only one value fails rather than using a partial identity. An +identity already present in `CODEMAN_APPDATA_PATH`'s `~/.gitconfig` overrides +the server image's system-level default. Run `bash docker/Start-Codeman.sh` after changing the server values. Rebuild an existing agent image with `node scripts/build-agent-image.mjs --no-cache` in the diff --git a/docker/agent.Dockerfile b/docker/agent.Dockerfile index 93cb1c3b..afc9c6e1 100644 --- a/docker/agent.Dockerfile +++ b/docker/agent.Dockerfile @@ -12,9 +12,6 @@ # writable even though the uid is not the baked 1000. FROM node:22-bookworm-slim -ARG GIT_USER_EMAIL= -ARG GIT_USER_NAME= - # Base toolchain. `curl` is needed for the hook callbacks (`curl -sk $CODEMAN_API_URL`), # `procps` for `ps`, `tmux` for the durable in-container session. RUN apt-get update \ @@ -30,17 +27,6 @@ RUN apt-get update \ openssh-client \ && rm -rf /var/lib/apt/lists/* -# Docker cases run with a fresh, container-owned home directory. Configure Git -# at the system level during the build so the identity supplied in docker/.env -# remains stable after an agent image rebuild. Refuse an incomplete identity. -RUN set -eux; \ - if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \ - test -n "${GIT_USER_NAME}"; \ - test -n "${GIT_USER_EMAIL}"; \ - git config --system user.name "${GIT_USER_NAME}"; \ - git config --system user.email "${GIT_USER_EMAIL}"; \ - fi - # GitHub CLI and Azure CLI (+ the azure-devops extension) with the same system # git credential helpers as docker/server.Dockerfile, so an agent in a Docker # case can clone and push to private GitHub / Azure DevOps repositories. The @@ -268,6 +254,21 @@ RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \ && chgrp -R 0 /home/agent \ && chmod -R g=u /home/agent +# Docker cases have a fresh, container-owned home directory. Declare the +# optional identity here so changing it invalidates only this final layer, then +# configure Git's system defaults. A user-level config still takes precedence. +ARG GIT_USER_EMAIL= +ARG GIT_USER_NAME= +RUN set -eux; \ + if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \ + if [ -z "${GIT_USER_NAME}" ] || [ -z "${GIT_USER_EMAIL}" ]; then \ + echo 'Git user name and email must both be set when configuring Git identity' >&2; \ + exit 1; \ + fi; \ + git config --system user.name "${GIT_USER_NAME}"; \ + git config --system user.email "${GIT_USER_EMAIL}"; \ + fi + USER agent WORKDIR /home/agent diff --git a/docker/docker-compose.yaml b/docker/docker-compose.yaml index 9b50e8f1..db6e17c3 100644 --- a/docker/docker-compose.yaml +++ b/docker/docker-compose.yaml @@ -36,8 +36,8 @@ services: CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH} # Passed through only so Codeman can use the same identity when it builds # the Docker-case agent image. - GIT_USER_EMAIL: ${GIT_USER_EMAIL:-} - GIT_USER_NAME: ${GIT_USER_NAME:-} + CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: ${GIT_USER_EMAIL:-} + CODEMAN_AGENT_IMAGE_GIT_USER_NAME: ${GIT_USER_NAME:-} # Extra Host-header allowlist entries for a reverse-proxied deployment # (docker/README.md, "Reverse-proxy host allowlist"). Optional, so it # defaults to empty rather than requiring a line in every .env. diff --git a/docker/server.Dockerfile b/docker/server.Dockerfile index 10f58ec0..fd2cdf65 100644 --- a/docker/server.Dockerfile +++ b/docker/server.Dockerfile @@ -25,8 +25,6 @@ RUN npm ci \ FROM node:22-bookworm-slim ARG CODEMAN_RUNTIME_USER=codeman -ARG GIT_USER_EMAIL= -ARG GIT_USER_NAME= ARG PUID=1000 ARG PGID=1000 @@ -50,18 +48,6 @@ RUN apt-get update \ tmux \ && rm -rf /var/lib/apt/lists/* -# A runtime home is normally a bind mount, so user-level Git configuration is -# not durable across a fresh deployment. Keep the operator-supplied identity in -# the image's system config instead. Both values are required together to avoid -# producing commits with a misleading partial identity. -RUN set -eux; \ - if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \ - test -n "${GIT_USER_NAME}"; \ - test -n "${GIT_USER_EMAIL}"; \ - git config --system user.name "${GIT_USER_NAME}"; \ - git config --system user.email "${GIT_USER_EMAIL}"; \ - fi - # The Docker CLI, taken from the official image rather than Debian's `docker.io`. # That package is the full ENGINE: with --no-install-recommends it still pulls 15 # packages including containerd, runc, dmsetup and iptables, none of which a @@ -313,6 +299,21 @@ EXPOSE 3000 COPY docker/entrypoint.sh /usr/local/bin/entrypoint.sh RUN chmod 0755 /usr/local/bin/entrypoint.sh +# Declare the optional identity immediately before configuring it so a change +# invalidates only this final layer. This is declarative setup: a persisted +# ~/.gitconfig in CODEMAN_APPDATA_PATH still overrides the system-level values. +ARG GIT_USER_EMAIL= +ARG GIT_USER_NAME= +RUN set -eux; \ + if [ -n "${GIT_USER_NAME}" ] || [ -n "${GIT_USER_EMAIL}" ]; then \ + if [ -z "${GIT_USER_NAME}" ] || [ -z "${GIT_USER_EMAIL}" ]; then \ + echo 'Git user name and email must both be set when configuring Git identity' >&2; \ + exit 1; \ + fi; \ + git config --system user.name "${GIT_USER_NAME}"; \ + git config --system user.email "${GIT_USER_EMAIL}"; \ + fi + ENTRYPOINT ["/usr/local/bin/entrypoint.sh"] CMD ["node", "dist/index.js", "web"] diff --git a/docs/docker-cases.md b/docs/docker-cases.md index f8a1459a..4b8bd7bc 100644 --- a/docs/docker-cases.md +++ b/docs/docker-cases.md @@ -83,6 +83,8 @@ Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.jso The image can also carry the GitHub CLI (`gh`) and the Azure CLI (`az` + the `azure-devops` extension, in `AZURE_EXTENSION_DIR=/opt/az-extensions` so it stays out of the seeded HOME), wired into the system git config as credential helpers for github.com and dev.azure.com / *.visualstudio.com, exactly as in `docker/server.Dockerfile`. Their sign-ins are seeded per-FILE like pi's: `~/.config/gh/{hosts.yml,config.yml}` and `~/.azure/{azureProfile.json,msal_token_cache.json,service_principal_entries.json,clouds.config,config}`, never `~/.azure`'s logs, command index or extensions. A token kept in a desktop keyring, or in the encrypted MSAL cache az uses on Windows/macOS, is not in those files and does not carry. None of the three is version-pinned; the `--no-cache` rebuild recommended above is also what refreshes them. Both CLIs are opt-in and OFF by default: `CODEMAN_AGENT_IMAGE_INSTALL_GH=1` / `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` in the environment of `scripts/build-agent-image.mjs`, or of the Codeman server for its own auto-build (in the Compose deployment, `environment:` in `docker-compose.override.yml`), become the `CODEMAN_INSTALL_GH` / `CODEMAN_INSTALL_AZ` build args and put that CLI, its extension and its helper entry into the image. Unset passes nothing, so a default build's argv is unchanged and the image has neither. The sign-in seeds follow the same switches, read when a case container is created: `.config/gh` only with `CODEMAN_AGENT_IMAGE_INSTALL_GH=1`, `.azure` only with `CODEMAN_AGENT_IMAGE_INSTALL_AZ=1` (`enabledByEnv` in `CRED_STORES`), never merely because the files exist. Seeds are create-time mounts and deliberately not part of the config hash (hashing them would trip the drift gate for every case), so an existing case container picks them up only when it is recreated. +Set `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` and `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` together to configure the agent image's Git identity. Rebuild an existing `codeman/agent:base` with `node scripts/build-agent-image.mjs --no-cache`, then recreate Docker-case containers so they use the rebuilt image. + ## Quickest path: one-click "Run in Docker" On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in. diff --git a/scripts/lib/cli-catalog.mjs b/scripts/lib/cli-catalog.mjs index 23edb136..fd9c53f6 100644 --- a/scripts/lib/cli-catalog.mjs +++ b/scripts/lib/cli-catalog.mjs @@ -66,8 +66,8 @@ export const GIT_HOST_CLI_BUILD_ARGS = [ /** Environment variables passed through to the agent image's system Git configuration. */ export const GIT_IDENTITY_BUILD_ARGS = [ - ['GIT_USER_NAME', 'GIT_USER_NAME'], - ['GIT_USER_EMAIL', 'GIT_USER_EMAIL'], + ['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'], + ['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'], ]; /** @@ -94,7 +94,7 @@ export function gitIdentityBuildArgPairs(env) { const configured = pairs.filter(([, value]) => value !== ''); if (configured.length === 0) return []; if (configured.length !== pairs.length) { - throw new Error('GIT_USER_NAME and GIT_USER_EMAIL must both be set when configuring Git identity'); + throw new Error('Git user name and email must both be set when configuring Git identity'); } return pairs; } diff --git a/src/docker-hosts.ts b/src/docker-hosts.ts index 269dd540..3b6f7744 100644 --- a/src/docker-hosts.ts +++ b/src/docker-hosts.ts @@ -617,8 +617,8 @@ export const GIT_HOST_CLI_BUILD_ARGS: ReadonlyArray = /** Environment variables passed through to the agent image's system Git configuration. */ export const GIT_IDENTITY_BUILD_ARGS: ReadonlyArray = [ - ['GIT_USER_NAME', 'GIT_USER_NAME'], - ['GIT_USER_EMAIL', 'GIT_USER_EMAIL'], + ['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'], + ['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'], ]; /** @@ -648,7 +648,7 @@ export function gitIdentityBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, const configured = pairs.filter(([, value]) => value !== ''); if (configured.length === 0) return []; if (configured.length !== pairs.length) { - throw new Error('GIT_USER_NAME and GIT_USER_EMAIL must both be set when configuring Git identity'); + throw new Error('Git user name and email must both be set when configuring Git identity'); } return pairs; } @@ -1228,7 +1228,7 @@ function buildAgentImage( try { buildArgPairs = agentImageBuildArgPairs(); } catch (err) { - // A malformed CODEMAN_AGENT_IMAGE_INSTALL_* value: report it like any other build failure. + // A malformed CODEMAN_AGENT_IMAGE_* value: report it like any other build failure. return Promise.resolve({ ok: false, built: false, alreadyPresent: false, error: String((err as Error).message) }); } const argv = dockerEngineArgv(docker); diff --git a/test/agent-image-build-args-parity.test.ts b/test/agent-image-build-args-parity.test.ts index 1f29a4a9..3ca358d0 100644 --- a/test/agent-image-build-args-parity.test.ts +++ b/test/agent-image-build-args-parity.test.ts @@ -154,13 +154,16 @@ describe('Git identity in the agent image: both producers pass the same settings it('maps the Git environment variables to matching Dockerfile ARGs', () => { expect(tsGitIdentityArgs).toEqual(mjsGitIdentityArgs); expect(tsGitIdentityArgs).toEqual([ - ['GIT_USER_NAME', 'GIT_USER_NAME'], - ['GIT_USER_EMAIL', 'GIT_USER_EMAIL'], + ['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'], + ['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'], ]); }); it('passes a complete identity and omits an absent identity', () => { - const identity = { GIT_USER_NAME: 'Ada Lovelace', GIT_USER_EMAIL: 'ada@example.com' }; + const identity = { + CODEMAN_AGENT_IMAGE_GIT_USER_NAME: 'Ada Lovelace', + CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: 'ada@example.com', + }; const expected: Array<[string, string]> = [ ['GIT_USER_NAME', 'Ada Lovelace'], ['GIT_USER_EMAIL', 'ada@example.com'], @@ -172,9 +175,12 @@ describe('Git identity in the agent image: both producers pass the same settings }); it('refuses a partial identity in both build paths', () => { - for (const identity of [{ GIT_USER_NAME: 'Ada Lovelace' }, { GIT_USER_EMAIL: 'ada@example.com' }]) { - expect(() => tsGitIdentityPairs(identity)).toThrow(/GIT_USER_NAME and GIT_USER_EMAIL/); - expect(() => mjsGitIdentityPairs(identity)).toThrow(/GIT_USER_NAME and GIT_USER_EMAIL/); + for (const identity of [ + { CODEMAN_AGENT_IMAGE_GIT_USER_NAME: 'Ada Lovelace' }, + { CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: 'ada@example.com' }, + ]) { + expect(() => tsGitIdentityPairs(identity)).toThrow(/Git user name and email/); + expect(() => mjsGitIdentityPairs(identity)).toThrow(/Git user name and email/); } }); From 9676e90133deebbc4fc3b3ad8426aa2accb9d7d5 Mon Sep 17 00:00:00 2001 From: timkjr Date: Fri, 25 Sep 2026 19:05:16 -0500 Subject: [PATCH 13/46] fix(terminal): let a Shell pane's scroll-up reach tmux history A burst of output leaves a Shell pane with about one screen of browser scrollback, because tmux repaints the burst instead of scrolling it, while tmux itself keeps every line. Shell declined the scroll-to-top re-pull other modes use, and the Load full history button renders only once a replay was truncated, so a Shell tab under 1 MiB could not scroll back at all. The scroll gesture now pulls ?full=1&tail=TERMINAL_TAIL_SIZE, the same bound a tab switch loads; the route's existing tail cut marks longer histories 'tail', so the banner still offers the unbounded pull. A window no longer than the browser's buffer is not rewritten. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/architecture-invariants.md | 2 +- docs/wiki/The-Dashboard.md | 5 +- src/web/public/app.js | 34 +++++- src/web/public/terminal-ui.js | 6 +- test/history-truncation-notice.test.ts | 9 +- test/shell-scroll-history-pull.test.ts | 142 +++++++++++++++++++++++++ 6 files changed, 184 insertions(+), 14 deletions(-) create mode 100644 test/shell-scroll-history-pull.test.ts diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 6a25c5b3..742419f6 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -205,7 +205,7 @@ Further detail: the `: ` form (`w3-myapp: fix the login redirect` ### Full-scrollback replay -**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. +**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. **A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, because a retry would capture the identical frame; and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **The clamp used to manufacture that equality, and no longer can** (#464). This paragraph previously explained it as the signature of a clamp: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` did not, so a terminal under 40 columns or 10 rows reported a pane permanently bigger than itself and would replay on every tab switch. That divergence is fixed at the source — `syncTerminalGeometry()` (terminal-ui.js) fits, floors and APPLIES in one step, so the browser terminal IS the size it reports. The only remaining reason the two can differ is a resize the server declined, which the server now reports back (`Session.ptyGeometry`, the `{"t":"zc"}` frame and the resize response) for the client to adopt by COLUMNS. The equality guard stays, for the plain case of a pane already at the requested size. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. diff --git a/docs/wiki/The-Dashboard.md b/docs/wiki/The-Dashboard.md index cf1f415e..323299fe 100644 --- a/docs/wiki/The-Dashboard.md +++ b/docs/wiki/The-Dashboard.md @@ -151,8 +151,9 @@ Worth knowing: - **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open. Shell sessions open from a bounded recent tail so a large transcript cannot stall tab - switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling - and automatic output recovery stay within the bounded browser buffer. + switching. Scrolling to the top of a Shell pane pulls the most recent 1 MiB of its tmux + history; press **Load full history** to pull the rest explicitly. Automatic output + recovery stays within the bounded browser buffer. - **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is always local scrollback. Other CLIs scroll locally. diff --git a/src/web/public/app.js b/src/web/public/app.js index 05577c34..1da813be 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -6320,10 +6320,15 @@ class CodemanApp { if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return; if (this.detachedSessions?.has(sessionId)) return; const session = this.sessions.get(sessionId); - // A shell's full capture can be many megabytes. Replaying it from an - // ordinary scroll gesture blocks xterm's main thread, so keep that cost - // behind the explicit "Load full history" button. - if (!force && session?.mode === 'shell') return; + // A shell's full capture can be many megabytes, and replaying all of it from + // an ordinary scroll gesture blocks xterm's main thread. So a shell scroll + // pulls a BOUNDED window of tmux's full history (the same 1 MiB a tab switch + // loads, but of the scrollback rather than the visible frame) and the + // unbounded pull stays behind the "Load full history" button. Declining + // outright left a shell pane about one screen of browser scrollback after any + // burst, and the button only renders once a replay was truncated, so a young + // shell tab had no way back to output tmux was still holding. + const boundedShellPull = !force && session?.mode === 'shell'; const now = Date.now(); // Momentum scrolling fires this dozens of times per flick, and a burst of new // output is the normal reason to want a re-pull, so cooldown rather than latch. @@ -6336,7 +6341,12 @@ class CodemanApp { this._fullHistoryRepullInFlight = true; try { const requestStartedAt = performance.now(); - const capture = await this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true }); + const capture = await this._fetchTerminalCapture( + boundedShellPull + ? `/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}` + : `/api/sessions/${sessionId}/terminal?full=1`, + { full: true } + ); const headersReceivedAt = capture.headersAt; const payload = capture.json?.data ?? {}; const bodyParsedAt = performance.now(); @@ -6368,6 +6378,20 @@ class CodemanApp { this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); return; } + // A bounded window no longer than the browser's buffer buys nothing, and + // resetting to rewrite it would jump the viewport on every scroll that + // outlasts the cooldown at the top. An untruncated window IS all of tmux's + // history, so nothing is missing; a window cut at the tail size would trade + // more old rows than it recovers, and its 'tail' truncation keeps the + // banner offering the unbounded pull. Not latched as useless: the next + // burst of output can put more history in tmux than the browser has. + if ( + boundedShellPull && + this._estimateReplayRows(buffer, this.terminal.cols) <= this.terminal.buffer.active.length + ) { + this._setHistoryTruncation(sessionId, payload); + return; + } this._setHistoryTruncation(sessionId, payload); this._fullHistoryRepullUseless?.delete(sessionId); const rowsBefore = this.terminal.buffer.active.length; diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index 903661e5..413fe99b 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -3304,9 +3304,9 @@ Object.assign(CodemanApp.prototype, { /** * Post-scroll companion to _noteTerminalUserScroll: hitting the TOP of the * buffer while scrolling up gives the app a chance to pull the rest of tmux's - * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions decline - * automatic pulls because their captures can be large; their banner button is - * the explicit path. Must be called AFTER scrollLines(), since the check is on + * scrollback (issue #205, see _maybeRefetchFullHistory). Shell sessions pull a + * bounded window because their captures can be large; their banner button is + * the unbounded path. Must be called AFTER scrollLines(), since the check is on * the resulting position, and it is deliberately not folded into * _noteTerminalUserScroll for exactly that reason. */ diff --git a/test/history-truncation-notice.test.ts b/test/history-truncation-notice.test.ts index b24fe3b6..50eb5735 100644 --- a/test/history-truncation-notice.test.ts +++ b/test/history-truncation-notice.test.ts @@ -122,7 +122,7 @@ describe('the in-terminal truncation line is gone (static guard)', () => { expect(app).not.toContain('earlier output truncated for performance'); }); - it('loads a bounded shell tail first and keeps full history user-triggered', () => { + it('loads a bounded shell tail first and keeps unbounded full history user-triggered', () => { const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); expect(app).toContain("session?.mode !== 'shell' && !this._fullHistoryLoaded.has(sessionId)"); expect(app).toContain("!restoredSnapshot && session?.mode !== 'shell'"); @@ -131,10 +131,13 @@ describe('the in-terminal truncation line is gone (static guard)', () => { // an abort deadline (a `?full=1` body can be megabytes and used to hang // indefinitely on a stalled mobile link). The URL and the full-vs-tail // decision this guard exists to pin are unchanged. - expect(app).toContain('this._fetchTerminalCapture(`/api/sessions/${sessionId}/terminal?full=1`, { full: true })'); + expect(app).toContain(': `/api/sessions/${sessionId}/terminal?full=1`,\n { full: true }'); expect(app).toContain("if (this.sessions.get(sessionId)?.mode !== 'shell')"); expect(app).toContain("if (session?.mode === 'shell')"); - expect(app).toContain("if (!force && session?.mode === 'shell') return;"); + // A shell scroll gesture pulls a BOUNDED window of full history; only the + // button pulls all of it (behaviour pinned in shell-scroll-history-pull.test.ts). + expect(app).toContain("const boundedShellPull = !force && session?.mode === 'shell';"); + expect(app).toContain('`/api/sessions/${sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`'); expect(app).toContain("trigger: force ? 'full-history-button' : 'full-history-scroll'"); }); diff --git a/test/shell-scroll-history-pull.test.ts b/test/shell-scroll-history-pull.test.ts new file mode 100644 index 00000000..229d8a0c --- /dev/null +++ b/test/shell-scroll-history-pull.test.ts @@ -0,0 +1,142 @@ +/** + * @fileoverview A shell pane's scroll-up must reach the history tmux still holds. + * + * tmux repaints a burst of output instead of scrolling it, so after `cat` of a + * file longer than the screen the browser holds about one screen of scrollback + * while tmux holds all of it. Other modes recover it by re-pulling `?full=1` + * when the wheel reaches the top (`_maybeRefetchFullHistory`, issue #205). Shell + * declined that gesture outright to keep a multi-megabyte capture off xterm's + * main thread, leaving only the "Load full history" button, and that button + * renders only once a replay was truncated. A young shell tab therefore had no + * way to scroll back at all. + * + * The gesture now pulls a BOUNDED window (`?full=1&tail=TERMINAL_TAIL_SIZE`), + * the button stays the unbounded path, and a window the browser already holds + * in full is not rewritten. + * + * The method is extracted from app.js and run in a `vm` against stubs (no jsdom + * on this box; see connection-indicator.test.ts), with the REAL row estimators + * from terminal-ui.js, which decide both the downgrade and the no-gain skip. + */ +import { readFileSync } from 'node:fs'; +import { performance } from 'node:perf_hooks'; +import { resolve } from 'node:path'; +import vm from 'node:vm'; +import { describe, expect, it, vi } from 'vitest'; + +const PUBLIC = resolve(import.meta.dirname, '../src/web/public'); +const TERMINAL_TAIL_SIZE = 1024 * 1024; + +function methodSource(source: string, method: string): string { + const start = source.search(new RegExp(`^ {2}(?:async )?${method}\\(`, 'm')); + expect(start, `${method} not found`).toBeGreaterThan(-1); + const next = /^ {2}(?:async )?[A-Za-z_$][\w$]*\(/m.exec(source.slice(start + 1)); + return next ? source.slice(start, start + 1 + next.index) : source.slice(start); +} + +/** Real terminal-ui.js mixin, for `_estimateReplayRows` / `_replayWouldShrinkBuffer`. */ +function loadTerminalMixin(): Record<string, unknown> { + const source = readFileSync(resolve(PUBLIC, 'terminal-ui.js'), 'utf8'); + const FakeCodemanApp = function () {} as unknown as { prototype: Record<string, unknown> }; + const context = vm.createContext({ + console, + performance, + setTimeout, + clearTimeout, + setInterval: vi.fn(), + clearInterval: vi.fn(), + requestAnimationFrame: vi.fn(), + CodemanApp: FakeCodemanApp, + window: { addEventListener: vi.fn(), removeEventListener: vi.fn() }, + document: { addEventListener: vi.fn() }, + }); + vm.runInContext(source, context); + return FakeCodemanApp.prototype; +} + +function loadRefetch(): (this: unknown, opts?: { force?: boolean }) => Promise<void> { + const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); + const body = methodSource(app, '_maybeRefetchFullHistory'); + const context = vm.createContext({ performance, TERMINAL_TAIL_SIZE, TERMINAL_CHUNK_SIZE: 32 * 1024 }); + return vm.runInContext(`({ ${body} })._maybeRefetchFullHistory`, context); +} + +const mixin = loadTerminalMixin(); +const refetch = loadRefetch(); +const lines = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\r\n'); + +function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; capture: string }) { + const urls: string[] = []; + const app = { + activeSessionId: 's1', + sessions: new Map([['s1', { mode }]]), + detachedSessions: new Set<string>(), + _fullHistoryRepullInFlight: false, + _isLoadingBuffer: false, + _fullHistoryRepullAt: new Map<string, number>(), + _fullHistoryRepullUseless: new Set<string>(), + terminalBufferCache: new Map<string, string>(), + terminal: { + cols: 80, + rows: 30, + buffer: { active: { length: bufferRows } }, + scrollToLine: vi.fn(), + scrollToTop: vi.fn(), + }, + _estimateReplayRows: mixin._estimateReplayRows, + _replayWouldShrinkBuffer: mixin._replayWouldShrinkBuffer, + _fetchTerminalCapture: vi.fn(async (url: string) => { + urls.push(url); + return { + headersAt: performance.now(), + headers: { get: () => '' }, + json: { data: { terminalBuffer: capture, source: 'mux-full-history' } }, + }; + }), + _recordTerminalLoadTiming: vi.fn(), + _logScrollRouting: vi.fn(), + _setHistoryTruncation: vi.fn(), + _resetTerminalForReplay: vi.fn(), + _bufferLoadFinishOpts: vi.fn(() => ({})), + chunkedTerminalWrite: vi.fn(async () => ({ parsedAt: performance.now(), bufferLength: 400, completed: true })), + _syncStickyScrollBaseline: vi.fn(), + }; + return { app, urls }; +} + +describe('shell scroll-up pulls a bounded window of tmux history', () => { + it('a shell scroll gesture requests full history bounded by the tail size', async () => { + const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app); + expect(urls).toEqual([`/api/sessions/s1/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`]); + // It then actually replays the recovered history. + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + expect(app.chunkedTerminalWrite).toHaveBeenCalledTimes(1); + }); + + it('the Load full history button stays unbounded for a shell', async () => { + const { app, urls } = makeApp('shell', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app, { force: true }); + expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']); + }); + + it('other modes keep the unbounded scroll pull', async () => { + const { app, urls } = makeApp('claude', { bufferRows: 40, capture: lines(300) }); + await refetch.call(app); + expect(urls).toEqual(['/api/sessions/s1/terminal?full=1']); + }); + + it('a bounded window the browser already holds is not rewritten', async () => { + // Browser already has every row the window carries: resetting to rewrite + // it would jump the viewport on every scroll that outlasts the cooldown. + const { app } = makeApp('shell', { bufferRows: 320, capture: lines(300) }); + await refetch.call(app); + expect(app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); + // Not latched as useless: more output can put more history in tmux. + expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); + // The truncation state is still recorded, so a window capped at the tail + // size keeps offering the button. + expect(app._setHistoryTruncation).toHaveBeenCalledTimes(1); + }); +}); From f6aa50239fd44426d47c713750fcd402b0ebee92 Mon Sep 17 00:00:00 2001 From: timkjr <timkjr@k-lab.lan> Date: Fri, 25 Sep 2026 20:17:22 -0500 Subject: [PATCH 14/46] fix(terminal): skip a bounded Shell window before the downgrade guard A window cut at the tail size can be smaller than the browser's buffer while tmux still holds more. The downgrade guard reads that as "tmux has nothing more to give", which is true of an unbounded capture only, so a bounded window reaching it marked the session exhausted and removed Load full history from the banner. The bounded skip now runs first, so such a window never reaches the exhausted path, and it no longer writes banner state: relabelling it from the bounded payload would call a terminal holding all of a Load full history pull "the most recent 1 MiB". A skipped window that came back truncated cannot reach anything older than the browser shows, and every ask costs the server a synchronous capture-pane of the whole history (tail is applied after the capture), so it puts the session on the 60 s cooldown. An untruncated one keeps 4 s. _replayWouldShrinkBuffer takes optional pre-estimated rows so a megabyte capture is not scanned twice. CLAUDE.md's Full-scrollback replay entry no longer says Shell never pulls on ordinary scroll. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- src/web/public/app.js | 42 ++++--- src/web/public/terminal-ui.js | 8 +- test/shell-scroll-history-pull.test.ts | 167 ++++++++++++++++++++++++- test/terminal-scroll-routing.test.ts | 10 +- 6 files changed, 205 insertions(+), 26 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index c4a910cc..dac96cc7 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -265,7 +265,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and Shell loads the rest only via **Load full history**, never on ordinary scroll. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 742419f6..e23dc232 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -205,7 +205,7 @@ Further detail: the `<prefix>: <title>` form (`w3-myapp: fix the login redirect` ### Full-scrollback replay -**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. +**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection and automatic drop recovery: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell FULL history is explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. That gesture instead pulls a BOUNDED window of it (`?full=1&tail=<TERMINAL_TAIL_SIZE>`, so a capture beyond 1 MiB comes back `truncationReason: 'tail'` and the banner offers the rest), and skips the replay when the window holds no more rows than the browser already has. ⚠️ **That skip runs BEFORE the downgrade guard, and it never touches the banner state**: the guard reads "smaller than the browser" as "tmux has nothing more to give", true of an unbounded capture and false of a window cut at the tail size, so a bounded window that reached it marked the session `exhausted` and removed **Load full history** while tmux still held the rest; and re-labelling the banner from a skipped payload would call a terminal that holds all of a Load full history pull "the most recent 1 MiB". A skipped window that came back truncated puts the session on the 60 s cooldown (`_fullHistoryRepullUseless`), because it can never reach anything older than the browser shows and `tail` is applied AFTER the server's synchronous `capture-pane` of the whole history, so each ask costs every client an event-loop stall (about 0.3 s at 100k lines); an untruncated one keeps the 4 s cooldown, since it is all of tmux's history and the next burst can add to it. Declining the gesture outright was a dead end: after any burst a shell pane held about one screen of browser scrollback, and the button renders only once a replay was truncated, so a young shell tab had no way back to history tmux still held. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, and the marker callback supplies accurate parse timing. ⚠️ **How the load ENDS depends on where the payload came from**, and `_bufferLoadFinishOpts` (app.js) is the one place that decides it for all four fetch-and-write paths. A payload built from the server's accumulated byte history is current up to the response, so the events queued during the load already appear in it and stay DISCARDED; replaying them would duplicate output, most visibly Ink's cursor-up redraws. A pane capture (`mux-visible` or `mux-full-history`) is current only up to CAPTURE time, so `_finishBufferLoad` replays the queue from the response's own arrival timestamp (`since`) and the pre-capture events stay dropped. ⚠️ **A path that then restores a scroll position must re-take the sticky-scroll baseline** (`_syncStickyScrollBaseline`): the replay runs inside `chunkedTerminalWrite` before its promise resolves, with the terminal freshly reset, so `batchTerminalWrite` samples `_wasAtBottomBeforeWrite` as true and the next `flushPendingWrites` would scroll to the bottom over the restore. ⚠️ The cutoff is a client-side timestamp and the server broadcasts on a batch timer (8ms WebSocket, 16-50ms SSE), so a batch pending when the capture ran arrives after the response and replays although the capture holds it — bounded by one batch interval, and closable only server side by flushing that batch before the capture. Tests for the three: `test/terminal-flush-budget.test.ts` pins which sources flush, `test/terminal-buffer-flush.test.ts` pins the `since` cutoff and the baseline re-take, and `test/capture-load-window.browser.test.ts` drives both against a live server. Live output is separately one-chunk-in-flight: xterm's callback releases each 32/64 KiB write before the next is submitted, keeping the remainder in the app queue where the 128 KiB cap can observe it instead of hiding an unbounded backlog in xterm's private WriteBuffer. While WebSocket owns terminal I/O, parallel SSE terminal/output-recovery events are discarded before JSON parsing; fallback recovery is single-flight per active session so backpressure cannot start overlapping reset+replay cycles. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The capture ENDS with a cursor move back to the pane's own caret position** (`formatCursorRestore`, from the same `display-message` query the visible-frame path uses). The linear replay otherwise leaves the caret wherever the last character landed — the bottom-most row carrying text, which for an agent CLI is the status line — so the caret sat on the composer's border instead of its input line and every cursor-relative update the CLI sent afterwards was measured from the wrong row, until its next full redraw silently repaired it (that self-repair is why the report read as "it fixes itself as soon as Claude writes a line"). ⚠️ **The move is RELATIVE — up `rows - 1 - cursor_y`, then `\r`, then right `cursor_x` — never `CUP`.** `\x1b[<row>;<col>H` numbers rows from the top of the browser's screen, so it lands correctly only while the browser's row count equals `pane_height`, and nothing guarantees that: `resizeWindow` issues its tmux resize fire-and-forget and returns immediately, so a capture can be taken before a requested resize has applied, and `_onSessionNeedsRefresh` sends no resize at all. Counting up from the last replayed row anchors to the content both ends share. Restoring the cursor makes ROW ALIGNMENT load-bearing on this path: **no transform that can DELETE A LINE may run over a full-history capture**, because every deletion shifts the frame out from under the restored position. Four had accumulated — trailing blank rows stripped by `\n+$`, `stripInkRedrawBloat`, the `CLAUDE_BANNER_PATTERN` trim that cuts everything above the banner, and `LEADING_WHITESPACE_PATTERN` — each correct for a byte stream of successive frames and each wrong for a single rendered frame. ⚠️ **Those skips key on `isFullCapture`, meaning a capture actually came back — never on `?full=1` alone.** When `captureActivePaneBuffer` returns null (ENOBUFS, a timeout, a vanished pane, or a session with no mux at all) the reply falls back to `session.terminalBuffer`, which IS a byte stream and must still be stripped; gating on the query flag returned it whole, and a direct-PTY session takes that path on every first selection rather than only during an outage. ⚠️ A capture holding nothing visible (`hasVisibleContent`) returns `''`, because the caller reads an empty capture as "unavailable" and keeps its byte history — retaining trailing blank rows made an all-blank pane non-empty, which would have replaced real history with a blank screen from the server side, where `_replayWouldShrinkBuffer` cannot see it. ⚠️ **"One line per screen row" holds only where no row was hard-wrapped**: `-J` joins a wrapped row into its logical line (measured: a 100-character line in a 40-column pane captures as 10 lines against a 12-row pane), and the counts reconcile only once the browser xterm re-wraps at the same width — the same assumption `_estimateReplayRows` already documents. Tests: `test/tmux-capture-full-history.test.ts` covers the cursor move, the trim pairing and `hasVisibleContent`; `test/routes/session-routes.test.ts` covers a surviving blank first row, an unstripped byte-history fallback, and an empty capture leaving history intact. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`, `test/terminal-flush-budget.test.ts`. **A capture reports the geometry it was taken at** (#435): a visible frame repaints each row at an absolute position, counting up to the pane's height and out to the pane's width, so a terminal smaller than that pane damages it two ways at once. Too short and every address past the browser's own height clamps onto the last line, overwriting the rows underneath (measured: against a 50-row pane, a 30-row terminal rendered 28 of a 45-line command and drew the survivors twice). Too narrow and each row is painted out to the pane's width, so the browser wraps every painted row and the wrap on the last one scrolls the whole frame up by one. Nothing in the response used to say what geometry the frame was built for, so the client could not see either case. `PaneCaptureOptions.capturedGeometry` carries it out, and the terminal response publishes it as `captureCols`/`captureRows`. ⚠️ **Both fields are ABSENT unless a frame was really positioned**, and every consumer must test `Number.isFinite` rather than truthiness: `mux-visible` is necessary but not sufficient, because when the `display-message` cursor query fails `capturePaneBuffer` skips the snapshot repaint and returns the raw capture, and the route still labels that non-empty body `mux-visible`. A body that positioned nothing has no geometry to describe and nothing to repair, so a comparison that fires there buys a second capture, a reset plus chunked rewrite, a dropped and reopened WebSocket and a discarded xterm snapshot for no gain. ⚠️ **The comparison runs on `mux-visible` ONLY.** A full-history body is linear scrollback closed by a RELATIVE cursor move, which is relative precisely so the browser's row count need not match the pane's, and a byte-history body carries no row alignment at all, so a size mismatch damages neither and a replay repairs neither. That gate matters because the first select of every non-shell session per page takes the full-history path, where an ungated comparison would fire most often on the one response it cannot help, at the price of a second whole-scrollback capture. ⚠️ **The replay is capped at one attempt and latches per session when it cannot converge.** `resizeRetry` stops two competing fits trading replays forever; a pane already drawing at the size just requested is left alone, because a retry would capture the identical frame; and a pass that still does not converge joins `_geometryRetryUseless`, so the case `Session.resize` declines outright (a small viewport while a desktop viewport's size claim is live, where the retry re-sends the same declined resize and captures the same pane) costs one attempt per session per page load instead of one per select. ⚠️ **The clamp used to manufacture that equality, and no longer can** (#464). This paragraph previously explained it as the signature of a clamp: `getTerminalDimensions()` floors at 40x10 while `fitAddon.fit()` did not, so a terminal under 40 columns or 10 rows reported a pane permanently bigger than itself and would replay on every tab switch. That divergence is fixed at the source — `syncTerminalGeometry()` (terminal-ui.js) fits, floors and APPLIES in one step, so the browser terminal IS the size it reports. The only remaining reason the two can differ is a resize the server declined, which the server now reports back (`Session.ptyGeometry`, the `{"t":"zc"}` frame and the resize response) for the client to adopt by COLUMNS. The equality guard stays, for the plain case of a pane already at the requested size. ⚠️ **A retry pass must not re-arm `_fullHistoryLoaded`**: it did not consume the full-history pull, and re-arming it would spend a whole-scrollback capture on the next select. That branch is currently unreachable by construction, since reaching it needs `source === 'mux-visible'` while a `full=1` pass is answered `mux-full-history` or `history`; a static test over the source is the habit this repo uses for an invariant nothing can execute. Tests: `test/capture-geometry-retry.browser.test.ts` (eight cases, five of which fail against the merge base), `test/tmux-capture-full-history.test.ts`, `test/routes/session-routes.test.ts`. diff --git a/src/web/public/app.js b/src/web/public/app.js index 1da813be..a8bda99f 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -6367,7 +6367,33 @@ class CodemanApp { // Bail on a tab switch mid-fetch: writing here would paint another session's // history into the terminal the user is now looking at. if (!buffer || this.activeSessionId !== sessionId) return; - if (this._replayWouldShrinkBuffer(buffer)) { + const windowRows = this._estimateReplayRows(buffer, this.terminal.cols); + // A bounded window no longer than the browser's buffer buys nothing, and + // resetting to rewrite it would jump the viewport on every scroll that + // outlasts the cooldown at the top. This runs BEFORE the downgrade guard + // on purpose: that guard reads "smaller than the browser" as "tmux has + // nothing more to give", which is true of an unbounded capture but not of a + // window cut at the tail size, so a bounded window must never reach the + // exhausted path, which would take Load full history off the banner while + // tmux still holds the rest. Nothing was written here, so the banner state + // is left as the load that produced it set it: re-labelling it from this + // payload would call a terminal that holds ALL of a Load full history pull + // "the most recent 1 MiB". + if (boundedShellPull && windowRows <= this.terminal.buffer.active.length) { + // An untruncated window IS all of tmux's history, so nothing is missing, + // and the next burst of output can put more in tmux than the browser has: + // keep the normal 4 s cooldown. A truncated one is the opposite case, since + // the gesture can never reach anything older than what the browser already + // shows, and every ask costs the server a synchronous capture-pane of the + // whole history (`tail` is applied after the capture): back off to 60 s. + // Trade-off: only a successful replay clears that latch, so a tab switch or + // burst that shrinks the browser's buffer below the window can leave a + // scroll-to-top inert for up to a minute. Load full history (`force`) + // bypasses the cooldown, and the latch is bounded, never permanent. + if (payload.truncated) (this._fullHistoryRepullUseless ||= new Set()).add(sessionId); + return; + } + if (this._replayWouldShrinkBuffer(buffer, windowRows)) { timing.refused = true; timing.totalMs = performance.now() - requestStartedAt; this._recordTerminalLoadTiming(timing); @@ -6378,20 +6404,6 @@ class CodemanApp { this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); return; } - // A bounded window no longer than the browser's buffer buys nothing, and - // resetting to rewrite it would jump the viewport on every scroll that - // outlasts the cooldown at the top. An untruncated window IS all of tmux's - // history, so nothing is missing; a window cut at the tail size would trade - // more old rows than it recovers, and its 'tail' truncation keeps the - // banner offering the unbounded pull. Not latched as useless: the next - // burst of output can put more history in tmux than the browser has. - if ( - boundedShellPull && - this._estimateReplayRows(buffer, this.terminal.cols) <= this.terminal.buffer.active.length - ) { - this._setHistoryTruncation(sessionId, payload); - return; - } this._setHistoryTruncation(sessionId, payload); this._fullHistoryRepullUseless?.delete(sessionId); const rowsBefore = this.terminal.buffer.active.length; diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index 413fe99b..e969e1e3 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -3354,13 +3354,17 @@ Object.assign(CodemanApp.prototype, { * below the last line, and _estimateReplayRows can only approximate wrapping. * Only a capture that is worse by more than a full screen counts as a * downgrade, which leaves every genuine recovery case untouched. + * + * A caller that already estimated the capture's rows passes them as + * `estimatedRows`, so a megabyte capture is not scanned twice. */ - _replayWouldShrinkBuffer(capture) { + _replayWouldShrinkBuffer(capture, estimatedRows) { const term = this.terminal; const rowsNow = term?.buffer?.active?.length || 0; if (!rowsNow) return false; const screen = term?.rows || 24; - return this._estimateReplayRows(capture, term?.cols) + screen < rowsNow; + const rows = estimatedRows ?? this._estimateReplayRows(capture, term?.cols); + return rows + screen < rowsNow; }, /** diff --git a/test/shell-scroll-history-pull.test.ts b/test/shell-scroll-history-pull.test.ts index 229d8a0c..ac9de9b0 100644 --- a/test/shell-scroll-history-pull.test.ts +++ b/test/shell-scroll-history-pull.test.ts @@ -14,6 +14,14 @@ * the button stays the unbounded path, and a window the browser already holds * in full is not rewritten. * + * ORDER MATTERS: that skip must run BEFORE the downgrade guard. The guard reads + * "smaller than the browser" as "tmux has nothing more to give", which is true of + * an unbounded capture and false of a window cut at the tail size, so a bounded + * window that reached it marked the session exhausted and took Load full history + * off the banner while tmux still held the rest. The second block below drives + * the real `_setHistoryTruncation` and the real `computeHistoryTruncationNotice` + * to pin what the user is actually told. + * * The method is extracted from app.js and run in a `vm` against stubs (no jsdom * on this box; see connection-indicator.test.ts), with the REAL row estimators * from terminal-ui.js, which decide both the downgrade and the no-gain skip. @@ -61,11 +69,43 @@ function loadRefetch(): (this: unknown, opts?: { force?: boolean }) => Promise<v return vm.runInContext(`({ ${body} })._maybeRefetchFullHistory`, context); } +/** The REAL `_setHistoryTruncation`, so the banner state a pull leaves behind is what production would hold. */ +function loadSetHistoryTruncation(): (this: unknown, sessionId: string, payload?: Record<string, unknown>) => void { + const app = readFileSync(resolve(PUBLIC, 'app.js'), 'utf8'); + const body = methodSource(app, '_setHistoryTruncation'); + return vm.runInContext(`({ ${body} })._setHistoryTruncation`, vm.createContext({})); +} + +/** The REAL banner decision from constants.js: what the user is told, and whether Load full history is offered. */ +function loadNotice() { + const context = vm.createContext({ console, window: {}, document: {}, navigator: { userAgent: 'test' } }); + vm.runInContext( + `${readFileSync(resolve(PUBLIC, 'constants.js'), 'utf8')}\n;globalThis.__notice = computeHistoryTruncationNotice;`, + context, + { filename: 'constants.js' } + ); + return (context as { __notice: (s: Record<string, unknown>) => { visible: boolean; canLoadMore: boolean } }).__notice; +} + const mixin = loadTerminalMixin(); const refetch = loadRefetch(); +const setHistoryTruncation = loadSetHistoryTruncation(); +const computeNotice = loadNotice(); const lines = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\r\n'); -function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; capture: string }) { +/** A `?full=1&tail=` answer whose window was CUT at the tail size: tmux holds ~3 MiB, the window carries 1 MiB. */ +const TAIL_CUT = { + truncated: true, + truncationReason: 'tail', + fullSize: 3 * 1024 * 1024, + retainedBytes: TERMINAL_TAIL_SIZE, + source: 'mux-full-history', +}; + +function makeApp( + mode: string, + { bufferRows, capture, payload = {} }: { bufferRows: number; capture: string; payload?: Record<string, unknown> } +) { const urls: string[] = []; const app = { activeSessionId: 's1', @@ -90,12 +130,18 @@ function makeApp(mode: string, { bufferRows, capture }: { bufferRows: number; ca return { headersAt: performance.now(), headers: { get: () => '' }, - json: { data: { terminalBuffer: capture, source: 'mux-full-history' } }, + json: { data: { terminalBuffer: capture, source: 'mux-full-history', ...payload } }, }; }), _recordTerminalLoadTiming: vi.fn(), _logScrollRouting: vi.fn(), - _setHistoryTruncation: vi.fn(), + // The real method behind a spy, so a test sees both what it was called with + // and the banner state (`_historyTruncation`) it leaves behind. + _historyTruncation: new Map<string, unknown>(), + _renderHistoryTruncationBanner: vi.fn(), + _setHistoryTruncation: vi.fn((sessionId: string, p?: Record<string, unknown>): void => { + setHistoryTruncation.call(app, sessionId, p); + }), _resetTerminalForReplay: vi.fn(), _bufferLoadFinishOpts: vi.fn(() => ({})), chunkedTerminalWrite: vi.fn(async () => ({ parsedAt: performance.now(), bufferLength: 400, completed: true })), @@ -135,8 +181,117 @@ describe('shell scroll-up pulls a bounded window of tmux history', () => { expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); // Not latched as useless: more output can put more history in tmux. expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); - // The truncation state is still recorded, so a window capped at the tail - // size keeps offering the button. - expect(app._setHistoryTruncation).toHaveBeenCalledTimes(1); + // Nothing was written, so the banner state is left exactly as it was. + expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + }); +}); + +describe('a skipped bounded window never damages the Load full history banner', () => { + it('a tail-cut window smaller than the browser is not replayed and never marks the session exhausted', async () => { + // The browser holds far more rows than a 1 MiB window carries, and tmux holds + // ~3 MiB. The downgrade guard reads that as "tmux has nothing more to give", + // which is true of an unbounded capture and false of a window cut at the tail. + const { app } = makeApp('shell', { bufferRows: 5000, capture: lines(300), payload: TAIL_CUT }); + // The tab load that put this session on screen left it truncated and recoverable. + app._setHistoryTruncation('s1', TAIL_CUT); + app._setHistoryTruncation.mockClear(); + + await refetch.call(app); + + expect(app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(app.chunkedTerminalWrite).not.toHaveBeenCalled(); + // Not even a relabel: a skipped window writes nothing, banner state included. + expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + const notice = computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>); + expect(notice.visible).toBe(true); + // "Earlier output is no longer kept" would be a lie: tmux still holds ~2 MiB more. + expect(notice.canLoadMore).toBe(true); + }); + + it('a skip right after Load full history leaves the banner as that load set it', async () => { + // Load full history replayed everything, so nothing is truncated any more. + const afterLoadFullHistory = { + truncated: false, + fullSize: 3 * 1024 * 1024, + retainedBytes: 3 * 1024 * 1024, + source: 'mux-full-history', + }; + const { app } = makeApp('shell', { bufferRows: 5000, capture: lines(300), payload: TAIL_CUT }); + app._setHistoryTruncation('s1', afterLoadFullHistory); + const before = structuredClone(app._historyTruncation.get('s1')); + app._setHistoryTruncation.mockClear(); + + await refetch.call(app); + + // Relabelling it from the bounded payload would call a terminal that holds ALL + // of the history "the most recent 1.0 MB". + expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + expect(app._historyTruncation.get('s1')).toEqual(before); + expect(computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>).visible).toBe(false); + }); + + it('backs off for a minute after a truncated skip, and keeps the 4 s cooldown after an untruncated one', async () => { + // Truncated: the gesture cannot reach anything older than the browser shows, and + // every ask costs the server a synchronous capture of the whole history. + // Only just larger than the window, so the downgrade guard does not fire here: + // the back-off has to come from the skip itself. + const cut = makeApp('shell', { bufferRows: 320, capture: lines(300), payload: TAIL_CUT }); + await refetch.call(cut.app); + expect(cut.app._fullHistoryRepullUseless.has('s1')).toBe(true); + expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(1); + + // Well past 4 s, still inside the minute: no second capture. + cut.app._fullHistoryRepullAt.set('s1', Date.now() - 10_000); + await refetch.call(cut.app); + expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(1); + + cut.app._fullHistoryRepullAt.set('s1', Date.now() - 61_000); + await refetch.call(cut.app); + expect(cut.app._fetchTerminalCapture).toHaveBeenCalledTimes(2); + + // Untruncated: it IS all of tmux's history, and the next burst can add to it. + const whole = makeApp('shell', { bufferRows: 320, capture: lines(300) }); + await refetch.call(whole.app); + expect(whole.app._fullHistoryRepullUseless.has('s1')).toBe(false); + whole.app._fullHistoryRepullAt.set('s1', Date.now() - 5000); + await refetch.call(whole.app); + expect(whole.app._fetchTerminalCapture).toHaveBeenCalledTimes(2); + }); + + it('the downgrade guard still refuses an unbounded capture smaller than the browser, and still marks it exhausted', async () => { + // Reordering must not weaken the guard it moved above: a repaint-mode pane's + // capture really is one frame, and rewriting with it would destroy history. + const oneFrame = makeApp('claude', { bufferRows: 300, capture: lines(36) }); + await refetch.call(oneFrame.app); + expect(oneFrame.app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(oneFrame.app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ exhausted: true })); + expect(oneFrame.app._fullHistoryRepullUseless.has('s1')).toBe(true); + + // The button is unbounded too, so the same guard governs it for a shell. + const button = makeApp('shell', { bufferRows: 5000, capture: lines(300) }); + await refetch.call(button.app, { force: true }); + expect(button.app._resetTerminalForReplay).not.toHaveBeenCalled(); + expect(button.app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ exhausted: true })); + }); +}); + +describe('_replayWouldShrinkBuffer takes rows the caller already estimated', () => { + const shrink = mixin._replayWouldShrinkBuffer as (this: unknown, capture: string, rows?: number) => boolean; + const make = () => ({ + terminal: { cols: 80, rows: 30, buffer: { active: { length: 200 } } }, + _estimateReplayRows: vi.fn(mixin._estimateReplayRows as (t: string, c: number) => number), + }); + + it('does not scan the capture again when handed the estimate', () => { + const ctx = make(); + expect(shrink.call(ctx, lines(300), 300)).toBe(false); + expect(shrink.call(ctx, lines(300), 5)).toBe(true); + expect(ctx._estimateReplayRows).not.toHaveBeenCalled(); + }); + + it('still estimates for itself when called the old way', () => { + const ctx = make(); + expect(shrink.call(ctx, lines(300))).toBe(false); + expect(ctx._estimateReplayRows).toHaveBeenCalledTimes(1); }); }); diff --git a/test/terminal-scroll-routing.test.ts b/test/terminal-scroll-routing.test.ts index b3cb888b..d2dd206a 100644 --- a/test/terminal-scroll-routing.test.ts +++ b/test/terminal-scroll-routing.test.ts @@ -111,12 +111,20 @@ describe('full-history re-pull downgrade guard (issue #205 round 2)', () => { // Anchor on the open paren, not the full empty signature: the method takes // options since #258 ({ force }) and this guard is about ORDER, not arity. const start = source.indexOf('async _maybeRefetchFullHistory('); - const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer)', start); + // Also anchored on the open paren: the guard is handed the rows the caller + // already estimated, and this test is about ORDER, not the argument list. + const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer', start); + const boundedSkip = source.indexOf('boundedShellPull && windowRows <=', start); const reset = source.indexOf('this._resetTerminalForReplay()', start); expect(start).toBeGreaterThan(-1); expect(guard).toBeGreaterThan(start); expect(guard).toBeLessThan(reset); // refuse first, only then reset+rewrite + // A bounded shell window is skipped BEFORE the guard sees it: the guard reads + // "smaller than the browser" as "tmux has nothing more", which a window cut at + // the tail size does not mean (see shell-scroll-history-pull.test.ts). + expect(boundedSkip).toBeGreaterThan(start); + expect(boundedSkip).toBeLessThan(guard); // A hollow pane must also stop re-fetching megabytes on every scroll-up. expect(source).toContain('this._fullHistoryRepullUseless'); expect(source).toContain('this._fullHistoryRepullUseless?.has(sessionId) ? 60000 : 4000'); From 1da2fa2529b305c23d1d41d4db2fffe85dbd8c5b Mon Sep 17 00:00:00 2001 From: JD <jd@jds.haus> Date: Sat, 26 Sep 2026 15:33:44 -0400 Subject: [PATCH 15/46] fix(terminal): only forward scroll to Claude while it tracks the mouse Claude 2.1.280 renders inline by default: no alt screen, no mouse tracking, transcript in real scrollback. The version-only gate still sent every wheel tick and touch swipe as SGR reports, which Claude ignores, so scrolling a Claude session was dead while codex (routed locally) worked. Gate forwarding on the server-recorded cliMouseTracking flag, which fullscreen mode (CLAUDE_CODE_NO_FLICKER=1) sets. --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- docs/wiki/The-Dashboard.md | 5 +++-- docs/wiki/Troubleshooting.md | 5 +++-- src/web/public/terminal-ui.js | 8 ++++++++ test/terminal-touch-tap.test.ts | 25 ++++++++++++++++++++++--- 6 files changed, 38 insertions(+), 9 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index c4a910cc..e681e200 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -275,7 +275,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Ctrl+V paste trap** (`image-input.js`): `Ctrl+V` routes through `_handleImagePaste()`, which focuses a hidden `contenteditable` trap and reads the clipboard from the paste event landing there; images upload and their paths are typed in, text goes through `terminal.paste()` so bracketed-paste markers survive. ⚠️ **The trap must consume exactly ONE paste event** (Firefox delivers two per keypress: the `execCommand('paste')` event and the keydown's default action); the one-shot flag lives on the trap, never on a browser check. ⚠️ Do not remove the `execCommand('paste')` call: on some mobile engines it is the only route into the trap, and the trap is the only place image blobs are read. Tests: `test/image-paste-trap.test.ts`. → [architecture-invariants#terminal-paste-ctrlv](docs/architecture-invariants.md#terminal-paste-ctrlv) -**Terminal scrollback strip + wheel/touch forwarding**: codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only). ⚠️ Gated on `useMux`: direct-PTY sessions must keep the alt screen. Wheel and touch forward to the CLI for **claude ≥ 2.1.187 ONLY**; ⚠️ never re-add codex without a fresh measurement (it ignores SGR wheel reports). ⚠️ `getClaudeCliVersion()` must never cache a FAILED probe. ⚠️ Hand-report clicks only while the CLI has mouse tracking on: `_shouldReportMouseToCli()` gates all three report sites on `cliMouseTracking` (from `_recordStrippedMouseMode()`, session.ts), or a plain shell prints the reports as literal text. Read `_logScrollRouting()` before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding) +**Terminal scrollback strip + wheel/touch forwarding**: codex/claude/gemini get the FULL strip (alt-screen, `3J`, mouse DECSETs); tmux-backed shell/opencode/antigravity/omp get a NARROW strip (alt-screen toggles only). ⚠️ Gated on `useMux`: direct-PTY sessions must keep the alt screen. Wheel and touch forward to the CLI for **claude ≥ 2.1.187 ONLY, and only while it has mouse tracking on** (`cliMouseTracking`: fullscreen claude sets it, its default inline renderer does not and scrolls locally like codex); ⚠️ never re-add codex without a fresh measurement (it ignores SGR wheel reports). ⚠️ `getClaudeCliVersion()` must never cache a FAILED probe. ⚠️ Hand-report clicks only while the CLI has mouse tracking on: `_shouldReportMouseToCli()` gates all three report sites on `cliMouseTracking` (from `_recordStrippedMouseMode()`, session.ts), or a plain shell prints the reports as literal text. Read `_logScrollRouting()` before diagnosing a scroll report. → [architecture-invariants#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding](docs/architecture-invariants.md#terminal-scrollback-strip-flavors-and-wheeltouch-forwarding) **Detached start + service install**: `codeman web -d` relaunches the same entry script `detached:true` (setsid); `nohup` is not what makes it survive. ⚠️ Both `-d` and `service install` must REFUSE when a server is already up on this data dir (pidfile + `/api/status` probe), or a second instance attaches to the first one's live sessions. ⚠️ Never report success not observed: poll `/api/status` until the child answers or dies. `--stop` must verify the pid still looks like Codeman (`ps -o command=`) before signalling. Unit/label names live only in `config/service-names.ts`. `service install` bakes the installing shell's PATH into the unit and never writes `CODEMAN_PASSWORD` into it. → [architecture-invariants#detached-start-and-service-install](docs/architecture-invariants.md#detached-start-and-service-install) **Self-update** (App Settings → System → Updates): in-app updater for git-clone installs under a supervisor (`systemd`, `launchd`, `launchd-daemon`, `docker-compose`, else `none`). The work runs in a DETACHED `scripts/self-update.sh` writing `update-status.json`, polled across the restart; pure helpers in `src/web/self-update.ts`. ⚠️ Compose: the restart kills the script, so nothing may be appended after the `restarting` marker; the repo must stay a host bind mount over `/opt/codeman` and the image must keep devDependencies + toolchain. ⚠️ `evaluateEnvironmentGate()` refuses releases that change `server.Dockerfile`/`docker-compose.yaml` or add `.env.example` keys, re-evaluated on `POST /api/system/update`; unknowns fail OPEN, but the exit-to-restart needs `--restart-by-exit 1` (`CODEMAN_RESTART_BY_EXIT=1` only in the Compose file). ⚠️ Keep the agent CLIs in `server.Dockerfile` pinned. → [docs/docker-self-update.md](docs/docker-self-update.md), [architecture-invariants#self-update](docs/architecture-invariants.md#self-update) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 6a25c5b3..96908039 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -217,7 +217,7 @@ Further detail: the `<prefix>: <title>` form (`w3-myapp: fix the login redirect` ⚠️ **What the full strip removes, it must REMEMBER.** Stripping the mouse DECSETs means xterm's `modes.mouseTrackingMode` is permanently `'none'` for those modes, so the browser hand-encodes click reports to compensate (`_sendSyntheticSgrTap`). With no state to consult it did that on EVERY click, which delivered mouse reports to programs that never asked for them: the same pane runs a plain shell whenever the CLI has exited or a `shell` was started inside a claude-mode session, and a shell prints the report as literal text (`[<0;88;20M`), garbling the next line typed. `_recordStrippedMouseMode()` therefore records each stripped sequence as it goes and publishes `cliMouseTracking` through `toState()`, and `_shouldReportMouseToCli()` requires it. ⚠️ Only the TRACKING modes count (1000/1001/1002/1003): 1005/1006 select an ENCODING and 1007 is alt-scroll, and counting those would put the stray reports straight back. ⚠️ The change broadcasts IMMEDIATELY rather than through `broadcastSessionStateDebounced`, because the flag flips when a dialog opens and the user can click that dialog inside the 500ms debounce window. Measured on a live claude 2.x: the CLI holds a tracking mode on continuously (so clicks keep being reported exactly as before), while a bash prompt in the same stripped mode reports nothing. Fails toward silence: after a server restart the flag is false until the CLI re-emits, which tmux does at client attach. -**Only claude ≥ 2.1.187 forwards the wheel; every other mode scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS are gated by `_shouldReportMouseToCli()` (strip mode AND the server-observed `cliMouseTracking` flag, recorded by `_recordStrippedMouseMode` in session.ts as it strips): codex never enables mouse tracking, so since #325 no tap report is sent there at all — click-to-position was already a measured no-op in codex, and a pane that has fallen back to a shell no longer receives `[<0;88;20M` junk. +**Only claude ≥ 2.1.187 with mouse tracking on forwards the wheel; everything else scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. Claude repeats this exactly in its default INLINE renderer (measured on 2.1.280: `alternate_on=0`, `mouse_any_flag=0`, `history_size` grows), and swipes on iOS Safari were dead there while codex scrolled; only fullscreen claude (`CLAUDE_CODE_NO_FLICKER=1`: alt screen plus modes 1003/1006) pages its transcript on wheel reports. So claude forwards only while the server-observed `cliMouseTracking` flag is true; a stale-false flag after a server restart falls through to the PageUp/PageDown fallback, never a dead wheel. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS are gated by `_shouldReportMouseToCli()` (strip mode AND the server-observed `cliMouseTracking` flag, recorded by `_recordStrippedMouseMode` in session.ts as it strips): codex never enables mouse tracking, so since #325 no tap report is sent there at all — click-to-position was already a measured no-op in codex, and a pane that has fallen back to a shell no longer receives `[<0;88;20M` junk. **Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`. diff --git a/docs/wiki/The-Dashboard.md b/docs/wiki/The-Dashboard.md index cf1f415e..eea5cca6 100644 --- a/docs/wiki/The-Dashboard.md +++ b/docs/wiki/The-Dashboard.md @@ -153,8 +153,9 @@ Worth knowing: Shell sessions open from a bounded recent tail so a large transcript cannot stall tab switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling and automatic output recovery stay within the bounded browser buffer. -- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude - versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is +- **Wheel and touch scrolling** are forwarded into Claude's own transcript when a recent + Claude runs fullscreen, so the wheel scrolls the conversation rather than the terminal. + Claude's default inline view keeps its history in the terminal and scrolls locally. `Shift+Wheel` is always local scrollback. Other CLIs scroll locally. - **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not. `Ctrl+Shift+C` always copies. diff --git a/docs/wiki/Troubleshooting.md b/docs/wiki/Troubleshooting.md index 7eb279b3..39b6754c 100644 --- a/docs/wiki/Troubleshooting.md +++ b/docs/wiki/Troubleshooting.md @@ -164,8 +164,9 @@ Scrollback behaviour depends on the CLI, and Codeman adjusts what it strips per Things to try: - `Shift+Wheel` always scrolls the local buffer, whatever else is going on. -- On Claude sessions with a recent CLI, the wheel is forwarded into Claude's own transcript, - so it scrolls the conversation rather than the terminal buffer. That is intended. +- On Claude sessions running fullscreen (recent CLI with mouse tracking on), the wheel is + forwarded into Claude's own transcript, so it scrolls the conversation rather than the + terminal buffer. That is intended. Claude's default inline view scrolls locally. - Scrolling to the very top pulls the full tmux scrollback again on demand. ### The wheel does nothing in a Codex session diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index 903661e5..6a8ea3d9 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -5151,6 +5151,14 @@ Object.assign(CodemanApp.prototype, { const sessionMode = session?.mode || 'claude'; if (sessionMode !== 'claude') return false; if (!this._cliVersionAtLeast(session?.cliVersion, '2.1.187')) return false; + // Only while Claude is actually listening for the mouse. In its default + // inline renderer (2.1.280 measured: alternate_on=0, mouse_any_flag=0) the + // transcript lives in real scrollback, like codex, and SGR wheel reports are + // ignored, so forwarding made every swipe and wheel tick dead. Fullscreen + // (CLAUDE_CODE_NO_FLICKER=1) turns on alt-screen + mode 1003/1006, which the + // server records as cliMouseTracking. A stale-false flag after a server + // restart falls through to _maybePageCliTranscript, so it never goes dead. + if (session?.cliMouseTracking !== true) return false; // Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so // that leaving the bottom handed the wheel back to local scrollback and both // histories stayed reachable without a mode switch. In practice that inverted diff --git a/test/terminal-touch-tap.test.ts b/test/terminal-touch-tap.test.ts index 13bba1dc..eacd815d 100644 --- a/test/terminal-touch-tap.test.ts +++ b/test/terminal-touch-tap.test.ts @@ -585,7 +585,7 @@ describe('terminal touch tap mouse guard', () => { it('wheel: forwards to the app for verified sessions without Shift, at ANY scroll position', () => { const { app } = loadTerminalUiHarness(); app.activeSessionId = 'sess-1'; - app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187' }]]); + app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187', cliMouseTracking: true }]]); app.terminal = { modes: { mouseTrackingMode: 'none' }, buffer: { active: { viewportY: 50, baseY: 50 } }, @@ -638,7 +638,7 @@ describe('terminal touch tap mouse guard', () => { buffer: { active: { viewportY: 50, baseY: 50 } }, }; const withVersion = (cliVersion?: string) => { - app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion }]]); + app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion, cliMouseTracking: true }]]); return app._shouldForwardWheelToApp({ shiftKey: false }); }; @@ -650,6 +650,25 @@ describe('terminal touch tap mouse guard', () => { expect(withVersion('garbage')).toBe(false); // unparseable → assume older }); + it('wheel: inline claude (no mouse tracking) keeps the local wheel', () => { + // Claude 2.1.280's default inline renderer never enables mouse tracking and + // keeps its transcript in real scrollback, so SGR wheel reports are ignored. + // Forwarding there made every swipe dead on iOS Safari while codex scrolled. + const { app } = loadTerminalUiHarness(); + app.activeSessionId = 'sess-1'; + app.terminal = { + modes: { mouseTrackingMode: 'none' }, + buffer: { active: { viewportY: 50, baseY: 50 } }, + }; + app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.280' }]]); + expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); + app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.280', cliMouseTracking: false }]]); + expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); + // Fullscreen (CLAUDE_CODE_NO_FLICKER=1) turns tracking on → forwarding resumes. + app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.280', cliMouseTracking: true }]]); + expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(true); + }); + it('wheel: only claude forwards — codex and gemini keep the local wheel', () => { const { app } = loadTerminalUiHarness(); app.activeSessionId = 'sess-1'; @@ -674,7 +693,7 @@ describe('terminal touch tap mouse guard', () => { it('wheel: the local-scrollback opt-out pins the plain wheel to local scrollback (issue #154)', () => { const { app } = loadTerminalUiHarness(); app.activeSessionId = 'sess-1'; - app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187' }]]); + app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: '2.1.187', cliMouseTracking: true }]]); app.terminal = { modes: { mouseTrackingMode: 'none' }, buffer: { active: { viewportY: 50, baseY: 50 } }, From 1d85909a0667d6d1c0ac9102403deaeadc96c2e0 Mon Sep 17 00:00:00 2001 From: Aamer Akhter <aakhter@gmail.com> Date: Sat, 26 Sep 2026 18:45:17 -0400 Subject: [PATCH 16/46] build: add a browser-test exclusion check and a pre-push static-check hook npm run check:browser-excludes finds tests that import a browser driver and asks `vitest list` whether the CI config still collects them; wired into CI. npm install now also installs a marker-owned pre-push hook that runs the static CI checks (~15s). Skip with CODEMAN_SKIP_PREPUSH=1; hand-written hooks are left alone. --- .github/CONTRIBUTING.md | 7 +- .github/workflows/ci.yml | 7 + CLAUDE.md | 4 +- docs/wiki/Contributing.md | 5 + package.json | 1 + scripts/check-browser-test-excludes.mjs | 149 ++++++++++++ scripts/git-hooks.mjs | 172 ++++++++++++++ scripts/postinstall.js | 20 +- test/check-browser-test-excludes.test.ts | 97 ++++++++ test/git-hooks.test.ts | 283 +++++++++++++++++++++++ 10 files changed, 738 insertions(+), 7 deletions(-) create mode 100644 scripts/check-browser-test-excludes.mjs create mode 100644 scripts/git-hooks.mjs create mode 100644 test/check-browser-test-excludes.test.ts create mode 100644 test/git-hooks.test.ts diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md index 71221862..1468cff6 100644 --- a/.github/CONTRIBUTING.md +++ b/.github/CONTRIBUTING.md @@ -28,12 +28,15 @@ The frontend is plain JS served from `src/web/public/` with no bundler in dev: e CI runs all of these, so save yourself a round trip: ```bash -npm run typecheck # tsc --noEmit, strict mode +npm run typecheck # tsc --noEmit, strict mode npm run lint npm run format:check -npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules +npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules +npm run check:browser-excludes # every browser-driven test is kept out of `npm test` ``` +`npm install` also installs a `pre-push` git hook that runs these static checks (~15s) and blocks the push if one fails. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; it never replaces a `pre-push` hook of your own. + ### Tests ```bash diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 8eceb330..fe86f671 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -34,6 +34,13 @@ jobs: - name: Frontend JS syntax check run: npm run check:frontend-syntax + # Asks `vitest list` what CI would actually collect, rather than matching + # filenames: a browser-driven test missing from BROWSER_TEST_GLOBS + # (config/test-suites.ts) passes locally and dies in the test job with + # "browserType.launch: Executable doesn't exist". + - name: Browser-test exclusion check + run: npm run check:browser-excludes + - name: Format check run: npm run format:check diff --git a/CLAUDE.md b/CLAUDE.md index c4a910cc..a262f074 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -112,6 +112,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | Gesture playground | `npm run dev` **in** `packages/gesture-control/` (standalone vite demo, fake tabs) | | Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) | | Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) | +| Browser-test exclusion check | `npm run check:browser-excludes` (`scripts/check-browser-test-excludes.mjs`; runs in CI, <1s). Fails if a test importing playwright/puppeteer is still collected by `config/vitest.ci.config.ts`; add it to `BROWSER_TEST_GLOBS` in `config/test-suites.ts` | +| Pre-push hook | Installed by `npm install` (`scripts/git-hooks.mjs`, via postinstall): runs the static CI checks (~15s) before `git push`. Skip once: `CODEMAN_SKIP_PREPUSH=1 git push`. Marker-owned, so a hand-written `pre-push` is never overwritten; hooks dir resolved via `git rev-parse --git-path hooks` (worktree-safe) | | Excluded-suite runners | `npm run test:browser` · `npm run test:mobile` · `npm run test:perf` · `npm run test:all` (everything, environmental failures included) — see Testing | | Production start | `npm run start` | | Production logs | `journalctl --user -u codeman-web -f` | @@ -120,7 +122,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | Dependency doctor | `codeman doctor` (alias `check-deps`; `--json`, `--category core\|office\|other`). Probes Node/Claude CLI/tmux/LibreOffice/MS Office against `config/dependency-registry.ts`; engine is pure given an injectable `ProbeHost` | | Multi-user accounts | `codeman users add <name>` / `passwd <name>` / `list` / `rm <name>` (writes `~/.codeman/users.json`, mode 0600; see Multi-user mode) | -**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 9 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`). +**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `check:browser-excludes`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 9 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`). **Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`) — config lives in the **`"prettier"` key of `package.json`**, not a `.prettierrc` (keeps the repo root short; editors read it natively). `.prettierignore` stays at the root because Prettier resolves it relative to cwd. ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`. diff --git a/docs/wiki/Contributing.md b/docs/wiki/Contributing.md index 892e594f..87ce1a0a 100644 --- a/docs/wiki/Contributing.md +++ b/docs/wiki/Contributing.md @@ -42,10 +42,15 @@ npm run typecheck npm run lint npm run format:check npm run check:frontend-syntax +npm run check:browser-excludes npm test -- test/<file>.test.ts # one file, the normal way npm run test:ci # the full CI sweep ``` +`npm install` installs a `pre-push` git hook that runs the static checks above (~15s) and +blocks a push that would fail them. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; a +`pre-push` hook of your own is never overwritten. + **Never run bare `npm test`.** The default configuration includes browser-driven Playwright suites that need a live server, Chromium, and environment-specific baselines; they hang or fail on a normal machine. `test:ci` is the honest "run everything". diff --git a/package.json b/package.json index c51e42cc..0a57e311 100644 --- a/package.json +++ b/package.json @@ -28,6 +28,7 @@ "pretest:mobile": "node scripts/prepare-test-vendor.mjs", "test:mobile": "vitest run --config test/mobile/vitest.config.ts", "check:frontend-syntax": "node scripts/check-frontend-syntax.mjs", + "check:browser-excludes": "node scripts/check-browser-test-excludes.mjs", "fix:node-pty": "node scripts/fix-node-pty.mjs", "typecheck": "tsc --noEmit && tsc -p config/tsconfig.scripts.json", "lint": "eslint --config config/eslint.config.js 'src/**/*.ts'", diff --git a/scripts/check-browser-test-excludes.mjs b/scripts/check-browser-test-excludes.mjs new file mode 100644 index 00000000..209160ce --- /dev/null +++ b/scripts/check-browser-test-excludes.mjs @@ -0,0 +1,149 @@ +#!/usr/bin/env node +/** + * Browser-test exclusion check. + * + * `npm run test:ci` must never try to drive a real browser: CI runners (and any + * clean checkout) have no chromium, so such a file dies with + * `browserType.launch: Executable doesn't exist` and takes the whole suite with + * it. `config/vitest.ci.config.ts` therefore excludes every browser-driven test + * via `BROWSER_TEST_GLOBS` in `config/test-suites.ts`. That list is maintained + * BY HAND, and a new browser test simply does not appear in it unless someone + * remembers. The omission is invisible on a developer machine that has run + * `npx playwright install`, where the test passes, and only shows up on a clean + * runner. + * + * Two deliberate design choices: + * + * 1. **Detection is by CONTENT, not filename.** Matching `*.browser.test.ts` + * would miss the browser tests that predate that convention + * (`inline-rename`, `opencode-resize`, `webgl-fallback`, + * `terminal-copy-shortcut`, `codex-predictive-echo`). What actually makes a + * file dangerous is importing a browser driver, so that is what is tested. + * + * 2. **The exclusion side is answered by vitest itself**, via + * `vitest list --filesOnly`, rather than by re-implementing glob matching + * against the config's `exclude` array. Patterns there include `test/mobile/**` + * and `perf-*`; a hand-rolled matcher that disagreed with vitest by even one + * edge case would report a gap that does not exist, or miss one that does. + * Asking the real resolver cannot drift from the real behaviour. + * + * The pure pieces are exported for test/check-browser-test-excludes.test.ts; the + * check itself only runs when this file is executed directly. + */ +import { readdirSync, readFileSync } from 'node:fs'; +import { join, dirname, relative, sep, resolve } from 'node:path'; +import { fileURLToPath } from 'node:url'; +import { execFileSync } from 'node:child_process'; + +const ROOT = join(dirname(fileURLToPath(import.meta.url)), '..'); +const CI_CONFIG = join('config', 'vitest.ci.config.ts'); +const SUITES_FILE = join('config', 'test-suites.ts'); + +/** Importing any one of these means the test needs a real browser binary. */ +const BROWSER_DRIVER = + /\bfrom\s+['"](?:playwright|playwright-core|@playwright\/test|puppeteer|puppeteer-core)['"]|\b(?:require|import)\(\s*['"](?:playwright|playwright-core|@playwright\/test|puppeteer|puppeteer-core)['"]\s*\)/; + +/** @param {string} source */ +export function importsBrowserDriver(source) { + return BROWSER_DRIVER.test(source); +} + +/** @param {string} dir @returns {string[]} */ +function walk(dir) { + const out = []; + for (const entry of readdirSync(dir, { withFileTypes: true })) { + const path = join(dir, entry.name); + if (entry.isDirectory()) out.push(...walk(path)); + else if (entry.isFile() && entry.name.endsWith('.test.ts')) out.push(path); + } + return out; +} + +/** + * Every `*.test.ts` under `<root>/test` that imports a browser driver, as sorted + * repo-relative POSIX paths (the form `vitest list` prints). + * + * @param {string} root + * @returns {string[]} + */ +export function findBrowserTests(root) { + return walk(join(root, 'test')) + .filter((file) => importsBrowserDriver(readFileSync(file, 'utf8'))) + .map((file) => relative(root, file).split(sep).join('/')) + .sort(); +} + +/** + * Parse `vitest list --filesOnly` output into a set of repo-relative paths. Stray + * blank or decorative lines are ignored rather than assuming the format is pristine. + * + * @param {string} output + * @returns {Set<string>} + */ +export function parseVitestFileList(output) { + return new Set( + output + .split('\n') + .map((line) => line.trim()) + .filter((line) => line.endsWith('.test.ts')) + .map((line) => line.replace(/^\.\//, '')) + ); +} + +/** + * @param {string[]} browserTests + * @param {Set<string>} ciFiles + * @returns {string[]} browser-driven files that the CI config would still collect + */ +export function findLeaks(browserTests, ciFiles) { + return browserTests.filter((file) => ciFiles.has(file)); +} + +function main() { + const browserTests = findBrowserTests(ROOT); + + let collected; + try { + collected = execFileSync('npx', ['vitest', 'list', '--config', CI_CONFIG, '--filesOnly'], { + cwd: ROOT, + encoding: 'utf8', + stdio: ['ignore', 'pipe', 'pipe'], + }); + } catch (err) { + console.error('✗ could not enumerate the CI test set via `vitest list`.'); + console.error(err.stderr ? err.stderr.toString() : String(err)); + process.exit(1); + } + + const ciFiles = parseVitestFileList(collected); + if (ciFiles.size === 0) { + // An empty list would make every browser test look excluded: fail rather than pass vacuously. + console.error('✗ `vitest list` reported no test files; refusing to pass on an empty CI set.'); + process.exit(1); + } + + const leaked = findLeaks(browserTests, ciFiles); + + if (leaked.length > 0) { + console.error(`✗ ${leaked.length} browser-driven test file(s) are NOT excluded from ${CI_CONFIG}:\n`); + for (const file of leaked) console.error(` ${file}`); + console.error(` +These import a browser driver, so on a runner with no chromium they fail with +"browserType.launch: Executable doesn't exist" and take the suite down. Add each +to BROWSER_TEST_GLOBS in ${SUITES_FILE} (${CI_CONFIG} derives its excludes from +it, and \`npm run test:browser\` its includes). + +They may well pass on this machine; that is the trap. To reproduce a clean +runner locally: + PLAYWRIGHT_BROWSERS_PATH=\$(mktemp -d) PUPPETEER_CACHE_DIR=\$(mktemp -d) npm run test:ci`); + process.exit(1); + } + + console.log( + `✓ all ${browserTests.length} browser-driven test files are excluded from the CI suite (${ciFiles.size} files collected)` + ); +} + +if (process.argv[1] && resolve(process.argv[1]) === fileURLToPath(import.meta.url)) { + main(); +} diff --git a/scripts/git-hooks.mjs b/scripts/git-hooks.mjs new file mode 100644 index 00000000..ba2ade8e --- /dev/null +++ b/scripts/git-hooks.mjs @@ -0,0 +1,172 @@ +/** + * @fileoverview Git hook bodies + install policy, shared by scripts/postinstall.js and + * pinned by test/git-hooks.test.ts. + * + * Why a pre-push hook: the static CI job (lockfile, typecheck, lint, format, frontend + * syntax, ...) fails often on things a contributor could have caught locally in seconds, + * and finding out after a push costs a full CI round-trip plus a fix-up commit. Running + * the same checks before the push surfaces those failures in ~15s instead. + * + * Why pre-PUSH and not pre-commit: a commit is cheap and local, a push is what CI and + * reviewers pick up. And why the STATIC tier only: the unit/integration suite takes + * minutes, which nobody tolerates per push, so a hook that ran it would be bypassed + * within a day. The checks below mirror the static CI job and measured ~15s total. + * + * ⚠️ This installer is deliberately MARKER-OWNED, unlike the older pre-commit installer in + * postinstall.js which overwrites whatever it finds. A developer's own pre-push hook must + * survive `npm install`. + */ + +import { execFileSync } from 'node:child_process'; +import { chmodSync, existsSync, mkdirSync, readFileSync, realpathSync, writeFileSync } from 'node:fs'; +import { isAbsolute, join } from 'node:path'; + +/** Ownership marker. Bump the version suffix when the body changes meaningfully. */ +export const PRE_PUSH_MARKER = '# codeman-managed-hook: pre-push v1'; + +/** + * Checks that make up the fast tier, cheapest first so failures surface sooner. Each entry + * is the argument list for `npm run`, and each is a step of the static job in + * .github/workflows/ci.yml (test/git-hooks.test.ts pins that every script exists). + */ +export const PRE_PUSH_CHECKS = [ + ['check:lockfile'], + ['generate:cli-catalog', '--', '--check'], + ['check:browser-excludes'], + ['check:frontend-syntax'], + ['format:check'], + ['lint'], + ['typecheck'], +]; + +/** + * Render the pre-push hook script. + * + * POSIX sh, not bash: this ships to whatever shell the contributor's git uses. + */ +export function renderPrePushHook() { + const runs = PRE_PUSH_CHECKS.map((args) => `run_check ${args.join(' ')}`).join('\n'); + + return `#!/bin/sh +${PRE_PUSH_MARKER} +# Installed by scripts/postinstall.js. Edit scripts/git-hooks.mjs, not this file: +# it is regenerated on npm install. Delete the marker line above to take ownership +# and the installer will leave your version alone. +# +# Skip once: CODEMAN_SKIP_PREPUSH=1 git push +# Skip always: remove this file. + +[ "$CODEMAN_SKIP_PREPUSH" = "1" ] && exit 0 + +repo_root=$(git rev-parse --show-toplevel 2>/dev/null) || exit 0 +cd "$repo_root" || exit 0 + +# Nothing to check without dependencies (fresh clone, or a worktree that never ran +# npm install). Warn rather than blocking the push on a setup detail. +if [ ! -d node_modules ]; then + echo "pre-push: node_modules missing, skipping checks (run 'npm install' to enable them)." + exit 0 +fi + +# git feeds us "<localref> <localsha> <remoteref> <remotesha>" per ref. A deletion has an +# all-zero local sha and no tree worth checking; if every ref is a deletion, skip. +has_content=0 +while read -r _localref localsha _remoteref _remotesha; do + [ -z "$localsha" ] && continue + case "$localsha" in + 0000000000000000000000000000000000000000) ;; + *) has_content=1 ;; + esac +done +[ "$has_content" = "0" ] && exit 0 + +log=$(mktemp "\${TMPDIR:-/tmp}/codeman-prepush.XXXXXX") || exit 0 +trap 'rm -f "$log"' EXIT + +failed='' +run_check() { + if ! npm run --silent "$@" >"$log" 2>&1; then + echo "" + echo "pre-push: FAILED npm run $*" + tail -n 25 "$log" + failed="$failed $1" + fi +} + +echo "pre-push: running static checks (~15s)..." +${runs} + +if [ -n "$failed" ]; then + echo "" + echo "pre-push: blocked by:$failed" + echo "Fix, or push anyway with: CODEMAN_SKIP_PREPUSH=1 git push" + exit 1 +fi + +echo "pre-push: static checks passed." +exit 0 +`; +} + +/** + * Decide what to do with an existing hook file. + * + * @param {{ existing: string | null | undefined, next: string }} args + * @returns {'write' | 'up-to-date' | 'skip-foreign'} + */ +export function planHookInstall({ existing, next }) { + if (existing === null || existing === undefined || existing.trim() === '') return 'write'; + if (!existing.includes(PRE_PUSH_MARKER)) return 'skip-foreign'; + return existing === next ? 'up-to-date' : 'write'; +} + +/** @param {string} cwd @param {string[]} args */ +function git(cwd, args) { + return execFileSync('git', args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'] }).trim(); +} + +/** + * Resolve the hooks directory for the checkout rooted at `repoRoot`, or null when there + * is nothing to install into. + * + * Asks git (`--git-path hooks`) rather than assuming `<root>/.git/hooks`: in a worktree + * `.git` is a FILE pointing at the parent repo, and `core.hooksPath` can move it anywhere. + * + * Returns null unless `repoRoot` is itself the top of a work tree. Without that guard, a + * copy of this package sitting inside SOMEONE ELSE's repository (e.g. under their + * node_modules) would resolve to their hooks directory and install Codeman's hook there. + * + * @param {string} repoRoot + * @returns {string | null} + */ +export function resolveGitHooksDir(repoRoot) { + try { + const top = git(repoRoot, ['rev-parse', '--show-toplevel']); + if (!top || realpathSync(top) !== realpathSync(repoRoot)) return null; + const hooks = git(repoRoot, ['rev-parse', '--git-path', 'hooks']); + if (!hooks) return null; + return isAbsolute(hooks) ? hooks : join(repoRoot, hooks); + } catch { + return null; + } +} + +/** + * Install (or refresh) the managed pre-push hook in `hooksDir`, honouring + * {@link planHookInstall}: a hook without the marker is never touched. + * + * @param {string} hooksDir + * @returns {'write' | 'up-to-date' | 'skip-foreign'} + */ +export function installPrePushHook(hooksDir) { + const path = join(hooksDir, 'pre-push'); + const next = renderPrePushHook(); + const existing = existsSync(path) ? readFileSync(path, 'utf8') : null; + const action = planHookInstall({ existing, next }); + if (action === 'write') { + mkdirSync(hooksDir, { recursive: true }); + writeFileSync(path, next, { mode: 0o755 }); + chmodSync(path, 0o755); // `mode` only applies when the file is created + } + return action; +} diff --git a/scripts/postinstall.js b/scripts/postinstall.js index 02d23f59..6878d9f6 100644 --- a/scripts/postinstall.js +++ b/scripts/postinstall.js @@ -356,14 +356,17 @@ if (!isGlobalInstall) { } // ---------------------------------------------------------------------------- -// 5. Install git pre-commit hook (format check) +// 5. Install git hooks (pre-commit format check, pre-push static checks) // ---------------------------------------------------------------------------- if (!isGlobalInstall) { try { const { writeFileSync, mkdirSync } = await import('fs'); - const gitHooksDir = join(import.meta.dirname, '..', '.git', 'hooks'); - if (existsSync(join(import.meta.dirname, '..', '.git'))) { + const { resolveGitHooksDir, installPrePushHook } = await import('./git-hooks.mjs'); + // Resolved through git, not `../.git/hooks`: in a worktree `.git` is a file. + // null when this directory is not the top of a git checkout. + const gitHooksDir = resolveGitHooksDir(join(import.meta.dirname, '..')); + if (gitHooksDir) { mkdirSync(gitHooksDir, { recursive: true }); const hook = `#!/bin/bash # Auto-installed by postinstall — prevents CI format failures @@ -379,9 +382,18 @@ fi const hookPath = join(gitHooksDir, 'pre-commit'); writeFileSync(hookPath, hook, { mode: 0o755 }); console.log(colors.green('✓ Git pre-commit hook installed (prettier check)')); + + // Unlike the pre-commit hook above, this one is marker-owned: a pre-push + // hook the developer wrote themselves is left alone. + const action = installPrePushHook(gitHooksDir); + if (action === 'write') { + console.log(colors.green('✓ Git pre-push hook installed') + colors.dim(' (static CI checks, ~15s)')); + } else if (action === 'skip-foreign') { + console.log(colors.dim(' Existing pre-push hook left untouched (not Codeman-managed)')); + } } } catch { - // Non-critical — git hook is a convenience + // Non-critical — git hooks are a convenience } } diff --git a/test/check-browser-test-excludes.test.ts b/test/check-browser-test-excludes.test.ts new file mode 100644 index 00000000..7dc34ad8 --- /dev/null +++ b/test/check-browser-test-excludes.test.ts @@ -0,0 +1,97 @@ +/** + * @fileoverview scripts/check-browser-test-excludes.mjs: the detection side (which test + * files need a real browser) and the leak computation. The exclusion side is vitest's own + * `vitest list`, which `npm run check:browser-excludes` exercises for real in CI. + */ + +import { afterAll, beforeAll, describe, expect, it } from 'vitest'; +import { mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join, resolve } from 'node:path'; +import { + findBrowserTests, + findLeaks, + importsBrowserDriver, + parseVitestFileList, +} from '../scripts/check-browser-test-excludes.mjs'; +import { BROWSER_TEST_GLOBS } from '../config/test-suites'; + +const repoRoot = resolve(import.meta.dirname, '..'); + +// Fixture sources are assembled from the module name at runtime, so THIS file never contains +// a literal driver import and is not itself flagged by the checker it tests. +const fromImport = (mod: string) => `import { chromium, type Browser } from '${mod}';\n`; + +describe('importsBrowserDriver', () => { + it.each([ + fromImport('playwright'), + fromImport('playwright-core').replace(/'/g, '"'), + fromImport('@playwright/test'), + fromImport('puppeteer'), + `import type { Page } from '${'playwright'}';`, + `const { chromium } = require('${'playwright'}');`, + `const pw = await import('${'playwright'}');`, + ])('flags %s', (src) => { + expect(importsBrowserDriver(src)).toBe(true); + }); + + it.each([ + "import { describe } from 'vitest';", + "// needs ms-playwright's cache dir\nconst dir = '.cache/ms-playwright';", + fromImport('./playwright-helpers'), + fromImport('playwright-extra-thing'), + ])('ignores %s', (src) => { + expect(importsBrowserDriver(src)).toBe(false); + }); +}); + +describe('findBrowserTests (fixture tree)', () => { + let root: string; + beforeAll(() => { + root = mkdtempSync(join(tmpdir(), 'codeman-browser-excludes-')); + const put = (rel: string, src: string) => { + mkdirSync(join(root, rel, '..'), { recursive: true }); + writeFileSync(join(root, rel), src); + }; + put('test/unit.test.ts', "import { it } from 'vitest';\n"); + put('test/legacy-name.test.ts', fromImport('playwright')); + put('test/new.browser.test.ts', fromImport('playwright')); + put('test/nested/deep.test.ts', fromImport('puppeteer')); + put('test/helpers/browser.ts', fromImport('playwright')); // not a test file + }); + afterAll(() => rmSync(root, { recursive: true, force: true })); + + it('finds driver imports by content, recursively, as sorted repo-relative paths', () => { + expect(findBrowserTests(root)).toEqual([ + 'test/legacy-name.test.ts', + 'test/nested/deep.test.ts', + 'test/new.browser.test.ts', + ]); + }); +}); + +describe('parseVitestFileList + findLeaks', () => { + it('keeps only test paths and normalizes a leading ./', () => { + const out = '\n./test/a.test.ts\ntest/b.test.ts\nsome banner line\n test/c.test.ts \n'; + expect([...parseVitestFileList(out)].sort()).toEqual(['test/a.test.ts', 'test/b.test.ts', 'test/c.test.ts']); + }); + + it('reports exactly the browser tests the CI set still collects', () => { + const ci = new Set(['test/unit.test.ts', 'test/legacy-name.test.ts']); + expect(findLeaks(['test/legacy-name.test.ts', 'test/new.browser.test.ts'], ci)).toEqual([ + 'test/legacy-name.test.ts', + ]); + expect(findLeaks(['test/new.browser.test.ts'], ci)).toEqual([]); + }); +}); + +describe('against this repository', () => { + it('detects every file already listed in BROWSER_TEST_GLOBS', () => { + // If detection stopped recognising a known browser test, the checker would go blind to + // exactly the class of file it exists for. + const detected = new Set(findBrowserTests(repoRoot)); + const literals = BROWSER_TEST_GLOBS.filter((g) => !/[*?[{]/.test(g)); + expect(literals.length).toBeGreaterThan(0); + for (const file of literals) expect(detected, file).toContain(file); + }); +}); diff --git a/test/git-hooks.test.ts b/test/git-hooks.test.ts new file mode 100644 index 00000000..04caf9a7 --- /dev/null +++ b/test/git-hooks.test.ts @@ -0,0 +1,283 @@ +/** + * @fileoverview The pre-push hook that scripts/postinstall.js installs (scripts/git-hooks.mjs). + * + * Two properties matter more than the hook's contents, because the older pre-commit + * installer gets both wrong and this one must not copy it: + * 1. It is MARKER-OWNED: a hook the developer wrote by hand is never overwritten. + * 2. The hooks directory is resolved via `git rev-parse --git-path hooks`, since in a + * worktree `.git` is a FILE and `<root>/.git/hooks` does not exist. + * + * ⚠️ Every filesystem/git test here runs against THROWAWAY repositories under a temp dir. + * Never point the installer at this checkout: its hooks directory is shared with every + * worktree of it, including whatever the developer is running right now. + */ + +import { afterAll, beforeAll, describe, expect, it } from 'vitest'; +import { execFileSync, spawnSync } from 'node:child_process'; +import { + chmodSync, + mkdirSync, + mkdtempSync, + readFileSync, + realpathSync, + rmSync, + statSync, + writeFileSync, +} from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join, resolve } from 'node:path'; +import { + PRE_PUSH_CHECKS, + PRE_PUSH_MARKER, + installPrePushHook, + planHookInstall, + renderPrePushHook, + resolveGitHooksDir, +} from '../scripts/git-hooks.mjs'; + +const repoRoot = resolve(import.meta.dirname, '..'); +const read = (rel: string) => readFileSync(resolve(repoRoot, rel), 'utf8'); + +/** git with no user/system config leaking in (a global core.hooksPath would redirect everything). */ +const GIT_ENV = { + ...process.env, + GIT_CONFIG_NOSYSTEM: '1', + GIT_CONFIG_GLOBAL: '/dev/null', + GIT_AUTHOR_NAME: 'test', + GIT_AUTHOR_EMAIL: 'test@example.invalid', + GIT_COMMITTER_NAME: 'test', + GIT_COMMITTER_EMAIL: 'test@example.invalid', + CODEMAN_SKIP_PREPUSH: '', +}; + +function git(cwd: string, args: string[], env: NodeJS.ProcessEnv = GIT_ENV): string { + return execFileSync('git', args, { cwd, env, encoding: 'utf8', stdio: ['ignore', 'pipe', 'pipe'] }).trim(); +} + +let scratch: string; +beforeAll(() => { + scratch = realpathSync(mkdtempSync(join(tmpdir(), 'codeman-git-hooks-'))); +}); +afterAll(() => { + rmSync(scratch, { recursive: true, force: true }); +}); + +let counter = 0; +function newRepo(): string { + const dir = join(scratch, `repo-${++counter}`); + mkdirSync(dir, { recursive: true }); + git(dir, ['init', '-q', '-b', 'main']); + git(dir, ['commit', '-q', '--allow-empty', '-m', 'init']); + return dir; +} + +describe('pre-push hook body', () => { + const hook = renderPrePushHook(); + + it('carries the ownership marker', () => { + expect(hook).toContain(PRE_PUSH_MARKER); + }); + + it('runs every configured check through npm, and nothing slow', () => { + for (const args of PRE_PUSH_CHECKS) { + expect(hook).toContain(`run_check ${args.join(' ')}`); + } + expect(hook).toContain('npm run --silent "$@"'); + // The whole point of the tier: the minutes-long suites stay out of a per-push hook. + expect(hook).not.toMatch(/\btest:(ci|browser|mobile|perf|all)\b/); + }); + + it('is POSIX sh', () => { + expect(hook.startsWith('#!/bin/sh\n')).toBe(true); + const r = spawnSync('sh', ['-n'], { input: hook }); + expect(r.status).toBe(0); + }); +}); + +describe('pre-push checks match the static CI job', () => { + const scripts = JSON.parse(read('package.json')).scripts as Record<string, string>; + const ci = read('.github/workflows/ci.yml'); + + it.each(PRE_PUSH_CHECKS.map((args) => [args.join(' ')] as const))('%s is a real script that CI runs', (joined) => { + const [name] = joined.split(' '); + expect(scripts[name], `package.json has no "${name}" script`).toBeTypeOf('string'); + expect(ci).toContain(`npm run ${joined}`); + }); +}); + +describe('planHookInstall', () => { + const hook = renderPrePushHook(); + + it('writes when no hook exists', () => { + expect(planHookInstall({ existing: null, next: hook })).toBe('write'); + }); + + it('refuses to clobber a hook it does not own', () => { + expect(planHookInstall({ existing: '#!/bin/sh\nmake lint\n', next: hook })).toBe('skip-foreign'); + }); + + it('refreshes its own hook when the body changed', () => { + expect(planHookInstall({ existing: `#!/bin/sh\n${PRE_PUSH_MARKER}\necho old\n`, next: hook })).toBe('write'); + }); + + it('is idempotent when already current', () => { + expect(planHookInstall({ existing: hook, next: hook })).toBe('up-to-date'); + }); + + it('treats an empty file as absent rather than foreign', () => { + expect(planHookInstall({ existing: ' \n', next: hook })).toBe('write'); + }); +}); + +describe('resolveGitHooksDir (temp repos)', () => { + it('resolves <root>/.git/hooks in a plain checkout', () => { + const repo = newRepo(); + expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks')); + }); + + it('resolves the SHARED hooks dir from a worktree, where .git is a file', () => { + const repo = newRepo(); + const wt = join(scratch, `wt-${counter}`); + git(repo, ['worktree', 'add', '-q', wt, '-b', 'wt-branch']); + expect(statSync(join(wt, '.git')).isFile()).toBe(true); + expect(resolveGitHooksDir(wt)).toBe(join(repo, '.git', 'hooks')); + }); + + it('returns null outside any git checkout', () => { + const dir = join(scratch, `plain-${++counter}`); + mkdirSync(dir); + expect(resolveGitHooksDir(dir)).toBeNull(); + }); + + it("returns null for a copy nested inside someone else's repo (e.g. under node_modules)", () => { + const repo = newRepo(); + const nested = join(repo, 'node_modules', 'aicodeman'); + mkdirSync(nested, { recursive: true }); + expect(resolveGitHooksDir(nested)).toBeNull(); + }); +}); + +describe('installPrePushHook (temp repos)', () => { + it('writes an executable hook into a fresh repo', () => { + const hooks = join(newRepo(), '.git', 'hooks'); + expect(installPrePushHook(hooks)).toBe('write'); + const path = join(hooks, 'pre-push'); + expect(readFileSync(path, 'utf8')).toBe(renderPrePushHook()); + expect(statSync(path).mode & 0o111).not.toBe(0); + expect(installPrePushHook(hooks)).toBe('up-to-date'); + }); + + it('leaves a foreign pre-push hook byte-identical', () => { + const hooks = join(newRepo(), '.git', 'hooks'); + const path = join(hooks, 'pre-push'); + const mine = '#!/bin/sh\n# my own hook\nexit 0\n'; + writeFileSync(path, mine, { mode: 0o755 }); + expect(installPrePushHook(hooks)).toBe('skip-foreign'); + expect(readFileSync(path, 'utf8')).toBe(mine); + }); + + it('refreshes a stale managed hook and keeps it executable', () => { + const hooks = join(newRepo(), '.git', 'hooks'); + const path = join(hooks, 'pre-push'); + writeFileSync(path, `#!/bin/sh\n${PRE_PUSH_MARKER}\necho old\n`, { mode: 0o644 }); + expect(installPrePushHook(hooks)).toBe('write'); + expect(readFileSync(path, 'utf8')).toBe(renderPrePushHook()); + expect(statSync(path).mode & 0o111).not.toBe(0); + }); +}); + +/** + * Drive the rendered hook through a real `git push` to a local bare remote. The repo gets a + * stub package.json whose check scripts only record that they ran, so this exercises the + * hook's control flow (ref parsing, skips, blocking) without running the real checks. + */ +describe('the installed hook on a real push (temp repos)', () => { + function setup(opts: { failing?: string; nodeModules?: boolean } = {}) { + const repo = newRepo(); + const remote = join(scratch, `remote-${counter}.git`); + git(scratch, ['init', '-q', '--bare', remote]); + git(repo, ['remote', 'add', 'origin', remote]); + const log = join(repo, 'ran.log'); + const scripts: Record<string, string> = {}; + for (const [name] of PRE_PUSH_CHECKS) { + scripts[name] = + name === opts.failing ? `echo ${name} >> ran.log && echo boom-${name} && exit 1` : `echo ${name} >> ran.log`; + } + writeFileSync(join(repo, 'package.json'), JSON.stringify({ name: 'hook-fixture', private: true, scripts })); + writeFileSync(join(repo, '.gitignore'), 'node_modules/\nran.log\n'); + git(repo, ['add', 'package.json', '.gitignore']); + git(repo, ['commit', '-q', '-m', 'fixture']); + if (opts.nodeModules !== false) mkdirSync(join(repo, 'node_modules')); + installPrePushHook(join(repo, '.git', 'hooks')); + chmodSync(join(repo, '.git', 'hooks', 'pre-push'), 0o755); + const ran = () => { + try { + return readFileSync(log, 'utf8').trim().split('\n').filter(Boolean); + } catch { + return []; + } + }; + const push = (args: string[], env: NodeJS.ProcessEnv = {}) => + spawnSync('git', ['push', ...args], { cwd: repo, env: { ...GIT_ENV, ...env }, encoding: 'utf8' }); + return { repo, remote, ran, push }; + } + + /** What the stubs record: npm appends the args after `--` to the script, so they prove forwarding. */ + const expectedRuns = PRE_PUSH_CHECKS.map((args) => args.filter((a) => a !== '--').join(' ')); + + it('runs every check before a push, in order', () => { + const { ran, push } = setup(); + const r = push(['-q', 'origin', 'main']); + expect(r.status, r.stderr + r.stdout).toBe(0); + expect(ran()).toEqual(expectedRuns); + }); + + it('blocks the push when a check fails, but still runs the rest', () => { + const { ran, push, remote } = setup({ failing: 'lint' }); + const r = push(['origin', 'main']); + expect(r.status).not.toBe(0); + expect(r.stdout + r.stderr).toContain('pre-push: FAILED npm run lint'); + expect(r.stdout + r.stderr).toContain('boom-lint'); + expect(ran()).toEqual(expectedRuns); + expect(spawnSync('git', ['rev-parse', '--verify', '-q', 'refs/heads/main'], { cwd: remote }).status).not.toBe(0); + }); + + it('CODEMAN_SKIP_PREPUSH=1 skips every check', () => { + const { ran, push } = setup({ failing: 'lint' }); + const r = push(['-q', 'origin', 'main'], { CODEMAN_SKIP_PREPUSH: '1' }); + expect(r.status, r.stderr).toBe(0); + expect(ran()).toEqual([]); + }); + + it('a delete-only push skips the checks', () => { + const { ran, push, repo } = setup({ failing: 'lint' }); + expect(push(['-q', 'origin', 'main'], { CODEMAN_SKIP_PREPUSH: '1' }).status).toBe(0); + git(repo, ['branch', 'doomed']); + expect(push(['-q', 'origin', 'doomed'], { CODEMAN_SKIP_PREPUSH: '1' }).status).toBe(0); + const r = push(['-q', 'origin', '--delete', 'doomed']); + expect(r.status, r.stderr).toBe(0); + expect(ran()).toEqual([]); + }); + + it('skips (never blocks) when node_modules is absent', () => { + const { ran, push } = setup({ failing: 'lint', nodeModules: false }); + const r = push(['origin', 'main']); + expect(r.status, r.stderr).toBe(0); + expect(r.stdout + r.stderr).toContain('node_modules missing'); + expect(ran()).toEqual([]); + }); +}); + +describe('postinstall wiring', () => { + const postinstall = read('scripts/postinstall.js'); + + it('installs the pre-push hook through the shared module', () => { + expect(postinstall).toContain("import('./git-hooks.mjs')"); + expect(postinstall).toContain('installPrePushHook(gitHooksDir)'); + }); + + it('resolves the hooks dir through git, so worktrees work', () => { + expect(postinstall).toContain('resolveGitHooksDir('); + expect(postinstall).not.toContain("join(import.meta.dirname, '..', '.git', 'hooks')"); + }); +}); From e60b5a8a2c0ab72c6053914c0c8438b5bffe37da Mon Sep 17 00:00:00 2001 From: Aamer Akhter <aakhter@gmail.com> Date: Sat, 26 Sep 2026 22:44:14 -0400 Subject: [PATCH 17/46] build: address review on the pre-push hook and hooks-dir resolution resolveGitHooksDir now returns a directory only when it is the repo's own <git-common-dir>/hooks (compared on canonical paths), so a core.hooksPath elsewhere, global or repo-local, is never written to by postinstall, while a core.hooksPath pointing back at the repo's own .git/hooks still resolves. The pre-push hook skips with a one-line notice when a pushed ref is not the checked-out HEAD (tags peeled) or when git status shows uncommitted or untracked changes under a path the checks read (src, config, scripts, test, package.json, package-lock.json, install.sh), since the checks read the working tree rather than the pushed commit. Also: honest timing (~10-40s instead of ~15s), CLAUDE.md Session Safety note on CODEMAN_SKIP_PREPUSH for another session's WIP, 14 (not 9) Playwright tests, and a note that the browser-excludes check only sees direct imports. --- .github/CONTRIBUTING.md | 2 +- CLAUDE.md | 5 +- docs/wiki/Contributing.md | 6 +- scripts/check-browser-test-excludes.mjs | 3 + scripts/git-hooks.mjs | 90 +++++++++++++-- scripts/postinstall.js | 2 +- test/git-hooks.test.ts | 143 +++++++++++++++++++++++- 7 files changed, 232 insertions(+), 19 deletions(-) diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md index 1468cff6..2ed1d8f1 100644 --- a/.github/CONTRIBUTING.md +++ b/.github/CONTRIBUTING.md @@ -35,7 +35,7 @@ npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules npm run check:browser-excludes # every browser-driven test is kept out of `npm test` ``` -`npm install` also installs a `pre-push` git hook that runs these static checks (~15s) and blocks the push if one fails. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; it never replaces a `pre-push` hook of your own. +`npm install` also installs a `pre-push` git hook that runs these static checks (about 10-40s, machine-dependent) and blocks the push if one fails. It skips itself when you push something other than the checked-out HEAD, or when the tree has uncommitted changes the checks would read. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; it never replaces a `pre-push` hook of your own. ### Tests diff --git a/CLAUDE.md b/CLAUDE.md index a262f074..9ae9febc 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -34,6 +34,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co - To land a commit on master **without** switching branches (which would yank the tree out from under the other session): `git push origin HEAD:master` then `git branch -f master HEAD`. Never `git checkout master` to "fix" it. - **Never `git add -A`/`git add .`** — stage explicit paths. A sweep will pick up another session's WIP. - Another session's broken WIP can block `npm run build`, since `tsc` is the first step and the build gates on it. That is not your bug to fix. ⚠️ `tsc` still EMITS on type errors, so a failed `npm run build` leaves a rebuilt `dist/index.js` compiled from their tree; check what it pulled in before restarting the service. To deploy frontend-only changes past a blocked `tsc`, run the asset stage of `scripts/build.mjs` (everything after the `tsc`/`chmod` lines is independent of it). +- **A pre-push failure in a file you did not touch is another session's WIP.** Push with `CODEMAN_SKIP_PREPUSH=1 git push` and leave it alone. (The hook already skips itself when the tree has uncommitted changes in a path it checks, so this mostly happens once the other session has committed.) ## CRITICAL: Always Test Before Deploying @@ -113,7 +114,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | Check public-asset formatting | `npm run check:public-assets` (prettier-checks `src/web/public/**` text assets; `scripts/check-public-assets.mjs`) | | Frontend JS syntax check | `npm run check:frontend-syntax` (`scripts/check-frontend-syntax.mjs`; runs in CI) | | Browser-test exclusion check | `npm run check:browser-excludes` (`scripts/check-browser-test-excludes.mjs`; runs in CI, <1s). Fails if a test importing playwright/puppeteer is still collected by `config/vitest.ci.config.ts`; add it to `BROWSER_TEST_GLOBS` in `config/test-suites.ts` | -| Pre-push hook | Installed by `npm install` (`scripts/git-hooks.mjs`, via postinstall): runs the static CI checks (~15s) before `git push`. Skip once: `CODEMAN_SKIP_PREPUSH=1 git push`. Marker-owned, so a hand-written `pre-push` is never overwritten; hooks dir resolved via `git rev-parse --git-path hooks` (worktree-safe) | +| Pre-push hook | Installed by `npm install` (`scripts/git-hooks.mjs`, via postinstall): runs the static CI checks (~10-40s) before `git push`. Skip once: `CODEMAN_SKIP_PREPUSH=1 git push`. Skips itself with a notice when HEAD is not the pushed commit or the tree has uncommitted changes the checks would read. Marker-owned, so a hand-written `pre-push` is never overwritten; installs ONLY into the repo's own `<git-common-dir>/hooks` (worktree-safe; a `core.hooksPath` elsewhere, e.g. a global one, is left alone) | | Excluded-suite runners | `npm run test:browser` · `npm run test:mobile` · `npm run test:perf` · `npm run test:all` (everything, environmental failures included) — see Testing | | Production start | `npm run start` | | Production logs | `journalctl --user -u codeman-web -f` | @@ -122,7 +123,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph | Dependency doctor | `codeman doctor` (alias `check-deps`; `--json`, `--category core\|office\|other`). Probes Node/Claude CLI/tmux/LibreOffice/MS Office against `config/dependency-registry.ts`; engine is pure given an injectable `ProbeHost` | | Multi-user accounts | `codeman users add <name>` / `passwd <name>` / `list` / `rm <name>` (writes `~/.codeman/users.json`, mode 0600; see Multi-user mode) | -**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `check:browser-excludes`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 9 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`). +**CI**: `.github/workflows/ci.yml` (push to master/main + PRs, Node 22) runs two jobs: **(1)** `check:lockfile`, `typecheck`, `lint`, `check:frontend-syntax`, `check:browser-excludes`, `format:check`, then a **server boot smoke test** (`tsx src/index.ts web --port 3151` must answer `/api/status` within 30s); **(2)** the **unit/integration test suite** via `npm run test:ci` (`config/vitest.ci.config.ts` — excludes the browser-driven `test/mobile/**` suite, `perf-*` benchmarks, and 14 Playwright tests; globs live in `config/test-suites.ts`), followed by the **`packages/xterm-zerolag-input` package tests** (a bare `npx vitest run` in that directory; its vitest is hoisted by the root `npm ci`, so no separate install, and `npm test` at the root does NOT run them). `npm test` runs this same config, so local green == CI green. Tests are tmux-safe in CI: `TmuxManager` no-ops all shell commands under `VITEST` (see Testing). A third workflow, `wiki-sync.yml`, fires only on master pushes touching `docs/wiki/**` and mirrors that directory to the GitHub wiki (browser edits to the wiki are overwritten by the next sync, so fix pages via `docs/wiki/`). **Code style**: Prettier (`singleQuote: true`, `printWidth: 120`, `trailingComma: "es5"`) — config lives in the **`"prettier"` key of `package.json`**, not a `.prettierrc` (keeps the repo root short; editors read it natively). `.prettierignore` stays at the root because Prettier resolves it relative to cwd. ESLint flat config (`config/eslint.config.js`) allows `no-console`, warns on `@typescript-eslint/no-explicit-any`. Ignores: `app.js`, `scripts/**/*.mjs`, `src/web/public/vendor/**`, `scripts/remotion/**`. diff --git a/docs/wiki/Contributing.md b/docs/wiki/Contributing.md index 87ce1a0a..c08229ee 100644 --- a/docs/wiki/Contributing.md +++ b/docs/wiki/Contributing.md @@ -47,8 +47,10 @@ npm test -- test/<file>.test.ts # one file, the normal way npm run test:ci # the full CI sweep ``` -`npm install` installs a `pre-push` git hook that runs the static checks above (~15s) and -blocks a push that would fail them. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; a +`npm install` installs a `pre-push` git hook that runs the static checks above (about 10-40s, +machine-dependent) and blocks a push that would fail them. It skips itself when you push +something other than the checked-out HEAD, or when the tree has uncommitted changes the +checks would read. Skip it once with `CODEMAN_SKIP_PREPUSH=1 git push`; a `pre-push` hook of your own is never overwritten. **Never run bare `npm test`.** The default configuration includes browser-driven Playwright diff --git a/scripts/check-browser-test-excludes.mjs b/scripts/check-browser-test-excludes.mjs index 209160ce..3d3b6233 100644 --- a/scripts/check-browser-test-excludes.mjs +++ b/scripts/check-browser-test-excludes.mjs @@ -19,6 +19,9 @@ * (`inline-rename`, `opencode-resize`, `webgl-fallback`, * `terminal-copy-shortcut`, `codex-predictive-echo`). What actually makes a * file dangerous is importing a browser driver, so that is what is tested. + * ⚠️ Only a DIRECT import is seen: a test that reaches playwright through a + * helper module (e.g. `test/mobile/helpers/browser.ts`) is not detected, so + * such a test still has to be added to `BROWSER_TEST_GLOBS` by hand. * * 2. **The exclusion side is answered by vitest itself**, via * `vitest list --filesOnly`, rather than by re-implementing glob matching diff --git a/scripts/git-hooks.mjs b/scripts/git-hooks.mjs index ba2ade8e..a0b5770d 100644 --- a/scripts/git-hooks.mjs +++ b/scripts/git-hooks.mjs @@ -5,12 +5,19 @@ * Why a pre-push hook: the static CI job (lockfile, typecheck, lint, format, frontend * syntax, ...) fails often on things a contributor could have caught locally in seconds, * and finding out after a push costs a full CI round-trip plus a fix-up commit. Running - * the same checks before the push surfaces those failures in ~15s instead. + * the same checks before the push surfaces those failures in ~10-40s instead (12s on a fast + * workstation, ~35s measured elsewhere; typecheck, format:check and lint dominate). * * Why pre-PUSH and not pre-commit: a commit is cheap and local, a push is what CI and * reviewers pick up. And why the STATIC tier only: the unit/integration suite takes * minutes, which nobody tolerates per push, so a hook that ran it would be bypassed - * within a day. The checks below mirror the static CI job and measured ~15s total. + * within a day. The checks below mirror the static CI job. + * + * ⚠️ The checks read the WORKING TREE, not the commits being pushed. So the hook skips + * (with a one-line notice) whenever the two can differ: when HEAD is not the commit being + * pushed, and when `git status` shows uncommitted or untracked changes in a path a check + * reads ({@link PRE_PUSH_WATCHED_PATHS}). In a checkout shared by several agent sessions + * the second case is usually another session's WIP, which must not block this push. * * ⚠️ This installer is deliberately MARKER-OWNED, unlike the older pre-commit installer in * postinstall.js which overwrites whatever it finds. A developer's own pre-push hook must @@ -19,7 +26,7 @@ import { execFileSync } from 'node:child_process'; import { chmodSync, existsSync, mkdirSync, readFileSync, realpathSync, writeFileSync } from 'node:fs'; -import { isAbsolute, join } from 'node:path'; +import { basename, dirname, join, resolve } from 'node:path'; /** Ownership marker. Bump the version suffix when the body changes meaningfully. */ export const PRE_PUSH_MARKER = '# codeman-managed-hook: pre-push v1'; @@ -39,6 +46,25 @@ export const PRE_PUSH_CHECKS = [ ['typecheck'], ]; +/** + * Paths whose uncommitted state would leak into a check, so a dirty one makes the hook skip. + * Derived from what each check reads: src/ (format:check, lint, typecheck, + * check:frontend-syntax), config/ (eslint + vitest configs, test-suites.ts, the CLI + * catalogue), scripts/ (every check is a script there, and typecheck's second pass compiles + * one), test/ (check:browser-excludes scans it and runs `vitest list` over it), + * package.json + package-lock.json (check:lockfile) and install.sh (generate:cli-catalog + * --check diffs its generated block). + */ +export const PRE_PUSH_WATCHED_PATHS = [ + 'src', + 'config', + 'scripts', + 'test', + 'package.json', + 'package-lock.json', + 'install.sh', +]; + /** * Render the pre-push hook script. * @@ -46,6 +72,7 @@ export const PRE_PUSH_CHECKS = [ */ export function renderPrePushHook() { const runs = PRE_PUSH_CHECKS.map((args) => `run_check ${args.join(' ')}`).join('\n'); + const watched = PRE_PUSH_WATCHED_PATHS.join(' '); return `#!/bin/sh ${PRE_PUSH_MARKER} @@ -70,16 +97,36 @@ fi # git feeds us "<localref> <localsha> <remoteref> <remotesha>" per ref. A deletion has an # all-zero local sha and no tree worth checking; if every ref is a deletion, skip. +# The checks below read the working tree, so they only say something about a pushed commit +# that IS the checked-out HEAD (tags are peeled to their commit first). +head=$(git rev-parse -q --verify HEAD 2>/dev/null) has_content=0 -while read -r _localref localsha _remoteref _remotesha; do +not_head='' +while read -r localref localsha _remoteref _remotesha; do [ -z "$localsha" ] && continue case "$localsha" in 0000000000000000000000000000000000000000) ;; - *) has_content=1 ;; + *) + has_content=1 + commit=$(git rev-parse -q --verify "$localsha^{commit}" 2>/dev/null) + [ -n "$head" ] && [ "$commit" = "$head" ] || not_head="$localref" + ;; esac done [ "$has_content" = "0" ] && exit 0 +if [ -n "$not_head" ]; then + echo "pre-push: skipping static checks: $not_head is not the checked-out HEAD, and the checks read the working tree." + exit 0 +fi + +# Uncommitted or untracked changes in a path a check reads would be judged instead of the +# pushed commit. In a checkout shared by several sessions that is usually someone else's WIP. +if [ -n "$(git --no-optional-locks status --porcelain -- ${watched} 2>/dev/null)" ]; then + echo "pre-push: skipping static checks: uncommitted changes under ${watched} would be checked instead of the pushed commit." + exit 0 +fi + log=$(mktemp "\${TMPDIR:-/tmp}/codeman-prepush.XXXXXX") || exit 0 trap 'rm -f "$log"' EXIT @@ -93,7 +140,7 @@ run_check() { fi } -echo "pre-push: running static checks (~15s)..." +echo "pre-push: running static checks (~10-40s)..." ${runs} if [ -n "$failed" ]; then @@ -125,15 +172,33 @@ function git(cwd, args) { return execFileSync('git', args, { cwd, encoding: 'utf8', stdio: ['ignore', 'pipe', 'ignore'] }).trim(); } +/** + * realpath() that tolerates a missing leaf: a fresh `.git` may have no `hooks/` yet, so + * canonicalize the parent and re-append the name. Throws if the parent is missing too. + * + * @param {string} path + */ +function canonicalPath(path) { + return existsSync(path) ? realpathSync(path) : join(realpathSync(dirname(path)), basename(path)); +} + /** * Resolve the hooks directory for the checkout rooted at `repoRoot`, or null when there * is nothing to install into. * * Asks git (`--git-path hooks`) rather than assuming `<root>/.git/hooks`: in a worktree - * `.git` is a FILE pointing at the parent repo, and `core.hooksPath` can move it anywhere. + * `.git` is a FILE pointing at the parent repo, so the hooks live under + * `--git-common-dir`. * - * Returns null unless `repoRoot` is itself the top of a work tree. Without that guard, a - * copy of this package sitting inside SOMEONE ELSE's repository (e.g. under their + * ⚠️ Returns a directory ONLY when it is this repository's own `<git-common-dir>/hooks`. + * `--git-path hooks` also reports `core.hooksPath`, and that setting is often GLOBAL (a + * shared hooks directory used by every repo on the machine); installing there would + * overwrite the user's own hooks and run Codeman's checks on unrelated repos. A + * `core.hooksPath` that points back at the repo's own hooks dir still resolves, because + * the comparison is on canonical paths rather than on whether the setting exists. + * + * Also returns null unless `repoRoot` is itself the top of a work tree. Without that guard, + * a copy of this package sitting inside SOMEONE ELSE's repository (e.g. under their * node_modules) would resolve to their hooks directory and install Codeman's hook there. * * @param {string} repoRoot @@ -143,9 +208,12 @@ export function resolveGitHooksDir(repoRoot) { try { const top = git(repoRoot, ['rev-parse', '--show-toplevel']); if (!top || realpathSync(top) !== realpathSync(repoRoot)) return null; + // Both are printed relative to the cwd (repoRoot) unless already absolute. const hooks = git(repoRoot, ['rev-parse', '--git-path', 'hooks']); - if (!hooks) return null; - return isAbsolute(hooks) ? hooks : join(repoRoot, hooks); + const common = git(repoRoot, ['rev-parse', '--git-common-dir']); + if (!hooks || !common) return null; + const own = join(realpathSync(resolve(repoRoot, common)), 'hooks'); + return canonicalPath(resolve(repoRoot, hooks)) === own ? own : null; } catch { return null; } diff --git a/scripts/postinstall.js b/scripts/postinstall.js index 6878d9f6..414ac2e3 100644 --- a/scripts/postinstall.js +++ b/scripts/postinstall.js @@ -387,7 +387,7 @@ fi // hook the developer wrote themselves is left alone. const action = installPrePushHook(gitHooksDir); if (action === 'write') { - console.log(colors.green('✓ Git pre-push hook installed') + colors.dim(' (static CI checks, ~15s)')); + console.log(colors.green('✓ Git pre-push hook installed') + colors.dim(' (static CI checks, ~10-40s)')); } else if (action === 'skip-foreign') { console.log(colors.dim(' Existing pre-push hook left untouched (not Codeman-managed)')); } diff --git a/test/git-hooks.test.ts b/test/git-hooks.test.ts index 04caf9a7..8ba70e87 100644 --- a/test/git-hooks.test.ts +++ b/test/git-hooks.test.ts @@ -4,8 +4,10 @@ * Two properties matter more than the hook's contents, because the older pre-commit * installer gets both wrong and this one must not copy it: * 1. It is MARKER-OWNED: a hook the developer wrote by hand is never overwritten. - * 2. The hooks directory is resolved via `git rev-parse --git-path hooks`, since in a - * worktree `.git` is a FILE and `<root>/.git/hooks` does not exist. + * 2. The hooks directory is resolved through git, since in a worktree `.git` is a FILE + * and `<root>/.git/hooks` does not exist, and it is ONLY ever the repo's own + * `<git-common-dir>/hooks`: a `core.hooksPath` elsewhere (typically a global one) is + * never written to. * * ⚠️ Every filesystem/git test here runs against THROWAWAY repositories under a temp dir. * Never point the installer at this checkout: its hooks directory is shared with every @@ -29,6 +31,7 @@ import { join, resolve } from 'node:path'; import { PRE_PUSH_CHECKS, PRE_PUSH_MARKER, + PRE_PUSH_WATCHED_PATHS, installPrePushHook, planHookInstall, renderPrePushHook, @@ -149,6 +152,64 @@ describe('resolveGitHooksDir (temp repos)', () => { expect(resolveGitHooksDir(dir)).toBeNull(); }); + it('returns null when a repo-local core.hooksPath points outside the repo', () => { + const repo = newRepo(); + const outside = join(scratch, `shared-hooks-${counter}`); + mkdirSync(outside); + git(repo, ['config', 'core.hooksPath', outside]); + expect(resolveGitHooksDir(repo)).toBeNull(); + }); + + it('returns null when core.hooksPath points at a directory that does not exist yet', () => { + const repo = newRepo(); + git(repo, ['config', 'core.hooksPath', join(scratch, `missing-${counter}`, 'hooks')]); + expect(resolveGitHooksDir(repo)).toBeNull(); + }); + + it("still resolves when core.hooksPath points at the repo's OWN .git/hooks", () => { + const repo = newRepo(); + git(repo, ['config', 'core.hooksPath', join(repo, '.git', 'hooks')]); + expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks')); + }); + + it('resolves before .git/hooks exists (compares the would-be path)', () => { + const repo = newRepo(); + rmSync(join(repo, '.git', 'hooks'), { recursive: true, force: true }); + expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks')); + }); + + it('returns null under a GLOBAL core.hooksPath, from a checkout and from a worktree', () => { + const repo = newRepo(); + const wt = join(scratch, `wt-global-${counter}`); + git(repo, ['worktree', 'add', '-q', wt, '-b', 'wt-global']); + const globalHooks = join(scratch, `global-hooks-${counter}`); + mkdirSync(globalHooks); + const globalConfig = join(scratch, `gitconfig-${counter}`); + writeFileSync(globalConfig, `[core]\n\thooksPath = ${globalHooks}\n`); + // resolveGitHooksDir runs git with the ambient environment, so scope the fake global + // config to this test through process.env (never the developer's real ~/.gitconfig). + const saved = { + GIT_CONFIG_GLOBAL: process.env.GIT_CONFIG_GLOBAL, + GIT_CONFIG_NOSYSTEM: process.env.GIT_CONFIG_NOSYSTEM, + }; + process.env.GIT_CONFIG_GLOBAL = globalConfig; + process.env.GIT_CONFIG_NOSYSTEM = '1'; + try { + expect(git(repo, ['rev-parse', '--git-path', 'hooks'], { ...GIT_ENV, GIT_CONFIG_GLOBAL: globalConfig })).toBe( + globalHooks + ); + expect(resolveGitHooksDir(repo)).toBeNull(); + expect(resolveGitHooksDir(wt)).toBeNull(); + } finally { + for (const [k, v] of Object.entries(saved)) { + if (v === undefined) delete process.env[k]; + else process.env[k] = v; + } + } + // Control: the same repo resolves again once the global setting is gone. + expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks')); + }); + it("returns null for a copy nested inside someone else's repo (e.g. under node_modules)", () => { const repo = newRepo(); const nested = join(repo, 'node_modules', 'aicodeman'); @@ -259,6 +320,84 @@ describe('the installed hook on a real push (temp repos)', () => { expect(ran()).toEqual([]); }); + it('skips when the pushed ref is not the checked-out HEAD', () => { + const { ran, push, repo } = setup({ failing: 'lint' }); + git(repo, ['branch', 'other']); + git(repo, ['commit', '-q', '--allow-empty', '-m', 'only on main']); + git(repo, ['checkout', '-q', 'other']); + // HEAD is `other`; pushing `main` would check a working tree that is not main's. + const r = push(['origin', 'main']); + expect(r.status, r.stderr).toBe(0); + expect(r.stdout + r.stderr).toContain( + 'pre-push: skipping static checks: refs/heads/main is not the checked-out HEAD' + ); + expect(ran()).toEqual([]); + }); + + it('skips when any one of several pushed refs is not HEAD', () => { + const { ran, push, repo } = setup({ failing: 'lint' }); + git(repo, ['branch', 'behind']); + git(repo, ['commit', '-q', '--allow-empty', '-m', 'ahead']); + const r = push(['origin', 'main', 'behind']); + expect(r.status, r.stderr).toBe(0); + expect(r.stdout + r.stderr).toContain('is not the checked-out HEAD'); + expect(ran()).toEqual([]); + }); + + it('still checks an annotated tag that points at HEAD (the tag is peeled)', () => { + const { ran, push, repo } = setup(); + git(repo, ['tag', '-a', 'v1', '-m', 'v1']); + const r = push(['-q', 'origin', 'v1']); + expect(r.status, r.stderr + r.stdout).toBe(0); + expect(ran()).toEqual(expectedRuns); + }); + + it.each(['src/wip.ts', 'config/wip.json', 'scripts/wip.mjs', 'test/wip.test.ts', 'install.sh'])( + 'skips when %s is untracked (another session may own it)', + (rel) => { + const { ran, push, repo } = setup({ failing: 'lint' }); + mkdirSync(join(repo, rel, '..'), { recursive: true }); + writeFileSync(join(repo, rel), 'wip\n'); + const r = push(['origin', 'main']); + expect(r.status, r.stderr).toBe(0); + expect(r.stdout + r.stderr).toContain('pre-push: skipping static checks: uncommitted changes under'); + expect(ran()).toEqual([]); + } + ); + + it('skips when a tracked package.json has an unstaged edit', () => { + const { ran, push, repo } = setup({ failing: 'lint' }); + const pkg = join(repo, 'package.json'); + writeFileSync(pkg, readFileSync(pkg, 'utf8') + '\n'); + const r = push(['origin', 'main']); + expect(r.status, r.stderr).toBe(0); + expect(r.stdout + r.stderr).toContain('uncommitted changes under'); + expect(ran()).toEqual([]); + }); + + it('still checks when the only uncommitted changes are outside the watched paths', () => { + const { ran, push, repo } = setup({ failing: 'lint' }); + mkdirSync(join(repo, 'docs')); + writeFileSync(join(repo, 'docs', 'notes.md'), 'draft\n'); + writeFileSync(join(repo, 'README.md'), 'draft\n'); + const r = push(['origin', 'main']); + expect(r.status).not.toBe(0); + expect(r.stdout + r.stderr).toContain('pre-push: FAILED npm run lint'); + expect(ran()).toEqual(expectedRuns); + }); + + it('watches exactly the paths the checks read', () => { + expect(PRE_PUSH_WATCHED_PATHS).toEqual([ + 'src', + 'config', + 'scripts', + 'test', + 'package.json', + 'package-lock.json', + 'install.sh', + ]); + }); + it('skips (never blocks) when node_modules is absent', () => { const { ran, push } = setup({ failing: 'lint', nodeModules: false }); const r = push(['origin', 'main']); From e71971cab4afd9f8bf714199a148e68579547866 Mon Sep 17 00:00:00 2001 From: Aamer Akhter <aakhter@gmail.com> Date: Sat, 26 Sep 2026 22:45:35 -0400 Subject: [PATCH 18/46] fix(webview): route webview:changed only to its owner in multi-user mode webview:changed carried only {action, id} and the SSE routing hint had no webview: branch, so every connected client received it: in multi-user mode any user saw the ids of other users' web-tab creates, edits and deletes. The event now carries the web tab's owner (from the stored record) and is routed to that owner plus admins. Single-user delivery is unchanged. --- src/web/routes/webview-routes.ts | 18 ++- src/web/server.ts | 7 + src/web/sse-events.ts | 6 +- src/web/webview-sse.ts | 6 + test/webview-sse.test.ts | 217 +++++++++++++++++++++++++++++++ 5 files changed, 249 insertions(+), 5 deletions(-) create mode 100644 src/web/webview-sse.ts create mode 100644 test/webview-sse.test.ts diff --git a/src/web/routes/webview-routes.ts b/src/web/routes/webview-routes.ts index 16204ba1..99dc8420 100644 --- a/src/web/routes/webview-routes.ts +++ b/src/web/routes/webview-routes.ts @@ -165,7 +165,11 @@ function registerCrudRoutes(app: FastifyInstance, ctx: EventPort & TabLayoutPort }); throw error; } - ctx.broadcast(SseEvent.WebviewChanged, { action: 'created', id: created.id }); + ctx.broadcast(SseEvent.WebviewChanged, { + action: 'created', + id: created.id, + owner: ownerLayoutKey(created.owner), + }); return { success: true, data: created }; }); @@ -195,7 +199,11 @@ function registerCrudRoutes(app: FastifyInstance, ctx: EventPort & TabLayoutPort // Any edit invalidates the outstanding capability. Otherwise a token minted // against the OLD url keeps proxying to it after the user repointed the tab. webviewCapabilities.revokeWebview(id); - ctx.broadcast(SseEvent.WebviewChanged, { action: 'updated', id }); + ctx.broadcast(SseEvent.WebviewChanged, { + action: 'updated', + id, + owner: ownerLayoutKey(updated.owner), + }); return { success: true, data: updated }; }); @@ -231,7 +239,11 @@ function registerCrudRoutes(app: FastifyInstance, ctx: EventPort & TabLayoutPort webviewCapabilities.revokeWebview(id); socketCounts.delete(id); - ctx.broadcast(SseEvent.WebviewChanged, { action: 'deleted', id }); + ctx.broadcast(SseEvent.WebviewChanged, { + action: 'deleted', + id, + owner: ownerLayoutKey(result.owner), + }); return { success: true, data: { id } }; }); diff --git a/src/web/server.ts b/src/web/server.ts index c9de5895..d70577e6 100644 --- a/src/web/server.ts +++ b/src/web/server.ts @@ -96,6 +96,7 @@ import { PushSubscriptionStore } from '../push-store.js'; import webpush from 'web-push'; import { SseStreamManager } from './sse-stream-manager.js'; import { deriveTabLayoutSseHint } from './tab-layout-sse.js'; +import { deriveWebviewSseHint } from './webview-sse.js'; import { type SessionListenerRefs, createSessionListeners, @@ -2461,6 +2462,12 @@ export class WebServer extends EventEmitter { if (event.startsWith('tab:')) { return deriveTabLayoutSseHint(data); } + // Saved-webview invalidations carry the trusted resource owner. Route them to + // that owner (plus admins), so an admin editing a user's web tab notifies the + // user, and no other user learns the ids of someone else's web tabs. + if (event.startsWith('webview:')) { + return deriveWebviewSseHint(data); + } // Session-scoped families: resolve the owner from the payload's session id. const SESSION_PREFIXES = [ 'session:', diff --git a/src/web/sse-events.ts b/src/web/sse-events.ts index 6359ed4f..f18fbf6d 100644 --- a/src/web/sse-events.ts +++ b/src/web/sse-events.ts @@ -481,8 +481,10 @@ export const AuthPasswordChangeRequired = 'auth:passwordChangeRequired' as const export const SessionOrderChanged = 'session:orderChanged' as const; /** A saved web tab (dashboard URL) was created, updated or deleted. - * Payload: `{ action: 'created' | 'updated' | 'deleted', id }`. The client - * re-fetches the list rather than patching from the payload. */ + * Payload: `{ action: 'created' | 'updated' | 'deleted', id, owner }`. The client + * re-fetches the list rather than patching from the payload. `owner` is the web + * tab's owner (`'@single'` when multi-user mode is off); in multi-user mode the + * event is delivered only to that owner and admins (`deriveWebviewSseHint`). */ export const WebviewChanged = 'webview:changed' as const; /** Owner-scoped layout invalidation. Payload contains only `{ owner, version }`. */ export const TabLayoutChanged = 'tab:layoutChanged' as const; diff --git a/src/web/webview-sse.ts b/src/web/webview-sse.ts new file mode 100644 index 00000000..12707cc3 --- /dev/null +++ b/src/web/webview-sse.ts @@ -0,0 +1,6 @@ +/** @fileoverview Trusted owner routing metadata for saved-webview invalidations. */ +import type { SseRoutingHint } from './sse-stream-manager.js'; + +export function deriveWebviewSseHint(data: unknown): SseRoutingHint { + return { username: (data as { owner?: string }).owner, sessionScoped: true }; +} diff --git a/test/webview-sse.test.ts b/test/webview-sse.test.ts new file mode 100644 index 00000000..bcc47cf9 --- /dev/null +++ b/test/webview-sse.test.ts @@ -0,0 +1,217 @@ +/** + * @fileoverview Owner routing of the `webview:changed` SSE event. + * + * Saved web tabs are owner-scoped in multi-user mode (`canAccessOwned` on every + * CRUD route), but their invalidation event used to carry no owner and fell + * through `deriveSseHint`'s global branch, so every connected user learned the + * ids of every other user's web-tab creates, edits and deletes. The event now + * carries the resource owner and routes to that owner plus admins; single-user + * mode (no SSE identity) still delivers it to every client. + */ +import fs from 'node:fs/promises'; +import os from 'node:os'; +import path from 'node:path'; +import Fastify, { type FastifyInstance, type FastifyReply } from 'fastify'; +import fastifyCookie from '@fastify/cookie'; +import fastifyWebsocket from '@fastify/websocket'; +import { afterEach, beforeEach, describe, expect, it } from 'vitest'; +import { CleanupManager } from '../src/utils/index.js'; +import type { AuthUser } from '../src/types.js'; +import { installRouteErrorHandler } from '../src/web/route-error-handler.js'; +import { registerWebviewRoutes } from '../src/web/routes/webview-routes.js'; +import { WebServer } from '../src/web/server.js'; +import { SseEvent } from '../src/web/sse-events.js'; +import { SseStreamManager, type SseRoutingHint } from '../src/web/sse-stream-manager.js'; +import { deriveWebviewSseHint } from '../src/web/webview-sse.js'; + +function client() { + const writes: string[] = []; + return { + writes, + reply: { raw: { write: (chunk: string) => (writes.push(chunk), true) } } as unknown as FastifyReply, + }; +} + +/** The server's real event → routing-hint derivation, without starting the server. */ +function serverHint(event: string, data: unknown): SseRoutingHint | undefined { + const server = new WebServer(3999, false, true) as unknown as { + deriveSseHint(event: string, data: unknown): SseRoutingHint | undefined; + }; + return server.deriveSseHint(event, data); +} + +describe('deriveWebviewSseHint', () => { + it('routes to the exact owner and fails closed when the owner is missing', () => { + expect(deriveWebviewSseHint({ action: 'updated', id: 'w1', owner: 'alice' })).toEqual({ + username: 'alice', + sessionScoped: true, + }); + expect(deriveWebviewSseHint({ action: 'updated', id: 'w1' })).toEqual({ + username: undefined, + sessionScoped: true, + }); + }); + + it('is what the server derives for the webview: family (never the global branch)', () => { + const payload = { action: 'created', id: 'w1', owner: 'alice' }; + expect(serverHint(SseEvent.WebviewChanged, payload)).toEqual({ username: 'alice', sessionScoped: true }); + expect(serverHint(SseEvent.WebviewChanged, { action: 'deleted', id: 'w1' })).not.toBeUndefined(); + }); +}); + +describe('webview:changed delivery', () => { + it('reaches the owner and admins, never another ordinary user', () => { + const cleanup = new CleanupManager(); + const manager = new SseStreamManager({ getSessionStateWithRespawn: () => null }, cleanup); + const alice = client(); + const bob = client(); + const admin = client(); + manager.addClient(alice.reply, null, false, undefined, { username: 'alice', role: 'user' }); + manager.addClient(bob.reply, null, false, undefined, { username: 'bob', role: 'user' }); + manager.addClient(admin.reply, null, false, undefined, { username: 'root', role: 'admin' }); + const payload = { action: 'deleted', id: 'w1', owner: 'alice' }; + + manager.broadcast(SseEvent.WebviewChanged, payload, serverHint(SseEvent.WebviewChanged, payload)); + + expect(alice.writes).toEqual(['event: webview:changed\ndata: {"action":"deleted","id":"w1","owner":"alice"}\n\n']); + expect(admin.writes).toEqual(alice.writes); + expect(bob.writes).toEqual([]); + cleanup.dispose(); + }); + + it('single-user mode: clients without an identity all still receive it', () => { + const cleanup = new CleanupManager(); + const manager = new SseStreamManager({ getSessionStateWithRespawn: () => null }, cleanup); + const tabA = client(); + const tabB = client(); + manager.addClient(tabA.reply, null, false, undefined, undefined); + manager.addClient(tabB.reply, null, false, undefined, undefined); + const payload = { action: 'created', id: 'w1', owner: '@single' }; + + manager.broadcast(SseEvent.WebviewChanged, payload, serverHint(SseEvent.WebviewChanged, payload)); + + expect(tabA.writes).toHaveLength(1); + expect(tabB.writes).toEqual(tabA.writes); + cleanup.dispose(); + }); +}); + +describe('webview routes → SSE, end to end', () => { + let tmpDir: string; + let savedDataDir: string | undefined; + let savedMode: string | undefined; + let cleanup: CleanupManager; + let manager: SseStreamManager; + const apps: FastifyInstance[] = []; + + beforeEach(async () => { + tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), 'codeman-webview-sse-')); + savedDataDir = process.env.CODEMAN_DATA_DIR; + savedMode = process.env.CODEMAN_MULTIUSER; + process.env.CODEMAN_DATA_DIR = tmpDir; + cleanup = new CleanupManager(); + manager = new SseStreamManager({ getSessionStateWithRespawn: () => null }, cleanup); + }); + + afterEach(async () => { + for (const app of apps.splice(0)) await app.close(); + cleanup.dispose(); + if (savedDataDir === undefined) delete process.env.CODEMAN_DATA_DIR; + else process.env.CODEMAN_DATA_DIR = savedDataDir; + if (savedMode === undefined) delete process.env.CODEMAN_MULTIUSER; + else process.env.CODEMAN_MULTIUSER = savedMode; + await fs.rm(tmpDir, { recursive: true, force: true }).catch(() => {}); + }); + + /** A route app acting as `authUser`, whose broadcasts go through the server's routing. */ + async function appAs(authUser: AuthUser | undefined): Promise<FastifyInstance> { + const app = Fastify({ logger: false }); + await app.register(fastifyCookie); + await app.register(fastifyWebsocket); + app.decorateRequest('authUser', undefined); + app.addHook('onRequest', async (req) => { + req.authUser = authUser; + }); + registerWebviewRoutes(app, { + broadcast: (event: string, data: unknown) => manager.broadcast(event, data, serverHint(event, data)), + tabLayouts: { webviewCreated: async () => {}, webviewDeleted: async () => {} }, + } as never); + installRouteErrorHandler(app); + await app.ready(); + apps.push(app); + return app; + } + + it("multi-user: another user's SSE stream never sees a web-tab create, edit or delete", async () => { + process.env.CODEMAN_MULTIUSER = '1'; + const aliceApp = await appAs({ username: 'alice', role: 'user' }); + const alice = client(); + const bob = client(); + const admin = client(); + manager.addClient(alice.reply, null, false, undefined, { username: 'alice', role: 'user' }); + manager.addClient(bob.reply, null, false, undefined, { username: 'bob', role: 'user' }); + manager.addClient(admin.reply, null, false, undefined, { username: 'root', role: 'admin' }); + + const created = await aliceApp.inject({ + method: 'POST', + url: '/api/webviews', + payload: { name: 'Grafana', url: 'http://127.0.0.1:4000/' }, + }); + expect(created.statusCode).toBe(200); + const id = created.json().data.id as string; + const patched = await aliceApp.inject({ method: 'PATCH', url: `/api/webviews/${id}`, payload: { name: 'G2' } }); + expect(patched.statusCode).toBe(200); + expect((await aliceApp.inject({ method: 'DELETE', url: `/api/webviews/${id}` })).statusCode).toBe(200); + + const expected = ['created', 'updated', 'deleted'].map( + (action) => `event: webview:changed\ndata: ${JSON.stringify({ action, id, owner: 'alice' })}\n\n` + ); + expect(alice.writes).toEqual(expected); + expect(admin.writes).toEqual(expected); + expect(bob.writes).toEqual([]); + }); + + it("multi-user: an admin editing a user's web tab notifies that user, not a bystander", async () => { + process.env.CODEMAN_MULTIUSER = '1'; + const aliceApp = await appAs({ username: 'alice', role: 'user' }); + const adminApp = await appAs({ username: 'root', role: 'admin' }); + const id = ( + await aliceApp.inject({ + method: 'POST', + url: '/api/webviews', + payload: { name: 'G', url: 'http://127.0.0.1:4000/' }, + }) + ).json().data.id as string; + const alice = client(); + const bob = client(); + manager.addClient(alice.reply, null, false, undefined, { username: 'alice', role: 'user' }); + manager.addClient(bob.reply, null, false, undefined, { username: 'bob', role: 'user' }); + + expect((await adminApp.inject({ method: 'DELETE', url: `/api/webviews/${id}` })).statusCode).toBe(200); + + expect(alice.writes).toEqual([ + `event: webview:changed\ndata: ${JSON.stringify({ action: 'deleted', id, owner: 'alice' })}\n\n`, + ]); + expect(bob.writes).toEqual([]); + }); + + it('single-user: every client still receives the event', async () => { + const soloApp = await appAs(undefined); + const tabA = client(); + const tabB = client(); + manager.addClient(tabA.reply, null, false, undefined, undefined); + manager.addClient(tabB.reply, null, false, undefined, undefined); + + const created = await soloApp.inject({ + method: 'POST', + url: '/api/webviews', + payload: { name: 'G', url: 'http://127.0.0.1:4000/' }, + }); + expect(created.statusCode).toBe(200); + + expect(tabA.writes).toEqual([ + `event: webview:changed\ndata: ${JSON.stringify({ action: 'created', id: created.json().data.id, owner: '@single' })}\n\n`, + ]); + expect(tabB.writes).toEqual(tabA.writes); + }); +}); From 5e27043bf730896c16a72054041f00d87ac0e897 Mon Sep 17 00:00:00 2001 From: JD <jd@jds.haus> Date: Mon, 28 Sep 2026 01:30:05 -0400 Subject: [PATCH 19/46] feat(files): render markdown in the File Viewer, with Lines/Wrap toggles Clicking a .md in the Files panel showed wrapped source with an Edit pencil and no way to see it rendered, although marked + DOMPurify were already on the page for the Response Viewer. The viewer now renders .md/.markdown through that same pipeline (one parser, one click delegate) with an MD pill back to source, and the plain-text view gains Lines (CSS-counter gutter) and Wrap toggles. All three persist per device in their own localStorage keys. - Relative images are rebased onto the workspace-confined file-raw route under the document's directory, built inside a <template> so no fetch fires before the rewrite; a failed load degrades to alt text. Relative links become a.rv-path so the existing delegate opens them in the viewer; fragment and http(s) links are untouched. - The rendered container carries data-i18n-skip so the translator does not rewrite the document's prose. - Markdown fetches the route's 10000-line ceiling; other text keeps 500. - avif renders inline (file-content image set, file-raw MIME map), and avif/ico printed paths open the viewer instead of tailing bytes. .md deliberately stays with the tail viewer for printed paths. --- CLAUDE.md | 2 + docs/architecture-invariants.md | 13 ++ docs/wiki/Working-With-Files.md | 3 +- src/web/public/constants.js | 4 +- src/web/public/index.html | 3 + src/web/public/panels-ui.js | 211 ++++++++++++++++++++- src/web/public/styles.css | 65 +++++++ src/web/routes/file-routes.ts | 3 +- test/file-preview-markdown.test.ts | 292 +++++++++++++++++++++++++++++ test/routes/file-routes.test.ts | 25 +++ 10 files changed, 609 insertions(+), 12 deletions(-) create mode 100644 test/file-preview-markdown.test.ts diff --git a/CLAUDE.md b/CLAUDE.md index c4a910cc..9dd64299 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -290,6 +290,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md` +**File Viewer text view: rendered markdown + Lines/Wrap toggles** (`_renderFilePreviewText()` in panels-ui.js): a `.md`/`.markdown` opens RENDERED by default with an `MD` pill back to source; the plain-text view has `Lines` (CSS-counter gutter) and `Wrap` toggles. ⚠️ ONE markdown pipeline: the viewer calls `_renderMarkdown()` (marked + the DOMPurify allowlist, the Response Viewer's) and binds the Response Viewer's click delegate (`_bindResponseViewerInteractions`) on the preview body for code-copy buttons and path links; never a second parser or handler. ⚠️ The document is built inside a `<template>` (a detached div with `innerHTML` set starts fetching every `<img src>` before the rewrite), then `_rebaseFilePreviewMarkdownRefs()` points relative images at the workspace-confined `file-raw` under the document's directory (never a widened route; a failed load degrades to alt text) and turns relative links into `a.rv-path` for the delegate, stripping the `target` marked gave them. ⚠️ The container carries `data-i18n-skip` or the translator rewrites the document's prose. ⚠️ Toggles are per-device localStorage keys (`codeman:filePreview*`), never `SettingsUpdateSchema`; Lines/Wrap are class flips on the ONE `<pre>`, with rules scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's code blocks. Markdown fetches `lines=10000` (the route ceiling), other text keeps 500. ⚠️ `md` stays OUT of `FILE_PREVIEW_EXTENSIONS`: a printed `.md` path keeps the tail viewer (live follow); the rendered view is the Files panel's. Tests: `test/file-preview-markdown.test.ts`. → [architecture-invariants#file-viewer-text-view-rendered-markdown-and-text-toggles](docs/architecture-invariants.md#file-viewer-text-view-rendered-markdown-and-text-toggles) + **Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure) compiles the query into a predicate the server-side walk prunes with; a query returns a FLAT match list and the walk recurses past non-matching directories. An empty, whitespace-only or overlong (`MAX_QUERY_LENGTH`, 256) query compiles to `null`, keeping the default tree response byte-identical. ⚠️ **Never compile a glob into a RegExp** (`*a*a*a…` backtracks and freezes the event loop for the whole server): `globMatch()` is a two-pointer wildcard walk. → [architecture-invariants#files-panel-search](docs/architecture-invariants.md#files-panel-search) **Raw file bodies are streamed and range-aware**: `file-raw`, the attachments `/raw` route and `GET /api/download` share `sendFileBody()`, advertise `Accept-Ranges: bytes` and answer `Range` with `206` + `Content-Range` (single-range, parser in `src/web/http-range.ts`); without it `<video>` cannot seek. The size cap (`MAX_FILE_DOWNLOAD_BYTES`, default 2GB, env `CODEMAN_MAX_DOWNLOAD_BYTES`, `0` = unlimited) is a sanity bound, not memory protection; never reintroduce a whole-file buffer. ⚠️ Bodies go out via `reply.hijack()`, so `sendRawStream` must copy the status onto `reply.raw` by hand or a partial body ships as `200`. ⚠️ Closing the preview must pause and unload media (`_stopFilePreviewMedia`), since a detached `HTMLMediaElement` keeps playing. → [architecture-invariants#raw-file-bodies-streamed-and-range-aware](docs/architecture-invariants.md#raw-file-bodies-streamed-and-range-aware) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 6a25c5b3..d9f47f70 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -384,6 +384,19 @@ The general rule: **any new endpoint that turns a caller-supplied `sessionId` in Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write-routes.test.ts` (deliberately **unmocked fs** against a real temp workspace — symlink/TOCTOU/mode behavior must be exercised for real). +### File Viewer text view: rendered markdown and text toggles + +**Rendered markdown + Lines/Wrap** (`_renderFilePreviewText()` and its helpers in `panels-ui.js`, buttons in the `.file-preview-actions` row): a `.md`/`.markdown` opened in the File Viewer renders as a document by default, with an `MD` pill back to source; the plain-text view has `Lines` (a CSS-counter gutter) and `Wrap` toggles. Codeman already had `marked` + DOMPurify behind `_renderMarkdown()` for the Response Viewer, so the viewer reuses that and the codebase keeps ONE markdown pipeline. + +- ⚠️ **One pipeline, one delegate.** The viewer calls `_renderMarkdown()` (marked + the `sanitize-html.js` allowlist) and binds `_bindResponseViewerInteractions()` on `#filePreviewBody` (container-bound and idempotent, so once per page) for the code-block copy buttons and `a.rv-path` opening. Never a second parser, never a second click handler for the same markup. +- ⚠️ **Build inside a `<template>`, then rebase.** A detached div whose `innerHTML` is set starts fetching every `<img src>` at once, so the document's relative image paths would hit the server as `/docs/img.png` 404s before being rewritten. `_rebaseFilePreviewMarkdownRefs()` runs on the template content: relative images go to the workspace-confined `file-raw` under the document's directory (the server refuses escapes, so `..` is forwarded as-is), and one `error` handler per image degrades it to alt text, which covers a remote image the page CSP blocks, a 404 for a document outside the workspace, and an SVG that `file-raw` serves as a download. Never widen a route for this. Relative links become `a.rv-path` with `data-path` and lose the `target`/`rel` that `_renderMarkdown` gives every link, which would otherwise open `<origin>/docs/x.md` in a new tab; fragment and http(s) links are untouched. +- ⚠️ **`data-i18n-skip` on the container.** The translator's MutationObserver translates inserted headings and paragraphs, and the `.file-preview-content` entry in its skip list matches nothing (no element has that class), so the attribute is what keeps a Chinese UI from rewriting a README. +- ⚠️ **Toggles are per-device, in their own localStorage keys** (`codeman:filePreviewMdRendered` / `LineNumbers` / `Wrap`), for the same reason as the Files panel's show-hidden toggle: the app-settings object is rebuilt from the settings modal on every save, and they are display state, not synced settings (`SettingsUpdateSchema` is `.strict()`). MD re-renders from the kept source (`filePreviewContent`, which is also what Copy copies) without a refetch; Lines/Wrap are class flips on the one `<pre>`, whose rules are scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's own code blocks. Lines are one inline `<span class="fp-line">` per line joined by real newlines, the counter in `::before` with `user-select: none`, so select and copy return the exact text. +- ⚠️ **Caps.** Markdown fetches `lines=10000` (the route's `MAX_LINES_LIMIT`), because a rendered document cut at 500 lines reads as the whole document; other text keeps 500, which is what stops a huge log locking the tab in one `<pre>`. The attachment (out-of-workspace) branch keeps its 512 KB Range read and skips the 500-line clip for markdown. Edit mode is unchanged and still re-fetches `edit=1`; the three toggles hide while editing and for images, media and PDFs. +- ⚠️ **`md` is NOT in `FILE_PREVIEW_EXTENSIONS`.** A `.md` path printed in the terminal or chat still opens the tail viewer (see File-path links above: in-workspace text keeps live follow, which is what Ralph's `fix_plan.md` needs); the rendered view is reached from the Files panel. `avif` and `ico` were added there (a printed `favicon.ico` used to tail binary noise), with `avif` also in file-content's image set and file-raw's MIME map; out-of-workspace avif/ico stay unregistrable, like svg/bmp. + +Tests: `test/file-preview-markdown.test.ts` (jsdom-in-vm, pins every rule above), `test/routes/file-routes.test.ts` (avif). + ### Clone a repository as a case **Clone Repo tab** (issue #236, proposed by @DodgyBadger): `POST /api/cases/clone` clones a public repository into the caller's case space and registers it as a normal local case; `POST /api/cases/clone-preflight` answers "can this be cloned anonymously, and what refs does it have?" while the user is still typing. Core in `src/git-clone.ts`, split into a PURE half (URL parse, argv/env, `ls-remote` parse, stderr classification) and a thin IO half (`probeGitRemote`, `cloneRepository`). diff --git a/docs/wiki/Working-With-Files.md b/docs/wiki/Working-With-Files.md index d9ebaf48..d8c8b357 100644 --- a/docs/wiki/Working-With-Files.md +++ b/docs/wiki/Working-With-Files.md @@ -13,7 +13,8 @@ It renders what it can: | Kind | Behaviour | | ------------------------ | ------------------------------------------------------------------------- | -| Text and code | Syntax-aware preview. Long files are truncated in plain preview. | +| Text and code | Plain preview with Lines (line numbers) and Wrap toggles in the header. Long files are truncated in plain preview. | +| Markdown | Rendered by default: headings, tables, code blocks with copy buttons, images and links relative to the file. The MD pill in the header flips to source. | | Images | Inline. | | Audio and video | Inline with a working scrub bar, because range requests are supported. | | PDF and Office documents | Converted for preview when a converter is available. | diff --git a/src/web/public/constants.js b/src/web/public/constants.js index 6f9d85c8..6a0ff942 100644 --- a/src/web/public/constants.js +++ b/src/web/public/constants.js @@ -1458,7 +1458,7 @@ function computeRewriteScrollLine(input) { * a `/g` regex, so {@link absoluteFilePathPattern} mints a fresh one per call. */ const FILE_PATH_LINK_PATTERN = - /(\/(?:home|Users|tmp|var|private|opt|mnt|srv|media|data|workspace)\/[^\s"'<>|;&\n\x00-\x1f]*\.(?:log|txt|json|md|ya?ml|csv|xml|sh|py|tsx|ts|jsx|js|mjs|cjs|css|html|toml|ini|sql|png|jpe?g|gif|webp|bmp|svg|pdf|docx|pptx|mp4|webm|mov|mp3|wav))\b/g; + /(\/(?:home|Users|tmp|var|private|opt|mnt|srv|media|data|workspace)\/[^\s"'<>|;&\n\x00-\x1f]*\.(?:log|txt|json|md|ya?ml|csv|xml|sh|py|tsx|ts|jsx|js|mjs|cjs|css|html|toml|ini|sql|png|jpe?g|gif|webp|avif|bmp|ico|svg|pdf|docx|pptx|mp4|webm|mov|mp3|wav))\b/g; /** A fresh, zero-state instance of {@link FILE_PATH_LINK_PATTERN}. */ function absoluteFilePathPattern() { @@ -1476,7 +1476,7 @@ function absoluteFilePathPattern() { * file in /tmp played fine. test/media-extension-parity.test.ts pins the sync. */ const FILE_PREVIEW_EXTENSIONS = new Set( - ('png jpg jpeg gif webp bmp svg pdf docx pptx mp4 webm mov m4v ogv mp3 wav ogg oga m4a aac flac opus').split(' ') + ('png jpg jpeg gif webp avif bmp ico svg pdf docx pptx mp4 webm mov m4v ogv mp3 wav ogg oga m4a aac flac opus').split(' ') ); /** Whether a path's extension is one {@link FILE_PREVIEW_EXTENSIONS} covers. */ diff --git a/src/web/public/index.html b/src/web/public/index.html index a5009b53..40d19750 100644 --- a/src/web/public/index.html +++ b/src/web/public/index.html @@ -573,6 +573,9 @@ <div class="file-preview-header"> <span class="file-preview-title" id="filePreviewTitle">file.ts</span> <div class="file-preview-actions"> + <button class="btn-icon-sm file-preview-pill" id="filePreviewMdBtn" onclick="app.toggleFilePreviewMd()" title="Rendered markdown" aria-label="Rendered markdown" aria-pressed="true" hidden>MD</button> + <button class="btn-icon-sm" id="filePreviewLinesBtn" onclick="app.toggleFilePreviewLines()" title="Line numbers" aria-label="Line numbers" aria-pressed="false" hidden><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><line x1="10" y1="6" x2="21" y2="6"/><line x1="10" y1="12" x2="21" y2="12"/><line x1="10" y1="18" x2="21" y2="18"/><path d="M4 6h1v4"/><path d="M4 10h2"/><path d="M6 18H4c0-1 2-2 2-3s-1-1.5-2-1"/></svg></button> + <button class="btn-icon-sm" id="filePreviewWrapBtn" onclick="app.toggleFilePreviewWrap()" title="Wrap lines" aria-label="Wrap lines" aria-pressed="true" hidden><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><line x1="3" y1="6" x2="21" y2="6"/><path d="M3 12h15a3 3 0 1 1 0 6h-4"/><polyline points="16 16 14 18 16 20"/><line x1="3" y1="18" x2="10" y2="18"/></svg></button> <button class="btn-icon-sm file-preview-edit-btn" id="filePreviewEditBtn" onclick="app.enterFilePreviewEdit()" title="Edit file" aria-label="Edit file" hidden><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M17 3a2.85 2.83 0 1 1 4 4L7.5 20.5 2 22l1.5-5.5z"/></svg></button> <button class="btn-icon-sm" onclick="app.copyFilePreviewContent()" title="Copy content">⎘</button> <button class="btn-icon-sm" id="filePreviewDetachBtn" onclick="app.detachFilePreview()" title="Open in new tab" aria-label="Open in new tab" hidden><svg width="13" height="13" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M18 13v6a2 2 0 0 1-2 2H5a2 2 0 0 1-2-2V8a2 2 0 0 1 2-2h6"/><polyline points="15 3 21 3 21 9"/><line x1="10" y1="14" x2="21" y2="3"/></svg></button> diff --git a/src/web/public/panels-ui.js b/src/web/public/panels-ui.js index 8bdf5df8..c8949d75 100644 --- a/src/web/public/panels-ui.js +++ b/src/web/public/panels-ui.js @@ -20,6 +20,20 @@ const FILE_BROWSER_SHOW_HIDDEN_KEY = 'codeman:fileBrowserShowHidden'; // a huge log is a partial read rather than a download the viewer throws away. const TEXT_PREVIEW_MAX_BYTES = 512 * 1024; const TEXT_PREVIEW_MAX_LINES = 500; +// A markdown document gets the route's ceiling instead of the 500-line preview +// cap: a rendered README cut mid-way reads as the whole document. +const MARKDOWN_PREVIEW_MAX_LINES = 10000; +const MARKDOWN_EXTS = new Set(['md', 'markdown']); +// File Viewer text-view prefs: per-device, in their own localStorage keys for +// the same reason as FILE_BROWSER_SHOW_HIDDEN_KEY (the app-settings object is +// rebuilt from the settings modal on save, so a key toggled from the viewer +// would be dropped on the next save). +const FILE_PREVIEW_PREF_KEYS = { + mdRendered: 'codeman:filePreviewMdRendered', + lineNumbers: 'codeman:filePreviewLineNumbers', + wrap: 'codeman:filePreviewWrap', +}; +const FILE_PREVIEW_PREF_DEFAULTS = { mdRendered: true, lineNumbers: false, wrap: true }; const AWAY_DIGEST_SECTIONS = [ ['needsAttention', 'Needs Attention'], ['completed', 'Completed'], @@ -4012,6 +4026,10 @@ Object.assign(CodemanApp.prototype, { this.filePreviewDetachUrl = ''; const detachBtn = this.$('filePreviewDetachBtn'); if (detachBtn) detachBtn.hidden = true; + // Same for the text-view toggles: they act on the text this load has not + // fetched yet, and an image or PDF has nothing for them to toggle. + this.filePreviewText = null; + this._updateFilePreviewToolbar('none'); // Show overlay with loading state overlay.classList.add('visible'); @@ -4095,12 +4113,17 @@ Object.assign(CodemanApp.prototype, { const text = await res.text(); const clippedByBytes = res.status === 206 && text.length >= TEXT_PREVIEW_MAX_BYTES; const lines = text.split('\n'); - const clippedByLines = lines.length > TEXT_PREVIEW_MAX_LINES; - const shown = clippedByLines ? lines.slice(0, TEXT_PREVIEW_MAX_LINES).join('\n') : text; - bodyEl.innerHTML = `<pre><code>${escapeHtml(shown)}</code></pre>`; + // Markdown keeps every line the Range read returned: the byte bound is + // what protects the tab, and a rendered document cut at 500 lines + // reads as the whole document. + const lineCap = MARKDOWN_EXTS.has(ext) ? Infinity : TEXT_PREVIEW_MAX_LINES; + const clippedByLines = lines.length > lineCap; + const shown = clippedByLines ? lines.slice(0, lineCap).join('\n') : text; this.filePreviewContent = shown; + this.filePreviewText = { ext, sessionId, filePath }; + this._renderFilePreviewText(); if (clippedByLines || clippedByBytes) { - const note = clippedByLines ? `showing first ${TEXT_PREVIEW_MAX_LINES} lines` : 'showing the start of the file'; + const note = clippedByLines ? `showing first ${lineCap} lines` : 'showing the start of the file'; footerEl.textContent = `${footerEl.textContent} (${note})`; } } catch (err) { @@ -4146,8 +4169,13 @@ Object.assign(CodemanApp.prototype, { return; } + // 500 lines is what keeps a huge log from locking the tab in one <pre>; + // markdown is rendered as a document and takes the route's ceiling instead. + const lineCap = MARKDOWN_EXTS.has(ext) ? MARKDOWN_PREVIEW_MAX_LINES : TEXT_PREVIEW_MAX_LINES; try { - const res = await fetch(`/api/sessions/${sessionId}/file-content?path=${encodeURIComponent(filePath)}&lines=500`); + const res = await fetch( + `/api/sessions/${sessionId}/file-content?path=${encodeURIComponent(filePath)}&lines=${lineCap}` + ); if (!res.ok) throw new Error('Failed to load file'); const result = await res.json(); @@ -4172,10 +4200,11 @@ Object.assign(CodemanApp.prototype, { bodyEl.innerHTML = `<div class="binary-message">Binary file (${this.formatFileSize(data.size)})<br>Cannot preview<br><a href="${escapeHtml(downloadHref)}" download>Download</a></div>`; footerEl.textContent = data.extension || 'binary'; } else { - // Text content + // Text content: rendered markdown or plain text, per the viewer's toggles. this.filePreviewContent = data.content; - bodyEl.innerHTML = `<pre><code>${escapeHtml(data.content)}</code></pre>`; - const truncNote = data.truncated ? ` (showing 500/${data.totalLines} lines)` : ''; + this.filePreviewText = { ext, sessionId, filePath }; + this._renderFilePreviewText(); + const truncNote = data.truncated ? ` (showing ${lineCap}/${data.totalLines} lines)` : ''; footerEl.textContent = `${data.totalLines} lines \u2022 ${this.formatFileSize(data.size)}${truncNote}`; // Edit affordance only when the server says an edit=1 re-fetch would // succeed (workspace text file inside the allowlist and size cap). @@ -4203,6 +4232,8 @@ Object.assign(CodemanApp.prototype, { // audible and keeps streaming from the server. Closing has to stop it. this._stopFilePreviewMedia(); this.filePreviewContent = ''; + this.filePreviewText = null; + this._updateFilePreviewToolbar('none'); this.filePreviewDetachUrl = ''; const detachBtn = this.$('filePreviewDetachBtn'); if (detachBtn) detachBtn.hidden = true; @@ -4251,6 +4282,168 @@ Object.assign(CodemanApp.prototype, { bodyEl.innerHTML = ''; }, + // ═══════════════════════════════════════════════════════════════ + // File Viewer text view: rendered markdown, line numbers, wrap + // ═══════════════════════════════════════════════════════════════ + + _filePreviewPref(name) { + try { + const stored = localStorage.getItem(FILE_PREVIEW_PREF_KEYS[name]); + if (stored === '1') return true; + if (stored === '0') return false; + } catch { + /* private mode: fall through to the default */ + } + return FILE_PREVIEW_PREF_DEFAULTS[name]; + }, + + _setFilePreviewPref(name, on) { + try { + localStorage.setItem(FILE_PREVIEW_PREF_KEYS[name], on ? '1' : '0'); + } catch { + /* private mode: the toggle still applies for this page load */ + } + }, + + /** + * Paint the loaded text (filePreviewContent) into the preview body: a + * rendered document for .md/.markdown while the MD toggle is on, otherwise + * plain text with one span per line so the Lines toggle can number them. + * The MD toggle re-runs this without a refetch. + * + * Markdown goes through the same pipeline as the Response Viewer + * (`_renderMarkdown`: marked + the DOMPurify allowlist), never a second + * parser, and is built inside a <template>: a detached div with innerHTML + * already set starts fetching every <img src>, so the document's relative + * image paths would hit the server as /docs/img.png 404s before + * `_rebaseFilePreviewMarkdownRefs` rewrote them. + */ + _renderFilePreviewText() { + const info = this.filePreviewText; + const bodyEl = this.$('filePreviewBody'); + if (!info || !bodyEl) return; + const isMarkdown = MARKDOWN_EXTS.has(info.ext); + const rendered = isMarkdown && this._filePreviewPref('mdRendered'); + if (rendered) { + // data-i18n-skip: the translator's MutationObserver would otherwise + // rewrite the document's own headings and paragraphs. + const tmpl = document.createElement('template'); + tmpl.innerHTML = `<div class="rv-text file-preview-md" data-i18n-skip>${this._renderMarkdown(this.filePreviewContent)}</div>`; + const doc = tmpl.content.firstElementChild; + this._rebaseFilePreviewMarkdownRefs(doc, info); + this._linkifyFilePaths(doc); + bodyEl.replaceChildren(tmpl.content); + // The Response Viewer's click delegate (path links, code-block copy + // buttons, loopback links): container-bound and idempotent, so binding it + // on the body once serves every preview. + this._bindResponseViewerInteractions(bodyEl); + } else { + const pre = document.createElement('pre'); + pre.className = 'file-preview-text'; + pre.classList.toggle('wrap', this._filePreviewPref('wrap')); + pre.classList.toggle('show-lines', this._filePreviewPref('lineNumbers')); + const code = document.createElement('code'); + // One span per line joined by real newlines: empty lines survive, select + // and copy return the exact text, and the gutter counter hangs off the + // spans' ::before so the numbers are never part of the text. + code.innerHTML = this.filePreviewContent + .split('\n') + .map((line) => `<span class="fp-line">${escapeHtml(line)}</span>`) + .join('\n'); + pre.appendChild(code); + bodyEl.replaceChildren(pre); + } + this._updateFilePreviewToolbar(rendered ? 'markdown' : 'text'); + }, + + /** + * Point a rendered document's relative references at the file it came from. + * + * Images are rebased onto the workspace-confined file-raw route under the + * document's directory (the server refuses escapes, so `..` is safe to + * forward). Whatever fails to load degrades to its alt text with one error + * handler: a remote image the page CSP blocks, a 404 for a document outside + * the workspace, an SVG that file-raw serves as a download. Relative links + * take the `a.rv-path` shape the Response Viewer delegate already opens in + * this overlay, minus the target/rel `_renderMarkdown` gave them, which + * would otherwise open <origin>/docs/x.md in a new tab. + */ + _rebaseFilePreviewMarkdownRefs(root, { sessionId, filePath }) { + const dir = filePath.includes('/') ? filePath.slice(0, filePath.lastIndexOf('/') + 1) : ''; + // Relative = no scheme, not root-relative (which includes //host), not a fragment. + const isRelative = (ref) => !!ref && !/^[a-z][a-z0-9+.-]*:/i.test(ref) && !ref.startsWith('/') && !ref.startsWith('#'); + // GitHub-style `img.png#gh-dark-mode-only` and `doc.md#section`: the + // fragment is not part of the path. `.` and `..` segments are collapsed so + // the title reads `README.md`, not `docs/../README.md`; a `..` that climbs + // past the start is kept and left for the server to refuse. + const resolveRef = (ref) => { + const parts = []; + for (const seg of (dir + ref.split('#')[0]).split('/')) { + if (seg === '.' || (seg === '' && parts.length)) continue; + if (seg === '..' && parts.length && parts[parts.length - 1] !== '..' && parts[parts.length - 1] !== '') parts.pop(); + else parts.push(seg); + } + return parts.join('/'); + }; + for (const img of root.querySelectorAll('img[src]')) { + const src = img.getAttribute('src') || ''; + if (isRelative(src)) { + const path = resolveRef(src); + img.setAttribute('src', CodemanBase.url(`/api/sessions/${sessionId}/file-raw?path=${encodeURIComponent(path)}`)); + } + img.addEventListener('error', () => img.replaceWith(img.getAttribute('alt') || src), { once: true }); + } + for (const a of root.querySelectorAll('a[href]')) { + const href = a.getAttribute('href') || ''; + if (!isRelative(href)) continue; + a.className = 'rv-path'; + a.dataset.path = resolveRef(href); + a.setAttribute('href', '#'); + a.removeAttribute('target'); + a.removeAttribute('rel'); + } + }, + + /** + * Show the toggles that apply to the current view: MD for a markdown file in + * either view, Lines/Wrap for the plain-text view only; `'none'` (loading, + * image, media, PDF, edit mode) hides all three. + */ + _updateFilePreviewToolbar(view) { + const set = (id, shown, pressed) => { + const btn = this.$(id); + if (!btn) return; + btn.hidden = !shown; + if (shown) btn.setAttribute('aria-pressed', String(pressed)); + }; + const isMarkdown = view !== 'none' && MARKDOWN_EXTS.has(this.filePreviewText?.ext || ''); + set('filePreviewMdBtn', isMarkdown, view === 'markdown'); + set('filePreviewLinesBtn', view === 'text', this._filePreviewPref('lineNumbers')); + set('filePreviewWrapBtn', view === 'text', this._filePreviewPref('wrap')); + }, + + toggleFilePreviewMd() { + this._setFilePreviewPref('mdRendered', !this._filePreviewPref('mdRendered')); + this._renderFilePreviewText(); + }, + + toggleFilePreviewLines() { + this._toggleFilePreviewTextClass('lineNumbers', 'show-lines'); + }, + + toggleFilePreviewWrap() { + this._toggleFilePreviewTextClass('wrap', 'wrap'); + }, + + /** Lines and Wrap are pure class flips on the <pre>; no re-render needed. */ + _toggleFilePreviewTextClass(pref, className) { + const on = !this._filePreviewPref(pref); + this._setFilePreviewPref(pref, on); + const pre = this.$('filePreviewBody')?.querySelector(':scope > pre.file-preview-text'); + if (pre) pre.classList.toggle(className, on); + this._updateFilePreviewToolbar('text'); + }, + // ═══════════════════════════════════════════════════════════════ // File Viewer edit mode (issue #212 — docs/file-viewer-edit-plan.md) // ═══════════════════════════════════════════════════════════════ @@ -4318,6 +4511,8 @@ Object.assign(CodemanApp.prototype, { textarea.addEventListener('input', () => this._onFilePreviewEditInput()); bodyEl.innerHTML = ''; bodyEl.appendChild(textarea); + // The MD/Lines/Wrap toggles act on the text view this textarea replaced. + this._updateFilePreviewToolbar('none'); // Deliberately no autofocus: on phones that would pop the OS keyboard // before the user has scrolled to the line they want to change. diff --git a/src/web/public/styles.css b/src/web/public/styles.css index 050671a4..7e530ce2 100644 --- a/src/web/public/styles.css +++ b/src/web/public/styles.css @@ -10984,6 +10984,19 @@ kbd { .file-preview-actions { display: flex; gap: 0.25rem; + /* The title yields on a phone, not the buttons. */ + flex-shrink: 0; +} + +.file-preview-actions .btn-icon-sm[aria-pressed='true'] { + color: var(--accent, #4ea1ff); + background: var(--bg-hover, rgba(255, 255, 255, 0.08)); +} + +.file-preview-actions .file-preview-pill { + font-size: 0.7rem; + font-weight: 600; + letter-spacing: 0.02em; } .file-preview-body { @@ -11008,6 +11021,58 @@ kbd { font-family: inherit; } +/* ---- File Viewer text view: Lines / Wrap toggles ---- + Child-combinator scoped so none of this leaks into the code blocks of a + rendered markdown document, which are <pre>s too. */ + +.file-preview-body > pre.file-preview-text { + counter-reset: fp-line; + tab-size: 4; +} + +.file-preview-body > pre.file-preview-text:not(.wrap) { + white-space: pre; + word-break: normal; + overflow-x: auto; +} + +/* Numbers hug the left edge (4px, left-aligned) instead of sitting behind the + pre's own padding right-aligned in a 4ch column, where "1" landed 40px in. */ +.file-preview-body > pre.file-preview-text.show-lines { + padding-left: 4px; +} + +.file-preview-body > pre.file-preview-text.show-lines .fp-line::before { + counter-increment: fp-line; + content: counter(fp-line); + display: inline-block; + min-width: 4ch; + margin-right: 1ch; + text-align: left; + color: var(--text-muted); + user-select: none; +} + +/* ---- File Viewer rendered markdown ---- + Styling comes from the Response Viewer's .rv-text rules; .rv-text itself + carries no padding or base font (the chat card supplies those). */ + +.file-preview-body > .file-preview-md { + padding: 1rem 1.25rem 2rem; + font-size: 15px; + line-height: 1.55; + max-width: 960px; +} + +/* Relative links in the document are rebased onto a.rv-path so the Response + Viewer delegate opens them here; keep them reading as prose, not as paths. + Three-class selector on purpose: `.rv-text a.rv-path` (the monospace path + style) is declared later in this file and ties on specificity otherwise. */ +.file-preview-body .file-preview-md a.rv-path { + font: inherit; + word-break: normal; +} + .file-preview-body img { max-width: 100%; max-height: 100%; diff --git a/src/web/routes/file-routes.ts b/src/web/routes/file-routes.ts index 6157ce48..bd8220d4 100644 --- a/src/web/routes/file-routes.ts +++ b/src/web/routes/file-routes.ts @@ -1724,7 +1724,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even // breadth of formats the attachments viewer renders (image/audio/video/pdf) // so the file viewer can open the same files. const ext = filePath.split('.').pop()?.toLowerCase() || ''; - const imageExts = new Set(['png', 'jpg', 'jpeg', 'gif', 'webp', 'svg', 'bmp', 'ico']); + const imageExts = new Set(['png', 'jpg', 'jpeg', 'gif', 'webp', 'avif', 'svg', 'bmp', 'ico']); // Shared with the attachment registry so a video plays the same whether it // sits in the workspace or is reached by id from outside it. const videoExts = VIDEO_ATTACHMENT_EXTENSIONS; @@ -2045,6 +2045,7 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even jpeg: 'image/jpeg', gif: 'image/gif', webp: 'image/webp', + avif: 'image/avif', ico: 'image/x-icon', bmp: 'image/bmp', mp4: 'video/mp4', diff --git a/test/file-preview-markdown.test.ts b/test/file-preview-markdown.test.ts new file mode 100644 index 00000000..553e2363 --- /dev/null +++ b/test/file-preview-markdown.test.ts @@ -0,0 +1,292 @@ +/** + * @fileoverview File Viewer text view: rendered markdown plus Lines/Wrap toggles. + * + * Clicking a `.md` in the Files panel showed wrapped source with no way to see + * it rendered, although the Response Viewer's marked + DOMPurify pipeline + * (`_renderMarkdown`) was already on the page. The viewer now renders markdown + * through that same pipeline, with an MD toggle back to source, and the + * plain-text view gained Lines and Wrap toggles. Pinned here: + * + * 1. `.md` renders into `.rv-text.file-preview-md[data-i18n-skip]` while the + * pref is on and into a `<pre>` of per-line spans while it is off; the MD + * toggle re-renders WITHOUT a second fetch and persists per device. + * 2. Relative image refs are rebased onto the workspace-confined file-raw + * route under the document's directory and a failed load degrades to alt + * text; relative links become `a.rv-path` for the Response Viewer delegate + * and lose the `target` marked gave them, while fragment and http(s) links + * stay untouched. + * 3. Markdown fetches the route's line ceiling; other text keeps 500. + * 4. Lines/Wrap flip classes on the <pre> and persist, and the text the <pre> + * holds is byte-identical to the file; every toggle is hidden for an image + * and while editing. + * 5. `FILE_PREVIEW_EXTENSIONS` gained avif/ico and still has no `md` + * (in-workspace text keeps the tail viewer, see architecture-invariants). + * + * Loaded via `vm` with a jsdom document injected (the technique from + * response-viewer-file-links.test.ts): constants.js + panels-ui.js only, with + * the app.js markdown pipeline stubbed to a fixed fragment. + */ + +import { readFileSync } from 'node:fs'; +import { resolve } from 'node:path'; +import vm from 'node:vm'; +import { JSDOM } from 'jsdom'; +import { describe, expect, it, vi } from 'vitest'; + +const PUBLIC = resolve(import.meta.dirname, '../src/web/public'); +const constantsJs = readFileSync(resolve(PUBLIC, 'constants.js'), 'utf8'); +const panelsJs = readFileSync(resolve(PUBLIC, 'panels-ui.js'), 'utf8'); + +// A real origin: vitest's equality walker reaches the window through a node's +// ownerDocument, and jsdom's localStorage getter throws on an opaque one. +const dom = new JSDOM('<!DOCTYPE html><html><body></body></html>', { url: 'http://localhost/' }); +const { document } = dom.window; + +/** What the stubbed `_renderMarkdown` hands back: every ref shape the rebase pass must classify. */ +const MARKDOWN_HTML = + '<h1>Title</h1><p>x</p>' + + '<img src="img/a.png#gh-dark-mode-only" alt="Alt A">' + + '<img src="https://cdn.example.com/r.png" alt="remote">' + + '<a href="guide/x.md#sec" target="_blank" rel="noopener noreferrer">x</a>' + + '<a href="../CHANGELOG.md" target="_blank" rel="noopener noreferrer">up</a>' + + '<a href="#top">t</a>' + + '<a href="https://e.com" target="_blank" rel="noopener noreferrer">e</a>'; + +const MD_CONTENT = '# Title\n\nx\n'; +const TXT_CONTENT = 'one\n\n three\tfour\n'; + +function jsonResponse(body: unknown) { + return { ok: true, status: 200, json: async () => body, text: async () => JSON.stringify(body) }; +} + +/** Answer file-content like the route does: text as JSON, an image as metadata. */ +function fetchStub(url: string) { + const path = decodeURIComponent(new URL(url, 'http://x').searchParams.get('path') || ''); + const ext = path.split('.').pop() || ''; + if (ext === 'png') { + return jsonResponse({ + success: true, + data: { type: 'image', url: `/file-raw?path=${path}`, size: 5, extension: ext }, + }); + } + const content = ext === 'md' ? MD_CONTENT : TXT_CONTENT; + if (url.includes('edit=1')) { + return jsonResponse({ + success: true, + data: { content, hash: 'h', eol: 'lf', totalLines: 3, size: content.length }, + }); + } + return jsonResponse({ + success: true, + data: { path, content, totalLines: 3, size: content.length, truncated: false, extension: ext, editable: true }, + }); +} + +function loadApp(prefs: Record<string, string> = {}) { + const store = new Map(Object.entries(prefs)); + const CodemanApp = function CodemanApp(this: unknown) {} as unknown as new () => Record<string, any>; + const fetchMock = vi.fn(async (url: string) => fetchStub(url)); + const context = vm.createContext({ + CodemanApp, + console: { ...console, warn: vi.fn(), error: vi.fn() }, + localStorage: { + getItem: (k: string) => (store.has(k) ? store.get(k) : null), + setItem: (k: string, v: string) => store.set(k, v), + removeItem: (k: string) => store.delete(k), + }, + document, + window: { addEventListener: vi.fn(), removeEventListener: vi.fn(), open: vi.fn() }, + MobileDetection: {}, + setTimeout, + clearTimeout, + confirm: () => true, + fetch: fetchMock, + }); + vm.runInContext(`${constantsJs}\n${panelsJs}\nglobalThis.__exts = FILE_PREVIEW_EXTENSIONS;`, context, { + filename: 'panels-ui.js', + }); + + document.body.innerHTML = ` + <div id="filePreviewOverlay"></div><span id="filePreviewTitle"></span> + <button id="filePreviewMdBtn" hidden></button> + <button id="filePreviewLinesBtn" hidden></button> + <button id="filePreviewWrapBtn" hidden></button> + <button id="filePreviewEditBtn" hidden></button> + <button id="filePreviewDetachBtn" hidden></button> + <div id="filePreviewBody"></div><div id="filePreviewFooter"></div>`; + + const app = new CodemanApp(); + app.$ = (id: string) => document.getElementById(id); + app._resetFilePreviewEdit = () => {}; + app._isExternalPreviewPath = () => false; + app.formatFileSize = (n: number) => `${n} B`; + app.showToast = vi.fn(); + app.filePreviewContent = ''; + app._renderMarkdown = vi.fn(() => MARKDOWN_HTML); + app._linkifyFilePaths = vi.fn(); + app._bindResponseViewerInteractions = vi.fn(); + + const byId = (id: string) => document.getElementById(id) as HTMLButtonElement; + return { + app, + fetchMock, + store, + body: byId('filePreviewBody'), + exts: (context as { __exts: Set<string> }).__exts, + btn: { md: byId('filePreviewMdBtn'), lines: byId('filePreviewLinesBtn'), wrap: byId('filePreviewWrapBtn') }, + }; +} + +describe('file viewer rendered markdown', () => { + it('renders .md through the shared markdown pipeline, inert to i18n, with the viewer delegate bound', async () => { + const { app, body, btn } = loadApp(); + + await app.openFilePreview('docs/README.md', 's1'); + + const doc = body.firstElementChild as HTMLElement; + expect(doc.matches('.rv-text.file-preview-md[data-i18n-skip]')).toBe(true); + expect(doc.querySelector('h1')?.textContent).toBe('Title'); + expect(app._renderMarkdown).toHaveBeenCalledWith(MD_CONTENT); + // Identity, not deep equality: DOM nodes are compared by reference here. + expect(app._linkifyFilePaths.mock.calls[0][0]).toBe(doc); + expect(app._bindResponseViewerInteractions.mock.calls[0][0]).toBe(body); + // The source stays what Copy copies. + expect(app.filePreviewContent).toBe(MD_CONTENT); + // MD is the only toggle that applies to a rendered document; Edit still offered. + expect(btn.md.hidden).toBe(false); + expect(btn.md.getAttribute('aria-pressed')).toBe('true'); + expect(btn.lines.hidden).toBe(true); + expect(btn.wrap.hidden).toBe(true); + expect(document.getElementById('filePreviewEditBtn')!.hidden).toBe(false); + }); + + it('fetches the route ceiling for markdown and the 500-line cap for other text', async () => { + const { app, fetchMock } = loadApp(); + + await app.openFilePreview('docs/README.md', 's1'); + await app.openFilePreview('notes.txt', 's1'); + + const urls = fetchMock.mock.calls.map((c) => c[0]); + expect(urls[0]).toContain('lines=10000'); + expect(urls[1]).toContain('lines=500'); + }); + + it('rebases relative images and links onto the document directory and leaves the rest alone', async () => { + const { app, body } = loadApp(); + + await app.openFilePreview('docs/README.md', 's1'); + + const local = body.querySelector('img[alt="Alt A"]')!; + expect(local.getAttribute('src')).toBe(`/api/sessions/s1/file-raw?path=${encodeURIComponent('docs/img/a.png')}`); + expect(body.querySelector('img[alt="remote"]')!.getAttribute('src')).toBe('https://cdn.example.com/r.png'); + + const rel = body.querySelector('a.rv-path')!; + expect(rel.getAttribute('data-path')).toBe('docs/guide/x.md'); + expect(rel.getAttribute('href')).toBe('#'); + expect(rel.hasAttribute('target')).toBe(false); + expect(rel.hasAttribute('rel')).toBe(false); + + const anchors = Array.from(body.querySelectorAll('a')); + // `..` is collapsed against the document directory, so the title reads + // CHANGELOG.md rather than docs/../CHANGELOG.md. + const up = anchors.find((a) => a.textContent === 'up')!; + expect(up.classList.contains('rv-path')).toBe(true); + expect(up.getAttribute('data-path')).toBe('CHANGELOG.md'); + const fragment = anchors.find((a) => a.textContent === 't')!; + expect(fragment.getAttribute('href')).toBe('#top'); + expect(fragment.classList.contains('rv-path')).toBe(false); + const external = anchors.find((a) => a.textContent === 'e')!; + expect(external.getAttribute('href')).toBe('https://e.com'); + expect(external.getAttribute('target')).toBe('_blank'); + }); + + it('degrades an image that fails to load to its alt text', async () => { + const { app, body } = loadApp(); + + await app.openFilePreview('docs/README.md', 's1'); + const remote = body.querySelector('img[alt="remote"]')!; + remote.dispatchEvent(new dom.window.Event('error')); + + expect(body.querySelector('img[alt="remote"]')).toBeNull(); + expect(body.textContent).toContain('remote'); + }); + + it('MD toggle flips to per-line source and back without refetching, and persists', async () => { + const { app, body, btn, fetchMock, store } = loadApp(); + + await app.openFilePreview('docs/README.md', 's1'); + app.toggleFilePreviewMd(); + + const pre = body.firstElementChild as HTMLElement; + expect(pre.matches('pre.file-preview-text')).toBe(true); + expect(pre.querySelectorAll('.fp-line')).toHaveLength(MD_CONTENT.split('\n').length); + expect(pre.textContent).toBe(MD_CONTENT); + expect(store.get('codeman:filePreviewMdRendered')).toBe('0'); + expect(btn.md.getAttribute('aria-pressed')).toBe('false'); + expect(btn.lines.hidden).toBe(false); + expect(btn.wrap.hidden).toBe(false); + + app.toggleFilePreviewMd(); + expect((body.firstElementChild as HTMLElement).matches('.file-preview-md')).toBe(true); + expect(store.get('codeman:filePreviewMdRendered')).toBe('1'); + expect(fetchMock).toHaveBeenCalledTimes(1); + }); + + it('opens as source when the device pref says so', async () => { + const { app, body, btn } = loadApp({ 'codeman:filePreviewMdRendered': '0' }); + + await app.openFilePreview('docs/README.md', 's1'); + + expect((body.firstElementChild as HTMLElement).matches('pre.file-preview-text')).toBe(true); + expect(btn.md.hidden).toBe(false); + expect(btn.md.getAttribute('aria-pressed')).toBe('false'); + }); +}); + +describe('file viewer Lines and Wrap toggles', () => { + it('flip classes on the <pre>, persist, and never alter the text', async () => { + const { app, body, btn, store } = loadApp(); + + await app.openFilePreview('notes.txt', 's1'); + + const pre = body.firstElementChild as HTMLElement; + expect(pre.matches('pre.file-preview-text.wrap:not(.show-lines)')).toBe(true); + expect(pre.textContent).toBe(TXT_CONTENT); + expect(btn.md.hidden).toBe(true); + + app.toggleFilePreviewLines(); + expect(pre.classList.contains('show-lines')).toBe(true); + expect(store.get('codeman:filePreviewLineNumbers')).toBe('1'); + expect(btn.lines.getAttribute('aria-pressed')).toBe('true'); + + app.toggleFilePreviewWrap(); + expect(pre.classList.contains('wrap')).toBe(false); + expect(store.get('codeman:filePreviewWrap')).toBe('0'); + expect(btn.wrap.getAttribute('aria-pressed')).toBe('false'); + // Same element, no re-render: the counter gutter is CSS, not text. + expect(body.firstElementChild).toBe(pre); + expect(pre.textContent).toBe(TXT_CONTENT); + }); + + it('are hidden for an image and while editing', async () => { + const { app, btn, body } = loadApp(); + + await app.openFilePreview('shot.png', 's1'); + expect(btn.md.hidden && btn.lines.hidden && btn.wrap.hidden).toBe(true); + + await app.openFilePreview('notes.txt', 's1'); + expect(btn.lines.hidden).toBe(false); + await app.enterFilePreviewEdit(); + expect(body.querySelector('textarea.file-preview-editor')).not.toBeNull(); + expect(btn.md.hidden && btn.lines.hidden && btn.wrap.hidden).toBe(true); + }); +}); + +describe('FILE_PREVIEW_EXTENSIONS', () => { + it('routes avif and ico paths to the viewer and leaves .md with the tail viewer', () => { + const { exts } = loadApp(); + expect(exts.has('avif')).toBe(true); + expect(exts.has('ico')).toBe(true); + expect(exts.has('md')).toBe(false); + }); +}); diff --git a/test/routes/file-routes.test.ts b/test/routes/file-routes.test.ts index 24604e45..daa89066 100644 --- a/test/routes/file-routes.test.ts +++ b/test/routes/file-routes.test.ts @@ -681,6 +681,18 @@ describe('file-routes', () => { expect(body.data.url).toContain('file-raw'); }); + it('classifies avif as an image so the viewer renders it instead of dumping bytes', async () => { + mockedStat.mockResolvedValue({ size: 1024 } as never); + + const res = await harness.app.inject({ + method: 'GET', + url: `/api/sessions/${harness.ctx._sessionId}/file-content?path=photo.avif`, + }); + expect(res.statusCode).toBe(200); + const body = JSON.parse(res.body); + expect(body.data.type).toBe('image'); + }); + it('returns audio metadata for audio files', async () => { mockedStat.mockResolvedValue({ size: 2048 } as never); @@ -817,6 +829,19 @@ describe('file-routes', () => { expect(res.headers['content-type']).toBe('image/png'); }); + it('serves avif with its image type, since <img> refuses an octet-stream', async () => { + const content = Buffer.from('fake avif data'); + mockedCreateReadStream.mockReturnValue(Readable.from([content]) as never); + mockedStat.mockResolvedValue({ size: content.length } as never); + + const res = await harness.app.inject({ + method: 'GET', + url: `/api/sessions/${harness.ctx._sessionId}/file-raw?path=photo.avif`, + }); + expect(res.statusCode).toBe(200); + expect(res.headers['content-type']).toBe('image/avif'); + }); + it('serves workspace SVG as an untrusted attachment instead of inline image/svg+xml', async () => { const content = Buffer.from('<svg><script>alert("xss")</script></svg>'); mockedCreateReadStream.mockReturnValue(Readable.from([content]) as never); From 61037082d1420d294ded5e469820c513413fd384 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 15:23:22 +0200 Subject: [PATCH 20/46] fix(mobile): make the phone header tab strip read as live tabs On a phone every inactive tab rendered transparent: grey 11px text floating in unmarked gaps, a boxed Alt+N digit in each tab (a phone has no Alt key), names capped at 50px so a shared `w1-` prefix was most of what showed, and the tab that did not fit was chopped mid-word against the connection dot. The strip looked like a row of disabled labels. Phone block of mobile.css only: - Every header tab is a chip, filled and bordered from the skin's --control-* tokens, name in --text at weight 500. Written `:where(.header) .session-tab` so it stays at (0,1,0): the per-colour left border still wins, and sidebar layout (where the list leaves the header) is untouched. - The Alt+N digit is hidden in the header; inactive tabs drop their empty .tab-actions container, which padded the chip's right side. - Name cap 50px -> 80px, status dot 4px -> 6px, strip gap 2px -> 6px. - Scroll-driven edge fade: a mask on the strip whose widths follow its own inline scroll timeline (registered @property lengths), so the clipped tab dissolves into the edge. No JS; a strip that does not overflow gets no mask, and browsers without scroll timelines keep the old hard edge. The tap-zone arithmetic comment is updated for the numberless phone tabs and the bigger dot (the required reserve drops from 38px to 36px; the 44px min-width stays). test/mobile-tab-strip-chips.test.ts pins the (0,1,0) selector, the top-level @property registration and the timeline-after-shorthand order, each of which fails silently otherwise. test/mobile/tabs.test.ts follows the new name cap. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- src/web/public/mobile.css | 133 +++++++++++++++++++--- test/mobile-tab-strip-chips.test.ts | 166 ++++++++++++++++++++++++++++ test/mobile/tabs.test.ts | 7 +- 3 files changed, 287 insertions(+), 19 deletions(-) create mode 100644 test/mobile-tab-strip-chips.test.ts diff --git a/src/web/public/mobile.css b/src/web/public/mobile.css index 1a5136f4..5d47f317 100644 --- a/src/web/public/mobile.css +++ b/src/web/public/mobile.css @@ -332,6 +332,39 @@ html.mobile-init .file-browser-panel { } } +/* Edge fade for the phone's header tab strip (used in the block below). + Registered so the keyframes can interpolate them as lengths; @property is + only valid at the top level, hence out here. */ +@property --tab-strip-fade-start { + syntax: '<length>'; + inherits: false; + initial-value: 0px; +} + +@property --tab-strip-fade-end { + syntax: '<length>'; + inherits: false; + initial-value: 0px; +} + +/* Driven by the strip's own scroll position, not by time: at the start only the + end edge fades, at the end only the start edge, and in between both. */ +@keyframes tab-strip-edge-fade { + 0% { + --tab-strip-fade-start: 0px; + --tab-strip-fade-end: 28px; + } + 10%, + 90% { + --tab-strip-fade-start: 28px; + --tab-strip-fade-end: 28px; + } + 100% { + --tab-strip-fade-start: 28px; + --tab-strip-fade-end: 0px; + } +} + /* ============================================================================ Phone Breakpoint (<600px) ============================================================================ */ @@ -652,7 +685,7 @@ html.mobile-init .file-browser-panel { overscroll-behavior-x: contain; scrollbar-width: none; max-height: 36px; - gap: 2px; + gap: 6px; padding: 0; } @@ -660,6 +693,36 @@ html.mobile-init .file-browser-panel { display: none; } + /* Fade the strip's edges while there is more to scroll to, so the tab that + does not fit dissolves into the edge instead of being cut mid-word against + the connection dot. Scroll-driven, no JS: the timeline is the strip's own + inline scroll (keyframes + registered properties above this block). A strip + that does not overflow has an INACTIVE timeline, so the animation applies + nothing and both widths stay at their registered 0px, which is no mask at + all. Browsers without scroll timelines skip the block and keep the hard + edge. Header only: in sidebar layout the same list scrolls vertically. */ + @supports (animation-timeline: scroll()) { + .header .session-tabs { + -webkit-mask-image: linear-gradient( + to right, + transparent, + #000 var(--tab-strip-fade-start), + #000 calc(100% - var(--tab-strip-fade-end)), + transparent + ); + mask-image: linear-gradient( + to right, + transparent, + #000 var(--tab-strip-fade-start), + #000 calc(100% - var(--tab-strip-fade-end)), + transparent + ); + /* The shorthand resets animation-timeline, so the timeline comes after. */ + animation: tab-strip-edge-fade linear both; + animation-timeline: scroll(self inline); + } + } + /* Smaller tabs for mobile */ .session-tab { flex-shrink: 0; @@ -671,10 +734,46 @@ html.mobile-init .file-browser-panel { border-radius: 4px; } - /* Smaller status indicator on mobile */ + /* Every tab in the header strip is a chip, not only the active one. Left + transparent, the strip read as a row of disabled labels: grey 11px text + floating in unmarked gaps, with nothing saying "tap me". Fill and border + come from the skin's control tokens, so the four light skins (which repaint + the header with --glass-bg) get a matching chip with no override block, and + the active tab's !important fill and border in styles.css still win. + `:where(.header)` keeps this at (0,1,0): the per-colour left border + (`.session-tab[data-color="red"]`, (0,2,0)) must still outrank the + border-color here, and in sidebar layout the list leaves the header, so + its rows are untouched. */ + :where(.header) .session-tab { + border-radius: 8px; + background: var(--control-bg-hover); + border-color: var(--control-border-hover); + color: var(--text); + } + + :where(.header) .session-tab .tab-name { + font-weight: 500; + } + + /* Only the active tab shows its action icons on a phone (see below), so on + every other tab the container is empty but still a flex item, and its gap + made the chip visibly wider on the right than on the left. */ + :where(.header) .session-tab:not(.active) .tab-actions { + display: none; + } + + /* The boxed digit is the Alt+1..9 shortcut hint. A phone has no Alt key, so + here it was only a second grey box inside every tab, and 20px of the name's + width. Every header tab is therefore numberless on a phone, which is the + case the active-tab reserve below is already sized for. */ + :where(.header) .session-tab .tab-number { + display: none; + } + + /* Status dot: 6px so an idle green reads at arm's length (4px was a speck). */ .session-tab .tab-status { - width: 4px; - height: 4px; + width: 6px; + height: 6px; } /* The working dot is the one glance-state a phone needs: keep idle tiny, but @@ -710,9 +809,12 @@ html.mobile-init .file-browser-panel { opacity: 0.5; } - /* Truncate tab names more aggressively on mobile */ + /* Truncate tab names on mobile. 80px, not the old 50px: session names share + a `w1-` style prefix, and at 50px "w1-ingest-pipeline" became "w1-inge…" + and a clipped tab just "w1-", which says nothing about which session it is. + The 20px the hidden tab number gave back pays for most of the difference. */ .session-tab .tab-name { - max-width: 50px; + max-width: 80px; overflow: hidden; text-overflow: ellipsis; } @@ -726,20 +828,19 @@ html.mobile-init .file-browser-panel { difference instead, which costs a little strip space on exactly one tab and keeps tap-to-switch the majority of it. - ⚠️ The floor is set by the 10th tab onward, NOT by the numbered tabs you - are looking at. `.tab-number` is rendered only for `_tabIdx < 9` (app.js), - so tab 10 loses 16px + a 4px gap off its left and its centre sits 10px - further right. The centre clears the icons when + ⚠️ The floor is set by a NUMBERLESS tab. `.tab-number` is rendered only + for `_tabIdx < 9` (app.js), and the header hides it on phones altogether + (above), so every phone tab is that case now; a numbered one would sit 10px + further left and hide the problem. The centre clears the icons when reserved > icons + rightEdge - leftRunUp - gap - = 50 + 9 - 17 - 4 = 38px + = 50 + 9 - 19 - 4 = 36px with icons = gear 32 + close 20 - close's -2px margin, leftRunUp = border 1 - + padding 8 + status dot 4 + gap 4, and rightEdge = padding 8 + border 1. - Hit testing snaps to whole pixels, so 39px still lands on the gear: the - practical floor is 40px and 44px keeps 4px of headroom. A NUMBERED tab - clears it at 20px, so reasoning from the tabs on screen is exactly what - would put the centre back on the gear. Pinned by + + padding 8 + status dot 6 + gap 4, and rightEdge = padding 8 + border 1. + Hit testing snaps to whole pixels, so a centre half a pixel short still + lands on the gear: the practical floor was measured at 40px (with the + older 4px dot) and 44px keeps headroom. Pinned by test/mobile-tab-tap-zones.test.ts. */ .session-tab.active .tab-name { min-width: 44px; diff --git a/test/mobile-tab-strip-chips.test.ts b/test/mobile-tab-strip-chips.test.ts new file mode 100644 index 00000000..4c491394 --- /dev/null +++ b/test/mobile-tab-strip-chips.test.ts @@ -0,0 +1,166 @@ +/** + * @fileoverview The phone header's tab strip must read as live tabs. + * + * It used to render every inactive tab transparent: grey 11px text floating in + * unmarked gaps, a boxed Alt+N digit in each (a phone has no Alt key), names cut + * to 50px so a shared `w1-` prefix was most of what showed, and the tab that did + * not fit chopped mid-word against the connection dot. On a phone it looked like + * a row of disabled labels. + * + * The fix is four small rules in the phone block of mobile.css, and each has a + * way to be silently undone, which is what this file fences: + * + * - The chip rule is written `:where(.header) .session-tab` so it stays at + * (0,1,0). Written `.header .session-tab` it would be (0,2,0), tie with the + * per-colour `.session-tab[data-color="red"]` left border in styles.css, and + * win on source order (mobile.css loads later): every colour-tagged tab would + * lose its identity stripe. + * - The edge fade is scroll-DRIVEN (no JS). Its two widths must be registered + * with @property to interpolate, and @property is only valid at the top + * level: nested inside the phone @media it is dropped, the keyframes stop + * interpolating, and the fade snaps between states instead of following the + * scroll position. + * - `animation` is a shorthand that resets `animation-timeline`, so the + * timeline must be declared AFTER it or the fade silently becomes a 0s time + * animation. + * + * Parsed with postcss because the declarations live in nested at-rules. The + * rendered result (chips on dark and light skins, the fade at both scroll ends) + * was checked in a browser; this is the cheap regression fence. Port: N/A. + */ +import { readFileSync } from 'node:fs'; +import { resolve } from 'node:path'; +import postcss, { type AtRule, type Declaration, type Rule } from 'postcss'; +import { describe, expect, it } from 'vitest'; + +const CSS = readFileSync(resolve(import.meta.dirname, '../src/web/public/mobile.css'), 'utf8'); +const ROOT = postcss.parse(CSS); +const PHONE_QUERY = '(max-width: 599px)'; + +/** Declarations of the rule matching `selector` inside the phone block (later rules win). */ +function phoneDeclarations(selector: string): Record<string, string> { + const found: Record<string, string> = {}; + ROOT.walkAtRules('media', (atRule) => { + if (atRule.params !== PHONE_QUERY) return; + atRule.walkRules((rule: Rule) => { + if (!rule.selectors.map((s) => s.trim()).includes(selector)) return; + rule.walkDecls((decl: Declaration) => { + found[decl.prop] = decl.value.trim(); + }); + }); + }); + return found; +} + +/** The `@property` rule for `name`, wherever it sits. */ +function propertyRule(name: string): AtRule | undefined { + let hit: AtRule | undefined; + ROOT.walkAtRules('property', (atRule) => { + if (atRule.params.trim() === name) hit = atRule; + }); + return hit; +} + +function declsOf(node: AtRule | Rule): Record<string, string> { + const out: Record<string, string> = {}; + node.each((child) => { + if (child.type === 'decl') out[child.prop] = child.value.trim(); + }); + return out; +} + +describe('phone header tab strip', () => { + describe('chips', () => { + const chip = phoneDeclarations(':where(.header) .session-tab'); + + it('gives every header tab a fill and border from the skin control tokens', () => { + // Tokens, not literals: the four light skins repaint the header with + // --glass-bg and define their own --control-* values. + expect(chip.background).toMatch(/^var\(--control-bg/); + expect(chip['border-color']).toMatch(/^var\(--control-border/); + expect(chip.color).toBe('var(--text)'); + }); + + it('keeps the chip selector at (0,1,0) so per-colour borders still win', () => { + // The lookup above only matches the exact `:where(.header)` spelling, so a + // rewrite to `.header .session-tab` leaves it empty and fails here. + expect(Object.keys(chip).length).toBeGreaterThan(0); + expect(phoneDeclarations('.header .session-tab')).toEqual({}); + }); + + it('hides the Alt+N digit, which a phone has no key for', () => { + expect(phoneDeclarations(':where(.header) .session-tab .tab-number').display).toBe('none'); + }); + + it('drops the empty action container on inactive tabs only', () => { + // The active tab's gear and close live in .tab-actions, so the rule must + // stay scoped to :not(.active). + expect(phoneDeclarations(':where(.header) .session-tab:not(.active) .tab-actions').display).toBe('none'); + expect(phoneDeclarations(':where(.header) .session-tab .tab-actions')).toEqual({}); + }); + + it('leaves enough name to get past a shared w1- prefix', () => { + const maxWidth = Number.parseFloat(phoneDeclarations('.session-tab .tab-name')['max-width'] ?? ''); + expect(maxWidth).toBeGreaterThanOrEqual(72); + }); + }); + + describe('scroll-driven edge fade', () => { + it('registers both fade widths at the top level, as lengths starting at 0px', () => { + for (const name of ['--tab-strip-fade-start', '--tab-strip-fade-end']) { + const rule = propertyRule(name); + expect(rule, `${name} is not registered`).toBeDefined(); + // Nested in @media it is invalid and silently ignored. + expect(rule!.parent?.type, `${name} must be top level`).toBe('root'); + const d = declsOf(rule!); + expect(d.syntax).toBe("'<length>'"); + expect(d['initial-value']).toBe('0px'); + } + }); + + it('fades only the far edge at the start and only the near edge at the end', () => { + let frames: Record<string, Record<string, string>> = {}; + ROOT.walkAtRules('keyframes', (atRule) => { + if (atRule.params.trim() !== 'tab-strip-edge-fade') return; + frames = {}; + atRule.each((node) => { + if (node.type !== 'rule') return; + for (const sel of node.selectors) frames[sel.trim()] = declsOf(node); + }); + }); + expect(frames['0%']?.['--tab-strip-fade-start']).toBe('0px'); + expect(Number.parseFloat(frames['0%']?.['--tab-strip-fade-end'] ?? '0')).toBeGreaterThan(0); + expect(frames['100%']?.['--tab-strip-fade-end']).toBe('0px'); + expect(Number.parseFloat(frames['100%']?.['--tab-strip-fade-start'] ?? '0')).toBeGreaterThan(0); + }); + + it('masks the header strip behind a scroll-timeline feature check, timeline after the shorthand', () => { + let strip: Rule | undefined; + ROOT.walkAtRules('media', (media) => { + if (media.params !== PHONE_QUERY) return; + media.walkAtRules('supports', (supports) => { + if (!/animation-timeline:\s*scroll\(\)/.test(supports.params)) return; + supports.walkRules((rule) => { + if (rule.selectors.map((s) => s.trim()).includes('.header .session-tabs')) strip = rule; + }); + }); + }); + expect(strip, 'no @supports-gated .header .session-tabs rule in the phone block').toBeDefined(); + + const props: string[] = []; + const d: Record<string, string> = {}; + strip!.each((node) => { + if (node.type !== 'decl') return; + props.push(node.prop); + d[node.prop] = node.value.replace(/\s+/g, ' ').trim(); + }); + for (const prop of ['mask-image', '-webkit-mask-image']) { + expect(d[prop]).toContain('var(--tab-strip-fade-start)'); + expect(d[prop]).toContain('var(--tab-strip-fade-end)'); + } + expect(d.animation).toContain('tab-strip-edge-fade'); + expect(d['animation-timeline']).toBe('scroll(self inline)'); + expect(props.indexOf('animation-timeline')).toBeGreaterThan(props.indexOf('animation')); + }); + }); +}); diff --git a/test/mobile/tabs.test.ts b/test/mobile/tabs.test.ts index 33533090..25270e99 100644 --- a/test/mobile/tabs.test.ts +++ b/test/mobile/tabs.test.ts @@ -88,9 +88,10 @@ describe('Tab Navigation', () => { if (tabNameExists) { const maxWidth = await getCSSProperty(page, SELECTORS.TAB_NAME, 'max-width'); const maxWidthPx = parseFloat(maxWidth); - // Should be 50px on mobile - expect(maxWidthPx).toBeLessThanOrEqual(60); - expect(maxWidthPx).toBeGreaterThan(0); + // 80px on phones: wide enough to get past a shared `w1-` prefix, + // still short enough that several tabs fit the strip. + expect(maxWidthPx).toBeLessThanOrEqual(96); + expect(maxWidthPx).toBeGreaterThanOrEqual(72); } }); From 7659ca8b44c23fb764e3222c45d9cbb42058022a Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:18:27 +0200 Subject: [PATCH 21/46] fix(session): keep a tab working while Claude waits for its own workers MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit When Claude hands work to an ultracode workflow or background agents, it ends its own turn and closes it with `✻ Waiting for 1 dynamic workflow to finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when the workers report back. The pane sits quiet with the composer up, so the idle probe called the session idle for the whole wait. At phone width the workflow's progress row also drops its ticking timer, so nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and `_probePaneWorking()` counts it as work. Claude renders the row once from a snapshot and never redraws it, so the same words stay on screen after the workers finish. `isAwaitingWorkers()` therefore tests only the newest column-0 row directly above the composer, never the whole pane and never the PTY stream; a follow-up turn always puts rows of its own there. The column-0 anchor also keeps an agent from holding its own tab busy by printing the sentence. Verified against the live Mac mini pane that reported the bug (2.1.283), and end to end on an isolated instance: an ultracode session running a 90 s workflow at 46 columns stayed busy through the wait and the follow-up turn, then went idle 6 s after that turn closed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CLAUDE.md | 4 +- docs/architecture-invariants.md | 4 +- docs/cli-registry.md | 19 ++- src/config/cli-registry/schema.ts | 9 + src/config/cli-registry/stock.ts | 8 + src/config/cli-registry/types.ts | 13 ++ src/session-activity.ts | 56 +++++++ src/session.ts | 23 ++- test/cli-registry-schema.test.ts | 18 ++ test/session-awaiting-workers.test.ts | 233 ++++++++++++++++++++++++++ 10 files changed, 381 insertions(+), 6 deletions(-) create mode 100644 test/session-awaiting-workers.test.ts diff --git a/CLAUDE.md b/CLAUDE.md index c4a910cc..6c741be6 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -203,6 +203,8 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph ⚠️ **A `❯` sighting is NOT the end of a turn, and neither is silence.** Claude redraws the composer all through a turn, and its working line (`✻ Actualizing… (13m 23s · …)`) is invisible to `SPINNER_PATTERN`, keyword lists and the raw stream. `_confirmIdle()` (session.ts) requires the pane to go quiet AND the SCREEN (`capturePaneText()` + the working-line pattern) to agree; a sustained run of repaints (`session-activity.ts`) marks a turn as started. ⚠️ The composer glyph and working line are per-CLI registry DATA (`capabilities.workDetect`), never Claude constants; a CLI declaring neither falls back to Claude's pair. ⚠️ `workingLine` is config-supplied and runs on the PTY hot path, so it must compile through `compileVersionRegex()` in BOTH the schema refine and `_workingLinePattern()` (ReDoS guard; null, not throw). → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line) +⚠️ **A turn that ENDED waiting for its own workers is working, not idle.** Claude closes such a turn with `✻ Waiting for 1 dynamic workflow to finish` (background agents / ultracode) and resumes by itself; `capabilities.workDetect.awaitingLine` makes the idle probe count it as work. ⚠️ Claude never redraws that row, so it stays on screen after the workers finish: test it ONLY as the newest column-0 row above the composer (`isAwaitingWorkers()`), never pane-wide and never on the stream. → [architecture-invariants#idle-detection-composer-glyph-and-working-line](docs/architecture-invariants.md#idle-detection-composer-glyph-and-working-line) + ⚠️ **A quiet pane is not always a pane that wants you.** A CLI can declare an optional `capabilities.workDetect.watchingLine` (a monitor, background shell or cloud hand-off it is still running); the idle probe reads it into `Session.watching` and `notePrompt()` opens that idle item ALREADY acknowledged, so no surface alerts. Only `idle` is eligible, and the label is pane-derived and prompt-injectable, so a pattern must anchor on chrome only that CLI draws. → [architecture-invariants#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you](docs/architecture-invariants.md#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). Tests: `test/session-watching.test.ts`, `test/watching-no-alert.test.ts`. **An exited agent in a live pane** (`paneExit`, #446): panes use `remain-on-exit on`, so `/exit` leaves a pane, session and pid that look alive; `TmuxManager.startPaneExitWatcher()` publishes `SessionState.paneExit` via `session:updated`. ⚠️ Never set `status: 'error'` or null the `pid` for it; the field is TRI-STATE (absent = UNKNOWN, never alive, scoped by `Session.paneExitApplies`); an absent `#{pane_dead_status}` is not 0; a path that starts a command in a pane must clear the record AND persist. A clean exit is CLOSED via `cleanupSession()` (`pane-exit-sweep.ts`): only an explicit numeric status 0 with no signal, confirmed by 2 reads, with no start/attach in flight (`paneLifecycleInFlight`) and not within 10 s of one (a startup error keeps its row); a crashed agent keeps its row. → [architecture-invariants#an-exited-agent-in-a-live-pane-paneexit](docs/architecture-invariants.md#an-exited-agent-in-a-live-pane-paneexit) @@ -229,7 +231,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md` -**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `<Mode>Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries. It is WRITTEN only by the opt-in CLI management routes (`cliManagementEnabled`, default OFF; `/api/clis`, `cli-registry-routes.ts`), and only through `mutateRegistryFile()` in `registry-writer.ts`, which serializes mutations and refuses (409) a file that does not parse or has group/world permission bits rather than overwriting it. Importing the registry still writes nothing. → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md` +**CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`, `capabilities.workDetect.awaitingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `<Mode>Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries. It is WRITTEN only by the opt-in CLI management routes (`cliManagementEnabled`, default OFF; `/api/clis`, `cli-registry-routes.ts`), and only through `mutateRegistryFile()` in `registry-writer.ts`, which serializes mutations and refuses (409) a file that does not parse or has group/world permission bits rather than overwriting it. Importing the registry still writes nothing. → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md` **External CLI modes (OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek, OMP)**: `isExternalCliMode()` in `session.ts` gates Claude-specific behavior off (Ralph tracker, BashToolParser, token parsing, ❯ readiness; readiness is output stabilization); work detection is per-CLI `capabilities.workDetect` data, not this gate. All eight **require tmux, no direct PTY fallback** (secrets go via socket-scoped `tmux setenv`, never the command line). ⚠️ `run*()` in `session-ui.js` MUST unwrap the `{success,data}` envelope. ⚠️ **Codex uses predictive write-through echo, never the buffer overlay**: `_predictHookOnData` must never `return` (wire path stays byte-identical), and flushed text and a bracketed paste must go out as separate delayed writes. ⚠️ **Pi**: no bypass flag, never invent one; `approveProjectTrust` executes repo code, so it is in the clamp's **materialize** branch; never wire `--api-key`. ⚠️ **Grok**: `alwaysApprove` is stripped for non-granted owners (only-if-sent). ⚠️ **DeepSeek**: the agent is a PROFILE (Run gates on `isDeepSeekRunnable()`); the permission switch is the `DSH_PERMISSION_MODE` env var, so `clampEnvOverridesForOwner()` must DROP `DSH_PERMISSION_MODE`, `DSH_HOME` and `DEEPSEEK_BASE_URL` for non-granted owners; `hooksAvailableForMode()` is per-SESSION for it (pass `sessionHookOptions(session)`) and is never a stand-in for `mode === 'claude'`; answers come from `deepseek-transcript.ts`, paired by header `cwd` + boot window, never newest-mtime. ⚠️ **OMP**: `OMP_AUTH_BROKER_URL`/`_TOKEN` are clamped the same way. → [architecture-invariants#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp](docs/architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek-omp) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 6a25c5b3..3e3ae185 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -105,7 +105,7 @@ Further detail: ⚠️ An ADOPTED container may back SEVERAL cases at different ⚠️ `external`, `hooks` and `altScreen` are three INDEPENDENT capabilities on purpose; deriving one from another shipped the `until=stop`-hangs-on-shell bug. -⚠️ Three capability fields carry a REGEX from config (`discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`) and all three must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. +⚠️ Four capability fields carry a REGEX from config (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine` and `capabilities.workDetect.awaitingLine`) and all four must compile through `compileVersionRegex()`, which caps length and refuses nested quantifiers; `workingLine` is the one that runs on the PTY hot path. ⚠️ **`param` is TWO namespaces.** `launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a LAUNCH PARAM; the legacy `<Mode>Config` wire field is a separate namespace, bridged only by `launch.legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no load error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. Codex is the entry where the two names differ (`bypassApprovals` vs `dangerouslyBypassApprovals`) and therefore the one that catches a regression. @@ -266,6 +266,8 @@ So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks t ⚠️ A third, optional `workDetect` field, `watchingLine` (plus `watchingLines`), tells a quiet pane that is still RUNNING something apart from one that wants a human; it is read from the same idle-confirmation capture. See [The watching signal](#the-watching-signal-a-quiet-pane-that-is-not-waiting-for-you). +⚠️ A fourth, optional `workDetect` field, `awaitingLine`, makes a turn that ENDED waiting for workers the CLI started count as WORKING: with background agents or an ultracode workflow still running at turn end, Claude closes the turn with `✻ Waiting for 1 dynamic workflow to finish` instead of `✻ Brewed for 1m 18s`, then resumes by itself when they report back, so the pane is quiet with the composer up and every other signal read it as idle (reported 2026-09-28 on a live 2.1.283 session). `_probePaneWorking()` ORs it into the same capture. ⚠️ It must never be searched across the pane like `workingLine`: Claude renders the row from a `useState` snapshot taken at turn end and never redraws it, so the words stay on screen after the workers finish and a pane-wide match would pin the tab busy until they scroll off. `isAwaitingWorkers()` (pure, `session-activity.ts`) finds the composer (the LAST row carrying `promptGlyph`, bare or boxed), walks up at most `AWAITING_SEARCH_ROWS` past blank, box-drawn and indented rows, and tests only the first column-0 row. ⚠️ Keep the pattern anchored on column 0 (`^✻`): Claude's own rows start there and the agent's prose never does, which is what stops an agent from holding its own session busy. It is deliberately NOT tested on the PTY stream, where redraws of the stale row would re-mark working. Tests: `test/session-awaiting-workers.test.ts`. + ### The watching signal (a quiet pane that is not waiting for you) **A session that armed a monitor, backgrounded a shell or started a background terminal ends its turn and goes quiet, and a minute later Claude Code's idle notification arrives.** Before this existed, that prompt became an approval item like any other, so every surface filed the session under NEEDS YOU with nothing for a human to answer. The CLI says which kind of quiet it is on its own screen, and reading that row is the whole mechanism: `capabilities.workDetect.watchingLine` (optional, per CLI) plus `watchingLines` (how many non-blank rows at the foot of the screen may hold it, default `WATCHING_TAIL_LINES` = 1). `_confirmIdle()` already captures the pane at the moment a turn ends, so `_readWatching()` runs `watchingLabel()` (pure, `session-activity.ts`) over that same capture; the label lands on `Session.watching` and rides `toLightDetailedState()` out to every payload. ⚠️ It is cached BESIDE `_lastPaneProbeWorking` and goes stale with it, because the probe returns its cached boolean without re-capturing inside `PANE_PROBE_MIN_INTERVAL_MS` and a label from a capture nobody took is a guess. ⚠️ A capture that FAILS clears the label (and announces the change) rather than keeping the last one: a stale label opens the next idle prompt already acknowledged, so keeping it would turn a failed `capture-pane` into a missed alert, while clearing it costs at most an alert the next readable capture takes back. ⚠️ It then FREEZES once `_confirmIdle()` concludes — nothing looks at the pane again until it produces output — which is correct rather than tolerable, since work ending repaints the pane either way (a monitor firing wakes the agent; codex drops its background-terminal row by itself); a timer to keep it fresh would spend a `capture-pane` per idle session per tick to learn nothing. ⚠️ A server restart looks like a hole in that and is not one: the field is live state and starts empty, but reconciliation re-attaches the pane and the attach repaint arms the idle confirmation, which probes and re-reads the label with no input from anyone (measured 2026-09-23, back within ~20 s). A restored session showing no label has no chip on its screen. diff --git a/docs/cli-registry.md b/docs/cli-registry.md index 48cfee37..205514dc 100644 --- a/docs/cli-registry.md +++ b/docs/cli-registry.md @@ -46,8 +46,9 @@ interface CliEntry { launch: CliLaunch; // the structured argv template env: CliEnv; // exports, tmux setenv keys, the env-override allowlist capabilities: CliCapabilities; // what every call site reads instead of the id - // .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines? } — how - // this CLI's pane shows work, and how it shows work it started in the background + // .workDetect?: { promptGlyph, workingLine, watchingLine?, watchingLines?, awaitingLine? } + // — how this CLI's pane shows work, work it started in the background, and a turn + // that ended waiting for workers it will resume from overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store } ``` @@ -56,7 +57,7 @@ interface CliEntry { ### Regexes that come from config -Three capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine` and `capabilities.workDetect.watchingLine`. All three go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. +Four capability fields carry a regular expression an override file can set: `discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine` and `capabilities.workDetect.awaitingLine`. All four go through `compileVersionRegex()`, which caps the source at 200 characters, refuses the nested-quantifier shapes that cause catastrophic backtracking, and returns `null` rather than throwing so every caller degrades instead of crashing. `workingLine` is the one that matters most, because it is compiled once per session and then run against every accumulated PTY chunk and every pane capture. A nested quantifier there is a ReDoS against the event loop for the whole server, not just that session. The guard therefore runs in two places, and neither is redundant: `schema.ts` rejects the entry at LOAD time so a bad pattern never reaches a session, and `_workingLinePattern()` in `session.ts` compiles through the same helper so the runtime cannot end up with a pattern the schema would have refused. @@ -76,6 +77,18 @@ entry declares `watchingLines: 3` and matches that row end to end. Both were mea against live panes rather than read out of a binary, which is the standard for adding a third. +`awaitingLine` covers the quiet pane that is neither idle nor watching: a turn that ENDED +to wait for workers the CLI will resume from by itself. When background agents or an +ultracode workflow are still running at turn end, Claude closes the turn with +`✻ Waiting for 1 dynamic workflow to finish` instead of `✻ Brewed for 1m 18s`, and a pane +showing that row counts as working. ⚠️ Claude renders the row once and never redraws it, so +the words are still on screen after the workers report back and the follow-up turn ends. +The pattern is therefore never run over the whole pane: `isAwaitingWorkers()` +(`session-activity.ts`) walks up from the composer past blank, framed and indented rows and +tests only the first row that starts in column 0, which is the newest transcript row. Claude +starts its own rows in column 0 and the agent's prose never does, so the anchor also keeps an +agent from holding its own session busy. + That label is the one value in the registry that an AGENT can influence, because it comes off the agent's own screen. Two things keep it honest, and both belong to whoever adds a pattern for a new CLI. `watchingLabel()` in `session-activity.ts` searches only the last few diff --git a/src/config/cli-registry/schema.ts b/src/config/cli-registry/schema.ts index 9ff73a59..b72c58fd 100644 --- a/src/config/cli-registry/schema.ts +++ b/src/config/cli-registry/schema.ts @@ -338,6 +338,15 @@ const capabilitiesSchema = z // Bounded hard: this is how far up the screen a config file may push the search, // and every row it adds is one more row the agent itself may be able to write. watchingLines: z.number().int().min(1).max(8).optional(), + // Same guard again: tested against a pane row every time a session settles. + awaitingLine: z + .string() + .min(1) + .refine( + (src) => compileVersionRegex(src) !== null, + 'awaitingLine must be a regex compileVersionRegex() accepts: at most 200 characters, no nested quantifiers' + ) + .optional(), }) .strict() // A window with nothing to search is a typo, not a configuration. Refused at LOAD diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index 94d2e3b7..b562bbd8 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -238,6 +238,14 @@ const CLAUDE: CliEntry = { // watching rather than open that door. See `watchingLabel()` in // `session-activity.ts`. watchingLine: String.raw`·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?|Artifact comment monitors?))`, + // When a turn ends while background agents or an ultracode workflow are still + // running, Claude swaps its `✻ Brewed for 1m 18s` closing row for + // `✻ Waiting for 2 background agents and 1 dynamic workflow to finish` and resumes + // by itself when they report back. Read from the 2.1.283 bundle (the turn-duration + // renderer) and a live pane on 2026-09-28. The row is a snapshot taken at turn end + // and never redrawn, which is why only the newest row above the composer counts. + // Anchored on column 0: Claude's own rows start there, the agent's prose never does. + awaitingLine: String.raw`^✻ Waiting for \d+ (?:background agents?|dynamic workflows?)\b`, }, requiresMux: false, // Claude installs Codeman's own hooks block into every workspace it runs in, so its diff --git a/src/config/cli-registry/types.ts b/src/config/cli-registry/types.ts index e8a20b3e..c39007cd 100644 --- a/src/config/cli-registry/types.ts +++ b/src/config/cli-registry/types.ts @@ -362,6 +362,19 @@ export interface CliCapabilities { * alert. See `watchingLabel()` in `session-activity.ts`. */ watchingLines?: number; + /** + * Source of a regex matching the row this CLI closes a turn with when it ended that + * turn to WAIT for workers it started and will resume on its own once they finish, + * e.g. Claude's `✻ Waiting for 1 dynamic workflow to finish`. A pane showing it counts + * as working, not idle: nothing is being asked of the user, and the next turn starts + * without them. + * + * Unlike `workingLine` this is never searched across the pane. The CLI prints the row + * once and never updates it, so the copy from an earlier turn is still on screen after + * the workers are done. Only the newest transcript row directly above the composer is + * tested. See `isAwaitingWorkers()` in `session-activity.ts`. + */ + awaitingLine?: string; }; /** * How many columns this CLI indents its transcript body by, so a copy taken from its diff --git a/src/session-activity.ts b/src/session-activity.ts index c49cf4e3..599e9662 100644 --- a/src/session-activity.ts +++ b/src/session-activity.ts @@ -158,3 +158,59 @@ export function watchingLabel( } return null; } + +/** + * How many rows above the composer the turn's closing row may sit. Between the two Claude + * draws only its composer border and, sometimes, a right-aligned hint + * (`new task? /clear to save 169.1k tokens`), so this leaves room for a blank row or two + * and no more. A bound, not a tuning knob: the walk must never reach far enough up the + * transcript to find an old turn's row. + */ +export const AWAITING_SEARCH_ROWS = 6; + +/** A row that opens with a box-drawing character is the composer's frame, not transcript. */ +const COMPOSER_FRAME_ROW = /^[─-╿]/; + +/** + * Whether the pane's newest turn ended by handing off to workers the CLI will wait for, + * e.g. Claude's `✻ Waiting for 1 dynamic workflow to finish`. + * + * Such a pane is quiet and shows its composer, so every other signal calls it idle, yet + * nothing is being asked of the user: the CLI resumes by itself when the workers report + * back. That is why a session in this state counts as working. + * + * ⚠️ The row is a snapshot. Claude renders it once, at the end of the turn, and never + * updates it, so after the workers finish the same words are still on screen above the + * follow-up turn. Matching them anywhere on the pane would pin the session busy for as + * long as they stay visible. Only the newest transcript row counts: the walk starts at + * the composer (the LAST row carrying `promptGlyph`), steps up past blank rows, the + * composer's frame and anything indented (a right-aligned hint, a wrapped continuation), + * and tests the first row that starts in column 0. A follow-up turn always puts rows of + * its own there, so the stale copy is never the one tested. + * + * @param promptGlyph the CLI's composer glyph (`capabilities.workDetect.promptGlyph`) + * @returns false when the screen shows no composer, which is no evidence either way + */ +export function isAwaitingWorkers(paneText: string | null | undefined, pattern: RegExp, promptGlyph: string): boolean { + if (!paneText) return false; + const rows = stripAnsi(paneText) + .split('\n') + .map((row) => row.trimEnd()); + let composer = -1; + for (let i = rows.length - 1; i >= 0; i--) { + // Claude has drawn its composer both bare (`❯ …` between rules) and boxed (`│ ❯ … │`). + if (rows[i].replace(/^[\s│]+/, '').startsWith(promptGlyph)) { + composer = i; + break; + } + } + if (composer < 0) return false; + for (let i = composer - 1; i >= Math.max(0, composer - AWAITING_SEARCH_ROWS); i--) { + const row = rows[i]; + if (row === '' || /^\s/.test(row) || COMPOSER_FRAME_ROW.test(row)) continue; + // Same reasoning as watchingLabel(): a caller's `g` flag must not make this flap. + pattern.lastIndex = 0; + return pattern.test(row); + } + return false; +} diff --git a/src/session.ts b/src/session.ts index e8d1c4a3..e5da2586 100644 --- a/src/session.ts +++ b/src/session.ts @@ -85,6 +85,7 @@ import { isSustainedActivity, isPaneQuiet, watchingLabel, + isAwaitingWorkers, WATCHING_TAIL_LINES, IDLE_RECHECK_MS, PANE_PROBE_MIN_INTERVAL_MS, @@ -531,6 +532,8 @@ export class Session extends EventEmitter { private _watchingLineRe: RegExp | null | undefined = undefined; /** Resolved with the pattern above: how many rows at the foot of the screen to search. */ private _watchingWindow = WATCHING_TAIL_LINES; + /** Lazily compiled `capabilities.workDetect.awaitingLine`. See _awaitingLinePattern(). */ + private _awaitingLineRe: RegExp | null | undefined = undefined; private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up) private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read @@ -3093,7 +3096,10 @@ export class Session extends EventEmitter { if (now - this._lastPaneProbeAt < PANE_PROBE_MIN_INTERVAL_MS) return this._lastPaneProbeWorking; this._lastPaneProbeAt = now; const text = this._mux.capturePaneText?.(this._muxSession.muxName) ?? null; - this._lastPaneProbeWorking = text === null ? null : this._workingLinePattern().test(text); + // A turn that ended by handing off to workers the CLI waits for is work too: the + // composer is up and the pane is quiet, but the next turn starts without the user. + this._lastPaneProbeWorking = + text === null ? null : this._workingLinePattern().test(text) || this._paneAwaitsWorkers(text); this._readWatching(text); return this._lastPaneProbeWorking; } @@ -3151,6 +3157,21 @@ export class Session extends EventEmitter { return this._watchingLineRe; } + /** + * Whether the newest turn on this screen ended waiting for workers the CLI started + * (Claude's `✻ Waiting for 1 dynamic workflow to finish`). False for a CLI whose + * registry entry declares no `awaitingLine`. See `isAwaitingWorkers()`. + */ + private _paneAwaitsWorkers(paneText: string): boolean { + if (this._awaitingLineRe === undefined) { + const src = getCli(this.mode)?.capabilities.workDetect?.awaitingLine; + this._awaitingLineRe = src ? compileVersionRegex(src) : null; + } + if (!this._awaitingLineRe) return false; + const glyph = getCli(this.mode)?.capabilities.workDetect?.promptGlyph ?? '❯'; + return isAwaitingWorkers(paneText, this._awaitingLineRe, glyph); + } + /** * The regex matching this CLI's "a turn is running" status line. * diff --git a/test/cli-registry-schema.test.ts b/test/cli-registry-schema.test.ts index abd31ff3..7feb42cb 100644 --- a/test/cli-registry-schema.test.ts +++ b/test/cli-registry-schema.test.ts @@ -165,6 +165,24 @@ describe('workDetect.workingLine is guarded like every other config regex', () = expect(compileVersionRegex(src), `${entry.id} declares a watchingLine the guard refuses`).not.toBeNull(); } }); + + it('holds the optional awaitingLine to the same guard', () => { + expectRejected((e) => { + (e.capabilities as Record<string, unknown>).workDetect = { + promptGlyph: '>', + workingLine: 'working', + awaitingLine: '(a+)+b', + }; + }, 'it is tested against a pane row every time a session settles'); + }); + + it('accepts every shipped awaitingLine', () => { + for (const entry of STOCK_CLIS) { + const src = entry.capabilities.workDetect?.awaitingLine; + if (!src) continue; + expect(compileVersionRegex(src), `${entry.id} declares an awaitingLine the guard refuses`).not.toBeNull(); + } + }); }); describe('no shell text can reach the command line', () => { diff --git a/test/session-awaiting-workers.test.ts b/test/session-awaiting-workers.test.ts new file mode 100644 index 00000000..792d6cb0 --- /dev/null +++ b/test/session-awaiting-workers.test.ts @@ -0,0 +1,233 @@ +/** + * A session whose turn ended waiting for workers it started counts as working. + * + * The bug this pins: when Claude hands work to background agents or an ultracode + * workflow, it ends its turn and closes it with `✻ Waiting for 1 dynamic workflow to + * finish` instead of `✻ Brewed for 1m 18s`. The pane goes quiet with the composer up, so + * every other signal called the session idle while it was plainly busy, and it resumes + * on its own the moment the workers report back. + * + * The chrome rows below (the closing row, the right-aligned hint, the composer rules, the + * footer, the workflow progress row) are verbatim from a live Claude Code 2.1.283 pane on + * 2026-09-28 (`tmux -L codeman capture-pane -p`, 64 columns). The prose and the names are + * invented. + */ +import { describe, expect, it, vi, afterEach } from 'vitest'; +import { Session } from '../src/session.js'; +import { getCli } from '../src/config/cli-registry/index.js'; +import { compileVersionRegex } from '../src/config/cli-registry/patterns.js'; +import { isAwaitingWorkers, AWAITING_SEARCH_ROWS, IDLE_SILENCE_MS } from '../src/session-activity.js'; + +/** The registry's own pattern, which is what every consumer runs. */ +const CLAUDE_AWAITING = compileVersionRegex(getCli('claude')!.capabilities.workDetect!.awaitingLine!)!; + +const RULE = '────────────────────────────────────────────────────────────────'; +const NAMED_RULE = '──────────────────────────────────────────────────── w1-demo ─'; + +/** Everything Claude draws from the composer down while a workflow runs. */ +const COMPOSER_AND_FOOTER = [ + NAMED_RULE, + '❯ sounds good, go ahead', + RULE, + ' Opus 5.5 (1M context) in:285,618 out:581 ctx:29%', + ' ⏵⏵ bypass permissions on (shift+tab to cycle) · ← 2 agents', + '', + ' ◯ docs-research ▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱▱ ↓ 1.1m', +]; + +/** A pane whose newest turn closed with `closing`, then the optional hint row. */ +function frame(closing: string, { hint = true, body = [] as string[] } = {}): string { + return [ + '⏺ The research is running now: three tracks, each checked by', + ' a second agent.', + '', + ' When it is done I will rewrite the plan.', + '', + ...body, + closing, + ...(hint ? [' 286199 tokens'] : []), + ...COMPOSER_AND_FOOTER, + '', + ].join('\n'); +} + +const WAITING_WORKFLOW = frame('✻ Waiting for 1 dynamic workflow to finish'); +const WAITING_BOTH = frame('✻ Waiting for 2 background agents and 1 dynamic workflow to finish'); +const WAITING_AGENT = frame('✻ Waiting for 1 background agent to finish', { hint: false }); +const DONE = frame('✻ Brewed for 1m 18s · done 3:04 PM'); + +/** + * The same screen after the workers reported back and the follow-up turn ended. The + * waiting row is a snapshot Claude never redraws, so it is STILL on screen, just no longer + * the newest row. + */ +const FOLLOW_UP_DONE = frame('✻ Cooked for 12s', { + body: [ + '✻ Waiting for 1 dynamic workflow to finish', + '', + '⏺ All three tracks are back. The plan is rewritten and pushed.', + '', + ], +}); + +describe('isAwaitingWorkers', () => { + it('reads the closing row of a turn that handed off to a workflow', () => { + expect(isAwaitingWorkers(WAITING_WORKFLOW, CLAUDE_AWAITING, '❯')).toBe(true); + }); + + it('reads every form Claude builds the row in', () => { + expect(isAwaitingWorkers(WAITING_BOTH, CLAUDE_AWAITING, '❯')).toBe(true); + expect(isAwaitingWorkers(WAITING_AGENT, CLAUDE_AWAITING, '❯')).toBe(true); + expect(isAwaitingWorkers(frame('✻ Waiting for 3 background agents to finish'), CLAUDE_AWAITING, '❯')).toBe(true); + expect(isAwaitingWorkers(frame('✻ Waiting for 2 dynamic workflows to finish'), CLAUDE_AWAITING, '❯')).toBe(true); + }); + + it('leaves an ordinary turn end alone', () => { + expect(isAwaitingWorkers(DONE, CLAUDE_AWAITING, '❯')).toBe(false); + }); + + it('ignores a stale waiting row once a newer turn has closed below it', () => { + // The trap the whole positional walk exists for: matching the words anywhere on the + // screen would pin the session busy until they scrolled away. + expect(FOLLOW_UP_DONE).toContain('Waiting for 1 dynamic workflow to finish'); + expect(isAwaitingWorkers(FOLLOW_UP_DONE, CLAUDE_AWAITING, '❯')).toBe(false); + }); + + it('refuses the words when the agent wrote them', () => { + // Claude's own rows start in column 0; the agent's prose sits behind `⏺ ` or is + // indented, so an agent cannot keep itself busy by printing the sentence. + expect(isAwaitingWorkers(frame('⏺ ✻ Waiting for 1 background agent to finish'), CLAUDE_AWAITING, '❯')).toBe(false); + expect(isAwaitingWorkers(frame(' ✻ Waiting for 1 background agent to finish'), CLAUDE_AWAITING, '❯')).toBe(false); + }); + + it('says nothing about a screen with no composer on it', () => { + const noComposer = WAITING_WORKFLOW.replace('❯ sounds good, go ahead', ' 1. Yes 2. No'); + expect(isAwaitingWorkers(noComposer, CLAUDE_AWAITING, '❯')).toBe(false); + expect(isAwaitingWorkers('', CLAUDE_AWAITING, '❯')).toBe(false); + expect(isAwaitingWorkers(null, CLAUDE_AWAITING, '❯')).toBe(false); + }); + + it('finds the composer in the boxed layout too', () => { + const boxed = [ + '✻ Waiting for 1 dynamic workflow to finish', + '╭──────────────────────────────────────╮', + '│ ❯ │', + '╰──────────────────────────────────────╯', + ' ⏵⏵ bypass permissions on · ← 1 agent', + ].join('\n'); + expect(isAwaitingWorkers(boxed, CLAUDE_AWAITING, '❯')).toBe(true); + }); + + it('reads a coloured capture', () => { + const coloured = WAITING_WORKFLOW.replace( + '✻ Waiting for 1 dynamic workflow to finish', + '\u001b[2m✻\u001b[0m \u001b[2mWaiting for \u001b[1m1\u001b[22m dynamic workflow to finish\u001b[0m' + ); + expect(isAwaitingWorkers(coloured, CLAUDE_AWAITING, '❯')).toBe(true); + }); + + it('stops looking a few rows above the composer', () => { + const farAway = [ + '✻ Waiting for 1 dynamic workflow to finish', + ...Array.from({ length: AWAITING_SEARCH_ROWS }, () => ''), + ...COMPOSER_AND_FOOTER, + ].join('\n'); + expect(isAwaitingWorkers(farAway, CLAUDE_AWAITING, '❯')).toBe(false); + }); + + it('survives a pattern handed to it with the global flag set', () => { + const global = new RegExp(CLAUDE_AWAITING.source, 'g'); + expect(isAwaitingWorkers(WAITING_WORKFLOW, global, '❯')).toBe(true); + expect(isAwaitingWorkers(WAITING_WORKFLOW, global, '❯')).toBe(true); + }); +}); + +/** A composer repaint: the frame Claude ships roughly once a second while working. */ +const COMPOSER_REPAINT = + '\x1b[31;1H\x1b[38;5;246m❯\xa0\x1b[39m\x1b[0m\x1b[33;1H \x1b[38;5;246mOpus 5 in:143,699 out:669 ctx:14%\x1b[39m'; + +type SessionInternals = { + _handleTerminalOutput(data: string): void; + _detectInteractiveActivity(data: string): void; +}; + +function feed(session: Session, data: string): void { + const internals = session as unknown as SessionInternals; + internals._handleTerminalOutput(data); + internals._detectInteractiveActivity(data); +} + +/** A session whose mux reports a scripted screen for the pane probe to read. */ +function withFakePane(read: () => string, mode: 'claude' | 'codex' = 'claude'): Session { + const mux = { + isAvailable: () => true, + capturePaneText: () => read(), + } as unknown as NonNullable<ConstructorParameters<typeof Session>[0]>['mux']; + return new Session({ + workingDir: '/tmp', + mode, + mux, + muxSession: { muxName: 'codeman-test', sessionId: 'test', createdAt: Date.now() }, + } as ConstructorParameters<typeof Session>[0]); +} + +/** Run one turn and let it end, which is when the probe reads the screen. */ +function runAndSettle(session: Session, repaint: string = COMPOSER_REPAINT): void { + for (let i = 0; i < 3; i++) { + feed(session, repaint); + vi.advanceTimersByTime(1000); + } + vi.advanceTimersByTime(IDLE_SILENCE_MS + 2000); +} + +describe('Session status while its workers run', () => { + afterEach(() => { + vi.useRealTimers(); + }); + + it('stays working when the turn ends waiting for a workflow', () => { + vi.useFakeTimers(); + const session = withFakePane(() => WAITING_WORKFLOW); + + runAndSettle(session); + + expect(session.status).toBe('busy'); + expect(session.isWorking).toBe(true); + }); + + it('goes idle once the follow-up turn closes, although the old row is still on screen', () => { + vi.useFakeTimers(); + const screen = { text: WAITING_WORKFLOW }; + const session = withFakePane(() => screen.text); + runAndSettle(session); + expect(session.status).toBe('busy'); + + screen.text = FOLLOW_UP_DONE; + // The probe keeps re-reading the screen on its own slow cadence while it says busy, + // with no PTY output needed to trigger it. + vi.advanceTimersByTime(IDLE_SILENCE_MS + 10_000); + + expect(session.status).toBe('idle'); + expect(session.isWorking).toBe(false); + }); + + it('reports an ordinary turn end as idle, as before', () => { + vi.useFakeTimers(); + const session = withFakePane(() => DONE); + + runAndSettle(session); + + expect(session.status).toBe('idle'); + }); + + it('does not apply to a CLI whose registry entry declares no awaitingLine', () => { + vi.useFakeTimers(); + expect(getCli('codex')?.capabilities.workDetect?.awaitingLine).toBeUndefined(); + const codexScreen = ['✻ Waiting for 1 dynamic workflow to finish', '', '› Ask Codex to do anything', ''].join('\n'); + const session = withFakePane(() => codexScreen, 'codex'); + + runAndSettle(session, '\x1b[31;1H\x1b[38;5;246m›\xa0\x1b[39m\x1b[0m'); + + expect(session.status).toBe('idle'); + }); +}); From 627b76739cddc49f63855575e13af8eb97b9577d Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:24:24 +0200 Subject: [PATCH 22/46] fix(docker): merge-time fixes for #490 - test: every ENV PATH= line in server.Dockerfile must start $PATH:, and the ~/.local/bin append is pinned alongside /opt/codeman-cli/bin - invariants + CLAUDE.md: the append-only PATH rule names ~/.local/bin too - docker-compose.md: Settings-installed CLIs live in ~/.local on the app-data mount; reinstall once after upgrading; hand-run npm installs need --prefix ~/.local - installEnv() JSDoc describes the in-container npm prefix redirect Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- docs/docker-compose.md | 2 ++ src/web/routes/cli-registry-routes.ts | 4 +++- test/docker-entrypoint.test.ts | 7 +++---- 5 files changed, 10 insertions(+), 7 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 719a59e6..5076d9eb 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -232,7 +232,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Docker cases**: a case can point at a **container** running any CLI mode inside it, a **LOCATION OVERLAY on cases, never a `SessionMode`**. One long-lived container **per case**, shared by its sessions: killing a session kills only its in-container tmux, **never** `docker stop` while siblings remain. The workspace is bind-mounted at the **same absolute path**. Credentials are **seeded**, never shared RW. **NEVER a create-time `-e` for secrets, NEVER `--privileged`, NEVER the docker socket.** A drifted config (label hash) REFUSES the launch. ⚠️ An **adopted** container (`DockerCase.owned === false`) is only `exec`ed into: never create, start, stop, restart, remove, `docker commit` or `docker pause` it; fail closed. Test `owned === false`, never truthiness. ⚠️ Apply `owned` AFTER `dockerConfigHash`. ⚠️ Run modes come from the CONTAINER (`availableModes`), and a failed probe is normal for an OWNED case. ⚠️ Root exec user drops the bypass flag via the registry's `overlays.docker.rootCommand`, never a branch. ⚠️ Adoption is admin-only in multi-user mode. ⚠️ Loopback prod needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for in-container hooks. → [architecture-invariants#docker-cases](docs/architecture-invariants.md#docker-cases), `docs/docker-cases.md` -**Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md` +**Docker Compose deployment** (`docker/`): Codeman runs in a container and spawns Docker cases as **SIBLING** containers via the host socket, never nested. `resolveDockerDaemonMountSource()` maps HOME bind sources into the daemon's namespace (`CODEMAN_DOCKER_HOST_HOME`); `CODEMAN_CASES_PATH` makes workspaces resolve to the same absolute path on both sides. ⚠️ `CODEMAN_CASES_PATH` must move every consumer: resolve it only via `config/cases-dir.ts`. ⚠️ `.dockerignore` matches whole paths: keep `**/.env` or `docker/.env` secrets ship in the image. ⚠️ Long-form binds create missing sources ROOT-OWNED: `Start-Codeman.sh` pre-creates them, and `docker/entrypoint.sh` (root, `cap_add: [CHOWN, DAC_OVERRIDE, KILL, SETGID, SETUID]` against `cap_drop: ALL`, `KILL` for tini; pinned by the test) fixes ownership then drops to `PUID:PGID` via `setpriv`, never re-owning foreign dirs. ⚠️ Append `/opt/codeman-cli` and `~/.local/bin` to `PATH`, never prepend. ⚠️ `server.Dockerfile`, the compose file and `.env.example` feed the self-updater's environment gate (`docs/docker-self-update.md`). → [architecture-invariants#docker-compose-deployment](docs/architecture-invariants.md#docker-compose-deployment), `docs/docker-compose.md` **CLI registry** (`src/config/cli-registry/`): every run mode is a `CliEntry` (discovery, launch argv template, env handling, `capabilities`, and the `overlays` behind remote/docker pane commands). **No code outside `stock.ts` may branch on a CLI id**: use a capability field or a NAMED PROFILE (`profiles.ts`); `test/cli-registry-no-id-branching.test.ts` and `test/frontend-cli-no-id-branching.test.ts` enforce it. ⚠️ Config holds typed argv tokens, never shell text; literals are validated at LOAD time and a bad one rejects the whole entry. ⚠️ Keep `external`, `hooks` and `altScreen` independent; never derive one from another. ⚠️ Config regexes (`discovery.version.regex`, `capabilities.workDetect.workingLine`, `capabilities.workDetect.watchingLine`, `capabilities.workDetect.awaitingLine`) must compile through `compileVersionRegex()`. ⚠️ `privilegedParams[].param` names a LAUNCH PARAM, not the legacy `<Mode>Config` field (bridged only by `launch.legacyConfigAliases`); a wrong name silently clamps nothing. ⚠️ Resolve the registry AT CALL TIME, never in a module-level const. ⚠️ Remote claude/omp arms of `buildRemoteLaunchCommand` are not covered by the pane-command golden. `~/.codeman/clis.json` overrides entries. It is WRITTEN only by the opt-in CLI management routes (`cliManagementEnabled`, default OFF; `/api/clis`, `cli-registry-routes.ts`), and only through `mutateRegistryFile()` in `registry-writer.ts`, which serializes mutations and refuses (409) a file that does not parse or has group/world permission bits rather than overwriting it. Importing the registry still writes nothing. → [architecture-invariants#cli-registry](docs/architecture-invariants.md#cli-registry), `docs/cli-registry.md` diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index ec52e24b..2a05ea57 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -514,7 +514,7 @@ Adding the header changed the preamble, so `CODEMAN_PREAMBLE` was bumped (1.22.0 ⚠️ `KILL` is in that list for tini, not the entrypoint: `init: true` keeps tini as root while the server runs as PUID, and without CAP_KILL its SIGTERM forward fails and the server is SIGKILLed on every `compose down`/`restart` instead of flushing state. -⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) is APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). +⚠️ `/opt/codeman-cli` (the runtime-owned CLI prefix) and `~/.local/bin` (where Settings installs CLIs, on the home bind mount so they survive a rebuild) are APPENDED to `PATH`, never prepended, and the entrypoint pins its own `PATH` to the system dirs: the root part of the start resolves `setpriv` by bare name, and a prefix ahead of `/usr/bin` let a planted `setpriv` run as uid 0 (measured). `~/.local/bin` is writable from the host as well, and root `docker exec` and the healthcheck inherit the image `PATH` without the entrypoint's pin; `test/docker-entrypoint.test.ts` requires every `ENV PATH=` line to start `$PATH:`. `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` drops `--memory-swap` (and filters only that one kernel warning) for hosts without swap accounting; `--memory` still applies. diff --git a/docs/docker-compose.md b/docs/docker-compose.md index d1c73721..8a838d6d 100644 --- a/docs/docker-compose.md +++ b/docs/docker-compose.md @@ -6,6 +6,8 @@ For the Compose configuration, environment settings, storage migration, and macv The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image. +CLIs installed from **App Settings → Agents & CLIs → CLI management** (DeepSeek Harness, Pi, and the other npm-based ones) go to `~/.local` on the `CODEMAN_APPDATA_PATH` mount, so they survive an image rebuild and a container recreate. Releases up to 1.33.1 installed them into the image instead, so a CLI installed from Settings on one of those has to be installed again once after the rebuild. The same applies to a hand-run `npm install -g` inside a session: it writes to the image prefix (`/opt/codeman-cli`) and is lost on the next rebuild, so use `npm install -g --prefix ~/.local <package>` instead. + It can also include the GitHub CLI (`gh`) and the Azure CLI (`az`) with the `azure-devops` extension, wired in as Git credential helpers, so Clone Repo and `git clone` reach private GitHub and Azure DevOps repositories once they are signed in. Both are off by default; [Turning them on](../docker/README.md#turning-them-on) shows the `docker-compose.override.yml` settings. ## Prerequisites diff --git a/src/web/routes/cli-registry-routes.ts b/src/web/routes/cli-registry-routes.ts index 62a54e4f..a4d4b136 100644 --- a/src/web/routes/cli-registry-routes.ts +++ b/src/web/routes/cli-registry-routes.ts @@ -210,7 +210,9 @@ const installsInFlight = new Set<string>(); /** * The server's environment minus every `CODEMAN_*` variable. An install script is third-party * code, and those variables carry Codeman's own secrets and wiring (`CODEMAN_PASSWORD`, the - * data dir, the tmux socket), none of which an installer needs. + * data dir, the tmux socket), none of which an installer needs. Inside the Docker Compose + * container (`CODEMAN_IN_CONTAINER=1`) it also points `NPM_CONFIG_PREFIX` at `$HOME/.local`, + * so an `npm install -g` lands on the persistent home mount instead of the image. */ export function installEnv(source: NodeJS.ProcessEnv = process.env): NodeJS.ProcessEnv { const env: NodeJS.ProcessEnv = {}; diff --git a/test/docker-entrypoint.test.ts b/test/docker-entrypoint.test.ts index b46a0ec0..417b67fe 100644 --- a/test/docker-entrypoint.test.ts +++ b/test/docker-entrypoint.test.ts @@ -110,15 +110,14 @@ describe('docker-compose.yaml cap_add covers what entrypoint.sh and init:true ne }); describe('the runtime-owned CLI prefix never shadows root commands', () => { - it('server.Dockerfile appends /opt/codeman-cli/bin to PATH rather than prepending it', () => { + it('server.Dockerfile appends the runtime-writable CLI dirs to PATH rather than prepending them', () => { const pathLines = dockerfile.split('\n').filter((l) => /^ENV PATH=/.test(l)); expect(pathLines.length).toBeGreaterThan(0); for (const line of pathLines) { - expect(line, 'a writable prefix ahead of $PATH lets a planted setpriv run as root').not.toMatch( - /^ENV PATH=\/opt\/codeman-cli/ - ); + expect(line, 'a writable prefix ahead of $PATH lets a planted setpriv run as root').toMatch(/^ENV PATH=\$PATH:/); } expect(pathLines).toContain('ENV PATH=$PATH:/opt/codeman-cli/bin'); + expect(pathLines).toContain('ENV PATH=$PATH:/home/${CODEMAN_RUNTIME_USER}/.local/bin'); }); it('entrypoint.sh pins PATH to the system directories before its first command', () => { From 1645ef5f5c930d68581f0e6a0c386ccb0071656a Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:25:33 +0200 Subject: [PATCH 23/46] fix(docker): merge-time fixes for #492 - test: the complete-identity case now checks the combined agentImageBuildArgPairs() argv on both producers, so the manual build-agent-image.mjs path cannot drop the identity unnoticed - both producers: GIT_IDENTITY_BUILD_ARGS carries the mirror/parity warning its gh/az neighbour has - the partial-identity error names CODEMAN_AGENT_IMAGE_GIT_USER_NAME and CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL; test regex follows - wiki Docker-Cases: mention the identity variables next to the gh/az switches Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- docs/wiki/Docker-Cases.md | 6 ++++++ scripts/lib/cli-catalog.mjs | 8 ++++++-- src/docker-hosts.ts | 9 +++++++-- test/agent-image-build-args-parity.test.ts | 8 ++++++-- 4 files changed, 25 insertions(+), 6 deletions(-) diff --git a/docs/wiki/Docker-Cases.md b/docs/wiki/Docker-Cases.md index 5298bd11..30812cca 100644 --- a/docs/wiki/Docker-Cases.md +++ b/docs/wiki/Docker-Cases.md @@ -147,6 +147,12 @@ in `docker-compose.override.yml`), then rebuild the image with `--no-cache`. `docker/README.md` ("Private repositories") has the details and the matching switches for the server image. +To give agents a fixed Git commit identity, set `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` and +`CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` together in that same environment (in the Docker +deployment, set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` instead, which feeds +both images). An existing `codeman/agent:base` only picks it up after a `--no-cache` rebuild +and recreated case containers; `docker/README.md` ("Git commit identity") has the details. + ## Isolation Every container runs hardened by default: diff --git a/scripts/lib/cli-catalog.mjs b/scripts/lib/cli-catalog.mjs index fd9c53f6..614a56fb 100644 --- a/scripts/lib/cli-catalog.mjs +++ b/scripts/lib/cli-catalog.mjs @@ -64,7 +64,10 @@ export const GIT_HOST_CLI_BUILD_ARGS = [ ['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'], ]; -/** Environment variables passed through to the agent image's system Git configuration. */ +/** + * Environment variable → Dockerfile ARG for the image's system Git identity. + * ⚠️ Mirrored by `GIT_IDENTITY_BUILD_ARGS` in `src/docker-hosts.ts`; the parity test pins them. + */ export const GIT_IDENTITY_BUILD_ARGS = [ ['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'], ['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'], @@ -94,7 +97,8 @@ export function gitIdentityBuildArgPairs(env) { const configured = pairs.filter(([, value]) => value !== ''); if (configured.length === 0) return []; if (configured.length !== pairs.length) { - throw new Error('Git user name and email must both be set when configuring Git identity'); + const names = GIT_IDENTITY_BUILD_ARGS.map(([envName]) => envName).join(' and '); + throw new Error(`${names} must both be set when configuring Git identity`); } return pairs; } diff --git a/src/docker-hosts.ts b/src/docker-hosts.ts index 3b6f7744..fbafb460 100644 --- a/src/docker-hosts.ts +++ b/src/docker-hosts.ts @@ -615,7 +615,10 @@ export const GIT_HOST_CLI_BUILD_ARGS: ReadonlyArray<readonly [string, string]> = ['CODEMAN_AGENT_IMAGE_INSTALL_AZ', 'CODEMAN_INSTALL_AZ'], ]; -/** Environment variables passed through to the agent image's system Git configuration. */ +/** + * Environment variable → Dockerfile ARG for the image's system Git identity. + * ⚠️ Mirrors `GIT_IDENTITY_BUILD_ARGS` in `scripts/lib/cli-catalog.mjs`; the parity test pins them. + */ export const GIT_IDENTITY_BUILD_ARGS: ReadonlyArray<readonly [string, string]> = [ ['CODEMAN_AGENT_IMAGE_GIT_USER_NAME', 'GIT_USER_NAME'], ['CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL', 'GIT_USER_EMAIL'], @@ -642,13 +645,15 @@ export function gitHostCliBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, /** * The `--build-arg` pairs for a configured Git identity. An absent pair leaves * Git unconfigured, preserving existing deployments; a partial pair is refused. + * ⚠️ Mirrors `gitIdentityBuildArgPairs()` in `scripts/lib/cli-catalog.mjs`; the parity test pins them. */ export function gitIdentityBuildArgPairs(env: NodeJS.ProcessEnv): Array<[string, string]> { const pairs = GIT_IDENTITY_BUILD_ARGS.map(([envName, argName]) => [argName, env[envName] ?? ''] as [string, string]); const configured = pairs.filter(([, value]) => value !== ''); if (configured.length === 0) return []; if (configured.length !== pairs.length) { - throw new Error('Git user name and email must both be set when configuring Git identity'); + const names = GIT_IDENTITY_BUILD_ARGS.map(([envName]) => envName).join(' and '); + throw new Error(`${names} must both be set when configuring Git identity`); } return pairs; } diff --git a/test/agent-image-build-args-parity.test.ts b/test/agent-image-build-args-parity.test.ts index 3ca358d0..bbf315c6 100644 --- a/test/agent-image-build-args-parity.test.ts +++ b/test/agent-image-build-args-parity.test.ts @@ -172,6 +172,9 @@ describe('Git identity in the agent image: both producers pass the same settings expect(mjsGitIdentityPairs(identity)).toEqual(expected); expect(tsGitIdentityPairs({})).toEqual([]); expect(mjsGitIdentityPairs({})).toEqual([]); + // The combined argv, not just the helper: the manual build path could drop the identity otherwise. + expect(tsPairs(identity)).toEqual(mjsPairs(CATALOG, identity)); + expect(tsPairs(identity)).toEqual(expect.arrayContaining(expected)); }); it('refuses a partial identity in both build paths', () => { @@ -179,8 +182,9 @@ describe('Git identity in the agent image: both producers pass the same settings { CODEMAN_AGENT_IMAGE_GIT_USER_NAME: 'Ada Lovelace' }, { CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL: 'ada@example.com' }, ]) { - expect(() => tsGitIdentityPairs(identity)).toThrow(/Git user name and email/); - expect(() => mjsGitIdentityPairs(identity)).toThrow(/Git user name and email/); + const named = /CODEMAN_AGENT_IMAGE_GIT_USER_NAME and CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL must both be set/; + expect(() => tsGitIdentityPairs(identity)).toThrow(named); + expect(() => mjsGitIdentityPairs(identity)).toThrow(named); } }); From 272b56d47bfd812addde0ee8b176ffc4fa00b20f Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:26:47 +0200 Subject: [PATCH 24/46] fix(session): merge-time fixes for #491 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - claude watchingLine: the lookahead keys on "Artifact" alone, so a footer truncated mid-chip ("1 Artifact…", "1 Artifact comm…") is still refused instead of reporting the shell beside it; comment follows - test: both truncations return no watching label - invariants: a chip that waits on a human never counts as watching, and the ^ anchor is what stops the retry past the chip Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- docs/architecture-invariants.md | 2 ++ src/config/cli-registry/stock.ts | 7 ++++--- test/session-watching.test.ts | 3 +++ 3 files changed, 9 insertions(+), 3 deletions(-) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 2a05ea57..810708a7 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -274,6 +274,8 @@ So: `_confirmIdle()` (session.ts) requires the pane to go quiet, and then asks t **The fix is the alert that does not fire; the badge is cosmetic.** `hook-event-routes` passes the label to `notePrompt()`, which opens the idle item ALREADY acknowledged (`acknowledgedAt` + `acknowledgedReason`). Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", so the item stays pending, answerable and available as Read My Mind context, and a wrong label costs a card that does not blink rather than an alert that was never created. Every surface follows from that one flag — the broadcast carries `acknowledgedReason` so a live page declines to arm (`_onHookIdlePrompt`, settings-ui.js), the push is skipped, a reloading page reads `acknowledgedAt` in `seedApprovals()` as it always did, `classifySession()` and `pendingApprovalCount()` (tui-model.ts, tui-render.ts) ignore an acknowledged item, and the TUI card drops to the `info` tone and says why. It re-arms for free: the next idle prompt supersedes the item and is built fresh. ⚠️ Only `idle` is eligible, so a permission or question dialog still goes red whatever else the agent started — but a prose question is NOT a dialog, so an agent that arms a monitor and then asks "which branch?" in plain text is silenced along with the false alarms. That is the accepted cost of the design and the reason the kind gate sits at the single place items are created. +⚠️ **A chip that waits on a HUMAN must never count as watching**: Claude's Artifact comment monitor (an agent that published a page and hears nothing until somebody comments) is refused for the whole row by a `^`-anchored negative lookahead keyed on "Artifact" alone, so a footer cut off mid-chip is refused too, and the `^` is what stops the engine retrying past the chip and reporting a shell beside it. + **⚠️ The label is pane-derived, so the window and the anchor are a trust boundary, not formatting.** An agent that gets its own text matched silences its own alert. Two things prevent it, and BOTH belong to whoever adds a pattern for a new CLI: the window must cover only rows that CLI draws, and the pattern must anchor on chrome only that CLI can produce. Claude satisfies both — its chip is the LAST row, so the default window of one row excludes even the status line directly above it, whose content comes from a `statusLine` command a bypassed session can write into its own `.claude/settings.json`. Codex does not: its row sits above the composer, and the slot it occupies holds the last row of the TRANSCRIPT whenever no terminal is running, so matching the complete row raises the bar without closing it. What contains that is `hooks: 'none'` — no hook event from a codex session reaches `notePrompt()`, so a forged label costs a wrong badge and nothing else, and a CLI that gains hook signals must not keep a pattern that soft. The label is also ANSI-stripped and capped (`MAX_WATCHING_LABEL_CHARS`) at the source, and every interpolation of it into markup goes through `escapeHtml()`, since a config-supplied capture group decides what it holds. Tests: `test/session-watching.test.ts` (the label and both CLI patterns), `test/watching-no-alert.test.ts` (the negative claim across all four surfaces). ### Workspace-trust dialog auto-accept diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index f470b69b..d529a1da 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -242,10 +242,11 @@ const CLAUDE: CliEntry = { // lookahead refuses the whole row while that chip is on it, whatever else is // running beside it. The `^` is what makes the lookahead judge the row once: // without it the engine retries from each later position, and a start past the - // chip reports the shell beside it. The lookahead stops short of "monitor", so a - // footer cut off mid-chip still counts. Counting the chip as watching kept the + // chip reports the shell beside it. The lookahead keys on "Artifact" alone, so a + // footer cut off mid-chip (`· 1 Artifact…`, `· 1 Artifact comm…`) is still refused; + // no other chip on this row says "Artifact". Counting the chip as watching kept the // idle alert quiet for a session that was waiting for a human. - watchingLine: String.raw`^(?!.*Artifact comment).*?·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?))`, + watchingLine: String.raw`^(?!.*Artifact).*?·\s*(\d+ (?:monitors?|shells?|teams?|local agents?|cloud sessions?|MCP tasks?|background tasks?|(?:background|remote) dynamic workflows?))`, // When a turn ends while background agents or an ultracode workflow are still // running, Claude swaps its `✻ Brewed for 1m 18s` closing row for // `✻ Waiting for 2 background agents and 1 dynamic workflow to finish` and resumes diff --git a/test/session-watching.test.ts b/test/session-watching.test.ts index 44b940c4..dbdbbad7 100644 --- a/test/session-watching.test.ts +++ b/test/session-watching.test.ts @@ -174,6 +174,9 @@ describe('watchingLabel', () => { expect( watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comment moni…'), CLAUDE_WATCHING) ).toBeNull(); + // Cut before "comment" is complete: the lookahead keys on "Artifact" alone for these. + expect(watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact comm…'), CLAUDE_WATCHING)).toBeNull(); + expect(watchingLabel(pane('⏵⏵ bypass permissions on · 1 shell · 1 Artifact…'), CLAUDE_WATCHING)).toBeNull(); }); it('says nothing about a pane that is running nothing', () => { From a0fbd1d28d92230941ee75d4fbb70698c045f807 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:27:05 +0200 Subject: [PATCH 25/46] docs(registry): note the cliMouseTracking half of claude's wheel rule (#498) - claude's declared-for-later wheelForward says the live rule in _shouldForwardWheelToApp is the version AND the server-published cliMouseTracking flag, so whoever wires the field up needs both Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- src/config/cli-registry/stock.ts | 2 ++ 1 file changed, 2 insertions(+) diff --git a/src/config/cli-registry/stock.ts b/src/config/cli-registry/stock.ts index d529a1da..465da26a 100644 --- a/src/config/cli-registry/stock.ts +++ b/src/config/cli-registry/stock.ts @@ -263,6 +263,8 @@ const CLAUDE: CliEntry = { transcript: 'claude-jsonl', altScreen: 'strip-full', echo: { policy: 'buffer', anchor: { kind: 'glyph', glyph: '❯', offset: 2 } }, + // Declared-for-later: the live rule (`_shouldForwardWheelToApp`, terminal-ui.js) is this version + // AND the server-published `cliMouseTracking` flag (#498), so wiring this field up needs both. wheelForward: { mode: 'version-gated', minVersion: '2.1.187' }, keyboardAccessory: 'agent', privilegedCommandGate: false, From dfd3df82891426b6a5ac3c42caf4cc4245c81e33 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:28:20 +0200 Subject: [PATCH 26/46] fix(build): merge-time fixes for #500 - pre-push hook: skip with a notice when npm is not on PATH (GUI git clients and IDEs often run hooks with a minimal PATH), instead of blocking every push on "npm: not found"; real-push test with a stripped PATH - test/git-hooks.test.ts: pin GIT_CONFIG_NOSYSTEM=1 and GIT_CONFIG_GLOBAL=/dev/null around the resolveGitHooksDir tests, so an exported global or a system core.hooksPath no longer fails them - watch tsconfig.json, .prettierignore and .editorconfig too: typecheck and format:check read them - check:browser-excludes: fail loudly when the vitest list output and the walked test/**/*.test.ts tree share no path (format drift would otherwise pass vacuously) - Reword the PRE_PUSH_MARKER comment: bumping its version would make every installed v1 hook read as foreign and never refresh again. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- scripts/check-browser-test-excludes.mjs | 45 +++++++++++++-- scripts/git-hooks.mjs | 19 ++++++- test/check-browser-test-excludes.test.ts | 24 ++++++++ test/git-hooks.test.ts | 70 ++++++++++++++++++++---- 4 files changed, 137 insertions(+), 21 deletions(-) diff --git a/scripts/check-browser-test-excludes.mjs b/scripts/check-browser-test-excludes.mjs index 3d3b6233..44617f53 100644 --- a/scripts/check-browser-test-excludes.mjs +++ b/scripts/check-browser-test-excludes.mjs @@ -63,17 +63,26 @@ function walk(dir) { } /** - * Every `*.test.ts` under `<root>/test` that imports a browser driver, as sorted - * repo-relative POSIX paths (the form `vitest list` prints). + * Every `*.test.ts` under `<root>/test`, as sorted repo-relative POSIX paths (the form + * `vitest list` prints). + * + * @param {string} root + * @returns {string[]} + */ +export function findTestFiles(root) { + return walk(join(root, 'test')) + .map((file) => relative(root, file).split(sep).join('/')) + .sort(); +} + +/** + * The subset of {@link findTestFiles} that imports a browser driver. * * @param {string} root * @returns {string[]} */ export function findBrowserTests(root) { - return walk(join(root, 'test')) - .filter((file) => importsBrowserDriver(readFileSync(file, 'utf8'))) - .map((file) => relative(root, file).split(sep).join('/')) - .sort(); + return findTestFiles(root).filter((file) => importsBrowserDriver(readFileSync(join(root, file), 'utf8'))); } /** @@ -93,6 +102,19 @@ export function parseVitestFileList(output) { ); } +/** + * Whether the `vitest list` paths and the walked tree name at least one file in common. + * False means the two sides are not speaking the same path format (absolute paths, backslashes + * or a new prefix after a vitest upgrade), and then {@link findLeaks} would find nothing + * against a perfectly non-empty listing. + * + * @param {Set<string>} ciFiles + * @param {string[]} testFiles + */ +export function listingMatchesTree(ciFiles, testFiles) { + return testFiles.some((file) => ciFiles.has(file)); +} + /** * @param {string[]} browserTests * @param {Set<string>} ciFiles @@ -103,6 +125,7 @@ export function findLeaks(browserTests, ciFiles) { } function main() { + const testFiles = findTestFiles(ROOT); const browserTests = findBrowserTests(ROOT); let collected; @@ -124,6 +147,16 @@ function main() { console.error('✗ `vitest list` reported no test files; refusing to pass on an empty CI set.'); process.exit(1); } + // Same vacuous pass, one step removed: a listing whose paths never match the tree. This guard, + // not `vitest list --json`, is the answer to format drift: the JSON form prints absolute paths + // that would need canonicalizing against ROOT (symlinked checkouts), and its shape can drift too. + if (!listingMatchesTree(ciFiles, testFiles)) { + const sample = [...ciFiles].slice(0, 3).join(', '); + console.error( + `✗ none of the ${ciFiles.size} paths \`vitest list\` reported (e.g. ${sample}) is one of the ${testFiles.length} test/**/*.test.ts files; its output format has probably changed.` + ); + process.exit(1); + } const leaked = findLeaks(browserTests, ciFiles); diff --git a/scripts/git-hooks.mjs b/scripts/git-hooks.mjs index a0b5770d..4727a5fe 100644 --- a/scripts/git-hooks.mjs +++ b/scripts/git-hooks.mjs @@ -28,7 +28,11 @@ import { execFileSync } from 'node:child_process'; import { chmodSync, existsSync, mkdirSync, readFileSync, realpathSync, writeFileSync } from 'node:fs'; import { basename, dirname, join, resolve } from 'node:path'; -/** Ownership marker. Bump the version suffix when the body changes meaningfully. */ +/** + * Ownership marker. ⚠️ Never bump the version suffix: ownership is matched on this exact + * string, so a `v2` would read every installed `v1` hook as foreign and never refresh it. + * A changed body still reaches installed hooks, because the refresh compares the whole file. + */ export const PRE_PUSH_MARKER = '# codeman-managed-hook: pre-push v1'; /** @@ -52,8 +56,10 @@ export const PRE_PUSH_CHECKS = [ * check:frontend-syntax), config/ (eslint + vitest configs, test-suites.ts, the CLI * catalogue), scripts/ (every check is a script there, and typecheck's second pass compiles * one), test/ (check:browser-excludes scans it and runs `vitest list` over it), - * package.json + package-lock.json (check:lockfile) and install.sh (generate:cli-catalog - * --check diffs its generated block). + * package.json + package-lock.json (check:lockfile), install.sh (generate:cli-catalog + * --check diffs its generated block), tsconfig.json (typecheck, and + * config/tsconfig.scripts.json extends it) and .prettierignore + .editorconfig + * (format:check; the Prettier CLI honours .editorconfig by default). */ export const PRE_PUSH_WATCHED_PATHS = [ 'src', @@ -63,6 +69,9 @@ export const PRE_PUSH_WATCHED_PATHS = [ 'package.json', 'package-lock.json', 'install.sh', + 'tsconfig.json', + '.prettierignore', + '.editorconfig', ]; /** @@ -95,6 +104,10 @@ if [ ! -d node_modules ]; then exit 0 fi +# GUI git clients and IDEs often run hooks with a minimal PATH that lacks an nvm or +# Homebrew Node. Every check would then fail with "npm: not found", so skip instead. +command -v npm >/dev/null 2>&1 || { echo "pre-push: npm not on PATH, skipping checks."; exit 0; } + # git feeds us "<localref> <localsha> <remoteref> <remotesha>" per ref. A deletion has an # all-zero local sha and no tree worth checking; if every ref is a deletion, skip. # The checks below read the working tree, so they only say something about a pushed commit diff --git a/test/check-browser-test-excludes.test.ts b/test/check-browser-test-excludes.test.ts index 7dc34ad8..de522068 100644 --- a/test/check-browser-test-excludes.test.ts +++ b/test/check-browser-test-excludes.test.ts @@ -11,7 +11,9 @@ import { join, resolve } from 'node:path'; import { findBrowserTests, findLeaks, + findTestFiles, importsBrowserDriver, + listingMatchesTree, parseVitestFileList, } from '../scripts/check-browser-test-excludes.mjs'; import { BROWSER_TEST_GLOBS } from '../config/test-suites'; @@ -68,6 +70,15 @@ describe('findBrowserTests (fixture tree)', () => { 'test/new.browser.test.ts', ]); }); + + it('lists every test file, browser-driven or not, in the same form', () => { + expect(findTestFiles(root)).toEqual([ + 'test/legacy-name.test.ts', + 'test/nested/deep.test.ts', + 'test/new.browser.test.ts', + 'test/unit.test.ts', + ]); + }); }); describe('parseVitestFileList + findLeaks', () => { @@ -83,6 +94,19 @@ describe('parseVitestFileList + findLeaks', () => { ]); expect(findLeaks(['test/new.browser.test.ts'], ci)).toEqual([]); }); + + it('flags a non-empty listing whose paths never match the tree instead of passing vacuously', () => { + const tree = ['test/legacy-name.test.ts', 'test/unit.test.ts']; + // e.g. a vitest upgrade that starts printing absolute paths: nothing leaks, but only + // because nothing matches, so the checker must refuse rather than report success. + const drifted = parseVitestFileList('/repo/test/legacy-name.test.ts\n/repo/test/unit.test.ts\n'); + expect(drifted.size).toBe(2); + expect(findLeaks(['test/legacy-name.test.ts'], drifted)).toEqual([]); + expect(listingMatchesTree(drifted, tree)).toBe(false); + + const healthy = parseVitestFileList('test/legacy-name.test.ts\ntest/unit.test.ts\n'); + expect(listingMatchesTree(healthy, tree)).toBe(true); + }); }); describe('against this repository', () => { diff --git a/test/git-hooks.test.ts b/test/git-hooks.test.ts index 8ba70e87..3e00bc1e 100644 --- a/test/git-hooks.test.ts +++ b/test/git-hooks.test.ts @@ -24,6 +24,7 @@ import { realpathSync, rmSync, statSync, + symlinkSync, writeFileSync, } from 'node:fs'; import { tmpdir } from 'node:os'; @@ -133,6 +134,24 @@ describe('planHookInstall', () => { }); describe('resolveGitHooksDir (temp repos)', () => { + // resolveGitHooksDir runs git with process.env, so an exported GIT_CONFIG_GLOBAL or a system + // gitconfig carrying core.hooksPath would otherwise redirect every expectation below. + // test/setup.ts swaps HOME, which only covers ~/.gitconfig. + const ambient = { + GIT_CONFIG_GLOBAL: process.env.GIT_CONFIG_GLOBAL, + GIT_CONFIG_NOSYSTEM: process.env.GIT_CONFIG_NOSYSTEM, + }; + beforeAll(() => { + process.env.GIT_CONFIG_NOSYSTEM = '1'; + process.env.GIT_CONFIG_GLOBAL = '/dev/null'; + }); + afterAll(() => { + for (const [k, v] of Object.entries(ambient)) { + if (v === undefined) delete process.env[k]; + else process.env[k] = v; + } + }); + it('resolves <root>/.git/hooks in a plain checkout', () => { const repo = newRepo(); expect(resolveGitHooksDir(repo)).toBe(join(repo, '.git', 'hooks')); @@ -352,18 +371,24 @@ describe('the installed hook on a real push (temp repos)', () => { expect(ran()).toEqual(expectedRuns); }); - it.each(['src/wip.ts', 'config/wip.json', 'scripts/wip.mjs', 'test/wip.test.ts', 'install.sh'])( - 'skips when %s is untracked (another session may own it)', - (rel) => { - const { ran, push, repo } = setup({ failing: 'lint' }); - mkdirSync(join(repo, rel, '..'), { recursive: true }); - writeFileSync(join(repo, rel), 'wip\n'); - const r = push(['origin', 'main']); - expect(r.status, r.stderr).toBe(0); - expect(r.stdout + r.stderr).toContain('pre-push: skipping static checks: uncommitted changes under'); - expect(ran()).toEqual([]); - } - ); + it.each([ + 'src/wip.ts', + 'config/wip.json', + 'scripts/wip.mjs', + 'test/wip.test.ts', + 'install.sh', + 'tsconfig.json', + '.prettierignore', + '.editorconfig', + ])('skips when %s is untracked (another session may own it)', (rel) => { + const { ran, push, repo } = setup({ failing: 'lint' }); + mkdirSync(join(repo, rel, '..'), { recursive: true }); + writeFileSync(join(repo, rel), 'wip\n'); + const r = push(['origin', 'main']); + expect(r.status, r.stderr).toBe(0); + expect(r.stdout + r.stderr).toContain('pre-push: skipping static checks: uncommitted changes under'); + expect(ran()).toEqual([]); + }); it('skips when a tracked package.json has an unstaged edit', () => { const { ran, push, repo } = setup({ failing: 'lint' }); @@ -395,6 +420,9 @@ describe('the installed hook on a real push (temp repos)', () => { 'package.json', 'package-lock.json', 'install.sh', + 'tsconfig.json', + '.prettierignore', + '.editorconfig', ]); }); @@ -405,6 +433,24 @@ describe('the installed hook on a real push (temp repos)', () => { expect(r.stdout + r.stderr).toContain('node_modules missing'); expect(ran()).toEqual([]); }); + + it('skips (never blocks) when npm is not on PATH, as under a GUI git client', () => { + const { ran, push } = setup({ failing: 'lint' }); + // A PATH holding only what git and the hook need, and no npm/node. Symlinks rather than + // the real directories, since /usr/bin usually holds npm right next to git. + const bin = join(scratch, `bin-${counter}`); + mkdirSync(bin); + for (const tool of ['git', 'sh', 'mktemp', 'tail', 'rm', 'cat']) { + const found = spawnSync('sh', ['-c', `command -v ${tool}`], { encoding: 'utf8' }).stdout.trim(); + expect(found, `${tool} not found on the test PATH`).toMatch(/^\//); + symlinkSync(found, join(bin, tool)); + } + expect(spawnSync('sh', ['-c', 'command -v npm'], { env: { PATH: bin } }).status).not.toBe(0); + const r = push(['origin', 'main'], { PATH: bin }); + expect(r.status, r.stderr + r.stdout).toBe(0); + expect(r.stdout + r.stderr).toContain('pre-push: npm not on PATH, skipping checks.'); + expect(ran()).toEqual([]); + }); }); describe('postinstall wiring', () => { From 714050fe8a1164c76af511f2763d2801e638492a Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:26:00 +0200 Subject: [PATCH 27/46] fix(terminal): merge-time fixes for #494 - Skip and latch a bounded Shell window once the browser is at xterm's scrollback cap (scrollback + rows): a 1 MiB window of short lines can carry more rows than the browser can ever hold, so it replayed and re-captured on every scroll-to-top with no 60 s back-off. - Label a replayed bounded window 'tail' even when the capture was byte-capped, so the banner keeps offering Load full history instead of calling the rest unrecoverable. - Pin GET /terminal?full=1&tail=<n> in the route tests: full-history source, truncationReason 'tail', and the closing relative cursor move survive the cut. - Log the bounded skip via _logScrollRouting('repull-skipped-bounded'). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- src/web/public/app.js | 24 +++++++-- test/routes/session-routes.test.ts | 47 ++++++++++++++++++ test/shell-scroll-history-pull.test.ts | 67 +++++++++++++++++++++++++- test/terminal-scroll-routing.test.ts | 2 +- 4 files changed, 134 insertions(+), 6 deletions(-) diff --git a/src/web/public/app.js b/src/web/public/app.js index a8bda99f..56997f21 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -6379,18 +6379,29 @@ class CodemanApp { // is left as the load that produced it set it: re-labelling it from this // payload would call a terminal that holds ALL of a Load full history pull // "the most recent 1 MiB". - if (boundedShellPull && windowRows <= this.terminal.buffer.active.length) { + // + // A browser already at xterm's cap buys nothing either. xterm keeps at most + // `scrollback + rows` rows (DEFAULT_SCROLLBACK 50k) while tmux keeps 100k + // lines by default, so a 1 MiB window of short lines can render to more rows + // than the browser can ever hold, and `windowRows <= rowsNow` then never + // comes true: without this every scroll-to-top would reset and re-parse it. + const rowsNow = this.terminal.buffer.active.length; + const scrollbackCap = this.terminal.options?.scrollback || 0; + const browserFull = scrollbackCap > 0 && rowsNow >= scrollbackCap + this.terminal.rows; + if (boundedShellPull && (windowRows <= rowsNow || browserFull)) { // An untruncated window IS all of tmux's history, so nothing is missing, // and the next burst of output can put more in tmux than the browser has: // keep the normal 4 s cooldown. A truncated one is the opposite case, since // the gesture can never reach anything older than what the browser already // shows, and every ask costs the server a synchronous capture-pane of the // whole history (`tail` is applied after the capture): back off to 60 s. + // A full browser backs off too, since no window can ever fit in it. // Trade-off: only a successful replay clears that latch, so a tab switch or // burst that shrinks the browser's buffer below the window can leave a // scroll-to-top inert for up to a minute. Load full history (`force`) // bypasses the cooldown, and the latch is bounded, never permanent. - if (payload.truncated) (this._fullHistoryRepullUseless ||= new Set()).add(sessionId); + if (payload.truncated || browserFull) (this._fullHistoryRepullUseless ||= new Set()).add(sessionId); + this._logScrollRouting?.('repull-skipped-bounded'); return; } if (this._replayWouldShrinkBuffer(buffer, windowRows)) { @@ -6404,7 +6415,14 @@ class CodemanApp { this._setHistoryTruncation(sessionId, { ...payload, exhausted: true }); return; } - this._setHistoryTruncation(sessionId, payload); + // A bounded window that was cut is always recoverable: a capture over the + // byte cap keeps `truncationReason: 'capped'` through the tail cut, and that + // would tell the user the rest "cannot be recovered" and drop Load full + // history, whose unbounded pull returns up to the cap itself. + this._setHistoryTruncation( + sessionId, + boundedShellPull && payload.truncated ? { ...payload, truncationReason: 'tail' } : payload + ); this._fullHistoryRepullUseless?.delete(sessionId); const rowsBefore = this.terminal.buffer.active.length; const replayStartedAt = performance.now(); diff --git a/test/routes/session-routes.test.ts b/test/routes/session-routes.test.ts index 2db9dd1d..cc9c9422 100644 --- a/test/routes/session-routes.test.ts +++ b/test/routes/session-routes.test.ts @@ -920,6 +920,53 @@ describe('session-routes', () => { expect(res.headers['server-timing']).toMatch(/^capture;dur=\d+\.\d, prepare;dur=\d+\.\d, total;dur=\d+\.\d$/); }); + it('full reload with a tail (?full=1&tail=) cuts the full capture to its newest bytes, cursor restore intact', async () => { + // A Shell scroll-to-top asks for exactly this (`_maybeRefetchFullHistory`): + // tmux's whole scrollback, bounded to the tab-switch tail size. The client + // relies on all three answers below, so a refactor that dropped the tail on + // a full capture (an unbounded pull from an ordinary scroll) or cut off the + // closing cursor move (a caret parked below the prompt) must fail here. + const tail = 1024 * 1024; + const oldestMarker = 'BOUNDED_OLDEST_LINE_00001'; + const newestMarker = 'BOUNDED_NEWEST_LINE_40000'; + const rows: string[] = [oldestMarker]; + for (let i = 2; i < 40_000; i++) rows.push(`shell history line ${String(i).padStart(5, '0')} lorem ipsum`); + rows.push(newestMarker); + // What formatCursorRestore appends: up from the last row, then the column. + const cursorRestore = '\x1b[3A\r\x1b[2C'; + const fullHistoryCapture = `${rows.join('\r\n')}${cursorRestore}`; + expect(fullHistoryCapture.length).toBeGreaterThan(tail); + + harness.ctx._session.mode = 'shell'; + harness.ctx._session.terminalBuffer = ''; + const captureSpy = vi.fn((_name: string, opts?: { fullHistory?: boolean }) => + opts?.fullHistory ? fullHistoryCapture : 'only the visible frame' + ); + (harness.ctx.mux as { captureActivePaneBuffer?: unknown }).captureActivePaneBuffer = captureSpy; + + const res = await harness.app.inject({ + method: 'GET', + url: `/api/sessions/${harness.ctx._sessionId}/terminal?full=1&tail=${tail}`, + }); + + expect(res.statusCode).toBe(200); + const body = JSON.parse(res.body); + // Still the scrollback, not the visible frame a plain `?tail=` gets. + expect(captureSpy).toHaveBeenCalledWith( + harness.ctx._session.muxName, + expect.objectContaining({ fullHistory: true }) + ); + expect(body.data.source).toBe('mux-full-history'); + // Recoverable, not 'capped': Load full history can still bring the rest back. + expect(body.data.truncated).toBe(true); + expect(body.data.truncationReason).toBe('tail'); + expect(body.data.fullSize).toBe(fullHistoryCapture.length); + expect(body.data.terminalBuffer.length).toBeLessThanOrEqual(tail); + expect(body.data.terminalBuffer).toContain(newestMarker); + expect(body.data.terminalBuffer).not.toContain(oldestMarker); + expect(body.data.terminalBuffer.endsWith(`${newestMarker}${cursorRestore}`)).toBe(true); + }); + it('full reload (?full=1) returns the tmux capture ALONE — byte history is not duplicated', async () => { // The full-history capture is the rendered form of everything already in // the byte buffer; prepending the byte history would replay the whole diff --git a/test/shell-scroll-history-pull.test.ts b/test/shell-scroll-history-pull.test.ts index ac9de9b0..10b730aa 100644 --- a/test/shell-scroll-history-pull.test.ts +++ b/test/shell-scroll-history-pull.test.ts @@ -12,7 +12,9 @@ * * The gesture now pulls a BOUNDED window (`?full=1&tail=TERMINAL_TAIL_SIZE`), * the button stays the unbounded path, and a window the browser already holds - * in full is not rewritten. + * in full is not rewritten. Neither is one a browser at xterm's scrollback cap + * could never hold, and a window cut from a byte-capped capture is still labelled + * recoverable, since Load full history can reach past it. * * ORDER MATTERS: that skip must run BEFORE the downgrade guard. The guard reads * "smaller than the browser" as "tmux has nothing more to give", which is true of @@ -104,7 +106,12 @@ const TAIL_CUT = { function makeApp( mode: string, - { bufferRows, capture, payload = {} }: { bufferRows: number; capture: string; payload?: Record<string, unknown> } + { + bufferRows, + capture, + payload = {}, + scrollback = 0, + }: { bufferRows: number; capture: string; payload?: Record<string, unknown>; scrollback?: number } ) { const urls: string[] = []; const app = { @@ -119,6 +126,8 @@ function makeApp( terminal: { cols: 80, rows: 30, + // xterm's scrollback option; 0 leaves the browser-cap check out of a test. + options: { scrollback }, buffer: { active: { length: bufferRows } }, scrollToLine: vi.fn(), scrollToTop: vi.fn(), @@ -183,6 +192,60 @@ describe('shell scroll-up pulls a bounded window of tmux history', () => { expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); // Nothing was written, so the banner state is left exactly as it was. expect(app._setHistoryTruncation).not.toHaveBeenCalled(); + // …but the skip is visible to someone diagnosing "scroll-to-top does nothing". + expect(app._logScrollRouting).toHaveBeenCalledWith('repull-skipped-bounded'); + }); + + it("a browser at xterm's scrollback cap stops replaying a window it can never hold, and backs off", async () => { + // xterm keeps at most `scrollback + rows` rows while tmux keeps 100k lines, so + // a 1 MiB window of short lines can carry more rows than the browser ever will. + // `windowRows <= rows held` then never comes true, and every scroll-to-top + // past the cooldown reset and re-parsed the window. Untruncated on purpose: + // the back-off has to come from the full browser, not from `truncated`. + const { app } = makeApp('shell', { bufferRows: 40, capture: lines(2000), scrollback: 1000 }); + + // The first pull has room to grow, so it replays. + await refetch.call(app); + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + expect(app._fullHistoryRepullUseless.has('s1')).toBe(false); + // xterm kept only the last `scrollback + rows` of the 2000 rows written. + app.terminal.buffer.active.length = 1000 + 30; + + // Past the 4 s cooldown: the browser is full, so nothing is replayed and the + // session backs off for a minute. + app._fullHistoryRepullAt.set('s1', Date.now() - 5000); + await refetch.call(app); + expect(app._fetchTerminalCapture).toHaveBeenCalledTimes(2); + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + expect(app.chunkedTerminalWrite).toHaveBeenCalledTimes(1); + expect(app._fullHistoryRepullUseless.has('s1')).toBe(true); + + // So a scroll 10 s later does not even ask the server for another capture. + app._fullHistoryRepullAt.set('s1', Date.now() - 10_000); + await refetch.call(app); + expect(app._fetchTerminalCapture).toHaveBeenCalledTimes(2); + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + }); + + it('a replayed window cut from a byte-capped capture still offers Load full history', async () => { + // The route keeps `truncationReason: 'capped'` through the tail cut when the + // full capture exceeded the byte cap. On a bounded window that is not "gone for + // good": the unbounded pull behind the button returns up to the cap itself. + const capped = { ...TAIL_CUT, truncationReason: 'capped', fullSize: 40 * 1024 * 1024 }; + const { app } = makeApp('shell', { bufferRows: 40, capture: lines(300), payload: capped }); + + await refetch.call(app); + + expect(app._resetTerminalForReplay).toHaveBeenCalledTimes(1); + expect(app._setHistoryTruncation).toHaveBeenCalledWith('s1', expect.objectContaining({ truncationReason: 'tail' })); + const notice = computeNotice(app._historyTruncation.get('s1') as Record<string, unknown>); + expect(notice.visible).toBe(true); + expect(notice.canLoadMore).toBe(true); + + // The button's own unbounded pull is the one place 'capped' is the truth. + const button = makeApp('shell', { bufferRows: 40, capture: lines(300), payload: capped }); + await refetch.call(button.app, { force: true }); + expect(computeNotice(button.app._historyTruncation.get('s1') as Record<string, unknown>).canLoadMore).toBe(false); }); }); diff --git a/test/terminal-scroll-routing.test.ts b/test/terminal-scroll-routing.test.ts index d2dd206a..511736a8 100644 --- a/test/terminal-scroll-routing.test.ts +++ b/test/terminal-scroll-routing.test.ts @@ -114,7 +114,7 @@ describe('full-history re-pull downgrade guard (issue #205 round 2)', () => { // Also anchored on the open paren: the guard is handed the rows the caller // already estimated, and this test is about ORDER, not the argument list. const guard = source.indexOf('this._replayWouldShrinkBuffer(buffer', start); - const boundedSkip = source.indexOf('boundedShellPull && windowRows <=', start); + const boundedSkip = source.indexOf('boundedShellPull && (windowRows <= rowsNow || browserFull)', start); const reset = source.indexOf('this._resetTerminalForReplay()', start); expect(start).toBeGreaterThan(-1); From 1f4c390e124b6f6707782c96225ced3b2ca4e36c Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:30:13 +0200 Subject: [PATCH 28/46] fix(terminal): merge-time fixes for #498 - _logScrollRouting() reports cliMouseTracking, the gate's new input, in both the de-dup signature and the console line (xterm's own mouseTracking stays 'none' for Claude, so it gave no reason for a no). - Restore two guard tests the new gate made vacuous: the local-scrollback opt-out footgun test and the codex/gemini "no version rescues it" fixtures now set cliMouseTracking: true, so removing the opt-out or re-adding codex to the gate fails again. - Update the comments and architecture-invariants lines that still described the version-only rule (wheel handler header, gate doc, the false paths of _maybePageCliTranscript, "holds a tracking mode on continuously"). - Name both fullscreen switches (CLAUDE_CODE_NO_FLICKER=1 and "tui": "fullscreen" in ~/.claude/settings.json) in the code comment, the invariants and the two wiki pages. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- docs/architecture-invariants.md | 8 +++--- docs/wiki/The-Dashboard.md | 3 +- docs/wiki/Troubleshooting.md | 4 ++- src/web/public/terminal-ui.js | 42 +++++++++++++++++----------- test/terminal-scroll-routing.test.ts | 18 ++++++++++-- test/terminal-touch-tap.test.ts | 6 ++-- 6 files changed, 54 insertions(+), 27 deletions(-) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 810708a7..848ae8f1 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -215,13 +215,13 @@ Further detail: the `<prefix>: <title>` form (`w3-myapp: fix the login redirect` **Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity/pi) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend mirror (`_shouldReportMouseToCli()`) stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`. -⚠️ **What the full strip removes, it must REMEMBER.** Stripping the mouse DECSETs means xterm's `modes.mouseTrackingMode` is permanently `'none'` for those modes, so the browser hand-encodes click reports to compensate (`_sendSyntheticSgrTap`). With no state to consult it did that on EVERY click, which delivered mouse reports to programs that never asked for them: the same pane runs a plain shell whenever the CLI has exited or a `shell` was started inside a claude-mode session, and a shell prints the report as literal text (`[<0;88;20M`), garbling the next line typed. `_recordStrippedMouseMode()` therefore records each stripped sequence as it goes and publishes `cliMouseTracking` through `toState()`, and `_shouldReportMouseToCli()` requires it. ⚠️ Only the TRACKING modes count (1000/1001/1002/1003): 1005/1006 select an ENCODING and 1007 is alt-scroll, and counting those would put the stray reports straight back. ⚠️ The change broadcasts IMMEDIATELY rather than through `broadcastSessionStateDebounced`, because the flag flips when a dialog opens and the user can click that dialog inside the 500ms debounce window. Measured on a live claude 2.x: the CLI holds a tracking mode on continuously (so clicks keep being reported exactly as before), while a bash prompt in the same stripped mode reports nothing. Fails toward silence: after a server restart the flag is false until the CLI re-emits, which tmux does at client attach. +⚠️ **What the full strip removes, it must REMEMBER.** Stripping the mouse DECSETs means xterm's `modes.mouseTrackingMode` is permanently `'none'` for those modes, so the browser hand-encodes click reports to compensate (`_sendSyntheticSgrTap`). With no state to consult it did that on EVERY click, which delivered mouse reports to programs that never asked for them: the same pane runs a plain shell whenever the CLI has exited or a `shell` was started inside a claude-mode session, and a shell prints the report as literal text (`[<0;88;20M`), garbling the next line typed. `_recordStrippedMouseMode()` therefore records each stripped sequence as it goes and publishes `cliMouseTracking` through `toState()`, and `_shouldReportMouseToCli()` requires it. ⚠️ Only the TRACKING modes count (1000/1001/1002/1003): 1005/1006 select an ENCODING and 1007 is alt-scroll, and counting those would put the stray reports straight back. ⚠️ The change broadcasts IMMEDIATELY rather than through `broadcastSessionStateDebounced`, because the flag flips when a dialog opens and the user can click that dialog inside the 500ms debounce window. Measured on a live claude 2.x: the CLI holds a tracking mode on continuously in fullscreen (so clicks keep being reported exactly as before; the default inline renderer holds none, measured on 2.1.283 even with `/model` open), while a bash prompt in the same stripped mode reports nothing. Fails toward silence: after a server restart the flag is false until the CLI re-emits, which tmux does at client attach. -**Only claude ≥ 2.1.187 with mouse tracking on forwards the wheel; everything else scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. Claude repeats this exactly in its default INLINE renderer (measured on 2.1.280: `alternate_on=0`, `mouse_any_flag=0`, `history_size` grows), and swipes on iOS Safari were dead there while codex scrolled; only fullscreen claude (`CLAUDE_CODE_NO_FLICKER=1`: alt screen plus modes 1003/1006) pages its transcript on wheel reports. So claude forwards only while the server-observed `cliMouseTracking` flag is true; a stale-false flag after a server restart falls through to the PageUp/PageDown fallback, never a dead wheel. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS are gated by `_shouldReportMouseToCli()` (strip mode AND the server-observed `cliMouseTracking` flag, recorded by `_recordStrippedMouseMode` in session.ts as it strips): codex never enables mouse tracking, so since #325 no tap report is sent there at all — click-to-position was already a measured no-op in codex, and a pane that has fallen back to a shell no longer receives `[<0;88;20M` junk. +**Only claude ≥ 2.1.187 with mouse tracking on forwards the wheel; everything else scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. Claude repeats this exactly in its default INLINE renderer (measured on 2.1.280: `alternate_on=0`, `mouse_any_flag=0`, `history_size` grows), and swipes on iOS Safari were dead there while codex scrolled; only fullscreen claude (`CLAUDE_CODE_NO_FLICKER=1` or `"tui": "fullscreen"` in `~/.claude/settings.json`: alt screen plus modes 1003/1006) pages its transcript on wheel reports. So claude forwards only while the server-observed `cliMouseTracking` flag is true; a stale-false flag after a server restart falls through to the PageUp/PageDown fallback, never a dead wheel. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS are gated by `_shouldReportMouseToCli()` (strip mode AND the server-observed `cliMouseTracking` flag, recorded by `_recordStrippedMouseMode` in session.ts as it strips): codex never enables mouse tracking, so since #325 no tap report is sent there at all — click-to-position was already a measured no-op in codex, and a pane that has fallen back to a shell no longer receives `[<0;88;20M` junk. -**Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`. +**Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (claude ≥ 2.1.187 while `cliMouseTracking` is true, i.e. fullscreen; version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`. -**A false gate on a Claude session must not mean a DEAD gesture** (#205 round 2, `_maybePageCliTranscript`): every way `_shouldForwardWheelToApp()` returns false leaves a repaint-mode pane scrolling a buffer that has nothing in it (`baseY === 0`) — the version probe came back empty, the CLI really is older than 2.1.187, or the user turned on `terminalWheelLocalScrollback`. The 1.12.0 retest reported exactly that: a wheel that did nothing at all while Fn+Up (PageUp) paged back through intact text, which is the proof that the CLI's own history and the PTY input path were both fine. So under the triple guard (claude mode + gate false + `baseY === 0`) wheel and touch travel is translated into coalesced `\x1b[5~` / `\x1b[6~` through the same 40ms queue as the SGR reports, at half a screen of travel per page key (the key jumps a whole screen; a 1:1 mapping was unusably slow with a discrete wheel). ⚠️ Shift is excluded on purpose — it is the explicit "give me local scrollback" gesture and must keep that meaning. ⚠️ `terminalWheelLocalScrollback` is deliberately NOT scoped away from repaint-mode CLIs even though it is a footgun there: that would silently override an explicit user choice, so the fallback catches it instead. **Server-side counterpart**: `getClaudeCliVersion()` caches SUCCESS for the process lifetime but must never cache FAILURE — it used to, so one timed-out or PATH-starved probe at the first Claude session start disabled wheel-forwarding for every Claude session until the server restarted (a dead wheel on phone, tablet and laptop at once, the signature of a server-side cause). Failures now retry with a 1/2/4…15min backoff; the policy is the pure `resolveClaudeCliVersion()`. Tests: `test/terminal-scroll-routing.test.ts`, `test/claude-cli-version-cache.test.ts`. +**A false gate on a Claude session must not mean a DEAD gesture** (#205 round 2, `_maybePageCliTranscript`): every way `_shouldForwardWheelToApp()` returns false leaves a repaint-mode pane scrolling a buffer that has nothing in it (`baseY === 0`) — the version probe came back empty, the CLI really is older than 2.1.187, the `cliMouseTracking` flag is unset (inline claude, or fullscreen right after a server restart), or the user turned on `terminalWheelLocalScrollback`. The 1.12.0 retest reported exactly that: a wheel that did nothing at all while Fn+Up (PageUp) paged back through intact text, which is the proof that the CLI's own history and the PTY input path were both fine. So under the triple guard (claude mode + gate false + `baseY === 0`) wheel and touch travel is translated into coalesced `\x1b[5~` / `\x1b[6~` through the same 40ms queue as the SGR reports, at half a screen of travel per page key (the key jumps a whole screen; a 1:1 mapping was unusably slow with a discrete wheel). ⚠️ Shift is excluded on purpose — it is the explicit "give me local scrollback" gesture and must keep that meaning. ⚠️ `terminalWheelLocalScrollback` is deliberately NOT scoped away from repaint-mode CLIs even though it is a footgun there: that would silently override an explicit user choice, so the fallback catches it instead. **Server-side counterpart**: `getClaudeCliVersion()` caches SUCCESS for the process lifetime but must never cache FAILURE — it used to, so one timed-out or PATH-starved probe at the first Claude session start disabled wheel-forwarding for every Claude session until the server restarted (a dead wheel on phone, tablet and laptop at once, the signature of a server-side cause). Failures now retry with a 1/2/4…15min backoff; the policy is the pure `resolveClaudeCliVersion()`. Tests: `test/terminal-scroll-routing.test.ts`, `test/claude-cli-version-cache.test.ts`. **Why the wheel went where it went is LOGGED** (`_logScrollRouting`): one console line per session per distinct decision — `[scroll] <id> → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…, cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. #205 ran two rounds of remote guesswork over questions this line answers directly; keep it when touching the routing. diff --git a/docs/wiki/The-Dashboard.md b/docs/wiki/The-Dashboard.md index 501d85b5..408294ae 100644 --- a/docs/wiki/The-Dashboard.md +++ b/docs/wiki/The-Dashboard.md @@ -155,7 +155,8 @@ Worth knowing: history; press **Load full history** to pull the rest explicitly. Automatic output recovery stays within the bounded browser buffer. - **Wheel and touch scrolling** are forwarded into Claude's own transcript when a recent - Claude runs fullscreen, so the wheel scrolls the conversation rather than the terminal. + Claude runs fullscreen (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in + `~/.claude/settings.json`), so the wheel scrolls the conversation rather than the terminal. Claude's default inline view keeps its history in the terminal and scrolls locally. `Shift+Wheel` is always local scrollback. Other CLIs scroll locally. - **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not. diff --git a/docs/wiki/Troubleshooting.md b/docs/wiki/Troubleshooting.md index 39b6754c..41c64ef4 100644 --- a/docs/wiki/Troubleshooting.md +++ b/docs/wiki/Troubleshooting.md @@ -166,7 +166,9 @@ Things to try: - `Shift+Wheel` always scrolls the local buffer, whatever else is going on. - On Claude sessions running fullscreen (recent CLI with mouse tracking on), the wheel is forwarded into Claude's own transcript, so it scrolls the conversation rather than the - terminal buffer. That is intended. Claude's default inline view scrolls locally. + terminal buffer. That is intended. Claude's default inline view scrolls locally; turn + fullscreen on with `CLAUDE_CODE_NO_FLICKER=1` or `"tui": "fullscreen"` in + `~/.claude/settings.json`. - Scrolling to the very top pulls the full tmux scrollback again on demand. ### The wheel does nothing in a Codex session diff --git a/src/web/public/terminal-ui.js b/src/web/public/terminal-ui.js index e7aab353..657e1624 100644 --- a/src/web/public/terminal-ui.js +++ b/src/web/public/terminal-ui.js @@ -650,8 +650,9 @@ Object.assign(CodemanApp.prototype, { } // Mouse wheel: forward to the TUI only for sessions verified to handle SGR - // wheel reports (claude 2.1.187+ — see _shouldForwardWheelToApp), local - // scrollback otherwise. Claude Code 2.1.187+ scrolls its own + // wheel reports (claude 2.1.187+ while it tracks the mouse, which only its + // fullscreen renderer does; see _shouldForwardWheelToApp), local scrollback + // otherwise. Claude Code 2.1.187+ scrolls its own // transcript on SGR wheel reports — scrolled-away tool blocks re-render // live and stay clickable — and its select menus no longer capture wheel // as option navigation (verified against 2.1.202: /model menu highlight @@ -723,7 +724,7 @@ Object.assign(CodemanApp.prototype, { // phone/tablet swipe scrolls the local buffer of stale repaint frames and // drags the CLI's pinned input box off the screen (issue #205's mobile // half). Same gate, so Shift has no touch analog but the local-scrollback - // opt-out setting and the CLI-version gate apply to touch exactly as they + // opt-out setting and the version/tracking gate apply to touch exactly as they // do to the wheel — including the PageUp/PageDown fallback the wheel uses // when that gate is false and there is no local scrollback to scroll // (_maybePageCliTranscript), which is what keeps a swipe from being a @@ -5086,9 +5087,10 @@ Object.assign(CodemanApp.prototype, { // Wheel forwarding gate for the container wheel handler: no Shift override, // xterm's own encoder dormant, viewport at the bottom, and a TUI VERIFIED to - // scroll its transcript on SGR wheel reports — which today is claude 2.1.187+ - // and nothing else (older Claude Code captures wheel as select-menu option - // navigation; an unknown version is treated as older). Gemini and codex are + // scroll its transcript on SGR wheel reports, which today is claude 2.1.187+ + // with mouse tracking on (fullscreen) and nothing else (older Claude Code + // captures wheel as select-menu option navigation, an unknown version is + // treated as older, and inline Claude ignores it). Gemini and codex are // strip modes too but keep the local wheel — taps/clicks are still forwarded // for them (harmless no-ops at worst). // @@ -5159,9 +5161,10 @@ Object.assign(CodemanApp.prototype, { // inline renderer (2.1.280 measured: alternate_on=0, mouse_any_flag=0) the // transcript lives in real scrollback, like codex, and SGR wheel reports are // ignored, so forwarding made every swipe and wheel tick dead. Fullscreen - // (CLAUDE_CODE_NO_FLICKER=1) turns on alt-screen + mode 1003/1006, which the - // server records as cliMouseTracking. A stale-false flag after a server - // restart falls through to _maybePageCliTranscript, so it never goes dead. + // (CLAUDE_CODE_NO_FLICKER=1, or "tui": "fullscreen" in ~/.claude/settings.json) + // turns on alt-screen + mode 1003/1006, which the server records as + // cliMouseTracking. A stale-false flag after a server restart falls through + // to _maybePageCliTranscript, so it never goes dead. if (session?.cliMouseTracking !== true) return false; // Deliberately NOT gated on _terminalViewportAtBottom(). It used to be, so // that leaving the bottom handed the wheel back to local scrollback and both @@ -5236,11 +5239,13 @@ Object.assign(CodemanApp.prototype, { * * The rescue path for every way `_shouldForwardWheelToApp` can come back false * on a Claude session that has no local history to fall back on: the CLI - * version probe failed or is genuinely older than 2.1.187, or the user turned - * on "Wheel scrolls local history" (which pins the wheel to a buffer that, - * for a repaint-mode CLI, is empty — the setting's footgun). Before this, all - * of those produced a completely dead gesture; the #205 reporter proved the - * keyboard route works by paging back through intact text with Fn+Up. + * version probe failed or is genuinely older than 2.1.187, the CLI's mouse + * tracking flag is unset (the inline renderer, or fullscreen right after a + * server restart), or the user turned on "Wheel scrolls local history" (which + * pins the wheel to a buffer that, for a repaint-mode CLI, is empty: the + * setting's footgun). Before this, all of those produced a completely dead + * gesture; the #205 reporter proved the keyboard route works by paging back + * through intact text with Fn+Up. * * Triple-guarded (claude mode + gate false + `baseY === 0`), so a session with * real local scrollback is never touched. Shift is excluded on purpose: it is @@ -5284,14 +5289,19 @@ Object.assign(CodemanApp.prototype, { const session = this.sessions?.get(sessionId); const optOut = !!this.loadAppSettingsFromStorage?.()?.terminalWheelLocalScrollback; const tracking = this.terminal?.modes?.mouseTrackingMode || 'none'; + // xterm's own mode above stays 'none' for a strip mode (the server removes + // the DECSETs), so the CLI's real tracking state is reported separately. + const cliTracking = session?.cliMouseTracking === true; const baseY = this.terminal?.buffer?.active?.baseY ?? -1; - const signature = `${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${baseY > 0}`; + const signature = + `${decision}|${session?.mode}|${session?.cliVersion}|${optOut}|${tracking}|${cliTracking}|${baseY > 0}`; if (!this._scrollRoutingLogged) this._scrollRoutingLogged = new Map(); if (this._scrollRoutingLogged.get(sessionId) === signature) return; this._scrollRoutingLogged.set(sessionId, signature); console.log( `[scroll] ${sessionId} → ${decision} (mode=${session?.mode || '?'}, cliVersion=${session?.cliVersion || 'unknown'}, ` + - `localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, localScrollbackRows=${baseY})` + `localScrollbackOptOut=${optOut}, mouseTracking=${tracking}, cliMouseTracking=${cliTracking}, ` + + `localScrollbackRows=${baseY})` ); }, diff --git a/test/terminal-scroll-routing.test.ts b/test/terminal-scroll-routing.test.ts index 511736a8..2321d8fc 100644 --- a/test/terminal-scroll-routing.test.ts +++ b/test/terminal-scroll-routing.test.ts @@ -47,11 +47,13 @@ function loadTerminalUiHarness() { } /** A Claude session whose local buffer holds exactly one screen (baseY 0). */ -function hollowClaudeApp(overrides: { cliVersion?: string; rows?: number } = {}) { +function hollowClaudeApp(overrides: { cliVersion?: string; rows?: number; cliMouseTracking?: boolean } = {}) { const { app, logs } = loadTerminalUiHarness(); const sent: Array<{ id: string; data: string }> = []; app.activeSessionId = 'sess-1'; - app.sessions = new Map([['sess-1', { mode: 'claude', cliVersion: overrides.cliVersion }]]); + app.sessions = new Map([ + ['sess-1', { mode: 'claude', cliVersion: overrides.cliVersion, cliMouseTracking: overrides.cliMouseTracking }], + ]); app._sendInputEphemeral = (id: string, data: string) => sent.push({ id, data }); app.terminal = { cols: 80, @@ -194,7 +196,8 @@ describe('PageUp/PageDown fallback for a hollow local buffer (issue #205 round 2 // "Wheel scrolls local history" ON pins the wheel to a buffer that, for a // repaint-mode CLI, is empty — a user who flipped it while hunting for a fix // on 1.11.x would have ended up with a completely dead wheel on 1.12.0. - const { app, sent } = hollowClaudeApp({ cliVersion: '2.1.223' }); // gate would forward… + // Version and tracking both qualify, so the opt-out is the only thing saying no. + const { app, sent } = hollowClaudeApp({ cliVersion: '2.1.223', cliMouseTracking: true }); // gate would forward… app.loadAppSettingsFromStorage = () => ({ terminalWheelLocalScrollback: true }); expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); // …but the opt-out wins @@ -236,10 +239,19 @@ describe('scroll routing diagnostic (issue #205 round 2)', () => { expect(logs[0]).toContain('cliVersion=2.1.100'); expect(logs[0]).toContain('localScrollbackOptOut=false'); expect(logs[0]).toContain('mouseTracking=none'); + // The gate's real tracking input: xterm's own mode above is always 'none' + // for Claude, since the server strips the DECSETs. + expect(logs[0]).toContain('cliMouseTracking=false'); app._logScrollRouting('page-keys'); // a changed route still prints expect(logs).toHaveLength(2); expect(logs[1]).toContain('page-keys'); + + // The CLI turning tracking on changes the gate, so it prints again. + app.sessions.get('sess-1').cliMouseTracking = true; + app._logScrollRouting('page-keys'); + expect(logs).toHaveLength(3); + expect(logs[2]).toContain('cliMouseTracking=true'); }); it('reports an unknown CLI version, the false-path that disables forwarding', () => { diff --git a/test/terminal-touch-tap.test.ts b/test/terminal-touch-tap.test.ts index eacd815d..f35106b4 100644 --- a/test/terminal-touch-tap.test.ts +++ b/test/terminal-touch-tap.test.ts @@ -683,10 +683,12 @@ describe('terminal touch tap mouse guard', () => { // (the codex transcript lives there — inline viewport, no in-app pager) sat unused. app.sessions = new Map([['sess-1', { mode: 'codex' }]]); expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); - app.sessions = new Map([['sess-1', { mode: 'codex', cliVersion: '9.9.9' }]]); // no version rescues it + // Tracking on and a high version, so only the mode check can say no: without + // them the gate is false for claude too and this would pin nothing. + app.sessions = new Map([['sess-1', { mode: 'codex', cliVersion: '9.9.9', cliMouseTracking: true }]]); // no version rescues it expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); - app.sessions = new Map([['sess-1', { mode: 'gemini', cliVersion: '9.9.9' }]]); // unverified TUI + app.sessions = new Map([['sess-1', { mode: 'gemini', cliVersion: '9.9.9', cliMouseTracking: true }]]); // unverified TUI expect(app._shouldForwardWheelToApp({ shiftKey: false })).toBe(false); }); From e439cf0ef3229b880c58e856db53847d3e05b758 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:30:56 +0200 Subject: [PATCH 29/46] chore: changeset for the 2026-09-28 landing Folds the #490 and #492 contributor changesets (the latter said minor) into one patch changeset with the Thanks section, one paragraph per change and the fixes applied while landing. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- .changeset/cli9a1d.md | 5 ----- .changeset/land-2026-09-28.md | 31 +++++++++++++++++++++++++++++++ .changeset/static-git-identity.md | 5 ----- 3 files changed, 31 insertions(+), 10 deletions(-) delete mode 100644 .changeset/cli9a1d.md create mode 100644 .changeset/land-2026-09-28.md delete mode 100644 .changeset/static-git-identity.md diff --git a/.changeset/cli9a1d.md b/.changeset/cli9a1d.md deleted file mode 100644 index 0ecbf8df..00000000 --- a/.changeset/cli9a1d.md +++ /dev/null @@ -1,5 +0,0 @@ ---- -"aicodeman": patch ---- - -Docker Compose: CLIs installed from Settings (DeepSeek, Pi and any other npm-based CLI) survive `Update-Codeman.sh`. The image's `NPM_CONFIG_PREFIX` (`/opt/codeman-cli`) is image content and was discarded when the container was recreated; `POST /api/clis/:id/install` now installs into `~/.local` on the persistent home mount when running in the container, and `~/.local/bin` is on the image PATH. diff --git a/.changeset/land-2026-09-28.md b/.changeset/land-2026-09-28.md new file mode 100644 index 00000000..f6e28fd7 --- /dev/null +++ b/.changeset/land-2026-09-28.md @@ -0,0 +1,31 @@ +--- +"aicodeman": patch +--- + +### Thanks + +- @aakhter for keeping web-tab events private to their owner in multi-user mode (#501), with end-to-end isolation tests that fail without the fix, and for the browser-test exclusion check and pre-push hook (#500), including the hooks-dir resolution that never writes outside the repo's own `.git/hooks`. +- @opticon454 for keeping CLIs installed from Settings across Docker container updates (#490) and for the static Git identity for the Docker images (#492). +- @timkjr for letting a Shell pane's scroll-to-top reach tmux history (#494), with tests that fail on the commit before each fix. +- @JDProfresh for tracking down why wheel and touch scrolling did nothing in Claude's default inline view (#498), with the tmux measurements that proved it. +- @irisitymichaelgrundberg for the follow-up that makes an agent waiting on artifact comments raise its alert again (#491). + +**A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends. + +**Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback. + +**Shell panes scroll back into tmux history (#494).** Scrolling to the top of a Shell pane now pulls the most recent 1 MiB of its tmux history, so output that arrived in a burst is reachable without pressing **Load full history**, which still loads the rest. + +**An agent waiting on artifact comments alerts again (#491).** A session whose agent published an artifact and is waiting for somebody to comment on it now raises the normal idle alert and lands in NEEDS YOU, instead of being treated as busy with background work. + +**Web-tab changes stay private in multi-user mode (#501).** The `webview:changed` event reached every connected user, exposing the ids of other users' web-tab creates, edits and deletes. It now carries the tab's owner and reaches that owner plus admins only. Single-user mode is unchanged apart from a new optional `owner` field on the event. + +**Phone header tabs look like tabs (#504).** On phones every header tab is now a chip with a fill and a border, the Alt+N digit (a keyboard hint a phone cannot use) is hidden, names get 80px instead of 50px, and the strip fades at whichever edge still has tabs scrolled out of view. + +**Docker: CLIs installed from Settings survive container updates (#490).** On the Compose deployment, CLIs installed from App Settings (DeepSeek, Pi and other npm-based CLIs) now go to `~/.local` on the persistent home mount, and `~/.local/bin` is on the image PATH, so recreating the container no longer discards them. Anything installed from Settings before this release has to be installed once more after the rebuild. + +**Docker: a static Git identity for the server and agent images (#492).** Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` (or `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` / `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` on a bare-host install) and the identity is written to `/etc/gitconfig` in the server image and the Docker-case agent image. A half-set pair is refused on every build path. Both Docker changes edit `server.Dockerfile`, so the in-app updater asks Compose deployments to rebuild with `docker/Start-Codeman.sh` instead of updating in place. + +**Contributor tooling (#500).** `npm run check:browser-excludes`, now a CI step, fails when a test that drives a real browser is still collected by `npm test`. `npm install` also installs a pre-push hook that runs the static CI checks before a push; it steps aside when the pushed ref is not HEAD or the tree has uncommitted changes the checks would read, and `CODEMAN_SKIP_PREPUSH=1 git push` skips it once. + +**Fixes applied while landing.** A Shell pane's scroll-to-top (#494) no longer re-pulls the same window on every gesture once the browser's 50,000-row scrollback is full, and a pull that hit the byte cap no longer claims the older history is gone. The scroll-routing diagnostics (#498) now log whether Claude has mouse tracking on. The artifact-comment check (#491) also refuses a footer cut off in the middle of the chip. The pre-push hook (#500) steps aside when `npm` is not on PATH, as in some GUI git clients, instead of blocking every push. The Docker Git identity error (#492) names the two variables to set. New tests pin the image PATH order, the identity on both agent-image build paths, and the `?full=1&tail=` terminal route. diff --git a/.changeset/static-git-identity.md b/.changeset/static-git-identity.md deleted file mode 100644 index 345e124a..00000000 --- a/.changeset/static-git-identity.md +++ /dev/null @@ -1,5 +0,0 @@ ---- -"aicodeman": minor ---- - -Docker deployments can configure a static Git commit identity for server and Docker-case agent images. From c2dfc775a35381c66d184b7fcd19a7117eb34721 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:44:59 +0200 Subject: [PATCH 30/46] fix(input): deliver API prompts through tmux so their Enter is not lost A prompt posted to /api/sessions/:id/input without `useMux` was written into the pane in one piece. Claude Code (measured on 2.1.283) takes a `<text>\r` burst of about a hundred characters or more as a paste, so the trailing `\r` landed as a newline in the composer and the prompt sat there unsent while the route answered 200. A later raw `\r` did not recover it; a tmux `send-keys Enter` did. Short prompts submitted, which is why it looked random. The same stranding was seen with Codex and OpenCode. A plain prompt (printable text plus exactly one trailing `\r`, detected by `isPlainPromptInput()`) now goes through `writeViaMux` even without `useMux`: the text is typed, Enter is pressed as its own key, and the SubmitVerifier re-presses it while the prompt is still on the composer. The write is awaited, since the browser's POST fallback sends frames one at a time and a following keystroke must not overtake the Enter. Raw frames (escape sequences, bracketed paste, a line feed, a bare `\r`) and an explicit `useMux: false` keep the direct write. Verified on an isolated instance: the 239- and 104-character prompts that stranded (at +1 s, at +50 s on ultracode, and on a warm session) all submitted on the first Enter with no `useMux`. The phone's local-echo flow (a burst, then its `\r` as a separate write) was measured unaffected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/api-reference.md | 9 ++ src/web/route-helpers.ts | 20 ++++ src/web/routes/session-routes.ts | 21 +++- test/routes/session-input-delivery.test.ts | 106 ++++++++++++++++++++- 5 files changed, 155 insertions(+), 3 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 6c741be6..cf395318 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -128,7 +128,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph ## Common Gotchas -- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`. ⚠️ **Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints** while still taking the typed text (measured 2026-09-19: an Enter at 28 s stranded the prompt, one at 51 s submitted it), so text+`\r` sent at readiness sits unsent with `0 tokens` and a `wait` burns its timeout. So the SERVER verifies every programmatic write that carried a `\r`: `SubmitVerifier` (`session-submit-verifier.ts`, armed from `writeViaMux`) reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the LAST composer line (the CLI's own `promptGlyph`) verifiably still holds the head of what was sent; an empty composer, other text, or no composer line at all (a shell, a direct-PTY session) ends it, and a newer write replaces the schedule. The skill's `sendwait` keeps its own copy of the loop (`_composer_text` in `skills/codeman/preamble.sh`) for servers that predate this. The `shift+tab` footer only means the composer painted, never that Enter is accepted +- **Single-line prompts only** — `writeViaMux()` sends text+Enter separately; multi-line breaks Ink. ⚠️ **Input must END with `\r` or Enter is never sent**: `sendInput()` only issues `send-keys Enter` when the payload contains a carriage return, a `\r`-less `POST /api/sessions/:id/input` still succeeds (send-and-wait even reports `delivered:true`) while the text sits unsubmitted on the composer, and any `wait` burns its whole timeout on a turn that never started. Embedded newlines are stripped, not rejected, so `"echo A\necho B\r"` runs the joined `echo Aecho B`. ⚠️ **Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints** while still taking the typed text (measured 2026-09-19: an Enter at 28 s stranded the prompt, one at 51 s submitted it), so text+`\r` sent at readiness sits unsent with `0 tokens` and a `wait` burns its timeout. So the SERVER verifies every programmatic write that carried a `\r`: `SubmitVerifier` (`session-submit-verifier.ts`, armed from `writeViaMux`) reads the pane on a 2 s to 60 s schedule and re-sends Enter only while the LAST composer line (the CLI's own `promptGlyph`) verifiably still holds the head of what was sent; an empty composer, other text, or no composer line at all (a shell, a direct-PTY session) ends it, and a newer write replaces the schedule. The skill's `sendwait` keeps its own copy of the loop (`_composer_text` in `skills/codeman/preamble.sh`) for servers that predate this. The `shift+tab` footer only means the composer painted, never that Enter is accepted. ⚠️ **A prompt must never be written into the pane as ONE burst**: Claude Code 2.1.283 takes a `<text>\r` burst of ~100+ chars as a paste, its `\r` lands as a NEWLINE and the prompt strands (a later raw `\r` does not recover it, a tmux `send-keys Enter` does). So `POST .../input` routes a plain prompt (`isPlainPromptInput()`, route-helpers.ts: printable text + exactly one trailing `\r`) through `writeViaMux` even without `useMux`, AWAITED so the browser's serialized POST fallback keeps frame order; raw frames and an explicit `useMux:false` keep the direct write - **ESM only** — Never `require()`, use `await import()`. `tsx` masks CJS/ESM issues in dev but production breaks - **Package ≠ product name** — npm: `aicodeman`, product: **Codeman**. Release renames tags accordingly. Both `aicodeman` and `codeman` bin aliases are installed (`package.json` `bin`) - **Global regex `lastIndex`** — Shared `g`-flag patterns in loops must reset `lastIndex = 0` first, or use the `execPattern()` helper in `utils/regex-patterns.ts` (resets automatically) diff --git a/docs/api-reference.md b/docs/api-reference.md index bc520de2..c0f7235d 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -311,6 +311,15 @@ worker's prompt but never submitted, and the wait then runs its full timeout on turn that never started. Verified live; this is the most common silent failure on this endpoint. +A **plain prompt** (printable text followed by exactly one `\r`, nothing else) is +delivered through tmux even without `useMux`: the text is typed, Enter is pressed as +a separate key, and the server re-presses Enter while the prompt is still visibly +sitting on the composer. Written straight into the pane in one piece, a prompt of +about a hundred characters or more is taken as a paste by Claude Code, its `\r` +becomes a newline, and the prompt stays unsent (measured on 2.1.283). Any other +input (escape sequences, a bracketed-paste frame, a line feed, a bare `\r`) keeps +the raw write, and an explicit `"useMux": false` forces it. + ```bash curl -s -X POST "$API/api/v1/sessions/$SID/input" \ -H 'Content-Type: application/json' \ diff --git a/src/web/route-helpers.ts b/src/web/route-helpers.ts index 9acd6f71..08e1f683 100644 --- a/src/web/route-helpers.ts +++ b/src/web/route-helpers.ts @@ -437,6 +437,26 @@ export function parseBody<T>(schema: z.ZodType<T>, body: unknown, errorMessage?: return result.data; } +/** + * Whether an input body is a plain prompt: printable text followed by exactly one + * carriage return, and nothing else. + * + * That is the shape a script, a bot or a curl call sends to submit a prompt, and the + * one that must NOT be written into the pane in one piece. Measured on Claude Code + * 2.1.283 (2026-09-28): a direct write of `<text>\r` arrives as a single burst, and a + * burst of about a hundred characters or more is taken as a paste, so its trailing + * `\r` lands as a NEWLINE in the composer and the prompt sits there unsent. A later + * bare `\r` written the same way does not recover it; a tmux `send-keys Enter` does. + * Short bursts (tens of characters) submit, which is why the failure looked random. + * + * Anything with another control character (escape sequences, a bracketed-paste frame, + * a line feed, a tab, C1 controls) is raw terminal input and keeps the direct write. + */ +export function isPlainPromptInput(input: string): boolean { + // eslint-disable-next-line no-control-regex -- matching control characters is the point + return /^[^\x00-\x1f\x7f-\x9f]+\r$/.test(input); +} + /** * Persist session state and broadcast a SessionUpdated event. * Replaces the repeated two-line pattern across route handlers. diff --git a/src/web/routes/session-routes.ts b/src/web/routes/session-routes.ts index fe0d05e8..6a7b78ad 100644 --- a/src/web/routes/session-routes.ts +++ b/src/web/routes/session-routes.ts @@ -91,6 +91,7 @@ import { findSessionOrFail, getAuthUser, isAdmin, + isPlainPromptInput, isWorkingDirAllowed, ownerFor, parseBody, @@ -1835,10 +1836,18 @@ export function registerSessionRoutes( // the wrong recovery — wait longer, when the truth is "restart the worker". let delivered = false; + // A plain prompt (`<text>\r`) goes through the mux even when the caller did not + // ask for it: written straight into the pane it arrives as one burst, and Claude + // Code takes a long burst as a paste whose `\r` becomes a newline, so the prompt + // sat unsent (see isPlainPromptInput). The mux path types the text, presses Enter + // separately and arms the SubmitVerifier. An explicit `useMux: false` keeps the + // raw write for a caller that really wants it. + const autoMux = useMux === undefined && isPlainPromptInput(inputStr); + if (duplicate) { // Redelivery of an already-applied input: skip the write, but still honor the // wait, since the caller's question ("tell me when this settles") is unanswered. - } else if (useMux && waitPromise) { + } else if ((useMux || autoMux) && waitPromise) { // The response is already staying open for the wait, so the tmux write can be // awaited here. This is the ONE path where a writeViaMux failure is observable. const ok = await session.writeViaMux(inputStr, { fromUser: true }).catch(() => false); @@ -1849,6 +1858,16 @@ export function registerSessionRoutes( delivered = session.write(inputStr, { fromUser: true }); if (!delivered) undoOnFailure(); } + } else if (autoMux) { + // Awaited, unlike the explicit useMux branch below. This shape also reaches here + // from the browser's POST fallback, which sends its frames one at a time and + // waits for each 2xx; answering only once Enter has gone out is what keeps the + // next keystroke from overtaking it. + const ok = await session.writeViaMux(inputStr, { fromUser: true }).catch(() => false); + if (!ok) { + console.warn(`[Server] writeViaMux failed for session ${id}, falling back to direct write`); + if (!session.write(inputStr, { fromUser: true })) undoOnFailure(); + } } else if (useMux) { // Fire-and-forget: don't block the HTTP response on a tmux child process. // Fallback to a direct write on failure. Unchanged from before send-and-wait. diff --git a/test/routes/session-input-delivery.test.ts b/test/routes/session-input-delivery.test.ts index 2d655ffd..248a359f 100644 --- a/test/routes/session-input-delivery.test.ts +++ b/test/routes/session-input-delivery.test.ts @@ -14,11 +14,12 @@ import fastifyCookie from '@fastify/cookie'; import Fastify, { type FastifyInstance } from 'fastify'; -import { afterEach, beforeEach, describe, expect, it } from 'vitest'; +import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'; import { Session } from '../../src/session.js'; import { ApiErrorCode, httpStatusForErrorCode } from '../../src/types.js'; import { installRouteErrorHandler } from '../../src/web/route-error-handler.js'; +import { isPlainPromptInput } from '../../src/web/route-helpers.js'; import { registerSessionRoutes } from '../../src/web/routes/session-routes.js'; import { createMockRouteContext, type MockRouteContext } from '../mocks/index.js'; @@ -155,3 +156,106 @@ describe('POST /api/sessions/:id/input rollback wiring', () => { expect(session.shouldApplyInput('c2', 5)).toBe(true); }); }); + +/** + * A plain prompt goes through the mux even when the caller did not say `useMux`. + * + * Measured on Claude Code 2.1.283: a direct write of `<text>\r` arrives as one burst, + * a burst of about a hundred characters is taken as a paste, and its `\r` lands as a + * newline in the composer, so a script's prompt sat there unsent while the route + * answered 200. The mux path types the text, presses Enter on its own, and arms the + * submit verifier. + */ +describe('POST /api/sessions/:id/input plain-prompt routing', () => { + let harness: { app: FastifyInstance; ctx: MockRouteContext }; + + beforeEach(async () => { + harness = await createEnvelopeHarness(); + }); + afterEach(async () => { + await harness.app.close(); + }); + + const post = (body: Record<string, unknown>) => + harness.app.inject({ method: 'POST', url: '/api/sessions/test-session-1/input', payload: body }); + const spies = () => { + const session = harness.ctx.sessions.get('test-session-1')!; + return { session, viaMux: vi.spyOn(session, 'writeViaMux'), direct: vi.spyOn(session, 'write') }; + }; + const LONG_PROMPT = + 'Reply with only the word ok and nothing else, this sentence is padding to reach about one hundred chars.\r'; + + it('sends a prompt with no useMux through the mux, not as one burst', async () => { + const { viaMux, direct } = spies(); + + const res = await post({ input: LONG_PROMPT }); + + expect(res.statusCode).toBe(200); + expect(viaMux).toHaveBeenCalledWith(LONG_PROMPT, { fromUser: true }); + expect(direct).not.toHaveBeenCalled(); + }); + + it('answers only once the mux write is done, so the next frame cannot overtake it', async () => { + // The browser's POST fallback sends frames one at a time and waits for each 2xx; + // a fire-and-forget write here would let its next keystroke land before the Enter. + const { viaMux } = spies(); + let finished = false; + viaMux.mockImplementation(async () => { + await new Promise((r) => setTimeout(r, 30)); + finished = true; + return true; + }); + + await post({ input: 'ok\r', clientId: 'browser-1', seq: 1 }); + + expect(finished).toBe(true); + }); + + it('falls back to the direct write when the mux write fails', async () => { + const { viaMux, direct } = spies(); + viaMux.mockResolvedValue(false); + + await post({ input: 'hello\r' }); + + expect(direct).toHaveBeenCalledWith('hello\r', { fromUser: true }); + }); + + it('keeps the raw write for an explicit useMux: false', async () => { + const { viaMux, direct } = spies(); + + await post({ input: LONG_PROMPT, useMux: false }); + + expect(direct).toHaveBeenCalledWith(LONG_PROMPT, { fromUser: true }); + expect(viaMux).not.toHaveBeenCalled(); + }); + + it.each([ + ['a bare Enter', '\r'], + ['text with no Enter', 'hello'], + ['an arrow key', '\x1b[A'], + ['a bracketed paste frame', '\x1b[200~line one\nline two\x1b[201~'], + ['a line feed inside', 'line one\nline two\r'], + ['two Enters', 'hello\r\r'], + ['a tab', 'a\tb\r'], + ])('leaves %s on the direct write', async (_label, input) => { + const { viaMux, direct } = spies(); + + await post({ input }); + + expect(direct).toHaveBeenCalledWith(input, { fromUser: true }); + expect(viaMux).not.toHaveBeenCalled(); + }); +}); + +describe('isPlainPromptInput', () => { + it('accepts printable text ending in exactly one carriage return', () => { + expect(isPlainPromptInput('run the tests\r')).toBe(true); + expect(isPlainPromptInput('ünïcødé and emoji 🚀\r')).toBe(true); + }); + + it('refuses anything carrying another control character', () => { + for (const input of ['\r', 'x', 'x\n', 'x\r\n', 'x\r\r', '\x1b[Ax\r', 'a\tb\r', 'x\x7f\r', 'x\u009b\r', '\rx']) { + expect(isPlainPromptInput(input), JSON.stringify(input)).toBe(false); + } + }); +}); From 4d165d3fb134a93cfb4c9d953a00547bc8c9b307 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 16:46:04 +0200 Subject: [PATCH 31/46] chore: add the input-delivery fix to the landing changeset Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- .changeset/land-2026-09-28.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/.changeset/land-2026-09-28.md b/.changeset/land-2026-09-28.md index f6e28fd7..079ba52b 100644 --- a/.changeset/land-2026-09-28.md +++ b/.changeset/land-2026-09-28.md @@ -12,6 +12,8 @@ **A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends. +**Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. + **Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback. **Shell panes scroll back into tmux history (#494).** Scrolling to the top of a Shell pane now pulls the most recent 1 MiB of its tmux history, so output that arrived in a burst is reachable without pressing **Load full history**, which still loads the rest. From 2eece4f8f900a38dcf96fec487db04e66458d272 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 17:10:12 +0200 Subject: [PATCH 32/46] fix(cron): send a paste-mode prompt's Enter as its own write A cron job in "Paste (direct)" input mode wrote `<text>\r` into the pane in one piece. Claude Code (measured on 2.1.283) takes a burst of about a hundred characters as a paste, so the `\r` landed as a newline and the prompt sat unsent on the composer while the run reported `prompt_sent`. Delivery now lives in `deliverCronPrompt()`. Paste mode writes the text raw, waits CRON_PASTE_ENTER_DELAY_MS (300 ms), sends `\r` as a separate write down the same PTY (so it cannot overtake the text), and arms the session's composer check through the new public `Session.verifySubmitted()`, which re-presses Enter while the prompt is still visibly unsent. A session with nothing to write to now fails the run instead of reporting the prompt as sent. Typed mode is unchanged. Verified on an isolated instance: a paste-mode job with a 104-character prompt submitted on the first Enter and Claude answered. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- src/config/server-timing.ts | 8 ++++ src/cron/cron-service.ts | 47 +++++++++++++++++---- src/session.ts | 10 +++++ test/cron-service.test.ts | 82 ++++++++++++++++++++++++++++++++++++- 4 files changed, 137 insertions(+), 10 deletions(-) diff --git a/src/config/server-timing.ts b/src/config/server-timing.ts index d4e98ff2..7a9c1858 100644 --- a/src/config/server-timing.ts +++ b/src/config/server-timing.ts @@ -99,3 +99,11 @@ export const STALE_DATA_MAX_AGE_MS = 60 * 60 * 1000; /** Standard 5-minute inactivity timeout for streams and caches (ms) */ export const INACTIVITY_TIMEOUT_MS = 5 * 60 * 1000; + +/** + * Gap between a paste-mode cron prompt's text and its Enter (ms). The two must be + * separate writes: Claude Code takes a raw `<text>\r` burst of about a hundred + * characters as a paste and turns its `\r` into a newline. A separate `\r` 80 ms + * after the text was measured to submit; this leaves room for a longer prompt. + */ +export const CRON_PASTE_ENTER_DELAY_MS = 300; diff --git a/src/cron/cron-service.ts b/src/cron/cron-service.ts index cce014ee..7a57424c 100644 --- a/src/cron/cron-service.ts +++ b/src/cron/cron-service.ts @@ -20,7 +20,7 @@ import { getErrorMessage, createErrorResponse, ApiErrorCode } from '../types/api import { MAX_CONCURRENT_SESSIONS, MAX_CRON_JOBS, MAX_CRON_RUN_HISTORY } from '../config/map-limits.js'; import { canUsernameRunPrivilegedCommands, resolveClaudeModeForUsername } from '../user-store.js'; import { sessionCapacityState, isWorkingDirAllowedForUsername } from '../web/route-helpers.js'; -import { CRON_READY_MAX_ATTEMPTS, CRON_READY_SETTLE_MS } from '../config/server-timing.js'; +import { CRON_PASTE_ENTER_DELAY_MS, CRON_READY_MAX_ATTEMPTS, CRON_READY_SETTLE_MS } from '../config/server-timing.js'; import { DEFAULT_BLOCKED_TREES, isBlockedAttachmentPath, @@ -100,6 +100,41 @@ const CRON_WORKING_DIR_BLOCKED_TREES: readonly string[] = [...DEFAULT_BLOCKED_TR /** Prompt delivery is single-line only (writeViaMux/Ink constraint). */ const HAS_NEWLINE = /[\r\n]/; +/** The three session calls prompt delivery needs, so it can be tested without a PTY. */ +type CronPromptTarget = Pick<Session, 'write' | 'writeViaMux' | 'verifySubmitted'>; + +/** + * Send a cron job's (single-line) prompt into its session and press Enter. + * + * `typed` goes through the mux: the text is typed, Enter is its own key, and the + * session re-presses it while the prompt is still on the composer. + * + * `paste` writes the text straight into the PTY, and must send its Enter as a + * SEPARATE write. It used to send `<text>\r` in one piece, and Claude Code (measured + * on 2.1.283) takes a burst of about a hundred characters as a paste, so the `\r` + * landed as a newline and the prompt sat unsent while the run reported + * `prompt_sent`. The Enter goes down the same PTY as the text, so it cannot overtake + * it, and the same composer check then covers a CLI that was not taking Enter yet. + * + * @returns false when the session had no PTY or mux to write to + */ +export async function deliverCronPrompt( + target: CronPromptTarget, + prompt: string, + inputMode: CronJob['inputMode'], + wait: (ms: number) => Promise<void> = delay +): Promise<boolean> { + if (inputMode !== 'paste') { + return target.writeViaMux(prompt.endsWith('\r') ? prompt : `${prompt}\r`); + } + const text = prompt.replace(/[\r\n]+$/, ''); + if (!target.write(text)) return false; + await wait(CRON_PASTE_ENTER_DELAY_MS); + if (!target.write('\r')) return false; + target.verifySubmitted(text); + return true; +} + /** Order-insensitive equality for the weekly-days arrays. */ function sameDays(a: number[] | undefined, b: number[] | undefined): boolean { const x = [...(a ?? [])].sort((p, q) => p - q); @@ -623,15 +658,9 @@ export class CronService { const s = this.deps.sessions.get(sessionId); if (!s) return; try { - const payload = prompt.endsWith('\r') ? prompt : `${prompt}\r`; - let delivered = true; - if (job.inputMode === 'paste') { - s.write(payload); - } else { - delivered = await s.writeViaMux(payload); - } + const delivered = await deliverCronPrompt(s, prompt, job.inputMode); if (!delivered) { - this.failRun(job, run, 'Failed to send prompt: mux write failed'); + this.failRun(job, run, 'Failed to send prompt: the session could not be written to'); return; } run.status = 'prompt_sent'; diff --git a/src/session.ts b/src/session.ts index e5da2586..3a26e90e 100644 --- a/src/session.ts +++ b/src/session.ts @@ -4207,6 +4207,16 @@ export class Session extends EventEmitter { return false; } + /** + * Arm the composer check for a prompt that went out some other way than + * `writeViaMux`, e.g. cron's paste mode, which writes the body raw and its Enter + * separately. `text` is what the composer line starts with while the prompt is still + * unsent; the check re-presses Enter only while that holds. + */ + verifySubmitted(text: string): void { + this._verifySubmitted(`${text}\r`); + } + /** * Arm the composer check for a write that carried Enter (session-submit-verifier.ts): * Claude Code 2.1.277+ ignores Enter for the first 30-50 s after the composer paints, diff --git a/test/cron-service.test.ts b/test/cron-service.test.ts index cf10f37d..8b36bd32 100644 --- a/test/cron-service.test.ts +++ b/test/cron-service.test.ts @@ -16,7 +16,13 @@ import { describe, it, expect, beforeEach, vi } from 'vitest'; import { existsSync, mkdtempSync, mkdirSync, writeFileSync, symlinkSync } from 'node:fs'; import { tmpdir } from 'node:os'; import { join } from 'node:path'; -import { CronService, clampCronExternalCliConfigs, type CronDeps } from '../src/cron/cron-service.js'; +import { + CronService, + clampCronExternalCliConfigs, + deliverCronPrompt, + type CronDeps, +} from '../src/cron/cron-service.js'; +import { CRON_PASTE_ENTER_DELAY_MS } from '../src/config/server-timing.js'; import { CronJobSchema } from '../src/web/schemas.js'; import { MAX_CRON_JOBS } from '../src/config/map-limits.js'; import type { CronJob, CronJobRun } from '../src/types/cron.js'; @@ -705,3 +711,77 @@ describe('clampCronExternalCliConfigs', () => { } }); }); + +/** + * Paste mode used to write `<text>\r` in one piece. Claude Code takes a raw burst of + * about a hundred characters as a paste and turns its `\r` into a newline, so the + * prompt sat unsent on the composer while the run said `prompt_sent`. + */ +describe('deliverCronPrompt', () => { + const PROMPT = + 'Reply with only the word ok and nothing else, this sentence is padding to reach about one hundred chars.'; + + function fakeTarget(ok = true) { + const calls: string[] = []; + const target = { + write: vi.fn((d: string) => { + calls.push(`write:${JSON.stringify(d)}`); + return ok; + }), + writeViaMux: vi.fn(async (d: string) => { + calls.push(`mux:${JSON.stringify(d)}`); + return ok; + }), + verifySubmitted: vi.fn((t: string) => { + calls.push(`verify:${JSON.stringify(t)}`); + }), + }; + return { target, calls }; + } + const noWait = async (): Promise<void> => {}; + + it('paste mode writes the text and its Enter separately, then arms the composer check', async () => { + const { target, calls } = fakeTarget(); + const waits: number[] = []; + + const ok = await deliverCronPrompt(target, PROMPT, 'paste', async (ms) => { + waits.push(ms); + calls.push('wait'); + }); + + expect(ok).toBe(true); + expect(calls).toEqual([ + `write:${JSON.stringify(PROMPT)}`, + 'wait', + 'write:"\\r"', + `verify:${JSON.stringify(PROMPT)}`, + ]); + expect(waits).toEqual([CRON_PASTE_ENTER_DELAY_MS]); + expect(target.writeViaMux).not.toHaveBeenCalled(); + }); + + it('never puts the Enter in the same write as the text', async () => { + const { target } = fakeTarget(); + + await deliverCronPrompt(target, `${PROMPT}\r`, 'paste', noWait); + + for (const [data] of target.write.mock.calls) { + expect(data === '\r' || !data.includes('\r')).toBe(true); + } + }); + + it('typed mode is unchanged: one mux write that carries the Enter', async () => { + const { target, calls } = fakeTarget(); + + await deliverCronPrompt(target, PROMPT, 'typed', noWait); + + expect(calls).toEqual([`mux:${JSON.stringify(`${PROMPT}\r`)}`]); + }); + + it('reports a session it could not write to, instead of claiming the prompt went out', async () => { + const { target } = fakeTarget(false); + + expect(await deliverCronPrompt(target, PROMPT, 'paste', noWait)).toBe(false); + expect(target.verifySubmitted).not.toHaveBeenCalled(); + }); +}); From 0b106b03eb1b214e6388f9744919626ea38a9a3b Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 17:25:00 +0200 Subject: [PATCH 33/46] chore: add the cron paste-mode fix to the landing changeset Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- .changeset/land-2026-09-28.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/.changeset/land-2026-09-28.md b/.changeset/land-2026-09-28.md index 079ba52b..6085253e 100644 --- a/.changeset/land-2026-09-28.md +++ b/.changeset/land-2026-09-28.md @@ -12,7 +12,7 @@ **A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends. -**Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. +**Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. The same fix reaches cron jobs in "Paste (direct)" input mode, which reported `prompt_sent` for a prompt that never left the composer: the text is written raw, Enter follows as its own write 300 ms later, and the session presses it again while the prompt is still unsent. A cron run with no session to write to now fails instead of reporting the prompt as sent. **Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback. From 848ab48b0afcc9006df5090da309317f6a76b455 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Mon, 28 Sep 2026 17:35:45 +0200 Subject: [PATCH 34/46] chore: version packages (1.33.2) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- .changeset/land-2026-09-28.md | 33 ---------------------- .claude-plugin/marketplace.json | 2 +- CHANGELOG.md | 33 ++++++++++++++++++++++ CLAUDE.md | 2 +- package-lock.json | 4 +-- package.json | 2 +- plugins/codeman/.claude-plugin/plugin.json | 2 +- 7 files changed, 39 insertions(+), 39 deletions(-) delete mode 100644 .changeset/land-2026-09-28.md diff --git a/.changeset/land-2026-09-28.md b/.changeset/land-2026-09-28.md deleted file mode 100644 index 6085253e..00000000 --- a/.changeset/land-2026-09-28.md +++ /dev/null @@ -1,33 +0,0 @@ ---- -"aicodeman": patch ---- - -### Thanks - -- @aakhter for keeping web-tab events private to their owner in multi-user mode (#501), with end-to-end isolation tests that fail without the fix, and for the browser-test exclusion check and pre-push hook (#500), including the hooks-dir resolution that never writes outside the repo's own `.git/hooks`. -- @opticon454 for keeping CLIs installed from Settings across Docker container updates (#490) and for the static Git identity for the Docker images (#492). -- @timkjr for letting a Shell pane's scroll-to-top reach tmux history (#494), with tests that fail on the commit before each fix. -- @JDProfresh for tracking down why wheel and touch scrolling did nothing in Claude's default inline view (#498), with the tmux measurements that proved it. -- @irisitymichaelgrundberg for the follow-up that makes an agent waiting on artifact comments raise its alert again (#491). - -**A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends. - -**Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. The same fix reaches cron jobs in "Paste (direct)" input mode, which reported `prompt_sent` for a prompt that never left the composer: the text is written raw, Enter follows as its own write 300 ms later, and the session presses it again while the prompt is still unsent. A cron run with no session to write to now fails instead of reporting the prompt as sent. - -**Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback. - -**Shell panes scroll back into tmux history (#494).** Scrolling to the top of a Shell pane now pulls the most recent 1 MiB of its tmux history, so output that arrived in a burst is reachable without pressing **Load full history**, which still loads the rest. - -**An agent waiting on artifact comments alerts again (#491).** A session whose agent published an artifact and is waiting for somebody to comment on it now raises the normal idle alert and lands in NEEDS YOU, instead of being treated as busy with background work. - -**Web-tab changes stay private in multi-user mode (#501).** The `webview:changed` event reached every connected user, exposing the ids of other users' web-tab creates, edits and deletes. It now carries the tab's owner and reaches that owner plus admins only. Single-user mode is unchanged apart from a new optional `owner` field on the event. - -**Phone header tabs look like tabs (#504).** On phones every header tab is now a chip with a fill and a border, the Alt+N digit (a keyboard hint a phone cannot use) is hidden, names get 80px instead of 50px, and the strip fades at whichever edge still has tabs scrolled out of view. - -**Docker: CLIs installed from Settings survive container updates (#490).** On the Compose deployment, CLIs installed from App Settings (DeepSeek, Pi and other npm-based CLIs) now go to `~/.local` on the persistent home mount, and `~/.local/bin` is on the image PATH, so recreating the container no longer discards them. Anything installed from Settings before this release has to be installed once more after the rebuild. - -**Docker: a static Git identity for the server and agent images (#492).** Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` (or `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` / `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` on a bare-host install) and the identity is written to `/etc/gitconfig` in the server image and the Docker-case agent image. A half-set pair is refused on every build path. Both Docker changes edit `server.Dockerfile`, so the in-app updater asks Compose deployments to rebuild with `docker/Start-Codeman.sh` instead of updating in place. - -**Contributor tooling (#500).** `npm run check:browser-excludes`, now a CI step, fails when a test that drives a real browser is still collected by `npm test`. `npm install` also installs a pre-push hook that runs the static CI checks before a push; it steps aside when the pushed ref is not HEAD or the tree has uncommitted changes the checks would read, and `CODEMAN_SKIP_PREPUSH=1 git push` skips it once. - -**Fixes applied while landing.** A Shell pane's scroll-to-top (#494) no longer re-pulls the same window on every gesture once the browser's 50,000-row scrollback is full, and a pull that hit the byte cap no longer claims the older history is gone. The scroll-routing diagnostics (#498) now log whether Claude has mouse tracking on. The artifact-comment check (#491) also refuses a footer cut off in the middle of the chip. The pre-push hook (#500) steps aside when `npm` is not on PATH, as in some GUI git clients, instead of blocking every push. The Docker Git identity error (#492) names the two variables to set. New tests pin the image PATH order, the identity on both agent-image build paths, and the `?full=1&tail=` terminal route. diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index ba10fa84..e5283924 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -10,7 +10,7 @@ "name": "codeman", "source": "./plugins/codeman", "description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", - "version": "1.33.1", + "version": "1.33.2", "author": { "name": "Ark0N", "url": "https://github.com/Ark0N" diff --git a/CHANGELOG.md b/CHANGELOG.md index 7357dc1f..53aeb9c8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,38 @@ # aicodeman +## 1.33.2 + +### Patch Changes + +- e439cf0: ### Thanks + - @aakhter for keeping web-tab events private to their owner in multi-user mode (#501), with end-to-end isolation tests that fail without the fix, and for the browser-test exclusion check and pre-push hook (#500), including the hooks-dir resolution that never writes outside the repo's own `.git/hooks`. + - @opticon454 for keeping CLIs installed from Settings across Docker container updates (#490) and for the static Git identity for the Docker images (#492). + - @timkjr for letting a Shell pane's scroll-to-top reach tmux history (#494), with tests that fail on the commit before each fix. + - @JDProfresh for tracking down why wheel and touch scrolling did nothing in Claude's default inline view (#498), with the tmux measurements that proved it. + - @irisitymichaelgrundberg for the follow-up that makes an agent waiting on artifact comments raise its alert again (#491). + + **A tab stays busy while Claude waits for its own workers.** When Claude hands work to an ultracode workflow or background agents, it ends its turn with `✻ Waiting for 1 dynamic workflow to finish` and resumes by itself when they report back. The idle probe used to call that session idle for the whole wait, and at phone width nothing on screen changes for minutes. A new optional registry field, `capabilities.workDetect.awaitingLine`, names that closing row, and only the newest column-0 row directly above the composer counts, so the session goes idle normally once the follow-up turn ends. + + **Prompts sent through the API are no longer left unsent.** A prompt posted to `POST /api/sessions/:id/input` without `useMux` was written into the pane in one piece, and Claude Code (measured on 2.1.283) takes a burst of about a hundred characters or more as a paste, so the trailing `\r` became a newline and the prompt sat on the composer while the route answered 200. Short prompts went through, which is why it looked random; Codex and OpenCode showed the same thing. A plain prompt (printable text plus exactly one trailing `\r`) now goes through tmux: the text is typed, Enter is pressed as its own key, and the server presses it again while the prompt is still on the composer. Raw frames (escape sequences, a bracketed paste, a line feed, a bare `\r`) and an explicit `"useMux": false` keep the direct write. The same fix reaches cron jobs in "Paste (direct)" input mode, which reported `prompt_sent` for a prompt that never left the composer: the text is written raw, Enter follows as its own write 300 ms later, and the session presses it again while the prompt is still unsent. A cron run with no session to write to now fails instead of reporting the prompt as sent. + + **Scrolling works again in Claude's default inline view (#498).** Wheel and touch gestures were forwarded to every Claude 2.1.187+ session as mouse reports, but only Claude's fullscreen renderer (`CLAUDE_CODE_NO_FLICKER=1`, or `"tui": "fullscreen"` in `~/.claude/settings.json`) listens for them, so in the default view scrolling did nothing. Codeman now forwards them only while Claude has mouse tracking switched on, and otherwise scrolls the terminal's own scrollback. + + **Shell panes scroll back into tmux history (#494).** Scrolling to the top of a Shell pane now pulls the most recent 1 MiB of its tmux history, so output that arrived in a burst is reachable without pressing **Load full history**, which still loads the rest. + + **An agent waiting on artifact comments alerts again (#491).** A session whose agent published an artifact and is waiting for somebody to comment on it now raises the normal idle alert and lands in NEEDS YOU, instead of being treated as busy with background work. + + **Web-tab changes stay private in multi-user mode (#501).** The `webview:changed` event reached every connected user, exposing the ids of other users' web-tab creates, edits and deletes. It now carries the tab's owner and reaches that owner plus admins only. Single-user mode is unchanged apart from a new optional `owner` field on the event. + + **Phone header tabs look like tabs (#504).** On phones every header tab is now a chip with a fill and a border, the Alt+N digit (a keyboard hint a phone cannot use) is hidden, names get 80px instead of 50px, and the strip fades at whichever edge still has tabs scrolled out of view. + + **Docker: CLIs installed from Settings survive container updates (#490).** On the Compose deployment, CLIs installed from App Settings (DeepSeek, Pi and other npm-based CLIs) now go to `~/.local` on the persistent home mount, and `~/.local/bin` is on the image PATH, so recreating the container no longer discards them. Anything installed from Settings before this release has to be installed once more after the rebuild. + + **Docker: a static Git identity for the server and agent images (#492).** Set `GIT_USER_NAME` and `GIT_USER_EMAIL` in `docker/.env` (or `CODEMAN_AGENT_IMAGE_GIT_USER_NAME` / `CODEMAN_AGENT_IMAGE_GIT_USER_EMAIL` on a bare-host install) and the identity is written to `/etc/gitconfig` in the server image and the Docker-case agent image. A half-set pair is refused on every build path. Both Docker changes edit `server.Dockerfile`, so the in-app updater asks Compose deployments to rebuild with `docker/Start-Codeman.sh` instead of updating in place. + + **Contributor tooling (#500).** `npm run check:browser-excludes`, now a CI step, fails when a test that drives a real browser is still collected by `npm test`. `npm install` also installs a pre-push hook that runs the static CI checks before a push; it steps aside when the pushed ref is not HEAD or the tree has uncommitted changes the checks would read, and `CODEMAN_SKIP_PREPUSH=1 git push` skips it once. + + **Fixes applied while landing.** A Shell pane's scroll-to-top (#494) no longer re-pulls the same window on every gesture once the browser's 50,000-row scrollback is full, and a pull that hit the byte cap no longer claims the older history is gone. The scroll-routing diagnostics (#498) now log whether Claude has mouse tracking on. The artifact-comment check (#491) also refuses a footer cut off in the middle of the chip. The pre-push hook (#500) steps aside when `npm` is not on PATH, as in some GUI git clients, instead of blocking every push. The Docker Git identity error (#492) names the two variables to set. New tests pin the image PATH order, the identity on both agent-image build paths, and the `?full=1&tail=` terminal route. + ## 1.33.1 ### Patch Changes diff --git a/CLAUDE.md b/CLAUDE.md index bb9cee13..539f9b32 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -78,7 +78,7 @@ When user says "COM": CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed. -**Version**: 1.33.1 (must match `package.json`) +**Version**: 1.33.2 (must match `package.json`) ## Project Overview diff --git a/package-lock.json b/package-lock.json index 14841794..d929379e 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "aicodeman", - "version": "1.33.1", + "version": "1.33.2", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "aicodeman", - "version": "1.33.1", + "version": "1.33.2", "hasInstallScript": true, "license": "MIT", "workspaces": [ diff --git a/package.json b/package.json index 0a57e311..3f7aa92e 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "aicodeman", - "version": "1.33.1", + "version": "1.33.2", "description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence", "type": "module", "main": "dist/index.js", diff --git a/plugins/codeman/.claude-plugin/plugin.json b/plugins/codeman/.claude-plugin/plugin.json index 7d21a17e..93cb5488 100644 --- a/plugins/codeman/.claude-plugin/plugin.json +++ b/plugins/codeman/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "codeman", "description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", - "version": "1.33.1", + "version": "1.33.2", "author": { "name": "Ark0N", "url": "https://github.com/Ark0N" From 612c69d57a5b14ccb461d44cdd8fbcfa7c44c20f Mon Sep 17 00:00:00 2001 From: JD <jd@jds.haus> Date: Mon, 28 Sep 2026 15:22:37 -0400 Subject: [PATCH 35/46] fix(files): decode markdown refs, scope links to the preview session, drop name= from the sanitizer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review follow-up on #503. marked percent-encodes link and image destinations, and the rebase pass encoded them a second time, so a space or a CJK character in a file name made file-raw look for a file literally named my%20image.png; refs are now decoded once (a malformed escape is kept as written) and stripped of ?query along with #fragment. Root-relative refs resolve from the workspace root as on GitHub instead of falling through as Codeman URLs. Rebased links carry the preview's own session id and the response-viewer delegate prefers it, so a document opened from another session's attachment card opens its links in that workspace rather than the active tab's. The sanitizer no longer allows name=: marked never emits it, and <img name="app"> made document.app that image, which every inline onclick="app.…()" handler resolves before the global, so one rendered README broke every viewer button until a reload. Adds the zh-CN strings for the three toolbar titles. --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- docs/wiki/Working-With-Files.md | 2 +- src/web/public/app.js | 4 ++- src/web/public/i18n.js | 3 ++ src/web/public/panels-ui.js | 43 +++++++++++++++++-------- src/web/public/sanitize-html.js | 6 ++-- test/file-preview-markdown.test.ts | 36 ++++++++++++++++++++- test/markdown-sanitizer.test.ts | 10 ++++++ test/response-viewer-file-links.test.ts | 2 +- 10 files changed, 88 insertions(+), 22 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 9dd64299..a373ef36 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -290,7 +290,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md` -**File Viewer text view: rendered markdown + Lines/Wrap toggles** (`_renderFilePreviewText()` in panels-ui.js): a `.md`/`.markdown` opens RENDERED by default with an `MD` pill back to source; the plain-text view has `Lines` (CSS-counter gutter) and `Wrap` toggles. ⚠️ ONE markdown pipeline: the viewer calls `_renderMarkdown()` (marked + the DOMPurify allowlist, the Response Viewer's) and binds the Response Viewer's click delegate (`_bindResponseViewerInteractions`) on the preview body for code-copy buttons and path links; never a second parser or handler. ⚠️ The document is built inside a `<template>` (a detached div with `innerHTML` set starts fetching every `<img src>` before the rewrite), then `_rebaseFilePreviewMarkdownRefs()` points relative images at the workspace-confined `file-raw` under the document's directory (never a widened route; a failed load degrades to alt text) and turns relative links into `a.rv-path` for the delegate, stripping the `target` marked gave them. ⚠️ The container carries `data-i18n-skip` or the translator rewrites the document's prose. ⚠️ Toggles are per-device localStorage keys (`codeman:filePreview*`), never `SettingsUpdateSchema`; Lines/Wrap are class flips on the ONE `<pre>`, with rules scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's code blocks. Markdown fetches `lines=10000` (the route ceiling), other text keeps 500. ⚠️ `md` stays OUT of `FILE_PREVIEW_EXTENSIONS`: a printed `.md` path keeps the tail viewer (live follow); the rendered view is the Files panel's. Tests: `test/file-preview-markdown.test.ts`. → [architecture-invariants#file-viewer-text-view-rendered-markdown-and-text-toggles](docs/architecture-invariants.md#file-viewer-text-view-rendered-markdown-and-text-toggles) +**File Viewer text view: rendered markdown + Lines/Wrap toggles** (`_renderFilePreviewText()` in panels-ui.js): a `.md`/`.markdown` opens RENDERED by default with an `MD` pill back to source; the plain-text view has `Lines` (CSS-counter gutter) and `Wrap` toggles. ⚠️ ONE markdown pipeline: the viewer calls `_renderMarkdown()` (marked + the DOMPurify allowlist, the Response Viewer's) and binds the Response Viewer's click delegate (`_bindResponseViewerInteractions`) on the preview body for code-copy buttons and path links; never a second parser or handler. ⚠️ The document is built inside a `<template>` (a detached div with `innerHTML` set starts fetching every `<img src>` before the rewrite), then `_rebaseFilePreviewMarkdownRefs()` points relative images at the workspace-confined `file-raw` under the document's directory and root-relative ones under the workspace root (never a widened route; a failed load degrades to alt text), after `decodeURIComponent`ing the ref and dropping `?query`/`#fragment` (marked percent-encodes destinations, and the route encodes again), and turns workspace links into `a.rv-path` carrying `data-session-id` for the delegate, stripping the `target` marked gave them. ⚠️ The container carries `data-i18n-skip` or the translator rewrites the document's prose. ⚠️ Toggles are per-device localStorage keys (`codeman:filePreview*`), never `SettingsUpdateSchema`; Lines/Wrap are class flips on the ONE `<pre>`, with rules scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's code blocks. Markdown fetches `lines=10000` (the route ceiling), other text keeps 500. ⚠️ `md` stays OUT of `FILE_PREVIEW_EXTENSIONS`: a printed `.md` path keeps the tail viewer (live follow); the rendered view is the Files panel's. Tests: `test/file-preview-markdown.test.ts`. → [architecture-invariants#file-viewer-text-view-rendered-markdown-and-text-toggles](docs/architecture-invariants.md#file-viewer-text-view-rendered-markdown-and-text-toggles) **Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure) compiles the query into a predicate the server-side walk prunes with; a query returns a FLAT match list and the walk recurses past non-matching directories. An empty, whitespace-only or overlong (`MAX_QUERY_LENGTH`, 256) query compiles to `null`, keeping the default tree response byte-identical. ⚠️ **Never compile a glob into a RegExp** (`*a*a*a…` backtracks and freezes the event loop for the whole server): `globMatch()` is a two-pointer wildcard walk. → [architecture-invariants#files-panel-search](docs/architecture-invariants.md#files-panel-search) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index d9f47f70..aa379258 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -389,7 +389,7 @@ Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write **Rendered markdown + Lines/Wrap** (`_renderFilePreviewText()` and its helpers in `panels-ui.js`, buttons in the `.file-preview-actions` row): a `.md`/`.markdown` opened in the File Viewer renders as a document by default, with an `MD` pill back to source; the plain-text view has `Lines` (a CSS-counter gutter) and `Wrap` toggles. Codeman already had `marked` + DOMPurify behind `_renderMarkdown()` for the Response Viewer, so the viewer reuses that and the codebase keeps ONE markdown pipeline. - ⚠️ **One pipeline, one delegate.** The viewer calls `_renderMarkdown()` (marked + the `sanitize-html.js` allowlist) and binds `_bindResponseViewerInteractions()` on `#filePreviewBody` (container-bound and idempotent, so once per page) for the code-block copy buttons and `a.rv-path` opening. Never a second parser, never a second click handler for the same markup. -- ⚠️ **Build inside a `<template>`, then rebase.** A detached div whose `innerHTML` is set starts fetching every `<img src>` at once, so the document's relative image paths would hit the server as `/docs/img.png` 404s before being rewritten. `_rebaseFilePreviewMarkdownRefs()` runs on the template content: relative images go to the workspace-confined `file-raw` under the document's directory (the server refuses escapes, so `..` is forwarded as-is), and one `error` handler per image degrades it to alt text, which covers a remote image the page CSP blocks, a 404 for a document outside the workspace, and an SVG that `file-raw` serves as a download. Never widen a route for this. Relative links become `a.rv-path` with `data-path` and lose the `target`/`rel` that `_renderMarkdown` gives every link, which would otherwise open `<origin>/docs/x.md` in a new tab; fragment and http(s) links are untouched. +- ⚠️ **Build inside a `<template>`, then rebase.** A detached div whose `innerHTML` is set starts fetching every `<img src>` at once, so the document's relative image paths would hit the server as `/docs/img.png` 404s before being rewritten. `_rebaseFilePreviewMarkdownRefs()` runs on the template content: relative images go to the workspace-confined `file-raw` under the document's directory, root-relative ones (`/docs/x.png`) under the workspace root as on GitHub (the server refuses escapes, so `..` is forwarded as-is), and one `error` handler per image degrades it to alt text, which covers a remote image the page CSP blocks, a 404 for a document outside the workspace, and an SVG that `file-raw` serves as a download. Never widen a route for this. ⚠️ marked percent-encodes destinations (`my image.png` arrives as `my%20image.png`, CJK names as `%E5…`), so the ref is `decodeURIComponent`ed (a malformed escape is kept as written) and stripped of `#fragment` and `?query` BEFORE the route encodes it again; without that `file-raw` looks for a file literally named `my%20image.png`. Workspace links become `a.rv-path` with `data-path` AND `data-session-id` (the preview's session, which the `app.js` delegate prefers over `activeSessionId`, since a preview opened from another session's attachment card must resolve links against that workspace) and lose the `target`/`rel` that `_renderMarkdown` gives every link, which would otherwise open `<origin>/docs/x.md` in a new tab; fragment, protocol-relative and http(s) links are untouched. - ⚠️ **`data-i18n-skip` on the container.** The translator's MutationObserver translates inserted headings and paragraphs, and the `.file-preview-content` entry in its skip list matches nothing (no element has that class), so the attribute is what keeps a Chinese UI from rewriting a README. - ⚠️ **Toggles are per-device, in their own localStorage keys** (`codeman:filePreviewMdRendered` / `LineNumbers` / `Wrap`), for the same reason as the Files panel's show-hidden toggle: the app-settings object is rebuilt from the settings modal on every save, and they are display state, not synced settings (`SettingsUpdateSchema` is `.strict()`). MD re-renders from the kept source (`filePreviewContent`, which is also what Copy copies) without a refetch; Lines/Wrap are class flips on the one `<pre>`, whose rules are scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's own code blocks. Lines are one inline `<span class="fp-line">` per line joined by real newlines, the counter in `::before` with `user-select: none`, so select and copy return the exact text. - ⚠️ **Caps.** Markdown fetches `lines=10000` (the route's `MAX_LINES_LIMIT`), because a rendered document cut at 500 lines reads as the whole document; other text keeps 500, which is what stops a huge log locking the tab in one `<pre>`. The attachment (out-of-workspace) branch keeps its 512 KB Range read and skips the 500-line clip for markdown. Edit mode is unchanged and still re-fetches `edit=1`; the three toggles hide while editing and for images, media and PDFs. diff --git a/docs/wiki/Working-With-Files.md b/docs/wiki/Working-With-Files.md index d8c8b357..f28458df 100644 --- a/docs/wiki/Working-With-Files.md +++ b/docs/wiki/Working-With-Files.md @@ -14,7 +14,7 @@ It renders what it can: | Kind | Behaviour | | ------------------------ | ------------------------------------------------------------------------- | | Text and code | Plain preview with Lines (line numbers) and Wrap toggles in the header. Long files are truncated in plain preview. | -| Markdown | Rendered by default: headings, tables, code blocks with copy buttons, images and links relative to the file. The MD pill in the header flips to source. | +| Markdown | Rendered by default: headings, tables, code blocks with copy buttons, images and links relative to the file (root-relative ones resolve from the workspace root, as on GitHub). The MD pill in the header flips to source. | | Images | Inline. | | Audio and video | Inline with a working scrub bar, because range requests are supported. | | PDF and Office documents | Converted for preview when a converter is available. | diff --git a/src/web/public/app.js b/src/web/public/app.js index 05577c34..4f5716fd 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -2358,7 +2358,9 @@ class CodemanApp { ev.preventDefault(); ev.stopPropagation(); const filePath = pathLink.dataset.path; - if (filePath) this.openFilePreview(filePath, this.activeSessionId); + // A rendered document's links name the session the preview was opened + // for (_rebaseFilePreviewMarkdownRefs), which need not be the active tab. + if (filePath) this.openFilePreview(filePath, pathLink.dataset.sessionId || this.activeSessionId); return; } diff --git a/src/web/public/i18n.js b/src/web/public/i18n.js index 82d6e6bb..bfa8bd7a 100644 --- a/src/web/public/i18n.js +++ b/src/web/public/i18n.js @@ -727,6 +727,9 @@ 'Source type filter': '来源类型筛选', 'Copy content': '复制内容', 'Edit file': '编辑文件', + 'Rendered markdown': '渲染 Markdown', + 'Line numbers': '行号', + 'Wrap lines': '自动换行', 'Unsaved changes': '未保存的更改', Saved: '已保存', 'Export as JSON': '导出为 JSON', diff --git a/src/web/public/panels-ui.js b/src/web/public/panels-ui.js index c8949d75..16fc2ee7 100644 --- a/src/web/public/panels-ui.js +++ b/src/web/public/panels-ui.js @@ -4357,28 +4357,42 @@ Object.assign(CodemanApp.prototype, { }, /** - * Point a rendered document's relative references at the file it came from. + * Point a rendered document's workspace references at the file it came from. * * Images are rebased onto the workspace-confined file-raw route under the - * document's directory (the server refuses escapes, so `..` is safe to - * forward). Whatever fails to load degrades to its alt text with one error - * handler: a remote image the page CSP blocks, a 404 for a document outside - * the workspace, an SVG that file-raw serves as a download. Relative links + * document's directory, root-relative ones (`/docs/x.png`) under the + * workspace root as on GitHub (the server refuses escapes, so `..` is safe + * to forward). Whatever fails to load degrades to its alt text with one + * error handler: a remote image the page CSP blocks, a 404 for a document + * outside the workspace, an SVG that file-raw serves as a download. Links * take the `a.rv-path` shape the Response Viewer delegate already opens in * this overlay, minus the target/rel `_renderMarkdown` gave them, which - * would otherwise open <origin>/docs/x.md in a new tab. + * would otherwise open <origin>/docs/x.md in a new tab, and carry the + * preview's own session so a document opened from another session's + * attachment card resolves against that workspace, not the active tab's. */ _rebaseFilePreviewMarkdownRefs(root, { sessionId, filePath }) { const dir = filePath.includes('/') ? filePath.slice(0, filePath.lastIndexOf('/') + 1) : ''; - // Relative = no scheme, not root-relative (which includes //host), not a fragment. - const isRelative = (ref) => !!ref && !/^[a-z][a-z0-9+.-]*:/i.test(ref) && !ref.startsWith('/') && !ref.startsWith('#'); - // GitHub-style `img.png#gh-dark-mode-only` and `doc.md#section`: the - // fragment is not part of the path. `.` and `..` segments are collapsed so - // the title reads `README.md`, not `docs/../README.md`; a `..` that climbs + // Workspace ref = no scheme, not protocol-relative (//host), not a fragment. + const isWorkspaceRef = (ref) => + !!ref && !/^[a-z][a-z0-9+.-]*:/i.test(ref) && !ref.startsWith('//') && !ref.startsWith('#'); + // GitHub-style `img.png#gh-dark-mode-only`, `doc.md#section` and + // `img.png?raw=true`: neither fragment nor query is part of the path. + // marked percent-encodes destinations (`my image.png` arrives as + // `my%20image.png`), so decode before the route encodes again, or file-raw + // looks for a file literally named `my%20image.png`; a malformed escape + // keeps the ref as written. `.` and `..` segments are collapsed so the + // title reads `README.md`, not `docs/../README.md`; a `..` that climbs // past the start is kept and left for the server to refuse. const resolveRef = (ref) => { + let rel = ref.split('#')[0].split('?')[0]; + try { + rel = decodeURIComponent(rel); + } catch { + /* malformed escape: keep the ref as written */ + } const parts = []; - for (const seg of (dir + ref.split('#')[0]).split('/')) { + for (const seg of (rel.startsWith('/') ? rel.slice(1) : dir + rel).split('/')) { if (seg === '.' || (seg === '' && parts.length)) continue; if (seg === '..' && parts.length && parts[parts.length - 1] !== '..' && parts[parts.length - 1] !== '') parts.pop(); else parts.push(seg); @@ -4387,7 +4401,7 @@ Object.assign(CodemanApp.prototype, { }; for (const img of root.querySelectorAll('img[src]')) { const src = img.getAttribute('src') || ''; - if (isRelative(src)) { + if (isWorkspaceRef(src)) { const path = resolveRef(src); img.setAttribute('src', CodemanBase.url(`/api/sessions/${sessionId}/file-raw?path=${encodeURIComponent(path)}`)); } @@ -4395,9 +4409,10 @@ Object.assign(CodemanApp.prototype, { } for (const a of root.querySelectorAll('a[href]')) { const href = a.getAttribute('href') || ''; - if (!isRelative(href)) continue; + if (!isWorkspaceRef(href)) continue; a.className = 'rv-path'; a.dataset.path = resolveRef(href); + a.dataset.sessionId = sessionId; a.setAttribute('href', '#'); a.removeAttribute('target'); a.removeAttribute('rel'); diff --git a/src/web/public/sanitize-html.js b/src/web/public/sanitize-html.js index 78f7aba7..be7aea28 100644 --- a/src/web/public/sanitize-html.js +++ b/src/web/public/sanitize-html.js @@ -84,7 +84,10 @@ /** * Attributes allowed on the tags above. `style` is intentionally absent (CSS-based vectors). * `class`/`id` survive because the response viewer adds wrapper classes downstream and code - * blocks may carry `language-*` classes from marked. + * blocks may carry `language-*` classes from marked. `name` is absent on purpose: marked never + * emits it, and `<img name="app">` would make `document.app` that image, which every inline + * `onclick="app.…()"` handler resolves before the global (DOM clobbering), so one rendered + * README could break every button until a reload. */ var ALLOWED_ATTR = [ 'href', @@ -93,7 +96,6 @@ 'title', 'class', 'id', - 'name', 'colspan', 'rowspan', 'align', diff --git a/test/file-preview-markdown.test.ts b/test/file-preview-markdown.test.ts index 553e2363..7c279204 100644 --- a/test/file-preview-markdown.test.ts +++ b/test/file-preview-markdown.test.ts @@ -50,7 +50,15 @@ const MARKDOWN_HTML = '<a href="guide/x.md#sec" target="_blank" rel="noopener noreferrer">x</a>' + '<a href="../CHANGELOG.md" target="_blank" rel="noopener noreferrer">up</a>' + '<a href="#top">t</a>' + - '<a href="https://e.com" target="_blank" rel="noopener noreferrer">e</a>'; + '<a href="https://e.com" target="_blank" rel="noopener noreferrer">e</a>' + + // marked percent-encodes destinations; a query rides along on GitHub-style refs. + '<img src="my%20image.png" alt="space">' + + '<img src="raw.png?raw=true" alt="raw">' + + '<img src="bad%zz.png" alt="bad">' + + '<img src="/assets/root.png" alt="root">' + + '<img src="//cdn.example.com/p.png" alt="protorel">' + + '<a href="%E5%9B%BE%E7%89%87/%E6%88%AA%E5%9B%BE.md" target="_blank" rel="noopener noreferrer">cjk</a>' + + '<a href="/docs/root.md" target="_blank" rel="noopener noreferrer">rootlink</a>'; const MD_CONTENT = '# Title\n\nx\n'; const TXT_CONTENT = 'one\n\n three\tfour\n'; @@ -200,6 +208,32 @@ describe('file viewer rendered markdown', () => { expect(external.getAttribute('target')).toBe('_blank'); }); + it('decodes percent-encoded refs, drops the query, and resolves root-relative refs against the workspace', async () => { + const { app, body } = loadApp(); + + await app.openFilePreview('docs/README.md', 's1'); + const src = (alt: string) => body.querySelector(`img[alt="${alt}"]`)!.getAttribute('src'); + const raw = (path: string) => `/api/sessions/s1/file-raw?path=${encodeURIComponent(path)}`; + + // Decoded once here, encoded once for the route: never `my%2520image.png`. + expect(src('space')).toBe(raw('docs/my image.png')); + expect(src('raw')).toBe(raw('docs/raw.png')); + // A malformed escape keeps the ref as written. + expect(src('bad')).toBe(raw('docs/bad%zz.png')); + // Root-relative is the workspace root, as on GitHub; protocol-relative is remote. + expect(src('root')).toBe(raw('assets/root.png')); + expect(src('protorel')).toBe('//cdn.example.com/p.png'); + + const anchors = Array.from(body.querySelectorAll('a')); + expect(anchors.find((a) => a.textContent === 'cjk')!.getAttribute('data-path')).toBe('docs/图片/截图.md'); + expect(anchors.find((a) => a.textContent === 'rootlink')!.getAttribute('data-path')).toBe('docs/root.md'); + // Every rebased link names the preview's session, so the delegate opens it + // in that workspace even when another tab is active. + const rebased = body.querySelectorAll('a.rv-path'); + expect(rebased.length).toBe(4); + for (const a of rebased) expect(a.getAttribute('data-session-id')).toBe('s1'); + }); + it('degrades an image that fails to load to its alt text', async () => { const { app, body } = loadApp(); diff --git a/test/markdown-sanitizer.test.ts b/test/markdown-sanitizer.test.ts index b9213c97..6aae1a96 100644 --- a/test/markdown-sanitizer.test.ts +++ b/test/markdown-sanitizer.test.ts @@ -165,6 +165,16 @@ describe('COD-56 markdown sanitizer (DOMPurify allowlist)', () => { expect(sanitize(html).toLowerCase()).not.toContain(tag); }); } + + // DOM clobbering: <img name="app"> makes document.app that image, and inline + // onclick="app.…()" handlers resolve `app` on the document before the global, + // so a rendered README could break every button until a reload. + it('drops name= (marked never emits it; it clobbers document.<name>)', () => { + const out = sanitize('<img name="app" src="https://example.com/x.png" alt="x"><a name="app" href="#a">a</a>'); + expect(out).not.toMatch(/\sname\s*=/i); + expect(out).toContain('src="https://example.com/x.png"'); + expect(out).toContain('href="#a"'); + }); }); describe('legitimate markdown-rendered HTML survives', () => { diff --git a/test/response-viewer-file-links.test.ts b/test/response-viewer-file-links.test.ts index 9667edf6..43e46833 100644 --- a/test/response-viewer-file-links.test.ts +++ b/test/response-viewer-file-links.test.ts @@ -145,6 +145,6 @@ describe('response viewer file-path linkifier', () => { // either leaves inert paths (no linkify) or dead links (no handler). expect(APP_SOURCE).toContain('this._linkifyFilePaths(renderedText)'); expect(APP_SOURCE).toMatch(/closest\('a\.rv-path'\)/); - expect(APP_SOURCE).toMatch(/openFilePreview\(filePath, this\.activeSessionId\)/); + expect(APP_SOURCE).toMatch(/openFilePreview\(filePath, pathLink\.dataset\.sessionId \|\| this\.activeSessionId\)/); }); }); From 17232b01f6b555de1346c59d0539777423d8575c Mon Sep 17 00:00:00 2001 From: timkjr <timkjr@k-lab.lan> Date: Fri, 25 Sep 2026 19:52:31 -0500 Subject: [PATCH 36/46] fix(split-pane): let a Shell Pane B's scroll-up reach tmux history tmux repaints a burst of output instead of scrolling it, so a shell pane's xterm keeps about one screen of scrollback while tmux holds every line. The primary pane goes back for it when the wheel reaches the top; Pane B is a separate xterm that loaded history once at connect and never again, so after a `cat` its earlier output was unreachable. Pane B now does the same for a shell session: wheel-up at the top of the normal screen pulls ?full=1&tail=TERMINAL_TAIL_SIZE and holds the reader's place across the replay. The wheel listener is capture-phase because xterm stopPropagation()s the events it consumes. It follows the primary pane's rules from #494 and its 1.33.2 merge-time fixes: a window holding no more rows than the pane (which covers a downgrade), or a pane already at its `scrollback + rows` cap, is skipped without a rewrite. That skip backs off to 60 s when the window was truncated or the pane is full, since each ask costs the server a whole-history capture-pane; an untruncated window keeps the 4 s cooldown. There is no truncation banner in Pane B, so the 'tail' relabel does not apply. Live frames, a {t:'c'} clear included, are held with their arrival time while the replay runs and applied in order only if they arrived after the capture. The fetch has a 10 s deadline since it holds live output while it runs. The tail of _loadBuffer() becomes _endBufferLoad() so the pull shares its single-flight bookkeeping. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- src/web/public/terminal-split.js | 159 +++++++++- test/split-pane-terminal-unit.test.ts | 436 +++++++++++++++++++++++++- 4 files changed, 584 insertions(+), 15 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 539f9b32..1bfee139 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -270,7 +270,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. A Shell split-pane Pane B has its own copy of the bounded pull against its own xterm (`SplitTerminalPane._pullHistory`, terminal-split.js); keep the two in step. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 848ae8f1..3a7832b3 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -767,7 +767,7 @@ Further detail: with many sessions the horizontal strip stops being scannable, w ### Split-pane sessions -**Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. Design: `docs/split-pane-sessions-plan.md`. +**Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. ⚠️ A SHELL Pane B pulls scrollback itself when the wheel goes up at the top of its buffer (`_maybeLoadMoreHistory`/`_pullHistory`): tmux repaints a burst of output instead of scrolling it, so Pane B's own xterm holds about one screen of scrollback while tmux holds every line, and it loaded history exactly once at connect and never again. It is the same bounded pull as Pane A's (`?full=1&tail=TERMINAL_TAIL_SIZE`, no rewrite when the window holds no more rows than the pane already has or the pane is at its `scrollback + rows` cap, and a 60 s back-off instead of 4 s when that skipped window was truncated or the pane is full, since each ask costs the server a whole-history `capture-pane`), against Pane B's OWN terminal rather than `app.terminal`, so it cannot share `_maybeRefetchFullHistory`. The wheel listener is capture-phase because xterm `stopPropagation()`s the events it consumes; the alternate screen (nano, vim, less) is skipped; live frames arriving mid-replay, a `{t:'c'}` clear frame included, are held with their arrival time (`_liveQueue`) and replayed in order only if they arrived after the capture (the response's arrival stands in for the capture instant, as in `_finishBufferLoad`, so a frame inside that one round trip can be lost or doubled); the fetch has a 10 s deadline because it holds the pane's live output while it runs. There is no "Load full history" banner in Pane B, so a shell history past that 1 MiB window stays out of reach there. Non-shell Pane B is unchanged: it already loads `full=1`, and a repaint-mode CLI keeps no tmux history to recover. Design: `docs/split-pane-sessions-plan.md`. ### Gesture control: the setting diff --git a/src/web/public/terminal-split.js b/src/web/public/terminal-split.js index 3c18dd81..cc9fedb2 100644 --- a/src/web/public/terminal-split.js +++ b/src/web/public/terminal-split.js @@ -14,6 +14,9 @@ */ (function (global) { + // How long a scroll-to-top history pull may hold Pane B's live output. + const HISTORY_PULL_TIMEOUT_MS = 10000; + /** * Minimal chunked write for Pane B's own xterm instance — write() in * TERMINAL_CHUNK_SIZE slices, yielding a frame between each, instead of one @@ -72,6 +75,13 @@ // Single-flight state for _loadBuffer()/_refreshBuffer() below. this._bufferLoading = false; this._bufferRefreshPending = false; + // Scroll-to-top history pull (shell panes only), see _maybeLoadMoreHistory(). + // `_liveQueue` is non-null exactly while a pull is replaying: live frames + // are held there with their arrival time instead of written under it. + this._historyPullAt = 0; + this._historyPullUseless = false; + this._liveQueue = null; + this._onWheel = null; } async connect() { @@ -95,6 +105,8 @@ this.terminal.open(this.mountEl); this.fitAddon.fit(); + this._installWheelListener(); + this.terminal.onData((data) => { if (this.ws && this.ws.readyState === WebSocket.OPEN) { this.ws.send(JSON.stringify({ t: 'i', d: data })); @@ -254,9 +266,9 @@ try { const msg = JSON.parse(event.data); if (msg.t === 'o') { - this.terminal.write(msg.d); + this._onLiveOutput(msg.d); } else if (msg.t === 'c') { - this.terminal.clear(); + this._onLiveClear(); } else if (msg.t === 'r') { // Server-triggered refresh (SSE backpressure cleared, terminal // data was dropped). The primary pane routes this to @@ -324,14 +336,151 @@ } catch { /* Best-effort — live output still arrives once the socket connects. */ } finally { - this._bufferLoading = false; + this._endBufferLoad(); } + } + + // Ends a single-flight load (initial, refresh or history pull): clears the + // flag, then runs the ONE trailing refresh that arrived while it was busy. + _endBufferLoad() { + this._bufferLoading = false; if (this._bufferRefreshPending && !this._destroyed) { this._bufferRefreshPending = false; this._refreshBuffer(); } } + // Live terminal output. Written straight through, except while a history + // pull is replaying: a capture is current only up to the instant tmux took + // it, so a frame arriving mid-replay is held with its arrival time and + // replayed behind the snapshot by _pullHistory() (the primary pane's + // _finishBufferLoad `since` rule), never written underneath it. + _onLiveOutput(data) { + if (this._liveQueue) this._liveQueue.push({ at: performance.now(), data }); + else this.terminal?.write(data); + } + + // The server's `{t:'c'}` clear frame takes the same route as output, for the + // same reason: clearing straight away, mid-replay, would wipe the half-written + // snapshot and leave _pullHistory() measuring a buffer that is no longer the + // one it is restoring. Queued, it lands in order with the frames around it. + _onLiveClear() { + if (this._liveQueue) this._liveQueue.push({ at: performance.now(), clear: true }); + else this.terminal?.clear(); + } + + // Capture phase, because xterm's own wheel handler stopPropagation()s every + // event it consumes, so a bubbling listener here would never see the wheel + // while the pane still has scrollback to scroll. Passive: this only observes, + // xterm keeps doing the scrolling. + _installWheelListener() { + this._onWheel = (ev) => { + if (ev.deltaY < 0) this._maybeLoadMoreHistory(); + }; + this.mountEl.addEventListener('wheel', this._onWheel, { capture: true, passive: true }); + } + + // Wheel-up at the top of a SHELL pane's scrollback. tmux repaints a burst of + // output (`cat` of a file longer than the screen) instead of scrolling it, + // so this pane's xterm ends up with about one screen of scrollback while + // tmux holds every line — and nothing here ever went back to ask, so the + // history was unreachable. The primary pane has the same pull + // (app.js _maybeRefetchFullHistory); Pane B is a separate xterm and needs its + // own. Shell only: a repaint-mode agent CLI keeps no tmux history to recover, + // and its load already takes `full=1`. Skipped on the alternate screen + // (nano, vim, less), where the wheel belongs to the app, not the scrollback. + _maybeLoadMoreHistory() { + if (this.sessionMode !== 'shell' || this._destroyed || !this.terminal) return; + if (this._bufferLoading) return; + const active = this.terminal.buffer.active; + if (active.type !== 'normal' || active.viewportY !== 0) return; + // Momentum scrolling fires this dozens of times per flick, so cooldown + // rather than latch; a pull that could only have downgraded the pane + // waits far longer. + const cooldown = this._historyPullUseless ? 60000 : 4000; + const now = Date.now(); + if (now - this._historyPullAt < cooldown) return; + this._historyPullAt = now; + void this._pullHistory(); + } + + // Pulls a BOUNDED window of tmux's full history (the same TERMINAL_TAIL_SIZE + // a tab switch loads, so a multi-megabyte capture never lands on xterm's + // main thread) and replays it under the reader's current place. Holds the + // single-flight flag across the fetch AND the replay, like _loadBuffer(). + async _pullHistory() { + this._bufferLoading = true; + this._liveQueue = []; + let replayed = false; + let capturedAt = 0; + try { + // A deadline, because live output is held for as long as this runs: a + // request that hangs would otherwise freeze the whole pane. Aborting + // lands in the catch below, which releases the flag and the queue. It + // covers the body read too, not just the headers. + const res = await fetch(`/api/sessions/${this.sessionId}/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`, { + signal: global.AbortSignal?.timeout?.(HISTORY_PULL_TIMEOUT_MS), + }); + // The cutoff below is the response's arrival, the same `since` rule the + // primary pane uses (_finishBufferLoad). It is a client clock standing in + // for the instant tmux took the capture, which lies somewhere in the + // round trip, so a frame in that window can be lost or doubled. Bounded + // by one round trip and not closable without a server-side capture time. + capturedAt = performance.now(); + const payload = (await res.json())?.data; + const buffer = payload?.terminalBuffer; + const term = this.terminal; + if (!buffer || !term || this._destroyed) return; + const rowsBefore = term.buffer.active.length; + const rowsIncoming = global.app?._estimateReplayRows?.(buffer, term.cols) ?? buffer.split('\n').length; + // xterm keeps at most `scrollback + rows` rows while tmux keeps far more + // lines, so a window of short lines can carry more rows than this pane + // can ever hold, and `rowsIncoming <= rowsBefore` would never come true. + const scrollbackCap = term.options?.scrollback || 0; + const paneFull = scrollbackCap > 0 && rowsBefore >= scrollbackCap + term.rows; + // Nothing to gain (this also covers a downgrade, which would delete + // history mid-scroll), and a reset+rewrite would jump the viewport. An + // untruncated window IS all of tmux's history and the next burst can add + // more, so keep the 4 s cooldown. A truncated window can never reach past + // what the pane shows, and every ask costs the server a capture-pane of + // the whole history (`tail` is cut after it): back off to 60 s, as the + // primary pane does (app.js _maybeRefetchFullHistory). A full pane backs + // off too, since no window can ever fit in it. + if (rowsIncoming <= rowsBefore || paneFull) { + if (payload.truncated || paneFull) this._historyPullUseless = true; + return; + } + this._historyPullUseless = false; + term.write('\x1bc'); + replayed = true; + await writeChunked(term, buffer, () => this._destroyed); + if (this._destroyed || !this.terminal) return; + // xterm parses asynchronously: an empty write's callback fires only + // after everything before it, so the row count below is the settled one. + await new Promise((resolve) => this.terminal.write('', resolve)); + if (this._destroyed || !this.terminal) return; + // The replay grew the buffer UPWARD, so what was row 0 is now `delta` + // rows down; land there and the recovered history sits above it. + const delta = this.terminal.buffer.active.length - rowsBefore; + if (delta > 0) this.terminal.scrollToLine(delta); + else this.terminal.scrollToTop(); + } catch { + /* Best-effort — live output keeps arriving whatever happens here. */ + } finally { + const queued = this._liveQueue ?? []; + this._liveQueue = null; + // After a replay, only frames that arrived after the capture are news; + // earlier ones are already in it. With no replay, every held frame is. + const cutoff = replayed ? capturedAt : 0; + for (const entry of queued) { + if (entry.at < cutoff) continue; + if (entry.clear) this.terminal?.clear(); + else this.terminal?.write(entry.data); + } + this._endBufferLoad(); + } + } + // The `{t:'r'}` server-refresh path: clear, then replay. Two refresh // frames in a row used to start two concurrent replays, each clearing // the terminal under the other's chunked write. A refresh that arrives @@ -381,6 +530,10 @@ destroy() { this._destroyed = true; + if (this._onWheel) { + this.mountEl?.removeEventListener('wheel', this._onWheel, { capture: true }); + this._onWheel = null; + } if (this.ws) { this.ws.onopen = null; this.ws.onmessage = null; diff --git a/test/split-pane-terminal-unit.test.ts b/test/split-pane-terminal-unit.test.ts index 8ce88461..17021c3f 100644 --- a/test/split-pane-terminal-unit.test.ts +++ b/test/split-pane-terminal-unit.test.ts @@ -11,17 +11,31 @@ // refresh arriving mid-replay is now coalesced into ONE trailing re-run rather // than dropped, because the in-flight fetch may predate the drop the new frame // reports and no further frame comes to correct stale content. +// +// The last block covers the scroll-to-top history pull: a burst of output leaves +// a shell pane's xterm with about one screen of scrollback while tmux holds every +// line, and Pane B (a separate xterm from the primary pane) never went back to +// ask. See _maybeLoadMoreHistory / _pullHistory in terminal-split.js. import { readFileSync } from 'node:fs'; import { resolve } from 'node:path'; import vm from 'node:vm'; import { beforeEach, describe, expect, it, vi } from 'vitest'; const TERMINAL_CHUNK_SIZE = 32 * 1024; +const TERMINAL_TAIL_SIZE = 1024 * 1024; +/** The pane's `performance.now()`, so frame arrival vs. capture time is set by hand, not raced. */ +let clock = 0; type FakeTerminal = { write: ReturnType<typeof vi.fn>; clear: ReturnType<typeof vi.fn>; dispose: ReturnType<typeof vi.fn>; + scrollToLine: ReturnType<typeof vi.fn>; + scrollToTop: ReturnType<typeof vi.fn>; + cols: number; + rows: number; + options: { scrollback: number }; + buffer: { active: { type: string; viewportY: number; length: number } }; }; type FakeSocket = { onopen: unknown; @@ -36,43 +50,72 @@ type PaneUnderTest = { _destroyed: boolean; _bufferLoading: boolean; _bufferRefreshPending: boolean; + _historyPullAt: number; + _historyPullUseless: boolean; + _liveQueue: unknown[] | null; + _onWheel: unknown; destroy(): void; _loadBuffer(): Promise<void>; _refreshBuffer(): void; + _maybeLoadMoreHistory(): void; + _pullHistory(): Promise<void>; + _onLiveOutput(data: string): void; + _onLiveClear(): void; + _installWheelListener(): void; }; const fetchMock = vi.fn(); /** requestAnimationFrame stand-in: chunked writes queue here and are drained by hand. */ const rafQueue: Array<() => void> = []; +const SOURCE = readFileSync(resolve(import.meta.dirname, '../src/web/public/terminal-split.js'), 'utf8'); function loadSplitTerminalPane() { - const dir = resolve(import.meta.dirname, '../src/web/public'); - const src = readFileSync(resolve(dir, 'terminal-split.js'), 'utf8'); const context = vm.createContext({ console: { ...console, log: vi.fn(), warn: vi.fn(), error: vi.fn() }, - window: {}, + // The primary pane's row estimator, reduced to a line count: the pull only + // compares it with the pane's own row count. + window: { + app: { _estimateReplayRows: (text: string) => text.split('\n').length }, + AbortSignal: { timeout: (ms: number) => ({ timeoutMs: ms }) }, + }, + performance: { now: () => clock }, fetch: (...args: unknown[]) => fetchMock(...args), requestAnimationFrame: (fn: () => void) => rafQueue.push(fn), // The constants.js globals the module reads at call time. TERMINAL_CHUNK_SIZE, - TERMINAL_TAIL_SIZE: 1024 * 1024, + TERMINAL_TAIL_SIZE, }); // The module's tail patches CodemanApp.prototype; nothing on it runs here. - vm.runInContext(`class CodemanApp { _onSessionDeleted() {} selectSession() {} }\n${src}`, context); + vm.runInContext(`class CodemanApp { _onSessionDeleted() {} selectSession() {} }\n${SOURCE}`, context); return (context.window as { SplitTerminalPane: new (id: string, mount: unknown, opts?: object) => PaneUnderTest }) .SplitTerminalPane; } const SplitTerminalPane = loadSplitTerminalPane(); -function makePane(mode = 'claude'): PaneUnderTest & { terminal: FakeTerminal } { - const pane = new SplitTerminalPane('s1', {}, { mode }); - pane.terminal = { write: vi.fn(), clear: vi.fn(), dispose: vi.fn() }; +function makePane(mode = 'claude', mount: unknown = {}): PaneUnderTest & { terminal: FakeTerminal } { + const pane = new SplitTerminalPane('s1', mount, { mode }); + pane.terminal = { + // xterm invokes a write's callback once everything before it is parsed. + write: vi.fn((_data: string, done?: () => void) => done?.()), + clear: vi.fn(), + dispose: vi.fn(), + scrollToLine: vi.fn(), + scrollToTop: vi.fn(), + cols: 80, + rows: 30, + // xterm keeps at most `scrollback + rows` rows; small here so a test can fill it. + options: { scrollback: 1000 }, + // A pane sitting at the top of a 40-row buffer on the normal screen. + buffer: { active: { type: 'normal', viewportY: 0, length: 40 } }, + }; return pane as PaneUnderTest & { terminal: FakeTerminal }; } -function jsonResponse(terminalBuffer: string) { - return { json: async () => ({ data: { terminalBuffer } }) }; +const rowsOf = (n: number) => Array.from({ length: n }, (_, i) => `line ${i}`).join('\n'); + +function jsonResponse(terminalBuffer: string, extra: Record<string, unknown> = {}) { + return { json: async () => ({ data: { terminalBuffer, ...extra } }) }; } function deferred<T>() { @@ -89,6 +132,7 @@ const settle = () => new Promise((r) => setTimeout(r, 0)); beforeEach(() => { fetchMock.mockReset(); rafQueue.length = 0; + clock = 0; }); describe('SplitTerminalPane.destroy()', () => { @@ -232,3 +276,375 @@ describe('SplitTerminalPane server-refresh single-flight', () => { expect(pane.terminal.write).toHaveBeenCalledWith('back'); }); }); + +describe('SplitTerminalPane scroll-to-top history pull', () => { + it('a shell pane at the top pulls a bounded window of full history and replays it', async () => { + const pane = makePane('shell'); + const term = pane.terminal; + // The replay grows the buffer once xterm has parsed it (the empty write's callback). + term.write.mockImplementation((data: string, done?: () => void) => { + if (data === '' && done) term.buffer.active.length = 140; + done?.(); + }); + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(100))); + + pane._maybeLoadMoreHistory(); + await settle(); + + // With a deadline: live output is held for as long as the pull runs, so a + // request that never answers would freeze the pane. + expect(fetchMock).toHaveBeenCalledWith(`/api/sessions/s1/terminal?full=1&tail=${TERMINAL_TAIL_SIZE}`, { + signal: { timeoutMs: 10_000 }, + }); + expect(term.write).toHaveBeenCalledWith('\x1bc'); + expect(term.write).toHaveBeenCalledWith(rowsOf(100)); + // What was row 0 is now 100 rows down (140 - 40): the reader keeps their + // place with the recovered history above it, instead of being dropped at the bottom. + expect(term.scrollToLine).toHaveBeenCalledWith(100); + expect(pane._bufferLoading).toBe(false); + expect(pane._liveQueue).toBeNull(); + }); + + it('does nothing away from the top, for other modes, or on the alternate screen', async () => { + const midScroll = makePane('shell'); + midScroll.terminal.buffer.active.viewportY = 12; + midScroll._maybeLoadMoreHistory(); + + // A repaint-mode agent CLI keeps no tmux history to recover. + makePane('claude')._maybeLoadMoreHistory(); + + // nano/vim/less own the wheel; their screen is not scrollback. + const fullScreenApp = makePane('shell'); + fullScreenApp.terminal.buffer.active.type = 'alternate'; + fullScreenApp._maybeLoadMoreHistory(); + + await settle(); + expect(fetchMock).not.toHaveBeenCalled(); + }); + + it('a flick fires once: overlapping triggers are dropped, then the cooldown holds', async () => { + const pane = makePane('shell'); + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + pane._maybeLoadMoreHistory(); + const startedAt = pane._historyPullAt; + // The cooldown is cleared between triggers on purpose, so that only the + // in-flight guard can be what drops the overlapping ones. + pane._historyPullAt = 0; + pane._maybeLoadMoreHistory(); + pane._historyPullAt = 0; + pane._maybeLoadMoreHistory(); + expect(fetchMock).toHaveBeenCalledTimes(1); + pane._historyPullAt = startedAt; + + response.resolve(jsonResponse(rowsOf(100))); + await settle(); + expect(pane._bufferLoading).toBe(false); + + // Nothing in flight any more, so now it is the 4s cooldown alone. + pane._maybeLoadMoreHistory(); + await settle(); + expect(fetchMock).toHaveBeenCalledTimes(1); + + // Once the cooldown lapses a later scroll-to-top may pull again. + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(100))); + pane._historyPullAt = Date.now() - 5000; + pane._maybeLoadMoreHistory(); + await settle(); + expect(fetchMock).toHaveBeenCalledTimes(2); + }); + + it('a window the pane already holds in full is not rewritten, and is not latched as useless', async () => { + const pane = makePane('shell'); + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(30))); + + pane._maybeLoadMoreHistory(); + await settle(); + + // A reset+rewrite here would jump the viewport for no new rows. + expect(pane.terminal.write).not.toHaveBeenCalledWith('\x1bc'); + expect(pane.terminal.scrollToLine).not.toHaveBeenCalled(); + expect(pane.terminal.scrollToTop).not.toHaveBeenCalled(); + // The next burst can put more history in tmux than the pane has. + expect(pane._historyPullUseless).toBe(false); + expect(pane._bufferLoading).toBe(false); + }); + + it('refuses a downgrade, keeping the 4s cooldown when the window is all of tmux history', async () => { + const pane = makePane('shell'); + pane.terminal.buffer.active.length = 500; + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(5))); + + pane._maybeLoadMoreHistory(); + await settle(); + + expect(pane.terminal.write).not.toHaveBeenCalledWith('\x1bc'); + // Untruncated: tmux has nothing older, but the next burst can add history. + expect(pane._historyPullUseless).toBe(false); + }); + + it('a truncated window that fits in the pane backs off for a minute', async () => { + // Every ask costs the server a capture-pane of the WHOLE history (`tail` is + // cut after the capture), and a window cut at the tail size can never reach + // anything older than what the pane already shows. + const pane = makePane('shell'); + pane.terminal.buffer.active.length = 500; + // Within a screen of what the pane holds, so the old downgrade guard never + // latched it: only the truncated-skip rule can back this off. + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(480), { truncated: true, truncationReason: 'tail' })); + + pane._maybeLoadMoreHistory(); + await settle(); + + expect(pane.terminal.write).not.toHaveBeenCalledWith('\x1bc'); + expect(pane._historyPullUseless).toBe(true); + + // Inside the 60s back-off, well past the normal 4s cooldown. + pane._historyPullAt = Date.now() - 10_000; + pane._maybeLoadMoreHistory(); + await settle(); + expect(fetchMock).toHaveBeenCalledTimes(1); + }); + + it('a pane already at its scrollback cap skips the window and backs off for a minute', async () => { + // A 1 MiB window of short lines can carry more rows than xterm will ever hold + // (`scrollback + rows`), so `incoming <= rows held` never comes true and every + // scroll-to-top would reset and re-parse it. + const pane = makePane('shell'); + pane.terminal.buffer.active.length = 1030; + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(5000))); + + pane._maybeLoadMoreHistory(); + await settle(); + + expect(pane.terminal.write).not.toHaveBeenCalledWith('\x1bc'); + expect(pane.terminal.write).not.toHaveBeenCalledWith(rowsOf(5000)); + expect(pane._historyPullUseless).toBe(true); + }); + + it('a successful replay clears the one-minute back-off', async () => { + const pane = makePane('shell'); + pane._historyPullUseless = true; + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(100), { truncated: true, truncationReason: 'tail' })); + + void pane._pullHistory(); + await settle(); + + expect(pane.terminal.write).toHaveBeenCalledWith(rowsOf(100)); + expect(pane._historyPullUseless).toBe(false); + }); + + it('holds live output during the replay and replays only what arrived after the capture', async () => { + const pane = makePane('shell'); + const term = pane.terminal; + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + pane._maybeLoadMoreHistory(); + expect(pane._liveQueue).toEqual([]); + + // Arrives before the response does: it is IN the capture already. + clock = 1; + pane._onLiveOutput('early'); + expect(term.write).not.toHaveBeenCalledWith('early'); + await settle(); + + // 200 rows (more than the pane holds, so it replays) of 400 columns each: + // three chunks, which leaves the replay mid-write once the fetch lands. + const bigReplay = Array.from({ length: 200 }, () => 'y'.repeat(400)).join('\n'); + expect(bigReplay.length).toBeGreaterThan(TERMINAL_CHUNK_SIZE * 2); + clock = 2; // the response arrives: this is the cutoff + response.resolve(jsonResponse(bigReplay)); + await settle(); + expect(rafQueue).toHaveLength(1); + + // Arrives while the snapshot is still being written: must not land under it. + clock = 3; + pane._onLiveOutput('late'); + expect(term.write).not.toHaveBeenCalledWith('late'); + + rafQueue.shift()!(); + rafQueue.shift()!(); + await settle(); + + const written = term.write.mock.calls.map((call) => call[0]); + expect(written).not.toContain('early'); + expect(written.at(-1)).toBe('late'); + expect(pane._liveQueue).toBeNull(); + expect(pane._bufferLoading).toBe(false); + }); + + it('writes every held frame when the pull ends without replaying', async () => { + const pane = makePane('shell'); + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + pane._maybeLoadMoreHistory(); + pane._onLiveOutput('held'); + await settle(); + response.resolve(jsonResponse(rowsOf(30))); // nothing to gain: no replay + await settle(); + + // Nothing replaced the terminal, so the frame is news even though it + // arrived before the response did. + expect(pane.terminal.write).toHaveBeenCalledWith('held'); + }); + + it('a failed fetch releases the flag and the queue, so live output flows again', async () => { + const pane = makePane('shell'); + fetchMock.mockRejectedValueOnce(new Error('offline')); + + pane._maybeLoadMoreHistory(); + pane._onLiveOutput('held'); + await settle(); + + expect(pane._bufferLoading).toBe(false); + expect(pane._liveQueue).toBeNull(); + expect(pane.terminal.write).toHaveBeenCalledWith('held'); + pane._onLiveOutput('after'); + expect(pane.terminal.write).toHaveBeenLastCalledWith('after'); + }); + + it('a refresh frame during the pull runs once behind it', async () => { + const pane = makePane('shell'); + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise).mockResolvedValueOnce(jsonResponse('refreshed')); + + pane._maybeLoadMoreHistory(); + pane._refreshBuffer(); + expect(pane.terminal.clear).not.toHaveBeenCalled(); + expect(pane._bufferRefreshPending).toBe(true); + + response.resolve(jsonResponse(rowsOf(30))); + await settle(); + + expect(pane.terminal.clear).toHaveBeenCalledTimes(1); + expect(fetchMock).toHaveBeenCalledTimes(2); + expect(pane.terminal.write).toHaveBeenCalledWith('refreshed'); + }); + + it('a clear frame during the pull is queued in order, never applied under the replay', async () => { + const pane = makePane('shell'); + const term = pane.terminal; + const order: string[] = []; + term.write.mockImplementation((data: string, done?: () => void) => { + order.push(`write:${data}`); + done?.(); + }); + term.clear.mockImplementation(() => order.push('clear')); + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + pane._maybeLoadMoreHistory(); + pane._onLiveOutput('before'); + pane._onLiveClear(); + pane._onLiveOutput('after'); + // Held: clearing now would wipe a half-written snapshot. + expect(order).toEqual([]); + + response.resolve(jsonResponse(rowsOf(30))); // nothing to gain: no replay + await settle(); + + expect(order).toEqual(['write:before', 'clear', 'write:after']); + expect(pane._liveQueue).toBeNull(); + + // With nothing in flight a clear frame applies straight away. + pane._onLiveClear(); + expect(order.at(-1)).toBe('clear'); + }); + + it('a clear that arrived before the capture is not replayed after it', async () => { + const pane = makePane('shell'); + const term = pane.terminal; + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + pane._maybeLoadMoreHistory(); + clock = 1; + pane._onLiveClear(); // already reflected in the capture + clock = 2; + response.resolve(jsonResponse(rowsOf(100))); + await settle(); + + expect(term.write).toHaveBeenCalledWith('\x1bc'); + expect(term.clear).not.toHaveBeenCalled(); + }); + + it('destroy() mid-pull leaves nothing running and nothing written to the dead terminal', async () => { + const pane = makePane('shell'); + const term = pane.terminal; + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + pane._maybeLoadMoreHistory(); + pane._onLiveOutput('held'); + pane.destroy(); + response.resolve(jsonResponse(rowsOf(100))); + await settle(); + + expect(pane._bufferLoading).toBe(false); + expect(pane._liveQueue).toBeNull(); + expect(pane.terminal).toBeNull(); + expect(term.write).not.toHaveBeenCalledWith('\x1bc'); + expect(term.write).not.toHaveBeenCalledWith('held'); + }); + + it('a pull whose request is aborted (the deadline) frees the pane', async () => { + const pane = makePane('shell'); + fetchMock.mockRejectedValueOnce(new Error('The operation timed out')); + + pane._maybeLoadMoreHistory(); + pane._onLiveOutput('held'); + await settle(); + + expect(pane._bufferLoading).toBe(false); + expect(pane._liveQueue).toBeNull(); + expect(pane.terminal.write).toHaveBeenCalledWith('held'); + }); + + it('the wheel listener is capture-phase, and only a wheel UP can trigger a pull', async () => { + const mount = { addEventListener: vi.fn(), removeEventListener: vi.fn() }; + const pane = makePane('shell', mount); + fetchMock.mockResolvedValue(jsonResponse(rowsOf(100))); + + pane._installWheelListener(); + + // Capture phase: xterm's own wheel handler stopPropagation()s the events it + // consumes, so a bubbling listener would never fire while the pane still has + // scrollback to scroll, and the pull would work only from the exact top row. + const [type, listener, options] = mount.addEventListener.mock.calls[0]; + expect(type).toBe('wheel'); + expect(options).toEqual({ capture: true, passive: true }); + + listener({ deltaY: 120 }); // wheel down + listener({ deltaY: 0 }); + await settle(); + expect(fetchMock).not.toHaveBeenCalled(); + + listener({ deltaY: -120 }); // wheel up, at the top + await settle(); + expect(fetchMock).toHaveBeenCalledTimes(1); + }); + + it('destroy() detaches exactly the wheel listener it registered', () => { + const mount = { addEventListener: vi.fn(), removeEventListener: vi.fn() }; + const pane = makePane('shell', mount); + pane._installWheelListener(); + const registered = mount.addEventListener.mock.calls[0][1]; + + pane.destroy(); + + expect(mount.removeEventListener).toHaveBeenCalledWith('wheel', registered, { capture: true }); + expect(pane._onWheel).toBeNull(); + }); + + it('connect() installs the wheel listener (static guard)', () => { + // connect() needs a whole xterm to run, so its wiring is pinned by source + // rather than executed; the listener's behaviour is exercised above. + const connect = SOURCE.slice(SOURCE.indexOf('async connect()'), SOURCE.indexOf('async _loadBuffer()')); + expect(connect).toContain('this._installWheelListener();'); + expect(connect).toContain('this._onLiveClear();'); + expect(connect).not.toContain('this.terminal.clear();'); + }); +}); From 140ca35e2de0b52697c0adc60dc5f3513991fbc0 Mon Sep 17 00:00:00 2001 From: timkjr <timkjr@k-lab.lan> Date: Mon, 28 Sep 2026 19:59:22 -0500 Subject: [PATCH 37/46] fix(split-pane): keep the disconnected marker visible, skip detached sessions MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Address Ark0N's review on #506: - The history pull's own `\x1bc` reset erased the "Pane B disconnected" marker onclose wrote, painting a fresh, current-looking history while onData kept silently dropping every keystroke on the dead socket — a Codeman restart drops the socket while the tmux session (and so the HTTP pull) survives, making this easy to hit. onclose now tracks the closure via `_wsClosed` in addition to writing the marker (extracted into `_writeDisconnectedMarker()`), and a replay re-stamps it in the pull's `finally` block, after the live-frame flush, whichever order the close and the pull land in. - `_maybeLoadMoreHistory()` now stands aside for a detached session, mirroring `_sendResize()`'s existing check and app.js's `_maybeRefetchFullHistory()` — its own window already owns its PTY size and scrollback. - Wording: a non-shell CLI's history is out of scope for this pull, not absent (codex and Claude's inline renderer do grow tmux history); the alternate-screen skip only matters for a direct-PTY shell, since tmux never surfaces the alt buffer to the browser xterm. CLAUDE.md points at the invariants heading directly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- src/web/public/terminal-split.js | 30 ++++++++-- test/split-pane-terminal-unit.test.ts | 80 ++++++++++++++++++++++++++- 4 files changed, 106 insertions(+), 8 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 1bfee139..36bfcd40 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -270,7 +270,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. A Shell split-pane Pane B has its own copy of the bounded pull against its own xterm (`SplitTerminalPane._pullHistory`, terminal-split.js); keep the two in step. ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. A Shell split-pane Pane B has its own copy of the bounded pull against its own xterm (`SplitTerminalPane._pullHistory`, terminal-split.js); keep the two in step. → invariants: "Split-pane sessions" ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 3a7832b3..436821ec 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -767,7 +767,7 @@ Further detail: with many sessions the horizontal strip stops being scannable, w ### Split-pane sessions -**Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. ⚠️ A SHELL Pane B pulls scrollback itself when the wheel goes up at the top of its buffer (`_maybeLoadMoreHistory`/`_pullHistory`): tmux repaints a burst of output instead of scrolling it, so Pane B's own xterm holds about one screen of scrollback while tmux holds every line, and it loaded history exactly once at connect and never again. It is the same bounded pull as Pane A's (`?full=1&tail=TERMINAL_TAIL_SIZE`, no rewrite when the window holds no more rows than the pane already has or the pane is at its `scrollback + rows` cap, and a 60 s back-off instead of 4 s when that skipped window was truncated or the pane is full, since each ask costs the server a whole-history `capture-pane`), against Pane B's OWN terminal rather than `app.terminal`, so it cannot share `_maybeRefetchFullHistory`. The wheel listener is capture-phase because xterm `stopPropagation()`s the events it consumes; the alternate screen (nano, vim, less) is skipped; live frames arriving mid-replay, a `{t:'c'}` clear frame included, are held with their arrival time (`_liveQueue`) and replayed in order only if they arrived after the capture (the response's arrival stands in for the capture instant, as in `_finishBufferLoad`, so a frame inside that one round trip can be lost or doubled); the fetch has a 10 s deadline because it holds the pane's live output while it runs. There is no "Load full history" banner in Pane B, so a shell history past that 1 MiB window stays out of reach there. Non-shell Pane B is unchanged: it already loads `full=1`, and a repaint-mode CLI keeps no tmux history to recover. Design: `docs/split-pane-sessions-plan.md`. +**Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. ⚠️ A SHELL Pane B pulls scrollback itself when the wheel goes up at the top of its buffer (`_maybeLoadMoreHistory`/`_pullHistory`): tmux repaints a burst of output instead of scrolling it, so Pane B's own xterm holds about one screen of scrollback while tmux holds every line, and it loaded history exactly once at connect and never again. It is the same bounded pull as Pane A's (`?full=1&tail=TERMINAL_TAIL_SIZE`, no rewrite when the window holds no more rows than the pane already has or the pane is at its `scrollback + rows` cap, and a 60 s back-off instead of 4 s when that skipped window was truncated or the pane is full, since each ask costs the server a whole-history `capture-pane`), against Pane B's OWN terminal rather than `app.terminal`, so it cannot share `_maybeRefetchFullHistory`. The wheel listener is capture-phase because xterm `stopPropagation()`s the events it consumes; the alternate-screen skip (nano, vim, less) only matters for a direct-PTY shell, since under tmux the browser xterm never enters the alternate buffer; skipped too for a detached session (mirrors `_sendResize()`'s own check and app.js's `_maybeRefetchFullHistory`), since its own window already owns its PTY size and scrollback. Live frames arriving mid-replay, a `{t:'c'}` clear frame included, are held with their arrival time (`_liveQueue`) and replayed in order only if they arrived after the capture (the response's arrival stands in for the capture instant, as in `_finishBufferLoad`, so a frame inside that one round trip can be lost or doubled); the fetch has a 10 s deadline because it holds the pane's live output while it runs. ⚠️ A replay's own `\x1bc` would otherwise wipe the "Pane B disconnected" marker `onclose` wrote and paint a fresh, current-looking history while `onData` keeps silently dropping every keystroke on the dead socket (a Codeman restart drops the socket while the tmux session, and so the HTTP pull, still succeeds) — `_pullHistory()` re-stamps the marker after the live-frame flush when the socket closed in either order (before the pull started, or mid-fetch), tracked via `_wsClosed` rather than routed through `_onLiveOutput()`, since a close landing before the response is stamped before the cutoff and would be dropped with the rest of the pre-capture queue. There is no "Load full history" banner in Pane B, so a shell history past that 1 MiB window stays out of reach there. Non-shell Pane B is unchanged: it already loads `full=1`, and its history is out of scope for this pull (codex and Claude's inline renderer do grow tmux history; this just isn't how they recover it). Design: `docs/split-pane-sessions-plan.md`. ### Gesture control: the setting diff --git a/src/web/public/terminal-split.js b/src/web/public/terminal-split.js index cc9fedb2..fe031b4a 100644 --- a/src/web/public/terminal-split.js +++ b/src/web/public/terminal-split.js @@ -71,6 +71,7 @@ this.fitAddon = null; this.ws = null; this._wsReady = false; + this._wsClosed = false; this._destroyed = false; // Single-flight state for _loadBuffer()/_refreshBuffer() below. this._bufferLoading = false; @@ -294,7 +295,8 @@ // user's place in Pane B's scrollback for a transient blip. this.ws.onclose = () => { this._wsReady = false; - this.terminal?.write('\r\n\x1b[2m[Pane B disconnected — close and reopen the split to reconnect]\x1b[0m\r\n'); + this._wsClosed = true; + this._writeDisconnectedMarker(); }; this.ws.onerror = () => { @@ -302,6 +304,12 @@ }; } + // Extracted so both onclose and a history-pull replay that lands on an + // already-closed socket can write it (see _pullHistory()'s finally block). + _writeDisconnectedMarker() { + this.terminal?.write('\r\n\x1b[2m[Pane B disconnected — close and reopen the split to reconnect]\x1b[0m\r\n'); + } + // Fetches and writes the session's current scrollback. Used both by // connect() (initial load) and by the `{t:'r'}` server-refresh frame // (below) — the primary pane's own _onSessionNeedsRefresh (app.js) is @@ -386,12 +394,18 @@ // tmux holds every line — and nothing here ever went back to ask, so the // history was unreachable. The primary pane has the same pull // (app.js _maybeRefetchFullHistory); Pane B is a separate xterm and needs its - // own. Shell only: a repaint-mode agent CLI keeps no tmux history to recover, - // and its load already takes `full=1`. Skipped on the alternate screen - // (nano, vim, less), where the wheel belongs to the app, not the scrollback. + // own. Shell only: a non-shell CLI's history is out of scope for this pull + // (its load already takes `full=1`; codex and Claude's inline renderer do + // grow tmux history, this just isn't how they recover it). The alternate- + // screen skip (nano, vim, less) only matters for a direct-PTY shell — under + // tmux the browser xterm never enters the alternate buffer. _maybeLoadMoreHistory() { if (this.sessionMode !== 'shell' || this._destroyed || !this.terminal) return; if (this._bufferLoading) return; + // Mirrors app.js _maybeRefetchFullHistory and this pane's own + // _sendResize(): a detached session's own window already owns its PTY + // size and scrollback, so Pane B has nothing of its own to reconcile. + if (this.detachedSessions?.has(this.sessionId)) return; const active = this.terminal.buffer.active; if (active.type !== 'normal' || active.viewportY !== 0) return; // Momentum scrolling fires this dozens of times per flick, so cooldown @@ -477,6 +491,14 @@ if (entry.clear) this.terminal?.clear(); else this.terminal?.write(entry.data); } + // A replay's own `\x1bc` wipes the disconnected marker onclose wrote, + // painting a fresh, current-looking history while onData keeps + // silently dropping every keystroke on the dead socket. Re-stamp it + // if the socket closed in either order (before the pull started, or + // while the fetch was in flight) — checked after the queue flush so + // it is the last thing on screen, matching what onclose would have + // left had the pull never run. + if (replayed && this._wsClosed) this._writeDisconnectedMarker(); this._endBufferLoad(); } } diff --git a/test/split-pane-terminal-unit.test.ts b/test/split-pane-terminal-unit.test.ts index 17021c3f..90de98cc 100644 --- a/test/split-pane-terminal-unit.test.ts +++ b/test/split-pane-terminal-unit.test.ts @@ -54,6 +54,8 @@ type PaneUnderTest = { _historyPullUseless: boolean; _liveQueue: unknown[] | null; _onWheel: unknown; + _wsClosed: boolean; + detachedSessions: Set<string> | undefined; destroy(): void; _loadBuffer(): Promise<void>; _refreshBuffer(): void; @@ -62,6 +64,7 @@ type PaneUnderTest = { _onLiveOutput(data: string): void; _onLiveClear(): void; _installWheelListener(): void; + _writeDisconnectedMarker(): void; }; const fetchMock = vi.fn(); @@ -93,8 +96,12 @@ function loadSplitTerminalPane() { const SplitTerminalPane = loadSplitTerminalPane(); -function makePane(mode = 'claude', mount: unknown = {}): PaneUnderTest & { terminal: FakeTerminal } { - const pane = new SplitTerminalPane('s1', mount, { mode }); +function makePane( + mode = 'claude', + mount: unknown = {}, + opts: { detachedSessions?: Set<string> } = {} +): PaneUnderTest & { terminal: FakeTerminal } { + const pane = new SplitTerminalPane('s1', mount, { mode, ...opts }); pane.terminal = { // xterm invokes a write's callback once everything before it is parsed. write: vi.fn((_data: string, done?: () => void) => done?.()), @@ -322,6 +329,17 @@ describe('SplitTerminalPane scroll-to-top history pull', () => { expect(fetchMock).not.toHaveBeenCalled(); }); + it('stands aside for a detached session, mirroring _sendResize()', async () => { + // A detached session's own window already owns its PTY size and + // scrollback (buildSplitPickerSessions() already refuses to open one). + const pane = makePane('shell', {}, { detachedSessions: new Set(['s1']) }); + + pane._maybeLoadMoreHistory(); + await settle(); + + expect(fetchMock).not.toHaveBeenCalled(); + }); + it('a flick fires once: overlapping triggers are dropped, then the cooldown holds', async () => { const pane = makePane('shell'); const response = deferred<ReturnType<typeof jsonResponse>>(); @@ -647,4 +665,62 @@ describe('SplitTerminalPane scroll-to-top history pull', () => { expect(connect).toContain('this._onLiveClear();'); expect(connect).not.toContain('this.terminal.clear();'); }); + + it('re-stamps the disconnected marker after a replay if the socket closed before the pull started', async () => { + // onclose already wrote the marker once; a replay's own `\x1bc` would wipe + // it and paint a fresh, current-looking history while onData keeps + // silently dropping every keystroke on the dead socket. + const pane = makePane('shell'); + pane._wsClosed = true; + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(100))); + + void pane._pullHistory(); + await settle(); + + const marker = expect.stringContaining('Pane B disconnected'); + const writes = pane.terminal.write.mock.calls.map((c) => c[0]); + expect(writes.at(-1)).toEqual(expect.stringMatching(/Pane B disconnected/)); + expect(pane.terminal.write).toHaveBeenCalledWith(marker); + }); + + it('re-stamps the disconnected marker after a replay if the socket closes mid-fetch', async () => { + // The other order Ark0N's review called out: the close lands while the + // capture is in flight, so the HTTP pull still succeeds (a Codeman + // restart drops the WS while the tmux session, and so the pull, survives). + const pane = makePane('shell'); + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + const pull = pane._pullHistory(); + pane._wsClosed = true; // the close arrives mid-fetch, before the response + response.resolve(jsonResponse(rowsOf(100))); + await pull; + + expect(pane.terminal.write.mock.calls.at(-1)?.[0]).toEqual(expect.stringMatching(/Pane B disconnected/)); + }); + + it('does not re-stamp the marker when the socket is still open', async () => { + const pane = makePane('shell'); + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(100))); + + void pane._pullHistory(); + await settle(); + + for (const call of pane.terminal.write.mock.calls) { + expect(call[0]).toEqual(expect.not.stringMatching(/Pane B disconnected/)); + } + }); + + it('does not re-stamp the marker when the pull never replayed (skip/downgrade path)', async () => { + // Nothing erased the marker in this path, so re-stamping it would be a + // second, redundant write. + const pane = makePane('shell'); + pane._wsClosed = true; + fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(30))); // held in full already: no replay + + void pane._pullHistory(); + await settle(); + + expect(pane.terminal.write).not.toHaveBeenCalled(); + }); }); From 01eb8ef08a38da590b02b981d0628f9e4b6e3e38 Mon Sep 17 00:00:00 2001 From: Michael Grundberg <michael.grundberg@irisity.com> Date: Tue, 29 Sep 2026 18:03:55 +0200 Subject: [PATCH 38/46] feat(web): select a dashboard session from a #session=<id> link A page that keeps one Codeman window open, such as a task board, could only show a session by sending that window to /session/<id>, which loads the whole app again for every click. The dashboard now reads a #session=<id> fragment when it loads and on hashchange, selects that session, and removes the fragment with history.replaceState so the next identical link is still a change. Re-pointing a window that already shows the dashboard changes only the fragment, so the page stays loaded and the switch is a tab change. A link can name a session the dashboard does not list yet, because the page that created it may link before session:created arrives. The id waits until that event names it, and picking another tab yourself retires it. Following a link is an app selection (`auto: true`). The page that set the fragment may be a script, so it must not spend the session's idle alert. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- docs/architecture-invariants.md | 2 +- docs/extending-codeman.md | 27 ++++ src/web/public/app.js | 54 ++++++++ src/web/public/constants.js | 17 +++ test/session-select-ack-gate.test.ts | 10 +- test/url-session-fragment.test.ts | 179 +++++++++++++++++++++++++++ 6 files changed, 284 insertions(+), 5 deletions(-) create mode 100644 test/url-session-fragment.test.ts diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 848ae8f1..2751de8c 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -580,7 +580,7 @@ Either way the actual launch (`runCustomModelEntry`) routes through `run()` itse ⚠️ **Viewing a session ACKNOWLEDGES its idle item, it does not resolve it** (`POST /api/approvals/session/:sessionId/viewed` → `acknowledgedAt` → `approval:updated`): the item stays pending (still answerable, still Read My Mind context) and only stops arming the yellow tab alert. That flag is what makes the clear durable, since the view-clears-idle rule used to live in one browser's memory and `seedApprovals()` re-armed the alert on the next reload while other devices never heard about it at all; the local half is `markIdleAlertSeen()` (app.js), called from BOTH `selectSession` paths, including the already-active early return, where a click could otherwise never clear the alert. -⚠️ **Only a HUMAN opening a session acknowledges**: `selectSession(id, { auto: true })` marks the three selections the APP makes (boot restore, a solo window opening its target, the fallback after the active session is closed) and skips the acknowledgement, so a page load cannot silently spend an alert the user never saw. The flag defaults to user-initiated, so an untagged call site fails toward acknowledging rather than toward an alert nothing can clear; `test/session-select-ack-gate.test.ts` pins both the gate and the tagged call sites. Idle-only by construction (`acknowledge()` defaults to `['idle']`): looking at a permission/question dialog does not answer it. +⚠️ **Only a HUMAN opening a session acknowledges**: `selectSession(id, { auto: true })` marks the four selections the APP makes (boot restore, a solo window opening its target, a `#session=<id>` link from another page, the fallback after the active session is closed) and skips the acknowledgement, so a page load cannot silently spend an alert the user never saw. The flag defaults to user-initiated, so an untagged call site fails toward acknowledging rather than toward an alert nothing can clear; `test/session-select-ack-gate.test.ts` pins both the gate and the tagged call sites. Idle-only by construction (`acknowledge()` defaults to `['idle']`): looking at a permission/question dialog does not answer it. ⚠️ Same rule on the input path: `_ackDelivery` (app.js) spends the IDLE alert only, via that same `markIdleAlertSeen()`. It used to `clearPendingHooks(sessionId)` with no kind, so one keystroke wiped a RED alert on that device while the dialog was still up, the other devices stayed red, and a reload re-seeded it. diff --git a/docs/extending-codeman.md b/docs/extending-codeman.md index 2d440485..5e0bb605 100644 --- a/docs/extending-codeman.md +++ b/docs/extending-codeman.md @@ -338,6 +338,33 @@ codeman ralph start|stop|status|reset codeman users add|passwd|list codeman status | list | attach <path> codeman doctor ``` +### Opening a session from your own page + +To send someone from your page to one session, link to the dashboard with the +session id in the fragment, as in `http://127.0.0.1:3000/#session=<id>`. The +dashboard selects that tab when it loads. It also removes the fragment from its +own URL, so a later link to the same session still counts as a change. + +Keep reusing one named window to make later links fast: + +```js +window.open(`${codeman}/#session=${encodeURIComponent(id)}`, 'codeman'); +``` + +When that window already shows the dashboard, only the fragment differs. The +browser therefore keeps the page loaded, and the dashboard switches tabs without +reloading it. A session the window has shown before appears at once. A session +your page has only just created may not be listed yet, so the dashboard waits +for its `session:created` event and selects it then. + +Following a link does not count as someone looking at the session, so it +leaves the session's idle alert in place. The alert clears when the person +clicks the tab or types into the session. A link to a session that is popped +out into its own window asks that window to come forward, as clicking its tab does. + +A link to `/session/<id>` opens a page showing that session alone, and that +page loads from scratch for every link. + ## Seam 4: Hooks Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session. diff --git a/src/web/public/app.js b/src/web/public/app.js index 56997f21..c6662d90 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -602,6 +602,9 @@ class CodemanApp { // service-worker shell loads), with the server-injected global as a fallback. this.soloSessionId = this._detectSoloSessionId(); this.isSoloWindow = !!this.soloSessionId; + // A session another page asked for with a `#session=<id>` link. It waits + // here until the session list has that id (see _selectUrlSession). + this._urlSessionId = this.isSoloWindow ? null : this._takeUrlSession(); this.detachedSessions = new Set(); // dashboard-side: ids currently popped out this.detachedWindows = new Map(); // dashboard-side: id -> WindowProxy this._detachWatchTimers = new Map(); // dashboard-side: id -> setInterval handle @@ -997,6 +1000,16 @@ class CodemanApp { // strip never flashes before handleInit selects the target session. this._initWindowChannel(); if (this.isSoloWindow) document.body.classList.add('solo-mode'); + // A page holding this window switches its tab by changing only the + // fragment, which keeps the page loaded (see sessionIdFromFragment). + if (!this.isSoloWindow) { + window.addEventListener('hashchange', () => { + const id = this._takeUrlSession(); + if (!id) return; + this._urlSessionId = id; + this._selectUrlSession(); + }); + } // Initialize mobile handlers KeyboardHandler.init(); SwipeHandler.init(); @@ -1368,6 +1381,33 @@ class CodemanApp { } catch { return null; } } + /** Read a `#session=<id>` link off the URL and drop the fragment. The next + * link to the same session is then a change the browser reports, even + * after you have clicked away to another tab. Returns the id or null. */ + _takeUrlSession() { + const id = window.CodemanUrlSession?.sessionIdFromFragment(location.hash) ?? null; + if (id) { + try { history.replaceState(history.state, '', location.pathname + location.search); } catch {} + } + return id; + } + + /** Show the session a `#session=<id>` link asked for, once the session list + * has it. A page that has just created a session can link to it before + * session:created arrives here, so an unknown id stays pending and + * _onSessionCreated tries again. + * + * ⚠️ The selection is `auto`. The page that set the fragment may be a + * script, and this window may not even be in front, so following a link is + * not a human looking at the session and must not spend its idle alert. */ + _selectUrlSession() { + const id = this._urlSessionId; + if (!id || !this.sessions.has(id)) return false; + this._urlSessionId = null; + this.selectSession(id, { auto: true }); + return true; + } + /** * Pop a session out into its own browser window. SINGLE, idempotent entry * point: the tab's pop-out icon calls this, and a future gesture layer @@ -1949,6 +1989,7 @@ class CodemanApp { this.updateCost(); // Start stats polling when first session appears if (this.sessions.size === 1) this.startSystemStatsPolling(); + if (this._urlSessionId === data.id) this._selectUrlSession(); } _onSessionUpdated(data) { @@ -4367,6 +4408,13 @@ class CodemanApp { return; } + // A `#session=<id>` link wins over restoring the last active tab. + if (this._urlSessionId && this.sessions.has(this._urlSessionId)) { + this.activeSessionId = null; + this._selectUrlSession(); + return; + } + const previousActiveId = this.activeSessionId; if (this.sessionOrder.length === 0) { this.activeSessionId = null; @@ -6560,6 +6608,12 @@ class CodemanApp { } async selectSession(sessionId, options = {}) { + // Picking another tab yourself retires a `#session=<id>` link still + // waiting for its session, which would otherwise take the tab from you + // whenever that session turned up (see _selectUrlSession). + if (options?.auto !== true && this._urlSessionId && this._urlSessionId !== sessionId) { + this._urlSessionId = null; + } // If this session is popped out into its own window, raise that window // instead of showing it inline (focus-on-click for detached tabs). If we // owned a now-closed window, _raiseDetached re-docks and returns false so diff --git a/src/web/public/constants.js b/src/web/public/constants.js index 6f9d85c8..9bb29beb 100644 --- a/src/web/public/constants.js +++ b/src/web/public/constants.js @@ -1853,10 +1853,27 @@ function reconcilePtyGeometry(local, pty) { return { adopt: true, cols: pty.cols }; } +/** + * Which session does a dashboard URL's fragment ask for? Another page that + * holds the dashboard's window, such as a task board, points it at + * `/#session=<id>`. Only the fragment changes between two such links, so the + * browser keeps the page loaded and fires `hashchange`, and the dashboard + * switches tabs without reloading. Any other fragment asks for nothing. + * + * @param {string} hash - `location.hash`, with or without its leading `#` + * @returns {string|null} the session id, or null + */ +function sessionIdFromFragment(hash) { + const params = new URLSearchParams(String(hash || '').replace(/^#/, '')); + const id = params.get('session'); + return id && id.trim() ? id.trim() : null; +} + if (typeof window !== 'undefined') { window.CodemanHistoryFormat = { formatHistoryBytes, computeHistoryTruncationNotice, computeRewriteScrollLine }; window.CodemanFilePaths = { absoluteFilePathPattern, previewsInFileViewer, FILE_PREVIEW_EXTENSIONS }; window.CodemanTerminalLines = { terminalLogicalLine }; + window.CodemanUrlSession = { sessionIdFromFragment }; window.CodemanSplitPane = { clampDividerPercent, buildSplitPickerSessions, diff --git a/test/session-select-ack-gate.test.ts b/test/session-select-ack-gate.test.ts index 8a93b167..904b5de9 100644 --- a/test/session-select-ack-gate.test.ts +++ b/test/session-select-ack-gate.test.ts @@ -5,9 +5,10 @@ * `selectSession()` acknowledges the session's idle approval item server-side * (`markIdleAlertSeen` → `POST /api/approvals/session/:id/viewed`), which is * what makes "I checked it" survive a reload and reach the user's other - * devices. Three call sites are the APP choosing a session rather than the - * user: the boot restore, a solo (popped-out) window opening its target, and - * the fallback after the active session is deleted. Those pass `auto: true` + * devices. Four call sites are the APP choosing a session rather than the + * user: the boot restore, a solo (popped-out) window opening its target, a + * `#session=<id>` link from another page, and the fallback after the active + * session is deleted. Those pass `auto: true` * and must not spend the alert, or a yellow tab would clear itself every time * the page loaded and the user would never see it. * @@ -110,7 +111,7 @@ describe('selectSession acknowledgement gate', () => { }); describe('the call sites the app drives itself', () => { - // Source guard: these three are the reason the flag exists. If a refactor + // Source guard: these call sites are the reason the flag exists. If a refactor // moves or reformats them, fail loudly rather than silently going back to // "every page load clears the user's yellow tab". it.each([ @@ -118,6 +119,7 @@ describe('selectSession acknowledgement gate', () => { ['boot restore, first tab fallback', 'this.selectSession(this.sessionOrder[0], { auto: true });'], ['solo window opening its target', 'this.selectSession(this.soloSessionId, { auto: true });'], ['fallback after the active session is removed', 'this.selectSession(nextSessionId, { auto: true });'], + ['a #session=<id> link from another page', 'this.selectSession(id, { auto: true });'], ])('%s passes auto: true', (_label, call) => { expect(APP_SOURCE).toContain(call); }); diff --git a/test/url-session-fragment.test.ts b/test/url-session-fragment.test.ts new file mode 100644 index 00000000..c75a8fe7 --- /dev/null +++ b/test/url-session-fragment.test.ts @@ -0,0 +1,179 @@ +// test/url-session-fragment.test.ts +// Port: N/A (no server/browser — loads constants.js and app.js via `vm`, like session-select-ack-gate.test.ts). +// +// A page that holds the dashboard's window switches its tab with a +// `#session=<id>` link, and sessionIdFromFragment() is what reads the link. +import { readFileSync } from 'node:fs'; +import { resolve } from 'node:path'; +import vm from 'node:vm'; +import { performance } from 'node:perf_hooks'; +import { describe, expect, it, vi } from 'vitest'; + +function loadHelper() { + const context = vm.createContext({ window: {}, globalThis: {}, URLSearchParams }); + const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/constants.js'), 'utf8'); + vm.runInContext(source, context, { filename: 'constants.js' }); + return (context.window as { CodemanUrlSession: { sessionIdFromFragment: (hash: unknown) => string | null } }) + .CodemanUrlSession; +} + +describe('CodemanUrlSession.sessionIdFromFragment', () => { + const { sessionIdFromFragment } = loadHelper(); + + it('reads the id from a #session= fragment', () => { + expect(sessionIdFromFragment('#session=76763752-fa3a-40aa-a025-e1684c82d00e')).toBe( + '76763752-fa3a-40aa-a025-e1684c82d00e' + ); + }); + + it('accepts the fragment without its leading #', () => { + expect(sessionIdFromFragment('session=abc')).toBe('abc'); + }); + + it('decodes an encoded id', () => { + expect(sessionIdFromFragment('#session=' + encodeURIComponent('w1 my/app'))).toBe('w1 my/app'); + }); + + it('finds the id beside other fragment parameters', () => { + expect(sessionIdFromFragment('#tab=2&session=abc')).toBe('abc'); + }); + + it('asks for nothing when the fragment names no session', () => { + expect(sessionIdFromFragment('')).toBeNull(); + expect(sessionIdFromFragment('#')).toBeNull(); + expect(sessionIdFromFragment('#settings')).toBeNull(); + expect(sessionIdFromFragment('#session=')).toBeNull(); + expect(sessionIdFromFragment('#session=%20')).toBeNull(); + expect(sessionIdFromFragment(undefined)).toBeNull(); + }); +}); + +// The dashboard side: reading the link, holding an id it does not list yet, +// and handing the selection over. Loaded like session-select-ack-gate.test.ts, +// on a bare instance whose DOM-touching methods are stubbed. +function loadApp() { + const constants = readFileSync(resolve(import.meta.dirname, '../src/web/public/constants.js'), 'utf8'); + const app = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8'); + const location = { hash: '', pathname: '/', search: '' }; + const history = { + state: null, + replaceState: vi.fn((_state: unknown, _title: string, url: string) => { + location.hash = url.includes('#') ? url.slice(url.indexOf('#')) : ''; + }), + }; + const context = vm.createContext({ + console: { ...console, log: vi.fn(), warn: vi.fn(), error: vi.fn() }, + performance, + setInterval: vi.fn(), + clearInterval: vi.fn(), + setTimeout, + clearTimeout, + requestAnimationFrame: vi.fn(), + HTMLCanvasElement: class HTMLCanvasElement {}, + WebSocket: { OPEN: 1 }, + fetch: vi.fn(), + URLSearchParams, + location, + history, + document: { addEventListener: vi.fn(), getElementById: () => null, querySelector: () => null }, + localStorage: { length: 0, key: vi.fn(), getItem: vi.fn(), setItem: vi.fn(), removeItem: vi.fn() }, + window: { addEventListener: vi.fn(), removeEventListener: vi.fn() }, + MobileDetection: { isTouchDevice: () => false }, + }); + vm.runInContext(`${constants}\n${app}\nglobalThis.__CodemanApp = CodemanApp;`, context); + const CodemanApp = (context as { __CodemanApp: { prototype: object } }).__CodemanApp; + const make = (ids: string[]) => { + const inst = Object.create(CodemanApp.prototype) as Record<string, any>; + inst.sessions = new Map(ids.map((id) => [id, { id, name: id }])); + inst.sessionOrder = [...ids]; + inst.detachedSessions = new Set(); + inst.detachedWindows = new Map(); + inst.isSoloWindow = false; + inst._urlSessionId = null; + inst.selectSession = vi.fn(); + for (const stub of [ + 'saveSessionOrder', + 'markSessionTabEntering', + 'markTerminalEntering', + 'renderSessionTabs', + 'updateCost', + 'startSystemStatsPolling', + ]) { + inst[stub] = vi.fn(); + } + return inst; + }; + return { make, location, history, CodemanApp }; +} + +describe('dashboard handling of a #session=<id> link', () => { + it('reads the link and removes the fragment, so the same link counts as a change next time', () => { + const { make, location, history } = loadApp(); + const app = make(['a']); + location.hash = '#session=a'; + expect(app._takeUrlSession()).toBe('a'); + expect(history.replaceState).toHaveBeenCalledWith(null, '', '/'); + expect(location.hash).toBe(''); + }); + + it('leaves a URL without a session link alone', () => { + const { make, location, history } = loadApp(); + location.hash = '#settings'; + expect(make([])._takeUrlSession()).toBeNull(); + expect(history.replaceState).not.toHaveBeenCalled(); + }); + + it('selects a listed session as an app selection, which leaves its idle alert armed', () => { + const { make } = loadApp(); + const app = make(['a']); + app._urlSessionId = 'a'; + expect(app._selectUrlSession()).toBe(true); + expect(app.selectSession).toHaveBeenCalledWith('a', { auto: true }); + expect(app._urlSessionId).toBeNull(); + }); + + it('holds an unlisted id until session:created names it', () => { + const { make } = loadApp(); + const app = make([]); + app._urlSessionId = 'new'; + expect(app._selectUrlSession()).toBe(false); + expect(app.selectSession).not.toHaveBeenCalled(); + app._onSessionCreated({ id: 'other', name: 'other' }); + expect(app.selectSession).not.toHaveBeenCalled(); + app._onSessionCreated({ id: 'new', name: 'new' }); + expect(app.selectSession).toHaveBeenCalledWith('new', { auto: true }); + expect(app._urlSessionId).toBeNull(); + }); + + it('retires a waiting link when you pick another tab yourself', async () => { + const { make, CodemanApp } = loadApp(); + const app = make(['a', 'b']); + app.selectSession = (CodemanApp.prototype as Record<string, any>).selectSession; + app._urlSessionId = 'later'; + await app.selectSession('b').catch(() => {}); + expect(app._urlSessionId).toBeNull(); + }); + + it('keeps a waiting link through a selection the app makes itself', async () => { + const { make, CodemanApp } = loadApp(); + const app = make(['a', 'b']); + app.selectSession = (CodemanApp.prototype as Record<string, any>).selectSession; + app._urlSessionId = 'later'; + await app.selectSession('b', { auto: true }).catch(() => {}); + expect(app._urlSessionId).toBe('later'); + }); + + it('puts the link ahead of restoring the last active tab when the page loads', () => { + const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8'); + const link = source.indexOf('if (this._urlSessionId && this.sessions.has(this._urlSessionId))'); + const restore = source.indexOf("restoreId = localStorage.getItem('codeman-active-session')"); + expect(link).toBeGreaterThan(-1); + expect(link).toBeLessThan(restore); + }); + + it('never reads the link in a solo window', () => { + const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8'); + expect(source).toContain('this._urlSessionId = this.isSoloWindow ? null : this._takeUrlSession();'); + expect(source).toMatch(/if \(!this\.isSoloWindow\) \{\s*window\.addEventListener\('hashchange'/); + }); +}); From af4cc45cfd9ee1d6421479c2a2aa86ff4d7a7f4c Mon Sep 17 00:00:00 2001 From: d fei <dignfei@gmail.com> Date: Tue, 29 Sep 2026 18:02:05 -0700 Subject: [PATCH 39/46] fix(ui): stop Escape from stranding the keyboard after closing an overlay MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both the Command Palette and the Session Manager call `search.focus()` on open, and both closed by removing the `active` class and nothing else. Hiding a focused input does not hand focus back to anyone — the browser drops it on `<body>` — so after Escape closed the overlay every keystroke went nowhere and the user had to click the terminal before they could type again. Measured in headless chromium against a real shell session, one overlay at a time: overlay activeElement after Esc can type afterwards App Settings XTERM yes Session Options XTERM yes Token Stats XTERM yes Monitor Panel XTERM yes Session Manager BODY no <- fixed here Command Palette BODY no <- fixed here The four that worked did so because they use `FocusTrap`, whose `deactivate()` restores focus to whatever held it before. These two never got one. Every close path has the same hole — Escape, the close method, picking an item — so the restore lives in the close functions rather than in the global Escape chain. Deliberately only the save/restore half of `FocusTrap`, not the whole thing: `FocusTrap.activate()` moves focus to the first focusable element, which in neither overlay is the search box, so adopting it wholesale would trade "type a filter the moment it opens" for "focus survives the close" — and the former is the reason Cmd+K exists. The terminal fallback is gated on there being an active session: an overlay opened from the welcome screen has no terminal to return to, and focusing one on a phone summons the on-screen keyboard over a screen with no input on it. The five new cases were checked against the unfixed code first: four of them fail without this change. --- src/web/public/panels-ui.js | 50 +++++++++++++++++++++ test/command-palette-ui.test.ts | 78 +++++++++++++++++++++++++++++++++ 2 files changed, 128 insertions(+) diff --git a/src/web/public/panels-ui.js b/src/web/public/panels-ui.js index 8bdf5df8..81c7c990 100644 --- a/src/web/public/panels-ui.js +++ b/src/web/public/panels-ui.js @@ -339,6 +339,50 @@ Object.assign(CodemanApp.prototype, { return true; }, + /** + * Save what had focus before a search-first overlay takes it. + * + * The Command Palette and the Session Manager both call `search.focus()` on + * open, and both used to close by removing the `active` class and nothing + * else. Hiding the focused input does not hand focus back to anyone — the + * browser drops it on `<body>` — so after Escape closed the overlay every + * keystroke went nowhere and the user had to click the terminal to type + * again (measured: `document.activeElement` is BODY afterwards and the + * terminal emits no onData at all). Every close path has the same hole, so + * the restore lives in the close functions, not in the global Escape chain. + * + * ⚠️ Deliberately only the save/restore half of {@link FocusTrap}, not the + * whole thing. `FocusTrap.activate()` moves focus to the first focusable + * element, which in both of these overlays is not the search box — adopting + * it wholesale would fix the focus loss by breaking the thing Cmd+K exists + * for, typing a filter the moment it opens. + */ + _rememberOverlayFocus(key) { + this[key] = (typeof document !== 'undefined' && document.activeElement) || null; + }, + + /** + * Hand focus back to whatever {@link _rememberOverlayFocus} saved. + * + * ⚠️ The terminal fallback is gated on there being an active session: an + * overlay opened from the welcome screen has no terminal to return to, and + * focusing one on a phone summons the on-screen keyboard over a screen that + * has no input on it. + */ + _restoreOverlayFocus(key) { + const prev = this[key]; + this[key] = null; + const body = typeof document !== 'undefined' ? document.body : null; + // `isConnected === false` means the element was removed while the overlay + // was open (a re-render of the tab strip, say); anything else — including + // a stub with no such property — is treated as still focusable. + if (prev && prev !== body && prev.isConnected !== false && typeof prev.focus === 'function') { + prev.focus(); + return; + } + if (this.activeSessionId) this.terminal?.focus?.(); + }, + openCommandPalette() { const modal = document.getElementById('commandPaletteModal'); const search = document.getElementById('commandPaletteSearch'); @@ -351,6 +395,9 @@ Object.assign(CodemanApp.prototype, { this._wireCommandPalette(); this.renderCommandPalette(); + // BEFORE the steal, not after: `search.focus()` below is what loses the + // caller's focus, so the read has to happen while it is still there. + this._rememberOverlayFocus('_commandPalettePrevFocus'); search.focus(); search.select?.(); }, @@ -358,6 +405,7 @@ Object.assign(CodemanApp.prototype, { closeCommandPalette() { const modal = document.getElementById('commandPaletteModal'); if (modal) modal.classList.remove('active'); + this._restoreOverlayFocus('_commandPalettePrevFocus'); }, _wireCommandPalette() { @@ -618,6 +666,7 @@ Object.assign(CodemanApp.prototype, { }); } search.value = ''; + this._rememberOverlayFocus('_sessionManagerPrevFocus'); search.focus(); } await this._loadSessionManagerList(''); @@ -626,6 +675,7 @@ Object.assign(CodemanApp.prototype, { closeSessionManager() { const modal = document.getElementById('sessionManagerModal'); if (modal) modal.classList.remove('active'); + this._restoreOverlayFocus('_sessionManagerPrevFocus'); }, /** Replace the Session Manager list body with a single status line. */ diff --git a/test/command-palette-ui.test.ts b/test/command-palette-ui.test.ts index bbb6813d..70e19ced 100644 --- a/test/command-palette-ui.test.ts +++ b/test/command-palette-ui.test.ts @@ -480,6 +480,84 @@ describe('Session Manager unified list', () => { }); }); +describe('overlay focus restoration (Escape must not strand the keyboard)', () => { + /** + * Both overlays focus their search box on open. Closing them used to leave + * focus on <body>, so after Escape every keystroke went nowhere until the + * user clicked the terminal — measured in a real browser against a shell + * session: activeElement BODY, zero onData for anything typed afterwards. + */ + function focusHarness() { + const terminalTextarea = { focus: vi.fn(), isConnected: true }; + const priorElement = { focus: vi.fn(), isConnected: true, tagName: 'TEXTAREA' }; + const body = { tagName: 'BODY' }; + let active: any = priorElement; + const { app, elements } = loadPaletteHarness({ + document: { + getElementById: (id: string) => (globalThis as any).__els?.[id] ?? null, + get activeElement() { + return active; + }, + body, + }, + }); + (globalThis as any).__els = elements; + app.terminal = { focus: terminalTextarea.focus }; + app.activeSessionId = 'sess-beta'; + // The overlay's own focus() is what moves focus in a real browser; the + // fake document needs the same transition or the test proves nothing. + elements.commandPaletteSearch.focus = vi.fn(() => { + active = elements.commandPaletteSearch; + }); + return { app, elements, priorElement, terminalTextarea, body, setActive: (v: any) => (active = v) }; + } + + it('returns focus to whatever had it when the command palette closes', () => { + const { app, priorElement } = focusHarness(); + app.openCommandPalette(); + expect(priorElement.focus).not.toHaveBeenCalled(); + app.closeCommandPalette(); + expect(priorElement.focus).toHaveBeenCalledTimes(1); + }); + + it('falls back to the terminal when the prior element is gone, but only with a live session', () => { + const { app, priorElement, terminalTextarea } = focusHarness(); + app.openCommandPalette(); + priorElement.isConnected = false; + app.closeCommandPalette(); + expect(priorElement.focus).not.toHaveBeenCalled(); + expect(terminalTextarea.focus).toHaveBeenCalledTimes(1); + }); + + it('never focuses the terminal from the welcome screen (a phone would pop the keyboard)', () => { + const { app, priorElement, terminalTextarea } = focusHarness(); + app.activeSessionId = null; + app.openCommandPalette(); + priorElement.isConnected = false; + app.closeCommandPalette(); + expect(terminalTextarea.focus).not.toHaveBeenCalled(); + }); + + it('does not restore focus to <body>, which is the bug itself', () => { + const { app, terminalTextarea, body, setActive } = focusHarness(); + setActive(body); + app.openCommandPalette(); + app.closeCommandPalette(); + expect(terminalTextarea.focus).toHaveBeenCalledTimes(1); + }); + + it('restores focus on the session manager too, not just the palette', async () => { + const { app, elements, priorElement } = focusHarness(); + elements.sessionManagerModal = { classList: { add: vi.fn(), remove: vi.fn() }, addEventListener: vi.fn() }; + elements.sessionManagerSearch = { value: '', focus: vi.fn(), addEventListener: vi.fn() }; + elements.sessionManagerList = { replaceChildren: vi.fn(), appendChild: vi.fn() }; + app._loadSessionManagerList = vi.fn(); + await app.openSessionManager(); + app.closeSessionManager(); + expect(priorElement.focus).toHaveBeenCalledTimes(1); + }); +}); + describe('panel close helpers', () => { it('closes panels when the mobile header helper is unavailable', () => { const CodemanApp = function CodemanApp(this: any) {}; From 57f7a775737e7ee51277b81dbeb432f36565edc1 Mon Sep 17 00:00:00 2001 From: d fei <dignfei@gmail.com> Date: Wed, 30 Sep 2026 08:35:06 -0700 Subject: [PATCH 40/46] fix(ui): restore focus only when the overlay was actually open MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Review feedback. The global Escape handler in app.js calls both `closeSessionManager()` and `closeCommandPalette()` on every Escape, whether or not either overlay is open, in the capture phase. Nothing was saved in that case, so `_restoreOverlayFocus()` fell through to `terminal.focus()` and moved focus before the focused element's own Escape handler ran: - split view: with focus in Pane B, keys typed after Escape went to Pane A - any text field (File Viewer editor, search and history filters, case picker): keys typed after Escape went into the terminal - inline tab rename: the capture-phase focus fired the input's blur (which commits) before its own Escape handler (which cancels), so Escape committed the rename instead of cancelling it Both close methods now bail out on `classList.contains('active')`. Separately, gating the terminal fallback on `activeSessionId` alone only covered the welcome screen. On a touch device with the keyboard down, focus sits on `<body>`, so closing the Session Manager focused the terminal and brought the keyboard up — `selectSession()` deliberately skips that focus, and this overrode it. It now goes through `_shouldFocusTerminalForTabSwitch()`. Tests: the Session Manager case's modal stub now uses the harness's `makeClassList()` (without `contains` the new guard reads it as "not open" and skips the restore the case is about), plus two new cases — closing either overlay without opening it first with an active session asserts the terminal was not focused, which is the path the global Escape chain takes and none of the five existing cases covered, and a touch device with the keyboard down asserts the same. Each was checked against the unguarded code: removing either guard turns exactly its own case red. --- src/web/public/panels-ui.js | 24 ++++++++++++++++++++--- test/command-palette-ui.test.ts | 34 ++++++++++++++++++++++++++++----- 2 files changed, 50 insertions(+), 8 deletions(-) diff --git a/src/web/public/panels-ui.js b/src/web/public/panels-ui.js index 81c7c990..0e5fdfec 100644 --- a/src/web/public/panels-ui.js +++ b/src/web/public/panels-ui.js @@ -380,7 +380,15 @@ Object.assign(CodemanApp.prototype, { prev.focus(); return; } - if (this.activeSessionId) this.terminal?.focus?.(); + // ⚠️ `activeSessionId` alone only covers the welcome screen. On a touch device + // with the keyboard down, focus sits on `<body>`, so focusing the terminal here + // would summon the on-screen keyboard — `selectSession()` deliberately skips the + // focus for exactly that reason, and this would override it. The app's own + // predicate already encodes the rule (true on desktop, on touch only while the + // keyboard is open); the optional call keeps the vm test harness working. + if (this.activeSessionId && this._shouldFocusTerminalForTabSwitch?.() !== false) { + this.terminal?.focus?.(); + } }, openCommandPalette() { @@ -404,7 +412,15 @@ Object.assign(CodemanApp.prototype, { closeCommandPalette() { const modal = document.getElementById('commandPaletteModal'); - if (modal) modal.classList.remove('active'); + // ⚠️ Bail out when it was not open. The global Escape handler calls this on + // EVERY Escape (app.js), in the CAPTURE phase, so an unconditional restore + // runs before the focused element's own Escape handler and steals focus into + // the terminal: keys typed after Escape in split Pane B land in Pane A, keys + // typed in any text field land in the terminal, and the inline tab rename's + // Escape fires the input's blur (which commits) before its own handler + // (which cancels), turning a cancel into a rename. + if (!modal?.classList?.contains('active')) return; + modal.classList.remove('active'); this._restoreOverlayFocus('_commandPalettePrevFocus'); }, @@ -674,7 +690,9 @@ Object.assign(CodemanApp.prototype, { closeSessionManager() { const modal = document.getElementById('sessionManagerModal'); - if (modal) modal.classList.remove('active'); + // Same guard as closeCommandPalette — see the note there. + if (!modal?.classList?.contains('active')) return; + modal.classList.remove('active'); this._restoreOverlayFocus('_sessionManagerPrevFocus'); }, diff --git a/test/command-palette-ui.test.ts b/test/command-palette-ui.test.ts index 70e19ced..00048010 100644 --- a/test/command-palette-ui.test.ts +++ b/test/command-palette-ui.test.ts @@ -111,7 +111,7 @@ function loadPaletteHarness(overrides: Record<string, any> = {}) { app.getSessionName = (session: any) => session.name || session.workingDir?.split('/').pop() || app.getShortId(session.id); - return { app, elements, listeners }; + return { app, elements, listeners, makeClassList }; } describe('Command-K session palette', () => { @@ -492,7 +492,7 @@ describe('overlay focus restoration (Escape must not strand the keyboard)', () = const priorElement = { focus: vi.fn(), isConnected: true, tagName: 'TEXTAREA' }; const body = { tagName: 'BODY' }; let active: any = priorElement; - const { app, elements } = loadPaletteHarness({ + const { app, elements, makeClassList } = loadPaletteHarness({ document: { getElementById: (id: string) => (globalThis as any).__els?.[id] ?? null, get activeElement() { @@ -509,7 +509,7 @@ describe('overlay focus restoration (Escape must not strand the keyboard)', () = elements.commandPaletteSearch.focus = vi.fn(() => { active = elements.commandPaletteSearch; }); - return { app, elements, priorElement, terminalTextarea, body, setActive: (v: any) => (active = v) }; + return { app, elements, priorElement, terminalTextarea, body, makeClassList, setActive: (v: any) => (active = v) }; } it('returns focus to whatever had it when the command palette closes', () => { @@ -546,9 +546,33 @@ describe('overlay focus restoration (Escape must not strand the keyboard)', () = expect(terminalTextarea.focus).toHaveBeenCalledTimes(1); }); + it('leaves focus alone when neither overlay was open — the path every Escape takes', () => { + // app.js's global Escape handler calls both close methods on EVERY Escape, + // in the capture phase. Nothing was saved, so an unguarded restore would fall + // through to the terminal and steal focus from split Pane B, from any text + // field, and turn the inline rename's Escape into a commit. + const { app, elements, terminalTextarea, priorElement, makeClassList } = focusHarness(); + elements.sessionManagerModal = { classList: makeClassList(), addEventListener: vi.fn() }; + app.closeCommandPalette(); + app.closeSessionManager(); + expect(terminalTextarea.focus).not.toHaveBeenCalled(); + expect(priorElement.focus).not.toHaveBeenCalled(); + }); + + it('does not focus the terminal on touch while the keyboard is down', () => { + const { app, terminalTextarea, body, setActive } = focusHarness(); + app._shouldFocusTerminalForTabSwitch = () => false; + setActive(body); + app.openCommandPalette(); + app.closeCommandPalette(); + expect(terminalTextarea.focus).not.toHaveBeenCalled(); + }); + it('restores focus on the session manager too, not just the palette', async () => { - const { app, elements, priorElement } = focusHarness(); - elements.sessionManagerModal = { classList: { add: vi.fn(), remove: vi.fn() }, addEventListener: vi.fn() }; + const { app, elements, priorElement, makeClassList } = focusHarness(); + // A real classList: the close guard reads `contains('active')`, and a stub + // without it reports "not open" and skips the restore this test is about. + elements.sessionManagerModal = { classList: makeClassList(), addEventListener: vi.fn() }; elements.sessionManagerSearch = { value: '', focus: vi.fn(), addEventListener: vi.fn() }; elements.sessionManagerList = { replaceChildren: vi.fn(), appendChild: vi.fn() }; app._loadSessionManagerList = vi.fn(); From 988f111cd0dfc6bc35d852829538d8f755c158e5 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Thu, 1 Oct 2026 11:13:44 +0200 Subject: [PATCH 41/46] fix(ui): keep a focus the row menu moved when closing the Session Manager (#509 review) - _restoreOverlayFocus(key, modal) now leaves focus alone when something outside the overlay already holds it (not <body>, not inside the modal). The Session Manager's "Switch to session" and "Open folder" call selectSession() before closeSessionManager(), and the restore was pulling focus back from the terminal to the header button. Both close methods pass their modal; a regression test drives that order. - Test harness: focusHarness() routes getElementById through a local binding instead of leaking globalThis.__els, and its modal stubs report their own search box as contained, as the real DOM does. - CLAUDE.md and docs/architecture-invariants.md: record that the global Escape handler calls every close method on every Escape (capture phase), so a close method with side effects must return early when not open. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- src/web/public/panels-ui.js | 15 ++++++-- test/command-palette-ui.test.ts | 62 +++++++++++++++++++++++++++------ 4 files changed, 66 insertions(+), 15 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index a5ed52bc..2e5ffb1b 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -335,7 +335,7 @@ Frontend JS modules have `@fileoverview` with `@dependency`/`@loadorder` tags. L **Welcome "Resume Conversation" list** (terminal-ui.js): `loadHistorySessions()` fetches once and caches the corpus on `_historyAll`/`_historyCases`; every subsequent view (filter box, sort select, expand, the periodic refresh in panels-ui.js) goes through `_renderHistoryList()`, so never append rows to `#historyList` directly or re-fetch to re-sort. ⚠️ The box height is **class-driven**: expanding the list without `.history-list.expanded` leaves the collapsed `max-height` in place and just deepens a scroll well, which is the bug #260 reported (35 sessions in a ~4-row box). ⚠️ The A–Z sort keys off `_historyRowLabel()`, the SAME string the row renders (`name || firstPrompt || path`), most rows are transcript-backed and have no session name, so sorting on `name` alone silently does nothing. ⚠️ A filter implies expansion, and `_renderSearch()` hides `#historyHeader` (title + controls) as one unit while a search is active. Tests: `test/history-list-controls.test.ts`. -**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)**: with no selection it must `return true` without `preventDefault()` or the interrupt is lost; keep `copyTerminalSelection` out of `SHORTCUT_ACTIONS`. The gate tests the CLEANED selection (`CodemanCopySelection.clean`: trailing padding, plus a LEADING margin only up to the width the CLI declares in `capabilities.transcriptGutter`, never one derived from the pane); the strip is not idempotent, so clean once and pass the RAW selection on, and leave Alt+drag column selections untouched. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) +**Command palette + shortcut registry**: `Ctrl/Cmd/Alt+K` opens the session palette; shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js, overrides in `settings.shortcutOverrides`). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte into the PTY. ⚠️ The global Escape handler (app.js) calls EVERY close method on every Escape, in the capture phase, so a close method that does more than hide (the palette and Session Manager restore focus) must return early when its overlay is not open. ⚠️ `saveAppSettings()` rebuilds settings from the DOM, so keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. ⚠️ **Smart copy (`Ctrl+C`)**: with no selection it must `return true` without `preventDefault()` or the interrupt is lost; keep `copyTerminalSelection` out of `SHORTCUT_ACTIONS`. The gate tests the CLEANED selection (`CodemanCopySelection.clean`: trailing padding, plus a LEADING margin only up to the width the CLI declares in `capabilities.transcriptGutter`, never one derived from the pane); the strip is not idempotent, so clean once and pass the RAW selection on, and leave Alt+drag column selections untouched. → [architecture-invariants#command-palette-and-shortcut-registry](docs/architecture-invariants.md#command-palette-and-shortcut-registry) **Per-device vs synced settings**: the `displayKeys` set in settings-ui.js is a **client-side merge policy**, not a wire filter. A display key seeds from the server only when localStorage has no value for it, which is what prevents one device overwriting another; `showPlanUsageLimits` is additionally `delete`d from the incoming payload outright. Separately, `SettingsUpdateSchema` is `.strict()` and simply **does not declare** `skin`, `showFileViewerButton`, `showCronButton`, `webglRendererEnabled`, `localEchoEnabled`, `cjkInputEnabled`, or `extendedKeyboardBar`, so sending one of those is a validation error. The rest (`showResponseViewer`, `showPlanUsageLimits`, `language`, and most `show*` keys) ARE in the schema and do persist server-side; they are per-device by client policy only. ⚠️ Adding a new per-device setting means deciding **both** questions: membership in `displayKeys`, and presence in the schema. diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index b3b1e9a3..8acb6ba3 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -653,7 +653,7 @@ Matching semantics: a query containing `/` matches the relative path, otherwise ### Command palette and shortcut registry -**Command palette + shortcut registry** (COD-151/153/157/192, #146): `Ctrl/Cmd/Alt+K` opens the session palette (fuzzy search over live sessions; "Browse all sessions" → the Session Manager modal backed by `GET /api/sessions/unified`); the quick-start case `<select>` is fronted by a searchable picker (`buildCasePickerOptions`/`formatCasePickerLabel` — remote cases render `name @ hostId`). Shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js; overrides persist under `settings.shortcutOverrides` via `saveAppSettingsToStorage`); App Settings → Shortcuts renders capture/disable rows; `Ctrl+?` opens the registry-driven overlay (footer links to the full `#helpModal` reference). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ `saveAppSettings()` rebuilds settings from the DOM — keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. +**Command palette + shortcut registry** (COD-151/153/157/192, #146): `Ctrl/Cmd/Alt+K` opens the session palette (fuzzy search over live sessions; "Browse all sessions" → the Session Manager modal backed by `GET /api/sessions/unified`); the quick-start case `<select>` is fronted by a searchable picker (`buildCasePickerOptions`/`formatCasePickerLabel` — remote cases render `name @ hostId`). Shortcuts live in a rebindable registry (`DEFAULT_SHORTCUTS`/`getShortcutRegistry()`/`matchesShortcutEvent()` in app.js; overrides persist under `settings.shortcutOverrides` via `saveAppSettingsToStorage`); App Settings → Shortcuts renders capture/disable rows; `Ctrl+?` opens the registry-driven overlay (footer links to the full `#helpModal` reference). ⚠️ Palette-chord keys must ALSO be swallowed in `attachCustomKeyEventHandler` (terminal-ui.js) or xterm writes the control byte (0x0B) into the PTY. ⚠️ The global Escape handler (app.js) calls every overlay close method on every Escape, in the capture phase and so before the focused element's own Escape handler, which means a close method with side effects beyond hiding must return early when its overlay is not open: `closeCommandPalette()` and `closeSessionManager()` hand focus back on close (`_restoreOverlayFocus()`, panels-ui.js), and run unconditionally that restore stole focus from split Pane B and from any text field, and turned the inline tab rename's Escape into a commit. ⚠️ `saveAppSettings()` rebuilds settings from the DOM — keys edited elsewhere (`shortcutOverrides`, `showTokenCount`, `showCost`) need explicit `_prev` carry-over. **Terminal smart copy** (#211): `Ctrl+C` copies the selection when there is one and stays the interrupt when there isn't. Three rules keep that split honest, and breaking any of them silently costs the user their interrupt key: diff --git a/src/web/public/panels-ui.js b/src/web/public/panels-ui.js index a5aee874..fe5ada63 100644 --- a/src/web/public/panels-ui.js +++ b/src/web/public/panels-ui.js @@ -382,11 +382,20 @@ Object.assign(CodemanApp.prototype, { * overlay opened from the welcome screen has no terminal to return to, and * focusing one on a phone summons the on-screen keyboard over a screen that * has no input on it. + * + * ⚠️ A focus that already left the overlay is kept, not overridden. The + * Session Manager's row menu ("Switch to session", "Open folder") calls + * selectSession(), which focuses the terminal, BEFORE closeSessionManager(); + * restoring there would pull focus back to the header button that opened the + * modal. Only a focus still inside `modal`, or one dropped on `<body>`, is + * the overlay's to hand back. */ - _restoreOverlayFocus(key) { + _restoreOverlayFocus(key, modal) { const prev = this[key]; this[key] = null; const body = typeof document !== 'undefined' ? document.body : null; + const current = typeof document !== 'undefined' ? document.activeElement : null; + if (current && current !== body && modal?.contains?.(current) === false) return; // `isConnected === false` means the element was removed while the overlay // was open (a re-render of the tab strip, say); anything else — including // a stub with no such property — is treated as still focusable. @@ -435,7 +444,7 @@ Object.assign(CodemanApp.prototype, { // (which cancels), turning a cancel into a rename. if (!modal?.classList?.contains('active')) return; modal.classList.remove('active'); - this._restoreOverlayFocus('_commandPalettePrevFocus'); + this._restoreOverlayFocus('_commandPalettePrevFocus', modal); }, _wireCommandPalette() { @@ -707,7 +716,7 @@ Object.assign(CodemanApp.prototype, { // Same guard as closeCommandPalette — see the note there. if (!modal?.classList?.contains('active')) return; modal.classList.remove('active'); - this._restoreOverlayFocus('_sessionManagerPrevFocus'); + this._restoreOverlayFocus('_sessionManagerPrevFocus', modal); }, /** Replace the Session Manager list body with a single status line. */ diff --git a/test/command-palette-ui.test.ts b/test/command-palette-ui.test.ts index 00048010..d93fbcda 100644 --- a/test/command-palette-ui.test.ts +++ b/test/command-palette-ui.test.ts @@ -492,16 +492,19 @@ describe('overlay focus restoration (Escape must not strand the keyboard)', () = const priorElement = { focus: vi.fn(), isConnected: true, tagName: 'TEXTAREA' }; const body = { tagName: 'BODY' }; let active: any = priorElement; + // The harness builds its element map internally, so getElementById reads it + // through this binding, filled in once the harness returns. + let els: Record<string, any> = {}; const { app, elements, makeClassList } = loadPaletteHarness({ document: { - getElementById: (id: string) => (globalThis as any).__els?.[id] ?? null, + getElementById: (id: string) => els[id] ?? null, get activeElement() { return active; }, body, }, }); - (globalThis as any).__els = elements; + els = elements; app.terminal = { focus: terminalTextarea.focus }; app.activeSessionId = 'sess-beta'; // The overlay's own focus() is what moves focus in a real browser; the @@ -509,7 +512,36 @@ describe('overlay focus restoration (Escape must not strand the keyboard)', () = elements.commandPaletteSearch.focus = vi.fn(() => { active = elements.commandPaletteSearch; }); - return { app, elements, priorElement, terminalTextarea, body, makeClassList, setActive: (v: any) => (active = v) }; + // The restore keeps a focus that already left the overlay, so the modal has + // to know its own search box is inside it, as the real DOM does. + elements.commandPaletteModal.contains = (el: any) => el === elements.commandPaletteSearch; + // Same wiring for the Session Manager. A real classList: the close guard + // reads `contains('active')`, and a stub without it reports "not open" and + // skips the restore. + const installSessionManager = () => { + const search: any = { value: '', addEventListener: vi.fn() }; + search.focus = vi.fn(() => { + active = search; + }); + elements.sessionManagerSearch = search; + elements.sessionManagerModal = { + classList: makeClassList(), + addEventListener: vi.fn(), + contains: (el: any) => el === search, + }; + elements.sessionManagerList = { replaceChildren: vi.fn(), appendChild: vi.fn() }; + app._loadSessionManagerList = vi.fn(); + }; + return { + app, + elements, + priorElement, + terminalTextarea, + body, + makeClassList, + installSessionManager, + setActive: (v: any) => (active = v), + }; } it('returns focus to whatever had it when the command palette closes', () => { @@ -569,17 +601,27 @@ describe('overlay focus restoration (Escape must not strand the keyboard)', () = }); it('restores focus on the session manager too, not just the palette', async () => { - const { app, elements, priorElement, makeClassList } = focusHarness(); - // A real classList: the close guard reads `contains('active')`, and a stub - // without it reports "not open" and skips the restore this test is about. - elements.sessionManagerModal = { classList: makeClassList(), addEventListener: vi.fn() }; - elements.sessionManagerSearch = { value: '', focus: vi.fn(), addEventListener: vi.fn() }; - elements.sessionManagerList = { replaceChildren: vi.fn(), appendChild: vi.fn() }; - app._loadSessionManagerList = vi.fn(); + const { app, priorElement, installSessionManager } = focusHarness(); + installSessionManager(); await app.openSessionManager(); app.closeSessionManager(); expect(priorElement.focus).toHaveBeenCalledTimes(1); }); + + it('keeps the terminal focus the row menu gave it when the session manager closes', async () => { + // "Switch to session" and "Open folder" (terminal-ui.js) call selectSession(), + // which focuses the terminal on desktop, and only THEN closeSessionManager(). + // Opened from its header button, the saved focus is that button, so an + // unconditional restore pulled focus off the session the user just picked. + const { app, priorElement, terminalTextarea, installSessionManager, setActive } = focusHarness(); + installSessionManager(); + await app.openSessionManager(); + app.selectSession = vi.fn(() => setActive(terminalTextarea)); + app.selectSession('sess-alpha'); + app.closeSessionManager(); + expect(priorElement.focus).not.toHaveBeenCalled(); + expect(terminalTextarea.focus).not.toHaveBeenCalled(); + }); }); describe('panel close helpers', () => { From 3af1ff6fae737e73bac2b848a3df9b9ad7b5d94a Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Thu, 1 Oct 2026 11:14:35 +0200 Subject: [PATCH 42/46] fix(split-pane): keep Pane B's disconnected marker last when the socket closes mid-pull (#506 review) - terminal-split.js: move the socket's close into _onSocketClosed(), which defers the marker while a history pull holds live output (_liveQueue); _pullHistory() records closedBefore and its finally writes the marker after the queue flush when the socket closed during the pull, replayed or not, so it never lands above held frames or between replay chunks - tests: drive the real close path for a close mid-fetch ending in a skip, a downgrade or a failed fetch, a close during the chunked replay, and a close with no pull running; pin the onclose wiring in the static guard; describe the mid-fetch case on its own - CLAUDE.md: turn the plain-text split-pane pointer into a link - architecture-invariants.md: describe the deferred marker Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 2 +- src/web/public/terminal-split.js | 36 +++++++----- test/split-pane-terminal-unit.test.ts | 84 +++++++++++++++++++++++++-- 4 files changed, 103 insertions(+), 21 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 2e5ffb1b..450bfcbe 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -270,7 +270,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **Circuit breakers**: the Ralph breaker prevents respawn thrashing (`CLOSED` → `HALF_OPEN` → `OPEN`; reset via `/api/sessions/:id/ralph-circuit-breaker/reset`). **Distinct: the PTY-exit breaker** (`session-pty-exit-breaker.ts`) trips after repeated rapid PTY exits and blocks auto-restarts. ⚠️ It resets ONLY via an explicit `{clearBreaker:true}` body on `POST /api/sessions/:id/interactive`; the frontend's auto-reattach in `selectSession()` sends no body and must never clear it. → [architecture-invariants#circuit-breakers-ralph--pty-exit](docs/architecture-invariants.md#circuit-breakers-ralph-and-pty-exit) -**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. A Shell split-pane Pane B has its own copy of the bounded pull against its own xterm (`SplitTerminalPane._pullHistory`, terminal-split.js); keep the two in step. → invariants: "Split-pane sessions" ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) +**Full-scrollback replay**: `GET /api/sessions/:id/terminal?full=1` returns the whole tmux scrollback ALONE (`source='mux-full-history'`), superseding the byte buffer. First load of each non-shell TUI session requests it (`_fullHistoryLoaded`); Shell selection and drop recovery use a bounded 1 MiB `?tail=`, and a Shell scroll-to-top pulls a bounded `?full=1&tail=` window (a window no longer than the browser's buffer is skipped before the downgrade guard, so it never marks the session exhausted); the unbounded pull stays behind **Load full history**. A Shell split-pane Pane B has its own copy of the bounded pull against its own xterm (`SplitTerminalPane._pullHistory`, terminal-split.js); keep the two in step. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) ⚠️ The capture ends with a RELATIVE cursor move back to the pane's caret (never `CUP`), so no line-deleting transform may run over it; those skips key on `isFullCapture`, never on `?full=1` alone. ⚠️ A re-pull must never shrink the buffer (`_replayWouldShrinkBuffer()`). ⚠️ `captureCols`/`captureRows` are absent when no frame was positioned: test `Number.isFinite`, never truthiness. ⚠️ A frame dropped at the 128 KiB render cap MUST be recovered, and the recovery verifies itself: `_scheduleDroppedOutputRecovery` re-arms (bounded by `DROP_RECOVERY_MAX_ATTEMPTS`) while `_onSessionNeedsRefresh` reports no repaint, but never after a capture-fetch `'deadline'`. → [architecture-invariants#full-scrollback-replay](docs/architecture-invariants.md#full-scrollback-replay) **Split-pane sessions** (`showSplitButton`, header button, default OFF, desktop-only, per-device): a second live session ("Pane B") beside the active one, in its own `SplitTerminalPane` (terminal-split.js) with its own xterm + WebSocket, resizable via a draggable divider. Deliberately plainer than the primary pane — no local-echo overlay, CJK IME, or touch handlers — and NOT persisted across reloads. → [architecture-invariants#split-pane-sessions](docs/architecture-invariants.md#split-pane-sessions) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 8acb6ba3..32f6c172 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -780,7 +780,7 @@ Further detail: with many sessions the horizontal strip stops being scannable, w ### Split-pane sessions -**Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. ⚠️ A SHELL Pane B pulls scrollback itself when the wheel goes up at the top of its buffer (`_maybeLoadMoreHistory`/`_pullHistory`): tmux repaints a burst of output instead of scrolling it, so Pane B's own xterm holds about one screen of scrollback while tmux holds every line, and it loaded history exactly once at connect and never again. It is the same bounded pull as Pane A's (`?full=1&tail=TERMINAL_TAIL_SIZE`, no rewrite when the window holds no more rows than the pane already has or the pane is at its `scrollback + rows` cap, and a 60 s back-off instead of 4 s when that skipped window was truncated or the pane is full, since each ask costs the server a whole-history `capture-pane`), against Pane B's OWN terminal rather than `app.terminal`, so it cannot share `_maybeRefetchFullHistory`. The wheel listener is capture-phase because xterm `stopPropagation()`s the events it consumes; the alternate-screen skip (nano, vim, less) only matters for a direct-PTY shell, since under tmux the browser xterm never enters the alternate buffer; skipped too for a detached session (mirrors `_sendResize()`'s own check and app.js's `_maybeRefetchFullHistory`), since its own window already owns its PTY size and scrollback. Live frames arriving mid-replay, a `{t:'c'}` clear frame included, are held with their arrival time (`_liveQueue`) and replayed in order only if they arrived after the capture (the response's arrival stands in for the capture instant, as in `_finishBufferLoad`, so a frame inside that one round trip can be lost or doubled); the fetch has a 10 s deadline because it holds the pane's live output while it runs. ⚠️ A replay's own `\x1bc` would otherwise wipe the "Pane B disconnected" marker `onclose` wrote and paint a fresh, current-looking history while `onData` keeps silently dropping every keystroke on the dead socket (a Codeman restart drops the socket while the tmux session, and so the HTTP pull, still succeeds) — `_pullHistory()` re-stamps the marker after the live-frame flush when the socket closed in either order (before the pull started, or mid-fetch), tracked via `_wsClosed` rather than routed through `_onLiveOutput()`, since a close landing before the response is stamped before the cutoff and would be dropped with the rest of the pre-capture queue. There is no "Load full history" banner in Pane B, so a shell history past that 1 MiB window stays out of reach there. Non-shell Pane B is unchanged: it already loads `full=1`, and its history is out of scope for this pull (codex and Claude's inline renderer do grow tmux history; this just isn't how they recover it). Design: `docs/split-pane-sessions-plan.md`. +**Split-pane sessions** (`showSplitButton`, header button, default OFF, per-device like `showFileViewerButton` — not in `SettingsUpdateSchema`, `displayKeys` in settings-ui.js): shows two live sessions side-by-side in one Codeman window. Pane A is the untouched, existing singleton terminal (`this.terminal`/`this._ws` in terminal-ui.js); Pane B is a new, independent `SplitTerminalPane` (terminal-split.js) with its own xterm instance and its own `/ws/sessions/:id/terminal` WebSocket. ⚠️ **Pane B is deliberately plainer than Pane A** — no local-echo overlay, no CJK IME, no touch/mobile handlers, no keyboard accessory bar — since this is a desktop-only feature (a split view needs a wide viewport) and those features exist for mobile/touch input; `.btn-split` is hard-hidden below 1180px regardless of the setting by the `@media (max-width: 1179px)` rule in styles.css (mobile.css only carries a comment pointing at it: that file loads up to 1023px, so it cannot cover the 1024-1179px tablet range the feature also needs to stay off), and the per-device setting means turning it on at a desk can never sync it onto a phone in the first place. No persistence: closing the browser tab or reloading always returns to the normal single-pane view; there is no localStorage key for split state. ⚠️ Splitting a session against itself is disallowed (the picker excludes the active session), **as is splitting against a popped-out (detached) session** — `buildSplitPickerSessions()` excludes `detachedSessions` because a detached session's own window is already claiming its PTY size, and `MAX_WS_PER_SESSION` needs no change since Pane A/B are always two different sessions. ⚠️ Either pane's session ending (deleted locally or from another client) auto-collapses the split — Pane A's session ending promotes Pane B to the new single pane via `selectSession(id, { auto: true })` (an app-driven selection, so it must not spend the session's idle alert — see the Approvals Inbox note above), never by trying to hot-swap the lightweight `SplitTerminalPane` object into the primary singleton state. ⚠️ Pane B refits on every window/sidebar/tab-rail resize via the SAME trailing-edge `ResizeObserver` callback that resizes Pane A (`throttledResize` in terminal-ui.js) — it only ever measured Pane A's own container, so without an explicit `this._splitPane?.fit()` call there Pane B silently kept its stale PTY size through every resize that did not happen to be a divider drag. ⚠️ A dropped WebSocket leaves Pane B visibly dead (a message written into its own xterm buffer) rather than silently swallowing keystrokes with nothing on screen to explain why — there is no reconnect logic for v1, matching the "deliberately plainer than Pane A" design. Related but distinct: `detachSession()` already opens one session in a separate OS-level browser window (`isSoloWindow`) — that is prior art for "two sessions visible at once" but not for one window with a draggable in-page divider, which is what this feature adds. ⚠️ Pane B installs its own `attachCustomKeyEventHandler` gating the same app-level chords Pane A's own handler gates (command palette, Alt+1-9/[/] tab nav, Alt+B sidebar toggle, Ctrl+Z suspend, Shift/Ctrl+Enter newline, smart-copy Ctrl+C) — without it the document capture-phase handler's `preventDefault()` (which never stops xterm) let each chord ALSO write its raw byte/escape sequence into Pane B's live PTY on top of whatever the app action did to Pane A (COD-153). Ctrl+Z is swallowed unless Pane B's own session is `mode === 'shell'`, mirroring terminal-ui.js's reasoning: in a plain shell it is the user's own job-control tool, everywhere else it silently suspends an unattended agent loop. Shift/Ctrl+Enter POSTs to `/api/sessions/:id/send-key` (`{key:'S-Enter'|'C-Enter'}`, tmux `send-keys -H` for a real 0x0a) targeting THIS pane's own `sessionId` rather than the primary pane's `activeSessionId` — without it xterm's plain `\r` would submit an incomplete prompt instead of adding a line to it. Smart-copy Ctrl+C/Ctrl+Shift+C is re-implemented against `this.terminal` (Pane B's own) rather than reusing `app.copyTerminalSelection()`, which reads Pane A's terminal and would copy the wrong pane's selection; Ctrl+Shift+C never falls through even with nothing to copy, mirroring terminal-ui.js's own `ev.shiftKey` branch. ⚠️ **This is a UX-parity fix, not an interrupt-safety one** — verified live in a real browser: xterm's `evaluateKeyboardEvent` routes a shifted ctrl-letter into a branch that assigns `c.key` only for two special cases (`_`→US, `@`→NUL), so it emits no data for Ctrl+Shift+C at all regardless of any application gate; a synthetic keydown with the gate removed produces zero WS frames, proving no accidental interrupt reaches the PTY either way. What gating the whole copy block on `hasSelection()` (an earlier draft) actually cost: with no selection, a selection-less Ctrl+Shift+C fell straight to `return true`, silently ceding the keystroke to the BROWSER's own handling (e.g. Chrome's Inspect-Element binding) with no feedback and no copy attempt — Pane A always intercepts it. Ctrl+V stays on xterm's own default paste, since Pane B has no image-paste trap to route it to. ⚠️ `buildSplitPickerSessions()` also excludes any session with `pid === null` (an exited CLI, a crash-looped session whose breaker tripped, a restore that never re-attached): Pane B has no equivalent of `selectSession()`'s auto re-attach POST, so a pane opened onto one has nothing reading its tmux pane — no `terminal` events ever arrive, and `Session.write()` silently drops every keystroke with no ack either way, so the loss is invisible behind a socket that reports healthy. Pane B's input frames deliberately carry no `cid`/`seq` (`ws-routes.ts` supports that), matching the no-overlay/no-IME "deliberately plainer" list above, since it has no exactly-once delivery layer to key them against. ⚠️ A SHELL Pane B pulls scrollback itself when the wheel goes up at the top of its buffer (`_maybeLoadMoreHistory`/`_pullHistory`): tmux repaints a burst of output instead of scrolling it, so Pane B's own xterm holds about one screen of scrollback while tmux holds every line, and it loaded history exactly once at connect and never again. It is the same bounded pull as Pane A's (`?full=1&tail=TERMINAL_TAIL_SIZE`, no rewrite when the window holds no more rows than the pane already has or the pane is at its `scrollback + rows` cap, and a 60 s back-off instead of 4 s when that skipped window was truncated or the pane is full, since each ask costs the server a whole-history `capture-pane`), against Pane B's OWN terminal rather than `app.terminal`, so it cannot share `_maybeRefetchFullHistory`. The wheel listener is capture-phase because xterm `stopPropagation()`s the events it consumes; the alternate-screen skip (nano, vim, less) only matters for a direct-PTY shell, since under tmux the browser xterm never enters the alternate buffer; skipped too for a detached session (mirrors `_sendResize()`'s own check and app.js's `_maybeRefetchFullHistory`), since its own window already owns its PTY size and scrollback. Live frames arriving mid-replay, a `{t:'c'}` clear frame included, are held with their arrival time (`_liveQueue`) and replayed in order only if they arrived after the capture (the response's arrival stands in for the capture instant, as in `_finishBufferLoad`, so a frame inside that one round trip can be lost or doubled); the fetch has a 10 s deadline because it holds the pane's live output while it runs. ⚠️ The "Pane B disconnected" marker must be the LAST thing on screen. A replay's own `\x1bc` would otherwise wipe a marker written before the pull and paint a fresh, current-looking history while `onData` keeps silently dropping every keystroke on the dead socket (a Codeman restart drops the socket while the tmux session, and so the HTTP pull, still succeeds), so `_pullHistory()` re-stamps it after the live-frame flush. A close DURING a pull writes nothing: `_onSocketClosed()` skips the marker while `_liveQueue` is live, since written there it would sit above the held frames the pull flushes after a skip, a downgrade, a failed fetch or the deadline, or land mid-way through a chunked replay; the pull's `finally` then writes it once, replayed or not (`closedBefore` tells the two closes apart). Both are tracked via `_wsClosed` rather than routed through `_onLiveOutput()`, since a close landing before the response is stamped before the cutoff and would be dropped with the rest of the pre-capture queue. There is no "Load full history" banner in Pane B, so a shell history past that 1 MiB window stays out of reach there. Non-shell Pane B is unchanged: it already loads `full=1`, and its history is out of scope for this pull (codex and Claude's inline renderer do grow tmux history; this just isn't how they recover it). Design: `docs/split-pane-sessions-plan.md`. ### Gesture control: the setting diff --git a/src/web/public/terminal-split.js b/src/web/public/terminal-split.js index fe031b4a..f8d352c7 100644 --- a/src/web/public/terminal-split.js +++ b/src/web/public/terminal-split.js @@ -293,19 +293,26 @@ // normal while it quietly ate everything typed into it. v1 scope is // "say so", not reconnect — collapsing the split would lose the // user's place in Pane B's scrollback for a transient blip. - this.ws.onclose = () => { - this._wsReady = false; - this._wsClosed = true; - this._writeDisconnectedMarker(); - }; + this.ws.onclose = () => this._onSocketClosed(); this.ws.onerror = () => { // onclose fires after onerror — cleanup happens there. }; } - // Extracted so both onclose and a history-pull replay that lands on an - // already-closed socket can write it (see _pullHistory()'s finally block). + // The socket's close, split out of connect() so the tests can drive it. + // While a history pull is running the marker waits for the pull's finally + // block: written now, it would sit above the output the pull is still + // holding (flushed after it on a skip, a downgrade or a failed fetch) or + // land in the middle of a chunked replay. + _onSocketClosed() { + this._wsReady = false; + this._wsClosed = true; + if (!this._liveQueue) this._writeDisconnectedMarker(); + } + + // Extracted so both _onSocketClosed() and a history pull that ends on a + // closed socket can write it (see _pullHistory()'s finally block). _writeDisconnectedMarker() { this.terminal?.write('\r\n\x1b[2m[Pane B disconnected — close and reopen the split to reconnect]\x1b[0m\r\n'); } @@ -423,6 +430,8 @@ // main thread) and replays it under the reader's current place. Holds the // single-flight flag across the fetch AND the replay, like _loadBuffer(). async _pullHistory() { + // A close before the pull already wrote its marker; one during it did not. + const closedBefore = this._wsClosed; this._bufferLoading = true; this._liveQueue = []; let replayed = false; @@ -491,14 +500,15 @@ if (entry.clear) this.terminal?.clear(); else this.terminal?.write(entry.data); } - // A replay's own `\x1bc` wipes the disconnected marker onclose wrote, + // A replay's own `\x1bc` wipes a marker written before the pull, // painting a fresh, current-looking history while onData keeps - // silently dropping every keystroke on the dead socket. Re-stamp it - // if the socket closed in either order (before the pull started, or - // while the fetch was in flight) — checked after the queue flush so - // it is the last thing on screen, matching what onclose would have + // silently dropping every keystroke on the dead socket, so re-stamp it + // after a replay. A close DURING the pull wrote no marker at all + // (_onSocketClosed() defers it while the queue is live), so write it + // whether or not this pull replayed. Checked after the queue flush so + // it is the last thing on screen, matching what the close would have // left had the pull never run. - if (replayed && this._wsClosed) this._writeDisconnectedMarker(); + if (this._wsClosed && (replayed || !closedBefore)) this._writeDisconnectedMarker(); this._endBufferLoad(); } } diff --git a/test/split-pane-terminal-unit.test.ts b/test/split-pane-terminal-unit.test.ts index 90de98cc..9f86c384 100644 --- a/test/split-pane-terminal-unit.test.ts +++ b/test/split-pane-terminal-unit.test.ts @@ -65,6 +65,7 @@ type PaneUnderTest = { _onLiveClear(): void; _installWheelListener(): void; _writeDisconnectedMarker(): void; + _onSocketClosed(): void; }; const fetchMock = vi.fn(); @@ -133,6 +134,8 @@ function deferred<T>() { return { promise, resolve }; } +const isMarker = (data: unknown) => typeof data === 'string' && data.includes('Pane B disconnected'); + /** Lets every microtask the vm-side promise chain queued run. */ const settle = () => new Promise((r) => setTimeout(r, 0)); @@ -664,6 +667,18 @@ describe('SplitTerminalPane scroll-to-top history pull', () => { expect(connect).toContain('this._installWheelListener();'); expect(connect).toContain('this._onLiveClear();'); expect(connect).not.toContain('this.terminal.clear();'); + // The tests below drive the close through _onSocketClosed() directly. + expect(connect).toContain('this.ws.onclose = () => this._onSocketClosed();'); + }); + + it('a close with no pull running writes the marker straight away', () => { + const pane = makePane('shell'); + + pane._onSocketClosed(); + + expect(pane._wsClosed).toBe(true); + expect(pane.terminal.write).toHaveBeenCalledTimes(1); + expect(isMarker(pane.terminal.write.mock.calls[0][0])).toBe(true); }); it('re-stamps the disconnected marker after a replay if the socket closed before the pull started', async () => { @@ -683,20 +698,77 @@ describe('SplitTerminalPane scroll-to-top history pull', () => { expect(pane.terminal.write).toHaveBeenCalledWith(marker); }); - it('re-stamps the disconnected marker after a replay if the socket closes mid-fetch', async () => { - // The other order Ark0N's review called out: the close lands while the - // capture is in flight, so the HTTP pull still succeeds (a Codeman - // restart drops the WS while the tmux session, and so the pull, survives). + it('writes the disconnected marker once, after the replay, if the socket closes mid-fetch', async () => { + // The close lands while the capture is in flight, so the HTTP pull still + // succeeds (a Codeman restart drops the WS while the tmux session, and so + // the pull, survives) and the replay that follows is what the marker must + // end up below. const pane = makePane('shell'); const response = deferred<ReturnType<typeof jsonResponse>>(); fetchMock.mockReturnValueOnce(response.promise); const pull = pane._pullHistory(); - pane._wsClosed = true; // the close arrives mid-fetch, before the response + pane._onSocketClosed(); // the close arrives mid-fetch, before the response + expect(pane.terminal.write).not.toHaveBeenCalled(); response.resolve(jsonResponse(rowsOf(100))); await pull; - expect(pane.terminal.write.mock.calls.at(-1)?.[0]).toEqual(expect.stringMatching(/Pane B disconnected/)); + const writes = pane.terminal.write.mock.calls.map((c) => c[0]); + expect(writes.filter(isMarker)).toHaveLength(1); + expect(isMarker(writes.at(-1))).toBe(true); + }); + + it.each([ + ['a skip', 40, () => fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(30)))], + ['a downgrade', 500, () => fetchMock.mockResolvedValueOnce(jsonResponse(rowsOf(5)))], + ['a failed fetch', 40, () => fetchMock.mockRejectedValueOnce(new Error('offline'))], + ])( + 'a close mid-fetch that ends in %s writes the marker last, after the held frames', + async (_label, rowsHeld, mockFetch) => { + // No replay ever runs here, so nothing would wipe a marker written at the + // close; written straight away it sat ABOVE the output the pull was still + // holding, which the finally block then flushed underneath it. + const pane = makePane('shell'); + pane.terminal.buffer.active.length = rowsHeld; + mockFetch(); + + const pull = pane._pullHistory(); + pane._onLiveOutput('frame-A'); + pane._onLiveOutput('frame-B'); + pane._onSocketClosed(); + expect(pane.terminal.write).not.toHaveBeenCalled(); + await pull; + + const writes = pane.terminal.write.mock.calls.map((c) => c[0]); + expect(writes.slice(0, 2)).toEqual(['frame-A', 'frame-B']); + expect(writes).toHaveLength(3); + expect(isMarker(writes[2])).toBe(true); + expect(pane._liveQueue).toBeNull(); + } + ); + + it('a close during the chunked replay writes exactly one marker, at the end', async () => { + const pane = makePane('shell'); + const response = deferred<ReturnType<typeof jsonResponse>>(); + fetchMock.mockReturnValueOnce(response.promise); + + const pull = pane._pullHistory(); + // Three chunks, so the replay is still mid-write once the fetch lands. + const bigReplay = Array.from({ length: 200 }, () => 'y'.repeat(400)).join('\n'); + response.resolve(jsonResponse(bigReplay)); + await settle(); + expect(rafQueue).toHaveLength(1); + + // Written now, the marker would land between two chunks of recovered history. + pane._onSocketClosed(); + rafQueue.shift()!(); + rafQueue.shift()!(); + await pull; + + const writes = pane.terminal.write.mock.calls.map((c) => c[0]); + expect(writes[0]).toBe('\x1bc'); + expect(writes.filter(isMarker)).toHaveLength(1); + expect(isMarker(writes.at(-1))).toBe(true); }); it('does not re-stamp the marker when the socket is still open', async () => { From 73c0bfccc42625ef30072c9910d8a011fe3fc09f Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Thu, 1 Oct 2026 11:19:00 +0200 Subject: [PATCH 43/46] fix(files): keep attachment markdown refs from resolving into the workspace, and render files without chat line breaks (#503 review) - A markdown preview opened by attachment id under a bare file name (attachment cards, history drawer) no longer resolves relative refs against the workspace root: filePreviewText carries attachmentId, and the rebase pass turns those images into their alt text and unwraps those links. Absolute-path and workspace previews are unchanged. - _renderMarkdown(text, { breaks = true } = {}): the File Viewer passes breaks: false, so a hard-wrapped paragraph renders as one paragraph; the Response Viewer keeps a <br> per newline. - Absolute paths linkified inside a rendered document now carry the preview's data-session-id. - CLAUDE.md, architecture-invariants and the Working-With-Files wiki page now say that only an in-workspace path clicked in the terminal keeps the tail viewer. - Tests in test/file-preview-markdown.test.ts for all three fixes, including an end-to-end run of the shipping app.js + marked + DOMPurify. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- CLAUDE.md | 2 +- docs/architecture-invariants.md | 6 +- docs/wiki/Working-With-Files.md | 7 +- src/web/public/app.js | 11 +- src/web/public/panels-ui.js | 31 ++++- test/file-preview-markdown.test.ts | 177 ++++++++++++++++++++++++++--- 6 files changed, 205 insertions(+), 29 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 450bfcbe..380559ae 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -295,7 +295,7 @@ Codeman is a Claude Code session manager with web interface and autonomous Ralph **File Viewer edit mode** (issue #212): the file-preview overlay edits workspace text files in place — `GET .../file-content?edit=1` + `PUT /api/sessions/:id/file-content`, policy in `src/config/file-editing.ts`. This is a **third file surface and the only one that WRITES**: read-path confinement (realpath + workspace + ownership) plus sensitive/blocked/`.git` denies and an extension **allowlist**; writes are `wx`-temp + rename (no `O_CREAT` anywhere = edit-in-place is structural); optimistic concurrency via sha256 `baseHash` → 409. ⚠️ `edit=1` never truncates and the client must never save a plain-preview buffer (the 500-line truncation would silently delete the rest). ⚠️ CRLF/UTF-8 guards: EOL re-applied server-side, non-UTF-8 refused via round-trip compare. → [architecture-invariants#file-viewer-edit-mode](docs/architecture-invariants.md#file-viewer-edit-mode), `docs/file-viewer-edit-plan.md` -**File Viewer text view: rendered markdown + Lines/Wrap toggles** (`_renderFilePreviewText()` in panels-ui.js): a `.md`/`.markdown` opens RENDERED by default with an `MD` pill back to source; the plain-text view has `Lines` (CSS-counter gutter) and `Wrap` toggles. ⚠️ ONE markdown pipeline: the viewer calls `_renderMarkdown()` (marked + the DOMPurify allowlist, the Response Viewer's) and binds the Response Viewer's click delegate (`_bindResponseViewerInteractions`) on the preview body for code-copy buttons and path links; never a second parser or handler. ⚠️ The document is built inside a `<template>` (a detached div with `innerHTML` set starts fetching every `<img src>` before the rewrite), then `_rebaseFilePreviewMarkdownRefs()` points relative images at the workspace-confined `file-raw` under the document's directory and root-relative ones under the workspace root (never a widened route; a failed load degrades to alt text), after `decodeURIComponent`ing the ref and dropping `?query`/`#fragment` (marked percent-encodes destinations, and the route encodes again), and turns workspace links into `a.rv-path` carrying `data-session-id` for the delegate, stripping the `target` marked gave them. ⚠️ The container carries `data-i18n-skip` or the translator rewrites the document's prose. ⚠️ Toggles are per-device localStorage keys (`codeman:filePreview*`), never `SettingsUpdateSchema`; Lines/Wrap are class flips on the ONE `<pre>`, with rules scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's code blocks. Markdown fetches `lines=10000` (the route ceiling), other text keeps 500. ⚠️ `md` stays OUT of `FILE_PREVIEW_EXTENSIONS`: a printed `.md` path keeps the tail viewer (live follow); the rendered view is the Files panel's. Tests: `test/file-preview-markdown.test.ts`. → [architecture-invariants#file-viewer-text-view-rendered-markdown-and-text-toggles](docs/architecture-invariants.md#file-viewer-text-view-rendered-markdown-and-text-toggles) +**File Viewer text view: rendered markdown + Lines/Wrap toggles** (`_renderFilePreviewText()` in panels-ui.js): a `.md`/`.markdown` opens RENDERED by default with an `MD` pill back to source; the plain-text view has `Lines` (CSS-counter gutter) and `Wrap` toggles. ⚠️ ONE markdown pipeline: the viewer calls `_renderMarkdown(text, { breaks: false })` (marked + the DOMPurify allowlist, the Response Viewer's; chat keeps the default `breaks: true`, a file must not turn every hard wrap into a `<br>`) and binds the Response Viewer's click delegate (`_bindResponseViewerInteractions`) on the preview body for code-copy buttons and path links; never a second parser or handler. ⚠️ The document is built inside a `<template>` (a detached div with `innerHTML` set starts fetching every `<img src>` before the rewrite), then `_rebaseFilePreviewMarkdownRefs()` points relative images at the workspace-confined `file-raw` under the document's directory and root-relative ones under the workspace root (never a widened route; a failed load degrades to alt text), after `decodeURIComponent`ing the ref and dropping `?query`/`#fragment` (marked percent-encodes destinations, and the route encodes again), and turns workspace links into `a.rv-path` carrying `data-session-id` for the delegate (the linkifier's absolute paths get it too), stripping the `target` marked gave them. ⚠️ A preview opened by attachment id under a bare file name (attachment cards, history drawer: the registry keeps no relative path) has no directory, so its workspace refs degrade (images to alt text, links to their text), never resolve against the workspace root. ⚠️ The container carries `data-i18n-skip` or the translator rewrites the document's prose. ⚠️ Toggles are per-device localStorage keys (`codeman:filePreview*`), never `SettingsUpdateSchema`; Lines/Wrap are class flips on the ONE `<pre>`, with rules scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's code blocks. Markdown fetches `lines=10000` (the route ceiling), other text keeps 500. ⚠️ `md` stays OUT of `FILE_PREVIEW_EXTENSIONS`: only an IN-WORKSPACE `.md` path clicked in the TERMINAL keeps the tail viewer (live follow); the Files panel, chat paths, out-of-workspace terminal paths and attachment cards all reach `openFilePreview()` and render it. Tests: `test/file-preview-markdown.test.ts`. → [architecture-invariants#file-viewer-text-view-rendered-markdown-and-text-toggles](docs/architecture-invariants.md#file-viewer-text-view-rendered-markdown-and-text-toggles) **Files panel search** (COD-236, the `q` param on `GET /api/sessions/:id/files`): `compileFileQuery()` (`utils/file-query.ts`, pure) compiles the query into a predicate the server-side walk prunes with; a query returns a FLAT match list and the walk recurses past non-matching directories. An empty, whitespace-only or overlong (`MAX_QUERY_LENGTH`, 256) query compiles to `null`, keeping the default tree response byte-identical. ⚠️ **Never compile a glob into a RegExp** (`*a*a*a…` backtracks and freezes the event loop for the whole server): `globMatch()` is a two-pointer wildcard walk. → [architecture-invariants#files-panel-search](docs/architecture-invariants.md#files-panel-search) diff --git a/docs/architecture-invariants.md b/docs/architecture-invariants.md index 32f6c172..cbfdf3f2 100644 --- a/docs/architecture-invariants.md +++ b/docs/architecture-invariants.md @@ -392,12 +392,12 @@ Tests: `test/file-editing-policy.test.ts` (pure policy), `test/routes/file-write **Rendered markdown + Lines/Wrap** (`_renderFilePreviewText()` and its helpers in `panels-ui.js`, buttons in the `.file-preview-actions` row): a `.md`/`.markdown` opened in the File Viewer renders as a document by default, with an `MD` pill back to source; the plain-text view has `Lines` (a CSS-counter gutter) and `Wrap` toggles. Codeman already had `marked` + DOMPurify behind `_renderMarkdown()` for the Response Viewer, so the viewer reuses that and the codebase keeps ONE markdown pipeline. -- ⚠️ **One pipeline, one delegate.** The viewer calls `_renderMarkdown()` (marked + the `sanitize-html.js` allowlist) and binds `_bindResponseViewerInteractions()` on `#filePreviewBody` (container-bound and idempotent, so once per page) for the code-block copy buttons and `a.rv-path` opening. Never a second parser, never a second click handler for the same markup. -- ⚠️ **Build inside a `<template>`, then rebase.** A detached div whose `innerHTML` is set starts fetching every `<img src>` at once, so the document's relative image paths would hit the server as `/docs/img.png` 404s before being rewritten. `_rebaseFilePreviewMarkdownRefs()` runs on the template content: relative images go to the workspace-confined `file-raw` under the document's directory, root-relative ones (`/docs/x.png`) under the workspace root as on GitHub (the server refuses escapes, so `..` is forwarded as-is), and one `error` handler per image degrades it to alt text, which covers a remote image the page CSP blocks, a 404 for a document outside the workspace, and an SVG that `file-raw` serves as a download. Never widen a route for this. ⚠️ marked percent-encodes destinations (`my image.png` arrives as `my%20image.png`, CJK names as `%E5…`), so the ref is `decodeURIComponent`ed (a malformed escape is kept as written) and stripped of `#fragment` and `?query` BEFORE the route encodes it again; without that `file-raw` looks for a file literally named `my%20image.png`. Workspace links become `a.rv-path` with `data-path` AND `data-session-id` (the preview's session, which the `app.js` delegate prefers over `activeSessionId`, since a preview opened from another session's attachment card must resolve links against that workspace) and lose the `target`/`rel` that `_renderMarkdown` gives every link, which would otherwise open `<origin>/docs/x.md` in a new tab; fragment, protocol-relative and http(s) links are untouched. +- ⚠️ **One pipeline, one delegate.** The viewer calls `_renderMarkdown(text, { breaks: false })` (marked + the `sanitize-html.js` allowlist) and binds `_bindResponseViewerInteractions()` on `#filePreviewBody` (container-bound and idempotent, so once per page) for the code-block copy buttons and `a.rv-path` opening. Never a second parser, never a second click handler for the same markup. `breaks` is the one option that differs: chat keeps the default `true` (a newline is the agent's line break), while a file passes `false`, because a README hard-wrapped at 80 columns would otherwise render every wrap as a `<br>`, unlike GitHub's file view. +- ⚠️ **Build inside a `<template>`, then rebase.** A detached div whose `innerHTML` is set starts fetching every `<img src>` at once, so the document's relative image paths would hit the server as `/docs/img.png` 404s before being rewritten. `_rebaseFilePreviewMarkdownRefs()` runs on the template content: relative images go to the workspace-confined `file-raw` under the document's directory, root-relative ones (`/docs/x.png`) under the workspace root as on GitHub (the server refuses escapes, so `..` is forwarded as-is), and one `error` handler per image degrades it to alt text, which covers a remote image the page CSP blocks, a 404 for a document outside the workspace, and an SVG that `file-raw` serves as a download. Never widen a route for this. ⚠️ marked percent-encodes destinations (`my image.png` arrives as `my%20image.png`, CJK names as `%E5…`), so the ref is `decodeURIComponent`ed (a malformed escape is kept as written) and stripped of `#fragment` and `?query` BEFORE the route encodes it again; without that `file-raw` looks for a file literally named `my%20image.png`. Workspace links become `a.rv-path` with `data-path` AND `data-session-id` (the preview's session, which the `app.js` delegate prefers over `activeSessionId`, since a preview opened from another session's attachment card must resolve links against that workspace) and lose the `target`/`rel` that `_renderMarkdown` gives every link, which would otherwise open `<origin>/docs/x.md` in a new tab; fragment, protocol-relative and http(s) links are untouched. The absolute paths `_linkifyFilePaths()` then finds in the prose get the same `data-session-id`, set on every `a.rv-path` still lacking one. ⚠️ **An attachment opened under a bare file name has no directory.** Attachment cards and the history drawer call `openFilePreview(relativePath || fileName, sessionId, attachmentId)`, and a registered attachment's `relativePath` is always `''` (`attachmentRecordToEvent`), so `filePath` is just `report.md`: resolving against it sent `docs/report.md`'s `img/chart.png` to the workspace root's `img/chart.png` (a missing image) and its `CONTRIBUTING.md` link to the root's (a silently different file), and an out-of-workspace attachment's refs all landed in the workspace. `filePreviewText` therefore carries `attachmentId`, and when it is set with a non-absolute `filePath` the rebase degrades every workspace ref instead: images become their alt text (as a text node), links are unwrapped to their text. An absolute-path attachment (one `_registerExternalPreview` minted for a clicked path) still resolves against its own directory, as does every workspace preview. - ⚠️ **`data-i18n-skip` on the container.** The translator's MutationObserver translates inserted headings and paragraphs, and the `.file-preview-content` entry in its skip list matches nothing (no element has that class), so the attribute is what keeps a Chinese UI from rewriting a README. - ⚠️ **Toggles are per-device, in their own localStorage keys** (`codeman:filePreviewMdRendered` / `LineNumbers` / `Wrap`), for the same reason as the Files panel's show-hidden toggle: the app-settings object is rebuilt from the settings modal on every save, and they are display state, not synced settings (`SettingsUpdateSchema` is `.strict()`). MD re-renders from the kept source (`filePreviewContent`, which is also what Copy copies) without a refetch; Lines/Wrap are class flips on the one `<pre>`, whose rules are scoped `.file-preview-body > pre.file-preview-text` so they never leak into the document's own code blocks. Lines are one inline `<span class="fp-line">` per line joined by real newlines, the counter in `::before` with `user-select: none`, so select and copy return the exact text. - ⚠️ **Caps.** Markdown fetches `lines=10000` (the route's `MAX_LINES_LIMIT`), because a rendered document cut at 500 lines reads as the whole document; other text keeps 500, which is what stops a huge log locking the tab in one `<pre>`. The attachment (out-of-workspace) branch keeps its 512 KB Range read and skips the 500-line clip for markdown. Edit mode is unchanged and still re-fetches `edit=1`; the three toggles hide while editing and for images, media and PDFs. -- ⚠️ **`md` is NOT in `FILE_PREVIEW_EXTENSIONS`.** A `.md` path printed in the terminal or chat still opens the tail viewer (see File-path links above: in-workspace text keeps live follow, which is what Ralph's `fix_plan.md` needs); the rendered view is reached from the Files panel. `avif` and `ico` were added there (a printed `favicon.ico` used to tail binary noise), with `avif` also in file-content's image set and file-raw's MIME map; out-of-workspace avif/ico stay unregistrable, like svg/bmp. +- ⚠️ **`md` is NOT in `FILE_PREVIEW_EXTENSIONS`.** Only an in-workspace `.md` path clicked in the terminal still opens the tail viewer (see File-path links above: in-workspace text keeps live follow, which is what Ralph's `fix_plan.md` needs). Everything else reaches `openFilePreview()` and renders it: the Files panel, a path clicked in the Response Viewer (its delegate calls `openFilePreview()` directly), an out-of-workspace terminal path, and attachment cards. `avif` and `ico` were added to the set (a printed `favicon.ico` used to tail binary noise), with `avif` also in file-content's image set and file-raw's MIME map; out-of-workspace avif/ico stay unregistrable, like svg/bmp. Tests: `test/file-preview-markdown.test.ts` (jsdom-in-vm, pins every rule above), `test/routes/file-routes.test.ts` (avif). diff --git a/docs/wiki/Working-With-Files.md b/docs/wiki/Working-With-Files.md index f28458df..597ac026 100644 --- a/docs/wiki/Working-With-Files.md +++ b/docs/wiki/Working-With-Files.md @@ -14,7 +14,7 @@ It renders what it can: | Kind | Behaviour | | ------------------------ | ------------------------------------------------------------------------- | | Text and code | Plain preview with Lines (line numbers) and Wrap toggles in the header. Long files are truncated in plain preview. | -| Markdown | Rendered by default: headings, tables, code blocks with copy buttons, images and links relative to the file (root-relative ones resolve from the workspace root, as on GitHub). The MD pill in the header flips to source. | +| Markdown | Rendered by default: headings, tables, code blocks with copy buttons, images and links relative to the file (root-relative ones resolve from the workspace root, as on GitHub). Opened from an attachment card, where the file's folder is unknown, relative images show their alt text and relative links show as plain text. The MD pill in the header flips to source. | | Images | Inline. | | Audio and video | Inline with a working scrub bar, because range requests are supported. | | PDF and Office documents | Converted for preview when a converter is available. | @@ -86,8 +86,9 @@ File paths in a session are links. That works in two places: render as underlined monospace links. Clicking one opens it in the preview: images and PDFs render, video and audio play with a -working scrub bar, documents convert, text and Markdown show inline. Log-shaped files open in -the tail viewer instead, which follows a file that is still being written. +working scrub bar, documents convert, text shows inline and Markdown renders. The exception is +a text or Markdown file inside the workspace clicked in the terminal: that opens in the tail +viewer instead, which follows a file that is still being written. Paths **outside** the session's workspace work too, which matters because that is where most of an agent's output lands: a screenshot in `/tmp`, a capture in its own scratchpad, a file in diff --git a/src/web/public/app.js b/src/web/public/app.js index 86f8f670..fcfa641d 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -2296,13 +2296,18 @@ class CodemanApp { return processed.replace(/__CODEMAN_FENCE_(\d+)__/g, (_m, i) => placeholders[Number(i)]); } - /** Render markdown to sanitized HTML, falling back to plain text if marked.js unavailable */ - _renderMarkdown(text) { + /** + * Render markdown to sanitized HTML, falling back to plain text if marked.js unavailable. + * `breaks` turns every source newline into a <br>: right for chat, where a + * newline is the agent's line break, wrong for a file (the File Viewer passes + * false), where a README hard-wrapped at 80 columns would break at every wrap. + */ + _renderMarkdown(text, { breaks = true } = {}) { const src = text || ''; if (typeof marked !== 'undefined' && marked.parse) { try { const prepared = this._preprocessAsciiArt(src); - let html = this._sanitizeHtml(marked.parse(prepared, { breaks: true, gfm: true })); + let html = this._sanitizeHtml(marked.parse(prepared, { breaks, gfm: true })); // Wrap tables in a horizontal-scroll container so they overflow gracefully // on mobile without collapsing into block-level cells. html = html.replace(/<table>/g, '<div class="rv-table-wrap"><table>') diff --git a/src/web/public/panels-ui.js b/src/web/public/panels-ui.js index fe5ada63..49a582ee 100644 --- a/src/web/public/panels-ui.js +++ b/src/web/public/panels-ui.js @@ -4197,7 +4197,9 @@ Object.assign(CodemanApp.prototype, { const clippedByLines = lines.length > lineCap; const shown = clippedByLines ? lines.slice(0, lineCap).join('\n') : text; this.filePreviewContent = shown; - this.filePreviewText = { ext, sessionId, filePath }; + // attachmentId: a card's filePath is the bare file name, so the + // rebase pass must know there is no directory to resolve against. + this.filePreviewText = { ext, sessionId, filePath, attachmentId }; this._renderFilePreviewText(); if (clippedByLines || clippedByBytes) { const note = clippedByLines ? `showing first ${lineCap} lines` : 'showing the start of the file'; @@ -4393,7 +4395,9 @@ Object.assign(CodemanApp.prototype, { * parser, and is built inside a <template>: a detached div with innerHTML * already set starts fetching every <img src>, so the document's relative * image paths would hit the server as /docs/img.png 404s before - * `_rebaseFilePreviewMarkdownRefs` rewrote them. + * `_rebaseFilePreviewMarkdownRefs` rewrote them. `breaks: false` because a + * file is not a chat message: a paragraph hard-wrapped in the source is one + * paragraph, as on GitHub. */ _renderFilePreviewText() { const info = this.filePreviewText; @@ -4405,10 +4409,13 @@ Object.assign(CodemanApp.prototype, { // data-i18n-skip: the translator's MutationObserver would otherwise // rewrite the document's own headings and paragraphs. const tmpl = document.createElement('template'); - tmpl.innerHTML = `<div class="rv-text file-preview-md" data-i18n-skip>${this._renderMarkdown(this.filePreviewContent)}</div>`; + tmpl.innerHTML = `<div class="rv-text file-preview-md" data-i18n-skip>${this._renderMarkdown(this.filePreviewContent, { breaks: false })}</div>`; const doc = tmpl.content.firstElementChild; this._rebaseFilePreviewMarkdownRefs(doc, info); this._linkifyFilePaths(doc); + // The linkifier's absolute paths name no session; give them the + // preview's, like the rebased links, or they open in the active tab's. + for (const a of doc.querySelectorAll('a.rv-path:not([data-session-id])')) a.dataset.sessionId = info.sessionId; bodyEl.replaceChildren(tmpl.content); // The Response Viewer's click delegate (path links, code-block copy // buttons, loopback links): container-bound and idempotent, so binding it @@ -4447,9 +4454,17 @@ Object.assign(CodemanApp.prototype, { * would otherwise open <origin>/docs/x.md in a new tab, and carry the * preview's own session so a document opened from another session's * attachment card resolves against that workspace, not the active tab's. + * + * A preview opened by attachment id under a bare file name (attachment + * cards and the history drawer: the registry keeps no relative path) has no + * directory to resolve against, and the workspace root is the wrong one for + * docs/report.md and for a file outside the workspace alike. Its workspace + * refs degrade instead: images to their alt text, links to their text, + * rather than a missing image or a silently different file. */ - _rebaseFilePreviewMarkdownRefs(root, { sessionId, filePath }) { + _rebaseFilePreviewMarkdownRefs(root, { sessionId, filePath, attachmentId }) { const dir = filePath.includes('/') ? filePath.slice(0, filePath.lastIndexOf('/') + 1) : ''; + const unresolvable = !!attachmentId && !filePath.startsWith('/'); // Workspace ref = no scheme, not protocol-relative (//host), not a fragment. const isWorkspaceRef = (ref) => !!ref && !/^[a-z][a-z0-9+.-]*:/i.test(ref) && !ref.startsWith('//') && !ref.startsWith('#'); @@ -4478,6 +4493,10 @@ Object.assign(CodemanApp.prototype, { }; for (const img of root.querySelectorAll('img[src]')) { const src = img.getAttribute('src') || ''; + if (unresolvable && isWorkspaceRef(src)) { + img.replaceWith(img.getAttribute('alt') || src); + continue; + } if (isWorkspaceRef(src)) { const path = resolveRef(src); img.setAttribute('src', CodemanBase.url(`/api/sessions/${sessionId}/file-raw?path=${encodeURIComponent(path)}`)); @@ -4487,6 +4506,10 @@ Object.assign(CodemanApp.prototype, { for (const a of root.querySelectorAll('a[href]')) { const href = a.getAttribute('href') || ''; if (!isWorkspaceRef(href)) continue; + if (unresolvable) { + a.replaceWith(...a.childNodes); + continue; + } a.className = 'rv-path'; a.dataset.path = resolveRef(href); a.dataset.sessionId = sessionId; diff --git a/test/file-preview-markdown.test.ts b/test/file-preview-markdown.test.ts index 7c279204..7aa51c1d 100644 --- a/test/file-preview-markdown.test.ts +++ b/test/file-preview-markdown.test.ts @@ -21,21 +21,31 @@ * and while editing. * 5. `FILE_PREVIEW_EXTENSIONS` gained avif/ico and still has no `md` * (in-workspace text keeps the tail viewer, see architecture-invariants). + * 6. A preview opened by attachment id under a bare file name has no + * directory to resolve against, so its relative images degrade to alt text + * and its relative links to plain text instead of landing on the workspace + * root's files; an absolute-path attachment keeps resolving. + * 7. A file renders without chat line breaks (`breaks: false`): a paragraph + * hard-wrapped in the source is one paragraph, while the Response Viewer + * keeps a <br> per newline. * * Loaded via `vm` with a jsdom document injected (the technique from * response-viewer-file-links.test.ts): constants.js + panels-ui.js only, with - * the app.js markdown pipeline stubbed to a fixed fragment. + * the app.js markdown pipeline stubbed to a fixed fragment, except for rule 7, + * which runs the shipping app.js + vendored marked + DOMPurify end to end. */ import { readFileSync } from 'node:fs'; +import { performance } from 'node:perf_hooks'; import { resolve } from 'node:path'; import vm from 'node:vm'; import { JSDOM } from 'jsdom'; import { describe, expect, it, vi } from 'vitest'; const PUBLIC = resolve(import.meta.dirname, '../src/web/public'); -const constantsJs = readFileSync(resolve(PUBLIC, 'constants.js'), 'utf8'); -const panelsJs = readFileSync(resolve(PUBLIC, 'panels-ui.js'), 'utf8'); +const publicFile = (name: string) => readFileSync(resolve(PUBLIC, name), 'utf8'); +const constantsJs = publicFile('constants.js'); +const panelsJs = publicFile('panels-ui.js'); // A real origin: vitest's equality walker reaches the window through a node's // ownerDocument, and jsdom's localStorage getter throws on an opaque one. @@ -69,6 +79,8 @@ function jsonResponse(body: unknown) { /** Answer file-content like the route does: text as JSON, an image as metadata. */ function fetchStub(url: string) { + // An attachment's by-id raw route answers the bytes themselves. + if (url.includes('/attachments/')) return { ok: true, status: 200, text: async () => MD_CONTENT }; const path = decodeURIComponent(new URL(url, 'http://x').searchParams.get('path') || ''); const ext = path.split('.').pop() || ''; if (ext === 'png') { @@ -90,6 +102,18 @@ function fetchStub(url: string) { }); } +/** The file-preview overlay's elements, as index.html ships them. */ +function mountPreviewDom() { + document.body.innerHTML = ` + <div id="filePreviewOverlay"></div><span id="filePreviewTitle"></span> + <button id="filePreviewMdBtn" hidden></button> + <button id="filePreviewLinesBtn" hidden></button> + <button id="filePreviewWrapBtn" hidden></button> + <button id="filePreviewEditBtn" hidden></button> + <button id="filePreviewDetachBtn" hidden></button> + <div id="filePreviewBody"></div><div id="filePreviewFooter"></div>`; +} + function loadApp(prefs: Record<string, string> = {}) { const store = new Map(Object.entries(prefs)); const CodemanApp = function CodemanApp(this: unknown) {} as unknown as new () => Record<string, any>; @@ -114,14 +138,7 @@ function loadApp(prefs: Record<string, string> = {}) { filename: 'panels-ui.js', }); - document.body.innerHTML = ` - <div id="filePreviewOverlay"></div><span id="filePreviewTitle"></span> - <button id="filePreviewMdBtn" hidden></button> - <button id="filePreviewLinesBtn" hidden></button> - <button id="filePreviewWrapBtn" hidden></button> - <button id="filePreviewEditBtn" hidden></button> - <button id="filePreviewDetachBtn" hidden></button> - <div id="filePreviewBody"></div><div id="filePreviewFooter"></div>`; + mountPreviewDom(); const app = new CodemanApp(); app.$ = (id: string) => document.getElementById(id); @@ -154,7 +171,8 @@ describe('file viewer rendered markdown', () => { const doc = body.firstElementChild as HTMLElement; expect(doc.matches('.rv-text.file-preview-md[data-i18n-skip]')).toBe(true); expect(doc.querySelector('h1')?.textContent).toBe('Title'); - expect(app._renderMarkdown).toHaveBeenCalledWith(MD_CONTENT); + // A file, not a chat message: source newlines inside a paragraph are not breaks. + expect(app._renderMarkdown).toHaveBeenCalledWith(MD_CONTENT, { breaks: false }); // Identity, not deep equality: DOM nodes are compared by reference here. expect(app._linkifyFilePaths.mock.calls[0][0]).toBe(doc); expect(app._bindResponseViewerInteractions.mock.calls[0][0]).toBe(body); @@ -210,6 +228,14 @@ describe('file viewer rendered markdown', () => { it('decodes percent-encoded refs, drops the query, and resolves root-relative refs against the workspace', async () => { const { app, body } = loadApp(); + // An absolute path in the document's prose, linked by the Response + // Viewer's linkifier, which knows nothing of the preview's session. + app._linkifyFilePaths.mockImplementation((root: HTMLElement) => { + const a = root.ownerDocument.createElement('a'); + a.className = 'rv-path'; + a.dataset.path = '/tmp/out/run.log'; + root.appendChild(a); + }); await app.openFilePreview('docs/README.md', 's1'); const src = (alt: string) => body.querySelector(`img[alt="${alt}"]`)!.getAttribute('src'); @@ -227,13 +253,50 @@ describe('file viewer rendered markdown', () => { const anchors = Array.from(body.querySelectorAll('a')); expect(anchors.find((a) => a.textContent === 'cjk')!.getAttribute('data-path')).toBe('docs/图片/截图.md'); expect(anchors.find((a) => a.textContent === 'rootlink')!.getAttribute('data-path')).toBe('docs/root.md'); - // Every rebased link names the preview's session, so the delegate opens it - // in that workspace even when another tab is active. + // Every rebased link, and every path the linkifier found in the prose, + // names the preview's session, so the delegate opens it in that workspace + // even when another tab is active. const rebased = body.querySelectorAll('a.rv-path'); - expect(rebased.length).toBe(4); + expect(rebased.length).toBe(5); for (const a of rebased) expect(a.getAttribute('data-session-id')).toBe('s1'); }); + it('degrades relative refs of an attachment opened by bare file name instead of resolving them in the workspace', async () => { + const { app, body, fetchMock } = loadApp(); + + // An attachment card passes the registry's bare file name: the document's + // directory is unknown, so `img/a.png` must not become the workspace root's. + await app.openFilePreview('report.md', 's1', 'att-1'); + + expect(fetchMock.mock.calls[0][0]).toContain('/attachments/att-1/raw'); + expect(body.innerHTML).not.toContain('file-raw'); + // Relative and root-relative images are their alt text, as a text node. + for (const alt of ['Alt A', 'space', 'raw', 'bad', 'root']) { + expect(body.querySelector(`img[alt="${alt}"]`)).toBeNull(); + expect(body.textContent).toContain(alt); + } + // Remote images and links keep today's handling. + expect(body.querySelector('img[alt="remote"]')!.getAttribute('src')).toBe('https://cdn.example.com/r.png'); + expect(body.querySelector('img[alt="protorel"]')!.getAttribute('src')).toBe('//cdn.example.com/p.png'); + // Relative links are unwrapped to their text; fragment and http(s) links stay. + expect(body.querySelectorAll('a.rv-path')).toHaveLength(0); + const anchors = Array.from(body.querySelectorAll('a')).map((a) => a.textContent); + expect(anchors).toEqual(['t', 'e']); + for (const text of ['x', 'up', 'cjk', 'rootlink']) expect(body.textContent).toContain(text); + }); + + it('keeps resolving refs of an absolute-path attachment against its own directory', async () => { + const { app, body } = loadApp(); + + await app.openFilePreview('/tmp/out/report.md', 's1', 'att-2'); + + const raw = (path: string) => `/api/sessions/s1/file-raw?path=${encodeURIComponent(path)}`; + expect(body.querySelector('img[alt="Alt A"]')!.getAttribute('src')).toBe(raw('/tmp/out/img/a.png')); + const rel = Array.from(body.querySelectorAll('a.rv-path')).find((a) => a.textContent === 'x')!; + expect(rel.getAttribute('data-path')).toBe('/tmp/out/guide/x.md'); + expect(rel.getAttribute('data-session-id')).toBe('s1'); + }); + it('degrades an image that fails to load to its alt text', async () => { const { app, body } = loadApp(); @@ -316,6 +379,90 @@ describe('file viewer Lines and Wrap toggles', () => { }); }); +/** A vendored UMD build (or sanitize-html.js), evaluated as CommonJS the way the other suites do. */ +function loadCommonJs<T>(name: string): T { + const module: { exports: unknown } = { exports: {} }; + // eslint-disable-next-line @typescript-eslint/no-implied-eval, no-new-func + new Function('module', 'exports', publicFile(name))(module, module.exports); + return module.exports as T; +} + +/** + * The SHIPPING pipeline end to end: app.js (`_renderMarkdown` and the Response + * Viewer's message builder) with panels-ui.js mixed in, the vendored marked, + * and DOMPurify behind the real sanitize-html.js config. `content` is what the + * file-content route answers for every path. + */ +function loadShippingApp(content: string) { + const createDOMPurify = loadCommonJs<(win: unknown) => unknown>('vendor/dompurify.min.js'); + const { createMarkdownSanitizer } = loadCommonJs<{ createMarkdownSanitizer: (dp: unknown) => unknown }>( + 'sanitize-html.js' + ); + const context = vm.createContext({ + console: { ...console, warn: vi.fn(), error: vi.fn() }, + performance, + setInterval: vi.fn(), + clearInterval: vi.fn(), + setTimeout, + clearTimeout, + requestAnimationFrame: vi.fn(), + HTMLCanvasElement: class HTMLCanvasElement {}, + document, + NodeFilter: dom.window.NodeFilter, + localStorage: { length: 0, key: vi.fn(), getItem: () => null, setItem: vi.fn(), removeItem: vi.fn() }, + // _sanitizeHtml fails closed without the page's sanitizer, which would make + // every assertion below vacuous. + window: { + addEventListener: vi.fn(), + removeEventListener: vi.fn(), + sanitizeMarkdownHtml: createMarkdownSanitizer(createDOMPurify(dom.window)), + }, + marked: loadCommonJs('vendor/marked.min.js'), + MobileDetection: {}, + confirm: () => true, + fetch: vi.fn(async () => + jsonResponse({ + success: true, + data: { content, totalLines: 2, size: content.length, truncated: false, extension: 'md' }, + }) + ), + }); + vm.runInContext( + `${constantsJs}\n${publicFile('app.js')}\n${panelsJs}\nglobalThis.__CodemanApp = CodemanApp;`, + context, + { filename: 'app.js' } + ); + const CodemanApp = (context as { __CodemanApp: { prototype: object } }).__CodemanApp; + + mountPreviewDom(); + const app = Object.create(CodemanApp.prototype) as Record<string, any>; + app.$ = (id: string) => document.getElementById(id); + app.sessions = new Map(); + app.filePreviewContent = ''; + return app; +} + +describe('file viewer markdown line breaks', () => { + // A README hard-wrapped at the column limit: one paragraph in the source. + const WRAPPED = 'A paragraph hard-wrapped\nat the column limit.'; + + it('renders a hard-wrapped paragraph as one paragraph in the file view, while chat keeps a break per newline', async () => { + const app = loadShippingApp(`${WRAPPED}\n`); + + await app.openFilePreview('docs/README.md', 's1'); + const para = document.querySelector('#filePreviewBody .file-preview-md p')!; + expect(para, 'the document rendered through marked').not.toBeNull(); + expect(para.querySelector('br')).toBeNull(); + expect(para.textContent).toBe(WRAPPED); + + // The Response Viewer renders the same text the chat way, a <br> per newline. + const message = app._buildResponseViewerMessage(WRAPPED, 'assistant', 'Claude') as HTMLElement; + const chatPara = message.querySelector('.rv-text p')!; + expect(chatPara.querySelectorAll('br')).toHaveLength(1); + expect(chatPara.textContent).toBe(WRAPPED.replace('\n', '')); + }); +}); + describe('FILE_PREVIEW_EXTENSIONS', () => { it('routes avif and ico paths to the viewer and leaves .md with the tail viewer', () => { const { exts } = loadApp(); From 846c62fbf7595eeda0e7e1379a47a32e789e1f05 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Thu, 1 Oct 2026 11:19:41 +0200 Subject: [PATCH 44/46] fix(web): bound a pending #session= link and retire it on Home or a web tab (#507 review) - A #session=<id> link whose session never appears (closed, a typo, or another user's session in multi-user mode) is dropped after URL_SESSION_WAIT_MS (30 s) with a "Session not found" toast instead of waiting forever. One stored timer per link, cleared whenever the link is followed, replaced by a newer link, or retired. - goHome() and opening a web tab now retire a waiting link, so a session that turns up later no longer takes the screen. App-made web tab opens (frame self-recovery, the fallback after the active web tab closes) pass auto: true and keep it, as selectSession() does. - zh-CN translation for the new toast. - selectSession's auto: true comment now lists the #session=<id> link. - docs: the 30 s bound, a win.location.replace() tip that avoids piling up history entries, and the fragment declared a stable SemVer surface in versioning-policy.md. - Tests: timeout drops and toasts, an early arrival is still selected, the wait does not restart, goHome and a web tab retire it, an auto web tab open keeps it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- docs/extending-codeman.md | 11 ++- docs/versioning-policy.md | 4 + src/web/public/app.js | 76 +++++++++++++--- src/web/public/i18n.js | 2 + src/web/public/webview-tabs.js | 19 +++- test/url-session-fragment.test.ts | 141 +++++++++++++++++++++++++++++- 6 files changed, 233 insertions(+), 20 deletions(-) diff --git a/docs/extending-codeman.md b/docs/extending-codeman.md index 5e0bb605..9d19ff0d 100644 --- a/docs/extending-codeman.md +++ b/docs/extending-codeman.md @@ -355,7 +355,16 @@ When that window already shows the dashboard, only the fragment differs. The browser therefore keeps the page loaded, and the dashboard switches tabs without reloading it. A session the window has shown before appears at once. A session your page has only just created may not be listed yet, so the dashboard waits -for its `session:created` event and selects it then. +for its `session:created` event and selects it then. That wait lasts at most 30 +seconds: a link whose session never appears (a closed session, a typo, or in +multi-user mode another user's session) is dropped with a "Session not found" +notice. Clicking another tab, going Home or opening a web tab also ends the +wait, so a session that turns up later never takes the screen from the person. + +When your page holds the window reference (`const win = window.open(...)`), +prefer `win.location.replace(url)` for later links: it still fires `hashchange` +without a reload, but adds no history entry, so Back in the dashboard window +does not turn into a silent no-op. Following a link does not count as someone looking at the session, so it leaves the session's idle alert in place. The alert clears when the person diff --git a/docs/versioning-policy.md b/docs/versioning-policy.md index 44845689..4e635d07 100644 --- a/docs/versioning-policy.md +++ b/docs/versioning-policy.md @@ -40,6 +40,10 @@ A **MAJOR** bump is required to break any of these after 1.0: optional fields, new error codes, new SSE events) are non-breaking; breaking changes ship under a new prefix (`/api/v2`). The unversioned `/api/...` alias is kept working for the bundled UI. +5. **The dashboard's `#session=<id>` link.** Opening the dashboard URL with a + `#session=<id>` fragment selects that session if this client can see it. The + fragment name and that meaning are stable; see + [Opening a session from your own page](extending-codeman.md#opening-a-session-from-your-own-page). ## What SemVer does NOT cover (internal surfaces — may change in any release) diff --git a/src/web/public/app.js b/src/web/public/app.js index fcfa641d..f1d9da4a 100644 --- a/src/web/public/app.js +++ b/src/web/public/app.js @@ -567,6 +567,14 @@ const DEFAULT_SHORTCUTS = [ */ const SIDEBAR_RICH_CLOCK_MS = 20000; +/** + * How long a `#session=<id>` link waits for the session list to name its id + * before the dashboard drops it with a "Session not found" toast (see + * _armUrlSessionWait). Long enough for a page that has just created the + * session to see its session:created land here. + */ +const URL_SESSION_WAIT_MS = 30000; + class CodemanApp { constructor() { this.sessions = new Map(); @@ -605,6 +613,7 @@ class CodemanApp { // A session another page asked for with a `#session=<id>` link. It waits // here until the session list has that id (see _selectUrlSession). this._urlSessionId = this.isSoloWindow ? null : this._takeUrlSession(); + this._urlSessionWaitTimer = null; // bounds that wait (_armUrlSessionWait) this.detachedSessions = new Set(); // dashboard-side: ids currently popped out this.detachedWindows = new Map(); // dashboard-side: id -> WindowProxy this._detachWatchTimers = new Map(); // dashboard-side: id -> setInterval handle @@ -1006,6 +1015,8 @@ class CodemanApp { window.addEventListener('hashchange', () => { const id = this._takeUrlSession(); if (!id) return; + // A new link replaces one still waiting, and gets a wait of its own. + this._retireUrlSession(); this._urlSessionId = id; this._selectUrlSession(); }); @@ -1394,20 +1405,56 @@ class CodemanApp { /** Show the session a `#session=<id>` link asked for, once the session list * has it. A page that has just created a session can link to it before - * session:created arrives here, so an unknown id stays pending and - * _onSessionCreated tries again. + * session:created arrives here, so an unknown id stays pending (for at most + * URL_SESSION_WAIT_MS) and _onSessionCreated tries again. * * ⚠️ The selection is `auto`. The page that set the fragment may be a * script, and this window may not even be in front, so following a link is * not a human looking at the session and must not spend its idle alert. */ _selectUrlSession() { const id = this._urlSessionId; - if (!id || !this.sessions.has(id)) return false; - this._urlSessionId = null; + if (!id) return false; + if (!this.sessions.has(id)) { + this._armUrlSessionWait(id); + return false; + } + this._retireUrlSession(); this.selectSession(id, { auto: true }); return true; } + /** Bound the wait for a link whose id the session list does not have. A + * stale link (that session is closed), a typo, or in multi-user mode another + * user's session (never in this client's list) would otherwise wait with + * nothing on screen, and take the tab whenever a matching session turned up. + * One timer per link: handleInit running again (an SSE reconnect) does not + * restart it, and every way a link ends goes through _retireUrlSession. */ + _armUrlSessionWait(id) { + if (this._urlSessionWaitTimer) return; + this._urlSessionWaitTimer = setTimeout(() => { + this._urlSessionWaitTimer = null; + if (this._urlSessionId !== id) return; + // Listed by a path other than session:created (a session:updated): select it. + if (this.sessions.has(id)) { + this._selectUrlSession(); + return; + } + this._retireUrlSession(); + this.showToast?.('Session not found', 'warning'); + }, URL_SESSION_WAIT_MS); + } + + /** Drop a waiting `#session=<id>` link and its timer: the link was followed, + * replaced by a newer one, timed out, or the user chose something else + * (another tab, Home, a web tab). */ + _retireUrlSession() { + this._urlSessionId = null; + if (this._urlSessionWaitTimer) { + clearTimeout(this._urlSessionWaitTimer); + this._urlSessionWaitTimer = null; + } + } + /** * Pop a session out into its own browser window. SINGLE, idempotent entry * point: the tab's pop-out icon calls this, and a future gesture layer @@ -4421,6 +4468,9 @@ class CodemanApp { this._selectUrlSession(); return; } + // Not listed yet: its wait starts now that the list has loaded, and the + // last active tab is restored meanwhile. + if (this._urlSessionId) this._armUrlSessionWait(this._urlSessionId); const previousActiveId = this.activeSessionId; if (this.sessionOrder.length === 0) { @@ -6619,7 +6669,7 @@ class CodemanApp { // waiting for its session, which would otherwise take the tab from you // whenever that session turned up (see _selectUrlSession). if (options?.auto !== true && this._urlSessionId && this._urlSessionId !== sessionId) { - this._urlSessionId = null; + this._retireUrlSession(); } // If this session is popped out into its own window, raise that window // instead of showing it inline (focus-on-click for detached tabs). If we @@ -6630,12 +6680,13 @@ class CodemanApp { } const forceReload = options?.forceReload === true; // ⚠️ `auto: true` marks a selection the APP made rather than the human: - // the boot restore, a solo window opening its target, the fallback after - // the active session is deleted. Those must NOT spend a pending idle alert - // (the yellow survives until a real tap), because "the app put this on - // screen" is not "I checked it". The DEFAULT is user-initiated, so a call - // site nobody tagged fails toward acknowledging rather than toward an - // alert that can never be cleared. + // the boot restore, a solo window opening its target, a `#session=<id>` + // link from another page, the fallback after the active session is + // deleted. Those must NOT spend a pending idle alert (the yellow survives + // until a real tap), because "the app put this on screen" is not "I + // checked it". The DEFAULT is user-initiated, so a call site nobody tagged + // fails toward acknowledging rather than toward an alert that can never be + // cleared. const userInitiated = options?.auto !== true; if (this.activeSessionId === sessionId && !forceReload) { // Tapping the tab you are already on is still "I checked it". The alert @@ -7598,6 +7649,9 @@ class CodemanApp { // ═══════════════════════════════════════════════════════════════ goHome() { + // Going Home is choosing something else, so a `#session=<id>` link still + // waiting for its session must not take the screen later. + this._retireUrlSession(); // Deselect active session and show welcome screen this.activeSessionId = null; try { localStorage.removeItem('codeman-active-session'); } catch {} diff --git a/src/web/public/i18n.js b/src/web/public/i18n.js index bfa8bd7a..bf6deb26 100644 --- a/src/web/public/i18n.js +++ b/src/web/public/i18n.js @@ -571,6 +571,8 @@ 'Task Complete': '任务完成', 'Copied to clipboard': '已复制到剪贴板', 'Nothing to copy': '没有可复制的内容', + // A `#session=<id>` link whose session never appeared (app.js _armUrlSessionWait). + 'Session not found': '未找到会话', // Terminal touch-selection bar (long-press to select). The bar is a sibling of // `.xterm`, not a descendant, so SKIP_SELECTOR does not cover it and these apply. Copy: '复制', diff --git a/src/web/public/webview-tabs.js b/src/web/public/webview-tabs.js index 98d40188..01682157 100644 --- a/src/web/public/webview-tabs.js +++ b/src/web/public/webview-tabs.js @@ -274,7 +274,10 @@ Object.assign(CodemanApp.prototype, { // one `/`, and refuse whatever still opens a second one. The proxied form // is refused server-side as well (resolveUpstreamUrl). const path = data.path.replace(/[\t\n\r]/g, '').replace(/^[/\\]+/, '/'); - void this.openWebview(id, { path: path.startsWith('/') && !/^\/[/\\]/.test(path) ? path : '/' }); + void this.openWebview(id, { + path: path.startsWith('/') && !/^\/[/\\]/.test(path) ? path : '/', + auto: true, + }); return; } }; @@ -356,15 +359,23 @@ Object.assign(CodemanApp.prototype, { */ /** * @param {string} id - * @param {{path?: string}} [options] `path` (pathname+search+hash) opens a + * @param {{path?: string, auto?: boolean}} [options] `path` (pathname+search+hash) opens a * deep link inside the dashboard: appended to the proxy prefix, or resolved * against the real URL in direct mode. A mounted frame is navigated there - * rather than left on whatever page it was showing. + * rather than left on whatever page it was showing. `auto: true` marks an + * open the APP made (a frame recovering itself, the fallback after the + * active web tab closes), as on selectSession(). */ async openWebview(id, options = {}) { const webview = this.webviews.get(id); if (!webview) return; + // Opening a web tab yourself is choosing something else, so a + // `#session=<id>` link still waiting for its session must not take the + // screen from this tab later. Retired before the await below, which a + // session:created could otherwise land inside. + if (options.auto !== true) this._retireUrlSession?.(); + if (!this.webviewOrder.includes(id)) { this.webviewOrder.push(id); this._persistWebviewOrder(); @@ -513,7 +524,7 @@ Object.assign(CodemanApp.prototype, { this.activeWebviewId = null; const next = this.webviewOrder[0]; if (next) { - this.openWebview(next); + this.openWebview(next, { auto: true }); } else { this._hideWebviewLayer(); // Fall back to whatever session was last shown, or the welcome screen. diff --git a/test/url-session-fragment.test.ts b/test/url-session-fragment.test.ts index c75a8fe7..40e8ec6a 100644 --- a/test/url-session-fragment.test.ts +++ b/test/url-session-fragment.test.ts @@ -7,7 +7,7 @@ import { readFileSync } from 'node:fs'; import { resolve } from 'node:path'; import vm from 'node:vm'; import { performance } from 'node:perf_hooks'; -import { describe, expect, it, vi } from 'vitest'; +import { afterEach, describe, expect, it, vi } from 'vitest'; function loadHelper() { const context = vm.createContext({ window: {}, globalThis: {}, URLSearchParams }); @@ -50,10 +50,12 @@ describe('CodemanUrlSession.sessionIdFromFragment', () => { // The dashboard side: reading the link, holding an id it does not list yet, // and handing the selection over. Loaded like session-select-ack-gate.test.ts, -// on a bare instance whose DOM-touching methods are stubbed. +// on a bare instance whose DOM-touching methods are stubbed. webview-tabs.js +// rides along because opening a web tab is one of the ways a waiting link ends. function loadApp() { const constants = readFileSync(resolve(import.meta.dirname, '../src/web/public/constants.js'), 'utf8'); const app = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8'); + const webviewTabs = readFileSync(resolve(import.meta.dirname, '../src/web/public/webview-tabs.js'), 'utf8'); const location = { hash: '', pathname: '/', search: '' }; const history = { state: null, @@ -80,8 +82,12 @@ function loadApp() { window: { addEventListener: vi.fn(), removeEventListener: vi.fn() }, MobileDetection: { isTouchDevice: () => false }, }); - vm.runInContext(`${constants}\n${app}\nglobalThis.__CodemanApp = CodemanApp;`, context); + vm.runInContext( + `${constants}\n${app}\n${webviewTabs}\nglobalThis.__CodemanApp = CodemanApp;\nglobalThis.__waitMs = URL_SESSION_WAIT_MS;`, + context + ); const CodemanApp = (context as { __CodemanApp: { prototype: object } }).__CodemanApp; + const waitMs = (context as { __waitMs: number }).__waitMs; const make = (ids: string[]) => { const inst = Object.create(CodemanApp.prototype) as Record<string, any>; inst.sessions = new Map(ids.map((id) => [id, { id, name: id }])); @@ -90,7 +96,9 @@ function loadApp() { inst.detachedWindows = new Map(); inst.isSoloWindow = false; inst._urlSessionId = null; + inst._urlSessionWaitTimer = null; inst.selectSession = vi.fn(); + inst.showToast = vi.fn(); for (const stub of [ 'saveSessionOrder', 'markSessionTabEntering', @@ -103,7 +111,7 @@ function loadApp() { } return inst; }; - return { make, location, history, CodemanApp }; + return { make, location, history, CodemanApp, waitMs }; } describe('dashboard handling of a #session=<id> link', () => { @@ -163,6 +171,15 @@ describe('dashboard handling of a #session=<id> link', () => { expect(app._urlSessionId).toBe('later'); }); + it('starts the wait for an unlisted link once the page has loaded its session list', () => { + const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8'); + const link = source.indexOf('if (this._urlSessionId && this.sessions.has(this._urlSessionId))'); + const wait = source.indexOf('if (this._urlSessionId) this._armUrlSessionWait(this._urlSessionId);'); + const restore = source.indexOf("restoreId = localStorage.getItem('codeman-active-session')"); + expect(wait).toBeGreaterThan(link); + expect(wait).toBeLessThan(restore); + }); + it('puts the link ahead of restoring the last active tab when the page loads', () => { const source = readFileSync(resolve(import.meta.dirname, '../src/web/public/app.js'), 'utf8'); const link = source.indexOf('if (this._urlSessionId && this.sessions.has(this._urlSessionId))'); @@ -177,3 +194,119 @@ describe('dashboard handling of a #session=<id> link', () => { expect(source).toMatch(/if \(!this\.isSoloWindow\) \{\s*window\.addEventListener\('hashchange'/); }); }); + +// A link whose session never turns up: a stale link (the session is closed), a +// typo, or in multi-user mode another user's session, which is never in this +// client's list. It must not wait forever with nothing on screen, and choosing +// something else must end it, or a session turning up later takes the screen. +describe('a #session=<id> link that is still waiting', () => { + afterEach(() => { + vi.useRealTimers(); + }); + + it('waits 30 seconds', () => { + expect(loadApp().waitMs).toBe(30_000); + }); + + it('is dropped with a toast when its session has not appeared in time', () => { + vi.useFakeTimers(); + const { make, waitMs } = loadApp(); + const app = make([]); + app._urlSessionId = 'gone'; + expect(app._selectUrlSession()).toBe(false); + vi.advanceTimersByTime(waitMs - 1); + expect(app._urlSessionId).toBe('gone'); + expect(app.showToast).not.toHaveBeenCalled(); + vi.advanceTimersByTime(1); + expect(app._urlSessionId).toBeNull(); + expect(app._urlSessionWaitTimer).toBeNull(); + expect(app.showToast).toHaveBeenCalledWith('Session not found', 'warning'); + // Retired for good: the session turning up afterwards does not take the tab. + app._onSessionCreated({ id: 'gone', name: 'gone' }); + expect(app.selectSession).not.toHaveBeenCalled(); + }); + + it('still selects a session that arrives before the wait runs out, and stops the timer', () => { + vi.useFakeTimers(); + const { make, waitMs } = loadApp(); + const app = make([]); + app._urlSessionId = 'new'; + app._selectUrlSession(); + vi.advanceTimersByTime(waitMs - 1); + app._onSessionCreated({ id: 'new', name: 'new' }); + expect(app.selectSession).toHaveBeenCalledWith('new', { auto: true }); + expect(app._urlSessionWaitTimer).toBeNull(); + vi.advanceTimersByTime(waitMs); + expect(app.showToast).not.toHaveBeenCalled(); + }); + + it('does not restart its wait when handleInit asks again', () => { + vi.useFakeTimers(); + const { make, waitMs } = loadApp(); + const app = make([]); + app._urlSessionId = 'gone'; + app._armUrlSessionWait('gone'); + vi.advanceTimersByTime(waitMs - 1000); + app._armUrlSessionWait('gone'); + vi.advanceTimersByTime(1000); + expect(app._urlSessionId).toBeNull(); + expect(app.showToast).toHaveBeenCalledTimes(1); + }); + + it('is retired by going Home', () => { + vi.useFakeTimers(); + const { make, waitMs } = loadApp(); + const app = make([]); + app.terminal = { clear: vi.fn() }; + app.showWelcome = vi.fn(); + app.renderRalphStatePanel = vi.fn(); + app._urlSessionId = 'later'; + app._selectUrlSession(); + app.goHome(); + expect(app._urlSessionId).toBeNull(); + expect(app._urlSessionWaitTimer).toBeNull(); + app._onSessionCreated({ id: 'later', name: 'later' }); + expect(app.selectSession).not.toHaveBeenCalled(); + vi.advanceTimersByTime(waitMs); + expect(app.showToast).not.toHaveBeenCalled(); + }); + + function withWebTab(app: Record<string, any>) { + const webview = { id: 'dash', name: 'Dash', url: 'http://127.0.0.1:8080/' }; + app.webviews = new Map([['dash', webview]]); + app.webviewOrder = ['dash']; + app._persistWebviewOrder = vi.fn(); + app._apiJson = vi.fn(async () => ({ webview, embedUrl: '/webview/cap/' })); + app._mountWebviewFrame = vi.fn(); + app.hideWelcome = vi.fn(); + app._updateActiveWebviewTab = vi.fn(); + app.closeSessionSidebarOnHandheld = vi.fn(); + return app; + } + + it('is retired by opening a web tab', async () => { + vi.useFakeTimers(); + const { make, waitMs } = loadApp(); + const app = withWebTab(make([])); + app._urlSessionId = 'later'; + app._selectUrlSession(); + const opening = app.openWebview('dash'); + // Before the open's await: a session:created landing inside it finds no link. + expect(app._urlSessionId).toBeNull(); + expect(app._urlSessionWaitTimer).toBeNull(); + await opening; + expect(app.activeWebviewId).toBe('dash'); + app._onSessionCreated({ id: 'later', name: 'later' }); + expect(app.selectSession).not.toHaveBeenCalled(); + vi.advanceTimersByTime(waitMs); + expect(app.showToast).not.toHaveBeenCalled(); + }); + + it('survives a web tab the app opens itself', async () => { + const { make } = loadApp(); + const app = withWebTab(make([])); + app._urlSessionId = 'later'; + await app.openWebview('dash', { auto: true }); + expect(app._urlSessionId).toBe('later'); + }); +}); From f776ad87b6a4108d432b945503421f9db0316a18 Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Thu, 1 Oct 2026 11:20:39 +0200 Subject: [PATCH 45/46] chore: changeset for the 1.33.3 landing (#503, #506, #507, #509) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- .changeset/d919356e.md | 20 ++++++++++++++++++++ 1 file changed, 20 insertions(+) create mode 100644 .changeset/d919356e.md diff --git a/.changeset/d919356e.md b/.changeset/d919356e.md new file mode 100644 index 00000000..29352ef7 --- /dev/null +++ b/.changeset/d919356e.md @@ -0,0 +1,20 @@ +--- +"aicodeman": patch +--- + +### Thanks + +- @JDProfresh for rendering Markdown in the File Viewer (#503) through the chat's existing markdown pipeline and sanitizer rather than a second one, plus the Lines and Wrap toggles and the sanitizer fix that stops a document from clobbering `document.app`. +- @timkjr for bringing the Shell scroll-to-top history pull to the split view's second pane (#506), following #494's rules down to the back-off, with tests that fail on the code before each fix. +- @irisitymichaelgrundberg for the `#session=<id>` dashboard link (#507), so a page that keeps one Codeman window open can switch it between sessions without reloading it. +- @dignfei for handing focus back when the Command Palette or the Session Manager closes (#509), and for the six-overlay measurement that showed exactly which two were broken. + +**Markdown files render in the File Viewer (#503).** Opening a `.md` or `.markdown` file now shows it as a document: headings, tables, code blocks with the same copy buttons as the chat, images relative to the file, and links to other documents that open inside the viewer. An `MD` pill switches back to the source, and Edit works from either view. Plain text gets a `Lines` gutter (never part of a copy) and a `Wrap` toggle, all three remembered per device. `.avif` images preview inline, and printed `.avif`/`.ico` paths open the viewer instead of the tail view. An in-workspace file path clicked in the terminal still opens the live tail view. + +**Link a dashboard window to a session (#507).** An outside page, such as a task board, that keeps one Codeman window open can now switch it to a session by pointing it at `/#session=<id>`. Only the fragment changes, so the page stays loaded and the switch is an ordinary tab selection. A link to a session the dashboard does not list yet waits up to 30 seconds for it to appear and then shows "Session not found"; picking another tab, going Home or opening a web tab cancels the wait. Following a link does not count as looking at the session, so its idle alert stays armed. The fragment is documented in `docs/extending-codeman.md` and is now a stable surface under `docs/versioning-policy.md`. + +**Escape no longer strands the keyboard (#509).** Closing the Command Palette or the Session Manager now hands focus back to whatever held it before they opened, usually the terminal, so you can keep typing without clicking first. An Escape pressed while neither is open changes nothing. + +**Split view: a Shell Pane B scrolls back into tmux history (#506).** Wheel up at the top of a Shell session in the split view's second pane now pulls the most recent 1 MiB of its tmux history and keeps your place, the same as the primary pane since 1.33.2. + +**Fixes applied while landing.** Markdown opened from an attachment card no longer resolves relative images and links against the workspace root, where they could show a missing image or open a different file of the same name; they render as their alt text and link text instead. Rendered files no longer turn every source line break into a hard break the way chat messages do, so a README wrapped at 80 columns reads as flowing paragraphs. Absolute-path links inside a rendered document open in that document's session. A disconnected Pane B keeps its "disconnected" marker as the last line even when the socket closes in the middle of a history pull. Closing the Session Manager through a row's "Switch to session" or "Open folder" no longer pulls focus back from the terminal to the header button. From 9240493c43b59d233bf66263592b47a7b024763e Mon Sep 17 00:00:00 2001 From: Codeman maintainer <noreply@anthropic.com> Date: Thu, 1 Oct 2026 23:51:32 +0200 Subject: [PATCH 46/46] chore: version packages (1.33.3) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --- .changeset/d919356e.md | 20 -------------------- .claude-plugin/marketplace.json | 2 +- CHANGELOG.md | 20 ++++++++++++++++++++ CLAUDE.md | 2 +- package-lock.json | 4 ++-- package.json | 2 +- plugins/codeman/.claude-plugin/plugin.json | 2 +- 7 files changed, 26 insertions(+), 26 deletions(-) delete mode 100644 .changeset/d919356e.md diff --git a/.changeset/d919356e.md b/.changeset/d919356e.md deleted file mode 100644 index 29352ef7..00000000 --- a/.changeset/d919356e.md +++ /dev/null @@ -1,20 +0,0 @@ ---- -"aicodeman": patch ---- - -### Thanks - -- @JDProfresh for rendering Markdown in the File Viewer (#503) through the chat's existing markdown pipeline and sanitizer rather than a second one, plus the Lines and Wrap toggles and the sanitizer fix that stops a document from clobbering `document.app`. -- @timkjr for bringing the Shell scroll-to-top history pull to the split view's second pane (#506), following #494's rules down to the back-off, with tests that fail on the code before each fix. -- @irisitymichaelgrundberg for the `#session=<id>` dashboard link (#507), so a page that keeps one Codeman window open can switch it between sessions without reloading it. -- @dignfei for handing focus back when the Command Palette or the Session Manager closes (#509), and for the six-overlay measurement that showed exactly which two were broken. - -**Markdown files render in the File Viewer (#503).** Opening a `.md` or `.markdown` file now shows it as a document: headings, tables, code blocks with the same copy buttons as the chat, images relative to the file, and links to other documents that open inside the viewer. An `MD` pill switches back to the source, and Edit works from either view. Plain text gets a `Lines` gutter (never part of a copy) and a `Wrap` toggle, all three remembered per device. `.avif` images preview inline, and printed `.avif`/`.ico` paths open the viewer instead of the tail view. An in-workspace file path clicked in the terminal still opens the live tail view. - -**Link a dashboard window to a session (#507).** An outside page, such as a task board, that keeps one Codeman window open can now switch it to a session by pointing it at `/#session=<id>`. Only the fragment changes, so the page stays loaded and the switch is an ordinary tab selection. A link to a session the dashboard does not list yet waits up to 30 seconds for it to appear and then shows "Session not found"; picking another tab, going Home or opening a web tab cancels the wait. Following a link does not count as looking at the session, so its idle alert stays armed. The fragment is documented in `docs/extending-codeman.md` and is now a stable surface under `docs/versioning-policy.md`. - -**Escape no longer strands the keyboard (#509).** Closing the Command Palette or the Session Manager now hands focus back to whatever held it before they opened, usually the terminal, so you can keep typing without clicking first. An Escape pressed while neither is open changes nothing. - -**Split view: a Shell Pane B scrolls back into tmux history (#506).** Wheel up at the top of a Shell session in the split view's second pane now pulls the most recent 1 MiB of its tmux history and keeps your place, the same as the primary pane since 1.33.2. - -**Fixes applied while landing.** Markdown opened from an attachment card no longer resolves relative images and links against the workspace root, where they could show a missing image or open a different file of the same name; they render as their alt text and link text instead. Rendered files no longer turn every source line break into a hard break the way chat messages do, so a README wrapped at 80 columns reads as flowing paragraphs. Absolute-path links inside a rendered document open in that document's session. A disconnected Pane B keeps its "disconnected" marker as the last line even when the socket closes in the middle of a history pull. Closing the Session Manager through a row's "Switch to session" or "Open folder" no longer pulls focus back from the terminal to the header button. diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index e5283924..13c50a7a 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -10,7 +10,7 @@ "name": "codeman", "source": "./plugins/codeman", "description": "Drive Codeman from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", - "version": "1.33.2", + "version": "1.33.3", "author": { "name": "Ark0N", "url": "https://github.com/Ark0N" diff --git a/CHANGELOG.md b/CHANGELOG.md index 53aeb9c8..9624582d 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,25 @@ # aicodeman +## 1.33.3 + +### Patch Changes + +- f776ad8: ### Thanks + - @JDProfresh for rendering Markdown in the File Viewer (#503) through the chat's existing markdown pipeline and sanitizer rather than a second one, plus the Lines and Wrap toggles and the sanitizer fix that stops a document from clobbering `document.app`. + - @timkjr for bringing the Shell scroll-to-top history pull to the split view's second pane (#506), following #494's rules down to the back-off, with tests that fail on the code before each fix. + - @irisitymichaelgrundberg for the `#session=<id>` dashboard link (#507), so a page that keeps one Codeman window open can switch it between sessions without reloading it. + - @dignfei for handing focus back when the Command Palette or the Session Manager closes (#509), and for the six-overlay measurement that showed exactly which two were broken. + + **Markdown files render in the File Viewer (#503).** Opening a `.md` or `.markdown` file now shows it as a document: headings, tables, code blocks with the same copy buttons as the chat, images relative to the file, and links to other documents that open inside the viewer. An `MD` pill switches back to the source, and Edit works from either view. Plain text gets a `Lines` gutter (never part of a copy) and a `Wrap` toggle, all three remembered per device. `.avif` images preview inline, and printed `.avif`/`.ico` paths open the viewer instead of the tail view. An in-workspace file path clicked in the terminal still opens the live tail view. + + **Link a dashboard window to a session (#507).** An outside page, such as a task board, that keeps one Codeman window open can now switch it to a session by pointing it at `/#session=<id>`. Only the fragment changes, so the page stays loaded and the switch is an ordinary tab selection. A link to a session the dashboard does not list yet waits up to 30 seconds for it to appear and then shows "Session not found"; picking another tab, going Home or opening a web tab cancels the wait. Following a link does not count as looking at the session, so its idle alert stays armed. The fragment is documented in `docs/extending-codeman.md` and is now a stable surface under `docs/versioning-policy.md`. + + **Escape no longer strands the keyboard (#509).** Closing the Command Palette or the Session Manager now hands focus back to whatever held it before they opened, usually the terminal, so you can keep typing without clicking first. An Escape pressed while neither is open changes nothing. + + **Split view: a Shell Pane B scrolls back into tmux history (#506).** Wheel up at the top of a Shell session in the split view's second pane now pulls the most recent 1 MiB of its tmux history and keeps your place, the same as the primary pane since 1.33.2. + + **Fixes applied while landing.** Markdown opened from an attachment card no longer resolves relative images and links against the workspace root, where they could show a missing image or open a different file of the same name; they render as their alt text and link text instead. Rendered files no longer turn every source line break into a hard break the way chat messages do, so a README wrapped at 80 columns reads as flowing paragraphs. Absolute-path links inside a rendered document open in that document's session. A disconnected Pane B keeps its "disconnected" marker as the last line even when the socket closes in the middle of a history pull. Closing the Session Manager through a row's "Switch to session" or "Open folder" no longer pulls focus back from the terminal to the header button. + ## 1.33.2 ### Patch Changes diff --git a/CLAUDE.md b/CLAUDE.md index 380559ae..2ba2a201 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -78,7 +78,7 @@ When user says "COM": CI runs `npm run check:lockfile` on every push/PR, so lockfile drift fails the build even if the `version-packages` script is bypassed. -**Version**: 1.33.2 (must match `package.json`) +**Version**: 1.33.3 (must match `package.json`) ## Project Overview diff --git a/package-lock.json b/package-lock.json index d929379e..271350e4 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "aicodeman", - "version": "1.33.2", + "version": "1.33.3", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "aicodeman", - "version": "1.33.2", + "version": "1.33.3", "hasInstallScript": true, "license": "MIT", "workspaces": [ diff --git a/package.json b/package.json index 3f7aa92e..da2a097b 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "aicodeman", - "version": "1.33.2", + "version": "1.33.3", "description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence", "type": "module", "main": "dist/index.js", diff --git a/plugins/codeman/.claude-plugin/plugin.json b/plugins/codeman/.claude-plugin/plugin.json index 93cb5488..37304657 100644 --- a/plugins/codeman/.claude-plugin/plugin.json +++ b/plugins/codeman/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "codeman", "description": "Drive Codeman, the self-hosted session manager for AI coding agents, from inside a Claude Code session: spawn worker sessions, prompt them, wait for them, read their answers, clean up. Acts only inside a Codeman-managed session.", - "version": "1.33.2", + "version": "1.33.3", "author": { "name": "Ark0N", "url": "https://github.com/Ark0N"