mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
120d780267ac2f133c0b735313fcf751231d3511
2461
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
120d780267 |
chore: version packages (#474)
* chore: version packages * docs: sync CLAUDE.md version to 1.32.1 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Codeman maintainer <noreply@anthropic.com>codeman@1.32.1 |
||
|
|
0af925fe82 |
Merge remote-tracking branch 'origin/master' into land/1.32.1
# Conflicts: # CLAUDE.md # docs/architecture-invariants.md |
||
|
|
b404dacfde |
chore: changeset for the merge-time fixes and thanks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
cbd1fa639d |
fix(tmux): merge-time fixes for the exited-agent report (#466)
- docs/wiki/The-Dashboard.md: the tab-appearance table gains the exited state (muted dot plus an `exited (137)` badge) and explains the bare `exited` variant. - The detailed sidebar and rail no longer pair the muted dot with an "idle" pill: an exited session's pill reads "exited" (neutral styling) and its since stamp measures from the observed exit. This is a label override on the row model, not a new state, so SESSION_ACTIVITY_RANK and the home screen order are untouched, and a pending alert still keeps its own pill. The row signature includes the flag so the incremental path repaints it. - The exited badge is aria-hidden like its sibling badges, and the exit is appended to the tab's aria-label in both render paths through one helper. - test/tmux-manager.test.ts re-adds the junk-trailing-field parser case against parsePaneRows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
43d4be8eeb |
fix(session): merge-time fixes for the dead-pane resume pin (#467)
- test/setup.ts strips CLAUDE_CONFIG_DIR (pinned in test-env-isolation), so transcript-fixture tests such as session-custom-model-restart no longer go red on a machine that exports it for a separate Claude account (#255). - The vanished-tmux-session branch of _setupOrAttachMuxSession() relaunches the CLI through createSession() just like a failed respawn, so it now takes the same resume pin. A genuinely new session is unaffected. - After a dead-pane respawn of a fallback-chain CLI, _claudeSessionId names the conversation the walk actually pinned instead of the chain tail, which the walk may have passed over for lack of a transcript. - _claudeConfigDir() trims the override like claudeProjectsDir() does. - The remote-reattach test is labelled as documentation, since the pin builder's own remote guard would make it pass either way. - CLAUDE.md: the create-path pin persists through toState() as resumeSessionId, and the end of the walk adds no pin rather than clearing the launch seed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
f4d1ee8027 |
fix(terminal): merge-time fixes for dropped-output recovery (#470)
- The TERMINAL DROP crash-trail line moves behind the scheduler's debounce guard, so it is written once per window rather than once per dropped frame. At the server's 8ms batching, one second of drops evicted the whole 50-entry trail, including the recovery lines that explain it. - A refresh that failed at the capture fetch deadline now returns 'deadline', and the scheduler does not retry it: that is a stalled link, not contention, and each retry was another ?full=1 capture waiting out a deadline of up to two minutes. The early-return retries are unchanged. CLAUDE.md and the code comments no longer claim every skip reason is transient contention. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
7a30a31430 |
fix(terminal): merge-time fixes for the silent-failure paths (#431)
- While another device holds the pane width (_paneWidthRefused), a resize now asks for the container's width without applying it locally (_geometryForResizeRequest: rows follow the container, columns stay at the PTY's). Fitting first re-wrapped the whole buffer to the container and back on every 30s mobile retry, and throttledResize ran the scrollback clear for a resize that brings no redraw. selectSession clears the flag, since it belongs to the previous pane. New unit tests run the real mixin against a fake terminal and fail without the fix. - Session seeds _ptyCols/_ptyRows at spawn (_notePtySpawnGeometry), so a reattached pane reports its tmux window's real size through ptyGeometry. - Session.resize's declined-branch comment names ptyGeometry, not the deleted ptyCols/ptyRows getters. - Delete the dead terminalGeometryAgrees() and its window export. - test/xterm-private-api.test.ts header: it pins the exact locked version, so any bump fails, not only a major. - The main-terminal fit sweep also matches fitAddon?.fit?.(), and CLAUDE.md names the modules it actually covers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
fdfcc15c10 |
docs(docker): merge-time note for the gh/az sign-in in multi-user mode (#472)
Clone Repo clears the credential helpers for a non-admin, but a non-admin's Docker case with credential seeding on still receives a copy of the server account's gh/az sign-in when the agent-image switches are on, the same as the Claude and Codex credentials. Say so in the multi-user notes so the docs do not read as a stronger guarantee than they are. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
697b05b118 |
fix(docker): merge-time fixes for Update-Codeman.sh (#465)
- Remove exactly the codeman-node-modules/codeman-dist volumes by Compose label after a plain `down`, instead of `down --volumes` (which also takes any volume an override file declares while the message named two). `down --volumes` remains only as a warned fallback when the project name cannot be resolved. - Report a failing first `docker compose config --format json` call with a clear error instead of exiting silently under `set -e`. - Filter empty label lines in the collision guard so an unlabelled container cannot hide a real collision; name the moved-checkout exit in its error. - Comments no longer cite a guard or incident in Start-Codeman.sh that does not exist; the README states the real gap (a Node base-image bump leaves codeman-node-modules stale because the lockfile did not move). - docs: Update-Codeman.sh in the docker-self-update.md short-version table and a mention in docker-compose.md; "Major updates" moved under "Updating" in docker/README.md. - test: smoke test covers the new sequence, the config failure and the empty-line case; quiet stdio; @fileoverview names the fourth concern. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
0462a5d5a0 |
fix(approvals): merge-time fixes for the watching badge (#473)
- session.ts: a pane capture that fails now CLEARS the watching label (and emits watchingChanged so pages drop the badge) instead of keeping the last one, so a failed capture degrades toward an alert rather than pre-acknowledging the next real idle prompt. Test updated; invariant noted in architecture-invariants. - approvals-ui.js: the header bell counts only unacknowledged items (pendingApprovalsCount), matching codeman tui's pendingApprovalCount(); pinned in watching-no-alert.test.ts. - mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back onto MOBILE_OVERVIEW_PILL_LABEL. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
da6fa663e7 |
fix(terminal): merge-time fixes for the copy gutter strip (#469)
- stock.ts: claude is no longer the only entry declaring transcriptGutter; codex declares it too. - architecture-invariants: the strip applies when the session's CLI declares a margin (not detection), and a note that it keys on the session's launch mode, not on what is running in the pane (a claude pane dropped to a shell still loses up to two columns; copyStripMargin is the escape hatch). - render-index-html test: the gutter map is injected for a solo /session/:id render as well. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
10f87428c3 |
fix(cli-registry): merge-time fixes for the run-button accents (#463)
- mobile.css: gemini and antigravity run/gear rules get `!important` like pi/omp/grok/deepseek, so the gear half no longer keeps the skin accent while the body takes the mode colour (two-tone button on the default skin). - test/skin-themes.test.ts: static guard that every run mode with a base `.btn-toolbar.btn-run.mode-<id>` rule also has a resting rule inside the `html:not([data-skin="og"])` block; ids are derived from the stylesheet. - stock.ts: grok's accent comment names zinc-300 (border/badge colour); gemini's accent is #8ab4f8 to match its tab badge and run-mode dot, noted as the one exception to the border-colour method. - types.ts: "(below)" -> "(above)". - docs/cli-registry.md, CLAUDE.md: `accent` is now measured, not transcribed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
13e652e43f |
chore: changesets for the 1.32.1 batch
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
2afb1c2c2e |
docs: trim CLAUDE.md from 265 KB to 142 KB, detail moved to architecture-invariants
CLAUDE.md loads into every session, and its Architecture section had grown feature write-ups (history, measurements, rationale) that belong in docs/architecture-invariants.md per the file's own header. Each long block now keeps what the feature is, where it lives, its setting/default and the rules that prevent real bugs, and links to its invariants section. Everything removed was moved there: 29 new sections, extra facts appended to the existing ones. Also: hard-coded counts (SSE events, route handlers, module/file counts, device profiles) replaced by pointers to the source of truth, and the Debugging commands fixed to use the codeman tmux socket and HTTPS for prod. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
2fb744f865 |
Merge pull request #470 from rounakdatta/fix/dropped-output-recovery
fix(terminal): recover a dropped output frame, do not merely schedule it |
||
|
|
de4b1db490 |
Merge pull request #431 from rounakdatta/feat/mobile-terminal-resilience
fix(terminal): four silent-failure paths — renderer freeze, replay race, reconnect gap, unbounded fetches |
||
|
|
8536aaef7b |
Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468) # Conflicts: # src/config/cli-registry/stock.ts |
||
|
|
94b093b617 |
Merge pull request #469 from irisitymichaelgrundberg/feat/copy-dedent-pane-margin
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares |
||
|
|
6a01412af9 |
Merge pull request #466 from irisitymichaelgrundberg/feat/pane-exit-reporting
feat(tmux): report that a pane's agent has exited (#446, part 1) |
||
|
|
bc04b6457e |
Merge pull request #467 from irisitymichaelgrundberg/fix/respawn-session-id-collision
fix(session): resume the conversation when respawning a dead pane |
||
|
|
5b5e932ec4 |
Merge pull request #465 from opticon454/chore/docker-major-update-script
chore(docker): add Update-Codeman.sh for scripted major-update rebuilds # Conflicts: # docker/README.md |
||
|
|
f1dfbcdd65 |
Merge pull request #472 from opticon454/feature/git-host-auth-clis
feat(docker): opt-in gh + az CLIs with git credential helpers so Clone Repo and Docker cases can reach private repos |
||
|
|
233af33dac |
Merge pull request #463 from opticon454/fix/cli-accent-colours
fix(cli-registry): correct accent colours, and a real gemini/antigravity/omp rendering bug |
||
|
|
993e5e021c |
Merge pull request #471 from DodgyBadger/fix/mobile-blank-long-press
fix(mobile): swallow blank-space terminal long presses |
||
|
|
02e40f506b |
fix(docker): gate gh/az seeding on its switch; no shared git sign-in for non-admin clones
Addresses the review on #472. - CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv` (CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts skips a store unless that variable is exactly `1`, read at container create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL cache no longer copies them into every case container. Tests: the default environment seeds neither even with the files present, and each store follows only its own switch. - Multi-user mode: a non-admin's Clone Repo clone and preflight run with `git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the subcommand), so the server account's helpers are never lent to them. Verified against a real private repo that it also clears the URL-scoped credential.<url>.helper entries, and that public clones still work. Tests: the argv in test/git-clone.test.ts, and the route decision (non-admin cleared; admin and single-user kept) in test/routes/case-clone-credential-helpers.test.ts. - Docs: recreate the case container to pick up seeds (docker/README.md, Docker-Cases wiki, docker-cases.md); the multi-user behaviour in docker/README.md and security-architecture.md; "functionally unchanged" instead of "unchanged" for an image built with both switches off (server.Dockerfile comment, README, changeset). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw |
||
|
|
e558264977 |
chore: leave the changeset to the maintainer
CONTRIBUTING says releases are handled by the maintainer via changesets after merge, and every `.changeset/*.md` on master was written by him or by the release bot — including the ones covering other people's pull requests. The summary this file carried moves to the pull-request description, where it is the maintainer's to reuse or rewrite. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9a48c43aa1 |
docs(watching): a restart is not a gap, and here is the measurement
Claimed after a manual test that a session comes back from a server restart without its badge until it next produces output. Measured instead of assumed, and it is wrong: a codex session with a background terminal still running had its label back within about 20 seconds of the restart, with no input from anyone. Reconciliation re-attaches the pane, the attach repaint carries the composer glyph, the idle confirmation arms on it, and the probe re-reads the label — the ordinary path, doing the ordinary thing. What produced the false claim was a session whose monitor had simply expired while it sat there. Its footer carries no chip, so `watching: null` was the right answer and there was nothing missing to restore. Recorded at the field and in the invariants, because the shape of this invites exactly one wrong fix: a polling timer to keep a value fresh that the pane already refreshes by itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cf5a45438 |
feat(docker): opt-in gh + az CLIs with git credential helpers for private repos
Add Case -> Clone Repo could only reach public repositories in the Docker deployment. This lets a deployment opt in to the GitHub CLI and the Azure CLI (+ azure-devops extension) as git credential helpers. Codeman itself still collects no credentials. - server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH / CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the build). Off leaves no apt repository, package, extension, helper script or credential entry, so a default build is unchanged. On installs from the vendors' apt repositories and configures system gitconfig helpers: github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com / *.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT). A helper whose CLI is not signed in prints nothing, so a private clone still fails fast. - The extension lives in AZURE_EXTENSION_DIR outside HOME (/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0 group-writable in the agent image). - Hosts turn them on in docker-compose.override.yml: `build: args:` for the server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the agent image. build-agent-image.mjs and the in-app auto-build share one env -> ARG table (pinned by the parity test) and pass nothing when unset. docker-compose.yaml is untouched; .env.example only gains a comment, so the self-updater's environment gate sees no new keys. - Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and the az sign-in files from ~/.azure per file, read-only, like pi/grok. - The Clone Repo AUTH_REQUIRED message says how to sign the server's git in instead of claiming private repositories cannot be cloned. - Docs: docker/README.md "Private repositories", docker-compose.md, docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki pages, security-architecture.md, architecture-invariants.md, changeset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw |
||
|
|
110c4696ad | fix(mobile): swallow blank-space long presses | ||
|
|
ac6236b268 |
fix(terminal): clean a copy once, and reach every pane that copies
Review fixes for #469. The Ctrl+C branch cleaned the selection to decide whether to copy and then passed that cleaned string to copyTerminalSelection(), which cleans again. The trailing trim is a fixed point, so that was safe until this PR; the margin strip is not, because it takes the lesser of the declared width and the run every line shares, so a second pass takes up to `margin` columns more. The branch now gates on the cleaned string and hands the raw one on. Verified in chromium with a real drag, a real Ctrl+C and a real clipboard read on a live claude pane: an on-screen ` fix(terminal): trim it` reaches the clipboard as ` fix(terminal): trim it`, and reverting the branch reproduces the reported ` fix(terminal): trim it`. Pane B of a split resolves its own width. `_cliGutterColumns()` and `_normalisedSelectionRange()` take the session and the terminal to read, defaulting to the primary pane's, so Pane B looks its own run mode up instead of keeping a margin Pane A drops on the same keystroke. Verified live with two claude panes open side by side. A detached session window (`/session/:id`) receives the gutter map. The injection sat inside the block that skips the run menu's payloads for a solo window, so the toggle worked in the main window and did nothing in the popup on the same device. It needs no availability probe, so it moved below that block and the solo window still carries none of the payloads it skipped before. The settings description said the width is measured and named Codex as exempt. Nothing is measured, and Codex is one of the two panes that are stripped. docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing else" one sentence before the leading-margin rule, and both it and docs/architecture-invariants.md record that the strip is not idempotent. Two round-trip tests run on a mode that declares a gutter, which the existing copyTerminalSelection cases could not, since they all use the harness default mode that declares none. The Ctrl+C branch itself is pinned at the source, because it lives inside initTerminal's attachCustomKeyEventHandler closure over a real xterm the vm harness cannot build. Both pins fail on the reintroduced bug. Gate: 7865 passed, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00f022ccf8 |
fix(terminal): recover a dropped output frame, do not merely schedule it
`_onSessionTerminal` drops an incoming frame once the app-owned render queues already hold 128KB. That is the right call — the alternative is an unbounded backlog — but a hole in a TUI byte stream is a desynced cursor, and a desynced cursor is muffled text (#464). The drop was only half of it. The recovery was a fire-and-forget timer: it nulled its own handle and then called `_onSessionNeedsRefresh()`, which opens with four early returns. Two of them — a buffer load already in flight, a refresh already owning this session — are MOST likely to be true during exactly the output burst that caused the drop. So the recovery was skipped precisely when it was needed, with nothing left to retry it, and the dropped bytes were never replayed. `_onSessionNeedsRefresh` reports whether it actually repainted now, and `_scheduleDroppedOutputRecovery` re-arms while it has not. Bounded by `DROP_RECOVERY_MAX_ATTEMPTS`, because every reason the refresh can be skipped is transient contention that clears in seconds and a permanently failing refresh must not become a loop against the API; giving up at the cap leaves exactly what the old code left, so the floor is no worse than before. The same 2s debounce still collapses a burst of drops into one attempt. This is the principle Ark0N established reviewing #431 for the WebSocket output-gap marker — only a repaint that actually happened settles the recovery — applied to the one recovery path that still trusted a timer having fired. The retry decision is a pure function in constants.js so the gate can reach it, and the scheduler itself is driven from app.js under a fake clock. The retry case and the no-retry case only pin the fix AS A PAIR: either alone passes against something wrong, one against the old fire-and-forget timer and the other against retrying forever. Checked by reverting app.js to the old shape, where three of the twelve fail. Two harness details that would otherwise have made the tests lie. The vm context baked in the real `setTimeout`, so `vi.useFakeTimers()` could not reach the scheduler and every case reported zero calls; it delegates lazily now. And app.js reached `CodemanDroppedOutput` as a bare global, which resolves in a browser but not in the vm — worth fixing beyond the test, because that call sits inside a timer where a ReferenceError is swallowed and would take the recovery with it. It reads through `window.` like terminal-ui.js does with its own constants. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b13672596f |
fix(watching): tell the page when the badge goes away
Reported from a manual test: a Codex session went on showing the watching badge after its background terminal had finished. The server was right and the page was stale — `Session.watching` changes while the session's status does not, and nothing broadcast it. The label is usually SET on the idle transition, which broadcasts anyway, so the badge always appeared correctly. It CLEARS when the work ends, and a CLI can end background work without taking a turn: codex repaints its background-terminal row away and stays idle, so `_confirmIdle()` concludes without emitting `idle` (that emit is guarded by `wasWorking || isInitialReady`) and no other event fires. Every open page kept drawing a badge the server had already dropped. `_readWatching()` now emits `watchingChanged` when, and only when, the label really changes, and the wiring pushes the session state on it. No new SSE event: the badge reads off the session payload every surface already has. A/B measured on an isolated beta with the page loaded and then left untouched. Without this commit the server dropped the label at t+50s and the page still showed the badge at t+100s; with it, page and server cleared in the same ten-second window. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8d45b92eba |
docs(codex): record that a sub-agent leaves no row to read
Measured on the same beta, codex-cli 0.154.0: a sub-agent started without waiting outlives the turn exactly as a background terminal does — the sandboxed process was still running — and codex shows nothing for it. The last rows of the pane are the composer and the status line, and `Sub-agents running` belongs to the on-demand `/subagents` panel rather than to the row above the composer. So there is no second codex label to add. A codex session waiting on a sub-agent reads as plainly idle, which misfiles nothing (codex raises no idle prompts) and simply leaves that one kind of quiet unexplained until codex pins a row for it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
05c788ce9d |
fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief) found the trust boundary weaker than the comments around it claimed. Eleven findings, all applied. The two blockers were both about who can write the row the label is read from. Claude's window covered two rows, and the second one is the status line, whose command a session running with permissions bypassed can write into its own `.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of its own and silence its own idle alert. The default window is one row now, which is the footer and nothing else, and the constant says why. Separately, the label reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML, which is an injection sink for any config-supplied pattern whose capture group is permissive; it goes through escapeHtml() like every other untrusted string in that file. The Codex entry could not be fixed the same way, and now says so. Its row is third from the bottom only while a terminal runs; with none running that slot holds the last row of the transcript, so matching the complete row (with the `/stop to close` tail, window narrowed to three) raises the bar without closing it. What contains it is `hooks: 'none'`: no hook event from a codex session reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an alert. The registry comment, `docs/cli-registry.md` and the test all state that rather than claiming a guarantee the code does not have. Also from the review: the TUI header badge no longer counts an acknowledged item, which was the same gate the classifier fix already went through and was wrong for human acknowledgement too; the TUI approval card reads the quiet reason and drops to a new `info` tone instead of asking for a reply; the badge carries an aria-label, because the phone it was built for has no hover target; the schema refuses `watchingLines` without a `watchingLine`; and the pattern and its window are resolved together rather than one memoized and one not. Documentation moved with it. The mechanism now lives in `docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer, `docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped buzzing, and both that page and the changeset name the limitation neither did before: a question asked in plain prose is not a dialog, so it is silenced along with the false alarms while background work runs. Verified live again after the narrowing, on an isolated beta: a Claude session reported `1 monitor` and took its idle prompt acknowledged, and a Codex session reported `1 background terminal` against the full-row anchor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9c286eeddf |
fix(session): persist an exit retraction, and let tests reach the watcher
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.
A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.
The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.
The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.
Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
e1e7dc5bd8 |
fix(terminal): Ark0N's read of the #464 geometry work
Five items, two of which he could only see by running it, plus six smaller ones. Taking the two blockers first, because both were wrong in ways the existing tests could not catch. **Adopting the PTY's rows put the CLI's input line off-screen.** A phone that took a desktop's 43 rows into a viewport with room for 18 painted an `.xterm-screen` far taller than its container; xterm's own viewport then had nothing to scroll, so the bottom of the frame sat below the container with no gesture able to reach it. Output visible, typing invisible, for as long as the desktop kept the claim hot. `reconcilePtyGeometry` adopts COLUMNS ONLY now: width is the axis Ink's wrap and `eraseLines` arithmetic depend on, and keeping the local row count keeps the composer at the bottom of a viewport that scrolls. Measured at his geometry — a 360x300 container against a 198x43 pane now keeps 13 rows, takes 198 columns, paints 202px into a 210px container, and the input line is inside the box. **`capture-geometry-retry.browser.test.ts` failed, and CI could not see it** because the file is in `BROWSER_TEST_GLOBS`. Its premise WAS the clamp — `getTerminalDimensions()` floored while `fitAddon.fit()` did not — which this work removes at the source, so it can never hold again at any viewport. The case survives on its own terms: a pane already drawing at the requested size must not be replayed. Its premise is now the #464 invariant itself, that the floored report and the terminal agree, which is a stronger guard because the clamp coming back fails it here rather than silently restoring the replay loop. The helper docblock that repeated the old premise is corrected too. **A session with no pane reported 120x40 and the client adopted it.** `resize()` writes `_ptyCols`/`_ptyRows` only when `ptyProcess` is set and nothing seeds them from the spawn geometry, so a dead-pane session still held the constructor defaults — clicking that tab resized the browser terminal to 120x40 and, on anything narrower, claimed another device owned the pane when none existed. `Session.ptyGeometry` returns null without a pane, the HTTP route answers `{}` and the socket sends no frame at all. The raw `ptyCols`/`ptyRows` getters are deleted rather than left available to be misused again. **The 40-column floor clipped the pane with nothing able to reach it.** The affordance keyed on a PTY mismatch, and the floor produces no mismatch — xterm and the PTY agree throughout, the terminal is simply wider than the box. It keys on what does not FIT now, MEASURED (`.xterm-screen` against the container, on the next frame, because the screen takes its width with the render) rather than derived from cell arithmetic. Measured at 360px: font 24 applies 40 columns and paints 560px, and all 200px of the overhang is reachable. `.pty-oversized` is renamed `.term-overflows-x`, because after this the old name describes only one of the two causes. **"Scroll sideways" did not work on touch for the sessions it targets.** `touch-action: pan-x` is cancelled before it starts by the `preventDefault()` `touchstart` calls on every 'content' tap. The terminal's own touchmove handler pans the container now, with the axis locked once per gesture so a diagonal cannot pan and scroll at once, and the CSS grants no `touch-action` at all — handing the browser a pan AS WELL would move the pane twice for one finger on the taps where that preventDefault does not run. Measured under real touch dispatch: a 140px swipe reaches `scrollLeft` 140 where it reached 0 before, the buffer does not move with it, and a vertical swipe still scrolls the scrollback. Three defects in the above, found while checking it rather than by being told: - `canPanHorizontally` first tested `scrollWidth > clientWidth` alone, which is true of a container that is not a scroller — a sideways swipe would have locked the axis, done nothing, AND suppressed the vertical scroll it should have been. Gated on the class as well. - The notice advised scrolling sideways whenever the PTY was wider, including when it still fitted and nothing scrolled. It is gated on measured overflow, and on a comparison against the width this container WOULD request rather than the one it currently holds — once adopted those are equal, so the second question answers itself false while the condition is still true. - `_syncTerminalOverflowAffordance` could throw out of `document.getElementById` before reaching its try block. It runs off every geometry change, so a cosmetic affordance could have taken the resize down with it. The smaller items: - `docs/architecture-invariants.md` no longer explains the equality guard as a clamp signature; it records what the clamp used to do and why it cannot any more. Edited by hand — that file is outside the Prettier glob, and letting Prettier near it rewrote eleven unrelated emphasis markers. - `throttledResize`'s HTTP fallback reads the reply. It is the path where a declined resize is least likely to be noticed, because no socket means no `{"t":"zc"}` frame either. - The changeset covers the whole release: the geometry work, the queued replay clear, the renderer watchdog, the body-covering fetch deadline, the WebSocket output-gap reconcile, the build-generated service-worker precache and per-build cache key, and the crash-trail hygiene. - `@xterm/headless` is declared in the root devDependencies instead of being reached through workspace hoisting. - The output-gap marker is cleared after any response arrives, not only when the capture was non-empty: a server that answers with an empty capture HAS reconciled us, and leaving the marker set refetched on every reconnect. - `e587d845`'s message claimed a test asserted the failed-load copy against the built asset. It did not — that assertion lived in a probe deleted with the other scratch scripts, so the claim was false when it was written. There is a real test now, and it reads the source rather than `dist/`, because `dist/` is not committed and a test that skips when it is absent would pass for the wrong reason in CI. `Session.ptyGeometry` gets behavioural coverage against the real class in `session-resize-arbitration.test.ts` rather than a source guard, including the contrast — a pane that does exist still reports, and still follows a resize — so "always null" would fail it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ce80b7a212 |
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own two-column transcript gutter on the clipboard, so every pasted line arrives indented. #451 shipped the trailing half of the copy clean and left the leading half out, because deriving the width from the selection fires on 73% of ordinary indented text and cannot tell a margin from content. The width is DECLARED rather than derived. `capabilities.transcriptGutter` on the CLI registry is a bounded integer; claude and codex each declare 2, measured on live panes, and no other stock entry declares any, so a CLI whose transcript layout nobody has measured is never touched. The server publishes the map as `window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the capability rather than by listing ids, and `_activeCliGutterColumns()` looks the active session's mode up in it. The copy path reads no terminal buffer at all. The declared width is a CEILING, not the answer: `clean()` strips the lesser of it and the run every selected line shares. A block can therefore only shift as a unit, the structure inside a selection survives by construction, and a selection reaching column 0 loses nothing. That is what keeps a `git log` body at its own four-space indent inside an agent's two-column gutter. Codex was measured separately, because it renders nothing like Claude: it draws boxes narrower than the pane and pushes its transcript into ordinary scrollback. On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose continuations sit at 2, and a nested YAML block the model wrote rendered at 2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact. Two derived versions were built and measured first, and both are recorded in the code because both looked correct: - Painted trailing padding — a full-screen TUI writes real spaces across the unused part of a row, a shell leaves them never-written for xterm to trim — has no false positives and never over-stripped. It is also a function of pane WIDTH: the padding exists only while a rendered line stops short of the CLI's own layout width, and Claude's prose wraps to fill it. Dragging the same two prose rows of one live transcript at five window sizes, the share of padded rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the strip silently did nothing at every ordinary size while a corpus captured entirely at 282 columns said it worked. - Taking the narrowest indent on the rows around the selection fires at every width and over-strips about 1% of selections, because a file listing inside the transcript can be the narrowest thing on screen. Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235 and 282 columns — the declared width over-strips none, breaks no relative indent and alters no text, and serves 100% of the selections whose own indent covers the gutter. Verified end to end in a browser with a real mouse drag and a real Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell pane is untouched at every one. The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard), per-device and default ON: a display key, absent from the .strict() SettingsUpdateSchema, read as `!== false` because the desktop branch of getDefaultSettings() returns {}. The toggle is checked before the map. Two review findings from #451, handled: - The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the first selected line, the one whose margin the mousedown genuinely cut off, so the same three rows no longer produce three different clipboard results. - The reversed-drag finding does not reproduce on the pinned xterm. `getSelectionPosition()` reads `_selectionService.selectionStart`, whose getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair when `areSelectionValuesReversed()` says so. A real upward mouse drag through chromium against xterm 6.0 reports the same range as the downward drag. `_normalisedSelectionRange()` keeps the ordering as a guard, because the model one layer down exposes the unnormalised fields under the same two names. Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected script stripped in test/server-index-title.test.ts. Every guard is pinned: removing any one of seven reds at least one test, including declaring the wrong gutter width. Full suite green, 7,861 passed, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e587d84590 |
fix(terminal): make the failed-load notice fit the narrowest terminal
A third pass in a real browser, at the widths this app actually renders at. The notice a failed history load writes into the blanked pane was one 70-character sentence. At 430px that exactly filled the line; at 320px it wrapped and left a lone '.' on a line of its own. The floor this app will render at is 40 columns — reachable today by raising the font on a phone — so the notice is three lines now, none over 25 columns, one fact each: what failed, that the session is still alive, and what to do. It says RELOAD rather than "reopen the tab" because `selectSession` early-returns when the session is already active, so clicking the tab you are already on retries nothing. The earlier wording named no next step at all, which left a mostly-empty terminal and no way out of it. CLAUDE.md no longer cites "758px reachable to the right" as evidence: that figure is a property of the test content, not of the fix, and the file's value is that a reader can trust a claim without re-deriving it. What is pinned instead is the invariant that survives any content — the full pane width is reachable, and removing the class returns scrollLeft to 0, so a resolved mismatch cannot leave the pane parked off-screen. Verified at 430, 360 and 320px against the shipped bundle, with the test asserting the built asset carries the copy so an edit that never reached the build fails rather than passing on the source's wording. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d9fa9ba1eb |
test(terminal): follow the existing suites to the one geometry owner
The gate caught fourteen failures the focused tests could not: every harness that builds a partial app out of cherry-picked mixin methods, and every source guard that named `fitAddon.fit()` by hand. Most are wiring — `syncTerminalGeometry`, `_refitAfterCellSizeChange` and `_resizeTerminalTo` added to the fakes so the real chain runs rather than a stub of it. `file-browser-search` is the one that shows why it matters: without the method on the fake, selectSession's unconditional call threw into its own catch and every later assertion in the file measured a load that never happened. Two are not wiring. `detached-session-pane-sizing` pinned the behaviour this change deliberately reverses. It asserted the LOCAL fit still runs for a session owned by its own window — "withhold the send, never the reflow" — so the assertion is restated rather than patched, with the reason beside it and in the file's docblock: a reflow the PTY is never told about leaves this xterm rendering a CLI's frames against a shape that does not exist, and the popup that owns the pane is drawing for its own width regardless. The old rule bought a garbled frame, not a correct one. `mobile-prompt-composer` sliced `_cleanupSessionData` as a fixed 1200-character window, so the assertion depended on how much unrelated code sat above the line it cared about. It reads the whole method now. `terminal-scroll-intent` records `syncTerminalGeometry` rather than `fit`, under its own name: recording a bare fit there would name the very thing the subject was changed to stop doing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
abf1d1f1ca |
fix(terminal): the PTY and the browser terminal must never disagree about size
Issue #464, "text gets muffled sometimes, in both TUI default and fullscreen". The screenshot is not a dropped frame or a frozen renderer — it is arithmetic. Claude Code's TUI wraps its frame at the width the PTY reported and erases the previous frame by walking the cursor up the rows it believes that frame took. A browser terminal of a different width makes each logical line occupy more physical rows than Ink counted, so `eraseLines(n)` clears too few and the new frame paints over rows nothing erased: doubled lines, and short tool summaries sitting inside longer prose rows with the prose's tail still visible. Reproduced against this repo's own xterm before changing anything — a 120-column PTY against a 62-column terminal renders every wrapped line twice. `test/ terminal-pty-geometry.test.ts` pins that, and pins the clean render at matching widths beside it, so the assertion cannot be satisfied by code that fixes nothing. Four ways the two drifted apart, none of them observable from either end: 1. `fitAddon.fit()` resizes xterm to `proposeDimensions()` RAW while every server-facing path reported those floored at 40x10. Measured in Chrome at 430px: font size 44 proposed 13 columns, the server was told 40, and xterm stayed at 13. Three call sites each did their own fit-then-floor, and two re-read the proposal after the fit — `_shrinkPaddingToFit()` runs exactly there, so the container had moved. 2. `throttledResize` (keyboard up) and `sendResize` (session detached into its own window) reflowed locally and withheld only the SIGWINCH. That is the one combination that cannot be right: a reflow nothing is rendering for buys nothing and costs correctness. Both now withhold everything, and the keyboard's settle timer still sends the one resize that stops the PTY going stale. 3. `setFontSize`/`setFontFamily`/`setFontWeight` move the cell size — a geometry change — and told the server nothing at all, so raising the font on a phone left the CLI wrapping at the old column count. 4. `Session.resize` DECLINES a small-viewport request while a desktop connection holds an active sizing claim, and said nothing, because resize was write-only. `syncTerminalGeometry()` is now the one function that may change the terminal's size: it fits, floors and applies as a single step, so the numbers xterm holds are the numbers the server is told. A test sweeps every module for a bare `fit()` on the main terminal, and finds exactly one — the owner's own. For (4) the client cannot win, so it is told the truth instead: both transports answer a resize with `session.ptyCols`/`ptyRows` (`{"t":"zc"}` on the socket, the body of the resize POST) and `_onPtyGeometryReport` adopts them. A terminal that keeps a shape the PTY refused does not render "too narrow", it renders garbled. Adopting can leave the pane wider than the screen and the container is `overflow: hidden`, so `.pty-oversized` grants horizontal reach for exactly as long as the mismatch lasts: correct-and-reachable beats correct-and-clipped beats garbled. That rule sets both overflow axes and its own `touch-action` because mobile.css loads later and sets `.terminal-container { overflow: visible; touch-action: none }` — a bare `overflow-x` would leave overflow-y computing to `auto` and hand the browser a vertical scroll container the terminal's touch handler knows nothing about. Verified in Chrome at 430px against a live server, with a desktop client holding the claim: the phone adopts 198x43, gets `overflow-x: auto` / `overflow-y: hidden` / `touch-action: pan-x`, 758px of reach to the right, and keeps its own vertical scrolling. The pre-fix build was measured in the same harness for the control. Two things this deliberately does not do. It does not change who owns the pane size — the desktop still wins, and `_startMobileResizeRetry` still takes it back once that goes idle. And `throttledResize` still holds the PTY's shape for the whole keyboard animation rather than sending a SIGWINCH per step; that decision predates this and was not re-tested here. Also in this commit, Ark0N's third-pass review items on #431: - The response viewer's byte-buffer fallback and `_onSessionClearTerminal` both used the no-param `/terminal` form, capped only by `terminalBufferMaxBytes` (32MB) — the largest body the frontend asks for anywhere. One carried no deadline at all and the other got the 15s tail budget. Both now take the full-history budget. - A `?full=1` capture that outruns its deadline falls back to the bounded tail. The pane is blanked before that fetch, so an abort used to leave a black rectangle, discard the queued live output and never reach `_connectWs`. A failed load now still opens the socket, says one dim line where the content would have been, and clears the tab's spinner — which nothing did, so a failed select left `aria-busy="true"` set forever. - `_wsOutputGapSession` is cleared at the repaint that settles it, not in a `finally` that also ran on the catch. A reconcile that threw, or hit the new deadline — the flaky link the marker exists for — dropped the gap with nothing to retry it. `ws.onopen` no longer clears it up front either. - The replay-clear invariant is pinned in the gate, which is the drift this PR exists to fix: `_resetTerminalForReplay` must be a queued write and nothing else, and no module may blank the terminal with a `clear()+reset()` pair. - `DIAG_ENTRY_MAX_CHARS` replaces the hardcoded 300, bound through a local first: `CodemanDiag?.x` still throws a ReferenceError when the identifier was never declared, and that is the one function in the app that must not throw. - panels-ui's two kill-all clears route through the same helper, and the xterm-version guard's comment says "resolved lockfile version" rather than "dependency RANGE", which is what it has pinned since the last round. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
90fd0a5a15 |
fix(tmux): gate the pane-exit read, and mute the dot on the rich rail too
Four changes the maintainer asked for on Ark0N/Codeman#446 before merging. The pane-exit watcher stays always-on, but a tick now costs nothing when there is nothing to observe. `hasObservablePaneSession()` skips the tmux exec while every session on the manager is one of the shapes `Session.paneExitApplies` already forces to UNKNOWN: a remote SSH session (its local pane holds the ssh client), a docker case (a `docker exec` into the container's own tmux), and a record rebuilt from the socket (no provenance at all). The timer is untouched. Skipping retracts nothing, for the same reason a failed read does not: the map still holds the last real reading, and every path that puts a new command in a pane calls `clearPaneExit()` itself. The two copies of that rule are pinned against each other in `test/session-pane-exit.test.ts`, because drift between them is silent in both directions. `DEFAULT_PANE_EXIT_INTERVAL_MS` was already a constant beside the stats and remote-reconnect intervals; its comment now says why the watcher owns its own cadence and why the number is what it is. The never-default-an-absent-status rule is written where `PaneExit` is declared. It names `status ?? 0` as the thing never to write, and says that an agent the OOM killer took would otherwise read as a user typing `/exit` — which is what absent-stays-absent keeps a later clean-exit sweep away from. Nothing fails when somebody adds that `??`, which is why the sentence is there rather than a test. Checking the dot's specificity found a second fight, and it was losing. On the tab strip the alert rules win as intended: a session that exits with a permission dialog pending still renders red, and yellow for an idle alert. On the rich vertical tab rail they did not — that rail's own `tab-state-*` dot rules are (0,9,1) against the strip's mute at (0,5,0), so an exited session there kept a full green dot AND the working halo beside a badge reading "exited". The rail twin matches that specificity exactly and therefore must stay below those rules in source order; it clears the halo as well, which the strip's rule never had to think about. `test/session-pane-exit-ui.test.ts` now resolves the real stylesheet in jsdom rather than matching selector text: postcss collects every rule that paints `.tab-status`, a real engine decides, and the tests read back the answer. Two mutations were run against it to prove it has teeth — dropping the hand-written alert exclusions fails three cases, and moving the rail twin above the state rules fails one. Refs Ark0N/Codeman#446. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
64c288a683 |
feat(codex): read Codex's own background-terminal row
Codex states background work too, and it says so in a different place. Claude writes `· 1 monitor ·` on the last row of the screen; Codex pins `1 background terminal running · /ps to view · /stop to close` ABOVE its composer, which puts that row third from the bottom once the status line and the composer are counted. So how far up the screen to look is now per-CLI data as well: `capabilities.workDetect.watchingLines`, bounded to 1..8 by the schema, and defaulting to Claude's two. That bound is the point. The window is half the injection guard, since every row it adds is another row the agent itself may be able to write, and the label is what silences an idle alert. The other half is the anchor, and Codex's is ` · /ps to view`: chrome naming a slash command only the CLI can offer, so a session that writes "I left 1 background terminal running for you" into its own output matches nothing. Measured against a live codex-cli 0.154.0 pane rather than read out of a binary. The row appears when the terminal starts, follows the composer down as the conversation grows, and is gone after `/stop`. Verified end to end on an isolated beta: the session payload carried `watching: "1 background terminal"` and the badge rendered with it, and both cleared when the terminal stopped. The fixtures in the tests are that capture verbatim. Codex has no hook signals, so no idle prompt and no false NEEDS YOU row: for a Codex session this is the badge alone, which is the case the maintainer said a registry field could cover and a hook never could. Cross-CLI tests pin that neither pattern fires on the other's screen, and that a CLI declaring nothing still reports nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
abd39318e6 |
fix(terminal): deadline must cover the body, precache must ignore the cache-bust query
Review fixes. Two of these are defects in the previous commit.
1. The fetch deadline only covered time-to-headers. `await fetch()` settles on
response headers, so clearing the abort timer in a finally around it left the
body — the multi-megabyte `?full=1` capture the deadline exists for —
completely unbounded; it only ever bounded a server that accepts a connection
and never replies. Measured against a server that sends headers immediately
and stalls the body 4s under a 1s deadline: fetch resolved at 30ms, timer
cleared there, body completed at 4026ms unaborted. Now the body is read
inside `_fetchTerminalCapture`, which returns {json, headers, headersAt} —
headers because two callers read server-timing, headersAt because those same
callers measure header-vs-body time and can no longer observe that moment.
`_terminalCaptureInflight` is scoped the same way, so a body still streaming
counts toward a capture starting beside it. Same test now aborts at 1005ms.
2. The precache could never be hit, and the previous commit made that expensive
rather than free. `renderIndexHtml` runs `cacheBustAssets`, which appends
`?v=<mtime>` to every same-origin .js/.css reference INCLUDING content-hashed
names — confirmed against a running instance:
`vendor/xterm-zerolag-input.6fee72f2.js?v=1789402869101`. `caches.match` is
query-sensitive, so entries keyed on the bare hashed path were unreachable;
deriving the list from the manifest turned cheap 404s into ~1.3MB downloaded
at every install that nothing could read back, once per deploy now that
CACHE_NAME rotates. The fallback match takes `{ ignoreSearch: true }`, which
also lets runtime-cached entries survive an mtime change.
3. `_wsOutputGapSession` was only cleared in ws.onopen, so paths that already
repaint the buffer left it set and the socket replayed everything a second
time. `selectSession` loads the buffer and only THEN calls `_connectWs`, so
neither the _isLoadingBuffer nor the _terminalRefreshOwner guard applied.
`_markTerminalBufferReconciled()` is now called from _onSessionNeedsRefresh's
finally, from selectSession after its load, and from _cleanupSessionData.
The scope claim was also wrong and is corrected in the comment: when the
network drops, SSE drops with it and handleInit's keepTerminal branch already
reconciles. The genuinely uncovered case is the WS dying while SSE stays up,
where _onSSETerminal discards SSE terminal frames until _wsReady flips in
onclose — up to the ping+pong window of output nothing writes.
4. CLAUDE.md said "all of them measured rather than reasoned", which the PR's
own "not verified" section contradicted. Split explicitly: the replay race is
measured, the watchdog mechanism is verified against xterm 6.0.0 under jsdom
(field path resolves, a forced stale handle makes refreshRows a no-op, the
kick schedules a fresh frame), and the iOS rAF-discard premise is reasoned
and still wants a device. Adds the two missing entries — the WebSocket
reconcile and the sw.js/build.mjs "keep these in sync or the build throws"
contract.
Also: test/xterm-private-api.test.ts pins the RESOLVED lockfile version instead
of the declared `^6.0.0` range, which was the wrong assertion in both directions
— a real upgrade to 6.4.0 can rename a private field while resolving inside the
range, and an innocuous range edit failed while changing nothing installed. And
test/sw-precache-manifest.test.ts now parses HASHABLE out of scripts/build.mjs
rather than hand-copying it, which was the same drift this PR exists to fix; the
parse is guarded against silently matching nothing.
The deadline fix has a behavioural test against a real socket plus a source
guard asserting `await res.json()` precedes the finally — verified to fail when
the helper is reverted to the old shape, so it is not vacuous.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
c0422c4e21 |
feat(terminal): renderer watchdog, atomic replay clear, fetch deadlines, reconnect recovery
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.
1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
when a PWA backgrounds, and xterm's RenderDebouncer only clears its
`_animationFrame` handle from inside that callback — so one drop leaves it
permanently set and every later refresh() early-returns. Parsing is
decoupled from rendering, so bytes keep filling the buffer correctly while
nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
single backgrounding wedges it until a reload. Adds a 2s liveness poll and
`_kickRenderer()`, which does what the dropped `_innerRefresh` would have.
2. Replay clears raced live output. xterm's write() is async-queued while
reset() is synchronous and, per upstream, "does not clear input buffers and
does not reset the parser" — so bytes queued before a reset are parsed after
it and fuse into the snapshot. Verified against the real xterm 6 here:
write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
path was already safe via a queued erase; the needsRefresh and clearTerminal
paths were not. All three now share one queued `\x1bc` (RIS), which unlike
3J/H/2J also resets modes, charsets, scroll regions and SGR state.
3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
and flushes queued input, and needsRefresh only fires on external-CLI
startup and SSE backpressure drain — never on reconnect. Output produced
while offline was simply absent afterwards. Interim fix: reaching onclose
means the drop was unintentional, so the session is marked and the next open
reconciles from the server buffer. Sequencing output is the follow-up.
4. Terminal captures had no deadline. No AbortController anywhere in the
frontend, including `?full=1`, which the code itself calls "unbounded-ish
work: at the default history limit it can be megabytes". Adds a budget that
scales with full-vs-tail and with captures in flight, degrading to a plain
fetch where AbortController is missing.
Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.
The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.
Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.
Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
74884a20eb |
feat(approvals): let a watching session keep quiet, and fix the tui gate
The badge alone left the row in NEEDS YOU, which is the thing the issue was about. The fix is the alert that does not fire. An idle prompt from a session that is watching its own background work now opens ALREADY acknowledged. `hook-event-routes` passes `Session.watching` to `notePrompt()`, which sets `acknowledgedAt` and records why in a new `acknowledgedReason`. Nothing new suppresses anything: `acknowledge()` has always meant "the alert this prompt armed is spent", and the prompt itself stays pending, answerable and available as Read My Mind context. A wrong label therefore costs a card that does not blink, never an alert that was never created. Every surface follows from that. The broadcast carries the reason, so a live page declines to arm the tab alert and raises no desktop notification. The push is skipped, since a false alarm is hardest to ignore on a phone. A reloading page reads `acknowledgedAt` in `seedApprovals()`, which it already did. And `classifySession()` now reads it too, which is a pre-existing bug fixed here: acknowledging on one device cleared the alert everywhere except `codeman tui`. It re-arms for free, because the next idle prompt supersedes the item and is built fresh. Only `idle` is eligible, so a dialog that blocks the agent still goes red whatever else it started. The label is pane-derived and therefore prompt-injectable, so it is now read from the last two rows of the screen only, with Claude's pattern anchored on the `·` its footer joins items with, ANSI-stripped and length-capped at the source. An agent that prints `· 1 monitor ·` into its own output finds no match. Verified on an isolated beta: a session that armed a monitor took its idle prompt acknowledged with no alert on any surface, wore the badge, and showed "quiet, watching 1 monitor" on its still-answerable card; the same session with the monitor killed alerted normally on the next prompt. `test/watching-no-alert.test.ts` pins both directions across all four surfaces. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1cb0441bd8 |
fix(session): degrade the resume pin to the session id, not to nothing
A single pin that failed its transcript gate returned the options untouched, so `resumeSessionId` fell back to `_resumeSessionId` — undefined for an ordinary session — and the renderer emitted the bare `claude --dangerously-skip-permissions --session-id "<this.id>"`. Every session prompted before its first `/clear` owns a transcript under that id, so the dropped pin handed back exactly the refusal this branch removes, with no `||` branch to catch it. It was also a regression against master on the `restartCli()` path, which pinned `_claudeSessionId ?? this.id` and, since the constructor seeds that field, could never land unpinned. The pin now walks three candidates in priority order — the conversation chain's tail, the launch seed, then the session's own id — and takes the first one a transcript backs. A candidate that misses is passed over rather than ending the walk. Falling off the end pins nothing, which also settles the second half of the problem: the old code skipped the transcript check whenever the pin was the session's own id, so a genuinely new pane rendered the two-branch form after all. That costs a brand-new session claude's "No conversation found" line in its scrollback, and `wrapWithNice()` prefixes only the first branch of the rendered `a || b`, so the branch that actually runs loses its priority for the life of the session. With no transcript anywhere the bare `--session-id` is the correct command, so the comment claiming an unchanged shape is now true. The transcript lookup reads the server process's own `CLAUDE_CONFIG_DIR` when a session declares none. A pane inherits the server environment through tmux, so on an install that exports it the CLI writes its transcripts there and every lookup under `~/.claude` was a false negative — which under the old code meant the colliding command. `claudeCredentialsPath()` and `realClaudeConfigDir()` resolve the same directory the same way. The header sentence calling a skipped resume "the safe direction" described the opposite of what happens at this call site, and says so now. The create-path fallback writes `_resumeSessionId` alongside the create options. That branch leaves `isRestored` false, so `_claudeSessionId` is recomputed from the launch fields and settled on `this.id` while the CLI resumed the chain tail; the response viewer, Read My Mind and the unified-list alias map read that field until the next first-hand hook. Four new tests: a chain tail with no transcript while the session id has one, no transcript anywhere, the create path's alias, and the process-env lookup. All four fail against the previous commit. Two existing tests move with the gate — the guess-refusal test now backs the session's own id, and the custom-model restart test gives its working pane the transcript that makes `--session-id` collide in the first place, alongside a new one pinning the no-transcript case. CLAUDE.md described the pin as a `restartCli()`-only thing sourced from the live conversation id. All three halves of that moved here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3f2cde2db7 |
feat(session): say when a session is watching its own background work
An agent that arms a monitor, backgrounds a shell or hands a task to a cloud session is told to end its turn. The pane then falls quiet, Claude Code's idle_prompt notification arrives a minute later, and every surface files the session under NEEDS YOU with nothing for a human to answer. Claude states what it is still running on the last row of its screen (`⏵⏵ bypass permissions on · 1 monitor · ← for agents`). That row is now `capabilities.workDetect.watchingLine` in the CLI registry, guarded by compileVersionRegex() like every other config regex, and the idle probe reads it off the capture it already takes: `watchingLabel()` in session-activity.ts searches the last five lines only, so a session that PRINTS "1 monitor" is not mistaken for one running it. The label lands on Session.watching and rides toLightDetailedState() out to every surface. The phone overview, the desktop home rail and the rich sidebar rows wear it as a `watching` badge in the accent colour, beside the state pill and never in place of it: an agent can arm a monitor and ask a question in the same breath, and only the pill says which. Verified end to end against a throwaway session on an isolated beta instance: the payload carried `watching: "1 monitor"` once the turn ended, the badge rendered next to a yellow `waiting` pill, and both cleared when the monitor died. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
47e7935274 |
fix(session): resume the conversation when respawning a dead pane
A CLI that launches with `--session-id <id>` refuses an id that is already in use (claude: `Error: Session ID ... is already in use.`), and every session whose agent has been prompted owns a transcript under that id. The dead-pane respawn in `_setupOrAttachMuxSession()` passed the bare launch line, so recovering such a session relaunched a CLI that died on startup, the pane went dead again at once, and the conversation was stranded behind a tab that looked merely idle. `restartCli()` has pinned a resume id against this since the custom-model work, and its comment states the assumption that made the other path look safe: "Unlike the dead-pane respawn, this one kills a WORKING pane whose conversation already has a transcript". A pane whose agent exited has a transcript too. Both relaunch paths now build options through `_buildRespawnPaneOptionsWithResumePin()`, and so does the create-path fallback after a failed respawn, which otherwise met the same refusal that made it the fallback. Four gates guard the pin, each standing for a way of resuming the WRONG conversation or of making a working relaunch fail. A remote or docker session is never pinned. Unlike `restartCli()`, whose route refuses both, the dead-pane respawn is reached by every session shape. Their pane commands already render a self-healing `--session-id || --resume`, and both flip to resume-first once the resume id differs; the conversation lives on the far side, so a local id resolves to nothing there and the `--session-id` fallback then collides with the transcript the far side does hold. The id comes from the conversation CHAIN rather than `_claudeSessionId`, which also holds history-correlated guesses keyed on the working directory. `_recordClaudeSessionInChain()` refuses those so they cannot "write a foreign conversation into this pane's permanent record", and launching from one is worse than the display bug that rule prevents. The chain tail also outranks the launch seed, which is written once at construction and never moves off a `/clear`. A pin no transcript backs is dropped, because the fallback branch keeps `--session-id <this.id>` and would collide. A synthetic `restored-<fragment>` id from socket discovery is dropped too, and logged: it fails claude's `uuid` token pattern, so the renderer would emit the unpinned command while the caller believed otherwise. Tests cover each gate and the rendered command. Four of them fail against the unfixed source; the remote and docker ones were separately checked against a build with only that guard removed, since they pass on master for the wrong reason. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a2dcc91ddf |
fix(docker): add and correct the cross-checkout collision guard for Update-Codeman.sh
Same guard as Start-Codeman.sh's own (docs/docker-self-update.md-adjacent incident, 2026-09-21): docker-compose.yaml hard-codes `name: codeman`, so a second checkout run without COMPOSE_PROJECT_NAME resolves to the SAME Compose project as any other checkout on the host. It has to live here too, not just in Start-Codeman.sh: this script's own --no-cache build and its `down`/`down --volumes` both run BEFORE the handoff at the bottom of the file, so Start-Codeman.sh's copy of the guard would only fire after this script's own destructive calls already ran — and its default `down --volumes` is more destructive than Start-Codeman.sh's own targeted refresh, clearing every named volume the resolved project has. Also fixes a real bug the same guard shipped with: under `set -o pipefail`, `grep -v` legitimately exits 1 when nothing survives the filter (the ordinary, no-collision case), and without `|| true` on the pipeline that non-zero status propagates through the command substitution and `set -e` aborts the WHOLE script at the guard — every time, collision or not. Caught only by actually executing the guard end-to-end against a stub `docker` (the existing smoke-test harness), never by a static text/regex check on the source; the stub's `config --format json` response was also fixed to pretty-print like real Compose does, since a compact one-liner silently resolved project_name to empty and exercised neither script's guard the way production output does. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n |
||
|
|
c67c130caa |
feat(web): mark a session tab whose agent has exited
The tab now reads "exited (137)" beside the session name, drawn from the
`paneExit` field the server publishes. `applyPaneExitBadge()` owns the DOM
work, called from the incremental render path — the only path a live session
ever takes, since going from live to exited adds and removes no tab and so
never reaches the full rebuild.
An unknown answer draws nothing. A death tmux could not explain reads "exited"
with no number rather than "exited (0)", so an unexplained death and a clean
exit do not look alike. A signal death reads "exited (signal 9)".
The badge carries `data-i18n-skip`, like the status pills: it is generated
text, `i18n.js` walks inserted content, and a dictionary entry added later
would fight the renderer, whose in-place comparison is against English.
The tab also carries a `tab-agent-exited` class that mutes the status dot. That
dot is drawn from `status`, which stays `idle` or `busy` for an exited pane as
the issue requires, so without this a green or pulsing dot sits beside a badge
saying the agent is gone — the first thing a tester asked about. `status`
itself is untouched, so this is a rendering rule only. The CSS excludes the two
alert classes by hand, following the convention the rich-rail dot rules
document: a dot turning red or yellow because a session is blocked on a human
outranks "the agent exited".
The tab keeps its click behavior. X still closes it, and nothing here closes,
sweeps or restarts anything.
`docs/architecture-invariants.md` gains the mechanism under "Session data and
lifecycle", where every comparable one already lives: what the tri-state means,
the four shapes it is absent for, why the watcher cannot ride the stats
collector, why an absent `#{pane_dead_status}` is not 0, and the three things
that must never happen to an exited pane.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|