mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
0a52a99ca918cf8341446672eaa22b2f655fd413
140
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a52a99ca9 |
feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (
|
||
|
|
0af925fe82 |
Merge remote-tracking branch 'origin/master' into land/1.32.1
# Conflicts: # CLAUDE.md # docs/architecture-invariants.md |
||
|
|
0462a5d5a0 |
fix(approvals): merge-time fixes for the watching badge (#473)
- session.ts: a pane capture that fails now CLEARS the watching label (and emits watchingChanged so pages drop the badge) instead of keeping the last one, so a failed capture degrades toward an alert rather than pre-acknowledging the next real idle prompt. Test updated; invariant noted in architecture-invariants. - approvals-ui.js: the header bell counts only unacknowledged items (pendingApprovalsCount), matching codeman tui's pendingApprovalCount(); pinned in watching-no-alert.test.ts. - mobile-overview.js: move the orphaned "Pill copy per state" JSDoc back onto MOBILE_OVERVIEW_PILL_LABEL. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
da6fa663e7 |
fix(terminal): merge-time fixes for the copy gutter strip (#469)
- stock.ts: claude is no longer the only entry declaring transcriptGutter; codex declares it too. - architecture-invariants: the strip applies when the session's CLI declares a margin (not detection), and a note that it keys on the session's launch mode, not on what is running in the pane (a claude pane dropped to a shell still loses up to two columns; copyStripMargin is the escape hatch). - render-index-html test: the gutter map is injected for a solo /session/:id render as well. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
2afb1c2c2e |
docs: trim CLAUDE.md from 265 KB to 142 KB, detail moved to architecture-invariants
CLAUDE.md loads into every session, and its Architecture section had grown feature write-ups (history, measurements, rationale) that belong in docs/architecture-invariants.md per the file's own header. Each long block now keeps what the feature is, where it lives, its setting/default and the rules that prevent real bugs, and links to its invariants section. Everything removed was moved there: 29 new sections, extra facts appended to the existing ones. Also: hard-coded counts (SSE events, route handlers, module/file counts, device profiles) replaced by pointers to the source of truth, and the Debugging commands fixed to use the codeman tmux socket and HTTPS for prod. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> |
||
|
|
de4b1db490 |
Merge pull request #431 from rounakdatta/feat/mobile-terminal-resilience
fix(terminal): four silent-failure paths — renderer freeze, replay race, reconnect gap, unbounded fetches |
||
|
|
8536aaef7b |
Merge pull request #473 from irisitymichaelgrundberg/feat/session-watching-badge
feat(approvals): let a session watching its own background work keep quiet (#468) # Conflicts: # src/config/cli-registry/stock.ts |
||
|
|
94b093b617 |
Merge pull request #469 from irisitymichaelgrundberg/feat/copy-dedent-pane-margin
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares |
||
|
|
6a01412af9 |
Merge pull request #466 from irisitymichaelgrundberg/feat/pane-exit-reporting
feat(tmux): report that a pane's agent has exited (#446, part 1) |
||
|
|
02e40f506b |
fix(docker): gate gh/az seeding on its switch; no shared git sign-in for non-admin clones
Addresses the review on #472. - CRED_STORES: `.config/gh` and `.azure` now carry `enabledByEnv` (CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ), and resolveDockerCredentialArtifacts skips a store unless that variable is exactly `1`, read at container create. A host that merely has ~/.config/gh/hosts.yml or a plaintext MSAL cache no longer copies them into every case container. Tests: the default environment seeds neither even with the files present, and each store follows only its own switch. - Multi-user mode: a non-admin's Clone Repo clone and preflight run with `git -c credential.helper=` (GIT_NO_CREDENTIAL_HELPERS, placed before the subcommand), so the server account's helpers are never lent to them. Verified against a real private repo that it also clears the URL-scoped credential.<url>.helper entries, and that public clones still work. Tests: the argv in test/git-clone.test.ts, and the route decision (non-admin cleared; admin and single-user kept) in test/routes/case-clone-credential-helpers.test.ts. - Docs: recreate the case container to pick up seeds (docker/README.md, Docker-Cases wiki, docker-cases.md); the multi-user behaviour in docker/README.md and security-architecture.md; "functionally unchanged" instead of "unchanged" for an image built with both switches off (server.Dockerfile comment, README, changeset). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw |
||
|
|
9a48c43aa1 |
docs(watching): a restart is not a gap, and here is the measurement
Claimed after a manual test that a session comes back from a server restart without its badge until it next produces output. Measured instead of assumed, and it is wrong: a codex session with a background terminal still running had its label back within about 20 seconds of the restart, with no input from anyone. Reconciliation re-attaches the pane, the attach repaint carries the composer glyph, the idle confirmation arms on it, and the probe re-reads the label — the ordinary path, doing the ordinary thing. What produced the false claim was a session whose monitor had simply expired while it sat there. Its footer carries no chip, so `watching: null` was the right answer and there was nothing missing to restore. Recorded at the field and in the invariants, because the shape of this invites exactly one wrong fix: a polling timer to keep a value fresh that the pane already refreshes by itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5cf5a45438 |
feat(docker): opt-in gh + az CLIs with git credential helpers for private repos
Add Case -> Clone Repo could only reach public repositories in the Docker deployment. This lets a deployment opt in to the GitHub CLI and the Azure CLI (+ azure-devops extension) as git credential helpers. Codeman itself still collects no credentials. - server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH / CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the build). Off leaves no apt repository, package, extension, helper script or credential entry, so a default build is unchanged. On installs from the vendors' apt repositories and configures system gitconfig helpers: github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com / *.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT). A helper whose CLI is not signed in prints nothing, so a private clone still fails fast. - The extension lives in AZURE_EXTENSION_DIR outside HOME (/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0 group-writable in the agent image). - Hosts turn them on in docker-compose.override.yml: `build: args:` for the server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the agent image. build-agent-image.mjs and the in-app auto-build share one env -> ARG table (pinned by the parity test) and pass nothing when unset. docker-compose.yaml is untouched; .env.example only gains a comment, so the self-updater's environment gate sees no new keys. - Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and the az sign-in files from ~/.azure per file, read-only, like pi/grok. - The Clone Repo AUTH_REQUIRED message says how to sign the server's git in instead of claiming private repositories cannot be cloned. - Docs: docker/README.md "Private repositories", docker-compose.md, docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki pages, security-architecture.md, architecture-invariants.md, changeset. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw |
||
|
|
ac6236b268 |
fix(terminal): clean a copy once, and reach every pane that copies
Review fixes for #469. The Ctrl+C branch cleaned the selection to decide whether to copy and then passed that cleaned string to copyTerminalSelection(), which cleans again. The trailing trim is a fixed point, so that was safe until this PR; the margin strip is not, because it takes the lesser of the declared width and the run every line shares, so a second pass takes up to `margin` columns more. The branch now gates on the cleaned string and hands the raw one on. Verified in chromium with a real drag, a real Ctrl+C and a real clipboard read on a live claude pane: an on-screen ` fix(terminal): trim it` reaches the clipboard as ` fix(terminal): trim it`, and reverting the branch reproduces the reported ` fix(terminal): trim it`. Pane B of a split resolves its own width. `_cliGutterColumns()` and `_normalisedSelectionRange()` take the session and the terminal to read, defaulting to the primary pane's, so Pane B looks its own run mode up instead of keeping a margin Pane A drops on the same keystroke. Verified live with two claude panes open side by side. A detached session window (`/session/:id`) receives the gutter map. The injection sat inside the block that skips the run menu's payloads for a solo window, so the toggle worked in the main window and did nothing in the popup on the same device. It needs no availability probe, so it moved below that block and the solo window still carries none of the payloads it skipped before. The settings description said the width is measured and named Codex as exempt. Nothing is measured, and Codex is one of the two panes that are stripped. docs/wiki/Settings-Reference.md gains the row every Terminal and Input toggle carries. CLAUDE.md no longer says the clean touches trailing runs "and nothing else" one sentence before the leading-margin rule, and both it and docs/architecture-invariants.md record that the strip is not idempotent. Two round-trip tests run on a mode that declares a gutter, which the existing copyTerminalSelection cases could not, since they all use the harness default mode that declares none. The Ctrl+C branch itself is pinned at the source, because it lives inside initTerminal's attachCustomKeyEventHandler closure over a real xterm the vm harness cannot build. Both pins fail on the reintroduced bug. Gate: 7865 passed, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
05c788ce9d |
fix(watching): close the review findings on the label and its window
A dual review (Codex CLI and Claude's code-reviewer, same diff, same brief) found the trust boundary weaker than the comments around it claimed. Eleven findings, all applied. The two blockers were both about who can write the row the label is read from. Claude's window covered two rows, and the second one is the status line, whose command a session running with permissions bypassed can write into its own `.claude/settings.json` — so an agent could print `· 1 monitor ·` onto a row of its own and silence its own idle alert. The default window is one row now, which is the footer and nothing else, and the constant says why. Separately, the label reached `data-tab-meta-sig` unescaped while the row is installed with innerHTML, which is an injection sink for any config-supplied pattern whose capture group is permissive; it goes through escapeHtml() like every other untrusted string in that file. The Codex entry could not be fixed the same way, and now says so. Its row is third from the bottom only while a terminal runs; with none running that slot holds the last row of the transcript, so matching the complete row (with the `/stop to close` tail, window narrowed to three) raises the bar without closing it. What contains it is `hooks: 'none'`: no hook event from a codex session reaches notePrompt(), so a forged label costs a wrong badge and cannot quiet an alert. The registry comment, `docs/cli-registry.md` and the test all state that rather than claiming a guarantee the code does not have. Also from the review: the TUI header badge no longer counts an acknowledged item, which was the same gate the classifier fix already went through and was wrong for human acknowledgement too; the TUI approval card reads the quiet reason and drops to a new `info` tone instead of asking for a reply; the badge carries an aria-label, because the phone it was built for has no hover target; the schema refuses `watchingLines` without a `watchingLine`; and the pattern and its window are resolved together rather than one memoized and one not. Documentation moved with it. The mechanism now lives in `docs/architecture-invariants.md` with CLAUDE.md keeping the rule and a pointer, `docs/wiki/Notifications-And-Approvals.md` tells users why a session stopped buzzing, and both that page and the changeset name the limitation neither did before: a question asked in plain prose is not a dialog, so it is silenced along with the false alarms while background work runs. Verified live again after the narrowing, on an isolated beta: a Claude session reported `1 monitor` and took its idle prompt acknowledged, and a Codex session reported `1 background terminal` against the full-row anchor. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
9c286eeddf |
fix(session): persist an exit retraction, and let tests reach the watcher
Ten findings from a two-model review of this branch. Both reviewers cleared the
detection logic itself; everything here is a gap around it.
A route that starts a command in a pane now PERSISTS as well as broadcasts.
`/interactive` and `/shell` did neither before, and the pane-exit watcher cannot
cover for them: its next tick finds `paneExit` already cleared in memory,
reports no change and writes nothing, so `state.json` kept saying the agent had
exited for as long as the session stayed quiet. Nothing reads that record for a
decision yet, which is exactly why it had to be fixed now — part 2 is designed
to read it. The `clearPaneExitForNewPane()` docstring claimed its callers
already persisted; that claim was false for these two, and now says what the
caller owes instead.
The watcher's four guards were unreachable by any test. `refreshPaneExits()`
opened with `if (IS_TEST_MODE) return;`, so the read gate, the in-flight
suppression, the generation counter and the empty-read rule could each be
deleted with the whole suite green. The tmux call moves into `readPaneRows()`,
which a test subclass overrides — the shape `runRemoteReconnectTick` already
uses in this file for the same reason — and the test-mode gate moves with it, so
what a test cannot do is spawn a process rather than exercise the bookkeeping.
Each of the four guards now has a test that fails when it is deleted.
The muted status dot turned out to be a specificity fight on three surfaces, not
two. `.tab-status.error` was not excluded, so a session whose agent exited and
whose PTY-exit breaker then tripped lost its red dot to the mute — the state the
browser answers with a "restart it?" confirm, and a needs-you colour by the same
argument that protects the two alert classes. And mobile.css gives a `busy` dot
a 9px size and a green glow with `!important`, while `status` stays `busy` for a
pane whose agent died mid-turn, so a phone rendered a grey dot still wearing the
green halo beside a badge reading "exited". Both measured against the real
stylesheets, both now excluded, and the CSS test reads mobile.css too instead of
being structurally blind to half the problem.
Six comments said things that were not true. Two named the stats collector as
what replaces a restored reading, which is the opposite of the design. The
interval constant argued that 2000 ms keeps a read inside a tick, when the
5000 ms exec timeout means it cannot — which is why the in-flight guard exists.
`MuxSession.discovered` did not say the flag is permanent, though `saveSessions()`
serializes it. The empty-read docstring claimed a distinction that `|| true`
makes impossible. The invariants doc promised more than its drift test delivers.
And CLAUDE.md had no pointer at all, leaving its two hardest prohibitions
("never set `status: 'error'`", "never null the pid") only in the file it is
meant to route people to.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
e1e7dc5bd8 |
fix(terminal): Ark0N's read of the #464 geometry work
Five items, two of which he could only see by running it, plus six smaller ones. Taking the two blockers first, because both were wrong in ways the existing tests could not catch. **Adopting the PTY's rows put the CLI's input line off-screen.** A phone that took a desktop's 43 rows into a viewport with room for 18 painted an `.xterm-screen` far taller than its container; xterm's own viewport then had nothing to scroll, so the bottom of the frame sat below the container with no gesture able to reach it. Output visible, typing invisible, for as long as the desktop kept the claim hot. `reconcilePtyGeometry` adopts COLUMNS ONLY now: width is the axis Ink's wrap and `eraseLines` arithmetic depend on, and keeping the local row count keeps the composer at the bottom of a viewport that scrolls. Measured at his geometry — a 360x300 container against a 198x43 pane now keeps 13 rows, takes 198 columns, paints 202px into a 210px container, and the input line is inside the box. **`capture-geometry-retry.browser.test.ts` failed, and CI could not see it** because the file is in `BROWSER_TEST_GLOBS`. Its premise WAS the clamp — `getTerminalDimensions()` floored while `fitAddon.fit()` did not — which this work removes at the source, so it can never hold again at any viewport. The case survives on its own terms: a pane already drawing at the requested size must not be replayed. Its premise is now the #464 invariant itself, that the floored report and the terminal agree, which is a stronger guard because the clamp coming back fails it here rather than silently restoring the replay loop. The helper docblock that repeated the old premise is corrected too. **A session with no pane reported 120x40 and the client adopted it.** `resize()` writes `_ptyCols`/`_ptyRows` only when `ptyProcess` is set and nothing seeds them from the spawn geometry, so a dead-pane session still held the constructor defaults — clicking that tab resized the browser terminal to 120x40 and, on anything narrower, claimed another device owned the pane when none existed. `Session.ptyGeometry` returns null without a pane, the HTTP route answers `{}` and the socket sends no frame at all. The raw `ptyCols`/`ptyRows` getters are deleted rather than left available to be misused again. **The 40-column floor clipped the pane with nothing able to reach it.** The affordance keyed on a PTY mismatch, and the floor produces no mismatch — xterm and the PTY agree throughout, the terminal is simply wider than the box. It keys on what does not FIT now, MEASURED (`.xterm-screen` against the container, on the next frame, because the screen takes its width with the render) rather than derived from cell arithmetic. Measured at 360px: font 24 applies 40 columns and paints 560px, and all 200px of the overhang is reachable. `.pty-oversized` is renamed `.term-overflows-x`, because after this the old name describes only one of the two causes. **"Scroll sideways" did not work on touch for the sessions it targets.** `touch-action: pan-x` is cancelled before it starts by the `preventDefault()` `touchstart` calls on every 'content' tap. The terminal's own touchmove handler pans the container now, with the axis locked once per gesture so a diagonal cannot pan and scroll at once, and the CSS grants no `touch-action` at all — handing the browser a pan AS WELL would move the pane twice for one finger on the taps where that preventDefault does not run. Measured under real touch dispatch: a 140px swipe reaches `scrollLeft` 140 where it reached 0 before, the buffer does not move with it, and a vertical swipe still scrolls the scrollback. Three defects in the above, found while checking it rather than by being told: - `canPanHorizontally` first tested `scrollWidth > clientWidth` alone, which is true of a container that is not a scroller — a sideways swipe would have locked the axis, done nothing, AND suppressed the vertical scroll it should have been. Gated on the class as well. - The notice advised scrolling sideways whenever the PTY was wider, including when it still fitted and nothing scrolled. It is gated on measured overflow, and on a comparison against the width this container WOULD request rather than the one it currently holds — once adopted those are equal, so the second question answers itself false while the condition is still true. - `_syncTerminalOverflowAffordance` could throw out of `document.getElementById` before reaching its try block. It runs off every geometry change, so a cosmetic affordance could have taken the resize down with it. The smaller items: - `docs/architecture-invariants.md` no longer explains the equality guard as a clamp signature; it records what the clamp used to do and why it cannot any more. Edited by hand — that file is outside the Prettier glob, and letting Prettier near it rewrote eleven unrelated emphasis markers. - `throttledResize`'s HTTP fallback reads the reply. It is the path where a declined resize is least likely to be noticed, because no socket means no `{"t":"zc"}` frame either. - The changeset covers the whole release: the geometry work, the queued replay clear, the renderer watchdog, the body-covering fetch deadline, the WebSocket output-gap reconcile, the build-generated service-worker precache and per-build cache key, and the crash-trail hygiene. - `@xterm/headless` is declared in the root devDependencies instead of being reached through workspace hoisting. - The output-gap marker is cleared after any response arrives, not only when the capture was non-empty: a server that answers with an empty capture HAS reconciled us, and leaving the marker set refetched on every reconnect. - `e587d845`'s message claimed a test asserted the failed-load copy against the built asset. It did not — that assertion lived in a probe deleted with the other scratch scripts, so the claim was false when it was written. There is a real test now, and it reads the source rather than `dist/`, because `dist/` is not committed and a test that skips when it is absent would pass for the wrong reason in CI. `Session.ptyGeometry` gets behavioural coverage against the real class in `session-resize-arbitration.test.ts` rather than a source guard, including the contrast — a pane that does exist still reports, and still follows a resize — so "always null" would fail it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ce80b7a212 |
feat(terminal): take the transcript gutter off a copy, at the width the CLI declares
Copying a paragraph out of a Claude Code or Codex pane puts that pane's own two-column transcript gutter on the clipboard, so every pasted line arrives indented. #451 shipped the trailing half of the copy clean and left the leading half out, because deriving the width from the selection fires on 73% of ordinary indented text and cannot tell a margin from content. The width is DECLARED rather than derived. `capabilities.transcriptGutter` on the CLI registry is a bounded integer; claude and codex each declare 2, measured on live panes, and no other stock entry declares any, so a CLI whose transcript layout nobody has measured is never touched. The server publishes the map as `window.__codemanTranscriptGutter`, built by filtering `enabledClis()` on the capability rather than by listing ids, and `_activeCliGutterColumns()` looks the active session's mode up in it. The copy path reads no terminal buffer at all. The declared width is a CEILING, not the answer: `clean()` strips the lesser of it and the run every selected line shares. A block can therefore only shift as a unit, the structure inside a selection survives by construction, and a selection reaching column 0 loses nothing. That is what keeps a `git log` body at its own four-space indent inside an agent's two-column gutter. Codex was measured separately, because it renders nothing like Claude: it draws boxes narrower than the pane and pushes its transcript into ordinary scrollback. On a live 0.154.0 answer its `•`/`›`/`⚠` markers sit in the gutter, prose continuations sit at 2, and a nested YAML block the model wrote rendered at 2/4/6/8 for its own 0/2/4/6. Replayed at 100, 120, 160, 198, 235 and 282 columns its indents were 0, 2, 4, 6 and 8 at every one, never 1. Copying that YAML out of a live Codex pane now yields 0/2/4/6: gutter gone, nesting intact. Two derived versions were built and measured first, and both are recorded in the code because both looked correct: - Painted trailing padding — a full-screen TUI writes real spaces across the unused part of a row, a shell leaves them never-written for xterm to trim — has no false positives and never over-stripped. It is also a function of pane WIDTH: the padding exists only while a rendered line stops short of the CLI's own layout width, and Claude's prose wraps to fill it. Dragging the same two prose rows of one live transcript at five window sizes, the share of padded rows ran 44%, 6%, 6%, 7% and 87% at 123, 160, 198, 235 and 298 columns, so the strip silently did nothing at every ordinary size while a corpus captured entirely at 282 columns said it worked. - Taking the narrowest indent on the rows around the selection fires at every width and over-strips about 1% of selections, because a file listing inside the transcript can be the narrowest thing on screen. Measured over 1,392,281 selections — every 1, 2, 3, 5, 10 and 20-row window of real Claude screens replayed from live PTY streams at 100, 120, 160, 198, 235 and 282 columns — the declared width over-strips none, breaks no relative indent and alters no text, and serves 100% of the selections whose own indent covers the gutter. Verified end to end in a browser with a real mouse drag and a real Ctrl+C: Claude and Codex panes paste flush at 123, 198 and 298 columns, a shell pane is untouched at every one. The strip sits behind `copyStripMargin` (App Settings, Selection & clipboard), per-device and default ON: a display key, absent from the .strict() SettingsUpdateSchema, read as `!== false` because the desktop branch of getDefaultSettings() returns {}. The toggle is checked before the map. Two review findings from #451, handled: - The mid-row flag governs ONE line now. `range.start.x > 0` excludes only the first selected line, the one whose margin the mousedown genuinely cut off, so the same three rows no longer produce three different clipboard results. - The reversed-drag finding does not reproduce on the pinned xterm. `getSelectionPosition()` reads `_selectionService.selectionStart`, whose getter returns `SelectionModel.finalSelectionStart`, and that swaps the pair when `areSelectionValuesReversed()` says so. A real upward mouse drag through chromium against xterm 6.0 reports the same range as the downward drag. `_normalisedSelectionRange()` keeps the ordering as a guard, because the model one layer down exposes the unnormalised fields under the same two names. Tests: test/terminal-copy-clean.test.ts (64, up from 31), plus the injected script stripped in test/server-index-title.test.ts. Every guard is pinned: removing any one of seven reds at least one test, including declaring the wrong gutter width. Full suite green, 7,861 passed, 0 failed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
90fd0a5a15 |
fix(tmux): gate the pane-exit read, and mute the dot on the rich rail too
Four changes the maintainer asked for on Ark0N/Codeman#446 before merging. The pane-exit watcher stays always-on, but a tick now costs nothing when there is nothing to observe. `hasObservablePaneSession()` skips the tmux exec while every session on the manager is one of the shapes `Session.paneExitApplies` already forces to UNKNOWN: a remote SSH session (its local pane holds the ssh client), a docker case (a `docker exec` into the container's own tmux), and a record rebuilt from the socket (no provenance at all). The timer is untouched. Skipping retracts nothing, for the same reason a failed read does not: the map still holds the last real reading, and every path that puts a new command in a pane calls `clearPaneExit()` itself. The two copies of that rule are pinned against each other in `test/session-pane-exit.test.ts`, because drift between them is silent in both directions. `DEFAULT_PANE_EXIT_INTERVAL_MS` was already a constant beside the stats and remote-reconnect intervals; its comment now says why the watcher owns its own cadence and why the number is what it is. The never-default-an-absent-status rule is written where `PaneExit` is declared. It names `status ?? 0` as the thing never to write, and says that an agent the OOM killer took would otherwise read as a user typing `/exit` — which is what absent-stays-absent keeps a later clean-exit sweep away from. Nothing fails when somebody adds that `??`, which is why the sentence is there rather than a test. Checking the dot's specificity found a second fight, and it was losing. On the tab strip the alert rules win as intended: a session that exits with a permission dialog pending still renders red, and yellow for an idle alert. On the rich vertical tab rail they did not — that rail's own `tab-state-*` dot rules are (0,9,1) against the strip's mute at (0,5,0), so an exited session there kept a full green dot AND the working halo beside a badge reading "exited". The rail twin matches that specificity exactly and therefore must stay below those rules in source order; it clears the halo as well, which the strip's rule never had to think about. `test/session-pane-exit-ui.test.ts` now resolves the real stylesheet in jsdom rather than matching selector text: postcss collects every rule that paints `.tab-status`, a real engine decides, and the tests read back the answer. Two mutations were run against it to prove it has teeth — dropping the hand-written alert exclusions fails three cases, and moving the rail twin above the state rules fails one. Refs Ark0N/Codeman#446. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c67c130caa |
feat(web): mark a session tab whose agent has exited
The tab now reads "exited (137)" beside the session name, drawn from the
`paneExit` field the server publishes. `applyPaneExitBadge()` owns the DOM
work, called from the incremental render path — the only path a live session
ever takes, since going from live to exited adds and removes no tab and so
never reaches the full rebuild.
An unknown answer draws nothing. A death tmux could not explain reads "exited"
with no number rather than "exited (0)", so an unexplained death and a clean
exit do not look alike. A signal death reads "exited (signal 9)".
The badge carries `data-i18n-skip`, like the status pills: it is generated
text, `i18n.js` walks inserted content, and a dictionary entry added later
would fight the renderer, whose in-place comparison is against English.
The tab also carries a `tab-agent-exited` class that mutes the status dot. That
dot is drawn from `status`, which stays `idle` or `busy` for an exited pane as
the issue requires, so without this a green or pulsing dot sits beside a badge
saying the agent is gone — the first thing a tester asked about. `status`
itself is untouched, so this is a rendering rule only. The CSS excludes the two
alert classes by hand, following the convention the rich-rail dot rules
document: a dot turning red or yellow because a session is blocked on a human
outranks "the agent exited".
The tab keeps its click behavior. X still closes it, and nothing here closes,
sweeps or restarts anything.
`docs/architecture-invariants.md` gains the mechanism under "Session data and
lifecycle", where every comparable one already lives: what the tri-state means,
the four shapes it is absent for, why the watcher cannot ride the stats
collector, why an absent `#{pane_dead_status}` is not 0, and the three things
that must never happen to an exited pane.
Refs Ark0N/Codeman#446.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
299a21d5f5 |
fix(split-pane): merge-time fixes for split-pane sessions (#453)
The maintainer's promised merge-time fixes from the final review of #453: 1. closeSplitPane() tears down a divider drag still in progress, so a split that collapses mid-drag no longer leaves body.split-pane-resizing (the page-wide col-resize cursor and user-select lock) set until a reload. 2. openSplitPane() re-applies the picker's own exclusions (detached session, pid === null, no session record) for a row that went stale while the menu sat open, refusing silently like its neighbouring gates. 3. architecture-invariants: the hard-hide of .btn-split is the @media (max-width: 1179px) rule in styles.css, not mobile.css. 4. SplitTerminalPane.destroy() nulls onclose (and onerror) beside onopen and onmessage. 5. Picker rows drop the data-session-id attribute nothing read. 6. The Pane-A-ends branch collapses with skipPrimaryResize, so the closing resize is no longer aimed at the session the server just removed. 7. The {t:'r'} refresh path is single-flight across the fetch and the chunked write, coalescing a mid-replay refresh into one trailing re-run. Tests: split-pane-auto-collapse-unit gains the drag-teardown, exclusion and skip-resize cases; the new split-pane-terminal-unit covers destroy() and the refresh single-flight. All were run against the pre-fix module to confirm they fail there. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> (cherry picked from commit dbd39aed015ae5ae5870aba398bf4b4ab5118e47) |
||
|
|
dcf9437308 |
Merge pull request #460 from Ark0N/feat/installer-v2
feat(install): three questions up front, an unattended build, and a URL you can scan |
||
|
|
72d437ab63 |
fix(install): fold in both reviews of #460
The two reviews on the PR (DeepSeek Harness, then Claude) found one class of
bug twice and a list of smaller ones; all of them land here, each pinned in
test/install-sh-invariants.test.ts and, where it is bash logic, driven in the
bash:3.2 CI step as well.
The Start line the done screen prints is now composed in one place
(start_command_hint) from every non-default value, the same five the exec
branch exports through export_bind_env, so "do not start" under a sub-path or
a custom port no longer prints a bare `codeman web`. The --lan / --tailscale /
env preset paths read ${CODEMAN_PASSWORD:-$EXISTING_PASSWORD}: a flag re-run on
a unit that carried a password used to rewrite it without the password and
with the unauthenticated ack. --password and --port flip RECONFIGURE so they
reach the unit instead of taking the quiet update path, and `install.sh name`
re-syncs the unit's base URL after the mapping is re-added.
Also: the sudo keepalive is ended before the exec into the foreground server
(exec skips the EXIT trap, and the loop keys on $$); Ctrl+C in the HTTPS-toggle
poll is trapped for the poll only and skips Tailscale for the run instead of
killing the installer; uninstall asks before removing a LaunchDaemon this
installer never wrote; a foreign daemon gets a launchctl kickstart hint and the
done screen stops claiming the new build is running; the preflight summary
reads the Tailscale state with a line grep when node is not installed yet; the
LAN security notice uses the configured port; a bare re-run ends on the done
screen; a build failure after a rename names the install.sh tailscale
recovery; TS_JOINED_HERE (written, never read) is gone; the plan doc and
architecture-invariants say what the code does. A minor changeset is included.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
46d8b92049 |
fix(split-pane): port Ctrl+Shift+C's never-falls-through guarantee to Pane B
The smart-copy gate only entered its selection-check block behind hasSelection(), so a selection-less Ctrl+Shift+C skipped straight to `return true` and ceded the keystroke to the browser's own handling (e.g. Chrome's Inspect-Element binding) instead of matching Pane A's "never falls through" contract for that chord. Verified live in a real browser that this is a UX-parity fix, not an interrupt-safety one: xterm's evaluateKeyboardEvent never emits PTY data for a shifted ctrl-letter regardless of any gate (only "_" and "@" get special-cased), so no accidental 0x03 was ever at risk. The regression test added here asserts on the dispatched event's defaultPrevented rather than the absence of a WS frame, since the frame-count check passes vacuously for this exact key combo whether or not the gate fires. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
0b3e086334 |
fix(split-pane): address Ark0N's fourth pass — PTY-less picker exclusion, hollow chord test, remaining key gates
- buildSplitPickerSessions() now excludes any session with pid === null (exited CLI, tripped PTY-exit breaker, a restore that never re-attached). Pane B has no equivalent of selectSession()'s auto re-attach POST, so a split opened onto one had nothing reading its tmux pane: no terminal events ever arrived and Session.write() silently dropped every keystroke with no ack either way, while the socket itself reported healthy. - Fixed the hollow chord regression test: the synthetic keydowns carried no keyCode, which is what xterm's evaluateKeyboardEvent switches on to produce a data frame at all, so the assertion held regardless of whether the gate fired. Adding real keyCodes surfaced a second, real bug in the Alt+B case: the event bubbles to app.js's own document-level shortcut dispatcher, which really toggles the sidebar and resets the layout attribute the gate reads before Pane B's own (later, non-capture) handler ever sees it — fixed by driving the app's real settings cache instead of only the DOM attribute. - Ported the two remaining primary-pane gates with real consequences: Ctrl+Z (SIGTSTP) is swallowed for every non-shell session, matching terminal-ui.js's reasoning (an Ink/TUI agent loop stops dead with no visible output otherwise), and Shift/Ctrl+Enter now POSTs to /api/sessions/:id/send-key for THIS pane's own session instead of letting xterm send a bare \r, which used to submit an incomplete prompt instead of inserting a newline. Smart-copy Ctrl+C is re-implemented against Pane B's own terminal (copying app.copyTerminalSelection() would have copied Pane A's selection instead). - Updated docs/architecture-invariants.md and docs/split-pane-sessions-plan.md to match, and added CLAUDE.md's missing .split-picker-menu z-index entry. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
fafef0aa00 |
fix(split-pane): gate app-level chords out of Pane B, address Ark0N's third pass
Pane B had no attachCustomKeyEventHandler of its own, so the document
capture-phase shortcut handler's preventDefault() (which does not stop
xterm) left Ctrl+K/Alt+1/Alt+B ALSO writing their raw byte/escape
sequence into Pane B's live PTY on top of whatever the app action did
to Pane A. Pane B now installs the same registry-aware gates the
primary pane's own attachCustomKeyEventHandler uses. Ctrl+V is left on
xterm's default paste — no image-paste trap to route it to.
Plus the rest of the review's smaller items:
- Narrowing the window past the desktop gate now closes an open split
instead of leaving it stranded on screen.
- Split is refused while a web tab is active (activeWebviewId), which
used to open Pane B's socket behind a hidden container.
- Pane B now handles the server's `{t:'r'}` refresh frame via a shared
_loadBuffer() helper (also used by connect()), instead of ignoring it.
- The divider drag now uses pointer events + setPointerCapture (mirrors
tab-rail-resize.js), a button!==0 guard, preventDefault, and a
body.split-pane-resizing cursor/selection lock — a plain mousedown
drag selected the text under the cursor as it crossed both terminals.
- Pane B's close control and the picker rows are real <button>s now
(keyboard-reachable), with matching CSS chrome resets.
- Dropped the redundant CodemanBase.base prefix on the buffer fetch
(the global fetch wrapper already applies it).
- data-preview-order for the Split settings chip moved from a collision
with Ultracode Agents (both 15/12) to 11.5, matching its real
position between Multi-monitor and Ultracode Agents in the header;
widened test/app-settings-structure.test.ts's regex to allow the
decimal (Number() already parses it fine for the preview sort).
- Added zh-CN i18n entries for the Split button and empty-picker text.
- Dropped the stray unused `vi` import Ark0N flagged as unrelated to
this feature (vitest's `globals: true` makes it ambient anyway).
- Documented the fix and the deliberate no-cid/seq choice in the
split-pane-sessions architecture-invariants entry.
Added a real-Chromium regression test asserting Ctrl+K/Alt+1/Alt+B
dispatched at Pane B's own textarea send no `{t:'i'}` frame over its
WebSocket. Full CI gate green (409 files, 7736 tests) plus all 8
split-pane browser tests.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
||
|
|
6c8bd6c606 |
fix(split-pane): gate the Split button to desktop, make it per-device
Ark0N's PR #453 review: nothing gated this feature to desktop even though the design called for it (two 240px min-width panes plus the divider need ~486px, and the divider has no touch handlers), and showSplitButton was a SYNCED setting, so turning it on at a desk also put the button in the phone header. - Hard-hide .btn-split on phones in mobile.css regardless of the setting, matching the other desktop-oriented header buttons in the same @media (max-width: 599px) block. - Move showSplitButton into settings-ui.js's per-device displayKeys set and drop it from SettingsUpdateSchema entirely, matching the showFileViewerButton/skin precedent (CLAUDE.md's "per-device keys ... must NOT be added to SettingsUpdateSchema" rule) — a desktop opt-in must never sync onto a phone that never asked for it. Removes the now-invalid server-round-trip test for the setting. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
7fc66e8161 |
docs(split-pane): keep the design spec, drop the task-plan scaffolding
Per Ark0N's review on PR #453: rename the design spec to docs/split-pane-sessions-plan.md, matching every other feature's *-plan.md convention, and drop the 957-line implementation task plan (docs/superpowers/plans/2026-09-15-split-pane-sessions.md) — workflow scaffolding for the subagent-driven-development run, not repo documentation. Fixes the now-dangling link in architecture-invariants.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
f7852081b7 |
docs(split-pane): fix orphaned Session list layout section
The new "Split-pane sessions" section was inserted between the "Session list layout (header strip vs. left sidebar)" heading and that section's own body paragraphs, orphaning the heading from its content. Move "Split-pane sessions" to after the Session list layout section's full body, before "Gesture control: the setting" — no change to the Session list layout prose itself, only where the new section sits relative to it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
ba7b8b7bef | docs: add split-pane sessions architecture-invariants entry | ||
|
|
4205f6930f |
fix(release): the seven findings from the pre-release review of the whole tree
A full review of the release tree found seven things, and four of them were mine. **The gate was red, and I put it there.** Splitting `confirmed` into `confirmedContext` and `confirmedSwap` changed the wire field without moving three assertions that check it: `custom-model-one-shot-launch.test.ts` and two in `custom-model-run-menu-ui.test.ts` (the swap modal and the context modal, each of which already receives exactly the right per-question flag). Moved, with the titles. **Worse, my own tests for the split never ran.** The four cases in `session-custom-model.test.ts` that exist specifically to pin it call `mockRunning()`, which was declared inside a sibling `describe`, so they threw a ReferenceError during setup. The split would have shipped with no passing server-side coverage while the gate reported the failure as four broken tests rather than as four tests that were never written. `mockRunning` is hoisted to the outer describe. **The submit verifier pressed Enter into shell panes.** `#455`'s SubmitVerifier resolved its composer glyph as `promptGlyph ?? '❯'`, and only claude and codex declare one, so the other eight modes fell back to claude's `❯`. That is also starship's default shell prompt, and pure's, and spaceship's, and p10k lean's. On such a shell the line `❯ npm run build` sits on screen for as long as the command runs, the verifier reads it as an unsubmitted prompt, and re-presses Enter into the running program's stdin up to nine times on its 2s..60s schedule. Mostly a stray newline; not harmless against a y/N prompt, `read -p`, an installer or a pager, where it takes the default. The module's own fileoverview already stated the rule this broke. Now `?? ''`, which `promptStillInComposer()` already treats as inert, so the verifier runs only for a CLI that actually declares a composer. **My #451 dedent removal left a count behind**: "Two rules keep it honest" introducing three numbered rules. The rest is documentation the split outran. `confirmedContext`/`confirmedSwap` appeared in no doc at all, while `docs/api-reference.md` (the SemVer-covered contract) still told an integrator to retry with `confirmed: true` for both questions, which is precisely the thing the split exists to stop. Documented there, in `docs/custom-model-endpoints.md` and in CLAUDE.md. The custom-model changeset gained the split and the `CLAUDE_CONFIG_DIR` multi-user consequence, both user-visible and both previously absent, and #454's gained the one exception to its own claim: a Custom Endpoints launch ignores the Instance count stepper and always starts one session. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
12de3c5164 |
docs(custom-model): record the CLAUDE_CONFIG_DIR clamp in architecture-invariants
CLAUDE.md gained the admin-only note when the key joined claude's privilegedEnvKeys; architecture-invariants, which is where the exact-key allowlist rule is documented in depth, still described the pre-change world. The reboot-restore half is the one worth writing down: a non-granted owner's already-persisted CLAUDE_CONFIG_DIR is stripped on restore, which moves that session back to the default Claude account with no error. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4c705094f7 |
fix(terminal): ship the copy clean as a trailing trim, without the shared dedent
#451 cleaned two things on copy. The trailing trim is right and every native terminal does it. The shared leading-indent strip is this project's own rule, and it is dropped here rather than shipped. Measured against the shipped transform over 401,445 three-row windows across 1,010 tracked files in this repo, it fired on 73% of them: 92% inside a YAML workflow, 76% over `git log` output, 48% in a TypeScript source. No width threshold separates a margin from content because they are the same widths, a live Claude Code pane's own margins measuring 2 and 5 columns while the most common non-TUI shared run is 4. The failure modes are not symmetric either: a wrong trailing trim costs nothing, while a wrong dedent silently deletes information that was on the screen, with nothing in the clipboard to hint at it, on git log bodies, on indented code read out of cat (semantic in Python), on git diff context rows where the leading space is the marker, and on stack traces. It also could not be made self-consistent cheaply. Whether the first row joined the measurement depended on the mousedown COLUMN, which the user never sees, so one block of three rows produced three different clipboard results; and the flag read getSelectionPosition().start, which is xterm's mousedown anchor and is never normalised, so dragging UP through a block read it off the bottom row. The PR's test stub hardcoded a downward drag, so its suite could not express that case. The transform, the wiring, the tests, the invariants, CLAUDE.md, the wiki page and the changeset all move together. The test block now pins the ABSENCE as a contract, with the git log, Python and git diff cases as its examples, so this is not re-derived later. If it is ever revisited, the one qualification that measured clean is painted trailing padding: zero false positives over all 401,445 windows. Also from the review: the comments and invariant rule justifying the padding-only clear described the pre-change code (the Ctrl+C gate reads the CLEANED selection now, so such a selection falls through to the PTY on its own and the clear is feedback rather than protection), the new 'Nothing to copy' toast gained its zh-CN entry, and the invariants paragraph no longer repeats its own opening sentence. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c376534a50 |
fix(run,terminal): merge-time fixes for the Instance count stepper and capture geometry
#454: the behaviour the PR adds had no test, so a regression test drives runGrok() at tabCount 3 and asserts three quick-start POSTs with sequential w<n>-<case> names (verified to fail against master's session-ui.js). Each caller now reads the count BEFORE its opening banner and announces it there, the way runClaude() already did, so a launch no longer prints two headers and a launch with another session already active still says how many are starting. runClaude() calls the shared _readTabCount() instead of its own copy of the 1..20 clamp, and that helper optional-chains the element read, since hoisting it above each caller's try block would otherwise let a missing #tabCount throw where the launch-error path cannot report it. #435: sizeMovedUnderLoad derived from data.source alone. `mux-visible` is not sufficient: a failed display-message cursor query makes capturePaneBuffer skip the snapshot repaint and return the raw capture, which the route still labels mux-visible, so a size that moved during such a load bought a full forced reload to repair a frame that was never positioned. It now tests Number.isFinite(data.captureRows) like its two siblings. Plus the invariants and CLAUDE.md lines promised on #435: a visible capture reports its geometry and omits it when nothing was positioned, the comparison runs on mux-visible only, and the replay is capped at one attempt and latches per session when it cannot converge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2c3ccdf030 |
Merge pull request #439
feat(remote): wake a sleeping host (Wake-on-LAN) from input, banner and native magic packet |
||
|
|
613b774bf1 |
Merge pull request #451
fix(terminal): trim the padding and shared indent out of a copied selection |
||
|
|
1040f6c489 |
fix(remote): a proxied host is reachability-unknown; scope remote: SSE per session
Review round 2 on #439. 1. The bare TCP probe connects to host:port, which a host behind a jump host or SOCKS proxy does not answer even while ssh works. Acting on that verdict drew a permanent banner over a healthy session, replaced a real "needs tmux" error with "not reachable" in quick-start, and - with a wake target - buffered every HTTP input for the life of the session, since the readiness poll could never succeed. `WakeableRemote` now carries `jumpHost`/`socksProxy`/`extraSshOptions`, and `isProbeable()` turns such a host into reachability-UNKNOWN: input is delivered, `checkReachable` / `checkHostReachable` answer `null` (never `false`), `ensureHostAwake` returns `'unprobeable'` (handled like `'no-target'`), the quick-start gate fires on `=== false` only, and `GET …/reachability` reports `reachable: null, probeable: false` so the banner has nothing to key on. A wake target can still be fired for it, blind: no readiness poll, no reattach, no toast - the response says only whether the packet went out. 2. `'remote:'` joins the session-scoped SSE prefixes. The create/attach wake has no session yet, so the registry names the requesting user (`ensureHostAwake({ requestedBy })` -> `username` in the payload) and `deriveSseHint` routes on it; with neither it fails closed to admins. Single-user mode is unaffected. Smaller, from the same review: - A flush write that fails now drops the remaining buffer (logged) instead of retaining it: the wake still resolved and marked the host reachable, so the retained chunk waited for the NEXT wake and was replayed hours later, after everything typed since. Same policy as the oversized paste. - The banner polls on tab activation (a user action) and on its 30 s timer only for a host with a wake target; a timer connecting to a host Codeman cannot wake is the traffic invariant #2 rejects keepalives for. A proxied host is never polled. - `probeRemoteHostReachable`, `runRemoteWakeCommand` and the default UDP socket refuse under VITEST, as remote-files.ts does. The guard caught a leak on the spot: `createDefaultRemoteWakeDeps({ probe })` overrode the probe but still polled readiness with the real one, so the shutdown test had been connecting to a production address. The poll now uses the injected probe. - docs/remote-sessions.md is additions only again (the reformatting is gone); the architecture-invariants overlap resolved itself in the merge. Live, against a throwaway instance with a non-routable ghost host: proxied -> no probe, no wake, the genuine ssh error after 10 s; direct (control) -> probe, magic packet, "did not come back" after the 40 s budget. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
e271a65e79 |
Merge origin/master into feat/remote-host-wake
Resolves CLAUDE.md count tables (route counts recounted on the merged tree: 235 handlers, sessions 37) and keeps both the host-wake and the reboot-restore banner in index.html. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QdGP4jUTjc9J2RYYykDrCG |
||
|
|
492f8d8ddf |
Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain |
||
|
|
f9edb33d15 |
fix(terminal): trim the padding and shared indent out of a copied selection
xterm hands back whole screen rows and trims only the cells that were never written to, so the real spaces a full-screen TUI paints across the unused part of a row count as content and reach the clipboard. Measured against Claude Code in a 282-column pane, single lines arrived carrying 138 trailing spaces, and every line carried the two-space transcript indent as well. Windows Terminal, iTerm2 and GNOME Terminal all trim that for you, decideAutoCopy already calls a wall of spaces "never what the gesture meant", and _selectTouchSelectionLine already treats those cells as padding — the mouse and keyboard paths never had the same rule. CodemanCopySelection.clean lives in constants.js beside decideAutoCopy, its pure sibling. It drops the trailing run from each line, and removes the leading run only where every selected row shares one. A selection of a single row keeps its run, because one row shares nothing with anything and stripping it would silently reindent one line of `git log` body text or one line out of `less`. A drag that began inside a row keeps its partial first line untouched and out of the measurement, which otherwise pins the shared run to zero and leaves every following row indented. Every pass over a line is a scan rather than a regex. `/[ \t]+(\r?)$/` is quadratic on a line whose spaces are followed by a non-space character, which is what right-aligned or centred TUI content looks like: measured over 50 000 rows with a 280-column run it took 2.9s, against 1.3ms for the scan, and a 2 000-column run took 16s. The scan is also the faster of the two on an ordinary padded row. cleanedTerminalSelection in terminal-ui.js is the half that needs the live terminal. It returns a COLUMN selection untouched: Alt+drag makes one, and a rectangle's rows lining up is the point of the gesture, so both halves of the clean would destroy it. xterm exposes the mode nowhere public, so the check reads terminal._core._selectionService, the way this file already reads terminal._core for cell dimensions, and cleans normally if a future xterm renames the field. A test pins that assumption against the library rather than against a stub repeating the literal. The Ctrl+C chord decides on the cleaned selection, not the raw one. A drag across the blank part of a row selects real padding spaces, so the raw text is truthy, and testing it would spend that press on a copy of nothing and make the user press again to interrupt. A padding-only selection is now dropped and the press falls through to the PTY, while Ctrl+Shift+C still never falls through. copyTerminalSelection gates on trim() for the same reason, since a multi-row drag across padding cleans to line breaks alone and a bare newline pasted into a chat composer submits it. All four of the main terminal's copy paths go through it: the Ctrl+C chord, right-click, the phone selection button and Auto Copy. The browser's own Edit menu copy, a disabled copy shortcut and the subagent windows still copy raw rows, as they did before, and the invariants doc now says so rather than claiming every copy is cleaned. Auto Copy resolves its own toggle before it reads the selection, since it is off by default and a selection can run to the 50 000-row scrollback ceiling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
75a028e825 |
fix(terminal): re-take the sticky-scroll baseline after a replay
A capture load now replays its queued tail, and that replay runs through `batchTerminalWrite`, which samples `_wasAtBottomBeforeWrite` before it queues. It runs inside `chunkedTerminalWrite`, before that promise resolves, with the terminal freshly reset and rewritten — so the sample is always true. The caller then restored the reader's position and the next `flushPendingWrites` scrolled straight back to the bottom off the latched flag, undoing it. The only thing in the way was `_hasRecentUserScrollUp()`, a 1500ms window a server-triggered refresh is usually past. `_syncStickyScrollBaseline()` re-takes the flag from wherever the viewport now sits, and the two paths that restore a position call it right after doing so: `_onSessionNeedsRefresh` and `_maybeRefetchFullHistory`. Those are the paths #259 and #205 exist for, and they are also where a non-empty queue is most likely, since a needsRefresh fires when output is flooding. Re-taking rather than suppressing the sampling: suppressing leaves whatever stale value the flag held from before the load, which on the full-history re-pull has no reason to be false. `selectSession` and `_onSessionClearTerminal` deliberately end at the bottom, so the sampled true is already the truth there and they do not call it. `_bufferLoadFinishOpts` gains the coverage the CI gate can see: both mux sources flush, `history` does not, and a payload naming no source does not. Its only coverage was the browser suite, which CI does not run. The JSDoc and the changeset now record the one duplicate window this cutoff cannot close. The server appends output to the byte buffer in the same tick it emits, but broadcasts on a batch timer — 8ms over WebSocket, 16 to 50ms over SSE — so a batch pending when `capture-pane` ran leaves the server after the reply and is replayed although the capture holds it. It is one batch interval wide against a recovery window spanning the whole chunked write, and closing it means flushing that batch server side before the capture. The second browser test asserts its session was created, so a failed create fails it instead of passing with zero hits. docs/architecture-invariants.md no longer claims the replay leaves the queued-event discard window alone. That clause now describes what decides how a load ends, the baseline rule, the batch window, and the three covering tests. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
acb8d4b0aa |
docs(remote): correct what the wake PR moved
- `host-wake-ui.js` joins the documented load order (12.2) and gets its `@dependency`/`@loadorder` tags; the frontend module count is 33, not 32. - `remote-wake` is not "(pure)" — the module uses `dgram`/`net`/`child_process`. - SSE counts: 160 constants, and the category is "Remote auto-reconnect / wake (5)"; the route table's per-file counts are refreshed (sessions 37, cases 34). - The CLAUDE.md wake rule now names the create/attach wake, the 40 s request budget, the whole-chunk paste drop, the registry's lifetime (drop on cleanup, stop on shutdown) and the deliberately non-wake-aware WebSocket keystroke path — that paragraph is what the next person reads. - Reverted the eight lines of unrelated Prettier markdown churn in `docs/architecture-invariants.md` (docs/ is not in the format glob, so it was an editor): only the new wake paragraph remains in the diff. |
||
|
|
7b947fa3f1 |
fix(remote): close the wake-state leaks and the dishonest wake budget
Review follow-up on the wake-on-LAN PR (five findings, all of them about the state the feature keeps and the budgets it inherits): - Wake state is dropped by `WebServer.cleanupSession` instead of the two delete routes, so it now goes with the session on EVERY cleanup path (cron, admin, scheduled-run teardown, error paths) instead of surviving with up to 4 KB of the user's buffered keystrokes. `registerSessionRoutes` returns the registry so the server can own its lifetime without the wake-capable code living in `server.ts`; the wiring guard is updated to allow that and gains a second assertion that `server.ts` calls nothing but `drop`/`stop` on it. - `_effectiveRemote` returns before `_state`, so a LOCAL session no longer gets a wake-state entry — the input gate runs on every keystroke, so that entry used to be allocated for every session the user types in. - An input chunk larger than the 4 KB cap is dropped OUTRIGHT instead of being head-trimmed and then written as a fragment: one paste is one `input` value and was never typed character by character, so its tail is a partial command the user never sent. The drop is logged. - The manual wake button passes `REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS` (40 s) like the create/attach paths, instead of inheriting the 90 s session default that the dashboard's reverse proxy cuts off at 60 s. - `RemoteWakeRegistry.stop()` aborts in-flight readiness polls (abortable sleep) and refuses new wakes, and `WebServer.stop()` calls it, so a restart during a wake no longer waits the poll out. - The banner/toast wording keys off a new `queuedInput` flag on the two SSE events, which is true only when the server actually holds bytes: browser keystrokes travel over the WebSocket, which never passes through the registry, so the wake BUTTON must not promise queued input. The failed-wake path also stops pattern-matching the error message (it re-asks the reachability route) and the WoL dialog says "admin-only" instead of "host not found" for a non-admin in multi-user mode. |
||
|
|
d0a5a583cd |
feat(remote): wake a sleeping host when a session is created or attached
Pressing Run on a remote case whose host was asleep failed with `could not verify tmux on remote host 192.168.50.137: …` — an ssh error that blames tmux for a machine that is merely suspended. The only wake paths were typed input on an established session and the banner's Wake button, so OPENING a session (the moment the user actually decides to use that host) had none. `RemoteWakeRegistry.ensureHostAwake()` reuses the existing probe/wake/readiness machinery for a host that has no session yet, and is wired into the two user-initiated create paths: `POST /api/quick-start` for a remote case (before the tmux prereq probe, which is what surfaced the misleading error) and `POST /api/sessions` with `attachRemoteSession`. A host without a wake target is not even probed, so its behavior and latency are byte-identical. The wake is blocking — the caller gets the session or an error — but bounded by REMOTE_WAKE_REQUEST_READY_TIMEOUT_MS (40 s) instead of the 90 s session default, because the dashboard sits behind a reverse proxy whose default `proxy_read_timeout` is 60 s: a longer wait would be cut off at the proxy while the session was still being created. The budget has to cover the whole request (40 s wake + 1.5 s probe + the tmux probe's own 15 s = 56.5 s worst case), which is why it is 40 s and not 45. A timeout now says the host did not come back, and an unreachable host without a wake target says so instead of pointing at tmux. The wiring is deliberately in the HTTP ROUTE, never in the shared session service: `cron-service.ts` builds sessions there with nobody waiting on the answer, and a wake on that path would power the host on for every schedule — the timer-driven re-wake invariant #1 exists to prevent. Both halves are asserted (importers of `remote-wake`, and `ensureHostAwake` having exactly one caller file), so a future caller has to come through the guard test. A rejection from the wake IO is caught too: a broken target must fail the wake, not the route. `remote:hostWaking`/`remote:hostWakeFailed` now carry `forNewSession` for the session-less case, where "input is queued" would be untrue; the toast then reads "the session starts when it is back". Live wake numbers are unchanged (this reuses the measured ~12 s S3 path); the route behavior is covered by new tests in session-routes.test.ts with an injected registry, so no test opens a real socket or ssh. |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
8b5a13435a |
feat(remote): host-unreachable banner, manual wake, and native MAC wake-on-LAN
The reactive wake (typing into a session whose host slept) left the state invisible: nothing told the user the machine was asleep, and with no wake target configured there was nothing to do about it. Adds: - RemoteHost.wakeMac (comma-separated) - Codeman builds and broadcasts the magic packet itself (UDP port 9), so the common case needs no external script. The existing wakeCommand stays as the explicit override. - GET /api/sessions/:id/reachability - probes (throttled, cached, and it never wakes) and reports HOW the host can be woken, or that nothing is configured. - POST /api/sessions/:id/wake - wakes, waits, reattaches the pane and flushes buffered input; 400 with a routable message when no target is configured. - The amber host-unreachable banner + its 'Wake' / 'Configure WoL' action, and a small config dialog that saves via PUT /api/remote-hosts/:id. - RemoteWakeDeps.resolveRemote: host config is re-resolved for LIVE sessions (throttled + cached), so saving the dialog takes effect without a restart. |
||
|
|
3f0bfde54a |
docs(remote): document the wake-on-LAN invariants; drop wake state on bulk delete
Self-review pass: the input-ladder's two 'buffer' branches were the same three lines, and bulk delete left a session's (bounded, per-random-uuid) wake state behind. Documents the design where the code refers to it - remote-sessions.md section, the architecture invariant, and the CLAUDE.md key pattern. |
||
|
|
70fc6b32d5 |
docs: record the dup/last input ACK, Shift+drag and right-click copy, and multi-case adopted containers
Three behaviours landed from #375 without their doc entries: the duplicate input ACK now carries `dup:true` and the server's watermark (`docs/reliable-input-delivery.md` still described a bare ACK), Shift+drag and right-click copy in the terminal (the shortcut list did not know them), and one adopted container backing several cases at different in-container directories (the Docker cases paragraph still implied one case per container for adopted containers too). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
e49c48145b |
fix(files): fail closed on remote symlinks, guard PUT for remote cases, bound ssh fan-out
Follow-up to #421 (remote-case file reads over ssh), addressing the review. Symlink escape on a host without `readlink -f` (blocker). The probe's portable fallback canonicalized only the directory chain and returned the final component unresolved, so on macOS < 12.3 `ws/notes.txt -> ~/.ssh/id_rsa` came back as `.../ws/notes.txt` (with the target's size), passed every containment and blocklist check that runs on `realPath`, and `cat` followed the link. The fallback now walks the directory chain with `cd -P`/`pwd -P` and follows the LAST component with plain `readlink` for a bounded number of hops, and anything it cannot fully resolve (a loop, a readlink failure, the hop cap) is reported with an `x` marker that parses as null, i.e. 404. It never returns the unresolved string. Measured on a real /bin/sh with `readlink -f` shadowed: the pre-fix script reports `/ws/notes.txt`, the fixed one `/secret/id_rsa`; both branches (native and fallback) now agree. `PUT /api/sessions/:id/file-content` never had the remote guard the PR described. It sits ahead of `validateSessionFilePath`, which resolves against the LOCAL filesystem, because with a same-named directory on the Codeman host (an sshfs mount of the remote tree, the documented stop-gap) the write landed on the local twin while the viewer believed it edited the remote file. ssh fan-out is bounded. `src/remote-ssh-limiter.ts` is a document-conversion-limiter-shaped semaphore (default 4, env `CODEMAN_MAX_REMOTE_FILE_SSH`) around every probe and buffered read; the attachment-history list resolves its whole history in ONE batched probe (`probeRemoteAttachmentHistory`, threaded into `registerExternalAttachment({remoteProbes})` so the guards run unchanged) instead of one handshake per entry; and probes chunk at 40 paths because the whole script is one argv string. Terminal output in a remote session is written on the remote host, so a prompt-injected agent printing hundreds of `codeman://attach` links forked one ssh per link, each holding a 20 s timeout, and a 100-entry history re-listed on every attachment:detected tripped OpenSSH's default MaxStartups. Streams are deliberately not counted (one per browser request, held for a whole playback, and gated behind a counted probe anyway). Smaller items from the same review: probe records are NUL-terminated and index-keyed after a leading NUL (a newline in a filename can no longer shift the alignment, and the banner is fenced off without last-N-lines guessing); size comes from `stat -c %s || stat -f %z`; the three IO functions refuse under VITEST instead of opening a connection; an unreachable host now reads as unknown (missing: false) for detected AND external history entries, where external used to fold its 502 into missing; a client that aborted during the guard probe has its body's ssh child reaped (`reply.raw.destroyed` is checked before the close listener is attached); `describeExecError` never returns Node's `Command failed: <ssh line>` message, which carried the identity path and the probe script into a 502 body; and the docs note that `isSensitivePath`'s three home-anchored entries resolve against the Codeman host's home, not the remote one. Tests: the probe script runs on a real /bin/sh with a `readlink` shim that rejects `-f` (the escape, a relative chain through a symlinked directory, a loop, a newline filename, banner chatter that itself looks like a record), the limiter's cap and FIFO order, and route tests for the PUT guard (local twin untouched, no connection), the single batched history probe, the unreachable-host alignment and the aborted-client reap. All four route tests fail against the pre-fix file-routes.ts. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
792a251e35 |
Merge pull request #421 from Randalix/fix/remote-file-access
fix(files): read remote-case previews, downloads and attachments over ssh |
||
|
|
63aafdf274 |
fix(files): serve remote-case attachments, the path a click takes outside the case
A clicked path that points OUTSIDE the case directory goes through the attachment routes (the frontend's `_isExternalPreviewPath` sends every absolute path not under `workingDir` to `POST /attachments`), and those had the same local-`fs` assumption as file-raw: `realpathSync`/`fs.stat` on a path that only exists on the remote host, so the file never opened — the case the #415 report was actually about. - `registerExternalAttachment()` accepts `remote` and resolves through `remoteProbePaths` (canonical path, size/mtime, kind, plus the workspace root for the confinement check). Everything around it — blocklist, extension allowlist, workspace confinement, registry/dedupe — is now shared by both branches, so the remote path cannot drift from the local one. - The by-id routes (`raw`, `preview`, `thumbnail`), the metadata poll and the attachment history list resolve over ssh too. `raw` streams with the same Range contract as file-raw; `preview` (office) and `thumbnail` answer 400 for a remote record; an unreachable host answers 502, a vanished file 404. - Which host a record is read from follows the SESSION, never the path string: the same absolute path is a different file on each host, and a remote session never falls back to a local file with that name. - Codex generated artifacts keep force-workspace confinement for a remote case: the well-known artifact directories are anchored at THIS host's home, so only a file inside the remote workspace is trusted. Still local-only by design: writes, office conversion, thumbnails, the file tree/picker and tail-file. |