mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
1da2fa2529b305c23d1d41d4db2fffe85dbd8c5b
21
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0a52a99ca9 |
feat(cli-registry): CLI management write API + Settings UI (Phases 1-6) (#476)
* feat(cli-registry): add cliManagementEnabled flag and GET /api/clis
Phases 1-2 of docs/cli-enable-disable-plan.md ("PR C" from the #343
review): a synced, default-OFF master flag gating the upcoming CLI
management surface, plus a read-only GET /api/clis endpoint listing
every registry entry (stock + custom, enabled or not) for the
Settings UI. Non-admins in multi-user mode see an empty list rather
than a 403. Write endpoints, auto-install, custom entry CRUD and the
Settings UI list itself land in later phases.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* feat(cli-registry): Phases 3-6 - write API + custom entries + Settings UI
Completes docs/cli-enable-disable-plan.md ("PR C" from the #343 review).
Phase 3: PUT /api/clis/:id toggles enabled for any EXISTING entry (stock or
custom) via a shallow merge onto its clis.json override; shell/claude are
structurally un-disableable (Decision 4), an unknown id 404s rather than
becoming a creation backdoor.
Phase 4: POST /api/clis/:id/install runs a STOCK entry's already-vetted
install command (shell:true, bounded by timeout, process-group killed on
expiry, output captured, audit-logged). A custom entry's id is refused
outright, independent of anything Phase 5 does (Decision 3: a custom
entry's install text is display-only, never executed).
Phase 5: POST /api/clis (create) / PUT /api/clis/custom/:id (update) /
DELETE /api/clis/:id (custom only) — a deliberately minimal request shape
(id/label/shortBadge/binaries/a simple launch variant), assembled into a
full CliEntry with conservative capability defaults and re-validated
through CliEntrySchema before writing, never a relaxed path for
UI-originated entries. Stock-id collisions, duplicate custom ids, and
edits/deletes against a stock id are all rejected explicitly.
Phase 6: the Settings UI section (App Settings -> Agents & CLIs), gated
independently on cliManagementEnabled AND admin-in-multi-user-mode
(Decision 5), fetching/rendering GET /api/clis and wiring every write
endpoint above.
Every write endpoint answers the same way when the feature is off: 403
FORBIDDEN via one shared requireCliManagementGate() (Phase 1's own
checklist item). registry-writer.ts is a new, deliberately separate write
module so registry.ts itself stays import-side-effect-free, same tmp+
rename+0600 shape as custom-model-hosts.ts.
27 new/updated route tests covering every gate, collision, and cleanup
path; full CI gate green (415/416 files, 7854 tests).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GuHtuPiHXdykq9T6rKQJ9n
* fix(cli-registry): toggling a CLI off in Settings never hid it anywhere else
window.__codemanCliAvailable — the flag isCliAvailable() reads client-side
to gate the welcome-screen buttons, the Run-menu dropdown and the mobile
overview — was built purely from each CLI's own installed-on-PATH resolver
(isClaudeAvailable() etc.), with no reference to the registry's `enabled`
flag at all. So disabling a CLI via the new Settings UI (or a hand-edited
clis.json) updated the settings row and nothing else: every launch surface
kept offering it, both live and after a full page reload, since even a
fresh render never consulted the registry.
Fixed in two places:
- server.ts: after building `available`, intersect the nine real
SessionMode ids against `enabledClis()`. git/cloudflared (utility
binaries, not CLI registry entries) and deepseekBinary (a secondary
installed-only flag for the "add a profile" affordance) are deliberately
left alone.
- settings-ui.js: `toggleCliEnabled()` now patches
`window.__codemanCliAvailable` in place and refreshes the welcome screen,
the mobile overview and an already-open Run menu, mirroring the existing
`installDeepSeekProfile()` pattern for the same "injected once, needs an
explicit patch" reason — without this half, the server-side fix alone
still left every surface stale until the next reload.
New test in test/render-index-html.test.ts: an installed-but-disabled CLI
(codex, forced via clis.json + reloadCliRegistry()) reads as unavailable,
while an installed-and-enabled one (claude) is unaffected by the override.
Verified on the Debian devbox (codeman-devbox, real tmux — this sandbox has
none and WebServer's constructor hard-requires it): typecheck clean, the
new test passes (17/17 in render-index-html.test.ts), the CLI-registry
suites pass (86/86), and the full CI gate is green (415 test files, 7855
tests, 0 failures).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N6eadpRyqpA9PD3i139cSD
* docs(cli-registry): update the CLI-management plan with status, gotchas, and the Run-menu gap
Phases 1-6 were implemented across two commits (
|
||
|
|
8c73128cd6 | feat(split-pane): auto-collapse split when either session ends | ||
|
|
2abf328db8 |
feat(split-pane): add open/close orchestration, picker, and divider drag
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
97a1238c85 |
feat(split-pane): add SplitTerminalPane class for Pane B
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
3edf9aae2f |
fix(terminal): replay a pane capture at the geometry it was taken at
A visible-frame capture repaints each row at an absolute position, counting up to the pane's height. A terminal shorter than that clamps every address past its own height onto its last line. The overflow rows then overwrite one another, and the rows underneath are lost. Replaying a real 50-row capture into a 30-row terminal rendered 28 lines of a 45-line command and drew the frame twice. Nothing in the response said what height the frame was built for, so the client could not detect this. A capture now reports the geometry it was really taken at through `capturedGeometry` on `PaneCaptureOptions`, and the terminal response carries it as `captureCols` and `captureRows`. When the captured pane is taller than the terminal, or the size that produced the capture did not survive the load, `selectSession` replays once at the size that stuck. `resizeRetry` caps that at one attempt, so two competing fits cannot trade replays forever. The retry re-arms the full-history flag only when the pass that ran had consumed it. A tab switch takes the bounded tail, so its retry takes the tail too: clearing the flag unconditionally would upgrade that switch into a fresh scrollback capture the user never asked for, which the route's own comments put at tens of megabytes. What this repairs is a capture that won a race against the resize meant to precede it. It does not repair a capture whose pane was too tall because `Session.resize` declined the resize outright, which it does for a small viewport while a desktop viewport's size claim is live. The retry re-sends the same declined resize and captures the same pane, and `resizeRetry` then stops it. Repairing that means changing who owns the pane size, which is a policy question this does not touch. The reported geometry still helps there, because the client can see the mismatch at all rather than being blind to it. Follows #395, #396 and #397, which fixed the other ways the replayed frame and the terminal could disagree. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
492f8d8ddf |
Merge pull request #436 from irisitymichaelgrundberg/fix/replay-output-that-arrived-after-the-capture
fix(terminal): keep the output a pane capture could not contain |
||
|
|
c9515b1d4c |
fix(terminal): keep the output a pane capture could not contain
Live terminal events are queued while a buffer load runs, and the load discards that queue when it ends. That is right when the loaded buffer is the server's accumulated byte history. The route appends to that history right up to the moment it serializes the response, so a queued event already appears in it and replaying it would duplicate output, most visibly Ink's cursor-up redraws. A tmux pane capture is a photograph, current only as of the instant `capture-pane` ran. Output printed afterwards was queued and then dropped, and nothing scheduled a re-fetch to recover it: `_onSessionNeedsRefresh` is wired only to the 128KB overflow path. The CLI's next partial redraw then landed on a frame the terminal never received. How much went missing depended on which capture the route served. A `?full=1` load returns the capture alone, with no history in front of it, so it lost everything from the capture to the end of the chunked write. A `?tail=` load returns history, a clear, and then the capture, and the route reads that history after the capture, so it lost everything from the response to the end of that write. The chunked write dominates either way. An agent CLI hides the loss on its next full redraw; a shell session does not, because its output is linear and nothing repaints it. Queue entries now carry their arrival time, and `_finishBufferLoad` takes a `since` cutoff, so a capture load replays exactly the tail that arrived after the response headers. The earlier events stay dropped, because a payload that carries history does hold those. All four paths that fetch a terminal buffer and write it now decide this the same way, through one `_bufferLoadFinishOpts` helper, so they cannot drift apart: `selectSession`, `_onSessionNeedsRefresh`, `_onSessionClearTerminal` and `_maybeRefetchFullHistory`. The second of those is the one that stings. It exists to restore output the client already dropped once under backpressure, and it was dropping more output while performing that recovery. The cache-hit write inside `selectSession` stays on discard deliberately: it runs before the fetch, so its queue holds only events the capture that follows already contains. Two further things had to change for that tail to still exist when the load ends, and a browser test is what found both. `chunkedTerminalWrite` is what ends the load for every non-empty buffer, so the flush policy travels to its own finish calls; the call in `selectSession` runs only when the write was skipped. `_beginBufferLoad` no longer empties the queue when one load re-enters it, which it does on every write, because that reset discarded the whole fetch window before anything could replay it. The response already distinguishes the sources. `source` reads `mux-visible` or `mux-full-history` for a capture and `history` for the byte stream. Follows #395, #396 and #397, which fixed the ways the replayed frame itself could disagree with the terminal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
e6e5a62d9b |
Merge pull request #380 from opticon454/feature/cli-catalog-consumers
feat(cli-registry): drive install.sh and the Docker agent image from the CLI catalogue |
||
|
|
c2d019d956 |
chore: move the maintainer PR bot out of this repository
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the npm package: a Telegram bot that reviews open pull requests in Codeman sessions and reports to the maintainer. It now lives in its own private repository and keeps running unchanged, as a client of Codeman's HTTP API like any other. It moved because it grew a second watcher, for GitHub Discussions, and shipping that here would mean publishing the briefs it hands its review sessions, the judgement calls in them and its safety model. None of that helps anyone installing Codeman, and all of it is easier to change when it is not a public interface. The move cost nothing structurally: the whole tree depended on one external package plus Node builtins. What this removes from the repo, and nothing else: the sources, their three test files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script, the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md keeps a short pointer in place of the section, because the bot still constrains work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads into `refs/pr-bot/*` of this checkout, which it must never check out or reset. The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history rather than drift, and is left alone. Verified after the removal: typecheck, lint and format:check clean, and the suite passes 6843 tests across 357 files, which is the previous run minus exactly the 70 tests that moved out with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
7d6f612ef5 |
feat(cli-registry): generate a CLI catalogue for install.sh and the Docker build
Two consumers of the registry cannot import TypeScript: `install.sh`, which runs via `curl | bash` before any checkout exists, and `scripts/build-agent-image.mjs`. Both currently hand-maintain their own CLI lists, and both have already drifted. `scripts/generate-cli-catalog.mts` (`npm run generate:cli-catalog`, plus a `--check` mode) emits from `STOCK_CLIS`: - `config/clis.stock.json` for the `.mjs` and the tests. It carries `enabled` — the field the earlier attempt omitted, which is how a disabled CLI's npm package still got baked into every agent image. - a marker-delimited block inside `install.sh`, embedded rather than fetched. The embedded copy is the FULL catalogue on purpose: the earlier design fetched it and fell back to a hardcoded two-CLI list, degrading silently on an empty response. There is no degraded mode to fall into now. The block is bash 3.2 safe: parallel indexed arrays, no associative arrays, no namerefs, no mapfile. Variable-length lists use OFFSET/LENGTH windows into one flat array rather than a delimiter, so a $HOME containing a space needs no IFS handling and `shell` (no binaries) gets length 0 and is never iterated. Search paths are emitted dir-major, matching the probe order the hand-written arrays use and `test/install-sh-detection-parity.test.ts` pins. Only fields the two consumers need are exported. `launch`/`env`/`capabilities`/ `overlays` are spawn-time concerns the server alone interprets, and a test asserts they never leak into the artifact. `main()` sits behind an `isMainModule()` guard so the sync test can import the renderers. Without it, importing the module would rewrite the artifacts as a side effect of checking them — passing always, guarding never. This commit adds the block; it does not yet delete the hand-written arrays, so the detection pin keeps measuring both against each other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
e8a93ada1f |
fix(terminal): forward the orphaned input event instead of replaying a guessed key
The previous shape guessed the character from `event.key` on keydown, re-emitted it, and then tried to suppress a late canonical copy with a 250 ms character-keyed dedupe. Review found three defects in that, all reproducible: the dedupe matched on the character alone with nothing scoping a candidate to the keydown that created it, so the same character typed twice inside the window had its second, real byte swallowed; anything whose committed text differed from `event.key` (Enter, IME punctuation) was delivered twice, because the dedupe could never match it; and the trigger ignored `key === 'Unidentified'`, which is what a soft keyboard reports, so it may never have fired where it was needed. The input event already carries the committed text in `ev.data` — exactly what xterm itself would have forwarded — so nothing has to be guessed. The controller now only decides WHETHER to forward, by asking whether xterm produced canonical data since the keydown that began the keystroke. No character-keyed matching survives, so the first two defects are structurally impossible rather than defended against, and nothing reads `key`/`keyCode`, so the third cannot recur. Three details are load-bearing and each has a test that fails without it: - The "did xterm speak?" snapshot is taken at KEYDOWN, not at the input event. `_keyPress` emits and sets `_keyPressHandled` before `input` fires, so a snapshot read at input time already contains that emission, reads it as silence, and delivers the character twice. - Our `input` listener is registered with `capture: true`. The target is visited twice in the event path, so a capture listener calling `stopPropagation()` stops later BUBBLE listeners on that same target; xterm's `cancel()` runs exactly in the branch where it handled the input, so on bubble we would never observe handled events, and whether we observed them at all would hang off `options.cancelEvents`. Measured in jsdom and headless chromium; the table is in the module header. - Enter is deliberately no longer special-cased. That mapping is what made the committed text differ from the re-emitted value in the first place. The scope is also narrower than the old name suggests, and the browser test now proves it rather than assuming it. For a keydown that reports keyCode 229 xterm ALREADY self-rescues, via `CompositionHelper._handleAnyTextareaChanges()` diffing the helper textarea on a 0 ms timer. A test asserting "we recovered it" there passes while xterm does all the work, so the browser tests assert WHO delivered the byte: zero canonical emissions for the genuinely orphaned case, exactly one delivery for the case xterm rescues itself. Also addresses review notes: the module gains an `@fileoverview` with `@dependency`/`@loadorder` and an entry in the load-order list and module inventory, and the wiring test moves out of the Ctrl+C smart-copy file into its own. The keydown hook deliberately still runs for every key event rather than moving behind the 229 gate: gating it would reinstate exactly the blindness described above, and it is now a single counter assignment. |
||
|
|
f33b37c008 |
feat(pr-bot): review open PRs in Codeman sessions and report over Telegram
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.
Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.
typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
e3a2fb767f | feat(tabs): COD-358 add resizable vertical session rail | ||
|
|
947ff6f6fa |
chore(test): make npm test the CI gate and give each excluded suite a runner
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.
`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.
The suites it leaves out are not abandoned; each has a command:
test:browser 5 Playwright files (chromium + a live server; codex-predictive-echo
also needs a real codex binary)
test:mobile unchanged — the above plus per-machine PNG baselines
test:perf 2 wall-clock benchmarks; need an otherwise idle machine
test:all the old everything-behaviour, kept reachable
test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.
The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.
So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:
gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo
⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.
Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fa02bd4503 |
test(predictive-echo): E2E suite against a real codex TUI (Layer 5)
Out-of-process lab server (VITEST markers stripped so tmux/codex are real), CODEMAN_INSTANCE=codexlab on port 3222, throwaway CODEX_HOME with a fake key. Ten scenarios: bundle smoke, predict+converge typing, the #218 arrow retest (submitted text exact), the #222 live picker, the #219 paste order, the #220 wrap, the trust-modal ghost eliminator, the localEchoEnabled kill switch, the end-to-end byte-identity trace (predictor active vs null), and a display-delayed 300ms-RTT run pinning instant spans with exact pixel geometry plus arrow-edit correctness under lag. Live-TUI hardening learned the hard way: codex Ctrl+U kills only to line start (End first), a fake-key submit leaves a Reconnecting loop that can kill codex seconds later (retry-cancel + composer stability probe; the submitting scenario runs after all composer-state ones), and typing must wait for the predictWhen gate itself, not merely a rendered composer. CI-excluded like the other Playwright suites; skips cleanly when codex is not installed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d9123de9eb |
feat(terminal): Ctrl+C copies the selection, interrupts when nothing is selected
Closes #211. Copying from the terminal only worked through the browser context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown, so the muscle-memory copy failed silently and read as "no copy-paste at all". With a selection, Ctrl+C now copies it, toasts, clears the selection and sends nothing to the PTY. With no selection it falls through unchanged, so the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never falls through: an explicit copy that interrupts a running agent because the selection happened to be empty would be a footgun. Three details that keep the interrupt safe: - The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the no-selection path returns true WITHOUT preventDefault. xterm calls the custom handler before its own cancel(), so returning false alone does not cancel the event; the copy path therefore calls preventDefault explicitly, or the browser would run its native copy on top of ours. - copy-selection is a full registry entry (rebindable and disableable in App Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the same trick command-palette uses: the generic capture loop preventDefaults every match it dispatches, which would cost the user the interrupt key. - The gate is keydown-only, since the custom handler also runs for keypress and keyup. Copy goes through _copyText (Clipboard API, then hidden-textarea + execCommand) rather than raw navigator.clipboard, because install.sh's LAN option serves plain HTTP where navigator.clipboard is undefined; the fallback steals focus, so the terminal is refocused afterwards. Tests: test/terminal-copy-selection.test.ts pins the gate and the SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real key presses in chromium and asserts on the clipboard plus the bytes xterm emitted (browser-driven, so excluded from test:ci like the other Playwright suites). Verified manually on an isolated beta instance before landing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
da7a095e33 |
chore: move knip config into config/ and Prettier config into package.json
Continues trimming the repo root so the README is reached with less scrolling. Root files: 19 -> 15 across both passes. - knip.json -> config/knip.json, joining eslint.config.js and the vitest configs. `npm run knip` now passes --config explicitly. Verified by A/B: the run from the new location produces byte-identical findings and the same five configuration hints as from the root, so knip resolves its globs relative to cwd rather than the config file. Those hints are pre-existing, not caused by the move. - .prettierrc -> the "prettier" key in package.json, a config source Prettier reads natively, so editor format-on-save keeps working with no --config flag anywhere. Verified live: `npm run format:check` still passes across src/**, which it could not if the config had been lost (Prettier's defaults are double quotes at 80 columns and would flag nearly every file). .prettierignore deliberately stays at the root: Prettier resolves it relative to cwd, so moving it would require threading --ignore-path through every script and would break editor integration. Everything else in the root is load-bearing: .editorconfig (walks up from the edited file), .nvmrc/.npmrc (read from the project root), tsconfig.json (bare `tsc` discovers it), LICENSE (GitHub license detection), install.sh (its raw URL is the published one-liner in the README and cannot move without breaking every copy in the wild), plus the five documented .md files. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d5f91e4cd7 |
test(ci): run the unit suite in CI + frontend-syntax gate; green pre-existing test debt
- CI: add a 'test' job running the unit suite via config/vitest.ci.config.ts. Excludes browser (Playwright/chromium) and perf tests (timing-flaky), like the existing test/mobile suite. Safe in CI: TmuxManager no-ops shell commands under VITEST (test/setup.ts). - Add scripts/check-frontend-syntax.mjs (node --check on src/web/public/*.js), wired into the lint job — catches a class of frontend SyntaxError that passes lint today (lint globs only TS). - Add test/security-regression.test.ts (wired Host/Origin guard, self-update CSRF, CSP/security headers, text/plain raw body, WS anti-CSWSH) + test/sse-registry-parity.test.ts (backend<->frontend SSE registry parity). - Green pre-existing test debt surfaced by the new gate: stale 'Session not found' asserts -> 'not found' substring; drop tests for removed helpers (isError now internal; createSuccessResponse deleted); file-stream-manager: mock realpathSync + fix stale /tmp assertion; sse-subscription-filter: lifecycle events broadcast to all clients (only terminal stream filtered); session.test.ts: mkdir /tmp/test; skip one interactive-respawn test needing a real PTY (covered by respawn-controller.test.ts). - Full non-mobile suite verified green locally (2680 passed, 12 skipped). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7101e64800 |
refactor: restructure repo for cleaner GitHub landing page
Reduce visible top-level items from 21 to 14: - Untrack test-results/, tmp/, public symlink (added to .gitignore) - Move agent-teams/ → docs/agent-teams/ - Move mobile-test/ → test/mobile/ - Move tools/remotion/ → scripts/remotion/ - Move eslint.config.js, vitest.config.ts → config/ All path references updated across CLAUDE.md, package.json, .prettierignore, vitest configs, and capture scripts. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |