mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-09-30 12:39:42 +02:00
1306f731cfb2895993162964f0efac0b3cd49f26
13
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e6e5a62d9b |
Merge pull request #380 from opticon454/feature/cli-catalog-consumers
feat(cli-registry): drive install.sh and the Docker agent image from the CLI catalogue |
||
|
|
c2d019d956 |
chore: move the maintainer PR bot out of this repository
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the npm package: a Telegram bot that reviews open pull requests in Codeman sessions and reports to the maintainer. It now lives in its own private repository and keeps running unchanged, as a client of Codeman's HTTP API like any other. It moved because it grew a second watcher, for GitHub Discussions, and shipping that here would mean publishing the briefs it hands its review sessions, the judgement calls in them and its safety model. None of that helps anyone installing Codeman, and all of it is easier to change when it is not a public interface. The move cost nothing structurally: the whole tree depended on one external package plus Node builtins. What this removes from the repo, and nothing else: the sources, their three test files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script, the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md keeps a short pointer in place of the section, because the bot still constrains work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads into `refs/pr-bot/*` of this checkout, which it must never check out or reset. The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history rather than drift, and is left alone. Verified after the removal: typecheck, lint and format:check clean, and the suite passes 6843 tests across 357 files, which is the previous run minus exactly the 70 tests that moved out with it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a0628a40e8 |
fix(cli-registry): address maintainer review on #380
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
|
||
|
|
7d6f612ef5 |
feat(cli-registry): generate a CLI catalogue for install.sh and the Docker build
Two consumers of the registry cannot import TypeScript: `install.sh`, which runs via `curl | bash` before any checkout exists, and `scripts/build-agent-image.mjs`. Both currently hand-maintain their own CLI lists, and both have already drifted. `scripts/generate-cli-catalog.mts` (`npm run generate:cli-catalog`, plus a `--check` mode) emits from `STOCK_CLIS`: - `config/clis.stock.json` for the `.mjs` and the tests. It carries `enabled` — the field the earlier attempt omitted, which is how a disabled CLI's npm package still got baked into every agent image. - a marker-delimited block inside `install.sh`, embedded rather than fetched. The embedded copy is the FULL catalogue on purpose: the earlier design fetched it and fell back to a hardcoded two-CLI list, degrading silently on an empty response. There is no degraded mode to fall into now. The block is bash 3.2 safe: parallel indexed arrays, no associative arrays, no namerefs, no mapfile. Variable-length lists use OFFSET/LENGTH windows into one flat array rather than a delimiter, so a $HOME containing a space needs no IFS handling and `shell` (no binaries) gets length 0 and is never iterated. Search paths are emitted dir-major, matching the probe order the hand-written arrays use and `test/install-sh-detection-parity.test.ts` pins. Only fields the two consumers need are exported. `launch`/`env`/`capabilities`/ `overlays` are spawn-time concerns the server alone interprets, and a test asserts they never leak into the artifact. `main()` sits behind an `isMainModule()` guard so the sync test can import the renderers. Without it, importing the module would rewrite the artifacts as a side effect of checking them — passing always, guarding never. This commit adds the block; it does not yet delete the hand-written arrays, so the detection pin keeps measuring both against each other. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12 |
||
|
|
e8a93ada1f |
fix(terminal): forward the orphaned input event instead of replaying a guessed key
The previous shape guessed the character from `event.key` on keydown, re-emitted it, and then tried to suppress a late canonical copy with a 250 ms character-keyed dedupe. Review found three defects in that, all reproducible: the dedupe matched on the character alone with nothing scoping a candidate to the keydown that created it, so the same character typed twice inside the window had its second, real byte swallowed; anything whose committed text differed from `event.key` (Enter, IME punctuation) was delivered twice, because the dedupe could never match it; and the trigger ignored `key === 'Unidentified'`, which is what a soft keyboard reports, so it may never have fired where it was needed. The input event already carries the committed text in `ev.data` — exactly what xterm itself would have forwarded — so nothing has to be guessed. The controller now only decides WHETHER to forward, by asking whether xterm produced canonical data since the keydown that began the keystroke. No character-keyed matching survives, so the first two defects are structurally impossible rather than defended against, and nothing reads `key`/`keyCode`, so the third cannot recur. Three details are load-bearing and each has a test that fails without it: - The "did xterm speak?" snapshot is taken at KEYDOWN, not at the input event. `_keyPress` emits and sets `_keyPressHandled` before `input` fires, so a snapshot read at input time already contains that emission, reads it as silence, and delivers the character twice. - Our `input` listener is registered with `capture: true`. The target is visited twice in the event path, so a capture listener calling `stopPropagation()` stops later BUBBLE listeners on that same target; xterm's `cancel()` runs exactly in the branch where it handled the input, so on bubble we would never observe handled events, and whether we observed them at all would hang off `options.cancelEvents`. Measured in jsdom and headless chromium; the table is in the module header. - Enter is deliberately no longer special-cased. That mapping is what made the committed text differ from the re-emitted value in the first place. The scope is also narrower than the old name suggests, and the browser test now proves it rather than assuming it. For a keydown that reports keyCode 229 xterm ALREADY self-rescues, via `CompositionHelper._handleAnyTextareaChanges()` diffing the helper textarea on a 0 ms timer. A test asserting "we recovered it" there passes while xterm does all the work, so the browser tests assert WHO delivered the byte: zero canonical emissions for the genuinely orphaned case, exactly one delivery for the case xterm rescues itself. Also addresses review notes: the module gains an `@fileoverview` with `@dependency`/`@loadorder` and an entry in the load-order list and module inventory, and the wiring test moves out of the Ctrl+C smart-copy file into its own. The keydown hook deliberately still runs for every key event rather than moving behind the 229 gate: gating it would reinstate exactly the blindness described above, and it is now a single counter assignment. |
||
|
|
f33b37c008 |
feat(pr-bot): review open PRs in Codeman sessions and report over Telegram
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.
Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.
typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
||
|
|
e3a2fb767f | feat(tabs): COD-358 add resizable vertical session rail | ||
|
|
947ff6f6fa |
chore(test): make npm test the CI gate and give each excluded suite a runner
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.
`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.
The suites it leaves out are not abandoned; each has a command:
test:browser 5 Playwright files (chromium + a live server; codex-predictive-echo
also needs a real codex binary)
test:mobile unchanged — the above plus per-machine PNG baselines
test:perf 2 wall-clock benchmarks; need an otherwise idle machine
test:all the old everything-behaviour, kept reachable
test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.
The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.
So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:
gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo
⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.
Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
fa02bd4503 |
test(predictive-echo): E2E suite against a real codex TUI (Layer 5)
Out-of-process lab server (VITEST markers stripped so tmux/codex are real), CODEMAN_INSTANCE=codexlab on port 3222, throwaway CODEX_HOME with a fake key. Ten scenarios: bundle smoke, predict+converge typing, the #218 arrow retest (submitted text exact), the #222 live picker, the #219 paste order, the #220 wrap, the trust-modal ghost eliminator, the localEchoEnabled kill switch, the end-to-end byte-identity trace (predictor active vs null), and a display-delayed 300ms-RTT run pinning instant spans with exact pixel geometry plus arrow-edit correctness under lag. Live-TUI hardening learned the hard way: codex Ctrl+U kills only to line start (End first), a fake-key submit leaves a Reconnecting loop that can kill codex seconds later (retry-cancel + composer stability probe; the submitting scenario runs after all composer-state ones), and typing must wait for the predictWhen gate itself, not merely a rendered composer. CI-excluded like the other Playwright suites; skips cleanly when codex is not installed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d9123de9eb |
feat(terminal): Ctrl+C copies the selection, interrupts when nothing is selected
Closes #211. Copying from the terminal only worked through the browser context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown, so the muscle-memory copy failed silently and read as "no copy-paste at all". With a selection, Ctrl+C now copies it, toasts, clears the selection and sends nothing to the PTY. With no selection it falls through unchanged, so the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never falls through: an explicit copy that interrupts a running agent because the selection happened to be empty would be a footgun. Three details that keep the interrupt safe: - The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the no-selection path returns true WITHOUT preventDefault. xterm calls the custom handler before its own cancel(), so returning false alone does not cancel the event; the copy path therefore calls preventDefault explicitly, or the browser would run its native copy on top of ours. - copy-selection is a full registry entry (rebindable and disableable in App Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the same trick command-palette uses: the generic capture loop preventDefaults every match it dispatches, which would cost the user the interrupt key. - The gate is keydown-only, since the custom handler also runs for keypress and keyup. Copy goes through _copyText (Clipboard API, then hidden-textarea + execCommand) rather than raw navigator.clipboard, because install.sh's LAN option serves plain HTTP where navigator.clipboard is undefined; the fallback steals focus, so the terminal is refocused afterwards. Tests: test/terminal-copy-selection.test.ts pins the gate and the SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real key presses in chromium and asserts on the clipboard plus the bytes xterm emitted (browser-driven, so excluded from test:ci like the other Playwright suites). Verified manually on an isolated beta instance before landing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
da7a095e33 |
chore: move knip config into config/ and Prettier config into package.json
Continues trimming the repo root so the README is reached with less scrolling. Root files: 19 -> 15 across both passes. - knip.json -> config/knip.json, joining eslint.config.js and the vitest configs. `npm run knip` now passes --config explicitly. Verified by A/B: the run from the new location produces byte-identical findings and the same five configuration hints as from the root, so knip resolves its globs relative to cwd rather than the config file. Those hints are pre-existing, not caused by the move. - .prettierrc -> the "prettier" key in package.json, a config source Prettier reads natively, so editor format-on-save keeps working with no --config flag anywhere. Verified live: `npm run format:check` still passes across src/**, which it could not if the config had been lost (Prettier's defaults are double quotes at 80 columns and would flag nearly every file). .prettierignore deliberately stays at the root: Prettier resolves it relative to cwd, so moving it would require threading --ignore-path through every script and would break editor integration. Everything else in the root is load-bearing: .editorconfig (walks up from the edited file), .nvmrc/.npmrc (read from the project root), tsconfig.json (bare `tsc` discovers it), LICENSE (GitHub license detection), install.sh (its raw URL is the published one-liner in the README and cannot move without breaking every copy in the wild), plus the five documented .md files. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d5f91e4cd7 |
test(ci): run the unit suite in CI + frontend-syntax gate; green pre-existing test debt
- CI: add a 'test' job running the unit suite via config/vitest.ci.config.ts. Excludes browser (Playwright/chromium) and perf tests (timing-flaky), like the existing test/mobile suite. Safe in CI: TmuxManager no-ops shell commands under VITEST (test/setup.ts). - Add scripts/check-frontend-syntax.mjs (node --check on src/web/public/*.js), wired into the lint job — catches a class of frontend SyntaxError that passes lint today (lint globs only TS). - Add test/security-regression.test.ts (wired Host/Origin guard, self-update CSRF, CSP/security headers, text/plain raw body, WS anti-CSWSH) + test/sse-registry-parity.test.ts (backend<->frontend SSE registry parity). - Green pre-existing test debt surfaced by the new gate: stale 'Session not found' asserts -> 'not found' substring; drop tests for removed helpers (isError now internal; createSuccessResponse deleted); file-stream-manager: mock realpathSync + fix stale /tmp assertion; sse-subscription-filter: lifecycle events broadcast to all clients (only terminal stream filtered); session.test.ts: mkdir /tmp/test; skip one interactive-respawn test needing a real PTY (covered by respawn-controller.test.ts). - Full non-mobile suite verified green locally (2680 passed, 12 skipped). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7101e64800 |
refactor: restructure repo for cleaner GitHub landing page
Reduce visible top-level items from 21 to 14: - Untrack test-results/, tmp/, public symlink (added to .gitignore) - Move agent-teams/ → docs/agent-teams/ - Move mobile-test/ → test/mobile/ - Move tools/remotion/ → scripts/remotion/ - Move eslint.config.js, vitest.config.ts → config/ All path references updated across CLAUDE.md, package.json, .prettierignore, vitest configs, and capture scripts. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |