* fix(self-update): stop a stalled status from blocking every later update
A Homebrew node upgrade under a long-running server deletes the versioned
Cellar path the server passes as --node, so every status write from the
updater failed. The update itself still built and restarted (npm and the
build use node from PATH), but update-status.json stayed "queued" forever.
The boot reconcile ran one minute after the restart, inside its 15 min
window, and isInFlight() had no age limit, so "An update is already in
progress." blocked every later update until the next server restart.
- self-update.sh falls back to node on PATH when --node is not executable.
- expireStalledStatus() (pure) fails an in-flight status whose last write
is older than the stale window; applied on every read (start + status
poll) and persisted. The live updater heartbeats every few seconds, so a
running update never trips it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(self-update): a hung graceful shutdown no longer leaves a LaunchDaemon install down
On a KeepAlive LaunchDaemon (headless macOS) the updater restarts by sending
the server SIGTERM and letting launchd respawn it. launchd only respawns once
the process EXITS, and nothing escalates a stuck stop (systemd would SIGKILL
after TimeoutStopSec). Observed after an update to 1.32.1: the server closed
port 3000, server.stop() never resolved, the process stayed alive and the
service stayed down until it was killed by hand.
- cli.ts: the signal handler arms an unref'd 10s timer that force-exits if
server.stop() hangs.
- self-update.sh (launchd-daemon): wait up to 30s for the server pid to exit,
then SIGKILL it. tmux sessions live outside the server and survive.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Codeman maintainer <noreply@anthropic.com>
Add Case -> Clone Repo could only reach public repositories in the Docker
deployment. This lets a deployment opt in to the GitHub CLI and the Azure
CLI (+ azure-devops extension) as git credential helpers. Codeman itself
still collects no credentials.
- server.Dockerfile / agent.Dockerfile: CODEMAN_INSTALL_GH /
CODEMAN_INSTALL_AZ build args (0 or 1, default 0; anything else stops the
build). Off leaves no apt repository, package, extension, helper script
or credential entry, so a default build is unchanged. On installs from
the vendors' apt repositories and configures system gitconfig helpers:
github.com / gist.github.com -> `gh auth git-credential`, dev.azure.com /
*.visualstudio.com -> new docker/git-credential-azure-cli (an Entra ID
token from `az account get-access-token`, or AZURE_DEVOPS_EXT_PAT).
A helper whose CLI is not signed in prints nothing, so a private clone
still fails fast.
- The extension lives in AZURE_EXTENSION_DIR outside HOME
(/opt/codeman-az-extensions, runtime-owned; /opt/az-extensions, gid-0
group-writable in the agent image).
- Hosts turn them on in docker-compose.override.yml: `build: args:` for the
server image, `environment:` CODEMAN_AGENT_IMAGE_INSTALL_GH / _AZ for the
agent image. build-agent-image.mjs and the in-app auto-build share one
env -> ARG table (pinned by the parity test) and pass nothing when unset.
docker-compose.yaml is untouched; .env.example only gains a comment, so
the self-updater's environment gate sees no new keys.
- Docker cases seed the gh sign-in (~/.config/gh/hosts.yml, config.yml) and
the az sign-in files from ~/.azure per file, read-only, like pi/grok.
- The Clone Repo AUTH_REQUIRED message says how to sign the server's git
in instead of claiming private repositories cannot be cloned.
- Docs: docker/README.md "Private repositories", docker-compose.md,
docker-cases.md, the Quick-Start / Core-Concepts / Docker-Cases wiki
pages, security-architecture.md, architecture-invariants.md, changeset.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0167CiuzLrmjYWxwKp3rMWjw
Four ways the terminal can silently stop being correct — in each case the
buffer keeps updating, nothing throws, and the only recourse is a reload.
1. Renderer freeze after backgrounding. iOS DISCARDS scheduled rAF callbacks
when a PWA backgrounds, and xterm's RenderDebouncer only clears its
`_animationFrame` handle from inside that callback — so one drop leaves it
permanently set and every later refresh() early-returns. Parsing is
decoupled from rendering, so bytes keep filling the buffer correctly while
nothing paints. Codeman has exactly ONE xterm for the whole page load, so a
single backgrounding wedges it until a reload. Adds a 2s liveness poll and
`_kickRenderer()`, which does what the dropped `_innerRefresh` would have.
2. Replay clears raced live output. xterm's write() is async-queued while
reset() is synchronous and, per upstream, "does not clear input buffers and
does not reset the parser" — so bytes queued before a reset are parsed after
it and fuse into the snapshot. Verified against the real xterm 6 here:
write('p8'); reset(); write('rmissions') renders "p8rmissions". The main
path was already safe via a queued erase; the needsRefresh and clearTerminal
paths were not. All three now share one queued `\x1bc` (RIS), which unlike
3J/H/2J also resets modes, charsets, scroll regions and SGR state.
3. Output lost on WebSocket reconnect. Input frames carry seq+cid and are
delivered exactly once; output frames carry nothing. ws.onopen re-sends dims
and flushes queued input, and needsRefresh only fires on external-CLI
startup and SSE backpressure drain — never on reconnect. Output produced
while offline was simply absent afterwards. Interim fix: reaching onclose
means the drop was unintentional, so the session is marked and the next open
reconciles from the server buffer. Sequencing output is the follow-up.
4. Terminal captures had no deadline. No AbortController anywhere in the
frontend, including `?full=1`, which the code itself calls "unbounded-ish
work: at the default history limit it can be megabytes". Adds a budget that
scales with full-vs-tail and with captures in flight, degrading to a plain
fetch where AbortController is missing.
Also: the service-worker precache was dead — the build content-hashes assets
but sw.js listed pre-hash names, so 15 of 23 entries 404'd (verified against a
running instance) and cache.add().catch() hid it. Offline still worked via
runtime caching, but CACHE_NAME was a constant so activate's cleanup never
deleted anything and every past release's assets accumulated. Both are now
derived from the build manifest. Crash-trail entries are flattened and capped,
since they are joined with \n into one value and one call site interpolates a
server-controlled WS close reason.
The watchdog reads xterm privates — there is no public API. Every access is
optional-chained so a shape change degrades to a no-op. `_renderService` only
exists after open(), which needs a real DOM, so the gate cannot assert the
field path; test/xterm-private-api.test.ts pins the dependency range instead.
Tests: 23 new (terminal-resilience, sw-precache-manifest, xterm-private-api),
all pure/static so they run in the gate, which excludes the mobile suite. One
static source guard in history-truncation-notice updated for the renamed call;
the behaviour it pins is unchanged.
Not verified: no browser available, so no runtime reproduction of the freeze
and no real-device test of the reconnect path. Both warrant a device pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The caveat #429 added ends with "see the docs above", and "the docs
above" is CLI_DOCS[$i], which for DeepSeek is the upstream harness repo.
Per docs/deepseek-integration.md the harness ships only the web,
headless and base profiles, so following that link and running
`npm install -g @deepseek-ai/dsh` leaves the reader exactly where the
caveat is warning them about: a dsh that cannot drive a pane. What
actually resolves it is Codeman's own Run dropdown, which offers
"DeepSeek: add a terminal profile..." and installs one in a click.
The new wording stays generic for any future launcherProfile entry,
since Codeman is the thing being installed at all three call sites.
Also flips one word in the generator: the comment said "see
installCommandFor below" and that function is defined above it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses the "left as they are"/"worth knowing" items Ark0N named when
merging #380 (the CLI-catalogue-driven install.sh + Docker agent image
PR), none of which were correctness-blocking but all of which were real:
- Removed install.sh's dead _cli_index/check_cli/get_cli_path helpers:
the catalogue-driven menu and hints stopped calling them and nothing
else ever did.
- The generator no longer emits CLI_KIND/CLI_NPM, two bash arrays
install.sh never read (the .mjs/docker-hosts.ts producers already
read the JSON catalogue's kind/npmPackage fields directly, so only
the bash copies were dead).
- detect_all_clis now skips a disabled entry's probe entirely instead
of running it and filtering the result downstream. No stock entry
ships disabled today, so this closes a latent inefficiency before it
is a latent bug rather than fixing an observed one.
- The install hint for a launcherProfile entry (DeepSeek today) now
explains in one line why it's a docs link and not a command: its own
docs page documents `npm install -g @deepseek-ai/dsh`, which installs
the launcher only and can't drive a pane, the exact trap the menu
already avoids by withholding the command. Driven by a new generated
CLI_LAUNCHER_ONLY array (from discovery.launcherProfile), not an id
check, so any future launcherProfile entry gets the same caveat free.
- Corrected the non-interactive-default comment: on a wget-only host,
Claude's curl one-liner is filtered out of the offered list first, so
the default becomes whichever npm-based entry sorts earliest instead
(Codex today), not always Claude. Behaviour is unchanged — it was
already printed, never silent — only the comment overclaimed.
Tests: extended test/install-sh-invariants.test.ts with a positive
guard for the new array and the trimmed array list, a negative guard
that CLI_KIND/CLI_NPM/the three dead helpers cannot come back, and two
real-bash tests (driven the same way the existing skip-menu tests are)
proving a disabled entry is genuinely never probed rather than merely
filtered after the fact.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011WzDjJnbK7zug8iQWnCc9z
Custom Model Endpoint Profiles (#393) let a session point its CLI at a
custom OpenAI-compatible endpoint by injecting env vars or a config file
and restarting the CLI in place. Review of the apply path found four
things, two of them destructive. This lands all four plus the smaller
items from the same review.
1. Clearing a selection did not clear it. The injected vars reach the CLI
via `tmux setenv`, which persists at the tmux-session level and is
inherited by `respawn-pane` (measured: `setenv FOO bar` survived two
successive `respawn-pane -k`), so deleting the keys from the session's
envOverrides relaunched the CLI still pointed at the old endpoint, and
for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just
been deleted. `Session.setCustomModel()` now reports the removed keys,
queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys`
carries them into `applyEnvOverrides()`, which `setenv -u`s them before
re-applying the live overrides, on the same path that already unsets
the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket
that `setenv -u HOME` hands the next respawn the global HOME back.
2. Applying a model to a local claude session killed the pane. The
relaunch was `claude --session-id <id>` and Claude refuses an id that
already has a transcript, and unlike the dead-pane respawn this one
kills a working pane first. `restartCli()` now pins the live
conversation id as the resume id for that respawn when the CLI's launch
declares a `fallback` chain, which renders the same
`--resume <id> || --session-id <id>` shape the docker and remote pane
commands use. Gated on the registry shape, not the CLI id: an entry
whose resume id is minted by the CLI itself never declares that chain.
3. pi, omp and grok wrote their config file and then launched without the
`--model` that selects it, so the file was ignored. The registry entry
now declares `customModelInjection.launchModel` (`custom/{modelId}` for
pi and omp, grok's `[model.codeman-custom]` block name), the builder
renders it, and `_withCustomModelLaunchModel()` applies it onto the
respawn options through `legacyConfigField`, leaving the stored
<Mode>Config untouched so a clear falls back to the user's own model.
A model id the CLI's `model` token pattern cannot carry is refused
with a 400 rather than silently dropped by the argv engine.
4. Remote (SSH) and Docker sessions reported `restarted: true` and changed
nothing: their `restartCli()` reattaches the durable tmux rather than
relaunching the agent, and the env lands on the local pane. Both are
refused with a 400 until those paths are plumbed.
Smaller items from the same review:
- The selection survives a Codeman restart as the disk-only `__customModel`
bookkeeping (endpoint, model, injected key NAMES, config dir, launch
model; never the values, which carry the API key). Recovery re-derives
the values from the endpoint store through the same apply path the route
uses and keeps the bookkeeping even when the endpoint is gone, so a
later clear still has keys to unset.
- Discovery goes through `webviewFetch()`, so the RESOLVED address is
judged by the same egress guard the web-tab proxy uses, and `baseUrl`
reuses `webviewUrlSchema` (http(s) only, no embedded credentials,
link-local and cloud-metadata addresses refused). undici's `fetch failed`
wrapper is unwrapped so the user sees the ECONNREFUSED underneath.
- `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session
config dir 0700/0600 (pi and omp embed the key literally), and that dir
is removed with the session.
- `PR.md` is gone from the repo root and the design doc moved to
`docs/custom-model-endpoints-plan.md` with the LAN address and the
personal name scrubbed; every reference follows. The guide's `authStyle`
text matches the shipped schema (`bearer | api-key`, default `bearer`)
and says that `customModelEndpointsEnabled` is read by nothing until
the picker lands.
- `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts`
(four real type errors fixed). It is not yet wired into `npm run typecheck`
because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json`
there is the one-line follow-up.
Tests: `test/session-custom-model-restart.test.ts` drives a real Session and
fails on the unfixed code for items 1 to 3; the route suite covers item 4
and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets
run before the overrides and that a shell-metachar key never reaches tmux.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`scripts/pr-bot/` was maintainer tooling, not part of the server, the CLI or the
npm package: a Telegram bot that reviews open pull requests in Codeman sessions
and reports to the maintainer. It now lives in its own private repository and
keeps running unchanged, as a client of Codeman's HTTP API like any other.
It moved because it grew a second watcher, for GitHub Discussions, and shipping
that here would mean publishing the briefs it hands its review sessions, the
judgement calls in them and its safety model. None of that helps anyone
installing Codeman, and all of it is easier to change when it is not a public
interface. The move cost nothing structurally: the whole tree depended on one
external package plus Node builtins.
What this removes from the repo, and nothing else: the sources, their three test
files, `config/tsconfig.pr-bot.json`, `docs/pr-bot.md`, the `pr-bot` npm script,
the bot's globs in the typecheck/lint/format scripts, and its knip entry. CLAUDE.md
keeps a short pointer in place of the section, because the bot still constrains
work in here: it takes the `prbot-<n>` and `dscbot-<n>` session names on the local
Codeman, holds clones under `~/.codeman/pr-bot/`, and fetches pull-request heads
into `refs/pr-bot/*` of this checkout, which it must never check out or reset.
The CHANGELOG entries from 1.25.0 and earlier still describe it. That is history
rather than drift, and is left alone.
Verified after the removal: typecheck, lint and format:check clean, and the suite
passes 6843 tests across 357 files, which is the previous run minus exactly the
70 tests that moved out with it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With the repo root as the plugin root, `claude plugin install codeman@codeman`
copied the whole checkout into its cache and, because that root carries a
package.json, ran an npm install there: 832 MB, 511 packages and this repo's
postinstall build on every installer's machine (measured from a clean worktree
of the previous commit). A plugin root must be a directory without one.
The plugin is now `plugins/codeman/`: its manifest, a README, and a MIRROR of
`skills/codeman/`. A mirror rather than a symlink because the install copies
the plugin directory and a link pointing outside it would dangle; a mirror
rather than the source because every install path, injector and doc already
names `skills/codeman/`. `scripts/sync-plugin.mjs` (replacing
sync-plugin-version.mjs) mirrors the skill and syncs both manifest versions
inside `version-packages`; `test/plugin-manifest.test.ts` pins byte-identity,
the versions, the absence of a package.json in the plugin root and that the
repo root `.claude-plugin/` holds only the marketplace manifest.
`claude plugin validate --strict` now passes for both the plugin and the repo
root. Install commands are unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`.claude-plugin/marketplace.json` at the repo root makes
`/plugin marketplace add Ark0N/Codeman` work, and the one plugin it lists is
the repo itself (`source: "./"`), whose one component is `skills/codeman/`.
So `/plugin install codeman@codeman` is a third install route next to
`npx skills add` and `codeman skill install`, and the skill shows up in the
plugin directories that index Claude Code marketplaces.
Both manifests carry package.json's version: `scripts/sync-plugin-version.mjs`
rewrites them inside `version-packages`, right after `changeset version`, and
`test/plugin-manifest.test.ts` pins the equality, the skill's frontmatter name
(without it the installed skill would be named after a versioned cache dir),
and that no other plugin component (`commands/`, `agents/`, `hooks/`,
`.mcp.json`, `settings.json`) appears at the repo root, since an install would
silently ship it.
Verified with `claude plugin validate` (one expected warning: CLAUDE.md at a
plugin root is not plugin context) and a local marketplace add, install,
details, uninstall cycle against a clean checkout of this commit.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Rebased onto current master (the one real conflict was the import line
in docker-hosts.ts Ark0N flagged; kept both), then addressed every
point from the review:
**1. Rebase.** Done — this branch now sits on current upstream/master.
**2. Agent-image special cases are data now, not an id-keyed table
outside stock.ts.** `AGENT_IMAGE_SPECIAL_CASE_IDS`/`AGENT_IMAGE_SPECIAL_CASES`
are gone. `CliDiscovery.install.agentImageLayer?: { kind: 'dedicated';
reason: string }` is a field on the registry entry itself (pi,
deepseek), `reason` is required by schema.ts, both producers
(docker-hosts.ts and cli-catalog.mjs) filter on its presence instead
of an id, and the coverage test reads it from the generated catalogue.
Also added the npm-package-name validation to the TS producer, which
only the .mjs one had — same SAFE_PACKAGE regex, duplicated
(necessarily, one side can't import the other) and now pinned
byte-identical by a new parity test.
**3. Changeset said five, it's eight.** (Not nine — see the DeepSeek
point below, which changes the true count.) Reworded to state it
structurally rather than pin a number that will go stale again.
Then the four behavior-changing findings:
- **DeepSeek was offered as a normal install option but can't actually
drive a pane.** `npm install -g @deepseek-ai/dsh` installs the
launcher only; DeepSeek ships no profile that can run standalone.
The generator now emits an empty install command for any
`launcherProfile` entry, so install.sh's menu (which requires a
non-empty command) skips it and falls through to its docs URL hint
instead — matching what the old hand-written code did before this
PR replaced it.
- **wget-only hosts lost every automatic install, including the npm
ones that never needed curl.** The menu-building loop now filters
PER ENTRY (only a command starting with `curl ` is held back) rather
than wiping the whole menu when DOWNLOADER != curl.
- **The DISPLAY/TRUSTED split and the catalogue refresh didn't hold up
under review** (refresh's only real write was the label; it ran
before the Node existence check; its own eval-detection test was
tripped by the word "eval'd" in a comment). Dropped entirely per
your own recommendation — embedded catalogue only, no network
fetch, no second array. install-sh-invariants.test.ts now asserts
the refresh/DISPLAY machinery does not exist rather than testing its
internals.
The three take-or-leave items, applied:
- `dsh_banner_probe`'s bash 3.2 empty-array bug: `${runner[@]}` →
`${runner[@]+"${runner[@]}"}`. Verified live in a real `bash:3.2.57`
container with `timeout` removed from PATH — crashed before, clean
now, full `detect_all_clis` path exercised end to end.
- `docker-agent-image-coverage.test.ts` now anchors on each layer's
`<binary> --version` proof line instead of `Dockerfile.includes(binary)`,
which stayed true if a layer were deleted but its comment survived.
- Doc drift: docs/docker-cases.md (four → five, and now describes the
data field), docker/agent.Dockerfile's "other four CLIs" comment (no
longer a magic number — CLI_NPM_PACKAGES is generated and can grow),
CLAUDE.md's install.sh size (104KB → ~112KB) and its stale mention of
the now-dropped refresh.
Verified: tsc clean, prettier clean, the full targeted suite (142
tests across the 8 affected files) green, and the full `npm test` gate
diffed BY TEST NAME against a clean upstream/master baseline run on
this same machine — identical 201-name failure set both sides (168
tests / 67 files, all pre-existing Windows-environment noise: symlinks,
PTY spawning, POSIX permission bits — none of it touching anything
this PR changes), zero new failures either side of the diff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
docker/agent.Dockerfile hardcoded the four npm-published CLIs it installs, one
of the several lists that had to be kept in step with the registry by hand.
It now takes them as `ARG CLI_NPM_PACKAGES`, supplied by
scripts/build-agent-image.mjs from config/clis.stock.json, with the default set
to today's list so a bare `docker build` still produces the same image. The arg
is expanded unquoted because word splitting is what turns the list into several
arguments, which is exactly why every token is validated against
^[@A-Za-z0-9][@A-Za-z0-9/._-]*$ on the producing side; a package name carrying a
space or a metacharacter is refused rather than reaching the RUN line. Verified
by building the layer: four packages in, four arguments out, and the default
still applies with no arg.
The list is filtered on each entry's `enabled` flag — the field whose absence
was the maintainer's §3 finding, where a CLI shipping disabled still got baked
into every image. No stock entry is disabled today, so that assertion would pass
vacuously; a unit test feeds the pure helper a fabricated disabled entry so the
fix is covered now rather than the first time someone ships one.
⚠️ It reads the STOCK catalogue, never the merged registry. A user's
~/.codeman/clis.json must not change what is inside an image tagged
codeman/agent:base, or two machines holding that tag hold different images.
Four CLIs keep hand-written layers because the registry cannot describe what
makes them special: pi's --ignore-scripts, deepseek's pnpm companion and dsh-tui
profile, and the three standalone installers. Rather than extend the schema for
a Docker-only benefit, the coverage test requires each to carry a written reason
AND still be present, so an exclusion cannot quietly become an omission.
There are two producers of this command line and there have to be — the .mjs
cannot import TypeScript, and src/docker-hosts.ts builds the same argv for the
in-app auto-build — so a parity test pins them together, package list, arg pairs
and rendered argv. Their order is pinned too: a different order is a different
RUN string and so a needless cache miss between the two build paths.
docker/server.Dockerfile is deliberately NOT edited (PRs #373 and #377 both
modify it); its narrower list is asserted as a declared omission list instead, so
the divergence is reviewable without touching the file.
Also fixes the in-app hint at index.html, which the new coverage test caught
still omitting omp.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
Two consumers of the registry cannot import TypeScript: `install.sh`, which runs
via `curl | bash` before any checkout exists, and `scripts/build-agent-image.mjs`.
Both currently hand-maintain their own CLI lists, and both have already drifted.
`scripts/generate-cli-catalog.mts` (`npm run generate:cli-catalog`, plus a
`--check` mode) emits from `STOCK_CLIS`:
- `config/clis.stock.json` for the `.mjs` and the tests. It carries `enabled` —
the field the earlier attempt omitted, which is how a disabled CLI's npm
package still got baked into every agent image.
- a marker-delimited block inside `install.sh`, embedded rather than fetched.
The embedded copy is the FULL catalogue on purpose: the earlier design fetched
it and fell back to a hardcoded two-CLI list, degrading silently on an empty
response. There is no degraded mode to fall into now.
The block is bash 3.2 safe: parallel indexed arrays, no associative arrays, no
namerefs, no mapfile. Variable-length lists use OFFSET/LENGTH windows into one
flat array rather than a delimiter, so a $HOME containing a space needs no IFS
handling and `shell` (no binaries) gets length 0 and is never iterated. Search
paths are emitted dir-major, matching the probe order the hand-written arrays
use and `test/install-sh-detection-parity.test.ts` pins.
Only fields the two consumers need are exported. `launch`/`env`/`capabilities`/
`overlays` are spawn-time concerns the server alone interprets, and a test
asserts they never leak into the artifact.
`main()` sits behind an `isMainModule()` guard so the sync test can import the
renderers. Without it, importing the module would rewrite the artifacts as a
side effect of checking them — passing always, guarding never.
This commit adds the block; it does not yet delete the hand-written arrays, so
the detection pin keeps measuring both against each other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015EMxQreQUZX5ZyybxAGh12
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi,
Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its
native cloud backend, for a given session. Covers local hardware (llama.cpp,
Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter).
Off by default (customModelEndpointsEnabled, synced, default OFF).
- Registry: capabilities.customModelInjection per CLI entry (env /
configContentEnv / configDir / unsupported kinds)
- Pure injection builder (custom-model-injection.ts) turning an endpoint +
model id into the real env vars / config content per CLI
- Endpoint store + CRUD routes (custom-model-hosts.ts,
custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded
- Session integration: Session.setCustomModel()/restartCli()
(POST /api/sessions/:id/custom-model), reusing the existing
respawn-pane -k primitive to restart the CLI process with new env
- Multi-user hardening: every new redirect-capable env var added to its
CLI's privilegedEnvKeys, closing a pre-existing gap where several were
already reachable via the generic envOverrides field's prefix allowlist
- Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries
against a real endpoint outside the web UI, independent of tmux/sessions
- Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying
every CLI's injected values through a real HTTP shape
Real end-to-end validation against a live llama-swap server (inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries) found and
fixed three real bugs before they shipped:
- Codex's config.toml schema was wrong ([model].default table instead of
a top-level model string + [model_providers.custom]); fixing it then
surfaced a genuine, documented protocol incompatibility (Codex only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement)
- Claude Code's async session-title-generation call validates
ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and
hangs the whole -p invocation on an unrecognized name; documented for
chunk 6, worked around in the standalone script only (--bare is NOT
safe for a real interactive session, which needs hooks)
- The discovery route's authStyle: 'both' option (send both Authorization
and api-key headers) reliably hung a real server; removed the option
entirely rather than just changing the default
Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see
PR.md and deployment_plan.md for the full chunk breakdown and confidence
table.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
codeman-node-modules and codeman-dist (docker-compose.yaml) are seeded
from the image only while empty, so a rebuilt image's fresh dist/
node_modules sat unused behind old volume content until something
cleared it. The in-app self-updater never hit this (it rebuilds INSIDE
the running container, into the very volume already in use), but a
`docker compose build` triggered from outside it — Start-Codeman.sh,
after a manual `git pull` — did: the container came back up looking
unchanged, serving stale compiled routes against current source.
Start-Codeman.sh now compares the checkout's HEAD commit and
package-lock.json hash against a recorded marker
(docker-build-source.json) and clears just the affected volume(s)
before its own --build when either moved.
The in-place self-update path writes that same marker after a
successful build, so the two mechanisms agree on what the volumes
currently reflect — without it, the next plain Start-Codeman.sh run
would see the HEAD self-update just checked out, not recognise it as
already accounted for, and wipe the volumes self-update just correctly
rebuilt right back to the older baked image.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
Claude Code answers an exhausted model budget INSIDE the turn ("You've
reached your Fable limit. Run /usage-credits to continue or switch models
with /model.") and then sits there with nothing to write. The reviewer never
produces a report, so `runTurn` waited out its full 40-minute deadline and
reported a bare "timed out after 40 min without a report", which reads as a
hung reviewer rather than an account that needs attention.
Measured on 2026-09-08: #388, #393, #394 and #377 each lost 40 minutes this
way, and because every attempt counted, all four reached MAX_AUTO_RETRIES and
would NOT have been picked up again once the budget returned. One spent
afternoon quietly took the whole queue out of service.
`findModelLimitNotice()` reads the notice off the pane and `runTurn` returns
a new `limit` outcome instead of waiting. It is consulted in exactly two
places, both of which mean "the turn produced nothing": on a stop where
`isDone()` is still false, and on each timed-out wait slice. A review that
merely discusses usage limits in its own findings therefore cannot be
mistaken for one that hit the wall, and the pattern matches neither the model
name nor a straight apostrophe, since the pane renders a typographic one and
every model prints the same sentence.
A spent budget is an account condition, not a bad PR, so it no longer spends
the per-head retry budget: the queue resumes by itself when the budget does.
Telegram now names the cause and the file to change.
Tests use the pane captured verbatim off the run that lost the 40 minutes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Android/GBoard-style keyboards fire keydown with keyCode 229 and, on some
paths, never mutate xterm's helper textarea. xterm has nothing to diff, so
it emits no data and the typed character is silently dropped: it never
reaches the PTY and never appears on screen.
terminal-keycode229-recovery.js is a standalone controller that re-emits
exactly those keys, and only once. xterm stays authoritative throughout:
- Only an explicit keyCode 229 keydown carrying a single printable key (or
Enter) is eligible; Process/Unidentified/Dead, modifiers, AltGraph and a
live composition are all left alone.
- The re-emit is scheduled from a microtask and then a zero-delay timer, so
xterm's own textarea diff always gets the first opportunity; canonical
data for the same key cancels the pending fallback.
- compositionstart and blur drop every pending candidate, so a real IME
composition lifecycle is never second-guessed.
- After a recovery, one late canonical value attributed to that key token
(via beforeinput/input on the helper textarea) is suppressed so the
character cannot be delivered twice; the record expires after 250ms and
an unattributed byte is never suppressed.
terminal-ui.js wires it at the two existing choke points — the custom key
handler and the onData registration, the latter now a named handler so the
recovery path can re-enter it — with both hooks wrapped so a failure in the
fallback can never break canonical input.
Unit coverage drives the module directly in a vm; the wiring itself is
covered end-to-end in the (browser-only) terminal-copy-shortcut suite.
Observed on the first live merge (#383 via the Telegram button): runConfirmed
announced the merge and the scan five seconds later announced it again as a
closed PR. The scan now stays quiet for PRs the bot itself merged or closed,
and a merge of a `merge-with-fixes` verdict reminds that merging applies none
of the listed fixes.
A failed review used to be re-queued on every scan with no limit (two PRs
failed once each and were retried fine, but a head that keeps failing would
cost a session every ten minutes forever): three failures on one head now
stop the automatic retries until /review N or a new push. The failure notice
carries the reviewer's last message, so "finished without writing
report.json" says what it wrote instead.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Maintainer tooling in scripts/pr-bot/: a daemon (systemd user unit
codeman-pr-bot) that lists open PRs with gh, reviews each head commit once in
a Codeman claude session (`prbot-<n>`) running in a private `git clone
--shared`, and sends the verdict, ranked findings, checks and a recommendation
to Telegram with action buttons. Merge, close, post-comment and approve-CI
happen only from a Telegram command or button plus a confirmation tap; the
bot never writes to GitHub on its own. The Telegram token and chat id come
from the existing notifier bot's env file.
Verified live: three PRs reviewed end to end (383, 363, 368), reports
delivered with buttons, reviewer sessions on the pinned model. Findings
along the way, each fixed and documented: a linked worktree inherits the
main checkout's model pin (hence the shared clone), undici's 5-minute header
timeout cut off the first review, gh was missing from the service PATH, and
the periodic scan orphaned an in-flight review's record.
typecheck/lint/format now cover scripts/pr-bot; tests in
test/pr-bot-{report,state,commands}.test.ts; guide in docs/pr-bot.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.
The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.
Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.
Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.
Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:
- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
container on the new dist/. This is the one supervisor whose updater does NOT
outlive the restart, which is safe only because the terminal "restarting"
marker is written first.
- node_modules and dist are named volumes over the bind mount, so
container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
`npm run build` is tsc + esbuild and node-pty has no Linux prebuild.
An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.
The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.
Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.
Documented in docs/docker-self-update.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
scripts/tmux-chooser.sh is deleted. codeman tui replaces it and does the
job better: sc numbered its entries globally but only accepted a single
[1-9] keypress, so sessions 10+ were listed and unselectable, and it
inferred nothing about what an agent was doing. The tui carries the
server's real states, answers permission dialogs, and leaves an attach
with one key.
install.sh no longer creates the tmux-chooser symlink or the sc alias.
It now sweeps both up instead, on update AND uninstall, so an update
cannot leave a symlink pointing at a script this version stopped
shipping. The alias removal is marker-owned: it matches the exact line
the installer wrote, so someone's own 'alias sc=' for another tool is
never touched, and it rewrites through 'cat >' so the profile keeps its
mode and ownership. Verified against three profile shapes.
BREAKING CHANGE: the 'sc' command and the 'tmux-chooser' symlink are
gone. Use 'codeman tui' (and 'codeman tui --list' / 'codeman tui <n>').
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits,
authorship preserved) and adapts it across the 211 commits master gained
since the branch was cut:
- App Settings control re-authored for the set-* surface (PR #278): a
set-row in Layout -> Tabs, replacing the old settings-item markup the
branch targeted. i18n description synced.
- Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout:
computeLineagePath()'s U-bridge geometry hangs from the horizontal
strip's bottom edge and has no meaning against a vertical list. The
lineage strip-scroll listener now also redraws subagent/ultracode
connectors while the sidebar scrolls vertically.
- The desktop home tab rail (post-branch) defers to the sidebar: both dock
the session list flush left, and the rail would render z-ordered under it.
- Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on
master after the branch): sidebar mode branches to scrollIntoView
block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside
the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll
sidebar back to the top.
- Mobile active-tab hoisting the branch guarded against no longer exists on
master (removed by #257); kept master's order-stable render.
Verified: typecheck, lint, format:check, check:frontend-syntax,
check:public-assets, PostCSS parse of both merged stylesheets, the 26 new
jsdom tests, the structural guard suites, and the headless-Chromium harness
(scripts/verify-session-sidebar.mts) green across all seven layout states
at 1600/1000/393px against current master.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Second review round on the #241 follow-ups. Three defects in my own previous commit,
each reproduced before and after.
1. The temp path was shared between runs (`${dest}.tmp`), so two concurrent runs
fought over it: 4 of 4 concurrent pairs had one run die. Worse than a crash, a
sibling's cleanup landing between the esbuild and the alias append makes
`appendFileSync` CREATE the file, so the rename publishes a bundle-less file
containing only the alias tail, which still satisfies the content check and
would be blessed by the cache forever. The name now carries the owning pid.
8 concurrent pairs afterwards: no failures, no strays, aliases intact.
2. The content check only covered the bundle, so a truncated xterm.min.js with a
fresh mtime stayed truncated. This script can no longer produce one, but
postinstall.js writes the same directory in place, so a Ctrl+C during
`npm install` does, and a 200-byte xterm.min.js means `Terminal` is undefined
and every mobile test dies on a null. A copy must now match its source byte for
byte, and a derived output must clear a floor far below the real ratios
(measured 0.97-1.00 minified, 0.51 for the bundle) while a truncation misses by
orders of magnitude. Verified: 200-byte and 50-byte poisonings both repaired.
3. The try block ended before the append and rename, so a rename failure leaked its
temp behind a raw stack. It now covers both and reports which asset failed.
Per-pid names mean a killed run's temp is never reclaimed by a later rebuild, so
startup sweeps temps whose owning process is gone, and only those: `kill(pid, 0)`
throwing ESRCH. Deleting a live run's temp would recreate the collision fix 1
removes. Verified both directions, plus SIGKILL mid-build leaving no litter. The
sweep swallows its own errors, because reclaiming litter must never fail the run:
a directory named like a dead temp otherwise crashed the whole prepare step.
Security-reviewed: no shell (execFileSync with an array, `shell` unset), every
argument from the static asset table plus a numeric pid, all writes confined to the
vendor dir under strace, `process.kill` only ever with signal 0 (and pid 0 skipped,
since to kill(2) it means this process group), no new dependencies, no network, no
eval, nothing published. The emitted browser bundle is byte-identical to the one
scripts/build.mjs ships, tail included.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Follow-ups to #241 (thanks @Lint111), from an independent review of that PR. The
script is a real fix for a real gap; these are the four defects the review found,
each reproduced before and after.
1. A wrong-but-fresh output was never repaired. The zerolag bundle is finished by a
SECOND step (the alias append), so anything landing between esbuild and the
append is permanent: the file looks complete, carries a current mtime, and the
mtime-only cache reports "up to date" forever while the suite dies on
`LocalEchoOverlay is not defined`. Reproduced by replaying #241's own two
commits: running the first and then pulling the second kept the broken bundle.
Fixed twice over, because the two halves address different cases. Builds now go
to a temp file and `renameSync` into place, so this script can never publish a
half-written output (that also covers an interrupted esbuild or copy, and two
concurrent runs). And `isFresh` verifies the bundle actually contains its alias
tail, which is what repairs a file an EARLIER version already poisoned; a rename
alone cannot fix what is already on disk.
2. Freshness compared against the entry file only, but esbuild bundles its four
siblings too, so editing overlay-renderer.ts left the suite testing a stale
overlay while reporting "up to date". Editing those siblings is exactly the
single-source workflow CLAUDE.md mandates. It now stats every `.ts` in the
package source dir. A full rebuild is ~2s, so the cache was not buying much.
3. `execFileSync('npx', ...)` passed no cwd, unlike scripts/build.mjs, so a run from
another directory missed the repo's pinned esbuild and would fetch an unpinned
one from the registry. Both calls now pass `cwd: ROOT`.
4. Every invocation in test/mobile/README.md was a bare `npx vitest`, which skips
the `pretest:mobile` hook npm only fires for `npm run test:mobile`, so the
documented commands all bypassed the fix. Rewritten, with a note on why.
Also: an esbuild failure printed a raw stack; it now names the asset and its input,
matching the missing-input message. And the header comment no longer implies the
vendor dir is always empty: scripts/postinstall.js already writes these same seven
outputs, so what this script adds is freshness and independence from install time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
With a real codex login now available, record the one shape the fake-key
lab could never produce: a genuine model reply streaming above the pinned
composer, pushing lines into history (baseY grows) while keystrokes land
mid-stream. The recorder gains an opt-in CODEX_RECORD_REAL=1 scenario
using the user's own ~/.codex (fixture secret-scanned for key/JWT
material before writing; scanned clean). The replay test pins: baseY > 0,
mid-stream predictions painted, exact convergence to the typed text.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The zerolag bundle exports only `XtermZerolagInput`, but app.js constructs
`new LocalEchoOverlay(terminal)` directly. scripts/build.mjs appends global
aliases after esbuild (build.mjs:53-66); the first version of this script
omitted that step.
Without them initTerminal() throws `LocalEchoOverlay is not defined` at the
line that builds the overlay — and because that is midway through the function,
EVERY later step silently never runs, including the mobile touch handlers on
#terminalContainer. The page still had a terminal, so the failure looked like a
tap-routing bug rather than a boot error.
Verified: boot errors none, and all four terminalContainer touch listeners
(touchstart/touchmove/touchend/touchcancel) now register.
The mobile suite drives a real browser against a WebServer started from
TypeScript source, so fastify-static serves join(__dirname, 'public') =
src/web/public — not dist/web/public, where `npm run build` puts the vendor
bundles. Every /vendor/xterm* request 404s, so `Terminal` is never defined,
initTerminal() never runs, and any test touching app.terminal dies with
"Cannot read properties of null".
Measured in one worktree, toggling only the vendor files:
before: 404s=5 Terminal=undefined app.terminal=null 8 failed | 26 passed
after: 404s=0 Terminal=function app.terminal=live 6 failed | 28 passed
The 6 remaining failures are genuine pre-existing bugs (stale layout and
accessory-bar expectations, a CJK timeout) and are left alone here.
This went unnoticed because config/vitest.ci.config.ts excludes test/mobile/**,
so CI never ran the suite. `npm run test:mobile` now runs it, with a pretest
hook that builds the bundles.
The asset list was derived from the actual 404s rather than from build.mjs —
which is how xterm-addon-unicode11 and xterm-zerolag-input got included; reading
the build file alone would have missed both. Outputs go to the gitignored
src/web/public/vendor/, so they stay build artifacts. The script is idempotent
(skips outputs newer than their source) and does not touch the normal build.
Full CI suite unchanged: 4368 passed.
postinstall + build.mjs build vendor/xterm-predictive-echo.js as a SEPARATE
IIFE (window.PredictiveEchoAddon + self-activating PredictiveEchoOverlay);
the zerolag bundle command is untouched and its output verified
sha256-identical. index.html loads it after the zerolag tag (cacheBustAssets
covers it), sw.js precaches it, build.mjs HASHABLE content-hashes it.
A missing or broken bundle degrades codex to plain PTY echo.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Recorder (scripts/dev/record-codex-frames.mjs) captures real codex 0.147
TUI output through the production pipeline (tmux status-off + the codex-mode
full strip from session.ts) into JSONL fixtures with keystroke injection
points; analyzer replays them through @xterm/headless for the measurements
in docs/predictive-echo-plan.md. Composer signature /^> /-style (U+203A),
modal and wrapped rows correctly rejected, echo is unstyled default-fg,
tmux delivers echo as minimal in-place deltas.
Package: types.ts gains optional cursorX/cursorY, getCell, onWriteParsed,
onResize (all additive); prediction-renderer.ts renders per-glyph spans
keyed by prediction seq; @xterm/headless@^6.0.0 devDep for replay tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The synthetic session names end up in the harness screenshots, so shipping one
contributor's project list into everyone else's review reads oddly. The mix of
CLI modes is what the fixture actually needs — each renders a different badge —
and that is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The script carried two absolute paths from the machine it was written on: a full
scratchpad path including a session UUID, and /home/chaberl/projects as the
synthetic sessions' working directory. This branch is pushed to a public fork, so
they were visible to anyone.
Screenshot output now defaults to tmpdir() and is overridable via
SIDEBAR_SHOTS_DIR; the synthetic working directories are tmpdir()-based too, which
also makes the harness run for anyone who checks the branch out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The header tab strip stops working past roughly a dozen sessions: it wraps
into two or three rows, eats vertical space and still cannot be scanned.
This adds a vertical session list in a left <aside> as an ALTERNATIVE
layout — a filter box, a live count, and a 44px collapsed rail that keeps
the ambient signal (status dot, task badge) visible.
The strip is not removed. Settings -> Display -> Tab Bar -> Session List
Layout switches between them and the default stays 'header', so existing
users see no change until they opt in.
Structure: one #sessionTabs element, two mount points. applySessionListLayout()
re-parents the SAME node between #sessionTabsHost and #sessionSidebarList,
which is why there is no second renderer and no duplicated wiring — app.$()
caches getElementById results and never invalidates them, so a moved node
keeps every existing consumer (settings-ui, webview-tabs, the generated
gesture bundle, the mobile tests) working untouched.
Notable integration points:
- Below 1024px the sidebar is an off-canvas drawer overlaying the terminal;
closed it gets inert + aria-hidden so it cannot be tabbed into, and touch
swipes over it no longer switch sessions.
- Subagent and ultracode windows anchor to the right edge of a sidebar row
instead of its bottom, connector curves follow.
- Alt+B toggles; the chord is gated out of the PTY so xterm cannot also
write ESC b into a live session.
- Collapse state lives in its own localStorage key (the settings blob is
rebuilt from DOM controls on every save) and falls back to in-memory
intent where storage throws.
Verified: frontend syntax + public asset checks, tsc, eslint, 26 new jsdom
tests, and a headless-Chromium harness (scripts/verify-session-sidebar.mts)
that renders a synthetic 25-session fleet in both layouts at 1600/1000/393px
and asserts mount point, widths, inert/aria state and row count.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).
Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.
Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The default desktop header right cluster is now: WS, CPU, MEM, File Viewer
folder button, 5H/7D plan-usage chips, gear. The token-count chip and the
lifecycle-log document button default OFF (both still honor stored prefs),
and the File Viewer button defaults ON (phones keep hiding it via mobile.css).
Templates ship the hidden/shown state so nothing flashes before settings load.
Capture scripts seed showTokenCount:false so screenshots match regardless of
server defaults.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Docker cases: seamless Claude auth (seed ~/.claude.json instead of the
corruption-prone single-file mount), full credential-store isolation for
claude + codex/gemini/gcloud/opencode (share only transcripts/rollouts,
seed the rest), auto-build the base image on first use, C.UTF-8 locale
(fixes box-drawing), collapsed/shortened Create-Case UI + short "(docker)"
case-menu tags, and w<n>-<case> tab naming for docker/remote sessions.
Also: opt-in File Viewer header button; fix a TZ-boundary flaky test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude Code changed its subagent on-disk format (~2026-06-14): TUI Task
subagents now write `agent-{id}.meta.json` ({agentType,description,toolUseId})
into the session's `subagents/` dir and no longer reliably write a per-agent
`agent-{id}.jsonl` transcript there. The watcher discovered agents ONLY by
`.jsonl`, so it tracked zero — subagent windows and the monitor's "N TRACKED"
showed nothing.
- Add `registerAgentMeta()`: discover from the meta sidecar (description from
meta.description/agentType), prefer a sibling `.jsonl` transcript when present
(richer), never tail a meta file.
- Initial scan + directory watcher now handle `.meta.json` alongside `.jsonl`.
- Tests: 2 new cases (meta-only discovery; prefer-.jsonl-when-present).
Verified e2e against a real ~/.claude/projects fixture.
Known follow-ups (not in scope): meta-only agents have no per-agent transcript
to tail (no live tool-call feed, status stays 'active'); workflow agents under
`subagents/workflows/{wf}/agent-*.jsonl` are still missed by the flat scan.
Also adds the README screenshot tooling used to surface this:
- capture-real-overview.mjs: DSF=2 + ?nowebgl crisp path (DOM renderer avoids
the WebGL glyph-doubling at deviceScaleFactor>1).
- capture-readme-real.mjs: real-instance desktop-scene capture (dashboard/
monitor/subagent) for an isolated beta seeded from prod settings.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
scripts/capture-real-overview.mjs:
- Default deviceScaleFactor to 1 (DSF=2 makes xterm's headless WebGL renderer
draw console glyphs at ~2x while reporting nominal cell dims — invisible to
cols/cell measurement, only the pixels reveal it; HTML chrome is unaffected so
only the terminal font looks oversized)
- Mint a unique timestamped filename per run so a viewer/HTTP cache can't shadow
a fresh capture with a stale render of a fixed path
- Seed per-device localStorage (skin, codeman-font-size, codeman-app-settings)
so the capture reflects a real device: plan-usage chip shown (per-device key,
deleted from server payload), side panels closed for a full-width terminal
- Support prod's self-signed HTTPS (ignoreHTTPSErrors), env-configurable viewport
CLAUDE.md:
- Document the DSF=1 / unique-filename screenshot gotcha (incl. the real
Codeman-side immutable-static-asset cache footgun)
- Add the sanitize-html.js infra module (DOMPurify mXSS allowlist, COD-56) to the
frontend module list and load order (was missing)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The previous _sanitizeHtml was a denylist over agent/transcript markdown
rendered via innerHTML; it missed style attributes and the svg/math mXSS
namespaces — e.g. <svg><style><img src=x onerror=alert(1)></style></svg>
re-serialized into a live <img onerror>.
Vendor DOMPurify 3.4.8 (allowlist) following the existing marked.min.js
vendor pattern (same-origin, CSP script-src 'self'; not in package.json so
no lockfile drift). New sanitize-html.js wires a hardened allowlist config
(FORBID style/svg/math/script/iframe/object/embed/form; no data attrs);
app.js _sanitizeHtml delegates to it with a fail-closed escape-all fallback.
index.html loads dompurify -> sanitize-html -> app.js (defer); build.mjs
minifies + content-hashes sanitize-html.js.
Test: test/markdown-sanitizer.test.ts (jsdom, real shipping artifacts) —
mXSS payloads neutralized + legit markdown preserved.
Switching away from a session and back replayed only the server's byte
history. For TUI modes (codex especially) that shows just the latest
repaint — the idle banner — because the TUI drops earlier conversation
from its current frame. This restores the actual on-screen view.
Two complementary mechanisms:
- Client: load xterm's SerializeAddon and snapshot the rendered state
(viewport + scrollback + colors) per session on switch-away, restoring
it for an instant first paint on switch-back. The snapshot is only the
first paint — the canonical /terminal frame is still fetched and
reconciled (restoredSnapshot/clearedForBusy force the replay). Snapshots
are LRU-bounded in memory (<=20) and persisted to localStorage
(<=256KB each, <=10 sessions, stale-pruned) so they survive tab discard.
- Server: GET /api/sessions/:id/terminal prepends the live tmux pane
buffer (via the existing captureActivePaneBuffer) ahead of the byte
history, cleared between, so replay reflects the current frame.
Also fix formatPaneSnapshot dropping the rightmost column of every
captured row: it painted to cols - 1 out of caution about last-column
autowrap, but every row is followed by an absolute cursor-position CSI
that cancels xterm's pending-wrap, so painting the full width is safe.
The SerializeAddon is built from @xterm/addon-serialize (new dependency)
into the vendor bundle by postinstall.js (dev) and build.mjs (prod),
matching how the other xterm addons are vendored.