Compare commits

...
Author SHA1 Message Date
Codeman maintainer 6f7add7ce4 chore: version packages 2026-09-04 20:46:19 +02:00
Codeman maintainer eeb5f9d0b2 docs: web-tab egress guard, capability revocation and referrer policy
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:15 +02:00
Codeman maintainer 2ab21c1b32 fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring
claiming logout called it and no caller at all. The capability is a bearer
credential exempt from cookie auth with a rolling TTL refreshed on every use, so
a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose
referrer policy) stayed valid for as long as anything kept polling it.

- POST /api/logout revokes the caller's capabilities (all of them in single-user
  mode), the admin forced logout revokes the target user's, and user deletion
  revokes whatever that user had open. revokeOwner returns the count for the
  admin audit line.
- Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own
  policy is dropped: every URL inside the frame carries the capability, and a
  dashboard on no-referrer-when-downgrade or unsafe-url handed it to any
  third-party host it linked. Verified with Playwright that a sandboxed frame
  under an upstream `unsafe-url` sends no Referer to a third party while the
  root-absolute fetch and the CSS-triggered 404 fallback still reach the
  dashboard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:13 +02:00
Codeman maintainer 550e08a791 fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.

Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:

- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
  DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
  judges the RESOLVED addresses of a name and refuses when any is blocked. This
  is what closes DNS rebinding, which a hostname-string check cannot.

Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.

Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:12 +02:00
Codeman maintainer 99ad9cb236 fix(docker): never exit the server unless something is known to restart it
#373 restarts the Compose container by exiting the server, which is right for
the shipped deployment: `restart: unless-stopped` relaunches it. The updater
verified that policy through the Docker socket and, when it could not (no
socket mounted), failed open and exited anyway. Failing open is the correct
choice for the GATE, where refusing would block every install without a
socket, but not for the kill: a container the daemon does not restart goes
down for good, with no UI left to recover it from. That is exactly the case a
plain `docker run` of this image without `--restart` produces, and the image
sets CODEMAN_IN_CONTAINER=1 itself, so it takes the container path.

The decision now happens server-side, where both the socket and the Compose
env are reachable, and rides down to the script as `--restart-by-exit 0|1`.
It is 1 when the Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (added there
and only there, since that file is what sets the restart policy; the image ENV
deliberately does not) or when the daemon confirmed an auto-restart policy.
Otherwise the build still lands, the status becomes
`completed-needs-manual-restart` with the `docker restart` hint, and the
server keeps running. The shipped deployment is unchanged in effect: with the
socket it was already confirmed, and without it the declaration now covers it.

Also: a root-run `Start-Codeman.sh` (common on Unraid) created the
fingerprint baseline's `.codeman` directory before the container's first start
and left it root-owned, which the unprivileged server could then never write
its own state into. It is chowned to PUID:PGID when running as root.

Verified with a real image build of the merged tree (classic builder; this
box's BuildKit lacks buildx): runs as uid 1000, tsc/esbuild and the toolchain
present, the four CLIs at their pins, docker/.env absent, and `docker inspect
$HOSTNAME` returns the restart policy through the mounted socket as that user.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:36:35 +02:00
Codeman maintainer 823f56a243 Merge pull request #373 from opticon454/feature/docker-self-update
feat(docker): restore in-app self-update in the Compose dep
2026-09-04 14:25:56 +02:00
Codeman maintainer 72fd231d11 test(setup): one answer for CODEMAN_DATA_DIR, the strip from #371
#356 and #371 fixed the same leak two ways. #356 pointed CODEMAN_DATA_DIR at a
second throwaway directory and cleaned it up in afterAll and on exit; #371
deletes the variable along with CODEMAN_INSTANCE and CODEMAN_TMUX_SOCKET, so
`getDataDir()` falls back to `homedir()`, which the temp HOME already redirects.
Merged as they were, setup.ts set the variable and deleted it a few lines
later, and the second directory was created for nothing.

The strip wins: same protection, one tree to clean up, and the isolation test
#371 adds pins the list statically. The extra directory, its restore and its
two rmSync calls go, the vitest config `env` entries that set the same variable
go (they were documented as inert and would now be contradicted by the setup
file either way), the two test comments that described the old mechanism are
reworded, and CLAUDE.md's testing paragraph names the three stripped variables
and why CODEMAN_INSTANCE has to be stripped in the setup file rather than a hook.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 14:20:11 +02:00
Codeman maintainer 65d19c725e Merge pull request #371 from opticon454/fix/test-env-instance-isolation
fix(test): strip the instance-selection env vars in test/setup.ts
2026-09-04 14:20:04 +02:00
Codeman maintainer 80626567b2 chore: version packages 2026-09-04 14:01:20 +02:00
Codeman maintainer a81e87f440 fix(cli-registry): log why clis.json was ignored, and say 0600 when that is the rule
The loader refuses a `clis.json` with any group/world permission bit, read bits
included, so a file created with a normal umask (0644) is ignored. That is a
defensible posture for a file that chooses the binaries Codeman spawns, but two
things around it made the override feature look dead: the warning said
"group/world-writable", which a 0644 file is not, and `LoadResult.warnings` was
returned to a caller nobody wired up, so nothing anywhere printed it. A user
following the docs got silence.

The message now names the rule and the command that satisfies it, the loader
logs every warning once on first load (the result is memoized, so once per
process), the module header stops claiming that nothing ever writes (the
quarantine rename of a malformed file is a write, on first use) and the
registry doc gains a short section on the override file with the 0600
requirement in it. Whether the check should relax to writable bits only is a
separate decision; this keeps the shipped behaviour and makes it visible.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:51:20 +02:00
Codeman maintainer 2e0129f1f8 docs(test): name the real reason the suite could reach ~/.codeman
#356 stopped a bare suite run from overwriting the production
`remote-hosts.json` by pointing `CODEMAN_DATA_DIR` at a throwaway dir, and it
gated every case-tree delete on the temp HOME. Both changes are right; the
explanation written next to them is not. It says `os.homedir()` reads
/etc/passwd rather than `$HOME` on Linux, which would mean the temp HOME in
test/setup.ts never worked. It does: libuv checks the env var before the passwd
entry (measured: `HOME=/tmp/x node -e 'console.log(os.homedir())'` prints
/tmp/x), and CLAUDE.md's testing section relies on exactly that.

What bypasses the temp HOME is `CODEMAN_DATA_DIR` itself. `getDataDir()` reads
it as an absolute override before it looks at `homedir()`, so one inherited from
the shell (a second instance, a beta run) sends the whole suite at the real data
dir. That is the case setup.ts now closes, and #371 names the same variable from
the other direction.

The comments in setup.ts, the `safeRmHomeTree` helper, the voice-routes and
case-clone tests now say that, and the containment gate is described as what it
is: defense in depth. CLAUDE.md's testing paragraph gets the same note so the
next reader does not chase a homedir() bug that does not exist.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Codeman maintainer 28b44237ae fix(remote): classify the has-session probe by exit status, and forget it once the pane is back
#355 made the remote auto-reconnect watcher revive a dead pane only when the
durable remote tmux session is verifiably still alive, which is the right rule:
a clean Ctrl-C / Ctrl-D / exit tears that session down and must never relaunch
a fresh agent. Its probe, though, read `has-session`'s stdout and treated an
empty string as "gone". `tmux has-session` prints NOTHING on success (measured
on a scratch socket: exit 0, empty stdout, the failure message goes to stderr),
so every live remote session classified as gone and transport-drop reconnects
were silently disabled along with the clean-exit revives.

The probe now goes by exit status through a pure, unit-tested mapping
(`classifyRemoteAliveExit`): 0 is alive; ssh's own 255, a timeout (`killed`,
no numeric code) and a spawn failure are unknown, which the watcher already
treats as do-not-revive; any other status is the remote command's and means
gone (tmux's 1 for a missing session, 127 when tmux is not installed there).

Two smaller things in the same area:

- The cached answer was never invalidated, so after one successful reattach a
  stale `true` would have revived the NEXT clean exit (the original bug back
  after the first transport drop), and a cached `false` from a clean exit would
  have left a manually restarted session with auto-reconnect permanently off.
  The tick now forgets the cache entry whenever the pane is seen alive.
- The fire-and-forget probe has a 15s timeout against a 5s tick, so an
  unreachable host stacked up to three ssh processes per dead session. An
  in-flight set caps it at one.

The probe command is pinned as a literal string, and the reattach-then-clean-exit
sequence is driven through the watcher in the tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qg6bcATm1pNNY4kQWGwzgu
2026-09-04 13:50:22 +02:00
Ark0N 96960785d2 Merge pull request #356 from timkjr/pr/test-isolation
fix(test): isolate route tests from the production ~/.codeman data dir
2026-09-04 13:50:03 +02:00
Ark0N ee6a7af1d1 Merge pull request #355 from timkjr/pr/remote-exit
fix(remote): never auto-revive a remote session after a clean agent exit
2026-09-04 13:49:49 +02:00
Ark0N 850b00572c Merge pull request #347 from opticon454/feature/cli-registry-core
PR A: CLI registry core as a pure internal refactor
2026-09-04 13:49:35 +02:00
timkjr 4a63ab1604 test: extend CASES_DIR containment guard to the rest of the suite
#356 introduced safeRmHomeTree/isUnderTestHome to stop tests from deleting
the PRODUCTION ~/codeman-cases tree on platforms where os.homedir() ignores
the $HOME override -- but only applied it to the one file caught doing it
live. CASES_DIR has no CODEMAN_DATA_DIR-style env override at all, so every
other test file's raw rmSync(join(CASES_DIR, ...)) was the same unguarded
pattern, just not yet triggered.

Routes every CASES_DIR delete in these 10 files through safeRmHomeTree:
cli-skill-target, edge-cases, integration-flows, operation-lightspeed,
ralph-integration, routes/case-clone-routes, routes/voice-routes,
session-cleanup, sse-events, sse-subscription-filter.

Also fixes one instance in case-clone-routes.test.ts that mkdirSync'd then
rmSync'd a CASES_DIR path directly with no guard at all -- the exact
clobbering pattern #356 exists to prevent, found by extending the sweep.

Held as a separate commit (and intended as a separate PR once #356 merges)
rather than folding into #356 -- keeps the already-checked skinny fix
reviewable on its own; this is the same bug class applied broadly, not new
functionality.

Verified: all 10 files pass (180 tests), npm run typecheck clean.
2026-09-02 20:41:24 -05:00
timkjr 4068c02b9e fix(test): write the remote-hosts fixture where the route actually reads it
The "never writes hooks for a remote attach" test stubbed CODEMAN_DATA_DIR
to a separate throwaway dir just for this write, but session-routes.ts's
CODEMAN_CONFIG_DIR is a module-load-time constant frozen at test/setup.ts's
sandboxed dir before this test ever runs. The fixture landed somewhere the
route handler could never read, so the remote-host lookup silently failed
(NOT_FOUND) and the test passed for the wrong reason -- createErrorResponse
never sets reply.code(), so Fastify's default 200 made the NOT_FOUND branch
and the intended success branch indistinguishable by status code alone.

Write straight to getDataDir() instead, matching the docker-hosts fixture
convention already used elsewhere in this file. Verified the fix actually
exercises the success path (host resolves, 200 with a real session), not
just an accidental 200 from the error branch.
2026-09-02 20:41:24 -05:00
timkjr ff88b6957e fix(test): guard the CASES_DIR delete + harden the data-dir teardown
PR #356 stopped the remote-hosts.json fixture write from clobbering prod.
Two holes in the same file remain:

1. The quick-start afterEach still ran rmSync(CASES_DIR, recursive).
   CASES_DIR is join(homedir(), 'codeman-cases'), and on Linux builds
   where os.homedir() reads /etc/passwd instead of $HOME it resolves to
   the PROD case tree - so a full-suite run deleted the real
   ~/codeman-cases. Add a shared safeRmHomeTree() containment gate that
   only deletes a path under the redirected test HOME.

2. setup.ts teardown did rmSync(process.env.CODEMAN_DATA_DIR ?? '') AFTER
   restoring the env - if a pre-existing prod CODEMAN_DATA_DIR was set,
   that deleted prod. Capture the throwaway dir in a const and clean that.

A broader test-isolation sweep (10 files: cli-skill-target, edge-cases,
integration-flows, operation-lightspeed, ralph-integration,
case-clone-routes, voice-routes, session-cleanup, sse-events,
sse-subscription-filter) also applies the same containment gates to every
per-case delete. It is intentionally NOT included here to keep this PR
skinny; it is identified and available on request.
2026-09-02 20:41:24 -05:00
timkjr 2694d3f74a fix(test): isolate route tests from the production ~/.codeman data dir
session-routes-workspace-hooks.test.ts wrote its h1/box/10.0.0.5 host
fixture into getDataDir()/remote-hosts.json. getDataDir() resolves via
homedir() → ~/.codeman (INSTANCE_SUFFIX='' by default), and overriding
HOME in test/setup.ts does NOT change os.homedir() on Linux — so every
full-suite run silently overwrote the PRODUCTION remote-hosts.json,
wiping user-defined remote hosts, emptying the launch-case dropdown and
breaking remote session creation (found live 2026-08-29).

The vitest v4 test.env config key is ignored (probe confirmed the
worker still saw CODEMAN_DATA_DIR=undefined), so the reliable fix is
stubbing the env inside the test: the fixture write now goes to a
throwaway /tmp dir via vi.stubEnv + finally unstub. Verified: prod
remote-hosts.json hash is identical before and after the suite run.
2026-09-02 20:40:58 -05:00
DevvynandClaude Opus 5 66eb01ba8f feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.

Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:

- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
  update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
  container on the new dist/. This is the one supervisor whose updater does NOT
  outlive the restart, which is safe only because the terminal "restarting"
  marker is written first.
- node_modules and dist are named volumes over the bind mount, so
  container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
  `npm run build` is tsc + esbuild and node-pty has no Linux prebuild.

An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.

The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.

Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.

Documented in docs/docker-self-update.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
2026-09-02 19:33:32 +08:00
Codeman maintainer 1e24817b51 chore: version packages 2026-09-02 10:49:36 +02:00
Ark0N f7cf15485e feat(models): offer Fable 5.1 in the model picker and task routing (#372)
Adds `claude-fable-5-1` to the App Settings model picker and the five task-routing selects, mirroring how Fable 5 is already offered: a base option with data-ctx="1" plus its [1m] companion row. No settings-ui.js logic change, since the cards and the 1M switch are built from those options.
2026-09-02 10:48:51 +02:00
DevvynandClaude Opus 5 1125f7c1c5 fix(test): strip the instance-selection env vars in test/setup.ts
`test/setup.ts` gives every test file a temp HOME so the suite cannot touch the
real Codeman tree, and strips the env vars that would leak past it — but the
list only covered auth and the gesture flag. The three vars
`src/config/instance.ts` derives the data dir and tmux socket from were missing,
and they reach past the temp HOME:

- **`CODEMAN_DATA_DIR` is the one that matters.** It is an ABSOLUTE override
  read in `getDataDir()`, so it bypasses HOME entirely: a developer who exports
  it — or a shell left over from `codeman web -d` — has the suite reading and
  WRITING their real `state.json`, `users.json`, `intents.json` and
  `hook-secret`.
- **`CODEMAN_INSTANCE`** moves the data dir to `~/.codeman-<name>` and the
  socket to `codeman-<name>`. Inside the temp HOME that is not data loss, but it
  silently changes the paths tests assert on — and `scripts/run-beta.sh` exports
  it, so any shell that has run a beta carries it.
- **`CODEMAN_TMUX_SOCKET`** renames the socket `resolveTmuxSocketName()`
  returns. `TmuxManager` no-ops its shell commands under vitest, so this is
  assertion drift rather than a stray `tmux -L` against prod — same class of
  leak, same one-line fix.

They are deleted in the setup file rather than in a hook because
`CODEMAN_INSTANCE` is captured into a module-level const the first time
`config/instance.ts` is imported; a `beforeEach` would already be too late.

`test/test-env-isolation.test.ts` pins the whole list in two halves, because the
obvious half is not enough: asserting the vars are unset passes trivially on a
machine that never set them, so a removed `delete` line would sail through on
almost every box and on CI. The static half reads `setup.ts` and asserts each
name is deleted there, which fails everywhere. An anti-drift check catches the
other direction — a var stripped in `setup.ts` but never given a reason in the
list — and is scoped to the strip section so the teardown's restores are not
mistaken for strips.

Verified by demonstrating the leak: with the `CODEMAN_DATA_DIR` line removed and
the var exported, the runtime assertion fails; with the line restored it passes.
Full suite: no new failures against an upstream/master baseline.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 09:49:09 +08:00
DevvynandClaude Opus 5 c5b84fb5f4 docs(cli-registry): annotate overlays.credStore as declared-for-later
Review item 4 named THREE live tables duplicating registry data. Two are now
read from the entry (`defaultRemoteCommandForMode`, `defaultDockerCommandForMode`);
the third, `resolveDockerCredentialArtifacts`, is not — and it was left neither
wired nor annotated, which is the state that item explicitly rules out.

It is not wired because the shape cannot express the live table: `credStore` is
ONE store per CLI, and `CRED_STORES` needs two for gemini (`.gemini` for the
CLI's own auth plus `.config/gcloud` for Vertex), while deepseek's entry declares
none at all even though `.dsh` is seeded. Wiring it means making the field an
array and correcting those two entries — a change to credential seeding, which
is at once the worst thing in that file to get wrong and the least covered by
tests, since every docker IO path is no-op'd under vitest. It belongs in its own
change, measured against a real container.

So it is annotated instead, at the field, in the type's declared-for-later
header, in docs/cli-registry.md, and in the pinned DECLARED_FOR_LATER list — the
last of which means wiring it later makes a test fail rather than leaving a
stale comment behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 09:35:55 +08:00
DevvynandClaude Opus 5 6acf0dea0f fix(cron): scope the launch pre-flight to launcher CLIs, not every mode
CI caught three cron-service failures. Both are mine, from converting cron's
per-mode ladders to capability reads without checking what each ladder's scope
actually was.

**The pre-flight.** cron only ever pre-flighted `deepseek` — dsh is a profile
LAUNCHER, so "installed" is not "runnable" and a bare `dsh` can boot a profile
that cannot drive a pane. I replaced that with an unscoped
`resolveCliLaunchError(mode)`, which pre-flights EVERY mode, so a claude cron
job on a box with no claude binary now failed with "Claude CLI not found"
instead of reaching tmux-manager's own throw. Three tests assert the latter.
It is now gated on `discovery.launcherProfile !== undefined`, which is
byte-identical to the `mode === 'deepseek'` check it replaces and generalises to
the next launcher. The equivalent HTTP-route conversion was already scoped (to
`capabilities.external`, matching what that route has always pre-flighted); I
simply failed to carry the same reasoning across.

**The model.** cron's ladder was `mode !== 'shell' && mode !== 'deepseek'`, and
I read it as `capabilities.model.source === 'claude-settings-file'` — which is
the HTTP route's question, not cron's. There, every external CLI reads its model
from its own config object earlier in the chain, so only claude reaches the
global default; cron has no such config, so the same expression silently
narrowed the default model from eight modes to one. Now `!== 'none'`, which is
exactly the two entries the ladder excluded. Not caught by a test — found by
re-deriving each ladder's scope after the first failure.

Also names a fourth deliberate behaviour change in the changeset, found while
tracing these: `session.ts` carried a hand-written list of modes with no
direct-PTY fallback and OMP was missing from it, though CLAUDE.md's own text
says "all eight require tmux". `requiresMux` comes off the entry now, so an omp
session whose mux creation fails refuses instead of silently starting outside
tmux.

Verified by diffing failing tests BY NAME against an upstream/master baseline,
rather than by file as before — which is how the regression slipped through: the
three new failures landed inside a file already failing for unrelated
Windows-path reasons, and the aggregate count happened to collide.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:49:04 +08:00
DevvynandClaude Opus 5 4830e662f9 refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.

Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.

Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.

OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.

Guard rails:

- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
  branching reappears outside `stock.ts`, in any of its four shapes (`===`,
  `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
  negated forms, which is how 36 of them survived an earlier pass. Every
  allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
  deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
  `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
  wire field is separate, bridged only by `legacyConfigAliases`. Getting
  `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
  clamp's only handle on a CLI's privilege switch, and a wrong name clamps
  nothing with no error and no failing test — so `schema.ts` rejects an entry
  naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
  `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
  thunks). A module-level const freezes at first import, so a CLI enabled while
  the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
  (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
  `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
  rather than measured. A test pins the list so it cannot quietly grow.

Three user-visible changes, all deliberate and named:

- `probeDockerCliVersion()` derives the in-container binary from the registry
  rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
  hardcoded map it replaces omitted while its own comment said the rule was
  "every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
  install hint is the install command rather than a docs URL, five CLIs gain
  hints they never had, and the row order follows the catalog.

Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:26:45 +08:00
Codeman maintainer 71ffbf18e4 chore: version packages 2026-09-01 21:55:00 +02:00
Codeman maintainer 826ddaa9aa build(docker): ship the Docker CLI in the Compose image, not the whole engine
docker/server.Dockerfile installed Debian's `docker.io` to get a client for the
socket mounted by Compose. That package is the full ENGINE: even with
--no-install-recommends it pulls 15 packages including containerd, runc, dmsetup
and iptables, none of which a container that only talks to a mounted socket can
use, and it ships Docker 20.10.24 (2023).

Copy the CLI and the buildx plugin from the official docker:29-cli image
instead. Measured on the same node:22-bookworm-slim base: 266 MB -> 108 MB, so
158 MB smaller with a current CLI (29.7.2) in place of a two-year-old one.

Three things verified rather than assumed, by building the real image and
running it:

- docker:cli is an ALPINE image, so copying a binary into this Debian one is
  only safe because the binaries are static Go builds (ldd: "Not a valid dynamic
  program"). In the built image, `docker --version`, `docker ps` and
  `docker build` all work against a mounted host socket as the unprivileged
  runtime user.
- buildx is copied on purpose. scripts/build-agent-image.mjs shells out to
  `docker build` and Codeman auto-builds the agent image on the first Docker
  case. Without the plugin that still works today — CLI 29 falls back to the
  classic builder, tested — but that builder is deprecated and will be dropped,
  so the plugin keeps the path supported.
- docker-compose is NOT copied: Codeman never shells out to it.

Pinned to the 29 major, matching how the base images here are pinned.
2026-09-01 21:54:56 +02:00
Codeman maintainer e2b72aafd7 chore: version packages 2026-09-01 11:32:50 +02:00
Codeman maintainer b15cc0eb1a fix(ui): keep the plan-usage chip's 5h slot when no session window is open
The header chip silently shrank from "5h 4% · 7d 52%" to a lone "7d 52%", which
reads as half the feature breaking rather than as an idle window.

Nothing was broken. Claude Code documents `rate_limits.five_hour` as "present
only while the API reports it and its resets_at has not passed", so between
5-hour session windows the key simply leaves the statusline payload. Codeman's
snapshot replaces the Claude half wholesale on every sample, so the segment
disappeared until usage opened a new window. Confirmed against a live 2.1.252
session by capturing real statusline payloads on an isolated tmux socket: the
boot render carries no `rate_limits` at all, and the first post-response render
carries both windows.

The slot now stays, with a dimmed em dash. Claude only: a missing CODEX bucket
means that plan has no such limit rather than an idle window, so those stay
omitted (pinned by the existing test). The placeholder can never stand alone
either — hasWindows() still gates the row, so a provider reporting nothing
renders nothing rather than a row of dashes. The tooltip says "5-hour limit: no
active session window" instead of dropping the line.

Verified in a real browser against a dev instance: the idle chip renders
"5h — · 7d 52%" with the dash at opacity 0.55 in --text-dim while the live
value keeps its green, and the chip holds its shape (100px idle vs 107px with
both windows).
2026-09-01 11:32:18 +02:00
Codeman maintainer 2a32b5064a Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).

Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.

Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:

- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
  the whole context-relative path, so `docker/.env` — which the deployment's own
  README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
  was picked up by `COPY . .` and baked into the image at
  /opt/codeman/docker/.env. Verified in both directions against a real build
  context: with a canary secret in docker/.env, the unfixed ignore file lets
  /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
  the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
  still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
  reported "Case not found" on exactly the deployment the override exists for.
  Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
  literal on purpose: that one migrates the historical ~/claudeman-cases
  directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
  three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
  joins the documented list of files that genuinely belong in the repo root.

The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.

Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
2026-09-01 11:32:03 +02:00
Codeman maintainer 0da0c8219d chore: version packages
1.24.2. Also corrects two numbers in the CLAUDE.md trust-dialog paragraph that
was written while the fix was still uncommitted: the keystroke cap is 6, not 3,
and the scan now schedules its own follow-up read rather than waiting on PTY
output that a static dialog never produces.
2026-09-01 02:31:31 +02:00
Codeman maintainer aaa93d4252 fix(session): answer Claude Code 2.1.252's reversed folder-trust dialog
Every claude session in a directory claude had not seen before died about six
seconds after it started (`Pane is dead (status 1)`), before the agent drew a
composer. Reproduced on a fresh case and measured.

Claude Code 2.1.252 rewrote the dialog. It used to be

  ❯ 1. Yes, I trust this folder
    2. No, exit

and is now unnumbered, reversed, and highlights the option that quits:

  ❯ No, exit
    Yes, I trust this folder

Detection still worked (the confirm affordance carries the match once the
numbered option text is gone), so the failure was entirely in the answer: the
auto-accept pressed Enter on the highlighted default, which is now exit.

trustDialogNextKey() reads the ❯ marker off the rendered pane and returns ONE
keystroke at a time: an arrow while the cursor is on the wrong option, Enter
only once the screen shows it on the trust option, and null for a frame that
does not say. Both layouts are handled, and which way the trust option lies is
read from the frame rather than assumed, so a further reordering costs a
repaint instead of a session. The last marked option wins, because the
direct-PTY fallback reads an append-only buffer where an older frame must not
out-vote the freshest one.

Two things only a live pane showed:

- The scan ran solely from the PTY onData handler. The arrow that moves the
  cursor is the last output the pane produces, so the first fix parked every
  session with the cursor sitting on the right option and no Enter ever sent.
  It now schedules its own follow-up read (_trustDialogTimer, cleared in
  _clearAllTimers()), offset past the scan throttle so the chain cannot break
  on a boundary.
- The keystroke cap goes 3 -> 6, since answering is no longer one press.

The bundled codeman skill had the same blind \r as its bounded fallback, so
preamble 1.21.0 replaces it with _trust_key/_accept_trust: read
terminal?full=1, steer onto the trust option, re-read, then confirm. Those
keystrokes go out under their own clientId, because input sequence numbers are
monotonic per client and spending prompt numbers on dialog keys would make the
next send-and-wait look like a stale duplicate and vanish while reporting
success. The readiness recipes in docs/extending-codeman.md,
docs/api-reference.md and the skill's own reference carry the corrected answer,
plus a symptom-table entry for a worker whose pane is dead seconds after spawn.

Verified live on an isolated instance (own data dir and tmux socket): fresh
case -> arrow at 5 s -> Enter at 7 s -> composer, with hasTrustDialogAccepted
recorded. With the server-side auto-accept disabled in a throwaway copy, the
skill's fallback cleared a genuinely parked dialog in 1.1 s and spawn_worker
took a brand-new case to a live composer in 7.2 s; spawn_workers + sendwait +
last_text then ran end to end.
2026-09-01 02:31:26 +02:00
Codeman maintainer 3518af3a9f docs: correct CLAUDE.md drift and document four undocumented subsystems
Audit of CLAUDE.md against the tree. Verified still accurate: the 31-module
frontend load order (matches index.html exactly), SSE registry parity at
157 = 157 (confirmed by running the parity test), config/ 21 files, types/ 22
domain files, 136 mobile device profiles, the version line, and every Quick
Reference command.

Drift corrected: 24 route modules to 25, ~220 handlers to ~227, system-routes
51 to 56, app.js ~5K lines to ~6.7K and 30 modules to 31, install.sh 92KB to
104KB. Completed the CLI resolver inventory, which was missing
deepseek-cli-resolver and omp-cli-resolver even though both modes are
documented, and named the shared cli-executable-resolver lookup chain.

Filled the gaps found by sweeping every src module against the file:

- Owner tab layouts (COD-359) had 6 source modules, 7 test files, 2 routes, an
  SSE event and a state.json key, with zero mentions anywhere in CLAUDE.md or
  docs/. The paragraph records the four things a reader would otherwise get
  wrong: it is backend-only as of 1.24.1 with no frontend consumer, the service
  is the sole mutation boundary, it projects onto PUT /api/session-order rather
  than replacing it, and reconciliation is gated on a successful restore.
- codeman doctor and codeman users were undocumented top-level CLI commands.
- Four subsystems whose invariants lived only in their @fileoverview:
  the workspace-trust dialog recognizer, proc-tree's bounded walk (the
  2026-07-30 incident that took a machine down), deepseek-web-server (one
  child process, deliberately not a shell session), and the Files panel
  search matcher (globs are never compiled to a RegExp).

Also fixes a stale "156 event types" comment in constants.js (actual: 157) and
a contradiction in AGENTS.md, which still carried the retired "never run the
full suite inside a managed tmux session" rule against CLAUDE.md's current
"npm test is the gate and is safe to run bare".

Note: the trust-dialog paragraph documents trustDialogNextKey(), which is part
of a sibling session's in-flight fix for the Claude Code 2.1.252 layout change
(unnumbered, reversed options with "No, exit" highlighted, so a blind carriage
return picks exit and kills the pane). That fix was uncommitted in the shared
tree when this landed, so the doc leads the code until it is committed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NMN8UuvdBim3iM87reuQ9Z
2026-09-01 02:10:09 +02:00
Codeman maintainer e5c5d890aa chore: version packages 2026-08-31 22:34:57 +02:00
Codeman maintainer d5b5f8f618 fix(docker): make the dsh profile install survive pnpm's build-script gate
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.

That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.

Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
2026-08-31 22:26:09 +02:00
Codeman maintainer 7762809202 Merge pull request #350 from opticon454/bugfix/dsh-pnpm
fix(docker): install pnpm for DeepSeek profile
2026-08-31 22:22:30 +02:00
Codeman maintainer 02bbf13b3c chore: version packages 2026-08-30 16:29:14 +02:00
Codeman maintainer da91b4353b Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend
2026-08-30 16:16:26 +02:00
timkjr da5f5447d0 fix(remote): never auto-revive a remote session after a clean agent exit
The COD-108 reconnect watcher treated any dead local pane as a dropped
transport and re-ran the pane command — so a normal ctrl-c/ctrl-d on a
remote claude/opencode/omp auto-spawned a FRESH agent (claude only
looked correct because its '--session-id || --resume' fallback resumed,
with a loud 'already in use' error first).

Distinguish a transport drop from an intentional exit: only reconnect
when the durable remote tmux session (codeman-ssh-*) is verifiably
still alive on the remote host. A clean exit tears that session down;
the watcher now probes it via ssh has-session and skips (remote-gone)
when it is gone OR unknown (fail closed). The probe is cached
per-session and fired async so the 5s tick never blocks on ssh.

Tests: 3 new cases pinning remote-gone / unknown / alive decisions.
Verified live: all remote CLIs stay dead after ctrl-c/ctrl-d.
2026-08-29 17:43:07 -05:00
timkjr b6d0f1fa32 fix(omp): wire OMP into install.sh's CLI detection (it had none)
Every other CLI (claude/opencode/codex/gemini/antigravity/pi/grok/dsh) has
a check_*/get_*_path pair wired into install.sh's detection loop and the
"no AI CLI found" aggregate checks. OMP had neither -- a user with only
omp installed would be told no CLI was found and offered to install
Claude Code or OpenCode.

Added OMP_SEARCH_PATHS (mirrors src/utils/omp-cli-resolver.ts's
OMP_SEARCH_DIRS) and check_omp()/get_omp_path(), wired into both
aggregate conditions (the interactive install-menu trigger and the
end-of-run reminder) and added omp's real vendor curl one-liner to the
reminder block. The DeepSeek Harness line was never in that reminder to
begin with -- confirmed it has no vendor one-liner (dsh installs via
Codeman's own API after the server is already up), so it stays out, with
an explanatory line instead.

Also fixed the "Skip" menu text, which was missing Gemini and DeepSeek
Harness from its example list independent of the omp gap, and the same
stale sibling-CLI-list bug (missing DeepSeek Harness and OMP, "the
eight"/"这七个") in README.md and the repo's existing README.zh-CN.md.
2026-08-28 15:19:19 -05:00
timkjr 65e994d29a fix(omp): correct docs/counts/URLs, resolver install-path order, stray comment + CSS
Small cleanup items from upstream review (Ark0N/Codeman#353):

- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
  installer target (~/.omp/bin was an earlier unverified guess, confirmed
  wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
  can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
  SessionMode incl. shell -- not eighth), matched the install-path guidance
  to the resolver fix, updated the version example to the actually-tested
  18.0.8, and added a Docker-section caveat: --resume pinning does not
  currently reach an in-container omp process, since Docker panes never see
  ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
  already linked to the -omp anchor, so the link was dead) and added an OMP
  specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
  Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
  and split two CSS lines that had two declarations jammed onto one line.
2026-08-28 14:37:28 -05:00
timkjr f18dccace1 fix: don't discard codex/gemini/antigravity conversations on Resume; fix DELETE ownership dup + missing broadcast
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.

DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
2026-08-28 14:03:07 -05:00
timkjr 2ee2eacb4b fix(omp): clamp OMP_AUTH_BROKER_URL/TOKEN, correct the env-allowlist docs
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.

The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
2026-08-28 13:45:12 -05:00
timkjr c4f6eb1e5e fix(omp): resolve and pin the respawn session id only at actual respawn time
findLatestOmpSessionId()'s newest-mtime pin ran eagerly inside
_buildRespawnPaneOptions(), which startInteractive() calls unconditionally
on every boot-recovery reattach — before anything checks whether the pane
is actually dead. With two omp tabs in the same case dir, this could pin
an ALIVE pane's session onto whichever sibling's file happened to be
newest on disk, purely as a side effect of building options that might
never lead to a respawn (reported in Ark0N/Codeman#353 review).

Move resolution out of the eager builder into _pinOmpRespawnId(), called
explicitly only where a respawn is actually confirmed: the dead-pane
branch in _setupOrAttachMuxSession() and reattachRemote(). Add
resolveAndClaimOmpSessionId(), which verifies each candidate's own file
header (cwd) rather than trusting the mangled-directory match alone, and
tracks claimed ids in a process-wide registry so two ambiguous resolutions
can't both pick the same sibling's conversation.
2026-08-28 13:18:25 -05:00
timkjrandClaude Sonnet 5 ab83d8ffec fix(omp): a fresh "Run OMP" click no longer silently resumes an old conversation
Found live 2026-08-27 by Tim: clicking Run OMP to start a brand-new session
in a case directory with prior omp history launched --resume <old-id>
instead of a clean `omp` invocation.

Root cause: Session._resolvedOmpRespawnConfig() resolves-and-pins the
newest on-disk omp conversation as a side effect on this._ompConfig. That
is correct when reattaching to an ALREADY-TRACKED mux session (a dead-pane
respawn, or a boot-recovery reattach - the constructor sets _muxSession
from persisted state before startInteractive() ever runs there), but it
ran unconditionally. startInteractive() computes
`respawnPaneOptions: this._buildRespawnPaneOptions()` eagerly in the same
object literal that builds `createSessionOptions.ompConfig: this._ompConfig`,
so for a genuinely brand-new session (no muxSession in its create config,
_muxSession still null) the resolve-and-pin side effect ran and poisoned
this._ompConfig before that field was even read.

Fix: gate the resolve-and-pin logic on `this._muxSession` already being
set. A fresh session has no muxSession yet and now passes through
untouched; a real reattach (muxSession present since construction) keeps
resolving and pinning exactly as before.

Verified live in production against the exact reported scenario (a fresh
omp session in a case dir with 8+ hours of prior omp history) - confirmed
both via the API (ompConfig stays empty, claudeSessionId equals the
session's own id) and visually in the GUI. Regression test constructs a
real Session + TmuxManager to exercise the actual private-method
interaction directly, since no existing test called startInteractive() at
all.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 1829fe91af docs(omp): add docs/omp-integration.md, matching sibling CLI docs
OMP was the one external CLI mode with no dedicated user-guide doc, unlike
opencode/pi/grok/deepseek which each have one. Covers install, auth (omp
owns its own entirely - no Codeman-side login flow or bypass switch),
what Codeman wires up (OmpConfig), the exact-id pinning mechanism and the
directory-mangling bug behind it, kill-survival via transcript scanning,
terminal behavior, Docker/remote-SSH cases, and known gaps (no idle hook,
mid-turn kill data loss, unverified symlinked-$HOME behavior).

Cross-referenced from README.md's Multi-CLI doc list and docs/docker-cases.md's
credential-seeding summary (which now also documents OMP's sessions/-is-shared
exception to the seed-everything pattern the other CLIs use).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 d74cde759b feat(omp): install omp in the docker agent image, isolate its credentials
OMP had full routing at the Docker layer (default pane command, schema) but
was never actually installed in docker/agent.Dockerfile, and had no
credential-isolation entry in docker-hosts.ts's CRED_STORES - a Docker-mode
OMP session would have failed with "omp: command not found", and even with
the binary present would have had no config/auth seeded, despite the README
already claiming OMP has "seamless auth, isolated credentials" in Docker.

- docker/agent.Dockerfile: install omp via its own installer (standalone
  binary, same shape as grok/antigravity - not on npm). Verified against a
  real --no-cache build: the installer actually targets ~/.local/bin, not
  ~/.omp/bin as the resolver's OMP_SEARCH_DIRS ordering would suggest -
  confirmed omp/18.0.8 installs and runs correctly inside the image.
- src/docker-hosts.ts: add a .omp/agent CRED_STORES entry. Unlike every
  sibling CLI in this family, sessions/ is SHARED (RW), not seeded: Codeman
  reads ~/.omp/agent/sessions/**/*.jsonl host-side for history recovery and
  --resume pinning (omp-transcript.ts, omp-session-resolver.ts), the same
  reason codex's sessions/ is shared rather than seeded. Seeding it instead
  would silently break the kill-survival feature for Docker cases. Only the
  small config files (config.yml/mcp.json/models.yml/settings.yml) are
  seeded; the SQLite caches and terminal-sessions/ stay container-local.
- test/docker-hosts.test.ts: pin the new CRED_STORES entry's behavior.

Found in passing (NOT fixed here, unrelated and pre-existing on master): the
agent image's DeepSeek (dsh) plugin-install step currently fails on a fresh
build ("pnpm not found on PATH"), confirmed via git diff against
origin/master that this line is untouched by this branch. Worth a separate
issue/PR.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 853681f970 harden(omp): resume-path test coverage, silent-fallback logging, cwd validation
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:

- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
  exported to make it testable: the exact "resume this OMP row from
  history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
  zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
  continuation silently degrades to omp's own ambiguous --continue,
  in both call sites (session create and respawn pinning) - previously
  silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
  directory in omp-transcript.ts's parser, so a corrupted/malformed
  session file can't point a downstream resume at a relative or empty
  path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
  case in mangleOmpWorkingDir(): the review's suggested realpath() fix
  assumes omp itself resolves symlinks before mangling, which is
  unconfirmed - guessing wrong there would trade one silent mismatch
  for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
  session-routes.ts (antigravity/opencode dynamic import line-wrapping)
  that was blocking the pre-commit formatting gate on this file.

Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 ed983f898b fix(omp): resolve claudeSessionId alias on boot-recovery reattach
Two bugs compounded to break continuation pinning on every real OMP
case (only /tmp-based manual testing happened to work by coincidence):

1. startInteractive() had a second, unconditional claudeSessionId
   assignment after the mux branch that clobbered its correctly
   resolved value back to the session's own id on every mux path.

2. mangleOmpWorkingDir() assumed omp mirrors Claude Code's directory
   naming (home prefix kept), but omp actually strips $HOME first.
   findLatestOmpSessionId() was silently returning null for every
   case under ~/codeman-cases/, so resumeSessionId never resolved for
   any real case dir - only /tmp paths (outside $HOME) worked, which
   is every dir this feature was previously tested against.

Verified live: killed and relaunched the omp-verify server process
mid-session (plain reattach, pane stayed alive) and confirmed
claudeSessionId now resolves to the real omp transcript uuid instead
of the Codeman session's own id.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjr 54a930c80e feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.

Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.

Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.

Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.

Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.

Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
2026-08-28 11:32:30 -05:00
timkjr 4c332c6141 fix(omp): retire the old row on resume, and let DELETE remove persisted-only sessions
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.

That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).

Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
2026-08-28 11:32:30 -05:00
timkjr 253599ce9c fix(omp): wire ompConfig into respawnPane and default to --continue there
respawnPane() -- the path used when a session's pane died (crash, idle
respawn, or the user's own /exit) but the Codeman session object is
still tracked -- never had ompConfig wired through at all, in either
its options destructure or its inner buildSpawnCommand() call. This is
a gap in the original OMP patch, distinct from the resumeHistorySession
fix (which only covers a session that has been fully closed and shows
up as a history row): reselecting a tab whose CLI process just exited
goes through this path instead, and always launched a bare, contextless
`omp` no matter what.

Beyond the wiring, respawning a dead pane is semantically different
from creating a brand-new session: the conversation is still "this
session" to the user, so _buildRespawnPaneOptions() now defaults
ompConfig to continueSession:true unless the session already carries
an explicit resumeSessionId (which still wins in buildOmpCommand).

Verified live: told a session a secret, exited OMP so the pane died
(session and tmux both left alone), forced the exact dead-pane-respawn
path, and the new process replied with the secret -- confirming
`omp --continue` fired instead of a blank omp.
2026-08-28 11:32:30 -05:00
timkjr 3e1a0e679f fix(omp): resume by mode, not silently as claude, and support --continue
resumeHistorySession() never sent mode when recreating a session from a
history/session-manager row, so the server default silently opened a
plain Claude session for every non-claude row -- reproduced live: OMP
rows spawned Claude sessions on click. Thread the row's mode through
every call site (welcome list, session manager, mobile overview) and
only send the Claude-specific resumeSessionId for claude rows.

Codeman has no live PTY-reattach outside server boot, and it's moot for
OMP anyway (exiting it kills the pane's only process), so route the
non-claude relaunch through each CLI's own continue-most-recent flag
instead of a context-free fresh start. OMP never got one: buildOmpCommand
only implemented --model/--resume despite omp --help documenting
-c/--continue. Added continueSession to OmpConfig end-to-end (type,
schema, builder) mirroring the existing opencode/pi/grok/deepseek
fields, and wired resumeHistorySession to use it.

Verified live: told a real omp session a secret, exited it, closed the
tab without killing tmux, relaunched with --continue in the same
directory, and had it recall the secret.
2026-08-28 11:32:30 -05:00
timkjr 7ec48adcc8 fix(omp): keep external-CLI mode enumerations complete in skill docs
Two prose lists in skills/codeman/ named some but not all external CLI
modes after the omp-mode rebase, which is exactly the drift
test/agent-skill-mode-lists.test.ts exists to catch: SKILL.md's
no-hook-signals list was missing omp, and endpoints.md's version-probe
sentence named pi/grok/omp as a bare 3-mode run with no matching class.
2026-08-28 11:32:30 -05:00
timkjr 9841f4ffb9 refactor(omp): align omp resolver + doctor with upstream shared CLI resolver
- omp-cli-resolver.ts already uses createCliExecutableResolver; add dedicated
  test/omp-cli-resolver.test.ts mirroring pi's (version-probe accept/reject,
  negative-cache backoff, VITEST hermeticity gate)
- dependency-registry omp entry now requires OMP_VERSION_REGEX match like pi,
  so codeman doctor and the run-mode resolver agree on what counts as installed
- system-routes /api/omp/status surfaces version
2026-08-28 11:32:30 -05:00
Codeman maintainer d8688dc143 fix(web): drop the provider label from the plan-usage chip when there is only one
The chip prefixes every row with the provider name, so a machine that only
has Claude limits renders "CLAUDE 5H 60% 7D 23%" — a 46px label naming the
only thing it could possibly be. The name exists to tell two rows apart, so
it should only appear when there are two.

updatePlanUsageChip() now checks whether both Claude and Codex actually have
windows before building the rows, and emits the .pu-provider span only in
that case. The tooltip keeps naming the provider in both cases: it has the
room, and the chip no longer does.

Verified in a browser on an isolated beta instance: Claude-only renders bare
windows with no .pu-provider in the DOM, Codex-only the same, and the
two-provider chip is byte-identical to before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 13:59:40 +02:00
Codeman maintainer 23fae0c5af chore: version packages
Codex plan usage in the header chip (#346), a visible inline rename in
the session sidebar (#345), and the install.sh Tailscale re-run fix plus
the README network-access prompt description.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:25:32 +02:00
Ark0N e4699159e9 Merge pull request #346 from JackStuart/codex/show-codex-usage-limits
feat(web): show Codex plan usage in header
2026-08-28 00:14:54 +02:00
Ark0N da085f5f7f Merge pull request #345 from fibr/fix/sidebar-inline-rename
fix(ui): show inline rename text in session sidebar
2026-08-28 00:14:48 +02:00
Codeman maintainer 23e32b22d5 docs(readme): describe the actual three-way network-access prompt
The installer bullet still described a two-way choice with 0.0.0.0 as "the
default", which predates the Tailscale option. The prompt has offered three
choices for a while (Tailscale / any device on your network / this machine
only), and the default is computed from what is already on the machine rather
than being fixed at 0.0.0.0.

Now states all three options, that the Tailscale one is a loopback bind
fronted by `tailscale serve` with the tailnet as the login, and how the
highlighted default is chosen. Line 220 already documented the Tailscale
option correctly; this was the only stale spot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 00:02:33 +02:00
Devvyn 26b4ffbb0f fix(docker): install pnpm for DeepSeek profile 2026-08-27 20:47:57 +08:00
Devvyn e2179bd530 chore(docker): remove local handover references 2026-08-27 19:42:37 +08:00
Devvyn b85f7659b7 feat(docker): add Compose deployment support 2026-08-27 19:38:38 +08:00
timkjr e82380e14a fix(ui): close unclosed CSS blocks that killed the stylesheet tail
The rebase hand-repair dropped the closing brace of .welcome-btn-pi:hover
and .btn-toolbar.btn-run.mode-pi:hover before the inserted OMP rules.
The browser CSS parser drops every rule after an unclosed block, so the
deployed UI rendered as unstyled text bars (only ~456 of ~2583 rules
applied). Verified clean via esbuild --minify (no css-syntax-error) and
rebuilt dist.
2026-08-26 20:10:34 -05:00
timkjr c0423bf560 fix(omp): complete omp wiring in UI files, skill docs, and tests after rebase 2026-08-26 20:10:34 -05:00
timkjr 4f5678fac4 feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) 2026-08-26 20:05:48 -05:00
Codeman maintainer 7dfb4acf24 fix(install): offer Tailscale setup on re-run instead of losing it to a failed build
The network-access prompt, where Tailscale serve is configured, runs AFTER
the build step. A build failure therefore exits before the question is ever
asked, and a user who then finishes the build by hand (rather than re-running
install.sh) ends up with a healthy loopback-only Codeman, a connected
Tailscale, and no serve mapping — with nothing anywhere pointing at
`install.sh tailscale`, the command that fixes it. Reported from a fresh
Ubuntu 24 install that died on the node-pty compile.

- maybe_offer_tailscale_repair(): on the update/re-run path, detect exactly
  that state (loopback bind + tailscale Running + no serve mapping fronting
  Codeman) and offer the retrofit. Silent for a deliberate non-loopback bind,
  silent once a mapping exists, silent when tailscale is absent, and prints
  the command instead of prompting when non-interactive. Returns 0 even when
  setup fails so it can never abort an update.
- print_security_notice(): the loopback branch now names
  `install.sh tailscale` when Tailscale is installed on the box, rather than
  the generic "tailscale serve / cloudflared tunnel" advice.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:29:15 +02:00
Codeman maintainer d3f851a5e5 chore: version packages
install.sh installs a build toolchain on Linux (node-pty has no Linux
prebuild, so a stock Ubuntu 24 server died inside node-gyp with
"not found: make"), plus review hardening for #339: the write-queue
reset paths now release the one-chunk-in-flight gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-26 18:04:44 +02:00
Ark0N 5f8d4de443 Merge pull request #340 from aakhter/pr/cod-341-file-viewer-search
feat(file-viewer): COD-341 search the full workspace
2026-08-26 18:03:21 +02:00
Ark0N 00b32ad2b8 Merge pull request #339 from dignfei/fix/terminal-live-write-backpressure
fix(terminal): bound live xterm backpressure
2026-08-26 18:03:14 +02:00
Jack Stuart b00ab3ceea feat(web): show Codex plan usage in header 2026-08-26 18:43:52 +08:00
Sergei Lupashin 134e200aec fix(ui): show inline rename text in session sidebar 2026-08-25 19:41:48 +02:00
Codeman maintainer a51563ce1f chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:14:41 +02:00
Ark0N ca5fe1ab3e Merge pull request #338 from Ark0N/feat/vertical-rail-detailed-rows
Vertical tab rail: detailed rows (created / working / status), plus a rename-cancel fix
2026-08-25 19:12:57 +02:00
Ark0N 9cd10afdc9 Merge pull request #341 from Ark0N/feat/deepseek-agent-workers
Spawn and drive DeepSeek Harness workers from the codeman agent skill
2026-08-25 19:12:45 +02:00
Ark0N 975705ad87 Merge pull request #337 from Ark0N/feat/deepseek-harness
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
2026-08-25 19:12:01 +02:00
Codeman maintainer 93a1042bb3 Merge remote-tracking branch 'origin/feat/deepseek-harness' into feat/deepseek-agent-workers
# Conflicts:
#	CLAUDE.md
2026-08-25 19:02:34 +02:00
Codeman maintainer a628737d1f fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first:

- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
  _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
  every dsh pane and applyEnvOverrides() lands after it, so a non-granted
  owner who could redirect the base URL would have the operator's key sent
  as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
  dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
  The HERDR triple is set via LOCAL tmux setenv, which crosses neither
  docker exec nor ssh, so such a session can never post a hook event and
  the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
  route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
  option parser cannot read a third-party TUI's frames, so an answer was a
  blind keystroke into a foreign composer), and the push notification
  carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
  reports inside a 60s window (the TUI retries with backoff, so a retried
  'working' could land after 'blocked' and resolve an approval whose
  dialog was still on screen); 4xx responses exit 0 instead of retrying,
  so one misconfigured session cannot feed the auth rate-limit bucket
  until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
  racing POSTs used to pick the same port and orphan the winner), and the
  readiness poll / timeout paths only clear or stop the singleton while it
  is still theirs. First click actually opens the tab now
  (refreshWebviews, not the nonexistent loadWebviews). DELETE
  /api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
  create paths (impl moved into the resolver so all three share it) and no
  longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
  the remote branch; the Ralph auto-enable list gained deepseek;
  HookEventType gained agent_working; the phone overview run menu filters
  managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
  child that reads stdin eats the rest of the script), bounds the exec
  with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
  identity (it rendered as an unstyled UA-grey button); stale markup
  comment about the web shortcut rewritten; clamp docs updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:01:29 +02:00
Codeman maintainer 015b865f56 fix(deepseek): review fixes for the transcript reader — docker/remote gate, poll memo, honest pairing docs
Four review findings on the worker-transcript feature:

- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
  transcripts live in the container's / remote host's own ~/.dsh, which
  the local reader can never see, so the transcript path returned
  'nothing said yet' forever and an agent polling such a worker starved
  on an answer that existed. Gated on !session.docker && !session.remote
  (statically pinned) and documented in the integration guide.

- last-response reads are memoized on (path, mtime, size, blocks): the
  skill's last_text polls once per second, and each poll decompressed and
  reparsed the whole file on the event loop even when nothing had been
  appended. An unchanged poll now costs one stat.

- The pairing ladder's comment claimed /new is served by step 2; in truth
  the boot-window transcript wins for as long as it exists (deliberately:
  preferring newest-eligible would hand a worker its busier sibling's
  reply). The comment now states the real tradeoff instead of the
  aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
  corrupt frame truncates the decode there, which is the safe behavior.

- stripReasoningPrefix no longer runs on user prompt text, so a prompt
  containing a literal </think> renders whole in blocks view.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:30:41 +02:00
Codeman maintainer 33f77c4680 fix(tabs): review fixes for the detailed rail — width-dialog default, compact wrap pass, rich-aware resets
Three review findings on the detailed-rows feature, all in its edge cases:

- The App Settings width select consulted the handheld defaults blob
  (tabRailWidth: 256) BEFORE the rich-aware default, which the renderer
  never reads — so a tablet's unsized rich rail rendered 320 while the
  dialog said 256, and a routine Save persisted the 256 (below the 288px
  tight threshold, permanently). The chain now mirrors
  applyTabRailWidth()'s actual resolution.

- _setTabRailWidth() re-rendered on a compact flip but never re-ran
  applyTabWrapSettings(), the one owner of the folder line, whose railRich
  input reads the compact class this function just toggled. A rich rail
  dragged below 240px kept emitting folder rows — persistently, for a
  stored width < 240, since the boot wrap pass runs before the class is
  first applied. The wrap pass now re-runs on the flip, with exactly one
  render either way.

- Both reset affordances (handle dblclick, Enter on the handle) reset to
  the hardcoded 256 even on a rich rail, landing it below the tight
  threshold; both now resolve the rich-aware default (320), via a new
  optional defaultWidth input on resolveTabRailKeyboardWidth().

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:26:07 +02:00
Codeman maintainer 6261b6f655 feat(skill): spawn and drive DeepSeek Harness workers
The agent skill could spawn a worker in any mode, but it could only
DRIVE a claude one: every other CLI has neither a real end-of-turn
signal nor an answer to read, so the recipes route them through output
markers.

dsh has both halves now -- its harness reports idle/working/blocked to
Codeman, and the previous commit reads its transcript -- so it joins
claude as a mode the four verbs work on unchanged. `spawn_workers alpha
beta:deepseek` is a mixed fleet in one call, and `sendwait` / `last_text`
/ `delete_session` need no per-mode variant.

Preamble 1.20.0 (SKILL.md's §0 heredoc regenerated from it):

- `spawn_worker` grows a deepseek branch that gates on the harness
  composer. ⚠️ Readiness there is NOT the stop signal: the harness
  reports idle at BOOT ~300 ms before its composer paints (measured
  2.26 s vs 2.56 s after spawn), so a send-and-wait fired straight after
  quick-start resolves on the boot edge, reports a turn that never ran,
  and strands the prompt in a pane not yet taking input. Waiting for the
  composer also spends that edge, since signals are edge-triggered.
- `spawn_workers` takes `name[:mode]`, so a mixed fleet stays one
  concurrent call. Case names still have to be unique -- the mode never
  disambiguates two workers that would share a directory.
- `sendwait` asks for `wait:"stop,exit"` instead of the `wait:true`
  default set. That set also carries `idle`, which for an external CLI is
  inferred from output stabilization: on a dsh worker whose TUI repaints
  rarely, the re-wait resolved in 0 ms with `signal:"idle"` on a turn
  with three minutes left to run. It also makes a wrong mode loud -- the
  modes that cannot deliver `stop` answer 400 before writing anything,
  instead of resolving on a flap.
- The self-heal resend carries `delivered:true` forward. The resend is a
  tagged duplicate, so the server truthfully reports `delivered:false`
  about a write it skipped, and §1's cleanup then read a completed turn
  as an undelivered one and kept a finished worker forever.
- dsh workers spawn with the permission posture the Run button sends,
  because the harness default still asks and a worker parked on an
  approval row cannot finish a fan-out. The multi-user clamp still
  applies.

Docs: a worked dsh flow in recipes.md, readiness and the signal rules in
verbs.md, and the corrections this makes necessary -- `stop`/`blocked`
are no longer claude-only, and `last-response` is no longer permanently
empty for deepseek. The integration guide gains a section on reading a
session back and driving one as a worker; its web-UI section was also
stale (that server moved out of a shell session).

The static guard that keeps those lists from naming some external CLIs but
not others is extended rather than exempted: it now knows the three real
classes inside that family (no transcript, no hook signals, and the
positive twin -- the modes whose answers can be read), with the hook class
derived from `hooksAvailableForMode()` so the predicate and the prose
cannot drift apart. Any other partial list still fails, and a new backend
belongs to none of the classes until someone says so.
2026-08-25 04:17:48 +02:00
Codeman maintainer d1bc0c517d feat(deepseek): read dsh session transcripts for last-response
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.

dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:

1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
   (one-shot and streaming alike) stops at the first frame end: a real
   56-line transcript decoded as 1 line / 158 bytes -- the session header
   alone, i.e. a silent truncation that reads as "nothing said yet"
   forever. `zstdFrameRanges()` walks frame and block headers to find
   exact boundaries; splitting on the 4-byte magic would corrupt
   everything after a magic sequence occurring inside compressed data.
   zstd is resolved at RUNTIME because it landed in Node 22.15 while the
   project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
   context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
   surfaced as `Turn error: …` (and a non-error early stop as
   `Turn ended: …`) rather than as an empty string, which an agent reads
   as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
   the streamed deltas fill in only for a step that never finalized, so a
   partial answer is readable mid-turn and never doubled. "Finalized" is
   tracked as a set of steps rather than as non-empty text, because a
   step whose whole reply was reasoning strips to '' at the `</think>`
   boundary and would otherwise resurrect the raw deltas in its place.

Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.

An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
2026-08-25 04:14:26 +02:00
Codeman maintainer c30dfaf0e7 fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.

The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:

- Exactly one server. A second click reuses the running one instead of racing
  it for a port; the session flow could not do this at all, because two clicks
  were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
  own /api against the browser authority, and a Codeman reachable at both
  loopback and a tailnet name has two. Reusing a server fenced for the other
  origin renders a page whose every call 403s, which reads as a broken
  dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
  signalled at once, which also means it would outlive Codeman and hold its
  port against the next start - the exact EADDRINUSE this feature already got
  wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
  for a stack trace to land, so a failed spawn reports its own tail.

The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.

`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.

Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
2026-08-25 03:08:15 +02:00
Codeman maintainer 15ae5f5d81 fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.

1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
   precisely the port a DeepSeek user is most likely to be serving on already,
   so the launch died with EADDRINUSE against the user's own server. The port
   now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
   free loopback port by BINDING it (a connect probe cannot tell "free" from
   "listening but not answering yet").

2. It opened the tab unconditionally. The crashed server left a saved dashboard
   pointing at nothing, with the failure only visible in a shell tab nobody had
   a reason to look at. The launch now polls the existing webview probe until
   the URL answers, and on timeout reports the error naming the shell tab
   instead of persisting a dead dashboard.

3. The saved tab was untrusted, so the frame was sandboxed without
   `allow-same-origin` and the dashboard was broken twice over: the dsh
   client-runtime reads `localStorage` while loading its plugins and died there
   ("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
   opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
   every `/api` call no matter which authority `--trusted-host` named. Passing
   `location.host` only means anything once the frame actually carries that
   origin, so `--trusted-host` had never once done its job. The managed tab is
   now created `trusted: true`.

   That trade is real and deliberate: a trusted proxied frame is same-origin
   with Codeman and can reach Codeman's API. It is defensible only because this
   dashboard is an agent harness Codeman just started itself, on loopback, which
   can already run code as the user. It is not a precedent for trusting
   third-party dashboards, which is why it is set at this one call site rather
   than defaulted.

Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.

`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.

Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
2026-08-25 02:39:57 +02:00
Codeman maintainer 14de2b7012 fix(tabs): keep the created stamp reachable on a tight rail, and do not skip the first render
Two review nits on the vertical rail's detailed rows.

1. The tab-rail-tight rule (below 288px) hides `.tab-meta-created`, and its
   comment claimed the value "survives in the row's title attribute either way".
   It did not: the only title carrying it lived ON that element, and a
   `display: none` element has no hover target, so the created stamp was not
   shrunk but gone with no way to ask for it. Rather than just correcting the
   comment, `_sidebarRichMetaHTML()` now puts BOTH absolute stamps on the
   `.tab-meta` line itself, so the pill and the gaps around the stamps remain as
   hover targets. An item's own title still wins where the item is visible.

2. applyTabOrientation() decided whether applyTabWrapSettings() had already
   re-rendered by comparing `_tallTabsEnabled` before and after. That reads an
   UNDEFINED previous value as "it rendered", but applyTabWrapSettings()
   deliberately renders nothing on its first call ever (it only establishes the
   baseline: `prevTallTabs !== undefined && prevTallTabs !== showFolder`). So on
   a first call that also flips the folder row, neither function rendered and the
   rows stayed stale. Reachable when the pre-paint script throws and leaves the
   layout attributes on their catch-branch fallbacks for applyTabOrientation() to
   correct. The guard now mirrors applyTabWrapSettings()'s own condition.

Both new tests were run against the unfixed code first and fail there, which is
the only thing that makes them regression tests. (The third, "does not render
twice", passes either way by design: it pins that fix 2 did not introduce a
double rebuild.)

Verified in a real browser against a live server with two sessions, driving the
narrowing through _setTabRailWidth() the way the resize drag does: at the 320
default the row reads "CREATED 2m ago · IDLE <1m" with the created element
displayed; at 256 the tight class is on, the created element computes to
display:none, the visible text drops to "IDLE <1m", and the meta line's title
still reads "First created: ...". At 220 the compact threshold drops rich rows
entirely. Screenshots confirm no truncation artifacts in either state.

Full gate green (6104 passed), typecheck, lint, format, frontend-syntax and
public-assets all clean.
2026-08-24 22:58:28 +02:00
Codeman maintainer cdceede33d fix(deepseek): atomic shim write, honest attribution comment, name-fallback profile classifier
The three smaller review nits, plus the first real test coverage for the status
shim (it had none: it is emitted as a STRING, so tsc never sees it).

1. The shim was written with a plain writeFileSync. The TUI can be exec'ing that
   exact path while an upgraded Codeman refreshes it, and a reader catching a
   half-written file gets a syntax error, exits non-zero, and is retried four
   times per state change for a file that will never parse. Now temp + rename
   (atomic within the directory), with the temp chmod'ed before the rename since
   writeFileSync's mode only applies on create, and removed if the write throws.
   SHIM_VERSION bumped to 2, because SHIM_SOURCE changed and an existing v1 shim
   would otherwise keep matching the embedded marker and never be refreshed.

2. The pane-id comment claimed the ambient env "cannot be spoofed by an argument
   the agent itself could influence". The agent runs IN that pane and can invoke
   the shim with CODEMAN_SESSION_ID unset and any argv it likes. It buys nothing
   it did not already have (the hook-secret file is readable from the same pane,
   so it can POST /api/hook-event directly), but the comment read like a security
   boundary. Rewritten to say what the preference actually buys: correct
   attribution when a TUI mangles or re-uses the pane argument. Accidents, not
   adversaries.

3. classifyProfile() folded the directory name into the same haystack as the
   bundles, but only the TUI arm could match a bare name, so a stock profile
   whose package.json has no dsh.profile.bundles (hand-edited, older layout,
   mid-install) classified as `unknown` -> launchable -> eligible as the DEFAULT
   pick, which is exactly the pane-dies-on-arrival failure the two-part
   availability gate exists to prevent. The stock names are now a LAST-resort
   fallback consulted after the bundle patterns, so real bundle evidence still
   wins over a name the user chose. The loose `tui` arm gained word boundaries:
   it decides which profile boots by default, and matching the middle of
   `intuition` is not a rule anyone could predict.

New test/deepseek-status-shim.test.ts runs the generated script the way the
harness does -- real node process, real argv, real env, real listener -- and
covers the exit-code contract that makes the retry behaviour safe: mapped states
post and exit 0, an unknown verb or unmapped state exits 0 WITHOUT posting (a
non-zero there would be four HTTP requests per state change forever), a rejecting
server or an unreachable one exits non-zero so the caller retries, the hook secret
is read at execution time, and `node --check` parses the file (a template-literal
typo in SHIM_SOURCE is invisible to tsc).

Trap worth recording, hit while writing it: the tests must spawn the shim
ASYNCHRONOUSLY. The listener lives in the test process, so spawnSync blocks the
event loop that has to accept the connection, the shim waits out its own 1500ms
socket timeout and exits 1, and it reads exactly like a broken shim (measured:
Socket._onTimeout in its --trace-exit output, server logging nothing).

Verified: full gate green (6142 passed, +10), typecheck/lint/format clean.
2026-08-24 18:00:52 +02:00
Codeman maintainer 2034719d61 fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed.

1. The multi-user clamp was bypassable by a sibling field on the same request.
   clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
   DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
   _configureDeepSeek(), so a non-granted owner sending
   envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
   instance: a session created with permissionMode "read-only" and that override
   ran with DSH_PERMISSION_MODE=danger-full-access in its pane.

   Every other CLI's bypass is a command-line flag reachable only through the
   per-CLI config, which is why the config clamp alone is the whole gate for
   them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
   owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
   what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
   that list because it aims the launcher at a profile tree whose plugin code
   runs at boot, before any approval row can apply. Verified end to end in real
   multi-user mode: a non-granted user sending both now gets workspace-write and
   no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.

2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
   signals only the direct child, and a plugin install fans out into
   package-manager children that keep the inherited stdio pipes open, so `close`
   never fires and the held-open request leaks with no route-level deadline.
   Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
   6s and both fan-out children were alive. Now detached: true plus negative-pid
   SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
   last-resort reap for a grandchild that escaped the group. Same probe after the
   change: close fires, direct child and both grandchildren dead.

3. hooksAvailableForMode() promised more than a dsh session can deliver.
   deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
   triple is the only reason a dsh session posts hook events, so `until=stop` was
   accepted and then blocked for the caller's whole timeout: the exact
   infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
   takes HookCapabilityOptions and every call site passes sessionHookOptions(),
   with the deepseek arm reading `!== false` so a forgotten one degrades to the
   old behaviour. The refusal names the setting rather than saying "no Claude
   Code hooks", which would send the caller hunting a bug that is really a
   setting they chose. Profile conformance stays unknowable at request time and
   is documented as such. The stale "True for `claude` and nothing else" docblock
   is corrected.

4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
   claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
   capture read Claude's own transcript, and adding deepseek silently widened
   both to a mode that has none. They compare mode === 'claude' directly now, and
   a static check pins them there.

Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
2026-08-24 16:01:02 +02:00
d fei 7c62b16e5f fix(terminal): bound live xterm backpressure 2026-08-24 19:06:50 +08:00
Aamer Akhter d15d979a33 fix: COD-341 correct changeset package name 2026-08-23 22:55:10 -04:00
Aamer Akhter 858b15e3f5 chore: COD-341 add File Viewer search release note 2026-08-23 22:46:58 -04:00
Codeman maintainer b330f1d9e8 feat(tabs): give the vertical rail the home screen's per-session detail
The vertical tab rail (tabOrientation 'vertical') listed names and nothing
else, while the rich sidebar and both home screens already answered the
question a docked column exists to answer: which of these sessions wants me
next, and how long has it been like that. The rail is a docked column too, so
it now draws the same row.

- New per-device setting tabRailDetail ('rich' | 'simple', default rich),
  App Settings -> Appearance -> Tabs, in SettingsUpdateSchema + displayKeys and
  stamped as data-tab-rail-detail by the pre-paint script, so a detailed rail
  does not flash through simple rows on every load.
- ONE gate for both vertical surfaces: isRichTabRows() =
  isSessionSidebarRich() || isTabRailRich(). The row model, the markup and the
  20s in-place clock are the existing rich-sidebar ones, classified by
  _mobileOverviewState/_mobileOverviewSince, so the rail, the sidebar, the
  desktop home rail and the phone overview cannot disagree about what
  "working" means or which stamp measures it.
- Detail rides on its OWN attribute, exactly as the sidebar's does, so every
  existing [data-tab-orientation='vertical'] rule keeps matching both variants
  untouched. A flip of detail ALONE still forces a full render (the stamps line
  is emitted by the row template, not toggled by CSS) and re-runs
  applyTabWrapSettings(), which owns the folder line and is now rail-aware.
- CSS: every rich paint rule gains a rail twin as a COMMA-GROUPED selector,
  never :is() - an :is() list takes its most specific argument, which would
  lift the sidebar arm from (0,3,1) to the rail's (0,5,1) and let these rules
  outrank things they never used to.
- Width is why there are thresholds. At 256px the stamps line ellipsizes
  mid-word, the same reason the rich sidebar is 300px, so a rail that has never
  been sized defaults to 320 (RICH_DEFAULT_WIDTH, the existing Wide preset,
  which also keeps the settings select on a named choice). A width the user has
  chosen is never overridden: below 288px the created stamp is dropped rather
  than truncated (tab-rail-tight, CSS only) and below 240px the rows go back to
  simple (tab-rail-compact, which re-renders).
- The rich clock is armed and disarmed by applyTabOrientation() as well as
  applySessionListLayout(); a leaked interval would rewrite stamps in a list
  that no longer has any.

Also fixes a data-loss bug in the inline tab rename that predates the rail and
reproduces in every layout, header strip included: Escape set the input to ''
and blurred it, and the blur handler commits - so cancelling a rename PUT an
empty name, and the tab fell back to its folder label (measured against a live
server: ["rail-alpha","","rail-gamma"]). Escape now calls cancelRename(), which
invalidates the edit so the blur that follows the input's removal is a no-op.

Tests: rail-detail gate, the three ways it turns back off (simple, compact,
horizontal), sidebar-wins, render-on-detail-flip and the plumbing/CSS guards in
test/session-list-layout.test.ts; the rename cancel in test/inline-rename.test.ts
(browser suite), pinned by running it against the old code first. Verified live
against a real server on an isolated instance: detailed/simple/compact/header/
sidebar variants, click-select, the ... menu, inline rename, Alt+N, the in-place
stamp tick and a full settings-picker round-trip including reload.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 04:45:02 +02:00
Aamer Akhter c14171b534 fix(file-viewer): COD-341 finalize deferred navigation 2026-08-23 22:42:10 -04:00
Aamer Akhter acd9ffedc8 fix(file-viewer): COD-341 complete search transitions 2026-08-23 22:27:26 -04:00
Aamer Akhter 921933775b test(file-viewer): COD-341 execute session lifecycle path 2026-08-23 22:08:29 -04:00
Aamer Akhter f6a1f06633 fix(file-viewer): COD-341 synchronize session lifecycle 2026-08-23 21:58:45 -04:00
Aamer Akhter dab8e6643c fix(file-viewer): COD-341 deduplicate normal tree loads 2026-08-23 21:40:10 -04:00
Codeman maintainer 4cda150493 feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.

DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:

1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
   $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
   `base` -- the interactive terminal front door is always a third-party
   plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
   and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
   button gates on the latter, because reporting only the binary would spawn a
   pane that dies on arrival. When the binary is present but no profile is,
   the run menu offers to install one (POST /api/deepseek/install-profile).

2. The permission switch is an env var, not a flag. The harness has no
   command-line permission option; its sandbox/approval rows read
   DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
   Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
   own workspace-write, which still asks, so the multi-user clamp is the
   only-if-sent branch and clamps to workspace-write, never read-only.

3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
   earned that. The terminal front door reports idle/working/blocked to a
   supervising process over a generic env-gated contract; a generated shim
   (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
   report to /api/hook-event as stop / agent_working / permission_prompt. So a
   dsh session gets definitive respawn triggers, real wait-endpoint signals and
   real Approvals Inbox items instead of output-stabilization guesswork.
   `agent_working` is new (157th SSE constant) and joins
   APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
   alert at once.

The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.

Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.

Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.

Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 03:37:56 +02:00
Aamer Akhter 3af36f7c34 fix(file-viewer): COD-341 preserve search row layout 2026-08-23 21:23:17 -04:00
Aamer Akhter 49797e37dd fix(file-viewer): COD-341 gate stale search results 2026-08-23 21:09:52 -04:00
Aamer Akhter c614331d60 feat(file-viewer): COD-341 add server-side search 2026-08-23 21:01:39 -04:00
Codeman maintainer 9cfd8e8989 fix(docker): survive xAI installer's own /usr/local/bin/grok symlink
The agent-image grok step copied /root/.grok/bin/grok onto /usr/local/bin/grok
with cp -L. Newer versions of xAI's install.sh already create
/usr/local/bin/grok as a symlink to that same binary, so the copy failed with
'same file' and the --no-cache rebuild died at the grok layer (2026-08-24).
Stage the copy under a temp name, drop whatever the installer left at the
destination, then move into place - correct against both old and new
installers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 01:12:38 +02:00
Codeman maintainer 8fe393826b chore: version packages 2026-08-24 01:00:51 +02:00
Codeman maintainer 7a340fe7bc fix(tabs): review-driven hardening for the tab-layout foundation and vertical rail
Post-merge follow-ups from the deep review of #334 and #335, so they ship in
the same release as the features.

Tab-layout foundation (#335):
- PUT /api/session-order drops unknown/foreign ids again instead of 400ing
  the whole write, in both the owner and the admin path (single-user requests
  are the synthetic admin, so that path is the one the browser hits). The
  frontend debounces its reorder push and swallows errors, so a session
  deleted inside the debounce window silently cost the user the entire
  reorder - and the endpoint sits on the stable /api/v1 surface, where the
  pre-layout server merged leniently.
- A failed mux restore no longer locks explicit deletions into 500s for the
  process lifetime: runSessionDeletion and webviewDeleted degrade to
  best-effort without layout coordination, while the automated stale sweep
  (runStaleSessionCleanup) stays fail-closed.
- sse-events doc comment: no 'suppressed' hook event exists; hooks stay 8.
- registerSessionWithLayout resolves its owner through ownerLayoutKey()
  instead of a hardcoded '@single'.

Vertical rail (#334) - all rail-awareness gaps in sidebar-only predicates,
unified behind the new _isVerticalTabList() (sidebar OR rail):
- Drag-reorder read the insertion side from clientX in the rail, so
  before/after was effectively arbitrary on vertical rows; the drag-over
  indicators now draw as top/bottom edges there like the sidebar's.
- The active tab is scrolled into view in the rail (Alt+N/palette selection
  used to leave the row below the fold).
- Floating subagent/ultracode windows anchor to the RIGHT of rail tabs, and
  the connector redraw gates (render tail + strip scroll) cover the rail.
- Server-seeded tabOrientation is applied when the async settings load
  resolves, not only at boot, so a fresh device shows the rail immediately.
- The pre-paint script stamps data-tab-orientation and --tab-rail-width
  (sidebar-wins and solo carve-outs included), removing the flash of the
  header strip on every vertical-mode load.
- The session name font defaults to 12px, the sidebar's historical 0.75rem
  size, so installs that never touch the new slider are not restyled.

Also documents the rail in CLAUDE.md (second #sessionTabs host, mover
ordering, the axis-predicate rule) and gives tab-rail-resize.js its
@dependency/@loadorder header. Full gate green (6093 tests); the excluded
browser suite was run by hand - only the known environmental failures
(opencode/codex binaries) remain, identical to pristine master.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 00:59:48 +02:00
Codeman maintainer f3c615b669 fix(files): give the file preview a working detach button
The button next to the file preview's close icon was Copy Content, whose
overlapping-pages glyph reads as a pop-out control - and for a PDF or any
media/binary preview it was completely dead: those branches never fill
filePreviewContent, so the click hit an empty-content guard and did nothing,
with no feedback.

There is now a real detach button that opens the previewed file in a browser
tab (raw route for PDFs/images/media/text, the server-converted PDF preview
for docx/pptx), severs window.opener by hand so a blocked pop-up stays
detectable, closes the overlay on success (which also stops any playing
media), and disarms on close so it can never open a stale file. The copy
button now toasts 'Nothing to copy in this preview' instead of staying
silent.

Verified live with Playwright against an isolated instance: button visible
and armed on a PDF preview, file-raw answers 200, clicking opens the URL and
tears the overlay down, text previews keep a working copy buffer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 00:59:27 +02:00
Ark0N 82f81d21c4 Merge pull request #333 from Ark0N/feat/grok-mode
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
2026-08-24 00:43:04 +02:00
Codeman maintainer c173ae0264 Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts:
#	src/web/public/app.js
2026-08-24 00:32:25 +02:00
Ark0N e9dd55e5fd Merge pull request #334 from aakhter/pr/cod-358-vertical-rail
feat(tabs): add a resizable vertical session rail
2026-08-24 00:14:10 +02:00
Ark0N dd96f252ea Merge pull request #335 from aakhter/pr/cod-359-tab-layout
feat(tabs): add owner-scoped tab layout foundation
2026-08-24 00:12:20 +02:00
Aamer Akhter 74194e4fc0 feat(tabs): COD-359 add owner-scoped tab layouts 2026-08-23 14:46:10 -04:00
Aamer Akhter c45c6c3846 merge upstream master into COD-358 2026-08-23 14:15:07 -04:00
Aamer Akhter 17b141dc25 test(workflows): keep recent-run fixture clock-independent 2026-08-23 14:12:38 -04:00
Codeman maintainer 57f326ab8f docs(grok): fix the mode counts and line refs the seventh mode invalidated
The agent skill's endpoints.md is what other agents read as ground truth, and
three of its facts went stale when grok landed:

- `/api/v1/grok/status` was added to the probe list, but the sentence after it
  still said only Pi's response carries `.data.version`. Grok's carries it for
  the same reason (a squatted binary name), and an agent that trusts the old
  wording has no way to tell a misresolved grok from an absent one.
- the `active-tools` bullet listed grok among the modes it stays empty for, then
  claimed in the same breath that `isExternalCliMode` "lists only those five".
- its three source line refs had all drifted: `isExternalCliMode` is now
  session.ts:174-183 (it was already wrong before this branch), the external-CLI
  early return is session.ts:2261, and TEXT_COMMAND_PATTERN is
  bash-tool-parser.ts:89.

CLAUDE.md and architecture-invariants.md counted modes in their Docker-cases and
Web-tabs paragraphs ("any of the five CLI backends", "never a sixth
SessionMode"). Both numbers were already stale before grok (antigravity and pi
had made it seven) and grok is now in the agent image, so the counts are gone
rather than incremented: the invariant those sentences carry is that Docker and
web tabs are not modes at all, which no number has ever helped state. The two
plan docs keep their original wording, being historical design records.
2026-08-23 19:57:13 +02:00
Codeman maintainer 6b0b6d10ad fix(test): anchor the workflow-run fixture to now instead of a pinned epoch
test/workflow-run-watcher.test.ts pinned its fixture's newest activity at
2026-06-14T20:06:40Z and then asked getRecentRunSummaries(100000) to return
it. That argument is MINUTES, so the window is 69.4 days: the assertion
expired at 2026-08-23T06:46:40Z and the file has failed on every branch
since, on a suite nobody had touched. The last green CI run finished at
06:47:43Z, about a minute inside the boundary, which is why it landed as a
surprise rather than a bisectable regression.

The fixture epochs now hang off a RUN_ANCHOR of Date.now() - 601s with every
offset preserved verbatim, so the parsed durations, the ordering and the
live-vs-done discriminators are all unchanged, and the recency filter is
still the thing under test. It just cannot rot again.
2026-08-23 19:54:15 +02:00
Aamer Akhter c9ea8bbac5 fix(tabs): COD-358 re-query tab after rename cancel 2026-08-23 13:40:42 -04:00
Aamer Akhter 1795a138b3 test(workflows): document clock-independent fixture fix 2026-08-23 12:47:39 -04:00
Aamer Akhter e3a2fb767f feat(tabs): COD-358 add resizable vertical session rail 2026-08-23 12:21:02 -04:00
Codeman maintainer 3f8c8e99d1 feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.

Grok mixes two existing shapes and the wiring follows from that:

- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
  (--always-approve, grok's bypassPermissions mode; config-level deny rules
  still apply on top). The Run button sends it true, like runAntigravity(),
  and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
  a bare grok spawn is grok's own ask-mode default, which is already safe,
  so only a sent config needs the flag forced off. Cron needs nothing for
  the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
  with mouse support, so it stays OUT of isAltScreenStripMode() and lands
  on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
  (unmeasured against an authenticated composer; documented fallback is the
  'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
  also installs a grok bin), so grok-cli-resolver.ts version-probes every
  candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
  GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
  shared with the dependency registry so doctor and run mode cannot drift.

Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.

Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.

Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).

Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:39:03 +02:00
Codeman maintainer 88bb98de43 chore: version packages 2026-08-22 14:45:08 +02:00
Codeman maintainer 687e9d7565 feat!: retire the sc tmux chooser in favour of codeman tui
scripts/tmux-chooser.sh is deleted. codeman tui replaces it and does the
job better: sc numbered its entries globally but only accepted a single
[1-9] keypress, so sessions 10+ were listed and unselectable, and it
inferred nothing about what an agent was doing. The tui carries the
server's real states, answers permission dialogs, and leaves an attach
with one key.

install.sh no longer creates the tmux-chooser symlink or the sc alias.
It now sweeps both up instead, on update AND uninstall, so an update
cannot leave a symlink pointing at a script this version stopped
shipping. The alias removal is marker-owned: it matches the exact line
the installer wrote, so someone's own 'alias sc=' for another tool is
never touched, and it rewrites through 'cat >' so the profile keeps its
mode and ownership. Verified against three profile shapes.

BREAKING CHANGE: the 'sc' command and the 'tmux-chooser' symlink are
gone. Use 'codeman tui' (and 'codeman tui --list' / 'codeman tui <n>').
2026-08-22 14:44:14 +02:00
Codeman maintainer 5d81cc01ca docs: point terminal users at codeman tui instead of the sc chooser
codeman tui supersedes the sc bash chooser: it reaches sessions 10+,
carries the server's real states instead of a static list, and leaves an
attach with one key. Every place that told a user to run sc now names the
tui equivalent, including the two wiki pages and install.sh's next-steps
banner. Both wiki pages also carried the wrong detach chord (Ctrl+A D;
the socket's prefix is C-b), which the tui makes moot.

Source comments that explained themselves as "the sc -l replacement" now
just say what they do. docs/tui-plan.md and CHANGELOG.md are historical
records and keep their references.
2026-08-22 14:36:23 +02:00
Ark0N f49249fb2f Merge pull request #312 from Ark0N/feat/tui
feat: codeman tui, a terminal dashboard with live agent states
2026-08-22 14:32:51 +02:00
Codeman maintainer bc3f9f8a37 docs: record the socket-resolution rule where the mechanism lives
CLAUDE.md gained "any new `tmux -L` caller through `resolveTmuxSocketName()`"
but architecture-invariants.md, which that bullet points at for the
mechanism, still described only the dataPath() half. Say why the rule
exists there too: the TUI is the first non-server process to shell out
to tmux.
2026-08-22 14:25:23 +02:00
Codeman maintainer 97acfc61c3 docs: stop telling sc users to press a key only the TUI binds
The sc chooser runs a plain `tmux attach-session` and binds nothing, so
F1 does not detach from it; only `codeman tui`'s attach claims that key,
and only for its own duration. The line it replaced was wrong too (the
socket's prefix is C-b, not C-a), so name the real chord and say which
command gives you the single key instead.
2026-08-22 14:19:31 +02:00
Codeman maintainer c27363459d docs(tui): stop telling people to press a key that does not work
The guide and the README both said to detach with `Ctrl+B D`. Beta testing
proved that wrong twice over: tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`, and even the correct letter fails for anyone who
keeps Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. A
tester followed the documented instruction, stayed attached, and exited the
agent to escape.

Both now say `F1`, and the attach section describes what actually happens: the
session strip across the top of the pane, `Alt+1`..`Alt+9` switching without
returning to the dashboard, and `r` to resume a session whose pane has died.
Also corrected: `1-9` switches rather than jump-attaches, `x` confirms with `y`
rather than a typed name, and a new session opens straight into its pane.

`docs/tui-plan.md` is deliberately untouched — it is the design record of what
was planned, not a description of what shipped.
2026-08-22 14:13:58 +02:00
Codeman maintainer 737c2ed7f8 fix(tui): escape the separator in the switch binding, closing the sizing leak
The loose end from 777f974, now explained. Sessions came back from a detach on
`window-size latest` instead of `manual`, and the restore primitive round-tripped
correctly in isolation, so the corruption had to be upstream of it. It was: the
snapshot was taken from state this code had already broken.

`bindSwitchKey` passed a bare `;` between the two commands it wanted in one
binding. That is a command separator to tmux's OWN parser, not an argument: it
ended the `bind-key` and executed what followed immediately. So the binding kept
only `switch-client`, and `set-window-option ... window-size latest` RAN against
every switchable session at attach time — before the sizing snapshot was taken.
Every session was therefore snapshotted as `latest` and faithfully restored to
`latest`.

Proven against real tmux both ways before fixing: a bare `;` leaves the session
on `latest` and stores a one-command binding, while `\;` leaves it `manual` and
stores both commands.

Verified end to end: 7 sessions manual before, 1 latest + 6 manual during the
attach (the attached one follows the terminal, the rest are pre-sized), no dot
padding on a switch, and all 7 back to 120x40 manual after the detach.

This also means the "follow the terminal after switching" half of 777f974 never
actually worked — it was never in the binding.
2026-08-22 14:13:58 +02:00
Codeman maintainer bb24d2c256 fix(tui): stop the preview stacking every repaint of a session
The overview showed the same session twice, one frame above another, after
switching sessions (reported from the beta with a screenshot).

Claude repaints by ABSOLUTE CURSOR POSITIONING, not by clearing: a 198KB pane
tail carries 1142 `CSI r;c H` and exactly one `CSI 2J`. The replay honoured the
COLUMN of those sequences and ignored the ROW, so a repaint could never
overwrite what came before and was appended instead. That same tail replayed as
FIFTY stacked copies of one frame. The preview shows the last N lines, so on a
short terminal you saw the newest frame by luck and on a tall one you saw the
end of the previous frame above it.

A cursor HOME now starts the buffer over. That is not a heuristic but the
line-based equivalent of what a home means: a full-screen app announcing it is
repainting from the top, with everything on screen about to be overwritten in
place. Only row 1 column 1 counts — any other address is a write position
inside the frame being painted, and resetting on those would erase live
content.

Measured on the real tail that produced the screenshot: 198599 bytes and 50
copies of the welcome frame collapse to 40 lines carrying exactly one.

The old test pinned the append behaviour, including a spurious leading empty
line that the initial CUP produced; both are gone.
2026-08-22 14:13:58 +02:00
Codeman maintainer bc1821661f fix(tui): say alt+1-9 on the bar, and stop the dot grid when switching
The bar now reads "alt+1-9 switch · F1 back to the codeman dashboard", so the
switch keys are discoverable instead of secret. Shown only when those keys were
actually claimed, the same rule the way-out key follows: a bar naming a key
that does nothing is the bug this series started with.

THE DOT GRID. Switching landed in a pane occupying part of the terminal with
tmux's dot fill everywhere else. It was never a size mismatch — the window was
already the right size. `window-size latest` only resizes a window while a
client is ON it, and the sessions behind the tab strip have none until you
switch, so the resize happened AT the switch: tmux painted the newly-available
area with dots and an idle claude had no reason to redraw into it. Every
switchable session is now pre-sized to the attaching terminal, which moves that
repaint to attach time while the user is still looking at the first session,
and the switch binding restores `window-size latest` on arrival so a mid-attach
terminal resize still follows. Measured: 14 consecutive switches across 7
sessions, zero dot-padded rows, against 1-in-6 before.

⚠️ Known loose end, deliberately not papered over: after a detach the window
SIZE is restored exactly but the window-size MODE can come back as `latest`
rather than `manual`. The restore primitive round-trips correctly in isolation
(manual -> presize -> latest -> restore = manual) and no call site in the TUI
or the server sets `latest` afterwards, so the cause is not yet identified. The
practical effect is nil: the remaining client keeps the window at its own size
and Codeman re-pins `manual` on the browser's next resize.
2026-08-22 14:13:58 +02:00
Codeman maintainer 74c9879359 fix(tui): size every switchable session, not just the one being attached
Switching with Alt+N landed in a pane that filled part of the terminal with
tmux padding the rest as a dot grid — reported from the beta with a screenshot
showing the pane in the left half and dots everywhere else.

Codeman pins every window `window-size manual` at the BROWSER's size
(tmux-manager.ts), so no attaching client can resize it. The attach already
lifted that for the session it opened, which is why a plain attach looked
right; `switch-client` then moved the user into a session that had never been
lifted, and the old pin reasserted itself. `window-size latest` now goes on
every session the strip can reach, alongside the bar those sessions already
get, and each one's original sizing is snapshotted and restored on detach.

Verified by round-tripping a session pinned at 120x40 manual: latest 190x49
while attached, back to 120x40 manual after, with no dot rows at either step
and the bar intact at full width after a switch.
2026-08-22 14:13:58 +02:00
Codeman maintainer 499d3d6e4d fix(tui): finish the 1-9 rename in the fallback footer
The renderer's own FOOTER_KEYS table still said 'jump'. It is only reached when
the app layer supplies no footerKeys, so nothing visible was wrong, but a
fallback that contradicts the live footer is exactly the kind of drift that
turns into a bug report later.
2026-08-22 14:13:58 +02:00
Codeman maintainer 9b29666e03 fix(tui): keep the way out on the bar, and make Alt+1..9 actually switch
Four faults, all reported at once, and three of them were mine from the last
two commits.

THE HINT VANISHED. Two independent causes. First, a leaked F1 binding: an
attach whose TUI was killed leaves `F1 -> detach-client` in tmux's root table,
and the claim treated "already bound" as someone else's key, so every later
attach fell back to advertising the tmux chord — the bar stopped saying F1
while F1 still worked. A key already bound to `detach-client` now counts as
ours. Second, width: tmux truncates a status line that overflows and drops the
RIGHT-aligned segment, which is the hint. The strip now gets a budget measured
from the terminal's width minus the hint, and it drops tabs from the far end
until it fits. ⚠️ Measured on VISIBLE columns, not format bytes: `#[reverse]`
costs zero columns, and counting it made a strip that "fitted" still truncate
the hint at 80, 100, 120 and 176 columns on a real terminal.

ALT+N DID NOT SWITCH. On the dashboard, a bare digit meant jump AND ATTACH, and
a terminal sends Alt+N as ESC then N: when those land in separate reads —
routine over SSH — the chord decodes as Escape plus a bare digit, so "switch to
tab 2" threw the user into tab 2's pane. A digit now SELECTS, matching what
Alt+N means in the web UI; Enter is how you go in. Inside a pane the keys never
reached the TUI at all, since tmux owns the terminal, so the attach now binds
Alt+1..9 in tmux's root table to `switch-client` — the strip is usable rather
than decorative. ⚠️ The bar is applied to every session the strip can reach,
each highlighting its own tab: with it on the attached session only, switching
landed the user in a pane with no strip and no way out on screen.

⚠️ The leaked-state sweep was missing `status-position`, so it removed the
marker and left the position behind — and with no marker the leftover no longer
matched, making it permanently unsweepable. Found by diffing every session's
options after a detach.
2026-08-22 14:13:58 +02:00
Codeman maintainer aec6516638 feat(tui): keep the session tabs visible inside a pane, and move the way out to F1
Attaching made every other session disappear: the dashboard is gone, tmux owns
the terminal, and there is nothing left saying what else is running. The attach
bar now carries the session strip, numbered exactly as the dashboard numbers
them, with the session you are in inverted, and it sits at the TOP of the pane
where the web UI keeps its tabs.

The strip is a WINDOW around the active tab, not the whole list, with ellipses
marking each end that is actually cut. The bar is one line shared with the way
out, and that hint is the only instruction a user gets while tmux has the
terminal, so it must never be crowded off; a test drives 20 long-named sessions
through the bar and asserts it survives.

⚠️ The strip is a snapshot taken at attach time and never refreshed. The TUI is
blocked in `spawnSync` for the whole attach so there is no loop to update from,
and tmux's own format language cannot map a `codeman-<hex>` session name back
to a label a human recognises. Slightly stale beats absent.

The way out moves from F12 to F1, which sits beside Esc where a hand backing
out already goes. Verified against BOTH encodings a terminal sends for it:
xterm's SS3 (ESC O P) and PuTTY's default (ESC [ 1 1 ~).

`status-position` joins the snapshot, so a session that had its bar at the
bottom gets it back there on detach along with everything else.
2026-08-22 14:13:58 +02:00
Codeman maintainer c59f006bb6 fix(tui): stop drawing from the unicode blocks a plain terminal font lacks
Three separate "why are there boxes" reports, and I fixed them one glyph at a
time instead of as a class, so the next one was always waiting. Grouping the
tester's terminal by unicode block made the rule obvious:

  RENDERS   Latin-1 (·), Box Drawing (─ │), Block Elements (█ ▛ ▐),
            Geometric Shapes (○ ▶), General Punctuation (…), Arrows
  TOFU      Miscellaneous Technical (⏎ U+23CE, ⏵ U+23F5), the sparse end
            of Dingbats (❯ U+276F)

That is an ordinary font, not a broken one, so it is the profile to design
against. The working spinner moves off Dingbats and Math Operators onto
quadrant blocks (▖▘▝▗) — the same block as the `▛█▐` art claude itself draws,
which that font renders fine — and the blocked marker moves off `⚠`
(Misc Symbols, emoji presentation on many terminals) onto `▲`, the block that
already gives us `▶` and `○`.

The preview fold gains claude's own spinner dingbats (✢ ✳ ∗ ✻ ✽ ✴ → `*`) and
`⚠` → `!`. Its animated status line is exactly where a reader looks, so tofu
there is the most visible kind there is.

A test now enforces this as a CLASS: no glyph in the unicode set may come from
Misc Technical, Misc Symbols or Dingbats, with U+2714 the single documented
exception because it was observed rendering on the very font that failed the
others. Verified by scanning a live frame driven with the tester's exact
environment: zero glyphs from any of the three blocks.
2026-08-22 14:13:58 +02:00
Codeman maintainer b2ac6c1bd9 fix(tui): confirm a kill with y, and make the dialog say what it would kill
Killing demanded the session's NAME typed out in full. That is the right
ceremony for dropping a production database and the wrong one for closing a
pane you are looking at; the beta tester's verdict was "thats stupid, just make
me type Y to confirm". `x` then `y` is already two deliberate keystrokes on a
row the user selected, and the conversation lives in its transcript, which a
kill does not touch.

Everything that is not `y` CANCELS rather than being ignored, so a stray key
closes the dialog instead of leaving a destructive prompt armed and waiting for
whatever gets typed next. Enter cancels too: it is the key most likely to be
hit by reflex, and this is the one dialog that destroys something.

⚠️ Found while verifying the new dialog: it did not name the session. The label
was computed as `row.session.name ?? id.slice(0, 8)`, and `??` falls back only
on null or undefined, so every session the server left with an EMPTY name — all
of them, until the TUI started naming its own — sailed through and the box read
"Kill ?". A destructive prompt that cannot say what it will destroy is worse
than no prompt, and it is now a single keystroke. The caller passes the same
label the LIST shows, so the dialog names the row in front of the user.

The typed-name machinery goes with it: TuiConfirmState.typed, setConfirmInput(),
confirmAccepts() and the 'typing'/'reject' steps are all removed rather than
left as unreachable branches.
2026-08-22 14:13:58 +02:00
Codeman maintainer 7b1150ca4f fix(tui): start a session straight into it, and drop two unsafe glyphs
Two reports from the same beta screenshot.

Starting a session left the user on the dashboard next to the row they had just
asked for, which reads as the create having silently failed. Starting a session
is a request to WORK in it, so the terminal now goes there as soon as the pane
exists, and the CLI booting is worth watching. If the pane is slow the notice
says so and the row is left selected, exactly as the resume path does.

The footer's `↵` was drawing as an empty box: `⏎` (U+23CE) has poor font
coverage, on the same terminal that renders `·`, `─`, `│`, `○`, `▶` and `✔`
perfectly. It is now U+21B5, from the Arrows block every monospace font ships.

`✋` (U+270B) was worse than a coverage problem: it is East Asian WIDE, so the
renderer, which addresses cells by column, was reserving two cells for it. The
golden frames had the age column shifted a space left to match, which is how
long that had been wrong. It is now `!`, and the frames align correctly.

A test walks the whole unicode glyph set and fails on any entry wider than one
cell, so a glyph that shifts the layout cannot be added again. The comment on
the table spells out both bars a glyph has to clear, because the tier check
answers neither: it asks whether the LOCALE is UTF-8, which says nothing about
whether a font has the glyph or how wide it draws.
2026-08-22 14:13:58 +02:00
Codeman maintainer 95a1f540b5 feat(tui): switch sessions with the web UI's shortcuts
Alt+1..9 switches to that session, and `[` / `]` / Tab step through them, so the
muscle memory from the web UI carries over.

Alt+N SELECTS rather than attaches, which is what the web UI's Alt+N does:
switching which tab you look at is cheap and reversible, and the terminal
equivalent is moving the selection and its preview, not handing the whole
terminal to a pane. Bare 1-9 keeps its documented jump-and-attach meaning.

Two of the web UI's chords cannot cross into a terminal, so the nearest
transmittable keys carry them instead:

  Alt+[ / Alt+]  ESC+[ IS the CSI introducer every arrow key arrives on, and
                 ESC+] is OSC, so neither chord is distinguishable from a
                 sequence. Bare `[` and `]` do the job.
  Ctrl+Tab       a terminal cannot report the Ctrl, so plain Tab carries it.

⚠️ The parser now decodes ESC + a printable character in ONE read as an Alt
chord, and the app replays every chord it does not claim as `escape` then that
character. That fallback is load-bearing, not tidiness: a real Esc landing in
the same read as the next keystroke is byte-identical to a chord, and without
the replay "Esc then q" typed quickly decoded as Alt+Q, matched nothing and was
swallowed. The e2e suite caught exactly that as the dashboard refusing to quit.
A lone Esc is still held and flushed on the caller's timer, which is what keeps
the two separable at all.
2026-08-22 14:13:58 +02:00
Codeman maintainer e7b7e90a1b fix(tui): fold rare prompt glyphs in the preview so they stop rendering as boxes
A beta tester photographed claude's `❯` prompt and its `⏵⏵` bypass-permissions
marker rendering as empty boxes in the preview pane. Their font has no coverage
for those codepoints while drawing `·`, `─`, `│` and `▶` perfectly.

The glyph TIER cannot help here. It answers "can this terminal do Unicode at
all", which is a locale question, and it correctly says yes for exactly the
terminals this affects. Coverage is per-glyph and undetectable from inside the
process, so the handful of rare glyphs CLIs use as chrome are folded to the
ASCII arrows they already look like, and everything a plain font does render is
left alone.

Scoped tightly: the preview only, never the TUI's own chrome, and skipped
entirely at the `nerd` tier where the user has declared a font that can draw
anything. The table is short and every entry was seen as tofu in a real
terminal rather than guessed at. The fold is length-preserving, so the preview
pane's column arithmetic is unaffected.
2026-08-22 14:13:58 +02:00
Codeman maintainer 84f8a8a2fe feat(tui): offer r to resume a session whose pane has died
Refusing the attach stopped the freeze but told the user to throw the session
away (`x` to close, `n` for new), which loses the conversation. tmux's own
dead-pane screen already says what to do instead: `claude --resume "<name>"`.

The Error card now offers `r` when the row can actually be resumed (claude,
with a conversation id and a working directory), and the footer says so. One
press resumes into a fresh pane and attaches to it, so a dead end becomes
recovery.

⚠️ Three things keep this from becoming the resume runaway that once spawned 35
sessions in 40 seconds. The offer holds a session ID, not a row, and is
re-resolved from the model when the key is pressed: a row captured when the
card opened is stale by then. It disarms BEFORE anything async, so a second `r`
cannot start a second resume. And it routes through resumeSelected(), which
owns the `resuming` flag and ends in attachToSession() rather than the group
dispatch.

⚠️ The `r` branch has to run BEFORE the generic dismiss, because a message
overlay is dismissed by ANY key: without that ordering the offer is consumed as
"some key was pressed" and the card merely closes. `help` keeps the any-key
behaviour, so the two modes no longer share a case.

Verified end to end against a genuinely dead claude pane: card, footer, one
press, one new session, and F12 back to the dashboard.
2026-08-22 14:13:58 +02:00
Codeman maintainer 6ad9145417 feat(tui): leave an attach with ONE key, F12, and no modifier
Three beta rounds died on tmux's native way out, and the last one died on the
instruction rather than the mechanism: "press Ctrl+B, release Ctrl, then d" is,
in the tester's words, very unclear, and holding the modifier through both keys
silently does nothing.

So the way out stops being a chord. The attach claims F12 in tmux's prefix-less
`root` table for its own duration, and the bar reads "press F12 to get back to
the codeman dashboard" — one keystroke, nothing to hold, nothing to release,
no order to get right. F12 because stock tmux ships an empty root table apart
from mouse bindings, and none of the CLIs that run in these panes want the key.

⚠️ The bar names the one key ONLY when the claim succeeded, and falls back to
the chord wording otherwise. A bar advertising a key that does nothing is the
bug this whole series started with, and it must not come back in a new costume.
Same claim rules as the prefix alias: taken only when tmux reports the key
unbound, given back only while it still means `detach-client`.

The chord and the held-Ctrl alias both keep working; they are simply no longer
what the user is told to press.
2026-08-22 14:13:58 +02:00
Codeman maintainer ab4a868688 fix(tui): make the detach chord work when Ctrl is never released
Reported three times as "Ctrl+B and d is still not working", on a build whose
bar already named the right key. Measured against a live pane: of the three
ways a person types this, only one worked.

  Ctrl+B, release Ctrl, then d   detaches
  Ctrl+B then Ctrl+D (held)      nothing happens
  Ctrl+B then Shift+D            nothing happens

Holding Ctrl through both keys sends 0x02 then 0x04, and tmux ships `C-d`
unbound in the prefix table, so the keystroke is swallowed in silence and the
attach looks frozen. That is not a user error worth documenting around: holding
the modifier is how most people type a two-key chord.

The attach now claims the held-Ctrl form of whatever key detaches (`d` → `C-d`)
for its own duration and gives it back on restore, and the bar advertises it
only once the claim succeeded, so it can never name a key that does nothing.
⚠️ The key is claimed ONLY when tmux reports it unbound, and released only
while it still means `detach-client`, so a binding of the user's own is never
shadowed or removed. The alias is deliberately excluded from the leaked-state
sweep: key tables are server-global, so the sweep cannot tell a leak from a
second TUI's live claim, and a stray `C-d`→detach is harmless either way.

Ruled out along the way, with evidence rather than assumption: the encoding.
tmux negotiates no extended-key mode upstream on attach (no kitty CSI-u, no
modifyOtherKeys, no DECSET 2017), so Ctrl+B does arrive as a plain 0x02 even
from a Claude pane, which has its own keyboard protocol.
2026-08-22 14:13:58 +02:00
Codeman maintainer 35b2c1baa5 fix(tui): refuse to attach to a dead pane, and stop naming sessions after CLI noise
Two more from the same beta round, both reported as "basic things are broken".

Attaching to a DEAD pane trapped the user. Codeman sets `remain-on-exit on`, so
a session whose agent has exited does not disappear: the row looks ordinary,
the server still reports it idle, and Enter handed the terminal to a pane that
reads no input. With the detach chord also wrong at the time, that was a hard
freeze with no way out. Enter now probes `#{pane_dead}` first and refuses with
an Error card naming the session and what to do instead. The probe fails OPEN,
so it can never block an attach to a live pane. ⚠️ It also has to paint: the
keypress that reaches attachToSession() has already painted by the time an
awaited probe resolves, so message() alone left the refusal invisible and Enter
looked inert, which is the bug it was added to fix.

A session started from the TUI came out unnamed, because startSession() sent no
sessionName and rowLabel() then fell back to the transcript's first line. A
brand-new session has no prompt to be named after, so the list showed a
perfectly healthy session called "Login interrupted" — the CLI's startup
output, reading like a failure report. Sessions the TUI starts are now named
`w<n>-<case>` like the web UI's, and rowLabel() prefers the case directory over
a scraped prompt for any row with a mux name, since a LIVE pane is identified
by where it runs while a history row genuinely is its prompt.
2026-08-22 14:13:58 +02:00
Codeman maintainer bf860382a2 fix(tui): sweep an attach status bar a killed terminal left behind
restore() runs after spawnSync returns, which covers detaching and the agent
exiting inside the pane, but not the terminal dying while attached. Closing the
window or dropping the SSH kills the TUI where it stands, and the bar it
installed stays pinned on the session: the next attach wears a stale bar naming
a different session, and the pane is a row shorter for good. Seen on the beta,
where the tester closed the window instead of detaching.

One sweep at startup, fire-and-forget so it can neither delay the first frame
nor fail a start. Only a bar carrying our own marker is touched, and the marker
is now the single source of the bar's own wording so the two cannot drift; a
user's hand-written status bar on the same session is left exactly as it is.
The session goes back to `status off`, which is how Codeman creates every pane
it owns and the only state this bar is ever applied over.
2026-08-22 14:13:58 +02:00
Codeman maintainer aa487f13ce fix(tui): advertise the key that actually detaches, and stop tmux painting it green
Two things the attach status bar got wrong, both found in a beta test.

The bar read `Ctrl+B D`. tmux key tables are case-sensitive: lowercase `d` is
`detach-client`, capital `D` is `choose-client`. Pressing what the bar said
opened a client chooser and left the tester attached, with the way out on
screen and inert. The key is now READ from `list-keys -T prefix` the same way
the prefix already was, rather than hardcoded, so a rebound tmux is followed
too and the label cannot drift from the binding again. It never goes through
formatPrefixKey(), which uppercases.

The bar also rendered as a full-width bright green slab. Only `status-format[0]`
was styled, so tmux's stock `status-style` (`bg=green,fg=black`) stayed
underneath it and won; `#[reverse]` on top could not undo it. `status-style` is
now set explicitly to `bg=default,fg=default` and snapshotted/restored with the
rest, so the bar sits on the terminal's own background and reads as a hint
line.

Tests pin both: that the chord ends in lowercase `d` and never ` D`, that a
rebound key prints verbatim, that `status-style` is part of the banner, and
that parseDetachKey() picks `d` out of verbatim tmux 3.4 `list-keys` output
while ignoring `detach-client -a`/`-P`, which act on other clients.
2026-08-22 14:13:58 +02:00
Codeman maintainer 0919f9da62 fix(tui): make an attach fit the terminal, show the way out, and resume history
Three things the first beta test surfaced.

1. Attaching from a terminal of a different shape showed the pane clipped to the
   browser's size, with tmux's dot padding filling the rest. Codeman pins every
   window it owns to `window-size manual` at whatever the web client reports
   (tmux-manager.ts), so no attaching client can resize it. The handoff now
   brackets the attach with `window-size latest` and restores the snapshot on
   detach. `latest`, rather than a one-off resize to our own size, is also what
   lets a terminal resized MID-attach follow along: tmux recomputes on every
   SIGWINCH while the TUI is blocked in spawnSync and cannot.

2. Nothing on screen said how to get back out, because Codeman keeps the status
   bar off on its panes (the web UI carries that information around the terminal
   instead). The tester exited the agent looking for the exit, leaving a dead
   pane. An attach now wears a `status-format[0]` bar reading "<prefix> D
   detach, back to the codeman dashboard", with the prefix READ from tmux rather
   than assumed, and the session's options are put back exactly as they were on
   detach. One option, not status-left/status-right, so tmux draws no window
   list beside it; `reverse` so it inherits the terminal's own theme. Restoring
   an array option unsets the BASE name, since dropping the `[0]` index leaves
   an empty array, which renders as a blank bar on a session that had one. The
   help overlay names the chord, and the dashboard confirms the detach.

3. Enter on a RECENT row said resuming was not wired up. It now creates a
   session carrying that conversation (`resumeSessionId` plus `/interactive`,
   the path the web UI's Resume Conversation list already uses), in the
   directory it ran in and under its old name, then attaches to it.

   The attach mechanics deliberately sit in a method the group dispatch cannot
   reach, plus a re-entrancy flag: routing resume back through the Enter handler
   re-dispatched on "this row is RECENT" and spawned one session per pass, 35 in
   about 40 seconds on the beta before it was killed. test/tui/tui-e2e.test.ts
   pins one press to one session with a pane that never appears, which is
   exactly the case that looped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 75b272ff0a perf: pace the refetch and back the tail poll off a quiet pane
Both of the dashboard's periodic reads hit endpoints that are far more
expensive than their cadence assumed, and the cost lands on the SERVER's
event loop, so it is paid by every browser client too.

`GET /api/sessions/unified` is ~550ms against 11 live sessions: it scans
every Claude transcript plus the lifecycle log, uncached, and republishes
the search index. `scheduleRefresh()` was a 250ms trailing debounce with no
floor, and a queued refresh re-ran the instant the previous one returned
(by recursing, which also chained one pending promise per iteration), so a
stream of events paced the refetches at the endpoint's own latency: with
`session:updated` broadcast per session per 500ms while anything is
working, the scans ran back to back. `resyncDelayMs()` now keeps ambient
refetches 3s apart, measured start-to-start. The user's own actions call
`refresh()` directly and are unaffected, so what this paces is only
"notice what changed elsewhere".

`GET /api/sessions/:id/terminal` is ~80-100ms: two `execSync` tmux calls,
then the whole byte buffer normalized before the tail is taken. It was
polled every second for as long as a live row was selected. It now backs
off 1s, 2s, 4s, 5s while consecutive reads change nothing, and resets to 1s
on any change, when the selection moves, when this dashboard sends input or
answers a dialog, and on return from an attach. A pane that is printing is
still read every second; a pane at its composer is not.

The poll also kept running in three places it had nothing to draw for: the
whole time the user was attached in tmux (an attach can last hours), and
behind the message overlays that an async action opens (answered, killed,
started), which are not keystroke-driven and so never reached the
`afterInput()` path that stops it. `setInterval` becomes a chained
`setTimeout`, since the delay now varies.

Measured against the live server, same idle row selected, 25s window:
22 tail reads before, 5 after. With a working pane selected it stays at 22,
which is the intended cadence for a pane whose output you are watching.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 1385415e53 refactor: drop the two store members nothing consults
`TuiModelStore.confirmSatisfied()` and `approvalFor()` had no caller
outside their own tests. The first one mattered: it answered "does the
typed text authorize this kill?" with an exact name match, while the rule
actually consulted (`confirmAccepts()` in tui-app) also accepts the
8-character id prefix a mux name carries. Two divergent answers to one
question, the stricter one unreachable and waiting to be picked up by
mistake. knip cannot see class members, so the dead-code sweep never
flagged either.

The tests they existed for now assert observable state instead, and the
approvals one got stronger on the way: it checks that a session id coming
back does not inherit the dead session's dialog, which is the invariant
`removeSession()` is actually keeping.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 954a9ac26a fix: date a working row by its turn, not by the session age
`TuiSessionRow` declared `lastSubmitAt`/`inputTokens`/`outputTokens`,
`stateSince()` ordered the WORKING group by the first of them and
`renderRowLines()` painted the other two, but nothing ever filled any of
them in: the unified list carries none, and the `session:updated` payload
that does was discarded (an event only schedules a refetch).

So a running turn was dated by its SESSION's creation instead. Measured
against the live server before the fix: w65 (created 21h ago, turn started
one minute earlier) outranked w67 (created 15 minutes ago, turn started
five minutes earlier), the reverse of the rule docs/tui.md states, and the
elapsed column read `21h` for a turn a minute old. The token column was
unreachable code for the same reason.

`fetchLiveSessionMetrics()` reads the three fields from `GET /api/sessions`
and `applyLiveMetrics()` folds them onto the rows. That route answers from
the server's cached LIGHT state (no terminal buffers): 10-20ms measured,
against the ~550ms the unified list in the same `Promise.all` already
costs, so it is cheap enough to ride every refresh. It is best-effort like
the approvals and tmux reads beside it, because losing the anchor is
better than losing the list.

A ZERO is treated as unknown rather than merged: `stateSince()` reads
`lastSubmitAt ?? createdAt` and 0 is not nullish, so a merged 0 would date
every never-submitted session to the epoch.

The snapshot path gets the same merge, or `codeman tui --list` would number
the WORKING group differently from the dashboard that `codeman tui <n>`
indexes into.

Verified live: working rows now show 28m/8m (turn age, tokens 280.5k/65.2k)
where they showed 21h/34m and no tokens. The e2e assertion fails on master's
wiring with `[*] 10m` against a session that pressed Enter one minute ago.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 5fd6c5dd44 docs: extend the instance-isolation rule to tmux socket resolution
The data-dir half was already spelled out; the socket half only lived in
a function docstring, and the TUI is the first code that shells out to
`tmux -L` from a process that is not the server.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 8b5fc974b0 fix: cover tui in the CLI inventory and drop the em-dashes it printed
The inventory test predates the `tui` command, so a rename or an
accidental removal would have gone unnoticed: it now asserts the command,
its `-l`/`--list` flag and its optional position operand.

The digest and search-result lines joined their halves with an em-dash,
which the repo's own convention rules out, so both now use the middle dot
the surrounding lines already use. The one em-dash left in `src/tui/` is
load-bearing: `search-service.ts` builds a session snippet with it, and
the pattern that strips the repeated label has to match it.

Also moves `buildSearchEntries`'s doc comment back onto
`buildSearchEntries`; it had ended up stacked above a helper.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer a52abd9f96 fix: drop the two keymap and style entries nothing reaches
`mark()` had no callers (knip's only finding on this branch), and the
renderer's fallback help list advertised `r` resume, which is deferred
with the rest of phase 3: a help screen naming a verb the build does not
implement is worse than no help.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 008dfddc23 docs: document codeman tui
The user guide covers what the dashboard is (and is not), the two
non-interactive fast paths, the four groups and their ordering, the full
keymap, what answering an approval does server-side, and the SSH/narrow
and degraded cases. The example frame is a real 100x30 capture against
the E2E fake server, not a drawing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 32549789c7 fix: keep the plan-usage chip across a degraded-to-connected upgrade
A server that comes up mid-run was upgrading the header's hostname and
version but not its chip, which then stayed blank until the next telemetry
event. Also swaps a typographic apostrophe out of a preview error, which is
not renderable on the ASCII glyph tier.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 29d9a55eb6 fix: drop stale approvals and re-check the preview when the world changes
Two small honesty fixes at the edges: a server that goes down leaves the
dashboard holding prompts nothing can classify any more and whose answer
route is unreachable, so degraded mode clears them; and a resize can cross
the narrow breakpoint, where there is no preview pane to poll for.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer ef812236b0 fix: read a row-addressed repaint as lines in the preview
Measured against a live Claude pane: an Ink TUI paints by ROW and emits
almost no newlines, so dropping cursor-position sequences collapsed a whole
screen into one unreadable line, and a tail cut mid-sequence printed the
remains of it (";1H") as text. Now a jump to column 1 starts a display line,
a jump inside a row moves the write position (capped, since a stream may
address a column no terminal has), and a severed CSI head is dropped before
parsing.

The preview is readable against a real session as a result: tool calls, the
working line and the composer all land where they belong.

Also drop the repeated session name from a search row, whose snippet opens
with the name the row already shows in its first column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer b51abe2c27 test: drive the phase-2 verbs end to end under a pty
The fake API server grows the routes the dashboard now calls (terminal tail,
input, approvals answer, search, away digest, plan usage on status), and the
new cases assert on what the server RECEIVED rather than on the frame: the
prompt arrives as one line ending in a carriage return, and the answers as
the exact action and option digit.

Also covered: the tail refreshing in place, the search overlay selecting a
live session, the digest rendering, one bell for an item announced twice,
and the 409 path reported as "no longer on screen".

The plan-usage chip is punctuated with the glyph tier's separator, so an
ASCII terminal no longer gets a stray middle dot in the header.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer eb8958ddd0 feat: answer approvals and send prompts from the dashboard
The dashboard stops being read-only. The selected session's tail is polled
once a second while the plain list has focus and the layout is wide, and an
unchanged tail never reaches the model, so a quiet session costs no repaint.
A row with no live buffer says so instead of polling forever.

Keys: y/n and the parsed digits answer the selected session's dialog through
`POST /api/approvals/:id/answer` (never a blind keystroke: that route
re-captures the pane and 409s when the dialog has moved on, which the TUI
reports as "no longer on screen"); `p` opens a one-line composer aimed at
the selected session; `/` searches with a 250ms debounce and Enter switches
to a live session result; `g` shows the away digest. A new prompt rings the
bell exactly once, tracked by item id so a repaint or a refetch cannot
stutter, and the plan-usage chip rides `GET /api/status` plus its telemetry
event.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 953a560eee feat: render the approval card, composer, search and digest
The preview pane now leads with the pending dialog when the selected session
has one: the question, the options with their digits, and the keys that
answer them, red for a dialog and yellow for a waiting prompt. The card is
capped at half the pane, because the tail is why the pane exists.

Around it: a header badge counting prompts that need a human, a preview
title that sacrifices the path rather than the state word, the footer
becoming the composer line while one is open (with the cell the terminal
cursor belongs in, so it can be shown there and hidden everywhere else), and
the search and digest panels as overlays with a stable width.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer dc89f05b14 feat: hold composer, search and digest state in the TUI model
The store gains the three overlays phase 2 needs, each taking the keyboard
when it is set and all of them cleared together by closeOverlay(), plus the
pure flattening of `GET /api/search`'s typed groups into rows a cursor can
move over: headers are chrome, and only a session that is on the list counts
as selectable, since a history hit has no row to move the cursor to.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer bb73400afa feat: add the TUI's editor, approval and digest pure cores
Three small pure modules the phase-2 verbs are built on:

- tui-composer: the single-line editor behind `p` and `/`, holding text as
  code points so a cursor can never split a surrogate pair, with the scroll
  window derived from the width rather than remembered.
- tui-approvals: what an approvals-inbox item's card says, which keys are
  live for it (a digit answers only when the server parsed that option, and
  an idle prompt answers to none of them), and which ids the bell has not
  rung for yet.
- tui-digest: the away digest as compact lines, counts first and one line
  per entry, with a capped tail per section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 9d7dd2ab62 test: drive codeman tui end to end under a pty
Spawns the real command in a pseudo-terminal against a fake API server
(canned status/unified/approvals plus an SSE stream the test pushes
into), which is the only way to cover raw-mode key decoding, frames
reaching a terminal, SSE-driven refresh and the exit sequence that has to
restore the user's screen.

Two details the assertions depend on: frames are addressed absolutely
rather than newline-separated, so the parser takes the last COMPLETE
frame (the pty delivers one in several chunks, and reading a half-written
frame would be racy), and it reads the sidebar column only, or a name
echoed in the preview pane could answer for a row.

The child gets its own data dir and a tmux socket name nothing runs on,
so nothing here can see or touch the machine's real sessions.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 3f88226d50 feat: register the tui command with its two fast paths
`codeman tui` opens the dashboard, `codeman tui --list` prints the
numbered list and exits (the `sc -l` replacement, plain when piped) and
`codeman tui <n>` attaches straight to a row (the `sc 2` replacement).
Both fast paths short-circuit before any screen setup, and both refuse
the numbers path without a terminal instead of half-opening a UI.

Bare `codeman` still prints help: the web UI stays the primary surface.
The TUI module is imported lazily so the other commands do not pay for it
at startup.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer e5ae2826c6 feat: add the codeman tui dashboard
The IO half of src/tui: it owns the terminal, the timers, stdin and the
tmux handoff, and every decision it makes that is a function of its
inputs is an exported pure helper with unit tests (attach planning, the
typed kill confirmation, keymap selection, the repaint test, degraded
rows).

What it does: live session list over the unified API with SSE-driven
resync (debounced, with a 2s poll fallback the client asks for), cursor
and 1-9 navigation, attach and return, kill behind a typed confirmation
that refuses history rows and the session hosting the TUI, a new-session
case and CLI picker over quick-start, and degraded mode straight from
tmux when no server answers, re-probing so a server that starts upgrades
the dashboard in place.

Restoring the terminal is the part that has to be bulletproof: leave() is
idempotent and runs from normal quit, SIGINT/SIGTERM, a process exit hook
and prepended fatal handlers (src/index.ts already handles those by
exiting, so a listener registered after it would never run).

Attach is a handoff, never a proxy: the screen is restored and tmux gets
the real terminal. Inside tmux on the same socket there is nothing to
hand off to, so it issues switch-client and exits.

The preview pane, approvals answering, the prompt composer, search and
the digest are the next step; the region renders a placeholder rather
than pretending to load something.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 85b69e0923 feat: render the TUI picker overlay and a caller-supplied keymap
The footer and the help overlay held the plan's full keymap, which would
advertise verbs (prompt, search, digest, answer, resume) that the build
does not implement yet and teach users that the TUI ignores keys. Both
now take their entries from the render options when the caller passes
them; the built-in lists stay as the fallback.

The picker overlay windows its items around the cursor rather than
clipping them, so the selected case stays visible in a long list.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 6682231d68 feat: give the TUI model a revision signal and picker state
The app layer repaints on state change, so the store has to be able to
say that something changed: `revision` is bumped by every mutating
method, and the repaint test compares it against the last painted frame.
Without it an idle dashboard would either redraw on a timer or go stale.

Three additions come with it, all optional so nothing existing changes
shape: `TuiSessionRow.muxName` (the unified list carries no mux name, so
the app fills it in from the local tmux enumeration and a row without one
cannot be attached), a `new-session` UI mode, and `TuiPickerState`, the
one-column chooser behind `n`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer e140132e45 feat: add the TUI's API, SSE and degraded-mode client
Everything the dashboard needs from outside the process, behind one typed
surface, so the app loop stays a loop. It is a client of the running server and
nothing else: rows come from the unified list, blocked states from the
approvals inbox, and answering goes through the endpoint that re-captures the
pane and refuses with a 409 when the dialog has already been answered in tmux.
That refusal is a typed result rather than an exception, because a human
beating you to a prompt is normal operation.

Discovery mirrors the daemon probe (`CODEMAN_API_URL`, else loopback on
`CODEMAN_PORT`, self-signed TLS accepted) and credentials come from where
`codeman attach` already reads them. An explicit port outranks the ambient
`CODEMAN_API_URL`, which every managed session exports: a caller that named a
port must not be redirected at whatever server owns its shell.

Input is single-line and `\r`-terminated at this layer, so no caller can strand
text on an unsubmitted composer, and each send is tagged for the server's
exactly-once path. The event stream defaults to a `?sessions=` filter that
matches nothing, which drops the terminal firehose while lifecycle, hook and
approval events still arrive. A silent-but-open stream is caught by a watchdog
rather than a socket error, since that failure mode reports nothing at all.

With no server answering, sessions are listed from tmux on the instance socket
(argv, never a shell string) and decorated from a read-only peek at state.json,
which keeps the "the server died, get me to my sessions" path alive.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:58 +02:00
Codeman maintainer 6475c010a6 feat: decode the SSE wire format for the TUI
Node has no EventSource, so the live-update stream is read as raw bytes and
decoded here. Three details are what the parser exists for: a TCP read can end
between the CR and the LF of a CRLF, so a trailing CR is held back rather than
dispatched; the tunnel padding the server appends after a frame is a comment
with no blank line after it and must not split anything; and the keepalive is a
NAMED event, because an SSE comment is invisible to a browser client by spec.

Event classification lives here too, as a set rather than a prefix test:
`session:terminal` is most of the stream and the preview pane pulls its own
tail, so it is deliberately not a resync trigger.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 64c8048dda refactor: resolve the tmux socket from the instance config
The socket name was computed inside tmux-manager, which the TUI cannot import
just to learn which `-L` name its degraded-mode listing belongs on (that module
is the server's tmux driver, not a lookup table). The resolver moves next to
`dataPath()`, where the other half of the instance identity already lives, so
both processes agree by construction instead of by a copied default.

Behaviour is unchanged: the override still wins only when it is a name that can
be passed to `tmux -L` safely, and TmuxManager keeps warning about one that
cannot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 0566ea3453 docs: record why the key parser reads LF as Enter
Ctrl+J is unbindable as a result, which is worth knowing before someone tries
to bind it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 6ef3b2ba2e feat: render TUI frames from the model and layout
One absolutely-addressed line per row, each closed with an erase-to-end, so
nothing scrolls and a repaint cannot leave the previous frame's tail behind.
The caller wraps the result in synchronized-output brackets; that is an IO
decision and stays out of the renderer.

Color is passed in rather than detected. chalk's detection is right for the
one-shot CLI but would make a frame non-deterministic, so the palette is raw
SGR in the same semantic roles cli-style uses, and `color: false` emits nothing
but the cursor addressing, the session's own colors in the preview included.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 74fe2cad9f feat: add the TUI responsive layout math
Below 72 columns the preview pane is dropped and rows take two lines, the
constraint the `sc` chooser was built around and the reason it is still usable
on a phone; above it a clamped sidebar carries the list and the preview takes
the rest.

Every region is clamped to a non-negative size, so a 5x5 terminal degrades to a
header instead of handing the renderer negative widths.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer aa2deea73e feat: add the TUI session model, classification and cursor
Rows are the ones GET /api/sessions/unified already returns and blocked states
are the items the approvals inbox already parsed, both imported as types only
so a CLI process pulls in neither the server nor node-pty. Classification
speaks the web UI's language (red blocked, yellow waiting, green working) so a
user with both surfaces open never has to translate between them.

Groups order by how long a session has been in its state, which is why WORKING
anchors on the pane's last Enter: a working pane repaints about once a second,
so its last-activity stamp always says "now".

Selection is tracked by session id, never by row index: rows re-sort under the
cursor whenever a session starts working or an approval lands, and an
index-tracked cursor would quietly move the selection to another session
between two keystrokes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 5d7fdb528b feat: add the TUI raw-mode key parser
Decodes printable UTF-8, the control keys, arrows in both CSI and SS3 forms and
SGR mouse reports out of a byte stream that can tear anywhere, so a sequence
split across two reads decodes the same as one that arrives whole.

A lone ESC cannot be told from the start of an arrow key by looking at bytes,
so the parser holds it and the caller resolves it with flush() once its
disambiguation timer fires. Unknown sequences are swallowed: a stray CSI must
never reach a prompt composer as typed text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 64cf8384f2 feat: add the TUI's SGR-aware preview helpers
The preview pane shows a session's raw terminal stream, so it needs the tail
reconstructed rather than emulated: SGR survives, cursor steering and OSC do
not, and a carriage return returns to column 0 so a spinner that repaints its
line 200 times contributes one line instead of 200.

Widths count East Asian Wide characters as two columns, which the clip and pad
helpers rely on to never cut a wide character, a code point or an escape
sequence in half.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 596c08d20c chore: stop ignoring src/tui
The entry dates from an abandoned prototype (0.1427) and would have kept the
real TUI modules untracked while `git status` stayed silent about it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 1d0c3650f9 docs: fix the codeman attach description and the detach prefix
`codeman attach <path>` posts an attachment card for a local file; it
was described as attaching a Claude hook context. And Codeman never
overrides the tmux prefix for local sessions (only remote-SSH and docker
panes get C-q), so the detach hint is Ctrl+B D, matching the chooser.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer b9afd5a57e test: derive the CLI inventory from the real commander program
The file asserted against a hand-written fixture array with its own
argument parser, so it could not see a command being renamed, losing an
alias or disappearing, and it described a `tui` command that does not
exist. It now walks program.commands: names, aliases, subcommands,
option flags, operands, descriptions, and a guard against registering a
name or alias twice at one level.

Assertions are "at least this exists", so a new command (including the
tui one this plan adds later) passes without editing the test.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 14a911b4f6 fix: color the server startup line and its security warning
The startup banner is now the only one (the CLI printed a duplicate) and
is painted like the rest of the CLI. The non-loopback-without-password
warning was plain console.warn while the CLI's copy of the same warning
was yellow; chalk degrades off a TTY, so journald and web.log stay free
of escape codes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 9b1d269943 feat: wire the CLI through the style kit (doctor colors, spinners, confirm)
- doctor is colorized through the ReportStyle hook: verdict glyph and
  failing status text painted, paths and hints muted, versions left
  alone. `doctor --json` still prints raw JSON.
- `codeman web -d`, `web --stop` and `service install` block for up to
  30s polling /api/status; each now runs under a spinner instead of a
  silent terminal.
- `codeman reset` asks a real y/N question on a TTY. Non-interactive
  callers keep the old "Use --force to confirm." refusal, so no script
  can be answered by a question it cannot see.
- `codeman list` was a drifted copy of `codeman session list`; both now
  call one renderer, with the shorthand opting out of the stopped and
  web-server sections.
- `web` no longer prints its own "running at" line: the server prints
  one, and unlike this one it also covers the daemon and service paths.
- every chalk call goes through the palette, so the CLI has one place
  where colors are decided.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 4e5d0dcbd6 fix: measure the doctor table columns and let the CLI paint them
"Antigravity CLI" is 15 characters and the hardcoded padEnd(14) pushed
that whole row one column right. Widths now come from the widest cell.

The header always said the CLI layer may colorize, but there was no way
to: renderTable now takes an optional ReportStyle whose hooks are
identity by default, so the module still decides nothing about color and
its output stays byte-stable. Padding is applied outside the paint, so a
row with no path detail ends at its status text instead of trailing
spaces inside a color run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer f9d6c4f0c3 feat: shared CLI style kit
One vocabulary for everything the codeman CLI prints: semantic palette,
the glyph set the commands already used, heading/rule/kv, width-aware
table layout, a stderr spinner and a y/N confirm.

Color detection stays chalk's, so NO_COLOR and non-TTY degradation keep
working with no second detector to disagree with it. The layout math and
glyph selection are pure and exported, which is what lets the dependency
report reuse them while staying color-free.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 09d6bb9eb0 docs: TUI rework plan (codeman tui, herdr research)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer a922d301b1 chore: version packages 2026-08-21 20:24:38 +02:00
Ark0N abca552676 Merge pull request #327 from dignfei/fix/terminal-ime-punctuation
fix(terminal): preserve IME punctuation input
2026-08-21 20:23:26 +02:00
Ark0N 12a996b107 Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay
2026-08-21 20:23:17 +02:00
d fei 458e751a33 fix(terminal): keep shell history loading explicit 2026-08-22 01:55:50 +08:00
d fei dab432b3fd fix(terminal): bound shell history replay 2026-08-21 08:23:31 -04:00
Codeman maintainer 79a0399552 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:47:10 +02:00
Codeman maintainer 61251c0b94 fix(cli-resolvers): negative-result caching, SIGKILL on probes, restored VITEST hermeticity, wired not-found diagnostics
Post-merge follow-ups for PR #329 (shared CLI executable resolution):

- Negative-cache resolution misses with a doubling backoff (1min -> 5min
  cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
  shared resolver cached success only, so a missing CLI re-ran the whole
  chain - ending in a synchronous interactive login-shell spawn bounded by
  the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
  attempt, stalling the event loop each time, forever. Success still caches
  for the process lifetime, so an installed CLI is picked up within minutes
  without a restart. Tests drive the backoff via an injectable clock
  (createCliExecutableResolver `now` option, threaded through the
  createPiResolverForTest / createAntigravityResolverForTest wrappers).

- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
  pi/claude --version probes: execFileSync's timeout only SENDS the kill
  signal and then keeps waiting for the child to exit, and interactive bash
  ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
  survived the timeout and blocked the server permanently.

- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
  test pinned the deletion): under vitest the production resolver host now
  replaces un-injected IO primitives with inert stubs - no real PATH
  scanning, no login-shell spawns - and probePiVersion never executes a
  `pi` candidate again (`pi` is a generic binary name, so route tests
  hitting /api/pi/status executed whatever binary the machine carried).
  Tests opt in through the runCommand/isExecutableFile injection hooks or
  allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
  test is replaced by behavioral pins, including a real-executable fixture
  in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
  is ever removed again.

- Wire the six get*NotFoundMessage() exports (previously dead) into their
  intended call sites: the createSession throws in tmux-manager and the
  availability gates on POST /api/sessions and POST /api/quick-start in
  session-routes, replacing a third hardcoded copy of the text. A not-found
  error now names where resolution looked (server PATH, login shell,
  checked directories). npm run knip no longer reports any unused export
  from the resolver modules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:37:58 +02:00
Codeman maintainer bb4ba79791 fix(repo-status): async git, single-flight TTL cache, credential redaction, local-upstream parse
Post-merge follow-ups for #328 (GET /api/system/repo-status):

- Event-loop blocking: every git invocation in repo-status.ts is now async
  (promisified execFile), never execFileSync — the per-remote ls-remote +
  fetch could hold the event loop (SSE, PTY streaming) for up to ~60s per
  request. The whole computation is single-flight with a 45s TTL cache
  (createSingleFlightCache): concurrent requests share one in-flight
  promise, a fresh result is served without spawning git, and a rejected
  compute is never cached. Route handler shape and response fields
  unchanged; remotes still processed sequentially (concurrent fetches in
  one repo contend on ref locks).

- Credential disclosure: the redaction from git-clone.ts is extracted as
  exported redactGitCredentials() (sanitizeGitOutput now uses it) and
  applied via redactRemoteStatus() to every remote card's url and error
  string, so a scheme://user:token@host remote URL (or git stderr echoing
  it) never reaches a client.

- Non-interactive env: runGit() now uses the shared gitNonInteractiveEnv()
  instead of a partial GIT_TERMINAL_PROMPT/BatchMode env, also closing the
  GIT_ASKPASS/SSH_ASKPASS/SSH_ASKPASS_REQUIRE/DISPLAY/GCM_INTERACTIVE
  prompt paths.

- Upstream parse bug: a local-branch upstream (@{upstream} with no slash,
  e.g. after `git branch -u otherbranch`) made slice(0, indexOf('/')) into
  slice(0, -1) and yielded garbage like "maste". parseTrackingRemote()
  (pure, unit-tested) returns null for it, and the bare ref is dropped so
  it cannot be mistaken for a remote-tracking ref downstream.

Tests extended in test/repo-status.test.ts (parseTrackingRemote,
redactGitCredentials/redactRemoteStatus, createSingleFlightCache
single-flight/TTL/rejection semantics).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Codeman maintainer d7ad73bc9b fix(response-viewer): role on full-context blocks, divider ReDoS, pi mode (#326 follow-up)
Three post-merge fixes for the external-CLI response viewer:

- ?context=full blocks now carry role ('user' for prompts, 'assistant'
  for response/status/tool). The frontend's loadFullContext() renders
  via msg.role, so the roleless blocks lost the "You" badge and every
  turn rendered as the agent. kind/label/text are unchanged and the
  frontend needs no change.

- normalizeDividerStatusLine() dropped its backtracking regex
  (/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
  a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
  minutes at 10,000), and pane text is agent-controlled with buffers up
  to 32MB. Replaced by a linear counter walk with the identical accept
  set and captured content, pinned char-for-char against the old regex
  by a brute-force corpus test plus a hostile-input regression test
  that fails by timeout with the RegExp version (same approach as the
  glob-matcher hardening in 68ae9a8).

- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
  empty-viewer symptom the transcript branch exists to fix. The list
  stays a local duplicate of isExternalCliMode() (importing session.ts
  would drag node-pty into the pure module); a new exhaustive parity
  test asserts the two mode sets can no longer drift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Codeman maintainer 96ee8b536d docs: update the tap-report gate description after #325
#325 renamed _sessionUsesServerMouseStrip to _shouldReportMouseToCli and
added the server-observed cliMouseTracking half of the gate, which also
turned codex tap reports from measured no-ops into not-sent-at-all. The
invariants paragraph still described the old name and the old behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:13:26 +02:00
Ark0N b7f3b07c79 Merge pull request #330 from aakhter/ralph-loop-reschedule
Ralph loop silently stops polling after two ticks
2026-08-21 02:10:41 +02:00
Ark0N 2073a1b185 Merge pull request #328 from aakhter/repo-status-panel
Report repository status for git-clone installs (GET /api/system/repo-status)
2026-08-21 02:10:38 +02:00
Ark0N f9a8493823 Merge pull request #329 from aakhter/cli-login-shell-resolution
CLIs installed via nvm/Homebrew are not found when Codeman runs as a service
2026-08-21 02:10:35 +02:00
Ark0N d2711ef092 Merge pull request #326 from aakhter/response-viewer-external-cli
Response viewer is empty for OpenCode / Gemini / Antigravity sessions
2026-08-21 02:10:29 +02:00
Ark0N 30a15adbd6 Merge pull request #325 from Ark0N/feat/auto-copy-selection
feat(terminal): Auto Copy, put a finished selection on the clipboard
2026-08-21 02:10:19 +02:00
Aamer Akhter a35438ba34 fix(ralph): loop stops rescheduling after two ticks
The reschedule guard is `this._status === 'running' && this.loopTimer === null`,
but the timer callback never nulls `loopTimer`. So the handle stays non-null from
the first fire onward, the guard is false on every subsequent pass, and the Ralph
loop silently stops polling after exactly two ticks.

It stops without changing status: `status` stays `running`, `stop()` is never
called, and no error is raised — the loop just quietly never runs again, which is
what makes it hard to notice on a long autonomous run.

Null the handle inside the callback before re-entering `runLoop()`, which is the
pattern `orchestrator-loop.ts` already uses for its own reschedule.

Test: a regression case in test/ralph-loop.test.ts that runs a real 5ms-interval
loop for ~16 intervals and asserts it ticks at least 3 times. Against the unfixed
source it reports exactly 2.
2026-08-20 12:58:17 -04:00
Aamer Akhter fef903df98 fix(cli-resolvers): find CLIs installed via nvm/Homebrew when running as a service
A CLI installed by nvm, Homebrew or a user-level npm prefix lives on a PATH that
only a login shell sets up. Codeman running under systemd or launchd does not get
that PATH — launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` — so every
resolver reported the CLI as unavailable on installs where it is plainly there
and works from a terminal.

Each of the six resolvers had its own hand-rolled copy of the same PATH walk, so
the fix is factored into one shared `createCliExecutableResolver()` with an
explicit lookup order: the server process PATH, then common install directories in
order, then an interactive login shell as the last resort. Only the last step
spawns anything, and only when the cheap lookups have already missed.

Also adds `formatCliNotFoundMessage()`, so a failure explains where it looked
instead of just asserting the CLI is missing. Its diagnostics are bounded and
control characters are flattened, so a not-found message cannot dump arbitrary
environment data.

Success is cached and failure is retried, so installing a CLI while the server is
running is picked up without a restart.

Net -103 lines across the six resolvers. Behaviour is unchanged wherever the CLI
was already on the process PATH: that remains the first thing checked.

Tests: 20 cases in test/cli-executable-resolver.test.ts covering the precedence
order, login-shell-only resolution, the caching rule, unsafe-name rejection, and
the bounded diagnostics.
2026-08-20 12:47:42 -04:00
Aamer Akhter 02e7d3fcba feat(system): report repository status for git-clone installs
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.

Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.

Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.

Tests: 24 cases in test/repo-status.test.ts.
2026-08-20 12:39:28 -04:00
d fei f744719650 fix(terminal): preserve IME punctuation input 2026-08-20 10:35:23 -04:00
Aamer Akhter 63c5ba89da fix(response-viewer): populate the viewer for OpenCode/Gemini/Antigravity panes
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.

For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.

Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.

The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.

Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
2026-08-20 09:24:37 -04:00
Codeman maintainer 7fc4784d0f fix(approvals): clear the red tab alert when a dialog is answered in the terminal
Confirming an AskUserQuestion left its tab flowing red for the rest of
the turn (owner report: ~8 minutes on a running session, with no dialog
anywhere on screen). Two separate bugs, both live-verified.

The re-capture erased the evidence the staleness check runs on. Claude
Code fires the Notification behind the dialog (measured 6-7s on v2.1.237,
documented up to ~30s), so the 600ms re-capture routinely lands on a
frame the user has ALREADY answered, parses nothing, and applyCapture
overwrote item.options with undefined. A MISSING options is how "we never
could read this dialog" is expressed, and those items stay answerable by
design, so a cleared field was indistinguishable from a never-parsed one
and the item became permanently unsweepable: it survived every
GET /api/approvals and every page reload, cleared only on `stop`, and
still accepted an answer, sending a bare `1` into a composer with no
dialog under it. applyCapture is now ADD-ONLY for options.

Nothing ran the staleness check while a page was open. It lived only in
GET /api/approvals, which seedApprovals() calls on init and reconnect, so
`stop` was the first thing that ever cleared an answered dialog. The
`working` signal now runs the pane-VERIFIED variant (resolveIfDialogGone
-> verifyStillAnswerable): the heuristic only decides when to look, the
screen decides the outcome, so the existing "working can flap" rule is
respected.

A frame that parses no options is now conclusive in two cases, and only
those, so an unreadable capture still keeps the alert: the item once
parsed options, or the frame shows Claude actively running a turn. A
modal dialog BLOCKS the turn, so the two cannot coexist - measured, a
live-dialog frame carries neither the elapsed-timer spinner nor the
"esc to interrupt" footer, which the dialog replaces with "Enter to
select". That second signal is reached by a delayed staleness pass (3s)
scheduled alongside the re-capture, which closes the late-hook case where
the prompt is answered before the hook lands: nothing ever parses, `stop`
may have gone by already, and the alert outlived reloads until the 12h
TTL. The pass is deliberately later than RECAPTURE_DELAY_MS, whose whole
reason for existing is that the hook can beat Ink to the screen.

Frontend: _onHookElicitationComplete cleared only the elicitation entry,
but an AskUserQuestion arrives as permission_prompt, so it was clearing
the wrong alert; it now clears both, matching the server's kind-agnostic
APPROVAL_RESOLVING_EVENTS.

Verified end to end on an isolated beta instance, not just in unit tests:
before, resolution could only come from the stop route (approval:resolved
always immediately preceding hook:stop); after, it arrives from the new
paths, and a simulated late hook resolves at +3.12s with no stop, no
working signal and no GET, while the pane is still working. Tests use
frames captured off a live pane and each new one was confirmed to fail
against the old behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 12:18:16 +02:00
Codeman maintainer fa7e700834 fix(terminal): report a click to the CLI only when it asked for the mouse
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:

    $ [<0;88;20Mecho hello
    bash: 0: No such file or directory

The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.

What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.

Details that are easy to get wrong:

* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
  an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
  without turning tracking on is not asking about clicks, and counting
  those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
  cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
  broadcastSessionStateDebounced: the flag flips when a dialog opens, and
  the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
  the CLI re-emits its DECSET, which tmux does at client attach.

Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 11:21:29 +02:00
Codeman maintainer 7936a75e28 feat(terminal): Auto Copy, put a finished selection on the clipboard
App Settings > Terminal & Input > Selection & clipboard > Auto Copy
Selection (`autoCopySelection`, per-device, default OFF). With it on,
highlighting text in the terminal copies it: mouse drag, double-click
word, triple-click line, and the phone long-press selection. Ctrl+C is
untouched and still copies on demand.

Three things decide the shape of it:

* It fires at the END of a gesture, never in onSelectionChange. That
  callback runs for every cell a drag crosses, so copying there would be
  one clipboard write per mouse move. It only arms a pending flag; a
  document-level mouseup listener flushes, and the touch path calls the
  flush itself because it preventDefaults its touchend and no mouseup
  ever arrives there.
* The flush is synchronous inside the handler, because both clipboard
  paths need user activation: Firefox gates navigator.clipboard
  .writeText on it, and execCommand('copy'), the fallback the plain-HTTP
  LAN install lands on, has to run in the gesture's own task. A timer or
  a wait for onSelectionChange loses it, invisibly in Chrome.
* It deliberately does NOT do what copyTerminalSelection() does. That
  one clears the selection (so a second Ctrl+C is an interrupt) and
  focuses the terminal. Clearing would make text vanish under the cursor
  that just highlighted it, and focusing opens the on-screen keyboard
  over it on a phone. Focus is instead restored to whatever held it,
  which only matters for the execCommand fallback.

Guards are pure in decideAutoCopy() (constants.js): off, blank or
whitespace-only text, and a 1M-char cap, since a drag off the top of the
viewport autoscrolls and one gesture can sweep the whole 50k-line
scrollback. Past the cap the copy is refused rather than truncated, with
a toast pointing at Ctrl+C.

Feedback is silent on success except once per page load, so a feature
that works by doing nothing visible can still be told from a dead
toggle; failures and refusals toast, throttled to 10s.

Per-device on both counts the settings rule requires: in `displayKeys`
and absent from the .strict() SettingsUpdateSchema, because clipboard
access differs by device and by origin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 10:46:11 +02:00
Codeman maintainer 07b9c7fd7b fix(terminal): remove the unreachable copyTerminal(), closing out #322
The last two items of #322: copyTerminal() copied the entire buffer but
was wired to no button, shortcut or call site anywhere, and it wrote
through navigator.clipboard directly, which is undefined on the
plain-HTTP LAN install, so it would have failed there even if it were
reachable. Everything that actually copies goes through
copyTerminalSelection() and _copyText's execCommand fallback; whole-
buffer copy, should anyone want it, is a selectAll() away from that
same working path.

Closes #322

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 00:06:20 +02:00
294 changed files with 48861 additions and 4174 deletions
+18
View File
@@ -0,0 +1,18 @@
.git
.agents
.claude
.codex
# `**/` matters: a .dockerignore pattern is matched against the WHOLE
# context-relative path, so a bare `.env` excludes ONLY the root file and
# `COPY . .` would bake docker/.env -- CODEMAN_PASSWORD and any provider API
# keys -- into the published image at /opt/codeman/docker/.env (verified).
**/.env
**/.env.*
!**/.env.example
node_modules
dist
coverage
out
test-results
tmp
*.log
+12 -4
View File
@@ -69,10 +69,18 @@ shared-host, multi-user, or tunneled deployments.
- **Multi-instance tmux socket is process-wide.** Two Codeman instances on the same `CODEMAN_INSTANCE` share a tmux socket and can attach each other's live sessions — isolate with distinct `CODEMAN_INSTANCE` values.
- **The live log-tail route reads `/var/log` and `~/logs`** in addition to the session working directory (read-only) — a deliberate choice for tailing system/app logs. On a password-protected remote deployment an authenticated user can therefore read those roots outside their session. See `docs/security-architecture.md` §5.
Recent hardening (this release): web-push subscription endpoints are restricted
to https public hosts (SSRF guard — rejects internal/metadata IPs, validated at
subscribe and send time), and tmux session names discovered on the shared socket
are validated against the safe-name pattern before reaching any shell call site.
- **The web-tab proxy fetches from the server's network position.** Any authenticated user can save a dashboard URL on loopback or a private range and have Codeman relay to it; that is the feature. Link-local and cloud-metadata addresses are the only refused targets (see below). On a shared host, restrict who holds an account.
Recent hardening (2026-09-04): the web-tab proxy, its "Test" probe and its
WebSocket relay refuse link-local and cloud-metadata targets (`169.254.0.0/16`,
`fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`,
`metadata.google.internal`), judged on the RESOLVED address so a DNS name pointing
there is refused too; proxy capabilities are revoked on logout, admin logout and
user deletion; proxied responses carry `Referrer-Policy: same-origin`. Earlier:
web-push subscription endpoints are restricted to https public hosts (SSRF guard,
rejects internal/metadata IP literals, validated at subscribe and send time), and
tmux session names discovered on the shared socket are validated against the
safe-name pattern before reaching any shell call site.
For the detailed rationale, defenses, and recommended secure setups, see
[`docs/security-architecture.md`](../docs/security-architecture.md).
+3 -2
View File
@@ -2,6 +2,9 @@
.agents/
skills-lock.json
# In-session decision scratchpad (context-survival mechanism, not a deliverable)
DECISIONS.md
# Written by install.sh into end-user clones when setup finishes
.install-complete
@@ -93,8 +96,6 @@ packages/gesture-control/.vite/
# Claude Code plan tracking
plan.json
# Unfinished TUI (local development only)
src/tui/
.claude/
media-assets/
commands
+2 -1
View File
@@ -2,7 +2,8 @@
Canonical agent/contributor guidance for this repository lives in [CLAUDE.md](CLAUDE.md) —
project structure, build/test/lint commands, code style, testing safety rules
(never run the full suite inside a managed tmux session), security notes, and
(`npm test` is the CI gate and is safe to run bare; the three excluded suites
have their own runners), security notes, and
the deployment workflow are all maintained there. Please read it before making
changes, and keep it the single source of truth rather than duplicating
sections here.
+326
View File
@@ -1,5 +1,331 @@
# aicodeman
## 1.24.7
### Patch Changes
- The web-tab proxy refuses link-local and cloud-metadata targets. Its Test probe, the proxy itself and the WebSocket relay accepted any http(s) host, so a saved dashboard URL could reach `169.254.169.254` (in decimal, hex, IPv6-mapped or DNS-name form) through a capability and no cookie. Loopback and RFC1918 addresses stay allowed on purpose, since a localhost Grafana is the feature; only link-local and the fixed cloud-metadata addresses are refused, at the schema, at every connect site, and through a DNS lookup hook that judges the resolved addresses, which is what closes DNS rebinding. Adds `undici` so the proxy runs its fetch through its own agent.
Proxy capabilities are revoked on logout. `revokeOwner()` had shipped with no caller, so a leaked proxy URL stayed valid for as long as anything kept polling it. `POST /api/logout`, the admin forced logout and user deletion now revoke the capabilities they should, and proxied responses carry `Referrer-Policy: same-origin` with the upstream's own policy dropped, so a dashboard on a loose referrer policy cannot hand the capability to a third-party host it links to.
The Docker Compose deployment updates itself from App Settings again (#373, @opticon454). The checkout Compose builds from is bind-mounted at `/opt/codeman`, so an update's `git checkout` and rebuild land on the host and survive container recreation; build artefacts live in named volumes so container-compiled native modules never enter the host checkout; the image keeps devDependencies and a build toolchain; and the restart is the server exiting under `restart: unless-stopped`. An in-place update applies code only, so the updater refuses a release that changes `server.Dockerfile` or `docker-compose.yaml`, or that adds keys to `.env.example` the user's `.env` has no value for (Compose interpolates an unset variable to the empty string and starts anyway), and points at `docker/Start-Codeman.sh` on the host instead. The four global agent CLIs in the image are pinned. A follow-up makes the final step fail safe: the server exits only when the Compose file declares `CODEMAN_RESTART_BY_EXIT=1` or the daemon confirms an auto-restart policy, and otherwise the build is staged for a manual restart, so a container nothing would restart is never taken down. Details in `docs/docker-self-update.md`.
The test suite strips `CODEMAN_INSTANCE`, `CODEMAN_DATA_DIR` and `CODEMAN_TMUX_SOCKET` before any application module loads (#371, @opticon454), with a two-half test whose static half reads `test/setup.ts` so a dropped line fails everywhere. This replaces the throwaway data dir #356 had set for the same variable.
### Thanks
- @opticon454 for the Compose self-update (#373) and the test isolation fix (#371).
## 1.24.6
### Patch Changes
- CLI backends are now a data-driven registry (#347, @opticon454). Every run mode (Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP) is a `CliEntry` in `src/config/cli-registry/`: binary discovery (search dirs, version and identity probes), the launch argv template, environment handling, the multi-user privileged-parameter and privileged-env-key clamps, the remote and Docker pane commands, and the capability flags the rest of the app reads instead of branching on a CLI's name. `~/.codeman/clis.json` can override any stock entry or add a custom CLI; it is read-only in this release, must be mode 0600, and every reason it was ignored is now logged once on first load (`docs/cli-registry.md`). Config never contains shell text: entries declare typed argv tokens, literals are validated at load time, and values resolve through patterns named in code. Registry data resolves at call time rather than at module import, so a CLI enabled while the server runs moves every surface at once, and a guard test fails the build if per-CLI-id branching reappears outside the stock catalog.
This is an internal refactor. The spawn command every CLI receives is byte-identical to the previous hand-written builders, verified by pinned golden strings in the test suite and by diffing both implementations across 11,602 option combinations for all ten modes. Five small deliberate changes ride along: the in-container version probe derives the binary from the registry (`antigravity` runs `agy`), the remote version probe now covers Grok and DeepSeek, `codeman doctor`'s CLI rows are generated from the registry (Claude's install hint is the install command, five CLIs gain hints, the row order follows the catalog), OMP now requires tmux like its siblings instead of silently falling back to a direct PTY, and an OMP session's attach client now receives `COLORTERM=truecolor` like the other truecolor CLIs.
Remote sessions are no longer auto-revived after a clean agent exit (#355, @timkjr). The reconnect watcher could not tell a transport drop from a Ctrl-C, Ctrl-D or `exit` inside the remote CLI, so a clean exit relaunched a fresh agent (OpenCode and OMP started a new conversation every time; Claude only looked fine because its `--resume` fallback masked it). The watcher now revives a dead pane only when the durable remote tmux session is verifiably still alive, via a `has-session` probe over ssh, and an unreachable host means do not revive. A follow-up classifies that probe by exit status, since `tmux has-session` prints nothing on success and reading its stdout had marked every live session as gone, forgets the cached answer whenever the pane is seen alive again so a stale result cannot revive a later clean exit, and caps the probe at one in flight per session.
The test suite can no longer reach the production `~/.codeman` data dir (#356, @timkjr). `test/setup.ts` now points `CODEMAN_DATA_DIR` at a throwaway directory, which is the absolute override that bypasses the suite's temporary HOME when inherited from the shell, and every test that deletes a case tree goes through a containment gate that refuses paths outside the temporary HOME. A bare suite run had overwritten a real `remote-hosts.json` with a route test's fixture. The comments around it and CLAUDE.md's testing section now name that variable as the cause; `os.homedir()` itself does follow `$HOME`.
### Thanks
- @opticon454 for the CLI registry (#347), the phased resubmission of #343, and the review rounds that hardened it.
- @timkjr for the remote auto-revive fix (#355) and the test-isolation sweep (#356).
## 1.24.5
### Patch Changes
- Fable 5.1 is selectable in App Settings.
`claude-fable-5-1` is in Claude Code's model catalog (display name "Fable 5.1", June 2026 knowledge cutoff), but the model picker only went up to Fable 5, so pinning it meant hand-editing a case's `.claude/settings.local.json`. It now appears as a card under **App Settings -> Models -> New Claude sessions**, and as an option in **Task routing** (Default for tasks, plus the Explore / Implement / Test / Review overrides).
It is offered exactly the way Fable 5 already is: the "1M capable" badge, the 1M context window switch stays live for it, and base + switch compose into `claude-fable-5-1[1m]`. Both strings are accepted by the CLI.
Deliberately not claimed: that a 1M window is what sets Fable 5.1 apart. The CLI's model catalog marks both fable entries as natively 1M with the same window, so an always-on window for 5.1 next to a switchable one for 5 would encode a difference the models do not have.
### Thanks
- @shenlvkang-collab for #370, which surfaced that Fable 5.1 was missing from the picker.
## 1.24.4
### Patch Changes
- The Compose deployment image ships the Docker CLI instead of the whole Docker engine.
`docker/server.Dockerfile` installed Debian's `docker.io` to get a client for the mounted
host socket. That package is the full **engine**: even with `--no-install-recommends` it
pulls 15 packages including containerd, runc, dmsetup and iptables, none of which a
container that only talks to a socket can use. It also ships Docker 20.10.24, from 2023.
The CLI and the buildx plugin are now copied from the official `docker:29-cli` image
instead. Measured on the same `node:22-bookworm-slim` base: **266 MB → 108 MB**, a 158 MB
saving, with the current CLI (29.7.2) in place of a two-year-old one.
Verified by building the real image and running it: the binaries are static Go builds, so
they work on this glibc image even though they come from an Alpine one, and `docker
--version`, `docker ps` and `docker build` all succeed against a mounted host socket as
the unprivileged runtime user. buildx is copied deliberately — `scripts/build-agent-image.mjs`
shells out to `docker build` and Codeman auto-builds the agent image on the first Docker
case, which without the plugin falls back to the classic builder Docker has deprecated.
`docker-compose` is not copied; Codeman never shells out to it.
## 1.24.3
### Patch Changes
- Docker Compose deployment, and the plan-usage chip stops losing its 5-hour window.
**Run Codeman itself in a container** (#349, @opticon454). `docker/` now carries a
local-image Compose deployment: copy `docker/.env.example` to `docker/.env`, set
`CODEMAN_PASSWORD`, run `bash docker/Start-Codeman.sh`. Docker cases then start as
**sibling** containers through the mounted host socket rather than nested ones, which
inverts an assumption the bare-host path takes for granted: the daemon no longer shares
Codeman's filesystem, so a bind source that is valid inside Codeman means nothing to it.
`CODEMAN_DOCKER_HOST_HOME` translates sources under HOME into the daemon's namespace and
`CODEMAN_CASES_PATH` points the cases dir at a host-absolute bind mount, so a workspace
resolves to the same absolute path on both sides. `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1`
drops `--memory-swap` for hosts without swap accounting (`--memory` still applies) and
filters only that one kernel warning. Guides: `docs/docker-compose.md`, `docker/README.md`.
Three things were fixed while landing it:
- **`docker/.env` was being baked into the image.** A `.dockerignore` pattern matches the
whole context-relative path, so the bare `.env` line excluded only the root file while
`COPY . .` picked up `docker/.env` — the file the deployment's own README tells you to
fill with `CODEMAN_PASSWORD` and provider API keys — and left it at
`/opt/codeman/docker/.env`. Now excluded via `**/.env`, verified in both directions
against a real build context with a canary secret.
- **`codeman skill install --case <name>` could not find a case under Compose.**
`CODEMAN_CASES_PATH` moved the server's cases dir but not the CLI's, which still
hardcoded `~/codeman-cases`. Both now resolve through one place.
- **A Docker case handed its Claude conversation id to every other CLI.** `resumeOnStart`
seeded `dockerResumeId` from `lastClaudeSessionId` regardless of mode, and
`appendResumeFlag()` maps a resume id onto codex/gemini/pi/grok/deepseek/omp/antigravity.
This one is a plain master bug, unrelated to Compose.
**The plan-usage chip keeps its 5-hour slot.** It silently shrank from `5h 4% · 7d 52%`
to a lone `7d 52%`, which reads as half the feature breaking. Nothing was broken: Claude
Code ships `rate_limits.five_hour` "only while the API reports it and its resets_at has
not passed", so between 5-hour session windows the key simply leaves the statusline
payload. The slot now stays with a dimmed em dash and the tooltip says "no active session
window". Claude only — a missing Codex bucket means that plan has no such limit, so those
stay omitted.
### Thanks
- @opticon454 for #349, and for a write-up that made an infrastructure PR quick to review
## 1.24.2
### Patch Changes
- Fix every new claude session dying on Claude Code 2.1.252's rewritten folder-trust dialog.
That dialog used to offer `❯ 1. Yes, I trust this folder` / `2. No, exit`, so Codeman
answered it by pressing Enter on the highlighted default. 2.1.252 dropped the numbers,
reversed the options and highlights `No, exit`, so the same Enter now answers _exit_: a
session in any directory claude had not seen before died (`Pane is dead (status 1)`)
about six seconds after it started, before the agent ever drew a composer.
- `trustDialogNextKey()` (`src/session-trust-dialog.ts`) now reads the `❯` marker off
the rendered pane and returns ONE keystroke at a time: an arrow while the cursor is on
the wrong option, Enter only once the screen shows it on the trust option. A frame it
cannot read presses nothing. Both the 2.1.252 and the older numbered layout are
handled, and the direction is derived from the frame rather than assumed, so a further
reordering costs a repaint instead of a session.
- The scan schedules its own follow-up read. It had only ever run from the PTY data
handler, which was enough while one Enter answered the dialog; the arrow that moves the
cursor is the last output the pane produces, so a two-keystroke answer would otherwise
stall with the cursor sitting on the right option forever. The keystroke cap goes from
3 to 6 for the same reason.
- The bundled `codeman` agent skill gets the same treatment (preamble 1.21.0): its
`_accept_trust` fallback reads `terminal?full=1`, steers onto the trust option and
confirms only after re-reading, instead of posting a blind `\r`. It sends those
keystrokes under its own `clientId`, because input sequence numbers are monotonic per
client and spending prompt numbers on dialog keys would make the next send-and-wait
look like a stale duplicate and vanish silently.
- Readiness recipes in `docs/extending-codeman.md`, `docs/api-reference.md` and the
skill's own reference carry the corrected answer and a new symptom-table entry for a
worker whose pane is dead seconds after the spawn.
Also included: a CLAUDE.md audit against the tree, correcting counted drift (route
modules, handler counts, frontend module count and app.js size, install.sh size) and
documenting several subsystems that had no entry.
## 1.24.1
### Patch Changes
- The Docker agent base image builds again.
**`docker/agent.Dockerfile` could not be built from a fresh checkout** (#352, fix in #350): the DeepSeek Harness step died with `dsh: pnpm not found on PATH` and exit 127, which took the whole image with it and, because Codeman auto-builds this image on the first Docker case, left Docker mode unusable on a clean host. `dsh plugin` does not bundle a package manager; it spawns a literal `pnpm` with no npm fallback, so pnpm is now installed alongside `dsh` and the layer proves it with `pnpm --version`.
The profile install also passes `--config.dangerouslyAllowAllBuilds=true`, because pnpm, unlike npm, refuses dependency lifecycle scripts by default and fails the install over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1). Which packages that hits moves between rebuilds, since the terminal profile is resolved by dist-tag rather than pinned: the tree that broke the build in August pulled `@google/genai`, today's does not. An allowlist of those names would have gone stale rather than prevented the next break, and running those scripts is the same exposure the image already accepts three layers up, where `npm install -g` runs the install scripts of every transitive dependency of the five CLIs above it with no gate at all.
Documentation caught up with two things it had wrong: the image smoke test in `docs/docker-cases.md` now covers `dsh` and `omp`, and checks the dsh **profile** rather than only the binary (`dsh` is a launcher, so `dsh --version` says nothing about whether a session can start), and `docs/deepseek-integration.md` names pnpm as a prerequisite for installing a terminal profile at all, by hand or through the UI button. A comment in the `/api/deepseek/install-profile` route claimed the opposite of what this bug proved, and is corrected; the route's behaviour was already right, surfacing dsh's own "pnpm not found on PATH" line as the install error.
### Thanks
- @opticon454 for #350, with a reproduction that made this a confirmation rather than a hunt
- @timkjr for reporting #352, and for finding it while verifying Docker support for someone else's PR
## 1.24.0
### Minor Changes
- OMP (Oh My Pi) as a tenth run mode, mode-faithful Resume for external CLIs, and a cleaner plan-usage chip.
**OMP (`omp`) run mode** (#353): Oh My Pi joins Claude Code, shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok Build and DeepSeek Harness as a run mode, in local, Docker and remote-SSH sessions: toolbar dropdown, welcome button, phone overview, command palette, clone-repo brain picker, cron agent types, tab badges and per-mode colours, plus `GET /api/omp/status`, a `codeman doctor` entry, install.sh detection and the docker agent image. The resolver leads with `~/.local/bin` (the upstream installer's real target) and demands `omp/<semver>` from `--version`, so an unrelated binary with the same three-letter name is never spawned. Past omp conversations appear in Past Sessions, read from omp's own session files (the header line carries the real working directory, so nothing has to reverse-engineer omp's directory mangling), and a respawned or resumed omp session is pinned to an exact conversation with `--resume <id>` instead of omp's newest-file `--continue`. Review hardening before merge: the pin is resolved only at the moment a respawn is actually confirmed (an eager resolve on boot recovery used to alias two omp tabs in one case directory onto one conversation), candidates are verified against their own header `cwd` and claimed process-wide so siblings cannot double-pin; `OMP_*` joins the env-override allowlist and `OMP_AUTH_BROKER_URL`/`OMP_AUTH_BROKER_TOKEN` are clamped for non-granted owners in multi-user mode, the same shape as `DEEPSEEK_BASE_URL`. Known and documented: omp's own knobs are mostly `PI_*` (it is a pi fork), its default `tools.approvalMode` is `yolo`, and in-container `--resume` pinning does not reach a Docker omp pane.
**Resume keeps the row's own CLI** (#353): clicking Resume on an OpenCode, Pi, Grok, DeepSeek or OMP row used to create a plain Claude session, since the create request never carried the row's mode. Resume now relaunches in the row's own mode with that CLI's continue flag, and retires the stale row it came from so three clicks no longer leave three copies of the same name. Codex, Gemini and Antigravity rows have no continuation wired yet, so their rows are deliberately left in place. `DELETE /api/sessions/:id` accepts a persisted-only session (ownership enforced through the same helper as live lookups, 404 rather than 403 so nothing leaks) and broadcasts `session_deleted` so other tabs drop the row too.
**Plan-usage chip drops the provider label when there is only one**: a machine with only Claude limits rendered `CLAUDE 5H 60% 7D 23%`, a 46px label naming the only thing it could be. The name exists to tell two rows apart, so it now appears only when both Claude and Codex have windows; the tooltip still names the provider either way.
### Thanks
- @timkjr for #353, and for turning every review finding around within a day
## 1.23.2
### Patch Changes
- Codex plan usage in the header chip, a visible inline rename in the session sidebar, and an installer that no longer loses Tailscale access on a re-run.
**Codex plan usage in the header chip** (#346): the plan-usage chip used to show Claude's 5-hour and weekly limits without saying they were Claude's, which stops being a detail the moment you run more than one CLI. It now renders one compact row per provider, Claude above Codex, each labelled and colour-coded by how much is used up. Claude's numbers still come from Codeman's marked `statusLine.command` exporter; Codex's come from the signed-in host CLI's read-only `account/rateLimits/read` app-server request at startup and every five minutes, so credentials stay inside the CLI and no auth material reaches the browser. Only the main `codex` bucket is read (model-specific buckets such as Spark are separate limits and are deliberately excluded), and the Codex row is omitted entirely when no 5-hour or weekly window is available, rather than inventing one.
**Inline rename is visible in the session sidebar** (#345): starting a rename on a sidebar row opened a focused input you could not see. The row's ellipsis clamp was still painting over the live editor, so text and caret went in blind. The sidebar now gets the same unclamped editor layout the vertical tab rail already had. Covered by a Chromium regression test that asserts the painted `overflow` and the input's measured width, not just the class name.
**install.sh keeps Tailscale access on a re-run**: a re-run whose build failed could drop a working Tailscale binding instead of preserving it. The installer now offers Tailscale setup again on re-run rather than losing it, and the README describes the three-way network-access prompt (Tailscale / LAN / local-only) as it actually behaves.
### Thanks
- @JackStuart for #346
- @fibr for #345
- @tailong-wu for #342, whose analysis of the terminal refresh replay loop matched a fix that had landed on master a few hours earlier
## 1.23.1
### Patch Changes
- Fix a fresh-Linux install failure, and bound the browser terminal's live write queue.
**install.sh now installs a build toolchain.** Reported against a stock Ubuntu 24 server: node-pty publishes prebuilt binaries for darwin and win32 only, so on Linux it is always compiled from source during `npm install`. The installer set up Node, tmux and git but never a compiler, so a machine without `build-essential` died deep inside node-gyp with `not found: make` — which reads like an npm bug rather than a missing system package. `make`, a C++ compiler and `python3` are now checked up front exactly like git and tmux, installed per distro (apt / dnf / pacman / apk / zypper) behind the same consent prompt, and re-verified afterwards rather than assumed. If `npm install` fails anyway — including on `install.sh update` — it now names the missing tools and the command that installs them instead of leaving a node-gyp stack trace as the last word.
**Bounded live xterm backpressure** (#339): live output is now one chunk in flight at a time, released by xterm's own parse callback, so xterm's private WriteBuffer can no longer hide an unbounded backlog behind the browser's 128 KiB render cap; queued, loading and incoming bytes all count against that cap. Automatic drop recovery for a shell stays on the bounded 1 MiB tail — a 100k-line shell capture is tens of MiB, and parsing it on the main thread is the freeze the cap exists to prevent — while TUI modes still recover full history behind the existing downgrade guard. Duplicate SSE terminal events are dropped before JSON parsing while WebSocket owns terminal I/O, and recovery is single-flight per active session. Follow-up hardening: the three write-queue reset paths now also release the in-flight gate, so a parse callback that never lands cannot leave live output permanently stalled.
**File Viewer searches the workspace** (#340): the File Viewer search box now queries the server-side file search endpoint with a 250 ms debounce and strict response validation, instead of filtering only the part of the tree already loaded. Tree and search state are scoped to the active session, the hidden-file preference and independent request epochs, so a stale response cannot repaint the panel; cached-tree restoration, directory results and reset behaviour survive session switches and both panel-hide paths.
### Thanks
- @dignfei for #339
- @aakhter for #340
- 858b15e: Search the full session workspace from File Viewer while keeping results scoped to the active session and hidden-file preference.
## 1.23.0
### Minor Changes
- DeepSeek Harness as a ninth run mode, DeepSeek agent workers, and detailed rows for the vertical tab rail.
**DeepSeek Harness (`dsh`) run mode** (#337): DeepSeek's plugin-native agent framework joins Claude Code, shell, OpenCode, Codex, Gemini, Antigravity, Pi and Grok as a run mode. The harness is a profile launcher rather than an agent, so availability is two questions (binary AND a pane-capable profile): the Run button gates on both, a missing terminal profile is offered as a one-click install (`POST /api/deepseek/install-profile`, the only endpoint in Codeman that installs third-party code, fenced accordingly), and the resolver demands the harness's own help banner so Debian's unrelated `dsh` (dancer's shell) can never be spawned. Its permission switch is the `DSH_PERMISSION_MODE` env export (the harness has no bypass flag), injected via tmux setenv and clamped for non-granted owners in multi-user mode, including the env-override path. The community TUI's supervisor-reporting contract makes deepseek the first non-Claude mode with REAL lifecycle signals: a generated status shim turns its idle/working/blocked reports into definitive `stop`/`permission_prompt`/`agent_working` hook events, so dsh sessions get real respawn triggers, real wait signals and red "needs you" alerts instead of output-stabilization guesswork. The vendor's browser UI opens as a managed web tab through a background `dsh web` fenced to Codeman's origin. Docker image support included.
**DeepSeek agent workers** (#341): the codeman agent skill can spawn and drive dsh workers like claude ones — tasked, waited on and read with the same calls. `GET /api/sessions/:id/last-response` reads the harness's real zstd transcript (one frame per append; the reader walks frame boundaries itself, since a naive decode silently truncates to the first frame), distinguishes real prompts from plugin-injected context, and reports a failed turn's provider error instead of an empty answer.
**Vertical tab rail: detailed rows** (#338): the vertical rail can now show the home screen's per-session line (created stamp, state duration, status pill) via the new per-device `tabRailDetail` setting (default detailed; `simple` restores the 1.22.0 rows). One shared row model and one gate (`isRichTabRows()`) keep the rail, the rich sidebar and both home screens in agreement about what "working" means. A never-sized rail opens at the 320px Wide preset; below 288px the created stamp is dropped, below 240px rows fall back to simple. Also fixes Escape during an inline tab rename committing an empty name (the session then displayed its folder name).
**Review hardening across all three** (post-review commits on each PR): multi-user owners without the bypass grant can no longer redirect the server's forwarded `DEEPSEEK_API_KEY` via a `DEEPSEEK_BASE_URL` override; waits on `stop`/`blocked` are refused for docker/remote dsh sessions (their status bridge cannot reach the harness) and docker/remote dsh sessions keep the pane reader (their transcripts are not local); dsh approvals are alerts answered in the terminal, never blind keystrokes into a third-party TUI; the status shim forwards the contract's `--seq` token (stale retried reports are dropped server-side) and treats 4xx as permanent so a misconfigured session cannot rate-limit the hook endpoint for the whole instance; concurrent DeepSeek web-UI starts are serialized; cron deepseek jobs run the same launch gate as the HTTP paths; the installer's dsh identity probe is stdin-closed, bounded and memoized; transcript reads are memoized per (path, mtime, size) so 1s polling stops decoding unchanged files; the rail's width dialog, compact-threshold folder rows and reset affordances are rich-aware.
### Patch Changes
- b330f1d: Vertical tab rail: detailed rows, and a rename cancel that no longer wipes the name.
The vertical rail (Tab Orientation → Vertical) now draws the same per-session
line the home screen and the rich sidebar draw — when the session was created,
how long it has been in the state it is in, the folder it runs in, and a status
pill — instead of just the name. New per-device setting **Vertical Rail Rows**
(`tabRailDetail`, App Settings → Appearance → Tabs) with `Detailed` as the
default and `Simple (name only)` as the opt-out. A rail that has never been
sized now opens at 320px (the existing Wide preset) so the line fits; a narrower
rail sheds the created stamp below 288px and falls back to simple rows below
240px.
Also fixes a data-loss bug in the inline tab rename that predates the rail:
pressing Escape cleared the input and blurred it, and the blur handler commits —
so cancelling a rename stored an EMPTY session name and the tab fell back to its
folder label. Escape now cancels without a request, in every layout.
## 1.22.0
### Minor Changes
- 3f8c8e9: Add Grok Build (xAI `grok`) as a seventh CLI run mode. SessionMode gains 'grok', with its own resolver (version-probed, since the name has npm squatters; GET /api/grok/status surfaces path + version), GrokConfig (model, alwaysApprove -> --always-approve, resume/continue), GROK*\*/XAI*\* env allowlist entries, the multi-user only-if-sent bypass clamp, Docker (own image step + per-file credential seeding) and remote-SSH command defaults, cron agentType, run-mode/welcome/tab UI with a charcoal identity, and docs (grok-integration.md + plan). Verified end to end against grok 1.0.5 on an isolated instance.
- 74194e4: Add the owner-scoped tab-layout model, persistence, API, lifecycle repair, and synchronized legacy ordering foundation.
- e3a2fb7: Add an optional resizable vertical session rail with responsive layout, complete labels, accessible controls, and stable inline rename.
### Patch Changes
- Fix the file preview's dead pop-out control: a real detach button now opens the previewed file in a browser tab (raw route for PDFs/images/media/text, converted-PDF preview for docx/pptx) and the copy button reports when a preview has no text to copy instead of silently doing nothing. Review-driven hardening for the new tab features: PUT /api/session-order drops unknown ids again instead of rejecting the whole write (a session deleted inside the browser's debounce window could silently lose the user's reorder), a failed mux restore no longer blocks explicit session/webview deletion for the process lifetime (the automated stale sweep stays fail-closed), and the vertical rail gains the axis-awareness the sidebar-only predicates missed: correct drag-reorder insertion, active-tab scroll-into-view, floating windows anchored beside rail tabs, connector redraws on rail scroll, server-seeded orientation applied on first load, a pre-paint stamp so vertical mode no longer flashes through the header strip, and a 12px session-name default matching the sidebar's historical size so untouched installs are not restyled.
## 1.21.0
### Minor Changes
- **`codeman tui`: a terminal dashboard for your sessions.** For the times you are in SSH or Termius instead of a browser. The web UI remains the primary surface and bare `codeman` still prints help, so the dashboard itself is strictly additive.
Sessions are grouped NEEDS YOU / WORKING / IDLE / RECENT in the same status language as the web tabs and the phone overview, and the states come from the server (hooks, idle confirmation, the approvals inbox) over the existing HTTP/SSE API rather than being screen-scraped. That is what lets the dashboard answer a permission dialog instead of only reporting one.
- `↑↓`/`j`/`k` select; `1`-`9`, `[`/`]` and `Tab` switch between sessions
- `Enter` attaches and hands the terminal to tmux; **`F1` comes back**, one key, no modifier. Inside the pane a bar across the top carries the session strip and `Alt+1`..`Alt+9` switch without returning to the dashboard first
- `Enter` on a RECENT row resumes that conversation; on a session whose pane has died it refuses and offers `r` to resume it in a fresh pane
- `y`/`n`/digits answer the selected session's pending permission or question card (the server re-captures the pane first, so a keystroke can never land in the composer)
- `p` sends a one-line prompt without attaching, `x` kills (`y` confirms), `n` starts a session and opens straight into it
- `/` cross-session search, `g` away digest, `?` help, live preview pane, plan-usage chip in the header, a terminal bell when a new approval arrives
- `codeman tui --list` and `codeman tui <n>` are scriptable fast paths; with no server running it lists panes straight from the instance's tmux socket, attach-only, and upgrades live when the server comes back
Narrow terminals (under 72 columns, a phone SSH client) drop the preview and get a single-column layout. `NO_COLOR`, non-UTF-8 glyph fallback and a non-TTY refusal are all handled. Zero new dependencies: hand-rolled ANSI over chalk and commander. User guide: `docs/tui.md`.
**Breaking: the `sc` tmux chooser is retired.** `scripts/tmux-chooser.sh` is deleted and `install.sh` no longer creates the `tmux-chooser` symlink or the `sc` alias; it sweeps both up instead, on update and on uninstall. `codeman tui` replaces it and does the job better: `sc` numbered its entries globally but only accepted a single `[1-9]` keypress, so sessions 10+ were listed and could not be selected, and it inferred nothing about what an agent was doing. The alias cleanup is marker-owned, matching the exact line the installer wrote, so a user's own `alias sc=` for another tool is untouched.
**CLI polish that came with it.**
- New shared style kit (`src/cli-style.ts`) used across the CLI: semantic palette, glyphs, width-aware table, spinner, confirm.
- `codeman doctor` is colorized and its table is measured, so the "Antigravity CLI" label no longer pushes its row out of column. `--json` output is unchanged.
- `codeman web` no longer prints its "running at" line twice, and the server's non-loopback security warning is painted like the CLI's (chalk degrades off a TTY, so journald and `web.log` stay free of escape codes).
- Spinners on the silent up-to-30s waits in `codeman web -d`, `codeman web --stop` and `codeman service install`.
- `codeman reset` asks a real y/N confirmation on a TTY; non-interactive callers keep the old `--force` refusal.
- `codeman list` and `codeman session list` share one renderer instead of drifting copies.
- `codeman attach` is described correctly in the README (it shows an attachment card for a local file).
- `test/cli-commands.test.ts` now derives its inventory from the real commander program instead of a hand-written fixture that had drifted.
**Internal.** New `tmux -L` callers resolve the socket through `resolveTmuxSocketName()`, now exported from `config/instance.ts`, so a second process can never point a beta instance at prod's panes. CLAUDE.md and `docs/architecture-invariants.md` both record the rule.
### Thanks
The TUI went through seven rounds of beta testing over PuTTY/SSH by **@Ark0N**, which is where the way out of an attach, the session strip, the preview repaint handling and the glyph set all came from.
## 1.20.1
### Patch Changes
- Terminal input and scrollback fixes (PRs #327, #331):
- IME punctuation preserved (#327): keyCode 229 / `Process` key events are now delegated to xterm's CompositionHelper instead of being suppressed, so an active Chinese IME committing numbers and full-width punctuation (,。!? and friends) reaches the terminal correctly. The CJK input field sends the browser's committed text instead of guessing from `KeyboardEvent.key`, and the redundant Android orphan-input fallback is removed so xterm is the single input owner.
- Shell history replay bounded (#331): selecting a Shell session loads a bounded 1 MiB tail instead of replaying the entire multi-megabyte tmux scrollback on xterm's main thread; full history stays available via the explicit "Load full history" action. tmux history limits now apply correctly on both legacy tmux (global default set in the same command queue before pane creation) and tmux 3.7+ (per-pane targeting that never resizes or trims unrelated live panes). Also adds `Server-Timing` and `[TERMINAL-PERF]` timing stages for terminal loads, fixes `scrollToLastNonEmptyLine` double-counting scrollback rows, and keeps live output ordered behind snapshot replays.
### Thanks
- @dignfei for both fixes: the IME punctuation root-cause fix (#327) and the bounded shell history replay with the tmux history-limit correctness work (#331).
## 1.20.0
### Minor Changes
- Response viewer for OpenCode, Gemini, Antigravity and Pi sessions (#326). External CLIs render their own TUIs, so the viewer used to come up empty for them; a new transcript parser (`response-viewer-transcript.ts`) reconstructs the conversation from the pane text instead, and the `?context=full` view now tags every block with a role so prompts render as "You" and agent output as the assistant. The divider normalizer was rewritten as a linear scan after review found catastrophic backtracking on agent-controlled input (minutes of stall on a long dash run), with an equivalence corpus pinning the old accept set.
CLIs installed via nvm or Homebrew are now found when Codeman runs as a service (#329). A shared resolver falls back to a login-shell probe when the direct PATH lookup misses, so systemd and LaunchAgent installs no longer report every CLI as missing. Review hardening on top: a failed resolution is negative-cached with doubling backoff instead of re-spawning a login shell on every request, all probes pass `killSignal: 'SIGKILL'` (interactive bash shrugs off SIGTERM, and a blocking `.bash_profile` could have hung the server indefinitely), the resolvers are inert under vitest again so test suites cannot execute binaries found on the dev box, and the improved not-found guidance is wired into both the session-create errors and the per-CLI status endpoints.
`GET /api/system/repo-status` reports branch, upstream, ahead/behind and remote reachability for git-clone installs (#328). Review hardening: the git network calls moved off the synchronous path onto a single-flight 45s cache (one slow remote could previously freeze the whole server for up to a minute per request), remote URLs and git stderr are credential-redacted before they leave the server, the spawns use the same non-interactive git env as the clone path, and a local-branch upstream no longer parses into garbage.
Auto Copy for the terminal (#325, opt-in, per-device): a finished selection (mouse drag, double or triple click, or a phone long-press) lands on the clipboard by itself, so select-then-copy becomes select. Alongside it, hand-encoded tap reports are now gated on the server-observed `cliMouseTracking` state, so a pane that has fallen back to a plain shell no longer receives `[<0;88;20M` junk on tap.
The Ralph loop no longer stops polling after two ticks (#330): the reschedule guard read a stale timer handle that the timer callback never cleared, so the loop silently died while its status stayed `running`. The handle is now nulled as the callback's first statement, and a regression test pins the bug.
The red "needs you" tab alert clears when a dialog is answered in the terminal instead of surviving until the end of the turn: the post-hook re-capture could erase the parsed dialog options that the staleness sweep relies on (`applyCapture` is now add-only for options), and a delayed staleness pass now runs while a page is open. The unreachable `copyTerminal()` was removed, closing out #322.
### Thanks
- @aakhter contributed the external-CLI response viewer (#326), the repo-status endpoint (#328), the login-shell CLI resolution (#329) and the Ralph reschedule fix (#330)
- @rounakdatta reported the mobile copy gap (#322) closed out in this release
## 1.19.7
### Patch Changes
+54 -34
View File
File diff suppressed because one or more lines are too long
+26 -20
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; OMP &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -27,7 +27,7 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, six CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, eight CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or OMP](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -61,14 +61,14 @@ The installer asks before every system change, and re-running the same line upda
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
This installs Node.js, tmux and a build toolchain if missing (node-pty ships no Linux prebuilds, so it compiles from source), clones Codeman to `~/.codeman/app`, and builds it. A few things worth knowing:
- **It asks first.** Every system change (package installs, AI CLI download) is prompted, and a menu at the end lets you choose: run Codeman in this terminal, install it as a background service (systemd/launchd, auto-start on boot), or don't start yet. Nothing runs in the background unless you pick it.
- **Network or local-only, your choice.** The installer asks whether the dashboard should be reachable from other devices on your network (`0.0.0.0`, the default, with a strongly recommended password prompt) or from this machine only (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. A bare `codeman web` started by hand still defaults to loopback.
- **How it's reachable, your choice.** The installer offers three ways to reach the dashboard: **Tailscale** (loopback bind fronted by `tailscale serve`, so you get `https://<machine>.<tailnet>.ts.net` with a real certificate and your tailnet as the login, no password needed), **any device on your network** (`0.0.0.0`, with a strongly recommended password prompt), or **this machine only** (`127.0.0.1`, safest). Skipping the password on a network bind requires an explicit confirmation and ends with a loud warning. The highlighted default reflects what is already on the machine (Tailscale when it is already in use, your existing binding on a re-run), and a bare Enter never pulls in new software. A bare `codeman web` started by hand still defaults to loopback.
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), or [Pi](https://pi.dev) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the six is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the nine is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -82,6 +82,8 @@ codeman users add alice --admin # create the first admin account
codeman web --multiuser # named logins + per-user case spaces
```
**Prefer Docker Compose?** A local-image Compose deployment ships in `docker/`: copy `docker/.env.example` to `docker/.env`, set `CODEMAN_PASSWORD`, then run `bash docker/Start-Codeman.sh` on Linux. Codeman runs in a container and spawns Docker cases as sibling containers through the host socket. See the [Docker deployment guide](docker/README.md) for direct Compose commands, storage and networking options.
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
<details>
@@ -171,7 +173,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), or [Pi](https://pi.dev)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev), [Grok Build](https://github.com/xai-org/grok-build), [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), or [OMP](https://github.com/can1357/oh-my-pi)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
@@ -253,7 +255,7 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. **Add Case** creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (**Clone Repo**). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, or `Terminal` (plain shell). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, `Grok`, `OMP`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Models → New Claude sessions). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
@@ -285,7 +287,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
- **SSH** — `codeman tui` is a full-screen dashboard in the terminal (`codeman tui --list` to list, `codeman tui 2` to attach straight to one).
### 7. Operate & maintain
@@ -437,7 +439,7 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, or **Pi** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md) and [`docs/pi-integration.md`](docs/pi-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, **Pi**, **Grok**, or **OMP** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*` vs `GROK_*`/`XAI_*` vs `OMP_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md), [`docs/pi-integration.md`](docs/pi-integration.md), [`docs/grok-integration.md`](docs/grok-integration.md) and [`docs/omp-integration.md`](docs/omp-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
@@ -460,7 +462,7 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / OMP logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: just Docker (or Podman). The agent base image builds itself automatically on first use, with progress streamed to the UI (or pre-build it with `node scripts/build-agent-image.mjs`). Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
@@ -658,17 +660,19 @@ These run for **every** request — before auth, even on the default no-password
---
## SSH Alternative (`sc`)
## Terminal UI (`codeman tui`)
If you prefer SSH (Termius, Blink, etc.), the `sc` command is a thumb-friendly session chooser:
A full-screen dashboard for your sessions, in the terminal. Same states as the web UI, because it is a client of the same server:
```bash
sc # Interactive chooser
sc 2 # Quick attach to session 2
sc -l # List sessions
codeman tui # the dashboard
codeman tui --list # numbered session list, then exit (scriptable)
codeman tui 2 # attach straight to session 2 of that list
```
Single-digit selection (1-9), color-coded status, token counts, auto-refresh. Detach with `Ctrl+A D`.
Sessions are grouped **NEEDS YOU → WORKING → IDLE → RECENT**, longest-waiting first. `↑↓`/`j`/`k` select, `1`-`9` and `[`/`]` switch between sessions, `Enter` attaches into the tmux pane (**`F1`** to come back). Inside a pane the bar across the top keeps the session strip visible and `Alt+1`-`Alt+9` switch without leaving. `y`/`n`/digit answer a pending permission dialog right from the list, `p` sends a one-line prompt, `n` starts a session and opens straight into it, `x` kills one (`y` confirms), `/` searches, `g` shows the away digest, `?` is help, `q` quits. Below 72 columns it drops the preview pane and becomes a single-column list, so it stays usable in Termius on a phone. With no server running it still starts in attach-only degraded mode.
The web UI remains the primary surface; see **[docs/tui.md](docs/tui.md)** for the full guide.
---
@@ -793,7 +797,7 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/omp) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
@@ -893,7 +897,9 @@ codeman session start -d /path/to/repo # (s) start a session
codeman session list # list sessions
codeman session logs <id> # tail output
codeman task add "fix the failing test" # (t) queue a task
codeman attach <path> # attach a Claude hook context
codeman attach <path> # show an attachment card for a local file
codeman tui --list # numbered session list (plain text when piped)
codeman tui 3 # attach to session 3 of that list
```
### Hooks (events flowing _back_ to Codeman)
@@ -1007,7 +1013,7 @@ flowchart TB
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / OMP</small>"] BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
+6 -20
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Grok &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -58,7 +58,7 @@ curl -fsSL https://getcodeman.com/install | bash
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这六个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli)、[Pi](https://pi.dev)、[Grok Build](https://github.com/xai-org/grok-build)、[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 或 [OMP](https://github.com/can1357/oh-my-pi)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这九个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
@@ -141,7 +141,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli)、[Pi](https://pi.dev)、[Grok Build](https://github.com/xai-org/grok-build)、[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) 或 [OMP](https://github.com/can1357/oh-my-pi))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
@@ -221,7 +221,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi` 或 `Terminal`(普通 shell)。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi`、`Grok` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
@@ -253,7 +253,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
- **手机/平板** —— UI 完全触控优化;扫描桌面上的**二维码**即可免密码登录。
- **网络之外** —— `./scripts/tunnel.sh start` 打开一条 Cloudflare 隧道(先设置 `CODEMAN_PASSWORD`)。
- **SSH** —— `sc` 选择器可从终端附着任意会话(`sc` 交互式,`sc 2` 快速附着,`sc -l` 列表)。
- **SSH** —— `codeman tui` 是终端里的全屏会话面板(`codeman tui --list` 列出,`codeman tui 2` 直接附着到某个会话)。
### 7. 运维与维护
@@ -394,7 +394,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini** 或 **Pi**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md) 与 [`docs/pi-integration.md`](docs/pi-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini**、**Pi** 或 **Grok**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*`、`GROK_*`/`XAI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)、[`docs/pi-integration.md`](docs/pi-integration.md) 与 [`docs/grok-integration.md`](docs/grok-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
@@ -615,20 +615,6 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
---
## SSH 替代方案(`sc`)
如果你更喜欢 SSH(Termius、Blink 等),`sc` 命令是一个便于拇指操作的会话选择器:
```bash
sc # 交互式选择器
sc 2 # 快速附着到会话 2
sc -l # 列出会话
```
单数字选择(1–9)、颜色编码的状态、token 计数、自动刷新。用 `Ctrl+A D` 分离。
---
## 键盘快捷键
> Ctrl 绑定在 macOS 上也接受 Cmd。
+3
View File
@@ -19,6 +19,9 @@
* why these are a runnable suite (`npm run test:browser`) rather than skipped.
*/
export const BROWSER_TEST_GLOBS = [
'test/tab-rail-resize.browser.test.ts',
'test/session-sidebar-ux.browser.test.ts',
'test/session-options-responsive.browser.test.ts',
'test/inline-rename.test.ts',
'test/opencode-resize.test.ts',
'test/webgl-fallback.test.ts',
+76
View File
@@ -0,0 +1,76 @@
# =============================================================================
# Codeman Docker Compose environment template
# Copy this file to .env and set the values for the Docker host.
# =============================================================================
TZ=Australia/Perth
# Optional overrides for direct `docker compose` use. The Bash start script
# detects these values from CODEMAN_APPDATA_PATH automatically. Compose uses
# 1000:1000 when the variables are omitted.
# PUID=1000
# PGID=1000
# Name of the account that runs Codeman and all local CLI sessions. Changing
# this value rebuilds the image with a matching account.
CODEMAN_RUNTIME_USER=opencode
# Required. Persistent Codeman application data, CLI credentials, and session
# state are stored here on the host and mounted at the runtime account's home
# directory in the container.
CODEMAN_APPDATA_PATH=/mnt/user/appdata/Coding/codeman
# Optional. Absolute host path of this Codeman checkout, mounted at
# /opt/codeman so App Settings -> Updates can update Codeman in place. The Bash
# start script detects it from the compose file's own location, so it only needs
# setting for direct `docker compose` use or a checkout kept elsewhere. Point it
# at a directory that is not a git checkout and in-app updates are unavailable.
# CODEMAN_REPO_PATH=/mnt/user/appdata/Coding/codeman/app
# Required for Docker cases. This must be an absolute path on the Docker host.
# Codeman and each isolated case use this same path, so it cannot be a
# container-only path such as /home/opencode/codeman-cases.
CODEMAN_CASES_PATH=/mnt/user/appdata/Coding/codeman/codeman-cases
# Required. Network bind address, host port, and local image tag.
CODEMAN_HOST=0.0.0.0
CODEMAN_PORT=3000
CODEMAN_IMAGE=codeman:local
# Required for any network-accessible Codeman instance. Use a unique, strong
# password. This file is safe to commit; copy it to .env and set the value.
CODEMAN_PASSWORD=changeme
# Required. Username for Codeman HTTP Basic authentication.
CODEMAN_USERNAME=admin
# Optional: authenticate Gemini CLI without an interactive login.
GEMINI_API_KEY=
# Linux default. On Docker Desktop, use the socket path supported by your
# Docker installation when it differs from /var/run/docker.sock.
DOCKER_SOCKET=/var/run/docker.sock
# Optional override for direct `docker compose` use. The Bash start script
# detects this from DOCKER_SOCKET automatically. The direct Compose default is
# 999, but the correct value depends on the Docker host.
# DOCKER_SOCKET_GID=999
# Set to 1 only when Docker-case hook callbacks are required.
CODEMAN_DOCKER_BRIDGE_HOOKS=0
# Set to 1 when `docker info` reports `SwapLimit=false`. The case memory limit
# remains active; Codeman omits --memory-swap and filters the daemon's exact
# unsupported-swap warning while preserving all other Docker create errors.
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=0
# Required only when applying the macvlan example in README.md.
CODEMAN_MACVLAN_NETWORK=br0.11
CODEMAN_IPV4_ADDRESS=10.10.11.236
CODEMAN_MAC_ADDRESS=02:10:11:00:00:EC
# Required only when creating a new managed macvlan network, rather than using
# the external-network macvlan example.
CODEMAN_MACVLAN_PARENT=br0.11
CODEMAN_MACVLAN_SUBNET=10.10.11.0/24
CODEMAN_MACVLAN_GATEWAY=10.10.11.1
+108
View File
@@ -0,0 +1,108 @@
# Codeman Docker deployment
This folder contains the Compose configuration, server image Dockerfile, and environment template for a locally built Codeman server.
## Start
From the repository root, create the runtime environment file and set the required values, especially `CODEMAN_PASSWORD`.
```sh
cp docker/.env.example docker/.env
bash docker/Start-Codeman.sh
```
On PowerShell, use the following command instead.
```powershell
Copy-Item docker/.env.example docker/.env
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
Every required value is defined and explained in `.env.example`. `GEMINI_API_KEY` is intentionally optional and may remain blank.
On Linux, `Start-Codeman.sh` stops with an error when required paths are missing. It creates the application-data directory when safe, detects its numeric owner as `PUID:PGID`, and detects `DOCKER_SOCKET_GID` from the configured Docker socket. It rejects a root-owned application-data directory because Codeman and its local CLI sessions must remain unprivileged.
Codeman, Claude, OpenCode, and other local sessions run as the unprivileged account named by `CODEMAN_RUNTIME_USER`, which defaults to `opencode`. When Compose is run directly, `PUID` and `PGID` default to `1000:1000`; set them in `.env` when the application-data directory has a different owner. The Bash start script determines them automatically instead.
To retain Docker-case support without root when running Compose directly, set `DOCKER_SOCKET_GID` to the numeric group ID of the host socket. On a standard Linux Docker host, obtain it with `stat -c '%g' /var/run/docker.sock`. The Bash start script detects it automatically.
## Updating
Use **App Settings → Updates** in the web UI. The checkout Compose builds from is
also mounted at `/opt/codeman`, so an update's `git checkout` and rebuild persist
on the host, and the server exiting is what restarts the container onto the new
build.
Releases that change `server.Dockerfile`, `docker-compose.yaml`, or add a key to
`.env.example` cannot be applied that way — the updater detects them, names what
changed, and asks you to run `Start-Codeman.sh` here on the host instead. Details:
[`../docs/docker-self-update.md`](../docs/docker-self-update.md).
## Application data storage
The default configuration uses a host-folder bind mount:
```yaml
volumes:
- type: bind
source: ${CODEMAN_APPDATA_PATH}
target: /home/${CODEMAN_RUNTIME_USER}
```
Set `CODEMAN_APPDATA_PATH` in `.env` to a directory that the Docker daemon can access. The example value is `/mnt/user/appdata/Coding/codeman`.
`CODEMAN_CASES_PATH` is the separate host directory for managed case workspaces. It is mounted into Codeman at the same absolute path, allowing the host Docker daemon to bind it into an isolated case container. Set it to a child directory of `CODEMAN_APPDATA_PATH` unless you deliberately store workspaces elsewhere.
Compose also exposes `CODEMAN_APPDATA_PATH` to Codeman as `CODEMAN_DOCKER_HOST_HOME`. This lets Docker case seed files, CLI credentials and the hook secret be mounted using paths that exist in the host daemon's filesystem. Direct host installations do not set this variable and retain their existing behaviour.
Set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1` when `docker info` reports `SwapLimit=false`. Codeman continues to apply the configured case memory limit, omits Docker's unsupported `--memory-swap` option, and filters only the daemon's exact swap-capability warning. Every other Docker create error and its exit status remain visible.
For an existing installation created by a root-running image, change ownership of the application-data directory before upgrading so the configured `PUID` and `PGID` can read the saved credentials and state:
```sh
chown -R 99:100 /mnt/user/appdata/Coding/codeman
```
Replace `99:100` and the path with the values from your `.env` file.
Do not replace this bind mount with a Docker-managed named volume when Docker cases are enabled. Codeman passes seed, credential, transcript and hook-secret bind sources to the host Docker daemon, so their source files must have stable paths in the daemon's filesystem. A named volume does not provide the required host path mapping.
## Static macvlan networking
The default configuration publishes a host port. It does not use `network_mode: host`. To attach Codeman directly to an existing external macvlan network with a static IP address and MAC address, remove the `ports:` section and add the following to the `codeman` service:
```yaml
mac_address: ${CODEMAN_MAC_ADDRESS}
networks:
codeman_lan:
ipv4_address: ${CODEMAN_IPV4_ADDRESS}
```
Then add this top-level network declaration:
```yaml
networks:
codeman_lan:
external: true
name: ${CODEMAN_MACVLAN_NETWORK}
```
Set `CODEMAN_MACVLAN_NETWORK`, `CODEMAN_IPV4_ADDRESS`, and `CODEMAN_MAC_ADDRESS` in `.env`. The values in `.env.example` match the supplied Unraid example network and should be changed for other hosts.
### Create a managed macvlan network
If an external macvlan network does not already exist, use this top-level declaration instead. Do not use it together with the external-network declaration.
```yaml
networks:
codeman_lan:
driver: macvlan
driver_opts:
parent: ${CODEMAN_MACVLAN_PARENT}
ipam:
config:
- subnet: ${CODEMAN_MACVLAN_SUBNET}
gateway: ${CODEMAN_MACVLAN_GATEWAY}
```
Macvlan containers are ordinarily not reachable from their Docker host without additional host-network routing. Confirm the selected address, MAC address, parent interface, and subnet are reserved and valid for the target network before starting the stack.
+129
View File
@@ -0,0 +1,129 @@
#!/usr/bin/env bash
set -euo pipefail
script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
env_file="$script_dir/.env"
compose_file="$script_dir/docker-compose.yaml"
if [[ ! -f "$env_file" ]]; then
printf 'Error: Docker environment file is missing: %s\n' "$env_file" >&2
printf 'Create it from %s/.env.example before starting Codeman.\n' "$script_dir" >&2
exit 1
fi
compose_command=(docker compose --env-file "$env_file" -f "$compose_file")
appdata_path=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "CODEMAN_APPDATA_PATH" { sub(/^[^=]*=/, ""); print; exit }'
)
docker_socket=$(
"${compose_command[@]}" config --environment |
awk -F= '$1 == "DOCKER_SOCKET" { sub(/^[^=]*=/, ""); print; exit }'
)
if [[ -z "$appdata_path" ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is not set in %s\n' "$env_file" >&2
exit 1
fi
if [[ ! -d "$appdata_path" ]]; then
if [[ "$EUID" == '0' ]]; then
printf 'Error: Refusing to create CODEMAN_APPDATA_PATH as root: %s\n' "$appdata_path" >&2
printf 'Create it as the unprivileged account that should run Codeman, then retry.\n' >&2
exit 1
fi
mkdir -p -- "$appdata_path"
fi
if owner_ids=$(stat -c '%u:%g' -- "$appdata_path" 2>/dev/null); then
:
elif owner_ids=$(stat -f '%u:%g' "$appdata_path" 2>/dev/null); then
:
else
printf 'Error: Cannot determine the owner of CODEMAN_APPDATA_PATH: %s\n' "$appdata_path" >&2
exit 1
fi
export PUID=${owner_ids%%:*}
export PGID=${owner_ids##*:}
if [[ "$PUID" == '0' ]]; then
printf 'Error: CODEMAN_APPDATA_PATH is owned by root: %s\n' "$appdata_path" >&2
printf 'Change the directory ownership to the unprivileged account that should run Codeman.\n' >&2
exit 1
fi
if [[ -z "$docker_socket" || ! -S "$docker_socket" ]]; then
printf 'Error: DOCKER_SOCKET is not a Unix socket: %s\n' "${docker_socket:-<unset>}" >&2
exit 1
fi
if socket_ids=$(stat -c '%u:%g' -- "$docker_socket" 2>/dev/null); then
:
elif socket_ids=$(stat -f '%u:%g' "$docker_socket" 2>/dev/null); then
:
else
printf 'Error: Cannot determine the owner of DOCKER_SOCKET: %s\n' "$docker_socket" >&2
exit 1
fi
export DOCKER_SOCKET_GID=${socket_ids##*:}
repo_path=${CODEMAN_REPO_PATH:-$(cd -- "$script_dir/.." && pwd)}
if [[ ! -d "$repo_path" ]]; then
printf 'Error: CODEMAN_REPO_PATH is not a directory: %s\n' "$repo_path" >&2
exit 1
fi
export CODEMAN_REPO_PATH="$repo_path"
# The in-app updater runs `git checkout` and `npm install` against this checkout
# as PUID:PGID. If the directory belongs to someone else, git refuses outright
# ("detected dubious ownership") and the update fails at the first step — so warn
# here, where the fix is obvious, rather than in a failed update hours later.
if repo_owner=$(stat -c '%u' -- "$repo_path" 2>/dev/null || stat -f '%u' "$repo_path" 2>/dev/null); then
if [[ "$repo_owner" != "$PUID" ]]; then
printf 'Warning: %s is owned by UID %s but Codeman runs as UID %s.\n' "$repo_path" "$repo_owner" "$PUID" >&2
printf 'In-app updates will fail until the ownership matches. Codeman itself still starts.\n' >&2
fi
fi
if [[ ! -d "$repo_path/.git" ]]; then
printf 'Note: %s is not a git checkout, so in-app updates are unavailable.\n' "$repo_path" >&2
fi
# Record what the container is about to be built and created FROM. The in-app
# updater compares these against the release it wants to apply: a release that
# changes either file cannot be applied by the container restarting itself (a
# restart reuses the existing image and config), so it is refused and the user
# is sent back here. Written on every start, so the baseline always describes
# the container that is actually running. See docs/docker-self-update.md.
if command -v sha256sum >/dev/null 2>&1; then
sha256_of() { sha256sum -- "$1" | cut -d' ' -f1; }
elif command -v shasum >/dev/null 2>&1; then
sha256_of() { shasum -a 256 -- "$1" | cut -d' ' -f1; }
else
sha256_of() { printf ''; }
fi
dockerfile_sha=$(sha256_of "$script_dir/server.Dockerfile")
compose_sha=$(sha256_of "$compose_file")
if [[ -n "$dockerfile_sha" && -n "$compose_sha" ]]; then
# $CODEMAN_APPDATA_PATH is mounted at the runtime account's home, so this is
# dataPath('docker-env-applied.json') as the server inside the container sees it.
state_dir="$appdata_path/.codeman"
mkdir -p -- "$state_dir"
printf '{\n "dockerfileSha256": "%s",\n "composeSha256": "%s"\n}\n' \
"$dockerfile_sha" "$compose_sha" >"$state_dir/docker-env-applied.json.tmp"
mv -- "$state_dir/docker-env-applied.json.tmp" "$state_dir/docker-env-applied.json"
# A root-run start (common on Unraid) would otherwise leave a root-owned
# `.codeman` on a FIRST start, before the container has created it as PUID,
# and the unprivileged server could then never write its own state there.
if [[ "$EUID" == '0' ]]; then
chown -- "$PUID:$PGID" "$state_dir" "$state_dir/docker-env-applied.json"
fi
else
printf 'Warning: no sha256 tool found; in-app updates will not detect environment changes.\n' >&2
fi
exec docker compose --env-file "$env_file" -f "$compose_file" up --build -d
+80 -2
View File
@@ -51,6 +51,58 @@ RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
# Grok Build (`grok`, xAI) is NOT on npm: a standalone ~160MB Rust binary through
# xAI's installer, which targets $HOME/.grok/bin with no --dir override. At build
# time that is root's home and unreachable by the `agent` user, so copy the binary
# into /usr/local/bin and drop root's ~/.grok in the same layer so the image does
# not carry the download twice. The staging cp -T is what makes this survive the
# installer's own behavior EITHER way: newer installers already symlink
# /usr/local/bin/grok -> /root/.grok/bin/grok, and a direct `cp -L` onto that
# symlink fails with "same file" (2026-08-24 rebuild), while removing the link
# first and copying fresh works for both old and new installers.
RUN curl -fsSL https://x.ai/cli/install.sh | bash \
&& cp -L /root/.grok/bin/grok /usr/local/bin/grok.real \
&& rm -f /usr/local/bin/grok \
&& mv /usr/local/bin/grok.real /usr/local/bin/grok \
&& chmod 755 /usr/local/bin/grok \
&& rm -rf /root/.grok /root/.local/bin/grok /root/.local/bin/agent \
&& grok --version
# DeepSeek Harness (`dsh`). A normal npm package, but the ONLY entry here whose
# binary runs nothing on its own: `dsh` is a profile launcher, and DeepSeek ships
# only `web` and `headless`, so without an interactive profile a
# `mode: 'deepseek'` container would start a pane that dies on arrival. The
# profile itself is installed further down, into the `agent` HOME, because
# Codeman deliberately does NOT seed `profiles/` from the host: it is a
# per-profile node_modules tree, host-arch-specific and far too large to copy on
# every container start.
# ⚠️ `pnpm` is a HARD dependency of `dsh plugin`, not optional tooling: the
# subcommand is a thin forwarder that `spawnSync`s a literal `pnpm` with no
# fallback to npm, so on an image without it the profile install below dies
# with `dsh: pnpm not found on PATH` / exit 127 and takes the whole build with
# it (issue #352). It stays on PATH at runtime too, so a container user can run
# `dsh plugin add` themselves.
RUN npm install -g @deepseek-ai/dsh pnpm \
&& npm cache clean --force \
&& dsh --version \
&& pnpm --version
# OMP (Oh My Pi) is NOT on npm: a standalone binary via omp.sh's installer, which
# targets $HOME/.local/bin with no --dir override (verified 2026-08-27 — the
# resolver's OMP_SEARCH_DIRS lists ~/.omp/bin first, which turned out to be the
# WRONG guess for the installer's actual target; build this step for real
# rather than trust that ordering). At build time $HOME is root's home and
# unreachable by the `agent` user, so copy the binary into /usr/local/bin and
# drop root's ~/.local/bin/omp in the same layer so the image does not carry
# the download twice.
RUN curl -fsSL https://omp.sh/install | sh \
&& cp -L /root/.local/bin/omp /usr/local/bin/omp.real \
&& rm -f /usr/local/bin/omp \
&& mv /usr/local/bin/omp.real /usr/local/bin/omp \
&& chmod 755 /usr/local/bin/omp \
&& rm -f /root/.local/bin/omp \
&& omp --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
@@ -68,11 +120,37 @@ ENV HOME=/home/agent
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
# `.pi/agent` IS pre-created: pi is seeded per-FILE (auth/settings/trust/models), and a
# `.pi/agent` and `.grok` ARE pre-created: both are seeded per-FILE (pi:
# auth/settings/trust/models; grok: auth.json/config.toml/pager.toml), and a
# per-file seed copy, unlike a whole-dir one, does not create its parent directory.
# `.dsh` is pre-created for the same per-file reason (.env/settings.yaml/
# cordis.patch.yml), and the interactive profile is built into it HERE rather than
# after `USER agent`: this layer's closing chgrp/chmod is what makes the whole tree
# writable by the arbitrary uid the container actually runs as, and a profile
# installed after it would miss that fixup. DSH_HOME points the launcher at the
# agent's dir while this still runs as root.
# ⚠️ `dangerouslyAllowAllBuilds` is what keeps that profile install from becoming
# the next #352. pnpm (unlike npm) blocks dependency lifecycle scripts by default
# and FAILS the install over it — `ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on
# pnpm 11.24 — so any package in the tui's tree that ships one stops the build
# dead. An allowlist of the offenders rots: `@deepseek-harness-tui/dsh-tui` is
# resolved by dist-tag, not pinned, and 0.9.3 pulled `@google/genai` (a
# `preinstall: no-op`) where 0.10.0-beta.x does not, so the names to allow move
# under us between rebuilds. Allowing them wholesale is also the SAME exposure
# this image already accepts three layers up: `npm install -g` runs the install
# scripts of every transitive dep of the five CLIs above it, with no gate at all.
# `.omp/agent` is pre-created for the same reason `.codex` is: it is a MIXED
# store (per-file config seeds PLUS a shared `sessions/` RW bind mount for
# Codeman's own host-side history/resume reads), and neither kind of artifact
# creates its own parent directory.
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent /home/agent/.grok \
/home/agent/.dsh /home/agent/.omp/agent \
&& DSH_HOME=/home/agent/.dsh HOME=/home/agent \
dsh plugin --profile dsh-tui add --config.dangerouslyAllowAllBuilds=true \
@deepseek-harness-tui/dsh-tui \
&& test -f /home/agent/.dsh/profiles/dsh-tui/package.json \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
+111
View File
@@ -0,0 +1,111 @@
name: codeman
services:
codeman:
build:
context: ..
dockerfile: docker/server.Dockerfile
args:
CODEMAN_RUNTIME_USER: ${CODEMAN_RUNTIME_USER}
PGID: ${PGID:-1000}
PUID: ${PUID:-1000}
image: ${CODEMAN_IMAGE}
init: true
restart: unless-stopped
ports:
- "${CODEMAN_PORT}:${CODEMAN_PORT}"
environment:
# Tells the self-updater to restart by exiting (the restart policy below
# relaunches it) rather than by looking for an init system that is not
# here. Also set in the image; repeated so a container started without the
# image default still self-identifies.
CODEMAN_IN_CONTAINER: "1"
# This file sets `restart: unless-stopped` below, so the updater may restart
# the server by EXITING. Declared here and only here, never in the image: a
# container started by plain `docker run` has no restart policy unless the
# operator gave it one, and there the updater asks the daemon instead and
# stages the update for a manual restart when it cannot get an answer.
CODEMAN_RESTART_BY_EXIT: "1"
CODEMAN_DOCKER_BRIDGE_HOOKS: ${CODEMAN_DOCKER_BRIDGE_HOOKS}
# Host-side equivalent of the runtime user's HOME. Docker case seed,
# credential and hook mounts are translated into the daemon namespace.
CODEMAN_DOCKER_HOST_HOME: ${CODEMAN_APPDATA_PATH}
CODEMAN_DOCKER_DISABLE_SWAP_LIMIT: ${CODEMAN_DOCKER_DISABLE_SWAP_LIMIT}
CODEMAN_CASES_PATH: ${CODEMAN_CASES_PATH}
CODEMAN_HOST: ${CODEMAN_HOST}
CODEMAN_PASSWORD: ${CODEMAN_PASSWORD}
CODEMAN_PORT: ${CODEMAN_PORT}
CODEMAN_USERNAME: ${CODEMAN_USERNAME}
GEMINI_API_KEY: ${GEMINI_API_KEY}
PGID: ${PGID:-1000}
PUID: ${PUID:-1000}
TZ: ${TZ}
group_add:
# Retain access to the host Docker socket without running as root.
- ${DOCKER_SOCKET_GID:-999}
volumes:
# Application data and CLI credentials persist on the configured host
# path, rather than in a Docker-managed volume.
- type: bind
source: ${CODEMAN_APPDATA_PATH}
target: /home/${CODEMAN_RUNTIME_USER}
# Docker cases are sibling containers on the host daemon. Their workspace
# must be visible to Codeman at the same absolute path used by that daemon.
- type: bind
source: ${CODEMAN_CASES_PATH}
target: ${CODEMAN_CASES_PATH}
# Codeman uses the host daemon to create isolated Docker cases. This is
# Docker-outside-of-Docker, not Docker-in-Docker.
- type: bind
source: ${DOCKER_SOCKET}
target: /var/run/docker.sock
# The application source, so App Settings -> Updates can update in place.
# This is the SAME checkout used as the build context above, mounted over
# the image's baked copy: a `git checkout` performed inside the container
# then lands on the host and survives the container being recreated.
# Without it the pull would go to the container's writable layer and be
# silently discarded by the next `up`. See docs/docker-self-update.md.
# Defaults to `..` — the build context above — which Compose resolves
# against the project directory, so plain `docker compose up` works with
# no extra configuration. Set CODEMAN_REPO_PATH only to point elsewhere.
- type: bind
source: ${CODEMAN_REPO_PATH:-..}
target: /opt/codeman
# Build artefacts live in named volumes layered OVER the repo bind mount,
# so `npm install` and `npm run build` inside the container never write
# into the host checkout. That keeps container-compiled native modules
# (node-pty is built from source here) out of a checkout that may also be
# used to run Codeman natively, and keeps `git status` clean. Docker seeds
# an EMPTY named volume from the image, so the first start inherits the
# image's already-built node_modules and dist rather than paying for a
# bootstrap build.
- type: volume
source: codeman-node-modules
target: /opt/codeman/node_modules
- type: volume
source: codeman-dist
target: /opt/codeman/dist
extra_hosts:
- "host.docker.internal:host-gateway"
security_opt:
- no-new-privileges:true
cap_drop:
- ALL
healthcheck:
test:
- CMD-SHELL
- >-
node -e "fetch('http://127.0.0.1:${CODEMAN_PORT}/api/status').then((response) => process.exit(response.status < 500 ? 0 : 1)).catch(() => process.exit(1))"
interval: 30s
timeout: 5s
retries: 3
start_period: 30s
volumes:
# Container-owned build artefacts. They persist across container recreation,
# so an in-app update's `npm install` output is not thrown away by the next
# `up`, and they are seeded from the image on first use. Removing them (or
# `docker compose down -v`) is the supported reset: the next start rebuilds
# from the image.
codeman-node-modules:
codeman-dist:
+142
View File
@@ -0,0 +1,142 @@
# syntax=docker/dockerfile:1
# Build the application from the checkout supplied as the Docker build context.
# No published Codeman application image is required.
FROM node:22-bookworm-slim AS build
RUN apt-get update \
&& apt-get install -y --no-install-recommends python3 make g++ \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /opt/codeman
COPY . .
# devDependencies are deliberately KEPT (no `npm prune --omit=dev`). The in-app
# updater rebuilds from inside this container, and `npm run build` is tsc +
# esbuild — both devDependencies. Pruning them saves image size and takes the
# self-updater with it. See docs/docker-self-update.md.
RUN npm ci \
&& npm run build \
&& npm cache clean --force
# The Docker CLI talks to the host daemon through the socket mounted by
# docker/docker-compose.yaml. It does not run a Docker daemon in this container.
FROM node:22-bookworm-slim
ARG CODEMAN_RUNTIME_USER=opencode
ARG PUID=1000
ARG PGID=1000
# python3/make/g++ are here for the SELF-UPDATER, not for this build. An update
# runs `npm install` inside the running container, and node-pty ships no Linux
# prebuild, so a release that bumps it compiles from source right here. Without
# a toolchain that install fails and the update rolls back — every time, on the
# releases that need it most. Same reason install.sh installs one on bare hosts.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
ca-certificates \
curl \
g++ \
git \
make \
openssh-client \
procps \
python3 \
ripgrep \
tmux \
&& rm -rf /var/lib/apt/lists/*
# The Docker CLI, taken from the official image rather than Debian's `docker.io`.
# That package is the full ENGINE: with --no-install-recommends it still pulls 15
# packages including containerd, runc, dmsetup and iptables, none of which a
# client that only talks to a mounted socket can use. Measured on top of this
# base image: `docker.io` costs 266 MB and ships Docker 20.10.24 (2023), while
# these two files cost 108 MB and ship the current CLI (493 MB vs 335 MB total).
#
# The binaries are STATIC Go builds, so they run on this glibc image even though
# the image they come from is Alpine (verified: `docker --version`, `docker ps`
# and `docker build` all work here against a mounted host socket).
#
# buildx is copied on purpose. `scripts/build-agent-image.mjs` shells out to
# `docker build` — Codeman auto-builds the agent image on the first Docker case —
# and without the plugin that silently falls back to the CLASSIC builder, which
# Docker has deprecated and will eventually drop. `docker-compose` is NOT copied:
# Codeman never shells out to it.
COPY --from=docker:29-cli /usr/local/bin/docker /usr/local/bin/docker
COPY --from=docker:29-cli \
/usr/local/libexec/docker/cli-plugins/docker-buildx \
/usr/local/libexec/docker/cli-plugins/docker-buildx
# Keep credentials out of the image. Users authenticate these CLIs at runtime
# through Codeman sessions, and the configured host bind mount retains state.
#
# ⚠️ PINNED ON PURPOSE. Unpinned, the agent CLI versions a user ends up with are
# a function of WHEN their image was built, not of any commit — so a Codeman
# release that depends on newer CLI behaviour (the trust-dialog handling is
# pinned to Claude Code 2.1.252's layout; wheel forwarding to >= 2.1.187) breaks
# on an older image with no diff anywhere to explain why. In-app updates make
# rebuilds RARER, which makes that drift worse. Pinning turns "this release needs
# a newer CLI" into a Dockerfile change, which the updater's environment gate
# already detects and refuses (docs/docker-self-update.md).
#
# Bump these deliberately, in a release. `--no-cache` is still needed to rebuild
# this layer when only the pins change upstream.
RUN npm install --global \
@anthropic-ai/claude-code@2.1.258 \
@google/gemini-cli@0.58.0 \
@openai/codex@0.152.1 \
opencode-ai@1.18.26 \
&& npm cache clean --force
# Keep the web server and every local Codeman session unprivileged. PUID and
# PGID match the host-owned application-data directory mounted by Compose. The
# requested GID may not exist in the base image, and a host UID such as 1000 may
# already belong to the baked `node` account, so handle both cases explicitly.
RUN set -eux; \
case "${PUID}" in ''|*[!0-9]*) echo "PUID must be numeric" >&2; exit 1;; esac; \
case "${PGID}" in ''|*[!0-9]*) echo "PGID must be numeric" >&2; exit 1;; esac; \
if [ "${PUID}" -eq 0 ]; then \
echo "PUID must identify an unprivileged account, not root" >&2; \
exit 1; \
fi; \
if ! getent group "${PGID}" >/dev/null; then \
groupadd --gid "${PGID}" codeman-runtime; \
fi; \
existing_user="$(getent passwd "${PUID}" | cut -d: -f1 || true)"; \
if [ -n "${existing_user}" ]; then \
usermod \
--login "${CODEMAN_RUNTIME_USER}" \
--gid "${PGID}" \
--home "/home/${CODEMAN_RUNTIME_USER}" \
--move-home \
--shell /bin/bash \
"${existing_user}"; \
else \
useradd \
--uid "${PUID}" \
--gid "${PGID}" \
--create-home \
--home-dir "/home/${CODEMAN_RUNTIME_USER}" \
--shell /bin/bash \
"${CODEMAN_RUNTIME_USER}"; \
fi
WORKDIR /opt/codeman
COPY --from=build /opt/codeman /opt/codeman
# CODEMAN_IN_CONTAINER tells the self-updater it must restart by exiting rather
# than by asking an init system that is not here (src/web/self-update.ts).
# NODE_ENV stays `production`; the updater passes `npm install --include=dev`
# explicitly, since that value would otherwise omit the build toolchain.
ENV CODEMAN_IN_CONTAINER=1 \
CODEMAN_PORT=3000 \
HOME=/home/${CODEMAN_RUNTIME_USER} \
NODE_ENV=production
EXPOSE 3000
USER ${CODEMAN_RUNTIME_USER}
CMD ["node", "dist/index.js", "web"]
+10 -5
View File
@@ -204,11 +204,16 @@ turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
dialog only as the bounded fallback.
⚠️ **The fallback is not a bare `\r`.** Claude Code 2.1.252 unnumbered the dialog's
options, reversed them and highlights `No, exit`, so an Enter sent blind quits the
CLI and the pane dies seconds after the spawn. Read the `❯` marker off the current
frame (`GET /api/v1/sessions/:id/terminal?full=1`), send `ESC [ B` while it is on
`No, exit`, re-read, and confirm only once it is on `Yes, I trust this folder`.
Reading the current frame is also what keeps this correct on later runs: the dialog
text stays in the terminal buffer for the life of the session, so a `trust` probe
with `from=buffer` keeps matching long after the dialog is gone. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
File diff suppressed because one or more lines are too long
+130
View File
@@ -0,0 +1,130 @@
# The CLI registry
Every run mode Codeman can launch — Claude Code, Terminal/Shell, OpenCode, Codex, Gemini, Antigravity, Pi, Grok, DeepSeek Harness and OMP — is a `CliEntry`: a data record describing how to find the binary, how to build its command line, what environment it needs, and what it can do. Code that used to ask "which CLI is this?" asks the entry instead.
## Where it lives
| File | What it holds |
| ------------- | ------------------------------------------------------------------------------------------------- |
| `types.ts` | The `CliEntry` interface and everything under it. Read this first. |
| `stock.ts` | The shipped catalog. **The only file allowed to name a CLI id.** |
| `schema.ts` | Zod validation, including the cross-field checks that reject an incoherent entry at LOAD time. |
| `argv.ts` | The argv engine: the only code that turns typed tokens into a command string. |
| `patterns.ts` | The NAMED value patterns (`model`, `uuid`, `path-segment`, …) and the regex-compilation guard. |
| `profiles.ts` | The names of behaviours that genuinely need code, kept import-free so `schema.ts` can validate one. |
| `registry.ts` | Loading, merging `~/.codeman/clis.json`, and the accessors (`getCli`, `enabledClis`). |
`src/session-cli-registry-bridge.ts` maps the legacy per-mode option bag onto the engine, and `src/utils/cli-resolver.ts` / `src/utils/cli-launcher.ts` do registry-driven binary resolution and launcher-profile dispatch.
## The override file
`~/.codeman/clis.json` (instance-scoped through `dataPath()`) holds overrides and custom entries only, never a copy of the stock catalog: `{ "clis": { "<id>": { ...partial entry... } } }`. Objects merge key-wise onto the stock entry, arrays replace wholesale. **The file must be mode 0600**; the loader refuses any group/world permission bit, read bits included, so a file created with a normal umask (0644) is ignored until you `chmod 600` it. Every reason a file was ignored or an entry dropped is logged once, prefixed `[cli-registry]`, on the first load. A stock entry whose override fails validation falls back to the shipped definition; a custom entry that fails is dropped. The file is read once per process and re-read only on restart.
## The shape of an entry
```ts
interface CliEntry {
id: CliId; // 'codex'
label: string; // 'Codex' — shown in menus
shortBadge: string; // tab badge, e.g. 'CX'
accent: string; // single hex colour
enabled: boolean;
stock: boolean; // set by the loader; a custom entry can never claim it
order: number;
kind: 'agent' | 'shell';
discovery: CliDiscovery; // how to find and prove the binary
launch: CliLaunch; // the structured argv template
env: CliEnv; // exports, tmux setenv keys, the env-override allowlist
capabilities: CliCapabilities; // what every call site reads instead of the id
overlays: CliOverlays; // remote-SSH / Docker pane commands, credential store
}
```
`capabilities` is the important part. It is what `isExternalCliMode()`, `isAltScreenStripMode()`, `hooksAvailableForMode()` and every other former per-mode branch actually read.
### Three capabilities that must stay independent
`external`, `hooks` and `altScreen` describe three different, deliberately unequal sets, and deriving any one from another has already shipped a bug. `shell` has no hooks but is **not** an external CLI, so a hooks predicate written as `!isExternalCliMode()` accepted `until=stop` on a shell session and then blocked the caller for their entire timeout. `deepseek` is the mirror image: it IS external and it DOES have hooks.
`test/cli-capability-predicates.test.ts` asserts that no two of the three are equivalent across the catalog, so collapsing them fails the build rather than a user's session.
## Arg-template safety
The composed command line is interpolated into `bash -c "…"` inside tmux, which makes command construction a security boundary. Four independent layers keep config out of it:
1. **Config contains no shell text.** There is no `command: "..."` field anywhere in the schema. An entry declares a sequence of typed tokens; `argv.ts` is the only place that turns them into a string, and it owns every separator itself — one space between tokens, ` || ` between fallback variants. Neither can originate from config, because config has no field that could hold either.
2. **Every literal is validated at LOAD time** against a safe-word pattern (no space, quote, backtick, `$`, `;`, `&`, `|`, redirection, parens, braces, newline or backslash). A bad literal **rejects the whole entry** rather than being dropped, because a silently dropped flag would change security-relevant behaviour — losing `--no-approve` is not a cosmetic difference.
3. **Values resolve through NAMED patterns.** A value placeholder selects a `TokenPattern` (`model`, `uuid`, `slug`, `path-segment`, `tool-list`, …) from `patterns.ts`; config can never supply its own regex for a value, so a `clis.json` structurally cannot widen its own validation. A value that fails its pattern drops the whole argument, exactly as the hand-written builders did: an invalid `--model` omits `--model`, it never substitutes something else.
4. **Escaping is independent of validation.** `renderToken()` re-checks the resolved value before emitting it unquoted, and single-quotes anything else — so even a value that somehow bypassed validation is quoted, never concatenated raw.
The only config-supplied regexes are `discovery.version.regex` and `discovery.identity.regex`. Both run against **command output** rather than a shell token, both are compiled through `compileVersionRegex()` (length cap, nested-quantifier rejection, never the `g` flag), and the output they see is truncated first.
## Named profiles: the escape hatch
Some differences genuinely need to run code rather than be described. Those are **named profiles**: a capability field holds a profile NAME, and the implementation lives in one place keyed by that name — never by CLI id.
- `discovery.launcherProfile` — for a CLI whose binary is not the agent. `dsh` boots `$DSH_HOME/profiles/<name>`, so "installed" and "runnable" have different answers; the profile answers both, plus why a specifically-named target will not work. Implemented in `utils/cli-launcher.ts`.
- `env.setenvProfile` — per-CLI environment setup that is more than a list of keys, such as DeepSeek's status bridge.
- `capabilities.transcript` — which on-disk history reader understands this CLI (`claude-jsonl`, `codex-rollout`, `deepseek-zstd`, `omp-jsonl`, `none`).
- `capabilities.echo.predictProfile` — the predictive-echo model a composer needs.
The names live in `profiles.ts`, which is kept free of imports so `schema.ts` can validate a name at load time. A profile this build does not implement is a load-time error naming the field, rather than a CLI that silently looks permanently uninstalled.
## DeepSeek: the four assumptions it breaks
DeepSeek is worth reading before assuming an entry looks like its siblings — the schema carries four extensions because of it.
| What it breaks | How the registry expresses it |
| ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------ |
| `dsh` is a profile LAUNCHER, not the agent, so "installed" is not "runnable". | `discovery.launcherProfile` + `discovery.launcherTargetParam`. |
| Its permission switch is the **`DSH_PERMISSION_MODE` env var**, not a flag — the harness has none. | `env.configSetenv` (so the ordinary `privilegedParams` clamp still reaches it) **and** `capabilities.privilegedEnvKeys`. |
| It is the only non-claude mode with real hook signals, and for it that is a per-SESSION question. | `capabilities.hooks: 'supervised'` — a third state, not a boolean. |
| Its transcript is zstd session files, one frame per write. | `capabilities.transcript: 'deepseek-zstd'`. |
## Identity probes
`discovery.identity` asks the binary whether it is the program we meant, and it runs **before** the version probe, because a version probe cannot tell an impostor from the real thing. Debian ships an unrelated `dsh` (dancer's shell) that answers `--version` perfectly happily, and npm carries squatters for both `pi` and `grok`.
`discovery.version.requireVersionMatch` is the weaker companion: a binary whose version output has the wrong shape counts as ABSENT rather than present-with-unknown-version. That is what a short, generic binary name needs, and it is what keeps `codeman doctor` and the run mode from telling the user opposite things about the same binary — both read the same regex off the same entry.
## The no-id-branching rule
`test/cli-registry-no-id-branching.test.ts` fails the build if a CLI id comparison appears outside the stock catalog. It builds its id list from the live catalog, blanks comment lines before scanning (comments legitimately quote the pattern to explain why a branch was removed, and blanking rather than dropping is what keeps reported line numbers pointing at the real file), and keeps an allowlist in which **every entry carries its reason**.
It matches four shapes, not one: `mode === '<id>'`, `mode !== '<id>'`, `case '<id>':`, and `['<id>', …].includes(mode)`. The first version matched `===` only, and that gap was not academic — the refactor it guards converted the `===` sites and left the negated ones, so 36 `!==` branches survived it, including a seven-mode chain auto-enabling Ralph under a comment asking the next person to keep it in step with a predicate by hand while the sibling code path already read the capability. A guard that sees half the shapes reports a count measured over the half it happens to catch.
The allowlist is not a formality. If a branch is about what a CLI can DO it belongs in `CliCapabilities`; the entries that remain are things that are not CLI-behaviour branches at all — chiefly the legacy per-mode `<Mode>Config` objects on `POST /api/sessions`, which are a fact about the public HTTP API rather than about any CLI, plus a few documented cases where `mode === 'claude'` is genuinely the right question (Read My Mind reads Claude's _own_ transcript, so a capability there would be actively wrong).
## Two namespaces called `param`
`launch.params` keys, `env.configSetenv[].fromParam` and `capabilities.privilegedParams[].param` all name a **launch param**. The **legacy wire field** a param arrives as is a separate namespace, and `launch.legacyConfigAliases` is the only bridge between the two.
This matters because it is invisible when it is wrong. `capabilities.privilegedParams[].param` is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a name from the wrong namespace clamps **nothing**: no load error, no failing test, the clamp simply stops running. Codex is the entry where the two names differ (`bypassApprovals` as the param, `dangerouslyBypassApprovals` on the wire), so it is the one that catches a regression. `schema.ts` rejects any entry naming a param it never declared, on both `configSetenv.fromParam` and `privilegedParams.param`.
## Fields declared for later
`shortBadge`, `accent`, `capabilities.echo`, `capabilities.wheelForward`, `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` are **declared but not yet read**. They all describe frontend behaviour, and the frontend is deliberately untouched here: `app.js`, `terminal-ui.js` and `styles.css` keep their own hand-authored per-CLI rules, and moving them is its own piece of work verified by a browser/mobile suite the CI gate cannot see.
Treat those values as **transcribed, not authoritative** — nothing enforces that `echo.policy` matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches the gradient CSS paints, so re-measure before wiring one up. A field that is both wrong and unread is worse than an absent one, because the next reader trusts it; `test/cli-registry-no-id-branching.test.ts` pins the list so it cannot quietly grow, and wiring one up makes its line there fail, which is the direction you want.
`overlays.credStore` is in the same category, for a sharper reason: the Docker credential-seeding path still reads its own `CRED_STORES` table, because this shape allows ONE store per CLI and the live table needs two for gemini (`.gemini` for the CLI's own auth plus `.config/gcloud` for Vertex), while deepseek declares none here even though `.dsh` is seeded. Wiring it means making the field an array and correcting those two entries — a change to credential seeding, which is simultaneously the worst thing here to get wrong and the least covered by tests, since every docker IO path is no-op'd under vitest.
Everything else in the interface is live, including `overlays.remote` / `overlays.docker`, which back `defaultRemoteCommandForMode()` and `defaultDockerCommandForMode()` directly. Those two used to be hardcoded `Record<…CommandMode, string>` tables duplicating the registry with nothing keeping the two in step; `test/location-overlay-commands.test.ts` pins every resulting command as a literal string.
## Resolve at call time, never at import
Anything reading the registry must resolve it when it is asked, not when its module is first imported. `sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()` and each resolver's `searchDirs` thunk all re-read the catalog per call.
A module-level const freezes at first import, and the failure is asymmetric: a CLI enabled while the server is running moved the run menu but not the frozen surface, so validation rejected a mode the menu offered, or `codeman doctor` reported a catalog nobody had any more.
## Adding a CLI
1. Add a `CliEntry` to `stock.ts`.
2. Add a golden spawn-command pin to `test/cli-registry-spawn-golden.test.ts`, a row to `test/cli-capability-predicates.test.ts`, and its remote/docker commands to `test/location-overlay-commands.test.ts`.
3. That is usually all. If you find yourself wanting to add an `if` somewhere, the guard test will tell you — and the answer is a capability field, or a named profile if it genuinely needs to run code.
## See also
- [Agent CLIs](wiki/Agent-CLIs.md) — the user-facing per-CLI guide.
- `docs/architecture-invariants.md` — the mechanics and the history behind the rules above.
- `docs/deepseek-integration.md` — why DeepSeek is shaped the way it is.
+1 -1
View File
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` job's readiness poll looks for `❯`/a token count, neither of which pi prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` \| `grok` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` or `grok` job's readiness poll looks for `❯`/a token count, which neither CLI prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
+178
View File
@@ -0,0 +1,178 @@
# DeepSeek Harness (`dsh`) integration plan
> **Status**: Executed. This document records the plan, the decision behind each
> wiring point, and what was and was not verified. The user-facing guide is
> [`deepseek-integration.md`](./deepseek-integration.md); the per-decision
> invariants live in
> [`architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek`](./architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok-deepseek).
> Template: the grok integration ([`grok-integration-plan.md`](./grok-integration-plan.md)),
> itself calibrated against pi. Every fact below was measured against a live
> **dsh 0.1.1-rc.2** install and **@deepseek-harness-tui/dsh-tui 0.9.0**, not read
> from documentation.
## 1. What the DeepSeek Harness is
[deepseek-ai/deepseek-harness](https://github.com/deepseek-ai/deepseek-harness)
(open-sourced 2026-08-13, MIT) is a plugin-native agent framework: tools, skills,
sessions, sandboxes and whole APPS are Cordis plugins composed into *profiles*.
`dsh` is the launcher — `dsh --profile <name>` boots
`$DSH_HOME/profiles/<name>`, an ordered stack of plugin-bundle patch layers under
the user's own overrides. State lives in `~/.dsh` (`.env` 0600, `settings.yaml`,
`cordis.patch.yml`, `profiles/`, `sessions/`, `storages/`).
## 2. Shape decisions (why DeepSeek is wired the way it is)
DeepSeek is a ninth run mode. Never a location overlay, never a web tab (the
browser UI is handled separately, §3). Three of its decisions have no precedent
in the six external CLIs before it.
| Question | Decision | Why |
| --- | --- | --- |
| What does a pane run? | `dsh --profile <name>`, profile discovered | **The decision that shapes everything else.** DeepSeek ships `web`, `headless` and `base` — no terminal agent. The interactive front door is always a third-party plugin, so Codeman resolves a binary AND a profile inventory, and "available" means both. `resolveDefaultDeepSeekProfile()` prefers a recognized TUI, then an UNRECOGNIZED profile (anyone can publish an app bundle; a classifier that has not heard of one must not hide it), and refuses `web`/`headless`, which cannot occupy a pane. |
| Which TUI? | none blessed; default for BOOTSTRAP only | `POST /api/deepseek/install-profile` defaults to `@deepseek-harness-tui/dsh-tui` (~27.5k weekly downloads, ~4x the next, MIT, and it speaks the status contract in §2.3), but accepts any npm name and the resolver never assumes that profile exists. Codeman offers a default; it does not pick a winner. |
| Permission bypass | `DSH_PERMISSION_MODE` env export, no flag | The harness has NO command-line permission option; its sandbox/approval rows read one env var with three presets (`read-only` / `workspace-write` / `danger-full-access`, read off `dsh --dump-default-config`). This is the one legitimate exception to the `CLAUDE_CODE_EFFORT_LEVEL` ban: that var hard-locks in-session switching, whereas the harness reads this with `??` as a boot-time DEFAULT, so it stays soft. Exported via `tmux setenv`, never on the command line. The Run button sends `danger-full-access`, matching every sibling Run button. |
| Multi-user clamp branch | only-if-sent, clamped to `workspace-write`, **plus an env-var half** | Omitting the export leaves the harness on `workspace-write`, which still ASKS, so an absent config is already safe (the codex/antigravity/grok shape, not pi's materialize). Clamping to `workspace-write` rather than `read-only` is deliberate: the clamp removes privilege, it must not break a session's ability to edit its own workspace. ⚠️ Unlike every sibling, clamping the CONFIG is only half the gate: the switch is an env var, `DSH_*` is an allowlisted `envOverrides` prefix, and `applyEnvOverrides()` runs AFTER `_configureDeepSeek()`, so `envOverrides: {DSH_PERMISSION_MODE: 'danger-full-access'}` on the same request would land last and win. `clampEnvOverridesForOwner()` drops `DSH_PERMISSION_MODE` and `DSH_HOME` for a non-granted owner (dropping falls through to the clamped export). `DSH_HOME` because it aims the launcher at a profile tree whose plugin code runs at BOOT, before any approval row. |
| `hooksAvailableForMode()` granularity | per SESSION for deepseek, per mode for everything else | `deepSeekConfig.statusReporting: false` disarms the `HERDR_*` export, and the triple is the only reason a dsh session posts anything, so a mode-only answer would accept `until=stop` where nothing can send one — the infinite-wait the predicate exists to prevent. Call sites pass `sessionHookOptions(session)`; the default stays permissive so a forgotten one degrades to the old behaviour. ⚠️ Profile conformance stays unknowable at request time (an unrecognized profile is deliberately launchable), so a non-conforming TUI still times out on an explicit `stop`; the default set keeps `idle`/`exit` for that. ⚠️ The predicate is NOT "is this claude": Read My Mind and intent capture read Claude's transcript and were silently widened by this change, so they compare `mode === 'claude'` directly now. |
| Profile install spawn | own process group, hand-rolled timeout | `dsh plugin add` fans out into package-manager children, and spawn's built-in `timeout` signals only the direct child: survivors keep the inherited stdio pipes open, `close` never fires, and the held-open request leaks with no route-level deadline. `detached: true` + negative-pid SIGTERM→SIGKILL, the same escalation `runGit()` uses for the same reason, plus a last-resort reap for a grandchild that escaped the group. |
| Idle detection | **real hook events via a status shim** | The standout decision. The TUI already reports its lifecycle to a supervising process through a generic env-gated contract inherited from Herdr: `HERDR_ENV=1` + `HERDR_BIN_PATH` + `HERDR_PANE_ID` make it run `<bin> pane report-agent <id> --state idle\|working\|blocked …` on every state change, exit 0 = delivered. `deepseek-status-shim.ts` generates a script into the data dir and points `HERDR_BIN_PATH` at it. So deepseek is the only non-claude mode that passes `hooksAvailableForMode()` — earned by emitting definitive signals, not granted. An interface implementation, not an impersonation: no real `herdr` binary is ever executed, and a TUI that ignores the contract simply falls back to output stabilization. |
| `agent_working` event | new, 157th SSE constant | The one hook event with no Claude Code hook behind it. A harness turn cannot run while its own modal approval is on screen, so "started working" proves a dialog was answered in the terminal. Without it a dsh red alert would survive until the next `stop` — the exact stuck-alert bug the claude path already fixed once, and its pane-capture staleness sweep is Claude-dialog-shaped and cannot help here. |
| Resolver | identity probe THEN version probe | Strictest of the family, and not by preference. `dsh` is not merely a squattable npm name: Debian ships an unrelated `dsh` (dancer's shell, `apt install dsh`) which would answer a version probe convincingly and then be handed a spawn line. `dsh --help` must match `DeepSeek Harness` first. `DEEPSEEK_VERSION_REGEX` keeps the prerelease tail (`0.1.1-rc.2`), since truncating it would report an rc as a release. |
| Env allowlist | `DSH_*` + `DEEPSEEK_*` | `DSH_*` covers the launcher's documented inputs (`DSH_HOME`, `DSH_PERMISSION_MODE`, `DSH_TELEMETRY_MODE`, the `DSH_TUI_*` knobs); `DEEPSEEK_*` is the vendor namespace holding `DEEPSEEK_API_KEY`/`DEEPSEEK_BASE_URL`, same reasoning that admitted `XAI_*` for grok. ⚠️ Pi's lesson repeats exactly: a dsh `settings.yaml` can nominate ANY env var as a provider credential (`apiKeyEnv`), and the allowlist is one GLOBAL list, so admitting those would widen every mode at once. They stay out. |
| Model | NOT a session field | The model is a composition entry (`agent-default-model`) in the profile's config tree, set in `~/.dsh/settings.yaml` + `cordis.patch.yml`. Both create paths deliberately resolve no model for this mode rather than inventing a flag. |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Third-party fullscreen TUIs with their own scrollback and mouse handling — the opencode case, not the Ink case. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against a live authenticated session (see §5), same honest gap grok shipped with. The leading TUI's composer supports `@` completion and history search, which *may* make it per-keystroke reactive like codex; if so the fallback is the `'off'` branch. |
| Docker | image installs dsh AND a profile | Profiles are deliberately NOT seeded from the host: each is a per-profile `node_modules` tree, host-arch-specific and far too large to copy per container start. Only `~/.dsh/.env`, `settings.yaml`, `cordis.patch.yml` are seeded (auth + model composition). The profile install rides the `useradd` layer so the closing `chgrp`/`chmod g=u` covers it, which is what keeps it usable under the arbitrary uid the container runs as. |
| Remote SSH | `exec "$SHELL" -i -l -c 'dsh'` | Boots the remote box's default profile; a remote with several needs the per-host `commands.deepseek` override, since `deepSeekConfig` does not cross ssh. |
## 3. The web profile
The browser UI is the only interactive surface DeepSeek ships itself, so it gets
a **shortcut, not a run mode**: `Run ▸ DeepSeek web UI…` starts
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-authority>`
as a background process and opens the URL as an ordinary web tab.
The server was a **shell session** first, on the reasoning that Codeman already
supervises those (visible, scrollable, killable, dies with its tab) so nothing
new had to own a long-lived HTTP server. That version worked and was still
wrong in use: clicking "open the DeepSeek web UI" put a terminal tab on screen
next to the web tab actually asked for, every single time, and after the first
launch the terminal was pure noise. Opening a dashboard should open one tab.
So `POST /api/deepseek/web` owns it instead (`src/deepseek-web-server.ts`), and
what the session gave away for free is now explicit: exactly one server, reused
rather than raced on a second click; restarted when the requested authority
changes; killed on Codeman shutdown (a detached child would otherwise hold its
port against the next start — the very EADDRINUSE this feature already got
wrong once); and boot output captured, since with no shell tab there is nowhere
else for a stack trace to land. It is fenced at the same bar as the profile
installer: booting a dsh profile executes the plugin code in it, so it requires
the privileged grant in multi-user mode.
`--trusted-host` is load-bearing — dsh fences its `/api` behind a browser-trust
check on the request authority, and a Codeman web tab reaches it through
Codeman's own origin via the webview proxy, not directly. The authority comes
from the CLIENT (`location.host`) because only the browser knows which of a
multi-homed Codeman's origins is actually in play.
Three things about this shortcut are load-bearing and each came from it failing
in exactly that way against a real install:
- **The port is chosen, never hardcoded.** `GET /api/deepseek/web-port` walks
3080..3119 for a free loopback port. 3080 is dsh's own default, which makes it
precisely the port a DeepSeek user is most likely to already be serving on:
binding it unconditionally killed the launch with `EADDRINUSE` against the
user's own `dsh web`.
- **The tab is opened only after the server answers.** The launch polls
`POST /api/webviews/probe` until the URL responds, so a server that dies on
startup reports the failure and points at its shell tab, instead of silently
persisting a dashboard aimed at nothing.
- **The saved tab is `trusted: true`, and must be.** An untrusted webview is
sandboxed without `allow-same-origin`, which breaks this dashboard twice: the
dsh client-runtime reads `localStorage` while loading plugins and dies there,
and an opaque-origin frame sends `Origin: null`, so dsh's trust check 403s
every `/api` call regardless of what `--trusted-host` names. Passing
`location.host` only means anything once the frame actually carries that
origin. The trade is real — a trusted proxied frame is same-origin with
Codeman and can reach Codeman's API — and is defensible only because this
particular dashboard is an agent harness Codeman just started itself on
loopback, which can already run code as the user. It is not a precedent for
trusting third-party dashboards generally.
The record is marked `managed: 'deepseek-web'`, which keeps it out of the
saved-dashboard list: the shortcut that maintains it is already a menu entry, so
listing both showed the same dashboard twice. Being managed is also what lets a
relaunch repoint the existing row instead of stacking one dead dashboard per
restart, since the port is now chosen per launch.
The authority baked into `--trusted-host` is the one the launch was clicked
from, and reuse is conditional on it: a running server fenced for a *different*
origin is stopped and restarted rather than reused, because reusing it renders a
page whose every API call 403s — which reads as a broken dashboard rather than a
misconfigured one.
## 4. Touch points (the checklist)
Backend: `types/session.ts` (SessionMode + `DeepSeekConfig` + SessionState),
`utils/deepseek-cli-resolver.ts` (new) + barrel, `deepseek-status-shim.ts` (new),
`tmux-manager.ts` (`buildDeepSeekCommand`, dispatch, resume flag, PATH export,
truecolor, `_configureDeepSeek`, availability error, plumbing), `session.ts`
(external-mode gate, label, config plumbing, tmux-required error, attach env),
`mux-interface.ts`, `schemas.ts` (prefixes, `DeepSeekConfigSchema`,
`DeepSeekInstallProfileSchema`, both mode enums, remote overrides, cron agentType,
`agent_working`), `session-wait-registry.ts` (`hooksAvailableForMode`),
`hook-event-routes.ts` (`APPROVAL_RESOLVING_EVENTS`), `session-routes.ts` (clamp +
both create paths + `resolveDeepSeekLaunchError`), `system-routes.ts`
(`GET /api/deepseek/status`, `POST /api/deepseek/install-profile`), `server.ts`
(availability inject + mux restore), `sse-events.ts`, `docker-hosts.ts`,
`remote-hosts.ts`, `config/dependency-registry.ts`,
`response-viewer-transcript.ts`, `cron/cron-service.ts` (comment),
`tui/tui-client.ts` + `tui-app.ts`.
Frontend: `index.html` (welcome button, run-mode entry, install affordance, web-UI
shortcut, cron option, clone Brain option), `session-ui.js` (`runDeepSeek()`,
`runDeepSeekWeb()`, `installDeepSeekProfile()`, dispatch, availability, "Run DS"
label, external-CLI gates), `app.js` (label, `ds` tab badge, kill-menu, SSE map),
`settings-ui.js` (welcome gate + `_onHookAgentWorking`), `constants.js`,
`mobile-overview.js`, `home-sessions.js`, `panels-ui.js`, `i18n.js`,
`terminal-ui.js`, `styles.css` + `mobile.css` (brand-indigo identity; the non-og
skin block and the mobile `!important` pair are both load-bearing).
Meta: `docker/agent.Dockerfile`, `install.sh`, `package.json` keyword,
`skills/codeman/reference/*`, CLAUDE.md, `architecture-invariants.md`.
Tests: `test/deepseek-mode.test.ts` + `test/deepseek-cli-resolver.test.ts` (new);
`run-mode-ui`, `render-index-html`, `mobile-overview`, `agent-skill-mode-lists`
(extended).
## 5. Verification performed
See the summary at the end of the implementing session for the live run. In
short: the CI gate green; the resolver, profile inventory, spawn-line and clamp
behaviour covered by 31 new unit tests; and an isolated instance used to exercise
`GET /api/deepseek/status` and a real session against the live dsh install.
**Not verified (honest gaps):**
- The local-echo `'buffer'` policy against the TUI's real composer (§2). If it
turns out per-keystroke reactive like codex's, flip it to the `'off'` branch;
teaching `PredictiveEchoAddon` its composer row is the larger follow-up.
- Scrollback/repaint behaviour of a third-party fullscreen TUI under the narrow
strip during a long session.
- A Docker case with `mode: 'deepseek'` (needs a `--no-cache` agent-image
rebuild — see the `--no-cache` rule in CLAUDE.md).
- A remote-SSH deepseek case.
- The web-UI shortcut against a tunnel authority. Loopback and a tailnet name are
both verified end to end through the webview proxy (dashboard renders, its
`/api` calls succeed, no shell session created).
## 6. Follow-ups
- **Response viewer**: read `~/.dsh/sessions/**` (JSONL) the way codex rollouts
are read back. Highest-value follow-up, and very achievable.
- **`headless` as an execution backend** for Codeman's own AI checks
(`ai-idle-checker`, `ai-plan-checker`), today Claude-only.
- **Profile/model picker in Session Options**, reading `GET /api/deepseek/status`
`.profiles`.
- **`--patch` overlays per session**, which is the harness-native way to change
agent composition without touching the user's profile.
- Measure the local-echo policy and pin the result the way pi did.
+305
View File
@@ -0,0 +1,305 @@
# DeepSeek Harness (`dsh`) in Codeman
Codeman can run [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness)
as a session backend, alongside Claude Code, OpenCode, Codex, Gemini,
Antigravity, Pi and Grok. It is the ninth run mode, and the one that is wired
least like the others, for two reasons worth understanding before you use it.
## 1. The agent is a profile, not the binary
`dsh` is a **launcher**, not an agent. It boots a *profile*: an ordered stack of
plugin-bundle patch layers under `$DSH_HOME/profiles/<name>` (`$DSH_HOME`
defaults to `~/.dsh`). DeepSeek ships three bundles and none of them is a
terminal agent:
| Profile | What it is | Can Codeman run it in a tab? |
| ------------ | --------------------------------- | ---------------------------- |
| `web` | the browser UI, served on :3080 | no — but see §6 |
| `headless` | answers one task and exits | no |
| (`base`) | the shared core, no app at all | no |
The interactive terminal front door is **always a third-party plugin**. So
"DeepSeek is installed" and "Codeman can start a DeepSeek session" are different
questions, and Codeman answers both separately:
```bash
curl -s localhost:3000/api/deepseek/status | jq
{
"available": true, # the `dsh` binary resolved and proved its identity
"runnable": false, # ...but nothing installed can drive a pane
"path": "/home/you/.local/bin",
"version": "0.1.1-rc.2",
"dshHome": "/home/you/.dsh",
"defaultProfile": null,
"profiles": [ { "name": "web", "kind": "web", "bundles": [...] } ]
}
```
### Installing a terminal profile
From the UI: open the **Run** dropdown. When `dsh` is installed but no
pane-capable profile is, the menu shows **DeepSeek — add a terminal profile…**.
One click installs one and the normal DeepSeek entry appears.
By hand, or to pick a different front door:
```bash
dsh plugin --profile dsh-tui add @deepseek-harness-tui/dsh-tui
```
⚠️ **`pnpm` has to be on PATH for either route.** `dsh plugin` is a thin forwarder
that spawns a literal `pnpm` with no npm fallback, so without one it exits 127 with
`dsh: pnpm not found on PATH` — both by hand and behind the UI button, which
surfaces that same line as the install error. `npm install -g pnpm` (or
`corepack enable pnpm`) is the fix. This is what broke the Docker agent image in
[#352](https://github.com/Ark0N/Codeman/issues/352); the image now installs pnpm
alongside `dsh`.
Codeman's default is `@deepseek-harness-tui/dsh-tui` because it is by a wide
margin the most used community TUI, it is MIT, and it implements the status
contract described in §3. It is a **default, not a requirement**: any profile
under `$DSH_HOME/profiles` that is not `web` or `headless` shows up in the
inventory and can be launched, including one you compose yourself. The endpoint
accepts any npm package name:
```bash
curl -sX POST localhost:3000/api/deepseek/install-profile \
-H 'Content-Type: application/json' \
-d '{"profile":"my-tui","package":"@someone/dsh-tui"}'
```
Installing a plugin is arbitrary code execution on the host, so in multi-user
mode this endpoint requires the can-bypass-permissions grant (the same bar as a
`shell` session). The request is held open while the package manager runs and is
bounded at five minutes; the install runs in its own process group, so hitting
that bound kills the whole tree rather than just the launcher.
> **`dsh` is also a Debian program.** `apt install dsh` gives you "dancer's
> shell", a distributed shell, which would answer `--version` convincingly.
> Codeman's resolver therefore demands the harness's own help banner before it
> will point a spawn line at a candidate, and `GET /api/deepseek/status` reports
> `path` and `version` so a misresolution is diagnosable rather than presenting
> as "the mode just doesn't work".
## 2. Permissions are an env var, not a flag
The harness has **no `--dangerously-skip-permissions` equivalent**. Its sandbox
and approval rows are configuration, driven by one documented input,
`DSH_PERMISSION_MODE`, with three presets (read off `dsh --dump-default-config`):
| `DSH_PERMISSION_MODE` | sandbox | approvals | notes |
| --------------------- | -------------------- | --------- | ------------------------- |
| `read-only` | `read-only` | ask | |
| `workspace-write` | `workspace-write` | ask | the harness's own default |
| `danger-full-access` | `danger-full-access` | **never** | what the Run button sends |
Codeman exports it via `tmux setenv`, never on the command line. Because the
harness reads it with `??`, it is a **soft default**: it sets the boot-time
preset and you can still change permission mode inside the session.
Omitting it entirely leaves the harness on `workspace-write`, which still asks —
which is why the multi-user clamp only needs to force a *sent* value down. A
non-granted owner's `danger-full-access` becomes `workspace-write`, not
`read-only`: the clamp removes privilege without breaking the session's ability
to edit its own workspace.
Because the switch is an env var rather than a flag, that clamp has a second half
no other CLI needs. `DSH_*` is an allowlisted `envOverrides` prefix (it has to be:
that is also how you set the harness's ordinary knobs), and env overrides are
applied *after* the permission export, so in multi-user mode a non-granted owner
sending
```json
{ "mode": "deepseek", "envOverrides": { "DSH_PERMISSION_MODE": "danger-full-access" } }
```
would otherwise hand back the privilege the config clamp just removed. For a
non-granted owner Codeman therefore **drops `DSH_PERMISSION_MODE` and `DSH_HOME`
from `envOverrides`**; dropping them falls through to the clamped config and the
server's own `DSH_HOME`. `DSH_HOME` is in that list because it points the
launcher at a profile tree, and a profile's plugin code runs at boot, before any
approval row can apply. Single-user installs and granted owners are unaffected.
## 3. Real idle detection (the interesting part)
Every other external CLI mode in Codeman is **readiness-guessed**: Codeman
watches the PTY go quiet and infers that a turn ended. Claude is the exception,
because Claude Code fires hooks.
DeepSeek is the second exception. The community terminal front door already
reports its own lifecycle to a supervising process through a generic,
env-var-gated contract (inherited from [Herdr](https://herdr.dev)): when
`HERDR_ENV=1`, `HERDR_BIN_PATH` and `HERDR_PANE_ID` are set, it shells out on
every state change with
```
"$HERDR_BIN_PATH" pane report-agent "$HERDR_PANE_ID" \
--source custom:dsh-tui --agent dsh-tui \
--state idle|working|blocked [--message ...] --seq N
```
Codeman points `HERDR_BIN_PATH` at a small generated shim
(`~/.codeman/dsh-status-shim.mjs`, written at session create) which forwards each
report to `POST /api/hook-event`. The mapping:
| Harness state | Codeman hook event | What you get |
| ------------- | ------------------ | -------------------------------------------------------- |
| `blocked` | `permission_prompt`| red "needs you" tab alert + an Approvals Inbox item |
| `idle` | `stop` | definitive end-of-turn: respawn triggers, `wait` returns |
| `working` | `agent_working` | clears an alert answered in the terminal, at once |
So a DeepSeek session gets Claude-grade signals: `GET /api/sessions/:id/wait`
really can block on `stop` and `blocked` for it, and it is the only non-Claude
mode for which that is true (`hooksAvailableForMode`).
That is a per-*session* answer, not a per-mode one. Turning the bridge off with
`deepSeekConfig.statusReporting: false` means nothing will ever post a hook event
for that session, so an explicit `until=stop` is refused up front (with a message
naming the setting) rather than blocking for your whole timeout. Omitting `until`
never fails: the hook-only signals are dropped from the default set and you still
get `idle` and `exit`.
One limit worth knowing: whether the *profile* implements the contract cannot be
known at request time (Codeman deliberately treats an unrecognized profile as
launchable). A dsh session running a non-conforming TUI therefore still accepts
`until=stop` and will time out on it. `idle`/`exit` are the reliable pair there.
This is an interface implementation, not an impersonation — nothing on your
machine executes a real `herdr` binary. If you use a terminal profile that does
*not* implement the contract, the shim is simply never called and the mode falls
back to output-stabilization readiness like its siblings. Turn it off per session
with `deepSeekConfig.statusReporting: false`.
## 4. Starting a session
From the UI, pick **DeepSeek** in the Run dropdown (or the **Run DeepSeek**
welcome button) and press Run. Over the API:
```bash
curl -sX POST localhost:3000/api/quick-start \
-H 'Content-Type: application/json' \
-d '{
"caseName": "myproject",
"mode": "deepseek",
"deepSeekConfig": {
"profile": "dsh-tui",
"permissionMode": "danger-full-access"
}
}'
```
`deepSeekConfig` fields: `profile`, `permissionMode`, `resumeSession`,
`resumeSessionId`, `statusReporting`. Resume prefers an explicit id over the
most-recent form, and both are passed through to the profile's app, which is
where `--resume` is understood.
**Models are not a session field.** The model is a composition entry in the
profile's config tree (`agent-default-model`), not a CLI flag, so Codeman does
not try to set one. Configure it where the harness does: `~/.dsh/settings.yaml`
plus a home-level `~/.dsh/cordis.patch.yml`, or a `--patch` overlay on the
profile. That is also how you point dsh at a local or third-party provider.
**Environment.** `DSH_*` and `DEEPSEEK_*` are allowlisted for `envOverrides`
(so `DSH_HOME`, `DSH_PERMISSION_MODE`, `DEEPSEEK_API_KEY`, `DEEPSEEK_BASE_URL`
all flow through). Provider keys with *other* names are deliberately not: a dsh
`settings.yaml` can nominate any env var as a credential via `apiKeyEnv`, and
Codeman's allowlist is global, so admitting them would widen it for every mode at
once. Authenticate those the way dsh does, from the file or the server's own
environment.
## 5. Reading a session back, and driving one as a worker
dsh writes a real transcript — `$DSH_HOME/sessions/<mangled-cwd>/<id>/session.jsonl.zstd`
— so `GET /api/sessions/:id/last-response` reads that rather than segmenting the
pane, and the Response Viewer shows a dsh conversation the way it shows a claude
or codex one (`?context=full` returns prompt / response / tool blocks).
Reading the pane instead is not merely coarse for this mode, it is wrong: dsh-TUI
paints a full-screen splash, so the segmenter answered a `last-response` call for
a fresh dsh session with its ASCII-art logo — which anything polling for a
worker's first answer reads as an answer. Three things about the file shaped the
reader (`src/deepseek-transcript.ts`):
- **It is one zstd FRAME per append, not one zstd stream.** `zstd -dc` decodes all
of them, Node's `zlib` zstd decoder stops at the first: a real 56-line
transcript came back as 1 line. The reader walks frame headers itself. On a Node
older than 22.15 (no zstd at all) the mode falls back to the pane, as before.
- **Not every `user/message` is the user.** Each turn also records a
plugin-sourced runtime-context snapshot; only `source.kind === 'user'` is a
prompt.
- **A failed turn is not an empty one.** `turn/end` carries the provider's error,
which is returned as `Turn error: …` (and an early stop such as `max-tokens` as
`Turn ended: …`) instead of an empty string that reads as "still thinking".
The transcript reader applies to **local** dsh sessions only. A Docker case's
harness writes its transcript inside the container's own `~/.dsh` (the workspace
bind mount does not cover it), and a remote-SSH case's lives on the remote host,
so the local reader could never find those files — such sessions keep the pane
segmenter, coarse but real. The splash caveat above applies to them accordingly.
### As an agent worker
Because dsh has both halves — a real end-of-turn signal and a real transcript — an
agent can drive a dsh session the same way it drives a claude one, and the bundled
`codeman` agent skill does. Spawning `beta:deepseek` in its worker list gives a
worker that is tasked, waited on and read with the same calls as its claude
siblings; no other external CLI mode qualifies. Two edges are worth repeating here:
- **Readiness is not the stop signal.** The harness reports `idle` at boot roughly
300 ms *before* the composer paints (measured 2.26 s vs 2.56 s after spawn), so a
send-and-wait fired immediately after create resolves on that boot report,
reports a turn that never ran, and leaves the prompt in a pane that was not yet
accepting input. Wait for the composer (`❯`) instead.
- **Wait on `stop`, not on the default signal set.** That set also carries `idle`,
which for every external CLI is inferred from output stabilization; a dsh TUI
that repaints rarely reads as idle mid-turn.
## 6. The web UI as a tab
The browser UI is the one interactive surface DeepSeek ships itself, so it gets a
shortcut rather than a run mode: **Run ▸ DeepSeek web UI…** starts
`dsh web --no-open --host 127.0.0.1 --port <free> --trusted-host <codeman-host>`
as a background child process (`src/deepseek-web-server.ts`, behind
`POST/GET/DELETE /api/deepseek/web`) and opens it as a Codeman web tab once the
server actually answers.
It is a child process rather than a shell session because the session version
opened a terminal tab nobody asked for on every click. What the session gave for
free is therefore explicit here: one instance with reuse, a restart when the
requested `--trusted-host` authority differs from the running one, a kill on
server stop, and captured boot output. The `--trusted-host` flag is load-bearing —
dsh fences its `/api` behind a browser-trust check on the request authority, and a
Codeman web tab reaches it through Codeman's own origin via the webview proxy, not
directly. Without it the page renders and every API call fails.
## 7. Docker and remote cases
Docker cases work: the agent image installs `dsh` and bootstraps a `dsh-tui`
profile into the container. Profiles are deliberately **not** seeded from the
host (each is a per-profile `node_modules` tree, host-arch-specific and far too
large to copy on every container start); only `~/.dsh/.env`, `settings.yaml` and
`cordis.patch.yml` are seeded, which is what carries auth and model composition
in. As with pi and grok, in-container sessions are invisible host-side:
`~/.dsh/sessions` inside a container is that container's own.
Remote SSH cases default to `dsh` through a login shell, which boots the remote
box's default profile. If the remote has several, name one with the per-host
`commands.deepseek` override — the local `deepSeekConfig` does not cross ssh.
## 8. What is not wired
Deliberately minimal, on the same reasoning as the grok integration: the harness
is a fast-moving developer preview and every flag added is a flag validated
forever.
- `--patch` overlays per session (the profile's own layers apply as normal).
- `dsh plugin` management beyond first-time profile install.
- The `headless` profile as a one-shot execution backend for Codeman's own
internal AI checks (today those are Claude-only).
- Model/provider selection from Session Options.
## Verified against
`dsh 0.1.1-rc.2` and `@deepseek-harness-tui/dsh-tui 0.9.0`. The permission
presets, the profile layout, and the supervisor contract above were all read off
the live install rather than from documentation.
+16 -4
View File
@@ -2,7 +2,7 @@
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` / `grok` / `deepseek` / `omp` all work inside the container.
## One-time setup: build the base image
@@ -25,12 +25,24 @@ A zero exit code only proves the layers ran, not that the toolchain works. Verif
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
'for c in claude codex gemini opencode agy pi grok dsh omp; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other four npm CLIs install.
⚠️ `dsh --version` is the one line above that answers a different question than the
others: `dsh` is a profile launcher, so a working binary says nothing about whether
the image can actually run a DeepSeek session. Check the profile the Dockerfile
installs into the agent's HOME as well, or a `mode: 'deepseek'` case starts a pane
that dies on arrival:
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md).
```bash
docker run --rm codeman/agent:base ls ~/.dsh/profiles/dsh-tui/package.json
```
Building that profile is also why `pnpm` is in the image: `dsh plugin` forwards straight to a literal `pnpm` and exits 127 without it (issue #352), and pnpm — unlike npm — blocks dependency lifecycle scripts by default and fails the install over it, so the profile step passes `--config.dangerouslyAllowAllBuilds=true`.
Antigravity (`agy`) and Grok (`grok`) are the two CLIs not installed from npm (Google and xAI ship standalone binaries), so each has its own Dockerfile step, adding roughly 190MB and 160MB respectively. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other npm CLIs install.
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md). Grok is seeded per-file for the same reason (`auth.json`, `config.toml`, `pager.toml` out of `~/.grok`, which also holds `sessions/`, `memory/` and the ~160MB binary under `downloads/`), with the same consequence for `grok -c`. See [`grok-integration.md`](./grok-integration.md). OMP is the one CLI in this family where `sessions/` is the EXCEPTION rather than the rule: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded per-file (the dir also holds SQLite caches and `terminal-sessions/`), but `~/.omp/agent/sessions/` is shared RW like codex's, not seeded, because Codeman reads it host-side for history recovery and `--resume` pinning. See [`omp-integration.md`](./omp-integration.md).
## Quickest path: one-click "Run in Docker"
+78
View File
@@ -0,0 +1,78 @@
# Docker Compose deployment
This configuration builds the Codeman application image locally from this checkout. It does not download or depend on a pre-built Codeman image.
For the Compose configuration, environment settings, storage migration, and macvlan networking examples, see the [Docker deployment guide](../docker/README.md).
The image includes Claude Code, Codex, Gemini CLI, and OpenCode. Authenticate a CLI from its Codeman session; credentials are never baked into the image.
## Prerequisites
- Docker Engine or Docker Desktop with Docker Compose v2
- A reachable Docker daemon
The application container mounts the Docker daemon socket so Codeman can create and manage its isolated Docker cases. Treat anyone who can administer this Compose project as having Docker-host-equivalent access.
## Start
Copy the environment template, set a strong password, and confirm `CODEMAN_APPDATA_PATH`. The example maps `/mnt/user/appdata/Coding/codeman` on the host to `/home/${CODEMAN_RUNTIME_USER}` in the container, preserving Codeman state and CLI credentials outside Docker-managed volumes.
```sh
cp docker/.env.example docker/.env
```
On PowerShell, use the following command instead.
```powershell
Copy-Item docker/.env.example docker/.env
```
On Linux, run the stack with the start script. It determines `PUID` and `PGID` from the owner of `CODEMAN_APPDATA_PATH`, and `DOCKER_SOCKET_GID` from the configured Docker socket, before invoking Compose. A root-owned application-data directory is rejected so the runtime account cannot become UID 0.
```sh
bash docker/Start-Codeman.sh
```
On other platforms, run Compose directly. `PUID` and `PGID` default to `1000:1000`; set them in `docker/.env` when the application-data directory has a different owner.
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml up --build -d
```
Open `http://localhost:3000` and sign in with the username and password from `docker/.env`.
## Operations
The local image is tagged `codeman:local` by default. Change `CODEMAN_IMAGE` in `docker/.env` if a different local tag suits your environment.
```sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml logs -f codeman
bash docker/Start-Codeman.sh
docker compose --env-file docker/.env -f docker/docker-compose.yaml down
```
`CODEMAN_APPDATA_PATH` holds Codeman state and survives container recreation. Remove that host directory only when deliberately resetting the installation.
`CODEMAN_CASES_PATH` must be an absolute path on the Docker host. Compose mounts it at the same path inside Codeman, so the host daemon can bind the managed workspace into isolated Docker cases. Do not set it to `/home/${CODEMAN_RUNTIME_USER}/codeman-cases`.
Compose passes `CODEMAN_APPDATA_PATH` into Codeman as `CODEMAN_DOCKER_HOST_HOME`. Codeman uses that value to translate generated Docker seed, credential and hook-secret bind sources from the container's home path into paths visible to the host Docker daemon.
If `docker info` reports `SwapLimit=false`, set `CODEMAN_DOCKER_DISABLE_SWAP_LIMIT=1`. Isolated cases retain their configured memory limit. Codeman omits the unsupported swap-limit option and filters only the daemon's exact swap-capability warning while retaining every other Docker create error.
If that directory was created by an earlier root-running image, change its ownership to the configured `PUID:PGID` before starting this version. This preserves existing CLI credentials and session state while allowing the unprivileged runtime account to use them.
## Updating
Codeman updates itself from **App Settings → Updates**, as it does on a bare host. The checkout mounted at `/opt/codeman` is the same directory Compose builds from, so the update's `git checkout` and rebuild land on the host and survive container recreation; the restart is the server exiting, which `restart: unless-stopped` turns into a relaunch on the new build.
That applies application code only. A release that changes `docker/server.Dockerfile`, `docker/docker-compose.yaml`, or adds a key to `docker/.env.example` needs the image rebuilt or the container recreated, which a container cannot do to itself. The updater detects each case and refuses with a message naming what changed; run `docker/Start-Codeman.sh` on the host to apply those.
`CODEMAN_REPO_PATH` overrides which checkout is mounted. It defaults to the compose project's parent directory, so it normally needs no setting. Point it at a directory that is not a git checkout and in-app updates are reported as unavailable.
Full detail, including the fingerprint baseline and the troubleshooting table: [`docker-self-update.md`](docker-self-update.md).
## Docker cases
The default socket path is `/var/run/docker.sock`, which works with a standard Linux Docker Engine. The Bash start script detects its numeric group ID. When running Compose directly, set `DOCKER_SOCKET_GID`, for example using `stat -c '%g' /var/run/docker.sock`, so the unprivileged `CODEMAN_RUNTIME_USER` account can create Docker cases. Docker Desktop users should set `DOCKER_SOCKET` in `docker/.env` only when their Docker installation exposes a different compatible socket path.
Codeman Docker cases are sibling containers on the host daemon, not children of the application container. The Compose configuration handles their workspace bind mount through `CODEMAN_CASES_PATH`; the `/home/${CODEMAN_RUNTIME_USER}` application-data mapping is for Codeman state and ordinary in-container sessions, not sibling-case workspaces.
+218
View File
@@ -0,0 +1,218 @@
# Self-update in the Docker Compose deployment
Codeman running as a container updates itself from **App Settings → Updates**, the
same place and the same button as a bare-host install. This document explains how
that works, what it deliberately refuses to do, and how to recover when it stops.
The bare-host updater is documented in
[`architecture-invariants.md#self-update`](architecture-invariants.md#self-update);
this file covers only what the container changes.
## The short version
| Change in the release | Applied by |
| -------------------------------- | ------------------------------------------------ |
| Application code | The in-app updater |
| `docker/server.Dockerfile` | `docker/Start-Codeman.sh` on the host |
| `docker/docker-compose.yaml` | `docker/Start-Codeman.sh` on the host |
| New key in `docker/.env.example` | Add it to `docker/.env`, then `Start-Codeman.sh` |
The in-app updater detects all three of the bottom rows itself and refuses with a
message naming what changed, so you never have to work out which case you are in.
## Why the container needs its own path
The bare-host updater does `git checkout <tag> && npm install && npm run build`,
then asks systemd or launchd to restart the service. Two of those assumptions are
false in a container:
1. **There is no init system.** A container's supervisor is the Docker daemon,
which acts on the container, not on processes inside it.
2. **The image is immutable.** A `git pull` into the image's baked `/opt/codeman`
would land in the container's writable layer, survive `docker restart`, and be
silently discarded by the next `docker compose up`.
Both are solved by configuration rather than by a second updater:
- **The checkout is a host bind mount.** `docker-compose.yaml` mounts the repo
(the same directory used as the build context) over `/opt/codeman`, so the
updater's `git checkout` writes to the host filesystem and survives the
container being recreated.
- **The restart is the server exiting.** `restart: unless-stopped` relaunches the
container whenever its main process ends, including on a clean exit — so the
updater's final step is to signal the server, and Docker starts it again on the
freshly built `dist/`.
Everything else — the release-tag channel, the auto-stash, the atomic
`update-status.json` the browser polls across the connection drop, the boot-time
reconcile that flips `restarting` to `completed` — is the existing machinery,
unchanged. The container path is a new `SupervisorKind`, not a new updater.
## What the pieces are
| Piece | Role |
| ---------------------------------------------- | ------------------------------------------------------------------- |
| Repo bind mount at `/opt/codeman` | Makes the pull persistent. Without it, self-update is unavailable. |
| `codeman-node-modules`, `codeman-dist` volumes | Container-owned build artefacts, layered over the bind mount. |
| `CODEMAN_IN_CONTAINER=1` | Tells `detectSupervisor()` to restart by exiting. |
| `restart: unless-stopped` | Turns that exit into a restart. Verified before every update. |
| `CODEMAN_RESTART_BY_EXIT=1` | The Compose file's declaration of that policy, so the updater may exit even with no Docker socket. |
| Toolchain + devDependencies in the image | Lets `npm install` and `npm run build` run inside the container. |
| `docker-env-applied.json` | Fingerprint baseline, written by `Start-Codeman.sh` on every start. |
### Why build artefacts are in named volumes
`node_modules` and `dist` are mounted as named volumes **on top of** the repo bind
mount. Without that, an update's `npm install` would write into the host checkout,
leaving container-compiled native modules (node-pty builds from source here) in a
directory that may also be used to run Codeman natively, and leaving `git status`
permanently noisy.
Docker seeds an empty named volume from the image, so the first start inherits the
image's already-built `node_modules` and `dist` and pays no bootstrap cost.
`docker compose down -v` is the supported reset: the next start re-seeds them.
### Why the runtime image carries a build toolchain
`npm run build` is `tsc` plus `esbuild`, both devDependencies, so the image no
longer runs `npm prune --omit=dev`. And `npm install` may rebuild node-pty, which
ships no Linux prebuild, so `python3`, `make` and `g++` are installed as well.
This is the real cost of in-place updates: a noticeably larger image than a
runtime-only one. It buys an update that takes about a minute instead of a full
image rebuild, and it is why `NODE_ENV=production` is paired with an explicit
`npm install --include=dev` in the updater.
## The environment gate
An in-place update applies **code only**. A restarted container reuses its existing
image and configuration, so a release that changes the environment cannot take
effect that way — and would half-apply: new code against an old environment. The
updater therefore checks the **target release's own files**, read straight out of
git with `git show <tag>:<path>` before anything is checked out.
### 1. `server.Dockerfile` changed, so the image must be rebuilt
Compared by sha256 against the fingerprint `Start-Codeman.sh` recorded when the
running container was built.
### 2. `docker-compose.yaml` changed, so the container must be recreated
Same mechanism. A restart cannot pick up a new mount, port or environment
variable; only recreating the container can.
### 3. `.env.example` gained keys your `.env` has no value for
The check that matters most, because **Compose will not tell you**. An unset
`${VAR}` interpolates to the empty string; Compose prints a warning to a terminal
nobody is watching and starts anyway. A new required setting therefore arrives as
a silently blank environment variable and misbehaves later, far from the cause.
The updater names the missing keys instead.
Commented-out lines in `.env.example` are deliberately *not* keys — that is how
the file marks optional overrides such as `# PUID=1000`, and counting them would
block updates on settings you are meant to leave alone.
### 4. A restart policy that would not bring the container back
Before signalling the server, the updater asks the Docker daemon for its own
container's restart policy. If it is `no`, the update is refused: applying it
would take Codeman down and leave no UI to recover from.
If the policy cannot be read at all (no Docker socket mounted) the update is
still allowed, but the final step changes: the server exits only when the
Compose file declared `CODEMAN_RESTART_BY_EXIT=1` (the shipped one does, because
it is the file that sets `restart: unless-stopped`) or the daemon confirmed an
auto-restart policy. Otherwise the build completes and the panel asks you to
restart the container by hand. A container started by plain `docker run` with no
restart policy therefore gets a staged update, never an outage.
### What the gate deliberately does not do
Every unknown fails **open**:
- A missing fingerprint baseline (a container started before this feature existed)
is not treated as a change, or those installs could never update at all.
- An unreadable `.env`, an unreachable Docker socket, or a target tag whose files
cannot be read all yield "no blocker" rather than a refusal.
The one place an unknown does NOT fail open is the kill itself: with neither the
Compose declaration nor a daemon answer, the updater stages the build and asks
for a manual restart rather than exiting a server nothing may bring back.
The gate catches a specific, detectable class of mistake; it is not a last line of
defence. It is also re-evaluated server-side on `POST /api/system/update`, so
hiding the button in the UI is a courtesy rather than the control.
## The one residual risk
The gate is derived from the diff, so it cannot see a release that needs a newer
environment **without changing any of those files** — for example, code that
depends on newer agent-CLI behaviour.
That is why the four global CLIs in `server.Dockerfile` are **pinned**. Unpinned,
the versions a user ends up with are a function of when their image was built
rather than of any commit, and in-app updates make rebuilds rarer, which makes
that drift worse over time. Pinned, "this release needs a newer CLI" becomes a
Dockerfile change, which check 1 already detects. Bump them deliberately, as part
of a release.
The complementary merge-side guard is `test/docker-compose-env-parity.test.ts`,
which fails CI when a variable is added to `docker-compose.yaml` without an entry
in `.env.example`, or the reverse.
## Sequence of an in-place update
1. **Check** — `GET /api/system/update/check` finds the latest release tag, fetches
that one ref so the gate can read the target's files, and returns any blockers.
2. **Start** — `POST /api/system/update` re-evaluates the gate, writes `queued` to
`update-status.json`, stages `self-update.sh` outside the repo and runs it.
3. **Apply** — stash if dirty, fetch the tag, check it out, `npm install
--include=dev`, `npm run build`. A failure at any step rolls back to the
previous commit, rebuilds it and reports `failed`; the server is never
restarted into a broken build.
4. **Restart** — write the terminal `restarting` marker, then signal the server.
The container exits and Docker restarts it.
5. **Reconcile** — the rebooted server compares its own version against the target
and flips the status to `completed` or `failed`. The browser, still polling,
picks that up.
Step 4 kills the updater script along with the container — unlike the systemd
path, it does not outlive the restart. That is safe only because the terminal
marker is written first, which is why nothing may be appended after the kill.
## Troubleshooting
**"This install can't update itself (unknown)"** — the repo bind mount is missing,
so the container is running the baked image copy. Check `CODEMAN_REPO_PATH` and
confirm the mounted directory really contains `.git`.
**The update fails immediately with a git ownership or permission error** — the
mounted checkout belongs to a different user than the one Codeman runs as
(`PUID`), so git refuses it as "dubious ownership". `Start-Codeman.sh` warns
about this at start; fix it by chowning the checkout to the same account that
owns `CODEMAN_APPDATA_PATH`.
**A rebuild is reported as required every time** — the fingerprint baseline does
not match the checkout. `Start-Codeman.sh` writes it on every start, so start
through that script rather than a bare `docker compose up` after either file
changes.
**Codeman does not come back after an update** — the build succeeded, since the
updater gates the restart on it, so read the container logs with `docker compose
logs codeman`. To roll back, check out the previous tag in the host checkout and
run `docker/Start-Codeman.sh`.
**The update failed during `npm install`** — most likely a native rebuild with no
toolchain, meaning the image predates the toolchain being added. Rebuild once from
the host and the in-app path works from then on.
**Resetting the build artefacts** — `docker compose down -v`, then
`Start-Codeman.sh`. This discards the named volumes and re-seeds them from a fresh
image.
## Disabling it
Set `CODEMAN_DISABLE_SELF_UPDATE=1` in `docker/.env` and pass it through in the
compose file's `environment:` block. The Updates panel then reports that in-app
updates are disabled, and the host-side script is the only way to update.
+31 -11
View File
@@ -228,6 +228,14 @@ window: the wait resolves on `idle` in a couple of seconds with `timedOut: false
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
⚠️ **Answering it is not "press Enter".** Claude Code 2.1.252 dropped the options'
numbers, reversed them, and highlights `No, exit` by default, so a blind `\r` quits
the CLI and the pane is dead seconds after the spawn. Read the `❯` marker off the
rendered pane (`GET .../terminal?full=1`), send `ESC [ B` while it sits on `No, exit`,
re-read, and confirm only once the marker is on `Yes, I trust this folder`. Codeman's
own auto-accept (`trustDialogNextKey()` in `src/session-trust-dialog.ts`) does exactly
this, inside a 90 s startup window and a 6-keystroke cap.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
@@ -244,24 +252,36 @@ SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
# Skip this and step 3 reports a turn that never ran. Match single tokens only:
# TUI text can arrive without its spaces. Stage 1 is short on purpose (an
# already-trusted case matches in <1 s; a first-run case can never pass it and
# pays it in full).
# ⚠️ NEVER answer the dialog with a bare \r. Its highlighted option is `No, exit`
# (claude-cli 2.1.252), so a blind Enter quits the CLI; and the dialog text stays
# in the buffer for the life of the session, so a `from=buffer` probe for `trust`
# keeps matching long after it is gone. Read the CURRENT pane instead and steer.
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
ESC=$(printf '\033') # \x1b is GNU-sed only; this form also works on macOS
for _ in 1 2 3 4 5 6; do
# Which option the ❯ marker sits on, read off the CURRENT frame.
K=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$ESC\[[0-9;?]*[a-zA-Z]//g" -e "s/$ESC[()][AB0]//g" | tr -d ' \t' \
| grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/')
[ -n "$K" ] || break # no dialog on screen: nothing to answer
[ "$K" = confirm ] && IN="\r" || IN="$ESC[B"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
-H 'Content-Type: application/json' \
-d "$(jq -nc --arg i "$IN" '{input:$i,useMux:true}')" >/dev/null
[ "$K" = confirm ] && break
sleep 1 # re-read: confirm the arrow landed
done
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
+106
View File
@@ -0,0 +1,106 @@
# Grok Build (xAI) integration plan
> **Status**: Executed. This document records the plan, the decision behind each wiring
> point, and what was and was not verified. The user-facing guide is
> [`grok-integration.md`](./grok-integration.md); the per-decision invariants live in
> [`architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok`](./architecture-invariants.md#external-cli-modes-opencode-codex-gemini-antigravity-pi-grok).
> Template: the pi integration (`c5b5963`, [`pi-integration-plan.md`](./pi-integration-plan.md)),
> which was itself calibrated against the four follow-up commits the antigravity
> integration needed. All of grok's facts below were verified against **grok 1.0.5**
> (`grok 1.0.5 (5115b46bc9)`), installed live during the work.
## 1. What Grok Build is
[xai-org/grok-build](https://github.com/xai-org/grok-build) is xAI's coding agent: a
Rust fullscreen-TUI binary named `grok`, installed by
`curl -fsSL https://x.ai/cli/install.sh | bash` into `~/.grok/bin` (with symlinks into
`~/.local/bin`; the installer also ships an `agent` alias). Config lives in
`~/.grok/config.toml`, TUI appearance in `~/.grok/pager.toml`, credentials in
`~/.grok/auth.json` (0600), sessions under `~/.grok/sessions/`. Auth is browser OAuth
on first launch, `grok login --device-auth` for SSH boxes, or `XAI_API_KEY` for
headless use. It has Claude-style permission modes (`default`/`acceptEdits`/`auto`/
`dontAsk`/`bypassPermissions`/`plan`), allow/deny rules, hooks, MCP, subagents, and a
headless `-p` mode.
## 2. Shape decisions (why grok is wired the way it is)
Grok is a seventh run mode, alongside Claude Code, shell, OpenCode, Codex, Gemini,
Antigravity and Pi. Never a location overlay, never a web tab. Its wiring mixes two
existing shapes:
| Question | Decision | Why |
| --- | --- | --- |
| Permission bypass | `GrokConfig.alwaysApprove` -> `--always-approve` | Grok's real flag (verified via `--help`): "Auto-approve all tool executions", i.e. its `bypassPermissions` mode. Config-level deny rules still apply on top. The Run button sends `true`, matching `runAntigravity()` and Claude's own `--dangerously-skip-permissions` default: Codeman sessions exist for autonomous work. |
| Multi-user clamp branch | only-if-sent (codex/antigravity branch) | A bare `grok` spawn is grok's own ask-mode default, which is already safe, so the clamp only needs to force a SENT `alwaysApprove` off. Contrast pi, whose absent default is an answerable prompt and therefore needs the materialize branch. Cron needs nothing for grok for the same reason (`clampCronExternalCliConfigs`). |
| Alt-screen strip | OUT of `isAltScreenStripMode()` | Grok is a fullscreen alternate-screen TUI with mouse support (its own scrollback pane, `pager.toml [terminal] alt_screen`), i.e. the opencode case, not the Ink repaint case. It falls through to the narrow tmux-attach strip like opencode/antigravity/pi. |
| Resolver | version probe, like pi | `grok` has npm squatters (the unrelated `@vibe-kit/grok-cli` installs a `grok` bin). Candidates must pass `grok --version`; `GROK_VERSION_REGEX` is exported and shared with the dependency registry so doctor and run mode cannot disagree. The probe cannot tell two version-printing `grok`s apart, so `GET /api/grok/status` surfaces path AND version. Search dirs: `~/.grok/bin` first (installer target), then `~/.local/bin`, `/usr/local/bin`, `~/bin`. |
| Env allowlist | `GROK_*` + `XAI_*` prefixes | `GROK_*` covers grok's documented inputs (`GROK_HOME`, `GROK_CONFIG`/`GROK_CONFIG_PATH`, `GROK_MEMORY`, `GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, `GROK_AUTH_PROVIDER_COMMAND`). `XAI_*` is xAI's vendor namespace and carries `XAI_API_KEY`, grok's documented headless auth var: the same narrow-vendor-namespace reasoning that admitted `GOOGLE_*` for gemini. Foreign provider keys stay out, as always. |
| Resume | `--resume <id>` / `--continue`, id-regexed | Grok's `--resume` also matches session TITLES (arbitrary user strings, case-insensitive). The `^[a-zA-Z0-9._-]+$` regex doubles as the no-titles rule, so nothing free-form can reach the `bash -c` spawn line. A valid explicit id wins over `-c`, mirroring pi. |
| Local echo | `'buffer'` via the `_updateLocalEchoState` fallthrough | UNMEASURED against an authenticated session (see §4). If grok's composer turns out per-keystroke reactive like codex's, the fallback is one `'off'` branch; teaching `PredictiveEchoAddon` grok's composer row is the larger follow-up. |
| Truecolor | `COLORTERM=truecolor` + `unset NO_COLOR` | Rust TUI with themes; joins the codex/gemini/antigravity/pi list in `buildEnvExports()` and `buildMuxAttachEnv()`. |
| Docker credentials | per-file seed: `auth.json`, `config.toml`, `pager.toml` | `~/.grok` also holds `sessions/`, `memory/`, `completions/`, `docs/` and the ~160MB binary under `downloads/`; a whole-dir seed would copy all of it on every container start. Same trade-off as pi: in-container sessions are invisible host-side, so `grok -c` in a Docker case sees only that container's history. |
| Docker install | own Dockerfile step | Not an npm package. xAI's installer has no `--dir` override, so the step copies `/root/.grok/bin/grok` (through the symlink, `cp -L`) into `/usr/local/bin` and removes root's `~/.grok` in the same layer. |
| Remote SSH | `exec "$SHELL" -i -l -c 'grok'` | sshd's remote-command PATH does not include `~/.grok/bin`; same login-shell fix as every other agent CLI. |
| What is NOT wired | `--permission-mode`, `--allow`/`--deny`, `-p` headless, `--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`, `--fork-session`, `--agent`, `--output-format` | Follow-ups. The flag surface is kept minimal on purpose; grok is pre-1.0-style fast-moving and every flag added is a flag validated forever. |
## 3. Touch points (the checklist)
Backend: `types/session.ts` (SessionMode + GrokConfig + SessionState), `utils/grok-cli-resolver.ts` (new)
+ barrel, `tmux-manager.ts` (`buildGrokCommand`, dispatch, resume flag, PATH export, truecolor,
availability error, plumbing), `session.ts` (external-mode gate, label, config plumbing,
tmux-required error, attach env), `mux-interface.ts`, `schemas.ts` (prefixes, `GrokConfigSchema`,
both mode enums, remote command overrides, cron agentType), `session-routes.ts` (clamp + both
create paths), `system-routes.ts` (`GET /api/grok/status`), `server.ts` (availability inject +
mux restore), `docker-hosts.ts`, `remote-hosts.ts`, `config/dependency-registry.ts`,
`cron/cron-service.ts` (comment), `response-viewer-transcript.ts`, `tui/tui-client.ts` + `tui-app.ts`.
Frontend: `index.html` (welcome button, run-mode entry, cron option, clone Brain option),
`session-ui.js` (`runGrok()`, dispatch, availability, "Run GK" label, external-CLI gates,
runMode setter), `app.js` (label, `gk` tab badge, kill-menu), `settings-ui.js`,
`mobile-overview.js`, `home-sessions.js`, `panels-ui.js`, `i18n.js`, `styles.css` +
`mobile.css` (charcoal monochrome identity; the non-og skin block and the mobile
`!important` pair are both load-bearing, see the pi plan's §2.9 cascade trap).
Meta: `docker/agent.Dockerfile`, `install.sh`, `package.json` keyword, changeset,
`skills/codeman/reference/*`, CLAUDE.md, READMEs, `architecture-invariants.md`,
`remote-sessions.md`, `security-architecture.md`, `docker-cases.md`, `cron-guide.md`.
Tests: `test/grok-mode.test.ts` + `test/grok-cli-resolver.test.ts` (new);
`external-cli-bypass-clamp`, `system-routes`, `render-index-html`, `run-mode-ui`,
`mobile-overview`, `local-echo-codex-gating` (extended).
## 4. Verification performed
On this box, with grok 1.0.5 really installed and an isolated
`CODEMAN_INSTANCE=grokwt` server (own data dir, own tmux socket, port 5077):
1. `npm test` (the CI gate): green, 5900+ tests. `typecheck`, `lint`, `format:check`,
`check:frontend-syntax`, `check:public-assets`, `check:lockfile`: green.
2. `GET /api/grok/status` -> `{available: true, path: "/home/arkon/.local/bin", version: "1.0.5"}`
through the real resolver and probe.
3. `POST /api/quick-start {mode: "grok", grokConfig: {alwaysApprove: true}}` -> session
created, tmux pane spawned, real spawn line verified to end in `grok --always-approve`,
and the actual grok TUI rendered its OAuth device-approval screen in the pane
(unauthenticated box, so sign-in is exactly where a first run lands).
4. `grokConfig` persisted into the instance's `state.json`.
5. Session deleted by exact id; instance data dir and throwaway case removed.
**Not verified (honest gaps, all requiring an xAI account or more hardware):**
an authenticated conversation end to end; the local-echo buffer policy against grok's
real composer (§2); scrollback/repaint behavior of the fullscreen TUI under the narrow
strip during a long session; a Docker case with `mode: 'grok'` (needs a `--no-cache`
agent-image rebuild); a remote-SSH grok case; cron readiness degradation (expected:
same slow-start-then-send as pi, documented in `cron-guide.md`).
## 5. Follow-ups
- Idle/completion signal: grok has a hooks system (user-guide `10-hooks.md`); a hook
POSTing to `/api/hook-event` could give grok sessions real idle detection instead of
output-stabilization. Highest-value follow-up, same slot as pi's `agent_settled` idea.
- Response viewer: sessions are ACP JSONL under `~/.grok/sessions/<encoded-cwd>/<id>/updates.jsonl`;
`grok -p ... --output-format json | jq -r '.sessionId'` exists for correlation.
- Permission-mode picker (`--permission-mode`, `--allow`/`--deny`) in Session Options.
- Measure the local-echo policy and the fullscreen-TUI scrollback behavior against an
authenticated session; pin the result in `local-echo-codex-gating` the way pi did.
- `grok doctor` is a built-in terminal-support check worth pointing users at when a
pane renders oddly.
+133
View File
@@ -0,0 +1,133 @@
# Grok Build (xAI) sessions
Codeman can drive [Grok Build](https://github.com/xai-org/grok-build) (xAI's `grok`
CLI, the agent behind docs.x.ai/build) as a session backend, alongside Claude Code,
OpenCode, Codex, Gemini, Antigravity and Pi. `grok` is a seventh **run mode**: its own
PTY, its own tmux session, its own tab identity (monochrome charcoal, `gk` badge). It
is not a location overlay like Docker or remote-SSH cases, and it is not a web tab.
The design rationale behind each decision below lives in
[`grok-integration-plan.md`](./grok-integration-plan.md). Everything here was verified
against grok 1.0.5.
## Install
```bash
curl -fsSL https://x.ai/cli/install.sh | bash
```
The installer places the binary in `~/.grok/bin` and symlinks it into `~/.local/bin`
(it also installs an `agent` alias Codeman ignores). `grok update` self-updates.
Codeman resolves the binary via the server PATH and then the usual install locations,
`~/.grok/bin` first. **`grok` is a name with known squatters** (the unrelated
`@vibe-kit/grok-cli` npm package also installs a `grok` bin), so like `pi` the
resolver does not trust a PATH hit on its own: it runs `grok --version` once and
requires version-shaped output (`grok 1.0.5 (5115b46bc9)`). Check what it resolved:
```bash
curl -s localhost:3000/api/grok/status | jq
# { "available": true, "path": "/home/you/.grok/bin", "version": "1.0.5" }
```
The endpoint carries `version` on top of the sibling `/api/*/status` shape precisely
so a misresolution is visible rather than presenting as "the mode just doesn't work".
## Authenticate
- **Browser OAuth (default)**: the first `grok` run opens a sign-in flow; in a
Codeman pane you get the device-code screen with a URL to open elsewhere.
Credentials land in `~/.grok/auth.json` (0600) and refresh automatically.
- **Device code**: `grok login --device-auth`, made for SSH boxes and headless hosts.
- **API key**: `export XAI_API_KEY="xai-..."` (console.x.ai). Used as a fallback when
no session token exists. As a per-session Codeman `envOverride` it flows through
socket-scoped `tmux setenv`, never the spawn command line.
- **Enterprise OIDC**: `GROK_OIDC_ISSUER` / `GROK_OIDC_CLIENT_ID`.
## What Codeman wires up
`GrokConfig` (per session, persisted in `state.json`, round-trips through respawn):
| Field | Flag | Notes |
| ----------------- | --------------------------- | --------------------------------------------------------------------- |
| `model` | `--model <v>` | e.g. `grok-4.5`, or a custom `[model.<name>]` from `config.toml` |
| `alwaysApprove` | `--always-approve` | Grok's `bypassPermissions` mode; deny rules still apply on top |
| `continueSession` | `--continue` | Most recent session for the working directory; skipped when resuming |
| `resumeSessionId` | `--resume <v>` | Ids only, never titles (grok's own `--resume` also matches titles) |
Every value is regex-validated and **dropped** (not escaped) if it fails, because the
result is interpolated into the pane's `bash -c "..."` command.
The Run button sends `grokConfig: { alwaysApprove: true }`, the same product decision
as Claude's `--dangerously-skip-permissions` default and Antigravity's
`--dangerously-skip-permissions`: Codeman sessions exist for autonomous work. Keep
hard limits as `deny` rules in `~/.grok/config.toml` (they apply in every mode), and
in **multi-user mode** a non-granted owner's `alwaysApprove` is forced off
server-side; a bare `grok` spawn is grok's own ask-mode default.
Env overrides: the `GROK_*` prefix (`GROK_HOME`, `GROK_CONFIG`, `GROK_MEMORY`,
`GROK_WORKFLOWS`, `GROK_SANDBOX`, `GROK_OIDC_*`, ...) plus the `XAI_*` vendor
namespace (`XAI_API_KEY`) are allowlisted. Foreign provider keys are not, as ever.
## What Codeman deliberately does NOT wire up
- **`--permission-mode`, `--allow`/`--deny`.** The boolean covers the autonomous
case; the full rule surface is a follow-up with UI.
- **`-p`/headless, `--output-format`, `--json-schema`.** Codeman drives the TUI.
- **`--worktree`, `--sandbox`, `--reasoning-effort`, `-s/--session-id`,
`--fork-session`, `--agent`/`--agents`.** Tracked as follow-ups in the plan doc.
## Terminal behavior
Grok renders a **fullscreen alternate-screen TUI** (scrollback pane + prompt, mouse
supported). Under Codeman it runs inside tmux like every external CLI, so the
fullscreen rendering stays inside the pane and the browser terminal shows tmux's
repaints; grok stays out of the alt-screen strip list on purpose (the opencode case,
not the Ink case). If a pane renders oddly, `grok doctor` checks terminal, color and
input support without starting a session, and `~/.grok/pager.toml` can force
`alt_screen = "inline"`.
On touch devices grok currently gets the buffered local-echo overlay like Claude,
Gemini, OpenCode and Pi. This is the fallthrough default and has not been measured
against an authenticated grok composer; if grok turns out per-keystroke reactive the
way codex was (issues #218/#219/#220/#222), the fix is the `'off'` branch in
`_updateLocalEchoState` (terminal-ui.js).
## Docker cases
The agent image installs grok in its own Dockerfile step (not npm; xAI's installer
targets `$HOME/.grok/bin` with no `--dir` override, so the binary is copied to
`/usr/local/bin`). Rebuild with the mandatory `--no-cache`:
```bash
node scripts/build-agent-image.mjs --no-cache
```
Credentials are **seeded**, not shared: `auth.json`, `config.toml` and `pager.toml`
are copied into the container's own `~/.grok`, so an in-container grok never writes
refreshed OAuth tokens back to the host and `docker commit` exports stay secret-free.
Only those three files, because `~/.grok` also holds `sessions/`, `memory/` and the
~160MB binary under `downloads/`. Trade-off, same as pi: in-container sessions are
invisible host-side, so `grok -c` inside a Docker case only sees that container's own
history.
## Remote SSH cases
`grok` mode is routed through an interactive login shell
(`exec "$SHELL" -i -l -c 'grok'`), because sshd's remote-command PATH does not include
`~/.grok/bin`. Per-session config and `envOverrides` do not cross ssh and are rejected
rather than silently ignored; use the per-host command override instead. For auth on
the remote host, `grok login --device-auth` exists for exactly this.
## Known gaps
- **No idle/completion hook yet.** Idle detection falls back to output-stabilization
like the other external CLIs. Grok has a hooks system, so a Codeman hook POSTing to
`/api/hook-event` is the highest-value follow-up.
- **No response viewer.** Grok writes ACP JSONL sessions under
`~/.grok/sessions/<encoded-cwd>/<session-id>/updates.jsonl`; nothing reads them yet.
- **Cron jobs mis-detect readiness.** The readiness poll looks for `❯` or a token
count, neither of which grok prints, so a grok cron job burns its poll budget and
then sends the prompt anyway. It works; it is just slower to start.
- **Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe are
off** for grok, as for every external CLI.
+171
View File
@@ -0,0 +1,171 @@
# OMP (Oh My Pi) sessions
Codeman can drive [OMP](https://github.com/can1357/oh-my-pi) (`omp`, Oh My Pi) as a session
backend, alongside Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi, Grok and
DeepSeek Harness. `omp` is the ninth CLI backend (tenth `SessionMode`, counting
`shell`): its own PTY, its own tmux session, its own tab identity. It is not a
location overlay like Docker or remote-SSH cases, and it is not a web tab.
## Install
```bash
curl -fsSL https://omp.sh/install | sh
```
The installer places the binary in `~/.local/bin` (verified against a real
`--no-cache` Docker build — see `docker/agent.Dockerfile`; an earlier guess of
`~/.omp/bin` was wrong). Codeman resolves the binary via the server PATH and then
the usual install locations (`~/.local/bin` first, then `~/.omp/bin`,
`/usr/local/bin`, `~/.bun/bin`, `~/.npm-global/bin`, `~/bin`).
**`omp` is a short name**, so like `pi` and `grok` the resolver does not trust a PATH
hit on its own: it runs `omp --version` and requires `omp/<semver>`-shaped output
(e.g. `omp/18.0.8`) before accepting a candidate. Check what it resolved:
```bash
curl -s localhost:3000/api/omp/status | jq
# { "available": true, "path": "/home/you/.local/bin", "version": "18.0.8" }
```
## Authenticate
OMP owns its own auth and provider configuration entirely in `~/.omp` — there is
no Codeman-side login flow, API key field, or bypass switch to configure. Run `omp`
directly once outside Codeman to complete whatever onboarding the CLI itself asks
for; every session started through Codeman afterward inherits that config.
## What Codeman wires up
`OmpConfig` (per session, persisted in `state.json`, round-trips through respawn):
| Field | Flag | Notes |
| ------------------ | --------------- | ---------------------------------------------------------- |
| `model` | `--model <v>` | Regex-validated (`[a-zA-Z0-9._-/]+`); `provider/model` forms like `crof/glm-5.2` pass |
| `continueSession` | `--continue` | omp's own "most recent conversation in this directory" heuristic |
| `resumeSessionId` | `--resume <id>` | Ids only, id-regexed; wins over `--continue` when both are present |
Every value is regex-validated and **dropped** (not escaped) if it fails, because the
result is interpolated into the pane's spawn command.
**omp reads its own model routing and hooks from `~/.omp`, so no trust or
permission flags are needed** — unlike every sibling CLI in this family, there is no
bypass-permissions equivalent to wire up, so `buildOmpCommand()` only ever passes
`--model`/`--resume`/`--continue`. ⚠️ That does NOT mean omp is unrestricted: its
documented default `tools.approvalMode` is `yolo`, so an omp pane auto-approves exec
with no flag from Codeman — the CLI's own config, not Codeman, is what would need to
change that.
Env overrides: the `OMP_*` prefix is allowlisted, and per omp's own
`docs/environment-variables.md` it is not the narrow surface it looks like. omp reads
roughly 40 provider keys from the environment (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`XAI_API_KEY`, `HF_TOKEN`, ...) — pi's 34-key problem in the same shape — which is why
none of those get a dedicated allowlist entry; a session authenticates from `~/.omp`
config or the server process's own env instead, like pi. omp's own documented knobs
are mostly `PI_*`, not `OMP_*` (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`,
`PI_CODING_AGENT_SESSION_DIR`, `PI_SUBPROCESS_CMD`, `PI_SHELL_PREFIX`,
`OMP_PROFILE`/`PI_PROFILE`), and `PI_*` is already allowlisted globally because pi
mode needs it — so an omp session today already accepts all of those. The first three
also move the tree `omp-session-resolver.ts` and `omp-transcript.ts` hardcode
(`resolveOmpHome()` assumes `~/.omp` unconditionally), so pinning and history quietly
stop working under a redirected config root; this is a known gap, not fixed here.
The `OMP_` prefix itself brings in `OMP_AUTH_BROKER_URL` / `OMP_AUTH_BROKER_TOKEN`,
where omp resolves credentials from — the same shape `DEEPSEEK_BASE_URL` is dropped
for in `clampEnvOverridesForOwner()` (session-routes.ts), so both are clamped there
for a non-granted owner in multi-user mode. None of this matters in single-user mode.
## Exact-id pinning: why `--resume`, not just `--continue`
`--continue` alone is ambiguous the moment **any** other omp conversation has
touched the same working directory more recently — it just picks the newest session
file on disk, silently. That happens routinely: a closed-then-resumed Codeman row
plus a still-running duplicate, two Codeman sessions pointed at the same case, or a
plain reattach after a server restart.
`src/utils/omp-session-resolver.ts` resolves and **pins** the exact conversation id
once (`findLatestOmpSessionId()` reads `~/.omp/agent/sessions/<mangled-workingDir>/`,
the newest `.jsonl` file's embedded uuid), then every later respawn reuses that
pinned id via `--resume` instead of re-guessing with `--continue`.
⚠️ **The directory mangling is NOT a straight `/` → `-` replace.** Unlike Claude
Code's `~/.claude/projects/*` convention (which keeps the full path, e.g.
`-home-user-codeman-cases-foo`), omp strips the `$HOME` prefix FIRST and only then
dash-replaces (`/home/user/codeman-cases/foo` → `-codeman-cases-foo`; a path outside
`$HOME`, like `/tmp/...`, is dash-replaced as-is with no stripping). Getting this
wrong doesn't error — `findLatestOmpSessionId()` just silently returns null for
every case under `$HOME` (virtually all real Codeman cases), so pinning quietly
degrades to omp's own ambiguous `--continue`. This was found and fixed 2026-08-27
after months of testing had only ever exercised `/tmp`-based working directories,
where the bug's wrong output happened to coincidentally match the right one.
## Surviving a full session kill
`src/omp-transcript.ts` scans `~/.omp/agent/sessions/**/*.jsonl` directly — a second,
independent history source alongside Codeman's own state. This means an OMP
conversation's history (working directory, first/last prompt, size) is recoverable
in the Past Sessions list even when **both** the Codeman session record and the
underlying tmux pane are gone — verified live against a full OS reboot, not just a
"Kill Tmux" button click.
## Terminal behavior
OMP renders inside tmux like every external CLI (narrow scrollback strip — alt-screen
toggles only, not the full Claude/Codex/Gemini strip). It stays out of the
alt-screen-strip list and lands on the `'buffer'` local-echo policy via the
`_updateLocalEchoState` fallthrough, same as grok and pi.
## Docker cases
The agent image installs omp in its own Dockerfile step (not npm; omp's installer
targets `$HOME/.local/bin` with no `--dir` override, the same shape as grok's
installer). Rebuild with the mandatory `--no-cache`:
```bash
node scripts/build-agent-image.mjs --no-cache
```
⚠️ **`--resume` pinning does not currently reach an in-container omp process.**
Docker panes are built from `defaultDockerCommandForMode`, which never sees
`ompConfig` — `appendResumeFlag()`'s `case 'omp'` keys off the top-level
`resumeSessionId` field, which nothing populates for omp today. Host-side history
recovery still works (the shared `sessions/` mount below), but a respawned
in-container omp pane falls back to its own ambiguous `--continue`, not a pinned
id. Flagged in upstream review, not yet fixed.
Credentials are **mostly seeded**, but `sessions/` is the one exception in this CLI
family: `~/.omp/agent/{config.yml,mcp.json,models.yml,settings.yml}` are seeded
(read-only mount, copied into the container's own `~/.omp/agent` once), so an
in-container omp never writes refreshed config back to the host and `docker commit`
exports stay secret-free. But `~/.omp/agent/sessions/` is **shared (RW)**, not
seeded — the same treatment as codex's `sessions/`, and for the identical reason:
Codeman reads it host-side (`omp-transcript.ts`, `omp-session-resolver.ts`) for
history recovery and `--resume` pinning. Seeding it instead of sharing it would make
an in-container OMP conversation invisible to Codeman's own history/resume logic,
silently breaking Docker support for the kill-survival feature above. The rest of
`~/.omp/agent` (`agent.db`/`history.db`/`models.db` SQLite caches,
`terminal-sessions/`, `blobs/`, `cache/`) stays container-local and is neither
shared nor seeded.
## Remote SSH cases
`omp` mode is routed through an interactive login shell
(`exec "$SHELL" -i -l -c 'omp'`), because sshd's remote-command PATH does not
include `~/.local/bin`. Per-session config and `envOverrides` do not cross ssh and are
rejected rather than silently ignored; use the per-host command override instead.
## Known gaps
- **No idle/completion hook.** Idle detection falls back to output-stabilization
like every other external CLI. If omp ever ships a hooks system, a Codeman hook
POSTing to `/api/hook-event` would be the highest-value follow-up.
- **Killing a pane mid-turn loses the conversation for real.** `tmux kill-session`
before an in-TUI `/exit` beats omp's own session-file flush — confirmed by direct
testing (kill after a clean `/exit` resumes correctly; kill without `/exit` first
does not). This is not something Codeman can compensate for from outside the
process; it would need an upstream omp fix (e.g. flush-on-SIGTERM).
- **Unverified: `$HOME` as a symlink.** The directory-mangling fix above compares
against the literal `homedir()` string, not a `realpath()`-resolved one. Whether
omp itself canonicalizes symlinks before mangling is unconfirmed — this has not
been tested against a symlinked-home setup.
- Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe are
off for omp, as for every external CLI.
+1 -1
View File
@@ -139,7 +139,7 @@ set -g extended-keys-format csi-u
Codeman's browser input path sends `\r` for submit, so basic use works
unconfigured — what degrades is newline-in-editor, mostly when you attach to the
pane directly (`sc`).
pane directly (`codeman tui`).
⚠️ Upstream notes the setting may need a full `tmux kill-server` to take effect.
**Never run `tmux kill-server` on Codeman's socket** — it would kill every live
+2 -2
View File
@@ -1,7 +1,7 @@
# Remote Sessions (SSH)
Codeman can run a session's agent on a **remote host over SSH** instead of the
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, Pi, or a plain shell)
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, or a plain shell)
runs inside a `tmux` server **on the remote host**, so it survives the SSH
connection dropping; Codeman attaches to it the same way it attaches to a local
managed session.
@@ -30,7 +30,7 @@ Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi'>` — the modes that can run remotely. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi' \| 'grok'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
+2 -2
View File
@@ -489,7 +489,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, and five seeded files from `~/.pi/agent`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, five seeded files from `~/.pi/agent`, and three from `~/.grok`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
@@ -518,7 +518,7 @@ A saved dashboard URL renders as a tab, served through Codeman's own origin at `
- **The proxy is exempt from cookie auth and the Origin/CSRF guard, and that is deliberate.** The iframe is sandboxed without `allow-same-origin`, so it is opaque‑origin: its requests are cross‑site, meaning the `SameSite=lax` session cookie is never attached and its writes and WS upgrades arrive with `Origin: null`. The credential is instead a 192‑bit capability in the path, minted only by an authenticated `POST /api/webviews/:id/open`, held in memory (a restart invalidates every one), rolling TTL, bound to the minting user, and granting nothing but "relay bytes to this one saved URL". ⚠️ **The Host allowlist is NOT bypassed**, so DNS‑rebinding protection is unaffected. A second `Referer`‑keyed form exists for root‑absolute assets and is the only exemption decided by a request‑supplied header, so it is fenced to safe methods on non‑`/api`, non‑`/ws`, non‑`/q` paths. Edges pinned by `test/webview-auth-exemption.test.ts`.
- **Sandboxed by default; `allow-same-origin` is an explicit per‑dashboard opt‑in.** A proxied page is same‑origin with Codeman, so without the sandbox its JavaScript could read the Codeman document and call the agent‑spawning API. ⚠️ In BOTH modes the `Authorization` header and the `codeman_session` cookie are stripped before the upstream request, because a trusted (same‑origin) frame makes the browser attach Codeman's own Basic‑auth credentials to every proxied request; forwarding them would hand `CODEMAN_PASSWORD` to the dashboard.
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from).
- **Not an open relay, and not a privilege boundary.** `resolveUpstreamUrl()` refuses anything leaving the saved origin, and cross‑origin redirects are handed back unchanged rather than followed. The proxy does reach whatever the SERVER can reach, which is not an escalation for someone who already commands `--dangerously-skip-permissions` agents, but in multi‑user mode it means a non‑admin's dashboard is fetched from the server's network position. Saved URLs are validated to plain http(s) with no embedded credentials, and there is deliberately **no magic‑link path**: terminal output can never create a webview (the mistake the attachment scanner had to be walled off from). The one refused destination class is link‑local and cloud‑metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`, `168.63.129.16`, `100.100.100.200`, `metadata.google.internal`): `webview-egress-policy.ts` refuses them at save time, and `webview-egress.ts` re‑judges the RESOLVED address at connect time through a `lookup` hook on the proxy's undici Agent and on its WebSocket client, so a DNS name pointing into those ranges is refused as well. Loopback and RFC1918 stay allowed on purpose. Capabilities are revoked on logout, admin logout and user deletion, and proxied responses carry `Referrer-Policy: same-origin` so a dashboard cannot hand the capability‑bearing URL to a third‑party host it links.
---
+219
View File
@@ -0,0 +1,219 @@
# Codeman TUI Rework Plan
Status: **phases 0-2 implemented** on `feat/tui`; phases 3-4 remain follow-ups. The user guide is [`docs/tui.md`](tui.md); this document stays the design record.
- Phase 0: `src/cli-style.ts` (palette, glyphs, `heading`/`kv`/`table`/`spinner`/`confirm`) plus the mechanical fixes of §5, and `test/cli-commands.test.ts` now derives its inventory from the real commander `program` instead of parsing a fixture.
- Phases 1-2: `src/tui/`. `tui-app.ts` (main loop, attach handoff, verbs) and `tui-client.ts` (API, SSE, degraded enumeration) are the only IO; `tui-model`, `tui-layout`, `tui-render`, `tui-keys`, `tui-ansi`, `tui-composer`, `tui-approvals`, `tui-digest`, `tui-sse` and `tui-types` are pure and unit-tested, with an E2E suite driving the real binary under node-pty.
- Deferred with the rest of phase 3: `r` (resume a RECENT row) is not wired up, so the help overlay does not advertise it.
- Not started: phase 3 (mouse, `--pick` popup switcher, opt-in attach status line, OSC 9) and phase 4 (retiring the bash choosers).
The goal: replace Codeman's scattered terminal surfaces with one first-class TUI, `codeman tui`, that gives SSH/terminal users the same at-a-glance awareness the web UI gives browsers. The reference point is herdr (herdr.dev), the trending Rust "agent multiplexer" whose defining feature is a live agent-state sidebar. Codeman can match and beat that sidebar in the terminal because the states herdr infers from screen-scraping heuristics are states our server already computes from hooks, pane probing, and the approvals inbox.
---
## 1. What we have today (inventory)
Three disconnected surfaces, three visual idioms, two data sources:
| Surface | What it is | Data source | Idiom |
| --- | --- | --- | --- |
| `codeman` CLI (`src/cli.ts`, 1214 lines) | commander + chalk, ~20 commands | HTTP API + state files | `✓`/`✗` line-per-fact, no interactivity |
| `sc` (`scripts/tmux-chooser.sh`, 663 lines) | bash number-menu chooser, mobile-tuned (44 cols) | `tmux -L codeman` + `state.json` via jq | 256-color, numbered, full repaint per key |
| `scripts/tmux-manager.sh` (529 lines) | bash cursor TUI with kill/info | `mux-sessions.json` (and writes it back) | 8-color, box-drawn, arrow keys |
Weaknesses found in the audit (file:line refs verified 2026-08-16):
1. **No interactive picker in the Node CLI at all.** Every `session stop`, `task status`, `session logs` requires a pasted UUID prefix. There is no `codeman attach <session>`; `codeman attach` is actually the attachment-card command (and `README.md:895` describes it wrongly).
2. **`sc` cannot reach sessions 10+ interactively**: entries are numbered globally (`tmux-chooser.sh:343`) but input accepts a single `[1-9]` keypress (`:487-493`). Page 2 shows items 8-14 that mostly cannot be selected.
3. **No cursor/selection concept in `sc`** (`BG_SEL` at `:90` is dead code); arrows only page.
4. The two bash tools can disagree about which sessions exist (different data files), and only `sc` is on PATH.
5. **Zero live feedback anywhere**: `codeman web -d` and `service install` block silently up to 30s (`daemon-control.ts:395-412`); no spinner exists in the codebase.
6. Styling drift: `doctor` is the only table and is deliberately monochrome with a colorize hook nobody wired up (`dependency-report.ts:5-7`); `codeman web` prints its "running at" line twice (colored `cli.ts:934`, plain `server.ts:2366`); the server's security warning is colorless `console.warn` while the CLI's version of the same warning is yellow; `tmux-manager.sh`'s header box is visibly misaligned; `padEnd(14)` overflows on "Antigravity CLI".
7. Bash TUIs emit raw escapes unconditionally (no TTY/NO_COLOR gate); `install.sh` and `postinstall.js` do it right.
8. Detach hint inconsistency: chooser says Ctrl+B D, `README.md:671` says Ctrl+A D.
9. Inside an attached session there is **no chrome at all**: Codeman turns the tmux status bar off (`tmux-manager.ts:1978`), so an SSH user in a pane has no session identity, no state, no way back to a picker except detach.
10. `test/cli-commands.test.ts` asserts against a hand-written fixture, not the real `program`, and that fixture already lists a `tui` command that does not exist (`:57-61`). The name is pre-approved by our own test file.
## 2. Research: how herdr does it
herdr (github.com/herdrdev/herdr, ~30k stars, single Rust binary, pre-1.0) is a background terminal multiplexer "your coding agents live on". What matters for us:
- **The agent-state sidebar is the product.** Every pane is classified live as `working` / `blocked` / `done` / `idle` and grouped in a sidebar, so you see who needs you without switching tabs. Reviews unanimously call this "the killer feature tmux can't match".
- **Detection is heuristic-first**: process-name matching + screen-manifest TOML rules parsing the visible frame; optional per-agent "integration install" adds lifecycle hooks over JSON-RPC on a unix socket for accurate states. Claude Code there is on the heuristic path and reviewers note blocked-state lag.
- **Model**: workspaces → tabs → panes, tmux-style prefix keys (Ctrl+B V split, arrows navigate, D detach), mouse-first (click select, drag resize, right-click menus, touch over SSH), adapts to narrow widths.
- **Agent-shaped API**: socket API with `pane read` (visible/recent/detection), `send-text`/`send-keys`/`run`, `agent start|prompt|wait|explain`, `pane wait-output` with regex, plugins placed as overlay/split/tab/popup.
- **Persistence**: sessions survive disconnects, reattach from any terminal / SSH.
- Weaknesses reviewers cite: pre-1.0 churn, bus factor 1, no session resurrection, rendering lag with many panes.
What is striking is how much of herdr Codeman already has, server-side: our hooks give exact `permission_prompt`/`stop`/`idle_prompt` events (herdr's "integration" path, but installed by default), `_confirmIdle()` does the screen-probe fallback, the approvals inbox parses the actual dialog options, and the agent skill + wait primitives are our socket API. What we lack is purely the presentation layer in the terminal.
Prior art for the architecture we want: **agent-deck** (Bubble Tea + tmux) proves the "TUI list + attach into tmux" model works great: session list with live glyphs (● ◐ ○ ✕), Enter attaches into a tmux pane, status polling, groups, fuzzy search. We take the shape, not the code.
Licensing note: herdr is reported variously as Apache-2.0/AGPL-3.0. Irrelevant either way: we copy concepts, never code.
### What we take / what we skip
Take: the four-state sidebar as the organizing principle; grouping by "needs you first"; narrow-width adaptation; mouse support; tmux-familiar keys; the "attention at a glance" framing.
Skip: being a multiplexer. tmux already backs every Codeman session and is a hard dependency; herdr had to build pane management because it owns terminals, we do not. Also skip (for now): plugin marketplace, split layouts, pane drag. Our TUI is a **dashboard + switchboard over tmux**, not a tmux replacement.
## 3. Design: `codeman tui`
One command, one full-screen client of the existing HTTP/SSE API.
**Positioning (owner decision, 2026-08-16): the web UI remains THE primary surface.** The TUI is strictly additive, for users who want a terminal workflow (SSH, Termius, tmux die-hards). Bare `codeman` keeps printing help; nothing existing changes behavior. The `sc` bash chooser also stays untouched for now; flipping its alias to `codeman tui` is deferred to a follow-up release once the TUI has mileage.
### Layout (≥100 cols)
```
codeman tnode · v1.19.0 · 6 sessions · 5h ▂▂▅ 32% wk 61% ? help q quit
────────────────────────────────────────────────────────────────────────────────────────────
NEEDS YOU ──────────────────────────┐ ┌ w4-api-refactor ── claude · ~/dev/api ────────────
▶ 1 w4-api-refactor ⚠ approval 2m │ │ ✻ Actualizing… (2m 14s · ↓ 12.3k tokens)
2 w6-docs ✋ waiting 11m │ │
│ │ ⚠ Claude requests: Bash(git push origin main)
WORKING ────────────────────────────┤ │ 1. Yes 2. Yes, don't ask again 3. No
3 w1-codeman ✻ 17m 45.2k │ │
4 w2-gallery ✻ 3m 8.1k │ │ [y] approve [n] deny [Enter] attach
IDLE ───────────────────────────────┤ │
5 w3-promo ○ 2h │ │ …live tail of the selected session's
RECENT ─────────────────────────────┤ │ terminal (ANSI colors preserved),
· api-hotfix ✔ done Fri │ │ updating while you browse the list…
────────────────────────────────────────────────────────────────────────────────────────────
↑↓ select · ⏎ attach · 1-9 jump · y/n answer · p prompt · n new · x kill · / search · g digest
```
- **Header**: hostname/instance, server version, session count, plan-usage chip (same telemetry that feeds the web chip, when available). Degrades gracefully when the server is down (see §3.6).
- **Sidebar**: sessions grouped `NEEDS YOU` → `WORKING` → `IDLE` → `RECENT` (past sessions from the unified list, resumable). Within groups, reuse the activity ordering already built for the home screens in PR #303 (blocked first, running longest, quiet newest); that logic is pure and shared.
- **Preview pane**: live tail of the selected session, SGR colors preserved, cursor-movement stripped. When the selected session has a pending approval, the parsed dialog is rendered as a card above the tail with one-key answer bindings.
- **Footer**: contextual keymap (changes when a dialog/confirm is active).
### States and vocabulary
Exactly the web's language so the two surfaces read the same:
| Group | Glyph | Color | Source |
| --- | --- | --- | --- |
| NEEDS YOU (question/permission) | `⚠` | red, blinking row | approvals inbox / `permission_prompt` |
| NEEDS YOU (waiting for input) | `✋` | yellow | `idle_prompt` / waiting classification |
| WORKING | `✻` animating through `· ✢ ✳ ∗ ✻ ✽` at 2Hz | green | working classification (the same glyph family Claude itself draws, a deliberate nod) |
| IDLE | `○` | muted | idle |
| RECENT / done | `✔` | muted green | unified list history rows |
Nerd-font/glyph fallback exactly like `sc` does today (`[!] [w] [*] [-] [ok]` when the terminal is not known-capable), plus full NO_COLOR / `tput colors` degradation (8-color and mono renderings are designed, not accidental).
### Keymap
- `↑/↓` or `j/k` select · `Enter` attach · `1-9` jump-attach (parity with `sc`, but now the cursor covers 10+)
- `y`/`n` (or the digit keys) answer the selected session's pending approval right from the dashboard, via `POST /api/approvals/:id/answer`. The server already re-captures the pane and 409s if the dialog is gone, so this is safe by construction.
- `p` send a one-line prompt to the selected session without attaching (`POST /input` with `\r`, the composer opens in the footer)
- `n` new session (case picker → mode picker, drives `POST /api/quick-start`) · `x` kill with typed confirm (never bulk; refuses the session hosting the TUI itself, like tmux-manager.sh does)
- `/` fuzzy search across sessions/history/attachments (`GET /api/search`) · `g` away digest (`GET /api/away-digest`) rendered as a panel
- `r` resume selected RECENT row (unified list `resume-session` flow) · `?` help overlay · `q` quit
- Mouse (phase 3): SGR mouse reporting, click selects, wheel scrolls list/preview, click on footer keys triggers them. Works over SSH, same as herdr's touch story.
### Responsive behavior
The `sc` design constraint survives: below ~72 cols (Termius, iPhone portrait) the preview pane drops and the TUI is a single-column list with two-line rows, nearly identical to today's `sc` but with a cursor, live states, and the answer/prompt/new/kill verbs. The layout switch is width-driven at draw time, no mode flag.
### Attach model
Enter suspends the TUI (restore main screen + cooked mode), then hands the terminal to `tmux -L <socket> attach-session -t <name>` with `stdio: inherit`. On tmux exit/detach, the TUI resumes and refreshes. Full fidelity (mouse, paste, colors) is tmux's, we never proxy bytes.
- Inside tmux already: same socket → `switch-client -t`; different socket → warn about nesting and offer detach-first. `$TMUX` + `CODEMAN_MUX` detection.
- **Return path**: a tmux binding installed for codeman sessions (opt-in) runs `codeman tui --pick` inside `tmux display-popup -E`, a minimal picker-only mode (list + jump, no preview) so switching sessions from inside a pane is one keystroke, fzf-style.
- Optional per-attach chrome (opt-in setting, default off since `status off` at `tmux-manager.ts:1978` is deliberate): a minimal codeman-styled tmux status line showing `name · state · alert`, set on attach, restored on detach.
### Notifications
While the TUI is open and a session flips to NEEDS YOU: flash the row, ring BEL, and optionally emit OSC 9 (desktop notification in kitty/WezTerm/iTerm2, and it traverses SSH). This is the herdr sidebar promise delivered even when the terminal is backgrounded.
### Degraded mode (server down)
`sc` works without the server today and the TUI must too: when no server answers, enumerate `tmux -L codeman list-sessions` + read `state.json` (read-only), show a "server not running" header line, and offer attach only (no states, no approvals). This keeps the "web server crashed, get me to my sessions" path alive.
## 4. Architecture
### A client of the server, not a second brain
Everything live comes from the API the web UI already uses:
| Need | Endpoint |
| --- | --- |
| Session list + history | `GET /api/sessions/unified` |
| Live updates | SSE `GET /api/events` (heartbeat `sse:heartbeat` already exists; fall back to 2s polling) |
| Pending approvals + parsed options | `GET /api/approvals`, answer via `POST /api/approvals/:id/answer` |
| Preview tail | `GET /api/sessions/:id/terminal?tail=N` (throttled to the selected session only) |
| Prompt send | `POST /api/sessions/:id/input` (single line + `\r`, per the composer contract) |
| New session | `POST /api/quick-start` (routes remote/docker cases correctly) |
| Search | `GET /api/search` |
| Away digest | `GET /api/away-digest` |
| Plan usage chip | latest status-telemetry snapshot (`plan-usage-latest`) |
Server discovery and auth reuse what exists: instance config from `src/config/instance.ts` (`CODEMAN_INSTANCE`, `CODEMAN_PORT`), the probe logic from `daemon-control.ts`, credentials from `~/.codeman/.env` (the established `codeman attach` pattern), self-signed HTTPS accepted for loopback probes (the hooks-on-HTTPS lesson). Multi-user scoping comes free: the API only returns what the authenticated user owns.
### Renderer: hand-rolled, zero new dependencies (decision)
Options considered:
- **Ink (React for CLIs)**: what Claude Code uses. Pros: layout engine, ecosystem. Cons: pulls React into a CLI that today ships only commander+chalk; rerender model fights the two things we care most about (a raw-ANSI preview region and 2Hz glyph animation without flicker); version-pins React for every `npm i -g aicodeman`.
- **blessed/neo-blessed**: unmaintained, skip.
- **Hand-rolled screen core** (recommended): this repo hand-rolls ANSI everywhere already and has the expertise (regex-patterns, stripAnsi, the xterm work). The core is small and boring: alt screen + raw mode + cursor-home full-frame repaint from an off-screen string buffer, throttled to state changes and the 2Hz animation tick, wrapped in DECSET 2026 (synchronized output) where supported so repaints are atomic in modern terminals (tmux, kitty, WezTerm, iTerm2). No diffing needed at these frame rates.
The one genuinely tricky pure function: SGR-aware line clipping for the preview (keep colors, strip cursor movement/OSC/DECSET, clip to width while carrying SGR state, reset at EOL). That is a pure module with exhaustive unit tests, and it is exactly the kind of function Ink would not have given us anyway.
### Module layout
```
src/tui/
tui-app.ts entry + main loop + attach handoff (IO)
tui-client.ts API + SSE client, degraded-mode enumeration (IO)
tui-model.ts pure: state store, grouping, ordering (reuses PR #303 helpers)
tui-layout.ts pure: responsive layout math, row building
tui-render.ts pure: model+layout -> frame string (palette, glyphs, fallbacks)
tui-keys.ts pure: byte stream -> key/mouse events (incl. SGR mouse decode)
tui-ansi.ts pure: SGR-aware clip/filter for the preview
```
Pure modules unit-test with no TTY. `cli.ts` gains one thin `tui` command registration (and `--list`/`<n>` fast paths for `sc -l` / `sc 2` parity, which must stay fast: they short-circuit before any screen setup).
## 5. CLI-wide polish (the rest of "make it much nicer")
A shared style kit, `src/cli-style.ts`: one palette (mirroring the web's status colors), one glyph set with fallback, `heading()`, `kv()`, `table()` (width-aware, fixes the Antigravity overflow), `spinner()` (finally: the 30s silent daemon/service waits get a live line), `confirm()` (used by `reset --force`'s missing prompt and `x` in the TUI). Then the mechanical fixes from §1: colorize `doctor` through the hook that already exists for it, dedupe the `codeman web` startup line, colorize the server's security warning, fix the README `codeman attach` description and the Ctrl+B/Ctrl+A detach drift, TTY/NO_COLOR gates everywhere.
## 6. Phasing
| Phase | Contents | Size |
| --- | --- | --- |
| 0 | `cli-style.ts` + mechanical fixes (§5), real CLI tests (retire the fixture parser in `test/cli-commands.test.ts`) | S |
| 1 | `codeman tui` core: list + states via SSE, cursor + 1-9, attach/return loop, kill w/ confirm, new session, narrow mode, degraded mode, `sc` alias flip + `--list`/`<n>` parity | M/L |
| 2 | Preview pane (SGR clip), approvals answering, prompt composer, search, digest, resume, plan-usage header | M |
| 3 | Mouse support, `--pick` popup switcher + tmux binding, opt-in attach status line, BEL/OSC 9 notifications | M |
| 4 | Retire `tmux-chooser.sh`/fold `tmux-manager.sh` (keep as thin wrappers for one release), docs/README/wiki, screenshots for promo | S |
Phases 0-1 are the useful minimum; 2 is where it beats herdr's sidebar (answering approvals from the dashboard); 3 is delight.
## 7. Testing
- Pure modules (`tui-model/layout/render/keys/ansi`): plain vitest, frame snapshots as stripped strings plus targeted ANSI assertions.
- Interactive E2E: spawn the built TUI under `node-pty` (already a dependency), feed keys, assert on captured frames; the vitest tmux mock (`IS_TEST_MODE`) keeps attach paths inert. Port rules per CLAUDE.md (3150+, `app.inject()` where possible by testing `tui-client` against injected routes).
- Manual: Termius/iPhone portrait (the 44-col case), tmux nesting, server-down mode, NO_COLOR, non-nerd-font terminal.
## 8. Invariants this plan respects
- tmux socket and data dir always via instance config (`dataPath()`, `-L codeman`); a beta instance TUI sees only its own world.
- Never bulk kill, always confirm, never touch another session implicitly, refuse killing the session the TUI runs in (w1/w2/w3 are sacred).
- Input is single-line with `\r`, via the server (never raw tmux send-keys from the TUI while the server owns the session).
- Approvals answering goes through the server's re-capture + 409 path, never blind keystrokes.
- `status off` on panes stays the default; any chrome is opt-in.
- No new runtime dependencies; the npm package stays light.
## 9. Decisions (resolved 2026-08-16)
1. **Bare `codeman` does NOT open the TUI** (owner decision): the web UI is the main thing, the TUI is additional. `codeman tui` only.
2. **`sc` stays the bash chooser for now**; the alias flip is a follow-up once the TUI has mileage. `codeman tui --list` / `codeman tui <n>` provide the same fast paths for people who want to switch.
3. Opt-in tmux status line: deferred to phase 3 along with the `--pick` popup switcher.
4. Preview tail goes over the API (auth/multi-user/remote-consistent); previews are simply unavailable in degraded server-down mode.
5. Name is `codeman tui` (the test fixture historically expected it).
Initial PR scope: phases 0-2. Phase 3 (mouse, popup switcher, status line, OSC 9) and phase 4 (bash chooser retirement) are follow-ups.
+278
View File
@@ -0,0 +1,278 @@
# Terminal UI (`codeman tui`)
`codeman tui` is a full-screen dashboard for your Codeman sessions, in the terminal.
It shows every session grouped by whether it needs you, lets you answer a permission
dialog or send a prompt without switching anywhere, and puts you inside a session's
tmux pane with one keystroke.
It is **additional, not a replacement**: the web UI stays the primary surface and
gets every feature first. The TUI exists for the terminal workflow (SSH, Termius,
a tmux window you keep open all day), and it is a *client* of the running server,
so the two surfaces can never disagree about what a session is doing. It is also
not a multiplexer: tmux still owns every pane, and attaching hands the terminal to
tmux rather than proxying bytes.
## Starting it
```bash
codeman tui # the dashboard
codeman tui --list # print the numbered session list and exit
codeman tui 2 # attach straight to session 2 of that list
```
The two fast paths are the scriptable ones.
Neither sets up a screen, so both are as quick as the one API call they make, and
`--list` prints plain text when piped, so it composes with `grep`/`awk`.
What it needs:
| Needs | What you get |
| --- | --- |
| **Full features** | A running Codeman server (states, approvals, preview, prompts, search, digest). The TUI finds it the way `codeman attach` does: `CODEMAN_API_URL`, else loopback on `CODEMAN_PORT` for this `CODEMAN_INSTANCE`. The self-signed certificate an `--https` install generates is accepted, as it is everywhere else in the CLI. |
| **Server down** | It still starts, in **degraded mode**: sessions are enumerated straight from `tmux -L codeman` plus a read-only peek at `state.json`, and attach is the only verb. See [Troubleshooting](#troubleshooting). |
| **A terminal** | `codeman tui` refuses to run when stdin/stdout are not a TTY, and says to use `--list` instead. A cron job or a pipe therefore fails loudly rather than emitting escape codes into a log. |
## What it looks like
A real frame at 100x30 (`NO_COLOR`, trailing blank rows trimmed). The selected
session has a pending permission dialog, so the preview pane leads with the card:
```
codeman ⚠ 2 tnode · v1.19.0 · 5 sessions · 5h 32% · wk 61% ? help q quit
NEEDS YOU ─────────────────────────│ w4-api-refactor · claude · /home/you/dev/api · blocked
1 w6-docs ✋ 11m│ ⚠ requests: Bash(git push origin main)
▶ 2 w4-api-refactor ⚠ 2m│ 1. Yes
WORKING ───────────────────────────│ 2. Yes, and do not ask again
3 w1-codeman ∗ 1h│ 3. No, tell Claude what to do
4 w2-gallery ∗ 15m│ y approve · n deny · digit chooses
IDLE ──────────────────────────────│
5 w3-promo shell ○ 2h│ > refactor the api routes onto the shared port interface
RECENT ────────────────────────────│
6 api-hotfix ✔ 3d│ Read src/web/ports/session-port.ts (48 lines)
│ Read src/api/routes.ts (312 lines)
│ Edit src/api/routes.ts
│ 1 -import { SessionManager } from "../session-manager.js";
│ 2 +import type { SessionPort } from "../web/ports/session-
│
│ Bash(npm run typecheck)
│ └ tsc --noEmit: no errors
│
│ ✻ Actualizing… (2m 14s · ↓ 12.3k tokens)
↑↓ select · ⏎ attach · y approve · n deny · 1-9 option · p prompt · x kill · / search · g digest ·
```
- **Header**: the machine, the server version, how many sessions are live, and the
plan-usage chip (the same statusLine telemetry that feeds the web chip, when the
server has a snapshot). A `⚠ n` badge counts pending approvals.
- **Sidebar**: every session, grouped and numbered.
- **Preview**: a live tail of the selected session, its own colors preserved, with
the parsed dialog card on top when that session is blocked.
- **Footer**: only the keys that work right now. `n` reads `n new` normally and
`n deny` when the selected session has a dialog, because it cannot be both.
The same world through `--list`:
```
1 waiting w6-docs /home/you/dev/docs
2 blocked w4-api-refactor /home/you/dev/api
3 working w1-codeman /home/you/dev/codeman
4 working w2-gallery /home/you/dev/gallery
5 idle w3-promo /home/you/dev/promo
6 done api-hotfix /home/you/dev/api
```
The numbers are the same on both surfaces, so `codeman tui --list` then
`codeman tui 4` is one thought.
## The four groups
Groups are always in this order, and a session is in exactly one of them:
| Group | Glyph | Means | Comes from |
| --- | --- | --- | --- |
| **NEEDS YOU** | `⚠` | A permission or question dialog is blocking the agent | The approvals inbox (`permission_prompt` hooks, with the on-screen options parsed) |
| | `✋` | Waiting for your next instruction, or errored | `idle_prompt`, or an errored session (equally something only a human clears) |
| **WORKING** | `✻` animating | A turn is running | The same working classification the web dashboard uses |
| **IDLE** | `○` | Live, but sitting there | |
| **RECENT** | `✔` | A past session from the unified list | History rows, no live pane |
Ordering inside a group is "the one that has waited longest, first": blocked
sessions sort by how long the dialog has been up, working sessions by when their
turn started (the pane's last Enter, since a working pane repaints every second
and would otherwise always look freshly started), and quiet ones by last activity.
That is the ordering the web home screens already use.
The cursor sticks to a **session**, not a row number, so a session that jumps to
NEEDS YOU does not drag your selection with it. The number beside each row is what
`1-9` and `codeman tui <n>` mean, and it is renumbered on every re-sort.
When a new dialog appears, the terminal bell rings once, for that dialog only: the
same item announced twice does not ring twice.
## Keymap
| Key | Does |
| --- | --- |
| `↑` `↓` or `j` `k` | Move the cursor. PageUp/PageDown jump five rows. |
| `Enter` | Attach to the selected session (see [Attaching](#attaching)) |
| `1`-`9` | Jump to that row and attach. When a dialog is on screen, a digit answers it instead (see below). |
| `y` | Approve the selected session's dialog |
| `n` | Deny it, or **start a new session** when there is no dialog |
| `p` | Send one line to the selected session without attaching |
| `x` | Kill the selected session; `y` confirms, any other key cancels |
| `/` | Search sessions, events and files |
| `g` | Away digest: what happened while you were gone |
| `?` | Help overlay |
| `Esc` | Close whatever overlay is open |
| `q` or `Ctrl+C` | Quit, restoring the screen you started with |
Inside the `p` composer and the `/` query: `←` `→` `Home` `End` `Delete`
`Backspace` plus `Ctrl+A` / `Ctrl+E` / `Ctrl+U` / `Ctrl+W`, `Enter` to send or open,
`Esc` (or `Ctrl+C`) to cancel. In the kill confirmation you retype the session name;
anything else cancels. In the `n` pickers, type to filter, `Enter` chooses.
Verbs that need the server (`y`/`n`/`p`/`x`/`/`/`g`) say so in degraded mode
instead of failing silently; `Enter` and `1-9` keep working.
### `p` sends exactly one line
The composer is a single line by design, ending in a carriage return: that is the
input contract every Codeman path follows, because multi-line text breaks the
agent's own composer. Pasted newlines become spaces rather than being rejected, so
a paste cannot silently run a different command than the one you read.
## Answering approvals
This is the thing the terminal could not do before. Select a blocked session and:
- `y` approves.
- `n` picks the parsed "No" option, or sends Esc when the dialog did not parse one.
- A digit picks that numbered option, **but only a digit the dialog actually
offers**. A digit with no matching option falls through to the list's own
jump-and-attach binding, so it can never be typed at whatever has focus.
The answer goes through `POST /api/approvals/:id/answer`, which **re-captures the
pane before it types anything**. If the dialog is no longer on screen (you answered
it in tmux a moment ago, or the agent moved on), the server refuses with a 409 and
the TUI says `that dialog is no longer on screen` rather than pressing a key into a
live composer. The answer is scoped to the options the server parsed off the actual
frame, never to a guess.
An idle prompt (`✋`) is not a dialog: there is nothing to approve, so `p` is the
reply path and the footer says `p reply` instead of `p prompt`.
## Attaching
`Enter` suspends the dashboard (main screen back, cooked mode back) and hands the
terminal to tmux with `stdio: inherit`. Colors, mouse and paste are tmux's, at full
fidelity.
**Press `F1` to come back.** One key, no modifier to hold or release, nothing to
type in a particular order. tmux's own way out is a chord — press the prefix, let
go, then a letter — and beta testing showed that is genuinely hard to convey: the
bar first named the wrong letter (tmux binds lowercase `d` to `detach-client` and
capital `D` to `choose-client`), and once corrected it still failed for anyone who
kept Ctrl held, because that sends `Ctrl+D`, which tmux leaves unbound. So the TUI
claims `F1` in tmux's prefix-less key table for the length of the attach and gives
it back afterwards. The chord still works; it is simply not what you are told to
press.
You do not have to remember any of it. For as long as the attach lasts the pane
wears a bar across the top:
```
1 w3-codeman-… 2 w4-codeman-… 3 testcase … alt+1-9 switch · F1 back to the codeman dashboard
```
That is the **session strip**: the other sessions stay visible from inside a pane,
numbered exactly as the dashboard numbers them, with the one you are in inverted.
`Alt+1`..`Alt+9` switch between them without going back to the dashboard first. With
more sessions than fit, the strip shows a window around the current one and marks
each cut end with `…`; the way-out hint is measured first and always keeps its space.
Codeman keeps the status bar off on its panes (the web UI carries that information
around the terminal instead), so the TUI turns it on for the attach and puts it back
exactly as it was on detach, along with each window's size. Every session the strip
can switch to is dressed and sized the same way, so switching is instant and lands
in a pane that already fills your terminal.
Detaching leaves the agent running; typing `exit` or pressing `Ctrl+D` would end it,
which is the difference the bar exists to make obvious. If an agent does exit, its
pane stays as a corpse: the TUI refuses to attach to a dead pane and offers `r` to
resume the conversation in a fresh one instead.
Three cases:
| Where you are | What happens |
| --- | --- |
| Not in tmux | `tmux -L codeman attach-session` |
| Already in tmux on Codeman's socket | `switch-client`, so you do not nest |
| In tmux on a **different** socket | Refused, with an explanation: detach from that tmux first, then run `codeman tui` again |
A direct-PTY session has no pane to attach to, and says so.
**`Enter` on a RECENT row resumes that conversation** instead: there is no pane to
attach to, so the TUI creates a new claude session carrying the old transcript
(`resumeSessionId`, exactly what the web UI's "Resume Conversation" list does), in
the directory it originally ran in and under its old name, then attaches to it. It
is claude-only, and a row with no working directory or no conversation id says why
rather than resuming something else.
`x` never bulk-kills: it kills one session, only after you retype its name, never a
history row, and never the session the TUI itself is running in.
## Over SSH, and on a phone
The TUI is an ordinary terminal program with no local dependencies beyond tmux, so
`ssh box` then `codeman tui` works exactly like running it locally. There is no
separate remote mode.
Below 72 columns (Termius, an iPhone in portrait) the preview pane is dropped and
rows take two lines each, keeping the cursor, the live states and the
answer/prompt/kill verbs. The switch is
width-driven at draw time, so unfolding a foldable or resizing a window re-lays out
immediately; there is no mode flag to set.
## Troubleshooting
**"The Codeman server rejected these credentials."** The server has
`CODEMAN_PASSWORD` set. Export `CODEMAN_PASSWORD` (and `CODEMAN_USERNAME` if it is
not `admin`), or put them in the data dir's `.env` (`~/.codeman/.env`), which is
where `codeman attach` already reads them from.
**`server not running: attach only`** in a yellow banner. Nothing answered on the
expected port, so the TUI fell back to enumerating tmux. You get names and attach;
you do not get states, approvals or previews, because those only exist on the
server. Start the server (`codeman web -d`, or `systemctl --user start codeman-web`)
and the banner clears on its own: the TUI keeps re-probing.
**It found the wrong server, or none.** Discovery is instance-scoped. A beta
instance (`CODEMAN_INSTANCE=beta`) has its own data dir *and* its own tmux socket,
so its TUI sees only its own sessions. Set `CODEMAN_PORT` or `CODEMAN_API_URL`
explicitly when you run more than one.
**"this terminal is already inside tmux on socket ..."** You are in a tmux session
on a socket that is not Codeman's, so attaching would nest two multiplexers whose
prefix keys collide. Detach from that tmux and run `codeman tui` from outside.
**Boxes and glyphs render as garbage.** The TUI picks a glyph tier from the
environment: no `TERM` (or `dumb`), or a non-UTF-8 locale, gets the ASCII set
(`[!] [w] [*] [-]`, `+`/`-`/`|` frames). Force it either way with
`CODEMAN_TUI_GLYPHS=ascii|unicode|nerd`.
**Colors.** Standard `NO_COLOR` / `FORCE_COLOR` handling (chalk's, the same as the
rest of the CLI). Under `NO_COLOR` the frame is cursor addressing and text only,
and the preview's own colors are stripped too, so a session's output cannot repaint
the dashboard.
**It refuses to open at all**, saying it needs an interactive terminal. stdout or
stdin is not a TTY. That is the guard: use `codeman tui --list`.
## Related
- [`docs/tui-plan.md`](tui-plan.md): the design record. Why hand-rolled ANSI, why a
client and not a second brain, and what is deliberately deferred.
- [`docs/approvals-inbox-plan.md`](approvals-inbox-plan.md): where the parsed
dialogs and the answer endpoint come from.
- [`docs/remote-sessions.md`](remote-sessions.md): remote-SSH cases, which the TUI
lists like any other session.
+15 -4
View File
@@ -161,10 +161,21 @@ then every API call fails, which looks like the dashboard being broken.
then streams the body without any time bound; a header timeout is logged
server-side and answered as a 502 that names the limit. WebSocket handshakes use
the separate `CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS` (default 30s).
- **Not a security boundary.** The proxy reaches whatever the Codeman server can
reach. That is not an escalation for someone who already commands
`--dangerously-skip-permissions` agents, but in multi-user mode it does mean a
non-admin user's dashboard is fetched from the server's network position.
- **Not a security boundary, with one carve-out.** The proxy reaches whatever the
Codeman server can reach (a `localhost` dashboard is the point), so it is not an
escalation for someone who already commands `--dangerously-skip-permissions`
agents, but in multi-user mode it does mean a non-admin user's dashboard is
fetched from the server's network position. The carve-out: link-local and
cloud-metadata addresses (`169.254.0.0/16`, `fe80::/10`, `fd00:ec2::254`,
Azure's `168.63.129.16`, Alibaba's `100.100.100.200`, the
`metadata.google.internal` alias) are refused at save time AND at connect
time, judged on the address a name actually resolves to. Nothing anyone embeds
as a dashboard lives there; an instance's IAM credentials do.
- **The proxy URL is a bearer credential.** `/webview/<cap>/...` needs no cookie,
so treat it like a password. It is revoked when you log out, when an admin logs
you out, and when your account is deleted, and it expires after 12 hours
without use. Proxied responses carry `Referrer-Policy: same-origin`, so a
dashboard that links to third-party sites does not hand them the URL.
## Where the code lives
+1 -1
View File
@@ -88,7 +88,7 @@ would. That indirection buys:
- **Real scrollback.** History is held by tmux, so reconnecting replays what happened while
you were gone instead of starting from blank.
- **Attach from anywhere else.** The same session is reachable from a terminal over SSH
with the `sc` chooser, or plain `tmux -L codeman attach`.
with `codeman tui`, or plain `tmux -L codeman attach`.
- **Secrets off the command line.** Environment overrides are injected with socket-scoped
`tmux setenv` rather than being visible in the spawn command.
+4 -2
View File
@@ -73,8 +73,10 @@ features are Claude-only; [Agent CLIs](Agent-CLIs) lists exactly which.
### Can I attach to a session from a terminal instead of the browser?
Yes. `sc` is an interactive chooser (`sc 2` attaches directly, `sc -l` lists), or use tmux
directly on the `codeman` socket. Detach with `Ctrl+A D`.
Yes. `codeman tui` is a full-screen dashboard of your sessions, with the same
NEEDS YOU / WORKING / IDLE grouping the web UI uses. `codeman tui --list` prints the
numbered list and exits, and `codeman tui 2` attaches straight to session 2. `Enter`
attaches, `F1` comes back. You can also use tmux directly on the `codeman` socket.
## Running unattended
+1 -1
View File
@@ -119,7 +119,7 @@ Shell and external CLI sessions accept `idle`, `working`, and `exit`.
## SSE
`GET /api/events` is the live event stream. 155 event names, kept in sync between server and
`GET /api/events` is the live event stream. 156 event names, kept in sync between server and
client with a test that fails on drift.
The heartbeat is a **named** `sse:heartbeat` event rather than an SSE comment, because
+9 -6
View File
@@ -179,16 +179,19 @@ the same IP, which matters because all tunnel traffic arrives from one loopback
## Terminal alternatives
You do not have to use a browser. `sc` is a thumb-friendly session chooser for SSH clients
like Termius or Blink:
You do not have to use a browser. `codeman tui` is a full-screen session dashboard that
works well in SSH clients like Termius or Blink:
```bash
sc # interactive chooser
sc 2 # attach to session 2
sc -l # list
codeman tui # the dashboard
codeman tui 2 # attach straight to session 2
codeman tui --list # numbered list, then exit
```
Detach with `Ctrl+A D`. The sessions are the same ones the dashboard shows.
`Enter` attaches into the pane and `F1` comes back. Under 72 columns it drops the preview
and becomes a single-column list, so it stays usable on a phone. The sessions are the same
ones the dashboard shows. See [docs/tui.md](https://github.com/Ark0N/Codeman/blob/master/docs/tui.md)
for the full guide.
## Common problems
+1
View File
@@ -45,6 +45,7 @@ supervised by systemd or launchd; npm installs report as non-updatable. See
| CJK Input | Off | IME composition through a dedicated text field. |
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
+4 -2
View File
@@ -133,8 +133,10 @@ TUIs render correctly.
Worth knowing:
- **Scrollback.** The first time you open a session, Codeman pulls the entire tmux
scrollback, not just the recent tail. Scrolling to the very top pulls again on demand.
- **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open.
Shell sessions open from a bounded recent tail so a large transcript cannot stall tab
switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling
and automatic output recovery stay within the bounded browser buffer.
- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude
versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is
always local scrollback. Other CLIs scroll locally.
+353 -27
View File
@@ -125,6 +125,22 @@ PI_SEARCH_PATHS=(
"$HOME/bin/pi"
)
# DeepSeek Harness search paths (from src/utils/deepseek-cli-resolver.ts)
DSH_SEARCH_PATHS=(
"$HOME/.local/bin/dsh"
"/usr/local/bin/dsh"
"$HOME/.npm-global/bin/dsh"
"$HOME/bin/dsh"
)
# Grok CLI search paths (from src/utils/grok-cli-resolver.ts)
GROK_SEARCH_PATHS=(
"$HOME/.grok/bin/grok"
"$HOME/.local/bin/grok"
"/usr/local/bin/grok"
"$HOME/bin/grok"
)
# Antigravity CLI search paths (from src/utils/antigravity-cli-resolver.ts)
ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/.local/bin/agy"
@@ -133,6 +149,17 @@ ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/bin/agy"
)
# OMP CLI search paths (from src/utils/omp-cli-resolver.ts's OMP_SEARCH_DIRS —
# ~/.local/bin leads, omp.sh's installer target; ~/.omp/bin is a fallback only)
OMP_SEARCH_PATHS=(
"$HOME/.local/bin/omp"
"$HOME/.omp/bin/omp"
"/usr/local/bin/omp"
"$HOME/.bun/bin/omp"
"$HOME/.npm-global/bin/omp"
"$HOME/bin/omp"
)
# ============================================================================
# Color Output
# ============================================================================
@@ -229,7 +256,11 @@ print_security_notice() {
echo -e " ${YELLOW}${BOLD}Security:${NC}"
echo -e " Codeman binds ${BOLD}127.0.0.1${NC} (this machine only) — no password needed by default."
echo -e " To reach it from another device, do ONE of:"
echo -e " ${CYAN}•${NC} tailscale serve / cloudflared tunnel ${DIM}(recommended)${NC}, or"
if check_tailscale; then
echo -e " ${CYAN}•${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC} ${DIM}(Tailscale is installed here; HTTPS, recommended)${NC}, or"
else
echo -e " ${CYAN}•${NC} tailscale serve / cloudflared tunnel ${DIM}(recommended)${NC}, or"
fi
echo -e " ${CYAN}•${NC} ${CYAN}codeman web --host 0.0.0.0${NC} AND set ${CYAN}CODEMAN_PASSWORD${NC}"
echo -e " A non-loopback bind without a password still starts, but warns loudly."
echo -e " ${DIM}Details: docs/security-architecture.md${NC}"
@@ -396,6 +427,27 @@ check_tmux() {
command -v tmux &>/dev/null
}
# node-pty ships prebuilt binaries for darwin and win32 ONLY, so on Linux it is
# always compiled from source during `npm install`. Without a toolchain that
# fails deep inside node-gyp with `not found: make`, which reads like an npm bug
# rather than a missing system package (issue: fresh Ubuntu 24 server install).
# So the toolchain is checked up front, exactly like git and tmux.
#
# Returns a human-readable list of what is missing, empty when all present.
missing_build_tools() {
local missing=""
command -v make &>/dev/null || missing="make"
if ! command -v c++ &>/dev/null && ! command -v g++ &>/dev/null && ! command -v clang++ &>/dev/null; then
missing="${missing:+$missing, }a C++ compiler (g++)"
fi
command -v python3 &>/dev/null || missing="${missing:+$missing, }python3"
printf '%s' "$missing"
}
check_build_tools() {
[[ -z "$(missing_build_tools)" ]]
}
check_claude() {
# Check PATH first
if command -v claude &>/dev/null; then
@@ -569,6 +621,119 @@ get_pi_path() {
done
}
# `grok` has known squatters too (the unrelated @vibe-kit/grok-cli), so the
# server-side resolver additionally probes `grok --version`. Detection here only
# feeds the "you have no AI CLI" hint, so a plain executable test is enough.
check_grok() {
if command -v grok &>/dev/null; then
return 0
fi
for path in "${GROK_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
# `dsh` is the hardest name of the lot: Debian ships an unrelated `dsh`
# (dancer's shell). The server-side resolver settles it by demanding the
# harness's own help banner; detection here only feeds the "you have no AI CLI"
# hint, so the same banner grep is enough — but unlike every sibling probe it
# EXECUTES the candidate, so it must be bounded. </dev/null is load-bearing
# twice over: a foreign binary that blocks on stdin would hang the install, and
# under `curl | bash` a child that reads stdin EATS THE REST OF THIS SCRIPT.
# The timeout (where coreutils ships one; stock macOS has none) bounds a binary
# that ignores EOF, mirroring the server resolver's own EXEC_TIMEOUT_MS.
dsh_banner_probe() {
local runner=()
if command -v timeout &>/dev/null; then runner=(timeout 5); fi
"${runner[@]}" "$1" --help </dev/null 2>/dev/null | grep -qi "DeepSeek Harness"
}
# Resolved ONCE and memoized: the probe executes a possibly-foreign binary, and
# the check/get/reminder call sites together used to re-run the whole scan many
# times per install.
DSH_RESOLVE_DONE=""
DSH_RESOLVED_PATH=""
resolve_dsh() {
[[ -n "$DSH_RESOLVE_DONE" ]] && return 0
DSH_RESOLVE_DONE=1
local candidate path
if command -v dsh &>/dev/null; then
candidate="$(command -v dsh)"
if dsh_banner_probe "$candidate"; then
DSH_RESOLVED_PATH="$candidate"
return 0
fi
fi
for path in "${DSH_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]] && dsh_banner_probe "$path"; then
DSH_RESOLVED_PATH="$path"
return 0
fi
done
return 0
}
check_dsh() {
resolve_dsh
[[ -n "$DSH_RESOLVED_PATH" ]]
}
get_dsh_path() {
resolve_dsh
echo "$DSH_RESOLVED_PATH"
}
get_grok_path() {
if command -v grok &>/dev/null; then
command -v grok
return
fi
for path in "${GROK_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
# `omp` is a short name too, so like grok/pi the server-side resolver
# additionally probes `omp --version`. Detection here only feeds the
# "you have no AI CLI" hint, so a plain executable test is enough.
check_omp() {
if command -v omp &>/dev/null; then
return 0
fi
for path in "${OMP_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
get_omp_path() {
if command -v omp &>/dev/null; then
command -v omp
return
fi
for path in "${OMP_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
check_cloudflared() {
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
@@ -836,6 +1001,50 @@ install_git_suse() {
run_as_root zypper install -y git
}
# Build toolchain for node-pty's source compile (see missing_build_tools).
install_buildtools_debian() {
info "Installing build tools via apt (build-essential, python3)..."
ensure_sudo
run_as_root apt-get update -qq
run_as_root apt-get install -y -qq build-essential python3
}
install_buildtools_fedora() {
info "Installing build tools (gcc, gcc-c++, make, python3)..."
ensure_sudo
if command -v dnf &>/dev/null; then
run_as_root dnf install -y gcc gcc-c++ make python3
else
run_as_root yum install -y gcc gcc-c++ make python3
fi
}
install_buildtools_arch() {
info "Installing build tools via pacman (base-devel, python)..."
ensure_sudo
run_as_root pacman -Sy --noconfirm base-devel python
}
install_buildtools_alpine() {
info "Installing build tools via apk (build-base, python3)..."
ensure_sudo
run_as_root apk add --no-cache build-base python3
}
install_buildtools_suse() {
info "Installing build tools via zypper..."
ensure_sudo
run_as_root zypper install -y gcc gcc-c++ make python3
}
install_buildtools_macos() {
# macOS normally never gets here: node-pty ships darwin prebuilds. Only a
# forced source build needs a compiler, and Xcode CLT is its only supplier.
info "Requesting Xcode Command Line Tools..."
xcode-select --install 2>/dev/null || true
die "Finish the Xcode Command Line Tools install in the dialog, then re-run this installer."
}
install_cloudflared_macos() {
info "Installing cloudflared via Homebrew..."
ensure_homebrew
@@ -1072,21 +1281,28 @@ add_to_path() {
success "Added to $profile"
}
setup_sc_alias() {
# The `sc` bash chooser was retired in favour of `codeman tui`, which reaches
# sessions 10+, carries the server's real states and leaves an attach with one
# key. Older installers wrote this alias, so take it back out.
#
# Marker-owned on purpose: it matches the exact line WE wrote, so a user's own
# `alias sc=` for something entirely different is never touched. The rewrite
# goes through `cat >` rather than `mv` so the profile keeps its own mode and
# ownership.
remove_sc_alias() {
local profile
profile=$(detect_shell_profile)
[[ -f "$profile" ]] || return 0
grep -qE "^alias sc='tmux-chooser'\$" "$profile" 2>/dev/null || return 0
# Check if alias already exists
if [[ -f "$profile" ]] && grep -qE "^alias sc=" "$profile" 2>/dev/null; then
info "Alias 'sc' already configured in $profile"
return 0
local tmp
tmp=$(mktemp 2>/dev/null) || return 0
if sed -e "/^alias sc='tmux-chooser'\$/d" \
-e '/^# Codeman tmux session shortcut$/d' "$profile" > "$tmp" 2>/dev/null; then
cat "$tmp" > "$profile"
info "Removed the retired 'sc' alias from $profile (use: codeman tui)"
fi
echo "" >> "$profile"
echo "# Codeman tmux session shortcut" >> "$profile"
echo "alias sc='tmux-chooser'" >> "$profile"
info "Added 'sc' alias for tmux-chooser"
rm -f "$tmp"
}
# ============================================================================
@@ -1704,6 +1920,41 @@ setup_tailscale_access() {
return 0
}
# A loopback install with Tailscale already connected but nothing fronting
# Codeman is one command away from working remote access — and that is exactly
# where a user lands when the first install died BEFORE the network-access
# prompt (it runs after the build, so any build failure costs the network step
# too) or when they finished a broken build by hand instead of re-running the
# installer. Detect that state on re-run and offer the retrofit, rather than
# leaving them to discover `install.sh tailscale` on their own. Never nags a
# deliberate network bind, and never nags once a serve mapping already exists.
maybe_offer_tailscale_repair() {
# A non-loopback bind already has network access; leave that choice alone.
if [[ "$EXISTING_FOUND" == "1" && -n "$EXISTING_HOST" && "$EXISTING_HOST" != "127.0.0.1" ]]; then
return 0
fi
check_tailscale || return 0
command -v node &>/dev/null || return 0
[[ "$(ts_status_field 's.BackendState')" == "Running" ]] || return 0
# Already fronting Codeman: nothing to repair.
[[ -z "$(detect_tailscale_serve_url)" ]] || return 0
echo ""
info "Tailscale is connected here, but no serve mapping fronts Codeman yet."
if [[ "$NONINTERACTIVE" == "1" ]] || ! has_tty; then
echo -e " ${DIM}Enable HTTPS access from your tailnet with:${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC}"
return 0
fi
if ! prompt_yes_no "Set up Tailscale HTTPS access now? (your tailnet is the login; no password needed)" "y"; then
echo -e " ${DIM}Any time later:${NC} ${CYAN}bash $INSTALL_DIR/install.sh tailscale${NC}"
return 0
fi
if setup_tailscale_access; then
verify_tailscale_access || true
fi
return 0
}
# `install.sh tailscale`: retrofit Tailscale access onto an existing install
# (also the target of every "set it up later" hint above).
setup_tailscale_subcommand() {
@@ -1956,6 +2207,29 @@ setup_tunnel_service() {
# Installation Helpers
# ============================================================================
# npm install with an actionable message for the failure that actually happens
# on a fresh Linux box: no toolchain, so node-pty cannot compile.
npm_install_deps() {
if npm install --quiet --no-fund --no-audit 2>/dev/null; then
return 0
fi
if npm install --no-fund --no-audit; then
return 0
fi
error "npm install failed."
if [[ "$(detect_os)" == "linux" ]] && ! check_build_tools; then
error "Missing native build tools: $(missing_build_tools)"
error "node-pty has no Linux prebuilds, so it must compile from source."
error "Install them and re-run this installer:"
error " Debian/Ubuntu: sudo apt-get install -y build-essential python3"
error " Fedora/RHEL: sudo dnf install -y gcc gcc-c++ make python3"
error " Arch: sudo pacman -S --noconfirm base-devel python"
error " Alpine: sudo apk add build-base python3"
fi
exit 1
}
install_dependency() {
local dep_name="$1"
local os="$2"
@@ -2069,6 +2343,31 @@ main() {
fi
fi
# Native build toolchain. node-pty compiles from source on Linux, so this is
# a hard requirement there, not a nicety.
if [[ "$os" == "linux" ]]; then
info "Checking build tools (node-pty compiles from source on Linux)..."
local missing_tools
missing_tools="$(missing_build_tools)"
if [[ -z "$missing_tools" ]]; then
success "Build tools are installed"
else
warn "Missing build tools: $missing_tools"
headless_guard "install build tools (system package via sudo)"
if prompt_yes_no "Install the build tools now?"; then
install_dependency "buildtools" "$os" "$distro"
hash -r 2>/dev/null || true
missing_tools="$(missing_build_tools)"
if [[ -n "$missing_tools" ]]; then
die "Build tools still missing after install: $missing_tools. Install them manually and re-run."
fi
success "Build tools installed"
else
die "A build toolchain (make, g++, python3) is required: node-pty has no Linux prebuilds and compiles from source."
fi
fi
fi
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi)
local has_claude=false
local has_opencode=false
@@ -2076,6 +2375,9 @@ main() {
local has_gemini=false
local has_antigravity=false
local has_pi=false
local has_grok=false
local has_dsh=false
local has_omp=false
info "Checking AI CLI tools..."
if check_claude; then
@@ -2102,17 +2404,29 @@ main() {
has_pi=true
success "Pi CLI found at $(get_pi_path)"
fi
if check_grok; then
has_grok=true
success "Grok CLI found at $(get_grok_path)"
fi
if check_dsh; then
has_dsh=true
success "DeepSeek Harness found at $(get_dsh_path)"
fi
if check_omp; then
has_omp=true
success "OMP CLI found at $(get_omp_path)"
fi
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" && "$has_pi" == "false" ]]; then
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" && "$has_pi" == "false" && "$has_grok" == "false" && "$has_dsh" == "false" && "$has_omp" == "false" ]]; then
echo ""
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi."
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, Gemini, Pi, Grok, DeepSeek Harness, or OMP."
headless_guard "install an AI CLI (curl | bash from its vendor)"
echo ""
echo -e " ${BOLD}Which AI CLI would you like to install?${NC}"
echo -e " ${CYAN}1)${NC} Claude Code (Anthropic)"
echo -e " ${CYAN}2)${NC} OpenCode (open-source)"
echo -e " ${CYAN}3)${NC} Both"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex, Antigravity or Pi)"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex, Antigravity, Gemini, Pi, Grok, DeepSeek Harness or OMP)"
echo ""
local cli_choice=""
@@ -2160,6 +2474,7 @@ main() {
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity)"
info " or: npm install -g --ignore-scripts @earendil-works/pi-coding-agent (Pi)"
info " or: curl -fsSL https://x.ai/cli/install.sh | bash (Grok)"
elif [[ "$has_claude" == "false" ]] && [[ "$has_opencode" == "false" ]]; then
die "The selected AI CLI failed to install. Install one manually and re-run the installer."
fi
@@ -2225,7 +2540,7 @@ main() {
# ========================================================================
info "Installing dependencies..."
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm_install_deps
info "Building..."
npm run build --quiet 2>/dev/null || npm run build
@@ -2243,13 +2558,14 @@ main() {
ln -sf "$INSTALL_DIR/dist/index.js" "$symlink_dir/codeman"
info "Created symlink: $symlink_dir/codeman"
# Install tmux-chooser as 'tmux-chooser' command
if [[ -f "$INSTALL_DIR/scripts/tmux-chooser.sh" ]]; then
ln -sf "$INSTALL_DIR/scripts/tmux-chooser.sh" "$symlink_dir/tmux-chooser"
info "Created symlink: $symlink_dir/tmux-chooser"
# Add 'sc' alias for quick access
setup_sc_alias
# tmux-chooser/`sc` is retired; `codeman tui` replaces it. Sweep up what
# an older installer left behind, so an update does not leave a symlink
# pointing at a script this version no longer ships.
if [[ -L "$symlink_dir/tmux-chooser" ]]; then
rm -f "$symlink_dir/tmux-chooser"
info "Removed the retired tmux-chooser symlink (use: codeman tui)"
fi
remove_sc_alias
# Add ~/.local/bin to PATH if not already there
if [[ ":$PATH:" != *":$symlink_dir:"* ]]; then
@@ -2450,23 +2766,28 @@ main() {
echo -e " ${BOLD}Mobile Access (Termius/SSH):${NC}"
echo ""
echo -e " ${CYAN}sc${NC} # Interactive tmux session chooser"
echo -e " ${CYAN}sc 2${NC} # Quick attach to session 2"
echo -e " ${CYAN}sc -h${NC} # Help"
echo -e " ${CYAN}codeman tui${NC} # Full-screen session dashboard"
echo -e " ${CYAN}codeman tui 2${NC} # Attach straight to session 2"
echo -e " ${CYAN}codeman tui -l${NC} # Numbered list, then exit"
echo ""
echo -e " ${BOLD}Documentation:${NC}"
echo -e " https://github.com/Ark0N/Codeman"
echo ""
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity && ! check_pi; then
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity && ! check_pi && ! check_grok && ! check_dsh && ! check_omp; then
echo -e " ${YELLOW}${BOLD}Reminder:${NC} Install at least one AI CLI to start using Codeman:"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}curl -fsSL https://antigravity.google/cli/install.sh | bash${NC} # Antigravity"
echo -e " ${CYAN}npm install -g --ignore-scripts @earendil-works/pi-coding-agent${NC} # Pi"
echo -e " ${CYAN}curl -fsSL https://x.ai/cli/install.sh | bash${NC} # Grok"
echo -e " ${CYAN}curl -fsSL https://omp.sh/install | sh${NC} # OMP"
echo ""
echo -e " DeepSeek Harness has no vendor one-liner — install it from within Codeman"
echo -e " once the server is up (Run dropdown → Install DeepSeek Profile, or see"
echo -e " docs/deepseek-integration.md)."
fi
# Security notice — last informational block so it stays visible (when not
@@ -2519,7 +2840,7 @@ update() {
git fetch --quiet origin
git reset --hard "origin/$BRANCH" --quiet
npm install --quiet --no-fund --no-audit 2>/dev/null || npm install --no-fund --no-audit
npm_install_deps
npm run build --quiet 2>/dev/null || npm run build
date -u +%Y-%m-%dT%H:%M:%SZ > "$INSTALL_DIR/.install-complete"
success "Updated to $(node -e "console.log(require('./package.json').version)")"
@@ -2556,6 +2877,10 @@ update() {
BIND_ACK="$EXISTING_ACK"
fi
# An update is the only place a half-configured install gets a second
# chance at remote access; the fresh-install path asks outright.
maybe_offer_tailscale_repair
print_security_notice
}
@@ -2620,6 +2945,7 @@ uninstall() {
rm -f "$symlink_dir/tmux-chooser"
success "Removed symlink: $symlink_dir/tmux-chooser"
fi
remove_sc_alias
# Remove install directory
if [[ -d "$INSTALL_DIR" ]]; then
+12 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.19.7",
"version": "1.24.7",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.19.7",
"version": "1.24.7",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -32,6 +32,7 @@
"jpeg-js": "^0.4.4",
"node-pty": "^1.1.0",
"qrcode": "^1.5.4",
"undici": "^6.28.0",
"uuid": "^14.0.0",
"web-push": "^3.6.7",
"ws": "^8.21.0",
@@ -11550,6 +11551,15 @@
"dev": true,
"license": "MIT"
},
"node_modules/undici": {
"version": "6.28.0",
"resolved": "https://registry.npmjs.org/undici/-/undici-6.28.0.tgz",
"integrity": "sha512-LIY910g9TI13YS95lrMFrs8Rm/u/irgHeTWoKCoteeJ04CUJ92eEfj0rVn+7VKMPBpUPiUoBKfhNyLI23EE/KA==",
"license": "MIT",
"engines": {
"node": ">=18.17"
}
},
"node_modules/undici-types": {
"version": "6.21.0",
"resolved": "https://registry.npmjs.org/undici-types/-/undici-types-6.21.0.tgz",
+4 -1
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.19.7",
"version": "1.24.7",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -62,6 +62,8 @@
"codex",
"antigravity",
"pi",
"grok",
"deepseek",
"gemini-cli",
"ai-agents",
"agent",
@@ -100,6 +102,7 @@
"jpeg-js": "^0.4.4",
"node-pty": "^1.1.0",
"qrcode": "^1.5.4",
"undici": "^6.28.0",
"uuid": "^14.0.0",
"web-push": "^3.6.7",
"ws": "^8.21.0",
+2
View File
@@ -86,6 +86,7 @@ run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --
run('minify i18n.js', 'npx esbuild dist/web/public/i18n.js --minify --outfile=dist/web/public/i18n.js --allow-overwrite');
run('minify sanitize-html.js', 'npx esbuild dist/web/public/sanitize-html.js --minify --outfile=dist/web/public/sanitize-html.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify tab-rail-resize.js', 'npx esbuild dist/web/public/tab-rail-resize.js --minify --outfile=dist/web/public/tab-rail-resize.js --allow-overwrite');
run('minify terminal-ui.js', 'npx esbuild dist/web/public/terminal-ui.js --minify --outfile=dist/web/public/terminal-ui.js --allow-overwrite');
run('minify respawn-ui.js', 'npx esbuild dist/web/public/respawn-ui.js --minify --outfile=dist/web/public/respawn-ui.js --allow-overwrite');
run('minify ralph-panel.js', 'npx esbuild dist/web/public/ralph-panel.js --minify --outfile=dist/web/public/ralph-panel.js --allow-overwrite');
@@ -111,6 +112,7 @@ console.log('\n[build] content-hash cache busting');
'input-cjk.js',
'sanitize-html.js',
'app.js',
'tab-rail-resize.js',
'terminal-ui.js',
'respawn-ui.js',
'ralph-panel.js',
+55 -7
View File
@@ -7,18 +7,25 @@
# the repo (the server stages it at ~/.codeman/self-update-runner.sh) — `git
# checkout` rewrites the in-repo copy and bash reads scripts lazily.
#
# ⚠️ The `docker-compose` supervisor is the exception to "outlives": there the
# restart IS the container exiting, which kills this script too. That is safe
# because the terminal "restarting" marker is written before the kill and the
# rebooted server reconciles it — but nothing may be added after that kill.
#
# Reports progress by writing ~/.codeman/update-status.json atomically; the
# browser polls GET /api/system/update/status across the restart drop. The
# freshly-booted server reconciles the final "restarting" → "completed"/"failed".
#
# Cross-platform: restarts via systemd (Linux), launchd (macOS), or prints a
# manual command (foreground installs). Linux launches inside a transient
# systemd scope so `systemctl restart codeman-web` can't kill it mid-build.
# Cross-platform: restarts via systemd (Linux), launchd (macOS), a container exit
# under Docker Compose (the restart policy relaunches it), or prints a manual
# command (foreground installs). Linux launches inside a transient systemd scope
# so `systemctl restart codeman-web` can't kill it mid-build.
#
# Args (all from the server, never user input — tag is validated server-side):
# --repo <dir> --tag <codeman@X.Y.Z> --supervisor <systemd|launchd|none>
# --repo <dir> --tag <codeman@X.Y.Z> --supervisor <systemd|launchd|docker-compose|none>
# --status-file <path> --update-id <uuid> --from-version <ver> --node <path>
# --log <path> [--prev-sha <sha>] [--stash]
# --log <path> [--prev-sha <sha>] [--stash] [--server-pid <pid>]
# [--restart-by-exit 0|1] (docker-compose only: may we exit the server?)
#
set -uo pipefail
@@ -32,6 +39,7 @@ REPO=""
TAG=""
SUPERVISOR="none"
SERVER_PID=""
RESTART_BY_EXIT="0"
STATUS_FILE=""
UPDATE_ID=""
FROM_VERSION=""
@@ -52,6 +60,7 @@ while [[ $# -gt 0 ]]; do
--log) LOG="$2"; shift 2 ;;
--prev-sha) PREV_SHA="$2"; shift 2 ;;
--server-pid) SERVER_PID="$2"; shift 2 ;;
--restart-by-exit) RESTART_BY_EXIT="$2"; shift 2 ;;
--stash) DO_STASH=1; shift ;;
*) shift ;;
esac
@@ -144,7 +153,7 @@ rollback_and_fail() {
echo "[self-update] $msg — rolling back to ${PREV_SHA:-<none>}"
if [[ -n "$PREV_SHA" ]]; then
git checkout --force "$PREV_SHA" >/dev/null 2>&1 || true
npm install --no-fund --no-audit >/dev/null 2>&1 || true
npm install --no-fund --no-audit --include=dev >/dev/null 2>&1 || true
npm run build >/dev/null 2>&1 || true
fi
fail "$msg — rolled back to the previous version" "$msg"
@@ -176,7 +185,9 @@ write_status "checkout" "Checking out $TAG…"
git -c advice.detachedHead=false checkout --force "$TAG" || rollback_and_fail "Could not check out $TAG"
# 4) Install dependencies (heartbeat keeps the UI live during this slow step).
run_step "installing" "Installing dependencies" npm install --no-fund --no-audit \
# --include=dev: tsc and esbuild are devDependencies, and the Compose image sets
# NODE_ENV=production, which would otherwise omit them and fail the build below.
run_step "installing" "Installing dependencies" npm install --no-fund --no-audit --include=dev \
|| rollback_and_fail "Dependency install failed"
# 5) Build (gate the restart on success — never restart into a torn dist/).
@@ -200,6 +211,43 @@ case "$SUPERVISOR" in
|| fail "Build succeeded but launchd restart failed" "launchctl"
}
;;
docker-compose)
# In the Compose deployment there is no init system to ask: the "restart" is
# the server EXITING, so the container's `restart: unless-stopped` policy
# relaunches it on the dist/ we just built. The repo and dist/ live on host
# mounts, so the new build survives the container being replaced.
#
# ⚠️ This script dies WITH the container it is restarting — it is a child of
# the server process, not a survivor like the systemd-scope path. That is
# fine, and load-bearing: the terminal "restarting" marker is already written
# above, and the freshly-booted server reconciles it. Nothing may be appended
# after the kill that the update depends on.
#
# ⚠️ The server is signalled by PID rather than `docker restart`: this
# container's own Docker CLI talks to the HOST daemon, and a self-directed
# restart there races the client's own death. Exiting is the one path that
# needs no cooperation from anything outside the container.
#
# ⚠️ Only when the SERVER said the container comes back (`--restart-by-exit 1`:
# the Compose file declared it, or the daemon reported an auto-restart policy).
# An unknown policy stages the build and asks for a restart instead. Exiting
# blind would take a container the daemon does not restart down for good,
# with no UI left to recover it from.
if [[ "$RESTART_BY_EXIT" != "1" ]]; then
MANUAL_CMD="docker restart \$(hostname) # from the Docker host"
write_status "completed-needs-manual-restart" "Update built — restart the Codeman container to apply v$TO_VERSION."
echo "[self-update] docker-compose: restart-by-exit not confirmed — not exiting; manual restart required"
exit 0
fi
if [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null; then
: # container exit + restart policy take it from here
else
MANUAL_CMD="docker restart \$(hostname) # from the Docker host"
write_status "completed-needs-manual-restart" "Update staged — restart the Codeman container to apply v$TO_VERSION."
echo "[self-update] docker-compose: could not signal server pid '$SERVER_PID' — manual restart required"
exit 0
fi
;;
launchd-daemon)
# System-level KeepAlive LaunchDaemon (headless Mac): kickstarting the system
# domain needs root, but we don't need it — kill the server and launchd
-662
View File
@@ -1,662 +0,0 @@
#!/bin/bash
# ============================================================================
# Codeman Sessions - Mobile-friendly Tmux Session Chooser
# Optimized for iPhone/Termius (portrait ~45 chars, landscape ~95 chars)
# ============================================================================
#
# Design principles:
# - Single-digit selection (1-9) for fast thumb typing
# - Compact display, no wasted space
# - Color-coded status for quick scanning
# - Names pulled from Codeman state.json
# - Minimal keystrokes to attach
#
# Usage:
# tmux-chooser # Interactive chooser
# tmux-chooser 1 # Quick attach to session 1
# tmux-chooser -l # List only (non-interactive)
# tmux-chooser -h # Help
#
# Alias: alias sc='tmux-chooser'
# Then: sc (interactive)
# sc 2 (attach session 2)
#
# ============================================================================
set -e
# ============================================================================
# Configuration
# ============================================================================
CODEMAN_STATE="$HOME/.codeman/state.json"
CODEMAN_SESSIONS="$HOME/.codeman/mux-sessions.json"
# Dedicated tmux socket all Codeman sessions live on. MUST match
# DEFAULT_CODEMAN_TMUX_SOCKET / CODEMAN_TMUX_SOCKET in src/tmux-manager.ts —
# otherwise list-sessions would enumerate the user's default tmux server
# (missing the real Codeman sessions, surfacing unrelated ones).
CODEMAN_TMUX_SOCKET="${CODEMAN_TMUX_SOCKET:-codeman}"
TMUX_CMD=(tmux -L "$CODEMAN_TMUX_SOCKET")
# iPhone 17 Pro portrait width (conservative)
MAX_WIDTH=44
MAX_NAME_LEN=28
# Page size for pagination (leave room for header/footer)
PAGE_SIZE=7
# Auto-refresh timeout (seconds) - 0 to disable
AUTO_REFRESH=60
# ============================================================================
# Icon Detection (Nerd Fonts vs ASCII)
# ============================================================================
detect_icons() {
if [[ "$TERM_PROGRAM" == "iTerm"* ]] || \
[[ "$TERM" == "xterm-kitty" ]] || \
[[ -n "$WEZTERM_PANE" ]] || \
[[ "$LC_TERMINAL" == "iTerm2" ]]; then
ICON_SESSION="󰆍"
ICON_ATTACHED="●"
ICON_DETACHED="○"
ICON_UNKNOWN="◌"
else
ICON_SESSION="[T]"
ICON_ATTACHED="*"
ICON_DETACHED="-"
ICON_UNKNOWN="?"
fi
}
detect_icons
# ============================================================================
# Colors - ANSI 256 for better Termius compatibility
# ============================================================================
R='\033[0m' # Reset
B='\033[1m' # Bold
D='\033[2m' # Dim
GREEN='\033[38;5;82m'
YELLOW='\033[38;5;220m'
BLUE='\033[38;5;75m'
CYAN='\033[38;5;87m'
RED='\033[38;5;203m'
GRAY='\033[38;5;245m'
WHITE='\033[38;5;255m'
BG_SEL='\033[48;5;236m'
# ============================================================================
# Utilities
# ============================================================================
truncate() {
local str="$1"
local max="$2"
local len=${#str}
if [ "$len" -le "$max" ]; then
echo "$str"
return
fi
if [[ "$str" == *"/"* ]]; then
echo "..${str: -$((max-2))}"
else
echo "${str:0:$((max-1))}…"
fi
}
find_full_session_id() {
local short_id="$1"
if [ -f "$CODEMAN_STATE" ]; then
local full_id
full_id=$(jq -r --arg short "$short_id" '
.sessions | keys[] | select(startswith($short))
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$full_id" ]; then
echo "$full_id"
return
fi
fi
if [ -f "$CODEMAN_SESSIONS" ]; then
local full_id
full_id=$(jq -r --arg short "$short_id" '
.[] | select(.sessionId | startswith($short)) | .sessionId
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ -n "$full_id" ]; then
echo "$full_id"
return
fi
fi
echo "$short_id"
}
get_session_name() {
local session_id="$1"
local name=""
local workdir=""
if [ -f "$CODEMAN_SESSIONS" ]; then
local result
result=$(jq -r --arg id "$session_id" '
.[] | select(.sessionId | startswith($id)) | "\(.name // "")\t\(.workingDir // "")"
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ -n "$result" ]; then
name="${result%% *}"
workdir="${result#* }"
fi
fi
if [ -z "$name" ] && [ -f "$CODEMAN_STATE" ]; then
local result
result=$(jq -r --arg id "$session_id" '
.sessions | to_entries[] | select(.key | startswith($id)) | "\(.value.name // "")\t\(.value.workingDir // "")"
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$result" ]; then
name="${result%% *}"
[ -z "$workdir" ] && workdir="${result#* }"
fi
fi
if [ -n "$name" ]; then
echo "$name"
return
fi
if [ -n "$workdir" ]; then
echo "${workdir##*/}"
return
fi
echo "${session_id:0:8}"
}
get_working_dir() {
local session_id="$1"
if [ -f "$CODEMAN_SESSIONS" ]; then
local dir
dir=$(jq -r --arg id "$session_id" '
.[] | select(.sessionId | startswith($id)) | .workingDir // empty
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ -n "$dir" ] && [ "$dir" != "null" ]; then
echo "${dir/#$HOME/~}"
return
fi
fi
if [ -f "$CODEMAN_STATE" ]; then
local dir
dir=$(jq -r --arg id "$session_id" '
.sessions | to_entries[] | select(.key | startswith($id)) | .value.workingDir // empty
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$dir" ] && [ "$dir" != "null" ]; then
echo "${dir/#$HOME/~}"
return
fi
fi
echo ""
}
get_tokens() {
local session_id="$1"
if [ -f "$CODEMAN_STATE" ]; then
local tokens
tokens=$(jq -r --arg id "$session_id" '
.sessions | to_entries[] | select(.key | startswith($id)) |
((.value.inputTokens // 0) + (.value.outputTokens // 0))
' "$CODEMAN_STATE" 2>/dev/null | head -1)
if [ -n "$tokens" ] && [ "$tokens" != "null" ] && [ "$tokens" -gt 0 ] 2>/dev/null; then
if [ "$tokens" -gt 1000 ]; then
echo "$((tokens / 1000))k"
else
echo "${tokens}"
fi
return
fi
fi
echo ""
}
get_respawn_status() {
local session_id="$1"
if [ -f "$CODEMAN_SESSIONS" ]; then
local respawn_enabled
respawn_enabled=$(jq -r --arg id "$session_id" '
.[] | select(.sessionId | startswith($id)) | .respawnConfig.enabled // false
' "$CODEMAN_SESSIONS" 2>/dev/null | head -1)
if [ "$respawn_enabled" = "true" ]; then
echo "R"
return
fi
fi
echo ""
}
check_deps() {
if ! command -v jq &>/dev/null; then
echo -e "${YELLOW}Note: Install jq for session names${R}"
echo ""
fi
}
# ============================================================================
# Tmux Session Parser
# ============================================================================
declare -a SESSION_PIDS
declare -a MUX_NAMES
declare -a SESSION_STATES
declare -a SESSION_IDS
declare -a DISPLAY_NAMES
declare -a WORKING_DIRS
declare -a TOKEN_COUNTS
declare -a RESPAWN_STATUS
parse_sessions() {
SESSION_PIDS=()
MUX_NAMES=()
SESSION_STATES=()
SESSION_IDS=()
DISPLAY_NAMES=()
WORKING_DIRS=()
TOKEN_COUNTS=()
RESPAWN_STATUS=()
local i=0
# Parse tmux list-sessions output
while IFS= read -r line; do
local session_name="${line%%:*}"
# Only show codeman sessions
if [[ "$session_name" != codeman-* ]]; then
continue
fi
# Check if attached
local state="Detached"
if [[ "$line" == *"(attached)"* ]]; then
state="Attached"
fi
# Get PID from tmux
local pid
pid=$("${TMUX_CMD[@]}" display-message -t "$session_name" -p '#{pane_pid}' 2>/dev/null || echo "0")
SESSION_PIDS+=("$pid")
MUX_NAMES+=("$session_name")
SESSION_STATES+=("$state")
# Extract session ID from codeman session name
local session_id=""
local cm_regex='^codeman-(.+)$'
if [[ "$session_name" =~ $cm_regex ]]; then
session_id="${BASH_REMATCH[1]}"
fi
SESSION_IDS+=("$session_id")
# Get display name and metadata
if [ -n "$session_id" ]; then
DISPLAY_NAMES+=("$(get_session_name "$session_id")")
WORKING_DIRS+=("$(get_working_dir "$session_id")")
TOKEN_COUNTS+=("$(get_tokens "$session_id")")
RESPAWN_STATUS+=("$(get_respawn_status "$session_id")")
else
DISPLAY_NAMES+=("$session_name")
WORKING_DIRS+=("")
TOKEN_COUNTS+=("")
RESPAWN_STATUS+=("")
fi
i=$((i + 1))
done < <("${TMUX_CMD[@]}" list-sessions 2>/dev/null || true)
}
# ============================================================================
# Display Functions
# ============================================================================
clear_screen() {
printf '\033[2J\033[H'
}
print_header() {
local count=${#SESSION_PIDS[@]}
echo -e "${B}${CYAN}Codeman Sessions${R} ${D}($count)${R}"
echo -e "${D}$(printf '%.0s─' {1..32})${R}"
}
print_entry() {
local idx="$1"
local num=$((idx + 1))
local name="${DISPLAY_NAMES[$idx]}"
local state="${SESSION_STATES[$idx]}"
local dir="${WORKING_DIRS[$idx]}"
local tokens="${TOKEN_COUNTS[$idx]}"
local respawn="${RESPAWN_STATUS[$idx]}"
local name_max=$MAX_NAME_LEN
[ -n "$respawn" ] && name_max=$((name_max - 2))
[ -n "$tokens" ] && name_max=$((name_max - 4))
name=$(truncate "$name" $name_max)
local status_icon status_color
if [[ "$state" == *"Attached"* ]]; then
status_icon="$ICON_ATTACHED"
status_color="$GREEN"
elif [[ "$state" == *"Detached"* ]]; then
status_icon="$ICON_DETACHED"
status_color="$GRAY"
else
status_icon="$ICON_UNKNOWN"
status_color="$YELLOW"
fi
local num_str="${B}${WHITE}${num})${R}"
local name_str="${B}${WHITE}${name}${R}"
local status_str="${status_color}${status_icon}${R}"
local respawn_str=""
if [ -n "$respawn" ]; then
respawn_str=" ${GREEN}${respawn}${R}"
fi
local token_str=""
if [ -n "$tokens" ]; then
token_str=" ${D}${tokens}${R}"
fi
echo -e " ${num_str} ${name_str} ${status_str}${respawn_str}${token_str}"
if [ -n "$dir" ]; then
dir=$(truncate "$dir" $((MAX_NAME_LEN - 2)))
echo -e " ${D}${dir}${R}"
fi
}
print_footer() {
local page="$1"
local total_pages="$2"
echo ""
echo -e "${D}────────────────────────────────${R}"
if [ "$total_pages" -gt 1 ]; then
echo -e " ${D}Page $((page+1))/$total_pages${R} ${GRAY}[${WHITE}n${GRAY}]ext [${WHITE}p${GRAY}]rev${R}"
fi
echo -e " ${GRAY}[${WHITE}1-9${GRAY}]attach [${WHITE}r${GRAY}]efresh [${WHITE}q${GRAY}]uit${R}"
}
print_no_sessions() {
clear_screen
echo -e "${B}${CYAN}Codeman Sessions${R}"
echo -e "${D}$(printf '%.0s─' {1..32})${R}"
echo ""
echo -e " ${YELLOW}No tmux sessions found${R}"
echo ""
echo -e " ${D}Start one with:${R}"
echo -e " ${WHITE}codeman web${R}"
echo ""
echo -e "${D}$(printf '%.0s─' {1..32})${R}"
echo -e " ${GRAY}[${WHITE}r${GRAY}]efresh [${WHITE}q${GRAY}]uit${R}"
}
# ============================================================================
# Main Display Loop
# ============================================================================
current_page=0
render() {
clear_screen
parse_sessions
local count=${#SESSION_PIDS[@]}
if [ "$count" -eq 0 ]; then
print_no_sessions
return
fi
local total_pages=$(( (count + PAGE_SIZE - 1) / PAGE_SIZE ))
if [ "$current_page" -ge "$total_pages" ]; then
current_page=$((total_pages - 1))
fi
if [ "$current_page" -lt 0 ]; then
current_page=0
fi
local start=$((current_page * PAGE_SIZE))
local end=$((start + PAGE_SIZE))
if [ "$end" -gt "$count" ]; then
end=$count
fi
print_header
echo ""
for ((i = start; i < end; i++)); do
print_entry $i
done
print_footer $current_page $total_pages
}
attach_session() {
local idx="$1"
local mux_name="${MUX_NAMES[$idx]}"
if [ -z "$mux_name" ]; then
return 1
fi
clear_screen
echo -e "${GREEN}Attaching to ${B}${DISPLAY_NAMES[$idx]}${R}${GREEN}...${R}"
echo -e "${D}(Ctrl+B D to detach)${R}"
sleep 0.3
"${TMUX_CMD[@]}" attach-session -t "$mux_name"
return 0
}
# ============================================================================
# Input Handler
# ============================================================================
handle_input() {
local key="$1"
local count=${#SESSION_PIDS[@]}
local total_pages=$(( (count + PAGE_SIZE - 1) / PAGE_SIZE ))
case "$key" in
[1-9])
local idx=$((key - 1))
if [ "$idx" -lt "$count" ]; then
attach_session "$idx"
return 0
fi
;;
$'\e')
read -rsn2 -t 0.1 seq 2>/dev/null || true
case "$seq" in
'[A'|'[D')
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page - 1 + total_pages) % total_pages ))
fi
;;
'[B'|'[C')
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page + 1) % total_pages ))
fi
;;
esac
;;
n|N|j|J)
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page + 1) % total_pages ))
fi
;;
p|P|k|K)
if [ "$total_pages" -gt 1 ]; then
current_page=$(( (current_page - 1 + total_pages) % total_pages ))
fi
;;
r|R)
;;
q|Q)
clear_screen
exit 0
;;
'')
if [ "$count" -eq 1 ]; then
attach_session 0
return 0
fi
;;
esac
return 0
}
# ============================================================================
# List Mode
# ============================================================================
list_mode() {
parse_sessions
local count=${#SESSION_PIDS[@]}
if [ "$count" -eq 0 ]; then
echo "No tmux sessions"
exit 0
fi
for ((i = 0; i < count; i++)); do
local num=$((i + 1))
local name="${DISPLAY_NAMES[$i]}"
local state="${SESSION_STATES[$i]}"
local respawn="${RESPAWN_STATUS[$i]}"
local indicator="-"
[[ "$state" == *"Attached"* ]] && indicator="*"
[ -n "$respawn" ] && indicator="${indicator}R"
echo "$num) $name [$indicator]"
done
}
# ============================================================================
# Quick Attach
# ============================================================================
quick_attach() {
local num="$1"
parse_sessions
local count=${#SESSION_PIDS[@]}
local idx=$((num - 1))
if [ "$idx" -lt 0 ] || [ "$idx" -ge "$count" ]; then
echo -e "${RED}Invalid session: $num${R}"
echo "Available: 1-$count"
exit 1
fi
attach_session "$idx"
}
# ============================================================================
# Help
# ============================================================================
show_help() {
cat << 'EOF'
Codeman Sessions - Mobile-friendly Tmux Session Chooser
USAGE:
sc Interactive chooser
sc <number> Quick attach to session N
sc -l List sessions (non-interactive)
sc -h Show this help
INTERACTIVE KEYS:
1-9 Attach to session
n/j/↓ Next page
p/k/↑ Previous page
r Refresh
q Quit
INDICATORS:
* / ● Attached (someone connected)
- / ○ Detached (available)
R Respawn enabled
45k Token count
TIPS:
- Detach from tmux: Ctrl+B D
- Session names from Codeman state
- Optimized for Termius/iPhone
EOF
}
# ============================================================================
# Main
# ============================================================================
main() {
case "${1:-}" in
-h|--help)
show_help
exit 0
;;
-l|--list)
list_mode
exit 0
;;
[1-9]|[1-9][0-9])
quick_attach "$1"
exit $?
;;
esac
check_deps
render
while true; do
local timeout_opt=""
if [ "$AUTO_REFRESH" -gt 0 ]; then
timeout_opt="-t $AUTO_REFRESH"
fi
if read -rsn1 $timeout_opt key 2>/dev/null; then
handle_input "$key"
fi
render
done
}
trap 'clear_screen; exit 0' INT
main "$@"
+165 -44
View File
@@ -47,7 +47,7 @@ later call opens with, and your first REAL call performs them anyway:
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null
[ "${CODEMAN_PREAMBLE:-}" = 1.19.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.21.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
```
⚠️ **Never spend a Bash call on this check alone.** §1's block opens with this same
@@ -75,8 +75,8 @@ PRE="${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh"
mkdir -p "$(dirname "$PRE")"
# Rewrite unless the file already ends with THIS version's stamp, so a stale or a
# half-written file self-heals here instead of costing you a round trip to rm it.
grep -qs '^CODEMAN_PREAMBLE=1.19.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.19.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
grep -qs '^CODEMAN_PREAMBLE=1.21.0$' "$PRE" || (umask 077; cat > "$PRE" <<'PREAMBLE'
# ---- Codeman agent preamble 1.21.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -121,23 +121,90 @@ _composer_up() { # <sid> <timeoutMs> -> "true"/"false". `shift+tab` is the one
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
_dsh_up() { # <sid> <timeoutMs> -> "true"/"false". The DeepSeek Harness TUI's
# composer glyph. Override with DSH_READY_MARK for a profile that draws another one.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
--data-urlencode "match=${DSH_READY_MARK:-❯}" --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
# ---- the workspace-trust dialog: READ the screen, never press Enter blind ----
# Claude Code 2.1.252 dropped the option numbers, REVERSED them, and highlights
# "No, exit" by default:
# Security guide
# ❯ No, exit
# Yes, I trust this folder
# Enter to confirm . Esc to cancel
# so the bare \r that answered the old layout now answers *exit* and the pane is
# dead (`status 1`) seconds after the spawn -- measured on a live 2.1.252 case.
# These two read the rendered pane and steer onto the trust option instead.
_trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
# full=1 returns the RENDERED pane; a claude pane keeps no tmux history, so that
# is the current frame rather than every repaint since launch. tail -1 anyway,
# because the freshest marked row is the only one still true.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
k=$(_trust_key "$sid")
[ -n "$k" ] || return 1 # no dialog on screen, or a layout this cannot read
# A SEPARATE clientId for these keys. seq is monotonic per clientId, so
# spending prompt numbers here would make the next sendwait -- whose default
# seq is the epoch second -- look like a stale duplicate and vanish silently.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg k "$([ "$k" = confirm ] && printf '\r' || printf '\033[B')" \
--arg c "$CID-trust-$sid" --argjson s "$i" \
'{input:$k,useMux:true,clientId:$c,seq:$s}')" >/dev/null
[ "$k" = confirm ] && return 0
sleep 1; i=$((i+1)) # re-read: the arrow is CONFIRMED before Enter goes out
done
return 1
}
# spawn_worker <caseName> [mode] -> session id on stdout, diagnostics on stderr.
# quick-start AND readiness in one call, with a strict contract: NON-EMPTY stdout means
# a READY claude worker in a hook-carrying case. Anything less is rc 1 with EMPTY
# stdout, and the half-spawned session is deleted here rather than handed back, because
# a worker that never drew its composer would eat the task prompt with its trust
# dialog. There is deliberately no pid poll: wait-output already blocks until the
# composer draws, and pid!=null proved startup, never readiness.
# a READY worker whose end-of-turn signal can be trusted -- a claude worker in a
# hook-carrying case, or a `deepseek` worker whose harness TUI drew its composer.
# Anything less is rc 1 with EMPTY stdout, and the half-spawned session is deleted here
# rather than handed back, because a worker that never drew its composer would eat the
# task prompt with its trust dialog. There is deliberately no pid poll: wait-output
# already blocks until the composer draws, and pid!=null proved startup, never readiness.
spawn_worker() {
local name="${1:?spawn_worker needs a case name}" mode="${2:-claude}" q sid cp r
# parentSessionId doubles the CURL header, so a spawn_worker copied off the shared
# curl (or a body someone rebuilt from this recipe) still carries its lineage.
# deepseek: ask for the same permission posture the Run button sends, because the
# harness's own default (`workspace-write`) still ASKS, and a worker that stops on
# an approval row is a worker no fan-out can finish. It is not an escalation --
# claude workers already spawn with permissions skipped, and in multi-user mode the
# server clamps this back to `workspace-write` for an owner without the grant.
# Spawn by hand (§5.1) when you want a worker that asks.
q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" '{caseName:$n,mode:$m,parentSessionId:$p}')")
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" \
'{caseName:$n,mode:$m,parentSessionId:$p}
+ (if $m == "deepseek" then {deepSeekConfig:{permissionMode:"danger-full-access"}} else {} end)')")
sid=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$q")
# NOT retryable in a loop: every quick-start failure code is terminal (§5.1).
[ -n "$sid" ] || { jq -c '{error,errorCode}' <<<"$q" >&2; return 1; }
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # only claude draws a composer
if [ "$mode" = deepseek ]; then
# The one non-claude mode with REAL end-of-turn signals: its TUI reports
# idle/working/blocked to Codeman, so sendwait, until=stop and the Approvals
# Inbox all work here exactly as they do for claude. No hook file to vet
# (the bridge is env-injected, not a workspace file) and no trust dialog.
# ⚠️ Readiness is still not optional, and NOT interchangeable with the stop
# signal: the harness's boot report lands ~300ms BEFORE the composer paints
# (measured 2.26s vs 2.56s after spawn), so a sendwait fired straight after
# quick-start returns on that BOOT signal, reports a turn that never ran, and
# strands the prompt in a pane that was not yet taking input.
r=$(_dsh_up "$sid" 45000)
[ "$r" = true ] || { echo "dsh worker $sid never drew a composer: no pane-capable profile, a profile whose composer is not '${DSH_READY_MARK:-❯}' (set DSH_READY_MARK), or a harness that failed to boot -- check GET /api/v1/deepseek/status. Deleted it" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"; return 0
fi
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # no other mode draws a composer to wait on
# The server installs hooks into every claude workspace now, so this grep normally
# passes; it stays because the install is gated on a setting the operator can turn
# off, remote sessions never get hooks, and a session created by an older server
@@ -147,38 +214,41 @@ spawn_worker() {
grep -qs '/api/hook-event' "$cp/.claude/settings.local.json" || {
echo "case '$name' resolved to '$cp', which has no Codeman hooks (workspaceHooksEnabled off, remote, or an older server?): turn the setting on, or work §5.1+§5.5 by hand with markers" >&2
delete_session "$sid" >/dev/null; return 1; }
# Short composer wait FIRST, then the trust-dialog probe: a case still showing the
# dialog can never pass the composer wait, so probing early keeps a cold case from
# Short composer wait FIRST, then the trust dialog: a case still showing the
# dialog can never pass the composer wait, so acting early keeps a cold case from
# paying the whole long wait before the fallback even runs (§5.2). A warm case
# matches in under a second and never reaches the probe.
# matches in under a second and never reaches it, and _accept_trust returns in a
# blink when there is no dialog, so this costs nothing in the ordinary slow case.
r=$(_composer_up "$sid" 5000)
if [ "$r" != true ]; then
if "${CURL[@]}" -G "$API/api/v1/sessions/$sid/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000' \
| jq -e '.data.wait.matched' >/dev/null; then
# Codeman's own auto-accept gives up after 90 s / 3 tries; this is that bounded fallback.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" '{input:"\r",useMux:true,clientId:$c,seq:1}')" >/dev/null
fi
# Codeman answers this dialog itself and normally wins the race; this is the
# bounded fallback for when its 90 s window / 6-keystroke cap has run out.
_accept_trust "$sid"
r=$(_composer_up "$sid" 45000)
fi
[ "$r" = true ] || { echo "worker $sid never drew a composer; deleted it. Retry by hand via the §5.2 ladder (its billed stage-4 probe included)" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"
}
# spawn_workers <caseName>... -> one "<caseName> <sessionId>" line per worker, in order;
# the sessionId column is EMPTY for a spawn that failed (stderr has why). CONCURRENT:
# N workers cost about what one costs. Spawning them one Bash call at a time is the
# single biggest avoidable delay in this skill. Names must be UNIQUE: two workers in
# one case directory co-edit the same tree (§4), so a repeat is an error here, not a race.
# spawn_workers <caseName[:mode]>... -> one "<caseName> <sessionId>" line per worker, in
# order; the sessionId column is EMPTY for a spawn that failed (stderr has why).
# CONCURRENT: N workers cost about what one costs. Spawning them one Bash call at a time
# is the single biggest avoidable delay in this skill. A bare name is a claude worker;
# `beta:deepseek` makes that one a DeepSeek Harness worker, and a mixed fleet is one
# call. Case names must be UNIQUE: two workers in one case directory co-edit the same
# tree (§4), so a repeat is an error here, not a race (the mode never disambiguates two
# workers, since they would still share the directory).
spawn_workers() {
local d n i=0
local d spec n m i=0
[ "$#" -gt 0 ] || { echo "spawn_workers: no case names given" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sed 's/:.*//' | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
d=$(mktemp -d "${TMPDIR:-/tmp}/codeman-spawn.XXXXXX") || return 1
for n in "$@"; do ( spawn_worker "$n" > "$d/$i" ) & i=$((i+1)); done
for spec in "$@"; do
n=${spec%%:*}; m=${spec#*:}; [ "$m" = "$spec" ] && m=claude
( spawn_worker "$n" "$m" > "$d/$i" ) & i=$((i+1))
done
wait
i=0; for n in "$@"; do printf '%s %s\n' "$n" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
i=0; for spec in "$@"; do printf '%s %s\n' "${spec%%:*}" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
rm -rf "$d"
}
# sendwait <sid> <prompt> [seq] -> blocks until that worker's turn ENDS (~10 min ceiling
@@ -194,27 +264,46 @@ spawn_workers() {
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy only for a claude
# worker spawn_worker handed back (hooks vetted); hook-less workspaces and other modes
# resolve on flapping idle: markers instead (§5.5).
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
# implement the status contract is the one case that LOOKS like claude but is not:
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
# is still answering, and the re-wait below then resolved in 0 ms with
# `signal:"idle"` on a turn that had another three minutes to run (measured).
# A wait named after the end of a turn should only end with the turn, or with
# the worker. ⚠️ This is also what makes a wrong mode LOUD: the modes that
# cannot deliver `stop` answer 400 (before writing anything) instead of
# resolving on a flap, which is the answer that sends you to markers (§5.5).
body=$(jq -nc --arg p "$p" --arg c "$CID-$sid" --argjson s "$seq" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:20000}')
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:"stop,exit",waitTimeout:20000}')
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")")
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
fi
printf '%s\n' "$r"
}
# last_text <sid> [prev] -> that worker's last assistant message. Polled, because the
# transcript write LAGS the stop signal, and "some text exists" is not "THIS turn's
# text exists": right after a SECOND turn on the same worker the endpoint still serves
# last_text <sid> [prev] -> that worker's last assistant message (claude, codex and
# deepseek write a real transcript; the other modes have none, so read the terminal
# instead -- §5.4). Polled, because the transcript write LAGS the stop signal, and
# "some text exists" is not "THIS turn's text exists": right after a SECOND turn on the same worker the endpoint still serves
# the previous answer for a beat (observed live). When reading consecutive turns, pass
# the previous answer as [prev]: the poll then holds out for text that differs from it,
# falling back to whatever it last saw if the budget runs dry, so an honestly repeated
@@ -233,10 +322,10 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.19.0
CODEMAN_PREAMBLE=1.21.0
PREAMBLE
)
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.19.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
. "$PRE"; [ "${CODEMAN_PREAMBLE:-}" = 1.21.0 ] || { echo "preamble at $PRE is stale or truncated: rm it and re-run this block"; exit 1; }
```
Every later Bash call that touches the API starts with the same two loader lines from
@@ -287,8 +376,9 @@ and no per-call body to hand-build.
```bash
. "${XDG_CACHE_HOME:-$HOME/.cache}/codeman-agent-$CODEMAN_SESSION_ID.sh" 2>/dev/null # §0 loader
[ "${CODEMAN_PREAMBLE:-}" = 1.19.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
[ "${CODEMAN_PREAMBLE:-}" = 1.21.0 ] || { echo "preamble missing or stale; run the full §0 block"; exit 1; }
N=(alpha beta) # INVENT one fresh case name per worker; never list cases first
# (a name may carry a mode: `beta:deepseek`, see below)
T=('reply with one line: the absolute path of your working directory'
'reply with one line: your model name') # tasks, same order as N
@@ -353,6 +443,37 @@ Four things this block leans on, each one link away, no detour needed to run it:
- Each `sendwait` costs that worker one billed turn, as does every prompt you send it.
- Deleting the sessions does **not** remove the case directories: §5.14.
### DeepSeek Harness workers
The block above spawns claude workers. Any entry in `N` may instead name a mode
(`beta:deepseek`), and **a `deepseek` worker is driven by the same four verbs, with no
change to the rest of the block**: `spawn_workers` waits for its composer, `sendwait`
blocks on its real end-of-turn signal, `last_text` reads its answer, `delete_session`
removes it.
That is true of no other non-claude mode, and it is worth knowing why: the DeepSeek
Harness TUI reports `idle`/`working`/`blocked` to Codeman over the supervisor contract it
implements, so dsh is the one external CLI with definitive `stop`/`blocked` signals
instead of guessed-from-silence ones — and it writes a structured transcript, which is
what `last-response` reads for it. `shell`, `opencode`, `codex`, `gemini`, `antigravity`,
`pi`, `grok` and `omp` have neither and still need markers ([§5.5](reference/verbs.md#55-markers-for-hook-less-workers)).
Three things to know before you spawn one:
- **It needs a pane-capable profile.** `dsh` ships only `web`/`headless`, so the terminal
agent is always an installed profile. `GET /api/v1/deepseek/status` answers both
questions separately (`available` = the binary, `runnable` = a profile that can drive a
pane); a spawn without one fails with `OPERATION_FAILED` rather than falling back.
- **Do not task it on the strength of a `stop` alone.** The harness reports `idle` at
boot ~300 ms *before* its composer paints (measured 2.26 s vs 2.56 s), so a `sendwait`
fired straight after `quick-start` resolves on that boot signal, reports a turn that
never ran, and leaves the prompt in a pane that was not yet taking input. Letting
`spawn_worker` gate on readiness is what steps past that edge; it is not optional.
- **A profile that does not implement the contract looks like a hang.** Codeman cannot
know at spawn time whether one does. The tell is a `sendwait` that times out on a
worker whose pane clearly finished: that profile is one of them, so drive it with
markers instead.
## 2. What do you want to do?
One row per job. Acting on this table alone is correct; the §5 links are the detail.
@@ -360,10 +481,10 @@ One row per job. Acting on this table alone is correct; the §5 links are the de
| I want to | Call | Detail |
|-----------|------|--------|
| start a worker **where the work is** | `POST /api/v1/quick-start {"caseName":…}`, which **creates** `~/codeman-cases/<name>` unless the name is already a case. Any other path (a git worktree): `POST /api/v1/sessions {"workingDir":…}` then `POST /api/v1/sessions/:id/interactive`. Both install hooks by default, so expect full signals in either, and **verify** rather than assume. N workers means N worktrees | [§5.1](reference/verbs.md#51-where-to-spawn) |
| know a new worker can accept a prompt | `GET .../wait-output?match=shift+tab&from=buffer` (urlencode the `+`) | [§5.2](reference/verbs.md#52-readiness) |
| deliver a task **and** know when it finished | `POST .../input` with `"input":"…\r"`, `clientId`, `seq`, `"wait":true`. Resolves on `stop`, so it is trustworthy only where the workspace **has hooks** (claude mode; installed by default, but the operator can disable it and remote sessions never get them). Costs the worker one billed turn | [§5.3](reference/verbs.md#53-send-a-task-and-wait) |
| know a new worker can accept a prompt | `GET .../wait-output?match=shift+tab&from=buffer` (urlencode the `+`); a `deepseek` worker draws `❯` instead, and its boot `stop` fires ~300 ms BEFORE that, so never read the signal as readiness | [§5.2](reference/verbs.md#52-readiness) |
| deliver a task **and** know when it finished | `POST .../input` with `"input":"…\r"`, `clientId`, `seq`, `"wait":true`. Resolves on `stop`, so it is trustworthy where the signal is real: claude mode with hooks (installed by default, but the operator can disable it and remote sessions never get them) and `deepseek` mode through its status bridge. Costs the worker one billed turn | [§5.3](reference/verbs.md#53-send-a-task-and-wait) |
| know a hook-less worker finished | it has no `stop`, and `wait:true` there resolves on flapping `idle` **without erroring**: make it print a split, unique marker and `wait-output` on that instead | [§5.5](reference/verbs.md#55-markers-for-hook-less-workers) |
| read the answer | `GET .../last-response`, **polled** (claude/codex only; empty for the other modes) | [§5.4](reference/verbs.md#54-read-the-answer) |
| read the answer | `GET .../last-response`, **polled** (claude, codex and deepseek write a transcript; empty for the other modes) | [§5.4](reference/verbs.md#54-read-the-answer) |
| know if it is alive | `GET .../wait?until=exit&timeout=1000`: an immediate `signal:"exit"` means dead. `status` and `pid` both lie | [§5.6](reference/verbs.md#56-alive-and-stuck) |
| know if it is stuck | `GET .../active-tools` and `GET .../run-summary` are structured and free; two `terminal?tail=` samples are the crude fallback | [§5.6](reference/verbs.md#56-alive-and-stuck) |
| make a runaway worker stop | `POST .../input {"input":"\u001b"}` (ESC, **no** `\r`). Deleting the session would destroy the conversation instead | [§5.7](reference/verbs.md#57-interrupt-without-destroying) |
@@ -464,7 +585,7 @@ these**; open the one row you actually hit.
| [5.1 Where to spawn](reference/verbs.md#51-where-to-spawn) | the work is **not** a fresh scratch case: a linked case, a git worktree, any path that already existed. Hooks are absent there, which silently breaks send-and-wait. The costliest mistake in this skill |
| [5.2 Readiness](reference/verbs.md#52-readiness) | a worker never drew its composer, or you need the trust-dialog ladder by hand |
| [5.3 Send a task and wait](reference/verbs.md#53-send-a-task-and-wait) | the `sendwait` body, its signals, and the duplicate-resend loop |
| [5.4 Read the answer](reference/verbs.md#54-read-the-answer) | `last_text` came back empty, or the mode is not claude/codex |
| [5.4 Read the answer](reference/verbs.md#54-read-the-answer) | `last_text` came back empty, or the mode is not claude/codex/deepseek |
| [5.5 Markers for hook-less workers](reference/verbs.md#55-markers-for-hook-less-workers) | the worker has no `stop` hook: synchronize on a split, unique printed marker |
| [5.6 Alive and stuck](reference/verbs.md#56-alive-and-stuck) | is it dead or just slow? `status` and `pid` both lie |
| [5.7 Interrupt without destroying](reference/verbs.md#57-interrupt-without-destroying) | a runaway worker you want to stop but keep |
+125 -36
View File
@@ -1,4 +1,4 @@
# ---- Codeman agent preamble 1.19.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
# ---- Codeman agent preamble 1.21.0 (seeded by Codeman at session spawn; the SKILL.md §0 bootstrap rewrites it when missing or stale) ----
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set; refusing to guess}"
SELF="${CODEMAN_SESSION_ID:?CODEMAN_SESSION_ID not set}"
# Credentials, cheapest first. Your session has usually INHERITED the server's
@@ -43,23 +43,90 @@ _composer_up() { # <sid> <timeoutMs> -> "true"/"false". `shift+tab` is the one
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
_dsh_up() { # <sid> <timeoutMs> -> "true"/"false". The DeepSeek Harness TUI's
# composer glyph. Override with DSH_READY_MARK for a profile that draws another one.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/wait-output" \
--data-urlencode "match=${DSH_READY_MARK:-❯}" --data-urlencode 'from=buffer' \
--data-urlencode "timeout=$2" | jq -r '.data.wait.matched // false'
}
# ---- the workspace-trust dialog: READ the screen, never press Enter blind ----
# Claude Code 2.1.252 dropped the option numbers, REVERSED them, and highlights
# "No, exit" by default:
# Security guide
# ❯ No, exit
# Yes, I trust this folder
# Enter to confirm . Esc to cancel
# so the bare \r that answered the old layout now answers *exit* and the pane is
# dead (`status 1`) seconds after the spawn -- measured on a live 2.1.252 case.
# These two read the rendered pane and steer onto the trust option instead.
_trust_key() { # <sid> -> "confirm" | "move" | "" (nothing safe to press)
# full=1 returns the RENDERED pane; a claude pane keeps no tmux history, so that
# is the current frame rather than every repaint since launch. tail -1 anyway,
# because the freshest marked row is the only one still true.
"${CURL[@]}" -G "$API/api/v1/sessions/$1/terminal" --data-urlencode 'full=1' \
| jq -r '.data.terminalBuffer // empty' \
| sed -e "s/$(printf '\033')\[[0-9;?]*[a-zA-Z]//g" -e "s/$(printf '\033')[()][AB0]//g" \
| tr -d ' \t' | grep -i '❯[0-9.]*\(yes,itrustthisfolder\|no,exit\)' | tail -1 \
| sed -e 's/.*[Yy]es,.*/confirm/' -e 's/.*[Nn]o,.*/move/'
}
_accept_trust() { # <sid> -> 0 once it has answered the dialog, 1 if it could not
local sid="$1" k i=1
while [ "$i" -le 6 ]; do
k=$(_trust_key "$sid")
[ -n "$k" ] || return 1 # no dialog on screen, or a layout this cannot read
# A SEPARATE clientId for these keys. seq is monotonic per clientId, so
# spending prompt numbers here would make the next sendwait -- whose default
# seq is the epoch second -- look like a stale duplicate and vanish silently.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg k "$([ "$k" = confirm ] && printf '\r' || printf '\033[B')" \
--arg c "$CID-trust-$sid" --argjson s "$i" \
'{input:$k,useMux:true,clientId:$c,seq:$s}')" >/dev/null
[ "$k" = confirm ] && return 0
sleep 1; i=$((i+1)) # re-read: the arrow is CONFIRMED before Enter goes out
done
return 1
}
# spawn_worker <caseName> [mode] -> session id on stdout, diagnostics on stderr.
# quick-start AND readiness in one call, with a strict contract: NON-EMPTY stdout means
# a READY claude worker in a hook-carrying case. Anything less is rc 1 with EMPTY
# stdout, and the half-spawned session is deleted here rather than handed back, because
# a worker that never drew its composer would eat the task prompt with its trust
# dialog. There is deliberately no pid poll: wait-output already blocks until the
# composer draws, and pid!=null proved startup, never readiness.
# a READY worker whose end-of-turn signal can be trusted -- a claude worker in a
# hook-carrying case, or a `deepseek` worker whose harness TUI drew its composer.
# Anything less is rc 1 with EMPTY stdout, and the half-spawned session is deleted here
# rather than handed back, because a worker that never drew its composer would eat the
# task prompt with its trust dialog. There is deliberately no pid poll: wait-output
# already blocks until the composer draws, and pid!=null proved startup, never readiness.
spawn_worker() {
local name="${1:?spawn_worker needs a case name}" mode="${2:-claude}" q sid cp r
# parentSessionId doubles the CURL header, so a spawn_worker copied off the shared
# curl (or a body someone rebuilt from this recipe) still carries its lineage.
# deepseek: ask for the same permission posture the Run button sends, because the
# harness's own default (`workspace-write`) still ASKS, and a worker that stops on
# an approval row is a worker no fan-out can finish. It is not an escalation --
# claude workers already spawn with permissions skipped, and in multi-user mode the
# server clamps this back to `workspace-write` for an owner without the grant.
# Spawn by hand (§5.1) when you want a worker that asks.
q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" '{caseName:$n,mode:$m,parentSessionId:$p}')")
-d "$(jq -nc --arg n "$name" --arg m "$mode" --arg p "$SELF" \
'{caseName:$n,mode:$m,parentSessionId:$p}
+ (if $m == "deepseek" then {deepSeekConfig:{permissionMode:"danger-full-access"}} else {} end)')")
sid=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$q")
# NOT retryable in a loop: every quick-start failure code is terminal (§5.1).
[ -n "$sid" ] || { jq -c '{error,errorCode}' <<<"$q" >&2; return 1; }
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # only claude draws a composer
if [ "$mode" = deepseek ]; then
# The one non-claude mode with REAL end-of-turn signals: its TUI reports
# idle/working/blocked to Codeman, so sendwait, until=stop and the Approvals
# Inbox all work here exactly as they do for claude. No hook file to vet
# (the bridge is env-injected, not a workspace file) and no trust dialog.
# ⚠️ Readiness is still not optional, and NOT interchangeable with the stop
# signal: the harness's boot report lands ~300ms BEFORE the composer paints
# (measured 2.26s vs 2.56s after spawn), so a sendwait fired straight after
# quick-start returns on that BOOT signal, reports a turn that never ran, and
# strands the prompt in a pane that was not yet taking input.
r=$(_dsh_up "$sid" 45000)
[ "$r" = true ] || { echo "dsh worker $sid never drew a composer: no pane-capable profile, a profile whose composer is not '${DSH_READY_MARK:-❯}' (set DSH_READY_MARK), or a harness that failed to boot -- check GET /api/v1/deepseek/status. Deleted it" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"; return 0
fi
[ "$mode" = claude ] || { printf '%s\n' "$sid"; return 0; } # no other mode draws a composer to wait on
# The server installs hooks into every claude workspace now, so this grep normally
# passes; it stays because the install is gated on a setting the operator can turn
# off, remote sessions never get hooks, and a session created by an older server
@@ -69,38 +136,41 @@ spawn_worker() {
grep -qs '/api/hook-event' "$cp/.claude/settings.local.json" || {
echo "case '$name' resolved to '$cp', which has no Codeman hooks (workspaceHooksEnabled off, remote, or an older server?): turn the setting on, or work §5.1+§5.5 by hand with markers" >&2
delete_session "$sid" >/dev/null; return 1; }
# Short composer wait FIRST, then the trust-dialog probe: a case still showing the
# dialog can never pass the composer wait, so probing early keeps a cold case from
# Short composer wait FIRST, then the trust dialog: a case still showing the
# dialog can never pass the composer wait, so acting early keeps a cold case from
# paying the whole long wait before the fallback even runs (§5.2). A warm case
# matches in under a second and never reaches the probe.
# matches in under a second and never reaches it, and _accept_trust returns in a
# blink when there is no dialog, so this costs nothing in the ordinary slow case.
r=$(_composer_up "$sid" 5000)
if [ "$r" != true ]; then
if "${CURL[@]}" -G "$API/api/v1/sessions/$sid/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000' \
| jq -e '.data.wait.matched' >/dev/null; then
# Codeman's own auto-accept gives up after 90 s / 3 tries; this is that bounded fallback.
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" '{input:"\r",useMux:true,clientId:$c,seq:1}')" >/dev/null
fi
# Codeman answers this dialog itself and normally wins the race; this is the
# bounded fallback for when its 90 s window / 6-keystroke cap has run out.
_accept_trust "$sid"
r=$(_composer_up "$sid" 45000)
fi
[ "$r" = true ] || { echo "worker $sid never drew a composer; deleted it. Retry by hand via the §5.2 ladder (its billed stage-4 probe included)" >&2
delete_session "$sid" >/dev/null; return 1; }
printf '%s\n' "$sid"
}
# spawn_workers <caseName>... -> one "<caseName> <sessionId>" line per worker, in order;
# the sessionId column is EMPTY for a spawn that failed (stderr has why). CONCURRENT:
# N workers cost about what one costs. Spawning them one Bash call at a time is the
# single biggest avoidable delay in this skill. Names must be UNIQUE: two workers in
# one case directory co-edit the same tree (§4), so a repeat is an error here, not a race.
# spawn_workers <caseName[:mode]>... -> one "<caseName> <sessionId>" line per worker, in
# order; the sessionId column is EMPTY for a spawn that failed (stderr has why).
# CONCURRENT: N workers cost about what one costs. Spawning them one Bash call at a time
# is the single biggest avoidable delay in this skill. A bare name is a claude worker;
# `beta:deepseek` makes that one a DeepSeek Harness worker, and a mixed fleet is one
# call. Case names must be UNIQUE: two workers in one case directory co-edit the same
# tree (§4), so a repeat is an error here, not a race (the mode never disambiguates two
# workers, since they would still share the directory).
spawn_workers() {
local d n i=0
local d spec n m i=0
[ "$#" -gt 0 ] || { echo "spawn_workers: no case names given" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
[ -z "$(printf '%s\n' "$@" | sed 's/:.*//' | sort | uniq -d)" ] || { echo "spawn_workers: duplicate case names" >&2; return 1; }
d=$(mktemp -d "${TMPDIR:-/tmp}/codeman-spawn.XXXXXX") || return 1
for n in "$@"; do ( spawn_worker "$n" > "$d/$i" ) & i=$((i+1)); done
for spec in "$@"; do
n=${spec%%:*}; m=${spec#*:}; [ "$m" = "$spec" ] && m=claude
( spawn_worker "$n" "$m" > "$d/$i" ) & i=$((i+1))
done
wait
i=0; for n in "$@"; do printf '%s %s\n' "$n" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
i=0; for spec in "$@"; do printf '%s %s\n' "${spec%%:*}" "$(cat "$d/$i" 2>/dev/null)"; i=$((i+1)); done
rm -rf "$d"
}
# sendwait <sid> <prompt> [seq] -> blocks until that worker's turn ENDS (~10 min ceiling
@@ -116,27 +186,46 @@ spawn_workers() {
# (observed live). So the first wait is short; on its timeout a bare \r goes out (the
# missing Enter when the prompt is stranded, a no-op when the turn is genuinely
# running), then the ORIGINAL frame is resent unchanged, which the server takes as a
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy only for a claude
# worker spawn_worker handed back (hooks vetted); hook-less workspaces and other modes
# resolve on flapping idle: markers instead (§5.5).
# tagged duplicate: it re-waits without retyping (§5.3). Trustworthy for a worker
# spawn_worker handed back -- claude (hooks vetted) or deepseek (status bridge) --
# and for those only. Hook-less workspaces and the other modes resolve on flapping
# idle: markers instead (§5.5). ⚠️ A dsh worker running a profile that does not
# implement the status contract is the one case that LOOKS like claude but is not:
# it accepts the send and then burns both waits. One timeout on a dsh worker whose
# pane clearly finished means that profile, so switch that worker to markers.
sendwait() {
local sid="${1:?}" p="${2:?}" seq="${3:-$(date +%s)}" body r
# `wait:"stop,exit"`, never the `wait:true` default set: that set also carries
# `idle`, which is INFERRED from output stabilization and flaps mid-turn. On a
# dsh worker whose TUI repaints rarely the session reads `idle` while the model
# is still answering, and the re-wait below then resolved in 0 ms with
# `signal:"idle"` on a turn that had another three minutes to run (measured).
# A wait named after the end of a turn should only end with the turn, or with
# the worker. ⚠️ This is also what makes a wrong mode LOUD: the modes that
# cannot deliver `stop` answer 400 (before writing anything) instead of
# resolving on a flap, which is the answer that sends you to markers (§5.5).
body=$(jq -nc --arg p "$p" --arg c "$CID-$sid" --argjson s "$seq" \
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:true,waitTimeout:20000}')
'{input:($p+"\r"),useMux:true,clientId:$c,seq:$s,wait:"stop,exit",waitTimeout:20000}')
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$body")
if jq -e '.data.delivered and .data.wait.timedOut' <<<"$r" >/dev/null 2>&1; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" -H 'Content-Type: application/json' \
-d "$(jq -nc --arg c "$CID-$sid" --argjson s "$(date +%s)" \
'{input:"\r",useMux:true,clientId:$c,seq:$s}')" >/dev/null
# The resend is a tagged DUPLICATE, so the server skips the write and reports
# `delivered:false` for it -- truthfully, but about the wrong send. The first
# one delivered, so carry that forward, or §1's cleanup reads a completed turn
# as an undelivered one and keeps a finished worker forever.
r=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$sid/input" \
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")")
-H 'Content-Type: application/json' --data-binary "$(jq -c '.waitTimeout=580000' <<<"$body")" \
| jq -c 'if .success and (.data.wait.ended | not) then .data.delivered = true else . end')
fi
printf '%s\n' "$r"
}
# last_text <sid> [prev] -> that worker's last assistant message. Polled, because the
# transcript write LAGS the stop signal, and "some text exists" is not "THIS turn's
# text exists": right after a SECOND turn on the same worker the endpoint still serves
# last_text <sid> [prev] -> that worker's last assistant message (claude, codex and
# deepseek write a real transcript; the other modes have none, so read the terminal
# instead -- §5.4). Polled, because the transcript write LAGS the stop signal, and
# "some text exists" is not "THIS turn's text exists": right after a SECOND turn on the same worker the endpoint still serves
# the previous answer for a beat (observed live). When reading consecutive turns, pass
# the previous answer as [prev]: the poll then holds out for text that differs from it,
# falling back to whatever it last saw if the budget runs dry, so an honestly repeated
@@ -155,4 +244,4 @@ last_text() {
# The stamp is the LAST line on purpose (a truncated write leaves it unset) and is kept
# bare on purpose: the write condition above anchors on it with $, so an inline comment
# here would fail that match and rewrite this file on every single bootstrap.
CODEMAN_PREAMBLE=1.19.0
CODEMAN_PREAMBLE=1.21.0
+29 -19
View File
@@ -237,7 +237,10 @@ minutes, never retry the credential.
flushed slightly *after* the `stop` hook fires, so a read taken the instant the wait
returns is too early (verified live: empty on the first call, full prose seconds later).
It is also `""` before the worker's first completed turn, and permanently `""` for
`shell`, `opencode`, `gemini`, `antigravity` and `pi`, which write no Claude transcript.
`shell`, `opencode`, `gemini`, `antigravity`, `pi`, `grok` and `omp`, which write no transcript at
all. `deepseek` is NOT one of those — it is read from `$DSH_HOME/sessions/**` and lags
for the same reason claude does (the harness finalizes the assistant message just after
it reports `idle`), so poll it the same way.
**Fix** Poll it, bounded (10 tries, 1 s apart). If it is still empty on a hook-less mode,
that is expected, not a failure: read `terminal?tail=` and strip ANSI instead.
@@ -279,7 +282,7 @@ than into an existing checkout.
| start case + session in one call | `POST /api/v1/quick-start` |
| create a session in an arbitrary directory (no case, **no PTY**, id at `.data.session.id`) | `POST /api/v1/sessions`, then `POST /api/v1/sessions/:id/interactive` or `.../shell` to start it, see [Starting a worker](#starting-a-worker) |
| send input | `POST /api/v1/sessions/:id/input` |
| **read a worker's answer** (claude/codex) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}`, clean transcript text, no TUI noise. ⚠️ **Poll it**, see [symptom 7](#7-last-response-returns-an-empty-string-right-after-stop) |
| **read a worker's answer** (claude/codex/deepseek) | `GET /api/v1/sessions/:id/last-response` → `.data.{text,timestamp}`, clean transcript text, no TUI noise. ⚠️ **Poll it**, see [symptom 7](#7-last-response-returns-an-empty-string-right-after-stop) |
| read terminal (tail is in **BYTES**, raw ANSI) | `GET /api/v1/sessions/:id/terminal?tail=3000` → `.data.terminalBuffer`, for *diagnosis* (unsubmitted prompt?), not for reading answers |
| full tmux scrollback (context bomb; post-mortems only) | `GET /api/v1/sessions/:id/terminal?full=1` |
| background agents, one session | `GET /api/v1/sessions/:id/subagents` |
@@ -321,7 +324,7 @@ on signals and markers for exactly this reason.
⚠️ `GET /api/v1/sessions/:id/output` → `.data.textOutput` looks like the obvious read
but stays **empty for interactive tmux-backed sessions** (it is fed only by the legacy
JSON-stream path). Verified empty on live claude and shell sessions. Use
`last-response` for claude/codex answers; only fall back to `terminal?tail=` for
`last-response` for claude/codex/deepseek answers; only fall back to `terminal?tail=` for
hook-less modes, or to diagnose a prompt that was never submitted, and strip ANSI:
```bash
@@ -336,19 +339,20 @@ ESC=$(printf '\033')
`POST /api/v1/quick-start` body (all optional):
`{"caseName":"worker-1","mode":"claude","sessionName":"w9-worker","effort":"high"}`
, `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity|pi`; response is
, `mode` ∈ `claude|shell|opencode|codex|gemini|antigravity|pi|grok|deepseek|omp`; response is
`.data.{sessionId, caseName, casePath}`. Creates the case directory (a real directory
on the user's disk) if missing, do not retry it in a loop, and remember the name.
⚠️ A `mode` whose CLI is **not installed on the server** fails the spawn with
`OPERATION_FAILED`; it never falls back to claude. Probe first whenever you did not pick
the mode yourself: `GET /api/v1/claude/status`, `GET /api/v1/opencode/status`,
`GET /api/v1/codex/status`, `GET /api/v1/gemini/status`, `GET /api/v1/antigravity/status`
and `GET /api/v1/pi/status` each return `.data.{available, path}` (no session needed).
Pi's also carries `.data.version`, because `pi` is a short generic name that an unrelated
binary on `$PATH` can shadow: the resolver rejects one whose `--version` is not
semver-shaped, so `available:false` there can mean "a different `pi` is in front" rather
than "nothing is installed". `shell` has no CLI to probe.
`GET /api/v1/codex/status`, `GET /api/v1/gemini/status`, `GET /api/v1/antigravity/status`, `GET /api/v1/grok/status`, `GET /api/v1/deepseek/status`,
`GET /api/v1/pi/status` and `GET /api/v1/omp/status` each return `.data.{available, path}` (no session needed).
Pi's, grok's and OMP's also carry `.data.version`, because `pi` is a short generic name,
`grok` is a name with npm squatters, and `omp` is a similarly short name, so an unrelated
binary on `$PATH` can shadow any of them: the resolver rejects one whose `--version` is
not version-shaped, so `available:false` there can mean "a different program of the same
name is in front" rather than "nothing is installed". `shell` has no CLI to probe.
⚠️ **Branch on `.success` before reading `.data.sessionId`.** On any failure the field
is absent, `jq -r` prints the literal string `null`, and every later call then targets
@@ -462,10 +466,10 @@ Quirks that will bite you:
session answers with an empty timeline rather than a 404.
- ⚠️ **`active-tools` proves presence, never absence.** It is fed by the BashToolParser,
which reads Claude's rendered `● Bash(…)` lines, and `_processExpensiveParsers`
returns early for every external CLI mode (`session.ts:2136`), so it is permanently
`[]` on `opencode`/`codex`/`gemini`/`antigravity`/`pi`. ⚠️ **`shell` is NOT one of those**
(`isExternalCliMode`, `session.ts:165-167`, lists only those five), so the parser does
run on a shell worker, and `TEXT_COMMAND_PATTERN` (`bash-tool-parser.ts:88`) matches
returns early for every external CLI mode (`session.ts:~2225`), so it is permanently
`[]` on `opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`deepseek`/`omp`. ⚠️ **`shell` is NOT one of those**
(`isExternalCliMode`, `session.ts:176-187`, lists only those seven), so the parser does
run on a shell worker, and `TEXT_COMMAND_PATTERN` (`bash-tool-parser.ts:89`) matches
bare `tail|cat|head|less|grep|watch|multitail <path>` lines with no `● Bash(` wrapper:
a shell worker running `cat build.log` really does populate this. In practice it stays
empty for most shell work. It also never sees non-Bash
@@ -639,10 +643,14 @@ block, so a linked case or a raw `workingDir` had no hooks at all. `POST
session-create path installs hooks regardless of how the directory got there. See
[symptom 8](#8-send-and-wait-resolves-instantly-with-signalidle-and-the-answer-is-last-turns).
Default `until` set: `stop,idle,exit`. On non-claude modes the server silently drops
`stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
Default `until` set: `stop,idle,exit`. On modes with no hook signals the server silently
drops `stop`/`blocked` from the *default* set (echoed back as `wait.until`, e.g.
`["idle","exit"]` on shell); requesting them *explicitly* there is a 400 naming the
mode. ⚠️ That 400 is about **mode**, so a hooks-less *claude* session accepts
mode. ⚠️ `deepseek` is not one of those: its harness reports its own lifecycle, so it
keeps the full default set and accepts an explicit `until=stop`. ⚠️ For dsh the answer is
per-SESSION rather than per-mode — a session created with `statusReporting: false` has no
bridge, and an explicit `until=stop` there is a 400 naming that setting. ⚠️ That 400 is
otherwise about **mode**, so a hooks-less *claude* session accepts
`until=stop` happily and then never resolves it. ⚠️ On hook-less modes the lifecycle
signals are also **coarse in practice**: a
short shell command produced **no** `idle` transition within 60 s (verified live), so
@@ -790,8 +798,10 @@ for environment and setup problems.
| `CODEMAN_MUX` unset but you seem to be in a session | remote-SSH case: the env vars are not exported there. Fail closed, refuse to act |
| connection refused from inside a container | a loopback-bound server is unreachable from a container, and `CODEMAN_DOCKER_BRIDGE_HOOKS=1` does **not** fix that: it opens a hooks-only listener, so hook events start flowing but `/api/v1/*` stays refused. Driving the API from inside a Docker case needs a reachable bind (an operator decision); report it, don't retry |
| wait routes 404 on a valid session id | read the `.error` text: a `Route ...` prefix means the server predates the wait endpoints (< 1.13.0; a dev build can serve them while reporting an older version, so probe, never version-compare), poll `terminal?tail=` and say so. `Session ... not found` means your id is wrong, not the server |
| wait on `stop` never resolves | non-claude mode, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept did not fire (it is bounded by a 90 s window and an attempt cap); use the readiness recipe in SKILL.md, wait for `shift+tab` first, accept the dialog only as the bounded fallback |
| wait on `stop` never resolves | a mode with no hook signals, or hooks not reaching the server (Docker/remote), or a case created by Codeman < 1.13.0 against an `--https` install (its hook curls lacked `-k` and TLS-failed silently; a 1.13.0+ server rewrites them the next time a session starts in that case). Use markers or `idle,exit` |
| wait on `stop` never resolves, on a **dsh** worker whose pane clearly finished | that profile does not implement the harness's supervisor contract, which Codeman cannot detect at request time (an unrecognized profile is treated as launchable on purpose). The wait is accepted and then times out. Drive that worker with markers, or switch to a profile that reports — `@deepseek-harness-tui/dsh-tui` does |
| new claude worker ignores its first prompt | it was showing the first-run trust dialog and Codeman's auto-accept did not fire (it is bounded by a 90 s window and a keystroke cap); use the readiness recipe in SKILL.md, wait for `shift+tab` first, answer the dialog only as the bounded fallback |
| a brand-new claude worker's pane is DEAD (`status 1`) seconds after the spawn | something pressed Enter at the first-run trust dialog. Since claude-cli 2.1.252 its options are unnumbered, reversed, and the highlighted default is `No, exit`, so a blind `\r` — an up-front Enter, or a task prompt typed into the dialog — quits the CLI. Answer it by reading the `❯` marker off `terminal?full=1` and arrowing onto `Yes, I trust this folder` first: `_accept_trust` in the §0 preamble |
| readiness burns its whole budget, then the worker answers fine anyway | you matched `bypass`, which is the statusline of ONE permission mode. Codeman spawns `--dangerously-skip-permissions` by default, but the server's `claudeMode` setting also has `auto` (`auto mode on`), `allowedTools` and `normal` (both `don't ask on`), and the effective per-session value is not exposed on `GET /api/v1/sessions/:id`. Match **`shift+tab`** instead: every mode's status bar ends `(shift+tab to cycle)` (measured per mode against claude-cli 2.1.226). Expect `blocked` signals mid-turn on the non-default modes |
| ANSI escapes survive the strip pipeline | `sed -e 's/\x1b…'` on macOS: `\x1b` is GNU-only, BSD sed matches nothing and strips nothing. Use the `ESC=$(printf '\033')` form above |
| `wait-output` times out although the pane shows the text | multi-word match against a TUI screen; the stream has no spaces there, match one token |
+2 -2
View File
@@ -56,7 +56,7 @@ own head: the worker enforcing the cap is the one who has to be told about it.
| synchronize on end of turn | HTTP `wait until=stop` (fires for message-initiated turns too, verified live) |
| liveness / death check | HTTP `wait?until=exit` |
| interrupt a running turn (break-glass) | HTTP input, a bare `\x1b` with no `\r` |
| non-claude modes (`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`) | HTTP only (no other CLI has messaging) |
| non-claude modes (`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`deepseek`/`omp`) | HTTP only (no other CLI has messaging) |
| delete | HTTP, via SKILL.md's `delete_session` guard |
## Availability: probe, never assume
@@ -347,7 +347,7 @@ Without a break-glass, a pair with a bad brief is a token bonfire with no off sw
### Mixed fleets: the pairing matrix
Non-claude workers (`shell`, `opencode`, `codex`, `gemini`, `antigravity`, `pi`) cannot be peers
Non-claude workers (`shell`, `opencode`, `codex`, `gemini`, `antigravity`, `pi`, `grok`, `deepseek`, `omp`) cannot be peers
at all; no other CLI has this feature. Their tasks route over HTTP, and you never mention
messaging in their briefs. The claude half of the fleet can use messaging among itself,
subject to the namespace rule: **messaging works between two sessions that share one
+70 -16
View File
@@ -69,9 +69,14 @@ SEQ=1 # $CID is the fixed literal from the preamble; never rebuild
# plus a two-marker screen match in session-trust-dialog.ts), not the output stream.
# It still misses two ways, and both leave the dialog up until someone answers it:
# it only scans in the first 90 s after the pane started (TRUST_DIALOG_WINDOW_MS),
# and it gives up after 3 Enter presses (TRUST_DIALOG_MAX_ATTEMPTS). So: composer
# marker first, dialog only as the bounded fallback (a blind Enter up front would
# land in an already-ready composer).
# and it gives up after 6 keystrokes (TRUST_DIALOG_MAX_ATTEMPTS). So: composer
# marker first, dialog only as the bounded fallback.
# ⚠️ The dialog is NOT answered with Enter. Since claude-cli 2.1.252 the options
# lost their numbers, swapped places, and the highlighted one is `No, exit`, so a
# blind \r quits the CLI and the pane is dead seconds after the spawn (measured).
# _accept_trust (§0 preamble) reads the ❯ marker off the rendered pane, arrows onto
# `Yes, I trust this folder`, re-reads to confirm the move landed, and only then
# presses Enter.
# Stage 1 is SHORT on purpose: an already-trusted case matches in <1 s, while a
# virgin case can never pass it (the dialog is up) and always pays it in full,
# the long budget belongs to stage 3, after the dialog is answered.
@@ -92,13 +97,8 @@ done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
_accept_trust "$SID" # reads the marker and steers; never a blind \r. Own clientId,
# so it spends none of $SEQ's numbers.
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
fi
@@ -106,8 +106,8 @@ if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# stage 4, mode-agnostic and bounded: answering a trivial prompt IS readiness.
# COSTS THE WORKER ONE BILLED TURN, so it only runs when the fast marker missed.
# Split token (the typed line echoes into the stream) and unique per call. Must stay
# AFTER the dialog fallback: free text plus \r into a trust dialog still up answers
# it blind, the same footgun as an up-front Enter.
# AFTER the dialog fallback: the select widget swallows the text and the \r answers
# whatever is highlighted, which on a live dialog is `No, exit`.
TOK="${RANDOM}_$$"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"reply with the word READY immediately followed by _'"$TOK"' and nothing else\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
@@ -188,7 +188,7 @@ for _ in $(seq 1 10); do
done
printf '%s\n' "$TXT"
# (.data is {text,timestamp}; text is also "" before the first completed turn and
# always "" for shell/opencode/gemini/antigravity/pi, which have no transcript, use
# always "" for shell/opencode/gemini/antigravity/pi/grok/omp, which have no transcript, use
# the terminal tail there, and here only to diagnose an unsubmitted prompt.)
# 6. clean up: exact id, own list only, through the fail-closed preamble helper
@@ -198,6 +198,59 @@ delete_session "$SID"
Increment `SEQ` for every *new* input to the same worker. Reuse the same `SEQ` only to
re-ask about the same delivery (the duplicate-wait loop above).
## Flow 1b: DeepSeek Harness worker, end to end
A `deepseek` worker is driven with the same four verbs as a claude one, because the
harness reports its own lifecycle: its `stop` is a real end-of-turn signal, and its
answer comes from a real transcript. The differences are all at the edges.
```bash
# 0. Is there anything to spawn? `available` is the binary, `runnable` is a profile
# that can drive a pane -- dsh ships only web/headless, so the two differ.
"${CURL[@]}" "$API/api/v1/deepseek/status" | jq -c '{available:.data.available,runnable:.data.runnable,profile:.data.defaultProfile}'
# 1. Spawn. `deepSeekConfig` is optional: an absent profile picks the first
# pane-capable one, and an absent permissionMode leaves the harness on its own
# workspace-write default, which still ASKS before it acts.
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"dsh-worker","mode":"deepseek","deepSeekConfig":{"permissionMode":"danger-full-access"}}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
[ -n "$SID" ] || { jq -c '{error, errorCode}' <<<"$Q"; exit 1; } # OPERATION_FAILED = no runnable profile
CREATED+=("$SID")
# 2. Readiness, and ONLY readiness. ⚠️ Do not use the stop signal for this: the
# harness reports idle at BOOT, ~300 ms before the composer paints.
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=❯' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' \
| jq -e '.data.wait.matched' >/dev/null || { echo "no composer"; delete_session "$SID"; exit 1; }
# 3. Task it. Identical to a claude worker, including the \r and the (clientId, seq).
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"Read calc.py and tell me in one sentence whether add() is correct.\r","useMux":true,"clientId":"codeman-dsh-1","seq":1,"wait":"stop,exit","waitTimeout":300000}' \
| jq -c '{delivered:.data.delivered,signal:.data.wait.signal,timedOut:.data.wait.timedOut}'
# 4. Read it. From $DSH_HOME/sessions/**, not the pane -- scraping a dsh pane returns
# its ASCII-art splash. Poll: the harness finalizes the message just after it
# reports idle. Two answers are not the model's words and say so:
# "Turn error: …" (the provider or harness failed) and "Turn ended: …" (early stop).
for _ in $(seq 1 15); do
TXT=$("${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# 5. Full conversation, if you need the tool calls too:
# "${CURL[@]}" "$API/api/v1/sessions/$SID/last-response?context=full" | jq -r '.data.messages[]|"[\(.label)] \(.text)"'
delete_session "$SID"
```
⚠️ **`wait:"stop,exit"`, not `wait:true`.** The default set also carries `idle`, which
for an external CLI is inferred from output stabilization: a dsh TUI that repaints
rarely reads as idle mid-turn, and a wait carrying `idle` then resolves in 0 ms on a
turn with minutes left to run (measured). The same reason the preamble's `sendwait`
asks for `stop,exit` on every mode.
## Flow 2: shell worker, marker-synchronized
`shell` sessions have no hooks (`stop`/`blocked` are a 400 there), and their lifecycle
@@ -474,9 +527,10 @@ done
circuit breaker, which exists to stop a worker that crashes on every start from being
restarted in a loop; clearing it unasked re-arms that loop.
- Then run **Flow 1's readiness stages 1-3** on each SID. A path claude has never been
run in shows the trust dialog, and typing your task into a dialog answers it blind and
loses the task. Stages 1-3 cost no turn; stage 4, if it fires, costs that worker one
billed turn.
run in shows the trust dialog, and typing your task into it does not just lose the
task: the select widget swallows the text and the trailing `\r` answers the
highlighted option, which since claude-cli 2.1.252 is `No, exit`. Stages 1-3 cost no
turn; stage 4, if it fires, costs that worker one billed turn.
### 4. Hand out the tasks: markers, not send-and-wait
+71 -24
View File
@@ -152,17 +152,46 @@ It is **decoration, and resolved rather than trusted**, so treat it accordingly:
### 5.2 Readiness
A new session reports `idle` before its CLI has spawned, and a brand-new case shows a
**dsh workers first**, because their trap is the opposite of claude's: they have no
trust dialog and boot straight into a composer (`❯`, matched `from=buffer`), but the
harness reports `idle` — which reaches you as a `stop` signal — about 300 ms BEFORE that
composer paints (measured 2.26 s vs 2.56 s after spawn, twice). So the signal that means
"this worker finished its turn" is also the first thing it emits at boot, and a
send-and-wait fired straight after `quick-start` resolves on it, reports a turn that
never ran, and leaves the prompt in a pane that was not yet taking input. Wait for the
composer, not for the signal; `spawn_worker` does exactly that, and by the time it
returns the boot edge is spent (signals are edge-triggered, so nothing can catch it
later). A profile whose composer is not `❯` needs `DSH_READY_MARK` set to whatever it
does draw.
For claude: a new session reports `idle` before its CLI has spawned, and a brand-new case shows a
**trust dialog** first, so neither "wait for idle" nor "wait for ❯" means ready (the
trust dialog contains `❯` too, observed live). Codeman auto-accepts that dialog
itself, reliably enough that stage 1 usually just works: `_maybeAcceptTrustDialog()`
reads the **rendered pane** via `capturePaneText()` rather than the arriving chunk
(the per-chunk `includes()` version could never match, because tmux repaints the row
with cursor-forward escapes in place of spaces, and it is documented in-source as the
historical bug). The remaining miss modes are structural: the auto-accept only runs
inside a 90 s window after interactive start and gives up after 3 attempts. So keep
the dialog handling as a bounded fallback, and never send a blind Enter up front (if
auto-accept already fired, it lands in the composer).
historical bug).
⚠️ **The answer is no longer "press Enter".** Claude Code 2.1.252 dropped the option
numbers, reversed the two options, and highlights the one that quits:
```
❯ No, exit
Yes, I trust this folder
Enter to confirm · Esc to cancel
```
so a blind `\r` answers *exit*: the pane is dead (`Pane is dead (status 1)`) about six
seconds after the spawn, measured on a fresh case. Read the marker off the rendered
pane (`GET .../terminal?full=1`), send `ESC [ B` while it sits on `No, exit`, re-read,
and press Enter only once the marker is on the trust option. `_accept_trust` in the
§0 preamble is exactly that, and `trustDialogNextKey()` is the server-side twin.
The remaining miss modes are structural: the auto-accept only runs inside a 90 s window
after interactive start and gives up after 6 keystrokes. So keep the dialog handling as
a bounded fallback, and never send a blind Enter up front — landing in an already-ready
composer only wastes a turn, landing in this dialog ends the worker.
Stage 1 is short on purpose: an already-trusted case matches `shift+tab` in under a
second, while a case still showing the dialog cannot pass stage 1 at all and always
@@ -220,14 +249,11 @@ SEQ=1 # $CID came from the §0 preamble; do NOT rebuild it from $$
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# composer never appeared, so the trust dialog is probably still up; accept it once
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
if jq -e '.data.wait.matched' <<<"$T" >/dev/null; then
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
SEQ=$((SEQ+1))
fi
# Composer never appeared, so the trust dialog is probably still up. NEVER a blind
# Enter here: the highlighted option is "No, exit". _accept_trust (§0 preamble) reads
# the marker off the pane, arrows onto the trust option, re-reads, then confirms. It
# carries its OWN clientId, so it spends none of $SEQ's numbers.
_accept_trust "$SID"
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=shift+tab' --data-urlencode 'from=buffer' --data-urlencode 'timeout=45000')
fi
@@ -236,8 +262,10 @@ if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
# of a broken worker, and answering is proof that it works. Split the token (your
# keystrokes echo into the stream) and keep it unique per call. This costs the worker
# one billed turn, so it runs only after the fast path missed. It must stay AFTER
# stage 2, which is the only thing that clears the trust dialog: free text plus \r
# into a dialog still up answers it blind, the same footgun as the up-front Enter.
# stage 2, which is the only thing that clears the trust dialog: the typed text is
# swallowed by the select widget and the \r then answers whatever is highlighted,
# which since 2.1.252 is "No, exit" -- the same footgun as the up-front Enter, except
# that it kills the worker rather than wasting a turn.
TOK="${RANDOM}_$$"
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"reply with the word READY immediately followed by _'"$TOK"' and nothing else\r","useMux":true,"clientId":"'"$CID"'","seq":'$SEQ'}' >/dev/null
@@ -341,17 +369,35 @@ If the loop exhausts its cap, do not keep looping: read the terminal, report wha
see, and remember that a still-typed-but-unsubmitted prompt (missing `\r`) can only be
recovered by submitting it with `{"input":"\r"}`.
⚠️ `stop` and `blocked` fire for `claude` sessions only (they are Claude Code hooks,
and only when the workspace actually has them, see [§5.1](#51-where-to-spawn)). On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`, requesting them explicitly is a
⚠️ `stop` and `blocked` fire for `claude` sessions (they are Claude Code hooks, and
only when the workspace actually has them, see [§5.1](#51-where-to-spawn)) **and for
`deepseek`** — the one external CLI that reports its own lifecycle, so its `stop` is a
real end-of-turn signal rather than a guess. On
`shell`/`opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`omp`, requesting them explicitly is a
400, and lifecycle transitions there are coarse (a short shell command may emit **no**
`idle` transition at all, verified live), so synchronize those with markers.
⚠️ A dsh session can still refuse them for a per-SESSION reason: `statusReporting:
false` at create time disarms the bridge, and an explicit `until=stop` is then a 400
naming that setting. And a `stop` that is *accepted* is not proof it will ever fire —
whether the installed profile implements the supervisor contract cannot be known at
request time, so a non-conforming one accepts the wait and times out on it. One timeout
on a dsh worker whose pane clearly finished identifies that profile; switch it to
markers.
### 5.4 Read the answer
For `claude` and `codex` workers this is the read path: `last-response` returns the
agent's final message as clean text, taken from the transcript rather than the screen,
so it carries none of the TUI's box-drawing or repaint noise.
For `claude`, `codex` and `deepseek` workers this is the read path: `last-response`
returns the agent's final message as clean text, taken from the transcript rather than
the screen, so it carries none of the TUI's box-drawing or repaint noise.
⚠️ For `deepseek` it reads `$DSH_HOME/sessions/**`, and reading it is the ONLY way to
get that answer: dsh-TUI paints a full-screen splash, so scraping its pane returns the
ASCII-art logo (that is what `last-response` itself used to return for dsh). Two dsh
answers are not the model's words and say so: `Turn error: …` (the provider or the
harness failed the turn) and `Turn ended: …` (an early stop such as `max-tokens`). A
turn still streaming reads back as the partial answer so far, so a non-empty read is
not by itself proof the turn ended — that is what the `stop` signal is for.
```bash
for _ in $(seq 1 10); do # the transcript write LAGS the stop signal
@@ -369,9 +415,10 @@ from the transcript file, which is flushed slightly *after* the `stop` hook fire
single read taken the instant send-and-wait returns comes back `""` even though the
turn finished (verified live: empty on the first call, full text seconds later). `text`
is also `""` before the worker's first completed turn, and always `""` for modes with
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, `pi`; the first four
no transcript (`shell`, `opencode`, `gemini`, `antigravity`, `pi`, `grok`, `omp`; the first four
verified live, pi from the same source path), which is
why the loop above is bounded rather than open-ended. Fall back to the terminal buffer
why the loop above is bounded rather than open-ended. A dsh worker lags too, for its own
reason: the harness finalizes the assistant message just after it reports `idle`. Fall back to the terminal buffer
there, tail in **bytes** (`textOutput` in `GET .../output` stays empty for interactive
sessions; don't use it):
@@ -454,7 +501,7 @@ turn), and both better than diffing terminal samples:
```
⚠️ `active-tools` is parsed out of Claude's own output format, so it is **empty for
`opencode`/`codex`/`gemini`/`antigravity`/`pi`** (those parsers are skipped wholesale) and
`opencode`/`codex`/`gemini`/`antigravity`/`pi`/`grok`/`deepseek`/`omp`** (those parsers are skipped wholesale) and
in practice empty for `shell`. Source-verified, not measured live.
Only if neither helps: sample `terminal?tail=` twice a few seconds apart. A changing
+299
View File
@@ -0,0 +1,299 @@
/**
* @fileoverview One style vocabulary for everything the `codeman` CLI prints:
* palette, glyphs, the small block helpers (heading/rule/kv), width-aware table
* layout, a stderr spinner and a y/N confirm.
*
* Color detection is chalk's alone. It already honors NO_COLOR, FORCE_COLOR,
* TERM=dumb and TTY-ness, and a second detector here would disagree with it on
* some terminal with no way to tell which one was right.
*
* The layout math is pure and exported separately from anything that touches a
* terminal, which is what lets it be unit-tested with no TTY and reused by
* `utils/dependency-report.ts` while that file stays color-free.
*
* @module cli-style
*/
import chalk, { type ChalkInstance } from 'chalk';
import { createInterface } from 'node:readline';
// Direct import, not the `utils` barrel: the barrel pulls in node-pty and every
// CLI resolver, which a style module has no business loading.
import { stripAnsi } from './utils/regex-patterns.js';
// ─────────────────────────────────────────────────────────────────────────────
// Palette and glyphs
// ─────────────────────────────────────────────────────────────────────────────
/** Semantic roles, mirroring the web UI's status language (green fine, yellow waiting, red blocked). */
export const palette = {
ok: chalk.green,
warn: chalk.yellow,
err: chalk.red,
info: chalk.cyan,
muted: chalk.gray,
emph: chalk.bold,
accent: chalk.magenta,
} as const satisfies Record<string, ChalkInstance>;
/** The glyph vocabulary the CLI already used, in one place. */
export const GLYPH = {
ok: '✓',
fail: '✗',
warn: '⚠',
idle: '○',
dot: '●',
arrow: '→',
} as const;
/** Spinner frames (braille, one cell wide in every terminal we support). */
export const SPINNER_FRAMES = ['⠋', '⠙', '⠹', '⠸', '⠼', '⠴', '⠦', '⠧', '⠇', '⠏'] as const;
/** What a line is reporting, independent of how it is painted. */
export type Tone = 'ok' | 'warn' | 'err' | 'idle' | 'info';
const TONE_GLYPH: Record<Tone, string> = {
ok: GLYPH.ok,
warn: GLYPH.warn,
err: GLYPH.fail,
idle: GLYPH.idle,
info: GLYPH.dot,
};
const TONE_STYLE: Record<Tone, ChalkInstance> = {
ok: palette.ok,
warn: palette.warn,
err: palette.err,
idle: palette.muted,
info: palette.info,
};
/** Glyph for a tone. Pure, so the mapping is testable without a terminal. */
export function glyphFor(tone: Tone): string {
return TONE_GLYPH[tone];
}
/** Paint text in a tone's color. */
export function tint(tone: Tone, text: string): string {
return TONE_STYLE[tone](text);
}
// ─────────────────────────────────────────────────────────────────────────────
// Blocks
// ─────────────────────────────────────────────────────────────────────────────
/** Section heading. The blank line above it is part of the existing block idiom. */
export function heading(text: string): string {
return `\n${palette.emph(text)}`;
}
/** Horizontal rule under a title. */
export function rule(width = 40): string {
return palette.muted('─'.repeat(Math.max(0, width)));
}
/**
* Indented `Label: value` line. `pad` aligns the values of a block by padding
* the label column (including its colon), for blocks whose labels differ in
* length.
*/
export function kv(label: string, value: string, pad = 0): string {
const key = pad > 0 ? padCell(`${label}:`, pad) : `${label}:`;
return ` ${key} ${value}`;
}
// ─────────────────────────────────────────────────────────────────────────────
// Width-aware layout (pure)
// ─────────────────────────────────────────────────────────────────────────────
/** Printed width of a cell: ANSI sequences take no columns. */
export function displayWidth(text: string): number {
return stripAnsi(text).length;
}
export type CellAlign = 'left' | 'right';
/** Pad to `width` columns, measuring by display width so colored cells still align. */
export function padCell(text: string, width: number, align: CellAlign = 'left'): string {
const fill = ' '.repeat(Math.max(0, width - displayWidth(text)));
return align === 'right' ? `${fill}${text}` : `${text}${fill}`;
}
/**
* Pad AFTER the paint, so the fill stays outside the color run and a trailing
* empty column can be trimmed away instead of ending in a reset sequence with
* invisible spaces before it.
*/
export function padStyled(text: string, width: number, paint: (t: string) => string): string {
return `${paint(text)}${' '.repeat(Math.max(0, width - displayWidth(text)))}`;
}
/** Widest cell per column. Short rows count as empty cells, never as narrower columns. */
export function columnWidths(rows: readonly (readonly string[])[]): number[] {
const widths: number[] = [];
for (const row of rows) {
for (let i = 0; i < row.length; i++) {
widths[i] = Math.max(widths[i] ?? 0, displayWidth(row[i] ?? ''));
}
}
return widths;
}
export interface TableOptions {
/** Per-column alignment; missing entries are left-aligned. */
align?: readonly CellAlign[];
/** Spaces between columns. */
gap?: number;
/** Prefix for every row. */
indent?: string;
}
/**
* Lay rows out in columns sized to their widest cell. The last cell of a row is
* never padded, so no line carries trailing whitespace.
*/
export function layoutTable(rows: readonly (readonly string[])[], options: TableOptions = {}): string[] {
const { align = [], gap = 1, indent = '' } = options;
const widths = columnWidths(rows);
const separator = ' '.repeat(Math.max(0, gap));
return rows.map((row) => {
const cells = row.map((cell, i) => (i === row.length - 1 ? cell : padCell(cell, widths[i], align[i] ?? 'left')));
return `${indent}${cells.join(separator)}`;
});
}
/** `layoutTable()` as one printable block. */
export function table(rows: readonly (readonly string[])[], options: TableOptions = {}): string {
return layoutTable(rows, options).join('\n');
}
// ─────────────────────────────────────────────────────────────────────────────
// Spinner
// ─────────────────────────────────────────────────────────────────────────────
const HIDE_CURSOR = '\x1b[?25l';
const SHOW_CURSOR = '\x1b[?25h';
const CLEAR_LINE = '\x1b[K';
/** The slice of a stream a spinner needs; `process.stderr` satisfies it. */
export interface SpinnerStream {
isTTY?: boolean;
write(chunk: string): unknown;
}
export interface Spinner {
start(): Spinner;
/** Change the text mid-flight. Silent on a non-TTY, which prints once and stops. */
setText(text: string): void;
/** Clear the line, restore the cursor and optionally print a final line. */
stop(finalLine?: string): void;
}
export interface SpinnerOptions {
stream?: SpinnerStream;
intervalMs?: number;
}
/**
* In-place progress line on stderr, for the calls that block for tens of seconds
* (daemon start, service install). Only a TTY gets the animation: piped output
* and journald get the text once, so a log file never fills with `\r` frames.
*/
export function spinner(text: string, options: SpinnerOptions = {}): Spinner {
const stream = options.stream ?? process.stderr;
const intervalMs = options.intervalMs ?? 90;
const animated = Boolean(stream.isTTY);
let label = text;
let frame = 0;
let timer: NodeJS.Timeout | null = null;
let started = false;
let stopped = false;
const restoreCursor = () => {
if (animated) stream.write(SHOW_CURSOR);
};
const render = () => {
stream.write(`\r${palette.info(SPINNER_FRAMES[frame % SPINNER_FRAMES.length])} ${label}${CLEAR_LINE}`);
frame++;
};
const handle: Spinner = {
start() {
if (started || stopped) return handle;
started = true;
if (!animated) {
stream.write(`${label}\n`);
return handle;
}
stream.write(HIDE_CURSOR);
// A hidden cursor left behind by a Ctrl+C outlives the process, so the
// exit hook is not optional.
process.once('exit', restoreCursor);
render();
// Unref'd: a spinner must never be the reason the process stays alive.
timer = setInterval(render, intervalMs);
timer.unref();
return handle;
},
setText(next: string) {
label = next;
if (animated && started && !stopped) render();
},
stop(finalLine?: string) {
if (stopped) return;
stopped = true;
if (timer) {
clearInterval(timer);
timer = null;
}
if (animated && started) {
stream.write(`\r${CLEAR_LINE}`);
restoreCursor();
process.off('exit', restoreCursor);
}
if (finalLine && animated) stream.write(`${finalLine}\n`);
},
};
return handle;
}
/** Run `work` with a spinner up, stopping it however `work` ends. */
export async function withSpinner<T>(text: string, work: () => Promise<T>, options?: SpinnerOptions): Promise<T> {
const handle = spinner(text, options).start();
try {
return await work();
} finally {
handle.stop();
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Confirm
// ─────────────────────────────────────────────────────────────────────────────
/** Is there a human on the other end of both halves of the terminal? */
export function isInteractive(): boolean {
return Boolean(process.stdin.isTTY && process.stdout.isTTY);
}
/**
* y/N prompt. Answers `false` immediately when stdin is not a TTY (a script
* piping into the CLI must never hang on an invisible question), so callers
* that support a `--force` flag can branch on `isInteractive()` to keep printing
* their "pass --force" hint instead.
*/
export async function confirm(question: string): Promise<boolean> {
if (!isInteractive()) return false;
const rl = createInterface({ input: process.stdin, output: process.stdout });
try {
const answer = await new Promise<string>((resolve) => {
rl.once('SIGINT', () => resolve(''));
rl.question(`${question} ${palette.muted('[y/N]')} `, resolve);
});
return /^y(es)?$/i.test(answer.trim());
} finally {
rl.close();
// readline resumes stdin; a still-flowing stdin keeps the process alive.
process.stdin.pause();
}
}
+277 -210
View File
@@ -8,7 +8,6 @@
*/
import { Command } from 'commander';
import chalk from 'chalk';
import { createRequire } from 'module';
import http from 'node:http';
import https from 'node:https';
@@ -16,6 +15,7 @@ import { existsSync, readFileSync } from 'node:fs';
import { isAbsolute, join } from 'node:path';
import { homedir } from 'node:os';
import { dataPath } from './config/instance.js';
import { casePath } from './config/cases-dir.js';
import { installAgentSkillInto, removeAgentSkillFrom, type AgentSkillApplyResult } from './hooks-config.js';
import { getSessionManager } from './session-manager.js';
import { getTaskQueue } from './task-queue.js';
@@ -26,6 +26,9 @@ import { isSupportedAttachmentExtension } from './attachment-registry.js';
import { daemonStatus, startDaemon, stopDaemon, type WebLaunchOptions } from './daemon-control.js';
import { installService, serviceStatus, uninstallService } from './service-installer.js';
import { isLoopbackBindHost, isUnauthenticatedNetworkAcknowledged } from './web/network-auth-policy.js';
import { confirm, heading, isInteractive, kv, palette, rule, tint, withSpinner, type Tone } from './cli-style.js';
import type { ToolResult } from './utils/dependency-checker.js';
import type { ReportStyle } from './utils/dependency-report.js';
const require = createRequire(import.meta.url);
const pkg = require('../package.json') as { version: string };
@@ -107,14 +110,14 @@ program
.action(async (filePath, options) => {
const extension = String(filePath).split('.').pop()?.toLowerCase() || '';
if (!isAbsolute(filePath) || !isSupportedAttachmentExtension(extension)) {
console.error(chalk.red('✗ attach requires an absolute path to a png, pdf, docx, pptx, md, or txt file'));
console.error(palette.err('✗ attach requires an absolute path to a png, pdf, docx, pptx, md, or txt file'));
process.exit(1);
}
const sessionId = options.session || process.env.CODEMAN_SESSION_ID;
const apiUrl = options.url || process.env.CODEMAN_API_URL || 'https://127.0.0.1:3000';
if (sessionId && (await postAttachment(apiUrl, sessionId, filePath))) {
console.log(chalk.green('✓ Attachment card requested'));
console.log(palette.ok('✓ Attachment card requested'));
return;
}
@@ -144,7 +147,9 @@ export function resolveCliCasePath(name: string): string {
} catch {
// no registry yet, or unreadable/invalid JSON: fall through to the cases dir
}
return join(homedir(), 'codeman-cases', name);
// Same resolver the server uses, so CODEMAN_CASES_PATH (Docker Compose) moves
// the CLI's idea of a case with it instead of leaving it on the home default.
return casePath(name);
}
/**
@@ -174,7 +179,7 @@ export function resolveSkillTargetPath(options: {
function resolveSkillTarget(options: { case?: string }): string {
const resolved = resolveSkillTargetPath(options);
if (resolved.missingCase !== undefined) {
console.error(chalk.red(`✗ Case not found: ${resolved.missingCase}`));
console.error(palette.err(`✗ Case not found: ${resolved.missingCase}`));
process.exit(1);
}
return resolved.target;
@@ -199,9 +204,9 @@ function reportSkillResult(result: AgentSkillApplyResult, target: string): void
};
const message = messages[result];
if (message.ok) {
console.log(chalk.green(`✓ ${message.text}`));
console.log(palette.ok(`✓ ${message.text}`));
} else {
console.error(chalk.red(`✗ ${message.text}`));
console.error(palette.err(`✗ ${message.text}`));
process.exit(1);
}
}
@@ -220,7 +225,7 @@ skillCmd
const target = resolveSkillTarget(options);
reportSkillResult(await installAgentSkillInto(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to install agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -235,7 +240,7 @@ skillCmd
const target = resolveSkillTarget(options);
reportSkillResult(await removeAgentSkillFrom(target), target);
} catch (err) {
console.error(chalk.red(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to remove agent skill: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -252,11 +257,11 @@ sessionCmd
try {
const manager = getSessionManager();
const session = await manager.createSession(options.dir);
console.log(chalk.green(`✓ Session started: ${session.id}`));
console.log(palette.ok(`✓ Session started: ${session.id}`));
console.log(` Working directory: ${session.workingDir}`);
console.log(` PID: ${session.pid}`);
} catch (err) {
console.error(chalk.red(`✗ Failed to start session: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to start session: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -268,70 +273,81 @@ sessionCmd
try {
const manager = getSessionManager();
await manager.stopSession(id);
console.log(chalk.green(`✓ Session stopped: ${id}`));
console.log(palette.ok(`✓ Session stopped: ${id}`));
} catch (err) {
console.error(chalk.red(`✗ Failed to stop session: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to stop session: ${getErrorMessage(err)}`));
process.exit(1);
}
});
/** Session status in the shared vocabulary: idle is fine, busy is working, anything else is a problem. */
function sessionStatusLabel(status: string): string {
if (status === 'idle') return palette.ok('idle');
if (status === 'busy') return palette.warn('busy');
return palette.err(status);
}
/**
* The one session listing. `codeman list` used to be a copy of this that had
* drifted (it lost the stopped and web-server sections), so it now calls the
* same renderer and only opts out of those two sections.
*/
function printSessionList(options: { includeStored: boolean }): void {
const manager = getSessionManager();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
if (sessions.length === 0 && Object.keys(stored).length === 0) {
console.log(palette.warn('No sessions found'));
return;
}
console.log(heading('Active Sessions:'));
if (sessions.length === 0) {
console.log(' (none)');
} else {
for (const session of sessions) {
console.log(
` ${palette.info(session.id.slice(0, 8))} ${sessionStatusLabel(session.status)} ${session.workingDir}`
);
}
}
if (options.includeStored) {
const stoppedSessions = Object.values(stored).filter((s) => s.status === 'stopped');
if (stoppedSessions.length > 0) {
console.log(heading('Stopped Sessions:'));
for (const session of stoppedSessions) {
const name = session.name ? ` (${session.name})` : '';
console.log(
` ${palette.muted(session.id.slice(0, 8))} ${palette.muted('stopped')}${name} ${session.workingDir}`
);
}
}
// Sessions the web server owns: this process has no PTY for them, so they
// only exist in the shared state file.
const activeSessions = Object.values(stored).filter((s) => s.status !== 'stopped');
if (sessions.length === 0 && activeSessions.length > 0) {
console.log(heading('Active Sessions (from web server):'));
for (const session of activeSessions) {
const name = session.name ? ` (${session.name})` : '';
const mode = session.mode === 'shell' ? palette.muted(' [shell]') : '';
const cost = session.totalCost ? palette.muted(` $${session.totalCost.toFixed(4)}`) : '';
console.log(
` ${palette.info(session.id.slice(0, 8))} ${sessionStatusLabel(session.status)}${name}${mode}${cost} ${session.workingDir}`
);
}
}
}
console.log('');
}
sessionCmd
.command('list')
.alias('ls')
.description('List all sessions')
.action(() => {
const manager = getSessionManager();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
if (sessions.length === 0 && Object.keys(stored).length === 0) {
console.log(chalk.yellow('No sessions found'));
return;
}
console.log(chalk.bold('\nActive Sessions:'));
if (sessions.length === 0) {
console.log(' (none)');
} else {
for (const session of sessions) {
const status =
session.status === 'idle'
? chalk.green('idle')
: session.status === 'busy'
? chalk.yellow('busy')
: chalk.red(session.status);
console.log(` ${chalk.cyan(session.id.slice(0, 8))} ${status} ${session.workingDir}`);
}
}
const stoppedSessions = Object.values(stored).filter((s) => s.status === 'stopped');
if (stoppedSessions.length > 0) {
console.log(chalk.bold('\nStopped Sessions:'));
for (const session of stoppedSessions) {
const name = session.name ? ` (${session.name})` : '';
console.log(` ${chalk.gray(session.id.slice(0, 8))} ${chalk.gray('stopped')}${name} ${session.workingDir}`);
}
}
// Show active sessions from state (when web server manages them)
const activeSessions = Object.values(stored).filter((s) => s.status !== 'stopped');
if (sessions.length === 0 && activeSessions.length > 0) {
console.log(chalk.bold('\nActive Sessions (from web server):'));
for (const session of activeSessions) {
const status =
session.status === 'idle'
? chalk.green('idle')
: session.status === 'busy'
? chalk.yellow('busy')
: chalk.red(session.status);
const name = session.name ? ` (${session.name})` : '';
const mode = session.mode === 'shell' ? chalk.gray(' [shell]') : '';
const cost = session.totalCost ? chalk.gray(` $${session.totalCost.toFixed(4)}`) : '';
console.log(` ${chalk.cyan(session.id.slice(0, 8))} ${status}${name}${mode}${cost} ${session.workingDir}`);
}
}
console.log('');
});
.action(() => printSessionList({ includeStored: true }));
sessionCmd
.command('logs <id>')
@@ -342,12 +358,12 @@ sessionCmd
const output = options.errors ? manager.getSessionError(id) : manager.getSessionOutput(id);
if (output === null) {
console.log(chalk.yellow(`Session ${id} not found or not active`));
console.log(palette.warn(`Session ${id} not found or not active`));
return;
}
if (output === '') {
console.log(chalk.gray('(no output)'));
console.log(palette.muted('(no output)'));
return;
}
@@ -374,7 +390,7 @@ taskCmd
completionPhrase: options.completion,
timeoutMs: options.timeout ? parseInt(options.timeout, 10) : undefined,
});
console.log(chalk.green(`✓ Task added: ${task.id}`));
console.log(palette.ok(`✓ Task added: ${task.id}`));
console.log(` Prompt: ${prompt.slice(0, 50)}${prompt.length > 50 ? '...' : ''}`);
console.log(` Priority: ${task.priority}`);
});
@@ -393,26 +409,28 @@ taskCmd
}
if (tasks.length === 0) {
console.log(chalk.yellow('No tasks found'));
console.log(palette.warn('No tasks found'));
return;
}
const statusColors = {
pending: chalk.gray,
running: chalk.yellow,
completed: chalk.green,
failed: chalk.red,
pending: palette.muted,
running: palette.warn,
completed: palette.ok,
failed: palette.err,
};
console.log(chalk.bold('\nTasks:'));
console.log(palette.emph('\nTasks:'));
for (const task of tasks) {
const color = statusColors[task.status];
const prompt = task.prompt.slice(0, 40) + (task.prompt.length > 40 ? '...' : '');
console.log(` ${chalk.cyan(task.id.slice(0, 8))} ${color(task.status.padEnd(10))} [${task.priority}] ${prompt}`);
console.log(
` ${palette.info(task.id.slice(0, 8))} ${color(task.status.padEnd(10))} [${task.priority}] ${prompt}`
);
}
const counts = queue.getCount();
console.log(chalk.bold('\nSummary:'));
console.log(palette.emph('\nSummary:'));
console.log(
` Pending: ${counts.pending}, Running: ${counts.running}, Completed: ${counts.completed}, Failed: ${counts.failed}`
);
@@ -427,11 +445,11 @@ taskCmd
const task = queue.getTask(id);
if (!task) {
console.log(chalk.red(`Task ${id} not found`));
console.log(palette.err(`Task ${id} not found`));
return;
}
console.log(chalk.bold('\nTask Details:'));
console.log(palette.emph('\nTask Details:'));
console.log(` ID: ${task.id}`);
console.log(` Status: ${task.status}`);
console.log(` Priority: ${task.priority}`);
@@ -441,10 +459,10 @@ taskCmd
console.log(` Session: ${task.assignedSessionId}`);
}
if (task.error) {
console.log(` Error: ${chalk.red(task.error)}`);
console.log(` Error: ${palette.err(task.error)}`);
}
if (task.output) {
console.log(chalk.bold('\nOutput:'));
console.log(palette.emph('\nOutput:'));
console.log(task.output.slice(0, 500) + (task.output.length > 500 ? '...' : ''));
}
console.log('');
@@ -457,9 +475,9 @@ taskCmd
.action((id) => {
const queue = getTaskQueue();
if (queue.removeTask(id)) {
console.log(chalk.green(`✓ Task removed: ${id}`));
console.log(palette.ok(`✓ Task removed: ${id}`));
} else {
console.log(chalk.red(`Task ${id} not found`));
console.log(palette.err(`Task ${id} not found`));
}
});
@@ -474,13 +492,13 @@ taskCmd
if (options.all) {
count = queue.clearAll();
console.log(chalk.green(`✓ Cleared ${count} tasks`));
console.log(palette.ok(`✓ Cleared ${count} tasks`));
} else if (options.failed) {
count = queue.clearFailed();
console.log(chalk.green(`✓ Cleared ${count} failed tasks`));
console.log(palette.ok(`✓ Cleared ${count} failed tasks`));
} else {
count = queue.clearCompleted();
console.log(chalk.green(`✓ Cleared ${count} completed tasks`));
console.log(palette.ok(`✓ Cleared ${count} completed tasks`));
}
});
@@ -503,38 +521,38 @@ ralphCmd
}
if (loop.isRunning()) {
console.log(chalk.yellow('Ralph loop is already running'));
console.log(palette.warn('Ralph loop is already running'));
return;
}
loop.on('taskAssigned', (taskId, sessionId) => {
console.log(chalk.cyan(`→ Task ${taskId.slice(0, 8)} assigned to session ${sessionId.slice(0, 8)}`));
console.log(palette.info(`→ Task ${taskId.slice(0, 8)} assigned to session ${sessionId.slice(0, 8)}`));
});
loop.on('taskCompleted', (taskId) => {
console.log(chalk.green(`✓ Task ${taskId.slice(0, 8)} completed`));
console.log(palette.ok(`✓ Task ${taskId.slice(0, 8)} completed`));
});
loop.on('taskFailed', (taskId, error) => {
console.log(chalk.red(`✗ Task ${taskId.slice(0, 8)} failed: ${error}`));
console.log(palette.err(`✗ Task ${taskId.slice(0, 8)} failed: ${error}`));
});
loop.on('stopped', () => {
console.log(chalk.yellow('\nRalph loop stopped'));
console.log(palette.warn('\nRalph loop stopped'));
printStats(loop.getStats());
process.exit(0);
});
await loop.start();
console.log(chalk.green('✓ Ralph loop started'));
console.log(palette.ok('✓ Ralph loop started'));
if (options.minHours) {
console.log(` Minimum duration: ${options.minHours} hours`);
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
console.log(palette.muted(' Press Ctrl+C to stop\n'));
// Keep process running
process.on('SIGINT', () => {
console.log(chalk.yellow('\nStopping Ralph loop...'));
console.log(palette.warn('\nStopping Ralph loop...'));
loop.stop();
});
});
@@ -545,11 +563,11 @@ ralphCmd
.action(() => {
const loop = getRalphLoop();
if (!loop.isRunning()) {
console.log(chalk.yellow('Ralph loop is not running'));
console.log(palette.warn('Ralph loop is not running'));
return;
}
loop.stop();
console.log(chalk.green('✓ Ralph loop stopped'));
console.log(palette.ok('✓ Ralph loop stopped'));
});
ralphCmd
@@ -562,9 +580,10 @@ ralphCmd
});
function printStats(stats: ReturnType<ReturnType<typeof getRalphLoop>['getStats']>) {
const statusColor = stats.status === 'running' ? chalk.green : stats.status === 'paused' ? chalk.yellow : chalk.gray;
const statusColor =
stats.status === 'running' ? palette.ok : stats.status === 'paused' ? palette.warn : palette.muted;
console.log(chalk.bold('\nRalph Loop Status:'));
console.log(palette.emph('\nRalph Loop Status:'));
console.log(` Status: ${statusColor(stats.status)}`);
console.log(` Elapsed: ${stats.elapsedHours.toFixed(2)} hours`);
if (stats.minDurationMs) {
@@ -574,14 +593,14 @@ function printStats(stats: ReturnType<ReturnType<typeof getRalphLoop>['getStats'
);
}
console.log(chalk.bold('\nTasks:'));
console.log(palette.emph('\nTasks:'));
console.log(` Pending: ${stats.pending}`);
console.log(` Running: ${stats.running}`);
console.log(` Completed: ${stats.completed} (${stats.tasksCompleted} this session)`);
console.log(` Failed: ${stats.failed}`);
console.log(` Generated: ${stats.tasksGenerated}`);
console.log(chalk.bold('\nSessions:'));
console.log(palette.emph('\nSessions:'));
console.log(` Active: ${stats.activeSessions}`);
console.log(` Idle: ${stats.idleSessions}`);
console.log(` Busy: ${stats.busySessions}`);
@@ -695,20 +714,20 @@ program
}
}
console.log(chalk.bold('\nCodeman Status'));
console.log('─'.repeat(40));
console.log(heading('Codeman Status'));
console.log(rule(40));
console.log(chalk.bold('\nWeb Server:'));
console.log(heading('Web Server:'));
if (probe.reachable) {
const version = probe.version ? ` (v${probe.version})` : '';
console.log(` Status: ${chalk.green('running')}${version} at ${probe.url}`);
console.log(kv('Status', `${palette.ok('running')}${version} at ${probe.url}`));
if (probe.authRequired) {
console.log(chalk.gray(' (answers 401: set CODEMAN_PASSWORD/CODEMAN_USERNAME to see session details)'));
console.log(palette.muted(' (answers 401: set CODEMAN_PASSWORD/CODEMAN_USERNAME to see session details)'));
}
} else {
console.log(` Status: ${chalk.red('not reachable')} at ${candidates.join(' or ')}`);
console.log(kv('Status', `${palette.err('not reachable')} at ${candidates.join(' or ')}`));
console.log(
chalk.gray(' (start it with `codeman web`, or check your service: systemctl --user status codeman-web)')
palette.muted(' (start it with `codeman web`, or check your service: systemctl --user status codeman-web)')
);
}
@@ -716,26 +735,26 @@ program
// as such, so the numbers are never silently a different thing.
if (probe.sessions) {
const live = probe.sessions;
console.log(chalk.bold('\nSessions (live, from the server):'));
console.log(` Total: ${live.length}`);
console.log(` Idle: ${live.filter((s) => s.status === 'idle').length}`);
console.log(` Busy: ${live.filter((s) => s.status === 'busy').length}`);
console.log(heading('Sessions (live, from the server):'));
console.log(kv('Total', String(live.length)));
console.log(kv('Idle', String(live.filter((s) => s.status === 'idle').length)));
console.log(kv('Busy', String(live.filter((s) => s.status === 'busy').length)));
} else {
const manager = getSessionManager();
const storedValues = Object.values(manager.getStoredSessions());
console.log(chalk.bold('\nSessions (from saved state):'));
console.log(` Active: ${storedValues.filter((s) => s.status !== 'stopped').length}`);
console.log(` Idle: ${storedValues.filter((s) => s.status === 'idle').length}`);
console.log(` Busy: ${storedValues.filter((s) => s.status === 'busy').length}`);
console.log(heading('Sessions (from saved state):'));
console.log(kv('Active', String(storedValues.filter((s) => s.status !== 'stopped').length)));
console.log(kv('Idle', String(storedValues.filter((s) => s.status === 'idle').length)));
console.log(kv('Busy', String(storedValues.filter((s) => s.status === 'busy').length)));
}
const taskCounts = getTaskQueue().getCount();
console.log(chalk.bold('\nTasks:'));
console.log(` Total: ${taskCounts.total}`);
console.log(` Pending: ${taskCounts.pending}`);
console.log(` Running: ${taskCounts.running}`);
console.log(` Completed: ${taskCounts.completed}`);
console.log(` Failed: ${taskCounts.failed}`);
console.log(heading('Tasks:'));
console.log(kv('Total', String(taskCounts.total)));
console.log(kv('Pending', String(taskCounts.pending)));
console.log(kv('Running', String(taskCounts.running)));
console.log(kv('Completed', String(taskCounts.completed)));
console.log(kv('Failed', String(taskCounts.failed)));
console.log('');
});
@@ -745,9 +764,17 @@ program
.option('-f, --force', 'Skip confirmation')
.action(async (options) => {
if (!options.force) {
console.log(chalk.yellow('This will stop all sessions and clear all state.'));
console.log(chalk.yellow('Use --force to confirm.'));
return;
console.log(palette.warn('This will stop all sessions and clear all state.'));
// Non-interactive callers keep the old refusal: a script piping into the
// CLI must never be able to reset state by hanging on an unseen question.
if (!isInteractive()) {
console.log(palette.warn('Use --force to confirm.'));
return;
}
if (!(await confirm('Reset all Codeman state?'))) {
console.log(palette.muted('○ Cancelled, nothing was changed'));
return;
}
}
const manager = getSessionManager();
@@ -756,7 +783,7 @@ program
await manager.stopAllSessions();
store.reset();
console.log(chalk.green('✓ All state reset'));
console.log(palette.ok('✓ All state reset'));
});
// Shorthand commands at root level
@@ -767,38 +794,45 @@ program
.action(async (options) => {
const manager = getSessionManager();
const session = await manager.createSession(options.dir);
console.log(chalk.green(`✓ Session started: ${session.id}`));
console.log(palette.ok(`✓ Session started: ${session.id}`));
});
program
.command('list')
.alias('ls')
.description('List all sessions (shorthand)')
.action(() => {
const manager = getSessionManager();
const sessions = manager.getAllSessions();
const stored = manager.getStoredSessions();
.description('List active sessions (shorthand; `codeman session list` also shows stopped ones)')
.action(() => printSessionList({ includeStored: false }));
if (sessions.length === 0 && Object.keys(stored).length === 0) {
console.log(chalk.yellow('No sessions found'));
// ============ TUI ============
program
.command('tui')
.argument('[n]', 'attach straight to the nth session of `codeman tui --list`')
.description('Terminal dashboard for your sessions (the web UI remains the primary surface)')
.option('-l, --list', 'Print the numbered session list and exit, instead of opening the dashboard')
.action(async (position: string | undefined, options: { list?: boolean }) => {
// Imported here, not at the top: the dashboard pulls in the whole TUI core,
// and every other command would pay for it at startup.
const { runTui, runTuiAttach, runTuiList } = await import('./tui/tui-app.js');
if (options.list) {
process.exitCode = await runTuiList();
return;
}
console.log(chalk.bold('\nActive Sessions:'));
if (sessions.length === 0) {
console.log(' (none)');
} else {
for (const session of sessions) {
const status =
session.status === 'idle'
? chalk.green('idle')
: session.status === 'busy'
? chalk.yellow('busy')
: chalk.red(session.status);
console.log(` ${chalk.cyan(session.id.slice(0, 8))} ${status} ${session.workingDir}`);
if (position !== undefined) {
const n = Number.parseInt(position, 10);
if (!Number.isSafeInteger(n) || n < 1) {
console.error(palette.err(`"${position}" is not a session number.`));
console.error(`Run ${palette.info('codeman tui --list')} to see them.`);
process.exitCode = 1;
return;
}
process.exitCode = await runTuiAttach(n);
return;
}
console.log('');
// The dashboard owns the terminal until it quits; exiting explicitly keeps a
// stray handle (a socket mid-close) from stranding the user's shell.
process.exit(await runTui());
});
// ============ Web / daemon / service Commands ============
@@ -831,7 +865,7 @@ function toWebLaunchOptions(options: {
}): WebLaunchOptions {
const port = parseInt(options.port, 10);
if (!Number.isInteger(port) || port <= 0 || port > 65535) {
console.error(chalk.red(`✗ Invalid port: ${options.port}`));
console.error(palette.err(`✗ Invalid port: ${options.port}`));
process.exit(1);
}
return {
@@ -852,11 +886,11 @@ function warnIfUnauthenticatedNetwork(launch: WebLaunchOptions): void {
if (isLoopbackBindHost(launch.host)) return;
if (isUnauthenticatedNetworkAcknowledged(launch.allowUnauthenticatedNetwork)) return;
console.log(
chalk.yellow(
palette.warn(
`⚠ Binding ${launch.host} without CODEMAN_PASSWORD: anyone who can reach this port gets terminal control.`
)
);
console.log(chalk.yellow(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
console.log(palette.warn(' Set CODEMAN_PASSWORD, or bind 127.0.0.1 and front it with tailscale serve.'));
}
// Web interface command
@@ -872,17 +906,18 @@ webCmd.action(async (options) => {
const launch = toWebLaunchOptions(options);
if (options.stop) {
const result = await stopDaemon(launch);
// stopDaemon waits for the process to actually exit (up to 15s).
const result = await withSpinner('Stopping Codeman...', () => stopDaemon(launch));
if (result.ok && result.reason === 'not-running') {
console.log(chalk.gray(`○ ${result.message}`));
console.log(palette.muted(`○ ${result.message}`));
return;
}
if (result.ok) {
console.log(chalk.green(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(chalk.gray(' Your agents keep running in tmux.'));
console.log(palette.ok(`✓ ${result.message ?? `Stopped Codeman (pid ${result.pid})`}`));
console.log(palette.muted(' Your agents keep running in tmux.'));
return;
}
console.error(chalk.red(`✗ ${result.message ?? 'Could not stop the server'}`));
console.error(palette.err(`✗ ${result.message ?? 'Could not stop the server'}`));
process.exit(1);
}
@@ -890,31 +925,33 @@ webCmd.action(async (options) => {
const status = await daemonStatus(launch);
if (status.responding) {
const version = status.version ? ` (v${status.version})` : '';
console.log(chalk.green(`✓ Responding at ${status.url}${version}`));
console.log(palette.ok(`✓ Responding at ${status.url}${version}`));
} else {
console.log(chalk.yellow(`○ Nothing answering at ${status.url}`));
console.log(palette.warn(`○ Nothing answering at ${status.url}`));
}
console.log(` Daemon pid: ${status.running ? chalk.green(String(status.pid)) : chalk.gray('not running')}`);
console.log(chalk.gray(` Pidfile: ${status.pidFile}`));
console.log(chalk.gray(` Log: ${status.logPath}`));
console.log(kv('Daemon pid', status.running ? palette.ok(String(status.pid)) : palette.muted('not running'), 11));
console.log(palette.muted(kv('Pidfile', status.pidFile, 11)));
console.log(palette.muted(kv('Log', status.logPath, 11)));
if (!status.running && status.responding) {
console.log(chalk.gray(' (running, but not started with --daemon: probably a service or a foreground run)'));
console.log(palette.muted(' (running, but not started with --daemon: probably a service or a foreground run)'));
}
return;
}
if (options.daemon) {
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Starting Codeman in the background...'));
const result = await startDaemon(launch);
// The start polls /api/status for up to 30s; without this the shell just sits there.
const result = await withSpinner('Starting Codeman in the background, waiting for it to answer...', () =>
startDaemon(launch)
);
if (result.ok) {
console.log(chalk.green(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(chalk.gray(` Logs: ${result.logPath}`));
console.log(chalk.gray(' Stop it with: codeman web --stop'));
console.log(chalk.gray(' Want it back after a reboot? codeman service install'));
console.log(palette.ok(`\n✓ Codeman is running at ${result.url} (pid ${result.pid})`));
console.log(palette.muted(` Logs: ${result.logPath}`));
console.log(palette.muted(' Stop it with: codeman web --stop'));
console.log(palette.muted(' Want it back after a reboot? codeman service install'));
return;
}
console.error(chalk.red(`\n✗ ${result.message ?? 'Failed to start'}`));
console.error(palette.err(`\n✗ ${result.message ?? 'Failed to start'}`));
process.exit(1);
}
@@ -924,29 +961,29 @@ webCmd.action(async (options) => {
const https = launch.https;
const titleHostname = options.titleHostname;
const allowUnauthenticatedNetwork = launch.allowUnauthenticatedNetwork ?? false;
const protocol = https ? 'https' : 'http';
const displayHost = host === '0.0.0.0' ? 'localhost' : host;
console.log(chalk.cyan(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
console.log(palette.info(`Starting Codeman web interface on ${displayHost}:${port}${https ? ' (HTTPS)' : ''}...`));
try {
// The server prints its own "running at" line (it also covers the daemon and
// service launch paths), so this one used to be a duplicate of it.
const server = await startWebServer(port, https, false, host, titleHostname, allowUnauthenticatedNetwork);
console.log(chalk.green(`\n✓ Web interface running at ${protocol}://${displayHost}:${port}`));
if (https) {
console.log(chalk.yellow(' Note: Accept the self-signed certificate in your browser on first visit'));
console.log(palette.warn(' Note: Accept the self-signed certificate in your browser on first visit'));
}
console.log(chalk.gray(' Press Ctrl+C to stop\n'));
console.log(palette.muted(' Press Ctrl+C to stop\n'));
// Graceful shutdown handler — flush state and clean up on SIGTERM/SIGINT
let shuttingDown = false;
const shutdown = async (signal: string) => {
if (shuttingDown) return;
shuttingDown = true;
console.log(chalk.yellow(`\n${signal} received, shutting down gracefully...`));
console.log(palette.warn(`\n${signal} received, shutting down gracefully...`));
try {
await server.stop();
} catch (err) {
console.error(chalk.red(`Error during shutdown: ${getErrorMessage(err)}`));
console.error(palette.err(`Error during shutdown: ${getErrorMessage(err)}`));
}
process.exit(0);
};
@@ -954,7 +991,7 @@ webCmd.action(async (options) => {
process.on('SIGINT', () => shutdown('SIGINT'));
process.on('SIGHUP', () => shutdown('SIGHUP'));
} catch (err) {
console.error(chalk.red(`✗ Failed to start web server: ${getErrorMessage(err)}`));
console.error(palette.err(`✗ Failed to start web server: ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -970,20 +1007,23 @@ addWebLaunchOptions(
).action(async (options) => {
const launch = toWebLaunchOptions(options);
warnIfUnauthenticatedNetwork(launch);
console.log(chalk.cyan('Installing the Codeman service...'));
const result = await installService(launch);
for (const warning of result.warnings ?? []) console.log(chalk.yellow(`⚠ ${warning}`));
// Install polls the new unit's /api/status for up to 30s before it can honestly
// report success, so the wait needs a visible heartbeat.
const result = await withSpinner('Installing the Codeman service, waiting for it to answer...', () =>
installService(launch)
);
for (const warning of result.warnings ?? []) console.log(palette.warn(`⚠ ${warning}`));
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
console.error(palette.err(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(chalk.gray(` Unit: ${result.unitPath}`));
console.log(palette.ok(`✓ ${result.message}`));
console.log(palette.muted(` Unit: ${result.unitPath}`));
if (process.env.CODEMAN_PASSWORD) {
console.log(
chalk.yellow(
palette.warn(
' Note: CODEMAN_PASSWORD was NOT copied into the unit file. Add it there yourself if the service needs auth.'
)
);
@@ -996,10 +1036,10 @@ serviceCmd
.action(() => {
const result = uninstallService();
if (!result.ok) {
console.error(chalk.red(`✗ ${result.message}`));
console.error(palette.err(`✗ ${result.message}`));
process.exit(1);
}
console.log(chalk.green(`✓ ${result.message}`));
console.log(palette.ok(`✓ ${result.message}`));
});
addWebLaunchOptions(
@@ -1007,15 +1047,15 @@ addWebLaunchOptions(
).action(async (options) => {
const status = await serviceStatus(toWebLaunchOptions(options));
if (!status.kind) {
console.log(chalk.yellow(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
console.log(palette.warn(`No supported supervisor on ${process.platform}. Use \`codeman web -d\` instead.`));
return;
}
console.log(` Supervisor: ${status.kind} (${status.name})`);
console.log(` Unit file: ${status.installed ? chalk.green(status.unitPath) : chalk.gray('not installed')}`);
console.log(` Loaded: ${status.loaded ? chalk.green('yes') : chalk.gray('no')}`);
console.log(` Unit file: ${status.installed ? palette.ok(status.unitPath) : palette.muted('not installed')}`);
console.log(` Loaded: ${status.loaded ? palette.ok('yes') : palette.muted('no')}`);
const version = status.version ? ` (v${status.version})` : '';
console.log(
` Responding: ${status.responding ? chalk.green(`yes at ${status.url}${version}`) : chalk.gray(`no at ${status.url}`)}`
` Responding: ${status.responding ? palette.ok(`yes at ${status.url}${version}`) : palette.muted(`no at ${status.url}`)}`
);
});
@@ -1085,7 +1125,7 @@ usersCmd
.action(async (name, options) => {
const { createUser, isValidUsername } = await import('./user-store.js');
if (!isValidUsername(name)) {
console.error(chalk.red('✗ Username must be lowercase, start alphanumeric, 2-32 chars ([a-z0-9_-])'));
console.error(palette.err('✗ Username must be lowercase, start alphanumeric, 2-32 chars ([a-z0-9_-])'));
process.exit(1);
}
try {
@@ -1096,18 +1136,18 @@ usersCmd
password = await promptHiddenPassword('New password: ');
const confirm = await promptHiddenPassword('Confirm password: ');
if (password !== confirm) {
console.error(chalk.red('✗ Passwords do not match'));
console.error(palette.err('✗ Passwords do not match'));
process.exit(1);
}
}
if (!password || password.length < 8) {
console.error(chalk.red('✗ Password must be at least 8 characters'));
console.error(palette.err('✗ Password must be at least 8 characters'));
process.exit(1);
}
const user = await createUser({ username: name, role: options.admin ? 'admin' : 'user', password });
console.log(chalk.green(`✓ Created ${user.role} "${user.username}"`));
console.log(palette.ok(`✓ Created ${user.role} "${user.username}"`));
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
console.error(palette.err(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -1126,14 +1166,14 @@ usersCmd
password = await promptHiddenPassword('New password: ');
const confirm = await promptHiddenPassword('Confirm password: ');
if (password !== confirm) {
console.error(chalk.red('✗ Passwords do not match'));
console.error(palette.err('✗ Passwords do not match'));
process.exit(1);
}
}
await setPassword(name, password, { mustChangePassword: false });
console.log(chalk.green(`✓ Password updated for "${name}"`));
console.log(palette.ok(`✓ Password updated for "${name}"`));
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
console.error(palette.err(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
@@ -1146,17 +1186,17 @@ usersCmd
const { readUsers } = await import('./user-store.js');
const users = await readUsers(true);
if (users.length === 0) {
console.log(chalk.yellow('No users defined (run: codeman users add <name> --admin)'));
console.log(palette.warn('No users defined (run: codeman users add <name> --admin)'));
return;
}
console.log(chalk.bold('\nUsers:'));
console.log(palette.emph('\nUsers:'));
for (const u of users) {
const role = u.role === 'admin' ? chalk.magenta('admin') : chalk.cyan('user ');
const state = u.disabled ? chalk.red('disabled') : chalk.green('enabled ');
const role = u.role === 'admin' ? palette.accent('admin') : palette.info('user ');
const state = u.disabled ? palette.err('disabled') : palette.ok('enabled ');
const flags = [u.mustChangePassword ? 'must-change-pw' : '', u.canBypassPermissions ? 'can-bypass' : '']
.filter(Boolean)
.join(' ');
console.log(` ${role} ${state} ${u.username}${flags ? chalk.gray(` [${flags}]`) : ''}`);
console.log(` ${role} ${state} ${u.username}${flags ? palette.muted(` [${flags}]`) : ''}`);
}
console.log('');
});
@@ -1171,16 +1211,43 @@ usersCmd
await deleteUser(name);
if (options.deleteSpace) {
await deleteUserSpace(name);
console.log(chalk.green(`✓ Deleted user "${name}" and their space`));
console.log(palette.ok(`✓ Deleted user "${name}" and their space`));
} else {
console.log(chalk.green(`✓ Deleted user "${name}" (space left on disk)`));
console.log(palette.ok(`✓ Deleted user "${name}" (space left on disk)`));
}
} catch (err) {
console.error(chalk.red(`✗ ${getErrorMessage(err)}`));
console.error(palette.err(`✗ ${getErrorMessage(err)}`));
process.exit(1);
}
});
/**
* Missing REQUIRED tools are failures; a missing optional one or a skipped check
* is just absence, so it stays muted rather than shouting red at everyone
* without LibreOffice installed.
*/
function dependencyTone(result: ToolResult): Tone {
if (result.status === 'ok') return 'ok';
if (result.status === 'skipped') return 'idle';
return result.required ? 'err' : 'idle';
}
/**
* The colorize hook `dependency-report.ts` was written for. Versions stay in the
* default color (they are data, not a verdict); everything that IS a verdict is
* painted, and the supporting detail is muted so the glyph column reads first.
*/
const DOCTOR_STYLE: ReportStyle = {
title: (text) => palette.emph(text),
heading: (text) => palette.emph(palette.info(text)),
glyph: (result, glyph) => tint(dependencyTone(result), glyph),
label: (text) => text,
status: (result, text) => (result.status === 'ok' ? text : tint(dependencyTone(result), text)),
path: (text) => palette.muted(text),
meta: (text) => palette.muted(text),
summary: (text) => palette.emph(text),
};
program
.command('doctor')
.alias('check-deps')
@@ -1190,7 +1257,7 @@ program
.action(async (options) => {
const { createRealHost, checkAll } = await import('./utils/dependency-checker.js');
const { renderTable, renderJson, computeExitCode } = await import('./utils/dependency-report.js');
const { DEPENDENCY_REGISTRY, TOOL_CATEGORIES } = await import('./config/dependency-registry.js');
const { dependencyRegistry, TOOL_CATEGORIES } = await import('./config/dependency-registry.js');
if (options.category && !(TOOL_CATEGORIES as readonly string[]).includes(options.category)) {
console.error(`Unknown category "${options.category}". Valid categories: ${TOOL_CATEGORIES.join(', ')}`);
@@ -1198,15 +1265,15 @@ program
}
const host = createRealHost();
const registry = options.category
? DEPENDENCY_REGISTRY.filter((t) => t.category === options.category)
: DEPENDENCY_REGISTRY;
const allTools = dependencyRegistry();
const registry = options.category ? allTools.filter((t) => t.category === options.category) : allTools;
const results = checkAll(registry, host);
if (options.json) {
// Raw JSON, never styled: this output is parsed, not read.
console.log(JSON.stringify(renderJson(results, host.environment), null, 2));
} else {
console.log(renderTable(results, host.environment));
console.log(renderTable(results, host.environment, DOCTOR_STYLE));
}
process.exit(computeExitCode(results));
});
+37
View File
@@ -0,0 +1,37 @@
/**
* @fileoverview Where case (project) folders live.
*
* Deliberately NOT instance-scoped, unlike `dataPath()`: `~/codeman-cases` is
* shared by every Codeman on the machine, the same way `~/codeman-users/<u>`
* user spaces are, so a beta instance sees the same projects as prod.
*
* `CODEMAN_CASES_PATH` overrides the location. Docker Compose deployments set
* it to a host-absolute bind mount so a Docker case's workspace resolves to the
* SAME absolute path inside Codeman and on the host daemon that mounts it.
*
* ⚠️ **One resolver, every caller.** This started life as three hardcoded
* `join(homedir(), 'codeman-cases')` copies. When only the web server's copy
* learned the override, `codeman skill install --case <name>` still looked in
* the home default and reported "Case not found" on exactly the deployment the
* override exists for. A new cases-dir consumer imports this; it does not
* rebuild the path.
*
* (`state-store.ts` keeps its own literal on purpose: that one migrates the
* historical `~/claudeman-cases` directory to `~/codeman-cases` by name, and is
* about the old default location rather than the active one.)
*
* @module config/cases-dir
*/
import { homedir } from 'node:os';
import { join } from 'node:path';
/** Absolute path to the shared cases directory. */
export function getCasesDir(): string {
return process.env.CODEMAN_CASES_PATH || join(homedir(), 'codeman-cases');
}
/** Absolute path to one case folder inside it. */
export function casePath(name: string): string {
return join(getCasesDir(), name);
}
+190
View File
@@ -0,0 +1,190 @@
/**
* @fileoverview The argv rendering engine — turns a `CliLaunch` spec plus a set of resolved
* parameter values into the shell command string that goes into `bash -c "..."`.
*
* SECURITY MODEL (read before touching this file):
*
* 1. Config contains no shell text. There is no `command: "..."` field anywhere in the
* schema. An entry declares a sequence of typed tokens (`ArgSpec`); this module is the
* ONLY place that turns them into a string, and it owns every separator itself: a single
* space between tokens, and ` || ` between fallback variants. Neither can originate from
* config, because config has no field that could hold either.
* 2. Every literal (`lit`, `flag`, `value`) is validated against `SAFE_BARE_TOKEN` — no
* space, quote, backtick, `$`, `;`, `&`, `|`, `<`, `>`, parens, braces, newline or
* backslash — at LOAD time (see schema.ts), so a bad literal fails registry validation
* rather than reaching this renderer.
* 3. Every `valueFrom` resolves through a declared `ParamSpec`, whose `token` variant names
* a PATTERN rather than accepting one — see patterns.ts. A value that fails its pattern
* causes the WHOLE ArgSpec to be dropped, exactly like the hand-written builders this
* replaces (an invalid `--model` value silently omits `--model`, it does not substitute
* something else).
* 4. Escaping and validation are independent. `renderToken()` always re-checks the resolved
* value against `SAFE_BARE_TOKEN` before emitting it unquoted; anything else is
* single-quote-escaped. So even a value that somehow bypassed pattern validation is still
* quoted, never concatenated raw.
*
* @module config/cli-registry/argv
*/
import type { ArgSpec, CliEntry, CliLaunch, Cond, EngineValue, ParamSpec, QuoteStyle } from './types.js';
import { matchesPattern } from './patterns.js';
import { SAFE_BARE_TOKEN } from './patterns.js';
/** Resolved parameter values, keyed by the name declared in `CliLaunch.params`. */
export type ParamValues = Record<string, string | boolean | undefined>;
/** Values the caller supplies for the reserved engine params. */
export type EngineValues = Partial<Record<EngineValue, string>>;
/**
* POSIX single-quote escaping: end-quote, escaped-literal-quote, restart-quote. Identical in
* shape to the three copies already in the codebase (tmux-manager.ts, remote-hosts.ts,
* docker-hosts.ts) — kept local rather than importing one of them so this module has no
* dependency on the files it is replacing.
*/
function singleQuoteEscape(value: string): string {
return `'${value.replace(/'/g, `'\\''`)}'`;
}
function doubleQuoteEscape(value: string): string {
// Escape the characters that are special inside a double-quoted bash string. SAFE_BARE_TOKEN
// already excludes all of them, so in practice this never fires; kept as defense in depth.
return `"${value.replace(/([$`"\\])/g, '\\$1')}"`;
}
/**
* Render a single resolved value per its requested quote style. `auto` (the default) emits
* bare only when the value is provably safe; every other case single-quotes.
*/
function renderToken(value: string, style: QuoteStyle | undefined): string {
const safe = SAFE_BARE_TOKEN.test(value);
switch (style) {
case 'double':
return doubleQuoteEscape(value);
case 'single':
return singleQuoteEscape(value);
case 'bare':
return safe ? value : singleQuoteEscape(value);
case 'auto':
default:
return safe ? value : singleQuoteEscape(value);
}
}
/** Resolve one parameter to a plain string, or undefined if it is unset / invalid. */
function resolveParam(
name: string,
spec: ParamSpec | undefined,
params: ParamValues,
engineValues: EngineValues
): string | undefined {
if (!spec) return undefined;
if (spec.type === 'engine') return engineValues[spec.source];
const raw = params[name];
if (raw === undefined) return spec.type === 'enum' ? spec.default : undefined;
if (spec.type === 'bool') return typeof raw === 'boolean' ? String(raw) : undefined;
if (spec.type === 'enum') {
const s = String(raw);
return spec.values.includes(s) ? s : spec.default;
}
// token
const s = String(raw);
return matchesPattern(spec.pattern, s) ? s : undefined;
}
/** Is the resolved value "set" for the purposes of a `state` condition? */
function isSet(name: string, params: ParamValues, resolved: (n: string) => string | undefined): boolean {
if (name in params) {
const raw = params[name];
if (typeof raw === 'boolean') return true; // a bool param is always "set" once declared
}
return resolved(name) !== undefined;
}
function evalCond(
cond: Cond | undefined,
params: ParamValues,
resolved: (n: string) => string | undefined,
gatesPassed: ReadonlySet<string>
): boolean {
if (!cond) return true;
if ('allOf' in cond) return cond.allOf.every((c) => evalCond(c, params, resolved, gatesPassed));
if ('anyOf' in cond) return cond.anyOf.some((c) => evalCond(c, params, resolved, gatesPassed));
if ('not' in cond) return !evalCond(cond.not, params, resolved, gatesPassed);
if ('capabilityGate' in cond) return gatesPassed.has(cond.capabilityGate);
if ('state' in cond) {
const set = isSet(cond.param, params, resolved);
return cond.state === 'set' ? set : !set;
}
// { param, is }
const raw = params[cond.param];
if (typeof cond.is === 'boolean') return raw === cond.is;
return resolved(cond.param) === cond.is;
}
function renderArg(
spec: ArgSpec,
params: ParamValues,
resolved: (n: string) => string | undefined,
gatesPassed: ReadonlySet<string>
): string | null {
if (!evalCond(spec.when, params, resolved, gatesPassed)) return null;
if ('lit' in spec) return spec.lit;
if ('flag' in spec && !('value' in spec) && !('valueFrom' in spec)) return spec.flag;
if ('flag' in spec && 'value' in spec) return `${spec.flag} ${renderToken(spec.value, spec.quote)}`;
if ('flag' in spec && 'valueFrom' in spec) {
const v = resolved(spec.valueFrom);
return v === undefined ? null : `${spec.flag} ${renderToken(v, spec.quote)}`;
}
// bare positional
const v = resolved((spec as { valueFrom: string }).valueFrom);
return v === undefined ? null : renderToken(v, (spec as { quote?: QuoteStyle }).quote);
}
/**
* Render one CLI's launch command. Returns the full `bash -c` payload — never a shell
* fragment with embedded newlines or unescaped separators, by construction (see file header).
*
* `gatesPassed` — the set of `capabilities.gates` keys whose version requirement is
* currently satisfied. Callers compute this once per spawn (it depends on a version probe),
* never inside the renderer, keeping this function pure and easy to test byte-for-byte.
*/
export function renderLaunch(
launch: CliLaunch,
params: ParamValues,
engineValues: EngineValues,
gatesPassed: ReadonlySet<string> = new Set()
): string {
const cache = new Map<string, string | undefined>();
const resolved = (name: string): string | undefined => {
if (cache.has(name)) return cache.get(name);
const v = resolveParam(name, launch.params[name], params, engineValues);
cache.set(name, v);
return v;
};
const passing = launch.variants.filter((variant) => evalCond(variant.when, params, resolved, gatesPassed));
const chosen = launch.chain === 'fallback' ? passing : passing.slice(0, 1);
const rendered = chosen.map((variant) =>
variant.args
.map((arg) => renderArg(arg, params, resolved, gatesPassed))
.filter((tok): tok is string => tok !== null)
.join(' ')
);
return rendered.join(' || ');
}
/** Convenience: render an entry's launch command straight from a `CliEntry`. */
export function renderCliCommand(
entry: CliEntry,
params: ParamValues,
engineValues: EngineValues,
gatesPassed?: ReadonlySet<string>
): string {
return renderLaunch(entry.launch, params, engineValues, gatesPassed);
}
+61
View File
@@ -0,0 +1,61 @@
/**
* @fileoverview Barrel for the CLI registry module.
* @module config/cli-registry
*/
export type {
ArgSpec,
CliCapabilities,
CliCredStore,
CliDiscovery,
CliEntry,
CliEnv,
CliId,
CliIdentityProbe,
CliLaunch,
CliOverlays,
CliRegistryFile,
CliVariant,
CliVersionProbe,
Cond,
EngineValue,
ParamSpec,
QuoteStyle,
} from './types.js';
export {
matchesPattern,
TOKEN_PATTERNS,
SAFE_BARE_TOKEN,
compileVersionRegex,
MAX_VERSION_OUTPUT,
} from './patterns.js';
export type { TokenPattern } from './patterns.js';
export { renderLaunch, renderCliCommand } from './argv.js';
export type { EngineValues, ParamValues } from './argv.js';
export { CliEntrySchema } from './schema.js';
export type { ValidatedCliEntry } from './schema.js';
export { STOCK_CLIS } from './stock.js';
export {
asCliId,
cliIds,
enabledCliIds,
enabledClis,
getCli,
listClis,
loadCliRegistry,
reloadCliRegistry,
resolveInstallCommandForPlatform,
resolveRegistry,
} from './registry.js';
export type { LoadResult } from './registry.js';
export {
COMPOSER_ANCHOR_KINDS,
isKnownLauncherProfile,
isKnownPredictProfile,
isKnownSetenvProfile,
LAUNCHER_PROFILE_NAMES,
PREDICT_PROFILES,
SETENV_PROFILE_NAMES,
TRANSCRIPT_READER_NAMES,
} from './profiles.js';
export type { LauncherProfileName, SetenvProfileName } from './profiles.js';
+121
View File
@@ -0,0 +1,121 @@
/**
* @fileoverview Named value patterns for the CLI registry's argv engine.
*
* Config entries select a pattern BY NAME; the regexes themselves live here, in code.
* That is deliberate and is the reason a user-editable `clis.json` cannot widen its own
* validation: there is no field anywhere in the schema that accepts a raw regex for a
* shell token, so no entry can supply `.*` (nor a catastrophically backtracking one).
*
* The sole user-supplied regex in the whole registry is `discovery.version.regex`, which
* is applied to `--version` OUTPUT rather than to a shell token, and goes through
* `compileVersionRegex()` below.
*
* Every pattern here is transcribed from the builder it replaces in tmux-manager.ts, so
* the argv engine accepts and rejects exactly the values the hand-written builders did.
*
* @module config/cli-registry/patterns
*/
/** Names a value pattern. Config may only reference these. */
export type TokenPattern =
| 'model'
| 'model-claude'
| 'model-pi'
| 'id'
| 'id-dotted'
| 'uuid'
| 'slug'
| 'path-segment'
| 'tool-list'
| 'config-kv';
/**
* The patterns, each traced to the builder it came from.
*
* ⚠️ These are ALLOWLISTS (`^...$` over a safe character class), never blocklists — with
* one deliberate exception, `tool-list`, which mirrors the existing `--allowedTools`
* sanitizer. That one is a metacharacter REJECTION because tool specs legitimately contain
* `(`, `)`, `*`, `:` and spaces (`Bash(git:*), Read`), so an allowlist of safe words cannot
* express it. Keeping it byte-identical to the original matters more than making it uniform.
*/
const PATTERNS: Record<TokenPattern, RegExp> = {
// buildOpenCodeCommand / buildCodexCommand / buildGeminiCommand / buildAntigravityCommand
model: /^[a-zA-Z0-9._\-/]+$/,
// buildSpawnCommand's claude branch — `[` and `]` for bracketed model aliases
'model-claude': /^[a-zA-Z0-9._\-[\]]+$/,
// buildPiCommand — `:` for a thinking suffix (`sonnet:high`), `/` for `provider/id`
'model-pi': /^[a-zA-Z0-9._\-/:]+$/,
// opencode --session, codex resume
id: /^[a-zA-Z0-9_-]+$/,
// gemini --resume, antigravity --conversation, pi --session
'id-dotted': /^[a-zA-Z0-9._-]+$/,
// claude --resume / --session-id
uuid: /^[a-f0-9-]+$/,
// pi --provider
slug: /^[a-z0-9-]+$/,
// dsh --profile. Deliberately STRICTER than `id-dotted`: a profile name is both
// interpolated into the shell line AND joined into a filesystem path, so it must be a
// single path segment. Requiring a leading alphanumeric is what rules out `.`, `..` and
// dotfile names, which `id-dotted` would happily accept.
'path-segment': /^[a-zA-Z0-9][a-zA-Z0-9._-]*$/,
// codex --config tui.animations=false
'config-kv': /^[A-Za-z0-9._-]+=[A-Za-z0-9._-]+$/,
// Placeholder; `tool-list` is handled by isSafeToolList() below, not by a match.
'tool-list': /^$/,
};
/**
* Shell metacharacters rejected in an `--allowedTools` value. Transcribed verbatim from
* buildClaudePermissionFlags so the accepted set does not move.
*/
const TOOL_LIST_DANGEROUS = /[;&|$`\\{}<>'"[\]\n\r]/;
/** Does `value` satisfy the named pattern? */
export function matchesPattern(pattern: TokenPattern, value: string): boolean {
if (pattern === 'tool-list') return value.length > 0 && !TOOL_LIST_DANGEROUS.test(value);
return PATTERNS[pattern].test(value);
}
/** Every pattern name, for schema validation and error messages. */
export const TOKEN_PATTERNS = Object.keys(PATTERNS) as TokenPattern[];
/**
* Characters a token may contain and still be emitted UNQUOTED into the `bash -c "..."`
* command string. Intentionally narrower than "what bash tolerates": anything outside it
* gets single-quoted, so the classification can only ever err toward more quoting.
*/
export const SAFE_BARE_TOKEN = /^[A-Za-z0-9._:@=+/,-]+$/;
/**
* Longest `--version` output we will run a user-supplied regex over. A version banner is a
* line or two; anything larger is a misconfiguration, and capping the input is what keeps a
* sloppy (not necessarily malicious) regex from becoming a stall.
*/
export const MAX_VERSION_OUTPUT = 200;
/** Longest permitted `discovery.version.regex` source. */
const MAX_VERSION_REGEX_SOURCE = 200;
/**
* Nested quantifiers — `(a+)+`, `(a*)*`, `(a+)*` and friends — the classic catastrophic
* backtracking shape. Rejected outright rather than analysed: this field exists to pull a
* semver out of a banner, and nothing legitimate for that job needs a nested quantifier.
*/
const NESTED_QUANTIFIER = /\([^)]*[+*][^)]*\)\s*[+*{]/;
/**
* Compile a user-supplied version regex, or return null if it is not one we are willing to
* run. Returning null (rather than throwing) lets the caller degrade to "version unknown",
* which every consumer already handles.
*/
export function compileVersionRegex(source: string): RegExp | null {
if (source.length > MAX_VERSION_REGEX_SOURCE) return null;
if (NESTED_QUANTIFIER.test(source)) return null;
try {
// No `g`: a global regex carries lastIndex state across calls, which is a documented
// footgun in this codebase (see utils/regex-patterns.ts).
return new RegExp(source);
} catch {
return null;
}
}
+100
View File
@@ -0,0 +1,100 @@
/**
* @fileoverview The NAMES of code profiles a `CliEntry` field may select, and the helpers
* that validate them.
*
* A profile is the escape hatch for behaviour that is genuinely code-shaped and cannot be
* expressed as data — codex's predictive write-through echo, deepseek's profile-launcher
* runnability check, deepseek's status bridge — without letting any of that code branch on
* a CLI's id. A registry field names a profile; the implementation lives beside whatever it
* needs, and looks its name up here.
*
* ⚠️ This module is PURE and must stay that way: names, types and predicates only, no
* imports outside this directory. The implementations pull in resolvers and the status
* shim, which in turn reach back into the registry, so holding them here would close an
* import cycle (profiles → deepseek-cli-resolver → cli-resolver → registry → schema →
* profiles). Keeping the names here and the implementations at their call sites is what
* lets `schema.ts` validate a profile name at LOAD time — a custom entry naming a profile
* this build does not implement fails loudly instead of silently failing closed later.
*
* The rule all of this enforces: `test/cli-registry-no-id-branching.test.ts` fails on any
* `mode === '<stock id>'` comparison outside `stock.ts`, so a NEW behavioural special case
* must be added here, named, and referenced from a registry field — never inlined as an id
* check at the call site.
*
* ⚠️ A profile is a LAST resort, not a convenience. Reach for one only when the behaviour
* needs to run code (a side effect, a computed value, a probe); anything that is a list, a
* flag, or a string belongs in the entry as data, where a custom CLI can also use it.
*
* @module config/cli-registry/profiles
*/
/**
* Predictive local-echo profiles, selected via `capabilities.echo.predictProfile`.
*
* Implementation: packages/xterm-zerolag-input/src/predictive-echo-addon.ts.
*
* ⚠️ Unlike the other two registries, an unknown name here degrades to the 'buffer' policy
* rather than failing. Echo is a comfort feature — a worse-but-working overlay beats a
* refused session — which is why `predictProfile` alone is not schema-validated below.
*/
export const PREDICT_PROFILES: Record<string, true> = {
codex: true,
};
/**
* Launcher profiles, selected via `discovery.launcherProfile`.
*
* For a CLI whose binary launches some further target, and so cannot answer two questions
* from the binary alone: is it RUNNABLE (stricter than "is the binary on disk?"), and what
* is the DEFAULT target when the caller names none? A CLI naming no profile is runnable
* exactly when its binary resolves, and has no default target.
*
* Implementation: `src/utils/cli-launcher.ts`.
*/
export const LAUNCHER_PROFILE_NAMES = [
// `dsh` is a launcher over $DSH_HOME/profiles/<name>, and the profiles DeepSeek itself
// ships (web, headless) cannot drive a terminal pane. Binary AND a pane-capable profile.
'deepseek-profile',
] as const;
/**
* Extra `tmux setenv` work, selected via `env.setenvProfile`.
*
* Implementation: `src/tmux-manager.ts`, which already owns every setenv call.
*
* ⚠️ Anything that is merely "forward this name from the server's own env" belongs in
* `env.tmuxSetenvKeys` as data and must NOT be given a profile.
*/
export const SETENV_PROFILE_NAMES = [
// DeepSeek's terminal front door reports idle/working/blocked to a supervisor over the
// generic env-gated Herdr contract; this makes Codeman that supervisor. It needs a
// profile rather than key names because it writes an executable shim to disk and then
// exports that shim's path along with the session's own pane id.
'deepseek-status-bridge',
] as const;
export type LauncherProfileName = (typeof LAUNCHER_PROFILE_NAMES)[number];
export type SetenvProfileName = (typeof SETENV_PROFILE_NAMES)[number];
/**
* Transcript readers, selected via `capabilities.transcript`. Unlike the profile registries
* above this one is closed over the schema enum itself rather than an open string, since
* transcript format is a small, genuinely fixed set — see CliCapabilities['transcript'].
*/
export const TRANSCRIPT_READER_NAMES = ['claude-jsonl', 'codex-rollout', 'deepseek-zstd', 'none'] as const;
/** Composer-row finders, selected via `capabilities.echo.anchor.kind`. Also schema-closed. */
export const COMPOSER_ANCHOR_KINDS = ['glyph', 'cursor', 'none'] as const;
/** True when `name` is a predictive-echo profile this build actually implements. */
export function isKnownPredictProfile(name: string | undefined): boolean {
return name !== undefined && Object.prototype.hasOwnProperty.call(PREDICT_PROFILES, name);
}
export function isKnownLauncherProfile(name: string): name is LauncherProfileName {
return (LAUNCHER_PROFILE_NAMES as readonly string[]).includes(name);
}
export function isKnownSetenvProfile(name: string): name is SetenvProfileName {
return (SETENV_PROFILE_NAMES as readonly string[]).includes(name);
}
+235
View File
@@ -0,0 +1,235 @@
/**
* @fileoverview Loads, merges and re-validates the CLI registry.
*
* `~/.codeman/clis.json` holds OVERRIDES and CUSTOM entries only — never a full copy of the
* stock catalog — so a shipped fix to a stock definition actually reaches an existing
* install, and the file stays small enough to hand-edit.
*
* Resolution: start from `STOCK_CLIS` → deep-merge each override by id (objects merge
* key-wise, arrays replace wholesale) → validate every resulting entry. A stock entry that
* fails validation after merge falls back to its pristine stock definition (a fat-fingered
* override cannot brick a shipped CLI); a custom entry that fails is dropped with a warning
* rather than failing the whole load. Stock entries are always emitted, so `shell` and
* `claude` can be disabled but can never go missing — large parts of the app assume at
* minimum that a shell fallback exists.
*
* ⚠️ READ-ONLY. Nothing in this module writes, creates or migrates the file. That is a
* deliberate property, not a missing feature: there is no settings UI and no write API yet,
* so there is nothing to persist, and it means importing the registry — which
* `src/web/schemas.ts` does, transitively, just to validate a request — performs no
* filesystem writes. A `seededStockIds` ratchet belongs with the write API that needs it.
* The one exception is the quarantine RENAME of a file that fails to parse (see
* `readRegistryFile`), which happens on first use rather than at import.
*
* ⚠️ The file must be mode 0600. `isUnsafePermissions` refuses ANY group/world bit, read
* bits included, so a file created with a normal umask (0644) is ignored. Every reason the
* file was ignored, or an entry in it dropped, is logged ONCE on first load: the warnings
* used to be returned to a caller that nobody wired up, so a normally-created file was
* ignored with no feedback anywhere (found reviewing #347).
*
* @module config/cli-registry/registry
*/
import { existsSync, readFileSync, renameSync, statSync } from 'node:fs';
import { dataPath } from '../instance.js';
import type { CliEntry, CliId, CliRegistryFile } from './types.js';
import { CliEntrySchema } from './schema.js';
import { STOCK_CLIS } from './stock.js';
/** Construct a validated CliId. Throws if `raw` is not a well-formed id — call at API boundaries. */
export function asCliId(raw: string): CliId {
if (!/^[a-z][a-z0-9-]{0,23}$/.test(raw)) {
throw new Error(`invalid CLI id: ${JSON.stringify(raw)}`);
}
return raw as CliId;
}
function filePath(): string {
return dataPath('clis.json');
}
/**
* Keys that must never be merged out of a hand-editable JSON file.
*
* `JSON.parse` produces `__proto__` as an ORDINARY own property, but `result[key] = …` on a
* plain object walks the setter chain and would set the merged object's PROTOTYPE instead.
* Not exploitable today — every merged entry is spread into `{ ...merged, id, stock }` and
* then Zod-parsed before anything reads it, which drops the effect — but "not exploitable
* because of what a caller happens to do afterwards" is a property that quietly stops
* holding. A `continue` in the loop that reads the file is the cheap end of that trade.
*/
const UNMERGEABLE_KEYS = new Set(['__proto__', 'constructor', 'prototype']);
/** Plain-object deep merge: nested objects merge key-wise, arrays and primitives replace. */
function deepMerge<T>(base: T, override: unknown): T {
if (override === null || typeof override !== 'object' || Array.isArray(override)) {
return (override === undefined ? base : (override as T)) ?? base;
}
if (base === null || typeof base !== 'object' || Array.isArray(base)) {
return override as T;
}
const result: Record<string, unknown> = { ...(base as Record<string, unknown>) };
for (const [key, value] of Object.entries(override as Record<string, unknown>)) {
if (UNMERGEABLE_KEYS.has(key)) continue;
result[key] = deepMerge((base as Record<string, unknown>)[key], value);
}
return result as T;
}
export interface LoadResult {
entries: CliEntry[];
warnings: string[];
}
/**
* Refuse a registry file with any group/world permission bit — same posture as the ssh-key
* discipline, so 0600 is the only accepted mode. This file selects the binaries Codeman
* spawns, so a writable one is a way to redirect every session.
*
* POSIX only: Windows has no meaningful group/world bits on NTFS (Node reports every file
* as mode 0o666 there regardless of its actual ACL), so this check would flag every file on
* Windows and silently ignore all user config. `win32` relies on NTFS ACLs instead, which
* this check cannot see and does not attempt to.
*/
function isUnsafePermissions(path: string): boolean {
if (process.platform === 'win32') return false;
try {
const mode = statSync(path).mode & 0o777;
return (mode & 0o077) !== 0;
} catch {
return false;
}
}
function readRegistryFile(path: string, warnings: string[]): CliRegistryFile | null {
if (!existsSync(path)) return null;
if (isUnsafePermissions(path)) {
warnings.push(
`${path} must be mode 0600 (no group/world permission bits; run \`chmod 600 ${path}\`); ignoring it and falling back to stock CLIs.`
);
return null;
}
let raw: string;
try {
raw = readFileSync(path, 'utf-8');
} catch (err) {
warnings.push(`Failed to read ${path}: ${(err as Error).message}. Falling back to stock CLIs.`);
return null;
}
try {
const parsed = JSON.parse(raw) as CliRegistryFile;
if (typeof parsed !== 'object' || parsed === null || typeof parsed.clis !== 'object') {
throw new Error('missing "clis" object');
}
return parsed;
} catch (err) {
// QUARANTINE, never overwrite: the file is hand-editable, so a syntax error is far more
// likely to be a half-finished edit than junk. Renaming keeps the user's work.
const quarantined = `${path}.invalid-${Date.now()}`;
try {
renameSync(path, quarantined);
warnings.push(`${path} was not valid JSON (${(err as Error).message}); moved to ${quarantined}.`);
} catch {
warnings.push(
`${path} was not valid JSON (${(err as Error).message}); left in place, falling back to stock CLIs.`
);
}
return null;
}
}
/**
* Merge the stock catalog with a (possibly absent) registry file. PURE — no IO, which is
* what lets the load tests drive every merge case directly.
*/
export function resolveRegistry(stock: CliEntry[], file: CliRegistryFile | null, warnings: string[]): LoadResult {
const stockById = new Map(stock.map((e) => [e.id as string, e]));
const overrides = file?.clis ?? {};
const entries: CliEntry[] = [];
for (const stockEntry of stock) {
const id = stockEntry.id as string;
const override = overrides[id];
const merged = override ? deepMerge(stockEntry, override) : stockEntry;
// `stock: true` is forced here rather than read from the merged object, so an override
// can never flip a custom entry's provenance or vice versa.
const parsed = CliEntrySchema.safeParse({ ...merged, id, stock: true });
if (parsed.success) {
entries.push(parsed.data as CliEntry);
} else {
warnings.push(
`Override for stock CLI "${id}" failed validation; using the shipped definition. ${parsed.error.message}`
);
entries.push(stockEntry);
}
}
for (const [id, raw] of Object.entries(overrides)) {
if (stockById.has(id)) continue; // already merged above
// Same forcing in the other direction: a custom entry claiming `stock: true` cannot
// shadow or impersonate a shipped one.
const parsed = CliEntrySchema.safeParse({ ...(raw as object), id, stock: false });
if (parsed.success) {
entries.push(parsed.data as CliEntry);
} else {
warnings.push(`Custom CLI "${id}" failed validation and was dropped. ${parsed.error.message}`);
}
}
entries.sort((a, b) => a.order - b.order);
return { entries, warnings };
}
let cache: LoadResult | null = null;
/**
* Load the effective registry (stock + user overrides). Memoized for the process lifetime;
* `reloadCliRegistry()` invalidates.
*/
export function loadCliRegistry(): LoadResult {
if (cache) return cache;
const warnings: string[] = [];
const existing = readRegistryFile(filePath(), warnings);
cache = resolveRegistry(STOCK_CLIS, existing, warnings);
// Once per process (the result is memoized): silence here is what made a 0644 file look
// like "the override feature does nothing".
for (const warning of warnings) console.warn(`[cli-registry] ${warning}`);
return cache;
}
/** Drop the memoized registry so the next `loadCliRegistry()` re-reads the file. */
export function reloadCliRegistry(): void {
cache = null;
}
export function listClis(): CliEntry[] {
return loadCliRegistry().entries;
}
export function enabledClis(): CliEntry[] {
return listClis().filter((e) => e.enabled);
}
export function getCli(id: string): CliEntry | undefined {
return listClis().find((e) => (e.id as string) === id);
}
export function cliIds(): string[] {
return listClis().map((e) => e.id as string);
}
/** Every enabled entry's id, in registry order. */
export function enabledCliIds(): string[] {
return enabledClis().map((e) => e.id as string);
}
/**
* Resolve the install command for the current platform, falling back to the linux one (the
* common case for a `curl | bash` or `npm install -g` line) and then to whatever is
* declared. Display text only — never executed. See CliDiscovery.install.command.
*/
export function resolveInstallCommandForPlatform(entry: CliEntry): string | undefined {
const { command } = entry.discovery.install;
const platform = process.platform as 'linux' | 'darwin' | 'win32';
return command[platform] ?? command.linux ?? Object.values(command)[0];
}
+428
View File
@@ -0,0 +1,428 @@
/**
* @fileoverview Zod validation for CLI registry entries.
*
* Every object here is `.strict()`: an unknown key is a hard validation error, not a
* silently-ignored one. That matters for a security-relevant schema — a typo in a field name
* must never degrade to "field absent, so the permissive default applies".
*
* The load-bearing rule enforced here is `SHELL_TOKEN`: it is what makes it impossible for a
* `clis.json` entry to smuggle shell metacharacters into the eventual `bash -c "..."` string
* (see argv.ts's file header for the full model).
*
* @module config/cli-registry/schema
*/
import { z } from 'zod';
import { TOKEN_PATTERNS } from './patterns.js';
import { isKnownLauncherProfile, isKnownSetenvProfile } from './profiles.js';
/** A bare CLI id: lowercase, starts with a letter, at most 24 chars. Also used as a CSS/URL token. */
const cliId = z
.string()
.regex(/^[a-z][a-z0-9-]{0,23}$/, 'id must be lowercase, start with a letter, and be at most 24 chars');
/** An env var name. */
const envName = z
.string()
.regex(/^[A-Z_][A-Z0-9_]*$/, 'env var name must be UPPER_SNAKE_CASE')
.max(64);
/**
* A shell-safe bare word: no space, quote, backtick, `$`, `;`, `&`, `|`, `<`, `>`, parens,
* braces, newline or backslash. Every LITERAL in the launch spec (base command, flag names,
* fixed values) must satisfy this — see argv.ts's file header.
*/
const shellToken = z
.string()
.min(1)
.max(256)
.regex(/^[A-Za-z0-9._:@=+/,-]+$/, 'must be a plain word with no shell metacharacters');
const flagToken = z.string().regex(/^--?[A-Za-z0-9][A-Za-z0-9-]*$/, 'must look like -x or --long-flag');
const quoteStyle = z.enum(['auto', 'bare', 'double', 'single']);
const condSchema: z.ZodType<import('./types.js').Cond> = z.lazy(() =>
z.union([
z.object({ param: z.string(), is: z.union([z.string(), z.boolean()]) }).strict(),
z.object({ param: z.string(), state: z.enum(['set', 'unset']) }).strict(),
z.object({ allOf: z.array(condSchema).min(1).max(8) }).strict(),
z.object({ anyOf: z.array(condSchema).min(1).max(8) }).strict(),
z.object({ not: condSchema }).strict(),
z.object({ capabilityGate: z.string() }).strict(),
])
);
const paramSpecSchema = z.union([
z
.object({ type: z.literal('enum'), values: z.array(z.string()).min(1).max(16), default: z.string().optional() })
.strict(),
z.object({ type: z.literal('bool') }).strict(),
z.object({ type: z.literal('token'), pattern: z.enum(TOKEN_PATTERNS as [string, ...string[]]) }).strict(),
z
.object({
type: z.literal('engine'),
source: z.enum([
'sessionId',
'sessionName',
'muxName',
'effortLevel',
'effortSettingsJson',
'codemanPrefixedSessionId',
'launcherDefaultTarget',
]),
})
.strict(),
]);
const argSpecSchema = z.union([
z.object({ lit: shellToken, when: condSchema.optional() }).strict(),
z.object({ flag: flagToken, when: condSchema.optional() }).strict(),
z.object({ flag: flagToken, value: shellToken, quote: quoteStyle.optional(), when: condSchema.optional() }).strict(),
z
.object({ flag: flagToken, valueFrom: z.string(), quote: quoteStyle.optional(), when: condSchema.optional() })
.strict(),
z.object({ valueFrom: z.string(), quote: quoteStyle.optional(), when: condSchema.optional() }).strict(),
]);
const variantSchema = z
.object({
id: z.string().min(1).max(40),
when: condSchema.optional(),
// min(0): the `shell` entry declares a variant with no args — tmux-manager resolves the
// real login shell in code, since it varies per remote user's /etc/passwd entry.
args: z.array(argSpecSchema).max(32),
})
.strict();
const launchSchema = z
.object({
params: z.record(z.string(), paramSpecSchema),
chain: z.enum(['first', 'fallback']).optional(),
variants: z.array(variantSchema).min(1).max(4),
legacyConfigAliases: z.record(z.string(), z.string()).optional(),
legacyConfigField: z.string().min(1).max(40).optional(),
resumeAppend: z
.union([
z.object({ style: z.literal('flag'), flag: flagToken }).strict(),
z.object({ style: z.literal('positional'), token: shellToken }).strict(),
])
.optional(),
})
.strict()
.superRefine((launch, ctx) => {
const paramNames = new Set(Object.keys(launch.params));
const checkValueFrom = (name: string, path: (string | number)[]) => {
if (!paramNames.has(name)) {
ctx.addIssue({ code: 'custom', message: `valueFrom "${name}" is not a declared param`, path });
}
};
launch.variants.forEach((variant, vi) => {
variant.args.forEach((arg, ai) => {
if ('valueFrom' in arg) checkValueFrom(arg.valueFrom, ['variants', vi, 'args', ai, 'valueFrom']);
});
});
if (launch.chain === 'fallback') {
const last = launch.variants.at(-1);
if (last?.when) {
ctx.addIssue({
code: 'custom',
message: 'the last variant of a fallback chain must have no `when` (it must be the guaranteed terminal case)',
path: ['variants', launch.variants.length - 1, 'when'],
});
}
}
if (launch.legacyConfigAliases) {
for (const paramName of Object.keys(launch.legacyConfigAliases)) {
if (!paramNames.has(paramName)) {
ctx.addIssue({
code: 'custom',
message: `legacyConfigAliases key "${paramName}" is not a declared param`,
path: ['legacyConfigAliases', paramName],
});
}
}
}
});
const versionProbeSchema = z
.object({
arg: shellToken,
regex: z.string().max(200).optional(),
requireVersionMatch: z.boolean().optional(),
retryOnTransientFailure: z.boolean().optional(),
})
.strict();
const identityProbeSchema = z
.object({
arg: shellToken,
// Same 200-char cap as version.regex, and compiled through the same compileVersionRegex()
// guard at use time. This is the second and last config-supplied regex in the registry.
regex: z.string().min(1).max(200),
})
.strict();
const discoverySchema = z
.object({
// min(0): the `shell` entry has no binary of its own (it resolves the login shell in code).
binaries: z.array(shellToken).max(4),
searchDirs: z.array(z.string().max(300)).max(16),
version: versionProbeSchema.optional(),
identity: identityProbeSchema.optional(),
launcherProfile: z.string().max(40).optional(),
launcherTargetParam: z.string().max(40).optional(),
install: z
.object({
// z.record with an enum key type requires every enum member in Zod v4; the install
// command legitimately varies by platform and most entries only need one or two, so
// this is a plain object of optional platform keys instead.
command: z
.object({
linux: z.string().max(500).optional(),
darwin: z.string().max(500).optional(),
wsl: z.string().max(500).optional(),
win32: z.string().max(500).optional(),
})
.strict(),
npmPackage: z.string().max(200).optional(),
docsUrl: z.url().optional(),
})
.strict(),
})
.strict();
const envExportSchema = z
.object({
name: envName,
value: z.union([
shellToken,
z
.object({
engine: z.enum([
'sessionId',
'sessionName',
'muxName',
'effortLevel',
'effortSettingsJson',
'codemanPrefixedSessionId',
'launcherDefaultTarget',
]),
})
.strict(),
]),
when: condSchema.optional(),
})
.strict();
const envSchema = z
.object({
exports: z.array(envExportSchema).max(16),
unset: z.array(envName).max(16),
tmuxSetenvKeys: z.array(envName).max(32),
dockerExecEnvNames: z.array(envName).max(32),
configSetenv: z
.array(z.object({ name: envName, fromParam: z.string().min(1).max(40) }).strict())
.max(8)
.optional(),
allowedPrefixes: z
.array(
z
.string()
.min(3)
.max(32)
.regex(/^[A-Z][A-Z0-9_]*_$/)
)
.max(8),
allowedKeys: z.array(envName).max(8),
configContentVar: envName.optional(),
setenvProfile: z.string().max(40).optional(),
})
.strict();
const echoSchema = z
.object({
policy: z.enum(['buffer', 'predict', 'off']),
anchor: z.union([
z
.object({ kind: z.literal('glyph'), glyph: z.string().min(1).max(4), offset: z.number().int().min(0).max(16) })
.strict(),
z.object({ kind: z.literal('cursor') }).strict(),
z.object({ kind: z.literal('none') }).strict(),
]),
predictProfile: z.string().max(40).optional(),
})
.strict();
const capabilitiesSchema = z
.object({
external: z.boolean(),
requiresMux: z.boolean(),
hooks: z.enum(['none', 'always', 'supervised']),
transcript: z.enum(['claude-jsonl', 'codex-rollout', 'deepseek-zstd', 'omp-jsonl', 'none']),
altScreen: z.enum(['strip-full', 'strip-mux-only', 'preserve']),
echo: echoSchema,
wheelForward: z
.object({ mode: z.enum(['never', 'version-gated']), minVersion: z.string().max(20).optional() })
.strict(),
keyboardAccessory: z.enum(['agent', 'shell']),
privilegedCommandGate: z.boolean(),
startMode: z.enum(['interactive', 'shell']),
stripInkBloat: z.boolean(),
ralph: z.boolean(),
respawn: z.boolean(),
effort: z.boolean(),
agentSkillInjection: z.boolean(),
statusLineTelemetry: z.boolean(),
model: z
.object({ source: z.enum(['flag', 'claude-settings-file', 'none']), param: z.string().optional() })
.strict(),
privilegedParams: z
.array(
z
.object({
param: z.string(),
clampTo: z.union([z.boolean(), z.string()]),
materializeWhenAbsent: z.boolean().optional(),
})
.strict()
)
.max(8),
// Exact env var NAMES, not prefixes: this list is a targeted deny, and a prefix here
// would let one entry silently strip a whole namespace off every owner's overrides.
privilegedEnvKeys: z.array(envName).max(8),
gates: z.record(z.string(), z.object({ minVersion: z.string().max(20), failClosed: z.boolean() }).strict()),
maxFrameBytes: z.number().int().positive().optional(),
})
.strict();
const credStoreSchema = z
.object({
rel: z.string().min(1).max(100),
shareDirs: z.array(z.string().max(100)).optional(),
shareFiles: z.array(z.string().max(100)).optional(),
seedFiles: z.array(z.string().max(100)).optional(),
seedWhole: z.boolean().optional(),
})
.strict();
/**
* A remote/docker default pane command: space-separated bare words from the SAME safe
* charset as `shellToken` (no shell metacharacters), so `claude --dangerously-skip-permissions`
* is expressible while still excluding `;`, `|`, `$`, backticks and quotes — this is not an
* escape hatch into arbitrary shell text, it is one bare command plus bare flags.
*/
const commandLine = z
.string()
.min(1)
.max(200)
.regex(
/^[A-Za-z0-9._:@=+/,-]+( [A-Za-z0-9._:@=+/,-]+)*$/,
'must be space-separated bare words with no shell metacharacters'
);
const overlayTargetSchema = z.union([
z.object({ command: commandLine.optional() }).strict(),
z.object({ disabled: z.literal(true) }).strict(),
]);
const overlaysSchema = z
.object({
remote: overlayTargetSchema.optional(),
docker: overlayTargetSchema.optional(),
credStore: credStoreSchema.optional(),
})
.strict();
export const CliEntrySchema = z
.object({
id: cliId,
label: z.string().min(1).max(60),
shortBadge: z.string().min(1).max(6),
accent: z.string().regex(/^#[0-9a-fA-F]{6}$/, 'accent must be a 6-digit hex colour'),
enabled: z.boolean(),
stock: z.boolean(),
order: z.number().int(),
kind: z.enum(['agent', 'shell']),
discovery: discoverySchema,
launch: launchSchema,
env: envSchema,
capabilities: capabilitiesSchema,
overlays: overlaysSchema,
})
.strict()
.superRefine((entry, ctx) => {
const gateNames = new Set(Object.keys(entry.capabilities.gates));
const walkConds = (cond: import('./types.js').Cond | undefined) => {
if (!cond) return;
if ('capabilityGate' in cond && !gateNames.has(cond.capabilityGate)) {
ctx.addIssue({
code: 'custom',
message: `capabilityGate "${cond.capabilityGate}" is not declared in capabilities.gates`,
});
}
if ('allOf' in cond) cond.allOf.forEach(walkConds);
if ('anyOf' in cond) cond.anyOf.forEach(walkConds);
if ('not' in cond) walkConds(cond.not);
};
for (const variant of entry.launch.variants) {
walkConds(variant.when);
for (const arg of variant.args) walkConds(arg.when);
}
// Reject a profile name this build does not implement, rather than letting it fail
// closed at use time. An unimplemented `launcherProfile` would make the CLI look
// permanently uninstalled, and an unimplemented `setenvProfile` would silently skip
// setup the CLI needs; both are far easier to diagnose as a load-time error naming the
// field. (`echo.predictProfile` is deliberately NOT checked here — see profiles.ts.)
const { launcherProfile } = entry.discovery;
if (launcherProfile !== undefined && !isKnownLauncherProfile(launcherProfile)) {
ctx.addIssue({
code: 'custom',
message: `discovery.launcherProfile "${launcherProfile}" is not a profile this build implements`,
path: ['discovery', 'launcherProfile'],
});
}
// An env var exported from a param that does not exist would silently export nothing,
// and for DSH_PERMISSION_MODE that means silently losing a permission clamp.
const declaredParams = new Set(Object.keys(entry.launch.params));
entry.env.configSetenv?.forEach((mapping, i) => {
if (!declaredParams.has(mapping.fromParam)) {
ctx.addIssue({
code: 'custom',
message: `configSetenv fromParam "${mapping.fromParam}" is not a declared launch param`,
path: ['env', 'configSetenv', i, 'fromParam'],
});
}
});
// Same class of silent failure on the OTHER privileged surface, and this one is a
// security control: `privilegedParams[].param` is the multi-user bypass clamp's only
// handle on a CLI's privilege switch, and a name that is not a declared param clamps
// NOTHING — no load error, no failing test, the clamp simply stops running. The clamp
// resolves the name through `legacyConfigAliases`, so this check is what keeps the two
// in ONE namespace rather than two that merely coincide today: they do not for codex
// (`bypassApprovals` vs `dangerouslyBypassApprovals`), and giving deepseek's
// `permissionMode` an alias later would otherwise have removed its clamp with nothing
// saying so.
entry.capabilities.privilegedParams.forEach((clamp, i) => {
if (!declaredParams.has(clamp.param)) {
ctx.addIssue({
code: 'custom',
message: `privilegedParams param "${clamp.param}" is not a declared launch param`,
path: ['capabilities', 'privilegedParams', i, 'param'],
});
}
});
const { setenvProfile } = entry.env;
if (setenvProfile !== undefined && !isKnownSetenvProfile(setenvProfile)) {
ctx.addIssue({
code: 'custom',
message: `env.setenvProfile "${setenvProfile}" is not a profile this build implements`,
path: ['env', 'setenvProfile'],
});
}
});
export type ValidatedCliEntry = z.infer<typeof CliEntrySchema>;
File diff suppressed because it is too large Load Diff
+528
View File
@@ -0,0 +1,528 @@
/**
* @fileoverview Type definitions for the CLI registry — the single source of truth for
* which agent CLIs Codeman supports and how each one is discovered, launched and treated.
*
* This replaces the hard-coded `SessionMode` union and the ~123 per-mode branches that grew
* out of it. The guiding rule: NO code may branch on a CLI's id. Behaviour that genuinely
* differs between CLIs is expressed either as data here, or as a named PROFILE selected by
* a capability field (see profiles.ts) — never as `mode === 'codex'`.
*
* @module config/cli-registry/types
*/
import type { TokenPattern } from './patterns.js';
/**
* A CLI identifier. Branded so an arbitrary string cannot be passed where a validated id is
* expected; construct with `asCliId()` at the API boundary.
*/
export type CliId = string & { readonly __cliId: unique symbol };
// ---------------------------------------------------------------------------
// Launch argv DSL
// ---------------------------------------------------------------------------
/** Values the ENGINE supplies. Config may reference these by name but never author them. */
export type EngineValue =
| 'sessionId'
| 'sessionName'
| 'muxName'
| 'effortLevel'
| 'effortSettingsJson'
/** `sessionId` prefixed `codeman_<id>` — codex's unique per-pane rollout originator. */
| 'codemanPrefixedSessionId'
/**
* For a launcher CLI (`discovery.launcherProfile`), the target to launch when the caller
* named none — deepseek's default `dsh` profile. Resolved at spawn time, never frozen
* into config, because it depends on what is installed on this machine right now.
*/
| 'launcherDefaultTarget';
/**
* A declared launch parameter. `token` params carry caller-supplied data and are therefore
* the only ones that need a pattern; `engine` params are produced in code.
*/
export type ParamSpec =
| { type: 'enum'; values: string[]; default?: string }
| { type: 'bool' }
| { type: 'token'; pattern: TokenPattern }
| { type: 'engine'; source: EngineValue };
/** A boolean guard over parameter state. */
export type Cond =
| { param: string; is: string | boolean }
| { param: string; state: 'set' | 'unset' }
| { allOf: Cond[] }
| { anyOf: Cond[] }
| { not: Cond }
/** Names an entry in `capabilities.gates`. Fail-closed gates omit when version is unknown. */
| { capabilityGate: string };
/**
* How a token is quoted when emitted into the bash command string.
*
* This exists ONLY to preserve byte-identical output with the hand-written builders being
* replaced (claude wraps its values in double quotes; the other builders emit bare words).
* It is never a safety lever: `renderToken()` verifies the value is metacharacter-free
* before honouring an explicit style, and falls back to single-quote escaping if it is not.
* So the worst a wrong `quote` can do is make output uglier, never unsafe.
*/
export type QuoteStyle = 'auto' | 'bare' | 'double' | 'single';
/** One argv element. */
export type ArgSpec =
/** A bare literal word, e.g. the base binary or codex's `resume` subcommand. */
| { lit: string; when?: Cond }
/** A valueless flag, e.g. `--no-approve`. */
| { flag: string; when?: Cond }
/** A flag with a fixed literal value. */
| { flag: string; value: string; quote?: QuoteStyle; when?: Cond }
/** A flag whose value comes from a declared param. */
| { flag: string; valueFrom: string; quote?: QuoteStyle; when?: Cond }
/** A bare positional value from a param, e.g. codex's `resume <id>`. */
| { valueFrom: string; quote?: QuoteStyle; when?: Cond };
/** One alternative command form. */
export interface CliVariant {
/** Stable name for diagnostics and tests, e.g. 'resume' / 'new'. */
id: string;
when?: Cond;
args: ArgSpec[];
}
export interface CliLaunch {
params: Record<string, ParamSpec>;
/**
* 'first' — emit the first variant whose `when` passes (the usual case).
* 'fallback' — emit EVERY passing variant joined by the engine's own ` || `, which is how
* claude's `--resume X || --session-id Y` shell fallback is expressed without
* config ever containing shell text. The engine owns the operator.
*/
chain?: 'first' | 'fallback';
variants: CliVariant[];
/**
* Maps a declared param name to the field name it arrives under on the legacy
* `POST /api/sessions` wire shape (`OpenCodeConfig.continueSession`, etc — the per-mode
* config objects predate this registry and stay on the wire for compatibility). A param
* with no entry here is looked up under its own name. This is what lets the spawn-command
* bridge (`session-cli-registry-bridge.ts`) stay generic: it reads the raw legacy config
* object through this DATA-declared alias table instead of a per-mode `if (mode === ...)`.
*/
legacyConfigAliases?: Record<string, string>;
/**
* The field on the legacy spawn option bag holding this CLI's `<Mode>Config` object
* (`openCodeConfig`, `codexConfig`, …). Those per-mode objects predate this registry and
* stay on the wire for API compatibility, so SOMETHING has to know which one to read —
* declaring it here as data is what keeps the bridge a generic reader instead of a
* `switch (mode)`.
*
* ABSENT means this CLI's launch fields live at the TOP LEVEL of the option bag rather
* than nested in a config object. That is claude, whose discrete `claudeMode` /
* `allowedTools` / `model` / `resumeSessionId` fields predate the `<Mode>Config` pattern
* entirely — so "read the option bag itself" is not a special case for it, it is just
* the other shape.
*/
legacyConfigField?: string;
/**
* How to APPEND a resume id onto an already-built base command, for the docker in-container
* "tmux was re-created, resume the surviving transcript" path (`appendResumeFlag` in
* tmux-manager.ts) — a narrower, append-only sibling of the full `variants` shape above,
* which builds a whole command from scratch. Absent = this CLI has no resume flag to
* append (shell, opencode: opencode's docker resume goes through its own config object).
*/
resumeAppend?: { style: 'flag'; flag: string } | { style: 'positional'; token: string };
}
// ---------------------------------------------------------------------------
// Discovery
// ---------------------------------------------------------------------------
export interface CliVersionProbe {
arg: string;
/** Serialized regex, applied to `--version` output only. See compileVersionRegex(). */
regex?: string;
/**
* Treat a binary whose version output does not match as ABSENT rather than as
* present-with-unknown-version. For CLIs with short, generic binary names (`pi`), where a
* `which` hit is not by itself evidence the right program is installed.
*/
requireVersionMatch?: boolean;
/** Retry a failed probe with backoff instead of caching the failure (claude's behaviour). */
retryOnTransientFailure?: boolean;
}
/**
* An identity probe: proof that the binary we found is the program we meant, not an
* unrelated one that happens to share the name.
*
* A version probe is not enough on its own. Debian ships a `dsh` (dancer's shell) that
* answers `--version` perfectly happily, and npm carries squatters for `pi` and `grok`.
* `requireVersionMatch` catches a binary whose version output has the WRONG SHAPE; this
* catches one whose output has the right shape but names the wrong program.
*
* Ordering matters and belongs to the resolver, not to config: identity is checked FIRST,
* so an impostor is rejected before its version string is ever parsed.
*/
export interface CliIdentityProbe {
/** Argument that makes the binary describe itself, e.g. `--help`. */
arg: string;
/**
* Serialized regex the output must match. Compiled through `compileVersionRegex()`, so
* it inherits the same length cap and nested-quantifier rejection — this is the second
* (and last) config-supplied regex in the registry, and it runs against truncated
* command output exactly like the first.
*/
regex: string;
}
export interface CliDiscovery {
/**
* Binary name(s), first hit wins.
*
* This is why the registry fixes a live bug: the mode name is NOT always the binary
* name (`antigravity` runs `agy`), and `probeDockerCliVersion` assumed it was.
*/
binaries: string[];
/** Extra directories probed after `which`. A leading `~` expands to homedir; nothing else. */
searchDirs: string[];
version?: CliVersionProbe;
/** Proof the binary is the right program, checked BEFORE the version probe. */
identity?: CliIdentityProbe;
/**
* Names a LAUNCHER profile (profiles.ts): this CLI's binary is a launcher over some
* further target, so two questions the registry normally answers from the binary alone
* have to be asked of that target instead.
*
* - Is it RUNNABLE? Stricter than "is the binary on disk?".
* - What is the DEFAULT target, when the caller names none?
*
* DeepSeek is why this exists and is its only user. `dsh` launches a profile from
* `$DSH_HOME/profiles/<name>`, and the profiles DeepSeek itself ships (`web`,
* `headless`) cannot drive a terminal pane — so a perfectly-installed `dsh` with no
* third-party TUI profile is installed-but-NOT-runnable. The Run button gates on
* runnability while the "add a profile" affordance gates on mere availability;
* collapsing the two would either hide the affordance that fixes the problem or offer a
* run that always fails.
*
* The default target reaches the launch spec as the `launcherDefaultTarget` engine
* value, so it stays a runtime lookup rather than a value frozen into config.
*
* Absent (the normal case) means the binary IS the program, and its presence IS
* runnability.
*/
launcherProfile?: string;
/**
* The launch param naming the target a caller asked for, so the launcher profile can say
* why THAT specific target will not start rather than only whether any will. Meaningless
* without `launcherProfile`.
*/
launcherTargetParam?: string;
install: {
/**
* DISPLAY TEXT ONLY. Shown verbatim in "CLI not found. Install with: ...".
*
* ⚠️ NEVER executed by the server. That is a documented invariant, not an oversight:
* running it would turn a config file into a code-execution surface. A proposal to
* execute this on enable is deliberately deferred to its own change so the trust
* model can be decided on its own merits rather than inside a refactor.
*/
command: Partial<Record<'linux' | 'darwin' | 'wsl' | 'win32', string>>;
/** Package name for an npm-installable CLI. Display/tooling metadata only. */
npmPackage?: string;
docsUrl?: string;
};
}
// ---------------------------------------------------------------------------
// Environment
// ---------------------------------------------------------------------------
export interface CliEnv {
/** `export K=V` in the bash prelude. Values are literals or engine values, never secrets. */
exports: Array<{ name: string; value: string | { engine: EngineValue }; when?: Cond }>;
/** `unset K` — e.g. claude's CLAUDECODE, the truecolor CLIs' NO_COLOR. */
unset: string[];
/**
* NAMES ONLY. Values are read from the server's own process.env and pushed via
* `tmux setenv`, so a secret is structurally unable to reach the command line.
*/
tmuxSetenvKeys: string[];
/** NAMES ONLY, forwarded as `docker exec -e NAME`. */
dockerExecEnvNames: string[];
/**
* Env vars set via `tmux setenv` from a LAUNCH PARAM rather than from the server's own
* environment — for a CLI whose switch is an env var instead of a flag.
*
* DeepSeek's `DSH_PERMISSION_MODE` is the case this exists for. Routing it through a
* declared param (rather than a bespoke configure step) is what lets the ordinary
* `privilegedParams` clamp apply to it: the clamp rewrites the param, and whatever the
* param ends up as is what gets exported.
*
* ⚠️ Values are read from a declared, schema-validated param, never from free text, and
* they reach the pane through `tmux setenv` rather than the command line.
*/
configSetenv?: Array<{ name: string; fromParam: string }>;
/** This entry's contribution to the env-override allowlist. Never widens BLOCKED_ENV_KEYS. */
allowedPrefixes: string[];
allowedKeys: string[];
/**
* Env var carrying a JSON config blob pushed via `tmux setenv` (opencode's
* OPENCODE_CONFIG_CONTENT). Generic so it is not an opencode special case.
*/
configContentVar?: string;
/**
* Names an entry in `SETENV_PROFILES` (profiles.ts): extra `tmux setenv` work that is
* genuinely code-shaped rather than a list of key names.
*
* DeepSeek's status bridge is the only current user. It has to write an executable shim
* to disk (`ensureDeepSeekStatusShim()`), then export the shim's path and this session's
* pane id — a side effect and two computed values, none of which `tmuxSetenvKeys` (a
* list of names forwarded from the server's own env) can express.
*
* Plain secret forwarding stays in `tmuxSetenvKeys` and must NOT move here.
*/
setenvProfile?: string;
}
// ---------------------------------------------------------------------------
// Capabilities
// ---------------------------------------------------------------------------
/**
* The closed set of behavioural switches. Each field replaces an id-check somewhere.
*
* `hooks`, `transcript` and `altScreen` are INDEPENDENT on purpose. The three predicates
* they back (`hooksAvailableForMode`, `isExternalCliMode`, `isAltScreenStripMode`) describe
* three different, deliberately unequal sets, and deriving any one from another has already
* caused a real bug — a `shell` session has no hooks but is not an "external CLI", so
* `!isExternalCliMode()` wrongly accepted `until=stop` on it and hung for the full timeout.
* Keeping them as separate fields makes that invariant structural rather than commented.
*/
export interface CliCapabilities {
/**
* Non-Claude run mode that uses its own TUI and output format (`isExternalCliMode`):
* no Claude transcript, no hooks, no Claude-format token/BashTool parsing. An explicit
* field rather than derived from `hooks`/`kind`, precisely because it must stay
* independent — see this interface's own doc comment.
*/
external: boolean;
/** No direct-PTY fallback: the CLI must run inside tmux (secrets ride tmux setenv). */
requiresMux: boolean;
/**
* Whether `stop`/`blocked` wait signals can ever fire for this CLI.
*
* ⚠️ A TRI-STATE, not a boolean, because for one CLI this is a per-SESSION question:
* 'none' — no hook signals, ever (every external CLI, and `shell`).
* 'always' — the CLI installs Codeman's hooks (claude).
* 'supervised' — the CLI REPORTS its own idle/working/blocked state to a supervisor
* over a generic env-gated contract, and Codeman is that supervisor
* (deepseek, via deepseek-status-shim.ts). Definitive rather than
* inferred, so it earns real signals — but the session can disarm the
* bridge (`deepSeekConfig.statusReporting: false`), and a docker or
* remote session cannot reach it at all.
*
* That last case is why `hooksAvailableForMode()` takes per-session options and why
* every call site must pass `sessionHookOptions(session)`. Answering from the mode alone
* would promise a `stop` that never arrives, which is the infinite-wait-dressed-as-a-
* timeout the predicate exists to prevent.
*/
hooks: 'none' | 'always' | 'supervised';
/**
* Which transcript reader, if any, understands this CLI's on-disk history.
*
* `deepseek-zstd` is the odd one out: dsh writes zstd-compressed session files and
* appends ONE FRAME PER WRITE, so it needs a reader that walks frame headers itself
* rather than the stock decoder. It exists because the pane segmenter served dsh's
* ASCII-art splash as the worker's first answer.
*/
transcript: 'claude-jsonl' | 'codex-rollout' | 'deepseek-zstd' | 'omp-jsonl' | 'none';
/**
* 'strip-full' — alt-screen + erase-scrollback + mouse DECSETs stripped (Ink TUIs).
* 'strip-mux-only' — only tmux's own attach-time smcup (the safe default).
* 'preserve' — leave everything (a direct-PTY shell running vim/less/htop).
*/
altScreen: 'strip-full' | 'strip-mux-only' | 'preserve';
echo: {
policy: 'buffer' | 'predict' | 'off';
/** How the local-echo overlay locates the composer row. */
anchor: { kind: 'glyph'; glyph: string; offset: number } | { kind: 'cursor' } | { kind: 'none' };
/** Names a PREDICT_PROFILES key. Unknown or absent degrades to 'buffer', never to broken. */
predictProfile?: string;
};
/** Forwarding the wheel to the CLI's own transcript. 'never' keeps local scrollback. */
wheelForward: { mode: 'never' | 'version-gated'; minVersion?: string };
keyboardAccessory: 'agent' | 'shell';
/** Multi-user: this CLI is a raw shell, so its commands need the privileged gate. */
privilegedCommandGate: boolean;
startMode: 'interactive' | 'shell';
stripInkBloat: boolean;
ralph: boolean;
respawn: boolean;
effort: boolean;
agentSkillInjection: boolean;
statusLineTelemetry: boolean;
/** Where a model override is delivered. Claude uniquely writes settings.local.json. */
model: { source: 'flag' | 'claude-settings-file' | 'none'; param?: string };
/**
* Params a non-granted multi-user owner may not set freely, and what they are forced to.
* Data-driven so a CUSTOM CLI's bypass flag is clampable exactly like codex's.
*
* `materializeWhenAbsent` distinguishes two real shapes, not one:
* - only-if-sent (false/omitted; codex, antigravity, grok): the CLI's own
* absent-config default already spawns safe, so the clamp should only touch
* a config the caller actually sent.
* - materialize (true; gemini, pi): the absent-config default is ITSELF unsafe
* for a non-granted owner (gemini defaults to `yolo`; pi's absent default is
* an interactive trust prompt the session user could just answer "yes" to),
* so the clamp must CREATE a config object even when none was sent.
*
* ⚠️ `param` names the LAUNCH PARAM, like every other `param` in this file — never the
* legacy wire field. The clamp translates it through `legacyConfigAliases` on the way out,
* the same hop `env.configSetenv` makes. The two names coincide for most entries and
* DELIBERATELY do not for codex (`bypassApprovals` here, `dangerouslyBypassApprovals` on
* the wire), which is what keeps the distinction visible. `schema.ts` rejects an entry
* naming a param it never declared, because getting this wrong is a SILENT no-op: no load
* error, no failing test, the clamp just stops clamping.
*/
privilegedParams: Array<{ param: string; clampTo: boolean | string; materializeWhenAbsent?: boolean }>;
/**
* Env var names a non-granted multi-user owner may not set at all, DROPPED from
* `envOverrides` before spawn.
*
* ⚠️ This is a second, structurally different privileged surface from `privilegedParams`
* above, and one cannot substitute for the other. `privilegedParams` clamps a field on a
* per-CLI config object, which reaches the CLI as an argv flag. These clamp env vars,
* which reach it through `tmux setenv` — a path no argv clamp can see.
*
* DeepSeek is why this exists. Its permission switch IS an env var
* (`DSH_PERMISSION_MODE`), not a flag, so a config-level clamp alone leaves a real
* multi-user control with nothing enforcing it. Worse, `DSH_*` is an allowlisted
* `envOverrides` prefix and `applyEnvOverrides()` runs AFTER the per-CLI env configure
* step, so a non-granted owner sending that key on the SAME request would land last and
* hand back exactly the privilege the config clamp just removed.
*
* Dropping (rather than rewriting) is deliberate: the value then falls through to what
* the CLI's own env configuration exports, which is already the clamped one.
*
* The other two DeepSeek keys are here for reasons worth keeping written down:
* - `DSH_HOME` points the launcher at a profile tree whose plugin code runs at BOOT,
* before any approval row could apply.
* - `DEEPSEEK_BASE_URL` would redirect the server's OWN forwarded `DEEPSEEK_API_KEY`
* to a host of the caller's choosing.
*
* Every other CLI's bypass is a command-line flag reachable only through its config
* object, which is why `privilegedParams` alone is the whole gate for them.
*/
privilegedEnvKeys: string[];
/** Version gates referenced by `capabilityGate` conditions. */
gates: Record<string, { minVersion: string; failClosed: boolean }>;
/** Cap on a single terminal frame, when this CLI needs a tighter one than the default. */
maxFrameBytes?: number;
}
// ---------------------------------------------------------------------------
// Location overlays (remote SSH / docker)
// ---------------------------------------------------------------------------
/** Docker credential seeding policy — which host dirs are copied or shared into a container. */
export interface CliCredStore {
rel: string;
shareDirs?: string[];
shareFiles?: string[];
seedFiles?: string[];
seedWhole?: boolean;
}
export interface CliOverlays {
/**
* The remote/docker DEFAULT pane command: just the CLI invocation (e.g. `claude
* --dangerously-skip-permissions`), independent of each location's own wrapping
* (remote: login-shell `-c`; docker: `exec`). Absent `command` = the bare
* `discovery.binaries[0]`. `disabled: true` = this location has no story for this CLI at
* all (docker for `shell`) — distinct from "no override", which still gets a default.
*/
remote?: { command?: string } | { disabled: true };
docker?: { command?: string } | { disabled: true };
/**
* ⚠️ DECLARED-FOR-LATER, unlike `remote`/`docker` above, which are live.
*
* The Docker credential-seeding path still reads its own `CRED_STORES` table in
* `docker-hosts.ts`, because this shape cannot yet express that table: it allows ONE store
* per CLI, and the live table needs two for gemini (`.gemini` for the CLI's own auth plus
* `.config/gcloud` for Vertex), while deepseek's entry here declares none at all even
* though `.dsh` is seeded. Wiring it therefore means making this an ARRAY and correcting
* those two entries — a change to credential seeding, which is both the highest-consequence
* thing in this file to get wrong and the least covered by tests, since every docker IO
* path is no-op'd under vitest. It belongs in its own change, measured against a real
* container.
*/
credStore?: CliCredStore;
}
// ---------------------------------------------------------------------------
// The entry
// ---------------------------------------------------------------------------
/**
* ⚠️ DECLARED-FOR-LATER: fields no code reads yet.
*
* `shortBadge`, `accent`, `overlays.credStore`, `capabilities.echo`, `capabilities.wheelForward`,
* `capabilities.keyboardAccessory` and `capabilities.maxFrameBytes` all describe FRONTEND
* behaviour, and the frontend is deliberately untouched by the change that introduced this
* registry — `app.js`, `terminal-ui.js`, `styles.css` and friends keep their own
* hand-authored per-CLI rules, and moving them is its own piece of work with its own way of
* being verified (a mobile/browser suite the CI gate cannot see).
*
* They are declared now because each entry should describe its CLI completely, and because
* transcribing them while the hand-written source is still on screen is when the values are
* actually known. But an unread field is a promise, not a fact: nothing enforces that
* `echo.policy` here matches `_updateLocalEchoState`'s fallthrough, or that `accent` matches
* the gradient CSS paints. Treat every value in this group as TRANSCRIBED, not authoritative,
* and re-measure against the frontend before wiring one up.
*
* The rest of the interface is live: something reads it, and `test/cli-registry-*.test.ts`
* pins what it does with it.
*/
export interface CliEntry {
id: CliId;
label: string;
/** Two-ish character tab badge, e.g. 'OC'. */
shortBadge: string;
/** Single hex colour. CSS derives every per-CLI gradient from it via --cli-accent. */
accent: string;
enabled: boolean;
/** Set by the loader from the shipped catalog; a user entry can never claim it. */
stock: boolean;
order: number;
/** 'shell' unlocks the raw-shell code paths; everything else is an agent CLI. */
kind: 'agent' | 'shell';
discovery: CliDiscovery;
launch: CliLaunch;
env: CliEnv;
capabilities: CliCapabilities;
overlays: CliOverlays;
}
/**
* The on-disk shape of ~/.codeman/clis.json — overrides and custom entries only, never the
* full catalog. Small and hand-readable by design.
*
* ⚠️ READ-ONLY in this build. Nothing here writes this file: there is no settings UI and no
* write API yet, so there is nothing to persist. That also means importing the registry
* (and therefore `schemas.ts`, which validates against it) performs no filesystem writes —
* an import side effect worth not having.
*/
export interface CliRegistryFile {
schemaVersion: number;
/**
* Stock ids already introduced to this install — the ratchet that lets one file both gain
* newly-shipped CLIs on upgrade AND remember that the user disabled one.
*
* Read and IGNORED here, and never written: the ratchet only earns its keep once a CLI
* can be disabled, which needs the write API. Declared now purely so a file written by a
* later version still loads cleanly under this one instead of failing `.strict()`.
*/
seededStockIds?: string[];
/** Keyed by id: a partial override of a stock entry, or a complete custom entry. */
clis: Record<string, unknown>;
}
+159 -128
View File
@@ -7,7 +7,8 @@
* @module config/dependency-registry
*/
import { PI_VERSION_REGEX } from '../utils/pi-cli-resolver.js';
import { enabledClis } from './cli-registry/registry.js';
import { compileVersionRegex } from './cli-registry/patterns.js';
export type ProbeEnvironment = 'linux' | 'darwin' | 'win32' | 'wsl';
@@ -56,134 +57,164 @@ export interface ToolDependency {
const ALL: ProbeEnvironment[] = ['linux', 'darwin', 'wsl', 'win32'];
export const DEPENDENCY_REGISTRY: ToolDependency[] = [
{
id: 'node',
label: 'Node.js',
category: 'core',
required: true,
minVersion: '22.0.0',
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['node'], versionArg: '--version' } }],
installHint: { linux: 'https://nodejs.org', darwin: 'brew install node', wsl: 'https://nodejs.org' },
},
{
id: 'claude',
label: 'Claude CLI',
category: 'core',
required: false,
usedBy: ['Claude Code sessions (default backend)'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['claude'], versionArg: '--version' } }],
installHint: { linux: 'https://docs.claude.com/claude-code', darwin: 'https://docs.claude.com/claude-code' },
},
{
id: 'tmux',
label: 'tmux',
category: 'core',
required: true,
resolvers: [{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['tmux'], versionArg: '-V' } }],
installHint: { linux: 'sudo apt install tmux', darwin: 'brew install tmux', wsl: 'sudo apt install tmux' },
},
{
id: 'opencode',
label: 'OpenCode CLI',
category: 'core',
required: false,
usedBy: ['OpenCode sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['opencode'], versionArg: '--version' } }],
},
{
id: 'codex',
label: 'Codex CLI',
category: 'core',
required: false,
usedBy: ['Codex sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['codex'], versionArg: '--version' } }],
},
{
id: 'gemini',
label: 'Gemini CLI',
category: 'core',
required: false,
usedBy: ['Gemini sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['gemini'], versionArg: '--version' } }],
},
{
id: 'antigravity',
label: 'Antigravity CLI',
category: 'core',
required: false,
usedBy: ['Antigravity sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['agy'], versionArg: '--version' } }],
},
{
id: 'pi',
label: 'Pi CLI',
category: 'core',
required: false,
usedBy: ['Pi sessions'],
// The only entry that requires a version match, for the same reason
// pi-cli-resolver.ts probes: `pi` is a short generic name (Raspberry Pi tooling,
// personal scripts), so a `which pi` hit alone is not the coding agent. Both sides
// share PI_VERSION_REGEX, so the doctor and the run mode cannot drift into telling
// the user opposite things about the same binary.
resolvers: [
{
match: ALL,
resolver: {
kind: 'path',
bins: ['pi'],
versionArg: '--version',
versionRegex: PI_VERSION_REGEX,
requireVersionMatch: true,
/**
* The doctor's ROW IDENTITY for a CLI, where it differs from the registry id.
*
* These are two separate contracts and they have never been the same thing: `codeman doctor`
* prints a tool table whose ids predate the registry, and `dsh` names the BINARY while the
* run mode is `deepseek`. Keeping the historical id here means the doctor's output does not
* shift under a refactor that was supposed to change nothing a user can see.
*
* `usedBy` is likewise preserved verbatim rather than generated, because the strings are
* shown to the user and claude's does not follow the pattern.
*/
const DOCTOR_ROW_OVERRIDES: Record<string, { id?: string; label?: string; usedBy: string[] }> = {
claude: { usedBy: ['Claude Code sessions (default backend)'] },
opencode: { usedBy: ['OpenCode sessions'] },
codex: { usedBy: ['Codex sessions'] },
gemini: { usedBy: ['Gemini sessions'] },
antigravity: { usedBy: ['Antigravity sessions'] },
pi: { usedBy: ['Pi sessions'] },
grok: { usedBy: ['Grok sessions'] },
// Both the id and the label are historical: `dsh` names the binary, and the doctor has
// always spelled this row out in full rather than as `${label} CLI`.
deepseek: { id: 'dsh', label: 'DeepSeek Harness CLI', usedBy: ['DeepSeek sessions'] },
};
/**
* Build one `codeman doctor` row per enabled CLI, straight from its registry entry.
*
* This replaces eight hand-written rows that had to be kept in step with the run modes by
* hand — and were not: an earlier draft of this refactor silently dropped the Grok and
* DeepSeek rows, so `codeman doctor` stopped reporting two shipped CLIs at all. Deriving
* the list makes that class of omission impossible.
*
* ⚠️ The version regex is compiled through `compileVersionRegex()`, NOT `new RegExp()`. It
* is a config-supplied pattern, so it goes through the same length cap and
* nested-quantifier rejection the argv engine applies; the doctor runs it over command
* output exactly like the resolver does, and skipping the guard here would leave one
* unguarded path into a user-supplied regex.
*
* ⚠️ Sharing the entry's regex with the resolver is what stops the doctor and the run mode
* telling the user opposite things about the same binary — the Dependencies panel reporting
* "Pi CLI ✓" on a box where Run Pi stays hidden.
*/
function cliDependencyEntries(): ToolDependency[] {
const rows: ToolDependency[] = [];
for (const cli of enabledClis()) {
// `shell` has no binary of its own (the login shell is resolved at spawn time), so
// there is nothing for the doctor to probe.
const bin = cli.discovery.binaries[0];
if (!bin) continue;
const override = DOCTOR_ROW_OVERRIDES[cli.id as string];
const version = cli.discovery.version;
const versionRegex = version?.regex ? (compileVersionRegex(version.regex) ?? undefined) : undefined;
rows.push({
id: override?.id ?? (cli.id as string),
label: override?.label ?? `${cli.label} CLI`,
category: 'core',
required: false,
usedBy: override?.usedBy ?? [`${cli.label} sessions`],
resolvers: [
{
match: ALL,
resolver: {
kind: 'path',
bins: [bin],
versionArg: version?.arg ?? '--version',
versionRegex,
// Only meaningful for a CLI whose binary name is short, generic or squatted
// (pi, grok, dsh): a bare `which` hit there is not evidence of the right
// program, so a version mismatch means MISSING rather than unknown-version.
requireVersionMatch: version?.requireVersionMatch,
},
},
},
],
},
{
id: 'libreoffice',
label: 'LibreOffice',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['linux', 'darwin', 'wsl'],
resolver: { kind: 'path', bins: ['libreoffice', 'soffice'], versionArg: '--version' },
},
],
installHint: { linux: 'sudo apt install libreoffice', darwin: 'brew install --cask libreoffice' },
},
{
id: 'pdftoppm',
label: 'pdftoppm',
category: 'office',
required: false,
usedBy: ['document preview', 'PDF/Office first-page thumbnails'],
// poppler's pdftoppm prints its version to stderr; presence is what matters here.
resolvers: [
{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['pdftoppm'], versionArg: '-v' } },
],
installHint: {
linux: 'sudo apt install poppler-utils',
darwin: 'brew install poppler',
wsl: 'sudo apt install poppler-utils',
],
installHint: cli.discovery.install.command,
});
}
return rows;
}
/**
* The tools `codeman doctor` probes, resolved AT CALL TIME.
*
* ⚠️ A FUNCTION, not a module-level const, and for the same reason `sessionModeSchema()` and
* `allowedEnvPrefixes()` are functions: `cliDependencyEntries()` reads the CLI registry, and
* a const would have frozen the doctor's rows at first import while every schema resolved
* per parse. A CLI enabled while the server was running — or a `reloadCliRegistry()` — then
* moved the run menu and the validation but never the doctor, which would keep reporting the
* catalog as it stood when something first imported this module. Building the array per call
* costs a handful of object literals on a command that shells out to probe binaries anyway.
*/
export function dependencyRegistry(): ToolDependency[] {
return [
{
id: 'node',
label: 'Node.js',
category: 'core',
required: true,
minVersion: '22.0.0',
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['node'], versionArg: '--version' } }],
installHint: { linux: 'https://nodejs.org', darwin: 'brew install node', wsl: 'https://nodejs.org' },
},
},
{
id: 'msoffice',
label: 'MS Office',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['wsl', 'win32'],
resolver: {
kind: 'windows-side',
appDirs: ['Microsoft Office/root/Office16'],
exes: ['WINWORD.EXE', 'POWERPNT.EXE', 'EXCEL.EXE'],
{
id: 'tmux',
label: 'tmux',
category: 'core',
required: true,
resolvers: [{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['tmux'], versionArg: '-V' } }],
installHint: { linux: 'sudo apt install tmux', darwin: 'brew install tmux', wsl: 'sudo apt install tmux' },
},
...cliDependencyEntries(),
{
id: 'libreoffice',
label: 'LibreOffice',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['linux', 'darwin', 'wsl'],
resolver: { kind: 'path', bins: ['libreoffice', 'soffice'], versionArg: '--version' },
},
],
installHint: { linux: 'sudo apt install libreoffice', darwin: 'brew install --cask libreoffice' },
},
{
id: 'pdftoppm',
label: 'pdftoppm',
category: 'office',
required: false,
usedBy: ['document preview', 'PDF/Office first-page thumbnails'],
// poppler's pdftoppm prints its version to stderr; presence is what matters here.
resolvers: [
{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['pdftoppm'], versionArg: '-v' } },
],
installHint: {
linux: 'sudo apt install poppler-utils',
darwin: 'brew install poppler',
wsl: 'sudo apt install poppler-utils',
},
],
},
];
},
{
id: 'msoffice',
label: 'MS Office',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['wsl', 'win32'],
resolver: {
kind: 'windows-side',
appDirs: ['Microsoft Office/root/Office16'],
exes: ['WINWORD.EXE', 'POWERPNT.EXE', 'EXCEL.EXE'],
},
},
],
},
];
}
+15
View File
@@ -40,6 +40,21 @@ const INSTANCE_SUFFIX = CODEMAN_INSTANCE ? `-${CODEMAN_INSTANCE}` : '';
/** Default tmux socket for this instance. `CODEMAN_TMUX_SOCKET` still overrides. */
export const DEFAULT_TMUX_SOCKET = `codeman${INSTANCE_SUFFIX}`;
/** Characters tmux accepts in a `-L` socket name. */
export const SAFE_TMUX_SOCKET_PATTERN = /^[a-zA-Z0-9_.-]+$/;
/**
* This instance's tmux socket: the `CODEMAN_TMUX_SOCKET` override when it is a
* safe name, else the instance default. Every process that runs `tmux -L` has
* to resolve it through here (the server via TmuxManager, the TUI for its
* degraded-mode listing), or a beta instance ends up driving prod's sessions.
*/
export function resolveTmuxSocketName(): string {
const raw = process.env.CODEMAN_TMUX_SOCKET;
if (raw !== undefined && SAFE_TMUX_SOCKET_PATTERN.test(raw)) return raw;
return DEFAULT_TMUX_SOCKET;
}
let _ensured = false;
/**
+2 -1
View File
@@ -8,7 +8,8 @@
* src/web/public/constants.js and deliberately stays at 50k — 100k xterm lines per tab
* is a mobile-memory hazard — so DEFAULT_TERMINAL_SCROLLBACK_LINES stays 50,000 to match.
* The terminalScrollbackLines/terminalBufferMaxBytes/terminalBufferTrimBytes settings keys
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is live.
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is wired.
* tmux <3.7 applies it to new panes; tmux 3.7+ can also resize live panes.
* All values remain env- and settings-overridable and bounds-clamped via
* resolveTerminalHistoryConfig().
*/
+56 -7
View File
@@ -10,6 +10,8 @@
import { v4 as uuidv4 } from 'uuid';
import { readFile } from 'node:fs/promises';
import { statSync, realpathSync } from 'node:fs';
import { getCli } from '../config/cli-registry/registry.js';
import { resolveCliLaunchError } from '../utils/cli-launcher.js';
import { Session } from '../session.js';
import { applyWorkspaceHooks } from '../hooks-config.js';
import { SseEvent } from '../web/sse-events.js';
@@ -48,7 +50,9 @@ const delay = (ms: number): Promise<void> => new Promise((r) => setTimeout(r, ms
* answer "yes" to, which then loads and EXECUTES repo-local `.pi/extensions` TypeScript,
* so `approveProjectTrust: false` (`--no-approve`) is materialized. Omitting `--approve`
* is NOT a clamp.
* Codex and antigravity need nothing here: their absent config already spawns safe.
* Codex, antigravity, grok and deepseek need nothing here: their absent config already spawns safe
* (grok's bare spawn is its own ask-mode default and deepseek's omits DSH_PERMISSION_MODE
* entirely, leaving the harness on workspace-write, which asks; both switches are only ever sent).
* Granted/admin/single-user get undefined for both, i.e. upstream defaults untouched.
*/
export function clampCronExternalCliConfigs(
@@ -56,9 +60,26 @@ export function clampCronExternalCliConfigs(
ownerGranted: boolean
): { geminiConfig: GeminiConfig | undefined; piConfig: PiConfig | undefined } {
if (ownerGranted) return { geminiConfig: undefined, piConfig: undefined };
// A cron job carries no per-CLI config at all, so ONLY the materialize-when-absent params
// can apply here — an only-if-sent clamp has nothing to clamp. Reading them off the
// registry rather than naming gemini and pi means a future CLI whose bare spawn is unsafe
// is covered the moment its entry says so, instead of silently missing this path.
const entry = getCli(mode);
const aliases = entry?.launch.legacyConfigAliases ?? {};
const materialized: Record<string, unknown> = {};
for (const { param, clampTo, materializeWhenAbsent } of entry?.capabilities.privilegedParams ?? []) {
// Same registry-param → legacy-wire-field hop the HTTP clamp makes. Neither gemini's
// `approvalMode` nor pi's `approveProjectTrust` is aliased today, so this changes nothing
// now — but the two are DIFFERENT namespaces, and writing the raw param here would make
// this path stop clamping the moment one of them gained an alias, silently.
if (materializeWhenAbsent) materialized[aliases[param] ?? param] = clampTo;
}
const has = Object.keys(materialized).length > 0;
const field = entry?.launch.legacyConfigField;
return {
geminiConfig: mode === 'gemini' ? { approvalMode: 'auto_edit' } : undefined,
piConfig: mode === 'pi' ? { approveProjectTrust: false } : undefined,
geminiConfig: has && field === 'geminiConfig' ? (materialized as GeminiConfig) : undefined,
piConfig: has && field === 'piConfig' ? (materialized as PiConfig) : undefined,
};
}
@@ -385,7 +406,7 @@ export class CronService {
// Section 6.3: re-resolve the owner's grant at FIRE time (it may have been revoked
// since create). Gates shell/launchCommand AND clamps the external-CLI bypass below.
const ownerGranted = await canUsernameRunPrivilegedCommands(job.owner);
if ((job.agentType === 'shell' || job.launchCommand) && !ownerGranted) {
if ((getCli(job.agentType)?.capabilities.privilegedCommandGate || job.launchCommand) && !ownerGranted) {
return this.failRun(job, run, 'Owner lacks the can-bypass-permissions grant for shell/launchCommand jobs');
}
@@ -393,11 +414,37 @@ export class CronService {
let session: Session;
try {
const mode = job.agentType;
// A LAUNCHER CLI's binary is not its agent, so "installed" is not "runnable": without
// this, a job on a box carrying only dsh's stock web/headless profiles spawns a bare
// `dsh` that boots a profile unable to drive a pane, and the prompt is typed into a
// logging server or a dead pane instead of failing the run with an actionable message.
//
// ⚠️ Scoped to `discovery.launcherProfile`, which is byte-identical to the
// `mode === 'deepseek'` check this replaces (dsh is the only launcher today) and
// generalises to the next one. Deliberately NOT every CLI: cron has never pre-flighted
// a merely-missing binary, and doing so replaces tmux-manager's own not-found throw
// ("Session launch failed") with a different message for claude and shell. An earlier
// draft of this line was unscoped and did exactly that — three cron tests caught it.
if (getCli(mode)?.discovery.launcherProfile !== undefined) {
const cronLaunchError = await resolveCliLaunchError(mode);
if (cronLaunchError) return this.failRun(job, run, cronLaunchError);
}
const globalNice = await this.deps.getGlobalNiceConfig();
const modelConfig = await this.deps.getModelConfig();
const claudeModeConfig = await this.deps.getClaudeModeConfig();
const effectiveClaudeMode = await resolveClaudeModeForUsername(claudeModeConfig.claudeMode, job.owner);
const model = mode !== 'shell' ? modelConfig?.defaultModel || undefined : undefined;
// Cron carries no per-CLI config object, so the only model it can supply is the global
// default — and only to a CLI that takes a model at all.
//
// ⚠️ `!== 'none'` is the faithful reading of the `mode !== 'shell' && mode !== 'deepseek'`
// ladder this replaces: those two are exactly the entries declaring `model.source: 'none'`
// (shell has no model; deepseek's is a profile composition entry, not a session flag).
// NOT `=== 'claude-settings-file'`, which is the HTTP route's question — there, every
// external CLI reads its model from its own config object earlier in the chain, so only
// claude reaches the global default. Cron has no such config, so the same expression
// means something different here.
const model =
getCli(mode)?.capabilities.model.source !== 'none' ? modelConfig?.defaultModel || undefined : undefined;
// Section 6.3: materialize the safe default for a non-granted owner (see
// clampCronExternalCliConfigs — cron sends no per-CLI config, so the CLI's own
// spawn default is what would otherwise apply).
@@ -425,7 +472,7 @@ export class CronService {
piConfig,
owner: job.owner,
});
this.deps.addSession(session);
await this.deps.addSession(session);
this.store.incrementSessionsCreated();
this.deps.persistSessionState(session);
await this.deps.setupSessionListeners(session);
@@ -546,7 +593,9 @@ export class CronService {
private sendPromptWhenReady(sessionId: string, prompt: string, job: CronJob, run: CronJobRun): void {
setImmediate(() => {
const poll = async (): Promise<void> => {
if (job.agentType !== 'shell') {
// A shell pane is ready the moment it exists; an agent CLI has a TUI to paint
// first. That is the `kind` the registry already records, not a fact about shell.
if (getCli(job.agentType)?.kind !== 'shell') {
for (let attempt = 0; attempt < CRON_READY_MAX_ATTEMPTS; attempt++) {
await delay(500);
const s = this.deps.sessions.get(sessionId);
+271
View File
@@ -0,0 +1,271 @@
/**
* @fileoverview The DeepSeek Harness -> Codeman status bridge.
*
* ## Why this exists
*
* Every external CLI mode before this one (opencode, codex, gemini, antigravity,
* pi, grok) is READINESS-GUESSED: Codeman watches the PTY go quiet and infers a
* turn ended. Claude is the exception, because Claude Code fires real hooks. The
* DeepSeek Harness TUI gives us a third option, and a much better one than
* guessing: the community terminal front door already reports its own lifecycle
* to an owning supervisor, and it does so through a fully GENERIC, env-var-gated
* contract it inherited from Herdr (herdr.dev).
*
* When all three of `HERDR_ENV=1`, `HERDR_BIN_PATH` and `HERDR_PANE_ID` are set,
* the TUI shells out on every state change:
*
* "$HERDR_BIN_PATH" pane report-agent "$HERDR_PANE_ID" \
* --source custom:dsh-tui --agent dsh-tui \
* --state idle|working|blocked [--message <text>] --seq <n>
*
* and treats exit code 0 as "delivered" (retrying with backoff otherwise). So
* Codeman points `HERDR_BIN_PATH` at the script below and gets DEFINITIVE
* idle/working/blocked signals for dsh sessions: real respawn triggers, real
* `wait`/`wait-output` stop+blocked signals, and real Approvals Inbox items,
* on par with Claude's hooks rather than with output stabilization.
*
* This is an interface implementation, not an impersonation: we implement the
* one verb (`pane report-agent`) that the contract defines, and nothing on the
* machine ever executes a real `herdr` binary — `HERDR_BIN_PATH` is our own
* script, in our own data dir. `HERDR_ENV=1` is the flag the TUI checks to know
* a supervisor is present; a supervisor IS present, it is Codeman.
*
* ## Why it is generated rather than committed
*
* The shim must be an executable file at a stable absolute path in every
* install shape: a git clone (where `scripts/` exists), an `npm i -g aicodeman`
* (where `files` ships only `dist` plus two named scripts), and any
* `CODEMAN_INSTANCE`. Writing it into the data dir at session-create time makes
* one code path cover all of them, single-sources the content here in TS, and
* follows the precedent of `self-update-runner.sh`. It is rewritten whenever the
* embedded version marker changes, so an upgraded Codeman refreshes a stale shim
* without the user knowing it exists.
*
* @module deepseek-status-shim
*/
import { chmodSync, mkdirSync, readFileSync, renameSync, rmSync, writeFileSync } from 'node:fs';
import { dirname } from 'node:path';
import { dataPath } from './config/instance.js';
/**
* Bumped whenever SHIM_SOURCE changes. The marker is embedded in the generated
* file, so `ensureDeepSeekStatusShim()` can tell a current shim from one written
* by an older Codeman and rewrite only when needed (rather than rewriting on
* every session create, or — worse — leaving a stale one in place forever).
*/
const SHIM_VERSION = 3;
const SHIM_MARKER = `codeman-dsh-status-shim v${SHIM_VERSION}`;
/**
* Mapping from the harness's three lifecycle states to Codeman hook events.
*
* - `blocked` -> `permission_prompt`: the TUI reports blocked when a tool
* approval or an `ask_user_question` questionnaire is on screen, which is
* exactly the red "needs you" alert and an answerable Approvals Inbox item.
* - `idle` -> `stop`: the definitive end-of-turn signal, the one respawn and the
* wait endpoints care about.
* - `working` -> `agent_working`: a turn STARTED. Codeman infers "working" from
* PTY output well enough on its own, but the event is what RESOLVES a pending
* approval when the user answers a dialog in the terminal instead of in the
* inbox. Without it a dsh session's red alert would survive until the next
* `stop`, which is the exact stuck-alert bug the claude path already had to
* fix once (and the pane-capture staleness sweep that fixed it there is
* Claude-dialog-shaped, so it cannot help here).
*/
export const DEEPSEEK_STATE_TO_HOOK_EVENT: Readonly<Record<string, string>> = Object.freeze({
idle: 'stop',
blocked: 'permission_prompt',
working: 'agent_working',
});
/**
* The generated script.
*
* Constraints it must satisfy, each learned from an existing Codeman hook bug:
* - **TLS**: `CODEMAN_API_URL` is loopback HTTPS with a self-signed cert on
* `--https`/tailscale installs, so certificate verification is disabled for
* the request. Without this the whole bridge dies silently, exactly as the
* claude hook curls did before they grew `-k`.
* - **Secret**: the hook-secret file is read AT EXECUTION TIME, never baked in,
* so rotation needs no respawn and the value never lands on a command line.
* - **Exit codes**: 0 means delivered. Anything else makes the TUI retry with
* backoff, so transport failures self-heal, but an unknown verb or an
* unmapped state exits 0 to avoid a pointless retry storm over something that
* will never succeed.
* - **Timeout**: bounded below the caller's own 2s budget, so we lose the race
* deliberately rather than being killed mid-flight.
*/
const SHIM_SOURCE = `#!/usr/bin/env node
// ${SHIM_MARKER}
// GENERATED BY CODEMAN — do not edit. Rewritten from src/deepseek-status-shim.ts
// whenever its version marker changes.
//
// Implements the one verb the DeepSeek Harness TUI's supervisor contract uses:
// pane report-agent <paneId> --state <idle|working|blocked> [--message <t>] ...
// and forwards it to this Codeman instance as a hook event.
import { readFileSync } from 'node:fs'
import http from 'node:http'
import https from 'node:https'
const STATE_TO_EVENT = ${JSON.stringify(DEEPSEEK_STATE_TO_HOOK_EVENT)}
const TIMEOUT_MS = 1500
const argv = process.argv.slice(2)
const flag = (name) => {
const i = argv.indexOf(name)
return i >= 0 && i + 1 < argv.length ? argv[i + 1] : undefined
}
// Unknown verb: succeed silently. Retrying could never make it succeed, and a
// non-zero exit here would make the caller retry four times per state change.
if (argv[0] !== 'pane' || argv[1] !== 'report-agent') process.exit(0)
const event = STATE_TO_EVENT[String(flag('--state') ?? '')]
if (!event) process.exit(0)
// The pane id we hand the TUI IS the Codeman session id, but prefer the ambient
// env: it is set by the same code that set HERDR_PANE_ID, so a TUI that mangles,
// truncates or re-uses the pane argument still reports against the right session.
// NOT a security boundary, and do not read it as one: the agent runs IN this pane
// and can invoke the shim with CODEMAN_SESSION_ID unset and any argv it likes.
// That buys it nothing it did not already have, since the hook-secret file is
// readable from the same pane and any process there can POST /api/hook-event
// directly. Attribution here is about accidents, not adversaries.
const sessionId = process.env.CODEMAN_SESSION_ID || argv[2]
const apiUrl = process.env.CODEMAN_API_URL
if (!sessionId || !apiUrl) process.exit(1)
let secret = ''
try {
secret = readFileSync(process.env.CODEMAN_HOOK_SECRET_FILE || '', 'utf-8').trim()
} catch {
// Missing file: the loopback bypass still applies when no tunnel is running.
}
// The contract's ordering token: the TUI retries failed deliveries with
// backoff, so a stale report can land AFTER a newer one. Forwarded so the
// server can drop out-of-order arrivals instead of, say, resolving an
// approval with a retried 'working' while the harness sits blocked.
const seq = Number(flag('--seq'))
const body = JSON.stringify({
event,
sessionId,
data: {
source: 'dsh-status-shim',
agent: flag('--agent') || 'dsh',
...(Number.isFinite(seq) ? { seq } : {}),
...(flag('--message') ? { message: flag('--message') } : {}),
},
})
let url
try {
url = new URL('/api/hook-event', apiUrl)
} catch {
process.exit(1)
}
const transport = url.protocol === 'https:' ? https : http
const req = transport.request(
{
protocol: url.protocol,
hostname: url.hostname,
port: url.port,
path: url.pathname,
method: 'POST',
timeout: TIMEOUT_MS,
headers: {
'Content-Type': 'application/json',
'Content-Length': Buffer.byteLength(body),
'X-Codeman-Hook-Secret': secret,
},
// Loopback HTTPS with a self-signed cert (--https / tailscale installs).
rejectUnauthorized: false,
},
(res) => {
res.resume()
const status = res.statusCode ?? 0
// 2xx: delivered. 4xx: PERMANENT — a 401 (missing/rotated secret) or 429
// can never be fixed by retrying, and each retry feeds the auth-failure
// rate-limit bucket, so a single misconfigured dsh session could 429 the
// hook endpoint for the whole instance (killing every claude session's
// real hooks). Exit 0 so the TUI does not retry; only transport errors
// and 5xx stay retryable.
process.exit(status >= 200 && status < 500 ? 0 : 1)
}
)
req.on('timeout', () => {
req.destroy()
process.exit(1)
})
req.on('error', () => process.exit(1))
req.end(body)
`;
/** Absolute path of the generated shim for this instance. */
export function deepSeekStatusShimPath(): string {
return dataPath('dsh-status-shim.mjs');
}
let ensuredThisProcess = false;
/**
* Write the shim if it is missing or stale, and return its path.
*
* Idempotent and cheap: after the first call in a process it does nothing, and
* even the first call only rewrites when the on-disk marker differs. Never
* throws — a data dir that cannot be written is a degraded status bridge, not a
* failed session start, so callers fall back to output-stabilization readiness
* by receiving null.
*/
export function ensureDeepSeekStatusShim(): string | null {
const path = deepSeekStatusShimPath();
if (ensuredThisProcess) return path;
try {
let current = '';
try {
current = readFileSync(path, 'utf-8');
} catch {
// Missing — fall through to the write.
}
if (!current.includes(SHIM_MARKER)) {
mkdirSync(dirname(path), { recursive: true });
// Temp + rename, not a plain write: the TUI can be executing this exact
// path at the moment an upgraded Codeman refreshes it (every state change
// runs it, and session create is when the rewrite happens), and a reader
// that catches a half-written file gets a syntax error, exits non-zero,
// and is retried four times per state change for a file that will never
// parse. rename(2) is atomic within the directory, so a concurrent exec
// sees either the old shim or the new one, never a truncated one.
// Same reasoning as the state-store writes; pid-suffixed so two instances
// sharing a data dir cannot collide on the temp name.
const tempPath = `${path}.${process.pid}.tmp`;
try {
writeFileSync(tempPath, SHIM_SOURCE, { mode: 0o700 });
// The mode argument only applies when writeFileSync CREATES the file, so
// a leftover temp from a crashed run would keep its old permissions.
chmodSync(tempPath, 0o700);
renameSync(tempPath, path);
} catch (err) {
rmSync(tempPath, { force: true });
throw err;
}
}
// Re-assert the mode even when the content matched: a shim that lost its
// executable bit (a restored backup, a copied data dir) would make every
// report fail, and the TUI would retry four times per state change forever.
chmodSync(path, 0o700);
ensuredThisProcess = true;
return path;
} catch (err) {
console.warn(`[DeepSeek] Could not install the status shim at ${path}: ${(err as Error).message}`);
return null;
}
}
/** Test seam: forget the per-process memo so a fresh temp HOME is re-provisioned. */
export function resetDeepSeekStatusShimForTest(): void {
ensuredThisProcess = false;
}
+696
View File
@@ -0,0 +1,696 @@
/**
* @fileoverview Reading a DeepSeek Harness (`dsh`) session transcript off disk.
*
* ## Why this exists
*
* `GET /api/sessions/:id/last-response` is how an agent (and the Response
* Viewer) reads what a worker actually said. For Claude it comes from
* `~/.claude/projects/**`, for Codex from `~/.codex/sessions/**`, and for every
* other external CLI it comes from segmenting the terminal buffer, because
* those CLIs write nothing a reader could open.
*
* dsh is not in that last group: it writes a complete, structured JSONL
* transcript per session. Falling back to the pane for it was measurably wrong
* rather than merely coarse — dsh-TUI paints a full-screen splash, so the pane
* segmenter answered a `last-response` call for a fresh dsh session with the
* ASCII-art logo:
*
* {"text":"✦dsh-TUI v0.8.8█▀▀▀▄█▀▀▀▀█▀▀▀▀█▀▀▀▄█▀▀▀▀…","hasContext":true}
*
* which an agent polling for a worker's answer reads as an answer. This module
* is the real source: it locates the session's transcript, decodes it, and
* returns the last turn's text.
*
* ## The three things that make dsh transcripts unlike codex rollouts
*
* **1. One zstd FRAME per append, not one zstd stream.** The file is
* `session.jsonl.zstd`, and dsh appends by compressing each batch of lines into
* its own frame and writing it at the end. `zstd -dc` handles that (frames
* concatenate by definition), but Node's `zlib.zstdDecompress()` and
* `createZstdDecompress()` both stop at the first frame end: measured on a real
* 56-line transcript, Node returned 158 bytes / 1 line where the CLI returned
* 43,747 bytes / 56 lines. That is a silent truncation to the session header —
* every call would have reported "no answer yet" forever. `decodeZstdFrames()`
* below walks the frame headers itself and decompresses each frame, and
* `test/deepseek-transcript.test.ts` pins it against multi-frame fixtures.
*
* **2. The user's prompts are mixed with injected context.** Every turn also
* writes a `user/message` whose source is a plugin (the runtime-context
* snapshot: sandbox policy, approval policy, cwd). Those are `source.kind ===
* 'plugin'`; a real prompt is `source.kind === 'user'`. Rendering the plugin
* ones would show the agent its own boilerplate back as the user's words.
*
* **3. A failed turn is not an empty turn.** `turn/end` carries
* `reason.kind === 'error'` with the provider's message. Returning `""` there
* makes an agent poll `last-response` fifteen times and conclude the worker
* never answered, when the truth ("the provider rejected the request") was on
* disk the whole time. A turn that ends in an error and produced no text
* answers with that error, prefixed so it can never be mistaken for the model's
* own words.
*
* Verified against `dsh 0.1.1-rc.2` + `@deepseek-harness-tui/dsh-tui 0.8.8`.
*/
import { promises as fs } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
import * as zlib from 'node:zlib';
/**
* One rendered block, in the shape the Response Viewer already speaks (see
* `web/response-viewer-transcript.ts`). Imported as a type only — this module
* must stay usable from the session layer without dragging web/ into it.
*/
export interface DeepSeekTranscriptBlock {
kind: 'prompt' | 'response' | 'status' | 'tool';
label: 'Prompt' | 'Response' | 'Status' | 'Tool';
role: 'user' | 'assistant';
text: string;
}
export interface DeepSeekTranscriptResult {
/** Last turn's answer (or its error, prefixed). Empty before the first turn. */
text: string;
/** ISO timestamp of the event `text` came from, or '' when unknown. */
timestamp: string;
/** Rendered blocks, oldest first. Only built when the caller asks for them. */
blocks: DeepSeekTranscriptBlock[];
/** dsh's own session id, from the header line. */
sessionId?: string;
/** Workspace the harness recorded for the session. */
cwd?: string;
}
/**
* zstd decompression is a RUNTIME capability here, not an import.
*
* Node grew `zlib` zstd support in 22.15 (and `@types/node` still does not
* declare it), while Codeman's floor is Node 22.0. So it is resolved through a
* narrow cast and checked before use: on an older 22.x a dsh session keeps the
* pane-segmenter behaviour it had before this module existed instead of
* throwing on every `last-response` call.
*/
type ZstdDecompressSync = (buf: Buffer) => Buffer;
const zstdDecompressSync: ZstdDecompressSync | undefined = (
zlib as unknown as { zstdDecompressSync?: ZstdDecompressSync }
).zstdDecompressSync;
/** Whether this Node can decode the compressed transcripts dsh writes. */
export function zstdSupported(): boolean {
return typeof zstdDecompressSync === 'function';
}
/** zstd frame magic (RFC 8878 §3.1.1). */
const ZSTD_MAGIC = 0xfd2fb528;
/** Skippable-frame magic range: 0x184D2A50..0x184D2A5F. */
const ZSTD_SKIPPABLE_LO = 0x184d2a50;
const ZSTD_SKIPPABLE_HI = 0x184d2a5f;
const DID_FIELD_SIZE = [0, 1, 2, 4];
const FCS_FIELD_SIZE = [0, 2, 4, 8];
/**
* Byte ranges of the zstd frames in `buf`, in order.
*
* Walks frame headers and block headers only — no decompression — so the cost
* is proportional to the number of blocks, not to the content. Stops (rather
* than throws) at the first thing it cannot parse, so a transcript still being
* appended to mid-write yields every whole frame before the torn tail instead
* of failing the whole read.
*
* ⚠️ Splitting on the magic bytes instead would be wrong: the 4-byte sequence
* can occur inside compressed data, and a false split corrupts everything after
* it. The block walk is what makes the boundaries exact.
*/
export function zstdFrameRanges(buf: Buffer): Array<[number, number]> {
const ranges: Array<[number, number]> = [];
let offset = 0;
while (offset + 4 <= buf.length) {
const magic = buf.readUInt32LE(offset);
if (magic >= ZSTD_SKIPPABLE_LO && magic <= ZSTD_SKIPPABLE_HI) {
if (offset + 8 > buf.length) break;
const end = offset + 8 + buf.readUInt32LE(offset + 4);
if (end > buf.length || end <= offset) break;
offset = end;
continue;
}
if (magic !== ZSTD_MAGIC) break;
let p = offset + 4;
if (p >= buf.length) break;
const descriptor = buf[p] as number;
p += 1;
const fcsFlag = descriptor >> 6;
const singleSegment = (descriptor >> 5) & 1;
const hasChecksum = (descriptor >> 2) & 1;
const dictIdFlag = descriptor & 3;
if (!singleSegment) p += 1; // window descriptor
p += DID_FIELD_SIZE[dictIdFlag] as number;
// FCS is absent for flag 0 UNLESS Single_Segment is set, where it is 1 byte.
p += fcsFlag === 0 ? (singleSegment ? 1 : 0) : (FCS_FIELD_SIZE[fcsFlag] as number);
if (p > buf.length) break;
let lastBlock = false;
let torn = false;
while (!lastBlock) {
if (p + 3 > buf.length) {
torn = true;
break;
}
const header = (buf[p] as number) | ((buf[p + 1] as number) << 8) | ((buf[p + 2] as number) << 16);
p += 3;
lastBlock = (header & 1) === 1;
const blockType = (header >> 1) & 3;
const blockSize = header >> 3;
if (blockType === 3) {
torn = true; // reserved: refuse rather than guess
break;
}
p += blockType === 1 ? 1 : blockSize; // RLE stores a single byte
if (p > buf.length) {
torn = true;
break;
}
}
if (torn) break;
if (hasChecksum) p += 4;
if (p > buf.length) break;
ranges.push([offset, p]);
offset = p;
}
return ranges;
}
/**
* Decode a possibly multi-frame zstd buffer. A buffer that does not start with
* a zstd magic is passed through unchanged, which is what lets the same reader
* open a plain `session.jsonl` (dsh writes one when compression is off).
*
* A frame that fails to decompress truncates the decode THERE rather than
* failing it: everything decoded before it is kept, so a half-written tail
* frame does not cost the caller the whole conversation. (Not "skipped" — a
* frame after a corrupt one is never reached, which is the safe reading: dsh
* appends, so a bad frame means everything after it is suspect too.)
*/
export function decodeZstdFrames(buf: Buffer): string {
if (buf.length < 4) return buf.toString('utf8');
const magic = buf.readUInt32LE(0);
if (magic !== ZSTD_MAGIC && (magic < ZSTD_SKIPPABLE_LO || magic > ZSTD_SKIPPABLE_HI)) {
return buf.toString('utf8');
}
if (!zstdDecompressSync) return '';
const parts: Buffer[] = [];
for (const [start, end] of zstdFrameRanges(buf)) {
try {
parts.push(zstdDecompressSync(buf.subarray(start, end)));
} catch {
// Torn or corrupt frame: keep what decoded before it.
break;
}
}
return Buffer.concat(parts).toString('utf8');
}
interface DshEvent {
type?: string;
seq?: number | null;
time?: number;
data?: Record<string, unknown>;
}
function asRecord(value: unknown): Record<string, unknown> | undefined {
return value && typeof value === 'object' && !Array.isArray(value) ? (value as Record<string, unknown>) : undefined;
}
function asArray(value: unknown): unknown[] {
return Array.isArray(value) ? value : [];
}
/**
* Strip a leaked reasoning prefix.
*
* Some providers stream reasoning into the same text block and close it with
* `</think>` without ever opening it (measured on a local deepseek-v4-flash
* route: `"I'll read the file first.</think>\n\nThe add function is…"`). The
* closing tag is the only reliable boundary, so everything up to the LAST one
* goes. A block with no tag is returned untouched.
*/
function stripReasoningPrefix(text: string): string {
const close = text.lastIndexOf('</think>');
return close === -1 ? text : text.slice(close + '</think>'.length);
}
/** `stripReasoning` is for ASSISTANT content only: a user prompt containing a
* literal `</think>` (someone pasting a transcript, say) must render whole. */
function textOfContent(content: unknown, stripReasoning = true): string {
const parts: string[] = [];
for (const entry of asArray(content)) {
const block = asRecord(entry);
if (!block) continue;
if (block.type === 'text' && typeof block.text === 'string') {
parts.push(stripReasoning ? stripReasoningPrefix(block.text) : block.text);
}
}
return parts.join('').trim();
}
function toolCallsOfContent(content: unknown): string[] {
const calls: string[] = [];
for (const entry of asArray(content)) {
const block = asRecord(entry);
if (!block || block.type !== 'tool-call') continue;
const name = typeof block.name === 'string' ? block.name : 'tool';
const args = typeof block.arguments === 'string' ? block.arguments : JSON.stringify(block.arguments ?? {});
calls.push(`${name}(${args})`);
}
return calls;
}
/** Flatten a `tool/result` message down to its text payload. */
function textOfToolResult(message: unknown): string {
const parts: string[] = [];
for (const entry of asArray(asRecord(message)?.content)) {
const block = asRecord(entry);
if (!block) continue;
if (block.type === 'text' && typeof block.text === 'string') parts.push(block.text);
if (block.type === 'tool-result') {
for (const inner of asArray(block.content)) {
const innerBlock = asRecord(inner);
if (innerBlock?.type === 'text' && typeof innerBlock.text === 'string') parts.push(innerBlock.text);
}
}
}
return parts.join('\n').trim();
}
function isoTime(time: unknown): string {
return typeof time === 'number' && Number.isFinite(time) ? new Date(time).toISOString() : '';
}
interface TurnAccumulator {
/** Finalized `assistant/message` text, in step order. */
finalized: Map<number, string>;
/** Steps that produced a finalized message AT ALL. ⚠️ Not the same as a
* non-empty entry in `finalized`: a step whose whole reply was reasoning
* strips to `''`, and without this the deltas — which are NOT stripped at
* write time — would be resurrected in its place, putting the model's raw
* `</think>` monologue in front of the caller (measured). */
finalizedSteps: Set<number>;
/** Streamed deltas per step, used only where no finalized message landed. */
streamed: Map<number, string>;
/** Step order as encountered, so a reply reads in the order it was produced. */
steps: number[];
timestamp: string;
/** Pre-rendered "Turn error: …" / "Turn ended: …" line, when the turn did not
* end with `completed`. */
ending?: string;
}
function ensureStep(turn: TurnAccumulator, step: number): void {
if (!turn.steps.includes(step)) turn.steps.push(step);
}
function turnText(turn: TurnAccumulator): string {
const parts: string[] = [];
for (const step of turn.steps) {
// Deltas are only consulted for a step the model never finalized — a step
// that has both would otherwise render its text twice.
const text = turn.finalizedSteps.has(step)
? (turn.finalized.get(step) ?? '')
: stripReasoningPrefix(turn.streamed.get(step) ?? '');
if (text.trim()) parts.push(text.trim());
}
return parts.join('\n\n').trim();
}
/**
* Parse a decoded dsh transcript.
*
* `text` is the LAST TURN's answer, not the last assistant message anywhere in
* the file: a turn that errored after an earlier turn answered must not hand
* back the earlier turn's text as though it were this turn's reply.
*/
export function parseDeepSeekTranscript(raw: string, options: { blocks?: boolean } = {}): DeepSeekTranscriptResult {
const wantBlocks = options.blocks === true;
const blocks: DeepSeekTranscriptBlock[] = [];
const turns = new Map<number, TurnAccumulator>();
const turnOrder: number[] = [];
let sessionId: string | undefined;
let cwd: string | undefined;
const getTurn = (n: number): TurnAccumulator => {
let turn = turns.get(n);
if (!turn) {
turn = { finalized: new Map(), finalizedSteps: new Set(), streamed: new Map(), steps: [], timestamp: '' };
turns.set(n, turn);
turnOrder.push(n);
}
return turn;
};
for (const line of raw.split('\n')) {
if (!line.trim()) continue;
let event: DshEvent;
try {
event = JSON.parse(line) as DshEvent;
} catch {
continue; // a torn tail line, or a frame we could not decode
}
const data = asRecord(event.data) ?? {};
const turnNo = typeof data.turn === 'number' ? data.turn : 0;
const stepNo = typeof data.step === 'number' ? data.step : 0;
switch (event.type) {
case 'session': {
const header = event as unknown as Record<string, unknown>;
if (typeof header.id === 'string') sessionId = header.id;
if (typeof header.cwd === 'string') cwd = header.cwd;
break;
}
case 'user/message': {
// ⚠️ Only a real prompt. The plugin-sourced twin is the runtime-context
// snapshot dsh injects every turn (sandbox policy, approvals, cwd).
if (asRecord(data.source)?.kind !== 'user') break;
if (!wantBlocks) break;
const text = textOfContent(data.content, false);
if (text) blocks.push({ kind: 'prompt', label: 'Prompt', role: 'user', text });
break;
}
case 'assistant/message': {
const message = asRecord(data.message);
const turn = getTurn(turnNo);
ensureStep(turn, stepNo);
const text = textOfContent(message?.content);
if (message) turn.finalizedSteps.add(stepNo);
if (text) {
turn.finalized.set(stepNo, text);
turn.timestamp = isoTime(event.time) || turn.timestamp;
if (wantBlocks) blocks.push({ kind: 'response', label: 'Response', role: 'assistant', text });
}
if (wantBlocks) {
for (const call of toolCallsOfContent(message?.content)) {
blocks.push({ kind: 'tool', label: 'Tool', role: 'assistant', text: call });
}
}
break;
}
case 'assistant/chunk': {
const chunk = asRecord(data.chunk);
if (chunk?.type !== 'text-delta' || typeof chunk.text !== 'string') break;
const turn = getTurn(turnNo);
ensureStep(turn, stepNo);
turn.streamed.set(stepNo, (turn.streamed.get(stepNo) ?? '') + chunk.text);
break;
}
case 'text-chunks': {
// The batched form of the same deltas (dsh coalesces once a stream gets
// going). ⚠️ These carry `seq: null`, so file order is the only order.
const turn = getTurn(turnNo);
ensureStep(turn, stepNo);
const texts = asArray(data.texts)
.filter((t): t is string => typeof t === 'string')
.join('');
if (texts) turn.streamed.set(stepNo, (turn.streamed.get(stepNo) ?? '') + texts);
break;
}
case 'tool/result': {
if (!wantBlocks) break;
const text = textOfToolResult(data.message);
if (text) blocks.push({ kind: 'tool', label: 'Tool', role: 'assistant', text });
break;
}
case 'turn/end': {
const turn = getTurn(turnNo);
const reason = asRecord(data.reason);
if (reason && reason.kind !== 'completed') {
// Two different things wear this field: a provider failure
// (`kind:'error'` with a message) and an ordinary early stop
// (`kind:'max-tokens'`, measured live). Calling the second one an
// error would misreport a truncated but real answer.
const error = asRecord(reason.error);
const message = typeof error?.message === 'string' ? error.message : undefined;
const kind = typeof reason.kind === 'string' ? reason.kind : 'unknown';
turn.ending = message ? `Turn error: ${message}` : `Turn ended: ${kind}`;
if (wantBlocks) {
blocks.push({ kind: 'status', label: 'Status', role: 'assistant', text: turn.ending });
}
}
turn.timestamp = isoTime(event.time) || turn.timestamp;
break;
}
default:
break;
}
}
const lastTurn = turnOrder.length > 0 ? turns.get(turnOrder[turnOrder.length - 1] as number) : undefined;
let text = lastTurn ? turnText(lastTurn) : '';
// A turn that failed and said nothing answers with its failure, labelled so
// it can never read as the model's own words. Without this an agent polls
// `last-response` fifteen times and concludes the worker never answered.
if (!text && lastTurn?.ending) text = lastTurn.ending;
return { text, timestamp: lastTurn?.timestamp ?? '', blocks, sessionId, cwd };
}
/**
* `$DSH_HOME` for one session: a per-session override wins (`DSH_HOME` is an
* allowlisted `envOverrides` prefix, and pointing a worker at its own profile
* tree is a documented thing to do), then the server's own environment, then
* `~/.dsh`. Reading the wrong tree does not fail loudly — it silently finds no
* transcript — so this must resolve exactly the way the spawn did.
*/
/* ⚠️ The override is EPHEMERAL: `envOverrides` is applied at spawn and exported
* through `tmux setenv`, but is deliberately not persisted to state.json (it can
* carry provider keys). A session that overrode `DSH_HOME` and then outlived a
* server restart therefore resolves to the default tree and finds no transcript
* — it reads as "nothing said yet" rather than as another session's answer,
* because every candidate is matched on its recorded `cwd`. */
export function resolveDeepSeekHome(session: { deepSeekHomeOverride?: string }): string {
const override = session.deepSeekHomeOverride;
if (override && override.trim()) return override.trim();
const fromEnv = process.env.DSH_HOME;
if (fromEnv && fromEnv.trim()) return fromEnv.trim();
return join(homedir(), '.dsh');
}
/**
* How far apart a session's start and its transcript's `createdAt` may be and
* still be the same session. dsh writes the header within ~2 s of pane start
* (measured); 60 s absorbs a cold profile boot without ever reaching a sibling
* started minutes later.
*/
const PAIRING_WINDOW_MS = 60_000;
/** Transcript file names dsh has used, newest convention first. */
const TRANSCRIPT_FILES = ['session.jsonl.zstd', 'session.jsonl'];
/**
* Locate the transcript for a session.
*
* dsh buckets sessions by a mangled cwd (`--home-you-code-app--`) and then by
* its own session id, and the id form has changed between versions (`<uuid>`
* and `session-<uuid>` both exist on disk here). ⚠️ So the mangling is NOT
* reproduced: every candidate's own header line carries `cwd`, which is
* authoritative, and matching on it is immune to the next naming change.
*
* Pairing a Codeman session with ITS transcript then has one hard rule and one
* ladder. The rule: a transcript created BEFORE this session started belongs to
* an earlier conversation in the same directory and is never eligible. Measured
* cost of getting that wrong — a freshly spawned worker answered its very first
* `last-response` with the PREVIOUS session's reply, which is worse than saying
* nothing, because an agent cannot tell a stale answer from a fresh one.
*
* The ladder, once the older ones are out:
*
* 1. a transcript whose header `createdAt` sits within `PAIRING_WINDOW_MS` of
* this session's start — that is this pane's own boot, and it stays right
* even when a sibling session is running in the same case directory;
* 2. otherwise the newest transcript created after this session started;
* 3. otherwise nothing.
*
* ⚠️ The boot transcript wins for as long as it exists on disk — deliberately,
* and even over a LATER transcript in the same workspace. Step 2 cannot tell a
* `/new` from a sibling session that started later in the same directory, so
* preferring newest-eligible would hand a worker its busier sibling's reply
* (the exact bug the hard rule above was measured against, one seat over).
* The cost of that choice: after an interactive `/new` in a dsh tab, this
* reader keeps serving the pre-`/new` conversation (the same session's own
* earlier turns — stale, never foreign); step 2 is reached only when no
* boot-window transcript exists. Worker fleets never `/new`, so they only
* ever see step 1.
*/
export async function findDeepSeekTranscript(options: {
dshHome: string;
workingDir: string;
startedAt?: number;
}): Promise<string | null> {
const sessionsDir = join(options.dshHome, 'sessions');
let buckets: string[];
try {
buckets = (await fs.readdir(sessionsDir, { withFileTypes: true }))
.filter((entry) => entry.isDirectory())
.map((entry) => entry.name);
} catch {
return null;
}
const candidates: Array<{ path: string; mtimeMs: number }> = [];
for (const bucket of buckets) {
const bucketPath = join(sessionsDir, bucket);
let sessions: string[];
try {
sessions = (await fs.readdir(bucketPath, { withFileTypes: true }))
.filter((entry) => entry.isDirectory())
.map((entry) => entry.name);
} catch {
continue;
}
for (const sessionDir of sessions) {
for (const file of TRANSCRIPT_FILES) {
const path = join(bucketPath, sessionDir, file);
const stat = await fs.stat(path).catch(() => null);
if (!stat || !stat.isFile() || stat.size === 0) continue;
candidates.push({ path, mtimeMs: stat.mtimeMs });
break;
}
}
}
if (candidates.length === 0) return null;
candidates.sort((a, b) => b.mtimeMs - a.mtimeMs);
const startedAt = options.startedAt ?? 0;
// Slack in both directions: the harness writes its header a beat after the
// pane starts, and mtimes on a shared clock are not worth trusting to the ms.
const floor = startedAt > 0 ? startedAt - PAIRING_WINDOW_MS : 0;
let laterMatch: string | null = null;
for (const candidate of candidates) {
const header = await readTranscriptHeader(candidate.path);
if (!header || header.cwd !== options.workingDir) continue;
// No usable header timestamp: fall back to the file's own mtime, which is
// still enough to keep a pre-session transcript out.
const createdAt = header.createdAt ?? candidate.mtimeMs;
if (createdAt < floor) continue;
if (startedAt > 0 && Math.abs(createdAt - startedAt) <= PAIRING_WINDOW_MS) return candidate.path;
if (!laterMatch) laterMatch = candidate.path;
}
return laterMatch;
}
/**
* Read only the first frame of a transcript, which is where the header line
* lives. Bounded: a candidate scan must never decompress every conversation on
* the box to answer one `last-response` call.
*/
async function readTranscriptHeader(path: string): Promise<{ cwd?: string; id?: string; createdAt?: number } | null> {
let handle;
try {
handle = await fs.open(path, 'r');
} catch {
return null;
}
try {
const head = Buffer.alloc(65536);
const { bytesRead } = await handle.read(head, 0, head.length, 0);
if (bytesRead === 0) return null;
const text = decodeZstdFrames(head.subarray(0, bytesRead));
const firstLine = text.split('\n').find((line) => line.trim());
if (!firstLine) return null;
const parsed = JSON.parse(firstLine) as { type?: string; cwd?: string; id?: string; createdAt?: number };
if (parsed.type !== 'session') return null;
return {
cwd: parsed.cwd,
id: parsed.id,
createdAt: typeof parsed.createdAt === 'number' ? parsed.createdAt : undefined,
};
} catch {
return null;
} finally {
await handle.close().catch(() => {});
}
}
/** Hard ceiling on a transcript read. A long agent run is a few hundred KB; a
* file past this is pathological and is not worth a synchronous decode. */
const MAX_TRANSCRIPT_BYTES = 64 * 1024 * 1024;
/**
* Memo of the last few decoded transcripts, keyed on (path, mtime, size,
* blocks). The skill's `last_text` polls once per second, and each poll used
* to zstdDecompressSync + reparse the WHOLE file on the event loop even when
* nothing had been appended — a multi-MB transcript made that a repeated
* ~100ms-class stall on the single-threaded server. A poll that finds the
* file unchanged now costs one stat. Insertion-order eviction; tiny, because
* an entry only earns its keep while a session is being actively polled.
*/
const parseMemo = new Map<string, DeepSeekTranscriptResult>();
const PARSE_MEMO_MAX = 16;
/** Test seam: a fixture that rewrites one path in place inside a single mtime
* tick would otherwise read its predecessor back out of the memo. */
export function resetDeepSeekTranscriptMemoForTest(): void {
parseMemo.clear();
}
/**
* Read one dsh session's last answer.
*
* ⚠️ The two empty outcomes are deliberately different, because the caller must
* treat them differently:
*
* - `null` means **this reader cannot run here** (a Node without zstd), and is
* the signal to fall back to the pane segmenter.
* - an empty `text` means **read fine, nothing said yet** — no transcript for
* this workspace, or a turn still in flight.
*
* Collapsing the two would put the ASCII-art splash back in front of an agent
* that is polling for a worker's first answer.
*/
export async function readDeepSeekLastResponse(
session: { workingDir: string; createdAt?: Date | number; deepSeekHomeOverride?: string },
options: { blocks?: boolean } = {}
): Promise<DeepSeekTranscriptResult | null> {
const createdAt = session.createdAt instanceof Date ? session.createdAt.getTime() : session.createdAt;
// dsh compresses by default, so a Node without zstd can read nothing here.
// That is the one case the pane is still the better answer.
if (!zstdSupported()) return null;
const empty: DeepSeekTranscriptResult = { text: '', timestamp: '', blocks: [] };
const path = await findDeepSeekTranscript({
dshHome: resolveDeepSeekHome(session),
workingDir: session.workingDir,
startedAt: typeof createdAt === 'number' ? createdAt : undefined,
});
if (!path) return empty;
const stat = await fs.stat(path).catch(() => null);
if (!stat || stat.size > MAX_TRANSCRIPT_BYTES) return empty;
const memoKey = `${path}|${stat.mtimeMs}|${stat.size}|${options.blocks ? 1 : 0}`;
const memoized = parseMemo.get(memoKey);
if (memoized) return memoized;
let buf: Buffer;
try {
buf = await fs.readFile(path);
} catch {
return empty;
}
const result = parseDeepSeekTranscript(decodeZstdFrames(buf), options);
if (parseMemo.size >= PARSE_MEMO_MAX) {
const oldest = parseMemo.keys().next().value;
if (oldest !== undefined) parseMemo.delete(oldest);
}
parseMemo.set(memoKey, result);
return result;
}
+283
View File
@@ -0,0 +1,283 @@
/**
* @fileoverview Supervises the one background `dsh web` process behind the Run
* menu's "DeepSeek web UI..." entry.
*
* The shortcut originally started the server inside an ordinary SHELL SESSION,
* on the reasoning that Codeman already knows how to supervise those: it was
* visible, scrollable, killable, and died with its tab, and nothing new had to
* own a long-lived HTTP server. That reasoning was sound and the result was
* still wrong in use — clicking "open the DeepSeek web UI" spawned a terminal
* tab the user never asked for, next to the web tab they did, and the terminal
* was noise every time after the first.
*
* So the server moves here instead: one child process, no session, no tab.
* What that buys back has to be paid for explicitly, which is what this module
* is:
*
* - **Exactly one.** A second click reuses the running server rather than
* racing it for a port. The old shell-session flow could not do this at all,
* because two clicks were simply two sessions.
* - **Restarted when the authority changes.** `--trusted-host` fences dsh's
* `/api` against the browser authority, and a Codeman reachable at both
* loopback and a tailnet name has two. Whoever asks last wins, because the
* asker is by definition the origin about to load the page.
* - **Killed on shutdown.** A detached child that outlived Codeman would hold
* its port against the next start, which is exactly the EADDRINUSE this
* feature already got wrong once.
* - **Failures reported, not swallowed.** The shell tab used to be where the
* stack trace landed. With no tab, the spawn's own output is captured and
* handed back to the caller instead.
*/
import { spawn, type ChildProcess } from 'node:child_process';
import { createServer } from 'node:net';
import { join } from 'node:path';
import { getErrorMessage } from './types.js';
/**
* Where the port search starts, and how far it walks.
*
* 3080 is `dsh web`'s own default, so it is the friendly first choice — and
* emphatically not a fixed port. DeepSeek's web UI is a thing users run
* themselves, which makes the default precisely the port most likely to be
* taken already; hardcoding it made this feature die with EADDRINUSE against
* the user's own server.
*/
const PORT_BASE = 3080;
const PORT_SPAN = 40;
/** How long a freshly spawned server gets to answer before we call it failed. */
const READY_TIMEOUT_MS = 30_000;
const READY_POLL_MS = 400;
/** Grace between SIGTERM and SIGKILL when stopping the tree. */
const KILL_GRACE_MS = 3_000;
/** Bound on captured child output, so a chatty boot cannot grow without limit. */
const OUTPUT_CAP = 16_384;
export interface DeepSeekWebStatus {
running: boolean;
port: number | null;
url: string | null;
/** Browser authority this server was started to trust (`--trusted-host`). */
authority: string | null;
}
interface RunningServer {
child: ChildProcess;
port: number;
authority: string;
output: () => string;
}
let current: RunningServer | null = null;
/**
* True when nothing holds `port` on loopback.
*
* Binding is the only honest test: a connect probe cannot tell "free" from
* "listening but not answering yet", and this runs moments before `dsh web`
* binds the same port. It is inherently racy, which is why the caller still
* waits for the server to actually answer before reporting success.
*/
async function isLoopbackPortFree(port: number): Promise<boolean> {
return new Promise((resolve) => {
const probe = createServer();
probe.once('error', () => resolve(false));
probe.once('listening', () => probe.close(() => resolve(true)));
probe.listen(port, '127.0.0.1');
});
}
async function findFreePort(): Promise<number | null> {
for (let port = PORT_BASE; port < PORT_BASE + PORT_SPAN; port++) {
if (await isLoopbackPortFree(port)) return port;
}
return null;
}
/** Does the server answer HTTP yet? Any status counts: dsh may 4xx a bare GET. */
async function answersHttp(port: number): Promise<boolean> {
try {
await fetch(`http://127.0.0.1:${port}/`, { signal: AbortSignal.timeout(2_000) });
return true;
} catch {
return false;
}
}
/**
* Signal the whole process group.
*
* `dsh web` boots a plugin tree and fans out, so signalling only the direct
* child leaves survivors holding the port. Same negative-pid escalation as
* `runGit()` in git-clone.ts and the profile installer.
*/
function killTree(child: ChildProcess, signal: NodeJS.Signals): void {
try {
if (child.pid) process.kill(-child.pid, signal);
} catch {
try {
child.kill(signal);
} catch {
/* already gone */
}
}
}
export function getDeepSeekWebStatus(): DeepSeekWebStatus {
if (!current) return { running: false, port: null, url: null, authority: null };
return {
running: true,
port: current.port,
url: `http://127.0.0.1:${current.port}`,
authority: current.authority,
};
}
/** Stop the background server, if one is running. Safe to call when none is. */
export async function stopDeepSeekWeb(): Promise<void> {
const running = current;
current = null;
if (!running) return;
await new Promise<void>((resolve) => {
let done = false;
const finish = () => {
if (done) return;
done = true;
clearTimeout(hard);
resolve();
};
running.child.once('exit', finish);
killTree(running.child, 'SIGTERM');
const hard = setTimeout(() => {
killTree(running.child, 'SIGKILL');
finish();
}, KILL_GRACE_MS);
});
}
type StartResult = { ok: true; port: number; url: string; reused: boolean } | { ok: false; error: string };
/**
* Serializes concurrent starts. Two POSTs racing (two devices, or a double
* click while the first boots) used to both see `current === null`, pick the
* SAME free port, and spawn twice: the loser died on EADDRINUSE while its exit
* handler nulled the singleton out from under the winner, leaving a live
* `dsh web` nothing tracked or killed — the exact orphan this module exists to
* prevent. The second caller now simply waits and reuses the first's server.
*/
let startLock: Promise<unknown> = Promise.resolve();
/**
* Start (or reuse) the background `dsh web` for `authority`.
*
* @param dshDir directory holding the resolved `dsh` binary.
* @param authority browser authority to pass as `--trusted-host`.
*/
export function startDeepSeekWeb(dshDir: string, authority: string): Promise<StartResult> {
const run = startLock.then(
() => startDeepSeekWebLocked(dshDir, authority),
() => startDeepSeekWebLocked(dshDir, authority)
);
startLock = run.then(
() => undefined,
() => undefined
);
return run;
}
async function startDeepSeekWebLocked(dshDir: string, authority: string): Promise<StartResult> {
// Reuse only when the running server is BOTH healthy and fenced for the
// authority now asking. A server trusting the other origin renders a page
// whose every API call 403s, which looks like a broken dashboard rather than
// a misconfigured one.
if (current) {
if (current.authority === authority && (await answersHttp(current.port))) {
return { ok: true, port: current.port, url: `http://127.0.0.1:${current.port}`, reused: true };
}
await stopDeepSeekWeb();
}
const port = await findFreePort();
if (port === null) {
return { ok: false, error: `No free port for the DeepSeek web UI in ${PORT_BASE}-${PORT_BASE + PORT_SPAN - 1}` };
}
let child: ChildProcess;
try {
child = spawn(
join(dshDir, 'dsh'),
['web', '--no-open', '--host', '127.0.0.1', '--port', String(port), '--trusted-host', authority],
{
stdio: ['ignore', 'pipe', 'pipe'],
// Own process group so the whole plugin tree can be signalled at once.
detached: true,
env: process.env,
}
);
} catch (err) {
return { ok: false, error: `Failed to start dsh web: ${getErrorMessage(err)}` };
}
// The pipes must be drained whether or not anyone reads them: a full pipe
// blocks the child. Storage is capped; draining is not.
let output = '';
const capture = (chunk: Buffer) => {
if (output.length < OUTPUT_CAP) output += chunk.toString('utf-8');
};
child.stdout?.on('data', capture);
child.stderr?.on('data', capture);
let exited = false;
child.once('exit', () => {
exited = true;
// Only clear if this is still the current server: a restart may have
// already replaced it, and clearing then would drop the live one.
if (current?.child === child) current = null;
});
child.once('error', () => {
exited = true;
if (current?.child === child) current = null;
});
const running: RunningServer = { child, port, authority, output: () => output };
current = running;
const deadline = Date.now() + READY_TIMEOUT_MS;
while (Date.now() < deadline) {
if (exited) {
// Guarded like the exit/error handlers: a concurrent stop (DELETE route,
// shutdown) may already have cleared or replaced the singleton, and an
// unconditional null here would drop a server this call does not own.
if (current === running) current = null;
const tail = output.trim().slice(-800);
return { ok: false, error: tail ? `dsh web exited during startup: ${tail}` : 'dsh web exited during startup' };
}
if (await answersHttp(port)) {
return { ok: true, port, url: `http://127.0.0.1:${port}`, reused: false };
}
await new Promise((r) => setTimeout(r, READY_POLL_MS));
}
// Timeout: kill OUR child. Only route through stopDeepSeekWeb() while the
// singleton is still ours — signalling `current` unconditionally here could
// SIGTERM a healthy server a concurrent actor now owns.
if (current === running) {
await stopDeepSeekWeb();
} else {
killTree(running.child, 'SIGKILL');
}
const tail = output.trim().slice(-800);
return {
ok: false,
error: tail
? `dsh web did not answer on port ${port} within ${READY_TIMEOUT_MS / 1000}s: ${tail}`
: `dsh web did not answer on port ${port} within ${READY_TIMEOUT_MS / 1000}s`,
};
}
/** Test seam: forget any tracked child without signalling it. */
export function resetDeepSeekWebForTest(): void {
current = null;
}
+89 -20
View File
@@ -23,7 +23,8 @@
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join, dirname } from 'node:path';
import { dirname, isAbsolute, join, relative, resolve } from 'node:path';
import { getCli } from './config/cli-registry/registry.js';
import { fileURLToPath } from 'node:url';
import { homedir } from 'node:os';
import { createHash } from 'node:crypto';
@@ -32,7 +33,6 @@ import { promisify } from 'node:util';
import { dataPath } from './config/instance.js';
import type {
DockerCase,
DockerCommandMode,
DockerEngine,
DockerHost,
DockerNetworkMode,
@@ -134,19 +134,23 @@ export function dockerContainerName(caseName: string): string {
return `${CONTAINER_NAME_PREFIX}${caseName}`;
}
/** Default pane command per CLI mode (mirror of defaultRemoteCommandForMode). */
/**
* Default in-container pane command per CLI mode (mirror of defaultRemoteCommandForMode).
*
* ⚠️ Read from the registry (`overlays.docker`), not from a hardcoded
* `Record<DockerCommandMode, string>`. That table duplicated the registry exactly with
* nothing keeping the two in step. `shell` is the one arm still written here, because it is
* the entry that declares `docker: { disabled: true }` — a container has no per-user login
* shell to resolve, so it gets a plain `bash -l` rather than a CLI invocation.
*/
export function defaultDockerCommandForMode(mode: SessionMode): string {
const commands: Record<DockerCommandMode, string> = {
shell: 'exec bash -l',
// Mirror the LOCAL claude default so the in-container agent runs non-interactively.
claude: 'exec claude --dangerously-skip-permissions',
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
antigravity: 'exec agy',
pi: 'exec pi',
};
return commands[mode as DockerCommandMode] || commands.shell;
const entry = getCli(mode);
const overlay = entry?.overlays.docker;
if (!entry || (overlay && 'disabled' in overlay)) return 'exec bash -l';
// Mirrors the LOCAL default for each CLI; claude's carries
// `--dangerously-skip-permissions` so the in-container agent runs non-interactively.
const cli = overlay?.command ?? entry.discovery.binaries[0];
return cli ? `exec ${cli}` : 'exec bash -l';
}
/** `container:/workdir` display string (mirror of remoteDisplayPath's `user@host:path`). */
@@ -274,6 +278,24 @@ export interface DockerMount {
readonly?: boolean;
}
/**
* Resolve a bind source into the Docker daemon's filesystem namespace.
*
* A bare-host Codeman process and its Docker daemon see the same HOME, so the
* source is returned unchanged. In Docker-outside-of-Docker deployments,
* `runtimeHome` is the path inside Codeman while `daemonHome` is the host path
* bind-mounted there. Sources beneath HOME must therefore be translated before
* they are sent through the Docker socket.
*/
export function resolveDockerDaemonMountSource(source: string, runtimeHome: string, daemonHome?: string): string {
const configuredDaemonHome = daemonHome?.trim();
if (!configuredDaemonHome) return source;
const relativeSource = relative(resolve(runtimeHome), resolve(source));
if (relativeSource.startsWith('..') || isAbsolute(relativeSource)) return source;
return resolve(configuredDaemonHome, relativeSource);
}
/**
* Resolved, IO-free context for buildDockerCreateArgs. The caller (tmux-manager)
* resolves the environment-dependent bits (host uid, existing cred mounts, the
@@ -297,6 +319,8 @@ export interface DockerCreateContext {
addHostGateway: boolean;
/** Engine host-gateway alias (host.docker.internal / host.containers.internal). */
gatewayAlias: string;
/** Omit --memory-swap when the host kernel cannot enforce swap limits. */
disableSwapLimit?: boolean;
}
/**
@@ -314,12 +338,15 @@ function mountSpec(m: DockerMount): string {
return `type=bind,src=${m.src},dst=${m.dst}${m.readonly ? ',readonly' : ''}`;
}
function resourceFlags(resources?: DockerResourceLimits): string[] {
function resourceFlags(resources?: DockerResourceLimits, disableSwapLimit = false): string[] {
if (!resources) return [];
const flags: string[] = [];
if (resources.memory) {
// memory-swap == memory disables swap, making --memory a REAL OOM cap.
flags.push('--memory', resources.memory, '--memory-swap', resources.memory);
flags.push('--memory', resources.memory);
// memory-swap == memory disables swap where the daemon supports swap
// accounting. Some kernels, including the deployed Unraid host, do not;
// requesting it there emits a warning and Docker ignores the value.
if (!disableSwapLimit) flags.push('--memory-swap', resources.memory);
}
if (resources.cpus) flags.push('--cpus', resources.cpus);
if (resources.pidsLimit) flags.push('--pids-limit', String(resources.pidsLimit));
@@ -384,7 +411,7 @@ export function buildDockerCreateArgs(ctx: DockerCreateContext): string[] {
if (addHostGateway) args.push('--add-host', `${gatewayAlias}:host-gateway`);
args.push(
...resourceFlags(docker.resources),
...resourceFlags(docker.resources, ctx.disableSwapLimit),
// GPU passthrough (needs the NVIDIA container toolkit on the host). No storage
// cap is set, so the container's writable layer + volumes grow elastically as
// data flows in (bounded only by host disk).
@@ -614,8 +641,45 @@ const CRED_STORES: CredStorePolicy[] = [
rel: '.pi/agent',
seedFiles: ['auth.json', 'settings.json', 'trust.json', 'models.json', 'models-store.json'],
},
// Grok (xAI) keeps auth + config in `~/.grok`, but that dir ALSO holds
// `sessions/`, `memory/`, `downloads/` (the ~100MB binary itself) and `bin/`,
// so seedWhole would copy all of it into every container start. Seed only what
// grok needs to authenticate and behave consistently. Same trade-off as pi:
// in-container grok sessions are invisible host-side, so `grok -c` inside a
// Docker case only sees that container's own history.
{
rel: '.grok',
seedFiles: ['auth.json', 'config.toml', 'pager.toml'],
},
// DeepSeek Harness keeps credentials in `~/.dsh/.env` (0600) and composition in
// `settings.yaml` / `cordis.patch.yml`. `profiles/` is deliberately NOT seeded:
// it is a pnpm workspace holding a full node_modules tree per profile, which is
// both enormous and host-arch-specific. An in-container dsh therefore needs its
// profile installed IN the image (see docker/agent.Dockerfile), and the seeded
// files only supply auth and model composition. Same host-invisibility trade-off
// as pi and grok: `~/.dsh/sessions` inside a container is that container's own.
{
rel: '.dsh',
seedFiles: ['.env', 'settings.yaml', 'cordis.patch.yml'],
},
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
// OMP keeps its config in `~/.omp/agent` (config.yml/mcp.json/models.yml/
// settings.yml — small, no bigger than grok's config.toml/pager.toml), but
// that dir ALSO holds agent.db/history.db/models.db (SQLite caches) and
// terminal-sessions/blobs/cache (large, regenerable), so seed only the
// config files. UNLIKE pi/grok, `sessions/` is SHARED (RW), not
// host-invisible: Codeman reads `~/.omp/agent/sessions/**/*.jsonl`
// HOST-SIDE for history recovery and --resume pinning
// (omp-transcript.ts, omp-session-resolver.ts) — the same reason codex's
// `sessions/` is shared rather than seeded. Without this, an in-container
// OMP conversation would be invisible to Codeman's own history-scan/resume
// logic, silently breaking the kill-survival feature for Docker cases.
{
rel: '.omp/agent',
shareDirs: ['sessions'],
seedFiles: ['config.yml', 'mcp.json', 'models.yml', 'settings.yml'],
},
];
/**
@@ -1054,8 +1118,13 @@ export async function probeDockerCliVersion(
mode: SessionMode
): Promise<string | undefined> {
if (IS_TEST_MODE) return undefined;
const bin = mode === 'shell' ? null : mode;
if (!bin) return undefined;
// ⚠️ The MODE NAME IS NOT ALWAYS THE BINARY NAME — `antigravity` runs `agy`. This used
// to pass the mode straight through as the command, which would have probed a binary that
// does not exist. Only claude reaches this today (it is the one CLI with a version gate),
// so nothing was actually broken, but the registry is what makes it correct for the next
// CLI that needs a version.
const bin = getCli(mode)?.discovery.binaries[0];
if (!bin) return undefined; // `shell` has no binary of its own
const argv = dockerEngineArgv(docker);
try {
const { stdout } = await execFileAsync(
+11 -2
View File
@@ -479,14 +479,23 @@ export function gitNonInteractiveEnv(base: NodeJS.ProcessEnv = process.env): Nod
// ─── Pure: output handling ───────────────────────────────────────────────────
/**
* Redact any `scheme://user:secret@host` credential pair embedded in text — a
* remote URL stored with an inline token, or git stderr echoing such a URL
* back. Shared by the clone error path (`sanitizeGitOutput`) and the
* repository-status card fields (`web/repo-status.ts`).
*/
export function redactGitCredentials(text: string): string {
return text.replace(/([a-zA-Z][a-zA-Z0-9+.-]*:\/\/)[^/@\s]*:[^/@\s]*@/g, '$1***:***@');
}
/**
* Make git's stderr safe to show in the browser: strip ANSI/control bytes,
* redact any `scheme://user:secret@host` that a credential helper echoed back,
* and keep only the tail (the last lines are the ones that say why it failed).
*/
export function sanitizeGitOutput(text: string, maxBytes = MAX_STDERR_BYTES): string {
const redacted = text
.replace(/([a-zA-Z][a-zA-Z0-9+.-]*:\/\/)[^/@\s]*:[^/@\s]*@/g, '$1***:***@')
const redacted = redactGitCredentials(text)
// eslint-disable-next-line no-control-regex -- deliberate: strip C0/C1 and DEL.
.replace(/[\u0000-\u0008\u000b\u000c\u000e-\u001f\u007f-\u009f]/g, '')
.trim();
+12 -3
View File
@@ -19,6 +19,9 @@ import type {
GeminiConfig,
AntigravityConfig,
PiConfig,
GrokConfig,
DeepSeekConfig,
OmpConfig,
SessionRemote,
SessionDocker,
} from './types.js';
@@ -78,13 +81,16 @@ export interface CreateSessionOptions {
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
piConfig?: PiConfig;
grokConfig?: GrokConfig;
deepSeekConfig?: DeepSeekConfig;
ompConfig?: OmpConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Ephemeral — not written to disk. */
envOverrides?: Record<string, string>;
/** Claude CLI effort level, injected as a `--settings` soft default (overridable via /effort in-session) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session. */
/** tmux history-limit (scrollback lines) allocated when this session is created. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
@@ -110,13 +116,16 @@ export interface RespawnPaneOptions {
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
piConfig?: PiConfig;
grokConfig?: GrokConfig;
deepSeekConfig?: DeepSeekConfig;
ompConfig?: OmpConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session after respawn. */
/** Original tmux history-limit retained for config parity; respawn cannot resize the existing pane. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
@@ -216,7 +225,7 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Update Ralph enabled state for a session */
updateRalphEnabled(sessionId: string, enabled: boolean): void;
/** Apply a tmux history-limit to all tracked sessions. */
/** Apply history-limit to live panes where tmux supports it, otherwise to future panes. */
setHistoryLimit(limit: number): Promise<void>;
// ========== Discovery ==========
+175
View File
@@ -0,0 +1,175 @@
/**
* @fileoverview Scan `~/.omp/agent/sessions/*&#47;*.jsonl` for Past Sessions rows,
* the omp analog of what `scanProjectDir()` (session-routes.ts) does for
* Claude's own `~/.claude/projects` transcripts.
*
* Without this, an omp conversation exists ONLY as a Codeman-level live/
* persisted session record — delete that (a "Kill Tmux" close, or any other
* cleanup) and the conversation vanishes from Past Sessions entirely, even
* though `omp` itself never forgot it. Claude conversations don't have that
* problem because Codeman already reads them back from Claude's own
* transcript files independent of its own session bookkeeping; this gives
* omp conversations the same treatment.
*
* Each omp session file's SECOND line is a `{"type":"session","id":...,
* "cwd":...}` header carrying the real (unmangled) working directory and the
* session's own id directly — no need to reverse-engineer the mangled
* directory name the way Claude Code's own scanner has to (see
* `decodeProjectKey()` in session-routes.ts and its "lossy" caveat). Prompt
* text comes from each `{"type":"message","message":{"role":"user",...}}`
* entry, giving a real first-message title instead of a bare case name.
*
* Unlike Claude's transcripts (which can run to tens of MB of tool-call
* output), an omp session file is the conversation only, so this reads each
* file whole rather than doing head/tail windows — bounded by a size cap so
* one unexpectedly huge file can't blow up memory.
*
* @module omp-transcript
*/
import { readFileSync, readdirSync, statSync } from 'node:fs';
import { homedir } from 'node:os';
import { join } from 'node:path';
function ompSessionsRoot(): string {
return join(homedir(), '.omp', 'agent', 'sessions');
}
/** Skip anything absurdly large rather than parsing it whole into memory. */
const MAX_OMP_SESSION_FILE_BYTES = 2 * 1024 * 1024;
/** Defensive cap on total files scanned across every directory, mirroring
* the Claude scanner's own instinct not to let one pathological tree stall
* a request — a real omp install has, at most, a few hundred of these. */
const MAX_OMP_SESSION_FILES = 2000;
export interface OmpHistorySession {
sessionId: string;
workingDir: string;
sizeBytes: number;
/** ISO timestamp, from the file's own mtime. */
lastModified: string;
firstPrompt?: string;
lastPrompt?: string;
}
function extractUserPromptText(message: unknown): string | undefined {
if (!message || typeof message !== 'object') return undefined;
const m = message as { role?: unknown; content?: unknown };
if (m.role !== 'user' || !Array.isArray(m.content)) return undefined;
const parts: string[] = [];
for (const block of m.content) {
if (block && typeof block === 'object' && (block as { type?: unknown }).type === 'text') {
const text = (block as { text?: unknown }).text;
if (typeof text === 'string') parts.push(text);
}
}
const joined = parts.join(' ').trim();
return joined || undefined;
}
/** Parse one omp session `.jsonl` file, or null when it's unreadable, empty, or has no session header. */
function parseOmpSessionFile(filePath: string): OmpHistorySession | null {
let stat: ReturnType<typeof statSync>;
try {
stat = statSync(filePath);
} catch {
return null;
}
if (stat.size === 0 || stat.size > MAX_OMP_SESSION_FILE_BYTES) return null;
let raw: string;
try {
raw = readFileSync(filePath, 'utf-8');
} catch {
return null;
}
let sessionId: string | undefined;
let workingDir: string | undefined;
let firstPrompt: string | undefined;
let lastPrompt: string | undefined;
for (const line of raw.split('\n')) {
if (!line) continue;
let entry: unknown;
try {
entry = JSON.parse(line);
} catch {
continue;
}
if (!entry || typeof entry !== 'object') continue;
const e = entry as Record<string, unknown>;
if (e.type === 'session' && typeof e.id === 'string' && typeof e.cwd === 'string' && e.cwd.startsWith('/')) {
// A corrupted or malformed session file could carry a relative or empty
// cwd; requiring an absolute path keeps a downstream resume attempt
// from being pointed at a nonsense working directory.
sessionId = e.id;
workingDir = e.cwd;
} else if (e.type === 'message') {
const prompt = extractUserPromptText(e.message);
if (prompt) {
if (!firstPrompt) firstPrompt = prompt;
lastPrompt = prompt;
}
}
}
if (!sessionId || !workingDir) return null;
return {
sessionId,
workingDir,
sizeBytes: stat.size,
lastModified: stat.mtime.toISOString(),
firstPrompt,
lastPrompt,
};
}
/**
* Scan every omp conversation on disk into Past-Sessions rows. Best-effort
* throughout: a missing `~/.omp` (never installed/used), an unreadable
* directory, or one corrupt file yields fewer rows rather than throwing —
* this feeds the same unified merge the Claude transcript scanner does, and
* one broken source must never blank the whole Past Sessions list.
*/
export function scanOmpSessionsHistory(): OmpHistorySession[] {
const root = ompSessionsRoot();
let dirEntries: string[];
try {
dirEntries = readdirSync(root);
} catch {
return [];
}
const out: OmpHistorySession[] = [];
for (const dirName of dirEntries) {
if (out.length >= MAX_OMP_SESSION_FILES) break;
const dirPath = join(root, dirName);
let dirStat: ReturnType<typeof statSync>;
try {
dirStat = statSync(dirPath);
} catch {
continue;
}
if (!dirStat.isDirectory()) continue;
let files: string[];
try {
files = readdirSync(dirPath);
} catch {
continue;
}
for (const file of files) {
if (out.length >= MAX_OMP_SESSION_FILES) break;
if (!file.endsWith('.jsonl')) continue;
try {
const parsed = parseOmpSessionFile(join(dirPath, file));
if (parsed) out.push(parsed);
} catch {
// One bad file must not sink the whole scan.
}
}
}
return out;
}
+8 -1
View File
@@ -281,7 +281,14 @@ export class RalphLoop extends EventEmitter {
// Guard: only reschedule if still running AND no timer is pending
// (prevents race where stop() clears timer between our check and setTimeout)
if (this._status === 'running' && this.loopTimer === null) {
this.loopTimer = setTimeout(() => this.runLoop(), this.pollIntervalMs);
// Null the handle when the timer fires, BEFORE re-entering runLoop —
// otherwise the `loopTimer === null` guard above stays false on the
// next pass and the loop stops rescheduling after 2 ticks.
// Mirrors the orchestrator-loop reschedule pattern.
this.loopTimer = setTimeout(() => {
this.loopTimer = null;
this.runLoop();
}, this.pollIntervalMs);
}
});
}
+142 -39
View File
@@ -4,9 +4,9 @@ import { join } from 'node:path';
import { homedir } from 'node:os';
import { exec } from 'node:child_process';
import { promisify } from 'node:util';
import { getCli } from './config/cli-registry/registry.js';
import type {
RemoteCase,
RemoteCommandMode,
RemoteHost,
RemoteSessionInfo,
RemoteSshOptions,
@@ -89,33 +89,54 @@ export function remoteLoginShellCommand(command: string): string {
return `exec ${REMOTE_LOGIN_SHELL} -i -l -c ${shellescape(command)}`;
}
/**
* The CLI text a location overlay should launch for `mode`, or null when this build has no
* entry for it. `overlays.<location>.command` when the entry names one, otherwise the bare
* binary — which is what every non-claude CLI wants, and why only claude declares a command.
*
* ⚠️ This returns the CLI INVOCATION only. Each location wraps it its own way (remote: a
* login-shell `-c`; docker: `exec`), which is exactly why the overlay stores the unwrapped
* form rather than a ready-made line.
*/
function overlayCliCommand(mode: SessionMode, location: 'remote' | 'docker'): string | null {
const entry = getCli(mode);
if (!entry) return null;
const overlay = entry.overlays[location];
if (overlay && 'disabled' in overlay) return null;
return overlay?.command ?? entry.discovery.binaries[0] ?? null;
}
/**
* The default remote pane command for `mode`.
*
* Agent CLIs (claude/opencode/codex/gemini/antigravity/…) are typically installed under
* per-user paths like ~/.local/bin or ~/.opencode/bin, added to PATH only by the remote
* user's interactive-login shell startup files (~/.zshrc etc.). ssh's remote-command
* execution is neither interactive nor login, so a bare `exec claude` sees only sshd's
* minimal default PATH and fails with "command not found" (exit 127) — confirmed via
* `tmux capture-pane` on the remain-on-exit-preserved dead pane. Route through
* `$SHELL -i -l -c`, the same fix shell mode uses, so PATH is fully resolved first.
*
* ⚠️ The per-CLI half is now READ FROM THE REGISTRY (`overlays.remote`), not from a
* hardcoded `Record<RemoteCommandMode, string>`. The table it replaces duplicated the
* registry exactly, with nothing keeping the two in step — a capability that is both wrong
* and unread is worse than an absent one, because the next person trusts it. Notes that were
* attached to individual rows and are still true:
* - claude carries `--dangerously-skip-permissions` so the remote agent runs
* non-interactively (no trust-folder prompt nothing on the remote can answer);
* `overlays.remote.command` on the claude entry is where that now lives.
* - `dsh` alone boots nothing — the launcher needs a profile, and the remote box's profile
* inventory is unknown here. The per-host `commands.deepseek` override names one.
* The per-host `commands.*` override remains the escape hatch for every mode.
*/
export function defaultRemoteCommandForMode(mode: SessionMode): string {
// Agent CLIs (claude/opencode/codex/gemini/antigravity) are typically installed
// under per-user paths like ~/.local/bin or ~/.opencode/bin, added to PATH only by
// the remote user's interactive-login shell startup files (~/.zshrc etc.). ssh's
// remote-command execution is neither interactive nor login, so a bare `exec
// claude` sees only sshd's minimal default PATH and fails with "command not
// found" (exit 127) — confirmed via `tmux capture-pane` on the
// remain-on-exit-preserved dead pane. Route through `$SHELL -i -l -c`, the same
// fix already used for shell mode below, so PATH is fully resolved before the
// CLI name is looked up.
const commands: Record<RemoteCommandMode, string> = {
// $SHELL, not a hardcoded bash: sshd sets it from the remote user's
// /etc/passwd entry, so this launches their actual login shell (zsh,
// fish, etc.). -i -l so it sources rc files (~/.zshrc etc.), matching
// the local shell-mode launch.
shell: `exec ${REMOTE_LOGIN_SHELL} -i -l`,
// Mirror the LOCAL claude default so the remote agent runs non-interactively
// (no trust-folder/permission prompt that nothing on the remote answers). The
// per-host `commands.claude` override stays the escape hatch.
claude: remoteLoginShellCommand('claude --dangerously-skip-permissions'),
opencode: remoteLoginShellCommand('opencode'),
codex: remoteLoginShellCommand('codex'),
gemini: remoteLoginShellCommand('gemini'),
antigravity: remoteLoginShellCommand('agy'),
pi: remoteLoginShellCommand('pi'),
};
return commands[mode as RemoteCommandMode] || commands.shell;
// $SHELL, not a hardcoded bash: sshd sets it from the remote user's /etc/passwd entry, so
// this launches their actual login shell (zsh, fish, …). `-i -l` so it sources rc files,
// matching the local shell-mode launch. Not templatable as overlay data: the shell is
// whatever the REMOTE passwd says, which is why `shell` is the one arm still written here.
const shellCommand = `exec ${REMOTE_LOGIN_SHELL} -i -l`;
const cli = overlayCliCommand(mode, 'remote');
return cli === null ? shellCommand : remoteLoginShellCommand(cli);
}
export function remoteSshTarget(host: Pick<RemoteHost, 'username' | 'host'>): string {
@@ -258,18 +279,22 @@ export async function checkRemoteTmuxAvailable(
}
/**
* The CLI binary each session mode runs on the remote host. Antigravity's
* binary is `agy` (the mode name is not the command); shell has no CLI to
* probe, so it is absent.
* The CLI binary a session mode runs on the remote host, read from the registry rather than
* from a hardcoded map. `shell` has no CLI to probe and resolves to undefined, which is what
* makes the probe return null for it.
*
* ⚠️ Deriving this CHANGES BEHAVIOUR, deliberately and in one direction. The map it replaces
* listed claude/opencode/codex/gemini/antigravity/pi/omp and simply omitted `grok` and
* `deepseek` — its own comment said the rule was "every mode except shell", so the two were
* an oversight from when those CLIs were added, not a decision. A remote grok or deepseek
* session therefore reported no version at all. It now probes `grok --version` /
* `dsh --version` through the same login-shell wrapper as its siblings.
*
* (`antigravity` is why this cannot be the mode name: its binary is `agy`.)
*/
const REMOTE_CLI_BIN: Partial<Record<SessionMode, string>> = {
claude: 'claude',
opencode: 'opencode',
codex: 'codex',
gemini: 'gemini',
antigravity: 'agy',
pi: 'pi',
};
function remoteCliBin(mode: SessionMode): string | undefined {
return getCli(mode)?.discovery.binaries[0];
}
/**
* Build the SSH command that reads the remote CLI's version (`claude --version`
@@ -285,7 +310,7 @@ export function buildRemoteCliVersionProbeCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
mode: SessionMode
): string | null {
const bin = REMOTE_CLI_BIN[mode];
const bin = remoteCliBin(mode);
if (!bin) return null;
return [
...buildSshConnectionArgs(host),
@@ -321,6 +346,84 @@ export async function probeRemoteCliVersion(
}
}
/**
* COD-108 — build the SSH command that asks whether THIS Codeman's durable
* remote tmux session (`-L codeman-remote -s codeman-ssh-<id>`) is still alive
* on the remote host.
*
* `has-session` exits 0 when the session exists, non-zero otherwise (and
* stderr is swallowed). Connection options come from the shared
* `buildSshConnectionArgs` so this probe reaches exactly the hosts the launch
* can reach — same port/identity/proxy/jump-host as `buildRemoteLaunchCommand`.
*/
export function buildRemoteSessionAliveCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
remoteSessionName: string
): string {
const [ssh, ...connectionArgs] = buildSshConnectionArgs(host);
const remoteCmd = `tmux -L codeman-remote has-session -t ${shellescape(remoteSessionName)} 2>/dev/null`;
return [ssh, ...connectionArgs, remoteSshTarget(host), shellescape(remoteCmd)].join(' ');
}
/**
* COD-108 — resolve whether THIS Codeman's durable remote tmux session is still
* alive on the remote host, for the auto-reconnect watcher.
*
* Returns:
* - `true` → the remote tmux session exists (the agent is still running
* on the remote; the LOCAL pane died from a transport drop →
* safe to auto-reconnect).
* - `false` → the remote session is gone (the agent exited cleanly and
* the remote tmux tore down; reviving would relaunch a fresh
* agent — must NOT auto-reconnect).
* - `undefined` → probe failed (host unreachable, ssh error, tmux missing).
* Callers MUST treat this as "do not reconnect": an
* unreachable host is not a reason to relaunch the agent.
*
* VITEST guard — returns `true` under test so a real ssh never runs; the
* command construction is covered by `buildRemoteSessionAliveCommand`.
*/
export async function remoteTmuxSessionAlive(
remote: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions,
remoteSessionName: string
): Promise<boolean | undefined> {
if (process.env.VITEST) return true;
const command = buildRemoteSessionAliveCommand(remote, remoteSessionName);
try {
await execAsync(command, { timeout: 15_000 });
return classifyRemoteAliveExit(0, false);
} catch (err) {
const e = err as { code?: unknown; killed?: boolean };
return classifyRemoteAliveExit(typeof e.code === 'number' ? e.code : null, e.killed === true);
}
}
/**
* Map the `has-session` probe's exit status onto the tri-state the watcher
* reads. Pure, so the mapping is unit-tested even though the probe itself is
* VITEST-guarded.
*
* ⚠️ `tmux has-session` prints NOTHING on success (measured: exit 0, empty
* stdout; the failure message goes to stderr), so the exit status is the ONLY
* signal. An earlier version read stdout and therefore classified every live
* remote session as gone, which silently disabled transport-drop reconnects.
*
* - exit 0 → the durable remote session exists → `true`.
* - exit 255 is ssh's own failure (unreachable host, auth, proxy/jump error)
* and a timeout arrives as `killed` with no numeric code: we learned
* nothing about the session → `undefined`, which the watcher treats as
* "do not revive".
* - any other non-zero status is the REMOTE command's: tmux's 1 for a missing
* session, or 127 when tmux is not installed there (no durable session can
* exist without it) → `false`.
*/
export function classifyRemoteAliveExit(code: number | null, killed: boolean): boolean | undefined {
if (killed) return undefined;
if (code === 0) return true;
if (code === null || code === 255) return undefined;
return false;
}
/**
* COD-105 — build the SSH command that lists `codeman-*` tmux sessions on a
* remote host's canonical `-L codeman` socket.
+19 -1
View File
@@ -108,6 +108,15 @@ export interface ReconnectSessionView {
isRemote: boolean;
/** Result of `isPaneDead(muxName)` for this session. */
paneDead: boolean;
/**
* Whether the DURABLE remote tmux session is still alive on the remote host.
* Tri-state: `true` = transport drop with the agent still running (safe to
* reattach); `false` = the remote session is gone (the agent exited cleanly
* via ctrl-c/ctrl-d/exit and the remote tmux tore down); `undefined` =
* unknown/unresolvable. The watcher must NOT revive when the remote session
* is gone or unknown — a clean exit must never auto-relaunch the agent.
*/
remoteAlive: boolean | undefined;
}
/**
@@ -130,7 +139,8 @@ export type ReconnectSkipReason =
| 'in-flight'
| 'not-due'
| 'exhausted'
| 'disabled';
| 'disabled'
| 'remote-gone';
export interface DecideReconnectInput {
session: ReconnectSessionView;
@@ -166,6 +176,14 @@ export function decideReconnect(input: DecideReconnectInput): ReconnectAction {
if (!session.paneDead) return { kind: 'skip', reason: 'pane-alive' };
// Intentional kill / detach must NEVER be auto-revived.
if (guarded) return { kind: 'skip', reason: 'guarded' };
// A clean exit tears down the durable remote tmux (the session's only pane
// exiting destroys it). Reviving is ONLY correct for a transport drop: the
// agent is still running on the remote, so the durable session must still
// exist. When it is gone (or status is unknown — probe failed/unreachable),
// the agent exited intentionally and must not be auto-relaunched (found
// live 2026-08-29: remote omp/opencode ctrl-c/ctrl-d auto-respawned fresh
// sessions; only claude's `|| --resume` accidentally masked it).
if (session.remoteAlive !== true) return { kind: 'skip', reason: 'remote-gone' };
const s = state ?? freshReconnectState();
+11
View File
@@ -99,6 +99,13 @@ export type HistoryInput = {
gitBranch?: string;
worktreeName?: string;
worktreeRepo?: string;
/**
* Set only by a non-claude transcript source (currently omp); the Claude
* scanner never stamps this; the meaningfulness floor below still counts a
* row with a `mode` as real, since that also signals "not claude" — see
* where it's read below for the isReal check this touches.
*/
mode?: string;
};
/** Mux process-stat view. */
@@ -175,6 +182,10 @@ export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionIte
overwrite(item, 'gitBranch', h.gitBranch);
overwrite(item, 'worktreeName', h.worktreeName);
overwrite(item, 'worktreeRepo', h.worktreeRepo);
// Claude rows never set this (they're implicitly claude); a non-claude
// transcript source (currently only omp) does, so a history-only row
// still gets a mode badge instead of reading as claude by default.
overwrite(item, 'mode', h.mode);
const ms = Date.parse(h.lastModified);
if (!Number.isNaN(ms) && item.lastActivityAt === undefined) item.lastActivityAt = ms;
}
+202
View File
@@ -0,0 +1,202 @@
/**
* @fileoverview Bridges the legacy per-mode spawn options (`buildSpawnCommand`'s option bag
* in tmux-manager.ts, unchanged on the wire since before this registry existed) onto the CLI
* registry's generic argv engine (`renderLaunch`).
*
* The per-mode `<Mode>Config` objects on `POST /api/sessions` predate the registry and stay
* on the wire for API compatibility (`docs/versioning-policy.md`), so SOMETHING has to know
* which field holds which CLI's config. That knowledge is DATA — `launch.legacyConfigField`
* and `launch.legacyConfigAliases`, declared once per entry in `config/cli-registry/stock.ts`
* — which is what lets this file stay a generic reader rather than a `switch (mode)`.
*
* An entry declaring NO `legacyConfigField` reads its params straight off the top-level
* option bag. That is claude, whose discrete `claudeMode`/`allowedTools`/`model`/
* `resumeSessionId` fields predate the `<Mode>Config` pattern — not a special case for
* claude, just the other of the two shapes the wire has always had.
*
* @module session-cli-registry-bridge
*/
import type { CliEntry } from './config/cli-registry/types.js';
import { renderLaunch, type EngineValues, type ParamValues } from './config/cli-registry/argv.js';
import { matchesPattern } from './config/cli-registry/patterns.js';
import { buildEffortCliArgs, sanitizeCliSessionName } from './session-cli-builder.js';
import { compareVersions } from './utils/dependency-checker.js';
import { getClaudeCliVersion } from './utils/claude-cli-resolver.js';
import { launcherDefaultTarget } from './utils/cli-launcher.js';
import { getCli } from './config/cli-registry/registry.js';
import type {
AntigravityConfig,
ClaudeMode,
CodexConfig,
DeepSeekConfig,
EffortLevel,
GeminiConfig,
GrokConfig,
OmpConfig,
OpenCodeConfig,
PiConfig,
} from './types/session.js';
export interface SpawnBridgeOptions {
mode: string;
sessionId: string;
model?: string;
claudeMode?: ClaudeMode;
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
antigravityConfig?: AntigravityConfig;
piConfig?: PiConfig;
grokConfig?: GrokConfig;
deepSeekConfig?: DeepSeekConfig;
ompConfig?: OmpConfig;
resumeSessionId?: string;
effort?: EffortLevel;
sessionName?: string;
claudeCliVersion?: string | null;
}
/**
* The raw legacy config object this entry's params should be read from: the declared
* `<Mode>Config` field, or the option bag itself when none is declared.
*/
function legacyConfigFor(entry: CliEntry, options: SpawnBridgeOptions): Record<string, unknown> | undefined {
const field = entry.launch.legacyConfigField;
if (field === undefined) return options as unknown as Record<string, unknown>;
return (options as unknown as Record<string, unknown>)[field] as Record<string, unknown> | undefined;
}
/**
* Same lookup, addressed by mode rather than by entry, for callers holding only a mode and an
* option bag (tmux-manager's env configuration). Returns undefined for an unregistered mode.
*/
export function legacyConfigForMode(
mode: string,
options: Record<string, unknown>
): Record<string, unknown> | undefined {
const entry = getCli(mode);
if (!entry) return undefined;
return legacyConfigFor(entry, options as unknown as SpawnBridgeOptions);
}
/**
* Build `ParamValues` for every declared `token`/`bool`/`enum` param by reading it out of the
* legacy config object through `legacyConfigAliases` (falling back to the param's own name).
* `engine`-sourced params are skipped — those come from `EngineValues`, never legacy config.
*/
function buildParamsFromLegacyConfig(entry: CliEntry, rawConfig: Record<string, unknown> | undefined): ParamValues {
const params: ParamValues = {};
if (!rawConfig) return params;
const aliases = entry.launch.legacyConfigAliases ?? {};
for (const [paramName, spec] of Object.entries(entry.launch.params)) {
if (spec.type === 'engine') continue;
const legacyKey = aliases[paramName] ?? paramName;
const value = rawConfig[legacyKey];
if (value === undefined) continue;
// Anything that is not already a string or boolean is DROPPED rather than coerced: the
// wire shape is Zod-validated upstream, so a surprise here means something is wrong,
// and `String({})` would happily produce a token nobody intended.
if (typeof value === 'string' || typeof value === 'boolean') {
params[paramName] = value;
}
}
return params;
}
/**
* The env vars this CLI declares in `env.configSetenv`, resolved from its legacy config
* object — i.e. the ones whose value comes from the CALLER rather than the server's own
* environment.
*
* ⚠️ Re-validated here against the declared `ParamSpec` even though the wire shape is already
* Zod-checked upstream. These values reach `tmux setenv`, and for DeepSeek the value IS a
* permission level: a builder must never trust its caller on a security-relevant field, and
* the cost of re-checking an enum is nothing.
*
* A value that fails validation is DROPPED, not defaulted — which is the safe direction: the
* var goes unset, and the CLI falls back to its own default (for dsh, `workspace-write`,
* which asks) rather than to something we guessed.
*/
export function configSetenvValues(
entry: CliEntry,
rawConfig: Record<string, unknown> | undefined
): Record<string, string> {
const out: Record<string, string> = {};
const mappings = entry.env.configSetenv;
if (!mappings || !rawConfig) return out;
const aliases = entry.launch.legacyConfigAliases ?? {};
for (const { name, fromParam } of mappings) {
const spec = entry.launch.params[fromParam];
if (!spec) continue; // schema-validated at load; belt and braces
const raw = rawConfig[aliases[fromParam] ?? fromParam];
if (typeof raw !== 'string') continue;
if (spec.type === 'enum' && !spec.values.includes(raw)) continue;
if (spec.type === 'token' && !matchesPattern(spec.pattern, raw)) continue;
out[name] = raw;
}
return out;
}
/**
* Which `capabilities.gates` are currently satisfied. `resolveVersion` is called AT MOST
* ONCE, and only when the entry actually declares a gate — a `--version` subprocess probe
* has no reason to run for an entry with none.
*/
function resolveGatesPassed(entry: CliEntry, resolveVersion: () => string | null): Set<string> {
const passed = new Set<string>();
const gateEntries = Object.entries(entry.capabilities.gates);
if (gateEntries.length === 0) return passed;
const cliVersion = resolveVersion();
if (!cliVersion) return passed; // fail-closed: an unknown version satisfies no gate
for (const [name, gate] of gateEntries) {
if (compareVersions(cliVersion, gate.minVersion) >= 0) passed.add(name);
}
return passed;
}
/**
* Render the spawn command for `entry` from the legacy option bag. Returns `undefined` for a
* `shell`-kind entry (or any entry declaring no launch variants), which callers take as "fall
* back to the local login-shell resolution" — shell has no CLI to template.
*/
export function buildSpawnCommandFromRegistry(entry: CliEntry, options: SpawnBridgeOptions): string | undefined {
if (entry.kind === 'shell' || entry.launch.variants.length === 0) return undefined;
const params = buildParamsFromLegacyConfig(entry, legacyConfigFor(entry, options));
const engineValues: EngineValues = {
sessionId: options.sessionId,
// Allowlist-sanitized (Unicode letters/digits + ` . _ : -`, 64 chars), matching
// buildNameCliArgs exactly — sanitizeCliSessionName is the injection guard for this
// value, NOT the `quote: 'double'` escaping on the --name arg (which only makes an
// unsafe value inert, it does not launder one into something meaningful).
sessionName: sanitizeCliSessionName(options.sessionName),
};
// Only a launcher CLI has one, and resolving it means a filesystem scan of the launcher's
// profile tree, so skip the lookup entirely for the eight entries that declare no profile.
if (entry.discovery.launcherProfile !== undefined) {
engineValues.launcherDefaultTarget = launcherDefaultTarget(entry) ?? undefined;
}
// Mirrors buildEffortCliArgs exactly: ultracode carries a fixed settings blob, every other
// level rides a plain `--effort <level>` flag. Reusing the canonical builder here (rather
// than re-deriving the ultracode special case) keeps the EFFORT_LEVELS allowlist and the
// settings-JSON shape single-sourced in session-cli-builder.ts.
const [effortFlag, effortValue] = buildEffortCliArgs(options.effort);
if (effortFlag === '--settings') engineValues.effortSettingsJson = effortValue;
else if (effortFlag === '--effort') engineValues.effortLevel = effortValue;
// Preserves buildSpawnCommand's original fallback exactly: an EXPLICIT `undefined` probes
// the local claude CLI (getClaudeCliVersion, null under vitest); an explicit `null` means
// "known to be unresolvable" and must not probe. The probe only ever runs from
// resolveGatesPassed, and only for an entry that actually declares a gate, so this stays
// generic without spawning a stray `claude --version` for every other CLI's launch.
const gatesPassed = resolveGatesPassed(entry, () =>
options.claudeCliVersion !== undefined ? options.claudeCliVersion : getClaudeCliVersion()
);
return renderLaunch(entry.launch, params, engineValues, gatesPassed);
}
+93 -9
View File
@@ -1,16 +1,32 @@
/**
* @fileoverview Recognizing Claude Code's workspace-trust dialog on screen.
* @fileoverview Recognizing Claude Code's workspace-trust dialog on screen, and
* working out which keystroke answers it.
*
* Claude asks once per directory before it will read or edit anything:
* Claude asks once per directory before it will read or edit anything. The
* layout has changed under us at least twice; both of these are live shapes:
*
* Quick safety check: Is this a project you created or one you trust? ...
* Quick safety check: Is this a project you created or one you trust? ... (<= 2.1.220)
* ❯ 1. Yes, I trust this folder
* 2. No, exit
* Enter to confirm · Esc to cancel
*
* Quick safety check: Is this a project you created or one you trust? ... (2.1.252)
* Security guide
* ❯ No, exit
* Yes, I trust this folder
* Enter to confirm · Esc to cancel
*
* Codeman sessions run permission-skipping or classifier-guarded modes, so the
* answer is always yes, and a session parked on this dialog is simply stuck.
*
* ⚠️ **Never press Enter without reading the selection.** The options are now
* unnumbered, REVERSED, and the highlighted default is "No, exit" — so the blind
* `\r` that answered the old layout picks *exit* on the new one and the pane
* dies (`Pane is dead (status 1)`) seconds after the session starts, which is
* exactly what a fresh case did on Claude Code 2.1.252. `trustDialogNextKey()`
* reads the `❯` marker instead and moves the cursor onto the trust option before
* it confirms anything.
*
* **Why the text has to be compacted.** tmux repaints a row by writing each word
* and then a cursor-forward (`\x1b[C`) instead of a space, and Ink colours each
* word separately, so the wire carries `I\x1b[Ctrust\x1b[Cthis\x1b[Cfolder`.
@@ -39,6 +55,25 @@ const TRUST_PHRASES = [
/** The dialog's own affordances. Prose that quotes the question will not have these. */
const CONFIRM_PHRASES = ['entertoconfirm', 'esctocancel', '2.no,exit'];
/** The option that answers yes, compacted. Identical text in both layouts. */
const YES_OPTION = 'yes,itrustthisfolder';
/** The option that quits Claude. It is the highlighted DEFAULT since 2.1.252. */
const NO_OPTION = 'no,exit';
/** Ink's selection marker. The only marked row while the dialog is up. */
const SELECTION_MARK = '❯';
/** A numbered option's `1.` / `2.` prefix, which the 2.1.220 layout put after the marker. */
const OPTION_NUMBER_PREFIX = /^\d+\./;
/** Move the selection one row down / up. Literal, so `send-keys -l` carries them. */
export const TRUST_KEY_DOWN = '\x1b[B';
export const TRUST_KEY_UP = '\x1b[A';
/** Confirm the highlighted option. */
export const TRUST_KEY_CONFIRM = '\r';
/**
* Charset-select sequences (`ESC ( B`), which tmux emits around styled runs and
* `stripAnsi` does not cover. Left in, they would land inside a phrase as a
@@ -65,6 +100,50 @@ export function isTrustDialogScreen(text: string): boolean {
return TRUST_PHRASES.some((p) => compact.includes(p)) && CONFIRM_PHRASES.some((p) => compact.includes(p));
}
/**
* Which option the `❯` marker sits on, or null when this text does not say.
*
* The LAST marked option wins. A pane capture holds exactly one frame and so
* exactly one marker, but the direct-PTY fallback reads an append-only buffer
* where every repaint since launch is still present — there the freshest frame
* is the one at the end, and an older one must not out-vote it.
*/
function selectedTrustOption(compact: string): { at: number; option: 'yes' | 'no' } | null {
let selected: { at: number; option: 'yes' | 'no' } | null = null;
for (let at = compact.indexOf(SELECTION_MARK); at >= 0; at = compact.indexOf(SELECTION_MARK, at + 1)) {
const after = compact.slice(at + SELECTION_MARK.length).replace(OPTION_NUMBER_PREFIX, '');
if (after.startsWith(YES_OPTION)) selected = { at, option: 'yes' };
else if (after.startsWith(NO_OPTION)) selected = { at, option: 'no' };
}
return selected;
}
/**
* The single keystroke that moves this dialog one step closer to "yes", or null
* when the screen does not show clearly enough to touch.
*
* One step per call on purpose: the caller re-reads the screen between
* keystrokes, so a moved cursor is CONFIRMED before Enter is pressed rather than
* assumed. Firing arrow+Enter together would re-create the failure this exists
* to prevent whenever the arrow is dropped (Ink drops keystrokes while it is
* still mounting a widget) — the Enter would then land on "No, exit".
*
* Returning null is the safe answer, not a failure: an unreadable frame means
* wait for the next repaint, and a layout whose options this cannot name means
* leave the dialog to the human. The caller's startup window bounds the waiting.
*/
export function trustDialogNextKey(text: string): string | null {
const compact = compactScreenText(text);
if (!compact.includes(YES_OPTION)) return null; // no trust option to steer onto
const selected = selectedTrustOption(compact);
if (!selected) return null; // marker missing, or not on an option we recognize
if (selected.option === 'yes') return TRUST_KEY_CONFIRM;
// On "No, exit". Which way the trust option lies is read from THIS frame — it
// sits below in 2.1.252 and above in the numbered layout before it — so the
// order flipping again costs a repaint, not a killed session.
return compact.includes(YES_OPTION, selected.at) ? TRUST_KEY_DOWN : TRUST_KEY_UP;
}
/**
* How long after the pane starts the dialog is still plausible. It renders
* before the main UI, so this only has to cover a slow first launch; leaving it
@@ -72,16 +151,21 @@ export function isTrustDialogScreen(text: string): boolean {
*/
export const TRUST_DIALOG_WINDOW_MS = 90_000;
/** Minimum gap between two Enter presses, and between two screen reads. */
/** Minimum gap between two keystrokes, and between two screen reads. */
export const TRUST_DIALOG_RETRY_MS = 1500;
/**
* Attempts before giving up and leaving the dialog to the user. A keystroke can
* land while Ink is still mounting the widget and be dropped, which is the other
* half of why sessions got stuck here; retrying costs nothing, but retrying
* forever would hammer Enter into whatever came next.
* Keystrokes before giving up and leaving the dialog to the user. A keystroke
* can land while Ink is still mounting the widget and be dropped, which is the
* other half of why sessions got stuck here; retrying costs nothing, but
* retrying forever would hammer Enter into whatever came next.
*
* Six rather than three because answering is no longer one press: the 2.1.252
* layout needs an arrow onto the trust option and then Enter, each confirmed
* against a re-read of the screen, so a cap of three left only one dropped
* keystroke of slack.
*/
export const TRUST_DIALOG_MAX_ATTEMPTS = 3;
export const TRUST_DIALOG_MAX_ATTEMPTS = 6;
/**
* How much of the append-only terminal buffer to read on a direct-PTY session,
+352 -62
View File
@@ -51,9 +51,13 @@ import {
type GeminiConfig,
type AntigravityConfig,
type PiConfig,
type GrokConfig,
type DeepSeekConfig,
type OmpConfig,
type SessionRemote,
type SessionDocker,
} from './types.js';
import { resolveAndClaimOmpSessionId } from './utils/omp-session-resolver.js';
import { probeDockerCliVersion } from './docker-hosts.js';
import { probeRemoteCliVersion } from './remote-hosts.js';
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
@@ -62,6 +66,8 @@ import { RalphTracker } from './ralph-tracker.js';
import { BashToolParser } from './bash-tool-parser.js';
import {
isTrustDialogScreen,
trustDialogNextKey,
TRUST_KEY_CONFIRM,
TRUST_DIALOG_WINDOW_MS,
TRUST_DIALOG_RETRY_MS,
TRUST_DIALOG_MAX_ATTEMPTS,
@@ -99,6 +105,8 @@ import {
} from './config/buffer-limits.js';
import { DEFAULT_TMUX_HISTORY_LIMIT } from './config/terminal-history.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import { getCli } from './config/cli-registry/registry.js';
import { resolveSessionCliVersion } from './utils/cli-resolver.js';
import {
buildInteractiveArgs,
buildPromptArgs,
@@ -169,28 +177,50 @@ const CTRL_L_PATTERN = /\x0c/g;
/** Pattern to split by newlines (CR or LF) */
const NEWLINE_SPLIT_PATTERN = /\r?\n/;
/** True for external-CLI run modes (non-Claude) that use their own TUI and output format. */
/**
* True for external-CLI run modes (non-Claude) that use their own TUI and output format:
* no Claude transcript, no hooks, no Claude-format token/BashTool parsing.
*
* ⚠️ Reads its OWN capability flag rather than being derived from `hooks` or `kind`, and
* that independence is load-bearing. `shell` has no hooks but is NOT external, so a
* predicate derived from hooks would sweep it in here; `deepseek` HAS hooks but IS
* external. Deriving one of these three predicates from another has already shipped a bug
* (see CliCapabilities' own doc comment), which is why they are three separate fields.
*
* An UNREGISTERED mode is treated as external — the conservative answer, since it disables
* Claude-specific parsing rather than pointing it at output that was never Claude's.
*/
export function isExternalCliMode(mode: SessionMode): boolean {
return mode === 'opencode' || mode === 'codex' || mode === 'gemini' || mode === 'antigravity' || mode === 'pi';
return getCli(mode)?.capabilities.external ?? true;
}
/** Display name for a run mode. Falls back to the raw id for an unregistered one. */
function getModeLabel(mode: SessionMode): string {
switch (mode) {
case 'opencode':
return 'OpenCode';
case 'codex':
return 'Codex';
case 'gemini':
return 'Gemini';
case 'antigravity':
return 'Antigravity';
case 'pi':
return 'Pi';
case 'shell':
return 'Shell';
case 'claude':
return 'Claude';
}
return getCli(mode)?.label ?? mode;
}
/**
* Does this CLI's launch spec gate anything on its own version?
*
* Only such a CLI needs its version probed at session start — probing one with no gates
* would spawn a `--version` subprocess whose answer nothing reads. Today that is claude
* (the `--name` flag, gated at 2.1.224), which is why the probe used to be written as
* `mode === 'claude'`.
*/
function cliNeedsVersionProbe(mode: SessionMode): boolean {
return Object.keys(getCli(mode)?.capabilities.gates ?? {}).length > 0;
}
/**
* Does this CLI ask for `COLORTERM=truecolor`?
*
* Read off the SAME `env.exports` list that `buildEnvExports()` emits into the tmux
* session, so the attach client and the pane cannot disagree about colour depth. These
* used to be two hand-maintained lists of mode names in two files that had to be edited
* together, with a comment in each asking the next person to remember.
*/
function cliExportsTruecolor(mode: SessionMode): boolean {
return (getCli(mode)?.env.exports ?? []).some((entry) => entry.name === 'COLORTERM' && entry.value === 'truecolor');
}
/**
@@ -202,8 +232,9 @@ function getModeLabel(mode: SessionMode): string {
* repaint via cursor positioning, so dropping the alt-screen switch is safe —
* content stays in the normal buffer. Excluded: `shell` (arbitrary programs like
* vim/less/htop legitimately need the alt screen), `opencode` (renders its own
* TUI that may rely on it) and `pi` (below). Keep parity with the replay-side
* strip in session-routes.ts.
* TUI that may rely on it), `pi` (below) and `grok` (a fullscreen alt-screen TUI
* with mouse support, i.e. the opencode case, not the Ink case). Keep parity
* with the replay-side strip in session-routes.ts.
*
* ⚠️ Being excluded here does NOT preserve the alt screen. Every excluded mode
* falls through to isMuxAltScreenOnlyStripMode(), which strips the alt-screen
@@ -218,7 +249,7 @@ function getModeLabel(mode: SessionMode): string {
* vim inside a tmux `shell` session.
*/
export function isAltScreenStripMode(mode: SessionMode): boolean {
return mode === 'codex' || mode === 'claude' || mode === 'gemini';
return getCli(mode)?.capabilities.altScreen === 'strip-full';
}
/**
@@ -415,6 +446,15 @@ export class Session extends EventEmitter {
// sequences split across PTY chunks can't slip past the alt-screen/scrollback
// strip (see _handleTerminalOutput / isAltScreenStripMode)
private _altScreenSeqCarry: string = '';
/**
* Mouse-tracking DECSET modes the CLI currently has ON, as observed while
* STRIPPING them out of the stream below. Kept as a set rather than a boolean
* because a TUI may enable 1002 and later disable 1000 (a mode it never
* enabled); tracking is on while any of them is.
*/
private _cliMouseModes = new Set<number>();
private _cliMouseTracking = false;
private resolvePromise: ((value: { result: string; cost: number }) => void) | null = null;
private rejectPromise: ((reason: Error) => void) | null = null;
private _promptResolved: boolean = false; // Guard against race conditions in runPrompt
@@ -426,8 +466,9 @@ export class Session extends EventEmitter {
private _lastPaneProbeAt = 0; // Throttle for the tmux screen probe
private _lastPaneProbeWorking: boolean | null = null; // Its last verdict (null = could not read)
private _trustDialogAccepted: boolean = false; // Stops the trust-dialog scan (answered, or given up)
private _trustDialogAttempts = 0; // Enter presses sent at the trust dialog
private _trustDialogAttempts = 0; // Keystrokes sent at the trust dialog
private _lastTrustDialogScanAt = 0; // Throttle for the trust-dialog screen read
private _trustDialogTimer: NodeJS.Timeout | null = null; // Re-read after a keystroke (see below)
private _interactiveStartedAt = 0; // When the interactive pane launched (bounds that scan)
private _taskTracker: TaskTracker;
@@ -499,6 +540,13 @@ export class Session extends EventEmitter {
private _antigravityConfig: AntigravityConfig | undefined;
// Pi configuration (only for mode === 'pi')
private _piConfig: PiConfig | undefined;
// Grok configuration (only for mode === 'grok')
private _grokConfig: GrokConfig | undefined;
// DeepSeek Harness configuration (only for mode === 'deepseek')
private _deepSeekConfig: DeepSeekConfig | undefined;
// OMP configuration (only for mode === 'omp')
private _ompConfig: OmpConfig | undefined;
private _resumeSessionId: string | undefined;
// Ephemeral env overrides (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Exported by tmux
@@ -510,7 +558,7 @@ export class Session extends EventEmitter {
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
private _effort: EffortLevel | undefined;
// tmux history-limit (scrollback lines) applied to this session's pane.
// tmux history-limit (scrollback lines) allocated when this session's pane is created.
private readonly _tmuxHistoryLimit: number;
// Remote execution metadata, present when this session runs over SSH through local tmux.
@@ -594,13 +642,19 @@ export class Session extends EventEmitter {
antigravityConfig?: AntigravityConfig;
/** Pi configuration (only for mode === 'pi') */
piConfig?: PiConfig;
/** Grok configuration (only for mode === 'grok') */
grokConfig?: GrokConfig;
/** DeepSeek Harness configuration (only for mode === 'deepseek') */
deepSeekConfig?: DeepSeekConfig;
/** OMP configuration (only for mode === 'omp') */
ompConfig?: OmpConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
/** Extra env vars exported to the CLI at spawn time (no disk persistence) */
envOverrides?: Record<string, string>;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) for this session's pane. */
/** tmux history-limit (scrollback lines) allocated when this session's pane is created. */
tmuxHistoryLimit?: number;
/** Restored per-session attachment history. May include server-private external paths. */
attachmentHistory?: SessionAttachmentHistoryItem[];
@@ -649,7 +703,13 @@ export class Session extends EventEmitter {
this._wireActivityAt = config.lastActivityAt || Date.now();
this._wireActivitySettleUntil = config.lastActivityAt ? Date.now() + WIRE_ACTIVITY_SETTLE_MS : 0;
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = config.resumeSessionId || this.id;
// For omp, `claudeSessionId` doubles as the generic "external transcript id"
// alias key mergeUnifiedSessions() folds a history row into its owning
// session by: omp mints its OWN uuid, unrelated to this Codeman id, so
// without this an omp conversation's Past-Sessions row (keyed by omp's
// id) would never merge with its own live/persisted row (keyed by this
// id) — it would just show up a second time.
this._claudeSessionId = config.resumeSessionId || config.ompConfig?.resumeSessionId || this.id;
// Restored from state.json on boot recovery. start() resets _claudeSessionId
// to the launch id even when re-attaching to a mux session whose CLI has
// moved on (a `/clear` before the restart), so this anchor is what lets the
@@ -702,6 +762,20 @@ export class Session extends EventEmitter {
if (config.piConfig) {
this._piConfig = config.piConfig;
}
// Apply OMP configuration
if (config.ompConfig) {
this._ompConfig = config.ompConfig;
}
// Apply DeepSeek Harness configuration
if (config.deepSeekConfig) {
this._deepSeekConfig = config.deepSeekConfig;
}
// Apply Grok configuration
if (config.grokConfig) {
this._grokConfig = config.grokConfig;
}
// Apply env overrides (exported at spawn, not persisted to disk).
// Legacy migration: pre-0.7.2 carried effort as the CLAUDE_CODE_EFFORT_LEVEL env var,
@@ -850,6 +924,34 @@ export class Session extends EventEmitter {
return this._remote;
}
/**
* `deepSeekConfig.statusReporting` verbatim: `undefined` when the caller sent
* none (i.e. ON), `false` when the user disarmed the status bridge for this
* session.
*
* Exposed because whether a dsh session can deliver `stop`/`blocked` is a
* per-SESSION fact, not a per-mode one, and `hooksAvailableForMode()` is pure
* and holds no `Session` reference by design. Undefined for every other mode,
* where the flag is meaningless.
*/
get deepSeekStatusReporting(): boolean | undefined {
return this._deepSeekConfig?.statusReporting;
}
/**
* This session's `DSH_HOME` override, if it set one.
*
* Deliberately ONE key rather than an `envOverrides` getter: the map can hold
* provider credentials (`DEEPSEEK_API_KEY`, `GEMINI_API_KEY`, …) and is
* kept off the public `SessionState` for exactly that reason. The transcript
* reader needs the profile tree's location and nothing else, so that is all
* this exposes.
*/
get deepSeekHomeOverride(): string | undefined {
const value = this._envOverrides?.DSH_HOME;
return value && value.trim() ? value.trim() : undefined;
}
/** Owning username in multi-user mode, else undefined. */
get owner(): string | undefined {
return this._owner;
@@ -1285,6 +1387,7 @@ export class Session extends EventEmitter {
niceValue: this._niceConfig.niceValue,
color: this._color,
flickerFilterEnabled: this._flickerFilterEnabled,
cliMouseTracking: this._cliMouseTracking || undefined,
cliVersion: this._cliVersion || undefined,
cliModel: this._cliModel || undefined,
cliAccountType: this._cliAccountType || undefined,
@@ -1294,6 +1397,9 @@ export class Session extends EventEmitter {
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
piConfig: this._piConfig,
grokConfig: this._grokConfig,
deepSeekConfig: this._deepSeekConfig,
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
effort: this._effort,
// COD-118: runtime-only — surfaced so the frontend can require explicit user
@@ -1420,7 +1526,11 @@ export class Session extends EventEmitter {
let needsNewSession = false;
if (this._muxSession && mux.isPaneDead(this._muxSession.muxName)) {
console.log('[Session] Dead pane detected, respawning:', this._muxSession.muxName);
const newPid = await mux.respawnPane(options.respawnPaneOptions);
// Confirmed dead — safe to resolve/pin now (see `_pinOmpRespawnId()`).
// `options.respawnPaneOptions` was built eagerly before this dead-pane
// check ran, so it still carries the pre-pin ompConfig; rebuild it.
this._pinOmpRespawnId();
const newPid = await mux.respawnPane(this._buildRespawnPaneOptions());
if (!newPid) {
console.error('[Session] Failed to respawn pane, will create new session');
needsNewSession = true;
@@ -1463,11 +1573,11 @@ export class Session extends EventEmitter {
cols: ptyCols,
rows: ptyRows,
cwd: resolveMuxAttachCwd(this.workingDir, this._remote, this._docker),
// COD-75: codex/gemini/antigravity/pi get COLORTERM=truecolor — mirrors buildEnvExports()
// in tmux-manager.ts so the attach client and the tmux session agree.
env: buildMuxAttachEnv(
this.mode === 'codex' || this.mode === 'gemini' || this.mode === 'antigravity' || this.mode === 'pi'
),
// COD-75: a CLI that declares `export COLORTERM=truecolor` gets it on the ATTACH
// client too. Both sides read the same registry entry, which is what stops the
// attach client and the tmux session from disagreeing — they used to be two
// hand-maintained lists of mode names that had to be edited in lockstep.
env: buildMuxAttachEnv(cliExportsTruecolor(this.mode)),
})
);
} catch (spawnErr) {
@@ -1506,6 +1616,9 @@ export class Session extends EventEmitter {
return false;
}
// Confirmed the mux session (and thus the pane) exists but this reattach
// is about to respawn it — safe to resolve/pin now.
this._pinOmpRespawnId();
const newPid = await mux.respawnPane(this._buildRespawnPaneOptions());
if (!newPid) {
console.error('[Session] reattachRemote: respawnPane failed for', this._muxSession.muxName);
@@ -1536,6 +1649,18 @@ export class Session extends EventEmitter {
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
piConfig: this._piConfig,
grokConfig: this._grokConfig,
deepSeekConfig: this._deepSeekConfig,
// OMP resolution/pinning does NOT happen here. This object is built
// EAGERLY — including on every boot-recovery reattach, before anyone
// knows whether the pane is actually dead — so resolving here mutated
// `_ompConfig`/`_claudeSessionId` even for a pane that was simply being
// reattached to, not respawned; with two omp tabs in the same case dir
// that mis-pinned the ALIVE session onto whichever file happened to be
// newest on disk (reported live in the Ark0N/Codeman#353 review). The
// real pin now happens in `_pinOmpRespawnId()`, called by callers ONLY
// once they've confirmed an actual respawn is about to happen.
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
@@ -1546,6 +1671,87 @@ export class Session extends EventEmitter {
};
}
/**
* OMP-only: resolve and PIN the exact conversation to continue when
* respawning a dead pane, so every later respawn reuses the same id
* instead of re-resolving (and re-risking picking up a DIFFERENT
* conversation that happened to touch this directory more recently). See
* the comment at the call site in {@link _buildRespawnPaneOptions} for why
* "newest file on disk" is safe here specifically. Non-omp modes and a
* session that already carries an explicit id pass through untouched.
*/
private _pinOmpRespawnId(): void {
// The omp-jsonl transcript reader is what this pin exists to feed, so ask for the
// reader rather than for the CLI's name.
if (getCli(this.mode)?.capabilities.transcript !== 'omp-jsonl') return;
if (this._ompConfig?.resumeSessionId) return;
// Callers MUST call this only immediately before an ACTUAL respawn (a
// confirmed-dead pane, or a genuine remote reattach) — never while merely
// building options that might not lead to a respawn. A fresh "Run OMP"
// click has no _muxSession yet and must never inherit whatever omp
// conversation happens to be newest on disk for this working directory
// (reported live 2026-08-27, fixed in 13a19f79); this guard keeps that
// fix intact now that resolution has moved out of the eager options build.
if (!this._muxSession) return;
const resolvedId = resolveAndClaimOmpSessionId(this.workingDir);
if (resolvedId) {
this._ompConfig = { ...this._ompConfig, resumeSessionId: resolvedId };
// Alias omp's own session uuid to this Codeman id — see the
// constructor's claudeSessionId comment for why this field is the
// (generically-named) mechanism that folds a Past-Sessions row back
// into its live/persisted session instead of duplicating it.
this._claudeSessionId = resolvedId;
return;
}
// Nothing unclaimed on disk (the dying process never got far enough to
// write a session file, or a sibling already claimed the only candidate)
// — fall back to the CLI's own "most recent" heuristic.
console.warn(
`[Session] OMP: no session file found under ${this.workingDir} to pin --resume on respawn; falling back to ambiguous --continue`
);
this._ompConfig = { ...this._ompConfig, continueSession: true };
}
/**
* Remember whether the CLI currently wants to be told about mouse clicks.
*
* The strip in {@link _handleTerminalOutput} is the ONLY place these sequences
* exist. After it, neither the browser nor xterm can ever learn that the CLI
* asked for mouse tracking, so `terminal.modes.mouseTrackingMode` is
* permanently 'none' for a stripped mode. The browser hand-encodes SGR reports
* to compensate (`_sendSyntheticSgrTap` in terminal-ui.js), and with no state
* to consult it had to do that on EVERY click, delivering mouse reports to a
* CLI that never asked for them. Publishing this through `toState()` is what
* lets the browser report a click only when the CLI is listening.
*
* Only the TRACKING modes count. 1005/1006 select an encoding and 1007 is
* alt-scroll; a CLI that picks SGR encoding without turning a tracking mode on
* is not asking about clicks, and counting those would put the stray reports
* straight back.
*
* This must stay in lockstep with the strip regex that calls it: a sequence
* removed from the stream but not recorded here is one the browser can neither
* see nor be told about.
*/
private _recordStrippedMouseMode(seq: string): void {
// eslint-disable-next-line no-control-regex
const match = /\x1b\[\?(\d+)([hl])$/.exec(seq);
if (!match) return;
const mode = Number(match[1]);
if (mode !== 1000 && mode !== 1001 && mode !== 1002 && mode !== 1003) return;
if (match[2] === 'h') this._cliMouseModes.add(mode);
else this._cliMouseModes.delete(mode);
this._syncCliMouseTracking();
}
/** Emit only on a real transition: a TUI re-emitting its enable on every repaint costs nothing. */
private _syncCliMouseTracking(): void {
const active = this._cliMouseModes.size > 0;
if (active === this._cliMouseTracking) return;
this._cliMouseTracking = active;
this.emit('mouseTrackingChanged', active);
}
private _handleTerminalOutput(data: string): void {
// Codex AND Claude Code emit sequences that wipe xterm.js scrollback, plus
// mouse-tracking enables that hijack the scroll wheel so the user can't reach
@@ -1598,7 +1804,10 @@ export class Session extends EventEmitter {
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, (seq) => {
this._recordStrippedMouseMode(seq);
return '';
});
}
}
@@ -1607,7 +1816,11 @@ export class Session extends EventEmitter {
// `Saved to: file://...` — that scanner (and its relaxed trust policy) is
// only enabled for codex-mode sessions. The web server applies the trust
// boundary for each request source.
const attachmentRequests = parseTerminalAttachmentRequests(data, { codexArtifacts: this.mode === 'codex' });
// Codex is the only CLI that announces generated artifacts in its pane output, and it
// is also the only one whose transcript is a rollout file — one implies the other.
const attachmentRequests = parseTerminalAttachmentRequests(data, {
codexArtifacts: getCli(this.mode)?.capabilities.transcript === 'codex-rollout',
});
for (const request of attachmentRequests) {
const seenKey = `${request.source}:${request.path}`;
if (this._attachmentMagicSeen.has(seenKey)) continue;
@@ -1641,6 +1854,10 @@ export class Session extends EventEmitter {
this._interactiveStartedAt = Date.now();
this._trustDialogAttempts = 0;
this._lastTrustDialogScanAt = 0;
if (this._trustDialogTimer) {
clearTimeout(this._trustDialogTimer);
this._trustDialogTimer = null;
}
// COD-118: if the PTY exit breaker has tripped (repeated non-zero exits in a
// short window), refuse to respawn. This is the uniform choke point that stops
@@ -1667,8 +1884,8 @@ export class Session extends EventEmitter {
// repaint/alt-screen mode; issue #154). Remote sessions run claude on
// another host, so a local probe wouldn't reflect their version; they get
// their own over-ssh probe below. Cached process-wide, best-effort.
if (this.mode === 'claude' && !this._remote && !this._docker && !this._cliVersion) {
const probedVersion = getClaudeCliVersion();
if (cliNeedsVersionProbe(this.mode) && !this._remote && !this._docker && !this._cliVersion) {
const probedVersion = resolveSessionCliVersion(this.mode);
if (probedVersion) {
this._cliVersion = probedVersion;
this.emit('cliInfoUpdated', {
@@ -1684,7 +1901,7 @@ export class Session extends EventEmitter {
// reports the HOST claude (wrong version, and leaving cliVersion undefined
// silently disables wheel-forwarding, #154). Probe the IN-CONTAINER version
// instead — deferred so the container is up after the mux attach below.
if (this.mode === 'claude' && this._docker && !this._cliVersion) {
if (cliNeedsVersionProbe(this.mode) && this._docker && !this._cliVersion) {
const dockerMeta = this._docker;
setTimeout(() => {
if (this._isStopped || this._cliVersion) return;
@@ -1710,7 +1927,7 @@ export class Session extends EventEmitter {
// is the unreliable path #154 was filed for, so remote Claude cases silently
// never got wheel-forwarding (noted in the #205 analysis). Probe over ssh,
// deferred so session start never waits on the ssh round-trip.
if (this.mode === 'claude' && this._remote && !this._cliVersion) {
if (cliNeedsVersionProbe(this.mode) && this._remote && !this._cliVersion) {
const remoteMeta = this._remote;
setTimeout(() => {
if (this._isStopped || this._cliVersion) return;
@@ -1751,6 +1968,9 @@ export class Session extends EventEmitter {
geminiConfig: this._geminiConfig,
antigravityConfig: this._antigravityConfig,
piConfig: this._piConfig,
grokConfig: this._grokConfig,
deepSeekConfig: this._deepSeekConfig,
ompConfig: this._ompConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
@@ -1762,8 +1982,14 @@ export class Session extends EventEmitter {
spawnErrLabel: 'mux attachment',
});
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
// Set claudeSessionId — when resuming, the Claude conversation ID is the
// resumed one. `_pinOmpRespawnId()` (called just above, inside
// `_setupOrAttachMuxSession()`'s dead-pane branch) may have JUST aliased
// this to omp's own session uuid — that already-resolved id must win
// over the generic `this.id` fallback, or this line clobbers it back
// to the Codeman id
// on every single respawn.
this._claudeSessionId = this._resumeSessionId || this._ompConfig?.resumeSessionId || this.id;
// For NEW mux sessions: wait for readiness then clean buffer
// For RESTORED mux sessions: don't do anything - client will fetch buffer on tab switch
@@ -1820,25 +2046,16 @@ export class Session extends EventEmitter {
// Fallback to direct PTY if mux is not used
if (!this.ptyProcess) {
// OpenCode sessions require tmux for env var injection (API keys via setenv)
if (this.mode === 'opencode') {
throw new Error('OpenCode sessions require tmux. Direct PTY fallback is not supported.');
}
// Codex sessions require tmux for OPENAI_API_KEY injection via setenv
if (this.mode === 'codex') {
throw new Error('Codex sessions require tmux. Direct PTY fallback is not supported.');
}
// Gemini sessions require tmux for Gemini/Google auth env injection via setenv
if (this.mode === 'gemini') {
throw new Error('Gemini sessions require tmux. Direct PTY fallback is not supported.');
}
// Antigravity sessions require tmux for env override injection via setenv
if (this.mode === 'antigravity') {
throw new Error('Antigravity sessions require tmux. Direct PTY fallback is not supported.');
}
// Pi sessions require tmux for env override injection via setenv
if (this.mode === 'pi') {
throw new Error('Pi sessions require tmux. Direct PTY fallback is not supported.');
// Every external CLI requires tmux and has NO direct-PTY fallback, because its
// secrets are injected with socket-scoped `tmux setenv` and so must never touch a
// spawn command line. DeepSeek additionally needs it for the HERDR_* status-bridge
// triple, without which the mode silently loses its definitive idle/blocked signals.
//
// Refusing is the only safe answer: falling back to a direct PTY would start the CLI
// unauthenticated (or, worse, tempt a future change into passing the key as an
// argument, where every process on the box can read it).
if (getCli(this.mode)?.capabilities.requiresMux) {
throw new Error(`${getModeLabel(this.mode)} sessions require tmux. Direct PTY fallback is not supported.`);
}
try {
// Pass --session-id to use the SAME ID as the Codeman session
@@ -1871,7 +2088,12 @@ export class Session extends EventEmitter {
}
// Set claudeSessionId — when resuming, the Claude conversation ID is the resumed one.
this._claudeSessionId = this._resumeSessionId || this.id;
// Mirrors the mux branch above and must not clobber it: this line runs
// unconditionally after both the mux and direct-PTY paths, so it also needs
// the ompConfig fallback or it stomps the mux branch's correctly-resolved
// OMP alias back to this.id on every mux/plain-reattach boot recovery
// (the "third reset point" — see DECISIONS.md).
this._claudeSessionId = this._resumeSessionId || this._ompConfig?.resumeSessionId || this.id;
this._pid = this.ptyProcess.pid;
console.log('[Session] Interactive PTY spawned with PID:', this._pid);
@@ -2005,6 +2227,15 @@ export class Session extends EventEmitter {
* makes a retry safe, since the terminal buffer is append-only and keeps the
* dialog in its tail long after it has been answered.
*
* ⚠️ **The keystroke is read off the screen, never assumed.** Claude Code
* 2.1.252 dropped the option numbers, put "No, exit" first, and highlights IT
* by default, so the bare `\r` this used to send now answers *exit*: a fresh
* case died (`Pane is dead (status 1)`) about six seconds after spawning.
* `trustDialogNextKey()` returns one step at a time — an arrow while the
* cursor is on the wrong option, Enter only once the screen shows it on the
* trust option — and this method re-reads the pane between the two, so a
* dropped arrow costs a repaint instead of the session.
*
* Three guards keep an Enter press off a live session: a startup-only window,
* a two-marker match (isTrustDialogScreen), and an attempt cap.
*/
@@ -2025,17 +2256,39 @@ export class Session extends EventEmitter {
this._terminalBuffer.value.slice(-TRUST_DIALOG_SCAN_BYTES);
if (!isTrustDialogScreen(screen)) return;
// Null means the frame does not say which option is highlighted. Waiting for
// the next repaint is the safe move; pressing Enter blind is the bug.
const key = trustDialogNextKey(screen);
if (key === null) return;
this._trustDialogAttempts++;
if (this._trustDialogAttempts > TRUST_DIALOG_MAX_ATTEMPTS) {
this._trustDialogAccepted = true; // leave it to the user rather than keep typing
console.warn(`[Session] Workspace trust dialog did not clear after retries: ${this.id}`);
return;
}
const step = key === TRUST_KEY_CONFIRM ? 'confirming' : 'moving to the trust option';
console.log(
`[Session] Auto-accepting workspace trust dialog for: ${this.id} (attempt ${this._trustDialogAttempts})`
`[Session] Auto-accepting workspace trust dialog for: ${this.id} (attempt ${this._trustDialogAttempts}, ${step})`
);
// Enter confirms the highlighted default, "1. Yes, I trust this folder".
this.writeViaMux('\r');
this.writeViaMux(key);
// ⚠️ Schedule the next read; do NOT wait for more PTY output. This scan only
// ever ran from `onData`, which was enough while one Enter answered the
// dialog. It is not enough now: the arrow that moves the cursor is the LAST
// output the pane produces, so a dialog left sitting on the trust option
// never gets its Enter and the worker stays parked on it forever (measured
// on a live 2.1.252 spawn: cursor moved at 6 s, then nothing). The timer is
// one-shot and self-rearming through this same path, and every exit route
// goes through _clearAllTimers().
// The +100ms puts the re-entry OUTSIDE the scan throttle above; firing at
// exactly the throttle boundary would let the scan return early and break
// the chain with the dialog still on screen.
if (this._trustDialogTimer) clearTimeout(this._trustDialogTimer);
this._trustDialogTimer = setTimeout(() => {
this._trustDialogTimer = null;
this._maybeAcceptTrustDialog();
}, TRUST_DIALOG_RETRY_MS + 100);
}
/**
@@ -2162,10 +2415,36 @@ export class Session extends EventEmitter {
this._isWorking = false;
this._status = 'idle';
this._lastPromptTime = Date.now();
if (wasWorking) this._maybeCaptureOmpSessionId();
this.emit('idle');
}
}
/**
* A brand-new omp session (never yet respawned, so
* {@link _pinOmpRespawnId} has never run) has no captured
* omp-native session id: `_claudeSessionId` still defaults to this
* session's OWN Codeman id from the constructor. Until something aliases
* it, the omp history scan's row for this exact conversation (keyed by
* omp's own uuid) merges with nothing and shows up a second time. The
* first turn going idle is the first moment omp has definitely written
* its session file, so resolve and alias it here — best-effort, and only
* once (skips once `_claudeSessionId` differs from `this.id`, whether from
* this capture or a resume/respawn that already resolved one).
*/
private _maybeCaptureOmpSessionId(): void {
if (getCli(this.mode)?.capabilities.transcript !== 'omp-jsonl' || this._claudeSessionId !== this.id) return;
try {
const resolvedId = resolveAndClaimOmpSessionId(this.workingDir);
if (resolvedId) {
this._claudeSessionId = resolvedId;
this._ompConfig = { ...this._ompConfig, resumeSessionId: resolvedId };
}
} catch {
// Best-effort: a failed capture just means the next respawn tries again.
}
}
/**
* Process expensive parsers (ANSI strip, Ralph, bash tool, token, CLI info, task descriptions).
* Called on a throttled schedule (every EXPENSIVE_PROCESS_INTERVAL_MS) instead of on every
@@ -2525,10 +2804,21 @@ export class Session extends EventEmitter {
this._messages = [];
this._lineBuffer = '';
this._altScreenSeqCarry = '';
// A restarted pane starts with no mouse mode: the new program has not asked
// for one yet, and carrying the old CLI's state over would report clicks
// into a program that never enabled tracking.
this._cliMouseModes.clear();
this._syncCliMouseTracking();
this._markActivity(true);
}
private _clearAllTimers(): void {
// Clear the workspace-trust follow-up read
if (this._trustDialogTimer) {
clearTimeout(this._trustDialogTimer);
this._trustDialogTimer = null;
}
// Clear activity timeout to prevent memory leak
if (this.activityTimeout) {
clearTimeout(this.activityTimeout);
+60 -14
View File
@@ -40,6 +40,8 @@ import {
} from './types.js';
import { Debouncer, MAX_SESSION_TOKENS } from './utils/index.js';
import { dataPath, CODEMAN_INSTANCE } from './config/instance.js';
import { normalizeSessionOrder } from './session-order.js';
import { validateTabLayout, type TabLayout } from './tab-layout.js';
/** Debounce delay for batching state writes (ms) */
const SAVE_DEBOUNCE_MS = 500;
@@ -281,6 +283,9 @@ export class StateStore {
if (this.state.sessionOrder) {
parts.push(`"sessionOrder":${JSON.stringify(this.state.sessionOrder)}`);
}
if (this.state.tabLayouts !== undefined) {
parts.push(`"tabLayouts":${JSON.stringify(this.state.tabLayouts)}`);
}
return `{${parts.join(',')}}`;
}
@@ -514,22 +519,28 @@ export class StateStore {
*/
cleanupStaleSessions(activeSessionIds: Set<string>): {
count: number;
cleaned: Array<{ id: string; name?: string }>;
cleaned: Array<{ id: string; name?: string; owner?: string }>;
} {
const allSessionIds = Object.keys(this.state.sessions);
const cleaned: Array<{ id: string; name?: string }> = [];
const staleIds = new Set(Object.keys(this.state.sessions).filter((sessionId) => !activeSessionIds.has(sessionId)));
return this.cleanupSessionsByIds(staleIds);
}
for (const sessionId of allSessionIds) {
if (!activeSessionIds.has(sessionId)) {
if (this.state.sessions[sessionId]?.pinned === true) continue; // COD-142: pinned records persist even with no live session
const name = this.state.sessions[sessionId]?.name;
cleaned.push({ id: sessionId, name });
delete this.state.sessions[sessionId];
this.cachedSessionJsons.delete(sessionId);
this.dirtySessions.delete(sessionId);
// Also clean up Ralph state for this session
this.ralphStates.delete(sessionId);
}
/** Deletes only confirmed stale session IDs, retaining records pinned after confirmation. */
cleanupSessionsByIds(sessionIds: ReadonlySet<string>): {
count: number;
cleaned: Array<{ id: string; name?: string; owner?: string }>;
} {
const cleaned: Array<{ id: string; name?: string; owner?: string }> = [];
for (const sessionId of sessionIds) {
const session = this.state.sessions[sessionId];
if (!session || session.pinned === true) continue; // COD-142: pinned records persist even with no live session
cleaned.push({ id: sessionId, name: session.name, owner: session.owner });
delete this.state.sessions[sessionId];
this.cachedSessionJsons.delete(sessionId);
this.dirtySessions.delete(sessionId);
// Also clean up Ralph state for this session
this.ralphStates.delete(sessionId);
}
if (cleaned.length > 0) {
@@ -664,6 +675,41 @@ export class StateStore {
this.save();
}
/** Returns an owner layout, or null before that owner has been migrated. */
getTabLayout(owner: string): TabLayout | null {
const layouts = this.state.tabLayouts;
return layouts && Object.hasOwn(layouts, owner) ? layouts[owner] : null;
}
/** Returns a defensive snapshot of every stored owner layout. */
getTabLayouts(): Record<string, TabLayout> {
return structuredClone(this.state.tabLayouts ?? {});
}
/** Validates and atomically persists one owner layout. */
setTabLayout(owner: string, layout: TabLayout): void {
const validated = validateTabLayout(layout);
this.state.tabLayouts = { ...(this.state.tabLayouts ?? {}), [owner]: validated };
this.save();
}
/** Atomically publishes validated owner layouts and their latest global compatibility projection. */
commitTabLayoutProjection(
layouts: Readonly<Record<string, TabLayout>>,
projectOrder: (latest: readonly string[]) => readonly string[]
): { layouts: Record<string, TabLayout>; sessionOrder: string[] } {
const validated = Object.fromEntries(
Object.entries(layouts).map(([owner, layout]) => [owner, validateTabLayout(layout)])
);
const sessionOrder = normalizeSessionOrder(projectOrder([...(this.state.sessionOrder ?? [])]));
const nextLayouts = { ...(this.state.tabLayouts ?? {}), ...validated };
this.state.tabLayouts = nextLayouts;
this.state.sessionOrder = sessionOrder;
this.save();
return { layouts: structuredClone(validated), sessionOrder: [...sessionOrder] };
}
/** Resets all state to initial values and saves immediately. */
reset(): void {
this.state = createInitialState();
+81
View File
@@ -0,0 +1,81 @@
/**
* @fileoverview Pure compatibility translation between legacy session order and owner tab layouts.
*/
import { mergeSessionOrder, normalizeSessionOrder } from './session-order.js';
import {
normalizeTabLayout,
validateTabLayout,
type TabLayout,
type TabRef,
type TabRefMetadata,
} from './tab-layout.js';
export interface OwnerOrderProjection {
owner: string;
ownedIds: readonly string[];
order: readonly string[];
}
export function applyLegacySessionRank(
input: TabLayout,
requestedOrder: readonly string[],
metadata: readonly TabRefMetadata[]
): TabLayout {
const layout = validateTabLayout(input);
const requestedRank = new Map(normalizeSessionOrder(requestedOrder).map((id, index) => [id, index]));
const sessionMetadata = new Map<string, TabRefMetadata>();
for (const item of metadata) {
if (item.kind !== 'session' || !item.ownerValid || !item.visible || sessionMetadata.has(item.id)) continue;
sessionMetadata.set(item.id, item);
}
const isRanked = (ref: TabRef): boolean =>
ref.kind === 'session' && sessionMetadata.has(ref.id) && requestedRank.has(ref.id);
const prepare = (ref: TabRef): TabRef => {
if (ref.kind !== 'session') return { ...ref };
const item = sessionMetadata.get(ref.id);
const ownerValidParent = item?.parentSessionId && sessionMetadata.has(item.parentSessionId);
return ownerValidParent ? { ...ref, placement: 'manual' } : { ...ref };
};
const rankContainer = (refs: readonly TabRef[]): TabRef[] => {
const ranked = refs
.filter(isRanked)
.map(prepare)
.sort((a, b) => requestedRank.get(a.id)! - requestedRank.get(b.id)!);
let rankedIndex = 0;
return refs.map((ref) => (isRanked(ref) ? ranked[rankedIndex++] : { ...ref }));
};
const transformed: TabLayout = {
...layout,
groups: layout.groups.map((group) => ({ ...group, refs: rankContainer(group.refs) })),
ungrouped: rankContainer(layout.ungrouped),
};
return normalizeTabLayout(transformed, metadata);
}
export function recomposeGlobalSessionOrder(
current: readonly string[],
projections: readonly OwnerOrderProjection[],
preferred?: readonly string[]
): string[] {
let result = mergeSessionOrder([...(preferred ?? current)], [...current]);
for (const projection of projections) {
const ownedIds = normalizeSessionOrder(projection.ownedIds);
const owned = new Set(ownedIds);
const canonical = normalizeSessionOrder(projection.order).filter((id) => owned.has(id));
const canonicalSet = new Set(canonical);
for (const id of ownedIds) {
if (canonicalSet.has(id)) continue;
canonicalSet.add(id);
canonical.push(id);
}
let canonicalIndex = 0;
const recomposed = result.map((id) => (owned.has(id) ? canonical[canonicalIndex++] : id));
recomposed.push(...canonical.slice(canonicalIndex));
result = normalizeSessionOrder(recomposed);
}
return result;
}
+144
View File
@@ -0,0 +1,144 @@
/**
* @fileoverview Owner-scoped tab-layout persistence and legacy migration primitives.
*
* This module is deliberately independent of routes and runtime managers. Callers
* provide persisted/live session facts plus saved webviews in server-store order.
*/
import { normalizeTabLayout, type TabLayout, type TabRef, type TabRefMetadata } from './tab-layout.js';
export const SINGLE_USER_LAYOUT_OWNER = '@single';
export interface TabLayoutSessionRecord {
id: string;
owner?: string;
createdAt: number;
parentSessionId?: string;
}
export interface TabLayoutWebviewRecord {
id: string;
owner?: string;
}
export interface TabLayoutMigrationInput {
owner: string;
layouts?: Readonly<Record<string, TabLayout>>;
sessionOrder?: readonly string[];
persistedSessions: readonly TabLayoutSessionRecord[];
liveSessions: readonly TabLayoutSessionRecord[];
/** Saved webviews in authoritative server-store order. */
webviews: readonly TabLayoutWebviewRecord[];
/** Required only when creating a layout, making migration deterministic in tests. */
updatedAt?: string;
}
export interface TabLayoutMigrationResult {
layout: TabLayout;
layouts: Record<string, TabLayout>;
created: boolean;
}
/** Resolve the persistence key without accepting an owner key from a client. */
export function ownerLayoutKey(username?: string): string {
return username || SINGLE_USER_LAYOUT_OWNER;
}
function recordOwner(record: { owner?: string }): string {
return record.owner ?? SINGLE_USER_LAYOUT_OWNER;
}
function compareSessions(a: TabLayoutSessionRecord, b: TabLayoutSessionRecord): number {
return a.createdAt - b.createdAt || (a.id < b.id ? -1 : a.id > b.id ? 1 : 0);
}
function collectSessions(input: TabLayoutMigrationInput): Map<string, TabLayoutSessionRecord> {
const sessions = new Map<string, TabLayoutSessionRecord>();
for (const record of input.persistedSessions) sessions.set(record.id, { ...record });
// A matching live record is authoritative as a whole. In particular, absent
// optional owner/parent fields mean single-user ownership and root lineage;
// retaining those fields from a stale persisted copy changes their semantics.
for (const record of input.liveSessions) sessions.set(record.id, { ...record });
return sessions;
}
function buildMetadata(
input: TabLayoutMigrationInput,
sessions: ReadonlyMap<string, TabLayoutSessionRecord>
): TabRefMetadata[] {
const ownerSessions = [...sessions.values()]
.filter((record) => recordOwner(record) === input.owner)
.sort(compareSessions);
const sessionOrder = new Map(ownerSessions.map((record, index) => [record.id, index]));
const metadata: TabRefMetadata[] = [...sessions.values()].map((record) => ({
kind: 'session',
id: record.id,
ownerValid: recordOwner(record) === input.owner,
visible: true,
order: sessionOrder.get(record.id) ?? record.createdAt,
parentSessionId: record.parentSessionId,
}));
const webviewOffset = ownerSessions.length;
input.webviews.forEach((record, index) => {
metadata.push({
kind: 'webview',
id: record.id,
ownerValid: recordOwner(record) === input.owner,
visible: true,
order: webviewOffset + index,
});
});
return metadata;
}
/**
* Normalize an existing owner layout, or idempotently migrate legacy flat order.
* Unknown stored refs remain unknown to metadata and are therefore preserved.
* No input object is mutated; validation/capacity failure is atomic.
*/
export function normalizeOrMigrateOwnerTabLayout(input: TabLayoutMigrationInput): TabLayoutMigrationResult {
const sessions = collectSessions(input);
const metadata = buildMetadata(input, sessions);
const existing = input.layouts && Object.hasOwn(input.layouts, input.owner) ? input.layouts[input.owner] : undefined;
if (existing) {
const layout = normalizeTabLayout(existing, metadata);
return { layout, layouts: { ...(input.layouts ?? {}), [input.owner]: layout }, created: false };
}
const ownerSessions = [...sessions.values()].filter((record) => recordOwner(record) === input.owner);
const ownerSessionById = new Map(ownerSessions.map((record) => [record.id, record]));
const liveOwnerIds = new Set(
input.liveSessions.filter((record) => recordOwner(record) === input.owner).map((record) => record.id)
);
const seen = new Set<string>();
const orderedSessions: TabLayoutSessionRecord[] = [];
for (const id of input.sessionOrder ?? []) {
const record = ownerSessionById.get(id);
if (!record || seen.has(id)) continue;
seen.add(id);
orderedSessions.push(record);
}
for (const record of ownerSessions.filter((item) => !seen.has(item.id)).sort(compareSessions)) {
seen.add(record.id);
orderedSessions.push(record);
}
const refs: TabRef[] = orderedSessions.map((record) => {
const manual = record.parentSessionId !== undefined && liveOwnerIds.has(record.parentSessionId);
return manual ? { kind: 'session', id: record.id, placement: 'manual' } : { kind: 'session', id: record.id };
});
for (const webview of input.webviews) {
if (recordOwner(webview) === input.owner) refs.push({ kind: 'webview', id: webview.id });
}
const layout = normalizeTabLayout(
{
version: 0,
groups: [],
ungrouped: refs,
updatedAt: input.updatedAt ?? new Date().toISOString(),
},
metadata
);
return { layout, layouts: { ...(input.layouts ?? {}), [input.owner]: layout }, created: true };
}
+678
View File
@@ -0,0 +1,678 @@
/**
* @fileoverview Owner-scoped authoritative tab-layout coordination.
*
* This is the single mutation boundary between the pure layout model, persisted
* state, live sessions, saved webviews, and SSE. Lifecycle callers describe one
* completed server action; this service performs at most one versioned write.
*/
import type { StateStore } from './state-store.js';
import { mergeSessionOrder, normalizeSessionOrder } from './session-order.js';
import { applyLegacySessionRank, recomposeGlobalSessionOrder } from './tab-layout-legacy-order.js';
import {
flattenOwnerSessionOrder,
materializeOrphans,
normalizeTabLayout,
TabLayoutValidationError,
validateTabLayout,
type TabLayout,
type TabRef,
type TabRefMetadata,
} from './tab-layout.js';
import {
normalizeOrMigrateOwnerTabLayout,
SINGLE_USER_LAYOUT_OWNER,
type TabLayoutSessionRecord,
type TabLayoutWebviewRecord,
} from './tab-layout-persistence.js';
import { SseEvent } from './web/sse-events.js';
export interface TabLayoutSessionLike {
id: string;
owner?: string;
createdAt: number;
parentSessionId?: string;
}
interface TabLayoutServiceDeps {
store: Pick<
StateStore,
'getTabLayout' | 'getTabLayouts' | 'getSessions' | 'getSessionOrder' | 'commitTabLayoutProjection'
>;
sessions: ReadonlyMap<string, TabLayoutSessionLike>;
readWebviews(): Promise<readonly TabLayoutWebviewRecord[]>;
broadcast(event: string, data: unknown): void;
broadcastSessionOrder(change: SessionOrderProjectionChange): void;
now?: () => string;
}
export type TabLayoutPutResult = { status: 'updated'; layout: TabLayout } | { status: 'conflict'; layout: TabLayout };
export interface LegacyOrderActor {
owner: string;
isAdmin: boolean;
}
export interface SessionOrderProjectionChange {
changedOwnerOrders: Record<string, string[]>;
globalOrder: string[];
globalChanged: boolean;
}
export interface LegacyOrderPutResult extends SessionOrderProjectionChange {
order: string[];
}
export interface RemovedTabLayoutSession {
id: string;
owner?: string;
}
interface PreparedOwnerLayout {
current: TabLayout | null;
authoritative: TabLayout;
metadata: TabRefMetadata[];
needsReconciliationCommit: boolean;
}
interface OwnerProjectionPublication {
owner: string;
previous: TabLayout | null;
next: TabLayout;
metadata: readonly TabRefMetadata[];
excludedSessionIds?: ReadonlySet<string>;
}
interface PreparedOrderProjection {
owner: string;
previousOrder: string[];
authoritativeBeforeIds: string[];
excludedIds: string[];
currentIds: string[];
order: string[];
}
const ownerOf = (record: { owner?: string }): string => record.owner ?? SINGLE_USER_LAYOUT_OWNER;
const refKey = (ref: Pick<TabRef, 'kind' | 'id'>): string => `${ref.kind}\u0000${ref.id}`;
const sameLayout = (a: TabLayout, b: TabLayout): boolean => JSON.stringify(a) === JSON.stringify(b);
const sameOrder = (a: readonly string[], b: readonly string[]): boolean =>
a.length === b.length && a.every((id, index) => id === b[index]);
export class TabLayoutService {
private restorationState: 'pending' | 'complete' | 'failed' | 'skipped' = 'pending';
private readonly ownerQueues = new Map<string, Promise<void>>();
constructor(private readonly deps: TabLayoutServiceDeps) {}
private async withOwner<T>(owner: string, task: () => Promise<T>): Promise<T> {
const previous = this.ownerQueues.get(owner) ?? Promise.resolve();
const run = previous.catch(() => undefined).then(task);
const tail = run.then(
() => undefined,
() => undefined
);
this.ownerQueues.set(owner, tail);
try {
return await run;
} finally {
if (this.ownerQueues.get(owner) === tail) this.ownerQueues.delete(owner);
}
}
/** Acquire multiple owner queues in stable order so overlapping bulk cleanups cannot deadlock. */
private async withOwners<T>(owners: readonly string[], task: () => Promise<T>, index = 0): Promise<T> {
if (index >= owners.length) return task();
return this.withOwner(owners[index], () => this.withOwners(owners, task, index + 1));
}
markRestorationComplete(): void {
this.restorationState = 'complete';
}
markRestorationFailed(): void {
this.restorationState = 'failed';
}
markRestorationSkipped(): void {
this.restorationState = 'skipped';
}
assertDeletionReady(): void {
if (this.restorationState === 'complete' || this.restorationState === 'skipped') return;
throw new Error(`Tab layout restoration is ${this.restorationState}; destructive deletion is unavailable`);
}
/** Repair/migrate every owner visible after startup restoration. */
async reconcileAfterRestoration(): Promise<void> {
if (this.restorationState !== 'complete') return;
const { persisted, live } = this.sessionRecords();
const webviews = await this.deps.readWebviews();
const owners = new Set<string>();
for (const record of [...persisted, ...live, ...webviews]) owners.add(ownerOf(record));
for (const owner of owners) await this.get(owner);
}
private sessionRecords(): { persisted: TabLayoutSessionRecord[]; live: TabLayoutSessionRecord[] } {
const persisted = Object.entries(this.deps.store.getSessions()).map(([id, record]) => ({
id,
owner: record.owner,
createdAt: record.createdAt,
parentSessionId: record.parentSessionId,
}));
const live = [...this.deps.sessions.values()].map((record) => ({
id: record.id,
owner: record.owner,
createdAt: record.createdAt,
parentSessionId: record.parentSessionId,
}));
return { persisted, live };
}
private async facts(owner: string): Promise<{
persisted: TabLayoutSessionRecord[];
live: TabLayoutSessionRecord[];
webviews: readonly TabLayoutWebviewRecord[];
metadata: TabRefMetadata[];
}> {
const { persisted, live } = this.sessionRecords();
const webviews = await this.deps.readWebviews();
const sessions = new Map<string, TabLayoutSessionRecord>();
for (const record of persisted) sessions.set(record.id, record);
for (const record of live) sessions.set(record.id, record);
const ownedSessions = [...sessions.values()]
.filter((record) => ownerOf(record) === owner)
.sort((a, b) => a.createdAt - b.createdAt || (a.id < b.id ? -1 : a.id > b.id ? 1 : 0));
const sessionOrder = new Map(ownedSessions.map((record, index) => [record.id, index]));
const metadata: TabRefMetadata[] = [...sessions.values()].map((record) => ({
kind: 'session',
id: record.id,
ownerValid: ownerOf(record) === owner,
visible: true,
order: sessionOrder.get(record.id) ?? record.createdAt,
parentSessionId: record.parentSessionId,
}));
const offset = ownedSessions.length;
webviews.forEach((record, index) =>
metadata.push({
kind: 'webview',
id: record.id,
ownerValid: ownerOf(record) === owner,
visible: true,
order: offset + index,
})
);
return { persisted, live, webviews, metadata };
}
private prepareCommit(base: TabLayout, next: TabLayout): TabLayout {
return validateTabLayout({
...next,
version: base.version + 1,
updatedAt: (this.deps.now ?? (() => new Date().toISOString()))(),
});
}
private prepareOrderProjection(item: OwnerProjectionPublication): PreparedOrderProjection {
const excluded = item.excludedSessionIds ?? new Set<string>();
const authoritativeBeforeIds = item.metadata
.filter((fact) => fact.kind === 'session' && fact.ownerValid && fact.visible)
.map((fact) => fact.id);
const facts = authoritativeBeforeIds.filter((id) => !excluded.has(id));
const visible = new Set(facts);
const rawPrevious = item.previous ? flattenOwnerSessionOrder(item.previous) : [];
const rawNext = flattenOwnerSessionOrder(item.next);
const previousOrder = rawPrevious.filter((id) => visible.has(id) || excluded.has(id));
const order = rawNext.filter((id) => visible.has(id) && !excluded.has(id));
const excludedIds = normalizeSessionOrder([...excluded]);
return {
owner: item.owner,
previousOrder,
authoritativeBeforeIds: normalizeSessionOrder([...authoritativeBeforeIds, ...excluded]),
excludedIds,
currentIds: normalizeSessionOrder([...order, ...facts]),
order,
};
}
private projectOrder(
latest: readonly string[],
projections: readonly PreparedOrderProjection[],
preferred?: readonly string[]
): string[] {
const before = normalizeSessionOrder(latest);
const removed = new Set(
projections.flatMap((projection) => projection.excludedIds.filter((id) => !projection.currentIds.includes(id)))
);
return recomposeGlobalSessionOrder(
before.filter((id) => !removed.has(id)),
projections.map((projection) => ({
owner: projection.owner,
ownedIds: projection.currentIds,
order: projection.order,
})),
preferred
);
}
private publish(
layouts: Readonly<Record<string, TabLayout>>,
publications: readonly OwnerProjectionPublication[],
preferred?: readonly string[]
): SessionOrderProjectionChange {
const projections = publications.map((item) => this.prepareOrderProjection(item));
let beforeOrder: string[] = [];
const accepted = this.deps.store.commitTabLayoutProjection(layouts, (latest) => {
beforeOrder = normalizeSessionOrder(latest);
return this.projectOrder(beforeOrder, projections, preferred);
});
const changedEntries: Array<[string, string[]]> = [];
for (const projection of projections) {
const beforeIds = new Set(projection.authoritativeBeforeIds);
const currentIds = new Set(projection.currentIds);
const persistedBefore = beforeOrder.filter((id) => beforeIds.has(id));
const persistedAfter = accepted.sessionOrder.filter((id) => currentIds.has(id));
const layoutOrderChanged = !sameOrder(projection.previousOrder, projection.order);
const persistedOwnerSliceChanged = !sameOrder(persistedBefore, persistedAfter);
if (layoutOrderChanged || persistedOwnerSliceChanged) {
changedEntries.push([projection.owner, persistedAfter]);
}
}
const change: SessionOrderProjectionChange = {
changedOwnerOrders: Object.fromEntries(changedEntries),
globalOrder: [...accepted.sessionOrder],
globalChanged: !sameOrder(beforeOrder, accepted.sessionOrder),
};
for (const [owner, layout] of Object.entries(accepted.layouts)) {
this.deps.broadcast(SseEvent.TabLayoutChanged, { owner, version: layout.version });
}
if (changedEntries.length > 0 || change.globalChanged) this.deps.broadcastSessionOrder(change);
return change;
}
private commit(
owner: string,
base: TabLayout,
next: TabLayout,
metadata: readonly TabRefMetadata[],
previous: TabLayout | null = base.version < 0 ? null : base
): TabLayout {
const stored = this.prepareCommit(base, next);
this.publish({ [owner]: stored }, [{ owner, previous, next: stored, metadata }]);
return stored;
}
private async prepareUnlocked(owner: string): Promise<PreparedOwnerLayout> {
const facts = await this.facts(owner);
const current = this.deps.store.getTabLayout(owner);
const authoritative = normalizeOrMigrateOwnerTabLayout({
owner,
layouts: current ? { [owner]: current } : undefined,
sessionOrder: this.deps.store.getSessionOrder(),
persistedSessions: facts.persisted,
liveSessions: facts.live,
webviews: facts.webviews,
updatedAt: (this.deps.now ?? (() => new Date().toISOString()))(),
}).layout;
return {
current,
authoritative,
metadata: facts.metadata,
needsReconciliationCommit: !current || !sameLayout(current, authoritative),
};
}
private async getUnlocked(owner: string): Promise<TabLayout> {
const prepared = await this.prepareUnlocked(owner);
if (!prepared.needsReconciliationCommit) {
const publication = {
owner,
previous: prepared.current,
next: prepared.authoritative,
metadata: prepared.metadata,
};
const latest = this.deps.store.getSessionOrder();
const projected = this.projectOrder(latest, [this.prepareOrderProjection(publication)]);
if (!sameOrder(normalizeSessionOrder(latest), projected)) this.publish({}, [publication]);
return prepared.authoritative;
}
const base = prepared.current ?? { ...prepared.authoritative, version: -1 };
return this.commit(owner, base, prepared.authoritative, prepared.metadata);
}
async get(owner: string): Promise<TabLayout> {
return this.withOwner(owner, () => this.getUnlocked(owner));
}
async put(owner: string, desired: unknown, baseVersion: number): Promise<TabLayoutPutResult> {
return this.withOwner(owner, async () => {
const prepared = await this.prepareUnlocked(owner);
if (baseVersion !== prepared.authoritative.version) return { status: 'conflict', layout: prepared.authoritative };
const validated = validateTabLayout(desired);
const owned = new Set(prepared.metadata.filter((item) => item.ownerValid && item.visible).map(refKey));
const refs = [...validated.groups.flatMap((group) => group.refs), ...validated.ungrouped];
const invalid = refs.find((ref) => !owned.has(refKey(ref)));
if (invalid)
throw new TabLayoutValidationError(`ref is not owned by layout owner: ${invalid.kind}:${invalid.id}`);
const normalized = normalizeTabLayout(
{ ...validated, version: prepared.authoritative.version },
prepared.metadata
);
return {
status: 'updated',
layout: this.commit(owner, prepared.authoritative, normalized, prepared.metadata, prepared.current),
};
});
}
async putLegacyOrder(actor: LegacyOrderActor, requested: readonly string[]): Promise<LegacyOrderPutResult> {
return actor.isAdmin ? this.putAdminLegacyOrder(requested) : this.putOwnerLegacyOrder(actor.owner, requested);
}
private async putOwnerLegacyOrder(owner: string, requested: readonly string[]): Promise<LegacyOrderPutResult> {
return this.withOwner(owner, async () => {
const prepared = await this.prepareUnlocked(owner);
const normalized = normalizeSessionOrder(requested);
const visible = new Set(
prepared.metadata
.filter((item) => item.kind === 'session' && item.ownerValid && item.visible)
.map((item) => item.id)
);
// Unknown or foreign ids are DROPPED, never a 400: the browser debounces
// its reorder push (and swallows errors), so a session deleted inside
// that window would otherwise cost the user the whole reorder — and the
// endpoint sits on the stable /api/v1 surface, where the pre-layout
// server merged leniently. Same philosophy as resolveParentSessionId.
const requestedVisible = normalized.filter((id) => visible.has(id));
const currentKnown = flattenOwnerSessionOrder(prepared.authoritative).filter((id) => visible.has(id));
const effective = mergeSessionOrder(requestedVisible, currentKnown);
const ranked = applyLegacySessionRank(prepared.authoritative, effective, prepared.metadata);
const needsLayout = prepared.needsReconciliationCommit || !sameLayout(prepared.authoritative, ranked);
const base = prepared.current ?? { ...prepared.authoritative, version: -1 };
const next = needsLayout ? this.prepareCommit(base, ranked) : prepared.authoritative;
const change = this.publish(needsLayout ? { [owner]: next } : {}, [
{ owner, previous: prepared.current, next, metadata: prepared.metadata },
]);
return { order: flattenOwnerSessionOrder(next).filter((id) => visible.has(id)), ...change };
});
}
private async putAdminLegacyOrder(requested: readonly string[]): Promise<LegacyOrderPutResult> {
const discoverOwners = (): string[] => {
const owners = new Set(Object.keys(this.deps.store.getTabLayouts()));
const { persisted, live } = this.sessionRecords();
for (const record of [...persisted, ...live]) owners.add(ownerOf(record));
return [...owners].sort();
};
for (;;) {
const owners = discoverOwners();
const result = await this.withOwners(owners, async (): Promise<LegacyOrderPutResult | null> => {
if (!sameOrder(owners, discoverOwners())) return null;
const normalized = normalizeSessionOrder(requested);
const knownOwners = new Map<string, string>();
const { persisted, live } = this.sessionRecords();
for (const record of persisted) knownOwners.set(record.id, ownerOf(record));
for (const record of live) knownOwners.set(record.id, ownerOf(record));
// Unknown ids are DROPPED, never a 400 — see putOwnerLegacyOrder. In
// single-user mode every request is the synthetic admin, so this path
// IS the one the browser's debounced (error-swallowing) push hits.
const known = normalized.filter((id) => knownOwners.has(id));
const publications: OwnerProjectionPublication[] = [];
const updates: Record<string, TabLayout> = Object.create(null) as Record<string, TabLayout>;
for (const owner of owners) {
const prepared = await this.prepareUnlocked(owner);
const visible = new Set(
prepared.metadata
.filter((item) => item.kind === 'session' && item.ownerValid && item.visible)
.map((item) => item.id)
);
const requestedOwner = known.filter((id) => visible.has(id));
const currentKnown = flattenOwnerSessionOrder(prepared.authoritative).filter((id) => visible.has(id));
const effective = mergeSessionOrder(requestedOwner, currentKnown);
const ranked = applyLegacySessionRank(prepared.authoritative, effective, prepared.metadata);
const needsLayout = prepared.needsReconciliationCommit || !sameLayout(prepared.authoritative, ranked);
const base = prepared.current ?? { ...prepared.authoritative, version: -1 };
const next = needsLayout ? this.prepareCommit(base, ranked) : prepared.authoritative;
if (needsLayout) updates[owner] = next;
publications.push({ owner, previous: prepared.current, next, metadata: prepared.metadata });
}
const change = this.publish(updates, publications, known);
return { order: [...change.globalOrder], ...change };
});
if (result) return result;
}
}
/** Reconcile one completed session creation into one versioned mutation. */
async sessionCreated(owner: string): Promise<TabLayout> {
return this.get(owner);
}
/** Reconcile one completed saved-webview creation into one versioned mutation. */
async webviewCreated(owner: string): Promise<TabLayout> {
return this.get(owner);
}
async sessionsRemoved(removed: readonly RemovedTabLayoutSession[]): Promise<void> {
if (this.restorationState !== 'complete' || removed.length === 0) return;
const byOwner = new Map<string, string[]>();
for (const item of removed) {
const owner = ownerOf(item);
const ids = byOwner.get(owner) ?? [];
ids.push(item.id);
byOwner.set(owner, ids);
}
const owners = [...byOwner.keys()].sort();
await this.withOwners(owners, async () => {
const publications: OwnerProjectionPublication[] = [];
const updates: Record<string, TabLayout> = Object.create(null) as Record<string, TabLayout>;
for (const owner of owners) {
const ids = byOwner.get(owner) ?? [];
const prepared = await this.prepareUnlocked(owner);
const current = prepared.current;
// Normalize and prune together so stale cleanup, orphan materialization,
// and missing-ref repair remain one versioned server mutation.
const next = normalizeTabLayout(
materializeOrphans(prepared.authoritative, ids, prepared.metadata),
prepared.metadata
);
const stored = current && !sameLayout(current, next) ? this.prepareCommit(current, next) : null;
if (stored) updates[owner] = stored;
publications.push({
owner,
previous: current,
next: stored ?? next,
metadata: prepared.metadata,
excludedSessionIds: new Set(ids),
});
}
if (publications.length > 0) this.publish(updates, publications);
});
}
/**
* Hold the owner mutation lock across an irreversible session deletion.
* All failure-prone normalization happens before `action`; the prepared layout
* commits only after the resource cleanup finishes.
*/
async runSessionDeletion<T>(removed: readonly RemovedTabLayoutSession[], action: () => Promise<T>): Promise<T> {
// A failed restoration must not lock the user out of explicitly closing a
// tab for the rest of the process lifetime: degrade to best-effort deletion
// without layout coordination. Only the AUTOMATED stale sweep stays
// fail-closed on 'failed' (runStaleSessionCleanup), because that one picks
// its victims itself from state a failed restore may have left incomplete.
if (this.restorationState === 'failed') return action();
this.assertDeletionReady();
if (this.restorationState === 'skipped' || removed.length === 0) return action();
const owners = new Set(removed.map(ownerOf));
if (owners.size !== 1) throw new Error('A session deletion transaction must contain exactly one owner');
const owner = owners.values().next().value as string;
const ids = removed.map((item) => item.id);
return this.withOwner(owner, async () => {
const prepared = await this.prepareUnlocked(owner);
const current = prepared.current;
// Prepare while the soon-to-be-deleted sessions are still known, so
// direct children can be materialized before their parent ref is removed.
const next = materializeOrphans(prepared.authoritative, ids, prepared.metadata);
const stored = current && !sameLayout(current, next) ? this.prepareCommit(current, next) : null;
const result = await action();
this.publish(stored ? { [owner]: stored } : {}, [
{
owner,
previous: current,
next: stored ?? next,
metadata: prepared.metadata,
excludedSessionIds: new Set(ids),
},
]);
return result;
});
}
/**
* Prepare every affected owner layout before bulk stale-state deletion.
* The StateStore action remains synchronous in production, so the candidate
* snapshot cannot change between successful preparation and resource removal.
*/
async runStaleSessionCleanup<T>(
activeSessionIds: ReadonlySet<string>,
action: (ids: ReadonlySet<string>) => T | Promise<T>
): Promise<T> {
this.assertDeletionReady();
const candidates = Object.entries(this.deps.store.getSessions())
.filter(([id, record]) => !activeSessionIds.has(id) && record.pinned !== true)
.map(([id, record]) => ({ id, owner: record.owner }));
if (this.restorationState === 'skipped') return action(new Set(candidates.map((item) => item.id)));
if (candidates.length === 0) return action(new Set());
const byOwner = new Map<string, string[]>();
for (const item of candidates) {
const owner = ownerOf(item);
const ids = byOwner.get(owner) ?? [];
ids.push(item.id);
byOwner.set(owner, ids);
}
const owners = [...byOwner.keys()].sort();
return this.withOwners(owners, async () => {
const webviews = await this.deps.readWebviews();
const persistedState = this.deps.store.getSessions();
const persisted = Object.entries(persistedState).map(([id, record]) => ({
id,
owner: record.owner,
createdAt: record.createdAt,
parentSessionId: record.parentSessionId,
}));
const liveIds = new Set(this.deps.sessions.keys());
const confirmed = candidates.filter((candidate) => {
const record = persistedState[candidate.id];
return (
record !== undefined &&
ownerOf(record) === ownerOf(candidate) &&
record.pinned !== true &&
!activeSessionIds.has(candidate.id) &&
!liveIds.has(candidate.id)
);
});
const confirmedByOwner = new Map<string, string[]>();
for (const item of confirmed) {
const owner = ownerOf(item);
const ids = confirmedByOwner.get(owner) ?? [];
ids.push(item.id);
confirmedByOwner.set(owner, ids);
}
const prepared: Array<{
owner: string;
current: TabLayout | null;
next: TabLayout;
stored: TabLayout | null;
metadata: TabRefMetadata[];
excludedSessionIds: ReadonlySet<string>;
}> = [];
for (const owner of owners) {
const ids = confirmedByOwner.get(owner) ?? [];
if (ids.length === 0) continue;
const current = this.deps.store.getTabLayout(owner);
const sessions = new Map<string, TabLayoutSessionRecord>();
for (const record of persisted) sessions.set(record.id, record);
for (const record of this.deps.sessions.values()) sessions.set(record.id, record);
const ownedSessions = [...sessions.values()]
.filter((record) => ownerOf(record) === owner)
.sort((a, b) => a.createdAt - b.createdAt || (a.id < b.id ? -1 : a.id > b.id ? 1 : 0));
const sessionOrder = new Map(ownedSessions.map((record, index) => [record.id, index]));
const metadata: TabRefMetadata[] = [...sessions.values()].map((record) => ({
kind: 'session',
id: record.id,
ownerValid: ownerOf(record) === owner,
visible: true,
order: sessionOrder.get(record.id) ?? record.createdAt,
parentSessionId: record.parentSessionId,
}));
const offset = ownedSessions.length;
webviews.forEach((record, index) =>
metadata.push({
kind: 'webview',
id: record.id,
ownerValid: ownerOf(record) === owner,
visible: true,
order: offset + index,
})
);
const authoritative = normalizeOrMigrateOwnerTabLayout({
owner,
layouts: current ? { [owner]: current } : undefined,
sessionOrder: this.deps.store.getSessionOrder(),
persistedSessions: persisted,
liveSessions: [...this.deps.sessions.values()],
webviews,
updatedAt: (this.deps.now ?? (() => new Date().toISOString()))(),
}).layout;
const next = materializeOrphans(authoritative, ids, metadata);
prepared.push({
owner,
current,
next,
stored: current && !sameLayout(current, next) ? this.prepareCommit(current, next) : null,
metadata,
excludedSessionIds: new Set(ids),
});
}
const result = await action(new Set(confirmed.map((item) => item.id)));
if (prepared.length > 0) {
this.publish(
Object.fromEntries(prepared.filter((item) => item.stored).map((item) => [item.owner, item.stored!])),
prepared.map((item) => ({
owner: item.owner,
previous: item.current,
next: item.stored ?? item.next,
metadata: item.metadata,
excludedSessionIds: item.excludedSessionIds,
}))
);
}
return result;
});
}
async webviewDeleted(owner: string, id: string): Promise<void> {
// Same explicit-user-action escape hatch as runSessionDeletion: a failed
// restore skips layout coordination instead of failing the delete.
if (this.restorationState === 'failed') return;
this.assertDeletionReady();
if (this.restorationState === 'skipped') return;
await this.withOwner(owner, async () => {
const current = this.deps.store.getTabLayout(owner);
if (!current) return;
const strip = (refs: readonly TabRef[]): TabRef[] =>
refs.filter((ref) => ref.kind !== 'webview' || ref.id !== id).map((ref) => ({ ...ref }));
const stripped: TabLayout = {
...current,
groups: current.groups.map((group) => ({ ...group, refs: strip(group.refs) })),
ungrouped: strip(current.ungrouped),
};
const { metadata } = await this.facts(owner);
const next = normalizeTabLayout(stripped, metadata);
if (!sameLayout(current, next)) this.commit(owner, current, next, metadata);
});
}
}
+547
View File
@@ -0,0 +1,547 @@
/**
* @fileoverview Framework-independent tab layout model.
*
* Callers provide owner-scoped session/webview metadata. This module deliberately
* has no dependency on session runtime, persistence, routes, or browser state.
*/
export const MAX_TAB_GROUPS = 32;
export const MAX_TAB_GROUP_NAME_LENGTH = 60;
export const MAX_TAB_REFS = 512;
export type TabRefKind = 'session' | 'webview';
export interface TabRef {
kind: TabRefKind;
id: string;
placement?: 'manual';
}
export interface TabGroup {
id: string;
name: string;
refs: TabRef[];
}
export interface TabLayout {
version: number;
groups: TabGroup[];
ungrouped: TabRef[];
updatedAt: string;
}
/** Owner and lineage facts supplied by the server or browser integration. */
export interface TabRefMetadata {
kind: TabRefKind;
id: string;
/** False for missing, foreign-owned, or otherwise invalid refs. */
ownerValid: boolean;
/** False when the owner is not permitted to see/store this ref. */
visible: boolean;
/** Stable creation/sibling order. Ties fall back to kind and id. */
order: number;
/** Session-only lineage hint. Ignored for webviews. */
parentSessionId?: string;
}
export interface TabMoveTarget {
/** Null denotes the real ungrouped container. */
groupId: string | null;
/** Zero-based insertion index after removing the moved block. */
index: number;
}
export interface CreateTabGroupInput {
id: string;
name: string;
index?: number;
}
export interface VisibleTabProjectionOptions {
liveSessionIds: ReadonlySet<string>;
openWebviewIds: ReadonlySet<string>;
collapsedGroupIds?: ReadonlySet<string>;
highlighted?: TabRef;
}
export class TabLayoutValidationError extends Error {
constructor(message: string) {
super(message);
this.name = 'TabLayoutValidationError';
}
}
const keyOf = (ref: Pick<TabRef, 'kind' | 'id'>): string => `${ref.kind}\u0000${ref.id}`;
function assertRecord(value: unknown, label: string): asserts value is Record<string, unknown> {
if (value === null || typeof value !== 'object' || Array.isArray(value)) {
throw new TabLayoutValidationError(`${label} must be an object`);
}
}
function parseNonEmptyString(value: unknown, label: string): string {
if (typeof value !== 'string' || value.length === 0) {
throw new TabLayoutValidationError(`${label} must be a non-empty string`);
}
return value;
}
function parseName(value: unknown, label: string): string {
if (typeof value !== 'string') throw new TabLayoutValidationError(`${label} must be a string`);
const trimmed = value.trim();
if (trimmed.length === 0 || trimmed.length > MAX_TAB_GROUP_NAME_LENGTH) {
throw new TabLayoutValidationError(`${label} must be 1-${MAX_TAB_GROUP_NAME_LENGTH} trimmed characters`);
}
return trimmed;
}
function parseRef(value: unknown, label: string): TabRef {
assertRecord(value, label);
if (value.kind !== 'session' && value.kind !== 'webview') {
throw new TabLayoutValidationError(`${label}.kind must be session or webview`);
}
const id = parseNonEmptyString(value.id, `${label}.id`);
if (value.placement !== undefined && value.placement !== 'manual') {
throw new TabLayoutValidationError(`${label}.placement must be manual when present`);
}
return value.placement === 'manual' ? { kind: value.kind, id, placement: 'manual' } : { kind: value.kind, id };
}
function parseTabLayout(input: unknown, repairDuplicates: boolean): TabLayout {
assertRecord(input, 'layout');
if (!Number.isSafeInteger(input.version) || (input.version as number) < 0) {
throw new TabLayoutValidationError('layout.version must be a non-negative safe integer');
}
if (!Array.isArray(input.groups)) throw new TabLayoutValidationError('layout.groups must be an array');
if (input.groups.length > MAX_TAB_GROUPS) {
throw new TabLayoutValidationError(`layout.groups cannot exceed ${MAX_TAB_GROUPS}`);
}
if (!Array.isArray(input.ungrouped)) throw new TabLayoutValidationError('layout.ungrouped must be an array');
const updatedAt = parseNonEmptyString(input.updatedAt, 'layout.updatedAt');
const groupIds = new Set<string>();
const refKeys = new Set<string>();
let refCount = input.ungrouped.length;
const parseStoredRef = (entry: unknown, label: string): TabRef => {
const ref = parseRef(entry, label);
const key = keyOf(ref);
if (!repairDuplicates && refKeys.has(key)) {
throw new TabLayoutValidationError(`duplicate ref: ${ref.kind}:${ref.id}`);
}
refKeys.add(key);
return ref;
};
const groups = input.groups.map((rawGroup, groupIndex): TabGroup => {
const label = `layout.groups[${groupIndex}]`;
assertRecord(rawGroup, label);
const id = parseNonEmptyString(rawGroup.id, `${label}.id`);
if (groupIds.has(id)) throw new TabLayoutValidationError(`duplicate group id: ${id}`);
groupIds.add(id);
if (!Array.isArray(rawGroup.refs)) throw new TabLayoutValidationError(`${label}.refs must be an array`);
refCount += rawGroup.refs.length;
return {
id,
name: parseName(rawGroup.name, `${label}.name`),
refs: rawGroup.refs.map((entry, refIndex) => parseStoredRef(entry, `${label}.refs[${refIndex}]`)),
};
});
if (refCount > MAX_TAB_REFS) {
throw new TabLayoutValidationError(`layout cannot contain more than ${MAX_TAB_REFS} refs`);
}
return {
version: input.version as number,
groups,
ungrouped: input.ungrouped.map((entry, index) => parseStoredRef(entry, `layout.ungrouped[${index}]`)),
updatedAt,
};
}
/** Validate and defensively clone a layout. Group names are normalized by trimming. */
export function validateTabLayout(input: unknown): TabLayout {
return parseTabLayout(input, false);
}
function validMetadata(metadata: readonly TabRefMetadata[]): TabRefMetadata[] {
const byKey = new Map<string, TabRefMetadata>();
for (const item of metadata) {
if ((item.kind !== 'session' && item.kind !== 'webview') || typeof item.id !== 'string' || item.id.length === 0) {
throw new TabLayoutValidationError('metadata contains an invalid ref identity');
}
if (!Number.isFinite(item.order)) throw new TabLayoutValidationError(`metadata order is invalid for ${item.id}`);
if (!item.ownerValid || !item.visible) continue;
const key = keyOf(item);
if (!byKey.has(key)) byKey.set(key, { ...item });
}
const compareText = (a: string, b: string): number => (a < b ? -1 : a > b ? 1 : 0);
const result = [...byKey.values()].sort(
(a, b) => a.order - b.order || compareText(a.kind, b.kind) || compareText(a.id, b.id)
);
if (result.length > MAX_TAB_REFS) {
throw new TabLayoutValidationError(`owner layout cannot exceed ${MAX_TAB_REFS} refs`);
}
return result;
}
interface LocatedRef {
ref: TabRef;
container: string | null;
position: number;
}
function locations(layout: TabLayout): LocatedRef[] {
const result: LocatedRef[] = [];
let position = 0;
for (const group of layout.groups) {
for (const ref of group.refs) result.push({ ref, container: group.id, position: position++ });
}
for (const ref of layout.ungrouped) result.push({ ref, container: null, position: position++ });
return result;
}
function withContainers(layout: TabLayout, refsByContainer: ReadonlyMap<string | null, TabRef[]>): TabLayout {
return {
...layout,
groups: layout.groups.map((group) => ({ ...group, refs: [...(refsByContainer.get(group.id) ?? [])] })),
ungrouped: [...(refsByContainer.get(null) ?? [])],
};
}
/**
* Reconcile a layout against owner-valid metadata and session lineage.
* First stored occurrence wins; missing valid refs append to ungrouped.
*/
export function normalizeTabLayout(input: TabLayout, metadata: readonly TabRefMetadata[]): TabLayout {
const layout = parseTabLayout(input, true);
const valid = validMetadata(metadata);
const metadataByKey = new Map(valid.map((item) => [keyOf(item), item]));
const knownMetadataKeys = new Set(metadata.map((item) => keyOf(item)));
const seen = new Set<string>();
const dedupedByContainer = new Map<string | null, TabRef[]>();
for (const group of layout.groups) dedupedByContainer.set(group.id, []);
dedupedByContainer.set(null, []);
for (const located of locations(layout)) {
const key = keyOf(located.ref);
// Missing metadata is unknown rather than invalid (for example, during
// restoration). Preserve it until an explicit invalid/deletion fact arrives.
if ((knownMetadataKeys.has(key) && !metadataByKey.has(key)) || seen.has(key)) continue;
seen.add(key);
dedupedByContainer.get(located.container)!.push({ ...located.ref });
}
for (const item of valid) {
const key = keyOf(item);
if (seen.has(key)) continue;
seen.add(key);
dedupedByContainer.get(null)!.push({ kind: item.kind, id: item.id });
}
if (seen.size > MAX_TAB_REFS) {
throw new TabLayoutValidationError(`normalized layout cannot exceed ${MAX_TAB_REFS} refs`);
}
let working = withContainers(layout, dedupedByContainer);
const located = locations(working);
const refByKey = new Map(located.map((item) => [keyOf(item.ref), item.ref]));
const sessionById = new Map(valid.filter((item) => item.kind === 'session').map((item) => [item.id, item]));
const manualCycleEdges = new Set<string>();
const state = new Map<string, 'visiting' | 'done'>();
const visit = (id: string): void => {
if (state.get(id) === 'done') return;
state.set(id, 'visiting');
const item = sessionById.get(id);
const stored = refByKey.get(keyOf({ kind: 'session', id }));
if (item?.parentSessionId && stored?.placement !== 'manual') {
const parent = sessionById.get(item.parentSessionId);
const parentStored = refByKey.get(keyOf({ kind: 'session', id: item.parentSessionId }));
if (parent && parentStored) {
if (state.get(parent.id) === 'visiting') manualCycleEdges.add(id);
else visit(parent.id);
}
}
state.set(id, 'done');
};
for (const item of located)
if (item.ref.kind === 'session' && state.get(item.ref.id) === undefined) visit(item.ref.id);
if (manualCycleEdges.size > 0) {
working = {
...working,
groups: working.groups.map((group) => ({
...group,
refs: group.refs.map((ref) =>
ref.kind === 'session' && manualCycleEdges.has(ref.id) ? { ...ref, placement: 'manual' } : ref
),
})),
ungrouped: working.ungrouped.map((ref) =>
ref.kind === 'session' && manualCycleEdges.has(ref.id) ? { ...ref, placement: 'manual' } : ref
),
};
}
const ordered = locations(working);
const updatedRefByKey = new Map(ordered.map((item) => [keyOf(item.ref), item.ref]));
const parentOf = new Map<string, string>();
const children = new Map<string, string[]>();
for (const item of ordered) {
if (item.ref.kind !== 'session' || item.ref.placement === 'manual') continue;
const info = sessionById.get(item.ref.id);
const parentId = info?.parentSessionId;
if (!parentId || !sessionById.has(parentId) || !updatedRefByKey.has(keyOf({ kind: 'session', id: parentId })))
continue;
parentOf.set(item.ref.id, parentId);
const siblings = children.get(parentId) ?? [];
siblings.push(item.ref.id);
children.set(parentId, siblings);
}
const emitted = new Set<string>();
const output = new Map<string | null, TabRef[]>();
for (const group of working.groups) output.set(group.id, []);
output.set(null, []);
const emitSubtree = (root: TabRef, container: string | null): void => {
const rootKey = keyOf(root);
if (emitted.has(rootKey)) return;
emitted.add(rootKey);
output.get(container)!.push({ ...root });
if (root.kind !== 'session') return;
for (const childId of children.get(root.id) ?? []) {
const child = updatedRefByKey.get(keyOf({ kind: 'session', id: childId }));
if (child) emitSubtree(child, container);
}
};
for (const item of ordered) {
if (item.ref.kind === 'session' && parentOf.has(item.ref.id)) continue;
emitSubtree(item.ref, item.container);
}
return withContainers(working, output);
}
function cloneForEdit(input: TabLayout): TabLayout {
return validateTabLayout(input);
}
function boundedIndex(index: number, length: number, label: string): number {
if (!Number.isSafeInteger(index) || index < 0 || index > length) {
throw new TabLayoutValidationError(`${label} index must be between 0 and ${length}`);
}
return index;
}
export function createGroup(input: TabLayout, group: CreateTabGroupInput): TabLayout {
const layout = cloneForEdit(input);
if (layout.groups.length >= MAX_TAB_GROUPS)
throw new TabLayoutValidationError(`cannot exceed ${MAX_TAB_GROUPS} groups`);
const id = parseNonEmptyString(group.id, 'group.id');
if (layout.groups.some((entry) => entry.id === id)) throw new TabLayoutValidationError(`duplicate group id: ${id}`);
const index = boundedIndex(group.index ?? layout.groups.length, layout.groups.length, 'group');
const groups = [...layout.groups];
groups.splice(index, 0, { id, name: parseName(group.name, 'group.name'), refs: [] });
return { ...layout, groups };
}
export function renameGroup(input: TabLayout, groupId: string, name: string): TabLayout {
const layout = cloneForEdit(input);
if (!layout.groups.some((group) => group.id === groupId))
throw new TabLayoutValidationError(`unknown group: ${groupId}`);
return {
...layout,
groups: layout.groups.map((group) =>
group.id === groupId ? { ...group, name: parseName(name, 'group.name') } : group
),
};
}
export function deleteGroup(input: TabLayout, groupId: string): TabLayout {
const layout = cloneForEdit(input);
const group = layout.groups.find((entry) => entry.id === groupId);
if (!group) throw new TabLayoutValidationError(`unknown group: ${groupId}`);
return {
...layout,
groups: layout.groups.filter((entry) => entry.id !== groupId),
ungrouped: [...layout.ungrouped, ...group.refs.map((ref) => ({ ...ref }))],
};
}
export function reorderGroup(input: TabLayout, groupId: string, index: number): TabLayout {
const layout = cloneForEdit(input);
const from = layout.groups.findIndex((group) => group.id === groupId);
if (from < 0) throw new TabLayoutValidationError(`unknown group: ${groupId}`);
const groups = [...layout.groups];
const [group] = groups.splice(from, 1);
groups.splice(boundedIndex(index, groups.length, 'group'), 0, group);
return { ...layout, groups };
}
function mapRef(input: TabLayout, target: TabRef, transform: (ref: TabRef) => TabRef): TabLayout {
const layout = cloneForEdit(input);
let found = false;
const apply = (ref: TabRef): TabRef => {
if (keyOf(ref) !== keyOf(target)) return ref;
found = true;
return transform(ref);
};
const result = {
...layout,
groups: layout.groups.map((group) => ({ ...group, refs: group.refs.map(apply) })),
ungrouped: layout.ungrouped.map(apply),
};
if (!found) throw new TabLayoutValidationError(`unknown ref: ${target.kind}:${target.id}`);
return result;
}
export function setManualPlacement(input: TabLayout, target: TabRef, manual: boolean): TabLayout {
if (!manual) {
throw new TabLayoutValidationError('manual placement can only be cleared through followParent');
}
return mapRef(input, target, (ref) => ({ ...ref, placement: 'manual' }));
}
export function followParent(input: TabLayout, target: TabRef, metadata: readonly TabRefMetadata[]): TabLayout {
const normalized = normalizeTabLayout(input, metadata);
if (target.kind !== 'session') {
throw new TabLayoutValidationError('only a session ref can follow a parent');
}
const valid = validMetadata(metadata);
const targetMetadata = valid.find((item) => item.kind === 'session' && item.id === target.id);
if (!targetMetadata?.parentSessionId) {
throw new TabLayoutValidationError(`session has no owner-valid parent: ${target.id}`);
}
const parentMetadata = valid.find((item) => item.kind === 'session' && item.id === targetMetadata.parentSessionId);
if (!parentMetadata) {
throw new TabLayoutValidationError(`session parent is not owner-valid: ${targetMetadata.parentSessionId}`);
}
const storedKeys = new Set(locations(normalized).map((item) => keyOf(item.ref)));
if (!storedKeys.has(keyOf(target))) {
throw new TabLayoutValidationError(`unknown ref: ${target.kind}:${target.id}`);
}
const parentRef: TabRef = { kind: 'session', id: targetMetadata.parentSessionId };
if (!storedKeys.has(keyOf(parentRef))) {
throw new TabLayoutValidationError(`session parent is not represented: ${targetMetadata.parentSessionId}`);
}
const cleared = mapRef(normalized, target, (ref) => ({ kind: ref.kind, id: ref.id }));
return normalizeTabLayout(cleared, metadata);
}
function descendantKeys(root: TabRef, layout: TabLayout, metadata: readonly TabRefMetadata[]): Set<string> {
const valid = validMetadata(metadata);
const stored = new Map(locations(layout).map((item) => [keyOf(item.ref), item.ref]));
const children = new Map<string, string[]>();
for (const item of valid) {
if (item.kind !== 'session' || !item.parentSessionId) continue;
const child = stored.get(keyOf(item));
if (!child || child.placement === 'manual' || !stored.has(keyOf({ kind: 'session', id: item.parentSessionId })))
continue;
const siblings = children.get(item.parentSessionId) ?? [];
siblings.push(item.id);
children.set(item.parentSessionId, siblings);
}
const result = new Set<string>();
const add = (ref: TabRef): void => {
const key = keyOf(ref);
if (result.has(key)) return;
result.add(key);
if (ref.kind !== 'session') return;
for (const childId of children.get(ref.id) ?? []) add({ kind: 'session', id: childId });
};
add(root);
return result;
}
export function moveRef(
input: TabLayout,
target: TabRef,
destination: TabMoveTarget,
metadata: readonly TabRefMetadata[]
): TabLayout {
let layout = normalizeTabLayout(input, metadata);
const targetKey = keyOf(target);
if (!locations(layout).some((item) => keyOf(item.ref) === targetKey)) {
throw new TabLayoutValidationError(`unknown ref: ${target.kind}:${target.id}`);
}
if (destination.groupId !== null && !layout.groups.some((group) => group.id === destination.groupId)) {
throw new TabLayoutValidationError(`unknown group: ${destination.groupId}`);
}
const blockKeys = descendantKeys(target, layout, metadata);
const block = locations(layout)
.filter((item) => blockKeys.has(keyOf(item.ref)))
.map((item) => ({ ...item.ref }));
const metadataItem = validMetadata(metadata).find((item) => keyOf(item) === targetKey);
if (target.kind === 'session' && metadataItem?.parentSessionId) block[0] = { ...block[0], placement: 'manual' };
const remaining = new Map<string | null, TabRef[]>();
for (const group of layout.groups)
remaining.set(
group.id,
group.refs.filter((ref) => !blockKeys.has(keyOf(ref)))
);
remaining.set(
null,
layout.ungrouped.filter((ref) => !blockKeys.has(keyOf(ref)))
);
const destinationRefs = remaining.get(destination.groupId)!;
const index = boundedIndex(destination.index, destinationRefs.length, 'destination');
destinationRefs.splice(index, 0, ...block);
layout = withContainers(layout, remaining);
return normalizeTabLayout(layout, metadata);
}
/**
* Remove explicitly deleted session parents and pin their direct inherited
* children at their current stored positions so a later reused ID cannot adopt them.
*/
export function materializeOrphans(
input: TabLayout,
removedParentIds: readonly string[],
metadata: readonly TabRefMetadata[]
): TabLayout {
const layout = cloneForEdit(input);
const removed = new Set(removedParentIds);
const directChildren = new Set(
validMetadata(metadata)
.filter((item) => item.kind === 'session' && item.parentSessionId && removed.has(item.parentSessionId))
.map((item) => item.id)
);
const transform = (refs: readonly TabRef[]): TabRef[] =>
refs
.filter((ref) => ref.kind !== 'session' || !removed.has(ref.id))
.map((ref) =>
ref.kind === 'session' && directChildren.has(ref.id) && ref.placement !== 'manual'
? { ...ref, placement: 'manual' }
: { ...ref }
);
return {
...layout,
groups: layout.groups.map((group) => ({ ...group, refs: transform(group.refs) })),
ungrouped: transform(layout.ungrouped),
};
}
/** Session-only compatibility order; collapse and webviews do not affect it. */
export function flattenOwnerSessionOrder(input: TabLayout): string[] {
return locations(validateTabLayout(input))
.map((item) => item.ref)
.filter((ref): ref is TabRef & { kind: 'session' } => ref.kind === 'session')
.map((ref) => ref.id);
}
/** Locally renderable order used by tab painting and Alt-number consumers. */
export function flattenVisibleRefs(input: TabLayout, options: VisibleTabProjectionOptions): TabRef[] {
const layout = validateTabLayout(input);
const collapsed = options.collapsedGroupIds ?? new Set<string>();
const renderable = (ref: TabRef): boolean =>
ref.kind === 'session' ? options.liveSessionIds.has(ref.id) : options.openWebviewIds.has(ref.id);
const highlightedKey = options.highlighted ? keyOf(options.highlighted) : undefined;
const result: TabRef[] = [];
for (const group of layout.groups) {
for (const ref of group.refs) {
if (!renderable(ref)) continue;
if (collapsed.has(group.id) && keyOf(ref) !== highlightedKey) continue;
result.push({ ...ref });
}
}
for (const ref of layout.ungrouped) if (renderable(ref)) result.push({ ...ref });
return result;
}
+384 -486
View File
File diff suppressed because it is too large Load Diff
+598
View File
@@ -0,0 +1,598 @@
/**
* @fileoverview Pure ANSI helpers for the TUI preview pane.
*
* The preview shows the tail of a session's raw terminal stream, which is
* xterm-bound bytes: SGR colors, cursor jumps, OSC titles, DECSET modes and
* carriage-return repaints. This is NOT a terminal emulator. It reconstructs a
* readable, color-preserving tail: SGR survives, everything else that steers a
* cursor is dropped, and a `\r` is honored as "back to column 0" so a spinner
* that repaints its line 200 times contributes one line instead of 200.
*
* CURSOR ADDRESSING (`ESC [ r ; c H`) is honored too, and it has to be: an Ink
* TUI like Claude Code repaints by ROW and emits almost no newlines, so
* dropping those sequences collapses a whole screen into one unreadable line
* (measured against a live pane, 2026-08-16). A jump to column 1 starts a new
* display line, a jump within a row moves the write position, which is the same
* reading `normalizeCapturedFrame` in `web/approval-inbox.ts` takes of the same
* kind of frame.
*
* Two approximations are deliberate, because the alternative is an emulator:
* a carriage-return overwrite counts CODE POINTS, not display columns (so a
* repaint over CJK text can land one cell off), and tab stops are counted the
* same way. Neither can corrupt output, they only shift a repaint's alignment.
* Absolute ROW numbers are ignored as well: rows arrive in the order they are
* painted, which for a tail is the order worth reading.
*
* @module tui/tui-ansi
*/
const ESC = 0x1b;
const BEL = 0x07;
const ST_C1 = 0x9c;
const DEL = 0x7f;
/** SGR reset, appended by `clipStyledLine` so a clipped line cannot bleed. */
export const SGR_RESET = '\x1b[0m';
const TAB_WIDTH = 8;
/** Cap on remembered SGR sequences per cell, so a pathological stream cannot grow one unboundedly. */
const MAX_ACTIVE_SGR = 32;
/** Ceiling on a display line's cells: a stream may address column 99999, a terminal has none. */
const MAX_LINE_CELLS = 1000;
// ─────────────────────────────────────────────────────────────────────────────
// Escape-sequence scanning
// ─────────────────────────────────────────────────────────────────────────────
interface EscapeScan {
/** Index just past the sequence; `text.length` for a truncated one. */
next: number;
/** The sequence itself, only when it is SGR (`CSI ... m`) and therefore kept. */
sgr?: string;
/** 1-based column of a cursor-position sequence (`CSI r ; c H` or `f`). */
column?: number;
/** 1-based row of that same sequence. Row 1 means a repaint is starting. */
row?: number;
}
/** The row and column a `CSI r ; c H` addresses. Both parameters default to 1. */
function cursorPosition(params: string): { row: number; column: number } {
const parts = params.split(';');
const read = (index: number): number => {
const value = Number.parseInt(parts[index] ?? '', 10);
return Number.isSafeInteger(value) && value > 0 ? value : 1;
};
return { row: read(0), column: read(1) };
}
/** Scan a CSI body starting at `from` (params, then intermediates, then a final byte). */
function readCsi(text: string, start: number, from: number, keepSgr: boolean): EscapeScan {
let j = from;
while (j < text.length && text.charCodeAt(j) >= 0x30 && text.charCodeAt(j) <= 0x3f) j++;
while (j < text.length && text.charCodeAt(j) >= 0x20 && text.charCodeAt(j) <= 0x2f) j++;
if (j >= text.length) return { next: text.length };
const next = j + 1;
if (keepSgr && text[j] === 'm') return { next, sgr: text.slice(start, next) };
if (keepSgr && (text[j] === 'H' || text[j] === 'f')) {
return { next, ...cursorPosition(text.slice(from, j)) };
}
return { next };
}
/** Scan an OSC/DCS/PM/APC body: everything up to BEL, C1 ST or `ESC \`. */
function readStringSequence(text: string, from: number): number {
let j = from;
while (j < text.length) {
const code = text.charCodeAt(j);
if (code === BEL || code === ST_C1) return j + 1;
if (code === ESC && text[j + 1] === '\\') return j + 2;
j++;
}
return text.length;
}
/** Scan the escape sequence starting at `i` (which must be an ESC). */
function readEscape(text: string, i: number): EscapeScan {
const second = text[i + 1];
if (second === undefined) return { next: text.length };
if (second === '[') return readCsi(text, i, i + 2, true);
if (second === ']' || second === 'P' || second === 'X' || second === '^' || second === '_') {
return { next: readStringSequence(text, i + 2) };
}
// Charset / character-set selection: one more byte belongs to the sequence.
if (second === '(' || second === ')' || second === '*' || second === '+' || second === '#' || second === '%') {
return { next: Math.min(text.length, i + 3) };
}
return { next: i + 2 };
}
/** Scan a single-byte C1 control at `i` (0x80-0x9f). */
function readC1(text: string, i: number): number {
const code = text.charCodeAt(i);
if (code === 0x9b) return readCsi(text, i, i + 1, false).next;
if (code === 0x90 || code === 0x9d || code === 0x9e || code === 0x9f) return readStringSequence(text, i + 1);
return i + 1;
}
function isC1(code: number): boolean {
return code >= 0x80 && code <= 0x9f;
}
/** `CSI 0 m`, `CSI m` and `CSI 0;0 m` all mean "back to plain". */
function isSgrReset(seq: string): boolean {
const params = seq.slice(2, -1);
return params === '' || /^0(?:;0)*$/.test(params);
}
/**
* Fold one SGR sequence into the active set. Sequences accumulate in arrival
* order (a later color simply wins when replayed), a reset clears them, and a
* repeat moves rather than duplicates.
*/
function applySgr(active: string[], seq: string): string[] {
if (isSgrReset(seq)) return [];
const next = active.filter((s) => s !== seq);
next.push(seq);
return next.length > MAX_ACTIVE_SGR ? next.slice(-MAX_ACTIVE_SGR) : next;
}
// ─────────────────────────────────────────────────────────────────────────────
// Display width
// ─────────────────────────────────────────────────────────────────────────────
/**
* Combining marks, variation selectors and other zero-advance code points.
* Pragmatic, not exhaustive: enough that accents and emoji modifiers do not
* inflate a measured width.
*/
const ZERO_WIDTH_RANGES: ReadonlyArray<readonly [number, number]> = [
[0x0300, 0x036f],
[0x0483, 0x0489],
[0x0591, 0x05bd],
[0x05bf, 0x05bf],
[0x0610, 0x061a],
[0x064b, 0x065f],
[0x0670, 0x0670],
[0x06d6, 0x06dc],
[0x0e31, 0x0e31],
[0x0e34, 0x0e3a],
[0x0e47, 0x0e4e],
[0x200b, 0x200f],
[0x2028, 0x202e],
[0x2060, 0x2064],
[0x20d0, 0x20f0],
[0xfe00, 0xfe0f],
[0xfe20, 0xfe2f],
[0xfeff, 0xfeff],
];
/**
* East Asian Wide + Fullwidth, plus the standalone code points UAX #11 marks
* Wide because they are emoji-presentation by default. This repo ships a zh-CN
* locale, so CJK correctness is the point; exhaustive Unicode is not required,
* but the scattered BMP entries below are not optional either: `✋` (U+270B) is
* one of them and it is a glyph this TUI draws in every waiting row, so getting
* it wrong mis-pads a column on every frame.
*/
const WIDE_RANGES: ReadonlyArray<readonly [number, number]> = [
[0x1100, 0x115f],
[0x231a, 0x231b],
[0x23e9, 0x23ec],
[0x23f0, 0x23f0],
[0x23f3, 0x23f3],
[0x25fd, 0x25fe],
[0x2614, 0x2615],
[0x2648, 0x2653],
[0x267f, 0x267f],
[0x2693, 0x2693],
[0x26a1, 0x26a1],
[0x26aa, 0x26ab],
[0x26bd, 0x26be],
[0x26c4, 0x26c5],
[0x26ce, 0x26ce],
[0x26d4, 0x26d4],
[0x26ea, 0x26ea],
[0x26f2, 0x26f3],
[0x26f5, 0x26f5],
[0x26fa, 0x26fa],
[0x26fd, 0x26fd],
[0x2705, 0x2705],
[0x270a, 0x270b],
[0x2728, 0x2728],
[0x274c, 0x274c],
[0x274e, 0x274e],
[0x2753, 0x2755],
[0x2757, 0x2757],
[0x2795, 0x2797],
[0x27b0, 0x27b0],
[0x27bf, 0x27bf],
[0x2b1b, 0x2b1c],
[0x2b50, 0x2b50],
[0x2b55, 0x2b55],
[0x2e80, 0x303e],
[0x3041, 0x33ff],
[0x3400, 0x4dbf],
[0x4e00, 0x9fff],
[0xa000, 0xa4cf],
[0xa960, 0xa97f],
[0xac00, 0xd7a3],
[0xf900, 0xfaff],
[0xfe10, 0xfe19],
[0xfe30, 0xfe6f],
[0xff00, 0xff60],
[0xffe0, 0xffe6],
[0x1f004, 0x1f004],
[0x1f0cf, 0x1f0cf],
[0x1f18e, 0x1f18e],
[0x1f191, 0x1f19a],
[0x1f200, 0x1f320],
[0x1f32d, 0x1f335],
[0x1f337, 0x1f37c],
[0x1f37e, 0x1f393],
[0x1f3a0, 0x1f3ca],
[0x1f3cf, 0x1f3d3],
[0x1f3e0, 0x1f3f0],
[0x1f3f4, 0x1f3f4],
[0x1f3f8, 0x1f43e],
[0x1f440, 0x1f440],
[0x1f442, 0x1f4fc],
[0x1f4ff, 0x1f53d],
[0x1f54b, 0x1f54e],
[0x1f550, 0x1f567],
[0x1f57a, 0x1f57a],
[0x1f595, 0x1f596],
[0x1f5a4, 0x1f5a4],
[0x1f5fb, 0x1f64f],
[0x1f680, 0x1f6c5],
[0x1f6cc, 0x1f6cc],
[0x1f6d0, 0x1f6d2],
[0x1f6eb, 0x1f6ec],
[0x1f6f4, 0x1f6fc],
[0x1f7e0, 0x1f7eb],
[0x1f90c, 0x1f93a],
[0x1f93c, 0x1f945],
[0x1f947, 0x1f9ff],
[0x1fa70, 0x1faff],
[0x20000, 0x2fffd],
[0x30000, 0x3fffd],
];
function inRanges(cp: number, ranges: ReadonlyArray<readonly [number, number]>): boolean {
for (const [lo, hi] of ranges) {
if (cp < lo) return false;
if (cp <= hi) return true;
}
return false;
}
/** Columns one code point advances the cursor by: 0, 1 or 2. */
export function charWidth(codePoint: number): number {
if (codePoint < 0x20 || (codePoint >= DEL && codePoint <= 0x9f)) return 0;
if (inRanges(codePoint, ZERO_WIDTH_RANGES)) return 0;
if (inRanges(codePoint, WIDE_RANGES)) return 2;
return 1;
}
/** Display width of a string: escape sequences take no columns, CJK takes two. */
export function visibleWidth(text: string): number {
let width = 0;
let i = 0;
while (i < text.length) {
const code = text.charCodeAt(i);
if (code === ESC) {
i = readEscape(text, i).next;
continue;
}
if (isC1(code)) {
i = readC1(text, i);
continue;
}
if (code < 0x20 || code === DEL) {
i++;
continue;
}
const cp = text.codePointAt(i) as number;
i += cp > 0xffff ? 2 : 1;
width += charWidth(cp);
}
return width;
}
// ─────────────────────────────────────────────────────────────────────────────
// Raw stream to display lines
// ─────────────────────────────────────────────────────────────────────────────
/** One printed code point (plus any combining marks) and the SGR state under it. */
interface Cell {
text: string;
sgr: string;
}
/**
* Replay cells into a string, emitting an SGR change only where the state
* actually changes and closing the line so it is self-contained.
*/
function renderCells(cells: Cell[]): string {
let out = '';
let active = '';
for (const cell of cells) {
if (cell.sgr !== active) {
if (active !== '') out += SGR_RESET;
out += cell.sgr;
active = cell.sgr;
}
out += cell.text;
}
if (active !== '') out += SGR_RESET;
return out;
}
/**
* Turn a raw terminal stream into display lines: SGR preserved, every other
* escape sequence dropped, `\r` treated as a return to column 0 (the following
* text overwrites what is there), tabs expanded, other control characters
* dropped.
*
* Splitting matches `String.split('\n')`, so `''` yields `['']` and a trailing
* newline yields a trailing empty line.
*/
/**
* Glyphs a CLI draws as chrome that a plain terminal font very often has no
* coverage for, and the ASCII that means the same thing.
*
* ⚠️ This is NOT a substitute for the glyph TIER. The tier answers "can this
* terminal do Unicode at all", which is a locale question, and it says yes for
* exactly the terminals this table exists for: a beta tester's font rendered
* `·`, `─`, `│` and `▶` perfectly while drawing claude's `❯` prompt and its
* `⏵⏵` mode marker as empty boxes. Coverage is per-glyph and undetectable from
* here, so the rare ones are folded and the common ones are left alone.
*
* Kept deliberately SHORT. Every entry is a glyph seen rendering as tofu in a
* real terminal, not a guess, and each maps to the arrow it already looks like.
*/
const PREVIEW_GLYPH_FOLD: ReadonlyMap<string, string> = new Map([
['\u276F', '>'], // ❯ heavy right-pointing angle quotation mark (claude, starship, zsh prompts)
['\u276E', '<'], // ❮
['\u23F5', '>'], // ⏵ black medium right-pointing triangle (claude's bypass-permissions marker)
['\u23F4', '<'], // ⏴
['\u23F6', '^'], // ⏶
['\u23F7', 'v'], // ⏷
['\u2771', '>'], // ❱
['\u2770', '<'], // ❰
// claude's own working/done spinner cycles through these, and they are the
// same sparse-Dingbats class as `❯`: the animated line is exactly where a
// reader looks, so tofu there is the most visible kind.
['\u2722', '*'], // ✢
['\u2733', '*'], // ✳
['\u2217', '*'], // ∗
['\u273B', '*'], // ✻
['\u273D', '*'], // ✽
['\u2734', '*'], // ✴
['\u26A0', '!'], // ⚠ Misc Symbols, and emoji-presentation on many terminals
]);
/**
* Replace preview glyphs a plain font is likely to draw as an empty box.
*
* Applied to ANOTHER program's output on its way into the preview pane, never
* to the TUI's own chrome, and skipped at the `nerd` tier where the user has
* declared a font that can draw anything.
*/
export function foldPreviewGlyphs(line: string): string {
let out = '';
for (const char of line) out += PREVIEW_GLYPH_FOLD.get(char) ?? char;
return out;
}
export function toDisplayLines(raw: string): string[] {
const lines: string[] = [];
let cells: Cell[] = [];
let col = 0;
let active: string[] = [];
let sgr = '';
const endLine = (): void => {
lines.push(renderCells(cells));
cells = [];
col = 0;
};
/**
* Park the write position at a column, padding the gap so the cell array
* never grows a hole (a hole would crash the replay, and a stream can address
* any column it likes).
*/
const moveTo = (column: number): void => {
const target = Math.min(column, MAX_LINE_CELLS);
while (cells.length < target) cells.push({ text: ' ', sgr: '' });
col = target;
};
const write = (text: string, width: number): void => {
if (width === 0) {
// A combining mark belongs to the character it follows, never to a cell
// of its own: keeping them together is what stops a clip from severing
// an accent from its base letter.
if (col > 0) cells[col - 1].text += text;
return;
}
cells[col] = { text, sgr };
col++;
};
let i = 0;
while (i < raw.length) {
const code = raw.charCodeAt(i);
if (code === ESC) {
const scan = readEscape(raw, i);
if (scan.sgr !== undefined) {
active = applySgr(active, scan.sgr);
sgr = active.join('');
} else if (scan.row === 1 && scan.column === 1) {
// ⚠️ A HOME is a full-screen app announcing that it is repainting from
// the top, and everything already on screen is about to be overwritten
// in place. This replay is line-based and cannot overwrite, so the
// faithful equivalent is to start over — without it every repaint was
// APPENDED, and a claude pane's tail carried fifty stacked copies of
// the same frame. The preview then showed the last N lines, which on a
// tall terminal spanned two of them (reported from the beta as the
// overview showing the session twice).
lines.length = 0;
cells = [];
col = 0;
} else if (scan.column !== undefined) {
// Column 1 is a fresh row, which is the only thing a repainting TUI
// gives us to split lines on.
if (scan.column <= 1) endLine();
else moveTo(scan.column - 1);
}
i = scan.next;
continue;
}
if (isC1(code)) {
i = readC1(raw, i);
continue;
}
if (code === 0x0a) {
endLine();
i++;
continue;
}
if (code === 0x0d) {
col = 0;
i++;
continue;
}
if (code === 0x09) {
const stop = TAB_WIDTH - (col % TAB_WIDTH);
for (let n = 0; n < stop; n++) write(' ', 1);
i++;
continue;
}
if (code < 0x20 || code === DEL) {
i++;
continue;
}
const cp = raw.codePointAt(i) as number;
const text = String.fromCodePoint(cp);
i += text.length;
write(text, charWidth(cp));
}
endLine();
return lines;
}
/**
* The parameter bytes plus final byte of a CSI sequence whose `ESC [` was cut
* off. Requires at least one parameter byte, so ordinary text starting with a
* letter is never mistaken for one.
*/
const SEVERED_CSI = /^[0-9;?:<>=]+[A-Za-z]/;
/**
* Drop the remains of an escape sequence a byte-sliced tail begins in the
* middle of.
*
* `GET /api/sessions/:id/terminal?tail=N` cuts the buffer at a byte offset, so
* a tail can start inside `ESC [ 12 ; 1 H` and hand the parser `;1H` as text,
* which is exactly what it then prints (observed against a live Claude pane).
* Only the severed head is dropped, never a whole line.
*/
export function dropSeveredEscape(raw: string): string {
return raw.replace(SEVERED_CSI, '');
}
/**
* Drop every escape sequence, keeping the visible text. Needed because the
* preview carries the session's OWN colors: under NO_COLOR the frame must not
* smuggle them back in.
*/
export function stripStyles(text: string): string {
let out = '';
let i = 0;
while (i < text.length) {
const code = text.charCodeAt(i);
if (code === ESC) {
i = readEscape(text, i).next;
continue;
}
if (isC1(code)) {
i = readC1(text, i);
continue;
}
if (code < 0x20 || code === DEL) {
i++;
continue;
}
const cp = text.codePointAt(i) as number;
const size = cp > 0xffff ? 2 : 1;
out += text.slice(i, i + size);
i += size;
}
return out;
}
// ─────────────────────────────────────────────────────────────────────────────
// Clipping and padding
// ─────────────────────────────────────────────────────────────────────────────
/**
* Clip a line that carries SGR to `width` display columns, keeping the styling
* that is active up to the clip point and closing it with a reset. Never splits
* a code point, a combining sequence or an escape sequence, and never emits
* half of a double-width character (the cell is dropped instead).
*/
export function clipStyledLine(line: string, width: number): string {
if (width <= 0) return '';
let out = '';
let used = 0;
let active: string[] = [];
// Styles are emitted lazily, right before the character that wears them, so a
// sequence sitting exactly on the clip boundary is not carried into a line it
// no longer styles.
let emitted = '';
let i = 0;
while (i < line.length) {
const code = line.charCodeAt(i);
if (code === ESC) {
const scan = readEscape(line, i);
if (scan.sgr !== undefined) active = applySgr(active, scan.sgr);
i = scan.next;
continue;
}
if (isC1(code)) {
i = readC1(line, i);
continue;
}
if (code < 0x20 || code === DEL) {
i++;
continue;
}
const cp = line.codePointAt(i) as number;
const w = charWidth(cp);
if (used + w > width) break;
const style = active.join('');
if (style !== emitted) {
if (emitted !== '') out += SGR_RESET;
out += style;
emitted = style;
}
out += String.fromCodePoint(cp);
used += w;
i += cp > 0xffff ? 2 : 1;
}
return emitted !== '' ? out + SGR_RESET : out;
}
/**
* Pad or clip to exactly `width` display columns. A clip that lands on a
* double-width boundary leaves one column short, so the pad runs after it.
*/
export function padDisplay(text: string, width: number): string {
if (width <= 0) return '';
const w = visibleWidth(text);
if (w === width) return text;
if (w < width) return text + ' '.repeat(width - w);
const clipped = clipStyledLine(text, width);
return clipped + ' '.repeat(Math.max(0, width - visibleWidth(clipped)));
}
+2697
View File
File diff suppressed because it is too large Load Diff
+137
View File
@@ -0,0 +1,137 @@
/**
* @fileoverview Pure reading of an approvals-inbox item: what the card says,
* which keys are live for it, and which of them just appeared.
*
* This is the half of "answer the dialog from the dashboard" that can be stated
* as a function of the item. The IO half (`POST /api/approvals/:id/answer`)
* lives in `tui-client.ts`, and the server re-captures the pane before it aims
* any keystroke, so a card that went stale is refused rather than mis-answered.
*
* The key matrix is deliberately narrow, because the alternative is typing a
* digit into whatever now has focus:
*
* | kind | y | n | 1-9 |
* | ---------- | ------------ | ---------------------- | ------------------------- |
* | permission | approve | the parsed "No" option, | only digits the server |
* | question | approve | else Esc | actually parsed off screen |
* | idle | not a dialog: `p` (the composer) is the reply path |
*
* A digit that is not among the parsed options returns null, which is what lets
* the caller fall back to the list's own 1-9 jump instead of sending a keystroke
* the dialog has no answer for.
*
* PURE: no IO, no timers, no `process.*`.
*
* @module tui/tui-approvals
*/
import type { ApprovalItem, ApprovalOption } from '../web/approval-inbox.js';
import type { TuiApprovalAnswer } from './tui-client.js';
/** Card severity, in the same red/yellow vocabulary the web inbox uses. */
export type TuiApprovalTone = 'err' | 'warn';
export interface TuiApprovalCard {
tone: TuiApprovalTone;
/** One line: what is being asked. */
title: string;
/** Extra context, one entry per line, already trimmed. May be empty. */
detail: string[];
/** Numbered choices parsed off the pane, empty when the frame did not parse. */
options: ApprovalOption[];
/** What the user can press right now, in words. */
hint: string;
}
/** Longest single line the card contributes before the renderer clips it. */
const MAX_CARD_TEXT = 400;
function clean(text: string | undefined): string {
return (text ?? '').replace(/\s+/g, ' ').trim().slice(0, MAX_CARD_TEXT);
}
export function approvalTone(item: ApprovalItem): TuiApprovalTone {
return item.kind === 'idle' ? 'warn' : 'err';
}
/**
* What the card says. Permission prompts lead with the tool (that is the whole
* question), questions lead with their message, and an idle prompt says what it
* is, since there is nothing to approve.
*/
export function approvalCard(item: ApprovalItem): TuiApprovalCard {
const options = item.options ?? [];
const message = clean(item.message);
const summary = clean(item.toolSummary) || clean(item.toolName);
if (item.kind === 'idle') {
return {
tone: 'warn',
title: message || 'waiting for your reply',
detail: [],
options: [],
hint: 'p to reply',
};
}
const title =
item.kind === 'permission'
? `requests: ${summary || 'permission'}`
: message || `question: ${summary || 'Claude is asking'}`;
const detail: string[] = [];
if (item.kind === 'permission' && message && message !== summary) detail.push(message);
return {
tone: 'err',
title,
detail,
options,
hint: options.length > 0 ? 'y approve · n deny · digit chooses' : 'y approve · n deny',
};
}
/**
* The parsed option that means "no". Claude renders it as `3. No, tell Claude
* what to do (esc)`, and answering with its digit is the same keystroke the
* dialog itself is waiting for; without a parsed one the answer route's `deny`
* sends Esc, which every dialog understands.
*/
export function approvalDenyOption(item: ApprovalItem): number | null {
const match = (item.options ?? []).find((option) => /^no\b/i.test(option.label));
return match ? match.n : null;
}
/**
* The answer one key produces, or null when that key means nothing here (so the
* caller can let its normal binding through).
*/
export function approvalAnswerForKey(item: ApprovalItem, key: string): TuiApprovalAnswer | null {
// An idle prompt has no dialog on screen: a digit or a `1` would land in the
// composer as text. The card points at `p` instead.
if (item.kind === 'idle') return null;
if (key === 'y') return { action: 'approve' };
if (key === 'n') {
const deny = approvalDenyOption(item);
return deny === null ? { action: 'deny' } : { action: 'option', option: deny };
}
if (key >= '1' && key <= '9') {
const option = Number.parseInt(key, 10);
return (item.options ?? []).some((entry) => entry.n === option) ? { action: 'option', option } : null;
}
return null;
}
/**
* Ids in `items` that `seen` has not recorded. The bell rings for these and for
* nothing else, which is what keeps a repaint (or a refetch that returns the
* same pending item) silent.
*
* Answered ids stay in `seen` on purpose: the inbox restores an item under its
* ORIGINAL id when a write fails, and re-ringing for a prompt the user already
* heard about is worse than missing one.
*/
export function newApprovalIds(seen: ReadonlySet<string>, items: readonly ApprovalItem[]): string[] {
const fresh: string[] = [];
for (const item of items) if (!seen.has(item.id) && !fresh.includes(item.id)) fresh.push(item.id);
return fresh;
}
File diff suppressed because it is too large Load Diff
+205
View File
@@ -0,0 +1,205 @@
/**
* @fileoverview Pure single-line editor behind the TUI's prompt composer (`p`)
* and search query (`/`).
*
* Text is held as CODE POINTS rather than a string, because every operation
* here is index-based and a cursor that can land inside a surrogate pair
* eventually deletes half an emoji. Combining marks are their own entries: they
* are zero-width, so they neither move the cursor's column nor cost a cell, and
* backspace peeling one off a base letter is what a terminal editor does.
*
* Scrolling is derived, never remembered implicitly: `composerScroll()` takes
* the width and returns the state whose window holds the cursor, which is what
* keeps "what the footer shows" a function of the state plus the terminal width
* rather than of the order the user pressed keys in.
*
* PURE: no IO, no timers, no `process.*`. Enter and Escape are reported as
* `submit`/`cancel` rather than acted on, since only the caller knows whether
* Enter means "send this prompt" or "open the highlighted search result".
*
* @module tui/tui-composer
*/
import { charWidth } from './tui-ansi.js';
import type { TuiInputEvent } from './tui-keys.js';
export interface TuiComposerState {
/** Code points. `chars.join('')` is the text. */
readonly chars: readonly string[];
/** 0..chars.length. The cursor sits BEFORE `chars[cursor]`. */
readonly cursor: number;
/** First visible code point, as `composerScroll()` last resolved it. */
readonly scroll: number;
}
export function createComposer(text = ''): TuiComposerState {
const chars = [...text];
return { chars, cursor: chars.length, scroll: 0 };
}
export function composerText(state: TuiComposerState): string {
return state.chars.join('');
}
function withChars(chars: readonly string[], cursor: number, scroll: number): TuiComposerState {
const clampedCursor = Math.min(Math.max(0, cursor), chars.length);
return { chars, cursor: clampedCursor, scroll: Math.min(Math.max(0, scroll), chars.length) };
}
/** Insert typed text at the cursor. Newlines are stripped: this is one line. */
export function composerInsert(state: TuiComposerState, value: string): TuiComposerState {
const inserted = [...value.replace(/[\r\n]+/g, ' ')];
if (inserted.length === 0) return state;
const chars = [...state.chars.slice(0, state.cursor), ...inserted, ...state.chars.slice(state.cursor)];
return withChars(chars, state.cursor + inserted.length, state.scroll);
}
/** Delete the code point before the cursor. */
export function composerBackspace(state: TuiComposerState): TuiComposerState {
if (state.cursor === 0) return state;
const chars = [...state.chars.slice(0, state.cursor - 1), ...state.chars.slice(state.cursor)];
return withChars(chars, state.cursor - 1, state.scroll);
}
/** Delete the code point under the cursor (the Delete key). */
export function composerDelete(state: TuiComposerState): TuiComposerState {
if (state.cursor >= state.chars.length) return state;
const chars = [...state.chars.slice(0, state.cursor), ...state.chars.slice(state.cursor + 1)];
return withChars(chars, state.cursor, state.scroll);
}
/** Delete back to the start of the word before the cursor (Ctrl+W). */
export function composerDeleteWord(state: TuiComposerState): TuiComposerState {
let start = state.cursor;
while (start > 0 && state.chars[start - 1] === ' ') start--;
while (start > 0 && state.chars[start - 1] !== ' ') start--;
if (start === state.cursor) return state;
const chars = [...state.chars.slice(0, start), ...state.chars.slice(state.cursor)];
return withChars(chars, start, state.scroll);
}
export function composerMove(state: TuiComposerState, delta: number): TuiComposerState {
const cursor = Math.min(Math.max(0, state.cursor + Math.trunc(delta)), state.chars.length);
return cursor === state.cursor ? state : withChars(state.chars, cursor, state.scroll);
}
export function composerHome(state: TuiComposerState): TuiComposerState {
return state.cursor === 0 ? state : withChars(state.chars, 0, state.scroll);
}
export function composerEnd(state: TuiComposerState): TuiComposerState {
return state.cursor === state.chars.length ? state : withChars(state.chars, state.chars.length, state.scroll);
}
export function composerClear(state: TuiComposerState): TuiComposerState {
return state.chars.length === 0 ? state : { chars: [], cursor: 0, scroll: 0 };
}
/** Display columns of `chars[from..to)`. */
function widthOf(chars: readonly string[], from: number, to: number): number {
let width = 0;
for (let i = from; i < to; i++) width += charWidth(chars[i].codePointAt(0) ?? 0);
return width;
}
/**
* Resolve `scroll` so the cursor is inside a window `width` columns wide,
* scrolling the minimum needed. One column is reserved for the cursor itself,
* so a cursor at the end of the text still has a cell to sit in instead of
* hanging one past the edge where the terminal would wrap it.
*/
export function composerScroll(state: TuiComposerState, width: number): TuiComposerState {
const usable = Math.max(0, Math.trunc(width) - 1);
let scroll = Math.min(Math.max(0, state.scroll), state.cursor);
while (scroll < state.cursor && widthOf(state.chars, scroll, state.cursor) > usable) scroll++;
return scroll === state.scroll ? state : { chars: state.chars, cursor: state.cursor, scroll };
}
export interface TuiComposerWindow {
/** The visible slice of the text. */
text: string;
/** Cursor offset in display columns from the start of `text`. */
cursorColumn: number;
/** Resolved first visible code point (may differ from `state.scroll`). */
scroll: number;
}
/**
* The slice the footer draws plus where the terminal cursor belongs. The scroll
* is resolved here too, so a renderer that never writes state back still shows
* the cursor.
*/
export function composerWindow(state: TuiComposerState, width: number): TuiComposerWindow {
const columns = Math.max(1, Math.trunc(width));
const scrolled = composerScroll(state, columns);
const { chars, cursor, scroll } = scrolled;
let used = 0;
let end = scroll;
while (end < chars.length) {
const next = charWidth(chars[end].codePointAt(0) ?? 0);
if (used + next > columns) break;
used += next;
end++;
}
return {
text: chars.slice(scroll, Math.max(end, cursor)).join(''),
cursorColumn: widthOf(chars, scroll, cursor),
scroll,
};
}
export type TuiComposerStep =
| { kind: 'edit'; state: TuiComposerState }
| { kind: 'submit'; text: string }
| { kind: 'cancel' }
| { kind: 'ignore' };
/**
* One keystroke. Enter and Escape are REPORTED rather than applied: `p` sends
* the line while `/` opens the highlighted result, and only the caller knows
* which.
*/
export function composerStep(state: TuiComposerState, event: TuiInputEvent): TuiComposerStep {
switch (event.type) {
case 'char':
return { kind: 'edit', state: composerInsert(state, event.value) };
case 'backspace':
return { kind: 'edit', state: composerBackspace(state) };
case 'enter':
return { kind: 'submit', text: composerText(state) };
case 'escape':
return { kind: 'cancel' };
case 'key':
switch (event.name) {
case 'left':
return { kind: 'edit', state: composerMove(state, -1) };
case 'right':
return { kind: 'edit', state: composerMove(state, 1) };
case 'home':
return { kind: 'edit', state: composerHome(state) };
case 'end':
return { kind: 'edit', state: composerEnd(state) };
case 'delete':
return { kind: 'edit', state: composerDelete(state) };
default:
return { kind: 'ignore' };
}
case 'ctrl':
switch (event.key) {
case 'c':
return { kind: 'cancel' };
case 'a':
return { kind: 'edit', state: composerHome(state) };
case 'e':
return { kind: 'edit', state: composerEnd(state) };
case 'u':
return { kind: 'edit', state: composerClear(state) };
case 'w':
return { kind: 'edit', state: composerDeleteWord(state) };
default:
return { kind: 'ignore' };
}
default:
return { kind: 'ignore' };
}
}
+90
View File
@@ -0,0 +1,90 @@
/**
* @fileoverview Pure formatting of `GET /api/away-digest` into the lines the
* `g` overlay scrolls.
*
* The digest answers "what happened while I was away", so it is read top-down
* and never studied: every entry is one line (age, session, what happened), a
* long section is capped with a "… n more" tail rather than allowed to push the
* next section off screen, and the counts that matter live in the first line
* where they are visible without scrolling at all.
*
* PURE: no IO, no clock of its own (the caller passes `now`), no `process.*`.
*
* @module tui/tui-digest
*/
import { formatElapsed, formatTokens } from './tui-render.js';
import type { AwayDigestItem, AwayDigestResponse, AwayDigestSectionName } from '../web/away-digest.js';
/** Entries per section before the tail takes over. */
export const DIGEST_SECTION_LIMIT = 6;
const SECTION_ORDER: ReadonlyArray<readonly [AwayDigestSectionName, string]> = [
['needsAttention', 'NEEDS ATTENTION'],
['completed', 'COMPLETED'],
['stillRunning', 'STILL RUNNING'],
['idle', 'IDLE'],
['informational', 'INFO'],
];
const RANGE_WORDS: Record<string, string> = {
'since-last-visit': 'since your last visit',
'1h': 'the last hour',
today: 'today',
'24h': 'the last 24 hours',
custom: 'the selected window',
};
export interface TuiDigestOptions {
now: number;
sectionLimit?: number;
}
function ageColumn(item: AwayDigestItem, now: number): string {
const age = item.timestamp > 0 ? formatElapsed(now - item.timestamp) : '';
return age.padEnd(4);
}
function itemLine(item: AwayDigestItem, now: number): string {
const who = item.sessionName ?? item.sessionId?.slice(0, 8) ?? '';
const what = [item.title, item.detail].filter((part) => part && part.trim() !== '').join(' · ');
return ` ${ageColumn(item, now)} ${[who, what].filter((part) => part !== '').join(' ')}`.replace(/\s+$/, '');
}
/**
* The digest as display lines. The first line is the summary, then one block
* per non-empty section, then the token totals when the range had any.
*/
export function formatAwayDigest(digest: AwayDigestResponse, options: TuiDigestOptions): string[] {
const limit = Math.max(1, Math.trunc(options.sectionLimit ?? DIGEST_SECTION_LIMIT));
const { totals } = digest;
const lines: string[] = [
[
RANGE_WORDS[digest.range.range] ?? 'recently',
`${totals.sessionsCreated} started`,
`${totals.sessionsExited} exited`,
`${totals.activeSessions} running`,
].join(' · '),
];
let entries = 0;
for (const [key, label] of SECTION_ORDER) {
const items = digest.sections[key] ?? [];
if (items.length === 0) continue;
entries += items.length;
lines.push('', `${label} (${items.length})`);
for (const item of items.slice(0, limit)) lines.push(itemLine(item, options.now));
if (items.length > limit) lines.push(` … ${items.length - limit} more`);
}
if (entries === 0) lines.push('', 'nothing happened while you were away');
const tokens = [
formatTokens(totals.inputTokens ?? 0) ? `${formatTokens(totals.inputTokens ?? 0)} in` : '',
formatTokens(totals.outputTokens ?? 0) ? `${formatTokens(totals.outputTokens ?? 0)} out` : '',
typeof totals.estimatedCost === 'number' && totals.estimatedCost > 0 ? `$${totals.estimatedCost.toFixed(2)}` : '',
].filter((part) => part !== '');
if (tokens.length > 0) lines.push('', `tokens: ${tokens.join(' · ')}`);
return lines;
}
+240
View File
@@ -0,0 +1,240 @@
/**
* @fileoverview Pure byte-stream to input-event parser for raw-mode stdin.
*
* Stateful (a sequence can arrive split across reads, and a UTF-8 character can
* be split mid-code-point) but pure: it owns a byte buffer and nothing else, no
* stdin, no timers. The one timing decision a terminal forces on us stays with
* the caller: a lone ESC is indistinguishable from the start of an arrow key
* until something either follows it or does not, so the parser HOLDS a trailing
* ESC and the caller calls `flush()` after ~30ms of silence to turn it into an
* Escape event.
*
* Unknown sequences are swallowed rather than leaked as text: a stray
* `CSI 200~` must never end up typed into a prompt composer.
*
* @module tui/tui-keys
*/
/** Keys with a name rather than a character. */
export type TuiNamedKey =
| 'up'
| 'down'
| 'left'
| 'right'
| 'home'
| 'end'
| 'pageup'
| 'pagedown'
| 'delete'
| 'insert';
export type TuiMouseKind = 'press' | 'release' | 'wheel-up' | 'wheel-down';
/** Discriminated union, exhaustive-switch friendly (see `utils/assertNever`). */
export type TuiInputEvent =
| { type: 'char'; value: string }
| { type: 'enter' }
| { type: 'tab' }
| { type: 'backspace' }
| { type: 'escape' }
| { type: 'ctrl'; key: string }
| { type: 'alt'; value: string }
| { type: 'key'; name: TuiNamedKey }
| { type: 'mouse'; kind: TuiMouseKind; x: number; y: number; button: number };
export interface TuiKeyParser {
/** Decode a chunk. Incomplete tails are held for the next call. */
feed(chunk: Buffer | string): TuiInputEvent[];
/** Resolve a held ESC (the caller's disambiguation timer fired). */
flush(): TuiInputEvent[];
/** Bytes currently held back. Exposed for the ESC timer and for tests. */
pending(): number;
}
/**
* An unterminated sequence longer than this is not a sequence: the held bytes
* are dropped whole, so a garbage burst can neither wedge the parser nor leak
* its bytes into a prompt as typed characters.
*/
const MAX_PENDING_BYTES = 64;
/** Bytes in a UTF-8 sequence given its lead byte; 0 for a byte that cannot lead one. */
function utf8SequenceLength(lead: number): number {
if (lead < 0x80) return 1;
if (lead >= 0xc2 && lead <= 0xdf) return 2;
if (lead >= 0xe0 && lead <= 0xef) return 3;
if (lead >= 0xf0 && lead <= 0xf4) return 4;
return 0;
}
const CSI_FINAL_KEYS: Record<string, TuiNamedKey> = {
A: 'up',
B: 'down',
C: 'right',
D: 'left',
H: 'home',
F: 'end',
};
/** `CSI <n> ~` keys, by their first numeric parameter. */
const CSI_TILDE_KEYS: Record<number, TuiNamedKey> = {
1: 'home',
2: 'insert',
3: 'delete',
4: 'end',
5: 'pageup',
6: 'pagedown',
7: 'home',
8: 'end',
};
/** Result of trying to parse one sequence off the front of the buffer. */
type ParseStep = { consumed: number; events: TuiInputEvent[] } | 'incomplete';
const NOTHING: TuiInputEvent[] = [];
export function createKeyParser(): TuiKeyParser {
let buf: Buffer = Buffer.alloc(0);
/** Parse the CSI/SS3 sequence that starts at buf[0] === ESC. */
const parseEscape = (): ParseStep => {
if (buf.length < 2) return 'incomplete';
const second = buf[1];
// SS3 (`ESC O <final>`): the arrows/Home/End of application-cursor mode.
if (second === 0x4f) {
if (buf.length < 3) return 'incomplete';
const name = CSI_FINAL_KEYS[String.fromCharCode(buf[2])];
return { consumed: 3, events: name ? [{ type: 'key', name }] : NOTHING };
}
// ESC followed by a printable character IN THE SAME READ is Alt+that key:
// that is how every terminal sends a meta chord. A lone Esc cannot look
// like this, because a buffer holding only ESC returns 'incomplete' above
// and is flushed as `escape` when the read ends, which is the standard way
// to tell the two apart without a timer.
//
// ⚠️ Three characters are deliberately NOT treated as Alt chords, because
// the terminal uses them to introduce sequences and a chord is
// indistinguishable from one: `[` (CSI) and `O` (SS3) would swallow every
// arrow key, and `]` (OSC) would swallow a terminal's colour-query reply.
// Alt+[ and Alt+] therefore cannot exist in a terminal at all, which is why
// the list binds bare `[` and `]` for the same job.
if (second !== 0x5b) {
if (second >= 0x20 && second <= 0x7e && second !== 0x4f && second !== 0x5d) {
return { consumed: 2, events: [{ type: 'alt', value: String.fromCharCode(second) }] };
}
return { consumed: 1, events: [{ type: 'escape' }] };
}
let j = 2;
while (j < buf.length && buf[j] >= 0x30 && buf[j] <= 0x3f) j++;
while (j < buf.length && buf[j] >= 0x20 && buf[j] <= 0x2f) j++;
if (j >= buf.length) return 'incomplete';
const final = String.fromCharCode(buf[j]);
const params = buf.subarray(2, j).toString('latin1');
const consumed = j + 1;
// X10 mouse (`CSI M` + 3 raw bytes): swallowed, but its payload bytes must
// be consumed or they would surface as typed characters.
if (params === '' && final === 'M') {
if (buf.length < consumed + 3) return 'incomplete';
return { consumed: consumed + 3, events: NOTHING };
}
if (params.startsWith('<') && (final === 'M' || final === 'm')) {
return { consumed, events: parseSgrMouse(params.slice(1), final) };
}
if (final === '~') {
const name = CSI_TILDE_KEYS[Number.parseInt(params, 10)];
return { consumed, events: name ? [{ type: 'key', name }] : NOTHING };
}
// Modified arrows (`CSI 1;5A`) carry the same final byte; the modifier is
// dropped rather than exposed, since nothing in the keymap wants it yet.
const named = CSI_FINAL_KEYS[final];
return { consumed, events: named ? [{ type: 'key', name: named }] : NOTHING };
};
const parseSgrMouse = (params: string, final: string): TuiInputEvent[] => {
const parts = params.split(';');
if (parts.length < 3) return NOTHING;
const button = Number.parseInt(parts[0], 10);
const x = Number.parseInt(parts[1], 10);
const y = Number.parseInt(parts[2], 10);
if (!Number.isFinite(button) || !Number.isFinite(x) || !Number.isFinite(y)) return NOTHING;
if (button >= 64) {
// 64 = wheel up, 65 = wheel down (the low bit is the direction).
const kind: TuiMouseKind = (button & 1) === 1 ? 'wheel-down' : 'wheel-up';
return [{ type: 'mouse', kind, x, y, button }];
}
// Motion reports (bit 32) would fire on every pixel of a drag; nothing in
// the keymap consumes them, so they are swallowed here rather than upstream.
if ((button & 32) === 32) return NOTHING;
return [{ type: 'mouse', kind: final === 'M' ? 'press' : 'release', x, y, button }];
};
/** Parse one non-escape byte (or one UTF-8 character) off the front. */
const parseByte = (): ParseStep => {
const b = buf[0];
// LF counts as Enter because some terminals send it for Return; the cost is
// that Ctrl+J is not bindable, which no key in the plan's keymap wants.
if (b === 0x0d || b === 0x0a) return { consumed: 1, events: [{ type: 'enter' }] };
if (b === 0x09) return { consumed: 1, events: [{ type: 'tab' }] };
if (b === 0x7f || b === 0x08) return { consumed: 1, events: [{ type: 'backspace' }] };
if (b === 0x00) return { consumed: 1, events: [{ type: 'ctrl', key: '@' }] };
if (b >= 0x01 && b <= 0x1a) {
return { consumed: 1, events: [{ type: 'ctrl', key: String.fromCharCode(b + 0x60) }] };
}
if (b >= 0x1c && b <= 0x1f) {
return { consumed: 1, events: [{ type: 'ctrl', key: String.fromCharCode(b + 0x40) }] };
}
const length = utf8SequenceLength(b);
if (length === 0) return { consumed: 1, events: NOTHING };
if (buf.length < length) return 'incomplete';
const value = buf.subarray(0, length).toString('utf8');
// A lead byte followed by junk decodes to U+FFFD; that is corruption on the
// wire, not something to type into a composer. Only the bad lead byte is
// dropped, so whatever valid input followed it still decodes.
if (value.includes('�')) return { consumed: 1, events: NOTHING };
return { consumed: length, events: [{ type: 'char', value }] };
};
/** Drain the buffer, stopping at the first incomplete sequence. */
const drain = (events: TuiInputEvent[]): void => {
while (buf.length > 0) {
const step = buf[0] === 0x1b ? parseEscape() : parseByte();
if (step === 'incomplete') {
if (buf.length > MAX_PENDING_BYTES) buf = Buffer.alloc(0);
return;
}
for (const event of step.events) events.push(event);
buf = buf.subarray(step.consumed);
}
};
return {
feed(chunk: Buffer | string): TuiInputEvent[] {
const bytes = typeof chunk === 'string' ? Buffer.from(chunk, 'utf8') : chunk;
buf = buf.length === 0 ? Buffer.from(bytes) : Buffer.concat([buf, bytes]);
const events: TuiInputEvent[] = [];
drain(events);
return events;
},
flush(): TuiInputEvent[] {
const events: TuiInputEvent[] = [];
if (buf.length > 0 && buf[0] === 0x1b) {
events.push({ type: 'escape' });
buf = buf.subarray(1);
drain(events);
}
return events;
},
pending(): number {
return buf.length;
},
};
}
+137
View File
@@ -0,0 +1,137 @@
/**
* @fileoverview Pure responsive layout math for the TUI frame.
*
* One rule decides the shape: below 72 columns (Termius, iPhone portrait) the
* preview pane is gone and rows take two lines, which is the constraint the
* `sc` chooser was built around and the reason it is still usable on a phone.
* Above it, a clamped sidebar carries the session list and the preview takes
* the rest.
*
* Rectangles are 1-based (row 1, column 1 is the top-left cell) because that is
* what `ESC [ <row>;<col> H` takes, and every region is clamped to a
* non-negative size so a 5x5 terminal degrades instead of producing negative
* widths that would crash the renderer.
*
* @module tui/tui-layout
*/
import type { TuiConnectionStatus } from './tui-types.js';
/** Width at which the preview pane is dropped and rows become two lines. */
export const NARROW_BREAKPOINT = 72;
/** Sidebar clamp: narrower than this and a session name stops being readable. */
export const SIDEBAR_MIN_WIDTH = 34;
/** Sidebar clamp: wider than this is wasted on a list of short names. */
export const SIDEBAR_MAX_WIDTH = 44;
/** A preview thinner than this shows nothing useful, so the layout goes narrow instead. */
export const PREVIEW_MIN_WIDTH = 24;
/** Share of the width the sidebar aims for between the clamps. */
const SIDEBAR_RATIO = 0.36;
export interface TuiRect {
/** 1-based terminal row of the first line. */
row: number;
/** 1-based terminal column of the first cell. */
col: number;
width: number;
height: number;
}
export interface TuiLayoutOptions {
/**
* Reserve one line under the header for the connection banner. The caller
* decides with `needsBanner(model.connection)`, so layout stays pure math.
*/
banner?: boolean;
}
export interface TuiLayout {
cols: number;
rows: number;
/** No preview pane, two-line rows. */
narrow: boolean;
/** Terminal lines one session row occupies. */
rowHeight: 1 | 2;
header: TuiRect;
/** Connection banner, when the caller asked for one and there was room. */
banner: TuiRect | null;
/** Everything between header and footer, banner included. */
body: TuiRect;
/** The session list. */
list: TuiRect;
/** The one-column rule between list and preview; null in narrow mode. */
divider: TuiRect | null;
/** The preview pane; null in narrow mode. */
preview: TuiRect | null;
footer: TuiRect;
}
/** Which connection states get a banner line under the header. */
export function needsBanner(connection: TuiConnectionStatus): boolean {
return connection !== 'connected';
}
function clamp(value: number, min: number, max: number): number {
return Math.min(max, Math.max(min, value));
}
/**
* Rectangles for one frame at `cols` x `rows`.
*
* The header always exists; the footer appears from 2 rows up; the body is
* whatever is left, which may legitimately be zero lines high.
*/
export function computeLayout(cols: number, rows: number, options: TuiLayoutOptions = {}): TuiLayout {
const width = Math.max(1, Math.floor(cols) || 1);
const height = Math.max(1, Math.floor(rows) || 1);
const headerHeight = 1;
const footerHeight = height >= 2 ? 1 : 0;
const bodyHeight = Math.max(0, height - headerHeight - footerHeight);
const bodyRow = headerHeight + 1;
const header: TuiRect = { row: 1, col: 1, width, height: headerHeight };
const footer: TuiRect = { row: height, col: 1, width, height: footerHeight };
const body: TuiRect = { row: bodyRow, col: 1, width, height: bodyHeight };
const bannerHeight = options.banner === true && bodyHeight > 0 ? 1 : 0;
const banner: TuiRect | null = bannerHeight > 0 ? { row: bodyRow, col: 1, width, height: 1 } : null;
const contentRow = bodyRow + bannerHeight;
const contentHeight = Math.max(0, bodyHeight - bannerHeight);
const sidebarTarget = Math.floor(width * SIDEBAR_RATIO);
const sidebarWidth = clamp(sidebarTarget, SIDEBAR_MIN_WIDTH, SIDEBAR_MAX_WIDTH);
const previewWidth = width - sidebarWidth - 1;
const narrow = width < NARROW_BREAKPOINT || previewWidth < PREVIEW_MIN_WIDTH;
if (narrow) {
return {
cols: width,
rows: height,
narrow: true,
rowHeight: 2,
header,
banner,
body,
list: { row: contentRow, col: 1, width, height: contentHeight },
divider: null,
preview: null,
footer,
};
}
return {
cols: width,
rows: height,
narrow: false,
rowHeight: 1,
header,
banner,
body,
list: { row: contentRow, col: 1, width: sidebarWidth, height: contentHeight },
divider: { row: contentRow, col: sidebarWidth + 1, width: 1, height: contentHeight },
preview: { row: contentRow, col: sidebarWidth + 2, width: previewWidth, height: contentHeight },
footer,
};
}
+562
View File
@@ -0,0 +1,562 @@
/**
* @fileoverview Pure state, classification and grouping for the TUI dashboard.
*
* Classification speaks the web UI's language on purpose (red blocked, yellow
* waiting, green working, muted idle), because a user who has both surfaces
* open must never have to translate between them. The inputs are the ones the
* server already computes: a unified-list row and, when the session is blocked,
* the approvals-inbox item that blocks it. Nothing here screen-scrapes.
*
* Selection is tracked by session id, never by row index: rows re-sort under
* the cursor constantly (a session starts working, an approval lands), and an
* index-tracked cursor would silently move the selection to a different
* session between two keystrokes.
*
* PURE: no IO, no timers, no `process.*`. The store mutates its own state and
* nothing else.
*
* @module tui/tui-model
*/
import type { SearchResultGroup, SearchSourceType } from '../types/search.js';
import type { ApprovalItem } from '../web/approval-inbox.js';
import type {
TuiConfirmState,
TuiConnectionStatus,
TuiDigestState,
TuiGroup,
TuiGroupKey,
TuiHeaderInfo,
TuiMessage,
TuiPickerState,
TuiPreview,
TuiPromptState,
TuiRenderModel,
TuiRow,
TuiSearchEntry,
TuiSearchState,
TuiSessionRow,
TuiSessionState,
TuiUiMode,
} from './tui-types.js';
/** How many history rows the RECENT group shows before it stops being a dashboard. */
export const DEFAULT_RECENT_LIMIT = 8;
export const GROUP_ORDER: readonly TuiGroupKey[] = ['needs-you', 'working', 'idle', 'recent'];
export const GROUP_LABELS: Record<TuiGroupKey, string> = {
'needs-you': 'NEEDS YOU',
working: 'WORKING',
idle: 'IDLE',
recent: 'RECENT',
};
const STATE_GROUP: Record<TuiSessionState, TuiGroupKey> = {
'blocked-question': 'needs-you',
'blocked-permission': 'needs-you',
waiting: 'needs-you',
working: 'working',
idle: 'idle',
recent: 'recent',
};
/** A row is live when the unified merge saw it in the in-memory session map. */
export function isLiveRow(session: TuiSessionRow): boolean {
return Array.isArray(session.sources) && session.sources.includes('live');
}
/**
* Classify one row.
*
* Order matters and mirrors `_mobileOverviewState()` in the web UI: a pending
* prompt outranks everything (it is literally blocking the agent), and it
* outranks a stale `busy` status because the hook is the newer signal. An
* errored session has no state of its own here and joins the waiting tier,
* since it is equally something only a human can clear.
*/
export function classifySession(session: TuiSessionRow, approval?: ApprovalItem): TuiSessionState {
if (!isLiveRow(session)) return 'recent';
if (approval) {
if (approval.kind === 'permission') return 'blocked-permission';
if (approval.kind === 'question') return 'blocked-question';
return 'waiting';
}
if (session.status === 'error') return 'waiting';
if (session.isWorking === true || session.status === 'busy') return 'working';
return 'idle';
}
/**
* Epoch ms the session entered its current state, which is what the intra-group
* ordering sorts on. 0 when nothing usable is known.
*
* A WORKING pane repaints about once a second, so its `lastActivityAt` is
* always "now" and would report every running turn as freshly started; the
* turn's own start is the pane's last Enter.
*/
export function stateSince(state: TuiSessionState, session: TuiSessionRow, approval?: ApprovalItem): number {
if (approval) return approval.createdAt;
if (state === 'working') return session.lastSubmitAt ?? session.createdAt ?? 0;
return session.lastActivityAt ?? session.createdAt ?? 0;
}
/** Classify a batch of rows against the pending approvals, keyed by session id. */
export function buildRows(
sessions: readonly TuiSessionRow[],
approvals: ReadonlyMap<string, ApprovalItem> = new Map()
): TuiRow[] {
return sessions.map((session) => {
const approval = approvals.get(session.sessionId);
const state = classifySession(session, approval);
const row: TuiRow = {
session,
state,
group: STATE_GROUP[state],
since: stateSince(state, session, approval),
};
if (approval) row.approval = approval;
return row;
});
}
function compareIds(a: TuiRow, b: TuiRow): number {
if (a.session.sessionId < b.session.sessionId) return -1;
if (a.session.sessionId > b.session.sessionId) return 1;
return 0;
}
/** Longest first: the oldest anchor wins, and an unknown anchor sorts last. */
function compareLongestFirst(a: TuiRow, b: TuiRow): number {
const left = a.since || Number.MAX_SAFE_INTEGER;
const right = b.since || Number.MAX_SAFE_INTEGER;
return left !== right ? left - right : compareIds(a, b);
}
/** Newest first: the freshest anchor wins, and an unknown anchor sorts last. */
function compareNewestFirst(a: TuiRow, b: TuiRow): number {
const left = a.since || 0;
const right = b.since || 0;
return left !== right ? right - left : compareIds(a, b);
}
export interface GroupOptions {
/** RECENT is a tail, not a list: everything past this is dropped. */
recentLimit?: number;
}
/**
* Split classified rows into the four display groups.
*
* Always returns all four in display order (empty ones included) so callers
* never have to guess the shape; the renderer skips the empty ones.
*
* NEEDS YOU and WORKING are ordered by how long they have been in that state
* (longest first: the thing that has waited longest for you is the thing to
* look at). IDLE and RECENT are ordered by recency, newest first.
*/
export function groupSessions(rows: readonly TuiRow[], options: GroupOptions = {}): TuiGroup[] {
const recentLimit = Math.max(0, Math.floor(options.recentLimit ?? DEFAULT_RECENT_LIMIT));
const buckets: Record<TuiGroupKey, TuiRow[]> = {
'needs-you': [],
working: [],
idle: [],
recent: [],
};
for (const row of rows) buckets[row.group].push(row);
buckets['needs-you'].sort(compareLongestFirst);
buckets.working.sort(compareLongestFirst);
buckets.idle.sort(compareNewestFirst);
buckets.recent.sort(compareNewestFirst);
buckets.recent = buckets.recent.slice(0, recentLimit);
return GROUP_ORDER.map((key) => ({ key, label: GROUP_LABELS[key], rows: buckets[key] }));
}
/** The cursor's list: group headers are chrome, only sessions are selectable. */
export function flattenRows(groups: readonly TuiGroup[]): TuiRow[] {
const rows: TuiRow[] = [];
for (const group of groups) rows.push(...group.rows);
return rows;
}
/**
* Fold an incoming row into a known one. Defined fields win, `undefined` never
* clobbers (a live SSE payload carries no transcript fields, a unified refresh
* carries no token counters), but a non-empty `sources` list REPLACES rather
* than unions: a session that ended must be able to lose its `live` source and
* fall to RECENT.
*/
export function mergeSessionRow(existing: TuiSessionRow, incoming: TuiSessionRow): TuiSessionRow {
const merged: TuiSessionRow = { ...existing };
for (const [key, value] of Object.entries(incoming)) {
if (value === undefined) continue;
(merged as unknown as Record<string, unknown>)[key] = value;
}
merged.sources = incoming.sources?.length ? [...incoming.sources] : [...(existing.sources ?? [])];
return merged;
}
// ─────────────────────────────────────────────────────────────────────────────
// Search results (pure)
// ─────────────────────────────────────────────────────────────────────────────
const SEARCH_GROUP_LABELS: Record<SearchSourceType, string> = {
session: 'SESSIONS',
event: 'EVENTS',
file: 'FILES',
};
/**
* A session snippet opens with the session's own name, which the row already
* shows in its first column (`search-service.ts` builds it as
* `w1-alpha <em dash> /tmp/alpha`, hence the separator in the pattern).
* Dropping the repeat is what keeps a result row from reading as a stutter.
*/
function withoutLabelPrefix(snippet: string, label: string): string {
const rest = snippet.startsWith(label) ? snippet.slice(label.length) : snippet;
return rest === snippet ? snippet : rest.replace(/^\s*(?:[—:-]\s*)?/, '');
}
/**
* Flatten `GET /api/search`'s typed groups into the overlay's lines: a header
* per group, then its results. Only a result row carries a session id, which is
* what the cursor uses to skip headers.
*
* `isLive` decides which rows can hand the dashboard a session: a history hit
* has a session id too, but selecting it would move the cursor to a row that is
* not on the list.
*/
export function buildSearchEntries(
groups: readonly SearchResultGroup[],
isLive: (sessionId: string) => boolean
): TuiSearchEntry[] {
const entries: TuiSearchEntry[] = [];
for (const group of groups) {
if (group.results.length === 0) continue;
entries.push({ kind: 'header', text: SEARCH_GROUP_LABELS[group.type] ?? group.type.toUpperCase() });
for (const result of group.results) {
const live = result.jumpTo.kind === 'session' && isLive(result.sessionId);
const label = result.jumpTo.relativePath ?? result.sessionName ?? result.sessionId.slice(0, 8);
entries.push({
kind: 'result',
text: label,
detail: withoutLabelPrefix(result.snippet, label),
sessionId: result.sessionId,
live,
});
}
}
return entries;
}
/** First selectable row, or -1 when the list is all headers (or empty). */
export function firstSearchIndex(entries: readonly TuiSearchEntry[]): number {
return entries.findIndex((entry) => entry.kind === 'result');
}
/**
* Move the search cursor by `delta` result rows, skipping headers and stopping
* at both ends (wrapping a search result list scrolls past the answer the user
* was reading).
*/
export function moveSearchIndex(entries: readonly TuiSearchEntry[], index: number, delta: number): number {
const step = Math.trunc(delta);
if (step === 0) return index;
const direction = step > 0 ? 1 : -1;
let current = index;
for (let remaining = Math.abs(step); remaining > 0; remaining--) {
let next = current + direction;
while (next >= 0 && next < entries.length && entries[next].kind !== 'result') next += direction;
if (next < 0 || next >= entries.length) break;
current = next;
}
return current;
}
/**
* The dashboard's state. Update methods mutate in place (one store per TUI
* process, no subscribers) and every derived view is recomputed from scratch,
* which keeps "what is on screen" a pure function of the stored facts.
*/
export class TuiModelStore implements TuiRenderModel {
private sessionsById = new Map<string, TuiSessionRow>();
private approvalsBySession = new Map<string, ApprovalItem>();
private _revision = 0;
selectedId: string | null = null;
connection: TuiConnectionStatus = 'connected';
mode: TuiUiMode = 'list';
header: TuiHeaderInfo = {};
preview: TuiPreview | null = null;
message: TuiMessage | null = null;
confirm: TuiConfirmState | null = null;
picker: TuiPickerState | null = null;
prompt: TuiPromptState | null = null;
search: TuiSearchState | null = null;
digest: TuiDigestState | null = null;
recentLimit: number;
constructor(options: GroupOptions = {}) {
this.recentLimit = Math.max(0, Math.floor(options.recentLimit ?? DEFAULT_RECENT_LIMIT));
}
/**
* Bumped by every mutating method. The app layer repaints when this changed
* (plus on resize and on the animation tick), which is what keeps an idle
* dashboard from redrawing itself. Writing a public field directly bypasses
* it, so state changes go through the methods below.
*/
get revision(): number {
return this._revision;
}
private touch(): void {
this._revision++;
}
// ── Data ───────────────────────────────────────────────────────────────────
upsertSession(session: TuiSessionRow): void {
this.mutate(() => {
const existing = this.sessionsById.get(session.sessionId);
this.sessionsById.set(session.sessionId, existing ? mergeSessionRow(existing, session) : { ...session });
});
}
removeSession(sessionId: string): void {
this.mutate(() => {
this.sessionsById.delete(sessionId);
this.approvalsBySession.delete(sessionId);
});
}
/** Full refresh (a `GET /api/sessions/unified` poll): the server is authoritative. */
replaceSessions(sessions: readonly TuiSessionRow[]): void {
this.mutate(() => {
this.sessionsById.clear();
for (const session of sessions) this.sessionsById.set(session.sessionId, { ...session });
});
}
setApprovals(items: readonly ApprovalItem[]): void {
this.mutate(() => {
this.approvalsBySession.clear();
// One active item per session is an inbox invariant; the newest wins if
// that ever stops being true.
for (const item of items) this.approvalsBySession.set(item.sessionId, item);
});
}
sessions(): TuiSessionRow[] {
return [...this.sessionsById.values()];
}
// ── Chrome ─────────────────────────────────────────────────────────────────
setConnection(status: TuiConnectionStatus): void {
if (this.connection === status) return;
this.connection = status;
this.touch();
}
setHeader(header: TuiHeaderInfo): void {
this.header = { ...this.header, ...header };
this.touch();
}
setPreview(preview: TuiPreview | null): void {
this.preview = preview;
this.touch();
}
setMode(mode: TuiUiMode): void {
if (this.mode === mode) return;
this.mode = mode;
this.touch();
}
setMessage(message: TuiMessage | null): void {
this.message = message;
this.mode = message ? 'message' : 'list';
this.touch();
}
/** Show (or clear) the overlay chooser. Setting one takes the keyboard. */
setPicker(picker: TuiPickerState | null): void {
this.picker = picker;
this.mode = picker ? 'new-session' : 'list';
this.touch();
}
/** Open (or close) the one-line prompt composer. Setting one takes the keyboard. */
setPrompt(prompt: TuiPromptState | null): void {
this.prompt = prompt;
this.mode = prompt ? 'prompt' : 'list';
this.touch();
}
/** Replace the composer's editor state, keeping the target session. */
updatePrompt(composer: TuiPromptState['composer']): void {
if (!this.prompt || this.prompt.composer === composer) return;
this.prompt = { ...this.prompt, composer };
this.touch();
}
setSearch(search: TuiSearchState | null): void {
this.search = search;
this.mode = search ? 'search' : 'list';
this.touch();
}
/** Fold a partial update into the open search overlay. No-op when it is closed. */
updateSearch(patch: Partial<TuiSearchState>): void {
if (!this.search) return;
this.search = { ...this.search, ...patch };
this.touch();
}
setDigest(digest: TuiDigestState | null): void {
this.digest = digest;
this.mode = digest ? 'digest' : 'list';
this.touch();
}
/**
* Scroll the digest by `delta` lines. `capacity` is how many lines the box
* shows, so the last page cannot scroll into empty space.
*/
scrollDigest(delta: number, capacity: number): void {
if (!this.digest) return;
const room = Math.max(0, this.digest.lines.length - Math.max(1, Math.trunc(capacity)));
const offset = Math.min(Math.max(0, this.digest.offset + Math.trunc(delta)), room);
if (offset === this.digest.offset) return;
this.digest = { ...this.digest, offset };
this.touch();
}
/**
* Arm the typed-name confirmation for `x` (kill). Whether what the user typed
* AUTHORIZES the kill is `confirmAccepts()` in tui-app, which owns that rule
* for every caller: a second copy here answered the same question differently
* (it refused the id prefix a mux name carries) and nothing consulted it.
*/
beginConfirmKill(row: TuiRow, label: string): void {
this.confirm = {
sessionId: row.session.sessionId,
// ⚠️ Passed in, not derived here. `row.session.name ?? id.slice(0,8)`
// used to compute it, and `??` falls back only on null/undefined: a
// session whose name is the EMPTY STRING (every session the server did
// not name) sailed through it and the dialog read "Kill ?". A destructive
// prompt that cannot say what it is about to destroy is worse than no
// prompt, and it is now one keystroke. The caller passes the same label
// the LIST shows, so the dialog names the row the user is looking at.
name: label,
};
this.mode = 'confirm-kill';
this.touch();
}
/** Drop whatever overlay owns the keyboard and go back to the list. */
closeOverlay(): void {
this.confirm = null;
this.message = null;
this.picker = null;
this.prompt = null;
this.search = null;
this.digest = null;
this.mode = 'list';
this.touch();
}
// ── Derived views ──────────────────────────────────────────────────────────
groups(): TuiGroup[] {
return groupSessions(buildRows(this.sessions(), this.approvalsBySession), { recentLimit: this.recentLimit });
}
rows(): TuiRow[] {
return flattenRows(this.groups());
}
get sessionCount(): number {
let count = 0;
for (const session of this.sessionsById.values()) if (isLiveRow(session)) count++;
return count;
}
// ── Cursor ─────────────────────────────────────────────────────────────────
selectedSession(): TuiRow | null {
if (!this.selectedId) return null;
return this.rows().find((row) => row.session.sessionId === this.selectedId) ?? null;
}
/** Select a session by id. Returns false when it is not on screen. */
select(sessionId: string): boolean {
if (!this.rows().some((row) => row.session.sessionId === sessionId)) return false;
this.moveTo(sessionId);
return true;
}
/** Move by `delta` rows, skipping group headers and wrapping at both ends. */
moveCursor(delta: number): void {
const rows = this.rows();
if (rows.length === 0) {
this.moveTo(null);
return;
}
const current = this.indexOfSelected(rows);
if (current < 0) {
this.moveTo(rows[delta >= 0 ? 0 : rows.length - 1].session.sessionId);
return;
}
const step = Math.trunc(delta);
const next = (((current + step) % rows.length) + rows.length) % rows.length;
this.moveTo(rows[next].session.sessionId);
}
/** The 1-9 jump: `n` is the 1-based position in the flattened list. */
cursorToIndex(n: number): boolean {
const rows = this.rows();
const index = Math.trunc(n) - 1;
if (index < 0 || index >= rows.length) return false;
this.moveTo(rows[index].session.sessionId);
return true;
}
private moveTo(sessionId: string | null): void {
if (this.selectedId === sessionId) return;
this.selectedId = sessionId;
this.touch();
}
private indexOfSelected(rows: readonly TuiRow[] = this.rows()): number {
if (!this.selectedId) return -1;
return rows.findIndex((row) => row.session.sessionId === this.selectedId);
}
/**
* Run a data mutation and keep the cursor sane afterwards: the selected
* session stays selected wherever it moved to, and a session that vanished
* hands the cursor to whatever now occupies its place.
*/
private mutate(apply: () => void): void {
const previousIndex = this.indexOfSelected();
apply();
this.touch();
const rows = this.rows();
if (rows.length === 0) {
this.moveTo(null);
return;
}
if (this.selectedId !== null && rows.some((row) => row.session.sessionId === this.selectedId)) return;
const index = Math.min(Math.max(previousIndex, 0), rows.length - 1);
this.moveTo(rows[index].session.sessionId);
}
}
export function createTuiModel(options: GroupOptions = {}): TuiModelStore {
return new TuiModelStore(options);
}
+921
View File
@@ -0,0 +1,921 @@
/**
* @fileoverview Pure frame renderer: model + layout in, one string out.
*
* The frame is absolute-addressed, one `ESC [ <row>;1 H` per line followed by
* `ESC [ K`, so nothing ever scrolls and a repaint cannot leave debris. The
* caller wraps the result in synchronized-output brackets (DECSET 2026) where
* the terminal supports it; that is an IO decision and stays out of here.
*
* Color is decided by the caller and passed in, never detected here: chalk's
* auto-detection is the right answer for the one-shot CLI (see `cli-style.ts`)
* but it would make a frame non-deterministic, and "same inputs, same string"
* is what makes this module testable. The palette below is the same semantic
* vocabulary chalk gives `cli-style` (ok green, warn yellow, err red, info
* cyan, muted gray, emph bold), written as raw SGR so the mapping is fixed.
*
* With `color: false` the frame contains no escape sequences at all beyond the
* cursor addressing that puts each line in place.
*
* @module tui/tui-render
*/
import { clipStyledLine, padDisplay, stripStyles, visibleWidth } from './tui-ansi.js';
import { approvalCard } from './tui-approvals.js';
import { composerText, composerWindow } from './tui-composer.js';
import type { TuiLayout, TuiRect } from './tui-layout.js';
import type { ApprovalItem } from '../web/approval-inbox.js';
import type { StatusTelemetry } from '../usage-telemetry.js';
import type {
TuiDigestState,
TuiGlyphTier,
TuiGroup,
TuiPickerState,
TuiPromptState,
TuiRenderModel,
TuiRow,
TuiSearchState,
TuiSessionRow,
TuiSessionState,
} from './tui-types.js';
export interface TuiRenderOptions {
/** Emit SGR color. False is NO_COLOR: cursor addressing and nothing else. */
color: boolean;
glyphs: TuiGlyphTier;
/** Animation counter. The WORKING glyph cycles with it. */
tick: number;
/** Wall clock for elapsed times, passed in so a frame is reproducible. */
now: number;
/**
* Footer entries, already labelled, joined here with the separator glyph.
* The app layer passes the keys that actually do something right now (which
* verbs are wired up, whether a server is answering); omitting it falls back
* to the full keymap below.
*/
footerKeys?: readonly string[];
/**
* `[key, what it does]` pairs for the help overlay, same reasoning as
* `footerKeys`: the app layer knows which verbs are wired up. Omitting it
* falls back to the full keymap.
*/
helpKeys?: ReadonlyArray<readonly [string, string]>;
}
// ─────────────────────────────────────────────────────────────────────────────
// Palette and glyphs
// ─────────────────────────────────────────────────────────────────────────────
const SGR = {
reset: '\x1b[0m',
bold: '\x1b[1m',
dim: '\x1b[2m',
inverse: '\x1b[7m',
red: '\x1b[31m',
green: '\x1b[32m',
yellow: '\x1b[33m',
magenta: '\x1b[35m',
cyan: '\x1b[36m',
gray: '\x1b[90m',
} as const;
/**
* One word per state, shared by the preview title and the `--list` output so
* both surfaces call a session the same thing.
*/
export const STATE_WORDS: Record<TuiSessionState, string> = {
'blocked-permission': 'blocked',
'blocked-question': 'blocked',
waiting: 'waiting',
working: 'working',
idle: 'idle',
recent: 'done',
};
const STATE_COLOR: Record<TuiSessionState, string> = {
'blocked-permission': SGR.red,
'blocked-question': SGR.red,
waiting: SGR.yellow,
working: SGR.green,
idle: SGR.gray,
recent: SGR.gray,
};
export interface TuiGlyphSet {
blockedPermission: string;
blockedQuestion: string;
waiting: string;
/** WORKING animates through Claude's own glyph family, a deliberate nod. */
working: readonly string[];
idle: string;
recent: string;
cursor: string;
rule: string;
divider: string;
boxTopLeft: string;
boxTopRight: string;
boxBottomLeft: string;
boxBottomRight: string;
boxHorizontal: string;
boxVertical: string;
enter: string;
updown: string;
separator: string;
ellipsis: string;
}
/**
* ⚠️ Every glyph here must clear TWO bars that are easy to miss, and both were
* failed at once by the first version of this table.
*
* WIDTH: the renderer addresses cells by column, so a glyph the terminal draws
* two cells wide shifts everything after it. `east_asian_width` W or F is
* therefore disqualifying. `✋` (U+270B) was Wide, and being an emoji is also
* why fonts render it at emoji size in the middle of a text row.
*
* COVERAGE: a plain terminal font carries far less than the unicode TIER
* implies. The tier answers "is the locale UTF-8", which says nothing about
* whether a given codepoint has a glyph.
*
* One beta tester's font mapped the blocks like this, and it is the profile to
* design against because it is an ordinary terminal font, not a broken one:
*
* RENDERS Latin-1 (·), Box Drawing (─ │), Block Elements (█ ▛ ▐),
* Geometric Shapes (○ ▶), General Punctuation (…), Arrows
* TOFU Misc Technical (⏎ U+23CE, ⏵ U+23F5), the sparse end of
* Dingbats (❯ U+276F)
*
* So: draw from the blocks on the first line. Dingbats, Miscellaneous
* Technical, Miscellaneous Symbols and anything with emoji presentation are
* out — that class produced three separate "why are there boxes" reports, one
* per glyph, because each was fixed on its own instead of as a class.
*/
const UNICODE_GLYPHS: TuiGlyphSet = {
blockedPermission: '▲',
blockedQuestion: '▲',
waiting: '!',
// Quadrant blocks, which rotate as a spinner and live in the same block as
// the `▛█▐` art claude itself draws — proven to render on the font that
// failed the dingbats this used to use.
working: ['▖', '▘', '▝', '▗'],
idle: '○',
recent: '✔',
cursor: '▶',
rule: '─',
divider: '│',
boxTopLeft: '┌',
boxTopRight: '┐',
boxBottomLeft: '└',
boxBottomRight: '┘',
boxHorizontal: '─',
boxVertical: '│',
enter: '↵',
updown: '↑↓',
separator: '·',
ellipsis: '…',
};
/**
* The lowest tier, for terminals that are not known-capable. Every state token
* is three columns wide so rows still line up, mirroring what `sc` falls back
* to today.
*/
const ASCII_GLYPHS: TuiGlyphSet = {
blockedPermission: '[!]',
blockedQuestion: '[?]',
waiting: '[w]',
working: ['[*]', '[+]', '[x]', '[+]'],
idle: '[-]',
recent: '[v]',
cursor: '>',
rule: '-',
divider: '|',
boxTopLeft: '+',
boxTopRight: '+',
boxBottomLeft: '+',
boxBottomRight: '+',
boxHorizontal: '-',
boxVertical: '|',
enter: 'enter',
updown: 'up/dn',
separator: '-',
ellipsis: '..',
};
/**
* Glyphs for a tier. `nerd` currently renders like `unicode`: the tier exists
* so detection has somewhere to land and a nerd-font-only set has a home,
* without shipping glyphs nobody has reviewed on a real font.
*/
export function glyphsFor(tier: TuiGlyphTier): TuiGlyphSet {
return tier === 'ascii' ? ASCII_GLYPHS : UNICODE_GLYPHS;
}
/**
* Glyph tier from the environment. IO-ish by nature (it reads env), so it takes
* the env as an argument and the app layer calls it once at startup. The
* known-capable list is a TERM allowlist, plus a UTF-8 locale check and an
* explicit override.
*/
export function detectGlyphTier(env: Record<string, string | undefined>): TuiGlyphTier {
const override = env.CODEMAN_TUI_GLYPHS;
if (override === 'ascii' || override === 'unicode' || override === 'nerd') return override;
const term = env.TERM ?? '';
if (term === '' || term === 'dumb') return 'ascii';
const locale = env.LC_ALL || env.LC_CTYPE || env.LANG || '';
if (!/utf-?8/i.test(locale)) return 'ascii';
const termProgram = env.TERM_PROGRAM ?? '';
if (termProgram.startsWith('iTerm') || term === 'xterm-kitty' || env.WEZTERM_PANE || env.LC_TERMINAL === 'iTerm2') {
return 'nerd';
}
return 'unicode';
}
// ─────────────────────────────────────────────────────────────────────────────
// Formatting helpers (pure, exported for tests and for the app layer)
// ─────────────────────────────────────────────────────────────────────────────
/** Compact age: `45s`, `11m`, `2h`, `3d`. Empty when the anchor is unknown. */
export function formatElapsed(ms: number): string {
if (!Number.isFinite(ms) || ms < 0) return '';
const seconds = Math.floor(ms / 1000);
if (seconds < 60) return `${seconds}s`;
const minutes = Math.floor(seconds / 60);
if (minutes < 60) return `${minutes}m`;
const hours = Math.floor(minutes / 60);
if (hours < 24) return `${hours}h`;
return `${Math.floor(hours / 24)}d`;
}
function trimTrailingZero(value: string): string {
return value.endsWith('.0') ? value.slice(0, -2) : value;
}
/** Compact token count: `842`, `45.2k`, `1.2M`. Empty when there is nothing to show. */
export function formatTokens(total: number): string {
if (!Number.isFinite(total) || total <= 0) return '';
if (total < 1000) return String(Math.floor(total));
if (total < 1_000_000) return `${trimTrailingZero((total / 1000).toFixed(1))}k`;
return `${trimTrailingZero((total / 1_000_000).toFixed(1))}M`;
}
/**
* The header's plan-usage chip: `5h 32% · wk 61%`, the same two windows the web
* chip shows (the statusline telemetry carries no others). Empty when the
* account reports neither, so the header shows no placeholder for a fact that
* does not exist. The separator is passed in because the header's own comes
* from the glyph tier, and an ASCII terminal must not get a stray `·`.
*/
export function formatPlanUsage(usage: StatusTelemetry | null | undefined, separator = ' · '): string {
if (!usage) return '';
const parts: string[] = [];
if (typeof usage.fiveHour?.usedPercentage === 'number') {
parts.push(`5h ${Math.round(usage.fiveHour.usedPercentage)}%`);
}
if (typeof usage.sevenDay?.usedPercentage === 'number') {
parts.push(`wk ${Math.round(usage.sevenDay.usedPercentage)}%`);
}
return parts.join(separator);
}
/**
* What a row is called. Same rule as the web history rows, including the
* "(no content)" placeholder the transcript reader emits, which is not a title.
*/
export function rowLabel(session: TuiSessionRow): string {
if (session.name) return session.name;
const base = (session.workingDir ?? '').split('/').filter(Boolean).pop();
// ⚠️ A LIVE pane (it has a mux name) is identified by WHERE it runs, never by
// a line scraped out of its transcript. A session created before the user has
// typed anything has no prompt to be named after, so the fallback took
// whatever the CLI happened to print first: a beta tester's new session
// appeared in the list called "Login interrupted", which reads like a failure
// report and was in fact a healthy session. A history row is the opposite
// case, where the prompt IS the identity, so it keeps the old order.
if (session.muxName && base) return base;
const prompt = (session.firstPrompt ?? '').trim();
if (prompt && prompt !== '(no content)') return prompt;
return base || session.sessionId.slice(0, 8);
}
/** Keep the tail of a path: the last segments identify it, the root never does. */
function truncatePathLeft(path: string, width: number, ellipsis: string): string {
if (width <= 0) return '';
if (visibleWidth(path) <= width) return path;
const keep = Math.max(0, width - visibleWidth(ellipsis));
return ellipsis + path.slice(path.length - keep);
}
function tokensOf(session: TuiSessionRow): number {
return (session.inputTokens ?? 0) + (session.outputTokens ?? 0);
}
// ─────────────────────────────────────────────────────────────────────────────
// Painting
// ─────────────────────────────────────────────────────────────────────────────
type Painter = (text: string, code: string) => string;
function painterFor(enabled: boolean): Painter {
return enabled ? (text, code) => (text === '' ? text : `${code}${text}${SGR.reset}`) : (text) => text;
}
function stateGlyph(row: TuiRow, glyphs: TuiGlyphSet, tick: number): string {
switch (row.state) {
case 'blocked-permission':
return glyphs.blockedPermission;
case 'blocked-question':
return glyphs.blockedQuestion;
case 'waiting':
return glyphs.waiting;
case 'working': {
const frames = glyphs.working;
const index = ((Math.trunc(tick) % frames.length) + frames.length) % frames.length;
return frames[index];
}
case 'idle':
return glyphs.idle;
case 'recent':
return glyphs.recent;
}
}
function centered(text: string, width: number): string {
const pad = Math.max(0, Math.floor((width - visibleWidth(text)) / 2));
return padDisplay(`${' '.repeat(pad)}${text}`, width);
}
// ─────────────────────────────────────────────────────────────────────────────
// Rows and groups
// ─────────────────────────────────────────────────────────────────────────────
interface RowContext {
width: number;
/** 1-based position in the flattened list; only 1-9 get a jump digit. */
index: number;
selected: boolean;
twoLine: boolean;
glyphs: TuiGlyphSet;
opts: TuiRenderOptions;
}
function renderRowLines(row: TuiRow, ctx: RowContext): string[] {
// A selected row is one inverse-video block, so its parts are built unpainted:
// an inner reset would punch a hole in the highlight.
const inverse = ctx.selected && ctx.opts.color;
const paint = painterFor(ctx.opts.color && !inverse);
const { session } = row;
const marker = ctx.selected ? padDisplay(ctx.glyphs.cursor, 2) : ' ';
const digit = ctx.index >= 1 && ctx.index <= 9 ? `${ctx.index} ` : ' ';
const glyph = paint(stateGlyph(row, ctx.glyphs, ctx.opts.tick), STATE_COLOR[row.state]);
const elapsed = row.since > 0 ? formatElapsed(ctx.opts.now - row.since) : '';
const tokens = formatTokens(tokensOf(session));
const rightParts = [glyph, paint(elapsed, SGR.gray)];
if (!ctx.twoLine && tokens) rightParts.push(paint(tokens, SGR.gray));
const right = rightParts.filter((part) => part !== '').join(' ');
const mode = session.mode && session.mode !== 'claude' ? session.mode : '';
const nameWidth = Math.max(4, ctx.width - visibleWidth(marker + digit) - visibleWidth(right) - 1);
const label = rowLabel(session);
const name = mode ? `${label} ${paint(mode, SGR.magenta)}` : label;
const first = padDisplay(`${marker}${digit}${padDisplay(name, nameWidth)} ${right}`, ctx.width);
const lines = [first];
if (ctx.twoLine) {
const detail = [truncatePathLeft(session.workingDir ?? '', Math.max(0, ctx.width - 8), ctx.glyphs.ellipsis)];
if (mode) detail.push(mode);
if (tokens) detail.push(tokens);
const text = detail.filter((part) => part !== '').join(` ${ctx.glyphs.separator} `);
lines.push(padDisplay(` ${paint(text, SGR.gray)}`, ctx.width));
}
return inverse ? lines.map((line) => `${SGR.inverse}${line}${SGR.reset}`) : lines;
}
function renderGroupHeader(group: TuiGroup, width: number, glyphs: TuiGlyphSet, opts: TuiRenderOptions): string {
const paint = painterFor(opts.color);
const label = ` ${group.label} `;
const fill = Math.max(0, width - visibleWidth(label));
return padDisplay(`${paint(label, SGR.bold)}${paint(glyphs.rule.repeat(fill), SGR.gray)}`, width);
}
export interface TuiListEntry {
text: string;
/** Set on the lines that belong to a session row, so the window can chase the cursor. */
sessionId?: string;
}
function buildListEntries(model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): TuiListEntry[] {
const glyphs = glyphsFor(opts.glyphs);
const width = layout.list.width;
const entries: TuiListEntry[] = [];
let index = 0;
for (const group of model.groups()) {
if (group.rows.length === 0) continue;
entries.push({ text: renderGroupHeader(group, width, glyphs, opts) });
for (const row of group.rows) {
index++;
const ctx: RowContext = {
width,
index,
selected: row.session.sessionId === model.selectedId,
twoLine: layout.rowHeight === 2,
glyphs,
opts,
};
for (const text of renderRowLines(row, ctx)) entries.push({ text, sessionId: row.session.sessionId });
}
}
return entries;
}
/**
* First visible entry, scrolling the minimum needed to keep the selected row on
* screen. Deterministic on purpose: the window is derived, never remembered, so
* two identical models render identically.
*/
export function computeListWindow(
entries: readonly TuiListEntry[],
capacity: number,
selectedId: string | null
): number {
if (capacity <= 0 || entries.length <= capacity) return 0;
const maxStart = entries.length - capacity;
if (!selectedId) return 0;
const first = entries.findIndex((entry) => entry.sessionId === selectedId);
if (first < 0) return 0;
let last = first;
while (last + 1 < entries.length && entries[last + 1].sessionId === selectedId) last++;
let start = 0;
if (last >= capacity) start = Math.min(last - capacity + 1, maxStart);
if (first < start) start = first;
return start;
}
// ─────────────────────────────────────────────────────────────────────────────
// Preview
// ─────────────────────────────────────────────────────────────────────────────
/**
* The pending dialog, drawn above the tail: the question, the parsed options
* with their digits, and the keys that answer them. Red for a permission or
* question prompt, yellow for an idle one, the same severity vocabulary the web
* inbox uses.
*/
export function renderApprovalCard(
item: ApprovalItem,
width: number,
glyphs: TuiGlyphSet,
opts: TuiRenderOptions
): string[] {
const paint = painterFor(opts.color);
const card = approvalCard(item);
const color = card.tone === 'err' ? SGR.red : SGR.yellow;
const glyph = card.tone === 'err' ? glyphs.blockedPermission : glyphs.waiting;
const lines: string[] = [];
const push = (text: string, style: string): void => {
lines.push(padDisplay(paint(clipStyledLine(text, width), style), width));
};
push(` ${glyph} ${card.title}`, color);
for (const detail of card.detail) push(` ${detail}`, SGR.gray);
for (const option of card.options) push(` ${option.n}. ${option.label}`, '');
push(` ${card.hint}`, SGR.gray);
return lines;
}
/** The card may take half the pane at most: the tail is why the pane exists. */
function cardCapacity(height: number): number {
return Math.max(0, Math.floor((height - 1) / 2));
}
/**
* `name · mode · dir · state`, with the DIRECTORY absorbing the squeeze: the
* state word is the one fact the pane exists to confirm, so it must survive a
* narrow preview that a full path would push off the end.
*/
function previewTitle(row: TuiRow, width: number, glyphs: TuiGlyphSet): string {
const { session } = row;
const sep = ` ${glyphs.separator} `;
const head = ` ${rowLabel(session)}${sep}${session.mode ?? 'claude'}`;
const tail = `${sep}${STATE_WORDS[row.state]}`;
const dirBudget = width - visibleWidth(head) - visibleWidth(tail) - visibleWidth(sep);
const dir = session.workingDir ? truncatePathLeft(session.workingDir, Math.max(0, dirBudget), glyphs.ellipsis) : '';
return clipStyledLine(dir ? `${head}${sep}${dir}${tail}` : `${head}${tail}`, width);
}
function buildPreviewLines(model: TuiRenderModel, rect: TuiRect, opts: TuiRenderOptions): string[] {
const paint = painterFor(opts.color);
const glyphs = glyphsFor(opts.glyphs);
const lines: string[] = [];
const selected = model.selectedId
? (model
.groups()
.flatMap((group) => group.rows)
.find((row) => row.session.sessionId === model.selectedId) ?? null)
: null;
if (!selected) {
lines.push(padDisplay(paint(' no session selected', SGR.gray), rect.width));
} else {
lines.push(padDisplay(paint(previewTitle(selected, rect.width, glyphs), SGR.bold), rect.width));
}
const budget = cardCapacity(rect.height);
if (selected?.approval && budget > 0) {
for (const line of renderApprovalCard(selected.approval, rect.width, glyphs, opts).slice(0, budget)) {
lines.push(line);
}
if (lines.length < rect.height) lines.push(' '.repeat(rect.width));
}
const body = previewBody(model, selected, rect, opts, rect.height - lines.length);
for (const line of body) lines.push(line);
while (lines.length < rect.height) lines.push(' '.repeat(rect.width));
return lines.slice(0, Math.max(0, rect.height));
}
function previewBody(
model: TuiRenderModel,
selected: TuiRow | null,
rect: TuiRect,
opts: TuiRenderOptions,
capacity: number
): string[] {
const paint = painterFor(opts.color);
if (capacity <= 0) return [];
const hint = (text: string): string[] => [padDisplay(paint(` ${text}`, SGR.gray), rect.width)];
if (!selected) return [];
if (model.connection === 'degraded' || model.connection === 'down') {
return hint('preview unavailable while the server is down');
}
const preview = model.preview;
if (!preview || preview.sessionId !== selected.session.sessionId) return hint('loading preview…');
if (preview.note) return hint(preview.note);
if (preview.error) return hint(preview.error);
const trimmed = [...preview.lines];
while (trimmed.length > 0 && trimmed[trimmed.length - 1].trim() === '') trimmed.pop();
if (trimmed.length === 0) return hint('(no output yet)');
// The tail carries the session's OWN colors, which is the point of the pane,
// but under NO_COLOR they must go too.
return trimmed
.slice(-capacity)
.map((line) => padDisplay(` ${clipStyledLine(opts.color ? line : stripStyles(line), rect.width - 1)}`, rect.width));
}
// ─────────────────────────────────────────────────────────────────────────────
// Chrome
// ─────────────────────────────────────────────────────────────────────────────
/** Sessions with a prompt waiting on a human, which is what the badge counts. */
export function pendingApprovalCount(model: TuiRenderModel): number {
let count = 0;
for (const group of model.groups()) for (const row of group.rows) if (row.approval) count++;
return count;
}
function renderHeaderLine(model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): string {
const paint = painterFor(opts.color);
const glyphs = glyphsFor(opts.glyphs);
const { hostname, instance, version, planUsage } = model.header;
const facts = [
instance ? `${hostname ?? ''}:${instance}` : (hostname ?? ''),
version ? `v${version}` : '',
`${model.sessionCount} session${model.sessionCount === 1 ? '' : 's'}`,
planUsage ?? '',
].filter((part) => part !== '');
const pending = pendingApprovalCount(model);
const badge = pending > 0 ? `${paint(`${glyphs.blockedPermission} ${pending}`, SGR.red)} ` : '';
const left = ` ${paint('codeman', SGR.bold)} ${badge}${paint(facts.join(` ${glyphs.separator} `), SGR.gray)}`;
const right = paint('? help q quit ', SGR.gray);
const gap = layout.cols - visibleWidth(left) - visibleWidth(right);
if (gap < 1) return padDisplay(left, layout.cols);
return `${left}${' '.repeat(gap)}${right}`;
}
function renderBannerLine(model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): string {
const paint = painterFor(opts.color);
const glyphs = glyphsFor(opts.glyphs);
const [text, color] =
model.connection === 'degraded'
? ['server not running: attach only', SGR.yellow]
: model.connection === 'reconnecting'
? ['reconnecting to the server…', SGR.yellow]
: ['server unreachable', SGR.red];
return padDisplay(paint(` ${glyphs.blockedPermission} ${text}`, color), layout.cols);
}
const FOOTER_KEYS: Record<string, (glyphs: TuiGlyphSet) => string> = {
list: (g) =>
[
`${g.updown} select`,
`${g.enter} attach`,
'1-9 switch',
'y/n answer',
'p prompt',
'n new',
'x kill',
'/ search',
'g digest',
'? help',
'q quit',
].join(` ${g.separator} `),
help: (g) => `esc ${g.separator} ? close`,
'confirm-kill': (g) => `y kill ${g.separator} any other key cancels`,
message: () => 'esc dismiss',
prompt: (g) => `${g.enter} send ${g.separator} esc cancel`,
search: (g) => `${g.updown} results ${g.separator} ${g.enter} open ${g.separator} esc close`,
digest: (g) => `j/k ${g.separator} ${g.updown} scroll ${g.separator} esc close`,
'new-session': (g) => `${g.updown} select ${g.separator} ${g.enter} choose ${g.separator} esc cancel`,
};
/**
* The composer's prefix. Fixed width on purpose: the terminal cursor is placed
* by column arithmetic (`composerCursorCell`), and a prefix that changed with
* the session name would move the cursor with it.
*/
export const COMPOSER_PREFIX = ' > ';
function renderComposerLine(prompt: TuiPromptState, layout: TuiLayout, opts: TuiRenderOptions): string {
const paint = painterFor(opts.color);
const window = composerWindow(prompt.composer, Math.max(1, layout.cols - visibleWidth(COMPOSER_PREFIX)));
return padDisplay(`${paint(COMPOSER_PREFIX, SGR.cyan)}${window.text}`, layout.cols);
}
/**
* Where the terminal's own cursor belongs, or null when nothing is being typed
* into a single-line editor. The app shows the cursor there and hides it
* otherwise, because a blinking cursor parked in a dashboard reads as a bug.
*/
export function composerCursorCell(model: TuiRenderModel, layout: TuiLayout): { row: number; col: number } | null {
if (model.mode !== 'prompt' || !model.prompt || layout.footer.height <= 0) return null;
const prefix = visibleWidth(COMPOSER_PREFIX);
const window = composerWindow(model.prompt.composer, Math.max(1, layout.cols - prefix));
return { row: layout.footer.row, col: Math.min(layout.cols, prefix + 1 + window.cursorColumn) };
}
function renderFooterLine(model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): string {
const paint = painterFor(opts.color);
const glyphs = glyphsFor(opts.glyphs);
if (model.mode === 'prompt' && model.prompt) return renderComposerLine(model.prompt, layout, opts);
const text = opts.footerKeys
? opts.footerKeys.join(` ${glyphs.separator} `)
: (FOOTER_KEYS[model.mode] ?? FOOTER_KEYS.list)(glyphs);
return padDisplay(paint(clipStyledLine(` ${text}`, layout.cols), SGR.gray), layout.cols);
}
// ─────────────────────────────────────────────────────────────────────────────
// Overlays
// ─────────────────────────────────────────────────────────────────────────────
interface OverlayContent {
title: string;
lines: string[];
/**
* Floor for the box's inner width. The search and digest panels are lists
* people scan, so they keep a stable width instead of snapping around their
* longest current line.
*/
minWidth?: number;
}
function wrapText(text: string, width: number): string[] {
if (width <= 0) return [];
const out: string[] = [];
let line = '';
for (const word of text.split(/\s+/).filter((part) => part !== '')) {
const candidate = line === '' ? word : `${line} ${word}`;
if (visibleWidth(candidate) > width && line !== '') {
out.push(line);
line = word;
} else {
line = candidate;
}
}
if (line !== '') out.push(line);
return out.length > 0 ? out : [''];
}
function helpLines(glyphs: TuiGlyphSet, custom?: ReadonlyArray<readonly [string, string]>): string[] {
const pairs: ReadonlyArray<readonly [string, string]> = custom ?? [
[`${glyphs.updown} / j k`, 'select'],
[glyphs.enter, 'attach'],
['1-9', 'jump'],
['y / n', 'answer the pending approval'],
['p', 'send a prompt'],
['n', 'new session'],
['x', 'kill (typed confirmation)'],
['/', 'search'],
['g', 'away digest'],
['?', 'this help'],
['q', 'quit'],
];
const keyWidth = Math.max(...pairs.map(([key]) => visibleWidth(key)));
return pairs.map(([key, description]) => `${padDisplay(key, keyWidth)} ${description}`);
}
/** Longest item list a picker overlay shows, however tall the terminal is. */
const PICKER_MAX_ROWS = 10;
/**
* A picker's lines: hint, a window of items around the cursor, then the filter
* echo. Windowed rather than clipped, so the selected item is always visible in
* a long case list.
*/
function pickerLines(picker: TuiPickerState, glyphs: TuiGlyphSet, capacity: number): string[] {
const head: string[] = picker.hint ? [picker.hint, ''] : [];
const tail: string[] = picker.filter === undefined ? [] : ['', `filter: ${picker.filter}_`];
if (picker.items.length === 0) return [...head, '(nothing to choose)', ...tail];
const budget = Math.max(1, Math.min(PICKER_MAX_ROWS, capacity - head.length - tail.length));
const first = Math.max(0, Math.min(picker.index - Math.floor(budget / 2), picker.items.length - budget));
const rows = picker.items.slice(first, first + budget).map((item, i) => {
const marker = first + i === picker.index ? glyphs.cursor : ' '.repeat(visibleWidth(glyphs.cursor));
return `${marker} ${item.label}${item.detail ? ` ${item.detail}` : ''}`;
});
return [...head, ...rows, ...tail];
}
/**
* The `/` overlay: the query with a caret, one status line, then the results.
*
* The caret is a trailing `_` rather than the terminal's own cursor, and that is
* why the search keymap leaves left/right to the result list: a caret that
* cannot move is honest, an invisible one that can is not.
*/
function searchLines(state: TuiSearchState, glyphs: TuiGlyphSet, capacity: number): string[] {
const head = [`${composerText(state.composer)}_`];
if (state.note) head.push(state.note);
head.push('');
const budget = Math.max(1, capacity - head.length);
if (state.entries.length === 0) {
return [...head, state.status === 'searching' ? 'searching…' : '(type to search sessions, events and files)'];
}
const first = Math.max(0, Math.min(state.index - Math.floor(budget / 2), state.entries.length - budget));
const rows = state.entries.slice(first, first + budget).map((entry, i) => {
if (entry.kind === 'header') return entry.text;
const marker = first + i === state.index ? glyphs.cursor : ' '.repeat(visibleWidth(glyphs.cursor));
return `${marker} ${entry.text}${entry.detail ? ` ${entry.detail}` : ''}`;
});
return [...head, ...rows];
}
/** Lines an overlay box can show inside its border, given the body's height. */
function overlayCapacity(height: number): number {
return Math.max(1, height - 2);
}
/**
* How many digest lines fit. Exported because the app scrolls by pages and must
* not scroll the last page into empty space, which needs this exact number.
*/
export function digestCapacity(layout: TuiLayout): number {
return overlayCapacity(layout.body.height);
}
function digestLines(state: TuiDigestState, capacity: number): string[] {
const offset = Math.min(Math.max(0, state.offset), Math.max(0, state.lines.length - 1));
return state.lines.slice(offset, offset + capacity);
}
function overlayContent(
model: TuiRenderModel,
opts: TuiRenderOptions,
width: number,
height: number
): OverlayContent | null {
const glyphs = glyphsFor(opts.glyphs);
const panelWidth = Math.max(20, Math.min(width - 8, 72));
switch (model.mode) {
case 'help':
return { title: 'Keys', lines: helpLines(glyphs, opts.helpKeys) };
case 'search': {
if (!model.search) return null;
return {
title: 'Search',
lines: searchLines(model.search, glyphs, overlayCapacity(height)),
minWidth: panelWidth,
};
}
case 'digest': {
if (!model.digest) return null;
return {
title: model.digest.title,
lines: digestLines(model.digest, overlayCapacity(height)),
minWidth: panelWidth,
};
}
case 'new-session': {
if (!model.picker) return null;
return { title: model.picker.title, lines: pickerLines(model.picker, glyphs, Math.max(1, height - 2)) };
}
case 'confirm-kill': {
if (!model.confirm) return null;
return {
title: 'Kill session',
lines: [`Kill ${model.confirm.name}?`, '', 'press y to kill, any other key cancels'],
};
}
case 'message':
if (!model.message) return null;
return {
title: model.message.tone === 'err' ? 'Error' : model.message.tone === 'warn' ? 'Warning' : 'Notice',
lines: wrapText(model.message.text, Math.max(8, width - 8)),
};
default:
return null;
}
}
/** Paint an overlay box over the body, centered, replacing whole terminal rows. */
function applyOverlay(lines: string[], model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): void {
const body = layout.body;
if (body.height < 3 || body.width < 12) return;
const content = overlayContent(model, opts, body.width, body.height);
if (!content) return;
const paint = painterFor(opts.color);
const glyphs = glyphsFor(opts.glyphs);
const maxInner = body.width - 4;
const visible = content.lines.slice(0, Math.max(1, body.height - 2));
const inner = Math.min(
maxInner,
Math.max(content.minWidth ?? 0, visibleWidth(content.title) + 2, ...visible.map((line) => visibleWidth(line)))
);
const boxWidth = inner + 4;
const boxHeight = visible.length + 2;
const left = body.col + Math.max(0, Math.floor((body.width - boxWidth) / 2));
const top = body.row + Math.max(0, Math.floor((body.height - boxHeight) / 2));
const titleText = ` ${content.title} `;
const titleFill = Math.max(0, inner + 2 - visibleWidth(titleText));
const boxLines = [
`${glyphs.boxTopLeft}${titleText}${glyphs.boxHorizontal.repeat(titleFill)}${glyphs.boxTopRight}`,
...visible.map((line) => `${glyphs.boxVertical} ${padDisplay(line, inner)} ${glyphs.boxVertical}`),
`${glyphs.boxBottomLeft}${glyphs.boxHorizontal.repeat(inner + 2)}${glyphs.boxBottomRight}`,
];
for (let i = 0; i < boxLines.length; i++) {
const row = top + i - 1;
if (row < 0 || row >= lines.length) continue;
lines[row] = `${' '.repeat(left - 1)}${paint(boxLines[i], SGR.cyan)}`;
}
}
// ─────────────────────────────────────────────────────────────────────────────
// Frame
// ─────────────────────────────────────────────────────────────────────────────
function writeBody(lines: string[], model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): void {
const { list, preview, divider } = layout;
if (list.height <= 0) return;
const paint = painterFor(opts.color);
const glyphs = glyphsFor(opts.glyphs);
const entries = buildListEntries(model, layout, opts);
if (entries.length === 0) {
const hint = paint('No sessions. n to start one, q to quit.', SGR.gray);
const row = list.row + Math.floor((list.height - 1) / 2);
lines[row - 1] = centered(hint, layout.cols);
return;
}
const start = computeListWindow(entries, list.height, model.selectedId);
const previewLines = preview ? buildPreviewLines(model, preview, opts) : [];
for (let i = 0; i < list.height; i++) {
const left = entries[start + i]?.text ?? ' '.repeat(list.width);
if (!preview || !divider) {
lines[list.row - 1 + i] = left;
continue;
}
const right = previewLines[i] ?? ' '.repeat(preview.width);
lines[list.row - 1 + i] = `${left}${paint(glyphs.divider, SGR.gray)}${right}`;
}
}
/**
* The whole frame as one string: absolute cursor addressing per line, each line
* closed with an erase-to-end so a shorter line cannot leave the previous
* frame's tail behind.
*/
export function renderFrame(model: TuiRenderModel, layout: TuiLayout, opts: TuiRenderOptions): string {
const lines: string[] = new Array<string>(layout.rows).fill('');
lines[0] = renderHeaderLine(model, layout, opts);
if (layout.banner) lines[layout.banner.row - 1] = renderBannerLine(model, layout, opts);
writeBody(lines, model, layout, opts);
if (layout.footer.height > 0) lines[layout.footer.row - 1] = renderFooterLine(model, layout, opts);
applyOverlay(lines, model, layout, opts);
let frame = '';
for (let i = 0; i < lines.length; i++) {
frame += `\x1b[${i + 1};1H${clipStyledLine(lines[i], layout.cols)}\x1b[K`;
}
return frame;
}
+234
View File
@@ -0,0 +1,234 @@
/**
* @fileoverview Pure SSE wire parsing, event classification and reconnect math.
*
* Node has no `EventSource`, so the TUI reads `GET /api/events` as a raw stream
* and decodes the wire format here. Everything in this module is pure: bytes
* (as decoded strings) in, frames out. The socket, the timers and the backoff
* loop live in `tui-client.ts`.
*
* Three wire details this parser exists to get right:
*
* 1. **Frames split across chunk boundaries.** A TCP read can end anywhere,
* including between the `\r` and the `\n` of a CRLF, so a lone trailing
* `\r` is held back rather than treated as a line end.
* 2. **Comments are not frames.** The server appends a `:pppp…` padding line
* after a frame while a Cloudflare tunnel is up (it flushes the proxy
* buffer) and that line carries no blank line after it. Dispatch happens on
* a blank line and on nothing else, so padding cannot split a frame.
* 3. **The keepalive is a NAMED event** (`sse:heartbeat`), because an SSE
* comment is invisible to a browser `EventSource` by spec. We treat ANY
* inbound bytes as liveness, comments included, which is why comments need
* no representation in the returned frames.
*
* @module tui/tui-sse
*/
import {
ApprovalPending,
ApprovalResolved,
ApprovalUpdated,
Heartbeat,
Init,
MuxCreated,
MuxDied,
MuxKilled,
RemoteSessionDropped,
RemoteSessionReconnected,
SessionCliInfo,
SessionCompletion,
SessionCreated,
SessionDeleted,
SessionError,
SessionExit,
SessionIdle,
SessionInteractive,
SessionPinned,
SessionRunning,
SessionStatusTelemetry,
SessionUpdated,
SessionWorking,
} from '../web/sse-events.js';
/** One dispatched SSE frame. `event` defaults to `message` per the spec. */
export interface SseFrame {
event: string;
data: string;
id?: string;
retry?: number;
}
/**
* Ceiling on the unterminated tail the parser will hold. The `init` frame
* carries the whole light state and is legitimately large, so this is not a
* frame-size limit but a guard against a non-SSE endpoint streaming something
* with no line terminators at all.
*/
export const MAX_PENDING_BYTES = 8 * 1024 * 1024;
/** Incremental decoder. One instance per connection; `reset()` on reconnect. */
export class SseFrameParser {
private buffer = '';
private eventName = '';
private dataLines: string[] = [];
private lastId: string | undefined;
private retry: number | undefined;
/** Decode one chunk, returning every frame it completed (possibly none). */
feed(chunk: string): SseFrame[] {
this.buffer += chunk;
const frames: SseFrame[] = [];
let start = 0;
for (let i = 0; i < this.buffer.length; i++) {
const ch = this.buffer[i];
if (ch !== '\n' && ch !== '\r') continue;
// A trailing CR may be the first half of a CRLF the next chunk finishes.
if (ch === '\r' && i === this.buffer.length - 1) break;
const line = this.buffer.slice(start, i);
if (ch === '\r' && this.buffer[i + 1] === '\n') i++;
start = i + 1;
const frame = this.consumeLine(line);
if (frame) frames.push(frame);
}
this.buffer = this.buffer.slice(start);
if (this.buffer.length > MAX_PENDING_BYTES) this.reset();
return frames;
}
/** Drop every partial frame. Called when a connection is torn down. */
reset(): void {
this.buffer = '';
this.eventName = '';
this.dataLines = [];
this.lastId = undefined;
this.retry = undefined;
}
private consumeLine(line: string): SseFrame | null {
if (line === '') return this.dispatch();
if (line.startsWith(':')) return null;
const colon = line.indexOf(':');
const field = colon === -1 ? line : line.slice(0, colon);
let value = colon === -1 ? '' : line.slice(colon + 1);
if (value.startsWith(' ')) value = value.slice(1);
switch (field) {
case 'event':
this.eventName = value;
break;
case 'data':
this.dataLines.push(value);
break;
case 'id':
this.lastId = value;
break;
case 'retry': {
const ms = Number.parseInt(value, 10);
if (Number.isSafeInteger(ms) && ms >= 0) this.retry = ms;
break;
}
default:
break;
}
return null;
}
/**
* A blank line ends a frame. Per the spec an empty data buffer dispatches
* nothing (it still clears the event name), which is what makes a bare
* `event:` line or a stray blank line harmless.
*/
private dispatch(): SseFrame | null {
if (this.dataLines.length === 0) {
this.eventName = '';
return null;
}
const frame: SseFrame = {
event: this.eventName || 'message',
data: this.dataLines.join('\n'),
};
if (this.lastId !== undefined) frame.id = this.lastId;
if (this.retry !== undefined) frame.retry = this.retry;
this.eventName = '';
this.dataLines = [];
return frame;
}
}
/** What the app layer should do with a frame. */
export type SseEventClass = 'init' | 'heartbeat' | 'resync' | 'approval' | 'plan-usage' | 'ignore';
/**
* Events that change WHICH sessions exist or WHAT state they are in.
*
* The TUI never patches a single row from a payload: it re-fetches the unified
* list, which is the only source that also carries history rows, so this set
* only has to answer "is a refetch worth it". `session:terminal` is
* deliberately absent (it is the bulk of the stream and the preview pane pulls
* its own tail), as are the ralph/respawn/subagent/orchestrator families, which
* change nothing the dashboard draws.
*/
const RESYNC_EVENTS: ReadonlySet<string> = new Set<string>([
SessionCreated,
SessionUpdated,
SessionDeleted,
SessionExit,
SessionError,
SessionIdle,
SessionWorking,
SessionCompletion,
SessionInteractive,
SessionRunning,
SessionPinned,
SessionCliInfo,
MuxCreated,
MuxKilled,
MuxDied,
RemoteSessionDropped,
RemoteSessionReconnected,
]);
const APPROVAL_EVENTS: ReadonlySet<string> = new Set<string>([ApprovalPending, ApprovalUpdated, ApprovalResolved]);
/** Which approval event this is, or null when the name is not one. */
export function approvalEventKind(name: string): 'pending' | 'updated' | 'resolved' | null {
if (name === ApprovalPending) return 'pending';
if (name === ApprovalUpdated) return 'updated';
if (name === ApprovalResolved) return 'resolved';
return null;
}
/** Route one event name. Unknown names are ignored, never a resync. */
export function classifySseEvent(name: string): SseEventClass {
if (name === Init) return 'init';
if (name === Heartbeat) return 'heartbeat';
if (APPROVAL_EVENTS.has(name)) return 'approval';
if (name === SessionStatusTelemetry) return 'plan-usage';
if (RESYNC_EVENTS.has(name)) return 'resync';
return 'ignore';
}
/**
* Silence that means the stream is dead even though the socket never errored.
* The server heartbeats every 15s, so three missed beats is the signal.
*/
export const SSE_STALE_TIMEOUT_MS = 45_000;
/** Reconnect delay ceiling. A local server is back in milliseconds, not minutes. */
export const SSE_MAX_BACKOFF_MS = 15_000;
/** First reconnect delay; doubles per consecutive failure up to the ceiling. */
export const SSE_BASE_BACKOFF_MS = 500;
/**
* Delay before reconnect attempt `attempt` (1-based). Deterministic, with no
* jitter on purpose: one client talks to one loopback server, so there is no
* herd to spread out and a reproducible delay is testable.
*/
export function sseBackoffDelay(attempt: number, base = SSE_BASE_BACKOFF_MS, max = SSE_MAX_BACKOFF_MS): number {
const step = Math.max(1, Math.trunc(attempt));
const exponent = Math.min(step - 1, 30);
return Math.min(max, base * 2 ** exponent);
}
+209
View File
@@ -0,0 +1,209 @@
/**
* @fileoverview Shared types for the `codeman tui` pure core.
*
* The TUI is a client of the server, never a second brain: its rows are the
* rows `GET /api/sessions/unified` already returns (`UnifiedSessionItem`) and
* its blocked states are the items `GET /api/approvals` already parsed
* (`ApprovalItem`). Both are imported as TYPES only, so nothing here pulls the
* server, node-pty or the utils barrel into a CLI process.
*
* Everything in `src/tui/*` except `tui-app.ts` / `tui-client.ts` is pure:
* deterministic outputs from inputs, no `process.*`, no timers, no IO.
*
* @module tui/tui-types
*/
import type { UnifiedSessionItem } from '../services/unified-session-service.js';
import type { ApprovalItem } from '../web/approval-inbox.js';
import type { TuiComposerState } from './tui-composer.js';
/**
* A unified-list row plus the few live-only extras the dashboard shows.
*
* The unified list is the spine (it is the only source that carries history
* rows), but it has no token counters and no turn-start stamp, so the client
* merges those from the live session payload (`GET /api/sessions` /
* `session_updated` SSE) when a row is live. History rows simply lack them.
*/
export interface TuiSessionRow extends UnifiedSessionItem {
/**
* Wall-clock ms of the pane's last Enter (`SessionState.lastSubmitAt`). The
* only usable "working since" anchor: a working pane repaints about once a
* second, so its `lastActivityAt` is always "now".
*/
lastSubmitAt?: number;
inputTokens?: number;
outputTokens?: number;
/**
* tmux session name to attach to (`codeman-<first 8 of the id>`).
*
* The unified list does not carry it (no server view merges the mux name into
* a row), so the app layer fills it in from the local tmux enumeration, which
* is also the only thing that proves the pane really exists. A row without one
* cannot be attached: it is either history or a direct-PTY session.
*/
muxName?: string;
}
/**
* Row state, in the web UI's vocabulary so both surfaces read the same.
*
* There is deliberately no `error` member: an errored session is something a
* human has to look at, so it classifies as `waiting` and lands in NEEDS YOU
* rather than growing a fifth color nobody designed.
*/
export type TuiSessionState = 'blocked-question' | 'blocked-permission' | 'waiting' | 'working' | 'idle' | 'recent';
/** The four display groups, in display order. */
export type TuiGroupKey = 'needs-you' | 'working' | 'idle' | 'recent';
/** A classified session: what the cursor moves over and the renderer paints. */
export interface TuiRow {
session: TuiSessionRow;
state: TuiSessionState;
group: TuiGroupKey;
/** The pending prompt that blocks this session, when it has one. */
approval?: ApprovalItem;
/** Epoch ms the session entered `state`; the intra-group sort key. 0 when unknown. */
since: number;
}
export interface TuiGroup {
key: TuiGroupKey;
label: string;
rows: TuiRow[];
}
/** How the client currently sees the server. */
export type TuiConnectionStatus = 'connected' | 'reconnecting' | 'degraded' | 'down';
/** Which overlay (if any) owns the keyboard. */
export type TuiUiMode = 'list' | 'help' | 'confirm-kill' | 'prompt' | 'search' | 'digest' | 'message' | 'new-session';
/**
* Glyph capability tier. Detection is env-driven and therefore lives in a tiny
* function the app layer calls (`detectGlyphTier`); the renderer only ever
* takes the resolved tier as an input.
*/
export type TuiGlyphTier = 'nerd' | 'unicode' | 'ascii';
/** Header facts, all optional: the header degrades to just the product name. */
export interface TuiHeaderInfo {
hostname?: string;
instance?: string;
version?: string;
/** Plan-usage chip text, e.g. `5h 32% · wk 61%`. */
planUsage?: string;
}
/** The selected session's terminal tail, already run through `toDisplayLines()`. */
export interface TuiPreview {
sessionId: string;
/** Display lines, oldest first. */
lines: string[];
/** Set instead of lines when the tail could not be fetched. */
error?: string;
/**
* Set instead of lines when there is nothing to fetch (a history row has no
* live buffer). Distinct from `error`: nothing failed, so it must not read
* like something did.
*/
note?: string;
}
export interface TuiMessage {
text: string;
tone: 'info' | 'warn' | 'err';
}
/** Typed-confirmation state for `x` (kill): the user retypes the session name. */
export interface TuiConfirmState {
sessionId: string;
name: string;
}
export interface TuiPickerItem {
/** What choosing this item means to the caller; never shown. */
id: string;
label: string;
/** Second column, dimmed (a case path, a mode description). */
detail?: string;
}
/**
* A one-column chooser drawn as an overlay (the case and mode pickers behind
* `n`). Items are already filtered: the app owns the unfiltered list, the
* renderer only paints what it is given.
*/
export interface TuiPickerState {
title: string;
items: TuiPickerItem[];
/** Index into `items`; -1 when the list is empty. */
index: number;
/** Current filter text, when the picker filters as you type. */
filter?: string;
/** One line above the list: what is being chosen, or why the list is empty. */
hint?: string;
}
/** The `p` composer: one line aimed at one session. */
export interface TuiPromptState {
sessionId: string;
/** What the session is called on screen, for the footer prefix. */
label: string;
composer: TuiComposerState;
}
/**
* One line of the `/` overlay. Group headers are chrome (the API returns typed
* groups), so only `result` rows are selectable.
*/
export interface TuiSearchEntry {
kind: 'header' | 'result';
text: string;
detail?: string;
sessionId?: string;
/** The row can hand the dashboard a session that is open right now. */
live?: boolean;
}
export interface TuiSearchState {
composer: TuiComposerState;
/** The query `entries` answer. Lags the composer while a search is in flight. */
query: string;
entries: TuiSearchEntry[];
/** Index into `entries`, always a `result` row; -1 when none is selectable. */
index: number;
status: 'idle' | 'searching' | 'done' | 'error';
/** One line under the query: what happened, or why there is nothing. */
note?: string;
}
/** The `g` overlay: pre-formatted lines plus where the window starts. */
export interface TuiDigestState {
title: string;
lines: string[];
offset: number;
}
/**
* What `renderFrame()` reads. The store implements it; a test can hand-build
* one, which is what keeps the renderer testable without the model.
*/
export interface TuiRenderModel {
groups(): TuiGroup[];
readonly selectedId: string | null;
readonly connection: TuiConnectionStatus;
readonly mode: TuiUiMode;
readonly header: TuiHeaderInfo;
readonly preview: TuiPreview | null;
readonly message: TuiMessage | null;
readonly confirm: TuiConfirmState | null;
/** Optional so a test can hand-build a model without one. */
readonly picker?: TuiPickerState | null;
readonly prompt?: TuiPromptState | null;
readonly search?: TuiSearchState | null;
readonly digest?: TuiDigestState | null;
/** Live sessions only (RECENT rows are history, not sessions you have open). */
readonly sessionCount: number;
}
+5 -1
View File
@@ -109,7 +109,11 @@ export type HookEventType =
| 'elicitation_response'
| 'stop'
| 'teammate_idle'
| 'task_completed';
| 'task_completed'
// No Claude Code hook behind this one: it is the DeepSeek status bridge's
// "a turn STARTED" report (see deepseek-status-shim.ts). Keep in step with
// HookEventSchema in web/schemas.ts.
| 'agent_working';
// ========== API Response Types ==========

Some files were not shown because too many files have changed in this diff Show More