Compare commits

...
Author SHA1 Message Date
Codeman maintainer a922d301b1 chore: version packages 2026-08-21 20:24:38 +02:00
Ark0N abca552676 Merge pull request #327 from dignfei/fix/terminal-ime-punctuation
fix(terminal): preserve IME punctuation input
2026-08-21 20:23:26 +02:00
Ark0N 12a996b107 Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay
2026-08-21 20:23:17 +02:00
d fei 458e751a33 fix(terminal): keep shell history loading explicit 2026-08-22 01:55:50 +08:00
d fei dab432b3fd fix(terminal): bound shell history replay 2026-08-21 08:23:31 -04:00
Codeman maintainer 79a0399552 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:47:10 +02:00
Codeman maintainer 61251c0b94 fix(cli-resolvers): negative-result caching, SIGKILL on probes, restored VITEST hermeticity, wired not-found diagnostics
Post-merge follow-ups for PR #329 (shared CLI executable resolution):

- Negative-cache resolution misses with a doubling backoff (1min -> 5min
  cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
  shared resolver cached success only, so a missing CLI re-ran the whole
  chain - ending in a synchronous interactive login-shell spawn bounded by
  the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
  attempt, stalling the event loop each time, forever. Success still caches
  for the process lifetime, so an installed CLI is picked up within minutes
  without a restart. Tests drive the backoff via an injectable clock
  (createCliExecutableResolver `now` option, threaded through the
  createPiResolverForTest / createAntigravityResolverForTest wrappers).

- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
  pi/claude --version probes: execFileSync's timeout only SENDS the kill
  signal and then keeps waiting for the child to exit, and interactive bash
  ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
  survived the timeout and blocked the server permanently.

- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
  test pinned the deletion): under vitest the production resolver host now
  replaces un-injected IO primitives with inert stubs - no real PATH
  scanning, no login-shell spawns - and probePiVersion never executes a
  `pi` candidate again (`pi` is a generic binary name, so route tests
  hitting /api/pi/status executed whatever binary the machine carried).
  Tests opt in through the runCommand/isExecutableFile injection hooks or
  allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
  test is replaced by behavioral pins, including a real-executable fixture
  in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
  is ever removed again.

- Wire the six get*NotFoundMessage() exports (previously dead) into their
  intended call sites: the createSession throws in tmux-manager and the
  availability gates on POST /api/sessions and POST /api/quick-start in
  session-routes, replacing a third hardcoded copy of the text. A not-found
  error now names where resolution looked (server PATH, login shell,
  checked directories). npm run knip no longer reports any unused export
  from the resolver modules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:37:58 +02:00
Codeman maintainer bb4ba79791 fix(repo-status): async git, single-flight TTL cache, credential redaction, local-upstream parse
Post-merge follow-ups for #328 (GET /api/system/repo-status):

- Event-loop blocking: every git invocation in repo-status.ts is now async
  (promisified execFile), never execFileSync — the per-remote ls-remote +
  fetch could hold the event loop (SSE, PTY streaming) for up to ~60s per
  request. The whole computation is single-flight with a 45s TTL cache
  (createSingleFlightCache): concurrent requests share one in-flight
  promise, a fresh result is served without spawning git, and a rejected
  compute is never cached. Route handler shape and response fields
  unchanged; remotes still processed sequentially (concurrent fetches in
  one repo contend on ref locks).

- Credential disclosure: the redaction from git-clone.ts is extracted as
  exported redactGitCredentials() (sanitizeGitOutput now uses it) and
  applied via redactRemoteStatus() to every remote card's url and error
  string, so a scheme://user:token@host remote URL (or git stderr echoing
  it) never reaches a client.

- Non-interactive env: runGit() now uses the shared gitNonInteractiveEnv()
  instead of a partial GIT_TERMINAL_PROMPT/BatchMode env, also closing the
  GIT_ASKPASS/SSH_ASKPASS/SSH_ASKPASS_REQUIRE/DISPLAY/GCM_INTERACTIVE
  prompt paths.

- Upstream parse bug: a local-branch upstream (@{upstream} with no slash,
  e.g. after `git branch -u otherbranch`) made slice(0, indexOf('/')) into
  slice(0, -1) and yielded garbage like "maste". parseTrackingRemote()
  (pure, unit-tested) returns null for it, and the bare ref is dropped so
  it cannot be mistaken for a remote-tracking ref downstream.

Tests extended in test/repo-status.test.ts (parseTrackingRemote,
redactGitCredentials/redactRemoteStatus, createSingleFlightCache
single-flight/TTL/rejection semantics).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Codeman maintainer d7ad73bc9b fix(response-viewer): role on full-context blocks, divider ReDoS, pi mode (#326 follow-up)
Three post-merge fixes for the external-CLI response viewer:

- ?context=full blocks now carry role ('user' for prompts, 'assistant'
  for response/status/tool). The frontend's loadFullContext() renders
  via msg.role, so the roleless blocks lost the "You" badge and every
  turn rendered as the agent. kind/label/text are unchanged and the
  frontend needs no change.

- normalizeDividerStatusLine() dropped its backtracking regex
  (/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
  a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
  minutes at 10,000), and pane text is agent-controlled with buffers up
  to 32MB. Replaced by a linear counter walk with the identical accept
  set and captured content, pinned char-for-char against the old regex
  by a brute-force corpus test plus a hostile-input regression test
  that fails by timeout with the RegExp version (same approach as the
  glob-matcher hardening in 68ae9a8).

- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
  empty-viewer symptom the transcript branch exists to fix. The list
  stays a local duplicate of isExternalCliMode() (importing session.ts
  would drag node-pty into the pure module); a new exhaustive parity
  test asserts the two mode sets can no longer drift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Codeman maintainer 96ee8b536d docs: update the tap-report gate description after #325
#325 renamed _sessionUsesServerMouseStrip to _shouldReportMouseToCli and
added the server-observed cliMouseTracking half of the gate, which also
turned codex tap reports from measured no-ops into not-sent-at-all. The
invariants paragraph still described the old name and the old behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:13:26 +02:00
Ark0N b7f3b07c79 Merge pull request #330 from aakhter/ralph-loop-reschedule
Ralph loop silently stops polling after two ticks
2026-08-21 02:10:41 +02:00
Ark0N 2073a1b185 Merge pull request #328 from aakhter/repo-status-panel
Report repository status for git-clone installs (GET /api/system/repo-status)
2026-08-21 02:10:38 +02:00
Ark0N f9a8493823 Merge pull request #329 from aakhter/cli-login-shell-resolution
CLIs installed via nvm/Homebrew are not found when Codeman runs as a service
2026-08-21 02:10:35 +02:00
Ark0N d2711ef092 Merge pull request #326 from aakhter/response-viewer-external-cli
Response viewer is empty for OpenCode / Gemini / Antigravity sessions
2026-08-21 02:10:29 +02:00
Ark0N 30a15adbd6 Merge pull request #325 from Ark0N/feat/auto-copy-selection
feat(terminal): Auto Copy, put a finished selection on the clipboard
2026-08-21 02:10:19 +02:00
Aamer Akhter a35438ba34 fix(ralph): loop stops rescheduling after two ticks
The reschedule guard is `this._status === 'running' && this.loopTimer === null`,
but the timer callback never nulls `loopTimer`. So the handle stays non-null from
the first fire onward, the guard is false on every subsequent pass, and the Ralph
loop silently stops polling after exactly two ticks.

It stops without changing status: `status` stays `running`, `stop()` is never
called, and no error is raised — the loop just quietly never runs again, which is
what makes it hard to notice on a long autonomous run.

Null the handle inside the callback before re-entering `runLoop()`, which is the
pattern `orchestrator-loop.ts` already uses for its own reschedule.

Test: a regression case in test/ralph-loop.test.ts that runs a real 5ms-interval
loop for ~16 intervals and asserts it ticks at least 3 times. Against the unfixed
source it reports exactly 2.
2026-08-20 12:58:17 -04:00
Aamer Akhter fef903df98 fix(cli-resolvers): find CLIs installed via nvm/Homebrew when running as a service
A CLI installed by nvm, Homebrew or a user-level npm prefix lives on a PATH that
only a login shell sets up. Codeman running under systemd or launchd does not get
that PATH — launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` — so every
resolver reported the CLI as unavailable on installs where it is plainly there
and works from a terminal.

Each of the six resolvers had its own hand-rolled copy of the same PATH walk, so
the fix is factored into one shared `createCliExecutableResolver()` with an
explicit lookup order: the server process PATH, then common install directories in
order, then an interactive login shell as the last resort. Only the last step
spawns anything, and only when the cheap lookups have already missed.

Also adds `formatCliNotFoundMessage()`, so a failure explains where it looked
instead of just asserting the CLI is missing. Its diagnostics are bounded and
control characters are flattened, so a not-found message cannot dump arbitrary
environment data.

Success is cached and failure is retried, so installing a CLI while the server is
running is picked up without a restart.

Net -103 lines across the six resolvers. Behaviour is unchanged wherever the CLI
was already on the process PATH: that remains the first thing checked.

Tests: 20 cases in test/cli-executable-resolver.test.ts covering the precedence
order, login-shell-only resolution, the caching rule, unsafe-name rejection, and
the bounded diagnostics.
2026-08-20 12:47:42 -04:00
Aamer Akhter 02e7d3fcba feat(system): report repository status for git-clone installs
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.

Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.

Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.

Tests: 24 cases in test/repo-status.test.ts.
2026-08-20 12:39:28 -04:00
d fei f744719650 fix(terminal): preserve IME punctuation input 2026-08-20 10:35:23 -04:00
Aamer Akhter 63c5ba89da fix(response-viewer): populate the viewer for OpenCode/Gemini/Antigravity panes
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.

For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.

Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.

The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.

Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
2026-08-20 09:24:37 -04:00
Codeman maintainer 7fc4784d0f fix(approvals): clear the red tab alert when a dialog is answered in the terminal
Confirming an AskUserQuestion left its tab flowing red for the rest of
the turn (owner report: ~8 minutes on a running session, with no dialog
anywhere on screen). Two separate bugs, both live-verified.

The re-capture erased the evidence the staleness check runs on. Claude
Code fires the Notification behind the dialog (measured 6-7s on v2.1.237,
documented up to ~30s), so the 600ms re-capture routinely lands on a
frame the user has ALREADY answered, parses nothing, and applyCapture
overwrote item.options with undefined. A MISSING options is how "we never
could read this dialog" is expressed, and those items stay answerable by
design, so a cleared field was indistinguishable from a never-parsed one
and the item became permanently unsweepable: it survived every
GET /api/approvals and every page reload, cleared only on `stop`, and
still accepted an answer, sending a bare `1` into a composer with no
dialog under it. applyCapture is now ADD-ONLY for options.

Nothing ran the staleness check while a page was open. It lived only in
GET /api/approvals, which seedApprovals() calls on init and reconnect, so
`stop` was the first thing that ever cleared an answered dialog. The
`working` signal now runs the pane-VERIFIED variant (resolveIfDialogGone
-> verifyStillAnswerable): the heuristic only decides when to look, the
screen decides the outcome, so the existing "working can flap" rule is
respected.

A frame that parses no options is now conclusive in two cases, and only
those, so an unreadable capture still keeps the alert: the item once
parsed options, or the frame shows Claude actively running a turn. A
modal dialog BLOCKS the turn, so the two cannot coexist - measured, a
live-dialog frame carries neither the elapsed-timer spinner nor the
"esc to interrupt" footer, which the dialog replaces with "Enter to
select". That second signal is reached by a delayed staleness pass (3s)
scheduled alongside the re-capture, which closes the late-hook case where
the prompt is answered before the hook lands: nothing ever parses, `stop`
may have gone by already, and the alert outlived reloads until the 12h
TTL. The pass is deliberately later than RECAPTURE_DELAY_MS, whose whole
reason for existing is that the hook can beat Ink to the screen.

Frontend: _onHookElicitationComplete cleared only the elicitation entry,
but an AskUserQuestion arrives as permission_prompt, so it was clearing
the wrong alert; it now clears both, matching the server's kind-agnostic
APPROVAL_RESOLVING_EVENTS.

Verified end to end on an isolated beta instance, not just in unit tests:
before, resolution could only come from the stop route (approval:resolved
always immediately preceding hook:stop); after, it arrives from the new
paths, and a simulated late hook resolves at +3.12s with no stop, no
working signal and no GET, while the pane is still working. Tests use
frames captured off a live pane and each new one was confirmed to fail
against the old behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 12:18:16 +02:00
Codeman maintainer fa7e700834 fix(terminal): report a click to the CLI only when it asked for the mouse
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:

    $ [<0;88;20Mecho hello
    bash: 0: No such file or directory

The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.

What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.

Details that are easy to get wrong:

* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
  an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
  without turning tracking on is not asking about clicks, and counting
  those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
  cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
  broadcastSessionStateDebounced: the flag flips when a dialog opens, and
  the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
  the CLI re-emits its DECSET, which tmux does at client attach.

Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 11:21:29 +02:00
Codeman maintainer 7936a75e28 feat(terminal): Auto Copy, put a finished selection on the clipboard
App Settings > Terminal & Input > Selection & clipboard > Auto Copy
Selection (`autoCopySelection`, per-device, default OFF). With it on,
highlighting text in the terminal copies it: mouse drag, double-click
word, triple-click line, and the phone long-press selection. Ctrl+C is
untouched and still copies on demand.

Three things decide the shape of it:

* It fires at the END of a gesture, never in onSelectionChange. That
  callback runs for every cell a drag crosses, so copying there would be
  one clipboard write per mouse move. It only arms a pending flag; a
  document-level mouseup listener flushes, and the touch path calls the
  flush itself because it preventDefaults its touchend and no mouseup
  ever arrives there.
* The flush is synchronous inside the handler, because both clipboard
  paths need user activation: Firefox gates navigator.clipboard
  .writeText on it, and execCommand('copy'), the fallback the plain-HTTP
  LAN install lands on, has to run in the gesture's own task. A timer or
  a wait for onSelectionChange loses it, invisibly in Chrome.
* It deliberately does NOT do what copyTerminalSelection() does. That
  one clears the selection (so a second Ctrl+C is an interrupt) and
  focuses the terminal. Clearing would make text vanish under the cursor
  that just highlighted it, and focusing opens the on-screen keyboard
  over it on a phone. Focus is instead restored to whatever held it,
  which only matters for the execCommand fallback.

Guards are pure in decideAutoCopy() (constants.js): off, blank or
whitespace-only text, and a 1M-char cap, since a drag off the top of the
viewport autoscrolls and one gesture can sweep the whole 50k-line
scrollback. Past the cap the copy is refused rather than truncated, with
a toast pointing at Ctrl+C.

Feedback is silent on success except once per page load, so a feature
that works by doing nothing visible can still be told from a dead
toggle; failures and refusals toast, throttled to 10s.

Per-device on both counts the settings rule requires: in `displayKeys`
and absent from the .strict() SettingsUpdateSchema, because clipboard
access differs by device and by origin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 10:46:11 +02:00
Codeman maintainer 07b9c7fd7b fix(terminal): remove the unreachable copyTerminal(), closing out #322
The last two items of #322: copyTerminal() copied the entire buffer but
was wired to no button, shortcut or call site anywhere, and it wrote
through navigator.clipboard directly, which is undefined on the
plain-HTTP LAN install, so it would have failed there even if it were
reachable. Everything that actually copies goes through
copyTerminalSelection() and _copyText's execCommand fallback; whole-
buffer copy, should anyone want it, is a selectAll() away from that
same working path.

Closes #322

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 00:06:20 +02:00
Codeman maintainer c00e054e0e chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:43:22 +02:00
Codeman maintainer 68ae9a8c5f fix(files): match glob queries without regex so a hostile query cannot stall the server
The Files search compiled the user's query into a backtracking RegExp:
'*a*a*a...' became '^.*a.*a.*a...$', the classic blowup, evaluated
synchronously against every walked path — a pathological query could
freeze the event loop for the whole server (and every user of it in
multi-user mode). /api/search stays regex-free for exactly this reason.

Globs now match through a two-pointer wildcard walk, O(text · pattern)
worst case, with a 256-char query cap bounding the pattern side; an
overlong query compiles to null, the same answer as an empty one.
Semantics are unchanged (anchored, case-insensitive, * spans slashes)
and the existing tests pass untouched; the pathological pattern gets a
test that fails by timeout with the RegExp version.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:35:45 +02:00
Ark0N a49c30d173 Merge pull request #324 from aakhter/feat/files-panel-search
feat(files): search the Files panel by name or path
2026-08-19 23:33:08 +02:00
Codeman maintainer d871d1913f docs: restore the bullet PR #321 dropped off the xterm-zerolag-input gotcha
The new local-echo-overlay gotcha landed as a list item but left the
xterm-zerolag-input entry below it without its leading '- ', splitting
the Common Gotchas bullet list in two.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:25:09 +02:00
Ark0N ede3b05c10 Merge pull request #321 from rounakdatta/fix/mobile-link-taps
feat(mobile): links open from a tap, text can be copied, long prompts stay visible, wrapped links open whole
2026-08-19 23:23:51 +02:00
Ark0N c049de75db Merge pull request #320 from comzine/feat/custom-terminal-font
feat: Nerd Font prompt icons out of the box + configurable terminal font
2026-08-19 23:04:50 +02:00
Rounak DattaandClaude Opus 5 aae90599e5 fix(terminal): stitch a wrapped line through the indent its continuation carries
An agent's numbered list wraps its URL, and the link opened a PREFIX of it:

    1. https://github.com/users/someone/packages/container/p
       ackage/thing

opened `…/container/p`. The provider already stitched hard wraps — Ink emits a real
newline, so nothing is flagged `isWrapped` and a row that fills the last column is
taken as continuing — but it joined the row texts VERBATIM, and the continuation
carries the list's own three-space indent. That whitespace lands in the middle of
the token, which is exactly where the URL pattern stops. Flush-left wrapped URLs
(Claude Code's own `/login`) worked, which is why this survived.

The touch-selection helpers had the shallower version of the same bug: they walked
`isWrapped` only, so `Line` grabbed the single row on screen rather than the
logical line, and a long-press on a wrapped token selected only its visible half.

So the reconstruction now lives in ONE place, `terminalLogicalLine` in
constants.js, and both consumers use it — the link provider matching patterns over
its text and the selection helpers measuring words and lines with it. A link that
spans a wrap and a `Line` that stops at the screen edge were the same bug twice.

The helper drops the leading whitespace of a HARD continuation (the program's
indent) and keeps that of a SOFT one (the emulator inserts nothing, so it is real
content), records the dropped width per segment so the offset↔cell mapping stays
exact in both directions, trims only the final row so earlier offsets stay aligned
to cells, and keeps the 12-row bound that stops a screenful of full-width output
from being re-scanned on every hover.

⚠️ Selection spans are computed in CELLS, not text offsets: an xterm selection is
one contiguous run, so a token spanning a hard wrap also covers the indent cells
between its halves. A run that skipped them cannot be expressed, and would not
match what is highlighted.

Tests: `test/terminal-logical-line.test.ts` (8 cases: the indent drop, resolving
from either row, both mapping directions, soft continuations kept verbatim, no
over-reach past a short row, the row bound, final-row trimming, a missing row) and
5 in `terminal-touch-tap.test.ts` (the whole URL from either row, a token selected
across the wrap, `Line` spanning both rows, no reach into the next line). Removing
either half of the fix reds 5 and 8 of them respectively.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:41:43 +00:00
Rounak DattaandClaude Opus 5 2e58da7479 docs(mobile): document the phone gestures, and translate the selection bar
The three fixes in this branch change what a tap and a long-press MEAN on a
phone, and add a UI surface with its own z-index — all of which this repo keeps
written down rather than discoverable only by reading the handlers.

- `docs/wiki/Mobile-Guide.md` (the published user manual): a new "Tapping, links
  and copying" section, and the long-prompt behaviour in the keyboard section
  where the existing scroll/tap rules live.
- `CLAUDE.md`: the touch-gesture invariants next to the scrollback/wheel material
  (why the caret line is the boundary rather than the tap intent; why all three
  selection guards exist), the overlay's new bottom bound alongside the
  single-source note, and the selection bar in the z-index registry — 900, above
  terminal content and the local-echo overlay and deliberately below floating
  agent windows so it can never cover their controls.
- `i18n.js`: zh-CN for the bar's `Copy` / `Line` / `Clear selection`. The bar is a
  SIBLING of `.xterm`, not a descendant, so `SKIP_SELECTOR` does not cover it and
  the entries actually apply.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:19:36 +00:00
Rounak DattaandClaude Opus 5 ba843bb272 fix(mobile): keep a long prompt visible instead of hiding it behind the keyboard
Typing a prompt long enough to wrap ran the text off the bottom of the screen: the
tail — the part being typed, where the cursor is — sat behind the on-screen
keyboard, so the user was typing blind. Two independent causes.

**The overlay had no bottom bound.** On touch devices keystrokes are buffered in
the local-echo overlay and do not reach the PTY until Enter, so the CLI never
learns the prompt is long and nothing scrolls or reflows to make room. Meanwhile
the renderer lays its wrapped lines out straight DOWNWARD from the prompt row
(`top = promptRow * cellH`, each line at `i * cellH`) with nothing clamping it to
the visible rows — and with the keyboard up there are only a handful of those.

The block now grows UPWARD once it would pass the last visible row: it is lifted
so its final line lands ON that row. Every line div is opaque, so it covers
transcript above rather than vanishing under the keyboard below — the same thing a
real terminal does when a composer expands. A prompt taller than the whole
viewport keeps its TAIL, for the same reason the fix exists: the end is what the
user is looking at. `startCol` indents only the line that starts at the prompt
marker, so it is dropped along with that line when only the tail fits, and the
cursor follows the last VISIBLE line.

`rows` joins the render key: the layout depends on it, so a keyboard opening —
which changes rows without changing the text — must not be skipped as a redundant
render.

**`_shrinkPaddingToFit()` was reclaiming the bars' own space.** On phones the
toolbar and accessory bar are `position: fixed`, so they occupy no layout space
and `main`'s padding-bottom is the ONLY thing reserving room for them. Shrinking
it by the full sub-row slack pulled the terminal's bottom edge down underneath
them, and the row the following re-fit gained was painted behind them — clipping
the last line of a long prompt. The shrink now has a floor: the MEASURED height of
the currently-visible fixed bars, so genuine over-reservation of the hard-coded
84px is still reclaimed while a device that needs those pixels keeps them. The
floor is `Math.min(currentPadding, measured)`, so it can only ever prevent a
shrink, never cause a grow that would resize the terminal as a side effect.

Overlay behaviour lives in `packages/xterm-zerolag-input/` (single-source; the
vendor bundles are generated), so the fix is in the package with the row count
passed in as an optional `totalRows` — absent, the layout is exactly as before.

Tests: 7 cases in the package's `overlay-renderer.test.ts` (upward lift, tail
retention, indent drop, cursor on the last visible line, and the unclamped
fallbacks) and 7 in a new `test/mobile-keyboard-bottom-padding.test.ts` (reclaim,
floor, partial reclaim, no-grow, hidden bars, CJK strip, whole-row slack). 5 and 4
of them respectively fail without the fix. Package suite 238 pass, including the
codex byte-identity and replay tests.

Verified on Android + Chrome against a live instance: a ~460-character prompt
wrapping ~12 rows stays on screen while typing and arrives at the PTY intact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:14:52 +00:00
Rounak DattaandClaude Opus 5 756728e553 feat(mobile): long-press to select terminal text, tap to extend, Copy
There was no way to copy terminal text from a phone at all, and three layers
ruled it out independently: `user-select: none` across the whole terminal subtree
on touch devices (taps are cursor gestures there, so the OS callout had to go),
the WebGL renderer drawing glyphs as pixels with only the accessibility tree
behind them, and xterm's own selection being a mouse DRAG while the touch path
dispatches a zero-movement mousedown/mouseup pair — a click. `copyTerminal()`
exists but is wired to no button and calls `navigator.clipboard` directly, which
is undefined on the plain-HTTP LAN install the installer offers.

So the gesture drives xterm's `select()` directly: public API, renderer-
independent, and the highlight is drawn by xterm itself. Long-press is free real
estate — tap and swipe are taken, long-press and double-tap are used by nothing.

- **Long-press** (350ms, finger still within the shared tap slop) selects the
  run of non-whitespace under the finger. Whitespace is the only delimiter on
  purpose: every punctuation-aware word rule cuts a path, URL or hash in half,
  which is what you came to copy.
- **Drag** while held extends the selection; touchmove diverts from scrolling.
- **Tap** while the bar is up extends it too. That is the ergonomic core:
  picking up a 4px handle with a fingertip is a coin flip, tapping the other end
  is not. Dismissal stays explicit (✕ or Copy), so no tap is spent leaving a mode
  the user is still using.
- **Copy** goes through the existing `copyTerminalSelection()`, so it inherits
  the execCommand fallback that is the only route that works on plain HTTP.
- **Line** takes the whole logical line, wraps included, trailing pad trimmed.

Three guards are what make the gesture survive contact with a real phone, and
each fixes a symptom measured on Android Chrome:

1. **The compat mouse pair after touchend.** xterm focuses from its screen-element
   mousedown and SelectionService resets the model there, so lifting your finger
   popped the keyboard and dissolved the selection in one go. The tap path already
   had a guard for those events; the selection path simply never armed it. Armed
   now, and the touchend is `preventDefault`ed so the synthesis is stopped at the
   source (that listener is no longer passive).
2. **The platform's own long-press.** Android Chrome runs its handling at ~500ms
   and focuses the nearest editable element — xterm's helper textarea, parked at
   the cursor — which no touch handler can preventDefault because it never sees an
   event. A focus guard blurs the terminal input for the duration of the gesture,
   whatever focused it, bounded by a self-expiring deadline so a stuck flag can
   never leave the keyboard unreachable. `contextmenu` is suppressed for the same
   window, and the threshold sits at 350ms so it lands clear of the platform's.
3. **Copy re-focusing the terminal.** `copyTerminalSelection()` ends with
   `terminal.focus()`, which is right on a desktop and wrong on a phone: the
   keyboard covers what was just copied with nothing waiting to be typed.

The bar is built in JS because index.html is read once at server start, and its
styles live in styles.css rather than mobile.css because the gesture is
touch-driven, not width-driven — a touch tablet in landscape gets the gesture and
would otherwise have no bar to copy from.

12 tests in `terminal-touch-tap.test.ts` cover the word rule, forward and
backward extension, cross-row selection, Line, tap-to-extend, the copy path, and
each of the three guards including the focus guard's expiry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:14:52 +00:00
Rounak DattaandClaude Opus 5 f2d3a7e3c1 fix(mobile): links open in a new tab from a tap, in the terminal and the chat
On a phone no link was openable, on either surface, for two unrelated reasons.

**Terminal.** xterm resolves the link under the pointer on `mousemove` and
activates it on `mouseup` over its SCREEN element. A touch tap delivers neither:
`touch-action: none` on the terminal subtree plus touchstart's preventDefault for
a 'content' tap suppress the browser's compatibility mouse events,
`_installMobileTapMouseGuard` drops the trusted ones that still arrive inside the
450ms tap window, and the synthetic mousedown/mouseup pair dispatched for mouse
REPORTING goes to the `.xterm` root — an ancestor of the node the linkifier
listens on, so it cannot reach it — and carries no mousemove either way. Every
URL and file path in the terminal was therefore inert on phones and tablets,
Claude Code's own `/login` URL included.

The tap path now activates the link itself, through the SAME provider that feeds
the hover linkifier (`_terminalLinkAtPoint`), so a tap and a desktop click can
never disagree about what is a link or where it ends — containment mirrors
xterm's own `_linkAtPosition`. It runs synchronously inside the touchend handler,
which is what keeps the user gesture that lets `window.open` past the popup
blocker, and before any mouse report, exactly as `_handleDesktopTerminalClick`
already skips the SGR tap for a hovered link.

Two kinds of row keep their existing meaning: the caret's logical line, where a
tap places the cursor and a URL the user typed must stay editable, and TUI-owned
rows, where a numbered choice or an expandable readback is answering a dialog and
routinely carries the very path the tap would otherwise open. The caret line is
the boundary rather than the tap intent, because a plain shell classifies EVERY
tap as 'input' and gating on that would leave every URL in shell output inert.

**Chat.** `marked` emits a bare `<a href>` and the markdown sanitizer's allowlist
carries no `target`, so a tap in the response viewer navigated the current tab
away: on a phone that unloads the whole dashboard — SSE, terminal buffers, unsent
composer text — and there is no middle-click or open-in-new-tab affordance to
work around it. `_renderMarkdown` now decorates anchors in the template pass it
already makes for code blocks. That pass runs AFTER sanitizing, so it is the only
source of both attributes: an agent-authored `target`/`rel` is already stripped,
and `rel="noopener noreferrer"` is set on the same element in the same breath, so
no page Codeman opens gets a `window.opener` handle back. Fragment links stay
in-page; mailto:/tel: are left to the OS rather than stranding an empty tab.

Tests: 10 cases in `terminal-touch-tap.test.ts` (URL, file path, log path,
scrollback, no-double-report, composer, shell mode, dialog row, no provider) and
a new `response-viewer-external-links.test.ts` driving the shipped marked +
DOMPurify + app.js. 7 of them fail without the fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:14:52 +00:00
Aamer Akhter 5cc78669bd feat(files): search the Files panel by name or path
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.

compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.

The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.

Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
2026-08-19 09:17:20 -04:00
Ark0N d4ccff07ca Merge pull request #319 from Ark0N/fix/dep-advisories
fix(deps): clear production npm advisories, fix sw.js caching regression
2026-08-19 14:54:21 +02:00
Tobias WeberandClaude Fable 5 108c00e78d feat: bundled Nerd Font symbols fallback + per-device terminal font setting
Shell prompts using Nerd Font glyphs (powerline, p10k/starship folder and
git icons) rendered as missing-glyph boxes: the built-in xterm stack has no
private-use-area symbols, and phones have no Nerd Fonts installed at all.

- Bundle Symbols Nerd Font Mono (icons-only, MIT, 1.2MB woff2) served from
  fonts/ and appended to the terminal stack before monospace — browsers fall
  back per glyph, so icons render everywhere while text stays in the text
  fonts. font-display: block + preload keep tofu out of xterm's glyph atlas.
- New per-device terminalFontFamily setting (App Settings > Terminal &
  Input > Font): prepended to the built-in stack, never a replacement, so
  the symbols fallback and final monospace always survive. Applied live on
  save (refit + echo-overlay refreshFont, mirroring setFontSize).
- Single source for both xterm surfaces: TERMINAL_FONT_DEFAULT_STACK +
  resolveTerminalFontFamily() in constants.js, unit-tested.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VjnbbZRBuvR5E3SDouwXr9
2026-08-19 00:45:19 +02:00
Codeman maintainer 8a54b331e3 fix(deps): clear production npm advisories, fix sw.js caching regression
Resolves the four advisories that reach the production dependency tree. The
other 16 npm audit reports are devDependencies-only (Remotion, Puppeteer,
postcss, the eslint/tsx toolchain) and never ship to users.

- @fastify/static 9.1.3 -> 10.1.3  GHSA-8pvw-jcv7-9cmj (authz bypass via
  non-canonical URL paths). Covers <=10.1.1, so all of 9.x is affected and
  the fix exists only on the 10.x line.
- find-my-way 9.6.0 -> 9.8.0       GHSA-c96f-x56v-gq3h (HTTP/2 DDoS)
- fast-uri 3.1.2 -> 3.1.5          GHSA-v2hh-gcrm-f6hx (host confusion)
- brace-expansion -> 5.0.9/1.1.18  GHSA-3jxr-9vmj-r5cp (expansion DoS)

The last three are transitive and needed only a lockfile re-resolve, so no
overrides were introduced.

The @fastify/static major changes setHeaders' first argument from a Node
ServerResponse to a FastifyReply. Two consequences:

1. res.setHeader() -> reply.header(). The v9 body throws TypeError from
   inside the plugin on every static request.
2. Precedence flips, silently. The callback used to write to the raw
   response and lose to the route's staged reply headers; it now writes to
   the reply and wins. That gave /sw.js a year of immutable in place of the
   no-cache, no-store its route sets, pinning a service worker on every
   client with no server-side recovery. A route that already set
   Cache-Control now keeps it.

Verified against v9 to confirm the sw.js behaviour is a regression and not
a pre-existing bug.

ws appears in npm audit but production is on 8.21.0, outside the vulnerable
range; the only affected copy is bundled under @remotion/renderer (dev-only,
and remotion is pinned at 4.0.473 because the compositor refuses to start on
a version mismatch).

Adds test/static-cache-headers.test.ts, which drives a real server and covers
a caching contract that had no test at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:24:22 +02:00
Codeman maintainer 09bf00c815 chore: version packages 2026-08-18 21:17:09 +02:00
Ark0N 2e19fc0430 Merge pull request #315 from aakhter/fix/respawn-stop-race
fix(respawn): do not revive a stopped controller after a cycle-step write
2026-08-18 21:15:17 +02:00
Ark0N 30f35490f6 Merge pull request #314 from aakhter/fix/symlink-safe-workspace-confinement
fix(routes): canonicalize the workspace before comparing it to a resolved path
2026-08-18 21:15:11 +02:00
Ark0N 322801052b Merge pull request #316 from Ark0N/chore/test-script-split
chore(test): make `npm test` the CI gate and give each excluded suite a runner
2026-08-18 21:15:00 +02:00
Ark0N 736a6b8b7b Merge pull request #313 from Ark0N/feat/sidebar-rich
feat(sidebar): add a rich session sidebar that carries the home screen's row detail
2026-08-18 21:14:53 +02:00
Ark0N 6525ade530 Merge pull request #317 from Ark0N/feat/wiki-tapzones-lineage-colours
Wiki publishing, phone tab tap-zone fix, and per-parent lineage colours
2026-08-18 21:14:46 +02:00
Codeman maintainer 947ff6f6fa chore(test): make npm test the CI gate and give each excluded suite a runner
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.

`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.

The suites it leaves out are not abandoned; each has a command:

  test:browser  5 Playwright files (chromium + a live server; codex-predictive-echo
                also needs a real codex binary)
  test:mobile   unchanged — the above plus per-machine PNG baselines
  test:perf     2 wall-clock benchmarks; need an otherwise idle machine
  test:all      the old everything-behaviour, kept reachable

test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.

The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.

So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:

  gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo

⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.

Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 19:55:23 +02:00
Aamer Akhter ce405a4cff fix(respawn): do not revive a stopped controller after a cycle-step write
Each cycle step (kickstart, update, /clear, /init) checks for `stopped` before
`await session.writeViaMux(...)`, then emits `stepSent` and calls
`setState('waiting_*')` after it.

stop() is asynchronous with respect to that await. One that lands while the
write is in flight has already passed the guard that ran, so the post-await
setState() puts a stopped controller back into a waiting state — re-arming its
step timers against a session the user asked to stop.

Re-check after the await, before emitting and setting state.

The guard reads the public `state` getter rather than `_state` on purpose:
TypeScript narrows `_state` across the await from the pre-await check and cannot
see that stop() mutated it, so `this._state === 'stopped'` is rejected as a
comparison with no overlap (TS2367) at all four sites.

Adds test/respawn-stop-race.test.ts, which drives the interleaving
deterministically by calling stop() from inside the mocked write rather than
relying on timing. All four steps go red without these guards.
2026-08-18 10:59:44 -04:00
Aamer Akhter 8e5691b05c fix(routes): canonicalize the workspace before comparing it to a resolved path
validateSessionFilePath realpath-resolves the candidate path but compared it
against the raw sessionWorkingDir. When the workspace is itself reached through
a symlink the two sides live in different namespaces, so relative() reports a
spurious `../` and every file in that workspace is judged an escape — reads and
writes in the session are refused wholesale.

That is not an exotic setup: os.tmpdir() hands back a symlinked path on macOS
(/tmp -> /private/tmp), and symlinked project directories and bind-mounted case
paths hit it too.

Resolve both sides and compare canonical to canonical. This only makes the
comparison honest — it does not widen it. The candidate keeps its own realpath,
so a symlink pointing out of the workspace and a ../ traversal are still
refused, and a workspace that cannot be resolved now fails closed.

Three stubs in file-routes.test.ts used a blanket
realpathSync.mockReturnValue(escapeTarget), which answers the same path for the
workspace and the candidate; with both sides resolved that makes an escape look
contained. They now use the input-aware mockImplementation idiom the rest of
that file already uses, so the workspace resolves to itself and only the
candidate escapes. Verified they still bite: removing the confinement check
turns all of them red.

Adds test/route-helpers-symlink-confinement.test.ts, which exercises the
function against a real symlinked workspace on disk and pins the negative cases
(../ escape, symlink-out, missing file) alongside the fix.
2026-08-18 10:46:21 -04:00
Codeman maintainer 98e37bf895 feat(sidebar): add a rich session sidebar that carries the home screen's row detail
Session List Layout gains a third option. The old "Left sidebar" becomes
"Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar"
puts on each row what the desktop home rail and the phone overview already show:
when the session was first created, how long it has been in the state it is in,
and a status pill naming that state.

A docked column is not a tab strip. It has width to spare and a row per session
either way, and "name + folder" is the whole story a TAB can tell, not the whole
story there is. This is the information that was missing, and it already existed
one surface over.

Both sidebar values are the same layout, and both set data-session-list="sidebar";
the row detail rides on a separate data-sidebar-detail attribute. That split is
the load-bearing decision here: every one of the ~25 isSessionSidebarActive()
call sites and every html[data-session-list="sidebar"] rule in styles.css and
mobile.css keeps matching both variants without being touched. A third
data-session-list value would have meant auditing and editing all of them.

- Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on
  'sidebar' keeps exactly the layout they picked — the rename is label-only.
- State classification and the "how long has it been like this" anchor come from
  _mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three
  surfaces cannot disagree about what "working" means. A working pane repaints
  ~1/s, so its duration is measured from the turn's last Enter: a running turn
  reads "working 12m", not "0m".
- Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild
  would restart every load spinner and alert animation in the list, twice a
  minute. The clock runs only while rich rows are on screen, and is stopped from
  both render paths and from applySessionListLayout().
- The incremental render path updates the pill, the accent class and the since
  anchor; a tick alone cannot see a state change, and a new turn re-stamps
  lastSubmitAt without changing state.
- applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich
  leaves data-session-list on 'sidebar' both times, and the meta line is emitted
  by the row template rather than toggled by CSS, so the old layout-only test
  would have flipped the setting and repainted nothing.
- Width: 300px for the extra line. The collapsed 44px rail and the handheld
  drawer are both explicitly held back from it — the desktop rule is (0,3,1) and
  would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px
  phone's drawer to 300px.
- Missing/stale mobile-overview.js degrades to a row with no meta line rather
  than throwing and taking the whole tab strip down.

15 new tests cover the attribute split, the solo-window override, the
detail-change re-render, the row model, both render paths, the clock lifecycle
and the mobile width guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:26:34 +02:00
Codeman maintainer 5ded2ed1a3 docs: correct drifted counts in CLAUDE.md, declare postcss
The frontend load order omitted session-lineage.js (29 modules listed, 30
loaded), and several inventory counts had drifted from the tree: route handlers
~200 to ~217 with system, files and approvals each understated, src/config 20 to
21 files, install.sh 69KB to 92KB, and the Prettier exemption list, which also
never mentioned mobile.css. Two of the missing handlers are endpoints CLAUDE.md
already documents in prose but never counted.

postcss is imported by two tests but was only present transitively via vite, so
knip reported it as an unlisted dependency. Declared at the version already
resolved in the lockfile.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:44 +02:00
Codeman maintainer 76ea090a67 fix(ui): bind lineage line colours to the spawning tab
Lineage arcs were coloured per child, so one tab's own workers each got a
different colour, which is the distinction the colours exist to make. The colour
is now keyed on the parent: every arc leaving one tab is the same colour however
many workers it spawns, so the strip reads as "these five came from w1, those
two came from w2". A child that spawns in turn is a parent in its own right and
gets its own colour for the arcs below it, so a chain changes colour at each
generation while each generation's fan-out stays uniform.

The new tests drive the real _appendLineageConnectionLines() and assert the
painted custom property, because testing the colour function alone passes just
as happily with the child id passed back in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:43 +02:00
Codeman maintainer d30cac4440 fix(mobile): keep the active phone tab's centre off its action icons
The active tab is the only one that grows a gear and a close button, and with a
short session name they were eating it: "w1" rendered a 13px label while gear
plus close took 50px of a 116px tab, so the tab's geometric centre landed on the
gear and a thumb aiming at the middle of the tab opened Session Options instead
of switching sessions. Reserving a minimum label width on the active tab widens
the tab by the difference instead.

The floor is set by the 10th tab onward, which renders no number badge and so
sits 10px further right; a numbered tab clears the icons at 20px but a
numberless one needs 40px. The test recomputes that inequality from the
stylesheet rather than pinning the pixel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:43 +02:00
Codeman maintainer f3cb7696f0 docs: publish docs/wiki as the user manual, with a sync workflow
30 pages covering install, concepts, the dashboard, the agent CLIs, unattended
runs, remote and Docker cases, security and the HTTP API, plus a sidebar and a
footer. The wiki repo has no CI and no review, so docs/wiki is the source of
truth and .github/workflows/wiki-sync.yml mirrors it on every push to master.

The workflow refuses to mirror when docs/wiki is missing or holds no pages,
because it deletes before it copies and would otherwise publish the deletion of
every page. The footer carries a {{VERSION}} placeholder stamped at publish
time rather than a hand-written version, which went stale on every release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:43 +02:00
Codeman maintainer 5080390e2c chore: version packages
Give the active-session handoff one owner: closeSession captures wasActive before its await and the session_deleted handler stands down for a close this tab started.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 22:54:46 +02:00
Codeman maintainer f7e2975883 chore: version packages
Gate the idle-alert acknowledgement to human selections: the boot restore, a solo window opening its target, and the post-close fallback no longer spend a yellow tab alert.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 21:42:25 +02:00
Codeman maintainer f60bf93c99 chore: version packages
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:46:48 +02:00
Codeman maintainer f4ba4d2cb1 chore: version packages
Persist the 'I checked it' state of yellow idle tab alerts across reloads and devices.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 00:46:03 +02:00
Codeman maintainer fa8ebe0068 chore: version packages 2026-08-16 23:07:43 +02:00
Ark0N b0e493d462 Merge pull request #311 from Ark0N/fix/workspace-hooks-followups
Workspace hooks follow-ups: one decision core for every claude create path, docker shell gate, boot-sweep and statusLine guards
2026-08-16 20:44:49 +02:00
Ark0N 631913f04c Merge pull request #310 from Ark0N/fix/files-sidebar-followups
fix: file-link and session-sidebar review follow-ups from 1.19.0
2026-08-16 20:44:14 +02:00
Ark0N bb959c4aac Merge pull request #309 from Ark0N/fix/home-order-followups
Home-screen ordering follow-ups: live stamps, one numbering, restart-proof recency
2026-08-16 20:43:21 +02:00
Codeman maintainer 24ed43935c fix: file-link and session-sidebar review follow-ups from 1.19.0
Five post-merge review items from PRs #306 (clickable file paths) and
#307 (session sidebar):

- constants.js FILE_PREVIEW_EXTENSIONS gains the media extensions it was
  missing vs the single-source sets in attachment-registry.ts (m4v ogv
  ogg oga m4a aac flac opus), so an in-workspace .m4a opens the preview
  player instead of the log viewer; new test/media-extension-parity.test.ts
  pins all three copies (constants.js, panels-ui.js, attachment-registry.ts)
  against each other.
- FILE_PATH_LINK_PATTERN drops `etc` from its root alternation: /etc is
  unconditionally in DEFAULT_BLOCKED_TREES, so every /etc link 403'd.
  Negative cases added to the link-provider and response-viewer tests.
- updateSidebarCount() counts the rows actually on the sidebar list
  (session rows + web-tab rows, minus filtered-out ones) instead of
  this.sessions.size, and applySidebarFilter() refreshes it so the count
  follows the filter box per keystroke.
- The incremental-render connection-line gate now also fires in sidebar
  layout (this._lineageEdgeCount is permanently 0 there), matching the
  strip-scroll listener widened in #307, so a badge changing row heights
  redraws subagent/ultracode connectors.
- isSensitivePath() blocks ~/.claude.json, ~/.claude/settings.json and
  ~/.claude/settings.local.json (credential-bearing by schema), anchored
  to homedir() read at check time so case-level .claude/settings*.json
  files stay servable in the File Viewer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 20:34:34 +02:00
Codeman maintainer cb9149879d restore the activity stamp across restarts: the quiet ordering no longer flattens on deploy
Root cause of the reviewer's mass-bump measurement (17 of 17 sessions with an
identical lastActivityAt): every restart restamps all sessions in the
constructor loop, and the boot auto-attach's repaint re-bumps the rest within
the same second. A 12-minute steady-state sample shows NO ambient mass bump,
so restarts are the whole story, and Codeman restarts on every deploy.

The stamp now has a display twin: recovery threads the previous run's
lastActivityAt from state.json into the wire-visible stamp (getter + toState),
and a 15s settle window keeps the attach repaint from overwriting it. Real
actions (input, task assignment, respawn) always write through. The private
stamp keeps its boot-anchored semantics untouched, because the idle
confirmation reads it as how long the pane has been quiet.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 20:32:41 +02:00
Codeman maintainer f44d597450 review fixes: every claude create path routes through the workspace-hooks decision
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled,
default ON) moved from a session-routes-local helper into hooks-config.ts as
applyWorkspaceHooks(workspace, install?), and the claude session-create sites
that bypassed it now go through it: cron job fires (cron-service), legacy
scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research
and planner one-shots. A cron or scheduled run firing in a linked case that
never had an interactive session ran hook-blind (no stop for completion
detection, no tab alert on a blocking dialog).

The shared core also carries the two guards every caller needs: a workspace
that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot
recovery sweep used to resurrect a deleted repo as an empty tree holding only
.claude/settings.local.json), and all errors are swallowed since a create must
never fail on hooks. Route handlers keep resolving the setting through their
ConfigPort and pass it in; non-route callers omit it and the core reads
settings.json itself (absent key or unreadable file = ON).

Two adjacent gates tightened in session-routes:
- the docker quick-start hooks branch excluded the five external CLIs but let
  `shell` through, contradicting its own rule that only claude reads .claude
  hooks; it is now gated on mode === 'claude'
- the statusLine exporter call in POST /api/sessions got the same
  !remote && body.workingDir guard the hooks call got in 499d355 (it mkdirs the
  same way, so a remote attach created a junk user@host:session dir locally and
  a cwd-fallback create wrote into $HOME)

plan-routes' one-shot deliberately stays out: its workingDir is process.cwd(),
exactly the target 499d355 forbids writing into. restoreMuxSessions stays out
too: the boot sweep already covers recovered workspaces.

Tests: quick-start existing-case install, docker claude-installs/shell-does-not,
and the core directly (default-ON install, OFF add-nothing, OFF still heals a
stale block, malformed file untouched, vanished workspace skipped); the remote
and cwd-fallback regressions now also send statusLineTelemetry:true to pin the
statusLine guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 20:27:38 +02:00
Codeman maintainer cdbde9f36f home-order follow-ups: fresh stamps for the blocked group, sort/display agreement, live Alt+N projection
Three follow-ups from the 1.19.0 review of the activity-ordered home screens:

- Hook events now ride the same debounced session state broadcast the
  working/idle handlers use. The blocked group ranks on lastActivityAt, and
  without this a permission prompt raised after page load kept ranking by
  whatever stamp the browser loaded with.

- A working row with no submit stamp now shows the lastActivityAt fallback
  its sort anchor already uses: a row must never be ranked by a number it
  does not display.

- Alt+digit resolves through the live-session projection the render paints
  (sessionOrder minus dead ids), so a stale id cannot shift every painted
  number off its target, web tabs included. New tests pin both surfaces to
  one shared order and the numbering to the live projection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 20:21:54 +02:00
Codeman maintainer f07905b193 chore: version packages 2026-08-16 19:32:25 +02:00
Ark0N 05c94f5ac0 Merge pull request #304 from Ark0N/fix/workspace-hooks-install
fix: install Codeman hooks into every claude workspace, not just cases Codeman created
2026-08-16 19:31:06 +02:00
Ark0N aaf22909bc Merge pull request #306 from Ark0N/feat/file-path-links
fix(files): open the files agents print, wherever they wrote them
2026-08-16 19:30:41 +02:00
Ark0N 94908ffdb5 Merge pull request #303 from Ark0N/feat/overview-activity-order
Sort the home-screen session lists by activity, not tab order
2026-08-16 19:26:03 +02:00
Ark0N 82fe3cf684 Merge pull request #305 from Ark0N/docs/skill-hooks-rule
docs(skill): hooks are a setting now, not who created the directory
2026-08-16 19:23:17 +02:00
Ark0N 6946ca0b8a Merge pull request #307 from Ark0N/feat/session-sidebar
feat(web): optional collapsible left session sidebar
2026-08-16 19:23:14 +02:00
Codeman maintainer ea4b940cef review fixes: block Codeman's own credential-bearing JSON, make the inside-anchor test bite
Widening the servable extensions to EDITABLE_EXTENSIONS made ~/.codeman
JSON previewable for the first time, and the blocklist named only
state.json. But settings.json holds a credential BY SCHEMA
(voiceSettings.apiKey), push-keys.json holds the VAPID PRIVATE key, and
intents.json is written 0600 precisely because captured prompts can carry
secrets — all three were one authenticated click away once an agent
printed the path. Blocked alongside state.json, whose rule now also
catches state-* siblings.

The never-re-cuts-inside-an-anchor test used an unmatchable URL tail, so
it passed with the guard deleted; the fixture now carries a matchable
/tmp path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 19:21:43 +02:00
Codeman maintainer c6f428e687 review fixes: broadcast session state on working, guard the CodemanSessionOrder global
The running group sorts on lastSubmitAt, but nothing pushed a session:updated
when a turn STARTS — the browser kept whatever stamp it loaded with, so a
30-second-old turn could rank (and read) as an hour-long one. The working
handler now rides the same debounced state broadcast idle already uses.

And both call sites of window.CodemanSessionOrder now degrade to tab order
when the global is missing (iOS Safari's documented stale-cached-JS after a
deploy) instead of TypeErroring the whole home screen away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 19:18:49 +02:00
Codeman maintainer c8f3981b0c review fixes: 1.19.0 is the real version boundary, and the preamble stamp matches its bytes again
endpoints.md named 1.18.x as the version where workspace hooks became a
setting, but 1.18.x servers do NOT have this behavior — an agent driving
one would falsely conclude its workspace has hooks. The feature ships in
1.19.0. And preamble.sh changed content this PR without bumping its
CODEMAN_PREAMBLE stamp, so a cache stamped 1.18.3 would pass the
staleness check while holding old bytes; stamp bumped to 1.19.0 in
preamble.sh and the SKILL.md heredoc together (byte-identity pin).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 19:16:10 +02:00
Codeman maintainer 499d35566b review fixes: never install workspace hooks for a remote attach or a cwd-fallback create
A claude-mode attachRemoteSession create overwrites workingDir with the
user@host:session pseudo-path, which is a RELATIVE path locally — the old
refresh-only call no-op'd on it, but ensureCodemanHooks mkdirs, so it
created a junk local directory. And with workingDir omitted the cwd
fallback reaches the hooks write unvalidated; under installer-created
services cwd is $HOME, so hooks materialized in ~/.claude/settings.local.json.

Both guarded at the applyWorkspaceHooks call site; regression tests prove
the remote attach leaves no junk dir and the no-workingDir create leaves
the server cwd untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 19:15:08 +02:00
Codeman maintainer 210da991d5 Merge christianhaberl's session-sidebar branch, ported to current master
Brings in https://github.com/christianhaberl/Codeman/pull/4 (three commits,
authorship preserved) and adapts it across the 211 commits master gained
since the branch was cut:

- App Settings control re-authored for the set-* surface (PR #278): a
  set-row in Layout -> Tabs, replacing the old settings-item markup the
  branch targeted. i18n description synced.
- Lineage arcs (PR #291, post-branch) are SKIPPED in sidebar layout:
  computeLineagePath()'s U-bridge geometry hangs from the horizontal
  strip's bottom edge and has no meaning against a vertical list. The
  lineage strip-scroll listener now also redraws subagent/ultracode
  connectors while the sidebar scrolls vertically.
- The desktop home tab rail (post-branch) defers to the sidebar: both dock
  the session list flush left, and the rail would render z-ordered under it.
- Active-row reveal unified into _scrollActiveTabIntoView() (#257 landed on
  master after the branch): sidebar mode branches to scrollIntoView
  block:'nearest', and _fullRenderSessionTabs() restores scrollTop alongside
  the #257 scrollLeft restore so ambient rebuilds cannot yank a mid-scroll
  sidebar back to the top.
- Mobile active-tab hoisting the branch guarded against no longer exists on
  master (removed by #257); kept master's order-stable render.

Verified: typecheck, lint, format:check, check:frontend-syntax,
check:public-assets, PostCSS parse of both merged stylesheets, the 26 new
jsdom tests, the structural guard suites, and the headless-Chromium harness
(scripts/verify-session-sidebar.mts) green across all seven layout states
at 1600/1000/393px against current master.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-16 18:28:15 +02:00
Codeman maintainer da999b130e feat(files): preview text files from outside the workspace, and stop routing them at a viewer that cannot read them
A .json/.log/.yaml/code path outside the session workspace was refused as an
unsupported type, and clicking one in the terminal made it worse: text goes to
the log viewer, which spawns `tail -f` and allows only the workspace, /var/log
and ~/logs, so it answered "Path must be within working directory or allowed
log directories" while the same path clicked in the response viewer previewed
fine. Two surfaces, two answers, for a file the session can already cat.

- TEXT_ATTACHMENT_EXTENSIONS IS EDITABLE_EXTENSIONS (config/file-editing.ts),
  not a second curated list that would drift from it. The rule reads: if the
  viewer would open a file for editing inside the workspace, the same file
  outside it can be read. The suffix was never the confidentiality gate here,
  the path guard is (sensitive-file blocklist, /root and /etc trees, realpath
  before the check), and it still runs on every registration.
- Widening what can be READ must not widen what can RUN. html/htm join svg in
  serveRawFile's download-only branch, so markup is never served with a
  renderable type on our own origin; other text goes out as inert
  text/plain; charset=utf-8 with nosniff, matching what the path picker does.
  The preview reads through fetch(), which ignores the disposition, so a
  clicked .html still shows its source.
- ~/.codeman*/state.json joins isSensitivePath. It persists
  SessionState.envOverrides and the env allowlist admits key-shaped names
  (GEMINI_API_KEY, CLAUDE_CODE_*), so it can hold a live credential. Same
  treatment as hook-secret and users.json, and the rest of the tree stays
  attachable.
- The terminal sends an out-of-workspace path to the preview instead of the log
  viewer. In-workspace text keeps the tail viewer, which is the point of it, and
  file-stream-manager's allowlist is untouched: no `tail -f` on arbitrary host
  paths.
- The by-id text preview is bounded like the workspace one: a Range request for
  the first 512KB (a real partial read, not a discarded 50MB download) plus a
  500-line cap, with the footer saying so.

Verified on an isolated instance: a 1.1MB external log opens in ~1.8s showing
500 lines with "showing first 500 lines" in the footer; json, yaml and code
preview; an .html carrying a script tag renders as source and does not execute;
.svg is still refused; a terminal click on an external .yaml opens the preview
with no log viewer and no attachment card; an in-workspace .log still opens the
streaming tail viewer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 18:12:00 +02:00
Codeman maintainer cbc54fc98d feat(files): play video and audio from outside the workspace too
A clip an agent wrote inside the workspace played with a working scrub bar,
while the same file in /tmp was refused as an unsupported type. The workspace
preview classified media with its own inline extension sets and the attachment
allowlist had no media at all, so the two paths disagreed about what a video is.

- VIDEO_ATTACHMENT_EXTENSIONS and AUDIO_ATTACHMENT_EXTENSIONS now live in
  attachment-registry.ts and are imported by file-content's classification, so
  both paths answer the same. mp4/webm/mov/m4v/ogv and
  mp3/wav/ogg/oga/m4a/aac/flac/opus join the attachment allowlist.
- Real MIME types for those extensions. Without one the raw route falls back to
  application/octet-stream, which a <video> refuses to decode: the player
  renders and then does nothing.
- getAttachmentType() gained the video and audio members of
  AttachmentDetectedType. Attachment cards have no per-type CSS and their
  thumbnail falls back to the type label, since the thumbnailer has no media
  branch and answers 204 rather than spawning a converter.
- The preview overlay's by-id branch renders <video>/<audio> with the same
  markup as the workspace branch, playsinline included. Serving was already
  range-aware, so seeking works.

The image-watcher keeps its own narrow detection list (png/pdf/docx/pptx), so
this does not start popping cards for every video an agent writes. Text types
that are not md or txt (.json, .log, code files) remain out of the allowlist by
choice and still report what is previewable instead.

Verified on an isolated instance: an external mp4 and mp3 both play, seek, and
report the right duration, matching the in-workspace clip exactly, and a click
on an external mp4 in the terminal opens the player with no attachment card.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 17:35:43 +02:00
Codeman maintainer 4e2c1b9989 fix(files): open file paths agents print, from the terminal and the chat
A path an agent prints was already underlined in the terminal, but clicking
one opened the preview overlay on "File not found": file-content/file-raw
resolve against the session workingDir and refuse anything outside it, and the
paths agents print most (a /tmp capture, Claude's own scratchpad, another
checkout) are outside it by definition. In the response viewer those paths were
not links at all.

- openFilePreview() detects an out-of-workspace path and registers it through
  POST /api/sessions/:id/attachments first, rendering by attachment id. That is
  the surface built for live external files, so the server-side guard is
  unchanged: secret trees blocked, symlinks resolved, extension allowlist. The
  workspace routes keep refusing escapes exactly as before.
- New optional `notify` field on that route. `notify: false` suppresses only the
  attachment:detected broadcast, so a click does not also pop a card announcing
  the file already filling the screen. Default stays true for the CLI and
  publish callers.
- _linkifyFilePaths() links paths in rendered response-viewer markdown. It walks
  text nodes and builds anchors with DOM APIs (the source is model output; never
  a string rebuild of sanitized markup), skips subtrees already inside an <a>,
  and keeps the message text byte-identical so copy-code is unaffected.
- One path pattern in constants.js now feeds both the xterm link provider and
  the chat linkifier, a fresh instance per call since lastIndex is per-object
  state. It picks up /Users and /mnt roots (nothing was clickable on macOS or
  WSL), plus docx/pptx and video/audio extensions.
- .file-preview-overlay moves to z-index 5100, above the response viewer at
  5000. At its old 2000 a path clicked in the chat opened the overlay behind the
  panel it was launched from.

Verified end to end on an isolated instance, desktop and phone viewport: real
clicks in the terminal and the chat both render the image, external md and pdf
render, /etc/hosts is still refused, workspace previews unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 17:03:42 +02:00
Codeman maintainer 1c94995290 docs(skill): hooks are a setting now, not who created the directory
The workspace-hooks install makes the skill's central hooks rule wrong in the
cautious direction. Six places told a worker that a linked case or a raw
workingDir has no `stop`/`blocked` and that send-and-wait cannot be trusted
there, so an agent would hand-roll output-marker synchronization in exactly the
workspaces where `wait:true` now works.

Rewritten against the setting rather than directory provenance:

- verbs.md §5.1: the where-to-spawn table, the rule paragraph (now naming
  `workspaceHooksEnabled`, default ON, the add-only merge, and the boot sweep of
  recovered sessions), and the silent-failure warning. The three cases that stay
  hook-less regardless are called out: remote SSH sessions, docker cases that
  opted out, and a workspace Codeman cannot write to.
- verbs.md §5.3: the send-and-wait precondition is "the workspace has the hooks
  block", not "a case Codeman created".
- endpoints.md: the Signals-by-mode table is now keyed on the setting, with rows
  for OFF, for remote/docker-opt-out, and for a session from an older server.
  The old create-path grep list becomes a "before 1.18.x" note.
- SKILL.md §2 + the cost list, recipes.md Flow-1 contrast, messaging.md step 1.

"Check, do not assume" is kept and promoted to the load-bearing habit, because
the setting is not visible from the call and a session created by an older server
that has not restarted still has nothing.

The `spawn_worker` hooks grep STAYS: it guards the setting being off, remote
sessions, and older servers. Only its diagnostic changes, since "pick an unused
name" is no longer the fix. That text lives in both the §0 heredoc and
`preamble.sh`, which `test/agent-skill.test.ts` pins byte-identical, so both are
patched with the same bytes.

Docs only, no behavior change. 23 skill tests green, full test:ci 5109 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 07:30:10 +02:00
Codeman maintainer f485085174 feat: workspaceHooksEnabled setting as the opt-out for workspace hook installs
Installing hooks into any workspace a Claude session runs in is the right
default, but it takes a decision away from a user who deliberately removed
them: nothing on disk distinguishes "removed on purpose" from "never had any",
so they would come back on the next session create.

Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs ->
Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks
block that is already present is still refreshed when stale (COD-91), but one
is never added.

Every create path routes through one applyWorkspaceHooks() helper so the gate
cannot apply to some paths only, and the boot-time recovery sweep honours it too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 07:05:10 +02:00
Codeman maintainer 19aabe34d2 feat: sort both home-screen session lists by activity, not tab order
The phone overview and the desktop tab rail list the same sessions, so
they now share one order (CodemanSessionOrder in constants.js, pure and
unit-tested): blocked on you first (longest-blocked at the top), then
running longest-turn-first, then quiet most-recently-quiet first.

The tiebreak flips direction halfway down on purpose: for a state a
session is still in, longer is more urgent; for a state it has stopped
in, more recent is more relevant. The running group keys off the pane's
last Enter (lastSubmitAt), never lastActivityAt, because a working pane
repaints about once a second and would rank every turn as freshly
started. A 0 stamp means "unknown" and sorts last within its state.

The desktop rail was previously in raw tab order. Its number badge stays
the Alt+1..9 index, so on a sorted rail it deliberately no longer runs
1,2,3 downward: it names a shortcut, not a row position. Its second
stamp changes from "active 3m ago" to the state duration the order is
computed from ("created 1d ago . working 40m"), since both working rows
otherwise read "active just now" and the order looked arbitrary.

The tab strip itself is untouched: still user-ordered and drag-sortable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 05:11:00 +02:00
Codeman maintainer 98fa8c00d1 fix: install Codeman hooks into every claude workspace, not just cases Codeman created
A session in a linked case (or any pre-existing repo) ran with no hooks block
at all: writeHooksConfig only fires when Codeman CREATES the case directory,
and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven
surface was therefore dead in exactly the place most sessions run: no tab
alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no
Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn,
and no stop/blocked for the agent wait endpoints.

Both session-create paths and restoreMuxSessions() now call
ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and
leaves a malformed settings file untouched. Claude Code re-reads
settings.local.json, so a session already running in the workspace starts
firing hooks without a restart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-16 04:46:01 +02:00
Codeman maintainer 869a507482 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 07:18:41 +02:00
Codeman maintainer 854bcb99aa docs: README Community section + CONTRIBUTING guide
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 07:18:18 +02:00
Ark0N 9ee6bf113b Merge pull request #291 from Ark0N/feat/alerts-lineage-popout
Skill fast-path hardening, lineage retune + colors, per-tab pop-out, reliable tab alerts
2026-08-15 07:16:34 +02:00
Codeman maintainer 66d4c483c7 docs: tab alert screenshots and README glow gif
Captured live from an isolated instance running this branch: a regular
active tab beside a yellow waiting-for-input tab and a red needs-decision
tab. The gif covers one full 17.5s loop (LCM of the 2.5s red and 3.5s
yellow pulse cycles), so it loops cleanly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 07:03:06 +02:00
Codeman maintainer ff13234b3d review fixes: pin alert-overlay opacity against tab-enter's ::before, guard stripBottom against a non-finite strip.top
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 06:49:39 +02:00
Codeman maintainer 0af80b417c feat: skill fast-path hardening, lineage retune + colors, per-tab pop-out, reliable tab alerts
- SKILL.md: forbid the standalone preamble check and pre-spawn recon turns
  (measured: two wasted model turns cost ~12s of a 28s two-worker run; the
  hardened flow measured 20.2s cold / 12.8s warm end to end)
- Lineage lines: dip now hangs from the strip's bottom edge (cap 104 -> 64,
  no stacked row offsets), fixing the deep bow in wrapped strips and keeping
  row-1 arcs off row-2 tab labels; per-child color palette (skin blue first,
  then matrix green, pink, violet, red, turquoise, orange) via an inline
  --lineage-color custom property
- Session Options -> Session: per-TAB pop-out (open-in-window) button override
  on top of the general showTabDetachButton setting; per-device localStorage
  map rendered as the tab-show-detach class
- Tab alerts: seed the pending-hook state machine from GET /api/approvals
  regardless of the approvals-inbox setting (reloads used to lose the red tab
  entirely with the inbox off), clear unconditionally on approval_resolved,
  and repaint the alert as a steady red/yellow ring + glow + status dot on a
  ::before overlay so it stays visible on the selected (active) tab until the
  permission is actually resolved
- docs: worker warm-pool design sketch (verified numbers baked in)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-15 06:38:26 +02:00
Codeman maintainer 52d113ab12 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 15:02:54 +02:00
Codeman maintainer 74662dd788 fix(skill): stale user-level skill copy shadowed injections; seed the preamble
Two live failures from one root cause: Claude Code loads a same-named
user-level skill (~/.claude/skills/codeman, written once by `codeman skill
install`) over the fresh per-case copy, and nothing ever refreshed it. A
stale Aug-9 copy (pre fast-path, pre lineage header) made every agent-driven
spawn run the old recipes: workers spawned serially with pid polls and
without X-Codeman-Parent-Session, so the web UI drew no lineage arcs.

- refreshUserAgentSkill(): session create now refreshes a marker-owned
  user-level copy (refresh-only: absent copies are not installed,
  foreign/symlink copies stay untouched).
- seedAgentSessionPreamble(): local claude session create pre-seeds the
  skill's preamble into ${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh,
  single-sourced from the new skills/codeman/preamble.sh, so the skill's §0
  bootstrap collapses to a two-line loader instead of a ~150-line paste the
  model has to type out (measured ~47s of generation per run).
- SKILL.md: §0 now leads with the loader and keeps the full block as the
  stale/missing fallback; explicit verbatim-paste warning (a hand-assembled
  preamble is how the header and the fast-path functions got lost);
  spawn_worker also sends parentSessionId in the body as defense in depth;
  preamble stamp bumped to 1.18.3 so pre-fix cached preambles self-heal.
- test/agent-skill.test.ts pins preamble.sh byte-identical to the SKILL.md
  heredoc and covers seeding (XDG + HOME fallback, 0600) and the user-level
  refresh (absent/stale/foreign).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 14:46:23 +02:00
Codeman maintainer 0a89505358 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:59:04 +02:00
Ark0N 5387587a64 Merge pull request #287 from Ark0N/fix/lineage-line-blue
fix(ui): draw session lineage lines in blue for contrast
2026-08-14 13:58:21 +02:00
Ark0N 9c0a9bf8e3 Merge pull request #288 from Ark0N/feat/skill-fast-path
perf(skill): spawn workers instead of deliberating (codeman agent skill)
2026-08-14 13:58:18 +02:00
Codeman maintainer 210154f96f chore(skill): stamp the preamble 1.18.2 to match the patch release
The changeset ships this as 1.18.2, so the stamp, the bootstrap's grep/write
condition, both re-source guards and the recipes guard all carry 1.18.2 now
instead of a version that would never exist.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:51:30 +02:00
Codeman maintainer bbc960a8ff fix(skill): harden the fast path against the review findings
Fifteen review findings on the fast-path rewrite plus one caught live, all
verified against a real 1.18.1 server before landing:

- sendwait picks a fresh seq (the epoch second) instead of a fixed 2, so a
  second prompt to the same worker is typed instead of silently swallowed as
  an already-applied duplicate; explicit seq remains for deliberate resends
- sendwait self-heals stranded delivery: an Ink repaint occasionally eats the
  Enter (observed live), so a timed-out short first wait sends one bare \r and
  re-waits by resending the identical frame as a tagged duplicate
- spawn_worker verifies the resolved casePath carries Codeman hooks (the same
  /api/hook-event marker the server checks), refusing names that resolve to
  linked or pre-existing hook-less directories instead of running the job in
  what may be the user's real repo
- spawn_worker probes the trust dialog after a short 5s composer wait, not the
  full 45s, restoring the ladder staging verbs.md documents; on a readiness
  miss it deletes the half-spawned session and returns 1 with empty stdout,
  so a prompt can never be typed blind into a trust dialog
- spawn_workers refuses duplicate case names and empty argument lists, and
  keys result files by index
- section 1 is bash 3.2 compatible (indexed arrays, no declare -A), prints the
  full delivered/timedOut/signal tuple per worker with an explicit line for a
  missing result, deletes only workers whose turn really ended (a timeout
  means still working), cleans up spawned siblings when any spawn fails, and
  guards its mktemp
- last_text takes the previous answer as an optional second argument for
  consecutive-turn reads (the transcript briefly serves the prior answer
  after a stop, observed live)
- the stale duplicate bullets in section 1's closing list are gone
- reference/verbs.md joins the mode-list drift guard's file list
- README's skill inventory covers verbs.md and the new SKILL.md shape
- the changeset is minor so the shipped release matches the 1.19.0 stamp

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 13:40:42 +02:00
Codeman maintainer f18097cb23 perf(skill): make the codeman skill spawn workers instead of deliberating
Measured against a live 1.18.1 server, the API does the whole job in about ten
seconds: two cold claude workers spawned and ready in 6.3s, both tasked and both
answers read in 4.0s more. The slowness users reported was agent-side.

Three causes, all of them things the skill taught:

- It taught serial spawning. Nothing in the main document showed `&`/`wait`, so
  "spawn two workers" read as "do the readiness ladder twice", which is one model
  turn per worker.
- It had no spawn primitive. The happy path had to be reassembled on every run from
  where-to-spawn, a four-stage readiness ladder, send-and-wait, the fan-out caveats
  and a recipe with two variants. Each is a decision, and most carry a warning.
- It cost ~16k tokens before the first call, at 3.6:1 prose to code, with 25 warning
  glyphs and 55 occurrences of "never". A document that is mostly failure modes
  teaches caution, and caution bills as thinking tokens.

The preamble now defines the verbs rather than describing them: spawn_worker,
spawn_workers (concurrent), sendwait, last_text. Section 1 composes them into the
whole job in one Bash call and says to stop reading there.

Two ceremonies the measurements retired: the pid poll (one iteration, 33ms, and
wait-output already blocks on the composer) and reading settings.local.json to check
hooks for a case quick-start creates, which always has them. That check stays
required for linked cases and raw paths, where its absence silently breaks
send-and-wait.

The bootstrap's write condition now greps the version stamp, so a stale or truncated
preamble self-heals rather than failing and asking for a manual rm. The stamp line is
kept bare because the grep anchors on it with $; an inline comment there would rewrite
the file on every bootstrap.

Section 5 moved to reference/verbs.md behind an index, cutting the always-paid
SKILL.md from ~16.4k to ~7.6k tokens. Section numbers and anchor slugs are unchanged,
so existing references still resolve; all 201 anchors across the five files were
checked, with the checker positive-controlled against an injected bad link.

Verified by extracting the code blocks from the shipped file and running them against
the live server: bootstrap plus full fast path, two workers resolving on the
definitive stop signal, answers read and sessions deleted, in 6.8s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 10:20:41 +02:00
Codeman maintainer 62b0039dc5 fix(ui): draw session lineage lines in blue for contrast
Follow-up to #285. Violet sits close to the terminal's own dim foreground,
so the arcs lost contrast exactly where they cross text, which is most of
their length. Blue reads at a glance on the dark skins and on the light
ones.

Colour still comes from a token every skin block already defines and tunes
for its own background (--session-blue instead of --session-purple), so it
stays one rule for all seven skins with no per-skin override, and the two
blues are not even the same: --session-blue is per palette while the
subagent rule hardcodes #3b82f6.

Hue no longer separates this layer from the subagent lines, so the
separation now rests entirely on shape (a lineage arc hangs under the strip
and never reaches a window), weight and dash pattern. Noted in the rule.

CSS only: no geometry, no markup, no settings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 10:03:10 +02:00
Codeman maintainer 174976fc40 Merge origin/master (1.18.1 release) 2026-08-14 01:17:58 +02:00
Codeman maintainer 5ae54536cb chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 01:17:42 +02:00
Ark0N 2d2a455dd2 Merge pull request #286 from Ark0N/fix/terminal-history-scroll
fix(terminal): preserve scroll intent across keyboard resize, surface history truncation
2026-08-14 01:16:01 +02:00
Codeman maintainer 943f04ba53 Merge master into fix/terminal-history-scroll 2026-08-14 01:01:24 +02:00
Ark0N 69d8a9ea6f Merge pull request #285 from Ark0N/fix/lineage-line-visibility
fix(ui): make session lineage lines read as arcs, not straight threads
2026-08-14 01:01:05 +02:00
Ark0N 405eb50ba3 Merge pull request #284 from Ark0N/fix/file-viewer-video
fix(file-viewer): make previewed video seekable and stop it on close
2026-08-14 01:00:55 +02:00
Codeman maintainer b6f15b30c6 docs: correct the rewrite-anchor comment now that refresh pulls full history
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:56:01 +02:00
Codeman maintainer 736f35da7f fix(terminal): bail the backpressure refresh on a mid-fetch tab switch
The refresh can now issue two fetches (full history, then the tail as a
downgrade fallback), which widens an existing window where the user switches
tabs mid-flight and this session's history gets painted into the terminal they
are now looking at. Guard it the way _maybeRefetchFullHistory already does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:54:56 +02:00
Codeman maintainer 6866a617a8 fix(terminal): stop the backpressure refresh yanking and shrinking the buffer
Two further instances of the same root cause, both in _onSessionNeedsRefresh,
which is SERVER-triggered (it fires after SSE backpressure clears) so the user
has no gesture to blame the result on.

1. It ended in an unconditional scrollToBottom, so a user quietly reading
   scrollback was dropped to the live output by a background event. It now
   holds their place. The rewrite REPLACES the buffer, so an absolute viewportY
   captured beforehand is meaningless afterwards; distance from the bottom is
   the anchor that survives, via computeRewriteScrollLine().

2. It rebuilt the terminal from a 1MB TAIL. Measured end to end on a 900-line
   shell pane: an 869-row buffer came back as 158 rows, so the refresh meant to
   REPAIR the display was destroying most of the scrollback every time it ran.
   It now asks for full history, and falls back to the tail only when
   _replayWouldShrinkBuffer refuses the capture, which keeps repaint-mode panes
   (tmux holds roughly one frame for them) exactly as they were.

Also records truncation state here, so the #258 banner stops describing the
pre-refresh buffer.

Verified in a real browser against a live session: baseY 869 -> 869 where it
used to be 869 -> 158, a reader 200 lines up stays 200 lines up, and a follower
stays pinned to the bottom.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:51:58 +02:00
Codeman maintainer a415948736 fix(ui): make session lineage lines read as arcs, not straight threads
The lines that join a tab to the workers its codeman skill spawned were
drawn with numbers tuned against two tabs sitting side by side, and they
degraded in exactly the two situations the feature is actually used in.

1. A spawned worker is appended to the END of the strip, so the real span
   between a lead and its worker is 800-1500px. With the dip clamped at
   44px that is a 33px sag: the arc reads as a straight line drawn across
   the terminal instead of a bracket hanging under the strip. The dip now
   grows at 0.085/px and clamps at 104.

2. When the desktop strip wraps (tabs-two-rows / tabs-auto-wrap), a parent
   on row 1 and its child on row 2 are ~14px apart, and the cross-row
   branch drew parent-bottom to child-TOP: a flat line hidden inside the
   row gap, with siblings overprinting each other. Both ends now anchor on
   the tab BOTTOM with the control points below the LOWER row, so a wrapped
   pair gets the same bracket a flat strip gets. That deletes the branch:
   one shape covers both.

Visibility, at 1:1 rather than in a zoomed mockup: 2 -> 2.5px stroke,
4 4 -> 5 5 dashes (lineage-flow moves with them, -16 -> -20), opacity
.55 -> .72, and a second wider glow so the contrast comes from the halo
rather than from more weight, keeping the line under the subagent lines'
3px. A working child is bright (.95) outside the reduced-motion block, so
turning motion off no longer also dims every worker's arc. Sibling nesting
6 -> 8px and the direction dot 3 -> 3.5px to match the heavier stroke.

Verified at 1:1 in a harness driving the real styles.css and the real
computeLineagePath over three layouts (adjacent workers, workers at the
far end of a full strip, wrapped two-row strip) on a dark and a light
skin. test/session-lineage-lines.test.ts pins both regressions.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:28:04 +02:00
Codeman maintainer d68cba9432 fix(file-viewer): make previewed video seekable and stop it on close
Two bugs in the File Viewer's media player, both reproduced in a real
browser against an 18MB mp4 before and after the fix.

1. Closing the preview left the video playing. closeFilePreview() only
   dropped the overlay's `visible` class, which is display:none and
   nothing else, so the audio kept going with no visible player to pause.
   Detaching the element is not a fix either: a detached HTMLMediaElement
   plays on until it is garbage collected. _stopFilePreviewMedia() now
   pauses, drops src and load()s every media element (also on re-open,
   where overwriting innerHTML had the same effect), which additionally
   aborts the in-flight download.

2. The scrub bar was inert. file-raw read the whole file and answered
   200 with no Accept-Ranges, so Chrome reported video.seekable as
   [0, 0] and silently reverted `currentTime = x`; Safari refuses to
   start such media at all. Raw bodies are now streamed and range-aware:
   Accept-Ranges: bytes on every response, 206 + Content-Range for a
   Range request, 416 for one past EOF, and a malformed spec ignored
   (200) per RFC 9110. Parsing is pure in src/web/http-range.ts.

Measured on tmp/codeman-crt-v5-66s.mp4 (18MB, 66.6s):
  before  seekable [0, 0]     seek to 56.6s reverted to 3.9s   close: still playing
  after   seekable [0, 66.56] seek to 56.6s landed at 60.2s    close: paused, NETWORK_EMPTY

Range slices are byte-identical to `dd`, the full-file path is
byte-identical to the file, and the SVG octet-stream/attachment
hardening and the 50MB cap are unchanged (the cap is still checked
before the range).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:14:21 +02:00
Codeman maintainer 497cbe55bd docs(skill): fix the run-endpoint claim and the Flow cross-references
Four documentation defects found while analysing the agent skill against the
code it drives.

The lineage section attributed "deletes its session as soon as the one-shot
prompt returns" to `POST /api/v1/sessions/:id/run`. That is true of
`POST /api/v1/run`, which creates a throwaway session and calls cleanupSession
on both the success and the error path; the per-session route deletes nothing.
Name the right endpoint, and give the real reason the per-session one carries
no lineage: it is not a create call.

While verifying that, the per-session route turned out to be a sharper trap
than documented. `runPrompt()` rejects whenever a PTY already exists, which is
every interactive session, but the route has already returned `{}` with HTTP
200 by then and routes the rejection only to SSE. An agent calling it against
a live worker reads the 200 as delivery. Document it.

`Flow 3b` never existed in recipes.md. The real mapping is Flow 3 = shell
fan-out, Flow 4 = claude fan-out, Flow 5 = worker blocked on a prompt, so the
same sentence was also mislabelling Flow 4. Fixed in SKILL.md and in the
endpoints.md reference to it; every other Flow reference audited and correct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:10:44 +02:00
Codeman maintainer 9a0e665f72 fix(terminal): preserve scroll intent across keyboard resize, surface history truncation
Closes #259, closes #258. Both bottom out in the same gap: nothing tracked
whether the user was following live output or reading history.

#259 — the keyboard path forced the terminal to the bottom unconditionally
(onKeyboardShow/onKeyboardHide passed scrollToBottom:true, applied with no
check), so opening the keyboard while scrolled up yanked the user down. The
settle cycle now captures intent on its FIRST event, before any fit() has
reflowed the buffer, and returns to that anchor when the user was reading.
A later capture would read an already-moved viewportY, which is why the
capture point matters. The param is renamed restoreScroll to match.

Separately, flushPendingWrites gated viewport preservation on
_hasRecentUserScrollUp(), a 1500ms decay window, so a user who scrolled up and
then actually READ for longer lost protection mid-read. Being scrolled up IS
the intent however long ago it was expressed, so it now keys off position.
The recency window stays as a race guard on the sticky scroll-to-bottom.

The full-history repull already held the user's place and is unchanged.

#258 — truncation was reported by a grey line written INTO the terminal
("earlier output truncated"), which scrolls away with the output it describes,
cannot be acted on, and said the same thing whether the rest was one click away
or gone forever. The server set one `truncated` boolean at two sites meaning
opposite things, and the client discarded fullSize and source entirely.

The route now reports truncationReason ('tail' = intentional partial replay,
the rest is retained; 'capped' = the byte ceiling dropped it) plus
retainedBytes, and 'capped' is not downgraded by a later tail cut. The client
renders a dismissible banner outside terminal output with three honest states:
recoverable (offers Load full history), at-ceiling, and exhausted. The Load
button forces past the scroll cooldown but NOT past _replayWouldShrinkBuffer,
which still refuses a downgrade for repaint-mode panes.

The banner is an overlay, not a flex child: FitAddon derives rows/cols from the
terminal parent's computed height, so occupying real layout space would SIGWINCH
the CLI on every truncation-state change.

Verified in a real browser on the 7 skins: banner text and button clear 4.5:1
contrast on all of them, and terminal height is byte-identical with the banner
shown. The first cut used --bg-elevated and --accent-muted, which do not exist,
so light skins rendered a hardcoded dark bar under dark text; it now uses only
tokens every skin redefines.

test/terminal-scroll-intent.test.ts lives outside test/mobile/ deliberately —
that suite is excluded from test:ci, so a guard placed there is invisible to CI.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 00:00:59 +02:00
Codeman maintainer 4bbe2b7ff6 docs: add pi to the mode lists the sixth-backend sweep missed
PR #282 added pi across the prominent surfaces but left the enumerations
that read as exhaustive: the env-prefix allowlist (missing PI_*), the
external-CLI list for stop/blocked, cron's agent types (also missing
antigravity), the narrow-strip mode list, and the claude-only caveats in the
cron and Read My Mind guides. Both READMEs and the four affected docs now agree
with the schema.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 19:40:54 +02:00
Codeman maintainer a7928f5c64 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 18:13:57 +02:00
Codeman maintainer 4fa44f2e55 Merge branch 'fix/sse-stale-watchdog'
Heal a stalled SSE stream: the server's :keepalive comment becomes a named
sse:heartbeat event (comments are invisible to EventSource by spec), and the
client gains a staleness watchdog that forces a reconnect after three missed
beats. Also applies a confirmed rename locally instead of waiting on SSE.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 18:05:41 +02:00
Ark0N 829b202f51 Merge pull request #282 from Ark0N/feat/pi-mode
feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
2026-08-13 18:05:27 +02:00
Codeman maintainer 86234db1ef docs(skill): document the per-CLI availability probes, and guard the family
`GET /api/pi/status` shipped undocumented in the agent skill, and only a human
reading the doc noticed. Turns out none of its five siblings were documented
either, so this adds the whole family in one place: spawning with a mode whose
CLI is absent fails with OPERATION_FAILED rather than falling back, which is
exactly what an agent picking a backend it did not choose needs to know. Pi's
extra `.data.version` is called out, since a false `available:false` there means
an unrelated `pi` is in front on PATH.

On whether the endpoint scanner should also check registered-to-documented:
measured, and NO for the general case. The skill documents 34 of 217 registered
endpoints deliberately (it is an agent guide, not an API reference), so a blanket
reverse check needs a 183-entry allowlist that would fail CI on unrelated route
work and get appended to mechanically, which is worse than the gap it closes.
Grouping by path shape does not save it either: the families that yields are
things like `DELETE /api/<any>/:id`, lumping cases, webviews and docker hosts
together, and it would not have caught this gap anyway (the family had zero
documented members).

What IS cheap is a family the schema can enumerate with no allowlist: the new
assertion derives the agent modes from the Zod enum and requires each one's
`/api/<mode>/status` to be documented, so a seventh backend fails here until it
is. The sibling scanner still proves the other direction, that nothing documented
is a 404. Both mutation-checked: dropping pi's probe fails the new guard, and
documenting a nonexistent probe fails the old one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:55:43 +02:00
Codeman maintainer 86c78fece3 fix(pi): align the doctor with the pi resolver, correct the strip rationale, update the skill
Second review pass on #282, the three items left open after f4dcfbe.

1. `codeman doctor` and the run mode disagreed about pi. The registry entry
   accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
   `--version` output, so the Dependencies panel could report an installed Pi CLI
   on a box where Run Pi stays hidden, which reads as a broken mode rather than a
   missing install. Both sides now share one exported PI_VERSION_REGEX, and
   PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
   shape check is reported MISSING instead of installed-with-unknown-version.
   Only pi sets it; every other tool keeps its current behaviour.

2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
   is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
   is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
   strips the alt-screen toggles anyway. What exclusion actually preserves is
   `\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
   main screen and is mouse-aware). Comment and changeset now say that, and state
   the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
   shell session.

3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
   agents a backend does not exist and understating class-wide caveats by one
   mode. All updated, plus stale session.ts line references refreshed.

Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:41:44 +02:00
Codeman maintainer c790166564 feat(sse): heal a stalled SSE stream with a heartbeat + client watchdog
An EventSource that stops delivering does not always error. A proxy that
idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect:
`onerror` never fires, the header dot stays green, and every SSE-driven surface
(tab status dots, sessions created on another device, renames) freezes until the
user reloads. Nothing on the client tracked stream liveness at all.

The server already wrote a keepalive every 15s, but as an SSE `:keepalive`
COMMENT, and comments are invisible to `EventSource` by spec, so there was
nothing a client could observe.

Server:
- `sse:heartbeat` under a new Transport category in the event registry
  (155 constants now, both counts updated).
- `cleanupDeadClients()` writes that named frame (`{"t":<epoch ms>}`) instead of
  the comment. Interval, tunnel padding and dead-socket eviction are unchanged.
  The write stays per-client rather than going through `broadcast()`: the frame
  carries no session data, so it needs no multi-user owner routing.

Client:
- `computeSseStale()` in constants.js, a pure policy beside
  `computeConnectionLossUi`. Stale only when the transport believes it is
  `connected`, the device is online, and no frame has arrived for 45s (three
  missed heartbeats). The `connected`-only guard is also the loop breaker: a
  forced reconnect leaves that state immediately, so the watchdog cannot re-fire
  while one is in flight.
- The liveness stamp is applied inside `addListener` itself, so the
  `_SSE_HANDLER_MAP` wrappers and the directly-registered listeners all feed it
  from one place instead of three that can drift. The heartbeat's own listener
  is a no-op that exists only to be registered, since `EventSource` drops named
  events nobody listens for.
- A 5s watchdog forces `connectSSE()` when the policy says stale, and is cleared
  at the top of `connectSSE()` and nowhere else (its only teardown path).
  Recovery needs no new sync path: the reconnect re-runs `handleInit`, which
  already rebuilds from the server. `visibilitychange` -> visible checks too,
  riding the existing listener, since a background tab's timers are throttled
  and a wake is exactly when a stream comes back zombie.
- The forced reconnect logs one diagnostic line: if a middlebox ever strips or
  delays heartbeats, the failure mode is "silently reconnects every 45s", which
  is undebuggable from a field report without it.

Tests: `test/sse-staleness.test.ts` (node VM over constants.js, threshold
boundaries and every not-stale guard) and `test/sse-heartbeat.test.ts` (drives
`cleanupDeadClients()` with fake replies: named frame not a comment, parseable
payload, padding only with a tunnel, dead clients still evicted).

Verified end to end on an isolated instance: with the stream closed client-side
(no `onerror`), a rename sticks, an out-of-band session stays invisible, then
the watchdog reconnects on its own and it appears without a reload.

Event names are part of the stable API contract, so this is a MINOR bump.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:28:35 +02:00
Codeman maintainer f4dcfbe6ca fix(pi): close four mode-list gaps in the pi run mode
Review follow-ups on #282. All four are the same failure shape: a list that
enumerates run modes, missed by the sweep that added 'pi'.

1. Cron ignored pi's project-trust clamp. The PR widened CronJobBaseSchema's
   agentType to accept 'pi' but not the matching clamp beside gemini's, so a
   non-granted multi-user owner's cron pi job spawned bare `pi` (pi's own
   defaultProjectTrust, an interactive prompt they can answer "yes" to, which
   loads and EXECUTES repo-local .pi/extensions TypeScript) while the same
   user's UI/API launch was forced to --no-approve. The clamp is now a pure
   exported helper, clampCronExternalCliConfigs(), so both it and gemini's
   previously untested materialization are pinned.

2. POST /api/sessions/:id/interactive auto-enabled the Ralph tracker for pi:
   its denylist covered opencode/codex/gemini/antigravity only. The tracker is
   never fed for an external CLI (_processExpensiveParsers returns early), so a
   pi session reported ralphEnabled and Ralph UI state no sibling backend shows.

3. REMOTE_CLI_BIN had no pi entry, so buildRemoteCliVersionProbeCommand()
   returned null and Session.cliVersion stayed blank for every remote-SSH pi
   session, even though the PR wired the remote launch command and the
   per-mode override schema field.

4. The desktop home rail's badge map had no pi entry, and its lookup falls back
   to '', which is what claude renders. A pi session read as Claude there while
   the tab strip and phone overview badged it correctly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:23:27 +02:00
Codeman maintainer d19895651d fix(rename): apply the server's confirmed name locally instead of waiting on SSE
Renaming a tab appeared to do nothing: the new name only showed after a full
page reload. The PUT always succeeded; what was broken is how the tab strip
learns the result. `finishRename()` re-renders the strip from the client-side
`app.sessions` map, and nothing wrote the new name into that map, so the rename
depended on the `session:updated` SSE frame to carry its own write back. On a
page whose stream has gone quiet without erroring, that frame never lands and
the re-render repaints the stale label.

- `_applyLocalSessionName()` writes the confirmed name into `this.sessions` and
  refreshes cached subagent parent names, mirroring `_onSessionUpdated`.
- `_putSessionName()` returns the stored name or null. `_apiPut` turns a network
  error into a null Response and an API failure into a non-ok status, so a
  rejected rename previously read as success and silently dropped the edit (the
  old try/catch could never fire).
- Both surfaces use them: `startInlineRename()`'s `finishRename` and
  `saveSessionName()`.

Two regression tests: the commit applies the name with no SSE frame dispatched,
and a 500 restores the old label, leaves the map untouched, and toasts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:20:20 +02:00
Codeman maintainer c5b59633d8 feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.

Pi is a different shape of CLI from the other four, and three decisions
follow from that:

- It has NO permission prompts and no sandbox, so there is no
  --dangerously-skip-permissions analog and none was invented. The
  privilege-shaped knob is the tri-state approveProjectTrust, which makes
  pi load and EXECUTE repo-local .pi/extensions TypeScript and install
  missing project packages. clampExternalCliBypassForOwner() therefore
  puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
  --no-approve even when no config was sent, because pi's own default is
  a prompt the session user could answer themselves. That helper had zero
  test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
  share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
  context, so admitting them would widen the allowlist for every mode at
  once. Auth goes through pi's /login or the server's own environment.
  --api-key is deliberately never wired: it would put a provider secret on
  the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
  main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
  mode is runtime-switchable via /settings; that flip was measured to put
  the pane into the alt screen, which the strip would have corrupted.

pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.

Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).

Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:54:47 +02:00
Codeman maintainer f39beb3326 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 02:30:08 +02:00
Codeman maintainer cf3183abf7 chore: version packages
Release 1.16.6: phone overview started/idle stamps, plus fixes for the
selection-dialog keyboard lockout, the accessory bar arrows bypassing the
local-echo overlay, and recovered sessions being restamped as newly created
on every server restart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 23:14:57 +02:00
Codeman maintainer 15a43894f9 chore: version packages 2026-08-11 19:36:19 +02:00
Codeman maintainer e20aa1d4d8 style(settings): pair Save and Close into one tray in the phone sheet header
Below 860px Save moves into the header (a bottom action bar would cost 60px
of a phone sheet), which left the two ways OUT of the sheet sitting side by
side in mismatched shapes: a fat accent pill next to a bare 1.5rem glyph
with no box at all. They are the same decision (save-and-close vs
discard-and-close), hit in the same corner with the same thumb, so they now
share a recessed tray and matching pill geometry and read as one cluster.

- 36px on both, so the tray comes out at 44px including its 3px padding and
  1px border — the same height as the phone header it sits in.
- `.modal-close` gets a real box (36x36, radius 9) only inside the tray; its
  bare-glyph form is still right in a plain modal header.
- Tray colors come from skin tokens (--border/--bg-input). A hardcoded black
  alpha would render as a grey slab on the four light skins, the same trap
  the layout preview frame hit.
- `:has(.set-head-save)` keeps the tray off the sheets that carry a lone x:
  Session Options and Add Case save from inside their own forms.
- The shared focus ring offsets OUTWARD, which inside the tray would draw on
  top of the tray border, so it is inset to ring the button instead.

DOM order stays close-then-save so the focus trap still lands on Close;
row-reverse paints Save to its left.

Verified at 390x844: tray 44px tall, Save 36px, Close 36x36, both radius 9
inside a 12-radius tray. PostCSS-parsed (prettier does not catch an unclosed
CSS block, and styles.css is prettier-ignored by design).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 19:34:21 +02:00
Codeman maintainer aa28ef048c fix(mobile): reconcile the two keyboard-dismiss paths (#279 + #280)
#279 and #280 auto-merge cleanly, but the merged result was red: neither
branch could see the other, and CI cannot see either, because the only test
covering #279 lives in test/mobile/** which test:ci excludes.

Two problems, both in #279's test:

1. The in-terminal case tapped the terminal's top-left corner, i.e. an inert
   transcript row, and asserted focus was retained. That is precisely the
   gesture #280 redefines, so #280 turned it red. Aim it at the PROMPT row
   instead: the one in-terminal tap whose outcome neither PR claims, so it
   still proves the #terminalContainer exemption without asserting the
   toggle's behaviour.

2. The "a real control is exempt" case was VACUOUS. It picked the first
   button measuring >8px, which is .welcome-ralph-link inside the welcome
   overlay hideWelcome() had already hidden: the rect still measures, but
   elementFromPoint at that point returns .xterm-screen, so the case tapped
   the TERMINAL and passed for the wrong reason. It only surfaced because
   #280 changed what a terminal tap does. Require the sampled point to
   actually resolve to the button, and fail loudly when no control is
   usable rather than silently asserting nothing.

Mutation-checked: removing the install, the #terminalContainer exemption,
the control exemption or the `if (moved) return` scroll guard each turns
the test red on its own. The control exemption had no coverage before.

Also fold the duplicated tap slop into one constant: initTerminal's
TAP_THRESHOLD now reads MOBILE_KEYBOARD_DISMISS_TAP_SLOP instead of
re-declaring 8, since a drift between them is exactly the bug the second
#279 commit fixed. And restore the comment the slop constant was inserted
into the middle of, which left "Regions where a tap must NOT dismiss"
sitting above the slop rather than the selector it documents.

test/mobile/keyboard.test.ts: 5 failed | 47 passed (52). Master is
5 failed | 46 passed (51) — the same five pre-existing failures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 19:34:00 +02:00
Ark0N 67f6ed3168 Merge pull request #280 from Lint111/feat/mobile-tap-toggles-keyboard 2026-08-11 19:33:41 +02:00
Ark0N 2d4616f059 Merge pull request #279 from Lint111/feat/mobile-keyboard-dismiss 2026-08-11 19:33:33 +02:00
liorandClaude Opus 5 35f8f9d19f fix(mobile): let a second tap on inert transcript close the keyboard
Every terminal tap re-focuses the hidden textarea, so once the on-screen keyboard
is open the only way to close it is the accessory bar's dismiss chevron. Tapping
the transcript to get the screen back is the obvious gesture and it did nothing.

A tap on INERT content with the keyboard already up now dismisses it. Nothing
else claims that gesture: an inert row has no action to trigger, so by that point
the tap has already done its only other job (the mouse report).

Scoped to 'content' ON PURPOSE. The prompt row ('input') keeps
focus-then-position, so a second tap there still places the caret — that is real
capability and trading it away would be a worse deal than the bug. A separate
test pins it rather than leaving it to the reader.

Actionable rows are unchanged: readbacks, "esc to interrupt" status rows and menu
selections still blur via _isActionableMobileTerminalTap, which runs first.

`keeps the hidden keyboard input focused after an inert Claude transcript tap`
asserted the OLD behaviour and is renamed and inverted, since revising that
behaviour is the point of this change. Its setup already focused the terminal
before tapping, so it was always exercising the second-tap case.

test/terminal-touch-tap.test.ts: 28 tests. The two new ones fail on master —
`closes the keyboard on a second tap of INERT transcript content` behaviourally,
by asserting blur where master re-focuses.

test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same five
pre-existing failures as master, untouched here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 20:31:38 +03:00
liorandClaude Opus 5 c992784681 test(mobile): cover the scroll case in the keyboard-dismiss test
The dismiss handler fired on any touchend, so a scroll closed the keyboard too —
a regression the original test could not see, because it only ever dispatched a
stationary tap.

The helper now takes an optional travel distance and emits touchmove steps, and
the test asserts a 120px scroll leaves the terminal input focused. Removing the
`if (moved) return` guard fails this assertion, so it genuinely pins the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 18:31:09 +03:00
liorandClaude Opus 5 3a7be356ae fix(mobile): do not dismiss the keyboard when a scroll ends
Regression from the dismiss handler in #279: it fired on any touchend,
and a scroll ends in touchend too. Scrolling to read something while composing
closed the keyboard and dropped the composer — worse than the bug it fixed.

Track finger travel from touchstart and only treat a near-stationary gesture as
a tap, using the same 8px TAP_THRESHOLD the terminal's own touch handling uses
so both agree on tap-vs-scroll. Multi-touch is never a dismissing tap.

All three listeners stay passive; nothing calls preventDefault.

Measured on a Pixel-class viewport with a Firefox UA:
  tap                -> dismissed
  scroll (120px)     -> keyboard kept
  micro-drift (4px)  -> dismissed, so an imprecise tap still works

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 18:21:05 +03:00
liorandClaude Opus 5 a6a572e635 fix(mobile): close the on-screen keyboard when tapping outside the terminal
On a phone the terminal holds focus on a hidden textarea, and nothing ever
released it. Once the keyboard was up, tapping the header, the tab strip or any
empty page chrome left it up — covering roughly half the screen with no in-app
way to dismiss it.

Repro, iPhone-class viewport (390x844), claude-mode session, focus the terminal
then tap the header logo:

| | document.activeElement after the tap |
| --- | --- |
| master | textarea.xterm-helper-textarea (keyboard stays up) |
| this branch | body (keyboard closes) |

A document-level touchend handler blurs the terminal input, deliberately scoped
so focus is never stolen from something that wants it:

- only when the terminal input actually holds focus;
- never inside #terminalContainer — _handleMobileTerminalTap already classifies
  and routes those taps and owns that decision;
- never on a control. Anything focusable or clickable is about to take focus
  itself, and the keyboard accessory bar exists to be used WHILE the keyboard is
  open, so dismissing there would fight the user.

Bound to touchend rather than click: a tap meant to dismiss usually is not meant
to activate what sits underneath, and touchend fires before the synthesized
click so the blur lands first. The listener is passive — it never calls
preventDefault.

Test: `dismisses the on-screen keyboard when a tap lands outside the terminal`
in test/mobile/keyboard.test.ts. It fails on master with a BEHAVIOURAL assertion
(`expected 'xterm-helper-textarea' not to contain 'xterm-helper-textarea'`),
not a TypeError, and passes here. It drives real dispatched touch events rather
than calling the helper, because the handler is bound on document and a direct
call would bypass the routing under test.

test/mobile/keyboard.test.ts: 52 tests, 5 failed | 47 passed. Master is 51 tests,
5 failed | 46 passed — the same five pre-existing failures (stale layout and
accessory-bar expectations, a CJK timeout), untouched here.

Full suite: 4944 passed | 12 skipped, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 17:37:27 +03:00
Codeman maintainer 26416f98de chore: version packages 2026-08-10 13:23:45 +02:00
Codeman maintainer 084d7b7328 fix(run-menu): let recent-session rows use the width the menu was given
PR #274 lifted the Run menu's 250px cap to `calc(100vw - 24px)` so a
recent-session row would have room for its worktree pill and parent path.
The rows never took it: `.run-mode-history` is a block scroller, so its
<button> rows are shrink-to-fit and stayed at ~250px inside a 1376px menu,
leaving ~1100px of empty dropdown and no space for `.hist-dir`'s
`flex: 1` + `text-align: right` to expand into.

Rows now fill the menu, and the menu is capped at the 760px one full row
actually costs rather than the whole window.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 13:21:40 +02:00
Ark0N a4cdb352be Merge pull request #244 from Lint111/feat/mobile-terminal-taps
fix(mobile): route terminal taps without breaking keyboard focus
2026-08-10 13:09:29 +02:00
Ark0N d81454b6f9 Merge pull request #275 from Ark0N/feat/claude-voice-integration
feat(voice): dictate through the server's Claude Code login, no API key
2026-08-10 13:09:24 +02:00
Ark0N 00f1b9228a Merge pull request #274 from jordan8037310/fix/run-menu-recent-sessions
fix(run-menu): make Recent Sessions rows legible on macOS (home-prefix regex + width + worktree)
2026-08-10 13:09:18 +02:00
Codeman maintainer 13d069e1e5 Merge remote-tracking branch 'origin/master' into feat/claude-voice-integration
# Conflicts:
#	CLAUDE.md
2026-08-10 12:56:40 +02:00
Codeman maintainer fa4c36c2a5 Merge remote-tracking branch 'origin/master' into pr274-rebase
# Conflicts:
#	src/web/public/session-ui.js
2026-08-10 12:55:32 +02:00
Ark0N fe2c03b2cc Merge pull request #276 from Ark0N/fix/home-path-abbreviation
fix(paths): one home-prefix helper, so path labels abbreviate on Linux and macOS
2026-08-10 12:53:22 +02:00
Ark0N 4e3f7ac36b Merge pull request #277 from Ark0N/feat/readmymind-phase3-part2
feat(readmymind): rethink steer note (phase 3 part 2)
2026-08-10 12:53:19 +02:00
Ark0N 089283e0b3 Merge pull request #278 from Ark0N/appsettings-details
One settings surface: App Settings, Session Options and Add Case
2026-08-10 12:52:43 +02:00
liorandClaude Opus 5 3b85001fed fix(mobile): keep the keyboard reachable when the viewport is scrolled up
Addresses the review on #244.

BLOCKING (item 1). selectSession() ends with scrollToLastNonEmptyLine(), which
parks the viewport above the bottom for any session taller than the screen, so
after a tab switch every tap classified as 'history' — touchstart ran
preventDefault() + blur, and touchend's early return skipped focus. Both routes
to focus closed on one gesture, the same mechanism as #173.

Suppressing the mouse REPORT while scrolled up is right and is kept; suppressing
FOCUS is not. touchstart now only preventDefaults 'content' taps (a scrolled-up
viewport sends nothing, so there is no compatibility click worth cancelling), and
the 'history' branch focuses instead of blurring.

Verified against the maintainer's own test, which was already on master and red:
`keeps the terminal input focusable after a tab switch parks the viewport
off-bottom` fails without this change and passes with it.

Item 2: dropped both `terminal-action-pending` guards. The class exists nowhere
in the repo, so both branches were permanently false and the comment promised
coverage that did not exist.

Item 3: removed the `Working` literals. Live claude 2.1.226 prints
"Cooked for 2m 6s" with a different bullet and a randomised verb, so they were
dead code. The status row is matched by its affordance ("esc to interrupt")
instead, which is what makes it actionable. The affordance regex is also
tightened to require a key or gesture name, so prose like "click here to open
the file" no longer dismisses the keyboard.

Item 4: removed _shouldForwardTouchScrollToApp and its test. It was never called,
and wiring it as written would have restricted forwarding to claude only,
dropping gemini from the path #205 established — a behaviour change this PR has
no reason to make.

Smaller items: the touchstart classification is cached and reused for the
touchend of the same gesture (keyed on exact coordinates, so a moved finger
re-classifies), removing two of the three full-viewport scans per gesture; the
duplicated touchLastX assignment is gone; and the no-touch bail-out returns null
rather than claiming 'history'.

test/mobile/keyboard.test.ts: 51 tests, 5 failed | 46 passed — the same 5
pre-existing failures as master, unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 13:32:11 +03:00
Codeman maintainer 1513067a7f feat(settings): lead with version + update, tail the rest of System
App Settings opened on a System section that mixed the two things worth seeing
immediately (what this install runs, whether a newer release is waiting) with
three groups nobody sets twice (CLAUDE.md template path, default working
directory, image watcher, Cloudflare tunnel).

Split in two. **Updates** is now the first section and carries only the current
version and the update action, so the modal opens on it and the second thing in
reach is Terminal & Input, where Local Echo lives. **System** keeps Paths,
Automation and Remote access and tails the document, last in the rail.

Also fixes the admin-ui load-order test, which broke on this branch: it located
the modules with a bare `indexOf('session-ui.js')`, and the modal markup now
cites those modules in comments well above the script tags, so it was comparing
a comment against a `<script src>`. It matches the script tag itself now.
2026-08-10 12:29:00 +02:00
liorandClaude Opus 5 623fedf5b7 fix(mobile): keep the keyboard reachable on inert transcript taps
A mid-terminal tap on a claude-mode session left document.activeElement on
<body>, so the on-screen keyboard could not be raised and there was no way to
type — the blocker reduced upstream in #173.

_classifyMobileTerminalTap returns 'content' for any non-prompt row, and
_handleMobileTerminalTap blurred on every 'content' tap while touchstart's
preventDefault had already cancelled the compatibility click that would
otherwise focus xterm. Both routes to focus were closed on the same gesture.

Blur now applies only to rows that are actually TUI-owned. The distinguishing
signal is the affordance a CLI prints on or beside the row ("ctrl+r to expand",
"tap to collapse", "esc to interrupt"), not the row's title text — a readback's
title row carries no hint of its own, so the adjacent row is consulted too.
Keying on titles would recognise only the exact strings a fixture happens to
use and would let a real readback keep the keyboard open.

Measured with a real touchstart/touchend gesture, iPhone-class viewport,
claude-mode session, tapping mid-transcript:

  before  document.activeElement = body
  after   document.activeElement = xterm-helper-textarea

Note: upstream master already passes this assertion, so the added test is a
regression guard for this branch, not a test that fails on master.

test/mobile/keyboard.test.ts: 40 tests, 5 failed | 35 passed — the same 5
pre-existing failures as master (stale layout/accessory-bar expectations and a
CJK timeout), unchanged by this commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-10 13:22:59 +03:00
lior 1410362e5b fix(mobile): keep promptless terminal input focusable 2026-08-10 13:22:33 +03:00
lior 92ae46246c fix(mobile): route Claude terminal gestures 2026-08-10 13:21:15 +03:00
lior 6831d79127 fix(mobile): route terminal content taps to the CLI 2026-08-10 13:21:15 +03:00
lior b01ed611c4 fix(mobile): keep keyboard focus taps non-activating 2026-08-10 13:21:15 +03:00
Codeman maintainer 8d094b086c docs: document the settings surface and repoint the moved settings paths
A docs pass landed in this worktree while the preview was up (a respawn loop on
the throwaway session it was serving), and it is the documentation this work
needed, so it is reviewed and kept rather than thrown away.

- docs/architecture-invariants.md gains a "Settings surface" section: the one
  `:is()` scope and why the id-only list preserves specificity, the anatomy,
  the two meanings of the rail, the deliberate two sizes, the phone strip, the
  Add Case adapter, the flex-summary chevron trap, the Respawn ordering, the
  retired tab chrome, and the live preview's clone-the-chip-icon rule.
- Settings paths are repointed everywhere they moved: Display -> Header &
  Panels (header buttons, cron, multi-monitor, response viewer, file viewer),
  Settings -> App Settings -> System -> Updates, Panels -> Header & Panels ->
  Cross-session features (Read My Mind), Display -> Terminal & Input (gesture
  control), Claude Model -> Models -> New Claude sessions.
- Stale counts refreshed (route modules, frontend modules, type files, config
  files) and the typecheck script named.
- browser-testing-guide gains the three modal ids and the `set-*` selectors.
- The styles.css block comment covers all three modals.

Two claims it got wrong are corrected here: an external-CLI session opens
Session Options on the Session tab (`switchOptionsTab('context')`), not
Summary - measured in the browser - and the Cron toggle lives under Header &
Panels -> Scheduling, with no "Header Displays" step under it any more.
2026-08-10 12:18:08 +02:00
Codeman maintainer ecc6f30e24 fix(voice): move Language and Domain keywords into the Provider group
Both are read by every engine (the Claude path sends the language as its base
tag and the keyterms as a recognition hint), but they sat under the "Deepgram
Nova-3" heading, which read as if they only applied to Deepgram. That group now
holds just the API key.

Ids are unchanged, so the getElementById load/save contract in settings-ui.js is
untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 12:17:54 +02:00
Codeman maintainer 7da9fb4d53 fix(cases): give the collapsed Add Case blocks a disclosure chevron
`summary { display: flex }` in the Add Case adapter drops the browser's own
disclosure triangle, so Clone options, Container settings, Advanced SSH,
Discover existing sessions and Advanced container settings rendered as plain
uppercase headings with nothing to say they open. Reported as exactly that.

Each summary now carries an explicit chevron that rotates 180 degrees on
`[open]`, matching the Advanced group in App Settings, plus a hover state on
the row. The default marker is suppressed in both spellings (`list-style` and
`::-webkit-details-marker`) so a browser that would still paint one does not
end up with two.
2026-08-10 11:58:06 +02:00
Codeman maintainer b025047cbf feat(settings): size up the two task modals, lead Respawn with auto-resume
The shared surface is tuned for App Settings: a long, dense document you scan.
Add Case and Session Options are the opposite - a handful of short panels you
act on once - and at that density they read as a few small fields marooned in a
large empty frame, with rail entries too small to aim at.

Both now take the same size-up while App Settings stays tight: 900px wide, a
236px rail with 0.9rem entries and 19px icons, 0.88rem row labels, 0.82rem
fields, and `height: auto` between a 560px floor and 88vh - so the shell is as
tall as the panel showing instead of a fixed box the content rattles in
(Summary opened two thirds empty before).

Respawn is reordered around what people come to it for:

- Auto-resume is a CALLOUT again, not the first row of a list. It is what turns
  a limit-halted overnight run back on, so it gets an accent card, an icon, and
  a hit target covering the whole card (the label wraps its own switch - no
  `for`, since nesting already associates them and the pair has historically
  double-fired). The armed "resumes at HH:MM" note renders inside it.
- Loop control (status + Enable/Stop) moves ABOVE the loop configuration. A
  running loop is the thing you open this tab to see or stop, and Enable is the
  point of the tab either way; it was previously below three groups of config.
- Enable/Stop and the status pill scale with the rows around them.

The Context tab is renamed Session, since "context" only described one of its
three groups, and those groups become Identity / Context window / Behavior.
2026-08-10 11:54:10 +02:00
Codeman maintainer f11bee72f5 Merge branch 'master' into appsettings-details 2026-08-10 11:44:01 +02:00
Codeman maintainer 78356d7fd0 feat(settings): tighten the surface, put Add Case on it, retire the tab chrome
Three things, all on the same surface.

**Tighter.** The shell drops to 760x620 (was 840x700) and the density comes
down with it: rail 176px, doc padding 15px, row padding 5px 10px, group gaps
3px, section head 0.88rem, row label 0.76rem, description 0.645rem. The model
cards were the biggest block in the document and shrink the most (6px 8px
padding, 0.72rem name). The toggle switches keep their size on purpose - only
the space around them was the problem.

**Checkboxes stay checkboxes.** The respawn cycle steps go back to real
checkboxes in a row card (`.set-checks` / `.set-check`) rather than the chips
they briefly became: they are numbered steps of one sequence, not a set of
independent tags, and chips read as the latter.

**Add Case joins the surface.** Same shell, rail and sections; its rail
switches panels like Session Options'. The six panels keep their legacy
`.form-row` markup - every id in them is read back by session-ui.js, so
restructuring the forms would be a lot of risk for no visual gain. Instead an
adapter block scoped to `#createCaseModal .set-doc` maps the old primitives
onto the look: a form row paints as a row card, its label as a row label, its
`.form-hint` as a row description, `<details class="advanced-options">` as a
collapsed group head. `.form-row` everywhere else is untouched.

With that, `.modal-tabs` / `.modal-tab-btn` / `.modal-tab-content` have no
users left, so their CSS is deleted from both stylesheets and the guard in
test/app-settings-structure.test.ts flips from "the settings modal must not
steal these shared classes" to "nothing uses them any more" - a reappearance
now means a modal drifted back off the shared surface.
2026-08-10 11:43:55 +02:00
Codeman maintainer 0da7f652b4 fix(home): stop the desktop home screen clipping, show full tab names
The welcome column was 880px tall inside a 752px overlay on a 1470x842
window, so it ran off both ends (title above the top edge, "Or click Run
to start" below the bottom one) with no way to scroll to either.
.welcome-content is now a flex column bounded at the overlay height with
every child fixed except the Resume list, which shrinks and scrolls
internally. Short windows (<=900px tall) get a tighter rhythm as well, so
the list keeps usable height instead of collapsing to two rows.

The open-tabs rail drops its border-right (the gradient already reads as
docked) and widens 19vw -> 25vw, which stays inside the gutter at the
1180px gate (295px of 310px). The status pill moves from beside the name
down to the created/active stamps line, handing the full row width to the
session name: names render whole instead of ellipsizing
"w34-claudeman: mindreading" into "w34-claudeman: ...", and wrap to a
second line only when they still do not fit.

Verified against the live server with the edited files served into the
page: content fits the overlay at 1180x800 through 2560x1440 and on phone
widths, no clipped names or stamps, no page errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:34:57 +02:00
Codeman maintainer 6ccab925b1 feat(settings): put Session Options on the same surface as App Settings
Session Options was the last modal still wearing the old chrome: a strip of
top tabs over `.form-row` stacks, sitting next to a settings modal that had just
been rebuilt around a rail and grouped row cards. It now uses the same surface.

The `set-*` rules move from `#appSettingsModal` to
`:is(#appSettingsModal, #sessionOptionsModal)`. An `:is()` list takes the
specificity of its most specific argument, and both arguments are ids, so every
rule keeps exactly the weight it had - nothing downstream shifts in the cascade.

What the two modals do NOT share is what the rail means:

- App Settings stays a table of contents over one scrolling document.
- Session Options switches: one `.set-section` visible, `.hidden` on the rest.
  Summary owns its own scroller and Respawn is long, so stacking them into a
  single document would bury both. `switchOptionsTab` now queries
  `.set-rail-item` (it read `.modal-tab-btn` before) and resets the document
  scroll, so a switched-to section starts at its own top.

Phones get a horizontal, scrollable rail strip rather than App Settings' sticky
jump pill, which Session Options has no equivalent of. That is close to the tab
bar it replaces, so the phone gesture is unchanged.

Content is regrouped into the row language - label, description, control pinned
right - across all four sections: usage limits / respawn loop / cycle steps /
loop control, identity / token management / this session, tracker / limits, and
the summary timeline. The three cycle-step checkboxes became chips, which is why
`_syncSettingsChips` now covers both modals and Session Options registers one
delegated change listener per page for them.

Every id and handler the JS reads is preserved, and the component classes it
queries (`.duration-preset-btn`, `.duration-custom-input`, `.color-swatch`,
`.respawn-status-text`, `.run-summary-filters .filter-btn`) are untouched.
`data-claude-only` moved onto the rail entries, so external-CLI sessions still
lose Respawn and Ralph and land on Context.

`.modal-tabs`/`.modal-tab-btn`/`.modal-tab-content` now belong to
#createCaseModal alone. test/session-options-structure.test.ts pins the rail to
section pairing, the ids openSessionOptions reads, the one-visible-section
invariant and the Claude-only entries.
2026-08-10 11:19:56 +02:00
Codeman maintainer 4b51ba306e feat(voice): dictate through the server's Claude Code login, no API key
The mic button previously needed a Deepgram API key, or fell back to the
browser's Web Speech engine. It can now transcribe through the same
speech-to-text service Claude Code's own /voice mode uses, so anyone signed
in to Claude Code on the server gets dictation with no third-party account.

Claude Code's voice mode cannot be driven directly: it opens the HOST's
microphone (sox/arecord), and the CLI runs in a headless tmux pane while the
human is in a browser somewhere else. So capture stays in the browser and only
the transcription backend is borrowed.

Audio goes browser -> Codeman -> Anthropic. The OAuth token never reaches the
page: the browser sends PCM16 (16 kHz mono, produced by an AudioWorklet since
MediaRecorder cannot emit raw PCM) and receives text.

- GET /api/voice/status reports readiness and never the token
- GET /ws/voice/stream relays one dictation, with the same Host/Origin upgrade
  guard as the terminal socket, plus caps on concurrency, stream length and
  frame size
- credentials are read-only: Codeman never refreshes them, since a refresh
  rotates the refresh token and could sign the user out of their own CLI
- claudeVoiceEnabled (synced, default OFF) gates the whole server side
- voiceSettings.provider picks auto/claude/deepgram/webspeech; auto prefers
  Claude, then a configured Deepgram key, then the browser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:19:51 +02:00
Codeman maintainer aaad031510 fix(paths): one home-prefix helper, so labels abbreviate on both platforms
The rule "show ~/project rather than /home/<user>/project" had three
implementations in the frontend, two of them platform-specific in opposite
directions, so each looked correct to whoever wrote it.

- The Run menu's Recent Sessions rows matched /home/<user>/ only. On macOS
  nothing was stripped, so every row spent its first ~19 characters on an
  identical /Users/<user>/ prefix and the left-to-right ellipsis removed the
  tail that identifies the row. That is #273, reported by @jordan8037310, who
  also traced why the menu's 250px cap made it worse: the width was chosen on
  the assumption the abbreviation had run.
- The case-manage list matched /Users/<user> only, the mirror image, so on a
  Linux host no case path was ever abbreviated there. Unreported.

Both now call _shortenHomePath(), which was already correct for both layouts
and already used by the Resume list, Cmd+K, the desktop home rail and the phone
overview. Its regex collapses to one alternation with a lookahead, so a path
that is exactly $HOME renders "~" instead of being left raw, matching what the
case-manage list used to do on macOS.

test/home-path-abbreviation.test.ts pins the helper on both layouts and the
rendered case-manage label, and fails if a fourth copy of the pattern appears in
src/web/public. The Run-menu guard counts helper calls rather than pinning a
source line, so it survives the row restructure in #274.

test/run-mode-ui.test.ts gains a _shortenHomePath stub: its harness loads
session-ui.js without terminal-ui.js, which the real app never does.

Verified against an isolated instance with 27 real cases and 50 history rows:
27 of 27 case paths and 17 of 20 Run menu rows abbreviate, the other 3 are
/tmp paths that correctly stay raw, tooltips keep the full path, no page errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 11:16:58 +02:00
Codeman maintainer 29efd0e970 fix(readmymind): style the modal footer, point the empty-result copy at the steer note
The footer buttons shipped with class="btn btn-secondary/primary", but no
.btn or .btn-secondary rule exists in this codebase, so all four rendered
as unstyled UA buttons. Moved them to the btn-toolbar convention every
other modal footer uses, with a scoped flex-row footer rule (btn-toolbar
is display:flex, block-level) mirroring the runSummaryModal footer.

Send's accent needs a (0,4,0) re-assert: the skin block's bare
.btn-toolbar rule is (0,2,1) under html:not([data-skin="og"]) and beats
.btn-toolbar.btn-primary (0,2,0), the same specificity trap CLAUDE.md
documents for mobile.css. Scoped to this modal; the repo-wide greying of
btn-primary on non-OG skins is pre-existing and left as a design call.

The empty-result copy now points at the steer note sitting right below
it ("Add a steer note and Rethink to try again"), zh-CN updated.

Verified with the steer E2E (still green) plus desktop, phone (390px),
and error-phase screenshots; static guards extended to pin the footer
convention and the accent re-assert.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 11:15:47 +02:00
Codeman maintainer a6cf4c2b2a feat(settings): reorder App Settings, tighten the rows, add a live layout preview
The document side of the settings modal was wider than it needed to be: every
row is text on the left and a switch pinned to the right, so a 960px shell plus
a 62ch cap on the description left a dead gap of ~350px between the two. The
shell is now 840px, the rail 196px, and descriptions run to 78ch, which closes
the gap and makes the right side sit proportionally with the rail.

Section order now leads with what you look at first: System (the version this
install runs and whether an update is waiting, with Updates promoted above
Paths/Automation/Remote access), then Terminal & Input, then Header & Panels.
The modal opens scrolled to System instead of Terminal & Input.

Header & Panels gains two things:

- every chip carries the icon of the button it switches on, so the list reads
  as the header itself rather than as a column of names (File Viewer shows the
  folder button, Cron the clock, and so on);
- a live preview above the chips: a scale model of the app with a header bar,
  right-docked panels, a toolbar and floating windows, rebuilt on every chip
  change so "what does this add" is answered in place, before saving.

The preview owns no icons of its own - it CLONES `.set-chip-ico` out of the
chip - so each icon has exactly one copy in index.html and a chip can never
drift from the button it previews. A chip joins the preview by carrying
`data-preview` (which slot) and `data-preview-order` (where in it); readouts
that are not buttons (plan usage, CPU, font size) use `data-preview-text`
instead. The frame is painted from skin tokens only, since hardcoded black
alphas turned it into a grey slab on the four light skins, and it is marked
`data-i18n-skip`: the mock tab names are decoration, and the labels inside are
copies of chip text i18n has already translated.

Cron moved into its own Scheduling group (it is a toolbar button, not a header
one, and the preview places it accordingly).

test/app-settings-structure.test.ts pins the new contract: the rail and the
document agree on order, System leads with the version above the paths, and
every previewed chip has both an icon to clone and a slot that exists.
2026-08-10 10:58:05 +02:00
Codeman maintainer 831af88579 feat(readmymind): rethink steer note (phase 3 part 2)
Adds the optional free-text steer note to the Read My Mind modal: a
dashed input under the suggestions ("no, I meant the mobile bug") that
rides along as `steer` on every Rethink. The API already accepted it;
this wires the frontend end of the contract.

- Shown whenever Rethink is live (ready AND empty-result phases),
  hidden only while a prediction runs; typed text survives re-runs.
- Enter in the field triggers Rethink, mirroring the prompt field's
  Enter-to-send; a fresh open clears it with the rethink memory.
- Trimmed and capped to the schema's 2000 chars on the way out; a
  plain open still sends an empty body (neither steer nor rejected).
- zh-CN strings for the placeholder and aria-label, phone-sized
  touch target in mobile.css, static guards in the phase-3 test.

Verified with a browser E2E against a live dev server (stubbed predict
endpoint): payload contents, phase visibility, Enter wiring, and
reset-on-reopen all asserted with real keystrokes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 10:41:39 +02:00
Jordan RyanandClaude Opus 5 3a106bd048 fix(run-menu): make Recent Sessions rows legible on macOS
Closes #273. Every row in the Run dropdown's Recent Sessions list rendered as
`/Users/<user>/co…`, indistinguishable from every other row.

The width was the symptom. The cause is that the home-prefix abbreviation
matched `/home/<user>/` only:

    s.workingDir.replace(/^\/home\/[^/]+\//, '~/')

On macOS the prefix is `/Users/<user>/`, so nothing was stripped and every row
spent its first ~19 characters on an identical prefix, with left-to-right
ellipsis cutting the only part that identifies it. The 250px menu cap was
chosen, per its own comment, as "the width at which the common `~/<dir>/<repo>`
+ timestamp recent-session row still fits whole" — sizing that assumes the
abbreviation ran. On Linux it does. On macOS the menu was permanently too
narrow for content it was never actually shortening, which is why this reads
as fine on one platform and broken on the other.

Changes:

- the regex matches `/home/` and `/Users/`
- the row leads with the identifying folder in semibold, with the parent path
  trailing, dimmed and right-aligned, so truncation removes context instead of
  identity
- the menu goes full width above 769px and the history list grows 200px -> 320px.
  Phones keep the compact popover deliberately: mobile.css positions this menu
  itself and a viewport-wide drawer there would cover the composer
- a worktree pill renders from the fields /api/history/sessions already returns
  unprojected (#266/#269), since a worktree's directory basename is often just
  the worktree name and rows stayed ambiguous without it
- a trailing `/.claude/worktrees` is trimmed from the displayed parent path once
  the pill states it, so the repo name stays visible

Verified in a browser at 1440px against a real 38-session history: menu 1416px,
0 of 34 rows clip their project name (was: all of them), 9 worktree pills
render, no page errors.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3
2026-08-10 01:44:14 -04:00
Codeman maintainer 752374abc7 chore: version packages 2026-08-10 04:48:21 +02:00
Codeman maintainer adfc4fbb1c test(mobile): give the shell keyboard bar stub a classList.toggle
A semantic conflict between two PRs that were each green on their own:
#268 added this test with a fake bar element whose classList carries only
add/remove/contains, and #270 added syncReadMyMind() to init(), which
toggles the RMM marker class with an explicit force argument. Neither
branch saw the other, so the failure only appeared once both were on
master. Production is unaffected: init() builds a real element via
document.createElement, which has toggle.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 04:48:16 +02:00
Codeman maintainer b0b058891c docs(readme): document cloning a GitHub repo into a case
The Clone Repo tab shipped in 1.16.2 (#236) but only ever appeared in
docs/architecture-invariants.md, so nothing a user reads first mentioned
that a repository URL is a way to start a case. Adds it to More Features
and to the working-directory row of the create-a-session table, where the
question "how do I get a project in here" actually gets asked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 04:48:08 +02:00
Ark0N 4a1ad8d194 Merge pull request #270 from Ark0N/feat/readmymind-phase3-part1
Read My Mind phase 3 part 1: the modal grows up and reaches phones
2026-08-10 04:38:02 +02:00
Ark0N 193ce6348d Merge pull request #268 from Ark0N/feat/mobile-shell-keyboard-262
feat(mobile): shell keyboard bar with a one-shot Ctrl modifier
2026-08-10 04:33:22 +02:00
Codeman maintainer 8668b4b352 Merge remote-tracking branch 'origin/master' into feat/readmymind-phase3-part1
# Conflicts:
#	CLAUDE.md
#	src/web/public/home-sessions.js
2026-08-10 04:30:33 +02:00
Ark0N 312ca541e6 Merge pull request #271 from Ark0N/feat/app-settings-redesign
feat(settings): rebuild App Settings as a rail over one scrolling document
2026-08-10 04:29:13 +02:00
Ark0N 40ce91f098 Merge pull request #267 from Ark0N/fix/mobile-tab-scroll-257
fix(mobile): make every session tab reachable in the tab strip
2026-08-10 04:27:42 +02:00
Ark0N 250a53125a Merge pull request #264 from Ark0N/fix/history-search-260-261
fix(web): usable past-conversation list (#260) and search that finds past sessions (#261)
2026-08-10 04:27:12 +02:00
Ark0N 14ea9f630f Merge pull request #269 from jordan8037310/feat/session-worktree-label
feat(sessions): show the git worktree (name + branch) on session rows
2026-08-10 04:26:23 +02:00
Codeman maintainer c8ac04662d fix(mobile): apply the one-shot Ctrl on the CJK input path too
onData is not the only way keystrokes reach the PTY. With cjkInputEnabled
on, the CJK textarea owns the keyboard: onData returns early for
everything it swallows, and the focus router even redirects
terminal.focus() into the field, which is exactly where the accessory bar
sends focus after every key. So an armed modifier could neither fire NOR
be spent there — it survived until a session switch or keyboard dismissal
and then turned an innocent keystroke into a control byte, the failure
mode the whole disarm list exists to prevent.

`_handleCjkInput()` is that module's single choke point to the PTY, so
applying the modifier there covers typed characters, IME flushes, Enter,
backspace and arrows in one place, with the same policy as the onData
hook: the next single character is modified, anything longer merely
spends it. A committed CJK word therefore passes through untouched and
still clears the modifier.

Verified against a real shell session with the CJK field focused and
owning input (cjkActive true, focus in #cjkInput). Before: typing c left
a literal c in the pane, `sleep 300` kept running, and Ctrl stayed armed.
After: ^C in the pane, modifier disarmed, plain typing still literal.

Tests: 5 cases driving the real _handleCjkInput against the real bar,
both loaded into one vm scope (the bar is a const singleton, so a shared
script scope is what makes the bare reference resolve). Removing the fix
fails 3 of them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 04:25:31 +02:00
Codeman maintainer 7c2a49d432 fix(mobile): keep the one-shot Ctrl armed through terminal-generated reports
Review of #268 turned up two defects, both verified against a real shell
session on an isolated instance.

1. A tap spent the modifier. The onData hook consumed every chunk while
   armed, but not every chunk is a keystroke: a shell session keeps the
   narrow scrollback strip, so mouse DECSETs reach the browser, and with
   vim/htop running a tap arrives as `\x1b[<0;31;23M`. Measured in the
   real app: armed, one tap, disarmed, and the Ctrl button read as dead.
   The hook now skips mouse and focus reports via a new
   `CodemanTerminalInput.isTerminalFocusOrMouseReport()`; they still reach
   the PTY, they just no longer stand in for the next key. Focus reports
   are covered for the same reason even though FOCUS_ESCAPE_FILTER in
   session.ts strips DECSET 1004 today, since the bar refocuses the
   terminal after every key and would spend the modifier on its own
   `\x1b[I` the moment that filter changed.

2. The armed style did not land on the four light skins. The competing
   rule is (0,3,1), not (0,2,1) as the comments claimed: `:is()` takes the
   specificity of its most specific argument and that list holds
   `.btn-toolbar.btn-shell`, so it outranked the (0,3,0) armed rules in
   both stylesheets. Measured across all seven skins at 390px, armed and
   resting backgrounds were byte-identical on paper-gray, solarized-light,
   catppuccin-latte and rose-pine-dawn. The light-skin rule now excludes
   the state as `.accessory-btn:not(.armed)`, which fixes phone and tablet
   at once; adding another class to the armed rules would only have moved
   the tie.

Tests: 20 more cases in test/mobile-shell-keyboard.test.ts (the report
classifier, the gate's effect on the modifier, and a static guard on the
light-skin selector, since the existing E2E background assertion passes on
a light skin and the browser suite runs the dark default), plus a browser
regression that taps the terminal with mouse reporting on.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 04:07:19 +02:00
Codeman maintainer 8fcfdb1e6e feat(tabs): orbit the working ring around a busy session tab's dot
The desktop home rail and the phone overview both draw a spinning
`tab-load-spin` ring around their green dot while a session works; the
tab strip itself only pulsed. Same ring on the tab dot now, so "working"
reads identically on every surface.

Drawn as a ::after border circle rather than a halo: the skin block sets
`box-shadow: none` on .tab-status.busy to keep tab dots quiet and
outranks any plain class rule, and a pseudo-element sidesteps that
without reintroducing the glow. It is absolutely positioned, so it never
widens the tab or shifts the label, and it is disabled under
prefers-reduced-motion.

Phones keep their existing tell (a 9px dot with a glow) and suppress the
ring: a 15px ring inside a 32px tab would sit on top of the tab name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 04:00:05 +02:00
Codeman maintainer 45ad9de89e feat(settings): rebuild App Settings as a rail over one scrolling document
The modal had grown to 8 tabs that wrapped onto two rows on desktop and
became a horizontal scroller on phones, with a "Display" mega-tab holding
11 sections and ~35 controls. Local Echo sat 60% down it, and the model
settings were split across two tabs whose three controls fought each
other (the 1M Opus toggle's own hint said it was "ignored when a Claude
Model is selected above").

Replaced with a left rail that is a TABLE OF CONTENTS over one scrolling
document: every section stays mounted, the rail follows the scroll, and
find-in-page works across the whole thing. Nine sections:

  Terminal & Input (Local Echo is the first row of the first section)
  Appearance, Header & Panels, Models, Agents & CLIs,
  Notifications, Voice, Shortcuts, System

Models are now one page. The picker is a card grid of BASE models with a
single "1M context window" switch; context becomes a property of the
chosen model and composes back into `claudeModel` as `base + [1m]`, which
retires the precedence trap. Thinking effort is a segmented control on
the same page, and the old Models tab (task routing) becomes a collapsed
Advanced block under it.

The 12 header-button toggles and the 8 panel toggles become chip grids,
which is most of the old Display tab reclaimed. Rows now say whether a
setting is per-device or synced, stated once per group.

Phones drop the rail for a sticky jump pill that names the current
section and opens a jump list, move Save into the header (the bottom
action bar cost 60px), and render groups as one inset rounded list with
hairline dividers instead of a stack of bordered cards.

Load and save are untouched: every control keeps its id, so
openAppSettings()/saveAppSettings() work as before. Model cards and the
effort segment are views over hidden <select>s that stay the source of
truth. test/app-settings-structure.test.ts pins that contract, plus the
rail hooks admin-ui.js injects the multi-user Users section into.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 03:59:43 +02:00
Codeman maintainer a070fc43ea feat(readmymind): alternates row, phone accessory key, phone-sized modal (phase 3 part 1)
The Read My Mind modal grows up and reaches phones:

- Alternate suggestions (the predictor's verify/redirect kinds) now render
  as tappable rows below the main field. Tapping one swaps it into the
  editable field; the edit you were making folds back into the row you
  leave, so toggling between alternates never loses typing. Rethink now
  records the WHOLE shown set (main + alternates) as rejected.
- Phones get a 🧠 key on the keyboard accessory bar (both simple and
  extended layouts), gated on the same synced readMyMindEnabled setting
  via an rmm-enabled marker class on the BAR element: setMode() rebuilds
  the buttons' innerHTML, so per-key state would be wiped. Synced at init
  and re-synced by applyHeaderVisibilitySettings() on every settings
  apply, so a live toggle needs no reload. The header button stays off
  phones.
- On phones the modal renders as a small dialog (mirrors modal-sm) instead
  of the full-screen default, with wrap-friendly finger-sized footer
  buttons. Not modal-sm itself: that caps desktop width at 340px and this
  modal wants 560px there.
- On touch devices the ready/swap paths no longer focus the field, so the
  OS keyboard does not pop over the alternates that just rendered.
- New static guard test/readmymind-phone-key.test.ts pins the dual-template
  key, the marker-class gating, the phone-hidden header button, the
  small-dialog phone modal, and the no-innerHTML discipline.

Part 2 of phase 3 (rethink steering, the free-text steer note) is next;
the API already accepts steer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 03:57:06 +02:00
Jordan RyanandClaude Opus 5 aa35c1a0c4 feat(sessions): show the git worktree on session rows
Closes #266. Sessions from different worktrees of the same repo were
indistinguishable in the Resume list, Cmd+K and search — the row showed a
session name and a case label, nothing about which worktree it ran in.

Claude Code already stamps "cwd" and "gitBranch" on every user/assistant
record, and writes a worktree-state record naming the worktree when the
session was started through its own worktree feature. scanProjectDir()
already buffers the head of every transcript for prompt extraction, so
extractTranscriptGitInfo() parses buffers that are already in memory: no
extra file reads, no git subprocess. (Measured on this machine: a git
rev-parse per directory costs 482ms for 35 rows; parsing the existing
buffers costs nothing.)

cwd is taken from the first record that carries it, since a session's cwd
does not move. gitBranch is taken from the last, since a branch genuinely
changes mid-session.

The badge requires a worktree NAME. An earlier revision rendered whenever a
branch was known, which put a badge on all 35 rows of a real history --
"master" on every ordinary session, burying the ten rows the badge exists to
distinguish. A hand-made `git worktree add` therefore gets no badge rather
than a guessed name; Claude's own <repo>/.claude/worktrees/<name> layout is
recognised from the path when no worktree-state record is present.

worktreeName and gitBranch join the filterAndPaginate haystack so the session
manager can search by them. panels-ui re-projects the unified item into a
5-field record before rendering, so the new fields are carried there
explicitly -- omitting that silently drops them from Cmd+K only.

Also prefers the transcript cwd over decodeProjectKey()'s stat-walked guess,
which falls back to $HOME when nothing resolves (#265). Note that path is
currently LATENT, not active: on the install this was developed against,
every project key whose directory is gone has zero transcripts and so
produces no row at all. The transcript value is used because it is
authoritative and non-lossy, not because a live bug was reproduced.

Verified against a real 35-session history on an isolated CODEMAN_INSTANCE:
10 of 36 rows badged, history row count unchanged at 35 (nothing dropped),
no page errors. 129 tests pass across the new suite plus the unified service,
unified route and session route suites.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3
2026-08-09 21:32:04 -04:00
Codeman maintainer c13b3c55d3 style: drop em-dashes from the prose added in this branch
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 03:15:13 +02:00
Codeman maintainer 053a6d238d fix(web): adopt #263's fetch ceiling, persisted sort and numeric collation
@jordan8037310 opened #263 against the same two issues while this branch
was in flight. Three details there are better than what this had, so they
are folded in with credit:

- the Resume list pulls 200 unified sessions instead of 60, so the filter
  can reach a real backlog rather than stopping at an arbitrary ceiling
  (the endpoint clamps at 500),
- the sort choice persists per device in localStorage, like `codeman:skin`
  and the other display keys that stay out of the synced schema,
- alphabetical sorts collate with `{sensitivity:'base', numeric:true}`, so
  w2- sorts before w10- and case never splits one project's rows apart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 03:12:02 +02:00
Codeman maintainer 9b9f2c21e9 feat(mobile): shell keyboard bar with a one-shot Ctrl modifier (#262)
The mobile accessory bar was built around coding-agent commands, so a shell
session had no way to send Ctrl chords at all.

A shell-mode session now gets its own bar automatically: Ctrl, Esc, Tab,
four arrows, paste, dismiss. Agent sessions (claude, codex, opencode,
gemini, antigravity) keep the existing bar unchanged.

Ctrl is a one-shot modifier: tap it and it lights up, the next character
typed on the system keyboard is sent as its control byte, and Ctrl disarms.
Tapping it again cancels. That puts Ctrl+C/D/Z/R/L/A/E/W/U/K on a
nine-button bar without a button per chord.

Implementation notes:

* The interception lives in terminal.onData, not a keydown handler: a
  virtual keyboard reports no usable key events, so the character only
  exists as onData text. It sits after shouldSuppressTerminalQueryResponse
  (xterm answers DA/CPR queries through onData too, and letting one of those
  spend the modifier would silently eat the user's Ctrl) and before every
  send path, so the control byte follows the normal control-char route.
* ctrlByteFor() maps `code & 0x1f` over @A-Z[\]^_ and a-z, plus
  Ctrl+Space = NUL and Ctrl+? = DEL. Characters with no control equivalent
  pass through unchanged, like a hardware keyboard.
* The bar now separates the base layout (the extendedKeyboardBar setting)
  from the effective one, resolved per session by refreshForActiveSession().
  A settings save during a shell session cannot yank the bar away, and
  switching back to an agent tab restores the user's choice.
* Ctrl disarms on use, a second tap, any other accessory key, a session
  switch, keyboard dismissal and a layout swap.
* Ctrl joins the refocus set, so tapping it keeps the terminal focused and
  the keyboard open.
* The armed style needs three classes to outrank mobile.css's light-skin
  .accessory-btn rule at (0,2,1).

Verified end to end against a real shell session on an isolated instance:
tapping Ctrl then typing c interrupted a running `sleep 300` (^C in the
pane), the modifier disarmed, plain typing stayed literal, Ctrl+L cleared,
and a cancelled Ctrl typed a literal c.

Tests: test/mobile-shell-keyboard.test.ts (new, runs in CI) covers the
mapping table, layout selection per session mode, base-mode memory and every
disarm path; test/mobile/keyboard.test.ts adds nine browser regressions that
drive the real xterm with page.keyboard.type() and assert on the bytes that
would go out.

Closes #262

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 03:09:44 +02:00
Codeman maintainer a80eda8e4c fix(mobile): make every session tab reachable in the tab strip (#257)
With five tabs open on a phone, the right-hand tabs were effectively
unreachable. Selecting a tab only toggled the .active class, so the strip
never moved, and every full rebuild (a task badge appearing, a session
created elsewhere) replaced the strip's innerHTML, which resets scrollLeft
to 0 and yanked a mid-swipe strip back to the first tab.

Three changes, which only work together:

* computeTabScrollLeft() (pure, constants.js) decides the scroll target from
  measured rects, and _scrollActiveTabIntoView() applies it on selection.
  Rect math on the strip's own scrollLeft rather than scrollIntoView(), which
  also scrolls ancestors: on a phone that is the document, under a fixed
  header and possibly an open keyboard.
* _fullRenderSessionTabs() saves and restores scrollLeft across the rebuild,
  and re-reveals the active tab only when it actually changed
  (_lastRenderedActiveTabId), so a background render never undoes a manual
  swipe.
* Mobile no longer hoists the active session to the front of the strip. That
  reordering ran on full renders only, so tab order flipped depending on
  which render path fired, and it renumbered the Alt+N badges. Scrolling the
  active tab into view replaces it.

Also sets overscroll-behavior-x: contain on the strip so a swipe that runs
past the last tab stays in the strip instead of becoming the browser's back
gesture.

Tests: scroll-target math in test/tab-overflow.test.ts (runs in CI), plus
five browser regressions in test/mobile/tabs.test.ts covering reveal-on-
select in both directions, scroll preservation across an ambient rebuild,
sessionOrder rendering on phones, and a real touch drag reaching the last
tab.

Closes #257

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 03:06:23 +02:00
Codeman maintainer 5d42f64393 fix(web): usable past-conversation list, and search that finds past sessions
Two home-screen reports from @jordan8037310, both about history that is
present but unreachable.

#260 — "Resume Conversation" rendered 4 rows, then a button that appended
every remaining row into a `max-height: 240px` box, so 35 conversations
landed in a four-row scroll well with no ordering or filtering. Rendering
now goes through `_renderHistoryList()` over a cached corpus: 10 rows to
start, Show more/Show less that grows and shrinks the box (the height cap
is class-driven, `.history-list.expanded`), plus a filter box (name,
folder, #case label, prompts), a sort control (recent / name / folder,
pinned rows still first) and a shown-of-total count. A filter implies
expansion, so every match is visible, and the whole header hides as one
unit while a federated search is active. The A-Z sort keys off the same
string the row renders, since most rows are transcript-backed and carry
no session name at all.

#261 — the search box could not match a past project by folder name:
`harvestSources()` built its session corpus from the live in-memory map,
while past sessions come from `/api/sessions/unified` (lifecycle log +
transcript scan). Folding that scan into the request path would have cost
the search its no-filesystem-reads property, so the corpus arrives via a
bounded snapshot instead: `session-history-index.ts` is published as a
side effect of `/api/sessions/unified` (the home screen fetches it on
open, which is the same screen the search box lives on) and rebuilt
fire-and-forget, single-flight and TTL-guarded when a search finds it
stale. A result for a closed session now resumes the conversation rather
than selecting a tab that no longer exists, and is badged RESUME.

The snapshot is stored unscoped with a per-row owner and re-filtered
through canAccessOwned() on read, so multi-user sees exactly what
/api/sessions/unified exposes: own sessions only, host-wide transcript
history admin-only. Live rows are harvested first and win the dedupe.

Verified end-to-end against a real instance with 60 past sessions: cold
process answers its first search without history and its second with it;
folder-name queries return resume targets; clicking one posts the right
resumeSessionId + workingDir.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 03:02:11 +02:00
Codeman maintainer c891a8045d feat(home): dock the desktop tab list as a left rail with age stamps
The open-tabs list on the welcome screen was a fixed 256px card floating
vertically centered in the left gutter, which read as debris rather than
chrome and left 12px type stranded on a wide display.

- Dock it: left/top/bottom 0, full height, hairline right border and a soft
  background fade. The centered welcome content still does not move.
- Scale it off one knob: width clamp(250px, 19vw, 430px) plus a fluid
  font-size on .home-sessions, every child sized in em. Measured 250px/12.2px
  at the 1180px gate, 380px/15px at 2000px, 430px/17px at 2938px; the gap to
  the centered content never goes negative.
- Show when each session was first created and last active, on a full-width
  footer line so it does not fight the status pill, exact dates in the title.
  Both stamps refresh in place on a 20s clock (disarmed when the home screen
  goes away) rather than by re-rendering, which would restart every row's
  blink animation and working ring twice a minute.
- Mute idle green: dot and pill mix toward --text-muted, so idle reads as
  greyed-out next to the vivid green of a working session. Mixed rather than
  hardcoded, so every skin keeps its own green.

Verified in a browser at 1180/2000/2938px and on a light skin, plus
test/home-sessions.test.ts, frontend-syntax, public-assets, prettier and a
PostCSS parse of styles.css.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-10 02:56:35 +02:00
Codeman maintainer c942bb5dfb chore: version packages 2026-08-10 00:55:13 +02:00
Ark0N f98922063a Merge pull request #256 from Ark0N/feat/readmymind-phase2
Read My Mind phase 2: the predictor and the 🧠 button
2026-08-10 00:54:27 +02:00
Codeman maintainer 5671c20076 Merge remote-tracking branch 'origin/master' into feat/readmymind-phase2
# Conflicts:
#	CLAUDE.md
2026-08-10 00:45:59 +02:00
Ark0N d5375d7f0b Merge pull request #251 from Ark0N/feat/clone-repo-case
feat(cases): clone a Git repository as a new case (#236)
2026-08-10 00:45:13 +02:00
Codeman maintainer 1692238531 Merge remote-tracking branch 'origin/master' into pr251-review-fixes
# Conflicts:
#	CLAUDE.md
2026-08-10 00:29:19 +02:00
Codeman maintainer f9510f8a54 fix(clone): route EVERY settings writer through one safe-write gate
Round 2 of the #251 review: settingsWriteBlocker covered only
writeHooksConfig and updateCaseModel, while applyStatusLineConfig,
stripCaseEnvKeys, updateCaseEnvVars, refreshStaleCodemanHooks and
ensureCodemanHooks still wrote the same repository-controlled path
unguarded (applyStatusLineConfig was demonstrated writing through a
symlinked settings.local.json).

All seven writers now go through withSafeSettingsWrite(), which runs
the blocker check INSIDE the per-path settings lock and then hands the
writer its claudeDir/settingsPath; none of them touch the settings path
directly anymore. Test pins all seven against a symlinked
settings.local.json at once (link target must stay byte-identical).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 00:28:50 +02:00
Codeman maintainer 62ca7f1381 chore: retrigger CI (synchronize event was dropped) 2026-08-10 00:18:50 +02:00
Codeman maintainer 93df8188a5 fix(clone): harden per review: symlink-safe scaffolding, race-safe cleanup, decode guard, bounded git queue
Addresses all four findings from the #251 review:

- Scaffolding no longer writes through repository-controlled symlinks.
  The guard lives in hooks-config.ts (settingsWriteBlocker) so it also
  covers quick-start/docker/ralph writers, not just the clone route:
  refuses a symlinked .claude or settings.local.json, a .claude that is
  a file, or one resolving outside the case. The clone route surfaces
  the refusal as a user-visible warning, and the CLAUDE.md write checks
  presence via lstat so a BROKEN repo-shipped symlink counts as present
  (existsSync follows links and would have created the outside target).

- Failed-clone cleanup can no longer delete a concurrent winner's tree:
  git clones into an attempt-owned temp sibling (.<name>.cloning-<rand>)
  which is atomically renamed into place; the loser reports
  DESTINATION_EXISTS and only ever removes its own temp dir.

- decodeURIComponent(url.pathname) is guarded: malformed percent-escapes
  now come back as BAD_SYNTAX instead of an uncaught URIError 500.

- The git pool's waiter queue is bounded (CODEMAN_MAX_GIT_QUEUE, default
  16): overflow answers BUSY immediately (HTTP 429 via RATE_LIMITED),
  and queue time counts against the operation's own deadline.

Tests: hostile symlink fixture repo (route level), settingsWriteBlocker
units, concurrent same-destination race, temp-dir leak assertions,
percent-escape rejection, and a fake-git pool-bounds suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-10 00:07:01 +02:00
Codeman maintainer 94abcf29dc feat: Read My Mind phase 2, the predictor and the brain button
The feature as pitched in docs/readmymind-plan.md: pressing the header
brain button predicts the prompt you were about to type, from the case's
intent profile plus everything the session already knows.

Backend:
- readmymind-context.ts: pure budgeted context assembler (9 ranked
  sources: pending approval dialog, user goals, last assistant turn tail,
  recent prompts, tool activity, git workspace signals, away context,
  sibling sessions, rethink state; 30 KB budget, whole-section drop from
  the bottom of the ranking, trust tiers stated in the prompt)
- readmymind-collectors.ts: transcript tail reader (the live watcher
  keeps only a 500-char snippet) and git signal collection (execFile,
  2s timeout, skipped for remote-SSH cases)
- readmymind-predictor.ts: one-shot claude -p in a throwaway tmux
  session, opus by default (readMyMindModel setting), strict JSON
  contract with 1-3 suggestions (continue / verify / redirect), newline
  stripping, 90s timeout; mutable singleton so route tests can stub it
- POST /api/sessions/:id/readmymind: claude-mode only (400), one
  prediction in flight per session (409 CONFLICT), rethink body
  { steer, rejected }; ownership via findSessionOrFail

Frontend:
- readmymind-ui.js (loadorder 11.3): header brain button, marker-hidden
  until readMyMindEnabled is ON, desktop only (phone key is phase 3);
  modal with editable suggestion + rationale and Send / Insert /
  Rethink / Dismiss; suggestion text rendered via value/textContent only
  and nothing ever auto-sends
- App Settings -> Panels checkbox for readMyMindEnabled; en + zh-CN
  strings

Verified end to end against a live isolated instance: transcript
capture, a real opus prediction grounded in the stated goals, rethink
steering, the 409, and the browser modal incl. Insert leaving the text
unsubmitted on the composer. 41 new unit/route tests; full test:ci
sweep green (4680 tests).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 23:05:11 +02:00
Codeman maintainer 23d91a6ee1 feat(sessions): allow per-session CLAUDE_CONFIG_DIR env override (#255)
Adds an exact-key tier (ALLOWED_ENV_KEYS) beside ALLOWED_ENV_PREFIXES in
schemas.ts, admitting CLAUDE_CONFIG_DIR so a case can run on a separate
Claude subscription (client-billed accounts). Exact match only: other
CLAUDE_* keys and near-misses like CLAUDE_CONFIG_DIR_EXTRA stay rejected,
blocked keys stay blocked. The key also survives getEnvOverridesForPersist()
(a path, not a secret; dropping it would silently switch a rebuilt session
back to the default account after a reboot).

Docs cover the transcript caveat: a relocated config dir writes transcripts
outside ~/.claude/projects, so response viewer / subagent windows /
ultracode / Read My Mind go blind for that session unless projects is
symlinked back into the shared tree.

Design and spec contributed by @jordan8037310 in #255. Closes #255.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 22:39:07 +02:00
Codeman maintainer 8a6570e22d chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 18:35:29 +02:00
Codeman maintainer 0aafabd28d feat(mobile): 44px phone header, making the home button a true 44x44 target
The brand "C" got a 44px-wide hit box in the previous commit but was capped at
36px tall by the bar it sits in. The phone header is now 44px, so the one
control that gets you back to the home screen is square at the platform
minimum, and every other header control gains the same 8px.

Redefined as --header-height inside the phone media query rather than as a
literal, so the panels positioned off that token (file browser, project
insights, plan overlays) follow the bar instead of drifting 8px underneath it;
.app's top offset is derived from it for the same reason. The header also stops
top-aligning its children on phones: that read as centred in a 36px bar whose
contents were ~31px, and leaves a visible gap under everything at 44px.

Costs 8px of terminal height on a phone.

Verified on a real isolated instance at 390px: header 44px, button 44x44
spanning the bar, a touch tap at (4,41) - inside the new area, outside the old
one - reaches the home screen, tabs centred, and content still clears the fixed
header. Tablet (48px) and desktop are untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 18:34:35 +02:00
Codeman maintainer 4add38c4b1 feat(home): open-tab column on the desktop home screen; bigger phone home button
The welcome overlay centers ~560px of content in a ~1400px window, so both
gutters are dead space. The left one now carries the open tabs as a vertical
list (home-sessions.js): one row per live session plus saved web tabs, in TAB
order rather than by urgency, because the row badges are the Alt+1..9 indices.
Clicking a row enters that session.

Working state is deliberately the phone's, exactly: a pulsing green dot ringed
by the same tab-load-spin the tab strip uses while a tab loads, now with a green
halo added on both surfaces so "working" reads identically wherever you see it.

The column is position:absolute so the centered content never moves, which is
why it needs a width gate in two places (HOME_SESSIONS_MIN_WIDTH = 1180 in JS,
a max-width: 1179px media query as the backstop for a resize that outruns the
matchMedia listener). A test pins the two equal. State classification is reused
from mobile-overview.js rather than re-derived, so the two home screens cannot
disagree about what counts as needing you.

Phones keep the mobile overview, and their brand "C" was a 0.85rem inline span,
roughly a 12x13px target on the one control that gets you back to that screen.
It is now a 44px-wide button filling the full header height, with the glyph
scaled to match. 44 is horizontal only: the phone header is pinned to 36px and
clips overflow, so a true 44x44 would mean taking height off the terminal.

Verified end to end against a real isolated instance (own tmux socket + data
dir): 18 browser checks covering render, live update through the tab renderer,
the working dot's animation/glow/ring, row click, the narrow-window gate, the
phone fallback, and a real touch tap on the far corner of the new hit box.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 18:34:35 +02:00
Ark0N e6df0c4094 Merge pull request #253 from Ark0N/feat/readmymind
feat: Read My Mind phase 1, per-case intent profiles (opt-in)
2026-08-09 18:34:11 +02:00
Codeman maintainer 6bb3d66004 docs: Read My Mind user guide (enable, capture rules, privacy, API, troubleshooting)
docs/readmymind.md covers phase 1 as a user guide: how to enable the synced
readMyMindEnabled setting via the API (no UI checkbox until phase 2), exactly
what is and is not captured, the hooks dependency (Docker bridge / remote-SSH
caveats), storage and wipe paths, curl examples for the three endpoints, the
agent-skill ground rules, and a troubleshooting table. Cross-linked from the
CLAUDE.md Key Patterns entry and the api-reference section.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 18:22:52 +02:00
Codeman maintainer 161f1da2eb feat: Read My Mind phase 1, per-case intent profiles (capture + API + skill)
Per-case profiles of user intent (docs/readmymind-plan.md): user-stated goals
plus the user's recently submitted prompts, captured from the Claude session
transcript behind the new synced readMyMindEnabled setting (default OFF).

- intent-store.ts: keyed by owner + realpath(workingDir), FIFO/size caps,
  consecutive-dupe collapse, atomic 0600 writes to ~/.codeman/intents.json
- transcript-watcher.ts: new transcript:user_prompt event for typed user turns
  (tool_result-only entries stay silent); capture wiring in server.ts is
  claude-only and gated on the setting per event
- readmymind-routes.ts: GET/PUT/DELETE /api/sessions/:id/intent, ownership
  via findSessionOrFail, strict Zod schema
- agent skill: SKILL.md recipe + endpoints.md rows so agents can read and
  record intent (PUT replaces: read + merge; never delete unprompted)
- groundwork for the phase-2 predictor button; nothing is ever auto-sent

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 18:03:13 +02:00
Codeman maintainer 87e787e934 test(mobile): guard the phone keyboard against off-bottom tap routing
selectSession() ends with scrollToLastNonEmptyLine(), which parks the viewport
one row ABOVE the bottom for any session whose buffer is taller than the screen
and ends in blank rows, so that is the normal state after a tab switch. Nothing
pinned that a tap there still leaves the keyboard reachable.

The blocker reduced in #173 came back through exactly that gap in #244: a tap
classifier that treats "viewport is scrolled up" as a reason to blur, paired
with touchstart preventDefault cancelling the compatibility click, closes both
routes to focus on the same gesture and strands document.activeElement on
<body> with no way to type. The prompt row is no exception.

Measured on a 390x844 viewport, claude-mode session, dispatched touch gesture:
master leaves focus on textarea.xterm-helper-textarea, PR #244's terminal-ui.js
leaves it on body. Green here, red against that branch.

The test also pins the half that IS correct: SGR coordinates are meaningless
off-bottom, so the tap must send no mouse report.

It has to be a dispatched gesture. Calling the touchend handler directly
bypasses touchstart's preventDefault, which is half of what closes the focus
path, so a direct call reports the right intent and still misses the bug.

test/mobile/keyboard.test.ts: 4 failed | 32 passed (36), against 4 failed |
31 passed (35) without it. Same four pre-existing failures either way.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 17:52:49 +02:00
Codeman maintainer 3533c4332b feat(mobile): bigger green working dot on tabs; Tab key replaces /clear in the simple keyboard bar
The working dot is the one glance-state a phone needs: busy tabs now get a
9px pulsing dot with a green glow (idle stays 4px). The glow needs !important
because the skin block's no-halo rule outranks mobile.css.

The simple keyboard accessory bar swaps /clear for Tab (/clear and /compact
stay in the extended bar with their double-tap confirm). The tab action now
flushes locally-buffered prompt text to the PTY before sending \t, so
completion applies to what was just typed instead of an empty composer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 17:32:24 +02:00
Codeman maintainer 6fc772f697 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 17:10:29 +02:00
Ark0N 527ce10491 Merge pull request #248 from Ark0N/feat/offline-state
feat(web): make a dead connection unmistakable instead of a red dot
2026-08-09 17:08:15 +02:00
Codeman maintainer 26a4dd2879 Merge master into feat/offline-state (keep both offline overlay and approvals drawer) 2026-08-09 17:01:34 +02:00
Ark0N e087198056 Merge pull request #250 from Ark0N/feat/path-picker-show-hidden
feat(path-picker): show hidden files and folders, and harden the secret blocklist
2026-08-09 16:59:33 +02:00
Ark0N 2e266380f8 Merge pull request #247 from Ark0N/feat/file-viewer-show-hidden
feat(file-viewer): show hidden files and folders
2026-08-09 16:59:14 +02:00
Ark0N 3363d25876 Merge pull request #245 from Ark0N/feat/approvals-inbox
feat: Approvals Inbox, answer any session's pending prompt from one place (opt-in)
2026-08-09 16:58:56 +02:00
Ark0N b793ff3294 Merge pull request #249 from Ark0N/fix/trust-dialog-auto-accept
fix: workspace trust dialog auto-accept has been dead (tmux sends cursor-forwards, not spaces)
2026-08-09 16:58:30 +02:00
Ark0N a68b2c5bc5 Merge pull request #246 from Ark0N/fix/idle-detection-working-state
fix: sessions reported idle while working, plus a working state you can see
2026-08-09 16:58:04 +02:00
Ark0N 89f9e0becb Merge pull request #243 from Ark0N/feat/skill-cross-session-messaging
feat(skill): drive claude workers over Claude Code cross-session messaging
2026-08-09 16:57:25 +02:00
Codeman maintainer 6cc7b4328b feat(cases): clone a Git repository as a new case (#236)
Adds an Add Case -> "Clone Repo" tab plus two endpoints, implementing
@DodgyBadger's proposal in #236: clone a public repository straight into
codeman-cases/<name> and register it as a normal local case.

POST /api/cases/clone is synchronous by design (request held open, bounded
by GIT_CLONE_TIMEOUT_MS): no job store, no polling, no cancellation
surface. Success broadcasts the usual case:created event, so the case
still appears when a proxy idle-timeout kills the request mid-clone.

POST /api/cases/clone-preflight runs `git ls-remote --symref` so the UI can
say, while the user is still typing, whether the URL is cloneable without
credentials, what its default branch is, and which branches/tags exist.

Core lives in src/git-clone.ts, split into a pure half (URL parse, argv/env,
ls-remote parse, stderr classification) and a thin IO half, so every
security decision is unit-testable without spawning anything:

- `<name>::<payload>` transports are refused as a family, not by name:
  ext:: is the famous one, but any of them dispatches to git-remote-<name>
  and turns a clone into arbitrary command execution.
- A leading `-` is refused AND every spawn puts `--` before the operands.
  Either alone is one edit away from being a hole.
- argv arrays, never a shell. URLs carrying user:password@ are refused.
- gitNonInteractiveEnv() closes all four ways git can block on a prompt
  with no terminal attached (terminal prompt, askpass/GUI, ssh, GCM).
  HOME/PATH stay inherited, so a user's own credential helper or ssh agent
  keeps working; Codeman itself collects and stores nothing.
- The timeout signals the process GROUP, since clone fans out into
  git-remote-https/index-pack children that outlive a signal to the parent.
- Bounded output (redacted stderr tail, capped ls-remote stdout, 500 refs
  each) and a global 2-op pool, so N large clones cannot exhaust the host.

Repository contents beat scaffolding: an existing CLAUDE.md is kept, hooks
are merged into whatever .claude/settings.local.json the repo shipped, and
a repo that ships its own Claude settings is reported back as a warning
(those hooks run locally as soon as a session starts there). A failed clone
removes only the directory the attempt created, and refuses a pre-existing
destination outright, so it can never squat on a case name.

Not admin-gated in multi-user mode, unlike /api/cases/link: it writes only
inside the caller's own case space. Local-path/file:// sources are the
exception and stay admin-only there.

UI: live verdict under the URL field, case name filled from the parsed repo
until the user types their own, branch/tag as a datalist of the remote's
real refs, optional shallow clone, and a Brain picker (installed CLIs only)
that points the Run button at the chosen agent. Starting a session stays
opt-in. The tab hides itself when the server reports no git.

Tests: the pure half exhaustively (every refusal has a case), plus real git
against a real local bare repo for clone/ref/timeout/cleanup, and a
route-level suite with unmocked fs that clones through the endpoint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 16:33:01 +02:00
Codeman maintainer ce22c2a608 feat(path-picker): show hidden files and folders, and harden the secret blocklist
The picker behind Link Existing's "Browse" and the mobile keyboard's Path key
refused every path with a dot-prefixed segment, so `.github/workflows/ci.yml`
could not be selected and a hidden folder could not be opened at all. It gains
the same `.*` toggle as the File Viewer: default OFF, per-device, and applied to
both the listing and the preview endpoint, which re-resolves the path
independently.

That dotfile filter was quietly doing security work. The picker's roots include
Home, so with every hidden path unreachable the shared blocklist never had to
name the credentials that live in dot-directories. Lifting the filter removes
that accident, so `isSensitivePath` now covers them explicitly: SSH keys at any
depth rather than only under $HOME, GPG keyrings, AWS/GCloud/Azure/Docker/
Kubernetes credentials, npm, Yarn, git, gh, netrc, PyPI, RubyGems, Cargo and
Terraform tokens, .pgpass and .my.cnf, and the Claude and Codeman agent
credentials. `~/.codeman/` and `~/.claude/` stay attachable as trees, since the
publish skill and the review-card loop read from them; only their secret-bearing
members are named.

Everything else still applies with the toggle on: blocked trees, sensitive
files, root confinement, ownership scoping and symlink-escape checks. A hidden
entry whose realpath is a secret is dropped from the listing, and opening it is
refused.

Follows #221

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 16:16:54 +02:00
Codeman maintainer 338f0e460d feat(approvals): gate push Approve/Deny buttons on the opt-in setting too
One switch now governs the whole feature: with approvalsInboxEnabled off
(the default), sendPushNotifications strips the actions and approvalId
from permission push payloads, so the buttons no longer render at all
(pre-inbox they rendered and did nothing). The page-side action relay is
gated the same way for stale notifications sent before the toggle
flipped. Only the store and answer endpoints keep running, so enabling
the toggle surfaces anything already pending immediately.

sendPushNotifications is async now (cached settings read); all call
sites were already fire-and-forget. Covered by three new payload tests
alongside the existing hostTitle suite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 16:03:27 +02:00
Codeman maintainer 8595e84c56 fix(session): auto-accept the workspace trust dialog again
A session on a fresh directory sat on Claude's "Quick safety check: Is
this a project you created or one you trust?" dialog until a human
pressed Enter. Reproduced on a new case, then read off the wire:

  1.\x1b[C Yes,\x1b[C I\x1b[C trust\x1b[C this\x1b[C folder

tmux repaints a row by writing each word followed by a cursor-forward
escape instead of a space, and Ink colours each word separately, so
`data.includes('trust this folder')` could never match a chunk. The
spaces are not there to strip: they were never sent. The auto-accept has
been dead for every session that hit the dialog.

Match on whitespace-free, ANSI-free, lowercased text instead
(`compactScreenText`), which survives both that repaint style and the
spaced full-screen redraw.

Answering means pressing Enter into a session, so three guards bound it:

- Read the RENDERED SCREEN (capturePaneText), not the chunk. The terminal
  buffer is append-only and keeps the dialog in its tail long after it
  has been answered, so a retry driven off the buffer would type into a
  live session. Direct-PTY sessions, which have no pane, fall back to a
  short buffer tail.
- Require a trust phrase AND the dialog's own confirm affordance. One
  phrase is not enough, since an agent's transcript can quote it.
- Only look during the first 90s of the pane's life, and cap it at three
  attempts. Ink can drop a keystroke while it is still mounting the
  widget, which is the other half of why sessions got stuck, but a
  dialog that will not clear must not become an Enter loop.

Verified end to end on a fresh case: dialog answered on attempt 1, one
Enter sent in total, session went straight to the composer and answered a
prompt. Before the fix the same flow parked on the dialog indefinitely.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 16:01:48 +02:00
Codeman maintainer 696339fe12 feat(web): make a dead connection unmistakable instead of a red dot
Opening Codeman with nothing reachable (phone off the tailnet, VPN down,
server stopped) rendered a normal-looking UI: the service worker serves the
cached app shell, every /api call fails, and the only tell was an 8px red dot
in the header corner. On a phone that reads as "there are no sessions".

Two surfaces, chosen by whether there is anything worth looking at:

- Full-screen overlay while no server state has loaded this page load. It
  names the host, lists the three things to check (network, VPN/Tailscale,
  server), counts down to the next retry, and offers "Retry now" plus
  "Show cached view" to demote itself to the banner.
- Non-blocking banner once state HAS loaded, so a mid-session drop leaves the
  terminal scrollback readable.

A 2.5s grace keeps a COM deploy (SSE is back in ~200ms) from flashing the
banner every release; navigator.onLine === false skips the grace, since the
device saying "no network" is never a blip. Retry re-arms the terminal
WebSocket as well as SSE: planWsReconnect can give up outright, and the SSE
backoff caps at 30s, so waiting it out is not always an option.

The decision is pure (computeConnectionLossUi in constants.js, unit-tested in
a node VM like the WS reconnect policy); app.js only writes the DOM.
2026-08-09 15:56:44 +02:00
Codeman maintainer 6c744f8677 feat(approvals): make the inbox opt-in (default OFF) and drop em-dashes
Owner decision: every Approvals Inbox UI surface (header bell, drawer,
phone overview answer strips, reload seeding) now requires enabling
approvalsInboxEnabled in App Settings -> Panels; only an explicit true
turns it on. The store, endpoints, and push Approve/Deny actions keep
running regardless (the push buttons are already opt-in per subscription).

Also replaces em-dashes with plain punctuation across the newly authored
comments, docs, and strings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 15:55:51 +02:00
Codeman maintainer c50bb02e62 feat(file-viewer): show hidden files and folders
The tree endpoint has accepted `showHidden=true` since it was written; the
panel hardcoded `showHidden=false`, so dot-prefixed entries were unreachable
from the File Viewer and opening one meant guessing its path.

Adds a `.*` toggle to the panel header. It re-fetches instead of re-rendering
the cached tree (the filtering is server-side), preserves the expanded
directories so toggling does not collapse the tree, and persists per-device to
its own `codeman:fileBrowserShowHidden` key. That key is deliberately not part
of the app-settings object, which `saveAppSettings()` rebuilds from the
settings-modal DOM and would drop it on the next save.

Default is OFF, so an untouched install behaves exactly as before.

Closes #221

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 15:45:10 +02:00
Codeman maintainer 086ea4dd7c feat(mobile): make a working session look like one on the phone overview
The overview already had a `working` state; nothing ever reached it,
because the status it reads was wrong (see previous commit). Now that a
row can actually be in it, the state needed to look like something.

- The row gets a slow green breathing edge (2.2s). Deliberately calmer
  and slower than the red/yellow alert blinks, since working is not an
  alert and must not compete with the two states that do want you.
- The dot keeps its `pulse` and picks up a spinning ring: the same 2px
  ring with a bright leading edge that a tab shows while it loads,
  reusing the `tab-load-spin` keyframes from styles.css rather than
  re-declaring them, so the two cannot drift. Green rather than the tab's
  blue because here it means "running", not "loading": the motion is the
  shared part, the color still belongs to the state.
- The pill animates "working ...".

Reduced motion drops all three to static: a green edge, a full ring, a
static ellipsis.

Verified in headless Chromium at 390px against a live working session:
row breathe-green, dot pulse plus tab-load-spin ring, pill dots.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 15:31:16 +02:00
Codeman maintainer b03780dfd2 fix(session): decide working/idle from the pane, not the composer redraw
Every working Claude session reported `status: "idle"` about two seconds
into its turn. Measured on live workers: two sessions mid-tool-call at 13
and 17 minutes both read `idle` while their panes showed
`✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`.

Two things had drifted apart:

1. The working indicator changed. Claude animates the glyph through
   `· ✢ ✳ ∗ ✻ ✽` and randomizes the gerund per turn, so neither
   SPINNER_PATTERN (braille, no longer drawn) nor the keyword list
   (Thinking/Writing/Reading/Running) matches a turn anymore.
2. A `❯` sighting is not the end of a turn. Claude redraws the composer
   roughly once a second all the way through one, and that redraw armed
   the "2s later, call it idle" timer.

Matching the new status line in the STREAM does not fix it either: tmux
ships partial repaints, so the complete line reached the PTY about once
every 20 seconds while the `❯` arrived every second.

So the decision moves off the stream:

- An unbroken run of repaints marks a turn as started. Sampled once a
  second for 12s over six live sessions, the two working ones produced
  output in 12/12 windows and the four idle ones in 0/12. Pure helpers in
  session-activity.ts carry the thresholds.
- Idle now needs the pane to go quiet AND the screen to agree.
  `_confirmIdle()` asks tmux what is rendered (new `capturePaneText()`,
  one plain `capture-pane`, floored at 1.5s per session and only ever at
  a transition) and re-checks every 5s while the screen still shows work.
  A turn can sit silent for tens of seconds inside one tool call, so
  silence alone proves nothing.
- The same screen check vetoes keystroke echo, which is a steady stream
  of repaints too but is not work.

CLAUDE_WORKING_LINE_PATTERN matches the `… (elapsed)` shape rather than
the glyph, because the FINISHED line (`✻ Cooked for 2m 49s`) carries the
same glyph and would otherwise pin a session at working forever.

Claude mode only. An external CLI has no `❯`, so nothing would arm the
confirmation and such a session would latch busy.

respawn-patterns.hasWorkingPattern() had the same blind spot (its gerund
list cannot see "Actualizing"), so it takes the pattern as an extra
signal. That can only make respawn less eager, never more.

Idle now lands about 3 to 5 seconds after a turn ends instead of 2
seconds into one. Verified end to end against a live worker, sampled
against the CLI's own "esc to interrupt" footer as independent ground
truth: busy for all 25s of a turn, idle 3s after it ended.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 15:31:02 +02:00
Codeman maintainer ff10a50bc0 feat: Approvals Inbox, one cross-session queue for prompts waiting on a human
Permission dialogs, AskUserQuestion questions and idle prompts from every
session now land in a server-side inbox (web/approval-inbox.ts, one item per
session, claude-mode only) and are answerable in place: a header bell + drawer
on desktop, inline answer strips on the phone overview's NEEDS YOU rows, and
working push Approve/Deny buttons (previously dead ends, now answered straight
from sw.js with no tab open). Pending alerts survive reloads because the
frontend seeds from GET /api/approvals on init.

Answering sends the digit / Esc / prompt text through the existing tmux input
path; option digits are accepted only when they match options parsed from the
captured pane frame, and the answer path re-captures the pane first so a
dialog that already left the screen refuses with 409 instead of typing into
the composer. New elicitation_complete / elicitation_response hook matchers
resolve question items the moment they are answered in the terminal;
refreshStaleCodemanHooks heals existing cases.

Verified end-to-end against a live claude session: a real AskUserQuestion
dialog parsed into 5 option buttons and was answered from the drawer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 14:04:34 +02:00
Codeman maintainer 3e568511f8 style: drop em-dashes from new comments
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 13:23:52 +02:00
Codeman maintainer 64b33eb630 feat: pass --name to local claude spawns so workers carry their session names as peer names
Version-gated fail-closed at 2.1.224 (the cross-session-messaging release,
flag presence verified against that binary): an unknown or older CLI yields
a spawn command byte-identical to before, because claude aborts startup on
an unknown option and that would kill every session spawn. The value is
allowlist-sanitized ahead of the double-quoted interpolation, and only the
local command carries the flag; docker/remote builders never see it since
their CLI is not the probed binary. Verified E2E on an isolated instance:
cmdline shows --name, ListAgents lists the session name, replies arrive
tagged from-name.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 13:23:24 +02:00
Codeman maintainer 1e1db947c5 feat(skill): drive claude workers over Claude Code cross-session messaging
Claude Code v2.1.224+ gives sessions ListAgents/SendMessage and a per-session
inbox socket. Codeman's claude workers are ordinary local Claude Code sessions,
so the agent skill now teaches task delivery and result collection over
messaging where available (multi-line exactly-once messages, mid-turn steering,
latched replies), with the HTTP primitives keeping spawn, readiness,
synchronization, liveness and delete, and a bounded fallback to the HTTP
recipes whenever the feature is absent. All mechanics verified live against
claude-cli 2.1.226.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 13:06:06 +02:00
Codeman maintainer b1614e89fc chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:38:28 +02:00
Codeman maintainer 0aa16cd4d3 docs(skill): never branch on .status, it is wrong in both directions
Measured on a live claude worker: `GET /api/v1/sessions/:id` reported
`status: "idle"` while the worker was mid-turn and actively producing output, with
`lastActivityAt` equal to the moment of the call. The skill already warned that a
worker which dies inside its pane also reads `idle`, so the field is unreliable in
both directions and nothing an agent does should depend on it.

Synchronize on `stop` via send-and-wait or on an output marker. To judge from
outside, sample `terminal?tail=` twice a few seconds apart: a changing buffer is the
only cheap positive proof a worker is still working. `wait?until=exit` stays the
death check.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:29:15 +02:00
Codeman maintainer 4ed86aa0cd fix(test-vendor): private temp per run, integrity checks, reclaim dead temps
Second review round on the #241 follow-ups. Three defects in my own previous commit,
each reproduced before and after.

1. The temp path was shared between runs (`${dest}.tmp`), so two concurrent runs
   fought over it: 4 of 4 concurrent pairs had one run die. Worse than a crash, a
   sibling's cleanup landing between the esbuild and the alias append makes
   `appendFileSync` CREATE the file, so the rename publishes a bundle-less file
   containing only the alias tail, which still satisfies the content check and
   would be blessed by the cache forever. The name now carries the owning pid.
   8 concurrent pairs afterwards: no failures, no strays, aliases intact.

2. The content check only covered the bundle, so a truncated xterm.min.js with a
   fresh mtime stayed truncated. This script can no longer produce one, but
   postinstall.js writes the same directory in place, so a Ctrl+C during
   `npm install` does, and a 200-byte xterm.min.js means `Terminal` is undefined
   and every mobile test dies on a null. A copy must now match its source byte for
   byte, and a derived output must clear a floor far below the real ratios
   (measured 0.97-1.00 minified, 0.51 for the bundle) while a truncation misses by
   orders of magnitude. Verified: 200-byte and 50-byte poisonings both repaired.

3. The try block ended before the append and rename, so a rename failure leaked its
   temp behind a raw stack. It now covers both and reports which asset failed.

Per-pid names mean a killed run's temp is never reclaimed by a later rebuild, so
startup sweeps temps whose owning process is gone, and only those: `kill(pid, 0)`
throwing ESRCH. Deleting a live run's temp would recreate the collision fix 1
removes. Verified both directions, plus SIGKILL mid-build leaving no litter. The
sweep swallows its own errors, because reclaiming litter must never fail the run:
a directory named like a dead temp otherwise crashed the whole prepare step.

Security-reviewed: no shell (execFileSync with an array, `shell` unset), every
argument from the static asset table plus a numeric pid, all writes confined to the
vendor dir under strace, `process.kill` only ever with signal 0 (and pid 0 skipped,
since to kill(2) it means this process group), no new dependencies, no network, no
eval, nothing published. The emitted browser bundle is byte-identical to the one
scripts/build.mjs ships, tail included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:29:15 +02:00
Codeman maintainer a15b81db77 fix(test-vendor): repair a poisoned bundle, track all bundle inputs, pin esbuild
Follow-ups to #241 (thanks @Lint111), from an independent review of that PR. The
script is a real fix for a real gap; these are the four defects the review found,
each reproduced before and after.

1. A wrong-but-fresh output was never repaired. The zerolag bundle is finished by a
   SECOND step (the alias append), so anything landing between esbuild and the
   append is permanent: the file looks complete, carries a current mtime, and the
   mtime-only cache reports "up to date" forever while the suite dies on
   `LocalEchoOverlay is not defined`. Reproduced by replaying #241's own two
   commits: running the first and then pulling the second kept the broken bundle.
   Fixed twice over, because the two halves address different cases. Builds now go
   to a temp file and `renameSync` into place, so this script can never publish a
   half-written output (that also covers an interrupted esbuild or copy, and two
   concurrent runs). And `isFresh` verifies the bundle actually contains its alias
   tail, which is what repairs a file an EARLIER version already poisoned; a rename
   alone cannot fix what is already on disk.

2. Freshness compared against the entry file only, but esbuild bundles its four
   siblings too, so editing overlay-renderer.ts left the suite testing a stale
   overlay while reporting "up to date". Editing those siblings is exactly the
   single-source workflow CLAUDE.md mandates. It now stats every `.ts` in the
   package source dir. A full rebuild is ~2s, so the cache was not buying much.

3. `execFileSync('npx', ...)` passed no cwd, unlike scripts/build.mjs, so a run from
   another directory missed the repo's pinned esbuild and would fetch an unpinned
   one from the registry. Both calls now pass `cwd: ROOT`.

4. Every invocation in test/mobile/README.md was a bare `npx vitest`, which skips
   the `pretest:mobile` hook npm only fires for `npm run test:mobile`, so the
   documented commands all bypassed the fix. Rewritten, with a note on why.

Also: an esbuild failure printed a raw stack; it now names the asset and its input,
matching the missing-input message. And the header comment no longer implies the
vendor dir is always empty: scripts/postinstall.js already writes these same seven
outputs, so what this script adds is freshness and independence from install time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:13:18 +02:00
Codeman maintainer 341c7ccc59 test: cover the skill CLI, the injection call site, and endpoints.md drift
Three gaps found while auditing the agent skill.

`codeman skill install` / `uninstall` had no tests at all, including the linked-case
resolution that shipped in 1.14.2 with nothing guarding it. Covered now: global target
resolution, `--case` resolving through linked-cases.json, `--case` falling back to the
cases dir for an unlinked name, a missing or malformed registry degrading to the
fallback instead of throwing, and a nonexistent case being rejected. `resolveSkillTarget`
called `process.exit(1)` for a missing case, which would have killed the test runner, so
the pure resolution is split out and exported; CLI behavior is unchanged.

The `POST /api/sessions` injection call site was never exercised, because the shared
route mock hardcoded the gate off. The mock's gate is overridable per test now (default
still off, since other tests rely on that), and there is coverage that the path injects
when the setting is on, does not when it is off, and is claude-mode gated.

Nothing guarded skills/codeman/reference/endpoints.md against drifting from the routes
it documents, which is how it drifted in the first place. A static guard parses the
endpoints out of the markdown and asserts each is really registered, tolerating the
/api/v1 alias and path params.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:09:51 +02:00
Codeman maintainer c1719e04e5 docs: fix the zh-CN agent recipe, the \r gotcha, and the agent-control plan status
README.zh-CN.md taught a recipe that cannot work: its programmatic-input example had no
trailing `\r`, so Enter was never sent and the prompt sat unsubmitted forever, and its
read step used `/output`, whose `textOutput` is always empty for interactive tmux-backed
sessions. A reader following the Chinese README walked into both of the silent failures
the English one warns about. Its agent/automation section is now brought in line with
README.md: the `\r` rule and every example that needs it, and the correct read path.

CLAUDE.md's "Single-line prompts only" gotcha described the newline restriction but
never mentioned that input must end with `\r` or Enter is never sent, which is the most
common silent failure when driving the API.

docs/agent-control-plan.md asserted as still-open several things that shipped in 1.14.1
and 1.14.2 (the wait endpoints, the packaged skill, the install CLI, agentSkillEnabled).
The status header and the stale bullets now match reality; the historical design content
is untouched, since the document is a record.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:09:51 +02:00
Codeman maintainer d33f3803a1 fix(agent-skill): write the skill atomically and stop swallowing refusals
Two ways the injection could go wrong quietly.

`installAgentSkillInto()` wrote each file with a bare `writeFile`, no lock and no
temp+rename, while every sibling mutator in hooks-config.ts goes through
`withSettingsLock`. Two Claude sessions created concurrently in one repo both wrote the
same ~16KB SKILL.md, and any reader loading it mid-write could observe a truncated
file. Writes now go through a temp+rename helper under the same lock the neighbours
use, so a reader sees either the old file or the new one.

Both server call sites discarded the outcome with `.catch(() => {})`, so the two
refusal results were invisible: `foreign` (a user-authored skills/codeman is present,
so we declined to touch it) and `symlink` (the skill dir or its parent is a symlink, so
we declined to write through it). Turning `agentSkillEnabled` on, seeing nothing appear
and having no way to find out why was the reportable-as-a-bug outcome. Refusals are now
logged with the path and what to do about it. The boring outcomes stay silent, since
they happen on every session create. Injection remains best-effort: a refusal or a
thrown error still cannot fail session creation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:09:51 +02:00
Codeman maintainer 477e73039c fix(skill): match shift+tab for readiness, portable ANSI strip, endpoint gaps
The readiness gate matched `bypass`, which is the status bar of ONE permission mode.
`buildPermissionArgs()` also spawns `--permission-mode auto`, `--allowedTools` and
plain `normal`, and the mode is not exposed on `GET /api/v1/sessions/:id`, so an agent
cannot know which token to expect. A non-default worker was therefore reported broken
after burning the whole ladder.

Measured one pane per mode against claude-cli 2.1.226:

  --dangerously-skip-permissions  ->  "bypass permissions on"
  --permission-mode auto          ->  "auto mode on"
  --allowedTools Read,Grep        ->  "don't ask on"
  (none, normal)                  ->  "don't ask on"
  --permission-mode plan          ->  "plan mode on"

Every one ends `(shift+tab to cycle)`, so `shift+tab` is the single space-free token
that means "the composer is up" in every mode, and it is what the ladder matches now.
Verified live end to end on a virgin case: stage 1 misses while the trust dialog is up,
stage 2 accepts it, stage 3 matches in 623ms.

⚠️ `shift+tab` contains a `+`, so it only works through `--data-urlencode`. In a
hand-built query the `+` decodes to a space and the server searches for `shift tab`,
which never appears; the response echoes `match: "shift tab"`, which is how to spot it.
Measured both ways. The stage-4 fallback (make the worker echo a split token, proving
readiness by answering rather than by chrome) stays as the last resort, and is now also
verified live: it matched in 2.5s, with the token surviving the space-less TUI intact.

Also portable ANSI stripping: the read pipelines used `sed 's/\x1b...'`, and BSD sed
(the macOS default) has no `\xHH` escape, so on macOS the strip silently removed
nothing and handed the agent raw ANSI. They now build a real ESC with `printf`.

And endpoints.md gaps: the `FORBIDDEN` 403 row and which auth responses are plain text
rather than the JSON envelope, the input size cap, the undocumented `killMux` parameter
on DELETE, and the fact that zero/negative/non-integer timeouts are rejected with a 400
rather than clamped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 12:09:51 +02:00
Ark0N b374032699 Merge pull request #241 from Lint111/fix/mobile-test-vendor
test(mobile): serve the xterm vendor bundles the browser suite needs
2026-08-09 12:09:36 +02:00
Ark0N b6efdfccf4 Merge pull request #240 from Ark0N/feat/predictive-echo-codex
Zero-lag predictive echo for Codex sessions (mosh-style write-through)
2026-08-09 11:42:53 +02:00
Codeman maintainer b191f3c2c6 test(predictive-echo): real-auth streaming fixture pins baseY growth
With a real codex login now available, record the one shape the fake-key
lab could never produce: a genuine model reply streaming above the pinned
composer, pushing lines into history (baseY grows) while keystrokes land
mid-stream. The recorder gains an opt-in CODEX_RECORD_REAL=1 scenario
using the user's own ~/.codex (fixture secret-scanned for key/JWT
material before writing; scanned clean). The replay test pins: baseY > 0,
mid-stream predictions painted, exact convergence to the typed text.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 11:35:08 +02:00
lior 9fd856a918 fix(test-vendor): append the zerolag global aliases
The zerolag bundle exports only `XtermZerolagInput`, but app.js constructs
`new LocalEchoOverlay(terminal)` directly. scripts/build.mjs appends global
aliases after esbuild (build.mjs:53-66); the first version of this script
omitted that step.

Without them initTerminal() throws `LocalEchoOverlay is not defined` at the
line that builds the overlay — and because that is midway through the function,
EVERY later step silently never runs, including the mobile touch handlers on
#terminalContainer. The page still had a terminal, so the failure looked like a
tap-routing bug rather than a boot error.

Verified: boot errors none, and all four terminalContainer touch listeners
(touchstart/touchmove/touchend/touchcancel) now register.
2026-08-09 10:27:25 +03:00
lior be449e6e9e test(mobile): serve the xterm vendor bundles the browser suite needs
The mobile suite drives a real browser against a WebServer started from
TypeScript source, so fastify-static serves join(__dirname, 'public') =
src/web/public — not dist/web/public, where `npm run build` puts the vendor
bundles. Every /vendor/xterm* request 404s, so `Terminal` is never defined,
initTerminal() never runs, and any test touching app.terminal dies with
"Cannot read properties of null".

Measured in one worktree, toggling only the vendor files:

  before: 404s=5  Terminal=undefined  app.terminal=null   8 failed | 26 passed
  after:  404s=0  Terminal=function   app.terminal=live   6 failed | 28 passed

The 6 remaining failures are genuine pre-existing bugs (stale layout and
accessory-bar expectations, a CJK timeout) and are left alone here.

This went unnoticed because config/vitest.ci.config.ts excludes test/mobile/**,
so CI never ran the suite. `npm run test:mobile` now runs it, with a pretest
hook that builds the bundles.

The asset list was derived from the actual 404s rather than from build.mjs —
which is how xterm-addon-unicode11 and xterm-zerolag-input got included; reading
the build file alone would have missed both. Outputs go to the gitignored
src/web/public/vendor/, so they stay build artifacts. The script is idempotent
(skips outputs newer than their source) and does not touch the normal build.

Full CI suite unchanged: 4368 passed.
2026-08-09 10:00:18 +03:00
Codeman maintainer 04de943b7f chore(predictive-echo): changeset notes cover the anchor-hold review fix
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 05:30:07 +02:00
Codeman maintainer 9e7c537e14 fix(predictive-echo): anchor hold after unpredicted wire edits (review findings)
Independent post-build review found three gaps, all one family: input that
changes the composer without a prediction leaves the DISPLAYED cursor stale
for one RTT, and anchoring a new run on it painted ghosts one cell off
(blank-neutral, so they lived out the full TTL: "tehh" on
backspace-then-retype, exactly on the links the feature targets).

Fix: the addon now HOLDS new predictions after any such edit (backspace with
nothing outstanding = deleting echoed text, clearPredictions, and now also
IME/plain-paste 'text' commits, which the hook clears like 'clear') until
the next PARSED write releases the hold. The inline predictChar reconcile
deliberately does not count: only the emitter pass or the public
reconcile() is the display-caught-up contract. Worst case is exactly one
unpredicted keystroke, whose own echo releases the hold. Also patched the
one bypass path the PR had missed: _handleCjkInput now clears predictions
like insertTerminalText and the other bypass sends.

Package suite 230, vm gating 85, E2E 10/10 all green after the change.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 05:29:45 +02:00
Codeman maintainer 55bff4a4bf docs+ci(predictive-echo): CI package-suite step, invariants, changeset
ci.yml runs the xterm-zerolag-input suite (Layers 1-3) after the root
npm ci (workspaces hoisting; no separate install). CLAUDE.md and
architecture-invariants.md rewrite the codex echo story: predictive
write-through with the wire-neutrality, separate-bundle, composer-gate,
baseY and blank-neutral invariants spelled out; the single-source section
now covers both vendor bundles and why their entry points differ.
Changeset: minor for aicodeman + xterm-zerolag-input.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 05:11:18 +02:00
Codeman maintainer fa02bd4503 test(predictive-echo): E2E suite against a real codex TUI (Layer 5)
Out-of-process lab server (VITEST markers stripped so tmux/codex are real),
CODEMAN_INSTANCE=codexlab on port 3222, throwaway CODEX_HOME with a fake
key. Ten scenarios: bundle smoke, predict+converge typing, the #218 arrow
retest (submitted text exact), the #222 live picker, the #219 paste order,
the #220 wrap, the trust-modal ghost eliminator, the localEchoEnabled kill
switch, the end-to-end byte-identity trace (predictor active vs null), and
a display-delayed 300ms-RTT run pinning instant spans with exact pixel
geometry plus arrow-edit correctness under lag.

Live-TUI hardening learned the hard way: codex Ctrl+U kills only to line
start (End first), a fake-key submit leaves a Reconnecting loop that can
kill codex seconds later (retry-cancel + composer stability probe; the
submitting scenario runs after all composer-state ones), and typing must
wait for the predictWhen gate itself, not merely a rendered composer.
CI-excluded like the other Playwright suites; skips cleanly when codex is
not installed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 05:02:14 +02:00
Codeman maintainer 5bde897752 feat(predictive-echo): Codeman integration + Layer 4 vm tests
terminal-ui.js: _localEchoPolicy ('buffer'|'predict'|'off') computed at the
end of _updateLocalEchoState with _localEchoEnabled keeping its exact 1.12.2
values; _predictHookOnData called as a plain statement between the buffer
block and Normal Mode (visual-only, try/catch, never returns, never touches
_pendingInput); classifyPredictInput + isCodexComposerRow (baseY-based,
measured /^> /-signature gate) on CodemanTerminalInput; construction beside
the LocalEchoOverlay from the separate bundle with graceful absence;
insertTerminalText/clearTerminalInput/setFontSize/applyTerminalSkin clear or
refresh predictions. app.js: fields + tab-switch and SSE-reconnect clears.
voice-input '\r' branch and keyboard-accessory sendKey clear predictions
(both bypass onData). sendEnterKey needs no change: codex falls through to
the immediate-flush branch.

Layer 4 vm tests: classify truth table (20 cases), composer-row gate incl.
the baseY pin, policy matrix with the 1.12.2 invariants untouched, wire
neutrality + throwing-predictor pins. Stale mobile keyboard codex-buffering
tests repointed at claude; new codex twin asserts write-through streaming,
prediction spans and TTL self-heal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 04:33:09 +02:00
Codeman maintainer 6c55ce3f8d feat(predictive-echo): second vendor bundle wiring
postinstall + build.mjs build vendor/xterm-predictive-echo.js as a SEPARATE
IIFE (window.PredictiveEchoAddon + self-activating PredictiveEchoOverlay);
the zerolag bundle command is untouched and its output verified
sha256-identical. index.html loads it after the zerolag tag (cacheBustAssets
covers it), sw.js precaches it, build.mjs HASHABLE content-hashes it.
A missing or broken bundle degrades codex to plain PTY echo.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 04:26:39 +02:00
Codeman maintainer a30524060a feat(predictive-echo): PredictiveEchoAddon + Layers 1-3 test suites (0.2.0)
Mosh-style write-through prediction: the consumer sends every keystroke
unchanged; the addon paints predicted glyphs and reconciles against the
parsed buffer. Confirm = cell match + cursor advance (placeholder-safe,
repaint-safe); two-pass mismatch cascade with neutral blanks (measured:
codex clears its placeholder on first echo); TTL bound; baseY-based line
reads; scroll/resize/off-row clears. Zero edits to zerolag-input-addon.ts.

Tests: 30 addon-law specs + renderer geometry (fake performance clock for
TTL/grace), 6 replay suites running the real algorithm through a real
@xterm/headless parser fed by the recorded codex fixtures, and a
500-iteration seeded fuzz with per-op span/record + grid invariants.
227 total, the pre-existing 175 untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 04:25:24 +02:00
Codeman maintainer 00fb3b0908 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:18:23 +02:00
Codeman maintainer 5aa59c70cc feat(predictive-echo): Phase 0 codex fixtures, measurements, package scaffolding
Recorder (scripts/dev/record-codex-frames.mjs) captures real codex 0.147
TUI output through the production pipeline (tmux status-off + the codex-mode
full strip from session.ts) into JSONL fixtures with keystroke injection
points; analyzer replays them through @xterm/headless for the measurements
in docs/predictive-echo-plan.md. Composer signature /^> /-style (U+203A),
modal and wrapped rows correctly rejected, echo is unstyled default-fg,
tmux delivers echo as minimal in-place deltas.

Package: types.ts gains optional cursorX/cursorY, getCell, onWriteParsed,
onResize (all additive); prediction-renderer.ts renders per-glyph spans
keyed by prediction seq; @xterm/headless@^6.0.0 devDep for replay tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 04:10:18 +02:00
Codeman maintainer ffccde4f7d fix(cli): codeman status probes the running server (#230)
Reported by @mtiller.

`codeman status` runs in its own fresh process, and reported THAT process's
always-stopped Ralph loop under a bare "Status:", which reads as "the web server
is down" while the service is running fine and agents are reachable. It now probes
the real server first (`CODEMAN_API_URL`, else https then http on the local port,
overridable with `--url`) and reports reachability, version and live session
state. Any HTTP answer proves the server is up, including a 401 from a
password-protected install. The Ralph loop keeps its own `codeman ralph status`.

This complements `codeman web --status` from the daemon work: that answers "did I
start a daemon", this answers "is a server running at all", which is what the bare
command was already being used for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:24 +02:00
Codeman maintainer bec3da3d31 fix(ui): a described session tab shows just the description (#232)
Reported by @mtiller.

A session named `w2-foo-bar: some description` rendered both halves on the tab, so
the generated id ate the width that the part the user actually chose needed. The
tab now shows the description alone and the `w<n>-<case>` id moves to the tooltip,
where it stays available without being read every time. It is still shown in the
session settings modal. Undescribed tabs are unchanged.

`aria-label` deliberately keeps the FULL name, so screen readers still get the id.

Also fixes a re-render loop this exposed: the incremental update compared
`nameEl.textContent` against the full name, which for a described tab never
matched, so those tabs re-rendered on every pass. The compare now targets the
display label.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:24 +02:00
Codeman maintainer 6e89eb9ec1 fix(web-tabs): bound time-to-headers, not the whole proxied exchange (#237, #238)
Reported by @DodgyBadger.

#237: the proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which
bounded the ENTIRE exchange rather than the wait for response headers. A dashboard
endpoint doing model inference, and any actively streaming response, both died at
30s as a generic 502 that Codeman never logged, so it read as an intermittent
network error. The timeout now bounds time-to-headers only and is cleared the
moment headers arrive, so a slow endpoint and a long stream both survive. The
default moves to 300s because "the app is thinking" is normal for the dashboards
people proxy; abandoned upstreams are reclaimed by the client-hangup abort rather
than by this value.

A browser that navigates away mid-request now aborts the upstream fetch, guarded
by `writableFinished` for the same reason as `abortOnClientHangUp` in
session-routes: `close` also fires after a completed response and must not abort
anything. Header timeouts are logged as a warning with a sanitized identity
(method plus origin plus path, never the query string, which can carry the
dashboard's tokens), and a client hangup is deliberately not warned since nobody
is listening and it would read as the dashboard being broken.

The WebSocket handshake keeps its own 30s budget
(`CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`), decoupled from the request timeout:
a handshake is connection establishment, and waiting minutes on one only delays
the browser's reconnect logic.

#238: the web-tab guide covered sandboxed dashboards having no cookies, but not
cookie authentication in front of Codeman itself (Cloudflare Access and similar),
where a sandboxed frame's asset and API requests carry no auth cookie, bounce to
the login provider, and leave the embedded app looking unstyled or broken while
trusted mode works. Documented, and the Test button's result now says it probes
server-to-upstream reachability only, not how the page behaves in a sandboxed
frame.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 04:06:23 +02:00
Codeman maintainer 94aa53c65b chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:38:40 +02:00
Codeman maintainer e88b971bb7 feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a
repo-only reference, and fix six defects found while verifying it live.

Install layer:
- `codeman skill install [--case <name>]` / `codeman skill uninstall`.
  Case names resolve through linked-cases.json first, mirroring the
  server's resolveCasePath(), so a case linked in from outside
  ~/codeman-cases no longer fails with "Case not found".
- applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in
  hooks-config.ts. Copies are marker-owned, so an unmarked user-authored
  skill is never touched, and a symlinked skill dir is refused (this
  repo's own .claude/skills/codeman is a symlink to the source).
- Synced `agentSkillEnabled` setting, default OFF: schemas.ts,
  ports/config-port.ts, server.ts, session-routes.ts (add-only injection
  on Claude session create and quick-start), plus the App Settings toggle.

Skill content fixes, each reproduced before and after:
- Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`.
  Shell state does not survive between agent tool calls, and an undefined
  is_self exited 127, firing the `||` branch and deleting the caller's own
  session with the one guard bypassed. The request now lives inside the
  guard, so a lost preamble deletes nothing.
- clientId is a fixed literal instead of `agent-$$`. The pid changes per
  tool call, so the documented resend-identical-request loop stopped being
  a duplicate and retyped the prompt, submitting the turn twice.
- `last-response` is now the documented read path for claude and codex
  workers. It returns clean transcript text; the terminal scrape it
  replaces returns a wall of TUI repaint noise. Its transcript flush lags
  the stop signal, so the recipes poll it rather than reading once.
- quick-start examples branch on `.success`. Previously a failed spawn
  yielded the literal session id "null" and burned the whole readiness
  budget before reporting jq noise instead of the cause.
- Documented that turning `agentSkillEnabled` off sweeps nothing, and
  corrected the hooks-config comment that claimed a toggle-off sweep
  exists. Per-case cleanup is `codeman skill uninstall --case <name>`.
- Documented that SESSION_BUSY means the 50-session cap on quick-start,
  and that caseName resolves linked cases, so a generic name can land a
  worker in a real repo.

Tests: test/agent-skill.test.ts covers install, refresh, idempotence,
marker ownership and symlink refusal against the real packaged source;
test/quick-start.test.ts covers injection behind the setting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:30:15 +02:00
Codeman maintainer 8406c497e2 fix(terminal): stop forwarding the wheel to codex, it ignores SGR reports
DodgyBadger reported a completely dead wheel in codex tabs (#227 comment)
while the scrollbar drag worked, and the [scroll] line confirmed the
branch: forward-sgr with 967 rows of healthy local scrollback unused.

Measured against codex-cli 0.147.0 in a bare tmux: codex never enables
mouse tracking (mouse_any_flag=0), runs an inline viewport
(alternate_on=0) and pushes its transcript into the terminal's own
scrollback (history_size grows), and SGR wheel reports written to its
pane change nothing at all. Hand-encoded SGR taps are no-ops too, so
they stay (harmless), which means click-to-position is merely
unavailable there rather than damaging.

_shouldForwardWheelToApp now returns true for claude >= 2.1.187 and
nothing else; codex falls to the local-scrollback path like
shell/gemini/opencode, which is the same history the scrollbar drag was
already reaching. The claude-only PageUp fallback is untouched.

Verified in Chromium against a live codex session on an isolated
instance: routing logs local-scrollback, the viewport moves 39 -> 4 and
zero bytes go to the PTY.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 03:22:06 +02:00
Codeman maintainer 40b4aba043 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:35:59 +02:00
Codeman maintainer 4b44988bfc test: give daemon-control tests a unique port (3212 was already taken)
test/sse-subscription-filter.test.ts already binds 3212; sequential test
execution hid the clash. Moves the probeServer fixture to 3216 (3217 for
the nothing-listening case) per the unique-port convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:22:01 +02:00
Codeman maintainer 316d0a4c82 Merge pull request #233 from Lint111/feat/hooks-config
Conflict in refreshStaleCodemanHooks resolved by keeping every staleness
trigger: the master-side TLS-flagless curl check (hooks without -k) AND the
PR-side current-wake-marker (V3) + SubagentStop guard marker checks.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-09 02:21:52 +02:00
Ark0N 1184720648 Merge pull request #239 from Ark0N/feat/daemon-mode
feat(cli): codeman web -d and codeman service install (#231)
2026-08-09 02:19:50 +02:00
Ark0N b067aad9b6 Merge pull request #235 from Lint111/feat/deferred-terminal-flush
fix(terminal): drain deferred output without a wake event
2026-08-09 02:19:32 +02:00
Ark0N 19a3d7c773 Merge pull request #234 from Lint111/feat/ai-checker-stderr
fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
2026-08-09 02:19:10 +02:00
Codeman maintainer 085f4acb60 feat(cli): codeman web -d and codeman service install (#231)
Two ways to keep the server running, split by how long it should last.

`codeman web -d` relaunches the same entry script detached (setsid), with
`--stop` and `--status` alongside it. A pidfile and log live in the data
dir. `nohup` is not what makes this work: Node re-arms SIGHUP to its
default disposition even when it inherits "ignore", and cli.ts handles
SIGHUP with a graceful shutdown, so a delivered HUP still stops the
server. Removing the shell's ability to send one is the fix.

`codeman service install|uninstall|status` writes and loads the systemd
user unit or the LaunchAgent, with the installing shell's PATH baked in
(launchd hands a job /usr/bin:/bin:/usr/sbin:/sbin, which finds neither a
Homebrew/nvm node nor tmux/claude). install.sh already covers one-liner
installs; this is for npm globals.

Both refuse to start when a server is already up on the data dir, since a
second instance on the shared tmux socket attaches PTYs to the first
one's live sessions. Both poll /api/status until the child answers or
dies rather than reporting a success they have not seen. `--stop` checks
the pid still looks like a Codeman server before signalling it.

The systemd unit name and launchd label move to config/service-names.ts
so install.sh, detectSupervisor() and service install cannot drift into
supervising two copies. Instance-scoped, unchanged for the default
instance.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:34:55 +02:00
Codeman maintainer d26f26fe34 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 01:18:03 +02:00
lior 091df2b6d8 fix(terminal): drain deferred output without a wake event 2026-08-08 23:00:36 +03:00
lior 5f775b1ab1 fix(hooks): guard subagent stops and rewake from the parent transcript
Two defects in the background-task hook scripts.

SubagentStop had no handler at all. When a subagent launched background work and
one watcher ended while others were still running, Claude could publish the
worker's last progress sentence as its final result, abandoning the live tasks.
A new guard pairs launched task IDs against completed ones and confirms liveness
by scanning /proc/<pid>/fd for an open tasks/<id>.output handle, blocking the
stop only while genuinely-live work remains. It fails open — allowing the stop —
when /proc is unavailable, nothing was launched, or everything finished.

The rewake helper watched only input.transcript_path. A subagent has its own
transcript, but Claude writes the completion queue-operation to the PARENT
transcript, so the record it waited for never appeared and the wake never fired.
It now watches both paths, but only when the relationship is provable: the
transcript's parent directory is subagents/ and its grandparent basename equals
input.session_id. It also now requires operation === 'enqueue'.

The rewake marker moves V2 -> V3; refreshStaleCodemanHooks treats absence of the
current marker as stale, so existing cases self-heal on next launch (the same
mechanism as the V1 -> V2 bump). Ownership matches on marker PREFIXES, so a
future bump still recognises older Codeman handlers and never adopts a user's.

12 tests fail on unmodified master, e.g.
  expected '[{"matcher":"Bash",…' to contain 'CODEMAN_BACKGROUND_REWAKE_V3'
  expected 'Background command bg-report-1 comple…' to contain '<codeman-background-result>'
2026-08-08 22:31:38 +03:00
lior da51193264 fix(ai-checker): keep CLI stderr out of the verdict and surface it on failure
AiCheckerBase spawned the check with `> out 2>&1`, so anything the Claude CLI
wrote to stderr landed inside the same file the verdict parser reads. A CLI that
failed to start (corrupt settings, missing auth) produced either an empty verdict
or an unparseable one, and the actual cause was destroyed on the way through —
the user saw only "Empty output from AI idle check".

stderr now goes to its own temp file. When output is empty or the verdict cannot
be parsed, the first 200 characters of stderr are appended to the error message.
The file is cleaned up alongside the existing temp files, including on the error
paths.

Two tests, both failing on master:
  expected 'export PATH="…' to contain ' 2> "'
  expected 'Empty output from AI idle check' to contain 'Claude CLI failed to load settings'
2026-08-08 22:30:39 +03:00
Codeman maintainer fa18eeef35 feat: tab action icons on the active tab only, middle-click closes tabs
Rework of the previous hover-overlay approach after feedback: sliding the
title under incoming icons made names hard to read, and icons appearing
under the cursor caused accidental gear/close clicks while switching tabs.

Now the gear/pop-out/close icons expand in flow on the ACTIVE tab only.
Selection is a deliberate click, so the strip's geometry never changes
while the pointer is aiming at a tab; hovering a background tab changes
nothing (the full title stays readable) and a stray click can only switch
sessions. Middle-click closes any tab (session tabs via the existing
close-confirm modal, web tabs via closeWebviewTab), matching browser
muscle memory so background tabs still close in one action.

The pop-out button stays opt-in via App Settings -> Tab Bar (per-device
showTabDetachButton, default off), and a detached tab keeps its icon as
the re-focus affordance. Phone layouts already used the active-only
pattern; tablets keep their always-visible touch fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer a9f26bd03a feat: fixed-width tab hover with sliding title, pop-out button now opt-in
Hovering a session tab no longer grows it. The three per-tab icons now
live in a .tab-actions wrapper that overlays the tab's right edge on
hover-capable devices: the icons slide in while the title (and any
badges) slide left by a per-tab --tab-slide distance computed in
_applyTabHoverSlide(), clipped at the left edge of .tab-info so the
readable tail (the :comment suffix) stays visible. Keyboard focus
reveals the overlay via :has(:focus-visible), so a mouse click on the
gear does not pin it open. Touch devices keep the previous in-flow
behavior (the wrapper adds no width in flow, and the legacy tap-reveal
rules are preserved under @media (hover: none)).

The open-in-a-new-window (pop-out) button is now hidden by default and
opt-in via App Settings -> Tab Bar -> "Pop-out Button on Tabs"
(showTabDetachButton, per-device, absent from SettingsUpdateSchema like
the other display keys). A tab whose session is already detached keeps
its icon as the re-focus affordance regardless of the setting.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:58:10 +02:00
Codeman maintainer 8dc8b164a7 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 13:48:12 +02:00
Ark0N 2524759655 Merge pull request #229 from Lint111/feat/keyboard-viewport-settle
fix(mobile): coalesce keyboard viewport settling
2026-08-08 12:51:04 +02:00
Codeman maintainer 1f164bc8d2 fix(mobile): only arm the viewport settle on a real keyboard transition
A visualViewport resize event without a pending show/hide transition now
only pushes a pending settle back (_deferViewportSettle) instead of arming
fit + PTY-resize work of its own. Keyboard detection can miss a
fine-grained OS animation entirely (each step under 150px, with the
baseline chasing the animation down), while MobileDetection's own listener
still shrinks --app-height, so the per-event settle fitted xterm against a
mid-animation container with no keyboard CSS compensation and resized the
PTY to transient dims. The resulting SIGWINCH thrash (58 -> 10 -> 50 rows)
duplicated prompts and left tmux dot filler in the transcript on keyboard
close. Reproduced with a faked visualViewport driving the real handler;
master is unaffected because it never resized the PTY from this path.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 12:06:08 +02:00
lior 0a1439b1e9 test(mobile): make the coalescing test actually exercise the settle path
The suite never selects a session, so initTerminal() does not run and both
`app.terminal` and `app.fitAddon` are null at rest. `_scheduleViewportSettle`
returns early on a falsy terminal, so the coalescing assertions could not
reach the behavior they claimed to cover -- the test errored on
`Cannot read properties of null` rather than measuring anything.

Installs the minimum surface the settle callback touches and restores it
afterwards, so the coalescing path executes for real.

Adds a behavioral counterpart driven through the PUBLIC entry point
(`onKeyboardShow`) instead of the internal scheduler: three viewport steps
in quick succession must produce exactly ONE refit. On master that returns
3 (each show arms its own uncoalesced 150ms timeout), so this fails by
COUNT rather than by a missing method -- which is the failure mode that
actually demonstrates the bug.

Verified: `expected 3 to be 1` on unmodified master; passes here. The
remaining 8 failures in this file are pre-existing on master and unrelated
(same null-initialization limitation of the headless harness).
2026-08-08 08:41:04 +03:00
lior 66abe6c70a fix(mobile): coalesce keyboard viewport settling 2026-08-08 08:13:58 +03:00
Codeman maintainer fa1700da5b chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:38:35 +02:00
Ark0N aed1e59ee3 Merge pull request #227 from Ark0N/fix/scrollback-205-round2
fix(terminal): scrollback round 2 for #205 (re-pull downgrade guard, PageUp fallback, CLI version probe retry)
2026-08-08 01:36:03 +02:00
Ark0N 7f6d18b398 Merge pull request #226 from christianhaberl/fix/input-loss-on-failed-delivery
fix(api,ws): an input whose delivery fails can be retried instead of being lost
2026-08-08 01:31:10 +02:00
Ark0N 52571c7fd4 Merge pull request #225 from christianhaberl/fix/bound-the-process-tree-walk
fix(mux): bound the process-tree walk — unbounded pgrep recursion can take a machine down
2026-08-08 01:31:00 +02:00
Ark0N cb95a8562c Merge pull request #224 from christianhaberl/fix/raw-writehead-drops-security-headers
fix(http): raw writeHead routes drop every header the security hook set
2026-08-08 01:30:47 +02:00
Codeman maintainer 3cb7e30636 fix(ui): scope the wheel-opt-out tooltip's paging fallback to Claude
The reworded tooltip promised the PageUp/PageDown fallback for Claude and
Codex alike, but _localScrollbackIsHollow() gates it to claude mode only
(codex page-key handling is unverified, as the routing tests note). A codex
user reading the old text would flip the setting expecting a rescue and get
a dead wheel instead. Say plainly that Codex has no fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-08 01:29:48 +02:00
Claudia 0afd4e1cdc test: generic project names in the verification fixture
The synthetic session names end up in the harness screenshots, so shipping one
contributor's project list into everyone else's review reads oddly. The mix of
CLI modes is what the fixture actually needs — each renders a different badge —
and that is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 17:41:07 +02:00
Claudia b6293959d2 fix(test): drop hardcoded personal paths from the verification script
The script carried two absolute paths from the machine it was written on: a full
scratchpad path including a session UUID, and /home/chaberl/projects as the
synthetic sessions' working directory. This branch is pushed to a public fork, so
they were visible to anyone.

Screenshot output now defaults to tmpdir() and is overridable via
SIDEBAR_SHOTS_DIR; the synthetic working directories are tmpdir()-based too, which
also makes the harness run for anyone who checks the branch out.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:42:21 +02:00
Claudia[bot] c01edcbbb8 feat(web): optional collapsible left session sidebar
The header tab strip stops working past roughly a dozen sessions: it wraps
into two or three rows, eats vertical space and still cannot be scanned.
This adds a vertical session list in a left <aside> as an ALTERNATIVE
layout — a filter box, a live count, and a 44px collapsed rail that keeps
the ambient signal (status dot, task badge) visible.

The strip is not removed. Settings -> Display -> Tab Bar -> Session List
Layout switches between them and the default stays 'header', so existing
users see no change until they opt in.

Structure: one #sessionTabs element, two mount points. applySessionListLayout()
re-parents the SAME node between #sessionTabsHost and #sessionSidebarList,
which is why there is no second renderer and no duplicated wiring — app.$()
caches getElementById results and never invalidates them, so a moved node
keeps every existing consumer (settings-ui, webview-tabs, the generated
gesture bundle, the mobile tests) working untouched.

Notable integration points:
- Below 1024px the sidebar is an off-canvas drawer overlaying the terminal;
  closed it gets inert + aria-hidden so it cannot be tabbed into, and touch
  swipes over it no longer switch sessions.
- Subagent and ultracode windows anchor to the right edge of a sidebar row
  instead of its bottom, connector curves follow.
- Alt+B toggles; the chord is gated out of the PTY so xterm cannot also
  write ESC b into a live session.
- Collapse state lives in its own localStorage key (the settings blob is
  rebuilt from DOM controls on every save) and falls back to in-memory
  intent where storage throws.

Verified: frontend syntax + public asset checks, tsc, eslint, 26 new jsdom
tests, and a headless-Chromium harness (scripts/verify-session-sidebar.mts)
that renders a synthetic 25-session fleet in both layouts at 1600/1000/393px
and asserts mount point, widths, inert/aria state and row count.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:42:21 +02:00
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00
Claudia 9d27cc0bab docs: merge the stacked doc comments the previous commits left behind
Cosmetic, but the kind that quietly costs: JSDoc tooling attaches only the
nearest block, so a stacked second block silently hides the first.

- write() had two: the original description with @param and @example, then a
  @returns-only block added on top, which dropped the params and examples from
  hover. Merged into one. The @returns wording is also honest now — write() still
  discards the data without a PTY; what changed is that it says so.
- forgetInputSeq had been inserted BETWEEN shouldApplyInput's detailed doc comment
  and its declaration, leaving that function undocumented on hover and the doc
  attached to the wrong thing. Moved below.
- The mock kept an orphaned one-line comment above failWrites' own block.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 15:29:46 +02:00
Claudia 84132d3025 fix(ws): pin the withheld ACK with a test, and correct the changeset
Two blockers from the pre-submission gate, both reproduced before fixing.

1. The changeset claimed the non-mux POST branch answers OPERATION_FAILED. The
   code says the opposite in as many words ("NOT an error response,
   deliberately"), the commit message says response codes are unchanged, and the
   test asserts the 200. It was a leftover sentence from an earlier iteration that
   would have shipped into the CHANGELOG announcing an API contract change that
   does not exist — and errorCode values are SemVer-relevant per
   docs/versioning-policy.md.

2. The WebSocket half of the fix had no test protection: reverting ws-routes.ts to
   master left all 9 tests green, while the commit message sells "plus the whole
   WebSocket path" as part of the fix. Three tests added against the real WS
   route — ACK on delivery, ACK withheld and seq re-opened when the write did not
   land, and a deduplicated frame still ACKed so the client can drop it. Verified
   the other way round: with ws-routes.ts reverted, the middle one fails.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 14:52:44 +02:00
Codeman maintainer cc163792e5 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:43:47 +02:00
Ark0N 6f1ff17ccc Merge pull request #223 from Ark0N/fix/scrollback-shell-alt-screen
fix: terminal scrollback overhaul for shell and CLI sessions (#205)
2026-08-07 13:42:47 +02:00
Codeman maintainer f262b8cb69 feat(terminal): gentler glide start and fractional wheel accumulation
Two smoothness refinements on the local wheel path: the drain factor
drops from 35% to 22% per frame, so the first frame of a notch takes a
smaller step and the glide lasts longer; and local scrolling accumulates
FRACTIONAL lines (_wheelScrollLinesFloat) instead of rounding every
event, so a slow macOS trackpad drag no longer snaps a whole line per
tiny delta (the old ±1 fallback made slow drags scroll faster than the
finger). Sub-line residuals stay pending until further input crosses a
whole line. Forwarded SGR ticks keep the rounded integer path. Probe:
a 20-line notch now glides through 14 positions to an exact landing;
the 9-check scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:36:27 +02:00
Codeman maintainer 5f2b491d99 feat(terminal): ease-out smooth scrolling for the local wheel path
The capture-phase handler owns local scrolling (xterm's smooth scroller
is bypassed for the stale-dimensions reasons documented there), which
made every notch an instant multi-line jump. Wheel deltas now accumulate
into a pending line count drained ~35% per animation frame with a
one-line floor, so scrolling glides and extra notches mid-glide read as
acceleration. Pending momentum is dropped on session switch so it never
scrolls the tab the user just switched to. Verified on the beta: a
20-line notch eases over 9 frames to an exact landing, and the 9-check
scroll matrix still passes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:30:34 +02:00
Codeman maintainer c067167dbc fix(terminal): take the wheel in capture phase; xterm's scroller is deaf after reset
Measured on the live instance: xterm's vscode-style viewport scroller
consumes wheel events itself whenever it believes a scrollbar exists
(preventDefault + stopPropagation, attachCustomWheelEventHandler is not
consulted), so Codeman's bubble-phase handler never fired once local
scrollback existed. Forwarding, the deltaMode conversion and the
top-of-buffer history re-pull were all silently dead exactly on the
sessions that had history, which is the 'input box scrolls up then it
fights and hangs' report. Worse, that scroller's dimensions go stale
after terminal.reset(): following a tab switch or full-history replay it
neither scrolls nor propagates, which is the 'works at first, breaks
after reload and tab switch' report.

The container wheel listener now runs in capture phase, stops
propagation, and scrolls locally through buffer-level scrollLines(),
which keeps working after resets. Mouse-tracking sessions and the
alternate buffer (direct-PTY vim/less) are passed through untouched so
xterm's encoder and alt-scroll arrow conversion keep owning those.

Verified end to end against the beta: 9/9 matrix checks including the
exact reported flows (claude wheel with scrollback present stays pinned
and forwards, shell reaches full history by wheel alone, reload then tab
switch then back still works, SSE reconnect survives, Shift+wheel stays
local), plus the two prior E2E suites re-passing 10/10 and 6/6.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 13:03:17 +02:00
Codeman maintainer ad2ca9b575 docs: record the #205 scrollback mechanisms and the shipped fix plan
Update the full-scrollback replay invariant (per-session full=1 Set plus
the scroll-to-top re-pull), add a new invariants section covering the two
strip flavors and the wheel/touch forwarding rules, sync the CLAUDE.md
Key Patterns bullets, and commit the fix plan with a status header
describing what shipped and where it deliberately diverged (narrow strip
plus re-pull instead of tmux mouse on; viewport-at-bottom gate dropped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:42 +02:00
Codeman maintainer a7a1cef3d6 fix(session): probe the Claude CLI version over ssh for remote sessions
Remote Claude sessions were the one backend left relying on the
startup-banner scrape for cliVersion (the unreliable path #154 was filed
for: newer Claude Code builds print no banner and resumed sessions never
do), so wheel/touch forwarding silently stayed off for them. Mirror the
docker approach: a deferred best-effort probe at session start, running
claude --version on the remote host through the same
buildSshConnectionArgs + login-shell wrapper as the real launch, parsing
the first semver in stdout (an interactive login shell may echo rc-file
noise around it).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:41 +02:00
Codeman maintainer a1d7ec02e9 fix(terminal): forward touch scrolls to the CLI transcript on mobile
Touch drags and flick momentum on forwarding-capable sessions (codex,
claude >= 2.1.187) now go to the CLI as coalesced SGR wheel reports via
the shared _forwardScrollToApp helper, exactly like the desktop wheel:
snap the viewport home first, then encode. Before this, every phone or
tablet swipe scrolled the local buffer of stale repaint frames and
dragged the CLI's pinned input box off the screen (the mobile half of
issue #205). The _shouldForwardWheelToApp gate is shared, so the
local-scrollback opt-out setting and the CLI version gate apply to touch
exactly as they do to the wheel; shell and other local modes keep the
existing local touch scrolling and the scroll-to-top history re-pull.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-07 05:05:25 +02:00
Codeman maintainer dfa43928af docs: record the scrollback analysis and its measurements for #205 2026-08-07 04:33:15 +02:00
Codeman maintainer adbb74cd5a fix(terminal): keep the CLI's input box pinned when scrolling with the wheel
Reported against the beta: scrolling up in a Claude session drags the prompt
box and status line up the screen along with everything else, and only once
the local buffer hits its top does the CLI's own history start moving.

_shouldForwardWheelToApp() gated forwarding on the viewport being at the buffer
bottom, so that leaving the bottom handed the wheel back to local scrollback and
both histories stayed reachable. Two things make that the wrong default:

- A repaint-mode CLI keeps no terminal scrollback of its own (tmux reports
  history_size=0 for a Claude pane), so xterm's buffer holds only Codeman's
  REPLAYED repaint frames. Scrolling those locally moves the CLI's pinned
  furniture and shows stale frames underneath.
- scrollToLastNonEmptyLine() parks the viewport `rows - 2` above the last
  non-empty row, so any session with trailing blank rows was left off-bottom
  and every later wheel event went local without the user ever scrolling.

Forward unconditionally for the verified modes instead, and snap the viewport
back to the bottom before encoding the report (SGR coordinates address the live
screen, and forwarding while the user stares at stale scrollback looks dead).
Shift+wheel and the "Wheel scrolls local history" opt-out still reach local
scrollback.

Verified against a real Claude 2.1.223 session: wheel-up scrolls its transcript
back 48 lines (rows showing 85-92 -> 37-44) while the input box, separator and
status line stay fixed at the bottom.
2026-08-07 04:27:16 +02:00
Codeman maintainer eb8d11ffc3 fix(terminal): restore shell scrollback, recover history lost to tmux repaints
Four fixes for the scrollback reports in #205 (plus its follow-up comment).

1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's
   ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup
   (\x1b[?1049h) as its first bytes on attach, and the existing strip is gated
   to claude/codex/gemini, so it reached the browser verbatim. In the alternate
   buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op)
   and xterm's own wheel handler translates the wheel into \x1bOA cursor keys,
   which readline receives as shell history navigation. Both reported symptoms,
   one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those
   modes, but ONLY under tmux (the direct-PTY fallback still needs a program's
   own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and
   the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux
   never forwards a pane's alt-screen toggles to its client, it repaints;
   captured from a real attach, vim/less/htop emit zero.

2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window
   onto tmux's history, and tmux repaints the pane rectangle instead of emitting
   linefeeds whenever output outpaces its flush, OVERWRITING already-rendered
   scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the
   same 60 lines emitted slowly added all 60. Scrolling up at the top now
   re-pulls the full tmux scrollback and holds the user's place. Verified
   end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines.

3. The full-scrollback replay was gated on a single "first load after page load"
   flag, which whichever session auto-selected consumed, so every other tab
   started with one visible frame. Now tracked per session.

4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3
   per notch) scrolled one line where Chrome scrolls four or five, and capped
   the forwarded SGR report at one tick. Line and page deltas are now converted,
   and a pure horizontal swipe no longer falls through to a phantom -1.

Analysis and measurements: docs/scrollback-issues-analysis.md
2026-08-07 04:06:54 +02:00
Claudia ebfcac6ad1 fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the
frame BEFORE knowing whether the write had landed: the POST route because its mux
write is fire-and-forget so the response never waits on a tmux child, the
WebSocket handler because it ACKed unconditionally.

When the write then failed, the client dropped the frame from its durable queue
and the server rejected the retry as a duplicate. The reliable-delivery layer was
guaranteeing exactly-once delivery of something that had never been delivered —
and `Session.write()` returned void, so a session whose PTY was gone swallowed the
data with no signal at all.

- `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq
  is still the newest one; a later input has superseded it and must not re-open.
- The WebSocket handler withholds its ACK when the write did not land, so the
  client redelivers.
- `Session.write()` reports whether it reached a PTY.

Response codes are unchanged, deliberately: a session can legitimately have no PTY
yet, and turning that into a failure status would be a contract change of its own.

What this does NOT do: remove the root cause. The POST still answers 200 before
the mux write is attempted, so a client that treats any 2xx as final cannot learn
about that failure. What closes is the narrower window — the write failed AND the
ACK never reached the client — plus the whole WebSocket path. Closing the rest
would mean awaiting the tmux child inside the request.

9 tests. They drive the HTTP route, not only the Session primitives: with the
rollback removed from the route, 2 of them fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:36:33 +02:00
Claudia 2e69e28e71 fix(mux): bound the process-tree walk — it can take a machine down
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Two further sites forked a `pgrep` per session on
every stats tick.

Across ~28 adopted tmux trees the fan-out exploded, and because each `pgrep`
blocks in the kernel while reading `/proc/<pid>/cgroup` under WSL, none returned
while the walk kept spawning more. Observed: ~13,000 `pgrep` processes stuck in
D-state out of ~39,000 total, load average above 13,000, and a machine only
recoverable by restarting WSL — which cost every running session. Every diagnostic
command timed out too, because they read /proc as well.

- ONE `ps -eo pid=,ppid=` snapshot, cached briefly and refreshed asynchronously
  with a single-flight guard. Async matters: under the same procfs pathology,
  `execSync`'s timeout cannot return (spawnSync waits for the unkillable child),
  which would freeze the server where a hung async poll only costs staleness.
- The traversal moved to `proc-tree.ts` as a pure function — breadth-first, with a
  visited set (a stale snapshot can contain a cycle), a depth cap and a node cap,
  both reporting when they truncate. Pure so the regression tests can exercise the
  shipped code rather than a copy of it.
- The kill path forces a fresh snapshot: the wait between SIGTERM and the survivor
  re-scan (200ms) sits inside the cache TTL (2000ms), so reading the cache there
  would return pre-SIGTERM state and aim SIGKILL at stale PIDs. That wait is
  bounded, so a wedged `ps` cannot stop killSession from reaching its
  process-group and tmux fallbacks.
- Any `ps` error keeps the previous snapshot instead of caching partial output as
  fresh; a truncated table would make whole subtrees invisible to the kill path.

13 tests, including one that drives TmuxManager itself — with the caps bypassed at
the call site, 3 of them fail. The snapshot refresh is stubbed there, because
otherwise the manager runs a real `ps`, replaces the fixture, and the test
silently measures the machine's own process tree instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:34:47 +02:00
Claudia 1a32e63765 fix(http): raw writeHead routes lost every header the security hook set
`reply.raw.writeHead()` writes straight to the Node response and bypasses
Fastify's header store, so everything the `onRequest` security hook granted is
silently dropped on every route that answers that way.

The visible symptom is CORS. The hook emits `Access-Control-Allow-Origin` for
localhost origins, so a page served from a local dev server may call every `/api`
endpoint cross-origin — except the four below, whose requests fail. The security
headers (`X-Content-Type-Options`, `X-Frame-Options`, CSP) were being lost the
same way.

Affected: `GET /api/events`, and `file-raw` / `tail-file` / `download` in
file-routes.ts. Each now spreads the inherited headers first and lets its own
headers win over them.

Tests drive a real WebServer and compare `/api/events` against `/api/status` for
the same Origin — the point of the fix being that the SSE route stops being the
odd one out. Verified in both directions: with the fix removed, 3 of the 5 fail.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 01:31:50 +02:00
Codeman maintainer d41f28bc14 docs(docker): warn that a plain agent-image rebuild keeps stale CLIs
The CLIs live in one `RUN npm install -g` layer, so rebuilding without
--no-cache re-uses it and freezes them at the versions the image was FIRST
built with. Editing the Dockerfile does not help when the edit lands below
that line: the npm layer stays cached and only the new step runs.

That is not hypothetical. Adding the Antigravity step (which appends below
the npm line) produced a "successful" rebuild that silently kept a stale
@openai/codex@0.144.6 whose aliased platform binary had never installed, so
every codex docker case died with "Missing optional dependency
@openai/codex-linux-x64" while the build reported success. A --no-cache
rebuild fixed codex and also un-froze claude, gemini and opencode.

Documents the failure, makes --no-cache the recommended invocation in both
the guide and the CLAUDE.md quick-reference row, and adds a verify command
that actually executes each CLI, since a zero exit code only proves the
layers ran.

No changeset: docs-only, rides the next release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 09:02:28 +02:00
Codeman maintainer 322f21ef9f docs(extending): scope the "no sandbox" claim, point at Docker cases
The bullet read as a blanket "Codeman has no sandbox", which is wrong and
undersells a headline feature. Two different axes were conflated:

- Integration code cannot be sandboxed by Codeman because Codeman never
  launches it. It is the reader's own process, started by them.
- Agent workloads are sandboxed per case via Docker cases, which is the
  documented isolation story.

Scopes the claim to integration code and links docs/docker-cases.md, noting
that an integration driving a Docker-backed session inherits that isolation
because it is a property of the session, not the caller.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:46:22 +02:00
Codeman maintainer c2d973cb2d docs: link the integration guide from both READMEs, fix three inaccuracies
Adds a pointer to docs/extending-codeman.md at the end of the API section in
README.md and README.zh-CN.md, so the guide is reachable from where people
read about endpoints rather than only from CLAUDE.md.

Reading the README's programmatic guide alongside the new page surfaced three
errors in it, all now fixed:

- POST /api/sessions/:id/input takes `useMux`, not `useScreen`. The latter is
  a legacy name that no longer appears in the schema.
- The page told integrators to send `\r` to submit. With `useMux: true` the
  server delivers text and Enter as two separate writes (writeViaMux does
  send-keys -l then send-keys Enter), so appending `\r` is wrong.
- "Unwrap the envelope" was incomplete: a few legacy GETs put the payload at
  the top level, so the advice is now `body.data ?? body`.

Also cross-references the README's programmatic guide, which covers the
in-session case (CODEMAN_MUX, CODEMAN_API_URL, CODEMAN_SESSION_ID,
CODEMAN_HOOK_SECRET_FILE) that the new page deliberately does not duplicate,
and documents the optional clientId/seq exactly-once fields.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 08:34:58 +02:00
Codeman maintainer 84e31c0ee1 chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:31:02 +02:00
Codeman maintainer 0d0b772619 feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.

Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
  mode 'antigravity' died on command-not-found. agy is not on npm, so it
  gets its own installer step. --dir /usr/local/bin is load-bearing: the
  default $HOME/.local/bin resolves to root's home at build time and is
  unreachable by the `agent` user the container runs as. Verified inside
  codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
  ~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
  present like the other CLI buttons, with a cyan identity matching the
  toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
  it as a satisfying AI CLI, and recommends it over Gemini in the install
  hints. Detection only, no new auto-install path.

Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
  opencode/codex/gemini when the code has included antigravity for a
  while, said "all three modes", and omitted ANTIGRAVITY_ from the env
  prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
  remote-sessions' RemoteCommandMode were all stale.

Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.

test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.

Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.

isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:22:29 +02:00
Ark0N bfff20a093 Merge pull request #216 from shenlvkang-collab/fix/response-viewer-brief-format
fix(web): align brief Response Viewer formatting
2026-08-06 07:22:13 +02:00
codeman-local b982c5d0e0 fix(web): align brief response viewer formatting 2026-08-06 10:16:15 +08:00
Codeman maintainer f50c922240 docs: add extending-codeman.md, the third-party integration guide
Codeman has no plugin runtime by design: running third-party code inside
the process that spawns agents, on a server people expose over a tunnel,
would trade away the security posture that is a reason to use it. But it
already has four extension seams that work from any language with nothing
installed, and they were undocumented.

Documents web tabs (render your own UI as a tab), the SSE event channel
(react when an agent needs you), the HTTP API plus the codeman CLI (drive
it from a script), and hook events. Every endpoint, schema field, event
name and header in the page was read from source and then verified against
a running instance, including the localhost-only CORS behavior and the SSE
framing the example depends on.

Also corrects a stale line in CLAUDE.md: it claimed the HTTP/SSE API was
internal/unstable, which contradicts docs/versioning-policy.md, where the
API under /api/v1 was finalized as part of the stable surface for the 1.0
cut. No new stability commitment is made here; the page makes an existing
one discoverable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 02:03:39 +02:00
Codeman maintainer de5b048c3f docs(vm): VM cases plan + Apple virtualization stack reference
Two design/reference docs for the planned native-macOS VM isolation tier
("VM cases"), a location overlay on cases in the same shape as Docker and
remote-SSH cases, never a sixth SessionMode. Nothing is implemented; both
docs are marked PLANNED and are blocked on macOS 27 GA.

- vm-cases-plan.md: the Codeman-side design and phased plan. Swift helper
  CLI, DiskImageKit base + per-case overlay, sessions riding the existing
  remote-SSH machinery, VirtioFS workspace at the same absolute path, and
  seeded credentials, each mirroring an established Docker-cases rule.

- vm-subsystem-apple-stack.md: what the Apple stack actually provides,
  measured on the 27 beta rather than inferred from the WWDC session. Of
  note: the 2-concurrent-macOS-VM cap is a kernel quota (refused at 39%
  free RAM, so more hardware does not help), DiskImageKit has no flatten
  API so exports must ship the layer chain, and a macOS guest renders
  nothing without an attached view in an unlocked host session.

No credentials, hostnames, tailnet addresses or account names in either
file; every host/guest reference is a placeholder.

Also joins a table row that a stray blank line had split off into its own
malformed table.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 01:28:17 +02:00
Codeman maintainer 12a5f5919e chore: version packages 2026-08-05 22:36:51 +02:00
Codeman maintainer ecd3f3f32a harden(history): exclude automated transcripts by SDK shape, not by "not cli"
#215 filters non-interactive transcripts out of Past Sessions with
`entrypoint !== 'cli'`. That is an allowlist on a value, and the check
hides rows, so it fails CLOSED on anything Claude Code has not shipped
yet: the day it stamps a new interactive entrypoint (a rename, or a
second interactive host), no transcript matches 'cli' any more and the
entire Past Sessions list goes blank with nothing in the UI explaining
why.

Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`).
An automated entrypoint we do not recognize yet now costs a few noisy
rows, which is the annoyance the filter set out to fix, rather than a
dead feature. Matches the fail-open reasoning #215 already applied to a
MISSING entrypoint field; only the unknown-VALUE case was inverted.

Test fails against the pre-fix line and passes after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 21:47:06 +02:00
Ark0N c19d884a51 Merge pull request #215 from timkjr/fix/past-sessions-history-quality
fix(history): three Past Sessions data-quality bugs (automated-session noise, cross-contaminated previews, blank restart-heavy rows)
2026-08-05 21:44:47 +02:00
Ark0N 22e77a1827 Merge pull request #214 from timkjr/fix/mobile-overview-run-gating
fix(mobile): gate the phone overview's run picker on CLI availability
2026-08-05 21:44:42 +02:00
Ark0N b641560040 Merge pull request #203 from shenlvkang-collab/contrib/claude-viewer-session-pin
fix(web): pin the Claude response viewer to the pane's own conversation
2026-08-05 21:44:37 +02:00
timkjr 8300c15cbd fix(history): entrypoint detection was first-field-wins, plus a two-tier head read
extractTranscriptEntrypoint returned the FIRST entrypoint-bearing message's
value instead of scanning for any 'cli' occurrence, so a transcript that
started under an older Claude Code build (no entrypoint field) and later
picked up a non-'cli' entrypoint on some later message was wrongly excluded
from history — the opposite of the fail-open behavior the function's own
comment claimed. Now returns 'cli' the moment any scanned message carries it,
and only falls back to a non-cli value when nothing else qualifies. Head/tail
entrypoints are merged the same way (either side being 'cli' wins).

Also restructures scanProjectDir's head read into two tiers: try 16KB first
and escalate to 128KB only when that wasn't enough, instead of reading 128KB
for every file unconditionally. Measured against a real ~/.claude/projects
tree, the unconditional-128KB version roughly quadrupled scan cost to fix a
problem only a minority of files actually have; the two-tier version cuts
bytes read by ~36% and wall time by ~17% while producing identical output.
Also fixes a fallback regression where a failed head read (e.g. EMFILE) on a
file at or under the head buffer size no longer got a shot at the tail-read
fallback, silently dropping the session from history.
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 09f5f28017 docs(test): correct an overclaiming comment in the tail-fallback regression test
The comment implied the fallback could be "silently skipped" by the
stale hardcoded threshold, which isn't actually true -- the old
smaller numbers were always more eager to trigger the fallback, never
less (same correction as the commit this test belongs to). What the
test actually protects against is the fallback logic itself breaking
(e.g. a copy-paste slip dropping the check entirely), not the exact
threshold value. Reworded to say that.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 251706be3b harden: scope entrypoint detection to message lines, add fallback coverage
Two follow-ups from reviewing the entrypoint-filter and head-buffer
fixes before submitting them upstream:

1. extractTranscriptEntrypoint() scanned any line containing the
   substring "entrypoint", not specifically the first "type":"user"/
   "type":"assistant" message line (unlike its sibling
   extractFirstUserPrompt, which does scope to type). A transcript
   that started under an older Claude Code version (no entrypoint
   field) and got resumed under a newer one mid-conversation could
   pick up the field from a much later message than the true first
   one, misattributing the session's origin. Scoped it to match.

2. Added a regression test proving the tail-read fallback still
   engages correctly when bookkeeping accumulation exceeds even the
   new 128KB head window, not just the 16KB it previously blanked at.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 18b473f0e4 fix(history): raise the transcript head-read window to fit restart bookkeeping
Blank firstPrompt rows weren't all oversized messages -- traced one
directly: a session restarted many times (mux deaths, redeploys)
accumulates a batch of small bookkeeping lines (mode/permission-mode/
last-prompt/queue-operation, one batch per restart) ahead of the real
first message. With enough restarts these alone crossed the old 16KB
head-read window, so extraction found nothing even though the actual
first message was tiny (measured case: ~17.5KB of bookkeeping pushed a
189-byte real message just past the boundary).

Raise the head buffer from 16KB to 128KB (matching the existing
precedent at the codex-history head-read a few hundred lines up) and
fix three now-stale `> 16384`/`> 65536` fallback thresholds to
reference headBuf.length instead of hardcoded numbers, so the tail-read
fallbacks stay correctly scoped to "beyond what head already covered."

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 a2aed38073 fix(unified-sessions): stop the firstPrompt workingDir backfill from cross-contaminating history rows
COD-140's backfill was meant to cover live/persisted rows whose Codeman
id doesn't match an on-disk transcript UUID, guessing from the newest
transcript in the same workingDir as a last resort. It was also firing
for pure history rows whose OWN transcript scan already ran (and
genuinely found nothing, e.g. an oversized first message) -- those got
silently backfilled with the newest OTHER session's opening line from
the same directory. Not a blank row, but actively wrong: old sessions
displayed a completely unrelated (often today's live) conversation's
first prompt as if it were their own.

Skip the workingDir guess for any item that already has its own
'history' source -- it already had a real, direct attempt. Rows with
no history source at all (their transcript isn't linked/scanned under
their own id yet) still get the guess, matching the original intent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 e888c65c52 fix(history): exclude non-interactive (SDK-driven) transcripts from Past Sessions
Automated tools (CI review bots, etc.) invoke Claude Code via the SDK
and write their transcripts into the same ~/.claude/projects tree as
real interactive sessions, but were never something a user can resume
into -- no PTY, no running process. Their one-shot review prompts also
embed the full diff inline as a single message, often exceeding the
16KB head / 32KB tail windows this scanner reads, so they cluttered
Past Sessions two ways: as blank rows when the huge message couldn't
be parsed, or as N identical "Review this change for security
vulnerabilities..." rows when it could.

Claude Code stamps `entrypoint` on its own message records ('cli' for
a real interactive session, e.g. 'sdk-py' for an SDK invocation).
Exclude any transcript whose entrypoint isn't 'cli' from the history
list entirely, checked last so it reuses whatever head/tail the prompt
extraction already read. Missing entrypoint (older transcripts) reads
as interactive -- fail open, matching every other gating check in this
codebase. Shared by /api/history/sessions and /api/sessions/unified,
since both call the same scanProjectDir().

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:18 -05:00
timkjrandClaude Sonnet 5 1ea39de650 fix(mobile): gate the phone overview's run picker on CLI availability
MOBILE_OVERVIEW_RUN_MODES / _buildMobileOverviewRunMenu is a separate,
hardcoded duplicate of the toolbar's #runModeMenu (mobile-overview.js
is a newer feature that mirrors the toolbar menu's look/behavior
rather than reusing its render), so it never picked up #201's
isCliAvailable() gating and offered every backend regardless of what
the server actually has installed.

Gate it the same way: skip an entry unless isCliAvailable(mode),
shell always exempt. Added functional + static regression tests
mirroring the toolbar menu's own test pattern.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-05 11:11:15 -05:00
Codeman maintainer e2a644997e chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 09:01:45 +02:00
Ark0N cd5a101626 Merge pull request #213 from Ark0N/feat/file-viewer-edit-mode
File Viewer: edit mode for text files (edit + save in the viewer)
2026-08-05 09:00:22 +02:00
Codeman maintainer 4ea781c80f feat(file-viewer): edit mode for text files (edit + save in the viewer)
Closes #212. The file-preview overlay can now edit workspace text files in
place, phone-first: agent writes a file, you review it in the viewer, tweak
two lines, save, tell the agent to continue.

Backend (file-routes.ts, policy in src/config/file-editing.ts):
- GET file-content?edit=1: read-for-edit that never truncates (a truncated
  buffer must never become an edit buffer), 512KB cap (413 over it), and
  returns the sha256 hash + detected EOL the client echoes back on save.
- PUT /api/sessions/:id/file-content: edit-in-place only, with no O_CREAT
  anywhere in the handler. Confinement matches the read path (realpath +
  workspace boundary + ownership via findSessionOrFail), plus sensitive-path
  and attachment-guard blocklists, a .git subtree deny, and an extension
  allowlist (svg and env deliberately excluded). Optimistic concurrency via
  baseHash: mismatch is a 409 unless force. Writes are wx-temp + fchmod +
  fsync + rename, closing the validate-then-write TOCTOU window.
- Corruption guards: NUL sniff + UTF-8 round-trip compare (refuses binary
  and latin-1), and server-side EOL re-application so a textarea's LF
  normalization cannot rewrite every line of a CRLF file.
- Plain reads gain an additive editable flag the UI keys the button off.

Frontend (panels-ui.js + overlay markup/styles):
- Edit button on editable text previews; textarea editor with Save/Cancel,
  dirty indicator, discard-confirm on cancel/close, and a conflict dialog
  that offers overwrite (force) when the file changed on disk mid-edit.
- Phone: full-bleed window sized by --app-height so the editor and Save bar
  track the OS keyboard; 16px editor font (iOS zoom guard); no autofocus.
- zh-CN strings for the new chrome.

Tests: pure policy unit tests plus a route suite that deliberately does NOT
mock node:fs. It runs against a real temp workspace so symlink escapes,
write-through of in-workspace symlinks, mode preservation, CRLF round-trip,
409/force, and the no-create property are exercised for real. Also verified
end to end on an isolated beta instance: 39-check curl matrix, Playwright
desktop flow (real clicks and typing, bytes asserted on disk, live conflict
with an external rewrite), and a 393px phone profile.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 08:44:47 +02:00
Codeman maintainer d9123de9eb feat(terminal): Ctrl+C copies the selection, interrupts when nothing is selected
Closes #211. Copying from the terminal only worked through the browser
context menu, because xterm turns Ctrl+C into 0x03 and cancels the keydown,
so the muscle-memory copy failed silently and read as "no copy-paste at all".

With a selection, Ctrl+C now copies it, toasts, clears the selection and
sends nothing to the PTY. With no selection it falls through unchanged, so
the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never
falls through: an explicit copy that interrupts a running agent because the
selection happened to be empty would be a footgun.

Three details that keep the interrupt safe:

- The decision lives in attachCustomKeyEventHandler (terminal-ui.js) and the
  no-selection path returns true WITHOUT preventDefault. xterm calls the
  custom handler before its own cancel(), so returning false alone does not
  cancel the event; the copy path therefore calls preventDefault explicitly,
  or the browser would run its native copy on top of ours.
- copy-selection is a full registry entry (rebindable and disableable in App
  Settings) whose action is deliberately absent from SHORTCUT_ACTIONS, the
  same trick command-palette uses: the generic capture loop preventDefaults
  every match it dispatches, which would cost the user the interrupt key.
- The gate is keydown-only, since the custom handler also runs for keypress
  and keyup.

Copy goes through _copyText (Clipboard API, then hidden-textarea +
execCommand) rather than raw navigator.clipboard, because install.sh's LAN
option serves plain HTTP where navigator.clipboard is undefined; the
fallback steals focus, so the terminal is refocused afterwards.

Tests: test/terminal-copy-selection.test.ts pins the gate and the
SHORTCUT_ACTIONS invariant; test/terminal-copy-shortcut.test.ts drives real
key presses in chromium and asserts on the clipboard plus the bytes xterm
emitted (browser-driven, so excluded from test:ci like the other Playwright
suites). Verified manually on an isolated beta instance before landing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 02:38:57 +02:00
shenlvkang-collab ab7a703e90 fix(web): keep the viewer's conversation anchor across a Codeman restart
start() reassigns _claudeSessionId to `resumeSessionId || id` on every launch,
including the path that re-attaches to a mux session that outlived the restart.
A pane whose CLI had moved on via /clear therefore came back pointing the
response viewer at its pre-/clear transcript, and because Session.lastSubmitAt
lived only in memory, the history correlation had nothing to correct it with
until the user happened to type again — observed as hours of the eye showing a
conversation the pane had long since left.

Persist lastSubmitAt in SessionState, restore it in restoreMuxSessions(), and
flush it when the viewer adopts (a /clear emits no completion event, which is
the trigger that would otherwise have persisted it). Recovered panes now
re-derive their live conversation on the viewer's first poll.

Restoring a stale anchor is safe: the resolver already refuses a candidate
transcript older than the one the pane is currently on, which is the shape of a
respawn into a fresh conversation.
2026-08-03 21:22:33 +08:00
shenlvkang-collabandClaude Opus 5 73315bc351 fix(web): pin the Claude response viewer to the pane's own conversation
The viewer re-derived a pane's live conversation from the newest
~/.claude/history.jsonl entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` the user runs in their own terminal, so the eye followed whichever of
those was typed into last — and since the match was written back through
adoptClaudeSessionId(), the mispin stuck.

Credit a history entry to a pane only when it lands within 10s of that pane's
own Enter and no other pane on the same cwd submitted closer, reusing the
last-submit correlation the Codex locator already relies on. Submit tracking
moves from _codexLastSubmitAt to a mode-agnostic Session.lastSubmitAt. With no
correlated entry the pane keeps the id it has: a viewer one turn behind beats a
viewer showing someone else's conversation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-03 14:53:42 +08:00
369 changed files with 73812 additions and 3599 deletions
+89
View File
@@ -0,0 +1,89 @@
# Contributing to Codeman
Thanks for wanting to help! Codeman is a small project with a fast loop: issues usually get a response within a day, good PRs get reviewed quickly, and every release credits its contributors and bug reporters by name in the release notes. This guide gets you from clone to merged PR without stepping on the traps.
## The short version
1. **Bugs**: open an issue with your OS, install method (installer / npm / git clone), browser, and which CLI + version the session was running.
2. **Questions and ideas**: use [Discussions](https://github.com/Ark0N/Codeman/discussions), not issues.
3. **Small fixes** (docs, typos, a new skin, a translation): just send the PR.
4. **Anything bigger**: open an issue or Discussion first and get a nod before building. Codeman has strong architectural invariants, and a design chat up front is what turns a big idea into a merged PR instead of a stalled one. This flow works: features like Clone Repo (#236) went idea, then design discussion, then review, then shipped.
5. **Security issues**: never a public issue. See [SECURITY.md](SECURITY.md).
## Dev setup
Requirements: Node.js 22+ (see `.nvmrc`), tmux, and at least one supported agent CLI on your PATH (Claude Code is the primary one).
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
The frontend is plain JS served from `src/web/public/` with no bundler in dev: edit a `.js`/`.css` file and reload the page. The one exception is `index.html`, which is read once at server start, so markup changes need a server restart.
## Before you push
CI runs all of these, so save yourself a round trip:
```bash
npm run typecheck # tsc --noEmit, strict mode
npm run lint
npm run format:check
npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
```
### Tests
```bash
npm test # the gate — exactly what CI runs
npm test -- test/<file>.test.ts # one file
```
`npm test` is the same suite CI runs, so a green run locally means a green run there. It leaves out three suites that cannot pass on an arbitrary machine, each with its own command:
```bash
npm run test:browser # Playwright + chromium (+ a live server; codex-predictive-echo needs a real codex binary)
npm run test:mobile # the above plus environment-specific PNG baselines
npm run test:perf # wall-clock benchmarks — run on an otherwise idle machine
npm run test:all # literally everything, environmental failures included
```
Expect `test:browser`/`test:mobile`/`test:perf` to fail where the machine cannot provide what they need; read that as "not runnable here", not as a regression. `config/test-suites.ts` holds the globs, and both configs derive from it, so the exclusions and those runners cannot drift apart.
If you add a test that binds a port, pick a unique one at 3150 or above (search the repo for `const PORT =` first). Never 3000.
Tests are tmux-safe by design: under vitest, the tmux layer becomes an in-memory mock, so tests cannot touch real sessions.
## Finding your way around
- Every source file starts with a `@fileoverview` JSDoc block. Read it before diving into the file, it is the map.
- [`CLAUDE.md`](../CLAUDE.md) at the repo root is the densest architecture primer in the repo. It is written for AI coding agents, but the invariants and gotchas in it apply to humans exactly the same, and most review feedback on PRs traces back to something already written there.
- Deep mechanisms and the history behind each rule live in [`docs/architecture-invariants.md`](../docs/architecture-invariants.md).
- Third-party extension surfaces are documented in [`docs/extending-codeman.md`](../docs/extending-codeman.md).
## Great first contributions
These are well-fenced areas where a first PR is genuinely easy to get right:
- **A new theme skin.** A skin is four things kept in sync: the `html[data-skin="…"]` token block in `styles.css`, the xterm ANSI palette in `terminal-ui.js`, the pre-paint allowlist and the settings picker (both in `index.html`). `test/skin-themes.test.ts` statically checks the sync, so if the test passes, your skin works.
- **A new language.** `src/web/public/i18n.js` is dependency-free, English is the canonical source, and `zh-CN` is a complete example to copy. Add your language's entries and register it in `SUPPORTED_LANGUAGES`.
- **Docs.** If you got stuck on something and then figured it out, the sentence that would have unstuck you is a PR.
- Anything labeled [`good first issue`](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
Bigger extension points worth discussing first: new CLI backends (the pluggable resolver pattern has absorbed six CLIs so far; `docs/extending-codeman.md` and `docs/opencode-integration.md` show the shape), and real-device testing reports, especially mobile, which always find things emulation cannot.
## PR expectations
- **One change per PR.** Small and focused reviews fast; a grab-bag stalls.
- Target the `master` branch.
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all (GitHub quirk), so rebase or merge master when conflicts appear.
- Include or update tests when you change behavior. Route handlers have a lightweight pattern in `test/routes/` using `app.inject()` (no live server needed).
- Formatting is Prettier with a deliberately narrow scope (`npm run format`), several frontend files are hand-formatted on purpose and excluded via `.prettierignore`. Don't "fix" a file by adding it back into Prettier's scope.
- Don't bump versions or touch `CHANGELOG.md`; releases are handled by the maintainer via changesets after merge.
- AI-assisted contributions are welcome (much of Codeman is built that way), with one condition: you must understand what you're submitting and have actually run it. "The model said it works" is not a test.
## Conduct
Be kind, be direct, assume good faith. Report unacceptable behavior privately via the contact in [SECURITY.md](SECURITY.md).
+19 -4
View File
@@ -87,10 +87,25 @@ jobs:
fi
- name: Run unit & integration tests
# Excludes the browser-driven mobile suite (test/mobile/**); see config/vitest.ci.config.ts.
# Excludes the suites that need chromium, per-machine PNG baselines or a
# quiet machine — see config/test-suites.ts for the list and the reason
# behind each entry. Identical to what `npm test` runs locally.
# Safe in CI: TmuxManager no-ops all shell commands under VITEST (test/setup.ts).
run: npm run test:ci
# Note: The browser-driven mobile suite (test/mobile/**) is excluded from CI —
# it needs a live server + chromium + environment-specific PNG baselines.
# Run it locally/manually. All other tests run via the `test` job above.
- name: Run xterm-zerolag-input package tests
# Layers 1-3 of the predictive-echo suites (unit laws, fixture replay,
# seeded fuzz): deterministic, no browser, no live server. Depends on
# the ROOT `npm ci` above — workspaces hoist the package's vitest into
# the root node_modules; do not add a separate install here.
run: npx vitest run
working-directory: packages/xterm-zerolag-input
# Note: three suites are excluded from CI, each with its own local runner:
# npm run test:browser Playwright + chromium (+ a live server, and a real
# codex binary for codex-predictive-echo)
# npm run test:mobile the above plus environment-specific PNG baselines
# npm run test:perf wall-clock benchmarks; need an otherwise idle machine
# config/test-suites.ts holds the globs; the configs derive from it so the
# exclusions here and those runners cannot drift apart. Everything else runs in
# the `test` job above, which is the same thing `npm test` runs.
+109
View File
@@ -0,0 +1,109 @@
name: Sync Wiki
# Publishes docs/wiki/ to the repository's GitHub wiki.
#
# The wiki is a separate git repo with no CI and no review, so the source of truth
# lives in docs/wiki/ and this workflow mirrors it. Browser edits to the wiki are
# overwritten by the next sync; fix pages with a PR against docs/wiki/ instead.
#
# One-time setup: GitHub only creates <repo>.wiki.git once the first page has been
# saved in the browser. Save a stub page at /wiki/_new before the first run.
#
# Token: GITHUB_TOKEN can push to the wiki on most repos but not all. If a run fails
# with 403, add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret;
# it is preferred automatically when present. Note the 403 usually surfaces on the
# PUSH, not the clone: this repo is public, so a read-only token still clones the
# wiki fine. Both steps carry the hint.
on:
push:
branches: [master]
paths:
- 'docs/wiki/**'
- '.github/workflows/wiki-sync.yml'
workflow_dispatch:
concurrency: ${{ github.workflow }}
jobs:
sync:
name: Push docs/wiki to the wiki
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- name: Checkout repo
uses: actions/checkout@v6
- name: Clone wiki
env:
WIKI_TOKEN: ${{ secrets.WIKI_TOKEN || secrets.GITHUB_TOKEN }}
run: |
set -euo pipefail
if ! git clone "https://x-access-token:${WIKI_TOKEN}@github.com/${GITHUB_REPOSITORY}.wiki.git" wiki 2>"${RUNNER_TEMP}/clone-err.txt"; then
cat "${RUNNER_TEMP}/clone-err.txt"
echo "::error::Could not clone ${GITHUB_REPOSITORY}.wiki.git. If this says 'Repository not found', the wiki has never had a page: save one at https://github.com/${GITHUB_REPOSITORY}/wiki/_new and re-run. If it says 403, add a WIKI_TOKEN secret."
exit 1
fi
- name: Mirror pages
run: |
set -euo pipefail
# The mirror deletes before it copies, so an empty source would wipe
# every published page and the commit step would happily push that. A
# MISSING directory already fails safely (cp aborts under set -e); an
# empty one does not, so check explicitly. This is the one failure mode
# here that destroys something a browser edit cannot get back.
if [ ! -d docs/wiki ]; then
echo "::error::docs/wiki does not exist. Refusing to mirror, which would delete the entire published wiki."
exit 1
fi
pages=$(find docs/wiki -maxdepth 1 -name '*.md' | wc -l)
if [ "$pages" -eq 0 ]; then
echo "::error::docs/wiki contains no .md pages. Refusing to mirror, which would delete the entire published wiki."
exit 1
fi
echo "Mirroring ${pages} pages."
find wiki -mindepth 1 -maxdepth 1 ! -name '.git' -exec rm -rf {} +
cp -R docs/wiki/. wiki/
- name: Stamp the documented version
run: |
set -euo pipefail
# _Footer.md renders on every page and used to carry a hand-written
# version, which went stale on every release because nothing refreshed
# it. It carries {{VERSION}} instead and the series is stamped here.
series="$(node -p "require('./package.json').version.split('.').slice(0,2).join('.') + '.x'")"
# grep exits 1 when it matches nothing, which under `set -o pipefail`
# would fail the step instead of warning, so test before substituting.
if grep -rlq '{{VERSION}}' wiki/; then
grep -rlZ '{{VERSION}}' wiki/ | xargs -0 -r sed -i "s/{{VERSION}}/${series}/g"
else
echo "::warning::No {{VERSION}} placeholder found in docs/wiki. The published version line can no longer be refreshed automatically."
fi
if grep -rq '{{VERSION}}' wiki/; then
echo "::error::A {{VERSION}} placeholder survived substitution and would be published verbatim."
exit 1
fi
echo "Stamped version ${series}."
- name: Commit and push
run: |
set -euo pipefail
cd wiki
git config user.name 'github-actions[bot]'
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
git add -A
if git diff --quiet --cached; then
echo "Wiki already up to date."
exit 0
fi
git commit -m "docs: sync wiki from docs/wiki @ ${GITHUB_SHA:0:7}"
if ! git push 2>"${RUNNER_TEMP}/push-err.txt"; then
cat "${RUNNER_TEMP}/push-err.txt"
echo "::error::Could not push to ${GITHUB_REPOSITORY}.wiki.git. A 403 here means the token can read the wiki but not write it, which is the usual GITHUB_TOKEN case: add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret."
exit 1
fi
+1 -1
View File
@@ -10,7 +10,7 @@ sections here.
Quick pointers:
- Type check: `tsc --noEmit` · Lint: `npm run lint` · Format: `npm run format:check`
- Targeted tests only: `npm test -- test/<file>.test.ts` (bare `npm test` is unsafe in managed sessions)
- Tests: `npm test` (the CI gate, safe to run bare) or `npm test -- test/<file>.test.ts` for one file
- Route tests use `app.inject()`; new tests needing ports must pick a unique `const PORT =`
- Branch off `master` for all work; Conventional Commit-style messages (`fix(mobile): ...`)
- Never commit secrets or local state from `~/.codeman/`
+936
View File
@@ -1,5 +1,941 @@
# aicodeman
## 1.20.1
### Patch Changes
- Terminal input and scrollback fixes (PRs #327, #331):
- IME punctuation preserved (#327): keyCode 229 / `Process` key events are now delegated to xterm's CompositionHelper instead of being suppressed, so an active Chinese IME committing numbers and full-width punctuation (,。!? and friends) reaches the terminal correctly. The CJK input field sends the browser's committed text instead of guessing from `KeyboardEvent.key`, and the redundant Android orphan-input fallback is removed so xterm is the single input owner.
- Shell history replay bounded (#331): selecting a Shell session loads a bounded 1 MiB tail instead of replaying the entire multi-megabyte tmux scrollback on xterm's main thread; full history stays available via the explicit "Load full history" action. tmux history limits now apply correctly on both legacy tmux (global default set in the same command queue before pane creation) and tmux 3.7+ (per-pane targeting that never resizes or trims unrelated live panes). Also adds `Server-Timing` and `[TERMINAL-PERF]` timing stages for terminal loads, fixes `scrollToLastNonEmptyLine` double-counting scrollback rows, and keeps live output ordered behind snapshot replays.
### Thanks
- @dignfei for both fixes: the IME punctuation root-cause fix (#327) and the bounded shell history replay with the tmux history-limit correctness work (#331).
## 1.20.0
### Minor Changes
- Response viewer for OpenCode, Gemini, Antigravity and Pi sessions (#326). External CLIs render their own TUIs, so the viewer used to come up empty for them; a new transcript parser (`response-viewer-transcript.ts`) reconstructs the conversation from the pane text instead, and the `?context=full` view now tags every block with a role so prompts render as "You" and agent output as the assistant. The divider normalizer was rewritten as a linear scan after review found catastrophic backtracking on agent-controlled input (minutes of stall on a long dash run), with an equivalence corpus pinning the old accept set.
CLIs installed via nvm or Homebrew are now found when Codeman runs as a service (#329). A shared resolver falls back to a login-shell probe when the direct PATH lookup misses, so systemd and LaunchAgent installs no longer report every CLI as missing. Review hardening on top: a failed resolution is negative-cached with doubling backoff instead of re-spawning a login shell on every request, all probes pass `killSignal: 'SIGKILL'` (interactive bash shrugs off SIGTERM, and a blocking `.bash_profile` could have hung the server indefinitely), the resolvers are inert under vitest again so test suites cannot execute binaries found on the dev box, and the improved not-found guidance is wired into both the session-create errors and the per-CLI status endpoints.
`GET /api/system/repo-status` reports branch, upstream, ahead/behind and remote reachability for git-clone installs (#328). Review hardening: the git network calls moved off the synchronous path onto a single-flight 45s cache (one slow remote could previously freeze the whole server for up to a minute per request), remote URLs and git stderr are credential-redacted before they leave the server, the spawns use the same non-interactive git env as the clone path, and a local-branch upstream no longer parses into garbage.
Auto Copy for the terminal (#325, opt-in, per-device): a finished selection (mouse drag, double or triple click, or a phone long-press) lands on the clipboard by itself, so select-then-copy becomes select. Alongside it, hand-encoded tap reports are now gated on the server-observed `cliMouseTracking` state, so a pane that has fallen back to a plain shell no longer receives `[<0;88;20M` junk on tap.
The Ralph loop no longer stops polling after two ticks (#330): the reschedule guard read a stale timer handle that the timer callback never cleared, so the loop silently died while its status stayed `running`. The handle is now nulled as the callback's first statement, and a regression test pins the bug.
The red "needs you" tab alert clears when a dialog is answered in the terminal instead of surviving until the end of the turn: the post-hook re-capture could erase the parsed dialog options that the staleness sweep relies on (`applyCapture` is now add-only for options), and a delayed staleness pass now runs while a page is open. The unreachable `copyTerminal()` was removed, closing out #322.
### Thanks
- @aakhter contributed the external-CLI response viewer (#326), the repo-status endpoint (#328), the login-shell CLI resolution (#329) and the Ralph reschedule fix (#330)
- @rounakdatta reported the mobile copy gap (#322) closed out in this release
## 1.19.7
### Patch Changes
- Mobile catches up: links open from a tap, terminal text can be selected and copied, long prompts stay visible while you type. Plus Files panel search, a bundled Nerd Font symbols fallback, and a per-device terminal font setting.
- **Terminal and chat links work on phones** (#321): tapping a URL or file path in terminal output now opens it (new tab, file preview, or log viewer), resolved through the same provider desktop hover uses, so tap and click can never disagree about what is a link. Dialog rows and the composer keep their existing meaning. Response-viewer links open in a new tab with `rel="noopener noreferrer"` instead of navigating the dashboard away. Wrapped links open whole: the logical-line reconstruction now stitches hard wraps through the indent their continuation carries, which also fixes desktop hover-click truncating wrapped URLs.
- **Terminal text can be copied on touch devices** (#321): long-press selects the token under the finger, drag or tap the other end to extend, and a small bar offers Copy, Line (the whole logical line, wraps included) and dismiss. Copy works on plain-HTTP installs too. Three guards keep the keyboard down and the selection alive through the browser's own long-press handling.
- **A long prompt stays visible on phones** (#321): the local-echo overlay grows upward once it would run past the last visible row (a prompt taller than the screen keeps its tail, where the cursor is), and the keyboard-driven padding shrink can no longer reclaim the space the fixed toolbar and accessory bar stand in.
- **Files panel search** (#324): `GET /api/sessions/:id/files?q=...` answers a flat match list (name or path substring, `*`/`?` globs), recursing past non-matching directories with its own match cap on top of the existing bounds; without `q` the response is byte-identical to before. Glob queries are matched without regex so a pathological pattern cannot stall the server.
- **Nerd Font prompt glyphs out of the box, custom terminal font** (#320): a bundled icons-only Symbols Nerd Font Mono fallback renders powerlevel10k/starship/oh-my-posh glyphs on every device with no font install, and App Settings gains a per-device terminal font family that is prepended to the built-in stack.
### Thanks
Three contributor PRs in one release: thanks to @rounakdatta (#321), @aakhter (#324) and @comzine (#320).
- 8a54b33: Clear every production-reachable npm advisory, and fix a service-worker caching regression the upgrade exposed.
`npm audit` reported 20 advisories, but 16 were devDependencies-only (Remotion, Puppeteer, postcss, the eslint/tsx toolchain) and never reached anyone installing the package. Four reached production and are now resolved:
- **`@fastify/static` 9.1.3 to 10.1.3** — GHSA-8pvw-jcv7-9cmj, authorization bypass via non-canonical URL paths. The advisory covers `<=10.1.1`, so the entire 9.x line is affected and the fix only exists on 10.x.
- **`find-my-way` 9.6.0 to 9.8.0** — GHSA-c96f-x56v-gq3h (HTTP/2 DDoS). Not exploitable here since Codeman does not enable HTTP/2, fixed anyway.
- **`fast-uri` 3.1.2 to 3.1.5** — GHSA-v2hh-gcrm-f6hx, host confusion via a literal backslash authority delimiter.
- **`brace-expansion` to 5.0.9 / 1.1.18** — GHSA-3jxr-9vmj-r5cp, exponential-time expansion DoS.
The last three were transitive and only needed a lockfile re-resolve; no `overrides` were added.
The `@fastify/static` major changes the `setHeaders` callback's first argument from a Node `ServerResponse` to a `FastifyReply`, which required two fixes:
- `res.setHeader()` became `reply.header()`. A v9-style body throws `TypeError: res.setHeader is not a function` from inside the plugin on every static request.
- **That change also flips precedence, silently.** The callback used to write to the raw response and be overwritten by the route's staged reply headers; it now writes to the reply and wins instead. That handed `/sw.js` a year of `immutable` in place of the `no-cache, no-store` its route sets, which would pin a service worker on every client with no server-side way to recover. A route that already set `Cache-Control` now keeps it.
`ws` also appears in `npm audit` but production is already on 8.21.0, outside the vulnerable range; the only affected copy is bundled under `@remotion/renderer` and is dev-only.
Adds `test/static-cache-headers.test.ts`, which drives a real server and covers the caching contract that had no test at all, and moves the floors in `test/dependency-security.test.ts` up to the patched versions.
## 1.19.6
### Patch Changes
- Wiki user manual, a phone tab tap-zone fix, per-parent lineage colours, and two robustness fixes.
- **Wiki**: `docs/wiki/` is now a 30-page user manual (installation, quick start, the dashboard, agent CLIs, remote/Docker cases, hooks, security, HTTP API, troubleshooting and more), published to the GitHub wiki by a sync workflow on every push that touches it.
- **Phone tabs**: on a narrow phone the active tab's geometric centre could land on its gear icon, so a thumb aiming at the tab opened Session Options instead of switching. The active tab's name now reserves a minimum width, and a static test recomputes the clearance from the stylesheet so widening the icons fails there rather than on a phone.
- **Lineage lines**: the arcs between a tab and the tabs it spawned are now coloured per SPAWNING tab, so every arc leaving one tab shares a colour and the strip reads as "these came from w1, those from w2". A child that spawns in turn gets its own colour, so a chain changes colour at each generation.
- **File access**: `validateSessionFilePath()` now canonicalizes the workspace as well as the candidate path before comparing them. Resolving only the candidate made a workspace reached through a symlink (`/tmp` on macOS, symlinked project dirs, bind-mounted case paths) report a spurious escape and refuse every read and write in that session. Escapes are still refused.
- **Respawn**: a cycle step that is stopped mid-write no longer revives the state machine. `stop()` could land during the `await` on the kickstart / update / clear / init write, after which the controller set itself back to a waiting state and kept running.
### Thanks
- @aakhter for the symlink-safe workspace confinement fix (#314) and the respawn stop-race fix (#315).
- 98e37bf: Session List Layout gains a third option, "Left sidebar", whose rows carry the same per-session detail the home screen shows.
The sidebar previously had one row style: a name and a folder. That is the whole story a tab can tell, but a docked column is not a tab strip — it has width to spare and a row per session either way, and the information that was missing is exactly the information the desktop home rail and the phone overview already put on screen. So the new option lifts it onto the rows: when the session was first created, how long it has been in the state it is in, and a status pill naming that state.
- The old "Left sidebar" is now **"Left sidebar simple"** and is unchanged, down to the byte — the stored value stays `sidebar`, so anyone already using it keeps exactly the layout they chose. The new option is `sidebar-rich`.
- Both sidebar values are the SAME layout and both set `data-session-list="sidebar"`; row detail rides on a separate `data-sidebar-detail` attribute. That is deliberate: every `isSessionSidebarActive()` call site and every `html[data-session-list="sidebar"]` rule in styles.css and mobile.css keeps matching both, untouched.
- Which state a session is in, and which stamp measures it, come from `_mobileOverviewState()` / `_mobileOverviewSince()` rather than being re-derived — the sidebar, the home rail and the phone overview cannot disagree about what "working" means. A working row is measured from the turn's last Enter, not from its last repaint, so a running turn reads `working 12m` instead of `0m`.
- The stamps refresh in place on a 20s clock instead of re-rendering: a rebuild would restart every load spinner and alert animation in the list, twice a minute. The clock only runs while rich rows are on screen.
- The column widens to 300px for the extra line, and the collapsed 44px rail and the handheld drawer are explicitly held back from that width.
- 947ff6f: `npm test` is now the CI gate and is safe to run bare; the suites it cannot run each got their own command.
`npm test` ran the everything-config, which fails ~87 tests on a clean master on any machine without chromium, a free port and per-machine PNG baselines. That made the repo's most obvious command useless as a pass/fail signal, and the docs had accumulated "never run bare `npm test`" warnings in four files to work around it. It now runs `config/vitest.ci.config.ts` — exactly what CI runs — so local green means CI green.
- New: `test:browser` (5 Playwright files), `test:perf` (2 wall-clock benchmarks), `test:all` (the old everything-behaviour, kept reachable). `test:ci` and `test:mobile` are unchanged; `test:watch` and `test:coverage` follow `test` onto the gate's config.
- The exclusion list moved to `config/test-suites.ts`, with the reason each suite cannot run in CI. Every config derives from it, so the gate's excludes and the runners' includes cannot drift.
- That drift was a silent hole, not a tidiness problem: a file excluded from CI and added to no runner is tested by NOTHING, and every command stays green, because vitest counts "no files matched" as success. `test/test-suite-partition.test.ts` now fails if any test file is reachable by no runner or by two.
- ⚠️ A file filter must match its runner: `npm test -- test/mobile/keyboard.test.ts` matches nothing and exits green having run zero tests, because the gate excludes that path. Use `npm run test:mobile -- <file>`. Documented in CLAUDE.md, and the one place that recommended the old form was corrected.
- Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, both READMEs, and two ci.yml comments that claimed only `test/mobile/**` was excluded (it is three suites, and 5 Playwright files rather than 3).
## 1.19.5
### Patch Changes
- Closing the session you are looking at now always moves you to the next tab.
The delete request and its own `session_deleted` broadcast raced each other: the close path selected the next tab, while the broadcast handler cleared the active session and showed the home screen, and whichever ran first decided what you saw. On one build, closing a tab either switched sessions or dumped you on the welcome screen depending on timing. The close now owns that handoff from beginning to end, and the broadcast handler stays out of the way for a close started in that tab. A session deleted from somewhere else still returns you to the home screen, which is the honest answer when what you were looking at was taken away.
The next tab is also picked from sessions that still exist, so a stale entry in the tab order can no longer name a tab that is already gone.
## 1.19.4
### Patch Changes
- Only a human opening a session clears its yellow "waiting for input" tab alert.
1.19.2 made that clear durable and cross-device, which also meant the app itself could spend it: restoring your last session on page load, a popped-out window opening its target, and the fallback to another tab after you close the active one all counted as "I checked it", so a yellow tab could clear itself before you ever saw it. Those three app-driven selections are now marked and skip the acknowledgement, so the alert survives until you actually open the session.
Everything a human does still clears it, on every surface: tapping a tab, tapping a row on the phone home screen, the keyboard tab shortcuts, and submitting a prompt into the session. The flag defaults to user-initiated, so a selection path nobody marked keeps acknowledging rather than leaving an alert nothing can clear.
## 1.19.3
### Patch Changes
- Red "needs you" tab alerts now follow the dialog instead of the keyboard.
Typing in the terminal no longer clears a red alert. It used to clear every pending alert on the device you typed on, but a permission or question dialog ignores keystrokes that are not one of its options, so the dialog was still open and still blocking: the other devices stayed red and a reload brought the red back on the first one. Input now spends the yellow idle alert only, and it does that through the server-side acknowledgement added in 1.19.2, so the clear is durable and reaches every device.
A dialog answered in the terminal now clears by itself. Claude Code fires no "permission answered" hook, so the item stayed pending until the whole turn ended, and any page load in between re-armed a red alert for a dialog that was long gone. Listing approvals now re-captures the pane and resolves items whose dialog is no longer on screen, using the same conservative check the answer path already uses: only an item whose original frame parsed numbered options can be dropped this way, so an unreadable capture keeps the alert rather than losing a live one. Measured against a real AskUserQuestion dialog: the stale item cleared 5 seconds ahead of the stop hook that used to be the only signal, while a dialog still on screen survived 11 consecutive listings over 55 seconds untouched.
## 1.19.2
### Patch Changes
- Yellow "waiting for input" tab alerts now stay cleared once you have checked them, on every device.
Viewing a session used to clear its idle alert in that browser's memory only. The server-side approval store still held the prompt, so the next page load seeded the alert straight back and a tab you had already checked went yellow again, while your other devices never heard about the click at all. Opening a session now acknowledges its pending idle prompt server-side (`POST /api/approvals/session/:sessionId/viewed`, a new `acknowledgedAt` field on approval items, broadcast as `approval:updated`), so the clear survives reloads and reaches every connected client.
Acknowledgement is deliberately not resolution: the prompt is still unanswered, so the item stays in the Approvals Inbox, stays answerable, and stays available as Read My Mind context, it just stops arming the tab alert. Permission and question dialogs are never acknowledged this way, since looking at a dialog does not answer it, so the red "needs you" alert survives being viewed. Clicking the tab you are already on now clears the alert as well; that path returned early before, so an alert armed on the active tab could not be cleared by clicking at all.
## 1.19.1
### Patch Changes
- Follow-up hardening from the 1.19.0 reviews, across all three of that release's areas (#309, #310, #311).
Home screens: the activity ordering introduced in 1.19.0 now stays truthful. Hook events push a session state broadcast, so a blocked session ranks by a fresh stamp instead of whatever the page loaded with; a working row with no recorded submit shows the same stamp it sorts by; Alt+1..9 resolves through the live sessions the tabs actually paint, so a stale id in the saved order can no longer shift every number off its target; and the "most recently quiet" ordering survives restarts, since recovery now restores each session's previous activity stamp from state.json instead of restamping everything at boot (previously every deploy flattened the ordering to tab order).
Files and sidebar: playable media extensions are pinned to the attachment registry by a parity test, so an in-workspace .m4a/.flac/.opus opens the preview player instead of the log viewer; /etc paths no longer render as links that can only 403; the sidebar session count counts the rows actually on screen (web tabs included, filtered rows excluded) and follows the filter box; connectors re-anchor on incremental renders in sidebar layout; and ~/.claude.json plus ~/.claude/settings(.local).json are blocked from file serving, home-anchored only, so case-level .claude files stay viewable.
Workspace hooks: the install-vs-refresh decision is one shared core that every claude create path routes through, so the workspaceHooksEnabled setting now also applies to cron jobs, legacy scheduled runs, and plan-orchestrator one-shots; a shell session in a docker case no longer authors a hooks block; the boot sweep no longer resurrects a deleted workspace as an empty directory; and the statusLine exporter got the same remote-attach and cwd-fallback guards as the hooks install.
## 1.19.0
### Minor Changes
- c01edcb: Add an optional collapsible left session sidebar as an alternative to the header tab strip.
With many concurrent sessions the horizontal strip wraps into several rows and stops being scannable. The new layout puts the session list in a vertical `<aside>` with a filter box and a live session count, collapsible to a 44px rail that keeps the status dots and task badges visible.
Opt-in via Settings → Layout → Tabs → Session List Layout; the default stays the header strip, so nothing changes unless you switch. Both layouts share one `#sessionTabs` element that is re-parented between mount points, so every existing affordance (status, mode badge, alerts, drag-reorder, keyboard navigation, web tabs, subagent windows) behaves identically in both. Below 1024px the sidebar is an off-canvas drawer that overlays the terminal instead of shrinking it. Collapse state persists per device; `Alt+B` toggles it.
- Codeman hooks now install into every claude workspace at session create, not just cases Codeman created (#304). Linked cases and cloned repos previously ran hook-blind: tab alerts, the Approvals Inbox, and the agent skill's stop/blocked wait signals were silently dead there. The install is an add-only merge that preserves user-authored hooks and leaves malformed files untouched, and a boot sweep heals sessions recovered from a restart. Opt out with the new synced `workspaceHooksEnabled` setting. Note: a `.claude/settings.local.json` can now appear in repos you link as cases; it contains no secrets. Remote SSH attaches and creates without a `workingDir` never write hooks.
File paths an agent prints are now clickable in both the terminal and the response viewer, opening the file preview overlay, including paths outside the session workspace (#306). Out-of-workspace paths are served through the attachment routes' extension allowlist, realpath confinement, and sensitive-path blocklist; Codeman's own credential-bearing files (`settings.json`, `push-keys.json`, `intents.json`, `state*.json`) are blocked from serving.
Both home screens (the desktop home tab rail and the phone overview) sort sessions by activity instead of tab order (#303): blocked sessions first with the longest-blocked on top, then running sessions longest-running first, then quiet sessions most recently active first. A turn starting now pushes a session state broadcast so the ordering stays live after page load.
The codeman agent skill docs teach hook presence as a setting to check rather than a consequence of who created the workspace, and the §0 preamble stamp is bumped to 1.19.0 (#305).
### Thanks
- @christianhaberl designed and built the collapsible left session sidebar (#307)
## 1.18.4
### Patch Changes
- Faster agent-skill workers, retuned multi-color lineage arcs, a per-tab pop-out option, reliable tab alerts, and the community launch.
- Agent skill: SKILL.md now forbids the standalone preamble check and the pre-spawn reconnaissance turns that were costing whole model turns; the same two-worker spawn measured at 28.6s end to end now runs 20.2s cold and 12.8s warm, with the spawn machinery itself unchanged.
- Session lineage lines: arcs now hang from the tab strip's bottom edge (dip cap 104px to 64px, no stacked row offsets), fixing the deep bow on wrapped tab strips and keeping same-row arcs off the second row's tab labels; each spawned worker's arc gets its own color (skin blue first, then matrix green, pink, violet, red, turquoise, orange), assigned per child and stable across re-renders.
- Session Options > Session: new "Pop-out button on this tab" per-tab override on top of the general App Settings toggle (per-device).
- Tab alerts: pending permission/question alerts now survive page reloads regardless of the Approvals Inbox setting (the alert state machine seeds from the server-side approval store on every load), stay visible on the selected tab until the prompt is actually resolved (the alert paints on a ::before overlay the active tab's styling cannot bury), and render as a steady red/yellow ring with glow and a colored status dot instead of a blink that spent half of every cycle looking like a normal tab. The README carries a live capture of the new alerts.
- Community launch: README Community section, .github/CONTRIBUTING.md (dev setup, test safety, great first contributions, PR expectations), and GitHub Discussions.
- docs: worker warm-pool design sketch with the measured baselines.
## 1.18.3
### Patch Changes
- Fix skill-spawned workers losing their lineage arcs and spawning slowly: a stale user-level agent skill copy (`~/.claude/skills/codeman`, written once by `codeman skill install`) shadowed the fresh per-case injections, so agents ran old recipes (serial spawns with pid polls, no `X-Codeman-Parent-Session` header). Session create now refreshes a marker-owned user-level copy (refresh-only, never installs, foreign/symlink copies untouched) and pre-seeds the skill's preamble into `${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh` (0600, local claude sessions only), single-sourced from the new `skills/codeman/preamble.sh` and pinned byte-identical to the SKILL.md heredoc by test. The skill's bootstrap is now a two-line loader with the full block as fallback, cutting measured prompt-to-workers-spawned time from 35s to 10.6s; `spawn_worker` also sends `parentSessionId` in the request body as defense in depth, and the preamble stamp is bumped to 1.18.3 so pre-fix cached preambles self-heal.
## 1.18.2
### Patch Changes
- Draw session lineage lines in blue for contrast. The violet arcs sat close to the
terminal's own dim foreground, so they lost contrast exactly where they cross text;
the colour now comes from each skin's own `--session-blue` token, and the layer is
separated from subagent lines by shape, weight and dash pattern rather than hue.
- f18097c: Make the `codeman` agent skill spawn workers fast instead of deliberating first.
Measured against a live server, the API does the whole job (spawn two claude workers,
task them, read both answers) in about 10 seconds, so the delay users saw was
agent-side: the skill taught serial spawning, made the happy path something to
reassemble from five sections on every run, and cost ~16k tokens of mostly failure
modes before the first call.
- The §0 preamble now defines the verbs instead of describing them: `spawn_worker`,
`spawn_workers` (concurrent), `sendwait` and `last_text`. §1 composes them into the
whole job in one Bash call, and says to stop reading there.
- Dropped two ceremonies the measurements retired: the pid-poll loop (`wait-output`
already blocks on the composer) and the agent-driven hooks check, which is now folded
into `spawn_worker` itself as a single local grep of the resolved `casePath`, so a
name that resolves to a linked case or a hook-less pre-existing directory is refused
instead of silently running the job there. Linked cases and raw paths still require
the by-hand check, where its absence silently breaks send-and-wait.
- The bootstrap's write condition now greps the version stamp, so a stale or truncated
preamble file self-heals instead of failing and asking you to `rm` it by hand.
- `sendwait` picks a fresh `seq` per call (a fixed default made every second prompt to
the same worker a silently-swallowed duplicate) and self-heals stranded delivery: an
Ink repaint occasionally eats the Enter, leaving the prompt typed but unsubmitted
(observed live), so a timed-out first wait sends one bare `\r` and re-waits by
resending the identical frame as a tagged duplicate.
- §5 moved to `reference/verbs.md`, leaving an index. SKILL.md is the only part paid on
every load and drops from ~16.4k to roughly 9k tokens (~35KB); section numbers and
anchors are unchanged, so existing `§5.x` references still resolve.
## 1.18.1
### Patch Changes
- Terminal history and scroll position fixes, a seekable file-viewer video player, and clearer session lineage lines.
**Terminal scroll position (#259).** Three paths dragged the terminal to the bottom while the user was reading scrollback. Opening or closing the mobile keyboard forced it unconditionally; scroll intent is now captured before the keyboard reflow and restored afterwards. Live writes preserved the viewport only inside a 1500ms window, so a user who scrolled up and then actually read for longer was dragged along by the next repaint; that is now based on position rather than recency. The backpressure refresh, which is server-triggered and so has no gesture to blame, now holds the reader's place too.
**Terminal history loss (#259 follow-on).** The backpressure refresh rebuilt the terminal from a 1MB tail, which measured as an 869-row buffer coming back with 158 rows: the routine meant to repair the display was discarding most of the scrollback every time SSE backpressure cleared. It now restores full history, falling back to the tail only when the capture would shrink the buffer, so repaint-mode panes are unaffected. It also bails if the user switches tabs mid-fetch, which would otherwise paint one session's history into another's terminal.
**History truncation is now visible and recoverable (#258).** Truncation was reported by a grey line written into the terminal, which scrolled away with the output it described and read the same whether the rest was one click away or gone forever. `GET /api/sessions/:id/terminal` now reports `truncationReason` (`tail` for an intentional partial replay whose remainder is still retained, `capped` for the byte ceiling) plus `retainedBytes`, and the browser shows a dismissible banner outside terminal output with three honest states: recoverable, which offers a Load full history button, at-ceiling, and exhausted. The button bypasses the scroll cooldown but not the downgrade guard, so it cannot destroy history on a repaint-mode pane.
**File viewer video (#284).** Closing the preview left the video playing with audible audio and no visible player, since hiding the overlay does not stop a media element and detaching one does not either. Media is now paused, unsourced and reloaded on close and on re-open, which also aborts the in-flight download. The scrub bar was inert because raw file bodies were served as a single `200` with no `Accept-Ranges`, so Chrome reported `video.seekable` as `[0, 0]` and Safari refused to start the media at all. Raw bodies are now streamed and range-aware (`Accept-Ranges` on every response, `206` with `Content-Range` for a range request, `416` past EOF, malformed specs ignored per RFC 9110), with pure, unit-tested parsing in `src/web/http-range.ts`. The attachments raw route gets the same treatment.
**Session lineage lines (#285).** The arcs joining a tab to the workers it spawned were tuned for two adjacent tabs and flattened into a straight thread across the terminal at the 800-1500px spans they are actually used at, drew a flat overprinted line inside the row gap on a wrapped strip, and were too faint to see at 1:1. Every pair now uses one U-bridge shape anchored on both tabs' bottom edges, with a deeper span-scaled dip and heavier, higher-contrast strokes.
**Docs.** The pi run mode is now listed in the mode lists that the sixth-backend sweep missed.
## 1.18.0
### Minor Changes
- Heal a stalled SSE stream with a heartbeat and a client-side staleness watchdog, and make a tab rename apply immediately.
An `EventSource` that stops delivering does not always error. A proxy that idle-closed the connection, a laptop resumed from sleep, a tailnet reconnect: `onerror` never fires, the header dot stays green, and every SSE-driven surface (tab status dots, sessions created on another device, renames) freezes until the user reloads. Nothing on the client tracked stream liveness at all.
- **`sse:heartbeat` is a new named event** under a new Transport category in the registry (155 constants now, both the backend list and the frontend `SSE_EVENTS` copy updated). The server already wrote a keepalive every 15s, but as an SSE `:keepalive` **comment**, and comments are invisible to `EventSource` by spec, so there was nothing a client could observe. `cleanupDeadClients()` now writes the named frame (`{"t":<epoch ms>}`) instead; interval, tunnel padding and dead-socket eviction are unchanged, and the write stays per-client rather than going through `broadcast()` because the frame carries no session data and so needs no multi-user owner routing.
- **Client watchdog.** `computeSseStale()` in `constants.js` is a pure policy beside `computeConnectionLossUi`: stale only when the transport believes it is `connected`, the device is online, and no frame has arrived for 45s (three missed heartbeats). That `connected`-only guard doubles as the loop breaker, since a forced reconnect leaves the state immediately and the watchdog cannot re-fire while one is in flight. The liveness stamp is applied inside `addListener` itself so every registered listener feeds it from one place instead of three that can drift, and the heartbeat's own listener is a deliberate no-op that exists only to be registered (`EventSource` drops named events nobody listens for). A 5s watchdog forces `connectSSE()`, `visibilitychange` to visible checks too (a background tab's timers are throttled, and a wake is exactly when a stream comes back zombie), and the forced reconnect logs one diagnostic line so a middlebox that strips or delays heartbeats does not present as an undebuggable "silently reconnects every 45s".
- **Renaming a tab appeared to do nothing** until a full page reload. The `PUT` always succeeded; what was broken is how the tab strip learned the result. `finishRename()` re-renders from the client-side `app.sessions` map and nothing wrote the new name into it, so the rename depended on the `session:updated` SSE frame to carry its own write back, which is precisely what a quiet stream never delivers. `_applyLocalSessionName()` now writes the confirmed name locally and refreshes cached subagent parent names. A rejected rename also used to read as success and silently drop the edit, because `_apiPut` turns a network error into a null Response so the old `try`/`catch` could never fire; a failure now restores the old label and toasts.
Tests: `test/sse-staleness.test.ts` (node VM over `constants.js`, threshold boundaries and every not-stale guard), `test/sse-heartbeat.test.ts` (drives `cleanupDeadClients()` with fake replies: named frame not a comment, parseable payload, padding only with a tunnel, dead clients still evicted), and `test/inline-rename.test.ts` (the name applies with no SSE frame dispatched, and a 500 leaves the map untouched).
Event names are part of the stable `/api/v1` contract, so this is a minor bump.
- c5b5963: Add Pi (pi.dev) as a sixth CLI run mode (#206).
`SessionMode` gains `'pi'`, a first-class backend alongside Claude Code, OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose tab identity, welcome button, run-mode entry, cron `agentType`, Docker and remote-SSH command defaults, and clone-repo Brain option.
- **New resolver** `src/utils/pi-cli-resolver.ts`. Unlike the sibling resolvers it sanity-probes `pi --version` and requires semver-shaped output, because `pi` is a short generic name that a stray binary on `$PATH` can shadow; the rejected path is logged. `GET /api/pi/status` returns `{ available, path, version }` so a misresolution is diagnosable.
- **`PiConfig`** maps to `--model` (accepts `provider/id` and a `:thinking` suffix), `--provider`, `--thinking`, `--session`/`-c`, and the tri-state `--approve` / `--no-approve`. Every value is regex-allowlisted and dropped on failure. `--api-key` is deliberately never wired: it would put a provider secret on the spawn command line.
- **No bypass flag.** Pi has no permission prompts and no sandbox, so there is no `--dangerously-skip-permissions` analog. Its privilege-shaped knob is `approveProjectTrust`, which makes pi load and execute repo-local `.pi/extensions` TypeScript and install missing project packages. `clampExternalCliBypassForOwner()` therefore puts pi in the **materialize** branch: a non-granted multi-user owner gets `--no-approve` even when no config was sent, because pi's own default is an interactive prompt the session user could answer themselves. The same materialization applies to cron-fired jobs (`clampCronExternalCliConfigs`), which carry no per-CLI config and would otherwise launch on pi's own default. Both helpers had no test coverage at all; they now do, for every CLI.
- **Env allowlist gains only the `PI_*` prefix.** Pi's ~34 provider key vars share no prefix and `ALLOWED_ENV_PREFIXES` is one global list with no mode context, so admitting them would widen the allowlist for every mode at once. Users authenticate via pi's `/login` or the server process's own environment.
- **Pi stays out of `isAltScreenStripMode()`.** Its default TUI renders into the main screen with terminal-owned scrollback and is mouse-aware, so it consumes `\x1b[3J` and the mouse DECSETs that the full strip removes, unlike an Ink TUI repainting in place. Note what exclusion does NOT do: pi is tmux-backed, so it still falls through to the narrow `isMuxAltScreenOnlyStripMode()` strip and its alt-screen toggles are dropped either way. Pi's runtime-switchable fullscreen TUI therefore paints into the main buffer, exactly like vim inside a tmux `shell` session.
- **Docker**: pi installs in its own `--ignore-scripts` step so that flag cannot affect the other four CLIs, and its credentials are seeded per-file (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json`) rather than whole-dir, since `~/.pi/agent` also holds sessions, extensions and installed package trees.
- **Local echo**: pi lands on the buffer overlay. Verified that codex's per-keystroke starvation does not reproduce: pi's slash picker re-filters on the whole composer content, so a one-shot flush behaves identically to per-keystroke typing.
- **Mode-list parity**: pi is excluded from the Ralph tracker auto-enable on `POST /api/sessions/:id/interactive` (like every other external CLI, whose output the tracker never parses), carries a `REMOTE_CLI_BIN` entry so a remote-SSH pi session reports its CLI version, and gets its own badge in the desktop home rail instead of rendering like Claude. The packaged agent skill's mode enumerations list pi too, and it now documents the per-CLI availability probes (`GET /api/<mode>/status`) that agents should check before spawning a worker on a backend the server may not have installed. Both are pinned by a new guard that derives the mode set from the Zod schema instead of restating it.
- **`codeman doctor` and the run mode agree about pi.** The registry entry resolved a bare `which pi` while `pi-cli-resolver` demanded semver output, so the Dependencies panel could report an installed Pi CLI that sessions refuse to launch. Both now share one exported regex, and the registry's new `requireVersionMatch` reports a non-semver `pi` as missing rather than installed. Only pi sets it; every other tool keeps its existing behaviour.
- Installer detection, docs (`docs/pi-integration.md`), READMEs, and the architecture invariants are updated. Tests: `test/pi-mode.test.ts` and `test/routes/external-cli-bypass-clamp.test.ts`, plus extensions to the run-mode, mobile-overview, render-index-html, system-routes and local-echo suites.
## 1.17.0
### Minor Changes
- Agent skill rework, session lineage lines, and a sharper endpoint drift guard.
**The packaged agent skill is rewritten around learning it, not just being correct** (`skills/codeman/`, ~2000 lines changed across four files). It previously opened with about fifty lines of credential archaeology before a single working call, and interleaved every recipe with the rationale for its own warnings.
- `SKILL.md` is restructured into: a 12-line "Hello, worker" that runs as written, a verb table an agent can act correctly from without reading anything else, a ten-line rules digest, the safety rules, the recipes, and setup/credentials last.
- **The preamble is no longer re-pasted.** A bootstrap writes it once to a `$HOME`-derived 0600 file and later calls source it and check a version stamp. Shell state does not survive between tool calls, but the filesystem does. The stamp is the last line written, so a truncated file leaves it unset and the guard aborts instead of running a half-written preamble.
- **New: where to spawn.** The only documented spawn used to create a scratch case, so "spin up workers on this repo" led an agent to do correct-looking work in the wrong directory. The rule is now explicit: hooks (and therefore `stop`/`blocked`) exist only where Codeman created the directory, so a linked case or a raw `workingDir` must synchronize on output markers. `wait:true` is still accepted there and silently degrades to a heuristic `idle`, which is documented as its own trap.
- **New verbs**: interrupt a runaway worker with ESC instead of deleting it, `active-tools` and `run-summary` as structured liveness signals, `auto-resume` for usage limits, the workspace as a high-bandwidth channel, and `GET /api/events` as a fleet watcher.
- `reference/messaging.md` gains a fleet protocol for Claude Code cross-session messaging: peer refs are injected and never discovered (a worker calling `ListAgents` sees the user's real sessions), every message costs a billed turn in both sessions, plus review pairs, mid-task questions, relay chains, mixed fleets, and their failure modes.
- `reference/recipes.md` is renumbered to a flat Flow 1-7 and gains Flow 7, one whole job start to finish: worktree fleet, tasks, gather, a review pass, report, cleanup.
- `reference/endpoints.md` gains an auth section, a symptom gallery keyed on what you actually see in the JSON, and a consolidated limits table.
- **Corrections found by auditing the old text against source**: the input cap is 65536 characters and not 100000 (65537-100000 passes Zod then 400s at the route); `wait.ended` is returned by a _live_ session whose write did not land, so "the session is gone" was wrong recovery advice and `delivered:false` is the discriminator; `DELETE /api/subagents` clears the map rather than killing anything; the trust-dialog auto-accept reads the rendered pane, not the output stream; `claudeMode` is readable globally though not per session; `run-summary` is envelope-wrapped (`.data.summary`); `active-tools` is not empty for `shell` mode; and a session does inherit the server's `CODEMAN_PASSWORD`.
**Session lineage lines** (`sessionLineageLines`, per-device, desktop default on). A create request may name the session that spawned it, as a `parentSessionId` body field on `POST /api/sessions` and `POST /api/quick-start`, or as an `X-Codeman-Parent-Session` header, and the web UI draws an arc from the parent's tab to each child's. The skill's preamble sets the header once, so every spawn recipe carries it. The value is **resolved rather than trusted**: exact id or a unique prefix of at least eight characters (ids reach agents truncated), it must be a live session the caller can see with the same owner, and anything unresolvable is dropped rather than returning a 400, so a cosmetic field can never fail a worker spawn. It confers no permission and no lifecycle meaning. Rendering is an additional layer on the existing connection-line pass, sharing one batched reflow; desktop only, because the mobile header would bury the overlay.
**The endpoint drift guard now covers routes it silently could not see.** `test/agent-skill-endpoints-doc.test.ts` matched only bare `app.<method>('path')` registrations under `src/web/routes/`, so routes registered on the server itself (`/api/events`, `/api/events/subscribe`) and any registered with Fastify generics (the approvals routes) were unverifiable. It now scans `server.ts` too and tolerates generics, taking it from about 200 to 216 recognized routes.
## 1.16.6
### Patch Changes
- Phone home screen now shows session ages, plus three mobile input fixes.
**Phone overview: started / how long stamps.** Every live session row on the "C" home screen carries a third line: when the session first started, and how long it has been in the state it is in ("started 3d ago · idle 12m"). Idle, waiting, error and ended states measure from the pane's last output, which for a Claude pane sitting at its composer is exactly when the turn ended; a WORKING session measures from its last Enter instead, because a running pane repaints about once a second and would otherwise report every turn as 0m. A 20s clock rewrites the values in place rather than re-rendering, so no row's blink or pulse restarts.
**Fix: a recovered session was restamped as new on every restart.** Boot recovery never passed `createdAt`, so each server start reset it to `Date.now()` and a week-old pane reported "created 2m ago" (and sorted as the newest thing in the unified session list). It now comes from the tmux session's own birth time, which mux-sessions.json already carried. The desktop home rail's "created" stamp is fixed by the same change.
**Fix: a selection dialog locked the on-screen keyboard out of the terminal (regression in 1.16.5).** The check that decides whether a tap belongs to the TUI scanned the whole viewport for a numbered menu, so while a Claude question or permission dialog was on screen EVERY tap in the terminal counted as actionable and blurred the input. The keyboard could not be opened at all until the dialog was answered, which left tapping an option, the one gesture that commits an answer, as the only interaction a phone had. The menu test is now row-local: the dialog's own rows still report the tap and keep the keyboard down, while the question title, the transcript and blank space summon the keyboard so a digit can be typed at the dialog instead of aimed at it.
**Fix: the accessory bar's arrow keys bypassed the local-echo overlay.** On a phone the text you type is buffered in the browser and has never reached the PTY, so an arrow tapped on the bar arrived at a composer the CLI still considered empty: Up recalled a history entry into it while the overlay went on painting the draft over the same row and still believed it was pending, and the next Enter submitted the two mixed together. The four arrows now flush the draft first and hand the session to plain PTY echo, the same contract a nav key typed on a hardware keyboard has had since #218. The CLI stashes the flushed draft, so Down brings it back. Tab now shares that one flush helper instead of its own copy.
## 1.16.5
### Patch Changes
- Mobile keyboard dismissal, and a tidier Save/Close pair in the phone settings sheet.
**The on-screen keyboard can finally be closed from inside the app.** The terminal
keeps focus on a hidden textarea and nothing ever released it, so once the keyboard
was up it covered roughly half the screen with no way out but the OS back gesture.
Two gestures now dismiss it:
- **A tap outside the terminal** (header, tab strip, empty page chrome). Deliberately
narrow: it only fires while the terminal input actually holds focus, never inside
the terminal (tap classification owns that decision), and never on a control, since
anything focusable is about to take focus itself and the keyboard accessory bar
exists to be used _while_ the keyboard is open. A scroll ends in `touchend` too, so
finger travel is tracked from `touchstart` and only a near-stationary gesture counts
as a tap, sharing the terminal's own 8px threshold so both agree on tap-vs-scroll.
Scrolling to read something mid-compose no longer drops the composer.
- **A second tap on inert transcript content.** Every terminal tap used to re-focus,
which left the accessory bar's chevron as the only way out. Scoped to inert rows on
purpose: the prompt row keeps focus-then-position, so a second tap there still
places the caret, and actionable rows (readbacks, `esc to interrupt` status rows,
menu selections) still blur as before.
**Settings sheet header on phones.** Below 860px Save moves into the header, which
left the two ways out of the sheet as a fat accent pill beside a bare glyph. Save and
Close now share a recessed tray with matching 36px pill geometry, reading as one
44px cluster the height of the phone header. Tray colors come from skin tokens, so
the light skins keep their look, and the tray stays off the sheets that carry a lone
close button.
Also fixes a test that could never have caught a regression: the case asserting that
tapping a control does _not_ dismiss the keyboard was picking a button from the
hidden welcome overlay, whose rect still measures while the hit-test lands on the
terminal underneath, so it passed for the wrong reason and stayed green even with the
exemption deleted. All four guards in the dismiss handler are now individually
pinned.
## 1.16.4
### Patch Changes
- **Voice dictation through your Claude Code login (no API key).** The mic button can now transcribe using this machine's existing Claude Code subscription, via the same speech-to-text service the CLI's own `/voice` mode uses. Off by default (`claudeVoiceEnabled`, synced): turning it on spends the server owner's Claude subscription on transcription for anyone who can reach the UI. The OAuth token never leaves the server process, credentials are read-only (Codeman never refreshes them, which would rotate the refresh token out from under the CLI), streams are capped at 5 minutes and 4 concurrent, and the WebSocket carries the same allowed-Host + same-site Origin guard as the terminal socket. A new Speech engine picker (Auto / Claude / Deepgram / Browser) sits alongside the existing Deepgram and Web Speech paths, which are untouched.
**One settings surface.** Session Options and Add Case now use the same `set-*` chrome as App Settings instead of the old modal-tab chrome, with a left rail, grouped rows, per-group device/synced scope badges and a search box. App Settings leads with version + update; the Session Options rail stays a real switcher (one section at a time) because Summary and Respawn are each long enough to bury the other. Collapsed Add Case blocks gained a disclosure chevron.
**Read My Mind: rethink steer note (phase 3 part 2).** Rethink now carries an optional free-text note ("no, I meant the mobile bug") sent as `steer`, the highest-authority signal the predictor gets. It stays in the field across re-runs, clears on each open, and the empty-result copy points at it. The modal footer moved to the styled `btn-toolbar` convention; the bare `btn btn-*` classes it shipped with match no CSS in this codebase and rendered as unstyled browser buttons.
**Mobile terminal taps no longer fight the keyboard.** Taps on TUI-owned rows (expandable readbacks, tool results, decision menus, the working/status row) now act on the CLI without popping the keyboard, while a tap on inert transcript text keeps the keyboard reachable. Rows are told apart by the affordance the CLI prints (`ctrl+r to expand`, `tap to collapse`, `esc to interrupt`) rather than by row titles, which vary per CLI and per version. A tap with the viewport scrolled up sends no mouse report at all but still restores focus, so the keyboard is reachable after every tab switch. Thanks to @Lint111.
**Path labels abbreviate `$HOME` on both platforms.** The "show `~/project`" rule had three implementations and two were platform-specific in opposite directions: the Run menu's matched `/home/<user>/` only, so on macOS every Recent Sessions row spent its first ~19 characters on an identical `/Users/<user>/` prefix and ellipsized away the tail that identifies it (#273); the case-manage list's matched `/Users/<user>` only, so no Linux case path was ever abbreviated. Both now route through one helper, with a static guard against a fourth copy appearing.
**Run menu Recent Sessions rows are legible.** Rows now read as folder, worktree pill, dimmed parent path, timestamp, with only the parent path allowed to shrink, so truncation can never hide which project (or which worktree) a row refers to. `<repo>/.claude/worktrees` is dropped from the parent path as noise. Thanks to @jordan8037310. Follow-up fix: the widened menu was not actually usable by its rows, since `.run-mode-history` is a block scroller and its `<button>` rows stayed shrink-to-fit at ~250px inside a full-window-width menu; rows now fill the menu and it is capped at the 760px one full row costs.
**Desktop home screen** no longer clips, and shows full tab names.
## 1.16.3
### Patch Changes
- Session rows that name their worktree, a shell keyboard bar for phones, App Settings as one scrolling document, and the Read My Mind modal on phones.
- **#265 / #266**: a past session whose directory no longer exists used to report
`$HOME` as its working directory, because history rows reconstructed a path by
stat-walking the filesystem and fell back to `$HOME` when nothing resolved.
Deleting a worktree is the normal end of its life, so every past worktree
session collapsed onto the same indistinguishable row. History rows now read
the literal `cwd` Claude Code stamps on its own records, out of buffers the
scanner had already loaded, so it costs no extra file reads and survives the
directory being removed. Sessions that ran in a worktree also carry a
`⑂ name · branch` pill in the Resume list and the Cmd+K session manager, and
both are searchable by worktree name and branch. Measured on a real install:
the cwd was recoverable for 215 of 216 transcripts, 212 of them from the first
16KB, and 28 rows that previously read `$HOME` now report their real path.
Reported and implemented by @jordan8037310.
- **#262**: a shell session now gets its own mobile accessory bar
(`Ctrl · Esc · Tab · ↑ · ↓ · ← · → · Paste · ⌄`), with Ctrl as a one-shot
modifier: tap it, and the next character goes out as its control byte. That
puts Ctrl+C/D/Z/R/L/A/E/W/U/K on a nine-button bar without a button per chord.
The modifier is applied on the CJK input path too, where the textarea owns the
keyboard and an armed modifier could previously neither fire nor be spent, so
it survived until a later keystroke and turned that one into a control byte.
Agent sessions keep the existing bar unchanged. Proposed by @DodgyBadger.
- **#257**: with several tabs open on a phone, the rightmost ones could not be
reached. Selecting a tab never scrolled the strip, and every ambient rebuild
reset `scrollLeft` to 0, so a strip the user had just swiped snapped back a
moment later. Reported by @DodgyBadger.
- **App Settings** is now a left rail acting as a table of contents over one
scrolling document instead of 8 tabs that wrapped onto two rows. Nine sections,
all mounted at once, so find-in-page works across the whole thing. The model
controls stop contradicting each other: the base model lives on cards and "1M
context window" is a switch that composes onto it, retiring the old pair of
settings that each claimed precedence over the other.
- **Read My Mind** suggestions beyond the first are no longer discarded. The
alternates render as tappable rows with their kind badge, tapping one swaps it
into the editable field without losing an in-progress edit, and Rethink now
records the whole shown set as rejected. The modal is sized for phones and
reachable from the phone keyboard bar.
- The desktop welcome screen carries the open tabs as a rail docked to the left
edge, with created and last-active stamps refreshed in place.
- The README now documents cloning a GitHub repository straight into a case
(**Add Case → Clone Repo**), which shipped in 1.16.2 but was only described in
the architecture docs.
- 5d42f64: Home screen: make the past-conversation list usable, and let search find past sessions.
- **#260**: "Resume Conversation" showed 4 rows and then dumped every remaining
one into a fixed 240px box, with no ordering or filtering. The list now opens
with 10 rows, "Show more"/"Show less" grows and shrinks the box itself (the
height cap is class-driven instead of fixed), and the header carries a filter
box (matches name, folder, `#case` label and the conversation's prompts), a
sort control (recent / name A–Z / folder A–Z, pinned rows still first) and a
shown-of-total count. Filtering implies expansion, so every match is visible.
- **#261**: the search box could not match a past project by folder name: its
session corpus was the live in-memory map, while past sessions come from
`/api/sessions/unified`. Search now also harvests a bounded snapshot of that
unified list, refreshed OUTSIDE the request path (published by
`/api/sessions/unified`, plus a fire-and-forget rebuild when stale), so the
search path keeps its no-filesystem-reads property. Results for a closed
session resume the conversation instead of trying to select a tab that no
longer exists, and are badged `RESUME`. In multi-user mode the snapshot is
re-scoped per row on read, matching what `/api/sessions/unified` exposes.
Reported by @jordan8037310.
## 1.16.2
### Patch Changes
- Clone a Git repository straight into a case, predict the prompt you were about to type, and point a session at a separate Claude account.
**Clone Repo (#251, proposed by @DodgyBadger in #236)**: Add Case gains a **Clone Repo** tab that clones a repository into `codeman-cases/<name>` and registers it as a normal local case. A live verdict under the URL field answers, while you type, whether the URL is cloneable without credentials, what its default branch is, and which branches and tags exist (`POST /api/cases/clone-preflight` behind `git ls-remote --symref`). The case name fills in from the parsed repo, refs come from the remote as a datalist, shallow clone is optional, and a Brain picker (installed CLIs only) points the Run button at the agent you chose. Starting a session stays opt-in, and the tab hides itself when the server has no `git`.
**Every settings writer now refuses to write through a symlink (from the #251 review, affects existing cases too)**: case contents can be foreign, and a repository can ship `.claude` or `.claude/settings.local.json` as a symlink pointing anywhere on this machine. Since `writeFile` follows links, a scaffold write could land outside the case, up to and including replacing your own `~/.claude/settings.json`. All seven writers that touch a case's `settings.local.json` (`writeHooksConfig`, `ensureCodemanHooks`, `refreshStaleCodemanHooks`, `updateCaseModel`, `updateCaseEnvVars`, `stripCaseEnvKeys`, `applyStatusLineConfig`) now go through one `withSafeSettingsWrite()` gate that runs the symlink check inside the per-path settings lock. A refusal is a warning rather than a throw, so hooks degrade to output-based idle detection instead of failing the operation. If you have deliberately symlinked a case's `.claude` or its `settings.local.json`, Codeman will now decline to write there and say so; replace the link with a real file or directory to get hooks, model and statusLine writes back.
The clone endpoint (`POST /api/cases/clone`) is synchronous by design: no job store, no polling, bounded by `GIT_CLONE_TIMEOUT_MS` (default 5 minutes). Security decisions live in a pure half of `src/git-clone.ts` so each is unit-testable without spawning anything: `<name>::<payload>` transports are refused as a family (any of them dispatches to a `git-remote-<name>` helper, which turns a clone into arbitrary command execution), a leading `-` is refused and `--` precedes every operand, argv arrays are used rather than a shell, URLs carrying credentials are refused, and non-interactive means more than `GIT_TERMINAL_PROMPT=0` (empty `GIT_ASKPASS`/`SSH_ASKPASS`, `SSH_ASKPASS_REQUIRE=never`, empty `DISPLAY`, `GCM_INTERACTIVE=never`, `ssh -oBatchMode=yes`), since with the request held open any one of those left open is a hang instead of an error. Timeouts signal the process group, because `git clone` fans out into `git-remote-https`/`index-pack` and SIGTERM to the parent alone can leave the fetch running. Repository contents beat scaffolding: an existing `CLAUDE.md` is kept, hooks merge into whatever `.claude/settings.local.json` the repo shipped, and a repo shipping its own `.claude/settings*` is reported back as a warning, because those hooks run locally as soon as a session starts.
**Read My Mind phase 2 (#256)**: phase 1 (1.16.1) gave each case an intent profile; this turns it into the feature as pitched. Press 🧠 on a Claude session and Codeman predicts the prompt you were about to type, from your stated goals, your recent prompts in your own voice, the last assistant reply, tool activity, git state, away context, sibling sessions, and any dialog the session is waiting on. The context assembler is pure and budgeted with trust tiers, so user-stated intent outranks observed content and terminal output alone can never justify a suggestion. One shot at opus (`readMyMindModel` overrides), a strict JSON contract, and 1 to 3 suggestions typed continue / verify / redirect. The modal keeps the suggestion editable: Send, Insert (drops it on the composer without Enter), Rethink (rejections feed back into the next attempt), Dismiss. Nothing is ever auto-sent, the click is the boundary. Opt-in via App Settings, Panels (synced, default OFF), desktop header only. Agents get the same verb through the Codeman skill (`POST /api/sessions/:id/readmymind`).
**Per-session `CLAUDE_CONFIG_DIR` (#255, designed and specified by @jordan8037310)**: `schemas.ts` gains an exact-key tier (`ALLOWED_ENV_KEYS`) beside `ALLOWED_ENV_PREFIXES`, admitting `CLAUDE_CONFIG_DIR` so a case can run on a separate Claude subscription (client-billed accounts). Exact match only: other `CLAUDE_*` keys and near misses like `CLAUDE_CONFIG_DIR_EXTRA` stay rejected, blocked keys stay blocked. The key survives `getEnvOverridesForPersist()` because it is a path rather than a secret, and dropping it would silently switch a rebuilt session back to the default account after a reboot. Caveat worth knowing: a relocated config dir writes transcripts outside `~/.claude/projects`, so the response viewer, subagent windows, ultracode panel and Read My Mind go blind for that session unless `projects` is symlinked back into the shared tree.
## 1.16.1
### Patch Changes
- 161f1da: Read My Mind phase 1: per-case intent profiles (docs/readmymind-plan.md). Codeman can now capture the prompts a user actually submits (from the Claude session transcript, opt-in via the new synced readMyMindEnabled setting, default OFF) into a per-case intent profile alongside user-stated goals, stored in ~/.codeman/intents.json (mode 0600, never searched). New endpoints GET/PUT/DELETE /api/sessions/:id/intent (ownership-scoped, strict schemas), a transcript:user_prompt event on TranscriptWatcher, and agent-skill coverage (SKILL.md recipe + endpoints.md rows) so agents can read and record the user's intent. Groundwork for the phase-2 predictor button: nothing is ever auto-sent.
- Home screen and phone touch targets.
The desktop welcome screen now lists your open tabs as a vertical column down its left gutter, which was previously dead space: one row per live session plus any saved web tabs, in tab order so the row badges match Alt+1..9, with case, backend and state on each row. Clicking a row enters that session. The column is width-gated (1180px and up) and never moves the centered welcome content.
Working state now reads the same everywhere it appears. A busy session shows a pulsing green dot ringed by the same spinner a tab draws while it loads, with a green halo, on the desktop home column, the phone home screen and the tab strip alike. Phone tabs got the bigger 9px glowing dot for the same reason.
Phone touch targets: the brand "C" that returns you to the home screen was roughly a 12x13px hit area, well under the 44px minimum. It is now a real 44x44 button, and the phone header grew from 36px to 44px to make that possible, which gives every other header control the same 8px. The simple keyboard accessory bar also swaps /clear for Tab (/clear and /compact stay in the extended bar), flushing locally buffered text to the terminal first so completion applies to what you just typed.
## 1.16.0
### Minor Changes
- Approvals Inbox, truthful idle detection, a revived trust-dialog auto-accept, and an unmistakable offline state.
**Approvals Inbox (#245, opt-in, default OFF)**: one cross-session inbox for every prompt that is waiting on a human (permission dialogs, AskUserQuestion questions, idle prompts). Enable "Approvals Inbox" in App Settings -> Panels (synced setting `approvalsInboxEnabled`); until then no new UI renders anywhere. Desktop gets a header bell (visible only while something is pending, with a count badge) opening a drawer of cards answerable in place: session, tool/message summary, the captured dialog frame, and one button per parsed dialog option (fallback: Approve / Deny-Esc). The phone overview's NEEDS YOU rows gain compact answer strips, and push notification action buttons were fixed along the way.
**Sessions no longer report idle while working (#246)**: every working Claude session flipped to `status: "idle"` about two seconds into its turn, and tabs, notifications, respawn and the phone overview all read that bad value. The `❯` prompt redraws throughout a turn, so readiness now requires a sustained repaint streak plus a capture-pane probe that recognizes the live working line (`✻ ... (Xs)`), and the UI shows a working state you can actually see.
**Workspace trust dialog auto-accept has been dead and now works (#249)**: a session started in a directory Claude had not seen before sat on the workspace-trust dialog until a human pressed Enter, because tmux delivers cursor-forward sequences rather than spaces. Detection now goes through the capture-pane text added in #246 and the dialog is answered reliably.
**A dead connection is unmistakable instead of a red dot (#248)**: the service worker serves the cached app shell, so opening Codeman with nothing reachable rendered a normal-looking empty dashboard with only an 8px red header dot as a clue. Now a connection-loss overlay (retry button, server host, actionable hints) plus a persistent banner make the state obvious on desktop and phone, and clear the moment the server answers again.
- 1e1db94: Cross-session messaging integration, two halves. **Workers now carry their Codeman session names as messaging peer names**: local claude spawns pass `--name <session name>` when the installed CLI is 2.1.224+ (the cross-session-messaging release). The gate is fail-closed, since an older claude aborts startup on an unknown option: an unknown or older version yields a spawn command byte-identical to before, the value is allowlist-sanitized before shell interpolation, and docker/remote spawns never carry the flag (their CLI is not the probed binary). Verified end to end on an isolated instance: the worker lists as its session name in `ListAgents`, and its replies arrive tagged `from-name="<session name>"`.
**The Codeman agent skill teaches cross-session messaging**: drive claude workers over `ListAgents`/`SendMessage` where available, map rows to Codeman sessions via the `tmux codeman-<id8>` column, deliver multi-line exactly-once task messages (including mid-turn steering), collect results as latched replies instead of polling, and fall back to the HTTP recipes whenever the feature is absent (version, feature flag, telemetry-disabling env vars, Docker/remote cases, non-claude modes). Adds `reference/messaging.md` (ships automatically, the installer enumerates `reference/*.md`), fan-out Flow 5 in `reference/recipes.md`, troubleshooting rows in `reference/endpoints.md`, and safety rules for the shared peer namespace (message only workers you created, no permission laundering in either direction). All mechanics verified live against claude-cli 2.1.226.
### Patch Changes
- c50bb02: The File Viewer can show hidden files and folders.
`GET /api/sessions/:id/files` has always accepted `showHidden=true`, but the panel
hardcoded `showHidden=false`, so dot-prefixed entries were unreachable from the
tree: no `.gitignore`, no `.github/`, no `.env.example`, and nothing under them.
Opening one meant guessing its path.
The panel header gains a `.*` toggle. It re-fetches rather than re-rendering the
cached tree, because the filtering happens server-side, and it keeps the expanded
directories so toggling does not collapse the tree you just navigated. The state
is per-device (its own `codeman:fileBrowserShowHidden` key rather than the
app-settings object, which is rebuilt from the settings-modal DOM on save and
would drop a key toggled from outside it), defaults to OFF, and survives a reload.
Generated and version-control directories (`.git`, `node_modules`, `.next`,
`.venv`, ...) stay excluded either way: that list is about tree size, not about
hiding dotfiles.
Closes #221.
- ce22c2a: The filesystem path picker can show hidden files and folders, and the shared secret blocklist grew to make that safe.
The picker behind Link Existing's "Browse" and the mobile keyboard's `Path` key
refused every path with a dot-prefixed segment, so `.github/workflows/ci.yml`
could not be selected and a hidden folder could not even be opened. It now has
the same `.*` toggle as the File Viewer, default OFF, per-device, and it applies
to both the listing and the preview endpoint (which re-resolves the path
independently).
That filter was quietly doing security work. With every hidden path unreachable,
`isSensitivePath` never had to name the credentials that live in dot-directories,
because the picker's roots include Home. Lifting the filter removes that
accident, so the blocklist now covers them explicitly: SSH keys at any depth (not
only under `$HOME`), GPG keyrings, AWS/GCloud/Azure/Docker/Kubernetes
credentials, npm, Yarn, git, `gh`, netrc, PyPI, RubyGems, Cargo and Terraform
tokens, `.pgpass` and `.my.cnf`, and the Claude and Codeman agent credentials.
`~/.codeman/` and `~/.claude/` stay attachable as trees, since the publish skill
and the review-card loop read from them; only their secret-bearing members are
named.
Blocked trees, sensitive files, root confinement and symlink-escape checks are
all unchanged and still apply with the toggle on: a hidden entry that resolves
to a secret is dropped from the listing, and opening it is refused.
Follows #221.
## 1.15.0
### Minor Changes
- 55bff4a: Zero-lag predictive echo for Codex sessions (mosh-style write-through prediction).
Codex's per-keystroke composer forced 1.12.2 to disable the local-echo overlay (issues #218/#219/#220/#222), leaving Codex typing at full round-trip latency on remote links. This release adds a second echo mode instead of re-enabling the first: every keystroke still goes to the PTY exactly as before (byte-identical wire behavior, pinned by vm-level and end-to-end trace-equality tests), while the new `PredictiveEchoAddon` in `xterm-zerolag-input` 0.2.0 paints the predicted glyph at the predicted cell. When the real echo lands, the prediction is confirmed and its span removed (an invisible swap); mispredictions self-heal via a two-pass mismatch cascade and a TTL.
- Reconciliation reads the parsed terminal buffer, never the raw stream: full-line redraws, ECH gap painting and tmux's in-place deltas all converge to the same cells. Confirmation requires the cell match PLUS a cursor advance, so placeholder glyphs and identical repaints never false-confirm; blank cells are neutral (codex clears its placeholder on the first echo).
- Predictions paint only while the cursor sits on the measured Codex composer row (`/^› /`, codex-cli 0.147): trust/approval modals and wrapped continuation rows get no ghosts, deliberately falling back to real echo.
- Ships as a SEPARATE `vendor/xterm-predictive-echo.js` bundle: the existing zerolag bundle is byte-identical (sha256-verified), and a missing or broken bundle degrades Codex to exact 1.12.2 behavior. The per-device `localEchoEnabled` toggle is the kill switch.
- Claude/Gemini/OpenCode/Antigravity keep buffer mode untouched; shell stays off.
- A post-build adversarial review added the anchor-hold rule: after an unpredicted wire edit (backspace into echoed text, cleared input, IME text commits) new predictions hold until the next parsed write, so a stale displayed cursor can never mis-anchor a run.
- Tests: 55 new package tests including replay suites driven by fixtures recorded from a real codex TUI through the production tmux+strip pipeline (`scripts/dev/record-codex-frames.mjs`) and a 500-iteration seeded fuzz; new vm policy/wire-neutrality suites; a 10-scenario Playwright E2E against real codex covering the #218/#219/#220/#222 retests, byte-identity, and a simulated 300ms-RTT run. The package test suite now runs in CI.
### Patch Changes
- Agent-skill hardening, plus a fix for the mobile browser suite.
## The Codeman agent skill
Twelve issues found by auditing the skill against a live instance, and fixing them meant measuring things rather than reasoning about them.
**Readiness now works in every permission mode.** The ladder matched `bypass`, which is the status bar of only ONE mode. Measured one pane per mode against claude-cli 2.1.226:
| how Codeman spawned it | statusline | `shift+tab` | `bypass` |
| ------------------------------------------ | ----------------------- | ----------- | -------- |
| `--dangerously-skip-permissions` (default) | `bypass permissions on` | yes | yes |
| `--permission-mode auto` | `auto mode on` | yes | no |
| `--allowedTools …` | `don't ask on` | yes | no |
| neither (`normal`) | `don't ask on` | yes | no |
| `--permission-mode plan` | `plan mode on` | yes | no |
Every mode ends `(shift+tab to cycle)`, and the `claudeMode` setting is not exposed on `GET /api/v1/sessions/:id`, so there was nothing to branch on. The ladder matches `shift+tab` now: universal, and space-free, which is what makes it survive the TUI stream. A non-default worker used to be reported broken after burning the full budget. ⚠️ The `+` means it only works through `--data-urlencode`; a hand-built query silently searches for `shift tab`.
**`.status` is documented as unreliable in both directions.** Measured on a live worker reading `idle` while mid-turn and actively producing output, with `lastActivityAt` equal to the moment of the call. A worker that dies inside its pane also reads `idle`. Synchronize on `stop` or an output marker; to judge from outside, sample `terminal?tail=` twice and compare.
**The self-delete guard is fail-closed.** Documented in 1.14.2; the reference files and every recipe now route through it consistently.
**Reads work on macOS.** The ANSI-strip pipelines used `sed 's/\x1b…'`, and BSD sed has no `\xHH` escape, so on macOS they silently stripped nothing and handed the agent raw ANSI.
**Injection is atomic and no longer silent.** `installAgentSkillInto()` wrote each file with a bare `writeFile`, so two sessions created concurrently in one repo could leave a reader observing a truncated SKILL.md; writes now go through temp+rename under the same lock every sibling mutator uses. And both server call sites discarded the outcome, so a `foreign` refusal (a user-authored skill is present) or a `symlink` refusal was invisible: turning the setting on, seeing nothing, and having no way to find out why. Refusals are logged now; injection stays best-effort and still cannot fail session creation.
**Reference corrections**: the `FORBIDDEN` 403 row and which auth responses are plain text rather than the JSON envelope, the input size cap, the undocumented `killMux` parameter on DELETE, and the fact that zero, negative and non-integer timeouts are rejected with a 400 rather than clamped.
**README.zh-CN.md taught a recipe that could not work**: its input example had no trailing `\r`, so Enter was never sent and the prompt sat unsubmitted, and its read step used `/output`, whose `textOutput` is always empty for interactive sessions. Its agent section is now in line with the English one. CLAUDE.md's single-line gotcha also gained the `\r` rule.
**Tests**: the `codeman skill install`/`uninstall` CLI had none, including the linked-case resolution shipped in 1.14.2; the `POST /api/sessions` injection call site was never exercised because the shared route mock hardcoded the gate off; and nothing guarded `reference/endpoints.md` against drifting from the routes it documents. All three covered now.
## Mobile browser suite
The suite drives a real browser against a server started from TypeScript source, so it serves `src/web/public`, while `npm run build` puts the xterm vendor bundles in `dist/web/public`. Without them every `/vendor/xterm*` request 404s, `Terminal` is never defined, and every test touching `app.terminal` dies on a null. A `pretest:mobile` step now prepares them.
Hardened after two review rounds, each defect reproduced: the freshness cache trusted mtime alone, so a bundle left without its alias tail (or truncated by an interrupted `npm install`) was reported "up to date" forever while the suite died on `LocalEchoOverlay is not defined`; it now verifies content and size, and repairs what an earlier run poisoned. Builds go to a temp file private to the run and rename into place, so a partial write can never be published and two concurrent runs cannot corrupt each other. Temps whose owning process is gone are reclaimed, and only those. Freshness tracks every input the bundle derives from, not just the entry, so editing a sibling of the addon no longer leaves the suite testing a stale overlay. `npx` runs with the repo as cwd, so it uses the pinned esbuild instead of fetching an unpinned one.
## 1.14.2
### Patch Changes
- Four reported bugs fixed, and the Codeman agent skill from 1.14.1 gets its first published build with the fixes below alongside it.
## The Codeman agent skill
Introduced in 1.14.1 and the headline of this line. `skills/codeman` is a Claude Code skill that lets an agent running **inside** a Codeman session drive the HTTP API: start worker sessions, send them prompts, block until they finish, read their answers and clean up. It ships in the npm package and self-gates, so outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act and costs unrelated sessions nothing.
### Installing it
```bash
codeman skill install # ~/.claude/skills/codeman, every new Claude Code session sees it
codeman skill install --case myproject # just that case; linked cases resolve by name too
codeman skill uninstall # reverses either one
```
Or turn on **App Settings > Agent Skill** (`agentSkillEnabled`, synced, default off) and Codeman injects the skill into each case when a Claude session is created there.
Installs are marker-owned: a `skills/codeman` that Codeman did not write is never touched, a stale managed copy is refreshed in place, and a symlinked skill directory is refused rather than written through. Re-run `codeman skill install` after upgrading to refresh the copy. Turning `agentSkillEnabled` back off does **not** remove already-injected copies, because a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` directory; remove them per case with `codeman skill uninstall --case <name>`.
### Using it
Ask for orchestration in plain language ("spin up three workers, have them lint, typecheck and test in parallel, then report back") and the skill supplies the guard, the safety rules and the recipes. The flow it runs:
1. **Guard.** Re-runs a preamble on every shell call that refuses outside `CODEMAN_MUX=1`, reads `CODEMAN_API_URL` and `CODEMAN_SESSION_ID`, recovers a password from the data dir `.env` or the install's service definition if one is set, and defines a fail-closed `delete_session`. It re-runs it every call because shell state does not survive between an agent's tool calls.
2. **Start a worker** with `POST /api/v1/quick-start` (`mode` is any of `claude`, `shell`, `opencode`, `codex`, `gemini`, `antigravity`), checking `.success` before reading `.data.sessionId`.
3. **Wait until it is really ready.** A new session reports `idle` before its CLI has spawned, and a brand-new case shows a trust dialog first, so the skill waits for the composer's own status bar and treats the dialog as a bounded fallback.
4. **Send and wait in one call**: `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input`. It registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer. For `claude` workers it resolves on the `stop` hook, usually within seconds.
5. **Read the answer** from `GET /api/v1/sessions/:id/last-response`, which returns clean transcript text rather than a screen scrape.
6. **Clean up** with `delete_session`, for ids it created and nothing else.
Hook-less modes (`shell` and the external CLIs) have no `stop` signal and coarse lifecycle transitions, so the skill synchronizes those with a unique split marker and `wait-output ... from=buffer`. Worked fan-out flows, the per-mode signal table, error codes and the Docker/remote caveats live in the skill's `reference/` files, loaded on demand.
### The rules it encodes
Each of these silently wastes a run, which is why they are written down: every input must end with `\r` or Enter is never sent; input is single-line; a wait timeout is HTTP 200 with `wait.timedOut`, not an error; `stop` and `blocked` are `claude`-only; signals are edge-triggered with no history, so never fire-and-forget N prompts and then gather signal-waits one by one; a typed command echoes into the output stream, so markers must be split; a full-screen TUI stream is space-less, so match single tokens; and `pid != null` proves startup, not life, so `wait?until=exit` is the death check.
## Bug fixes
- **Web tabs: long-running proxied requests were aborted after 30 seconds with no server log (#237).** The proxy wrapped each upstream fetch in a 30s `AbortSignal.timeout`, which bounds the entire exchange rather than the wait for response headers, so a dashboard endpoint doing model inference and any actively streaming response both died at 30s as a generic unlogged 502 that read as an intermittent network error. The timeout now bounds time-to-headers only and is cleared the moment headers arrive, with the default raised to 300s (`CODEMAN_WEBVIEW_TIMEOUT_MS`). Header timeouts are logged with a sanitized identity (method plus origin plus path, never the query string, which can carry the dashboard's tokens). A browser that navigates away mid-request now aborts the upstream fetch, guarded by `writableFinished` so a completed response never triggers it. The WebSocket handshake keeps its own 30s budget via the new `CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS`, since a handshake is connection establishment and waiting minutes on one only delays the browser's reconnect logic.
- **Web tabs: sandbox incompatibility with cookie-authenticated reverse proxies documented (#238).** `docs/web-tabs.md` now covers cookie auth in front of Codeman itself (Cloudflare Access and similar), where a sandboxed frame's asset and API requests carry no auth cookie, bounce to the login provider, and leave the embedded app apparently unstyled while trusted mode works. The Test button's result now states its own scope: it verifies server-to-upstream reachability, not how the page behaves in a sandboxed frame.
- **A described session tab now shows just the description (#232).** A session named `w2-foo-bar: some description` rendered both halves, so the generated id ate the width the chosen part needed. The tab shows the description alone, the `w<n>-<case>` id moves to the tooltip and stays in the session settings modal, and `aria-label` deliberately keeps the full name so screen readers still get the id. Undescribed tabs are unchanged. Right-click a tab to rename it inline. This also fixed a re-render loop: the incremental update compared against the full name, which a described tab never matched, so those tabs re-rendered on every pass.
- **`codeman status` now probes the running server (#230).** The command runs in its own fresh process and reported that process's always-stopped Ralph loop under a bare "Status:", which reads as "the server is down" while the service is running fine and agents are reachable. It now probes the real server (`CODEMAN_API_URL`, else https then http on the local port, overridable with `--url`) and reports reachability, version and live session state; any HTTP answer proves the server is up, including a 401 from a password-protected install. The Ralph loop keeps its own `codeman ralph status`. This complements `codeman web --status` from the daemon work: that answers "did I start a daemon", this answers "is a server running at all".
## 1.14.1
### Patch Changes
- The Codeman agent skill is now installable, so an agent running inside a Codeman session can drive the API without you pasting docs into its prompt. Plus six fixes to the packaged skill, each found by running it live against a real instance.
## What the skill is
`skills/codeman` is a Claude Code skill that teaches an agent inside a Codeman session how to start worker sessions, send them prompts, block until they finish, read their answers and clean up. It ships in the npm package. It self-gates: outside a Codeman session (`CODEMAN_MUX` unset) it refuses to act, so installing it globally costs unrelated sessions nothing.
## Installing it
Three ways, pick one:
```bash
codeman skill install # ~/.claude/skills/codeman, every new Claude Code session sees it
codeman skill install --case myproject # just that case; linked cases resolve by name too
codeman skill uninstall # reverses either one
```
Or turn on **App Settings > Agent Skill** (`agentSkillEnabled`, synced, default off) and Codeman injects the skill into each case when a Claude session is created there.
Installs are marker-owned: a `skills/codeman` that Codeman did not write is never touched, a stale managed copy is refreshed in place, and a symlinked skill directory is refused rather than written through. Re-run `codeman skill install` after upgrading Codeman to refresh the copy.
Note that turning `agentSkillEnabled` back off does **not** remove already-injected copies, because a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` directory. Remove them per case with `codeman skill uninstall --case <name>`.
## Using it
Once installed, just ask: "spin up three workers and have them lint, typecheck and test in parallel, then report back". The skill supplies the guard, the safety rules and the recipes. What it does under the hood:
**1. Guard.** Every Bash call re-runs a preamble that refuses outside `CODEMAN_MUX=1`, reads `CODEMAN_API_URL` and `CODEMAN_SESSION_ID`, recovers a password from the data dir `.env` or the install's service definition if one is set, and defines a fail-closed `delete_session`. It re-runs it every call because shell state does not survive between an agent's tool calls.
**2. Start a worker.**
```bash
Q=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" -H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}')
SID=$(jq -r 'if .success then .data.sessionId else empty end' <<<"$Q")
```
`mode` is any of `claude`, `shell`, `opencode`, `codex`, `gemini`, `antigravity`.
**3. Wait until it is actually ready.** A new session reports `idle` before its CLI has spawned, and a brand-new case shows a trust dialog first, so the skill waits for the composer's own status bar and treats the dialog as a bounded fallback.
**4. Send a prompt and wait for the turn to end.**
```bash
BODY=$(jq -n --arg p "$PROMPT" '{input:($p+"\r"),useMux:true,clientId:"codeman-agent-1",seq:1,wait:true,waitTimeout:60000}')
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" -H 'Content-Type: application/json' --data-binary "$BODY"
```
Send-and-wait registers the waiter before typing, which closes the race where a separate wait reports the previous turn's idle state as this turn's answer. For `claude` workers it resolves on the `stop` hook, typically within seconds.
**5. Read the answer.**
```bash
"${CURL[@]}" "$API/api/v1/sessions/$SID/last-response" | jq -r '.data.text'
```
**6. Clean up.** `delete_session "$SID"`, for ids you created and nothing else.
Hook-less modes (`shell` and the external CLIs) have no `stop` signal and coarse lifecycle transitions, so the skill synchronizes those with a unique split marker and `wait-output ... from=buffer` instead. Worked fan-out flows, the per-mode signal table, error codes and the Docker/remote caveats live in the skill's `reference/` files, loaded on demand.
## The rules that bite
The skill documents these because each one silently wastes a run:
- **Every input must end with `\r`** or Enter is never sent and the text sits unsubmitted on the worker's prompt. `delivered:true` means "written to the pane", not "submitted".
- **Input is single-line.** Newlines are stripped.
- **A wait timeout is HTTP 200** with `wait.timedOut:true`, not an error. Loop over short waits; timeouts clamp to [1s, 600s] and the applied value comes back as `wait.timeoutMs`.
- **`stop` and `blocked` are `claude`-only.** Requesting them elsewhere is a 400.
- **Signals are edge-triggered with no history.** One that fires while no waiter is registered is unobservable afterwards, so never fire-and-forget N prompts and then gather signal-waits worker by worker.
- **Your typed command echoes into the output stream**, so a marker that appears verbatim in the input line matches before the command runs. Split it.
- **A full-screen TUI stream is space-less**, so match a single space-free token, never a phrase.
- **`pid != null` proves startup, not life.** A worker that dies inside its pane keeps `status:"idle"` and a pid. `wait?until=exit` is the death check.
## Fixes to the packaged skill
- **The self-delete guard failed open.** The old `is_self "$SID" || curl -X DELETE ...` shape meant an undefined `is_self` exited 127, the `||` branch fired, and the agent deleted its own session with the one guard bypassed. That is reachable because shell state does not survive between tool calls, so a partially re-pasted preamble was enough. The DELETE now lives inside a fail-closed `delete_session`, which also refuses an empty id and refuses when `$SELF` is unset or too short to prove the target is not the caller.
- **`clientId` was built from `$$`.** The pid changes between tool calls, so the documented "resend the identical request" loop stopped being recognized as a duplicate and retyped the prompt, submitting the turn twice. It is a fixed literal now.
- **`GET /api/v1/sessions/:id/last-response` was undocumented.** It returns the agent's final message as clean transcript text; the terminal scrape the skill previously recommended returns a wall of TUI repaint noise with the answer buried in it. It is now the documented read path for `claude` and `codex`, with the terminal buffer demoted to diagnosis and hook-less modes. Because the transcript flush lags the `stop` signal, the recipes poll it instead of reading once.
- **`quick-start` responses were never checked for `.success`.** On failure `.data.sessionId` is absent, `jq -r` prints the string `null`, and the flow burned its full readiness budget against `/api/v1/sessions/null` before reporting jq noise instead of the cause.
- **`codeman skill install --case <name>` could not resolve a linked case.** It hardcoded `~/codeman-cases/<name>` while the server resolves through `linked-cases.json` first, so it failed with "Case not found" for a case the web UI handled fine.
- **Documentation corrections**: `SESSION_BUSY` on `quick-start` is the 50-session cap rather than the waiter cap; `caseName` resolves linked cases, so a generic name can land a worker in a real repo; and the claim that a toggle-off sweep exists was wrong, so the per-case `skill uninstall` cleanup is now stated in both the README and the code.
## Also in this release
- **Terminal**: the wheel is no longer forwarded to codex, which ignores SGR mouse reports.
## 1.14.0
### Minor Changes
- Daemon mode and service install, plus subagent hook hardening and terminal/idle-checker fixes.
**New: run Codeman in the background without a terminal (#239, closes #231)**
- `codeman web -d` starts the server detached: it survives closing the shell, logs to `~/.codeman/web.log`, records a pidfile, and only reports success after the server actually answers `/api/status` (a port clash or missing dependency can never read as a clean start). `codeman web --status` and `codeman web --stop` manage it; `--stop` verifies the pid still looks like a Codeman server before signalling, so a recycled pid is never SIGTERMed.
- `codeman service install` / `status` / `uninstall`: installs a systemd user unit (Linux) or LaunchAgent (macOS) so the server comes back after reboots. The unit carries the installing shell's PATH (launchd's default PATH finds neither an nvm/Homebrew `node` nor `tmux`/`claude`), never contains `CODEMAN_PASSWORD`, and uses the same instance-scoped unit names as `install.sh` and the self-updater so no second copy can end up supervised.
- Both refuse to start a second server on one data dir (pidfile check plus a live probe): two servers on the shared tmux socket would attach to each other's sessions.
- Why `-d` exists at all: `nohup` does not protect a Node process, Node re-arms SIGHUP even when it inherits "ignore", so `nohup codeman web &` still dies on HUP. The detached relaunch (setsid) removes the controlling terminal instead.
**Subagent background-work hooks (#233, thanks @Lint111)**
- The background Bash rewake helper now also watches the top-level parent transcript when the hook fires inside a subagent: Claude records a subagent's Bash result in its own `subagents/agent-*.jsonl` but queues the completion in the lead session transcript, so subagents previously never woke. It can also inline a `CODEMAN_RESULT_BEGIN/END` marked report (up to 64 KiB) from the task output file into the wake feedback.
- New SubagentStop guard: a subagent that still owns live Monitor or background Bash processes is kept working instead of publishing an intermediate progress line as its final report. Ownership is verified against live process descriptors on `tasks/<id>.output`, so stale transcript text alone never blocks, and the guard fails open on systems without `/proc`.
- Existing cases self-heal to the new hooks on next launch.
**AI idle checker: stderr kept out of the verdict (#234, thanks @Lint111)**
The `claude -p` verdict command no longer merges stderr into the verdict file, where CLI warnings could turn a valid verdict into a parse error. On failures, the first 200 chars of stderr are attached to the diagnostic instead.
**Terminal: large final batches drain fully (#235, thanks @Lint111)**
A render-scheduling flag was cleared after the flush instead of before it, so when a large batch left a remainder behind, the remainder stayed unrendered until unrelated output arrived. This looked like truncated responses or shell commands that never finish. The flush now reschedules itself until the queue is empty.
**Docs and tests**
- README documents daemon mode and service install.
- Unique test port for the daemon-control suite.
## 1.13.0
### Minor Changes
- Agent wait primitives, the Codeman agent skill, a fix for hooks dying silently on HTTPS installs, and the tab-strip UX improvements from the previous batch.
**Agent wait primitives (new API surface, the reason this is a minor).** Three bounded long-polls let an agent driving Codeman from a shell block instead of poll:
- `GET /api/v1/sessions/:id/wait` blocks until a lifecycle signal fires (`until=stop,idle,working,blocked,exit`, `fresh=1` to require a new transition).
- `GET /api/v1/sessions/:id/wait-output` blocks until a literal substring appears in the session's output (`match=`, `nocase=`, `from=now|buffer`; never regex, by design).
- `wait`/`waitTimeout` on `POST /api/v1/sessions/:id/input` (send-and-wait) registers the waiter before typing, closing the race where a separate wait reports the previous turn's idle state as this turn's answer.
Shared semantics: a timeout is HTTP 200 with `wait.timedOut: true` (callers loop over short waits; tunnels cut idle connections), timeouts are clamped to [1s, 600s] and echoed back as `wait.timeoutMs`, all three nest the result under `data.wait`, and `status`/`limitPaused` ride along. `stop`/`blocked` exist for `claude` mode only: requesting them explicitly elsewhere is a 400, the default set silently narrows and echoes what it waited on. Capacity caps (16 waiters per session, 128 process-wide) answer 409/429, waiter slots release on client hang-up, and shutdown resolves parked waiters instead of stranding them. Bounds are operator-tunable via `CODEMAN_WAIT_*` env vars.
Reliability details that came out of three verification rounds: a worker that dies inside its tmux pane is now detected at the mux layer (pane-death probe, ~750ms cache, a 3s watcher for waits already parked), so a corpse answers `exit` instead of `idle` and send-and-wait rolls back its dedup seq when the write went nowhere; output matching normalizes charset-designation escapes (a stock bash prompt's `ESC ( B` no longer breaks `match=tnode:`) and holds back partial escapes at chunk boundaries, so matches straddling PTY chunks are found.
**Codeman agent skill (`skills/codeman`).** A packaged skill that teaches an agent running inside a Codeman session to drive the API safely: guard preamble (refuses outside `CODEMAN_MUX=1`, resolves credentials from the data dir `.env` or the install's service definition), self-protection (`is_self` prefix check in both directions), readiness for claude workers (composer-first, trust dialog as bounded fallback), send-and-wait loops that cannot report a never-submitted prompt as success, marker-synchronized shell flows, fan-out patterns, and cleanup discipline. Ships in the npm package via the `files` entry.
**Hooks were dying silently on every HTTPS install (bug fix).** The generated hook curls lacked `-k`, so on `--https` installs (self-signed cert) every hook event (`stop`, `permission_prompt`, `elicitation_dialog`, `idle_prompt`, `teammate_idle`, `task_completed`) failed TLS verification and the failure was swallowed, taking respawn's definitive idle signals with it. Hooks are now generated with `curl -sk`, and a staleness detector regenerates the on-disk hook config of already-created cases the next time a session starts in them. Relatedly, `CODEMAN_API_URL` is no longer exported with a guessed `http://localhost:3000` fallback (wrong scheme on HTTPS installs); it is omitted unless the server has stamped the real URL, so in-session guards fail closed.
**Tab strip (from the previous batch, reported by christianhaberl):** action icons (kill/pop-out) now appear on the active tab only, middle-click closes a tab, tab hover uses a fixed width with a sliding title instead of resizing the strip, and the pop-out button is opt-in (default off).
**Docs.** `docs/api-reference.md` gained the full long-polling contract (signals by mode, readiness, what the matcher sees, response discriminators); `docs/extending-codeman.md` and the README carry verified copy-paste orchestration recipes; `docs/architecture-invariants.md` records the load-bearing ordering, liveness, and edge-triggered-signal invariants. Net +163 tests (4300 passing in the CI sweep).
## 1.12.2
### Patch Changes
- Codex input fixes: all four bugs reported by @DodgyBadger traced to one root cause (the zero-lag local-echo overlay buffering keystrokes until Enter, which starves codex's per-keystroke composer) and fixed in terminal-ui.js:
- Slash command picker never appeared in codex sessions (#222): the "/" sat in the overlay until Enter, so codex never saw it. Codex-mode sessions now use plain PTY echo (same branch as shell), so the picker pops and live-filters as you type.
- Arrow keys dead while typing, backspace dead after Ctrl+Backspace (#218): arrows were forwarded to a still-empty composer while typed text sat pending, and after a control-char flush the overlay swallowed every backspace. Codex bypasses the overlay entirely now; the shared overlay branch (claude/gemini/opencode) additionally flushes pending text on composer nav keys, then hands the session to pass-through until Enter/Ctrl+C, and forwards backspace instead of swallowing it when the overlay has no state.
- Pasting displaced the typed prompt (#219): bracketed pastes (xterm terminal.paste with DECSET 2004 active) were forwarded without flushing pending typed text, so the paste landed first. The shared branch now flushes typed text first and delays the paste sequence by 80ms, because codex's paste-burst handling drops keystrokes that arrive in the same PTY read as a bracketed paste (verified against codex 0.147.0 at the byte level).
- Long prompts overflowed the bottom of the screen (#220): long typed prompts existed only in the overlay DOM so codex never grew its composer; with plain PTY echo the composer grows and rewraps normally.
Verified end to end against a real codex 0.147.0 TUI driven by a headless browser: the pre-fix build reproduces all four bugs, the fixed build passes 17/17 assertions. New CI test file test/local-echo-codex-gating.test.ts (41 tests) pins the nav-key classifier, per-mode overlay gating, the flush helper, and pass-through routing. Known upstream limitation: Ctrl+Backspace deletes one character, not a word (xterm.js sends 0x08; word-delete needs kitty CSI-u encoding that xterm.js 6.0.0 cannot emit).
Mobile keyboard viewport settling fixes by @Lint111 (#229): coalesce keyboard viewport settling so rapid visualViewport resize events during keyboard show/hide no longer thrash the terminal fit, and only arm the settle logic on a real keyboard transition instead of every viewport resize.
## 1.12.1
### Patch Changes
- Terminal scrollback fixes, round 2 of issue #205. A Claude pane's local buffer is hollow (tmux keeps no history for a repaint-mode pane), and both retest reports traced back to that fact. The scroll-to-top full-history re-pull now refuses to rewrite the terminal when the capture holds less than the browser already does, so it can no longer delete history mid-scroll on iPhone (a refused session also re-fetches far less often). When wheel-forwarding is unavailable on a Claude session (version probe failed, CLI older than 2.1.187, or the "Wheel Scrolls Local History" opt-out) and there is no local scrollback to scroll, wheel and touch now page the CLI's own transcript via coalesced PageUp/PageDown instead of doing nothing. The `claude --version` probe no longer caches a failed run for the server's lifetime (one timed-out probe used to silently disable wheel-forwarding on every device until restart); failures retry with backoff. Every scroll gesture now logs a one-line `[scroll]` routing decision to the browser console for direct diagnosis, and the opt-out setting's tooltip explains that the paging fallback is Claude-only (Codex has none).
- 2e69e28: Bound the process-tree walk that could take a machine down.
`getChildPids` ran `pgrep -P <pid>` per node and recursed with no visited set, no
depth limit and no node cap. Across ~28 adopted tmux trees the fan-out exploded,
and because each `pgrep` blocks in the kernel while reading `/proc/<pid>/cgroup`
under WSL, none returned while the walk kept spawning more — ~13,000 `pgrep`
processes stuck in D-state out of ~39,000 total, load average above 13,000,
recoverable only by restarting WSL.
Now: one `ps` snapshot, breadth-first with a visited set, a depth cap and a node
cap, in a pure module (`proc-tree.ts`) that the regression tests exercise
directly. The snapshot is refreshed asynchronously, and the kill path forces a
fresh one so the SIGKILL escalation cannot re-read pre-SIGTERM state.
- ebfcac6: An input whose delivery fails can be retried instead of being lost for good.
Both input paths recorded the `(clientId, seq)` pair as applied and acknowledged
the frame _before_ knowing whether the write had landed — the POST route because
its mux write is fire-and-forget, the WebSocket handler because it ACKed
unconditionally. When the write then failed, the client dropped the frame from its
durable queue and the server rejected the retry as a duplicate: the reliable
delivery layer was guaranteeing exactly-once delivery of something that had never
been delivered.
The bookkeeping is now rolled back on failure and the WebSocket ACK withheld, so
the client redelivers. `Session.write()` reports whether it reached a PTY at all
instead of silently swallowing the data.
Response codes are unchanged: a session can legitimately have no PTY yet (created
but not started), so turning that into a failure status would be a contract change
of its own.
Note this does not remove the root cause: the POST still answers 200 before the
mux write is attempted, so a client that treats any 2xx as final still cannot
learn about that failure. Closing that would mean awaiting the tmux child in the
request path.
- 1a32e63: Routes that answer with `reply.raw.writeHead()` no longer drop the headers the
security hook set.
`writeHead` writes straight to the Node response and bypasses Fastify's header
store, so everything the `onRequest` hook granted was silently lost — including the
`Access-Control-Allow-Origin` it emits for localhost origins, and the
`X-Content-Type-Options` / `X-Frame-Options` / CSP headers. A localhost page could
therefore call every other `/api` endpoint cross-origin while its EventSource
failed CORS.
Affects `GET /api/events` and the three raw-writing routes in `file-routes.ts`
(`file-raw`, `tail-file`, `download`).
## 1.12.0
### Minor Changes
- Terminal scrollback overhaul (issue #205), fixing every reported scroll failure across shell and CLI sessions, desktop and mobile:
- Shell, OpenCode and Antigravity sessions finally have working scrollback: tmux's own client-side alternate-screen switch is stripped for tmux-backed sessions (narrow strip: alt-screen toggles only, keeping `clear`'s 3J and mouse DECSETs), so xterm stays in the normal buffer instead of a scrollback-less alt buffer where the wheel turned into shell history cycling and touch scrolling did nothing. Direct-PTY fallback sessions are untouched so fullscreen apps (vim/less/htop) keep the alt screen there.
- The wheel listener now runs in capture phase and owns the scroll: xterm's internal vscode-style viewport scroller consumed wheel events whenever local scrollback existed (and goes deaf entirely after a tab switch or replay resets the terminal), which silently killed wheel forwarding, made scrolling break after reload/tab switches, and let the CLI's input box scroll away. Local scrolling goes through buffer-level scrollLines and keeps working after resets; mouse-tracking apps and alternate-buffer sessions are passed through untouched.
- Wheel AND touch scrolling now forward to the CLI's own transcript for Codex and Claude 2.1.187+, at any scroll position (the viewport snaps home first), so the input box stays pinned on desktop and phones alike. Shift+wheel and the "Wheel scrolls local history" setting still pin local scrollback.
- Smooth scrolling: local wheel scrolling glides with an ease-out animation (fractional line accumulation, so slow trackpad drags track the finger instead of running ahead).
- Full tmux history on demand: the full-scrollback replay is now per session instead of once per page load, and scrolling up at the top of the buffer re-pulls the complete tmux history, recovering everything tmux's repaint bursts or tab switches removed from the browser's copy.
- Firefox wheel speed: wheel deltas are normalized by deltaMode (Firefox reports line units, previously read as pixels and slowed ~4x).
- Remote SSH Claude sessions now probe the CLI version over ssh (same connection options and login-shell wrapper as the real launch), so wheel forwarding works for them too instead of silently staying off.
Docs: scrollback analysis and fix plan recorded in docs/, architecture invariants updated (strip flavors, capture-phase wheel ownership, per-session full-history replay); docker agent-image rebuild warning and integration-guide link fixes from the preceding docs commits.
## 1.11.2
### Patch Changes
- Make Antigravity (`agy`) a first-class CLI everywhere, and stop presenting Gemini CLI as a consumer product now that it is enterprise-only.
Antigravity was already wired into the session layer, schemas, run-mode menu and remote/Docker command maps, but the surfaces around it were never updated. Gemini keeps full support; Antigravity now sits beside it.
Fixes:
- **Docker cases with `mode: 'antigravity'` were broken.** `docker/agent.Dockerfile` installs its CLIs from npm, and `agy` is not an npm package, so the binary was never in the image and the container died on command-not-found. It now gets its own installer step. The `--dir /usr/local/bin` flag is load-bearing: the installer's default `$HOME/.local/bin` resolves to root's home at build time and would be unreachable by the `agent` user the container runs as. Note the binary is roughly 190MB, making it the largest layer in the image, so rebuild with `node scripts/build-agent-image.mjs` when convenient.
- **Welcome screen** gained a "Run Antigravity" action, gated on `agy` being present like the other CLI buttons, styled with the same cyan identity as the toolbar run button and run-mode dot.
- **`install.sh`** now detects `agy` (search paths mirroring `antigravity-cli-resolver.ts`), counts it as a satisfying AI CLI so an Antigravity-only box is not told it has none, and recommends it instead of Gemini in the install hints.
Documentation corrections where it had become factually wrong: `architecture-invariants.md` described `isExternalCliMode()` as opencode/codex/gemini when the code has included antigravity for some time, said "all three modes", and omitted `ANTIGRAVITY_*` from the env-prefix allowlist row; the `agentType` enum in `cron-guide.md`, `SessionMode` in `cron-discovery.md`, and `RemoteCommandMode` in `remote-sessions.md` were all stale.
Also updated both READMEs (five CLIs, Gemini marked enterprise-only), the `antigravity` npm keyword, and comment drift in eight places. Test coverage added for the new welcome button.
Antigravity stores its state under `~/.gemini/antigravity-cli/` rather than a `~/.antigravity` directory, so the existing `.gemini` Docker credential seed already covers it. That is now recorded in a code comment so no dead configuration gets added later.
- b982c5d: Keep the brief Response Viewer output inside the same message card and Markdown wrapper used by the full conversation view, so opening the viewer without clicking More preserves the same readable formatting.
## 1.11.1
### Patch Changes
- fix(history): Past Sessions data quality, and gate the phone run picker on CLI availability
**Past Sessions data quality (#215).** Three bugs in the transcript scanner behind
the Cmd+K Session Manager and the phone overview's PAST SESSIONS list:
- Automated/SDK-driven transcripts (CI review bots and other tooling, which Claude
Code stamps with a non-`cli` `entrypoint`) were listed alongside real interactive
sessions even though they were never resumable. They are now excluded. Detection
scans every entrypoint-bearing message rather than stopping at the first, so a
transcript that began under an older Claude Code build and only later picked up a
non-`cli` entrypoint is no longer wrongly hidden.
- A resumed session could show a same-directory sibling's preview text as its own.
The `workingDir` backfill in `mergeUnifiedSessions()` now only ever applies to rows
that have no history entry of their own, so it can no longer overwrite a row's real
content with another conversation's.
- Sessions restarted many times accumulated enough bookkeeping lines to push the real
first prompt past the scanner's 16KB head-read window, leaving a blank row. The read
is now two-tier: 16KB first, escalating to 128KB only when that was not enough, which
is both correct and cheaper than reading 128KB unconditionally (measured on a real
transcript tree: 36% fewer bytes read, roughly 17.5% faster than the unconditional
version). Also restores the tail-read fallback for a file whose head read failed
outright (for example `EMFILE` while scanning hundreds of files), which had been
silently dropping the session from history.
Follow-up hardening on top of the above: the automated-transcript exclusion now
blocklists the SDK entrypoint shape (`sdk`, `sdk-cli`, `sdk-py`) instead of allowlisting
the exact value `cli`. Because the check hides rows, an allowlist failed closed on any
value Claude Code has not shipped yet: a future rename of the interactive entrypoint,
or a second interactive host, would have blanked the entire Past Sessions list with
nothing in the UI to explain it. An unrecognized automated entrypoint now costs a few
noisy rows instead, which is the annoyance this filter set out to fix rather than a
broken feature.
**Phone overview run picker (#214).** The "C" logo home screen's Run picker listed all
six backends regardless of what was installed, so tapping an uninstalled one produced a
failed launch instead of the entry simply not being offered. It is now gated on
`isCliAvailable()` exactly like the desktop toolbar's run-mode dropdown (shell exempt,
since it has no external CLI dependency and keeps the menu from ever being empty). The
picker is a hardcoded duplicate of the toolbar menu rather than a shared render, which
is why it never picked up the earlier gating work; a test now asserts that every mode
the picker offers is gated, so a newly added backend cannot silently drift again.
- 73315bc: fix(web): stop the Claude response viewer from following another session's conversation
The viewer re-derived a pane's live conversation by taking the newest
`~/.claude/history.jsonl` entry for the pane's cwd. A cwd is shared with every
other Codeman tab on it, with tabs long since closed, and with any plain
`claude` run in the user's own terminal, so the eye followed whichever of those
was typed into last — and the adoption was written back to the session, so the
mispin persisted. Entries are now credited to a pane only when they land within
10s of that pane's own Enter and no other pane on the cwd submitted closer, the
same last-submit correlation the Codex locator already uses.
That correlation also has to survive a restart. `start()` resets
`claudeSessionId` to the launch id even when re-attaching to a mux session whose
CLI has since moved on via `/clear`, so a recovered pane pointed the viewer at
its pre-`/clear` transcript — and with the anchor itself living only in memory,
nothing corrected it until the user happened to type again. `lastSubmitAt` is
now persisted in `SessionState` and restored on boot recovery, so the viewer
re-derives the live conversation on its first poll.
## 1.11.0
### Minor Changes
- Two user-facing features since 1.10.0.
**Terminal: Ctrl+C copies the selection, interrupts when nothing is selected** (#211). Copying from the terminal previously worked only through the browser context menu: xterm turns Ctrl+C into 0x03 and cancels the keydown, so the muscle-memory copy failed silently and read as "no copy-paste at all". With a selection, Ctrl+C now copies it, shows the "Copied to clipboard" toast, clears the selection and sends nothing to the PTY; with no selection it falls through unchanged, so the interrupt is intact. Ctrl+Shift+C is an explicit copy chord that never interrupts. The shortcut is a normal registry entry (`copy-selection`), so it can be rebound or disabled in App Settings, and disabling it restores plain always-interrupt Ctrl+C. Copy goes through the Clipboard API with a hidden-textarea fallback, so it also works on plain-HTTP LAN installs.
**File Viewer: edit mode for text files** (#212). The file-preview overlay can now edit workspace text files in place, phone-first: `GET /api/sessions/:id/file-content?edit=1` reads for edit without the 500-line preview truncation (saving a truncated buffer would silently delete the rest) and returns a sha256 hash plus the detected EOL; `PUT /api/sessions/:id/file-content` saves. Edit-in-place only: there is no O_CREAT anywhere in the handler, so "never create, never delete" is structural. Confinement inherits the read path (realpath plus workspace boundary, ownership scoping) and adds sensitive-path and attachment-guard blocklists, a `.git/` subtree deny, and an extension allowlist (`svg` and `env` deliberately excluded). Optimistic concurrency is by content hash, so a file changed on disk mid-edit returns 409 with an overwrite option rather than clobbering. Writes are atomic (`wx` temp, fchmod, fsync, rename) which closes the validate-then-write TOCTOU window and cannot follow a pre-existing symlink. Binary and latin-1 content are refused via a NUL sniff plus a UTF-8 round-trip compare, and EOL is re-applied server-side so a textarea's LF normalization cannot turn a two-line edit of a CRLF file into a whole-file diff.
## 1.10.0
### Minor Changes
+109 -45
View File
File diff suppressed because one or more lines are too long
+204 -34
View File
@@ -5,7 +5,7 @@
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Gemini &bull; Terminal - One Dashboard &bull; Any Device</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -27,7 +27,7 @@
<img src="docs/images/subagent-demo-20260724.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, or Gemini CLI inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
**Codeman** is a self-hosted mission control for AI coding agents. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi inside persistent tmux sessions, streams the real terminal to any browser, and keeps agents productive after you walk away: it re-prompts on idle, resumes when a usage limit resets, runs scheduled jobs, and shows every background agent working in real time.
Get started in one line (macOS & Linux, Windows via WSL):
@@ -42,7 +42,7 @@ codeman web
The installer asks before every system change, and re-running the same line updates in place. Full details: [Quick Start - Installation](#quick-start---installation).
- **One dashboard, four CLIs** - run [Claude Code, OpenCode, Codex, or Gemini](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **One dashboard, six CLIs** - run [Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi](#more-features) per session (plus plain shell), locally, [in Docker](#isolated-docker-sessions), or [over SSH](#remote-ssh-sessions)
- **Truly phone-friendly** - a [touch-optimized terminal](#mobile-optimized-web-ui) with instant local echo, QR login, swipe navigation, and push notifications
- **Runs while you sleep** - [idle detection + respawn cycling](#respawn-controller) and auto-resume when a subscription limit resets, for 24+ hour unattended runs
- **See your agents think** - [live floating windows](#live-agent-visualization) for every subagent and teammate, with real-time transcripts
@@ -68,7 +68,7 @@ This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, a
- **Re-run to update.** The same one-liner updates a finished install in place: local changes in `~/.codeman/app` are stashed (never discarded), and a running service is restarted and verified. If a first install was interrupted, re-running resumes the full setup instead. `install.sh update` and `install.sh uninstall` also exist.
- **CI / headless:** without a terminal attached, steps that would change your system abort with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to approve them for automation.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli) (any combination works). The installer detects whichever of the four is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), or [Pi](https://pi.dev) (any combination works; Gemini CLI is enterprise-only since Google's consumer cutover, and Antigravity is its successor). The installer detects whichever of the six is present; if none is found, it offers to install Claude Code or OpenCode, or you can skip and install one yourself later. After install:
```bash
codeman web
@@ -85,9 +85,29 @@ codeman web --multiuser # named logins + per-user case spaces
Details in [Multi-User Mode](#multi-user-mode-opt-in) below.
<details>
<summary><strong>Run as a background service</strong></summary>
<summary><strong>Keep it running in the background</strong></summary>
The installer's final menu sets this up for you (option 2) and verifies the service actually comes up before claiming success. To configure it manually instead:
To outlive the shell you started it in, without setting anything up:
```bash
codeman web -d # detach; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful SIGTERM; agents keep running in tmux
```
`-d` waits until the server actually answers before reporting success, and refuses to start a second one on the same data dir (two servers sharing a tmux socket attach to each other's sessions).
To have it come back after a reboot, install it as a service instead. The installer's final menu does this for you (option 2); `codeman service` is the equivalent for an `npm i -g aicodeman` install:
```bash
codeman service install # systemd user unit (Linux) or LaunchAgent (macOS)
codeman service status
codeman service uninstall
```
`service install` writes the unit with your current PATH baked in, which matters more than it sounds: launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin`, so a Homebrew or nvm `node`, `tmux` or `claude` is invisible to a hand-written plist. It never copies `CODEMAN_PASSWORD` into the unit file; add that yourself if the service needs auth.
To write the unit by hand instead:
**Linux (systemd):**
@@ -151,7 +171,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), or [Gemini CLI](https://github.com/google-gemini/gemini-cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), or [Pi](https://pi.dev)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
@@ -220,6 +240,8 @@ codeman web # localhost:3000 (loopback only — safe defau
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
codeman web -d # detach: survives closing the shell (--status, --stop)
codeman service install # systemd/launchd service: comes back after reboots
```
Open the printed URL. The page is a single dashboard; everything below happens there.
@@ -230,9 +252,9 @@ Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. **Add Case** creates one from scratch, links an existing folder, or clones a GitHub repo straight into one (**Clone Repo**). |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Antigravity`, `Gemini`, `Pi`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Models → New Claude sessions). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.
@@ -256,7 +278,7 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------- |
| **Respawn** | Long unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: `solo-work`, `overnight-autonomous`, … | Respawn tab |
| **Orchestrator** | Turn one goal into a phased plan and drive it to completion across agents. | Orchestrator panel |
| **Cron** | Saved, named jobs on a schedule (`once`/`interval`/`daily`/`weekly`) that spawn a session and send a prompt when due. | ⏰ Cron button _(opt-in: App Settings → Display → Header Displays)_ |
| **Cron** | Saved, named jobs on a schedule (`once`/`interval`/`daily`/`weekly`) that spawn a session and send a prompt when due. | ⏰ Cron button _(opt-in: App Settings → Header & Panels → Scheduling)_ |
| **Auto-resume** | Automatically continue after a subscription rate-limit resets. | Respawn tab (top) |
### 6. Reach it from anywhere
@@ -268,7 +290,8 @@ Hit start — Codeman spawns the CLI via a real PTY and streams it to your brows
### 7. Operate & maintain
- **App Settings** — model, effort, permission startup mode, theme/skin, notifications, display toggles, per-CLI options, a synced custom display name, and per-device English/Simplified Chinese UI language.
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
- **Run it in the background** — `codeman web -d` detaches from your shell (`--status`, `--stop`); `codeman service install` makes it a systemd user unit / macOS LaunchAgent that survives reboots. Both verify the server actually answers before reporting success, and both refuse to start a second server on one data dir. See [Keep it running in the background](#quick-start---installation).
- **Self-update** — git-clone installs update in place from **App Settings → System → Updates**.
- **Deploy your own changes** — see [Development](#development).
> ⚠️ **Safety:** if you're working _inside_ a Codeman-managed session (`echo $CODEMAN_MUX` → `1`), never run `tmux kill-session` / `pkill claude` directly — use the web UI or `./scripts/tmux-manager.sh`.
@@ -383,6 +406,14 @@ The title is templated into the served HTML on first byte, so it's correct from
| **110k tokens** | Auto `/compact` | Context summarized, work continues |
| **140k tokens** | Auto `/clear` | Fresh start with `/init` |
### Tab Alerts
<p align="center">
<img src="docs/images/tab-alerts-glow-20260815.gif" alt="Session tabs: a regular active tab beside a yellow waiting-for-input tab and a red needs-decision tab, both with a breathing glow" width="900">
</p>
Every tab tells you its state at a glance. A running session keeps its green status dot. When a session stops and waits for input, its tab turns **yellow**: steady ring, tinted background, yellow dot, with a slow breathing glow on top. When a permission prompt or question is **blocking** the agent, the tab turns **red** with a faster pulse. The base tint never blinks off, so even a split-second glance (or a screenshot) reads the true state; the ring stays visible while the tab is selected, and a page reload re-arms pending alerts from the server, so a blocked session can never hide behind a fresh-looking tab.
### Notifications
Real-time desktop alerts when sessions need attention — `permission_prompt` and `elicitation_dialog` trigger critical red tab blinks, `idle_prompt` triggers yellow blinks. Click any notification to jump directly to the affected session. Hooks auto-configured per case directory.
@@ -403,16 +434,18 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, or **Gemini** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `GEMINI_*`/`GOOGLE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Background daemon & service install** — `codeman web -d` runs the server detached with a pidfile, `~/.codeman/web.log`, and verified startup (it polls the server until it answers, so a port clash never reads as success); `codeman service install` writes a systemd user unit (Linux) or LaunchAgent (macOS) with your shell's PATH baked in, so an nvm or Homebrew `node`, `tmux` and `claude` are actually found. Secrets are never written into unit files
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → System → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Clone a GitHub repo as a case** — paste a repository URL into **Add Case → Clone Repo** and Codeman clones it into `~/codeman-cases/<name>` and registers it as a normal case, ready to run an agent in. It preflights the URL while you type (tells you whether it can be cloned anonymously and offers the repo's real branches and tags for the optional branch/tag field), fills the case name in from the URL, and lets you pick which CLI the Run button should use. Public repositories over `https://`; Codeman never collects or stores credentials
- **Multi-CLI** — run **Claude Code**, **OpenCode**, **Codex**, **Antigravity**, **Gemini**, or **Pi** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*` vs `ANTIGRAVITY_*` vs `GEMINI_*`/`GOOGLE_*` vs `PI_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md) and [`docs/pi-integration.md`](docs/pi-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Remote SSH sessions** — point a case at another machine and run the agent there inside a durable remote tmux: survives SSH drops, auto-reconnects, and can discover + attach sessions already running on the host. See [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
- **Voice input** — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Image input** — paste or drag-and-drop images straight into a session
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Display
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Terminal & Input
- **Multi-monitor span** _(macOS)_ — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
- **File Viewer button** _(opt-in)_ — a header button that toggles the built-in file browser panel with one tap; enable under App Settings → Display → Header Displays
- **File Viewer button** _(opt-in)_ — a header button that toggles the built-in file browser panel with one tap; enable under App Settings → Header & Panels → Header buttons
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean
- **OS notifications & hostname-aware titles** — desktop alerts and tab titles are prefixed `codeman:<host>` so multi-host setups stay unambiguous
@@ -426,7 +459,7 @@ Run a case inside its own hardened Docker container instead of directly on your
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket; a **sealed** profile (no host credentials, network off) is one toggle away.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Gemini / OpenCode logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Seamless auth, isolated credentials** — your host Claude / Codex / Antigravity / Gemini / OpenCode / Pi logins work inside the container out of the box: credentials are seeded (copied) in at launch and onboarding/trust prompts are pre-answered, so no login wizard appears. The container keeps its own copies and never writes back to your host credential stores; only conversation transcripts are shared, and exports never capture secrets.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
@@ -496,7 +529,7 @@ The script auto-installs a systemd user service on first run. The tunnel URL is
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
# Or via the Codeman web UI: Settings → Tunnel → Toggle On
# Or via the Codeman web UI: App Settings → System → Remote access → Cloudflare Tunnel
```
</details>
@@ -598,7 +631,7 @@ By default Codeman launches sessions with `--dangerously-skip-permissions`, so t
- **Loopback by default** — the server binary binds `127.0.0.1`, reachable only from the same machine, so the no-password default is safe out of the box (the guided installer asks about network access and configures the binding + password for you). Binding a non-loopback host without `CODEMAN_PASSWORD` _starts but prints a loud warning_ with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with `--allow-unauthenticated-network`)
- **Optional auth, real sessions** — HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Success issues an opaque 256-bit `codeman_session` cookie (`randomBytes(32)`) — validated server-side, not client-signed, so it can't be forged offline (24h TTL, auto-extend, device-context audit log)
- **Per-IP rate limiting** — 10 failed attempts → `429` with `Retry-After` (15-min decay). A valid cookie or correct password recovers _immediately_ even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter
- **Configurable permission mode** - `--dangerously-skip-permissions` is only the default. **App Settings → Claude CLI → Startup Mode** can switch new sessions to Anthropic's classifier-guarded `auto` mode (low-prompt, needs Claude Code 2.1.207+), `normal` prompting, or an explicit allowed-tools list. In multi-user mode, non-granted users are forced to `auto`, and shell sessions / skip-permissions require an explicit per-user grant
- **Configurable permission mode** - `--dangerously-skip-permissions` is only the default. **App Settings → Agents & CLIs → Claude → Startup Mode** can switch new sessions to Anthropic's classifier-guarded `auto` mode (low-prompt, needs Claude Code 2.1.207+), `normal` prompting, or an explicit allowed-tools list. In multi-user mode, non-granted users are forced to `auto`, and shell sessions / skip-permissions require an explicit per-user grant
### Always-on browser hardening (v0.9.5)
@@ -612,7 +645,7 @@ These run for **every** request — before auth, even on the default no-password
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` env-prefix allowlist gates which settings each CLI can receive
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
@@ -650,7 +683,10 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
| `Ctrl/Cmd+Tab` | Next session |
| `Alt/Option+[` / `Alt/Option+]` | Previous / next session |
| `Alt/Option+1`-`Alt/Option+9` | Switch to tab N (physical keys, so macOS Option layouts work) |
| `Alt/Option+B` | Collapse / expand the session sidebar (sidebar layout only) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+C` | Copy selection, or interrupt when nothing is selected |
| `Ctrl+Shift+C` | Copy selection (never interrupts) |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
@@ -665,6 +701,78 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
### The agent skill (start here)
Everything in this section also ships as a **Claude Code skill** in [`skills/codeman`](skills/codeman/SKILL.md). Install it once and you never paste API docs into a prompt again. You ask for what you want in plain English, and the agent already sitting inside a Codeman session loads the recipes and drives the API itself.
#### Step 1: install it
| How | Command | Scope |
| -------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| Skills CLI | `npx skills add Ark0N/Codeman --skill codeman -g` | Global, works for any skills-aware agent |
| Bundled CLI | `codeman skill install` | Global (`~/.claude/skills/codeman`), for npm installs that never cloned the repo |
| Bundled CLI | `codeman skill install --case <name>` | One case only |
| Web UI | App Settings → Agents & CLIs → Claude → **Agent Skill** | Auto-injects into each case on Claude session create (`agentSkillEnabled`, SYNCED, default off) |
`codeman skill uninstall [--case <name>]` reverses the CLI installs, and never touches a `skills/codeman` you wrote yourself.
#### Step 2: ask for things
That is the entire interface. No curl, no endpoint names, no session ids. These prompts work as written:
| You say | The skill does |
| ------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------- |
| _"What sessions are running right now?"_ | Lists them with name, mode and status. Read-only, safe to ask anytime. |
| _"Start a shell worker on the `myapp` case, run the test suite, tell me if it passes."_ | Spawns, waits on a split completion marker, reads back the exit code, cleans up. |
| _"Spin up 3 workers for lint, typecheck and tests. Run them in parallel, report failures."_ | The fan-out flow: one session per task, all started first, then gathered as each finishes. |
| _"Have a claude worker on `refactor-auth` summarize `src/session.ts`, then close it."_ | Spawns, runs the readiness ladder (first-run trust dialog included), send-and-wait, reads the clean transcript answer, deletes. |
| _"Watch session w4 and tell me if it gets stuck on a permission prompt."_ | Blocks on the `blocked` signal and surfaces the question to **you**. It never answers another session's prompt itself. |
#### Step 3: nothing
The agent deletes every session it started. Watch the tabs appear and disappear in the dashboard while it works.
#### A real run, start to finish
> **You:** spin up 3 shell workers, run lint / typecheck / the frontend syntax check in parallel, and tell me which failed.
```text
lint -> 9f2d8e5f dispatched
typecheck -> aff9c691 dispatched 3 tabs appear in the dashboard
syntax -> be9f1f15 dispatched
lint DONE_lint_17909 rc=0
typecheck DONE_typecheck_3409 rc=0 gathered as each one finishes
syntax DONE_syntax_18501 rc=0
deleted 9f2d8e5f, aff9c691, be9f1f15 tabs disappear
```
Those `DONE_<task>_<random>` strings are the skill's **split marker** trick, and they are why the fan-out is reliable on hook-less `shell` sessions: the typed line contains `${M}_17909`, so only the command's real *output* ever contains `DONE_17909`. An unsplit marker would match the echo of your own keystrokes before the command had even run.
#### What's in the box
| File | Contents |
| --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- |
| [`SKILL.md`](skills/codeman/SKILL.md) | Safety rules, the ready-made fast path (spawn N workers, task them, collect), and the verb index. Always loaded. |
| [`reference/verbs.md`](skills/codeman/reference/verbs.md) | The 14 verbs in detail: readiness, send-and-wait, markers, interrupts, cleanup. On demand. |
| [`reference/recipes.md`](skills/codeman/reference/recipes.md) | 6 worked multi-worker flows (fan-out, blocked-worker watch, messaging fan-out). On demand. |
| [`reference/endpoints.md`](skills/codeman/reference/endpoints.md) | Full endpoint tables, error codes, per-mode signal table, capacity limits. On demand. |
| [`reference/messaging.md`](skills/codeman/reference/messaging.md) | Talking to claude workers directly via Claude Code cross-session messaging. On demand. |
Every recipe in there was verified against a live server, and the comments record the failure modes that were measured rather than guessed.
#### Two things worth knowing
- **It self-gates.** Outside a Codeman session (`CODEMAN_MUX` unset) the skill refuses to act and does not guess an API URL, so a global install costs an unrelated Claude Code session nothing.
- **It is deliberately conservative.** Unprompted, it may only spawn sessions, prompt them, and delete ones **it created in that same conversation, by exact id**, through a fail-closed guard that refuses to delete the agent's own session. Deleting a case (which erases a real directory of your code), bulk kills, respawn/ralph/cron/orchestrator changes and settings writes all require you to ask, naming the target.
⚠️ Turning `agentSkillEnabled` back off **does not remove already-injected copies** (a create-time sweep would yank the skill out from under other live sessions sharing that `.claude/` dir). Remove them per case with `codeman skill uninstall --case <name>`.
---
**The rest of this section is the manual path**: the same operations as raw HTTP, for a CI bot, a shell script, or any agent without skill support.
### Detect that you're inside Codeman
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
@@ -678,15 +786,21 @@ When a CLI runs in a Codeman-managed session, these environment variables are se
### Rules of the road (read before you POST)
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
1. **Single-line input, ending in `\r`.** Programmatic input is sent as literal text, and Enter fires **only when the input contains a carriage return**: `{"input":"run tests\r"}`. Without the `\r` the text sits on the session's prompt unsubmitted (and a combined `wait` runs its full timeout on a turn that never started). Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs the joined command `echo Aecho B`: send one line per call.
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard). ⚠️ A `401` replies with the bare string `Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error instead of showing the failure: check the status before parsing.
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5. **`/api/v1/*`** is a stable alias of `/api/*`.
6. **Wait instead of polling, and don't treat a timeout as an error.** The wait endpoints answer with HTTP `200` and `wait.timedOut: true` when nothing happened in time, so loop over short waits (60s is the default) rather than issuing one long call, because tunnels cut idle connections. `wait.timeoutMs` tells you the timeout the server actually applied after clamping (600s ceiling).
7. **Only `claude` sessions emit `stop` and `blocked`.** Those two come from Claude Code hooks; `shell` and the external CLIs (opencode/codex/gemini/antigravity/pi) accept only `idle`, `working` and `exit`. Asking for `stop` explicitly on those is a `400`; omitting `until` is always safe. ⚠️ On a `shell` session `idle` fires **once**, at startup, and never again, so send-and-wait there can only time out; synchronize hook-less sessions with a `wait-output` marker.
8. **Nothing reports "ready", so wait for it explicitly.** A new session answers `{"signal":"exit","immediate":true}` (that means *not started*, not *crashed*) until its PID exists, and a `claude` worker in a fresh case then sits on the CLI's trust dialog. Prompt it there and the wait resolves on `idle` in ~2s looking exactly like a finished turn, while the text sits stuck in the dialog. Recipe 2b below is the sequence that avoids it.
### Recipes
```bash
# CODEMAN_API_URL is auto-set inside every Codeman session, correct scheme included.
# The fallback below fits a stock install; on a --https install set the https:// URL
# yourself and add -k to each curl (self-signed cert).
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
@@ -698,18 +812,63 @@ curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. Wait until that worker is actually READY (see rule 8): composer marker first,
# first-run trust dialog only as the fallback. (Probing trust first and sending
# a blind Enter misfires on re-runs: the dialog text stays in the buffer forever,
# so the probe matches stale text and the Enter lands in a ready composer.)
# Match single tokens: TUI text can reach the matcher without its spaces.
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # accept the first-run trust dialog
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. Read the terminal back
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 4. Send a prompt and BLOCK until that turn is done (registers the wait before
# writing, so it can't answer with the previous turn's idle state)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` is the definitive end-of-turn hook. Adding `idle` makes it resolve on a
# spinner pause too, and on anything that redraws a ❯ prompt — like a dialog.)
# 5. Stream live events (session output, agent activity, status)
# 4b. Timed out? That's a 200, not a failure. Loop over short waits.
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. Or wait for a marker in the output (works for shell sessions too).
# ⚠️ Unique per call (tmux repaints replay old screen text), and SPLIT so the
# typed line never contains it: your own keystrokes echo into the output
# stream, so an unsplit marker matches before the command has run. from=buffer
# catches a marker that printed before the wait landed.
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. Read the terminal back. ⚠️ Use terminal?tail=, NOT /output: the latter's
# textOutput is empty for every tmux-backed (i.e. every interactive) session.
# tail counts BYTES, and what comes back is terminal data, ANSI included.
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events" # Server-Sent Events
# 6. Schedule recurring work (cron-style job)
# 7. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
@@ -717,11 +876,11 @@ curl -s -X POST "$API/api/cron/jobs" \
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. Inspect background sub-agents and their transcripts
# 8. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
# 9. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq
```
@@ -747,7 +906,7 @@ Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_pro
## API
REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
REST over Fastify — **~200 handlers across 21 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
@@ -755,8 +914,11 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | List all |
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
| `GET` | `/api/sessions/:id/output` | Read terminal output |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`: `clientId`+`seq` = exactly-once; `wait` blocks until the turn ends) |
| `GET` | `/api/sessions/:id/terminal` | Read terminal output (`?tail=<bytes>`, `?full=1`); the read path for interactive sessions |
| `GET` | `/api/sessions/:id/output` | Parsed one-shot output (`textOutput` is empty for tmux-backed sessions) |
| `GET` | `/api/sessions/:id/wait` | Block until a signal fires (`?until=stop,idle,exit&timeout=&fresh=`); a timeout is a `200` |
| `GET` | `/api/sessions/:id/wait-output` | Block until a literal string appears (`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | Unified live + history list (Session Manager) — `?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | Pin/unpin in the Session Manager (`{pinned}`) |
| `PUT` | `/api/session-order` | Sync tab order across devices (`{order: [ids]}`) |
@@ -810,6 +972,8 @@ REST over Fastify — **~190 handlers across 20 route modules**, plus an SSE str
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
> **Building something on top of Codeman?** [`docs/extending-codeman.md`](docs/extending-codeman.md) is the integration guide: render your own UI as a tab, subscribe to the SSE event stream to react when an agent needs you, drive Codeman from a script, and the traps worth knowing before you start. Codeman has no plugin runtime on purpose, so an integration is just your own process talking HTTP.
---
## Architecture
@@ -842,7 +1006,7 @@ flowchart TB
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
@@ -873,13 +1037,19 @@ flowchart TB
npm install
npx tsx src/index.ts web # Dev mode
npm run build # Production build
npm run test:ci # Run tests (the CI suite; browser suites need extra setup)
npm test # Run tests (same suite CI runs; browser/mobile/perf suites have their own commands)
```
See [CLAUDE.md](./CLAUDE.md) for full documentation.
---
## Community
Questions, setup help, and ideas live in [GitHub Discussions](https://github.com/Ark0N/Codeman/discussions): the [Q&A section](https://github.com/Ark0N/Codeman/discussions/categories/q-a) answers the most common ones (phone access, overnight runs, updating), and the roadmap gets decided in [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas). Bugs go to [issues](https://github.com/Ark0N/Codeman/issues); reports usually get a response within a day, and every release credits its reporters and contributors by name. Want to contribute? [CONTRIBUTING.md](.github/CONTRIBUTING.md) has the map: skins, translations, and docs make great first PRs, and bigger features start life as a Discussion. And if you're proud of your rig, post it in [Show and tell](https://github.com/Ark0N/Codeman/discussions/300).
---
## Codebase Quality
The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:
+103 -29
View File
@@ -5,7 +5,7 @@
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Gemini &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Antigravity &bull; Gemini &bull; Pi &bull; 终端 —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -58,7 +58,7 @@ curl -fsSL https://getcodeman.com/install | bash
- **重跑即更新。** 再次运行同一条命令即可原地更新已完成的安装:`~/.codeman/app` 中的本地改动会被 stash(绝不丢弃),运行中的服务会自动重启并校验。若首次安装中途失败,重跑会继续完成完整的安装流程。也可以使用 `install.sh update` 与 `install.sh uninstall`。
- **CI / 无终端环境:** 没有终端时,涉及系统改动的步骤会带着说明中止,而不是静默执行;在自动化场景设置 `CODEMAN_NONINTERACTIVE=1` 即可批准这些步骤。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli)(任意组合均可)。安装器会自动检测这四个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev)(任意组合均可;自 Google 面向消费者停售后,Gemini CLI 仅限企业版,Antigravity 是其继任者)。安装器会自动检测这六个中已安装的任意一个;若一个都没有,会提供安装 Claude Code 或 OpenCode 的选项,也可以选择跳过、稍后自行安装。安装完成后:
```bash
codeman web
@@ -141,7 +141,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli) 或 [Gemini CLI](https://github.com/google-gemini/gemini-cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai)、[Codex](https://developers.openai.com/codex/cli)、[Antigravity](https://antigravity.google)、[Gemini CLI](https://github.com/google-gemini/gemini-cli) 或 [Pi](https://pi.dev))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
@@ -221,7 +221,7 @@ codeman web -H 0.0.0.0 # 绑定局域网 —— 必须设置 CODEMAN_
| 字段 | 作用 |
| ---------------------- | ------------------------------------------------------------------------------------------- |
| **工作目录 / case** | 智能体操作的文件夹。「case」就是一个 Codeman 记住的命名工作目录。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Gemini` 或 `Terminal`(普通 shell)。 |
| **CLI / 运行模式** | `Claude`(默认)、`OpenCode`、`Codex`、`Antigravity`、`Gemini`、`Pi` 或 `Terminal`(普通 shell)。 |
| **模型** | 每会话模型(App Settings → Claude Model)。软默认值 —— 会话内 `/model` 依然有效。 |
| **Effort / Ultracode** | 推理力度(`low`–`max`),或用 `ultracode` 开启动态多智能体工作流。随时可用 `/effort` 切换。 |
@@ -394,7 +394,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex** 或 **Gemini**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode**、**Codex**、**Antigravity**、**Gemini** 或 **Pi**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*`、`CODEX_*`、`ANTIGRAVITY_*`、`PI_*` 与 `GEMINI_*`/`GOOGLE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md) 与 [`docs/pi-integration.md`](docs/pi-integration.md)
- **Docker 会话** —— 在隔离且加固的容器中运行案例。**Create New** 上勾选一个复选框即可用合理的默认值启动容器并在其中启动智能体;同一案例的多个会话共享一个容器;可将容器连同工作区导出为可移植的 `.tar.gz`,迁移到另一台机器。详见 [`docs/docker-cases.md`](docs/docker-cases.md)
- **远程 SSH 会话**:把案例指向另一台机器,让智能体在那里一个持久的远程 tmux 中运行:SSH 断连不中断任务、自动重连,还能发现并附着主机上已在运行的会话。详见 [`docs/remote-sessions.md`](docs/remote-sessions.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
@@ -416,7 +416,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
- **资源模板** —— 展开复选框可选 **Small / Medium / Large / GPU** 预设(内存、CPU、GPU),也可以完全自定义。**磁盘是弹性的** —— 存储随数据增长,没有固定上限。
- **按案例共享容器** —— 多个会话可以 `docker exec` 进同一个容器;结束某个会话绝不会影响其他会话所在的容器。
- **默认加固** —— 非 root、`--cap-drop ALL`、`no-new-privileges`、PID/内存上限,绝不使用 `--privileged` 或 docker socket;**密封(sealed)** 配置(不注入主机凭据、关闭网络)只需一个开关。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Gemini / OpenCode 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **无感认证、凭据隔离** —— 主机上的 Claude / Codex / Antigravity / Gemini / OpenCode / Pi 登录在容器内开箱即用:凭据在启动时以只读种子方式复制注入,onboarding/信任提示已预先答复,不会弹出登录向导。容器保留自己的副本,绝不回写主机的凭据存储;跨边界共享的只有对话转录,导出文件也绝不包含机密。
- **迁移到另一台机器** —— 把容器的完整环境(工具链 + 工作区)导出为可移植的 `.tar.gz`,在另一台机器上导入到新案例即可继续。
- **持久耐用** —— Codeman 重启后重连会回到同一个存活的智能体;容器停止/重启后则从绑定挂载的转录恢复对话。
@@ -602,7 +602,7 @@ Codeman 默认用 `--dangerously-skip-permissions` 启动会话,因此 Web UI
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `GEMINI_*` / `GOOGLE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` / `ANTIGRAVITY_*` / `GEMINI_*` / `GOOGLE_*` / `PI_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
@@ -641,6 +641,8 @@ sc -l # 列出会话
| `Alt/Option+[` / `Alt/Option+]` | 上一个 / 下一个会话 |
| `Alt/Option+1`–`Alt/Option+9` | 切换到第 N 个标签(按物理键位,macOS Option 布局也适用) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | 将当前标签左移 / 右移 |
| `Ctrl/Cmd+C` | 复制选中内容;未选中时中断代理 |
| `Ctrl+Shift+C` | 复制选中内容(永不中断) |
| `Ctrl/Cmd+L` | 清屏 |
| `Ctrl+Shift+R` | 恢复终端尺寸 |
| `Ctrl+Shift+V` | 切换语音输入 |
@@ -655,6 +657,16 @@ sc -l # 列出会话
面向不经浏览器控制 Codeman 的 AI 智能体与自动化:一个拉起工作会话的智能体、一个 CI 机器人,或是**运行在 Codeman 会话*内部*、编排其他会话的 Claude Code**。UI 能做的一切都是 HTTP + CLI,因此智能体也能做。
> **捷径:装上打包好的智能体技能。** 下面这一整套(外加多工作会话的实战配方)已经作为 Claude Code 技能随仓库发布在 [`skills/codeman`](skills/codeman/SKILL.md),会话内部的智能体不必等你把文档粘进提示词就能驱动 Codeman。三种获取方式:
>
> - `npx skills add Ark0N/Codeman --skill codeman -g`:全局安装,任何支持技能的智能体都能用
> - `codeman skill install`(全局)或 `codeman skill install --case <name>`:给那些从 npm 安装、从未克隆过仓库的用户;`codeman skill uninstall` 可撤销
> - **App Settings → Agent Skill**(`agentSkillEnabled`,默认关闭):开启后,Codeman 会在每次于某个 case 中创建 Claude 会话时把技能注入该 case;case 里用户自己写的 `skills/codeman` 永远不会被覆盖
>
> 全局安装(`codeman skill install` 或 `npx skills add`)会被**本机每一个新建的 Claude Code 会话**读到,无论它在不在 Codeman 里。技能自带门禁:不在 Codeman 会话中(`CODEMAN_MUX` 未设置)时它拒绝动作,所以全局装上它对无关会话没有代价。
>
> ⚠️ 把 `agentSkillEnabled` 关回去**不会删掉已经注入的副本**(在创建时做清扫,会把技能从共用同一个 `.claude/` 目录的其他活动会话脚下抽走)。要删就按 case 删:`codeman skill uninstall --case <name>`。
### 检测自己身处 Codeman 内部
当 CLI 运行在 Codeman 受管会话中时,以下环境变量会被设置 —— 读取它们,别硬编码任何东西:
@@ -668,15 +680,21 @@ sc -l # 列出会话
### 行路规则(POST 之前先读)
1. **只发单行输入。** 编程输入会作为字面文本 **+ Enter** 一次性发送。多行字符串会破坏智能体 TUI(Ink)—— 发送一行,或拆成多次调用。
1. **只发单行输入,而且必须以 `\r` 结尾。** 编程输入按字面文本发送,**只有当输入里含回车符时才会触发 Enter**:`{"input":"run tests\r"}`。少了 `\r`,文本就停在会话的输入框里不被提交(同一次调用里的 `wait` 还会在一个压根没开始的回合上耗满整个超时)。内嵌的换行会被剥掉而不是报错,因此 `"echo A\necho B\r"` 执行的是拼起来的 `echo Aecho B`:一次调用只发一行。
2. **让输入幂等。** 在 `POST …/input` 上带上稳定的 `clientId` 和按会话单调递增的 `seq`。服务端会去重,因此连接中断后的重试不会重复投递提示。
3. **认证。** 若设置了 `CODEMAN_PASSWORD`,发送 HTTP Basic 认证(用户 `admin` 或 `CODEMAN_USERNAME`)或 `codeman_session` cookie。默认的环回安装无密码。缺失的 `Origin` 头被允许,因此普通 `curl` 可用;跨站的浏览器 origin 会被拒绝(CSRF 防护)。
3. **认证。** 若设置了 `CODEMAN_PASSWORD`,发送 HTTP Basic 认证(用户 `admin` 或 `CODEMAN_USERNAME`)或 `codeman_session` cookie。默认的环回安装无密码。缺失的 `Origin` 头被允许,因此普通 `curl` 可用;跨站的浏览器 origin 会被拒绝(CSRF 防护)。⚠️ `401` 回的是裸字符串 `Unauthorized`,**不是** JSON 信封,直接喂给 `jq` 只会抛解析错误而看不到真正的失败原因:先看状态码,再解析。
4. **响应信封。** 多数端点返回 `{ "success": true, "data": … }`(错误:`{ "success": false, "error", "errorCode" }`)。少数遗留 GET 返回裸响应体 —— **两种都要处理**(`body.data ?? body`)。
5. **`/api/v1/*`** 是 `/api/*` 的稳定别名。
6. **用等待代替轮询,别把超时当成错误。** 等待类端点在没等到事情发生时也以 HTTP `200` 加 `wait.timedOut: true` 应答,所以要循环调用短等待(默认 60 秒),而不是发一个超长的调用:隧道会掐断空闲连接。`wait.timeoutMs` 告诉你服务端钳制之后真正采用的超时(上限 600 秒)。
7. **只有 `claude` 会话会发出 `stop` 与 `blocked`。** 这两个来自 Claude Code hook;`shell` 与外部 CLI(opencode/codex/gemini/antigravity/pi)只接受 `idle`、`working` 与 `exit`。在这些模式上显式索要 `stop` 会得到 `400`;不传 `until` 则永远安全。⚠️ `shell` 会话的 `idle` 只在启动时触发**一次**,此后再也不会,所以在那里用「发送并等待」只能等到超时:没有 hook 的会话请用 `wait-output` 标记来同步。
8. **没有任何东西会报告「就绪」,得自己显式等。** 新会话在 PID 出现之前一律回答 `{"signal":"exit","immediate":true}`(意思是*还没启动*,不是*崩了*),而全新 case 里的 `claude` 工作会话接着会停在 CLI 的信任对话框上。此时给它发提示,等待会在约 2 秒后因 `idle` 解除,看上去和一个跑完的回合一模一样,而文本其实卡在对话框里。下面的配方 2b 就是避开它的顺序。
### 常用配方
```bash
# 每个 Codeman 会话里都自动设好了 CODEMAN_API_URL,协议也是对的。
# 下面的兜底值适用于标准安装;在 --https 安装上请自己写 https:// 的地址,
# 并给每个 curl 加上 -k(自签名证书)。
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (若设置了密码,给每个调用加上 -u admin:"$CODEMAN_PASSWORD")
@@ -688,18 +706,69 @@ curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 2b. 等这个工作会话真正就绪(见规则 8):先探输入框的标记,信任对话框只作兜底。
# (反过来先探信任对话框、再盲发一个 Enter,在重复运行时会误伤:对话框的文字
# 会一直留在缓冲区里,探测因此匹配到旧文本,而那个 Enter 落进了已经就绪的输入框。)
# 匹配单个词:TUI 的文字到达匹配器时可能已经丢掉了词间空格。
until [ "$(curl -s "$API/api/sessions/$SID" | jq '.data.pid')" != null ]; do sleep 1; done
R=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=5000')
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$(curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=trust' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d '{"input":"\r","useMux":true}' # 接受首次运行的信任对话框
curl -sG "$API/api/sessions/$SID/wait-output" --data-urlencode 'match=bypass' \
--data-urlencode 'from=buffer' --data-urlencode 'timeout=45000' >/dev/null
fi
# 3. 向会话发送提示(精确一次:clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,"clientId":"agent-1","seq":1}'
# 4. 读回终端内容
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 4. 发送提示并阻塞到这一回合结束(先注册等待再写入,因此不会拿上一回合的状态来应答)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures\r","useMux":true,
"clientId":"agent-1","seq":2,"wait":"stop,exit","waitTimeout":60000}' \
| jq '.data.wait' # -> {"signal":"stop","timedOut":false,"waitedMs":41230,...}
# (`stop` 是回合结束的权威 hook。加上 `idle` 会让它在转圈停顿时也解除,
# 任何重画出 ❯ 提示符的东西同理,比如一个对话框。)
# 5. 流式接收实时事件(会话输出、智能体活动、状态)
# 4b. 超时了?那是 200,不是失败。循环调用短等待即可。
curl -s "$API/api/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq '.data.wait'
# 4c. 或者等输出里出现某个标记(shell 会话也适用)。
# ⚠️ 每次调用都要用不同的标记(tmux 重画会重放旧屏幕文字),并且把标记拆开写,
# 让敲进去的那一行本身不包含它:你自己的按键会回显进输出流,不拆开的标记会在
# 命令还没跑之前就匹配上。from=buffer 用来接住在等待落地之前就已打印的标记。
N=$RANDOM
curl -s -X POST "$API/api/sessions/$SID/input" -H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
curl -sG "$API/api/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
# 5. 读回答案。claude / codex 会话用 last-response:它取自 transcript 而不是屏幕,
# 因此不带 TUI 的画框与重画噪声。⚠️ 要轮询,别只读一次:transcript 落盘比 stop
# 信号稍晚,紧跟着「发送并等待」返回后立刻读,常常拿到空串。
for _ in $(seq 1 10); do
TXT=$(curl -s "$API/api/sessions/$SID/last-response" | jq -r '.data.text')
[ -n "$TXT" ] && break; sleep 1
done
printf '%s\n' "$TXT"
# 5b. 其他模式(shell/opencode/gemini/antigravity/pi)没有 transcript,读终端。
# ⚠️ 用 terminal?tail=,不要用 /output:后者的 textOutput 对每个由 tmux 承载的
# (也就是每个交互式)会话都是空的。tail 按字节计,返回的是含 ANSI 的终端数据。
curl -s "$API/api/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
# 6. 流式接收实时事件(会话输出、智能体活动、状态)
curl -sN "$API/api/events" # Server-Sent Events
# 6. 调度周期性工作(cron 风格任务)
# 7. 调度周期性工作(cron 风格任务)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
@@ -707,11 +776,11 @@ curl -s -X POST "$API/api/cron/jobs" \
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. 查看后台子智能体及其活动记录
# 8. 查看后台子智能体及其活动记录
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. 全系统快照(会话、设置、重生、统计)
# 9. 全系统快照(会话、设置、重生、统计)
curl -s "$API/api/status" | jq
```
@@ -737,20 +806,23 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
## API
基于 Fastify 的 REST —— **20 个路由模块中约 190 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。所有响应都使用 `ApiResponse<T>` 信封(`{success, data}` / `{success, error, errorCode}`);`/api/v1/*` 是稳定别名。以下是一个有代表性的子集:
基于 Fastify 的 REST —— **21 个路由模块中约 200 个处理器**,外加一条 SSE 流和一条 WebSocket 终端通道。所有响应都使用 `ApiResponse<T>` 信封(`{success, data}` / `{success, error, errorCode}`);`/api/v1/*` 是稳定别名。以下是一个有代表性的子集:
### 会话(Sessions)
| 方法 | 端点 | 说明 |
| -------- | -------------------------- | ------------------------------------------------------------------------------ |
| `GET` | `/api/sessions` | 列出全部 |
| `POST` | `/api/quick-start` | 创建 case + 启动会话(`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | 发送输入(`{input, useMux?, clientId?, seq?}` —— `clientId`+`seq` = 精确一次) |
| `GET` | `/api/sessions/:id/output` | 读取终端输出 |
| `GET` | `/api/sessions/unified` | 统一的活动 + 历史清单(会话管理器):`?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | 在会话管理器中置顶 / 取消置顶(`{pinned}`) |
| `PUT` | `/api/session-order` | 跨设备同步标签顺序(`{order: [ids]}`) |
| `DELETE` | `/api/sessions/:id` | 删除会话 |
| 方法 | 端点 | 说明 |
| -------- | ------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | 列出全部 |
| `POST` | `/api/quick-start` | 创建 case + 启动会话(`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | 发送输入(`{input, useMux?, clientId?, seq?, wait?, waitTimeout?}`:`clientId`+`seq` = 精确一次;`wait` 阻塞到这一回合结束) |
| `GET` | `/api/sessions/:id/terminal` | 读取终端输出(`?tail=<bytes>`、`?full=1`):交互式会话的读取路径 |
| `GET` | `/api/sessions/:id/output` | 一次性的解析输出(tmux 承载的会话里 `textOutput` 为空) |
| `GET` | `/api/sessions/:id/wait` | 阻塞到某个信号触发(`?until=stop,idle,exit&timeout=&fresh=`);超时是 `200` |
| `GET` | `/api/sessions/:id/wait-output` | 阻塞到某个字面串出现(`?match=&nocase=&from=now\|buffer&timeout=`) |
| `GET` | `/api/sessions/unified` | 统一的活动 + 历史清单(会话管理器):`?q=&limit=` |
| `POST` | `/api/sessions/:id/pin` | 在会话管理器中置顶 / 取消置顶(`{pinned}`) |
| `PUT` | `/api/session-order` | 跨设备同步标签顺序(`{order: [ids]}`) |
| `DELETE` | `/api/sessions/:id` | 删除会话 |
### 重生(Respawn)
@@ -800,6 +872,8 @@ Codeman 会注册 Claude Code hook,它们 `POST /api/hook-event`(`permission
| `POST` | `/api/clipboard` | 把文本推送到所有已连接浏览器(`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | 时间线 + 统计 |
> **想在 Codeman 之上做集成?**[`docs/extending-codeman.md`](docs/extending-codeman.md)(英文)是集成指南:把你自己的界面作为标签页嵌入、订阅 SSE 事件流以便在 agent 需要你时做出响应、用脚本驱动 Codeman,以及动手前值得先了解的那些坑。Codeman 刻意不提供插件运行时,所以一个集成就是你自己的进程在讲 HTTP。
---
## 架构
@@ -832,7 +906,7 @@ flowchart TB
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Gemini</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex / Antigravity / Gemini / Pi</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
@@ -863,7 +937,7 @@ flowchart TB
npm install
npx tsx src/index.ts web # 开发模式
npm run build # 生产构建
npm run test:ci # 运行测试(CI 套件;浏览器套件需要额外环境)
npm test # 运行测试(与 CI 相同;浏览器/移动端/性能套件另有独立命令)
```
完整文档见 [CLAUDE.md](./CLAUDE.md)。
+45
View File
@@ -0,0 +1,45 @@
/**
* The test suites that `npm test` deliberately does NOT run, in one place.
*
* Why this file exists: the exclusion list used to live only in
* config/vitest.ci.config.ts, as literals. Anything excluded there was
* therefore reachable only by running the everything-config by hand and reading
* past its failures — and a newly excluded file was reachable by nothing at
* all, silently, because nothing pointed at it. Both configs now derive their
* globs from the arrays below, so adding a suite here puts it in exactly one
* runner and takes it out of exactly one gate.
*
* Adding a new test that cannot run in CI: put its glob in the array that
* describes WHY it cannot, not in whichever one is shortest.
*/
/**
* Playwright-driven: needs chromium and, in most cases, a live Codeman server
* on a real port. Deterministic where the environment provides both, which is
* why these are a runnable suite (`npm run test:browser`) rather than skipped.
*/
export const BROWSER_TEST_GLOBS = [
'test/inline-rename.test.ts',
'test/opencode-resize.test.ts',
'test/webgl-fallback.test.ts',
'test/terminal-copy-shortcut.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
];
/**
* Wall-clock benchmarks. They assert on durations, so a loaded shared runner
* fails them for reasons that have nothing to do with the diff under test.
*/
export const PERF_TEST_GLOBS = ['test/perf-*.test.ts'];
/**
* Browser + visual regression: chromium AND environment-specific PNG baselines
* that are generated per machine. Has its own config
* (test/mobile/vitest.config.ts) because it needs serial execution, a longer
* timeout and the `pretest:mobile` vendor step — run it with
* `npm run test:mobile`, not through the configs here.
*/
export const MOBILE_TEST_GLOBS = ['test/mobile/**'];
/** Everything `npm test` skips. */
export const NON_CI_TEST_GLOBS = [...MOBILE_TEST_GLOBS, ...PERF_TEST_GLOBS, ...BROWSER_TEST_GLOBS];
+34
View File
@@ -0,0 +1,34 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
import { BROWSER_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The Playwright-driven suite `npm test` skips — `npm run test:browser`.
*
* Needs chromium and, for most of these, a live Codeman server on a real port;
* codex-predictive-echo also needs a real codex binary. Expect failures where
* the machine cannot provide those, and read them as "not runnable here", not
* as a regression.
*
* The mobile suite is NOT here: it needs per-machine PNG baselines, serial
* execution and the `pretest:mobile` vendor step, so it keeps its own config
* (test/mobile/vitest.config.ts) behind `npm run test:mobile`.
*
* fileParallelism stays off for the same reason as every other config in this
* directory: these bind real ports and drive real tmux sessions, and two files
* doing that at once fail each other rather than the code.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: BROWSER_TEST_GLOBS,
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 60000,
teardownTimeout: 60000,
},
});
+9 -12
View File
@@ -1,13 +1,17 @@
import { resolve } from 'node:path';
import { defineConfig, configDefaults } from 'vitest/config';
import { NON_CI_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* CI test config — same as vitest.config.ts but EXCLUDES the browser-driven
* mobile suite (test/mobile/**). Those are Playwright visual-regression tests
* that need a live server + chromium + environment-specific PNG baselines, so
* they are run/maintained separately and are not part of the CI gate.
* The default gate — what `npm test` and CI both run.
*
* Same as vitest.config.ts but EXCLUDES the suites that cannot pass on an
* arbitrary machine: browser-driven (Playwright + chromium), visual-regression
* (per-machine PNG baselines) and wall-clock perf. Those are not unmaintained;
* they have their own runners (`test:browser`, `test:mobile`, `test:perf`).
* See config/test-suites.ts for the list and the reason behind each entry.
*
* Keep the rest in sync with config/vitest.config.ts.
*/
@@ -17,14 +21,7 @@ export default defineConfig({
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
exclude: [
...configDefaults.exclude,
'test/mobile/**', // browser/visual (Playwright + chromium)
'test/perf-*.test.ts', // timing-sensitive perf benchmarks (flaky in CI)
'test/inline-rename.test.ts', // browser (Playwright)
'test/opencode-resize.test.ts', // browser (Playwright)
'test/webgl-fallback.test.ts', // browser (Playwright)
],
exclude: [...configDefaults.exclude, ...NON_CI_TEST_GLOBS],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 30000,
+11
View File
@@ -3,6 +3,17 @@ import { defineConfig } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
/**
* EVERY test in the repo, including the ones that cannot pass on an arbitrary
* machine — `npm run test:all`. Reach for it when you want the complete picture
* and are prepared to read past environmental failures.
*
* This is NOT what `npm test` runs. On a machine without chromium, a free port
* or per-machine PNG baselines this config fails ~87 tests on a clean master,
* which makes it useless as a pass/fail signal: the default gate is
* config/vitest.ci.config.ts, and the suites it leaves out each have their own
* runner (`test:browser`, `test:perf`, `test:mobile`). See config/test-suites.ts.
*/
export default defineConfig({
test: {
root,
+25
View File
@@ -0,0 +1,25 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
import { PERF_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The wall-clock benchmarks `npm test` skips — `npm run test:perf`.
*
* These assert on durations, so run them on an otherwise idle machine: a loaded
* runner fails them for reasons that have nothing to do with the diff under
* test, which is exactly why they are not part of the default gate.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: PERF_TEST_GLOBS,
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 60000,
teardownTimeout: 60000,
},
});
+23 -4
View File
@@ -26,8 +26,8 @@ RUN apt-get update \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The agent CLIs (all four backends Codeman supports). Pinning is left to the
# rebuild cadence (see docs/docker-cases-plan.md, user-decision 2).
# The npm-published agent CLIs. Pinning is left to the rebuild cadence (see
# docs/docker-cases-plan.md, user-decision 2).
RUN npm install -g \
@anthropic-ai/claude-code \
@openai/codex \
@@ -35,6 +35,22 @@ RUN npm install -g \
opencode-ai \
&& npm cache clean --force
# Antigravity (`agy`) is NOT on npm — Google ships a standalone binary through its
# own installer, so it needs its own step. `--dir /usr/local/bin` is load-bearing:
# the installer's default target is `$HOME/.local/bin`, which at build time is
# root's home and would be unreachable by the `agent` user the container runs as.
# ⚠️ This binary is ~190MB on its own; it is the single largest layer in the image.
RUN curl -fsSL https://antigravity.google/cli/install.sh | bash -s -- --dir /usr/local/bin \
&& chmod 755 /usr/local/bin/agy \
&& agy --version
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot silently change how the
# other four CLIs install.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
@@ -50,10 +66,13 @@ ENV HOME=/home/agent
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir.)
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir;
# Antigravity nests its state inside `.gemini/antigravity-cli`, so it rides that seed.)
# `.pi/agent` IS pre-created: pi is seeded per-FILE (auth/settings/trust/models), and a
# per-file seed copy, unlike a whole-dir one, does not create its parent directory.
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions \
/home/agent/.claude/projects /home/agent/.codex/sessions /home/agent/.pi/agent \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
+759
View File
@@ -0,0 +1,759 @@
# Agent Control Plan: skill packaging + wait primitives
**Status**: steps 1 to 8 DONE and RELEASED. The wait primitives and the skill itself
(steps 1 to 5) shipped in **1.13.0**; the `codeman skill install` CLI, per-case injection
and `agentSkillEnabled` (step 6) shipped in **1.14.1** and were republished with fixes in
**1.14.2**. Steps 1 to 5 were multi-round verified on 2026-08-08, step 6 on 2026-08-09;
see [§7 Build log](#7-build-log-what-actually-happened) for what shipped, what each
verification round found, and the two items that genuinely remain open (§2.4's footgun
guard and the Part 3 deferrals).
**Date**: 2026-08-08
**Scope**: Part 1 (agent skill) and Part 2 (wait primitives) were specified and built.
Parts 3 to 5 are captured so they are not lost, but remain deliberately deferred.
---
## 0. Where this came from: what herdr does
[herdr](https://github.com/herdrdev/herdr) (Rust, Apache-2.0, ~25.8k stars) is a terminal
multiplexer built around AI coding agents. Relevant findings from the research pass:
| Capability | How herdr does it |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Agent state | Four states (`idle`, `working`, `blocked`, `done`) that roll up pane to tab to workspace in a sidebar |
| Detection | Lifecycle hooks where the agent supports them (it names Pi and MastraCode), otherwise TOML manifests matched against a live bottom-buffer snapshot. Bundled manifests plus remote updates from herdr.dev, local overrides win |
| Control API | Newline-delimited JSON over a Unix socket (`~/.config/herdr/sessions/<name>/herdr.sock`), `{"id":"req_1","method":"pane.split","params":{}}`, dot-notation methods, plus long-lived event subscriptions |
| Discoverability | `herdr api schema` prints a machine-readable schema |
| Agent skill | `npx skills add herdrdev/herdr --skill herdr -g`, a SKILL.md wrapping the CLI, guarded by `test "${HERDR_ENV:-}" = 1` so an agent outside a herdr pane refuses to act |
| Persistence | Background server, detach with `ctrl+b q`, snapshot restore of workspaces/tabs/panes/cwd/layout, experimental screen-history replay, agent resume via native session ids, live PTY handoff across server replacement |
| Plugins | `herdr-plugin.toml` manifest, actions, event hooks, plugin panes, link handlers, GitHub-topic marketplace index |
The commands the skill teaches the agent:
| Group | Commands |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| workspace | `workspace list`, `workspace create` |
| tab | `tab list --workspace <id>`, `tab create` |
| pane | `pane current`, `pane list`, `pane layout`, `pane split --current --direction right --cwd <path> --no-focus`, `pane run <id> "<cmd>"`, `pane wait-output <id> --match/--regex <p> --timeout <ms>`, `pane read <id> --source visible\|recent\|detection` |
| agent | `agent list`, `agent start <name> --kind <type> --pane <id>`, `agent prompt <name> "<text>" --wait --timeout <ms>`, `agent wait <name> --until <state> --timeout <ms>`, `agent send-keys`, `agent get`, `agent read` |
### The honest comparison
herdr and Codeman are not the same product. herdr is a local, keyboard-first multiplexer with
no server, no web UI, and no autonomy layer. Codeman is a server with a browser and mobile UI,
remote and Docker cases, respawn, Ralph, cron, and the orchestrator, none of which herdr has.
What herdr genuinely does better is being **callable by the agent running inside it**. For
Codeman that is a packaging problem plus one missing primitive, not an architecture problem.
---
## 1. Gap analysis
| herdr capability | Codeman equivalent today | Gap |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------- |
| `pane split` + `agent start` | `POST /api/quick-start`, `POST /api/sessions` | none, already there |
| `agent prompt` | `POST /api/sessions/:id/input` with `clientId`+`seq` exactly-once | no `--wait` |
| `pane read` | `GET /api/sessions/:id/output`, `GET /api/sessions/:id/terminal?full=1` | none |
| `agent list` / `agent get` | `GET /api/sessions`, `GET /api/sessions/unified`, `GET /api/status` | none |
| `agent wait --until <state>` | SSE only (`/api/events`) | **missing**, and SSE is impractical from a shell tool |
| `pane wait-output --match` | nothing | **missing** |
| Skill file | README section "Driving Codeman from an Agent" | **not packaged**, an agent will never find it |
| Env guard `HERDR_ENV=1` | `CODEMAN_MUX=1`, `CODEMAN_API_URL`, `CODEMAN_SESSION_ID` already exported at spawn | none, the guard variables exist |
| `blocked` state | hook events (`permission_prompt`, `elicitation_dialog`) plus CSS classes plus the phone overview NEEDS YOU section | not in the wire contract (`SessionStatus = 'idle' \| 'busy' \| 'stopped' \| 'error'`) |
| `api schema` | hand-written `docs/api-reference.md` | no machine-readable schema |
| Detection manifests | hardcoded in `usage-limit-patterns.ts`, `respawn-*-patterns`, `regex-patterns.ts` | patterns are code, not data |
| Plugin runtime | deliberately refused, see `docs/extending-codeman.md` | not a gap, a decision |
| Session handoff on restart | tmux owns the PTYs, so they already survive a Codeman restart | not a gap, solved by architecture |
**Conclusion**: roughly 90% of the capability surface already exists. Parts 1 and 2 below close
the two real gaps.
The table is the 2026-08-08 snapshot that motivated the work, kept as written. The three rows
marked missing are closed since: `GET .../wait` and `GET .../wait-output` shipped in 1.13.0, and
the skill is packaged at `skills/codeman` (npm tarball included). `blocked` as a wire-contract
state, and the machine-readable schema, are still open (Parts 3 and 4).
---
## 2. Part 1: the Codeman agent skill
### 2.1 Goal
An agent running inside a Codeman session can discover and correctly drive Codeman without the
user pasting API docs into the prompt, and without inventing dangerous calls.
### 2.2 Layout and distribution
The `npx skills` CLI (vercel-labs/skills) clones a GitHub repo and looks for
`skills/<name>/SKILL.md`. Claude Code natively discovers `.claude/skills/<name>/SKILL.md` in a
project and `~/.claude/skills/` globally. Both are satisfied with one source of truth plus a
symlink, which is the pattern this repo already uses for `remotion-best-practices`.
```
skills/
codeman/
SKILL.md <- single source of truth
reference/
endpoints.md <- full endpoint tables, loaded on demand
recipes.md <- worked multi-session orchestration examples
.claude/skills/codeman -> ../../skills/codeman (symlink, dogfooding in this repo)
```
Adding a `skills/` directory to the repo root costs one entry in the GitHub listing. CLAUDE.md
keeps the root short on purpose, so this needs a conscious sign-off; the alternative is
`docs/skills/codeman/` with a `--skill` path argument, which breaks the one-liner install.
**Recommendation**: accept `skills/` at the root, because the install one-liner is the whole
point of shipping a skill.
Install paths, in order of how a user gets it:
1. `npx skills add Ark0N/Codeman --skill codeman -g` (global, any agent, matches the herdr flow).
2. `codeman skill install [--global | --case <name>]`, a new CLI subcommand writing the same
file. This is the path for users who installed via npm and never cloned the repo.
3. **Automatic per-case injection**, modeled exactly on `applyStatusLineConfig(casePath, enabled)`
in `hooks-config.ts`: write `<case>/.claude/skills/codeman/SKILL.md` at case creation,
gated on a new setting. Codeman already writes `<case>/.claude/settings.local.json` hooks
through `writeHooksConfig()`, so this is the same mechanism with the same lifecycle.
Setting name: `agentSkillEnabled`. Synced (not per-device), since it changes on-disk case
content rather than display. Default: **ON after the dogfooding phase, OFF in the first
release**. Rationale for starting OFF: Claude Code loads every skill's name and description
into context on every turn, so an always-on skill has a small permanent token cost, and we
should measure that we are buying something with it first.
### 2.3 SKILL.md content
Frontmatter, per the skills convention (`name` + `description` required):
```yaml
---
name: codeman
description: >-
Control Codeman, the session manager this agent is running inside: list sessions,
start worker sessions, send prompts, read terminal output, and wait for other agents
to finish. Only usable when CODEMAN_MUX=1.
---
```
Body sections, in order:
**1. Guard (first thing, non-negotiable).**
```bash
test "${CODEMAN_MUX:-}" = 1 || { echo "not inside a Codeman session"; exit 1; }
API="${CODEMAN_API_URL:?CODEMAN_API_URL not set, refusing to guess}"
SELF="${CODEMAN_SESSION_ID:-}"
```
If `CODEMAN_MUX` is not `1`, the agent must stop and say it is not running inside a
Codeman-managed session. Same shape as herdr's `HERDR_ENV` guard, and the variables are
already exported by `tmux-manager.buildEnvExports()`. No fallback URL when
`CODEMAN_API_URL` is unset: any guess is the wrong scheme on an HTTPS install (prod is
HTTPS with a self-signed cert, hence `curl -sk` throughout), and a server the agent
cannot identify is not one it should be driving.
**2. Rules of the road.** Lifted and tightened from README lines 666 to 745:
- Single-line input only. Multi-line breaks the agent TUI (Ink).
- Always send `clientId` + a monotonic `seq` on `POST .../input` so a retry cannot double-deliver.
- Envelope is `{success, data}`; a few legacy GETs are bare, so read `body.data ?? body`.
- Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set. Prod is HTTPS, so `curl -sk`.
- Prefer `/api/v1/*`, the stable alias.
**3. Safety rules (the section that does not exist anywhere today).**
- Never act on `$CODEMAN_SESSION_ID`. That is you.
- Only `DELETE` sessions **you created in this conversation**, by exact id. Keep the list.
- Never bulk-delete, never loop a `DELETE` over `/api/sessions`. There is no undo.
- Never `tmux kill-session`, `pkill tmux`, `pkill claude`. Use the API.
- Creating a session consumes a slot against the 50-session cap. Clean up what you start.
**4. Recipes**, each one a single copy-pasteable curl:
| Task | Call |
| -------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| list sessions | `GET /api/v1/sessions` |
| find yourself | match ids by PREFIX of `$CODEMAN_SESSION_ID` (Docker cases truncate it to 8 chars, so an equality check never fires there) |
| start a worker | `POST /api/v1/quick-start {caseName, mode, effort}` |
| send a prompt | `POST /api/v1/sessions/:id/input {input:"…\r", useMux:true, clientId, seq}` (the trailing `\r` is what sends Enter; without it the text sits on the prompt unsubmitted) |
| send prompt and wait | `POST /api/v1/sessions/:id/input {input:"…\r", wait:"stop", waitTimeout:600000}` (Part 2) |
| wait for a worker | `GET /api/v1/sessions/:id/wait?until=stop,blocked&timeout=300000` (Part 2) |
| wait for a marker | `GET /api/v1/sessions/:id/wait-output?match=DONE_<random>&timeout=120000` (Part 2; unique per call, per §3.3's repaint rule) |
| read output | `GET /api/v1/sessions/:id/output` |
| read full scrollback | `GET /api/v1/sessions/:id/terminal?full=1` |
| watch sub-agents | `GET /api/v1/subagents` |
| schedule work | `POST /api/v1/cron/jobs` |
| clean up | `DELETE /api/v1/sessions/:id` |
**5. Pointer to `reference/endpoints.md`** for anything not in the table, so the always-loaded
part of the skill stays small.
### 2.4 An ergonomics guard worth adding server-side
The skill will tell the agent not to act on itself, but a confused agent can still try. Propose:
the skill sends `X-Codeman-Caller-Session: $CODEMAN_SESSION_ID` on every request, and the server
refuses destructive operations (`DELETE /api/sessions/:id`, kill, respawn stop) when that header
equals the target id, with a clear error.
This is a **footgun guard, not a security control**: any caller can omit the header. Document it
as such so nobody mistakes it for a boundary. It costs about 10 lines in `route-helpers.ts`.
### 2.5 Verification
Per the always-end-to-end-test rule, "the skill exists" is not done. Done is:
1. Symlink it into `.claude/skills/`, start a real throwaway Codeman session, and ask that agent
to "start a worker session that runs the test suite and tell me when it finishes".
2. Confirm from the outside that exactly one new session appeared, got the prompt, and that the
lead agent waited rather than polling in a busy loop.
3. Confirm the guard: run the same prompt in a shell with `CODEMAN_MUX` unset and confirm refusal.
4. Confirm cleanup: the worker session is deleted by exact id and no other session was touched.
Never run this against `w1`/`w2`/`w3`.
### 2.6 Files touched
- `skills/codeman/SKILL.md` (new), `skills/codeman/reference/*.md` (new)
- `.claude/skills/codeman` symlink (new)
- `src/cli.ts` (new `skill install` subcommand)
- `src/hooks-config.ts` (new `applyAgentSkill(casePath, enabled)`, mirroring `applyStatusLineConfig`)
- `src/web/schemas.ts` (`agentSkillEnabled` in `SettingsUpdateSchema`, which is `.strict()`)
- `src/web/routes/system-routes.ts` (settings PUT must resolve the flag from `merged`, never
from the raw body, per the partial-PUT invariant)
- `src/web/public/settings-ui.js` + `index.html` (checkbox)
- `package.json` `files` array, so `skills/` ships to npm
- README pointer, `docs/extending-codeman.md` seam 3 pointer
---
## 3. Part 2: wait primitives
### 3.1 Goal
Make Codeman orchestratable from a shell tool. Today the only "tell me when" channel is SSE,
which a curl-driven agent cannot practically consume: it would have to hold a streaming
connection and parse events inline. herdr solves this with blocking CLI calls. Codeman should
solve it with bounded long-poll endpoints.
All three additions are **additive**, so the versioning policy stays intact (new endpoints and
new optional fields are non-breaking).
### 3.2 The signal model
A waiter resolves on the first of a set of signals. Sources that already exist:
| Signal | Source today |
| --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
| `idle` | `Session` emits `idle` (session.ts ~1775 for Claude, ~2101 for shell), wired at `session-listener-wiring.ts:402` |
| `working` | `Session` emits `working` (session.ts ~1788), wired at `session-listener-wiring.ts:401` |
| `stop` | `POST /api/hook-event` with `event: 'stop'`, the definitive "Claude finished responding" signal already used by `controller.signalStopHook()` |
| `blocked` | `POST /api/hook-event` with `permission_prompt` or `elicitation_dialog` |
| `exit` | `Session` emits `exit` |
`stop` is the highest-quality signal for "the turn is over" and should be the documented default
for orchestration. `idle` is heuristic: output stabilization plus prompt detection, and it can
flap mid-turn when a spinner pauses. External CLI modes (`isExternalCliMode()`) have no stop
hook at all, so for opencode/codex/gemini/antigravity only `idle`, `working` and `exit` are
available. **The skill and the docs must say which signals exist per mode**, otherwise an agent
waits forever on `stop` in a codex session.
### 3.3 Endpoint specs
#### A. `GET /api/sessions/:id/wait`
| Param | Type | Default | Notes |
| --------- | ---------------------------------------------- | ---------------- | ------------------------------------------------------------ |
| `until` | comma list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on first match |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` (600000) |
| `fresh` | `0`/`1` | `0` | `1` requires a _transition_, ignoring the state at call time |
Response (always 200 unless the session is missing or a cap is hit):
```json
{
"success": true,
"data": {
"signal": "stop",
"timedOut": false,
"immediate": false,
"ended": false,
"waitedMs": 8421,
"status": "idle",
"sessionId": "...",
"until": ["stop", "idle", "exit"],
"limitPaused": false
}
}
```
`until` is echoed back because the server may narrow it: `stop`/`blocked` are dropped
from the DEFAULT set for external CLI modes (asking for them EXPLICITLY is a 400
instead, since omitting `until` must never 400). `limitPaused` tells a caller that a
timeout was expected rather than a stall worth retrying hard.
**A timeout is not an error.** `{"timedOut": true, "signal": null}` with HTTP 200, so a caller
can loop without treating every poll boundary as a failure. Errors are reserved for
`NOT_FOUND` (unknown or not-owned session) and `SESSION_BUSY` (waiter cap exceeded).
`immediate: true` means the session was already in the requested state and `fresh` was not set.
#### B. `GET /api/sessions/:id/wait-output`
| Param | Type | Default | Notes |
| --------- | ------------------------------ | -------- | --------------------------------------------------------- |
| `match` | literal string, 1 to 200 chars | required | substring match against ANSI-stripped output |
| `nocase` | `0`/`1` | `0` | case-insensitive compare |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the existing text buffer first, then waits |
| `timeout` | ms | 60000 | clamped to `MAX_WAIT_MS` |
Response: `{ matched: true, timedOut: false, snippet: "...", waitedMs }`.
**No regex in v1, deliberately.** `search-service.ts` already avoids regex specifically so there
is no ReDoS surface, and this endpoint would be even more exposed since the pattern is attacker
supplied and the input is a live stream. herdr can offer `--regex` because Rust's regex crate is
linear-time with no backtracking; JS `RegExp` is not. If regex is wanted later, the honest
options are a length-capped subset compiled once with a match budget, or `re2`. Note it and move on.
Implementation detail that will bite if missed: a match can straddle two PTY chunks. Keep a
carry buffer of `match.length - 1` bytes from the previous chunk and test `carry + chunk`.
⚠️ **`from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, resize, or any TUI redraw, and a repaint arrives as ordinary `terminal`
data. Observed live: a marker echoed a minute earlier matched instantly on a fresh
`from=now` wait. This is inherent to a terminal multiplexer, not fixable in the registry,
so the contract is: **use a marker unique per call** (`echo DONE_$RANDOM`), never a
generic one like `BUILD OK`. The skill's recipes must show that.
The returned snippet is whitespace-collapsed (blank runs to a single newline) for
readability only; matching runs on the raw stripped text. Without it, a real pane's
`\r\n` padding between the prompt and the match fills the whole context window with
nothing, which was the first thing the live test showed.
#### C. `wait` on the existing input endpoint
`POST /api/sessions/:id/input` gains two optional fields:
```json
{ "input": "run the tests\r", "useMux": true, "clientId": "agent-1", "seq": 7, "wait": "stop", "waitTimeout": 600000 }
```
(The trailing `\r` is required on every input body: `sendInput` sends Enter only
when the input contains a carriage return.)
Response gains `"wait": { "signal": "stop", "timedOut": false, "waitedMs": 41230 }`.
This is the important one, because it closes a race the standalone `GET .../wait` cannot: between
"input delivered" and "session flips to working" there is a window where a naive
send-then-wait sees the _pre-existing_ idle state and returns instantly. The combined endpoint
**registers the waiter before writing**, so that window does not exist. This is exactly why herdr
ships `agent prompt --wait` as its own thing.
`wait` accepts `true` (the default signal set) or the same comma grammar as `until`.
Both new fields are `.nullish()`, not `.optional()`: a third-party caller building the
body with `JSON.stringify` keeps an explicit `null` on the wire, and `.optional()`
rejects that with `INVALID_INPUT`. That gotcha has shipped as a real bug twice.
Two behaviors to preserve carefully:
- **`useMux` is fire-and-forget today.** The handler responds without awaiting `writeViaMux`, on
purpose (a tmux child process must not block the HTTP response). With `wait` present the
handler already has to stay open, so it can await delivery, and a `writeViaMux` failure becomes
observable for the first time. The non-wait path must keep its current fire-and-forget shape
byte for byte.
- **Duplicate suppression.** A tagged redelivery (`clientId`+`seq` already applied) returns 200
without writing. With `wait` set it still waits, since the caller's intent is "tell me when
this settles". But it waits with `requireTransition: false`, unlike a fresh delivery: the
original turn may be long over, and requiring a new transition would block a redelivery until
timeout for no reason. Fresh delivery requires a transition, a duplicate answers from the
current state.
- **Capacity rollback.** `shouldApplyInput()` MUTATES (it records the seq), and it runs before
the waiter is registered. If registration then fails on a full pool, the handler must call
`forgetInputSeq` before returning `SESSION_BUSY`, or the caller's retry is rejected as a
duplicate and the input is lost by the very mechanism reliable delivery exists for.
### 3.4 Module design
New file `src/web/session-wait-registry.ts`, with the IO-free core unit-testable in isolation
(same split as `self-update.ts`):
```ts
type WaitSignal = 'idle' | 'working' | 'stop' | 'blocked' | 'exit';
waitForSignal(sessionId, { until: Set<WaitSignal>, timeoutMs, requireTransition }): Promise<WaitResult>
notifySignal(sessionId, signal: WaitSignal): void
waitForOutput(sessionId, { match, nocase, timeoutMs }): Promise<OutputWaitResult>
notifyOutput(sessionId, chunk: string): void
cancelAll(sessionId, reason): void
```
Wiring points, all existing:
- `src/web/session-listener-wiring.ts` around lines 190 and 200 already handles `working` and
`idle` and broadcasts them. Add a `notifySignal()` call next to each broadcast, plus `exit`.
- `src/web/routes/hook-event-routes.ts` already switches on `event` for the respawn controller.
Add `notifySignal(sessionId, 'stop' | 'blocked')` in the same switch.
- Output: `notifyOutput()` rides the ALREADY-attached `terminal` listener in
session-listener-wiring.ts. An earlier draft had the registry hand out attach/detach
callbacks so a listener could be added lazily; that was deleted once it was clear no
second listener is needed at all. The cost is one Map lookup per PTY chunk, which is why
the no-waiter check comes before the ANSI strip.
- Session deletion calls `notifySignal('exit')` then `cancelAll()`, so no promise is left
hanging. Both are required: `_doCleanupSession` detaches the session's listeners BEFORE
`session.stop()`, so on a delete the PTY exit event never reaches the registry, and an
`until=exit` caller would otherwise get a bare `ended` instead of its signal. Found by
live-testing the delete path, not by the unit tests.
Memory-leak discipline, per the 24-hour-session rules: every waiter owns a timer that is cleared
on resolve, the per-session waiter set is deleted when it empties, and the output listener is
removed with it. `test/memory-leak-prevention.test.ts` should grow a case for this.
Caps in a new `src/config/agent-wait.ts` (limits live in `src/config/`, env-overridable):
| Constant | Default | Why |
| ------------------------- | ------- | --------------------------------------- |
| `MAX_WAIT_MS` | 600000 | an unbounded long-poll is a socket leak |
| `DEFAULT_WAIT_MS` | 60000 | short enough to survive most proxies |
| `MAX_WAITERS_PER_SESSION` | 16 | |
| `MAX_WAITERS_TOTAL` | 128 | same reasoning as `MAX_SSE_CLIENTS` |
Exceeding a cap returns `SESSION_BUSY`, not a silent queue.
### 3.5 Transport concerns
Fastify is constructed with defaults in `server.ts:329-331`. `requestTimeout` defaults to 0
(disabled) and `keepAliveTimeout` (72s) applies between requests, not to an in-flight one, so a
10-minute in-process hold is fine. **Verify this on the real instance before relying on it.**
Intermediaries are the actual risk. Prod is reached through `tailscale serve`, and users also run
cloudflared tunnels; both can cut an idle connection. That is why `DEFAULT_WAIT_MS` is 60s and
why the documented pattern is a client-side loop over short waits rather than one 10-minute call.
The skill's recipes must show the loop.
### 3.6 Edge cases to get right
| Case | Behavior |
| ------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Session already idle, `fresh=0` | return immediately, `immediate: true` |
| Session already idle, `fresh=1` | wait for the next transition into a requested state |
| Session dies mid-wait | resolve with `signal: "exit"` if `exit` was requested, otherwise resolve `timedOut:false, signal:null, ended:true`. Never hang |
| Session deleted mid-wait | same, resolve, do not throw. Verified live: `until=exit` gets `signal:"exit"`, a concurrent `until=blocked` gets `ended:true`, both in ~0ms |
| Shutdown with a wait pending | `cancelEverything()` in `stop()`. Verified live: SIGTERM with a 300s wait in flight exits in 1s |
| External CLI mode | `stop` and `blocked` never fire. Reject `until=stop` for those modes with a clear `INVALID_INPUT` rather than hanging until timeout |
| Multi-user | goes through `findSessionOrFail(ctx, id, req)`, which already enforces ownership |
| Remote / Docker cases | signals originate from the same `Session` object, so no special casing. Docker hooks need `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for `stop`/`blocked` to arrive at all; without it, only `idle` works. Document it |
| Respawn `/clear` mid-wait | a respawn cycle emits `idle`. Callers waiting on `stop` are unaffected; callers on `idle` may resolve early. Documented, not fixed |
| Limit pause | if the session is paused on a usage limit, nothing will fire until the reset. The wait times out honestly. Consider surfacing `limitPaused: true` in the response so the caller can back off |
### 3.7 Tests
- `test/session-wait-registry.test.ts` (pure): immediate resolve, transition-required, multi-signal
first-wins, timeout, cap exceeded, cancel on session end, no listener leak after resolve,
chunk-straddling output match, case-insensitive match.
- `test/routes/session-wait-routes.test.ts` (`app.inject()`, no port): all three endpoints against
a `MockSession`, including the 200-with-`timedOut` contract and the ownership 404.
- `test/routes/session-input-wait.test.ts`: the send-and-wait race, plus proof that the non-wait
path is unchanged (still returns before `writeViaMux` settles).
- Live verification on a throwaway session before COM, per the always-end-to-end-test rule.
### 3.8 Files touched
- `src/config/agent-wait.ts` (new)
- `src/web/session-wait-registry.ts` (new)
- `src/web/session-listener-wiring.ts` (notify on idle/working/exit)
- `src/web/routes/hook-event-routes.ts` (notify on stop/blocked)
- `src/web/routes/session-routes.ts` (two new routes, `wait` fields on input)
- `src/web/schemas.ts` (`SessionWaitQuerySchema`, `SessionWaitOutputQuerySchema`, extend
`SessionInputWithLimitSchema`. Note: `.optional()` rejects `null`, so the frontend and any
generated client must send `undefined`, never `null`)
- `docs/api-reference.md`, `docs/extending-codeman.md`, README API table
- `skills/codeman/SKILL.md` recipes (Part 1 depends on this)
---
## 4. Deferred: parts 3 to 5
Not in scope now, kept here so they are not lost.
### Part 3: promote `blocked` to a first-class state
`SessionStatus` is `'idle' | 'busy' | 'stopped' | 'error'`. "Needs you" exists three times over:
hook events, the `tab-alert-action` CSS class, and the phone overview NEEDS YOU section, each
re-deriving it. herdr makes `blocked` a real state that rolls up.
Add `blocked` (and possibly `done`) to `SessionStatus`, set it from the same hook events that
Part 2 uses as wait signals, and clear it on the next `working`/`stop`. Then the tab strip, the
mobile overview, the wait endpoints, and any external agent read one field.
Cost: `SessionStatus` is a widely-consumed union, so every exhaustive `switch` (the codebase has
`assertNever` and `noFallthroughCasesInSwitch`) will need a branch. That is a feature, it makes
the compiler find every site. This is a **minor** bump, not a patch: it widens a public type in
the HTTP contract.
### Part 4: `GET /api/schema`
herdr ships `herdr api schema`. Every Codeman route is already Zod-validated, so
`zod-to-json-schema` over `schemas.ts` gives a self-describing API almost free. Value: third-party
tools and the skill stop drifting from hand-written docs. Open question: whether to emit full
OpenAPI (`@fastify/swagger` would need per-route schema registration, which is a much larger
change) or just dump the Zod schemas keyed by name (cheap, 80% of the value).
### Part 5: detection manifests instead of hardcoded patterns
CLI-specific readiness, blocked and usage-limit patterns live in code across
`usage-limit-patterns.ts`, the respawn pattern helpers and `regex-patterns.ts`. Externalizing the
per-CLI ones into data files would make adding a sixth CLI a data change instead of a code change.
**Do not copy the remote-update part.** herdr auto-fetches manifest updates from herdr.dev.
Codeman auto-pulling behavioral rules from a vendor server contradicts its security posture.
Bundled manifests plus local override only, no network.
### Explicit non-goals
- **Plugin runtime and marketplace.** `docs/extending-codeman.md` already argues this: a plugin
runtime means third-party code inside a process that spawns agents with your credentials, on a
server people expose over a tunnel. The reasoning still holds. If the marketplace _pattern_ is
wanted, apply it to data (web tabs, case templates, cron recipes), never to executable code.
- **Live PTY handoff on restart.** herdr needs it because it owns the terminals. Codeman
delegates to tmux, so PTYs already survive a self-update restart.
- **Socket API.** HTTP plus SSE is the existing, documented, stable contract. A second transport
would double the surface for no capability gain.
---
## 5. Sequencing
| Step | Work | Gate |
| ---- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 ✅ | `src/config/agent-wait.ts` + `session-wait-registry.ts` + unit tests | 48 tests green |
| 2 ✅ | `GET .../wait` + wiring in listener-wiring, hook-event-routes, server teardown | 15 route tests green; live-verified on an isolated `CODEMAN_INSTANCE=waittest` instance (immediate resolve, 400 on a bad signal, 200+`timedOut` on timeout, hook `stop` and `permission_prompt`→`blocked` waking an in-flight wait, delete delivering `exit`, SIGTERM not blocked); full `test:ci` sweep green |
| 3 ✅ | `GET .../wait-output` | 16 route tests green; live-verified on real PTY bytes (`echo MARKER` waking a blocked request in ~1s, `from=buffer` immediate hit, never-seen marker timing out at exactly 2001ms, nocase, `regex` refused with a 400); full `test:ci` sweep green |
| 4 ✅ | `wait` field on `POST .../input`, non-wait path proven unchanged | 16 route tests green; live-verified (no-wait returns in 26ms with the historical bare body; an idle session did NOT satisfy a `wait` request, blocking the full 2001ms, which is the race the endpoint exists to close; the stop hook resolved a send-and-wait at 1510ms and the input was confirmed in the tmux pane; `wait:null` accepted) |
| 5 ✅ | `skills/codeman/SKILL.md` + reference files + `.claude/skills` symlink | live dogfood: a real session orchestrates a worker end to end |
| 6 ✅ | `codeman skill install` CLI + `applyAgentSkill()` + `agentSkillEnabled` setting | 10 unit tests (`test/agent-skill.test.ts`) + real-server case-creation tests (`test/quick-start.test.ts`, incl. the settings PUT accepting the key) green; CLI verified live (install/uninstall, global + `--case`, foreign/symlink refusals) |
| 7 ✅ | Docs: api-reference, extending-codeman, README | plus `architecture-invariants.md` (§agent-wait-primitives), `CLAUDE.md` and the API reference's per-mode signal table |
| 8 ✅ | COM (minor bump: new endpoints, new setting, new optional fields) | released as 1.13.0 (wait primitives + skill); step 6 followed in 1.14.1 and was republished as 1.14.2 after live-testing the packaged skill |
Parts 1 and 2 are independent enough to land separately, but the skill is much less useful
without the wait endpoints, so the wait work goes first.
## 6. Open questions for the owner
1. ✅ `skills/` at the repo root: accepted (built that way; the install one-liner depends on it).
2. ✅ `agentSkillEnabled` default: **OFF** for the first release, per §2.2's rationale (skills
cost context on every turn; measure before defaulting on). Flip later if dogfooding earns it.
3. ✅ Both: global install via `npx skills add` / `codeman skill install`, AND per-case
auto-injection behind the (default-off) setting. Injection is add-only at session create and
marker-guarded, so a user-authored copy is never touched.
4. Is `X-Codeman-Caller-Session` self-protection worth the 10 lines, given it is a footgun guard
and not a security boundary? (Still open, not built with step 6.)
5. ✅ Regex support in `wait-output`: literal-only shipped, and a `regex` query param is
rejected with a 400 rather than ignored, so an agent that assumed otherwise cannot
silently wait on the wrong thing.
---
## 7. Build log: what actually happened
Written at the end of the build so the next person inherits the reasoning, not just the
diff. Process artifacts (per-agent briefs, findings, reports) live in the gitignored
`tmp/agent-wait-review/`; this section is the part worth keeping.
### What shipped
| Piece | Files |
| ------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Bounds + clamping | `src/config/agent-wait.ts` (new) |
| Blocking-wait registry | `src/web/session-wait-registry.ts` (new, IO-free, unit-tested) |
| `GET .../wait`, `GET .../wait-output`, `wait`/`waitTimeout` on `POST .../input` | `src/web/routes/session-routes.ts` |
| Signal wiring | `session-listener-wiring.ts` (idle/working/exit + output), `hook-event-routes.ts` (stop/blocked), `server.ts` (teardown, shutdown) |
| Agent skill | `skills/codeman/SKILL.md` + `reference/`, `.claude/skills/codeman` symlink, `package.json` `files` |
| Docs | `api-reference.md`, `extending-codeman.md`, `architecture-invariants.md`, `README.md`, `CLAUDE.md` |
| Tests | `test/session-wait-registry.test.ts`, three `test/routes/session-*wait*.test.ts`, `http-contract.test.ts`, `mock-session.ts` |
### Bugs found in ADJACENT code, not in the new feature
These are the highest-value output of the exercise and none were on the plan:
1. **Every Codeman hook was dead on HTTPS installs.** `hooks-config.ts` built the hook
curl as `curl -s` with no `-k` while the statusline exporter 300 lines below used
`curl -sk` and documented why. Proven with the real hook command: `curl exit=60`
without the flag, success with it, and the failure swallowed by the hook's own
`2>/dev/null || true`. This silently killed `stop`, `permission_prompt`,
`elicitation_dialog`, `idle_prompt`, `teammate_idle` and `task_completed`, taking
respawn's definitive idle signals with them. Fixed, **plus** a staleness detector in
`refreshStaleCodemanHooks` that regenerates the on-disk config of already-created
cases (23 of 26 local cases carried the broken form; fixing the generator alone would
have left every one of them broken).
2. **`buildEnvExports()` exported a wrong-scheme `CODEMAN_API_URL`** (`http://` fallback
on an HTTPS install). Now omitted rather than guessed, so in-session guards fail closed.
3. **Programmatic input is only submitted when it contains `\r`.** `sendInput` sends Enter
only if the payload has a carriage return; without it the text sits in the composer
forever. Bit this build repeatedly before it was diagnosed, and had leaked into the
docs' own examples.
### Design decisions worth not re-litigating
- **A timeout is HTTP 200** with `wait.timedOut`, never a 4xx: callers loop over short
waits because tunnels cut idle connections, and every poll boundary would otherwise be
indistinguishable from failure.
- **Send-and-wait must be one endpoint.** A separate POST-then-wait races: between the
write and the flip to `working`, a wait sees the stale `idle` and reports the PREVIOUS
turn as this one. The waiter is registered before the write.
- **`stop`/`blocked` exist for `claude` mode only.** They come from Claude Code hooks;
`shell` installs none either, so keying off `isExternalCliMode()` was wrong.
- **Literal matching only, never regex.** JS `RegExp` backtracks; herdr can offer
`--regex` because Rust's regex crate is linear-time.
- **Client-hangup abort listens on `reply.raw` guarded by `writableFinished`.** On
`req.raw`, `close` fires when the request BODY ends, which on a POST killed every
send-and-wait instantly, and no `app.inject()` test can see it (inject never emits
`close`).
- **Liveness cannot come from `session.pid`.** For a tmux session that is the local
`tmux attach` client, not the worker: a worker exiting inside its pane leaves
`pane_dead=1` with the client alive, so `pid` never goes null. Liveness is probed at
the mux layer, cached (~750 ms) and only on blocking waits, never on the input hot path.
### Verification rounds
Six agents across three rounds, each verifying the previous round's work rather than its
own. Findings that mattered, in order of severity, were: the dead-pane liveness gap; the
`reply.raw` abort regression; abandoned long-polls leaking waiter slots; a crashed session
reporting `idle`; `shell` accepting `until=stop`; and a documented recipe that reported
success without running its task. Two traps recurred often enough to name:
- **Vacuous passes.** `app.inject()` never emits `close`; a latched `cancelEverything()`
in `afterEach` silently killed the registry for every later test in a file; three test
files sharing one session id against the process-wide registry let one file's leftover
waiter fail another's assertion. Any new wait test needs care on all three.
- **HTTP-only test instances.** Every isolated instance used during the build was plain
HTTP, which is exactly why the HTTPS hook bug survived so long. Test the transport the
user actually runs.
### Resolved at wrap-up (2026-08-08, conclusion pass)
- **R2-A**: the fire-and-forget-then-gather-sequentially pattern was **removed from
the skill** rather than patched. Signals are edge-triggered with no history, so a
`stop` that fires before its waiter registers is unobservable afterwards; a
`fresh=0` gather was rejected because the only `until` set that current state can
satisfy answers `idle` for a prompt that never submitted, resurrecting the exact
false-success failure R2-B had just closed. Flow 3b's pattern B now gathers on
latched `wait-output` markers (`from=buffer`), the same mechanism that makes the
shell flows reliable; the limitation is recorded in
`architecture-invariants#agent-wait-primitives` and `endpoints.md`. The durable
fix, a latched last-signal-per-turn on the server, stays with deferred Part 3.
- Docs F7/F8, F4 and the false-`idle` attribution: `api-reference.md`,
`extending-codeman.md` and `architecture-invariants.md` rewritten to the post-fix
matcher (one normalized stream, chunk-straddling found, snippet as a rendering of
the matched window), the real no-PTY answer (`ended:true`, `aborted:false`,
`delivered:false`), and the startup-idle mechanism (a session parked on the trust
dialog emits no further `idle`; the false success is the startup transition).
- Orchestrate #12, #5/R2-B, #6, and R2-C..R2-E: fire-and-forget's empty `data`
documented; every send-and-wait retry loop now treats `duplicate:true` +
`immediate:true` as "no new turn ran" and reads the terminal before believing it;
claude fan-out is pattern A (backgrounded send-and-waits) or the marker gather;
readiness budgets rebalanced (5 s stage 1, 45 s stage 3) with the virgin-case
floor named; the auth fallback now also reads the supervisor definition
(`codeman-web.service` / launchd plist) and accepts `export`-prefixed `.env`
lines; `pid != null` is documented as startup-only, never liveness.
- Both public readiness recipes (extending-codeman.md, README) are bypass-first with
the trust probe as the bounded fallback; the worked recipe carries `-k` and fails
loudly on an empty SID; the hook `-k`/self-heal fix appears in every
"hooks go missing" list; the multi-word-TUI claim is "unreliable", not "never".
### Still open
Both release-checklist items that used to sit here are done: `skills/` is tracked and
ships through `package.json` `files` (published with 1.13.0, republished with 1.14.2),
and the changeset was consumed, committed and deployed. What is left:
- Deferred with Part 3: the latched last-signal-per-turn. Nice-to-haves from the
reviews: N2 (create the death-watcher inside its `try`, still built one line above
it in `GET .../wait`) and converting timeout-shaped test detections into fast
assertions.
- §2.4's `X-Codeman-Caller-Session` footgun guard: still not built (open question 4).
### Step 6 (2026-08-09): install command, per-case injection, the setting
Built to the §2.6 file list, mirroring the statusLine mechanism throughout:
| Piece | Where |
| ----- | ----- |
| `applyAgentSkill(casePath, enabled)` + `installAgentSkillInto` / `removeAgentSkillFrom` | `src/hooks-config.ts` |
| `codeman skill install` / `skill uninstall` (`--global` default, `--case <name>`) | `src/cli.ts` |
| `agentSkillEnabled` (SYNCED, default OFF) | `schemas.ts` (`SettingsUpdateSchema`), `getAgentSkillEnabled()` on `ConfigPort`/`server.ts`, checkbox in `index.html` + `settings-ui.js` |
| Injection call sites (Claude mode only) | `POST /api/sessions` next to `refreshStaleCodemanHooks`; `POST /api/quick-start` after the case-create/self-heal blocks (local + docker cases; remote skipped, its path lives on another host) |
| Tests | `test/agent-skill.test.ts` (10 unit), `test/quick-start.test.ts` (real server: default-off, PUT accepts key, injection on create, shell-mode skipped) |
Decisions worth keeping:
- **Ownership marker, prefix-matched.** The injected SKILL.md ends with
`<!-- codeman-managed-agent-skill: … -->`; install/refresh/remove all refuse a copy
without the marker (a user's own skill) and match on the PREFIX so a wording change
cannot disown older injected copies (the `BACKGROUND_WAKE_MARKER_PREFIX` pattern).
- **Symlink refusal.** This repo's own dogfooding layout
(`.claude/skills/codeman -> ../../skills/codeman`) means the injector must `lstat`
the skill dir AND its `skills/` parent and bail on a symlink, or enabling the
setting in the Codeman repo itself would overwrite the skill source through the link.
- **ADD-ONLY at session create**, same shared-`.claude` rationale as the statusLine:
a create while the setting is off must not yank the skill out from under other live
sessions in the repo. The remove path exists (CLI `skill uninstall`, tests); no
automatic sweep removes on toggle-off.
- **Removal is manifest-based, never `rm -rf`**: only files the packaged source would
have written are deleted, directories are pruned bottom-up only if they emptied, so
a user's extra notes in `reference/` survive an uninstall.
- **Source resolution**: `join(moduleDir, '..', 'skills', 'codeman')` works from
`src/` (tsx), `dist/` (tsc build), and the npm tarball alike, because all three sit
one level below the package root and `files` ships `skills/`.
- **Nothing acts on the setting at PUT time**: injection reads the merged persisted
settings at session create (`readSettings`, ~2s cache), so the partial-PUT invariant
(`toggleService` reading `merged`) is untouched by construction.
### 2026-08-09 addendum: cross-session messaging folded into the skill
Claude Code 2.1.224+ ships cross-session messaging: `ListAgents`/`SendMessage`
tools, a per-session Unix inbox socket, and a registry in
`~/.claude/sessions/<pid>.json`. Codeman's claude workers are ordinary local Claude
Code sessions, so the skill now routes task delivery and result collection over it
when available, while the HTTP primitives keep spawn, readiness, synchronization,
liveness and delete. New `skills/codeman/reference/messaging.md` (ships with zero
installer changes: `readAgentSkillSource()` enumerates `reference/*.md` from disk),
Flow 5 in recipes.md, and §4 in SKILL.md.
Verified live (claude-cli 2.1.226, Linux):
- A message to an idle worker starts a turn and that turn fires the normal `stop`
hook (8.3 s send-to-stop measured), so the HTTP wait primitives compose with
messaging unchanged; delivery to a busy session lands between tool calls.
- First contact needs the `name [ref]` form; the bare name errors with the exact
string to resend. The `uds:` reply address of an inbound message works as a `to`.
- The `tmux codeman-<id8>` column in `ListAgents` (and the registry's `tmux` field)
is the join key to Codeman session ids. The registry's `sessionId` field starts as
the Codeman id (we spawn `claude --session-id <id>`) but drifts after `/clear` or
resume, so it must never be the join key.
- The feature is flag-gated beyond the version: two 2.1.226 sessions on one machine,
one with an inbox socket and one without. Absence is a fallback case, not an error.
- Codeman's default `--dangerously-skip-permissions` spawn puts both ends in the
bypassing class, which delivers; mixed classes hold behind an approval dialog that
expires unattended (upstream default 5 min), which on a headless worker means the
message silently dies. The skill's backstop covers it.
Follow-up, landed in the same PR: local claude spawns now pass
`--name <session name>` so peers carry Codeman session names. The gate is
`buildNameCliArgs()` (session-cli-builder.ts), fail-closed at
`CLAUDE_NAME_FLAG_MIN_VERSION = 2.1.224`: that is the messaging release, the flag's
presence there was verified against the installed 2.1.224 binary, and the version
comes from `getClaudeCliVersion()` (null on probe failure and under vitest), so an
older or unknown CLI gets a command byte-identical to before. That matters because
claude aborts startup on an unknown option, which would kill every session spawn.
The value is allowlist-sanitized (Unicode letters/digits plus ` ._:-`, leading
dashes stripped so it cannot parse as another option, 64-char cap, empty result =
flag omitted) before the double-quoted interpolation in `buildSpawnCommand`, and
only the LOCAL command carries it: the docker/remote builders never see it, since
their CLI is not the binary the probe measured. E2E on an isolated instance
(`CODEMAN_INSTANCE`): process cmdline `claude ... --name w9-msgtest`, registry
`name: "w9-msgtest"`, `ListAgents` lists it under that name, a message round-trip
works, and its replies arrive tagged `from-name="w9-msgtest"` (a derived-name
worker's replies carry no `from-name`). A quick-start without `sessionName` has an
empty Codeman name, so the peer name stays derived: agents should name their
workers. Tests: `test/name-flag-injection.test.ts`.
+490
View File
@@ -46,6 +46,20 @@ payload return `{ "success": true, "data": {} }`.
> `GET /api/screenshots/:name`, `GET /q/:code` (QR redirect), and the
> `GET /ws/sessions/:id/terminal` WebSocket upgrade.
> The [agent wait endpoints](#long-polling-agent-wait) use the normal envelope but
> are the only JSON endpoints that deliberately **hold the connection open**, for up
> to 600 s. Proxy operators and HTTP clients with a global read timeout need to know
> that before pointing them at Codeman.
⚠️ **A `401` is the one status that is not an envelope.** Authentication is rejected
in a request hook, before any handler runs, and it replies with the bare string
`Unauthorized` (`Unauthorized: hook secret required` on the hook path) plus
`WWW-Authenticate: Basic realm="Codeman"`. There is no `success`, no `error`, and no
`errorCode`, because the wrapping hook only wraps object payloads. So a client that
pipes every response straight into a JSON parser dies with a parse error rather than
reporting an auth failure, which is a confusing way to discover that a password is
set. Branch on the HTTP status **before** parsing.
## Error codes → HTTP status
The single source of truth is `ErrorStatus` / `httpStatusForErrorCode()` in
@@ -66,6 +80,459 @@ the HTTP status.
Adding a new error code is non-breaking; removing or renaming one is a major change.
## Long-polling (agent wait)
Three calls block until something happens instead of answering immediately. They
exist because SSE is Codeman's only other "tell me when" channel, and an agent
driving the API from a shell tool cannot practically hold a stream and parse
events inline.
| Call | Blocks until |
|------|--------------|
| `GET /api/v1/sessions/:id/wait` | one of a set of lifecycle signals fires |
| `GET /api/v1/sessions/:id/wait-output` | a literal string appears in the session's output |
| `POST /api/v1/sessions/:id/input` with `wait` | the input is delivered **and then** a signal fires |
`POST .../input` with `wait` is not the same as a `POST` followed by a separate
`GET .../wait`. It registers the waiter **before** writing, which closes the window
in which a separate wait sees the session still idle from the previous turn and
answers instantly with the wrong turn's result. Use it whenever you send a prompt
and want to know when that prompt is done.
### Three semantics that break callers who assume otherwise
**1. A timeout is HTTP `200`, not an error.** A wait that ends without its signal
returns `{"success":true, ...,"wait":{"timedOut":true,"signal":null}}`. The
intended pattern is a client-side loop over short waits, because `tailscale serve`
and cloudflared can both cut an idle connection, and turning every poll boundary
into a `4xx` would make that loop indistinguishable from a real failure. `408` is
auto-retried by several clients (silently doubling the polling load), `504` is what
a genuine tunnel failure looks like, and `204` cannot carry `waitedMs` / `status` /
`limitPaused`. Reserve error handling for the four codes in the table below.
**2. `stop` and `blocked` fire only for `claude` sessions.** Both come from Claude
Code hooks, and no other mode installs them: `shell` runs no agent, and the external
CLIs (`opencode`, `codex`, `gemini`, `antigravity`, `pi`) render their own TUIs and post
no hooks. For every non-`claude` mode only `idle`, `working` and `exit` are
accepted, and of those only `exit` is dependable: see the caveats under
[Signals](#signals) before building on `idle`. Requesting `stop` or `blocked`
**explicitly** on such a session is a
`400`; omitting `until` never fails, the server just drops them from the default set
and echoes the narrowed set back as `wait.until`. Three more places hooks can go
missing even in `claude` mode: a **Docker case** needs
`CODEMAN_DOCKER_BRIDGE_HOOKS=1`, since a container cannot reach a loopback-bound
Codeman (without it, only `idle` / `working` / `exit` work); a **remote-SSH
case** runs the agent on another host, whose hooks may never reach this server at
all; and a case whose hook config was written by **Codeman < 1.13.0 against an
`--https` install** carries hook curls without `-k`, which TLS-fail silently (the
hook line ends in `|| true`). Codeman now writes `curl -sk` and repairs a stale
case config the next time a session starts in that case. When in doubt, ask for
`stop,idle,exit` so a session without hooks still resolves on the heuristic
signal.
**3. `from=now` does not mean "printed after you asked".** tmux repaints the visible
screen on attach, on resize, and on any TUI redraw, and a repaint arrives as
ordinary output, so text that was already on screen can satisfy a fresh wait. This
was observed live: a marker echoed a minute earlier matched instantly on a new
`from=now` wait. It is inherent to running the agent under a multiplexer, so the
contract is a **marker unique to each call** (`MARK="DONE_$RANDOM"`, send
`echo $MARK`, then wait on `$MARK`), never a generic string like `BUILD OK`.
### Signals
| Signal | Source | Actually fires for |
|--------|--------|--------------------|
| `idle` | the session's own `idle` event | `claude`: yes, on ❯-prompt detection after activity. `shell`: **once only**, ~500 ms after start, and never again. External CLIs: not guaranteed (they render their own TUIs and readiness is output stabilization) |
| `working` | the session's own `working` event | `claude` only in practice (spinner and work-keyword detection are Claude output formats) |
| `stop` | the Claude Code `stop` hook, the definitive end-of-turn signal | `claude` only |
| `blocked` | a `permission_prompt` or `elicitation_dialog` hook | `claude` only, and rarer than it looks: see below |
| `exit` | no process is behind the session | every mode |
`stop` is the signal to orchestrate on where it exists; `idle` is a heuristic
fallback that can flap mid-turn when a spinner pauses. The default set when `until`
is omitted is `stop,idle,exit` (`exit` is in there so a worker that crashes resolves
the wait promptly instead of burning the caller's whole timeout on something that
can no longer happen). On a `claude` worker, prefer an explicit `until=stop,exit`
once the session is up: the default set's `idle` also resolves on a spinner pause,
and on a fresh session the **startup** `idle` (emitted when the CLI first comes up)
can land inside your first wait window and report a turn that never ran. Measured:
a session parked on the trust dialog emits no *further* `idle`, so it is the
startup transition, not the dialog, that produces the false success below.
⚠️ **`exit` means "nothing is running", which includes "not started yet".** The
server answers from `pid === null` plus a mux-layer pane-death probe, and that
covers a session that exited — including a worker that died *inside* its tmux pane
while the local attach client (and therefore `pid`) lives on — one that was
detached, and one that was **created but never started**. So the first wait
after `POST /api/v1/sessions` returns `{"signal":"exit","immediate":true}` in
milliseconds, and reading that as "the worker died" is wrong: it means start it, or
wait for it to come up. `status` is carried alongside so nothing is hidden. The
alternative (trusting `status`) is worse, because a dead PTY parks the session at
`status: "idle"`, which would answer the default wait with `immediate: true` for a
worker that has crashed. A worker dying while a wait is parked resolves it within
a few seconds (a background death-watcher), not at the timeout.
⚠️ **`blocked` is reachable less often than the table suggests.** It fires on two
hooks, and the default configuration suppresses one of them: Codeman spawns claude
with `--dangerously-skip-permissions`, so permission prompts do not happen unless the
instance is switched to the `auto` Claude mode (App Settings), or the caller is a
multi-user account without the bypass grant, which is forced to `--permission-mode
auto`. What does still fire under the default is `elicitation_dialog`, the agent
asking the user a question. So `until=stop,blocked,exit` is a reasonable belt on a
long turn, but a worker that never comes back is far more likely to be working than
blocked, and polling `blocked` alone will sit at its timeout.
⚠️ **On a `shell` session, only `exit` and marker-matching are dependable.** A shell
session emits its one `idle` at startup and then stays `status: "idle"` forever,
whatever the pane is doing, so it never emits a *transition*. Since send-and-wait
requires a transition (and so does `fresh=1`), both can only time out there:
a documented default `wait` on a shell worker running `sleep 4` times out at the
full 25 s. Synchronize hook-less sessions with `wait-output` and a unique marker
instead. The same caution applies to the external CLIs.
### Readiness is not a signal
Nothing here reports "the agent is ready for a prompt", and no combination of
`until`/`fresh` synthesizes one. A freshly created session reads as `exit` (above),
and a `claude` worker in a brand-new case comes up on the CLI's **trust dialog**,
which contains a ❯ prompt of its own. Send-and-wait posted at that moment types the
prompt into the dialog, where the `\r` never gets past it, while the session's
startup `idle` lands inside the wait window: the wait resolves on `idle` in a
couple of seconds with `timedOut: false`, which looks exactly like a completed
turn.
The reliable sequence is: poll `GET /api/v1/sessions/:id` until `.data.pid` is
non-null, then `wait-output` for the composer's own marker (`bypass`, the status
bar of a CLI spawned in bypass mode) with a short timeout, handling the trust
dialog only as the bounded fallback (`trust` matched → send `\r` → wait for
`bypass` again). Do not probe `trust` first and Enter blindly: the dialog text
stays in the terminal buffer for the life of the session, so a `trust` probe with
`from=buffer` keeps matching on every later run and the Enter lands in a ready
composer. A worked version is in
[`extending-codeman.md`](extending-codeman.md#seam-3-http-api-and-cli).
### `GET /api/v1/sessions/:id/wait`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `until` | comma-separated list of `idle,working,stop,blocked,exit` | `stop,idle,exit` | resolves on the first to fire. An unknown token is a `400` naming it, never a silent fallback |
| `timeout` | positive integer ms | `60000` | **validated first, clamped second.** `0`, a negative value and a fractional value are all `400`s, not clamps; a valid value outside `[1000, 600000]` is clamped and echoed as `wait.timeoutMs` |
| `fresh` | `0` \| `1` \| `false` \| `true` | `0` | `1` requires an actual transition, ignoring the state at call time |
```bash
curl -s "$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000"
```
Both GET wait routes answer with `Cache-Control: no-store`, because the documented
pattern polls one identical URL in a loop and a cached `{"timedOut":true}` would
turn that loop into a busy spin. `POST .../input` sends no cache header (it is a
POST, which is not heuristically cacheable).
⚠️ **Unknown query parameters are ignored, not rejected**, with one exception
(`regex`, below). In particular `match=` on `/wait` is silently dropped and you get
a plain signal wait, so check the endpoint path before blaming the parameters.
### `GET /api/v1/sessions/:id/wait-output`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `match` | literal string, 1 to 200 chars | required | substring match against the PTY stream with ANSI escapes stripped. A match spanning two PTY chunks is found |
| `nocase` | `0` \| `1` \| `false` \| `true` | `0` | case-insensitive compare. The returned snippet keeps the terminal's original casing |
| `from` | `now` \| `buffer` | `now` | `buffer` scans the tail of the existing terminal buffer (bounded, 256 KB by default) before blocking |
| `timeout` | positive integer ms | `60000` | same validation and clamp as `/wait` |
**Matching is literal, never a pattern.** A `regex` parameter is rejected with a
`400` rather than ignored, so a caller that assumed otherwise finds out immediately
instead of waiting on the wrong thing. The reasoning is in
[`architecture-invariants.md`](architecture-invariants.md#agent-wait-primitives).
#### What the matcher actually sees
The matcher scans the raw PTY stream, **normalized**: ANSI escape sequences are
stripped — CSI, OSC, and the charset-designation escapes a stock bash prompt emits
on every line (`ESC ( B`), so `match=tnode:` matches a prompt that renders
`…@tnode:` — a partial escape arriving at a chunk boundary is held back until its
tail arrives, and a match may straddle PTY chunks: `printf STRAD; sleep 1; printf
DLEQQ` is matchable as `STRADDLEQQ` (all measured live). Three caveats remain:
⚠️ **It is still the byte stream, not the rendered pane.** `GET .../terminal`
answers from a tmux screen capture (`data.source: "mux-visible"`), the finished
picture; the matcher sees the stream that painted it. For linear output the two
agree once escapes are stripped, but a full-screen TUI composes its picture with
cursor positioning, so what the pane shows and what the stream carries can differ.
Seeing your string in `terminal?tail=` makes a match likely, not guaranteed.
⚠️ **A TUI's text can arrive without its spaces.** Claude Code positions words
with cursor moves rather than printing spaces, so screen text can reach the
matcher as `Quicksafetycheck:Isthisaprojectyoucreated...`. Whether a given phrase
keeps its spaces depends on how the TUI happened to draw it (measured: `I trust
this folder` matched, `Quick safety check` did not), so a multi-word `match`
against a TUI pane is unreliable rather than impossible. Match a **single
space-free token**, ideally one you printed yourself. Plain command output (a
shell worker, an `echo`) keeps its spaces.
⚠️ **The returned `snippet` is a rendering of the matched text, not a quotation of
it.** It is cut from the same normalized stream the match ran against, then
cleaned for display: remaining raw control bytes are removed (an agent pipes the
snippet into its own terminal, so a worker's bytes must not be able to reset that
display) and blank runs are collapsed. A printable needle that matched will appear
in it; a needle containing control bytes or a blank run may not survive verbatim.
```bash
MARK="DONE_$RANDOM"
curl -sG "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=$MARK" --data-urlencode 'timeout=120000'
```
Build the query with `-G --data-urlencode` rather than by hand: a `+` in a
hand-written query string decodes to a space.
### `POST /api/v1/sessions/:id/input` with `wait`
Two optional fields on the existing endpoint:
| Field | Type | Notes |
|-------|------|-------|
| `wait` | `true` or the same comma grammar as `until` | `true` means the default signal set. Omitted keeps the historical fire-and-forget behavior, unchanged. `null`, `false` and an empty string are all read as **absent**, not as an error and not as "wait for the default" |
| `waitTimeout` | positive integer ms | same validation **and** clamp as `timeout`: `0`, a negative and a fractional value are `400`s, anything valid is clamped into `[1000, 600000]` and echoed as `wait.timeoutMs` |
Both are `nullish`, so an explicit `null` from `JSON.stringify` is accepted as
"absent" rather than failing validation. That is deliberate: `.optional()` would
reject it, which has shipped as a real bug twice.
The input must end with `\r` (a real carriage return in the JSON string): Enter is
sent only when the input contains one, so text without it is typed onto the
worker's prompt but never submitted, and the wait then runs its full timeout on a
turn that never started. Verified live; this is the most common silent failure on
this endpoint.
```bash
curl -s -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true,"clientId":"agent-1","seq":1,
"wait":"stop","waitTimeout":600000}'
```
A **tagged duplicate** (a `clientId` + `seq` pair the server has already applied)
still honors `wait`, because the caller's question is unanswered, but it answers
from the session's current state rather than requiring a new transition: the
original turn may be long over. It comes back as
`"delivered": false, "duplicate": true`.
### Response
All three nest the wait result under `data.wait`, so one client helper works against
any of them:
```json
{ "success": true, "data": {
"sessionId": "28325fd3-caa7-4178-82bf-87dfebf0f464",
"status": "idle",
"limitPaused": false,
"wait": {
"signal": "stop", "until": ["stop", "idle", "exit"],
"timedOut": false, "immediate": false, "ended": false, "aborted": false,
"waitedMs": 8421, "timeoutMs": 60000
}
}}
```
`POST .../input` returns the same `wait` object alongside `delivered`, `duplicate`,
`status` and `limitPaused`. `POST .../input` **without** `wait` is unchanged and
still returns `{"success": true, "data": {}}`.
⚠️ `delivered: false` has **two** meanings, and they must be told apart by
`duplicate`: with `duplicate: true` the input was suppressed as an already-applied
redelivery (harmless, the turn it refers to may be long over), while with
`duplicate: false` the **write failed** (typically no PTY behind the session). A
client that reads `delivered === false` as "duplicate" silently treats a failed send
as a success.
| Field | Type | Meaning |
|-------|------|---------|
| `wait.signal` | signal \| `null` | the signal that fired (`/wait` and `/input` only) |
| `wait.until` | array of signals | what the server actually waited on, after narrowing the default set for the session's mode (`/wait` and `/input` only) |
| `wait.matched` | boolean | the string appeared (`/wait-output` only) |
| `wait.match` | string | the literal that was searched for (`/wait-output` only) |
| `wait.snippet` | string \| `null` | bounded window of output around the match, blank runs collapsed for readability (`/wait-output` only) |
| `wait.timedOut` | boolean | the wait hit its timeout. Still a `200` |
| `wait.immediate` | boolean | the condition already held at call time, so nothing was waited for (`waitedMs` is 0) |
| `wait.ended` | boolean | the session went away (deleted or torn down) before the condition was met |
| `wait.aborted` | boolean | the client hung up, so the waiter was released without resolving — and by that definition a client never reads `true`. When the **server** abandons a wait itself (send-and-wait against a session with no PTY), it answers in about a millisecond with `ended: true`, `delivered: false`, `duplicate: false` and `aborted: false`: `delivered`/`ended` carry that story, and `aborted` stays the transport flag. Present for completeness; treat a `true` as "this wait answered nothing", never as an outcome |
| `wait.waitedMs` | number | wall-clock ms actually spent waiting |
| `wait.timeoutMs` | number | the timeout **after clamping**, which is what was applied |
| `status` | `SessionStatus` | the session's status after the wait, so a caller that timed out still learns where things stand |
| `limitPaused` | boolean | the session is paused on a usage limit and will emit nothing until its reset, so a timeout here is expected rather than a stall worth retrying hard |
Read the outcome by discriminator, in this order:
1. `wait.signal !== null` (or `wait.matched === true`): the thing happened.
2. `wait.timedOut`: a poll boundary. Loop again.
3. `wait.ended` or `wait.aborted`: the wait answered nothing, because the session is
gone or was never running. Re-check the session instead of looping.
`wait.immediate` is not a fourth outcome: it rides along with the first one and
means the condition already held at call time, so nothing was actually waited for.
If that is not what you meant, you wanted `fresh=1` or the send-and-wait form. Note
that `{"signal":"exit","immediate":true}` on a session you just created is the
not-started-yet case, not a crash.
**The timeout is clamped, so read it back.** A request for 1800000 ms is silently
reduced to the server's ceiling (600000 ms by default, operator-tunable), and a
request for 1 ms is raised to 1000 ms. `wait.timeoutMs` is the value that was
applied. Without checking it, a caller that asked for 30 minutes and got 10 will
read the timeout as "the worker is wedged" and kill a session that was working fine.
### Errors
| `errorCode` | HTTP | When |
|-------------|------|------|
| `INVALID_INPUT` | 400 | unknown `until` / `wait` token; `stop` or `blocked` requested explicitly on a mode that installs no hooks (the message names the mode); `regex=` on `/wait-output`; `match` outside 1 to 200 chars; a non-numeric `timeout` |
| `NOT_FOUND` | 404 | no such session, or one this caller does not own |
| `SESSION_BUSY` | 409 | this session's waiter cap is full |
| `RATE_LIMITED` | 429 | a per-owner or process-wide waiter cap is full. Retry later; the session you named is not the problem |
The two capacity codes are deliberately different. A process-wide cap reported as
`SESSION_BUSY` would tell the caller to switch sessions, which cannot help. The
error message names the cap that was hit.
⚠️ A `401` is **not** in this table and is not an envelope at all (see
[Response envelope](#response-envelope)). It matters most here: a polling loop that
pipes each wait straight into `jq` fails with a parse error on every iteration
against a password-protected server, which reads as "the wait endpoints are broken".
Check the status first.
The per-session cap is a **combined** budget: signal waiters and output waiters
count against the same 16, not 16 of each. An abandoned request no longer holds its
slot, because the routes release the waiter when the client disconnects, but a
client that opens many concurrent waits against one session will still hit the cap.
## Session lineage (`parentSessionId`)
A create request may name the session that spawned it, which the web UI draws as a
line between the two tabs. Accepted on `POST /api/v1/sessions` and
`POST /api/v1/quick-start`, either way:
```bash
# as a body field
-d '{"caseName":"worker-1","mode":"claude","parentSessionId":"'"$CODEMAN_SESSION_ID"'"}'
# or as a header, which is what an agent driving many spawns should use: set it once
# on the curl invocation and every spawn call carries it
-H "X-Codeman-Parent-Session: $CODEMAN_SESSION_ID"
```
The body field wins if both are present. The value is resolved against live sessions
(exact id, or a unique prefix of at least 8 characters) and must belong to the same
owner as the session being created.
**It cannot fail your spawn.** An unknown, stale, foreign or malformed value is
silently dropped and the session is created without lineage — never a `400`. It is
also pure decoration: it confers no permission, and a child is unaffected by its
parent exiting. It appears on session state as `parentSessionId` (absent when
unresolved) and survives a server restart.
## Approvals Inbox
Cross-session queue of prompts waiting on a human (permission dialogs,
AskUserQuestion questions, idle prompts). Claude-mode sessions only; items are
in-memory (a server restart drops them; the next prompt re-fires the hook).
Design: [`approvals-inbox-plan.md`](approvals-inbox-plan.md).
- `GET /api/v1/approvals` → `{ approvals: ApprovalItem[] }`, oldest first,
ownership-scoped in multi-user mode. `ApprovalItem`: `{ id, sessionId,
sessionName, kind: 'permission'|'question'|'idle', createdAt, toolName?,
toolSummary?, message?, cwd?, context?, options?: {n, label}[],
acknowledgedAt? }`. `context` is the ANSI-stripped visible pane frame;
`options` is present only when the dialog's numbered choices parsed
confidently; `acknowledgedAt` marks an item a human has already looked at
(see `/viewed` below) and tells clients not to re-arm its tab alert. Listing
also runs a staleness sweep over the caller's own items: the pane is
re-captured, and an item whose dialog no longer parses is resolved as
`resolved_in_terminal` instead of being returned (only items whose original
frame parsed `options` can be dropped this way, so an unreadable capture
keeps the item).
- `POST /api/v1/approvals/:id/answer` with `{ action: 'approve' }` (sends the
digit `1`), `{ action: 'deny' }` (sends Esc), `{ action: 'option', option: n }`
(sends the digit; accepted only when `n` is among the item's parsed
`options`), or `{ action: 'text', text }` (idle prompts only; submits the
line as a prompt). `404 NOT_FOUND` when the item is no longer pending,
`409 CONFLICT` when the dialog left the screen or another actor answered
first, `422 OPERATION_FAILED` when the session refused input.
- `POST /api/v1/approvals/:id/dismiss` removes the item without keystrokes.
- `POST /api/v1/approvals/session/:sessionId/viewed` → `{ sessionId,
acknowledged: itemId | null }`. Marks the session's pending **idle** item as
seen by a human (the web UI calls it when you open the session's tab): the
item stays pending and answerable, but stops arming the yellow tab alert on
every client, including after a reload. Permission/question items are never
acknowledged this way, since looking at a dialog does not answer it. `404`
for an unknown or inaccessible session; acknowledging twice is a no-op
(`acknowledged: null`).
SSE events: `approval:pending` (full item), `approval:updated` (context/options
re-captured, or the item acknowledged), `approval:resolved` (`{ id, sessionId, kind, resolution }` with
`resolution` one of `answered | resolved_in_terminal | superseded |
session_ended | dismissed | expired`).
## Read My Mind intent profiles
Per-case profiles of what the user is trying to accomplish: user/agent-stated
goals plus the user's recently submitted prompts, captured from the Claude
session transcript while the opt-in `readMyMindEnabled` setting is on (default
OFF). Keyed by owner + workingDir, so the profile survives `/clear`, respawns,
and session churn. Stored in `~/.codeman/intents.json` (mode 0600); never fed
into `/api/v1/search`. Design: [`readmymind-plan.md`](readmymind-plan.md);
user guide: [`readmymind.md`](readmymind.md).
- `GET /api/v1/sessions/:id/intent` -> `{ intent: IntentProfile }` for the
session's case. `IntentProfile`: `{ key, workingDir, updatedAt, goals,
recentPrompts: { ts, sessionId, text }[] }` (prompts oldest first, FIFO cap
50, each <= 500 chars). A case with nothing recorded answers an empty
profile with `updatedAt: 0`; nothing is persisted by reads.
- `PUT /api/v1/sessions/:id/intent` with `{ goals }` (<= 8192 chars, strict
schema) replaces the goals text and answers the updated profile.
`400 INVALID_INPUT` on over-long or unknown fields.
- `DELETE /api/v1/sessions/:id/intent` -> `{ deleted: boolean }` forgets the
case's profile entirely.
- `POST /api/v1/sessions/:id/readmymind` predicts the user's next prompt:
a one-shot model call over the intent profile plus live session signals
(pending approval dialog, transcript tail, git state, run-summary events,
sibling sessions). Body is optional; the rethink flow passes
`{ steer?, rejected? }` (strict schema: `steer` <= 2000 chars, `rejected`
up to 10 strings <= 1000 chars). Answers
`{ suggestions: { prompt, why, kind }[], durationMs }` with 1-3 suggestions
(`kind`: `continue` | `verify` | `redirect`; prompts are single-line).
Claude-mode sessions only (`400 INVALID_INPUT` otherwise); one prediction in
flight per session (`409 CONFLICT`); predictor failures answer
`502 OPERATION_FAILED`. Takes 5-90 s and costs real tokens. Suggestions are
only ever returned, never sent: submitting one is the caller's explicit act.
All four enforce session ownership in multi-user mode; a foreign session id
answers `404 NOT_FOUND` (no existence leak), and profiles of two owners of the
same directory are distinct by construction.
## Voice dictation
Browser dictation transcribed through this server's Claude Code login, i.e. the
same speech-to-text service the CLI's own `/voice` mode uses. Gated on the synced
`claudeVoiceEnabled` setting (default OFF). Design:
[`claude-voice-plan.md`](claude-voice-plan.md).
- `GET /api/v1/voice/status` -> `{ available, reason?, subscriptionType?,
expiresAt? }`. `reason` is `disabled` (setting off), `no-credentials` (nobody
signed in to Claude Code on the server), `expired` (the access token elapsed;
running any Claude session refreshes it) or `malformed`. The OAuth token
itself is never returned by this or any other endpoint.
- `GET /ws/voice/stream?language=&keyterms=` (WebSocket, not under `/api`)
relays one dictation. Client sends binary frames of signed 16-bit
little-endian PCM, 16 kHz mono (<= 64 KB per frame), plus JSON control frames
`{"t":"finalize"}` (ask for the final transcript) and `{"t":"stop"}`. Server
sends `{"t":"ready"}`, `{"t":"transcript","text","final"}` (each frame is the
WHOLE running transcript, not a delta), `{"t":"error","message"}` and
`{"t":"closed"}`. Close codes: `4003` disallowed Host/Origin, `4004`
unavailable (reason in the close reason), `4008` too many concurrent streams.
Streams are capped in count and length (`src/config/voice.ts`).
## Authentication
Optional HTTP Basic (`CODEMAN_USERNAME`/`CODEMAN_PASSWORD`) → opaque
@@ -82,6 +549,29 @@ the stable contract — event names are not renamed without a major bump. An
optional `?sessions=<id,...>` filter suppresses only the high-volume terminal
stream; lifecycle/metadata events are delivered to all clients regardless.
### `sse:heartbeat` (liveness)
Every 15s the server writes a `sse:heartbeat` frame to every connected client:
```
event: sse:heartbeat
data: {"t":1755100000000}
```
`t` is the server's epoch-ms timestamp at write time. The frame carries no
application state and can be ignored for correctness. It exists so a client can
tell a live stream from a dead one: an `EventSource` whose connection has been
idle-closed by a proxy (or that resumed from sleep on a stale socket) keeps
delivering nothing without ever firing `onerror`. Clients that care should treat
silence longer than about three intervals as a dead stream and reconnect, which
is what the bundled frontend does.
This replaced a `:keepalive` SSE **comment**, which served the same
proxy-flushing purpose but is invisible to `EventSource` by spec and so could
never be observed by a client. Consumers written against the old behavior are
unaffected: `EventSource` dispatches only events that have a registered
listener, so an unknown event name is dropped.
## Consuming from JavaScript
The bundled frontend reads responses through `_apiJson()`
+107
View File
@@ -0,0 +1,107 @@
# Approvals Inbox (design)
One cross-session inbox for every prompt that is waiting on a human: permission dialogs, questions (AskUserQuestion / elicitation), and idle prompts. Cards are answerable in place (option digits, Esc, or a typed prompt) from desktop, phone overview, and push notification action buttons. Inspired by Cloudflare OS's Gatekeeper approval queue (https://github.com/cloudflare/cloudflare-os, asynchronous human-in-the-loop approvals): with a fleet of sessions the human is the bottleneck, and today answering means finding the right tab.
## Problems this fixes (all real today)
1. **No cross-session surface.** Pending prompts exist only as per-tab alert colors (`tab-alert-action`/`tab-alert-idle`) and NEEDS YOU rows on the phone overview. Answering means switching to the session and typing.
2. **Alerts die on reload.** `pendingHooks` lives only in `app.js` memory, fed by transient SSE `hook:*` events. A page reload (or a phone browser evicting the tab) silently loses every pending alert. There is no server-side record.
3. **Push Approve/Deny buttons are dead.** `PUSH_EVENT_MAP` already attaches `approve`/`deny` actions to permission pushes, and `sw.js` forwards `event.action` to the page, but the `notification-click` handler in settings-ui.js ignores it (and when no tab is open, the action is dropped entirely). The buttons render on the lock screen and do nothing.
4. **Card context is missing.** The frontend handlers read `data.question` / `data.message` / `data.tool`, but `sanitizeHookData` never forwards `message`, so notifications show generic fallback text.
## Scope
- Claude mode only (hooks fire only for `claude`; external CLIs keep their output-stabilization heuristics and get no inbox items). This mirrors the wait-primitive `stop`/`blocked` gating.
- Permission prompts occur for sessions running `ClaudeMode` `normal` / `auto` / `allowedTools` (and the trust-folder dialog even under skip-permissions). Question and idle prompts occur in every mode including `dangerously-skip-permissions`.
- In-memory store (plus the frontend seeding from it on load). Server restart drops items; hooks re-fire on the next prompt. No new state file in v1.
## Data model
At most **one active item per session**: the Claude TUI shows one dialog at a time, so a new prompt event supersedes the session's previous item (resolution `superseded`).
```ts
interface ApprovalItem {
id: string; // `${sessionId}:${seq}`
sessionId: string;
sessionName: string;
kind: 'permission' | 'question' | 'idle';
createdAt: number;
toolName?: string; // from sanitized hook data
toolSummary?: string; // command / file_path / description, already bounded
message?: string; // Notification hook `message` (newly allowlisted)
cwd?: string;
context?: string; // ANSI-stripped visible pane frame tail, ≤ 4000 chars
options?: { n: number; label: string }[]; // parsed from context when confident
}
```
Resolutions (server-emitted, item removed from pending): `answered` (via inbox), `resolved_in_terminal` (stop / elicitation_complete / elicitation_response / session went working), `superseded`, `session_ended`, `dismissed`, `expired` (12h TTL sweep).
## Backend
### Store: `src/approval-inbox.ts`
Module-level singleton in the style of `session-wait-registry.ts` (pure, no `Session` import, injected emit callback so there is no import cycle with the server):
- `notePrompt(info)` creates/supersedes the session's item; schedules ONE re-capture ~600ms later (the Notification hook can fire before the dialog finishes painting) which updates `context`/`options` and emits `approval:updated`.
- `resolveForSession(sessionId, reason)`, `dismiss(id)`, `answerable(id)`, `listPending()`, `stop()` (clears timers; tests).
- Option parsing (pure, unit-tested): consecutive `❯? N. label` lines, 2..6 options, labels ≤ 120 chars. Parsed options gate which digits the answer endpoint accepts; when parsing fails the card falls back to Approve(1)/Deny(Esc) only.
- TTL: items expire after 12h (checked on read + a lazy sweep; no standing interval).
### Wiring
- `hook-event-routes.ts`: on `permission_prompt` / `elicitation_dialog` / `idle_prompt`, call `notePrompt` with sanitized data + a pane capture callback (`mux.capturePaneBuffer(muxName)` visible frame, ANSI-stripped via existing utils; fall back to `session.terminalBuffer` tail). On `stop` / `elicitation_complete` / `elicitation_response`, `resolveForSession(id, 'resolved_in_terminal')`.
- `session-listener-wiring.ts`: `working` listener resolves **idle items only** (`working` is heuristic and can flap mid-turn, so it must never clear a pending permission/question dialog); `exit` resolves with `session_ended`. Same singleton-import pattern as `sessionWaits`.
- Session delete route: resolve with `session_ended`.
- **New hook matchers** `elicitation_complete` + `elicitation_response` added to `generateHooksConfig()`, `HookEventType`, `HookEventSchema`, and both SSE registries. `refreshStaleCodemanHooks` gets a staleness probe for them (`hooksJson.includes('elicitation_complete')`) so existing cases heal on next Claude spawn, exactly like the `-k`/secret/marker probes.
- `sanitizeHookData`: allowlist `message` (bounded 500 chars). This also un-deadens the existing notification text paths.
### Routes: `src/web/routes/approval-routes.ts`
Normal authed API (NOT the hook-secret bypass), `ApiResponse` envelope, Zod schemas in `schemas.ts`:
- `GET /api/approvals` → pending items, multi-user filtered by `canAccessOwned` (same policy as session lists). Also sweeps the caller's own items for staleness through `verifyStillAnswerable()`: Claude Code fires no "permission answered" hook, so a dialog answered in the terminal used to sit pending until `stop` and re-arm a red tab alert on the next page load. Only items whose original frame parsed options can be dropped this way, so an unreadable capture keeps the alert.
- `POST /api/approvals/:id/answer` body `{ action: 'approve' | 'deny' | 'option' | 'text', option?, text? }`:
- `approve` → `writeViaMux('1')` (option 1 is always plain Yes; no Enter, menus react to the digit).
- `deny` → `writeViaMux('\x1b')` (Esc is the official No/cancel; precedent: auto-resume sends Esc the same way).
- `option` → digit `String(n)`; accepted only when `n` is within the item's parsed options (prevents blind digit-poking at an unparsed dialog).
- `text` → `idle` items only: single line, embedded newlines stripped, sent as `text\r` (the `\r` discipline from CLAUDE.md).
- Guards: item still pending (404 otherwise), session exists + ownership via `findSessionOrFail`, session mode installs hooks. **Answer-time re-capture**: for items whose frame parsed options, the pane is re-captured before sending; if the dialog no longer parses, the item resolves and the answer is refused with 409 (the keystroke would land in whatever now has focus). Marks `answered` BEFORE the write so a double-tap cannot double-send; rolls back to pending if the write fails.
- `POST /api/approvals/:id/dismiss` → remove without keystrokes.
- `POST /api/approvals/session/:sessionId/viewed` → acknowledge the session's pending **idle** item (`acknowledgedAt`, emitted as `approval:updated`). Added after the owner reported that a yellow tab clicked and checked went yellow again on reload: the view-clears-idle rule lived in one browser's memory, so the seed re-armed it and other devices never saw the clear. Acknowledgement is deliberately **not** resolution (the prompt is still unanswered, so it stays in the inbox and stays available as Read My Mind context), and deliberately **idle-only** (looking at a permission/question dialog does not answer it, so the red alert survives being viewed).
### SSE
`approval:pending`, `approval:updated`, `approval:resolved` in `sse-events.ts` + `SSE_EVENTS` in constants.js (the parity test pins the sync). Broadcasts carry `sessionId`, so multi-user SSE scoping applies unchanged.
### Push
- `sendPushNotifications` payload gains `approvalId` for the three hook events. Both `approvalId` and the Approve/Deny `actions` are **gated on the opt-in setting**: with it off, permission pushes carry no buttons at all (pre-inbox they rendered and did nothing, so stripping them is the honest shape).
- `sw.js` `notificationclick`: when `event.action` is `approve`/`deny`, POST `/api/approvals/:id/answer` directly from the worker (same-origin, cookie credentials) so the buttons work **with no tab open**; on failure fall back to focusing/opening a tab. Non-action clicks keep today's behavior.
- Page-side `notification-click` handler: honor `action` instead of dropping it (also setting-gated, for stale notifications sent before the toggle flipped).
- Question/idle pushes keep no action buttons (options vary per dialog); tapping opens the inbox.
## Frontend
New module `approvals-ui.js` (@loadorder 11.2, after panels-ui.js), prettier-formatted (not added to `.prettierignore`).
- **Seed on connect**: `GET /api/approvals` on init and SSE reconnect; each pending item re-feeds `setPendingHook(...)` so tab alerts and the phone overview survive reload (fixes problem 2 with zero changes to the alert state machine). Items carrying `acknowledgedAt` are skipped, and `markIdleAlertSeen()` (app.js) is what sets it: viewing a session clears its yellow locally and POSTs `.../viewed`, so "I checked it" survives the reload and reaches the user's other devices through `approval:updated`.
- **Desktop**: header bell `btn-approvals` with count badge. Ships default-hidden via marker class `btn-approvals--hidden` (same policy as the attachments button, so `test/mobile-header-buttons-policy.test.ts` excludes it from the default-visible enumeration); JS shows it only while count > 0. Click toggles a drawer of cards: session name + kind, tool/message summary, mono context block, buttons rendered from parsed options (else Approve/Deny), plus Dismiss and Open session. Esc closes; existing z-index layers respected.
- **Phone**: header button stays hidden (`mobile.css`); the phone surface is the overview's NEEDS YOU section, whose rows gain inline ✓/✗ buttons for permission items (tap-through to the session remains the row's main action). Toolbar classes/status language rules from the mobile-overview section of CLAUDE.md apply.
- **i18n**: new strings registered in i18n.js (en + zh-CN); status words carry `data-i18n-skip` where they would collide (mirroring the overview pills).
- **Setting**: `approvalsInboxEnabled`, synced (in `SettingsUpdateSchema`), **default OFF** (owner decision: the entire feature is opt-in, meaning no bell, no drawer, no overview strips, no seeding, and no push action buttons until enabled in App Settings → Panels). Only the store and answer endpoints keep running regardless, so flipping the toggle ON surfaces anything already pending immediately, with no restart.
## Race honesty
The prompt can be answered in the terminal a moment before an inbox answer lands; then the keystroke would hit whatever now has focus (worst case: a digit typed into the composer, not submitted, since no `\r` is ever sent for menu answers). Mitigations, in order: answer-time re-capture (the dialog must still parse on screen or the answer is refused), answered-before-write marking, digit-only/Esc-only writes for menus, and the card's context block showing what the pane looked like when captured. This is the same class of risk `writeViaMux` automation (auto-resume, respawn) already accepts.
## Tests
- `test/approval-inbox.test.ts`: supersede per session, every resolution path, TTL, option parsing fixtures (2-option, 3-option with ❯, unparseable frame), re-capture update.
- `test/routes/approval-routes.test.ts` (`app.inject`, no port): list; hook event creates item; answer approve/deny/option writes the exact bytes (test-PTY echo asserts them); text answers restricted to idle; 404 unknown id; 409 answered twice; option out of range rejected; multi-user scoping.
- Existing suites extended: hook-event schema accepts the two new events; `sanitizeHookData` forwards bounded `message`; SSE parity + mobile-header policy pass as-is by construction.
## Docs
- CLAUDE.md: Key Patterns entry + SSE/route counts + frontend load order.
- `docs/api-reference.md`: the two endpoints + three SSE events (additive, fine under the 0.9.x contract).
File diff suppressed because one or more lines are too long
+5
View File
@@ -300,6 +300,11 @@ For reference when writing browser tests:
.xterm // Terminal container
#helpModal // Help modal
#appSettingsModal // Settings modal
#sessionOptionsModal // Session Options (same set-* surface)
#createCaseModal // Add Case (same set-* surface)
.set-rail-item // Rail entry: scrolls in App Settings, switches in the other two
.set-section // A settings section (`.hidden` on the inactive ones outside App Settings)
.set-row // One setting: label + description left, control right
.modal-content // Modal content
.modal-close // Modal close button
.header-brand .logo // Logo text
+21 -3
View File
@@ -149,9 +149,17 @@ to Claude as a system reminder. This implies `"async": true`; ordinary async
hooks do not wake an idle turn, and their output waits for the next interaction.
Codeman uses this on `PostToolUse(Bash)`: a self-contained Node helper extracts
the background task ID from the Bash result, watches the session transcript for
the matching completion notification, and exits 2. It does not send terminal
input, so it cannot submit a user's partially written prompt.
the background task ID from the Bash result, watches the originating transcript
and, for subagents, the top-level parent transcript for the matching completion
notification, and exits 2. Claude records a subagent's Bash result in its
`subagents/agent-*.jsonl` file but queues completion in the lead session JSONL.
The task ID keeps each wake targeted. The helper does not send terminal input,
so it cannot submit a user's partially written prompt.
For script-dispatched Codex work, `codex-run.sh` writes the final response
between `CODEMAN_RESULT_BEGIN/END` markers in the background task output. The
rewake helper includes a maximum of 64 KiB of that report in its feedback. UI
subagent discovery and dispatcher result delivery are separate contracts.
### Notification
@@ -219,6 +227,16 @@ Or to allow exit:
**Use Cases**: Control nested loops, verify subagent output.
The hook input includes `agent_id`, `agent_transcript_path`, and
`last_assistant_message`. Like `Stop`, a command hook can return
`{"decision":"block","reason":"..."}` to keep the subagent running and feed
the reason back to it.
Codeman uses this to prevent premature reports from workers that still own live
Monitor or background-Bash processes. It derives candidate task IDs from the
subagent transcript, but requires a matching live Linux process descriptor for
`tasks/<id>.output`; historical task text by itself is not treated as active.
### TeammateIdle
**When**: When an agent-team teammate is about to go idle.
+121
View File
@@ -0,0 +1,121 @@
# Claude voice dictation in Codeman
Wire Codeman's existing mic button to the same speech-to-text service Claude Code's own
`/voice` mode uses, so dictation works with **no third-party API key** for anyone already
signed in to Claude Code on the server.
## Why the CLI's own voice mode cannot be reused directly
Claude Code 2.1.x ships voice input: `/voice hold|tap|off` arms it, the CLI opens the
**host's** microphone (native `audio-capture-napi`, falling back to `sox`/`arecord` on Linux
after probing `/proc/asound/cards`), streams PCM upstream and types the transcript into its
own composer.
Every part of that is on the wrong machine for Codeman. The CLI runs inside a tmux pane on
the server, which is typically headless and has no sound card at all, while the human is in
a browser on a phone somewhere else. Toggling `/voice` in the pane from Codeman would arm a
microphone nobody is sitting in front of. So Codeman keeps capturing audio in the browser,
where the user actually is, and only borrows the CLI's **transcription backend**.
## The backend, as the CLI uses it
Extracted from the 2.1.226 binary (`connectVoiceStream`):
| | |
| --- | --- |
| URL | `wss://api.anthropic.com/api/ws/speech_to_text/voice_stream` |
| Query | `encoding=linear16`, `sample_rate=16000`, `channels=1`, `endpointing_ms=300`, `utterance_end_ms=1000`, `language=<lang>`, `use_conversation_engine=true`, `stt_provider=deepgram-nova3` |
| Headers | `Authorization: Bearer <Claude Code OAuth access token>`, `User-Agent`, `x-app: cli`, `anthropic-client-platform`, optional `x-config-keyterms` |
| Audio | raw binary frames, PCM signed 16-bit little-endian, 16 kHz, mono |
| Keepalive | `{"type":"KeepAlive"}` on open, then every 8 s |
| Finalize | `{"type":"CloseStream"}`, then wait for the endpoint frame |
| Downstream | `{"type":"TranscriptText"\|"TranscriptInterim","data":"…"}` (running interim), `{"type":"TranscriptEndpoint"}` (promotes the pending interim to final), `{"type":"TranscriptError",…}`, `{"type":"error","message":…}` |
Deepgram Nova-3 runs server-side, so the Deepgram-quality result arrives without a Deepgram
account. Verified against the live endpoint before this design was written: connect, stream
PCM, receive interims and an endpoint frame.
## Architecture
The browser cannot call that endpoint itself: it would need the OAuth bearer token in page
JavaScript (and CORS would refuse anyway). So the audio goes browser → Codeman → Anthropic,
and Codeman is the only thing that ever touches the token.
```
mic → AudioWorklet (Float32 → PCM16 @16 kHz)
→ wss://<codeman>/ws/voice/stream [cookie/basic auth, Origin+Host guarded]
→ VoiceStreamRelay (reads ~/.claude/.credentials.json per connect)
→ wss://api.anthropic.com/api/ws/speech_to_text/voice_stream
← {"t":"transcript","text":…,"final":…} → existing _insertText() path
```
Nothing about the insert path changes: the transcript lands in the same preview overlay,
the same direct/compose insert modes, the same green Send button.
### Server pieces
- **`src/claude-credentials.ts`** — locate and parse the Claude Code OAuth credentials.
`parseClaudeCredentials()` is pure (JSON string + `now` → status) and unit-tested;
`readClaudeOAuthToken()` wraps it with IO: `$CLAUDE_CONFIG_DIR/.credentials.json` or
`~/.claude/.credentials.json`, and on macOS the login keychain
(`security find-generic-password -s "Claude Code-credentials"`).
**Read-only, always.** Codeman never writes credentials and never refreshes the token: a
refresh rotates the refresh token, and racing Claude Code's own refresh could sign the
user out of their CLI. An expired token surfaces as a plain "run a Claude session to
refresh" error instead.
The token is never logged, never returned by any endpoint, and never sent to the browser.
- **`src/web/voice-stream.ts`** — pure `buildVoiceStreamUrl()` / `buildVoiceStreamHeaders()` /
`sanitizeKeyterms()` (ASCII-only, deduped, 1024-char cap, mirroring the CLI), plus
`VoiceStreamRelay`, which owns one upstream socket: keepalive timer, audio passthrough,
transcript translation, finalize, and the caps below.
- **`src/web/routes/voice-routes.ts`**
- `GET /api/voice/status` → `{ available, reason, subscriptionType?, expiresAt? }`. Never
the token. `available:false` with a machine-readable `reason` (`disabled`, `no-credentials`,
`expired`) is what the settings row and the provider resolver read.
- `GET /ws/voice/stream?language=&keyterms=` → the relay. Same upgrade guard as
`/ws/sessions/:id/terminal`: allowed Host, same-site Origin, and the global auth hook has
already run on the handshake.
Caps, because an open mic is an open pipe: one stream per connection, `MAX_VOICE_STREAMS`
concurrent server-wide, a hard `MAX_STREAM_MS` per stream, and a per-frame size cap. A tab
left recording cannot bill an unbounded amount of upstream audio.
### Frontend pieces
- **`voice-pcm-worklet.js`** — an `AudioWorkletProcessor` converting Float32 blocks to PCM16
and posting ~256 ms frames back. `MediaRecorder` cannot produce raw PCM, which is why the
existing Deepgram path (container audio, auto-detected) cannot be reused as-is. Falls back
to `ScriptProcessorNode` where AudioWorklet is unavailable.
- **`ClaudeVoiceProvider`** in `voice-input.js` — mirrors `DeepgramProvider`'s shape
(`start({language, keyterms, onStream, onResult, onError, onEnd})`) so `VoiceInput` treats
the three providers uniformly.
- **Provider resolution** — new `voiceSettings.provider`: `auto` (default) | `claude` |
`deepgram` | `webspeech`. `auto` picks Claude when `/api/voice/status` reports it
available, else Deepgram when a key is set, else Web Speech. Pinning a provider always
wins, so an existing Deepgram user can keep exactly what they have.
### Settings
- `claudeVoiceEnabled` — synced, **default OFF**, gating the whole server side. Off is the
honest default: turning it on means this machine's Claude subscription starts paying for
transcription for whoever can reach the UI, and the audio goes to Anthropic rather than to
wherever it went before. One switch in Settings → Voice, and the mic works with no key.
- `voiceSettings.provider` — per the resolution table above; joins the existing synced
`voiceSettings` object.
## Things worth knowing
- **This uses an undocumented endpoint with subscription credentials.** It is the user's own
token, on the user's own machine, driving the user's own Claude Code install, but it is not
a published API and Anthropic can change or restrict it. Default-OFF is deliberate; the
Deepgram and Web Speech paths stay untouched as the supported fallbacks.
- **Multi-user mode**: every user's dictation would run on the server owner's Claude
credentials, exactly as every user's *sessions* already run on them. Consistent, but worth
stating out loud in the settings copy.
- **Token lifetime** is about 8 hours, refreshed by Claude Code itself whenever it runs. The
relay re-reads the file on every connect rather than caching, so a refresh is picked up on
the next press of the mic.
- **HTTPS or localhost**: `getUserMedia` needs a secure context. Prod is HTTPS behind
`tailscale serve`, so this is already satisfied; the existing error copy covers the rest.
+2 -2
View File
@@ -44,9 +44,9 @@ records), kept distinct from the existing `ScheduledRun`.
## 2. Where agent/session types are defined
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' | 'antigravity' | 'pi'`
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
- CLI availability resolvers in `src/utils/{claude,codex,gemini,antigravity,opencode,pi}-cli-resolver.ts`.
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
## 3. Where input is sent into a session
+2 -2
View File
@@ -1,7 +1,7 @@
# Cron Jobs — User & Operator Guide
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
spin up a Claude (or shell / OpenCode / Codex / Antigravity / Gemini / Pi) session on a schedule and
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
@@ -91,7 +91,7 @@ These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` \| `antigravity` \| `pi` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. ⚠️ A `pi` job's readiness poll looks for `❯`/a token count, neither of which pi prints, so it burns the poll budget and then sends the prompt anyway (slower start, still works). |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
+18 -1
View File
@@ -2,7 +2,7 @@
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` all work inside the container.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` / `antigravity` / `pi` all work inside the container.
## One-time setup: build the base image
@@ -15,6 +15,23 @@ node scripts/build-agent-image.mjs # builds codeman/agent:base
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
⚠️ **Re-build with `--no-cache`, always.** The CLIs are installed in a single `RUN npm install -g` layer, so a plain rebuild re-uses it from the Docker layer cache and the CLIs stay frozen at whatever versions the image was **first** built with, however long ago that was. Editing the Dockerfile does not help unless the edit lands at or above that line: a change appended below it leaves the npm layer cached and only runs the new step. Observed 2026-08-06: a rebuild silently kept a stale `@openai/codex@0.144.6` whose aliased platform binary had not installed, so every `codex` docker case died with `Missing optional dependency @openai/codex-linux-x64` while the build itself reported success.
```bash
node scripts/build-agent-image.mjs --no-cache
```
A zero exit code only proves the layers ran, not that the toolchain works. Verify by actually executing each CLI in the image, and check the build log for `Using cache` lines:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
Antigravity (`agy`) is the one CLI not installed from npm (Google ships a standalone binary), so it has its own Dockerfile step and adds roughly 190MB; a full image lands near 1.6GB. Pi also gets its own step, because upstream documents installing it with `--ignore-scripts` and that flag must not silently change how the other four npm CLIs install.
Pi's credentials are seeded per-FILE rather than as a whole directory (`auth.json`, `settings.json`, `trust.json`, `models.json`, `models-store.json` out of `~/.pi/agent`), because that directory also holds `sessions/`, `extensions/`, `skills/` and the installed package trees — gigabytes on an active host. Consequence: in-container pi sessions are invisible host-side, so `pi -c` inside a Docker case only sees that container's own history. See [`pi-integration.md`](./pi-integration.md).
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
+417
View File
@@ -0,0 +1,417 @@
# Extending Codeman
Codeman has no plugin runtime, and that is a deliberate choice rather than a
missing feature. A plugin runtime means running third-party code inside a process
that spawns agents with your credentials, on a server people routinely expose
over a tunnel or Tailscale. Codeman's security model is one of its reasons to
exist, so it does not hand that away for an extension mechanism.
Instead there are four seams that already work, from any language, with nothing
installed:
| You want to | Use | Runs where |
| --- | --- | --- |
| Show your own UI inside Codeman | [Web tabs](#seam-1-web-tabs) | Your own process, rendered as a tab |
| React when an agent needs you | [SSE events](#seam-2-sse-events) | Anywhere that can hold an HTTP connection |
| Drive Codeman from a script | [HTTP API](#seam-3-http-api-and-cli) or the `codeman` CLI | Anywhere |
| React inside a Claude session | [Hooks](#seam-4-hooks) | The agent's own machine |
Everything below is covered by the stability promise in
[`versioning-policy.md`](versioning-policy.md): endpoint paths, the response
envelope, `errorCode` values, and SSE event names are stable. Additive changes
(new endpoints, new optional fields, new events) are non-breaking. Breaking
changes ship under a new prefix (`/api/v2`).
## Before you start
**Base URL.** `http://127.0.0.1:3000` by default. Prefer the versioned prefix
`/api/v1/...` for anything you publish; the unversioned `/api/...` is an alias.
**Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic on every request, or
authenticate once and keep the `codeman_session` cookie. With no password set,
Codeman is loopback-only and unauthenticated.
```bash
curl -u admin:$CODEMAN_PASSWORD http://127.0.0.1:3000/api/v1/sessions
```
**Envelope.** Every response is `{"success": true, "data": ...}` or
`{"success": false, "error": "...", "errorCode": "..."}`. Check the HTTP status
or `body.success`, then read `body.data`. The full `errorCode` to status mapping
is in [`api-reference.md`](api-reference.md).
⚠️ A few legacy GETs (`/api/away-digest` among them) return a bare-ish body with
the payload at the top level rather than under `data`. Read defensively with
`body.data ?? body`.
⚠️ A `401` is not an envelope at all: auth is rejected in a request hook that
replies with the bare string `Unauthorized`, so parsing it as JSON throws. Branch on
the status code before you parse, or a missing password looks like a broken endpoint.
**Already driving Codeman from an agent?** The README's
[Programmatic Guide](../README.md#driving-codeman-from-an-agent--programmatic-guide)
covers the in-session case: the `CODEMAN_MUX`, `CODEMAN_API_URL`,
`CODEMAN_SESSION_ID` and `CODEMAN_HOOK_SECRET_FILE` variables that let a CLI
running inside Codeman find the API and avoid acting on itself. This page is for
code running *outside* a session.
## Seam 1: Web tabs
The highest-leverage seam. Any web app you can serve locally becomes a tab beside
your agent sessions. You write a normal web page; Codeman handles embedding it.
```bash
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/webviews \
-H 'Content-Type: application/json' \
-d '{"name":"My Dashboard","url":"http://127.0.0.1:8787","icon":"📊"}'
```
Fields: `name` (1 to 60 chars), `url`, and optionally `icon` (a single glyph, max
8 code units), `embedMode` (`proxy` by default, or `direct`), and `trusted`.
Related endpoints: `GET /api/v1/webviews`, `PATCH /api/v1/webviews/:id`,
`DELETE /api/v1/webviews/:id`, `POST /api/v1/webviews/probe` (reachability and
framing check), `POST /api/v1/webviews/:id/open`.
### Why it is proxied
By default your page is served through Codeman's own origin at `/webview/:cap/*`
rather than framed directly. A direct iframe fails three ways at once: production
is HTTPS so `http://` targets are blocked as mixed content, many dashboards send
`X-Frame-Options: DENY`, and Codeman's own `default-src 'self'` CSP blocks
cross-origin frames. Proxying solves all three without weakening the CSP.
### The two things that will confuse you
A proxied frame is sandboxed and therefore **opaque-origin** unless you set
`trusted: true`. Two consequences look like bugs in your own app:
1. **Root-absolute URLs built at runtime** (`/assets/x.png` assembled in JS)
escape the injected `<base>` tag. Codeman injects a `runtimeUrlShim()` that
patches the common DOM sinks, but if you construct URLs in an unusual way,
prefer relative paths.
2. **Same-host `fetch` and `XHR` are CORS-checked with `Origin: null`.** Codeman
handles this with `buildProxyCorsHeaders()`, and the proxy is exempt from the
global `OPTIONS` short-circuit. If you see "Failed to fetch" while the page
itself renders fine, this is the area to look at.
⚠️ `trusted: true` opts out of the sandbox. A proxied page is served from
Codeman's origin, so `allow-same-origin` lets it read the Codeman page and call
the API that spawns agents. Only mark your own trusted code.
## Seam 2: SSE events
`GET /api/v1/events` is a Server-Sent Events stream. Each message is
`event: <name>` plus `data: <json>`. There are 149 event names following a
`domain:action` convention, registered in `src/web/sse-events.ts`.
The ones most integrations want:
| Event | Meaning |
| --- | --- |
| `session:created`, `session:deleted` | A session appeared or went away |
| `session:idle` | The agent stopped working |
| `session:completion` | A completion message was detected |
| `session:exit`, `session:error` | The session ended or failed |
| `hook:permission_prompt` | The agent is asking for permission |
| `hook:idle_prompt`, `hook:stop` | The agent is waiting on you, or stopped |
| `hook:task_completed`, `task:completed` | Work finished |
| `subagent:discovered`, `subagent:completed` | Background agent lifecycle |
| `mux:died` | A multiplexer session died unexpectedly |
| `cron:runCreated`, `cron:runUpdated` | Scheduled job activity |
### Filtering
`?sessions=id1,id2` suppresses only the high-volume `session:terminal` stream for
sessions you did not list. Lifecycle and metadata events are always delivered, so
you cannot accidentally filter away the thing you are listening for.
Pass `?clientId=<uuid>` to enable live filter updates through
`POST /api/v1/events/subscribe` without reconnecting the stream.
### Example: notify when any agent needs you
```js
const res = await fetch('http://127.0.0.1:3000/api/v1/events', {
headers: { Authorization: 'Basic ' + btoa(`admin:${process.env.CODEMAN_PASSWORD}`) },
});
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buf = '';
const WANTED = new Set(['hook:permission_prompt', 'hook:idle_prompt', 'session:idle']);
for (;;) {
const { value, done } = await reader.read();
if (done) break;
buf += decoder.decode(value, { stream: true });
const frames = buf.split('\n\n');
buf = frames.pop() ?? '';
for (const frame of frames) {
const name = frame.match(/^event: (.+)$/m)?.[1];
const data = frame.match(/^data: (.+)$/m)?.[1];
if (name && WANTED.has(name)) notify(name, JSON.parse(data ?? '{}'));
}
}
```
## Seam 3: HTTP API and CLI
Around 200 handlers across 21 route files cover sessions, cases, files, cron,
respawn, Ralph, the orchestrator, search, and admin. Each route module carries an
`@fileoverview` describing its endpoints.
If the caller is an agent running _inside_ a Codeman session, install the packaged
agent skill instead of teaching it these calls by hand: `skills/codeman` in the repo
(`npx skills add Ark0N/Codeman --skill codeman -g`, or `codeman skill install
[--case <name>]`, or the synced `agentSkillEnabled` App Setting for automatic
per-case injection on Claude session create). The skill carries the guard, the
safety rules, and verified wait/orchestration recipes.
The common ones:
```bash
# List sessions (live + persisted + transcript history, deduped)
curl -u admin:$PASS http://127.0.0.1:3000/api/v1/sessions/unified
# Create a session
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions \
-H 'Content-Type: application/json' \
-d '{"workingDir":"/home/me/project","mode":"claude"}'
# Send a prompt (single-line only, and it must end with \r: Enter is sent only
# when the input contains a carriage return; without it the text sits on the
# session's prompt unsubmitted)
curl -u admin:$PASS -X POST http://127.0.0.1:3000/api/v1/sessions/$ID/input \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","useMux":true}'
```
`POST .../input` also accepts `clientId` (stable per client, max 128 chars) and
`seq` (monotonic per session). Send both and the server applies each pair
at-most-once, so retrying after a dropped connection cannot type the prompt
twice. Omit them entirely rather than sending `null`.
It also accepts `wait` and `waitTimeout`, which hold the response open until the
session finishes the turn you just started. `wait` is `true` (the default signal
set) or a comma list of `idle,working,stop,blocked,exit`; the result comes back
under `data.wait`. Sending them changes nothing for callers that do not: without
`wait` the response is still `{"success": true, "data": {}}` and the write is still
fire-and-forget. The two interact with `clientId` / `seq` in one way worth knowing:
a **tagged duplicate** (a pair the server already applied) skips the write but still
waits, answering from the session's current state rather than blocking for a
transition that already happened. It reports `"delivered": false, "duplicate": true`.
### Waiting instead of polling
Three calls block until something happens: `GET /api/v1/sessions/:id/wait` (a
lifecycle signal), `GET /api/v1/sessions/:id/wait-output` (a literal string in the
output), and the `wait` field above. Full parameter and response tables are in
[`api-reference.md`](api-reference.md#long-polling-agent-wait). Four things decide
whether your integration works, and the last one is what actually bites:
- **A timeout is a `200` with `wait.timedOut: true`**, not an error. Loop over short
waits rather than issuing one long one, because `tailscale serve` and cloudflared
both cut idle connections and a single 10-minute call is the pattern most likely
to die in the field.
- **`wait.timeoutMs`** is the timeout after server-side clamping (600 s ceiling by
default). Read it rather than assuming you got what you asked for.
- **`stop` and `blocked` only exist for `claude` sessions**, and on a `shell` session
even `idle` fires only once at startup, so send-and-wait there can only time out.
See the Gotchas below.
⚠️ **There is no readiness signal, and skipping readiness is the failure that looks
like success.** A session reports `idle` before its CLI has spawned, and a `claude`
worker in a brand-new case comes up on the CLI's **trust dialog**, which has a ❯
prompt of its own. Prompt it at that moment and the text lands in the dialog, the
`\r` does not get past it, and the session's startup `idle` lands inside the wait
window: the wait resolves on `idle` in a couple of seconds with `timedOut: false`,
indistinguishable from a finished turn. Wait for the pid, then wait for the
composer, answering the dialog only as the bounded fallback.
A worked orchestration: start a worker, get it ready, prompt it, wait, clean up.
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}" # auto-set in-session, correct scheme included
AUTH=(-u "admin:$CODEMAN_PASSWORD") # omit entirely if no password is set
CURL=(curl -sk "${AUTH[@]}") # -k: harmless on http, required on --https installs (self-signed cert)
# 1. Start a worker session (creates the case if it does not exist yet).
# The guard matters: a TLS or auth failure otherwise leaves SID empty and every
# later step "succeeds" against nothing.
SID=$("${CURL[@]}" -X POST "$API/api/v1/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"worker-1","mode":"claude"}' | jq -r '.data.sessionId')
[ -n "$SID" ] && [ "$SID" != null ] || { echo "quick-start failed"; exit 1; }
# 2. READINESS: composer marker first, trust dialog only as the bounded fallback.
# Skip this and step 3 reports a turn that never ran. Do NOT probe trust first
# and Enter blindly: the dialog text stays in the buffer for the life of the
# session, so on every later run that probe matches stale text and the Enter
# lands in a ready composer. Match single tokens only: TUI text can arrive
# without its spaces. Stage 1 is short on purpose (an already-trusted case
# matches in <1 s; a first-run case can never pass it and pays it in full).
until [ "$("${CURL[@]}" "$API/api/v1/sessions/$SID" | jq '.data.pid')" != null ]
do sleep 1; done
R=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=5000') # composer's status bar = ready
if ! jq -e '.data.wait.matched' <<<"$R" >/dev/null; then
T=$("${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=trust' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=2000')
jq -e '.data.wait.matched' <<<"$T" >/dev/null && \
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' -d '{"input":"\r","useMux":true}' >/dev/null
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode 'match=bypass' --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=45000' >/dev/null
fi
# 3. Send the prompt AND register the wait in one call, so the answer cannot be
# the previous turn's idle state. Single line only, ending in \r (otherwise
# Enter is never sent and this wait times out on a turn that never started).
W=$("${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize the failures\r","useMux":true,
"clientId":"orchestrator","seq":1,"wait":"stop,exit","waitTimeout":60000}' \
| jq -c '.data.wait')
# 4. That first wait probably timed out (60 s). Keep going in SHORT waits.
for _ in $(seq 1 30); do
[ "$(jq -r '.timedOut' <<<"$W")" = 'true' ] || break # signal fired, or wait ended
W=$("${CURL[@]}" \
"$API/api/v1/sessions/$SID/wait?until=stop,exit&timeout=60000" | jq -c '.data.wait')
done
jq -r 'if .ended or .aborted then "worker is not running"
elif .timedOut then "still working after 30 waits"
else "signal: \(.signal)" end' <<<"$W"
# 5. Read what it produced, then delete the session YOU created, by exact id.
# ⚠️ NOT /output: its textOutput is empty for every tmux-backed session.
# `tail` counts BYTES, and the payload is terminal data with ANSI in it.
"${CURL[@]}" "$API/api/v1/sessions/$SID/terminal?tail=8000" | jq -r '.data.terminalBuffer'
"${CURL[@]}" -X DELETE "$API/api/v1/sessions/$SID"
```
Waiting on a marker instead of a signal is the form that works in **every** mode,
and the only one that works on a `shell` session:
```bash
# ⚠️ Split the marker so the typed line never contains it: your own keystrokes echo
# into the output stream, so an unsplit marker matches before the command has run.
# `from=buffer` also catches a marker that printed before the wait registered.
N=$RANDOM
"${CURL[@]}" -X POST "$API/api/v1/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d "{\"input\":\"M=DONE; npm test; echo \${M}_$N rc=\$?\r\",\"useMux\":true}"
"${CURL[@]}" -G "$API/api/v1/sessions/$SID/wait-output" \
--data-urlencode "match=DONE_$N" --data-urlencode 'from=buffer' \
--data-urlencode 'timeout=60000' | jq '.data.wait'
```
For shell scripting, the `codeman` CLI is the same surface without the HTTP
plumbing:
```
codeman session start|stop|list|logs codeman task add|list|status|remove|clear
codeman ralph start|stop|status|reset codeman users add|passwd|list
codeman status | list | attach <path> codeman doctor
```
## Seam 4: Hooks
Claude Code hooks post to `POST /api/v1/hook-event` from inside an agent session.
Codeman installs its own hooks automatically, but the endpoint is open to yours.
```json
{ "event": "task_completed", "sessionId": "abc123", "data": { "any": "json" } }
```
`event` must be one of `permission_prompt`, `elicitation_dialog`, `idle_prompt`,
`stop`, `teammate_idle`, `task_completed`. Each becomes the matching `hook:*` SSE
event.
⚠️ This endpoint skips Basic auth so hooks keep working, but when auth is active
the loopback bypass requires the `X-Codeman-Hook-Secret` header
(`~/.codeman/hook-secret`) unconditionally.
## Gotchas
Every one of these has cost somebody real time.
- **CORS is localhost-only.** `Access-Control-Allow-Origin` is echoed only for
`localhost`, `127.0.0.1`, and `::1`. A browser app on any other origin cannot
call the API. Integrate server-side.
- **A missing `Origin` header is allowed**, which is why curl, CLIs, and hooks
work. Cross-site origins are blocked by the CSRF guard.
- **Reverse-proxy domains are rejected** by the anti-DNS-rebinding Host allowlist
unless added via `CODEMAN_ALLOWED_HOSTS=host,.suffix`.
- **`null` is not `undefined`.** Request schemas use Zod `.optional()`, which
accepts `undefined` only. `JSON.stringify({ field: null })` keeps the null on
the wire and fails with `INVALID_INPUT`. Omit the key instead. This has caused
shipped bugs more than once.
- **`text/plain` bodies stay raw.** Auto-parsing them as JSON enabled
simple-request CSRF, so it is deliberate. Send `application/json`.
- **Prompts are single-line and must end with `\r`.** The server splits your text
and Enter into two separate tmux writes (Ink needs them apart), but it sends the
Enter **only when the input contains a carriage return**. Without it your text
sits on the prompt unsubmitted, which is the single most common "the wait
endpoints don't work" report: the wait runs its full timeout on a turn that never
started. Newlines inside the string are stripped rather than rejected, so
`"echo A\necho B\r"` runs the single joined command `echo Aecho B`: send one line
per call.
- **`wait-output`'s `from=now` is not "printed after you asked".** tmux repaints
the visible screen on attach, on resize, and on any TUI redraw, and a repaint
arrives as ordinary output, so text already on screen can satisfy a fresh wait.
Observed live: a marker echoed a minute earlier matched instantly. Use a marker
unique to each call, and build it so the typed line never contains it (your own
keystrokes echo into the stream). Matching is a literal substring, so `regex=` is
rejected with a `400` rather than ignored.
- **`wait-output` matches the normalized PTY stream, not the screen.** ANSI escape
sequences are stripped (the `ESC ( B` charset escape a bash prompt emits on every
line included), a partial escape at a chunk boundary is held back until its tail
arrives, and a match may straddle PTY chunks, so text you printed yourself
matches reliably (`printf STRAD; sleep 1; printf DLEQQ` is matchable as
`STRADDLEQQ`). What can still fail is TUI output: a full-screen TUI positions
words with cursor moves, so its text can reach the matcher **without spaces** and
a multi-word match is unreliable there. Match one short space-free token, ideally
one you printed yourself, and keep it out of the typed line (your own keystrokes
echo into the stream).
- **`stop` and `blocked` never fire for `shell`, `opencode`, `codex`, `gemini`,
`antigravity` or `pi` sessions.** They come from Claude Code hooks, which no other mode
installs, so only `idle`, `working` and `exit` exist there. Asking for them
explicitly is a `400`; omitting `until` is safe, since the server drops them from
the default set and echoes what it actually waited on as `wait.until`. Even in
`claude` mode, a Docker case needs `CODEMAN_DOCKER_BRIDGE_HOOKS=1` for hooks to
reach the server at all, a remote-SSH case's hooks may never arrive, and a case
written by Codeman < 1.13.0 against an `--https` install carries hook curls
without `-k` that TLS-fail silently — a 1.13.0+ server rewrites them the next
time a session starts in that case.
- **Unwrap the envelope** before reading fields. `data` is not the response body.
## Publishing your integration
There is no registry and no review queue. Add the GitHub topic
**`codeman-integration`** to your public repository so others can find it, and
link back to Codeman in your README.
If a real ecosystem of these appears, a manifest format and an install command
become worth building. Until then, these four seams are the contract, and they
require nothing of you but HTTP.
## What Codeman deliberately does not have
- **No in-process plugin runtime.** See the reasoning at the top of this page.
- **No build or startup hooks** for third-party code. Run your own process.
- **No per-plugin config or state directories.** Manage your own files.
- **No sandbox for integration code**, because Codeman never launches it. Your
integration is your own process, started by you, with your permissions,
talking HTTP.
That last point is about integration code specifically, not about Codeman.
Sandboxing lives on a different axis here: the thing worth isolating is the
**agent**, and you isolate it per case with
[Docker cases](docker-cases.md), which run the agent in a hardened container with
a bind-mounted workspace and seeded (not shared) credentials. An integration that
creates or drives a Docker-backed session inherits that isolation for free, since
it is a property of the session rather than of the caller.
+432
View File
@@ -0,0 +1,432 @@
# File Viewer edit mode (issue #212)
Plan only. No implementation yet.
Goal: close the loop "agent writes a file, you review it in the viewer, tweak two lines, save, tell the
agent to continue" without hopping into the terminal, with the phone as the primary target.
Scope from the issue: an Edit toggle on text previews, a write endpoint that inherits the read path's
confinement, text-only, edit-in-place (no create, no delete, no rename), no editing through the
Docker/remote overlays.
---
## 1. What exists today
**Read path (backend), all in `src/web/routes/file-routes.ts`:**
| Route | Line | Notes |
| ------------------------------------ | ------ | ------------------------------------------------------------------ |
| `GET /api/sessions/:id/files` | `741` | Tree scan of `session.workingDir`, hidden files off by default |
| `GET /api/sessions/:id/file-content` | `865` | The text/preview classifier. `findSessionOrFail` + `validateSessionFilePath` |
| `GET /api/sessions/:id/file-raw` | `1018` | Bytes, 50MB cap |
| `GET /api/sessions/:id/file-preview` | `1254` | DOCX/PPTX to PDF, everything else redirects to `file-raw` |
| `GET /api/download` | `1384` | The only read route that also runs `isSensitivePath()` |
`file-content` classification order (`file-routes.ts:881-1011`): extension buckets (image / video / audio /
known-binary) return metadata only; otherwise the bytes are read, sniffed for a NUL in the first 8KB, and
either reported as `type:'binary'` or decoded as UTF-8 and **truncated to `lines` (default 500, hard cap
10000)**. Caps: `MAX_TEXT_FILE_SIZE` 10MB.
Confinement is `validateSessionFilePath()` (`src/web/route-helpers.ts:67`): `resolve()` then `realpathSync()`
then reject if the result is not under `workingDir`. Because it realpaths the *full* path, a symlink whose
target escapes the workspace is already rejected. Ownership is `findSessionOrFail()` which runs
`canAccessOwned()` (`route-helpers.ts:102`), a no-op outside multi-user mode.
**Read path (frontend), `src/web/public/panels-ui.js`:**
- `loadFileBrowser()` `2947`, `renderFileBrowserTree()` `2978`, click to `openFilePreview()` `3056`.
- `openFilePreview(filePath, sessionId, attachmentId)` `3193`: attachment-id branch, then docx/pptx, pdf,
svg branches, then the generic `file-content` fetch at `3274` with **`&lines=500` hardcoded**, rendering
text as `<pre><code>${escapeHtml(...)}</code></pre>` at `3298` and stashing `this.filePreviewContent`.
- `closeFilePreview()` `3308`, `copyFilePreviewContent()` `3751`.
- Markup: `src/web/public/index.html:420-432` (`filePreviewOverlay` / `-Title` / `-Body` / `-Footer`, two
header buttons: copy and close).
- CSS: `src/web/public/styles.css:9320-9430`. Overlay `z-index: 2000`, window `80vw/80vh`, capped
`900x700`. There are **no `.file-preview-*` rules in `mobile.css` at all**.
**Reachability on phones.** The header File Viewer button is hidden below 430px
(`mobile.css:482`, locked by `KNOWN_PHONE_HIDDEN` in `test/mobile-header-buttons-policy.test.ts`), so on a
phone the preview overlay is reached through:
1. an attachment card's **Preview** button (`panels-ui.js:3451`), which is exactly the "agent just wrote a
file" path the issue describes,
2. the attachment-history drawer (`panels-ui.js:3709`),
3. App Settings to Panels to **File Browser** (`showFileBrowser`, applied in `settings-ui.js:2202`; the
panel is mobile-styled at `mobile.css:1868`).
So edit mode is reachable on a phone today via (1) and (2) without touching the header policy. Improving
the entry point is listed as an open decision in section 10, not assumed.
---
## 2. Threat model, stated honestly
Anyone who can call this API can already reach `POST /api/sessions/:id/input` and type an arbitrary prompt
into an agent running with `--dangerously-skip-permissions`. A workspace-confined write endpoint therefore
does not create a new privilege tier for an authenticated caller.
What it *would* create if built carelessly is a **new host-write primitive reachable by path**, so the
things this plan actually defends against are:
1. **Path traversal / symlink escape** writing outside the workspace.
2. **TOCTOU**: a path component that becomes a symlink between validation and write.
3. **Cross-user writes** in multi-user mode (`canAccessOwned`).
4. **Silent data loss**, which is the highest-probability real-world failure here and gets its own section.
CSRF is already covered: `registerHostGuard()` (`src/web/middleware/auth.ts:555-578`) rejects any
non-safe-method request whose `Origin` is cross-site. The webview-capability exemption at that gate is
fenced to `GET`/`HEAD` for the Referer form (`auth.ts:161`) and to `/webview/:cap/*` paths for the path
form, so a proxied dashboard cannot reach a new `PUT /api/...`. Using `PUT` + `application/json` also
forces a preflight for any cross-origin attempt.
---
## 3. Backend design
### 3.1 New policy module: `src/config/file-editing.ts`
Pure, unit-testable, no IO (config lives in `src/config/`, no barrel, import the file directly).
```ts
export const MAX_EDITABLE_BYTES = 512 * 1024; // content cap, both directions
export const EDITABLE_EXTENSIONS: ReadonlySet<string>; // ts,tsx,js,jsx,mjs,cjs,json,jsonc,md,mdx,txt,
// css,scss,less,html,htm,xml,svg?,yml,yaml,toml,
// ini,cfg,conf,env?,sh,bash,zsh,fish,py,rb,go,rs,
// java,kt,swift,c,h,cpp,hpp,cs,php,sql,graphql,
// proto,lua,pl,r,jl,tf,gradle,csv,tsv,log,diff,patch
export const EDITABLE_BASENAMES: ReadonlySet<string>; // Dockerfile, Makefile, LICENSE, .gitignore,
// .prettierignore, .editorconfig, .nvmrc, ...
export function isEditableFileName(fileName: string): boolean;
export function isDeniedEditRelativePath(rel: string): boolean; // `.git/` subtree
export function detectEol(text: string): 'lf' | 'crlf';
export function applyEol(text: string, eol: 'lf' | 'crlf'): string;
```
Decisions baked in:
- **Allowlist, not blocklist**, per the issue and per the existing attachment-guard precedent.
- `svg` and `env` are deliberately marked with `?` above: `svg` is served as an untrusted octet-stream on
the read side (`file-routes.ts:118`) so allowing an edit is defensible, but I recommend **excluding
both** in v1. `.env` files are matched by `isSensitivePath()` anyway and would be rejected downstream;
excluding them at the allowlist keeps a single obvious refusal.
- `isDeniedEditRelativePath` blocks the `.git/` subtree: `.git/hooks/*` is code execution and a corrupt
index is unrecoverable-looking to a user who only wanted to fix a typo. Other dotfiles stay allowed but
are not reachable from the tree UI anyway (`showHidden=false`).
### 3.2 Read-for-edit: extend the existing GET
`GET /api/sessions/:id/file-content?path=<rel>&edit=1`
When `edit=1`:
- skip line truncation entirely (a truncated buffer must never become an edit buffer, see section 4.1),
- enforce `MAX_EDITABLE_BYTES` instead of `MAX_TEXT_FILE_SIZE` and answer 413 over it (as a structured
throw with `statusCode: 413`, the `throwFilesystemPickerError` pattern, since the central errorCode-to-
status map has no 413 entry; see the error-mechanics note in 3.3),
- run the editability gate (`isEditableFileName`, `isDeniedEditRelativePath`, `isSensitivePath`,
`isBlockedAttachmentPath`) and the content gate (NUL sniff plus UTF-8 round-trip, see 4.3),
- return `{ content, size, mtimeMs, totalLines, truncated: false, extension, editable: true, hash, eol }`.
`hash` is `sha256` hex of the exact on-disk bytes.
Non-`edit` responses gain **only** `editable: boolean` (additive, no shape change for existing consumers),
which is all the UI needs to decide whether to show the Edit button. No `hash` on plain reads: the Edit
action re-fetches with `edit=1` anyway (section 4.1), which is where the hash comes from, and hashing every
casual 10MB preview would be pure waste.
### 3.3 Write: `PUT /api/sessions/:id/file-content`
Body (new `FileWriteSchema` in `src/web/schemas.ts`, Zod v4):
```ts
{ path: string, content: string, baseHash: string, eol?: 'lf'|'crlf', force?: boolean }
```
Registered with an explicit route option `{ bodyLimit: 4 * 1024 * 1024 }`. **Fastify's default `bodyLimit`
is 1MB and this repo configures none**, and JSON escaping expands content: 2x for a file full of quotes or
backslashes, up to 6x for control characters (each serialized as a `\uXXXX` escape), so 512KB of content
can legitimately exceed 1MB on the wire; blowing the limit produces a raw `FST_ERR_CTP_BODY_TOO_LARGE`, not an `ApiResponse` envelope. Two
related sizing notes: `z.string().max()` counts **UTF-16 code units, not bytes**, so the schema's `.max()`
is only a coarse pre-filter and the real cap is an explicit `Buffer.byteLength(content, 'utf8')` check in
the handler (step 7a below); and 4MB comfortably bounds the worst-case expansion of a 512KB file without
inviting multi-MB bodies elsewhere.
**Error mechanics** (matters for both prod behavior and testability): a handler that *returns* a
`{success:false, errorCode}` envelope gets its HTTP status assigned centrally by the preSerialization hook
in `server.ts` (`httpStatusForErrorCode()`, `src/types/api.ts`), but the route-test harness
(`test/routes/_route-test-utils.ts`) installs only `installRouteErrorHandler`, **not** that hook, so
returned envelopes surface as HTTP 200 in tests. The PUT handler should therefore use the same
structured-**throw** pattern as the filesystem picker (`throwFilesystemPickerError`, `file-routes.ts:411`):
thrown `{statusCode, body}` errors are rendered identically in prod and in the harness, and they allow the
one status the code map cannot express (413). The error envelope itself is strictly
`{success:false, error, errorCode}`, **it has no data arm**, so no error response may carry extra payload.
Handler order (each step is a test case):
1. `findSessionOrFail(ctx, id, req)` (live sessions only, matching the read route, and it carries the
multi-user ownership check).
2. `parseBody(FileWriteSchema, req.body)`, then `Buffer.byteLength(content, 'utf8') <= MAX_EDITABLE_BYTES`
or 413 (the schema `.max()` alone cannot enforce a byte cap, see the sizing note above).
3. `validateSessionFilePath(session.workingDir, path)` or 404 (do not distinguish "outside workspace" from
"missing", matching the read route).
4. `isSensitivePath(resolvedPath) || isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)` or 403.
5. `isDeniedEditRelativePath(relativePath)` or 403.
6. `isEditableFileName(basename(resolvedPath))` or 400.
7. `stat`: must be `isFile()`, size within `MAX_EDITABLE_BYTES`, else 400/413. **No `O_CREAT` anywhere in
this handler**, which is what enforces edit-in-place.
8. Read current bytes, compute `hash`, run the NUL sniff and the UTF-8 round-trip check, else 400.
9. `hash !== baseHash && !force` gives **409 CONFLICT** (`ApiErrorCode.CONFLICT`, plain envelope; the error
arm carries no data, see the error-mechanics note). The client's conflict dialog gets fresh state by
re-fetching `edit=1`, which it needs for its Reload action anyway.
10. Build the output buffer: `applyEol(content, eol ?? detected-from-original)`; re-check
`Buffer.byteLength` against the cap.
11. Write atomically in the resolved parent directory:
`fs.open(<dir>/.<name>.codeman-tmp-<rand>, 'wx', stat.mode & 0o777)`, then `fchmod(stat.mode & 0o777)`
(open's mode argument is masked by the process umask, so the chmod is what actually preserves an
unusual mode), write, `fsync`, close, `fs.rename(tmp, resolvedPath)`, unlink the temp on any failure.
12. Re-stat, return `{ success: true, data: { path, size, mtimeMs, hash, totalLines } }`.
Why `O_EXCL` temp plus rename rather than truncate-in-place:
- `wx` cannot follow a pre-existing symlink, which closes the TOCTOU window from step 3 to step 11 without
needing `O_NOFOLLOW` gymnastics.
- `rename()` does not follow a symlink in the final component, so even if `resolvedPath` were swapped for a
symlink after validation, the symlink itself is replaced and the swap target is untouched.
- A crash mid-write leaves the original intact.
Caveat to document in the code comment: rename replaces the inode, so hardlinks to the file keep the old
content. That is the same trade-off vim makes by default and is preferable to a truncate window here.
No SSE event in v1. Nothing else in the app needs to know: `image-watcher.ts` only reacts to
`.png/.jpg/.jpeg/.gif/.webp/.bmp/.svg/.pdf/.docx/.pptx` adds (`image-watcher.ts:23-25`), none of which are
editable text, and the temp filename does not match either.
---
## 4. The five traps
These are the parts that turn a "small write endpoint" into a bug report.
### 4.1 Truncation (the data-loss trap)
The frontend fetches `&lines=500` (`panels-ui.js:3274`). Saving that buffer back would **delete every line
past 500**. Worse, the content hash of the full file would still match, so an optimistic-concurrency check
cannot catch it.
Mitigations, all three:
- The Edit affordance is only offered when the loaded payload came from `edit=1` (which never truncates).
Tapping Edit on an already-rendered preview **re-fetches** with `edit=1` before swapping in the editor.
- The read-for-edit path 413s above `MAX_EDITABLE_BYTES` rather than truncating, so "too big to edit here"
is an explicit refusal with a message, never a silent partial buffer.
- A test asserts `edit=1` never returns `truncated: true`.
### 4.2 Line endings
A `<textarea>`'s `.value` normalizes to LF. Saving a CRLF file naively rewrites every line, producing a
whole-file diff for a two-line change. So: the read returns the detected `eol`, the client echoes it back
unchanged, and the server re-applies it. Mixed-EOL files use the dominant style, which is lossy for the
minority lines; call that out in the response and accept it in v1.
### 4.3 Encoding
`buf.toString('utf-8')` on a latin-1 or otherwise non-UTF-8 file yields U+FFFD replacement characters, and
writing that back **corrupts the file**. The check is a round-trip:
`Buffer.from(decoded, 'utf8').equals(buf)`. If it fails, `editable: false` and the write is refused. This
also catches binary content that the NUL sniff misses. A UTF-8 BOM survives because it round-trips as a
leading U+FEFF; do not strip it.
### 4.4 Concurrency with the agent
The whole use case is editing a file the agent just wrote and may write again. `baseHash` plus 409 is the
guard. Do not use mtime alone: agents rewrite files within a single filesystem timestamp tick, and an
identical rewrite should not be reported as a conflict.
### 4.5 Symlinks and TOCTOU
Covered by `validateSessionFilePath` (escape) plus `wx` temp and `rename` (post-validation swap). One
intentional allowance: a symlink whose target is *inside* the workspace is edited through to its target,
because `validateSessionFilePath` returns the realpath. That matches what a user tapping the file expects.
---
## 5. Frontend design
All in `panels-ui.js` (prettier-exempt, hand-formatted; match the surrounding style), `index.html`,
`styles.css`, `mobile.css`.
### 5.1 State
```js
filePreviewEdit = { active, sessionId, path, baseHash, eol, original, dirty }
```
Reset in `closeFilePreview()` and on every `openFilePreview()` entry.
### 5.2 Markup (`index.html:420-432`)
Add one header button (pencil, `btn-icon-sm`, `id="filePreviewEditBtn"`, hidden by default) next to the
copy button, and an edit bar inside the footer region holding Save / Cancel / a dirty dot. Keep the
existing footer text element; the edit bar is a sibling toggled by class so the read-mode footer is
untouched.
### 5.3 Behavior
- `openFilePreview()` shows the Edit button only when the response has `editable: true` and the render took
the text branch. Attachment-id previews, media, binary, pdf, docx/pptx and svg all leave it hidden.
- **Enter edit**: re-fetch with `edit=1`; on 413 or `editable:false`, toast the reason and stay in read
mode. This fetch must **parse the error envelope on non-ok responses**: the existing generic
`if (!res.ok) throw new Error('Failed to load file')` pattern (`panels-ui.js:3275`) would swallow the
specific "too large to edit here" message, since error envelopes arrive with real 4xx statuses in prod. On success replace the body with `<textarea class="file-preview-editor" spellcheck="false"
autocapitalize="off" autocorrect="off" autocomplete="off" wrap="off">` and assign `.value = content`
(never `innerHTML`, so no escaping question arises). Do **not** autofocus: on a phone that opens the
keyboard before the user has picked a line.
- `input` sets `dirty` and enables Save.
- **Save**: `PUT` with `baseHash`, `eol`, and `content`. On success update `baseHash`/`original` from the
response, leave edit mode, re-render the read view from the local editor value (the response carries
metadata only, not content), toast "Saved". On **409** offer `Reload (discard mine)` / `Overwrite`:
Reload re-fetches `edit=1` and replaces the buffer; Overwrite re-sends with `force: true`. The 409 body
itself carries no state (section 3.3, step 9).
- **Cancel / close / Escape while dirty**: `confirm('Discard unsaved changes?')`, consistent with the
existing `window.confirm` usage in this codebase (`panels-ui.js:4323`, `app.js:4176`). Note the global
Escape handler (`app.js:999-1007`) closes other panels via `closeAllPanels()` but does not touch this
overlay today; if Escape-to-close is wired up as part of this work it must go through the same dirty
guard.
- `copyFilePreviewContent()` copies the live editor value while editing.
⚠️ Repo gotcha to respect at the fetch call: **Zod `.optional()` rejects `null`**. Build the body with
`eol: eol ?? undefined` (or declare `.nullish()`), or the PUT fails `INVALID_INPUT`. This has shipped as a
real bug twice.
### 5.4 Mobile
- **Sizing.** The window is `80vw/80vh` centered with no mobile override, so when the keyboard opens on iOS
the lower half sits behind it. Add a `@media (max-width: 430px)` block using
`height: var(--app-height, 100vh)`, full width, no border radius. `--app-height` is already maintained
against `visualViewport` by `KeyboardHandler.handleViewportResize()` (`mobile-handlers.js:283-317`), so
the editor tracks the keyboard for free.
- **iOS zoom.** The editor font must be >= 16px on phones; there is an existing zoom-prevention block at
`mobile.css` under `@media (max-width: 768px)`. Verify it covers `textarea` and do not override it with a
smaller `rem` value.
- **Accessory bar.** Focusing any input fires `KeyboardHandler.onKeyboardShow()`, which calls
`KeyboardAccessoryBar.show()` and refits/resizes the terminal (`mobile-handlers.js:407+`). The bar's keys
target the **terminal**, not the editor, so an Esc or clear-input tap while editing goes to the agent.
The overlay's `z-index: 2000` covers the bar's `51`, so it is not visible, but confirm it is not
interactive underneath and consider an explicit `KeyboardAccessoryBar.hide()` while the editor holds
focus. This is the item most likely to look "fine on desktop, wrong on the phone".
- No header-policy change is needed (section 1), so
`test/mobile-header-buttons-policy.test.ts` stays untouched.
### 5.5 i18n
`i18n.js` already skips `textarea`, `pre`, `code` and `.file-preview-content` in its `SKIP_SELECTOR`
(`i18n.js:20-38`), so file content is never translated. Add zh-CN entries for the new chrome: Edit, Save,
Cancel, Unsaved changes, Discard unsaved changes?, File changed on disk, Reload, Overwrite, Saved,
Too large to edit here.
---
## 6. Docker and remote cases
Out of scope per the issue, and the current behavior already degrades correctly:
- **Docker cases**: the workspace is a host directory bind-mounted at the same absolute path, so a host-side
write is visible in the container immediately. Edit mode works and needs nothing special. Worth one line
in the docs.
- **Remote SSH cases**: `workingDir` is a path on the remote host. `validateSessionFilePath` realpaths it
locally, which fails, so the write returns 404 exactly like the read routes do today. Confirm the viewer
shows a clean empty/error state rather than an unexplained failure, and do not attempt an SFTP path.
---
## 7. Tests
| File | Kind | Covers |
| ------------------------------------------- | ----------- | ---------------------------------------------------------------------- |
| `test/file-editing-policy.test.ts` | pure unit | `isEditableFileName` (allow + deny + basenames), `isDeniedEditRelativePath`, `detectEol`/`applyEol` round-trip incl. mixed EOL, BOM preservation |
| `test/routes/file-write-routes.test.ts` | `app.inject` | The handler order in 3.3, against a **real temp dir** (do not `vi.mock('node:fs')` in this file; set `MockSession.workingDir`, `test/mocks/mock-session.ts:14`) |
| extend `test/routes/file-routes.test.ts` | `app.inject` | `edit=1` never truncates; `editable` present on the plain read |
Status-code caveat for all of these: the route-test harness does not install the server's preSerialization
envelope hook, so a handler that *returns* an error envelope answers 200 in tests. The statuses below are
only assertable because the plan has the handler **throw** structured errors (section 3.3, error
mechanics), which `installRouteErrorHandler` renders identically in prod and in the harness.
Route cases to assert explicitly:
1. happy path writes the bytes and returns a new hash
2. `../` and absolute paths give 404
3. symlink pointing outside the workspace gives 404
4. symlink pointing inside is written through to the target
5. non-allowlisted extension gives 400
6. `.git/config` gives 403
7. a `.env` in the workspace gives 403 (sensitive-path)
8. a file with a NUL byte gives 400
9. a latin-1 file that fails the UTF-8 round-trip gives 400
10. stale `baseHash` gives 409 (`CONFLICT` envelope, no data); `force:true` then succeeds
11. over `MAX_EDITABLE_BYTES` gives 413
12. a path that does not exist gives 404 and creates nothing (no `O_CREAT`)
13. multi-user: `authUser: {role:'user'}` against another user's session gives 404 (pass `authUser` to
`createRouteTestHarness`, otherwise the synthetic admin makes the test pass vacuously)
14. CRLF file edited and saved stays CRLF
15. file mode is preserved across the temp-plus-rename
Run with `npm test -- test/routes/file-write-routes.test.ts`, never bare `npm test`.
**End-to-end verification before any deploy** (unit tests passing is not sufficient here):
- `curl -sk https://localhost:3000/...` against a **throwaway** session created for the purpose, never
`w1`/`w2`/`w3`; delete it by exact id afterwards.
- Playwright on a phone profile: open a preview, tap Edit, type with `page.keyboard.type()`, Save, then
assert the bytes on disk changed. Assert real state, not HTTP 200.
---
## 8. Docs and release
- This plan lives at `docs/file-viewer-edit-plan.md`.
- `docs/architecture-invariants.md`: new anchor `#file-viewer-edit-mode` covering the write confinement
chain, the truncation invariant, and why temp-plus-rename.
- `CLAUDE.md`: one line under the **Filesystem path picker** neighborhood noting that the File Viewer now
has a **third** file surface and that it is the only one that writes, plus its confinement rules.
Remember `CLAUDE.md` is prettier-ignored on purpose.
- `docs/api-reference.md`: the new `PUT` and the `edit=1` query.
- Release: a normal COM applies (the 1.10.0 batch hold is over). This is a new user-facing feature plus an
additive API surface, so **COM minor** when it ships.
Formatting note: `panels-ui.js`, `styles.css`, `mobile.css`, `index.html` are all in `.prettierignore` and
are hand-formatted; new TypeScript (`src/config/file-editing.ts`, route + schema edits) is prettier-enforced
and must pass `npm run format:check`.
---
## 9. Implementation order
Each phase is independently reviewable and leaves the tree working.
1. **Policy module + tests.** `src/config/file-editing.ts` and `test/file-editing-policy.test.ts`. Pure, no
route wiring. (Small.)
2. **Read-for-edit.** `edit=1` (returning `hash`/`eol`) plus the additive `editable` flag on plain reads,
tests. Nothing consumes it yet. (Small.)
3. **Write endpoint.** `FileWriteSchema`, `PUT` handler, `test/routes/file-write-routes.test.ts`. Fully
testable by curl before any UI exists. (Medium, the security-relevant part.)
4. **Desktop UI.** Edit button, textarea swap, Save/Cancel, dirty guard, 409 flow. (Medium.)
5. **Mobile pass.** `mobile.css` sizing against `--app-height`, font size, accessory-bar interaction,
real-device check. (Small but the part that decides whether the feature is actually usable.)
6. **Docs, i18n strings, changeset.**
---
## 10. Open decisions
1. **Editor widget.** Recommend a plain `<textarea>` for v1: zero dependencies, no CSP question, no bundle
growth, and it is the only thing guaranteed to behave with the iOS keyboard. CodeMirror-light with
syntax highlighting is a clean follow-up once the write path is proven. The issue allows either.
2. **Phone entry point.** Edit mode is reachable on a phone through attachment cards and the history
drawer without changing anything. A dedicated toolbar or overview affordance for "browse this session's
files" would make it discoverable, but it is a separate UX change and would need a decision against the
deliberately minimal phone header policy. Recommend deferring it and revisiting after the feature ships.
3. **`svg` editability.** Recommend excluded in v1 (it is deliberately treated as untrusted on the read
side). Easy to add later.
4. **Create / delete / rename.** Explicitly out of scope per the issue. Note that keeping `O_CREAT` out of
the handler is what makes that a structural property rather than a convention.
Binary file not shown.

After

Width:  |  Height:  |  Size: 34 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 207 KiB

+681
View File
@@ -0,0 +1,681 @@
# Pi (pi.dev) Run Mode: Implementation Plan
Tracking issue: [#206 "Plans to support pi.dev?"](https://github.com/Ark0N/Codeman/issues/206)
Status: **IMPLEMENTED 2026-08-13** (see `docs/pi-integration.md` for the user-facing
guide). Everything below is the design record; the open questions were resolved
empirically against pi 0.84.1 and the answers are recorded inline as **RESULT**
notes. Originally reworked 2026-08-06; **rechecked 2026-08-13 against master @
`f39beb3` (v1.17.0)**, and every line anchor below was re-verified at that commit (the 1.11.2-era
anchors drifted heavily: six releases landed in between, including the settings-surface overhaul and
the codex predictive-echo work, both of which added new pi touchpoints, §2.10 and the Brain picker in
Phase 3). Upstream facts verified against `@earendil-works/pi-coding-agent` **v0.84.1** (npm latest,
published 2026-08-07) and the [`earendil-works/pi`](https://github.com/earendil-works/pi) repo (cite
that name: upstream docs still contain stale `pi-mono` links from a repo rename). Line numbers are
anchors for orientation, not contracts; they drift.
---
## 1. What Pi is
[Pi](https://pi.dev) (MIT) is a minimal, extensible coding-agent harness. Facts below are verified
against the upstream docs in `packages/coding-agent/docs/`.
| Property | Value |
| ---------------- | -------------------------------------------------------------------------------------------------- |
| Binary | `pi` (`bin: { pi: 'dist/cli.js' }`) |
| npm package | `@earendil-works/pi-coding-agent`, latest **0.84.1** (2026-08-07; 0.84.0 was 2026-08-06); `legacy-node20` dist-tag at 0.74.2 |
| Install | `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`, or `curl -fsSL https://pi.dev/install.sh \| sh` (the curl installer also goes through global npm, so both uninstall via npm) |
| Config dir | `~/.pi/agent` (override: `PI_CODING_AGENT_DIR`). Holds `auth.json`, `trust.json`, `settings.json`, `models.json` (user-defined providers), `models-store.json` (cached catalogs), `keybindings.json`, `extensions/`, `skills/`, `prompts/`, `themes/`, `AGENTS.md`, `SYSTEM.md`, and the package trees `npm/` + `git/` |
| Sessions | `~/.pi/agent/sessions/--<cwd with / replaced by ->--/<timestamp>_<uuid>.jsonl`, tree-structured (`id`/`parentId`), format v3. Overrides: `PI_CODING_AGENT_SESSION_DIR`, `--session-dir` |
| Credentials | `~/.pi/agent/auth.json` (OAuth subscriptions + API keys, auto-refresh), plus ~34 provider env vars with **no common prefix**. 0.84.1 adds `pi auth check` (auth preflight with optional credential output) |
| TUI | Default: **main screen with terminal-owned scrollback**. Since **0.84.0** an experimental fullscreen mode exists, selectable via `--tui-mode fullscreen` **or at runtime through `/settings`**; the default remains the main-screen mode |
| Providers | 15+ (Anthropic, OpenAI, Google, Azure, Bedrock, Mistral, Groq, xAI, OpenRouter, Copilot, Baseten since 0.84.0, ...). OAuth subscription login via `/login` for six: ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, xAI, OpenRouter, Radius |
| Permission model | **No permission prompts at all.** No built-in sandbox, no MCP (none planned), no sub-agents, no plan mode, no to-dos, no background bash. Tools run with the user's own permissions |
| Trust model | "Project trust" gates **loading** of project-local `.pi/` config/extensions/skills and **installing missing project packages**, not tool execution. Triggered only when the cwd (or an ancestor) contains `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`.pi/APPEND_SYSTEM.md`, or `.agents/skills`; a bare `.pi/` directory does NOT prompt. Global `defaultProjectTrust`: `ask` (default) / `always` / `never` |
Three consequences shape the whole integration:
1. **There is no `--dangerously-skip-permissions` analog and none is needed.** Pi never prompts for
tool approval. The Claude/Codex/Gemini/Antigravity pattern of "send the bypass flag so the session
is not stuck on a modal" does not apply. Codeman must not invent a flag here.
2. **The one privileged knob is `--approve` / `-a`** (trust project-local files for this run), which
makes pi load and execute project `.pi/extensions` TypeScript **and run an npm install of missing
project packages**. That is the field the multi-user clamp has to cover. Its explicit inverse
`-na` / `--no-approve` exists, which lets the clamp force-deny rather than merely omit (§3, §5.2).
3. **Provider keys cannot ride the env allowlist.** Pi's provider key vars (`ANTHROPIC_API_KEY`,
`OPENAI_API_KEY`, `DEEPSEEK_API_KEY`, `HF_TOKEN`, `BASETEN_API_KEY`, ...) share no prefix, so
there is no way to admit them through `ALLOWED_ENV_PREFIXES` without widening the list for every
mode (§2.4).
---
## 2. Design decisions
### 2.1 Mode identity
`SessionMode` gains `'pi'`. Not a location overlay (unlike Docker/remote-SSH cases), not a web tab:
a real sixth CLI backend with its own PTY, tmux session and respawn behaviour, exactly like
`antigravity`. Append `pi` after `antigravity` in every enum/list to keep ordering consistent.
| Surface | Value |
| ---------------- | --------------------------------------------------------------------- |
| `SessionMode` | `'pi'` |
| Display label | `Pi` |
| Tab badge | `pi` (two-letter lowercase, like `sh`/`oc`/`cx`/`gm`/`ag`) |
| Run button label | `Run PI` (short-label ternary in `_applyRunMode`, pattern `Run AG`) |
| Kill-menu label | `Kill Tmux & Pi` |
| Identity color | **`#f472b6` (rose-400)**. Verified free: live computed values on the default skin are claude `#38b6f0`, opencode `#44b993`, codex `#2b8fd9`, gemini `#8ab4f8`, antigravity `#22d3ee`, shell `#98a2b1`, web `#38bdf8`; purple is codex's base hex and amber reads as the shell tab badge, so pink/rose (or orange `#fb923c`) are the only genuinely free hues. No `pi` CSS identifier collides anywhere (`mode-pi`, `.tab-mode.pi`, `.run-mode-dot.pi` all grep clean, re-checked at f39beb3) |
| Env prefix | `PI_` |
| Dependency id | `pi` |
| Status endpoint | `GET /api/pi/status` |
### 2.2 `isExternalCliMode()` yes, `isAltScreenStripMode()` no
Pi joins `isExternalCliMode()` (`session.ts:164-167`): its own TUI, its own output format, so the
Ralph tracker, `BashToolParser`, token/CLI-info scraping and the `❯` readiness probe all stay off
(gates at `session.ts:1100`, `:1701`, `:2000`, `:2103`), and readiness falls back to the output
stabilization used by the other external CLIs.
Pi stays **out** of `isAltScreenStripMode()` (`session.ts:197-199`, currently codex/claude/gemini;
antigravity and opencode are deliberately excluded). Pi's default TUI renders into the main screen
with terminal-owned scrollback, so there is nothing to strip. The fullscreen mode **shipped in
0.84.0 and is runtime-switchable via `/settings`**, so Codeman cannot assume a pi session stays
main-screen for its lifetime; staying out of the strip list is exactly what makes that safe (the alt
screen is load-bearing when the user flips to fullscreen, as it is for `opencode`). Putting pi IN
the strip list would corrupt fullscreen sessions. Three mirrors must stay consistent (all unchanged
for pi, i.e. pi appears in none of them): the replay-side strip in `session-routes.ts:2275`, the
live-stream twin in `session.ts`, and the frontend `_sessionUsesServerMouseStrip()` in
`terminal-ui.js` (usages `:3432`, `:3697`).
### 2.3 tmux required, no direct-PTY fallback, no per-mode configurator
Same rule as the other external CLIs: `pi` mode throws if tmux is unavailable. Add a fourth block to
the guard chain at `session.ts:1751-1768` (antigravity's is `:1765-1768`).
**No `_configurePi()` is needed.** Opencode/codex/gemini each have a tmux-`setenv` configurator
(`tmux-manager.ts:1709-1727`), but antigravity has none: it relies entirely on the generic
`applyEnvOverrides()` (`tmux-manager.ts:1643`, `VALID_KEY = /^[A-Z_][A-Z0-9_]*$/`), which runs for
every mode in both create (`:1880`) and respawn (`:2107`) and injects via socket-scoped
`tmux setenv`, never the spawn command line. Pi follows the antigravity precedent: `PI_*` overrides
flow through `applyEnvOverrides()` and nothing else.
Pi joins the truecolor branches: `buildEnvExports()` (`tmux-manager.ts:1604-1609`,
`export COLORTERM=truecolor` + `unset NO_COLOR` for codex/gemini/antigravity) and the attach-env
condition at `session.ts:1400-1402` (`buildMuxAttachEnv(...)`, whose comment says it must mirror
`buildEnvExports`). Add `|| mode === 'pi'` to both, or the tmux session and the attach client
disagree about color depth.
### 2.4 Env prefix: `PI_` only
Add `'PI_'` to `ALLOWED_ENV_PREFIXES` (`schemas.ts:125`) and to the prose error message at `:163`
(two edits: the message hardcodes the list, and since 1.12+ it also names the exact-key allowlist,
currently `...ANTIGRAVITY_* keys and CLAUDE_CONFIG_DIR are allowed.`; there is now a separate
`ALLOWED_ENV_KEYS` exact-key set alongside the prefix list, which pi does not need to touch). That
covers every documented variable pi reads: `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`,
`PI_PACKAGE_DIR`, `PI_OFFLINE`, `PI_SKIP_VERSION_CHECK`, `PI_TELEMETRY`, `PI_CACHE_RETENTION`,
`PI_SHARE_VIEWER_URL`, `PI_HARDWARE_CURSOR`, `PI_EXPERIMENTAL` (whose meaning 0.84.0 extended to
strict JSON-schema tool sampling). (Pi also *sets* `PI_CODING_AGENT=true` and `AI_AGENT=pi` in child
processes; those are output markers, not inputs, and need nothing from us.)
**Deliberately not added:** `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `GEMINI_API_KEY`, `XAI_API_KEY`,
`GROQ_API_KEY`, `MISTRAL_API_KEY` and the other ~28 provider keys. `ALLOWED_ENV_PREFIXES` is a
single global list applied by one Zod refine with no mode context (`safeEnvOverridesSchema`,
`schemas.ts:153-165`), so allowlisting bare provider keys for pi would widen the allowlist for
**every** mode at once, violating the multi-CLI prefix discipline in CLAUDE.md. Users authenticate
pi through `/login` (stored in `~/.pi/agent/auth.json`, auto-refreshed) or by exporting the key in
the Codeman server process's own environment.
Making the allowlist mode-aware is the clean fix, listed as a follow-up in §9. Do not smuggle it
into this change.
### 2.5 Docker credential policy: seed files, not the whole dir
`CRED_STORES` (`docker-hosts.ts:597-605`; file unchanged since the 2026-08-06 verification) gets a
`.pi/agent` entry. Nested `rel` paths already work (`.config/gcloud` maps to seed name
`.config-gcloud` via the `replace(/\//g, '-')` at `:620`). Unlike antigravity, which needed **no**
entry (`agy` nests all state under `~/.gemini/antigravity-cli/`, already covered by the `.gemini`
policy, per the comment at `:599-602`), pi has its own top-level dir and needs its own entry. Use
`seedFiles`, **not** `seedWhole`:
```ts
{ rel: '.pi/agent', seedFiles: ['auth.json', 'settings.json', 'trust.json', 'models.json', 'models-store.json'] },
```
Rationale: `~/.pi/agent` also contains `sessions/`, `extensions/`, `skills/` and the installed
package trees (`npm/`, `git/`), which on an active host is easily gigabytes; `seedWhole` would
`cp -a` all of it into every container start. The five seeded files are what pi needs to
authenticate and behave consistently: `models.json` is in the list because it holds user-defined
custom providers, and omitting it would silently strip those inside containers. Seeding (RO mount
then copy) also means the in-container pi never writes refreshed OAuth tokens back to the host,
which is the whole point of the seeding policy, and bind mounts stay excluded from `docker commit`
so exports remain secret-free.
Trade-off to accept and document: in-container pi sessions are not visible host-side, so `pi -c`
inside a Docker case only sees that container's own history. Codex shares `sessions/` RW precisely
because Codeman reads it host-side for the response viewer; there is no such reader for pi yet
(the response-viewer follow-up in §9 would justify flipping this).
### 2.6 The `pi` binary name is generic
Unlike `agy`/`codex`/`gemini`, `pi` is a short, common name (Raspberry Pi tooling, personal scripts,
`$PATH` accidents). The resolver must not blindly trust a hit. None of the existing external-CLI
resolvers execute their binary (only `claude-cli-resolver.ts` does, via the cached
`getClaudeCliVersion()`, skipped under vitest), so the sanity check is new ground: model it on
`getClaudeCliVersion()`. Run `pi --version` once via `execFileSync`, cache the result module-level,
skip under `VITEST`, and require output matching `/^\d+\.\d+\.\d+/`; on mismatch treat the binary as
unavailable and log the rejected path. Surface `{ available, path, version }` from
`GET /api/pi/status` so a misresolution is diagnosable from the UI (additive relative to the sibling
endpoints' `{ available, path }`). The `dependency-registry` entry carries `versionArg: '--version'`
for `codeman doctor`.
### 2.7 tmux extended keys (a real pi-specific footgun)
Pi documents (`docs/tmux.md`, verified verbatim) that without
```tmux
set -g extended-keys on
set -g extended-keys-format csi-u
```
tmux collapses `Shift+Enter` and `Ctrl+Enter` into a plain `\r` (and `Alt+Enter` into `\x1b\r`), and
pi's editor uses those for newline vs submit. `extended-keys-format` requires tmux 3.5+; tmux
3.2-3.4 works with `extended-keys on` alone (pi then falls back to xterm `modifyOtherKeys`).
Codeman's own browser input path sends `\r` for submit, so basic use works unconfigured, but
newline-in-editor is degraded both for a user typing in an attached terminal (`sc`) and potentially
for the browser Shift+Enter path.
Upstream recommends `~/.tmux.conf` and notes the setting may need a full `tmux kill-server` restart
to take effect. **Codeman must NEVER run `kill-server` on its socket** (it would kill every live
session, including `w1`/`w2`/`w3`). Action: attempt to set both options **server-scoped on
Codeman's own socket only** (`tmux -L codeman set -s ...`, never `-g` on the user's default socket)
at the point the tmux server is first started, verify with `tmux -L codeman show-options -s` and an
empirical Shift+Enter test which scope actually takes for the installed tmux version, and fall back
to a documented manual step in `docs/pi-integration.md` (a `~/.tmux.conf` snippet plus the
kill-server caveat) if it cannot be applied safely to an already-running server. Upstream does not
discuss socket- or server-scoped configuration at all, so this verification is original work, not a
doc lookup.
**RESULT (measured, tmux 3.4 + pi 0.84.1):** `tmux -L <socket> set -s extended-keys on` takes effect
on an **already-running** server with **no `kill-server`** — pi's own startup warning
(`Warning: tmux extended-keys is off…`, a convenient in-band probe) disappears for the next session
started afterwards. `extended-keys-format` does **not exist on tmux 3.4** and errors with
`invalid option: extended-keys-format`, so the two options must be issued independently rather than
chained. Decision: Codeman does **not** set this itself — it is a server-wide tmux option affecting
every session of every backend, so silently changing key encoding is not Codeman's call. It is
documented as a user step in `docs/pi-integration.md` instead, carrying the measured facts.
### 2.8 The completeness trap: which mode tables fail loud vs silent
Adding `'pi'` to the `SessionMode` union makes some omissions compile errors and leaves others
silent. The plan calls this out so review can focus on the silent ones.
**Loud (typecheck fails until edited):** `getModeLabel()` (`session.ts:168-183`, exhaustive switch
with no default), `defaultDockerCommandForMode` and `defaultRemoteCommandForMode` (both typed
`Record<...CommandMode, string>`), **but only after** `RemoteCommandMode` (`types/session.ts:48-51`)
and `DockerCommandMode` (`:157-161`) are widened: both are `Extract<SessionMode, '...'>` with every
member spelled out, so forgetting the `Extract` lists keeps `tsc` green while docker/remote pi cases
silently fall back to `exec bash -l` via the `|| commands.shell` on the lookup. Edit union + both
`Extract` lists + both `Record` literals together.
**Silent (compiles clean, mode just doesn't work):**
- `appendResumeFlag()` (`tmux-manager.ts:1030-1042`) has a `default:` arm; a missing `case 'pi'`
silently drops docker resume.
- `buildSpawnCommand()` (`:770-825`) and `buildPathExport()` (`:1680-1707`) are if-chains with
fallthrough returns; a missing branch spawns pi as a login shell / with no PATH augmentation.
- `isExternalCliMode()` / `isAltScreenStripMode()` are boolean chains.
- The `runMode` accessor's **setter whitelist** (`session-ui.js:2949-2960`) coerces any unknown mode
to `'claude'`. Omitting `pi` there makes the mode **unselectable while every other edit appears to
work**: this is the single most deceptive omission in the frontend.
- `window.__codemanCliAvailable` (injected by `renderIndexHtml`, `server.ts:1375-1407`): the client
treats a **missing key as available** (`isCliAvailable` in settings-ui.js), so forgetting the
injection un-gates pi on boxes without the CLI instead of hiding it.
### 2.9 The Daylight skin cascade eats per-mode run-button colors
A finding that changes the CSS work (verified empirically with computed styles on the live
instance, re-confirmed at f39beb3): `styles.css:13681` opens a nested skin block,
`html:not([data-skin="og"]) { ... }`, and the **default skin is `daylight-blue`, not `og`**, so the
block is live for every default-skin user. Inside it, `.btn-toolbar.btn-run` is re-declared
generically and per-mode only for claude/opencode/codex (codex at `:13787`). CSS nesting adds the
wrapper's specificity (the nested rules resolve to (0,3,1) vs (0,3,0) for
`.btn-toolbar.btn-run.mode-X`), so **gemini's and antigravity's toolbar gradients are dead on the
default skin**: both render the generic claude gradient today, still unfixed as of f39beb3. The
base-sheet rules (gemini/antigravity at `:4406`/`:4420`) only ever render on the `og` skin. Since
1.12+ styles.css itself documents this trap in comments (`:9214`, `:11091`), which confirms the
mechanism.
Consequences for pi:
- The toolbar gradient needs **two** rules: one in the base sheet (`:4420` area, for `og`), and one
**inside** the `13681` block next to codex's (`:13787` area), using the block's own idiom
(or the color is invisible to the average user).
- `mobile.css` phone-toolbar colors need `!important` on `background`/`border-color`/`color`,
exactly as the CLAUDE.md gotcha prescribes. Antigravity's phone block (`mobile.css:895-910`,
inside the `@media (max-width: 430px)` opened at `:338`) has no `!important` and is dead on the
default skin; do not copy that mistake.
- Three surfaces work from base rules alone (verified): run-mode **dots** (list at `:4506-4516`;
the skin block overrides only claude/opencode/codex/shell dots, so a base-sheet
`.run-mode-dot.pi` renders as authored), **tab badges**, and the **welcome button** (the skin
block overrides only claude/opencode/tunnel welcome buttons).
- Optional, separate cleanup (not this change): gemini/antigravity could get the same in-block
treatment to resurrect their colors.
### 2.10 Local-echo policy: pi lands on the buffer overlay by default
New since the first draft of this plan: the codex predictive-echo work (1.13+) introduced a
per-session echo policy in `_updateLocalEchoState()` (terminal-ui.js, `_localEchoPolicy` set at
`:2837`): `codex → 'predict'` (write-through predictive echo), `shell → 'off'`, **everything else
→ 'buffer'** (the `LocalEchoOverlay` that buffers typed text until Enter). Pi therefore gets the
buffer overlay on touch devices with zero edits, via the fallthrough.
That default is a real open question, not a freebie: the codex history (issues #218/#219/#220/#222)
shows that a composer which re-renders per keystroke (live-filtering slash picker, server-side
cursor movement, wrap-as-you-type) is starved by buffer-until-Enter, and pi's editor is exactly
such a composer. Decision for v1: ship with the default `'buffer'` policy but make phone-profile
typing an explicit E2E gate (§7 step 4); if pi's editor mis-renders under the overlay, the cheap
fallback is forcing `'off'` for pi (one branch in `_updateLocalEchoState`), and teaching the
predict path pi's composer row is a follow-up, not a v1 requirement.
`test/local-echo-codex-gating.test.ts` pins the per-mode policy via
`it.each(['claude', 'gemini', 'opencode'])` lists (`:193`, `:376`); add `'pi'` to those lists once
the buffer decision is confirmed (or pin the `'off'` branch if that is the outcome).
**RESULT (measured, pi 0.84.1, iPhone 14 Pro profile + a PTY-level A/B):** the buffer policy
**holds**; codex's failure mode does **not** reproduce. Pi's slash picker re-filters on the **whole
composer content**, not on per-keystroke deltas: a one-shot literal write of `/set` (what the overlay
flush does) filters the picker to `settings` **identically** to sending `/ s e t` as five separate
keystrokes, and the delayed `\r` then selects it and opens the settings menu. Prose prompts buffer
correctly (`pendingText` right, nothing on the PTY before Enter), flush on Enter, and are accepted as
a single prompt. `'pi'` was added to both `it.each` lists. The `'off'` fallback stays documented but
unused.
---
## 3. Config surface: `PiConfig` to CLI flags
```ts
/** Pi CLI session configuration */
export interface PiConfig {
/** Model pattern or ID. Supports `provider/id` and a `:<thinking>` suffix (e.g. `sonnet:high`). Passed via --model. */
model?: string;
/** Provider name (anthropic, openai, google, ...). Passed via --provider. */
provider?: string;
/** Reasoning level. Passed via --thinking. */
thinking?: 'off' | 'minimal' | 'low' | 'medium' | 'high' | 'xhigh' | 'max';
/** Continue the most recent session (-c). Per-cwd scoping is strongly implied upstream but not documented; treat as probable. */
continueSession?: boolean;
/** Resume a specific session by ID or partial UUID (--session). Codeman deliberately accepts ids only, never paths. */
resumeSessionId?: string;
/**
* Tri-state project trust (repo-local `.pi/` settings/extensions/skills, plus installing
* missing project packages):
* true -> --approve (trust for this run; loads and EXECUTES repository TypeScript)
* false -> --no-approve (force-deny; the trust prompt never appears)
* absent -> pi's own defaultProjectTrust (ask).
* Multi-user: MATERIALIZED to false for non-granted owners (§5.2).
*/
approveProjectTrust?: boolean;
}
```
Flag mapping in `buildPiCommand()` (new, `tmux-manager.ts`, directly after `buildAntigravityCommand`
at `:718-736`; every builder there regex-allowlists each user value and silently drops failures
because the result lands in a `bash -c "..."` string):
| Field | Flag | Validation |
| --------------------- | ------------------------------- | --------------------------------------------------------------------------------- |
| `approveProjectTrust` | `--approve` / `--no-approve` / nothing | tri-state boolean, clamped (§5.2) |
| `model` | `--model <v>` | `/^[a-zA-Z0-9._\-/:]+$/` (`:` for `sonnet:high`, `/` for `openai/gpt-4o`) |
| `provider` | `--provider <v>` | `/^[a-z0-9-]+$/` |
| `thinking` | `--thinking <v>` | runtime allowlist of the 7 enum values (defense in depth beyond Zod) |
| `resumeSessionId` | `--session <v>` | `/^[a-zA-Z0-9._-]+$/` (same shape as `RESUME_ID_SAFE`, `:1021`; excludes paths on purpose) |
| `continueSession` | `-c` | boolean; **skipped when a valid `resumeSessionId` is present** (the two conflict) |
**Not** wired in v1, with reasons:
- `--api-key <key>`: ⚠️ **never wire this.** It puts a provider secret on the spawn command line,
which is exactly what the socket-scoped `tmux setenv` discipline exists to prevent (visible in
`ps`, tmux server state, and logs). Listed here so nobody "helpfully" adds it later.
- `--tui-mode` (released in 0.84.0): never passed by Codeman. The main-screen default is the
friendly case for the browser terminal, and fullscreen remains the user's own runtime choice via
`/settings` (§2.2 is designed for that). `--use-theme` (still unreleased) likewise.
- `--name <name>` (`-n`): nice for `/resume` readability, but names contain spaces and would be the
first user-controlled value needing real shell quoting in `buildSpawnCommand`. Defer.
- `--no-session`: ephemeral mode fights respawn/resume. Defer.
- `-p`/`--print`, `--mode json`, `--mode rpc`: non-interactive transports, a different product shape
(§9). Note upstream already shipped a breaking change to JSON-mode `message_update` framing, so
any future consumer must assemble deltas.
- `--tools` / `--exclude-tools` / `--no-tools` / `--no-builtin-tools` (`-t`/`-xt`/`-nt`/`-nbt`): a
genuinely useful "read-only session" affordance (0.84.0 also added a `defaultTools` setting), but
it needs UI design. Follow-up.
- `-r`/`--resume` (interactive picker), `--fork`, `-e`/`--extension`, `--skill`, `--system-prompt`,
`--append-system-prompt`, `--export`, `--models`, `--list-models`: not session-manager concerns in
v1. (`-e` matters later: §9's extension follow-up notes CLI extensions load before trust
resolution.)
---
## 4. Implementation phases
### Phase 1: Backend core
| File | Change |
| ----------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| `src/utils/pi-cli-resolver.ts` | **New**, mirror `antigravity-cli-resolver.ts` (65 lines: search-dir list, module-level cache with `''` negative sentinel, `which pi` first). Search dirs: `~/.local/bin`, `/usr/local/bin`, `~/.bun/bin`, `~/.npm-global/bin`, `~/bin`. Add the `pi --version` sanity probe from §2.6 (execFileSync, cached, vitest-skipped). Export `resolvePiDir()`, `isPiAvailable()`, `getPiCliVersion()` |
| `src/utils/index.ts` | Re-export the three (resolver block `:30-36`) |
| `src/types/session.ts` | `SessionMode` union `:46`; **both `Extract` lists**: `RemoteCommandMode` `:48-51`, `DockerCommandMode` `:157-161` (§2.8); new `PiConfig` after `AntigravityConfig` (`:325-333`); `SessionState.piConfig` after `:486`; `@fileoverview` mode list `:11` + config list `:17` |
| `src/mux-interface.ts` | `piConfig?: PiConfig` on `CreateSessionOptions` (config block ends `:78`) and `RespawnPaneOptions` (ends `:109`) |
| `src/session.ts` | `isExternalCliMode()` `:164-167` (+pi); `getModeLabel()` `:168-183` (+`'Pi'`); `_piConfig` field decl `:466-470`; ctor option `:556-563` + apply `:652-654`; `toState()` `:1227-1230`; `_buildRespawnPaneOptions()` `:1466-1469` (single source of truth shared by `startInteractive` and `reattachRemote`); `startInteractive()` createSessionOptions `:1680-1683`; COLORTERM attach-env condition `:1400-1402` (+pi); requires-tmux guard chain `:1751-1768` (new block: "Pi sessions require tmux for env override injection via setenv") |
| `src/tmux-manager.ts` | `buildPiCommand()` after `:736` per §3; `buildSpawnCommand()` signature `:770-779` + dispatch branch after `:822-825`; `appendResumeFlag()` `:1030-1042` (`case 'pi': return \`${modeCommand} --session ${resumeId}\`;`); `buildEnvExports()` truecolor branches `:1604-1609` (+pi); `buildPathExport()` `:1680-1707` (+pi branch calling `resolvePiDir()`); missing-CLI error chain in `createSession` `:1788-1806` (+pi, install hint `npm install -g --ignore-scripts @earendil-works/pi-coding-agent`; note `respawnPane` deliberately has no such check); `piConfig` threading at the four sites `:1748`, `:1817`, `:2041`, `:2080`. **No `_configurePi`** (§2.3) |
| `src/config/dependency-registry.ts` | New entry after antigravity's (`:101-108`; file unchanged since 2026-08-06): `{ id: 'pi', label: 'Pi CLI', category: 'core', required: false, usedBy: ['Pi sessions'], resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['pi'], versionArg: '--version' } }] }` |
| `src/docker-hosts.ts` | `defaultDockerCommandForMode` `:138-149`: `pi: 'exec pi'`. `CRED_STORES` `:597-605`: the `.pi/agent` seedFiles entry per §2.5 (nested `rel` already handled at `:613-645`). File unchanged since 2026-08-06 |
| `src/remote-hosts.ts` | `defaultRemoteCommandForMode` `:92-118`: `pi: remoteLoginShellCommand('pi')` (`remoteLoginShellCommand` at `:88-90`). Login-shell routing is mandatory (the #209/e803186 lesson: ssh remote-command exec sees only sshd's minimal PATH, and npm's global bin is usually only on PATH via rc files) |
### Phase 2: Web layer
| File | Change |
| ---------------------------------- | ----------------------------------------------------------------------------------------------- |
| `src/web/schemas.ts` | `'PI_'` in `ALLOWED_ENV_PREFIXES` `:125` **and** the prose error message `:163` (which now also names `CLAUDE_CONFIG_DIR`; the `ALLOWED_ENV_KEYS` exact-key set needs no change); new `PiConfigSchema` after `AntigravityConfigSchema` (`:256-271`), mirroring §3's regexes, `.optional()`, not `.strict()`; `piConfig` on `CreateSessionSchema` (`:299` area) and `QuickStartSchema` (`:712` area); `'pi'` in all three mode enums (`:285`, `:708`, cron `agentType` `:1214`; they are byte-identical and there is no fourth); `pi` key in `RemoteCommandOverridesSchema` `:426-436` (it is `.strict()`, so an unknown key is a hard error today; one edit covers both remote `:501` and docker `:577` reuse) |
| `src/web/routes/session-routes.ts` | Thread `piConfig` through create (`POST /api/sessions`): disk-strip exclusion chain `:705-712`, availability gate `:782-790` (+`isPiAvailable` with install-hint error), model resolution `:825-838` (`mode === 'pi' ? body.piConfig?.model : ...`), clamp call `:845`, Session ctor `:860` (`piConfig: mode === 'pi' ? gatedPiConfig : undefined`). Quick-start (`POST /api/quick-start`, handler `:2559`): remote-case config rejection `:2614-2621` and docker-case `:2645-2652` (+`piConfig`: per-CLI config does not cross ssh or the bind mount), hooks-scaffold exclusions `:2801`/`:2809`, availability gate `:2744-2752` (local-case branch only), env-strip chains `:2833`/`:2863`, model resolution `:2885`, clamp `:2897`, ctor `:2913`. **Extend `clampExternalCliBypassForOwner()`** (`:305-336`, doc comment above): fifth param + return field; pi joins the **materialize** branch per §5.2. Alt-screen replay-strip at `:2275` unchanged (pi not in it, §2.2) |
| `src/web/routes/system-routes.ts` | `GET /api/pi/status` after the antigravity handler (`:418-426`; file unchanged since 2026-08-06), same shape plus `version` (§2.6); update the "CLI Integrations" prose comment `:377` |
| `src/web/server.ts` | Restore path: `piConfig: muxSession.mode === 'pi' ? savedState?.piConfig : undefined` after `:2636`. **`renderIndexHtml` CLI-availability injection `:1375-1407`**: add `isPiAvailable` to the dynamic-import tuple (`:1382`) and a `pi` key to the injected object (`:1399`). Per §2.8 a missing key reads as *available*, so this is a correctness edit, not polish |
### Phase 3: Frontend
The antigravity touchpoints are the template. Since the first draft, the settings-surface overhaul
moved most anchors and added one **new touchpoint** (the clone-repo Brain picker below).
`constants.js`, `api-client.js`, `ralph-wizard.js`, `cron-ui.js`, `webview-tabs.js` and `sw.js`
still need **no** changes (re-verified zero mode coupling at f39beb3; cron-ui reads the `<select>`
generically and special-cases only `shell`).
| File | Change |
| ------------------- | ------------------------------------------------------------------------------------------------------ |
| `index.html` | Welcome button `welcomePiBtn` after Gemini's (antigravity's is `:347`; there is deliberately no codex welcome button), `display:none` default, `onclick="app.setRunMode('pi'); app.runPi()"`, text `Run Pi`; run-mode-option row with `.run-mode-dot.pi` after antigravity's (`:526-528`), before the `.run-mode-sep` `:529`; cron `<option value="pi">Pi</option>` after `:803`; **NEW: the clone-repo "Brain" picker** (`cloneCaseBrain`, `:2476-2486`): add `<option value="pi" data-cli="pi">Pi</option>` after the antigravity option `:2483` (gating is automatic: session-ui.js `:2107-2115` hides options whose `data-cli` fails `isCliAvailable`, and `:2250` reads the value at clone time); docker image hint `:2624` (`claude/codex/gemini/opencode/agy` + pi). No per-CLI remote-command override field needed (only codex has one, `:2559`) |
| `session-ui.js` | `@fileoverview` mode list `:2`; `run()` dispatch branch after `:400-402`; `_refreshRunModeAvailability` list `:468` (+`'pi'` as a quoted literal, the static test in §6 demands it); short-label ternary `:565` (+`'Run PI'`); **the `runMode` setter whitelist `:2949-2960`** (§2.8, the deceptive one); new `runPi()` modeled on `runAntigravity()` `:1170-1219`: same remote/docker skip, same `_beginSessionLaunchStatus` frame, probes `/api/pi/status` reading `(await res.json()).data.available` (envelope!), **sends no `piConfig` at all** (no bypass exists and trust defaults are pi's own; envOverrides still sent for local cases), install-hint error text matching Phase 1's; `isAltMode` `:1233` and `isExternalCli` `:1263` four-way comparisons (+pi) |
| `settings-ui.js` | `applyWelcomeCliVisibility()` `:1176-1191`: add `['welcomePiBtn', 'pi']` |
| `app.js` | Response-viewer agent label `:1998-2009` (+pi -> `'Pi'`); tab badge ternary `:3884` (`<span class="tab-mode pi" aria-hidden="true">pi</span>`; claude stays badge-less); kill-title ternary `:5046-5057` (`Kill Tmux & Pi`) |
| `panels-ui.js` | Command-palette `labels` map `:430` (+`pi: 'Pi'`; the `\|\| mode` fallback means this is cosmetic, not load-bearing) |
| `mobile-overview.js`| `MOBILE_OVERVIEW_RUN_MODES` `:55-62`: `{ mode: 'pi', label: 'Pi', short: 'Pi' }` after antigravity `:60`, before the shell entry. Nothing else: the Run-button badge (`:499`) and menu builder (`:554-556`) consume the list generically, and the buttons carry `btn-toolbar btn-run mode-pi`, which is exactly why they inherit the §2.9 cascade problem and its fix |
| `terminal-ui.js` | Badge-row comment `:1750` only (the badge itself is a raw `s.mode` passthrough, no list to extend). `_sessionUsesServerMouseStrip` unchanged (§2.2). `_updateLocalEchoState` unchanged for v1 (§2.10: pi lands on `'buffer'` via the fallthrough; only touch it if E2E forces the `'off'` fallback) |
| `i18n.js` | `'Run Pi': '运行 Pi'` in the zh-CN table (`:102-107`, matches the welcome-button text; short labels like `Run PI` are deliberately untranslated, as are the other modes') |
| `styles.css` | Tab badge `.session-tab .tab-mode.pi` after `:2157` (`background: rgba(244,114,182,0.2); color: #f472b6;`); add `.session-tab .tab-mode.pi` to the light-skin ink list `:325-336` (gemini + antigravity are its precedent, `:332`); welcome `.welcome-btn-pi` + `:hover` after antigravity's `:3366` block, rose family (e.g. base `linear-gradient(135deg, #33121f 0%, #9d174d 55%, #be185d 100%)`, border `rgba(244,114,182,0.4)`, text `#fce7f3`); toolbar gradient pair `.btn-toolbar.btn-run.mode-pi, .btn-toolbar.btn-run-gear.mode-pi` + `:hover` after `:4420`'s antigravity block; `.run-mode-dot.pi { background: #f472b6; }` in the dot list `:4506-4516`; **and the §2.9 rule inside the Daylight block** next to codex's `:13787` (e.g. `background: linear-gradient(135deg, #be185d, #f472b6); border-color: #be185d; color: #fff1f7;`). The dot needs no skin-block entry (the block overrides only claude/opencode/codex/shell dots; gemini/antigravity dots already fall through correctly) |
| `mobile.css` | Phone toolbar block after `:910` inside the `@media (max-width: 430px)` opened at `:338`: `mode-pi` base + `:active`, **with `!important` on background/border-color/color** (§2.9; antigravity's block `:895-910` omits it and is dead); light-skin override entry after `:2985` with the same four-skin `html:is(...)` prefix as its siblings |
### Phase 4: Docker image and installer
Both files are unchanged since the 2026-08-06 verification; all anchors stand.
- `docker/agent.Dockerfile`: a **separate** `RUN` step after the antigravity block (`:38-45`), not a
fifth line in the shared npm block (`:31-36`), because pi documents `--ignore-scripts` and that
flag must not silently change how the other four install:
```dockerfile
# Pi (pi.dev). Upstream documents --ignore-scripts (pi needs no lifecycle scripts);
# kept out of the shared npm block above so the flag cannot affect the other CLIs.
RUN npm install -g --ignore-scripts @earendil-works/pi-coding-agent \
&& npm cache clean --force \
&& pi --version
```
Implementation checklist item: the gid-0 pre-created dirs at `:64-68` include `.claude/projects`
and `.codex/sessions`; verify whether the cred-seed copy into `~/.pi/agent` creates its target
dir in a fresh container or whether `.pi/agent` must join that `mkdir` line. Rebuild with
`node scripts/build-agent-image.mjs --no-cache` (the script itself needs no change; nothing in it
is CLI-specific). The cached npm layer has silently frozen a CLI at a broken version before; see
`docs/docker-cases.md`.
- `install.sh` (six edit sites, all verified): `PI_SEARCH_PATHS` block after `:125` (mirror the
resolver's dirs); `check_pi` / `get_pi_path` pair inserted at `:531` (antigravity's pair spans
`:504-530`); the satisfying-AI-CLI chain `:2032-2063` (`has_pi` local at `:2037` area, detect
block after `:2059`, widen the five-way test at `:2061` and the warn text at `:2063`); the menu
option-4 text `:2070`; the skip-path hints `:2115-2116` (add
`npm install -g --ignore-scripts @earendil-works/pi-coding-agent (Pi)`); the final no-CLI
reminder `:2416-2423` (add `check_pi` to the condition and a pi line to the echo block).
Detection plus a hint only; do **not** add an auto-install path in this change.
### Phase 5: Docs
- `docs/pi-integration.md` (**new**, user-facing): install (both installers uninstall via npm), auth
(`/login` OAuth for six providers vs API keys; `pi auth check` for preflight; Claude Pro/Max
third-party harness usage bills as Anthropic "extra usage" per token, not plan limits; OpenRouter
login supports pasting the redirect URL, which matters over remote SSH), what Codeman wires up
and deliberately does not (§3, incl. never passing `--tui-mode`), the tmux extended-keys note
from §2.7 with the manual `~/.tmux.conf` fallback, Docker/remote behaviour (in-container sessions
invisible host-side), the trust model in §1 words, known gaps.
- `CLAUDE.md`: tech-stack line (six CLIs + `SessionMode` union), the env-prefix gotcha bullet, the
multi-CLI prefix-discipline bullet, the "External CLI modes" key-pattern paragraph (note it now
also carries the codex predictive-echo block; pi's echo-policy decision from §2.10 belongs in the
same paragraph), the `src/utils/` resolver list.
- `docs/architecture-invariants.md`: the external-CLI-modes section. ⚠️ Its anchor was already
renamed once to `#external-cli-modes-opencode-codex-gemini-antigravity` while CLAUDE.md's link
text still shows the old name; when renaming again for pi, update every inbound link (CLAUDE.md
and this file).
- `docs/docker-cases.md` (cred-seeding table + supported modes + image contents),
`docs/remote-sessions.md` (`RemoteCommandMode`), `docs/cron-guide.md` + `docs/cron-discovery.md`
(`agentType` enum; note the readiness caveat from §6's cron paragraph),
`docs/security-architecture.md` (env prefix allowlist row).
- `README.md` + `README.zh-CN.md`: six CLIs.
- `package.json` keywords: `pi`.
- Update the issue #206 thread when it ships.
---
## 5. Security checklist
1. **Command injection.** Every `PiConfig` value is regex-validated in `buildPiCommand()` before
entering the `bash -c "..."` string; anything failing validation is dropped, not escaped
(matching the four existing builders). No user string reaches the spawn line unvalidated. Pinned
by a "rejects unsafe values" test per field.
2. **Multi-user clamp, materialize branch.** `approveProjectTrust` is the privilege-shaped field: it
makes pi execute repository-supplied TypeScript and install project packages.
`clampExternalCliBypassForOwner()` (`session-routes.ts:305-336`) has two branches, and pi belongs
in the **gemini-style materialize branch**, not the codex/antigravity only-if-sent branch:
pi's absent-config default is an *interactive trust prompt the session user can answer
themselves in the terminal*, so merely omitting `--approve` is not a clamp. For a non-granted
owner, materialize `{ ...(piConfig ?? {}), approveProjectTrust: false }` so `buildPiCommand`
always emits `--no-approve` and the prompt never appears. Both call sites (`:845`, `:2897`)
widen. This helper still has **zero test coverage** (re-confirmed at f39beb3); §6 adds the first
tests.
3. **Secrets stay off the command line.** `PI_*` overrides flow through `applyEnvOverrides()` /
socket-scoped `tmux setenv`, never inlined into the spawn string. No `-e` at container create
time. And `--api-key` is never wired (§3): it would put a provider secret into `ps`/tmux state.
4. **Env allowlist not widened.** Only the `PI_` prefix is added; the provider keys stay out (§2.4)
and `ALLOWED_ENV_KEYS` is untouched. Pinned by a test that `PI_OFFLINE` passes and
`ANTHROPIC_API_KEY` still fails validation.
5. **Docker seeding, not sharing.** Per §2.5: RO mount then copy, so refreshed OAuth tokens never
write back to the host; bind mounts stay excluded from `docker commit` so exports remain
secret-free.
6. **Remote SSH.** `pi` mode goes through `defaultRemoteCommandForMode` and therefore
`buildSshConnectionArgs()`. No hand-built ssh line anywhere.
7. **No sandbox claims.** Pi documents that it has no sandbox and no permission prompts, and that
extensions run with the user's full permissions. Codeman docs must say plainly that a pi session
can read, write and execute anything the Codeman user can, and point at Docker cases as the
isolation story. Do not imply the trust prompt is a safety boundary (upstream itself says it is
not). Worth one doc sentence: `pi auth print-api-key` / `print-bearer-token` (0.83.0) and
`pi auth check` (0.84.1) mean a pi session can print its own provider credentials by design;
isolation, again, is Docker.
8. **Loud-vs-silent audit.** Before review, walk §2.8's silent list and confirm each site has its
pi branch; the loud ones the compiler already caught.
---
## 6. Test plan
- `test/pi-mode.test.ts` (**new**, modeled on `test/antigravity-mode.test.ts`, 125 lines, no port;
file unchanged since 2026-08-06 so its structure remains the template):
`CreateSessionSchema`/`QuickStartSchema` accept a pi config; unsafe `model`/`provider`/
`resumeSessionId` values are rejected (`'pi; rm -rf /'` shapes); `buildSpawnCommand({ mode: 'pi', ... })`
emits expected flags, drops invalid ones, emits `--no-approve` for `approveProjectTrust: false`
and `--approve` for `true`, and skips `-c` when a `resumeSessionId` is present;
`defaultDockerCommandForMode('pi') === 'exec pi'` and
`defaultRemoteCommandForMode('pi') === 'exec "${SHELL:-/bin/sh}" -i -l -c \'pi\''`;
`isExternalCliMode('pi') === true`, `isAltScreenStripMode('pi') === false`; the env pair
(`PI_OFFLINE` accepted, `ANTHROPIC_API_KEY` rejected), mirroring antigravity-mode `:49-63`.
- **First-ever coverage for `clampExternalCliBypassForOwner`** (still nothing in `test/` touches
it): cover pi's materialize branch (absent config still yields `approveProjectTrust: false` for a
non-granted owner; a sent `true` is forced to `false`; granted owner passes through) and, while
there, pin the three existing modes' behavior. Prefer exporting the helper for direct unit tests
over a heavier multi-user route fixture; either way it lives under `test/routes/`.
- `test/run-mode-ui.test.ts`: extend `loadUi()`'s stub lists (welcome-button ids, mode buttons,
`ALL_OFF`) and add pi welcome/dropdown gating cases; note the static parser test
`'gates every mode the run-mode menu actually offers'` (`:433-456`) picks up the new
`data-mode="pi"` from index.html automatically and **fails until** `_refreshRunModeAvailability`
contains a quoted `'pi'`, which is exactly the regression it exists for. Add a
`describe('Pi quick start')` modeled on the antigravity one (`:840`) driving `runPi()` against a
stubbed `/api/pi/status` + `/api/quick-start`, asserting the posted body has `mode: 'pi'` and
**no `piConfig`**, and that the envelope is unwrapped. (The short-label assertion pattern is at
`:82`, `'Run AG'`.)
- `test/render-index-html.test.ts` `:141`: the injected `window.__codemanCliAvailable` is asserted
with an exact `toEqual` and now carries **seven** keys (claude, opencode, codex, gemini,
antigravity, cloudflared, and since 1.12+ `git`), so it **must** gain the `pi` key (and the
resolver mock an `isPiAvailable`); its comment explains why: a dropped key silently un-gates
(§2.8).
- `test/routes/system-routes.test.ts`: `GET /api/pi/status` shape, modeled on the antigravity
describe (`:816-838`) + resolver mock (`:84-87`); file unchanged since 2026-08-06.
- `test/mobile-overview.test.ts`: `:375` is an exact-array `toEqual` over the run-menu modes and
**will fail until updated** to include `'pi'` (the second exact-array at `:366`,
`['claude', 'shell']`, is a gating case and stays as-is); the sibling static parser then covers
the new entry automatically. The no-hex-literals guard only scans `.mobile-overview*` rules, so
pi's `mode-pi` colors in mobile.css do not trip it.
- `test/local-echo-codex-gating.test.ts` (§2.10): once the buffer-policy decision is confirmed in
E2E, add `'pi'` to the `it.each(['claude', 'gemini', 'opencode'])` lists (`:193`, `:376`) so the
chosen policy is pinned.
- `test/skin-themes.test.ts`: will NOT trip (it enumerates skins, not modes); run it anyway since
styles.css is touched. `test/mobile-header-buttons-policy.test.ts`: trips only if a header
button is added; pi adds none (welcome button and run-menu rows are outside `header-right`).
- Cron: schema-level acceptance of `agentType: 'pi'` (the service consumes `SessionMode`
generically; `src/cron/` is unchanged since the first draft). Known, documented degradation: the
readiness poll (`cron-service.ts:515`) looks for `❯`/`tokens`, which pi never prints, so cron pi
jobs burn the ready-poll attempts and then send anyway. Acceptable for v1; note it in
`docs/cron-guide.md`.
- Sweep with `npm run test:ci`. Never bare `npm test`. No new ports needed (all new/extended suites
are portless).
---
## 7. End-to-end verification (required before COM)
Unit tests passing is not evidence the mode works (pi is not currently installed on the dev box, so
step 1 is a real step). Before shipping:
1. Install pi (`npm install -g --ignore-scripts @earendil-works/pi-coding-agent`), authenticate once
with `/login`.
2. `curl -sk https://localhost:3000/api/pi/status | jq` reports `available: true`, the right path,
and a sane `version`.
3. Create a **throwaway** case, launch a pi session from the Run dropdown, send a prompt from the
browser, confirm the reply renders and scrollback survives a tab switch. Do not touch
`w1`/`w2`/`w3`.
4. **Local-echo policy gate (§2.10):** on a phone profile, type into the pi editor through the
buffer overlay (drive with `page.keyboard.type()`, never `app.sendInput()`, and force
`app._localEchoEnabled = true`; headless Chromium reports touch as false) and confirm pi's
composer renders the flushed text correctly on Enter. If it mis-renders, flip pi to the `'off'`
branch in `_updateLocalEchoState` and pin that instead.
5. Visual pass on the **default skin** (the §2.9 finding makes this the load-bearing check, not a
formality): run-button gradient actually renders rose (not generic claude blue), dot, tab badge,
welcome button, kill-menu label; then a phone profile (toolbar `!important` colors and light-skin
overrides are the usual regressions).
6. Kill and respawn the session; confirm `piConfig` round-trips through `state.json` and the pane
comes back with the same flags. Then `/clear`-style respawn via the Respawn tab.
7. Extended keys (§2.7): in an attached terminal, verify whether Shift+Enter inserts a newline in
pi's editor with and without the socket-scoped options; record the outcome in
`docs/pi-integration.md` either way. While attached, also flip `/settings` to the fullscreen TUI
and back to confirm the no-strip decision holds (§2.2).
8. Trust model: point a throwaway case at a repo containing `.pi/extensions`, confirm the trust
prompt appears interactively and that a multi-user non-granted session instead launches with
`--no-approve` (prompt never shown, extensions not loaded).
9. **NOT RUN in this pass — an honest gap.** Docker case with `mode: 'pi'`: rebuild the agent image with `--no-cache`, confirm `pi --version`
inside the container **as the `agent` user**, confirm seeded auth works and a session starts
(this is exactly where the antigravity Docker path broke in 1.11.2: the CLI was never installed
in the image).
10. **NOT RUN in this pass — the other gap.** Remote SSH case with `mode: 'pi'`: confirm the
login-shell wrapper resolves the npm global bin.
11. Only then: changeset, `COM minor` (new capability, additive to the API surface).
**Verification actually performed** (2026-08-13, pi 0.84.1, isolated `CODEMAN_INSTANCE=pi-beta`
server on :5055 with its own tmux socket and data dir): steps 1-8 pass. Highlights:
`/api/pi/status` resolved through the **search-dir fallback** (pi installed to `~/.npm-global/bin`,
deliberately not on PATH) and reported
`{available:true, path:'/home/arkon/.npm-global/bin', version:'0.84.1'}`; the real spawn line came
out as `… COLORTERM=truecolor … && pi --approve --provider anthropic --thinking high`; `piConfig`
round-tripped through `state.json` across a **full server restart**; the trust prompt appeared for a
case containing `.pi/extensions` + `.pi/settings.json`, and `--no-approve` suppressed it
(`This project is not trusted. Project .pi resources and packages are ignored.`); on the **default
`daylight-blue` skin** the toolbar Run button computed to
`linear-gradient(135deg, rgb(190,24,93), rgb(244,114,182))` — genuinely rose and **distinct from
claude's blue**, so the §2.9 cascade trap is avoided; and flipping `/settings` to the fullscreen TUI
put the pane into the alt screen (`alternate_on=1`), **empirically confirming §2.2**: had pi been in
the strip list, Codeman would have stripped that switch and corrupted the session. Steps 9-10 need a
Docker daemon and a remote host respectively.
---
## 8. Effort estimate
Calibrated against the real antigravity history, which is the honest baseline: the feature commit
`26cbbe0` was 24 files, +638/-63, and it then took **four follow-up commits** (`e803186` login-shell
routing, `292ba2c` ownership helpers, `5d28999` CLI gating incl. tests, `0d0b772` docs/installer/UI
propagation) totaling roughly +600/-170 across ~43 file-touches to make the mode actually
first-class. Budgeting only the feature-commit shape under-scopes by ~40%. This plan folds all four
follow-up surfaces in from the start (login-shell routing in Phase 1, availability gating in Phases
2-3, installer/docs propagation in Phases 4-5), so expect the full footprint in one pass:
| Phase | Size |
| --------------------- | -------------------------------------------------------------------------- |
| 1. Backend core | ~260 lines across 9 files, one new file (resolver incl. version probe) |
| 2. Web layer | ~110 lines across 4 files (incl. the clamp widening + availability inject) |
| 3. Frontend | ~175 lines across 10 files (enumerations + CSS in two sheets + skin block + the Brain picker option) |
| 4. Docker + installer | ~45 lines, plus one `--no-cache` image rebuild |
| 5. Docs | one new doc, ~10 files touched |
| 6. Tests | one new test file, 6 extended (2 of which fail loudly until updated), plus the first clamp coverage |
---
## 9. Out of scope, tracked as follow-ups
- **A Codeman pi extension for real idle/completion events (highest value, now fully de-risked).**
Pi extensions are TypeScript modules with Node built-ins and npm deps available, so an HTTP POST
to `/api/hook-event` is trivial. The **`agent_settled`** event **shipped in 0.84.0** and is
documented for exactly this use case (fires only when pi will not continue on its own: after
auto-retries, auto-compaction and queued follow-ups; `ctx.isIdle()` is true inside the handler).
That is a genuine idle signal replacing output-silence heuristics, i.e. the same class of upgrade
hooks give Claude sessions. The bash tool exposes five env vars (`PI_SESSION_ID`,
`PI_SESSION_FILE`, `PI_PROVIDER`, `PI_MODEL`, `PI_REASONING_LEVEL`), injected per command. Bonus:
an extension can own the **`project_trust`** event (first yes/no wins, and CLI `-e` extensions
load *before* trust resolution), so Codeman could answer the trust prompt programmatically, a
cleaner mechanism than the `--approve` flag for both the single-user convenience case and the
multi-user deny case.
- **Response viewer for pi.** Sessions are JSONL v3 under
`~/.pi/agent/sessions/--<cwd-dashed>--/<timestamp>_<uuid>.jsonl` with an `id`/`parentId` tree and
typed content blocks (text, image, thinking, toolCall); the cwd-derived dir name is trivially
computable host-side. Feasible, and it would justify flipping the Docker cred policy to share
`sessions/` RW like Codex.
- **Mode-aware env allowlist.** Would let pi sessions accept provider keys without widening the
global list. Needs `ALLOWED_ENV_PREFIXES` to become a per-mode map plus mode context inside the
Zod refine.
- **`--tools` / `--exclude-tools` / `--no-tools` / `--no-builtin-tools` read-only sessions** (plus
the 0.84.0 `defaultTools` setting). Real product value, needs UI.
- **Predictive echo for pi's composer** if the §2.10 buffer decision does not hold up in practice:
teach `PredictiveEchoAddon` pi's composer row the way `isCodexComposerRow` handles codex's.
- **`--mode json` / `--mode rpc`, and upstream's experimental remote-session client APIs**
(transport-neutral `PiClient`, CBOR protocol, Unix-socket transport, `RemoteSession` controller,
still unreleased as of 0.84.1). A potential non-PTY integration path, a different architecture
from the tmux+PTY model. Note the already-shipped breaking change to `message_update` framing
(delta-only): any consumer must assemble deltas between `message_start`/`message_end`.
- **`--name` for session labels.** Blocked on shell-quoting a user string in `buildSpawnCommand`.
---
## 10. Risks
| Risk | Mitigation |
| ------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| `pi` resolves to an unrelated binary | `pi --version` + semver-shape check in the resolver (§2.6); path and version shown in `/api/pi/status` |
| Pi's TUI repaints in a way the browser terminal handles badly | Test scrollback and repaint early (step 3 of §7); pi's default is main-screen with terminal-owned scrollback, which is the friendly case |
| Fullscreen TUI mode (shipped 0.84.0, runtime-switchable) | Already designed for: pi stays OUT of the strip list, so a user flipping `/settings` to fullscreen gets opencode-like alt-screen behavior, not corruption. §7 step 7 tests the flip explicitly |
| The buffer local-echo overlay fights pi's live composer | §2.10: explicit E2E gate (§7 step 4) with the one-line `'off'` fallback; predictive echo for pi is a tracked follow-up, not a v1 blocker |
| Pi moves fast (pre-1.0; 9 releases in the 7 weeks before 0.84.1) | Keep the flag surface small; every flag validated and droppable; nothing pinned in the Dockerfile beyond the `--no-cache` rebuild cadence. Live example of the hazard: `--tui-mode` went from main-only docs to released between the two drafts of this plan |
| Docker image grows | Pi is an npm package; the layer is modest next to the ~190MB `agy` binary |
| Trust prompt blocks a session | Narrower than feared: only fires when `.pi/settings.json`, `.pi/extensions\|skills\|prompts\|themes`, `.pi/SYSTEM.md`/`APPEND_SYSTEM.md` or `.agents/skills` exists (bare `.pi/` does not). Documented; `approveProjectTrust` is the opt-in escape hatch; multi-user forces `--no-approve` (§5.2); the `project_trust` extension follow-up removes the prompt entirely |
| Interactive `/login` OAuth can't complete headlessly | Document: authenticate once interactively (or seed `auth.json`); `pi auth check` verifies credentials preflight; OpenRouter's paste-the-redirect-URL flow covers remote SSH |
| Provider auth is awkward without key prefixes in the allowlist | `/login` writes `~/.pi/agent/auth.json` once and Docker seeds it; the mode-aware allowlist follow-up removes the friction |
| Cron pi jobs mis-detect readiness | Known degradation, documented in §6; readiness falls through after the poll budget and the prompt still sends |
+235
View File
@@ -0,0 +1,235 @@
# Pi (pi.dev) sessions
Codeman can drive [Pi](https://pi.dev) (`@earendil-works/pi-coding-agent`, MIT) as a
session backend, alongside Claude Code, OpenCode, Codex, Gemini and Antigravity.
`pi` is a sixth **run mode**: its own PTY, its own tmux session, its own tab colour
(rose). It is not a location overlay like Docker or remote-SSH cases, and it is not
a web tab.
Tracking issue: [#206](https://github.com/Ark0N/Codeman/issues/206). The design
rationale behind each decision below lives in `docs/pi-integration-plan.md`.
## Install
```bash
npm install -g --ignore-scripts @earendil-works/pi-coding-agent
# or
curl -fsSL https://pi.dev/install.sh | sh
```
Both installers end up going through global npm, so either one uninstalls with
`npm uninstall -g @earendil-works/pi-coding-agent`.
Codeman finds the binary via `which pi` and then the usual global-bin locations
(`~/.local/bin`, `/usr/local/bin`, `~/.bun/bin`, `~/.npm-global/bin`, `~/bin`).
**`pi` is a short, generic name**, so unlike the other CLI resolvers Codeman does
not trust a `which` hit on its own: it runs `pi --version` once and requires
semver-shaped output. Anything else is rejected as "not installed" and the
rejected path is logged. Check what it resolved:
```bash
curl -s localhost:3000/api/pi/status | jq
# { "available": true, "path": "/home/you/.local/bin", "version": "0.84.1" }
```
That endpoint carries `version` on top of the shape the sibling `/api/*/status`
endpoints return, precisely so a misresolution is visible rather than presenting
as "the mode just doesn't work".
## Authenticate
Pi supports 15+ providers. Two ways in:
- **OAuth subscription login** — run `/login` inside a pi session. Six providers
support it: ChatGPT Plus/Pro, Claude Pro/Max, GitHub Copilot, xAI, OpenRouter
and Radius. Credentials land in `~/.pi/agent/auth.json` and pi refreshes them
itself. OpenRouter's flow accepts a pasted redirect URL, which is what makes it
workable over remote SSH.
- **API keys** — exported in the environment of the **Codeman server process**.
⚠️ **Provider API keys cannot be sent as per-session `envOverrides`.** Pi reads
about 34 provider variables (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`DEEPSEEK_API_KEY`, `HF_TOKEN`, `BASETEN_API_KEY`, …) that share no common prefix.
Codeman's env allowlist is a single global list applied to every mode at once, so
admitting bare provider keys for pi would widen the allowlist for Claude, Codex,
Gemini and everything else too. Only the **`PI_*`** prefix was added, which covers
every documented pi input: `PI_CODING_AGENT_DIR`, `PI_CODING_AGENT_SESSION_DIR`,
`PI_PACKAGE_DIR`, `PI_OFFLINE`, `PI_SKIP_VERSION_CHECK`, `PI_TELEMETRY`,
`PI_CACHE_RETENTION`, `PI_SHARE_VIEWER_URL`, `PI_HARDWARE_CURSOR`,
`PI_EXPERIMENTAL`.
`pi auth check` verifies credentials before you start a long run.
Note if you authenticate with a Claude Pro/Max subscription: third-party harness
usage bills as Anthropic "extra usage" per token rather than against plan limits.
## What Codeman wires up
`PiConfig` (per session, persisted in `state.json`, round-trips through respawn):
| Field | Flag | Notes |
| --------------------- | -------------------------------------- | ---------------------------------------------------------------- |
| `model` | `--model <v>` | Accepts `provider/id` and a `:<thinking>` suffix (`sonnet:high`) |
| `provider` | `--provider <v>` | `anthropic`, `openai`, `google`, … |
| `thinking` | `--thinking <v>` | `off`/`minimal`/`low`/`medium`/`high`/`xhigh`/`max` |
| `continueSession` | `-c` | Skipped when `resumeSessionId` is set (the two conflict) |
| `resumeSessionId` | `--session <v>` | Ids only, never paths |
| `approveProjectTrust` | `--approve` / `--no-approve` / nothing | Tri-state, see below |
Every value is regex-validated and **dropped** (not escaped) if it fails, because
the result is interpolated into the pane's `bash -c "…"` command.
The Run button sends **no `PiConfig` at all**: pi has no permission prompts to
bypass, and project trust is a decision the person at the terminal makes.
## What Codeman deliberately does NOT wire up
- **`--api-key`.** Never. It would put a provider secret on the spawn command
line, visible in `ps`, tmux server state and logs. `PI_*` overrides go through
socket-scoped `tmux setenv` for exactly this reason.
- **`--tui-mode`.** Pi's default main-screen TUI is the friendly case for a
browser terminal. The fullscreen mode (0.84.0) stays your own runtime choice via
`/settings`.
- **`--name`, `--no-session`, `-p`/`--print`, `--mode json`, `--mode rpc`,
`--tools`/`--exclude-tools`, `-e`/`--extension`, `--skill`,
`--system-prompt`.** Tracked as follow-ups in the plan doc.
## Permission and trust model — read this
**Pi has no permission prompts and no sandbox.** There is no
`--dangerously-skip-permissions` analog and none is needed: tools run with the
user's own permissions, always. A pi session can read, write and execute anything
the Codeman user can. If you need isolation, use a **Docker case** — that is the
isolation story, here as everywhere else in Codeman.
Pi's "project trust" prompt is **not** a safety boundary (upstream says so too).
It gates *loading* repo-local `.pi/` config, extensions and skills, and
*installing* missing project packages. It only appears when the cwd or an ancestor
contains `.pi/settings.json`, `.pi/extensions|skills|prompts|themes`,
`.pi/SYSTEM.md`/`.pi/APPEND_SYSTEM.md`, or `.agents/skills`. A bare `.pi/`
directory does not trigger it.
`approveProjectTrust: true` answers it with `--approve`, which means pi **loads
and executes repository-supplied TypeScript** and runs an npm install for missing
project packages. Treat it exactly as seriously as that sounds.
**Multi-user mode:** for an owner without the privileged-command grant, Codeman
materializes `approveProjectTrust: false` so the pane launches with
`--no-approve` and the prompt never appears. Merely *omitting* `--approve` would
not be a clamp, since pi's own default is to ask and the session user could just
answer yes.
Also worth knowing: `pi auth print-api-key` / `print-bearer-token` and
`pi auth check` mean a pi session can print its own provider credentials by
design. Isolation is Docker.
## tmux extended keys (Shift+Enter)
Pi's editor uses `Shift+Enter` / `Ctrl+Enter` for newline-vs-submit. Without
extended keys, tmux collapses both into a plain `\r`. Upstream recommends:
```tmux
set -g extended-keys on
set -g extended-keys-format csi-u
```
`extended-keys-format` needs tmux 3.5+; on 3.2–3.4 `extended-keys on` alone works
(pi falls back to xterm `modifyOtherKeys`).
Codeman's browser input path sends `\r` for submit, so basic use works
unconfigured — what degrades is newline-in-editor, mostly when you attach to the
pane directly (`sc`).
⚠️ Upstream notes the setting may need a full `tmux kill-server` to take effect.
**Never run `tmux kill-server` on Codeman's socket** — it would kill every live
session, `w1`/`w2`/`w3` included.
**Measured (tmux 3.4, pi 0.84.1): no `kill-server` is needed.** Setting the option
server-scoped on Codeman's own socket takes effect on the ALREADY-RUNNING server;
the next pi session starts without the warning. Existing sessions keep the old
setting until they respawn.
```bash
tmux -L codeman set -s extended-keys on
tmux -L codeman set -s extended-keys-format csi-u # tmux 3.5+ only, see below
tmux -L codeman show-options -s | grep extended # verify
```
On **tmux 3.4 and older, `extended-keys-format` does not exist** and the second
line fails with `invalid option: extended-keys-format`. That is harmless — pi
falls back to xterm `modifyOtherKeys` and `extended-keys on` alone silences the
warning. Run the two lines independently rather than chained.
Pi tells you which state it is in: an unconfigured session prints
`Warning: tmux extended-keys is off. Modified Enter keys may not work.` in its
startup banner, so you can verify the change by starting a new pi session.
⚠️ Use `-L <socket>` and `-s`, never `-g` on your default socket, and never
`kill-server`. Codeman does not set this for you: it is a server-wide tmux option
and silently changing key encoding for every session of every backend is not
Codeman's call to make.
## Typing from the browser (local echo)
On touch devices Codeman buffers typed characters in the `LocalEchoOverlay` and
flushes them to the PTY on Enter. Pi gets that `'buffer'` policy, the same as
Claude, Gemini and OpenCode.
This was an explicit open question, because that policy is exactly what broke
Codex (issues #218/#219/#220/#222): Codex's composer reacts per keystroke, so
buffer-until-Enter starved it. **Measured against pi 0.84.1: it does not
reproduce.** Pi's slash-command picker re-filters on the whole composer content
rather than on per-keystroke deltas, so a one-shot flush of `/set` filters the
picker down to `settings` identically to typing it character by character, and
the delayed `\r` then selects it. Prose prompts flush and submit correctly too.
If a future pi release changes that, the cheap fallback is one `'off'` branch in
`_updateLocalEchoState` (terminal-ui.js); teaching `PredictiveEchoAddon` pi's
composer row is the larger follow-up.
## Docker cases
The agent image (`docker/agent.Dockerfile`) installs pi in its own `RUN` step with
`--ignore-scripts`, kept out of the shared npm block so the flag cannot change how
the other four CLIs install. Rebuild with:
```bash
node scripts/build-agent-image.mjs --no-cache # --no-cache is mandatory
```
Credentials are **seeded**, not shared: `~/.pi/agent/auth.json`, `settings.json`,
`trust.json`, `models.json` and `models-store.json` are mounted read-only and
copied into the container's own `~/.pi/agent`. So an in-container pi never writes
refreshed OAuth tokens back to the host, and `docker commit` exports stay
secret-free. `models.json` is in the list because it holds user-defined custom
providers, which would otherwise silently vanish inside containers.
Only those five files are seeded because `~/.pi/agent` also holds `sessions/`,
`extensions/`, `skills/` and the installed package trees (`npm/`, `git/`), which
on an active host is easily gigabytes.
**Trade-off:** in-container pi sessions are invisible host-side, so `pi -c` inside
a Docker case only sees that container's own history.
## Remote SSH cases
`pi` mode is routed through an interactive login shell
(`exec "$SHELL" -i -l -c 'pi'`), because sshd's remote-command PATH does not
include npm's global bin on most hosts. Per-session config and `envOverrides` do
not cross ssh and are rejected rather than silently ignored; use the per-host
command override instead.
## Known gaps
- **No idle/completion hook.** Pi has no hook system Codeman can install into, so
idle detection falls back to output-stabilization like the other external CLIs.
Pi 0.84.0 shipped an `agent_settled` extension event that is a genuine idle
signal; a Codeman pi extension using it is the highest-value follow-up.
- **No response viewer.** Pi writes JSONL v3 session files under
`~/.pi/agent/sessions/`; nothing reads them yet.
- **Cron jobs mis-detect readiness.** The cron readiness poll looks for `❯` or a
token count, neither of which pi prints, so a pi cron job burns its poll budget
and then sends the prompt anyway. It works; it is just slower to start.
- **Ralph, respawn heuristics, token/CLI-info parsing and the `❯` readiness probe
are off** for pi, as for every external CLI.
+143
View File
@@ -0,0 +1,143 @@
# Predictive write-through echo for codex
Zero-lag local echo for codex sessions via a second, mosh-style mode in the
`xterm-zerolag-input` package: every keystroke goes to the PTY exactly as the
1.12.2 overlay-disabled path did (byte-identical wire behavior), while a
`PredictiveEchoAddon` simultaneously paints the predicted glyph at the predicted
cell. When the real echo lands, the prediction is confirmed and its span removed
(invisible swap: identical glyph beneath). Mispredictions drop via a mismatch
cascade + TTL. Visual-only, self-healing.
## Why this exists
Issues #218/#219/#220/#222 (one root cause) forced 1.12.2 to disable the
LocalEchoOverlay for codex: buffer-until-Enter starves codex's per-keystroke TUI
(live slash picker, arrows editing server-side composer state, composer
rewrap/growth, paste_burst classification). Buffer mode is structurally
incompatible with codex; write-through prediction is the only echo mode that
can coexist with it.
## The reconciliation lesson (do not regress this)
`docs/local-echo-overlay-plan.md` ("What NOT to Do") documented that matching
predictions against the raw output STREAM fails against Ink/TUI full-line
redraws. This design reads the parsed terminal BUFFER instead (cells after
xterm's parser ran), which converges to the same cells no matter how the bytes
arrived. The Phase 0 recordings prove the point twice over: tmux converts
codex's full-line redraws into minimal in-place deltas (an echo arrives as
`e\x1b[K\x1b[20;80H...`), and codex itself paints word gaps with ECH+cursor-forward
instead of spaces. Stream matching can never survive that; buffer diffing does
not care.
## Phase 0 measurements (codex-cli 0.147.0 via tmux, 100x30, 2026-08-09)
Recorded with `scripts/dev/record-codex-frames.mjs` (production pipeline:
codex inside tmux `status off`, chunks passed through the same full strip
`session.ts _handleTerminalOutput()` applies to codex mode). Fixtures in
`packages/xterm-zerolag-input/test/fixtures/codex/`; replay/measure with
`scripts/dev/analyze-codex-frames.mjs <fixture>`.
| Question | Measured answer |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Composer signature | Cursor row starts `"› "` (U+203A + space), text begins col 2. Present when empty (placeholder), while typing, and while the slash picker filters. `CODEX_COMPOSER_ROW_RE = /^› /` |
| Composer text color | Plain default foreground, zero SGR around echoed chars. Span `foregroundColor` default (theme fg) is an exact match |
| Placeholder | Cycling hint text ("Use /skills...", "Improve documentation in @filename", ...) rendered AT the cursor cell. First prediction lands over placeholder glyphs: covered by the snapshot + cursor-advance rules |
| Wrap | Word-wrap near `cols - 2`; continuation rows are indented 2 spaces WITHOUT `› `. The gate therefore suppresses predictions on wrapped lines: deliberate fallback to real echo, wrap was the #220 ghost zone. `edgeMarginCells = 4` |
| Modal (trust dialog) | Cursor parks on `" Press enter to continue"`: no `› ` prefix, gate false, zero predictions painted while keystrokes still reach the PTY (the ghost eliminator) |
| Streaming | Error/reconnect bursts render above a re-rendered composer that keeps the `› ` signature; end-of-frame cursor parks at the insertion point (col 2 of the composer row). Confirms the cursor-advance confirm rule and the no-drop-on-baseY rule |
| Echo shape under tmux | tmux emits minimal deltas for simple echoes and full repaints for busy frames; both converge in the parsed buffer |
| Slash picker | Picker rows render below; the cursor row keeps the composer signature and advances per filter char, so predictions stay active while filtering (#222 surface) |
Constants decided at the Phase 0 gate: `CODEX_COMPOSER_ROW_RE = /^› /`,
`ttlMs = 1000`, `maxPending = 32`, `cursorGraceMs = 150`, `edgeMarginCells = 4`,
span colors = theme defaults, `underlinePredictions = false`.
## Algorithm
See `PredictiveEchoAddon` in
`packages/xterm-zerolag-input/src/predictive-echo-addon.ts`. Summary of the
rules and why each exists:
- **State**: ordered `PredictionRecord[]` (`seq`, `char`, `width`, cumulative
`offsetCells`, `snapshot` of the cell at predict time, `sentAt`,
`mismatches`), plus a run `_anchor {row, col}` captured when the outstanding
count goes 0 -> 1. Positions are FIXED at predict time; confirmation deletes
spans and never re-lays-out, so partial confirmation causes zero jitter.
- **predictChar(ch)** runs an inline reconcile first and re-anchors whenever
outstanding drains to zero (absorbs the echo-landed-between-keystrokes race).
Guards: dims present, cursor numbers present, `viewportY === baseY`,
`predictWhen` gate, single codepoint >= 0x20 (not 0x7f), width <= 2,
`maxPending`, edge margin. Returns false = suppressed; the consumer sends the
keystroke regardless.
- **Coordinate base is `baseY`**: xterm's `cursorY` is baseY-relative, so
absolute buffer line = `baseY + row`. `viewportY` would only coincide while
the scrolled-to-bottom guards hold; the addon never relies on that.
- **reconcile()** (debounced `onWriteParsed` microtask, inline in predictChar,
TTL timer): clears everything when scrolled up; off-anchor-row cursor
tolerated for `cursorGraceMs` then clears; PREFIX-ONLY confirm loop requiring
cell match AND cursor advanced past the record (prevents false confirms
against placeholder glyphs and makes identical in-place tmux repaints a
no-op); TWO-PASS mismatch rule (a cell that is neither snapshot nor predicted
char must persist across two passes before cascading the drop: a half-parsed
row on pass N is fully redrawn a few ms later); TTL drop of the stale suffix.
- **No drop on baseY change**: codex streams push lines to history while the
composer stays viewport-pinned; predictions are row-relative to the pinned
composer and remain valid (measured above).
- **Anchor hold** (added by the independent post-build review): after any wire
input whose cursor effect the display has not shown yet (backspace with
nothing outstanding = deleting echoed text, every 'clear'-classified input,
an IME/plain-paste 'text' commit, and the bypass send paths), new
predictions are suppressed until the next PARSED write. Anchoring on the
stale cursor painted ghosts one cell off ("tehh" on backspace-then-retype
within RTT), blank-neutral and therefore TTL-lived. Worst case is exactly
one unpredicted keystroke: its own echo is a write, which releases the hold.
- **predictBackspace()** pops the newest outstanding record (informational
return; the consumer forwards `\x7f` unconditionally). Deleting already-echoed
text renders at RTT in v1.
- **CJK/wide**: 2-cell spans, stacking by cumulative visual width, leading-cell
confirm. In Codeman, IME input never reaches the hook (`window.cjkActive`
returns from onData first); package support exists for other consumers.
## Integration map (Codeman)
- Policy: `_localEchoPolicy` (`'buffer' | 'predict' | 'off'`) computed at the
end of `_updateLocalEchoState()`; codex + `localEchoEnabled` -> `'predict'`
while `_localEchoEnabled` stays false (every 1.12.2 consumer unchanged).
- onData hook sits between the buffer block and Normal Mode, classifies via
`classifyPredictInput()` (pure, on `window.CodemanTerminalInput`), never
returns, try/catch-wrapped: the wire path below is byte-identical with the
predictor active, absent, or throwing.
- Composer gate: `isCodexComposerRow()` set via `setPredictWhen()` at
construction (the vendor footer stays package-agnostic).
- Second vendor bundle `vendor/xterm-predictive-echo.js` (postinstall + build);
the zerolag bundle build command is untouched and its output byte-identical.
Missing/broken bundle = plain 1.12.2 echo (`typeof PredictiveEchoOverlay ===
'undefined'` guard).
- Prediction clears on: tab switch, SSE reconnect init, `insertTerminalText`,
`clearTerminalInput`, voice send, keyboard-accessory `sendKey`, resize, skin
and font changes re-read style via `refreshFont()`.
## Risk register
Eliminated structurally: other-mode regression (zero edits to buffer
addon/branches, byte-identical existing bundle, policy-matrix + byte-identity
tests); bundle breakage (separate bundle, graceful degradation); wire
corruption (no-return fall-through + try/catch + byte-identity pins at vm and
E2E level); modal ghosts (measured predictWhen gate); false confirms
(cursor-advance rule); mid-parse flicker drops (two-pass rule); wrap
misplacement (edge margin + continuation-row gate fallback + off-row grace).
Accepted residuals (visual-only, self-healing <= ttlMs, kill-switchable via
`localEchoEnabled` per device): no predictions on wrapped continuation lines
(gate false there, deliberate); brief dropout during composer growth; DOM-span
vs WebGL glyph rendering can differ subtly (same trade-off as the buffer
overlay, same font recipe); typing during an unsynchronized half-frame can
mis-anchor one run (mismatch/TTL cleans within 1s).
## Future work
RTT-adaptive TTL; mosh-style confidence gating (paint only after the link
proves laggy); predicted backspace into echoed text; predict mode for shell
prompts; unifying the small font/container duplication between the two addons
once predict mode has proven out; continuation-line prediction behind a
smarter composer-extent detector.
+140
View File
@@ -0,0 +1,140 @@
# Read My Mind (design)
A 🧠 button that predicts the prompt you were about to type. Codeman keeps a per-case **intent profile** (your stated goals plus the real prompts you recently sent), feeds it and the live pane tail to a one-shot `claude -p`, and shows the predicted next prompt in a plan-mode-style approval dialog: **Send** / **Rethink** (with an optional steer note) / **Insert** (drop it on the composer to edit) / **Dismiss**. It is also a skill surface: the agent can read the intent profile, record intentions, and request a prediction over the HTTP API. Suggestions are **never auto-sent**; the human click is the boundary.
## UX flow
1. User hits 🧠 (desktop header button; phone: keyboard-accessory key).
2. Modal opens with a spinner, then the top suggestion in an editable single-line field, rationale below it, up to 2 alternates as tappable rows.
3. Buttons: **Send** (submits with `\r`), **Insert** (sends without `\r`, so the text sits unsubmitted on the CLI composer for editing, a documented mechanism), **Rethink** (optional free-text steer, e.g. "no, I meant the mobile bug", re-runs with the rejected suggestions included), **Dismiss**.
4. Accepted prompts flow back into the intent history like any other sent prompt, so the profile self-corrects.
## Scope (v1)
- Claude mode only (capture rides Claude transcripts; external CLIs have no transcript watcher). Mirrors the approvals-inbox scoping.
- Opt-in: `readMyMindEnabled`, synced, default **OFF**. While OFF: no capture, no UI surfaces. Privacy first, and every press costs real tokens.
- One prediction in flight per session; the button disables while checking.
- Sync request/response (the predictor takes 5-30s; agent-wait long-polls already hold requests longer). No new SSE events in v1.
## Data model
Per case, not per session: intentions outlive `/clear` and respawns.
```ts
interface IntentProfile {
key: string; // sha256(owner + ':' + realpath(workingDir)).slice(0, 16)
workingDir: string;
updatedAt: number;
goals: string; // freeform markdown, user/agent editable, ≤ 8 KB
recentPrompts: { ts: number; sessionId: string; text: string }[]; // FIFO cap 50, each ≤ 500 chars
}
```
Storage: `dataPath('intents.json')`, written mode 0600 (prompts can contain secrets; same posture as `users.json`). Never enters the `/api/search` index. Add to the CLAUDE.md State Files list.
## Intent capture
**Source: the session transcript, not the input paths.** `POST /api/sessions/:id/input` sees only programmatic input, and the WS channel delivers raw keystrokes (`session.write(msg.d)`), so neither yields clean submitted prompts. Claude's own JSONL transcript records every user turn as structured text, and `transcript-watcher.ts` already tails it. Add a `userPrompt` event there:
- Emit for `type: 'user'` entries whose content is a string or contains a text block; skip entries that are only `tool_result` blocks (tool results are wrapped as user messages).
- Skip `<command-name>` / `<local-command-stdout>` tagged entries (local slash-command echo, not intent).
- Skip texts < 3 chars (menu digits, Esc artifacts), truncate to 500, drop consecutive duplicates ("continue" spam from auto-resume stays but dedupes).
`IntentStore` (new `src/intent-store.ts`, pure core + IO wrapper, in the style of `session-order.ts`) subscribes via session wiring, gated on the setting resolved from **merged** settings per the partial-PUT rule.
## Context assembly (how the mind reading actually works)
The quality of the suggestion is decided before the model ever runs, by what we put in front of it. A new pure function `buildPredictionContext()` (in `src/readmymind-context.ts`, unit-testable with fixtures, no IO of its own; collectors inject their data) assembles a budgeted, priority-ordered prompt from every signal Codeman already has:
| # | Source | What it contributes | Cap |
| - | ------ | ------------------- | --- |
| 1 | **Pending dialog** (approvals-inbox store, when present) | If the session is sitting on an AskUserQuestion / permission / idle prompt, the honest "next prompt" is an *answer*. The dialog text + parsed options go in first and the model is told to answer it. | 2 KB |
| 2 | **User goals** (`goals` from the intent profile) | The only fully-trusted statement of what the user wants. Highest authority in the trust ranking below. | 8 KB |
| 3 | **Last assistant turn** (transcript, not the pane) | Assistant replies usually *end* with the fork in the road ("Want me to X?", "Next steps: ..."), so keep the **tail** when truncating. The transcript has the full message; the pane is a repaint window full of spinner junk. | 6 KB |
| 4 | **Recent user prompts** (intent profile, with timestamps) | The conversation rhythm AND the user's prompting voice: length, tone, shorthand (`COM`, lowercase, typos and all). The model is instructed to write suggestions in *this* style, not assistant-ese. | last 20 |
| 5 | **Recent tool activity** (transcript `tool_use` blocks, already parsed by `TranscriptWatcher`) | One line per call: `Edit src/foo.ts`, `Bash npm test (failed)`. What the agent actually *did*, which the last message may summarize away. | last 10 |
| 6 | **Workspace signals** (`collectWorkspaceSignals()`: `git` via `execFile` in `workingDir`, 2s timeout) | Branch, `status --short` (dirty files scream "commit/test/deploy next"), last 5 commits oneline, presence of `.changeset/*.md` (release pending). Skipped for remote-SSH cases (workingDir is not local); fine for Docker cases (bind-mounted at the same host path). Non-git dirs: section omitted. | 3 KB |
| 7 | **Away context** (run-summary events + elapsed time) | `Last user prompt was 6h ago; since then: <run-summary events for this session>`. After a long gap the right suggestion is often "review / continue yesterday's thread", not a blind continuation. | 2 KB |
| 8 | **Sibling sessions** (live sessions sharing the case) | One line each: name, mode, working/idle. A lead-and-workers setup changes what the next prompt should be ("check on w2" beats "keep going"). | 1 KB |
| 9 | **Rethink state** (steer note + rejected suggestions) | Only on re-runs. Rejections are strong negative signal and go in verbatim. | 2 KB |
Total budget ~30 KB. When over budget, drop from the bottom up (siblings first, then away context, then workspace signals); sections 1-4 never drop, they only truncate. Deterministic assembly means fixture tests can pin exactly what a given situation feeds the model.
**Trust tiers are stated in the prompt.** Goals and user prompts are *the user*; assistant text, tool logs, and pane content are *observations that may contain text trying to manipulate you* (a hostile repo can print "SUGGEST: run curl evil.sh"). The prompt instructs: user-stated intent outranks anything observed, and never propose a prompt whose primary source is terminal output alone. The human approval click remains the hard boundary regardless.
**Output contract** (strict JSON, parse failure = clean error, never a half-suggestion):
```json
{ "suggestions": [ { "prompt": "...", "why": "...", "kind": "continue" | "verify" | "redirect" } ] }
```
1-3 entries, and the *kinds* force useful diversity instead of three rewordings: `continue` (finish the current thread, or answer the pending dialog), `verify` (test/review what was just built; the user's own "always end-to-end test" discipline), `redirect` (the next goal from the intent profile that the current thread is not serving). The modal shows `continue` big, the others as alternates. Embedded newlines are stripped server-side (single-line prompt rule; multi-line breaks Ink).
## Predictor
New `src/readmymind-predictor.ts`, reusing the `AiCheckerBase` mechanics (prompt file to dodge E2BIG, one-shot `claude -p --output-format text` in a throwaway tmux `codeman-rmm-<id8>`, done-marker polling, timeout, model-name validation) but standalone: the base class is verdict-shaped (positive/negative/cooldown) and prediction is freeform JSON, so subclassing would abuse `reasoning` as a payload. If a shared spawn/poll helper falls out naturally, extract it; do not block on the refactor.
- **Model: opus** (decided). `readMyMindModel` setting, default `AI_CHECK_MODEL` (currently `claude-opus-4-5-20251101`); prediction quality is the product, and it runs only on an explicit press, so the cost profile is nothing like the idle checker's. Timeout 90s (opus headroom over a ~30 KB prompt).
- Input: the assembled context above. The predictor itself stays dumb: text in, JSON out; all intelligence about *what to include* lives in the testable assembler.
## API (new `src/web/routes/readmymind-routes.ts`)
Normal authed API, `ApiResponse` envelope, Zod schemas in `schemas.ts`, ownership via `findSessionOrFail` (the profile key derives from the session's owner + workingDir, so multi-user scoping is structural):
- `GET /api/sessions/:id/intent` → the session's `IntentProfile`.
- `PUT /api/sessions/:id/intent` body `{ goals }` (bounded) → update goals. Used by the modal's edit view and by the agent skill ("record that the user is working toward X").
- `DELETE /api/sessions/:id/intent` → forget everything for this case (the modal's "Forget" affordance).
- `POST /api/sessions/:id/readmymind` body `{ steer?, rejected? }` → `{ suggestions }`. 409 `INVALID_STATE` while a prediction is already running for the session; claude-mode sessions only (400 otherwise, mirroring wait-signal gating).
## Frontend
New module `readmymind-ui.js` (@loadorder 11.3, after panels-ui.js), prettier-formatted.
- **Desktop**: header button `btn-readmymind`, default-hidden via marker class `btn-readmymind--hidden` (the `!important` display rules require the marker-class pattern), shown by `applyHeaderVisibilitySettings()` when the setting is ON. Off phones per `test/mobile-header-buttons-policy.test.ts`.
- **Phone**: a 🧠 key on the keyboard accessory bar (that bar is where input helpers live, and phones are where typing hurts most). Opens the same modal. Modal z-index respects the ≤768px layer rules (1300+).
- **Send** goes server-side: `POST /api/sessions/:id/input` with `\r` appended. Deliberately NOT the browser keystroke path, so the `sendEnterKey` / local-echo-overlay trap never applies (the modal is UI chrome, not terminal typing). **Insert** is the same POST without `\r`.
- i18n strings registered (en + zh-CN); suggestion text itself carries `data-i18n-skip`.
## Skill integration
The user-facing promise: the button is also a skill. Extend `skills/codeman`:
- New section "Read My Mind: intent + prediction" with the three intent verbs (read profile, append/replace goals, predict) and the guard notes (single-line prompts, never auto-send to another session without the user asking).
- Update `reference/endpoints.md` (the endpoints.md drift test pins this).
- The auto-injected case copy heals via the existing marker-owned `applyAgentSkill` mechanism; nothing new needed there.
Agent use cases this unlocks: a lead session records intentions as the user states them ("remember: shipping 1.16 is the goal"), and a returning user gets a prediction grounded in what the agent knew, not just raw prompt history.
## Security / privacy
- **The human gate is the injection mitigation**: pane output (attacker-influenceable) flows into the predictor, so its output is only ever *proposed*, rendered as text (`textContent`), and sent solely by an explicit user click. No auto-send path exists, including for the skill.
- Intent data: 0600 file, bounded fields, per-owner keys, endpoints ownership-checked, excluded from search, cleared via DELETE.
- Predictor spawns with the user's own credentials exactly like the AI idle/plan checkers; model name shell-validated the same way.
- Setting OFF stops capture immediately; existing data stays until DELETE (explicit, not silent).
## Tests
- `test/intent-store.test.ts`: key derivation, caps/FIFO, consecutive-dupe skip, tag/tool_result filtering fixtures, 0600 mode, multi-user key separation.
- `test/readmymind-context.test.ts`: fixture scenarios pinning the assembled prompt: pending-dialog-first ordering, tail-keeping truncation of the assistant turn, budget drop order (siblings before workspace signals), remote-case git skip, trust-tier framing present, rejected suggestions included only on rethink.
- `test/readmymind-predictor.test.ts`: strict JSON parse, garbage output → error result, newline stripping, `kind` validation, rejected-suggestions threading into the prompt.
- `test/routes/readmymind-routes.test.ts` (`app.inject`): CRUD round-trip, predict with a stubbed predictor, 409 while in flight, non-claude 400, ownership 404, Send/Insert byte assertions via the test-PTY echo (`\r` present vs absent).
- Transcript capture: extend the transcript-watcher fixtures with user-turn entries.
## Phases
1. **Intent store + capture + intent endpoints + skill docs.** Immediately useful to agents even before any UI exists.
2. **Context assembler + predictor + predict endpoint + desktop button/modal.** The feature as pitched. The assembler ships with all collectors it can serve from day one (transcript, intent, git, run-summary, siblings); the approvals collector activates when PR #245 lands.
3. **Phone accessory key, rethink steering, alternates row.** Part 1 (shipped): the alternates row (tappable, swap into the field without losing edits; Rethink rejects the whole shown set), the phone 🧠 keyboard-accessory key (both bar templates, `rmm-enabled` marker class on the bar), and a phone-sized modal (small dialog, not full-screen). Part 2 (shipped): rethink steering, the free-text steer note under the suggestions, sent as `steer`, visible whenever Rethink is live (ready and empty-result phases), cleared on each open; the empty-result copy points at the note, and the footer buttons moved to the styled `btn-toolbar` convention (the bare `btn btn-*` classes they shipped with match no CSS in this codebase and rendered as unstyled UA buttons).
4. Explicitly later: proactive predict-on-idle (ghost suggestion chip), auto-compaction of `recentPrompts` into `goals` via a cheap model, codex/gemini capture, cross-case "global" intent.
## Open questions
- Should Rethink's rejected-suggestion memory persist across modal closes, or reset each open?
- Is a composer-adjacent placement (next to the toolbar Run controls) better than the header for discoverability?
- Pending-dialog input (source #1) consumes the approvals-inbox store (PR #245, merged): the phase-2 collector reads pending items directly from `src/approval-inbox.ts`.
## Docs
- CLAUDE.md: Key Patterns entry, State Files (`intents.json`), frontend load order, route count.
- `docs/api-reference.md`: four endpoints (additive under the 0.9.x contract).
- `skills/codeman/reference/endpoints.md`: new rows (drift-test enforced).
+108
View File
@@ -0,0 +1,108 @@
# Read My Mind
Codeman's per-case memory of what you are trying to accomplish, and the 🧠 button that turns it into a predicted next prompt. Each case gets an **intent profile**: a freeform `goals` text (written by you or your agent) plus the prompts you actually submitted, captured automatically while the feature is on. Pressing 🧠 feeds that profile and the live session signals to a one-shot model call and shows the predicted prompt for you to send, edit, or rethink. Nothing is ever sent to a session automatically. Design doc: [`readmymind-plan.md`](readmymind-plan.md).
## What it does
- Captures the prompts you submit in Claude sessions into a per-case history (50 most recent, bounded).
- Lets you (or your agent) record explicit goals per case.
- Predicts your next prompt on demand (the 🧠 header button, or `POST .../readmymind` for agents): the suggestion arrives in a modal with Send / Insert / Rethink / Dismiss.
- Exposes the profile over the HTTP API, and to agents through the `codeman` skill, so an agent can ground its work in what you actually want instead of guessing from the last screenful.
## Turning it on
App Settings → Header & Panels → Cross-session features → **Read My Mind** (synced setting `readMyMindEnabled`, default **OFF**). It gates everything: capture, the header button, and nothing shows anywhere while it is off. The API equivalent:
```bash
curl -sk -X PUT https://localhost:3000/api/settings \
-H 'Content-Type: application/json' \
-d '{"readMyMindEnabled": true}'
```
Add `-u user:password` if your install has `CODEMAN_PASSWORD` set, and drop `-k`/use `http://` for a plain-HTTP dev server. Turning it OFF stops capture immediately; existing profiles stay until you delete them (below).
## The 🧠 button
On a Claude session, press the brain button in the header (desktop) or the 🧠 key on the keyboard accessory bar (phones and tablets; it appears when the setting is on). Codeman assembles everything it already knows: your goals, your recent prompts (with your voice: length, tone, shorthand), the tail of the last assistant reply, recent tool activity, git state (branch, dirty files, pending changesets), how long you have been away and what happened meanwhile, sibling sessions in the same case, and any dialog the session is currently waiting on. A one-shot model call (opus by default, `readMyMindModel` to override) turns that into 1-3 suggestions; the top one lands in an editable field with its rationale, and the others render as tappable alternate rows: tap one to swap it into the field (edits you already made are kept on the row you leave).
- **Send** submits it to the session (with Enter).
- **Insert** drops it on the CLI composer *without* Enter, so you can edit it in the terminal before sending.
- **Rethink** re-runs with everything shown (the field and the alternates) recorded as rejected. An optional steer note below the suggestions ("no, I meant the mobile bug") rides along as your own words, the highest-authority signal the predictor gets; it stays in the field across re-runs until you clear it or reopen the modal.
- **Dismiss** closes; nothing happens.
A prediction takes 5-90 seconds and costs real tokens; one runs per session at a time. If the session is sitting on a permission/question dialog, the suggestion is usually an answer to that dialog: that is intentional.
**Security note**: the prediction reads observable content (assistant output, tool logs, git output) which a hostile repo could try to steer. The predictor is told user-stated intent outranks anything observed, and, more importantly, a suggestion is only ever *proposed*: your click is the boundary. No auto-send path exists, including for agents.
## What gets captured, exactly
Capture reads the Claude session transcript, not your keystrokes: when a user turn lands in the transcript, its text is folded into the case's profile. Filters applied on the way in:
- **Claude-mode sessions only.** Shell, OpenCode, Codex, Gemini, Antigravity, and Pi sessions are never captured (they have no transcript watcher).
- Tool results, local slash-command echo (`/model` and friends), system wrappers, and interrupt markers are skipped.
- Entries shorter than 3 characters are skipped (menu digits, Esc artifacts).
- Consecutive duplicates collapse (auto-resume's "continue" spam counts once per run).
- Each prompt is stored as one line, truncated to 500 characters; the history caps at 50 prompts FIFO.
Because the transcript path arrives via Claude Code hooks, capture needs hooks to reach the server, the same condition as hook-based idle detection. Docker cases against a loopback-only server need `CODEMAN_DOCKER_BRIDGE_HOOKS=1`; remote-SSH cases do not capture.
## What is never captured
- Anything while `readMyMindEnabled` is OFF (capture is not retroactive).
- Terminal output, keystrokes, passwords typed into shells: only submitted Claude prompts are read.
- Nothing leaves the machine beyond the model call you explicitly trigger, and profiles are never fed into `/api/search`.
## Where it lives, and how to wipe it
Profiles live in `~/.codeman/intents.json`, written atomically at mode 0600 (captured prompts can contain secrets). The file is per Codeman instance. Keys derive from owner + the case's resolved working directory, so profiles survive `/clear`, respawn cycles, and session churn, and in multi-user mode two owners of the same directory get separate profiles.
Forget one case: `DELETE /api/sessions/:id/intent` (below). Forget everything: stop the server and delete `~/.codeman/intents.json`.
## The API
Four endpoints, session-scoped so ownership is enforced by the session itself (`/api/v1/` aliases work too; full spec in [`api-reference.md`](api-reference.md)):
```bash
# Read the profile for a session's case
curl -sk https://localhost:3000/api/sessions/$SID/intent | jq '.data.intent'
# Record goals (REPLACES the text: read + merge if you want to append)
curl -sk -X PUT https://localhost:3000/api/sessions/$SID/intent \
-H 'Content-Type: application/json' \
-d '{"goals":"ship 1.17; then mobile polish"}'
# Forget the case
curl -sk -X DELETE https://localhost:3000/api/sessions/$SID/intent
# Predict the next prompt (claude-mode only; takes 5-90 s)
curl -sk -X POST https://localhost:3000/api/sessions/$SID/readmymind \
-H 'Content-Type: application/json' -d '{}' | jq '.data.suggestions'
```
A case with nothing recorded answers an empty profile with `updatedAt: 0`; reads never persist anything. Goals cap at 8192 characters and the schema is strict, so unknown fields or over-long goals answer `400 INVALID_INPUT`. A session you do not own answers `404 NOT_FOUND`, indistinguishable from a nonexistent one. Predict answers `{ suggestions: [{ prompt, why, kind }], durationMs }` (`kind`: `continue` / `verify` / `redirect`), `409 CONFLICT` while one is already running, `400 INVALID_INPUT` on non-claude sessions, and `502 OPERATION_FAILED` when the model produced no usable JSON. The rethink flow passes `{"steer":"…","rejected":["…"]}`.
## For agents (the skill)
The `codeman` agent skill documents the same verbs (SKILL.md §3 plus `reference/endpoints.md`), with the ground rules: read the profile to understand what the user wants, record goals the user actually stated, merge instead of blind-writing (PUT replaces), never delete a profile unprompted, and never send a predicted suggestion into a session unless the user asked. It is the user's memory, not the agent's.
## What comes next
Explicitly later: proactive predict-on-idle, auto-compaction of the prompt history into goals, non-Claude capture. See the phases section of [`readmymind-plan.md`](readmymind-plan.md).
## Troubleshooting
| Symptom | Cause / fix |
| ------- | ----------- |
| No 🧠 button in the header | `readMyMindEnabled` is OFF (App Settings → Header & Panels → Cross-session features), you are on a phone (there it is a key on the keyboard accessory bar instead, visible while typing), or the active session is not claude-mode |
| Prediction feels generic | The profile is thin: record goals (PUT or ask your agent to), and let capture accumulate a few real prompts first |
| "A prediction is already running" (409) | One per session at a time; wait for the current one (up to 90 s) |
| Prediction fails (502) | The model returned no usable JSON, or the CLI could not start; retry. Check `readMyMindModel` if you overrode it |
| Profile stays empty although I am prompting | `readMyMindEnabled` was OFF at the time (capture is not retroactive), the session is not claude-mode, or hooks are not reaching the server (Docker case on a loopback bind without `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, or a remote-SSH case) |
| Short answers I typed are missing | Entries under 3 characters are filtered by design (menu digits, Esc artifacts) |
| My goals text vanished after an agent wrote to it | PUT replaces the whole text; the skill tells agents to read + merge, but a blind write wins. Re-state the goals; consider phrasing them in the session so capture keeps the evidence |
| Two profiles for what I think is one case | Different owners in multi-user mode, or genuinely different directories; paths are realpath-resolved, so symlink spellings converge but distinct checkouts do not |
| `400 INVALID_INPUT` on PUT | Goals over 8192 chars, or an extra field in the body (strict schema) |
## Where the code lives
`src/intent-store.ts` (store + pure helpers, singleton), the `transcript:user_prompt` event in `src/transcript-watcher.ts`, capture wiring in `src/web/server.ts` (`captureIntentPrompt`), context assembly in `src/readmymind-context.ts` (pure) + `src/readmymind-collectors.ts` (transcript tail + git IO), the predictor in `src/readmymind-predictor.ts`, routes in `src/web/routes/readmymind-routes.ts`, schemas in `src/web/schemas.ts`, frontend in `src/web/public/readmymind-ui.js`. Tests: `test/intent-store.test.ts`, `test/readmymind-context.test.ts`, `test/readmymind-collectors.test.ts`, `test/readmymind-predictor.test.ts`, `test/routes/readmymind-routes.test.ts`, and the capture cases in `test/transcript-watcher.test.ts`.
+3 -3
View File
@@ -1,7 +1,7 @@
# Remote Sessions (SSH)
Codeman can run a session's agent on a **remote host over SSH** instead of the
local machine. The agent (Claude, OpenCode, Codex, Gemini, or a plain shell)
local machine. The agent (Claude, OpenCode, Codex, Antigravity, Gemini, Pi, or a plain shell)
runs inside a `tmux` server **on the remote host**, so it survives the SSH
connection dropping; Codeman attaches to it the same way it attaches to a local
managed session.
@@ -30,7 +30,7 @@ Types live in `src/types/session.ts`; persistence in `src/remote-hosts.ts`.
| `RemoteHost` (extends `RemoteSshOptions`) | A saved host: `id`, `label`, `host`, `username`, `port?`, `commands?` (per-mode launch command override). |
| `RemoteCase` | A working directory on a host: `name`, `type: 'remote'`, `hostId`, `remotePath`. |
| `SessionRemote` (extends `RemoteSshOptions`) | The resolved bundle stamped onto a live session: host coordinates + `remotePath` + `commands`, plus **`owned?`** and **`remoteSessionName?`** (COD-105 — see [Ownership](#ownership-launched-vs-discovered-and-attached-cod-105)). Built by `toSessionRemote(host, case)` (sets `owned: true`) for the launch path, or `toAttachedSessionRemote(host, name, path)` (sets `owned: false`) for the attach path. Both copy the advanced SSH options through so every connection is identical. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini'>` — the modes that can run remotely. |
| `RemoteCommandMode` | `Extract<SessionMode, 'shell' \| 'claude' \| 'opencode' \| 'codex' \| 'gemini' \| 'antigravity' \| 'pi'>` — the modes that can run remotely. |
| `RemoteSessionInfo` (COD-105) | One discovered remote tmux session: `name` (always `codeman-*`), `attached` (a client is connected), `created` (epoch s), `windows`. Returned by `listRemoteCodemanSessions()`. |
Persistence is two flat JSON arrays in the instance data dir:
@@ -115,7 +115,7 @@ Key points:
- **`exec <cli>`** replaces the pane shell with the agent, so the pane PID *is*
the agent. The per-mode command comes from `remote.commands?.[mode]` or
`defaultRemoteCommandForMode(mode)` (`exec claude` / `exec opencode` /
`exec codex` / `exec gemini` / `exec bash -l`).
`exec codex` / `exec gemini` / `exec agy` / `exec bash -l`).
- The **whole tmux invocation is a single shell-quoted ssh argument**, and the
pane command is independently quoted, so a `remotePath` with spaces is safe.
- Connection options come from the **same `buildSshConnectionArgs(remote)`** as
+301
View File
@@ -0,0 +1,301 @@
# Scrollback fix plan (issue #205)
Status: IMPLEMENTED on `fix/scrollback-shell-alt-screen` (2026-08-07), with one deliberate
divergence from the recommendation below. Kept for the diagnosis record; the measured evidence
behind it is `docs/scrollback-issues-analysis.md`, and the mechanisms as shipped are documented
in `docs/architecture-invariants.md` (§ Full-scrollback replay, § Terminal scrollback: strip
flavors and wheel/touch forwarding).
What shipped vs. what this doc proposed:
- **Bug A (deltaMode)**: implemented as specified (`_wheelScrollLines()` normalizes
line/page/pixel units, Shift-axis trap kept).
- **Bug B (shell scrollback)**: implemented via the NARROW alt-screen strip for tmux-backed
shell/opencode/antigravity plus the scroll-to-top `full=1` re-pull, NOT the recommended
approach (a) `tmux mouse on`. The measurements in the analysis doc showed the alt buffer
comes from tmux's own client-side `smcup` at attach (tmux never forwards a pane program's
alt-screen toggles), so stripping that one sequence fixes both symptoms with no selection
tradeoff, keeps vim/less/htop untouched, and the re-pull also covers the repaint-burst
history loss that `mouse on` would not have addressed.
- **Invariant change**: the "viewport-at-bottom gate stays" invariant below was deliberately
DROPPED for forwarding modes: a repaint-mode CLI keeps no real terminal scrollback, so the
gate pinned users to a buffer of stale frames whenever the viewport parked off-bottom.
Forwarding now snaps to bottom first; Shift+wheel and the opt-out setting keep local
scrollback reachable. Touch forwards through the same gate (the mobile half of the fix).
- **Finding 5 (remote probe)**: implemented (`probeRemoteCliVersion` over ssh, deferred at
session start, same login-shell wrapper as the launch).
## RETEST FAILED (2026-08-07, after v1.12.0 shipped) — analysis round 2
mtiller retested on 1.12.0 and reports it is NOT fixed (issue #205 comment, 2026-08-07 12:12 UTC;
issue reopened same day with clarifying questions: mouse vs trackpad, Shift+scroll behavior,
Claude vs shell session on the phone, and an iOS full-tab-kill to rule out stale JS). Two
failure signatures, now analyzed against the SHIPPED 1.12.0 code (not the pre-fix code):
1. **iPhone Safari (Claude session assumed)**: touch scrollback goes back only a limited
amount and sometimes REPEATS blocks of text; unreliable.
2. **Firefox on macOS (mouse)**: wheel does NOTHING at all, while Fn+Up (= PageUp) pages back
through INTACT text.
### Ruled out by code reading
- deltaMode mishandling: `_wheelScrollLinesFloat` normalizes line/page/pixel units correctly;
a Firefox line-mode notch yields ±3 lines. Not the bug.
- Ephemeral transport: `_sendInputEphemeral` (app.js) has a POST fallback when WS is down.
- Service worker: sw.js is network-first with cache fallback; it serves stale JS only when the
fetch FAILS (flaky mobile connection can do this — relevant to "unreliable" on the phone,
and the fixed `CACHE_NAME = 'codeman-v1'` never invalidates that offline copy).
### The load-bearing observation: PageUp works, the wheel does not
Fn+Up is a KEYBOARD event: xterm encodes PageUp and Claude pages its own transcript (intact
text proves Claude-side history is fine and the PTY input path is fine). The wheel path is the
capture-phase handler, and for a Claude session it has exactly two branches:
- **Forwarding branch** (`_shouldForwardWheelToApp` true): snap-to-bottom + SGR reports. If
this branch ran, the user would see the same paging motion Fn+Up produces. They see nothing.
- **Local branch** (gate false): `_smoothScrollBy` over xterm's local buffer. For a Claude
pane in repaint mode, tmux keeps `history_size≈0`, so `?full=1` returns roughly one frame:
the local buffer is structurally HOLLOW, the top-of-buffer re-pull recovers nothing, and the
wheel looks completely dead. **This matches every observed detail on Firefox.**
So the working hypothesis is that mtiller's sessions evaluate the gate FALSE. The gate
(`_shouldForwardWheelToApp`) has exactly four false-paths worth checking, in likelihood order:
1. **`terminalWheelLocalScrollback` opt-out is ON.** Plausible: a user whose scrolling was
broken on 1.11.x may well have toggled "Wheel scrolls local history" while trying to fix
it. On 1.12.0 that setting now routes the wheel to a hollow local buffer = dead wheel on
desktop AND the stale-repaint-frames experience on the phone (see below). Ask, or check
what the setting does on their export.
2. **`cliVersion` missing — CONFIRMED BUG, independent of whether it is mtiller's**:
`getClaudeCliVersion()` (utils/claude-cli-resolver.ts:124-148) caches its result
process-wide including FAILURE: on any exception it sets `_claudeVersion = null`, and the
guard is `!== undefined`, so a single failed/timed-out probe (5s `EXEC_TIMEOUT_MS`; PATH
under systemd/launchd; transient fs hiccup) at the FIRST Claude session start disables
wheel forwarding for every Claude session until the server restarts. Fix: cache success
permanently, but let failure retry (retry on next call, or a short negative-cache TTL).
Note that mtiller sees identical breakage on phone + iPad + laptop, which points at a
SERVER-side/session-side cause exactly like this (cliVersion is shared by all devices)
rather than anything browser-specific.
3. **Claude Code genuinely < 2.1.187** on their machine: gate false BY DESIGN, but the
resulting UX is a dead-end (no local history to fall back on).
4. mouseTrackingMode non-none (a DECSET leaked past the strip, e.g. emitted before attach or
split across chunks in a way the carry missed): would also kill the container handler via
the early return. Least likely, checkable via `terminal.modes.mouseTrackingMode` in console.
### The iPhone symptoms fit the same gate-false story
Touch with gate false = local `scrollLines()` over whatever repaint frames accumulated:
"repeats blocks of text" is literally what a buffer of successive overlapping repaint frames
looks like; "limited amount" is its thinness; "unreliable" is burst-dependence (finding 2)
PLUS the new re-pull being actively DESTRUCTIVE for repaint panes: `_maybeRefetchFullHistory`
does `_resetTerminalForReplay()` then writes the fetched capture, and when that capture is
one frame (Claude pane, `history_size≈0`) it REPLACES a multi-frame buffer with less than the
user had, mid-scroll. Stale pre-1.12 JS on the phone (suspended Safari tab) remains possible
until they confirm the tab kill.
### Fix directions, ranked
1. **Make the re-pull refuse downgrades** (`_maybeRefetchFullHistory`, app.js): if the fetched
capture would yield FEWER buffer rows than currently present, skip the reset+rewrite and
keep the richer buffer (optionally cache-mark the session "re-pull useless"). Small, safe,
kills the "got worse after scrolling to top" class. Consider skipping the re-pull entirely
for forwarding-capable modes where tmux keeps no history.
2. **Rescue the gate-false Claude dead-end with PageUp forwarding**: when mode is `claude`,
the gate is false, AND the local buffer has no scrollback (`baseY === 0`), translate wheel
lines into coalesced PageUp/PageDown key sends (mtiller just proved Claude pages correctly
on PageUp even on their version). Zero regression risk under that triple guard: sessions
with real local history keep local scrolling; only the currently-dead path changes.
Caveat: older Claude menus may react to PageUp; acceptable against "completely dead".
3. **Audit `getClaudeCliVersion()` failure caching** (utils/claude-cli-resolver.ts): a cached
empty probe must retry (with backoff), not poison the process.
4. **Guard the opt-out setting's footgun**: if `terminalWheelLocalScrollback` is ON for a
repaint-mode CLI session, local history is hollow; either scope the setting's effect to
modes with real local scrollback, or pair it with fix 2's PageUp fallback so it still
scrolls SOMETHING.
5. **Add a one-line gate diagnostic**: log (once per session, console) WHY the wheel chose
local vs forward: `{mode, cliVersion, optOut, trackingMode}`. The #205 thread is now two
rounds deep on guesswork a single console line would have answered.
### What shipped for round 2 (branch `fix/scrollback-205-round2`)
All five directions above, implemented as ranked:
1. **Downgrade guard** — `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the rows a
capture will occupy (ANSI stripped, `capture-pane -J` re-wrapping accounted for) and
`_maybeRefetchFullHistory` (app.js) skips the reset+rewrite when that is more than one
screen short of what xterm already holds. A refused session goes on
`_fullHistoryRepullUseless`, which raises its re-pull cooldown from 4s to 60s so a hollow
pane stops re-fetching. Measured A/B on a live Claude pane, same gesture, same buffer:
guard off → 341 rows collapse to 42 and every seeded row is gone; guard on → 341 rows
preserved. The tab-switch recovery it must not break still runs (shell buffer 401 → 44 on
a tab switch → 401 again after scrolling to the top).
2. **PageUp/PageDown fallback** — `_maybePageCliTranscript()` translates wheel/touch travel
into coalesced `\x1b[5~` / `\x1b[6~` under the triple guard (claude mode, forwarding gate
false, `baseY === 0`), through the same 40ms queue as the SGR reports. Half a screen of
travel per page: the page key always jumps a whole screen, and a 1:1 mapping was
unusably slow with a discrete wheel. Shift is excluded — it keeps meaning "local
scrollback". Verified live: opt-out ON on a Claude session sends real PageUp/PageDown to
the PTY where the wheel previously did nothing.
3. **Probe caching** — `getClaudeCliVersion()` no longer caches failure. Success is kept for
the process lifetime; a failed probe retries with a 1/2/4…15min backoff. The cache policy
is a pure function (`resolveClaudeCliVersion`) so the retry semantics are unit-testable
without spawning `claude`. The VITEST short-circuit now records nothing, where before it
wrote a permanent null.
4. **Opt-out footgun** — handled by pairing rather than by scoping: the setting keeps meaning
exactly what it says (the wheel goes local), and fix 2 catches the case where "local" is
empty. Scoping the setting away from repaint-mode CLIs would have silently overridden an
explicit user choice. The App Settings tooltip now says to leave it off for Claude/Codex.
5. **Diagnostic** — `_logScrollRouting()` prints one line per session per distinct decision:
`[scroll] <id> → forward-sgr|page-keys|local-scrollback|repull-refused-downgrade (mode=…,
cliVersion=…, localScrollbackOptOut=…, mouseTracking=…, localScrollbackRows=…)`. That
single line answers every open question in the list below.
Still unanswered by code alone: whether mtiller's Claude Code is genuinely older than
2.1.187 (false-path 3), and whether the iPhone was running stale JS. The diagnostic makes
both self-reporting, so the retest ask is now "open the console and paste the `[scroll]` line".
### What to get from mtiller (some already asked)
- Shift+scroll behavior on Firefox (distinguishes hollow-local from handler-not-firing).
- `claude --version` on the Mac (decides false-paths 2 vs 3).
- App Settings → Input → "Wheel scrolls local history" state (false-path 1).
- iPhone: Claude or shell session, and whether a full tab kill changes anything.
- Browser console: `app.terminalUi?.terminal?.modes?.mouseTrackingMode` (false-path 4).
## ROUND 3 (2026-08-09): Codex wheel dead — CONFIRMED AND FIXED
DodgyBadger (Codex latest, Chrome, Windows 11): mouse wheel does nothing in a CODEX session
while working fine in shell and web tabs; DRAGGING THE SCROLLBAR WORKS, so xterm's local
buffer demonstrably has content for their codex pane. Analysis against the shipped code:
- `_shouldForwardWheelToApp` returns true UNCONDITIONALLY for `codex` (no version gate, unlike
claude's `>= 2.1.187`), so every plain wheel tick is sent as SGR reports to Codex.
- The "verified to scroll its transcript on SGR wheel reports" claim for codex predates
current Codex builds; if Codex latest ignores SGR wheel, forwarding eats the gesture while
the healthy local scrollback (proven by the working scrollbar) sits unused.
- The #227 PageUp fallback cannot rescue this: it is gated to `claude` mode AND `baseY === 0`,
and codex here has real local scrollback. The `[scroll]` diagnostic will still say
`forward-sgr (mode=codex, ...)`, confirming the branch, worth asking the reporter to paste.
**CONFIRMED by the reporter's `[scroll]` line (2026-08-09, PR #227 comment)**:
`forward-sgr (mode=codex, cliVersion=unknown, localScrollbackOptOut=false, mouseTracking=none,
localScrollbackRows=967)`. Forwarding branch active, 967 rows of healthy local scrollback
unused, Codex ignoring the SGR reports. Environment: Codex latest, Chrome, Windows 11.
**Measured against codex-cli 0.147.0** (isolated `tmux -L codexwheel`, fake `CODEX_HOME/auth.json`,
history built with 401ing prompts), which settles it without needing a version gate at all:
| Probe | Result |
| ---------------------------------------------- | ----------------------------------------------- |
| `#{mouse_any_flag}` once the TUI is up | `0`: codex never enables mouse tracking |
| `#{alternate_on}` | `0`: inline viewport, not an alt-screen pager |
| `#{history_size}` while prompting | grows 3 → 32: the transcript goes to scrollback |
| 6 × `\x1b[<64;10;10M` written to the pane | pane capture byte-identical, nothing happens |
| control: literal `zz` | pane changes, so the probe can see changes |
| `\x1b[<0;12;5M` + release (the click-tap path) | no change either: taps are no-ops, not garbage |
Codex has no in-app pager to drive: its history lives in the terminal's own scrollback, which is
exactly what forwarding was stealing the gesture from. A version gate would be the wrong fix (and
`cliVersion=unknown` means there is no codex probe to gate on anyway).
**Fix (shipped):** `_shouldForwardWheelToApp` now returns true for `claude >= 2.1.187` and nothing
else. Codex falls to the normal local-scrollback path like shell/gemini/opencode, so wheel and touch
scroll the same history the scrollbar drag was already scrolling. The claude-only PageUp fallback is
untouched: codex never needs it, its local buffer is real. Taps stay hand-encoded for codex
(`_sessionUsesServerMouseStrip`), measured harmless, so click-to-position is merely unavailable
there rather than damaging. Lesson for the next mode added to the forward list: "it is a strip mode"
proves nothing, write a real SGR report into a live pane and diff the capture first.
Verified end-to-end in Chromium against a live codex session on an isolated instance
(`CODEMAN_INSTANCE=codexwheel`, port 5055, `envOverrides.CODEX_HOME` pointing at the fake auth
dir): trusted `page.mouse.wheel` up now logs
`[scroll] … → local-scrollback (mode=codex, …, localScrollbackRows=43)`, moves the viewport
39 → 4 (back to the Codex banner), and sends ZERO bytes to the PTY. Unit coverage:
`test/terminal-touch-tap.test.ts` ("only claude forwards — codex and gemini keep the local wheel").
Original plan follows.
## Reports
- **Issue #205** (https://github.com/Ark0N/Codeman/issues/205), OPEN:
- **jonocodes** (author, 2026-08-03): SHELL session. Host Mac M4, brew tmux. On Android, touch-scrolling the terminal does nothing. On desktop, the mouse wheel cycles shell command history (acts like Up/Down arrows) instead of scrolling the screen.
- **mtiller** (comment, 2026-08-06): "similar issue just with scrolling backward to see agent output. This is with Firefox on MacOS." (Claude session implied.)
- **Reddit r/selfhosted** comment `p21x6ts` by mmtiller (= mtiller on GitHub): scrolling broken enough across phone/iPad/laptop that they fall back to Claude's own remote-control feature. Churn-risk user who otherwise loves the product; fixing this has promo value beyond the bug itself.
## How scrolling works today (read this before touching anything)
Three independent paths, all in `src/web/public/terminal-ui.js` unless noted:
1. **Desktop wheel** (container `wheel` listener, ~line 421): ALWAYS `preventDefault()`s, then either
- forwards synthetic SGR wheel reports to the app (`_sendSyntheticSgrWheel`, coalesced every 40ms, fire-and-forget) when `_shouldForwardWheelToApp(ev)` (~line 2823) passes: no Shift held, opt-out setting `terminalWheelLocalScrollback` off, xterm `mouseTrackingMode === 'none'`, session mode is `claude` with `cliVersion >= 2.1.187` or `codex`, and viewport is at bottom;
- otherwise scrolls xterm's LOCAL scrollback via `terminal.scrollLines(lines)`.
- `lines` comes from `_wheelScrollLines(ev)` (~line 2818): `delta / 25`, i.e. it assumes PIXEL deltas.
- NOTE: xterm.js's own internal wheel handler sits on an element INSIDE the container, so it runs FIRST (bubble order) and is not suppressed by the container's `preventDefault`.
2. **Touch** (touchstart/move/end, ~lines 441-585): converts touch deltas to `terminal.scrollLines()` with momentum. Touch is ALWAYS local-scrollback, never forwarded to the app. Tap-to-position (touchend, ~line 533) is separate and already handles both mouse-tracking-on and server-strip cases.
3. **Server-side strip** (`_handleTerminalOutput`, `src/session.ts:1384`): for modes in `isAltScreenStripMode()` (`src/session.ts:179` = `codex | claude | gemini`), strips alt-screen switches (`?47/?1047/?1049`), scrollback erase (`3J`), and mouse-tracking DECSETs (`?1000-?1007` except `?1004` focus) so content stays in xterm's normal buffer with scrollback intact. Includes a chunk-boundary carry so split sequences can't leak. `shell` and `opencode` (and `antigravity`) are deliberately EXCLUDED: arbitrary shell programs (vim/less/htop) legitimately need the alt screen. There is a parity copy of this strip on the replay path (`src/web/routes/session-routes.ts`, ~line 1697) and a frontend parity check `_sessionUsesServerMouseStrip()` (terminal-ui.js ~line 2751). All three must stay in sync.
4. Related: full-scrollback replay (`GET .../terminal?full=1` on first buffer load) fills xterm local scrollback; client scrollback is hardcoded 50k (`DEFAULT_SCROLLBACK`, constants.js) vs tmux 100k.
## Diagnosis
### Bug A: Firefox wheel deltas (mtiller's desktop case)
`_wheelScrollLines()` divides by 25 assuming `WheelEvent.deltaY` is pixels (`deltaMode === 0`, Chrome/Safari behavior). Firefox commonly fires `deltaMode === 1` (LINE units, deltaY around 1-3 per notch), so `Math.round(3/25) = 0` and the `|| ±1` fallback yields 1 line per event. With a discrete mouse wheel that is 1 line per notch: scrolling feels dead/broken. This hits BOTH the local-scroll path and the forwarded path, since both use the same function.
**Fix**: normalize by `ev.deltaMode` in `_wheelScrollLines()`:
- `deltaMode 0` (pixels): current behavior, `delta / 25`.
- `deltaMode 1` (lines): use the delta directly (round, keep sign fallback).
- `deltaMode 2` (pages): `delta * terminal.rows` (or a sane page size).
Keep the existing Shift-axis trap intact: on macOS trackpads Shift+two-finger scroll arrives as a HORIZONTAL wheel (deltaX carries the magnitude, deltaY ~0); that's why the function reads deltaX when Shift is held (issue #154). Don't lose it.
**Verify**: don't trust this diagnosis blindly. First reproduce in real Firefox on macOS and log `deltaMode`/`deltaY` (Firefox trackpad input can arrive as pixels; external mouse as lines). Also confirm the session's `cliVersion` probe succeeded (a failed probe disables forwarding entirely, which would point elsewhere). Unit-test by dispatching synthetic `WheelEvent`s with explicit `deltaMode` values; a Playwright `firefox` project pass is the end-to-end check.
### Bug B: shell mode has NO working scrollback at all (jonocodes)
Chain: shell mode is excluded from the alt-screen strip (correctly) → tmux attaches on the alternate screen → xterm's alt buffer has zero scrollback. Consequences:
- **Wheel**: xterm's own internal wheel handler runs first and, in the alt buffer, converts wheel ticks into Up/Down arrow keys (alternateScroll behavior). The shell receives arrows → command history cycles. That is jonocodes' exact desktop symptom. The container handler's `scrollLines()` afterwards is a no-op (no scrollback in alt buffer).
- **Touch**: the touch handler's `scrollLines()` is equally a no-op → "scrolling does nothing" on Android. Exact symptom two.
- The real history exists the whole time in tmux's 100k-line buffer; nothing exposes it.
**Fix, recommended approach (a): enable tmux `mouse on` for shell sessions.**
- Server-side, set `mouse on` scoped to shell sessions' tmux sessions (`tmux set-option -t <session> mouse on` at create + on attach of recovered sessions). Do NOT set it globally on the socket: claude/codex/gemini sessions rely on the DECSET strip and must not change.
- What this buys, all natively: tmux enables mouse tracking on the outer terminal → xterm `mouseTrackingMode` goes non-none → the container handler stands down (line ~2830 check) and xterm's own encoder forwards wheel as SGR reports → tmux scrolls its OWN copy-mode history on wheel-up, auto-exits at bottom. The alt-scroll arrow conversion disappears too (tracking mode takes precedence). Desktop is fully fixed with no new endpoints.
- **Touch**: still needs one small client change: in the touchmove path, when the active session is `shell` AND `mouseTrackingMode !== 'none'`, convert accumulated lines to `_sendSyntheticSgrWheel(x, y, lines)` instead of `scrollLines()`. The 40ms coalescing already prevents the tmux process storm (each send is a tmux send-keys server-side; unbatched flicks would spawn dozens of processes: this constraint is documented at `_sendSyntheticSgrWheel`, do not bypass it).
- **Selection tradeoff to verify**: with tracking on, xterm hands drag events to tmux instead of doing local browser selection. Shift+drag still does local selection (xterm shift-override). Verify this UX on desktop before shipping; if it's unacceptable, fall back to approach (b).
- **Also verify**: vim/less/htop inside the shell still behave (they'll now receive real mouse events via tmux, generally an improvement); remote shell sessions run tmux on the REMOTE host (`tmux -L codeman-remote`) and need the same option set there if remote shells are in scope (fine to defer, note it in the changeset if skipped).
**Fallback approach (b), only if (a)'s selection tradeoff fails testing**: keep mouse off; when a shell session is in the alt buffer, have the client send scroll intents to a small server endpoint that drives `tmux copy-mode -e -t <pane>` + `send-keys -X -N <n> scroll-up/down`. Preserves selection semantics exactly, but needs a new endpoint, server-side batching, AND suppression of xterm's native alt-scroll arrow conversion (capture-phase wheel listener with `stopPropagation`, or `attachCustomWheelEventHandler` if the vendored xterm version has it). More moving parts; (a) should be tried first.
**Not acceptable**: adding `shell` to `isAltScreenStripMode()`. vim/less/htop need the alt screen; that exclusion is deliberate and documented.
### Bug C: mtiller's phone/iPad case — UNREPRODUCED, do not guess
Touch is always-local by design, and Claude sessions keep content in the normal buffer (strip), so touch scrollback "should" work there. Before coding anything: build a repro matrix (iPhone Safari / iPad Safari / Android Chrome × claude / shell) on the current release. Plausible candidates if it does reproduce: auto-scroll-to-bottom fighting user scrolls (`_noteTerminalUserScroll`, ~line 2004), or they were in shell sessions on mobile too (then Bug B covers it). Ask mtiller on #205 for session mode + Codeman version if the matrix comes up clean.
## Invariants the implementation MUST respect
- Shift+wheel always scrolls local scrollback; the trackpad Shift-axis handling from #154 stays.
- The `terminalWheelLocalScrollback` opt-out setting keeps working (pins plain wheel to local).
- The viewport-at-bottom gate stays: once the user scrolled up locally, wheel stays local until they return to bottom.
- 40ms SGR coalescing: never send per-event writes to the server.
- Strip parity triangle: `session.ts` live strip ↔ `session-routes.ts` replay strip ↔ `_sessionUsesServerMouseStrip()` in the frontend. If you touch mode lists, update all three.
- Don't add `opencode`/`antigravity` to any strip/forward list; their TUI wheel behavior is unverified (documented at `_shouldForwardWheelToApp`).
- The chunk-boundary sequence carry in `_handleTerminalOutput` must not be weakened.
## Testing (per repo rules)
- `npm test -- test/<file>.test.ts` only; never bare `npm test`. New test ports 3150+, never 3000.
- Browser-test traps (documented in CLAUDE.md Testing): drive input/scroll through real events (`page.mouse.wheel`, real touch), not app internals; headless Chromium reports `isTouchDevice()` false even with `hasTouch: true`; assert on real state (xterm viewport position, `tmux -L codeman capture-pane`), not HTTP 200.
- Shell-mode E2E: create a throwaway shell session, `seq 1 500`, then (1) wheel up on desktop shows earlier lines, not history cycling; (2) touch-scroll on a phone shows earlier lines; (3) `vim` + `less` still enter/leave the alt screen cleanly; (4) Shift+drag still selects text.
- Firefox E2E: Playwright `firefox` project, wheel over a Claude session's finished output, assert viewport moved more than 1 line per notch.
- End-to-end against the REAL environment before claiming done (standing user rule). w1/w2/w3 tmux sessions are the user's live sessions: never send input to them; create your own throwaway session and DELETE it by exact id when done.
## Related observation (not a reported bug, worth a look while in there)
The `claude --version` probe that feeds the forwarding gate runs only for local and docker sessions (`src/session.ts:1490` gates `!this._remote`; docker handled at :1507). Remote Claude sessions therefore never get `cliVersion` and silently keep local-only wheel. Harmless (local scrollback works) but inconsistent; cheap to fix by probing over ssh, or document as intended.
## Rollout
1. Bug A (deltaMode) is small and independent: can ship alone as a patch.
2. Bug B (shell scrollback) is the headline fix for #205: patch or minor per COM flow.
3. After deploy + verification: comment on #205 (what was fixed, what needs their retest), then reply to the Reddit comment `p21x6ts` with the release version. Both reporters gave environment details; address them specifically.
+255
View File
@@ -0,0 +1,255 @@
# Scrollback issues: analysis and test evidence
Covers GitHub issue **#205** ("Scrollback in terminal not working", jonocodes, shell mode,
Android + macOS desktop) and the follow-up comment on it from **mtiller** (Firefox on macOS,
"scrolling backward to see agent output"). Related closed issue: **#154** (fixed in 1.3.3).
Status: **analysis only, nothing implemented.** Measured against the live 1.11.2 instance on
2026-08-06 with throwaway `zz-*` shell sessions (all deleted afterwards; the user's `w*`
sessions were never touched).
---
## TL;DR
Five distinct problems, not one. #205 is fully explained by finding 1; findings 2 and 3 are
independent and hit **every** mode including Claude, and are the likely substance of the
"similar issue" follow-up.
| # | Problem | Modes affected | Severity | Confirmed |
| - | ------- | -------------- | -------- | --------- |
| 1 | xterm parked in the **alternate buffer** for the whole session, so there is no scrollback at all and the wheel is translated into Up/Down arrow keys | `shell`, `opencode`, `antigravity` | High | Reproduced end to end |
| 2 | **Bursty output silently destroys a screenful** of the browser's scrollback and adds ~1 row | all | High | Measured |
| 3 | **Tab switch collapses scrollback** to roughly one screen (`full=1` fires once per page load) | all | Medium | Measured |
| 4 | `deltaMode` is never read, so Firefox scrolls ~4x slower per notch | all, Firefox | Low | Static, needs reporter data |
| 5 | **Remote SSH Claude cases get no `claude --version` probe**, so wheel forwarding silently stays off (residual #154) | `claude` + remote | Medium | Static |
---
## Finding 1: shell / opencode / antigravity are stuck in xterm's alternate buffer
### Root cause
The local tmux **client** (the `tmux attach` that node-pty spawns) emits `smcup` as its very
first bytes on attach. Captured from a real PTY:
```
b'\x1b[?1049h\x1b[22;0;0t\x1b[?1h\x1b=\x1b[H\x1b[2J\x1b[?12l\x1b[?25h\x1b[?1000l...'
^^^^^^^^^^ enter alternate screen ^^^^^ application cursor keys ON
```
`Session._handleTerminalOutput()` strips `\x1b[?1049h` from the live stream, but only when
`isAltScreenStripMode(mode)` is true, and that is `claude | codex | gemini` only
(`src/session.ts:179`). For `shell`, `opencode` and `antigravity` the sequence reaches the
browser verbatim and xterm switches to the alternate buffer, where:
1. `buffer.active.type === 'alternate'` and `baseY` is pinned at 0, so there is **no
scrollback to reach**. `terminal.scrollLines()` is a no-op, which is why touch scrolling
on Android "does nothing".
2. xterm's own wheel listener takes over. From the vendored bundle
(`src/web/public/vendor/xterm.min.js`):
```js
if (!this.buffer.hasScrollback) {
if (ev.deltaY === 0) return false;
if (coreMouseService.consumeWheelEvent(...) === 0) return this.cancel(ev, true);
const seq = ESC + (decPrivateModes.applicationCursorKeys ? 'O' : '[') + (ev.deltaY < 0 ? 'A' : 'B');
coreService.triggerDataEvent(seq, true);
return this.cancel(ev, true);
}
```
tmux also set `\x1b[?1h`, so the emitted sequence is `\x1bOA`, i.e. **Up arrow**, straight
into the shell's readline. That is exactly the reported "the mouse wheel scrolls back
through previous commands, like pressing up".
3. `cancel(ev, true)` calls `preventDefault()` **and `stopPropagation()`**, and xterm's
listener sits on `terminal.element` (a child of Codeman's container). So Codeman's own
container wheel handler, `_shouldForwardWheelToApp` and `_wheelScrollLines` included, is
**never reached** for these modes. That whole path is dead code for shell.
### Reproduction (live instance, real browser)
Create a shell session with the page already open, print 150 lines, then dispatch 8 wheel-up
events over `.xterm-screen`:
```
t+1500 after shell start {"type":"alternate","length":35,"baseY":0}
t+3000 after shell start {"type":"alternate","length":35,"baseY":0}
after 150 live lines {"type":"alternate","length":35,"baseY":0}
WHEEL on live shell: {"ptyBytes":["OA","OA","OA","OA",
"OA","OA","OA","OA"],
"before":0,"after":0,"type":"alternate"}
```
Both reported symptoms, one root cause.
### Why it looks intermittent
The alternate-screen sequence only ever reaches the browser through the **live stream at
attach**. Neither replay path carries it:
- `?full=1` returns `capture-pane` output (`source: mux-full-history`), verified 0 hits for
`\x1b[?1049h`.
- `?tail=` returns the visible pane frame (`source: mux-visible`), also 0 hits; the shell byte
buffer was empty in every probe.
- `_resetTerminalForReplay()` calls `terminal.reset()`, which returns xterm to the normal
buffer.
So: watching a shell from creation leaves you in the alternate buffer until you reload or
switch tabs, at which point it silently starts working again. Then the next PTY attach (a
restart, or the auto-reattach in `selectSession()`) puts you back.
### Is stripping safe for shell? Probably yes when tmux-backed, and the current code comment is wrong about why
`src/session.ts:1404` says *"shell must keep the alt screen for vim/less/htop"*. For a
**tmux-backed** shell that reasoning does not hold: tmux is a full terminal emulator and never
forwards a pane's alternate-screen toggles to its client, it repaints instead. Measured per
phase on a real attach:
| phase | bytes | `?1049h` | `?1049l` | `?47/1047` |
| ----- | ----: | -------: | -------: | ---------: |
| attach | 772 | **1** | 0 | 0 |
| `seq 1 60` echo | 1402 | 0 | 0 | 0 |
| `less` open / end / quit | 284 / 230 / 321 | 0 | 0 | 0 |
| `vim` open / quit | 2200 / 646 | 0 | 0 | 0 |
`vim` and `less` inside tmux emit **zero** alternate-screen sequences to the client.
The caveat that does matter: `startShell()` falls back to a **direct PTY with no tmux** when
mux creation fails (`src/session.ts:1961`, `this._useMux = false`). In that path the inner
app's own `?1049h` does reach xterm, and a blanket strip would break vim/less/htop for real.
Any fix has to be conditional on `_useMux`, which is known server-side.
Second caveat: stripping alone buys less than it looks like, because of finding 2. It fixes
the wheel (no more phantom Up arrows) and it makes the `full=1` replay reachable, but live
output still will not accumulate.
---
## Finding 2: bursty output silently overwrites a screenful of browser scrollback
Independent of the alternate buffer, and it hits Claude sessions too.
tmux decides per flush whether to emit real linefeeds (which push rows into the outer
terminal's scrollback) or to repaint the pane rectangle with cursor addressing (which
overwrites the visible rows in place). When output outpaces its flush interval it coalesces
into a repaint, and one screenful of the browser's history is **destroyed**.
Measured on one session, same page, `rows = 36`:
| step | `baseY` | rows containing SEED | BURST | SLOW |
| ---- | ------: | -------------------: | ----: | ---: |
| after `?full=1` replay (120 seeded lines) | 86 | 120 | 0 | 0 |
| after 60 lines emitted as fast as possible | **87** (+1) | **86** (-34) | 35 | 0 |
| after 60 lines at ~16/s (`sleep 0.06`) | **148** (+61) | 86 | 35 | 60 |
The burst added **one** row of scrollback and ate **34** rows of existing history. The slow
run behaved correctly. So "I printed a bunch of lines and now I cannot scroll back" reproduces
without the alternate buffer being involved at all, and it is rate dependent, which is exactly
the kind of thing that reads as random flakiness.
Consequence: the browser's scrollback is effectively frozen at whatever the last `?full=1`
replay produced, minus a screen per burst. tmux's own history is fine throughout
(`history_size` kept growing, `history-limit` 2000), so the data is never actually lost
server-side, it just never reaches the browser again until a reload.
---
## Finding 3: switching tabs collapses a session's scrollback
`_initialFullBufferLoad` is true for the **first buffer load after a page load only**
(`app.js:4374`). Everything after that uses `?tail=`, which returns byte history plus the
visible pane frame. Worse, the snapshot restore path deliberately throws away the restored
xterm snapshot (which does carry scrollback) and replaces it with that frame
(`app.js:4316-4328` plus `needsRewrite`).
Measured, switching away from session A and back:
```
A: initial full=1 load {"len":152,"baseY":116,"AAA":150}
A: after switch away and back {"len": 87,"baseY": 51,"AAA": 59}
```
150 lines of history down to 59. Note also that the page's single `full=1` is consumed by
whichever session auto-selects at load, so **every other tab starts life with one frame of
history**.
---
## Finding 4: `deltaMode` is never read (Firefox)
`grep -rn "deltaMode" src/web/public packages` returns nothing. `_wheelScrollLines()`
(`terminal-ui.js:2818`) treats `deltaY` as pixels unconditionally:
```js
return Math.round(delta / 25) || (delta > 0 ? 1 : -1);
```
Chrome/WebKit report `deltaMode: 0` with `deltaY` around 100 to 120 px per notch, so about 4
to 5 lines. Firefox reports `deltaMode: 1` (`DOM_DELTA_LINE`) with `deltaY` around 3, so
`Math.round(3/25) === 0` and the `|| ±1` fallback yields **1 line per notch**, roughly 4x
slower. In Claude mode the same value caps the forwarded SGR report at 1 tick per event
instead of 4, so the transcript crawls too.
This is sluggishness, not breakage, so it is a plausible but unproven contributor to the
mtiller report. No Firefox build is installed under `~/.cache/ms-playwright` (chromium and
webkit only), so this was not measured. Worth asking the reporter for `deltaMode` / `deltaY`
from a live wheel event before acting on it.
---
## Finding 5: remote SSH Claude cases still have no version probe
`src/session.ts:1490` deliberately skips the deterministic `claude --version` probe for
remote sessions and defers to the startup-banner scrape, which the same comment block
describes as unreliable ("newer Claude Code builds don't print the banner and resumed sessions
never show it"). That is precisely the condition #154 was filed for: `cliVersion` empty means
`_shouldForwardWheelToApp()` returns false, wheel forwarding is off, and the user is left with
local scrollback that (per finding 2) does not accumulate.
Local and Docker Claude sessions are fine; verified all 7 live sessions report
`cliVersion=2.1.223`, so the 1.3.3 fix is still working there.
---
## Candidate directions (not decided)
Roughly in order of value per unit of risk.
1. **Extend the alternate-screen strip to tmux-backed `shell` / `opencode` / `antigravity`.**
Gate on `_useMux` so the direct-PTY fallback keeps vim/less/htop working. Kills the phantom
Up arrows and makes replayed history reachable. `isAltScreenStripMode()` currently takes
only `mode`, so it would need the mux flag threaded in, and
`test/claude-scrollback-strip.test.ts:16-17` plus `test/antigravity-mode.test.ts:116` pin
the current answers and would need updating.
2. **Re-pull `?full=1` when the user scrolls to the top of the buffer.** Directly addresses
findings 2 and 3 with machinery that already exists and is already proven to return
complete history (200/200 lines in the probe). Needs a guard against refetch storms.
3. **Stop discarding the xterm snapshot on tab switch**, or request `full=1` on the first load
per session rather than per page. Cheaper partial fix for finding 3 alone.
4. **Read `ev.deltaMode`** in `_wheelScrollLines()` and normalise line/page deltas to lines.
Small, self-contained, worth doing regardless of whether it is mtiller's actual bug.
5. **Probe the CLI version over SSH for remote Claude cases**, mirroring the deferred
in-container probe that Docker cases already use.
Option 1 alone does not fix #205's "print a bunch of lines then scroll" complaint; that needs
2 as well.
## Reproduction assets
Scripts used, in the session scratchpad
(`/tmp/claude-1000/-home-arkon-default-claudeman/597ffc9f-.../scratchpad/`):
- `ptycap.py` / `ptycap2.py`: PTY-level capture of the tmux client stream, per phase counts of
alternate-screen and mouse-tracking sequences.
- `sim.mjs`: replays a captured stream through `@xterm/headless` with and without the strip.
- `browser-test*.mjs`: Playwright against the live instance, reports `buffer.active.type`,
`baseY`, row content and the exact bytes xterm sends to the PTY on a wheel event.
`@xterm/headless` was installed with `npm i --no-save`, so `package.json` and the lockfile are
untouched.
+2 -2
View File
@@ -312,7 +312,7 @@ TOCTOU window.
| Route | Cap | Notes |
|-------|-----|-------|
| `file-content` | 10 MB | text preview |
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses** |
| `file-raw` | 50 MB | inline MIME map; **`X-Content-Type-Options: nosniff` on all responses**; streamed, `Range`-aware (206 slices come from the same validated path, and the cap is checked before the range) |
| `POST /api/download` | 50 MB | forced `attachment`; sensitive‑path blocklist |
### SVG / content‑type XSS
@@ -489,7 +489,7 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini`, `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini` — which also carries Antigravity's `antigravity-cli/` state — `~/.config/{gcloud,opencode}`, and five seeded files from `~/.pi/agent`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
+246
View File
@@ -0,0 +1,246 @@
# Session lineage lines (spawn lines between tabs)
**Goal:** when a session spawns another session (the `codeman` agent skill starting a
worker, or anything else that says who it is), draw the same kind of glowing connection
line the subagent windows already use, but **tab → tab**, so a glance at the strip shows
which tab spawned which.
Status: PLAN. Nothing implemented yet.
---
## 1. The blocking fact: no parent relationship exists today
There is no spawn-parent link between sessions anywhere in the codebase:
- `SessionState` (`src/types/session.ts:388`) has no `parentSessionId` / `spawnedBy` /
`createdBy`.
- `POST /api/quick-start` and `POST /api/sessions` record only `owner = ownerFor(req)`,
which is the multi-user **human**, not the calling session.
- The only parent links that do exist are `TeamConfig.leadSessionId` (agent teams) and
`subagent-parents.json` (a frontend **window-layout** store for subagent windows).
Neither says "session A spawned session B".
- Nothing in the HTTP request identifies the caller: an agent's spawn call is plain
`curl` from inside a tmux pane, so there is no socket-level identity to recover
(`SO_PEERCRED` needs a unix socket; the API is TCP).
So the caller has to **tell** us. It already knows its own id: every managed pane gets
`CODEMAN_SESSION_ID` exported by `session-cli-builder.ts` (and the skill's §0 preamble
already binds it to `$SELF`).
## 2. Wire format
Two ways in, because they serve different callers. Body wins when both are present.
| Where | Shape | Who uses it |
| --- | --- | --- |
| body field | `"parentSessionId": "<uuid>"` | anything hand-writing one create call |
| request header | `X-Codeman-Parent-Session: <uuid>` | the skill: added **once** to the `CURL` array in the §0 preamble, so every present and future create call carries it with no per-recipe edit |
Rules, all of them deliberate:
- **Advisory decoration only.** It never grants access, never scopes anything, never
affects lifecycle. A child is not killed when its parent dies; the line just stops
being drawn once the parent tab is gone.
- **Never fails a spawn.** An unknown / stale / foreign parent id is silently dropped
(field ends up `undefined`), not a `400`. A cosmetic field must not be able to break
worker creation.
- **Resolved, not trusted.** The id must match a live session the caller can already
see (`canAccessOwned`), and the resolved parent's `owner` must equal the new
session's `owner`. Otherwise a user could staple their session under another user's
tab in multi-user mode.
- Exact id match first; a `>= 8`-char **unique** prefix match as a fallback (ids appear
truncated in mux names and UI surfaces; ambiguous prefixes resolve to nothing).
## 3. Server changes
| File | Change |
| --- | --- |
| `src/types/session.ts` | `SessionState.parentSessionId?: string` with a doc comment saying it is UI decoration and never a permission signal |
| `src/session.ts` | constructor option `parentSessionId` → `_parentSessionId`, public getter, emitted from `toState()` (~line 1170) |
| `src/web/schemas.ts` | `parentSessionId: z.string().max(100).optional()` on `CreateSessionSchema` (272) and `QuickStartSchema` (680). Neither is `.strict()`, so this is additive |
| `src/web/route-helpers.ts` | new `resolveParentSessionId(ctx, req, bodyValue, owner)` implementing §2's rules; returns `string \| undefined`, never throws |
| `src/web/routes/session-routes.ts` | pass it into the three `new Session({...})` sites: `POST /api/sessions` (846), `POST /api/run` (2522), `POST /api/quick-start` (2896) |
| `src/web/server.ts` | recovery path (~2617): `parentSessionId: savedState?.parentSessionId` so the link survives a restart |
**No new SSE event.** `session_created` / `session_updated` broadcast
`getSessionStateWithRespawn(session)`, which is `toState()`-derived, so the field rides
along to the browser for free — and the frontend already does
`this.sessions.set(data.id, data)`, so `session.parentSessionId` is simply there.
Optional follow-up: surface it on `/api/sessions/unified` rows so the Session Manager
and the home rails can show "spawned by w3-claudeman".
## 4. Frontend rendering
### 4.1 Where the code goes
`_updateConnectionLinesImmediate()` (`subagent-windows.js:242`) is a strict
**batched read → batched write** pass, and it already has an extension point:
ultracode appends its own layer via `_appendUltracodeConnectionLines(svg, rects)` at
the end, sharing the `rects` cache so no layer forces a second reflow.
Lineage lines follow that exactly: a new module `src/web/public/session-lineage.js`
(load order 15.6, after `ultracode-windows.js`) exporting
`_appendLineageConnectionLines(svg, rects)` onto `CodemanApp.prototype`, called from the
same tail. **The core function keeps ownership of the read/write split**; the new layer
only reads through the shared `rects` map and only appends paths.
The path math itself lives in `constants.js` as a pure
`computeLineagePath(parentRect, childRect, stripRect, depth)` — same treatment as
`computeTabScrollLeft`, so the geometry is unit-testable without a browser.
### 4.2 Geometry
Both endpoints are tabs in one horizontal strip, so the subagent shape (tab-bottom →
window-top) does not apply. **One case**, a **U-bridge hanging below the strip** that
touches both tabs on their bottom edge:
```
y0 = max(parent.bottom, child.bottom)
d = clamp(14 + |x2 - x1| * 0.085, 22, 104) + depth * 8 + |child.bottom - parent.bottom|
path: M x1 parent.bottom C x1 y0+d, x2 y0+d, x2 child.bottom
```
`depth` is the child's index among its siblings, so several children of one parent
**nest** instead of overprinting.
> **Superseded (2026-08-14): the two shapes this section used to specify.** The dip was
> `clamp(14 + span * 0.06, 16, 44) + depth * 6`, and a wrapped strip
> (`tabs-two-rows` / `tabs-auto-wrap`) got its own parent-bottom → child-**top** bezier.
> Both were tuned against two tabs side by side and failed at the distances the feature
> is used at:
>
> - a skill worker is appended to the **end** of the strip, so the real span is
> 800-1500px, where a 44px cap is a 33px sag, i.e. a line that reads as straight and
> crosses the terminal instead of bracketing under the strip;
> - and when the strip wraps, parent-bottom (34) to child-top (48) leaves **14px** to
> bend in, so the arc was a flat line hidden in the row gap, with siblings drawn on
> top of each other. Reported as *"they connect already, but the lines are straight
> and not easy visible"*.
>
> Anchoring both ends at the tab bottoms and hanging the control points below the
> **lower** row gives the wrapped case the same bracket as the flat one, and removes the
> branch. Pinned by `test/session-lineage-lines.test.ts`.
A small `<circle r="3.5">` at the child end marks direction (it breathes to 4.5 while that worker is busy) (an SVG `marker` would need a
`<defs>` block and fights `stroke-dasharray`).
Each path gets `class="connection-line lineage-line"`, `data-parent-tab`,
`data-child-tab`, and `data-agent-id="lineage:<childId>"` — that last one is what makes
the existing entrance machinery (`markConnectionLineEntering` / `_applyLineEntrances`,
keyed on `data-agent-id`) work on these lines with **zero** new animation code,
including the negative-`animation-delay` resume across the `svg.innerHTML = ''` rebuild.
### 4.3 Clipping
`.session-tabs` is `overflow-x: auto`, so a tab scrolled out of the strip still has a
rect — one that lies outside the strip box and would draw an arc across the logo or the
header buttons. **Skip any edge whose parent or child center falls outside
`stripRect` (4px tolerance).** Skipping is honest; clamping would draw a line to a tab
that is not there.
### 4.4 Redraw triggers
`updateConnectionLines()` already coalesces through `scheduleBackground`, so extra
callers are cheap. Needed:
- `_fullRenderSessionTabs()` — already calls it (app.js:3912). Free.
- `_renderSessionTabsImmediate()` — does **not**. A badge appearing widens a tab and
moves every tab after it, which slides the arcs off their anchors. Add the call,
guarded on `this._lineageEdgeCount > 0` so nobody pays for it without the feature.
- **strip `scroll`** (passive listener on `#sessionTabs`) — the arcs must track the
scroller. This is new; no existing line layer needed it.
- window `resize` — piggyback the throttled handler in `terminal-ui.js:930`.
- `_onSessionCreated` — `markConnectionLineEntering('lineage:' + data.id)` so a new
child draws in **if** the user has a line-entrance theme on (all entrance styles are
`legacy`/off by default, so this is a no-op for an untouched install).
### 4.5 Styling
`.connection-line.lineage-line`: blue stroke from the per-skin `--session-blue` token
(violet until 2026-08-14, changed because it lost contrast against the terminal's own
dim foreground the moment the arc crossed text),
`stroke-width: 2.5`, `dasharray 5 5`, `opacity: .72` (`.95` while the child works),
softer than the subagent lines so the two layers still read as different things now that
hue no longer separates them (shape does most of that work: a lineage arc hangs under the
strip and never reaches a window), but the contrast against the terminal comes from a
**second, wider glow** rather than more weight, because the first
cut (2px / `4 4` / `.55` / one 5px glow) disappeared into terminal text on a real 1080p
desktop. `lineage-flow` marches by two dash cycles, so it moves with the dash array
(`5 5` → `-20`). Trap to respect: the skin block nests under
`html:not([data-skin="og"])`, so a bare `.lineage-line` rule inside it would outrank the
base rule at higher specificity. **Define the color as a token per skin, keep exactly
one `.lineage-line` rule.** Light skins get a darker stroke.
Optional signal worth having: `.lineage-line--working` (a slow `stroke-dashoffset`
march) only while the **child** session is working, wrapped in
`prefers-reduced-motion: no-preference`. Static otherwise — a permanently marching line
per tab pair is noise and battery.
### 4.6 Desktop only, and why
The SVG overlay is `z-index: 999`. On desktop the header is `z-index: 100`, so arcs
paint **over** the header and can touch tab bottoms. Under 1024px `mobile.css` makes the
header `position: fixed; z-index: 1200`, which would **bury** the arcs — and the phone
strip is a scroller where both endpoints are rarely on screen together anyway. So the
layer returns early unless `MobileDetection.getDeviceType() === 'desktop'`.
Raising the SVG to ~1250 (above the fixed header, below modals at 1300) is a possible
phase 2, but it needs a real check against the mobile overview and the drawer.
### 4.7 Setting
`sessionLineageLines`, **per-device** — so it goes in the `displayKeys` set in
`settings-ui.js` and must **not** be added to `SettingsUpdateSchema` (`.strict()`;
sending an undeclared key fails the whole PUT). Rendered as a switch in
App Settings → Appearance, beside the entrance-animation pickers.
**Default: ON for desktop** (phones never render it). This is the one deliberate
departure from the "new visual surfaces ship OFF" convention — the feature is the
request, and a user with 12 unrelated tabs has a one-click off switch. Flag for the
owner if the convention should win instead.
## 5. Optional extras (call them separately, none are required)
1. **Order children after their parent** in `sessionOrder` on create, so arcs stay short
and the strip reads as a tree. Real cost: it renumbers the Alt+N badges and moves
tabs under the user's cursor, so it should be its own toggle, default OFF.
2. **Lineage hover focus**: hovering a tab dims unrelated arcs and brightens its own
subtree.
3. **"Spawned by" in the Session Manager / home rails**, once `parentSessionId` is on
the unified rows.
4. **Inherited tab tint**: children pick up a faded version of the parent's tab color.
## 6. Tests
- `test/session-lineage.test.ts` (route-level, `app.inject`): round-trips through
`POST /api/sessions` + `POST /api/quick-start`, header path, body-wins-over-header,
unknown id dropped without failing the spawn, cross-owner parent dropped in
multi-user, field present in `GET /api/sessions` and persisted state.
- `test/session-lineage-lines.test.ts` (jsdom, pure): `computeLineagePath` — same-row U,
wrapped-row bezier, sibling nesting depth, off-strip skip, degenerate zero-width rects.
- Browser check (not in `test:ci`): spawn two workers with a parent, assert two
`path.lineage-line` elements anchored to the right tabs, then scroll the strip and
assert they moved.
- Existing guards that must stay green: `test/mobile-header-buttons-policy.test.ts`
(nothing new on phones), `test/app-settings-structure.test.ts` (the new switch pairs
with its rail section).
## 7. Skill side (owned by the release session, not this plan)
One line in the `codeman` skill's §0 preamble covers every spawn recipe:
```bash
CURL=(curl -sk "${AUTH[@]}" -H "X-Codeman-Parent-Session: $SELF")
```
plus a `CODEMAN_PREAMBLE` version bump so stale cached preambles fail loudly instead of
silently spawning unparented workers. Recipes that build a create payload by hand can
alternatively send `"parentSessionId":"'"$SELF"'"`.
## 8. Docs to update when it lands
`CLAUDE.md` (a Key Patterns bullet), `docs/architecture-invariants.md` (new anchor: the
resolve-don't-trust rule, the desktop-only z-index reason, the shared `rects` pass),
`docs/api-reference.md` (the new field + header on the create endpoints).
+10 -28
View File
@@ -43,38 +43,22 @@ const syncData = DEC_SYNC_START + data + DEC_SYNC_END;
this.broadcast('session:terminal', { id: sessionId, data: syncData });
```
## Client-Side Implementation (`app.js`)
## Client-Side Implementation (`terminal-ui.js`)
### `batchTerminalWrite(data)`
1. Checks if flicker filter is enabled (optional, per-session)
2. If flicker filter active: buffers screen-clear patterns (`ESC[2J`, `ESC[H ESC[J`, `ESC[nA`)
3. Accumulates data in `pendingWrites`
4. Schedules `requestAnimationFrame` if not already scheduled
5. On rAF callback: checks for incomplete sync blocks (start without end)
6. If incomplete: waits up to 50ms via `syncWaitTimeout`
7. Calls `flushPendingWrites()` when complete
### `extractSyncSegments(data)`
- Parses DEC 2026 markers, returns array of content segments
- Content before sync blocks returned as-is
- Content inside sync blocks returned without markers
- Incomplete blocks (start without end) returned with marker for next chunk
4. Calls `_scheduleTerminalWriteFlush()` if no flush is pending
5. The yielded callback clears its scheduled flag before calling `flushPendingWrites()`
6. Large batches schedule their own next chunk until the queue is empty
### `flushPendingWrites()`
```javascript
const segments = extractSyncSegments(this.pendingWrites);
this.pendingWrites = ''; // Clear before writing
for (const segment of segments) {
if (segment && !segment.startsWith(DEC_SYNC_START)) {
terminal.write(segment); // Skip incomplete blocks (start with marker)
}
}
```
Note: Segments starting with `DEC_SYNC_START` are incomplete blocks awaiting more data. These are skipped (discarded if timeout forces flush).
- Joins the queued terminal data and passes DEC 2026 markers through to xterm.js 6, which handles synchronized output natively.
- Writes at most 32KB per yield for Codex and 64KB for other modes.
- Requeues the remainder and immediately schedules another safe yield. A final large response therefore drains without waiting for another SSE event.
### `chunkedTerminalWrite(buffer, chunkSize=128KB)`
@@ -116,17 +100,15 @@ When detected, buffers 50ms of subsequent output before flushing atomically.
## Edge Cases
- **Incomplete sync blocks**: 50ms timeout forces flush (content discarded to prevent freeze)
- **Incomplete sync blocks**: xterm.js retains synchronized output until its closing marker
- **Large buffers**: Chunked writing prevents UI freeze
- **Server shutdown**: Skips batching via `_isStopping` flag
- **Session switch**: Clears flicker filter state, pending writes, and sync timeout (prevents cross-session data bleed)
- **SSE reconnect**: `handleInit()` clears all pending write state
**Trade-off:** If a sync block is split across SSE packets and the end marker doesn't arrive within 50ms, the incomplete content is discarded. This prioritizes responsiveness over completeness. In practice this is rare since the server always sends complete `SYNC_START...SYNC_END` pairs and SSE typically delivers them atomically.
## DEC Mode 2026 Compatibility
Terminals that natively support DEC 2026 will buffer and render atomically. Terminals that don't support it ignore the escape sequences harmlessly. xterm.js doesn't support DEC 2026 natively, so the client implements its own buffering by parsing the markers.
Terminals that natively support DEC 2026 buffer and render atomically. Codeman uses xterm.js 6, so the client passes the markers through instead of parsing or discarding partial blocks.
**Supporting terminals:** WezTerm, Kitty, Ghostty, iTerm2 3.5+, Windows Terminal, VSCode terminal
@@ -135,4 +117,4 @@ Terminals that natively support DEC 2026 will buffer and render atomically. Term
| File | Key Functions |
|------|---------------|
| `src/web/server.ts` | `batchTerminalData()`, `flushTerminalBatches()`, `broadcast()` |
| `src/web/public/app.js` | `batchTerminalWrite()`, `extractSyncSegments()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
| `src/web/public/terminal-ui.js` | `batchTerminalWrite()`, `_scheduleTerminalWriteFlush()`, `flushPendingWrites()`, `flushFlickerBuffer()`, `chunkedTerminalWrite()` |
+303
View File
@@ -0,0 +1,303 @@
# Terminal smart copy (Ctrl+C) plan
Issue: [#211](https://github.com/Ark0N/Codeman/issues/211) "Terminal: Ctrl+C should copy when text is selected (interrupt otherwise)".
Origin: r/selfhosted feedback, "Biggest stumbling block is apparent lack of copy-paste in the terminal."
Status: **implemented and shipped** on 2026-08-05 (this document is kept as the rationale record). It was first served as an isolated beta over Tailscale for manual sign-off, then landed. Section 2 is the research that shaped the design, sections 4 to 6 describe what was built.
---
## 1. What the issue asks for
- Text selected in the terminal + `Ctrl+C` -> copy the selection, toast, clear the selection, do NOT send the byte to the PTY.
- No selection + `Ctrl+C` -> unchanged, the interrupt (`0x03`) reaches the PTY.
- `Ctrl+Shift+C` as an explicit copy chord.
- The selection check must run before the shortcut registry dispatch so a rebind cannot cost the user their interrupt key.
- Paste is out of scope (it already works via `Ctrl+V`, which terminal-ui.js routes to the image/text paste trap).
## 2. Verified current behavior
### 2.1 xterm cancels the Ctrl+C keydown, so no copy can happen
`src/web/public/vendor/xterm.min.js` (xterm 6.x), `_keyDown`:
```js
_keyDown(x){ if(this._keyDownHandled=!1, this._keyDownSeen=!0,
this._customKeyEventHandler && this._customKeyEventHandler(x)===!1) return !1;
... evaluateKeyboardEvent(...) ... this.cancel(x) ... }
```
Two consequences that shape the design:
1. The custom handler runs **first**, before xterm evaluates the key. Returning `false` exits before `cancel(x)`, so returning `false` does **not** call `preventDefault()` for us.
2. When the handler returns `true`, xterm turns Ctrl+C into `0x03` and cancels the event, which is why the browser's own copy command never runs.
Probe (headless chromium against an isolated server on port 3174, selection active, real focus on `.xterm-helper-textarea`, synthetic Ctrl+C keydown):
```json
{ "hasSelection": true, "defaultPrevented": true, "dataSeen": ["\"\\u0003\""],
"clipboardAfter": "SENTINEL-BEFORE", "stillHasSelection": false }
```
So today: interrupt byte sent, clipboard untouched, and xterm drops the selection anyway. The last point matters, "copy then clear the selection" is not a behavior change in how the selection feels, it is what already happens on any keypress.
### 2.2 Why right-click Copy works today
xterm registers a `copy` listener on its root element that substitutes the selection text:
```js
this._register(addDisposableListener(this.element,"copy",(k=>{ this.hasSelection() && copyHandler(k,this._selectionService) })))
```
Second probe (port 3175, real `page.keyboard.press('Control+c')`, custom handler patched to return `false` for Ctrl+C without `preventDefault`):
```json
{ "dataSeen": [], "copyEvents": ["xterm-element"],
"clipboardAfter": "native-copy-probe-line\n...", "stillHasSelection": true }
```
So a "return false and let the browser copy" implementation would also work in Chromium. It is rejected below (section 3.3) because it gives no toast, does not clear the selection, and leans on per-browser behavior of the copy command when the focused element is xterm's empty helper textarea.
### 2.3 The document-level capture handler will not interfere
`setupEventListeners()` in `src/web/public/app.js:989` runs on document capture, before xterm's textarea listener. Its registry loop skips any entry whose action is not in the local `SHORTCUT_ACTIONS` map:
```js
if (shortcut.disabled || !shortcut.action) continue;
const action = SHORTCUT_ACTIONS[shortcut.action];
if (!action) continue;
```
This is exactly how `command-palette` already behaves: it is a full registry entry (rebindable and disableable in App Settings) whose dispatch happens in a dedicated, focus-aware gate rather than the generic loop. The new copy entry follows that pattern, so the capture handler falls through untouched and the terminal handler owns the decision.
### 2.4 Registry matching rules that constrain the bindings
`matchesShortcutEvent()` (`app.js:4890`):
- Ctrl and Cmd are interchangeable as the primary modifier, so a `['ctrl']` binding also matches Cmd+C on macOS. That is fine here: with a selection it copies (same result the native macOS path gives today), without one it falls through.
- Every other modifier must be declared exactly: `if (mods.includes('shift') !== !!e.shiftKey) return false`. So `Ctrl+Shift+C` needs its own binding, a plain `ctrl+c` binding will never swallow it.
- `binding.code` wins when present, otherwise `binding.key` is compared case-insensitively.
### 2.5 Where selection is actually possible
- The server strips mouse-tracking DECSETs for `claude`, `codex`, and `gemini` (`isAltScreenStripMode`, `src/session.ts:179`), which is why plain drag-select works in those tabs even though the TUI has mouse tracking on.
- `shell`, `opencode`, and `antigravity` keep mouse reporting, so xterm requires `Shift`+drag to force a selection there. Worth one line in the docs, it is not a code change.
- Touch devices deliberately disable selection entirely (`body.touch-device .terminal-container .xterm{user-select:none !important}`, `styles.css:3196`), and phones have no Ctrl key. This feature is desktop and hardware-keyboard only, with no mobile regression surface.
### 2.6 Helpers that already exist and should be reused
| Need | Existing code |
| --- | --- |
| Clipboard write with an HTTP-safe fallback | `_copyText(text)` in `app.js:1887` (Clipboard API, then hidden textarea + `execCommand`) |
| Toast | `showToast(message, type)` in `panels-ui.js:4385` |
| Translated string | `'Copied to clipboard'` already in `i18n.js:453` |
| Focus-aware chord gate to copy the shape of | `shouldOpenCommandPaletteFromShortcut(e)` in `panels-ui.js:285` |
| Buffer-wide copy (currently unreferenced) | `copyTerminal()` in `terminal-ui.js:2615` |
`_copyText` matters more than it looks: `install.sh`'s LAN option serves plain HTTP, where `navigator.clipboard` is undefined. The issue's suggested `navigator.clipboard.writeText` alone would silently do nothing for those users, the `execCommand` fallback covers them.
## 3. Design
### 3.1 Behavior
| Chord | Selection present | No selection |
| --- | --- | --- |
| `Ctrl+C` (and Cmd+C, per registry equivalence) | copy, toast, clear selection, swallow the key | fall through, xterm sends `0x03` (interrupt) |
| `Ctrl+Shift+C` | copy, toast, clear selection, swallow the key | swallow, no-op (see 3.2) |
| Shortcut disabled in App Settings | never copies, `Ctrl+C` is always the interrupt | unchanged |
| Rebound to another chord | that chord copies when a selection exists | plain `Ctrl+C` is always the interrupt |
### 3.2 Why `Ctrl+Shift+C` with no selection is swallowed rather than forwarded
Today `Ctrl+Shift+C` produces `0x03` as well (the shift is irrelevant to the control byte), so forwarding would be "no regression". But once the chord is advertised as *the explicit copy key*, letting it interrupt a running agent when the selection happens to be empty is a footgun with no upside. Swallowing costs nothing: a user who wants to interrupt has `Ctrl+C` right there.
The rule in code is "no selection and the matched chord had Shift -> swallow", not a hardcoded key check, so it stays correct under rebinds.
### 3.3 Why an explicit clipboard write rather than falling through to the native copy
Probe 2 showed the native path works in Chromium, but the explicit write is chosen because it:
- gives the "Copied to clipboard" toast, which is the discoverability half of the issue,
- clears the selection so a second `Ctrl+C` interrupts (the smart-copy contract),
- works on plain-HTTP LAN installs through `_copyText`'s `execCommand` fallback,
- does not depend on how each browser treats a copy command issued while an empty textarea has focus.
### 3.4 Why no new app setting
Per-shortcut enable/disable and rebinding already exist in App Settings -> Shortcuts and are driven by the registry. A user who wants "Ctrl+C is always interrupt" unchecks one box. Adding a `terminalSmartCopy` setting would duplicate that and would drag in the per-device vs synced decision (`displayKeys` + `.strict()` `SettingsUpdateSchema`) for no gain.
## 4. Code changes, file by file
### 4.1 `src/web/public/app.js`, registry entry
Add to `DEFAULT_SHORTCUTS` (after the `clear-terminal` entry, ~line 351) so the Terminal group stays together:
```js
{
id: 'copy-selection',
group: 'Terminal',
label: 'Copy Selection',
bindings: [
{ modifiers: ['ctrl'], key: 'c' },
{ modifiers: ['ctrl', 'shift'], key: 'C' },
],
// Dispatched by shouldCopyTerminalSelectionFromShortcut() in terminal-ui.js,
// deliberately NOT in SHORTCUT_ACTIONS: the generic capture loop always
// preventDefaults, which would cost the user the interrupt key.
action: 'copyTerminalSelection',
},
```
Match on `key`, not `code`. xterm decides what byte to emit from the produced character, so intercepting the physical `KeyC` on a layout where it does not produce "c" would diverge from what xterm would have sent.
The `action` string is required for App Settings to render the row as configurable (`configurable = !!shortcut.action && Array.isArray(shortcut.bindings)`, `settings-ui.js:2624`). Do **not** add `copyTerminalSelection` to `SHORTCUT_ACTIONS`.
### 4.2 `src/web/public/terminal-ui.js`, the gate
New prototype method, modeled on `shouldOpenCommandPaletteFromShortcut`:
```js
shouldCopyTerminalSelectionFromShortcut(ev) {
if (!ev || ev.type !== 'keydown') return false; // the handler also runs for keypress/keyup
if (!ev.ctrlKey && !ev.metaKey && !ev.altKey) return false; // hot path: plain typing exits here
const registryAvailable =
typeof this.getShortcutRegistry === 'function' && typeof this.matchesShortcutEvent === 'function';
const entry = registryAvailable
? this.getShortcutRegistry().find((s) => s.id === 'copy-selection')
: null;
if (entry) return !entry.disabled && this.matchesShortcutEvent(ev, entry);
return (ev.key || '').toLowerCase() === 'c' && !ev.altKey; // fallback for isolated harnesses
}
```
### 4.3 `src/web/public/terminal-ui.js`, the branch
Inside `attachCustomKeyEventHandler` (`terminal-ui.js:133`), after the command-palette gate and before the `Ctrl+V` branch:
```js
// Smart copy (#211): with a selection, Ctrl+C copies instead of sending ^C.
// With no selection it MUST fall through (return true, no preventDefault) or
// the interrupt key is lost. Ctrl+Shift+C is the explicit chord and never
// falls through: an "explicit copy" that interrupts the agent is a footgun.
if (this.shouldCopyTerminalSelectionFromShortcut?.(ev)) {
const selection = this.terminal.hasSelection?.() ? this.terminal.getSelection() : '';
if (selection) {
ev.preventDefault();
void this.copyTerminalSelection(selection);
return false;
}
if (ev.shiftKey) {
ev.preventDefault();
return false;
}
return true;
}
```
`preventDefault()` is explicit because returning `false` alone does not cancel the event (section 2.1), and without it the browser would run its own copy on top of ours.
### 4.4 `src/web/public/terminal-ui.js`, the copy action
```js
async copyTerminalSelection(text) {
const selection = text ?? (this.terminal.hasSelection?.() ? this.terminal.getSelection() : '');
if (!selection) return false;
const ok = await this._copyText(selection);
if (ok) {
this.terminal.clearSelection?.();
this.showToast('Copied to clipboard', 'success');
} else {
this.showToast('Failed to copy', 'error');
}
// _copyText's execCommand fallback focuses a temp textarea; restore the
// terminal (this.terminal.focus is the CJK-aware router, not xterm's raw focus).
this.terminal.focus();
return ok;
}
```
The selection text is captured **before** the first `await`, and `navigator.clipboard.writeText` is reached in the same task as the keydown, so user activation still holds.
### 4.5 `src/web/public/i18n.js`
`'Copied to clipboard'` exists. Add `'Failed to copy': '复制失败'` (the error path is new to this surface).
### 4.6 Documentation
| File | Change |
| --- | --- |
| `README.md` shortcut table (~line 648) | `\| `Ctrl/Cmd+C` \| Copy selection (interrupts when nothing is selected) \|` and a `Ctrl+Shift+C` row |
| `src/web/public/index.html` help modal, Terminal section (~line 641) | `<div><kbd>Ctrl</kbd>+<kbd>C</kbd></div><div>Copy Selection / Interrupt</div>` plus the Ctrl+Shift+C row. Keep the existing negative assertion in `help-modal-shortcuts.test.ts` in mind (it forbids `Ctrl+K`, `C` is fine) |
| `CLAUDE.md` "Keyboard shortcuts" line | add `Ctrl+C` (copy selection, else interrupt) and `Ctrl+Shift+C` |
| `docs/architecture-invariants.md` -> "Command palette and shortcut registry" | append the invariant: the no-selection path must return `true` without `preventDefault`, the branch is keydown-only, and `copyTerminalSelection` must stay out of `SHORTCUT_ACTIONS` |
The shortcut overlay (`Ctrl+?`) and App Settings -> Shortcuts are registry-driven and pick the entry up with no edit.
## 5. Edge cases and risks
| Case | Handling |
| --- | --- |
| Handler also fires for `keypress`/`keyup` | gated on `ev.type === 'keydown'`. xterm's `_keyPress` bails on ctrl combos anyway, so no stray byte |
| CJK IME composing | the existing `isComposing || keyCode === 229` guard is the first line of the handler and stays first |
| Local echo overlay has unsent `pendingText` | the copy branch returns before `onData`, so `pendingText`, flushed offsets and the durable input queue are untouched. The no-selection path is byte-identical to today, including the "control char flushes buffered text then sends `0x03`" logic at `terminal-ui.js:895` |
| Plain HTTP (LAN install) | `_copyText` falls back to `execCommand`, then focus is restored |
| Clipboard write rejected (permissions policy, no gesture) | error toast, right-click Copy still available |
| Whitespace-only or empty selection | `getSelection()` empty string is treated as "no selection", so Ctrl+C still interrupts |
| macOS Cmd+C | registry treats ctrl/meta as interchangeable, so with a selection it takes our path (same visible result as today's native copy), without one it falls through |
| Chrome/Firefox `Ctrl+Shift+C` is the devtools inspect chord | browser-level and may still toggle devtools, our copy runs regardless. Document as a caveat, `Ctrl+C` is the primary path |
| Selection in a tab whose TUI owns the mouse (`shell`/`opencode`/`antigravity`) | unchanged, `Shift`+drag selects, then Ctrl+C copies |
| Web tab (iframe dashboard) focused | xterm handler never runs, browser-native copy inside the iframe |
| Teammate/subagent terminals (`panels-ui.js:2268`, `onData` wired) | same limitation exists there, out of scope for this PR (section 8) |
## 6. Test plan
New file `test/terminal-copy-selection.test.ts` (node env, `vm` harness in the style of `test/command-palette-ui.test.ts`), covering `shouldCopyTerminalSelectionFromShortcut` in isolation:
1. Ctrl+C keydown -> true, keyup/keypress of the same chord -> false.
2. Ctrl+Shift+C -> true, plain `c` -> false, Ctrl+K -> false.
3. Registry entry `disabled: true` -> false for every chord.
4. Rebound entry (for example Alt+Y) -> true for the rebind, false for Ctrl+C.
5. Missing registry (harness without `getShortcutRegistry`) -> falls back to the `c` check.
Static assertions appended to `test/keyboard-shortcuts.test.ts` (this suite already pins the xterm-handler chokepoint):
6. `DEFAULT_SHORTCUTS` contains `id: 'copy-selection'` and `SHORTCUT_ACTIONS` does **not** contain `copyTerminalSelection` (the interrupt-safety invariant).
7. `terminal-ui.js` contains the `shouldCopyTerminalSelectionFromShortcut` branch and a `return true` no-selection fall-through.
8. README + help modal rows exist (mirrors the existing palette/Alt-nav doc assertions).
`test/help-modal-shortcuts.test.ts`: add `expectShortcut(helpModal, ['Ctrl', 'C'], 'Copy Selection')`.
New browser test `test/terminal-copy-shortcut.test.ts` (Playwright, port **3174**, free per a scan of `test/`), following `test/webgl-fallback.test.ts`: boot `WebServer`, grant `clipboard-read`/`clipboard-write`, `terminal.write()` a known line, `selectLines()`, real `page.keyboard.press('Control+c')`, then assert clipboard content, empty `onData` capture, cleared selection and the toast. Second case: no selection, assert `onData` saw `\u0003` and the clipboard is unchanged.
Per repo convention, browser suites are excluded from CI, so add the filename to the exclude list in `config/vitest.ci.config.ts` and run it locally.
Regression runs: `npm test -- test/keyboard-shortcuts.test.ts`, `test/help-modal-shortcuts.test.ts`, `test/command-palette-ui.test.ts`, `test/input-send-order.test.ts`, then `npm run test:ci`.
## 7. Manual verification before COM (CLAUDE.md rule)
Against a throwaway session on the live instance (`curl -sk https://localhost:3000/...`, never w1/w2/w3):
1. Select output with the mouse, press Ctrl+C, confirm the toast, paste elsewhere, confirm the agent did not stop.
2. Press Ctrl+C again with nothing selected, confirm the agent interrupts.
3. Type a few characters with local echo on (phone or `localEchoEnabled` forced), press Ctrl+C with no selection, confirm buffered text plus interrupt behave as before.
4. Uncheck the shortcut in App Settings -> Shortcuts, confirm Ctrl+C always interrupts even with a selection.
5. Rebind it, confirm the new chord copies and Ctrl+C reverts to pure interrupt.
6. Repeat 1 and 2 in an `opencode` or `shell` tab using Shift+drag to select.
7. Load over plain HTTP (`--host` LAN or `http://127.0.0.1:<port>`) and confirm the `execCommand` fallback copies and focus returns to the terminal.
8. Mobile smoke: confirm nothing changed (selection is CSS-disabled, no Ctrl key).
## 8. Out of scope, follow-ups worth filing separately
- **Teammate/subagent terminals** (`panels-ui.js:2268`) have the same blocked-copy problem. One `attachCustomKeyEventHandler` reusing `copyTerminalSelection` would fix them, but it touches a different surface and deserves its own change.
- **A mobile copy affordance.** Selection is disabled on touch, so phones still cannot copy terminal text. The unreferenced `copyTerminal()` (whole buffer) plus a keyboard-accessory "Copy" button would be the cheapest answer.
- **Right-click context menu** with Copy/Paste, better discoverability than any chord, but a bigger UI surface.
- **`copyTerminal()` cleanup**: it uses raw `navigator.clipboard` rather than `_copyText`, so it would fail on plain HTTP if ever wired up.
## 9. PR mechanics
- Branch off `master` (verify with `git branch --show-current`, the tree is shared), stage explicit paths only.
- Files touched: `src/web/public/app.js`, `src/web/public/terminal-ui.js`, `src/web/public/i18n.js`, `src/web/public/index.html`, `README.md`, `CLAUDE.md`, `docs/architecture-invariants.md`, `docs/terminal-copy-shortcut-plan.md`, three test files, `config/vitest.ci.config.ts`.
- `index.html`, `app.js` and `terminal-ui.js` are `.prettierignore`d hand-formatted assets, match the surrounding style by hand. `npm run check:public-assets` and `npm run check:frontend-syntax` are the guards.
- No changeset in this PR: a merged, unconsumed changeset turns the Release workflow red until the next COM, and the COM flow writes release notes covering everything since the last tag (current version is 1.10.0).
- Close #211 from the PR body.
Rough size: about 60 lines of product code, most of the work is the tests and the four documentation surfaces.
+201
View File
@@ -0,0 +1,201 @@
<!-- Design doc drafted 2026-07-28 from WWDC26 session 224 research. STATUS: PLANNED, NOT IMPLEMENTED. Blocked on macOS 27 "Golden Gate" (beta now, GA expected fall 2026). -->
# VM Cases (macOS Virtualization framework), Implementation Plan
## Status
PLANNED, nothing implemented. This is the design + phased execution plan for a native-macOS VM isolation tier for cases ("the VM subsystem"), modeled on Docker cases (`docs/docker-cases-plan.md`). Testbed prerequisite: a macOS 27 host (see Section 8).
**⚠ DESIGN DIRECTION (owner, 2026-07-29): the subsystem is GUI-first.** Users want real macOS desktops, not headless SSH machines. Guests may be macOS (GUI-only in practice) or Linux (GUI or headless). Key decision 3 below carries the full consequences; anything in this doc that reads as "Linux-first / headless-first" predates this and has been revised.
**2026-07-29: Phase 0 substantially validated on the beta testbed; full Apple-stack reference now lives in [`docs/vm-subsystem-apple-stack.md`](vm-subsystem-apple-stack.md)** (API surfaces, beta bugs, our empirical results, and design implications). Plan-relevant corrections from that work: vmnet's topology/port-forwarding APIs are macOS 26 (only the loopback fix is 27); guest provisioning is macOS-guests-only (Linux stays cloud-init, proven working); DiskImageKit has NO flatten/merge, so the `export` subcommand ships the layer chain (or flattens in-guest) instead of flattening; seed ISOs are base-build-time only, never attached at case runtime; per-case EFI variable stores are mandatory; guest health checks read DHCP leases, never serial/ping.
## 1. Context and motivation
WWDC 2026 session 224 ("Expand the Capabilities of your Virtualization App", https://developer.apple.com/videos/play/wwdc2026/224/) shipped the missing pieces for programmatic, fleet-style VM management on macOS:
- **`VZMacGuestProvisioningOptions`**: automated first-boot setup of a macOS guest (user account, auto-login, SSH enabled) with zero interactive setup.
- **DiskImageKit**: stacked disk images on the Apple Sparse Image Format (ASIF): a read-only base layer plus cheap per-VM cache/overlay layers. Direct analog of Docker image layers + writable container layer.
- **vmnet framework**: custom network topologies and port forwarding from the host process.
- **`VZCustomVirtioDevice`**: custom low-latency host<->guest channels (Linux guests).
- **AccessoryAccess**: USB passthrough (not relevant to Codeman, out of scope).
Codeman's isolation story today is Docker cases. On macOS, Docker means Docker Desktop / a Linux VM anyway, with weaker fidelity and a heavyweight dependency. The Virtualization framework gives hardware-virtualized per-case sandboxes natively, with a layered-image story that mirrors what `scripts/build-agent-image.mjs` does for Docker. This is the premium native-macOS tier ON TOP of Docker cases, never a replacement (Docker remains the cross-platform story; the Linux prod box cannot use any of this).
## 2. Platform reality (hard constraints)
| Constraint | Detail |
| ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Host OS | macOS 27 "Golden Gate" required for the new APIs (dev beta since 2026-06-08, public beta since 2026-07-13, GA expected fall 2026) |
| Host hardware | Apple Silicon only (macOS 27 dropped Intel). Testbed: the owner's dedicated MacBook (Section 8); the M4 Mac mini (macOS 26.4, runs the second Codeman install) stays on stable + untouched |
| Guest provisioning | `VZMacGuestProvisioningOptions` needs macOS 27 on BOTH host and guest. Linux guests provision via cloud-init instead |
| macOS guest concurrency | **Hard kernel cap: 2 concurrent macOS VMs per host. MEASURED on 27 beta 4 (2026-07-29), not inferred**: the 3rd VM is refused instantly with `VZErrorDomain` code 6 while 39% of RAM is free, so more hardware does NOT raise it. Since macOS GUI guests are the headline use case, this is a real product capacity limit to schedule around and surface in the UI. Linux guests are uncapped (resource-bound only) |
| Language | Virtualization framework is Swift/ObjC only; Node cannot call it. Requires a Swift helper binary (Key decision 2) |
| Entitlement | Host process needs `com.apple.security.virtualization`. Fine for a locally built dev binary; distribution needs signing thought (Section 9) |
| Nested virtualization | Linux-guest-only on M3+. A macOS 27 VM cannot dependably host its own guests, so the host-side APIs must be tested on bare-metal 27 (dual-boot) |
| CI | Cannot run in CI (needs beta macOS on Apple Silicon). Same answer as tmux/docker: no-op all VM IO under `VITEST`, unit-test the pure parts |
## 3. Goal and user stories
Add "VM cases" to Codeman: a case can point at a per-case virtual machine on a macOS host, and any CLI backend runs inside it over the existing remote-SSH session machinery. A LOCATION OVERLAY on cases, exactly like remote-SSH and Docker cases, NEVER a sixth `SessionMode`.
- As a Mac user, I link a case to a VM so an autonomous run executes behind a hardware virtualization boundary (stronger than Docker's shared kernel) while file viewing, transcripts, and hooks keep working.
- Per-case VMs are instant and cheap: a shared provisioned base image plus a per-case overlay, not a full image copy per case.
- Killing a session kills only its in-guest tmux; the VM stays up while sibling sessions remain; case delete tears the VM down.
- I export a case's VM overlay as a portable artifact (mirror of `docker-exports/`), secrets excluded.
- On a non-mac host, or a Mac without the helper, the feature is invisible: zero UI, zero probes, zero errors.
Non-goals for the MVP: USB passthrough, custom Virtio channels (Phase 3 candidate), macOS-guest fleets (capped at 2 anyway), Kubernetes-style orchestration, Intel Macs.
## 4. Architecture
```
Codeman (Node, unchanged session layer)
| JSON over stdout (same pattern as shelling out to docker/tmux)
v
codeman-vm (Swift package: CLI + per-VM GUI runner app in the console session)
| Virtualization / DiskImageKit / vmnet
v
per-case VM (macOS or Linux)
|-- GUI mode: VZVirtualMachineView in a window --> guest screen sharing --> browser (noVNC)
|-- shell: SSH on vmnet IP --> existing remote-SSH tmux machinery
^ VirtioFS: host case dir mounted at the SAME absolute path
```
Note the runner is a **GUI app in the console user's session**, not a detached daemon: a daemon-launched VM cannot render, which is fatal for macOS guests and for Linux desktop cases.
### Key decision 1: location overlay, not a mode
Identical reasoning to Docker/remote-SSH (see CLAUDE.md): the session layer, respawn, Ralph, recovery, and quick-start plumbing all stay untouched. `SessionMode` stays five-valued. State mirrors the Docker pair: `~/.codeman/vm-hosts.json` + `vm-cases.json`, new `src/vm-hosts.ts` with the storage + pure helpers split.
### Key decision 2: Swift helper CLI (`codeman-vm`)
The framework is Swift-only, so all VM work lives in a SwiftPM package (`packages/codeman-vm/`), a CLI with a stable JSON contract:
- `create-base --guest linux|macos`: build the shared base image. Linux: boot an arm64 cloud image with EFI + cloud-init, install Node 22 + tmux + the four CLIs (same inventory as `docker/agent.Dockerfile`), seal as base ASIF. macOS: IPSW restore + `VZMacGuestProvisioningOptions` (agent user, SSH on), then **desktop-readiness baking**, which is mandatory for GUI guests: suppress the per-user first-login assistant (`com.apple.SetupAssistant` keys + the User Template), enable auto-login (`autoLoginUser` + `/etc/kcpassword`), disable screensaver/lock/display-sleep, and set a static wallpaper (animated "aerials" wallpaper is unusable over remote display). ⚠ Use RAW (not ASIF) for macOS guest disks until the beta's macOS-guest space-reclamation bug is fixed.
- `create <case>`: DiskImageKit stacked image: shared read-only base + fresh per-case overlay. Near-instant, space-efficient.
- `start <case>` / `stop` / `status` / `ip`: lifecycle + vmnet NAT; `ip` reports the guest SSH endpoint.
- `export <case>` / `import`: flatten overlay + workspace tar + manifest, credentials excluded (mirror of docker-export).
A VM dies with its owning process, so `start` spawns a DETACHED per-VM runner process (analog of the detached `scripts/self-update.sh` trick) rather than a monolithic daemon; `status` talks to it over a unix socket in the instance data dir (`dataPath()`, never a hardcoded `~/.codeman` path).
### Key decision 3: multi-guest, and GUI is a first-class mode (REVISED 2026-07-29 by the repo owner)
The subsystem supports both macOS and Linux guests, and a guest runs in one of two **display modes**:
| | macOS guest | Linux guest |
| --- | --- | --- |
| **GUI mode** | **the point of the feature**; a real macOS desktop. Mandatory: nothing renders without an attached `VZVirtualMachineView` in an unlocked host session | supported (EFI + virtio-gpu framebuffer) for desktop Linux cases |
| **Headless mode** | not offered: a macOS guest with no view renders nothing, so a "headless macOS desktop" is a contradiction. SSH-only macOS is possible but is not what this feature is for | supported and cheap; the natural mode for agent/CI work, driven over SSH |
Consequences that flow from GUI being first-class:
- VM processes are **GUI apps in the console user's session** (LaunchAgent / `launchctl asuser`), never daemons. A daemon-launched VM cannot render.
- **The host is part of the product surface**: it must auto-login, never lock, never sleep, and keep a live WindowServer. Host lock == every VM's screen goes black, so the screen lock is effectively a global kill switch for every VM display on the machine. The product must own these host settings rather than treat them as user preference.
- **FileVault conflicts with unattended GUI hosting** and the trade-off must be a deliberate choice: FileVault disables auto-login, so a full-disk-encrypted host needs a human at a keyboard (or a remote screen-sharing session) after every reboot before any VM can render. Options are (a) FileVault on, accept manual login per boot, (b) FileVault off on a dedicated VM host so it boots straight into a rendering session, or (c) FileVault on plus a remote-unlock runbook. Codeman should detect the state and tell the user which one they are in instead of silently serving black screens.
- **Guests must be desktop-ready, not just booted**: auto-login, no screensaver/lock, and the per-user first-login assistant pre-suppressed at base-image time (`com.apple.SetupAssistant` keys, plus the User Template so later accounts inherit it). Otherwise the user connects to a login prompt or a setup wizard, which is exactly what happened during the first hands-on run.
- **Capacity is capped for macOS**: at most 2 concurrent macOS VMs per host, confirmed by our own test on 27 beta 4 (3rd refused with `VZErrorDomain` 6 at 39% free RAM; it is a kernel quota, so bigger hardware does not help). Scheduling must queue or evict beyond 2, the UI must explain why, and the scheduler should tolerate the acknowledged slot-leak bug (a slot occupied with nothing running, host-reboot to clear). Linux guests are uncapped and bounded only by host resources, which is the lever for scaling case counts on one machine.
- **Access is via the guest's own screen**, viewable in a browser through the noVNC chain (see `docs/vm-subsystem-apple-stack.md` §8), so no client-version or client-install requirements land on the user.
Provisioning per guest type: `VZMacGuestProvisioningOptions` for macOS (needs 27-on-27, first-boot-only, and does NOT skip the per-user wizard), cloud-init NoCloud seed ISO for Linux (proven working).
### Key decision 3b: the GUI VM host profile, and supervision that catches black screens
GUI hosting only works if the host is configured for it and supervised. This profile was derived the hard way on the testbed (prototyped there 2026-07-30) and should be what `codeman-vm` installs and verifies:
**Host profile** (the product should own these, not leave them to preference):
1. **No login barrier.** Either FileVault off + auto-login (a dedicated VM host boots straight into a rendering session, fully unattended), or FileVault on and remote reboots done with `sudo fdesetup authrestart`, where the pre-boot unlock *is* the login so the machine returns already logged in with encryption intact. **`authrestart` is VERIFIED on the testbed (2026-07-30): the host rebooted remotely and came back with a live logged-in console session, FileVault still enabled, no password prompt** — this is the recommended pattern for an encrypted GUI VM host. Plain reboots on a FileVault host always need a human, so Codeman should detect that combination and warn instead of serving black screens.
2. **Never lock**: lock policy off (needs the account password, so it is a setup step, not a scriptable one) plus `caffeinate -d -i -m -u` re-armed per session.
3. **Never sleep**: `pmset -a sleep 0 displaysleep 0 disablesleep 1`; a physical display is NOT required (a lid-closed laptop renders fine, only an unlocked session matters). Note OS updates reset these.
4. **Session-independent control plane**: run VPN/remote access as a system service, never a session app, and keep the access chain (forwards, VNC proxies, web endpoints) in LaunchDaemons so a session restart cannot sever operator access.
**Supervision** must be a **root LaunchDaemon**, not a user LaunchAgent. This is the load-bearing detail: a user agent cannot launch a GUI app into the Aqua session, so its restart attempts fail *silently* (the child dies instantly, leaving an empty log while the supervisor cheerfully reports success). A root daemon can, via `launchctl asuser <uid> sudo -u <user> …`, and those launches persist. Prototyped and verified on the testbed 2026-07-30; a working supervisor runs on a short interval and:
- Restarts the runner when the process is gone **or when its log shows `WindowServer event port death`**, which means it is permanently blind while still looking alive.
- Defers restarts while the console is at the login window, and launches into whichever session actually exists (resolve the console user with `stat -f %Su /dev/console`, never a hardcoded one).
- Re-points the guest port-forward whenever the guest's NAT lease changes, which happens on **every guest boot** under plain NAT. A vmnet DHCP reservation for a stable per-case IP is the better long-term answer.
- **Re-applies host power settings**, because `pmset -a disablesleep 1` does NOT survive a reboot (caught on the supervisor's first run after a real reboot) and OS updates reset it too.
- Re-arms the keep-awake helper, which dies with its session.
- Ideally also samples the guest framebuffer for non-black content, since a black screen is the one symptom common to every failure mode here.
`pgrep` alone is worthless for health: every failure mode in this session presented as a healthy process.
### Key decision 4: sessions ride the existing remote-SSH machinery
A provisioned guest is literally an SSH host on a vmnet IP. Session launch = the remote-SSH flow with the host swapped in: durable remote `tmux -L codeman-remote`, session names failing `SAFE_MUX_NAME_PATTERN` on purpose, EVERY ssh command line through `buildSshConnectionArgs()` (command-injection invariant), run flows through `POST /api/quick-start` (never `POST /api/sessions`, which stat-validates `workingDir` locally). What is genuinely new is only lifecycle (create/start/stop/export) and the vm-hosts/vm-cases overlay state.
### Key decision 5: workspace via VirtioFS at the same absolute path
Mirror the Docker bind-mount invariant: the case workspace is a real host directory shared into the guest via VirtioFS and mounted at the SAME absolute path. That keeps file-routes/watchers on real host bytes and makes the in-guest transcript projHash match the host. Without this, transcripts/attachments/file viewer all silently degrade.
### Key decision 6: credentials seeded, hooks bridged
- Credentials are SEEDED (read-only share, copied into the guest once at create), never shared read-write, and excluded from exports: byte-for-byte the Docker cases rule and rationale.
- Hooks: on the loopback-only prod bind a guest cannot reach `127.0.0.1:3000`. Mirror `CODEMAN_DOCKER_BRIDGE_HOOKS` with a `CODEMAN_VM_BRIDGE_HOOKS` opt-in listener on the vmnet gateway IP; otherwise idle detection falls back to output-based, same as Docker.
### Key decision 7: drift and teardown copy Docker semantics verbatim
Config hash label on the VM (guest type, cpu/mem, share list); a drifted launch is REFUSED, never silently launched stale. One VM per case shared by all sessions; session kill = in-guest tmux kill only; case delete = stop + remove overlay; instance-scoped boot reaper for orphaned runner processes.
## 5. Implementation phases
**Phase 0, testbed (no repo code):** dedicated MacBook on the macOS 27 beta, remotely accessible over the tailnet (setup protocol in Section 8), Xcode 27 beta, then a throwaway Swift script proving the loop: create base -> overlay -> boot -> ssh in. This validates 80% of the design before any Codeman code.
**Phase 1, `codeman-vm` helper:** SwiftPM package, the six subcommands above, JSON contract doc, detached runner + unix-socket status, Linux base image build. Deliverable is testable entirely without Codeman.
**Phase 2, Codeman integration:** types (`VmHost`/`VmCase`/`SessionVm`), `src/vm-hosts.ts` (+ pure helpers: config hash, arg building, endpoint parsing), Zod schemas, `case-routes` link/unlink + listing, `quick-start` vm branch reusing the remote-SSH launch path, `Session` threading + recovery round-trip, `VITEST` no-op layer, unit tests. Feature-detect: darwin + arm64 + helper binary present, else invisible.
**Phase 3, polish:** export/import UI, frontend Create Case "VM" tab + case-picker labels, SSE `vm:*` events, macOS-guest opt-in with cap surfaced, custom-Virtio input channel exploration, CLAUDE.md Key Pattern + `docs/vm-cases.md` + COM.
## 6. Testing
- Pure helpers unit-tested (ports pattern from `docker-hosts.ts`: 26 tests there, aim similar).
- All helper-invoking IO no-ops under `VITEST` (the `IS_TEST_MODE` pattern in `tmux-manager.ts`).
- End-to-end verification happens ON the beta MacBook, per the always-end-to-end rule: real base build, real per-case overlay boot, real quick-start into the guest, workspace round-trip through VirtioFS, session-delete keeps VM up, case-delete removes it.
- CI never runs the real path; the static guards are type-level + unit-level only.
## 7. Risks
1. **Beta API churn**: everything here targets beta SDKs; symbol/behavior changes are likely before fall GA. Mitigation: Phase 0/1 are throwaway-tolerant; no Codeman-side commitment until the helper contract survives a beta cycle.
2. **New artifact class**: Codeman ships pure TypeScript today; a Swift binary changes build/distribution (build-on-install via `xcrun swift build` on macs with Xcode CLT? prebuilt signed binary per release?). Needs an owner decision; local dev build is fine for the whole beta period.
3. **Entitlement/signing**: `com.apple.security.virtualization` is trivial for local dev, real for distribution.
4. **Adoption gating**: users need macOS 27 + Apple Silicon for months after GA. Docker cases remain the default recommendation; VM cases ship dark (feature-detected) with zero cost to everyone else.
## 8. Beta testbed plan: dedicated MacBook (actionable now)
Testbed is a dedicated MacBook the owner sacrifices to the beta (after a full backup). This supersedes the earlier dual-boot-the-Mini idea (git history has it): a dedicated machine means no OS-switching, no downtime for the Mini's live Codeman, and no FileVault pre-boot headaches.
**Sequencing rule that makes it headless: configure ALL remote access on the CURRENT macOS first, THEN upgrade in place.** An in-place beta upgrade preserves Remote Login, Tailscale, user accounts, and auto-login, so there is no Setup Assistant and no post-install physical step. (A fresh install would boot into GUI-only Setup Assistant with no SSH, which on a headless box is a dead end.)
Confirmed hardware (2026-07-28): MacBook, M3, 16 GB RAM, 256 GB disk with ~100 GB free. Verdict: green. M3 = eligible + nested-virt capable; 16 GB = host + 2-3 concurrent Linux guests (macOS guest = one at a time); 100 GB = fits with discipline: install Xcode 27 beta with the macOS platform only (skipping iOS/watchOS/tvOS simulators saves 15-20 GB), and defer any macOS guest base (~30 GB) to an external SSD or until actually needed. Linux guests + sparse ASIF overlays are the comfortable path.
### Pre-upgrade checklist (owner, physical, once)
1. Full backup (Time Machine or clone); the machine should be considered beta-only afterwards.
2. Tailscale: install, sign into the tailnet, confirm it appears in `tailscale status` from another node.
3. System Settings -> General -> Sharing: **Remote Login ON** (SSH) and **Screen Sharing ON** (for the rare GUI-only moments: Xcode license, Apple Account dialogs).
4. **FileVault stays ON** (owner decision 2026-07-28, security over convenience). Consequences: auto-login is unavailable, but FileVault's pre-boot unlock doubles as login, so an unlocked boot still lands in a live GUI session; planned remote reboots go through `sudo fdesetup authrestart` (unlocks for exactly one restart); an UNPLANNED reboot (beta kernel panic, battery drain) parks the machine at the pre-boot screen, no SSH/Tailscale, until the password is typed physically. If the testbed goes silent, suspect this first. Keep it on AC so the battery absorbs power blips.
5. Beta enrollment (manual): sign into the Apple Account in System Settings; System Settings -> General -> Software Update -> **Beta Updates** -> select the **macOS 27 Developer Beta** (preferred: framework fixes land weeks earlier than public beta; free since 2023 after accepting the agreement once at developer.apple.com; public-beta alternative: enroll at beta.apple.com). Then run the offered upgrade: plugged in, lid open, trusted network.
6. Send over: tailnet name/IP, username, and a first-login password (key install + lockdown happens remotely right after).
### Post-upgrade setup (remote, over the tailnet)
1. Verify: `sw_vers` reports 27.x, SSH reachable.
2. Server-ize the laptop: `sudo pmset -a sleep 0 disksleep 0 disablesleep 1` (lid-closed operation without an external display), `womp 1` (wake on network), `sudo systemsetup -setrestartpowerfailure on`. Keep on AC power.
3. Install the controlling host's SSH key, then disable password auth.
4. Xcode 27 beta install (the one step needing the owner's Apple Account sign-in once, doable via Screen Sharing from anywhere); `xcode-select`, license accept, verify `swift --version` + the 27 SDK (`xcrun --show-sdk-version`).
5. Phase 0 prototype loop, all remote from here: Linux guest base image (no 27-on-27 provisioning dependency), DiskImageKit overlay, boot, vmnet NAT, ssh into the guest, run `claude --version` inside.
6. Only after that loop works: start Phase 1 in `packages/codeman-vm/`.
## 9. Open decisions (owner)
1. Linux base distro/image for the default guest (proposal: Ubuntu 24.04 arm64 cloud image, matching the docker agent image's userland).
2. Helper distribution for GA: build-on-install vs prebuilt signed binary vs "bring your own Xcode".
3. Ship dark behind `CODEMAN_VM_CASES=1` for the first release, or feature-detect only?
4. Export format parity with docker-exports (one manifest schema for both?).
## References
- Session 224: https://developer.apple.com/videos/play/wwdc2026/224/
- Fleet-angle writeup: https://bitrise.io/blog/post/wwdc26-the-virtualization-framework-updates-that-matter-for-large-mac-fleets
- Beta timeline: https://www.macworld.com/article/3189014/apple-july-2026-ios-ipados-macos-27-public-betas-tv-arcade-releases.html
- Internal analogs: `docs/docker-cases-plan.md` (architecture template), `docs/remote-sessions.md` (session transport), `docs/architecture-invariants.md#docker-cases`
+281
View File
@@ -0,0 +1,281 @@
<!-- Reference doc for the VM subsystem (Codeman VM cases). Compiled 2026-07-29 from: Apple DocC JSON backend, macOS 27 beta 4 SDK on the testbed, a multi-source web research sweep, and hands-on prototyping on a MacBook Air M3 running macOS 27.0 beta (26A5388g). Companion to vm-cases-plan.md (the Codeman integration plan). -->
# The VM Subsystem: Apple Virtualization Stack Reference (macOS 27 "Golden Gate")
"VM subsystem" is the working name for Codeman's native-macOS VM isolation tier and everything under it. This document is the single place for what the Apple stack actually provides, what we have verified ourselves on the beta, and what is known-broken. The Codeman-side design lives in `docs/vm-cases-plan.md`.
**Research method note:** Apple's HTML doc pages are JS-rendered and come back empty to fetchers. The working route is the DocC JSON backend: `https://developer.apple.com/tutorials/data/documentation/<path>.json` (page content) and `https://developer.apple.com/tutorials/data/index/<framework>` (full symbol tree with per-symbol `beta` flags). Everything below marked "Apple docs" was parsed from that backend directly.
## 1. Component map and minimum OS versions
| Component | What it is | Min host OS | Notes |
| --- | --- | --- | --- |
| Virtualization.framework core | VMs, EFI/Linux boot, virtio devices, VirtioFS | macOS 11-13 era | Unchanged basics; our prototype uses nothing newer than macOS 13 APIs except the DiskImageKit bridge |
| **DiskImageKit** | ASIF + raw disk images, layered stacks | **macOS 27** | Swift-only, no ObjC headers. Section 2 |
| **Guest provisioning** | First-boot account/SSH setup for macOS guests | **macOS 27 host AND guest** | Mac guests only as of beta 4. Section 3 |
| vmnet topology/port-forward/DHCP APIs | Custom networks, port forwarding | **macOS 26** (NOT 27) | 27 adds exactly one fix: loopback port forwarding. Section 4 |
| `VZVmnetNetworkDeviceAttachment` | In-process vmnet attach | macOS 26 | |
| **`VZCustomVirtioDevice`** family | Custom paravirt devices | **macOS 27** | Linux guests only, custom guest driver required. Section 5 |
| AccessoryAccess (USB passthrough) | USB claim + attach to VMs | macOS 27 | Requires paid-team provisioning profile, Dock app. Out of scope for Codeman. Section 6 |
Corrections to the WWDC-session framing we started with: vmnet's topology family is a macOS 26 story (129 symbols, zero beta-flagged in 27); provisioning does NOT currently extend beyond macOS guests despite the generic-looking `VZGuestProvisioningOptions` base class; DiskImageKit has no attach/mount API at all (it is a file-format library that hands `DiskImage` objects to Virtualization, no `/dev/diskN`, no root needed, no entitlement documented).
## 2. DiskImageKit (macOS 27, Swift-only)
Public framework, `/System/Library/Frameworks/DiskImageKit.framework`. No ObjC headers; the API surface lives in the `.swiftinterface`. Verified present in the CLT 27 beta 4 SDK, and our prototype compiled against it with plain `swiftc` on the first attempt.
### API surface (complete as of beta 4)
```swift
class DiskImage {
convenience init(creating: some DiskImage.CreationConfiguration) throws
convenience init(opening: some OpenConfigurationProtocol) throws
func appending(any DiskImage.CreationConfiguration & DiskImage.StackableLayer) throws -> any StackedImage
func appending(consuming DiskImage) throws -> any StackedImage // reattach an existing layer; validates parentUUID
func truncate(blockCount: Int) throws // stacked: affects top layer; does NOT resize guest fs
var blockCount, blockSize, format, layerType, layerUUID, parentUUID, openMode, size, url
}
protocol StackedImage: DiskImage { var layers: [DiskImage] }
struct OpenConfiguration { init(url:mode:); Mode = automatic | readOnly | readWrite }
// CreationConfiguration statics: .asif(url:blockCount:blockSize:), .asifLayer(url:type:), .raw(url:blockCount:)
// DiskImage.LayerType: .cache | .overlay | .overlay(blockCount:)
// DiskImage.BlockSize: .bytes512 | .bytes4096
// Errors: CorruptedImageError, IncompatibleStackingError(reason), InvalidBlockCountError, UnsupportedFormatError
```
Bridge into Virtualization is a new beta convenience init on the existing attachment class. Note there is no `readOnly:` parameter; read-only-ness comes from each layer's own `openMode`:
```swift
VZDiskImageStorageDeviceAttachment(diskImage: stack, cachingMode: .automatic, synchronizationMode: .full)
```
### Stacking rules (Apple docs, verbatim where quoted)
- ASIF works standalone or stacked. "You can only use RAW images as standalone images or as **base** images in stacked configurations." Upper layers are always ASIF.
- **One cache layer per stack**, any number of overlays conceptually, "shallow stacks perform better" (WWDC 224). No published max-depth guidance.
- "Layers are processed from bottom (base) to top. The **topmost layer determines the stack's size and receives all writes**." `.overlay(blockCount:)` therefore also grows the virtual disk.
- UUID chaining: appending sets the child's `parentUUID` to the parent's `layerUUID`. Raw bases have no UUID. "The layer UUID **changes if the layer is written to**", and reattaching a mismatched layer throws `IncompatibleStackingError`. This is the mechanism that makes a shared read-only base safe.
- Base sharing across multiple VMs is the stated design intent ("can be shared across multiple VMs"), with the WWDC caveat that per-VM auxiliary files (EFI variable store, macOS auxiliary storage) must be duplicated per VM, never shared.
- **There is no flatten/merge.** An overlay cannot be merged back into its base (confirmed by Howard Oakley's coverage plus an independent hands-on report). Export/move flows must ship the layer chain, or flatten inside a guest (dd to a fresh attached image).
### Known issues and adoption
- **ASIF space reclamation is broken for macOS guests on the beta** (deleted files never return space, survives reboots). Linux guests reclaim correctly on both raw and ASIF via `fstrim -av`. Single detailed field report, unrefuted. Since the VM subsystem targets macOS guests, the practical rule until this is fixed is: back macOS guest disks with RAW, and revisit ASIF stacking for macOS guests each beta (stacking still works, the disks just never shrink).
- **Zero shipping adopters anywhere.** tart has a design issue with no activity; nobody has published working DiskImageKit code. Everything must be treated as field-untested (and our own testing bears that out, Section 8).
- Framework binary grew every beta (588 → 598 across betas 1-4); expect churn until GA.
- Release notes list no DiskImageKit known issues in any beta, which given the above says more about the notes than the framework.
## 3. Guest provisioning (macOS guests only)
```swift
class VZGuestProvisioningOptions: NSObject { func validate() throws } // "use one of its subclasses"
class VZMacGuestProvisioningOptions: VZGuestProvisioningOptions {
var fullName, username, password: String
var logsInAutomatically: Bool
var enablesRemoteLogin: Bool // SSH
}
// Wiring: VZMacOSVirtualMachineStartOptions.guestProvisioningOptions (Mac-typed)
// .setGuestProvisioning(_:) throws (validating setter)
```
- **Requires macOS 27 on host AND guest.** Older guests **silently ignore** the options (no error).
- **First boot after restore only.** Cannot reconfigure an already-provisioned VM; property changes after start are no-ops.
- The base class is forward-looking scaffolding; its only subclass is Mac. A Linux/cloud-init analogue may come later; do not assume it lands in 27.0. For Linux guests, cloud-init NoCloud seed ISOs remain the provisioning path (proven working, Section 8).
- Field-verified behavior (third-party hands-on, beta 3): provisioned account gets full admin + sudo; Setup Assistant fully skipped; SSH reachable ~48 s after first boot. **Race**: the account is created late in first boot (~T+54 s), after LaunchDaemons start (~T+33 s), so anything at daemon-level must wait for the account to exist.
- Open Apple-acknowledged bug: provisioned users are invisible to `CSIdentityQueryExecute()` (FB23716201).
- IPSW acquisition gotcha for automation: `VZMacOSRestoreImage.latestSupported` tracks the latest *release* (returned 26.5.2), not the installed beta; beta IPSWs must be fetched from the seed CDN explicitly.
## 4. vmnet: a macOS 26 feature set, one macOS 27 fix
Everything interesting shipped in macOS 26: `vmnet_network_create`, `vmnet_network_configuration_create`, `..._add_port_forwarding_rule`, `..._add_dhcp_reservation`, subnet/prefix/MTU/external-interface setters, NAT44/NAT66/DHCP/DNS-proxy/RA disables, plus serialization (`vmnet_network_copy_serialization` / `_create_with_serialization`) for handing networks across processes. `VZVmnetNetworkDeviceAttachment` is macOS 26.
macOS 27's only change (beta 4 release notes, verbatim): "The vmnet port forwarding APIs now support port forwarding when communicating over loopback." That closes the old gap where the host could not reach its own forwarded ports via 127.0.0.1 (confirmed working by the original bug reporter). Directly relevant to Codeman's loopback-bound production server talking to per-case guests.
Gotchas:
- vmnet networks are **not persisted**; they die with the owning process. Persist settings yourself and recreate (or serialize across processes).
- The `com.apple.vm.networking` entitlement is still restricted ("contact your Apple representative", though DTS says most requests are approved). The plain `VZNATNetworkDeviceAttachment` needs no special entitlement and is what our prototype uses.
- Ecosystem signal: tart's maintainer is not adopting in-process vmnet (prefers their separate-process softnet), so field testing of these APIs is thin.
## 5. VZCustomVirtioDevice (macOS 27, Linux guests only)
14 new types (`VZCustomVirtioDevice(+Configuration/Delegate/Provider)`, `VZVirtioQueue(+Element)`, `VZVirtioFeatureSet`, shared-memory-region types, `VZGuestMemoryMapping`), wired via `VZVirtualMachineConfiguration.customVirtioDevices`. Mandatory for guest discovery: `deviceID`, `pciClassID`, `pciSubclassID`, `virtioQueueCount`. You must write the Linux guest driver (Virtio spec 1.3/1.4). Threading contract: the framework calls the device/delegate on a serial queue (`deviceQueue`, defaulting to the VM's queue). Zero public adopters. For the VM subsystem this is a Phase 3+ option for a low-latency host-guest channel; SSH over NAT is proven and sufficient for now.
## 6. Signing and entitlements
- **Core loop (VZ + DiskImageKit + provisioning): ad-hoc signing with only `com.apple.security.virtualization` suffices.** Verified by us on beta 4 (plain `codesign --entitlements ... -s -` on a `swiftc` binary) and independently by third parties on beta 3. DiskImageKit documents no entitlement at all.
- **Over-entitling is the actual trap.** Adding `com.apple.application-identifier`/team-identifier keys without an embedded provisioning profile hangs the process before `main` (watchdog kill); shipping `com.apple.vm.networking` unauthorized gets AMFI SIGKILL at exec (exit 137, no crash report, even for `--version`). Keep the entitlements plist to exactly the one key.
- **USB passthrough breaks the ad-hoc story**: `com.apple.developer.accessory-access.usb` is profile-restricted (any paid team, no ad-hoc), additionally requires `com.apple.security.device.usb`, and `AAUSBAccessoryManager` presents UI, so it wants a Dock app, not a headless CLI. Out of scope for Codeman.
- No Xcode required for any of the above: the CLT beta (~500 MB via `softwareupdate`) carries the full macOS 27 SDK including DiskImageKit and compiles/signs everything.
## 7. Ecosystem state (July 2026)
- **tart is now `openai/tart`** (moved from cirruslabs, mid-2026) and **relicensed to FSL-1.1-ALv2** (no longer permissive). Provisioning support shipped in 2.33.0. Old cirruslabs URLs and license assumptions are stale.
- VirtualBuddy shipped provisioning ("Skip Setup Assistant") in 2.2 betas; had to add account-detail validation and a workaround installer for the cross-version bug below.
- lima is deliberately waiting for GA before touching macOS 27 APIs.
- **Code-Hex/vz (Go bindings) is dormant** (no commits since Feb 2026, no macOS 27 APIs), so the entire Go ecosystem (podman-machine, colima) currently has no path to these APIs. Swift is the only realistic binding today, which validates the VM subsystem's Swift-helper design.
- Useful pattern if ever supporting older SDKs: resolve new classes via `NSClassFromString` at runtime (no link-time dependency), fail gracefully when absent.
- **Cross-version restore bug**: installing a macOS 27 guest from IPSW on a macOS 26 host fails at 77-78% (`VZErrorDomain 10007`); fixed in 26.6b3 + Xcode 27b4 era, with a nasty MobileDevice.pkg trap (installing it from Xcode 27 beta on a 26 host requires a full macOS reinstall to undo). Not relevant to our 27-host testbed, very relevant to anyone on a 26 host.
## 8. Our empirical results (beta 4, 26A5388g, MacBook Air M3, 2026-07-29)
Prototype tooling, all in `~/vm-lab/` on the testbed, compiled with CLT-only `swiftc` and ad-hoc signed with the single virtualization entitlement:
| Tool | Purpose |
| --- | --- |
| `vzboot.swift` | Linux guest: EFI boot + virtio disk/net/entropy + NAT + optional cloud-init seed ISO + serial on stdio |
| `vzstack.swift` | Same, but boots a DiskImageKit stack (read-only raw base + ASIF overlay) |
| `vzmac.swift` | macOS guest: `install` (IPSW restore into a bundle) and `run` (boot, `--provision` for first-boot account/SSH) |
| `vzmacgui.swift` | macOS guest in a real window via `VZVirtualMachineView` (required for the guest to render at all) |
| `setup-seed.sh` | Builds a cloud-init NoCloud seed ISO with `hdiutil makehybrid` (volume label `cidata`) |
| `vncproxy.py` | RFB proxy that advertises only security type 2, so version-skewed/browser clients can authenticate |
| noVNC + `websockify` | Browser access; `websockify --web noVNC-<ver> 0.0.0.0:<port> 127.0.0.1:<proxy>` |
| `vmwatchdog.sh` + `vmaccess.sh` | Supervision: root LaunchDaemon that restarts a blind/dead runner, re-points the forward, re-applies `pmset`, re-arms keep-awake; plus a keeper for the proxy/web endpoints |
Host-side diagnostics written during this work (in the session scratchpad, not on the testbed): `vnclogin.py` (Apple DH auth + session open, distinguishes "credentials rejected" from "authorized but session refused"), `vncshot.py` (decodes the raw framebuffer to PNG and reports non-black pixel counts, plus optional synthetic wake input), `relay.py` (plain TCP relay used to bridge a tailnet peer to a LAN-only host), `sshpw.py` (pty-driven password SSH for the one-time key bootstrap into a freshly provisioned guest).
### Proven working
1. **Boot**: Debian 12 arm64 cloud images (nocloud and genericcloud variants) boot under `VZEFIBootLoader` + `VZGenericPlatformConfiguration`.
2. **Networking**: `VZNATNetworkDeviceAttachment` gives the guest a `192.168.64.x` DHCP lease from the host's bootpd (leases visible in `/var/db/dhcpd_leases`, bridge is `bridge100`).
3. **cloud-init provisioning**: NoCloud seed ISO (built with `hdiutil makehybrid -iso -joliet -default-volume-name cidata`) created a `codeman` user with SSH key + passwordless sudo on first boot; `ssh codeman@<lease-ip>` from the host works with key auth.
4. **DiskImageKit stack mechanics**: opening a raw base `.readOnly`, appending an ASIF overlay (`ASIFCreationConfiguration.layer(url:type:.overlay)`), attaching via `init(diskImage:)`, and booting it. The overlay received ~44 MB of boot-time writes while the **base file's SHA-256 stayed bit-identical**, which is the write-isolation property the whole per-case design rests on.
5. **Reattach**: reopening an existing overlay and `appending(consuming:)` onto the same base passes UUID validation.
6. **macOS guest install (added later the same day)**: `VZMacOSInstaller` restore of the 27.0 IPSW (26A5388g, fetched from the seed CDN via appledb; same build as host) into a sparse 64 GiB raw disk + auxiliary storage: INSTALL-OK on the first attempt, ~25 minutes.
7. **Headless guest provisioning WORKS**: `VZMacGuestProvisioningOptions` via `setGuestProvisioning` (username, password, `enablesRemoteLogin`, `logsInAutomatically=false`) produced, with zero GUI interaction: an account with full admin (groups include `80(admin)`, `com.apple.access_ssh`), Remote Login on from first boot, port 22 reachable ~140 s after first-boot start, hostname auto-derived from the account ("Codemans-Virtual-Machine"). SSH password auth is on by default, so the bootstrap path is: pty-driven password login once to install `authorized_keys`, key auth thereafter. Note the provisioned account's sudo is NOT passwordless (`echo <pass> | sudo -S ...`), and provisioning is first-boot-only (later boots take no options and just boot).
8. **Slot-leak bug NOT reproduced on 26A5388g**: a guest-initiated `shutdown -h now` fired `guestDidStop` cleanly and an immediate relaunch started fine (SSH-ready again in ~75 s), so FB22967193 (VM slot leaked on guest-initiated shutdown, host reboot to recover) did not manifest after one cycle. Either fixed in beta 4 or needs more cycles to trigger.
### Unstable / under investigation (beta-quality territory)
Boot reliability degraded over a ~15-VM session on one host boot, ending with reproducible silent hangs (VM process alive, 0% CPU, no DHCP, no ARP, nothing on serial):
- A genericcloud base that had been booted read-write once (cloud-init first boot) subsequently hung on every boot **with the seed ISO still attached**, while booting **without** the seed succeeded, then later runs failed in both configurations. The seed correlation is strong but was observed while host state was already suspect, so it needs a retest from a clean baseline.
- The first stack-boot "success" that later wedged turned out (via DHCP lease timestamp arithmetic) never to have reached the network at all; its overlay growth was pre-network boot writes.
- Working hypothesis, matching a class of acknowledged beta bugs (e.g. the VM-slot counter that leaks on guest-initiated shutdown, FB22967193, where only a host reboot recovers): accumulated hypervisor/vmnet state on the host degrades boots. Requires a host reboot + a disciplined retest matrix to confirm.
### Display rendering: the single most important operational finding
**A VZ macOS guest renders nothing unless a `VZVirtualMachineView` is attached AND the host session is actually drawing.** Verified byte-for-byte: the guest's own screen sharing serves an all-zero framebuffer (0 non-black bytes across 400 KB samples, with a sane pixel format: `rmax/gmax/bmax = 255`, shifts 16/8/0), in-guest `screencapture` fails with "could not create image from display", and no `IODisplayWrangler` shows up in the guest's `ioreg`. Three distinct states all produce black:
1. **Headless** (VM run with no view attached).
2. **View attached, host session locked.** The lock screen suspends drawing and the guest's virtual GPU produces no frames.
3. **View attached, but the app lost its WindowServer connection** (see the incident below): black permanently until the app is restarted.
**Consequence for the VM subsystem: rendering is a first-class requirement, not an optional extra (owner decision 2026-07-29).** The product serves GUI desktops: mandatory for macOS guests, optional-but-supported for Linux guests (which can also run headless over SSH). Any VM in GUI mode must be launched by an app that attaches a `VZVirtualMachineView`, from inside a host GUI session that is logged in and unlocked. That makes the following non-negotiable parts of the design, not workarounds:
- VMs run as **GUI apps in the console user's session** (launched via a LaunchAgent or `launchctl asuser`), never as daemons.
- The **host must auto-login and never lock or sleep**; a locked host is equivalent to a powered-off display for every VM on it.
- The **guest must auto-login, never lock, and have its first-login assistant pre-suppressed**, or the "desktop" a user connects to is a password prompt or a setup wizard.
- A VM app that loses its WindowServer connection is **permanently blind** and must be restarted; supervision has to detect that, not just check that the process is alive.
- The **2-concurrent-macOS-VM cap** becomes a real capacity limit for the product, so it must be surfaced in the UI and tested (still untested worldwide as of this writing).
### Incident 2026-07-29: `killall -HUP loginwindow` (never do this on a remote Mac)
Applying a wallpaper change on the testbed with `killall -HUP loginwindow` restarted the host's login session. Three consequences:
1. **The Mac dropped off the tailnet entirely.** Tailscale's App Store build is a GUI app living in the user session, so killing the session killed the VPN; remote access was gone until someone logged in. Recovery came from a second machine on the same LAN: it could still SSH in, and then relay ports back over the tailnet (a plain TCP relay on a tailnet-connected LAN peer is a good out-of-band path worth keeping ready).
2. **The VM app lost its WindowServer connection** (`HIToolbox: received notification of WindowServer event port death`) while surviving as a process. Every later black screen traced to this, and nothing guest-side could fix it; only restarting the app restored rendering.
3. The session's `caffeinate` died, so the host resumed auto-locking.
Rule: on a remote Mac, never run session-level commands (`killall -HUP loginwindow`, `pkill -u <user>`, logout, fast user switching). `killall WallpaperAgent` alone is session-safe. Before any such command, enumerate what depends on that session: VPN, VM processes, port forwards, keep-awake helpers.
### Keeping host and guest usable unattended
- **Host**: `caffeinate -d -i -m -u` prevents display sleep but does NOT override the lock policy. "Require password after screen saver begins or display is turned off → Never" must be set in System Settings; it needs the account password, so a passwordless-sudo shell cannot script it, and turning it off does NOT dismiss a lock that is already engaged (one more unlock is always needed). `pmset -a disablesleep 1` keeps a lid-closed laptop awake but **does not survive a reboot**, and OS updates reset it too, so a supervisor should re-apply it rather than assume it sticks.
- **Rebooting an encrypted host**: use `sudo fdesetup authrestart`. FileVault's pre-boot unlock doubles as the login, so the machine returns with a **live logged-in console session** and encryption intact, no password prompt, and supervision can then bring the VMs back by itself. Verified 2026-07-30. A plain `reboot` parks at the lock screen and blacks out every VM until a human logs in.
- **Guest**: set `autoLoginUser` plus a valid `/etc/kcpassword` (XOR-obfuscated password file, key `7D 89 52 23 D2 BC DE A3`, payload zero-padded to a multiple of 12). `sysadminctl -autologin` fails with `SACSetAutoLoginPassword error:22` on provisioned accounts, and a fresh guest has no Python, so generate the bytes on the controlling host and copy them in. Then `pmset -a displaysleep 0 sleep 0 disablesleep 1`, `defaults -currentHost write com.apple.screensaver idleTime 0`, `defaults write com.apple.screensaver askForPassword 0`, and `caffeinate` inside the guest. ⚠ `autoLoginUser` was observed being wiped by failed `sysadminctl -autologin` attempts; verify it after each boot until stable.
- **Wallpaper**: animated "aerials" wallpaper is brutal over VNC. The provider lives in `~/Library/Application Support/com.apple.wallpaper/Store/Index.plist` under several keys (`AllSpacesAndDisplays:Desktop`, `:Idle`, and `SystemDefault:*` which is what the login/lock screen uses). Switch each `Provider` to `com.apple.wallpaper.choice.solid-color` with PlistBuddy and restart `WallpaperAgent`. The login-window copy is cached and only refreshes on a later login cycle.
### Remote GUI/SSH access to a guest (recipe, verified 2026-07-29)
The guest lives on the host-private NAT bridge, so remote access is guest-service + host-forward:
1. **In the macOS guest** (over ssh), use ONE mechanism, fully activated. The reliable form is Remote Management in a single kickstart call:
```
sudo .../RemoteManagement/ARDAgent.app/Contents/Resources/kickstart \
-activate -configure -access -on \
-clientopts -setvnclegacy -vnclegacy yes -setvncpw -vncpw <8-char-pw> \
-allowAccessFor -allUsers -privs -all -restart -agent -menu
```
⚠ **Half-configured states authenticate but refuse the session.** Loading `com.apple.screensharing` while Remote Management is deactivated (or vice versa) produces an Apple-client error that names the wrong culprit: *"Screen Sharing is not permitted on <host>. Disable and re-enable Screen Sharing or Remote Management in System Settings"*. A raw-protocol client can still authenticate AND open a framebuffer in that state, so protocol-level tests pass while every Apple client fails. The remedy is exactly what the dialog says, done over ssh: `launchctl unload -w …screensharing.plist`, `kickstart -deactivate -configure -access -off`, `pkill screensharingd`, then the single activate call above.
Notes: `launchctl enable system/com.apple.screensharing` fails with "Could not find service" on this build; `load -w` is the plain-Screen-Sharing path if you deliberately want it instead of Remote Management. Apple clients negotiate `RSA-SRP` (auth type 33) and the guest logs `Authentication: SUCCEEDED :: User Name: … :: Type: RSA-SRP` on success, which is the definitive server-side confirmation.
2. **On the host**: a gateway port-forward makes the guest's 5900 reachable from the whole tailnet without per-client tunnels: self-authorize the host's own key, then `ssh -N -g -L 0.0.0.0:5901:<guest-ip>:5900 <user>@localhost` (nohup'd).
⚠⚠ **NEVER forward on host port 5900.** If the host has Screen Sharing enabled (our testbed does, from the pre-upgrade checklist), launchd already owns 5900 socket-activated. The `ssh -L` bind then fails with "Address already in use" **while the tunnel process keeps running**, so every symptom of success is present (process alive, port answers, real RFB banner) yet **every connection reaches the HOST's login window, not the guest**. This cost us an hour: guest credentials failed against the host's screensharingd, which reads exactly like broken guest auth, and we chased the (real, but irrelevant) provisioned-account identity bug. Diagnostics that would have caught it instantly: `sudo lsof -nP -iTCP:5900 -sTCP:LISTEN` showing `launchd` rather than `ssh`, or the guest's own logs showing NO auth attempts during a failed login. Always use a distinct host port and verify with `lsof` that the forward owns it.
⚠ `-g` binds all interfaces, so the forward is also visible on the host's LAN; the VNC layer still requires the account or VNC password. ⚠ The forward pins the guest IP, which changes per boot under plain NAT; re-point it after a guest reboot (the proper fix is a vmnet DHCP reservation, macOS 26 API, once we move off plain `VZNATNetworkDeviceAttachment`).
Verified working: with the forward on 5901, both a provisioned account and a `sysadminctl`-created one authenticate successfully (RFB `SecurityResult` = 0) against the guest. The guest offers security types `[30, 33, 36, 2, 35]`, i.e. Apple DH/SRP **plus classic type 2**, so non-Apple VNC clients work with the legacy password once ARD's `-setvnclegacy` is set. (The host's screensharingd, by contrast, offered no type 2, which is itself a tell that you are talking to the wrong machine.)
3. **SSH from any tailnet device**: `ssh -J <host-user>@<host> codeman@<guest-ip>` (jump through the host), after adding the connecting machine's key to the guest's `authorized_keys`.
**Client-version incompatibility (macOS 27 servers vs older Screen Sharing clients)**: an older Mac's Screen Sharing client fails Apple's `RSA-SRP` handshake against macOS 27 servers, logging `Authentication: FAILED :: User Name: <user> :: Type: RSA-SRP` server-side, while a macOS 27 client authenticates against the same servers without issue. This was verified against BOTH a macOS 27 guest and a macOS 27 host with the operator's own account, so it is a client-side version skew, not configuration, and no server-side change fixes it. Same family as the documented "macOS 26 host cannot install a 27 guest" bug. Practical workaround: bypass Apple auth entirely with classic VNC auth (security type 2), which macOS offers only when Remote Management legacy VNC is enabled. Two ways to consume it: any third-party VNC client, or a browser via noVNC.
**Browser-based access chain (zero client install, version-proof)**, all hosted on the Mac:
```
browser --HTTP/WS--> websockify (+ noVNC static files)
--> type-2-only proxy # rewrites the server's security-type list to [2]
--> ssh -L forward # loopback hop; see the Local Network note below
--> guest:5900
```
Notes learned the hard way: (a) **never bind the forward on host port 5900** (see the launchd warning above); (b) a Python proxy cannot reach the guest subnet directly because macOS **Local Network privacy** denies headless CLI binaries, surfacing as `No route to host`, so point the proxy at a loopback `ssh -L` forward instead (Apple-signed `ssh` is unaffected); (c) noVNC needs `?resize=scale` or Scaling Mode → Local Scaling, otherwise a Retina host screen (2940x1912) is unusable in a browser window; (d) noVNC speaks security type 2 only, which is exactly why the proxy rewrite is needed.
**Debugging technique that settled all of this**: a ~80-line Python RFB client (scratchpad `vnclogin.py`) that implements Apple DH auth (security type 30) and continues through `ClientInit`/`ServerInit`. It reports the server's `SecurityResult` plus the framebuffer size and desktop name, which separates "credentials rejected" from "authorized but session refused" without any GUI client. Pair it with `log stream --predicate 'process == "screensharingd"'` inside the guest, and drive a REAL Apple client headlessly from the host with `sudo launchctl asuser <uid> sudo -u <user> osascript -e 'tell application "Screen Sharing" to open location "vnc://user:pass@host:port"'`, verifying the result via `lsof -nP -iTCP -a -p <pid>` (an ESTABLISHED socket to the target) since `screencapture` fails on a lid-closed laptop ("could not create image from display"). Tailscale was never implicated: both the raw client and Apple's client work over the tailnet address once the guest service is fully activated.
### Hard-won operational lessons (write these into any tooling)
- **Silent serial is normal, not failure.** Debian's GRUB/kernel log to the graphics console; nothing attaches a getty to hvc0 by default. The reliable boot signal is the DHCP lease (or passive `tcpdump -i bridge100`), never the serial port and never a quick ping (BSD ping's first packet often dies to ARP latency; passive capture showed "dead" guests alive).
- **DHCP lease entries carry truth**: `name=` shows the guest hostname, and the lease timestamps order events; stale entries linger, so compare timestamps before attributing a lease to a boot.
- **Never boot a base image read-write.** Every RW boot mutates it (dhclient lease cache, journal, cloud-init state) and destroys experiment reproducibility, exactly why the production design only ever boots bases under overlays. Provision INTO the base once at base-build time, or provision per-case overlays with the seed, then detach the seed.
- **A killed SSH client does not kill a remote `nohup`'d VM**, and the survivor holds the EFI variable store lock: "The EFI variable store is already in use" (`VZErrorDomain 50002`) means a zombie VM process, `pkill` it.
- **EFI variable stores are per-VM state.** Fresh stores boot reliably; reuse across different VM instances is at minimum suspect on this beta (Apple's own guidance for cloned VMs is one store per VM). Cheap policy: one store per case, created with the overlay, deleted with it.
- **Downloads from cloud.debian.org mirrors truncate silently**; always verify byte count against origin `Content-Length` and resume with `curl -C -`.
- The remote host's default shell is zsh: `=` -prefixed words (`echo ===`) explode via zsh's `=cmd` expansion; keep separators zsh-safe in automation.
### The 2-concurrent-macOS-VM cap: TESTED AND CONFIRMED on macOS 27 beta 4 (2026-07-29)
We measured it, which as far as we can tell nobody had published for macOS 27. Method: `cp -c -R` the guest bundle (APFS clonefile, instant and **zero additional disk**), regenerate the machine identifier per clone (`VZMacMachineIdentifier()` written to `machine.id`; the hardware model is reused), then launch VMs until one is refused.
Result: VM #1 (8 GB, GUI) and VM #2 (4 GB, headless) ran concurrently without complaint. VM #3 was refused **instantly** at `vm.start`:
```
VZErrorDomain Code=6 "The maximum supported number of active virtual machines has been reached."
NSLocalizedFailure = "The number of virtual machines exceeds the limit."
```
**This is a licensing/kernel quota, not a resource limit**: the refusal came with **39% of system memory free** on a 16 GB host, and adding RAM or CPU cannot raise it. It matches the pre-27 behavior (`hv_apple_isa_vm_quota`), so nothing changed in 27 despite the framework's other additions. Linux guests are unaffected and are bounded only by host resources.
Design consequences: macOS-guest capacity per host is **hard-capped at 2**, so a GUI-macOS-per-case product must schedule around it (queue, evict idle VMs, or scale across hosts) and surface it in the UI. Also relevant: the acknowledged slot-leak bug (a guest-initiated shutdown failing to release a slot, recoverable only by host reboot) is far more damaging under a cap of 2 than it sounds; we did not reproduce it on beta 4, but any scheduler should treat "slot appears used but nothing is running" as a real state.
### Not yet tested
- Cache layers (`LayerType.cache`), `.overlay(blockCount:)` disk growth, stack depth performance, VirtioFS + stack combination, `truncate`, ASIF disks for macOS guests (raw used so far; ASIF has the reclamation bug).
- One more scripting lesson from this session: inner `ssh` calls inside a piped `sh -s` script MUST use `-n`, or they consume the remainder of the script from stdin and it silently never runs.
### Session timeline (what was actually established, 2026-07-29)
Linux path: base image download (with resume, mirrors truncate) → `vzboot` compiles against the beta SDK first try → EFI boot → NAT DHCP lease → cloud-init seed provisions a user with the host's SSH key → `ssh` into the guest works → DiskImageKit stack boots with an ASIF overlay taking all writes while the base stays SHA-identical. Later Linux boots became unreliable on an un-rebooted host (silent hangs, 0% CPU, no DHCP); a clean-baseline retest is still pending.
macOS path: seed-CDN IPSW (matched to the host build) → `VZMacOSInstaller` restore, ~25 min, first try → first boot with `VZMacGuestProvisioningOptions` creates an admin account with Remote Login on, no interaction needed, SSH reachable ~140 s later → key bootstrap over a one-time password login → guest shutdown/relaunch clean (the slot-leak bug did not reproduce) → GUI access fought through a port collision, a client-version incompatibility, the rendering dependency, and a self-inflicted session kill, ending with a browser-based path plus a guest hardened to auto-login and never lock.
**Lifecycle verified (stop → start), 2026-07-30**: an in-guest `shutdown -h now` fires `guestDidStop` and the runner app exits on its own; relaunching from the same bundle boots the guest in ~2 minutes straight into an auto-logged-in desktop, and the VM slot is released cleanly (an immediate restart works, so the slot-leak bug did not bite). Two operational notes: the guest takes a **new NAT lease on every boot**, so any port-forward must be re-pointed (or use a vmnet DHCP reservation), and a host reboot resets `pmset -a disablesleep`.
⚠ **Provisioning does NOT skip the per-user first-login assistant.** `VZMacGuestProvisioningOptions` skips the initial Setup Assistant (account creation, region, Apple Account) so the machine is immediately reachable, but the first time anyone actually logs into a desktop, macOS still presents its per-user wizard (Apple Intelligence, Siri, privacy, appearance, Touch ID). The operator hit exactly this. For a GUI-first product this MUST be pre-suppressed during base-image creation by writing `com.apple.SetupAssistant` keys for every account that will log in, and into `/System/Library/User Template/English.lproj/Library/Preferences/` so accounts created later inherit it.
⚠ **A partial key list is worse than none**, because the wizard simply shows the panes you missed and the operator has to click through them again after every fresh login (we hit this twice). The set that finally silenced macOS 27 beta 4: `DidSeeCloudSetup`, `DidSeeSiriSetup`, `DidSeePrivacy`, `DidSeeAppearanceSetup`, `DidSeeTouchIDSetup`, `DidSeeAvatarSetup`, `DidSeeScreenTime`, `DidSeeApplePaySetup`, `DidSeeSafariImport`, `DidSeeAccessibility`, **`DidSeeActivationLock`, `DidSeeAppStore`, `DidSeeLockdownMode`** (the three easy to miss), plus the Express-Settings flags **`SkipExpressSettingsUpdating`** and **`SkipFirstLoginOptimization`**, and the version markers `LastSeenCloudProductVersion` / `LastSeenBuddyBuildVersion` / `PreviousSystemVersion` / `PreviousBuildVersion` matching the guest build. Verify afterwards by reading the domain back and checking that no `DidSee*` key is still `0`. Note these keys change between macOS releases, so base-image creation should re-verify per OS version rather than trust a hardcoded list.
## 9. Design implications for Codeman's VM subsystem
0. **GUI is a first-class mode, and for macOS guests it is the whole point (owner decision, 2026-07-29).** The subsystem serves real desktops, not only headless SSH boxes. macOS guests are GUI-only in practice (nothing renders without an attached view). Linux guests are supported in BOTH modes: GUI when the case wants a desktop, headless-over-SSH when it wants a cheap agent sandbox. The costs of the GUI path are in §8 "Display rendering": VMs as GUI apps in a live session, a host that never locks, guests that auto-login with their first-login wizard pre-suppressed, and the macOS concurrency cap as a real capacity limit.
1. **The macOS-specific liabilities are accepted costs, not reasons to avoid macOS guests**: provisioning is macOS-only and first-boot-only, ASIF space reclamation is broken for macOS guests on the beta (use RAW disks for macOS guests until fixed), and the 2-VM cap applies. Plan around each: RAW-backed macOS disks, provisioning baked into base-image creation, and capacity limits surfaced in the UI.
2. **Base immutability is not just hygiene, it is load-bearing**: DiskImageKit's UUID invalidation plus our sha-stability proof make a read-only shared base per image-generation the core artifact. Bases are built once (seed attached), then only ever opened `.readOnly` under per-case overlays.
3. **Seed ISOs are a base-build-time tool only.** Never attach a seed to a routine case boot (correlated with boot hangs on the beta, and semantically wrong anyway since cloud-init already ran).
4. **Per-case files**: overlay ASIF + EFI variable store live and die together with the case.
5. **Export = ship the layer chain** (base ref + overlay + manifest), not flatten; there is no flatten API. In-guest `dd` to a fresh image is the fallback for a true single-file export.
6. **Health checking must be lease/API based**, not serial/ping based, and Codeman's `codeman-vm status` should read `/var/db/dhcpd_leases` (or use vmnet DHCP reservations for deterministic per-case IPs, a macOS 26 API).
7. **Run `fstrim` periodically in Linux guests** (or mount with discard) so overlays stay sparse.
8. **Entitlements plist stays minimal** (exactly `com.apple.security.virtualization`) to dodge the AMFI/watchdog traps.
9. **Expect beta churn**: pin findings to build numbers (this doc: 26A5388g) and retest each beta; the framework binaries changed every beta so far.
10. **A macOS guest is only "ready" when its desktop is ready**, which is a stricter bar than "the VM booted". Readiness means: VM app running with a live WindowServer connection, guest auto-logged-in (not at a login or lock screen), first-login assistant suppressed, and the guest's screen sharing serving a non-black framebuffer. Health checks should sample the framebuffer for non-black content, because every failure mode in this session (headless run, locked host, dead WindowServer, locked guest, setup wizard) presents as a perfectly healthy-looking process with a black or useless screen.
10b. **Supervision must run as a root LaunchDaemon.** A user LaunchAgent cannot launch a GUI app into the Aqua session; its restarts fail silently (child dies instantly, empty log, supervisor reports success). Root + `launchctl asuser <uid> sudo -u <user> …` works and the launched process persists. This bit us on the first supervisor implementation and is easy to repeat.
11. **Remote-access plumbing belongs in the helper CLI, not in ad-hoc shell**: a `codeman-vm` implementation should own port selection (never 5900), forward lifecycle across guest IP changes (or better, vmnet DHCP reservations for stable per-case IPs), and a documented browser path, because every failure in this session came from hand-rolled plumbing rather than from the Virtualization APIs themselves.
12. **Never let control-plane connectivity depend on a GUI session** on a remote Mac host: prefer a Tailscale system service over the App Store app, and keep a LAN-adjacent peer able to relay as an out-of-band recovery path.
## Sources
Apple DocC JSON backend (diskimagekit, virtualization, vmnet trees; macOS 27 release notes) | WWDC26 session 224 https://developer.apple.com/videos/play/wwdc2026/224/ | eclecticlight.co ASIF/virtualization coverage | developer.apple.com/forums threads 839343 (CSIdentity bug), 830118 (cross-version restore), 830119 (VM-slot leak), 830383 (VM cap), 834822 + 831902 (USB entitlements), 822658 (vmnet loopback) | openai/tart issues 1261/1263/1268/1269/1285 | Spooky-Labs provisioning design doc | VirtualBuddy 2.2 release notes | lima-vm discussions | our own test transcripts on the testbed (`~/vm-lab/*.log`, this repo's session)
+26 -2
View File
@@ -1,7 +1,7 @@
# Web Tabs (dashboards as Codeman tabs)
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port
4000, as a tab beside your Claude/Codex/Gemini sessions. Codeman becomes one mission
4000, as a tab beside your Claude/Codex/Antigravity sessions. Codeman becomes one mission
control instead of Codeman plus a pile of browser tabs.
## Using it
@@ -46,7 +46,11 @@ Codeman, including a phone that is not on the tailnet.
`direct` mode (a plain cross-origin iframe) still exists and is cheaper, but it only
works for an HTTPS dashboard that permits framing. The **Test** button probes from
the server and tells you which mode applies.
the server and tells you which mode applies. Note what Test actually verifies:
**server-to-upstream reachability, nothing else**. It does not exercise the browser
sandbox, cookies, CORS, CSP, or any reverse proxy sitting in front of Codeman, so a
passing Test does not guarantee the embedded page will render (see the
cookie-authenticated reverse proxy caveat below).
## The sandbox, and when to turn it off
@@ -66,6 +70,17 @@ Even in trusted mode, Codeman never forwards its own credentials upstream: the
`Authorization` header and the `codeman_session` cookie are stripped on the way out,
so `CODEMAN_PASSWORD` cannot leak into a dashboard.
⚠️ **Sandboxed tabs may not work when Codeman itself is behind a
cookie-authenticated reverse proxy** (Cloudflare Access, Authelia, oauth2-proxy and
similar). The sandboxed frame is opaque-origin, so its stylesheet, script, and API
requests do not carry the proxy's authentication cookie; the proxy redirects them to
the login provider, where CORS/CSP kills them, and the embedded app renders
unstyled or broken while the Codeman page around it works fine. Trusted mode
(**Open sandboxed** off) keeps a real origin and the cookie, so it works. The
**Test** button cannot catch this: it checks that the Codeman *server* can reach the
upstream, not that a sandboxed *browser* frame can load assets through the public
authentication layer.
## How the proxy authenticates
A sandboxed iframe is opaque-origin, so every request it makes is cross-site: the
@@ -137,6 +152,15 @@ then every API call fails, which looks like the dashboard being broken.
- **Login-protected dashboards need trusted mode**, since a sandboxed frame has no
cookie jar. A server-side per-dashboard cookie jar would lift this and is the
natural next step if it becomes annoying.
- **Cookie-authenticated reverse proxies in front of Codeman break sandboxed tabs**
(#238). The sandboxed frame's requests carry no auth cookie, so the proxy bounces
them to its login provider and the app loads broken while Test reports reachable.
Use trusted mode behind Cloudflare Access and friends; see the warning above.
- **Slow endpoints and the upstream timeout** (#237). The proxy waits
`CODEMAN_WEBVIEW_TIMEOUT_MS` (default 300s) for the upstream's response *headers*,
then streams the body without any time bound; a header timeout is logged
server-side and answered as a 502 that names the limit. WebSocket handshakes use
the separate `CODEMAN_WEBVIEW_WS_HANDSHAKE_TIMEOUT_MS` (default 30s).
- **Not a security boundary.** The proxy reaches whatever the Codeman server can
reach. That is not an escalation for someone who already commands
`--dangerously-skip-permissions` agents, but in multi-user mode it does mean a
+207
View File
@@ -0,0 +1,207 @@
# Agent CLIs
Codeman drives seven run modes: six agent CLIs plus a plain shell. This page covers picking
one, setting it up, and the differences that actually change how you work.
## The seven modes
| Mode | CLI | Get it |
| -------------------- | ---------------------------- | ---------------------------------------------------------------------- |
| **Claude Code** | `claude` | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) |
| **OpenCode** | `opencode` | [opencode.ai](https://opencode.ai) |
| **Codex** | `codex` | [developers.openai.com/codex/cli](https://developers.openai.com/codex/cli) |
| **Gemini** | `gemini` | [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) |
| **Antigravity** | `agy` | [antigravity.google](https://antigravity.google) |
| **Pi** | `pi` | [pi.dev](https://pi.dev) |
| **Terminal / Shell** | your `$SHELL` | Already installed. |
Any combination works, including all of them. The run mode is chosen per session from the
arrow beside the **Run** button, so one case can have a Claude session and a Codex session
open side by side.
## Codeman does not manage your logins
Install each CLI yourself and log it in once by hand. Codeman never collects, stores, or
refreshes your CLI credentials. It launches the binary and attaches to the result.
The one place credentials are touched is [Docker Cases](Docker-Cases), where host
credentials are copied into a container read-only at launch so you do not have to log in
again inside it. Even there, the container keeps its own copies and never writes back to
your host credential stores.
## Making a CLI visible to Codeman
Codeman resolves each binary from the environment the **server** runs in, which is not
necessarily the shell you tested in.
```bash
codeman doctor # what Codeman can actually see
codeman doctor --json
```
If a CLI is installed but a Run button for it never appears:
1. Check `which <cli>` in a plain login shell, not just your interactive one.
2. If Codeman runs as a service, remember that launchd hands a job
`/usr/bin:/bin:/usr/sbin:/sbin`. `codeman service install` bakes your PATH into the unit
precisely to avoid this; a hand-written plist or unit will not.
3. Restart the server after installing a new CLI.
`pi` is additionally version-probed rather than trusted by name, because `pi` is a generic
enough command that something else on your PATH may answer to it.
## Claude is the reference mode
A number of Codeman features exist only for Claude sessions. This is structural, not a
backlog: they depend on Claude Code's hook system, or on parsing Claude's specific terminal
output. The other CLIs expose no equivalent.
| Feature | Claude | Other CLIs |
| ------------------------------------------------ | ------ | --------------------------------------------------- |
| Sessions, tabs, scrollback, exactly-once input | Yes | Yes |
| Respawn cycling and unattended runs | Yes | Yes |
| Cron jobs | Yes | Yes |
| Docker cases, remote SSH cases | Yes | Yes |
| Precise idle detection (hook-driven) | Yes | Output-stabilization fallback, coarser |
| Auto-resume when a usage limit resets | Yes | No |
| Plan usage chip | Yes | No |
| Approvals Inbox | Yes | No |
| Read My Mind | Yes | No |
| Ralph loop and its task tracker | Yes | No |
| Subagent and team windows | Yes | No |
| Model, effort, and ultracode controls | Yes | No |
| `stop` and `blocked` wait signals | Yes | 400 if you ask for them explicitly |
| The bundled agent skill | Yes | No |
Everything that makes a session a session works everywhere. What is Claude-only is mostly
the machinery that needs to know *what* the agent is doing rather than *that* it is doing
something.
## Per-CLI notes
### Claude Code
The defaults you will care about, all under **App Settings**:
- **Model** (Models section). Written into the case's `.claude/settings.local.json` as a
soft default, so `/model` still works mid-session. The 1M-context Opus variant is a
switch on the model card rather than a separate model.
- **Effort** (`low` through `max`) or **ultracode** for dynamic multi-agent workflows. Also
a soft default: `/effort` overrides it any time. Effort is deliberately not passed as an
environment variable, because that would hard-lock it and block in-session switching.
- **Startup permission mode** (Agents & CLIs section). The default is
`--dangerously-skip-permissions`, which is why the security model matters. You can switch
new sessions to Anthropic's classifier-guarded `auto` mode, normal prompting, or an
explicit allowed-tools list.
**Separate Claude accounts per session.** Set `CLAUDE_CONFIG_DIR` in a session's environment
overrides to point it at a different Claude config directory, which is how you run one
session on a client's subscription and another on your own. One caveat: a relocated config
directory writes transcripts outside `~/.claude/projects`, which blinds the response viewer,
subagent windows, ultracode panel, and Read My Mind for that session. Symlink `projects`
back into the shared tree to keep them working:
```bash
ln -s ~/.claude/projects <configDir>/projects
```
### OpenCode
Renders its own TUI, so Codeman treats readiness as output stabilization rather than
watching for a prompt marker. Requires tmux, with no direct-PTY fallback, because its
environment is injected through socket-scoped `tmux setenv` rather than the command line.
Integration detail: [`docs/opencode-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/opencode-integration.md).
### Codex
Two behaviours that are deliberate and worth knowing:
- **Predictive echo instead of buffered echo.** Codex's composer reacts to every keystroke,
a `/` opens a live-filtering picker, arrows edit server-side state. Buffering keystrokes
until Enter starved it, so Codex paints each keystroke at the predicted cell while the
bytes on the wire stay byte-identical to what you typed.
- **The wheel is not forwarded** into its transcript. Codex ignores the mouse reports
Codeman would send, so forwarding produced a dead wheel. Scrolling in a Codex session is
local scrollback.
### Gemini
Enterprise only, since Google's June 2026 consumer cutover. Its environment allowlist
includes the broad `GOOGLE_*` namespace, deliberately, because Vertex AI authentication
needs `GOOGLE_CLOUD_PROJECT`, `GOOGLE_APPLICATION_CREDENTIALS`, and
`GOOGLE_GENAI_USE_VERTEXAI`. That is the loosest allowlist entry in Codeman and it affects
only the CLI you spawned yourself.
### Antigravity
Google's successor to the consumer Gemini CLI, invoked as `agy`. It keeps all of its state
in `~/.gemini/antigravity-cli/`, so the credential handling that applies to Gemini applies
to it as well.
### Pi
Pi needs the opposite instincts from every other CLI here.
- **It has no permission prompts and no sandbox.** There is no bypass flag to send, and
Codeman does not invent one.
- **Its privileged setting is project trust**, a three-way `--approve` / `--no-approve` /
unset. Approving trust makes Pi **execute repo-local `.pi/extensions` TypeScript**, so
point it at a repository you trust. In multi-user mode, a user without an explicit grant
gets `--no-approve` even when no configuration exists.
- **Authentication is `/login` inside the session**, or the server process's own
environment. Pi's roughly 34 provider keys (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`HF_TOKEN`, and so on) share no common prefix, and the environment allowlist is global
rather than per mode, so admitting them for Pi would widen the allowlist for every mode at
once. They stay out.
Guide: [`docs/pi-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/pi-integration.md).
### Terminal / Shell
A plain shell in a tmux session. No agent, no hooks, no idle detection.
On phones a shell session automatically swaps the keyboard accessory bar for terminal
controls: Ctrl, Esc, Tab, arrows, paste. **Ctrl is a one-shot modifier**: tap it, then tap a
letter, and the control byte is sent. It disarms on use, on a second tap, on any other
accessory key, on a session switch, and when the keyboard closes. Details in
[Mobile Guide](Mobile-Guide).
## Environment overrides
Per-session environment variables are set when creating a session and persist across
respawns. Which variables are accepted depends on the mode:
| Mode | Allowed prefixes |
| ----------- | --------------------------------- |
| Claude | `CLAUDE_CODE_*`, plus the exact key `CLAUDE_CONFIG_DIR` |
| OpenCode | `OPENCODE_*` |
| Codex | `CODEX_*` |
| Gemini | `GEMINI_*`, `GOOGLE_*` |
| Antigravity | `ANTIGRAVITY_*` |
| Pi | `PI_*` |
Anything outside the allowlist is rejected at the schema. This is intentional: the allowlist
is one global list, so widening it for one CLI widens it for all of them.
Two things that deliberately do **not** travel as environment variables: **effort**, because
an environment variable hard-locks it and blocks `/effort`, and **model**, which is written
into the case's `.claude/settings.local.json` so that `/model` keeps working.
## Choosing a mode
- **Claude Code** if you want every Codeman feature. Unattended overnight runs, usage-limit
auto-resume, the Approvals Inbox, and subagent visualization all assume it.
- **Codex, OpenCode, Gemini, Antigravity** when you prefer that agent or that model. You get
the session layer, respawn, cron, Docker, and remote SSH; you do not get the hook-driven
features.
- **Pi** if you want a fast, unsandboxed agent and you understand what project trust does.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
## Read next
- [Core Concepts](Core-Concepts) - run modes versus location overlays.
- [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings.
- [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode.
- [Security](Security) - what skipping permission prompts actually means.
+97
View File
@@ -0,0 +1,97 @@
# Autonomous Loops
Two features that go further than "keep the session going": the **Ralph loop**, which works
a task list to completion in one session, and the **Orchestrator**, which turns a goal into
a phased plan and drives it across agents.
Both are Claude-only, both are off by default, and neither is where to start. If what you
want is an agent that keeps working overnight, that is
[Keeping Agents Running](Keeping-Agents-Running), and it is simpler, better understood, and
what most people actually use.
## Which one, if either
| You have | Use |
| ------------------------------------------------- | ------------------------------------------------------------ |
| A session that stops too early | [Respawn](Keeping-Agents-Running) |
| A written task list to grind through | Ralph loop |
| One large goal that needs planning and checkpoints | Orchestrator |
| Work that should start at a certain time | [Cron Jobs](Cron-Jobs) |
| Several workers to fan out and supervise | [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) |
## The Ralph loop
Named after the Ralph Wiggum pattern: keep feeding the agent its own task list until the
list is empty.
The shape of it:
- The task list lives in a plan file in the case, conventionally `fix_plan.md`.
- Each cycle the agent reads the plan, works the next incomplete task, and marks progress.
- Codeman watches the file, tracks todos, and detects stalls.
- The loop ends when the agent signals completion, when the iteration cap is reached, or
when you stop it.
Start it from **Session Options → Ralph / Todo**, or from the wizard on the welcome screen.
| Setting | What it does |
| ---------------------- | ------------------------------------------------------------------------ |
| Max iterations | Hard ceiling on cycles. |
| Max todos | Cap on tracked tasks, default 500, oldest evicted first. |
| Todo expiration | Auto-expiry for stale todos, default 60 minutes. |
| Plan file | Which file holds the task list. |
A **circuit breaker** sits behind it to stop respawn thrashing: it moves from closed to
half-open to open, and is reset explicitly from the session's Ralph controls.
Honest assessment: Ralph is functional but is not where development attention goes. It
predates the respawn presets, which cover most of what people originally used it for with
less ceremony. Treat it as a specialised tool rather than the headline feature.
Full background, including the upstream pattern it is based on:
[`docs/ralph-wiggum-guide.md`](https://github.com/Ark0N/Codeman/blob/master/docs/ralph-wiggum-guide.md).
## The Orchestrator
A state machine that turns one goal into a phased plan and drives it to completion:
```
idle → planning → approval → executing → verifying → (replanning) → completed / failed
```
- **Planning** turns your goal into phases.
- **Approval** is yours. You see the plan before anything runs.
- **Executing** runs each phase, using team agents and the task queue.
- **Verifying** gates each phase before the next one starts. A failed gate can send it back
to replanning rather than forward.
Open it from the Orchestrator panel in the toolbar. State persists in `state.json`, so a
server restart does not lose an in-flight plan.
Where it differs from Ralph: Ralph is one session grinding a list, the Orchestrator
coordinates phases and agents with verification between them. It suits work that has a
natural shape ("migrate this, then update callers, then update the tests") rather than a
flat backlog.
Architecture: [`docs/orchestrator-loop-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/orchestrator-loop-architecture.md).
## Running any of this safely
Autonomous loops are the features most able to spend money and change code while you are not
looking. Some habits that pay off:
- **Run them in a case that is a git repository**, on a branch you are willing to throw
away. Being able to read the diff afterwards is the whole safety net.
- **Consider a container.** [Docker Cases](Docker-Cases) gives the agent its own filesystem
and network, and one checkbox is all it costs.
- **Set the iteration cap deliberately.** It is the ceiling on the spend.
- **Turn on notifications** so a blocked loop reaches you: see
[Notifications And Approvals](Notifications-And-Approvals).
- **Read the run summary and lifecycle log afterwards**, not just the final diff. They show
where it went sideways and recovered.
## Read next
- [Keeping Agents Running](Keeping-Agents-Running) - the simpler feature that usually fits better.
- [Watching Agents Work](Watching-Agents-Work) - seeing what a loop is doing while it runs.
- [Docker Cases](Docker-Cases) - a sandbox for unattended work.
+118
View File
@@ -0,0 +1,118 @@
# Contributing
The full guide lives in
[CONTRIBUTING.md](https://github.com/Ark0N/Codeman/blob/master/.github/CONTRIBUTING.md).
This page is the short orientation, plus how to fix a page in this wiki.
## Where things go
| You have | Send it to |
| --------------------------- | ---------------------------------------------------------------------------------------------- |
| A bug | An [issue](https://github.com/Ark0N/Codeman/issues), with OS, install method, browser, and which CLI the session was running. |
| A question or setup problem | [Discussions](https://github.com/Ark0N/Codeman/discussions). |
| An idea | [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas), where it gets voted on. |
| A small fix | Straight to a PR. |
| A bigger feature | An issue or Discussion first, then build once the design has a nod. |
| A security problem | Never a public issue. See [SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md). |
Issues usually get a response within a day, and every release credits its contributors and
bug reporters by name.
## Dev setup
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # http://localhost:3000
```
Requirements: Node 22+, tmux, and at least one agent CLI on your PATH.
The frontend is plain JavaScript with no bundler in dev: edit a `.js` or `.css` file and
reload. The exception is `index.html`, which is read once at server start, so markup changes
need a restart.
## Before you push
CI runs all of these, so running them locally saves a round trip:
```bash
npm run typecheck
npm run lint
npm run format:check
npm run check:frontend-syntax
npm test -- test/<file>.test.ts # one file, the normal way
npm run test:ci # the full CI sweep
```
**Never run bare `npm test`.** The default configuration includes browser-driven Playwright
suites that need a live server, Chromium, and environment-specific baselines; they hang or
fail on a normal machine. `test:ci` is the honest "run everything".
Tests are tmux-safe by design: under vitest the tmux layer becomes an in-memory mock, so
tests cannot touch real sessions. If you add a test that binds a port, pick a unique one at
3150 or above, and never 3000.
## Finding your way around
- Every source file opens with a `@fileoverview` block. Read it before the file; it is the
map.
- [`CLAUDE.md`](https://github.com/Ark0N/Codeman/blob/master/CLAUDE.md) at the repo root is
the densest architecture primer there is. It is written for AI coding agents, but its
invariants apply identically to humans, and most review feedback traces back to something
already written there.
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md)
holds the deep mechanisms and the history behind each rule.
## Good first contributions
- **A theme skin.** A skin is four things kept in sync, and a static test checks the sync, so
if the test passes your skin works.
- **A language.** The i18n module is dependency-free, English is canonical, and Simplified
Chinese is a complete example to copy.
- **Docs.** If you got stuck and then figured it out, the sentence that would have unstuck
you is a pull request.
- Anything labelled
[good first issue](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
Worth discussing first: new CLI backends, and real-device testing reports, especially
mobile, which always find things emulation cannot.
## PR expectations
- One change per PR. Small and focused reviews fast; a grab bag stalls.
- Target `master`.
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all, which
is a GitHub quirk rather than a Codeman one. Rebase when conflicts appear.
- Include or update tests when you change behaviour.
- Do not bump versions or edit the changelog; releases are handled after merge.
- AI-assisted contributions are welcome, with one condition: understand what you are
submitting, and actually run it. "The model said it works" is not a test.
## Fixing this wiki
These pages are generated from
[`docs/wiki/`](https://github.com/Ark0N/Codeman/tree/master/docs/wiki) in the main
repository, and pushed here automatically when master changes.
**Editing a page in the browser will be overwritten by the next sync.** Send a pull request
against `docs/wiki/` instead. It is plain markdown, and a documentation PR is a genuinely
useful contribution.
Conventions for wiki pages:
- Links between pages use the wiki form: `[Remote Access](Remote-Access)`, no `.md`.
- Links into the repository are absolute `https://github.com/Ark0N/Codeman/blob/master/...`
URLs.
- Images are referenced from the main repository over raw URLs rather than being copied into
the wiki.
- Say what the default is, especially when it is off. Most of Codeman is opt-in.
- Label Claude-only behaviour every time it appears. Six of the seven run modes are not
Claude.
## Conduct
Be kind, be direct, assume good faith. Report unacceptable behaviour privately via the
contact in
[SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
+183
View File
@@ -0,0 +1,183 @@
# Core Concepts
The five ideas the rest of the manual assumes: cases, sessions, run modes, location
overlays, and tmux. Plus what actually persists, and where it lives on disk.
## Case
A **case** is a named working directory that Codeman remembers. It is the unit you pick in
the toolbar before hitting Run, and every session belongs to exactly one.
A case is not a container or a sandbox. It is a folder plus a name plus a little
Codeman-side configuration:
- Which CLI the Run button should default to.
- Per-case toggles (Agent Teams, 1M Opus context).
- Where it runs, if it is not the local filesystem: see [Location overlays](#location-overlays).
Three ways to get one, all under **+** next to the case picker:
| How | Result |
| ----------------- | ------------------------------------------------------------------------------------------------------ |
| **Create New** | A fresh `~/codeman-cases/<name>` with a scaffolded `CLAUDE.md`. |
| **Clone Repo** | A public repo cloned into `~/codeman-cases/<name>` and registered as a case. |
| **Link Existing** | An existing folder anywhere on disk, registered in place. Nothing is copied or moved. |
Linked cases keep living where they are. Deleting a case in Codeman removes the
registration, and for a linked case that is all it removes.
**Cases created from scratch are the only copy of that code.** Uninstalling Codeman does not
delete `~/codeman-cases/`, but treat that directory as real work, not scratch space.
## Session
A **session** is one CLI process running in one tmux session, streamed to your browser.
Sessions are named `w<n>-<case>`, so `w1-myproject` is the first worker in the `myproject`
case. Each has a stable id, and that id is what the API, the wait primitives, and every
event use.
Several sessions can share one case. That is the normal way to parallelize: three workers
in the same repo, three tabs, one case.
A session carries state the case does not:
- Its run mode, model, effort level, and environment overrides.
- Its respawn configuration and Ralph loop state.
- Its terminal scrollback.
- Its owner, in [Multi-User Mode](Multi-User-Mode).
## Run mode
The **run mode** is which CLI the session runs: `claude`, `opencode`, `codex`, `gemini`,
`antigravity`, `pi`, or `shell`. It is chosen at start and does not change afterwards; to
switch, start another session.
Claude is the reference mode. Six of the seven are not Claude, and a number of Codeman
features are Claude-only for structural reasons rather than missing effort: they depend on
Claude Code's hook system or on parsing its terminal output. Every such feature is labelled
Claude-only where it appears, and [Agent CLIs](Agent-CLIs) lists them in one place.
## Location overlays
Where a case runs is **separate from** which CLI it runs. There are three locations:
| Location | What happens |
| -------------- | ------------------------------------------------------------------------------------------------------------ |
| **Local** | The default. tmux and the CLI run on the Codeman host. |
| **Docker** | One long-lived container per case; sessions `docker exec` into it. See [Docker Cases](Docker-Cases). |
| **Remote SSH** | A durable tmux server on the remote host, fronted by a local pane running `ssh`. See [Remote SSH Sessions](Remote-SSH-Sessions). |
This matters because it is a common source of confusion: Docker is **not** an eighth run
mode. All seven run modes work in all three locations. A case is docker-backed or
ssh-backed; a session is claude or codex or shell.
**Web tabs** are the other thing that is not a session. A saved dashboard URL renders as a
tab beside your agents, but there is no PTY, no tmux, and no respawn behind it. See
[Web Tabs](Web-Tabs).
## Why tmux
tmux is a hard requirement, and it is the reason Codeman behaves the way it does.
The agent runs inside a tmux session. Codeman attaches to it, the same way your terminal
would. That indirection buys:
- **Survival.** The agent outlives your browser tab, your network, your laptop lid, and a
restart of the Codeman server itself.
- **Real scrollback.** History is held by tmux, so reconnecting replays what happened while
you were gone instead of starting from blank.
- **Attach from anywhere else.** The same session is reachable from a terminal over SSH
with the `sc` chooser, or plain `tmux -L codeman attach`.
- **Secrets off the command line.** Environment overrides are injected with socket-scoped
`tmux setenv` rather than being visible in the spawn command.
The socket is `tmux -L codeman`, separate from your personal tmux server, so Codeman
sessions never appear in a bare `tmux ls`.
## What persists
| Survives | Does not survive |
| -------------------------------------------- | --------------------------------------------------- |
| Closing the browser | `tmux -L codeman kill-server` |
| Losing the network | A machine reboot (tmux dies with it) |
| Restarting the Codeman server | Killing the session from the UI |
| `codeman web --stop` | |
| A dropped SSH link, for remote cases | |
| A container restart, for docker cases | |
Conversation history is a separate question: Claude transcripts live in `~/.claude/`, so a
conversation can be resumed even after the tmux session is gone. That is what the welcome
screen's **Resume Conversation** list offers.
## State on disk
Everything Codeman knows lives under `~/.codeman/`:
| File | Holds |
| ---------------------------------------- | -------------------------------------------------------------------- |
| `state.json` | Sessions, settings, respawn config, orchestrator state, cron jobs. |
| `settings.json` | User preferences that sync across your devices. |
| `mux-sessions.json` | tmux recovery data. |
| `session-lifecycle.jsonl` | Append-only audit log of session starts, exits, and kills. |
| `linked-cases.json` | Registered cases. |
| `remote-hosts.json`, `docker-hosts.json` | Location overlay configuration. |
| `webviews.json` | Saved dashboard URLs. |
| `users.json` | Multi-user accounts, mode 0600. |
| `push-*.json` | Web push keys and subscriptions. |
| `certs/` | Self-signed TLS for `--https`. |
None of it needs root, none of it leaves the machine, and deleting `~/.codeman/` resets
Codeman to a fresh install without touching your code.
## Instances
The data directory and the tmux socket are both **process wide**. Two Codeman servers
started on one machine share them, which means the second one discovers the first one's
live sessions and attaches to them, resizing and mutating sessions you did not expect it to
touch.
To run two on purpose, give each its own instance name:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
That scopes the data directory and the tmux socket together, which is the only safe way to
do it. `CODEMAN_DATA_DIR` and `CODEMAN_TMUX_SOCKET` can be set individually if you need
them apart, but setting only one of the two reproduces exactly the problem you were trying
to avoid.
## Hooks
For Claude sessions, Codeman writes a hooks configuration into the case so Claude Code can
report events back: a permission prompt appeared, the turn finished, the agent went idle, a
task completed. Those events drive tab alerts, the Approvals Inbox, notifications, and the
wait primitives.
This is why some features are Claude-only. The other CLIs have no equivalent hook system,
so for them Codeman falls back to watching terminal output, which is coarser: it can see
that something happened, not what it was.
See [Hooks And Integrations](Hooks-And-Integrations).
## Vocabulary
| Term | Means |
| --------------- | ---------------------------------------------------------------------------- |
| **Case** | Named working directory. |
| **Session** | One CLI in one tmux session. |
| **Run mode** | Which CLI: claude, opencode, codex, gemini, antigravity, pi, shell. |
| **Respawn** | Restarting the CLI on idle to keep an unattended run going. |
| **Ralph loop** | An autonomous single-session task loop. |
| **Orchestrator**| A phased plan driven across multiple agents. |
| **Subagent** | An agent the CLI spawned itself, shown live in its own window. |
| **Web tab** | A saved dashboard URL rendered as a tab. Not a session. |
| **Instance** | One Codeman server with its own data directory and tmux socket. |
## Read next
- [The Dashboard](The-Dashboard) - what the UI is showing you.
- [Agent CLIs](Agent-CLIs) - the seven run modes in detail.
- [Keeping Agents Running](Keeping-Agents-Running) - respawn, idle detection, usage limits.
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
+161
View File
@@ -0,0 +1,161 @@
# Cron Jobs
Saved, named jobs that start a session and send it a prompt on a schedule. Cron for agent
sessions: *every weekday at 03:00, open a Claude session in `~/proj` and tell it to update
dependencies and open a PR.*
The ⏰ **Cron** header button is opt-in. Turn it on in
**App Settings → Header & Panels**.
## Creating a job
1. Click **⏰ Cron**, then **+ New Job**.
2. Give it a name, pick the agent type and working directory.
3. Write the prompt, or point at a file containing it.
4. Choose a schedule and leave **Enabled** on.
5. **Save**. The job appears with its computed next run.
**Run Now** fires it immediately without touching the schedule, which is the fastest way to
find out whether the prompt does what you meant.
## The fields
| Field | Notes |
| ------------------------ | ------------------------------------------------------------------------------------------- |
| **Name** | Also used as the created session's name. |
| **Agent type** | Any run mode, including `shell`. |
| **Working directory** | Validated when you save **and** again when the job fires. Blocked system trees are refused. |
| **Launch command** | Shell jobs only. Sent as the first line once the shell is up, before the prompt. |
| **Prompt** | Inline text, or a path to a file read at fire time. |
| **Input mode** | `typed` behaves like a human typing. `paste` writes directly. |
| **Schedule** | `once`, `interval`, `daily`, or `weekly`. |
| **Enabled** | Disabled jobs never fire on their own. **Run Now** still works. |
| **Concurrency policy** | What to do if sessions of the same type are already running. |
| **Auto-close previous** | Recurring jobs only. Closes the session the previous run created. Default on. |
| **Notes** | Free text for you. |
## Schedules
All wall-clock times are in the **server's local timezone**, not your browser's. A job set
for 03:00 fires at 03:00 where the server is.
| Type | Behaviour |
| ---------- | -------------------------------------------------------------------------------------------------- |
| `once` | Fires at an absolute time, then disables itself. A job missed because the server was down still fires once on the next tick. |
| `interval` | Every N minutes, from 1 minute to a year. |
| `daily` | At `HH:MM` every day. If today's time has passed, the next run is tomorrow. |
| `weekly` | At `HH:MM` on the weekdays you pick. |
Interval jobs re-anchor to when they actually fired, not to an ideal cadence, so a slow tick
or a server restart shifts later runs slightly. That drift is accepted rather than corrected.
## Prompts are single line
This is the rule people trip over. Programmatic input into an agent session is single line
everywhere in Codeman, because the terminal UIs these CLIs use treat a newline as submit. A
multi-line prompt would be silently mangled, so it is **rejected** instead: the form refuses
it, and a prompt file whose contents are multi-line fails the run with a clear message.
For anything longer than a sentence, put the instructions in a file and make the prompt tell
the agent to read it:
```
read TASKS.md and work through it
```
That is also easier to edit than a job field.
### Prompt files
Reading the prompt from a file at fire time is useful when the instructions change more
often than the schedule. The path is confined to the job's working directory, symlinks are
resolved before the check, sensitive trees are refused, and the file has to be a regular
file under 1 MiB.
If any of that fails, the run is recorded as failed and **no session is created**.
## Concurrency
Applies to scheduled runs only, never to **Run Now**:
| Policy | Behaviour |
| ------------------------------- | -------------------------------------------------------------------------------------- |
| `warn_only` | Always launch. The count of live same-type sessions is shown but does not block. |
| `skip_if_same_agent_running` | Skip this fire if another live session of that mode exists. |
The skip policy has the details you would want it to have:
- Only **live** sessions block. A tab whose CLI already exited does not count.
- Sessions the job created on its own previous runs never block it, otherwise a recurring
job would deadlock on itself after the first fire.
- A skipped `once` job is not consumed. It stays armed and fires when the blocker goes away.
- Consecutive skips are collapsed into one record per streak, so a perpetually skipped job
cannot bloat your state file.
## Run history
Every fire is recorded per job, with a status:
| Status | Meaning |
| --------- | -------------------------------------------------------------------- |
| `created` | The run started and a session was created. |
| `skipped` | The concurrency policy blocked it. Not counted as a run. |
| `failed` | The prompt could not be resolved, or the working directory was gone. |
The schedule is advanced **before** the session launches, so a slow start cannot cause the
same job to re-trigger.
## Cron versus the other autonomy features
| Want | Use |
| --------------------------------------------------- | ------------------------------------------------------- |
| Start work at a specific time | Cron |
| Keep an existing session working | [Keeping Agents Running](Keeping-Agents-Running) |
| Drive one goal to completion across phases | [Autonomous Loops](Autonomous-Loops) |
There is also an older, deliberately separate `ScheduledRun` concept behind
`/api/scheduled`: a run-now, duration-bounded loop with no recurrence and no saved jobs. The
two systems never interact, and Cron is the one you want.
## From the API
```bash
API=http://localhost:3000
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{
"name": "nightly-deps",
"agentType": "claude",
"workingDir": "/home/me/proj",
"promptMode": "inline_text",
"promptText": "Update dependencies and open a PR",
"inputMode": "typed",
"scheduleType": "daily",
"dailyTime": "03:00",
"enabled": true,
"concurrencyPolicy": "warn_only"
}' | jq
curl -s "$API/api/cron/jobs" | jq
curl -s -X POST "$API/api/cron/jobs/<jobId>/run" | jq
curl -s "$API/api/cron/jobs/<jobId>/runs" | jq
```
Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set, and `-k` with the `https://` URL
on an HTTPS install.
## Gotchas
- **Times are the server's, not yours.** Obvious until you are travelling.
- **A `pi` job starts slowly.** The readiness poll looks for markers pi does not print, so it
burns its poll budget before sending the prompt. The job still works.
- **A deleted working directory fails the run**, by design, rather than creating a session
somewhere unexpected.
- **Auto-close only touches sessions this job created.** Your own tabs are never closed.
## Read next
- [Keeping Agents Running](Keeping-Agents-Running) - continuing work rather than starting it.
- [Notifications And Approvals](Notifications-And-Approvals) - hearing about a job that got stuck.
- [`docs/cron-guide.md`](https://github.com/Ark0N/Codeman/blob/master/docs/cron-guide.md) - the complete reference, including the API and SSE events.
+177
View File
@@ -0,0 +1,177 @@
# Docker Cases
Run a case inside its own container instead of directly on your host: for isolation, for a
reproducible toolchain, and for the ability to pick the whole environment up and move it to
another machine.
A docker case is a **location overlay**, not a run mode. All seven run modes work inside a
container. See [Core Concepts](Core-Concepts).
## One-time setup: the base image
The container needs an image carrying the agent toolchain (node, the CLIs, git, tmux). It
builds itself on first use with progress streamed to the UI, or you can build it ahead of
time:
```bash
node scripts/build-agent-image.mjs --no-cache
```
**Always pass `--no-cache`.** The CLIs are installed in a single `npm install -g` layer, so
a plain rebuild reuses that layer from the cache and the CLIs stay frozen at whatever
versions the image was *first* built with. This has shipped a broken CLI while reporting a
successful build.
A zero exit code proves the layers ran, not that the toolchain works. Verify:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
The image is secret-free. Credentials are delivered at runtime, never baked in, so exports
never leak them. A full image lands around 1.6GB.
Prerequisite: Docker or Podman with a reachable daemon.
## The quick way
On **Add Case → Create New**, tick **🐳 Run in an isolated Docker container**. That alone is
enough: Codeman creates the case folder, spins up a hardened container with sensible
defaults, and starts the session inside it.
Expanding **Container settings** offers a template:
| Template | Memory | CPUs | GPUs |
| ----------------- | ------ | ---- | ------------------------------------- |
| Small | 2 GB | 1 | none |
| Medium (default) | 4 GB | 2 | none |
| Large | 8 GB | 4 | none |
| GPU | 8 GB | 4 | all (needs the NVIDIA container toolkit) |
Disk is elastic: storage grows as data arrives, bounded only by host disk. Changing any
setting creates a dedicated host profile for that case, so it never mutates the shared
default.
## The full way
**Add Case → Docker** exposes everything:
| Field | Meaning |
| -------------------- | ----------------------------------------------------------------------------------------------- |
| **Case name** | As usual. |
| **Workspace path** | A real host directory, bind-mounted into the container at the **same absolute path**. |
| **Host ID** | A reusable profile (image, network, resources). Share one across cases to share settings. |
| **Network** | `bridge` (internet on, default), `none` (fully isolated), or a custom bridge. |
| **Advanced** | Memory and CPU caps, host credential seeding, and whether to resume the last conversation on relaunch. |
The same-absolute-path bind mount is what keeps the File Viewer, attachments, and watchers
operating on real host bytes rather than a copy.
## One container per case
Exactly one long-lived container per case, shared by every session in it.
- Killing one session kills only that session's in-container tmux. Siblings keep running and
the container stays up.
- Reconnecting after a Codeman restart lands back in the same live agent.
- A container stop or a host reboot restarts the container and **resumes the last
conversation** from the bind-mounted transcript.
- Deleting the case removes the container. The workspace on the host survives.
## Credentials
Your existing host logins work inside the container without logging in again. Credentials
are **seeded**: mounted read-only and copied in once at launch, so in-container CLIs never
write refreshed tokens back to your host credential stores. Onboarding and trust prompts are
pre-answered so no wizard appears.
Turn seeding **off** for a sealed sandbox: no host credentials, and with `network: none`, no
outbound access either. That is the profile for genuinely untrusted work; you log in inside
the container instead.
Bind mounts are excluded from image capture, so exports stay secret-free.
One consequence worth knowing: Pi's credentials are seeded per file rather than as a whole
directory, because that directory also holds sessions, extensions, and installed packages,
which can be gigabytes. So in-container Pi sessions are invisible from the host, and `pi -c`
inside a docker case sees only that container's history.
## Isolation
Every container runs hardened by default:
- `--cap-drop ALL`
- `--security-opt no-new-privileges`
- Non-root, running as your host uid so workspace files stay host-owned
- PID limit, memory cap with swap pinned to it, `--init`
- **Never** `--privileged`, and **never** the docker socket
Rootless engines without cgroup-v2 systemd delegation cannot enforce resource caps; linking
such a host warns that the caps are advisory.
## Configuration drift is refused, not ignored
Editing a docker host's configuration (image, memory, network) after a container exists is
detected on the next launch by comparing a configuration hash against the container's label.
A mismatch **refuses the launch** and offers to recreate rather than silently running with
stale configuration.
Recreating is refused while sessions of that case are live. The workspace and the
conversation both survive it.
## Moving a case to another machine
**Export**, from the Docker tab:
| Option | Contents |
| -------------------------- | ------------------------------------------------------------------------- |
| **Full image + workspace** | The whole toolchain, installed packages, and files, in one `.tgz`. |
| **Workspace only** | Just the project files. Fast and small. |
The container is paused across the capture so image and workspace are consistent, free space
is checked first, and the intermediate image is cleaned up. Exports run in the background
and notify you when the bundle is ready.
**Import** on the other machine: copy the `.tgz` into `~/.codeman/docker-exports/` and
import it into a new case. The manifest and per-member checksums are verified, the workspace
tar is extracted with a traversal guard, and the image is loaded under a **quarantined tag**
so it can never overwrite a local image. The destination supplies its own credentials, so
nothing secret crosses machines.
## Hooks need to reach the server
In-container hooks (permission events, idle and stop notifications) call back to Codeman
over the docker bridge gateway. If Codeman binds **loopback only**, which is the default and
the production configuration, the container cannot reach it and **in-container hooks do not
fire**.
The session still works fully: idle detection falls back to output-based detection through
the exec PTY, and with permission prompts skipped there is nothing to forward anyway.
To enable them:
```bash
CODEMAN_DOCKER_BRIDGE_HOOKS=1
```
Codeman then starts a second listener bound to the docker bridge gateway that serves **only**
the hook endpoints and rejects everything else with a 403. The bridge is host-internal, so
this does not widen your network exposure. Add it to the service unit and restart.
## Limits
- Per-session environment overrides, effort, and per-CLI configuration are **rejected** for
docker cases, because they do not cross into the container. Configure the container through
the docker host's per-mode command override instead.
- tmux must exist in the base image. It is a hard prerequisite and is probed when linking a
host.
- On macOS, Docker Desktop takes a dedicated uid path, and memory caps are subject to the
VM's own ceiling.
## Read next
- [Core Concepts](Core-Concepts) - why this is an overlay rather than a run mode.
- [Security](Security) - where containers fit in the model.
- [Remote SSH Sessions](Remote-SSH-Sessions) - the other overlay.
- [`docs/docker-cases.md`](https://github.com/Ark0N/Codeman/blob/master/docs/docker-cases.md) - the full reference.
+188
View File
@@ -0,0 +1,188 @@
# Driving Codeman From An Agent
Everything the dashboard does is HTTP, so an agent can do it too. This page is for the case
that makes Codeman interesting: **Claude Code running inside a Codeman session, spawning and
supervising other sessions.**
Two routes. Start with the skill.
## The agent skill
A Claude Code skill that teaches the agent the whole API, so you ask in plain English
instead of pasting endpoint documentation into prompts.
### Install it
| How | Command | Scope |
| ------------ | ----------------------------------------------------------- | ----------------------------------------------------------- |
| Skills CLI | `npx skills add Ark0N/Codeman --skill codeman -g` | Global, any skills-aware agent. |
| Bundled CLI | `codeman skill install` | Global, at `~/.claude/skills/codeman`. |
| Bundled CLI | `codeman skill install --case <name>` | One case. |
| Web UI | **App Settings → Agents & CLIs → Claude → Agent Skill** | Injects into each case when a Claude session is created. Off by default. |
`codeman skill uninstall [--case <name>]` reverses the CLI installs, and never touches a
`skills/codeman` you wrote yourself.
### Then just ask
| You say | What happens |
| --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| "What sessions are running right now?" | Lists them with name, mode, and status. Read-only. |
| "Start a shell worker on the `myapp` case, run the test suite, tell me if it passes." | Spawns, waits on a completion marker, reads the exit code, cleans up. |
| "Spin up 3 workers for lint, typecheck and tests, run them in parallel, report failures." | One session per task, all started first, then gathered as each finishes. |
| "Have a claude worker summarize `src/session.ts`, then close it." | Spawns, runs the readiness ladder, sends and waits, reads the answer, deletes the session. |
| "Watch session w4 and tell me if it gets stuck on a permission prompt." | Blocks on the `blocked` signal and surfaces the question to **you**. |
Sessions the agent creates get deleted when it is done. You can watch the tabs appear and
disappear in the dashboard while it works.
### What it will and will not do
- **It self-gates.** Outside a Codeman session it refuses to act and does not guess an API
URL, so a global install costs an unrelated Claude Code session nothing.
- **Unprompted, it may only** spawn sessions, prompt them, and delete ones **it created in
that conversation, by exact id**, behind a guard that refuses to delete the agent's own
session.
- **It will not** answer another session's permission prompt on your behalf. It surfaces the
question instead.
- **Deleting a case** (which erases a real directory of your code), bulk kills, respawn,
Ralph, cron, orchestrator, and settings writes all require you to ask, naming the target.
Turning the setting back off **does not remove already-injected copies**, because a
create-time sweep would yank the skill out from under other live sessions sharing that
directory. Remove them per case with `codeman skill uninstall --case <name>`.
The skill ships with the verb index always loaded, plus on-demand references for the verbs,
worked multi-worker recipes, endpoint tables, and cross-session messaging.
## The manual path
The same operations as raw HTTP, for a CI bot, a shell script, or an agent without skill
support.
### Detect that you are inside Codeman
These are set in every managed session. Read them rather than hardcoding anything:
| Variable | Meaning |
| -------------------------- | ------------------------------------------------------------------------ |
| `CODEMAN_MUX=1` | You are in a managed tmux session. Never `tmux kill-session`, `pkill claude`, or `pkill tmux`: you will kill yourself or a sibling. |
| `CODEMAN_API_URL` | Base URL, with the correct scheme. |
| `CODEMAN_SESSION_ID` | Your own session id. Use it to avoid acting on yourself. |
| `CODEMAN_HOOK_SECRET_FILE` | Path to the hook secret. |
### Rules of the road
Read these before writing any code. Each one has cost somebody an afternoon.
1. **Input is single line and must end with `\r`.** Enter fires only when the payload
contains a carriage return. Without it the text sits unsubmitted on the prompt, the
request still succeeds, and a combined wait burns its full timeout on a turn that never
started. Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs
the joined `echo Aecho B`. One line per call.
2. **Make input idempotent.** Send a stable `clientId` and a monotonic per-session `seq`. The
server deduplicates, so a retry after a dropped connection cannot double-deliver.
3. **Auth.** With `CODEMAN_PASSWORD` set, use HTTP Basic or the session cookie. A missing
`Origin` is allowed, so plain curl works. A `401` replies with the bare string
`Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error
instead of showing the failure. Check the status before parsing.
4. **Envelope.** Most endpoints return `{ "success": true, "data": ... }`. A few legacy GETs
return bare bodies, so handle both: `body.data ?? body`.
5. **Wait instead of polling, and a timeout is not an error.** The wait endpoints answer
`200` with `wait.timedOut: true`. Loop over short waits rather than one long call, because
tunnels cut idle connections.
6. **Only `claude` sessions emit `stop` and `blocked`.** They come from Claude Code hooks.
Shell and the external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
explicitly there is a `400`, while omitting `until` is always safe. On a shell session
`idle` fires **once at startup and never again**, so synchronize hook-less sessions with an
output marker instead.
7. **Nothing reports "ready", so wait for it explicitly.** A new session answers
`{"signal":"exit","immediate":true}` until its PID exists, and that means *not started*,
not *crashed*. A Claude worker in a fresh case then sits on the CLI's trust dialog; prompt
it there and the wait resolves on idle in about two seconds looking exactly like a finished
turn, while your text sits stuck in the dialog.
### Recipes
```bash
API="${CODEMAN_API_URL:-http://localhost:3000}"
# Add -u admin:"$CODEMAN_PASSWORD" if a password is set, and -k on an HTTPS install.
# What is running
curl -s "$API/api/sessions" | jq '.data[] | {id, name, mode, status}'
# Spawn a worker in a case
curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"myapp","mode":"shell"}' | jq
# Send a prompt (note the \r)
curl -s -X POST "$API/api/sessions/$ID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","clientId":"my-agent","seq":1}' | jq
# Send and block until the turn finishes (registers the wait BEFORE writing)
curl -s -X POST "$API/api/sessions/$ID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"summarize src/session.ts\r","wait":["stop"],"waitTimeout":120000}' | jq
# Or wait for a marker in the output, which works on shell sessions too
curl -s "$API/api/sessions/$ID/wait-output?contains=DONE_17909&from=buffer" | jq
# Read the terminal back
curl -s "$API/api/sessions/$ID/terminal?tail=4000" | jq -r '.data.output'
# Clean up, by exact id
curl -s -X DELETE "$API/api/sessions/$ID" | jq
```
Use `POST /api/quick-start` rather than `POST /api/sessions` when a case might be remote:
the plain create endpoint validates the working directory locally and has no case concept.
### The split-marker trick
For hook-less sessions, synchronize on a marker in the output. The catch: your own
keystrokes echo into the output stream, so an unsplit marker matches **before the command
has run**.
Split it so the typed line never contains the string you are waiting for:
```bash
M=DONE; R=17909
# typed: echo ${M}_${R} → output contains DONE_17909, the typed line does not
```
Make it unique per call, because tmux repaints replay old screen text.
### Reading output
Use `terminal?tail=`, not `/output`. The latter's text field is empty for every tmux-backed
session, which is every interactive session. `tail` counts **bytes**, and what comes back is
terminal data with ANSI sequences included.
## Fan-out, and why it needs care
Wait signals are **edge triggered with no history**. A signal that fires with no waiter
registered is unobservable afterwards.
So a fan-out must register its waits before or as it dispatches: use send-and-wait per
worker, or latched output markers. Dispatching all the workers and then waiting on them one
at a time loses the signals of everyone who finished early.
Send-and-wait registers the waiter **before** the write for the same reason. A separate POST
followed by a wait races, and reports the previous turn's state.
## Lineage
A create request can name the session that spawned it, through a body field or a header, and
the dashboard then draws a lineage arc from parent to child. The skill sets it automatically.
It is resolved rather than trusted: an unresolvable parent is dropped silently rather than
failing the spawn, because a cosmetic field must never break a worker.
## Read next
- [HTTP API](HTTP-API) - the endpoint map and the envelope.
- [Hooks And Integrations](Hooks-And-Integrations) - events flowing the other way.
- [Watching Agents Work](Watching-Agents-Work) - seeing the fan-out in the UI.
- [`skills/codeman/SKILL.md`](https://github.com/Ark0N/Codeman/blob/master/skills/codeman/SKILL.md) - the skill itself.
+250
View File
@@ -0,0 +1,250 @@
# FAQ
The questions that keep arriving in
[Discussions](https://github.com/Ark0N/Codeman/discussions) and issues. For "why is it
doing that", go to [Troubleshooting](Troubleshooting) instead.
## The basics
### What is Codeman, in one sentence?
A self-hosted dashboard that runs AI coding agents in persistent tmux sessions on your own
machine and lets you drive them from any browser, including a phone.
### Is it free? What is the licence?
MIT, free, and open source. There is no paid tier and no account.
### Do I need an API key?
No. Codeman drives agent CLIs you have already installed and logged in yourself. Whatever
subscription or key that CLI uses is what pays for the tokens. Codeman never collects,
stores, or refreshes your credentials.
### Does Codeman send my code or prompts anywhere?
No. There is no telemetry, no analytics, and no phone-home. The only network traffic
Codeman itself makes is between your browser and your server.
Your agent CLI is a separate matter: Claude Code talks to Anthropic, Codex talks to OpenAI,
and so on. That traffic is the CLI's, on your own account, exactly as it would be in a
terminal.
Two features do send data outward, both off by default and both stated where they appear:
voice dictation through your own Claude login, and the Read My Mind prediction call.
### Does it work on Windows?
Through WSL2. Codeman requires tmux. Install it inside WSL, run your agent CLI inside WSL,
and `http://localhost:3000` works from your Windows browser. Work in the Linux filesystem
rather than `/mnt/c/...`, which is dramatically slower for file watching and git.
### Is there a mobile app?
The web UI is built for phones and installs as a PWA. There is no App Store or Play Store
app.
## Sessions and persistence
### Do my agents keep running when I close the browser?
Yes. Agents run in tmux on the server, not in your browser. Close the tab, close the laptop,
lose the network. When you come back, the session is still there with its scrollback.
The same holds when the Codeman server itself restarts. What does end a session is killing
the tmux server or rebooting the machine.
### What happens after a reboot?
tmux dies with the machine, so the sessions are gone. Conversations are not: Claude
transcripts persist on disk, and the welcome screen's **Resume Conversation** list picks
them back up. Install Codeman as a service and the server itself comes back on boot.
### How many sessions can I run at once?
The design target is 20 sessions and 50 agent windows at 60fps. The hard cap is higher, and
what you will actually hit first is the CPU and memory of the machine running the agents.
### Can I run Claude Code and Codex side by side?
Yes, that is a normal setup. The run mode is per session, so one case can have a Claude tab,
a Codex tab, and a shell tab open at the same time, each with its own colour. Some Codeman
features are Claude-only; [Agent CLIs](Agent-CLIs) lists exactly which.
### Can I attach to a session from a terminal instead of the browser?
Yes. `sc` is an interactive chooser (`sc 2` attaches directly, `sc -l` lists), or use tmux
directly on the `codeman` socket. Detach with `Ctrl+A D`.
## Running unattended
### I hit my Claude usage limit overnight. Can Codeman resume automatically?
Yes, and it is the reason the feature exists. Turn on auto-resume at the top of the Respawn
tab for that session. When Claude halts on a subscription limit, Codeman parses the reset
time from the message, waits until two minutes past it, and continues the conversation.
Respawn cycles are blocked while a session is limit-paused, which is what stops a `/clear`
from wiping the conversation you are waiting to resume. Claude-only.
### Will it keep prompting my agent forever?
Only if you configure it to. Respawn cycling is per session and off unless you turn it on,
and it has presets ranging from a 60 minute solo session to an 8 hour overnight run. There
are circuit breakers to stop a thrashing session from spinning indefinitely. See
[Keeping Agents Running](Keeping-Agents-Running).
### Does an idle session cost tokens?
No. An idle agent is a process waiting for input. Tokens are spent when a turn runs, so what
costs money is the re-prompting you configured, not the session sitting there.
### Can I schedule work for a specific time?
Yes. [Cron Jobs](Cron-Jobs) saves named jobs on a `once`, `interval`, `daily`, or `weekly`
schedule; each spins up a session and sends a prompt when due, with per-job run history.
## Access
### How do I reach Codeman from my phone when I am away from home?
Tailscale is the recommended answer: your devices join a private network, Codeman keeps its
loopback bind, and you get real HTTPS. The installer sets it up, and `install.sh tailscale`
retrofits it onto an existing install.
A Cloudflare tunnel gives a public URL faster, and requires `CODEMAN_PASSWORD`. Full
comparison in [Remote Access](Remote-Access).
### Why can't other devices reach Codeman?
Because the default bind is `127.0.0.1`, on purpose. Codeman starts agents with permission
prompts skipped, so whoever reaches the dashboard can run code on your machine. Exposing it
is a deliberate step, and [Remote Access](Remote-Access) covers the safe ways.
### My reverse proxy domain is rejected with `403 host not allowed`
The always-on Host-header allowlist blocks DNS rebinding, and it does not know your domain.
Add it:
```bash
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
A leading dot matches subdomains. Also make sure the proxy forwards WebSocket upgrades.
### Do I have to type a password on my phone?
No. Scan the QR code shown on the desktop dashboard. Tokens are single use and rotate every
60 seconds. The password remains the fallback.
## Multiple people, multiple instances
### Can several people share one Codeman?
Yes, with `codeman web --multiuser`. Each person gets a login and their own case space, and
sessions, cases, search, and events are scoped to their owner.
Be clear about what that is: it separates **workspaces**, not operating system accounts.
Every session still runs as the same OS user, so a determined user's agent can reach another
user's files. For real isolation, pair users with Docker cases or run separate instances
under separate OS accounts. See [Multi-User Mode](Multi-User-Mode).
### How do I run a second instance, a beta beside my main one?
Give it its own instance name, which scopes the data directory and the tmux socket together:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
Do not skip this. The data directory and tmux socket are process wide, so a second server on
the defaults discovers and attaches your live sessions.
## Updating and maintenance
### What is the right way to update Codeman?
| Install route | Update with |
| ------------- | --------------------------------------------------------------------------- |
| Installer | Re-run the install one-liner, or **App Settings → System → Updates**. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart the service. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It stashes a
dirty tree rather than discarding it, and streams progress across the restart. npm installs
report as non-updatable.
### Will updating kill my running sessions?
No. Sessions live in tmux, so restarting the server reattaches to them.
### Where is my data?
Everything under `~/.codeman/`, with cases created from scratch in `~/codeman-cases/`.
Nothing needs root and nothing leaves the machine. Uninstalling does not delete either
directory.
## Features
### What is the difference between respawn, Ralph, and the orchestrator?
- **Respawn** restarts a session's CLI when it goes idle, to keep a long run going. It is the
one most people want.
- **Ralph loop** is an autonomous single-session task loop with its own tracker.
- **Orchestrator** turns one goal into a phased plan and drives it across agents.
[Keeping Agents Running](Keeping-Agents-Running) and [Autonomous Loops](Autonomous-Loops)
cover them properly.
### Can agents start and supervise other agents?
Yes. Codeman ships an agent skill that lets an agent inside a session drive the HTTP API:
list sessions, spawn workers, send prompts, and block until a worker's turn finishes. It is
off by default and enabled per case.
See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
### Can I run a case in a container?
Yes. One container per case, shared by all its sessions, non-root and capability-dropped by
default, with your host CLI logins seeded in so nothing asks you to log in again. You can
export a container plus its workspace and move it to another machine. See
[Docker Cases](Docker-Cases).
### Can the agent run on a different machine?
Yes. Point a case at a remote host over SSH and the agent runs there, inside a durable
remote tmux, so a dropped connection does not kill the run. See
[Remote SSH Sessions](Remote-SSH-Sessions).
### Can I put my Grafana or other dashboards in here?
Yes. Saved URLs render as tabs beside your sessions, proxied through Codeman's own origin so
that mixed content and frame-blocking headers do not break them. See [Web Tabs](Web-Tabs).
### Why is a feature I read about not on screen?
Most of Codeman's UI is opt-in and defaults to off, so a stock install stays small. Check
**App Settings → Header & Panels**. [Settings Reference](Settings-Reference) lists the
defaults.
## Contributing
### How do I request a feature?
Open an [Idea](https://github.com/Ark0N/Codeman/discussions/categories/ideas) and it gets
voted on. Roadmap decisions happen there.
### How do I contribute code?
[CONTRIBUTING.md](https://github.com/Ark0N/Codeman/blob/master/.github/CONTRIBUTING.md) has
the full map. Small fixes can go straight to a PR; anything larger starts as an issue or
Discussion so the design gets a nod first. Skins, translations, and docs are good first
contributions.
### How do I fix a mistake in this wiki?
These pages are generated from
[`docs/wiki/`](https://github.com/Ark0N/Codeman/tree/master/docs/wiki) in the main
repository. Editing a page in the browser gets overwritten on the next sync, so send a PR
against that directory instead.
+164
View File
@@ -0,0 +1,164 @@
# HTTP API
Codeman's HTTP and SSE API is a **stable contract**. Everything the dashboard does goes
through it, so anything the dashboard can do, a script can do.
This page is the orientation. The complete specification, including every wait semantic and
the SSE catalogue, is
[`docs/api-reference.md`](https://github.com/Ark0N/Codeman/blob/master/docs/api-reference.md).
## What is stable
Covered by semantic versioning: endpoint paths under `/api/v1`, the response envelope,
`errorCode` values, and SSE event names.
Not covered, and free to change in a patch release: on-disk state files, internal modules,
and anything marked experimental. The full statement is in
[Versioning](Versioning).
`/api/v1/*` is a versioned alias of `/api/*`. Prefer the versioned form in anything you
intend to keep.
## The envelope
```json
{ "success": true, "data": { } }
```
```json
{ "success": false, "error": "human readable", "errorCode": "NOT_FOUND" }
```
A few legacy GET handlers return bare bodies rather than the envelope, so a robust client
reads `body.data ?? body`.
Branch on `errorCode`, which is stable. The HTTP status is reliable too:
| `errorCode` | HTTP | Meaning |
| ------------------ | ---- | ------------------------------------------------ |
| `INVALID_INPUT` | 400 | Malformed request or failed validation. |
| `UNAUTHORIZED` | 401 | Authentication required or failed. |
| `NOT_FOUND` | 404 | No such resource. |
| `SESSION_BUSY` | 409 | The session is busy. |
| `CONFLICT` | 409 | Conflicts with current state. |
| `ALREADY_EXISTS` | 409 | Resource already exists. |
| `OPERATION_FAILED` | 422 | Well formed, could not be completed. |
| `RATE_LIMITED` | 429 | Too many requests. |
| `INTERNAL_ERROR` | 500 | Unexpected server error. |
New error codes are non-breaking. Removing or renaming one is a major change.
**A `401` is the bare string `Unauthorized`, not the envelope.** Piping it into `jq` throws
a parse error rather than showing the failure, so check the status first.
## Authentication
With no password set, and the default loopback bind, there is none. With `CODEMAN_PASSWORD`
set, use HTTP Basic or the session cookie:
```bash
curl -s -u admin:"$CODEMAN_PASSWORD" "$API/api/sessions"
```
A **missing** `Origin` header is allowed, so curl and CLI tools work unchanged. A
present-but-foreign origin is rejected by the CSRF guard. On an HTTPS install with the
self-signed certificate, add `-k`.
## Endpoint map
Roughly 200 handlers across 24 route modules. By domain:
| Domain | Handlers | Covers |
| ------------------- | -------- | --------------------------------------------------- |
| System | 45 | Status, settings, search, digest, updates. |
| Sessions | 34 | Create, input, terminal, wait, kill. |
| Cases | 29 | Create, link, clone, remote and docker cases. |
| Files | 16 | Preview, edit, raw, attachments, path picker. |
| Orchestrator | 10 | Plans and phases. |
| Ralph | 9 | Loop control and configuration. |
| Cron | 9 | Jobs and run history. |
| Admin | 8 | Multi-user administration. |
| Plan | 8 | Plan orchestration. |
| Respawn | 7 | Respawn configuration and presets. |
| Webviews | 6 | Saved dashboards, plus the proxy. |
| Mux | 5 | tmux operations. |
| Push | 4 | Web push subscriptions. |
| Read My Mind | 4 | Intent profiles and prediction. |
| Scheduled | 4 | The legacy scheduled-run concept. |
| Approvals | 3 | The inbox and answering. |
| Teams, me, search, hooks, clipboard, telemetry, voice, ws | 1-2 each | |
Each route module documents its own endpoints in its file header.
## Long-polling instead of polling
Three calls block until something happens, so an agent driving Codeman from a shell can wait
rather than spin:
| Call | Blocks until |
| ----------------------------------------- | -------------------------------------------------------- |
| `GET /api/v1/sessions/:id/wait` | One of a set of lifecycle signals fires. |
| `GET /api/v1/sessions/:id/wait-output` | A literal string appears in the session's output. |
| `POST /api/v1/sessions/:id/input` + `wait`| The input is delivered **and then** a signal fires. |
Three semantics that break callers who assume otherwise:
1. **A timeout is `200`, not an error.** It answers with `wait.timedOut: true`. Loop over
short waits; a single long call gets cut by tunnels and proxies.
2. **Send-and-wait is not a POST followed by a wait.** It registers the waiter *before*
writing, which closes the window where a separate wait sees the session still idle from
the previous turn and answers instantly about the wrong turn.
3. **Signals are edge triggered with no history.** One that fires with no waiter registered
is unobservable afterwards. Fan-outs must register their waits as they dispatch.
`wait-output` matches a **literal substring, never a regex.** That is deliberate: no regex
means no catastrophic backtracking on attacker-influenced output.
Only `claude` sessions emit `stop` and `blocked`, because those come from Claude Code hooks.
Shell and external CLI sessions accept `idle`, `working`, and `exit`.
## SSE
`GET /api/events` is the live event stream. 155 event names, kept in sync between server and
client with a test that fails on drift.
The heartbeat is a **named** `sse:heartbeat` event rather than an SSE comment, because
comments are invisible to `EventSource` by specification and a client could not observe
them. That is what lets the browser detect a stream that has silently stopped delivering.
```js
const es = new EventSource('/api/events');
es.addEventListener('session:created', (e) => console.log(JSON.parse(e.data)));
```
## Quick examples
```bash
API="${CODEMAN_API_URL:-http://localhost:3000}"
curl -s "$API/api/status" | jq # whole-system snapshot
curl -s "$API/api/sessions" | jq '.data[].name' # live sessions
curl -s "$API/api/sessions/unified" | jq # live + historical, deduped
curl -s "$API/api/subagents" | jq # background agents
curl -s "$API/api/search?q=deploy" | jq # cross-session search
```
## Limits
| Limit | Default |
| --------------------- | ------------------------------------------ |
| Max sessions | 50 |
| Max agent windows | 500 |
| Max SSE clients | 100 |
| Terminal buffer | 32 MB per session |
| Text payload | 1 MB |
| Wait timeout ceiling | 600 s, and the response tells you what was applied |
Most are environment-overridable. See `src/config/`.
## Read next
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the practical version, with recipes.
- [Hooks And Integrations](Hooks-And-Integrations) - events flowing back into Codeman.
- [Versioning](Versioning) - what the version number promises.
- [`docs/api-reference.md`](https://github.com/Ark0N/Codeman/blob/master/docs/api-reference.md) - the full specification.
+136
View File
@@ -0,0 +1,136 @@
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/codeman-title.svg" alt="Codeman" height="56">
</p>
<h3 align="center">Mission control for AI coding agents</h3>
Codeman runs your coding agents on your own machine and puts them behind one dashboard you
can open from any device. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or
Pi inside persistent tmux sessions, streams the real terminal to the browser, and keeps
working while you are away from the keyboard: it re-prompts idle agents, resumes when a
subscription limit resets, runs jobs on a schedule, and shows every background subagent
live.
This wiki is the manual. The [README](https://github.com/Ark0N/Codeman) is the overview,
and the deep internals live in
[`docs/`](https://github.com/Ark0N/Codeman/tree/master/docs).
```bash
curl -fsSL https://getcodeman.com/install | bash
codeman web # then open http://localhost:3000
```
---
## Start here
**New to Codeman**
1. [Installation](Installation) - requirements, the installer, npm and git clone routes, updating.
2. [Quick Start](Quick-Start) - from a running server to a working agent in five minutes.
3. [Core Concepts](Core-Concepts) - cases, sessions, run modes, and what survives a restart.
4. [The Dashboard](The-Dashboard) - reading the tab strip, the status dots, and the alerts.
**Already running it**
- [Agent CLIs](Agent-CLIs) - the seven run modes, their setup, and which features are Claude-only.
- [Mobile Guide](Mobile-Guide) - phone and tablet use, QR login, the touch keyboard bar.
- [Remote Access](Remote-Access) - Tailscale, Cloudflare tunnel, LAN plus password, QR login.
- [Keeping Agents Running](Keeping-Agents-Running) - idle detection, respawn cycling, auto-resume on usage limits.
- [Troubleshooting](Troubleshooting) - symptom-first index of things that actually break.
**Driving it from code**
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the bundled skill, worker sessions, wait primitives.
- [HTTP API](HTTP-API) - the envelope, auth, the endpoint map, SSE events.
- [Hooks And Integrations](Hooks-And-Integrations) - events flowing back into Codeman.
---
## Everything in the manual
### Getting started
| Page | What it answers |
| ------------------------------- | --------------------------------------------------- |
| [Installation](Installation) | How do I install it, update it, and remove it? |
| [Quick Start](Quick-Start) | How do I get one agent working right now? |
| [Core Concepts](Core-Concepts) | What is a case, a session, a run mode? |
### Using it
| Page | What it answers |
| ------------------------------------------ | ---------------------------------------------------------- |
| [The Dashboard](The-Dashboard) | What is the UI telling me? |
| [Agent CLIs](Agent-CLIs) | Which agent should this session run, and how do I set it up? |
| [Working With Files](Working-With-Files) | How do I read, edit, and attach files? |
| [Input And Voice](Input-And-Voice) | How do I talk to an agent, including by voice? |
| [Mobile Guide](Mobile-Guide) | How well does this work on a phone? |
| [Keyboard Shortcuts](Keyboard-Shortcuts) | What can I drive from the keyboard? |
| [Settings Reference](Settings-Reference) | What does this setting do, and why did it not follow me to my phone? |
### Keeping agents running
| Page | What it answers |
| ------------------------------------------------------------- | --------------------------------------------------- |
| [Keeping Agents Running](Keeping-Agents-Running) | How does it run unattended overnight? |
| [Notifications And Approvals](Notifications-And-Approvals) | How do I know an agent needs me, and answer from my phone? |
| [Cron Jobs](Cron-Jobs) | How do I run an agent on a schedule? |
| [Autonomous Loops](Autonomous-Loops) | What are the Ralph and Orchestrator loops for? |
| [Watching Agents Work](Watching-Agents-Work) | How do I see what the subagents are doing? |
### Where it runs
| Page | What it answers |
| --------------------------------------------- | -------------------------------------------- |
| [Docker Cases](Docker-Cases) | How do I sandbox a project in a container? |
| [Remote SSH Sessions](Remote-SSH-Sessions) | How do I run the agent on another machine? |
| [Web Tabs](Web-Tabs) | Can my Grafana live in here too? |
| [Multi-User Mode](Multi-User-Mode) | Can several people share one Codeman? |
### Access and security
| Page | What it answers |
| ------------------------------- | ---------------------------------------------------------- |
| [Remote Access](Remote-Access) | How do I reach it from outside this machine, safely? |
| [Security](Security) | What is exposed, what protects it, what do I have to do? |
### Automation and integration
| Page | What it answers |
| ----------------------------------------------------------------- | -------------------------------------------- |
| [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) | How does an agent spawn and drive workers? |
| [HTTP API](HTTP-API) | What can I call, and what comes back? |
| [Hooks And Integrations](Hooks-And-Integrations) | How do I wire Codeman into something else? |
### Operating it
| Page | What it answers |
| --------------------------------------------- | -------------------------------------------------- |
| [Running As A Service](Running-As-A-Service) | How do I keep it up across reboots, and update it? |
| [Troubleshooting](Troubleshooting) | Why is it doing that? |
| [FAQ](FAQ) | The questions that keep coming up. |
| [Contributing](Contributing) | How do I send a fix? |
| [Versioning](Versioning) | What does the version number promise? |
---
## Requirements at a glance
| Thing | Needed |
| ------------ | --------------------------------------------------------------------------- |
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
machine.
## Getting help
- **Questions and setup help**: [Discussions](https://github.com/Ark0N/Codeman/discussions), especially [Q&A](https://github.com/Ark0N/Codeman/discussions/categories/q-a).
- **Bugs**: [Issues](https://github.com/Ark0N/Codeman/issues). Include your OS, install method, browser, and which CLI the session was running.
- **Ideas and roadmap**: [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas).
- **Security**: never a public issue. See [SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
+97
View File
@@ -0,0 +1,97 @@
# Hooks and Integrations
Events flowing **back** into Codeman, and the four seams a third party can build against.
## Hooks
Claude Code can run a command when something happens in a session. Codeman writes a hooks
configuration into each Claude case so those events post back to it, which is what turns a
terminal into something that can notify you.
| Event | Fires when | Drives |
| ---------------------- | ----------------------------------------------- | --------------------------------------------- |
| `permission_prompt` | The agent asks for permission. | Red tab alert, Approvals Inbox, push. |
| `idle_prompt` | The agent is waiting for input. | Yellow tab alert, the `idle` wait signal. |
| `stop` | A turn ends. | The `stop` wait signal, idle detection. |
| `elicitation_dialog` | A dialog opens. | Approvals Inbox. |
| `elicitation_complete` | The dialog closes. | Clearing the alert. |
| `elicitation_response` | The dialog is answered. | Clearing the alert. |
| `teammate_idle` | An agent-team member goes idle. | Team surfaces. |
| `task_completed` | A task finishes. | Task tracking, run summary. |
This is why several Codeman features are Claude-only. The other CLIs have no hook system, so
for them Codeman watches terminal output, which reveals that something happened but not what
it was.
### How hooks get installed
Codeman writes them into the case when a Claude session is created. Hook blocks are
**marker-owned**: Codeman only ever updates a block it wrote, and never touches
configuration you added yourself.
If tab alerts and approvals never fire in a particular case, that case is missing its hook
block. Recreating the case rewrites it.
### The hook secret
`/api/hook-event` and `/api/status-telemetry` skip HTTP Basic authentication, because they
are called from localhost by the CLI itself. When authentication is on, that bypass
additionally requires a per-instance hook secret, because Codeman cannot tell a genuine
loopback call from a request arriving through your own loopback reverse proxy.
The secret lives in the data directory, and its path is exported into every managed session.
### Two things that break hooks
- **HTTPS.** Hook callbacks must accept the self-signed certificate. Recent versions
self-heal existing cases; older cases need recreating.
- **Docker cases on a loopback bind.** A container cannot reach `127.0.0.1` on the host, so
in-container hooks silently do not fire. Set `CODEMAN_DOCKER_BRIDGE_HOOKS=1` to open a
hooks-only listener on the bridge gateway. See [Docker Cases](Docker-Cases).
## Integration seams
Codeman has **no plugin runtime**, and that is a decision rather than a gap. A plugin runtime
means running third-party code inside a process that spawns agents with your credentials, on
a server people routinely expose over a tunnel. Codeman's security posture is one of its
reasons to exist, so it does not trade that away for an extension mechanism.
What exists instead is four documented seams.
### 1. Web tabs
Anything with a web UI can live inside Codeman as a tab, proxied through Codeman's own
origin. The lowest-effort integration by a wide margin: if your tool has a dashboard, it can
sit beside the agents with no code at all. See [Web Tabs](Web-Tabs).
### 2. SSE events
`GET /api/events` streams everything Codeman knows: session lifecycle, output, agent
activity, approvals, cron runs. 155 named events, stable under semantic versioning.
This is the seam for anything that reacts. A bot that pings your chat channel when an agent
needs a human is a short script over this stream.
### 3. HTTP API and CLI
Everything the dashboard does. Create sessions, send input, block on wait primitives, read
terminals, manage cron. See [HTTP API](HTTP-API) and
[Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
### 4. Hooks
The seam above, in the other direction: your own hook commands can run alongside Codeman's
in a case, as long as you leave Codeman's marker-owned block alone.
## Publishing an integration
There is no registry to submit to. Share it in
[Show and tell](https://github.com/Ark0N/Codeman/discussions/300), and if it needs a change
in Codeman to work properly, open an issue or a Discussion first.
## Read next
- [HTTP API](HTTP-API) - the endpoint map and envelope.
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the agent-facing path.
- [`docs/extending-codeman.md`](https://github.com/Ark0N/Codeman/blob/master/docs/extending-codeman.md) - the seams in full, with examples.
- [`docs/claude-code-hooks-reference.md`](https://github.com/Ark0N/Codeman/blob/master/docs/claude-code-hooks-reference.md) - upstream hook semantics.
+135
View File
@@ -0,0 +1,135 @@
# Input and Voice
Getting words into an agent: typing, dictating, and letting Codeman guess. Plus the input
machinery that only shows up when it goes wrong.
## Typing
Click into the terminal and type. It is a real terminal, so everything the CLI supports
works, slash commands included.
| Key | Effect |
| ---------------------------- | --------------------------------------------- |
| `Enter` | Send. |
| `Shift+Enter` / `Ctrl+Enter` | Newline without sending. |
| `Ctrl+C` | Copy if text is selected, otherwise interrupt. |
| `Ctrl+Shift+C` | Copy, never interrupts. |
| `Ctrl+L` | Clear the terminal. |
### Exactly-once delivery
Browser input goes through a durable layer rather than a plain socket write. Each prompt
carries a stable client id and a per-session sequence number, held in local storage until
the server acknowledges it.
The result is the property you want on a phone: a connection that drops mid-prompt never
loses the prompt and never delivers it twice. Two browser tabs on the same session coexist,
and only a reconnect from the *same* tab supersedes the old connection.
## Zero-lag local echo
On touch devices, keystrokes are painted in the terminal immediately and sent when you press
Enter, instead of waiting for each character to round-trip to the server and back. Over a
mobile connection that is the difference between usable and not.
![Zero-lag input](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/zerolag-demo-20260728.gif)
The consequence to remember: **text on screen has not necessarily reached the agent yet.**
It is flushed on Enter. If a prompt appears to have been ignored, press Enter, or the phone
toolbar's **Enter** button.
Default on for touch devices, off for desktop, and switchable in
**App Settings → Terminal & Input**.
### Codex is different on purpose
Codex's composer reacts to every keystroke: `/` opens a live-filtering picker, arrows edit
state on its side, the composer grows as text wraps. Buffering until Enter starved it, so
Codex sessions use **predictive echo** instead: each keystroke is painted at its predicted
position while the bytes actually sent stay identical to what you typed. Predictions
reconcile against the real buffer and only apply while the cursor is on the composer row.
## CJK input
Chinese, Japanese, and Korean input needs an IME, and an IME needs a real text field.
Turning on CJK input in **App Settings → Terminal & Input** puts an always-visible textarea
below the terminal that owns composition, then delivers the composed text to the session.
## Voice dictation
`Ctrl+Shift+V`, or the microphone button. There are three providers and the default is
`auto`, which prefers them in this order:
| Provider | Needs | Notes |
| ------------------ | ---------------------------------------------- | ------------------------------------------------------------ |
| **Claude** | Claude Code logged in on the server. Opt-in. | Uses this machine's existing Claude login. No extra key. |
| **Deepgram** | A Deepgram API key. | Nova-3, with automatic silence detection. |
| **Web Speech** | Nothing. | Browser-provided, quality varies. |
### Dictating through your Claude login
Off by default; enable it in **App Settings → Voice**.
Claude Code has its own voice mode, but it opens the **host's** microphone, and in Codeman
the CLI runs headless in a tmux pane while you are in a browser somewhere else entirely. So
Codeman captures audio in your browser and borrows only the backend: audio goes browser to
Codeman to Anthropic, and the page never sees the OAuth token.
Two deliberate limits:
- **Credentials are read only.** Codeman never refreshes your Claude token, because a
refresh rotates the refresh token and could sign you out of your own CLI. An expired token
is reported as expired rather than silently renewed.
- **Capture is raw PCM** at 16 kHz mono, which requires an AudioWorklet rather than the
usual browser recorder.
## Read My Mind
**Claude only, off by default.** Turn it on in **App Settings**, and a 🧠 button appears in
the header (on phones, in the keyboard bar instead).
It keeps a per-case **intent profile**: goals you or your agent write down, plus the prompts
you actually submitted in that case. Pressing 🧠 feeds that profile plus live session signals
to a single model call and shows a predicted next prompt.
What you can do with the result:
- **Send** it, **Insert** it into the composer, or edit it first.
- Pick one of the alternate suggestions, which swaps into the editable field without losing
your edits.
- **Rethink**, optionally with a steer note, to reject the whole set and try again.
**Nothing is ever sent automatically.** Every path requires a click.
Where the data lives: the profile is keyed by owner and the resolved working directory, so
it survives `/clear` and respawns. Prompts can contain secrets, so the store is written
0600 and is deliberately excluded from cross-session search.
Guide: [`docs/readmymind.md`](https://github.com/Ark0N/Codeman/blob/master/docs/readmymind.md).
## Programmatic input
Sending prompts over the API has one rule that catches everyone: **the payload must end with
`\r`** or Enter is never sent. The request still succeeds, the text sits unsubmitted in the
composer, and any wait burns its whole timeout on a turn that never started.
Input is also **single line**. Embedded newlines are stripped rather than rejected, so
`"echo A\necho B\r"` runs the joined `echo Aecho B`. Put multi-line content in a file and
tell the agent to read it.
See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
## Gotchas
- **Typed text sitting on screen has not been sent.** Press Enter.
- **`Ctrl+C` with a selection copies.** Clear the selection to interrupt.
- **Voice needs HTTPS.** Microphone access requires a secure context, same as push
notifications.
- **Read My Mind goes blind for sessions using a relocated Claude config directory**, along
with the other transcript-backed features. See [Agent CLIs](Agent-CLIs).
## Read next
- [Mobile Guide](Mobile-Guide) - the keyboard bar and touch input.
- [Keyboard Shortcuts](Keyboard-Shortcuts) - the full list.
- [Working With Files](Working-With-Files) - images and attachments as input.
+234
View File
@@ -0,0 +1,234 @@
# Installation
Getting Codeman onto a machine, verifying it works, updating it, and removing it.
## Requirements
| Requirement | Notes |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
## Route A: the installer (recommended)
```bash
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
end to end.
- **Your local network** (`0.0.0.0`): prompts for a password. Skipping the password takes
an explicit confirmation and ends on a loud warning.
- **This machine only** (`127.0.0.1`): the safest option, and the default for a bare
`codeman web` regardless of what you pick here.
Which one is preselected depends on what the installer finds. A fresh install defaults to
the local network, unless Tailscale is already connected, in which case it defaults to
Tailscale. An existing loopback install defaults to keeping loopback, or to Tailscale when
a serve mapping for Codeman is already there. A bare Enter never pulls in new software,
and a non-interactive run always keeps the safe loopback default.
3. **What to do when it finishes.** Run in this terminal, install as a background service
that starts on boot, or do nothing yet.
Re-running the same one-liner **updates an existing install in place**. Local changes in
`~/.codeman/app` are stashed rather than discarded, a running service is restarted and
verified, and your existing network binding is preserved. An interrupted first install
resumes instead of restarting.
Two other entry points exist:
```bash
install.sh update # update only
install.sh uninstall # remove
install.sh tailscale # retrofit Tailscale access onto an existing install
```
**Automation and CI**: with no terminal attached, any step that would change the system
aborts with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to
approve those steps. `CODEMAN_TAILSCALE=1` preselects the Tailscale answer, and never
installs Tailscale itself non-interactively.
## Route B: npm
```bash
npm install -g aicodeman
codeman web
```
The npm package is named `aicodeman`; the product is Codeman. Both `codeman` and
`aicodeman` are installed as commands.
The trade-off against Route A: no guided network setup, and the in-app self-updater does
not apply. npm installs report as non-updatable in **App Settings → System → Updates**, and
you update with `npm update -g aicodeman`.
## Route C: git clone
For contributing, or for running unreleased code.
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
For a production run from a clone:
```bash
npm run build
npm run start
```
`npm run dev` runs TypeScript directly through `tsx` with no build step. The frontend is
plain JavaScript served from `src/web/public/` with no bundler, so editing a `.js` or `.css`
file and reloading the page is enough. The one exception is `index.html`, which is read once
at server start, so markup changes need a restart.
See [Contributing](Contributing) for the rest of the development loop.
## Installing an agent CLI
Codeman drives CLIs, it does not bundle them. Install at least one:
| CLI | Install | Notes |
| --------------- | ------------------------------------------------------------------ | -------------------------------------------------------------------------- |
| **Claude Code** | `npm i -g @anthropic-ai/claude-code` | The primary target. Some Codeman features are Claude-only: see [Agent CLIs](Agent-CLIs). |
| **OpenCode** | See [opencode.ai](https://opencode.ai) | |
| **Codex** | See [developers.openai.com/codex/cli](https://developers.openai.com/codex/cli) | |
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
## Verify the install
```bash
codeman doctor # checks Node, tmux, the agent CLIs, document converters
codeman --version
codeman web # then open http://localhost:3000
```
`codeman doctor --json` gives machine-readable output, and `--category core` narrows it to
the things a session cannot start without.
If the dashboard loads and **+ New Session** opens, you are done. Continue to
[Quick Start](Quick-Start).
## Where things live
| Path | What |
| ----------------------- | ------------------------------------------------------------------------------------------------ |
| `~/.codeman/app` | The installed code (installer route only). |
| `~/.codeman/` | All state: `state.json`, settings, session history, push keys, TLS certs. See [Core Concepts](Core-Concepts). |
| `~/codeman-cases/` | Cases created from scratch. Linked cases stay wherever they already are. |
| `~/.codeman/web.log` | Log for a detached (`-d`) server. |
Everything is under your home directory, and nothing needs root.
## Keeping it running
A bare `codeman web` dies with the shell that started it. Two ways to outlive that:
```bash
codeman web -d # detached; --status and --stop manage it
codeman service install # systemd user unit or macOS LaunchAgent; survives reboots
```
Full detail, including logs and the self-updater, is in
[Running As A Service](Running-As-A-Service).
## Updating
| Install route | How to update |
| ------------- | ----------------------------------------------------------------- |
| Installer | Re-run the one-liner, or **App Settings → System → Updates** in the UI. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It restarts
the process that is running it, so the actual work happens in a detached script and the
browser polls across the restart. Progress appears in the UI.
## Uninstalling
```bash
install.sh uninstall # installer route
npm uninstall -g aicodeman # npm route
```
Neither removes `~/.codeman/` or `~/codeman-cases/`. Delete those by hand if you want the
state and your case folders gone as well, and check `~/codeman-cases/` first: linked cases
point at directories you already had, but cases created from scratch have their only copy
there.
Running tmux sessions are not killed by an uninstall. `tmux -L codeman kill-server` ends
them.
## Windows (WSL)
```powershell
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows runs it inside
[WSL2](https://learn.microsoft.com/en-us/windows/wsl/install). If you do not have WSL yet:
run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, and install your agent CLI
*inside* WSL. `http://localhost:3000` then works from your Windows browser.
Work inside the Linux filesystem (`~/project`), not `/mnt/c/...`. Filesystem watching and
git are both dramatically slower across the Windows mount, and agents notice.
## macOS notes
**`Error: posix_spawnp failed.` on every session start.** node-pty publishes its macOS
`spawn-helper` without the executable bit, and macOS launches every PTY through it. Codeman
detects this and repairs it automatically on the first failure. If you hit it on a clone
install and want to fix it by hand:
```bash
npm run fix:node-pty
```
This is a `chmod`, not a rebuild. Look in `prebuilds/darwin-<arch>/`, not
`build/Release/`, which does not exist on macOS. Linux cannot reproduce this.
**launchd and PATH.** A LaunchAgent gets `/usr/bin:/bin:/usr/sbin:/sbin`, which finds
neither a Homebrew or nvm `node` nor `tmux` or `claude`. `codeman service install` bakes
your current PATH into the unit for exactly this reason, so prefer it over a hand-written
plist.
## Gotchas
- **`tmux: command not found` after a successful install.** The installer asks before
installing packages, and a declined prompt is a valid answer it remembers. Install tmux
and re-run.
- **Port 3000 in use.** `codeman web --port 8080`, or set `CODEMAN_PORT`.
- **Two Codemans on one machine.** The data directory and the tmux socket are both process
wide, so a second instance discovers and attaches the first one's live sessions. Give each
a distinct `CODEMAN_INSTANCE` before starting a second. See [Core Concepts](Core-Concepts).
- **The dashboard is not reachable from your phone.** That is the default, not a fault. The
server binds `127.0.0.1`. See [Remote Access](Remote-Access).
## Read next
- [Quick Start](Quick-Start) - your first working session.
- [Agent CLIs](Agent-CLIs) - picking and setting up a run mode.
- [Remote Access](Remote-Access) - reaching it from another device.
- [Troubleshooting](Troubleshooting) - when the above did not go as written.
+159
View File
@@ -0,0 +1,159 @@
# Keeping Agents Running
Codeman exists for the hours you are not at the keyboard. This page covers how it notices an
agent has stopped, what it does about it, and how to run a session overnight without
babysitting it.
Everything here is **per session and off by default**. A session you never configure just
sits there when it finishes, which is usually what you want.
## How Codeman knows an agent is idle
Harder than it sounds, and worth understanding, because it is what every other feature here
is built on.
**For Claude sessions**, the naive signal does not work. Claude redraws its prompt marker
roughly once a second all the way through a turn, so "saw a prompt, waited two seconds,
called it idle" flipped working sessions to idle a couple of seconds into every turn. Its
real working indicator is an animated line whose glyph and wording both change, and terminal
repaints arrive in partial fragments, so matching it in the output stream does not work
either.
So Codeman waits for the pane to go quiet, then **asks the screen** what is on it before
believing the session is idle. Turn-start detection works the same way in reverse: a
sustained run of repaints marks a turn as started, with the same screen check vetoing mere
keystroke echo. Idle now lands a few seconds after a turn genuinely ends.
There are several layers stacked on that: a completion message from the CLI, an AI check,
output silence, and token stability.
**For every other CLI**, there are no hooks to lean on, so detection is output
stabilization: the session is idle when output stops changing. Coarser, and it is why the
features further down this page are Claude-only.
## The Respawn Controller
Respawn keeps a session working past the point where the agent would otherwise stop. When
the session goes idle, Codeman runs a cycle and starts it again.
A cycle is up to four steps, each optional:
1. **Update prompt.** Ask the agent to write down where it got to, so the next round can pick
it up.
2. **`/clear`.** Reset the context window.
3. **`/init`.** Re-read the project's `CLAUDE.md`.
4. **Kickstart prompt.** Tell it to continue.
Steps 2 and 3 are what make long runs possible: without a context reset, a multi-hour
session eventually spends its whole window on its own history.
Configure it in **Session Options → Respawn**, then press **Enable**. It repeats until the
duration you set runs out.
| Setting | What it controls |
| ---------------------- | ----------------------------------------------------------------------- |
| **Idle timeout** | How long the session must be quiet before a cycle starts. |
| **Duration** | How long the whole arrangement stays armed. |
| **Inter-step delay** | Pause between the steps above, so a step is not sent into a busy pane. |
| **`/clear` + `/init`** | Whether the context reset happens at all. |
| **Update prompt** | What the agent is asked to record before the reset. |
| **Kickstart prompt** | What starts the next round. |
| **Auto-accept prompts**| Answer routine confirmation dialogs automatically. |
### Presets
Five built-ins, and the numbers matter more than the names. The idle timeout is the main
difference: a lead session coordinating subagents is legitimately silent for a minute at a
time, and a three second timeout would interrupt it constantly.
| Preset | Idle timeout | Duration | Built for |
| -------------- | ------------ | -------- | --------------------------------------------------------------- |
| **Solo** | 3s | 60 min | One agent working alone, fast cycles with a context reset. |
| **Subagents** | 45s | 240 min | A lead session running Task subagents; tolerates their silences. |
| **Team** | 90s | 480 min | Leading an agent team; tolerates long silences. |
| **Ralph/Todo** | 8s | 480 min | Working through a task list with progress tracking. |
| **Overnight** | 10s | 480 min | Unattended overnight runs with a full reset between cycles. |
Start from the preset that matches your shape of work and adjust the idle timeout first.
Presets you build yourself can be saved alongside these.
### What it costs
Every cycle is real tokens: the update prompt, the reset, and the kickstart, plus whatever
work follows. An overnight run is a deliberate spend, not a background nicety. The duration
setting is the ceiling, and it is worth setting honestly.
## Auto-resume when a usage limit resets
**Claude only.** At the top of the Respawn tab.
When Claude halts on a subscription limit, the message names the time the limit resets.
Codeman parses it, arms a timer for two minutes after that, then sends Escape followed by
`continue`.
The important part is what it does **not** do: respawn cycles are blocked while a session is
limit-paused. Without that, the next cycle would fire `/clear` and wipe the conversation you
are waiting to resume. This is the single most useful setting for overnight runs on a
subscription plan.
## The plan usage chip
**Claude only.** A header chip showing live subscription usage, on by default on desktop and
off on phones.
It works by installing a status line exporter into Claude Code, which posts Claude's own
rate limit data back to Codeman. The exporter is marker-identified, so it only ever touches
a status line Codeman installed, never one you wrote yourself, and it prints your footer
through so the in-terminal status line still works.
The chip and the exporter are the same setting. Turning the chip on without the exporter
would leave it showing a dash forever, so resolve it in one place: **App Settings**.
## Circuit breakers
Two, and they are unrelated:
- **The Ralph breaker** stops respawn thrashing. It moves from closed to half-open to open,
and is reset from the session's Ralph controls.
- **The PTY-exit breaker** trips when a session's process exits repeatedly and quickly, and
blocks automatic restarts so a broken configuration cannot spin forever.
The PTY-exit breaker resets **only** on an explicit clear. Reattaching to the session does
not clear it, deliberately, so a UI reconnect cannot paper over a session that is genuinely
failing to start.
## A working overnight setup
1. Start a Claude session in the case you want worked on.
2. Give it a clear goal and let it start. Respawn continues work, it does not invent it.
3. **Session Options → Respawn → Overnight preset.**
4. Turn on **auto-resume on usage limit**.
5. Set the duration to how long you actually want it running.
6. Press **Enable**.
7. Optionally turn on push notifications so a blocking question reaches your phone: see
[Notifications And Approvals](Notifications-And-Approvals).
In the morning, the **Away Digest** summarizes what happened while you were gone, and the
run summary and lifecycle log carry the detail.
## Gotchas
- **Respawn without a context reset stalls eventually.** The window fills with history and
the agent gets less useful every cycle.
- **An idle timeout that is too short interrupts real work.** If the agent runs long tool
calls or coordinates subagents, raise it. That is what the Subagents and Team presets are.
- **The update prompt is what makes a reset survivable.** After `/clear`, everything the
agent knows comes from that summary and the project files. A vague update prompt produces
a vague next cycle.
- **Non-Claude sessions can respawn**, but with output-based idle detection and no
usage-limit auto-resume.
- **Do not run respawn on a session you are actively typing in.** It will send prompts
underneath you.
## Read next
- [Autonomous Loops](Autonomous-Loops) - Ralph and the orchestrator, for structured
autonomous work rather than "keep going".
- [Cron Jobs](Cron-Jobs) - starting work on a schedule instead of continuing it.
- [Notifications And Approvals](Notifications-And-Approvals) - being told when it needs you.
- [`docs/respawn-state-machine.md`](https://github.com/Ark0N/Codeman/blob/master/docs/respawn-state-machine.md) - the state machine itself.
+72
View File
@@ -0,0 +1,72 @@
# Keyboard Shortcuts
Every binding, and how to change them. `Ctrl` also accepts `Cmd` on macOS.
Press `Ctrl+?` in the app for the same list in a floating overlay.
## Sessions and tabs
| Shortcut | Action |
| ------------------------------- | --------------------------------------------------------------- |
| `Ctrl+K` (also `Cmd+K`, `Alt+K`)| Find an open session or start a new one. |
| `Ctrl+W` | Kill the active session. |
| `Ctrl+Tab` | Next session. |
| `Alt+[` / `Alt+]` | Previous / next tab. |
| `Alt+1` to `Alt+9` | Switch to tab N. Physical keys, so macOS Option layouts work. |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move the active tab left / right. |
| `Alt+B` | Collapse / expand the session sidebar, when that layout is on. |
## Terminal
| Shortcut | Action |
| ----------------------- | --------------------------------------------------------------- |
| `Enter` | Send. |
| `Shift+Enter` | Insert a newline without sending. |
| `Ctrl+Enter` | Same. |
| `Ctrl+C` | Copy the selection, or interrupt when nothing is selected. |
| `Ctrl+Shift+C` | Copy the selection. Never interrupts. |
| `Ctrl+L` | Clear the terminal. |
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl` `+` / `Ctrl` `-` | Font size. |
| `Shift+Wheel` | Scroll the local buffer, even where the wheel is forwarded to the CLI. |
## Everything else
| Shortcut | Action |
| -------------- | ------------------------------- |
| `Ctrl+Shift+V` | Toggle voice input. |
| `Ctrl+?` | Shortcut reference overlay. |
| `Escape` | Close panels and modals. |
## Rebinding
**App Settings → Shortcuts.** Bindings live in a registry with per-user overrides, so a
rebind is stored as an override on top of the default rather than replacing the table.
Two things are deliberately not rebindable:
- **`Ctrl+C` smart copy.** The generic dispatch loop calls `preventDefault()` on every
shortcut it handles, and doing that to `Ctrl+C` would swallow the interrupt when nothing
is selected. It is handled separately for that reason.
- **`Escape`**, which closes whatever is open.
## Why some chords behave oddly
The terminal sees keystrokes before the app does. Any chord the app claims has to also be
swallowed at the terminal layer, or xterm writes the control byte into the session as well
as triggering the action. If you rebind something to a chord the terminal cares about
(`Ctrl+D`, say), expect the CLI to see it too.
`Alt+1` through `Alt+9` are matched on **physical key position** rather than the character
produced, so macOS Option layouts that produce `¡™£` still switch tabs.
## On phones
There is no physical keyboard, so the equivalents live in the keyboard accessory bar: `Esc`,
`Ctrl` as a one-shot modifier, `Tab`, arrows, and quick actions. See
[Mobile Guide](Mobile-Guide).
## Read next
- [The Dashboard](The-Dashboard) - what the shortcuts are navigating.
- [Settings Reference](Settings-Reference) - where the overrides are stored.
+165
View File
@@ -0,0 +1,165 @@
# Mobile Guide
Codeman on a phone is not a shrunken desktop UI. It is the surface most of its design
attention has gone into, because checking on an agent from a bus is the thing this software
is for.
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/screenshots/mobile-session-keyboard-20260727.png" alt="Answering an agent prompt on a phone" width="300">
</p>
## Getting there
1. **Set up access.** Tailscale is the recommended route and gives you real HTTPS. See
[Remote Access](Remote-Access).
2. **Log in by QR.** Open the dashboard on your desktop and scan the code. No password
typing. Tokens are single use and rotate every 60 seconds.
3. **Install it to your home screen.** On iOS this is mandatory for push notifications;
Safari does not deliver push to tabs. On Android it makes the app full screen.
HTTPS matters for more than security here: microphone access and push notifications both
require a secure context.
## The layout
| Element | Where |
| -------------------- | --------------------------------------------------------------------- |
| Header | Fixed at the top, deliberately minimal. Desktop-only controls never appear. |
| Tab strip | Scrolls horizontally. The active tab is always scrolled into view. |
| Terminal | The rest of the screen. |
| Toolbar | Bottom: Run, Stop, **Enter**, case picker, voice, settings. |
| Keyboard bar | Above the on-screen keyboard when it is open. |
Layout respects notch and home-indicator safe areas, touch targets are 44px, and the case
picker is a bottom sheet rather than a dropdown.
**Swipe left and right** on the terminal to switch sessions.
## The home screen
Tapping the "C" logo gives a session overview rather than a welcome page:
1. **NEEDS YOU** first: sessions blocked on a question, with answer strips so you can
resolve them without opening the session.
2. **CURRENT SESSIONS** with live status.
3. **PAST SESSIONS**, resumable.
Row status uses the same language as the tabs: green when fine, pulsing while working,
yellow when waiting for input, red when a question is pending.
The split Run button carries the same per-backend colours as the desktop toolbar, and its
picker mirrors the desktop run-mode menu.
On by default; it can be turned off in settings.
## The keyboard accessory bar
A row of keys above the virtual keyboard, and what it contains depends on the session.
**Agent sessions** get quick actions: `/init`, `/clear`, `/compact`, a clipboard key, `Esc`,
a path picker, an image key, and 🧠 when Read My Mind is on. Destructive commands need a
double press, so you cannot fire `/clear` with a stray thumb.
**Shell sessions** automatically swap it for terminal controls: `Ctrl`, `Esc`, `Tab`, four
arrows, paste, and dismiss. Your normal preference is remembered and restored when you
switch back to an agent session, so a settings change during a shell session cannot strip
the bar away permanently.
### One-shot Ctrl
`Ctrl` on the shell bar is a **one-shot modifier**: tap `Ctrl`, then tap `c`, and the
control byte is sent. It disarms on use, on a second tap, on any other accessory key, on a
session switch, and when the keyboard closes.
That list matters. A modifier left armed turns your next innocent keystroke into a control
byte, so it is deliberately eager to disarm. Keys with no control equivalent pass through
unchanged, exactly like a hardware keyboard.
## The Enter button
The toolbar's dedicated **Enter** button exists because of local echo. On a phone, the
characters you type are painted locally and have not reached the agent yet; Enter flushes
them and then submits.
It replays the keypress through the terminal rather than sending a bare carriage return.
Sending a bare `\r` would submit an empty line and strand your typed text on screen, which
looks exactly like a dead button.
On phones this button replaces the desktop's **Run Shell** control; starting a shell moved
into the Run dropdown.
## Tapping, links and copying
- **Tap a link** in terminal output and it opens in a new tab. Same for a link in an agent's
answer in the response viewer — it opens a tab rather than navigating the dashboard away,
which on a phone would unload the whole session view.
- **Tap a file path** an agent printed and the file-preview overlay opens; a log path opens the
log viewer. Works in scrolled-up transcript too.
- A tap on the prose *beside* a link still places the cursor as usual, and a tap on a dialog's
numbered choice still answers the dialog even when the row contains a path — the dialog wins,
because on a phone it is the only interaction that matters.
- **Long-press to select text**, then drag, or tap the other end to extend the selection — no
hairline handles to grab. A small bar offers **Copy**, **Line** (the whole logical line,
wrapped rows included) and dismiss. Copy works on plain-HTTP installs too, where the browser
clipboard API is unavailable.
- A swipe is never mistaken for a long-press, and the keyboard stays down while you select.
## Scrolling and the keyboard
- The terminal and toolbar shift up when the keyboard opens, tracked through the browser's
visual viewport rather than guessed.
- **Two ways to dismiss the keyboard**: tap outside the terminal on inert space, or tap twice
on inert terminal content. Tapping a control never dismisses it, and tapping the prompt row
keeps focus so you can place the caret.
- A scroll is never mistaken for a tap: travel is measured from the start of the gesture, and
multi-touch never counts.
- **A long prompt stays visible.** Once what you are typing wraps past the last visible row it
grows upward over the transcript instead of sliding under the keyboard, so the end of the
sentence — where the cursor is — is always on screen. A prompt taller than the visible strip
shows its tail.
## Voice
The microphone button, or the keyboard bar. Providers and setup are covered in
[Input And Voice](Input-And-Voice). Dictating is often faster than typing a prompt on a
phone, and it is the main reason the feature exists.
## Notifications
Push notifications reach you with no tab open, and with the Approvals Inbox on they carry
**Approve** and **Deny** buttons handled by the service worker, so you can unblock an agent
from the lock screen.
Setup in [Notifications And Approvals](Notifications-And-Approvals).
## Reading long answers
The terminal viewport is small. **Last Response** (opt-in header button) renders the agent's
last answer as scrollable text instead, with a **More** button for additional context.
The [File Viewer](Working-With-Files) works on phones too, including edit mode, which is
enough to fix a typo an agent introduced while you are away from your desk.
## What is deliberately not on phones
- Extra header buttons. New header controls are kept off phones by policy, with a test that
enforces it.
- The Approvals bell. Phones get the NEEDS YOU strips on the home screen instead.
- The desktop home tab rail, which needs a wide window.
- Lineage arcs, which are a desktop overlay.
## Gotchas
- **Typed text sitting on screen has not been sent.** Press Enter.
- **iOS needs the home screen install for push**, not just a bookmark.
- **iOS Safari can serve stale JavaScript after an update** until the tab is fully closed.
Close it and reopen.
- **Plain HTTP over a LAN address disables voice and push.** Use HTTPS.
- **An armed `Ctrl` is visibly highlighted.** If it looks the same as a resting key, you are
on an old version, on a light skin.
## Read next
- [Remote Access](Remote-Access) - getting the phone connected in the first place.
- [Notifications And Approvals](Notifications-And-Approvals) - being told when you are needed.
- [Input And Voice](Input-And-Voice) - local echo, dictation, and the input rules.
+97
View File
@@ -0,0 +1,97 @@
# Multi-User Mode
Share one Codeman with a small trusted team. Each person gets their own login and workspace,
and sessions, cases, search, and live events are scoped to their owner.
**Off by default.** Without the flag, behaviour is identical to single-user Codeman, because
every scoping check short-circuits.
## Read this before enabling it
**Multi-user mode separates workspaces. It does not sandbox users from each other.**
Every session still runs as the **same operating system account**. A determined user's agent
can reach another user's files, because at the OS level they are the same user. This is a
convenience and organization feature, not a security boundary.
If you need real isolation:
- Pair each user with [Docker Cases](Docker-Cases), which gives their work its own
filesystem and network.
- Or run separate Codeman instances under separate OS accounts, each with its own
`CODEMAN_INSTANCE`.
"Small trusted team" is the honest description of who this is for.
## Enabling it
```bash
codeman users add alice --admin # create the first admin, prompts for a password
codeman web --multiuser # or CODEMAN_MULTIUSER=1
```
Then manage users from the CLI or the **Users** entry in App Settings:
```bash
codeman users add bob # a regular user
codeman users list
codeman users passwd bob # reset to a one-time password
codeman users rm bob
```
`--password-stdin` reads the password from standard input, for scripts.
Accounts live in `~/.codeman/users.json` with scrypt-hashed passwords, mode 0600.
Administrative actions are audited to `~/.codeman/admin-audit.jsonl`.
## What each user gets
| Thing | Scope |
| ------------------- | ---------------------------------------------------------------------------- |
| **Case space** | `~/codeman-users/<name>/cases`, their own. |
| **Sessions** | Only theirs are listed, reachable, or controllable. |
| **Events** | Live event routing is per owner, and fails closed. |
| **Search** | Scoped on read, including historical results. |
| **File previews** | Scoped to sessions they own. |
| **Path picker** | Only their own user space as a root, not the whole home directory. |
Admins see everything.
Ownership threads through every list endpoint, the session lookup helper, the WebSocket
layer, and file previews. A user cannot address another user's session even by id.
## Safer defaults for regular users
Non-admins get tighter defaults, and lifting them is an explicit per-user grant:
| Default | Meaning |
| ---------------------------------- | ------------------------------------------------------------------------ |
| Claude runs in `auto` permission mode | Anthropic's classifier-guarded mode instead of skip-prompts. |
| Raw shell sessions require a grant | A plain shell is unmediated machine access. |
| Skip-permissions requires a grant | Same reasoning. |
| Cron `launchCommand` requires a grant | It is an arbitrary command on a schedule. |
| Pi project trust defaults to off | Trust makes Pi execute repo-local TypeScript. |
These exist because the OS boundary is shared. They narrow what a normal account can do
casually; they do not make the account a sandbox.
## Accounts and sessions
Each user authenticates with their own name and password rather than the shared
`CODEMAN_PASSWORD`. Logins are individually revocable: disable, reset, or delete an account
at any time, and existing browser sessions can be revoked.
## Gotchas
- **Enabling it does not migrate existing cases** into a user space. They stay where they
are, owned by whoever the ownership rules resolve them to.
- **Admins see everything**, including other users' sessions. Choose admins accordingly.
- **The audit log is append-only and local.** Ship it somewhere if you care about it.
- **It is not a substitute for OS accounts.** Restating this because it is the one thing
people get wrong.
## Read next
- [Security](Security) - where this fits in the model, and what it does not cover.
- [Docker Cases](Docker-Cases) - the isolation story that actually isolates.
- [`docs/multi-user-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/multi-user-plan.md) - the design.
+147
View File
@@ -0,0 +1,147 @@
# Notifications and Approvals
An agent that stops to ask a question, with nobody watching, is a run that quietly wasted an
hour. This page covers every way Codeman tells you it needs you, and how to answer without
opening the session.
## The signals, cheapest first
| Surface | Reaches you | Default |
| ---------------------- | ------------------------------------------------- | ------- |
| Tab alert | While the dashboard is open | On |
| Browser title flash | Another tab in the same browser | On |
| Desktop notification | Another window on the same machine | Opt-in |
| Push notification | Anywhere, even with no tab open | Opt-in |
| Approvals Inbox | One queue across every session | Opt-in |
| Phone overview | Phone home screen, NEEDS YOU section | On |
| Away Digest | Afterwards, as a summary | Opt-in |
## Tab alerts
The tab itself changes state:
| State | Meaning |
| -------------------- | ---------------------------------------------------------- |
| Yellow, blinking | The agent is waiting for input from you. |
| Red, blinking | A question or permission prompt is blocking the session. |
These are a steady colour with a pulse layered on top, not a blink to transparent, so a tab
needing attention looks that way at every point in the cycle.
They survive a reload. The alert state is re-seeded from the server on page load, so
reloading the dashboard while a permission dialog is blocking a session does not leave you
with a normal-looking tab.
For Claude sessions, these come from Claude Code's hooks and are precise about *why* the
session stopped. For other CLIs there are no hooks, so you get the coarser output-based
signal.
## Window title and OS notifications
The browser tab title is prefixed `codeman:<host>`, so several Codeman instances across
several machines stay distinguishable at a glance. Override the hostname with
`codeman web --title-hostname <name>`.
Desktop notifications use the same prefix. Enable them in **App Settings → Notifications**.
## Push notifications
Push reaches your phone with **no Codeman tab open at all**, which is the only option that
works while you are actually away.
Setup:
1. Open Codeman over **HTTPS**. Web push requires a secure context. Tailscale gives you real
HTTPS; `--https` gives you a self-signed certificate; plain HTTP over a LAN address will
not work.
2. **App Settings → Notifications → Subscribe**, and accept the browser prompt.
3. On **iOS**, add Codeman to your home screen first. Safari only delivers web push to
installed web apps, not to tabs.
Once subscribed, a blocking prompt reaches your phone even from a locked screen.
## The Approvals Inbox
**Opt-in, off by default. Claude sessions only.**
One queue of every prompt currently waiting on a human, across all your sessions, answerable
in place. When you have eight workers running, this is the difference between checking eight
tabs and checking one list.
Turn it on in **App Settings**. Surfaces:
- **A header bell** with a count, hidden entirely while the count is zero. Never shown on
phones.
- **A drawer** listing each waiting card.
- **NEEDS YOU strips** at the top of the phone overview home screen.
Each card shows the session, the case, and the captured prompt with its options. Answering
sends the keystroke into the session for you: a digit for a menu choice, Escape to decline,
or free text for an idle prompt.
Behaviour worth knowing:
- **One item per session.** A newer prompt supersedes the older one, because the older one
is no longer on screen.
- **Menu answers are validated against the live screen.** Codeman re-captures the pane before
sending, and refuses with a conflict if the dialog is no longer there. Otherwise your
keystroke would land in the composer as stray text.
- **Permission and question items clear only on definitive signals**: the turn ending, the
dialog completing, an answer, a supersede, the session exiting, or a 12 hour timeout. They
do not clear on a heuristic "looks busy again" signal, because that signal is wrong often
enough to lose a real prompt.
- **In memory only.** Restarting the server clears the queue; the prompts themselves are
still sitting in the sessions.
### Approve and Deny from the notification
With the inbox enabled, push notifications carry **Approve** and **Deny** buttons. Those are
handled by the service worker directly, so they work with no tab open: tap Approve on a
locked phone and the agent continues.
With the inbox off, the buttons are stripped from the notification payload entirely rather
than being shown and failing.
## The phone overview
On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then
current sessions, then past ones. Rows use the same language as the tab strip: a green dot
when fine, pulsing while working, yellow when waiting for input, red when a question is
pending.
Answer strips let you resolve a prompt straight from the home screen without opening the
session.
## The Away Digest
Retrospective rather than live: what happened while you were gone, aggregated from the
lifecycle log, run summaries, live sessions, token statistics, and recent subagents.
It is the morning-after view for an overnight run. Enable its header button in
**App Settings → Header & Panels**.
## Recommended setup for unattended runs
1. HTTPS access, ideally Tailscale. See [Remote Access](Remote-Access).
2. Push notifications subscribed, with Codeman installed to the home screen on iOS.
3. Approvals Inbox on.
4. Auto-resume on usage limit on, for each session you leave running. See
[Keeping Agents Running](Keeping-Agents-Running).
That combination means a blocking question wakes your phone and can be answered in two taps
from the lock screen.
## Gotchas
- **No push over plain HTTP.** It is a browser requirement, not a Codeman one.
- **iOS needs the home screen install.** A Safari tab will never receive push.
- **The bell is invisible at zero.** That is deliberate, not a broken setting.
- **Approvals are Claude-only.** They are built on hook events the other CLIs do not emit.
- **A stale menu answer is refused, not sent.** If you answer a card for a dialog that has
since gone away, Codeman declines rather than typing a digit into the composer.
## Read next
- [Keeping Agents Running](Keeping-Agents-Running) - what to configure before walking away.
- [Mobile Guide](Mobile-Guide) - the phone surfaces in full.
- [Settings Reference](Settings-Reference) - where each of these toggles lives.
+153
View File
@@ -0,0 +1,153 @@
# Quick Start
From an installed Codeman to a working agent, in about five minutes. If you have not
installed yet, start at [Installation](Installation).
## 1. Start the server
```bash
codeman web
```
It prints a URL, `http://localhost:3000` by default. Open it.
The server binds `127.0.0.1` only, so this URL works from the machine running it and
nowhere else. That is deliberate: Codeman starts agents with permission prompts skipped by
default, so anyone who can reach the dashboard can run code on this machine. Reaching it
from your phone is a separate, deliberate step covered in [Remote Access](Remote-Access).
To keep it alive after you close the terminal, use `codeman web -d` instead, or install it
as a service. See [Running As A Service](Running-As-A-Service).
## 2. Meet the welcome screen
With no sessions running you get the welcome screen:
- **Run buttons** for each agent CLI Codeman found on your PATH. If you expected one and it
is missing, its binary is not visible to the server; see [Agent CLIs](Agent-CLIs).
- **A QR code**, if a password is set. Scanning it logs a phone in without typing anything.
- **Resume Conversation**, a list of past sessions, including Claude conversations started
outside Codeman. Empty on a fresh install.
- **Search**, across sessions, events, and files.
You can click a Run button right now and get a working agent in your current case. The rest
of this page is the deliberate version.
## 3. Pick or create a case
A **case** is a named working directory that Codeman remembers. Every session runs inside
one. The case picker is in the bottom toolbar.
To make a new one, click **+** next to the picker. The Add Case dialog has three tabs:
| Tab | Use it when |
| ----------------- | ------------------------------------------------------------------------------------------------------------------ |
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. |
| **Clone Repo** | Working on an existing public repo. Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Link Existing** | The code is already on disk. Point at the folder, with **Browse** if you would rather click than type. |
The gear next to the picker holds two per-case toggles: **Agent Teams** and
**1M Opus Context**. Both are off by default and both are safe to ignore for now.
**Create New** also has a checkbox for running the case inside a Docker container, and a
**Remote** panel for running it over SSH on another machine. Those are
[Docker Cases](Docker-Cases) and [Remote SSH Sessions](Remote-SSH-Sessions); skip them for
your first session.
## 4. Pick a run mode and hit Run
The **Run** button starts an agent in the selected case. The arrow next to it picks which
one:
| Mode | What starts |
| -------------------- | -------------------------------------------------------------- |
| **Claude Code** | The default, and the mode every Codeman feature supports. |
| **OpenCode** | |
| **Codex** | OpenAI's CLI. |
| **Gemini** | Enterprise only since Google's consumer cutover. |
| **Antigravity** | Google's successor to the consumer Gemini CLI. |
| **Pi** | No permission prompts and no sandbox by design. |
| **Terminal / Shell** | A plain shell, no agent. Also the **Run Shell** button. |
The dropdown also lists any saved dashboard URLs ([Web Tabs](Web-Tabs)) and your recent
sessions. Those do not change the run mode: Run always means "start an agent".
Click **Run**. A tab appears, and Codeman spawns the CLI on a real PTY inside a tmux
session and streams it to your browser.
The number spinner beside the button starts several sessions at once, up to 20. Useful for
fanning the same case out across parallel workers; unnecessary for a first run.
## 5. Talk to the agent
Click into the terminal and type. It is a real terminal (xterm.js over a real PTY), so full
TUIs render properly and everything the CLI supports works, slash commands included.
| Key | Effect |
| ---------------------------- | --------------------------------------------- |
| `Enter` | Send. |
| `Shift+Enter` / `Ctrl+Enter` | Newline without sending. |
| `Ctrl+C` | Copy if text is selected, otherwise interrupt. |
| `Ctrl+Shift+V` | Voice input. |
You can also paste or drag an image straight into the session, and register external files
as attachments. See [Working With Files](Working-With-Files) and
[Input And Voice](Input-And-Voice).
Input is delivered **exactly once**, even if your connection drops mid-prompt. A dropped
link never loses a prompt and never sends it twice.
## 6. Read the tab
The tab tells you what the session is doing without opening it:
| Signal | Meaning |
| --------------------- | ---------------------------------------------------------- |
| Green dot | Alive and idle. |
| Pulsing green dot | Working on a turn. |
| Yellow, blinking | Waiting for you to type something. |
| Red, blinking | A question or permission prompt is blocking the agent. |
Full tour in [The Dashboard](The-Dashboard). If you want a phone notification when an agent
needs you, that is [Notifications And Approvals](Notifications-And-Approvals).
## 7. Leave, and come back
Close the browser tab. Close the laptop. The agent keeps running, because it lives in tmux
and not in your browser.
Reopen the dashboard and the session is still there with its scrollback intact. First load
of a session pulls the full tmux scrollback, so you get the history, not just what arrived
after you reconnected.
This also survives restarting the Codeman server itself. What does not survive is killing
the tmux server or rebooting the machine.
## 8. Stop things
| To do this | Do that |
| ------------------------- | ------------------------------------------------------------------- |
| Interrupt the current turn | `Ctrl+C` with nothing selected, or the **Stop** button. |
| Close one session | `Ctrl+W`, or the tab's close control. |
| Stop the server, keep agents | `codeman web --stop`. The tmux sessions stay alive. |
| Stop everything | `tmux -L codeman kill-server`. |
If you are working *inside* a Codeman-managed session (`echo $CODEMAN_MUX` prints `1`),
never run `tmux kill-session` or `pkill claude` by hand. You will kill the session you are
sitting in, along with its siblings.
## Where to go next
**Make it run without you.** [Keeping Agents Running](Keeping-Agents-Running) covers idle
detection, respawn cycling, and auto-resume when a subscription limit resets. That is the
feature Codeman exists for.
**Get it on your phone.** [Remote Access](Remote-Access), then
[Mobile Guide](Mobile-Guide).
**Understand what you just used.** [Core Concepts](Core-Concepts) explains cases, sessions,
run modes, and what state lives where.
**Automate it.** [Cron Jobs](Cron-Jobs) for scheduled work,
[Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) for agents that spawn and
supervise other agents.
+209
View File
@@ -0,0 +1,209 @@
# Remote Access
Reaching your Codeman from a phone, a laptop on the other side of the house, or a hotel
network. This is the page to read carefully, because Codeman's dashboard is a
remote-code-execution surface by design: it starts agents with permission prompts skipped,
so whoever can reach it can run code on your machine.
## Start from the default
`codeman web` binds `127.0.0.1`. It is reachable from the machine running it and nothing
else, which is why the no-password default is safe out of the box. Every option below is a
deliberate step away from that.
Two rules that make the rest of this page simple:
1. **Never expose Codeman on a network without `CODEMAN_PASSWORD`.** Binding a non-loopback
host without one starts, but prints a loud warning with the fixes.
2. **Prefer keeping the loopback bind** and putting an authenticated tunnel in front of it,
over binding wide and relying on a password alone.
## Pick an approach
| Approach | Good for | Cost |
| --------------------- | ----------------------------------------------------- | --------------------------------------------------------- |
| **Tailscale** | Phone access, permanently. The recommended setup. | Install Tailscale on both devices. |
| **Cloudflare tunnel** | A public URL, quickly, from anywhere. | Public URL, so a password is mandatory. |
| **LAN + password** | Home network only, no extra software. | Every device on your LAN can reach the login page. |
| **SSH port forward** | You already SSH to the box. | Manual, per session, terminal-bound. |
## Tailscale (recommended)
Your devices join a private network, and Codeman stays bound to loopback. Nothing is
published to the internet, and you get real HTTPS with a real certificate.
The installer sets this up for you, including installing Tailscale, logging in, enabling
tailnet HTTPS, and verifying the result end to end. To retrofit it onto an existing
install:
```bash
install.sh tailscale
```
By hand:
```bash
tailscale serve --bg 3000
tailscale serve status
```
Then open `https://<machine>.<tailnet>.ts.net` from any device on your tailnet.
Notes:
- Keep the loopback bind. `tailscale serve` connects to `127.0.0.1:3000` locally, so
binding wider adds exposure and buys nothing.
- Your tailnet is the authentication boundary. Setting `CODEMAN_PASSWORD` as well is
reasonable defence in depth, especially if other people have devices on your tailnet.
- Codeman's Host-header allowlist already accepts `.ts.net`, so no extra configuration is
needed.
- The installer never resets or rewrites `serve` mappings other than the one pointing at
Codeman's port, so unrelated serve configuration is left alone.
## Cloudflare tunnel
A free [quick tunnel](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/)
gives you a public HTTPS URL with no port forwarding, no DNS, and no static IP:
```
Browser → Cloudflare edge (HTTPS) → cloudflared → localhost:3000
```
Prerequisites: [`cloudflared`](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/downloads/)
installed, and `CODEMAN_PASSWORD` set.
```bash
./scripts/tunnel.sh start # starts the tunnel, prints the public URL
./scripts/tunnel.sh url
./scripts/tunnel.sh status
./scripts/tunnel.sh stop
```
The quick-tunnel URL is a random `*.trycloudflare.com` address that changes every time the
tunnel restarts. For a stable hostname, `./scripts/tunnel.sh named setup` walks through a
named tunnel.
To survive reboots:
```bash
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
```
There is also a toggle in **App Settings → System → Remote access**.
**The tunnel refuses to start without a password.** That is on purpose: a public URL with no
authentication is a terminal on your machine handed to the internet. Acknowledging the risk
explicitly is possible from the UI toggle, and only from there; the API will not do it for
you.
## LAN plus password
```bash
export CODEMAN_PASSWORD='something long'
codeman web -H 0.0.0.0 --https
```
Every device on your local network can now reach the login page. `--https` generates a
self-signed certificate into `~/.codeman/certs/`, which your browser will warn about once.
`CODEMAN_USERNAME` defaults to `admin`.
The installer offers this path and prompts for the password. On re-runs it preserves
whichever binding you already chose.
## SSH port forward
No configuration at all, if you already have SSH access:
```bash
ssh -L 3000:localhost:3000 you@your-box
```
Then open `http://localhost:3000` on the local machine. Codeman keeps its loopback bind and
sees a local connection. Good for occasional access, awkward as a permanent arrangement
because it dies with the SSH session.
## Logging in from a phone
Typing a long password on a phone keyboard is miserable, so Codeman issues **single-use QR
tokens**. The desktop dashboard shows a QR code; scan it and the phone is authenticated.
How it behaves:
- The code rotates every 60 seconds, with a 90 second grace window so scanning during a
rotation still works.
- Each token is **single use**. The moment a phone consumes it, a new one is generated.
- The URL contains a 6-character lookup code, not the secret, so it does not leak through
browser history, `Referer` headers, or the tunnel provider's logs.
- The desktop shows a toast naming the device and browser that just authenticated, with a
one-click revoke.
- QR attempts are rate limited separately from password attempts, so a mistyped password
cannot lock out your QR login and vice versa.
Someone holding only the tunnel URL still meets the normal password prompt. The QR is the
fast path, not a bypass.
Design detail and the threat analysis it is built against:
[`docs/qr-auth-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/qr-auth-plan.md).
## Behind a reverse proxy
Codeman enforces a Host-header allowlist on every request to block DNS rebinding, and the
same allowlist gates the cross-site Origin check. It accepts `localhost`, IP literals, the
bind host, `.ts.net`, `.trycloudflare.com`, `.cfargotunnel.com`, and the active managed
tunnel.
**Your own domain is not on that list.** Add it:
```bash
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
A bare entry matches that exact host; a leading dot matches subdomains. Without this, a
correctly configured proxy still gets `403 host not allowed`, which reads like a proxy bug
and is not one.
Also make sure the proxy forwards WebSocket upgrades. The terminal is a WebSocket, and the
upgrade runs the same Host and Origin checks, closing with code `4003` on failure.
## Session cookies and rate limits
The first request prompts for HTTP Basic credentials. On success the server issues an opaque
`codeman_session` cookie (24 hour lifetime, extended on activity, validated server-side so
it cannot be forged offline). Ten failed attempts from one IP produce a `429` with a 15
minute decay.
A valid cookie or a correct password recovers immediately even while an attacker is hammering
the same IP, which matters because all tunnel traffic arrives from one loopback address.
## Terminal alternatives
You do not have to use a browser. `sc` is a thumb-friendly session chooser for SSH clients
like Termius or Blink:
```bash
sc # interactive chooser
sc 2 # attach to session 2
sc -l # list
```
Detach with `Ctrl+A D`. The sessions are the same ones the dashboard shows.
## Common problems
| Symptom | Cause and fix |
| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| `403 host not allowed` | Your domain is not in the allowlist. Set `CODEMAN_ALLOWED_HOSTS`. |
| Phone shows the login page but the terminal never connects | The proxy is not forwarding WebSocket upgrades. |
| Browser warns about the certificate | Expected with `--https` and its self-signed certificate. Tailscale gives you a real one instead. |
| LAN IP does not respond, but a tunnel to the same box works | The server is bound to loopback. That is the default. A tunnel reaches it; a LAN browser cannot. |
| Hooks stopped working after switching to HTTPS | Hook callbacks need `-k` for the self-signed certificate. Recent versions self-heal existing cases; if yours predates that, recreate the case's hooks. |
| Everything is slow over the tunnel | Quick tunnels route through Cloudflare's edge. Tailscale is usually a direct connection and much faster. |
## Read next
- [Security](Security) - the whole model, and the hardening checklist.
- [Mobile Guide](Mobile-Guide) - once you can reach it from the phone.
- [Running As A Service](Running-As-A-Service) - keeping server and tunnel up across reboots.
- [`docs/security-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/security-architecture.md) - the full model.
+101
View File
@@ -0,0 +1,101 @@
# Remote SSH Sessions
Point a case at another machine and the agent runs **there**, with the same dashboard,
mobile UI, and autonomy features. Your laptop becomes a window onto a session living on the
remote host.
Like Docker, this is a **location overlay** on a case, not a run mode. All seven run modes
work remotely. See [Core Concepts](Core-Concepts).
## Why bother
The agent runs where the work is: a build server, a NAS, a GPU box, a machine reachable only
through a jump host. Your laptop can sleep, change networks, or close, and the run continues.
## Setting it up
**Add Case → Remote**:
| Field | Notes |
| --------------------- | -------------------------------------------------------------------- |
| **Host** | Hostname or IP. |
| **Username** | The SSH user. |
| **Port** | Defaults to 22. |
| **Identity file** | `~` and `$HOME` are expanded for you. |
| **Jump host** | The `-J` equivalent, `[user@]host[:port]`. |
| **SOCKS proxy** | For hosts reachable only through a proxy. |
| **Extra SSH options** | Any `KEY=VALUE` options your normal connection needs. |
| **Remote path** | The working directory on that machine. |
Hosts are saved and reusable, so a second case on the same machine is just a path. Host
profiles can also carry per-run-mode launch command overrides, for when the binary lives
somewhere unusual on that host.
The remote host needs **tmux**. Codeman probes for it when you link the host rather than
failing later at launch.
## What actually runs
The agent lives inside a dedicated tmux server on the **remote** host, and Codeman fronts it
with a local tmux pane running `ssh`.
That two-layer arrangement is what makes it durable: a dropped SSH connection, a network
change, or a closed laptop kills the local pane, not the remote session. Reconnecting lands
back in the same live conversation.
The remote session name is deliberately chosen so that a Codeman **running on the target
host** will not adopt it as one of its own. Two Codemans, one host, no interference.
## Auto-reconnect
A watcher with bounded backoff notices a dead SSH pane and quietly reattaches to the still
running remote session. On by default; the kill switch is in
**App Settings → Agents & CLIs → Remote auto-reconnect**.
Intentional kills are never revived. Closing a session means closing it.
## Discover and attach
Codeman can list the `codeman-*` sessions already running on a host, whether that machine's
own Codeman started them or another operator did, and attach to one.
The distinction that matters:
| Session | On tab close |
| ------------ | ------------------------------------------------ |
| **Launched** | Killed, like any local session. |
| **Attached** | **Detached, never killed.** |
Attaching to someone else's session and closing your tab must not end their run, so it does
not. Several clients can attach the same remote session at different window sizes without
clamping each other, and discovery shows a shared badge with the client count.
## Security
Every SSH command line in Codeman flows through one builder that shell-escapes every
user-supplied field: identity paths, jump hosts, proxy commands, and extra options. That is
the entire injection surface, and it is deliberately a single function rather than string
concatenation spread across the codebase.
Host, path, and identity fields are schema-validated on top of that.
Codeman does not store SSH passwords. Use keys, as you would for any other automation.
## Gotchas
- **The remote host needs tmux.** Probed at link time, so you find out immediately.
- **The local working directory is meaningless** for a remote session, and is not used.
- **Run flows must go through the quick-start path** for remote cases. This matters if you
are driving Codeman over the API: the plain session-create endpoint validates the working
directory locally and has no case concept, so it will reject or misroute a remote case.
- **Latency is SSH latency.** Local echo helps the typing feel, but a slow link is a slow
link.
- **Transcript-backed features follow the transcript.** Subagent windows and similar surfaces
read files on the machine where the agent runs.
## Read next
- [Core Concepts](Core-Concepts) - overlays versus run modes.
- [Docker Cases](Docker-Cases) - the other overlay.
- [Security](Security) - the wider model.
- [`docs/remote-sessions.md`](https://github.com/Ark0N/Codeman/blob/master/docs/remote-sessions.md) - the full design.
+197
View File
@@ -0,0 +1,197 @@
# Running As A Service
Keeping Codeman up: past the shell you started it in, past a logout, past a reboot. Plus
logs, updates, and running more than one instance.
## Three levels
| Level | Survives | Command |
| -------------------- | ----------------------------------------- | ------------------------- |
| Foreground | Nothing. Dies with the terminal. | `codeman web` |
| Detached | Closing the shell and logging out. | `codeman web -d` |
| Service | Reboots. | `codeman service install` |
Agents themselves survive all three, because they live in tmux. Stopping the server never
stops the agents.
## Detached mode
```bash
codeman web -d # start detached; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful stop; agents keep running
```
`-d` waits until the server actually answers before reporting success, so a port clash never
reads as a successful start.
Two implementation details that explain the behaviour:
- It relaunches the same entry script detached, so there is no controlling terminal and no
shell job entry. `nohup` is **not** what makes this work: Node re-arms the hangup signal to
its default even when it inherits "ignore", and Codeman handles that signal with a graceful
shutdown, so a delivered hangup would still stop the server.
- `--stop` verifies the process still looks like a Codeman server before signalling it,
because process ids get recycled.
**It refuses to start a second server on the same data directory.** Two servers sharing a
tmux socket attach to each other's live sessions.
## Installing as a service
```bash
codeman service install # systemd user unit on Linux, LaunchAgent on macOS
codeman service status
codeman service uninstall
```
The installer's final menu offers this too.
Notable behaviours:
- **Your PATH is baked into the unit.** launchd hands a job
`/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew or nvm `node` nor `tmux`
or `claude`. This is the single most common cause of a hand-written unit that starts and
immediately dies.
- **`CODEMAN_PASSWORD` is never written into the unit file.** Add it yourself if the service
needs authentication.
- **It refuses when a server is already running** on that data directory, for the same reason
detached mode does.
- **It verifies rather than assumes.** `launchctl load` and a clean spawn are both silent
about a server that starts and immediately exits, so the parent polls until the child
answers or dies.
On Linux, if you want the service running while you are not logged in:
```bash
loginctl enable-linger $USER
```
### Writing the unit by hand
**Linux (systemd user unit):**
```bash
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER
```
**macOS (LaunchAgent):**
```bash
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.codeman.web</string>
<key>ProgramArguments</key>
<array>
<string>$(which node)</string>
<string>$HOME/.codeman/app/dist/index.js</string>
<string>web</string>
</array>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
Prefer `codeman service install` where you can. It handles the PATH problem for you.
## Logs
```bash
journalctl --user -u codeman-web -f # systemd
tail -f ~/.codeman/web.log # detached mode
log stream --predicate 'process == "node"' # macOS, noisy
```
## Updating
| Install route | Update with |
| ------------- | ------------------------------------------------------------------------ |
| Installer | Re-run the one-liner, or **App Settings → System → Updates**. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
### The in-app updater
**App Settings → System → Updates**, for git-clone installs supervised by systemd or
launchd. npm installs report as non-updatable, and an unsupervised install is told to
restart manually.
The interesting part is that the update restarts the very process running it. So the real
work runs in a **detached script that outlives the restart** and writes progress to a status
file, which the browser polls across the connection drop. A dirty tree is stashed rather
than discarded.
### After updating
Sessions are unaffected: they live in tmux and the server reattaches. If the UI looks stale,
reload; on iOS Safari, close the tab completely and reopen.
## Running two instances
The data directory and the tmux socket are process wide, so a second server on the defaults
will discover and attach the first one's sessions. Scope both together:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
Service unit names are instance-scoped too, so a beta instance can be installed as its own
service without colliding with the main one. `CODEMAN_DATA_DIR` and `CODEMAN_TMUX_SOCKET`
exist for the rare case where they need to differ, but setting only one of them recreates
exactly the problem you were avoiding.
## The tunnel as a service
```bash
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
```
Or the toggle in **App Settings → System → Remote access**. See
[Remote Access](Remote-Access).
## Health checks
```bash
curl -s localhost:3000/api/status | jq '.version, .uptime'
codeman web --status
codeman doctor
```
Add `-k` and the `https://` URL on an HTTPS install.
## Read next
- [Installation](Installation) - the routes and what each supports.
- [Remote Access](Remote-Access) - exposing it once it stays up.
- [Troubleshooting](Troubleshooting) - when it does not.
+120
View File
@@ -0,0 +1,120 @@
# Security
The honest version first: **Codeman's dashboard is a remote code execution surface, by
design.** It starts agents with permission prompts skipped by default, so anyone who can
reach it can run arbitrary code as your user, on your machine. Every protection in Codeman
exists to control who that is.
That is not a flaw to be fixed. It is what "run my coding agent for me" means. The job is to
make sure the set of people who can reach it is exactly the set you intended.
## The default is safe
A bare `codeman web` binds `127.0.0.1`. Only processes on that machine can reach it, which
is why shipping with no password by default is defensible. Everything risky starts when you
expose it.
## Hardening checklist
In order of how much they matter:
1. **Do not expose it without `CODEMAN_PASSWORD`.** Binding a non-loopback host without one
starts, but warns loudly. A tunnel refuses outright unless you acknowledge the exposure
in the UI.
2. **Prefer Tailscale over a public tunnel.** Keeping the loopback bind and putting a
private network in front of it removes the public attack surface entirely, and gives you
real HTTPS. See [Remote Access](Remote-Access).
3. **Use a long password.** It is the only thing between a public URL and your shell.
4. **Consider the permission mode.** **App Settings → Agents & CLIs → Claude → Startup
Mode** can switch new sessions from skip-prompts to Anthropic's classifier-guarded `auto`
mode, to normal prompting, or to an explicit allowed-tools list.
5. **Use Docker cases for untrusted work.** If you are pointing an autonomous loop at a repo
you did not write, [Docker Cases](Docker-Cases) gives it its own filesystem and network
for the cost of one checkbox.
6. **Keep it updated.** Browser-driven attack paths were closed in 0.9.x and hardening is
ongoing.
## What protects what
These run on **every** request, including on a default no-password loopback install:
| Layer | What it stops |
| ---------------------------- | ----------------------------------------------------------------------------------------------- |
| **Host-header allowlist** | DNS rebinding. A domain rebound to `127.0.0.1` is rejected before any handler runs. Add your own domains with `CODEMAN_ALLOWED_HOSTS`. |
| **Cross-site Origin guard** | CSRF on state-changing requests. A *missing* Origin is allowed so curl, the CLI, and hooks keep working; a foreign or opaque one is rejected. |
| **Raw `text/plain` bodies** | The CORS simple-request CSRF vector, where a cross-site form could smuggle JSON into a write route with no preflight. |
| **WebSocket origin check** | Cross-site WebSocket hijacking. The terminal upgrade closes with code `4003` on failure. |
| **Output escaping** | Stored XSS from agent-derived strings: tool names, command arguments, subagent descriptions. |
| **Security headers** | A strict content security policy, `nosniff`, frame options, and HSTS over HTTPS. CORS is reflected only for loopback origins. |
When authentication is enabled:
| Layer | Behaviour |
| ------------------- | ------------------------------------------------------------------------------------------------ |
| **HTTP Basic** | `CODEMAN_USERNAME` (default `admin`) and `CODEMAN_PASSWORD`. |
| **Session cookie** | A 256-bit opaque token validated server side, so it cannot be forged offline. 24 hours, extended on activity, with a device-context audit trail. |
| **Rate limiting** | Ten failed attempts per IP produce a `429` with a 15 minute decay. A correct password or valid cookie recovers immediately even under attack, which matters because all tunnel traffic shares one loopback address. |
| **QR auth** | Single-use 60-second tokens with their own separate rate limiter, so a mistyped password cannot lock out QR login. |
| **Hook endpoints** | The hook and telemetry endpoints skip Basic auth because they are called from localhost by the CLI, but when auth is on, that bypass additionally requires a per-instance hook secret. |
## File access
Three separate file surfaces, each confined differently, because a single shared rule would
be wrong for at least one of them:
| Surface | Rules |
| -------------------- | -------------------------------------------------------------------------------------------- |
| **File Viewer** | Real path resolution before boundary checks, so symlinks cannot escape. Sensitive trees blocked. Edit mode adds an extension allowlist, a size cap, `.git` denial, and optimistic concurrency. It never creates files. |
| **Attachments** | An id-based registry, so browser requests never carry absolute paths. The magic-link scanner is prompt-injectable by nature and is therefore force-confined to the session's workspace. Extension allowlist, not a blocklist. |
| **Path picker** | Its own root allowlist rather than the workspace confinement. In multi-user mode a non-admin gets only their own user space, because per-user spaces live inside the home directory. |
Downloads block sensitive paths outright (`.env`, credentials files, `~/.ssh`, AWS
credentials), and SVG and HTML are served as downloads with `nosniff` so they cannot execute
in the page.
## Supply chain and isolation
- Security-sensitive transitive dependencies are pinned to patched versions, and lockfile
integrity is checked on every push and pull request: every entry must resolve to the public
registry with a hash.
- Public assets are scanned for NUL bytes and syntax-checked in CI.
- `CODEMAN_INSTANCE` scopes the tmux socket and the data directory together, so two
instances never attach each other's live sessions.
## What Codeman does not protect against
Stated plainly, because a security page that only lists strengths is not useful:
- **Multi-user mode is not a sandbox.** It separates workspaces. Every session still runs as
the same OS account, so a determined user's agent can reach another user's files. For real
isolation, pair users with Docker cases or run separate instances under separate OS
accounts.
- **An agent you gave shell access can do anything you can.** Permission modes narrow this;
they do not remove it.
- **A tunnel makes your machine reachable from the internet.** The password is the whole
boundary. Treat it accordingly.
- **Codeman cannot detect your own loopback reverse proxy**, which is why the hook-endpoint
bypass requires a secret unconditionally when auth is on.
- **The agent CLIs have their own trust models.** Pi's project trust executes repo-local
TypeScript, for instance. See [Agent CLIs](Agent-CLIs).
## Privacy
No telemetry, no analytics, no phone-home. Codeman's only network traffic is between your
browser and your server. Your agent CLI's traffic is its own, on your account.
Two features send data outward, both off by default and both stated where they appear: voice
dictation through your Claude login, and the Read My Mind prediction call.
## Reporting a vulnerability
**Never in a public issue.**
[SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md) has the
private disclosure process and the current list of known limitations.
## Read next
- [Remote Access](Remote-Access) - the safe ways to expose it.
- [Multi-User Mode](Multi-User-Mode) - what it does and does not separate.
- [Docker Cases](Docker-Cases) - real isolation for untrusted work.
- [`docs/security-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/security-architecture.md) - the complete model.
+173
View File
@@ -0,0 +1,173 @@
# Settings Reference
Two settings surfaces, and the rule that explains why a setting you changed on your laptop
did not follow you to your phone.
| Surface | Scope | Opened from |
| ------------------- | ------------------------------ | ---------------------------- |
| **App Settings** | Global, this Codeman install. | The header gear. |
| **Session Options** | One session. | The session's tab. |
App Settings is a single scrolling document with a rail acting as a table of contents;
clicking a rail entry scrolls rather than switching. Session Options genuinely switches
panels.
## Per-device versus synced
Some settings live on the server and follow you to every device. Others are stored in the
browser and stay put. This is deliberate, not an oversight: your phone wants a different
font size, a different keyboard bar, and a different set of header buttons than your
desktop.
| Category | Examples |
| ----------------------- | ------------------------------------------------------------------------------- |
| **Per-device, local** | Skin, WebGL renderer, local echo, CJK input, extended keyboard bar, File Viewer and Cron header buttons. Never sent to the server at all. |
| **Per-device policy** | Most `show*` toggles, plan usage chip, language. Stored server-side, but a device only takes the server value when it has no local one of its own. |
| **Synced** | Models, effort, CLI options, notification preferences, voice settings, display name, the agent skill and approvals toggles. |
The practical rule: **appearance and input are per device, behaviour is shared.** If a change
did not follow you, it is in one of the first two rows, and you change it again on that
device.
## App Settings
### Updates
Current version, a manual check, and the in-app updater. Covers git-clone installs
supervised by systemd or launchd; npm installs report as non-updatable. See
[Running As A Service](Running-As-A-Service).
### Terminal & Input
| Setting | Default | Notes |
| ----------------------------- | -------------------- | --------------------------------------------------------------------- |
| Local Echo | On for touch devices | Paints keystrokes locally and flushes on Enter. See [Input And Voice](Input-And-Voice). |
| CJK Input | Off | IME composition through a dedicated text field. |
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
### Header & Panels
Chips for every optional header control, with a live preview of the resulting header:
Run, Font Size, System Stats, Redraw Terminal, Response Viewer, Away Digest, Session
Manager, Attachments, File Viewer, Multi-monitor, Plan Usage, Lifecycle Log, Monitor,
Project Insights, File Browser, Subagents, Approvals Inbox, Read My Mind, Ultracode Agents,
Ultracode Windows, Cron.
Most default to off. The stock desktop header is system stats, File Viewer, and the gear.
New header controls never appear on phones.
This section also holds background-agent tracking, including whether to track agents for
every session or only the active tab.
### Appearance
| Setting | Notes |
| ---------------------- | ----------------------------------------------------------------------------------------- |
| Skin | Theme palettes, light ones included. Applied before first paint, so no flash of the wrong theme. |
| Entrance Animations | Per-surface animation styles for tabs, terminals, windows, and lineage lines. All default to the legacy no-animation behaviour. |
| Display Name | Your name in the UI. Cosmetic only; it never renames the package, CLI, API, or storage. |
| Interface Language | English or Simplified Chinese. Per device. |
| Session List Layout | Header tab strip (default) or a collapsible left sidebar. See [The Dashboard](The-Dashboard#session-list-layout). |
| Tall Tabs | Taller tab strip. |
| Pop-out Button on Tabs | Adds the detach control to tabs, with a per-tab override. |
| Spawn Lineage Lines | Arcs from a parent tab to sessions it spawned. Desktop only, on by default. |
| Overview Home Screen | The phone home screen. On by default. |
### Models
Claude model cards, the 1M context window switch, and the thinking effort segment. The cards
and the switch compose into one model choice, so there is no separate "which one wins"
question.
Model and effort are both **soft defaults**: the model is written into the case's
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
### Agents & CLIs
| Setting | Notes |
| -------------------------------- | -------------------------------------------------------------------------------------------- |
| Startup Mode | Claude's permission mode for new sessions. Default skips prompts; `auto` uses Anthropic's classifier-guarded mode; `normal` prompts; or give an explicit allowed-tools list. |
| Allowed Tools | The list used by the explicit mode. |
| Ralph / Todo Tracker | Enables the Ralph loop surfaces. |
| Agent Teams | Experimental teams. Also needs the CLI's own environment flag. |
| Codeman Agent Skill | Injects the agent skill into new Claude sessions per case. Off by default. See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent). |
| Remote auto-reconnect | Reattaches dropped remote SSH sessions. On by default. |
| Nice priority / value | Runs agent processes at a lower CPU priority. |
| Bypass approvals and sandbox | Pi's project trust. Read [Agent CLIs](Agent-CLIs) before enabling. |
| Animated status effects | Cosmetic. |
### Notifications
Master toggle, browser notifications, push subscription, audio alerts, and the idle
threshold that decides when a quiet session counts as needing you. See
[Notifications And Approvals](Notifications-And-Approvals).
### Voice
Active provider and the engine behind it, insert mode, language, domain keywords to bias
recognition, the Deepgram API key, and the opt-in switch for transcribing through this
server's Claude login, with its live credential status. See
[Input And Voice](Input-And-Voice).
### Shortcuts
Rebinding for the shortcut registry. See [Keyboard Shortcuts](Keyboard-Shortcuts).
### System
`CLAUDE.md` template for new cases, default working directory, the image watcher, and
Cloudflare tunnel controls including the tunnel and upload URLs. In multi-user mode, the
**Users** administration entry is injected here.
## Session Options
Per session, from the tab.
| Panel | Contains |
| ---------------- | ------------------------------------------------------------------------------------------- |
| **Respawn** | Auto-resume on usage limit, the respawn cycle configuration, presets, duration. See [Keeping Agents Running](Keeping-Agents-Running). |
| **Session** | Name, working directory, environment overrides, per-tab pop-out override. |
| **Ralph / Todo** | Loop configuration, iteration and todo caps, circuit breaker reset. See [Autonomous Loops](Autonomous-Loops). |
| **Summary** | What this session has done: tokens, activity, run summary. |
Panels that only make sense for Claude are hidden for other run modes rather than shown and
failing.
## Environment variables
Some things are configured before the server starts, not in the UI:
| Variable | Effect |
| ----------------------------------- | ---------------------------------------------------------------------- |
| `CODEMAN_PORT` | Listen port. |
| `CODEMAN_HOST` | Bind address. Loopback by default. |
| `CODEMAN_PASSWORD` / `CODEMAN_USERNAME` | HTTP Basic credentials. Username defaults to `admin`. |
| `CODEMAN_ALLOWED_HOSTS` | Extra Host and Origin allowlist entries for a reverse proxy. |
| `CODEMAN_INSTANCE` | Scopes the data directory and tmux socket together. Required for a second instance. |
| `CODEMAN_MULTIUSER` | Enables multi-user mode. |
| `CODEMAN_GESTURE` | Makes gesture control available to be enabled. |
| `CODEMAN_DOCKER_BRIDGE_HOOKS` | Lets in-container hooks reach the host on a loopback bind. |
| `CODEMAN_FILE_PICKER_ROOTS` | Extra roots for the path picker. |
| `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledges exposing the server with no password. |
## Gotchas
- **A setting that did not sync is per device.** Change it again on that device.
- **The plan usage chip and its telemetry exporter are one setting.** Enabling the chip
without the exporter would leave it blank forever, so it is deliberately not separable.
- **Toggling a header button does nothing on a phone.** Phones deliberately ignore most of
the header chips.
- **Enabling a feature does not retroactively configure existing sessions.** The agent skill
injection, for instance, applies at session creation.
## Read next
- [The Dashboard](The-Dashboard) - what each control does once visible.
- [Keeping Agents Running](Keeping-Agents-Running) - the Respawn panel in depth.
- [Agent CLIs](Agent-CLIs) - model, effort, and permission modes.
+206
View File
@@ -0,0 +1,206 @@
# The Dashboard
What the interface is telling you, and which parts of it are hidden until you turn them on.
Most of Codeman's UI is **opt-in**. A stock install shows a deliberately small header, and a
feature you read about here may simply not be on screen yet. Where that is the case, this
page says so and names the setting.
![Codeman dashboard](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/codeman-tour-20260724.png)
## Layout
| Region | What lives there |
| ------------------ | -------------------------------------------------------------------------------------- |
| **Header, left** | The "C" logo (goes home) and the session list, unless you moved it to the sidebar. |
| **Header, right** | Status chips and panel buttons, most of them off by default. |
| **Center** | The terminal for the active session, or the home screen when nothing is selected. |
| **Bottom toolbar** | Run, Stop, Run Shell, the case picker, and the instance counters. |
| **Overlays** | Panels and modals: Respawn, Cron, Subagents, File Viewer, Settings. |
## Session list layout
The session list lives in the header as a horizontal strip by default. With a lot of
sessions open that strip stops being scannable, so **App Settings → Appearance → Tabs →
Session List Layout** can move it into a vertical sidebar on the left instead.
| Layout | Behaviour |
| -------------------- | --------------------------------------------------------------------------------- |
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
per device, so a sidebar on your desktop does not force one onto your phone.
## Session tabs
One tab per session, in your order, and that order syncs across your devices.
**Status is carried by the dot and the tab's own styling:**
| Look | Meaning |
| ----------------------------- | ----------------------------------------------------------------------- |
| Green dot | Alive, not currently working. |
| Pulsing green dot with a ring | Working on a turn. |
| Yellow tab, blinking | The agent is waiting for input from you. |
| Red tab, blinking | A question or permission prompt is blocking the session. |
| No dot | The session is not running. |
![Tab alerts](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/tab-alerts-20260815.png)
The alert states are steady colour with a pulse layered on top, not a blink between the
alert colour and nothing, so a tab that needs you looks like it needs you at every point in
the cycle. They survive a page reload: the state is re-seeded from the server on load, so
reloading while a permission prompt is blocking does not lose the red tab.
**Navigation:**
| Action | Keys |
| ------------------------------- | ------------------------------------------------------- |
| Jump to tab N | `Alt+1` to `Alt+9` (the number on the tab) |
| Next / previous | `Ctrl+Tab`, `Alt+[`, `Alt+]` |
| Move the active tab | `Ctrl+Shift+{`, `Ctrl+Shift+}` |
| Close | `Ctrl+W` |
| Find any session, open or past | `Ctrl+K` (also `Cmd+K` and `Alt+K`) |
Tabs can also be dragged to reorder.
On phones the strip scrolls horizontally instead of wrapping, and the active tab is always
scrolled into view. It is not reordered to the front, so the `Alt+N` numbering stays stable.
### Lineage arcs
When one session spawns another (an agent starting a worker through the API), Codeman draws
a coloured arc under the strip connecting parent to child, with one colour per child. It is
how a fan-out of eight workers stays readable.
Desktop only, and on by default. Turn it off in **App Settings → Appearance**. Arcs are
skipped for tabs scrolled out of the strip.
## Header controls
The right side of the header. Almost all of these are off until you enable them in
**App Settings → Header & Panels**.
| Control | Default | What it does |
| ---------------------- | ------------------ | ------------------------------------------------------------------------------- |
| Connection dot | Always on | SSE connection health. Green is connected. |
| Font size `-` / `+` | Always on | `Ctrl +` / `Ctrl -` do the same. |
| CPU / MEM bars | On | Server resource use. |
| File Viewer | On | Toggles the file browser panel. |
| Settings gear | Always on | App Settings. |
| Plan usage chip | On, desktop only | Live Claude subscription usage. Claude-only, and needs its telemetry exporter, which the same setting installs. |
| Session Manager | Off | The full session list, live and historical. |
| Approvals bell | Off | Cross-session queue of prompts waiting on a human. Appears only when the count is above zero. Never shown on phones. |
| Read My Mind 🧠 | Off | Predicts your next prompt for this case. Claude-only. |
| Attachments | Off | Registered external files. |
| Away Digest | Off | What happened while you were gone. |
| Last Response | Off | Readable view of the agent's last answer, useful on phones. |
| Ultracode / Workflow | Off | Live workflow-run agents. |
| Notifications | Off | Notification history and settings. |
| Lifecycle Log | Off | Session start, exit, and kill audit trail. |
| Cron ⏰ | Off | Scheduled jobs. |
| Multi-monitor | Off, macOS | Opens a window spanning every display. |
| Tunnel indicator | When a tunnel runs | Cloudflare tunnel status. |
| Admin panel | Multi-user only | User administration. |
New header controls never appear on phones. Phone layout is deliberately minimal and is
covered in [Mobile Guide](Mobile-Guide).
## Connection state
The dot in the header is the quick read. Two louder surfaces exist because a cached page
with no server behind it used to look identical to a page with no sessions:
- **A full-screen overlay** when the page has never loaded server state. There is nothing
behind it worth preserving.
- **A banner** when the connection drops after state had loaded, so your scrollback stays
readable.
Both wait about 2.5 seconds before appearing, so a deploy that restarts the server does not
flash a warning at you every time. If the browser reports itself offline, the grace period
is skipped.
There is also a watchdog for the case where the connection stops delivering without
erroring. If the server's heartbeat stops arriving, Codeman reconnects on its own rather
than sitting on a green dot showing frozen data.
## The terminal
A real terminal: xterm.js in the browser, a real PTY on the server, tmux in between. Full
TUIs render correctly.
Worth knowing:
- **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open.
Shell sessions open from a bounded recent tail so a large transcript cannot stall tab
switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling
stays within the bounded browser buffer so dragging upward remains responsive.
- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude
versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is
always local scrollback. Other CLIs scroll locally.
- **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not.
`Ctrl+Shift+C` always copies.
- **Zero-lag input.** On touch devices, keystrokes paint locally before the round trip. See
[Input And Voice](Input-And-Voice).
- **Renderer.** WebGL by default, with a watchdog that falls back to DOM rendering if the
GPU stalls. `?nowebgl` forces DOM rendering for one page load.
## The home screen
With no session selected you get the welcome screen: run buttons for the CLIs Codeman
found, a QR code when a password is set, cross-session search, and **Resume Conversation**,
which lists past sessions including Claude conversations started outside Codeman entirely.
Two extras depending on the device:
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in tab
order, with created and last-active stamps. It needs at least 1180px of width; below that
it is hidden so it cannot overlap the search panel.
- **Phones**: tapping the "C" logo gives a session overview instead: NEEDS YOU first, then
current sessions, then past ones. On by default.
## Panels
| Panel | Opened from | Covered in |
| ---------------- | --------------------------------- | ---------------------------------------------------------------- |
| Respawn | Session Options | [Keeping Agents Running](Keeping-Agents-Running) |
| Ralph | Session Options | [Autonomous Loops](Autonomous-Loops) |
| Orchestrator | Toolbar | [Autonomous Loops](Autonomous-Loops) |
| Cron | Header ⏰ (opt-in) | [Cron Jobs](Cron-Jobs) |
| Subagents | Automatic while agents run | [Watching Agents Work](Watching-Agents-Work) |
| Ultracode | Header (opt-in) | [Watching Agents Work](Watching-Agents-Work) |
| File Viewer | Header | [Working With Files](Working-With-Files) |
| Attachments | Header (opt-in) | [Working With Files](Working-With-Files) |
| Approvals | Header bell (opt-in) | [Notifications And Approvals](Notifications-And-Approvals) |
| App Settings | Header gear | [Settings Reference](Settings-Reference) |
Session-specific configuration lives in **Session Options**, reachable from the tab. App
Settings is global; Session Options is per session.
## Search and the session palette
`Ctrl+K` opens the session palette: every session, live or historical, filtered as you
type. Picking a past one resumes its conversation.
The search box on the home screen is wider in scope. It federates over session metadata,
run-summary events, and attachment history, filtered by type, case, status, and date. It
does substring matching over data already in memory, with no regex and no filesystem reads,
so it is fast and cannot be turned into a traversal.
## Appearance
**App Settings → Appearance** carries the theme skins, including light ones. The choice is
applied before the first paint, so there is no flash of the wrong theme on load.
The same section has the entrance animations for tabs, terminals, agent windows, and
lineage lines. All of them default to the legacy no-animation behaviour, so an untouched
install animates nothing.
## Read next
- [Keyboard Shortcuts](Keyboard-Shortcuts) - the full list, and how to rebind.
- [Settings Reference](Settings-Reference) - every setting, and why some follow you across devices and others do not.
- [Mobile Guide](Mobile-Guide) - what changes on a phone.
- [Watching Agents Work](Watching-Agents-Work) - subagent windows and workflow runs.
+290
View File
@@ -0,0 +1,290 @@
# Troubleshooting
Symptom first. Find the line that matches what you are seeing.
Before anything else, check what version you are on and whether the problem is already
fixed:
```bash
codeman --version
codeman doctor
```
## Installing and starting
### `Failed to start claude: error: posix_spawnp failed` on macOS
node-pty ships its macOS `spawn-helper` without the executable bit, and macOS launches
every PTY through it. Codeman detects this and repairs it on the first failure, so updating
usually fixes it outright. To repair by hand on a clone install:
```bash
npm run fix:node-pty
```
It is a `chmod`, not a rebuild, so it does not need Xcode command line tools. The helper
lives in `prebuilds/darwin-<arch>/`, not `build/Release/`, which does not exist on macOS.
Linux never sees this.
### `tmux: command not found`
The installer asks before installing packages and remembers a declined answer. Install tmux
and start again. There is no tmux-free mode: sessions live in tmux.
### The port is already in use
```bash
codeman web --port 8080 # or set CODEMAN_PORT
```
If you believe nothing is on 3000, check for a Codeman you already started:
```bash
codeman web --status
```
### The terminal area is blank, and the console mentions a missing vendor file
Clone installs build the vendored xterm addon bundles in `postinstall`. If `npm install`
was interrupted or run with `--ignore-scripts`, those bundles are missing:
```bash
npm install
```
They are intentionally not committed to the repository.
### `Case path not found` when clicking Run
The case points at a directory that no longer exists, usually because it was deleted or
moved outside Codeman. Re-link the case, or create it again.
### The server starts but nothing is reachable
That is the default behaviour, not a failure. Codeman binds `127.0.0.1`. See
[Remote Access](Remote-Access).
## Reaching the interface
### The dashboard will not load from another device
Check, in order: the bind (loopback by default), a firewall, and then
[Remote Access](Remote-Access) for a supported way to expose it.
### `403 host not allowed`
The Host header is not in the allowlist, which is the DNS-rebinding guard doing its job. Add
your domain:
```bash
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
A leading dot matches subdomains.
### The page loads but the terminal never connects
The terminal is a WebSocket. Behind a reverse proxy, the upgrade must be forwarded. The
upgrade also runs the Host and Origin checks and closes with code `4003` when they fail.
### The UI looks stale after updating
The app shell is cached by a service worker, and static assets are served with a long cache
lifetime. `index.html` is not cached, and every asset reference is version-stamped, so a
normal reload picks up a new build.
Two exceptions worth knowing:
- **iOS Safari** can keep serving old JavaScript until the tab is fully closed, not just
reloaded. Close the tab and reopen it.
- If you edit files in dev, changes to `index.html` need a server restart. Changes to `.js`
and `.css` do not.
### A full-screen "cannot reach the server" overlay appears
The server is genuinely unreachable, or the connection dropped. Codeman waits about 2.5
seconds before showing it, so a quick restart does not flash it. Retry re-arms both the
event stream and the terminal socket.
## Sessions
### A session shows idle while it is clearly working
Update. Claude redraws its prompt roughly once a second throughout a turn, and older idle
detection treated that as the end of the turn, flipping working sessions to idle a couple of
seconds in. Current versions confirm against the actual screen before believing it.
### A session is stuck showing busy
For non-Claude CLIs, idle detection is output-based and coarser by necessity: those CLIs
expose no hooks. A session that has genuinely gone quiet will settle. If it never does,
interrupt it (`Ctrl+C` with nothing selected).
### The agent asks about bypass permissions every time
That prompt comes from Claude Code, not Codeman. Codeman's default is to start with
permission prompts skipped, which is what the security model is built around. If you would
rather it prompted, change **App Settings → Agents & CLIs → Claude → Startup Mode**.
### Sessions vanished after a reboot
Expected. tmux does not survive a reboot, so the sessions are gone. Conversations are not:
Claude transcripts persist, so the welcome screen's **Resume Conversation** list can pick
them back up.
### A session restarts, then refuses to restart again
That is the PTY-exit circuit breaker. Repeated rapid PTY exits trip it, and it blocks
automatic restarts so a broken configuration does not spin forever. Reset it explicitly from
the session's controls. Reattaching does not clear it, deliberately.
### Sessions I did not create appeared, or my session resized itself
Two Codeman servers are running against the same data directory and tmux socket. The second
one discovers and attaches the first one's sessions. Give each instance its own scope:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
`codeman web -d` and `codeman service install` both refuse to start a second server on one
data directory for exactly this reason.
## The terminal
### I cannot scroll back through history
Scrollback behaviour depends on the CLI, and Codeman adjusts what it strips per mode.
Things to try:
- `Shift+Wheel` always scrolls the local buffer, whatever else is going on.
- On Claude sessions with a recent CLI, the wheel is forwarded into Claude's own transcript,
so it scrolls the conversation rather than the terminal buffer. That is intended.
- Scrolling to the very top pulls the full tmux scrollback again on demand.
### The wheel does nothing in a Codex session
Codex ignores the mouse reports that forwarding would send, so Codeman does not forward
there. Scrolling is local, and `Shift+Wheel` behaves the same way.
### `Ctrl+C` copies when I wanted to interrupt
With a selection, `Ctrl+C` copies. With no selection, it interrupts. Clear the selection
first, or use the **Stop** button. `Ctrl+Shift+C` always copies and never interrupts.
### I typed a prompt but nothing was sent
On touch devices, keystrokes are painted locally and flushed when you press Enter, so text
on screen has not necessarily reached the agent yet. Press Enter, or the phone toolbar's
**Enter** button.
If you are sending input over the API instead, your payload must end with `\r` or no Enter
is ever sent. The request still succeeds and the text sits unsubmitted in the composer. See
[Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
## Mobile
### The keyboard covers the terminal, or scroll position jumps
Update first; several rounds of fixes have gone into keyboard resize and scroll restoration.
### I cannot reach the rightmost tabs
The strip scrolls horizontally on phones and the active tab is scrolled into view
automatically. Swipe the strip itself. If a background render snaps you back, update.
### The space key does nothing on Android
A long-standing Android keyboard bug, fixed some time ago. Update.
### The keyboard will not close
Tap outside the terminal, or tap twice on inert terminal content. Tapping a control does not
dismiss it, by design.
## Agents and CLIs
### A CLI is installed but Codeman does not offer it
Codeman resolves binaries from the environment the **server** runs in.
```bash
codeman doctor
```
If it runs as a service, launchd gives the job a minimal PATH. `codeman service install`
bakes your PATH into the unit; a hand-written plist does not. Restart the server after
installing a new CLI.
### Hooks stopped working after switching to HTTPS
Hook callbacks have to accept the self-signed certificate. Recent versions self-heal
existing cases; if yours predates that, recreate the case so its hooks are rewritten.
### The model or effort I chose is not being used
Both are **soft defaults**, on purpose. The model is written into the case's
`.claude/settings.local.json` and effort is passed on the command line at start, so `/model`
and `/effort` inside the session override them at any time. Effort is deliberately never
passed as an environment variable, because that hard-locks it.
### Tab alerts and approvals never fire in one of my repos
That case is missing its hooks block. Recreating the case rewrites it.
## Docker and remote
### Docker sessions do not detect idle
On a loopback-only bind, a container cannot reach `127.0.0.1` on the host, so in-container
hooks have nothing to call. Set `CODEMAN_DOCKER_BRIDGE_HOOKS=1` to open a hooks-only
listener on the docker bridge gateway. Without it, idle detection falls back to output
watching.
### A rebuilt agent image still has old CLI versions
Always rebuild with `--no-cache`:
```bash
node scripts/build-agent-image.mjs --no-cache
```
A plain rebuild reuses the cached `npm install -g` layer and keeps the CLIs frozen at their
original versions while reporting success.
### A remote SSH session dropped and did not come back
A bounded-backoff watcher reattaches dropped sessions, and it is on by default. Intentional
kills are never revived. Check the host is reachable and that the remote tmux server is
still running.
## Gathering diagnostics
```bash
codeman doctor # dependency check
curl -s localhost:3000/api/status | jq # full app state
tmux -L codeman list-sessions # what tmux thinks is alive
journalctl --user -u codeman-web -f # service logs (Linux)
tail -f ~/.codeman/web.log # detached mode logs
```
On an HTTPS install, add `-k` to the curl commands and use the `https://` URL.
## Filing a good bug report
Open an [issue](https://github.com/Ark0N/Codeman/issues) with:
- OS and version.
- Install method: installer, npm, or git clone.
- `codeman --version`.
- Browser and version, if the problem is in the UI.
- Which CLI the session was running, and its version.
- What you did, what happened, what you expected.
Reports usually get a response within a day, and every release credits its reporters by
name.
Questions and setup help fit better in
[Discussions](https://github.com/Ark0N/Codeman/discussions). Security problems never go in a
public issue; see
[SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
+72
View File
@@ -0,0 +1,72 @@
# Versioning
Codeman follows [semantic versioning](https://semver.org/). This page says what the version
number actually promises, which matters if you are building anything against Codeman.
## Covered by the version number
Breaking any of these after 1.0 requires a **major** bump:
1. **The CLI.** Command names, documented flags, and their behaviour. The npm package is
`aicodeman` and installs both the `aicodeman` and `codeman` commands; renaming either is
breaking.
2. **The HTTP API and SSE channel**, served under `/api/v1` with the uniform envelope and
conventional status codes. Endpoint paths, the envelope, `errorCode` values, and SSE event
names are all stable.
3. **Documented deployment environment variables**: `CODEMAN_PASSWORD`, `CODEMAN_USERNAME`,
`CODEMAN_HOST`, `CODEMAN_PORT`, `CODEMAN_INSTANCE`, `CODEMAN_ALLOWED_HOSTS`,
`CODEMAN_DATA_DIR`, `CODEMAN_TMUX_SOCKET`, plus the `--host`, `--port`, and `--https`
flags.
4. **The published `xterm-zerolag-input` library**, on its own independent version line.
Codeman reaching 1.0 says nothing about that package's version.
Additive changes are **not** breaking: new endpoints, new optional fields, new error codes,
new SSE events. Genuinely breaking API changes would ship under a new prefix rather than
changing `/api/v1`.
## Not covered
These can change in a minor or even patch release:
1. **The `~/.codeman/` state file formats.** Migrations are made on a best-effort basis and
have been done across renames, but the on-disk shape is not a contract. Do not write
tooling against it.
2. **Internal TypeScript modules.** The npm package is CLI-only. There is no stable library
entry point, and importing it programmatically is unsupported.
3. **Experimental and opt-in features**, whatever the app's version: gesture control, agent
teams, and anything labelled experimental in the UI or docs.
## Deprecation
- Additive changes are preferred over breaking ones.
- A covered surface slated for removal is deprecated first: it keeps working for at least one
minor release, with a runtime warning and a changelog note pointing at the replacement,
then is removed in the next major.
- Backwards-compatibility shims are kept until a major boundary.
## Releases
Releases are managed with changesets. Every release:
- Bumps the version and updates
[`CHANGELOG.md`](https://github.com/Ark0N/Codeman/blob/master/CHANGELOG.md).
- Publishes to npm as `aicodeman`.
- Cuts a GitHub release, tagged `codeman@X.Y.Z`.
- **Credits its contributors and bug reporters by name** in the release notes.
There is no fixed cadence. Patches ship when fixes are ready, which in practice is often.
## Which version am I on?
```bash
codeman --version
```
Or **App Settings → Updates**, which also checks for a newer one and can install it. See
[Running As A Service](Running-As-A-Service).
## Read next
- [HTTP API](HTTP-API) - the stable API surface itself.
- [Contributing](Contributing) - how changes get made.
- [`docs/versioning-policy.md`](https://github.com/Ark0N/Codeman/blob/master/docs/versioning-policy.md) - the authoritative statement.
+112
View File
@@ -0,0 +1,112 @@
# Watching Agents Work
Modern agents fan out. A single Claude session can be running six subagents, and the parent
terminal shows you almost none of it. Codeman surfaces that hidden work as live windows,
panels, and after-the-fact summaries.
Everything on this page is Claude-only. It reads Claude Code's transcripts and team state;
the other CLIs expose no equivalent.
![Subagent windows](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/subagent-windows-20260724.png)
## Subagent windows
When a Claude session spawns subagents, each one gets its own floating window with a live
transcript: what it was asked to do, what it is doing, and what it returned.
- Windows are draggable and resizable, and their positions persist across reloads.
- A connection line links each window to the session tab that spawned it, so with four
sessions running you can still tell whose worker is whose.
- Closing a window does not stop the subagent. It only stops you watching it.
This is the feature that makes a fan-out legible. Without it, a lead session that spawned
eight workers looks like a stalled terminal for several minutes.
## Session lineage arcs
The tab strip draws a coloured arc from a parent tab to any tab it spawned, one colour per
child. That covers the other direction of fan-out: not subagents inside one session, but
whole sessions started by an agent through the API.
Desktop only, on by default, and toggled in **App Settings → Appearance**. Arcs are skipped
for tabs scrolled out of view.
See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) for the spawning side.
## Agent teams
Claude Code's experimental agent teams appear as teammates alongside subagents. Enable them
in the CLI's own environment:
```bash
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
and turn the per-case **Agent Teams** toggle on in the case settings gear.
Codeman watches the team directory and matches teammates to the session leading them.
Teammates are in-process threads rather than separate CLI processes, so they show up as
windows, not tabs.
Notes and the experiment log:
[`docs/agent-teams/`](https://github.com/Ark0N/Codeman/tree/master/docs/agent-teams).
## Ultracode and workflow runs
When Claude runs a Workflow, dozens of agents can be in flight at once. The completion
artifact for a run is only written at the **end**, so a live run would otherwise be
invisible until it finished. Codeman synthesizes the in-flight view from the transcripts and
lets the real artifact supersede it when it lands.
Two independent toggles, both off by default:
| Setting | Shows |
| ---------------------- | ----------------------------------------- |
| Ultracode panel | A docked panel listing the run's agents. |
| Ultracode windows | Floating windows, like subagents. |
Turning on either starts the watcher.
## Reading the answer, not the terminal
**Last Response** (header button, opt-in) renders the agent's last answer as scrollable text
rather than terminal output. It exists mostly for phones, where reading a long answer in a
terminal viewport is painful. **More** loads additional context.
## After the fact
| Surface | Answers |
| ------------------ | -------------------------------------------------------------- |
| **Away Digest** | What happened while I was gone? |
| **Run summary** | What did this run actually do? |
| **Lifecycle log** | When did sessions start, exit, or get killed, and why? |
| **Token stats** | What did it cost? |
The Away Digest aggregates the lifecycle log, run summary events, live sessions, token
statistics, and recent subagents into one view. It is the right first thing to open in the
morning after an overnight run.
All of these header buttons are opt-in: **App Settings → Header & Panels**.
## Performance
The design target is 20 sessions and 50 agent windows at 60fps. If you routinely run more
than that, expect the browser rather than the server to be the limit, and close windows you
are not reading.
## Gotchas
- **A session pointed at a relocated Claude config directory goes blind here.** Transcripts
written outside `~/.claude/projects` are invisible to the watchers, so subagent windows,
the ultracode panel, the response viewer, and Read My Mind all stop working for that
session. Symlink `projects` back into the shared tree to fix it. See
[Agent CLIs](Agent-CLIs).
- **Closing a window does not cancel the agent.** Nothing on this page controls agents; it
observes them.
- **Windows are opt-in for ultracode, automatic for subagents.**
## Read next
- [The Dashboard](The-Dashboard) - where these surfaces live.
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the other kind of fan-out.
- [Autonomous Loops](Autonomous-Loops) - the loops that generate this much activity.
+101
View File
@@ -0,0 +1,101 @@
# Web Tabs
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port 4000, as
a tab beside your agent sessions. Codeman becomes one mission control instead of Codeman
plus a pile of browser tabs.
A web tab is **not a session**. There is no PTY, no tmux, and no respawn behind it, the same
way a docker case is not a run mode.
## Adding one
1. Click the chevron next to **Run**.
2. Under **Web / URL**, pick **Add URL**.
3. Name it, paste the URL, optionally hit **Test**, and **Save**.
It opens immediately and appears in the dropdown from then on. Web tabs share the tab strip
with sessions, continue the same `Alt+1` to `Alt+9` numbering, and carry a globe icon so
they never read as a running agent.
**Closing a tab is not deleting it.** The tab's `x` closes; the `x` on its **dropdown row**
deletes the saved dashboard. Each dropdown row also has a gear for editing the URL.
Switching tabs does not reload a dashboard. Frames stay alive in the background, so one that
took a while to authenticate is still there when you come back. Past six live frames, the
least recently viewed is dropped to bound memory.
## Why dashboards are proxied
A plain cross-origin iframe fails three ways at once in the setup Codeman actually ships in:
| Blocker | What happens |
| ------------------- | ----------------------------------------------------------------------------------------------- |
| **Mixed content** | Production is HTTPS, and browsers hard-block `http://` iframes on an HTTPS page. No override, and none at all on iOS Safari. |
| **Framing refusal** | Grafana, Portainer, Home Assistant and many others send `X-Frame-Options: DENY`. |
| **Codeman's CSP** | `default-src 'self'` blocks a cross-origin frame before it starts. |
So by default the dashboard is served **through Codeman's own origin**: the browser loads a
path on Codeman, and Codeman relays to the dashboard, stripping the framing refusal,
rewriting redirects, cookies and root-absolute URLs, and relaying WebSockets so live panels
still update.
A useful side effect: the dashboard is fetched **by the Codeman server**, so a tailnet-only
or localhost-only dashboard works from any device that can reach Codeman, including a phone
that is not on your tailnet.
There is also a `direct` mode, a plain cross-origin iframe, which is cheaper but only works
for an HTTPS dashboard that permits framing.
## The Test button, and what it does not test
**Test** probes from the server and tells you which mode applies. It verifies
**server-to-upstream reachability and nothing else**. It does not exercise the browser
sandbox, cookies, CORS, CSP, or any reverse proxy in front of Codeman.
A passing Test does not guarantee the embedded page renders.
## The sandbox, and when to turn it off
Because a proxied dashboard is served from Codeman's own address, the browser considers it
same-origin with Codeman. Unchecked, its JavaScript could read the Codeman page and call the
API that spawns agents.
So the frame is sandboxed **without** same-origin access by default. The page runs in an
opaque origin: it cannot touch Codeman, and it gets no cookies or local storage of its own.
Unchecking **Open sandboxed** grants a real origin. Do that only for a dashboard you fully
trust, and only when you need it, which in practice means one with its own login that stores
a session in a cookie.
Either way, Codeman never forwards its own credentials upstream. The `Authorization` header
and the `codeman_session` cookie are stripped on the way out, so `CODEMAN_PASSWORD` cannot
leak into a dashboard.
## Known incompatibility: cookie-authenticated reverse proxies
If Codeman itself sits behind Cloudflare Access, Authelia, oauth2-proxy, or similar, a
**sandboxed** tab may render unstyled or broken while the Codeman page around it works fine.
The reason: an opaque-origin frame's stylesheet, script, and API requests do not carry the
proxy's authentication cookie. The proxy redirects them to the login provider, and CORS or
CSP kills them there.
Trusted mode keeps a real origin and the cookie, so it works. Test cannot catch this, because
it checks the server's reach, not the browser's.
## Security notes
The proxy authenticates on an in-memory capability embedded in the path, which is why it is
exempt from the cookie and Origin checks that every API route enforces. That exemption is
fenced to safe methods and non-API paths, and there is a test pinning it in place.
Two failure modes that only appear inside a sandboxed frame, and that curl can never
reproduce, are handled: runtime-built root-absolute URLs escaping the injected base, and
same-host requests being CORS-checked with a null origin. Both present as the dashboard's own
"Failed to fetch" while the page itself renders fine.
## Read next
- [The Dashboard](The-Dashboard) - the tab strip these share.
- [Security](Security) - why the sandbox default is what it is.
- [`docs/web-tabs.md`](https://github.com/Ark0N/Codeman/blob/master/docs/web-tabs.md) - the full reference.
+159
View File
@@ -0,0 +1,159 @@
# Working With Files
Reading, editing, attaching, and previewing files without leaving the dashboard. Useful on
a desktop; on a phone it is the difference between reviewing an agent's work and waiting
until you get home.
## The File Viewer
A panel that browses the active session's working directory. Its header button is on by
default; if it is missing, re-enable it in **App Settings → Header & Panels**.
It renders what it can:
| Kind | Behaviour |
| ------------------------ | ------------------------------------------------------------------------- |
| Text and code | Syntax-aware preview. Long files are truncated in plain preview. |
| Images | Inline. |
| Audio and video | Inline with a working scrub bar, because range requests are supported. |
| PDF and Office documents | Converted for preview when a converter is available. |
| Anything else | Download. |
Caps: 10 MB for text preview, 50 MB for raw and download. Sensitive paths (`.env`, anything
matching credentials, `~/.ssh`, AWS credentials) are blocked from download, and SVG and HTML
are served as downloads rather than rendered, so they cannot execute in the page.
Closing the preview pauses and unloads any playing media. A video that keeps playing after
you close the panel means you are on an old version.
## Editing in place
Text files can be edited and saved directly in the viewer. Click the pencil in the preview
header, edit, **Save**.
The guardrails are worth knowing, because they are what makes editing safe rather than
convenient:
- **Extension allowlist**, not a blocklist. Code, docs, config, and markup are editable.
Anything not on the list is not.
- **512 KB cap** on both read and write.
- **Edit mode never truncates.** The plain preview does truncate long files, and saving a
truncated buffer would silently delete the rest, so the editor loads the whole file or
refuses.
- **Optimistic concurrency.** The save carries a hash of what you started from. If the file
changed underneath you (likely, when an agent is working in the same repo), the save is
rejected rather than clobbering their work.
- **No file creation.** Writes go to a temporary file and are renamed over the original, and
the open never creates. Editing in place is structural, not a rule.
- **Line endings are preserved** server-side, so editing two lines of a CRLF file does not
produce a whole-file diff.
- **`.git/` is denied outright.** Hooks are executable code, and a corrupted index looks
unrecoverable to someone who wanted to fix a typo.
- **Non-UTF-8 content is refused**, verified by a round-trip comparison.
## Attachments
Attachments are live references to files **outside** the session's workspace: a spec on your
desktop, a PDF in Downloads, a design document elsewhere on the machine.
Register one from the CLI:
```bash
codeman attach /path/to/spec.pdf
```
An attachment card appears in the session, and the file can be previewed inline. The
attachment gets a stable id, and browser requests use that id rather than carrying absolute
paths around.
Agents can register attachments too, by emitting a `codeman://attach?...` link in their
output. That path is **prompt-injectable by nature**, so it is force-confined to the
session's workspace: a hostile prompt cannot use it to pull arbitrary host files into the
event stream. The gate is an extension allowlist rather than a blocklist.
Document conversion for previews is globally rate limited. Without that, ten large documents
detected at once would fork ten multi-minute converter processes.
## Clicking a path
File paths in a session are links. That works in two places:
- **In the terminal**, on any absolute path an agent prints.
- **In the response viewer**, where paths are usually written as prose or in backticks. They
render as underlined monospace links.
Clicking one opens it in the preview: images and PDFs render, video and audio play with a
working scrub bar, documents convert, text and Markdown show inline. Log-shaped files open in
the tail viewer instead, which follows a file that is still being written.
Paths **outside** the session's workspace work too, which matters because that is where most
of an agent's output lands: a screenshot in `/tmp`, a capture in its own scratchpad, a file in
another checkout. Those are served through the attachment routes rather than the workspace
ones, so the same rules apply as to any other attachment: secret trees are blocked, the
extension allowlist decides what can be opened, and symlinks are resolved before either check.
Outside the workspace the allowlist is images, video, audio, PDF, Office documents, and text
files, where "text" is the same list the viewer will let you edit: code, config, logs, csv,
markdown. The reasoning is that a session can already `cat` any of those, so the file suffix
was never what kept anything secret; the path guard is. Types outside the list (`.svg`,
`.bmp`) say so rather than failing silently, and `.html` previews as source rather than being
rendered, so nothing served this way can execute in the page.
Text previews are capped at the first 500 lines, fetched as a partial read, so clicking a
one-gigabyte log does not try to paint one.
Log-shaped files inside the workspace still open in the tail viewer, which follows a file as
it is written. Outside the workspace they open in the preview instead: the tail viewer runs
`tail -f`, and that is deliberately restricted to the workspace, `/var/log` and `~/logs`.
Nothing is registered until you click. Opening a file this way does not add an attachment card.
## The path picker
For choosing a path rather than typing one. It appears in two places:
- **Browse** in **Add Case → Link Existing**.
- The **📁 Path** key on the mobile keyboard bar.
It browses one directory at a time and can show hidden entries on request. The picker
inserts the path into your prompt **without** pressing Enter, so nothing is submitted by
accident. Its sibling **⌫ All** key clears the unsent prompt, and never sends the agent's
`/clear` command.
This is a separate file-serving surface from the viewer, with its own rules: it allowlists
your home directory, the cases directory, and anything in `CODEMAN_FILE_PICKER_ROOTS`, and
blocks sensitive trees. In multi-user mode a non-admin gets only their own user space as a
root, because per-user spaces live inside the home directory and a home-directory root would
expose everyone.
## Images into a session
Paste from the clipboard or drag and drop straight onto the terminal. The image is written
where the agent can read it and the reference is inserted into your prompt. On a phone, the
image key in the keyboard bar opens the camera or photo library.
HEIC images from an iPhone are converted to JPEG on the way in.
## Generated artifacts
When an agent produces a file the UI can show (a chart, a diagram, a document), it can
surface as an artifact attachment rather than a path you have to go and find.
## Gotchas
- **The viewer follows the active session's workspace.** Switching tabs changes what you are
browsing.
- **A save can be rejected, and that is the feature.** It means the agent edited the file
while you were typing. Re-open, re-apply, save again.
- **Attachments live outside the workspace on purpose.** For files inside it, just use the
viewer.
- **`.env` files are readable in the viewer if the extension policy allows the preview, but
never downloadable.** Do not treat the viewer as a secrets boundary; treat the machine as
the boundary.
## Read next
- [The Dashboard](The-Dashboard) - where the panels live.
- [Input And Voice](Input-And-Voice) - other ways to get content into a session.
- [Security](Security) - how the file surfaces are confined.
- [`docs/file-viewer-edit-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/file-viewer-edit-plan.md) - the edit-mode design.
+9
View File
@@ -0,0 +1,9 @@
Documents Codeman **{{VERSION}}**. Something wrong or missing on this page? These pages are
generated from [`docs/wiki/`](https://github.com/Ark0N/Codeman/tree/master/docs/wiki) in
the main repository, so browser edits here are overwritten on the next sync. Send a pull
request against that directory instead, or open a
[Discussion](https://github.com/Ark0N/Codeman/discussions).
<!-- {{VERSION}} is replaced with the current major.minor series by
.github/workflows/wiki-sync.yml at publish time. Do not hardcode a
version here: it went stale every release when it was hand-written. -->
+53
View File
@@ -0,0 +1,53 @@
### [Codeman Wiki](Home)
[README](https://github.com/Ark0N/Codeman)
**Getting started**
- [Installation](Installation)
- [Quick Start](Quick-Start)
- [Core Concepts](Core-Concepts)
**Using it**
- [The Dashboard](The-Dashboard)
- [Agent CLIs](Agent-CLIs)
- [Working With Files](Working-With-Files)
- [Input And Voice](Input-And-Voice)
- [Mobile Guide](Mobile-Guide)
- [Keyboard Shortcuts](Keyboard-Shortcuts)
- [Settings Reference](Settings-Reference)
**Keeping agents running**
- [Unattended Runs](Keeping-Agents-Running)
- [Notifications & Approvals](Notifications-And-Approvals)
- [Cron Jobs](Cron-Jobs)
- [Autonomous Loops](Autonomous-Loops)
- [Watching Agents Work](Watching-Agents-Work)
**Where it runs**
- [Docker Cases](Docker-Cases)
- [Remote SSH Sessions](Remote-SSH-Sessions)
- [Web Tabs](Web-Tabs)
- [Multi-User Mode](Multi-User-Mode)
**Access & security**
- [Remote Access](Remote-Access)
- [Security](Security)
**Automation**
- [Driving It From An Agent](Driving-Codeman-From-An-Agent)
- [HTTP API](HTTP-API)
- [Hooks & Integrations](Hooks-And-Integrations)
**Operating it**
- [Running As A Service](Running-As-A-Service)
- [Troubleshooting](Troubleshooting)
- [FAQ](FAQ)
- [Contributing](Contributing)
- [Versioning](Versioning)
+121
View File
@@ -0,0 +1,121 @@
# Warm worker pool: sub-second claude worker spawns
Design sketch. Status: **proposed**, not started. Opt-in (`workerPoolSize`, default 0 = off); a user who touches nothing sees no change at all.
---
## 1. Problem and numbers
Measured against prod 1.18.3 on 2026-08-15, AFTER the SKILL.md fast-path hardening
(no recon turns), on the identical "spawn two codeman workers" prompt:
- **Cold orchestrator** (fresh session, skill loaded from disk): **20.2 s** prompt to
final report. Breakdown: 3.9 s Skill-load turn, 6.4 s generating the one fused Bash
call, **4.4 s spawn call**, 5.5 s summary. Tabs appeared at 10.5 s.
- **Warm orchestrator** (skill already in context, no Skill turn): **12.8 s**, spawn
call 6.0 s.
- Inside the spawn call, session + tmux + case creation is cheap: the workers (and
their tabs) appeared 0.2-1.7 s in, both siblings within ~350 ms of each other. The
remaining **~4-5 s is claude CLI boot plus the composer-readiness wait**, paid again
on every cold spawn. That slice is the pool's entire target.
The honest framing after the hardening: model turns dominate the skill flow (~16 of
20 cold seconds) and no server feature can shrink those. The pool attacks the
tool-side floor, and it has two distinct beneficiaries:
- **Skill/API orchestration**: the spawn call drops from ~4.4-6 s to ~1 s. Cold runs
land ~16-17 s, warm ~8 s. Tab appearance barely moves for this consumer (it is
model-turn-bound at ~10 s cold / ~4 s warm).
- **The UI Run button and direct quick-start callers**: a click today waits the full
boot + readiness before the worker can take a prompt; a pooled claim makes the tab
appear and the worker READY sub-second. This is the most visible win, and it
involves no skill at all.
Target: hand out an already-ready worker in **under 1 s**.
## 2. Shape
A new `src/worker-pool.ts` singleton service, following the `CronService` pattern: it **reuses the existing session layer** (`SessionManager` create + the normal spawn path) and never rebuilds tmux logic.
A pool member is a real claude `Session`, pre-spawned in a reserved scratch case (`~/codeman-cases/.pool-<n>`, created with the standard scaffold + hooks), already past readiness: composer drawn, hooks installed, preamble file seeded. It sits idle at the composer costing no tokens.
The claim happens **transparently inside `POST /api/quick-start`**: when a request is pool-eligible (§3) and a healthy member is available, quick-start returns that member instead of cold-spawning. The agent skill, the UI Run button, and every existing caller change **nothing**. Ineligible or pool-empty requests cold-spawn exactly as today, so the pool is only ever a fast path, never a behavior change.
## 3. Eligibility gate
Claim only when ALL of these hold; otherwise fall through to a cold spawn:
- `mode === 'claude'` (external CLIs have different readiness semantics and inject secrets via `tmux setenv` at spawn; out of scope).
- No `envOverrides`, no `CLAUDE_CONFIG_DIR`, and `modelOverride`/`effort` unset or equal to what the pool member was spawned with. Env vars flow at spawn time and cannot be applied to a running CLI.
- The requested case is **fresh** (does not exist yet). A linked case, an existing directory, a remote-SSH case, or a Docker case means the caller wants a specific workspace; pool members cannot provide one.
- Single-user mode, or the requester owns the pool (v1 ships single-user only; §11).
## 4. What a claim does (~300 ms)
1. Pop a ready member (in-memory check-and-remove; Node's single thread makes this atomic, so two concurrent quick-starts cannot claim the same member).
2. Health-probe it: `isPaneDead` (the existing ~750 ms-cached mux probe) plus one `capturePaneText` asserting a clean composer. A dead, limit-paused, or dirty member is recycled, and the claim tries the next member or falls through to cold spawn.
3. Rename the session to the normal `w<n>-<case>` name, set `parentSessionId` via the existing `resolveParentSessionId()`, clear the pool flag, persist state.
4. Emit `session_created` **now** (it was suppressed at warm-spawn time, §5). The tab appears here, sub-second after the request.
5. Return the **pool case** as `casePath`/`workingDir` and do NOT create a directory under the requested name: an empty dir the worker's CLI does not run in is a trap (files written there are invisible to the worker at cwd), and the agent skill greps the RETURNED `casePath` for Codeman hooks before trusting the worker, so the response must point at the directory that really carries them.
6. Kick a background refill (§6).
**The identity wrinkle, stated honestly:** the session id, `CODEMAN_SESSION_ID` inside the pane, the seeded preamble file, and the CLI's cwd are all fixed at warm-spawn and survive the claim unchanged. So a claimed worker's `workingDir` is the pool dir, not `~/codeman-cases/<requested-name>`; the requested name is a **label**. The API must report the truthful `workingDir`. Transcript projHash, response viewer, subagent windows, and Read My Mind all key off the real path and keep working precisely because we do not lie about it. This is acceptable for the dominant use (ephemeral skill workers that are deleted after answering) and is documented in the skill; a caller that needs the real case as cwd is by definition not pool-eligible.
**Verified skill compatibility (zero preamble changes).** Checked against the shipped 1.18.3 preamble: `spawn_worker`'s readiness probe (`_composer_up`) is a `wait-output` call with `from=buffer`, which scans output that already scrolled past before blocking, so a pooled member's long-since-drawn composer matches instantly instead of stranding a fresh-stream wait. The trust-dialog fallback never fires (members passed the dialog at warm time), and the hooks grep passes because the pool case carries the standard scaffold. Pooled and cold spawns are indistinguishable to the skill except in speed and the additive `pooled: true`.
## 5. Hiding pre-claim members
Pool members must be invisible until claimed or they read as ghost tabs. `Session.isPoolWorker` gates, at minimum:
- `GET /api/sessions` and `GET /api/sessions/unified` (and therefore the Cmd+K palette and the session-history-index snapshot that feeds `/api/search`).
- `session_created` SSE at warm-spawn (deferred to claim time). All other per-session SSE for a hidden member is suppressed at the broadcast call sites it would reach.
- Push notifications and the Approvals Inbox (a warm member showing a trust dialog must recycle, not notify).
- The phone overview / home rail (both render from the session list, so the list filter covers them).
- The lifecycle log records `pool_warm` / `pool_claim` events rather than user-visible session history.
`maxSessions` (50) **counts** pool members, and the pool refuses to warm within `poolSize + 2` of the cap so it can never starve real session creation.
## 6. Refill, TTL, drain
- **Refill** after each claim, debounced, at most one warm spawn in flight (a claim burst falls back to cold spawns rather than forking N CLIs at once; same reasoning as the document-conversion limiter).
- **TTL ~30 min**: recycle members older than that so they cannot drift from settings, hooks config, or a self-updated CLI on disk.
- **Drain and respawn** on: `claudeModel` change, hooks-config regeneration, self-update, and `workerPoolSize` changes. On server shutdown, kill pool sessions (they are stateless and ours). On boot, kill any leftover `.pool-*` tmux sessions found via `mux-sessions.json` rather than adopting them; adoption buys nothing for stateless members.
## 7. Failure modes
| Failure | Handling |
| --- | --- |
| Member died idle (PTY exit, crash) | Health probe at claim catches it; recycle + try next; PTY-exit breaker applies unchanged |
| Member hit a usage limit while idle | `isLimitPaused` members are never handed out; recycle |
| Composer dirty (stray keystrokes, dialog) | `capturePaneText` probe refuses it; recycle |
| Claim race | Impossible by construction (synchronous in-memory pop) |
| Warm spawn itself fails | Log, back off, retry on next refill tick; pool empty just means cold spawns |
## 8. Cost
Each warm member is one tmux session + one idle claude process (order 150-300 MB RSS; **measure before defaulting the size above 0**, including whether an idle CLI makes any background requests via its statusline refresh). Zero token cost while idle. Suggested starting size for users who opt in: 2.
## 9. Settings and API surface
- `workerPoolSize` (int, 0-4, default 0): **synced** setting in `SettingsUpdateSchema`. The watcher that resizes the pool on `PUT /api/settings` must resolve from `merged`, never the raw body (the partial-PUT gotcha in CLAUDE.md).
- One internal status endpoint, `GET /api/worker-pool` (size, members' ages, claims served, fall-through count), for debugging. No new SSE events: the claim emits the existing `session_created`.
- No new public API semantics: `/api/quick-start`'s contract is unchanged apart from a `pooled: true` field in the response data, which is additive.
## 10. Considered and rejected
- **Renaming the pool case dir to the requested name at claim.** Linux keeps the process cwd working across the rename (inode-based), but claude computed its transcript projHash from the old path string at boot, so transcripts, subagent windows, and the response viewer go blind, the exact failure mode the `CLAUDE_CONFIG_DIR` docs warn about. Truthful label semantics (§4) beat a clever rename.
- **A new explicit claim endpoint.** Transparency inside quick-start means the skill, the UI, and every existing script get the speedup with zero changes; a new endpoint means new docs, new drift, and callers that must know the pool exists.
- **Pooling external CLI modes.** Readiness there is output stabilization, secrets ride `tmux setenv` at spawn, and codex/pi composer semantics differ per CLI. Claude-only until someone measures a need.
- **Returning quick-start at creation instead of readiness (no pool).** Would move tabs earlier on cold spawns too, but `sendwait` immediately after would then race the composer; readiness is what makes immediate tasking safe, and the pool makes the whole question moot for eligible spawns.
## 11. Phasing
1. **v1**: single-user, claude-only, fixed-size pool, transparent claim, status endpoint. Everything above.
2. **v2**: per-owner pools for multi-user mode (pool members must carry an owner because ownership scoping is structural); possibly model-matched pools (one warm set per configured `claudeModel`).
3. **Explicitly out**: warming linked/repo cases (spawning where the work is has no hooks and is the skill's documented costliest mistake; a warm pool must not make it faster to reach).
## 12. Testing
- Unit: pool manager logic pure and mock-driven (eligibility gate, TTL, refill debounce, drain triggers), `MockSession` from `test/mocks/`.
- Route: `app.inject` on quick-start asserting claim vs cold-spawn per eligibility row in §3, plus the double-claim race (two concurrent injects, one pool member: exactly one `pooled: true`).
- Live: re-run the pinned baselines against a warmed beta instance. Before (2026-08-15, prod 1.18.3, post-hardening): cold orchestrator **20.2 s** / warm **12.8 s** end to end, spawn call 4.4-6.0 s. Acceptance: spawn call under 1 s, cold ~16-17 s, warm ~8-9 s, and a UI Run click to a READY worker in under 1 s.
+99 -11
View File
@@ -116,6 +116,23 @@ GEMINI_SEARCH_PATHS=(
"$HOME/bin/gemini"
)
# Pi CLI search paths (from src/utils/pi-cli-resolver.ts)
PI_SEARCH_PATHS=(
"$HOME/.local/bin/pi"
"/usr/local/bin/pi"
"$HOME/.bun/bin/pi"
"$HOME/.npm-global/bin/pi"
"$HOME/bin/pi"
)
# Antigravity CLI search paths (from src/utils/antigravity-cli-resolver.ts)
ANTIGRAVITY_SEARCH_PATHS=(
"$HOME/.local/bin/agy"
"$HOME/.antigravity/bin/agy"
"/usr/local/bin/agy"
"$HOME/bin/agy"
)
# ============================================================================
# Color Output
# ============================================================================
@@ -493,6 +510,65 @@ get_gemini_path() {
done
}
check_antigravity() {
if command -v agy &>/dev/null; then
return 0
fi
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
get_antigravity_path() {
if command -v agy &>/dev/null; then
command -v agy
return
fi
for path in "${ANTIGRAVITY_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
# `pi` is a short, generic name (Raspberry Pi tooling, personal scripts), so the
# server-side resolver additionally probes `pi --version`. Detection here only feeds
# the "you have no AI CLI" hint, so a plain executable test is enough.
check_pi() {
if command -v pi &>/dev/null; then
return 0
fi
for path in "${PI_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
return 0
fi
done
return 1
}
get_pi_path() {
if command -v pi &>/dev/null; then
command -v pi
return
fi
for path in "${PI_SEARCH_PATHS[@]}"; do
if [[ -x "$path" ]]; then
echo "$path"
return
fi
done
}
check_cloudflared() {
# Check ~/.local/bin first (matches tunnel-manager.ts resolution order)
if [[ -x "$HOME/.local/bin/cloudflared" ]]; then
@@ -1993,11 +2069,13 @@ main() {
fi
fi
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini)
# AI CLI (Codeman drives one of: Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi)
local has_claude=false
local has_opencode=false
local has_codex=false
local has_gemini=false
local has_antigravity=false
local has_pi=false
info "Checking AI CLI tools..."
if check_claude; then
@@ -2016,17 +2094,25 @@ main() {
has_gemini=true
success "Gemini CLI found at $(get_gemini_path)"
fi
if check_antigravity; then
has_antigravity=true
success "Antigravity CLI found at $(get_antigravity_path)"
fi
if check_pi; then
has_pi=true
success "Pi CLI found at $(get_pi_path)"
fi
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" ]]; then
if [[ "$has_claude" == "false" && "$has_opencode" == "false" && "$has_codex" == "false" && "$has_gemini" == "false" && "$has_antigravity" == "false" && "$has_pi" == "false" ]]; then
echo ""
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, or Gemini."
warn "No AI CLI found. Codeman needs at least one: Claude Code, OpenCode, Codex, Antigravity, Gemini, or Pi."
headless_guard "install an AI CLI (curl | bash from its vendor)"
echo ""
echo -e " ${BOLD}Which AI CLI would you like to install?${NC}"
echo -e " ${CYAN}1)${NC} Claude Code (Anthropic)"
echo -e " ${CYAN}2)${NC} OpenCode (open-source)"
echo -e " ${CYAN}3)${NC} Both"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex or Gemini)"
echo -e " ${CYAN}4)${NC} Skip (I'll install one myself, e.g. Codex, Antigravity or Pi)"
echo ""
local cli_choice=""
@@ -2071,8 +2157,9 @@ main() {
if [[ "$cli_choice" == "4" ]]; then
warn "Skipping AI CLI install. Codeman will run, but sessions need a CLI to drive."
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: npm install -g @google/gemini-cli (Gemini)"
info "Install one later, e.g.: npm install -g @openai/codex (Codex)"
info " or: curl -fsSL https://antigravity.google/cli/install.sh | bash (Antigravity)"
info " or: npm install -g --ignore-scripts @earendil-works/pi-coding-agent (Pi)"
elif [[ "$has_claude" == "false" ]] && [[ "$has_opencode" == "false" ]]; then
die "The selected AI CLI failed to install. Install one manually and re-run the installer."
fi
@@ -2372,12 +2459,13 @@ main() {
echo -e " https://github.com/Ark0N/Codeman"
echo ""
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini; then
if ! check_claude && ! check_opencode && ! check_codex && ! check_gemini && ! check_antigravity && ! check_pi; then
echo -e " ${YELLOW}${BOLD}Reminder:${NC} Install at least one AI CLI to start using Codeman:"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}npm install -g @google/gemini-cli${NC} # Gemini"
echo -e " ${CYAN}curl -fsSL https://claude.ai/install.sh | bash${NC} # Claude Code"
echo -e " ${CYAN}curl -fsSL https://opencode.ai/install | bash${NC} # OpenCode"
echo -e " ${CYAN}npm install -g @openai/codex${NC} # Codex"
echo -e " ${CYAN}curl -fsSL https://antigravity.google/cli/install.sh | bash${NC} # Antigravity"
echo -e " ${CYAN}npm install -g --ignore-scripts @earendil-works/pi-coding-agent${NC} # Pi"
echo ""
fi
+58 -29
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.10.0",
"version": "1.20.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.10.0",
"version": "1.20.1",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -17,7 +17,7 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/multipart": "^10.0.0",
"@fastify/static": "^9.1.3",
"@fastify/static": "^10.1.3",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-serialize": "^0.14.0",
@@ -60,6 +60,7 @@
"pixelmatch": "^6.0.0",
"playwright": "^1.58.0",
"pngjs": "^7.0.0",
"postcss": "^8.5.15",
"prettier": "^3.4.0",
"puppeteer": "^24.36.0",
"remotion": "4.0.473",
@@ -1453,9 +1454,9 @@
}
},
"node_modules/@fastify/static": {
"version": "9.1.3",
"resolved": "https://registry.npmjs.org/@fastify/static/-/static-9.1.3.tgz",
"integrity": "sha512-aXrYtsiryLhRxRNaxNqsn7FUISeb7rB9q4eHUPIot5aeQBLNahnz1m6thzm7JWC1poSGXS9XrX8DvuMivp2hkQ==",
"version": "10.1.3",
"resolved": "https://registry.npmjs.org/@fastify/static/-/static-10.1.3.tgz",
"integrity": "sha512-W6jqajYS974XjPjB5hQWoxPM8NKM4+p8YmQT6G5IbCa4uhdWSVadZUv75siy1wEA/3ty8RYdpBydfWeu9AqAqQ==",
"funding": [
{
"type": "github",
@@ -1469,13 +1470,30 @@
"license": "MIT",
"dependencies": {
"@fastify/accept-negotiator": "^2.0.0",
"@fastify/error": "^4.0.0",
"@fastify/send": "^4.0.0",
"content-disposition": "^1.0.1",
"fastify-plugin": "^5.0.0",
"content-disposition": "^2.0.1",
"fastify-plugin": "^6.0.0",
"fastq": "^1.17.1",
"glob": "^13.0.0"
}
},
"node_modules/@fastify/static/node_modules/fastify-plugin": {
"version": "6.0.0",
"resolved": "https://registry.npmjs.org/fastify-plugin/-/fastify-plugin-6.0.0.tgz",
"integrity": "sha512-fZOty7z3O7vOliF6d8bHE3wiEh1KcNnKEQensSgTk9C1DvN6nRLS++XVd86v33Hw/8u9Un8A1zDrQ8ujcQDHEg==",
"funding": [
{
"type": "github",
"url": "https://github.com/sponsors/fastify"
},
{
"type": "opencollective",
"url": "https://opencollective.com/fastify"
}
],
"license": "MIT"
},
"node_modules/@fastify/websocket": {
"version": "11.2.0",
"resolved": "https://registry.npmjs.org/@fastify/websocket/-/websocket-11.2.0.tgz",
@@ -4135,16 +4153,16 @@
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/brace-expansion": {
"version": "5.0.6",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.6.tgz",
"integrity": "sha512-kLpxurY4Z4r9sgMsyG0Z9uzsBlgiU/EFKhj/h91/8yHu0edo7XuixOIH3VcJ8kkxs6/jPzoI6U9Vj3WqbMQ94g==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
"node": "20 || >=22"
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/minimatch": {
@@ -4547,6 +4565,16 @@
"integrity": "sha512-b3fMOsyLVuCeNJWxolACEUED0vm7qC0cy4wRvf3oURSzDTYVQiGPhTnhWZwIHdvC48Y+oLhvYXnY4XDXPoJo6A==",
"license": "MIT"
},
"node_modules/@xterm/headless": {
"version": "6.0.0",
"resolved": "https://registry.npmjs.org/@xterm/headless/-/headless-6.0.0.tgz",
"integrity": "sha512-5Yj1QINYCyzrZtf8OFIHi47iQtI+0qYFPHmouEfG8dHNxbZ9Tb9YGSuLcsEwj9Z+OL75GJqPyJbyoFer80a2Hw==",
"dev": true,
"license": "MIT",
"workspaces": [
"addons/*"
]
},
"node_modules/@xterm/xterm": {
"version": "6.0.0",
"resolved": "https://registry.npmjs.org/@xterm/xterm/-/xterm-6.0.0.tgz",
@@ -5067,9 +5095,9 @@
"license": "MIT"
},
"node_modules/brace-expansion": {
"version": "1.1.15",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.15.tgz",
"integrity": "sha512-EwOCDEex4quD37XhqM3omwtMoJjr//isUZz1JopUNWms+4Z2ViyM/k1YIRePpoVNnQhENnxtFjLaxNHrT7xIUg==",
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -5420,9 +5448,9 @@
}
},
"node_modules/content-disposition": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/content-disposition/-/content-disposition-1.1.0.tgz",
"integrity": "sha512-5jRCH9Z/+DRP7rkvY83B+yGIGX96OYdJmzngqnw2SBSxqCFPd0w2km3s5iawpGX8krnwSGmF0FW5Nhr0Hfai3g==",
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/content-disposition/-/content-disposition-2.0.1.tgz",
"integrity": "sha512-e+H0ZXHSWYrENhQzw1LPuP4oF5MzVKmDU6d3hxlvaPEYLLg62MxtQNPRx4SYSuYJSBUgnQIG4HIN2tEtNv7Dog==",
"license": "MIT",
"engines": {
"node": ">=18"
@@ -6541,9 +6569,9 @@
}
},
"node_modules/fast-uri": {
"version": "3.1.2",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.2.tgz",
"integrity": "sha512-rVjf7ArG3LTk+FS6Yw81V1DLuZl1bRbNrev6Tmd/9RaroeeRRJhAt7jg/6YFxbvAQXUCavSoZhPPj6oOx+5KjQ==",
"version": "3.1.5",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.5.tgz",
"integrity": "sha512-gHwA1O9LDIcKunMKhObS/HimwtehO1nPUECKAu5TpKgaO19fcWEl4bliWe1jWxVFvIXztJjjQ4L8XQ1EU9f7Jw==",
"funding": [
{
"type": "github",
@@ -6650,9 +6678,9 @@
}
},
"node_modules/find-my-way": {
"version": "9.6.0",
"resolved": "https://registry.npmjs.org/find-my-way/-/find-my-way-9.6.0.tgz",
"integrity": "sha512-Zf4Xve4RymLl7NgaavNebZ01joJ8MfVerOG43wy7SHLO+r+K0C6d/SE0BiR7AV5V1VOCFlOP7ecdo+I4qmiHrQ==",
"version": "9.8.0",
"resolved": "https://registry.npmjs.org/find-my-way/-/find-my-way-9.8.0.tgz",
"integrity": "sha512-JtyUgATO7qxRp2zKhrmWof74Mqxc1ikbwpwMY97p8ipuTj2QtreA4gK2JNAF6SOqqHnYYkwMUvsgQVi2AJxIyw==",
"license": "MIT",
"dependencies": {
"fast-deep-equal": "^3.1.3",
@@ -6894,15 +6922,15 @@
}
},
"node_modules/glob/node_modules/brace-expansion": {
"version": "5.0.6",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.6.tgz",
"integrity": "sha512-kLpxurY4Z4r9sgMsyG0Z9uzsBlgiU/EFKhj/h91/8yHu0edo7XuixOIH3VcJ8kkxs6/jPzoI6U9Vj3WqbMQ94g==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
"node": "20 || >=22"
}
},
"node_modules/glob/node_modules/minimatch": {
@@ -12333,9 +12361,10 @@
}
},
"packages/xterm-zerolag-input": {
"version": "0.1.8",
"version": "0.3.1",
"license": "MIT",
"devDependencies": {
"@xterm/headless": "^6.0.0",
"jsdom": "^24.1.3",
"tsup": "^8.5.1",
"typescript": "^5.5.0",
+14 -5
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.10.0",
"version": "1.20.1",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -17,10 +17,15 @@
"dev": "tsx src/index.ts web",
"web": "node dist/index.js web",
"clean": "rm -rf dist",
"test": "vitest run --config config/vitest.config.ts",
"test:watch": "vitest --config config/vitest.config.ts",
"test:coverage": "vitest run --config config/vitest.config.ts --coverage",
"test": "vitest run --config config/vitest.ci.config.ts",
"test:watch": "vitest --config config/vitest.ci.config.ts",
"test:coverage": "vitest run --config config/vitest.ci.config.ts --coverage",
"test:ci": "vitest run --config config/vitest.ci.config.ts",
"test:browser": "vitest run --config config/vitest.browser.config.ts",
"test:perf": "vitest run --config config/vitest.perf.config.ts",
"test:all": "vitest run --config config/vitest.config.ts",
"pretest:mobile": "node scripts/prepare-test-vendor.mjs",
"test:mobile": "vitest run --config test/mobile/vitest.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
"fix:node-pty": "node scripts/fix-node-pty.mjs",
"typecheck": "tsc --noEmit",
@@ -55,6 +60,8 @@
"anthropic",
"opencode",
"codex",
"antigravity",
"pi",
"gemini-cli",
"ai-agents",
"agent",
@@ -78,7 +85,7 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/multipart": "^10.0.0",
"@fastify/static": "^9.1.3",
"@fastify/static": "^10.1.3",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-serialize": "^0.14.0",
@@ -117,6 +124,7 @@
"pixelmatch": "^6.0.0",
"playwright": "^1.58.0",
"pngjs": "^7.0.0",
"postcss": "^8.5.15",
"prettier": "^3.4.0",
"puppeteer": "^24.36.0",
"remotion": "4.0.473",
@@ -158,6 +166,7 @@
"dist",
"scripts/postinstall.js",
"scripts/fix-node-pty.mjs",
"skills",
"LICENSE",
"README.md"
]
+40
View File
@@ -1,5 +1,45 @@
# xterm-zerolag-input
## 0.3.1
### Patch Changes
- Mobile catches up: links open from a tap, terminal text can be selected and copied, long prompts stay visible while you type. Plus Files panel search, a bundled Nerd Font symbols fallback, and a per-device terminal font setting.
- **Terminal and chat links work on phones** (#321): tapping a URL or file path in terminal output now opens it (new tab, file preview, or log viewer), resolved through the same provider desktop hover uses, so tap and click can never disagree about what is a link. Dialog rows and the composer keep their existing meaning. Response-viewer links open in a new tab with `rel="noopener noreferrer"` instead of navigating the dashboard away. Wrapped links open whole: the logical-line reconstruction now stitches hard wraps through the indent their continuation carries, which also fixes desktop hover-click truncating wrapped URLs.
- **Terminal text can be copied on touch devices** (#321): long-press selects the token under the finger, drag or tap the other end to extend, and a small bar offers Copy, Line (the whole logical line, wraps included) and dismiss. Copy works on plain-HTTP installs too. Three guards keep the keyboard down and the selection alive through the browser's own long-press handling.
- **A long prompt stays visible on phones** (#321): the local-echo overlay grows upward once it would run past the last visible row (a prompt taller than the screen keeps its tail, where the cursor is), and the keyboard-driven padding shrink can no longer reclaim the space the fixed toolbar and accessory bar stand in.
- **Files panel search** (#324): `GET /api/sessions/:id/files?q=...` answers a flat match list (name or path substring, `*`/`?` globs), recursing past non-matching directories with its own match cap on top of the existing bounds; without `q` the response is byte-identical to before. Glob queries are matched without regex so a pathological pattern cannot stall the server.
- **Nerd Font prompt glyphs out of the box, custom terminal font** (#320): a bundled icons-only Symbols Nerd Font Mono fallback renders powerlevel10k/starship/oh-my-posh glyphs on every device with no font install, and App Settings gains a per-device terminal font family that is prepended to the built-in stack.
### Thanks
Three contributor PRs in one release: thanks to @rounakdatta (#321), @aakhter (#324) and @comzine (#320).
## 0.3.0
### Minor Changes
- 55bff4a: Zero-lag predictive echo for Codex sessions (mosh-style write-through prediction).
Codex's per-keystroke composer forced 1.12.2 to disable the local-echo overlay (issues #218/#219/#220/#222), leaving Codex typing at full round-trip latency on remote links. This release adds a second echo mode instead of re-enabling the first: every keystroke still goes to the PTY exactly as before (byte-identical wire behavior, pinned by vm-level and end-to-end trace-equality tests), while the new `PredictiveEchoAddon` in `xterm-zerolag-input` 0.2.0 paints the predicted glyph at the predicted cell. When the real echo lands, the prediction is confirmed and its span removed (an invisible swap); mispredictions self-heal via a two-pass mismatch cascade and a TTL.
- Reconciliation reads the parsed terminal buffer, never the raw stream: full-line redraws, ECH gap painting and tmux's in-place deltas all converge to the same cells. Confirmation requires the cell match PLUS a cursor advance, so placeholder glyphs and identical repaints never false-confirm; blank cells are neutral (codex clears its placeholder on the first echo).
- Predictions paint only while the cursor sits on the measured Codex composer row (`/^› /`, codex-cli 0.147): trust/approval modals and wrapped continuation rows get no ghosts, deliberately falling back to real echo.
- Ships as a SEPARATE `vendor/xterm-predictive-echo.js` bundle: the existing zerolag bundle is byte-identical (sha256-verified), and a missing or broken bundle degrades Codex to exact 1.12.2 behavior. The per-device `localEchoEnabled` toggle is the kill switch.
- Claude/Gemini/OpenCode/Antigravity keep buffer mode untouched; shell stays off.
- A post-build adversarial review added the anchor-hold rule: after an unpredicted wire edit (backspace into echoed text, cleared input, IME text commits) new predictions hold until the next parsed write, so a stale displayed cursor can never mis-anchor a run.
- Tests: 55 new package tests including replay suites driven by fixtures recorded from a real codex TUI through the production tmux+strip pipeline (`scripts/dev/record-codex-frames.mjs`) and a 500-iteration seeded fuzz; new vm policy/wire-neutrality suites; a 10-scenario Playwright E2E against real codex covering the #218/#219/#220/#222 retests, byte-identity, and a simulated 300ms-RTT run. The package test suite now runs in CI.
## 0.2.0
### Minor Changes
- **New addon: `PredictiveEchoAddon`, mosh-style write-through prediction.** The second echo mode for per-keystroke TUIs (OpenAI Codex's composer, live pickers) that buffer-until-Enter starves. Every keystroke is sent by the consumer immediately and unchanged; the addon paints the predicted glyph at the predicted cell and reconciles against the PARSED terminal buffer: confirmation requires the cell match plus a cursor advance past the record, foreign non-blank content on two consecutive passes cascades a drop, blank cells are neutral, a TTL bounds everything, and scroll/resize/sustained cursor moves clear the run. Visual-only by construction; it cannot gate, delay or rewrite input.
- Anchor-hold rule: after an unpredicted wire edit (backspace into echoed text, cleared input, an IME text commit) new predictions hold until the next parsed write, so a stale displayed cursor can never mis-anchor a run (worst case: exactly one unpredicted keystroke).
- New exports: `PredictiveEchoAddon`, `PredictiveEchoOptions`, `PredictionState`, plus the long-intended `charCellWidth` / `stringCellWidth` helpers.
- `XtermTerminal` type gains OPTIONAL members (`buffer.active.cursorX/cursorY`, `getLine().getCell?`, `onWriteParsed?`, `onResize?`). Additive only: existing consumers and mocks are unaffected.
- IIFE build exposes `window.PredictiveEchoAddon` and a self-activating `window.PredictiveEchoOverlay`, alongside the unchanged `ZerolagInputAddon` / `LocalEchoOverlay` globals.
- Tests: 52 new (30 addon-law specs, renderer geometry, 6 replay suites driven by fixtures recorded from real codex 0.147 through tmux + the production strip, and a 500-iteration seeded fuzz with per-op invariants). `@xterm/headless` as a devDependency; runtime dependencies remain zero.
## 0.1.8
### Patch Changes
+115 -2
View File
@@ -9,14 +9,14 @@
<a href="https://opensource.org/licenses/MIT"><img src="https://img.shields.io/badge/License-MIT-1e3a5f?style=flat-square" alt="MIT"></a>
<img src="https://img.shields.io/badge/Dependencies-0-22c55e?style=flat-square" alt="Zero dependencies">
<img src="https://img.shields.io/badge/Size-6.1%20kB%20gzip-22c55e?style=flat-square" alt="6.1 kB gzipped">
<img src="https://img.shields.io/badge/Tests-175-22c55e?style=flat-square" alt="175 tests">
<img src="https://img.shields.io/badge/Tests-227-22c55e?style=flat-square" alt="175 tests">
<img src="https://img.shields.io/badge/xterm.js-v5%20%7C%20v7+-3b82f6?style=flat-square" alt="xterm.js v5 and v7+">
</p>
</p>
> ### Made for [**Codeman**](https://getcodeman.com)
>
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Gemini sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
> This overlay is the local echo engine of [**Codeman**](https://github.com/Ark0N/Codeman), mission control for AI coding agents: run and monitor a dozen Claude Code, Codex, OpenCode and Antigravity sessions at once, watch their subagents work in live floating windows, let them run autonomously overnight, and drive all of it from your phone.
>
> That last part is why this library exists. The demo below is a real Codeman session on two phones.
@@ -46,6 +46,15 @@ Same keystroke, same link. The only difference is who you wait for: the server,
**No backend changes. No protocol. No server support.** It is a client-side addon that never touches the wire.
Since 0.2.0 the package ships **two addons for two kinds of TUIs**:
| Addon | Model | Use when |
| --------------------- | -------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- |
| `ZerolagInputAddon` | **Buffer**: hold keystrokes locally, flush on Enter | The remote side is a line-oriented prompt (shells, REPLs, Claude Code's composer) that only needs the finished line |
| `PredictiveEchoAddon` | **Predictive write-through**: send every keystroke immediately, paint a prediction, confirm against the parsed buffer | The remote side is a per-keystroke TUI (OpenAI Codex's composer, live pickers) that buffering would starve |
`ZerolagInputAddon` is documented below; jump to [PredictiveEchoAddon](#predictiveechoaddon-write-through-prediction) for the second mode.
## Why this one
| | |
@@ -268,6 +277,110 @@ Finds text that exists after the prompt but was never typed through the overlay.
---
## `PredictiveEchoAddon` (write-through prediction)
Buffering is the wrong model for TUIs that react to every keystroke: a slash
command picker filters live, arrows edit server-side state, the composer
rewraps as it grows. For those, `PredictiveEchoAddon` works like
[mosh](https://mosh.org/): the keystroke goes to the PTY **immediately and
unchanged**, and the addon simultaneously paints the predicted glyph at the
predicted cell. When the real echo lands, the prediction is confirmed and its
span removed: an invisible swap, identical glyph beneath. Mispredictions
self-heal via a mismatch cascade and a TTL. It is visual-only by construction:
nothing it does can gate, delay, reorder or rewrite what you send.
```typescript
import { Terminal } from '@xterm/xterm';
import { PredictiveEchoAddon } from 'xterm-zerolag-input';
const terminal = new Terminal();
const predictor = new PredictiveEchoAddon({
// Optional: only predict when the cursor sits on a composer row
predictWhen: (t) => {
const buf = t.buffer.active;
const line = buf.getLine(buf.baseY + buf.cursorY);
return !!line && /^› /.test(line.translateToString(true));
},
});
terminal.loadAddon(predictor);
terminal.onData((data) => {
const cps = Array.from(data);
if (cps.length === 1) {
const cp = cps[0].codePointAt(0);
if (cp === 0x7f) predictor.predictBackspace();
else if (cp >= 0x20) predictor.predictChar(data);
else predictor.clearPredictions(); // Enter, Ctrl+C, ...
} else if (data.charCodeAt(0) === 0x1b) {
predictor.clearPredictions(); // nav keys, bracketed paste
}
pty.write(data); // ALWAYS, unconditionally
});
```
### How reconciliation works
Predictions are reconciled against the **parsed terminal buffer** (cells after
xterm's parser ran), never the raw output stream. That distinction is
load-bearing: TUIs redraw whole lines, paint gaps with `ECH` + cursor-forward
instead of spaces, and multiplexers like tmux rewrite everything into minimal
deltas. Stream matching breaks on all of that; buffer cells converge to the
same values no matter how the bytes arrived.
A prediction is **confirmed** only when its cell shows the predicted glyph AND
the cursor has advanced past it (so a placeholder that happens to match, or an
identical in-place repaint, never false-confirms). A cell showing foreign
non-blank content on two consecutive passes drops that prediction and all
later ones (one pass tolerates half-parsed frames). Blank cells are neutral:
they are what "not yet echoed" looks like. Whatever remains is dropped by TTL.
Scrolling up, resizing, or a sustained cursor move clears the run. After a
backspace into already-echoed text, a cleared input, or a multi-char commit,
the addon **holds** new predictions until the next parsed write: the displayed
cursor is stale for one round trip, and anchoring on it would paint ghosts one
cell off (worst case: exactly one unpredicted keystroke, whose own echo
releases the hold).
### API
```typescript
predictChar(ch: string): boolean; // false = suppressed (still SEND the key)
predictBackspace(): boolean; // pops the newest prediction (still send \x7f)
clearPredictions(): void;
reconcile(): void; // manual pass (no onWriteParsed available)
setPredictWhen(fn | null): void; // swap the gate at runtime
refreshFont(): void; // after font/theme changes
get hasPredictions(): boolean;
get state(): PredictionState; // { outstanding, confirmedTotal, droppedTotal, anchor }
```
### Options
```typescript
{
zIndex?: number, // Default: 7
underlinePredictions?: boolean, // Default: false (underline unconfirmed glyphs)
foregroundColor?: string, // Default: terminal theme / computed .xterm-rows style
backgroundColor?: string, // Default: terminal theme background
ttlMs?: number, // Default: 1000
maxPending?: number, // Default: 32
cursorGraceMs?: number, // Default: 150
edgeMarginCells?: number, // Default: 4 (suppress near the right edge)
predictWhen?: (t) => boolean, // Default: predict everywhere
}
```
### Which addon should I use?
- The remote program shows a **line prompt** and ignores partial input:
`ZerolagInputAddon`. You also get backspace-before-send and batching.
- The remote program **reacts per keystroke** (pickers, filters, composers
that rewrap): `PredictiveEchoAddon`. It never withholds bytes, so the TUI
behaves exactly as with no addon at all; you just stop waiting for the RTT.
- Both can be loaded on one terminal and toggled per session mode; that is
exactly what Codeman does (buffer for Claude Code, predict for Codex).
---
## Integration patterns
### Buffered input (hold until Enter)
+5 -2
View File
@@ -1,6 +1,6 @@
{
"name": "xterm-zerolag-input",
"version": "0.1.8",
"version": "0.3.1",
"description": "Instant keystroke feedback overlay for xterm.js: Mosh-inspired local echo that removes perceived input latency over SSH, tunnels and other high-RTT connections",
"type": "module",
"main": "dist/index.cjs",
@@ -37,7 +37,9 @@
"ssh",
"remote-terminal",
"overlay",
"addon"
"addon",
"predictive",
"write-through"
],
"license": "MIT",
"homepage": "https://github.com/Ark0N/Codeman/tree/master/packages/xterm-zerolag-input#readme",
@@ -50,6 +52,7 @@
"directory": "packages/xterm-zerolag-input"
},
"devDependencies": {
"@xterm/headless": "^6.0.0",
"jsdom": "^24.1.3",
"tsup": "^8.5.1",
"typescript": "^5.5.0",
@@ -11,37 +11,36 @@ import type { XtermTerminal, CellDimensions } from './types.js';
* unavailable.
*/
export function getCellDimensions(terminal: XtermTerminal): CellDimensions | null {
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const t = terminal as any;
const dpr = typeof devicePixelRatio === 'number' && devicePixelRatio > 0
? devicePixelRatio : 1;
// eslint-disable-next-line @typescript-eslint/no-explicit-any
const t = terminal as any;
const dpr = typeof devicePixelRatio === 'number' && devicePixelRatio > 0 ? devicePixelRatio : 1;
// Try v7+ public API first
if (t.dimensions?.css?.cell) {
const cellH = t.dimensions.css.cell.height;
return {
width: t.dimensions.css.cell.width,
height: cellH,
charTop: (t.dimensions?.device?.char?.top ?? 0) / dpr,
charHeight: (t.dimensions?.device?.char?.height ?? (cellH * dpr)) / dpr,
};
// Try v7+ public API first
if (t.dimensions?.css?.cell) {
const cellH = t.dimensions.css.cell.height;
return {
width: t.dimensions.css.cell.width,
height: cellH,
charTop: (t.dimensions?.device?.char?.top ?? 0) / dpr,
charHeight: (t.dimensions?.device?.char?.height ?? cellH * dpr) / dpr,
};
}
// Fall back to v5 private API
try {
const dims = t._core?._renderService?.dimensions;
if (dims?.css?.cell) {
const cellH = dims.css.cell.height;
return {
width: dims.css.cell.width,
height: cellH,
charTop: (dims.device?.char?.top ?? 0) / dpr,
charHeight: (dims.device?.char?.height ?? cellH * dpr) / dpr,
};
}
} catch {
// Private API may throw in some environments
}
// Fall back to v5 private API
try {
const dims = t._core?._renderService?.dimensions;
if (dims?.css?.cell) {
const cellH = dims.css.cell.height;
return {
width: dims.css.cell.width,
height: cellH,
charTop: (dims.device?.char?.top ?? 0) / dpr,
charHeight: (dims.device?.char?.height ?? (cellH * dpr)) / dpr,
};
}
} catch {
// Private API may throw in some environments
}
return null;
return null;
}
+10 -7
View File
@@ -1,10 +1,13 @@
export { ZerolagInputAddon } from './zerolag-input-addon.js';
export { PredictiveEchoAddon } from './predictive-echo-addon.js';
export { charCellWidth, stringCellWidth } from './overlay-renderer.js';
export type {
XtermTerminal,
XtermAddon,
ZerolagInputOptions,
ZerolagInputState,
PromptFinder,
PromptPosition,
CellDimensions,
XtermTerminal,
XtermAddon,
ZerolagInputOptions,
ZerolagInputState,
PromptFinder,
PromptPosition,
CellDimensions,
} from './types.js';
export type { PredictiveEchoOptions, PredictionState } from './predictive-echo-addon.js';
@@ -65,38 +65,71 @@ export function renderOverlay(container: HTMLDivElement, params: RenderParams):
charTop,
charHeight,
promptRow,
totalRows,
font,
showCursor,
cursorColor,
terminal,
} = params;
// Position container at prompt row.
// ── Keep what is being typed ON SCREEN ────────────────────────────
//
// The overlay lays its wrapped lines out DOWNWARD from the prompt row, and
// nothing past the last terminal row is visible. On a phone the strip left
// above the on-screen keyboard is only a handful of rows, so a prompt long
// enough to wrap ran off the bottom and the user was typing blind — the tail
// of their own sentence, the part they are actually looking at, hidden behind
// the keyboard.
//
// So the composer grows UPWARD once it reaches the last row, exactly as a real
// terminal's does: every line div is opaque (see makeLine), so the lines cover
// transcript rows above instead of vanishing under the keyboard below, and the
// newest text stays where the eye is. A prompt taller than the whole viewport
// keeps its TAIL for the same reason.
//
// `startCol` indents only the line that begins at the prompt marker, so it is
// dropped along with that line when the tail is all that fits.
const rows = totalRows && totalRows > 0 ? totalRows : terminal?.rows;
let visibleLines = lines;
let keepsPromptLine = true;
let topRow = promptRow;
if (rows && rows > 0) {
if (lines.length > rows) {
visibleLines = lines.slice(lines.length - rows);
keepsPromptLine = false;
topRow = 0;
} else if (promptRow + lines.length > rows) {
topRow = rows - lines.length;
}
}
topRow = Math.max(0, topRow);
container.style.left = '0px';
container.style.top = promptRow * cellH + 'px';
container.style.top = topRow * cellH + 'px';
// Clear and rebuild (typically 1-3 line divs, negligible cost)
container.innerHTML = '';
const fullWidthPx = totalCols * cellW;
for (let i = 0; i < lines.length; i++) {
const leftPx = i === 0 ? startCol * cellW : 0;
const widthPx = i === 0 ? fullWidthPx - leftPx : fullWidthPx;
for (let i = 0; i < visibleLines.length; i++) {
const indents = i === 0 && keepsPromptLine;
const leftPx = indents ? startCol * cellW : 0;
const widthPx = indents ? fullWidthPx - leftPx : fullWidthPx;
const topPx = i * cellH;
const lineEl = makeLine(lines[i], leftPx, topPx, widthPx, cellH, cellW, charTop, charHeight, font, terminal);
const lineEl = makeLine(visibleLines[i], leftPx, topPx, widthPx, cellH, cellW, charTop, charHeight, font, terminal);
container.appendChild(lineEl);
}
// Block cursor at end of last line (use visual width for CJK support)
if (showCursor) {
const lastLine = lines[lines.length - 1];
const lastLineLeft = lines.length === 1 ? startCol : 0;
const lastLine = visibleLines[visibleLines.length - 1];
const lastLineLeft = visibleLines.length === 1 && keepsPromptLine ? startCol : 0;
const cursorCol = lastLineLeft + stringCellWidth(terminal, lastLine);
if (cursorCol < totalCols) {
const cursor = document.createElement('span');
cursor.style.cssText = 'position:absolute;display:inline-block';
cursor.style.left = cursorCol * cellW + 'px';
cursor.style.top = (lines.length - 1) * cellH + 'px';
cursor.style.top = (visibleLines.length - 1) * cellH + 'px';
cursor.style.width = cellW + 'px';
cursor.style.height = cellH + 'px';
cursor.style.backgroundColor = cursorColor;
@@ -0,0 +1,59 @@
/**
* Incremental DOM renderer for PredictiveEchoAddon.
*
* Unlike overlay-renderer.ts (which paints whole lines with an opaque
* background out to totalCols), prediction spans cover ONLY the predicted
* glyph's own cells: anything wider would blank real echo arriving around
* a prediction. Spans are keyed by prediction seq for O(1) removal.
*/
import type { CellDimensions, FontStyle } from './types.js';
export interface PredictionSpanParams {
seq: number;
/** Viewport-relative row (0-based). */
row: number;
/** Column (0-based). */
col: number;
char: string;
/** Cell width of the glyph (1 or 2). */
width: 1 | 2;
dims: CellDimensions;
font: FontStyle;
underline: boolean;
}
export function addPredictionSpan(
container: HTMLElement,
map: Map<number, HTMLSpanElement>,
p: PredictionSpanParams
): void {
const span = document.createElement('span');
// cellH+1 height: covers the sub-pixel seam between rows (same trick the
// buffer overlay renderer ships with). Background covers only this glyph's
// cells, never a full row.
span.style.cssText =
`position:absolute;left:${p.col * p.dims.width}px;top:${p.row * p.dims.height}px;` +
`width:${p.width * p.dims.width}px;height:${p.dims.height + 1}px;line-height:${p.dims.height}px;` +
`text-align:center;pointer-events:none;` +
`font-family:${p.font.fontFamily};font-size:${p.font.fontSize};font-weight:${p.font.fontWeight};` +
(p.font.letterSpacing ? `letter-spacing:${p.font.letterSpacing};` : '') +
`color:${p.font.color};background-color:${p.font.backgroundColor};` +
`font-feature-settings:'liga' 0,'calt' 0;` +
(p.underline ? 'text-decoration:underline;' : '');
span.textContent = p.char;
map.set(p.seq, span);
container.appendChild(span);
}
export function removePredictionSpan(map: Map<number, HTMLSpanElement>, seq: number): void {
const span = map.get(seq);
if (span) {
span.remove();
map.delete(seq);
}
}
export function clearAllSpans(map: Map<number, HTMLSpanElement>): void {
for (const span of map.values()) span.remove();
map.clear();
}
@@ -0,0 +1,480 @@
/**
* PredictiveEchoAddon: mosh-style write-through local echo.
*
* The consumer sends every keystroke to the PTY unchanged (write-through);
* this addon simultaneously paints the predicted glyph at the predicted cell.
* When the real echo lands, the prediction is confirmed and its span removed
* (an invisible swap: identical glyph beneath). Mispredictions self-heal via
* a mismatch cascade and a TTL. Everything here is visual-only: no method
* gates, delays, or rewrites what the consumer sends.
*
* Reconciliation reads the parsed terminal BUFFER (cells after xterm's parser
* ran), never the raw output stream. Full-line redraws, ECH-based gap
* painting, and tmux's in-place deltas all converge to the same cells; stream
* matching cannot survive them (see docs/local-echo-overlay-plan.md's
* "What NOT to Do" in the consuming repo).
*
* Coordinate base: xterm's `cursorY` is relative to `baseY`, so the absolute
* buffer line for a viewport row is `baseY + row`. `viewportY` would only
* coincide while scrolled to the bottom; this file never relies on that.
*/
import { getCellDimensions } from './cell-dimensions.js';
import { charCellWidth } from './overlay-renderer.js';
import { addPredictionSpan, clearAllSpans, removePredictionSpan } from './prediction-renderer.js';
import type { FontStyle, XtermAddon, XtermTerminal } from './types.js';
export interface PredictiveEchoOptions {
/** Z-index of the span container. @default 7 (same layer as the buffer overlay) */
zIndex?: number;
/** Render predicted glyphs underlined (visual hedge on unreliable links). @default false */
underlinePredictions?: boolean;
/** Predicted glyph color. @default theme foreground / computed .xterm-rows color */
foregroundColor?: string;
/** Predicted glyph background. @default theme background */
backgroundColor?: string;
/** Drop predictions older than this. @default 1000 */
ttlMs?: number;
/** Maximum outstanding predictions per run. @default 32 */
maxPending?: number;
/** How long the cursor may sit off the anchor row before predictions clear. @default 150 */
cursorGraceMs?: number;
/** Suppress predictions that would land within this many cells of the right edge. @default 4 */
edgeMarginCells?: number;
/** Gate: return false to suppress prediction (e.g. cursor not on a composer row). */
predictWhen?: (terminal: XtermTerminal) => boolean;
}
export interface PredictionState {
outstanding: number;
confirmedTotal: number;
droppedTotal: number;
anchor: { row: number; col: number } | null;
}
interface PredictionRecord {
seq: number;
char: string;
/** Cells this glyph occupies. */
width: 1 | 2;
/** Cumulative cell offset from the anchor column BEFORE this char. */
offsetCells: number;
/** Cell content at predict time, '' normalized to ' '. */
snapshot: string;
sentAt: number;
/** Consecutive reconcile passes that saw foreign non-blank content. */
mismatches: number;
}
const DEFAULT_OPTIONS = {
zIndex: 7,
underlinePredictions: false,
ttlMs: 1000,
maxPending: 32,
cursorGraceMs: 150,
edgeMarginCells: 4,
} as const;
const DEFAULT_BG = '#000000';
const DEFAULT_FG = '#ffffff';
export class PredictiveEchoAddon implements XtermAddon {
private _terminal: XtermTerminal | null = null;
private _container: HTMLDivElement | null = null;
private _spans = new Map<number, HTMLSpanElement>();
private _outstanding: PredictionRecord[] = [];
private _anchor: { row: number; col: number } | null = null;
private _cursorOffRowSince: number | null = null;
private _seq = 0;
private _confirmedTotal = 0;
private _droppedTotal = 0;
private _ttlTimer: ReturnType<typeof setTimeout> | null = null;
/** Anchor hold: set after an unpredicted wire edit (backspace into echoed
* text, any cleared input, an IME text commit). While held, new
* predictions are suppressed: the displayed cursor is stale until the
* next parsed write, and anchoring on it paints ghosts one cell off
* (found by review: backspace-then-retype within RTT). Cleared by the
* onWriteParsed pass and by public reconcile(), never by the inline
* predictChar pass (which runs before the display could catch up). */
private _anchorHold = false;
private _reconcileScheduled = false;
private _disposables: Array<{ dispose(): void }> = [];
private _predictWhen: ((terminal: XtermTerminal) => boolean) | null;
private _options: Required<Omit<PredictiveEchoOptions, 'foregroundColor' | 'backgroundColor' | 'predictWhen'>> &
Pick<PredictiveEchoOptions, 'foregroundColor' | 'backgroundColor'>;
private _font: FontStyle = {
fontFamily: 'monospace',
fontSize: '14px',
fontWeight: 'normal',
color: DEFAULT_FG,
backgroundColor: DEFAULT_BG,
letterSpacing: '',
};
constructor(options?: PredictiveEchoOptions) {
this._options = {
zIndex: options?.zIndex ?? DEFAULT_OPTIONS.zIndex,
underlinePredictions: options?.underlinePredictions ?? DEFAULT_OPTIONS.underlinePredictions,
ttlMs: options?.ttlMs ?? DEFAULT_OPTIONS.ttlMs,
maxPending: options?.maxPending ?? DEFAULT_OPTIONS.maxPending,
cursorGraceMs: options?.cursorGraceMs ?? DEFAULT_OPTIONS.cursorGraceMs,
edgeMarginCells: options?.edgeMarginCells ?? DEFAULT_OPTIONS.edgeMarginCells,
foregroundColor: options?.foregroundColor,
backgroundColor: options?.backgroundColor,
};
this._predictWhen = options?.predictWhen ?? null;
}
// ─── Lifecycle ────────────────────────────────────────────────────
/** Called by `terminal.loadAddon()`. Do not call directly. */
activate(terminal: XtermTerminal): void {
this._terminal = terminal;
this._container = document.createElement('div');
this._container.setAttribute('data-predictive-echo', '');
this._container.style.cssText = `position:absolute;left:0;top:0;z-index:${this._options.zIndex};pointer-events:none`;
const screen = terminal.element?.querySelector('.xterm-screen');
if (screen) screen.appendChild(this._container);
this._readFontStyle();
// Debounced post-parse reconcile: xterm fires onWriteParsed after the
// parser finishes a write chunk, so buffer reads see consistent state.
// The microtask coalesces multi-chunk bursts into one pass.
if (typeof terminal.onWriteParsed === 'function') {
try {
this._disposables.push(
terminal.onWriteParsed(() => {
if (this._reconcileScheduled) return;
this._reconcileScheduled = true;
queueMicrotask(() => {
this._reconcileScheduled = false;
this._anchorHold = false; // a parse pass ran: the display caught up
this._safeReconcile();
});
})
);
} catch {
/* consumers without a working emitter fall back to manual reconcile() */
}
}
if (typeof terminal.onResize === 'function') {
try {
this._disposables.push(terminal.onResize(() => this.clearPredictions()));
} catch {
/* ignore */
}
}
}
dispose(): void {
this.clearPredictions();
for (const d of this._disposables) {
try {
d.dispose();
} catch {
/* ignore */
}
}
this._disposables = [];
this._container?.remove();
this._container = null;
this._terminal = null;
}
// ─── Public API ───────────────────────────────────────────────────
/**
* Predict a single typed character at the current insertion point.
* Returns false when suppressed; the consumer sends the keystroke to the
* PTY either way (the return value is informational, never a send gate).
*/
predictChar(ch: string): boolean {
try {
this._reconcile();
if (this._anchorHold) return false; // display has not caught up with a wire edit
const t = this._terminal;
if (!t || !this._container) return false;
const dims = getCellDimensions(t);
if (!dims) return false;
const buf = t.buffer.active;
if (typeof buf.cursorX !== 'number' || typeof buf.cursorY !== 'number') return false;
if (buf.viewportY !== buf.baseY) return false;
if (this._predictWhen && this._predictWhen(t) === false) return false;
const cps = Array.from(ch);
if (cps.length !== 1) return false;
const cp = cps[0].codePointAt(0)!;
if (cp < 0x20 || cp === 0x7f) return false;
const w = charCellWidth(t, cps[0]);
if (w !== 1 && w !== 2) return false;
if (w === 2 && !this._hasGetCell()) return false; // ASCII fallback misaligns on wide cols
if (this._outstanding.length >= this._options.maxPending) return false;
if (this._outstanding.length === 0) {
this._anchor = { row: buf.cursorY, col: buf.cursorX };
this._cursorOffRowSince = null;
}
const anchor = this._anchor!;
const last = this._outstanding[this._outstanding.length - 1];
const offset = last ? last.offsetCells + last.width : 0;
const col = anchor.col + offset;
if (col + w > t.cols - this._options.edgeMarginCells) return false;
const rec: PredictionRecord = {
seq: this._seq++,
char: cps[0],
width: w,
offsetCells: offset,
snapshot: this._readCell(anchor.row, col),
sentAt: performance.now(),
mismatches: 0,
};
this._outstanding.push(rec);
addPredictionSpan(this._container, this._spans, {
seq: rec.seq,
row: anchor.row,
col,
char: rec.char,
width: w,
dims,
font: this._font,
underline: this._options.underlinePredictions,
});
this._armTtl();
return true;
} catch {
return false;
}
}
/**
* Pop the newest outstanding prediction (visual only). Returns false when
* none are outstanding. The consumer forwards \x7f UNCONDITIONALLY either
* way; deleting already-echoed text renders at RTT.
*/
predictBackspace(): boolean {
try {
const rec = this._outstanding.pop();
if (!rec) {
// \x7f goes to the wire and will delete ECHOED text: the cursor is
// about to move in a way we cannot see yet
this._anchorHold = true;
return false;
}
removePredictionSpan(this._spans, rec.seq);
if (this._outstanding.length === 0) this._resetRun();
return true;
} catch {
return false;
}
}
/** Drop every outstanding prediction and its spans. Also arms the anchor
* hold: consumers clear on inputs (Enter, Esc, arrows, pastes) whose
* cursor effect is unknown until the next parsed write. */
clearPredictions(): void {
try {
this._anchorHold = true;
this._droppedTotal += this._outstanding.length;
this._outstanding = [];
clearAllSpans(this._spans);
this._resetRun();
} catch {
/* ignore */
}
}
/** Manual reconcile pass, for consumers without onWriteParsed. By contract
* it is called after writes parsed, so it also releases the anchor hold. */
reconcile(): void {
this._anchorHold = false;
this._safeReconcile();
}
/** Swap the prediction gate at runtime (mirrors the buffer addon's setPrompt). */
setPredictWhen(fn: ((terminal: XtermTerminal) => boolean) | null): void {
this._predictWhen = fn;
}
/** Re-read font/theme (call after skin or font-size changes). */
refreshFont(): void {
this._readFontStyle();
}
get hasPredictions(): boolean {
return this._outstanding.length > 0;
}
get state(): PredictionState {
return {
outstanding: this._outstanding.length,
confirmedTotal: this._confirmedTotal,
droppedTotal: this._droppedTotal,
anchor: this._anchor ? { ...this._anchor } : null,
};
}
// ─── Reconciliation ───────────────────────────────────────────────
private _safeReconcile(): void {
try {
this._reconcile();
} catch {
/* predictions may degrade, never break input */
}
}
private _reconcile(): void {
const t = this._terminal;
if (!t) return;
if (this._outstanding.length === 0) return; // streaming cost: one boolean
const buf = t.buffer.active;
if (buf.viewportY !== buf.baseY) {
this.clearPredictions(); // user scrolled up
return;
}
if (typeof buf.cursorX !== 'number' || typeof buf.cursorY !== 'number') return; // TTL will clean
const anchor = this._anchor!;
const now = performance.now();
// Off-row grace: transient cursor excursions (repaints park the cursor
// elsewhere mid-frame) are tolerated; a sustained move means the composer
// relocated or the user navigated, so predictions are stale.
if (buf.cursorY !== anchor.row) {
this._cursorOffRowSince ??= now;
if (now - this._cursorOffRowSince > this._options.cursorGraceMs) {
this.clearPredictions();
return;
}
} else {
this._cursorOffRowSince = null;
}
// Confirm loop: PREFIX-ONLY, and only with the cursor advanced past the
// record. Cell match alone is not enough: the predicted char may equal
// pre-existing content (placeholder glyphs), and an identical in-place
// tmux repaint must be a no-op (cells match snapshots, cursor unmoved).
while (this._outstanding.length > 0) {
const rec = this._outstanding[0];
const cell = this._readCell(anchor.row, anchor.col + rec.offsetCells);
if (cell === rec.char && buf.cursorY === anchor.row && buf.cursorX >= anchor.col + rec.offsetCells + rec.width) {
this._outstanding.shift();
removePredictionSpan(this._spans, rec.seq);
this._confirmedTotal++;
} else {
break;
}
}
// Mismatch scan (two-pass rule): a half-parsed row on pass N is fully
// redrawn a few ms later, so only content foreign on TWO consecutive
// passes cascades. Blank cells are NEUTRAL, not foreign: codex clears its
// placeholder on the first echo, and the blanks left under later
// predictions are what "not yet echoed" looks like, not evidence of a
// redraw (measured 2026-08-09; without this, fast typing over the
// placeholder cascades exactly when RTT is high). TTL still bounds them.
let dropFrom = -1;
for (let i = 0; i < this._outstanding.length; i++) {
const rec = this._outstanding[i];
const cell = this._readCell(anchor.row, anchor.col + rec.offsetCells);
if (cell !== rec.snapshot && cell !== rec.char && cell !== ' ') {
rec.mismatches++;
if (rec.mismatches >= 2) {
dropFrom = i;
break;
}
} else {
rec.mismatches = 0;
}
}
if (dropFrom !== -1) this._dropFrom(dropFrom);
// TTL: the first stale record drops itself and everything after it.
for (let i = 0; i < this._outstanding.length; i++) {
if (now - this._outstanding[i].sentAt > this._options.ttlMs) {
this._dropFrom(i);
break;
}
}
if (this._outstanding.length === 0) {
this._resetRun();
} else {
this._armTtl();
}
}
private _dropFrom(index: number): void {
const dropped = this._outstanding.splice(index);
for (const rec of dropped) removePredictionSpan(this._spans, rec.seq);
this._droppedTotal += dropped.length;
}
private _resetRun(): void {
this._anchor = null;
this._cursorOffRowSince = null;
if (this._ttlTimer !== null) {
clearTimeout(this._ttlTimer);
this._ttlTimer = null;
}
}
private _armTtl(): void {
if (this._ttlTimer !== null) return;
const oldest = this._outstanding[0];
if (!oldest) return;
const delay = Math.max(0, oldest.sentAt + this._options.ttlMs - performance.now()) + 1;
this._ttlTimer = setTimeout(() => {
this._ttlTimer = null;
this._safeReconcile();
this._armTtl();
}, delay);
}
// ─── Cell access ──────────────────────────────────────────────────
private _hasGetCell(): boolean {
const buf = this._terminal?.buffer.active;
if (!buf) return false;
const line = buf.getLine(buf.baseY + (buf.cursorY ?? 0));
return typeof line?.getCell === 'function';
}
/** Read one cell's chars at (viewport-relative row, col); '' -> ' '. */
private _readCell(row: number, col: number): string {
const buf = this._terminal!.buffer.active;
const line = buf.getLine(buf.baseY + row);
if (!line) return ' ';
if (typeof line.getCell === 'function') {
const chars = line.getCell(col)?.getChars() ?? '';
return chars === '' ? ' ' : chars;
}
// ASCII fallback: code-unit index, misaligns after wide columns, which is
// why width-2 predictions are suppressed without getCell.
const text = line.translateToString(true);
return text[col] ?? ' ';
}
// ─── Font ─────────────────────────────────────────────────────────
/** Same recipe as the buffer addon's _cacheFont (kept private on purpose:
* zerolag-input-addon.ts must stay untouched by this feature). */
private _readFontStyle(): void {
const t = this._terminal;
if (!t) return;
this._font.fontFamily = t.options.fontFamily || 'monospace';
this._font.fontSize = (t.options.fontSize || 14) + 'px';
this._font.fontWeight = String(t.options.fontWeight || 'normal');
this._font.backgroundColor = this._options.backgroundColor ?? t.options.theme?.background ?? DEFAULT_BG;
this._font.color = this._options.foregroundColor ?? t.options.theme?.foreground ?? DEFAULT_FG;
this._font.letterSpacing = '';
const rows = t.element?.querySelector('.xterm-rows');
if (rows) {
const cs = getComputedStyle(rows);
this._font.letterSpacing = cs.letterSpacing;
if (!this._options.foregroundColor && cs.color) this._font.color = cs.color;
}
}
}
@@ -6,55 +6,50 @@ import type { XtermTerminal, PromptFinder, PromptPosition } from './types.js';
*
* @returns The prompt position (viewport-relative), or `null` if not found.
*/
export function findPrompt(
terminal: XtermTerminal,
finder: PromptFinder,
): PromptPosition | null {
try {
const buffer = terminal.buffer.active;
const viewportTop = buffer.viewportY;
export function findPrompt(terminal: XtermTerminal, finder: PromptFinder): PromptPosition | null {
try {
const buffer = terminal.buffer.active;
const viewportTop = buffer.viewportY;
switch (finder.type) {
case 'character': {
for (let row = terminal.rows - 1; row >= 0; row--) {
const line = buffer.getLine(viewportTop + row);
if (!line) continue;
const text = line.translateToString(true);
const idx = text.lastIndexOf(finder.char);
if (idx >= 0) return { row, col: idx };
}
return null;
}
case 'regex': {
// Create a fresh non-global regex to avoid lastIndex mutation
// and ensure .match() returns a single result with .index
const pattern = finder.pattern;
const safePattern = pattern.global
? new RegExp(pattern.source, pattern.flags.replace('g', ''))
: pattern;
for (let row = terminal.rows - 1; row >= 0; row--) {
const line = buffer.getLine(viewportTop + row);
if (!line) continue;
const text = line.translateToString(true);
const match = text.match(safePattern);
if (match) {
const col = match.index ?? 0;
return { row, col };
}
}
return null;
}
case 'custom':
return finder.find(terminal);
default:
return null;
switch (finder.type) {
case 'character': {
for (let row = terminal.rows - 1; row >= 0; row--) {
const line = buffer.getLine(viewportTop + row);
if (!line) continue;
const text = line.translateToString(true);
const idx = text.lastIndexOf(finder.char);
if (idx >= 0) return { row, col: idx };
}
} catch {
return null;
}
case 'regex': {
// Create a fresh non-global regex to avoid lastIndex mutation
// and ensure .match() returns a single result with .index
const pattern = finder.pattern;
const safePattern = pattern.global ? new RegExp(pattern.source, pattern.flags.replace('g', '')) : pattern;
for (let row = terminal.rows - 1; row >= 0; row--) {
const line = buffer.getLine(viewportTop + row);
if (!line) continue;
const text = line.translateToString(true);
const match = text.match(safePattern);
if (match) {
const col = match.index ?? 0;
return { row, col };
}
}
return null;
}
case 'custom':
return finder.find(terminal);
default:
return null;
}
} catch {
return null;
}
}
/**
@@ -65,19 +60,15 @@ export function findPrompt(
* @param offset - Characters to skip after the prompt marker (e.g., 2 for "> ")
* @returns The text after the prompt, trimmed. Empty string if nothing found.
*/
export function readTextAfterPrompt(
terminal: XtermTerminal,
prompt: PromptPosition,
offset: number,
): string {
try {
const buffer = terminal.buffer.active;
const absRow = buffer.viewportY + prompt.row;
const line = buffer.getLine(absRow);
if (!line) return '';
const lineText = line.translateToString(true);
return lineText.slice(prompt.col + offset).trimEnd();
} catch {
return '';
}
export function readTextAfterPrompt(terminal: XtermTerminal, prompt: PromptPosition, offset: number): string {
try {
const buffer = terminal.buffer.active;
const absRow = buffer.viewportY + prompt.row;
const line = buffer.getLine(absRow);
if (!line) return '';
const lineText = line.translateToString(true);
return lineText.slice(prompt.col + offset).trimEnd();
} catch {
return '';
}
}
+17
View File
@@ -22,9 +22,15 @@ export interface XtermTerminal {
readonly active: {
readonly viewportY: number;
readonly baseY: number;
/** Cursor column (0-based). Used by PredictiveEchoAddon. */
readonly cursorX?: number;
/** Cursor row, relative to baseY (0-based). Used by PredictiveEchoAddon. */
readonly cursorY?: number;
getLine(y: number):
| {
translateToString(trimRight?: boolean): string;
/** Cell access (xterm public API). Optional: mocks/exotic hosts may omit it. */
getCell?(x: number): { getChars(): string; getWidth(): number } | undefined;
}
| undefined;
};
@@ -34,6 +40,10 @@ export interface XtermTerminal {
getStringCellWidth(str: string): number;
activeVersion?: string;
};
/** Fires after the parser finishes a write chunk. Used by PredictiveEchoAddon. */
onWriteParsed?(cb: () => void): { dispose(): void };
/** Fires on terminal resize. Used by PredictiveEchoAddon. */
onResize?(cb: (size: { cols: number; rows: number }) => void): { dispose(): void };
}
/**
@@ -162,6 +172,13 @@ export interface RenderParams {
/** Height of the character rendering area (px). */
charHeight: number;
promptRow: number;
/**
* Visible terminal rows. When given, the overlay is kept ON SCREEN: it grows
* upward instead of running off the bottom edge, and a wrapped prompt taller
* than the viewport keeps its tail. Omit to lay out straight down from
* `promptRow` (the historical behaviour).
*/
totalRows?: number;
font: FontStyle;
showCursor: boolean;
cursorColor: string;
@@ -565,7 +565,10 @@ export class ZerolagInputAddon implements XtermAddon {
// Skip redundant re-renders — include text content to detect
// same-length changes (e.g., setFlushed with different text)
const renderKey = `${displayText}:${startCol}:${activePrompt.row}:${activePrompt.col}:${totalCols}:${this._flushedOffset}`;
// `rows` is part of the key: the layout is clamped to the visible rows
// (see renderOverlay), so a keyboard opening — which changes rows without
// changing the text — must not be skipped as a redundant render.
const renderKey = `${displayText}:${startCol}:${activePrompt.row}:${activePrompt.col}:${totalCols}:${this._terminal.rows}:${this._flushedOffset}`;
if (renderKey === this._lastRenderKey && this._overlay.style.display !== 'none') return;
this._lastRenderKey = renderKey;
@@ -612,6 +615,7 @@ export class ZerolagInputAddon implements XtermAddon {
charTop,
charHeight,
promptRow: activePrompt.row,
totalRows: this._terminal.rows,
font: this._font,
showCursor: this._options.showCursor,
cursorColor,
@@ -6,122 +6,125 @@ import type { XtermTerminal } from '../src/types.js';
let cleanups: (() => void)[] = [];
afterEach(() => {
for (const fn of cleanups) fn();
cleanups = [];
for (const fn of cleanups) fn();
cleanups = [];
});
describe('getCellDimensions', () => {
describe('v5 private API (mock _core._renderService)', () => {
it('returns cell width and height from css.cell', () => {
const mock = createMockTerminal({ cellWidth: 8.4, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
expect(dims!.width).toBe(8.4);
expect(dims!.height).toBe(19);
});
it('returns charTop from device.char.top divided by DPR', () => {
const mock = createMockTerminal({
cellWidth: 8, cellHeight: 19,
deviceCharTop: 2,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// DPR=1 in jsdom, so charTop = 2 / 1 = 2
expect(dims!.charTop).toBe(2);
});
it('returns charHeight from device.char.height divided by DPR', () => {
const mock = createMockTerminal({
cellWidth: 8, cellHeight: 19,
deviceCharHeight: 16,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// DPR=1, so charHeight = 16 / 1 = 16
expect(dims!.charHeight).toBe(16);
});
it('defaults charTop to 0 when device.char not present', () => {
// Default mock has deviceCharTop=0
const mock = createMockTerminal({ cellWidth: 8, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims!.charTop).toBe(0);
});
it('defaults charHeight to cellH when device.char.height not set', () => {
// Default mock has deviceCharHeight=cellH
const mock = createMockTerminal({ cellWidth: 8, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims!.charHeight).toBe(19);
});
describe('v5 private API (mock _core._renderService)', () => {
it('returns cell width and height from css.cell', () => {
const mock = createMockTerminal({ cellWidth: 8.4, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
expect(dims!.width).toBe(8.4);
expect(dims!.height).toBe(19);
});
describe('DPR simulation', () => {
const originalDPR = globalThis.devicePixelRatio;
beforeEach(() => {
// Set DPR=2 to test division
Object.defineProperty(globalThis, 'devicePixelRatio', {
value: 2,
writable: true,
configurable: true,
});
});
afterEach(() => {
Object.defineProperty(globalThis, 'devicePixelRatio', {
value: originalDPR,
writable: true,
configurable: true,
});
});
it('divides device.char.top by DPR', () => {
const mock = createMockTerminal({
cellWidth: 16, cellHeight: 38,
deviceCharTop: 4,
deviceCharHeight: 32,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// charTop = 4 / 2 = 2
expect(dims!.charTop).toBe(2);
// charHeight = 32 / 2 = 16
expect(dims!.charHeight).toBe(16);
});
it('returns charTop from device.char.top divided by DPR', () => {
const mock = createMockTerminal({
cellWidth: 8,
cellHeight: 19,
deviceCharTop: 2,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// DPR=1 in jsdom, so charTop = 2 / 1 = 2
expect(dims!.charTop).toBe(2);
});
describe('null cases', () => {
it('returns null for terminal without _core', () => {
const terminal = {
element: document.createElement('div'),
cols: 80,
rows: 24,
options: {},
buffer: { active: { viewportY: 0, baseY: 0, getLine: () => undefined } },
} as unknown as XtermTerminal;
const dims = getCellDimensions(terminal);
expect(dims).toBeNull();
});
it('returns null for terminal with no dimensions', () => {
const terminal = {
element: document.createElement('div'),
cols: 80,
rows: 24,
options: {},
buffer: { active: { viewportY: 0, baseY: 0, getLine: () => undefined } },
_core: { _renderService: {} },
} as unknown as XtermTerminal;
const dims = getCellDimensions(terminal);
expect(dims).toBeNull();
});
it('returns charHeight from device.char.height divided by DPR', () => {
const mock = createMockTerminal({
cellWidth: 8,
cellHeight: 19,
deviceCharHeight: 16,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// DPR=1, so charHeight = 16 / 1 = 16
expect(dims!.charHeight).toBe(16);
});
it('defaults charTop to 0 when device.char not present', () => {
// Default mock has deviceCharTop=0
const mock = createMockTerminal({ cellWidth: 8, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims!.charTop).toBe(0);
});
it('defaults charHeight to cellH when device.char.height not set', () => {
// Default mock has deviceCharHeight=cellH
const mock = createMockTerminal({ cellWidth: 8, cellHeight: 19 });
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims!.charHeight).toBe(19);
});
});
describe('DPR simulation', () => {
const originalDPR = globalThis.devicePixelRatio;
beforeEach(() => {
// Set DPR=2 to test division
Object.defineProperty(globalThis, 'devicePixelRatio', {
value: 2,
writable: true,
configurable: true,
});
});
afterEach(() => {
Object.defineProperty(globalThis, 'devicePixelRatio', {
value: originalDPR,
writable: true,
configurable: true,
});
});
it('divides device.char.top by DPR', () => {
const mock = createMockTerminal({
cellWidth: 16,
cellHeight: 38,
deviceCharTop: 4,
deviceCharHeight: 32,
});
cleanups.push(mock.cleanup);
const dims = getCellDimensions(mock.terminal as unknown as XtermTerminal);
expect(dims).not.toBeNull();
// charTop = 4 / 2 = 2
expect(dims!.charTop).toBe(2);
// charHeight = 32 / 2 = 16
expect(dims!.charHeight).toBe(16);
});
});
describe('null cases', () => {
it('returns null for terminal without _core', () => {
const terminal = {
element: document.createElement('div'),
cols: 80,
rows: 24,
options: {},
buffer: { active: { viewportY: 0, baseY: 0, getLine: () => undefined } },
} as unknown as XtermTerminal;
const dims = getCellDimensions(terminal);
expect(dims).toBeNull();
});
it('returns null for terminal with no dimensions', () => {
const terminal = {
element: document.createElement('div'),
cols: 80,
rows: 24,
options: {},
buffer: { active: { viewportY: 0, baseY: 0, getLine: () => undefined } },
_core: { _renderService: {} },
} as unknown as XtermTerminal;
const dims = getCellDimensions(terminal);
expect(dims).toBeNull();
});
});
});
@@ -0,0 +1,188 @@
/**
* @vitest-environment jsdom
*
* Layer 2 (the load-bearing suite): the REAL algorithm against the REAL xterm
* parser, fed by fixtures recorded from real codex 0.147 through the
* production pipeline (tmux + the codex full strip). See
* scripts/dev/record-codex-frames.mjs in the consuming repo.
*
* Every replay ends with the convergence invariant: predictions never outlive
* their run (outstanding 0, span container empty).
*/
import { describe, expect, it } from 'vitest';
import { PredictiveEchoAddon } from '../src/predictive-echo-addon.js';
import {
CELL_H,
CELL_W,
classifyPredictInput,
codexComposerGate,
createReplayTerminal,
loadFixture,
type ReplayTerminal,
} from './replay-helpers.js';
async function flushMicrotasks() {
await Promise.resolve();
await Promise.resolve();
}
function sleep(ms: number) {
return new Promise((r) => setTimeout(r, ms));
}
interface KeyEvent {
key: string;
kind: ReturnType<typeof classifyPredictInput>;
painted: boolean;
spansAfter: number;
}
function assertSpansInGrid(rt: ReplayTerminal) {
for (const s of rt.spans()) {
const left = parseFloat(s.style.left);
const width = parseFloat(s.style.width);
const top = parseFloat(s.style.top);
expect(left + width).toBeLessThanOrEqual(rt.hybrid.cols * CELL_W);
expect(top).toBeLessThanOrEqual((rt.hybrid.rows - 1) * CELL_H);
expect(left).toBeGreaterThanOrEqual(0);
expect(top).toBeGreaterThanOrEqual(0);
}
}
async function replay(name: string) {
const { meta, lines } = loadFixture(name);
const rt = createReplayTerminal(meta.cols, meta.rows);
const addon = new PredictiveEchoAddon({ predictWhen: codexComposerGate });
addon.activate(rt.hybrid);
const events: KeyEvent[] = [];
for (const line of lines) {
if (line.keyAt) {
const kind = classifyPredictInput(line.data);
let painted = false;
if (kind === 'char') painted = addon.predictChar(line.data);
else if (kind === 'backspace') addon.predictBackspace();
else addon.clearPredictions(); // 'clear' AND 'text', like the terminal-ui hook
// Span/record parity and grid bounds hold at every step
expect(rt.spanCount()).toBe(addon.state.outstanding);
assertSpansInGrid(rt);
events.push({ key: line.data, kind, painted, spansAfter: rt.spanCount() });
} else {
await rt.write(line.data);
await flushMicrotasks();
}
}
return { rt, addon, events, meta };
}
/** Convergence invariant: after the last chunk + reconcile (+ TTL if needed),
* nothing outlives the run. */
async function converge(rt: ReplayTerminal, addon: PredictiveEchoAddon) {
addon.reconcile();
if (addon.state.outstanding > 0) {
await sleep(1100); // ttlMs default
addon.reconcile();
}
expect(addon.state.outstanding).toBe(0);
expect(rt.spanCount()).toBe(0);
}
describe('codex replay', () => {
it('type-hello: all 5 predictions confirm, zero drops, composer converges', async () => {
const { rt, addon, events } = await replay('type-hello');
const chars = events.filter((e) => e.kind === 'char');
expect(chars).toHaveLength(5);
expect(chars.every((e) => e.painted)).toBe(true);
await converge(rt, addon);
expect(addon.state.confirmedTotal).toBe(5);
expect(addon.state.droppedTotal).toBe(0);
expect(rt.cursorRowText()).toBe('› hello');
addon.dispose();
rt.cleanup();
}, 15000);
it('slash-picker: "/" and filter chars confirm; no ghosts while picker rows redraw', async () => {
const { rt, addon, events } = await replay('slash-picker');
const chars = events.filter((e) => e.kind === 'char');
expect(chars.map((e) => e.key)).toEqual(['/', 'm', 'o']);
expect(chars.every((e) => e.painted)).toBe(true);
await converge(rt, addon);
expect(addon.state.confirmedTotal).toBe(3);
expect(addon.state.droppedTotal).toBe(0);
addon.dispose();
rt.cleanup();
}, 15000);
it('wrap: predictions stay inside the grid, continuation rows fall back to real echo, buffer converges', async () => {
const { rt, addon, events } = await replay('wrap');
// The gate goes false once the cursor is on a wrapped continuation row
// (2-space indent, no "› "): a tail of keystrokes must be suppressed.
const chars = events.filter((e) => e.kind === 'char');
expect(chars.some((e) => !e.painted)).toBe(true);
expect(chars.some((e) => e.painted)).toBe(true);
await converge(rt, addon);
// The composer content is exactly what was typed (word-wrapped)
const b = rt.term.buffer.active;
const cursorRow = b.cursorY;
expect(rt.rowText(cursorRow).trim()).toBe('this line twice over');
expect(rt.rowText(cursorRow - 1)).toMatch(/^› the quick brown fox/);
addon.dispose();
rt.cleanup();
}, 15000);
it('streaming-burst: typed predictions confirm; the re-rendered composer keeps its signature', async () => {
const { rt, addon, events } = await replay('streaming-burst');
const chars = events.filter((e) => e.kind === 'char');
expect(chars).toHaveLength(5); // "hello" (the \r is kind 'clear')
await converge(rt, addon);
expect(addon.state.confirmedTotal).toBe(5);
expect(addon.state.droppedTotal).toBe(0);
// After the 401 burst codex re-renders a fresh composer at the cursor
expect(rt.cursorRowText()).toMatch(/^› /);
addon.dispose();
rt.cleanup();
}, 15000);
it('streaming-real: mid-stream typing survives real baseY growth (recorded with real auth)', async () => {
// The one shape the fake-key lab cannot produce: a genuine model reply
// streaming above the pinned composer pushes lines into history, so
// baseY GROWS while predictions are outstanding: the no-drop-on-baseY
// rule against reality instead of a synthetic scroll.
const { rt, addon, events } = await replay('streaming-real');
expect(rt.term.buffer.active.baseY).toBeGreaterThan(0); // history really grew
const midStream = events.filter((e) => e.kind === 'char' && ['a', 'b', 'c'].includes(e.key));
expect(midStream.length).toBe(3);
expect(midStream.some((e) => e.painted)).toBe(true); // predictions ran mid-stream
await converge(rt, addon);
expect(rt.cursorRowText()).toBe('› abc'); // the mid-stream chars landed intact
addon.dispose();
rt.cleanup();
}, 15000);
it('paste-bracketed: typed chars confirm, the paste clears predictions, content intact', async () => {
const { rt, addon, events } = await replay('paste-bracketed');
const paste = events.find((e) => e.key.startsWith('\x1b[200~'))!;
expect(paste.kind).toBe('clear');
expect(paste.spansAfter).toBe(0);
await converge(rt, addon);
expect(addon.state.confirmedTotal).toBe(2); // 'a', 'b'
expect(rt.cursorRowText()).toContain('abXYZpasted');
addon.dispose();
rt.cleanup();
}, 15000);
it('trust-modal: the predictWhen gate paints ZERO spans on the modal (ghost eliminator)', async () => {
const { rt, addon, events } = await replay('trust-modal');
const x = events.find((e) => e.key === 'x')!;
expect(x.painted).toBe(false);
expect(x.spansAfter).toBe(0);
expect(events.every((e) => e.spansAfter === 0)).toBe(true);
await converge(rt, addon);
expect(addon.state.confirmedTotal).toBe(0);
expect(addon.state.droppedTotal).toBe(0);
// The transition landed on the real composer afterwards
expect(rt.cursorRowText()).toMatch(/^› /);
addon.dispose();
rt.cleanup();
}, 15000);
});
@@ -0,0 +1,28 @@
{"scenario":"paste-bracketed","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T01:51:11.762Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":0,"data":"\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":45,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-bWhHjh\u001b(B\u001b[m$ "}
{"delayMs":638,"data":"exec codex\r\n"}
{"delayMs":420,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":182,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":5,"data":"\r\n\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2;30r\u001b[2;1H\u001bM\u001bM\u001bM\u001b[33m⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b[39m \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[1;30r\u001b[4;1H\u001b(B\u001b[m"}
{"delayMs":1,"data":" \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[6;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-bWhHjh\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[14;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mSummarize rec\u001b(B\u001b[m\u001b[2ment commits\u001b[16;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-bWhHjh\u001b[14;3H\u001b(B\u001b[m"}
{"delayMs":7,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;27H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":21,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;27H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":8,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;27H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":159,"data":"\u001b[6;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mSummarize recent commits\u001b[9;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-bWhHjh\u001b[7;3H\u001b(B\u001b[m"}
{"delayMs":21,"data":"\u001b[5;30r\u001b[5;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\r\n\u001b[2m╭─────────────────────────────────────────────────╮\u001b[1;30r\u001b[7;1H\u001b(B\u001b[m\u001b[2m│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-bWhHjh\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\r\n"}
{"delayMs":0,"data":" reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[18;27H\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[24C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[24C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":3487,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[32m\u001b[1m\u001b[Harkon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-bWhHjh\u001b(B\u001b[m$ exec codex\u001b[K\u001b[33m\r\n⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\u001b[39m\u001b[K\r\n \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\u001b[K\r\n \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[39m\u001b[K\r\n\u001b[K\u001b[2m\r\n╭─────────────────────────────────────────────────╮\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-bWhHjh\u001b[2m │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n╰─────────────────────────────────────────────────╯\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\u001b[K\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\u001b[K\r\n reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[1m\r\n›\u001b(B\u001b[m\u001b[1X\u001b[2m\u001b[CSummarize recent commits\u001b(B\u001b[m\u001b[K\r\n\u001b[K\u001b[20;2H\u001b[1K\u001b[38;5;223m\u001b[Cgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-bWhHjh\u001b[39m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[18;3H"}
{"keyAt":true,"data":"a"}
{"delayMs":207,"data":"a\u001b[K\u001b[20;80H\u001b[K\u001b[18;4H"}
{"keyAt":true,"data":"b"}
{"delayMs":91,"data":"b\u001b[K\u001b[20;80H\u001b[K\u001b[18;5H"}
{"keyAt":true,"data":"\u001b[200~XYZpasted\u001b[201~"}
{"delayMs":383,"data":"XYZpasted\u001b[K\u001b[20;80H\u001b[K\u001b[18;14H"}
@@ -0,0 +1,32 @@
{"scenario":"slash-picker","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T01:50:42.069Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":0,"data":"\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":37,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-bw9Uto\u001b(B\u001b[m$ "}
{"delayMs":647,"data":"exec codex\r\n"}
{"delayMs":437,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":183,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":8,"data":"\r\n\u001b[J\u001b[A\u001b[K\u001b[2;30r\u001b[2;1H\u001bM\u001bM\u001bM\u001b[33m⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b[39m \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\r\n\u001b[39m \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[6;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-bw9Uto\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[14;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[16;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-bw9Uto\u001b[1;30r\u001b[14;3H\u001b(B\u001b[m"}
{"delayMs":9,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;39H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":12,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;39H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":10,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;39H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":157,"data":"\u001b[6;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[9;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-bw9Uto\u001b[7;3H\u001b(B\u001b[m"}
{"delayMs":20,"data":"\u001b[5;30r\u001b[5;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\r\n\u001b[2m╭─────────────────────────────────────────────────╮\u001b[1;30r\u001b[7;1H\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2m│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-bw9Uto\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\r\n"}
{"delayMs":0,"data":" reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[18;39H\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[36C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[36C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":3476,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[32m\u001b[1m\u001b[Harkon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-bw9Uto\u001b(B\u001b[m$ exec codex\u001b[K\u001b[33m\r\n⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\u001b[39m\u001b[K\r\n \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\u001b[K\r\n \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[39m\u001b[K\r\n\u001b[K\u001b[2m\r\n╭─────────────────────────────────────────────────╮\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-bw9Uto\u001b[2m │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n╰─────────────────────────────────────────────────╯\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\u001b[K\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\u001b[K\r\n reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[1m\r\n›\u001b(B\u001b[m\u001b[1X\u001b[2m\u001b[CUse /skills to list available skills\u001b(B\u001b[m\u001b[K\r\n\u001b[K\u001b[20;2H\u001b[1K\u001b[38;5;223m\u001b[Cgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-bw9Uto\u001b[39m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[18;3H"}
{"keyAt":true,"data":"/"}
{"delayMs":207,"data":"\u001b[17;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":1,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m/\u001b[20;3H\u001b[36m\u001b[1m/model choose what model and reasoning effort to use\u001b[21;3H\u001b(B\u001b[m/fast\u001b[10C\u001b[2m1.5x speed, increased usage\u001b[22;3H\u001b(B\u001b[m/ide\u001b[11C\u001b[2minclude current selection, open files, and other context from your IDE\u001b[23;3H\u001b(B\u001b[m/permissions\u001b[3C\u001b[2mchoose what Codex is allowed to do\u001b[24;3H\u001b(B\u001b[m/keymap\u001b[8C\u001b[2mremap TUI shortcuts\u001b[25;3H\u001b(B\u001b[m/vim\u001b[11C\u001b[2mtoggle Vim mode for the composer\u001b[26;3H\u001b(B\u001b[m/experimental\u001b[2C\u001b[2mtoggle experimental features\u001b[27;3H\u001b(B\u001b[m/approve\u001b[7C\u001b[2mapprove one retry of a recent auto-review denial\u001b[18;4H\u001b(B\u001b[m"}
{"keyAt":true,"data":"m"}
{"delayMs":398,"data":"\u001b[17;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":1,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m/m\u001b[20;3H\u001b[36m\u001b[1m/model choose what model and reasoning effort to use\u001b[21;3H\u001b(B\u001b[m/\u001b[1mm\u001b(B\u001b[memories\u001b[2C\u001b[2mconfigure memory use and generation\u001b[22;3H\u001b(B\u001b[m/\u001b[1mm\u001b(B\u001b[mention\u001b[3C\u001b[2mmention a file\u001b[23;3H\u001b(B\u001b[m/\u001b[1mm\u001b(B\u001b[mcp\u001b[7C\u001b[2mlist configured MCP tools; use /mcp verbose for details\u001b[18;5H\u001b(B\u001b[m"}
{"keyAt":true,"data":"o"}
{"delayMs":148,"data":"\u001b[17;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m/mo\u001b[20;3H\u001b[36m\u001b[1m/model choose what model and reasoning effort to use\u001b[18;6H\u001b(B\u001b[m"}
{"keyAt":true,"data":"\u001b"}
@@ -0,0 +1,233 @@
{"scenario":"streaming-burst","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T01:51:03.828Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":0,"data":"\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":39,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b(B\u001b[m$ "}
{"delayMs":635,"data":"exec codex\r\n"}
{"delayMs":439,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":184,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":9,"data":"\r\n\u001b[J\u001b[A\u001b[K\u001b[2;30r\u001b[2;1H\u001bM\u001bM\u001bM\u001b[1;30r\u001b[2;1H\u001b[33m⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b[39m \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\r\n\u001b(B\u001b[m \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[6;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-ruT16A\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[14;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[16;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b[14;3H\u001b(B\u001b[m"}
{"delayMs":8,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;39H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":8,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;39H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":8,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;39H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":158,"data":"\u001b[6;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[9;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b[7;3H\u001b(B\u001b[m"}
{"delayMs":21,"data":"\u001b[5;30r\u001b[5;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\r\n\u001b[2m╭─────────────────────────────────────────────────╮\u001b[1;30r\u001b[7;1H\u001b(B\u001b[m\u001b[2m│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-ruT16A\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\r\n"}
{"delayMs":0,"data":" reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[18;39H\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[36C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":3486,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[32m\u001b[1m\u001b[Harkon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b(B\u001b[m$ exec codex\u001b[K\u001b[33m\r\n⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\u001b[39m\u001b[K\r\n \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\u001b[K\r\n \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[39m\u001b[K\r\n\u001b[K\u001b[2m\r\n╭─────────────────────────────────────────────────╮\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-ruT16A\u001b[2m │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n╰─────────────────────────────────────────────────╯\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\u001b[K\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\u001b[K\r\n reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[1m\r\n›\u001b(B\u001b[m\u001b[1X\u001b[2m\u001b[CUse /skills to list available skills\u001b(B\u001b[m\u001b[K\r\n\u001b[K\u001b[20;2H\u001b[1K\u001b[38;5;223m\u001b[Cgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b[39m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[18;3H"}
{"keyAt":true,"data":"h"}
{"delayMs":199,"data":"h\u001b[K\u001b[20;80H\u001b[K\u001b[18;4H"}
{"keyAt":true,"data":"e"}
{"delayMs":40,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;5H"}
{"keyAt":true,"data":"l"}
{"delayMs":40,"data":"l\u001b[K\u001b[20;80H\u001b[K\u001b[18;6H"}
{"keyAt":true,"data":"l"}
{"delayMs":40,"data":"l\u001b[K\u001b[20;80H\u001b[K\u001b[18;7H"}
{"keyAt":true,"data":"o"}
{"delayMs":40,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;8H"}
{"keyAt":true,"data":"\r"}
{"delayMs":281,"data":"\u001b[16;30r\u001b[16;1H\u001bM\u001bM\u001bM\u001bM\u001b[1;30r\u001b[18;1H"}
{"delayMs":0,"data":"\u001b[1m\u001b[2m› \u001b(B\u001b[mhello\r\n"}
{"delayMs":0,"data":"\u001b[22;3H\u001b[2mUse /skills to list available skills\u001b(B\u001b[m\u001b[K\u001b[24;80H\u001b[K\u001b[22;3H"}
{"delayMs":12,"data":"\u001b[36C\u001b[K\u001b[24;80H\u001b[K\u001b[22;3H"}
{"delayMs":6,"data":"\u001b[36C\u001b[K\u001b[24;80H\u001b[K\u001b[22;3H"}
{"delayMs":6,"data":"\u001b[36C\u001b[K\u001b[24;80H\u001b[K\u001b[22;3H"}
{"delayMs":118,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":1,"data":"\r\n•\u001b[C\u001b[2mWorking\u001b[C(0s • esc to interrupt)\u001b[24;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[26;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b[24;3H\u001b(B\u001b[m"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":34,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":34,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[3AW\u001b[30C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[21;1H\u001b[2m◦\u001b[C\u001b(B\u001b[m\u001b[1mW\u001b(B\u001b[mo\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;4H\u001b[1mo\u001b(B\u001b[mr\u001b[28C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":32,"data":"\u001b[21;5H\u001b[1mr\u001b(B\u001b[mk\u001b[27C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":34,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;6H\u001b[1mk\u001b(B\u001b[mi\u001b[26C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[21;7H\u001b[1mi\u001b(B\u001b[mn\u001b[25C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":34,"data":"\u001b[3AW\u001b[4C\u001b[1mn\u001b(B\u001b[mg\u001b[24C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":32,"data":"\u001b[21;34H\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;12H\u001b[2m1\u001b(B\u001b[m\u001b[21C\u001b[K\u001b[24;39H\u001b[K\u001b[26;80H\u001b[K\u001b[24;3H"}
{"delayMs":19,"data":"\u001b[21;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":1,"data":"\r\n\u001b[2m◦\u001b[CReconne\u001b(B\u001b[mc\u001b[1mting.\u001b(B\u001b[m.\u001b[2m. 2/5\u001b[C(1s • esc to interrupt)\r\n └ Unexpected status 401 Unauthorized: {\r\n \"error\": {\r\n \"message\": \"Incorre, url: wss://api.openai.com/v1/responses, cf-ray: a2831cf59baa039d-ZRH,…\u001b[27;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[29;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-ruT16A\u001b[27;3H\u001b(B\u001b[m"}
{"delayMs":33,"data":"\u001b[21;10H\u001b[2mc\u001b(B\u001b[mt\u001b[4C\u001b[1m.\u001b(B\u001b[m.\u001b[28C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;11H\u001b[2mt\u001b(B\u001b[mi\u001b[4C\u001b[1m.\u001b(B\u001b[m \u001b[27C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;12H\u001b[2mi\u001b(B\u001b[mn\u001b[4C\u001b[1m \u001b(B\u001b[m2\u001b[26C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H•\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;13H\u001b[2mn\u001b(B\u001b[mg\u001b[4C\u001b[1m2\u001b(B\u001b[m/\u001b[25C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;14H\u001b[2mg\u001b(B\u001b[m.\u001b[4C\u001b[1m/\u001b(B\u001b[m5\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":36,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[?25l\u001b[?12l\u001b[?25h\u001b[27;3H"}
{"delayMs":31,"data":"\u001b[21;15H\u001b[2m.\u001b(B\u001b[m.\u001b[4C\u001b[1m5\u001b(B\u001b[m\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;16H\u001b[2m.\u001b(B\u001b[m.\u001b[28C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;17H\u001b[2m.\u001b(B\u001b[m \u001b[27C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":35,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;18H\u001b[2m \u001b(B\u001b[m2\u001b[26C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;19H\u001b[2m2\u001b(B\u001b[m/\u001b[25C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;20H\u001b[2m/\u001b(B\u001b[m5\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[21;21H\u001b[2m5\u001b(B\u001b[m\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H\u001b[2m◦\u001b[27;3H\u001b(B\u001b[m"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":21,"data":"\u001b[21;19H\u001b[2m3\u001b(B\u001b[m\u001b[26C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;85H\u001b[2maca388822\u001b(B\u001b[m\u001b[6C\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;24H\u001b[2m2\u001b(B\u001b[m\u001b[21C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[6AR\u001b[42C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H•\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[6A\u001b[1mR\u001b(B\u001b[me\u001b[41C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":33,"data":"\u001b[21;4H\u001b[1me\u001b(B\u001b[mc\u001b[40C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;5H\u001b[1mc\u001b(B\u001b[mo\u001b[39C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;6H\u001b[1mo\u001b(B\u001b[mn\u001b[38C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;7H\u001b[1mn\u001b(B\u001b[mn\u001b[37C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[6AR\u001b[4C\u001b[1mn\u001b(B\u001b[me\u001b[36C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[6A\u001b[2mR\u001b(B\u001b[me\u001b[4C\u001b[1me\u001b(B\u001b[mc\u001b[35C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;4H\u001b[2me\u001b(B\u001b[mc\u001b[4C\u001b[1mc\u001b(B\u001b[mt\u001b[34C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;5H\u001b[2mc\u001b(B\u001b[mo\u001b[4C\u001b[1mt\u001b(B\u001b[mi\u001b[33C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;6H\u001b[2mo\u001b(B\u001b[mn\u001b[4C\u001b[1mi\u001b(B\u001b[mn\u001b[32C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;7H\u001b[2mn\u001b(B\u001b[mn\u001b[4C\u001b[1mn\u001b(B\u001b[mg\u001b[31C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H\u001b[2m◦\u001b[6Cn\u001b(B\u001b[me\u001b[4C\u001b[1mg\u001b(B\u001b[m.\u001b[8C\u001b[2m3\u001b[27;3H\u001b(B\u001b[m"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;9H\u001b[2me\u001b(B\u001b[mc\u001b[4C\u001b[1m.\u001b(B\u001b[m.\u001b[29C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;10H\u001b[2mc\u001b(B\u001b[mt\u001b[4C\u001b[1m.\u001b(B\u001b[m.\u001b[28C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;100H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":25,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;11H\u001b[2mt\u001b(B\u001b[mi\u001b[4C\u001b[1m.\u001b(B\u001b[m \u001b[2m4\u001b(B\u001b[m\u001b[26C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;27H\u001b[2m, url: ws\u001b[C:/\u001b[Capi.openai.com/v1/responses, cf-ray: a2831d0298dca625-ZRH,\u001b(B\u001b[m\u001b[C\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[24;98H\u001b[2m…\u001b[27;3H\u001b(B\u001b[m"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;12H\u001b[2mi\u001b(B\u001b[mn\u001b[4C\u001b[1m \u001b(B\u001b[m4\u001b[26C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;13H\u001b[2mn\u001b(B\u001b[mg\u001b[4C\u001b[1m4\u001b(B\u001b[m/\u001b[25C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;14H\u001b[2mg\u001b(B\u001b[m.\u001b[4C\u001b[1m/\u001b(B\u001b[m5\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;15H\u001b[2m.\u001b(B\u001b[m.\u001b[4C\u001b[1m5\u001b(B\u001b[m\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[21;16H\u001b[2m.\u001b(B\u001b[m.\u001b[28C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;17H\u001b[2m.\u001b(B\u001b[m \u001b[27C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;18H\u001b[2m \u001b(B\u001b[m4\u001b[26C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;19H\u001b[2m4\u001b(B\u001b[m/\u001b[25C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;20H\u001b[2m/\u001b(B\u001b[m5\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H•\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;21H\u001b[2m5\u001b(B\u001b[m\u001b[24C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":1,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;24H\u001b[2m4\u001b(B\u001b[m\u001b[21C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H\u001b[2m◦\u001b[27;3H\u001b(B\u001b[m"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[6AR\u001b[42C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[6A\u001b[1mR\u001b(B\u001b[me\u001b[41C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;4H\u001b[1me\u001b(B\u001b[mc\u001b[40C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;5H\u001b[1mc\u001b(B\u001b[mo\u001b[39C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;6H\u001b[1mo\u001b(B\u001b[mn\u001b[38C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;7H\u001b[1mn\u001b(B\u001b[mn\u001b[37C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[6AR\u001b[4C\u001b[1mn\u001b(B\u001b[me\u001b[36C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[6A\u001b[2mR\u001b(B\u001b[me\u001b[4C\u001b[1me\u001b(B\u001b[mc\u001b[35C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":34,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
{"delayMs":33,"data":"\u001b[21;46H\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[21;1H•\u001b[2C\u001b[2me\u001b(B\u001b[mc\u001b[4C\u001b[1mc\u001b(B\u001b[mt\u001b[27;3H"}
{"delayMs":32,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[21;5H\u001b[2mc\u001b(B\u001b[mo\u001b[4C\u001b[1mt\u001b(B\u001b[mi\u001b[33C\u001b[K\u001b[22;42H\u001b[K\u001b[23;17H\u001b[K\u001b[24;99H\u001b[K\u001b[27;39H\u001b[K\u001b[29;80H\u001b[K\u001b[27;3H"}
@@ -0,0 +1,154 @@
{"scenario":"streaming-real","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T09:31:57.351Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H"}
{"delayMs":1,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":35,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b(B\u001b[m$ "}
{"delayMs":647,"data":"exec codex\r\n"}
{"delayMs":479,"data":"\u001b[30d\n\u001b[K\u001b[2d\u001b[J\u001b[H\u001b[K\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":2,"data":">\u001b[C\u001b[1mYou are in \u001b(B\u001b[m/home/arkon/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[3;3H\u001b[33mNote: You’re in a subdirectory of a Git project. Trusting will apply to the repository root:\u001b[4;3H/home/arkon/default/claudeman\u001b[6;3H\u001b[39mDo\u001b[Cyou\u001b[Ctrust\u001b[Cthe\u001b[Ccontents\u001b[Cof\u001b[Cthis\u001b[Cdirectory?\u001b[CWorking\u001b[Cwith\u001b[Cuntrusted\u001b[Ccontents\u001b[Ccomes\u001b[Cwith\u001b[Chigher\u001b[7;3Hrisk\u001b[Cof\u001b[Cprompt\u001b[Cinjection.\u001b[CTrusting\u001b[Cthe\u001b[Cdirectory\u001b[Callows\u001b[Cproject-local\u001b[Cconfig,\u001b[Chooks,\u001b[Cand\u001b[Cexec\u001b[8;3Hpolicies\u001b[Cto\u001b[Cload.\u001b[10;1H\u001b[36m› 1. Yes, continue\u001b[11;3H\u001b[39m2.\u001b[CNo,\u001b[Cquit\u001b[13;3H\u001b[2mPress enter to continue\u001b[?25l\u001b(B\u001b[m"}
{"delayMs":3830,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[13;26H\u001b[?25l"}
{"delayMs":1,"data":"\u001b[H>\u001b[1X\u001b[1m\u001b[CYou are in \u001b(B\u001b[m/home/arkon/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[K\r\n\u001b[K\u001b[3;2H\u001b[1K\u001b[33m\u001b[CNote: You’re in a subdirectory of a Git project. Trusting will apply to the repository root:\u001b[39m\u001b[K\u001b[4;2H\u001b[1K\u001b[33m\u001b[C/home/arkon/default/claudeman\u001b[39m\u001b[K\r\n\u001b[K\u001b[6;2H\u001b[1K\u001b[CDo\u001b[1X\u001b[Cyou\u001b[1X\u001b[Ctrust\u001b[1X\u001b[Cthe\u001b[1X\u001b[Ccontents\u001b[1X\u001b[Cof\u001b[1X\u001b[Cthis\u001b[1X\u001b[Cdirectory?\u001b[1X\u001b[CWorking\u001b[1X\u001b[Cwith\u001b[1X\u001b[Cuntrusted\u001b[1X\u001b[Ccontents\u001b[1X\u001b[Ccomes\u001b[1X\u001b[Cwith\u001b[1X\u001b[Chigher\u001b[K\u001b[7;2H\u001b[1K\u001b[Crisk\u001b[1X\u001b[Cof\u001b[1X\u001b[Cprompt\u001b[1X\u001b[Cinjection.\u001b[1X\u001b[CTrusting\u001b[1X\u001b[Cthe\u001b[1X\u001b[Cdirectory\u001b[1X\u001b[Callows\u001b[1X\u001b[Cproject-local\u001b[1X\u001b[Cconfig,\u001b[1X\u001b[Chooks,\u001b[1X\u001b[Cand\u001b[1X\u001b[Cexec\u001b[K\u001b[8;2H\u001b[1K\u001b[Cpolicies\u001b[1X\u001b[Cto\u001b[1X\u001b[Cload.\u001b[K\r\n\u001b[K\u001b[36m\r\n› 1. Yes, continue\u001b[39m\u001b[K\u001b[11;2H\u001b[1K\u001b[C2.\u001b[1X\u001b[CNo,\u001b[1X\u001b[Cquit\u001b[K\r\n\u001b[K\u001b[13;2H\u001b[1K\u001b[2m\u001b[CPress enter to continue\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[13;26H"}
{"keyAt":true,"data":"\r"}
{"delayMs":235,"data":"\u001b[2;1H\u001b[J\u001b[H\u001b[K"}
{"delayMs":0,"data":"\u001bM\u001bM\u001bM\r\n\u001b[33m⚠\u001b[39m\u001b[1;3r\u001b[3;1H\n\u001b[1;2H\u001b[33m Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b[39m \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\r\n\u001b[K\u001b[1;30r\u001b[3;1H"}
{"delayMs":1,"data":" \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[5;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-SFpno1\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills t\u001b(B\u001b[m\u001b[2mo list available skills\u001b[15;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-terra default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[13;3H\u001b[?12l\u001b[?25h\u001b(B\u001b[m"}
{"delayMs":10,"data":"\u001b[5;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[13;39H\u001b[K\u001b[15;82H\u001b[K\u001b[13;3H"}
{"delayMs":12,"data":"\u001b[5;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[13;39H\u001b[K\u001b[15;82H\u001b[K\u001b[13;3H"}
{"delayMs":208,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[5;1H\u001b[J\u001b[A\u001b[K\u001b[4;30r\u001b[4;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001b[1;30r\u001b[5;1H"}
{"delayMs":0,"data":"\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2m│ model: \u001b(B\u001b[mgpt-5.6-terra\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-SFpno1\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[12;1H\u001b(B\u001b[m"}
{"delayMs":0,"data":" \u001b[1mTip:\u001b(B\u001b[m \u001b[3mNew\u001b(B\u001b[m For a limited time, Codex is included in your plan for free – let’s build together.\u001b[14;1H•\u001b[C\u001b[2mBooting MCP server: codex_apps\u001b[C(0s • esc to interrupt)\u001b[17;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[19;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-terra default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[17;3H\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":19,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":1,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":33,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":1,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":33,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":33,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":28,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":34,"data":"\u001b[14;57H\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[3AB\u001b[53C\u001b[K\u001b[17;39H\u001b[K\u001b[19;82H\u001b[K\u001b[17;3H"}
{"delayMs":2,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":5,"data":"\u001b[14;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":1,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[17;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-terra default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[15;3H\u001b(B\u001b[m"}
{"delayMs":279,"data":"\u001b[36C\u001b[K\u001b[17;82H\u001b[K\u001b[15;3H"}
{"delayMs":86,"data":"\u001b[36C\u001b[K\u001b[17;82H\u001b[K\u001b[15;3H"}
{"delayMs":71,"data":"\u001b[36C\u001b[K\u001b[17;82H\u001b[K\u001b[15;3H"}
{"keyAt":true,"data":"r"}
{"keyAt":true,"data":"e"}
{"keyAt":true,"data":"p"}
{"keyAt":true,"data":"l"}
{"keyAt":true,"data":"y"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"w"}
{"keyAt":true,"data":"i"}
{"keyAt":true,"data":"t"}
{"keyAt":true,"data":"h"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"t"}
{"keyAt":true,"data":"h"}
{"keyAt":true,"data":"e"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"s"}
{"keyAt":true,"data":"i"}
{"keyAt":true,"data":"n"}
{"keyAt":true,"data":"g"}
{"keyAt":true,"data":"l"}
{"keyAt":true,"data":"e"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"w"}
{"keyAt":true,"data":"o"}
{"keyAt":true,"data":"r"}
{"keyAt":true,"data":"d"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"h"}
{"keyAt":true,"data":"e"}
{"keyAt":true,"data":"l"}
{"keyAt":true,"data":"l"}
{"keyAt":true,"data":"o"}
{"delayMs":2010,"data":"reply with the single word hello\u001b[K\u001b[17;82H\u001b[K\u001b[15;35H"}
{"keyAt":true,"data":"\r"}
{"delayMs":382,"data":"\u001b[13;30r\u001b[13;1H\u001bM\u001bM\u001bM\u001bM\u001b[1;30r\u001b[15;1H"}
{"delayMs":0,"data":"\u001b[1m\u001b[2m› \u001b(B\u001b[mreply with the single word hello\r\n"}
{"delayMs":0,"data":"\u001b[19;3H\u001b[2mUse /skills to list available skills\u001b(B\u001b[m\u001b[K\u001b[21;82H\u001b[K\u001b[19;3H"}
{"delayMs":0,"data":"\u001b[36C\u001b[K\u001b[21;82H\u001b[K\u001b[19;3H"}
{"delayMs":20,"data":"\u001b[36C\u001b[K\u001b[21;82H\u001b[K\u001b[19;3H"}
{"delayMs":8,"data":"\u001b[36C\u001b[K\u001b[21;82H\u001b[K\u001b[19;3H"}
{"delayMs":78,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\r\n•\u001b[C\u001b[2mWor\u001b(B\u001b[mk\u001b[1ming\u001b[C\u001b(B\u001b[m\u001b[2m(0s • esc to interrupt)\u001b[21;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[23;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-terra default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[21;3H\u001b(B\u001b[m"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;6H\u001b[2mk\u001b(B\u001b[mi\u001b[26C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":35,"data":"\u001b[18;7H\u001b[2mi\u001b(B\u001b[mn\u001b[25C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;8H\u001b[2mn\u001b(B\u001b[mg\u001b[24C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;9H\u001b[2mg\u001b(B\u001b[m\u001b[24C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[18;1H\u001b[2m◦\u001b[21;3H\u001b(B\u001b[m"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":32,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;12H\u001b[2m1\u001b(B\u001b[m\u001b[21C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":32,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[18;1H•\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[?25l\u001b[?12l\u001b[?25h\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":34,"data":"\u001b[3AW\u001b[30C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[3A\u001b[1mW\u001b(B\u001b[mo\u001b[29C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;4H\u001b[1mo\u001b(B\u001b[mr\u001b[28C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;5H\u001b[1mr\u001b(B\u001b[mk\u001b[27C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[18;34H\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":33,"data":"\u001b[18;6H\u001b[1mk\u001b(B\u001b[mi\u001b[26C\u001b[K\u001b[21;39H\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":32,"data":"\u001b[18;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[17;30r\u001b[17;1H\u001bM\u001bM\u001b[1;30r\u001b[18;1H"}
{"delayMs":0,"data":"\u001b[2m• \u001b(B\u001b[mhello\u001b[21;1H\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mUse /skills to list available skills\u001b[23;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-terra default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-SFpno1\u001b[21;3H\u001b(B\u001b[m"}
{"delayMs":25,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":0,"data":"\u001b[36C\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"delayMs":6,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":3,"data":"\u001b[36C\u001b[K\u001b[23;82H\u001b[K\u001b[21;3H"}
{"keyAt":true,"data":"a"}
{"delayMs":2252,"data":"a\u001b[K\u001b[23;82H\u001b[K\u001b[21;4H"}
{"keyAt":true,"data":"b"}
{"delayMs":121,"data":"b\u001b[K\u001b[23;82H\u001b[K\u001b[21;5H"}
{"keyAt":true,"data":"c"}
{"delayMs":121,"data":"c\u001b[K\u001b[23;82H\u001b[K\u001b[21;6H"}
@@ -0,0 +1,28 @@
{"scenario":"trust-modal","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T01:51:20.960Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":0,"data":"\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":28,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-X4gHpE\u001b(B\u001b[m$ "}
{"delayMs":654,"data":"exec codex\r\n"}
{"delayMs":486,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":168,"data":"\u001b[30d\n\u001b[K\u001b[2d\u001b[J\u001b[H\u001b[K"}
{"delayMs":2,"data":">\u001b[C\u001b[1mYou are in \u001b(B\u001b[m/home/arkon/default/claudeman-predictive/tmp/codexrec-work-X4gHpE\u001b[3;3H\u001b[33mNote: You’re in a subdirectory of a Git project. Trusting will apply to the repository root:\u001b[4;3H/home/arkon/default/claudeman\u001b[6;3H\u001b[39mDo\u001b[Cyou\u001b[Ctrust\u001b[Cthe\u001b[Ccontents\u001b[Cof\u001b[Cthis\u001b[Cdirectory?\u001b[CWorking\u001b[Cwith\u001b[Cuntrusted\u001b[Ccontents\u001b[Ccomes\u001b[Cwith\u001b[Chigher\u001b[7;3Hrisk\u001b[Cof\u001b[Cprompt\u001b[Cinjection.\u001b[CTrusting\u001b[Cthe\u001b[Cdirectory\u001b[Callows\u001b[Cproject-local\u001b[Cconfig,\u001b[Chooks,\u001b[Cand\u001b[Cexec\u001b[8;3Hpolicies\u001b[Cto\u001b[Cload.\u001b[10;1H\u001b[36m› 1. Yes, continue\u001b[11;3H\u001b[39m2.\u001b[CNo,\u001b[Cquit\u001b[13;3H\u001b[2mPress enter to continue\u001b[?25l\u001b(B\u001b[m"}
{"delayMs":3659,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[13;26H\u001b[?25l"}
{"delayMs":0,"data":"\u001b[H>\u001b[1X\u001b[1m\u001b[CYou are in \u001b(B\u001b[m/home/arkon/default/claudeman-predictive/tmp/codexrec-work-X4gHpE\u001b[K\r\n\u001b[K\u001b[3;2H\u001b[1K\u001b[33m\u001b[CNote: You’re in a subdirectory of a Git project. Trusting will apply to the repository root:\u001b[39m\u001b[K\u001b[4;2H\u001b[1K\u001b[33m\u001b[C/home/arkon/default/claudeman\u001b[39m\u001b[K\r\n\u001b[K\u001b[6;2H\u001b[1K\u001b[CDo\u001b[1X\u001b[Cyou\u001b[1X\u001b[Ctrust\u001b[1X\u001b[Cthe\u001b[1X\u001b[Ccontents\u001b[1X\u001b[Cof\u001b[1X\u001b[Cthis\u001b[1X\u001b[Cdirectory?\u001b[1X\u001b[CWorking\u001b[1X\u001b[Cwith\u001b[1X\u001b[Cuntrusted\u001b[1X\u001b[Ccontents\u001b[1X\u001b[Ccomes\u001b[1X\u001b[Cwith\u001b[1X\u001b[Chigher\u001b[K\u001b[7;2H\u001b[1K\u001b[Crisk\u001b[1X\u001b[Cof\u001b[1X\u001b[Cprompt\u001b[1X\u001b[Cinjection.\u001b[1X\u001b[CTrusting\u001b[1X\u001b[Cthe\u001b[1X\u001b[Cdirectory\u001b[1X\u001b[Callows\u001b[1X\u001b[Cproject-local\u001b[1X\u001b[Cconfig,\u001b[1X\u001b[Chooks,\u001b[1X\u001b[Cand\u001b[1X\u001b[Cexec\u001b[K\u001b[8;2H\u001b[1K\u001b[Cpolicies\u001b[1X\u001b[Cto\u001b[1X\u001b[Cload.\u001b[K\r\n\u001b[K\u001b[36m\r\n› 1. Yes, continue\u001b[39m\u001b[K\u001b[11;2H\u001b[1K\u001b[C2.\u001b[1X\u001b[CNo,\u001b[1X\u001b[Cquit\u001b[K\r\n\u001b[K\u001b[13;2H\u001b[1K\u001b[2m\u001b[CPress enter to continue\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[13;26H"}
{"keyAt":true,"data":"x"}
{"delayMs":188,"data":"\u001b[1;79H\u001b[K\u001b[3;95H\u001b[K\u001b[4;32H\u001b[K\u001b[6;97H\u001b[K\u001b[7;96H\u001b[K\u001b[8;20H\u001b[K\u001b[10;19H\u001b[K\u001b[11;14H\u001b[K\u001b[13;26H\u001b[K\u001b[30;2H"}
{"keyAt":true,"data":"\r"}
{"delayMs":849,"data":"\u001b[2;1H\u001b[J\u001b[H\u001b[K"}
{"delayMs":0,"data":"\u001bM\u001bM\u001bM\r\n\u001b[33m⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b(B\u001b[m\u001b[1;3r\u001b[3;1H\n\u001b[A \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\r\n\u001b[K\u001b[1;30r\u001b[3;1H"}
{"delayMs":2,"data":" \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[5;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-X4gHpE\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mImprove docum\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2mentation in @filename\u001b[15;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-X4gHpE\u001b[13;3H\u001b[?12l\u001b[?25h\u001b(B\u001b[m"}
{"delayMs":7,"data":"\u001b[5;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[13;37H\u001b[K\u001b[15;80H\u001b[K\u001b[13;3H"}
{"delayMs":13,"data":"\u001b[5;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[13;37H\u001b[K\u001b[15;80H\u001b[K\u001b[13;3H"}
{"delayMs":165,"data":"\u001b[5;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mImprove documentation in @filename\u001b[8;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-X4gHpE\u001b[6;3H\u001b(B\u001b[m"}
{"delayMs":1,"data":"\u001b[34C\u001b[K\u001b[8;80H\u001b[K\u001b[6;3H"}
{"delayMs":24,"data":"\u001b[4;30r\u001b[4;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\r\n\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\u001b[1;30r\u001b[7;1H\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2m│ │\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-X4gHpE\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[12;1H\u001b(B\u001b[m \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\r\n"}
{"delayMs":0,"data":" reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[17;37H\u001b[K\u001b[19;80H\u001b[K\u001b[17;3H"}
{"delayMs":1,"data":"\u001b[34C\u001b[K\u001b[19;80H\u001b[K\u001b[17;3H"}
@@ -0,0 +1,33 @@
{"scenario":"type-hello","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T01:50:33.854Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H"}
{"delayMs":1,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":32,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-tXbGez\u001b(B\u001b[m$ "}
{"delayMs":651,"data":"exec codex\r\n"}
{"delayMs":403,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":189,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":7,"data":"\r\n\u001b[J\u001b[A\u001b[K"}
{"delayMs":4,"data":"\u001b[2;30r\u001b[2;1H\u001bM\u001bM\u001bM\u001b[1;30r\u001b[2;1H\u001b[33m⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b[39m \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\r\n\u001b[39m \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[6;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-tXbGez\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[14;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mImprove documentation in @filename\u001b[16;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-tXbGez\u001b[14;3H\u001b(B\u001b[m"}
{"delayMs":6,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;37H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":25,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;37H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":8,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;37H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":185,"data":"\u001b[6;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[5;30r\u001b[5;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001b[1;30r\u001b[5;1H"}
{"delayMs":0,"data":"\r\n\u001b[2m╭─────────────────────────────────────────────────╮\r\n\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2m│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2m│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-tXbGez\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\r\n"}
{"delayMs":0,"data":" produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\r\n"}
{"delayMs":0,"data":" reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[18;1H\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mImprove documentation in @filename\u001b[20;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-tXbGez\u001b[18;3H\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[34C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[34C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":3484,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[32m\u001b[1m\u001b[Harkon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-tXbGez\u001b(B\u001b[m$ exec codex\u001b[K\u001b[33m\r\n⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\u001b[39m\u001b[K\r\n \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\u001b[K\r\n \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[39m\u001b[K\r\n\u001b[K\u001b[2m\r\n╭─────────────────────────────────────────────────╮\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-tXbGez\u001b[2m │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n╰─────────────────────────────────────────────────╯\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\u001b[K\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\u001b[K\r\n reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[1m\r\n›\u001b(B\u001b[m\u001b[1X\u001b[2m\u001b[CImprove documentation in @filename\u001b(B\u001b[m\u001b[K\r\n\u001b[K\u001b[20;2H\u001b[1K\u001b[38;5;223m\u001b[Cgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-tXbGez\u001b[39m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[18;3H"}
{"keyAt":true,"data":"h"}
{"delayMs":208,"data":"h\u001b[K\u001b[20;80H\u001b[K\u001b[18;4H"}
{"keyAt":true,"data":"e"}
{"delayMs":93,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;5H"}
{"keyAt":true,"data":"l"}
{"delayMs":89,"data":"l\u001b[K\u001b[20;80H\u001b[K\u001b[18;6H"}
{"keyAt":true,"data":"l"}
{"delayMs":92,"data":"l\u001b[K\u001b[20;80H\u001b[K\u001b[18;7H"}
{"keyAt":true,"data":"o"}
{"delayMs":90,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;8H"}
@@ -0,0 +1,250 @@
{"scenario":"wrap","cols":100,"rows":30,"codexVersion":"codex-cli 0.147.0","recordedAt":"2026-08-09T01:50:52.462Z"}
{"delayMs":0,"data":"\u001b[22;0;0t\u001b[?1h\u001b=\u001b[H\u001b[2J\u001b[?12l\u001b[?25h\u001b[?2004h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[c\u001b[>c\u001b[>q\u001b]10;?\u001b\\\u001b]11;?\u001b\\\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":0,"data":"\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[1;1H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[H"}
{"delayMs":32,"data":"\u001b[32m\u001b[1markon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-VGU83J\u001b(B\u001b[m$ "}
{"delayMs":650,"data":"exec codex\r\n"}
{"delayMs":437,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":181,"data":"\u001b[?25l\u001b[?12l\u001b[?25h"}
{"delayMs":4,"data":"\r\n\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[2;30r\u001b[2;1H\u001bM\u001bM\u001bM\u001b[1;30r\u001b[2;1H"}
{"delayMs":0,"data":"\u001b[33m⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\r\n\u001b[39m \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\r\n\u001b(B\u001b[m"}
{"delayMs":1,"data":" \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[6;1H\u001b[39m\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n│ model: \u001b[3mloading\u001b(B\u001b[m\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-VGU83J\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[14;1H\u001b(B\u001b[m\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mWrite tests for @filename\u001b[16;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-VGU83J\u001b[14;3H\u001b(B\u001b[m"}
{"delayMs":6,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;28H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":19,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;28H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":6,"data":"\u001b[6;52H\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\n\u001b[K\u001b[14;28H\u001b[K\u001b[16;80H\u001b[K\u001b[14;3H"}
{"delayMs":160,"data":"\u001b[6;1H\u001b[J\u001b[A\u001b[K"}
{"delayMs":0,"data":"\u001b[5;30r\u001b[5;1H\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001bM\u001b[1;30r\u001b[5;1H"}
{"delayMs":0,"data":"\r\n\u001b[2m╭─────────────────────────────────────────────────╮\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\r\n│ │\r\n\u001b(B\u001b[m"}
{"delayMs":0,"data":"\u001b[2m│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-VGU83J\u001b[2m │\r\n╰─────────────────────────────────────────────────╯\u001b[13;1H\u001b(B\u001b[m \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\r\n"}
{"delayMs":0,"data":" reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[18;1H\u001b[1m›\u001b[C\u001b(B\u001b[m\u001b[2mWrite tests for @filename\u001b[20;3H\u001b(B\u001b[m\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-VGU83J\u001b[18;3H\u001b(B\u001b[m"}
{"delayMs":18,"data":"\u001b[25C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":1,"data":"\u001b[25C\u001b[K\u001b[20;80H\u001b[K\u001b[18;3H"}
{"delayMs":3478,"data":"\u001b[?7727h\u001b(B\u001b[m\u001b[?12l\u001b[?25h\u001b[1;1H\u001b[1;30r\u001b[18;3H"}
{"delayMs":0,"data":"\u001b[?25l\u001b[32m\u001b[1m\u001b[Harkon@tnode\u001b(B\u001b[m:\u001b[34m\u001b[1m~/default/claudeman-predictive/tmp/codexrec-work-VGU83J\u001b(B\u001b[m$ exec codex\u001b[K\u001b[33m\r\n⚠ Codex could not find bubblewrap on PATH. Install bubblewrap with your OS package manager. See the\u001b[39m\u001b[K\r\n \u001b[33msandbox prerequisites: https://developers.openai.com/codex/concepts/sandboxing#prerequisites.\u001b[39m\u001b[K\r\n \u001b[33mCodex will use the bundled bubblewrap in the meantime.\u001b[39m\u001b[K\r\n\u001b[K\u001b[2m\r\n╭─────────────────────────────────────────────────╮\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ >_ \u001b(B\u001b[m\u001b[1mOpenAI Codex\u001b(B\u001b[m\u001b[2m (v0.147.0) │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ model: \u001b(B\u001b[mgpt-5.6-sol\u001b[2m \u001b(B\u001b[m\u001b[36m/model\u001b[39m\u001b[2m to change │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n│ directory: \u001b(B\u001b[m~/default/…/tmp/codexrec-work-VGU83J\u001b[2m │\u001b(B\u001b[m\u001b[K\u001b[2m\r\n╰─────────────────────────────────────────────────╯\u001b(B\u001b[m\u001b[K\r\n\u001b[K\r\n \u001b[1mTip:\u001b(B\u001b[m Our most capable model yet. GPT-5.6 Sol can tackle complex code changes, dig into research,\u001b[K\r\n produce polished documents, and take on your most ambitious work. Sol is highly capable at lower\u001b[K\r\n reasoning efforts—try starting lower, then turn it up for harder jobs.\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[1m\r\n›\u001b(B\u001b[m\u001b[1X\u001b[2m\u001b[CWrite tests for @filename\u001b(B\u001b[m\u001b[K\r\n\u001b[K\u001b[20;2H\u001b[1K\u001b[38;5;223m\u001b[Cgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-VGU83J\u001b[39m\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\r\n\u001b[K\u001b[?12l\u001b[?25h\u001b[18;3H"}
{"keyAt":true,"data":"t"}
{"delayMs":232,"data":"t\u001b[K\u001b[20;80H\u001b[K\u001b[18;4H"}
{"keyAt":true,"data":"h"}
{"delayMs":27,"data":"h\u001b[K\u001b[20;80H\u001b[K\u001b[18;5H"}
{"keyAt":true,"data":"e"}
{"delayMs":27,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;6H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;7H"}
{"keyAt":true,"data":"q"}
{"delayMs":26,"data":"q\u001b[K\u001b[20;80H\u001b[K\u001b[18;8H"}
{"keyAt":true,"data":"u"}
{"delayMs":17,"data":"u\u001b[K\u001b[20;80H\u001b[K\u001b[18;9H"}
{"keyAt":true,"data":"i"}
{"delayMs":28,"data":"i\u001b[K\u001b[20;80H\u001b[K\u001b[18;10H"}
{"keyAt":true,"data":"c"}
{"delayMs":27,"data":"c\u001b[K\u001b[20;80H\u001b[K\u001b[18;11H"}
{"keyAt":true,"data":"k"}
{"delayMs":27,"data":"k\u001b[K\u001b[20;80H\u001b[K\u001b[18;12H"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"b"}
{"delayMs":27,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;13H"}
{"delayMs":16,"data":"b\u001b[K\u001b[20;80H\u001b[K\u001b[18;14H"}
{"keyAt":true,"data":"r"}
{"delayMs":30,"data":"r\u001b[K\u001b[20;80H\u001b[K\u001b[18;15H"}
{"keyAt":true,"data":"o"}
{"delayMs":27,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;16H"}
{"keyAt":true,"data":"w"}
{"delayMs":27,"data":"w\u001b[K\u001b[20;80H\u001b[K\u001b[18;17H"}
{"keyAt":true,"data":"n"}
{"delayMs":27,"data":"n\u001b[K\u001b[20;80H\u001b[K\u001b[18;18H"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"f"}
{"delayMs":45,"data":"\u001b[Cf\u001b[K\u001b[20;80H\u001b[K\u001b[18;20H"}
{"keyAt":true,"data":"o"}
{"delayMs":27,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;21H"}
{"keyAt":true,"data":"x"}
{"delayMs":26,"data":"x\u001b[K\u001b[20;80H\u001b[K\u001b[18;22H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;23H"}
{"keyAt":true,"data":"j"}
{"keyAt":true,"data":"u"}
{"delayMs":27,"data":"j\u001b[K\u001b[20;80H\u001b[K\u001b[18;24H"}
{"keyAt":true,"data":"m"}
{"delayMs":46,"data":"um\u001b[K\u001b[20;80H\u001b[K\u001b[18;26H"}
{"keyAt":true,"data":"p"}
{"delayMs":26,"data":"p\u001b[K\u001b[20;80H\u001b[K\u001b[18;27H"}
{"keyAt":true,"data":"s"}
{"delayMs":28,"data":"s\u001b[K\u001b[20;80H\u001b[K\u001b[18;28H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;29H"}
{"keyAt":true,"data":"o"}
{"delayMs":16,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;30H"}
{"keyAt":true,"data":"v"}
{"delayMs":31,"data":"v\u001b[K\u001b[20;80H\u001b[K\u001b[18;31H"}
{"keyAt":true,"data":"e"}
{"delayMs":26,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;32H"}
{"keyAt":true,"data":"r"}
{"delayMs":28,"data":"r\u001b[K\u001b[20;80H\u001b[K\u001b[18;33H"}
{"keyAt":true,"data":" "}
{"delayMs":26,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;34H"}
{"keyAt":true,"data":"t"}
{"keyAt":true,"data":"h"}
{"delayMs":45,"data":"th\u001b[K\u001b[20;80H\u001b[K\u001b[18;36H"}
{"keyAt":true,"data":"e"}
{"delayMs":27,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;37H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;38H"}
{"keyAt":true,"data":"l"}
{"delayMs":27,"data":"l\u001b[K\u001b[20;80H\u001b[K\u001b[18;39H"}
{"keyAt":true,"data":"a"}
{"keyAt":true,"data":"z"}
{"delayMs":28,"data":"a\u001b[K\u001b[20;80H\u001b[K\u001b[18;40H"}
{"keyAt":true,"data":"y"}
{"delayMs":26,"data":"z\u001b[K\u001b[20;80H\u001b[K\u001b[18;41H"}
{"delayMs":17,"data":"y\u001b[K\u001b[20;80H\u001b[K\u001b[18;42H"}
{"keyAt":true,"data":" "}
{"delayMs":29,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;43H"}
{"keyAt":true,"data":"d"}
{"delayMs":27,"data":"d\u001b[K\u001b[20;80H\u001b[K\u001b[18;44H"}
{"keyAt":true,"data":"o"}
{"delayMs":27,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;45H"}
{"keyAt":true,"data":"g"}
{"delayMs":27,"data":"g\u001b[K\u001b[20;80H\u001b[K\u001b[18;46H"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"a"}
{"delayMs":45,"data":"\u001b[Ca\u001b[K\u001b[20;80H\u001b[K\u001b[18;48H"}
{"keyAt":true,"data":"n"}
{"delayMs":26,"data":"n\u001b[K\u001b[20;80H\u001b[K\u001b[18;49H"}
{"keyAt":true,"data":"d"}
{"delayMs":28,"data":"d\u001b[K\u001b[20;80H\u001b[K\u001b[18;50H"}
{"keyAt":true,"data":" "}
{"delayMs":26,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;51H"}
{"keyAt":true,"data":"k"}
{"keyAt":true,"data":"e"}
{"delayMs":27,"data":"k\u001b[20;80H\u001b[K\u001b[18;52H"}
{"keyAt":true,"data":"e"}
{"delayMs":26,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;53H"}
{"delayMs":17,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;54H"}
{"keyAt":true,"data":"p"}
{"delayMs":28,"data":"p\u001b[K\u001b[20;80H\u001b[K\u001b[18;55H"}
{"keyAt":true,"data":"s"}
{"delayMs":27,"data":"s\u001b[K\u001b[20;80H\u001b[K\u001b[18;56H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;57H"}
{"keyAt":true,"data":"r"}
{"keyAt":true,"data":"u"}
{"delayMs":27,"data":"r\u001b[K\u001b[20;80H\u001b[K\u001b[18;58H"}
{"delayMs":17,"data":"u\u001b[K\u001b[20;80H\u001b[K\u001b[18;59H"}
{"keyAt":true,"data":"n"}
{"delayMs":29,"data":"n\u001b[K\u001b[20;80H\u001b[K\u001b[18;60H"}
{"keyAt":true,"data":"n"}
{"delayMs":26,"data":"n\u001b[K\u001b[20;80H\u001b[K\u001b[18;61H"}
{"keyAt":true,"data":"i"}
{"delayMs":28,"data":"i\u001b[K\u001b[20;80H\u001b[K\u001b[18;62H"}
{"keyAt":true,"data":"n"}
{"delayMs":26,"data":"n\u001b[K\u001b[20;80H\u001b[K\u001b[18;63H"}
{"keyAt":true,"data":"g"}
{"keyAt":true,"data":" "}
{"delayMs":45,"data":"g\u001b[K\u001b[20;80H\u001b[K\u001b[18;65H"}
{"keyAt":true,"data":"u"}
{"delayMs":28,"data":"u\u001b[K\u001b[20;80H\u001b[K\u001b[18;66H"}
{"keyAt":true,"data":"n"}
{"delayMs":27,"data":"n\u001b[K\u001b[20;80H\u001b[K\u001b[18;67H"}
{"keyAt":true,"data":"t"}
{"keyAt":true,"data":"i"}
{"delayMs":27,"data":"t\u001b[K\u001b[20;80H\u001b[K\u001b[18;68H"}
{"keyAt":true,"data":"l"}
{"delayMs":45,"data":"il\u001b[K\u001b[20;80H\u001b[K\u001b[18;70H"}
{"keyAt":true,"data":" "}
{"delayMs":28,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;71H"}
{"keyAt":true,"data":"t"}
{"delayMs":26,"data":"t\u001b[K\u001b[20;80H\u001b[K\u001b[18;72H"}
{"keyAt":true,"data":"h"}
{"keyAt":true,"data":"e"}
{"delayMs":28,"data":"h\u001b[K\u001b[20;80H\u001b[K\u001b[18;73H"}
{"keyAt":true,"data":" "}
{"delayMs":45,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;75H"}
{"keyAt":true,"data":"c"}
{"delayMs":27,"data":"c\u001b[K\u001b[20;80H\u001b[K\u001b[18;76H"}
{"keyAt":true,"data":"o"}
{"delayMs":28,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;77H"}
{"keyAt":true,"data":"m"}
{"keyAt":true,"data":"p"}
{"delayMs":27,"data":"m\u001b[K\u001b[20;80H\u001b[K\u001b[18;78H"}
{"delayMs":16,"data":"p\u001b[K\u001b[20;80H\u001b[K\u001b[18;79H"}
{"keyAt":true,"data":"o"}
{"delayMs":30,"data":"o\u001b[K\u001b[2B\u001b[K\u001b[2A"}
{"keyAt":true,"data":"s"}
{"delayMs":27,"data":"s\u001b[K\u001b[20;80H\u001b[K\u001b[18;81H"}
{"keyAt":true,"data":"e"}
{"delayMs":27,"data":"e\u001b[K\u001b[20;80H\u001b[K\u001b[18;82H"}
{"keyAt":true,"data":"r"}
{"delayMs":27,"data":"r\u001b[K\u001b[20;80H\u001b[K\u001b[18;83H"}
{"keyAt":true,"data":" "}
{"delayMs":16,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;84H"}
{"keyAt":true,"data":"b"}
{"delayMs":30,"data":"b\u001b[K\u001b[20;80H\u001b[K\u001b[18;85H"}
{"keyAt":true,"data":"o"}
{"delayMs":27,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;86H"}
{"keyAt":true,"data":"x"}
{"delayMs":27,"data":"x\u001b[K\u001b[20;80H\u001b[K\u001b[18;87H"}
{"keyAt":true,"data":" "}
{"delayMs":26,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;88H"}
{"keyAt":true,"data":"h"}
{"delayMs":17,"data":"h\u001b[K\u001b[20;80H\u001b[K\u001b[18;89H"}
{"keyAt":true,"data":"a"}
{"delayMs":28,"data":"a\u001b[K\u001b[20;80H\u001b[K\u001b[18;90H"}
{"keyAt":true,"data":"s"}
{"delayMs":27,"data":"s\u001b[K\u001b[20;80H\u001b[K\u001b[18;91H"}
{"keyAt":true,"data":" "}
{"delayMs":28,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;92H"}
{"keyAt":true,"data":"t"}
{"delayMs":26,"data":"t\u001b[K\u001b[20;80H\u001b[K\u001b[18;93H"}
{"keyAt":true,"data":"o"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"o\u001b[K\u001b[20;80H\u001b[K\u001b[18;94H"}
{"delayMs":16,"data":"\u001b[K\u001b[20;80H\u001b[K\u001b[18;95H"}
{"keyAt":true,"data":"w"}
{"delayMs":30,"data":"w\u001b[K\u001b[20;80H\u001b[K\u001b[18;96H"}
{"keyAt":true,"data":"r"}
{"delayMs":27,"data":"r\u001b[K\u001b[20;80H\u001b[K\u001b[18;97H"}
{"keyAt":true,"data":"a"}
{"delayMs":26,"data":"a\u001b[K\u001b[20;80H\u001b[K\u001b[18;98H"}
{"keyAt":true,"data":"p"}
{"delayMs":27,"data":"p\u001b[K\u001b[20;80H\u001b[K\u001b[18;99H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[17;1H\u001b[J\u001b[A\u001b[K"}
{"keyAt":true,"data":"t"}
{"delayMs":1,"data":"\u001b[2B\u001b[1m›\u001b[C\u001b(B\u001b[mthe\u001b[Cquick\u001b[Cbrown\u001b[Cfox\u001b[Cjumps\u001b[Cover\u001b[Cthe\u001b[Clazy\u001b[Cdog\u001b[Cand\u001b[Ckeeps\u001b[Crunning\u001b[Cuntil\u001b[Cthe\u001b[Ccomposer\u001b[Cbox\u001b[Chas\u001b[Cto\u001b[Cwrap\u001b[21;3H\u001b[38;5;223mgpt-5.6-sol default\u001b[39m\u001b[2m · \u001b(B\u001b[m\u001b[38;5;151m~/default/claudeman-predictive/tmp/codexrec-work-VGU83J\u001b[19;3H\u001b(B\u001b[m"}
{"keyAt":true,"data":"h"}
{"delayMs":42,"data":"\u001b[18;99H\u001b[K\u001b[19;3Hth\u001b[21;80H\u001b[K\u001b[19;5H"}
{"keyAt":true,"data":"i"}
{"delayMs":29,"data":"\u001b[18;99H\u001b[K\u001b[19;5Hi\u001b[K\u001b[21;80H\u001b[K\u001b[19;6H"}
{"keyAt":true,"data":"s"}
{"delayMs":27,"data":"\u001b[18;99H\u001b[K\u001b[19;6Hs\u001b[K\u001b[21;80H\u001b[K\u001b[19;7H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[18;99H\u001b[K\u001b[19;7H\u001b[K\u001b[21;80H\u001b[K\u001b[19;8H"}
{"keyAt":true,"data":"l"}
{"delayMs":26,"data":"\u001b[18;99H\u001b[K\u001b[19;8Hl\u001b[K\u001b[21;80H\u001b[K\u001b[19;9H"}
{"keyAt":true,"data":"i"}
{"keyAt":true,"data":"n"}
{"delayMs":45,"data":"\u001b[18;99H\u001b[K\u001b[19;9Hin\u001b[K\u001b[21;80H\u001b[K\u001b[19;11H"}
{"keyAt":true,"data":"e"}
{"delayMs":27,"data":"\u001b[18;99H\u001b[K\u001b[19;11He\u001b[K\u001b[21;80H\u001b[K\u001b[19;12H"}
{"keyAt":true,"data":" "}
{"delayMs":27,"data":"\u001b[18;99H\u001b[K\u001b[19;12H\u001b[K\u001b[21;80H\u001b[K\u001b[19;13H"}
{"keyAt":true,"data":"t"}
{"delayMs":27,"data":"\u001b[18;99H\u001b[K\u001b[19;13Ht\u001b[K\u001b[21;80H\u001b[K\u001b[19;14H"}
{"keyAt":true,"data":"w"}
{"keyAt":true,"data":"i"}
{"delayMs":45,"data":"\u001b[18;99H\u001b[K\u001b[19;14Hwi\u001b[K\u001b[21;80H\u001b[K\u001b[19;16H"}
{"keyAt":true,"data":"c"}
{"delayMs":28,"data":"\u001b[18;99H\u001b[K\u001b[19;16Hc\u001b[K\u001b[21;80H\u001b[K\u001b[19;17H"}
{"keyAt":true,"data":"e"}
{"delayMs":26,"data":"\u001b[18;99H\u001b[K\u001b[19;17He\u001b[K\u001b[21;80H\u001b[K\u001b[19;18H"}
{"keyAt":true,"data":" "}
{"keyAt":true,"data":"o"}
{"delayMs":28,"data":"\u001b[18;99H\u001b[K\u001b[19;18H\u001b[K\u001b[21;80H\u001b[K\u001b[19;19H"}
{"keyAt":true,"data":"v"}
{"delayMs":26,"data":"\u001b[18;99H\u001b[K\u001b[19;19Ho\u001b[K\u001b[21;80H\u001b[K\u001b[19;20H"}
{"keyAt":true,"data":"e"}
{"delayMs":26,"data":"\u001b[18;99H\u001b[K\u001b[19;20Hv\u001b[K\u001b[21;80H\u001b[K\u001b[19;21H"}
{"delayMs":17,"data":"\u001b[18;99H\u001b[K\u001b[19;21He\u001b[K\u001b[21;80H\u001b[K\u001b[19;22H"}
{"keyAt":true,"data":"r"}
{"delayMs":29,"data":"\u001b[18;99H\u001b[K\u001b[19;22Hr\u001b[K\u001b[21;80H\u001b[K\u001b[19;23H"}
+183 -108
View File
@@ -3,129 +3,204 @@
*
* Creates a minimal Terminal-like object that satisfies the addon's
* requirements without needing a real xterm.js instance or DOM renderer.
*
* PredictiveEchoAddon additions (all ADDITIVE, existing tests unchanged):
* mutable cursor via setCursor(), wide-char-aware getCell() on mock lines,
* onWriteParsed/onResize emitters with fire* triggers, and opt-outs for
* getCell support and the emitters (getCellSupport / emitters options).
*/
import { charCellWidth } from '../src/overlay-renderer.js';
interface MockLine {
translateToString(_trimRight?: boolean): string;
translateToString(_trimRight?: boolean): string;
getCell?(x: number): { getChars(): string; getWidth(): number } | undefined;
}
interface MockBufferOptions {
lines: string[];
viewportY?: number;
baseY?: number;
cursorX?: number;
cursorY?: number;
lines: string[];
viewportY?: number;
baseY?: number;
cursorX?: number;
cursorY?: number;
}
interface MockTerminalOptions {
buffer?: MockBufferOptions;
cols?: number;
rows?: number;
fontFamily?: string;
fontSize?: number;
fontWeight?: string | number;
theme?: {
background?: string;
foreground?: string;
cursor?: string;
};
cellWidth?: number;
cellHeight?: number;
/** Device-pixel char top offset (for charTop calculation). Default: 0 */
deviceCharTop?: number;
/** Device-pixel char height (for charHeight calculation). Default: cellHeight * dpr */
deviceCharHeight?: number;
buffer?: MockBufferOptions;
cols?: number;
rows?: number;
fontFamily?: string;
fontSize?: number;
fontWeight?: string | number;
theme?: {
background?: string;
foreground?: string;
cursor?: string;
};
cellWidth?: number;
cellHeight?: number;
/** Device-pixel char top offset (for charTop calculation). Default: 0 */
deviceCharTop?: number;
/** Device-pixel char height (for charHeight calculation). Default: cellHeight * dpr */
deviceCharHeight?: number;
/** Provide getCell() on mock lines (PredictiveEchoAddon). Default: true */
getCellSupport?: boolean;
/** Provide onWriteParsed/onResize emitters (PredictiveEchoAddon). Default: true */
emitters?: boolean;
}
/** Column-indexed cell access over a plain string, wide-char aware. */
function cellAt(text: string, col: number): { getChars(): string; getWidth(): number } {
let c = 0;
for (const ch of text) {
const w = charCellWidth(null, ch);
if (col === c) return { getChars: () => ch, getWidth: () => w };
if (w === 2 && col === c + 1) return { getChars: () => '', getWidth: () => 0 };
c += w;
}
return { getChars: () => '', getWidth: () => 1 };
}
export function createMockTerminal(opts: MockTerminalOptions = {}) {
const bufOpts = opts.buffer ?? { lines: ['$ '] };
const lines = bufOpts.lines;
const viewportY = bufOpts.viewportY ?? 0;
const baseY = bufOpts.baseY ?? viewportY;
const cols = opts.cols ?? 80;
const rows = opts.rows ?? Math.max(lines.length, 24);
const cellW = opts.cellWidth ?? 8.4;
const cellH = opts.cellHeight ?? 17;
const bufOpts = opts.buffer ?? { lines: ['$ '] };
const viewportY = bufOpts.viewportY ?? 0;
const baseY = bufOpts.baseY ?? viewportY;
const cols = opts.cols ?? 80;
const rows = opts.rows ?? Math.max(bufOpts.lines.length, 24);
const cellW = opts.cellWidth ?? 8.4;
const cellH = opts.cellHeight ?? 17;
const getCellSupport = opts.getCellSupport ?? true;
const emitters = opts.emitters ?? true;
const mockLines: MockLine[] = lines.map((text) => ({
translateToString: () => text,
}));
// Create minimal DOM structure
const element = document.createElement('div');
element.className = 'terminal xterm';
const viewport = document.createElement('div');
viewport.className = 'xterm-viewport';
const screen = document.createElement('div');
screen.className = 'xterm-screen';
screen.style.position = 'relative';
const xtermRows = document.createElement('div');
xtermRows.className = 'xterm-rows';
element.appendChild(viewport);
element.appendChild(screen);
screen.appendChild(xtermRows);
// Append to document so getComputedStyle works
document.body.appendChild(element);
const terminal = {
element,
cols,
rows,
options: {
fontFamily: opts.fontFamily ?? 'monospace',
fontSize: opts.fontSize ?? 14,
fontWeight: opts.fontWeight ?? 'normal',
theme: opts.theme ?? {},
},
buffer: {
active: {
viewportY,
baseY,
cursorX: bufOpts.cursorX ?? 0,
cursorY: bufOpts.cursorY ?? 0,
getLine: (absRow: number): MockLine | undefined => {
return mockLines[absRow - viewportY];
},
},
},
_core: {
_renderService: {
dimensions: {
css: {
cell: { width: cellW, height: cellH },
},
device: {
char: {
top: opts.deviceCharTop ?? 0,
height: opts.deviceCharHeight ?? cellH,
},
},
},
},
},
// Simulate loadAddon
loadAddon(addon: { activate: (t: unknown) => void }) {
addon.activate(this);
},
const makeLine = (text: string): { line: MockLine; set(t: string): void } => {
let current = text;
const line: MockLine = {
translateToString: () => current,
};
if (getCellSupport) {
line.getCell = (x: number) => cellAt(current, x);
}
return { line, set: (t: string) => (current = t) };
};
return {
terminal,
/** Update buffer lines for subsequent calls */
setLines(newLines: string[]) {
mockLines.length = 0;
for (const text of newLines) {
mockLines.push({ translateToString: () => text });
}
let mockLines = bufOpts.lines.map(makeLine);
// Create minimal DOM structure
const element = document.createElement('div');
element.className = 'terminal xterm';
const viewport = document.createElement('div');
viewport.className = 'xterm-viewport';
const screen = document.createElement('div');
screen.className = 'xterm-screen';
screen.style.position = 'relative';
const xtermRows = document.createElement('div');
xtermRows.className = 'xterm-rows';
element.appendChild(viewport);
element.appendChild(screen);
screen.appendChild(xtermRows);
// Append to document so getComputedStyle works
document.body.appendChild(element);
const writeParsedCbs = new Set<() => void>();
const resizeCbs = new Set<(s: { cols: number; rows: number }) => void>();
const terminal = {
element,
cols,
rows,
options: {
fontFamily: opts.fontFamily ?? 'monospace',
fontSize: opts.fontSize ?? 14,
fontWeight: opts.fontWeight ?? 'normal',
theme: opts.theme ?? {},
},
buffer: {
active: {
viewportY,
baseY,
cursorX: bufOpts.cursorX ?? 0,
cursorY: bufOpts.cursorY ?? 0,
getLine: (absRow: number): MockLine | undefined => {
return mockLines[absRow - viewportY]?.line;
},
/** Clean up DOM */
cleanup() {
element.remove();
},
},
_core: {
_renderService: {
dimensions: {
css: {
cell: { width: cellW, height: cellH },
},
device: {
char: {
top: opts.deviceCharTop ?? 0,
height: opts.deviceCharHeight ?? cellH,
},
},
},
};
},
},
...(emitters
? {
onWriteParsed(cb: () => void) {
writeParsedCbs.add(cb);
return { dispose: () => writeParsedCbs.delete(cb) };
},
onResize(cb: (s: { cols: number; rows: number }) => void) {
resizeCbs.add(cb);
return { dispose: () => resizeCbs.delete(cb) };
},
}
: {}),
// Simulate loadAddon
loadAddon(addon: { activate: (t: unknown) => void }) {
addon.activate(this);
},
};
return {
terminal,
/** Update buffer lines for subsequent calls */
setLines(newLines: string[]) {
mockLines = newLines.map(makeLine);
},
/** Update one line's text in place (PredictiveEchoAddon echo simulation) */
setLine(index: number, text: string) {
mockLines[index]?.set(text);
},
/** Move the mock cursor (PredictiveEchoAddon) */
setCursor(x: number, y: number) {
terminal.buffer.active.cursorX = x;
terminal.buffer.active.cursorY = y;
},
/** Set scroll state (viewportY / baseY) */
setScroll(newViewportY: number, newBaseY: number) {
terminal.buffer.active.viewportY = newViewportY;
terminal.buffer.active.baseY = newBaseY;
},
/** Fire the onWriteParsed emitter (PredictiveEchoAddon reconcile trigger) */
fireWriteParsed() {
for (const cb of [...writeParsedCbs]) cb();
},
/** Fire the onResize emitter */
fireResize(newCols = cols, newRows = rows) {
for (const cb of [...resizeCbs]) cb({ cols: newCols, rows: newRows });
},
/** Number of live onWriteParsed listeners (dispose assertions) */
writeParsedListenerCount() {
return writeParsedCbs.size;
},
/** Number of live onResize listeners (dispose assertions) */
resizeListenerCount() {
return resizeCbs.size;
},
/** Clean up DOM */
cleanup() {
element.remove();
},
};
}

Some files were not shown because too many files have changed in this diff Show More