Compare commits

...
Author SHA1 Message Date
Codeman maintainer a922d301b1 chore: version packages 2026-08-21 20:24:38 +02:00
Ark0N abca552676 Merge pull request #327 from dignfei/fix/terminal-ime-punctuation
fix(terminal): preserve IME punctuation input
2026-08-21 20:23:26 +02:00
Ark0N 12a996b107 Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay
2026-08-21 20:23:17 +02:00
d fei 458e751a33 fix(terminal): keep shell history loading explicit 2026-08-22 01:55:50 +08:00
d fei dab432b3fd fix(terminal): bound shell history replay 2026-08-21 08:23:31 -04:00
Codeman maintainer 79a0399552 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:47:10 +02:00
Codeman maintainer 61251c0b94 fix(cli-resolvers): negative-result caching, SIGKILL on probes, restored VITEST hermeticity, wired not-found diagnostics
Post-merge follow-ups for PR #329 (shared CLI executable resolution):

- Negative-cache resolution misses with a doubling backoff (1min -> 5min
  cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
  shared resolver cached success only, so a missing CLI re-ran the whole
  chain - ending in a synchronous interactive login-shell spawn bounded by
  the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
  attempt, stalling the event loop each time, forever. Success still caches
  for the process lifetime, so an installed CLI is picked up within minutes
  without a restart. Tests drive the backoff via an injectable clock
  (createCliExecutableResolver `now` option, threaded through the
  createPiResolverForTest / createAntigravityResolverForTest wrappers).

- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
  pi/claude --version probes: execFileSync's timeout only SENDS the kill
  signal and then keeps waiting for the child to exit, and interactive bash
  ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
  survived the timeout and blocked the server permanently.

- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
  test pinned the deletion): under vitest the production resolver host now
  replaces un-injected IO primitives with inert stubs - no real PATH
  scanning, no login-shell spawns - and probePiVersion never executes a
  `pi` candidate again (`pi` is a generic binary name, so route tests
  hitting /api/pi/status executed whatever binary the machine carried).
  Tests opt in through the runCommand/isExecutableFile injection hooks or
  allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
  test is replaced by behavioral pins, including a real-executable fixture
  in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
  is ever removed again.

- Wire the six get*NotFoundMessage() exports (previously dead) into their
  intended call sites: the createSession throws in tmux-manager and the
  availability gates on POST /api/sessions and POST /api/quick-start in
  session-routes, replacing a third hardcoded copy of the text. A not-found
  error now names where resolution looked (server PATH, login shell,
  checked directories). npm run knip no longer reports any unused export
  from the resolver modules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:37:58 +02:00
Codeman maintainer bb4ba79791 fix(repo-status): async git, single-flight TTL cache, credential redaction, local-upstream parse
Post-merge follow-ups for #328 (GET /api/system/repo-status):

- Event-loop blocking: every git invocation in repo-status.ts is now async
  (promisified execFile), never execFileSync — the per-remote ls-remote +
  fetch could hold the event loop (SSE, PTY streaming) for up to ~60s per
  request. The whole computation is single-flight with a 45s TTL cache
  (createSingleFlightCache): concurrent requests share one in-flight
  promise, a fresh result is served without spawning git, and a rejected
  compute is never cached. Route handler shape and response fields
  unchanged; remotes still processed sequentially (concurrent fetches in
  one repo contend on ref locks).

- Credential disclosure: the redaction from git-clone.ts is extracted as
  exported redactGitCredentials() (sanitizeGitOutput now uses it) and
  applied via redactRemoteStatus() to every remote card's url and error
  string, so a scheme://user:token@host remote URL (or git stderr echoing
  it) never reaches a client.

- Non-interactive env: runGit() now uses the shared gitNonInteractiveEnv()
  instead of a partial GIT_TERMINAL_PROMPT/BatchMode env, also closing the
  GIT_ASKPASS/SSH_ASKPASS/SSH_ASKPASS_REQUIRE/DISPLAY/GCM_INTERACTIVE
  prompt paths.

- Upstream parse bug: a local-branch upstream (@{upstream} with no slash,
  e.g. after `git branch -u otherbranch`) made slice(0, indexOf('/')) into
  slice(0, -1) and yielded garbage like "maste". parseTrackingRemote()
  (pure, unit-tested) returns null for it, and the bare ref is dropped so
  it cannot be mistaken for a remote-tracking ref downstream.

Tests extended in test/repo-status.test.ts (parseTrackingRemote,
redactGitCredentials/redactRemoteStatus, createSingleFlightCache
single-flight/TTL/rejection semantics).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Codeman maintainer d7ad73bc9b fix(response-viewer): role on full-context blocks, divider ReDoS, pi mode (#326 follow-up)
Three post-merge fixes for the external-CLI response viewer:

- ?context=full blocks now carry role ('user' for prompts, 'assistant'
  for response/status/tool). The frontend's loadFullContext() renders
  via msg.role, so the roleless blocks lost the "You" badge and every
  turn rendered as the agent. kind/label/text are unchanged and the
  frontend needs no change.

- normalizeDividerStatusLine() dropped its backtracking regex
  (/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
  a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
  minutes at 10,000), and pane text is agent-controlled with buffers up
  to 32MB. Replaced by a linear counter walk with the identical accept
  set and captured content, pinned char-for-char against the old regex
  by a brute-force corpus test plus a hostile-input regression test
  that fails by timeout with the RegExp version (same approach as the
  glob-matcher hardening in 68ae9a8).

- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
  empty-viewer symptom the transcript branch exists to fix. The list
  stays a local duplicate of isExternalCliMode() (importing session.ts
  would drag node-pty into the pure module); a new exhaustive parity
  test asserts the two mode sets can no longer drift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Codeman maintainer 96ee8b536d docs: update the tap-report gate description after #325
#325 renamed _sessionUsesServerMouseStrip to _shouldReportMouseToCli and
added the server-observed cliMouseTracking half of the gate, which also
turned codex tap reports from measured no-ops into not-sent-at-all. The
invariants paragraph still described the old name and the old behavior.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:13:26 +02:00
Ark0N b7f3b07c79 Merge pull request #330 from aakhter/ralph-loop-reschedule
Ralph loop silently stops polling after two ticks
2026-08-21 02:10:41 +02:00
Ark0N 2073a1b185 Merge pull request #328 from aakhter/repo-status-panel
Report repository status for git-clone installs (GET /api/system/repo-status)
2026-08-21 02:10:38 +02:00
Ark0N f9a8493823 Merge pull request #329 from aakhter/cli-login-shell-resolution
CLIs installed via nvm/Homebrew are not found when Codeman runs as a service
2026-08-21 02:10:35 +02:00
Ark0N d2711ef092 Merge pull request #326 from aakhter/response-viewer-external-cli
Response viewer is empty for OpenCode / Gemini / Antigravity sessions
2026-08-21 02:10:29 +02:00
Ark0N 30a15adbd6 Merge pull request #325 from Ark0N/feat/auto-copy-selection
feat(terminal): Auto Copy, put a finished selection on the clipboard
2026-08-21 02:10:19 +02:00
Aamer Akhter a35438ba34 fix(ralph): loop stops rescheduling after two ticks
The reschedule guard is `this._status === 'running' && this.loopTimer === null`,
but the timer callback never nulls `loopTimer`. So the handle stays non-null from
the first fire onward, the guard is false on every subsequent pass, and the Ralph
loop silently stops polling after exactly two ticks.

It stops without changing status: `status` stays `running`, `stop()` is never
called, and no error is raised — the loop just quietly never runs again, which is
what makes it hard to notice on a long autonomous run.

Null the handle inside the callback before re-entering `runLoop()`, which is the
pattern `orchestrator-loop.ts` already uses for its own reschedule.

Test: a regression case in test/ralph-loop.test.ts that runs a real 5ms-interval
loop for ~16 intervals and asserts it ticks at least 3 times. Against the unfixed
source it reports exactly 2.
2026-08-20 12:58:17 -04:00
Aamer Akhter fef903df98 fix(cli-resolvers): find CLIs installed via nvm/Homebrew when running as a service
A CLI installed by nvm, Homebrew or a user-level npm prefix lives on a PATH that
only a login shell sets up. Codeman running under systemd or launchd does not get
that PATH — launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` — so every
resolver reported the CLI as unavailable on installs where it is plainly there
and works from a terminal.

Each of the six resolvers had its own hand-rolled copy of the same PATH walk, so
the fix is factored into one shared `createCliExecutableResolver()` with an
explicit lookup order: the server process PATH, then common install directories in
order, then an interactive login shell as the last resort. Only the last step
spawns anything, and only when the cheap lookups have already missed.

Also adds `formatCliNotFoundMessage()`, so a failure explains where it looked
instead of just asserting the CLI is missing. Its diagnostics are bounded and
control characters are flattened, so a not-found message cannot dump arbitrary
environment data.

Success is cached and failure is retried, so installing a CLI while the server is
running is picked up without a restart.

Net -103 lines across the six resolvers. Behaviour is unchanged wherever the CLI
was already on the process PATH: that remains the first thing checked.

Tests: 20 cases in test/cli-executable-resolver.test.ts covering the precedence
order, login-shell-only resolution, the caching rule, unsafe-name rejection, and
the bounded diagnostics.
2026-08-20 12:47:42 -04:00
Aamer Akhter 02e7d3fcba feat(system): report repository status for git-clone installs
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.

Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.

Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.

Tests: 24 cases in test/repo-status.test.ts.
2026-08-20 12:39:28 -04:00
d fei f744719650 fix(terminal): preserve IME punctuation input 2026-08-20 10:35:23 -04:00
Aamer Akhter 63c5ba89da fix(response-viewer): populate the viewer for OpenCode/Gemini/Antigravity panes
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.

For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.

Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.

The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.

Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
2026-08-20 09:24:37 -04:00
Codeman maintainer 7fc4784d0f fix(approvals): clear the red tab alert when a dialog is answered in the terminal
Confirming an AskUserQuestion left its tab flowing red for the rest of
the turn (owner report: ~8 minutes on a running session, with no dialog
anywhere on screen). Two separate bugs, both live-verified.

The re-capture erased the evidence the staleness check runs on. Claude
Code fires the Notification behind the dialog (measured 6-7s on v2.1.237,
documented up to ~30s), so the 600ms re-capture routinely lands on a
frame the user has ALREADY answered, parses nothing, and applyCapture
overwrote item.options with undefined. A MISSING options is how "we never
could read this dialog" is expressed, and those items stay answerable by
design, so a cleared field was indistinguishable from a never-parsed one
and the item became permanently unsweepable: it survived every
GET /api/approvals and every page reload, cleared only on `stop`, and
still accepted an answer, sending a bare `1` into a composer with no
dialog under it. applyCapture is now ADD-ONLY for options.

Nothing ran the staleness check while a page was open. It lived only in
GET /api/approvals, which seedApprovals() calls on init and reconnect, so
`stop` was the first thing that ever cleared an answered dialog. The
`working` signal now runs the pane-VERIFIED variant (resolveIfDialogGone
-> verifyStillAnswerable): the heuristic only decides when to look, the
screen decides the outcome, so the existing "working can flap" rule is
respected.

A frame that parses no options is now conclusive in two cases, and only
those, so an unreadable capture still keeps the alert: the item once
parsed options, or the frame shows Claude actively running a turn. A
modal dialog BLOCKS the turn, so the two cannot coexist - measured, a
live-dialog frame carries neither the elapsed-timer spinner nor the
"esc to interrupt" footer, which the dialog replaces with "Enter to
select". That second signal is reached by a delayed staleness pass (3s)
scheduled alongside the re-capture, which closes the late-hook case where
the prompt is answered before the hook lands: nothing ever parses, `stop`
may have gone by already, and the alert outlived reloads until the 12h
TTL. The pass is deliberately later than RECAPTURE_DELAY_MS, whose whole
reason for existing is that the hook can beat Ink to the screen.

Frontend: _onHookElicitationComplete cleared only the elicitation entry,
but an AskUserQuestion arrives as permission_prompt, so it was clearing
the wrong alert; it now clears both, matching the server's kind-agnostic
APPROVAL_RESOLVING_EVENTS.

Verified end to end on an isolated beta instance, not just in unit tests:
before, resolution could only come from the stop route (approval:resolved
always immediately preceding hook:stop); after, it arrives from the new
paths, and a simulated late hook resolves at +3.12s with no stop, no
working signal and no GET, while the pane is still working. Tests use
frames captured off a live pane and each new one was confirmed to fail
against the old behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 12:18:16 +02:00
Codeman maintainer fa7e700834 fix(terminal): report a click to the CLI only when it asked for the mouse
Found while verifying Auto Copy in a browser: a plain left click in a
claude/codex/gemini pane sent a synthetic SGR mouse report into the PTY
whether or not the program in that pane had ever enabled mouse tracking.
When the pane holds a plain shell (the CLI exited, or a shell was started
inside a session of that mode) readline prints the report as literal text
and it garbles the next line typed:

    $ [<0;88;20Mecho hello
    bash: 0: No such file or directory

The cause is that the browser could not know. The full strip
(isAltScreenStripMode) removes the mouse DECSETs from the stream, so
xterm's modes.mouseTrackingMode is permanently 'none' for those modes and
_sendSyntheticSgrTap() hand-encodes reports to stand in for xterm's own
encoder. With no state to consult it had to do that on every click.

What the strip removes, the server now remembers.
_recordStrippedMouseMode() records each sequence as it is stripped,
toState() publishes it as cliMouseTracking, and the browser's
_shouldReportMouseToCli() (renamed from _sessionUsesServerMouseStrip)
requires it at all three report sites: the desktop click, the touchend
tap, and the mobile tap classifier.

Details that are easy to get wrong:

* Only the tracking modes count (1000/1001/1002/1003). 1005/1006 select
  an encoding and 1007 is alt-scroll; a CLI that picks SGR encoding
  without turning tracking on is not asking about clicks, and counting
  those would put the stray reports straight back.
* Modes are held in a Set, so a TUI disabling a mode it never enabled
  cannot clear the ones that are really on.
* The change broadcasts immediately instead of through
  broadcastSessionStateDebounced: the flag flips when a dialog opens, and
  the user can click that dialog well inside the 500ms debounce window.
* It fails toward silence. After a server restart the flag is false until
  the CLI re-emits its DECSET, which tmux does at client attach.

Verified against a live claude 2.x session: the CLI holds a tracking mode
on continuously, so its clicks are still reported byte for byte as
before, while a bash prompt in the same stripped mode now reports
nothing and types cleanly. The flag also propagates live over SSE in both
directions, checked by toggling ?1002h/?1002l from inside the pane.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 11:21:29 +02:00
Codeman maintainer 7936a75e28 feat(terminal): Auto Copy, put a finished selection on the clipboard
App Settings > Terminal & Input > Selection & clipboard > Auto Copy
Selection (`autoCopySelection`, per-device, default OFF). With it on,
highlighting text in the terminal copies it: mouse drag, double-click
word, triple-click line, and the phone long-press selection. Ctrl+C is
untouched and still copies on demand.

Three things decide the shape of it:

* It fires at the END of a gesture, never in onSelectionChange. That
  callback runs for every cell a drag crosses, so copying there would be
  one clipboard write per mouse move. It only arms a pending flag; a
  document-level mouseup listener flushes, and the touch path calls the
  flush itself because it preventDefaults its touchend and no mouseup
  ever arrives there.
* The flush is synchronous inside the handler, because both clipboard
  paths need user activation: Firefox gates navigator.clipboard
  .writeText on it, and execCommand('copy'), the fallback the plain-HTTP
  LAN install lands on, has to run in the gesture's own task. A timer or
  a wait for onSelectionChange loses it, invisibly in Chrome.
* It deliberately does NOT do what copyTerminalSelection() does. That
  one clears the selection (so a second Ctrl+C is an interrupt) and
  focuses the terminal. Clearing would make text vanish under the cursor
  that just highlighted it, and focusing opens the on-screen keyboard
  over it on a phone. Focus is instead restored to whatever held it,
  which only matters for the execCommand fallback.

Guards are pure in decideAutoCopy() (constants.js): off, blank or
whitespace-only text, and a 1M-char cap, since a drag off the top of the
viewport autoscrolls and one gesture can sweep the whole 50k-line
scrollback. Past the cap the copy is refused rather than truncated, with
a toast pointing at Ctrl+C.

Feedback is silent on success except once per page load, so a feature
that works by doing nothing visible can still be told from a dead
toggle; failures and refusals toast, throttled to 10s.

Per-device on both counts the settings rule requires: in `displayKeys`
and absent from the .strict() SettingsUpdateSchema, because clipboard
access differs by device and by origin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 10:46:11 +02:00
Codeman maintainer 07b9c7fd7b fix(terminal): remove the unreachable copyTerminal(), closing out #322
The last two items of #322: copyTerminal() copied the entire buffer but
was wired to no button, shortcut or call site anywhere, and it wrote
through navigator.clipboard directly, which is undefined on the
plain-HTTP LAN install, so it would have failed there even if it were
reachable. Everything that actually copies goes through
copyTerminalSelection() and _copyText's execCommand fallback; whole-
buffer copy, should anyone want it, is a selectAll() away from that
same working path.

Closes #322

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-20 00:06:20 +02:00
Codeman maintainer c00e054e0e chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:43:22 +02:00
Codeman maintainer 68ae9a8c5f fix(files): match glob queries without regex so a hostile query cannot stall the server
The Files search compiled the user's query into a backtracking RegExp:
'*a*a*a...' became '^.*a.*a.*a...$', the classic blowup, evaluated
synchronously against every walked path — a pathological query could
freeze the event loop for the whole server (and every user of it in
multi-user mode). /api/search stays regex-free for exactly this reason.

Globs now match through a two-pointer wildcard walk, O(text · pattern)
worst case, with a 256-char query cap bounding the pattern side; an
overlong query compiles to null, the same answer as an empty one.
Semantics are unchanged (anchored, case-insensitive, * spans slashes)
and the existing tests pass untouched; the pathological pattern gets a
test that fails by timeout with the RegExp version.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:35:45 +02:00
Ark0N a49c30d173 Merge pull request #324 from aakhter/feat/files-panel-search
feat(files): search the Files panel by name or path
2026-08-19 23:33:08 +02:00
Codeman maintainer d871d1913f docs: restore the bullet PR #321 dropped off the xterm-zerolag-input gotcha
The new local-echo-overlay gotcha landed as a list item but left the
xterm-zerolag-input entry below it without its leading '- ', splitting
the Common Gotchas bullet list in two.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:25:09 +02:00
Ark0N ede3b05c10 Merge pull request #321 from rounakdatta/fix/mobile-link-taps
feat(mobile): links open from a tap, text can be copied, long prompts stay visible, wrapped links open whole
2026-08-19 23:23:51 +02:00
Ark0N c049de75db Merge pull request #320 from comzine/feat/custom-terminal-font
feat: Nerd Font prompt icons out of the box + configurable terminal font
2026-08-19 23:04:50 +02:00
Rounak DattaandClaude Opus 5 aae90599e5 fix(terminal): stitch a wrapped line through the indent its continuation carries
An agent's numbered list wraps its URL, and the link opened a PREFIX of it:

    1. https://github.com/users/someone/packages/container/p
       ackage/thing

opened `…/container/p`. The provider already stitched hard wraps — Ink emits a real
newline, so nothing is flagged `isWrapped` and a row that fills the last column is
taken as continuing — but it joined the row texts VERBATIM, and the continuation
carries the list's own three-space indent. That whitespace lands in the middle of
the token, which is exactly where the URL pattern stops. Flush-left wrapped URLs
(Claude Code's own `/login`) worked, which is why this survived.

The touch-selection helpers had the shallower version of the same bug: they walked
`isWrapped` only, so `Line` grabbed the single row on screen rather than the
logical line, and a long-press on a wrapped token selected only its visible half.

So the reconstruction now lives in ONE place, `terminalLogicalLine` in
constants.js, and both consumers use it — the link provider matching patterns over
its text and the selection helpers measuring words and lines with it. A link that
spans a wrap and a `Line` that stops at the screen edge were the same bug twice.

The helper drops the leading whitespace of a HARD continuation (the program's
indent) and keeps that of a SOFT one (the emulator inserts nothing, so it is real
content), records the dropped width per segment so the offset↔cell mapping stays
exact in both directions, trims only the final row so earlier offsets stay aligned
to cells, and keeps the 12-row bound that stops a screenful of full-width output
from being re-scanned on every hover.

⚠️ Selection spans are computed in CELLS, not text offsets: an xterm selection is
one contiguous run, so a token spanning a hard wrap also covers the indent cells
between its halves. A run that skipped them cannot be expressed, and would not
match what is highlighted.

Tests: `test/terminal-logical-line.test.ts` (8 cases: the indent drop, resolving
from either row, both mapping directions, soft continuations kept verbatim, no
over-reach past a short row, the row bound, final-row trimming, a missing row) and
5 in `terminal-touch-tap.test.ts` (the whole URL from either row, a token selected
across the wrap, `Line` spanning both rows, no reach into the next line). Removing
either half of the fix reds 5 and 8 of them respectively.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:41:43 +00:00
Rounak DattaandClaude Opus 5 2e58da7479 docs(mobile): document the phone gestures, and translate the selection bar
The three fixes in this branch change what a tap and a long-press MEAN on a
phone, and add a UI surface with its own z-index — all of which this repo keeps
written down rather than discoverable only by reading the handlers.

- `docs/wiki/Mobile-Guide.md` (the published user manual): a new "Tapping, links
  and copying" section, and the long-prompt behaviour in the keyboard section
  where the existing scroll/tap rules live.
- `CLAUDE.md`: the touch-gesture invariants next to the scrollback/wheel material
  (why the caret line is the boundary rather than the tap intent; why all three
  selection guards exist), the overlay's new bottom bound alongside the
  single-source note, and the selection bar in the z-index registry — 900, above
  terminal content and the local-echo overlay and deliberately below floating
  agent windows so it can never cover their controls.
- `i18n.js`: zh-CN for the bar's `Copy` / `Line` / `Clear selection`. The bar is a
  SIBLING of `.xterm`, not a descendant, so `SKIP_SELECTOR` does not cover it and
  the entries actually apply.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:19:36 +00:00
Rounak DattaandClaude Opus 5 ba843bb272 fix(mobile): keep a long prompt visible instead of hiding it behind the keyboard
Typing a prompt long enough to wrap ran the text off the bottom of the screen: the
tail — the part being typed, where the cursor is — sat behind the on-screen
keyboard, so the user was typing blind. Two independent causes.

**The overlay had no bottom bound.** On touch devices keystrokes are buffered in
the local-echo overlay and do not reach the PTY until Enter, so the CLI never
learns the prompt is long and nothing scrolls or reflows to make room. Meanwhile
the renderer lays its wrapped lines out straight DOWNWARD from the prompt row
(`top = promptRow * cellH`, each line at `i * cellH`) with nothing clamping it to
the visible rows — and with the keyboard up there are only a handful of those.

The block now grows UPWARD once it would pass the last visible row: it is lifted
so its final line lands ON that row. Every line div is opaque, so it covers
transcript above rather than vanishing under the keyboard below — the same thing a
real terminal does when a composer expands. A prompt taller than the whole
viewport keeps its TAIL, for the same reason the fix exists: the end is what the
user is looking at. `startCol` indents only the line that starts at the prompt
marker, so it is dropped along with that line when only the tail fits, and the
cursor follows the last VISIBLE line.

`rows` joins the render key: the layout depends on it, so a keyboard opening —
which changes rows without changing the text — must not be skipped as a redundant
render.

**`_shrinkPaddingToFit()` was reclaiming the bars' own space.** On phones the
toolbar and accessory bar are `position: fixed`, so they occupy no layout space
and `main`'s padding-bottom is the ONLY thing reserving room for them. Shrinking
it by the full sub-row slack pulled the terminal's bottom edge down underneath
them, and the row the following re-fit gained was painted behind them — clipping
the last line of a long prompt. The shrink now has a floor: the MEASURED height of
the currently-visible fixed bars, so genuine over-reservation of the hard-coded
84px is still reclaimed while a device that needs those pixels keeps them. The
floor is `Math.min(currentPadding, measured)`, so it can only ever prevent a
shrink, never cause a grow that would resize the terminal as a side effect.

Overlay behaviour lives in `packages/xterm-zerolag-input/` (single-source; the
vendor bundles are generated), so the fix is in the package with the row count
passed in as an optional `totalRows` — absent, the layout is exactly as before.

Tests: 7 cases in the package's `overlay-renderer.test.ts` (upward lift, tail
retention, indent drop, cursor on the last visible line, and the unclamped
fallbacks) and 7 in a new `test/mobile-keyboard-bottom-padding.test.ts` (reclaim,
floor, partial reclaim, no-grow, hidden bars, CJK strip, whole-row slack). 5 and 4
of them respectively fail without the fix. Package suite 238 pass, including the
codex byte-identity and replay tests.

Verified on Android + Chrome against a live instance: a ~460-character prompt
wrapping ~12 rows stays on screen while typing and arrives at the PTY intact.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:14:52 +00:00
Rounak DattaandClaude Opus 5 756728e553 feat(mobile): long-press to select terminal text, tap to extend, Copy
There was no way to copy terminal text from a phone at all, and three layers
ruled it out independently: `user-select: none` across the whole terminal subtree
on touch devices (taps are cursor gestures there, so the OS callout had to go),
the WebGL renderer drawing glyphs as pixels with only the accessibility tree
behind them, and xterm's own selection being a mouse DRAG while the touch path
dispatches a zero-movement mousedown/mouseup pair — a click. `copyTerminal()`
exists but is wired to no button and calls `navigator.clipboard` directly, which
is undefined on the plain-HTTP LAN install the installer offers.

So the gesture drives xterm's `select()` directly: public API, renderer-
independent, and the highlight is drawn by xterm itself. Long-press is free real
estate — tap and swipe are taken, long-press and double-tap are used by nothing.

- **Long-press** (350ms, finger still within the shared tap slop) selects the
  run of non-whitespace under the finger. Whitespace is the only delimiter on
  purpose: every punctuation-aware word rule cuts a path, URL or hash in half,
  which is what you came to copy.
- **Drag** while held extends the selection; touchmove diverts from scrolling.
- **Tap** while the bar is up extends it too. That is the ergonomic core:
  picking up a 4px handle with a fingertip is a coin flip, tapping the other end
  is not. Dismissal stays explicit (✕ or Copy), so no tap is spent leaving a mode
  the user is still using.
- **Copy** goes through the existing `copyTerminalSelection()`, so it inherits
  the execCommand fallback that is the only route that works on plain HTTP.
- **Line** takes the whole logical line, wraps included, trailing pad trimmed.

Three guards are what make the gesture survive contact with a real phone, and
each fixes a symptom measured on Android Chrome:

1. **The compat mouse pair after touchend.** xterm focuses from its screen-element
   mousedown and SelectionService resets the model there, so lifting your finger
   popped the keyboard and dissolved the selection in one go. The tap path already
   had a guard for those events; the selection path simply never armed it. Armed
   now, and the touchend is `preventDefault`ed so the synthesis is stopped at the
   source (that listener is no longer passive).
2. **The platform's own long-press.** Android Chrome runs its handling at ~500ms
   and focuses the nearest editable element — xterm's helper textarea, parked at
   the cursor — which no touch handler can preventDefault because it never sees an
   event. A focus guard blurs the terminal input for the duration of the gesture,
   whatever focused it, bounded by a self-expiring deadline so a stuck flag can
   never leave the keyboard unreachable. `contextmenu` is suppressed for the same
   window, and the threshold sits at 350ms so it lands clear of the platform's.
3. **Copy re-focusing the terminal.** `copyTerminalSelection()` ends with
   `terminal.focus()`, which is right on a desktop and wrong on a phone: the
   keyboard covers what was just copied with nothing waiting to be typed.

The bar is built in JS because index.html is read once at server start, and its
styles live in styles.css rather than mobile.css because the gesture is
touch-driven, not width-driven — a touch tablet in landscape gets the gesture and
would otherwise have no bar to copy from.

12 tests in `terminal-touch-tap.test.ts` cover the word rule, forward and
backward extension, cross-row selection, Line, tap-to-extend, the copy path, and
each of the three guards including the focus guard's expiry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:14:52 +00:00
Rounak DattaandClaude Opus 5 f2d3a7e3c1 fix(mobile): links open in a new tab from a tap, in the terminal and the chat
On a phone no link was openable, on either surface, for two unrelated reasons.

**Terminal.** xterm resolves the link under the pointer on `mousemove` and
activates it on `mouseup` over its SCREEN element. A touch tap delivers neither:
`touch-action: none` on the terminal subtree plus touchstart's preventDefault for
a 'content' tap suppress the browser's compatibility mouse events,
`_installMobileTapMouseGuard` drops the trusted ones that still arrive inside the
450ms tap window, and the synthetic mousedown/mouseup pair dispatched for mouse
REPORTING goes to the `.xterm` root — an ancestor of the node the linkifier
listens on, so it cannot reach it — and carries no mousemove either way. Every
URL and file path in the terminal was therefore inert on phones and tablets,
Claude Code's own `/login` URL included.

The tap path now activates the link itself, through the SAME provider that feeds
the hover linkifier (`_terminalLinkAtPoint`), so a tap and a desktop click can
never disagree about what is a link or where it ends — containment mirrors
xterm's own `_linkAtPosition`. It runs synchronously inside the touchend handler,
which is what keeps the user gesture that lets `window.open` past the popup
blocker, and before any mouse report, exactly as `_handleDesktopTerminalClick`
already skips the SGR tap for a hovered link.

Two kinds of row keep their existing meaning: the caret's logical line, where a
tap places the cursor and a URL the user typed must stay editable, and TUI-owned
rows, where a numbered choice or an expandable readback is answering a dialog and
routinely carries the very path the tap would otherwise open. The caret line is
the boundary rather than the tap intent, because a plain shell classifies EVERY
tap as 'input' and gating on that would leave every URL in shell output inert.

**Chat.** `marked` emits a bare `<a href>` and the markdown sanitizer's allowlist
carries no `target`, so a tap in the response viewer navigated the current tab
away: on a phone that unloads the whole dashboard — SSE, terminal buffers, unsent
composer text — and there is no middle-click or open-in-new-tab affordance to
work around it. `_renderMarkdown` now decorates anchors in the template pass it
already makes for code blocks. That pass runs AFTER sanitizing, so it is the only
source of both attributes: an agent-authored `target`/`rel` is already stripped,
and `rel="noopener noreferrer"` is set on the same element in the same breath, so
no page Codeman opens gets a `window.opener` handle back. Fragment links stay
in-page; mailto:/tel: are left to the OS rather than stranding an empty tab.

Tests: 10 cases in `terminal-touch-tap.test.ts` (URL, file path, log path,
scrollback, no-double-report, composer, shell mode, dialog row, no provider) and
a new `response-viewer-external-links.test.ts` driving the shipped marked +
DOMPurify + app.js. 7 of them fail without the fix.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-19 19:14:52 +00:00
Aamer Akhter 5cc78669bd feat(files): search the Files panel by name or path
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.

compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.

The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.

Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
2026-08-19 09:17:20 -04:00
Ark0N d4ccff07ca Merge pull request #319 from Ark0N/fix/dep-advisories
fix(deps): clear production npm advisories, fix sw.js caching regression
2026-08-19 14:54:21 +02:00
Tobias WeberandClaude Fable 5 108c00e78d feat: bundled Nerd Font symbols fallback + per-device terminal font setting
Shell prompts using Nerd Font glyphs (powerline, p10k/starship folder and
git icons) rendered as missing-glyph boxes: the built-in xterm stack has no
private-use-area symbols, and phones have no Nerd Fonts installed at all.

- Bundle Symbols Nerd Font Mono (icons-only, MIT, 1.2MB woff2) served from
  fonts/ and appended to the terminal stack before monospace — browsers fall
  back per glyph, so icons render everywhere while text stays in the text
  fonts. font-display: block + preload keep tofu out of xterm's glyph atlas.
- New per-device terminalFontFamily setting (App Settings > Terminal &
  Input > Font): prepended to the built-in stack, never a replacement, so
  the symbols fallback and final monospace always survive. Applied live on
  save (refit + echo-overlay refreshFont, mirroring setFontSize).
- Single source for both xterm surfaces: TERMINAL_FONT_DEFAULT_STACK +
  resolveTerminalFontFamily() in constants.js, unit-tested.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VjnbbZRBuvR5E3SDouwXr9
2026-08-19 00:45:19 +02:00
Codeman maintainer 8a54b331e3 fix(deps): clear production npm advisories, fix sw.js caching regression
Resolves the four advisories that reach the production dependency tree. The
other 16 npm audit reports are devDependencies-only (Remotion, Puppeteer,
postcss, the eslint/tsx toolchain) and never ship to users.

- @fastify/static 9.1.3 -> 10.1.3  GHSA-8pvw-jcv7-9cmj (authz bypass via
  non-canonical URL paths). Covers <=10.1.1, so all of 9.x is affected and
  the fix exists only on the 10.x line.
- find-my-way 9.6.0 -> 9.8.0       GHSA-c96f-x56v-gq3h (HTTP/2 DDoS)
- fast-uri 3.1.2 -> 3.1.5          GHSA-v2hh-gcrm-f6hx (host confusion)
- brace-expansion -> 5.0.9/1.1.18  GHSA-3jxr-9vmj-r5cp (expansion DoS)

The last three are transitive and needed only a lockfile re-resolve, so no
overrides were introduced.

The @fastify/static major changes setHeaders' first argument from a Node
ServerResponse to a FastifyReply. Two consequences:

1. res.setHeader() -> reply.header(). The v9 body throws TypeError from
   inside the plugin on every static request.
2. Precedence flips, silently. The callback used to write to the raw
   response and lose to the route's staged reply headers; it now writes to
   the reply and wins. That gave /sw.js a year of immutable in place of the
   no-cache, no-store its route sets, pinning a service worker on every
   client with no server-side recovery. A route that already set
   Cache-Control now keeps it.

Verified against v9 to confirm the sw.js behaviour is a regression and not
a pre-existing bug.

ws appears in npm audit but production is on 8.21.0, outside the vulnerable
range; the only affected copy is bundled under @remotion/renderer (dev-only,
and remotion is pinned at 4.0.473 because the compositor refuses to start on
a version mismatch).

Adds test/static-cache-headers.test.ts, which drives a real server and covers
a caching contract that had no test at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 23:24:22 +02:00
Codeman maintainer 09bf00c815 chore: version packages 2026-08-18 21:17:09 +02:00
Ark0N 2e19fc0430 Merge pull request #315 from aakhter/fix/respawn-stop-race
fix(respawn): do not revive a stopped controller after a cycle-step write
2026-08-18 21:15:17 +02:00
Ark0N 30f35490f6 Merge pull request #314 from aakhter/fix/symlink-safe-workspace-confinement
fix(routes): canonicalize the workspace before comparing it to a resolved path
2026-08-18 21:15:11 +02:00
Ark0N 322801052b Merge pull request #316 from Ark0N/chore/test-script-split
chore(test): make `npm test` the CI gate and give each excluded suite a runner
2026-08-18 21:15:00 +02:00
Ark0N 736a6b8b7b Merge pull request #313 from Ark0N/feat/sidebar-rich
feat(sidebar): add a rich session sidebar that carries the home screen's row detail
2026-08-18 21:14:53 +02:00
Ark0N 6525ade530 Merge pull request #317 from Ark0N/feat/wiki-tapzones-lineage-colours
Wiki publishing, phone tab tap-zone fix, and per-parent lineage colours
2026-08-18 21:14:46 +02:00
Codeman maintainer 947ff6f6fa chore(test): make npm test the CI gate and give each excluded suite a runner
`npm test` ran config/vitest.config.ts, which includes the browser, visual and
perf suites. On any machine without chromium, a free port and per-machine PNG
baselines that fails ~87 tests on a clean master, so the repo's most obvious
command could not be used as a pass/fail signal. The workaround had spread into
four docs as "never run bare `npm test`" warnings.

`npm test` now runs config/vitest.ci.config.ts — byte-for-byte what CI runs — so
local green means CI green. Verified: 264 files, 5248 tests, exit 0.

The suites it leaves out are not abandoned; each has a command:

  test:browser  5 Playwright files (chromium + a live server; codex-predictive-echo
                also needs a real codex binary)
  test:mobile   unchanged — the above plus per-machine PNG baselines
  test:perf     2 wall-clock benchmarks; need an otherwise idle machine
  test:all      the old everything-behaviour, kept reachable

test:ci is untouched (CI still calls it). test:watch and test:coverage follow
test onto the gate's config.

The more important half is the hole this closes. The exclusion list lived as
literals in one config and pointed one way only: a file excluded from CI and
added to no runner would be tested by NOTHING, silently, with every command
still green — vitest counts "no files matched a filter" as success. That is the
same shape as the #279/#280 blind spot already documented in CLAUDE.md.

So the globs moved to config/test-suites.ts, one array per REASON a suite cannot
run in CI, and all three configs derive from it. test/test-suite-partition.test.ts
then checks the arithmetic against the files on disk: it fails if any test file
is reachable by no runner, or by two. Confirmed it fires by orphaning a file and
watching it name it. The partition is exact today:

  gate 264 + browser 5 + perf 2 + mobile 9 = 280 = every *.test.ts in the repo

⚠️ One sharp edge, deliberate and documented: a file filter must match its
runner. `npm test -- test/mobile/keyboard.test.ts` now matches nothing and exits
GREEN having run zero tests, because the gate's config excludes that path.
CLAUDE.md recommended exactly that command in the on-screen-keyboard note; that
line now says `npm run test:mobile -- <file>`, and the Testing section calls out
the trap, since a green run of zero tests is worse than a red one.

Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, README.md,
README.zh-CN.md, and two ci.yml comments that claimed only test/mobile/** was
excluded — it is three suites, and 5 Playwright files rather than 3.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 19:55:23 +02:00
Aamer Akhter ce405a4cff fix(respawn): do not revive a stopped controller after a cycle-step write
Each cycle step (kickstart, update, /clear, /init) checks for `stopped` before
`await session.writeViaMux(...)`, then emits `stepSent` and calls
`setState('waiting_*')` after it.

stop() is asynchronous with respect to that await. One that lands while the
write is in flight has already passed the guard that ran, so the post-await
setState() puts a stopped controller back into a waiting state — re-arming its
step timers against a session the user asked to stop.

Re-check after the await, before emitting and setting state.

The guard reads the public `state` getter rather than `_state` on purpose:
TypeScript narrows `_state` across the await from the pre-await check and cannot
see that stop() mutated it, so `this._state === 'stopped'` is rejected as a
comparison with no overlap (TS2367) at all four sites.

Adds test/respawn-stop-race.test.ts, which drives the interleaving
deterministically by calling stop() from inside the mocked write rather than
relying on timing. All four steps go red without these guards.
2026-08-18 10:59:44 -04:00
Aamer Akhter 8e5691b05c fix(routes): canonicalize the workspace before comparing it to a resolved path
validateSessionFilePath realpath-resolves the candidate path but compared it
against the raw sessionWorkingDir. When the workspace is itself reached through
a symlink the two sides live in different namespaces, so relative() reports a
spurious `../` and every file in that workspace is judged an escape — reads and
writes in the session are refused wholesale.

That is not an exotic setup: os.tmpdir() hands back a symlinked path on macOS
(/tmp -> /private/tmp), and symlinked project directories and bind-mounted case
paths hit it too.

Resolve both sides and compare canonical to canonical. This only makes the
comparison honest — it does not widen it. The candidate keeps its own realpath,
so a symlink pointing out of the workspace and a ../ traversal are still
refused, and a workspace that cannot be resolved now fails closed.

Three stubs in file-routes.test.ts used a blanket
realpathSync.mockReturnValue(escapeTarget), which answers the same path for the
workspace and the candidate; with both sides resolved that makes an escape look
contained. They now use the input-aware mockImplementation idiom the rest of
that file already uses, so the workspace resolves to itself and only the
candidate escapes. Verified they still bite: removing the confinement check
turns all of them red.

Adds test/route-helpers-symlink-confinement.test.ts, which exercises the
function against a real symlinked workspace on disk and pins the negative cases
(../ escape, symlink-out, missing file) alongside the fix.
2026-08-18 10:46:21 -04:00
Codeman maintainer 98e37bf895 feat(sidebar): add a rich session sidebar that carries the home screen's row detail
Session List Layout gains a third option. The old "Left sidebar" becomes
"Left sidebar simple" and is unchanged down to the byte; the new "Left sidebar"
puts on each row what the desktop home rail and the phone overview already show:
when the session was first created, how long it has been in the state it is in,
and a status pill naming that state.

A docked column is not a tab strip. It has width to spare and a row per session
either way, and "name + folder" is the whole story a TAB can tell, not the whole
story there is. This is the information that was missing, and it already existed
one surface over.

Both sidebar values are the same layout, and both set data-session-list="sidebar";
the row detail rides on a separate data-sidebar-detail attribute. That split is
the load-bearing decision here: every one of the ~25 isSessionSidebarActive()
call sites and every html[data-session-list="sidebar"] rule in styles.css and
mobile.css keeps matching both variants without being touched. A third
data-session-list value would have meant auditing and editing all of them.

- Stored values: 'header', 'sidebar' (simple), 'sidebar-rich'. Anyone already on
  'sidebar' keeps exactly the layout they picked — the rename is label-only.
- State classification and the "how long has it been like this" anchor come from
  _mobileOverviewState() / _mobileOverviewSince(), not re-derived, so the three
  surfaces cannot disagree about what "working" means. A working pane repaints
  ~1/s, so its duration is measured from the turn's last Enter: a running turn
  reads "working 12m", not "0m".
- Stamps refresh in place on a 20s clock rather than by re-rendering — a rebuild
  would restart every load spinner and alert animation in the list, twice a
  minute. The clock runs only while rich rows are on screen, and is stopped from
  both render paths and from applySessionListLayout().
- The incremental render path updates the pill, the accent class and the since
  anchor; a tick alone cannot see a state change, and a new turn re-stamps
  lastSubmitAt without changing state.
- applySessionListLayout() now re-renders on a DETAIL change too. simple <-> rich
  leaves data-session-list on 'sidebar' both times, and the meta line is emitted
  by the row template rather than toggled by CSS, so the old layout-only test
  would have flipped the setting and repainted nothing.
- Width: 300px for the extra line. The collapsed 44px rail and the handheld
  drawer are both explicitly held back from it — the desktop rule is (0,3,1) and
  would otherwise out-specify mobile.css's (0,2,1) drawer base and pin a 320px
  phone's drawer to 300px.
- Missing/stale mobile-overview.js degrades to a row with no meta line rather
  than throwing and taking the whole tab strip down.

15 new tests cover the attribute split, the solo-window override, the
detail-change re-render, the row model, both render paths, the clock lifecycle
and the mobile width guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:26:34 +02:00
Codeman maintainer 5ded2ed1a3 docs: correct drifted counts in CLAUDE.md, declare postcss
The frontend load order omitted session-lineage.js (29 modules listed, 30
loaded), and several inventory counts had drifted from the tree: route handlers
~200 to ~217 with system, files and approvals each understated, src/config 20 to
21 files, install.sh 69KB to 92KB, and the Prettier exemption list, which also
never mentioned mobile.css. Two of the missing handlers are endpoints CLAUDE.md
already documents in prose but never counted.

postcss is imported by two tests but was only present transitively via vite, so
knip reported it as an unlisted dependency. Declared at the version already
resolved in the lockfile.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:44 +02:00
Codeman maintainer 76ea090a67 fix(ui): bind lineage line colours to the spawning tab
Lineage arcs were coloured per child, so one tab's own workers each got a
different colour, which is the distinction the colours exist to make. The colour
is now keyed on the parent: every arc leaving one tab is the same colour however
many workers it spawns, so the strip reads as "these five came from w1, those
two came from w2". A child that spawns in turn is a parent in its own right and
gets its own colour for the arcs below it, so a chain changes colour at each
generation while each generation's fan-out stays uniform.

The new tests drive the real _appendLineageConnectionLines() and assert the
painted custom property, because testing the colour function alone passes just
as happily with the child id passed back in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:43 +02:00
Codeman maintainer d30cac4440 fix(mobile): keep the active phone tab's centre off its action icons
The active tab is the only one that grows a gear and a close button, and with a
short session name they were eating it: "w1" rendered a 13px label while gear
plus close took 50px of a 116px tab, so the tab's geometric centre landed on the
gear and a thumb aiming at the middle of the tab opened Session Options instead
of switching sessions. Reserving a minimum label width on the active tab widens
the tab by the difference instead.

The floor is set by the 10th tab onward, which renders no number badge and so
sits 10px further right; a numbered tab clears the icons at 20px but a
numberless one needs 40px. The test recomputes that inequality from the
stylesheet rather than pinning the pixel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:43 +02:00
Codeman maintainer f3cb7696f0 docs: publish docs/wiki as the user manual, with a sync workflow
30 pages covering install, concepts, the dashboard, the agent CLIs, unattended
runs, remote and Docker cases, security and the HTTP API, plus a sidebar and a
footer. The wiki repo has no CI and no review, so docs/wiki is the source of
truth and .github/workflows/wiki-sync.yml mirrors it on every push to master.

The workflow refuses to mirror when docs/wiki is missing or holds no pages,
because it deletes before it copies and would otherwise publish the deletion of
every page. The footer carries a {{VERSION}} placeholder stamped at publish
time rather than a hand-written version, which went stale on every release.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-18 16:20:43 +02:00
Codeman maintainer 5080390e2c chore: version packages
Give the active-session handoff one owner: closeSession captures wasActive before its await and the session_deleted handler stands down for a close this tab started.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 22:54:46 +02:00
130 changed files with 13992 additions and 807 deletions
+12 -3
View File
@@ -37,11 +37,20 @@ npm run check:frontend-syntax # syntax-checks the plain-JS frontend modules
### Tests
```bash
npm test -- test/<file>.test.ts # one file (the normal way)
npm run test:ci # the full CI sweep
npm test # the gate — exactly what CI runs
npm test -- test/<file>.test.ts # one file
```
**Never run bare `npm test`.** The default config includes browser-driven Playwright suites that need a live server, Chromium, and environment-specific baselines; they will hang or fail on a normal machine. `test:ci` is the honest "run everything" command, it is exactly what CI runs.
`npm test` is the same suite CI runs, so a green run locally means a green run there. It leaves out three suites that cannot pass on an arbitrary machine, each with its own command:
```bash
npm run test:browser # Playwright + chromium (+ a live server; codex-predictive-echo needs a real codex binary)
npm run test:mobile # the above plus environment-specific PNG baselines
npm run test:perf # wall-clock benchmarks — run on an otherwise idle machine
npm run test:all # literally everything, environmental failures included
```
Expect `test:browser`/`test:mobile`/`test:perf` to fail where the machine cannot provide what they need; read that as "not runnable here", not as a regression. `config/test-suites.ts` holds the globs, and both configs derive from it, so the exclusions and those runners cannot drift apart.
If you add a test that binds a port, pick a unique one at 3150 or above (search the repo for `const PORT =` first). Never 3000.
+11 -4
View File
@@ -87,7 +87,9 @@ jobs:
fi
- name: Run unit & integration tests
# Excludes the browser-driven mobile suite (test/mobile/**); see config/vitest.ci.config.ts.
# Excludes the suites that need chromium, per-machine PNG baselines or a
# quiet machine — see config/test-suites.ts for the list and the reason
# behind each entry. Identical to what `npm test` runs locally.
# Safe in CI: TmuxManager no-ops all shell commands under VITEST (test/setup.ts).
run: npm run test:ci
@@ -99,6 +101,11 @@ jobs:
run: npx vitest run
working-directory: packages/xterm-zerolag-input
# Note: The browser-driven mobile suite (test/mobile/**) is excluded from CI —
# it needs a live server + chromium + environment-specific PNG baselines.
# Run it locally/manually. All other tests run via the `test` job above.
# Note: three suites are excluded from CI, each with its own local runner:
# npm run test:browser Playwright + chromium (+ a live server, and a real
# codex binary for codex-predictive-echo)
# npm run test:mobile the above plus environment-specific PNG baselines
# npm run test:perf wall-clock benchmarks; need an otherwise idle machine
# config/test-suites.ts holds the globs; the configs derive from it so the
# exclusions here and those runners cannot drift apart. Everything else runs in
# the `test` job above, which is the same thing `npm test` runs.
+109
View File
@@ -0,0 +1,109 @@
name: Sync Wiki
# Publishes docs/wiki/ to the repository's GitHub wiki.
#
# The wiki is a separate git repo with no CI and no review, so the source of truth
# lives in docs/wiki/ and this workflow mirrors it. Browser edits to the wiki are
# overwritten by the next sync; fix pages with a PR against docs/wiki/ instead.
#
# One-time setup: GitHub only creates <repo>.wiki.git once the first page has been
# saved in the browser. Save a stub page at /wiki/_new before the first run.
#
# Token: GITHUB_TOKEN can push to the wiki on most repos but not all. If a run fails
# with 403, add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret;
# it is preferred automatically when present. Note the 403 usually surfaces on the
# PUSH, not the clone: this repo is public, so a read-only token still clones the
# wiki fine. Both steps carry the hint.
on:
push:
branches: [master]
paths:
- 'docs/wiki/**'
- '.github/workflows/wiki-sync.yml'
workflow_dispatch:
concurrency: ${{ github.workflow }}
jobs:
sync:
name: Push docs/wiki to the wiki
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- name: Checkout repo
uses: actions/checkout@v6
- name: Clone wiki
env:
WIKI_TOKEN: ${{ secrets.WIKI_TOKEN || secrets.GITHUB_TOKEN }}
run: |
set -euo pipefail
if ! git clone "https://x-access-token:${WIKI_TOKEN}@github.com/${GITHUB_REPOSITORY}.wiki.git" wiki 2>"${RUNNER_TEMP}/clone-err.txt"; then
cat "${RUNNER_TEMP}/clone-err.txt"
echo "::error::Could not clone ${GITHUB_REPOSITORY}.wiki.git. If this says 'Repository not found', the wiki has never had a page: save one at https://github.com/${GITHUB_REPOSITORY}/wiki/_new and re-run. If it says 403, add a WIKI_TOKEN secret."
exit 1
fi
- name: Mirror pages
run: |
set -euo pipefail
# The mirror deletes before it copies, so an empty source would wipe
# every published page and the commit step would happily push that. A
# MISSING directory already fails safely (cp aborts under set -e); an
# empty one does not, so check explicitly. This is the one failure mode
# here that destroys something a browser edit cannot get back.
if [ ! -d docs/wiki ]; then
echo "::error::docs/wiki does not exist. Refusing to mirror, which would delete the entire published wiki."
exit 1
fi
pages=$(find docs/wiki -maxdepth 1 -name '*.md' | wc -l)
if [ "$pages" -eq 0 ]; then
echo "::error::docs/wiki contains no .md pages. Refusing to mirror, which would delete the entire published wiki."
exit 1
fi
echo "Mirroring ${pages} pages."
find wiki -mindepth 1 -maxdepth 1 ! -name '.git' -exec rm -rf {} +
cp -R docs/wiki/. wiki/
- name: Stamp the documented version
run: |
set -euo pipefail
# _Footer.md renders on every page and used to carry a hand-written
# version, which went stale on every release because nothing refreshed
# it. It carries {{VERSION}} instead and the series is stamped here.
series="$(node -p "require('./package.json').version.split('.').slice(0,2).join('.') + '.x'")"
# grep exits 1 when it matches nothing, which under `set -o pipefail`
# would fail the step instead of warning, so test before substituting.
if grep -rlq '{{VERSION}}' wiki/; then
grep -rlZ '{{VERSION}}' wiki/ | xargs -0 -r sed -i "s/{{VERSION}}/${series}/g"
else
echo "::warning::No {{VERSION}} placeholder found in docs/wiki. The published version line can no longer be refreshed automatically."
fi
if grep -rq '{{VERSION}}' wiki/; then
echo "::error::A {{VERSION}} placeholder survived substitution and would be published verbatim."
exit 1
fi
echo "Stamped version ${series}."
- name: Commit and push
run: |
set -euo pipefail
cd wiki
git config user.name 'github-actions[bot]'
git config user.email '41898282+github-actions[bot]@users.noreply.github.com'
git add -A
if git diff --quiet --cached; then
echo "Wiki already up to date."
exit 0
fi
git commit -m "docs: sync wiki from docs/wiki @ ${GITHUB_SHA:0:7}"
if ! git push 2>"${RUNNER_TEMP}/push-err.txt"; then
cat "${RUNNER_TEMP}/push-err.txt"
echo "::error::Could not push to ${GITHUB_REPOSITORY}.wiki.git. A 403 here means the token can read the wiki but not write it, which is the usual GITHUB_TOKEN case: add a fine-grained PAT with wiki write access as the WIKI_TOKEN secret."
exit 1
fi
+1 -1
View File
@@ -10,7 +10,7 @@ sections here.
Quick pointers:
- Type check: `tsc --noEmit` · Lint: `npm run lint` · Format: `npm run format:check`
- Targeted tests only: `npm test -- test/<file>.test.ts` (bare `npm test` is unsafe in managed sessions)
- Tests: `npm test` (the CI gate, safe to run bare) or `npm test -- test/<file>.test.ts` for one file
- Route tests use `app.inject()`; new tests needing ports must pick a unique `const PORT =`
- Branch off `master` for all work; Conventional Commit-style messages (`fix(mobile): ...`)
- Never commit secrets or local state from `~/.codeman/`
+106
View File
@@ -1,5 +1,111 @@
# aicodeman
## 1.20.1
### Patch Changes
- Terminal input and scrollback fixes (PRs #327, #331):
- IME punctuation preserved (#327): keyCode 229 / `Process` key events are now delegated to xterm's CompositionHelper instead of being suppressed, so an active Chinese IME committing numbers and full-width punctuation (,。!? and friends) reaches the terminal correctly. The CJK input field sends the browser's committed text instead of guessing from `KeyboardEvent.key`, and the redundant Android orphan-input fallback is removed so xterm is the single input owner.
- Shell history replay bounded (#331): selecting a Shell session loads a bounded 1 MiB tail instead of replaying the entire multi-megabyte tmux scrollback on xterm's main thread; full history stays available via the explicit "Load full history" action. tmux history limits now apply correctly on both legacy tmux (global default set in the same command queue before pane creation) and tmux 3.7+ (per-pane targeting that never resizes or trims unrelated live panes). Also adds `Server-Timing` and `[TERMINAL-PERF]` timing stages for terminal loads, fixes `scrollToLastNonEmptyLine` double-counting scrollback rows, and keeps live output ordered behind snapshot replays.
### Thanks
- @dignfei for both fixes: the IME punctuation root-cause fix (#327) and the bounded shell history replay with the tmux history-limit correctness work (#331).
## 1.20.0
### Minor Changes
- Response viewer for OpenCode, Gemini, Antigravity and Pi sessions (#326). External CLIs render their own TUIs, so the viewer used to come up empty for them; a new transcript parser (`response-viewer-transcript.ts`) reconstructs the conversation from the pane text instead, and the `?context=full` view now tags every block with a role so prompts render as "You" and agent output as the assistant. The divider normalizer was rewritten as a linear scan after review found catastrophic backtracking on agent-controlled input (minutes of stall on a long dash run), with an equivalence corpus pinning the old accept set.
CLIs installed via nvm or Homebrew are now found when Codeman runs as a service (#329). A shared resolver falls back to a login-shell probe when the direct PATH lookup misses, so systemd and LaunchAgent installs no longer report every CLI as missing. Review hardening on top: a failed resolution is negative-cached with doubling backoff instead of re-spawning a login shell on every request, all probes pass `killSignal: 'SIGKILL'` (interactive bash shrugs off SIGTERM, and a blocking `.bash_profile` could have hung the server indefinitely), the resolvers are inert under vitest again so test suites cannot execute binaries found on the dev box, and the improved not-found guidance is wired into both the session-create errors and the per-CLI status endpoints.
`GET /api/system/repo-status` reports branch, upstream, ahead/behind and remote reachability for git-clone installs (#328). Review hardening: the git network calls moved off the synchronous path onto a single-flight 45s cache (one slow remote could previously freeze the whole server for up to a minute per request), remote URLs and git stderr are credential-redacted before they leave the server, the spawns use the same non-interactive git env as the clone path, and a local-branch upstream no longer parses into garbage.
Auto Copy for the terminal (#325, opt-in, per-device): a finished selection (mouse drag, double or triple click, or a phone long-press) lands on the clipboard by itself, so select-then-copy becomes select. Alongside it, hand-encoded tap reports are now gated on the server-observed `cliMouseTracking` state, so a pane that has fallen back to a plain shell no longer receives `[<0;88;20M` junk on tap.
The Ralph loop no longer stops polling after two ticks (#330): the reschedule guard read a stale timer handle that the timer callback never cleared, so the loop silently died while its status stayed `running`. The handle is now nulled as the callback's first statement, and a regression test pins the bug.
The red "needs you" tab alert clears when a dialog is answered in the terminal instead of surviving until the end of the turn: the post-hook re-capture could erase the parsed dialog options that the staleness sweep relies on (`applyCapture` is now add-only for options), and a delayed staleness pass now runs while a page is open. The unreachable `copyTerminal()` was removed, closing out #322.
### Thanks
- @aakhter contributed the external-CLI response viewer (#326), the repo-status endpoint (#328), the login-shell CLI resolution (#329) and the Ralph reschedule fix (#330)
- @rounakdatta reported the mobile copy gap (#322) closed out in this release
## 1.19.7
### Patch Changes
- Mobile catches up: links open from a tap, terminal text can be selected and copied, long prompts stay visible while you type. Plus Files panel search, a bundled Nerd Font symbols fallback, and a per-device terminal font setting.
- **Terminal and chat links work on phones** (#321): tapping a URL or file path in terminal output now opens it (new tab, file preview, or log viewer), resolved through the same provider desktop hover uses, so tap and click can never disagree about what is a link. Dialog rows and the composer keep their existing meaning. Response-viewer links open in a new tab with `rel="noopener noreferrer"` instead of navigating the dashboard away. Wrapped links open whole: the logical-line reconstruction now stitches hard wraps through the indent their continuation carries, which also fixes desktop hover-click truncating wrapped URLs.
- **Terminal text can be copied on touch devices** (#321): long-press selects the token under the finger, drag or tap the other end to extend, and a small bar offers Copy, Line (the whole logical line, wraps included) and dismiss. Copy works on plain-HTTP installs too. Three guards keep the keyboard down and the selection alive through the browser's own long-press handling.
- **A long prompt stays visible on phones** (#321): the local-echo overlay grows upward once it would run past the last visible row (a prompt taller than the screen keeps its tail, where the cursor is), and the keyboard-driven padding shrink can no longer reclaim the space the fixed toolbar and accessory bar stand in.
- **Files panel search** (#324): `GET /api/sessions/:id/files?q=...` answers a flat match list (name or path substring, `*`/`?` globs), recursing past non-matching directories with its own match cap on top of the existing bounds; without `q` the response is byte-identical to before. Glob queries are matched without regex so a pathological pattern cannot stall the server.
- **Nerd Font prompt glyphs out of the box, custom terminal font** (#320): a bundled icons-only Symbols Nerd Font Mono fallback renders powerlevel10k/starship/oh-my-posh glyphs on every device with no font install, and App Settings gains a per-device terminal font family that is prepended to the built-in stack.
### Thanks
Three contributor PRs in one release: thanks to @rounakdatta (#321), @aakhter (#324) and @comzine (#320).
- 8a54b33: Clear every production-reachable npm advisory, and fix a service-worker caching regression the upgrade exposed.
`npm audit` reported 20 advisories, but 16 were devDependencies-only (Remotion, Puppeteer, postcss, the eslint/tsx toolchain) and never reached anyone installing the package. Four reached production and are now resolved:
- **`@fastify/static` 9.1.3 to 10.1.3** — GHSA-8pvw-jcv7-9cmj, authorization bypass via non-canonical URL paths. The advisory covers `<=10.1.1`, so the entire 9.x line is affected and the fix only exists on 10.x.
- **`find-my-way` 9.6.0 to 9.8.0** — GHSA-c96f-x56v-gq3h (HTTP/2 DDoS). Not exploitable here since Codeman does not enable HTTP/2, fixed anyway.
- **`fast-uri` 3.1.2 to 3.1.5** — GHSA-v2hh-gcrm-f6hx, host confusion via a literal backslash authority delimiter.
- **`brace-expansion` to 5.0.9 / 1.1.18** — GHSA-3jxr-9vmj-r5cp, exponential-time expansion DoS.
The last three were transitive and only needed a lockfile re-resolve; no `overrides` were added.
The `@fastify/static` major changes the `setHeaders` callback's first argument from a Node `ServerResponse` to a `FastifyReply`, which required two fixes:
- `res.setHeader()` became `reply.header()`. A v9-style body throws `TypeError: res.setHeader is not a function` from inside the plugin on every static request.
- **That change also flips precedence, silently.** The callback used to write to the raw response and be overwritten by the route's staged reply headers; it now writes to the reply and wins instead. That handed `/sw.js` a year of `immutable` in place of the `no-cache, no-store` its route sets, which would pin a service worker on every client with no server-side way to recover. A route that already set `Cache-Control` now keeps it.
`ws` also appears in `npm audit` but production is already on 8.21.0, outside the vulnerable range; the only affected copy is bundled under `@remotion/renderer` and is dev-only.
Adds `test/static-cache-headers.test.ts`, which drives a real server and covers the caching contract that had no test at all, and moves the floors in `test/dependency-security.test.ts` up to the patched versions.
## 1.19.6
### Patch Changes
- Wiki user manual, a phone tab tap-zone fix, per-parent lineage colours, and two robustness fixes.
- **Wiki**: `docs/wiki/` is now a 30-page user manual (installation, quick start, the dashboard, agent CLIs, remote/Docker cases, hooks, security, HTTP API, troubleshooting and more), published to the GitHub wiki by a sync workflow on every push that touches it.
- **Phone tabs**: on a narrow phone the active tab's geometric centre could land on its gear icon, so a thumb aiming at the tab opened Session Options instead of switching. The active tab's name now reserves a minimum width, and a static test recomputes the clearance from the stylesheet so widening the icons fails there rather than on a phone.
- **Lineage lines**: the arcs between a tab and the tabs it spawned are now coloured per SPAWNING tab, so every arc leaving one tab shares a colour and the strip reads as "these came from w1, those from w2". A child that spawns in turn gets its own colour, so a chain changes colour at each generation.
- **File access**: `validateSessionFilePath()` now canonicalizes the workspace as well as the candidate path before comparing them. Resolving only the candidate made a workspace reached through a symlink (`/tmp` on macOS, symlinked project dirs, bind-mounted case paths) report a spurious escape and refuse every read and write in that session. Escapes are still refused.
- **Respawn**: a cycle step that is stopped mid-write no longer revives the state machine. `stop()` could land during the `await` on the kickstart / update / clear / init write, after which the controller set itself back to a waiting state and kept running.
### Thanks
- @aakhter for the symlink-safe workspace confinement fix (#314) and the respawn stop-race fix (#315).
- 98e37bf: Session List Layout gains a third option, "Left sidebar", whose rows carry the same per-session detail the home screen shows.
The sidebar previously had one row style: a name and a folder. That is the whole story a tab can tell, but a docked column is not a tab strip — it has width to spare and a row per session either way, and the information that was missing is exactly the information the desktop home rail and the phone overview already put on screen. So the new option lifts it onto the rows: when the session was first created, how long it has been in the state it is in, and a status pill naming that state.
- The old "Left sidebar" is now **"Left sidebar simple"** and is unchanged, down to the byte — the stored value stays `sidebar`, so anyone already using it keeps exactly the layout they chose. The new option is `sidebar-rich`.
- Both sidebar values are the SAME layout and both set `data-session-list="sidebar"`; row detail rides on a separate `data-sidebar-detail` attribute. That is deliberate: every `isSessionSidebarActive()` call site and every `html[data-session-list="sidebar"]` rule in styles.css and mobile.css keeps matching both, untouched.
- Which state a session is in, and which stamp measures it, come from `_mobileOverviewState()` / `_mobileOverviewSince()` rather than being re-derived — the sidebar, the home rail and the phone overview cannot disagree about what "working" means. A working row is measured from the turn's last Enter, not from its last repaint, so a running turn reads `working 12m` instead of `0m`.
- The stamps refresh in place on a 20s clock instead of re-rendering: a rebuild would restart every load spinner and alert animation in the list, twice a minute. The clock only runs while rich rows are on screen.
- The column widens to 300px for the extra line, and the collapsed 44px rail and the handheld drawer are explicitly held back from that width.
- 947ff6f: `npm test` is now the CI gate and is safe to run bare; the suites it cannot run each got their own command.
`npm test` ran the everything-config, which fails ~87 tests on a clean master on any machine without chromium, a free port and per-machine PNG baselines. That made the repo's most obvious command useless as a pass/fail signal, and the docs had accumulated "never run bare `npm test`" warnings in four files to work around it. It now runs `config/vitest.ci.config.ts` — exactly what CI runs — so local green means CI green.
- New: `test:browser` (5 Playwright files), `test:perf` (2 wall-clock benchmarks), `test:all` (the old everything-behaviour, kept reachable). `test:ci` and `test:mobile` are unchanged; `test:watch` and `test:coverage` follow `test` onto the gate's config.
- The exclusion list moved to `config/test-suites.ts`, with the reason each suite cannot run in CI. Every config derives from it, so the gate's excludes and the runners' includes cannot drift.
- That drift was a silent hole, not a tidiness problem: a file excluded from CI and added to no runner is tested by NOTHING, and every command stays green, because vitest counts "no files matched" as success. `test/test-suite-partition.test.ts` now fails if any test file is reachable by no runner or by two.
- ⚠️ A file filter must match its runner: `npm test -- test/mobile/keyboard.test.ts` matches nothing and exits green having run zero tests, because the gate excludes that path. Use `npm run test:mobile -- <file>`. Documented in CLAUDE.md, and the one place that recommended the old form was corrected.
- Docs synced: CLAUDE.md, AGENTS.md, .github/CONTRIBUTING.md, both READMEs, and two ci.yml comments that claimed only `test/mobile/**` was excluded (it is three suites, and 5 Playwright files rather than 3).
## 1.19.5
### Patch Changes
- Closing the session you are looking at now always moves you to the next tab.
The delete request and its own `session_deleted` broadcast raced each other: the close path selected the next tab, while the broadcast handler cleared the active session and showed the home screen, and whichever ran first decided what you saw. On one build, closing a tab either switched sessions or dumped you on the welcome screen depending on timing. The close now owns that handoff from beginning to end, and the broadcast handler stays out of the way for a close started in that tab. A session deleted from somewhere else still returns you to the home screen, which is the honest answer when what you were looking at was taken away.
The next tab is also picked from sessions that still exist, so a stale entry in the tab order can no longer name a tab that is already gone.
## 1.19.4
### Patch Changes
+48 -28
View File
File diff suppressed because one or more lines are too long
+1 -1
View File
@@ -1037,7 +1037,7 @@ flowchart TB
npm install
npx tsx src/index.ts web # Dev mode
npm run build # Production build
npm run test:ci # Run tests (the CI suite; browser suites need extra setup)
npm test # Run tests (same suite CI runs; browser/mobile/perf suites have their own commands)
```
See [CLAUDE.md](./CLAUDE.md) for full documentation.
+1 -1
View File
@@ -937,7 +937,7 @@ flowchart TB
npm install
npx tsx src/index.ts web # 开发模式
npm run build # 生产构建
npm run test:ci # 运行测试(CI 套件;浏览器套件需要额外环境)
npm test # 运行测试(与 CI 相同;浏览器/移动端/性能套件另有独立命令)
```
完整文档见 [CLAUDE.md](./CLAUDE.md)。
+45
View File
@@ -0,0 +1,45 @@
/**
* The test suites that `npm test` deliberately does NOT run, in one place.
*
* Why this file exists: the exclusion list used to live only in
* config/vitest.ci.config.ts, as literals. Anything excluded there was
* therefore reachable only by running the everything-config by hand and reading
* past its failures — and a newly excluded file was reachable by nothing at
* all, silently, because nothing pointed at it. Both configs now derive their
* globs from the arrays below, so adding a suite here puts it in exactly one
* runner and takes it out of exactly one gate.
*
* Adding a new test that cannot run in CI: put its glob in the array that
* describes WHY it cannot, not in whichever one is shortest.
*/
/**
* Playwright-driven: needs chromium and, in most cases, a live Codeman server
* on a real port. Deterministic where the environment provides both, which is
* why these are a runnable suite (`npm run test:browser`) rather than skipped.
*/
export const BROWSER_TEST_GLOBS = [
'test/inline-rename.test.ts',
'test/opencode-resize.test.ts',
'test/webgl-fallback.test.ts',
'test/terminal-copy-shortcut.test.ts',
'test/codex-predictive-echo.test.ts', // also needs a real codex binary
];
/**
* Wall-clock benchmarks. They assert on durations, so a loaded shared runner
* fails them for reasons that have nothing to do with the diff under test.
*/
export const PERF_TEST_GLOBS = ['test/perf-*.test.ts'];
/**
* Browser + visual regression: chromium AND environment-specific PNG baselines
* that are generated per machine. Has its own config
* (test/mobile/vitest.config.ts) because it needs serial execution, a longer
* timeout and the `pretest:mobile` vendor step — run it with
* `npm run test:mobile`, not through the configs here.
*/
export const MOBILE_TEST_GLOBS = ['test/mobile/**'];
/** Everything `npm test` skips. */
export const NON_CI_TEST_GLOBS = [...MOBILE_TEST_GLOBS, ...PERF_TEST_GLOBS, ...BROWSER_TEST_GLOBS];
+34
View File
@@ -0,0 +1,34 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
import { BROWSER_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The Playwright-driven suite `npm test` skips — `npm run test:browser`.
*
* Needs chromium and, for most of these, a live Codeman server on a real port;
* codex-predictive-echo also needs a real codex binary. Expect failures where
* the machine cannot provide those, and read them as "not runnable here", not
* as a regression.
*
* The mobile suite is NOT here: it needs per-machine PNG baselines, serial
* execution and the `pretest:mobile` vendor step, so it keeps its own config
* (test/mobile/vitest.config.ts) behind `npm run test:mobile`.
*
* fileParallelism stays off for the same reason as every other config in this
* directory: these bind real ports and drive real tmux sessions, and two files
* doing that at once fail each other rather than the code.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: BROWSER_TEST_GLOBS,
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 60000,
teardownTimeout: 60000,
},
});
+9 -14
View File
@@ -1,13 +1,17 @@
import { resolve } from 'node:path';
import { defineConfig, configDefaults } from 'vitest/config';
import { NON_CI_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* CI test config — same as vitest.config.ts but EXCLUDES the browser-driven
* mobile suite (test/mobile/**). Those are Playwright visual-regression tests
* that need a live server + chromium + environment-specific PNG baselines, so
* they are run/maintained separately and are not part of the CI gate.
* The default gate — what `npm test` and CI both run.
*
* Same as vitest.config.ts but EXCLUDES the suites that cannot pass on an
* arbitrary machine: browser-driven (Playwright + chromium), visual-regression
* (per-machine PNG baselines) and wall-clock perf. Those are not unmaintained;
* they have their own runners (`test:browser`, `test:mobile`, `test:perf`).
* See config/test-suites.ts for the list and the reason behind each entry.
*
* Keep the rest in sync with config/vitest.config.ts.
*/
@@ -17,16 +21,7 @@ export default defineConfig({
globals: true,
environment: 'node',
include: ['test/**/*.test.ts'],
exclude: [
...configDefaults.exclude,
'test/mobile/**', // browser/visual (Playwright + chromium)
'test/perf-*.test.ts', // timing-sensitive perf benchmarks (flaky in CI)
'test/inline-rename.test.ts', // browser (Playwright)
'test/opencode-resize.test.ts', // browser (Playwright)
'test/webgl-fallback.test.ts', // browser (Playwright)
'test/terminal-copy-shortcut.test.ts', // browser (Playwright)
'test/codex-predictive-echo.test.ts', // browser (Playwright) + real codex binary
],
exclude: [...configDefaults.exclude, ...NON_CI_TEST_GLOBS],
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 30000,
+11
View File
@@ -3,6 +3,17 @@ import { defineConfig } from 'vitest/config';
const root = resolve(import.meta.dirname, '..');
/**
* EVERY test in the repo, including the ones that cannot pass on an arbitrary
* machine — `npm run test:all`. Reach for it when you want the complete picture
* and are prepared to read past environmental failures.
*
* This is NOT what `npm test` runs. On a machine without chromium, a free port
* or per-machine PNG baselines this config fails ~87 tests on a clean master,
* which makes it useless as a pass/fail signal: the default gate is
* config/vitest.ci.config.ts, and the suites it leaves out each have their own
* runner (`test:browser`, `test:perf`, `test:mobile`). See config/test-suites.ts.
*/
export default defineConfig({
test: {
root,
+25
View File
@@ -0,0 +1,25 @@
import { resolve } from 'node:path';
import { defineConfig } from 'vitest/config';
import { PERF_TEST_GLOBS } from './test-suites';
const root = resolve(import.meta.dirname, '..');
/**
* The wall-clock benchmarks `npm test` skips — `npm run test:perf`.
*
* These assert on durations, so run them on an otherwise idle machine: a loaded
* runner fails them for reasons that have nothing to do with the diff under
* test, which is exactly why they are not part of the default gate.
*/
export default defineConfig({
test: {
root,
globals: true,
environment: 'node',
include: PERF_TEST_GLOBS,
setupFiles: ['./test/setup.ts'],
fileParallelism: false,
testTimeout: 60000,
teardownTimeout: 60000,
},
});
+22 -4
View File
@@ -92,13 +92,15 @@ Implementation detail extracted from `CLAUDE.md` so that file stays small enough
### Full-scrollback replay
**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load OF EACH SESSION per page load requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other tab one frame of history); later switches keep the cheap `?tail=` visible-frame path. On top of that, scrolling up while already at the TOP of the buffer re-pulls `full=1` on demand (`_maybeRefetchFullHistory`, 4s per-session cooldown, in-flight + tab-switch guards, viewport position held across the replay). The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`.
**Full-scrollback replay** (COD-164/#148, reworked for #205): `GET /api/sessions/:id/terminal?full=1` returns the ENTIRE tmux scrollback (capture-pane `-e -S -<lines>` bounded by the configured history limit, explicit `maxBuffer` from the terminal-history config, early byte-cap before normalization, CRLF-normalized for shell panes). On success the capture is returned ALONE (`source='mux-full-history'` — it supersedes the byte buffer; no duplication). The first load of each non-shell TUI session per page requests `full=1` (`_fullHistoryLoaded` Set in app.js — the old one-shot `_initialFullBufferLoad` flag was consumed by whichever tab auto-selected, leaving every other TUI tab one frame of history). Shell sessions instead load a bounded 1 MiB `?tail=` window on every selection: a 100k-line shell capture can be tens of MiB, and automatically parsing it makes tab-switch latency scale with the entire session. Shell full history is therefore explicit-button-only; reaching the top during an ordinary wheel/touch gesture must not reset xterm and replay the multi-megabyte capture on its main thread. Other modes may still re-pull `full=1` at the TOP, and pressing **Load full history** forces the request for any recoverably truncated session (`_maybeRefetchFullHistory`, 4s per-session gesture cooldown, in-flight + tab-switch guards, viewport position held across the replay); Shell full pulls are not retained in the tab cache, so the next switch stays bounded. Chunked replay enqueues 32 KiB pieces across safe yields, appends an xterm parse marker, then releases the live-output gate; output arriving after that release stays ordered behind the snapshot, while the marker callback supplies accurate parse timing without extending the pre-existing queued-event discard window. The route exposes capture/prepare totals in `Server-Timing`, while `[TERMINAL-PERF]` separates TTFB, body/JSON, reset+parse and total time for both selection and on-demand full pulls; parse completion is not a browser compositor/GPU paint measurement. The re-pull exists because xterm's buffer is only a WINDOW onto tmux's history and two things shrink it: tmux coalesces bursty output into pane REPAINTS that overwrite rows instead of emitting linefeeds (measured: a 60-line burst added 1 row of browser scrollback and destroyed 34), and a tab switch replays only the visible frame. tmux's own history is intact throughout — the browser just has to ask for it again. On-demand rather than automatic because at a 100k history limit the capture can be megabytes. ⚠️ **The re-pull must never DOWNGRADE the buffer** (#205 round 2): the same reasoning that makes it a win for a shell pane makes it destructive for a repaint-mode CLI pane, where tmux keeps no history of its own (`history_size≈0` measured for a Claude pane) and the capture is roughly ONE frame while xterm may hold hundreds of rows of replayed frames — `_resetTerminalForReplay()` + rewrite then deletes history mid-scroll ("goes back a bit, repeats blocks, gets worse the further up I go"; measured A/B on a live pane: 341 rows → 42 with the guard off). `_replayWouldShrinkBuffer()` (terminal-ui.js) estimates the capture's rendered rows — escape sequences stripped, `capture-pane -J` re-wrapping accounted for — and the pull is skipped when that is more than one screen short of `buffer.active.length`. The one-screen tolerance matters: both sides are estimates (the buffer length counts trailing blank rows), so only a clear downgrade is refused. A refused session joins `_fullHistoryRepullUseless`, raising its cooldown from 4s to 60s so a hollow pane stops re-fetching megabytes on every scroll-up. Tests: `test/tmux-capture-full-history.test.ts`, `test/tmux-scrollback-eol.test.ts`, `test/terminal-scroll-routing.test.ts`.
### Terminal scrollback: strip flavors and wheel/touch forwarding
**Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity/pi) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend `_sessionUsesServerMouseStrip()` mirror stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`.
**Two strip flavors, one carry** (#205, `session.ts:_handleTerminalOutput`): the FULL strip (`isAltScreenStripMode` = codex/claude/gemini) removes alt-screen toggles, `3J`, and mouse-tracking DECSETs. Every other mode (shell/opencode/antigravity/pi) gets the NARROW strip (`isMuxAltScreenOnlyStripMode`) — alt-screen toggles ONLY — and only when tmux-backed (`useMux`). Rationale: the tmux CLIENT emits `smcup` as its first bytes at attach, before any program runs, parking xterm in the scrollback-less alternate buffer for the whole session (touch scrolling no-ops; xterm's own wheel handler converts the wheel to Up/Down arrows = readline history cycling — both #205 symptoms). tmux never forwards a pane program's alt-screen toggles to its client (it repaints instead; measured — vim/less inside a pane emit zero to the client), so the only thing the narrow strip ever removes is tmux's own smcup. It keeps `3J` (a user's `clear` is a deliberate scrollback wipe) and the mouse DECSETs (tmux passes those through even with `mouse off`; stripping them would break htop/vim mouse support). ⚠️ The `useMux` gate is load-bearing: `startShell()`/`startInteractive()` fall back to a DIRECT PTY when mux creation fails, and there the inner program's own `?1049h` really does reach xterm — stripping it would break vim/less/htop for real. The replay path (`session-routes.ts`, via `session.usesMux`) applies the same narrow branch; the frontend mirror (`_shouldReportMouseToCli()`) stays claude/codex/gemini because only the FULL strip touches mouse DECSETs. The chunk-boundary carry (`_altScreenSeqCarry`) runs for both flavors. Tests: `test/claude-scrollback-strip.test.ts`.
**Only claude ≥ 2.1.187 forwards the wheel; every other mode scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS stay enabled for codex (`_sessionUsesServerMouseStrip`); measured, they are no-ops that insert nothing, so click-to-position is simply unavailable there rather than harmful.
⚠️ **What the full strip removes, it must REMEMBER.** Stripping the mouse DECSETs means xterm's `modes.mouseTrackingMode` is permanently `'none'` for those modes, so the browser hand-encodes click reports to compensate (`_sendSyntheticSgrTap`). With no state to consult it did that on EVERY click, which delivered mouse reports to programs that never asked for them: the same pane runs a plain shell whenever the CLI has exited or a `shell` was started inside a claude-mode session, and a shell prints the report as literal text (`[<0;88;20M`), garbling the next line typed. `_recordStrippedMouseMode()` therefore records each stripped sequence as it goes and publishes `cliMouseTracking` through `toState()`, and `_shouldReportMouseToCli()` requires it. ⚠️ Only the TRACKING modes count (1000/1001/1002/1003): 1005/1006 select an ENCODING and 1007 is alt-scroll, and counting those would put the stray reports straight back. ⚠️ The change broadcasts IMMEDIATELY rather than through `broadcastSessionStateDebounced`, because the flag flips when a dialog opens and the user can click that dialog inside the 500ms debounce window. Measured on a live claude 2.x: the CLI holds a tracking mode on continuously (so clicks keep being reported exactly as before), while a bash prompt in the same stripped mode reports nothing. Fails toward silence: after a server restart the flag is false until the CLI re-emits, which tmux does at client attach.
**Only claude ≥ 2.1.187 forwards the wheel; every other mode scrolls local scrollback** (#227 follow-up, `terminal-ui.js:_shouldForwardWheelToApp`). Codex was in the forward list until a reporter hit a completely dead wheel in codex tabs while the scrollbar drag worked. Measured against codex-cli 0.147.0 in a bare tmux: it never enables mouse tracking (`mouse_any_flag=0`) and SGR wheel reports fed to its PTY change nothing on screen, because it runs an INLINE viewport (`alternate_on=0`) and pushes its transcript into the terminal's own scrollback (tmux `history_size` grows) instead of paging in-app. So for codex, local scrollback IS the transcript and forwarding swallowed every tick. ⚠️ "The TUI is a strip mode" is NOT evidence that it consumes wheel reports — verify with a real `\x1b[<64;c;rM` write into a live pane before adding a mode here. Hand-encoded SGR TAPS are gated by `_shouldReportMouseToCli()` (strip mode AND the server-observed `cliMouseTracking` flag, recorded by `_recordStrippedMouseMode` in session.ts as it strips): codex never enables mouse tracking, so since #325 no tap report is sent there at all — click-to-position was already a measured no-op in codex, and a pane that has fallen back to a shell no longer receives `[<0;88;20M` junk.
**Wheel/touch forwarding is NOT gated on viewport-at-bottom** (#205, `terminal-ui.js:_shouldForwardWheelToApp`): for sessions verified to scroll their own transcript on SGR wheel reports (claude ≥ 2.1.187 — version via the local/docker/remote `--version` probes), the plain wheel AND touch drags forward as coalesced SGR reports (`_forwardScrollToApp` → `_sendSyntheticSgrWheel`, 40ms batches, 5-tick cap, 512-byte queue bound). It used to gate on the viewport being at the bottom so both scrollbacks stayed reachable, but a repaint-mode CLI keeps NO terminal scrollback of its own — xterm's buffer holds only replayed repaint frames, so local scrolling drags the CLI's pinned prompt box up the screen over stale frames; and `scrollToLastNonEmptyLine()` routinely parked the viewport off-bottom, silently pinning the wheel to local. Forwarding now snaps the viewport home first (SGR coordinates address the LIVE screen — a report computed from a scrolled-up viewport would hit-test the wrong row). Local scrollback remains on Shift+wheel and the `terminalWheelLocalScrollback` opt-out (both also cover touch via the shared gate; touch has no Shift, so the setting is its only local pin). `_wheelScrollLines()` normalizes `deltaMode` (Firefox fires LINE deltas ≈3/notch — read as pixels that rounded to 0 and fell to the ±1 fallback, ~4× too slow; PAGE deltas scale by `terminal.rows`) while keeping the #154 Shift-axis trap (macOS trackpads put Shift+scroll magnitude on deltaX). Tests: `test/terminal-touch-tap.test.ts`.
@@ -258,6 +260,22 @@ Invariants:
Copy goes through `_copyText()` (Clipboard API, then hidden-textarea + `execCommand`), not raw `navigator.clipboard`, because `install.sh`'s LAN option serves plain HTTP where `navigator.clipboard` is undefined; the fallback steals focus, so the terminal is refocused afterwards. Related: xterm registers its own `copy` listener on the terminal element gated on `hasSelection()`, which is why right-click → Copy has always worked. Selection itself is unavailable on touch devices by design (`user-select: none` on the terminal subtree), and in `shell`/`opencode`/`antigravity` tabs the TUI owns the mouse, so selecting there needs Shift+drag. Tests: `test/terminal-copy-selection.test.ts` (gate + wiring invariants), `test/terminal-copy-shortcut.test.ts` (browser, real key presses).
### Auto Copy (copy-on-select)
**Auto Copy** (`autoCopySelection`, per-device, default OFF) puts a finished terminal selection on the clipboard without a keystroke. It is a thin layer over the smart-copy machinery above and shares `_copyText()` with it, but the two paths differ in every decision that matters:
- **It fires at the END of a gesture, never on selection change.** `onSelectionChange` runs for every cell a drag crosses, so copying there would be one clipboard write per mouse move. The callback only ARMS `_autoCopyPending`; the flush is a document-level `mouseup` listener installed once in `initTerminal`, plus explicit calls from the touch selection path.
- **The flush is synchronous inside the handler.** Both clipboard paths need user activation: Firefox gates `navigator.clipboard.writeText` on it, and `document.execCommand('copy')` (the plain-HTTP fallback that `install.sh`'s LAN option lands on) has to run in the gesture's own task. Deferring to a timer or waiting for `onSelectionChange` loses it, and the failure is browser-specific and invisible in Chrome.
- **The listener is on `document`, not the terminal container**, because a drag that ends outside the terminal (sweeping up past the header) delivers its mouseup to the document. Unrelated mouseups elsewhere on the page are filtered by the decision helper, not by the listener's target.
- **Touch has its own entry point.** `_endTouchSelectionGesture()` and `_selectTouchSelectionLine()` call the flush directly, because the touch path `preventDefault()`s its touchend (that is what stops the compat mouse pair from stealing the selection back), so no mouseup ever reaches the document there. Without those two calls the toggle is simply dead on a phone.
- **It must NOT do what `copyTerminalSelection()` does.** That one clears the selection (so a second `Ctrl+C` is an interrupt) and focuses the terminal. Clearing would make text vanish from under the cursor that just highlighted it, and focusing opens the on-screen keyboard over it on a phone. Focus is instead RESTORED to whatever held it before the copy, which only matters for the `execCommand` fallback (it focuses a temp textarea on the way through); the Clipboard API path never moves focus at all.
`decideAutoCopy()` (constants.js, pure) holds the guards: setting off, blank or whitespace-only text (what a drag across empty cells produces), and a `AUTO_COPY_MAX_CHARS` (1M) cap. ⚠️ The cap is not decoration: a drag off the top of the viewport autoscrolls, so one gesture can sweep the whole 50k-line scrollback. Past it the copy is REFUSED rather than truncated, with a toast pointing at `Ctrl+C`, which still copies everything through the explicit path.
⚠️ **Two dedupe rules, and both earn their place.** A genuine selection change (`pending`) always copies, so re-selecting the same text after copying something else in between still works. Otherwise only text differing from the last auto-copy does, which is what stops an unrelated mouseup from re-copying a stale selection AND what makes the first copy of a drag work at all: xterm fires `onSelectionChange` from its own document `mouseup` handler, and listener order between the two is registration order, not something this code controls. Gating on `pending` alone silently drops that first copy.
Feedback is silent on success except ONCE per page load (a feature that works by doing nothing visible cannot otherwise be told from a dead toggle); failures and refusals toast, throttled to 10s so a permanently blocked clipboard cannot paint a toast on every drag. The setting is per-device on both counts required by the settings rule: it is in `displayKeys` AND absent from the `.strict()` `SettingsUpdateSchema` (clipboard access differs by device and by origin, and the plain-HTTP LAN install has no `navigator.clipboard` at all). Tests: `test/terminal-auto-copy.test.ts`.
### Settings surface: App Settings, Session Options, Add Case
**One visual language, three modals.** `#appSettingsModal`, `#sessionOptionsModal` and `#createCaseModal` share the `set-*` surface (left rail, sections of grouped row cards, label + description on the left, control pinned right) through a single `:is(#appSettingsModal, #sessionOptionsModal, #createCaseModal)` scope in `styles.css`. An `:is()` list takes the specificity of its **most specific argument**, and all three arguments are ids, so every rule kept exactly the weight it had when the block was `#appSettingsModal`-only: nothing downstream shifted in the cascade. That property is what let the surface absorb Session Options and then Add Case in two separate commits without a cascade audit each time.
@@ -359,7 +377,7 @@ Anatomy: `.set-shell` → `.set-shell-head` (title + `.set-head-actions`) + `.se
### Buffers, uploads, and terminal history
Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 32MB (see below), text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. **Terminal history** (`src/config/terminal-history.ts`, COD-80): tmux history-limit 100k lines, PTY buffer 32MB max / 24MB trim (env `CODEMAN_MAX_TERMINAL_BUFFER`/`CODEMAN_TRIM_TERMINAL_TO`; the env-derived trim is clamped ≤75% of max — trim ≥ max would disable `BufferAccumulator` trimming entirely = unbounded memory); browser xterm scrollback stays a separate hardcoded 50k (`DEFAULT_SCROLLBACK` in constants.js — 100k/tab is a mobile-memory hazard). Settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but inert (only `tmuxHistoryLimit` is wired live); `buffer-limits.ts` re-exports the defaults. Text/message limits are env-overridable too (`CODEMAN_MAX_TEXT_OUTPUT`/`CODEMAN_TRIM_TEXT_TO`/`CODEMAN_MAX_MESSAGES`). **Image upload** (`image-input.js` / `config/buffer-limits.ts`): up to `_maxBatchImages` 20 images/batch (bounded concurrency 3), per-file `MAX_PASTE_IMAGE_BYTES` 50MB (env `CODEMAN_MAX_PASTE_IMAGE_BYTES`); the mobile camera-roll picker auto-downscales to fit before upload. **HEIC paste uploads** (#151): converted server-side to JPEG in a `worker_threads` worker (`web/heic-jpeg-worker.ts`, resourceLimits + 30s timeout) gated by `runWithConversionLimit()`; detection is magic-byte based (covers Android/MIUI HEIFs mislabeled as JPEG); headers declaring > 64MP are rejected 415 BEFORE decode (decompression-bomb guard). Deps: `heic-decode` + `jpeg-js`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`.
Target: 20 sessions, 50 agent windows at 60fps. Limits in `src/config/`: terminal 32MB (see below), text 1MB, messages 1000, max agents 500, max sessions 50, max SSE clients 100. **Terminal history** (`src/config/terminal-history.ts`, COD-80): tmux history-limit 100k lines, PTY buffer 32MB max / 24MB trim (env `CODEMAN_MAX_TERMINAL_BUFFER`/`CODEMAN_TRIM_TERMINAL_TO`; the env-derived trim is clamped ≤75% of max — trim ≥ max would disable `BufferAccumulator` trimming entirely = unbounded memory); browser xterm scrollback stays a separate hardcoded 50k (`DEFAULT_SCROLLBACK` in constants.js — 100k/tab is a mobile-memory hazard). tmux <3.7 allocates history at pane creation, so `createSession()` sets the global default in the same command queue immediately before `new-session`; tmux 3.7+ instead creates the session and targets only that pane, because changing the global option can resize and trim unrelated live panes. A settings change resizes tracked panes only on 3.7+ and otherwise affects future panes; no version can recover lines already evicted. Settings keys `terminalScrollbackLines`/`terminalBufferMaxBytes`/`terminalBufferTrimBytes` are schema-validated but inert (only `tmuxHistoryLimit` is wired); `buffer-limits.ts` re-exports the defaults. Text/message limits are env-overridable too (`CODEMAN_MAX_TEXT_OUTPUT`/`CODEMAN_TRIM_TEXT_TO`/`CODEMAN_MAX_MESSAGES`). **Image upload** (`image-input.js` / `config/buffer-limits.ts`): up to `_maxBatchImages` 20 images/batch (bounded concurrency 3), per-file `MAX_PASTE_IMAGE_BYTES` 50MB (env `CODEMAN_MAX_PASTE_IMAGE_BYTES`); the mobile camera-roll picker auto-downscales to fit before upload. **HEIC paste uploads** (#151): converted server-side to JPEG in a `worker_threads` worker (`web/heic-jpeg-worker.ts`, resourceLimits + 30s timeout) gated by `runWithConversionLimit()`; detection is magic-byte based (covers Android/MIUI HEIFs mislabeled as JPEG); headers declaring > 64MP are rejected 415 BEFORE decode (decompression-bomb guard). Deps: `heic-decode` + `jpeg-js`. Use `LRUMap` for bounded caches, `StaleExpirationMap` for TTL cleanup. Anti-flicker pipeline: `docs/terminal-anti-flicker.md`.
## Local packages and build artifacts
+207
View File
@@ -0,0 +1,207 @@
# Agent CLIs
Codeman drives seven run modes: six agent CLIs plus a plain shell. This page covers picking
one, setting it up, and the differences that actually change how you work.
## The seven modes
| Mode | CLI | Get it |
| -------------------- | ---------------------------- | ---------------------------------------------------------------------- |
| **Claude Code** | `claude` | [docs.anthropic.com](https://docs.anthropic.com/en/docs/claude-code) |
| **OpenCode** | `opencode` | [opencode.ai](https://opencode.ai) |
| **Codex** | `codex` | [developers.openai.com/codex/cli](https://developers.openai.com/codex/cli) |
| **Gemini** | `gemini` | [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) |
| **Antigravity** | `agy` | [antigravity.google](https://antigravity.google) |
| **Pi** | `pi` | [pi.dev](https://pi.dev) |
| **Terminal / Shell** | your `$SHELL` | Already installed. |
Any combination works, including all of them. The run mode is chosen per session from the
arrow beside the **Run** button, so one case can have a Claude session and a Codex session
open side by side.
## Codeman does not manage your logins
Install each CLI yourself and log it in once by hand. Codeman never collects, stores, or
refreshes your CLI credentials. It launches the binary and attaches to the result.
The one place credentials are touched is [Docker Cases](Docker-Cases), where host
credentials are copied into a container read-only at launch so you do not have to log in
again inside it. Even there, the container keeps its own copies and never writes back to
your host credential stores.
## Making a CLI visible to Codeman
Codeman resolves each binary from the environment the **server** runs in, which is not
necessarily the shell you tested in.
```bash
codeman doctor # what Codeman can actually see
codeman doctor --json
```
If a CLI is installed but a Run button for it never appears:
1. Check `which <cli>` in a plain login shell, not just your interactive one.
2. If Codeman runs as a service, remember that launchd hands a job
`/usr/bin:/bin:/usr/sbin:/sbin`. `codeman service install` bakes your PATH into the unit
precisely to avoid this; a hand-written plist or unit will not.
3. Restart the server after installing a new CLI.
`pi` is additionally version-probed rather than trusted by name, because `pi` is a generic
enough command that something else on your PATH may answer to it.
## Claude is the reference mode
A number of Codeman features exist only for Claude sessions. This is structural, not a
backlog: they depend on Claude Code's hook system, or on parsing Claude's specific terminal
output. The other CLIs expose no equivalent.
| Feature | Claude | Other CLIs |
| ------------------------------------------------ | ------ | --------------------------------------------------- |
| Sessions, tabs, scrollback, exactly-once input | Yes | Yes |
| Respawn cycling and unattended runs | Yes | Yes |
| Cron jobs | Yes | Yes |
| Docker cases, remote SSH cases | Yes | Yes |
| Precise idle detection (hook-driven) | Yes | Output-stabilization fallback, coarser |
| Auto-resume when a usage limit resets | Yes | No |
| Plan usage chip | Yes | No |
| Approvals Inbox | Yes | No |
| Read My Mind | Yes | No |
| Ralph loop and its task tracker | Yes | No |
| Subagent and team windows | Yes | No |
| Model, effort, and ultracode controls | Yes | No |
| `stop` and `blocked` wait signals | Yes | 400 if you ask for them explicitly |
| The bundled agent skill | Yes | No |
Everything that makes a session a session works everywhere. What is Claude-only is mostly
the machinery that needs to know *what* the agent is doing rather than *that* it is doing
something.
## Per-CLI notes
### Claude Code
The defaults you will care about, all under **App Settings**:
- **Model** (Models section). Written into the case's `.claude/settings.local.json` as a
soft default, so `/model` still works mid-session. The 1M-context Opus variant is a
switch on the model card rather than a separate model.
- **Effort** (`low` through `max`) or **ultracode** for dynamic multi-agent workflows. Also
a soft default: `/effort` overrides it any time. Effort is deliberately not passed as an
environment variable, because that would hard-lock it and block in-session switching.
- **Startup permission mode** (Agents & CLIs section). The default is
`--dangerously-skip-permissions`, which is why the security model matters. You can switch
new sessions to Anthropic's classifier-guarded `auto` mode, normal prompting, or an
explicit allowed-tools list.
**Separate Claude accounts per session.** Set `CLAUDE_CONFIG_DIR` in a session's environment
overrides to point it at a different Claude config directory, which is how you run one
session on a client's subscription and another on your own. One caveat: a relocated config
directory writes transcripts outside `~/.claude/projects`, which blinds the response viewer,
subagent windows, ultracode panel, and Read My Mind for that session. Symlink `projects`
back into the shared tree to keep them working:
```bash
ln -s ~/.claude/projects <configDir>/projects
```
### OpenCode
Renders its own TUI, so Codeman treats readiness as output stabilization rather than
watching for a prompt marker. Requires tmux, with no direct-PTY fallback, because its
environment is injected through socket-scoped `tmux setenv` rather than the command line.
Integration detail: [`docs/opencode-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/opencode-integration.md).
### Codex
Two behaviours that are deliberate and worth knowing:
- **Predictive echo instead of buffered echo.** Codex's composer reacts to every keystroke,
a `/` opens a live-filtering picker, arrows edit server-side state. Buffering keystrokes
until Enter starved it, so Codex paints each keystroke at the predicted cell while the
bytes on the wire stay byte-identical to what you typed.
- **The wheel is not forwarded** into its transcript. Codex ignores the mouse reports
Codeman would send, so forwarding produced a dead wheel. Scrolling in a Codex session is
local scrollback.
### Gemini
Enterprise only, since Google's June 2026 consumer cutover. Its environment allowlist
includes the broad `GOOGLE_*` namespace, deliberately, because Vertex AI authentication
needs `GOOGLE_CLOUD_PROJECT`, `GOOGLE_APPLICATION_CREDENTIALS`, and
`GOOGLE_GENAI_USE_VERTEXAI`. That is the loosest allowlist entry in Codeman and it affects
only the CLI you spawned yourself.
### Antigravity
Google's successor to the consumer Gemini CLI, invoked as `agy`. It keeps all of its state
in `~/.gemini/antigravity-cli/`, so the credential handling that applies to Gemini applies
to it as well.
### Pi
Pi needs the opposite instincts from every other CLI here.
- **It has no permission prompts and no sandbox.** There is no bypass flag to send, and
Codeman does not invent one.
- **Its privileged setting is project trust**, a three-way `--approve` / `--no-approve` /
unset. Approving trust makes Pi **execute repo-local `.pi/extensions` TypeScript**, so
point it at a repository you trust. In multi-user mode, a user without an explicit grant
gets `--no-approve` even when no configuration exists.
- **Authentication is `/login` inside the session**, or the server process's own
environment. Pi's roughly 34 provider keys (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`,
`HF_TOKEN`, and so on) share no common prefix, and the environment allowlist is global
rather than per mode, so admitting them for Pi would widen the allowlist for every mode at
once. They stay out.
Guide: [`docs/pi-integration.md`](https://github.com/Ark0N/Codeman/blob/master/docs/pi-integration.md).
### Terminal / Shell
A plain shell in a tmux session. No agent, no hooks, no idle detection.
On phones a shell session automatically swaps the keyboard accessory bar for terminal
controls: Ctrl, Esc, Tab, arrows, paste. **Ctrl is a one-shot modifier**: tap it, then tap a
letter, and the control byte is sent. It disarms on use, on a second tap, on any other
accessory key, on a session switch, and when the keyboard closes. Details in
[Mobile Guide](Mobile-Guide).
## Environment overrides
Per-session environment variables are set when creating a session and persist across
respawns. Which variables are accepted depends on the mode:
| Mode | Allowed prefixes |
| ----------- | --------------------------------- |
| Claude | `CLAUDE_CODE_*`, plus the exact key `CLAUDE_CONFIG_DIR` |
| OpenCode | `OPENCODE_*` |
| Codex | `CODEX_*` |
| Gemini | `GEMINI_*`, `GOOGLE_*` |
| Antigravity | `ANTIGRAVITY_*` |
| Pi | `PI_*` |
Anything outside the allowlist is rejected at the schema. This is intentional: the allowlist
is one global list, so widening it for one CLI widens it for all of them.
Two things that deliberately do **not** travel as environment variables: **effort**, because
an environment variable hard-locks it and blocks `/effort`, and **model**, which is written
into the case's `.claude/settings.local.json` so that `/model` keeps working.
## Choosing a mode
- **Claude Code** if you want every Codeman feature. Unattended overnight runs, usage-limit
auto-resume, the Approvals Inbox, and subagent visualization all assume it.
- **Codex, OpenCode, Gemini, Antigravity** when you prefer that agent or that model. You get
the session layer, respawn, cron, Docker, and remote SSH; you do not get the hook-driven
features.
- **Pi** if you want a fast, unsandboxed agent and you understand what project trust does.
- **Shell** for the times you want a terminal on your phone with no agent at all. It is a
genuinely useful mode, not a fallback.
## Read next
- [Core Concepts](Core-Concepts) - run modes versus location overlays.
- [Settings Reference](Settings-Reference) - model, effort, and permission-mode settings.
- [Keeping Agents Running](Keeping-Agents-Running) - what idle detection does per mode.
- [Security](Security) - what skipping permission prompts actually means.
+97
View File
@@ -0,0 +1,97 @@
# Autonomous Loops
Two features that go further than "keep the session going": the **Ralph loop**, which works
a task list to completion in one session, and the **Orchestrator**, which turns a goal into
a phased plan and drives it across agents.
Both are Claude-only, both are off by default, and neither is where to start. If what you
want is an agent that keeps working overnight, that is
[Keeping Agents Running](Keeping-Agents-Running), and it is simpler, better understood, and
what most people actually use.
## Which one, if either
| You have | Use |
| ------------------------------------------------- | ------------------------------------------------------------ |
| A session that stops too early | [Respawn](Keeping-Agents-Running) |
| A written task list to grind through | Ralph loop |
| One large goal that needs planning and checkpoints | Orchestrator |
| Work that should start at a certain time | [Cron Jobs](Cron-Jobs) |
| Several workers to fan out and supervise | [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) |
## The Ralph loop
Named after the Ralph Wiggum pattern: keep feeding the agent its own task list until the
list is empty.
The shape of it:
- The task list lives in a plan file in the case, conventionally `fix_plan.md`.
- Each cycle the agent reads the plan, works the next incomplete task, and marks progress.
- Codeman watches the file, tracks todos, and detects stalls.
- The loop ends when the agent signals completion, when the iteration cap is reached, or
when you stop it.
Start it from **Session Options → Ralph / Todo**, or from the wizard on the welcome screen.
| Setting | What it does |
| ---------------------- | ------------------------------------------------------------------------ |
| Max iterations | Hard ceiling on cycles. |
| Max todos | Cap on tracked tasks, default 500, oldest evicted first. |
| Todo expiration | Auto-expiry for stale todos, default 60 minutes. |
| Plan file | Which file holds the task list. |
A **circuit breaker** sits behind it to stop respawn thrashing: it moves from closed to
half-open to open, and is reset explicitly from the session's Ralph controls.
Honest assessment: Ralph is functional but is not where development attention goes. It
predates the respawn presets, which cover most of what people originally used it for with
less ceremony. Treat it as a specialised tool rather than the headline feature.
Full background, including the upstream pattern it is based on:
[`docs/ralph-wiggum-guide.md`](https://github.com/Ark0N/Codeman/blob/master/docs/ralph-wiggum-guide.md).
## The Orchestrator
A state machine that turns one goal into a phased plan and drives it to completion:
```
idle → planning → approval → executing → verifying → (replanning) → completed / failed
```
- **Planning** turns your goal into phases.
- **Approval** is yours. You see the plan before anything runs.
- **Executing** runs each phase, using team agents and the task queue.
- **Verifying** gates each phase before the next one starts. A failed gate can send it back
to replanning rather than forward.
Open it from the Orchestrator panel in the toolbar. State persists in `state.json`, so a
server restart does not lose an in-flight plan.
Where it differs from Ralph: Ralph is one session grinding a list, the Orchestrator
coordinates phases and agents with verification between them. It suits work that has a
natural shape ("migrate this, then update callers, then update the tests") rather than a
flat backlog.
Architecture: [`docs/orchestrator-loop-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/orchestrator-loop-architecture.md).
## Running any of this safely
Autonomous loops are the features most able to spend money and change code while you are not
looking. Some habits that pay off:
- **Run them in a case that is a git repository**, on a branch you are willing to throw
away. Being able to read the diff afterwards is the whole safety net.
- **Consider a container.** [Docker Cases](Docker-Cases) gives the agent its own filesystem
and network, and one checkbox is all it costs.
- **Set the iteration cap deliberately.** It is the ceiling on the spend.
- **Turn on notifications** so a blocked loop reaches you: see
[Notifications And Approvals](Notifications-And-Approvals).
- **Read the run summary and lifecycle log afterwards**, not just the final diff. They show
where it went sideways and recovered.
## Read next
- [Keeping Agents Running](Keeping-Agents-Running) - the simpler feature that usually fits better.
- [Watching Agents Work](Watching-Agents-Work) - seeing what a loop is doing while it runs.
- [Docker Cases](Docker-Cases) - a sandbox for unattended work.
+118
View File
@@ -0,0 +1,118 @@
# Contributing
The full guide lives in
[CONTRIBUTING.md](https://github.com/Ark0N/Codeman/blob/master/.github/CONTRIBUTING.md).
This page is the short orientation, plus how to fix a page in this wiki.
## Where things go
| You have | Send it to |
| --------------------------- | ---------------------------------------------------------------------------------------------- |
| A bug | An [issue](https://github.com/Ark0N/Codeman/issues), with OS, install method, browser, and which CLI the session was running. |
| A question or setup problem | [Discussions](https://github.com/Ark0N/Codeman/discussions). |
| An idea | [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas), where it gets voted on. |
| A small fix | Straight to a PR. |
| A bigger feature | An issue or Discussion first, then build once the design has a nod. |
| A security problem | Never a public issue. See [SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md). |
Issues usually get a response within a day, and every release credits its contributors and
bug reporters by name.
## Dev setup
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # http://localhost:3000
```
Requirements: Node 22+, tmux, and at least one agent CLI on your PATH.
The frontend is plain JavaScript with no bundler in dev: edit a `.js` or `.css` file and
reload. The exception is `index.html`, which is read once at server start, so markup changes
need a restart.
## Before you push
CI runs all of these, so running them locally saves a round trip:
```bash
npm run typecheck
npm run lint
npm run format:check
npm run check:frontend-syntax
npm test -- test/<file>.test.ts # one file, the normal way
npm run test:ci # the full CI sweep
```
**Never run bare `npm test`.** The default configuration includes browser-driven Playwright
suites that need a live server, Chromium, and environment-specific baselines; they hang or
fail on a normal machine. `test:ci` is the honest "run everything".
Tests are tmux-safe by design: under vitest the tmux layer becomes an in-memory mock, so
tests cannot touch real sessions. If you add a test that binds a port, pick a unique one at
3150 or above, and never 3000.
## Finding your way around
- Every source file opens with a `@fileoverview` block. Read it before the file; it is the
map.
- [`CLAUDE.md`](https://github.com/Ark0N/Codeman/blob/master/CLAUDE.md) at the repo root is
the densest architecture primer there is. It is written for AI coding agents, but its
invariants apply identically to humans, and most review feedback traces back to something
already written there.
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md)
holds the deep mechanisms and the history behind each rule.
## Good first contributions
- **A theme skin.** A skin is four things kept in sync, and a static test checks the sync, so
if the test passes your skin works.
- **A language.** The i18n module is dependency-free, English is canonical, and Simplified
Chinese is a complete example to copy.
- **Docs.** If you got stuck and then figured it out, the sentence that would have unstuck
you is a pull request.
- Anything labelled
[good first issue](https://github.com/Ark0N/Codeman/issues?q=is%3Aissue+is%3Aopen+label%3A%22good+first+issue%22).
Worth discussing first: new CLI backends, and real-device testing reports, especially
mobile, which always find things emulation cannot.
## PR expectations
- One change per PR. Small and focused reviews fast; a grab bag stalls.
- Target `master`.
- **Keep your branch mergeable.** A PR with conflicts silently gets no CI runs at all, which
is a GitHub quirk rather than a Codeman one. Rebase when conflicts appear.
- Include or update tests when you change behaviour.
- Do not bump versions or edit the changelog; releases are handled after merge.
- AI-assisted contributions are welcome, with one condition: understand what you are
submitting, and actually run it. "The model said it works" is not a test.
## Fixing this wiki
These pages are generated from
[`docs/wiki/`](https://github.com/Ark0N/Codeman/tree/master/docs/wiki) in the main
repository, and pushed here automatically when master changes.
**Editing a page in the browser will be overwritten by the next sync.** Send a pull request
against `docs/wiki/` instead. It is plain markdown, and a documentation PR is a genuinely
useful contribution.
Conventions for wiki pages:
- Links between pages use the wiki form: `[Remote Access](Remote-Access)`, no `.md`.
- Links into the repository are absolute `https://github.com/Ark0N/Codeman/blob/master/...`
URLs.
- Images are referenced from the main repository over raw URLs rather than being copied into
the wiki.
- Say what the default is, especially when it is off. Most of Codeman is opt-in.
- Label Claude-only behaviour every time it appears. Six of the seven run modes are not
Claude.
## Conduct
Be kind, be direct, assume good faith. Report unacceptable behaviour privately via the
contact in
[SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
+183
View File
@@ -0,0 +1,183 @@
# Core Concepts
The five ideas the rest of the manual assumes: cases, sessions, run modes, location
overlays, and tmux. Plus what actually persists, and where it lives on disk.
## Case
A **case** is a named working directory that Codeman remembers. It is the unit you pick in
the toolbar before hitting Run, and every session belongs to exactly one.
A case is not a container or a sandbox. It is a folder plus a name plus a little
Codeman-side configuration:
- Which CLI the Run button should default to.
- Per-case toggles (Agent Teams, 1M Opus context).
- Where it runs, if it is not the local filesystem: see [Location overlays](#location-overlays).
Three ways to get one, all under **+** next to the case picker:
| How | Result |
| ----------------- | ------------------------------------------------------------------------------------------------------ |
| **Create New** | A fresh `~/codeman-cases/<name>` with a scaffolded `CLAUDE.md`. |
| **Clone Repo** | A public repo cloned into `~/codeman-cases/<name>` and registered as a case. |
| **Link Existing** | An existing folder anywhere on disk, registered in place. Nothing is copied or moved. |
Linked cases keep living where they are. Deleting a case in Codeman removes the
registration, and for a linked case that is all it removes.
**Cases created from scratch are the only copy of that code.** Uninstalling Codeman does not
delete `~/codeman-cases/`, but treat that directory as real work, not scratch space.
## Session
A **session** is one CLI process running in one tmux session, streamed to your browser.
Sessions are named `w<n>-<case>`, so `w1-myproject` is the first worker in the `myproject`
case. Each has a stable id, and that id is what the API, the wait primitives, and every
event use.
Several sessions can share one case. That is the normal way to parallelize: three workers
in the same repo, three tabs, one case.
A session carries state the case does not:
- Its run mode, model, effort level, and environment overrides.
- Its respawn configuration and Ralph loop state.
- Its terminal scrollback.
- Its owner, in [Multi-User Mode](Multi-User-Mode).
## Run mode
The **run mode** is which CLI the session runs: `claude`, `opencode`, `codex`, `gemini`,
`antigravity`, `pi`, or `shell`. It is chosen at start and does not change afterwards; to
switch, start another session.
Claude is the reference mode. Six of the seven are not Claude, and a number of Codeman
features are Claude-only for structural reasons rather than missing effort: they depend on
Claude Code's hook system or on parsing its terminal output. Every such feature is labelled
Claude-only where it appears, and [Agent CLIs](Agent-CLIs) lists them in one place.
## Location overlays
Where a case runs is **separate from** which CLI it runs. There are three locations:
| Location | What happens |
| -------------- | ------------------------------------------------------------------------------------------------------------ |
| **Local** | The default. tmux and the CLI run on the Codeman host. |
| **Docker** | One long-lived container per case; sessions `docker exec` into it. See [Docker Cases](Docker-Cases). |
| **Remote SSH** | A durable tmux server on the remote host, fronted by a local pane running `ssh`. See [Remote SSH Sessions](Remote-SSH-Sessions). |
This matters because it is a common source of confusion: Docker is **not** an eighth run
mode. All seven run modes work in all three locations. A case is docker-backed or
ssh-backed; a session is claude or codex or shell.
**Web tabs** are the other thing that is not a session. A saved dashboard URL renders as a
tab beside your agents, but there is no PTY, no tmux, and no respawn behind it. See
[Web Tabs](Web-Tabs).
## Why tmux
tmux is a hard requirement, and it is the reason Codeman behaves the way it does.
The agent runs inside a tmux session. Codeman attaches to it, the same way your terminal
would. That indirection buys:
- **Survival.** The agent outlives your browser tab, your network, your laptop lid, and a
restart of the Codeman server itself.
- **Real scrollback.** History is held by tmux, so reconnecting replays what happened while
you were gone instead of starting from blank.
- **Attach from anywhere else.** The same session is reachable from a terminal over SSH
with the `sc` chooser, or plain `tmux -L codeman attach`.
- **Secrets off the command line.** Environment overrides are injected with socket-scoped
`tmux setenv` rather than being visible in the spawn command.
The socket is `tmux -L codeman`, separate from your personal tmux server, so Codeman
sessions never appear in a bare `tmux ls`.
## What persists
| Survives | Does not survive |
| -------------------------------------------- | --------------------------------------------------- |
| Closing the browser | `tmux -L codeman kill-server` |
| Losing the network | A machine reboot (tmux dies with it) |
| Restarting the Codeman server | Killing the session from the UI |
| `codeman web --stop` | |
| A dropped SSH link, for remote cases | |
| A container restart, for docker cases | |
Conversation history is a separate question: Claude transcripts live in `~/.claude/`, so a
conversation can be resumed even after the tmux session is gone. That is what the welcome
screen's **Resume Conversation** list offers.
## State on disk
Everything Codeman knows lives under `~/.codeman/`:
| File | Holds |
| ---------------------------------------- | -------------------------------------------------------------------- |
| `state.json` | Sessions, settings, respawn config, orchestrator state, cron jobs. |
| `settings.json` | User preferences that sync across your devices. |
| `mux-sessions.json` | tmux recovery data. |
| `session-lifecycle.jsonl` | Append-only audit log of session starts, exits, and kills. |
| `linked-cases.json` | Registered cases. |
| `remote-hosts.json`, `docker-hosts.json` | Location overlay configuration. |
| `webviews.json` | Saved dashboard URLs. |
| `users.json` | Multi-user accounts, mode 0600. |
| `push-*.json` | Web push keys and subscriptions. |
| `certs/` | Self-signed TLS for `--https`. |
None of it needs root, none of it leaves the machine, and deleting `~/.codeman/` resets
Codeman to a fresh install without touching your code.
## Instances
The data directory and the tmux socket are both **process wide**. Two Codeman servers
started on one machine share them, which means the second one discovers the first one's
live sessions and attaches to them, resizing and mutating sessions you did not expect it to
touch.
To run two on purpose, give each its own instance name:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
That scopes the data directory and the tmux socket together, which is the only safe way to
do it. `CODEMAN_DATA_DIR` and `CODEMAN_TMUX_SOCKET` can be set individually if you need
them apart, but setting only one of the two reproduces exactly the problem you were trying
to avoid.
## Hooks
For Claude sessions, Codeman writes a hooks configuration into the case so Claude Code can
report events back: a permission prompt appeared, the turn finished, the agent went idle, a
task completed. Those events drive tab alerts, the Approvals Inbox, notifications, and the
wait primitives.
This is why some features are Claude-only. The other CLIs have no equivalent hook system,
so for them Codeman falls back to watching terminal output, which is coarser: it can see
that something happened, not what it was.
See [Hooks And Integrations](Hooks-And-Integrations).
## Vocabulary
| Term | Means |
| --------------- | ---------------------------------------------------------------------------- |
| **Case** | Named working directory. |
| **Session** | One CLI in one tmux session. |
| **Run mode** | Which CLI: claude, opencode, codex, gemini, antigravity, pi, shell. |
| **Respawn** | Restarting the CLI on idle to keep an unattended run going. |
| **Ralph loop** | An autonomous single-session task loop. |
| **Orchestrator**| A phased plan driven across multiple agents. |
| **Subagent** | An agent the CLI spawned itself, shown live in its own window. |
| **Web tab** | A saved dashboard URL rendered as a tab. Not a session. |
| **Instance** | One Codeman server with its own data directory and tmux socket. |
## Read next
- [The Dashboard](The-Dashboard) - what the UI is showing you.
- [Agent CLIs](Agent-CLIs) - the seven run modes in detail.
- [Keeping Agents Running](Keeping-Agents-Running) - respawn, idle detection, usage limits.
- [`docs/architecture-invariants.md`](https://github.com/Ark0N/Codeman/blob/master/docs/architecture-invariants.md) - the mechanisms behind all of this, for contributors.
+161
View File
@@ -0,0 +1,161 @@
# Cron Jobs
Saved, named jobs that start a session and send it a prompt on a schedule. Cron for agent
sessions: *every weekday at 03:00, open a Claude session in `~/proj` and tell it to update
dependencies and open a PR.*
The ⏰ **Cron** header button is opt-in. Turn it on in
**App Settings → Header & Panels**.
## Creating a job
1. Click **⏰ Cron**, then **+ New Job**.
2. Give it a name, pick the agent type and working directory.
3. Write the prompt, or point at a file containing it.
4. Choose a schedule and leave **Enabled** on.
5. **Save**. The job appears with its computed next run.
**Run Now** fires it immediately without touching the schedule, which is the fastest way to
find out whether the prompt does what you meant.
## The fields
| Field | Notes |
| ------------------------ | ------------------------------------------------------------------------------------------- |
| **Name** | Also used as the created session's name. |
| **Agent type** | Any run mode, including `shell`. |
| **Working directory** | Validated when you save **and** again when the job fires. Blocked system trees are refused. |
| **Launch command** | Shell jobs only. Sent as the first line once the shell is up, before the prompt. |
| **Prompt** | Inline text, or a path to a file read at fire time. |
| **Input mode** | `typed` behaves like a human typing. `paste` writes directly. |
| **Schedule** | `once`, `interval`, `daily`, or `weekly`. |
| **Enabled** | Disabled jobs never fire on their own. **Run Now** still works. |
| **Concurrency policy** | What to do if sessions of the same type are already running. |
| **Auto-close previous** | Recurring jobs only. Closes the session the previous run created. Default on. |
| **Notes** | Free text for you. |
## Schedules
All wall-clock times are in the **server's local timezone**, not your browser's. A job set
for 03:00 fires at 03:00 where the server is.
| Type | Behaviour |
| ---------- | -------------------------------------------------------------------------------------------------- |
| `once` | Fires at an absolute time, then disables itself. A job missed because the server was down still fires once on the next tick. |
| `interval` | Every N minutes, from 1 minute to a year. |
| `daily` | At `HH:MM` every day. If today's time has passed, the next run is tomorrow. |
| `weekly` | At `HH:MM` on the weekdays you pick. |
Interval jobs re-anchor to when they actually fired, not to an ideal cadence, so a slow tick
or a server restart shifts later runs slightly. That drift is accepted rather than corrected.
## Prompts are single line
This is the rule people trip over. Programmatic input into an agent session is single line
everywhere in Codeman, because the terminal UIs these CLIs use treat a newline as submit. A
multi-line prompt would be silently mangled, so it is **rejected** instead: the form refuses
it, and a prompt file whose contents are multi-line fails the run with a clear message.
For anything longer than a sentence, put the instructions in a file and make the prompt tell
the agent to read it:
```
read TASKS.md and work through it
```
That is also easier to edit than a job field.
### Prompt files
Reading the prompt from a file at fire time is useful when the instructions change more
often than the schedule. The path is confined to the job's working directory, symlinks are
resolved before the check, sensitive trees are refused, and the file has to be a regular
file under 1 MiB.
If any of that fails, the run is recorded as failed and **no session is created**.
## Concurrency
Applies to scheduled runs only, never to **Run Now**:
| Policy | Behaviour |
| ------------------------------- | -------------------------------------------------------------------------------------- |
| `warn_only` | Always launch. The count of live same-type sessions is shown but does not block. |
| `skip_if_same_agent_running` | Skip this fire if another live session of that mode exists. |
The skip policy has the details you would want it to have:
- Only **live** sessions block. A tab whose CLI already exited does not count.
- Sessions the job created on its own previous runs never block it, otherwise a recurring
job would deadlock on itself after the first fire.
- A skipped `once` job is not consumed. It stays armed and fires when the blocker goes away.
- Consecutive skips are collapsed into one record per streak, so a perpetually skipped job
cannot bloat your state file.
## Run history
Every fire is recorded per job, with a status:
| Status | Meaning |
| --------- | -------------------------------------------------------------------- |
| `created` | The run started and a session was created. |
| `skipped` | The concurrency policy blocked it. Not counted as a run. |
| `failed` | The prompt could not be resolved, or the working directory was gone. |
The schedule is advanced **before** the session launches, so a slow start cannot cause the
same job to re-trigger.
## Cron versus the other autonomy features
| Want | Use |
| --------------------------------------------------- | ------------------------------------------------------- |
| Start work at a specific time | Cron |
| Keep an existing session working | [Keeping Agents Running](Keeping-Agents-Running) |
| Drive one goal to completion across phases | [Autonomous Loops](Autonomous-Loops) |
There is also an older, deliberately separate `ScheduledRun` concept behind
`/api/scheduled`: a run-now, duration-bounded loop with no recurrence and no saved jobs. The
two systems never interact, and Cron is the one you want.
## From the API
```bash
API=http://localhost:3000
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{
"name": "nightly-deps",
"agentType": "claude",
"workingDir": "/home/me/proj",
"promptMode": "inline_text",
"promptText": "Update dependencies and open a PR",
"inputMode": "typed",
"scheduleType": "daily",
"dailyTime": "03:00",
"enabled": true,
"concurrencyPolicy": "warn_only"
}' | jq
curl -s "$API/api/cron/jobs" | jq
curl -s -X POST "$API/api/cron/jobs/<jobId>/run" | jq
curl -s "$API/api/cron/jobs/<jobId>/runs" | jq
```
Add `-u admin:"$CODEMAN_PASSWORD"` when a password is set, and `-k` with the `https://` URL
on an HTTPS install.
## Gotchas
- **Times are the server's, not yours.** Obvious until you are travelling.
- **A `pi` job starts slowly.** The readiness poll looks for markers pi does not print, so it
burns its poll budget before sending the prompt. The job still works.
- **A deleted working directory fails the run**, by design, rather than creating a session
somewhere unexpected.
- **Auto-close only touches sessions this job created.** Your own tabs are never closed.
## Read next
- [Keeping Agents Running](Keeping-Agents-Running) - continuing work rather than starting it.
- [Notifications And Approvals](Notifications-And-Approvals) - hearing about a job that got stuck.
- [`docs/cron-guide.md`](https://github.com/Ark0N/Codeman/blob/master/docs/cron-guide.md) - the complete reference, including the API and SSE events.
+177
View File
@@ -0,0 +1,177 @@
# Docker Cases
Run a case inside its own container instead of directly on your host: for isolation, for a
reproducible toolchain, and for the ability to pick the whole environment up and move it to
another machine.
A docker case is a **location overlay**, not a run mode. All seven run modes work inside a
container. See [Core Concepts](Core-Concepts).
## One-time setup: the base image
The container needs an image carrying the agent toolchain (node, the CLIs, git, tmux). It
builds itself on first use with progress streamed to the UI, or you can build it ahead of
time:
```bash
node scripts/build-agent-image.mjs --no-cache
```
**Always pass `--no-cache`.** The CLIs are installed in a single `npm install -g` layer, so
a plain rebuild reuses that layer from the cache and the CLIs stay frozen at whatever
versions the image was *first* built with. This has shipped a broken CLI while reporting a
successful build.
A zero exit code proves the layers ran, not that the toolchain works. Verify:
```bash
docker run --rm codeman/agent:base bash -lc \
'for c in claude codex gemini opencode agy pi; do printf "%-9s " $c; $c --version 2>&1 | head -1; done'
```
The image is secret-free. Credentials are delivered at runtime, never baked in, so exports
never leak them. A full image lands around 1.6GB.
Prerequisite: Docker or Podman with a reachable daemon.
## The quick way
On **Add Case → Create New**, tick **🐳 Run in an isolated Docker container**. That alone is
enough: Codeman creates the case folder, spins up a hardened container with sensible
defaults, and starts the session inside it.
Expanding **Container settings** offers a template:
| Template | Memory | CPUs | GPUs |
| ----------------- | ------ | ---- | ------------------------------------- |
| Small | 2 GB | 1 | none |
| Medium (default) | 4 GB | 2 | none |
| Large | 8 GB | 4 | none |
| GPU | 8 GB | 4 | all (needs the NVIDIA container toolkit) |
Disk is elastic: storage grows as data arrives, bounded only by host disk. Changing any
setting creates a dedicated host profile for that case, so it never mutates the shared
default.
## The full way
**Add Case → Docker** exposes everything:
| Field | Meaning |
| -------------------- | ----------------------------------------------------------------------------------------------- |
| **Case name** | As usual. |
| **Workspace path** | A real host directory, bind-mounted into the container at the **same absolute path**. |
| **Host ID** | A reusable profile (image, network, resources). Share one across cases to share settings. |
| **Network** | `bridge` (internet on, default), `none` (fully isolated), or a custom bridge. |
| **Advanced** | Memory and CPU caps, host credential seeding, and whether to resume the last conversation on relaunch. |
The same-absolute-path bind mount is what keeps the File Viewer, attachments, and watchers
operating on real host bytes rather than a copy.
## One container per case
Exactly one long-lived container per case, shared by every session in it.
- Killing one session kills only that session's in-container tmux. Siblings keep running and
the container stays up.
- Reconnecting after a Codeman restart lands back in the same live agent.
- A container stop or a host reboot restarts the container and **resumes the last
conversation** from the bind-mounted transcript.
- Deleting the case removes the container. The workspace on the host survives.
## Credentials
Your existing host logins work inside the container without logging in again. Credentials
are **seeded**: mounted read-only and copied in once at launch, so in-container CLIs never
write refreshed tokens back to your host credential stores. Onboarding and trust prompts are
pre-answered so no wizard appears.
Turn seeding **off** for a sealed sandbox: no host credentials, and with `network: none`, no
outbound access either. That is the profile for genuinely untrusted work; you log in inside
the container instead.
Bind mounts are excluded from image capture, so exports stay secret-free.
One consequence worth knowing: Pi's credentials are seeded per file rather than as a whole
directory, because that directory also holds sessions, extensions, and installed packages,
which can be gigabytes. So in-container Pi sessions are invisible from the host, and `pi -c`
inside a docker case sees only that container's history.
## Isolation
Every container runs hardened by default:
- `--cap-drop ALL`
- `--security-opt no-new-privileges`
- Non-root, running as your host uid so workspace files stay host-owned
- PID limit, memory cap with swap pinned to it, `--init`
- **Never** `--privileged`, and **never** the docker socket
Rootless engines without cgroup-v2 systemd delegation cannot enforce resource caps; linking
such a host warns that the caps are advisory.
## Configuration drift is refused, not ignored
Editing a docker host's configuration (image, memory, network) after a container exists is
detected on the next launch by comparing a configuration hash against the container's label.
A mismatch **refuses the launch** and offers to recreate rather than silently running with
stale configuration.
Recreating is refused while sessions of that case are live. The workspace and the
conversation both survive it.
## Moving a case to another machine
**Export**, from the Docker tab:
| Option | Contents |
| -------------------------- | ------------------------------------------------------------------------- |
| **Full image + workspace** | The whole toolchain, installed packages, and files, in one `.tgz`. |
| **Workspace only** | Just the project files. Fast and small. |
The container is paused across the capture so image and workspace are consistent, free space
is checked first, and the intermediate image is cleaned up. Exports run in the background
and notify you when the bundle is ready.
**Import** on the other machine: copy the `.tgz` into `~/.codeman/docker-exports/` and
import it into a new case. The manifest and per-member checksums are verified, the workspace
tar is extracted with a traversal guard, and the image is loaded under a **quarantined tag**
so it can never overwrite a local image. The destination supplies its own credentials, so
nothing secret crosses machines.
## Hooks need to reach the server
In-container hooks (permission events, idle and stop notifications) call back to Codeman
over the docker bridge gateway. If Codeman binds **loopback only**, which is the default and
the production configuration, the container cannot reach it and **in-container hooks do not
fire**.
The session still works fully: idle detection falls back to output-based detection through
the exec PTY, and with permission prompts skipped there is nothing to forward anyway.
To enable them:
```bash
CODEMAN_DOCKER_BRIDGE_HOOKS=1
```
Codeman then starts a second listener bound to the docker bridge gateway that serves **only**
the hook endpoints and rejects everything else with a 403. The bridge is host-internal, so
this does not widen your network exposure. Add it to the service unit and restart.
## Limits
- Per-session environment overrides, effort, and per-CLI configuration are **rejected** for
docker cases, because they do not cross into the container. Configure the container through
the docker host's per-mode command override instead.
- tmux must exist in the base image. It is a hard prerequisite and is probed when linking a
host.
- On macOS, Docker Desktop takes a dedicated uid path, and memory caps are subject to the
VM's own ceiling.
## Read next
- [Core Concepts](Core-Concepts) - why this is an overlay rather than a run mode.
- [Security](Security) - where containers fit in the model.
- [Remote SSH Sessions](Remote-SSH-Sessions) - the other overlay.
- [`docs/docker-cases.md`](https://github.com/Ark0N/Codeman/blob/master/docs/docker-cases.md) - the full reference.
+188
View File
@@ -0,0 +1,188 @@
# Driving Codeman From An Agent
Everything the dashboard does is HTTP, so an agent can do it too. This page is for the case
that makes Codeman interesting: **Claude Code running inside a Codeman session, spawning and
supervising other sessions.**
Two routes. Start with the skill.
## The agent skill
A Claude Code skill that teaches the agent the whole API, so you ask in plain English
instead of pasting endpoint documentation into prompts.
### Install it
| How | Command | Scope |
| ------------ | ----------------------------------------------------------- | ----------------------------------------------------------- |
| Skills CLI | `npx skills add Ark0N/Codeman --skill codeman -g` | Global, any skills-aware agent. |
| Bundled CLI | `codeman skill install` | Global, at `~/.claude/skills/codeman`. |
| Bundled CLI | `codeman skill install --case <name>` | One case. |
| Web UI | **App Settings → Agents & CLIs → Claude → Agent Skill** | Injects into each case when a Claude session is created. Off by default. |
`codeman skill uninstall [--case <name>]` reverses the CLI installs, and never touches a
`skills/codeman` you wrote yourself.
### Then just ask
| You say | What happens |
| --------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------ |
| "What sessions are running right now?" | Lists them with name, mode, and status. Read-only. |
| "Start a shell worker on the `myapp` case, run the test suite, tell me if it passes." | Spawns, waits on a completion marker, reads the exit code, cleans up. |
| "Spin up 3 workers for lint, typecheck and tests, run them in parallel, report failures." | One session per task, all started first, then gathered as each finishes. |
| "Have a claude worker summarize `src/session.ts`, then close it." | Spawns, runs the readiness ladder, sends and waits, reads the answer, deletes the session. |
| "Watch session w4 and tell me if it gets stuck on a permission prompt." | Blocks on the `blocked` signal and surfaces the question to **you**. |
Sessions the agent creates get deleted when it is done. You can watch the tabs appear and
disappear in the dashboard while it works.
### What it will and will not do
- **It self-gates.** Outside a Codeman session it refuses to act and does not guess an API
URL, so a global install costs an unrelated Claude Code session nothing.
- **Unprompted, it may only** spawn sessions, prompt them, and delete ones **it created in
that conversation, by exact id**, behind a guard that refuses to delete the agent's own
session.
- **It will not** answer another session's permission prompt on your behalf. It surfaces the
question instead.
- **Deleting a case** (which erases a real directory of your code), bulk kills, respawn,
Ralph, cron, orchestrator, and settings writes all require you to ask, naming the target.
Turning the setting back off **does not remove already-injected copies**, because a
create-time sweep would yank the skill out from under other live sessions sharing that
directory. Remove them per case with `codeman skill uninstall --case <name>`.
The skill ships with the verb index always loaded, plus on-demand references for the verbs,
worked multi-worker recipes, endpoint tables, and cross-session messaging.
## The manual path
The same operations as raw HTTP, for a CI bot, a shell script, or an agent without skill
support.
### Detect that you are inside Codeman
These are set in every managed session. Read them rather than hardcoding anything:
| Variable | Meaning |
| -------------------------- | ------------------------------------------------------------------------ |
| `CODEMAN_MUX=1` | You are in a managed tmux session. Never `tmux kill-session`, `pkill claude`, or `pkill tmux`: you will kill yourself or a sibling. |
| `CODEMAN_API_URL` | Base URL, with the correct scheme. |
| `CODEMAN_SESSION_ID` | Your own session id. Use it to avoid acting on yourself. |
| `CODEMAN_HOOK_SECRET_FILE` | Path to the hook secret. |
### Rules of the road
Read these before writing any code. Each one has cost somebody an afternoon.
1. **Input is single line and must end with `\r`.** Enter fires only when the payload
contains a carriage return. Without it the text sits unsubmitted on the prompt, the
request still succeeds, and a combined wait burns its full timeout on a turn that never
started. Embedded newlines are stripped rather than rejected, so `"echo A\necho B\r"` runs
the joined `echo Aecho B`. One line per call.
2. **Make input idempotent.** Send a stable `clientId` and a monotonic per-session `seq`. The
server deduplicates, so a retry after a dropped connection cannot double-deliver.
3. **Auth.** With `CODEMAN_PASSWORD` set, use HTTP Basic or the session cookie. A missing
`Origin` is allowed, so plain curl works. A `401` replies with the bare string
`Unauthorized`, **not** the JSON envelope, so piping it into `jq` throws a parse error
instead of showing the failure. Check the status before parsing.
4. **Envelope.** Most endpoints return `{ "success": true, "data": ... }`. A few legacy GETs
return bare bodies, so handle both: `body.data ?? body`.
5. **Wait instead of polling, and a timeout is not an error.** The wait endpoints answer
`200` with `wait.timedOut: true`. Loop over short waits rather than one long call, because
tunnels cut idle connections.
6. **Only `claude` sessions emit `stop` and `blocked`.** They come from Claude Code hooks.
Shell and the external CLIs accept only `idle`, `working`, and `exit`; asking for `stop`
explicitly there is a `400`, while omitting `until` is always safe. On a shell session
`idle` fires **once at startup and never again**, so synchronize hook-less sessions with an
output marker instead.
7. **Nothing reports "ready", so wait for it explicitly.** A new session answers
`{"signal":"exit","immediate":true}` until its PID exists, and that means *not started*,
not *crashed*. A Claude worker in a fresh case then sits on the CLI's trust dialog; prompt
it there and the wait resolves on idle in about two seconds looking exactly like a finished
turn, while your text sits stuck in the dialog.
### Recipes
```bash
API="${CODEMAN_API_URL:-http://localhost:3000}"
# Add -u admin:"$CODEMAN_PASSWORD" if a password is set, and -k on an HTTPS install.
# What is running
curl -s "$API/api/sessions" | jq '.data[] | {id, name, mode, status}'
# Spawn a worker in a case
curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"myapp","mode":"shell"}' | jq
# Send a prompt (note the \r)
curl -s -X POST "$API/api/sessions/$ID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"run the tests\r","clientId":"my-agent","seq":1}' | jq
# Send and block until the turn finishes (registers the wait BEFORE writing)
curl -s -X POST "$API/api/sessions/$ID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"summarize src/session.ts\r","wait":["stop"],"waitTimeout":120000}' | jq
# Or wait for a marker in the output, which works on shell sessions too
curl -s "$API/api/sessions/$ID/wait-output?contains=DONE_17909&from=buffer" | jq
# Read the terminal back
curl -s "$API/api/sessions/$ID/terminal?tail=4000" | jq -r '.data.output'
# Clean up, by exact id
curl -s -X DELETE "$API/api/sessions/$ID" | jq
```
Use `POST /api/quick-start` rather than `POST /api/sessions` when a case might be remote:
the plain create endpoint validates the working directory locally and has no case concept.
### The split-marker trick
For hook-less sessions, synchronize on a marker in the output. The catch: your own
keystrokes echo into the output stream, so an unsplit marker matches **before the command
has run**.
Split it so the typed line never contains the string you are waiting for:
```bash
M=DONE; R=17909
# typed: echo ${M}_${R} → output contains DONE_17909, the typed line does not
```
Make it unique per call, because tmux repaints replay old screen text.
### Reading output
Use `terminal?tail=`, not `/output`. The latter's text field is empty for every tmux-backed
session, which is every interactive session. `tail` counts **bytes**, and what comes back is
terminal data with ANSI sequences included.
## Fan-out, and why it needs care
Wait signals are **edge triggered with no history**. A signal that fires with no waiter
registered is unobservable afterwards.
So a fan-out must register its waits before or as it dispatches: use send-and-wait per
worker, or latched output markers. Dispatching all the workers and then waiting on them one
at a time loses the signals of everyone who finished early.
Send-and-wait registers the waiter **before** the write for the same reason. A separate POST
followed by a wait races, and reports the previous turn's state.
## Lineage
A create request can name the session that spawned it, through a body field or a header, and
the dashboard then draws a lineage arc from parent to child. The skill sets it automatically.
It is resolved rather than trusted: an unresolvable parent is dropped silently rather than
failing the spawn, because a cosmetic field must never break a worker.
## Read next
- [HTTP API](HTTP-API) - the endpoint map and the envelope.
- [Hooks And Integrations](Hooks-And-Integrations) - events flowing the other way.
- [Watching Agents Work](Watching-Agents-Work) - seeing the fan-out in the UI.
- [`skills/codeman/SKILL.md`](https://github.com/Ark0N/Codeman/blob/master/skills/codeman/SKILL.md) - the skill itself.
+250
View File
@@ -0,0 +1,250 @@
# FAQ
The questions that keep arriving in
[Discussions](https://github.com/Ark0N/Codeman/discussions) and issues. For "why is it
doing that", go to [Troubleshooting](Troubleshooting) instead.
## The basics
### What is Codeman, in one sentence?
A self-hosted dashboard that runs AI coding agents in persistent tmux sessions on your own
machine and lets you drive them from any browser, including a phone.
### Is it free? What is the licence?
MIT, free, and open source. There is no paid tier and no account.
### Do I need an API key?
No. Codeman drives agent CLIs you have already installed and logged in yourself. Whatever
subscription or key that CLI uses is what pays for the tokens. Codeman never collects,
stores, or refreshes your credentials.
### Does Codeman send my code or prompts anywhere?
No. There is no telemetry, no analytics, and no phone-home. The only network traffic
Codeman itself makes is between your browser and your server.
Your agent CLI is a separate matter: Claude Code talks to Anthropic, Codex talks to OpenAI,
and so on. That traffic is the CLI's, on your own account, exactly as it would be in a
terminal.
Two features do send data outward, both off by default and both stated where they appear:
voice dictation through your own Claude login, and the Read My Mind prediction call.
### Does it work on Windows?
Through WSL2. Codeman requires tmux. Install it inside WSL, run your agent CLI inside WSL,
and `http://localhost:3000` works from your Windows browser. Work in the Linux filesystem
rather than `/mnt/c/...`, which is dramatically slower for file watching and git.
### Is there a mobile app?
The web UI is built for phones and installs as a PWA. There is no App Store or Play Store
app.
## Sessions and persistence
### Do my agents keep running when I close the browser?
Yes. Agents run in tmux on the server, not in your browser. Close the tab, close the laptop,
lose the network. When you come back, the session is still there with its scrollback.
The same holds when the Codeman server itself restarts. What does end a session is killing
the tmux server or rebooting the machine.
### What happens after a reboot?
tmux dies with the machine, so the sessions are gone. Conversations are not: Claude
transcripts persist on disk, and the welcome screen's **Resume Conversation** list picks
them back up. Install Codeman as a service and the server itself comes back on boot.
### How many sessions can I run at once?
The design target is 20 sessions and 50 agent windows at 60fps. The hard cap is higher, and
what you will actually hit first is the CPU and memory of the machine running the agents.
### Can I run Claude Code and Codex side by side?
Yes, that is a normal setup. The run mode is per session, so one case can have a Claude tab,
a Codex tab, and a shell tab open at the same time, each with its own colour. Some Codeman
features are Claude-only; [Agent CLIs](Agent-CLIs) lists exactly which.
### Can I attach to a session from a terminal instead of the browser?
Yes. `sc` is an interactive chooser (`sc 2` attaches directly, `sc -l` lists), or use tmux
directly on the `codeman` socket. Detach with `Ctrl+A D`.
## Running unattended
### I hit my Claude usage limit overnight. Can Codeman resume automatically?
Yes, and it is the reason the feature exists. Turn on auto-resume at the top of the Respawn
tab for that session. When Claude halts on a subscription limit, Codeman parses the reset
time from the message, waits until two minutes past it, and continues the conversation.
Respawn cycles are blocked while a session is limit-paused, which is what stops a `/clear`
from wiping the conversation you are waiting to resume. Claude-only.
### Will it keep prompting my agent forever?
Only if you configure it to. Respawn cycling is per session and off unless you turn it on,
and it has presets ranging from a 60 minute solo session to an 8 hour overnight run. There
are circuit breakers to stop a thrashing session from spinning indefinitely. See
[Keeping Agents Running](Keeping-Agents-Running).
### Does an idle session cost tokens?
No. An idle agent is a process waiting for input. Tokens are spent when a turn runs, so what
costs money is the re-prompting you configured, not the session sitting there.
### Can I schedule work for a specific time?
Yes. [Cron Jobs](Cron-Jobs) saves named jobs on a `once`, `interval`, `daily`, or `weekly`
schedule; each spins up a session and sends a prompt when due, with per-job run history.
## Access
### How do I reach Codeman from my phone when I am away from home?
Tailscale is the recommended answer: your devices join a private network, Codeman keeps its
loopback bind, and you get real HTTPS. The installer sets it up, and `install.sh tailscale`
retrofits it onto an existing install.
A Cloudflare tunnel gives a public URL faster, and requires `CODEMAN_PASSWORD`. Full
comparison in [Remote Access](Remote-Access).
### Why can't other devices reach Codeman?
Because the default bind is `127.0.0.1`, on purpose. Codeman starts agents with permission
prompts skipped, so whoever reaches the dashboard can run code on your machine. Exposing it
is a deliberate step, and [Remote Access](Remote-Access) covers the safe ways.
### My reverse proxy domain is rejected with `403 host not allowed`
The always-on Host-header allowlist blocks DNS rebinding, and it does not know your domain.
Add it:
```bash
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
A leading dot matches subdomains. Also make sure the proxy forwards WebSocket upgrades.
### Do I have to type a password on my phone?
No. Scan the QR code shown on the desktop dashboard. Tokens are single use and rotate every
60 seconds. The password remains the fallback.
## Multiple people, multiple instances
### Can several people share one Codeman?
Yes, with `codeman web --multiuser`. Each person gets a login and their own case space, and
sessions, cases, search, and events are scoped to their owner.
Be clear about what that is: it separates **workspaces**, not operating system accounts.
Every session still runs as the same OS user, so a determined user's agent can reach another
user's files. For real isolation, pair users with Docker cases or run separate instances
under separate OS accounts. See [Multi-User Mode](Multi-User-Mode).
### How do I run a second instance, a beta beside my main one?
Give it its own instance name, which scopes the data directory and the tmux socket together:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
Do not skip this. The data directory and tmux socket are process wide, so a second server on
the defaults discovers and attaches your live sessions.
## Updating and maintenance
### What is the right way to update Codeman?
| Install route | Update with |
| ------------- | --------------------------------------------------------------------------- |
| Installer | Re-run the install one-liner, or **App Settings → System → Updates**. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart the service. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It stashes a
dirty tree rather than discarding it, and streams progress across the restart. npm installs
report as non-updatable.
### Will updating kill my running sessions?
No. Sessions live in tmux, so restarting the server reattaches to them.
### Where is my data?
Everything under `~/.codeman/`, with cases created from scratch in `~/codeman-cases/`.
Nothing needs root and nothing leaves the machine. Uninstalling does not delete either
directory.
## Features
### What is the difference between respawn, Ralph, and the orchestrator?
- **Respawn** restarts a session's CLI when it goes idle, to keep a long run going. It is the
one most people want.
- **Ralph loop** is an autonomous single-session task loop with its own tracker.
- **Orchestrator** turns one goal into a phased plan and drives it across agents.
[Keeping Agents Running](Keeping-Agents-Running) and [Autonomous Loops](Autonomous-Loops)
cover them properly.
### Can agents start and supervise other agents?
Yes. Codeman ships an agent skill that lets an agent inside a session drive the HTTP API:
list sessions, spawn workers, send prompts, and block until a worker's turn finishes. It is
off by default and enabled per case.
See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
### Can I run a case in a container?
Yes. One container per case, shared by all its sessions, non-root and capability-dropped by
default, with your host CLI logins seeded in so nothing asks you to log in again. You can
export a container plus its workspace and move it to another machine. See
[Docker Cases](Docker-Cases).
### Can the agent run on a different machine?
Yes. Point a case at a remote host over SSH and the agent runs there, inside a durable
remote tmux, so a dropped connection does not kill the run. See
[Remote SSH Sessions](Remote-SSH-Sessions).
### Can I put my Grafana or other dashboards in here?
Yes. Saved URLs render as tabs beside your sessions, proxied through Codeman's own origin so
that mixed content and frame-blocking headers do not break them. See [Web Tabs](Web-Tabs).
### Why is a feature I read about not on screen?
Most of Codeman's UI is opt-in and defaults to off, so a stock install stays small. Check
**App Settings → Header & Panels**. [Settings Reference](Settings-Reference) lists the
defaults.
## Contributing
### How do I request a feature?
Open an [Idea](https://github.com/Ark0N/Codeman/discussions/categories/ideas) and it gets
voted on. Roadmap decisions happen there.
### How do I contribute code?
[CONTRIBUTING.md](https://github.com/Ark0N/Codeman/blob/master/.github/CONTRIBUTING.md) has
the full map. Small fixes can go straight to a PR; anything larger starts as an issue or
Discussion so the design gets a nod first. Skins, translations, and docs are good first
contributions.
### How do I fix a mistake in this wiki?
These pages are generated from
[`docs/wiki/`](https://github.com/Ark0N/Codeman/tree/master/docs/wiki) in the main
repository. Editing a page in the browser gets overwritten on the next sync, so send a PR
against that directory instead.
+164
View File
@@ -0,0 +1,164 @@
# HTTP API
Codeman's HTTP and SSE API is a **stable contract**. Everything the dashboard does goes
through it, so anything the dashboard can do, a script can do.
This page is the orientation. The complete specification, including every wait semantic and
the SSE catalogue, is
[`docs/api-reference.md`](https://github.com/Ark0N/Codeman/blob/master/docs/api-reference.md).
## What is stable
Covered by semantic versioning: endpoint paths under `/api/v1`, the response envelope,
`errorCode` values, and SSE event names.
Not covered, and free to change in a patch release: on-disk state files, internal modules,
and anything marked experimental. The full statement is in
[Versioning](Versioning).
`/api/v1/*` is a versioned alias of `/api/*`. Prefer the versioned form in anything you
intend to keep.
## The envelope
```json
{ "success": true, "data": { } }
```
```json
{ "success": false, "error": "human readable", "errorCode": "NOT_FOUND" }
```
A few legacy GET handlers return bare bodies rather than the envelope, so a robust client
reads `body.data ?? body`.
Branch on `errorCode`, which is stable. The HTTP status is reliable too:
| `errorCode` | HTTP | Meaning |
| ------------------ | ---- | ------------------------------------------------ |
| `INVALID_INPUT` | 400 | Malformed request or failed validation. |
| `UNAUTHORIZED` | 401 | Authentication required or failed. |
| `NOT_FOUND` | 404 | No such resource. |
| `SESSION_BUSY` | 409 | The session is busy. |
| `CONFLICT` | 409 | Conflicts with current state. |
| `ALREADY_EXISTS` | 409 | Resource already exists. |
| `OPERATION_FAILED` | 422 | Well formed, could not be completed. |
| `RATE_LIMITED` | 429 | Too many requests. |
| `INTERNAL_ERROR` | 500 | Unexpected server error. |
New error codes are non-breaking. Removing or renaming one is a major change.
**A `401` is the bare string `Unauthorized`, not the envelope.** Piping it into `jq` throws
a parse error rather than showing the failure, so check the status first.
## Authentication
With no password set, and the default loopback bind, there is none. With `CODEMAN_PASSWORD`
set, use HTTP Basic or the session cookie:
```bash
curl -s -u admin:"$CODEMAN_PASSWORD" "$API/api/sessions"
```
A **missing** `Origin` header is allowed, so curl and CLI tools work unchanged. A
present-but-foreign origin is rejected by the CSRF guard. On an HTTPS install with the
self-signed certificate, add `-k`.
## Endpoint map
Roughly 200 handlers across 24 route modules. By domain:
| Domain | Handlers | Covers |
| ------------------- | -------- | --------------------------------------------------- |
| System | 45 | Status, settings, search, digest, updates. |
| Sessions | 34 | Create, input, terminal, wait, kill. |
| Cases | 29 | Create, link, clone, remote and docker cases. |
| Files | 16 | Preview, edit, raw, attachments, path picker. |
| Orchestrator | 10 | Plans and phases. |
| Ralph | 9 | Loop control and configuration. |
| Cron | 9 | Jobs and run history. |
| Admin | 8 | Multi-user administration. |
| Plan | 8 | Plan orchestration. |
| Respawn | 7 | Respawn configuration and presets. |
| Webviews | 6 | Saved dashboards, plus the proxy. |
| Mux | 5 | tmux operations. |
| Push | 4 | Web push subscriptions. |
| Read My Mind | 4 | Intent profiles and prediction. |
| Scheduled | 4 | The legacy scheduled-run concept. |
| Approvals | 3 | The inbox and answering. |
| Teams, me, search, hooks, clipboard, telemetry, voice, ws | 1-2 each | |
Each route module documents its own endpoints in its file header.
## Long-polling instead of polling
Three calls block until something happens, so an agent driving Codeman from a shell can wait
rather than spin:
| Call | Blocks until |
| ----------------------------------------- | -------------------------------------------------------- |
| `GET /api/v1/sessions/:id/wait` | One of a set of lifecycle signals fires. |
| `GET /api/v1/sessions/:id/wait-output` | A literal string appears in the session's output. |
| `POST /api/v1/sessions/:id/input` + `wait`| The input is delivered **and then** a signal fires. |
Three semantics that break callers who assume otherwise:
1. **A timeout is `200`, not an error.** It answers with `wait.timedOut: true`. Loop over
short waits; a single long call gets cut by tunnels and proxies.
2. **Send-and-wait is not a POST followed by a wait.** It registers the waiter *before*
writing, which closes the window where a separate wait sees the session still idle from
the previous turn and answers instantly about the wrong turn.
3. **Signals are edge triggered with no history.** One that fires with no waiter registered
is unobservable afterwards. Fan-outs must register their waits as they dispatch.
`wait-output` matches a **literal substring, never a regex.** That is deliberate: no regex
means no catastrophic backtracking on attacker-influenced output.
Only `claude` sessions emit `stop` and `blocked`, because those come from Claude Code hooks.
Shell and external CLI sessions accept `idle`, `working`, and `exit`.
## SSE
`GET /api/events` is the live event stream. 155 event names, kept in sync between server and
client with a test that fails on drift.
The heartbeat is a **named** `sse:heartbeat` event rather than an SSE comment, because
comments are invisible to `EventSource` by specification and a client could not observe
them. That is what lets the browser detect a stream that has silently stopped delivering.
```js
const es = new EventSource('/api/events');
es.addEventListener('session:created', (e) => console.log(JSON.parse(e.data)));
```
## Quick examples
```bash
API="${CODEMAN_API_URL:-http://localhost:3000}"
curl -s "$API/api/status" | jq # whole-system snapshot
curl -s "$API/api/sessions" | jq '.data[].name' # live sessions
curl -s "$API/api/sessions/unified" | jq # live + historical, deduped
curl -s "$API/api/subagents" | jq # background agents
curl -s "$API/api/search?q=deploy" | jq # cross-session search
```
## Limits
| Limit | Default |
| --------------------- | ------------------------------------------ |
| Max sessions | 50 |
| Max agent windows | 500 |
| Max SSE clients | 100 |
| Terminal buffer | 32 MB per session |
| Text payload | 1 MB |
| Wait timeout ceiling | 600 s, and the response tells you what was applied |
Most are environment-overridable. See `src/config/`.
## Read next
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the practical version, with recipes.
- [Hooks And Integrations](Hooks-And-Integrations) - events flowing back into Codeman.
- [Versioning](Versioning) - what the version number promises.
- [`docs/api-reference.md`](https://github.com/Ark0N/Codeman/blob/master/docs/api-reference.md) - the full specification.
+136
View File
@@ -0,0 +1,136 @@
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/codeman-title.svg" alt="Codeman" height="56">
</p>
<h3 align="center">Mission control for AI coding agents</h3>
Codeman runs your coding agents on your own machine and puts them behind one dashboard you
can open from any device. It spawns Claude Code, OpenCode, Codex, Antigravity, Gemini, or
Pi inside persistent tmux sessions, streams the real terminal to the browser, and keeps
working while you are away from the keyboard: it re-prompts idle agents, resumes when a
subscription limit resets, runs jobs on a schedule, and shows every background subagent
live.
This wiki is the manual. The [README](https://github.com/Ark0N/Codeman) is the overview,
and the deep internals live in
[`docs/`](https://github.com/Ark0N/Codeman/tree/master/docs).
```bash
curl -fsSL https://getcodeman.com/install | bash
codeman web # then open http://localhost:3000
```
---
## Start here
**New to Codeman**
1. [Installation](Installation) - requirements, the installer, npm and git clone routes, updating.
2. [Quick Start](Quick-Start) - from a running server to a working agent in five minutes.
3. [Core Concepts](Core-Concepts) - cases, sessions, run modes, and what survives a restart.
4. [The Dashboard](The-Dashboard) - reading the tab strip, the status dots, and the alerts.
**Already running it**
- [Agent CLIs](Agent-CLIs) - the seven run modes, their setup, and which features are Claude-only.
- [Mobile Guide](Mobile-Guide) - phone and tablet use, QR login, the touch keyboard bar.
- [Remote Access](Remote-Access) - Tailscale, Cloudflare tunnel, LAN plus password, QR login.
- [Keeping Agents Running](Keeping-Agents-Running) - idle detection, respawn cycling, auto-resume on usage limits.
- [Troubleshooting](Troubleshooting) - symptom-first index of things that actually break.
**Driving it from code**
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the bundled skill, worker sessions, wait primitives.
- [HTTP API](HTTP-API) - the envelope, auth, the endpoint map, SSE events.
- [Hooks And Integrations](Hooks-And-Integrations) - events flowing back into Codeman.
---
## Everything in the manual
### Getting started
| Page | What it answers |
| ------------------------------- | --------------------------------------------------- |
| [Installation](Installation) | How do I install it, update it, and remove it? |
| [Quick Start](Quick-Start) | How do I get one agent working right now? |
| [Core Concepts](Core-Concepts) | What is a case, a session, a run mode? |
### Using it
| Page | What it answers |
| ------------------------------------------ | ---------------------------------------------------------- |
| [The Dashboard](The-Dashboard) | What is the UI telling me? |
| [Agent CLIs](Agent-CLIs) | Which agent should this session run, and how do I set it up? |
| [Working With Files](Working-With-Files) | How do I read, edit, and attach files? |
| [Input And Voice](Input-And-Voice) | How do I talk to an agent, including by voice? |
| [Mobile Guide](Mobile-Guide) | How well does this work on a phone? |
| [Keyboard Shortcuts](Keyboard-Shortcuts) | What can I drive from the keyboard? |
| [Settings Reference](Settings-Reference) | What does this setting do, and why did it not follow me to my phone? |
### Keeping agents running
| Page | What it answers |
| ------------------------------------------------------------- | --------------------------------------------------- |
| [Keeping Agents Running](Keeping-Agents-Running) | How does it run unattended overnight? |
| [Notifications And Approvals](Notifications-And-Approvals) | How do I know an agent needs me, and answer from my phone? |
| [Cron Jobs](Cron-Jobs) | How do I run an agent on a schedule? |
| [Autonomous Loops](Autonomous-Loops) | What are the Ralph and Orchestrator loops for? |
| [Watching Agents Work](Watching-Agents-Work) | How do I see what the subagents are doing? |
### Where it runs
| Page | What it answers |
| --------------------------------------------- | -------------------------------------------- |
| [Docker Cases](Docker-Cases) | How do I sandbox a project in a container? |
| [Remote SSH Sessions](Remote-SSH-Sessions) | How do I run the agent on another machine? |
| [Web Tabs](Web-Tabs) | Can my Grafana live in here too? |
| [Multi-User Mode](Multi-User-Mode) | Can several people share one Codeman? |
### Access and security
| Page | What it answers |
| ------------------------------- | ---------------------------------------------------------- |
| [Remote Access](Remote-Access) | How do I reach it from outside this machine, safely? |
| [Security](Security) | What is exposed, what protects it, what do I have to do? |
### Automation and integration
| Page | What it answers |
| ----------------------------------------------------------------- | -------------------------------------------- |
| [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) | How does an agent spawn and drive workers? |
| [HTTP API](HTTP-API) | What can I call, and what comes back? |
| [Hooks And Integrations](Hooks-And-Integrations) | How do I wire Codeman into something else? |
### Operating it
| Page | What it answers |
| --------------------------------------------- | -------------------------------------------------- |
| [Running As A Service](Running-As-A-Service) | How do I keep it up across reboots, and update it? |
| [Troubleshooting](Troubleshooting) | Why is it doing that? |
| [FAQ](FAQ) | The questions that keep coming up. |
| [Contributing](Contributing) | How do I send a fix? |
| [Versioning](Versioning) | What does the version number promise? |
---
## Requirements at a glance
| Thing | Needed |
| ------------ | --------------------------------------------------------------------------- |
| OS | macOS or Linux. Windows works through WSL2. |
| Node.js | 22 or newer. |
| tmux | Required. Sessions live in tmux, which is what makes them survive restarts. |
| An agent CLI | At least one of Claude Code, OpenCode, Codex, Gemini, Antigravity, Pi. Plain shell sessions need none. |
| Network | Binds to `127.0.0.1` by default. Reaching it from another device is a deliberate step: see [Remote Access](Remote-Access). |
Codeman is MIT licensed, self-hosted, and sends no telemetry. Everything runs on your
machine.
## Getting help
- **Questions and setup help**: [Discussions](https://github.com/Ark0N/Codeman/discussions), especially [Q&A](https://github.com/Ark0N/Codeman/discussions/categories/q-a).
- **Bugs**: [Issues](https://github.com/Ark0N/Codeman/issues). Include your OS, install method, browser, and which CLI the session was running.
- **Ideas and roadmap**: [Ideas](https://github.com/Ark0N/Codeman/discussions/categories/ideas).
- **Security**: never a public issue. See [SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
+97
View File
@@ -0,0 +1,97 @@
# Hooks and Integrations
Events flowing **back** into Codeman, and the four seams a third party can build against.
## Hooks
Claude Code can run a command when something happens in a session. Codeman writes a hooks
configuration into each Claude case so those events post back to it, which is what turns a
terminal into something that can notify you.
| Event | Fires when | Drives |
| ---------------------- | ----------------------------------------------- | --------------------------------------------- |
| `permission_prompt` | The agent asks for permission. | Red tab alert, Approvals Inbox, push. |
| `idle_prompt` | The agent is waiting for input. | Yellow tab alert, the `idle` wait signal. |
| `stop` | A turn ends. | The `stop` wait signal, idle detection. |
| `elicitation_dialog` | A dialog opens. | Approvals Inbox. |
| `elicitation_complete` | The dialog closes. | Clearing the alert. |
| `elicitation_response` | The dialog is answered. | Clearing the alert. |
| `teammate_idle` | An agent-team member goes idle. | Team surfaces. |
| `task_completed` | A task finishes. | Task tracking, run summary. |
This is why several Codeman features are Claude-only. The other CLIs have no hook system, so
for them Codeman watches terminal output, which reveals that something happened but not what
it was.
### How hooks get installed
Codeman writes them into the case when a Claude session is created. Hook blocks are
**marker-owned**: Codeman only ever updates a block it wrote, and never touches
configuration you added yourself.
If tab alerts and approvals never fire in a particular case, that case is missing its hook
block. Recreating the case rewrites it.
### The hook secret
`/api/hook-event` and `/api/status-telemetry` skip HTTP Basic authentication, because they
are called from localhost by the CLI itself. When authentication is on, that bypass
additionally requires a per-instance hook secret, because Codeman cannot tell a genuine
loopback call from a request arriving through your own loopback reverse proxy.
The secret lives in the data directory, and its path is exported into every managed session.
### Two things that break hooks
- **HTTPS.** Hook callbacks must accept the self-signed certificate. Recent versions
self-heal existing cases; older cases need recreating.
- **Docker cases on a loopback bind.** A container cannot reach `127.0.0.1` on the host, so
in-container hooks silently do not fire. Set `CODEMAN_DOCKER_BRIDGE_HOOKS=1` to open a
hooks-only listener on the bridge gateway. See [Docker Cases](Docker-Cases).
## Integration seams
Codeman has **no plugin runtime**, and that is a decision rather than a gap. A plugin runtime
means running third-party code inside a process that spawns agents with your credentials, on
a server people routinely expose over a tunnel. Codeman's security posture is one of its
reasons to exist, so it does not trade that away for an extension mechanism.
What exists instead is four documented seams.
### 1. Web tabs
Anything with a web UI can live inside Codeman as a tab, proxied through Codeman's own
origin. The lowest-effort integration by a wide margin: if your tool has a dashboard, it can
sit beside the agents with no code at all. See [Web Tabs](Web-Tabs).
### 2. SSE events
`GET /api/events` streams everything Codeman knows: session lifecycle, output, agent
activity, approvals, cron runs. 155 named events, stable under semantic versioning.
This is the seam for anything that reacts. A bot that pings your chat channel when an agent
needs a human is a short script over this stream.
### 3. HTTP API and CLI
Everything the dashboard does. Create sessions, send input, block on wait primitives, read
terminals, manage cron. See [HTTP API](HTTP-API) and
[Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
### 4. Hooks
The seam above, in the other direction: your own hook commands can run alongside Codeman's
in a case, as long as you leave Codeman's marker-owned block alone.
## Publishing an integration
There is no registry to submit to. Share it in
[Show and tell](https://github.com/Ark0N/Codeman/discussions/300), and if it needs a change
in Codeman to work properly, open an issue or a Discussion first.
## Read next
- [HTTP API](HTTP-API) - the endpoint map and envelope.
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the agent-facing path.
- [`docs/extending-codeman.md`](https://github.com/Ark0N/Codeman/blob/master/docs/extending-codeman.md) - the seams in full, with examples.
- [`docs/claude-code-hooks-reference.md`](https://github.com/Ark0N/Codeman/blob/master/docs/claude-code-hooks-reference.md) - upstream hook semantics.
+135
View File
@@ -0,0 +1,135 @@
# Input and Voice
Getting words into an agent: typing, dictating, and letting Codeman guess. Plus the input
machinery that only shows up when it goes wrong.
## Typing
Click into the terminal and type. It is a real terminal, so everything the CLI supports
works, slash commands included.
| Key | Effect |
| ---------------------------- | --------------------------------------------- |
| `Enter` | Send. |
| `Shift+Enter` / `Ctrl+Enter` | Newline without sending. |
| `Ctrl+C` | Copy if text is selected, otherwise interrupt. |
| `Ctrl+Shift+C` | Copy, never interrupts. |
| `Ctrl+L` | Clear the terminal. |
### Exactly-once delivery
Browser input goes through a durable layer rather than a plain socket write. Each prompt
carries a stable client id and a per-session sequence number, held in local storage until
the server acknowledges it.
The result is the property you want on a phone: a connection that drops mid-prompt never
loses the prompt and never delivers it twice. Two browser tabs on the same session coexist,
and only a reconnect from the *same* tab supersedes the old connection.
## Zero-lag local echo
On touch devices, keystrokes are painted in the terminal immediately and sent when you press
Enter, instead of waiting for each character to round-trip to the server and back. Over a
mobile connection that is the difference between usable and not.
![Zero-lag input](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/zerolag-demo-20260728.gif)
The consequence to remember: **text on screen has not necessarily reached the agent yet.**
It is flushed on Enter. If a prompt appears to have been ignored, press Enter, or the phone
toolbar's **Enter** button.
Default on for touch devices, off for desktop, and switchable in
**App Settings → Terminal & Input**.
### Codex is different on purpose
Codex's composer reacts to every keystroke: `/` opens a live-filtering picker, arrows edit
state on its side, the composer grows as text wraps. Buffering until Enter starved it, so
Codex sessions use **predictive echo** instead: each keystroke is painted at its predicted
position while the bytes actually sent stay identical to what you typed. Predictions
reconcile against the real buffer and only apply while the cursor is on the composer row.
## CJK input
Chinese, Japanese, and Korean input needs an IME, and an IME needs a real text field.
Turning on CJK input in **App Settings → Terminal & Input** puts an always-visible textarea
below the terminal that owns composition, then delivers the composed text to the session.
## Voice dictation
`Ctrl+Shift+V`, or the microphone button. There are three providers and the default is
`auto`, which prefers them in this order:
| Provider | Needs | Notes |
| ------------------ | ---------------------------------------------- | ------------------------------------------------------------ |
| **Claude** | Claude Code logged in on the server. Opt-in. | Uses this machine's existing Claude login. No extra key. |
| **Deepgram** | A Deepgram API key. | Nova-3, with automatic silence detection. |
| **Web Speech** | Nothing. | Browser-provided, quality varies. |
### Dictating through your Claude login
Off by default; enable it in **App Settings → Voice**.
Claude Code has its own voice mode, but it opens the **host's** microphone, and in Codeman
the CLI runs headless in a tmux pane while you are in a browser somewhere else entirely. So
Codeman captures audio in your browser and borrows only the backend: audio goes browser to
Codeman to Anthropic, and the page never sees the OAuth token.
Two deliberate limits:
- **Credentials are read only.** Codeman never refreshes your Claude token, because a
refresh rotates the refresh token and could sign you out of your own CLI. An expired token
is reported as expired rather than silently renewed.
- **Capture is raw PCM** at 16 kHz mono, which requires an AudioWorklet rather than the
usual browser recorder.
## Read My Mind
**Claude only, off by default.** Turn it on in **App Settings**, and a 🧠 button appears in
the header (on phones, in the keyboard bar instead).
It keeps a per-case **intent profile**: goals you or your agent write down, plus the prompts
you actually submitted in that case. Pressing 🧠 feeds that profile plus live session signals
to a single model call and shows a predicted next prompt.
What you can do with the result:
- **Send** it, **Insert** it into the composer, or edit it first.
- Pick one of the alternate suggestions, which swaps into the editable field without losing
your edits.
- **Rethink**, optionally with a steer note, to reject the whole set and try again.
**Nothing is ever sent automatically.** Every path requires a click.
Where the data lives: the profile is keyed by owner and the resolved working directory, so
it survives `/clear` and respawns. Prompts can contain secrets, so the store is written
0600 and is deliberately excluded from cross-session search.
Guide: [`docs/readmymind.md`](https://github.com/Ark0N/Codeman/blob/master/docs/readmymind.md).
## Programmatic input
Sending prompts over the API has one rule that catches everyone: **the payload must end with
`\r`** or Enter is never sent. The request still succeeds, the text sits unsubmitted in the
composer, and any wait burns its whole timeout on a turn that never started.
Input is also **single line**. Embedded newlines are stripped rather than rejected, so
`"echo A\necho B\r"` runs the joined `echo Aecho B`. Put multi-line content in a file and
tell the agent to read it.
See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
## Gotchas
- **Typed text sitting on screen has not been sent.** Press Enter.
- **`Ctrl+C` with a selection copies.** Clear the selection to interrupt.
- **Voice needs HTTPS.** Microphone access requires a secure context, same as push
notifications.
- **Read My Mind goes blind for sessions using a relocated Claude config directory**, along
with the other transcript-backed features. See [Agent CLIs](Agent-CLIs).
## Read next
- [Mobile Guide](Mobile-Guide) - the keyboard bar and touch input.
- [Keyboard Shortcuts](Keyboard-Shortcuts) - the full list.
- [Working With Files](Working-With-Files) - images and attachments as input.
+234
View File
@@ -0,0 +1,234 @@
# Installation
Getting Codeman onto a machine, verifying it works, updating it, and removing it.
## Requirements
| Requirement | Notes |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------- |
| **macOS or Linux** | Windows works through WSL2. See [Windows](#windows-wsl) below. |
| **Node.js 22+** | The installer offers to install it if missing. |
| **tmux** | Not optional. Sessions live inside tmux, which is what makes them survive a server restart, a dropped connection, or a closed laptop. |
| **An agent CLI** | At least one of [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), [Codex](https://developers.openai.com/codex/cli), [Antigravity](https://antigravity.google), [Gemini CLI](https://github.com/google-gemini/gemini-cli), [Pi](https://pi.dev). Plain shell sessions need none. See [Agent CLIs](Agent-CLIs). |
Codeman itself sends no telemetry and phones no home. The only network traffic is your
browser to your server, and whatever the agent CLI you chose does on its own.
## Route A: the installer (recommended)
```bash
curl -fsSL https://getcodeman.com/install | bash
```
This installs Node.js and tmux if they are missing, clones Codeman into `~/.codeman/app`,
and builds it.
What it asks you:
1. **Permission for every system change.** Package installs and agent CLI downloads are
prompted individually. Nothing is installed silently.
2. **How the dashboard should be reachable.** Three choices:
- **Tailscale** (recommended for phone access): keeps the loopback bind and walks you
through `tailscale serve`, including the tailnet HTTPS toggle, then verifies the result
end to end.
- **Your local network** (`0.0.0.0`): prompts for a password. Skipping the password takes
an explicit confirmation and ends on a loud warning.
- **This machine only** (`127.0.0.1`): the safest option, and the default for a bare
`codeman web` regardless of what you pick here.
Which one is preselected depends on what the installer finds. A fresh install defaults to
the local network, unless Tailscale is already connected, in which case it defaults to
Tailscale. An existing loopback install defaults to keeping loopback, or to Tailscale when
a serve mapping for Codeman is already there. A bare Enter never pulls in new software,
and a non-interactive run always keeps the safe loopback default.
3. **What to do when it finishes.** Run in this terminal, install as a background service
that starts on boot, or do nothing yet.
Re-running the same one-liner **updates an existing install in place**. Local changes in
`~/.codeman/app` are stashed rather than discarded, a running service is restarted and
verified, and your existing network binding is preserved. An interrupted first install
resumes instead of restarting.
Two other entry points exist:
```bash
install.sh update # update only
install.sh uninstall # remove
install.sh tailscale # retrofit Tailscale access onto an existing install
```
**Automation and CI**: with no terminal attached, any step that would change the system
aborts with instructions instead of running silently. Set `CODEMAN_NONINTERACTIVE=1` to
approve those steps. `CODEMAN_TAILSCALE=1` preselects the Tailscale answer, and never
installs Tailscale itself non-interactively.
## Route B: npm
```bash
npm install -g aicodeman
codeman web
```
The npm package is named `aicodeman`; the product is Codeman. Both `codeman` and
`aicodeman` are installed as commands.
The trade-off against Route A: no guided network setup, and the in-app self-updater does
not apply. npm installs report as non-updatable in **App Settings → System → Updates**, and
you update with `npm update -g aicodeman`.
## Route C: git clone
For contributing, or for running unreleased code.
```bash
git clone https://github.com/Ark0N/Codeman.git
cd Codeman
npm install # postinstall builds the vendored xterm addon bundles
npm run dev # dev server on http://localhost:3000
```
For a production run from a clone:
```bash
npm run build
npm run start
```
`npm run dev` runs TypeScript directly through `tsx` with no build step. The frontend is
plain JavaScript served from `src/web/public/` with no bundler, so editing a `.js` or `.css`
file and reloading the page is enough. The one exception is `index.html`, which is read once
at server start, so markup changes need a restart.
See [Contributing](Contributing) for the rest of the development loop.
## Installing an agent CLI
Codeman drives CLIs, it does not bundle them. Install at least one:
| CLI | Install | Notes |
| --------------- | ------------------------------------------------------------------ | -------------------------------------------------------------------------- |
| **Claude Code** | `npm i -g @anthropic-ai/claude-code` | The primary target. Some Codeman features are Claude-only: see [Agent CLIs](Agent-CLIs). |
| **OpenCode** | See [opencode.ai](https://opencode.ai) | |
| **Codex** | See [developers.openai.com/codex/cli](https://developers.openai.com/codex/cli) | |
| **Antigravity** | See [antigravity.google](https://antigravity.google) | Google's successor to the consumer Gemini CLI. |
| **Gemini CLI** | See [github.com/google-gemini/gemini-cli](https://github.com/google-gemini/gemini-cli) | Enterprise only since Google's June 2026 consumer cutover. |
| **Pi** | See [pi.dev](https://pi.dev) | No permission prompts and no sandbox by design. Read [Agent CLIs](Agent-CLIs) before using it on a repo you care about. |
Log each CLI in once, by hand, before pointing Codeman at it. Codeman never collects or
stores your CLI credentials.
## Verify the install
```bash
codeman doctor # checks Node, tmux, the agent CLIs, document converters
codeman --version
codeman web # then open http://localhost:3000
```
`codeman doctor --json` gives machine-readable output, and `--category core` narrows it to
the things a session cannot start without.
If the dashboard loads and **+ New Session** opens, you are done. Continue to
[Quick Start](Quick-Start).
## Where things live
| Path | What |
| ----------------------- | ------------------------------------------------------------------------------------------------ |
| `~/.codeman/app` | The installed code (installer route only). |
| `~/.codeman/` | All state: `state.json`, settings, session history, push keys, TLS certs. See [Core Concepts](Core-Concepts). |
| `~/codeman-cases/` | Cases created from scratch. Linked cases stay wherever they already are. |
| `~/.codeman/web.log` | Log for a detached (`-d`) server. |
Everything is under your home directory, and nothing needs root.
## Keeping it running
A bare `codeman web` dies with the shell that started it. Two ways to outlive that:
```bash
codeman web -d # detached; --status and --stop manage it
codeman service install # systemd user unit or macOS LaunchAgent; survives reboots
```
Full detail, including logs and the self-updater, is in
[Running As A Service](Running-As-A-Service).
## Updating
| Install route | How to update |
| ------------- | ----------------------------------------------------------------- |
| Installer | Re-run the one-liner, or **App Settings → System → Updates** in the UI. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
The in-app updater covers git-clone installs supervised by systemd or launchd. It restarts
the process that is running it, so the actual work happens in a detached script and the
browser polls across the restart. Progress appears in the UI.
## Uninstalling
```bash
install.sh uninstall # installer route
npm uninstall -g aicodeman # npm route
```
Neither removes `~/.codeman/` or `~/codeman-cases/`. Delete those by hand if you want the
state and your case folders gone as well, and check `~/codeman-cases/` first: linked cases
point at directories you already had, but cases created from scratch have their only copy
there.
Running tmux sessions are not killed by an uninstall. `tmux -L codeman kill-server` ends
them.
## Windows (WSL)
```powershell
wsl bash -c "curl -fsSL https://getcodeman.com/install | bash"
```
Codeman requires tmux, so Windows runs it inside
[WSL2](https://learn.microsoft.com/en-us/windows/wsl/install). If you do not have WSL yet:
run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, and install your agent CLI
*inside* WSL. `http://localhost:3000` then works from your Windows browser.
Work inside the Linux filesystem (`~/project`), not `/mnt/c/...`. Filesystem watching and
git are both dramatically slower across the Windows mount, and agents notice.
## macOS notes
**`Error: posix_spawnp failed.` on every session start.** node-pty publishes its macOS
`spawn-helper` without the executable bit, and macOS launches every PTY through it. Codeman
detects this and repairs it automatically on the first failure. If you hit it on a clone
install and want to fix it by hand:
```bash
npm run fix:node-pty
```
This is a `chmod`, not a rebuild. Look in `prebuilds/darwin-<arch>/`, not
`build/Release/`, which does not exist on macOS. Linux cannot reproduce this.
**launchd and PATH.** A LaunchAgent gets `/usr/bin:/bin:/usr/sbin:/sbin`, which finds
neither a Homebrew or nvm `node` nor `tmux` or `claude`. `codeman service install` bakes
your current PATH into the unit for exactly this reason, so prefer it over a hand-written
plist.
## Gotchas
- **`tmux: command not found` after a successful install.** The installer asks before
installing packages, and a declined prompt is a valid answer it remembers. Install tmux
and re-run.
- **Port 3000 in use.** `codeman web --port 8080`, or set `CODEMAN_PORT`.
- **Two Codemans on one machine.** The data directory and the tmux socket are both process
wide, so a second instance discovers and attaches the first one's live sessions. Give each
a distinct `CODEMAN_INSTANCE` before starting a second. See [Core Concepts](Core-Concepts).
- **The dashboard is not reachable from your phone.** That is the default, not a fault. The
server binds `127.0.0.1`. See [Remote Access](Remote-Access).
## Read next
- [Quick Start](Quick-Start) - your first working session.
- [Agent CLIs](Agent-CLIs) - picking and setting up a run mode.
- [Remote Access](Remote-Access) - reaching it from another device.
- [Troubleshooting](Troubleshooting) - when the above did not go as written.
+159
View File
@@ -0,0 +1,159 @@
# Keeping Agents Running
Codeman exists for the hours you are not at the keyboard. This page covers how it notices an
agent has stopped, what it does about it, and how to run a session overnight without
babysitting it.
Everything here is **per session and off by default**. A session you never configure just
sits there when it finishes, which is usually what you want.
## How Codeman knows an agent is idle
Harder than it sounds, and worth understanding, because it is what every other feature here
is built on.
**For Claude sessions**, the naive signal does not work. Claude redraws its prompt marker
roughly once a second all the way through a turn, so "saw a prompt, waited two seconds,
called it idle" flipped working sessions to idle a couple of seconds into every turn. Its
real working indicator is an animated line whose glyph and wording both change, and terminal
repaints arrive in partial fragments, so matching it in the output stream does not work
either.
So Codeman waits for the pane to go quiet, then **asks the screen** what is on it before
believing the session is idle. Turn-start detection works the same way in reverse: a
sustained run of repaints marks a turn as started, with the same screen check vetoing mere
keystroke echo. Idle now lands a few seconds after a turn genuinely ends.
There are several layers stacked on that: a completion message from the CLI, an AI check,
output silence, and token stability.
**For every other CLI**, there are no hooks to lean on, so detection is output
stabilization: the session is idle when output stops changing. Coarser, and it is why the
features further down this page are Claude-only.
## The Respawn Controller
Respawn keeps a session working past the point where the agent would otherwise stop. When
the session goes idle, Codeman runs a cycle and starts it again.
A cycle is up to four steps, each optional:
1. **Update prompt.** Ask the agent to write down where it got to, so the next round can pick
it up.
2. **`/clear`.** Reset the context window.
3. **`/init`.** Re-read the project's `CLAUDE.md`.
4. **Kickstart prompt.** Tell it to continue.
Steps 2 and 3 are what make long runs possible: without a context reset, a multi-hour
session eventually spends its whole window on its own history.
Configure it in **Session Options → Respawn**, then press **Enable**. It repeats until the
duration you set runs out.
| Setting | What it controls |
| ---------------------- | ----------------------------------------------------------------------- |
| **Idle timeout** | How long the session must be quiet before a cycle starts. |
| **Duration** | How long the whole arrangement stays armed. |
| **Inter-step delay** | Pause between the steps above, so a step is not sent into a busy pane. |
| **`/clear` + `/init`** | Whether the context reset happens at all. |
| **Update prompt** | What the agent is asked to record before the reset. |
| **Kickstart prompt** | What starts the next round. |
| **Auto-accept prompts**| Answer routine confirmation dialogs automatically. |
### Presets
Five built-ins, and the numbers matter more than the names. The idle timeout is the main
difference: a lead session coordinating subagents is legitimately silent for a minute at a
time, and a three second timeout would interrupt it constantly.
| Preset | Idle timeout | Duration | Built for |
| -------------- | ------------ | -------- | --------------------------------------------------------------- |
| **Solo** | 3s | 60 min | One agent working alone, fast cycles with a context reset. |
| **Subagents** | 45s | 240 min | A lead session running Task subagents; tolerates their silences. |
| **Team** | 90s | 480 min | Leading an agent team; tolerates long silences. |
| **Ralph/Todo** | 8s | 480 min | Working through a task list with progress tracking. |
| **Overnight** | 10s | 480 min | Unattended overnight runs with a full reset between cycles. |
Start from the preset that matches your shape of work and adjust the idle timeout first.
Presets you build yourself can be saved alongside these.
### What it costs
Every cycle is real tokens: the update prompt, the reset, and the kickstart, plus whatever
work follows. An overnight run is a deliberate spend, not a background nicety. The duration
setting is the ceiling, and it is worth setting honestly.
## Auto-resume when a usage limit resets
**Claude only.** At the top of the Respawn tab.
When Claude halts on a subscription limit, the message names the time the limit resets.
Codeman parses it, arms a timer for two minutes after that, then sends Escape followed by
`continue`.
The important part is what it does **not** do: respawn cycles are blocked while a session is
limit-paused. Without that, the next cycle would fire `/clear` and wipe the conversation you
are waiting to resume. This is the single most useful setting for overnight runs on a
subscription plan.
## The plan usage chip
**Claude only.** A header chip showing live subscription usage, on by default on desktop and
off on phones.
It works by installing a status line exporter into Claude Code, which posts Claude's own
rate limit data back to Codeman. The exporter is marker-identified, so it only ever touches
a status line Codeman installed, never one you wrote yourself, and it prints your footer
through so the in-terminal status line still works.
The chip and the exporter are the same setting. Turning the chip on without the exporter
would leave it showing a dash forever, so resolve it in one place: **App Settings**.
## Circuit breakers
Two, and they are unrelated:
- **The Ralph breaker** stops respawn thrashing. It moves from closed to half-open to open,
and is reset from the session's Ralph controls.
- **The PTY-exit breaker** trips when a session's process exits repeatedly and quickly, and
blocks automatic restarts so a broken configuration cannot spin forever.
The PTY-exit breaker resets **only** on an explicit clear. Reattaching to the session does
not clear it, deliberately, so a UI reconnect cannot paper over a session that is genuinely
failing to start.
## A working overnight setup
1. Start a Claude session in the case you want worked on.
2. Give it a clear goal and let it start. Respawn continues work, it does not invent it.
3. **Session Options → Respawn → Overnight preset.**
4. Turn on **auto-resume on usage limit**.
5. Set the duration to how long you actually want it running.
6. Press **Enable**.
7. Optionally turn on push notifications so a blocking question reaches your phone: see
[Notifications And Approvals](Notifications-And-Approvals).
In the morning, the **Away Digest** summarizes what happened while you were gone, and the
run summary and lifecycle log carry the detail.
## Gotchas
- **Respawn without a context reset stalls eventually.** The window fills with history and
the agent gets less useful every cycle.
- **An idle timeout that is too short interrupts real work.** If the agent runs long tool
calls or coordinates subagents, raise it. That is what the Subagents and Team presets are.
- **The update prompt is what makes a reset survivable.** After `/clear`, everything the
agent knows comes from that summary and the project files. A vague update prompt produces
a vague next cycle.
- **Non-Claude sessions can respawn**, but with output-based idle detection and no
usage-limit auto-resume.
- **Do not run respawn on a session you are actively typing in.** It will send prompts
underneath you.
## Read next
- [Autonomous Loops](Autonomous-Loops) - Ralph and the orchestrator, for structured
autonomous work rather than "keep going".
- [Cron Jobs](Cron-Jobs) - starting work on a schedule instead of continuing it.
- [Notifications And Approvals](Notifications-And-Approvals) - being told when it needs you.
- [`docs/respawn-state-machine.md`](https://github.com/Ark0N/Codeman/blob/master/docs/respawn-state-machine.md) - the state machine itself.
+72
View File
@@ -0,0 +1,72 @@
# Keyboard Shortcuts
Every binding, and how to change them. `Ctrl` also accepts `Cmd` on macOS.
Press `Ctrl+?` in the app for the same list in a floating overlay.
## Sessions and tabs
| Shortcut | Action |
| ------------------------------- | --------------------------------------------------------------- |
| `Ctrl+K` (also `Cmd+K`, `Alt+K`)| Find an open session or start a new one. |
| `Ctrl+W` | Kill the active session. |
| `Ctrl+Tab` | Next session. |
| `Alt+[` / `Alt+]` | Previous / next tab. |
| `Alt+1` to `Alt+9` | Switch to tab N. Physical keys, so macOS Option layouts work. |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move the active tab left / right. |
| `Alt+B` | Collapse / expand the session sidebar, when that layout is on. |
## Terminal
| Shortcut | Action |
| ----------------------- | --------------------------------------------------------------- |
| `Enter` | Send. |
| `Shift+Enter` | Insert a newline without sending. |
| `Ctrl+Enter` | Same. |
| `Ctrl+C` | Copy the selection, or interrupt when nothing is selected. |
| `Ctrl+Shift+C` | Copy the selection. Never interrupts. |
| `Ctrl+L` | Clear the terminal. |
| `Ctrl+Shift+R` | Restore terminal size. |
| `Ctrl` `+` / `Ctrl` `-` | Font size. |
| `Shift+Wheel` | Scroll the local buffer, even where the wheel is forwarded to the CLI. |
## Everything else
| Shortcut | Action |
| -------------- | ------------------------------- |
| `Ctrl+Shift+V` | Toggle voice input. |
| `Ctrl+?` | Shortcut reference overlay. |
| `Escape` | Close panels and modals. |
## Rebinding
**App Settings → Shortcuts.** Bindings live in a registry with per-user overrides, so a
rebind is stored as an override on top of the default rather than replacing the table.
Two things are deliberately not rebindable:
- **`Ctrl+C` smart copy.** The generic dispatch loop calls `preventDefault()` on every
shortcut it handles, and doing that to `Ctrl+C` would swallow the interrupt when nothing
is selected. It is handled separately for that reason.
- **`Escape`**, which closes whatever is open.
## Why some chords behave oddly
The terminal sees keystrokes before the app does. Any chord the app claims has to also be
swallowed at the terminal layer, or xterm writes the control byte into the session as well
as triggering the action. If you rebind something to a chord the terminal cares about
(`Ctrl+D`, say), expect the CLI to see it too.
`Alt+1` through `Alt+9` are matched on **physical key position** rather than the character
produced, so macOS Option layouts that produce `¡™£` still switch tabs.
## On phones
There is no physical keyboard, so the equivalents live in the keyboard accessory bar: `Esc`,
`Ctrl` as a one-shot modifier, `Tab`, arrows, and quick actions. See
[Mobile Guide](Mobile-Guide).
## Read next
- [The Dashboard](The-Dashboard) - what the shortcuts are navigating.
- [Settings Reference](Settings-Reference) - where the overrides are stored.
+165
View File
@@ -0,0 +1,165 @@
# Mobile Guide
Codeman on a phone is not a shrunken desktop UI. It is the surface most of its design
attention has gone into, because checking on an agent from a bus is the thing this software
is for.
<p align="center">
<img src="https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/screenshots/mobile-session-keyboard-20260727.png" alt="Answering an agent prompt on a phone" width="300">
</p>
## Getting there
1. **Set up access.** Tailscale is the recommended route and gives you real HTTPS. See
[Remote Access](Remote-Access).
2. **Log in by QR.** Open the dashboard on your desktop and scan the code. No password
typing. Tokens are single use and rotate every 60 seconds.
3. **Install it to your home screen.** On iOS this is mandatory for push notifications;
Safari does not deliver push to tabs. On Android it makes the app full screen.
HTTPS matters for more than security here: microphone access and push notifications both
require a secure context.
## The layout
| Element | Where |
| -------------------- | --------------------------------------------------------------------- |
| Header | Fixed at the top, deliberately minimal. Desktop-only controls never appear. |
| Tab strip | Scrolls horizontally. The active tab is always scrolled into view. |
| Terminal | The rest of the screen. |
| Toolbar | Bottom: Run, Stop, **Enter**, case picker, voice, settings. |
| Keyboard bar | Above the on-screen keyboard when it is open. |
Layout respects notch and home-indicator safe areas, touch targets are 44px, and the case
picker is a bottom sheet rather than a dropdown.
**Swipe left and right** on the terminal to switch sessions.
## The home screen
Tapping the "C" logo gives a session overview rather than a welcome page:
1. **NEEDS YOU** first: sessions blocked on a question, with answer strips so you can
resolve them without opening the session.
2. **CURRENT SESSIONS** with live status.
3. **PAST SESSIONS**, resumable.
Row status uses the same language as the tabs: green when fine, pulsing while working,
yellow when waiting for input, red when a question is pending.
The split Run button carries the same per-backend colours as the desktop toolbar, and its
picker mirrors the desktop run-mode menu.
On by default; it can be turned off in settings.
## The keyboard accessory bar
A row of keys above the virtual keyboard, and what it contains depends on the session.
**Agent sessions** get quick actions: `/init`, `/clear`, `/compact`, a clipboard key, `Esc`,
a path picker, an image key, and 🧠 when Read My Mind is on. Destructive commands need a
double press, so you cannot fire `/clear` with a stray thumb.
**Shell sessions** automatically swap it for terminal controls: `Ctrl`, `Esc`, `Tab`, four
arrows, paste, and dismiss. Your normal preference is remembered and restored when you
switch back to an agent session, so a settings change during a shell session cannot strip
the bar away permanently.
### One-shot Ctrl
`Ctrl` on the shell bar is a **one-shot modifier**: tap `Ctrl`, then tap `c`, and the
control byte is sent. It disarms on use, on a second tap, on any other accessory key, on a
session switch, and when the keyboard closes.
That list matters. A modifier left armed turns your next innocent keystroke into a control
byte, so it is deliberately eager to disarm. Keys with no control equivalent pass through
unchanged, exactly like a hardware keyboard.
## The Enter button
The toolbar's dedicated **Enter** button exists because of local echo. On a phone, the
characters you type are painted locally and have not reached the agent yet; Enter flushes
them and then submits.
It replays the keypress through the terminal rather than sending a bare carriage return.
Sending a bare `\r` would submit an empty line and strand your typed text on screen, which
looks exactly like a dead button.
On phones this button replaces the desktop's **Run Shell** control; starting a shell moved
into the Run dropdown.
## Tapping, links and copying
- **Tap a link** in terminal output and it opens in a new tab. Same for a link in an agent's
answer in the response viewer — it opens a tab rather than navigating the dashboard away,
which on a phone would unload the whole session view.
- **Tap a file path** an agent printed and the file-preview overlay opens; a log path opens the
log viewer. Works in scrolled-up transcript too.
- A tap on the prose *beside* a link still places the cursor as usual, and a tap on a dialog's
numbered choice still answers the dialog even when the row contains a path — the dialog wins,
because on a phone it is the only interaction that matters.
- **Long-press to select text**, then drag, or tap the other end to extend the selection — no
hairline handles to grab. A small bar offers **Copy**, **Line** (the whole logical line,
wrapped rows included) and dismiss. Copy works on plain-HTTP installs too, where the browser
clipboard API is unavailable.
- A swipe is never mistaken for a long-press, and the keyboard stays down while you select.
## Scrolling and the keyboard
- The terminal and toolbar shift up when the keyboard opens, tracked through the browser's
visual viewport rather than guessed.
- **Two ways to dismiss the keyboard**: tap outside the terminal on inert space, or tap twice
on inert terminal content. Tapping a control never dismisses it, and tapping the prompt row
keeps focus so you can place the caret.
- A scroll is never mistaken for a tap: travel is measured from the start of the gesture, and
multi-touch never counts.
- **A long prompt stays visible.** Once what you are typing wraps past the last visible row it
grows upward over the transcript instead of sliding under the keyboard, so the end of the
sentence — where the cursor is — is always on screen. A prompt taller than the visible strip
shows its tail.
## Voice
The microphone button, or the keyboard bar. Providers and setup are covered in
[Input And Voice](Input-And-Voice). Dictating is often faster than typing a prompt on a
phone, and it is the main reason the feature exists.
## Notifications
Push notifications reach you with no tab open, and with the Approvals Inbox on they carry
**Approve** and **Deny** buttons handled by the service worker, so you can unblock an agent
from the lock screen.
Setup in [Notifications And Approvals](Notifications-And-Approvals).
## Reading long answers
The terminal viewport is small. **Last Response** (opt-in header button) renders the agent's
last answer as scrollable text instead, with a **More** button for additional context.
The [File Viewer](Working-With-Files) works on phones too, including edit mode, which is
enough to fix a typo an agent introduced while you are away from your desk.
## What is deliberately not on phones
- Extra header buttons. New header controls are kept off phones by policy, with a test that
enforces it.
- The Approvals bell. Phones get the NEEDS YOU strips on the home screen instead.
- The desktop home tab rail, which needs a wide window.
- Lineage arcs, which are a desktop overlay.
## Gotchas
- **Typed text sitting on screen has not been sent.** Press Enter.
- **iOS needs the home screen install for push**, not just a bookmark.
- **iOS Safari can serve stale JavaScript after an update** until the tab is fully closed.
Close it and reopen.
- **Plain HTTP over a LAN address disables voice and push.** Use HTTPS.
- **An armed `Ctrl` is visibly highlighted.** If it looks the same as a resting key, you are
on an old version, on a light skin.
## Read next
- [Remote Access](Remote-Access) - getting the phone connected in the first place.
- [Notifications And Approvals](Notifications-And-Approvals) - being told when you are needed.
- [Input And Voice](Input-And-Voice) - local echo, dictation, and the input rules.
+97
View File
@@ -0,0 +1,97 @@
# Multi-User Mode
Share one Codeman with a small trusted team. Each person gets their own login and workspace,
and sessions, cases, search, and live events are scoped to their owner.
**Off by default.** Without the flag, behaviour is identical to single-user Codeman, because
every scoping check short-circuits.
## Read this before enabling it
**Multi-user mode separates workspaces. It does not sandbox users from each other.**
Every session still runs as the **same operating system account**. A determined user's agent
can reach another user's files, because at the OS level they are the same user. This is a
convenience and organization feature, not a security boundary.
If you need real isolation:
- Pair each user with [Docker Cases](Docker-Cases), which gives their work its own
filesystem and network.
- Or run separate Codeman instances under separate OS accounts, each with its own
`CODEMAN_INSTANCE`.
"Small trusted team" is the honest description of who this is for.
## Enabling it
```bash
codeman users add alice --admin # create the first admin, prompts for a password
codeman web --multiuser # or CODEMAN_MULTIUSER=1
```
Then manage users from the CLI or the **Users** entry in App Settings:
```bash
codeman users add bob # a regular user
codeman users list
codeman users passwd bob # reset to a one-time password
codeman users rm bob
```
`--password-stdin` reads the password from standard input, for scripts.
Accounts live in `~/.codeman/users.json` with scrypt-hashed passwords, mode 0600.
Administrative actions are audited to `~/.codeman/admin-audit.jsonl`.
## What each user gets
| Thing | Scope |
| ------------------- | ---------------------------------------------------------------------------- |
| **Case space** | `~/codeman-users/<name>/cases`, their own. |
| **Sessions** | Only theirs are listed, reachable, or controllable. |
| **Events** | Live event routing is per owner, and fails closed. |
| **Search** | Scoped on read, including historical results. |
| **File previews** | Scoped to sessions they own. |
| **Path picker** | Only their own user space as a root, not the whole home directory. |
Admins see everything.
Ownership threads through every list endpoint, the session lookup helper, the WebSocket
layer, and file previews. A user cannot address another user's session even by id.
## Safer defaults for regular users
Non-admins get tighter defaults, and lifting them is an explicit per-user grant:
| Default | Meaning |
| ---------------------------------- | ------------------------------------------------------------------------ |
| Claude runs in `auto` permission mode | Anthropic's classifier-guarded mode instead of skip-prompts. |
| Raw shell sessions require a grant | A plain shell is unmediated machine access. |
| Skip-permissions requires a grant | Same reasoning. |
| Cron `launchCommand` requires a grant | It is an arbitrary command on a schedule. |
| Pi project trust defaults to off | Trust makes Pi execute repo-local TypeScript. |
These exist because the OS boundary is shared. They narrow what a normal account can do
casually; they do not make the account a sandbox.
## Accounts and sessions
Each user authenticates with their own name and password rather than the shared
`CODEMAN_PASSWORD`. Logins are individually revocable: disable, reset, or delete an account
at any time, and existing browser sessions can be revoked.
## Gotchas
- **Enabling it does not migrate existing cases** into a user space. They stay where they
are, owned by whoever the ownership rules resolve them to.
- **Admins see everything**, including other users' sessions. Choose admins accordingly.
- **The audit log is append-only and local.** Ship it somewhere if you care about it.
- **It is not a substitute for OS accounts.** Restating this because it is the one thing
people get wrong.
## Read next
- [Security](Security) - where this fits in the model, and what it does not cover.
- [Docker Cases](Docker-Cases) - the isolation story that actually isolates.
- [`docs/multi-user-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/multi-user-plan.md) - the design.
+147
View File
@@ -0,0 +1,147 @@
# Notifications and Approvals
An agent that stops to ask a question, with nobody watching, is a run that quietly wasted an
hour. This page covers every way Codeman tells you it needs you, and how to answer without
opening the session.
## The signals, cheapest first
| Surface | Reaches you | Default |
| ---------------------- | ------------------------------------------------- | ------- |
| Tab alert | While the dashboard is open | On |
| Browser title flash | Another tab in the same browser | On |
| Desktop notification | Another window on the same machine | Opt-in |
| Push notification | Anywhere, even with no tab open | Opt-in |
| Approvals Inbox | One queue across every session | Opt-in |
| Phone overview | Phone home screen, NEEDS YOU section | On |
| Away Digest | Afterwards, as a summary | Opt-in |
## Tab alerts
The tab itself changes state:
| State | Meaning |
| -------------------- | ---------------------------------------------------------- |
| Yellow, blinking | The agent is waiting for input from you. |
| Red, blinking | A question or permission prompt is blocking the session. |
These are a steady colour with a pulse layered on top, not a blink to transparent, so a tab
needing attention looks that way at every point in the cycle.
They survive a reload. The alert state is re-seeded from the server on page load, so
reloading the dashboard while a permission dialog is blocking a session does not leave you
with a normal-looking tab.
For Claude sessions, these come from Claude Code's hooks and are precise about *why* the
session stopped. For other CLIs there are no hooks, so you get the coarser output-based
signal.
## Window title and OS notifications
The browser tab title is prefixed `codeman:<host>`, so several Codeman instances across
several machines stay distinguishable at a glance. Override the hostname with
`codeman web --title-hostname <name>`.
Desktop notifications use the same prefix. Enable them in **App Settings → Notifications**.
## Push notifications
Push reaches your phone with **no Codeman tab open at all**, which is the only option that
works while you are actually away.
Setup:
1. Open Codeman over **HTTPS**. Web push requires a secure context. Tailscale gives you real
HTTPS; `--https` gives you a self-signed certificate; plain HTTP over a LAN address will
not work.
2. **App Settings → Notifications → Subscribe**, and accept the browser prompt.
3. On **iOS**, add Codeman to your home screen first. Safari only delivers web push to
installed web apps, not to tabs.
Once subscribed, a blocking prompt reaches your phone even from a locked screen.
## The Approvals Inbox
**Opt-in, off by default. Claude sessions only.**
One queue of every prompt currently waiting on a human, across all your sessions, answerable
in place. When you have eight workers running, this is the difference between checking eight
tabs and checking one list.
Turn it on in **App Settings**. Surfaces:
- **A header bell** with a count, hidden entirely while the count is zero. Never shown on
phones.
- **A drawer** listing each waiting card.
- **NEEDS YOU strips** at the top of the phone overview home screen.
Each card shows the session, the case, and the captured prompt with its options. Answering
sends the keystroke into the session for you: a digit for a menu choice, Escape to decline,
or free text for an idle prompt.
Behaviour worth knowing:
- **One item per session.** A newer prompt supersedes the older one, because the older one
is no longer on screen.
- **Menu answers are validated against the live screen.** Codeman re-captures the pane before
sending, and refuses with a conflict if the dialog is no longer there. Otherwise your
keystroke would land in the composer as stray text.
- **Permission and question items clear only on definitive signals**: the turn ending, the
dialog completing, an answer, a supersede, the session exiting, or a 12 hour timeout. They
do not clear on a heuristic "looks busy again" signal, because that signal is wrong often
enough to lose a real prompt.
- **In memory only.** Restarting the server clears the queue; the prompts themselves are
still sitting in the sessions.
### Approve and Deny from the notification
With the inbox enabled, push notifications carry **Approve** and **Deny** buttons. Those are
handled by the service worker directly, so they work with no tab open: tap Approve on a
locked phone and the agent continues.
With the inbox off, the buttons are stripped from the notification payload entirely rather
than being shown and failing.
## The phone overview
On phones, tapping the "C" logo gives a session overview with **NEEDS YOU** first, then
current sessions, then past ones. Rows use the same language as the tab strip: a green dot
when fine, pulsing while working, yellow when waiting for input, red when a question is
pending.
Answer strips let you resolve a prompt straight from the home screen without opening the
session.
## The Away Digest
Retrospective rather than live: what happened while you were gone, aggregated from the
lifecycle log, run summaries, live sessions, token statistics, and recent subagents.
It is the morning-after view for an overnight run. Enable its header button in
**App Settings → Header & Panels**.
## Recommended setup for unattended runs
1. HTTPS access, ideally Tailscale. See [Remote Access](Remote-Access).
2. Push notifications subscribed, with Codeman installed to the home screen on iOS.
3. Approvals Inbox on.
4. Auto-resume on usage limit on, for each session you leave running. See
[Keeping Agents Running](Keeping-Agents-Running).
That combination means a blocking question wakes your phone and can be answered in two taps
from the lock screen.
## Gotchas
- **No push over plain HTTP.** It is a browser requirement, not a Codeman one.
- **iOS needs the home screen install.** A Safari tab will never receive push.
- **The bell is invisible at zero.** That is deliberate, not a broken setting.
- **Approvals are Claude-only.** They are built on hook events the other CLIs do not emit.
- **A stale menu answer is refused, not sent.** If you answer a card for a dialog that has
since gone away, Codeman declines rather than typing a digit into the composer.
## Read next
- [Keeping Agents Running](Keeping-Agents-Running) - what to configure before walking away.
- [Mobile Guide](Mobile-Guide) - the phone surfaces in full.
- [Settings Reference](Settings-Reference) - where each of these toggles lives.
+153
View File
@@ -0,0 +1,153 @@
# Quick Start
From an installed Codeman to a working agent, in about five minutes. If you have not
installed yet, start at [Installation](Installation).
## 1. Start the server
```bash
codeman web
```
It prints a URL, `http://localhost:3000` by default. Open it.
The server binds `127.0.0.1` only, so this URL works from the machine running it and
nowhere else. That is deliberate: Codeman starts agents with permission prompts skipped by
default, so anyone who can reach the dashboard can run code on this machine. Reaching it
from your phone is a separate, deliberate step covered in [Remote Access](Remote-Access).
To keep it alive after you close the terminal, use `codeman web -d` instead, or install it
as a service. See [Running As A Service](Running-As-A-Service).
## 2. Meet the welcome screen
With no sessions running you get the welcome screen:
- **Run buttons** for each agent CLI Codeman found on your PATH. If you expected one and it
is missing, its binary is not visible to the server; see [Agent CLIs](Agent-CLIs).
- **A QR code**, if a password is set. Scanning it logs a phone in without typing anything.
- **Resume Conversation**, a list of past sessions, including Claude conversations started
outside Codeman. Empty on a fresh install.
- **Search**, across sessions, events, and files.
You can click a Run button right now and get a working agent in your current case. The rest
of this page is the deliberate version.
## 3. Pick or create a case
A **case** is a named working directory that Codeman remembers. Every session runs inside
one. The case picker is in the bottom toolbar.
To make a new one, click **+** next to the picker. The Add Case dialog has three tabs:
| Tab | Use it when |
| ----------------- | ------------------------------------------------------------------------------------------------------------------ |
| **Create New** | Starting a fresh project. Creates `~/codeman-cases/<name>` and scaffolds a `CLAUDE.md` into it. |
| **Clone Repo** | Working on an existing public repo. Paste the URL; Codeman preflights it as you type, offers the repo's real branches and tags, and fills in the case name. |
| **Link Existing** | The code is already on disk. Point at the folder, with **Browse** if you would rather click than type. |
The gear next to the picker holds two per-case toggles: **Agent Teams** and
**1M Opus Context**. Both are off by default and both are safe to ignore for now.
**Create New** also has a checkbox for running the case inside a Docker container, and a
**Remote** panel for running it over SSH on another machine. Those are
[Docker Cases](Docker-Cases) and [Remote SSH Sessions](Remote-SSH-Sessions); skip them for
your first session.
## 4. Pick a run mode and hit Run
The **Run** button starts an agent in the selected case. The arrow next to it picks which
one:
| Mode | What starts |
| -------------------- | -------------------------------------------------------------- |
| **Claude Code** | The default, and the mode every Codeman feature supports. |
| **OpenCode** | |
| **Codex** | OpenAI's CLI. |
| **Gemini** | Enterprise only since Google's consumer cutover. |
| **Antigravity** | Google's successor to the consumer Gemini CLI. |
| **Pi** | No permission prompts and no sandbox by design. |
| **Terminal / Shell** | A plain shell, no agent. Also the **Run Shell** button. |
The dropdown also lists any saved dashboard URLs ([Web Tabs](Web-Tabs)) and your recent
sessions. Those do not change the run mode: Run always means "start an agent".
Click **Run**. A tab appears, and Codeman spawns the CLI on a real PTY inside a tmux
session and streams it to your browser.
The number spinner beside the button starts several sessions at once, up to 20. Useful for
fanning the same case out across parallel workers; unnecessary for a first run.
## 5. Talk to the agent
Click into the terminal and type. It is a real terminal (xterm.js over a real PTY), so full
TUIs render properly and everything the CLI supports works, slash commands included.
| Key | Effect |
| ---------------------------- | --------------------------------------------- |
| `Enter` | Send. |
| `Shift+Enter` / `Ctrl+Enter` | Newline without sending. |
| `Ctrl+C` | Copy if text is selected, otherwise interrupt. |
| `Ctrl+Shift+V` | Voice input. |
You can also paste or drag an image straight into the session, and register external files
as attachments. See [Working With Files](Working-With-Files) and
[Input And Voice](Input-And-Voice).
Input is delivered **exactly once**, even if your connection drops mid-prompt. A dropped
link never loses a prompt and never sends it twice.
## 6. Read the tab
The tab tells you what the session is doing without opening it:
| Signal | Meaning |
| --------------------- | ---------------------------------------------------------- |
| Green dot | Alive and idle. |
| Pulsing green dot | Working on a turn. |
| Yellow, blinking | Waiting for you to type something. |
| Red, blinking | A question or permission prompt is blocking the agent. |
Full tour in [The Dashboard](The-Dashboard). If you want a phone notification when an agent
needs you, that is [Notifications And Approvals](Notifications-And-Approvals).
## 7. Leave, and come back
Close the browser tab. Close the laptop. The agent keeps running, because it lives in tmux
and not in your browser.
Reopen the dashboard and the session is still there with its scrollback intact. First load
of a session pulls the full tmux scrollback, so you get the history, not just what arrived
after you reconnected.
This also survives restarting the Codeman server itself. What does not survive is killing
the tmux server or rebooting the machine.
## 8. Stop things
| To do this | Do that |
| ------------------------- | ------------------------------------------------------------------- |
| Interrupt the current turn | `Ctrl+C` with nothing selected, or the **Stop** button. |
| Close one session | `Ctrl+W`, or the tab's close control. |
| Stop the server, keep agents | `codeman web --stop`. The tmux sessions stay alive. |
| Stop everything | `tmux -L codeman kill-server`. |
If you are working *inside* a Codeman-managed session (`echo $CODEMAN_MUX` prints `1`),
never run `tmux kill-session` or `pkill claude` by hand. You will kill the session you are
sitting in, along with its siblings.
## Where to go next
**Make it run without you.** [Keeping Agents Running](Keeping-Agents-Running) covers idle
detection, respawn cycling, and auto-resume when a subscription limit resets. That is the
feature Codeman exists for.
**Get it on your phone.** [Remote Access](Remote-Access), then
[Mobile Guide](Mobile-Guide).
**Understand what you just used.** [Core Concepts](Core-Concepts) explains cases, sessions,
run modes, and what state lives where.
**Automate it.** [Cron Jobs](Cron-Jobs) for scheduled work,
[Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) for agents that spawn and
supervise other agents.
+209
View File
@@ -0,0 +1,209 @@
# Remote Access
Reaching your Codeman from a phone, a laptop on the other side of the house, or a hotel
network. This is the page to read carefully, because Codeman's dashboard is a
remote-code-execution surface by design: it starts agents with permission prompts skipped,
so whoever can reach it can run code on your machine.
## Start from the default
`codeman web` binds `127.0.0.1`. It is reachable from the machine running it and nothing
else, which is why the no-password default is safe out of the box. Every option below is a
deliberate step away from that.
Two rules that make the rest of this page simple:
1. **Never expose Codeman on a network without `CODEMAN_PASSWORD`.** Binding a non-loopback
host without one starts, but prints a loud warning with the fixes.
2. **Prefer keeping the loopback bind** and putting an authenticated tunnel in front of it,
over binding wide and relying on a password alone.
## Pick an approach
| Approach | Good for | Cost |
| --------------------- | ----------------------------------------------------- | --------------------------------------------------------- |
| **Tailscale** | Phone access, permanently. The recommended setup. | Install Tailscale on both devices. |
| **Cloudflare tunnel** | A public URL, quickly, from anywhere. | Public URL, so a password is mandatory. |
| **LAN + password** | Home network only, no extra software. | Every device on your LAN can reach the login page. |
| **SSH port forward** | You already SSH to the box. | Manual, per session, terminal-bound. |
## Tailscale (recommended)
Your devices join a private network, and Codeman stays bound to loopback. Nothing is
published to the internet, and you get real HTTPS with a real certificate.
The installer sets this up for you, including installing Tailscale, logging in, enabling
tailnet HTTPS, and verifying the result end to end. To retrofit it onto an existing
install:
```bash
install.sh tailscale
```
By hand:
```bash
tailscale serve --bg 3000
tailscale serve status
```
Then open `https://<machine>.<tailnet>.ts.net` from any device on your tailnet.
Notes:
- Keep the loopback bind. `tailscale serve` connects to `127.0.0.1:3000` locally, so
binding wider adds exposure and buys nothing.
- Your tailnet is the authentication boundary. Setting `CODEMAN_PASSWORD` as well is
reasonable defence in depth, especially if other people have devices on your tailnet.
- Codeman's Host-header allowlist already accepts `.ts.net`, so no extra configuration is
needed.
- The installer never resets or rewrites `serve` mappings other than the one pointing at
Codeman's port, so unrelated serve configuration is left alone.
## Cloudflare tunnel
A free [quick tunnel](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/)
gives you a public HTTPS URL with no port forwarding, no DNS, and no static IP:
```
Browser → Cloudflare edge (HTTPS) → cloudflared → localhost:3000
```
Prerequisites: [`cloudflared`](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/downloads/)
installed, and `CODEMAN_PASSWORD` set.
```bash
./scripts/tunnel.sh start # starts the tunnel, prints the public URL
./scripts/tunnel.sh url
./scripts/tunnel.sh status
./scripts/tunnel.sh stop
```
The quick-tunnel URL is a random `*.trycloudflare.com` address that changes every time the
tunnel restarts. For a stable hostname, `./scripts/tunnel.sh named setup` walks through a
named tunnel.
To survive reboots:
```bash
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
```
There is also a toggle in **App Settings → System → Remote access**.
**The tunnel refuses to start without a password.** That is on purpose: a public URL with no
authentication is a terminal on your machine handed to the internet. Acknowledging the risk
explicitly is possible from the UI toggle, and only from there; the API will not do it for
you.
## LAN plus password
```bash
export CODEMAN_PASSWORD='something long'
codeman web -H 0.0.0.0 --https
```
Every device on your local network can now reach the login page. `--https` generates a
self-signed certificate into `~/.codeman/certs/`, which your browser will warn about once.
`CODEMAN_USERNAME` defaults to `admin`.
The installer offers this path and prompts for the password. On re-runs it preserves
whichever binding you already chose.
## SSH port forward
No configuration at all, if you already have SSH access:
```bash
ssh -L 3000:localhost:3000 you@your-box
```
Then open `http://localhost:3000` on the local machine. Codeman keeps its loopback bind and
sees a local connection. Good for occasional access, awkward as a permanent arrangement
because it dies with the SSH session.
## Logging in from a phone
Typing a long password on a phone keyboard is miserable, so Codeman issues **single-use QR
tokens**. The desktop dashboard shows a QR code; scan it and the phone is authenticated.
How it behaves:
- The code rotates every 60 seconds, with a 90 second grace window so scanning during a
rotation still works.
- Each token is **single use**. The moment a phone consumes it, a new one is generated.
- The URL contains a 6-character lookup code, not the secret, so it does not leak through
browser history, `Referer` headers, or the tunnel provider's logs.
- The desktop shows a toast naming the device and browser that just authenticated, with a
one-click revoke.
- QR attempts are rate limited separately from password attempts, so a mistyped password
cannot lock out your QR login and vice versa.
Someone holding only the tunnel URL still meets the normal password prompt. The QR is the
fast path, not a bypass.
Design detail and the threat analysis it is built against:
[`docs/qr-auth-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/qr-auth-plan.md).
## Behind a reverse proxy
Codeman enforces a Host-header allowlist on every request to block DNS rebinding, and the
same allowlist gates the cross-site Origin check. It accepts `localhost`, IP literals, the
bind host, `.ts.net`, `.trycloudflare.com`, `.cfargotunnel.com`, and the active managed
tunnel.
**Your own domain is not on that list.** Add it:
```bash
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
A bare entry matches that exact host; a leading dot matches subdomains. Without this, a
correctly configured proxy still gets `403 host not allowed`, which reads like a proxy bug
and is not one.
Also make sure the proxy forwards WebSocket upgrades. The terminal is a WebSocket, and the
upgrade runs the same Host and Origin checks, closing with code `4003` on failure.
## Session cookies and rate limits
The first request prompts for HTTP Basic credentials. On success the server issues an opaque
`codeman_session` cookie (24 hour lifetime, extended on activity, validated server-side so
it cannot be forged offline). Ten failed attempts from one IP produce a `429` with a 15
minute decay.
A valid cookie or a correct password recovers immediately even while an attacker is hammering
the same IP, which matters because all tunnel traffic arrives from one loopback address.
## Terminal alternatives
You do not have to use a browser. `sc` is a thumb-friendly session chooser for SSH clients
like Termius or Blink:
```bash
sc # interactive chooser
sc 2 # attach to session 2
sc -l # list
```
Detach with `Ctrl+A D`. The sessions are the same ones the dashboard shows.
## Common problems
| Symptom | Cause and fix |
| ----------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| `403 host not allowed` | Your domain is not in the allowlist. Set `CODEMAN_ALLOWED_HOSTS`. |
| Phone shows the login page but the terminal never connects | The proxy is not forwarding WebSocket upgrades. |
| Browser warns about the certificate | Expected with `--https` and its self-signed certificate. Tailscale gives you a real one instead. |
| LAN IP does not respond, but a tunnel to the same box works | The server is bound to loopback. That is the default. A tunnel reaches it; a LAN browser cannot. |
| Hooks stopped working after switching to HTTPS | Hook callbacks need `-k` for the self-signed certificate. Recent versions self-heal existing cases; if yours predates that, recreate the case's hooks. |
| Everything is slow over the tunnel | Quick tunnels route through Cloudflare's edge. Tailscale is usually a direct connection and much faster. |
## Read next
- [Security](Security) - the whole model, and the hardening checklist.
- [Mobile Guide](Mobile-Guide) - once you can reach it from the phone.
- [Running As A Service](Running-As-A-Service) - keeping server and tunnel up across reboots.
- [`docs/security-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/security-architecture.md) - the full model.
+101
View File
@@ -0,0 +1,101 @@
# Remote SSH Sessions
Point a case at another machine and the agent runs **there**, with the same dashboard,
mobile UI, and autonomy features. Your laptop becomes a window onto a session living on the
remote host.
Like Docker, this is a **location overlay** on a case, not a run mode. All seven run modes
work remotely. See [Core Concepts](Core-Concepts).
## Why bother
The agent runs where the work is: a build server, a NAS, a GPU box, a machine reachable only
through a jump host. Your laptop can sleep, change networks, or close, and the run continues.
## Setting it up
**Add Case → Remote**:
| Field | Notes |
| --------------------- | -------------------------------------------------------------------- |
| **Host** | Hostname or IP. |
| **Username** | The SSH user. |
| **Port** | Defaults to 22. |
| **Identity file** | `~` and `$HOME` are expanded for you. |
| **Jump host** | The `-J` equivalent, `[user@]host[:port]`. |
| **SOCKS proxy** | For hosts reachable only through a proxy. |
| **Extra SSH options** | Any `KEY=VALUE` options your normal connection needs. |
| **Remote path** | The working directory on that machine. |
Hosts are saved and reusable, so a second case on the same machine is just a path. Host
profiles can also carry per-run-mode launch command overrides, for when the binary lives
somewhere unusual on that host.
The remote host needs **tmux**. Codeman probes for it when you link the host rather than
failing later at launch.
## What actually runs
The agent lives inside a dedicated tmux server on the **remote** host, and Codeman fronts it
with a local tmux pane running `ssh`.
That two-layer arrangement is what makes it durable: a dropped SSH connection, a network
change, or a closed laptop kills the local pane, not the remote session. Reconnecting lands
back in the same live conversation.
The remote session name is deliberately chosen so that a Codeman **running on the target
host** will not adopt it as one of its own. Two Codemans, one host, no interference.
## Auto-reconnect
A watcher with bounded backoff notices a dead SSH pane and quietly reattaches to the still
running remote session. On by default; the kill switch is in
**App Settings → Agents & CLIs → Remote auto-reconnect**.
Intentional kills are never revived. Closing a session means closing it.
## Discover and attach
Codeman can list the `codeman-*` sessions already running on a host, whether that machine's
own Codeman started them or another operator did, and attach to one.
The distinction that matters:
| Session | On tab close |
| ------------ | ------------------------------------------------ |
| **Launched** | Killed, like any local session. |
| **Attached** | **Detached, never killed.** |
Attaching to someone else's session and closing your tab must not end their run, so it does
not. Several clients can attach the same remote session at different window sizes without
clamping each other, and discovery shows a shared badge with the client count.
## Security
Every SSH command line in Codeman flows through one builder that shell-escapes every
user-supplied field: identity paths, jump hosts, proxy commands, and extra options. That is
the entire injection surface, and it is deliberately a single function rather than string
concatenation spread across the codebase.
Host, path, and identity fields are schema-validated on top of that.
Codeman does not store SSH passwords. Use keys, as you would for any other automation.
## Gotchas
- **The remote host needs tmux.** Probed at link time, so you find out immediately.
- **The local working directory is meaningless** for a remote session, and is not used.
- **Run flows must go through the quick-start path** for remote cases. This matters if you
are driving Codeman over the API: the plain session-create endpoint validates the working
directory locally and has no case concept, so it will reject or misroute a remote case.
- **Latency is SSH latency.** Local echo helps the typing feel, but a slow link is a slow
link.
- **Transcript-backed features follow the transcript.** Subagent windows and similar surfaces
read files on the machine where the agent runs.
## Read next
- [Core Concepts](Core-Concepts) - overlays versus run modes.
- [Docker Cases](Docker-Cases) - the other overlay.
- [Security](Security) - the wider model.
- [`docs/remote-sessions.md`](https://github.com/Ark0N/Codeman/blob/master/docs/remote-sessions.md) - the full design.
+197
View File
@@ -0,0 +1,197 @@
# Running As A Service
Keeping Codeman up: past the shell you started it in, past a logout, past a reboot. Plus
logs, updates, and running more than one instance.
## Three levels
| Level | Survives | Command |
| -------------------- | ----------------------------------------- | ------------------------- |
| Foreground | Nothing. Dies with the terminal. | `codeman web` |
| Detached | Closing the shell and logging out. | `codeman web -d` |
| Service | Reboots. | `codeman service install` |
Agents themselves survive all three, because they live in tmux. Stopping the server never
stops the agents.
## Detached mode
```bash
codeman web -d # start detached; logs to ~/.codeman/web.log
codeman web --status # is it up, and on which pid
codeman web --stop # graceful stop; agents keep running
```
`-d` waits until the server actually answers before reporting success, so a port clash never
reads as a successful start.
Two implementation details that explain the behaviour:
- It relaunches the same entry script detached, so there is no controlling terminal and no
shell job entry. `nohup` is **not** what makes this work: Node re-arms the hangup signal to
its default even when it inherits "ignore", and Codeman handles that signal with a graceful
shutdown, so a delivered hangup would still stop the server.
- `--stop` verifies the process still looks like a Codeman server before signalling it,
because process ids get recycled.
**It refuses to start a second server on the same data directory.** Two servers sharing a
tmux socket attach to each other's live sessions.
## Installing as a service
```bash
codeman service install # systemd user unit on Linux, LaunchAgent on macOS
codeman service status
codeman service uninstall
```
The installer's final menu offers this too.
Notable behaviours:
- **Your PATH is baked into the unit.** launchd hands a job
`/usr/bin:/bin:/usr/sbin:/sbin`, which finds neither a Homebrew or nvm `node` nor `tmux`
or `claude`. This is the single most common cause of a hand-written unit that starts and
immediately dies.
- **`CODEMAN_PASSWORD` is never written into the unit file.** Add it yourself if the service
needs authentication.
- **It refuses when a server is already running** on that data directory, for the same reason
detached mode does.
- **It verifies rather than assumes.** `launchctl load` and a clean spawn are both silent
about a server that starts and immediately exits, so the parent polls until the child
answers or dies.
On Linux, if you want the service running while you are not logged in:
```bash
loginctl enable-linger $USER
```
### Writing the unit by hand
**Linux (systemd user unit):**
```bash
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
[Unit]
Description=Codeman Web Server
After=network.target
[Service]
Type=simple
ExecStart=$(which node) $HOME/.codeman/app/dist/index.js web
Restart=always
RestartSec=10
[Install]
WantedBy=default.target
EOF
systemctl --user daemon-reload
systemctl --user enable --now codeman-web
loginctl enable-linger $USER
```
**macOS (LaunchAgent):**
```bash
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
"http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>Label</key>
<string>com.codeman.web</string>
<key>ProgramArguments</key>
<array>
<string>$(which node)</string>
<string>$HOME/.codeman/app/dist/index.js</string>
<string>web</string>
</array>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>StandardOutPath</key>
<string>/tmp/codeman.log</string>
<key>StandardErrorPath</key>
<string>/tmp/codeman.log</string>
</dict>
</plist>
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
Prefer `codeman service install` where you can. It handles the PATH problem for you.
## Logs
```bash
journalctl --user -u codeman-web -f # systemd
tail -f ~/.codeman/web.log # detached mode
log stream --predicate 'process == "node"' # macOS, noisy
```
## Updating
| Install route | Update with |
| ------------- | ------------------------------------------------------------------------ |
| Installer | Re-run the one-liner, or **App Settings → System → Updates**. |
| npm | `npm update -g aicodeman` |
| git clone | `git pull && npm install && npm run build`, then restart. |
### The in-app updater
**App Settings → System → Updates**, for git-clone installs supervised by systemd or
launchd. npm installs report as non-updatable, and an unsupervised install is told to
restart manually.
The interesting part is that the update restarts the very process running it. So the real
work runs in a **detached script that outlives the restart** and writes progress to a status
file, which the browser polls across the connection drop. A dirty tree is stashed rather
than discarded.
### After updating
Sessions are unaffected: they live in tmux and the server reattaches. If the UI looks stale,
reload; on iOS Safari, close the tab completely and reopen.
## Running two instances
The data directory and the tmux socket are process wide, so a second server on the defaults
will discover and attach the first one's sessions. Scope both together:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
Service unit names are instance-scoped too, so a beta instance can be installed as its own
service without colliding with the main one. `CODEMAN_DATA_DIR` and `CODEMAN_TMUX_SOCKET`
exist for the rare case where they need to differ, but setting only one of them recreates
exactly the problem you were avoiding.
## The tunnel as a service
```bash
systemctl --user enable codeman-tunnel
loginctl enable-linger $USER
```
Or the toggle in **App Settings → System → Remote access**. See
[Remote Access](Remote-Access).
## Health checks
```bash
curl -s localhost:3000/api/status | jq '.version, .uptime'
codeman web --status
codeman doctor
```
Add `-k` and the `https://` URL on an HTTPS install.
## Read next
- [Installation](Installation) - the routes and what each supports.
- [Remote Access](Remote-Access) - exposing it once it stays up.
- [Troubleshooting](Troubleshooting) - when it does not.
+120
View File
@@ -0,0 +1,120 @@
# Security
The honest version first: **Codeman's dashboard is a remote code execution surface, by
design.** It starts agents with permission prompts skipped by default, so anyone who can
reach it can run arbitrary code as your user, on your machine. Every protection in Codeman
exists to control who that is.
That is not a flaw to be fixed. It is what "run my coding agent for me" means. The job is to
make sure the set of people who can reach it is exactly the set you intended.
## The default is safe
A bare `codeman web` binds `127.0.0.1`. Only processes on that machine can reach it, which
is why shipping with no password by default is defensible. Everything risky starts when you
expose it.
## Hardening checklist
In order of how much they matter:
1. **Do not expose it without `CODEMAN_PASSWORD`.** Binding a non-loopback host without one
starts, but warns loudly. A tunnel refuses outright unless you acknowledge the exposure
in the UI.
2. **Prefer Tailscale over a public tunnel.** Keeping the loopback bind and putting a
private network in front of it removes the public attack surface entirely, and gives you
real HTTPS. See [Remote Access](Remote-Access).
3. **Use a long password.** It is the only thing between a public URL and your shell.
4. **Consider the permission mode.** **App Settings → Agents & CLIs → Claude → Startup
Mode** can switch new sessions from skip-prompts to Anthropic's classifier-guarded `auto`
mode, to normal prompting, or to an explicit allowed-tools list.
5. **Use Docker cases for untrusted work.** If you are pointing an autonomous loop at a repo
you did not write, [Docker Cases](Docker-Cases) gives it its own filesystem and network
for the cost of one checkbox.
6. **Keep it updated.** Browser-driven attack paths were closed in 0.9.x and hardening is
ongoing.
## What protects what
These run on **every** request, including on a default no-password loopback install:
| Layer | What it stops |
| ---------------------------- | ----------------------------------------------------------------------------------------------- |
| **Host-header allowlist** | DNS rebinding. A domain rebound to `127.0.0.1` is rejected before any handler runs. Add your own domains with `CODEMAN_ALLOWED_HOSTS`. |
| **Cross-site Origin guard** | CSRF on state-changing requests. A *missing* Origin is allowed so curl, the CLI, and hooks keep working; a foreign or opaque one is rejected. |
| **Raw `text/plain` bodies** | The CORS simple-request CSRF vector, where a cross-site form could smuggle JSON into a write route with no preflight. |
| **WebSocket origin check** | Cross-site WebSocket hijacking. The terminal upgrade closes with code `4003` on failure. |
| **Output escaping** | Stored XSS from agent-derived strings: tool names, command arguments, subagent descriptions. |
| **Security headers** | A strict content security policy, `nosniff`, frame options, and HSTS over HTTPS. CORS is reflected only for loopback origins. |
When authentication is enabled:
| Layer | Behaviour |
| ------------------- | ------------------------------------------------------------------------------------------------ |
| **HTTP Basic** | `CODEMAN_USERNAME` (default `admin`) and `CODEMAN_PASSWORD`. |
| **Session cookie** | A 256-bit opaque token validated server side, so it cannot be forged offline. 24 hours, extended on activity, with a device-context audit trail. |
| **Rate limiting** | Ten failed attempts per IP produce a `429` with a 15 minute decay. A correct password or valid cookie recovers immediately even under attack, which matters because all tunnel traffic shares one loopback address. |
| **QR auth** | Single-use 60-second tokens with their own separate rate limiter, so a mistyped password cannot lock out QR login. |
| **Hook endpoints** | The hook and telemetry endpoints skip Basic auth because they are called from localhost by the CLI, but when auth is on, that bypass additionally requires a per-instance hook secret. |
## File access
Three separate file surfaces, each confined differently, because a single shared rule would
be wrong for at least one of them:
| Surface | Rules |
| -------------------- | -------------------------------------------------------------------------------------------- |
| **File Viewer** | Real path resolution before boundary checks, so symlinks cannot escape. Sensitive trees blocked. Edit mode adds an extension allowlist, a size cap, `.git` denial, and optimistic concurrency. It never creates files. |
| **Attachments** | An id-based registry, so browser requests never carry absolute paths. The magic-link scanner is prompt-injectable by nature and is therefore force-confined to the session's workspace. Extension allowlist, not a blocklist. |
| **Path picker** | Its own root allowlist rather than the workspace confinement. In multi-user mode a non-admin gets only their own user space, because per-user spaces live inside the home directory. |
Downloads block sensitive paths outright (`.env`, credentials files, `~/.ssh`, AWS
credentials), and SVG and HTML are served as downloads with `nosniff` so they cannot execute
in the page.
## Supply chain and isolation
- Security-sensitive transitive dependencies are pinned to patched versions, and lockfile
integrity is checked on every push and pull request: every entry must resolve to the public
registry with a hash.
- Public assets are scanned for NUL bytes and syntax-checked in CI.
- `CODEMAN_INSTANCE` scopes the tmux socket and the data directory together, so two
instances never attach each other's live sessions.
## What Codeman does not protect against
Stated plainly, because a security page that only lists strengths is not useful:
- **Multi-user mode is not a sandbox.** It separates workspaces. Every session still runs as
the same OS account, so a determined user's agent can reach another user's files. For real
isolation, pair users with Docker cases or run separate instances under separate OS
accounts.
- **An agent you gave shell access can do anything you can.** Permission modes narrow this;
they do not remove it.
- **A tunnel makes your machine reachable from the internet.** The password is the whole
boundary. Treat it accordingly.
- **Codeman cannot detect your own loopback reverse proxy**, which is why the hook-endpoint
bypass requires a secret unconditionally when auth is on.
- **The agent CLIs have their own trust models.** Pi's project trust executes repo-local
TypeScript, for instance. See [Agent CLIs](Agent-CLIs).
## Privacy
No telemetry, no analytics, no phone-home. Codeman's only network traffic is between your
browser and your server. Your agent CLI's traffic is its own, on your account.
Two features send data outward, both off by default and both stated where they appear: voice
dictation through your Claude login, and the Read My Mind prediction call.
## Reporting a vulnerability
**Never in a public issue.**
[SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md) has the
private disclosure process and the current list of known limitations.
## Read next
- [Remote Access](Remote-Access) - the safe ways to expose it.
- [Multi-User Mode](Multi-User-Mode) - what it does and does not separate.
- [Docker Cases](Docker-Cases) - real isolation for untrusted work.
- [`docs/security-architecture.md`](https://github.com/Ark0N/Codeman/blob/master/docs/security-architecture.md) - the complete model.
+173
View File
@@ -0,0 +1,173 @@
# Settings Reference
Two settings surfaces, and the rule that explains why a setting you changed on your laptop
did not follow you to your phone.
| Surface | Scope | Opened from |
| ------------------- | ------------------------------ | ---------------------------- |
| **App Settings** | Global, this Codeman install. | The header gear. |
| **Session Options** | One session. | The session's tab. |
App Settings is a single scrolling document with a rail acting as a table of contents;
clicking a rail entry scrolls rather than switching. Session Options genuinely switches
panels.
## Per-device versus synced
Some settings live on the server and follow you to every device. Others are stored in the
browser and stay put. This is deliberate, not an oversight: your phone wants a different
font size, a different keyboard bar, and a different set of header buttons than your
desktop.
| Category | Examples |
| ----------------------- | ------------------------------------------------------------------------------- |
| **Per-device, local** | Skin, WebGL renderer, local echo, CJK input, extended keyboard bar, File Viewer and Cron header buttons. Never sent to the server at all. |
| **Per-device policy** | Most `show*` toggles, plan usage chip, language. Stored server-side, but a device only takes the server value when it has no local one of its own. |
| **Synced** | Models, effort, CLI options, notification preferences, voice settings, display name, the agent skill and approvals toggles. |
The practical rule: **appearance and input are per device, behaviour is shared.** If a change
did not follow you, it is in one of the first two rows, and you change it again on that
device.
## App Settings
### Updates
Current version, a manual check, and the in-app updater. Covers git-clone installs
supervised by systemd or launchd; npm installs report as non-updatable. See
[Running As A Service](Running-As-A-Service).
### Terminal & Input
| Setting | Default | Notes |
| ----------------------------- | -------------------- | --------------------------------------------------------------------- |
| Local Echo | On for touch devices | Paints keystrokes locally and flushes on Enter. See [Input And Voice](Input-And-Voice). |
| CJK Input | Off | IME composition through a dedicated text field. |
| Extended Keyboard Bar | Per device | Which accessory bar phones get. Shell sessions override it while they are active. |
| Wheel Scrolls Local History | Off | Keeps the wheel on the local buffer instead of forwarding it to the CLI. |
| Auto Copy Selection | Off | Copies highlighted terminal text to the clipboard the moment you finish selecting it. Ctrl+C still copies on demand. |
| WebGL Renderer | On | With a GPU-stall watchdog that falls back to DOM rendering. |
| Gesture Control | Off | Camera hand tracking. Also needs `CODEMAN_GESTURE=1` on the server. |
### Header & Panels
Chips for every optional header control, with a live preview of the resulting header:
Run, Font Size, System Stats, Redraw Terminal, Response Viewer, Away Digest, Session
Manager, Attachments, File Viewer, Multi-monitor, Plan Usage, Lifecycle Log, Monitor,
Project Insights, File Browser, Subagents, Approvals Inbox, Read My Mind, Ultracode Agents,
Ultracode Windows, Cron.
Most default to off. The stock desktop header is system stats, File Viewer, and the gear.
New header controls never appear on phones.
This section also holds background-agent tracking, including whether to track agents for
every session or only the active tab.
### Appearance
| Setting | Notes |
| ---------------------- | ----------------------------------------------------------------------------------------- |
| Skin | Theme palettes, light ones included. Applied before first paint, so no flash of the wrong theme. |
| Entrance Animations | Per-surface animation styles for tabs, terminals, windows, and lineage lines. All default to the legacy no-animation behaviour. |
| Display Name | Your name in the UI. Cosmetic only; it never renames the package, CLI, API, or storage. |
| Interface Language | English or Simplified Chinese. Per device. |
| Session List Layout | Header tab strip (default) or a collapsible left sidebar. See [The Dashboard](The-Dashboard#session-list-layout). |
| Tall Tabs | Taller tab strip. |
| Pop-out Button on Tabs | Adds the detach control to tabs, with a per-tab override. |
| Spawn Lineage Lines | Arcs from a parent tab to sessions it spawned. Desktop only, on by default. |
| Overview Home Screen | The phone home screen. On by default. |
### Models
Claude model cards, the 1M context window switch, and the thinking effort segment. The cards
and the switch compose into one model choice, so there is no separate "which one wins"
question.
Model and effort are both **soft defaults**: the model is written into the case's
`.claude/settings.local.json` and effort is passed at start, so `/model` and `/effort`
inside a session override them at any time.
### Agents & CLIs
| Setting | Notes |
| -------------------------------- | -------------------------------------------------------------------------------------------- |
| Startup Mode | Claude's permission mode for new sessions. Default skips prompts; `auto` uses Anthropic's classifier-guarded mode; `normal` prompts; or give an explicit allowed-tools list. |
| Allowed Tools | The list used by the explicit mode. |
| Ralph / Todo Tracker | Enables the Ralph loop surfaces. |
| Agent Teams | Experimental teams. Also needs the CLI's own environment flag. |
| Codeman Agent Skill | Injects the agent skill into new Claude sessions per case. Off by default. See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent). |
| Remote auto-reconnect | Reattaches dropped remote SSH sessions. On by default. |
| Nice priority / value | Runs agent processes at a lower CPU priority. |
| Bypass approvals and sandbox | Pi's project trust. Read [Agent CLIs](Agent-CLIs) before enabling. |
| Animated status effects | Cosmetic. |
### Notifications
Master toggle, browser notifications, push subscription, audio alerts, and the idle
threshold that decides when a quiet session counts as needing you. See
[Notifications And Approvals](Notifications-And-Approvals).
### Voice
Active provider and the engine behind it, insert mode, language, domain keywords to bias
recognition, the Deepgram API key, and the opt-in switch for transcribing through this
server's Claude login, with its live credential status. See
[Input And Voice](Input-And-Voice).
### Shortcuts
Rebinding for the shortcut registry. See [Keyboard Shortcuts](Keyboard-Shortcuts).
### System
`CLAUDE.md` template for new cases, default working directory, the image watcher, and
Cloudflare tunnel controls including the tunnel and upload URLs. In multi-user mode, the
**Users** administration entry is injected here.
## Session Options
Per session, from the tab.
| Panel | Contains |
| ---------------- | ------------------------------------------------------------------------------------------- |
| **Respawn** | Auto-resume on usage limit, the respawn cycle configuration, presets, duration. See [Keeping Agents Running](Keeping-Agents-Running). |
| **Session** | Name, working directory, environment overrides, per-tab pop-out override. |
| **Ralph / Todo** | Loop configuration, iteration and todo caps, circuit breaker reset. See [Autonomous Loops](Autonomous-Loops). |
| **Summary** | What this session has done: tokens, activity, run summary. |
Panels that only make sense for Claude are hidden for other run modes rather than shown and
failing.
## Environment variables
Some things are configured before the server starts, not in the UI:
| Variable | Effect |
| ----------------------------------- | ---------------------------------------------------------------------- |
| `CODEMAN_PORT` | Listen port. |
| `CODEMAN_HOST` | Bind address. Loopback by default. |
| `CODEMAN_PASSWORD` / `CODEMAN_USERNAME` | HTTP Basic credentials. Username defaults to `admin`. |
| `CODEMAN_ALLOWED_HOSTS` | Extra Host and Origin allowlist entries for a reverse proxy. |
| `CODEMAN_INSTANCE` | Scopes the data directory and tmux socket together. Required for a second instance. |
| `CODEMAN_MULTIUSER` | Enables multi-user mode. |
| `CODEMAN_GESTURE` | Makes gesture control available to be enabled. |
| `CODEMAN_DOCKER_BRIDGE_HOOKS` | Lets in-container hooks reach the host on a loopback bind. |
| `CODEMAN_FILE_PICKER_ROOTS` | Extra roots for the path picker. |
| `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK` | Acknowledges exposing the server with no password. |
## Gotchas
- **A setting that did not sync is per device.** Change it again on that device.
- **The plan usage chip and its telemetry exporter are one setting.** Enabling the chip
without the exporter would leave it blank forever, so it is deliberately not separable.
- **Toggling a header button does nothing on a phone.** Phones deliberately ignore most of
the header chips.
- **Enabling a feature does not retroactively configure existing sessions.** The agent skill
injection, for instance, applies at session creation.
## Read next
- [The Dashboard](The-Dashboard) - what each control does once visible.
- [Keeping Agents Running](Keeping-Agents-Running) - the Respawn panel in depth.
- [Agent CLIs](Agent-CLIs) - model, effort, and permission modes.
+206
View File
@@ -0,0 +1,206 @@
# The Dashboard
What the interface is telling you, and which parts of it are hidden until you turn them on.
Most of Codeman's UI is **opt-in**. A stock install shows a deliberately small header, and a
feature you read about here may simply not be on screen yet. Where that is the case, this
page says so and names the setting.
![Codeman dashboard](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/codeman-tour-20260724.png)
## Layout
| Region | What lives there |
| ------------------ | -------------------------------------------------------------------------------------- |
| **Header, left** | The "C" logo (goes home) and the session list, unless you moved it to the sidebar. |
| **Header, right** | Status chips and panel buttons, most of them off by default. |
| **Center** | The terminal for the active session, or the home screen when nothing is selected. |
| **Bottom toolbar** | Run, Stop, Run Shell, the case picker, and the instance counters. |
| **Overlays** | Panels and modals: Respawn, Cron, Subagents, File Viewer, Settings. |
## Session list layout
The session list lives in the header as a horizontal strip by default. With a lot of
sessions open that strip stops being scannable, so **App Settings → Appearance → Tabs →
Session List Layout** can move it into a vertical sidebar on the left instead.
| Layout | Behaviour |
| -------------------- | --------------------------------------------------------------------------------- |
| **Header tab strip** | The default. Wraps to a second row on desktop, scrolls sideways on a phone. |
| **Left sidebar** | A vertical list with a filter box and a live session count. `Alt+B` collapses it to a narrow rail that keeps the status dots and task badges visible. On a phone it is an off-canvas drawer rather than a docked rail. |
It is the same list either way, just re-hosted: tab order, drag-to-reorder, the `Alt+1`
to `Alt+9` numbers and every status colour below behave identically in both. The setting is
per device, so a sidebar on your desktop does not force one onto your phone.
## Session tabs
One tab per session, in your order, and that order syncs across your devices.
**Status is carried by the dot and the tab's own styling:**
| Look | Meaning |
| ----------------------------- | ----------------------------------------------------------------------- |
| Green dot | Alive, not currently working. |
| Pulsing green dot with a ring | Working on a turn. |
| Yellow tab, blinking | The agent is waiting for input from you. |
| Red tab, blinking | A question or permission prompt is blocking the session. |
| No dot | The session is not running. |
![Tab alerts](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/tab-alerts-20260815.png)
The alert states are steady colour with a pulse layered on top, not a blink between the
alert colour and nothing, so a tab that needs you looks like it needs you at every point in
the cycle. They survive a page reload: the state is re-seeded from the server on load, so
reloading while a permission prompt is blocking does not lose the red tab.
**Navigation:**
| Action | Keys |
| ------------------------------- | ------------------------------------------------------- |
| Jump to tab N | `Alt+1` to `Alt+9` (the number on the tab) |
| Next / previous | `Ctrl+Tab`, `Alt+[`, `Alt+]` |
| Move the active tab | `Ctrl+Shift+{`, `Ctrl+Shift+}` |
| Close | `Ctrl+W` |
| Find any session, open or past | `Ctrl+K` (also `Cmd+K` and `Alt+K`) |
Tabs can also be dragged to reorder.
On phones the strip scrolls horizontally instead of wrapping, and the active tab is always
scrolled into view. It is not reordered to the front, so the `Alt+N` numbering stays stable.
### Lineage arcs
When one session spawns another (an agent starting a worker through the API), Codeman draws
a coloured arc under the strip connecting parent to child, with one colour per child. It is
how a fan-out of eight workers stays readable.
Desktop only, and on by default. Turn it off in **App Settings → Appearance**. Arcs are
skipped for tabs scrolled out of the strip.
## Header controls
The right side of the header. Almost all of these are off until you enable them in
**App Settings → Header & Panels**.
| Control | Default | What it does |
| ---------------------- | ------------------ | ------------------------------------------------------------------------------- |
| Connection dot | Always on | SSE connection health. Green is connected. |
| Font size `-` / `+` | Always on | `Ctrl +` / `Ctrl -` do the same. |
| CPU / MEM bars | On | Server resource use. |
| File Viewer | On | Toggles the file browser panel. |
| Settings gear | Always on | App Settings. |
| Plan usage chip | On, desktop only | Live Claude subscription usage. Claude-only, and needs its telemetry exporter, which the same setting installs. |
| Session Manager | Off | The full session list, live and historical. |
| Approvals bell | Off | Cross-session queue of prompts waiting on a human. Appears only when the count is above zero. Never shown on phones. |
| Read My Mind 🧠 | Off | Predicts your next prompt for this case. Claude-only. |
| Attachments | Off | Registered external files. |
| Away Digest | Off | What happened while you were gone. |
| Last Response | Off | Readable view of the agent's last answer, useful on phones. |
| Ultracode / Workflow | Off | Live workflow-run agents. |
| Notifications | Off | Notification history and settings. |
| Lifecycle Log | Off | Session start, exit, and kill audit trail. |
| Cron ⏰ | Off | Scheduled jobs. |
| Multi-monitor | Off, macOS | Opens a window spanning every display. |
| Tunnel indicator | When a tunnel runs | Cloudflare tunnel status. |
| Admin panel | Multi-user only | User administration. |
New header controls never appear on phones. Phone layout is deliberately minimal and is
covered in [Mobile Guide](Mobile-Guide).
## Connection state
The dot in the header is the quick read. Two louder surfaces exist because a cached page
with no server behind it used to look identical to a page with no sessions:
- **A full-screen overlay** when the page has never loaded server state. There is nothing
behind it worth preserving.
- **A banner** when the connection drops after state had loaded, so your scrollback stays
readable.
Both wait about 2.5 seconds before appearing, so a deploy that restarts the server does not
flash a warning at you every time. If the browser reports itself offline, the grace period
is skipped.
There is also a watchdog for the case where the connection stops delivering without
erroring. If the server's heartbeat stops arriving, Codeman reconnects on its own rather
than sitting on a green dot showing frozen data.
## The terminal
A real terminal: xterm.js in the browser, a real PTY on the server, tmux in between. Full
TUIs render correctly.
Worth knowing:
- **Scrollback.** Agent/TUI sessions pull their entire tmux scrollback on first open.
Shell sessions open from a bounded recent tail so a large transcript cannot stall tab
switching; press **Load full history** to pull the rest explicitly. Ordinary Shell scrolling
stays within the bounded browser buffer so dragging upward remains responsive.
- **Wheel and touch scrolling** are forwarded into Claude's own transcript on recent Claude
versions, so the wheel scrolls the conversation rather than the terminal. `Shift+Wheel` is
always local scrollback. Other CLIs scroll locally.
- **Selection copy.** `Ctrl+C` copies when text is selected and interrupts when it is not.
`Ctrl+Shift+C` always copies.
- **Zero-lag input.** On touch devices, keystrokes paint locally before the round trip. See
[Input And Voice](Input-And-Voice).
- **Renderer.** WebGL by default, with a watchdog that falls back to DOM rendering if the
GPU stalls. `?nowebgl` forces DOM rendering for one page load.
## The home screen
With no session selected you get the welcome screen: run buttons for the CLIs Codeman
found, a QR code when a password is set, cross-session search, and **Resume Conversation**,
which lists past sessions including Claude conversations started outside Codeman entirely.
Two extras depending on the device:
- **Desktop, wide windows**: your open tabs appear as a rail docked to the left edge, in tab
order, with created and last-active stamps. It needs at least 1180px of width; below that
it is hidden so it cannot overlap the search panel.
- **Phones**: tapping the "C" logo gives a session overview instead: NEEDS YOU first, then
current sessions, then past ones. On by default.
## Panels
| Panel | Opened from | Covered in |
| ---------------- | --------------------------------- | ---------------------------------------------------------------- |
| Respawn | Session Options | [Keeping Agents Running](Keeping-Agents-Running) |
| Ralph | Session Options | [Autonomous Loops](Autonomous-Loops) |
| Orchestrator | Toolbar | [Autonomous Loops](Autonomous-Loops) |
| Cron | Header ⏰ (opt-in) | [Cron Jobs](Cron-Jobs) |
| Subagents | Automatic while agents run | [Watching Agents Work](Watching-Agents-Work) |
| Ultracode | Header (opt-in) | [Watching Agents Work](Watching-Agents-Work) |
| File Viewer | Header | [Working With Files](Working-With-Files) |
| Attachments | Header (opt-in) | [Working With Files](Working-With-Files) |
| Approvals | Header bell (opt-in) | [Notifications And Approvals](Notifications-And-Approvals) |
| App Settings | Header gear | [Settings Reference](Settings-Reference) |
Session-specific configuration lives in **Session Options**, reachable from the tab. App
Settings is global; Session Options is per session.
## Search and the session palette
`Ctrl+K` opens the session palette: every session, live or historical, filtered as you
type. Picking a past one resumes its conversation.
The search box on the home screen is wider in scope. It federates over session metadata,
run-summary events, and attachment history, filtered by type, case, status, and date. It
does substring matching over data already in memory, with no regex and no filesystem reads,
so it is fast and cannot be turned into a traversal.
## Appearance
**App Settings → Appearance** carries the theme skins, including light ones. The choice is
applied before the first paint, so there is no flash of the wrong theme on load.
The same section has the entrance animations for tabs, terminals, agent windows, and
lineage lines. All of them default to the legacy no-animation behaviour, so an untouched
install animates nothing.
## Read next
- [Keyboard Shortcuts](Keyboard-Shortcuts) - the full list, and how to rebind.
- [Settings Reference](Settings-Reference) - every setting, and why some follow you across devices and others do not.
- [Mobile Guide](Mobile-Guide) - what changes on a phone.
- [Watching Agents Work](Watching-Agents-Work) - subagent windows and workflow runs.
+290
View File
@@ -0,0 +1,290 @@
# Troubleshooting
Symptom first. Find the line that matches what you are seeing.
Before anything else, check what version you are on and whether the problem is already
fixed:
```bash
codeman --version
codeman doctor
```
## Installing and starting
### `Failed to start claude: error: posix_spawnp failed` on macOS
node-pty ships its macOS `spawn-helper` without the executable bit, and macOS launches
every PTY through it. Codeman detects this and repairs it on the first failure, so updating
usually fixes it outright. To repair by hand on a clone install:
```bash
npm run fix:node-pty
```
It is a `chmod`, not a rebuild, so it does not need Xcode command line tools. The helper
lives in `prebuilds/darwin-<arch>/`, not `build/Release/`, which does not exist on macOS.
Linux never sees this.
### `tmux: command not found`
The installer asks before installing packages and remembers a declined answer. Install tmux
and start again. There is no tmux-free mode: sessions live in tmux.
### The port is already in use
```bash
codeman web --port 8080 # or set CODEMAN_PORT
```
If you believe nothing is on 3000, check for a Codeman you already started:
```bash
codeman web --status
```
### The terminal area is blank, and the console mentions a missing vendor file
Clone installs build the vendored xterm addon bundles in `postinstall`. If `npm install`
was interrupted or run with `--ignore-scripts`, those bundles are missing:
```bash
npm install
```
They are intentionally not committed to the repository.
### `Case path not found` when clicking Run
The case points at a directory that no longer exists, usually because it was deleted or
moved outside Codeman. Re-link the case, or create it again.
### The server starts but nothing is reachable
That is the default behaviour, not a failure. Codeman binds `127.0.0.1`. See
[Remote Access](Remote-Access).
## Reaching the interface
### The dashboard will not load from another device
Check, in order: the bind (loopback by default), a firewall, and then
[Remote Access](Remote-Access) for a supported way to expose it.
### `403 host not allowed`
The Host header is not in the allowlist, which is the DNS-rebinding guard doing its job. Add
your domain:
```bash
CODEMAN_ALLOWED_HOSTS='codeman.example.com,.internal.example.com'
```
A leading dot matches subdomains.
### The page loads but the terminal never connects
The terminal is a WebSocket. Behind a reverse proxy, the upgrade must be forwarded. The
upgrade also runs the Host and Origin checks and closes with code `4003` when they fail.
### The UI looks stale after updating
The app shell is cached by a service worker, and static assets are served with a long cache
lifetime. `index.html` is not cached, and every asset reference is version-stamped, so a
normal reload picks up a new build.
Two exceptions worth knowing:
- **iOS Safari** can keep serving old JavaScript until the tab is fully closed, not just
reloaded. Close the tab and reopen it.
- If you edit files in dev, changes to `index.html` need a server restart. Changes to `.js`
and `.css` do not.
### A full-screen "cannot reach the server" overlay appears
The server is genuinely unreachable, or the connection dropped. Codeman waits about 2.5
seconds before showing it, so a quick restart does not flash it. Retry re-arms both the
event stream and the terminal socket.
## Sessions
### A session shows idle while it is clearly working
Update. Claude redraws its prompt roughly once a second throughout a turn, and older idle
detection treated that as the end of the turn, flipping working sessions to idle a couple of
seconds in. Current versions confirm against the actual screen before believing it.
### A session is stuck showing busy
For non-Claude CLIs, idle detection is output-based and coarser by necessity: those CLIs
expose no hooks. A session that has genuinely gone quiet will settle. If it never does,
interrupt it (`Ctrl+C` with nothing selected).
### The agent asks about bypass permissions every time
That prompt comes from Claude Code, not Codeman. Codeman's default is to start with
permission prompts skipped, which is what the security model is built around. If you would
rather it prompted, change **App Settings → Agents & CLIs → Claude → Startup Mode**.
### Sessions vanished after a reboot
Expected. tmux does not survive a reboot, so the sessions are gone. Conversations are not:
Claude transcripts persist, so the welcome screen's **Resume Conversation** list can pick
them back up.
### A session restarts, then refuses to restart again
That is the PTY-exit circuit breaker. Repeated rapid PTY exits trip it, and it blocks
automatic restarts so a broken configuration does not spin forever. Reset it explicitly from
the session's controls. Reattaching does not clear it, deliberately.
### Sessions I did not create appeared, or my session resized itself
Two Codeman servers are running against the same data directory and tmux socket. The second
one discovers and attaches the first one's sessions. Give each instance its own scope:
```bash
CODEMAN_INSTANCE=beta CODEMAN_PORT=5000 codeman web
```
`codeman web -d` and `codeman service install` both refuse to start a second server on one
data directory for exactly this reason.
## The terminal
### I cannot scroll back through history
Scrollback behaviour depends on the CLI, and Codeman adjusts what it strips per mode.
Things to try:
- `Shift+Wheel` always scrolls the local buffer, whatever else is going on.
- On Claude sessions with a recent CLI, the wheel is forwarded into Claude's own transcript,
so it scrolls the conversation rather than the terminal buffer. That is intended.
- Scrolling to the very top pulls the full tmux scrollback again on demand.
### The wheel does nothing in a Codex session
Codex ignores the mouse reports that forwarding would send, so Codeman does not forward
there. Scrolling is local, and `Shift+Wheel` behaves the same way.
### `Ctrl+C` copies when I wanted to interrupt
With a selection, `Ctrl+C` copies. With no selection, it interrupts. Clear the selection
first, or use the **Stop** button. `Ctrl+Shift+C` always copies and never interrupts.
### I typed a prompt but nothing was sent
On touch devices, keystrokes are painted locally and flushed when you press Enter, so text
on screen has not necessarily reached the agent yet. Press Enter, or the phone toolbar's
**Enter** button.
If you are sending input over the API instead, your payload must end with `\r` or no Enter
is ever sent. The request still succeeds and the text sits unsubmitted in the composer. See
[Driving Codeman From An Agent](Driving-Codeman-From-An-Agent).
## Mobile
### The keyboard covers the terminal, or scroll position jumps
Update first; several rounds of fixes have gone into keyboard resize and scroll restoration.
### I cannot reach the rightmost tabs
The strip scrolls horizontally on phones and the active tab is scrolled into view
automatically. Swipe the strip itself. If a background render snaps you back, update.
### The space key does nothing on Android
A long-standing Android keyboard bug, fixed some time ago. Update.
### The keyboard will not close
Tap outside the terminal, or tap twice on inert terminal content. Tapping a control does not
dismiss it, by design.
## Agents and CLIs
### A CLI is installed but Codeman does not offer it
Codeman resolves binaries from the environment the **server** runs in.
```bash
codeman doctor
```
If it runs as a service, launchd gives the job a minimal PATH. `codeman service install`
bakes your PATH into the unit; a hand-written plist does not. Restart the server after
installing a new CLI.
### Hooks stopped working after switching to HTTPS
Hook callbacks have to accept the self-signed certificate. Recent versions self-heal
existing cases; if yours predates that, recreate the case so its hooks are rewritten.
### The model or effort I chose is not being used
Both are **soft defaults**, on purpose. The model is written into the case's
`.claude/settings.local.json` and effort is passed on the command line at start, so `/model`
and `/effort` inside the session override them at any time. Effort is deliberately never
passed as an environment variable, because that hard-locks it.
### Tab alerts and approvals never fire in one of my repos
That case is missing its hooks block. Recreating the case rewrites it.
## Docker and remote
### Docker sessions do not detect idle
On a loopback-only bind, a container cannot reach `127.0.0.1` on the host, so in-container
hooks have nothing to call. Set `CODEMAN_DOCKER_BRIDGE_HOOKS=1` to open a hooks-only
listener on the docker bridge gateway. Without it, idle detection falls back to output
watching.
### A rebuilt agent image still has old CLI versions
Always rebuild with `--no-cache`:
```bash
node scripts/build-agent-image.mjs --no-cache
```
A plain rebuild reuses the cached `npm install -g` layer and keeps the CLIs frozen at their
original versions while reporting success.
### A remote SSH session dropped and did not come back
A bounded-backoff watcher reattaches dropped sessions, and it is on by default. Intentional
kills are never revived. Check the host is reachable and that the remote tmux server is
still running.
## Gathering diagnostics
```bash
codeman doctor # dependency check
curl -s localhost:3000/api/status | jq # full app state
tmux -L codeman list-sessions # what tmux thinks is alive
journalctl --user -u codeman-web -f # service logs (Linux)
tail -f ~/.codeman/web.log # detached mode logs
```
On an HTTPS install, add `-k` to the curl commands and use the `https://` URL.
## Filing a good bug report
Open an [issue](https://github.com/Ark0N/Codeman/issues) with:
- OS and version.
- Install method: installer, npm, or git clone.
- `codeman --version`.
- Browser and version, if the problem is in the UI.
- Which CLI the session was running, and its version.
- What you did, what happened, what you expected.
Reports usually get a response within a day, and every release credits its reporters by
name.
Questions and setup help fit better in
[Discussions](https://github.com/Ark0N/Codeman/discussions). Security problems never go in a
public issue; see
[SECURITY.md](https://github.com/Ark0N/Codeman/blob/master/.github/SECURITY.md).
+72
View File
@@ -0,0 +1,72 @@
# Versioning
Codeman follows [semantic versioning](https://semver.org/). This page says what the version
number actually promises, which matters if you are building anything against Codeman.
## Covered by the version number
Breaking any of these after 1.0 requires a **major** bump:
1. **The CLI.** Command names, documented flags, and their behaviour. The npm package is
`aicodeman` and installs both the `aicodeman` and `codeman` commands; renaming either is
breaking.
2. **The HTTP API and SSE channel**, served under `/api/v1` with the uniform envelope and
conventional status codes. Endpoint paths, the envelope, `errorCode` values, and SSE event
names are all stable.
3. **Documented deployment environment variables**: `CODEMAN_PASSWORD`, `CODEMAN_USERNAME`,
`CODEMAN_HOST`, `CODEMAN_PORT`, `CODEMAN_INSTANCE`, `CODEMAN_ALLOWED_HOSTS`,
`CODEMAN_DATA_DIR`, `CODEMAN_TMUX_SOCKET`, plus the `--host`, `--port`, and `--https`
flags.
4. **The published `xterm-zerolag-input` library**, on its own independent version line.
Codeman reaching 1.0 says nothing about that package's version.
Additive changes are **not** breaking: new endpoints, new optional fields, new error codes,
new SSE events. Genuinely breaking API changes would ship under a new prefix rather than
changing `/api/v1`.
## Not covered
These can change in a minor or even patch release:
1. **The `~/.codeman/` state file formats.** Migrations are made on a best-effort basis and
have been done across renames, but the on-disk shape is not a contract. Do not write
tooling against it.
2. **Internal TypeScript modules.** The npm package is CLI-only. There is no stable library
entry point, and importing it programmatically is unsupported.
3. **Experimental and opt-in features**, whatever the app's version: gesture control, agent
teams, and anything labelled experimental in the UI or docs.
## Deprecation
- Additive changes are preferred over breaking ones.
- A covered surface slated for removal is deprecated first: it keeps working for at least one
minor release, with a runtime warning and a changelog note pointing at the replacement,
then is removed in the next major.
- Backwards-compatibility shims are kept until a major boundary.
## Releases
Releases are managed with changesets. Every release:
- Bumps the version and updates
[`CHANGELOG.md`](https://github.com/Ark0N/Codeman/blob/master/CHANGELOG.md).
- Publishes to npm as `aicodeman`.
- Cuts a GitHub release, tagged `codeman@X.Y.Z`.
- **Credits its contributors and bug reporters by name** in the release notes.
There is no fixed cadence. Patches ship when fixes are ready, which in practice is often.
## Which version am I on?
```bash
codeman --version
```
Or **App Settings → Updates**, which also checks for a newer one and can install it. See
[Running As A Service](Running-As-A-Service).
## Read next
- [HTTP API](HTTP-API) - the stable API surface itself.
- [Contributing](Contributing) - how changes get made.
- [`docs/versioning-policy.md`](https://github.com/Ark0N/Codeman/blob/master/docs/versioning-policy.md) - the authoritative statement.
+112
View File
@@ -0,0 +1,112 @@
# Watching Agents Work
Modern agents fan out. A single Claude session can be running six subagents, and the parent
terminal shows you almost none of it. Codeman surfaces that hidden work as live windows,
panels, and after-the-fact summaries.
Everything on this page is Claude-only. It reads Claude Code's transcripts and team state;
the other CLIs expose no equivalent.
![Subagent windows](https://raw.githubusercontent.com/Ark0N/Codeman/master/docs/images/subagent-windows-20260724.png)
## Subagent windows
When a Claude session spawns subagents, each one gets its own floating window with a live
transcript: what it was asked to do, what it is doing, and what it returned.
- Windows are draggable and resizable, and their positions persist across reloads.
- A connection line links each window to the session tab that spawned it, so with four
sessions running you can still tell whose worker is whose.
- Closing a window does not stop the subagent. It only stops you watching it.
This is the feature that makes a fan-out legible. Without it, a lead session that spawned
eight workers looks like a stalled terminal for several minutes.
## Session lineage arcs
The tab strip draws a coloured arc from a parent tab to any tab it spawned, one colour per
child. That covers the other direction of fan-out: not subagents inside one session, but
whole sessions started by an agent through the API.
Desktop only, on by default, and toggled in **App Settings → Appearance**. Arcs are skipped
for tabs scrolled out of view.
See [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) for the spawning side.
## Agent teams
Claude Code's experimental agent teams appear as teammates alongside subagents. Enable them
in the CLI's own environment:
```bash
CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1
```
and turn the per-case **Agent Teams** toggle on in the case settings gear.
Codeman watches the team directory and matches teammates to the session leading them.
Teammates are in-process threads rather than separate CLI processes, so they show up as
windows, not tabs.
Notes and the experiment log:
[`docs/agent-teams/`](https://github.com/Ark0N/Codeman/tree/master/docs/agent-teams).
## Ultracode and workflow runs
When Claude runs a Workflow, dozens of agents can be in flight at once. The completion
artifact for a run is only written at the **end**, so a live run would otherwise be
invisible until it finished. Codeman synthesizes the in-flight view from the transcripts and
lets the real artifact supersede it when it lands.
Two independent toggles, both off by default:
| Setting | Shows |
| ---------------------- | ----------------------------------------- |
| Ultracode panel | A docked panel listing the run's agents. |
| Ultracode windows | Floating windows, like subagents. |
Turning on either starts the watcher.
## Reading the answer, not the terminal
**Last Response** (header button, opt-in) renders the agent's last answer as scrollable text
rather than terminal output. It exists mostly for phones, where reading a long answer in a
terminal viewport is painful. **More** loads additional context.
## After the fact
| Surface | Answers |
| ------------------ | -------------------------------------------------------------- |
| **Away Digest** | What happened while I was gone? |
| **Run summary** | What did this run actually do? |
| **Lifecycle log** | When did sessions start, exit, or get killed, and why? |
| **Token stats** | What did it cost? |
The Away Digest aggregates the lifecycle log, run summary events, live sessions, token
statistics, and recent subagents into one view. It is the right first thing to open in the
morning after an overnight run.
All of these header buttons are opt-in: **App Settings → Header & Panels**.
## Performance
The design target is 20 sessions and 50 agent windows at 60fps. If you routinely run more
than that, expect the browser rather than the server to be the limit, and close windows you
are not reading.
## Gotchas
- **A session pointed at a relocated Claude config directory goes blind here.** Transcripts
written outside `~/.claude/projects` are invisible to the watchers, so subagent windows,
the ultracode panel, the response viewer, and Read My Mind all stop working for that
session. Symlink `projects` back into the shared tree to fix it. See
[Agent CLIs](Agent-CLIs).
- **Closing a window does not cancel the agent.** Nothing on this page controls agents; it
observes them.
- **Windows are opt-in for ultracode, automatic for subagents.**
## Read next
- [The Dashboard](The-Dashboard) - where these surfaces live.
- [Driving Codeman From An Agent](Driving-Codeman-From-An-Agent) - the other kind of fan-out.
- [Autonomous Loops](Autonomous-Loops) - the loops that generate this much activity.
+101
View File
@@ -0,0 +1,101 @@
# Web Tabs
Open any dashboard you run, Grafana, Uptime Kuma, Portainer, a status page on port 4000, as
a tab beside your agent sessions. Codeman becomes one mission control instead of Codeman
plus a pile of browser tabs.
A web tab is **not a session**. There is no PTY, no tmux, and no respawn behind it, the same
way a docker case is not a run mode.
## Adding one
1. Click the chevron next to **Run**.
2. Under **Web / URL**, pick **Add URL**.
3. Name it, paste the URL, optionally hit **Test**, and **Save**.
It opens immediately and appears in the dropdown from then on. Web tabs share the tab strip
with sessions, continue the same `Alt+1` to `Alt+9` numbering, and carry a globe icon so
they never read as a running agent.
**Closing a tab is not deleting it.** The tab's `x` closes; the `x` on its **dropdown row**
deletes the saved dashboard. Each dropdown row also has a gear for editing the URL.
Switching tabs does not reload a dashboard. Frames stay alive in the background, so one that
took a while to authenticate is still there when you come back. Past six live frames, the
least recently viewed is dropped to bound memory.
## Why dashboards are proxied
A plain cross-origin iframe fails three ways at once in the setup Codeman actually ships in:
| Blocker | What happens |
| ------------------- | ----------------------------------------------------------------------------------------------- |
| **Mixed content** | Production is HTTPS, and browsers hard-block `http://` iframes on an HTTPS page. No override, and none at all on iOS Safari. |
| **Framing refusal** | Grafana, Portainer, Home Assistant and many others send `X-Frame-Options: DENY`. |
| **Codeman's CSP** | `default-src 'self'` blocks a cross-origin frame before it starts. |
So by default the dashboard is served **through Codeman's own origin**: the browser loads a
path on Codeman, and Codeman relays to the dashboard, stripping the framing refusal,
rewriting redirects, cookies and root-absolute URLs, and relaying WebSockets so live panels
still update.
A useful side effect: the dashboard is fetched **by the Codeman server**, so a tailnet-only
or localhost-only dashboard works from any device that can reach Codeman, including a phone
that is not on your tailnet.
There is also a `direct` mode, a plain cross-origin iframe, which is cheaper but only works
for an HTTPS dashboard that permits framing.
## The Test button, and what it does not test
**Test** probes from the server and tells you which mode applies. It verifies
**server-to-upstream reachability and nothing else**. It does not exercise the browser
sandbox, cookies, CORS, CSP, or any reverse proxy in front of Codeman.
A passing Test does not guarantee the embedded page renders.
## The sandbox, and when to turn it off
Because a proxied dashboard is served from Codeman's own address, the browser considers it
same-origin with Codeman. Unchecked, its JavaScript could read the Codeman page and call the
API that spawns agents.
So the frame is sandboxed **without** same-origin access by default. The page runs in an
opaque origin: it cannot touch Codeman, and it gets no cookies or local storage of its own.
Unchecking **Open sandboxed** grants a real origin. Do that only for a dashboard you fully
trust, and only when you need it, which in practice means one with its own login that stores
a session in a cookie.
Either way, Codeman never forwards its own credentials upstream. The `Authorization` header
and the `codeman_session` cookie are stripped on the way out, so `CODEMAN_PASSWORD` cannot
leak into a dashboard.
## Known incompatibility: cookie-authenticated reverse proxies
If Codeman itself sits behind Cloudflare Access, Authelia, oauth2-proxy, or similar, a
**sandboxed** tab may render unstyled or broken while the Codeman page around it works fine.
The reason: an opaque-origin frame's stylesheet, script, and API requests do not carry the
proxy's authentication cookie. The proxy redirects them to the login provider, and CORS or
CSP kills them there.
Trusted mode keeps a real origin and the cookie, so it works. Test cannot catch this, because
it checks the server's reach, not the browser's.
## Security notes
The proxy authenticates on an in-memory capability embedded in the path, which is why it is
exempt from the cookie and Origin checks that every API route enforces. That exemption is
fenced to safe methods and non-API paths, and there is a test pinning it in place.
Two failure modes that only appear inside a sandboxed frame, and that curl can never
reproduce, are handled: runtime-built root-absolute URLs escaping the injected base, and
same-host requests being CORS-checked with a null origin. Both present as the dashboard's own
"Failed to fetch" while the page itself renders fine.
## Read next
- [The Dashboard](The-Dashboard) - the tab strip these share.
- [Security](Security) - why the sandbox default is what it is.
- [`docs/web-tabs.md`](https://github.com/Ark0N/Codeman/blob/master/docs/web-tabs.md) - the full reference.
+159
View File
@@ -0,0 +1,159 @@
# Working With Files
Reading, editing, attaching, and previewing files without leaving the dashboard. Useful on
a desktop; on a phone it is the difference between reviewing an agent's work and waiting
until you get home.
## The File Viewer
A panel that browses the active session's working directory. Its header button is on by
default; if it is missing, re-enable it in **App Settings → Header & Panels**.
It renders what it can:
| Kind | Behaviour |
| ------------------------ | ------------------------------------------------------------------------- |
| Text and code | Syntax-aware preview. Long files are truncated in plain preview. |
| Images | Inline. |
| Audio and video | Inline with a working scrub bar, because range requests are supported. |
| PDF and Office documents | Converted for preview when a converter is available. |
| Anything else | Download. |
Caps: 10 MB for text preview, 50 MB for raw and download. Sensitive paths (`.env`, anything
matching credentials, `~/.ssh`, AWS credentials) are blocked from download, and SVG and HTML
are served as downloads rather than rendered, so they cannot execute in the page.
Closing the preview pauses and unloads any playing media. A video that keeps playing after
you close the panel means you are on an old version.
## Editing in place
Text files can be edited and saved directly in the viewer. Click the pencil in the preview
header, edit, **Save**.
The guardrails are worth knowing, because they are what makes editing safe rather than
convenient:
- **Extension allowlist**, not a blocklist. Code, docs, config, and markup are editable.
Anything not on the list is not.
- **512 KB cap** on both read and write.
- **Edit mode never truncates.** The plain preview does truncate long files, and saving a
truncated buffer would silently delete the rest, so the editor loads the whole file or
refuses.
- **Optimistic concurrency.** The save carries a hash of what you started from. If the file
changed underneath you (likely, when an agent is working in the same repo), the save is
rejected rather than clobbering their work.
- **No file creation.** Writes go to a temporary file and are renamed over the original, and
the open never creates. Editing in place is structural, not a rule.
- **Line endings are preserved** server-side, so editing two lines of a CRLF file does not
produce a whole-file diff.
- **`.git/` is denied outright.** Hooks are executable code, and a corrupted index looks
unrecoverable to someone who wanted to fix a typo.
- **Non-UTF-8 content is refused**, verified by a round-trip comparison.
## Attachments
Attachments are live references to files **outside** the session's workspace: a spec on your
desktop, a PDF in Downloads, a design document elsewhere on the machine.
Register one from the CLI:
```bash
codeman attach /path/to/spec.pdf
```
An attachment card appears in the session, and the file can be previewed inline. The
attachment gets a stable id, and browser requests use that id rather than carrying absolute
paths around.
Agents can register attachments too, by emitting a `codeman://attach?...` link in their
output. That path is **prompt-injectable by nature**, so it is force-confined to the
session's workspace: a hostile prompt cannot use it to pull arbitrary host files into the
event stream. The gate is an extension allowlist rather than a blocklist.
Document conversion for previews is globally rate limited. Without that, ten large documents
detected at once would fork ten multi-minute converter processes.
## Clicking a path
File paths in a session are links. That works in two places:
- **In the terminal**, on any absolute path an agent prints.
- **In the response viewer**, where paths are usually written as prose or in backticks. They
render as underlined monospace links.
Clicking one opens it in the preview: images and PDFs render, video and audio play with a
working scrub bar, documents convert, text and Markdown show inline. Log-shaped files open in
the tail viewer instead, which follows a file that is still being written.
Paths **outside** the session's workspace work too, which matters because that is where most
of an agent's output lands: a screenshot in `/tmp`, a capture in its own scratchpad, a file in
another checkout. Those are served through the attachment routes rather than the workspace
ones, so the same rules apply as to any other attachment: secret trees are blocked, the
extension allowlist decides what can be opened, and symlinks are resolved before either check.
Outside the workspace the allowlist is images, video, audio, PDF, Office documents, and text
files, where "text" is the same list the viewer will let you edit: code, config, logs, csv,
markdown. The reasoning is that a session can already `cat` any of those, so the file suffix
was never what kept anything secret; the path guard is. Types outside the list (`.svg`,
`.bmp`) say so rather than failing silently, and `.html` previews as source rather than being
rendered, so nothing served this way can execute in the page.
Text previews are capped at the first 500 lines, fetched as a partial read, so clicking a
one-gigabyte log does not try to paint one.
Log-shaped files inside the workspace still open in the tail viewer, which follows a file as
it is written. Outside the workspace they open in the preview instead: the tail viewer runs
`tail -f`, and that is deliberately restricted to the workspace, `/var/log` and `~/logs`.
Nothing is registered until you click. Opening a file this way does not add an attachment card.
## The path picker
For choosing a path rather than typing one. It appears in two places:
- **Browse** in **Add Case → Link Existing**.
- The **📁 Path** key on the mobile keyboard bar.
It browses one directory at a time and can show hidden entries on request. The picker
inserts the path into your prompt **without** pressing Enter, so nothing is submitted by
accident. Its sibling **⌫ All** key clears the unsent prompt, and never sends the agent's
`/clear` command.
This is a separate file-serving surface from the viewer, with its own rules: it allowlists
your home directory, the cases directory, and anything in `CODEMAN_FILE_PICKER_ROOTS`, and
blocks sensitive trees. In multi-user mode a non-admin gets only their own user space as a
root, because per-user spaces live inside the home directory and a home-directory root would
expose everyone.
## Images into a session
Paste from the clipboard or drag and drop straight onto the terminal. The image is written
where the agent can read it and the reference is inserted into your prompt. On a phone, the
image key in the keyboard bar opens the camera or photo library.
HEIC images from an iPhone are converted to JPEG on the way in.
## Generated artifacts
When an agent produces a file the UI can show (a chart, a diagram, a document), it can
surface as an artifact attachment rather than a path you have to go and find.
## Gotchas
- **The viewer follows the active session's workspace.** Switching tabs changes what you are
browsing.
- **A save can be rejected, and that is the feature.** It means the agent edited the file
while you were typing. Re-open, re-apply, save again.
- **Attachments live outside the workspace on purpose.** For files inside it, just use the
viewer.
- **`.env` files are readable in the viewer if the extension policy allows the preview, but
never downloadable.** Do not treat the viewer as a secrets boundary; treat the machine as
the boundary.
## Read next
- [The Dashboard](The-Dashboard) - where the panels live.
- [Input And Voice](Input-And-Voice) - other ways to get content into a session.
- [Security](Security) - how the file surfaces are confined.
- [`docs/file-viewer-edit-plan.md`](https://github.com/Ark0N/Codeman/blob/master/docs/file-viewer-edit-plan.md) - the edit-mode design.
+9
View File
@@ -0,0 +1,9 @@
Documents Codeman **{{VERSION}}**. Something wrong or missing on this page? These pages are
generated from [`docs/wiki/`](https://github.com/Ark0N/Codeman/tree/master/docs/wiki) in
the main repository, so browser edits here are overwritten on the next sync. Send a pull
request against that directory instead, or open a
[Discussion](https://github.com/Ark0N/Codeman/discussions).
<!-- {{VERSION}} is replaced with the current major.minor series by
.github/workflows/wiki-sync.yml at publish time. Do not hardcode a
version here: it went stale every release when it was hand-written. -->
+53
View File
@@ -0,0 +1,53 @@
### [Codeman Wiki](Home)
[README](https://github.com/Ark0N/Codeman)
**Getting started**
- [Installation](Installation)
- [Quick Start](Quick-Start)
- [Core Concepts](Core-Concepts)
**Using it**
- [The Dashboard](The-Dashboard)
- [Agent CLIs](Agent-CLIs)
- [Working With Files](Working-With-Files)
- [Input And Voice](Input-And-Voice)
- [Mobile Guide](Mobile-Guide)
- [Keyboard Shortcuts](Keyboard-Shortcuts)
- [Settings Reference](Settings-Reference)
**Keeping agents running**
- [Unattended Runs](Keeping-Agents-Running)
- [Notifications & Approvals](Notifications-And-Approvals)
- [Cron Jobs](Cron-Jobs)
- [Autonomous Loops](Autonomous-Loops)
- [Watching Agents Work](Watching-Agents-Work)
**Where it runs**
- [Docker Cases](Docker-Cases)
- [Remote SSH Sessions](Remote-SSH-Sessions)
- [Web Tabs](Web-Tabs)
- [Multi-User Mode](Multi-User-Mode)
**Access & security**
- [Remote Access](Remote-Access)
- [Security](Security)
**Automation**
- [Driving It From An Agent](Driving-Codeman-From-An-Agent)
- [HTTP API](HTTP-API)
- [Hooks & Integrations](Hooks-And-Integrations)
**Operating it**
- [Running As A Service](Running-As-A-Service)
- [Troubleshooting](Troubleshooting)
- [FAQ](FAQ)
- [Contributing](Contributing)
- [Versioning](Versioning)
+47 -29
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "1.19.4",
"version": "1.20.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "1.19.4",
"version": "1.20.1",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -17,7 +17,7 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/multipart": "^10.0.0",
"@fastify/static": "^9.1.3",
"@fastify/static": "^10.1.3",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-serialize": "^0.14.0",
@@ -60,6 +60,7 @@
"pixelmatch": "^6.0.0",
"playwright": "^1.58.0",
"pngjs": "^7.0.0",
"postcss": "^8.5.15",
"prettier": "^3.4.0",
"puppeteer": "^24.36.0",
"remotion": "4.0.473",
@@ -1453,9 +1454,9 @@
}
},
"node_modules/@fastify/static": {
"version": "9.1.3",
"resolved": "https://registry.npmjs.org/@fastify/static/-/static-9.1.3.tgz",
"integrity": "sha512-aXrYtsiryLhRxRNaxNqsn7FUISeb7rB9q4eHUPIot5aeQBLNahnz1m6thzm7JWC1poSGXS9XrX8DvuMivp2hkQ==",
"version": "10.1.3",
"resolved": "https://registry.npmjs.org/@fastify/static/-/static-10.1.3.tgz",
"integrity": "sha512-W6jqajYS974XjPjB5hQWoxPM8NKM4+p8YmQT6G5IbCa4uhdWSVadZUv75siy1wEA/3ty8RYdpBydfWeu9AqAqQ==",
"funding": [
{
"type": "github",
@@ -1469,13 +1470,30 @@
"license": "MIT",
"dependencies": {
"@fastify/accept-negotiator": "^2.0.0",
"@fastify/error": "^4.0.0",
"@fastify/send": "^4.0.0",
"content-disposition": "^1.0.1",
"fastify-plugin": "^5.0.0",
"content-disposition": "^2.0.1",
"fastify-plugin": "^6.0.0",
"fastq": "^1.17.1",
"glob": "^13.0.0"
}
},
"node_modules/@fastify/static/node_modules/fastify-plugin": {
"version": "6.0.0",
"resolved": "https://registry.npmjs.org/fastify-plugin/-/fastify-plugin-6.0.0.tgz",
"integrity": "sha512-fZOty7z3O7vOliF6d8bHE3wiEh1KcNnKEQensSgTk9C1DvN6nRLS++XVd86v33Hw/8u9Un8A1zDrQ8ujcQDHEg==",
"funding": [
{
"type": "github",
"url": "https://github.com/sponsors/fastify"
},
{
"type": "opencollective",
"url": "https://opencollective.com/fastify"
}
],
"license": "MIT"
},
"node_modules/@fastify/websocket": {
"version": "11.2.0",
"resolved": "https://registry.npmjs.org/@fastify/websocket/-/websocket-11.2.0.tgz",
@@ -4135,16 +4153,16 @@
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/brace-expansion": {
"version": "5.0.6",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.6.tgz",
"integrity": "sha512-kLpxurY4Z4r9sgMsyG0Z9uzsBlgiU/EFKhj/h91/8yHu0edo7XuixOIH3VcJ8kkxs6/jPzoI6U9Vj3WqbMQ94g==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"dev": true,
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
"node": "20 || >=22"
}
},
"node_modules/@typescript-eslint/typescript-estree/node_modules/minimatch": {
@@ -5077,9 +5095,9 @@
"license": "MIT"
},
"node_modules/brace-expansion": {
"version": "1.1.15",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.15.tgz",
"integrity": "sha512-EwOCDEex4quD37XhqM3omwtMoJjr//isUZz1JopUNWms+4Z2ViyM/k1YIRePpoVNnQhENnxtFjLaxNHrT7xIUg==",
"version": "1.1.18",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-1.1.18.tgz",
"integrity": "sha512-Edep/X9fGqVNmzKBVsDYIOtD+z1tuezV70LBjdCst9Tqu76lsnvRiZ6oTic1n+/BIwX6QDGAO94PN4N2SADvtw==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -5430,9 +5448,9 @@
}
},
"node_modules/content-disposition": {
"version": "1.1.0",
"resolved": "https://registry.npmjs.org/content-disposition/-/content-disposition-1.1.0.tgz",
"integrity": "sha512-5jRCH9Z/+DRP7rkvY83B+yGIGX96OYdJmzngqnw2SBSxqCFPd0w2km3s5iawpGX8krnwSGmF0FW5Nhr0Hfai3g==",
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/content-disposition/-/content-disposition-2.0.1.tgz",
"integrity": "sha512-e+H0ZXHSWYrENhQzw1LPuP4oF5MzVKmDU6d3hxlvaPEYLLg62MxtQNPRx4SYSuYJSBUgnQIG4HIN2tEtNv7Dog==",
"license": "MIT",
"engines": {
"node": ">=18"
@@ -6551,9 +6569,9 @@
}
},
"node_modules/fast-uri": {
"version": "3.1.2",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.2.tgz",
"integrity": "sha512-rVjf7ArG3LTk+FS6Yw81V1DLuZl1bRbNrev6Tmd/9RaroeeRRJhAt7jg/6YFxbvAQXUCavSoZhPPj6oOx+5KjQ==",
"version": "3.1.5",
"resolved": "https://registry.npmjs.org/fast-uri/-/fast-uri-3.1.5.tgz",
"integrity": "sha512-gHwA1O9LDIcKunMKhObS/HimwtehO1nPUECKAu5TpKgaO19fcWEl4bliWe1jWxVFvIXztJjjQ4L8XQ1EU9f7Jw==",
"funding": [
{
"type": "github",
@@ -6660,9 +6678,9 @@
}
},
"node_modules/find-my-way": {
"version": "9.6.0",
"resolved": "https://registry.npmjs.org/find-my-way/-/find-my-way-9.6.0.tgz",
"integrity": "sha512-Zf4Xve4RymLl7NgaavNebZ01joJ8MfVerOG43wy7SHLO+r+K0C6d/SE0BiR7AV5V1VOCFlOP7ecdo+I4qmiHrQ==",
"version": "9.8.0",
"resolved": "https://registry.npmjs.org/find-my-way/-/find-my-way-9.8.0.tgz",
"integrity": "sha512-JtyUgATO7qxRp2zKhrmWof74Mqxc1ikbwpwMY97p8ipuTj2QtreA4gK2JNAF6SOqqHnYYkwMUvsgQVi2AJxIyw==",
"license": "MIT",
"dependencies": {
"fast-deep-equal": "^3.1.3",
@@ -6904,15 +6922,15 @@
}
},
"node_modules/glob/node_modules/brace-expansion": {
"version": "5.0.6",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.6.tgz",
"integrity": "sha512-kLpxurY4Z4r9sgMsyG0Z9uzsBlgiU/EFKhj/h91/8yHu0edo7XuixOIH3VcJ8kkxs6/jPzoI6U9Vj3WqbMQ94g==",
"version": "5.0.9",
"resolved": "https://registry.npmjs.org/brace-expansion/-/brace-expansion-5.0.9.tgz",
"integrity": "sha512-ScQ4IuvIEF1TMlP7Zt+vjJ//9zlPb2SDcxWxM3bk8s6t6GGdJ7KO1dCcTidOPJKePW30LE/2cT7wCyPho9/Wxg==",
"license": "MIT",
"dependencies": {
"balanced-match": "^4.0.2"
},
"engines": {
"node": "18 || 20 || >=22"
"node": "20 || >=22"
}
},
"node_modules/glob/node_modules/minimatch": {
@@ -12343,7 +12361,7 @@
}
},
"packages/xterm-zerolag-input": {
"version": "0.3.0",
"version": "0.3.1",
"license": "MIT",
"devDependencies": {
"@xterm/headless": "^6.0.0",
+9 -5
View File
@@ -1,6 +1,6 @@
{
"name": "aicodeman",
"version": "1.19.4",
"version": "1.20.1",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
@@ -17,10 +17,13 @@
"dev": "tsx src/index.ts web",
"web": "node dist/index.js web",
"clean": "rm -rf dist",
"test": "vitest run --config config/vitest.config.ts",
"test:watch": "vitest --config config/vitest.config.ts",
"test:coverage": "vitest run --config config/vitest.config.ts --coverage",
"test": "vitest run --config config/vitest.ci.config.ts",
"test:watch": "vitest --config config/vitest.ci.config.ts",
"test:coverage": "vitest run --config config/vitest.ci.config.ts --coverage",
"test:ci": "vitest run --config config/vitest.ci.config.ts",
"test:browser": "vitest run --config config/vitest.browser.config.ts",
"test:perf": "vitest run --config config/vitest.perf.config.ts",
"test:all": "vitest run --config config/vitest.config.ts",
"pretest:mobile": "node scripts/prepare-test-vendor.mjs",
"test:mobile": "vitest run --config test/mobile/vitest.config.ts",
"check:frontend-syntax": "node scripts/check-frontend-syntax.mjs",
@@ -82,7 +85,7 @@
"@fastify/compress": "^8.3.1",
"@fastify/cookie": "^11.0.2",
"@fastify/multipart": "^10.0.0",
"@fastify/static": "^9.1.3",
"@fastify/static": "^10.1.3",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-serialize": "^0.14.0",
@@ -121,6 +124,7 @@
"pixelmatch": "^6.0.0",
"playwright": "^1.58.0",
"pngjs": "^7.0.0",
"postcss": "^8.5.15",
"prettier": "^3.4.0",
"puppeteer": "^24.36.0",
"remotion": "4.0.473",
+15
View File
@@ -1,5 +1,20 @@
# xterm-zerolag-input
## 0.3.1
### Patch Changes
- Mobile catches up: links open from a tap, terminal text can be selected and copied, long prompts stay visible while you type. Plus Files panel search, a bundled Nerd Font symbols fallback, and a per-device terminal font setting.
- **Terminal and chat links work on phones** (#321): tapping a URL or file path in terminal output now opens it (new tab, file preview, or log viewer), resolved through the same provider desktop hover uses, so tap and click can never disagree about what is a link. Dialog rows and the composer keep their existing meaning. Response-viewer links open in a new tab with `rel="noopener noreferrer"` instead of navigating the dashboard away. Wrapped links open whole: the logical-line reconstruction now stitches hard wraps through the indent their continuation carries, which also fixes desktop hover-click truncating wrapped URLs.
- **Terminal text can be copied on touch devices** (#321): long-press selects the token under the finger, drag or tap the other end to extend, and a small bar offers Copy, Line (the whole logical line, wraps included) and dismiss. Copy works on plain-HTTP installs too. Three guards keep the keyboard down and the selection alive through the browser's own long-press handling.
- **A long prompt stays visible on phones** (#321): the local-echo overlay grows upward once it would run past the last visible row (a prompt taller than the screen keeps its tail, where the cursor is), and the keyboard-driven padding shrink can no longer reclaim the space the fixed toolbar and accessory bar stand in.
- **Files panel search** (#324): `GET /api/sessions/:id/files?q=...` answers a flat match list (name or path substring, `*`/`?` globs), recursing past non-matching directories with its own match cap on top of the existing bounds; without `q` the response is byte-identical to before. Glob queries are matched without regex so a pathological pattern cannot stall the server.
- **Nerd Font prompt glyphs out of the box, custom terminal font** (#320): a bundled icons-only Symbols Nerd Font Mono fallback renders powerlevel10k/starship/oh-my-posh glyphs on every device with no font install, and App Settings gains a per-device terminal font family that is prepended to the built-in stack.
### Thanks
Three contributor PRs in one release: thanks to @rounakdatta (#321), @aakhter (#324) and @comzine (#320).
## 0.3.0
### Minor Changes
+1 -1
View File
@@ -1,6 +1,6 @@
{
"name": "xterm-zerolag-input",
"version": "0.3.0",
"version": "0.3.1",
"description": "Instant keystroke feedback overlay for xterm.js: Mosh-inspired local echo that removes perceived input latency over SSH, tunnels and other high-RTT connections",
"type": "module",
"main": "dist/index.cjs",
@@ -65,38 +65,71 @@ export function renderOverlay(container: HTMLDivElement, params: RenderParams):
charTop,
charHeight,
promptRow,
totalRows,
font,
showCursor,
cursorColor,
terminal,
} = params;
// Position container at prompt row.
// ── Keep what is being typed ON SCREEN ────────────────────────────
//
// The overlay lays its wrapped lines out DOWNWARD from the prompt row, and
// nothing past the last terminal row is visible. On a phone the strip left
// above the on-screen keyboard is only a handful of rows, so a prompt long
// enough to wrap ran off the bottom and the user was typing blind — the tail
// of their own sentence, the part they are actually looking at, hidden behind
// the keyboard.
//
// So the composer grows UPWARD once it reaches the last row, exactly as a real
// terminal's does: every line div is opaque (see makeLine), so the lines cover
// transcript rows above instead of vanishing under the keyboard below, and the
// newest text stays where the eye is. A prompt taller than the whole viewport
// keeps its TAIL for the same reason.
//
// `startCol` indents only the line that begins at the prompt marker, so it is
// dropped along with that line when the tail is all that fits.
const rows = totalRows && totalRows > 0 ? totalRows : terminal?.rows;
let visibleLines = lines;
let keepsPromptLine = true;
let topRow = promptRow;
if (rows && rows > 0) {
if (lines.length > rows) {
visibleLines = lines.slice(lines.length - rows);
keepsPromptLine = false;
topRow = 0;
} else if (promptRow + lines.length > rows) {
topRow = rows - lines.length;
}
}
topRow = Math.max(0, topRow);
container.style.left = '0px';
container.style.top = promptRow * cellH + 'px';
container.style.top = topRow * cellH + 'px';
// Clear and rebuild (typically 1-3 line divs, negligible cost)
container.innerHTML = '';
const fullWidthPx = totalCols * cellW;
for (let i = 0; i < lines.length; i++) {
const leftPx = i === 0 ? startCol * cellW : 0;
const widthPx = i === 0 ? fullWidthPx - leftPx : fullWidthPx;
for (let i = 0; i < visibleLines.length; i++) {
const indents = i === 0 && keepsPromptLine;
const leftPx = indents ? startCol * cellW : 0;
const widthPx = indents ? fullWidthPx - leftPx : fullWidthPx;
const topPx = i * cellH;
const lineEl = makeLine(lines[i], leftPx, topPx, widthPx, cellH, cellW, charTop, charHeight, font, terminal);
const lineEl = makeLine(visibleLines[i], leftPx, topPx, widthPx, cellH, cellW, charTop, charHeight, font, terminal);
container.appendChild(lineEl);
}
// Block cursor at end of last line (use visual width for CJK support)
if (showCursor) {
const lastLine = lines[lines.length - 1];
const lastLineLeft = lines.length === 1 ? startCol : 0;
const lastLine = visibleLines[visibleLines.length - 1];
const lastLineLeft = visibleLines.length === 1 && keepsPromptLine ? startCol : 0;
const cursorCol = lastLineLeft + stringCellWidth(terminal, lastLine);
if (cursorCol < totalCols) {
const cursor = document.createElement('span');
cursor.style.cssText = 'position:absolute;display:inline-block';
cursor.style.left = cursorCol * cellW + 'px';
cursor.style.top = (lines.length - 1) * cellH + 'px';
cursor.style.top = (visibleLines.length - 1) * cellH + 'px';
cursor.style.width = cellW + 'px';
cursor.style.height = cellH + 'px';
cursor.style.backgroundColor = cursorColor;
@@ -172,6 +172,13 @@ export interface RenderParams {
/** Height of the character rendering area (px). */
charHeight: number;
promptRow: number;
/**
* Visible terminal rows. When given, the overlay is kept ON SCREEN: it grows
* upward instead of running off the bottom edge, and a wrapped prompt taller
* than the viewport keeps its tail. Omit to lay out straight down from
* `promptRow` (the historical behaviour).
*/
totalRows?: number;
font: FontStyle;
showCursor: boolean;
cursorColor: string;
@@ -565,7 +565,10 @@ export class ZerolagInputAddon implements XtermAddon {
// Skip redundant re-renders — include text content to detect
// same-length changes (e.g., setFlushed with different text)
const renderKey = `${displayText}:${startCol}:${activePrompt.row}:${activePrompt.col}:${totalCols}:${this._flushedOffset}`;
// `rows` is part of the key: the layout is clamped to the visible rows
// (see renderOverlay), so a keyboard opening — which changes rows without
// changing the text — must not be skipped as a redundant render.
const renderKey = `${displayText}:${startCol}:${activePrompt.row}:${activePrompt.col}:${totalCols}:${this._terminal.rows}:${this._flushedOffset}`;
if (renderKey === this._lastRenderKey && this._overlay.style.display !== 'none') return;
this._lastRenderKey = renderKey;
@@ -612,6 +615,7 @@ export class ZerolagInputAddon implements XtermAddon {
charTop,
charHeight,
promptRow: activePrompt.row,
totalRows: this._terminal.rows,
font: this._font,
showCursor: this._options.showCursor,
cursorColor,
@@ -418,3 +418,88 @@ describe('stringCellWidth', () => {
expect(stringCellWidth(null, '')).toBe(0);
});
});
describe('renderOverlay — staying on screen (totalRows)', () => {
// A phone with the keyboard up leaves only a handful of terminal rows. The
// overlay lays its wrapped lines out downward from the prompt row, so a long
// prompt used to run off the bottom edge and the user typed blind, with the
// tail of their own sentence behind the keyboard. With totalRows known, the
// composer grows UPWARD instead — the line divs are opaque, so they cover
// transcript above rather than disappearing below.
const linesOf = (n: number) => Array.from({ length: n }, (_, i) => `line${i}`);
const lineDivs = (container: HTMLDivElement) =>
Array.from(container.children).filter((el) => el.tagName === 'DIV') as HTMLDivElement[];
it('lifts the block so its last line lands on the last visible row', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: linesOf(5), promptRow: 10, totalRows: 12, cellH: 17 }));
// 10 + 5 would end on row 14 of a 12-row screen; the block starts at 7 instead.
expect(container.style.top).toBe(7 * 17 + 'px');
expect(lineDivs(container)).toHaveLength(5);
});
it('leaves the prompt row alone when the block already fits', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: linesOf(3), promptRow: 5, totalRows: 24, cellH: 17 }));
expect(container.style.top).toBe(5 * 17 + 'px');
});
it('keeps the TAIL when the prompt is taller than the whole viewport', () => {
// The end is where the cursor is, and where the user is looking.
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: linesOf(6), promptRow: 2, totalRows: 3, cellH: 20 }));
const divs = lineDivs(container);
expect(container.style.top).toBe('0px');
expect(divs).toHaveLength(3);
expect(divs.map((d) => d.textContent)).toEqual(['line3', 'line4', 'line5']);
});
it('drops the prompt indent once the prompt line is no longer shown', () => {
// startCol indents only the line that begins at the prompt marker.
const container = document.createElement('div');
renderOverlay(
container,
makeParams({ lines: linesOf(6), promptRow: 2, totalRows: 3, startCol: 5, cellW: 10, totalCols: 80 })
);
const first = lineDivs(container)[0];
expect(first.style.left).toBe('0px');
expect(first.style.width).toBe(80 * 10 + 'px');
});
it('rides the cursor on the last VISIBLE line', () => {
const container = document.createElement('div');
renderOverlay(
container,
makeParams({ lines: ['aaa', 'bbb', 'ccc', 'ddd'], promptRow: 9, totalRows: 3, cellH: 20, cellW: 10, startCol: 4 })
);
const cursor = Array.from(container.children).find((el) => el.tagName === 'SPAN') as HTMLSpanElement;
// Tail is the last 3 lines, so the cursor sits on row 2 (0-based) of the block…
expect(cursor.style.top).toBe(2 * 20 + 'px');
// …at column 3, NOT startCol + 3: the indented prompt line is not shown.
expect(cursor.style.left).toBe(3 * 10 + 'px');
});
it('lays out straight down when totalRows is absent (unchanged behaviour)', () => {
const container = document.createElement('div');
renderOverlay(container, makeParams({ lines: linesOf(9), promptRow: 20, cellH: 17 }));
expect(container.style.top).toBe(20 * 17 + 'px');
expect(lineDivs(container)).toHaveLength(9);
});
it('falls back to the terminal row count when totalRows is not passed', () => {
// The addon passes totalRows, but a stale bundle / third-party caller may not.
const container = document.createElement('div');
renderOverlay(
container,
makeParams({ lines: linesOf(4), promptRow: 8, cellH: 17, terminal: { rows: 10, cols: 80 } as never })
);
expect(container.style.top).toBe(6 * 17 + 'px');
});
});
+2 -1
View File
@@ -8,7 +8,8 @@
* src/web/public/constants.js and deliberately stays at 50k — 100k xterm lines per tab
* is a mobile-memory hazard — so DEFAULT_TERMINAL_SCROLLBACK_LINES stays 50,000 to match.
* The terminalScrollbackLines/terminalBufferMaxBytes/terminalBufferTrimBytes settings keys
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is live.
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is wired.
* tmux <3.7 applies it to new panes; tmux 3.7+ can also resize live panes.
* All values remain env- and settings-overridable and bounds-clamped via
* resolveTerminalHistoryConfig().
*/
+11 -2
View File
@@ -479,14 +479,23 @@ export function gitNonInteractiveEnv(base: NodeJS.ProcessEnv = process.env): Nod
// ─── Pure: output handling ───────────────────────────────────────────────────
/**
* Redact any `scheme://user:secret@host` credential pair embedded in text — a
* remote URL stored with an inline token, or git stderr echoing such a URL
* back. Shared by the clone error path (`sanitizeGitOutput`) and the
* repository-status card fields (`web/repo-status.ts`).
*/
export function redactGitCredentials(text: string): string {
return text.replace(/([a-zA-Z][a-zA-Z0-9+.-]*:\/\/)[^/@\s]*:[^/@\s]*@/g, '$1***:***@');
}
/**
* Make git's stderr safe to show in the browser: strip ANSI/control bytes,
* redact any `scheme://user:secret@host` that a credential helper echoed back,
* and keep only the tail (the last lines are the ones that say why it failed).
*/
export function sanitizeGitOutput(text: string, maxBytes = MAX_STDERR_BYTES): string {
const redacted = text
.replace(/([a-zA-Z][a-zA-Z0-9+.-]*:\/\/)[^/@\s]*:[^/@\s]*@/g, '$1***:***@')
const redacted = redactGitCredentials(text)
// eslint-disable-next-line no-control-regex -- deliberate: strip C0/C1 and DEL.
.replace(/[\u0000-\u0008\u000b\u000c\u000e-\u001f\u007f-\u009f]/g, '')
.trim();
+3 -3
View File
@@ -84,7 +84,7 @@ export interface CreateSessionOptions {
envOverrides?: Record<string, string>;
/** Claude CLI effort level, injected as a `--settings` soft default (overridable via /effort in-session) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session. */
/** tmux history-limit (scrollback lines) allocated when this session is created. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
@@ -116,7 +116,7 @@ export interface RespawnPaneOptions {
envOverrides?: Record<string, string>;
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session after respawn. */
/** Original tmux history-limit retained for config parity; respawn cannot resize the existing pane. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
@@ -216,7 +216,7 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Update Ralph enabled state for a session */
updateRalphEnabled(sessionId: string, enabled: boolean): void;
/** Apply a tmux history-limit to all tracked sessions. */
/** Apply history-limit to live panes where tmux supports it, otherwise to future panes. */
setHistoryLimit(limit: number): Promise<void>;
// ========== Discovery ==========
+8 -1
View File
@@ -281,7 +281,14 @@ export class RalphLoop extends EventEmitter {
// Guard: only reschedule if still running AND no timer is pending
// (prevents race where stop() clears timer between our check and setTimeout)
if (this._status === 'running' && this.loopTimer === null) {
this.loopTimer = setTimeout(() => this.runLoop(), this.pollIntervalMs);
// Null the handle when the timer fires, BEFORE re-entering runLoop —
// otherwise the `loopTimer === null` guard above stays false on the
// next pass and the loop stops rescheduling after 2 ticks.
// Mirrors the orchestrator-loop reschedule pattern.
this.loopTimer = setTimeout(() => {
this.loopTimer = null;
this.runLoop();
}, this.pollIntervalMs);
}
});
}
+20
View File
@@ -1624,6 +1624,11 @@ export class RespawnController extends EventEmitter {
const prompt = this.config.kickstartPrompt!;
this.logAction('command', `Sending kickstart: "${prompt.substring(0, 40)}..."`);
await this.session.writeViaMux(prompt + '\r'); // \r triggers key.return in Ink/Claude CLI
// COD-51: stop() may have run during the await; re-check before reviving the
// state machine. Reads the public getter, not `_state`: TypeScript narrows
// `_state` across the await from the guard above and cannot see that stop()
// mutated it, so the comparison would be flagged as impossible.
if (this.state === 'stopped') return;
this.emit('stepSent', 'kickstart', prompt);
this.setState('waiting_kickstart');
this.promptDetected = false;
@@ -2833,6 +2838,11 @@ export class RespawnController extends EventEmitter {
const input = updatePrompt + '\r'; // \r triggers Enter in Ink/Claude CLI
this.logAction('command', `Sending: "${updatePrompt.substring(0, 50)}..."`);
await this.session.writeViaMux(input);
// COD-51: stop() may have run during the await; re-check before reviving the
// state machine. Reads the public getter, not `_state`: TypeScript narrows
// `_state` across the await from the guard above and cannot see that stop()
// mutated it, so the comparison would be flagged as impossible.
if (this.state === 'stopped') return;
this.emit('stepSent', 'update', updatePrompt);
this.setState('waiting_update');
this.promptDetected = false;
@@ -2860,6 +2870,11 @@ export class RespawnController extends EventEmitter {
if (this._state === 'stopped') return;
this.logAction('command', 'Sending: /clear');
await this.session.writeViaMux('/clear\r'); // \r triggers Enter in Ink/Claude CLI
// COD-51: stop() may have run during the await; re-check before reviving the
// state machine. Reads the public getter, not `_state`: TypeScript narrows
// `_state` across the await from the guard above and cannot see that stop()
// mutated it, so the comparison would be flagged as impossible.
if (this.state === 'stopped') return;
this.emit('stepSent', 'clear', '/clear');
this.setState('waiting_clear');
this.promptDetected = false;
@@ -2902,6 +2917,11 @@ export class RespawnController extends EventEmitter {
if (this._state === 'stopped') return;
this.logAction('command', 'Sending: /init');
await this.session.writeViaMux('/init\r'); // \r triggers Enter in Ink/Claude CLI
// COD-51: stop() may have run during the await; re-check before reviving the
// state machine. Reads the public getter, not `_state`: TypeScript narrows
// `_state` across the await from the guard above and cannot see that stop()
// mutated it, so the comparison would be flagged as impossible.
if (this.state === 'stopped') return;
this.emit('stepSent', 'init', '/init');
this.setState('waiting_init');
this.promptDetected = false;
+61 -3
View File
@@ -415,6 +415,15 @@ export class Session extends EventEmitter {
// sequences split across PTY chunks can't slip past the alt-screen/scrollback
// strip (see _handleTerminalOutput / isAltScreenStripMode)
private _altScreenSeqCarry: string = '';
/**
* Mouse-tracking DECSET modes the CLI currently has ON, as observed while
* STRIPPING them out of the stream below. Kept as a set rather than a boolean
* because a TUI may enable 1002 and later disable 1000 (a mode it never
* enabled); tracking is on while any of them is.
*/
private _cliMouseModes = new Set<number>();
private _cliMouseTracking = false;
private resolvePromise: ((value: { result: string; cost: number }) => void) | null = null;
private rejectPromise: ((reason: Error) => void) | null = null;
private _promptResolved: boolean = false; // Guard against race conditions in runPrompt
@@ -510,7 +519,7 @@ export class Session extends EventEmitter {
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
private _effort: EffortLevel | undefined;
// tmux history-limit (scrollback lines) applied to this session's pane.
// tmux history-limit (scrollback lines) allocated when this session's pane is created.
private readonly _tmuxHistoryLimit: number;
// Remote execution metadata, present when this session runs over SSH through local tmux.
@@ -600,7 +609,7 @@ export class Session extends EventEmitter {
envOverrides?: Record<string, string>;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) for this session's pane. */
/** tmux history-limit (scrollback lines) allocated when this session's pane is created. */
tmuxHistoryLimit?: number;
/** Restored per-session attachment history. May include server-private external paths. */
attachmentHistory?: SessionAttachmentHistoryItem[];
@@ -1285,6 +1294,7 @@ export class Session extends EventEmitter {
niceValue: this._niceConfig.niceValue,
color: this._color,
flickerFilterEnabled: this._flickerFilterEnabled,
cliMouseTracking: this._cliMouseTracking || undefined,
cliVersion: this._cliVersion || undefined,
cliModel: this._cliModel || undefined,
cliAccountType: this._cliAccountType || undefined,
@@ -1546,6 +1556,46 @@ export class Session extends EventEmitter {
};
}
/**
* Remember whether the CLI currently wants to be told about mouse clicks.
*
* The strip in {@link _handleTerminalOutput} is the ONLY place these sequences
* exist. After it, neither the browser nor xterm can ever learn that the CLI
* asked for mouse tracking, so `terminal.modes.mouseTrackingMode` is
* permanently 'none' for a stripped mode. The browser hand-encodes SGR reports
* to compensate (`_sendSyntheticSgrTap` in terminal-ui.js), and with no state
* to consult it had to do that on EVERY click, delivering mouse reports to a
* CLI that never asked for them. Publishing this through `toState()` is what
* lets the browser report a click only when the CLI is listening.
*
* Only the TRACKING modes count. 1005/1006 select an encoding and 1007 is
* alt-scroll; a CLI that picks SGR encoding without turning a tracking mode on
* is not asking about clicks, and counting those would put the stray reports
* straight back.
*
* This must stay in lockstep with the strip regex that calls it: a sequence
* removed from the stream but not recorded here is one the browser can neither
* see nor be told about.
*/
private _recordStrippedMouseMode(seq: string): void {
// eslint-disable-next-line no-control-regex
const match = /\x1b\[\?(\d+)([hl])$/.exec(seq);
if (!match) return;
const mode = Number(match[1]);
if (mode !== 1000 && mode !== 1001 && mode !== 1002 && mode !== 1003) return;
if (match[2] === 'h') this._cliMouseModes.add(mode);
else this._cliMouseModes.delete(mode);
this._syncCliMouseTracking();
}
/** Emit only on a real transition: a TUI re-emitting its enable on every repaint costs nothing. */
private _syncCliMouseTracking(): void {
const active = this._cliMouseModes.size > 0;
if (active === this._cliMouseTracking) return;
this._cliMouseTracking = active;
this.emit('mouseTrackingChanged', active);
}
private _handleTerminalOutput(data: string): void {
// Codex AND Claude Code emit sequences that wipe xterm.js scrollback, plus
// mouse-tracking enables that hijack the scroll wheel so the user can't reach
@@ -1598,7 +1648,10 @@ export class Session extends EventEmitter {
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, (seq) => {
this._recordStrippedMouseMode(seq);
return '';
});
}
}
@@ -2525,6 +2578,11 @@ export class Session extends EventEmitter {
this._messages = [];
this._lineBuffer = '';
this._altScreenSeqCarry = '';
// A restarted pane starts with no mouse mode: the new program has not asked
// for one yet, and carrying the old CLI's state over would report clicks
// into a program that never enabled tracking.
this._cliMouseModes.clear();
this._syncCliMouseTracking();
this._markActivity(true);
}
+67 -42
View File
@@ -75,11 +75,17 @@ import {
SAFE_PATH_PATTERN,
findClaudeDir,
getClaudeCliVersion,
getClaudeNotFoundMessage,
resolveOpenCodeDir,
getOpenCodeNotFoundMessage,
resolveCodexDir,
getCodexNotFoundMessage,
resolveGeminiDir,
getGeminiNotFoundMessage,
resolveAntigravityDir,
getAntigravityNotFoundMessage,
resolvePiDir,
getPiNotFoundMessage,
resolveLocalShell,
loginShellArgs,
} from './utils/index.js';
@@ -1562,6 +1568,8 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
private reconnectGuard: Set<string> = new Set();
private trueColorConfigured = false;
/** tmux 3.7+ can resize pane history after creation; older releases cannot. */
private liveHistoryResizeSupported: boolean | null = null;
constructor() {
super();
@@ -1580,6 +1588,26 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
return tmuxCommand(this.tmuxSocket);
}
private supportsLiveHistoryResize(): boolean {
if (this.liveHistoryResizeSupported !== null) return this.liveHistoryResizeSupported;
try {
const output = execSync(`${this.tmux()} -V`, {
encoding: 'utf8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
});
const match = output.match(/(?:^|\D)(\d+)\.(\d+)/);
const major = match ? Number(match[1]) : 0;
const minor = match ? Number(match[2]) : 0;
this.liveHistoryResizeSupported = major > 3 || (major === 3 && minor >= 7);
} catch {
// Unknown versions take the legacy path required by tmux <3.7.
this.liveHistoryResizeSupported = false;
}
return this.liveHistoryResizeSupported;
}
// Load saved sessions from disk (NEVER called in test mode)
private loadSessions(): void {
if (IS_TEST_MODE) return;
@@ -1853,29 +1881,28 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
return session;
}
// Resolve CLI binary directory based on mode
// Resolve CLI binary directory based on mode. The not-found messages come
// from the resolvers (formatCliNotFoundMessage) so the error names WHERE it
// looked — server PATH, login shell, checked directories — instead of just
// asserting the CLI is missing (the classic systemd/launchd PATH trap).
const { pathExport, dir: cliDir } = this.buildPathExport(mode);
if (mode === 'claude' && !cliDir) {
throw new Error('Claude CLI not found. Install it with: curl -fsSL https://claude.ai/install.sh | bash');
throw new Error(getClaudeNotFoundMessage());
}
if (mode === 'opencode' && !cliDir) {
throw new Error('OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash');
throw new Error(getOpenCodeNotFoundMessage());
}
if (mode === 'codex' && !cliDir) {
throw new Error('Codex CLI not found. Install with: npm install -g @openai/codex');
throw new Error(getCodexNotFoundMessage());
}
if (mode === 'gemini' && !cliDir) {
throw new Error('Gemini CLI not found. Install with: npm install -g @google/gemini-cli');
throw new Error(getGeminiNotFoundMessage());
}
if (mode === 'antigravity' && !cliDir) {
throw new Error(
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash'
);
throw new Error(getAntigravityNotFoundMessage());
}
if (mode === 'pi' && !cliDir) {
throw new Error(
'Pi CLI not found. Install with: npm install -g --ignore-scripts @earendil-works/pi-coding-agent'
);
throw new Error(getPiNotFoundMessage());
}
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
@@ -1921,7 +1948,16 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// launched in TMUX_LAUNCH_CWD (/tmp) rather than the real workingDir: a FUSE/rclone
// mount that isn't ready yet makes `getcwd` fail and breaks the spawn (see #110). The
// pane cd's into workingDir below via respawn-pane.
execSync(`${this.tmux()} new-session -ds "${muxName}" -c ${TMUX_LAUNCH_CWD}`, {
// tmux <3.7 allocates history only at pane creation, so its global default
// must be set immediately BEFORE new-session. tmux 3.7+ can resize a pane
// after creation; target only the new session there because changing the
// global option can resize (and when lowered, trim) unrelated live panes.
const safeHistoryLimit =
Number.isSafeInteger(historyLimit) && historyLimit > 0 ? Math.trunc(historyLimit) : DEFAULT_TMUX_HISTORY_LIMIT;
const createSessionCommand = this.supportsLiveHistoryResize()
? `${this.tmux()} new-session -ds "${muxName}" -c ${TMUX_LAUNCH_CWD} \\; set-option -t "${muxName}" history-limit ${safeHistoryLimit}`
: `${this.tmux()} set-option -g history-limit ${safeHistoryLimit} \\; new-session -ds "${muxName}" -c ${TMUX_LAUNCH_CWD} \\; set-option -t "${muxName}" history-limit ${safeHistoryLimit}`;
execSync(createSessionCommand, {
cwd: TMUX_LAUNCH_CWD,
timeout: EXEC_TIMEOUT_MS,
stdio: 'ignore',
@@ -1986,16 +2022,6 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
.catch(() => {
/* Already set globally as fallback */
}),
// Raise tmux scrollback from its 2000-line default so re-attach preserves
// more context. Intentionally exceeds the xterm-side DEFAULT_SCROLLBACK (50k
// in constants.js), which stays lower to protect browser/mobile memory.
execAsync(`${this.tmux()} set-option -t "${muxName}" history-limit ${historyLimit}`, {
timeout: EXEC_TIMEOUT_MS,
})
.then(() => {})
.catch(() => {
/* Non-critical — falls back to tmux default */
}),
];
// Enable 24-bit true color passthrough — server-wide, set once per lifetime
@@ -2119,7 +2145,6 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
resumeSessionId,
envOverrides,
effort,
historyLimit = DEFAULT_TMUX_HISTORY_LIMIT,
remote,
docker,
name,
@@ -2130,16 +2155,6 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
if (!isValidMuxName(muxName) || !isValidPath(workingDir)) return null;
// Re-apply the configured tmux history-limit after respawn (kept in sync
// with the live setting via setHistoryLimit()).
if (!IS_TEST_MODE) {
await execAsync(`${this.tmux()} set-option -t ${shellescape(muxName)} history-limit ${historyLimit}`, {
timeout: EXEC_TIMEOUT_MS,
}).catch(() => {
/* Non-critical — keeps existing tmux history-limit */
});
}
// Resolve CLI binary directory based on mode
const { pathExport } = this.buildPathExport(mode);
@@ -2998,9 +3013,9 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
/**
* Apply a tmux history-limit to all tracked sessions (e.g. when the user
* changes the terminal-history setting). Invalid limits fall back to the
* default. Best-effort per session.
* Apply a tmux history limit. tmux 3.7+ safely targets tracked live sessions;
* older releases can only change the global default for future panes. Invalid
* limits fall back to the default.
*/
async setHistoryLimit(limit: number): Promise<void> {
const safeLimit = Number.isSafeInteger(limit) && limit > 0 ? Math.trunc(limit) : DEFAULT_TMUX_HISTORY_LIMIT;
@@ -3009,12 +3024,22 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
return;
}
const updates = Array.from(this.sessions.values()).map((session) =>
execAsync(`${this.tmux()} set-option -t ${shellescape(session.muxName)} history-limit ${safeLimit}`, {
timeout: EXEC_TIMEOUT_MS,
})
);
await Promise.allSettled(updates);
if (this.supportsLiveHistoryResize()) {
const updates = Array.from(this.sessions.values()).map((session) =>
execAsync(`${this.tmux()} set-option -t ${shellescape(session.muxName)} history-limit ${safeLimit}`, {
timeout: EXEC_TIMEOUT_MS,
})
);
await Promise.allSettled(updates);
return;
}
await execAsync(`${this.tmux()} set-option -g history-limit ${safeLimit}`, {
timeout: EXEC_TIMEOUT_MS,
}).catch(() => {
// No tmux server yet is fine: legacy createSession sets the same default
// immediately before it creates the first pane.
});
}
/**
+8
View File
@@ -500,6 +500,14 @@ export interface SessionState {
color?: SessionColor;
/** Flicker filter enabled (buffers output after screen clears) */
flickerFilterEnabled?: boolean;
/**
* True while the CLI in the pane has a mouse-tracking DECSET on, as observed
* by the server on its way out of the stream (those sequences are stripped for
* claude/codex/gemini, so the browser can never see them itself). The browser
* hand-encodes a click report ONLY when this is true; without it, every click
* sent mouse reports to a CLI that never asked for them.
*/
cliMouseTracking?: boolean;
/** Claude Code CLI version (parsed from terminal, e.g., "2.1.27") */
cliVersion?: string;
/** Claude model in use (parsed from terminal, e.g., "Opus 4.5") */
+49
View File
@@ -98,3 +98,52 @@ export interface UpdateCheckResult {
source: 'github-api' | 'git-ls-remote' | 'none';
error?: string;
}
/**
* Role of a remote in the repository-status view.
* - `tracking`: the current branch's `@{upstream}` remote (where `git pull` goes).
* - `upstream`: the canonical project (a remote named `origin`/`upstream` that is
* not the tracking remote).
* - `other`: anything else explicitly requested via `CODEMAN_UPDATE_REMOTES`.
*/
export type RepoRemoteRole = 'tracking' | 'upstream' | 'other';
/** A single incoming commit — present on the remote ref but not in local HEAD. */
export interface RepoIncomingCommit {
/** Abbreviated SHA. */
sha: string;
/** Commit subject (first line). */
subject: string;
}
/** Ahead/behind + incoming summary for local HEAD vs one remote's compare ref. */
export interface RepoRemoteStatus {
/** Remote name, e.g. `origin`, `bitbucket`. */
name: string;
/** Remote URL (best-effort; empty if unresolved). */
url: string;
role: RepoRemoteRole;
/** Ref HEAD is compared against, e.g. `origin/master`, `bitbucket/local`. */
compareRef: string;
/** Commits in local HEAD not on the remote ref (local-only / unpushed). */
ahead: number;
/** Commits on the remote ref not in local HEAD (incoming). */
behind: number;
/** Up to N most recent incoming commits (newest first). */
incoming: RepoIncomingCommit[];
/** Set when this remote could not be fetched/compared. */
error?: string;
}
/** Result of the repository-status check across the configured remotes. */
export interface RepositoryStatusResult {
/** epoch ms of the check. */
checkedAt: number;
/** False when this is not a git install (then `remotes` is empty + `error` set). */
isGit: boolean;
/** Current running version, for display. */
currentVersion: string;
remotes: RepoRemoteStatus[];
/** Top-level error (e.g. not a git install, or no remotes resolved). */
error?: string;
}
+26 -30
View File
@@ -7,22 +7,37 @@
* @module utils/antigravity-cli-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import {
createCliExecutableResolver,
formatCliNotFoundMessage,
type CliResolverHost,
} from './cli-executable-resolver.js';
/** Common directories where the Antigravity CLI binary may be installed */
const ANTIGRAVITY_SEARCH_DIRS = [
join(homedir(), '.local', 'bin'),
join(homedir(), '.antigravity', 'bin'),
'/usr/local/bin',
join(homedir(), '.bun', 'bin'),
join(homedir(), '.npm-global', 'bin'),
join(homedir(), 'bin'),
];
/** Cached directory containing the agy binary (empty string = searched but not found) */
let _antigravityDir: string | null = null;
const ANTIGRAVITY_NOT_FOUND =
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash';
function createAntigravityResolver(host?: CliResolverHost, now?: () => number) {
return createCliExecutableResolver({ binary: 'agy', searchDirs: ANTIGRAVITY_SEARCH_DIRS, now }, host);
}
/** Creates an isolated Antigravity wrapper around an injected resolver host and clock. */
export function createAntigravityResolverForTest(host: CliResolverHost, now?: () => number) {
return createAntigravityResolver(host, now);
}
const antigravityResolver = createAntigravityResolver();
/**
* Finds the directory containing the `agy` binary.
@@ -31,30 +46,7 @@ let _antigravityDir: string | null = null;
* @returns Directory path, or null if not found
*/
export function resolveAntigravityDir(): string | null {
if (_antigravityDir !== null) return _antigravityDir || null;
try {
const result = execSync('which agy', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
_antigravityDir = dirname(result);
return _antigravityDir;
}
} catch {
// agy not in PATH, will check common locations
}
for (const dir of ANTIGRAVITY_SEARCH_DIRS) {
if (existsSync(join(dir, 'agy'))) {
_antigravityDir = dir;
return _antigravityDir;
}
}
_antigravityDir = '';
return null;
return antigravityResolver.resolve()?.directory ?? null;
}
/**
@@ -63,3 +55,7 @@ export function resolveAntigravityDir(): string | null {
export function isAntigravityAvailable(): boolean {
return resolveAntigravityDir() !== null;
}
export function getAntigravityNotFoundMessage(): string {
return formatCliNotFoundMessage(ANTIGRAVITY_NOT_FOUND, antigravityResolver.diagnostics());
}
+15 -28
View File
@@ -8,11 +8,11 @@
* @module utils/claude-cli-resolver
*/
import { execSync, execFileSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { execFileSync } from 'node:child_process';
import { delimiter, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import { createCliExecutableResolver, formatCliNotFoundMessage } from './cli-executable-resolver.js';
/** Common directories where the Claude CLI binary may be installed */
const CLAUDE_SEARCH_DIRS = [
@@ -23,8 +23,8 @@ const CLAUDE_SEARCH_DIRS = [
join(homedir(), 'bin'),
];
/** Cached directory containing the claude binary (empty string = searched but not found) */
let _claudeDir: string | null = null;
const claudeResolver = createCliExecutableResolver({ binary: 'claude', searchDirs: CLAUDE_SEARCH_DIRS });
const CLAUDE_NOT_FOUND = 'Claude CLI not found. Install it with: curl -fsSL https://claude.ai/install.sh | bash';
/**
* Returns true if the Claude CLI binary can be located (via `which` or one of
@@ -43,29 +43,11 @@ export function isClaudeAvailable(): boolean {
* @returns Directory path, or null if not found
*/
export function findClaudeDir(): string | null {
if (_claudeDir !== null) return _claudeDir || null;
return claudeResolver.resolve()?.directory ?? null;
}
// Try `which` first (respects current PATH)
try {
const result = execSync('which claude', { encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS }).trim();
if (result && existsSync(result)) {
_claudeDir = dirname(result);
return _claudeDir;
}
} catch {
// Claude not in PATH, will check common locations
}
// Fallback: check common installation directories
for (const dir of CLAUDE_SEARCH_DIRS) {
if (existsSync(join(dir, 'claude'))) {
_claudeDir = dir;
return _claudeDir;
}
}
_claudeDir = ''; // mark as searched, not found
return null;
export function getClaudeNotFoundMessage(): string {
return formatCliNotFoundMessage(CLAUDE_NOT_FOUND, claudeResolver.diagnostics());
}
/**
@@ -99,7 +81,9 @@ export function getAugmentedPath(): string {
const currentPath = process.env.PATH || '';
const claudeDir = findClaudeDir();
if (claudeDir && !currentPath.split(delimiter).includes(claudeDir)) {
if (!claudeDir) return currentPath;
if (!currentPath.split(delimiter).includes(claudeDir)) {
_augmentedPath = `${claudeDir}${delimiter}${currentPath}`;
return _augmentedPath;
}
@@ -193,6 +177,9 @@ function probeClaudeCliVersion(): string | null {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
env: { ...process.env, PATH: getAugmentedPath() },
// execFileSync's timeout only SENDS the signal and then keeps waiting; a
// child that ignores SIGTERM would block the server thread permanently.
killSignal: 'SIGKILL',
});
const match = out.match(/(\d+\.\d+\.\d+)/);
return match ? match[1] : null;
+301
View File
@@ -0,0 +1,301 @@
/**
* @fileoverview Shared CLI executable resolution for the per-CLI resolvers.
*
* One lookup chain behind all six *-cli-resolver modules (claude, opencode,
* codex, gemini, antigravity, pi): the server process PATH first, then the
* CLI's common install directories in order, then — last, because it is the
* only step that spawns anything — an interactive login shell, which is what
* finds nvm/Homebrew/user-npm installs when Codeman runs as a systemd/launchd
* service with a minimal PATH (launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin`).
*
* Caching is asymmetric, same shape as `resolveClaudeCliVersion` in
* claude-cli-resolver.ts: a successful resolution is cached for the process
* lifetime, a MISS is negative-cached and retried only after a doubling backoff
* (`cliResolveRetryDelayMs`). The callers are request-facing (the per-CLI
* status endpoints in system-routes.ts, the availability gates in
* session-routes.ts, and tmux-manager's spawn path), and the login-shell probe
* is a SYNCHRONOUS spawn bounded by `EXEC_TIMEOUT_MS` — without the negative
* cache, a missing CLI re-ran the whole chain and stalled the event loop for up
* to 5s on every request, forever.
*
* Test hermeticity: under vitest (`process.env.VITEST`) the production host
* short-circuits — IO primitives that were not injected become inert stubs, so
* a suite can never scan the machine's PATH or spawn login shells (the same
* rule as `IS_TEST_MODE` in tmux-manager and the VITEST gate in
* `getClaudeCliVersion`). Tests opt back in through the injection hooks
* (`runCommand`/`isExecutableFile` fakes do no real IO by construction) or, for
* fixtures that need the real filesystem predicate against their own temp
* files, via `allowRealIoUnderVitest`.
*
* @module utils/cli-executable-resolver
*/
import { execFileSync } from 'node:child_process';
import { accessSync, constants, statSync } from 'node:fs';
import { basename, delimiter, dirname, isAbsolute, join } from 'node:path';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import { loginShellArgs, resolveLocalShell } from './shell-resolver.js';
const SAFE_BINARY_NAME = /^[a-z0-9][a-z0-9._-]*$/i;
const LOGIN_SHELL_BEGIN_MARKER = '__CODEMAN_CLI_RESOLVE_BEGIN__';
const LOGIN_SHELL_END_MARKER = '__CODEMAN_CLI_RESOLVE_END__';
/** Maximum rendered length of each bounded diagnostic field, excluding its label. */
const DIAGNOSTIC_FIELD_MAX_LENGTH = 1024;
/** First retry window after a full-chain resolution miss. */
const RESOLVE_RETRY_BASE_MS = 60_000;
/**
* Ceiling for the doubling backoff. Deliberately shorter than the 15min cap on
* the claude version probe: that one is cosmetic, while this gates the Run
* flow, and "installing a CLI while the server is running is picked up without
* a restart" should stay true within minutes.
*/
const RESOLVE_RETRY_MAX_MS = 5 * 60_000;
/**
* How long to wait before re-running the resolution chain after `failures`
* consecutive misses: 1min, 2min, 4min… capped at 5min. Mirrors
* `claudeVersionRetryDelayMs` in claude-cli-resolver.ts. Exported for tests.
*/
export function cliResolveRetryDelayMs(failures: number): number {
if (failures <= 0) return 0;
return Math.min(RESOLVE_RETRY_BASE_MS * 2 ** (failures - 1), RESOLVE_RETRY_MAX_MS);
}
export type CliResolutionSource = 'process-path' | 'common-directory' | 'login-shell';
export interface CliResolutionDiagnostics {
binary: string;
processPath: string;
shellPath: string;
shellArgs: string[];
searchDirs: string[];
}
export interface CliResolverHost {
processPath: string;
shellPath: string;
shellArgs: string[];
findOnProcessPath(binary: string): string | null;
findInLoginShell(binary: string): string | null;
exists(path: string): boolean;
}
export interface CandidateValidation<T> {
accepted: boolean;
metadata?: T;
}
export interface CliResolution<T = undefined> {
binaryPath: string;
directory: string;
source: CliResolutionSource;
metadata?: T;
}
export interface CliExecutableResolver<T = undefined> {
resolve(): CliResolution<T> | null;
diagnostics(): CliResolutionDiagnostics;
}
export interface CliResolverCommandOptions {
encoding: 'utf8';
timeout: number;
stdio: ['ignore', 'pipe', 'ignore'];
killSignal: 'SIGKILL';
}
export type CliResolverCommandRunner = (file: string, args: string[], options: CliResolverCommandOptions) => string;
export interface ProductionCliResolverHostOptions {
processPath?: string;
shellPath?: string;
shellArgs?: string[];
runCommand?: CliResolverCommandRunner;
isExecutableFile?: (path: string) => boolean;
/**
* Test-only escape hatch: keep the REAL IO primitives even under vitest.
* For tests that exercise `isExecutableRegularFile` against their own temp
* fixtures. Such a test must still inject `runCommand` if it can reach the
* login-shell step, or it would spawn a real interactive shell.
*/
allowRealIoUnderVitest?: boolean;
}
function isExecutableRegularFile(path: string): boolean {
try {
if (!statSync(path).isFile()) return false;
accessSync(path, constants.X_OK);
return true;
} catch {
return false;
}
}
function parseLoginShellResult(output: string, binary: string): string | null {
const lines = output.split(/\r?\n/).map((line) => line.trim());
const begin = lines.indexOf(LOGIN_SHELL_BEGIN_MARKER);
if (begin === -1) return null;
const end = lines.indexOf(LOGIN_SHELL_END_MARKER, begin + 1);
if (end === -1) return null;
for (const candidate of lines.slice(begin + 1, end)) {
if (isAbsolute(candidate) && basename(candidate) === binary) return candidate;
}
return null;
}
function loginShellCommand(binary: string): string {
return [
`printf '%s\\n' '${LOGIN_SHELL_BEGIN_MARKER}'`,
`command -v -- ${binary}`,
`printf '%s\\n' '${LOGIN_SHELL_END_MARKER}'`,
].join('; ');
}
export function createProductionCliResolverHost(options: ProductionCliResolverHostOptions = {}): CliResolverHost {
const shellPath = options.shellPath ?? resolveLocalShell();
const shellArgs = options.shellArgs ?? loginShellArgs(shellPath).trim().split(/\s+/).filter(Boolean);
const processPath = options.processPath ?? process.env.PATH ?? '';
// Hermeticity gate (see @fileoverview): under vitest, any IO primitive the
// caller did not inject is replaced by an inert stub. The suites must never
// depend on — or execute — whatever happens to be installed on the machine
// running them, and route tests hitting the per-CLI status endpoints would
// otherwise scan the real PATH and spawn real login shells on CI.
const inert = Boolean(process.env.VITEST) && options.allowRealIoUnderVitest !== true;
const isExecutableFile = options.isExecutableFile ?? (inert ? () => false : isExecutableRegularFile);
const runCommand: CliResolverCommandRunner =
options.runCommand ?? (inert ? () => '' : (file, args, commandOptions) => execFileSync(file, args, commandOptions));
const run = (file: string, args: string[]): string => {
try {
return runCommand(file, args, {
encoding: 'utf8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
// SIGKILL is load-bearing: execFileSync's `timeout` only SENDS the kill
// signal and then keeps waiting for the child to exit. Interactive bash
// ignores SIGTERM (the default), so a login shell stuck in a blocking
// .bash_profile would survive the timeout and block the server forever.
killSignal: 'SIGKILL',
});
} catch {
return '';
}
};
return {
processPath,
shellPath,
shellArgs: [...shellArgs],
findOnProcessPath: (binary) => {
if (!SAFE_BINARY_NAME.test(binary)) return null;
for (const directory of processPath.split(delimiter).filter(Boolean)) {
const candidate = join(directory, binary);
if (isAbsolute(candidate) && isExecutableFile(candidate)) return candidate;
}
return null;
},
findInLoginShell: (binary) => {
if (!SAFE_BINARY_NAME.test(binary)) return null;
const candidate = parseLoginShellResult(run(shellPath, [...shellArgs, '-c', loginShellCommand(binary)]), binary);
return candidate && isExecutableFile(candidate) ? candidate : null;
},
exists: isExecutableFile,
};
}
export function createCliExecutableResolver<T = undefined>(
options: {
binary: string;
searchDirs: string[];
validateCandidate?: (path: string) => CandidateValidation<T>;
/** Clock injection for tests driving the failure backoff. Defaults to `Date.now`. */
now?: () => number;
},
host: CliResolverHost = createProductionCliResolverHost()
): CliExecutableResolver<T> {
if (!SAFE_BINARY_NAME.test(options.binary)) {
throw new Error(`Unsafe CLI binary name: ${options.binary}`);
}
const now = options.now ?? Date.now;
/** Successful resolution, cached for the process lifetime. */
let cached: CliResolution<T> | null = null;
/** Consecutive full-chain misses (drives the retry backoff). */
let failures = 0;
/** Timestamp of the most recent miss. */
let lastFailureAt = 0;
const accept = (path: string | null, source: CliResolutionSource): CliResolution<T> | null => {
if (!path || !isAbsolute(path) || !host.exists(path)) return null;
const validation = options.validateCandidate?.(path) ?? ({ accepted: true } as CandidateValidation<T>);
if (!validation.accepted) return null;
return {
binaryPath: path,
directory: dirname(path),
source,
metadata: validation.metadata,
};
};
return {
resolve() {
if (cached) return cached;
// Negative cache: a miss is remembered and the chain — whose login-shell
// tail is a synchronous 5s-bounded spawn — is not re-run until the
// backoff elapses. Without this, every status poll and Run click against
// a missing CLI froze the event loop for the full probe, forever.
if (failures > 0 && now() - lastFailureAt < cliResolveRetryDelayMs(failures)) return null;
cached = accept(host.findOnProcessPath(options.binary), 'process-path');
if (!cached) {
for (const dir of options.searchDirs) {
cached = accept(join(dir, options.binary), 'common-directory');
if (cached) break;
}
}
if (!cached) {
cached = accept(host.findInLoginShell(options.binary), 'login-shell');
}
if (cached) {
failures = 0;
lastFailureAt = 0;
return cached;
}
failures += 1;
lastFailureAt = now();
return null;
},
diagnostics: () => ({
binary: options.binary,
processPath: host.processPath,
shellPath: host.shellPath,
shellArgs: [...host.shellArgs],
searchDirs: [...options.searchDirs],
}),
};
}
function sanitizeDiagnosticField(value: string, emptyMarker: string): string {
const flattened = Array.from(value, (character) => {
const codePoint = character.codePointAt(0) ?? 0;
const isControl = codePoint <= 0x1f || (codePoint >= 0x7f && codePoint <= 0x9f);
return isControl || codePoint === 0x2028 || codePoint === 0x2029 ? ' ' : character;
})
.join('')
.replace(/ +/g, ' ')
.trim();
if (!flattened) return emptyMarker;
if (flattened.length <= DIAGNOSTIC_FIELD_MAX_LENGTH) return flattened;
return `${flattened.slice(0, DIAGNOSTIC_FIELD_MAX_LENGTH - 1)}…`;
}
export function formatCliNotFoundMessage(base: string, diagnostics: CliResolutionDiagnostics): string {
const processPath = sanitizeDiagnosticField(diagnostics.processPath, '(empty)');
const shell = sanitizeDiagnosticField(
[diagnostics.shellPath, ...diagnostics.shellArgs].filter(Boolean).join(' '),
'(none)'
);
const dirs = sanitizeDiagnosticField(diagnostics.searchDirs.join(', '), '(none)');
return `${base}\nServer PATH: ${processPath}\nLogin shell: ${shell}\nChecked directories: ${dirs}`;
}
+9 -31
View File
@@ -7,11 +7,9 @@
* @module utils/codex-cli-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import { createCliExecutableResolver, formatCliNotFoundMessage } from './cli-executable-resolver.js';
/** Common directories where the Codex CLI binary may be installed */
const CODEX_SEARCH_DIRS = [
@@ -23,8 +21,8 @@ const CODEX_SEARCH_DIRS = [
join(homedir(), 'bin'), // User bin
];
/** Cached directory containing the codex binary (empty string = searched but not found) */
let _codexDir: string | null = null;
const codexResolver = createCliExecutableResolver({ binary: 'codex', searchDirs: CODEX_SEARCH_DIRS });
const CODEX_NOT_FOUND = 'Codex CLI not found. Install with: npm install -g @openai/codex';
/**
* Finds the directory containing the `codex` binary.
@@ -34,31 +32,7 @@ let _codexDir: string | null = null;
* @returns Directory path, or null if not found
*/
export function resolveCodexDir(): string | null {
if (_codexDir !== null) return _codexDir || null;
// Try `which` first (respects current PATH)
try {
const result = execSync('which codex', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
_codexDir = dirname(result);
return _codexDir;
}
} catch {
// Codex not in PATH, will check common locations
}
for (const dir of CODEX_SEARCH_DIRS) {
if (existsSync(join(dir, 'codex'))) {
_codexDir = dir;
return _codexDir;
}
}
_codexDir = ''; // mark as searched, not found
return null;
return codexResolver.resolve()?.directory ?? null;
}
/**
@@ -67,3 +41,7 @@ export function resolveCodexDir(): string | null {
export function isCodexAvailable(): boolean {
return resolveCodexDir() !== null;
}
export function getCodexNotFoundMessage(): string {
return formatCliNotFoundMessage(CODEX_NOT_FOUND, codexResolver.diagnostics());
}
+93
View File
@@ -0,0 +1,93 @@
/**
* @fileoverview Pure file-name/path query matcher for the Files panel search
* (COD-236). Compiles a user query string into a reusable predicate so the
* server-side file walk can prune to matching entries instead of streaming the
* whole tree.
*
* Semantics:
* - Empty / whitespace-only query → `compileFileQuery` returns `null` (the
* caller treats this as "no search", falling back to the full tree). A query
* longer than `MAX_QUERY_LENGTH` compiles to `null` too: no honest filename
* search is that long, and the glob walk below is O(text · pattern).
* - A query containing a glob metachar (`*` or `?`) matches anchored and
* case-insensitively, `*` spanning any run (slashes included) and `?` exactly
* one character; every other character matches literally.
* - Otherwise the query is a plain case-insensitive substring.
* - When the query contains a `/` it matches against the relative path; else it
* matches against the bare entry name.
*
* ⚠️ Globs are matched by `globMatch` below, never by compiling the query into
* a RegExp: `*a*a*a…` translated to `^.*a.*a.*a…$` is a classic backtracking
* blowup, evaluated synchronously against every walked path — a pathological
* query could freeze the event loop for the whole server (the same reason
* `search-service.ts` is regex-free). The two-pointer wildcard walk is
* O(text · pattern) worst case, with both operands short by construction.
*
* No fs / IO — safe to unit-test directly.
*/
export type FileQueryMatcher = (name: string, relativePath: string) => boolean;
// Longer than any honest file search; bounds the O(text · pattern) glob walk.
const MAX_QUERY_LENGTH = 256;
/**
* Anchored glob match, linear-space two-pointer walk (no RegExp — see the
* fileoverview). `pattern` must already be lowercased; `text` is lowercased
* here so one compiled matcher serves many entries.
*/
function globMatch(pattern: string, rawText: string): boolean {
const text = rawText.toLowerCase();
let p = 0;
let t = 0;
let starP = -1;
let starT = -1;
while (t < text.length) {
const pc = p < pattern.length ? pattern[p] : '';
if (pc === '?' || pc === text[t]) {
p++;
t++;
} else if (pc === '*') {
// Remember the star; try matching zero characters first, and on a later
// mismatch re-expand it one character at a time from here.
starP = p++;
starT = t;
} else if (starP !== -1) {
p = starP + 1;
t = ++starT;
} else {
return false;
}
}
while (p < pattern.length && pattern[p] === '*') p++;
return p === pattern.length;
}
/**
* Compile a query string into a matcher predicate, or `null` when the query is
* empty/whitespace or overlong (caller treats null as "no search").
*/
export function compileFileQuery(query: string): FileQueryMatcher | null {
const trimmed = query.trim();
if (trimmed === '' || trimmed.length > MAX_QUERY_LENGTH) return null;
const matchesPath = trimmed.includes('/');
const isGlob = trimmed.includes('*') || trimmed.includes('?');
if (isGlob) {
const pattern = trimmed.toLowerCase();
return (name, relativePath) => globMatch(pattern, matchesPath ? relativePath : name);
}
const needle = trimmed.toLowerCase();
return (name, relativePath) => (matchesPath ? relativePath : name).toLowerCase().includes(needle);
}
/**
* Convenience: compile the query and apply it in one call. Returns false when
* the query compiles to null (empty).
*/
export function matchFileQuery(query: string, name: string, relativePath: string): boolean {
const matcher = compileFileQuery(query);
return matcher ? matcher(name, relativePath) : false;
}
+9 -30
View File
@@ -7,11 +7,9 @@
* @module utils/gemini-cli-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import { createCliExecutableResolver, formatCliNotFoundMessage } from './cli-executable-resolver.js';
/** Common directories where the Gemini CLI binary may be installed */
const GEMINI_SEARCH_DIRS = [
@@ -23,8 +21,8 @@ const GEMINI_SEARCH_DIRS = [
join(homedir(), 'bin'),
];
/** Cached directory containing the gemini binary (empty string = searched but not found) */
let _geminiDir: string | null = null;
const geminiResolver = createCliExecutableResolver({ binary: 'gemini', searchDirs: GEMINI_SEARCH_DIRS });
const GEMINI_NOT_FOUND = 'Gemini CLI not found. Install with: npm install -g @google/gemini-cli';
/**
* Finds the directory containing the `gemini` binary.
@@ -33,30 +31,7 @@ let _geminiDir: string | null = null;
* @returns Directory path, or null if not found
*/
export function resolveGeminiDir(): string | null {
if (_geminiDir !== null) return _geminiDir || null;
try {
const result = execSync('which gemini', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
_geminiDir = dirname(result);
return _geminiDir;
}
} catch {
// Gemini not in PATH, will check common locations
}
for (const dir of GEMINI_SEARCH_DIRS) {
if (existsSync(join(dir, 'gemini'))) {
_geminiDir = dir;
return _geminiDir;
}
}
_geminiDir = '';
return null;
return geminiResolver.resolve()?.directory ?? null;
}
/**
@@ -65,3 +40,7 @@ export function resolveGeminiDir(): string | null {
export function isGeminiAvailable(): boolean {
return resolveGeminiDir() !== null;
}
export function getGeminiNotFoundMessage(): string {
return formatCliNotFoundMessage(GEMINI_NOT_FOUND, geminiResolver.diagnostics());
}
+18 -6
View File
@@ -28,10 +28,22 @@ export { stringSimilarity, fuzzyPhraseMatch, todoContentHash } from './string-si
export { assertNever } from './type-safety.js';
export { wrapWithNice } from './nice-wrapper.js';
export { resolveLocalShell, loginShellArgs } from './shell-resolver.js';
export { findClaudeDir, getAugmentedPath, getClaudeCliVersion, getClaudeBinaryPath } from './claude-cli-resolver.js';
export {
findClaudeDir,
getAugmentedPath,
getClaudeCliVersion,
getClaudeBinaryPath,
getClaudeNotFoundMessage,
} from './claude-cli-resolver.js';
export { spawnPtyWithHelperRepair } from './node-pty-repair.js';
export { resolveOpenCodeDir } from './opencode-cli-resolver.js';
export { resolveCodexDir, isCodexAvailable } from './codex-cli-resolver.js';
export { resolveGeminiDir, isGeminiAvailable } from './gemini-cli-resolver.js';
export { resolveAntigravityDir, isAntigravityAvailable } from './antigravity-cli-resolver.js';
export { resolvePiDir, isPiAvailable, getPiCliVersion } from './pi-cli-resolver.js';
export { resolveOpenCodeDir, getOpenCodeNotFoundMessage } from './opencode-cli-resolver.js';
export { resolveCodexDir, isCodexAvailable, getCodexNotFoundMessage } from './codex-cli-resolver.js';
export { resolveGeminiDir, isGeminiAvailable, getGeminiNotFoundMessage } from './gemini-cli-resolver.js';
export {
resolveAntigravityDir,
isAntigravityAvailable,
getAntigravityNotFoundMessage,
} from './antigravity-cli-resolver.js';
export { resolvePiDir, isPiAvailable, getPiCliVersion, getPiNotFoundMessage } from './pi-cli-resolver.js';
export { compileFileQuery, matchFileQuery } from './file-query.js';
export type { FileQueryMatcher } from './file-query.js';
+9 -32
View File
@@ -7,11 +7,9 @@
* @module utils/opencode-cli-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import { createCliExecutableResolver, formatCliNotFoundMessage } from './cli-executable-resolver.js';
/** Common directories where the OpenCode CLI binary may be installed */
const OPENCODE_SEARCH_DIRS = [
@@ -24,8 +22,8 @@ const OPENCODE_SEARCH_DIRS = [
join(homedir(), 'bin'), // User bin
];
/** Cached directory containing the opencode binary (empty string = searched but not found) */
let _openCodeDir: string | null = null;
const openCodeResolver = createCliExecutableResolver({ binary: 'opencode', searchDirs: OPENCODE_SEARCH_DIRS });
const OPENCODE_NOT_FOUND = 'OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash';
/**
* Finds the directory containing the `opencode` binary.
@@ -35,32 +33,7 @@ let _openCodeDir: string | null = null;
* @returns Directory path, or null if not found
*/
export function resolveOpenCodeDir(): string | null {
if (_openCodeDir !== null) return _openCodeDir || null;
// Try `which` first (respects current PATH)
try {
const result = execSync('which opencode', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
_openCodeDir = dirname(result);
return _openCodeDir;
}
} catch {
// OpenCode not in PATH, will check common locations
}
// Fallback: check common installation directories
for (const dir of OPENCODE_SEARCH_DIRS) {
if (existsSync(join(dir, 'opencode'))) {
_openCodeDir = dir;
return _openCodeDir;
}
}
_openCodeDir = ''; // mark as searched, not found
return null;
return openCodeResolver.resolve()?.directory ?? null;
}
/**
@@ -69,3 +42,7 @@ export function resolveOpenCodeDir(): string | null {
export function isOpenCodeAvailable(): boolean {
return resolveOpenCodeDir() !== null;
}
export function getOpenCodeNotFoundMessage(): string {
return formatCliNotFoundMessage(OPENCODE_NOT_FOUND, openCodeResolver.diagnostics());
}
+53 -53
View File
@@ -15,11 +15,15 @@
* @module utils/pi-cli-resolver
*/
import { execFileSync, execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { execFileSync } from 'node:child_process';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import {
createCliExecutableResolver,
formatCliNotFoundMessage,
type CliResolverHost,
} from './cli-executable-resolver.js';
/** Common directories where the Pi CLI binary may be installed */
const PI_SEARCH_DIRS = [
@@ -45,10 +49,7 @@ const PI_SEARCH_DIRS = [
*/
export const PI_VERSION_REGEX = /(?:^|\s)(\d+\.\d+\.\d+)/;
/** Cached directory containing the pi binary (empty string = searched but not found) */
let _piDir: string | null = null;
/** Cached version string reported by the resolved binary (empty string = probed, unusable) */
let _piVersion: string | null = null;
const PI_NOT_FOUND = 'Pi CLI not found. Install with: npm install -g --ignore-scripts @earendil-works/pi-coding-agent';
/**
* Run `pi --version` on a candidate path and return the trimmed version when it
@@ -57,7 +58,12 @@ let _piVersion: string | null = null;
* (which is how an unrelated `pi` on PATH gets rejected).
*
* Never runs under vitest: the suites must stay hermetic and must not depend on
* whether the dev box happens to have pi installed.
* whether the dev box happens to have pi installed — and since `pi` is a short
* GENERIC name, this probe would EXECUTE whatever binary of that name the
* machine carries. The shared resolver host is already inert under vitest, so
* this gate is defense in depth for any opted-in host that still carries the
* default probe; tests drive resolution via `createPiResolverForTest`, whose
* injected probe bypasses it. Pinned by test/pi-cli-resolver.test.ts.
*/
function probePiVersion(binPath: string): string | null {
if (process.env.VITEST) return null;
@@ -66,6 +72,9 @@ function probePiVersion(binPath: string): string | null {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
// A stuck or hostile `pi` that ignores SIGTERM would survive the timeout
// and block the server (execFileSync keeps waiting after the signal).
killSignal: 'SIGKILL',
}).trim();
// Upstream prints a bare version today; tolerate a `pi 0.84.1` style prefix too.
const candidate = PI_VERSION_REGEX.exec(out)?.[1];
@@ -77,6 +86,34 @@ function probePiVersion(binPath: string): string | null {
return null;
}
type PiVersionProbe = (binPath: string) => string | null;
function createPiResolver(host?: CliResolverHost, versionProbe: PiVersionProbe = probePiVersion, now?: () => number) {
return createCliExecutableResolver<string>(
{
binary: 'pi',
searchDirs: PI_SEARCH_DIRS,
validateCandidate: (binPath) => {
const version = versionProbe(binPath);
return version ? { accepted: true, metadata: version } : { accepted: false };
},
now,
},
host
);
}
/**
* Creates an isolated Pi wrapper around an injected host, version probe and
* clock. Omitting `versionProbe` keeps the ambient (VITEST-gated) probe, which
* is exactly what the hermeticity test exercises.
*/
export function createPiResolverForTest(host: CliResolverHost, versionProbe?: PiVersionProbe, now?: () => number) {
return createPiResolver(host, versionProbe ?? probePiVersion, now);
}
const piResolver = createPiResolver();
/**
* Finds the directory containing a verified `pi` binary.
* Checks `which pi` first, then falls back to common install locations. Every
@@ -86,46 +123,7 @@ function probePiVersion(binPath: string): string | null {
* @returns Directory path, or null if not found
*/
export function resolvePiDir(): string | null {
if (_piDir !== null) return _piDir || null;
const accept = (binPath: string): string | null => {
// Under vitest the probe never runs, so existence alone decides (keeps the
// suites hermetic and matches how the sibling resolvers behave there).
if (process.env.VITEST) {
_piDir = dirname(binPath);
_piVersion = '';
return _piDir;
}
const version = probePiVersion(binPath);
if (!version) return null;
_piDir = dirname(binPath);
_piVersion = version;
return _piDir;
};
try {
const result = execSync('which pi', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
const dir = accept(result);
if (dir) return dir;
}
} catch {
// pi not in PATH, will check common locations
}
for (const dir of PI_SEARCH_DIRS) {
const binPath = join(dir, 'pi');
if (!existsSync(binPath)) continue;
const accepted = accept(binPath);
if (accepted) return accepted;
}
_piDir = '';
_piVersion = '';
return null;
return piResolver.resolve()?.directory ?? null;
}
/**
@@ -135,12 +133,14 @@ export function isPiAvailable(): boolean {
return resolvePiDir() !== null;
}
export function getPiNotFoundMessage(): string {
return formatCliNotFoundMessage(PI_NOT_FOUND, piResolver.diagnostics());
}
/**
* Version reported by the resolved `pi` binary, or null when pi is unavailable
* (or when the probe was skipped, i.e. under vitest). Surfaced through
* `GET /api/pi/status` so a misresolution is diagnosable from the UI.
* Version reported by the resolved `pi` binary, or null when pi is unavailable.
* Surfaced through `GET /api/pi/status` so a misresolution is diagnosable from the UI.
*/
export function getPiCliVersion(): string | null {
resolvePiDir();
return _piVersion || null;
return piResolver.resolve()?.metadata ?? null;
}
+110 -26
View File
@@ -22,14 +22,14 @@
* - Acknowledgement (`acknowledge()`, idle items only) is NOT resolution: the
* item stays pending, it just stops arming the tab alert on every client.
*
* @dependencies utils (stripAnsi)
* @dependencies utils (stripAnsi, CLAUDE_WORKING_LINE_PATTERN)
* @consumedby web/routes/hook-event-routes (notePrompt/resolve), web/routes/approval-routes,
* web/session-listener-wiring (working/exit resolution), web/server (emit callbacks + stop)
*
* @module web/approval-inbox
*/
import { stripAnsi } from '../utils/index.js';
import { stripAnsi, CLAUDE_WORKING_LINE_PATTERN } from '../utils/index.js';
// ─── Types ───────────────────────────────────────────────────────────────────
@@ -106,6 +106,12 @@ const ITEM_TTL_MS = 12 * 60 * 60 * 1000;
* single delayed re-capture picks up the frame the immediate capture missed.
*/
const RECAPTURE_DELAY_MS = 600;
/**
* Delayed staleness pass for the late-hook case (see notePrompt). Comfortably
* clear of RECAPTURE_DELAY_MS so a dialog Ink has not painted yet is never
* mistaken for one that is gone.
*/
const STALE_CHECK_DELAY_MS = 3000;
/** Context kept per item: enough for a dialog plus a few lines above it. */
const MAX_CONTEXT_CHARS = 4000;
const MAX_CONTEXT_LINES = 30;
@@ -195,7 +201,8 @@ export function parseDialogOptions(context: string | undefined): ApprovalOption[
export class ApprovalInbox {
/** Keyed by sessionId; the one-active-item-per-session invariant lives here. */
private items = new Map<string, ApprovalItem>();
private recaptureTimers = new Map<string, ReturnType<typeof setTimeout>>();
/** Post-capture timers per item id (re-capture + the delayed staleness check). */
private itemTimers = new Map<string, ReturnType<typeof setTimeout>[]>();
/** Capture callbacks kept for answer-time re-verification; dropped on remove. */
private captures = new Map<string, () => string | null>();
private seq = 0;
@@ -229,31 +236,55 @@ export class ApprovalInbox {
if (args.capture) this.captures.set(args.sessionId, args.capture);
this.onPending?.(item);
if (args.capture && !this.stopped) {
const timer = setTimeout(() => {
this.recaptureTimers.delete(item.id);
// Only update the item if it is still the live one for the session.
if (this.items.get(args.sessionId)?.id !== item.id) return;
// Pass 1 (600ms): enrich the card with the painted frame.
this.scheduleForItem(item, RECAPTURE_DELAY_MS, () => {
this.applyCapture(item, args.capture);
this.onUpdated?.(item);
}, RECAPTURE_DELAY_MS);
this.recaptureTimers.set(item.id, timer);
});
// Pass 2: the late-hook staleness check. Claude Code fires the
// Notification behind the dialog, so a prompt answered before the hook
// lands creates an item for a dialog that is ALREADY gone: nothing ever
// parsed, so the "options vanished" test can never fire, `stop` may have
// gone by already, and the red alert then outlived reloads until the 12h
// TTL. This pass re-reads the pane and resolves when the frame proves no
// dialog is up. Deliberately LATER than the re-capture, whose whole
// reason for existing is that Ink may not have painted the dialog yet:
// resolving inside that window could clear the alert for a dialog that
// was about to appear.
this.scheduleForItem(item, STALE_CHECK_DELAY_MS, () => {
this.verifyStillAnswerable(item.id);
});
}
return item;
}
/**
* Answer-time guard: re-capture the pane and check the dialog is still on
* screen before keystrokes are sent at it. Only conclusive when the ORIGINAL
* frame parsed options: if a fresh capture then parses none, the dialog is
* gone (answered in the terminal moments ago), so the item resolves and the
* answer must be refused, because the digit would land in whatever now has
* focus. Unparseable-from-the-start items stay answerable (approve/deny
* only), same risk the terminal user already carries.
* screen before keystrokes are sent at it. If the dialog is gone (answered in
* the terminal moments ago) the item resolves and the answer is refused,
* because the digit would land in whatever now has focus.
*
* A fresh frame that parses NO options is conclusive in two cases, and only
* those; anything else stays answerable, so an unreadable capture keeps the
* alert rather than losing a live dialog:
*
* 1. The item HAD parsed options. They cannot vanish while the dialog is up.
* 2. The frame shows Claude actively running a turn. A modal dialog BLOCKS
* the turn, so a working line and a dialog cannot coexist — measured on
* v2.1.237: a live-dialog frame carries neither the `… (13s` timer nor
* even the `esc to interrupt` footer, which the dialog replaces with
* `Enter to select · ↑/↓ to navigate · Esc to cancel`.
*
* Case 2 is what closes the late-hook hole. Claude Code fires the
* Notification behind the dialog, so a prompt answered before the hook lands
* produces an item whose FIRST capture already has no dialog in it — never
* parsed, so case 1 can never fire, and the red alert then outlived even
* `stop` (which had already fired) and survived reloads until the 12h TTL.
*/
verifyStillAnswerable(id: string): boolean {
const item = this.getById(id);
if (!item) return false;
if (item.kind === 'idle' || !item.options) return true;
if (item.kind === 'idle') return true;
const capture = this.captures.get(item.sessionId);
if (!capture) return true;
let raw: string | null = null;
@@ -266,14 +297,39 @@ export class ApprovalInbox {
if (!context) return true;
const options = parseDialogOptions(context);
if (!options) {
this.remove(item, 'resolved_in_terminal');
return false;
if (item.options || CLAUDE_WORKING_LINE_PATTERN.test(context)) {
this.remove(item, 'resolved_in_terminal');
return false;
}
return true; // never parsed and the pane is not visibly working: unreadable, not gone
}
item.context = context;
item.options = options;
return true;
}
/**
* "This session's pane started moving again": re-verify its pending DIALOG
* item against the screen and resolve it if the dialog is gone.
*
* The staleness check itself lived only in `GET /api/approvals`, which
* nothing calls while a page is open (`seedApprovals()` runs on init and
* reconnect), so a dialog answered in the terminal kept its red tab alert for
* the whole rest of the turn. The `working` signal is exactly the moment an
* answer lands, and routing it through `verifyStillAnswerable` is what makes
* it safe to act on for a permission/question item: `working` is heuristic
* and can flap, but it only decides WHEN to look — the pane decides the
* outcome, and an unreadable capture keeps the alert.
*
* Cheap by construction: a Map miss unless a dialog item is actually pending,
* and the item is gone after the first successful resolve.
*/
resolveIfDialogGone(sessionId: string): void {
const item = this.getForSession(sessionId);
if (!item || item.kind === 'idle') return;
this.verifyStillAnswerable(item.id);
}
/** Pending item for a session, TTL-checked. */
getForSession(sessionId: string): ApprovalItem | undefined {
const item = this.items.get(sessionId);
@@ -359,12 +415,27 @@ export class ApprovalInbox {
/** Clear all timers (shutdown/tests). Items become inert; no events fire after this. */
stop(): void {
this.stopped = true;
for (const timer of this.recaptureTimers.values()) clearTimeout(timer);
this.recaptureTimers.clear();
for (const timers of this.itemTimers.values()) for (const timer of timers) clearTimeout(timer);
this.itemTimers.clear();
this.items.clear();
this.captures.clear();
}
/**
* Run `fn` after `delayMs` if the item is still the live one for its session,
* tracking the timer so `remove()`/`stop()` can cancel it.
*/
private scheduleForItem(item: ApprovalItem, delayMs: number, fn: () => void): void {
const timer = setTimeout(() => {
const timers = this.itemTimers.get(item.id)?.filter((t) => t !== timer) ?? [];
if (timers.length > 0) this.itemTimers.set(item.id, timers);
else this.itemTimers.delete(item.id);
if (this.items.get(item.sessionId)?.id !== item.id) return;
fn();
}, delayMs);
this.itemTimers.set(item.id, [...(this.itemTimers.get(item.id) ?? []), timer]);
}
private applyCapture(item: ApprovalItem, capture?: () => string | null): void {
if (!capture) return;
let raw: string | null = null;
@@ -377,17 +448,30 @@ export class ApprovalInbox {
if (!context) return;
item.context = context;
// Idle prompts are not dialogs; never offer digit answers for them.
if (item.kind !== 'idle') item.options = parseDialogOptions(context);
if (item.kind === 'idle') return;
const options = parseDialogOptions(context);
// ⚠️ ADD-ONLY: a re-capture that parses NOTHING must never erase options a
// previous capture found. Claude Code delays the Notification hook behind
// the dialog (measured 6s here, up to ~30s), so the 600ms re-capture very
// often lands AFTER the user has already answered in the terminal, on a
// frame with no dialog in it. Clearing the field there was the whole bug:
// `verifyStillAnswerable` reads a MISSING `options` as "never parsed" and
// keeps such an item answerable by design, so a cleared field made the item
// permanently unsweepable — the red "needs you" alert then survived every
// `GET /api/approvals` and every page reload and only went away on `stop`
// (owner report 2026-08-20: a confirmed question left a tab flowing red for
// ~8 minutes while the turn ran on), and the stale card still accepted an
// answer, typing a bare `1` into a composer with no dialog under it.
// Keeping the parse means a later capture is CONCLUSIVE: options present +
// fresh frame without them == answered in the terminal.
if (options) item.options = options;
}
private remove(item: ApprovalItem, resolution: ApprovalResolution): void {
this.items.delete(item.sessionId);
this.captures.delete(item.sessionId);
const timer = this.recaptureTimers.get(item.id);
if (timer) {
clearTimeout(timer);
this.recaptureTimers.delete(item.id);
}
for (const timer of this.itemTimers.get(item.id) ?? []) clearTimeout(timer);
this.itemTimers.delete(item.id);
if (!this.stopped) {
this.onResolved?.({ id: item.id, sessionId: item.sessionId, kind: item.kind, resolution });
}
+434 -58
View File
@@ -481,6 +481,14 @@ const DEFAULT_SHORTCUTS = [
// CodemanApp Class — constructor and global state
// ═══════════════════════════════════════════════════════════════
/**
* How often the rich sidebar rewrites its relative stamps in place. Matches the
* two home screens (mobile-overview.js, home-sessions.js). Deliberately a local
* const and not a constants.js export: an undefined interval would make
* setInterval fire on every frame, and constants.js is cached independently.
*/
const SIDEBAR_RICH_CLOCK_MS = 20000;
class CodemanApp {
constructor() {
this.sessions = new Map();
@@ -528,10 +536,10 @@ class CodemanApp {
this._initGeneration = 0; // dedup concurrent handleInit calls
this._initFallbackTimer = null; // fallback timer if SSE init doesn't arrive
this._selectGeneration = 0; // cancel stale selectSession loads
// Sessions whose full tmux scrollback has already been replayed this page load
// (COD-47). Tracked PER SESSION rather than as a single "first load" flag: the
// flag was consumed by whichever session auto-selected at page load, so every
// OTHER tab started life with one visible frame of history (issue #205).
// Non-shell sessions whose full tmux scrollback has already been replayed this
// page load (COD-47). Shells deliberately start from a bounded tail because
// their scrollback can be very large; full history stays available on demand.
// Tracked PER SESSION rather than as a single "first load" flag (issue #205).
this._fullHistoryLoaded = new Set();
// Cooldown per session for the scroll-to-top "load more history" re-pull.
this._fullHistoryRepullAt = new Map(); // Map<sessionId, timestamp>
@@ -652,6 +660,11 @@ class CodemanApp {
// Tracks pending hook events that need resolution (permission_prompt, elicitation_dialog, idle_prompt)
this.pendingHooks = new Map();
// Sessions THIS tab is closing right now. closeSession() owns the follow-up
// selection, so _onSessionDeleted must not race its own delete's SSE
// broadcast to the welcome screen. Set<sessionId>, cleared in a finally.
this._closingSessions = new Set();
// Approvals Inbox: Map<approvalId, ApprovalItem> (methods in approvals-ui.js)
this.approvals = new Map();
@@ -1813,7 +1826,13 @@ class CodemanApp {
// Dashboard: a detached session ended → clear its detached state/timers.
if (this.detachedSessions.has(data.id)) this._redock(data.id);
this._cleanupSessionData(data.id);
if (this.activeSessionId === data.id) {
// ⚠️ Skip the whole active-session handoff while THIS tab is closing that
// session: closeSession() owns the follow-up selection and moves you to the
// next tab, so acting here would race it and flash the welcome screen (or
// strand you on it) for a close the user initiated right here. A delete from
// anywhere else still lands on the home screen, which is the honest answer
// when the thing you were looking at was taken away.
if (this.activeSessionId === data.id && !this._closingSessions.has(data.id)) {
this.activeSessionId = null;
try { localStorage.removeItem('codeman-active-session'); } catch {}
this.terminal.clear();
@@ -2034,6 +2053,28 @@ class CodemanApp {
wrap.appendChild(actions);
wrap.appendChild(pre);
});
// Links open in a NEW tab.
//
// marked emits a bare `<a href>` and the sanitizer's allowlist has no
// `target`, so a tap in the chat NAVIGATED THE APP AWAY: on a phone that
// unloads the whole dashboard — SSE, terminal buffers, unsent composer
// text — and the OS back gesture reloads it from scratch, which is what
// "links don't open" reads as on mobile, with no middle-click or
// open-in-new-tab affordance to work around it.
//
// This pass runs AFTER sanitizing, so it is the only source of these two
// attributes: whatever an agent wrote is already gone, and `rel` is set on
// the same element in the same breath, so no page Codeman opens ever gets
// a `window.opener` handle back (reverse tabnabbing).
//
// A fragment link stays in-page, and mailto:/tel: are handed to the OS —
// giving those a target just strands an empty tab.
tmpl.content.querySelectorAll('a[href]').forEach((a) => {
const href = a.getAttribute('href') || '';
if (!href || href.startsWith('#') || /^(?:mailto|tel):/i.test(href)) return;
a.setAttribute('target', '_blank');
a.setAttribute('rel', 'noopener noreferrer');
});
return tmpl.innerHTML;
} catch { /* fall through */ }
}
@@ -3693,18 +3734,26 @@ class CodemanApp {
// ═══════════════════════════════════════════════════════════════
/**
* 'header' | 'sidebar'. Solo (detached single-session) windows are ALWAYS
* 'header': they show exactly one session, so a session list is noise — and
* #sessionTabs must never be parked inside the display:none <aside>, where
* updateTabOverflowMode() would measure 0/0 and the inline rename input would
* get zero geometry.
* 'header' | 'sidebar' | 'sidebar-rich'. Solo (detached single-session) windows
* are ALWAYS 'header': they show exactly one session, so a session list is
* noise — and #sessionTabs must never be parked inside the display:none
* <aside>, where updateTabOverflowMode() would measure 0/0 and the inline
* rename input would get zero geometry.
*
* The two sidebar values are the SAME layout — same docked column, same
* re-parented #sessionTabs, same filter box, same Alt+B toggle. They differ
* only in how much each row says, which is why the split rides on a separate
* attribute (see applySessionListLayout) instead of a third data-session-list
* value: every one of the ~25 isSessionSidebarActive() call sites, and every
* html[data-session-list="sidebar"] rule in styles.css and mobile.css, must
* keep matching both without being touched.
*/
getSessionListLayout() {
if (this.soloSessionId) return 'header';
const settings = this.loadAppSettingsFromStorage();
const defaults = this.getDefaultSettings();
const layout = settings.sessionListLayout ?? defaults.sessionListLayout ?? 'header';
return layout === 'sidebar' ? 'sidebar' : 'header';
return layout === 'sidebar' || layout === 'sidebar-rich' ? layout : 'header';
}
/**
@@ -3718,6 +3767,21 @@ class CodemanApp {
return document.documentElement.dataset.sessionList === 'sidebar';
}
/**
* True when the sidebar is showing the DETAILED rows: the home screen's
* per-session line ("created 3d ago · working 12m") plus a status pill.
*
* Read off <html> for the same reason as isSessionSidebarActive() — it is
* called once per tab in the render loop, and getSessionListLayout()
* re-parses localStorage on every call. Implies isSessionSidebarActive():
* data-sidebar-detail is only ever 'rich' while data-session-list is
* 'sidebar', both in applySessionListLayout() and in the pre-paint script.
*/
isSessionSidebarRich() {
const root = document.documentElement;
return root.dataset.sessionList === 'sidebar' && root.dataset.sidebarDetail === 'rich';
}
/**
* True where the sidebar is a MODAL off-canvas drawer over the terminal
* instead of a docked column.
@@ -3795,23 +3859,30 @@ class CodemanApp {
*/
applySessionListLayout() {
const mode = this.getSessionListLayout();
// 'sidebar' and 'sidebar-rich' are the same column; only row detail differs.
const sidebar = mode === 'sidebar' || mode === 'sidebar-rich';
const collapsed = this.isSessionSidebarCollapsed();
const prevMode = document.documentElement.dataset.sessionList;
const prevDetail = document.documentElement.dataset.sidebarDetail;
const tabsEl = document.getElementById('sessionTabs');
const headerHost = document.getElementById('sessionTabsHost');
const sidebarList = document.getElementById('sessionSidebarList');
if (!tabsEl || !headerHost || !sidebarList) return;
const host = mode === 'sidebar' ? sidebarList : headerHost;
const host = sidebar ? sidebarList : headerHost;
if (tabsEl.parentElement !== host) host.appendChild(tabsEl);
document.documentElement.dataset.sessionList = mode;
document.documentElement.dataset.sessionList = sidebar ? 'sidebar' : 'header';
// Detail is meaningless outside the sidebar, and must not linger as 'rich'
// there: the rows carry no meta line in the header strip, and a stale 'rich'
// would let the sidebar CSS style a strip that has nothing to style.
document.documentElement.dataset.sidebarDetail = mode === 'sidebar-rich' ? 'rich' : 'simple';
document.documentElement.dataset.sidebar = collapsed ? 'collapsed' : 'expanded';
tabsEl.setAttribute('aria-orientation', mode === 'sidebar' ? 'vertical' : 'horizontal');
tabsEl.setAttribute('aria-orientation', sidebar ? 'vertical' : 'horizontal');
const btn = document.getElementById('sidebarToggleBtn');
if (btn) {
btn.classList.toggle('btn-sidebar-toggle--hidden', mode !== 'sidebar');
btn.classList.toggle('btn-sidebar-toggle--hidden', !sidebar);
const label = collapsed ? 'Expand session sidebar' : 'Collapse session sidebar';
btn.setAttribute('aria-expanded', collapsed ? 'false' : 'true');
btn.setAttribute('aria-label', label);
@@ -3822,12 +3893,12 @@ class CodemanApp {
// "collapsed" means the drawer is closed.
const aside = document.getElementById('sessionSidebar');
if (aside) {
aside.classList.toggle('open', mode === 'sidebar' && !collapsed);
aside.classList.toggle('open', sidebar && !collapsed);
// A closed overlay drawer is only moved off screen by translateX(-100%);
// it keeps display:flex, so without this its filter box and ~4 tab stops
// per session stay in the Tab order and in the accessibility tree.
// NOT applied to the docked desktop rail — its rows are still clickable.
const hiddenDrawer = mode === 'sidebar' && collapsed && this._isSessionSidebarOverlay();
const hiddenDrawer = sidebar && collapsed && this._isSessionSidebarOverlay();
aside.toggleAttribute('inert', hiddenDrawer);
if (hiddenDrawer) aside.setAttribute('aria-hidden', 'true');
else aside.removeAttribute('aria-hidden');
@@ -3836,7 +3907,7 @@ class CodemanApp {
// The filter box only exists inside the sidebar; leaving a stale filter
// applied when the layout goes back to the header strip would hide sessions
// from the tab bar with no reachable control to clear it.
if (mode !== 'sidebar') {
if (!sidebar) {
this._sidebarFilter = '';
const filterInput = document.getElementById('sessionSidebarFilter');
if (filterInput) filterInput.value = '';
@@ -3852,13 +3923,21 @@ class CodemanApp {
// A layout flip alone still needs one render: the rows are rebuilt into the
// new host with the drag/keyboard handlers re-bound. Skipped when
// applyTabWrapSettings() already rendered for the folder-row change.
if (prevMode !== mode && prevTall === this._tallTabsEnabled) {
//
// The detail half of the test is not redundant: simple ⟷ rich leaves
// data-session-list on 'sidebar' both times, so comparing only that would
// flip the setting and repaint nothing until the next SSE tick — and the
// meta line is emitted by the row template, not toggled by CSS.
const layoutChanged =
prevMode !== document.documentElement.dataset.sessionList ||
prevDetail !== document.documentElement.dataset.sidebarDetail;
if (layoutChanged && prevTall === this._tallTabsEnabled) {
this._fullRenderSessionTabs();
}
// tabs-auto-wrap is measured, not derived from settings — updateTabOverflowMode()
// drops it in sidebar mode, but drop it here too so nothing paints wrapped
// for a frame before the next measure.
if (mode === 'sidebar') tabsEl.classList.remove('tabs-auto-wrap');
if (sidebar) tabsEl.classList.remove('tabs-auto-wrap');
// Collapse/expand changes whether the filter is reachable, so re-evaluate it
// here too — not only at the render tails.
this.applySidebarFilter(this._sidebarFilter);
@@ -3869,6 +3948,9 @@ class CodemanApp {
if (document.getElementById('welcomeOverlay')?.classList.contains('visible')) {
this.showHomeSessions?.();
}
// Only the rich rows carry stamps that go stale with no event behind them.
if (this.isSessionSidebarRich()) this._startSidebarRichClock();
else this._stopSidebarRichClock();
}
toggleSessionSidebar() {
@@ -3959,6 +4041,152 @@ class CodemanApp {
this.updateSidebarCount();
}
// ═══════════════════════════════════════════════════════════════
// Rich sidebar rows (sessionListLayout === 'sidebar-rich')
// ═══════════════════════════════════════════════════════════════
/**
* Pill copy per state, matching the desktop home rail and the phone overview
* word for word. Duplicated rather than imported for the same reason those two
* duplicate it from each other: it is six words, and constants.js is served
* from cache independently of app.js — a shared map there could arrive stale
* or missing while this file is new. What is NOT duplicated is the part that
* can actually disagree: which state a session is IN, and which stamp measures
* it, both of which come from mobile-overview.js below.
*/
_sidebarRichPillLabel(state) {
return {
needs: 'needs you',
error: 'error',
waiting: 'waiting',
working: 'working',
idle: 'idle',
done: 'done',
}[state] || state;
}
/**
* The per-row model for a rich sidebar row: which state the session is in,
* when it was first created, and how long it has been in that state.
*
* Classification is `_mobileOverviewState()` and the state duration is
* `_mobileOverviewSince()` (both mobile-overview.js), NOT re-derived here —
* the sidebar, the desktop home rail and the phone overview must never
* disagree about what "working" means or about which stamp measures it.
*
* Guarded like every other cross-file consumer in this app: a stale cached
* mobile-overview.js must degrade to a row with no meta line, not throw and
* take the whole tab strip down with it.
*/
_sidebarRichRow(id, session) {
if (typeof this._mobileOverviewState !== 'function') return null;
const state = this._mobileOverviewState(session, this.pendingHooks?.get(id));
return {
state,
pill: this._sidebarRichPillLabel(state),
createdAt: Number(session.createdAt) || 0,
since: this._mobileOverviewSince ? this._mobileOverviewSince(state, session) : null,
};
}
/**
* The "created 3d ago · working 12m" line plus the status pill, as the third
* child of `.tab-info` (already a flex column, so no row-level wrapping is
* needed — unlike the home rail, whose pill rides a wrapped full-width line).
*
* Both stamps keep their raw epoch-ms in `data-tab-ts` so
* `_tickSidebarRichTimes()` can rewrite the text without rebuilding the row:
* a rebuild would restart the load spinner and every alert animation in the
* list, twice a minute, for nothing.
*
* Returns '' when there is no model, which is what keeps the header strip and
* the simple sidebar byte-identical to before.
*/
_sidebarRichMetaHTML(row) {
if (!row) return '';
const stamp = (key, ts, fmt, cls) => {
const text = this._sidebarRichStampText(ts, fmt);
const title = ts
? ` title="${escapeHtml(`${key === 'created' ? 'First created' : key}: ${new Date(ts).toLocaleString()}`)}"`
: '';
return `<span class="tab-meta-item ${cls}"${title}><span class="tab-meta-key">${escapeHtml(key)}</span><span data-tab-ts="${ts || 0}" data-tab-fmt="${fmt}">${escapeHtml(text)}</span></span>`;
};
// data-i18n-skip: relative times are generated text, and "created"/"idle"
// are the same generic words that mean something else on other surfaces.
const parts = [stamp('created', row.createdAt, 'ago', 'tab-meta-created')];
if (row.since) {
parts.push('<span class="tab-meta-sep" aria-hidden="true">\u00B7</span>');
parts.push(stamp(row.since.key, row.since.at, 'for', 'tab-meta-since'));
}
parts.push(`<span class="tab-pill tab-pill--${escapeHtml(row.state)}">${escapeHtml(row.pill)}</span>`);
return `<span class="tab-meta" data-i18n-skip>${parts.join('')}</span>`;
}
/** Same formatter as both home screens, so a duration is written the same way everywhere. */
_sidebarRichStampText(timestamp, format) {
return this._mobileOverviewStampText ? this._mobileOverviewStampText(timestamp, format) : '\u2014';
}
/**
* Incremental-render counterpart of `_sidebarRichMetaHTML()`. The stamps move
* on the clock, but the STATE can change between renders (a session starts
* working, a permission prompt lands), and that flips the pill, the accent
* class and which stamp the second slot is even showing.
*
* Rebuilds the meta line only when something it displays actually changed,
* because this runs for every session on every SSE tick.
*/
_updateSidebarRichRow(tab, id, session) {
const row = this._sidebarRichRow(id, session);
if (!row) return;
const prev = tab.dataset.tabState;
// The since ANCHOR moves without the state changing (each new turn re-stamps
// lastSubmitAt), so key the compare on both.
const sig = `${row.state}:${row.since ? row.since.at : 0}:${row.createdAt}`;
if (tab.dataset.tabMetaSig === sig) return;
tab.dataset.tabMetaSig = sig;
tab.dataset.tabState = row.state;
if (prev) tab.classList.remove(`tab-state-${prev}`);
tab.classList.add(`tab-state-${row.state}`);
const info = tab.querySelector('.tab-info');
if (!info) return;
const html = this._sidebarRichMetaHTML(row);
const existing = info.querySelector('.tab-meta');
if (existing) existing.outerHTML = html;
else info.insertAdjacentHTML('beforeend', html);
}
/**
* Rewrites the relative stamps in place. A session that is just sitting there
* emits no event at all, so without this its "idle 2m" would still read 2m an
* hour later — the one number in the list that has to move on its own.
*/
_startSidebarRichClock() {
if (this._sidebarRichClock) return;
this._sidebarRichClock = setInterval(() => {
if (!this.isSessionSidebarRich()) {
this._stopSidebarRichClock();
return;
}
this._tickSidebarRichTimes();
}, SIDEBAR_RICH_CLOCK_MS);
}
_stopSidebarRichClock() {
if (!this._sidebarRichClock) return;
clearInterval(this._sidebarRichClock);
this._sidebarRichClock = null;
}
_tickSidebarRichTimes() {
const container = this.$('sessionTabs');
if (!container) return;
for (const node of container.querySelectorAll('[data-tab-ts]')) {
const text = this._sidebarRichStampText(Number(node.dataset.tabTs) || 0, node.dataset.tabFmt);
if (node.textContent !== text) node.textContent = text;
}
}
// ═══════════════════════════════════════════════════════════════
// Session Tabs
// ═══════════════════════════════════════════════════════════════
@@ -4149,6 +4377,9 @@ class CodemanApp {
webTabsUnchanged;
if (canIncremental) {
// Read once for the whole pass, like the full-rebuild path: this touches
// the DOM and the loop below runs for every session on every SSE tick.
const richRows = this.isSessionSidebarRich();
// Incremental update - only modify changed properties
for (const [id, session] of this.sessions) {
const tab = container.querySelector(`.session-tab[data-id="${id}"]`);
@@ -4221,6 +4452,21 @@ class CodemanApp {
statusEl.className = `tab-status ${status}`;
}
// Rich sidebar meta ("created 3d ago · working 12m" + pill). The stamps
// themselves move on _tickSidebarRichTimes(); this is here for the parts
// a tick cannot see — the state flipping, and with it the pill, the row
// accent and which stamp the second slot is measuring at all.
if (richRows) {
this._updateSidebarRichRow(tab, id, session);
} else if (tab.dataset.tabState) {
// Layout flipped away from rich without a full rebuild reaching this
// row yet: strip the line rather than leave a frozen stamp behind.
tab.querySelector('.tab-meta')?.remove();
tab.classList.remove(`tab-state-${tab.dataset.tabState}`);
delete tab.dataset.tabState;
delete tab.dataset.tabMetaSig;
}
// Update name if changed. #232: a description (the `: suffix` part of the
// name) is the whole tab label; the generated id lives in the tooltip. The
// compare targets the DISPLAY text, or a described tab would re-render on
@@ -4419,6 +4665,9 @@ class CodemanApp {
// into view replaces it.
const parts = [];
const tabOrder = this.sessionOrder;
// Read once, not per session: isSessionSidebarRich() touches the DOM and
// this loop runs for every tab on every full rebuild.
const richRows = this.isSessionSidebarRich();
let _tabIdx = 0;
for (const id of tabOrder) {
const session = this.sessions.get(id);
@@ -4460,7 +4709,17 @@ class CodemanApp {
? (session.workingDir ? `${parsedName.prefix} (${session.workingDir})` : parsedName.prefix)
: (session.workingDir || '');
parts.push(`<div class="session-tab ${isActive ? 'active' : ''}${alertClass}${loadState ? ' tab-loading' : ''}${this.hasTabDetachOverride(id) ? ' tab-show-detach' : ''}" data-id="${id}" data-color="${color}" ${loadState ? `data-load-phase="${escapeHtml(loadState.phase)}"` : ''} onclick="app.handleSessionTabClick(event, ${escapeHtml(JSON.stringify(id))})" oncontextmenu="event.preventDefault(); app.startInlineRename(${escapeHtml(JSON.stringify(id))})" tabindex="0" role="tab" aria-selected="${isActive ? 'true' : 'false'}" aria-busy="${loadState ? 'true' : 'false'}" aria-label="${escapeHtml(name)} session" ${tabTooltip ? `title="${escapeHtml(tabTooltip)}"` : ''}>
// Rich sidebar rows only: the home screen's created/state stamps and a
// status pill. richRow is null in every other layout, and both helpers
// below collapse to '' — the header strip's markup is unchanged.
const richRow = richRows ? this._sidebarRichRow(id, session) : null;
const richMeta = this._sidebarRichMetaHTML(richRow);
const richClass = richRow ? ` tab-state-${richRow.state}` : '';
const richData = richRow
? ` data-tab-state="${richRow.state}" data-tab-meta-sig="${richRow.state}:${richRow.since ? richRow.since.at : 0}:${richRow.createdAt}"`
: '';
parts.push(`<div class="session-tab ${isActive ? 'active' : ''}${alertClass}${richClass}${loadState ? ' tab-loading' : ''}${this.hasTabDetachOverride(id) ? ' tab-show-detach' : ''}"${richData} data-id="${id}" data-color="${color}" ${loadState ? `data-load-phase="${escapeHtml(loadState.phase)}"` : ''} onclick="app.handleSessionTabClick(event, ${escapeHtml(JSON.stringify(id))})" oncontextmenu="event.preventDefault(); app.startInlineRename(${escapeHtml(JSON.stringify(id))})" tabindex="0" role="tab" aria-selected="${isActive ? 'true' : 'false'}" aria-busy="${loadState ? 'true' : 'false'}" aria-label="${escapeHtml(name)} session" ${tabTooltip ? `title="${escapeHtml(tabTooltip)}"` : ''}>
${_tabIdx < 9 ? '<span class="tab-number">' + (_tabIdx + 1) + '</span>' : ''}
${loadState ? '<span class="tab-load-spinner" aria-hidden="true"></span>' : ''}
<span class="tab-status ${status}" aria-hidden="true"></span>
@@ -4471,6 +4730,7 @@ class CodemanApp {
<span class="tab-detached-badge" aria-hidden="true">detached</span>
</span>
${showFolder ? `<span class="tab-folder">\u{1F4C1} ${escapeHtml(folderName)}</span>` : ''}
${richMeta}
</span>
${hasRunningTasks ? `<span class="tab-badge" onclick="event.stopPropagation(); app.toggleTaskPanel()" aria-label="${taskStats.running} running tasks">${taskStats.running}</span>` : ''}
${subagentBadge}
@@ -4520,6 +4780,12 @@ class CodemanApp {
// innerHTML was rebuilt wholesale, so the sidebar filter classes are gone —
// re-apply them or filtered-out sessions flicker back on every SSE tick.
this.applySidebarFilter(this._sidebarFilter);
// Rows that carry self-staling stamps need the clock; rows that don't must
// not leave it running. Both directions matter — the layout can flip
// underneath a render, and a solo window forces 'header' regardless.
if (richRows) this._startSidebarRichClock();
else this._stopSidebarRichClock();
}
// Set up arrow key navigation for session tabs (accessibility)
@@ -5002,6 +5268,22 @@ class CodemanApp {
this.terminal.write('\x1b[3J\x1b[H\x1b[2J');
}
_recordTerminalLoadTiming(timing) {
this._lastTerminalLoadTiming = timing;
console.info('[TERMINAL-PERF]', timing);
const resetAndParseMs =
(timing.cacheResetAndParseMs || 0) +
(timing.freshResetAndParseMs || 0) +
(timing.resetAndParseMs || 0);
const totalMs = timing.selectDoneMs ?? timing.totalMs ?? timing.selectToReplayCompleteMs ?? 0;
_crashDiag.log(
`TERMINAL_LOAD: ${timing.trigger} ${timing.full ? 'full' : 'tail'} ${timing.chars} chars ` +
`ttfb=${timing.ttfbMs.toFixed(0)}ms body+json=${timing.bodyAndJsonMs.toFixed(0)}ms ` +
`reset+parse=${resetAndParseMs.toFixed(0)}ms total=${totalMs.toFixed(0)}ms ` +
`server="${timing.serverTiming}"${timing.refused ? ' refused-downgrade' : ''}`
);
}
/**
* "Load more history": re-pull the whole tmux scrollback when the user scrolls up
* while already at the top of what the browser has.
@@ -5030,6 +5312,11 @@ class CodemanApp {
const sessionId = this.activeSessionId;
if (!sessionId || this._fullHistoryRepullInFlight || this._isLoadingBuffer) return;
if (this.detachedSessions?.has(sessionId)) return;
const session = this.sessions.get(sessionId);
// A shell's full capture can be many megabytes. Replaying it from an
// ordinary scroll gesture blocks xterm's main thread, so keep that cost
// behind the explicit "Load full history" button.
if (!force && session?.mode === 'shell') return;
const now = Date.now();
// Momentum scrolling fires this dozens of times per flick, and a burst of new
// output is the normal reason to want a re-pull, so cooldown rather than latch.
@@ -5041,13 +5328,32 @@ class CodemanApp {
this._fullHistoryRepullAt.set(sessionId, now);
this._fullHistoryRepullInFlight = true;
try {
const requestStartedAt = performance.now();
const res = await fetch(`/api/sessions/${sessionId}/terminal?full=1`);
const headersReceivedAt = performance.now();
const payload = (await res.json())?.data ?? {};
const bodyParsedAt = performance.now();
const buffer = payload.terminalBuffer;
const timing = {
trigger: force ? 'full-history-button' : 'full-history-scroll',
mode: session?.mode || 'unknown',
full: true,
source: payload.source || 'unknown',
chars: buffer?.length || 0,
ttfbMs: headersReceivedAt - requestStartedAt,
bodyAndJsonMs: bodyParsedAt - headersReceivedAt,
resetAndParseMs: 0,
totalMs: 0,
serverTiming: res.headers?.get?.('server-timing') || '',
refused: false,
};
// Bail on a tab switch mid-fetch: writing here would paint another session's
// history into the terminal the user is now looking at.
if (!buffer || this.activeSessionId !== sessionId) return;
if (this._replayWouldShrinkBuffer(buffer)) {
timing.refused = true;
timing.totalMs = performance.now() - requestStartedAt;
this._recordTerminalLoadTiming(timing);
(this._fullHistoryRepullUseless ||= new Set()).add(sessionId);
this._logScrollRouting?.('repull-refused-downgrade');
// The browser already holds more than tmux can give back, so there is
@@ -5058,17 +5364,32 @@ class CodemanApp {
this._setHistoryTruncation(sessionId, payload);
this._fullHistoryRepullUseless?.delete(sessionId);
const rowsBefore = this.terminal.buffer.active.length;
const replayStartedAt = performance.now();
this._resetTerminalForReplay();
await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
if (this.activeSessionId !== sessionId) return;
this.terminalBufferCache.set(sessionId, buffer);
const {
parsedAt,
bufferLength: parsedBufferLength,
completed,
} = await this.chunkedTerminalWrite(buffer, TERMINAL_CHUNK_SIZE, sessionId);
timing.resetAndParseMs = parsedAt - replayStartedAt;
if (!completed || this.activeSessionId !== sessionId) return;
// Keep shell tab restores bounded too. A user-triggered full-history pull
// may be tens of MB; caching it would replay that whole payload again on
// the next tab switch before the normal 1MB tail fetch replaces it.
if (this.sessions.get(sessionId)?.mode !== 'shell') {
this.terminalBufferCache.set(sessionId, buffer);
} else {
this.terminalBufferCache.delete(sessionId);
}
// Hold the user's place. The replay is a superset that grew the buffer
// UPWARD, so what used to be row 0 (what they were looking at) is now `delta`
// rows down; scrolling there reveals the recovered history above it instead
// of teleporting them to the bottom the way a normal buffer load does.
const delta = this.terminal.buffer.active.length - rowsBefore;
const delta = parsedBufferLength - rowsBefore;
if (delta > 0) this.terminal.scrollToLine(delta);
else this.terminal.scrollToTop();
timing.totalMs = performance.now() - requestStartedAt;
this._recordTerminalLoadTiming(timing);
} catch {
// Transient (offline, 5xx) — the next scroll-up past the cooldown retries.
} finally {
@@ -5348,6 +5669,7 @@ class CodemanApp {
// COD-144: track whether the load painted nothing (empty fetch + no cache).
// For that just-created-session case we flush (not discard) queued SSE events.
let bufferWasEmpty = false;
let cacheResetAndParseMs = 0;
try {
// Fit terminal to container BEFORE writing any buffer data.
// If the browser was resized while viewing another session, the terminal
@@ -5425,23 +5747,30 @@ class CodemanApp {
// blank and rewrites with fresh data. Skip the cache and write the fresh
// buffer once for a single clean transition.
const cachedBuffer = this.terminalBufferCache.get(sessionId);
let clearedForBusy = false;
if (cachedBuffer && !sessionIsBusy && !restoredSnapshot) {
let clearedBeforeFresh = false;
if (cachedBuffer && !sessionIsBusy && !restoredSnapshot && session?.mode !== 'shell') {
_crashDiag.log(`CACHE_WRITE: ${(cachedBuffer.length/1024).toFixed(0)}KB`);
this._setTerminalLoadState(sessionId, selectGen, 'replaying');
const cacheReplayStartedAt = performance.now();
this._resetTerminalForReplay();
await this.chunkedTerminalWrite(cachedBuffer, TERMINAL_CHUNK_SIZE, bufferLoadOwner);
const { parsedAt: cacheParsedAt } = await this.chunkedTerminalWrite(
cachedBuffer,
TERMINAL_CHUNK_SIZE,
bufferLoadOwner
);
cacheResetAndParseMs = cacheParsedAt - cacheReplayStartedAt;
if (this._isStaleSelect(selectGen)) {
this._clearTerminalLoadState(sessionId, selectGen);
return;
}
this.terminal.scrollToBottom();
_crashDiag.log('CACHE_DONE');
} else if (sessionIsBusy) {
// Clear stale content immediately — fresh buffer is being fetched
} else if (sessionIsBusy || session?.mode === 'shell') {
// Busy sessions have stale caches. Shell sessions deliberately skip even
// an idle cache so a changed 1MB tail cannot cause two back-to-back parses.
this._resetTerminalForReplay();
clearedForBusy = true;
_crashDiag.log('CACHE_SKIP_BUSY');
clearedBeforeFresh = true;
_crashDiag.log(session?.mode === 'shell' ? 'CACHE_SKIP_SHELL' : 'CACHE_SKIP_BUSY');
}
// Give TUI sessions a short chance to redraw after resize before the
@@ -5459,26 +5788,29 @@ class CodemanApp {
this._setTerminalLoadState(sessionId, selectGen, 'fetching');
_crashDiag.log('FETCH_START');
// The first load OF EACH SESSION this page load requests the full tmux
// scrollback (?full=1, COD-47) so history that scrolled off the server's byte
// buffer comes back. Later switches to an already-replayed session keep the
// fast ?tail= frame path, which is why this is a Set and not a flag: the flag
// version gave the full replay to the auto-selected tab and one frame of
// history to every other one (issue #205).
const useFullHistory = !this._fullHistoryLoaded.has(sessionId);
// TUI sessions still get one canonical full replay per page (COD-47/#205).
// A shell can retain hundreds of thousands of plain scrollback lines, so
// automatically replaying all of them makes tab selection scale with the
// entire session. Load its bounded 1MB tail first; the existing truncation
// banner action fetches ?full=1 when the user explicitly asks for it.
const useFullHistory = session?.mode !== 'shell' && !this._fullHistoryLoaded.has(sessionId);
if (useFullHistory) this._fullHistoryLoaded.add(sessionId);
const fetchStartedAt = performance.now();
const res = await fetch(
useFullHistory
? `/api/sessions/${sessionId}/terminal?full=1`
: `/api/sessions/${sessionId}/terminal?tail=${TERMINAL_TAIL_SIZE}`
);
const headersReceivedAt = performance.now();
if (this._isStaleSelect(selectGen)) {
this._clearTerminalLoadState(sessionId, selectGen);
return;
}
const data = (await res.json())?.data ?? {};
const bodyParsedAt = performance.now();
_crashDiag.log(`FETCH_DONE: ${data.terminalBuffer ? (data.terminalBuffer.length/1024).toFixed(0) + 'KB' : 'empty'} truncated=${data.truncated}`);
let freshResetAndParseMs = 0;
if (data.terminalBuffer) {
// Skip rewrite if fresh buffer matches cache — avoids visible clear+rewrite flash.
// On slow connections (mobile 5G), the gap between clear() and chunkedWrite() is
@@ -5487,10 +5819,11 @@ class CodemanApp {
// something other than the cache, so the fetched buffer must be
// replayed even when it byte-matches the cache.
const needsRewrite =
restoredSnapshot || clearedForBusy || data.terminalBuffer !== cachedBuffer;
restoredSnapshot || clearedBeforeFresh || data.terminalBuffer !== cachedBuffer;
if (needsRewrite) {
_crashDiag.log(`REWRITE: ${(data.terminalBuffer.length/1024).toFixed(0)}KB`);
this._setTerminalLoadState(sessionId, selectGen, 'replaying');
const replayStartedAt = performance.now();
this._resetTerminalForReplay();
// Truncation is reported OUT OF BAND (#258). This used to write a grey
// "... earlier output truncated ..." line into the
@@ -5498,7 +5831,12 @@ class CodemanApp {
// cannot be actioned, and is indistinguishable from real CLI output.
this._setHistoryTruncation(sessionId, data);
// Use chunked write for large buffers to avoid UI jank
await this.chunkedTerminalWrite(data.terminalBuffer, TERMINAL_CHUNK_SIZE, bufferLoadOwner);
const { parsedAt: freshParsedAt } = await this.chunkedTerminalWrite(
data.terminalBuffer,
TERMINAL_CHUNK_SIZE,
bufferLoadOwner
);
freshResetAndParseMs = freshParsedAt - replayStartedAt;
if (this._isStaleSelect(selectGen)) {
this._clearTerminalLoadState(sessionId, selectGen);
return;
@@ -5507,22 +5845,42 @@ class CodemanApp {
this.terminal.scrollToBottom();
}
// Update cache (cap at 20 entries)
this.terminalBufferCache.set(sessionId, data.terminalBuffer);
if (this.terminalBufferCache.size > 20) {
// Evict oldest entry (first key in Map iteration order)
const oldest = this.terminalBufferCache.keys().next().value;
this.terminalBufferCache.delete(oldest);
// Shell selection always uses a fresh bounded tail, so retaining its
// payload only wastes memory and can evict useful TUI caches.
if (session?.mode === 'shell') {
this.terminalBufferCache.delete(sessionId);
} else {
// Update cache (cap at 20 entries)
this.terminalBufferCache.set(sessionId, data.terminalBuffer);
if (this.terminalBufferCache.size > 20) {
// Evict oldest entry (first key in Map iteration order)
const oldest = this.terminalBufferCache.keys().next().value;
this.terminalBufferCache.delete(oldest);
}
}
} else if (!cachedBuffer) {
// No fresh buffer and no cache — clear any stale content
this._resetTerminalForReplay();
} else if (!cachedBuffer || clearedBeforeFresh) {
// Nothing was painted. If this path was not already cleared above,
// clear stale content now; either way queued live output must be flushed.
if (!clearedBeforeFresh) this._resetTerminalForReplay();
bufferWasEmpty = true;
}
const terminalLoadTiming = {
trigger: 'session-select',
mode: session?.mode || 'unknown',
full: useFullHistory,
source: data.source || 'unknown',
chars: data.terminalBuffer?.length || 0,
ttfbMs: headersReceivedAt - fetchStartedAt,
bodyAndJsonMs: bodyParsedAt - headersReceivedAt,
cacheResetAndParseMs,
freshResetAndParseMs,
selectToReplayCompleteMs: performance.now() - _selStart,
serverTiming: res.headers?.get?.('server-timing') || '',
};
// Buffer load complete — unblock live SSE writes. chunkedTerminalWrite calls
// _finishBufferLoad internally (discarding queued events to prevent duplicate
// content); if we skipped the write (cache hit or empty), call it here.
// _finishBufferLoad after ordering the fetched snapshot in xterm; if we skipped
// the write (cache hit or empty), call it here.
// COD-144: when the load painted nothing, FLUSH the queued events instead of
// discarding — a new session's prompt arrives only as a queued SSE event.
if (this._isLoadingBuffer) {
@@ -5651,9 +6009,12 @@ class CodemanApp {
if (typeof KeyboardHandler !== 'undefined' && KeyboardHandler.keyboardVisible) {
KeyboardHandler.onKeyboardShow();
}
const selectDoneMs = performance.now() - _selStart;
terminalLoadTiming.selectDoneMs = selectDoneMs;
this._recordTerminalLoadTiming(terminalLoadTiming);
this._clearTerminalLoadState(sessionId, selectGen);
_crashDiag.log(`SELECT_DONE: ${(performance.now() - _selStart).toFixed(0)}ms`);
console.log(`[CRASH-DIAG] selectSession DONE: ${sessionId.slice(0,8)} in ${(performance.now() - _selStart).toFixed(0)}ms`);
_crashDiag.log(`SELECT_DONE: ${selectDoneMs.toFixed(0)}ms`);
console.log(`[CRASH-DIAG] selectSession DONE: ${sessionId.slice(0,8)} in ${selectDoneMs.toFixed(0)}ms`);
} catch (err) {
if (this._isLoadingBuffer) this._finishBufferLoad(bufferLoadOwner);
this._restoringFlushedState = false;
@@ -5721,18 +6082,31 @@ class CodemanApp {
}
async closeSession(sessionId, killMux = true) {
// ⚠️ Captured BEFORE the await, and the delete is announced to
// _onSessionDeleted through _closingSessions. The `session_deleted` SSE
// broadcast for THIS delete routinely lands while the request is still in
// flight, and that handler nulls activeSessionId and shows the welcome
// screen. Re-reading the field after the await therefore made the fallback
// below a coin flip: closing the tab you were on either moved you to the
// next session or dumped you on the home screen, depending on which path
// won the race (both outcomes measured on one build, 2026-08-17).
const wasActive = this.activeSessionId === sessionId;
this._closingSessions.add(sessionId);
try {
await this._apiDelete(`/api/sessions/${sessionId}?killMux=${killMux}`);
this._cleanupSessionData(sessionId);
if (this.activeSessionId === sessionId) {
if (wasActive) {
this.activeSessionId = null;
try { localStorage.removeItem('codeman-active-session'); } catch {}
// Select another session or show welcome (use sessionOrder for consistent ordering)
if (this.sessionOrder.length > 0 && this.sessions.size > 0) {
// Next tab in the user's own order, skipping ids the cleanup has not
// caught up with yet: sessionOrder can transiently hold a dead id
// (delete racing the order sync), which is the same reason Alt+N
// indexes a live-filtered list rather than sessionOrder directly.
const nextSessionId = this.sessionOrder.find((id) => id !== sessionId && this.sessions.has(id));
if (nextSessionId) {
// `auto`: this tab was chosen by the app because the previous one
// went away, so it must not spend that session's idle alert.
const nextSessionId = this.sessionOrder[0];
this.selectSession(nextSessionId, { auto: true });
} else {
this.terminal.clear();
@@ -5750,6 +6124,8 @@ class CodemanApp {
}
} catch (err) {
this.showToast('Failed to close session', 'error');
} finally {
this._closingSessions.delete(sessionId);
}
}
+200 -1
View File
@@ -243,7 +243,9 @@ const LINEAGE_DIP_MAX_PX = 64;
// apart bled into one thick band instead of reading as three separate lines.
const LINEAGE_SIBLING_STEP_PX = 8;
const LINEAGE_STRIP_TOLERANCE_PX = 4;
// Lineage palette, assigned per CHILD in first-seen order and cycled (session-lineage.js).
// Lineage palette, assigned per SPAWNING TAB in first-seen order and cycled
// (session-lineage.js). Every arc leaving one tab shares its colour however many
// workers it spawns; a child that spawns in turn gets its own for the arcs below it.
// The empty FIRST entry means "no override": the CSS then falls back to --session-blue,
// which every skin block tunes for its own background, so a lone arc keeps the
// skin-aware blue that shipped in 1.18.2. The fixed entries are deliberately vivid
@@ -525,6 +527,90 @@ function sortSessionsByActivity(rows) {
return (Array.isArray(rows) ? rows.slice() : []).sort(compareSessionActivity);
}
// Terminal font stack — the single source for every xterm surface (the main
// terminal in terminal-ui.js, the log-viewer terminal in panels-ui.js).
// "Symbols Nerd Font Mono" is a bundled icons-only webfont (fonts/ +
// @font-face in styles.css): browsers fall back PER GLYPH, so Nerd Font
// prompt icons (powerline segments, folder/git glyphs from p10k, starship,
// oh-my-posh) render even though the text fonts carry no private-use-area
// symbols — while all readable text keeps coming from the text fonts.
const TERMINAL_FONT_DEFAULT_STACK =
'"Fira Code", "Cascadia Code", "JetBrains Mono", "SF Mono", Monaco, "Symbols Nerd Font Mono", monospace';
/**
* Resolve the xterm fontFamily from the per-device `terminalFontFamily`
* setting. A user-set family (or comma-separated list) is PREPENDED to the
* built-in stack, never a replacement — the symbols fallback and a final
* `monospace` must survive whatever the user types. Blank input yields the
* default. Unquoted names that need quoting for CSS (spaces, digits leading,
* etc.) are quoted; embedded quotes are stripped rather than escaped, since
* a font name cannot contain them anyway.
*/
function resolveTerminalFontFamily(custom) {
const raw = typeof custom === 'string' ? custom.trim() : '';
if (!raw) return TERMINAL_FONT_DEFAULT_STACK;
const families = raw
.split(',')
.map((f) => f.trim().replace(/^["']|["']$/g, '').replace(/["']/g, '').trim())
.filter(Boolean)
// Drop generic families the user may append — the default stack already
// ends in `monospace`, and a duplicate earlier entry would shadow the
// symbols fallback behind it.
.filter((f) => !/^(monospace|serif|sans-serif|system-ui)$/i.test(f))
.map((f) => (/^[A-Za-z][A-Za-z0-9-]*$/.test(f) ? f : `"${f}"`));
if (!families.length) return TERMINAL_FONT_DEFAULT_STACK;
return `${families.join(', ')}, ${TERMINAL_FONT_DEFAULT_STACK}`;
}
// ---------------------------------------------------------------------------
// Auto Copy (copy-on-select). Pure decision, so every guard below is testable
// without a terminal, a clipboard, or a browser.
// ---------------------------------------------------------------------------
/**
* Upper bound on an AUTO-copied selection.
*
* A drag that runs off the top of the viewport autoscrolls, so one gesture can
* sweep the entire 50k-line scrollback (millions of characters), and writing
* that to the clipboard on every mouseup is a real hazard on a phone. Past the
* cap the copy is REFUSED rather than truncated (half a selection on the
* clipboard is worse than none) and the user is told to press Ctrl+C, which
* still copies the whole thing through the explicit path.
*/
const AUTO_COPY_MAX_CHARS = 1_000_000;
/**
* What an auto-copy attempt should do at the end of a selection gesture.
*
* `pending` is set by xterm's onSelectionChange and cleared on every flush;
* `lastCopied` is the text this surface auto-copied last. Either one alone is
* wrong, which is why both are here:
*
* - onSelectionChange does not reliably fire BEFORE the mouseup that ends the
* drag (xterm fires it from its own document-level mouseup handler, and
* listener order between the two is registration order, not something this
* code controls). Gating on `pending` alone would silently drop the first
* copy of a drag-selection.
* - Gating on `text !== lastCopied` alone drops a deliberate re-selection of
* the same text after the user copied something else in between, and it
* would let any unrelated mouseup on the page re-copy a stale selection.
*
* So: a genuine selection change (`pending`) always copies, and otherwise only
* text that differs from the last auto-copy does.
*
* @param {{enabled?: boolean, text?: string, lastCopied?: string, pending?: boolean}} params
* @returns {'copy'|'skip'|'too-large'}
*/
function decideAutoCopy({ enabled, text, lastCopied, pending } = {}) {
if (!enabled) return 'skip';
// Whitespace-only is what a drag across blank cells produces; putting a wall
// of spaces on the clipboard is never what the gesture meant.
if (typeof text !== 'string' || !text.trim()) return 'skip';
if (!pending && text === lastCopied) return 'skip';
if (text.length > AUTO_COPY_MAX_CHARS) return 'too-large';
return 'copy';
}
if (typeof window !== 'undefined') {
window.WEBGL_FALLBACK = WEBGL_FALLBACK;
window.evaluateWebGLLongTaskTrip = evaluateWebGLLongTaskTrip;
@@ -558,6 +644,14 @@ if (typeof window !== 'undefined') {
compare: compareSessionActivity,
sort: sortSessionsByActivity,
};
window.CodemanAutoCopy = {
decide: decideAutoCopy,
MAX_CHARS: AUTO_COPY_MAX_CHARS,
};
window.CodemanTerminalFont = {
DEFAULT_STACK: TERMINAL_FONT_DEFAULT_STACK,
resolve: resolveTerminalFontFamily,
};
}
// Scheduler API — prioritize terminal writes over background UI updates.
@@ -1033,7 +1127,112 @@ function previewsInFileViewer(filePath) {
return FILE_PREVIEW_EXTENSIONS.has(ext);
}
/**
* The LOGICAL line a terminal row belongs to — the rows it spans, its text as one
* string, and a two-way map between that string and terminal cells.
*
* One definition, two consumers: the link provider matches its patterns over this
* text (`registerFilePathLinkProvider`) and touch selection measures words and
* whole lines with it (`_touchSelectionLogicalLine`). They MUST agree — a link that
* spans a wrap and a "Line" that stops at the screen edge is the same bug twice.
*
* Two kinds of continuation, and handling only the first is not enough:
*
* 1. **Soft wrap** — the emulator ran out of columns and flags the next row
* `isWrapped`. It inserts nothing, so the row's text is joined verbatim.
* 2. **Hard wrap** — the program wrapped the text itself and emitted a real
* newline, so nothing is flagged. A row that fills the last column is taken
* as continuing into the next; that is the only trace a hard wrap leaves.
*
* ⚠️ A hard-wrapped continuation may carry the program's own INDENT, and joining
* that verbatim puts whitespace in the middle of the token being stitched. That is
* why an agent's numbered list —
*
* 1. https://github.com/users/someone/packages/container/p
* ackage/thing
*
* — opened only `…/container/p`: the URL pattern stops at the space the indent
* contributed. So the leading whitespace of a HARD continuation is dropped, and
* `colStart` on that segment records how much, keeping the cell mapping exact. A
* soft continuation keeps its leading whitespace, since the terminal never adds
* any and it is therefore real content.
*
* ⚠️ Only the final row is trimmed. Continuation rows are read UNTRIMMED so each
* contributes exactly `cols` cells; trimming one would shift every later offset.
*
* The row span is bounded by `maxRows` (12 by default): this runs on every hover,
* and a screenful of full-width output would otherwise re-scan the viewport each
* time.
*
* @param {{getLine: (row: number) => any, length: number}} buffer xterm buffer.
* @param {number} row 0-based ABSOLUTE buffer row to expand around.
* @param {number} cols Terminal width.
* @param {number} [maxRows] Row-span bound.
* @returns {{startRow: number, endRow: number, text: string,
* offsetToCell: (offset: number) => {row: number, col: number},
* cellToOffset: (row: number, col: number) => number} | null}
* 0-based rows and columns throughout; null when the row does not exist.
*/
function terminalLogicalLine(buffer, row, cols, maxRows) {
if (!buffer || typeof buffer.getLine !== 'function') return null;
const width = Math.max(1, cols || 1);
const bound = Math.max(1, maxRows || 12);
const lineAt = (r) => (r >= 0 ? buffer.getLine(r) : undefined);
if (!lineAt(row)) return null;
const continuesPrevious = (r) => {
if (r <= 0) return false;
if (lineAt(r)?.isWrapped) return true;
const prev = lineAt(r - 1);
return !!prev && (prev.translateToString(true) || '').length >= width;
};
let startRow = row;
while (startRow > 0 && row - startRow < bound && continuesPrevious(startRow)) startRow--;
let endRow = row;
const length = Number.isFinite(buffer.length) ? buffer.length : endRow + 1;
while (endRow + 1 < length && endRow - startRow < bound && continuesPrevious(endRow + 1)) endRow++;
const segments = [];
let text = '';
for (let r = startRow; r <= endRow; r++) {
const line = lineAt(r);
if (!line) break;
let rowText = line.translateToString(r === endRow) || '';
let colStart = 0;
if (r > startRow && !line.isWrapped) {
const indent = rowText.length - rowText.replace(/^\s+/, '').length;
colStart = indent;
rowText = rowText.slice(indent);
}
segments.push({ row: r, textStart: text.length, colStart, length: rowText.length });
text += rowText;
}
const offsetToCell = (offset) => {
for (let i = segments.length - 1; i >= 0; i--) {
const seg = segments[i];
if (offset >= seg.textStart || i === 0) {
return { row: seg.row, col: seg.colStart + (offset - seg.textStart) };
}
}
return { row: startRow, col: offset };
};
const cellToOffset = (targetRow, targetCol) => {
for (const seg of segments) {
if (seg.row !== targetRow) continue;
return seg.textStart + Math.max(0, targetCol - seg.colStart);
}
return -1;
};
return { startRow, endRow, text, offsetToCell, cellToOffset };
}
if (typeof window !== 'undefined') {
window.CodemanHistoryFormat = { formatHistoryBytes, computeHistoryTruncationNotice, computeRewriteScrollLine };
window.CodemanFilePaths = { absoluteFilePathPattern, previewsInFileViewer, FILE_PREVIEW_EXTENSIONS };
window.CodemanTerminalLines = { terminalLogicalLine };
}
@@ -0,0 +1,21 @@
The MIT License (MIT)
Copyright (c) 2014 Ryan L McIntyre
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
Binary file not shown.
+17 -2
View File
@@ -234,8 +234,9 @@
'Session List Layout': '会话列表布局',
'Header tab strip': '顶栏标签条',
'Left sidebar': '左侧边栏',
'Horizontal strip in the header, or a collapsible left sidebar (Alt+B).':
'会话列表显示为顶栏横向标签条,或左侧可折叠侧边栏(Alt+B)。',
'Left sidebar simple': '左侧边栏(简洁)',
'Horizontal strip in the header, or a collapsible left sidebar (Alt+B). The rich sidebar carries the same per-session detail as the home screen.':
'会话列表显示为顶栏横向标签条,或左侧可折叠侧边栏(Alt+B)。完整侧边栏为每个会话显示与主界面相同的详细信息。',
'Tall Tabs (Name + Folder)': '双行标签(名称 + 文件夹)',
'Pop-out Button on Tabs': '标签页弹出窗口按钮',
Panels: '面板',
@@ -325,11 +326,20 @@
// Input settings
Input: '输入',
Font: '字体',
'Terminal font': '终端字体',
'Prepended to the built-in stack, so fallbacks (including bundled Nerd Font symbols) keep working. Must be installed on this device. Leave empty for the default.':
'置于内置字体栈之前,回退字体(包括内置的 Nerd Font 图标)仍然生效。需已安装在本设备上。留空使用默认值。',
'Local Echo': '本地回显',
'CJK Input': '中日韩输入',
'Extended Keyboard Bar': '扩展键盘栏',
'Gesture Control (beta)': '手势控制(测试版)',
'Wheel Scrolls Local History': '滚轮滚动本地历史',
'Auto Copy Selection': '自动复制选中内容',
'Selection & clipboard': '选中与剪贴板',
'Auto Copy: selection copied': '自动复制:已复制选中内容',
'Auto Copy failed: the browser blocked clipboard access': '自动复制失败:浏览器阻止了剪贴板访问',
'Selection too large to copy automatically. Press Ctrl+C.': '选中内容过大,无法自动复制。请按 Ctrl+C。',
'Instant typing feedback with local echo': '通过本地回显即时显示输入',
'Dedicated IME input field for CJK languages': '为中日韩语言提供专用输入法文本框',
'Extra keys: Tab, Esc, arrows, Ctrl+O': '附加按键:Tab、Esc、方向键、Ctrl+O',
@@ -499,6 +509,11 @@
'Respawn Blocked': '重生已阻止',
'Task Complete': '任务完成',
'Copied to clipboard': '已复制到剪贴板',
// Terminal touch-selection bar (long-press to select). The bar is a sibling of
// `.xterm`, not a descendant, so SKIP_SELECTOR does not cover it and these apply.
Copy: '复制',
Line: '整行',
'Clear selection': '清除选择',
'Failed to copy': '复制失败',
'Checking…': '正在检查…',
'Starting…': '正在启动…',
+33 -3
View File
@@ -24,6 +24,9 @@
<!-- Preload critical resources — lets browser discover these during HTML parse
instead of waiting until <script> tags at bottom-of-body are reached. -->
<link rel="preload" href="vendor/xterm.min.js" as="script">
<!-- Symbols font before the terminal first paints — a late-loading icon font
leaves tofu in xterm's glyph atlas until something forces a re-render. -->
<link rel="preload" href="fonts/symbols-nerd-font-mono.woff2" as="font" type="font/woff2" crossorigin>
<link rel="preload" href="constants.js" as="script">
<link rel="preload" href="app.js" as="script">
<!-- Self-hosted xterm.js — eliminates CDN DNS/TLS latency (~100ms).
@@ -62,7 +65,7 @@
app.js, NOT the handheld storage-key test `m`. Use a different predicate
here and boot will contradict this value, animating the drawer open by
itself on every load between 768 and 1023px. -->
<script>try{var m=window.innerWidth<768||(('ontouchstart' in window||navigator.maxTouchPoints>0)&&window.innerWidth<1024);var k=m?'codeman-app-settings-mobile':'codeman-app-settings';var L=JSON.parse(localStorage.getItem(k)||'{}').sessionListLayout;var solo=/^\/session\//.test(location.pathname);var C=localStorage.getItem('codeman-sidebar-collapsed');document.documentElement.dataset.sessionList=(L==='sidebar'&&!solo)?'sidebar':'header';document.documentElement.dataset.sidebar=(C===null?window.innerWidth<1024:C==='1')?'collapsed':'expanded';}catch(e){document.documentElement.dataset.sessionList='header';document.documentElement.dataset.sidebar='expanded';}</script>
<script>try{var m=window.innerWidth<768||(('ontouchstart' in window||navigator.maxTouchPoints>0)&&window.innerWidth<1024);var k=m?'codeman-app-settings-mobile':'codeman-app-settings';var L=JSON.parse(localStorage.getItem(k)||'{}').sessionListLayout;var solo=/^\/session\//.test(location.pathname);var C=localStorage.getItem('codeman-sidebar-collapsed');var S=(L==='sidebar'||L==='sidebar-rich')&&!solo;document.documentElement.dataset.sessionList=S?'sidebar':'header';document.documentElement.dataset.sidebarDetail=(S&&L==='sidebar-rich')?'rich':'simple';document.documentElement.dataset.sidebar=(C===null?window.innerWidth<1024:C==='1')?'collapsed':'expanded';}catch(e){document.documentElement.dataset.sessionList='header';document.documentElement.dataset.sidebarDetail='simple';document.documentElement.dataset.sidebar='expanded';}</script>
<!-- Inline critical CSS for instant skeleton paint (before styles.css loads) -->
<style>
.loading-skeleton{display:flex;flex-direction:column;height:100vh;height:100dvh;background:var(--bg-dark,#11151c)}
@@ -1624,6 +1627,32 @@
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Selection &amp; clipboard</h4><span class="set-scope">device</span></div>
<div class="set-group-body">
<div class="set-row" data-search="auto copy selection clipboard highlight copy on select mouse">
<div class="set-row-text">
<span class="set-row-label">Auto Copy Selection</span>
<span class="set-row-desc">Put highlighted terminal text on the clipboard as soon as you finish selecting it, with mouse, double-click or long-press. Ctrl+C still copies on demand, and nothing outside the terminal is copied.</span>
</div>
<label class="switch switch-sm"><input type="checkbox" id="appSettingsAutoCopySelection"><span class="slider"></span></label>
</div>
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Font</h4><span class="set-scope">device</span></div>
<div class="set-group-body">
<div class="set-row has-field" data-search="terminal font family nerd custom typeface">
<div class="set-row-text">
<span class="set-row-label">Terminal font</span>
<span class="set-row-desc">Prepended to the built-in stack, so fallbacks (including bundled Nerd Font symbols) keep working. Must be installed on this device. Leave empty for the default.</span>
</div>
<input type="text" id="appSettingsTerminalFont" class="set-input" placeholder='e.g. JetBrainsMono Nerd Font'>
</div>
</div>
</div>
<div class="set-group">
<div class="set-group-head"><h4>Scrolling &amp; rendering</h4></div>
<div class="set-group-body">
@@ -1843,11 +1872,12 @@
<div class="set-row has-field" data-search="session list layout sidebar tab strip vertical">
<div class="set-row-text">
<span class="set-row-label">Session List Layout</span>
<span class="set-row-desc">Horizontal strip in the header, or a collapsible left sidebar (Alt+B).</span>
<span class="set-row-desc">Horizontal strip in the header, or a collapsible left sidebar (Alt+B). The rich sidebar carries the same per-session detail as the home screen.</span>
</div>
<select id="appSettingsSessionListLayout" class="set-select">
<option value="header">Header tab strip</option>
<option value="sidebar">Left sidebar</option>
<option value="sidebar">Left sidebar simple</option>
<option value="sidebar-rich">Left sidebar</option>
</select>
</div>
<div class="set-row" data-search="tall tabs folder name two rows">
+27 -25
View File
@@ -27,8 +27,8 @@
*
* Solution: outside composition, flush is DEBOUNCED (200ms). The entire
* delete→reinsert cycle collapses into one flush of the final textarea value.
* Keyboard typing of single printable characters still goes through the
* keydown handler (immediate, no debounce).
* Physical-keyboard commits are flushed immediately after the input event
* exposes the final browser/IME text; keydown never guesses that text.
*
* ## Phantom character for Android backspace
*
@@ -56,8 +56,7 @@ const CjkInput = (() => {
let _compositionFlushTimer = null;
let _dictationActive = false;
let _dictationDecayTimer = null;
let _keydownSentAt = 0;
let _keydownSentText = '';
let _printableKeydownAt = null;
const _listeners = {};
const PHANTOM = '​';
@@ -197,6 +196,7 @@ const CjkInput = (() => {
_send = send;
_composing = false;
_printableKeydownAt = null;
_flushTimer = null;
_textarea = document.getElementById('cjkInput');
if (!_textarea) return this;
@@ -234,6 +234,7 @@ const CjkInput = (() => {
};
_listeners.blur = () => {
_t(`blur composing=${_composing} ${_vdesc(_textarea.value)}`);
_printableKeydownAt = null;
// Keep cjkActive while CJK input is visible — iOS dictation and system
// UI may steal focus temporarily, and clearing the flag during that
// window lets xterm's onData process duplicated input.
@@ -253,6 +254,7 @@ const CjkInput = (() => {
_listeners.compositionstart = () => {
_t(`compstart ${_vdesc(_textarea.value)}`);
_composing = true;
_printableKeydownAt = null;
_cancelDebouncedFlush();
// Leave textarea.value untouched — programmatic changes during
// compositionstart cancel the IME composition on iOS Safari.
@@ -277,6 +279,7 @@ const CjkInput = (() => {
// ── Keydown: special keys work REGARDLESS of composition state ──
_listeners.keydown = (e) => {
_t(`keydown ${_kdesc(e.key)} kc=${e.keyCode} ic=${e.isComposing} c=${_composing}`);
_printableKeydownAt = null;
if (e.key === 'Enter') {
e.preventDefault();
_composing = false;
@@ -325,16 +328,11 @@ const CjkInput = (() => {
return;
}
// Single printable character: send immediately to PTY.
// Third-party IMEs on iOS may ignore preventDefault, so the char
// still enters the textarea and fires an input event — _keydownSentAt
// tells the input handler to skip that echo.
// A printable KeyboardEvent.key is the physical key, not necessarily
// the committed text. Let the browser/IME produce the input event so
// full-width punctuation and other layout transforms are preserved.
if (e.key.length === 1 && !e.ctrlKey && !e.altKey && !e.metaKey && _isEffectivelyEmpty()) {
e.preventDefault();
_send(e.key);
_keydownSentAt = performance.now();
_keydownSentText = e.key;
_resetToPhantom();
_printableKeydownAt = performance.now();
return;
}
};
@@ -343,6 +341,8 @@ const CjkInput = (() => {
// ── Input event: primary path for virtual keyboards + dictation ──
_listeners.input = (e) => {
_t(`input ${e.inputType || '?'} ic=${e.isComposing} c=${_composing} ${_vdesc(_textarea.value)}`);
const printableKeydownAt = _printableKeydownAt;
_printableKeydownAt = null;
// ── Stuck-composition recovery ──
// Some IMEs (WeChat/Sogou keyboards) fire compositionstart without a
// matching compositionend. A stale _composing=true blocks every flush
@@ -388,18 +388,18 @@ const CjkInput = (() => {
if (_composing) return;
// Keydown handler already sent this character — clear the textarea
// echo that the IME inserted despite preventDefault. Content-checked:
// only a value matching the sent char is an echo. Anything else (e.g.
// an IME committing CJK text right after a keydown-sent char) is real
// input and must flow through to the debounced flush, not be dropped.
if (performance.now() - _keydownSentAt < 100) {
const cur = _strip(_textarea.value);
if (cur === '' || cur === _keydownSentText) {
_t('echo-drop');
_resetToPhantom();
return;
}
// A recent physical printable key makes this insertText a keyboard
// commit, so keep the old zero-latency path. Send the textarea's final
// Unicode value, never KeyboardEvent.key, because the IME may have
// transformed punctuation or the active layout may differ.
if (
e.inputType === 'insertText' &&
printableKeydownAt !== null &&
performance.now() - printableKeydownAt < 100
) {
_cancelDebouncedFlush();
_flush();
return;
}
// Outside composition: keyboard typing or voice dictation.
@@ -425,6 +425,7 @@ const CjkInput = (() => {
clearTimeout(_compositionFlushTimer);
_compositionFlushTimer = null;
_composing = false;
_printableKeydownAt = null;
_resetToPhantom();
},
@@ -446,6 +447,7 @@ const CjkInput = (() => {
}
window.cjkActive = false;
_composing = false;
_printableKeydownAt = null;
for (const key of Object.keys(_listeners)) delete _listeners[key];
_initialized = false;
},
+38 -1
View File
@@ -573,6 +573,42 @@ const KeyboardHandler = {
* space below the last row. After fitAddon.fit(), measure the gap and
* reduce padding by that amount so the terminal sits flush against the bars.
*/
/**
* Combined height of the fixed bars that overlay the terminal's bottom edge.
*
* On phones the toolbar and the accessory bar are `position: fixed`, so they
* occupy no layout space of their own — `main`'s padding-bottom is the only
* thing reserving room for them, and any pixel taken out of it is a pixel of
* terminal painted underneath them.
*/
_fixedBottomBarsHeight() {
let px = 0;
for (const selector of ['.toolbar', '.keyboard-accessory-bar', '#cjkInput.cjk-input-visible']) {
const el = document.querySelector(selector);
if (!el) continue;
const style = window.getComputedStyle?.(el);
if (style && (style.display === 'none' || style.visibility === 'hidden')) continue;
px += el.offsetHeight || 0;
}
return px;
},
/**
* Reclaim sub-row slack at the bottom of the terminal — but never the space the
* fixed bars stand in.
*
* Shrinking the padding by the whole slack pulled the terminal's bottom edge
* DOWN under those bars, and the row the following re-fit then gained was
* painted behind them: on a long wrapped prompt the last line was clipped by
* the accessory bar, i.e. the bottom half of the text being typed. The floor is
* now the bars' MEASURED height, so a device where the hard-coded 84px
* over-reserves still reclaims the difference, while one that genuinely needs
* it keeps every pixel.
*
* ⚠️ The floor can only ever prevent a shrink, never cause a grow
* (`Math.min(currentPadding, …)`): a measured height LARGER than the current
* padding makes this a no-op rather than silently resizing the terminal.
*/
_shrinkPaddingToFit() {
try {
const container = document.getElementById('terminalContainer');
@@ -583,7 +619,8 @@ const KeyboardHandler = {
const gap = container.clientHeight - app.terminal.rows * cellH;
if (gap > 0 && gap < cellH) {
const currentPadding = parseInt(main.style.paddingBottom) || 0;
main.style.paddingBottom = Math.max(0, currentPadding - gap) + 'px';
const floor = Math.min(currentPadding, this._fixedBottomBarsHeight());
main.style.paddingBottom = Math.max(floor, currentPadding - gap) + 'px';
if (app.fitAddon)
try {
app.fitAddon.fit();
+41
View File
@@ -696,6 +696,34 @@ html.mobile-init .file-browser-panel {
text-overflow: ellipsis;
}
/* ⚠️ Reserve a tappable label on the ACTIVE tab, which is the only one that
grows action icons. With a short session name the icons were eating the
tab: "w1" rendered a 13px label while gear + close took 50px of a 116px
tab, so the tab's geometric CENTRE landed on the gear and a thumb aiming
at the tab opened Session Options instead of switching sessions (measured
at 360, 393 and 430px; only long names cleared it). The tab widens by the
difference instead, which costs a little strip space on exactly one tab
and keeps tap-to-switch the majority of it.
⚠️ The floor is set by the 10th tab onward, NOT by the numbered tabs you
are looking at. `.tab-number` is rendered only for `_tabIdx < 9` (app.js),
so tab 10 loses 16px + a 4px gap off its left and its centre sits 10px
further right. The centre clears the icons when
reserved > icons + rightEdge - leftRunUp - gap
= 50 + 9 - 17 - 4 = 38px
with icons = gear 32 + close 20 - close's -2px margin, leftRunUp = border 1
+ padding 8 + status dot 4 + gap 4, and rightEdge = padding 8 + border 1.
Hit testing snaps to whole pixels, so 39px still lands on the gear: the
practical floor is 40px and 44px keeps 4px of headroom. A NUMBERED tab
clears it at 20px, so reasoning from the tabs on screen is exactly what
would put the centre back on the gear. Pinned by
test/mobile-tab-tap-zones.test.ts. */
.session-tab.active .tab-name {
min-width: 44px;
}
/* Hide close/gear buttons on non-active tabs on mobile */
.session-tab .tab-close,
.session-tab .tab-gear {
@@ -3659,6 +3687,19 @@ html[data-session-list="sidebar"] .session-sidebar {
padding-left: var(--safe-area-left);
}
/* The rich variant's desktop column is 300px (--sidebar-width-rich, styles.css)
and its selector carries one attribute MORE than the drawer base above —
(0,3,1) vs (0,2,1) — so without this it would win here and pin the drawer of
a 320px phone to 300px, leaving 20px of terminal behind it. How wide a drawer
may be is a viewport decision, never a row-detail one: match the specificity
and hand the width back. Row detail itself is kept — the stamps are as useful
on a phone as anywhere, and the drawer is wider than the rail they were
designed against. */
html[data-session-list="sidebar"][data-sidebar-detail="rich"] .session-sidebar {
width: min(280px, 80vw);
flex: 0 0 auto;
}
html[data-session-list="sidebar"] .session-sidebar.open {
transform: translateX(0);
visibility: visible;
+1 -1
View File
@@ -2279,7 +2279,7 @@ Object.assign(CodemanApp.prototype, {
const terminal = new Terminal({
theme: { ...window.codemanCurrentXtermTheme() },
minimumContrastRatio: window.codemanCurrentSkinIsLight() ? 4.5 : 1,
fontFamily: '"Fira Code", "Cascadia Code", "JetBrains Mono", "SF Mono", Monaco, monospace',
fontFamily: window.CodemanTerminalFont.resolve(this.loadAppSettingsFromStorage?.().terminalFontFamily),
fontSize: 12,
lineHeight: 1.2,
cursorBlink: true,
+26 -15
View File
@@ -91,28 +91,38 @@ Object.assign(CodemanApp.prototype, {
},
/**
* Colour for one child's arc, from CodemanLineage.COLORS, assigned in FIRST-SEEN
* order and remembered per child id. First-seen rather than draw-index keeps a
* line's colour stable across re-renders, tab reorders and sibling closes (the
* SVG is wiped and rebuilt constantly, so an index-based colour would flicker).
* An empty string means "no override": the CSS falls back to --session-blue.
* Colour for one arc, from CodemanLineage.COLORS, keyed on the SPAWNING tab.
*
* ⚠️ Per PARENT, not per child: every arc leaving one tab is the same colour, no
* matter how many workers it spawns, so the strip reads as "these five came from
* w1, those two came from w2". Keying it per child instead gave one tab's own
* children a different colour each, which is the thing the colours exist to tell
* apart. A child that goes on to spawn its own workers is a parent in its turn and
* gets its own colour for the arcs BELOW it, so a chain changes colour at each
* generation while each generation's fan-out stays uniform.
*
* Assigned in FIRST-SEEN order and remembered per parent id. First-seen rather than
* draw-index keeps a colour stable across re-renders, tab reorders and sibling
* closes (the SVG is wiped and rebuilt constantly, so an index-based colour would
* flicker). An empty string means "no override": the CSS falls back to
* --session-blue, so the first spawning tab keeps the skin-aware blue.
*/
_lineageColorFor(childId) {
_lineageColorFor(parentId) {
const palette = (window.CodemanLineage && window.CodemanLineage.COLORS) || [];
if (palette.length === 0) return '';
if (!this._lineageColorByChild) {
this._lineageColorByChild = new Map();
if (!this._lineageColorByParent) {
this._lineageColorByParent = new Map();
this._lineageColorNext = 0;
}
let idx = this._lineageColorByChild.get(childId);
let idx = this._lineageColorByParent.get(parentId);
if (idx === undefined) {
idx = this._lineageColorNext++ % palette.length;
this._lineageColorByChild.set(childId, idx);
this._lineageColorByParent.set(parentId, idx);
// Bounded: entries for long-gone sessions are pruned once the map is clearly
// stale, so a day-long dashboard cannot grow it without limit.
if (this._lineageColorByChild.size > 200 && this.sessions) {
for (const key of this._lineageColorByChild.keys()) {
if (!this.sessions.has(key)) this._lineageColorByChild.delete(key);
if (this._lineageColorByParent.size > 200 && this.sessions) {
for (const key of this._lineageColorByParent.keys()) {
if (!this.sessions.has(key)) this._lineageColorByParent.delete(key);
}
}
}
@@ -174,9 +184,10 @@ Object.assign(CodemanApp.prototype, {
// the line itself. `status` is the CHILD's, which is the interesting end.
const working = edge.status === 'working' ? ' lineage-line--working' : '';
line.setAttribute('class', 'connection-line lineage-line' + working);
// Per-child colour rides a CSS custom property so the stylesheet keeps owning
// The PARENT's colour rides a CSS custom property so the stylesheet keeps owning
// opacity, glow and dash; an empty colour leaves the --session-blue fallback.
const color = this._lineageColorFor(edge.childId);
// Every arc out of one tab shares it — see _lineageColorFor().
const color = this._lineageColorFor(edge.parentId);
if (color) line.style.setProperty('--lineage-color', color);
// `data-agent-id` is what _applyLineEntrances() queries — see the file header.
line.setAttribute('data-agent-id', 'lineage:' + edge.childId);
+31
View File
@@ -41,8 +41,17 @@ Object.assign(CodemanApp.prototype, {
_onHookElicitationComplete(data) {
// Question answered in the terminal: clear the action alert without
// waiting for `stop` (the turn may keep running for a long time).
// ⚠️ BOTH action kinds, matching the server's APPROVAL_RESOLVING_EVENTS,
// which resolves a session's pending item whatever its kind. An
// AskUserQuestion dialog arrives as `permission_prompt` (only MCP
// elicitation is `elicitation_dialog`), so clearing just the elicitation
// entry left the red alert armed on exactly the dialog these events are
// most often about. Normally the server's `approval:resolved` broadcast
// clears it too; this is the path that still works when the store holds no
// item for the session (restart, superseded).
if (data.sessionId) {
this.clearPendingHooks(data.sessionId, 'elicitation_dialog');
this.clearPendingHooks(data.sessionId, 'permission_prompt');
}
},
@@ -381,6 +390,11 @@ Object.assign(CodemanApp.prototype, {
document.getElementById('appSettingsTunnelEnabled').checked = settings.tunnelEnabled ?? false;
this.loadTunnelStatus();
document.getElementById('appSettingsLocalEcho').checked = settings.localEchoEnabled ?? MobileDetection.isTouchDevice();
// Auto Copy (copy-on-select): per-device, default OFF everywhere. It quietly
// overwrites the system clipboard on a gesture the user may have meant only as
// a way to read, so it is opt-in rather than a default anyone has to discover.
document.getElementById('appSettingsAutoCopySelection').checked = settings.autoCopySelection === true;
document.getElementById('appSettingsTerminalFont').value = settings.terminalFontFamily || '';
document.getElementById('appSettingsTerminalWheelLocal').checked =
settings.terminalWheelLocalScrollback ?? defaults.terminalWheelLocalScrollback ?? false;
document.getElementById('appSettingsCjkInput').checked = settings.cjkInputEnabled ?? defaults.cjkInputEnabled ?? false;
@@ -2006,6 +2020,8 @@ Object.assign(CodemanApp.prototype, {
imageWatcherEnabled: document.getElementById('appSettingsImageWatcherEnabled').checked,
tunnelEnabled: document.getElementById('appSettingsTunnelEnabled').checked,
localEchoEnabled: document.getElementById('appSettingsLocalEcho').checked,
autoCopySelection: document.getElementById('appSettingsAutoCopySelection').checked,
terminalFontFamily: document.getElementById('appSettingsTerminalFont').value.trim(),
terminalWheelLocalScrollback: document.getElementById('appSettingsTerminalWheelLocal').checked,
cjkInputEnabled: document.getElementById('appSettingsCjkInput').checked,
webglRendererEnabled: document.getElementById('appSettingsWebglRenderer').checked,
@@ -2052,6 +2068,7 @@ Object.assign(CodemanApp.prototype, {
// Save to localStorage
this.saveAppSettingsToStorage(settings);
this._updateLocalEchoState();
this.applyTerminalFontFamily?.(settings.terminalFontFamily);
// A real OFF→ON flip of the WebGL toggle retires the GPU-stall auto-fallback
// marker so the next reload actually re-tries WebGL. Only the transition
@@ -2200,6 +2217,14 @@ Object.assign(CodemanApp.prototype, {
showFileViewerButton: _fvb,
webglRendererEnabled: _wgl,
terminalWheelLocalScrollback: _twls,
// Copy-on-select. Per-device (clipboard access differs by device and by
// origin: the plain-HTTP LAN install has no navigator.clipboard at all)
// and absent from SettingsUpdateSchema (.strict()), so sending it would
// 400 the whole settings PUT.
autoCopySelection: _acs,
// Per-device by nature (the font must exist on the device) and absent
// from SettingsUpdateSchema (.strict()) — sending it would 400 the PUT.
terminalFontFamily: _tff,
// Per-device header/toolbar button toggles — client-only, and absent from
// SettingsUpdateSchema (.strict()), so sending them would 400 the PUT.
showSessionButton: _ssb,
@@ -2649,6 +2674,10 @@ Object.assign(CodemanApp.prototype, {
// rules unchanged. Kept here rather than only in applySessionListLayout() so
// that a stray applyTabWrapSettings() call (this one is invoked from
// saveAppSettings and from the resize path) cannot leave the sidebar wrapped.
// Matches BOTH sidebar variants: isSessionSidebarActive() reads
// data-session-list, which applySessionListLayout() sets to 'sidebar' for
// 'sidebar' and 'sidebar-rich' alike. Row detail rides on a separate
// attribute and has no bearing on wrapping.
const sidebar = this.isSessionSidebarActive?.() === true;
// Two-row tabs disabled on mobile/tablet — not enough screen space
const twoRows = !sidebar && deviceType === 'desktop'
@@ -2870,8 +2899,10 @@ Object.assign(CodemanApp.prototype, {
'showMonitor', 'showProjectInsights', 'showFileBrowser', 'showSubagents',
'subagentActiveTabOnly', 'tabTwoRows', 'sessionListLayout', 'localEchoEnabled', 'cjkInputEnabled', 'extendedKeyboardBar',
'skin', 'showPlanUsageLimits', 'showAttachmentsButton', 'showFileViewerButton', 'webglRendererEnabled',
'terminalFontFamily',
'language',
'terminalWheelLocalScrollback',
'autoCopySelection',
'showSessionButton', 'showAwayDigestButton', 'showCronButton',
'showTabDetachButton',
'mobileOverviewEnabled',
+198
View File
@@ -16,6 +16,19 @@
font-weight: 400 700;
src: url('fonts/jetbrains-mono-variable.woff2') format('woff2');
}
/* Icons-only per-glyph fallback for the terminal (Symbols Nerd Font Mono, MIT,
fonts/LICENSE-nerd-fonts.txt). Sits BEHIND the text fonts in the xterm stack
(constants.js: TERMINAL_FONT_DEFAULT_STACK), so it only ever supplies the
private-use-area glyphs shell prompts draw (powerline, p10k/starship folder
and git icons) — text rendering is untouched. `block` display, not `swap`:
there is no fallback that CAN render these glyphs, so swapping in tofu first
would poison xterm's glyph atlas until the next re-render. */
@font-face {
font-family: 'Symbols Nerd Font Mono';
font-style: normal;
font-display: block;
src: url('fonts/symbols-nerd-font-mono.woff2') format('woff2');
}
:root {
/* Dark is the safe fallback. Light skins override this so native selects,
@@ -50,6 +63,7 @@
--header-height: 36px;
--toolbar-height: 42px;
--sidebar-width: 260px;
--sidebar-width-rich: 300px; /* detailed rows carry a stamps line as well */
--sidebar-width-collapsed: 44px; /* == --touch-target-min */
--sidebar-transition: 0.18s ease;
--glass-bg: rgba(31, 38, 48, 0.85);
@@ -3394,6 +3408,57 @@ body.touch-device .terminal-container .xterm .xterm-helper-textarea {
-webkit-touch-callout: none !important;
}
/* Touch text-selection bar (long-press → select → Copy).
Lives in styles.css, NOT mobile.css: the gesture is touch-driven, not
width-driven, and mobile.css is media-gated to ≤1023px — a touch tablet in
landscape would get the gesture with no bar to copy from.
Built in JS (index.html is read once at server start, so markup added there
would need a restart to appear). z-index 900 sits above terminal content and
the local-echo overlay (7) and deliberately BELOW floating agent windows
(1000), so it can never cover their controls. */
.term-select-bar {
position: absolute;
z-index: 900;
display: none;
gap: 2px;
padding: 3px;
background: var(--bg-card);
border: 1px solid var(--border);
border-radius: 8px;
box-shadow: 0 4px 14px rgba(0, 0, 0, 0.45);
}
.term-select-bar.visible {
display: flex;
}
.term-select-btn {
min-height: 38px;
min-width: 46px;
padding: 0 0.7rem;
border: none;
border-radius: 6px;
background: transparent;
color: var(--text);
font-family: inherit;
font-size: 0.82rem;
font-weight: 600;
cursor: pointer;
/* The bar is the one place in the terminal subtree a tap must land on a
control rather than a cell, so it opts out of the gesture styles above. */
touch-action: manipulation;
}
.term-select-btn:active {
background: var(--bg-hover);
}
.term-select-btn--close {
min-width: 38px;
padding: 0;
color: var(--text-muted);
}
/* Welcome Overlay */
.welcome-overlay {
position: absolute;
@@ -16675,6 +16740,14 @@ html[data-session-list="sidebar"] .session-sidebar {
dropdowns and the inline rename input paintable in place. */
}
/* Rich rows carry a stamps line the simple rows do not, and at 260px
"created 3d ago · working 12m" ellipsizes before it is finished. Overridden
by the collapsed rule below, which is more specific and comes after. */
html[data-session-list="sidebar"][data-sidebar-detail="rich"] .session-sidebar {
flex-basis: var(--sidebar-width-rich);
width: var(--sidebar-width-rich);
}
html[data-session-list="sidebar"][data-sidebar="collapsed"] .session-sidebar {
flex-basis: var(--sidebar-width-collapsed);
width: var(--sidebar-width-collapsed);
@@ -16797,6 +16870,131 @@ html[data-session-list="sidebar"] .session-tab.tab-filtered-out {
display: none !important;
}
/* --- Rich rows (sessionListLayout 'sidebar-rich') ----------------------- */
/* The detailed variant of the SAME sidebar: identical column, identical
re-parented #sessionTabs, identical filter and Alt+B toggle. The only
difference is that each row also carries the line the desktop home rail and
the phone overview carry — when the session was first created, how long it
has been in the state it is in, and a status pill.
Everything here is scoped to html[data-sidebar-detail="rich"], which
applySessionListLayout() only ever sets to 'rich' while data-session-list is
'sidebar'. `.tab-meta` is emitted by the row template exclusively in that
mode, so these rules have nothing to match anywhere else — the display:none
below is the second lock, not the mechanism. */
html[data-sidebar-detail="rich"] .session-sidebar .tab-meta {
display: flex;
align-items: center;
gap: 0.35em;
min-width: 0;
margin-top: 0.15em;
font-size: 0.62rem;
font-family: monospace;
line-height: 1.3;
color: var(--text-muted);
opacity: 0.8;
white-space: nowrap;
overflow: hidden;
}
/* A meta line can only be produced by the rich row template, but if one ever
survives into another layout (a render that lost a race with a settings flip)
it must not paint: the header strip has no room for it. */
.session-tab .tab-meta {
display: none;
}
html[data-sidebar-detail="rich"] .session-sidebar .tab-meta-item {
min-width: 0;
overflow: hidden;
text-overflow: ellipsis;
}
html[data-sidebar-detail="rich"] .session-sidebar .tab-meta-key {
margin-right: 0.35em;
opacity: 0.7;
text-transform: uppercase;
letter-spacing: 0.06em;
}
html[data-sidebar-detail="rich"] .session-sidebar .tab-meta-sep {
opacity: 0.45;
}
/* While a session is actually doing something, how long it has been doing it is
what the eye should land on — same emphasis the home rail gives it. */
html[data-sidebar-detail="rich"] .session-sidebar .session-tab.tab-state-working .tab-meta-since {
color: var(--green);
opacity: 0.95;
}
/* Pushed hard right and never shrinking, so the stamps ellipsize before the
status word does. */
html[data-sidebar-detail="rich"] .session-sidebar .tab-pill {
flex-shrink: 0;
margin-left: auto;
padding: 0.1em 0.5em;
border-radius: 999px;
background: var(--bg-input);
border: 1px solid var(--border);
color: var(--text-muted);
font-size: 0.95em;
font-weight: 700;
letter-spacing: 0.02em;
white-space: nowrap;
}
/* Same three colors as every other session surface: red means a question is
pending, yellow means it wants input, green means work is happening. */
html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--needs,
html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--error {
background: color-mix(in srgb, var(--red) 18%, transparent);
border-color: color-mix(in srgb, var(--red) 45%, transparent);
color: var(--red);
}
html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--waiting {
background: color-mix(in srgb, var(--yellow) 18%, transparent);
border-color: color-mix(in srgb, var(--yellow) 45%, transparent);
color: var(--yellow);
}
html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--working {
background: color-mix(in srgb, var(--green) 15%, transparent);
border-color: color-mix(in srgb, var(--green) 40%, transparent);
color: var(--green);
}
/* Muted one step further than the idle dot: the pill is a block of color, so it
reads louder than a 9px dot at the same mix. */
html[data-sidebar-detail="rich"] .session-sidebar .tab-pill--idle {
background: color-mix(in srgb, var(--green) 7%, transparent);
border-color: color-mix(in srgb, var(--green) 18%, var(--border));
color: color-mix(in srgb, var(--green) 45%, var(--text-muted));
}
/* Three lines of content per row instead of two, so give them room to breathe
and stop the row actions crowding the pill.
:not([data-sidebar="collapsed"]) is load-bearing, not decoration: the two
rules below are the only ones in this block that move geometry rather than
paint the meta line, and the collapsed 44px rail centres a row that is by
then just a status dot and its badges. Without the guard, `align-items:
flex-start` and a 0.15rem top margin on .tab-status would push that dot off
the centre line of every row in the rail. */
html[data-sidebar-detail="rich"]:not([data-sidebar="collapsed"]) .session-sidebar .session-tab {
align-items: flex-start;
padding: 0.45rem 0.5rem;
}
/* The gear/detach/close column is centred against a two-line row; against a
three-line one it drifts low, so pin it to the name it acts on. */
html[data-sidebar-detail="rich"]:not([data-sidebar="collapsed"]) .session-sidebar .session-tab .tab-actions,
html[data-sidebar-detail="rich"]:not([data-sidebar="collapsed"]) .session-sidebar .session-tab .tab-number,
html[data-sidebar-detail="rich"]:not([data-sidebar="collapsed"]) .session-sidebar .session-tab .tab-status {
margin-top: 0.15rem;
}
/* --- Collapsed rail ---------------------------------------------------- */
/* Collapsed is a 44px icon rail, not "hidden": the ambient signal (status dot,
task/subagent/ultracode badges) is the whole point of mission control and
File diff suppressed because it is too large Load Diff
+397
View File
@@ -0,0 +1,397 @@
/**
* @fileoverview Repository-status check (App Settings → Updates → "Repository
* status").
*
* INFORMATIONAL companion to the release-tag self-updater (`self-update.ts`).
* Where the updater answers "is there a newer published release tag, and do you
* want to `git checkout` it", this module answers "where does my local checkout
* sit relative to the upstream project AND my own fork" — by commit ahead/behind
* count plus a short list of the incoming commits.
*
* This needs the actual commits locally, so it does a READ-ONLY `git fetch` of
* the compared ref per remote (updates only remote-tracking refs under `.git`,
* never the working tree), then `git rev-list`/`git log`. Works uniformly for
* GitHub and non-GitHub remotes (e.g. Bitbucket) — no release tags required.
*
* Remote selection (generalizable): by default the union of `origin` and the
* current branch's `@{upstream}` tracking remote, deduped. Override with the
* `CODEMAN_UPDATE_REMOTES` env var (comma-separated remote names) for any other
* layout (e.g. the common `origin`=fork / `upstream`=canonical convention).
*
* Split PURE helpers (parsing + remote-set/role/compare-ref decisions, unit
* tested) from the IO wrapper `getRepositoryStatus()` (touches git).
*
* PERFORMANCE: every git invocation is ASYNC (`execFile`), never `execFileSync`
* — the fetch path can spend `2 × FETCH_TIMEOUT_MS` per remote on the network,
* and a synchronous version froze the whole event loop (SSE, PTY streaming) for
* up to a minute per request. The computation is additionally single-flight
* with a short TTL cache: concurrent requests share one in-flight promise, and
* a fresh-enough result is served without spawning git at all.
*
* SECURITY: remote URLs (and git stderr echoing them) can embed
* `scheme://user:token@host` credentials, so every `url`/`error` field is
* passed through `redactGitCredentials()` (shared with `git-clone.ts`) before
* it reaches a client.
*
* @module web/repo-status
*/
import { execFile } from 'node:child_process';
import { createRequire } from 'node:module';
import { promisify } from 'node:util';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import { gitNonInteractiveEnv, redactGitCredentials } from '../git-clone.js';
import { getInstallInfo } from './self-update.js';
import type { RepoIncomingCommit, RepoRemoteRole, RepoRemoteStatus, RepositoryStatusResult } from '../types/update.js';
const require = createRequire(import.meta.url);
const { version: APP_VERSION } = require('../../package.json') as { version: string };
const execFileAsync = promisify(execFile);
/** Network/git timeout for the fetch path (longer than EXEC_TIMEOUT_MS — hits network). */
const FETCH_TIMEOUT_MS = 15_000;
/** Max incoming commit subjects to list per remote. */
const MAX_INCOMING = 10;
/** Fresh-enough window for a cached status — repeated UI polls reuse it instead of re-running git. */
const STATUS_CACHE_TTL_MS = 45_000;
// ─────────────────────────────────────────────────────────────────────────────
// PURE helpers (unit tested, no IO)
// ─────────────────────────────────────────────────────────────────────────────
/**
* Parse `git rev-list --left-right --count HEAD...<ref>` output.
* Git prints two tab-separated counts: LEFT (commits in HEAD not in ref → ahead)
* and RIGHT (commits in ref not in HEAD → behind). Returns null on malformed input.
*/
export function parseAheadBehind(out: string): { ahead: number; behind: number } | null {
const m = out.trim().match(/^(\d+)\s+(\d+)$/);
if (!m) return null;
return { ahead: parseInt(m[1], 10), behind: parseInt(m[2], 10) };
}
/**
* Parse `git log --oneline` output into commits. Each non-empty line is
* `<sha> <subject>`; the first whitespace-delimited token is the SHA.
*/
export function parseLogLines(out: string): RepoIncomingCommit[] {
const commits: RepoIncomingCommit[] = [];
for (const line of out.split('\n')) {
const trimmed = line.trim();
if (!trimmed) continue;
const idx = trimmed.indexOf(' ');
if (idx === -1) {
commits.push({ sha: trimmed, subject: '' });
} else {
commits.push({ sha: trimmed.slice(0, idx), subject: trimmed.slice(idx + 1).trim() });
}
}
return commits;
}
/**
* Extract the default branch name from `git ls-remote --symref <remote> HEAD`
* output (a line like `ref: refs/heads/master\tHEAD`). Returns null if absent.
*
* SECURITY: this parses UNTRUSTED remote output, and the result flows into a
* later `git fetch <remote> <branch>` positional. The capture is constrained to
* start with an alphanumeric (no leading `-`) so a hostile remote can't return
* `ref: refs/heads/--upload-pack=<cmd>\tHEAD` and smuggle an argv flag (RCE) —
* see `isSafeGitPositional` for the defense-in-depth re-check at the call site.
*/
export function parseSymrefDefaultBranch(out: string): string | null {
const m = out.match(/^ref:\s+refs\/heads\/([A-Za-z0-9_./][A-Za-z0-9_./+-]*)\s+HEAD$/m);
return m ? m[1] : null;
}
/**
* Reject a value that would be unsafe as a git positional argument (remote name
* or branch). A leading `-` lets untrusted ls-remote/symref output or a stray
* `CODEMAN_UPDATE_REMOTES` entry inject an option (e.g. `--upload-pack=<cmd>`)
* into a subsequent `git fetch`. Empty values are rejected too.
*/
export function isSafeGitPositional(value: string | null | undefined): value is string {
return typeof value === 'string' && value.length > 0 && !value.startsWith('-');
}
/** Split a comma-separated env value into trimmed, non-empty names. */
export function parseRemotesEnv(value: string | null | undefined): string[] {
if (!value) return [];
return value
.split(',')
.map((s) => s.trim())
.filter(Boolean);
}
/**
* Decide which remote NAMES the status view covers, in display order.
*
* - If `envRemotes` is non-empty, use exactly those that actually exist (order
* preserved). This is the full-control escape hatch.
* - Otherwise the union of `origin` (if it exists) and the tracking remote (if
* any), deduped. The tracking remote is listed first when it isn't `origin`,
* so "your fork" leads and "upstream" follows.
*/
export function resolveRemoteSet(opts: {
existingRemotes: string[];
trackingRemote: string | null;
envRemotes: string[];
}): string[] {
const exists = new Set(opts.existingRemotes);
if (opts.envRemotes.length > 0) {
return dedupe(opts.envRemotes.filter((n) => exists.has(n)));
}
const out: string[] = [];
if (opts.trackingRemote && exists.has(opts.trackingRemote) && opts.trackingRemote !== 'origin') {
out.push(opts.trackingRemote);
}
if (exists.has('origin')) out.push('origin');
if (opts.trackingRemote && exists.has(opts.trackingRemote)) out.push(opts.trackingRemote);
return dedupe(out);
}
/** Classify a remote's role relative to the tracking remote. */
export function roleForRemote(name: string, trackingRemote: string | null): RepoRemoteRole {
if (trackingRemote && name === trackingRemote) return 'tracking';
if (name === 'origin' || name === 'upstream') return 'upstream';
return 'other';
}
function dedupe(names: string[]): string[] {
return [...new Set(names)];
}
/**
* Extract the remote name from an `@{upstream}` short ref like `origin/master`.
*
* A ref with NO slash is a LOCAL-branch upstream (`git branch -u otherbranch`):
* `rev-parse --abbrev-ref @{upstream}` prints just the branch name, and there
* is no remote in it — the old `slice(0, indexOf('/'))` became `slice(0, -1)`
* there and yielded garbage like `maste`. Returns null for that case (treated
* as "no tracking remote"), and for empty/degenerate refs.
*/
export function parseTrackingRemote(trackingRef: string | null | undefined): string | null {
if (!trackingRef) return null;
const idx = trackingRef.indexOf('/');
if (idx <= 0) return null;
return trackingRef.slice(0, idx);
}
/**
* Redact embedded `scheme://user:secret@host` credentials from the fields of
* one remote's status that can carry them: the remote URL itself, and the
* error string (which may include git stderr echoing that URL back).
*/
export function redactRemoteStatus(status: RepoRemoteStatus): RepoRemoteStatus {
const out: RepoRemoteStatus = { ...status, url: redactGitCredentials(status.url) };
if (out.error !== undefined) out.error = redactGitCredentials(out.error);
return out;
}
/**
* Single-flight TTL cache around an async compute: concurrent callers await
* the SAME in-flight promise, and a result younger than `ttlMs` is returned
* without recomputing. A rejected compute is never cached, so the next call
* retries. `now` is injectable for deterministic TTL tests.
*/
export function createSingleFlightCache<T>(
ttlMs: number,
compute: () => Promise<T>
): { get(now?: number): Promise<T> } {
let cachedAt = 0;
let cachedValue: T | undefined;
let hasValue = false;
let inFlight: Promise<T> | null = null;
return {
get(now = Date.now()): Promise<T> {
if (hasValue && now - cachedAt < ttlMs) return Promise.resolve(cachedValue as T);
if (inFlight) return inFlight;
inFlight = compute().then(
(value) => {
cachedValue = value;
hasValue = true;
cachedAt = Date.now();
inFlight = null;
return value;
},
(err: unknown) => {
inFlight = null;
throw err;
}
);
return inFlight;
},
};
}
// ─────────────────────────────────────────────────────────────────────────────
// IO wrapper
// ─────────────────────────────────────────────────────────────────────────────
interface GitResult {
ok: boolean;
stdout: string;
stderr: string;
}
/**
* Run git non-interactively (no credential or SSH prompts — a missing key/cred
* fails fast instead of hanging; env via the shared `gitNonInteractiveEnv()`,
* which also closes the askpass/GCM/DISPLAY prompt paths). ASYNC on purpose:
* the fetch path hits the network for up to `FETCH_TIMEOUT_MS`, and a sync
* spawn would block the event loop for the whole wait.
*/
async function runGit(args: string[], cwd: string, timeout = EXEC_TIMEOUT_MS): Promise<GitResult> {
try {
const { stdout } = await execFileAsync('git', args, {
cwd,
encoding: 'utf-8',
timeout,
env: gitNonInteractiveEnv(),
});
return { ok: true, stdout: stdout.trim(), stderr: '' };
} catch (err: unknown) {
const e = err as { stdout?: Buffer | string; stderr?: Buffer | string };
return {
ok: false,
stdout: e.stdout ? String(e.stdout).trim() : '',
stderr: e.stderr ? String(e.stderr).trim() : '',
};
}
}
/** First line of stderr, trimmed — a compact human-readable failure reason. */
function firstLine(s: string): string {
return (
s
.split('\n')
.find((l) => l.trim())
?.trim() ?? 'git command failed'
);
}
/** Compute ahead/behind + incoming commits for one already-selected remote. */
async function statusForRemote(
dir: string,
name: string,
trackingRemote: string,
trackingRef: string | null
): Promise<RepoRemoteStatus> {
const role = roleForRemote(name, trackingRemote);
const url = (await runGit(['remote', 'get-url', name], dir)).stdout || '';
const base: RepoRemoteStatus = { name, url, role, compareRef: '', ahead: 0, behind: 0, incoming: [] };
// SECURITY: the remote name reaches git as a positional; reject a `-` prefix
// (e.g. a stray CODEMAN_UPDATE_REMOTES entry) before it can act as a flag.
if (!isSafeGitPositional(name)) {
return { ...base, error: `Refusing unsafe remote name "${name}".` };
}
// Resolve the compare ref + the remote branch to fetch.
let branch: string | null;
if (role === 'tracking' && trackingRef) {
// e.g. trackingRef = "bitbucket/local" → branch = "local"
branch = trackingRef.slice(name.length + 1) || null;
} else {
const symref = await runGit(['ls-remote', '--symref', name, 'HEAD'], dir, FETCH_TIMEOUT_MS);
branch = symref.ok ? parseSymrefDefaultBranch(symref.stdout) : null;
if (!branch && !symref.ok) {
return { ...base, error: `Could not reach ${name}: ${firstLine(symref.stderr)}` };
}
branch = branch ?? 'master';
}
if (!branch) return { ...base, error: `Could not resolve a branch on ${name}.` };
// SECURITY: defense-in-depth — `branch` may come from untrusted symref output
// or a tracking-ref slice; never let a `-`-prefixed value reach `git fetch`.
if (!isSafeGitPositional(branch)) {
return { ...base, error: `Refusing unsafe branch name "${branch}" from ${name}.` };
}
const compareRef = `${name}/${branch}`;
// Read-only fetch of just that ref so the local rev-list/log can see it.
// `--` ends option parsing so neither `name` nor `branch` can be read as a flag.
const fetched = await runGit(['fetch', '--no-tags', name, '--', branch], dir, FETCH_TIMEOUT_MS);
if (!fetched.ok) {
return { ...base, compareRef, error: `Could not fetch ${compareRef}: ${firstLine(fetched.stderr)}` };
}
const counts = await runGit(['rev-list', '--left-right', '--count', `HEAD...${compareRef}`], dir);
if (!counts.ok) {
return { ...base, compareRef, error: `Could not compare against ${compareRef}: ${firstLine(counts.stderr)}` };
}
const ab = parseAheadBehind(counts.stdout);
if (!ab) return { ...base, compareRef, error: `Unexpected git output comparing ${compareRef}.` };
const log = await runGit(['log', '--oneline', '-n', String(MAX_INCOMING), `HEAD..${compareRef}`], dir);
const incoming = log.ok ? parseLogLines(log.stdout) : [];
return { ...base, compareRef, ahead: ab.ahead, behind: ab.behind, incoming };
}
/** The uncached computation behind `getRepositoryStatus()`. */
async function computeRepositoryStatus(): Promise<RepositoryStatusResult> {
const checkedAt = Date.now();
const info = getInstallInfo();
const base: RepositoryStatusResult = {
checkedAt,
isGit: info.installKind === 'git',
currentVersion: info.currentVersion || APP_VERSION,
remotes: [],
};
if (info.installKind !== 'git') {
return { ...base, error: 'Not a git install — repository status is unavailable.' };
}
const dir = info.installDir;
// Tracking ref of the current branch, e.g. "bitbucket/local" (empty if none).
const trackingRef =
(await runGit(['rev-parse', '--abbrev-ref', '--symbolic-full-name', '@{upstream}'], dir)).stdout || null;
const trackingRemote = parseTrackingRemote(trackingRef);
// A local-branch upstream (no slash) carries no remote — drop the ref too, so
// nothing downstream can mistake a bare branch name for a remote-tracking ref.
const remoteTrackingRef = trackingRemote ? trackingRef : null;
const remotesOut = await runGit(['remote'], dir);
const existingRemotes = remotesOut.ok
? remotesOut.stdout
.split('\n')
.map((s) => s.trim())
.filter(Boolean)
: [];
const selected = resolveRemoteSet({
existingRemotes,
trackingRemote,
envRemotes: parseRemotesEnv(process.env.CODEMAN_UPDATE_REMOTES),
});
if (selected.length === 0) {
return { ...base, error: 'No comparable remotes found (set CODEMAN_UPDATE_REMOTES to choose).' };
}
// Sequential on purpose: concurrent `git fetch` in one repo can contend on
// ref locks, and the event loop no longer cares how long this takes.
const remotes: RepoRemoteStatus[] = [];
for (const name of selected) {
// SECURITY: redact `scheme://user:token@host` credentials from the URL and
// any error string (git stderr echoes the URL back) before they leave the server.
remotes.push(redactRemoteStatus(await statusForRemote(dir, name, trackingRemote ?? '', remoteTrackingRef)));
}
return { ...base, remotes };
}
const statusCache = createSingleFlightCache(STATUS_CACHE_TTL_MS, computeRepositoryStatus);
/**
* Inspect how the local checkout sits relative to the configured remotes.
* Each remote is fetched + compared independently; a single unreachable remote
* surfaces as that card's `error` and never fails the whole call.
*
* Single-flight + TTL-cached: concurrent requests share one in-flight
* computation, and a result younger than `STATUS_CACHE_TTL_MS` is served
* without spawning git.
*/
export function getRepositoryStatus(): Promise<RepositoryStatusResult> {
return statusCache.get();
}
+361
View File
@@ -0,0 +1,361 @@
export type ResponseViewerTranscriptKind = 'prompt' | 'response' | 'status' | 'tool';
export interface ResponseViewerTranscriptBlock {
kind: ResponseViewerTranscriptKind;
label: 'Prompt' | 'Response' | 'Status' | 'Tool';
/** What the frontend renders by: 'user' gets the "You" badge, everything else the agent badge. */
role: 'user' | 'assistant';
text: string;
}
// Keep in lockstep with isExternalCliMode() in src/session.ts. Importing it here
// would drag node-pty and the whole session layer into this pure module, so the
// list is duplicated and test/response-viewer-transcript.test.ts pins the parity.
const EXTERNAL_CLI_MODES = new Set(['codex', 'gemini', 'opencode', 'antigravity', 'pi']);
function isPromptLine(line: string): boolean {
return /^\s*›\s*/.test(line);
}
function isDividerDashChar(ch: string): boolean {
return ch === '─' || ch === '-';
}
function isWhitespaceChar(ch: string): boolean {
return /\s/.test(ch);
}
// Linear-time equivalent of the old /^[─-]+\s*(.+?)\s*[─-]{3,}$/. The lazy
// middle of that pattern backtracked catastrophically on a long dash run that
// does NOT end in 3+ dashes (measured >2min at 8,000 chars) — and pane text is
// agent-controlled with buffers up to 32MB, so this ran on hostile input. Same
// accept set and same captured content, computed with counters; equivalence is
// pinned char-for-char against the old regex by the brute-force corpus test in
// test/response-viewer-transcript.test.ts.
function normalizeDividerStatusLine(line: string): string | null {
const s = line.trim();
const n = s.length;
// Minimum match: 1 leading dash + 1 content char + 3 trailing dashes.
if (n < 5) return null;
let lead = 0;
while (lead < n && isDividerDashChar(s.charAt(lead))) lead += 1;
if (lead === 0) return null;
let trail = 0;
while (trail < n && isDividerDashChar(s.charAt(n - 1 - trail))) trail += 1;
if (trail < 3) return null;
// The regex was greedy on the leading run but gave dashes back until at least
// one content char plus the 3-dash tail fit (an all-dash line matched with a
// single leftover dash as its "content"), so the content window starts at the
// end of the leading run, clamped to leave 4 chars.
const contentStart = Math.min(lead, n - 4);
let ws = 0;
while (contentStart + ws < n - 4 && isWhitespaceChar(s.charAt(contentStart + ws))) ws += 1;
const from = contentStart + ws;
// The lazy middle stopped at the first position from which "optional
// whitespace, then dashes to end-of-line" matches: the start of the
// whitespace padding in front of the trailing dash run (never before the
// first content char).
const tailStart = n - trail;
let padded = tailStart;
while (padded > 0 && isWhitespaceChar(s.charAt(padded - 1))) padded -= 1;
const end = Math.max(from + 1, padded);
return s.slice(from, end).trim() || null;
}
function isDividerOnlyLine(line: string): boolean {
return /^[\s─-]{8,}$/.test(line.trim());
}
// COD-227: unambiguous Codex tool-call markers. These only ever appear as internal
// activity, never as ordinary assistant prose, so they are always a Tool block.
function isToolActivityMarker(line: string): boolean {
return /^\s*[•*-]\s+(Calling|Called)\b/.test(line.trim());
}
// COD-227: action verbs that ALSO occur in ordinary assistant prose (e.g.
// "• Created COD-226: …"). These are a Tool header only when corroborated by a
// box-drawing result tree on the next non-blank line (see the caller); the verb
// alone is not sufficient.
function isToolVerbBullet(line: string): boolean {
return /^\s*[•*-]\s+(Explored|Viewed|Read|Edited|Updated|Created|Deleted|Ran|Searched|Opened|Listed|Found|Applied|Patched|Used|Wrote|Executed|Modified|Analyzed|Compared|Fetched|Installed)\b/.test(
line.trim()
);
}
// A genuine Codex tool block renders its result as a box-drawing tree (└ │ ├).
function isBoxDrawingLine(line: string): boolean {
return /^[│├└]/.test(line.trim());
}
// Look past blank lines from `fromIndex + 1` for the next non-blank line and report
// whether it is a box-drawing tool-result line — the signal that a verb bullet is a
// real tool block rather than assistant prose that happens to start with a verb.
function nextNonBlankIsBoxDrawing(lines: string[], fromIndex: number): boolean {
for (let j = fromIndex + 1; j < lines.length; j += 1) {
const trimmed = (lines[j] || '').trim();
if (!trimmed) continue;
return isBoxDrawingLine(trimmed);
}
return false;
}
function isToolContinuationLine(line: string, currentKind: ResponseViewerTranscriptKind | null): boolean {
if (currentKind !== 'tool') return false;
const trimmed = line.trimEnd();
if (!trimmed) return true;
return /^\s*[│├└]/.test(trimmed) || /^\s{2,}\S/.test(line);
}
function isStatusLine(line: string, mode: string): boolean {
if (!EXTERNAL_CLI_MODES.has(mode)) return false;
const trimmed = line.trim();
if (!trimmed) return false;
if (normalizeDividerStatusLine(trimmed)) return true;
if (/^(model|directory):\s+/i.test(trimmed)) return true;
if (/\bContext\b.*\bleft\b/i.test(trimmed)) return true;
if (/\b\/model to change\b/i.test(trimmed)) return true;
if (/\bReady\b/i.test(trimmed) && /·/.test(trimmed)) return true;
if (/^(gpt|o\d|claude|gemini)\b/i.test(trimmed) && /·/.test(trimmed)) return true;
if (/^Tip:/i.test(trimmed)) return true;
if (/^\s*[•*-]\s+(Working|Thinking|Loading|Starting\b|Waiting\b)/i.test(trimmed)) return true;
return false;
}
function isStandaloneMarkdownLine(line: string): boolean {
const trimmed = line.trim();
if (!trimmed) return false;
if (/^#{1,6}\s/.test(trimmed)) return true;
if (/^>\s/.test(trimmed)) return true;
if (/^(```|~~~)/.test(trimmed)) return true;
if (/^[-*+]\s/.test(trimmed)) return true;
if (/^\d+[.)]\s/.test(trimmed)) return true;
if (/^\|/.test(trimmed)) return true;
if (/^\s{4,}\S/.test(line)) return true;
return false;
}
function shouldJoinWrappedLine(previous: string, next: string): boolean {
const prev = previous.trimEnd();
const curr = next.trim();
if (!prev || !curr) return false;
if (/[.!?]$/.test(prev)) return false;
if (/[:;]$/.test(prev)) return false;
if (isStandaloneMarkdownLine(curr)) return false;
if (/^[a-z(]/.test(curr)) return true;
if (prev.length >= 72 && /^[A-Za-z0-9"'(]/.test(curr)) return true;
return false;
}
function normalizeWrappedText(lines: string[]): string {
const out: string[] = [];
let paragraph = '';
const flushParagraph = () => {
if (!paragraph) return;
out.push(paragraph);
paragraph = '';
};
for (const rawLine of lines) {
const line = rawLine.trimEnd();
const trimmed = line.trim();
if (!trimmed) {
flushParagraph();
if (out[out.length - 1] !== '') out.push('');
continue;
}
if (isStandaloneMarkdownLine(line)) {
flushParagraph();
out.push(trimmed);
continue;
}
if (!paragraph) {
paragraph = trimmed;
continue;
}
if (shouldJoinWrappedLine(paragraph, trimmed)) {
paragraph += ` ${trimmed}`;
continue;
}
flushParagraph();
paragraph = trimmed;
}
flushParagraph();
return out
.join('\n')
.replace(/\n{3,}/g, '\n\n')
.trim();
}
function cleanTerminalTranscript(buffer: string): string {
// Stripping ANSI/OSC/DCS escape sequences and stray C0/C1 control bytes
// legitimately requires control characters in these patterns.
/* eslint-disable no-control-regex */
return String(buffer || '')
.replace(/\x1b\[[\x30-\x3F]*[\x20-\x2F]*[\x40-\x7E]/g, '')
.replace(/\x1b\][^\x07\x1b]*(?:\x07|\x1b\\)/g, '')
.replace(/\x1b[PX^_][^\x1b]*\x1b\\/g, '')
.replace(/\x1b[NO()][A-Z0-9]?/g, '')
.replace(/\x1b[>=<78cDEHM]/g, '')
.replace(/[\x00-\x08\x0b\x0c\x0e-\x1f\x7f]/g, '')
.replace(/\r\n/g, '\n')
.replace(/\r/g, '\n')
.replace(/[ \t]+$/gm, '')
.trim();
/* eslint-enable no-control-regex */
}
function trimLeadingStartup(lines: string[]): string[] {
let index = 0;
while (index < lines.length) {
const line = lines[index] || '';
const trimmed = line.trim();
if (!trimmed) {
index += 1;
continue;
}
if (/^[╭╰│─].*[╮╯│]?$/.test(trimmed)) {
index += 1;
continue;
}
if (/^>_\s*OpenAI Codex/i.test(trimmed)) {
index += 1;
continue;
}
if (/^(model|directory):\s+/i.test(trimmed)) {
index += 1;
continue;
}
break;
}
return lines.slice(index);
}
function pushBlock(
blocks: ResponseViewerTranscriptBlock[],
kind: ResponseViewerTranscriptKind | null,
lines: string[]
): void {
if (!kind || lines.length === 0) return;
const normalizedLines =
kind === 'status'
? lines.map((line) => normalizeDividerStatusLine(line) || line.trim()).filter((line) => line.length > 0)
: lines;
const text =
kind === 'tool' || kind === 'status'
? normalizedLines
.join('\n')
.replace(/\n{3,}/g, '\n\n')
.trim()
: normalizeWrappedText(normalizedLines);
if (!text) return;
const label = (kind.charAt(0).toUpperCase() + kind.slice(1)) as ResponseViewerTranscriptBlock['label'];
// The frontend's loadFullContext() renders via msg.role — a block without it
// lost the "You" badge on prompts and rendered every block as the agent.
blocks.push({ kind, label, role: kind === 'prompt' ? 'user' : 'assistant', text });
}
export function isExternalCliTranscriptMode(mode: string | null | undefined): boolean {
return EXTERNAL_CLI_MODES.has(String(mode || ''));
}
export function parseExternalCliTranscript(
buffer: string,
mode: string | null | undefined
): ResponseViewerTranscriptBlock[] {
const resolvedMode = String(mode || '');
if (!isExternalCliTranscriptMode(resolvedMode)) return [];
const cleaned = cleanTerminalTranscript(buffer);
if (!cleaned) return [];
const lines = trimLeadingStartup(cleaned.split('\n'));
const blocks: ResponseViewerTranscriptBlock[] = [];
let currentKind: ResponseViewerTranscriptKind | null = null;
let currentLines: string[] = [];
const flush = () => {
pushBlock(blocks, currentKind, currentLines);
currentKind = null;
currentLines = [];
};
for (let i = 0; i < lines.length; i += 1) {
const line = lines[i] ?? '';
// COD-226: within a prompt, the 2-space Codex gutter is authoritative. Blank
// lines and gutter-indented (2+ leading spaces) continuation lines stay in the
// Prompt block ahead of every structural detector below, so multiline prompts —
// including bullets, indented dividers, and literal › lines — are not
// misclassified as responses. Only a non-blank, non-gutter line ends the prompt
// and falls through (a column-0 › then opens a NEW prompt).
if (currentKind === 'prompt') {
if (!line.trim()) {
currentLines.push('');
continue;
}
if (/^ {2,}\S/.test(line)) {
currentLines.push(line);
continue;
}
}
if (isDividerOnlyLine(line)) {
flush();
continue;
}
if (isPromptLine(line)) {
flush();
currentKind = 'prompt';
currentLines = [line.replace(/^\s*›\s*/, '').trim()];
continue;
}
// COD-227: • Calling / • Called are always tool markers; the other action verbs
// are a tool header only when a box-drawing result tree follows on the next
// non-blank line — otherwise the verb bullet is ordinary assistant prose.
if (isToolActivityMarker(line) || (isToolVerbBullet(line) && nextNonBlankIsBoxDrawing(lines, i))) {
if (currentKind !== 'tool') flush();
currentKind = 'tool';
currentLines.push(line.trimEnd());
continue;
}
if (isToolContinuationLine(line, currentKind)) {
currentLines.push(line.trimEnd());
continue;
}
if (isStatusLine(line, resolvedMode)) {
if (currentKind !== 'status') flush();
currentKind = 'status';
currentLines.push(line.trim());
continue;
}
if (!line.trim()) {
currentLines.push('');
continue;
}
if (currentKind !== 'response') flush();
currentKind = 'response';
currentLines.push(line);
}
flush();
return blocks.filter((block) => block.text.trim().length > 0);
}
export function getLastTranscriptResponse(blocks: ResponseViewerTranscriptBlock[]): string {
for (let i = blocks.length - 1; i >= 0; i -= 1) {
if (blocks[i]?.kind === 'response') return blocks[i].text;
}
return '';
}
+13 -3
View File
@@ -63,19 +63,29 @@ export async function readJsonConfig<T>(filePath: string, logLabel: string, defa
* Validates that a file path (possibly containing symlinks) resolves to a location
* within the given session working directory. Returns the resolved and relative paths,
* or null if the path escapes the directory or doesn't exist.
*
* BOTH sides are realpath-resolved before they are compared. Resolving only the
* candidate leaves the two paths in different namespaces whenever the workspace
* itself is reached through a symlink, and `relative()` then reports a spurious
* `../` for a file that is genuinely inside it — refusing every read and write in
* that session. A symlinked workspace is ordinary: `os.tmpdir()` returns one on
* macOS (`/tmp` -> `/private/tmp`), as do symlinked project dirs and bind-mounted
* case paths. Canonicalizing the base only makes the comparison honest; escapes
* are still refused, since the candidate keeps its own realpath.
*/
export function validateSessionFilePath(
sessionWorkingDir: string,
filePath: string
): { resolvedPath: string; relativePath: string } | null {
const fullPath = resolve(sessionWorkingDir, filePath);
let resolvedWorkingDir: string;
let resolvedPath: string;
try {
resolvedPath = realpathSync(fullPath);
resolvedWorkingDir = realpathSync(sessionWorkingDir);
resolvedPath = realpathSync(resolve(sessionWorkingDir, filePath));
} catch {
return null;
}
const relativePath = relative(sessionWorkingDir, resolvedPath);
const relativePath = relative(resolvedWorkingDir, resolvedPath);
if (relativePath.startsWith('..') || isAbsolute(relativePath)) {
return null;
}
+91 -1
View File
@@ -21,6 +21,7 @@ import type {
FileWriteData,
} from '../../types.js';
import { ApiErrorCode, createErrorResponse, getErrorMessage } from '../../types.js';
import { compileFileQuery } from '../../utils/file-query.js';
import { fileStreamManager } from '../../file-stream-manager.js';
import {
AUDIO_ATTACHMENT_EXTENSIONS,
@@ -937,12 +938,15 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
// File tree listing
app.get('/api/sessions/:id/files', async (req) => {
const { id } = req.params as { id: string };
const { depth, showHidden } = req.query as { depth?: string; showHidden?: string };
const { depth, showHidden, q } = req.query as { depth?: string; showHidden?: string; q?: string };
const session = findSessionOrFail(ctx, id, req);
const maxDepth = Math.min(parseInt(depth || '5', 10), 10);
const includeHidden = showHidden === 'true';
const workingDir = session.workingDir;
// null for an empty/whitespace query, which is what keeps the default
// tree response byte-identical when no search is requested.
const matcher = compileFileQuery(q ?? '');
// Default excludes - large/generated directories
const excludeDirs = new Set([
@@ -976,6 +980,92 @@ export function registerFileRoutes(app: FastifyInstance, ctx: SessionPort & Even
let truncated = false;
const maxFiles = 5000;
// ===== Search mode =====
// A query turns this endpoint into a FLAT match list rather than a nested
// tree. It recurses past non-matching directories on purpose — the whole
// point of searching is to reach a file whose ancestors do not match — so
// it is bounded independently by maxMatches on top of the shared maxFiles
// and maxDepth caps, and reports `truncated` when it stops early.
if (matcher) {
const matches: FileTreeNode[] = [];
const maxMatches = 1000;
const searchDirectory = async (dirPath: string, currentDepth: number): Promise<void> => {
if (currentDepth > maxDepth || totalFiles + totalDirectories > maxFiles || matches.length >= maxMatches) {
truncated = true;
return;
}
let entries: import('node:fs').Dirent[];
try {
entries = await fs.readdir(dirPath, { withFileTypes: true });
} catch {
// Can't read directory (permission denied, etc.)
return;
}
entries.sort((a, b) => {
if (a.isDirectory() && !b.isDirectory()) return -1;
if (!a.isDirectory() && b.isDirectory()) return 1;
return a.name.localeCompare(b.name);
});
for (const entry of entries) {
if (totalFiles + totalDirectories > maxFiles || matches.length >= maxMatches) {
truncated = true;
break;
}
if (!includeHidden && entry.name.startsWith('.')) continue;
if (entry.isDirectory() && excludeDirs.has(entry.name)) continue;
const fullPath = join(dirPath, entry.name);
const relativePath = relative(workingDir, fullPath);
if (entry.isDirectory()) {
totalDirectories++;
if (matcher(entry.name, relativePath)) {
matches.push({ name: entry.name, path: relativePath, type: 'directory' });
}
// Always recurse, even when this directory does not match.
await searchDirectory(fullPath, currentDepth + 1);
} else {
totalFiles++;
if (matcher(entry.name, relativePath)) {
let size: number | undefined;
try {
size = (await fs.stat(fullPath)).size;
} catch {
// Skip size if we can't stat the match.
}
matches.push({
name: entry.name,
path: relativePath,
type: 'file',
size,
extension: entry.name.includes('.') ? entry.name.split('.').pop()?.toLowerCase() : undefined,
});
}
}
}
};
await searchDirectory(workingDir, 1);
return {
success: true,
data: {
root: workingDir,
tree: [],
matches,
totalFiles,
totalDirectories,
truncated,
matchCount: matches.length,
query: (q ?? '').trim(),
mode: 'search' as const,
},
};
}
const scanDirectory = async (dirPath: string, currentDepth: number): Promise<FileTreeNode[]> => {
if (currentDepth > maxDepth || totalFiles + totalDirectories > maxFiles) {
truncated = true;
+67 -60
View File
@@ -12,6 +12,7 @@ import { existsSync, statSync, mkdirSync, writeFileSync } from 'node:fs';
import { execFile } from 'node:child_process';
import fs from 'node:fs/promises';
import { randomBytes } from 'node:crypto';
import { performance } from 'node:perf_hooks';
import {
ApiErrorCode,
createErrorResponse,
@@ -131,6 +132,11 @@ import {
toSessionDocker,
} from '../../docker-hosts.js';
import { LRUMap } from '../../utils/lru-map.js';
import {
getLastTranscriptResponse,
isExternalCliTranscriptMode,
parseExternalCliTranscript,
} from '../response-viewer-transcript.js';
// Path to linked-cases registry (same file used by case-routes resolveCasePath)
const LINKED_CASES_FILE = dataPath('linked-cases.json');
@@ -797,54 +803,42 @@ export function registerSessionRoutes(
}
}
// Check OpenCode availability if requested
// Check OpenCode availability if requested. The error text comes from the
// resolver (formatCliNotFoundMessage) so it names where resolution looked —
// server PATH, login shell, common directories — same for the modes below.
if (body.mode === 'opencode') {
const { isOpenCodeAvailable } = await import('../../utils/opencode-cli-resolver.js');
const { isOpenCodeAvailable, getOpenCodeNotFoundMessage } = await import('../../utils/opencode-cli-resolver.js');
if (!isOpenCodeAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getOpenCodeNotFoundMessage());
}
}
// Check Codex availability if requested
if (body.mode === 'codex') {
const { isCodexAvailable } = await import('../../utils/codex-cli-resolver.js');
const { isCodexAvailable, getCodexNotFoundMessage } = await import('../../utils/codex-cli-resolver.js');
if (!isCodexAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Codex CLI not found. Install with: npm install -g @openai/codex'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getCodexNotFoundMessage());
}
}
// Check Gemini availability if requested
if (body.mode === 'gemini') {
const { isGeminiAvailable } = await import('../../utils/gemini-cli-resolver.js');
const { isGeminiAvailable, getGeminiNotFoundMessage } = await import('../../utils/gemini-cli-resolver.js');
if (!isGeminiAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Gemini CLI not found. Install with: npm install -g @google/gemini-cli'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getGeminiNotFoundMessage());
}
}
if (body.mode === 'antigravity') {
const { isAntigravityAvailable } = await import('../../utils/antigravity-cli-resolver.js');
const { isAntigravityAvailable, getAntigravityNotFoundMessage } =
await import('../../utils/antigravity-cli-resolver.js');
if (!isAntigravityAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getAntigravityNotFoundMessage());
}
}
if (body.mode === 'pi') {
const { isPiAvailable } = await import('../../utils/pi-cli-resolver.js');
const { isPiAvailable, getPiNotFoundMessage } = await import('../../utils/pi-cli-resolver.js');
if (!isPiAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Pi CLI not found. Install with: npm install -g --ignore-scripts @earendil-works/pi-coding-agent'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getPiNotFoundMessage());
}
}
@@ -1898,6 +1892,22 @@ export function registerSessionRoutes(
return await readCodexLastResponse(session, codexQuery.context === 'full');
}
// OpenCode / Gemini / Antigravity / Pi render their own TUIs and write no
// Claude transcript, so the scan below finds nothing and the response viewer
// renders permanently empty for them. Segment the terminal buffer instead —
// the pane IS the transcript for these CLIs. Codex is already handled above,
// where a real rollout file is the better source.
if (isExternalCliTranscriptMode(session.mode)) {
const externalQuery = req.query as { context?: string };
const blocks = parseExternalCliTranscript(session.terminalBuffer, session.mode);
return {
text: getLastTranscriptResponse(blocks),
timestamp: '',
hasContext: blocks.length > 0,
messages: externalQuery.context === 'full' ? blocks : undefined,
};
}
// Scan ~/.claude/projects/*/ for the transcript file
const projectsDir = join(process.env.HOME || '/tmp', '.claude', 'projects');
@@ -2270,18 +2280,17 @@ export function registerSessionRoutes(
// Query params:
// tail=<bytes> - Only return last N bytes (faster initial load)
// full=1 - Full page reload: replay the entire tmux scrollback (COD-47)
app.get('/api/sessions/:id/terminal', async (req) => {
// full=1 - Explicitly request the entire tmux scrollback (COD-47)
app.get('/api/sessions/:id/terminal', async (req, reply) => {
const routeStartedAt = performance.now();
const { id } = req.params as { id: string };
const query = req.query as { tail?: string; full?: string };
const session = findSessionOrFail(ctx, id, req);
// `full=1` is the EXPLICIT full-reload signal (COD-47): the browser reloaded
// the page and wants the whole scroll history back, so we capture the ENTIRE
// tmux scrollback and the user gets back history that scrolled off Codeman's
// byte buffer. Requests WITHOUT it — tab switches (`tail=`) and the legacy
// no-param callers (response-viewer fallback, clearTerminal refresh) — keep
// the fast visible-frame capture.
// `full=1` is the EXPLICIT full-history signal (COD-47): capture the ENTIRE
// tmux scrollback so history beyond the server byte buffer can be recovered.
// Requests WITHOUT it — shell selection/tab switches (`tail=`) and legacy
// no-param callers — keep the fast visible-frame capture.
const tailBytes = query.tail ? parseInt(query.tail, 10) : 0;
const isFullReload = query.full === '1' || query.full === 'true';
const { tmuxHistoryLimit, terminalBufferMaxBytes } = await ctx.getTerminalHistoryConfig();
@@ -2294,6 +2303,7 @@ export function registerSessionRoutes(
// overlap. `captureActivePaneBuffer` is a no-op ('') under test mode and
// returns null when unavailable, in which case we fall back to history.
const muxName = session.muxName;
const captureStartedAt = performance.now();
const liveMuxBuffer =
muxName && typeof ctx.mux.captureActivePaneBuffer === 'function'
? ctx.mux.captureActivePaneBuffer(
@@ -2303,6 +2313,7 @@ export function registerSessionRoutes(
: undefined
)
: null;
const captureFinishedAt = performance.now();
const hasLiveMuxBuffer = liveMuxBuffer !== null && liveMuxBuffer.length > 0;
const source: 'history' | 'mux-visible' | 'mux-full-history' = hasLiveMuxBuffer
? isFullReload
@@ -2406,6 +2417,14 @@ export function registerSessionRoutes(
// Remove Ctrl+L and leading whitespace (cheap on tailed subset)
cleanBuffer = cleanBuffer.replace(CTRL_L_PATTERN, '').replace(LEADING_WHITESPACE_PATTERN, '');
const finishedAt = performance.now();
reply.header(
'Server-Timing',
`capture;dur=${(captureFinishedAt - captureStartedAt).toFixed(1)}, ` +
`prepare;dur=${(finishedAt - captureFinishedAt).toFixed(1)}, ` +
`total;dur=${(finishedAt - routeStartedAt).toFixed(1)}`
);
return {
terminalBuffer: cleanBuffer,
status: session.status,
@@ -2806,58 +2825,46 @@ export function registerSessionRoutes(
dockerResumeId = dockerCase.lastClaudeSessionId;
}
} else {
// Check OpenCode availability if requested
// Check OpenCode availability if requested. Error text comes from the
// resolver so it carries the resolution diagnostics; same for the modes below.
if (mode === 'opencode') {
const { isOpenCodeAvailable } = await import('../../utils/opencode-cli-resolver.js');
const { isOpenCodeAvailable, getOpenCodeNotFoundMessage } =
await import('../../utils/opencode-cli-resolver.js');
if (!isOpenCodeAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getOpenCodeNotFoundMessage());
}
}
// Check Codex availability if requested
if (mode === 'codex') {
const { isCodexAvailable } = await import('../../utils/codex-cli-resolver.js');
const { isCodexAvailable, getCodexNotFoundMessage } = await import('../../utils/codex-cli-resolver.js');
if (!isCodexAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Codex CLI not found. Install with: npm install -g @openai/codex'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getCodexNotFoundMessage());
}
}
// Check Gemini availability if requested
if (mode === 'gemini') {
const { isGeminiAvailable } = await import('../../utils/gemini-cli-resolver.js');
const { isGeminiAvailable, getGeminiNotFoundMessage } = await import('../../utils/gemini-cli-resolver.js');
if (!isGeminiAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Gemini CLI not found. Install with: npm install -g @google/gemini-cli'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getGeminiNotFoundMessage());
}
}
// Check Antigravity availability if requested
if (mode === 'antigravity') {
const { isAntigravityAvailable } = await import('../../utils/antigravity-cli-resolver.js');
const { isAntigravityAvailable, getAntigravityNotFoundMessage } =
await import('../../utils/antigravity-cli-resolver.js');
if (!isAntigravityAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Antigravity CLI not found. Install with: curl -fsSL https://antigravity.google/cli/install.sh | bash'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getAntigravityNotFoundMessage());
}
}
// Check Pi availability if requested
if (mode === 'pi') {
const { isPiAvailable } = await import('../../utils/pi-cli-resolver.js');
const { isPiAvailable, getPiNotFoundMessage } = await import('../../utils/pi-cli-resolver.js');
if (!isPiAvailable()) {
return createErrorResponse(
ApiErrorCode.OPERATION_FAILED,
'Pi CLI not found. Install with: npm install -g --ignore-scripts @earendil-works/pi-coding-agent'
);
return createErrorResponse(ApiErrorCode.OPERATION_FAILED, getPiNotFoundMessage());
}
}
+8 -1
View File
@@ -48,6 +48,7 @@ import {
} from '../route-helpers.js';
import { SseEvent } from '../sse-events.js';
import { getInstallInfo, checkForUpdate, startUpdate, getUpdateStatusForApi } from '../self-update.js';
import { getRepositoryStatus } from '../repo-status.js';
import type { SessionPort, EventPort, ConfigPort, InfraPort, AuthPort } from '../ports/index.js';
import { AUTH_COOKIE_NAME } from '../middleware/auth.js';
import { QR_AUTH_FAILURE_MAX } from '../../config/tunnel-config.js';
@@ -354,6 +355,11 @@ export function registerSystemRoutes(
// Poll target for update progress — survives the restart the update triggers.
app.get('/api/system/update/status', async () => getUpdateStatusForApi());
// Informational companion to the release-tag updater above: what this CHECKOUT
// looks like against its own remotes (ahead/behind, incoming commits), which a
// release tag cannot answer for a git install tracking a branch.
app.get('/api/system/repo-status', async () => getRepositoryStatus());
// Kick off a detached update to the latest release. Returns immediately; the
// browser then polls /api/system/update/status across the service restart.
app.post('/api/system/update', async (_req, reply) => {
@@ -722,7 +728,8 @@ export function registerSystemRoutes(
const merged = { ...existing, ...settingsToStore };
await fs.writeFile(SETTINGS_PATH, JSON.stringify(merged, null, 2));
// Apply a changed tmux history-limit to all live sessions immediately.
// tmux 3.7+ resizes tracked panes; older versions apply this to new panes.
// Already-evicted history cannot be recovered on either version.
if (settings.tmuxHistoryLimit !== undefined) {
await ctx.mux.setHistoryLimit(resolveTerminalHistoryConfig(merged).tmuxHistoryLimit);
}
+11 -2
View File
@@ -956,8 +956,17 @@ export const SettingsUpdateSchema = z
// CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK env var. Stripped before persisting.
acknowledgeUnauthTunnel: z.boolean().optional(),
tabTwoRows: z.boolean().optional(),
/** Session list layout: 'header' = horizontal tab strip, 'sidebar' = collapsible left sidebar. Display key (per-device). */
sessionListLayout: z.enum(['header', 'sidebar']).optional(),
/**
* Session list layout. Display key (per-device).
* 'header' = horizontal tab strip
* 'sidebar' = collapsible left sidebar, one compact row per session
* 'sidebar-rich' = same sidebar, each row carrying the home screen's detail
* (created/idle/working stamps + status pill)
* Both sidebar values render the SAME docked column and set
* data-session-list="sidebar"; they differ only in row detail, which rides
* on data-sidebar-detail. See applySessionListLayout() in app.js.
*/
sessionListLayout: z.enum(['header', 'sidebar', 'sidebar-rich']).optional(),
agentTeamsEnabled: z.boolean().optional(),
/** Model for new Claude sessions (e.g. "claude-fable-5[1m]", "opus[1m]"); takes precedence over opusContext1mEnabled */
claudeModel: z.string().max(50).optional(),
+16 -4
View File
@@ -781,19 +781,31 @@ export class WebServer extends EventEmitter {
// Serve static files — content-hashed assets (e.g. app.a3f8c2e1.js) are immutable, cache aggressively.
// HTML must revalidate every time so browsers pick up new hashed filenames after deploys.
// cacheControl disabled so setHeaders has full control (fastify-static's reply.headers() overwrites setHeaders otherwise).
// cacheControl disabled so setHeaders owns Cache-Control for plain static assets.
// preCompressed: serve pre-built .br/.gz files (from build step) to avoid per-request CPU compression
await this.app.register(fastifyStatic, {
root: join(__dirname, 'public'),
prefix: '/',
cacheControl: false,
preCompressed: true,
setHeaders: (res, path) => {
// ⚠️ @fastify/static v10 changed this callback's first argument from a Node
// `ServerResponse` to a `FastifyReply`, so it is `reply.header()` here and
// NOT `res.setHeader()`. A v9-style body throws TypeError on every static
// request, which is every page load. See the v10.0.0 release notes.
setHeaders: (reply, path) => {
// ⚠️ That same change ALSO flipped precedence, and silently. Under v9 this
// callback wrote to the raw response and Fastify's staged reply headers then
// overwrote it, so a route that set its own Cache-Control before .sendFile()
// won. Under v10 the callback writes to the reply itself and now wins instead,
// which handed `/sw.js` a year of `immutable` in place of the `no-cache,
// no-store` its route asks for — a service worker that can never update.
// So: a route that already decided keeps its answer.
if (reply.getHeader('Cache-Control') !== undefined) return;
// Use .includes() not .endsWith() — preCompressed serves .html.br/.html.gz
if (path.includes('.html')) {
res.setHeader('Cache-Control', 'no-cache');
reply.header('Cache-Control', 'no-cache');
} else {
res.setHeader('Cache-Control', 'public, max-age=31536000, immutable');
reply.header('Cache-Control', 'public, max-age=31536000, immutable');
}
},
});
+28 -3
View File
@@ -52,6 +52,7 @@ export interface SessionListenerRefs {
limitResumeCancelled: (data: { reason: string }) => void;
respawnBreakerTripped: (data: { count: number }) => void;
cliInfoUpdated: (data: { version?: string; model?: string; accountType?: string; latestVersion?: string }) => void;
mouseTrackingChanged: (active: boolean) => void;
ralphLoopUpdate: (state: RalphTrackerState) => void;
ralphTodoUpdate: (todos: RalphTodoItem[]) => void;
ralphCompletionDetected: (phrase: string) => void;
@@ -219,10 +220,18 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
// An idle-prompt inbox item means "composer is waiting"; any working
// transition means input arrived, so the item is moot. ONLY the idle
// kind: `working` is heuristic and can flap mid-turn, so clearing a
// pending permission/question dialog on it would false-clear real
// approvals (those resolve via stop / elicitation hooks / answer-time
// re-capture instead).
// pending permission/question dialog on the signal ALONE would
// false-clear real approvals.
approvalInbox.resolveForSession(session.id, 'resolved_in_terminal', ['idle']);
// A permission/question dialog gets the pane-VERIFIED variant instead:
// the signal only decides when to look, `verifyStillAnswerable` re-reads
// the screen and resolves only when the dialog is really gone. Without
// this, answering a dialog in the terminal left its red "needs you" alert
// armed for the rest of the turn, because the only other staleness check
// lives in `GET /api/approvals` and nothing calls that while a page is
// open. `stop` was the first thing to clear it, which on a long turn is
// minutes away.
approvalInbox.resolveIfDialogGone(session.id);
deps.broadcast(SseEvent.SessionWorking, { id: session.id });
// Full state ride-along: the home screens sort the running group on
// lastSubmitAt, and without this the browser keeps the stamp it loaded
@@ -342,6 +351,20 @@ export function createSessionListeners(session: Session, deps: SessionListenerDe
deps.broadcastSessionStateDebounced(session.id);
},
/**
* The CLI turned mouse tracking on or off (observed while stripping the
* DECSETs out of the stream). Rides the full session state so the browser
* learns it through the session object it already merges, with no new SSE
* event to keep in sync across the two registries.
*
* Broadcast IMMEDIATELY, not debounced: this flips when a dialog opens, and
* a user can click that dialog inside the 500ms debounce window, which is
* exactly the click that has to be reported.
*/
mouseTrackingChanged: () => {
deps.broadcast(SseEvent.SessionUpdated, { session: deps.getSessionStateWithRespawn(session) });
},
// ─── Ralph Tracking Events ──────────────────────────────
/** Broadcasts `session:ralphLoopUpdate` — Ralph tracker loop state changed (iteration, phase) */
@@ -453,6 +476,7 @@ export function attachSessionListeners(session: Session, refs: SessionListenerRe
session.on('limitResumeCancelled', refs.limitResumeCancelled);
session.on('respawnBreakerTripped', refs.respawnBreakerTripped);
session.on('cliInfoUpdated', refs.cliInfoUpdated);
session.on('mouseTrackingChanged', refs.mouseTrackingChanged);
session.on('ralphLoopUpdate', refs.ralphLoopUpdate);
session.on('ralphTodoUpdate', refs.ralphTodoUpdate);
session.on('ralphCompletionDetected', refs.ralphCompletionDetected);
@@ -487,6 +511,7 @@ export function detachSessionListeners(session: Session, refs: SessionListenerRe
session.off('limitResumeCancelled', refs.limitResumeCancelled);
session.off('respawnBreakerTripped', refs.respawnBreakerTripped);
session.off('cliInfoUpdated', refs.cliInfoUpdated);
session.off('mouseTrackingChanged', refs.mouseTrackingChanged);
session.off('ralphLoopUpdate', refs.ralphLoopUpdate);
session.off('ralphTodoUpdate', refs.ralphTodoUpdate);
session.off('ralphCompletionDetected', refs.ralphCompletionDetected);
+121
View File
@@ -0,0 +1,121 @@
/**
* @fileoverview Tests for the Antigravity CLI resolver wrapper.
*/
import { homedir } from 'node:os';
import { join } from 'node:path';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { createAntigravityResolverForTest, isAntigravityAvailable } from '../src/utils/antigravity-cli-resolver.js';
import {
cliResolveRetryDelayMs,
type CliResolution,
type CliResolverHost,
} from '../src/utils/cli-executable-resolver.js';
const availabilityResolution = vi.hoisted(() => ({ current: null as CliResolution | null }));
vi.mock('../src/utils/cli-executable-resolver.js', async (importOriginal) => {
const actual = await importOriginal<typeof import('../src/utils/cli-executable-resolver.js')>();
return {
...actual,
createCliExecutableResolver: (options: { binary: string; searchDirs: string[] }, host?: CliResolverHost) =>
host
? actual.createCliExecutableResolver(options, host)
: {
resolve: () => availabilityResolution.current,
diagnostics: () => ({
binary: options.binary,
processPath: '/service/bin',
shellPath: '/bin/zsh',
shellArgs: ['-l'],
searchDirs: [...options.searchDirs],
}),
},
};
});
function createHost(
options: {
processPathResult?: string | null;
loginShellResults?: Array<string | null>;
existingPaths?: string[];
} = {}
): CliResolverHost {
const loginShellResults = [...(options.loginShellResults ?? [])];
const existingPaths = new Set(options.existingPaths ?? []);
return {
processPath: '/service/bin',
shellPath: '/bin/zsh',
shellArgs: ['-l'],
findOnProcessPath: () => options.processPathResult ?? null,
findInLoginShell: () => loginShellResults.shift() ?? null,
exists: (path) => existingPaths.has(path),
};
}
describe('Antigravity CLI resolver', () => {
beforeEach(() => {
availabilityResolution.current = null;
});
it('resolves agy from the service PATH', () => {
const binaryPath = '/service/bin/agy';
const resolver = createAntigravityResolverForTest(
createHost({ processPathResult: binaryPath, existingPaths: [binaryPath] })
);
expect(resolver.resolve()?.directory).toBe('/service/bin');
});
it('falls back to a common install directory', () => {
const binaryPath = join(homedir(), '.local', 'bin', 'agy');
const resolver = createAntigravityResolverForTest(createHost({ existingPaths: [binaryPath] }));
expect(resolver.resolve()?.directory).toBe(join(homedir(), '.local', 'bin'));
});
it('resolves agy found only by the login shell', () => {
const binaryPath = '/login-shell/bin/agy';
const resolver = createAntigravityResolverForTest(
createHost({ loginShellResults: [binaryPath], existingPaths: [binaryPath] })
);
expect(resolver.resolve()?.directory).toBe('/login-shell/bin');
});
it('returns null when agy is unavailable', () => {
const resolver = createAntigravityResolverForTest(createHost());
expect(resolver.resolve()).toBeNull();
});
it('retries a failed lookup after the backoff and caches the first successful login-shell discovery', () => {
const binaryPath = '/late-login-shell/bin/agy';
let now = 0;
const resolver = createAntigravityResolverForTest(
createHost({ loginShellResults: [null, binaryPath], existingPaths: [binaryPath] }),
() => now
);
expect(resolver.resolve()).toBeNull();
// A miss is negative-cached: within the backoff window nothing re-runs the
// chain (its login-shell tail is a synchronous bounded spawn in production).
expect(resolver.resolve()).toBeNull();
now = cliResolveRetryDelayMs(1);
expect(resolver.resolve()?.binaryPath).toBe(binaryPath);
expect(resolver.resolve()?.binaryPath).toBe(binaryPath);
});
it('reports the public wrapper as available when agy resolves', () => {
availabilityResolution.current = {
binaryPath: '/service/bin/agy',
directory: '/service/bin',
source: 'process-path',
};
expect(isAntigravityAvailable()).toBe(true);
});
it('reports the public wrapper as unavailable when agy does not resolve', () => {
expect(isAntigravityAvailable()).toBe(false);
});
});
+213
View File
@@ -40,6 +40,39 @@ const ASK_USER_QUESTION_FRAME = [
'Enter to select · ↑/↓ to navigate · Esc to cancel',
].join('\n');
// Both frames captured off a live Claude Code v2.1.237 pane. They are the
// discriminator behind the late-hook resolution: a modal dialog BLOCKS the
// turn, so a working line and a dialog cannot coexist. Note the dialog frame
// carries no `esc to interrupt` footer either — the dialog replaces it.
const LIVE_DIALOG_FRAME = [
'● Bash(sleep 12; echo "slept 12s")',
' ⎿ slept 12s',
'────────────────────────────────────────',
' ☐ Proceed',
'',
'Proceed?',
'',
'❯ 1. Yes',
' Go ahead and proceed.',
' 2. No',
' Do not proceed.',
' 3. Type something.',
'────────────────────────────────────────',
' 4. Chat about this',
'',
'Enter to select · ↑/↓ to navigate · Esc to cancel',
].join('\n');
const WORKING_FRAME = [
'● Bash(sleep 25)',
' ⎿ Tip: Use git worktrees to run multiple Claude sessions in parallel.',
'✢ Clauding… (13s · ↓ 1.4k tokens)',
'────────────────────────────────────────',
'❯',
'────────────────────────────────────────',
' ⏵⏵ bypass permissions on (shift+tab to cycle) · esc to interrupt · ← for agents',
].join('\n');
function collect(inbox: ApprovalInbox) {
const pending: ApprovalItem[] = [];
const updated: ApprovalItem[] = [];
@@ -333,6 +366,186 @@ describe('ApprovalInbox', () => {
});
});
describe('answered-in-the-terminal staleness', () => {
// Claude Code delays the Notification hook behind the dialog (measured 6s
// on v2.1.237, documented up to ~30s), so the 600ms re-capture routinely
// lands on a frame the user has ALREADY answered. Erasing `options` there
// made the item permanently unsweepable, because a missing `options` is how
// "we never could read this dialog" is expressed, and such items stay
// answerable on purpose. The red tab alert then survived every
// `GET /api/approvals` and every reload, clearing only on `stop`.
it('a re-capture taken after the answer does not erase parsed options', () => {
const { updated } = collect(inbox);
let frame = ASK_USER_QUESTION_FRAME;
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => frame,
});
expect(item.options).toHaveLength(5);
// Answered in the terminal before the re-capture fires.
frame = "● User answered Claude's questions:\n ⎿ · Which color do you prefer? → Red\n\n✶ Cooking… (6s)";
vi.advanceTimersByTime(600);
expect(updated).toHaveLength(1);
expect(inbox.getById(item.id)?.context).toContain('User answered');
expect(inbox.getById(item.id)?.options).toHaveLength(5);
// ...which keeps the staleness check conclusive instead of inconclusive.
expect(inbox.verifyStillAnswerable(item.id)).toBe(false);
expect(inbox.getById(item.id)).toBeUndefined();
});
it('resolveIfDialogGone resolves a dialog answered in the terminal', () => {
const { resolved } = collect(inbox);
let frame = PERMISSION_FRAME;
const item = inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'permission', capture: () => frame });
frame = 'the dialog is gone, claude is typing';
inbox.resolveIfDialogGone('s1');
expect(inbox.getById(item.id)).toBeUndefined();
expect(resolved.at(-1)).toMatchObject({ id: item.id, resolution: 'resolved_in_terminal' });
});
it('resolveIfDialogGone keeps a dialog that is still on screen', () => {
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => PERMISSION_FRAME,
});
inbox.resolveIfDialogGone('s1');
expect(inbox.getById(item.id)).toBeDefined();
});
it('resolveIfDialogGone never touches an idle item (that is the working signal job)', () => {
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'idle',
capture: () => 'composer is empty',
});
inbox.resolveIfDialogGone('s1');
expect(inbox.getById(item.id)).toBeDefined();
});
// The late-hook hole: Claude Code fires the Notification behind the dialog,
// so a prompt answered before the hook lands produces an item whose FIRST
// capture already has no dialog in it. Nothing ever parsed, so "options
// vanished" can never fire, and `stop` had already gone by too — the red
// alert then survived reloads until the 12h TTL.
it('resolves an item that never parsed options once the pane is visibly working', () => {
const { resolved } = collect(inbox);
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => WORKING_FRAME,
});
expect(item.options).toBeUndefined();
expect(inbox.verifyStillAnswerable(item.id)).toBe(false);
expect(inbox.getById(item.id)).toBeUndefined();
expect(resolved.at(-1)).toMatchObject({ id: item.id, resolution: 'resolved_in_terminal' });
});
it('keeps an unreadable dialog answerable when the pane is NOT visibly working', () => {
// The conservative rule this fix must not loosen: no options and no proof
// the turn is running means "we cannot read it", not "it is gone".
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => 'some dialog shape we cannot parse',
});
expect(item.options).toBeUndefined();
expect(inbox.verifyStillAnswerable(item.id)).toBe(true);
expect(inbox.getById(item.id)).toBeDefined();
});
it('a real live-dialog frame carries no working line, so it is never false-resolved', () => {
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => LIVE_DIALOG_FRAME,
});
expect(item.options).toHaveLength(4);
expect(inbox.verifyStillAnswerable(item.id)).toBe(true);
expect(inbox.getById(item.id)).toBeDefined();
});
it('resolveIfDialogGone clears a late-hook item on the next working signal', () => {
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => WORKING_FRAME,
});
inbox.resolveIfDialogGone('s1');
expect(inbox.getById(item.id)).toBeUndefined();
});
it('the delayed staleness pass resolves a late-hook item with no working signal needed', () => {
const { resolved } = collect(inbox);
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => WORKING_FRAME,
});
expect(item.options).toBeUndefined();
vi.advanceTimersByTime(600); // re-capture: enrichment only, never resolves
expect(inbox.getById(item.id)).toBeDefined();
vi.advanceTimersByTime(2400); // the delayed staleness pass
expect(inbox.getById(item.id)).toBeUndefined();
expect(resolved.at(-1)).toMatchObject({ id: item.id, resolution: 'resolved_in_terminal' });
});
it('a dialog Ink paints late is NOT resolved by the delayed pass', () => {
// The paint race the re-capture exists for: the hook can beat Ink to the
// screen. Resolving inside that window would clear the alert for a dialog
// that was about to appear, so the frame is what decides, every time.
let frame = WORKING_FRAME;
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => frame,
});
frame = LIVE_DIALOG_FRAME; // Ink finishes painting
vi.advanceTimersByTime(600);
expect(inbox.getById(item.id)?.options).toHaveLength(4);
vi.advanceTimersByTime(2400);
expect(inbox.getById(item.id)).toBeDefined();
});
it('resolving cancels both pending timers', () => {
const { updated } = collect(inbox);
const item = inbox.notePrompt({
sessionId: 's1',
sessionName: 'w1',
kind: 'permission',
capture: () => PERMISSION_FRAME,
});
inbox.dismiss(item.id);
vi.advanceTimersByTime(5000);
expect(updated).toHaveLength(0);
expect(inbox.listPending()).toHaveLength(0);
});
it('resolveIfDialogGone is a no-op for a session with nothing pending', () => {
const { resolved } = collect(inbox);
expect(() => inbox.resolveIfDialogGone('nobody')).not.toThrow();
expect(resolved).toHaveLength(0);
});
});
it('stop() clears items and silences events', () => {
const { resolved } = collect(inbox);
inbox.notePrompt({ sessionId: 's1', sessionName: 'w1', kind: 'permission' });
+91
View File
@@ -147,3 +147,94 @@ describe('tmux-backed shell: strip tmux’s own client smcup, keep everything el
}
});
});
/**
* Whatever the strip removes, the server has to remember, because after it runs
* nothing downstream can ever see it. The browser hand-encodes click reports for
* these modes (`_sendSyntheticSgrTap`), and with no state to consult it did that
* on EVERY click, delivering mouse reports to a CLI that never asked for them.
*/
describe('stripped mouse-tracking state', () => {
const trackingOf = (session: Session) => session.toState().cliMouseTracking;
it('starts off, and stays off for output that never enables tracking', () => {
const session = new Session({ workingDir: '/tmp', mode: 'claude' });
expect(trackingOf(session)).toBeUndefined();
handleOutput(session, 'plain output\x1b[?1049h\x1b[3J');
expect(trackingOf(session)).toBeUndefined();
});
it('follows the CLI enabling and disabling a tracking mode', () => {
const session = new Session({ workingDir: '/tmp', mode: 'claude' });
const changes: boolean[] = [];
session.on('mouseTrackingChanged', (active: boolean) => changes.push(active));
handleOutput(session, '\x1b[?1002hdialog');
expect(trackingOf(session)).toBe(true);
handleOutput(session, '\x1b[?1002ldismissed');
expect(trackingOf(session)).toBeUndefined();
expect(changes).toEqual([true, false]);
});
it('ignores encoding and alt-scroll modes, which do not ask about clicks', () => {
// 1005/1006 pick an ENCODING and 1007 is alt-scroll. Counting them would put
// the stray reports straight back: a CLI can select SGR encoding without ever
// asking to be told where the user clicked.
const session = new Session({ workingDir: '/tmp', mode: 'claude' });
handleOutput(session, '\x1b[?1006h\x1b[?1005h\x1b[?1007h');
expect(trackingOf(session)).toBeUndefined();
});
it('stays on until the LAST tracking mode goes away', () => {
const session = new Session({ workingDir: '/tmp', mode: 'claude' });
handleOutput(session, '\x1b[?1000h\x1b[?1002h\x1b[?1006h');
expect(trackingOf(session)).toBe(true);
// A TUI may disable a mode it never enabled; that must not clear the rest.
handleOutput(session, '\x1b[?1003l');
expect(trackingOf(session)).toBe(true);
handleOutput(session, '\x1b[?1000l');
expect(trackingOf(session)).toBe(true);
handleOutput(session, '\x1b[?1002l');
expect(trackingOf(session)).toBeUndefined();
});
it('emits only on a real transition, so a repainting TUI costs nothing', () => {
const session = new Session({ workingDir: '/tmp', mode: 'claude' });
const changes: boolean[] = [];
session.on('mouseTrackingChanged', (active: boolean) => changes.push(active));
handleOutput(session, '\x1b[?1002h\x1b[?1002h\x1b[?1002h');
expect(changes).toEqual([true]);
});
it('sees a sequence split across PTY chunks, like the strip that carries it', () => {
const session = new Session({ workingDir: '/tmp', mode: 'claude' });
handleOutput(session, 'before\x1b[?100');
handleOutput(session, '2h after');
expect(session.terminalBuffer).toBe('before after');
expect(trackingOf(session)).toBe(true);
});
it('tracks nothing for a mode whose DECSETs are never stripped', () => {
// shell keeps its mouse DECSETs, so xterm sees them and owns the reporting.
// A flag set here would mean a SECOND, hand-encoded report on every click.
const session = new Session({ workingDir: '/tmp', mode: 'shell', useMux: true });
handleOutput(session, '\x1b[?1002hhtop');
expect(trackingOf(session)).toBeUndefined();
expect(session.terminalBuffer).toBe('\x1b[?1002hhtop');
});
});
+400
View File
@@ -0,0 +1,400 @@
import { chmodSync, mkdirSync, mkdtempSync, rmSync, writeFileSync } from 'node:fs';
import { tmpdir } from 'node:os';
import { join } from 'node:path';
import { afterEach, describe, expect, it, vi } from 'vitest';
import { EXEC_TIMEOUT_MS } from '../src/config/exec-timeout.js';
import {
cliResolveRetryDelayMs,
createCliExecutableResolver,
createProductionCliResolverHost,
formatCliNotFoundMessage,
type CliResolverHost,
} from '../src/utils/cli-executable-resolver.js';
// Pass-through spy on execFileSync so the vitest-hermeticity test below can
// PROVE the un-injected production host never spawns anything.
const { execFileSyncSpy } = vi.hoisted(() => ({ execFileSyncSpy: vi.fn() }));
vi.mock('node:child_process', async (importOriginal) => {
const actual = await importOriginal<typeof import('node:child_process')>();
execFileSyncSpy.mockImplementation(actual.execFileSync as (...args: unknown[]) => unknown);
return { ...actual, execFileSync: execFileSyncSpy };
});
const BEGIN_MARKER = '__CODEMAN_CLI_RESOLVE_BEGIN__';
const END_MARKER = '__CODEMAN_CLI_RESOLVE_END__';
const temporaryDirectories: string[] = [];
afterEach(() => {
for (const directory of temporaryDirectories.splice(0)) {
rmSync(directory, { recursive: true, force: true });
}
});
function host(overrides: Partial<CliResolverHost> = {}): CliResolverHost {
return {
processPath: '/usr/bin:/bin',
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
findOnProcessPath: vi.fn(() => null),
findInLoginShell: vi.fn(() => null),
exists: vi.fn(() => false),
...overrides,
};
}
describe('createCliExecutableResolver', () => {
it('prefers the server process PATH over common directories and the login shell', () => {
const h = host({
findOnProcessPath: vi.fn(() => '/process/bin/codex'),
findInLoginShell: vi.fn(() => '/shell/bin/codex'),
exists: vi.fn(() => true),
});
const resolver = createCliExecutableResolver({ binary: 'codex', searchDirs: ['/known/bin'] }, h);
expect(resolver.resolve()).toMatchObject({ binaryPath: '/process/bin/codex', source: 'process-path' });
expect(h.exists).toHaveBeenCalledTimes(1);
expect(h.findInLoginShell).not.toHaveBeenCalled();
});
it('prefers common directories in order over the login shell', () => {
const h = host({
findInLoginShell: vi.fn(() => '/shell/bin/codex'),
exists: vi.fn((path) => path === '/second/bin/codex' || path === '/shell/bin/codex'),
});
const resolver = createCliExecutableResolver({ binary: 'codex', searchDirs: ['/first/bin', '/second/bin'] }, h);
expect(resolver.resolve()).toMatchObject({ binaryPath: '/second/bin/codex', source: 'common-directory' });
expect(h.exists).toHaveBeenNthCalledWith(1, '/first/bin/codex');
expect(h.exists).toHaveBeenNthCalledWith(2, '/second/bin/codex');
expect(h.findInLoginShell).not.toHaveBeenCalled();
});
it('finds an executable exposed only by the interactive login shell', () => {
const h = host({
findInLoginShell: vi.fn(() => '/home/u/.nvm/versions/node/v22/bin/codex'),
exists: vi.fn((path) => path === '/home/u/.nvm/versions/node/v22/bin/codex'),
});
const resolver = createCliExecutableResolver({ binary: 'codex', searchDirs: ['/known/bin'] }, h);
expect(resolver.resolve()).toMatchObject({
binaryPath: '/home/u/.nvm/versions/node/v22/bin/codex',
directory: '/home/u/.nvm/versions/node/v22/bin',
source: 'login-shell',
});
});
it('continues after a validator rejects an earlier candidate', () => {
const h = host({
findOnProcessPath: vi.fn(() => '/usr/bin/pi'),
findInLoginShell: vi.fn(() => '/home/u/.npm/bin/pi'),
exists: vi.fn(() => true),
});
const resolver = createCliExecutableResolver(
{
binary: 'pi',
searchDirs: [],
validateCandidate: (path) =>
path.includes('.npm') ? { accepted: true, metadata: '0.84.1' } : { accepted: false },
},
h
);
expect(resolver.resolve()).toMatchObject({
binaryPath: '/home/u/.npm/bin/pi',
source: 'login-shell',
metadata: '0.84.1',
});
});
it('caches success, and retries a miss only after the backoff elapses', () => {
let now = 0;
const findInLoginShell = vi.fn<() => string | null>().mockReturnValueOnce(null).mockReturnValue('/new/bin/codex');
const h = host({ findInLoginShell, exists: vi.fn((path) => path === '/new/bin/codex') });
const resolver = createCliExecutableResolver({ binary: 'codex', searchDirs: [], now: () => now }, h);
expect(resolver.resolve()).toBeNull();
// Within the backoff window the miss is answered from the negative cache:
// the chain — whose login-shell tail is a synchronous 5s-bounded spawn —
// must NOT re-run per call, or a missing CLI stalls every status request.
now = cliResolveRetryDelayMs(1) - 1;
expect(resolver.resolve()).toBeNull();
expect(findInLoginShell).toHaveBeenCalledTimes(1);
// Once the backoff elapses the retry runs, so installing a CLI while the
// server is up is still picked up without a restart.
now = cliResolveRetryDelayMs(1);
expect(resolver.resolve()?.binaryPath).toBe('/new/bin/codex');
expect(resolver.resolve()?.binaryPath).toBe('/new/bin/codex');
expect(findInLoginShell).toHaveBeenCalledTimes(2);
});
it('doubles the retry delay per consecutive miss and caps it at five minutes', () => {
let now = 0;
const findInLoginShell = vi.fn(() => null);
const resolver = createCliExecutableResolver(
{ binary: 'codex', searchDirs: [], now: () => now },
host({ findInLoginShell })
);
expect(cliResolveRetryDelayMs(0)).toBe(0);
expect(cliResolveRetryDelayMs(1)).toBe(60_000);
expect(cliResolveRetryDelayMs(2)).toBe(120_000);
expect(cliResolveRetryDelayMs(3)).toBe(240_000);
expect(cliResolveRetryDelayMs(4)).toBe(300_000);
expect(cliResolveRetryDelayMs(60)).toBe(300_000);
// Consecutive misses stack: after the second miss the SECOND delay applies.
expect(resolver.resolve()).toBeNull();
now += cliResolveRetryDelayMs(1);
expect(resolver.resolve()).toBeNull();
expect(findInLoginShell).toHaveBeenCalledTimes(2);
now += cliResolveRetryDelayMs(2) - 1;
expect(resolver.resolve()).toBeNull();
expect(findInLoginShell).toHaveBeenCalledTimes(2);
now += 1;
expect(resolver.resolve()).toBeNull();
expect(findInLoginShell).toHaveBeenCalledTimes(3);
});
it('rejects unsafe binary names', () => {
const h = host();
expect(() => createCliExecutableResolver({ binary: 'codex;id', searchDirs: [] }, h)).toThrow(
'Unsafe CLI binary name'
);
expect(() => createCliExecutableResolver({ binary: '../codex', searchDirs: [] }, h)).toThrow(
'Unsafe CLI binary name'
);
});
it('rejects relative and nonexistent candidates', () => {
let now = 0;
const findInLoginShell = vi
.fn<() => string | null>()
.mockReturnValueOnce('relative/codex')
.mockReturnValue('/missing/codex');
const h = host({ findInLoginShell, exists: vi.fn(() => false) });
const resolver = createCliExecutableResolver({ binary: 'codex', searchDirs: [], now: () => now }, h);
expect(resolver.resolve()).toBeNull();
now = cliResolveRetryDelayMs(1);
expect(resolver.resolve()).toBeNull();
expect(h.exists).toHaveBeenCalledTimes(1);
expect(h.exists).toHaveBeenCalledWith('/missing/codex');
});
});
describe('formatCliNotFoundMessage', () => {
it('includes only the base install hint and bounded resolution diagnostics', () => {
const base = 'Codex CLI not found. Install with: npm install -g @openai/codex';
const diagnostics = {
binary: 'codex',
processPath: '/usr/bin:/bin',
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
searchDirs: ['/home/u/.local/bin', '/usr/local/bin'],
API_KEY: 'super-secret',
};
const message = formatCliNotFoundMessage(base, diagnostics);
expect(message).toContain(base);
expect(message).toContain('Server PATH: /usr/bin:/bin');
expect(message).toContain('Login shell: /bin/bash -i -l');
expect(message).toContain('Checked directories: /home/u/.local/bin, /usr/local/bin');
expect(message).not.toContain('API_KEY');
expect(message).not.toContain('super-secret');
});
it('marks empty diagnostic values without dumping arbitrary environment data', () => {
const message = formatCliNotFoundMessage('Missing CLI', {
binary: 'codex',
processPath: '',
shellPath: '',
shellArgs: [],
searchDirs: [],
});
expect(message).toBe('Missing CLI\nServer PATH: (empty)\nLogin shell: (none)\nChecked directories: (none)');
expect(message).not.toContain('HOME=');
expect(message).not.toContain('TOKEN=');
});
it('flattens control characters and bounds every diagnostic field', () => {
const pathological = `first\r\nforged label: value\u0000${'x'.repeat(10_000)}`;
const message = formatCliNotFoundMessage('Missing CLI', {
binary: 'codex',
processPath: pathological,
shellPath: pathological,
shellArgs: [pathological],
searchDirs: [pathological, pathological],
});
const lines = message.split('\n');
expect(lines).toHaveLength(4);
expect(lines[1]).toMatch(/^Server PATH: first forged label: value x+…$/);
expect(lines[2]).toMatch(/^Login shell: first forged label: value x+…$/);
expect(lines[3]).toMatch(/^Checked directories: first forged label: value x+…$/);
expect(lines.slice(1).every((line) => line.length <= 1_050)).toBe(true);
});
});
describe('createProductionCliResolverHost', () => {
// Hermeticity gate (the guards PR #329 deleted, restored shared): under
// vitest an un-injected host must neither scan the machine nor spawn a login
// shell — route tests hitting the per-CLI status endpoints would otherwise
// walk the real PATH and execute real binaries on whatever box runs the suite.
it('never scans the machine or spawns a login shell under vitest without injected IO', () => {
const root = mkdtempSync(join(tmpdir(), 'codeman-cli-vitest-gate-'));
temporaryDirectories.push(root);
writeFileSync(join(root, 'codex'), '#!/bin/sh\n');
chmodSync(join(root, 'codex'), 0o755);
execFileSyncSpy.mockClear();
const gatedHost = createProductionCliResolverHost({
processPath: root,
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
});
// The real, executable candidate is invisible: the filesystem predicate is inert.
expect(gatedHost.findOnProcessPath('codex')).toBeNull();
expect(gatedHost.exists(join(root, 'codex'))).toBe(false);
// The login-shell step yields nothing and never reaches execFileSync.
expect(gatedHost.findInLoginShell('codex')).toBeNull();
expect(execFileSyncSpy).not.toHaveBeenCalled();
// The same fixture through the test-only real-IO opt-in IS found, proving
// the nulls above come from the vitest gate rather than from the fixture.
const optedInHost = createProductionCliResolverHost({
processPath: root,
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
runCommand: () => '',
allowRealIoUnderVitest: true,
});
expect(optedInHost.findOnProcessPath('codex')).toBe(join(root, 'codex'));
});
it('resolves through injected IO hooks under vitest (injection is the opt-in)', () => {
const runCommand = vi.fn(() => `${BEGIN_MARKER}\n/home/u/.nvm/bin/codex\n${END_MARKER}`);
const productionHost = createProductionCliResolverHost({
processPath: '',
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
runCommand,
isExecutableFile: () => true,
});
const resolver = createCliExecutableResolver({ binary: 'codex', searchDirs: [] }, productionHost);
expect(resolver.resolve()).toMatchObject({
binaryPath: '/home/u/.nvm/bin/codex',
source: 'login-shell',
});
expect(runCommand).toHaveBeenCalledTimes(1);
});
it('searches the captured process PATH directly in directory order without running a command', () => {
const runCommand = vi.fn(() => '');
const isExecutableFile = vi.fn((path: string) => path === '/second/bin/codex');
const productionHost = createProductionCliResolverHost({
processPath: '/first/bin:/second/bin:/third/bin',
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
runCommand,
isExecutableFile,
});
expect(productionHost.findOnProcessPath('codex')).toBe('/second/bin/codex');
expect(isExecutableFile).toHaveBeenNthCalledWith(1, '/first/bin/codex');
expect(isExecutableFile).toHaveBeenNthCalledWith(2, '/second/bin/codex');
expect(runCommand).not.toHaveBeenCalled();
});
it('accepts only executable regular files with the production predicate', () => {
const root = mkdtempSync(join(tmpdir(), 'codeman-cli-resolver-'));
temporaryDirectories.push(root);
const executableDirectory = join(root, 'executable');
const plainDirectory = join(root, 'plain');
const directoryCandidate = join(root, 'directory');
mkdirSync(executableDirectory);
mkdirSync(plainDirectory);
mkdirSync(directoryCandidate);
writeFileSync(join(executableDirectory, 'codex'), '#!/bin/sh\n');
chmodSync(join(executableDirectory, 'codex'), 0o755);
writeFileSync(join(plainDirectory, 'codex'), '#!/bin/sh\n');
mkdirSync(join(directoryCandidate, 'codex'));
const productionHost = createProductionCliResolverHost({
processPath: [directoryCandidate, plainDirectory, executableDirectory].join(':'),
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
// This test exists to exercise the REAL executable-regular-file predicate
// against its own temp fixtures, so it opts out of the vitest inert-IO
// gate; the stubbed runCommand keeps the login-shell path inert anyway.
runCommand: () => '',
allowRealIoUnderVitest: true,
});
expect(productionHost.findOnProcessPath('codex')).toBe(join(executableDirectory, 'codex'));
});
it('uses the resolved shell, allowlisted args, tagged command, and bounded timeout', () => {
const runCommand = vi.fn(() =>
['/profile/absolute-noise', BEGIN_MARKER, '/home/u/.nvm/bin/codex', END_MARKER, '/exit-trap/absolute-noise'].join(
'\n'
)
);
const productionHost = createProductionCliResolverHost({
processPath: '',
shellPath: '/usr/bin/fish',
shellArgs: ['-i', '-l'],
runCommand,
isExecutableFile: () => true,
});
expect(productionHost.findInLoginShell('codex')).toBe('/home/u/.nvm/bin/codex');
expect(runCommand).toHaveBeenCalledWith(
'/usr/bin/fish',
['-i', '-l', '-c', `printf '%s\\n' '${BEGIN_MARKER}'; command -v -- codex; printf '%s\\n' '${END_MARKER}'`],
{
encoding: 'utf8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
// SIGKILL is load-bearing: interactive bash ignores SIGTERM, and
// execFileSync's timeout only sends the signal, then keeps waiting.
killSignal: 'SIGKILL',
}
);
});
it.each([
['mismatched basename', `${BEGIN_MARKER}\n/opt/bin/not-codex\n${END_MARKER}`],
['missing begin marker', `/opt/bin/codex\n${END_MARKER}`],
['missing end marker', `${BEGIN_MARKER}\n/opt/bin/codex`],
['absolute output outside markers', `/profile/codex\n${BEGIN_MARKER}\nrelative/codex\n${END_MARKER}\n/exit/codex`],
])('rejects malformed tagged shell output: %s', (_name, output) => {
const productionHost = createProductionCliResolverHost({
processPath: '',
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
runCommand: () => output,
isExecutableFile: () => true,
});
expect(productionHost.findInLoginShell('codex')).toBeNull();
});
it('returns null when the shell command throws', () => {
const productionHost = createProductionCliResolverHost({
processPath: '',
shellPath: '/bin/bash',
shellArgs: ['-i', '-l'],
runCommand: () => {
throw new Error('exit 1');
},
isExecutableFile: () => true,
});
expect(productionHost.findInLoginShell('codex')).toBeNull();
});
});

Some files were not shown because too many files have changed in this diff Show More