Compare commits

..
Author SHA1 Message Date
Ark0N d1928f300e Merge pull request #160 from Ark0N/feat/docker-session-mode
v1.4.1: Docker session mode hardening + File Viewer button
2026-07-20 01:41:58 +02:00
Codeman maintainer ca731c67b3 feat(docker): harden session mode + File Viewer button (v1.4.1)
Docker cases: seamless Claude auth (seed ~/.claude.json instead of the
corruption-prone single-file mount), full credential-store isolation for
claude + codex/gemini/gcloud/opencode (share only transcripts/rollouts,
seed the rest), auto-build the base image on first use, C.UTF-8 locale
(fixes box-drawing), collapsed/shortened Create-Case UI + short "(docker)"
case-menu tags, and w<n>-<case> tab naming for docker/remote sessions.
Also: opt-in File Viewer header button; fix a TZ-boundary flaky test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-20 01:36:21 +02:00
Ark0N a3fe0ae728 Merge pull request #159 from Ark0N/feat/docker-session-mode
docs: reflect shipped Docker session mode (1.4.0)
2026-07-19 22:28:41 +02:00
Codeman maintainer 82825cbfb3 docs: reflect shipped Docker session mode (1.4.0) across CLAUDE.md/README/security
- CLAUDE.md: rewrite the Docker cases Key Pattern to the shipped 1.4.0 state
  (removes the stale "not on master / Phases remaining" framing); add
  docker-quickcreate/templates/GPU/elastic-disk/export-import, the
  CODEMAN_DOCKER_BRIDGE_HOOKS listener, docker state files, env vars, route +
  SSE counts, and the build-agent-image command.
- README.md: new "Isolated Docker Sessions" section + a More Features bullet.
- docs/security-architecture.md: new §10 "Docker container isolation" (hardening,
  commit-safe creds, blast radius, untrusted-import safety, bridge-hooks) +
  Quick-reference env vars.

docs/docker-cases.md (user guide) and docs/docker-cases-plan.md (design) were
shipped with the feature.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 22:23:05 +02:00
Ark0N d1868516f7 Merge pull request #158 from Ark0N/feat/docker-session-mode
feat: Docker session mode (isolated per-case containers + export/import) — v1.4.0
2026-07-19 21:59:47 +02:00
Codeman maintainer 3e1272a675 feat(docker): resource templates, GPU, elastic disk, bridge-hooks listener
- One-click "Run in Docker" gains an expandable settings panel with a Template
  picker (Small 2G/1 · Medium 4G/2 default · Large 8G/4 · GPU 8G/4/all) plus
  memory/cpu/gpu/network/image/mount-creds overrides. Any tweak creates a dedicated
  per-case host; the plain checkbox keeps using the shared `default` host.
- GPU passthrough: `gpus` on DockerHost/SessionDocker -> `--gpus <value>` in create
  args (needs the NVIDIA container toolkit). Elastic disk: no `--storage-opt` cap,
  so container storage grows as data flows in.
- CODEMAN_DOCKER_BRIDGE_HOOKS=1: opt-in second listener on the docker bridge gateway
  (auto-detected 172.17.0.1, override CODEMAN_DOCKER_BRIDGE_HOST) that serves ONLY
  the hook endpoints and delegates into the secret-gated pipeline, so in-container
  hooks fire on a loopback-only server. Non-hook paths -> 403; host-internal, not LAN.

Verified live: Large template applies real 8GB/4CPU limits; a secret-authenticated
hook POST from inside a container now reaches the handler (was connection-refused);
non-hook paths return 403; template UI + GPU field verified via Playwright.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 21:35:38 +02:00
Codeman maintainer db6cd838b1 feat(docker): one-click "Run in Docker" case creation + Export button
- New POST /api/cases/docker-quickcreate: creates a normal case (folder in
  CASES_DIR, scaffolded CLAUDE.md + hooks) AND links it to a hardened container
  with default settings, auto-provisioning a shared `default` docker host — the
  user never touches host/image/network fields.
- Create New tab gains a "Run in isolated Docker container" checkbox; on submit it
  calls docker-quickcreate then auto-starts a claude session inside the container.
- Case Manage list gains an Export (full-image) button per docker case.
- SSE listeners for docker:exportComplete/exportFailed toast + refresh the exports
  list.

Verified end-to-end on the live instance: one-click create put the case in
~/codeman-cases/<name>, auto-created the default host, launched claude in the
container; export button produces a bundle; checkbox + button render (Playwright).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 19:28:13 +02:00
Codeman maintainer 66a41f5aa9 chore: version packages 2026-07-19 19:01:50 +02:00
Codeman maintainer a36c1f62db fix(docker): set CLAUDE_CODE_TMPDIR + document hook reachability limit
Found in live testing: claude refuses its default /tmp/claude-<uid> temp dir when
that path pre-exists root-owned (happens when the workspace bind-mount traverses
it, e.g. a workspace under /tmp/claude-<uid>). Set CLAUDE_CODE_TMPDIR to a
nonexistent HOME subpath the running uid creates+owns, so docker claude sessions
are robust to any workspace location.

Also document the hook-reachability constraint: in-container hooks POST to
host.docker.internal (the bridge gateway), so they only fire when Codeman is
reachable from the container (bind 0.0.0.0 + password); on a loopback-only bind
they don't fire and idle detection falls back to output-based (which works).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 18:19:50 +02:00
Codeman maintainer 8b2c857c3f feat(settings): wire session, away-digest, and cron button visibility toggles
Per-device App Settings > Header Displays toggles that show/hide the session
manager and away-digest header buttons (default OFF) and the cron footer
button (default ON). Adds the load/save/apply/default/displayKeys wiring in
settings-ui.js plus the marker CSS in styles.css. Client-only display keys,
stripped from the settings PUT so they never reach the strict server schema
(mirrors the showAttachmentsButton pattern); session/away stay hidden on
phones via the existing mobile.css rules. The button markup and checkbox
rows landed earlier in 5728b86.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 17:59:23 +02:00
Codeman maintainer 583678c950 docs(docker): add user-facing docs/docker-cases.md
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 17:55:21 +02:00
Codeman maintainer 5728b86a68 feat(docker): frontend Docker tab, run wiring, and export/import UI
- index.html: Create Case "Docker" tab (name/workspace/host/image/network +
  advanced memory/cpus/mountCredentials/resumeOnStart), and a Docker-exports
  section in the Manage tab
- session-ui.js: linkDockerCase (POST docker-host, PUT on conflict, then
  docker-link; omitted optionals as undefined not null), case-picker label
  "name @ container" + search fields, switchCaseModalTab/submitCaseModal docker
  branch, and export/import UI (refresh/export/import/delete). Docker cases route
  through /api/quick-start like remote (runClaude/runShell/runOpenCode/Codex/Gemini)
- verified in a real browser (Playwright): Docker tab renders, linking through the
  UI creates the case and it appears in the picker as "uitest @ codeman-case-uitest"

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 17:53:14 +02:00
Codeman maintainer 39ef17b6af feat(docker): export/import (move a container to another machine) + boot reaper
- src/docker-export.ts: full-image export (pause-consistent commit + save|stream +
  workspace tar + manifest -> one .codeman-container.tgz) and workspace-only; import
  validates manifest + per-member sha256, traversal-guards the workspace tar, docker
  load + quarantine re-tag (never overwrites a local tag). Bounded by
  runWithConversionLimit; free-space precheck; docker rmi in finally; sealed
  containers refuse full-image export.
- routes: POST /api/docker-cases/:name/export (background + SSE), GET/DELETE
  /api/docker-exports, GET download, POST /api/docker-cases/import (-> new host+case)
- instance-scoped boot reaper (docker-hosts.reapOrphanedDockerContainers) wired after
  restoreMuxSessions; never touches another instance's containers
- SSE docker:exportComplete/exportFailed/importComplete (both registries)
- fix: stream pipeline in saveImageToTar so the bundle isn't truncated

VERIFIED end-to-end on real docker: full export -> 326MB valid bundle -> delete
case -> import -> new container runs from the quarantined image with the workspace
file AND the in-image change both restored.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 17:38:01 +02:00
Codeman maintainer 814362b67b docs(docker): record implementation status (phases 0-5 done, e2e verified)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:38:53 +02:00
Codeman maintainer e9f9497259 feat(docker): allowlist container-to-host gateway aliases in host guard
An in-container hook curl carries Host: host.docker.internal:<port> (the derived
CODEMAN_API_URL), so the always-on host guard must allow host.docker.internal /
host.containers.internal or every in-container hook is blocked 403. Exact-match
only; not a browser DNS-rebinding surface (resolves to the host only from inside
a container netns). Verified end-to-end: quick-start launches claude/shell in a
real container with the workspace bind-mounted and hooks scaffolded.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:32:45 +02:00
Codeman maintainer 8768ca4a5a feat(docker): docker-hosts CRUD, docker-link, and quick-start branch
- case-routes: GET/POST/PUT/DELETE /api/docker-hosts, POST /api/cases/docker-link
  (creates workspace, probes daemon + tmux-in-image), docker listing in
  GET /api/cases, docker-unlink (best-effort docker rm -f) in DELETE, single GET
- session-routes: /api/quick-start docker branch (rejects envOverrides/effort/
  per-CLI config, probes availability + tmux, casePath=hostWorkspacePath, seeds
  resume id, scaffolds hooks+CLAUDE.md if missing, threads docker into Session,
  Ralph auto-config skipped for docker)
- CaseInfo gains location:'docker' + docker{} block
- typecheck clean; 157 route+docker tests pass

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:28:44 +02:00
Codeman maintainer df9214ba9a feat(docker): thread SessionDocker through Session + recovery
- Session: _docker field, constructor config, toState, createSessionOptions/
  respawnPaneOptions (both interactive + shell paths), docker getter
- resolveMuxAttachCwd returns /tmp for docker sessions (local wrapper only execs)
- skip the LOCAL claude version probe for docker; probe the IN-CONTAINER version
  instead (deferred) so wheel-forwarding stays enabled (#154)
- server restoreMuxSessions round-trips MuxSession.docker / SessionState.docker
- full CI suite green (3444 passed)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:22:38 +02:00
Codeman maintainer 5f4c89b990 fix(docker): auto-assign agent uid (node:22 already occupies uid 1000)
node:22-bookworm-slim ships a `node` user at uid 1000, so `useradd -u 1000`
failed. Auto-assign the uid and rely on gid-0 + group-writable HOME so any
runtime `--user <hostUid>:0` can write $HOME. Verified: image builds; toolchain
(node/tmux/claude/codex/gemini/opencode) present; `--user 1000:0` writes
/home/agent and `claude --version` runs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:14:44 +02:00
Codeman maintainer 54615e2371 feat(docker): agent base image + local build script
docker/agent.Dockerfile: node:22 + claude/codex/gemini/opencode CLIs + git/
tmux/ripgrep/curl, secret-free, OpenShift arbitrary-uid-writable HOME (gid 0).
scripts/build-agent-image.mjs: local build (decision "build locally on first
use"), docker/podman auto-detect, --engine/--image/--no-cache flags.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:11:21 +02:00
Codeman maintainer 828b1664f7 feat(docker): Docker session mode foundation (types, storage, tmux builders)
Phase 0-2 of the Docker cases feature (docs/docker-cases-plan.md). Docker is a
LOCATION OVERLAY on cases (not a 6th SessionMode), mirroring the remote-SSH
feature: a local tmux pane runs `docker exec -it` into a durable in-container
tmux server. The container is per-CASE, so multiple sessions share it.

- types: DockerHost/DockerCase/SessionDocker + docker? on SessionState/MuxSession
- src/docker-hosts.ts: storage, toSessionDocker, pure buildDockerBaseArgs/
  buildDockerCreateArgs (cap-drop, no-new-privileges, --pull=never, mem==swap,
  never privileged/socket), containerApiUrl, hostGatewayAlias, config-hash,
  credential-mount resolution, daemon probes (VITEST no-op)
- schemas: DockerHostSchema + DockerCaseLinkSchema (NO_SHELL_META guards)
- tmux-manager: buildDockerLaunchCommand (image-check -> ensure -> start -> exec,
  resume-aware), buildDockerKillCommand (in-container tmux only, multi-session
  safe), stop/remove; wired into createSession/respawnPane/killSession
- 40 unit tests (docker-hosts + docker-exec-options), typecheck clean

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-19 15:09:48 +02:00
Codeman maintainer 6f4b2b8a17 chore: version packages
Release 1.3.5. Consumes the changeset from PR #155: re-issue the
codeman_session cookie on every authenticated request so the browser cookie
lifetime tracks the server-side sliding TTL, fixing the recurring native Basic
Auth dialog during active use.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 00:15:29 +02:00
Codeman maintainer a531f48e17 chore(gitignore): ignore local screenshot and design capture dirs
screenshots-readme/, screenshots-readme-real/, screenshots-real/ and
design-explorations/ are local capture scratch that was untracked but not
ignored, so an unqualified `git add -A` during a COM could sweep them into a
release (this has happened before).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-17 00:15:29 +02:00
Ark0N 7d5ea0bd50 Merge PR #155 from dennisentruencer/fix/sliding-auth-cookie: slide the session cookie so active users aren't logged out
Re-issue the codeman_session cookie on every authenticated request so the browser cookie lifetime tracks the server-side sliding TTL (authSessions already used refreshOnGet: true). Fixes the recurring native Basic Auth dialog during active use.

Reviewed: no token rotation (same server-generated token re-issued, so no fixation vector), forged cookies are not blessed, logout still emits only the clearing cookie and server-side invalidation holds, cookie attributes identical to the Basic Auth path. Verified against the merge result: tsc --noEmit, lint, format:check, check:frontend-syntax, check:lockfile, and npm run test:ci (3404 passed) all green.
2026-07-16 23:49:10 +02:00
Codeman maintainer a9ae141eec chore: version packages 2026-07-16 23:34:03 +02:00
Codeman maintainer 7b79d4207c fix(terminal): restore Claude scroll-back on macOS trackpads (#154)
Deterministic claude --version probe seeds cliVersion so wheel-forwarding
to Claude's transcript engages (banner scrape was unreliable on 2.1.187+
and resumed sessions). Shift+wheel reads the dominant axis so a trackpad's
horizontal Shift-scroll reaches local scrollback. New per-device
"Wheel Scrolls Local History" opt-out. Wheel reports use a fire-and-forget
send path so they no longer flicker the pending-bytes indicator.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 09:33:39 +02:00
Codeman maintainer 28744a2761 chore: version packages
Make the Cron Jobs modal skin-aware + consistent with App Settings:
skin-variable selects (appearance:none, --bg-input fill, custom chevron),
color-scheme:dark for native controls, themed date/time inputs, and
btn-toolbar-sized toolbar/footer buttons. Bumps aicodeman to 1.3.2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 15:59:02 +02:00
Codeman maintainer cca07e2b11 chore: version packages
Redesign the Cron Jobs modal to match App Settings styling + fix the
create form never collapsing (scoped #cronModal .hidden rule). Bumps
aicodeman to 1.3.1.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-13 15:09:47 +02:00
Codeman maintainer 9806efdf0a chore: version packages 2026-07-13 00:49:49 +02:00
Codeman maintainer b00e7cf17c Merge PR #153 from aakhter/cod-161-session-manager-frontend: unified Session Manager — welcome list + searchable modal + live SSE refresh
Rebuilt on the merged #146 Session Manager: kept master's Command Palette + fixed
Session Manager implementation, dropped the PR's stale duplicate block (last-key-wins
regression), rebased _buildHistoryItem on master's onActivate contract, kept the new
projectKey plumbing + SSE live refresh + kebab menu/badges, phone-hid the header button.
2026-07-13 00:43:18 +02:00
Codeman maintainer efe2d8966a fix(review): rebuild Session Manager additions on the merged #146 implementation (PR #153)
- Hide the new btn-session-manager header button on phones: add it to the
  @media (max-width: 430px) display:none block in mobile.css (next to
  .btn-away-digest) and to KNOWN_PHONE_HIDDEN in the mobile-header policy
  test, closing the recurring phone-header-leak regression that was PR
  #153's red CI job.
- Put the session-manager header button on its own line in index.html
  (was crammed onto the away-digest line).
- app.js: drop session:updated from the unified-list SSE refresh trigger —
  it is batch-broadcast ~every 500ms per active session and would turn an
  open modal / visible welcome list into a sustained ~1 Hz full projects
  rescan loop; created/deleted (structural changes) are sufficient.
- terminal-ui.js _fetchUnifiedSessions: check the ApiResponse envelope and
  throw on failure so a 5xx surfaces via the caller's catch instead of
  rendering an empty history.
- terminal-ui.js _openSessionRowMenu: on re-entry, invoke the previous
  menu's close fn (stored as _openRowMenuClose) so its document/window
  listeners are detached rather than leaked; use claudeSessionId ||
  sessionId in the 'Resume session' menu item to match the main-row and
  Session Manager resume routing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-13 00:38:56 +02:00
Codeman maintainer f55f035690 Merge master into PR #153 (unified Session Manager)
Resolve the 4 conflicted files toward master's merged #146 work while
keeping PR #153's genuinely-new additions:

- app.js: keep the full Escape chain (closeSessionManager +
  closeCommandPalette + closeShortcutOverlay).
- index.html: keep master's Command Palette modal markup alongside the
  PR's Session Manager modal + header button.
- styles.css: keep master's Command Palette + COD-157 shortcut CSS AND
  the PR's COD-130 session-row kebab-menu CSS (both inserted at the same
  spot — reunited each with its own closing brace).
- terminal-ui.js: resolve _buildHistoryItem's main-row click handler to
  master's options.onActivate contract with a liveness + claudeSessionId
  -aware resume default, preserving the PR's two-shape/badges/kebab body.
- panels-ui.js: the PR's pre-#146 Session Manager block auto-merged as a
  duplicate AFTER master's fixed block (last-key-wins regression) — drop
  it, keep master's implementation plus the PR's new
  _onSessionListMaybeChanged.

Backend projectKey plumbing and the SSE live-refresh listeners in app.js
merge additively and are kept as-is.
2026-07-13 00:34:43 +02:00
Codeman maintainer 58fc5f874a docs(CLAUDE.md): accuracy audit fixes + document the 15 merged PRs
Audit (18 verified findings): COM step 6 watches BOTH CI+Release runs; hook-secret
is unconditional when auth is active (COD-91); env-prefix allowlist includes
GEMINI_*/GOOGLE_*; applySkin()/isWorkflowAgentTrackingEnabled() name fixes;
ultracode watcher completion-vs-live sources; harvestSources location; LRUMap
barrel exception; gemini-cli-resolver; config 15 files; state-files inventory;
terminal-history centralization; tunnel.sh named mode; test:watch row; shortcut
list corrections.

New feature docs: cron jobs, remote SSH cases, unified session list, command
palette + shortcut registry, PTY-exit breaker, full-scrollback replay, WS
resilience, Codex artifacts/response-viewer, HEIC conversion, WebGL toggle;
route/SSE/type counts refreshed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 20:16:07 +02:00
Codeman maintainer 1301b4b58c Merge PR #152 from pirronewantlux529-coder/codex-response-viewer: response-viewer (eye) support for Codex sessions
Includes review fixes: full route-test coverage for the rollout locator/parser (originator/uuid/pin resolution, dedup, injected-context filtering), LRU caches, multi-block text joins.

# Conflicts:
#	src/web/routes/session-routes.ts
2026-07-12 20:09:53 +02:00
Codeman maintainer 46493f374e Merge PR #151 from aakhter/cod-167-heic-jpeg-conversion: convert HEIC paste uploads to JPEG
Includes review fixes: worker-thread conversion with resourceLimits + timeout, global conversion-limiter cap, 64MP pre-decode bomb guard, magic-byte detection (covers mislabeled Android HEIF).
2026-07-12 20:08:43 +02:00
Codeman maintainer b20c00702a Merge PR #150 from aakhter/cod-166-codex-generated-artifact-attachments: Codex generated artifacts as attachment cards
Includes review fixes: source arg threaded through the deps lambda (was silently dropped), codex-mode gating, realpath-first trust decisions with homedir-anchored markers, image thumbnail passthrough, ANSI-stripped scanning.

# Conflicts:
#	src/session.ts
2026-07-12 20:08:32 +02:00
Codeman maintainer d55ebcb644 Merge PR #149 from aakhter/cod-165-ws-resilience: WebSocket durable-delivery resilience
Includes review fixes: real _wsState lifecycle (connecting/connected/disconnected), per-tab supersede identity (multi-tab coexistence), preserved reconnect backoff, connection-dot CSS for connected/fallback states.
2026-07-12 20:07:55 +02:00
Codeman maintainer e84a3834d0 Merge PR #148 from aakhter/cod-164-scrollback-replay-crlf: replay full tmux scrollback on terminal reload + CRLF normalization
Includes review fixes: explicit ?full=1 trigger wired from initial page load, capture maxBuffer sized from config with -S line bound, capture returned alone (no byte-buffer duplication), early byte-cap before normalization.
2026-07-12 20:07:35 +02:00
Codeman maintainer f89bc420ba Merge PR #147 from aakhter/cod-168-pty-exit-breaker: scrub TMUX vars + PTY-exit circuit breaker (COD-115/COD-118)
Includes review fixes: breaker reset only on explicit clearBreaker restarts (auto-reattach never clears), trip observability survives listener detach, push notification wired into PUSH_EVENT_MAP.

# Conflicts:
#	src/session.ts
2026-07-12 20:07:19 +02:00
Codeman maintainer 5a4e60dc8e fix(merge): reconcile cross-PR test seams after #141/#145/#146 merges
- help-modal extractor bounds at the next HTML comment (cron modal's 'Run At'
  text false-positived the stale-shortcut regex)
- remote-shell run test expects the wired /api/quick-start path (#145) — POST
  /api/sessions has no caseName in its schema

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 20:06:32 +02:00
Codeman maintainer 460972a50e Merge PR #146 from aakhter/cod-163-command-palette: searchable case picker + Command-K session palette
Includes review fixes: Session Manager aligned to the merged /api/sessions/unified contract with error states, Ctrl+K no longer leaks 0x0B into the PTY, shortcut registry finished (dispatch/persistence/rendering), shortcutOverrides preserved across settings saves, help modal kept reachable.

# Conflicts:
#	README.md
#	src/web/public/index.html
#	src/web/public/session-ui.js
2026-07-12 20:03:53 +02:00
Codeman maintainer 8a971c3935 Merge PR #145 from aakhter/cod-94-remote-host-ssh: remote host SSH cases
Includes review fixes: reachable Remote tab UI, remote metadata restore on recovery, quick-start routing for remote run flows, ssh-arg injection guards, dedicated remote socket/name (no cross-instance adoption), remote tmux kill on delete, wired tmux probe + ConnectTimeout, --dangerously-skip-permissions default.
2026-07-12 20:01:47 +02:00
Codeman maintainer 83779cab4d Merge PR #141 from chatgptkrylor/feat/scheduler: recurring cron-style scheduled jobs
Includes review fixes: multi-line prompt rejection, prompt-file confinement hardening (realpath + attachment-guard blocklist), per-job autoClosePreviousSession lifecycle, live-session-only concurrency counting, wired launchCommand.
2026-07-12 20:01:16 +02:00
Codeman maintainer a8e7669f5a fix(review): preserve shortcutOverrides across settings saves + keep help modal reachable (PR #146)
- saveAppSettings() rebuilds settings from the DOM; carry over shortcutOverrides
  like showTokenCount/showCost so rebinding survives unrelated saves
- shortcut overlay footer links to the full help modal (its only opener was the
  legacy Ctrl+? route this PR replaced)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 19:58:32 +02:00
Codeman maintainer 5deb0d4a4c fix(review): harden + wire remote-host SSH cases end-to-end (PR #145)
- UI: add the missing data-tab="case-remote" tab button; dispatch it through
  submitCaseModal()/switchCaseModalTab() to linkRemoteCase() (was dead code).
- Restore: restoreMuxSessions() now passes remote (muxSession.remote ??
  savedState.remote) into the Session constructor, so remote metadata round-trips
  on restart instead of reattaching from a local cwd / respawning LOCAL / being
  erased from state.json. Recovery tests added.
- Run flows: runClaude()/runShell() route remote cases through /api/quick-start
  (POST /api/sessions stat-validates workingDir locally); run*() skip the
  /api/*/status pre-check and omit inert config/env for remote cases.
- Quick-start: resolve the remote case BEFORE the local CLI availability gates and
  skip isCodex/Gemini/OpenCodeAvailable() when remote; REJECT
  envOverrides/effort/codex/gemini/openCode config for remote (they don't cross
  ssh) instead of silently dropping them.
- Injection: reject $, backtick, $( in remotePath + identityFile at the schema
  layer (they survive shellescape into the bash -c launch double-quote layer).
  Regression tests for $(...) and backtick payloads added.
- Remote socket/name: launch on a DEDICATED -L codeman-remote socket under a
  codeman-ssh-<id> name that fails a remote Codeman's SAFE_MUX_NAME_PATTERN, so a
  remote instance can't adopt the session; scope tmux set-options per-session
  (never -g) so they don't mutate other sessions.
- Kill: best-effort ssh 'tmux -L codeman-remote kill-session' on remote session
  kill (fire-and-forget, never blocks/throws the local kill) so the remote agent
  isn't orphaned forever.
- Probe: wire checkRemoteTmuxAvailable() into POST /api/quick-start (structured
  OPERATION_FAILED) and as courtesy validation in remote-link; add a default
  -o ConnectTimeout=10 to buildSshConnectionArgs (overridable via extraSshOptions).
- Command default: remote claude default is now
  'exec claude --dangerously-skip-permissions' (per-host override stays the escape
  hatch), mirroring local non-interactive semantics.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 19:49:58 +02:00
Codeman maintainer 84ab4ff07b fix(review): harden cron security, session lifecycle, skip policy (PR #141)
- Reject multi-line prompts end-to-end: schema refines on promptText/
  launchCommand, runtime check in resolvePrompt (prompt-file content;
  trailing newlines tolerated), matching cron-ui form validation — delivery
  is single-line only, so multi-line was silently corrupted (typed mode
  fused lines, paste mode submitted partials)
- Close the workingDir confinement bypass (arbitrary server-side file read,
  e.g. workingDir=/proc + /proc/self/environ): realpath-resolve workingDir
  before the containment check, reject '/' and blocked/pseudo-fs trees
  (/proc, /sys, /dev + the attachment-guard blocklist) at fire time AND at
  job create/update (workingDir must exist and be a directory)
- Session lifecycle: new per-job autoClosePreviousSession (default true,
  recurring schedules only; ignored for 'once') — the previous run's
  still-open session is closed via the normal cleanupSession path when the
  next run fires; UI switch added; 50-session cap math documented in
  docs/cron-guide.md §8
- skip_if_same_agent_running: count only live sessions (exclude
  stopped/error dead tabs), exclude sessions created by this job's own runs
  (fixes the fire-once-then-skip-forever self-deadlock), and a skipped
  'once' job stays armed and retries next tick instead of being consumed;
  liveness filter mirrored in cron-ui _countActiveAgents
- Wire launchCommand (was accepted+documented but dead): shell mode sends
  it via writeViaMux as the first input line after startShell readiness
  (single-line, schema-enforced); form field shown for shell agent type
- Record delivery failures: a false writeViaMux result now fails the run
  instead of recording a false 'prompt_sent'
- Cap saved jobs at MAX_CRON_JOBS (100) to bound state.json growth
- Surface field-specific schema messages (drop parseBody custom
  errorMessage on cron create/update)
- Tests: workingDir create/update validation, /proc bypass regression,
  single-line enforcement (schema+runtime+trailing-newline tolerance),
  live/own-session skip filtering, once-skip re-arm, auto-close on/off/once,
  job-count cap

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 19:25:42 +02:00
Codeman maintainer 88f47754ad fix(review): wire Session Manager to /api/sessions/unified contract, stop Ctrl+K PTY leak, finish shortcut registry (PR #146)
- Session Manager (COD-121/192): align _loadSessionManagerList() with the
  merged #139 endpoint — map UnifiedSessionItem fields (lastActivityAt
  epoch-ms → lastModified, optional sizeBytes/firstPrompt/name) to the
  history-record shape _buildHistoryItem renders; surface non-2xx /
  error-envelope responses as a visible message instead of a silent
  "No sessions found"; route clicks by liveness (live row → selectSession,
  history row → resumeHistorySession by conversation UUID) via a new
  onActivate option so a live session is never duplicate-resumed
- Ctrl+K double-dispatch: gate the palette chord in
  attachCustomKeyEventHandler (return false on keydown) so xterm never
  writes 0x0b kill-line into the PTY while the palette opens; gate is
  registry-aware so a rebound/disabled palette shortcut restores normal
  terminal Ctrl+K
- Shortcut registry (COD-157) finished per maintainer decision: document
  keydown now dispatches through getShortcutRegistry() +
  matchesShortcutEvent() (legacy SHORTCUTS table removed), honoring
  per-shortcut disable and rebinds incl. the palette chord; overrides
  persist via saveAppSettingsToStorage() (correct device key + cache
  coherence, was orphaned 'codeman:settings'); Shortcuts tab renders on
  open via switchSettingsTab hook; capture uses a persistent listener that
  ignores bare modifier keydowns (combos now capturable) and requires a
  Ctrl/Cmd/Alt chord; settings rows use delegated listeners instead of
  inline onclick (JS-string injection sink) and overrides can no longer
  clobber id/label/action; added the missing row + overlay CSS
- matchesShortcutEvent: reject undeclared extra modifiers (Ctrl+Shift+K
  no longer hijacked from Firefox devtools) while keeping Ctrl/Cmd
  interchangeable; match physical code OR produced key for layout parity
- Registry/dispatch gaps: added restore-terminal-size entry, documented
  Ctrl+Shift+R again in the help modal (test flipped to assert presence),
  Ctrl+?/Alt+? now really open the registry-driven shortcut overlay, and
  Escape closes it
- Palette new-session pick routes through selectQuickStartCase() so the
  searchable combobox, dir display, and lastUsedCase stay in sync
- Removed fork cherry-pick debris: dead _onSessionListMaybeChanged(),
  orphaned .session-row-menu CSS, nonexistent closeMobileHeaderUtilities
  calls
- Tests: functional vm-harness coverage for the unified-list field
  mapping + error state + liveness routing, palette chord shift/disable/
  rebind handling, override persistence round-trip, capture flow, tab
  render hook, and source guards for the PTY gate + registry dispatch

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 19:06:12 +02:00
Codeman maintainer 6e417d69dc fix(review): WS state machine, per-tab supersede key, backoff, dot CSS (PR #149)
- _wsState now transitions through the full lifecycle: _connectWs() sets
  'connecting', ws.onopen (inside the this._ws === ws guard) sets 'connected',
  _disconnectWs() resets to 'disconnected' — the connection chip's "WS" state
  was previously unreachable (stuck on "WS…"/"HTTP" forever).
- WS registry supersede is now keyed per TAB: the upgrade URL sends
  cid = clientId + ':' + per-page nonce (reusing the constructor's page UUID),
  while input frames keep the bare browser clientId for seq dedup — two
  tabs/windows on one session coexist instead of 4010-evicting each other in a
  perpetual 5s ping-pong; a genuine same-tab reconnect still supersedes.
- Exponential backoff engages: _disconnectWs() no longer zeroes
  _wsReconnectAttempts (it's called at the top of _connectWs, so every retry
  replanned at attempt 0 → ~0ms tight reconnect loop during outages); onopen
  resets the counter on success.
- styles.css: add .connection-dot.connected (green) and .connection-dot.fallback
  (yellow) — both states rendered an invisible dot (no rule existed).
- Remove smuggled dead code: resolveMonitorRowLabels/CodemanMonitorLabels
  (COD-122, no consumer, referenced test doesn't exist) and the never-written
  _wsLastClose/_wsInputSendCount/_httpFallbackSendCount diagnostics.
- Tests: new test/ws-state-lifecycle.test.ts drives the REAL
  _connectWs/onopen/onclose/timer cycle (state transitions, escalating backoff
  delays, composite cid on the upgrade URL); registry two-tab coexistence test;
  static check that every emitted connection-dot class has a styles.css rule.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 18:42:36 +02:00
Codeman maintainer f98d29b323 fix(review): wire full-scrollback replay to an explicit ?full=1, dedup + bound the capture (PR #148)
- Replace the 'missing ?tail means reload' overload with an explicit ?full=1
  query param: the frontend's first buffer load after a page load (selectSession)
  now requests full=1, tab switches keep ?tail=, and the legacy no-param callers
  (response-viewer fallback, clearTerminal refresh) keep the cheap visible-frame
  path — the COD-47 feature was previously unreachable from a real reload.
- When the full-history capture succeeds, return it ALONE instead of prepending
  the byte buffer + \x1b[H\x1b[2J: the capture is the rendered superset of the
  byte history, and ED2 clears only the viewport so the concat replayed the whole
  conversation twice in xterm scrollback. The history+clear+frame concat stays
  for the visible-frame/tab-switch path.
- Pass an explicit execSync maxBuffer for the full-history capture (configured
  terminalBufferMaxBytes + slack) — the 1MB Node default ENOBUFS-killed exactly
  the multi-MB captures the feature exists for; log ENOBUFS concisely instead of
  dumping the truncated stdout.
- Bound the capture itself via -S -<N> derived from the configured tmux
  history limit (was unbounded -S -), and add -J so lines hard-wrapped at the
  capture-time pane width reflow in the browser xterm.
- Cap the concatenated buffer to terminalBufferMaxBytes EARLY (before the
  regex normalization passes) so multi-MB captures don't stall the event loop
  normalizing bytes that get sliced away.
- Tests: route tests updated for ?full=1 semantics (capture-alone response,
  config-forwarded capture bounds, byte-history fallback, no-param requests
  stay on the visible-frame path); source-scan tests cover the bounded -J -S -<N>
  flags and explicit maxBuffer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 18:33:18 +02:00
Codeman maintainer 360d58ca4f fix(review): breaker reset semantics, trip observability, push template (PR #147)
- Breaker reset is now explicit-only: POST /api/sessions/:id/interactive no
  longer unconditionally resets the PTY-exit breaker (that endpoint IS the
  frontend's automatic re-attach path, so the breaker could never trip on the
  COD-115 crash loop and any tab click silently re-armed it). The route accepts
  a schema-validated optional body flag {clearBreaker:true}
  (InteractiveStartSchema) and resets only when it is sent.
- Frontend restart control: app.js selectSession keeps the bare auto-attach
  (no body, never clears); when the selected session has respawnBlocked it asks
  for explicit user confirmation and only then re-POSTs with clearBreaker:true.
  respawnBlocked is surfaced via SessionState/toState() (runtime-only, not
  restored on boot so recovery can re-attach).
- Trip observability: WebServer.setupSessionListeners() is now idempotent
  (skips while refs are attached) and the re-attach routes (/interactive,
  /interactive-respawn, /shell) re-run it, restoring the wiring that the exit
  handler detaches on every PTY exit — without this the 5th-exit trip had
  guaranteed zero listeners (no SSE, no push, no persist, no run-summary).
- Push notification: added SessionRespawnBreakerTripped to PUSH_EVENT_MAP
  ('Session crash loop stopped', urgency critical) with an exit-count body
  branch; previously sendPushNotifications silently no-oped.
- Minor: buildMuxAttachEnv() truecolor param is now actually passed
  (codex/gemini, mirrors buildEnvExports); buildClaudeEnv() uses delete for
  COLORTERM/CLAUDECODE (same node-pty "KEY=undefined" quirk as COD-115).
- Tests: route tests assert auto-reattach does NOT reset, clearBreaker resets,
  invalid flag rejected, and listener re-wiring on /interactive + /shell;
  real-wiring lifecycle tests (createSessionListeners/attach/detach) prove the
  exit-detach gap and that re-setup keeps the 5th-exit trip observable;
  PUSH_EVENT_MAP regression guard.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 18:21:42 +02:00
Aamer Akhter 6cac517fa6 COD-130 session-row ⋯ becomes a context menu
The per-row ⋯ in the session list was a details toggle that did nothing in
the Session Manager modal (swallowed by the modal's capture-phase
close-on-click). Replace it with a real kebab context menu.

- terminal-ui.js: ⋯ now opens _openSessionRowMenu() — a body-anchored popup
  (fixed-positioned, flips/clamps to viewport, z-index above the modal) with:
  Resume/Switch-to (live→select tab, closed→resume), Open folder in the file
  browser (live sessions only — the browser is session-scoped), Copy path
  (_copyText + toast, when workingDir present), and Show details (the old
  inline prompt/path panel). Closes on outside-click / Escape / scroll / resize.
- panels-ui.js: _loadSessionManagerList scopes its modal-close to the
  .history-item-main (resume) click, so the ⋯/menu no longer closes the modal.
- styles.css: .session-row-menu + .session-row-menu-item.

Verified in Chromium on an isolated beta: ⋯ opens the menu with the modal
still open; closed rows show Resume/Copy path/Show details, live rows add
Switch-to + Open folder; Show details expands inline (modal stays open),
Copy path copies the path, Resume closes the modal, Escape closes only the
menu. Gates: tsc 0, lint 0, frontend-syntax + public-asset format clean.
2026-07-12 12:19:10 -04:00
Aamer Akhter 2235f06ea5 COD-121 unified session list: live SSE refresh (slice A, unit 4)
The complete session list now updates live as sessions change, instead of
only on open/welcome-load.

- app.js: extra SSE listeners (session:created/updated/deleted) on the same
  EventSource (multiple listeners per event; existing handlers untouched;
  registered via addListener so they tear down on reconnect) call
  _onSessionListMaybeChanged().
- panels-ui.js: _onSessionListMaybeChanged() debounced-refreshes the Session
  Manager modal when it's open and the welcome list when its overlay is
  visible (no work when neither is showing). _loadSessionManagerList stores the
  active query so refreshes preserve the user's search.

Verified on an isolated beta instance (Playwright): dispatching a session
event refreshes the modal while open, does NOT while closed (gated), and
refreshes the welcome list while visible. Gates: tsc 0, frontend-syntax +
public-asset format clean, build clean.
2026-07-12 12:19:10 -04:00
Aamer Akhter 65b609b6db COD-121 unified session list: persistent Session Manager modal (slice A, unit 3)
Adds a header-reachable Session Manager so the complete session list is
available mid-session, not only on the welcome screen.

- index.html: always-on header button (.btn-session-manager) + #sessionManagerModal
  (mirrors the Away Digest modal) with a search box + results list.
- panels-ui.js: openSessionManager()/closeSessionManager()/_loadSessionManagerList()
  — loads GET /api/sessions/unified (limit 200), renders via the unit-2
  _buildHistoryItem (rich items, mode/LIVE badges, open->select / closed->resume),
  debounced search wired to the endpoint's q= param, empty/error states. A
  modal-scoped Escape listener closes it even when focus is in the search input;
  backdrop click and item click also close it.
- app.js: closeSessionManager() added to the global Escape chain.
- styles.css: modal + list styling (items reuse .history-item).

Verified on an isolated beta instance (Playwright): the header button opens the
modal, it lists 200 sessions from /api/sessions/unified, a no-match query issues
?q= to the server and yields 0 items, clearing restores the list, clicking an
item closes the modal and routes resume/select, and Escape closes it. Gates:
tsc 0, lint 0, frontend-syntax + public-asset format clean, 17 tests pass.
2026-07-12 12:19:10 -04:00
Aamer Akhter c9f37f2628 COD-121 unified session list: welcome list frontend (slice A, unit 2)
Backs the welcome-screen "Resume Conversation" list with the new
GET /api/sessions/unified endpoint instead of /api/history/sessions, so it
shows the COMPLETE set (live + persisted + non-Claude + closed history)
newest-first with richer context, rather than only Claude transcripts.

- terminal-ui.js: new _fetchUnifiedSessions(); loadHistorySessions() now uses
  it. _buildHistoryItem upgraded to the unified shape (kept backward-compatible
  with the folder-modal's old shape): title = name || firstPrompt || dir; a
  mode badge + a LIVE badge (sources includes 'live'); timestamp from
  lastActivityAt (falls back to lastModified); size only when present; detail
  panel + "View all in this folder" preserved (gated on projectKey). Resume
  branches: an open live session selects its tab, a closed one resumes.
- unified-session-service.ts + endpoint: pass projectKey through the history
  source so the folder drill-down survives.
- styles.css: .history-item-badges / -badge / -badge-live pills.

Verified: tsc 0, lint 0, frontend-syntax + public-asset format clean, service
tests 13/13 (+projectKey), route tests 4/4. Playwright on an isolated beta:
the welcome list renders real items from /api/sessions/unified, and the
renderer produces the tab-name title + codex mode badge + visible LIVE badge,
omits LIVE on closed items, keeps "View all in folder", and routes resume
correctly (open->select tab, closed->resume). Persistent panel + live SSE
status are later units.
2026-07-12 12:19:10 -04:00
Codeman maintainer 309959be27 fix(review): worker-thread HEIC conversion with bomb guard, concurrency cap, and magic-byte routing (PR #151)
- Event-loop blockage: HEIC decode/encode (CPU-synchronous libheif WASM +
  jpeg-js) now runs in a per-conversion worker_threads Worker
  (src/web/heic-jpeg-worker.ts, spawned by heic-jpeg-converter.ts) with
  resourceLimits and a 30s hard timeout that terminates the worker —
  verified end-to-end under tsx and against compiled dist/ output with a
  real iPhone HEIC (event-loop max stall 52ms during conversion).
- No server-side concurrency cap: conversions now acquire a slot from the
  existing global runWithConversionLimit() pool (document-conversion-limiter),
  bounding peak decode memory/CPU across simultaneous uploads.
- Decompression bomb: header-declared dimensions are read via heic-decode's
  allocation-free `.all` path and rejected above 64MP BEFORE decode() can
  allocate width*height*4 bytes (a <300-byte crafted file can declare
  30000x30000 = 3.6GB). Regression-tested with a crafted ISOBMFF fixture
  against the real heic-decode WASM (test/heic-jpeg-core.test.ts).
- Mislabeled HEIC (documented Android/MIUI case): conversion now routes on
  ftyp magic-byte sniff of the raw buffer regardless of declared
  ext/Content-Type, so a HEIF uploaded as image/jpeg converts instead of
  415ing; the magic-mismatch 415 only fires for genuinely unrecognized bytes.
- Brand allowlist narrowed to what heic-decode's isHeic() accepts
  (heim/heis/hevm/hevs dropped — they could only ever fail conversion).
- Converted-output size: the JPEG result is checked against
  MAX_PASTE_IMAGE_BYTES (jpeg-js can inflate a within-limit HEIC past the cap).
- Deps: heic-convert replaced with its underlying heic-decode + jpeg-js
  (the wrapper could not expose the pre-decode dimension check); lockfile
  synced, drops pngjs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 18:06:59 +02:00
Codeman maintainer 13c877f938 fix(review): harden Codex generated-artifact attachment pipeline (PR #150)
- Pass the attachment request `source` through the server deps lambda and make
  it a required param on SessionListenerDeps.registerAttachment + the wiring
  event type (the 2-arg lambda silently dropped `source`, force-confining every
  codex-generated artifact — the feature never worked outside the workspace);
  new test/session-listener-wiring.test.ts asserts the pass-through
- Gate the Codex `Saved to: file://` scanner on mode === 'codex' via a
  codexArtifacts option threaded from the session call site; magic links stay
  mode-agnostic; tests assert claude/shell sessions never emit codex-generated
  requests
- Decide the generated-artifact trust policy on the realpath-RESOLVED path
  (unresolvable → force-confined) and anchor the ~/.codex marker dirs to
  os.homedir() prefixes with startsWith instead of substring matching; symlink
  escape + unanchored-marker regression tests added
- Run the Codex scanner on stripAnsi'd data so trailing SGR sequences don't
  ride into the captured URL; styled 'Saved to:' test added
- Extend generateFirstPageThumbnail with jpg/jpeg/gif/webp passthrough and
  per-extension content types (mirrors the png passthrough) so the PR's new
  image formats render real thumbnails instead of 204 letter-tiles

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 17:50:21 +02:00
Codeman maintainer 895edfedb0 fix(review): add Codex last-response test coverage + minor hardening (PR #152)
- Add test/routes/session-routes-codex-last-response.test.ts (app.inject +
  temp CODEX_HOME fixture rollouts): originator match beats cwd fallback when
  two panes share a dir, cwd fallback excludes sibling-claimed/foreign-cwd
  rollouts, resume-uuid filename match, history.jsonl pin outranks originator,
  event_msg/legacy user-turn dedup keeps old-codex turns, injected-context
  filtering, image placeholder, envelope shape ({success:true,data:{text,
  timestamp[,messages]}}), and a Claude-mode regression guard (codex reader
  never consulted for claude sessions)
- Replace clear-at-cap Map caches (codexHistoryPinCache, codexRolloutMetaCache)
  with the repo-standard LRUMap so a full cache wipe can't thrash hot entries
  on large rollout collections
- Join multi-block assistant/user text with a blank line instead of no
  separator (extractCodexBlockText)
- Re-enable the terminal-buffer eye fallback for shell sessions (they have no
  transcript source at all); TUI modes keep the clear placeholder

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 17:42:39 +02:00
Ark0N 3f23621f8d Merge pull request #144 from TeigenZhang/pr/mouse-restore-decset-strip
feat(web): restore tap/click/wheel mouse interaction when server strips mouse DECSETs
2026-07-12 13:23:22 +02:00
Ark0N 5cca965aa4 Merge pull request #143 from TeigenZhang/pr/cjk-input-loss
fix(mobile): CJK input loss — IME state machine, focus routing, and Android InputConnection recovery
2026-07-12 13:23:19 +02:00
Ark0N b74a904b41 Merge pull request #142 from TeigenZhang/fix/mobile-response-viewer-typography
fix(mobile): improve response-viewer readability on phones
2026-07-12 13:23:17 +02:00
Codeman maintainer 05d366e405 Merge PR #140 from crawlsys/feat/webgl-renderer-toggle: WebGL renderer toggle in settings
Includes review fixes (per-device setting + sticky-marker semantics); merged locally because the org-owned fork rejects maintainer pushes.
2026-07-12 13:22:53 +02:00
Ark0N 9204e42812 Merge pull request #139 from aakhter/cod-160-unified-session-service
Unified session list: backend service + endpoint
2026-07-12 13:22:39 +02:00
Ark0N a9749ead6a Merge pull request #138 from aakhter/cod-80-raise-terminal-defaults
Raise terminal history/scrollback/buffer defaults (50k→100k, 2MB→32MB)
2026-07-12 13:22:37 +02:00
Codeman maintainer e510ab74ca fix(review): use dvh fallback pair so the response-viewer header stays on-screen on iOS (PR #142)
- 92vh on iOS Safari measures the large viewport; with browser chrome visible the
  panel top (header + close button) clipped off-screen. 88vh fallback + 92dvh
  matches the repo's established dvh idiom.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 13:20:16 +02:00
Codeman maintainer 7fa52cdcd6 fix(review): make WebGL toggle per-device and fix sticky-marker semantics (PR #140)
- saveAppSettings no longer sends webglRendererEnabled on the settings PUT:
  the key is absent from the .strict() SettingsUpdateSchema, so every save
  400'd with INVALID_INPUT, silently killing all server-side settings
  persistence. Stripped in the per-device destructure alongside
  localEchoEnabled/skin/etc.
- shouldSkipWebGL now treats a stored true like the untouched default w.r.t.
  the sticky marker: the checkbox defaults checked on desktop, so any
  unrelated save stored true and every page load then cleared the
  'codeman-webgl-disabled' marker, permanently defeating the GPU-stall
  auto-fallback. Only ?webgl=force clears the marker at init.
- The marker is instead retired on a real OFF->ON toggle flip detected at
  save time (mirrors the _prevGestureEnabled pattern in settings-ui.js).
- webglRendererEnabled added to the displayKeys per-device set in
  loadAppSettingsFromServer (renderer choice is device/GPU-specific; syncing
  would leak mobile's hidden-checkbox false onto desktop).
- Tests: stored true + sticky marker -> still skips WebGL; OFF->ON save
  clears the marker and keeps the key off the wire; default-checked save
  leaves the marker alone; ?webgl=force / ?nowebgl behavior unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 12:42:36 +02:00
Codeman maintainer 5ea424565d fix(review): content-free IME traces, guarded onData self-heal, Android-only retap recovery (PR #143)
- BLOCKER (privacy): the CJK diagnostic trace logged typed CONTENT — _esc(e.key)
  per keystroke, up to 24 chars of textarea value on focus/blur/compstart/
  compend/input, and the flushed text — mirrored into _crashDiag, which
  persists to localStorage and beacons to POST /api/crash-diag. Traces are now
  content-free: key CLASS via _kdesc (any single code point → 'printable',
  named keys pass through), value lengths + phantom presence via _vdesc
  (len=N[+ph]), and 'flush send len=N'. _esc removed.
- MAJOR: the onData self-heal refocused the CJK field whenever gated data
  arrived with focus elsewhere — but onData also fires for xterm's
  SELF-GENERATED query replies (DA/DSR/CPR/OSC during Ink redraws), so it
  stole focus from rename/search/settings inputs while output streamed. Now
  requires document.activeElement === this.terminal.textarea (genuine typed
  input) and bails when shouldSuppressTerminalQueryResponse(data) matches.
- MAJOR: the pointerdown blur→setTimeout(focus,0) wedged-IME recovery ran on
  ALL platforms; on iOS tapping the focused empty field is normal and the
  async refocus is outside the user-gesture stack. The listener is now only
  registered when /Android/i.test(navigator.userAgent).
- tests: trace-privacy test (no typed character or textarea value ever appears
  in the trace; lengths/key classes still recorded), iOS harness asserts the
  pointerdown recovery never cycles, self-heal source guard asserts both new
  conditions; vm harness gained a ua option (navigator injected, Android UA
  default so the existing wedged-IME test still exercises the recovery).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 12:42:14 +02:00
Codeman maintainer 0ad673794f fix(review): dedupe resumed-session rows + newest-wins lifecycle name/mode (PR #139)
- Duplicate rows: transcript-history rows are keyed by the Claude
  conversation UUID (.jsonl filename stem), which diverges from the Codeman
  session id for resumed (claudeSessionId = resumeSessionId != id) and
  /clear-respawned sessions, so one conversation surfaced as both a live row
  and a history-only row. mergeUnifiedSessions now builds an alias map
  (claudeSessionId -> Codeman id) from the live + persisted views and
  resolves history/lifecycle keys through it; the route feeds
  SessionState.resumeSessionId as the persisted alias.
- Inverted precedence: SessionLifecycleLog.query() returns entries
  NEWEST-first, but the merge loop unconditionally overwrote name/mode so
  the OLDEST entry in the window won (stale rename/mode). First-seen now
  wins, mirroring the existing lastActivityAt guard.
- Tests: resumed session yields ONE row (service unit + route end-to-end
  with a real transcript fixture); renamed-then-deleted session surfaces
  the NEWEST name/mode. All 4 new tests fail against the pre-fix code.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 12:41:48 +02:00
Codeman maintainer 4d3080aacc fix(review): gate wheel forwarding by CLI mode/version, stop link click double-fire (PR #144)
- _shouldForwardWheelToApp: claude sessions forward wheel to the TUI only
  when the banner-parsed cliVersion is known AND >= 2.1.187 (older/unknown
  Claude Code captures wheel as select-menu navigation → keep local
  scrollLines); new dependency-free _cliVersionAtLeast semver-ish compare
- gemini excluded from wheel forwarding entirely (TUI wheel behavior
  unverified); codex keeps forwarding (verified); taps/clicks still
  forwarded for all strip modes
- link double-fire: registerFilePathLinkProvider links now track hover
  state via ILink hover/leave callbacks (_linkHovered) and
  _handleDesktopTerminalClick bails while a link is hovered, so a link
  click no longer also sends a synthetic SGR press/release to the TUI
- help modal: document Shift+Wheel (scroll local history when mouse
  passthrough is active)
- tests: version gate (2.1.186/unknown/garbage no forward, 2.1.187+
  forwards), codex/gemini split, link-hover click suppression

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 12:40:25 +02:00
Codeman maintainer 246f7b532d fix(review): clamp env-path trim below max; revert unwired scrollback raise (PR #138)
- UNBOUNDED-MEMORY: DEFAULT_TERMINAL_BUFFER_TRIM_BYTES from CODEMAN_TRIM_TERMINAL_TO
  had no relation to DEFAULT_TERMINAL_BUFFER_MAX_BYTES — setting only
  CODEMAN_MAX_TERMINAL_BUFFER=2097152 left the 24MB trim default in force, making
  BufferAccumulator.trim() (slice(-trimSize)) a no-op: unbounded growth past the cap
  plus a full string re-join on every append (O(n²)). Trim default is now clamped to
  75% of the resolved max (the 24MB/32MB default ratio, preserved as hysteresis);
  regression test re-evaluates the module under the env via vi.resetModules.
- OVERCLAIM: reverted DEFAULT_TERMINAL_SCROLLBACK_LINES 100k -> 50k — it has zero
  consumers; browser xterm scrollback is the separate hardcoded DEFAULT_SCROLLBACK
  (50k) in constants.js and deliberately stays 50k (mobile-memory hazard). The tmux
  history-limit raise (50k -> 100k) and PTY 32MB/24MB raise remain (those are wired).
  Module docstring now claims only what is wired; fixed the stale tmux-manager.ts
  comment saying the tmux limit "matches the xterm-side default in constants.js".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 12:39:46 +02:00
codeman-localandClaude Fable 5 116db81002 feat: response-viewer support for Codex sessions
The response-viewer (eye) currently reads only ~/.claude/projects — for
Codex panes it falls back to a raw terminal-buffer dump. This adds a
Codex-aware reader with exact per-pane rollout attribution.

Locating THIS pane's rollout (~/.codex/sessions/**), in confidence order:

1. history match — Session tracks the pane's last Enter
   (codexLastSubmitAt); correlating it against ~/.codex/history.jsonl
   {session_id, ts} entries identifies the thread the pane is ACTUALLY
   on, surviving /resume, /new and /fork typed inside the codex TUI.
   An entry is credited to the pane whose Enter is closest, so menu
   keystrokes in other panes can't steal attribution.
2. originator match — codex panes are spawned with
   CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_<sessionId>, which codex
   (verified on 0.144.1) writes into session_meta.originator of every
   rollout it creates.
3. resume-id match — resumed rollouts keep their original session_meta
   (codex appends without rewriting), but the uuid is in the filename.
4. cwd+mtime heuristic — case-blind compare (codex records launch-time
   path case) and rollouts claimed by other panes are excluded.

Reader details: user turns come from event_msg/user_message (real input
only — AGENTS.md / environment_context injections never appear there),
deduped against legacy response_item rows per-text so mixed-version
rollouts keep full history; image inputs render an [image xN]
placeholder; session_meta identity is cached per path (write-once).

Frontend: thread role label follows session mode (Codex/Gemini/
OpenCode); the terminal-buffer fallback is Claude-only — TUI modes show
a clear placeholder instead of a repaint dump.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 10:39:42 +08:00
Aamer AkhterandSaqeb Akhter bb1d16e230 feat(image): convert HEIC paste uploads to JPEG
When a browser pastes an HEIC file without normalising it first, the
paste-image route now converts it to JPEG server-side via heic-convert
before writing to .claude-images/. Magic-byte validation confirms the
output is valid JPEG. Adds type declarations for the heic-convert package.

Co-authored-by: Saqeb Akhter <saqeb.akhter@gmail.com>
2026-07-11 16:32:49 -04:00
Saqeb Akhter 978ca57343 fix: COD-152 preserve generated artifact filenames 2026-07-10 20:58:47 -04:00
Saqeb Akhter f8aa93969b fix: COD-152 surface Codex generated artifacts 2026-07-10 20:54:41 -04:00
Aamer Akhter 584910f645 COD-144 flush queued SSE output on empty buffer-load so new shells paint immediately
A freshly created shell session rendered blank until a tab-switch. selectSession()
fetches the terminal buffer, but for a just-started shell that fetch resolves before
the PTY emits its prompt, so the buffer is empty; the prompt then arrives as a live
SSE event queued during the load and _finishBufferLoad() discarded it. The discard is
correct for an established session (its fetched buffer already contains that output),
but harmful when the load painted nothing.

_finishBufferLoad(owner, { flushQueued }) now REPLAYS the queued events through
batchTerminalWrite (after _isLoadingBuffer is cleared, so they write through, not
re-queue) instead of discarding. selectSession passes flushQueued only in the empty
branch (no fresh buffer + no cache), so the established-session de-dup path is
unchanged. TDD: test/terminal-buffer-flush.test.ts exercises the real begin/finish
mixin (vm-harness, no jsdom).
2026-07-10 14:39:19 -04:00
Aamer Akhter b86b132af5 COD-136 skip redundant connection-indicator DOM writes on hot input path
_updateConnectionIndicator() ran on every keystroke (_reliableSend) and
every ACK (_ackDelivery), unconditionally writing display/className/
textContent/title. During fast typing the rendered output is usually
identical between calls, so those were wasted main-thread DOM writes.

Extracted the branch logic into a pure DOM-free _computeConnectionDescriptor()
returning { display, dotClass, text, title } (every branch/string preserved
verbatim; hidden state normalizes the three non-display fields to '' so the
compare is well-defined). _updateConnectionIndicator() now computes the
descriptor, compares all four fields against a cached _lastIndicatorDescriptor,
and early-returns when unchanged — otherwise caches and writes the DOM exactly
as before (display always; dotClass/text/title only when shown). First call
renders (cache starts null). Perf only, no behavior change.

Tests: test/connection-indicator.test.ts — 9 descriptor cases pinning the
exact strings per state + 4 skip cases (first call writes; two identical calls
write DOM once via counting setters; state change and hidden->shown re-render).
31/31 with input-send-order regression; build, frontend-syntax, prettier clean.
2026-07-10 14:39:09 -04:00
Aamer Akhter 4ab89f9a4e COD-137 scope WS per-session limit by clientId (fix spurious 4008 on reconnect)
MAX_WS_PER_SESSION was gated by a bare Map<sessionId,number> counter,
incremented on upgrade and decremented only on the old socket's async
close. A client that dropped and immediately reconnected could land its
new upgrade before the old socket's close fired, briefly over-counting and
tripping a spurious 4008 (-> HTTP fallback). The limit also counted raw
sockets, so a reconnecting client consumed a new slot instead of its own.

Replace the counter with WsConnectionRegistry (new pure, unit-tested module)
that tracks live sockets per session keyed by clientId. A same-cid upgrade
SUPERSEDES its own socket (evicts the stale one with close 4010, reuses the
slot, no net count change) -> a reconnect can never be rejected by the cap.
The reliable-input protocol (shouldApplyInput(cid,seq)) already assumes one
logical client per cid per session, so same-cid eviction is principled, not
a regression of multi-tab (which already collides on seq). Slots are freed
EAGERLY on error/terminate, not just async close; close is identity-matched
so a superseded socket's late close is a no-op. cid-less upgrades are
admitted anonymously up to the cap and never evict (backward-compat).
Client sends cid on the WS upgrade URL (?cid=, encoded, omitted if absent).

Tests: ws-connection-registry.test.ts (reconnect-reclaim at cap, rejects
N+1th distinct, eager-terminate frees slot, cid-less up-to-limit + no-evict,
late-close-no-evict, per-session isolation) + route integration in
ws-routes.test.ts (real upgrade through the cap). 45/45 across registry +
ws-routes + input-send-order + ws-reconnect-plan; tsc 0, build, prettier,
frontend-syntax clean.
2026-07-10 14:39:01 -04:00
Aamer Akhter 20cb42d202 COD-135 re-drive lost input ACK on a live WebSocket (durable-delivery gap)
A reliable-input frame could be stranded forever if its server ACK
({t:'ia',seq}) was lost while the WebSocket kept delivering other output.
_drainSession's WS fast path skips records with sentAt!==0, and after
COD-134 the sweep only force-closes a *silent* socket -- so a lost ACK on
an otherwise-live socket (stale && !silent) was never re-sent.

_redeliverSweep now, for an active-WS session whose oldest unacked frame
is stale but the socket is NOT silent, resets sentAt=0 on every stale
unacked frame and lets the existing _drainSession re-drive them over the
live socket (server dedups by seq). The stale && silent force-close
remains the fallback for a genuinely half-open socket. Restores the
exactly-once recovery guarantee without reintroducing the flap.

Tests: new failing-first COD-135 cases in test/input-send-order.test.ts
(re-drive on live socket; leave not-yet-stale alone; keep stale+silent
force-close). 18/18 across input-send-order + reliable-input-dedup +
ws-reconnect-plan; tsc 0, frontend-syntax, build all clean.
2026-07-10 14:38:52 -04:00
Aamer Akhter 68fd6e8962 COD-134 fix WS flap loop (undefined onopen call) + reconnect resilience + logging
Root cause of the WS->HTTP->WS flapping: the v1.1.15 input-delivery merge left a
call to the now-undefined _flushHttpFallbackQueuesViaWs() in ws.onopen, so every
(re)connect threw a TypeError BEFORE _onWsReady() ran -- durable input was never
re-flushed over the fresh socket, the 2s redeliver sweep then saw stale unacked
frames and force-closed the socket, reconnect, throw again: a self-sustaining
flap loop. Remove the dead call (_onWsReady, 10 lines below, is its replacement).

Resilience + observability:
- Pure CodemanWsReconnect.plan(code, attempt) (constants.js, TDD, 6 tests):
  <4004 -> fast reconnect (immediate jittered first retry, faster backoff);
  4008/unknown->=4004 -> bounded retry-fallback (HTTP no longer sticks until a
  tab switch); 4004/4009 -> give up (session gone). Wired into onclose.
- Redeliver sweep force-closes only a SILENT socket (no recent recv), not one
  actively delivering output/ACKs -- stops self-inflicted flaps while typing.
- Client logs WS close code/reason to crash-diag; server logs [ws]
  open/close/terminate/4008 (console -> journald; Fastify runs logger:false).

Verified: 6/6 unit, tsc 0, frontend-syntax + prettier clean, build; beta WS
reaches connected with zero console errors (onopen TypeError gone),
_wsLastRecvAt tracked, server [ws] lines emit.
2026-07-10 14:38:48 -04:00
Aamer Akhter be4fecdad5 COD-133 fix header WS status indicator + typing lag from v1.1.15 merge
The upstream v1.1.15 merge spliced upstream's transport-object indicator
body onto local's _connectionStatus-based _updateConnectionIndicator()
without defining `transport`, so every transport.* reference threw
ReferenceError on any queued state. That hid the "WS" status and, because
_reliableSend() updates the indicator before _drainSession(), made every
keystroke skip immediate delivery (input flushed only on the 2s sweep =
typing lag).

- Rewrite _updateConnectionIndicator() to show the terminal WebSocket
  transport from _wsState (WS / HTTP / WS… / Offline), falling back to the
  SSE _connectionStatus only on the idle dashboard.
- Only annotate a backlog (· N queued) above 4 bytes so normal typing no
  longer flickers "sending 1B" on each key press.
- test/connection-indicator.test.ts (new): transport labels, the >4B
  threshold, an exhaustive never-throws guard for the ReferenceError, and
  the _reliableSend -> _drainSession invariant (typing-lag guard).
- test/input-send-order.test.ts: reconcile to local's durable input layer
  (the prior coalescing-fallback tests had been failing since 1255e28).
2026-07-10 14:36:21 -04:00
Aamer Akhter c7967d4b55 COD-138 normalize shell scrollback to CRLF so replay doesn't staircase
A shell terminal could render output diagonally (each line shifted one
column right) after a full page reload or a cursor-query-failure replay.

Root cause: capturePaneBuffer's full-history path (capture-pane -p -e -S -)
and its cursor-query-failure fallback returned raw scrollback, which tmux
joins with a BARE \n. The browser xterm uses convertEol:false (correct for
the live PTY stream, which carries real \r\n), so each bare \n dropped a
row without returning the cursor to column 0 -> staircase. The visible /
tab-switch path (formatPaneSnapshot) was immune because it repaints each
row with an absolute cursor CSI.

Fix: new pure helper normalizeScrollbackEol() (\r?\n -> \r\n, idempotent
on CRLF, leaves lone \r overwrites untouched, adds/removes no rows) applied
at both raw-return seams. The absolute-positioned snapshot path is unchanged.

Tests: test/tmux-scrollback-eol.test.ts pins the invariant (no LF without a
preceding CR) + CRLF idempotency + lone-CR preservation. 136/136 across
tmux-scrollback-eol + tmux-capture-full-history + tmux-manager +
routes/session-routes; build, tsc, prettier, frontend-syntax clean.
2026-07-10 11:14:26 -04:00
Aamer Akhter 8be83cd585 COD-47 replay full tmux scrollback on terminal reload
A full page reload (GET /api/sessions/:id/terminal with no ?tail=) now captures
the ENTIRE tmux scrollback via capture-pane -p -e -S -, so users get back history
that scrolled off Codeman's byte buffer. Tab switches (?tail=N) keep the fast
visible-frame capture.

- tmux-manager capturePaneBuffer/captureActivePaneBuffer take { fullHistory }:
  full-history returns raw linear scrollback (skips the single-screen
  formatPaneSnapshot repaint, which would clip multi-screen history).
- /terminal selects full-history on full reload, visible on tail; caps the
  payload at the configured terminalBufferMaxBytes (keeps most-recent bytes,
  line-aligned) and returns source/fullSize/truncated metadata.

Verified: tsc 0, tmux-capture-full-history 5/5, session-routes 68/68.
Caveat: lines tmux already evicted past its history-limit can't be recovered.
2026-07-10 11:08:57 -04:00
Aamer Akhter 09fd1e495f COD-118 fix(test): stub resetRespawnBreaker in MockSession
POST /api/sessions/:id/interactive calls session.resetRespawnBreaker()
before startInteractive(); mock missing the stub → route threw → 422.
2026-07-09 16:00:06 -04:00
Aamer Akhter 7efc6cd5a8 COD-94 fix(lint): remove useless escapes in jump-host regex character class
\[ inside [...] doesn't need backslash — ESLint no-useless-escape.
2026-07-09 12:18:21 -04:00
Aamer AkhterandClaude Opus 4.8 286cf0768d COD-118 feat: circuit breaker bounding repeated non-zero interactive-PTY exits
Defense-in-depth after COD-115. If the interactive PTY exits non-zero
repeatedly within a short window, recovery/reconnect paths recreate it
indefinitely (COD-115 saw 114 'exited with code: 1' events + orphans).

- New pure InteractivePtyExitBreaker (session-pty-exit-breaker.ts):
  injectable time, sliding window, clean-exit resets counter, stays
  tripped until reset(). Defaults: threshold 5, window 10s.
- Session records each interactive PTY exit in the breaker; on trip it
  flips _status to 'error', sets _respawnBlocked, emits
  respawnBreakerTripped. startInteractive() refuses to respawn while
  blocked, so all recovery/reconnect callers stop looping uniformly.
- Explicit user restart (POST /api/sessions/:id/interactive) calls
  resetRespawnBreaker() so intentional restarts are never blocked.
- New SSE event session:respawnBreakerTripped wired in sse-events.ts +
  constants.js (registries in sync) + session-listener-wiring.ts;
  minimal diagnostic toast in app.js.
- Tests: test/respawn-pty-breaker.test.ts (pure trip/reset/window +
  MockSession session-level trip/reset).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 11:49:23 -04:00
Aamer AkhterandClaude Opus 4.8 3c0e6286f6 COD-115 fix: scrub TMUX/TMUX_PANE so tmux-backed sessions don't crash-loop
When the web server is launched from inside a tmux pane it inherits TMUX/
TMUX_PANE. tmux's nesting guard then makes every new attach-bridge PTY
(`tmux attach-session`, used by codex/opencode/gemini and mux-wrapped claude)
exit code 1; the respawn controller recreates the dead bridge → infinite loop.

The existing guard in buildMuxAttachEnv() used `TMUX: undefined` on a
{...process.env} spread, which leaves the KEY present with value undefined —
node-pty serializes it as the literal string "TMUX=undefined", still tripping
the guard. (The working create path in tmux-manager.ts uses `delete`.)

Fix:
- Primary: delete process.env.TMUX / TMUX_PANE at web bootstrap (src/index.ts)
  so every downstream {...process.env} spread is clean regardless of launch
  context. `delete`, not `= undefined`.
- buildMuxAttachEnv(): build a copy and `delete` TMUX/TMUX_PANE/CLAUDECODE
  (and COLORTERM when not truecolor) instead of `: undefined` — same node-pty
  quirk affected all of them.
- Test: assert the keys are genuinely ABSENT (`'TMUX' in env === false`), not
  merely undefined — the prior test only checked `toBeUndefined()`, which is
  why the bug slipped through. Red→green confirmed.

Verified on isolated beta launched from inside tmux (inherited the poisonous
TMUX=codeman,980,7): created a codex session + triggered interactive attach —
the bridge `tmux -L codeman-beta attach-session` spawned with NO TMUX in its
env, attached successfully, zero "exited with code: 1", server healthy.

Circuit-breaker for repeated non-zero bridge exits (AC bullet 4, optional)
split to a follow-up. Deploy-pending (substrate): never auto-deployed.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 11:47:38 -04:00
Aamer AkhterandClaude Sonnet 4.6 8a133d083b fix: COD-163 implementation gaps — shortcut overlay, settings tab, remote-case shell
- app.js: getShortcutRegistry()/matchesShortcutEvent()/showShortcutOverlay()/
  renderShortcutOverlay()/closeShortcutOverlay() (needed for DEFAULT_SHORTCUTS
  action dispatch + shortcut-registry-overlay tests)
- settings-ui.js: renderShortcutSettingsList()/startShortcutCapture()/
  onShortcutCaptureKeydown()/resetShortcutOverride()/toggleShortcutEnabled()
  (Settings → Shortcuts tab, needed for shortcut-registry-overlay tests)
- index.html: Shortcuts modal tab + shortcut overlay modal; remove Ctrl+Enter
  hint text (help-modal-shortcuts test asserts absence)
- session-ui.js: remote-case detection in runShell() (caseName vs workingDir);
  saveLastUsedCase after deleting selected case
- test/command-palette-ui.test.ts: expect browse-sessions item (COD-192 adds it)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-09 11:26:24 -04:00
Aamer AkhterandClaude Sonnet 4.6 a0e26db1dc feat: COD-192 add "Browse all sessions" escape hatch to command palette
Pins a "Browse all sessions…" item at the bottom of the command palette
list (after "New session"). Activating it closes the palette and opens
the Session Manager modal, bridging the gap between the fast in-memory
switcher and the full server-side history browser.

The item gets a distinct visual treatment (≡ icon, muted title/icon
color, 4px top gap) so it reads as a secondary action separate from the
primary session rows.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-09 11:11:48 -04:00
Saqeb Akhter 596899e19b fix: COD-157 refine shortcut palette labels 2026-07-09 11:11:33 -04:00
Saqeb Akhter e8f5ac94f3 fix: COD-153 preserve matched case selection 2026-07-09 11:09:00 -04:00
Saqeb Akhter 03192d9980 fix: COD-153 match new session case from palette query 2026-07-09 11:08:53 -04:00
Saqeb Akhter 3d4444ad78 fix: COD-153 guard command palette escape close 2026-07-09 11:08:48 -04:00
Saqeb Akhter c45e456b0e fix: COD-153 support terminal-focused command palette shortcuts 2026-07-09 11:08:23 -04:00
Saqeb Akhter ad25e234f4 feat: COD-153 add command-k session palette 2026-07-09 11:08:17 -04:00
Saqeb Akhter 48fd2da6ce fix: COD-151 launch case picker selection on enter 2026-07-09 11:01:45 -04:00
Saqeb Akhter e29721046c feat: COD-151 add searchable case picker 2026-07-09 11:01:31 -04:00
Aamer AkhterandClaude Sonnet 4.6 3a03792009 fix: resolve cherry-pick conflicts for COD-24/COD-107 remote host integration
- src/remote-hosts.ts: add missing execAsync = promisify(exec) that was
  implied by intermediate commits not in the cherry-pick set
- src/web/routes/session-routes.ts: add getDataDir import and
  readRemoteCases/readRemoteHosts/toSessionRemote for remote case support
  in quick-start; narrow casePath string|null via resolvedCasePath cast
- test/routes/session-routes.test.ts: add vi.hoisted remoteStore mock for
  remote-hosts.js; fix 'creates session from remote case' test to use
  /api/quick-start (remote cases are not supported on /api/sessions)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-09 09:35:14 -04:00
Aamer AkhterandClaude Opus 4.8 e83ff72b61 COD-107 fix: shellescape -J jumpHost + structural validator (close command-injection)
buildSshConnectionArgs interpolated jumpHost raw while its siblings
(identityFile/socksProxy/extraSshOptions) were shellescaped. The token array is
joined and run via execAsync (/bin/sh -c), so a jumpHost like "x; touch /tmp/pwned"
executed. The Zod denylist only blocked backtick/newline/$( and let ;|& and spaces
through.

- shellescape jumpHost in buildSshConnectionArgs (primary fix)
- replace jumpHost denylist with a structural allowlist: [user@]host[:port],
  comma-separated multi-hop, bracketed IPv6; no shell metachar can appear
- update/extend tests: escaped -J assertion + injection-safety case

Verified: remote-ssh-options (11) + case-routes (33) pass, tsc --noEmit clean,
regex accepts valid forms / rejects 8 injection payloads.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-09 09:17:02 -04:00
Aamer Akhter 268a0bbdbd COD-107 remote SSH: custom port + advanced connection options (escape hatch)
The Remote case form could only reach port-22, default-identity, directly
SSH-able hosts. Add an escape-hatch set of SSH connection options so Codeman
can reach a host like aa-desktop (custom port 2222, ed25519 identity, cloudflared
SOCKS5 ProxyCommand) the way ssh-aa-desktop does — without shelling out to that
wrapper.

- Model (types/session.ts): new optional RemoteSshOptions (identityFile,
  socksProxy, jumpHost, extraSshOptions) on RemoteHost AND SessionRemote; all
  absent = today's behavior. toSessionRemote() carries them case->session.
- Shared buildSshConnectionArgs(remote) in remote-hosts.ts: pure, exported,
  ordered ssh connection tokens (-o BatchMode=yes, -p, -i <abs identity with
  ~/$HOME expanded + shellescaped>, -J, -o ProxyCommand=nc -X 5 -x <socks>
  %h %p emitted as ONE shellescaped token so %h %p reach ssh literally, then
  each extraSshOptions -o). Both buildRemoteLaunchCommand (tmux-manager.ts) and
  buildRemoteTmuxCheckCommand now use it, so the prereq probe and the real
  launch connect identically. checkRemoteTmuxAvailable widened to accept the
  options (callers already pass the full host).
- Validation (schemas.ts): identityFile (no newline/NUL), socksProxy
  (host:port), jumpHost (no shell metachars), extraSshOptions (KEY=VALUE,
  reject newline/NUL/backtick/$() — defense-in-depth on operator-entered config.
- UI (index.html + session-ui.js): SSH Port field + collapsible "Advanced SSH"
  section (identity, SOCKS proxy, jump host, extra -o options one per line);
  wired into the remote-host create payload.

Empty-options remotes emit byte-identical ssh to before (pinned by test).

Tests: test/remote-ssh-options.test.ts (buildSshConnectionArgs +
buildRemoteLaunchCommand + buildRemoteTmuxCheckCommand for the aa-desktop set,
escaping/%h %p/identity-~ expansion, byte-identical back-compat); case-routes
schema tests (advanced options round-trip; malformed extraSshOptions/socksProxy
rejected). tsc/eslint/frontend-syntax/prettier/build clean.

Acceptance (real remote, no wrapper): the emitted command connected to
aa-desktop through the cloudflared SOCKS proxy and created a durable remote
tmux session (verified independently via ssh-aa-desktop: CONNECTED_NO_WRAPPER,
STILL_ALIVE_AFTER_DETACH); checkRemoteTmuxAvailable over the proxy returned
{ok:true, tmuxPath:/usr/local/bin/tmux}; test session cleaned up.
2026-07-09 09:16:57 -04:00
Saqeb Akhter 26e78daf58 fix: COD-24 stabilize remote host sessions 2026-07-09 09:06:06 -04:00
Saqeb Akhter 568d93efb0 feat: add remote host case routes 2026-07-09 08:51:12 -04:00
Saqeb Akhter 3bf991d730 feat: add remote host case domain 2026-07-09 08:51:08 -04:00
Teigen 7fb58648ba feat(web): forward wheel + guard clicks for desktop stripped-mouse sessions
Desktop click-to-position-cursor died under the server's mouse-DECSET strip
(same root cause as the mobile touchend tap regression): xterm's native mouse
encoder only emits SGR while mouseTrackingMode is ON, but the server strips the
enabling DECSETs from claude/codex/gemini output. Hand-encode the report for
plain left-clicks (_handleDesktopTerminalClick), skipping every click that
already means something else (synthetic/compat, modified, double/triple,
drag-selection, off-grid, xterm encoder live).

Also widen forwarding to the wheel: Claude Code 2.1.187+ scrolls its own
transcript on SGR wheel reports and no longer captures wheel as select-menu
navigation (verified against 2.1.202), so forward the wheel to the TUI for
strip-mode sessions at the buffer bottom (40ms-coalesced to avoid a tmux
send-keys storm). Shift+wheel and any scrolled-up viewport stay on xterm's
local scrollback. Guard synthetic taps/clicks on viewport-at-bottom so a
scrolled-up report can't hit-test the wrong row.

Tests: 12 cases in test/terminal-touch-tap.test.ts. Verified E2E via Playwright
against the live instance (wheel up/down forward, Shift+wheel local, click).
2026-07-07 17:52:32 +08:00
Teigen 9535edc367 fix(mobile): restore tap-to-position cursor after master merge — hand-encode SGR when server strips mouse DECSETs
v1.1.7 (3172bef, arrived via the master merge) strips mouse-tracking DECSET
sequences from claude/codex/gemini output so the wheel keeps scrolling
scrollback. Side effect: the browser xterm's mouseTrackingMode is permanently
'none' for those sessions, and the mobile touchend tap branch gates its
synthetic click on exactly that mode — so tap-to-position-cursor silently died.

Fix: when tracking reads 'none' but the session mode is one the server strips
(claude/codex/gemini — the PTY-side TUI still has tracking ON), encode the SGR
press+release report directly from the touch point and send it to the PTY,
bypassing xterm's mouse encoder. No DOM click is dispatched, so xterm's local
selection cannot trigger either.

Tests: 3 new cases in test/terminal-touch-tap.test.ts (SGR encoding, grid
clamping, shell-mode exclusion); verified E2E via Playwright iPhone emulation
against both a stripped-stream instance and the production bundle.
2026-07-07 17:52:32 +08:00
Teigen 443b85c18e fix(mobile): CJK input loss — IME state machine, focus routing, and Android InputConnection recovery
Three independent root causes of intermittent Chinese character loss
(English was unaffected because it bypasses the composition path):

1. input-cjk.js state machine: stuck _composing when compositionend never
   fires (WeChat/Sogou IMEs) silently swallowed all input; the deferred
   compositionend flush could reset the textarea mid-next-composition
   (cancels the live IME composition on iOS); the 100ms keydown-echo
   window discarded ANY input regardless of content.

2. Focus stealing: session-select / SSE-reconnect paths call
   terminal.focus() (15+ call sites), landing focus on xterm's hidden
   textarea; with the CJK onData gate active, everything typed there was
   swallowed. Fix: focus router in initTerminal routes ALL
   terminal.focus() calls to the CJK field while it is visible, plus a
   self-healing onData gate that reclaims focus when it swallows input.

3. Android InputConnection wedge (9-key IMEs + Chromium): the keyboard
   composes in its own UI but delivers zero DOM events. Fix: skip
   redundant textarea value/selection writes (they race IME session
   setup), and re-tapping the focused empty field forces a blur→focus
   cycle that restarts the input session.

Diagnostics: input-cjk.js now traces every IME event/flush decision into
the crash-diag breadcrumbs; /api/crash-diag stores beacons per page-load
id (iOS PWA reloads no longer wipe the trail, concurrent clients no
longer clobber each other) and flushes on visibilitychange.

Tests: test/input-cjk.test.ts (vm-sandbox, 9 cases incl. regression
guards for all three root causes).
2026-07-07 17:52:10 +08:00
Teigen 66eaaf0da3 fix(mobile): improve response-viewer readability on phones
The mobile media query only overrode .response-viewer-body with a flat
font-size: 12px / padding: 12px, leaving the desktop response-viewer
typography system (--rv-content-max, .rv-text pre, heading scale) with no
mobile tuning. Bump body text to 14.5px/1.65, give code blocks phone-sized
padding and 11.5px code, scale headings (h1 1.35em / h2 1.2em / h3 1.08em),
let content span full width, and cap the panel at 92vh.

Layers cleanly on top of the existing response-viewer selectors in
styles.css; desktop rendering is unchanged.
2026-07-06 10:36:57 +08:00
Aamer Akhter ce4c5dd584 test(types): stop asserting Date.now() timestamps are deep-equal
createInitialRalphTrackerState() stamps lastActivity: Date.now(). The
'should create fresh instances each time' test deep-equaled two factory
results, so two calls straddling a millisecond boundary differed by 1ms
and failed intermittently (e.g. PR #139 CI: 1782927694581 vs ...580).

Exclude the dynamic lastActivity from the equality check and assert it
is a number separately, preserving the test's intent (distinct instances
with identical initial field values) without the timing race.
2026-07-05 21:49:31 -04:00
KrisandClaude Opus 4.8 e77af21107 docs(cron): add cron user guide + Claude speedrun protocol
- docs/cron-guide.md: comprehensive user/operator guide for the Cron
  feature (fields, schedule types, prompt security, execution flow, API,
  SSE, limits, troubleshooting), sourced from the implementation.
- SPEEDRUN.md: fast-execution protocol for Claude grounded in this repo's
  real commands and CLAUDE.md guardrails.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SSVnYek4nq4Ztmbb3SrCA5
2026-07-04 16:52:05 +05:30
DennisandClaude Opus 4.8 a842f2db4d fix(auth): re-issue session cookie on each request (sliding expiry)
The codeman_session cookie was only set on the Basic Auth path with a fixed
lifetime from login and never refreshed, while the server-side session store
slides its TTL (refreshOnGet). So the browser cookie expired mid-use, the next
request arrived cookie-less and fell through to Basic Auth, popping the native
username/password dialog — perceived as a random logout while actively working.

Re-issue the cookie on every authenticated (valid-cookie) request so the browser
lifetime tracks the server-side sliding TTL.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 18:24:24 +00:00
Kevin Crawley bf36eb0db4 feat(terminal): add WebGL renderer toggle in settings
Adds a 'WebGL Renderer' toggle to Settings > Appearance (desktop). WebGL
stays on by default; users can turn it off to force the DOM renderer when
they hit GPU glitches, without needing the ?nowebgl URL param. Explicit
opt-in (or ?webgl=force) clears a stale auto-fallback marker. Mobile skip
and the long-task auto-fallback safety net are unchanged.

The device/param/sticky/pref interaction is factored into a pure,
unit-tested shouldSkipWebGL() helper in constants.js.
2026-07-01 20:41:38 -05:00
Aamer Akhter 4dfdbcd100 COD-160 unified session list: backend service + endpoint
First increment of the read-only "complete + searchable session list".

- New src/services/unified-session-service.ts: mergeUnifiedSessions() combines
  live + persisted (state.json) + lifecycle + ~/.claude transcript history + mux
  stats into one list de-duped by sessionId, with precedence
  history < lifecycle < persisted < live, a meaningfulness floor that drops bare
  lifecycle/mux-only noise, and a stable newest-first sort. Plus
  filterAndPaginate() (case-insensitive q over name/firstPrompt/workingDir/
  sessionId; total before paging; limit clamped [1,500]). No IO — unit-testable.
- New GET /api/sessions/unified in session-routes.ts: gathers the five sources
  from ctx (sessions/store/lifecycle/scanProjectDir/mux, each try/caught), feeds
  the pure service, returns { sessions, total } (ApiResponse envelope). testMode
  short-circuits to empty.

Tests: unified-session-service.test.ts (12, pure) + unified-sessions-routes.test.ts
(4, app.inject).
2026-07-01 13:25:16 -04:00
Aamer Akhter ad71a92f29 COD-80 raise terminal history/scrollback/buffer defaults
Bump the centralized terminal-history defaults: tmux scrollback 50k->100k and
PTY buffer cap 2MB->32MB (trim 1.5MB->24MB). Both remain env/settings overridable
and bounds-clamped. Worst-case 20-session buffer budget rises 40MB->640MB.
Stacked on the terminal-history config commit.
2026-07-01 13:00:08 -04:00
Codeman maintainer 1fa88cd187 chore: version packages
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-01 09:08:33 +02:00
Ark0N 613eb25302 Merge PR #137: Centralize terminal history/scrollback/buffer limits into config (COD-80)
Introduces src/config/terminal-history.ts as the single source of truth for terminal scrollback lines, tmux history-limit, and PTY buffer byte caps. Behavior-neutral: defaults match prior hardcoded values; env overrides preserved. tmuxHistoryLimit is wired live (setHistoryLimit + respawn re-apply); the other three keys are scaffolding for a stacked follow-up. Reviewed: CI green (typecheck/lint + full test suite).
2026-07-01 09:06:15 +02:00
Aamer Akhter 8c0c94540c COD-80 centralize terminal history/scrollback/buffer limits into config
Introduce src/config/terminal-history.ts: one place for terminal scrollback,
tmux history-limit, and PTY buffer byte caps, each overridable via env var or
the settings object and bounds-clamped via resolveTerminalHistoryConfig().
Defaults match the prior hardcoded values, so this is behavior-neutral. Wires
the resolver through buffer-limits, tmux-manager (incl. a setHistoryLimit so a
settings change applies live), session, server, system-routes, session-routes,
schemas, and the config port. Adds 4 optional settings keys (terminalScrollback
Lines, tmuxHistoryLimit, terminalBufferMaxBytes, terminalBufferTrimBytes) with
bounds + a trim<=max cross-check.
2026-06-30 19:52:49 -04:00
KrisandClaude Opus 4.8 d9c2c6420d fix(cron): harden cron fixes against adversarial-review findings
Two blind adversarial reviewers found real holes in the prior cron commits:

SECURITY (was CRITICAL): the prompt-file guard was blocklist-only by default,
so promptFilePath:/proc/self/environ leaked the SERVER PROCESS's entire
environment (every secret) into the agent session, and /dev/zero or a FIFO
caused an unbounded readFile → OOM/hang DoS. A denylist is the wrong posture
for an exfil-into-LLM sink. resolveSafePromptPath now:
  - confines the realpath-resolved file to the job's working dir (ALLOWLIST) —
    closes /proc, /dev, other homes, modern cloud-cred paths, and symlink escapes
  - requires a regular file (rejects dirs/FIFOs/char devices)
  - caps the read at MAX_PROMPT_FILE_BYTES (1 MiB)
  - keeps the /etc,/root,secrets blocklist as defense-in-depth

LOGIC:
  - once-rearm (was MED, defeated in prod): the edit UI round-trips the full
    job, so the field-PRESENCE re-arm check always fired → a renamed fired
    once-job could be resurrected via edit→re-enable. Now compares schedule
    VALUES; an unchanged schedule never re-arms.
  - skipped-run history (was HIGH): recording a skip every tick was unbounded
    state.json growth. Now coalesces consecutive skips (one record per streak)
    and prunes global run history to MAX_CRON_RUN_HISTORY (500), covering the
    launch path too.
  - skip bookkeeping (was MED): a skip no longer advances lastRunAt (nothing
    ran); lastStatus still reflects 'skipped'.

Regression tests added/updated (43 pass): /proc/self/environ + outside-workspace
+ symlink-escape + non-regular + oversized all blocked, in-workspace file
passes; UI-path once resurrection blocked; consecutive skips coalesce to one
record; skip leaves lastRunAt null.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
2026-06-29 12:05:04 +05:30
KrisandClaude Opus 4.8 6082bceee6 fix(cron): close three MED cron-job defects
1. once-rearm on edit: editing any field of a finished one-time job reset
   completedOnce, silently resurrecting it. Now only a SCHEDULE edit
   (scheduleType/runAt/interval/daily/weekly) re-arms a completed once job;
   cosmetic edits (rename/notes) leave completedOnce intact.

2. update-validation gap: CronJobUpdateSchema = .partial() drops the cross-field
   superRefine, so a PUT switching scheduleType without its dependent field
   produced a dead enabled job (nextRunAt:null). updateJob now re-validates the
   MERGED job against the full CronJobSchema and throws 400 on inconsistency,
   leaving the stored job untouched.

3. concurrency-skip silent starvation: skip_if_same_agent_running advanced the
   schedule but wrote no run record, so a perpetually-skipped job had empty
   history. Now records a 'skipped' run (new CronJobRunStatus) + lastStatus.

Tests updated/added in cron-service.test.ts (37 pass): once non-schedule edit
preserves completedOnce, schedule edit re-arms, inconsistent partial update is
rejected with the stored job untouched, and the skip path records a skipped run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
2026-06-29 11:50:42 +05:30
KrisandClaude Opus 4.8 40e26c5422 fix(cron): confine cron prompt-file reads to block sensitive paths
A cron job's promptFilePath is user-supplied via the API and was read with an
unconfined readFile of any absolute path, so a hostile job config could exfil
arbitrary host files (e.g. /etc/passwd, SSH keys) into a Claude session.

Guard the read in resolvePrompt by mirroring the attachment-serving guard
(resolveServableAttachmentPath in file-routes): realpath-resolve the path, then
reject via the shared blocklist (/etc, /root, secret locations) plus the
optional workspace-confinement toggle before reading.

Regression tests in cron-service.test.ts: blocks /etc/passwd (the live repro)
and /root/*, fails cleanly on a missing file, and still allows an ordinary
prompt file outside the blocklist.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
2026-06-29 11:38:51 +05:30
KrisandClaude Opus 4.8 9feaa0d6e5 refactor(cron): rename scheduler feature to cron
Rename the recurring-jobs feature scheduler->cron to disambiguate from the
legacy ScheduledRun system (/api/scheduled), which is left untouched:

- ScheduledJob->CronJob, SchedulerService->CronService
- /api/scheduler/jobs -> /api/cron/jobs; SSE scheduler:* -> cron:*
- state keys cronJobs/cronJobRuns
- files moved to src/cron/, cron-routes.ts, cron-port.ts, types/cron.ts
- frontend cron-ui.js, #cronModal, menu "Cron"
- docs moved to docs/cron-discovery.md + docs/cron-build-brief.md, README guides
- new tests: cron-service.test.ts, cron-time.test.ts

Green: tsc, lint, frontend-syntax, format, 30 cron + 9 legacy scheduled-runs tests.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PmvZR12aX2v8K7YhqxPUAU
2026-06-29 11:35:56 +05:30
KrisandClaude Opus 4.8 2d2f4e592b feat(scheduler): add Scheduled Jobs UI
- scheduler-ui.js: job list + create/edit form + Run Now/Enable/Disable/Delete,
  reacting to scheduler:* SSE events; same-agent Run Now warning
- index.html: "⏰ Schedules" toolbar button + #schedulerModal + script include
- constants.js / app.js: frontend SSE event constants + handler map entries
- styles.css: scheduler row/badge/form styles

Follows Codeman's vanilla-JS mixin + .modal/.form-row conventions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rp7JhmQXcYJhmxFMdZuuah
2026-06-27 08:15:36 +05:30
KrisandClaude Opus 4.8 6ae86b53f6 feat(scheduler): add cron-style scheduled jobs (backend)
Adds a saved/named scheduling layer on top of Codeman's existing session
primitives. Distinct from the legacy run-now ScheduledRun concept.

- types/scheduler.ts: ScheduledJob + ScheduledJobRun
- state-store: persist scheduledJobs/scheduledJobRuns in ~/.codeman/state.json
- scheduler/scheduler-time.ts: pure once/interval/daily/weekly next-run math
- scheduler/scheduler-service.ts: CRUD, Run Now, due-checker tick, run history;
  reuses SessionPort (create -> start -> writeViaMux) for launches
- web/routes/scheduler-routes.ts: /api/scheduler/jobs CRUD + run + history
- web/schemas.ts: zod validation with schedule-type-aware refinements
- web/sse-events.ts: scheduler:* events
- server.ts: wire service into route context + 30s background tick loop
- test/scheduler-time.test.ts: 14 unit tests for next-run calculations

Phase 1 discovery recorded in SCHEDULER_DISCOVERY.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Rp7JhmQXcYJhmxFMdZuuah
2026-06-27 08:10:40 +05:30
Codeman maintainer abb6447f66 chore: version packages
Release 1.2.1: fix iOS Safari local echo on keyboard-up tab switches
(selectSession now runs the keyboard-show heal so typed input paints at
the prompt instead of staying invisible/mispositioned until a manual
keyboard toggle).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-26 01:52:32 +02:00
Codeman maintainer cc7c0e5dcb chore: version packages
Release 1.2.0: Gemini run mode, cross-session search, away digest, and
Ralph todo-config (PRs #133–#136), plus review fixes. Also refreshes CLAUDE.md
with the new-feature docs and several audit-verified drift corrections
(MockSession path, ultracode floating-window toggle, route counts, durable
input-delivery layer, mobile image-upload limits).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 00:27:24 +02:00
Codeman maintainer 368fc20fc2 fix: address PR review findings for Gemini run mode + Ralph todo-config
Gemini (PR #134) blockers:
- runGemini() now unwraps the {success,data} envelope: status check reads
  .data.available, quick-start reads data.data.sessionId (was reading the raw
  shape, so the Run-Gemini button could never start a session).
- setGeminiEnvVars() now uses the socket-scoped ${this.tmux()} setenv instead of
  bare tmux — Gemini/Google auth env vars were targeting the wrong tmux server
  and silently failing on every install.

Gemini parity polish:
- gemini tab-mode badge ('gm') + .tab-mode.gemini CSS; kill-dialog label
  'Kill Tmux & Gemini'; codeman doctor dependency-registry entry; export
  isGeminiAvailable from utils barrel; COLORTERM=truecolor + unset NO_COLOR;
  add gemini to isAltScreenStripMode (Ink TUI, repaints inline like Codex/Claude).
- Revert 4 system-routes.test.ts envelope assertions weakened to
  (body.message ?? body.error) back to (body.success === false).
- Add a runGemini() vm-sandbox test that drives the envelope path end-to-end.

Ralph todo-config (PR #135): maxTodos/todoExpirationMinutes are now persisted
and read back — surfaced via the loopState getter (RalphTrackerState) into
toState()/SSE broadcast and restored in restoreState(), mirroring maxIterations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 00:26:06 +02:00
Codeman maintainer 9cc310e843 Merge PR #134: Gemini run mode (third external-CLI mode alongside Codex/OpenCode) (COD-36) 2026-06-25 00:10:52 +02:00
Codeman maintainer aa991ece8f Merge PR #133: cross-session search (federated GET /api/search + history-panel search box) (COD-113) 2026-06-25 00:10:38 +02:00
Codeman maintainer 3b4106c349 Merge PR #136: Away digest feature (COD-41) 2026-06-25 00:10:38 +02:00
Codeman maintainer c2867be77f Merge PR #135: Ralph todo-config (maxTodos / todoExpirationMinutes) (COD-79) 2026-06-25 00:10:33 +02:00
Codeman maintainer a1b66f3510 chore: version packages 2026-06-23 23:25:18 +02:00
Codeman maintainer 98ba1fd49c fix(input): stop the connection indicator flashing "Sending 1B…" while typing
The reliable-delivery layer marks every keystroke as briefly pending until its
ACK lands a few ms later, which made the connection indicator flash
"Sending 1B…" on every character during normal typing. Hide the indicator
entirely while the connection is healthy (connected/connecting) — it now only
appears for an actual problem (reconnecting/offline), where the queued-byte
count reassures the user their input is safely buffered.

Verified in a real browser: hidden throughout connected typing, shows
"Offline (NB queued)" when offline, hides again after reconnect+delivery.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 23:24:49 +02:00
Codeman maintainer 9df310c30a chore: version packages 2026-06-23 23:13:55 +02:00
Codeman maintainer 50b8f1d9a0 feat(mobile): large + multi-image uploads from the camera-roll picker
The mobile copy/paste overlay's "🖼 Image" button (and drag-drop / paste)
now handles real-world photo batches:

- Up to 20 images per batch, uploaded with bounded concurrency (3) and a
  live "Uploading N/M…" progress toast; a final summary reports successes,
  any failures, and whether the 20-cap trimmed the selection (no silent
  truncation).
- Per-file upload limit raised 10MB → 50MB (MAX_PASTE_IMAGE_BYTES in
  buffer-limits.ts, env-overridable) so full-resolution phone photos and
  large screenshots aren't rejected.
- Very large images are downscaled to <=4096px longest edge before upload:
  fixes iOS Safari's ~16.7M-px <canvas> limit (which made huge photos fail
  to re-encode and fall back to an original that tripped the magic-byte
  check), and keeps batch uploads fast and small.
- Fix a latent concurrency bug the batch path exposed: the first parallel
  uploads to a session raced on `mkdir(.claude-images)` and the EEXIST
  losers 500'd. mkdir now treats an existing real directory as success
  (re-verifying it isn't a planted symlink), so concurrent uploads succeed.

Verified end-to-end in a real browser (Playwright): downscale, >10MB
server acceptance, 20-cap, 20/20 concurrent uploads landing on disk.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-23 23:12:41 +02:00
Aamer Akhter 11bacf67a0 fix(mobile): COD-8 hide away-digest header button on phones
The mobile-header-buttons-policy static guard requires every default-visible
header button to make an explicit phone-visibility decision. The new
.btn-away-digest button had none, failing CI. Hide it on phones alongside
.btn-settings / .btn-lifecycle-log — it's a secondary informational control
that doesn't belong on the cramped phone header.
2026-06-20 10:48:52 -04:00
Aamer Akhter 509595b837 COD-52 fix: wire maxTodos + todoExpirationMinutes through ralph-config to the tracker
The Ralph settings modal sent maxTodos/todoExpirationMinutes but RalphConfigSchema
(zod) stripped them and the ralph-config route never applied them, so the inputs
were silent no-ops.

Fix: add both as optional positive-int fields to RalphConfigSchema; destructure
and apply them in the ralph-config route (matching the maxIterations pattern).
RalphTracker had no setters (the values were module constants) — added per-instance
_maxTodos/_todoExpiryMs (defaulting to the same constants, behavior unchanged),
switched the eviction + expiry sites to read them, and added
setMaxTodos/setTodoExpirationMinutes (minutes→ms) + getters.

Test: route test POSTs the two fields and asserts the route applies them to the
tracker. Verified RED (setters not called — fields stripped) → GREEN. 34/34
ralph-routes tests pass; tsc + eslint(src) + prettier + build clean. Frontend
already sent the fields (no change).
2026-06-19 18:00:29 -04:00
Aamer Akhter c95e94e4cb COD-8 add away digest 2026-06-19 17:58:06 -04:00
Aamer Akhter 19139837e4 feat: add Gemini run mode 2026-06-19 13:10:17 -04:00
Aamer Akhter 9afaccc85d COD-9 add cross-session search frontend (history-panel search box) v1
Search box + grouped result cards + filters folded into the welcome/history
panel, wired to GET /api/search. Debounced query (250ms), type-filter chips
(session/event/file), client-side case/status/date filters, grouped cards
(badge, name, timestamp, snippet) with jump-to (session->selectSession,
run-summary->openRunSummary, file-preview->openFilePreview), empty-state +
truncated notice. All result text via textContent (no XSS surface).

Files: index.html (panel markup), terminal-ui.js (search mixin + initSearchPanel),
styles.css (.search-* styles).
2026-06-19 12:35:16 -04:00
Aamer Akhter 95df96e06a COD-9 add cross-session search backend (GET /api/search) v1
Bounded federated search over in-memory stores (sessions/cases, run-summary
events, file paths). Zod-validated query (q 1-200 chars, types csv, limit 1-60),
grouped session->event->file with exact-match-first + recency tiebreak, total
cap 60 + per-group cap 25, snippet cap 200, path-safety (relativePath only).
Frontend search box (history panel) deferred to next cycle; resume/history-prompt
text matching deferred to v1.1 (lives in large on-disk files, out of v1 bounded scope).

New: src/search-service.ts (pure core), src/types/search.ts, src/web/routes/search-routes.ts.
Tests: test/search-service.test.ts (14), test/routes/search-routes.test.ts (10).
2026-06-19 12:35:16 -04:00
Codeman maintainer 1255e28f6f fix(input): durable exactly-once input delivery so a dropped link can't lose a prompt
A "sent" prompt could vanish with no trace on a flaky connection (e.g. a train):
with local echo on, Enter cleared the overlay then sent over the WebSocket
fire-and-forget. On a half-open socket (readyState===OPEN, dead TCP) ws.send()
doesn't throw, so the frame was silently discarded, nothing was enqueued, and
navigator.onLine stayed true — the prompt was lost and never resent.

Replace the best-effort offline queue with a durable, acknowledged delivery layer:

- Client (app.js): every input frame is recorded with a stable clientId +
  monotonic per-session seq and persisted to localStorage BEFORE delivery, and
  only dropped on a server ACK. Delivered over WS (acked via {t:'ia',seq}) or,
  when the socket is down, POST in seq order (HTTP 2xx = ACK). A 2s sweep
  force-reconnects a WS whose oldest frame is unacked past 4s (half-open sockets
  never recover on their own); on reconnect/reload all pending frames re-deliver.
  Survives reconnects AND page reloads. Connection indicator shows pending count.
- Server: Session.shouldApplyInput(clientId, seq) applies each frame exactly once
  (bounded MRU map); ws-routes + POST /input dedup a redelivered seq but still ACK
  it (200 / {t:'ia'}), so an at-least-once resend can never type the prompt twice.
  Untagged input (curl/legacy) applies unconditionally — no behavior change.
- terminal-ui.js sendInput() (voice / keyboard-accessory / paste) now routes
  through the same durable layer.

Tests: test/reliable-input-dedup.test.ts (exactly-once semantics on the real
Session) + POST /input dedup route tests. Design: docs/reliable-input-delivery.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 16:58:40 +02:00
Codeman maintainer 9d12fc7f94 feat(gesture): hand-drag subagent & ultracode windows in the gesture beta
Pinch any floating subagent or ultracode run/transcript window with the
camera hand-tracking overlay and move it anywhere. Adds a 'window' grab
kind to entry.ts, slotted into the pinch priority chain
(cg-float panel → agent window → session tab → toolbar button). It moves
the window via its own style.left/top (matching app.js's mouse drag,
incl. bottom:'auto') and calls window.app.updateConnectionLines() so the
glowing connector line to the session tab tracks live — app.js redraws
from fresh rects, so no reach into its internals.

Hardening: el.isConnected guard (ultracode windows tear down mid-grab on
SSE reconnect / auto-close), all window.app calls optional-chained +
try/caught so the standalone playground still works, bring-to-front via
app.js's own z-counters, rAF-coalesced redraws cleared on drop so the
final placement always redraws.

Rebuilt the committed gesture-codeman.js bundle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 16:01:04 +02:00
Codeman maintainer 5d406c9705 chore: version packages 2026-06-19 15:31:57 +02:00
Codeman maintainer a8782b364f fix(security): harden remaining inline onclick handlers against XSS double-context
Extends PR #132 (ultracode handlers) to the rest of the frontend. The same
JS-string-in-HTML-attribute pattern — '${escapeHtml(value)}' — remained in 32
more inline handlers across app.js, panels-ui.js, session-ui.js,
subagent-windows.js, and notification-manager.js. The browser HTML-decodes the
attribute value before parsing the handler source, so escapeHtml's &#39; reverts
to ' and a quote-bearing id/path/name breaks out of the JS string literal into
executable code.

Switch all to escapeHtml(JSON.stringify(value)): JSON.stringify JS-encodes and
quote-wraps first, then escapeHtml handles the HTML-attribute layer, so the
value round-trips as one inert string argument.

Also fixes two non-escapeHtml variants of the same class:
- panels-ui.js: mux-session `sid` was pre-escaped with escapeHtml() then dropped
  into a single-quoted JS string (selectSession / killMuxSession). Now
  JSON.stringify'd at the source.
- orchestrator-panel.js: phase.id was interpolated raw (no escaping at all) into
  orchestratorSkipPhase / orchestratorRetryPhase. Now escapeHtml(JSON.stringify()).

The most realistic vector here is file paths (panels-ui openLogViewerWindow) —
filenames can legally contain a single quote.

Numeric interpolations (${i+1}, ${index}, ${item.version}) and the
developer-literal ${onclick} in orchestrator-panel are not user data and are
left as-is. Verified: 0 vulnerable patterns remain, all 22 frontend files parse
(check:frontend-syntax + node --check), and a runtime round-trip confirms the
injection that fired under the old pattern is now an inert string argument.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 15:26:42 +02:00
Ark0N d8da1bd3ff Merge pull request #132 from aakhter/cod-127-xss-ultracode-handlers
Harden ultracode inline onclick handlers against XSS
2026-06-19 15:10:28 +02:00
Aamer Akhter 06871eb7e3 Harden ultracode inline onclick handlers against XSS
The ultracode run/agent cards and minimized-tab badges built inline onclick
handlers by interpolating escapeHtml(value) inside single-quoted JavaScript
strings within an HTML attribute:

    onclick="app.openUltracodeAgentWindow('${escapeHtml(agentId)}', ...)"

escapeHtml maps ' -> &#39;, but the browser HTML-decodes the attribute value
before the handler source is parsed, so &#39; becomes a literal ' again and a
quote in a run/agent/session id breaks out of the string literal into
executable JS. escapeHtml alone is insufficient for the JS-string-within-HTML-
attribute double context.

Switch each handler to escapeHtml(JSON.stringify(value)): JSON.stringify
JS-encodes and quote-wraps the value, then escapeHtml handles the HTML
attribute layer, so the value round-trips as an inert string argument. This
matches the encoding already used by other handlers in these files.

Affected:
- ultracode-panel.js: selectWorkflowRun, openUltracodeAgentWindow
- ultracode-windows.js: restore/dismiss for minimized run and agent tabs
2026-06-19 08:55:24 -04:00
Codeman maintainer 5d59c1764d feat(ultracode): in-page agent transcript windows + minimize-to-tab (1.1.14)
Clicking an agent card opens its live transcript as an in-page connected
floating window instead of a detached browser popup. The "−" button on both
run and agent windows now minimizes into the originating session tab as a
restorable ULTRA badge (🧬 runs, 📄 transcripts). Removes the old
collapse-to-header behavior.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 21:26:33 +02:00
Codeman maintainer cfcd9d288b fix(mobile): keep /compact in extended accessory bar, only drop it from simple
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 10:29:49 +02:00
Codeman maintainer 9c22114b5a fix(mobile): remove /compact button from keyboard accessory bar (reintroduced in 1.1.10)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 10:13:23 +02:00
Codeman maintainer 98b2124d7e feat(ultracode): enrich live run tracking — real per-agent tokens/tools/state, readable title, blue connector line, click-to-open window
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-18 10:00:04 +02:00
Codeman maintainer bdaec320f5 chore: version packages
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:26:42 +02:00
Ark0N dfe20a3742 Merge PR #131: terminal touch tap interaction + forced redraw resize
feat(terminal): touch tap interaction + forced redraw resize
2026-06-17 18:08:07 +02:00
Ark0N de5216b83f Merge PR #130: mobile CJK input reliability + iPad keyboard accessory bar
fix(mobile): CJK input reliability + iPad keyboard accessory bar
2026-06-17 18:02:37 +02:00
Codeman maintainer 57eefd7aa5 fix(terminal): don't scroll/fling on a sub-threshold tap
The touchmove handler accumulated pixelAccum/velocity and could scrollLines
on every move — including micro-drift below the 8px tap threshold. A jittery
tap (<8px) stayed classified as a tap (didScroll=false, so tap-to-position
fired) yet still left a non-zero velocity, which touchend turned into a
momentum fling. Result: one tap both positioned the cursor and scrolled.

Gate the scroll/velocity accumulation behind didScroll so sub-threshold
movement is inert, matching the handler's stated tap-vs-scroll intent.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-17 18:02:21 +02:00
Teigen 2c81bbc08b feat(terminal): add forced redraw resize 2026-06-17 23:41:47 +08:00
Teigen b374121c18 fix(mobile): prevent terminal tap selection 2026-06-17 23:40:21 +08:00
Teigen b1c4330680 fix(mobile): add tap threshold to terminal touch handler
touchmove fires on any 1px finger drift, marking didScroll=true and
skipping the tap handler (which refocuses terminal/CJK input). On
iPad's large touch surface and phones with imprecise taps, this makes
terminal tap unreliable — cjkActive gets stuck true, blocking all
input (CJK and paste).

Add 8px TAP_THRESHOLD: finger movement under 8px is still a tap.
Also add touch-action:none on .touch-device .terminal-container
so the browser doesn't consume touch events before our JS handler.
2026-06-17 23:40:21 +08:00
Teigen 47359e4002 fix(iPad): enable terminal touch interaction on all touch devices
touch-action: none was only set inside @media (max-width: 430px),
so iPad's browser consumed touch events before the JS scroll/tap
handler could preventDefault. Move to .touch-device class in
styles.css so it applies at any screen width.
2026-06-17 23:40:21 +08:00
Teigen a8e7d60db4 fix(iPad): show stop button on touch devices 2026-06-17 23:40:21 +08:00
Teigen 8dc70a5f1d fix(mobile): restore /compact button to keyboard accessory bar
Reverts eb83148 which removed the /compact button from both simple
and extended accessory bar modes. Restores double-tap confirmation
and refocus guard for the compact action.
2026-06-17 23:40:04 +08:00
Teigen 4d129086d1 fix(iPad): raise toolbar z-index when case settings popover is open
backdrop-filter on the toolbar creates a stacking context that traps
the popover's z-index (1000) inside the toolbar. CJK input (z-index 52)
in the root stacking context always wins. Use :has() to raise the
toolbar above CJK only while the popover is visible.
2026-06-17 23:40:04 +08:00
Teigen 566c65c3c9 fix(iPad): accessory bar styling, positioning, and paste dialog
Move keyboard accessory bar and paste dialog CSS from mobile.css
(gated behind max-width: 1023px) to styles.css (always loaded).
iPad landscape (≥1024px) was getting unstyled white buttons.

- Add position:fixed via .touch-device class for accessory bar
- Fix dismiss button: gray-blue → blue, matching phone styling
- JS: position accessory bar above keyboard on iPad via direct bottom
- JS: position CJK above accessory bar (bottom: keyboardHeight + 44)
- Clear accessory bar bottom in resetLayout()
2026-06-17 23:40:04 +08:00
Teigen cd7d8c7329 fix(mobile): split CJK keyboard positioning by device size
Phones use translateY(-keyboardOffset) — CSS bottom is relative to layout
viewport and keyboardOffset reliably lifts it above the keyboard (iOS
doesn't auto-scroll the visual viewport for the CJK textarea on phones).

iPad uses direct bottom positioning from keyboard height — translateY
broke because iOS auto-scrolls the visual viewport when the CJK textarea
receives focus, making keyboardOffset approach 0.
2026-06-17 23:40:04 +08:00
Teigen c55af9ec39 fix(iPad): CJK input positioning, paste dialog, and voice dictation duplication
Three iPad-specific issues fixed:

1. CJK input hidden behind keyboard: updateLayoutForKeyboard() gate changed
   from screen-size to touch-device detection. On iPad, CJK textarea (always
   position:fixed) gets bottom offset computed from keyboard HEIGHT directly
   instead of keyboardOffset (which depends on visualViewport.offsetTop that
   iOS adjusts when the CJK textarea receives focus). Toolbar/accessory bar
   transforms remain phone-only (they're normal-flow on iPad).

2. Paste dialog invisible on iPad: paste overlay CSS was inside
   @media (max-width: 430px) phone breakpoint — iPad (≥768px) had no styling.
   Extracted to universal section alongside keyboard accessory bar styles.

3. Voice dictation character duplication (Doubao/third-party IME):
   iOS voice dictation does NOT fire composition events (WebKit Bug 261764).
   Text arrives as bare input events; refinement is a delete→reinsert cycle.
   Rewrote CJK input handler with two-tier debounce:
   - Keyboard typing (no delete/replacement events): 150ms debounce
   - Dictation mode (deleteContentBackward or insertReplacementText detected):
     1500ms debounce, persists 3s to cover multi-word dictation
   - Composition path (compositionend): immediate flush, unchanged
   - Keydown singles/Enter/Esc/Ctrl: immediate, unchanged
   Also: keep cjkActive=true on blur while CJK is visible (prevents xterm
   from processing duplicate input when iOS dictation UI steals focus);
   keydown single-char sends tracked via timestamp to suppress the echo
   input event that third-party IMEs fire despite preventDefault.
2026-06-17 23:40:04 +08:00
Teigen 1a54217bfb fix(mobile): don't clear textarea during compositionstart
Programmatic _textarea.value = '' during compositionstart cancels the
active IME composition on iOS Safari, breaking Chinese character input.
The phantom (U+200B) is invisible and _strip() already removes it
before sending to PTY — no need to clear it manually.
2026-06-17 23:40:04 +08:00
Teigen 70742d400a fix(mobile): restore real-time CJK input and terminal tap interaction
Root cause: the mobile-composer mode (02fa3f3) routed CJK text through
local-echo buffering, which accumulated characters until Enter instead
of sending each composed word to the PTY immediately. Additionally,
xtermFocusRedirect hijacked all terminal taps, preventing cursor
positioning and scroll interaction.

Changes:
- Remove mobile-composer accumulation mode from input-cjk.js — all
  platforms now use the same immediate-flush path (compositionend →
  flush → PTY)
- Bypass local-echo buffering in _handleCjkInput (terminal-ui.js) —
  the CJK textarea already provides visual feedback
- Remove xtermFocusRedirect so terminal taps work normally again
- Reduce CJK textarea height (34px min, 6px padding) for less
  screen intrusion
- Paste dialog now sends Enter after text so pasted content submits
- Hide CJK textarea on welcome screen (no active session)
- Add Opus 4.6 model options to selector
2026-06-17 23:40:04 +08:00
Codeman maintainer d5809d1808 docs(CLAUDE.md): note 1.1.9 tunnel opt-in (acknowledgeUnauthTunnel) in COD-55 line
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 18:50:58 +02:00
Codeman maintainer a0ac10a07c feat(tunnel,ui): purple tunnel button + opt-in unauthenticated tunnel with warning (v1.1.9)
- Daylight Blue: Cloudflare Tunnel welcome button is now purple (was orange),
  keeping Claude blue / Tunnel purple / OpenCode green distinct.
- Allow enabling the Cloudflare tunnel with no CODEMAN_PASSWORD via the UI: the
  toggle now pops a security confirm dialog and, on confirm, sends an explicit
  per-request acknowledgeUnauthTunnel:true (new action field, never persisted).
  Server logs a loud warning whenever a passwordless public tunnel starts.
  curl/API/CLI stay refused unless password/env/flag — no accidental exposure.

Tests: extend test/routes/system-routes-tunnel-guard.test.ts (ack allows + not
persisted; ack:false still refuses). Verified e2e on an isolated instance
(purple button, confirm dialog, retry carries the flag, no real tunnel opened).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 18:44:03 +02:00
Codeman maintainer f7814ad364 feat(ui): distinct colors for welcome action buttons on Daylight Blue (v1.1.8)
On the default daylight-blue skin the three welcome buttons all read blue.
Give each its own identity: Run Claude Code keeps the blue accent, Cloudflare
Tunnel takes Cloudflare brand orange, Run OpenCode takes emerald green (with
matching hover/active states + dark ink for contrast). Scoped to daylight-blue
only; daylight-green and OG unchanged. Verified in-browser (blue/orange/green).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 18:22:33 +02:00
Codeman maintainer 3172befd5d fix(terminal): keep Claude scrollback reachable — strip alt-screen/3J/mouse for claude mode (v1.1.7)
Terminal scroll-up intermittently broke for Claude sessions (most visible on
iPhone). Claude Code periodically emits alt-screen switches (?1049h/?47h/?1047h),
scrollback-erase (3J), and mouse-tracking enables for full-screen UIs, which move
xterm.js to the scrollback-less alt buffer / wipe saved lines / hijack the wheel.
Codeman stripped these but only for codex mode.

Share the strip via isAltScreenStripMode(mode) = codex || claude, applied at both
sites that were codex-only: the live PTY stream (Session._handleTerminalOutput,
incl. the chunk-boundary carry) and the /terminal buffer replay. shell stays
excluded (vim/less/htop need the alt screen); opencode unchanged.

Tests: test/claude-scrollback-strip.test.ts (8 new); codex strip tests unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 18:08:45 +02:00
Codeman maintainer 29ffc62536 fix(ultracode): pop floating windows on fresh devices loading mid-run (v1.1.6)
Re-run syncAllUltracodeFloatingWindows() after server settings load so a
first-time device whose getLightState run snapshot arrives before the async
settings fetch resolves still pops an already-active run's window immediately,
instead of waiting for the next ~10s watcher tick. Also fixes a stale
@fileoverview comment that named the wrong gating setting.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 14:27:10 +02:00
Codeman maintainer 4cb3a4aac8 fix(ultracode): (x) Close fully hides the Ultracode Agents panel
closeUltracodeAgentsPanel() only removed `open`, leaving the drawer in its
collapsed peek state (header strip still visible) — so (x) looked like a no-op.
Now also adds `hidden` (display:none), mirroring closeSubagentsPanel; does NOT
flip showUltracodeAgents (that gates the watcher + floating windows). Verified in
a real browser (post-close computed display:none). Bumps 1.1.4 -> 1.1.5.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 22:46:53 +02:00
Codeman maintainer b6531cbf79 fix(ultracode): floating windows pop for LIVE runs (watch transcript tree)
The Workflow runtime writes workflows/wf_<id>.json only at completion (always
terminal), so workflow-run-watcher never saw a run until it was already done and
the ACTIVE-gated floating window never popped. The watcher now also scans
subagents/workflows/wf_<id>/ and synthesizes a minimal running record (agentId
slots preserved for the transcript-click join, lastActivityAt from mtimes,
done/running from the journal), superseded by the real wf_<id>.json at
completion. Standalone (no subagent-watcher import). Verified e2e on a real
in-flight run; +6 unit tests. Bumps 1.1.3 -> 1.1.4.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 22:20:52 +02:00
Codeman maintainer d16bf34e34 feat(ultracode): floating run windows with tab connector lines + dedicated toggle
Auto-popping draggable window per active ultracode/Workflow run, connected by a
glowing line to its originating session tab (resolved via claudeSessionId ===
sessionUuid). Mirrors the live agent grid; auto-closes after a run finishes;
dismissals are remembered. Additional to the existing docked panel.

New "Ultracode Floating Windows" setting (default OFF), independent of the
"Ultracode Agents" panel toggle; either toggle starts the workflow-run watcher.

Also bumps version to 1.1.3 and brings CLAUDE.md up to date for the ultracode
subsystem.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 16:33:54 +02:00
Codeman maintainer e6989bdb40 chore: version packages 2026-06-15 11:04:46 +02:00
Codeman maintainer 6ab6bbbbd4 feat(ultracode): Phase 4 — click an agent card to open its live transcript
Each workflow agent card with an agentId is now clickable and opens that agent's
live transcript in a popup, reusing the existing GET /api/subagents/:agentId/
transcript route. The workflow agent's agentId is byte-identical to the
agent-<id>.jsonl stem that subagent-watcher already tracks (via w16's
watchWorkflowDirs), so this is a pure client-side join — ZERO subagent-watcher
edits.

Graceful degradation: 'start' (queued) agents have no agentId yet and stay
non-clickable; an aged-out/untracked agent (subagent-watcher's 4h startup window,
or tracking disabled) returns an empty transcript and shows a friendly note
instead of an empty popup.

Verified on a live isolated server: the subagent transcript route serves a
workflow agent's transcript (150 entries) and the runId's agents[] carries the
matching agentId; Playwright confirmed clicking a card opens the transcript popup
with no console errors. frontend-syntax / public-assets / CSS-parse clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 10:49:22 +02:00
Codeman maintainer c15c19fab7 feat(ultracode): master-detail tab for Workflow/ultracode run visualization
Opt-in (showUltracodeAgents, default OFF) panel that visualizes ultracode /
Workflow-tool runs like Claude Code's "working agents" TUI: LEFT = runs + phases
(selectable tasks), RIGHT = each run's agents with model, live state, tokens
burned, and tool calls.

Standalone — ZERO edits to subagent-watcher.ts. A new workflow-run-watcher.ts
singleton globs the run-state tree (~/.claude/projects/*/*/workflows/wf_*.json,
disjoint from the transcript tree), strips the heavy script/scriptPath/result/logs
fields (174KB -> ~25KB/run), and emits workflow:run_* SSE events. The LEFT list
ships lightweight summaries (getLightState replay + SSE); the RIGHT pane fetches
the full run (with agents[]) via GET /api/workflows/:runId on selection.

Backend: workflow-run-watcher.ts, types/workflow-run.ts, config/workflow-config.ts,
3 SSE events, getLightState workflowRuns replay, GET /api/workflows[/:runId],
showUltracodeAgents schema key + boot-gate (default OFF) + live toggleService.
Frontend: ultracode-panel.js (debounced master-detail render, run/phase select),
header launcher (btn-ultracode-agents--hidden marker -> mobile-guard-exempt),
App Settings toggle (SYNCED, deliberately not in displayKeys).

Agent states on disk are start|progress|done (start=queued; done has
durationMs/resultPreview). Tests: workflow-run-watcher (9), workflow-routes (3).
Verified: tsc/lint/prettier/frontend-syntax/public-assets/mobile-header-guard
clean; full test:ci green (2986 passed); live server + Playwright e2e against 25
real runs (28-agent grid, phase filter, OFF hides launcher).

Design: docs/ultracode-agent-viz-plan.md (rev. 3).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 08:40:05 +02:00
Codeman maintainer f6a30d7335 fix(subagent-watcher): discover workflow-nested agents + harden meta→transcript upgrade
Two follow-ups to db93491 (the 2026-06 CC meta.json format change), after
reverse-engineering the new on-disk layout with a live current-CC subagent +
1Hz fs poller:

(1) Workflow recursion — the Workflow tool nests its agents at
    subagents/workflows/{wf}/agent-{id}.jsonl, one level below the flat
    subagents/ scan, so they were never tracked. Add watchWorkflowDirs()
    (driven from scanForSubagents) to descend and watch each workflow dir
    (idempotent; fs.watch recursive is unsupported on Linux, so the ~5s
    periodic scan re-drives it — same latency as new-session discovery).
    Require the `agent-` prefix in the flat readdir + watch callback so a
    workflow dir's sibling journal.jsonl can't register a bogus "journal" agent.
    E2E verified against real ~/.claude/projects: 32 workflow-nested agents
    discovered (wf_fa35c1d8-4a9), 0 bogus journal agents.

(2) Transcript timing — empirically the per-agent .jsonl IS written at the
    standard subagents/ path and grows incrementally (tailable); the
    /tmp/.../tasks/<id>.output the prior probe found is just a symlink back to
    it. meta.json lands at spawn, the .jsonl a beat later. Add a meta→transcript
    upgrade in registerAgentFile: when an agent registered meta-only gets its
    sibling .jsonl, re-point filePath, drop the stale sidecar context, start
    tailing, and emit subagent:updated (not a duplicate discovered). Corrects the
    now-inaccurate "no transcript to tail" doc comment on registerAgentMeta.

Tests: 2 new cases (workflow-nested discovery; journal.jsonl not registered).
All 56 pass; tsc/lint/format clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 01:43:15 +02:00
Codeman maintainer db93491dd1 fix(subagent-watcher): discover subagents via agent-*.meta.json (CC format change)
Claude Code changed its subagent on-disk format (~2026-06-14): TUI Task
subagents now write `agent-{id}.meta.json` ({agentType,description,toolUseId})
into the session's `subagents/` dir and no longer reliably write a per-agent
`agent-{id}.jsonl` transcript there. The watcher discovered agents ONLY by
`.jsonl`, so it tracked zero — subagent windows and the monitor's "N TRACKED"
showed nothing.

- Add `registerAgentMeta()`: discover from the meta sidecar (description from
  meta.description/agentType), prefer a sibling `.jsonl` transcript when present
  (richer), never tail a meta file.
- Initial scan + directory watcher now handle `.meta.json` alongside `.jsonl`.
- Tests: 2 new cases (meta-only discovery; prefer-.jsonl-when-present).
  Verified e2e against a real ~/.claude/projects fixture.

Known follow-ups (not in scope): meta-only agents have no per-agent transcript
to tail (no live tool-call feed, status stays 'active'); workflow agents under
`subagents/workflows/{wf}/agent-*.jsonl` are still missed by the flat scan.

Also adds the README screenshot tooling used to surface this:
- capture-real-overview.mjs: DSF=2 + ?nowebgl crisp path (DOM renderer avoids
  the WebGL glyph-doubling at deviceScaleFactor>1).
- capture-readme-real.mjs: real-instance desktop-scene capture (dashboard/
  monitor/subagent) for an isolated beta seeded from prod settings.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 01:24:30 +02:00
Codeman maintainer b7ff54b2ec fix: file viewer opens audio/svg/binary like the attachments viewer
The File Browser preview and Attachments preview share openFilePreview(),
but the workspace branch (via /file-content) misclassified several types the
attachments viewer handled fine:

- SVG was reported as type:image, but file-raw serves SVG as octet-stream +
  attachment (XSS hardening), so the <img> broke. Now fetched and rendered via
  a same-origin image/svg+xml blob <img> (safe; <img> never runs SVG scripts).
  file-raw's SVG hardening is unchanged.
- Audio (mp3/wav/ogg/m4a/aac/flac/opus) was type:binary -> "Cannot preview".
  Now classified as audio and rendered with <audio controls>; file-raw gained
  the matching audio/video MIME types so playback works.
- Binary formats not in the hardcoded list (xlsx/doc/zip/...) were decoded as
  UTF-8 and dumped as mojibake. Replaced the static list with a NUL-byte
  content sniff that flags arbitrary binaries; the binary fallback now offers a
  Download link instead of dead-ending.

Adds route tests for audio, known-binary (xlsx), and NUL-sniff classification.
Verified end-to-end on an isolated instance + headless browser.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 00:42:12 +02:00
Codeman maintainer dc63d1f1a6 tools: harden real-overview screenshot capture + document DSF/cache gotchas
scripts/capture-real-overview.mjs:
- Default deviceScaleFactor to 1 (DSF=2 makes xterm's headless WebGL renderer
  draw console glyphs at ~2x while reporting nominal cell dims — invisible to
  cols/cell measurement, only the pixels reveal it; HTML chrome is unaffected so
  only the terminal font looks oversized)
- Mint a unique timestamped filename per run so a viewer/HTTP cache can't shadow
  a fresh capture with a stale render of a fixed path
- Seed per-device localStorage (skin, codeman-font-size, codeman-app-settings)
  so the capture reflects a real device: plan-usage chip shown (per-device key,
  deleted from server payload), side panels closed for a full-width terminal
- Support prod's self-signed HTTPS (ignoreHTTPSErrors), env-configurable viewport

CLAUDE.md:
- Document the DSF=1 / unique-filename screenshot gotcha (incl. the real
  Codeman-side immutable-static-asset cache footgun)
- Add the sanitize-html.js infra module (DOMPurify mXSS allowlist, COD-56) to the
  frontend module list and load order (was missing)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 23:59:40 +02:00
Codeman maintainer 7c5920d3b9 chore: version packages 2026-06-14 23:06:32 +02:00
Ark0N e1e670594b Merge PR #128: auto-wrap desktop session tabs on overflow + resize re-eval
Auto-wrap desktop session tabs to a second row on overflow
2026-06-14 22:43:43 +02:00
Ark0N 2e28e17834 Merge PR #123: hide CJK textarea on welcome screen + mobile test update
fix(cjk): hide CJK textarea on welcome screen and fix vertical centering
2026-06-14 22:43:17 +02:00
Ark0N 90f18438ff Merge PR #127: require hook-event secret unconditionally + stale-config self-heal
Require the hook-event secret unconditionally (drop managed-tunnel gating)
2026-06-14 22:43:13 +02:00
Ark0N 5b62f397ec Merge PR #129: macOS Option/physical-key session shortcuts + terminal-ui ESC-leak fix
Make Option/Alt session shortcuts work on macOS (physical key codes)
2026-06-14 22:43:08 +02:00
Ark0N 1e54ebcdf4 Merge PR #125: add codeman doctor dependency checker + accuracy review fixes
Add `codeman doctor` tool-dependency checker
2026-06-14 22:43:04 +02:00
Ark0N 0364bea166 Merge PR #126: harden markdown sanitizer with DOMPurify (mXSS) + allowlist/test review fixes
Harden markdown HTML sanitizer with vendored DOMPurify (mXSS)
2026-06-14 22:42:59 +02:00
Claude (Codeman maintainer) c7e8ff616f fix(tabs): re-evaluate auto-wrap on resize and on every full tab rebuild
Review polish on the desktop tab auto-wrap:

- Auto-wrap is purely width-driven, but updateTabOverflowMode() was only called at the
  tail of _renderSessionTabsImmediate (SSE content renders). Window resize — the primary
  trigger for tabs crossing the one-row overflow threshold — never re-evaluated it, so
  narrowing/widening the window left the wrap state stale until an unrelated status event
  fired a render. Call it from the debounced window-resize handler (no-op on
  mobile/tablet, where the method bails).

- Move the re-evaluation into _fullRenderSessionTabs() as well, so the incremental
  branch's two early `_fullRenderSessionTabs(); return;` paths (badge add/remove, which
  change tab width) and the manual two-rows toggle (applyTabWrapSettings → _fullRender…)
  re-evaluate too. The latter also fixes a transient where enabling manual two-rows while
  auto-wrap was on left both classes set (clipping folder tabs to 96px) until the next
  render.

- Add boundary cases to the policy test: exact fit and the +1 sub-pixel tolerance (no
  wrap), 2px over (wrap), and a single overflowing tab (no wrap).

Verified: tab-overflow test passes; tsc, check:frontend-syntax, check:public-assets,
prettier all clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 22:38:19 +02:00
Claude (Codeman maintainer) 21fbff4d8a fix(hooks): self-heal stale pre-secret hook configs so COD-91 doesn't 401 them
Making the hook-event secret unconditionally required closes the own-loopback-proxy gap,
but it would also silently 401 the hook curls baked into cases created BEFORE the secret
header existed (COD-54, 2026-06-10): writeHooksConfig only runs at case CREATION, so an
existing/linked case on a password-protected install keeps secret-less curls that the new
gate rejects (degrading idle/stop/teammate/task signalling with no error surfaced).
No-password installs are unaffected — the gate isn't registered without CODEMAN_PASSWORD.

Add `refreshStaleHookSecret(casePath)` and call it on Claude-mode spawns in
POST /api/sessions and POST /api/quick-start (existing-case branch). It regenerates the
hooks block ONLY when settings.local.json already holds Codeman's own hook curls (they
target /api/hook-event) that lack the X-Codeman-Hook-Secret header — a no-op when the
hooks are absent, not ours, or already current, so it never clobbers user customizations
and is cheap on every spawn. Fresh cases are unaffected (writeHooksConfig already wrote
the secret). withSettingsLock serializes it with the model/statusLine writers.

Verified: new test/hook-secret-selfheal.test.ts 5/5 (heal + key-preservation + no-op on
current/foreign/absent/malformed); the PR's cod54 + auth-security suites still pass
(36); tsc, lint, format:check, and npm run build all clean (symbol present in dist).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 22:35:29 +02:00
Claude (Codeman maintainer) 8ffb2b0644 test(cjk): update the mobile server-override test for the welcome-screen gate
The PR gates CJK textarea visibility on an active session
(`showCjk = cjkUserEnabled && !!activeSessionId`) so the fixed-position textarea no
longer floats over the welcome overlay. That intentionally changes the behavior the
existing `shows the CJK textarea on mobile only for server override` test asserted —
it set `_serverCjkOverride = true` on a fresh page (no active session) and expected the
textarea visible, which now (correctly) resolves to hidden. The test lives in
test/mobile/** (excluded from CI), so it wasn't caught by the PR's green CI.

Update the test to verify the new, intended behavior: with the server override on it
stays hidden on the welcome screen (no active session) and is revealed once a session
is active. This is a co-authored review fix; the original change is TeigenZhang's.

Verified: tsc, check:frontend-syntax, check:public-assets, prettier all clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 22:30:09 +02:00
Claude (Codeman maintainer) 80ebf8b549 fix(shortcuts): stop Alt/Option nav keys leaking ESC sequences into the terminal
The PR migrated the app.js tab-nav handler to physical e.code but left xterm's
pass-through gate (terminal-ui.js) matching ev.key digits. Consequences:

- Alt+[ / Alt+] (the new bindings) were never in the gate, so xterm sent ESC[ / ESC]
  to the PTY on every platform AS WELL AS switching the session.
- Alt+digit on a remapped macOS Option layout (Option+1 -> "¡") didn't match the
  ev.key '0'-'9' gate either, so xterm injected ESC<char> — on exactly the layouts
  this PR exists to fix.

Update the xterm gate to mirror app.js exactly: suppress when
`ev.altKey && !ctrl && !shift && /^(Digit[1-9]|BracketLeft|BracketRight)$/.test(ev.code)`.
Returning false there tells xterm not to write to the PTY, so the shortcut switches
the tab with no stray escape sequence.

Also: relabel the docs Alt/Option (the mechanism is layout/OS-independent, so the
shortcut works for Linux/Windows Alt users too — "Option" alone was Mac-only wording),
and add a keyboard-shortcuts test asserting terminal-ui.js gates on the same physical
codes so this desync can't regress (a grep the original test missed).

Verified: keyboard-shortcuts test 4/4, check:frontend-syntax, check:public-assets,
format:check all clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 22:27:36 +02:00
Claude (Codeman maintainer) c101cc8716 fix(doctor): correct Node minimum, drop phantom gemini, add pdftoppm, validate --category
Review fixes on top of the `codeman doctor` checker:

- Node minVersion 18.0.0 -> 22.0.0. package.json engines is ">=22.0.0" and the docs/CI
  require Node 22+, so doctor was green-lighting Node 18-21 (a false pass).
- Remove the phantom `gemini` registry entry. Codeman has no Gemini backend
  (SessionMode = 'claude' | 'shell' | 'opencode' | 'codex'); the entry advertised a
  dependency that nothing uses.
- Add `pdftoppm` (poppler) to the office group. document-thumbnailer.ts calls pdftoppm
  with no fallback as the sole PDF/Office first-page thumbnail renderer, yet it was
  absent from the registry, so doctor never reported it missing.
- Fix the `--category` mismatch: the help advertised `documents|media` categories that
  the ToolCategory type/registry never defined, and an unknown category silently
  produced an empty "all healthy" table. Introduce TOOL_CATEGORIES as the single source
  of truth (type + help + validation); an invalid `--category` now errors with the
  valid list and exits 2.

Verified: tsc, lint, format:check all clean; both dependency tests pass (20);
`doctor` runs correctly (Node 22.22 ok, pdftoppm detected, no gemini), `--category media`
errors with exit 2, `--category office` lists libreoffice/pdftoppm/msoffice.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 22:24:46 +02:00
Claude (Codeman maintainer) cceb24ed8f fix(sanitizer): enforce the curated allowlist + make the test run order-independently
Review fixes on top of the DOMPurify mXSS hardening:

- Remove `USE_PROFILES: { html: true }` from the sanitize-html.js config. DOMPurify
  treats USE_PROFILES and ALLOWED_TAGS/ALLOWED_ATTR as mutually exclusive — with a
  profile set it resets the allow-lists to the full HTML profile and silently ignores
  the curated lists, so the tight markdown-only allowlist was dead config (still
  XSS-safe via FORBID + core, but far broader than intended: <button>/<input>/
  <details>/<audio>/<select>/<label> all survived). Dropping USE_PROFILES puts the
  curated ALLOWED_TAGS/ALLOWED_ATTR back in force; FORBID_TAGS/FORBID_ATTR stay as
  defense-in-depth and DOMPurify keeps its default safe-URI handling.

- Rewrite test/markdown-sanitizer.test.ts to run in the default node environment with
  an in-test jsdom window instead of a per-file jsdom environment. That environment
  externalizes node:fs/node:path under vite, so the suite failed to load in isolation
  ("No such built-in module: node:") and only survived the full CI run because an
  earlier node-env test happened to pre-cache node:fs — order-dependent and fragile.
  The rewrite is order-robust and adds an "allowlist is actually enforced" block
  (non-markdown tags must be dropped) that fails if USE_PROFILES is reintroduced.

Verified: 25/25 tests pass standalone under config/vitest.ci.config.ts; tsc, lint,
format:check, check:frontend-syntax, check:public-assets, and npm run build all clean.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 22:21:13 +02:00
Aamer Akhter 60dab7ce3f Make Option/Alt session shortcuts work on macOS (physical key codes)
Tab-switch shortcuts matched e.key, so on macOS Option+1 emits a special
character ('¡', not '1') and the shortcut silently failed. Switch to physical
e.code (Digit1-9), which is layout-independent. Also adds Option+[ / Option+]
for previous / next session. Help modal + README updated.

Test: test/keyboard-shortcuts.test.ts.
2026-06-14 15:54:50 -04:00
Aamer Akhter a5263b3252 Auto-wrap desktop session tabs to a second row on overflow
When desktop session tabs overflow one row, wrap them to a second row instead
of horizontal scroll — unless the user has pinned the manual two-row layout
(tabTwoRows). Mobile/tablet keep horizontal scroll. The wrap policy
(shouldAutoWrapTabs) lives in constants.js as a pure, unit-testable function;
updateTabOverflowMode() measures overflow after each tab render and toggles
.tabs-auto-wrap.

Test: test/tab-overflow.test.ts (vm-loads constants.js, asserts the policy).
2026-06-14 15:49:12 -04:00
Claude (Codeman maintainer) 90cd481b9f chore: version packages
Release 1.1.0. Headline: opt-in Plan Usage Limits chip (per-device live
5h/weekly plan %), attachment history drawer + opt-in Attachments button,
Opus 4.6 model options, and mobile header regression guards.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 21:39:14 +02:00
Claude (Codeman maintainer) 787e5e2a03 feat(attachments): make the header attachments button opt-in (default OFF)
The COD-39 attachments button was hard-visible in the header — first on
mobile, then (after the mobile-only hide) still on desktop. Make it a
proper opt-in App Settings → Display toggle ("Attachments Button"),
default OFF everywhere, mirroring the Response Viewer button:

- index.html: button ships with the `btn-attachments-history--hidden`
  marker; new settings checkbox #appSettingsShowAttachmentsButton.
- styles.css: base `display:inline-flex !important` + a more-specific
  `--hidden` rule (same pattern as the response viewer).
- settings-ui.js: load/save/getDefaultSettings(false) + a live toggle in
  applyHeaderVisibilitySettings. Per-device and NON-leaking — added to
  displayKeys AND stripped from the server payload, so enabling it on
  desktop never makes it appear on mobile (or any other device). No
  server-side render step (purely client display, like the eye button).
- mobile.css: dropped the now-redundant phone-only hide — the opt-in
  marker hides it everywhere by default; the per-device toggle governs
  both desktop and phone.

Tests updated: the CI static guard drops btn-attachments-history from the
phone-hidden lock (it's opt-in now, excluded from the default-visible
enumeration — the guard still gates any NEW default-visible button); the
real-browser E2E now asserts default-hidden on a desktop-class viewport
and visible after enabling the setting.

Verified on a real desktop browser: hidden by default, the settings
toggle exists, enabling it shows the button. tsc + frontend-syntax +
prettier + public-asset checks + both test suites green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 21:21:06 +02:00
Claude (Codeman maintainer) 097433c86f docs: update CLAUDE.md for the attachments subsystem growth
Document the three PRs that grew attachments since the last update:
COD-37/#119 (registry + magic links) was already covered, but
COD-38/#120 (document previews/thumbnails) and COD-39/#121 (history
drawer) added four source files and several endpoints that weren't
documented. Split a dedicated Attachments row out of Infra, extend the
Attachments Key Pattern to cover the converter pipeline + concurrency
limiter + history drawer, and refresh the files-route handler count
(8 -> 14) and total (~140 -> ~146). Also carries the prior pending
app.js line-count (3.7K -> 3.9K) and config-file-count (10 -> 12) bumps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 21:13:36 +02:00
Claude (Codeman maintainer) e738c776c1 fix(settings): slim the Skin picker select to match its row
The skin picker inherited .form-select's 0.8rem font + 0.5rem vertical
padding, rendering bigger and taller than the settings row it sits in
(0.75rem / 0.45rem). The daylight skins' Manrope font exaggerated it,
so "Daylight Blue" looked oversized and the field too thick. Scope a
0.75rem font + 0.3rem vertical padding to .settings-item-skin .form-select
so the field text matches the row label.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 21:08:18 +02:00
Claude (Codeman maintainer) e10f0dabdb fix(mobile): hide attachments-history button on phones + regression guards
The COD-39 attachment-history header button was visible on the cramped
phone header. Hide it on phones alongside the settings gear and lifecycle
log (the mobile header is intentionally minimal — those controls live in
the toolbar). One-line addition to the existing @media (max-width: 430px)
display:none block in mobile.css.

This is the second time a header control leaked onto mobile (the
plan-usage chip was the first), so add two regression guards:

- test/mobile-header-buttons-policy.test.ts — a pure static analysis of
  index.html + mobile.css (no browser), so it runs in the normal CI sweep
  (the test/mobile/** Playwright suite is EXCLUDED from CI and never gated
  this). It enumerates every default-visible header button and fails when
  one has no phone-visibility decision — either a mobile.css hide rule or
  an explicit MOBILE_VISIBLE_ALLOWLIST entry. A new header button now
  forces that decision. Verified it fails on the pre-fix state and passes
  after.
- test/mobile/header-buttons.test.ts — real-browser E2E in the mobile
  suite: asserts the attachments/settings/lifecycle buttons are hidden on
  an emulated iPhone 14 Pro and the attachments button is visible on a
  desktop-class tablet.

tsc + lint + prettier + both new tests green. Only CSS + tests changed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 20:45:23 +02:00
Claude (Codeman maintainer) 661c89cefd fix(plan-usage): make the usage chip per-device, not synced
The plan-usage header chip (5h/7d %) was a SYNCED setting, so enabling
it on desktop turned it on for mobile too — even though the user never
enabled it there. Make the chip's DISPLAY purely per-device (default
OFF) like the response viewer / skin, while keeping telemetry COLLECTION
server-side.

Three leak sources fixed:
- server.ts renderIndexHtml force-revealed the chip from the synced
  value (pre-paint), pushing the desktop choice onto every device.
  Removed — the chip now ships hidden and the client reveals it
  per-device via applyHeaderVisibilitySettings.
- settings-ui.js load-merge let the server value win, writing desktop's
  `true` into the (separate) mobile settings blob. showPlanUsageLimits
  is now a displayKey AND is dropped from the server payload on load, so
  a stale server value is never seeded into a device that didn't enable
  it. It's also stripped from the save payload so a mobile "off" can't
  clobber the server.
- Collection was gated on the same synced flag. Decoupled via a new
  `statusLineTelemetry` ACTION field (schema + system-routes): sent on
  ENABLE only and never persisted, so the exporter is injected when a
  device turns the chip on but is never yanked when another device has
  it off (it's shared across sibling sessions). Session-create already
  reads the per-device blob, so that path was already correct.

One-time migration clears a stale synced `true` from the mobile blob so
existing mobile installs default to OFF without a manual toggle.

Verified end-to-end on an isolated server: with showPlanUsageLimits=true
persisted, the rendered HTML ships the chip hidden; a fresh browser
context (mobile case) keeps it hidden while a context that explicitly
enabled it shows it; the PUT accepts statusLineTelemetry and does not
persist it. tsc + frontend-syntax + system-routes/index tests green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 20:15:39 +02:00
Ark0N a122e867ef Merge PR #122: restore response-viewer eye button on mobile
Remove the dead mobile-collapsed header tray that hid the entire header-right cluster (incl. the opt-in response-viewer eye) on phones/tablets, and update the mobile test to assert inline reachability. Eye stays hidden by default (showResponseViewer).
2026-06-14 19:46:11 +02:00
Claude (Codeman maintainer) a68f23e647 test(mobile): assert header tray reachable inline, not collapsed (#122)
Removing the dead `mobile-collapsed` tray (this PR) means the test that
asserted the headerRight tray *stays collapsed* on mobile now contradicts
the code and would fail when run. Flip it: with the three-dot utility
toggle gone, the header-right utilities must flow inline and stay
reachable on small viewports. The response-viewer eye itself remains
hidden by default (showResponseViewer opt-in), so this only re-exposes
the already-default-visible utilities inline.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 19:40:51 +02:00
Aamer Akhter f0f43ddbad Require the hook-event secret unconditionally, not only under a managed tunnel
COD-54 gated the /api/hook-event + /api/status-telemetry localhost bypass
behind the shared X-Codeman-Hook-Secret only WHILE a managed tunnel was
running, keeping a plain localhost bypass otherwise. But Codeman can't detect
a user's OWN loopback reverse proxy (their own `cloudflared --url`,
`tailscale serve`, nginx -> 127.0.0.1), which proxies internet traffic into
the loopback origin with req.ip === 127.0.0.1 — so that setup kept the unsafe
plain bypass.

Require the secret on the loopback bypass unconditionally. Managed-session
hooks already always present it (X-Codeman-Hook-Secret from
$CODEMAN_HOOK_SECRET_FILE, generated for every instance), so the legitimate
hook channel is unaffected; only the previously-unguarded own-proxy path is
now rejected. Drops the now-unused getTunnelRunning param from
registerAuthMiddleware.

Tests: cod54-hook-event-auth (tunnel-down now also requires the secret, plus
a good-secret positive case); auth-security (hook tests present the secret to
reach schema validation).
2026-06-14 12:46:58 -04:00
Aamer Akhter ea53916adc Replace markdown denylist sanitizer with vendored DOMPurify (mXSS hardening)
The previous _sanitizeHtml was a denylist over agent/transcript markdown
rendered via innerHTML; it missed style attributes and the svg/math mXSS
namespaces — e.g. <svg><style><img src=x onerror=alert(1)></style></svg>
re-serialized into a live <img onerror>.

Vendor DOMPurify 3.4.8 (allowlist) following the existing marked.min.js
vendor pattern (same-origin, CSP script-src 'self'; not in package.json so
no lockfile drift). New sanitize-html.js wires a hardened allowlist config
(FORBID style/svg/math/script/iframe/object/embed/form; no data attrs);
app.js _sanitizeHtml delegates to it with a fail-closed escape-all fallback.
index.html loads dompurify -> sanitize-html -> app.js (defer); build.mjs
minifies + content-hashes sanitize-html.js.

Test: test/markdown-sanitizer.test.ts (jsdom, real shipping artifacts) —
mXSS payloads neutralized + legit markdown preserved.
2026-06-14 12:33:48 -04:00
Aamer Akhter 585127deb2 Add codeman doctor tool-dependency checker (COD-45)
Environment-aware dependency probe (linux|darwin|win32|wsl) with a static
registry, an injectable ProbeHost seam for testing, grouped table + `--json`
output, and a non-zero exit when a required dependency is missing/outdated.
Node and tmux are the only hard-required tools; the agent CLIs and document
converters (LibreOffice / MS Office via WSL interop) are optional. CI-safe
unit tests (no tmux, injected host).
2026-06-14 12:25:49 -04:00
Claude (Codeman maintainer) e742d00c98 Merge PR #121: attachment history drawer (COD-39)
Per-session attachment history with a slide-in drawer, unread badge, and
re-show. Rebased onto master (stacked on #120) + review hardening (malformed-
history recovery guard, resilient list route, badge positioning, debounce
cancel, stable re-show, Escape-to-close, CSS token fixes). See PR #121.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 09:34:41 +02:00
Claude (Codeman maintainer) 1a363a3e62 fix(attachments): address review findings on attachment history drawer (#121)
Follow-up fixes applied during review of PR #121 (all confirmed minor/nit;
no blockers). Security posture verified sound (externalPath never leaves
toState()/the list route; re-registration runs the guard).

- fix(recovery): restoreAttachmentHistory now skips malformed/legacy saved
  items (null, non-object, missing source/fileName) instead of throwing inside
  the Session constructor — a corrupt __attachmentHistory entry could otherwise
  abort the entire mux-recovery loop. (P1)
- fix(routes): the attachment-list route degrades a single failing entry to
  {missing:true} instead of failing the whole drawer. (INT-4)
- fix(ui): give the attachments header button a positioning context so the
  unread badge anchors to the icon, not the header bar. (F1/CSS-1)
- fix(ui): cancel the debounced history refresh on drawer close and guard it
  against a stale session/closed drawer. (F3)
- fix(ui): re-show ("Card") of a detected item now uses the item's own
  timestamp so the cardId is stable — focuses the existing card instead of
  stacking duplicates. (F4)
- fix(ui): Escape now closes the drawer, matching every other panel. (UX-1)
- fix(ui): badge shows "99+" past 99 (was an inconsistent 100/99 cap). (BADGE-1)
- style: drop the duplicate @keyframes notif-badge-pulse (dead CSS). (INT-1/CSS-3)
- style: empty-state used three undefined CSS custom properties
  (--text-primary/--border-color/--bg-tertiary) → use the defined
  --text/--border-light/--bg-input tokens. (CSS-2)
- test: add constructor restore round-trip + malformed-item resilience tests.

Deferred (noted for author): broadcasting the full 100-item history in every
session-state SSE event (payload bloat), "unread" badge semantics, making the
header button opt-in, and app.inject route tests for the two new endpoints.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 09:30:01 +02:00
Aamer Akhter 577b6d7384 COD-39 attachment history drawer
Stacks on COD-38: accumulates a per-session attachment history and exposes it
through a slide-in drawer with an unread badge, so attachments stay reachable
after their cards are dismissed.

Backend:
- session-attachment-history: history state — dedupe by source path / relative
  path, newest-first, 100-item cap, and externalPath sanitization (the absolute
  host path is server-private and never leaves toState()).
- session.ts: _attachmentHistory + getter (sanitized) / upsert / restore /
  getAttachmentHistoryForPersist; restored from saved state in the constructor.
- file-routes: GET /attachments (list — resolves each entry to live metadata +
  routes; external entries are re-registered) and GET /attachments/:id
  (metadata poll). The by-id route guards via the registry's TOCTOU-safe
  resolveServableAttachmentPath.
- server.ts: detected/registered attachments upsert into history and persist;
  the private (externalPath-bearing) history rides on disk under
  __attachmentHistory, separate from the sanitized public copy, and is restored
  on mux-session recovery.
- types/session.ts: SessionAttachmentHistoryItem + SessionState.attachmentHistory.

Frontend:
- panels-ui: the drawer (lazy-built), unread badge, list render with per-item
  preview/download/open/"Card" (reshow) actions, and live refresh of the open
  drawer on new detections.
- app.js: history state + per-session badge/cleanup wiring.
- index.html / styles.css / mobile.css: header button + badge and the drawer.

Verified: tsc / eslint / prettier / frontend-syntax / public-assets clean; new
history-module unit tests pass; full test:ci green (2866 passed); badge, drawer
open/render/reshow/close verified in-browser.
2026-06-14 09:05:17 +02:00
Claude (Codeman maintainer) 5eacb1cf03 Merge PR #120: document attachment previews + thumbnails (COD-38)
Adds attachment cards with first-page thumbnails and inline document
previews (PDF/Office via pdftoppm + LibreOffice), plus review hardening
(converter concurrency limiter, bounded preview cache, fixed detected-doc
preview routing). See PR #120.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

# Conflicts:
#	src/web/public/styles.css
2026-06-14 08:49:36 +02:00
Claude (Codeman maintainer) 5fbe451c26 fix(attachments): harden document preview/thumbnail path (review of #120)
Follow-up hardening applied during review of PR #120, addressing the
adversarial multi-agent findings:

- fix(preview): render auto-detected (workspace, unregistered) DOCX/PPTX via
  the file-preview route and PDFs via file-raw in openFilePreview. Previously
  the Preview button fell through to file-content, dumping the binary Office/PDF
  bytes as mojibake, and the new file-preview route was unreachable dead code.
  (MAJOR: file-preview-route-unreachable-detected-office)

- perf(convert): add a global converter-concurrency limiter
  (document-conversion-limiter.ts) wrapping every pdftoppm / soffice /
  powershell spawn, so N simultaneous preview/thumbnail requests can no longer
  fork unbounded converter processes. Default cap 3, CODEMAN_MAX_DOCUMENT_CONVERSIONS.
  (MAJOR: no-converter-concurrency-limit)

- fix(cache): bound the converted-PDF disk cache with LRU-by-mtime eviction
  (pruneDocumentPreviewCache, default 100 files, CODEMAN_MAX_PREVIEW_CACHE_FILES),
  run after each successful conversion. Was unbounded.
  (MAJOR/MINOR: preview-cache-unbounded-disk-growth)

Tests: document-conversion-limiter.test.ts, document-preview-cache-eviction.test.ts,
and route coverage for the four new endpoints in
routes/file-routes-preview-thumbnail.test.ts (closes the missing-route-test gap).
Verified end-to-end against real pdftoppm (thumbnail render + concurrency cap).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 08:43:57 +02:00
Tenggan ZhangandTeigen 99e537ef1a feat(settings): add Opus 4.6 model options to Claude Model picker (#124)
Add claude-opus-4-6[1m] (1M context) and claude-opus-4-6 to the model
selector dropdown.

Co-authored-by: Teigen <teigenzhang@gmail.com>
2026-06-14 08:13:13 +02:00
arkonandClaude Opus 4.8 67c7973aa5 docs: document plan-usage telemetry feature in CLAUDE.md
- New "Plan-usage chip" Key Pattern: statusLine telemetry (rate_limits) →
  injected statusLine exporter → POST /api/status-telemetry (auth-exempt) →
  usage-telemetry.ts parse → SSE session:statusTelemetry → opt-in header chip,
  with plan-usage-latest.ts replaying the last value in the SSE init snapshot.
- Architecture map: add src/usage-telemetry.ts + src/web/plan-usage-latest.ts;
  bump route modules 15→16 and handlers ~136→~140 (status-telemetry route,
  attachment file routes).
- Security: note /api/status-telemetry shares the hook auth-bypass path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 07:34:02 +02:00
arkonandClaude Opus 4.8 534712e50f fix(usage): address code-review findings in plan-usage telemetry
Review of the plan-usage chip feature (commits since 1.0.0) surfaced several
issues; this fixes all confirmed findings:

- HIGH: applyStatusLineConfig clobbered a user's hand-authored statusLine on
  the enable path (the isOurs guard only protected disable). Now bails out when
  an existing statusLine isn't ours, on both the enable and disable paths.
- MED: StatusTelemetrySchema used z.optional() (rejects null) on Claude's
  undocumented statusline fields — a single stray null 400'd the entire POST and
  silently killed the chip's data feed. Switched the modeled fields to .nullish().
- MED: dropping the Token Count / Show Cost header toggles left their features
  reading settings.showTokenCount/showCost, but saveAppSettings rebuilds settings
  fresh from the DOM, dropping those keys and resetting them to defaults on every
  save (re-enabling the token chip with no UI to turn it off). Preserve the prior
  stored preference.
- telemetrySignature keyed on contextUsedPercentage (never displayed) and the raw
  unrounded %, churning a redundant SSE broadcast + localStorage write + identical
  chip re-render on every assistant message. Now keys on the rounded displayed
  window values only.
- Plan-usage chip flashed hidden on load (no server-side reveal): renderIndexHtml
  now strips header-plan-usage--hidden when enabled, matching btn-multimonitor;
  fixes the FOUC and makes the "server renders initial state" comments accurate.
- Serialize all settings.local.json read-modify-write writers in hooks-config via
  a shared per-path mutex (previously lock-free; concurrent session-create +
  settings-toggle on the same repo could lose writes).
- Hardened the chip's innerHTML against any future string field; removed the dead
  _latestPlanUsage field; clamped ctx% in the footer formatter; corrected the
  session-create comment (the path is add-only by design — a per-repo settings
  file is shared by sibling sessions).
- Tests: new test/routes/status-telemetry-routes.test.ts (route behavior, dedup,
  null-tolerance) + NaN/Infinity/fractional and signature-churn unit tests; made
  server-index-title.test.ts deterministic against the ambient settings.json.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 07:33:52 +02:00
arkonandClaude Opus 4.8 f69cd4874c feat(settings): drop Token Count + Show Cost header toggles, move Plan Usage Limits to top
The header Token Count and Show Cost ($) display options are superseded by the
Plan Usage Limits chip, so remove both toggles from App Settings → Header
Displays along with their read (populate) and write (save payload) wiring in
settings-ui.js. Relocate the Plan Usage Limits toggle to the top of the section
for easier access.

Header token-chip render logic is left intact (toggles-only change): the chip
keeps its existing default behavior, it's just no longer user-toggleable.

Verified e2e against an isolated instance with Playwright: section now leads
with Plan Usage Limits; Token Count/Show Cost elements are gone; openAppSettings
(populate) and saveAppSettings (payload build) run with no console errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 06:44:42 +02:00
arkonandClaude Opus 4.8 1ac3c09054 docs(usage): update plan-usage design doc to match what shipped
Rewrite to the as-built design: header chip (account limits, green/yellow/red)
+ session-status footer split; fixed /api/status-telemetry endpoint; curl -sk;
add-only create injection + settings-toggle reconcile; no CASES_DIR gate; chip
robustness (live SSE + init-snapshot replay + localStorage); and the E2E bugs
that earlier builds hid. Status: shipped/pushed, not released.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 06:28:11 +02:00
arkonandClaude Opus 4.8 95fb5fc226 feat(usage): replay last-known plan usage in the SSE init snapshot
The header chip previously only repopulated on reload from per-browser
localStorage, so a fresh browser (or cleared storage) stayed blank until a
session next rendered telemetry. Store the latest broadcast telemetry
process-wide (plan-usage-latest.ts) and include it as `planUsage` in
getLightState — the per-connection SSE init snapshot — so handleInit paints
the chip immediately on every fresh load / reconnect, authoritative over the
localStorage restore. Null until the first telemetry of the process.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 06:15:34 +02:00
arkonandClaude Opus 4.8 eae225bf9a fix(usage): make plan-usage chip work for every user, not just on enable
Two changes so the feature works for any user the moment they enable it,
without manual steps or per-client state:

- Reconcile on settings change: PUT /api/settings now applies the statusLine
  exporter across all ACTIVE Claude sessions' working dirs when
  showPlanUsageLimits is toggled (inject on enable, remove on disable). This is
  server-side and authoritative, so existing sessions get the footer + feed the
  chip immediately — no need to create a new session, no dependency on a
  browser's synced localStorage.

- Create is now ADD-ONLY: never remove the statusLine on session create.
  Sessions in a repo share one settings.local.json, so a single create-with-false
  (e.g. a client whose synced setting hadn't loaded) was yanking the statusLine
  out from under all other live sessions in that repo, killing their footer and
  the chip's data feed. Removal now happens only via the explicit settings toggle.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 06:00:30 +02:00
arkonandClaude Opus 4.8 4d9d93dfff fix(usage): make plan-usage chip work end-to-end + session-status footer
End-to-end testing on the real install surfaced several issues the unit
tests missed:

- Injection gate excluded real sessions: gated on workingDir under CASES_DIR,
  but sessions run in linked cases / real repos. Drop the gate (match
  updateCaseModel, which writes settings.local.json unconditionally).
- statusLine curl failed on HTTPS: prod is loopback HTTPS with a self-signed
  cert; `curl -s` returns 000. Use `curl -sk` (loopback only). applyStatusLineConfig
  now also updates an out-of-date ours-command so the fix propagates.
- Footer hijacked by limits: the in-terminal statusline now shows CURRENT
  SESSION status — `Opus 4.8 (1M context)  in:562,411 out:1,188  ctx:56%` —
  while the account-wide plan limits live only in the header chip.
- Chip blank after reload: persist last-known to localStorage and restore on
  load (account-global, slow-moving; 12h freshness guard).
- Readability + color: per-window green/yellow/red by usage (<60 / 60–84 / ≥85),
  bolder labels and values.
- Drop the renderIndexHtml strip (client-side reveal only, response-viewer
  pattern) — fixes server-index-title test fragility to local settings.

Footer fields flow through context_window.total_input_tokens/total_output_tokens
(schema + parser). Tests updated; verified live (footer, chip, colors, reload
persistence) on the real install.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 05:44:54 +02:00
arkonandClaude Opus 4.8 c82f6c802e feat(usage): plan usage limits header chip via statusLine telemetry
Surface Claude subscription plan usage limits (5-hour rolling + 7-day
weekly: percent used + reset time) in the header, opt-in via App Settings
→ Display → "Plan Usage Limits" (default OFF, no behavior change when off).

A Codeman-managed Claude statusLine exporter forwards the rate_limits JSON
to a new auth-exempt POST /api/status-telemetry (same loopback + hook-secret
gate as /api/hook-event); parsed telemetry broadcasts over SSE
session:statusTelemetry to a header chip (amber >=80%, red >=95%, reset
times on hover). The exporter prints the same summary back as the
in-terminal footer (print-through).

- src/usage-telemetry.ts: pure parser/formatter (epoch-sec -> ms, clamp,
  change signature) + test/usage-telemetry.test.ts
- hooks-config.ts: generateStatusLineCommand + applyStatusLineConfig
  (add/remove; never clobbers a user's own statusLine)
- session-routes.ts: inject gate (Claude-only, Codeman-managed cases),
  driven by create-payload statusLineTelemetry (session-ui.js)
- schemas.ts: StatusTelemetrySchema + showPlanUsageLimits + payload field
- frontend: header chip, applyHeaderVisibilitySettings toggle,
  renderIndexHtml strip, _onSessionStatusTelemetry handler

Schema empirically confirmed against Claude Code 2.1.177 (Claude Max):
only five_hour/seven_day windows exist (no Opus-weekly field); rate_limits
is absent before the first API response and for non-subscriber auth. Design
+ verification method in docs/usage-limits-display-plan.md.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 05:00:28 +02:00
arkonandClaude Fable 5 0809f59f0f chore: version packages — Codeman 1.0.0
Bumps aicodeman 0.9.14 → 1.0.0 (theme skins + first stable release).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-13 23:30:34 +02:00
arkonandClaude Fable 5 eda95adaa9 feat(ui): theme skins — OG Codeman, Daylight Green, Daylight Blue
Add a per-device skin switcher in App Settings → Display:
- Three skins via html[data-skin]: og (original Codeman look),
  daylight-green, and daylight-blue (new default). Per-skin CSS-variable
  token blocks; the v1.0 "Carbon Aurora" component polish is scoped to
  non-og skins and parameterized so green/blue differ only by token values.
- Self-hosted Manrope (UI) + JetBrains Mono (terminal) variable fonts,
  served from /fonts (no external CDN, CSP-safe via font-src 'self').
- Per-skin xterm terminal theme with live re-theming of open terminals on
  skin change; skin-aware --term-bg so the terminal background fills cleanly
  (fixes the variable-height gap above the toolbar).
- Pre-paint inline script applies the saved skin before first paint (no
  flash); persisted per-device in localStorage + the settings blob, and
  kept out of the server settings payload (device-local).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-13 23:21:47 +02:00
Teigen 41a209e96d fix(cjk): hide CJK textarea on welcome screen and fix vertical centering
- Guard `_updateCjkInputState()` with `activeSessionId` check so the
  `position: fixed` CJK textarea doesn't float over the welcome overlay
- Call `_updateCjkInputState()` in `showWelcome()`/`hideWelcome()` to
  sync CJK visibility on session enter/leave
- Add `padding: 12px 10px` to `.cjk-input-visible textarea` for proper
  vertical centering of input text
2026-06-13 18:49:57 +08:00
Teigen 7102fdb23a fix(mobile): restore response-viewer eye button on phones
The header-right tray (02fa3f3) was reworked into a position:fixed
collapsible panel with a hamburger toggle, but the toggle button, its
JS, and CSS were later reverted on master while the container kept a
static `mobile-collapsed` class. With no expand mechanism left, the
header-right stayed display:none on mobile, so the response-viewer eye
icon was unreachable even with "Response Viewer" enabled — desktop was
fine because the media-query rule doesn't apply there.

Restore the simple inline always-visible header-right layout (dev's
known-good state). The showResponseViewer setting still controls the
eye's --hidden marker class.

Verified on iPhone viewport: eye visible (26x26) with setting on, hidden
with setting off, tap opens the viewer; desktop eye unaffected.
2026-06-13 18:49:11 +08:00
Aamer Akhter 49c92e4723 COD-38 document attachment previews + thumbnails (attachment cards)
Builds on the COD-37 registry: surfaces detected/registered attachments as
dismissible cards with a first-page thumbnail and an inline preview — the
consumer the registry PR deliberately deferred.

Backend:
- document-thumbnailer: first-page PNG thumbnails (PNG passthrough; PDF via
  pdftoppm; Office via the preview cache).
- document-preview-cache: disk-cached DOCX/PPTX -> PDF conversion (LibreOffice
  / PowerShell COM), in-flight dedup, multi-converter fallback.
- file-routes: serveConvertedPreview / serveThumbnail + four routes —
  GET .../attachments/:id/preview, .../thumbnail and the workspace-path
  file-preview / file-thumbnail. Reuses the registry's TOCTOU-safe
  resolveServableAttachmentPath, so previews stream the freshly-resolved path.
- server: enrich detected attachment events with a thumbnail route.
- image-watcher: .png now routes to attachment:detected — this PR adds the card
  consumer, so the screenshot popup is no longer its only handler.

Frontend:
- panels-ui: attachment cards (addAttachmentCard, lazy stack, Clear-all,
  per-session cleanup) plus a 3-arg openFilePreview that renders registered
  attachments inline (image/PDF) or via the server-converted PDF (docx/pptx).
- app.js: wire attachment:detected -> _onAttachmentDetected and card state.
- styles: attachment-card + stack styling.

Verified: tsc / eslint / prettier / frontend-syntax clean; new thumbnailer +
preview-cache unit tests pass; full test:ci green (2861 passed); card render +
preview overlay + dismiss verified in-browser.
2026-06-12 09:09:14 -04:00
Ark0N 3f2c23cb0f Merge pull request #119 from aakhter/pr/cod-37-attachments
Add server-side attachment pipeline (registry, magic-link, path guard)

Review fixes (2767e80): force-confine the terminal magic-link scan path to the
session workspace (closes a prompt-injectable arbitrary host-file read primitive
that broadcast over SSE); keep PNG on the image-popup path (the attachment UI
consumer is out of scope, so rerouting it broke the screenshot popup); serve the
re-resolved path (TOCTOU); 50MB raw cap; per-session registry cap; CLI .env via
dataPath(). Documented in security-architecture.md.
2026-06-11 10:32:34 +02:00
arkon f7ce8e4767 fix(attachments): harden registry + close magic-link injection vector
Security (MAJOR): the terminal-output codeman://attach scanner registered any
matching path server-side with no user confirmation and broadcast the rawUrl
over SSE. Terminal output is attacker-influenceable (a prompt-injected session
can print an arbitrary path), so on the default no-auth deployment this was an
arbitrary host-file (png/pdf/docx/pptx/md/txt) read primitive reachable by any
SSE client. Magic-link registration is now force-confined to the session
workspace (forceWorkspaceConfinement) regardless of the global confine setting;
deliberate cross-workspace attach still works through the explicit,
Origin-guarded POST /attachments route and 'codeman attach' (which POSTs
directly inside a managed session). Documented in security-architecture.md.

Regression (MAJOR): .png was rerouted from the image-popup path to
attachment:detected, which has no frontend consumer — silently breaking the
dropped/pasted-screenshot popup. PNG stays on image:detected; only pdf/docx/pptx
(which never had a popup) emit attachment:detected.

Also:
- raw route streams the freshly-resolved path, not the stored one, so a
  post-registration symlink swap can't redirect the stream (TOCTOU).
- 50MB cap on the attachment raw route, matching file-raw / download.
- per-session attachment registry cap (200) to bound the POST path.
- CLI reads creds via dataPath('.env'), honoring CODEMAN_INSTANCE.

Tests: forced-confinement reject/allow cases; PNG popup-path assertions updated.
2026-06-11 10:27:09 +02:00
Aamer Akhter f1c64994ad COD-37 add server-side attachment pipeline (registry, magic-link, path guard)
Adds the foundation for serving local files to the browser as live external
attachments with a stable id, so requests never carry arbitrary absolute paths.

- attachment-registry: in-memory, session-scoped registry. registerExternalAttachment
  validates an absolute path, resolves symlinks, enforces the path guard, and mints
  an `att_<uuid>` id; records are cleared when the session is removed.
- attachment path guard: a configurable blocklist (secret locations + /root,/etc
  trees, extendable via attachmentBlockedPaths / CODEMAN_ATTACHMENT_BLOCKED_PATHS)
  plus an optional, default-off workspace-confinement mode. Shares one
  sensitive-path blocklist (web/sensitive-path.ts) with /api/download, which is
  refactored to use the extracted module instead of an inline copy.
- terminal magic links: the session scans output for codeman://attach?path=... and
  emits `attachmentRequested`; the web server registers the file and broadcasts an
  `attachment:detected` SSE event. `codeman attach <path>` (CLI) prints the magic
  link or POSTs directly when a session id is known.
- image watcher: detects png/pdf/docx/pptx dropped into a session's working dir and
  emits `attachment:detected`.
- routes: POST /api/sessions/:id/attachments (register) and
  GET /api/sessions/:id/attachments/:attachmentId/raw (serve), both re-checking the
  guard before streaming.

Document previews/thumbnails and the attachment-history drawer build on this
foundation and land separately.

Verified: tsc --noEmit, lint, format, frontend-syntax, full test:ci (2846 passed),
and a server boot smoke (/api/status 200).
2026-06-11 10:27:09 +02:00
Ark0N 12c8e080c1 Merge pull request #118 from aakhter/pr/cod-81-snapshot
feat(terminal): snapshot-replay on tab switches (xterm serialize + live pane capture)

Review fixes (9893a7f): bounded/hardened xterm snapshot persistence — shell-session skip, true LRU eviction, localStorage quota-deadlock fix with evict-and-retry, OSC-strip regex tightened.
2026-06-11 10:22:59 +02:00
arkon 9893a7f64a fix(terminal): bound + harden xterm snapshot persistence
- Skip snapshot save for shell sessions (restore is gated on mode!=='shell',
  so they only burned a serialize() + cache slot + localStorage quota).
- In-memory cache: delete-before-set so eviction is true LRU, not FIFO that
  could drop the most-recently-used session.
- localStorage: extract _persistXtermSnapshot — evict to a fixed key budget
  regardless of session liveness (the old prune only dropped dead keys, so
  >10 live sessions at the 20-session target deadlocked the quota) and
  evict-and-retry on quota errors (the old prune ran only after a successful
  setItem, so a full quota permanently disabled persistence).
- Tighten the OSC-strip regex in _isUsableXtermSnapshot to stop at ST.
- Update the structural test's usability-gate assertion to not depend on a
  fixed byte window.
2026-06-11 10:16:47 +02:00
Aamer Akhter 5b2da424a1 feat(terminal): snapshot-replay on tab switches (xterm serialize + live pane capture)
Switching away from a session and back replayed only the server's byte
history. For TUI modes (codex especially) that shows just the latest
repaint — the idle banner — because the TUI drops earlier conversation
from its current frame. This restores the actual on-screen view.

Two complementary mechanisms:

- Client: load xterm's SerializeAddon and snapshot the rendered state
  (viewport + scrollback + colors) per session on switch-away, restoring
  it for an instant first paint on switch-back. The snapshot is only the
  first paint — the canonical /terminal frame is still fetched and
  reconciled (restoredSnapshot/clearedForBusy force the replay). Snapshots
  are LRU-bounded in memory (<=20) and persisted to localStorage
  (<=256KB each, <=10 sessions, stale-pruned) so they survive tab discard.

- Server: GET /api/sessions/:id/terminal prepends the live tmux pane
  buffer (via the existing captureActivePaneBuffer) ahead of the byte
  history, cleared between, so replay reflects the current frame.

Also fix formatPaneSnapshot dropping the rightmost column of every
captured row: it painted to cols - 1 out of caution about last-column
autowrap, but every row is followed by an absolute cursor-position CSI
that cancels xterm's pending-wrap, so painting the full width is safe.

The SerializeAddon is built from @xterm/addon-serialize (new dependency)
into the vendor bundle by postinstall.js (dev) and build.mjs (prod),
matching how the other xterm addons are vendored.
2026-06-10 19:56:30 -04:00
Ark0N aa84447899 Update README to include 'Terminal' in description 2026-06-11 00:38:34 +02:00
Ark0N a0e1a2e33b Update README.md 2026-06-11 00:36:51 +02:00
arkonandClaude Fable 5 6da22f0db0 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 23:06:12 +02:00
Ark0N dc9c4b3bda Merge pull request #117 from aakhter/pr/cod-86-codex-frontend
fix(codex): smaller first-frame write budget + scroll-up grace for codex
2026-06-10 22:59:40 +02:00
Ark0N 1cf5c8c8ad Merge pull request #116 from aakhter/pr/cod-35-codex-polish
fix(codex): strip alt-screen + scrollback-erase from the codex byte stream
2026-06-10 22:52:02 +02:00
arkonandClaude Fable 5 7eda39e7f7 fix(codex): reassemble chunk-split sequences before the strip; mouse parity on replay
Review fixes:

- Hold back a trailing partial CSI (digit-only intro, ≤7 chars) in
  _handleTerminalOutput and prepend it to the next chunk. PTY chunk
  boundaries are arbitrary, so '\x1b[?1049h' can arrive as '\x1b[?104' +
  '9h' — the per-chunk strip misses it, xterm obeys the reassembled toggle,
  and (with the matching ?1049l stripped) stays stuck in the scrollback-less
  alt buffer until the next replay. Complete sequences are never held; the
  carry resets with the other buffers in _resetBuffers.

- Replay path now also strips mouse-tracking enables (?1000-?1007), matching
  the live strip: buffers persisted BEFORE the live strip existed can still
  carry them, and a replayed ?1006h re-hijacks the scroll wheel.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 22:47:02 +02:00
Ark0N f0db5f827f Merge pull request #115 from aakhter/pr/cod-78-security
feat(security): hook-event auth secret + tunnel password guard
2026-06-10 22:36:47 +02:00
arkonandClaude Fable 5 aa4e1ce9cf fix(security): deliver the hook secret to hooks + isolate its rate-limit bucket
Review fixes for COD-54:

- Generated hook curl commands now present X-Codeman-Hook-Secret, read from
  the secret file AT EXECUTION TIME via $CODEMAN_HOOK_SECRET_FILE (exported
  into every managed session's env by tmux buildEnvExports / the direct-PTY
  env builders). Without this, every local hook 401'd the moment a managed
  tunnel came up — the enforcement existed but nothing presented the secret.
  Path-not-value keeps the secret off command lines and out of config files,
  and running sessions pick up a newly generated secret with no respawn;
  server.start() ensures the file exists up front.

- Hook-secret failures now count into a DEDICATED per-IP bucket
  (hookSecretFailures) instead of the shared authFailures map. Legacy
  (pre-secret) hook configs fire constantly from 127.0.0.1; counting their
  401s against the shared bucket would 429 every cookie-less loopback
  request — locking out the Basic-Auth login path (and, through a tunnel,
  every client, since tunneled traffic also arrives as 127.0.0.1).

- docs/security-architecture.md: secret-gated hook exemption, dedicated
  bucket, COD-55 refusal, and the residual caveat for EXTERNAL loopback
  proxies (user-run cloudflared / tailscale serve), which the
  managed-tunnel probe cannot see.

- test/cod54-hook-event-auth.test.ts: +3 tests — login path unaffected
  after hook-bucket exhaustion; generated hooks reference the header +
  $CODEMAN_HOOK_SECRET_FILE without embedding the value; env builders
  export the path only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 22:31:09 +02:00
arkonandClaude Fable 5 b8cb4670dd fix(mobile,respawn-ui): unbury the session-options modal on phones; regroup the Respawn tab
Mobile fixes (user-reported: stuck in Session Options with no way to close):
- Modals now stack at z-index 1300, above the fixed mobile/tablet header
  (z-index 1200) that was burying the modal header and its close button —
  the full-screen modal was undismissable on phones
- Duration presets collapse to one compact 24px row (was a 3-row grid)
- Hide the tab detach (open-in-new-window) button on viewports <=768px

Respawn tab regrouped so its two features read as separate options:
- New green-tinted "Respawn loop" box wraps duration, presets, cycle
  steps, and the status/Enable row — a visual sibling of the blue
  auto-resume box; includes a short explanation of the loop
- Enable/status row moved from the top of the tab to the bottom of the
  box, so it no longer reads as a modal-level confirm button
- Font sizes unified: feature titles match; step checkboxes (2./3.)
  match the step labels (1./4.); "Respawn Cycle" renamed "Cycle Steps"

CLAUDE.md: add usage-limit-patterns.ts to the Session row; app.js ~3.7K

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 21:37:44 +02:00
arkonandClaude Fable 5 4a33b91107 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 20:56:24 +02:00
arkonandClaude Fable 5 28b531fa5b revert(session): drop the cross-device needsRefresh buffer reload
The post-takeover/re-assert needsRefresh made multi-client redraws worse
in practice (fragmented mixed-width frames on the phone) — reverted to
the behavior the user verified as good: cross-device reflows rely on
Ink's own redraw, stale scrollback scrolls away with new output.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 20:50:24 +02:00
arkonandClaude Fable 5 68310619a7 feat(session,mobile): auto-resume on usage limit + mobile view fixes
Auto-resume on usage limit ("token pause" control, opt-in checkbox at the
top of the Respawn tab, off by default):
- usage-limit-patterns.ts (new, pure): detects all Claude Code limit
  messages (1.0.x-2.1.x eras incl. "5-hour limit reached - resets 8pm",
  "You've hit your limit - resets 1:40pm (TZ)", weekly date forms, raw
  "usage limit reached|<epoch>") and parses the reset time. Conservative:
  no parseable future reset time, no action.
- SessionAutoOps: arms a timer at reset+2min, sends Esc (dismisses the
  rate-limit dialog) + "continue"; dedups footer redraws, retries every
  5min on stale times, cancels when Claude starts working, persists and
  re-arms across Codeman restarts (SessionState.autoResumeEnabled/At).
- Respawn guard: cycles are blocked while limit-paused so /clear cannot
  wipe the paused conversation (respawnBlocked reason 'usage_limit').
- POST /api/sessions/:id/auto-resume; SSE session:limitPauseScheduled/
  limitResume/limitResumeCancelled; toasts + status line in the modal.
- Respawn tab tidied: single-row prompt fields, merged behavior row.

Mobile fixes (0.9.8 regressions, user-reported):
- Resize arbitration is now activity-based: a desktop sizing claim only
  blocks phone resizes while the desktop typed within 90s
  (Session.DESKTOP_CLAIM_IDLE_MS). Idle desktop -> phone takes the pane;
  next desktop keystroke re-asserts the desktop layout server-side
  (noteDesktopActivity via ws-routes input). Phones re-send dims every
  30s (visible tab only, skipped while the keyboard is open) so attaching
  under a hot claim self-corrects. Fixes the desktop-width-stream-in-
  narrow-xterm soup (mid-word wraps, tmux dot fill, Ink overdraw).
- Cross-device reflows (takeover/re-assert) emit a debounced needsRefresh
  so all clients reload the buffer instead of stacking ghost Ink frames.
- Keyboard accessory/toolbar lift restored: measure keyboardOffset
  against window.innerHeight (layout viewport), not the shrunken .app -
  on iOS the offset computed to 0, leaving both bars hidden behind the
  OS keyboard with a dead gap above.
- Removed the mobile header utility ("three dots") toggle entirely;
  the headerRight tray stays collapsed on small viewports.

Tests: usage-limit-patterns (36), session-auto-resume (21), resize
arbitration (+6), session routes (+4), respawn guard (+2); MockSession
auto-resume/sizing stubs; mobile tabs test updated for toggle removal.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 20:41:34 +02:00
Aamer AkhterandSaqeb Akhter fe821fb679 fix(codex): smaller first-frame write budget + scroll-up grace for codex
Two render-polish fixes for codex sessions in the terminal write pipeline:

- flushPendingWrites uses a 32KB first-frame budget for codex (vs 64KB for
  other modes). Codex's TUI emits dense synchronized redraws during
  thinking/high-effort phases; a smaller first frame keeps per-frame
  xterm/WebGL stalls short and avoids multi-second main-thread blocks.

- Sticky-scroll now honours a short grace window after a manual scroll-up
  gesture (USER_SCROLL_STICKY_SUPPRESS_MS = 1500ms). High-frequency codex
  "Working (Ns)" status ticks were snapping the viewport back to the bottom
  while the user tried to read earlier output. The wheel/touch scroll
  handlers record the gesture (_noteTerminalUserScroll); flushPendingWrites
  suppresses the auto-scroll-to-bottom and restores the preserved viewport
  via scrollToLine while the grace window is active.

Adds test/terminal-flush-budget.test.ts (vm-sandbox harness over
terminal-ui.js): codex vs non-codex first-frame budget, buffer-load
ownership, and the scroll-up suppression / viewport restore.

Co-Authored-By: Saqeb Akhter <saqeb.akhter@gmail.com>
2026-06-10 13:34:45 -04:00
Aamer AkhterandSaqeb Akhter d7606366a2 fix(codex): strip alt-screen + scrollback-erase from the codex byte stream
Codex's TUI emits alternate-screen toggles (DECSET/DECRST 47/1047/1049),
scrollback-erase (CSI 3 J), and mouse-tracking enables (?1000-1007) during
startup and on every repaint. xterm.js obeys them: it switches to the
scrollback-less alternate buffer, wipes saved lines, and forwards the scroll
wheel to codex — so the user's conversation history both disappears and
becomes unreachable on each tab switch / pane refresh.

Strip these sequences in two places, leaving the visible-viewport erases
(2J / J) intact so codex can still repaint its own rows:

- Session._handleTerminalOutput: filter the live SSE/WS stream and the
  persisted terminal buffer at the source, for mode === 'codex'.
- GET /api/sessions/:id/terminal: apply the same strip to the replayed
  buffer (ALT_SCREEN_TOGGLE_PATTERN / ERASE_SCROLLBACK_PATTERN) so a
  tab-switch replay keeps full scrollback.

Adds test/codex-terminal-output.test.ts covering the strip (alt-screen and
3J removed, 2J/J preserved, Ctrl+L redraws preserved) and confirming codex
output passes through without Ink row-repair mangling.

Co-Authored-By: Saqeb Akhter <saqeb.akhter@gmail.com>
2026-06-10 13:18:45 -04:00
Aamer Akhter 42f0b28c75 feat(security): hook-event auth secret + tunnel password guard
Two hardening fixes for the public-tunnel exposure path (COD-54 / COD-55).

COD-54 — gate the /api/hook-event localhost bypass when a tunnel is up:
`cloudflared --url http://127.0.0.1:port` proxies internet traffic INTO the
loopback origin, so a tunneled hook request arrives with req.ip === 127.0.0.1
and the old bare-localhost bypass would pass it unauthenticated. Now:
- tunnel running  → bypass requires a shared per-instance hook secret
  (X-Codeman-Hook-Secret header; constant-time compare) + per-IP rate limiting
- tunnel not running (loopback-only, the normal case) → unchanged, so
  already-deployed credential-less hooks keep working.
New src/config/hook-secret.ts; auth middleware takes a getTunnelRunning probe
(wired from server.ts via tunnelManager.isRunning()).

COD-55 — refuse starting the Cloudflare tunnel without auth:
enabling the tunnel publishes full terminal control to a public URL; with no
CODEMAN_PASSWORD the auth middleware is inactive and the bind guard never trips
(tunnel binds loopback). PUT /api/settings now refuses tunnelEnabled:true with a
403 (before persisting) unless CODEMAN_PASSWORD is set or
CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1 is acknowledged. New
isUnauthenticatedNetworkAcknowledged() in network-auth-policy; settings-ui
surfaces the refusal as an error toast and reverts the toggle.

Scope: the always-on CSRF/Origin guard, Host-header allowlist, and
network-auth-policy itself are already upstream (#113) and not re-proposed here.

Verification: tsc, eslint, prettier, check:frontend-syntax clean; full test:ci
green (2723 passed), incl. test/cod54-hook-event-auth and
test/routes/system-routes-tunnel-guard.
2026-06-10 12:25:26 -04:00
arkonandClaude Fable 5 055f18fb66 chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 18:20:28 +02:00
arkonandClaude Fable 5 cf2a7f54bf docs(readme): final header tagline — One Dashboard • Any Device (en + zh)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 18:09:22 +02:00
arkonandClaude Fable 5 beeec63f72 fix(terminal): linear-time link-provider regex; always allow blob workers in CSP
cmdPattern's empty-matchable unbounded arg group backtracked exponentially
on wrapped heredoc/table lines — hovering one froze the tab for minutes.
Non-empty tokens + bounded reps make it O(n); regression test extracts the
shipped patterns and pins timing on the real killer shapes.

worker-src 'self' blob: is now unconditional so terminal-ui's _safeYield
tick worker (throttling escape) isn't CSP-blocked on non-gesture installs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 18:06:45 +02:00
arkonandClaude Fable 5 fad32eeaab chore: version packages
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 17:16:13 +02:00
arkonandClaude Fable 5 c1458d8ab8 feat(self-update): launchd-daemon supervisor — rootless restart on headless Macs
A KeepAlive system-level LaunchDaemon (the right setup for headless Macs,
where no GUI login means LaunchAgents never start) is now detected as
supervisor 'launchd-daemon': the updater kills the server PID (passed via
--server-pid) and launchd respawns it on the new dist/ — no root needed.
Detection requires the daemon plist to be bootstrapped AND KeepAlive=true.

Also: on boot, a 'completed-needs-manual-restart' status auto-completes
when the running version matches the staged target, so the stale
'restart Codeman to apply' instruction no longer lingers in the UI.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-06-10 17:08:36 +02:00
231 changed files with 44268 additions and 2438 deletions
+8
View File
@@ -52,8 +52,16 @@ Thumbs.db
# Generated output
out/
screenshots-echo-diag/
screenshots-readme/
screenshots-readme-real/
screenshots-real/
scripts/remotion/out/
# Local UI/README capture scratch (screenshot runs, design mockups). Not build
# output, but never meant for git — an unqualified `git add -A` during a COM has
# swept dirs like these into a release before.
design-explorations/
# Artifacts that should not be tracked
test-results/
tmp/
+432
View File
@@ -1,5 +1,437 @@
# aicodeman
## 1.4.1
### Patch Changes
- **Docker session mode** hardening + fixes, plus a File Viewer header button.
**What Docker session mode is** (recap): a case can run inside an isolated, hardened Docker container instead of on the host, and any of the CLI backends (Claude, Codex, Gemini, OpenCode, or a plain shell) runs inside it. It is a location overlay on cases — not a new session mode — and the container analog of remote-SSH cases: a local tmux pane `docker exec`s into a durable in-container tmux, with exactly one long-lived container per case that multiple sessions share. The workspace, credentials, and conversation transcripts are bind-mounted so the agent is authenticated and resumable; containers are hardened by default (`--cap-drop ALL`, `--security-opt no-new-privileges`, non-root, pids/memory caps, `--init`, never `--privileged` or the docker socket) and export-safe. Start one with the one-click "Run in Docker" checkbox on Create Case, or the Docker tab for full control.
This release fixes the rough edges found running it for real:
Docker cases:
- **Seamless Claude auth in containers**: `~/.claude.json` is no longer bind-mounted as a single file (a mount point that broke Claude's atomic-rename config writes — forcing re-auth and, via failed in-place writes, corrupting the host `~/.claude.json`). It is now seeded as a writable, onboarding-complete copy, so a docker session boots straight to the prompt (no theme picker, login, or folder-trust prompt).
- **Claude-state isolation**: containers no longer bind-mount the whole `~/.claude` directory (which wrote backups/tasks/teams/settings back into the host). Only `~/.claude/projects` transcripts are shared (host watchers + `--resume`); credentials, settings, and stats-cache are seeded as writable copies; everything else stays container-local.
- **Codex/Gemini/gcloud/opencode isolation**: same treatment — codex shares `sessions/` + `history.jsonl` (response-viewer + resume) and seeds `auth.json`/`config.toml`; gemini/gcloud/opencode are whole seed-copies. Containers never write their credential state back into the host dirs.
- **Base image auto-builds on first use**: a missing `codeman/agent:base` no longer blocks case creation or launch; it builds locally on first use (concurrency-safe, with SSE progress toasts).
- **UTF-8 locale**: containers set `LANG`/`LC_ALL=C.UTF-8` so tmux renders Claude's box-drawing correctly (fixes `qqqq` line artifacts).
- **Create Case UI**: larger, collapsed-by-default "Run in Docker" settings with a shorter hint; dockerized cases show a short `(docker)` tag (or the custom host id) in the case menus.
- **Tab naming**: docker/remote (and codex/gemini/opencode) sessions now follow the `w<n>-<case>` convention instead of `codeman-<id>`.
Other:
- **File Viewer header button** (opt-in via App Settings, Header Displays): toggle the file browser panel from the header.
- Fixed a timezone-boundary flaky test in the away-digest route suite.
## 1.4.0
### Minor Changes
- Add **Docker session mode**: a case can now run inside an isolated Docker container instead of on the host, with configurable network / resource / credential settings, multiple sessions sharing one per-case container, and one-click export to move a container (toolchain + workspace) to another machine.
- Docker is a location overlay on cases (not a new session mode), mirroring the remote-SSH feature: a local tmux pane runs `docker exec -it` into a durable in-container tmux server. The container is scoped to the case (`codeman-case-<name>`), so multiple sessions share it; killing one session never stops the shared container.
- New `/api/docker-hosts` CRUD, `/api/cases/docker-link`, and a `/api/quick-start` docker branch. Create Case gains a **Docker** tab. Base image is built locally via `scripts/build-agent-image.mjs` (node + claude/codex/gemini/opencode + tmux, secret-free, arbitrary-uid-writable HOME).
- Hardened by default: `--cap-drop ALL`, `--security-opt no-new-privileges`, non-root, `--pids-limit`, `--memory`==`--memory-swap`, `--init`; never `--privileged` or the docker socket. Convenient credential default bind-mounts host `~/.claude` etc. read-write (never captured by `docker commit`); a sealed profile is opt-in.
- Two-layer durability: reconnect after a Codeman restart reattaches the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript via `--resume`.
- Export / import: full-image (`docker commit` + `save` + workspace tar + manifest) or workspace-only, to one portable `.codeman-container.tgz`; import validates checksums, guards path traversal, and re-tags the loaded image into a quarantined namespace. Instance-scoped boot reaper cleans orphaned containers. New `docker:*` SSE events. Docs in `docs/docker-cases.md`.
- Robustness: sets `CLAUDE_CODE_TMPDIR` in the container so claude launches regardless of workspace path. In-container hooks require the server to be reachable from the container (documented); on a loopback-only bind, idle detection falls back to output-based.
Also wire session, away-digest, and cron header-button visibility toggles in App Settings.
## 1.3.5
### Patch Changes
- a842f2d: fix(auth): slide the session cookie so active users aren't logged out
Re-issue the `codeman_session` cookie on every authenticated request so the
browser cookie lifetime tracks the server-side sliding TTL (the session store
already uses `refreshOnGet`). Previously the cookie was only set on the Basic
Auth path with a fixed 24h lifetime from login, so the browser dropped it
mid-use; the next request arrived cookie-less, fell through to Basic Auth and
popped the native username/password dialog, perceived as a random logout while
actively working.
## 1.3.4
### Patch Changes
- Fix "Run Shell" not switching the terminal to the newly created shell session. Clicking Run Shell created the shell tab but left the previous session's terminal on screen, so you had to manually click the new tab to actually enter it. Root cause: `runShell()` pre-set `activeSessionId` to the new session's id right before calling `selectSession()`, and `selectSession()` early-returns when the requested id already matches the active one, so it skipped the terminal buffer load, tab activation, and focus. Removed the premature assignment in both the local and remote-SSH shell branches so `selectSession()` runs to completion (matching `runClaude`/`runCodex`/`runGemini`/`runOpenCode`, which already avoid this). Verified end-to-end in a real browser with a negative/positive control.
## 1.3.3
### Patch Changes
- Fix terminal scroll-back in Claude sessions, especially on macOS trackpads (#154).
- **Deterministic CLI version detection.** `cliVersion` was often `undefined` because it was scraped from the `Claude Code vX.Y.Z` startup banner, which newer Claude Code builds (2.1.187+) don't reliably print and resumed sessions never show. With the version unknown, wheel-forwarding to Claude's transcript was silently disabled — and since repaint-mode Claude keeps no local terminal scrollback, scrolling up reached nothing. A new `getClaudeCliVersion()` probe (`claude --version`, cached, local-only) seeds the version at session start so forwarding engages. Restored sessions pick it up on restart.
- **Trackpad Shift+scroll.** The wheel handler now reads the dominant axis, so a macOS trackpad's Shift+two-finger scroll — which the browser reports as horizontal `deltaX` — reaches xterm's local scrollback instead of collapsing to a fixed one line per tick.
- **Opt-out setting.** New per-device App Settings → Input → "Wheel Scrolls Local History" (default off) pins the plain wheel to local scrollback (the pre-#144 behavior) for shell and other non-repaint sessions.
- **No more "queued bytes" flicker on scroll.** Wheel-scroll reports now use a fire-and-forget send path (seq-less input frame) instead of the durable exactly-once input queue, so they no longer appear in the pending-bytes connection indicator or churn localStorage. Keystrokes, taps, and clicks still use the durable queue.
## 1.3.2
### Patch Changes
- Make the Cron Jobs modal fully skin-aware and consistent with App Settings' design language.
- **Fix white dropdowns:** `.form-select` had no `appearance` reset and the app set no `color-scheme`, so native `<select>` fields rendered as white OS widgets that ignored the active skin. Selects now use `appearance: none` with an opaque `var(--bg-input)` fill, a `var(--border)` outline, and a custom chevron, so they follow the skin (daylight `#202833`, OG `#1a1a1f`). This is on the shared `.form-select` class, so App Settings, Cron, and every other select match and are fixed together.
- Set `color-scheme: dark` on `:root` so native select option popups, date/time pickers, and scrollbars render dark across all three (dark) skins instead of flashing white.
- Themed the Cron date/time inputs with `var(--bg-input)` / `var(--border)` instead of hardcoded values.
- Fixed the Cron toolbar: "+ New Job" / "Refresh" and the footer Save / Cancel now use the full `btn-toolbar` size (matching the App Settings footer), with a wider gap and a divider under the toolbar for better spacing.
## 1.3.1
### Patch Changes
- Redesign the Cron Jobs modal to match the App Settings styling, and fix a bug that left its create form fully expanded.
- **Fix:** the cron modal's "New Cron Job" form and all of its conditional rows (Launch Command, Prompt File Path, and the once/interval/daily/weekly schedule fields) never actually collapsed — there is no global `.hidden` utility in the stylesheet and the cron modal never scoped its own, so the form opened fully expanded with every field visible at once. Added a scoped `#cronModal .hidden` rule; the form now stays collapsed until "+ New Job" and only shows the fields relevant to the selected agent type, prompt source, and schedule type.
- Sectioned the create/edit form into Basics / Prompt / Schedule / Options with the same section-header dividers used in App Settings, and increased row spacing.
- Styled the agent-type / prompt-source / input-mode / schedule-type dropdowns and the datetime-local / time inputs to share the bordered, rounded, focus-ringed field look.
- Converted the "Auto-close previous run's session" and "Enabled" toggles into App-Settings-style cards (label + description on the left, compact switch on the right).
- Replaced the raw weekday checkboxes with pill toggles that fill with the accent color when selected.
- Restyled the job list rows as hover-highlighted cards with pill badges (agent type, schedule, disabled) and right-aligned actions, and gave the modal a divider-topped Cancel / Save footer.
## 1.3.0
### Minor Changes
- Community release: 16 contributor PRs reviewed (multi-agent adversarial review), fixed, and merged. Thanks to @aakhter, @TeigenZhang, @chatgptkrylor, @kvncrw, and @pirronewantlux529-coder!
**New features**
- **Cron jobs** (#141, @chatgptkrylor): recurring scheduled jobs (once/interval/daily/weekly) that spawn a session and send a prompt when due — CRUD + run history (`/api/cron/*`), ⏰ modal UI, per-job concurrency policy and `autoClosePreviousSession` lifecycle, pure unit-tested next-run math. Distinct from the legacy `ScheduledRun`.
- **Remote host SSH cases** (#145, @aakhter): link cases on remote hosts (`remote-hosts.json`/`remote-cases.json`), launch sessions over ssh into a durable remote tmux (dedicated `-L codeman-remote` socket; adoption-safe naming), per-host command overrides, injection-guarded schemas, remote tmux probe + ConnectTimeout, remote kill on delete, recovery-safe persistence.
- **Command-K session palette + searchable case picker + shortcut registry** (#146, @aakhter): Ctrl/Cmd/Alt+K fuzzy session palette with "Browse all sessions" Session Manager; searchable quick-start case picker (remote-aware labels); rebindable shortcut registry with App Settings → Shortcuts tab and Ctrl+? overlay.
- **Unified session list** (#139, @aakhter): `GET /api/sessions/unified` merges live/persisted/lifecycle/transcript sessions into one deduped list (resumed sessions fold via claudeSessionId alias map).
- **Unified Session Manager UX** (#153, @aakhter): unified welcome list with mode/LIVE badges + per-row kebab menu, `projectKey` plumbing for "View all in this folder", SSE-driven live list refresh, desktop Session Manager header button.
- **Full-scrollback replay** (#148, @aakhter): page reload replays the entire tmux scrollback (`?full=1`, bounded capture with proper maxBuffer) with CRLF normalization for shell panes.
- **WebSocket resilience** (#149, @aakhter): reconnect with preserved exponential backoff, per-tab connection identity (multi-tab safe), ACK re-drive, and a truthful connection chip (WS/HTTP/reconnecting states).
- **PTY-exit circuit breaker + TMUX scrub** (#147, @aakhter): rapid PTY crash-loops trip a breaker (SSE + critical push notification; explicit-restart-only reset); inherited TMUX vars are scrubbed so Codeman-in-tmux doesn't nest.
- **Codex generated-artifact attachments** (#150, @aakhter): codex sessions surface `Saved to: file://…` outputs as attachment cards (realpath-anchored trust, codex-mode-gated, jpg/gif/webp thumbnails).
- **Codex response viewer** (#152, @pirronewantlux529-coder): the eye button now works for Codex sessions via 4-layer rollout resolution (history pin → originator → resume-UUID → cwd) with dedup + injected-context filtering.
- **HEIC paste conversion** (#151, @aakhter): iPhone HEIC pastes convert to JPEG server-side in a worker thread (concurrency-capped, 64MP decompression-bomb guard, magic-byte detection for mislabeled Android HEIFs). Deps: heic-decode + jpeg-js.
- **WebGL renderer toggle** (#140, @kvncrw): per-device setting to switch xterm between WebGL and DOM renderers, cooperating with the GPU-stall auto-fallback marker.
- **Raised terminal history defaults** (#138, @aakhter): tmux history-limit 50k→100k lines, PTY buffer 2MB/1.5MB→32MB/24MB (env-clamped so trim always stays below max).
**Mobile & input fixes**
- CJK input loss fixes: IME state machine, focus routing, Android InputConnection recovery — with content-free diagnostics (#143, @TeigenZhang).
- Tap/click/wheel restored when the server strips mouse DECSETs — version-gated wheel passthrough (claude ≥ 2.1.187), link-click double-fire fix, Shift+wheel documented (#144, @TeigenZhang).
- Response-viewer readability on phones + iOS dvh viewport fix (#142, @TeigenZhang).
**Docs**: CLAUDE.md accuracy audit (18 verified fixes: security hook-bypass description, env-prefix allowlist, state-file inventory, watcher/function names, counts) + documentation for all new subsystems. README gains a user walkthrough (#141).
All PRs went through adversarial multi-agent review; ~60 verified findings (including 12 blockers) were fixed on the contributors' branches before merge. Full test suite green: 3,400+ tests.
### Patch Changes
- bf36eb0: Add a **WebGL Renderer** toggle to Settings → Appearance (desktop). WebGL stays on by default; turning it off forces the DOM renderer for users who hit GPU glitches, without needing the `?nowebgl` URL param. Turning it back on (or `?webgl=force`) clears any stale auto-fallback marker. The existing mobile skip and long-task auto-fallback safety net are unchanged. The skip decision is factored into a pure, unit-tested `shouldSkipWebGL()` helper.
## 1.2.2
### Patch Changes
- Centralize terminal history/scrollback/buffer retention limits into config (PR #137, COD-80).
New `src/config/terminal-history.ts` is now the single source of truth for the terminal scrollback lines, tmux `history-limit`, and server PTY buffer byte caps that were previously scattered as hardcoded literals across `buffer-limits.ts`, `tmux-manager.ts`, and `session.ts`. Each value is overridable (env var or the settings object) and bounds-clamped via a pure `resolveTerminalHistoryConfig()`.
This change is behavior-neutral: the defaults intentionally match the prior hardcoded values (tmux history-limit 50,000; terminal scrollback 50,000; PTY buffer max 2 MB; trim 1.5 MB) and the existing `CODEMAN_MAX_TERMINAL_BUFFER` / `CODEMAN_TRIM_TERMINAL_TO` env overrides are preserved, so runtime behavior is unchanged on its own. It is the mechanism half of a stacked change; a follow-up raises the defaults.
- `buffer-limits.ts` sources `MAX_TERMINAL_BUFFER_SIZE` / `TRIM_TERMINAL_TO` from the resolver.
- `tmux-manager.ts` uses `DEFAULT_TMUX_HISTORY_LIMIT` in place of the hardcoded `history-limit 50000`, gains `setHistoryLimit()` (mux-interface + impl) so a settings change applies to live sessions, and re-applies the limit on `respawnPane` so it survives a respawn.
- `session.ts` threads a per-session `tmuxHistoryLimit` into the tmux spawn calls; `server.ts` exposes `getTerminalHistoryConfig()` on the route ctx and `system-routes.ts` applies a changed `tmuxHistoryLimit` to live sessions immediately.
- `schemas.ts` adds four optional, bounds-clamped settings keys (`terminalScrollbackLines`, `tmuxHistoryLimit`, `terminalBufferMaxBytes`, `terminalBufferTrimBytes`) with a `trim <= max` cross-field check.
- New tests: `test/terminal-history.test.ts` (resolver defaults / clamping / trim<=max / non-number fallback) and `test/terminal-history-schema.test.ts` (settings-schema validation).
## 1.2.1
### Patch Changes
- Fix local echo on iOS Safari when switching into a tab whose session already has output. The on-screen-keyboard "heal" (refit + scroll-to-bottom + overlay re-render + one-shot resize) only ran on a keyboard visibility transition, so switching into a tab while the keyboard was already up never triggered it — leaving the local-echo overlay rendering against stale, off-bottom terminal state. Typed characters were invisible (or mispositioned at the cursor row, far below the actual `❯` prompt) until the user manually hid and re-showed the keyboard. `selectSession` now replicates that heal when the keyboard is already visible, so local echo paints correctly on the first keystroke after a keyboard-up tab switch.
## 1.2.0
### Minor Changes
- Merge four feature PRs and harden them for release.
**Gemini run mode (PR #134, COD-36)** — a third external-CLI backend alongside Codex and OpenCode (`SessionMode` adds `'gemini'`). New `gemini-cli-resolver.ts`, `buildGeminiCommand()` (`--skip-trust`, `--approval-mode {default|auto_edit|yolo|plan}` defaulting to `yolo`, `--model`, `--resume`), `setGeminiEnvVars()` (socket-scoped `tmux setenv` of `GEMINI_*`/`GOOGLE_*` auth incl. Vertex AI), `GET /api/gemini/status` with an install hint (`npm install -g @google/gemini-cli`), run-mode dropdown + welcome "Run Gemini" button + "Run GM" label, `GeminiConfigSchema`, and `GEMINI_*`/`GOOGLE_*` added to the env-override allowlist. Requires tmux (no PTY fallback), like Codex.
**Cross-session search (PR #133, COD-113)** — `GET /api/search?q=&types=&limit=` federates an in-memory search across session metadata, run-summary events, and attachment-history file entries (substring match, hard caps, no FS reads); history-panel search box in the frontend.
**Away digest (PR #136, COD-41)** — `GET /api/away-digest` aggregates "what happened while you were away" (lifecycle log, run summaries, live sessions, daily token stats, recent subagents) into categorized sections behind a header-button modal (hidden on phones).
**Ralph todo-config (PR #135, COD-79)** — per-session `maxTodos` and `todoExpirationMinutes` via `POST /api/sessions/:id/ralph-config`; now persisted in `RalphTrackerState` and read back into the Session Options modal (mirrors `maxIterations` round-trip).
**Review fixes applied on merge:**
- Gemini: fixed two `{success,data}` envelope bugs in `runGemini()` (status check and new-session selection) that made the Run-Gemini button non-functional; fixed `setGeminiEnvVars()` to use the socket-scoped tmux command so Google-auth env injection actually reaches the session.
- Gemini parity: tab-mode badge, kill-dialog label, `codeman doctor` registry entry, `isGeminiAvailable` barrel export, `COLORTERM=truecolor`, and alt-screen/scrollback stripping (Ink TUI, like Codex/Claude).
- Restored four envelope-shape test assertions weakened during the Gemini PR; added a `runGemini()` regression test covering the envelope path.
- Ralph todo-config values now persist across restart and read back correctly instead of always reverting to defaults.
## 1.1.17
### Patch Changes
- Fix the connection indicator flashing "Sending 1B…" on every keystroke. The reliable input-delivery layer (1.1.16) marks each keystroke as briefly pending until its ACK arrives a few milliseconds later, which made the indicator flash on every character while typing on a healthy connection. The indicator is now hidden whenever the connection is healthy and only appears for an actual problem (reconnecting/offline), where it still shows the queued byte count so you know buffered input will be sent.
## 1.1.16
### Patch Changes
- Mobile image uploads, reliable input delivery, and gesture window dragging.
**Mobile image uploads (camera-roll picker / drag-drop / paste).** The "🖼 Image" button now handles real photo batches: up to 20 images per batch uploaded with bounded concurrency and a live "Uploading N/M…" progress toast (with a summary of successes, failures, and whether the 20-cap trimmed the selection). The per-file limit is raised from 10MB to 50MB (`MAX_PASTE_IMAGE_BYTES`, env-overridable via `CODEMAN_MAX_PASTE_IMAGE_BYTES`) so full-resolution phone photos and large screenshots are accepted. Very large images are downscaled to ≤4096px on the longest edge before upload, fixing iOS Safari's ~16.7M-px `<canvas>` limit that previously made huge photos fail to re-encode. Also fixes a latent concurrency bug the batch path exposed where the first parallel uploads to a session raced on creating `.claude-images/` and failed with EEXIST.
**Reliable, exactly-once input delivery.** A "sent" prompt could be silently lost on a flaky connection (e.g. a train): a half-open WebSocket accepts `ws.send()` without error while discarding the frame, and nothing was queued or resent. Input is now recorded durably (localStorage) with a stable clientId + monotonic per-session sequence before delivery, and only dropped once the server ACKs it — delivered over the WebSocket (acked via `{t:'ia',seq}`) or, when the socket is down, over POST in order. A 2s sweep force-reconnects a half-open socket; pending input survives reconnects and page reloads. The server applies each `(clientId, seq)` at most once (`Session.shouldApplyInput`), so an at-least-once resend can never type the prompt twice. Untagged input (curl/legacy) is unchanged. See `docs/reliable-input-delivery.md`.
**Gesture beta: drag agent windows.** With the camera hand-tracking overlay, you can now pinch and move the floating subagent and ultracode run/transcript windows. They keep their glowing connector line to the session tab while moving and can travel across a multi-monitor seam.
## 1.1.15
### Patch Changes
- Security: harden all frontend inline `onclick`/`ondblclick` handlers against a stored-XSS double-context bug.
Many inline handlers interpolated values as `'${escapeHtml(value)}'` — a JavaScript string literal sitting inside an HTML attribute. The browser HTML-decodes the attribute value _before_ parsing the handler source, so `escapeHtml`'s `&#39;` reverts to a literal `'` and a quote-bearing id/name/path/URL breaks out of the JS string into executable code. `escapeHtml` alone is insufficient for this JS-string-within-HTML-attribute context.
All affected handlers now use `escapeHtml(JSON.stringify(value))`: `JSON.stringify` JS-encodes and quote-wraps the value, then `escapeHtml` handles the HTML-attribute layer, so the value round-trips as a single inert string argument.
- ultracode run/agent cards and minimized-tab badges (`ultracode-panel.js`, `ultracode-windows.js`) — PR #132.
- Session tabs (click/rename/gear/detach/close), notifications, subagent windows + dropdowns, the agents/tools/log-viewer/image-popup panels, mux-session monitor rows, and case-management buttons (`app.js`, `notification-manager.js`, `subagent-windows.js`, `panels-ui.js`, `session-ui.js`).
- Two non-`escapeHtml` variants of the same class: a pre-escaped mux-session id in `panels-ui.js` (`selectSession`/`killMuxSession`) and a fully raw, unescaped `phase.id` in `orchestrator-panel.js` (`orchestratorSkipPhase`/`orchestratorRetryPhase`).
The most realistic exploitation vector was file paths in the project-insights log-viewer link, since filenames can legally contain a single quote. Purely numeric interpolations and developer-literal handler strings were left unchanged.
## 1.1.14
### Patch Changes
- Ultracode (Workflow-tool) floating windows — agent transcripts in-page, and minimize-to-tab.
- **Agent transcripts open in-page, connected, instead of a detached browser popup.** Clicking an agent card (in a run window or the dock panel) now opens the agent's live transcript as its own draggable floating window, tied by a connector line to its parent run window (falling back to the run's session tab if that window has since closed) — the same line idiom the run windows use. Re-clicking a card focuses the existing window; closing it removes the window and its line. (Previously this spawned a separate `window.open` browser popup.)
- **The window "−" button now minimizes into the originating session tab**, mirroring the subagent-window idiom. The window genie-animates into its tab and is tracked there; the tab shows an `ULTRA` badge whose hover/click dropdown lists each minimized item (🧬 run windows, 📄 agent transcripts). Click an item to restore its floating window, or dismiss it with ×. A run minimized while still active keeps tracking in the background and its badge auto-clears shortly after the run finishes. Both run windows and agent-transcript windows minimize into the same merged badge.
- Removed the old collapse-to-header behavior that the "−" button previously triggered (now superseded by minimize-to-tab).
## 1.1.13
### Patch Changes
- Keep the `/compact` button in the extended (full) mobile keyboard accessory bar; only the simple bar drops it. (1.1.12 had removed it from both.)
## 1.1.12
### Patch Changes
- Remove the `/compact` button from the mobile keyboard accessory bar. It had been reintroduced in 1.1.10; this removes the button from both the simple and full accessory-bar layouts (the underlying command handler is left in place as inert plumbing).
## 1.1.11
### Patch Changes
- Ultracode (Workflow-tool) run visualization — much better live tracking.
While a run is in flight, the watcher previously showed empty agent slots ("agent N", 0 tokens, raw `wf_…` id as the title) because the detailed completion JSON only lands when the run finishes. The live path now enriches in-flight runs directly from the on-disk transcript tree:
- **Real per-agent stats mid-run** — tokens and tool-call counts are parsed from each `agent-<id>.jsonl` transcript (tool counts match the final accounting exactly; token totals land within ~1% of the completion value), with model and a prompt preview. All mtime-cached (transcripts, journal, and script meta) so idle polls do no extra reads.
- **Readable window/run title** — workflow name, summary, and phases are derived from the persisted `workflows/scripts/<name>-<runId>.js` instead of showing the raw run id.
- **Agent status colors** — done agents show green, working agents show yellow (this also fixes the run/agent status badges, which referenced undefined `--success`/`--warning` CSS variables and were rendering with no color).
- **Connector line** — the floating-window → session-tab line now uses the session-tab accent blue (was purple).
- **Click a run to open its floating window** — clicking a workflow in the dock panel opens (or focuses) its floating window with the connector line, in addition to the auto-popped windows.
- Agents are ordered by journal launch order; concurrent run-detail fetches are de-duplicated.
## 1.1.10
### Patch Changes
- Mobile CJK input, iPad keyboard accessory bar, and terminal touch interaction fixes (PRs #130, #131).
Mobile / CJK (#130):
- Restore reliable real-time CJK (e.g. Pinyin) composition in the always-visible textarea, and refocus input when the terminal is tapped.
- Stop clearing the textarea during `compositionstart` — some IMEs include existing text in the composition region, and clearing it mid-composition corrupted input.
- iPad-specific fixes: `#cjkInput` positioning, paste-dialog placement, and duplicated voice-dictation output.
- Split CJK keyboard positioning by device size (phones vs iPad use different keyboard offsets).
- iPad accessory-bar styling/positioning: moved the accessory-bar and paste-overlay base styles out of the `max-width:1023px`-gated mobile stylesheet so iPad landscape (≥1024px) renders them correctly.
- Raise the toolbar stacking context while the case-settings popover is open so the popover is no longer hidden behind the toolbar.
- Restore the `/compact` button to the keyboard accessory bar (with double-tap confirmation, like `/clear`); the paste dialog now submits pasted text on "Send".
Terminal touch + forced redraw (#131):
- Enable terminal touch interaction on all touch devices and show the stop button on touch devices.
- Add an 8px tap threshold so micro-drift is treated as a tap, not a scroll, fixing cases where a tap failed to register.
- Tap-to-position the cursor via a synthesized mouse report, gated on the live mouse-tracking mode so it never triggers local text selection when tracking is off; let SGR mouse reports through to the PTY even while the CJK input field owns focus.
- Suppress the cursor/momentum side effects of a sub-threshold tap so a jittery tap no longer both positions the cursor and starts a momentum fling.
- New opt-in, per-device "Redraw Terminal" header button (`showRedrawButton`, default off) that forces an xterm redraw via a resize jitter to clear occasional rendering glitches; the resize path now accepts a `force` flag (threaded through the session, HTTP, and WebSocket resize routes) that guarantees a SIGWINCH/redraw at the current device's size without bypassing multi-client resize arbitration.
## 1.1.9
### Patch Changes
- Two welcome-screen tunnel changes:
- **UI (Daylight Blue skin):** the **Cloudflare Tunnel** button is now purple (was orange/yellow), keeping the three welcome buttons visually distinct — Claude blue, Tunnel purple, OpenCode green.
- **Enable a tunnel without `CODEMAN_PASSWORD`, with a warning.** Previously enabling the Cloudflare tunnel with no password set was hard-refused unless you set `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1`. Now you can opt in straight from the browser: clicking the tunnel toggle without a password pops a **security confirm dialog** ("publishes this machine to a public URL with no login — effectively remote code execution; set CODEMAN_PASSWORD instead"), and only on confirm does it enable, sending an explicit per-request `acknowledgeUnauthTunnel:true`. The server logs a loud warning whenever a passwordless public tunnel starts. curl/API/CLI callers are unchanged — still refused unless they set a password, set the env var, or pass `acknowledgeUnauthTunnel:true` — so nothing gets exposed accidentally. The acknowledgment is an action field and is never persisted to settings.json.
## 1.1.8
### Patch Changes
- UI (Daylight Blue skin): give the welcome-screen action buttons distinct colors instead of all reading blue. **Run Claude Code** keeps the blue accent, **Cloudflare Tunnel** now uses Cloudflare's brand orange, and **Run OpenCode** uses an emerald green — so the three are visually distinguishable at a glance. Scoped to the default `daylight-blue` skin only (daylight-green and OG are unchanged), with matching hover/active states and dark ink for contrast. Verified in a real browser: the three buttons compute to blue / orange / green gradients on the welcome overlay.
## 1.1.7
### Patch Changes
- Fix: terminal scroll-up (scrollback) intermittently breaking for **Claude** sessions — most visible on iPhone, where you suddenly "can't scroll up the Claude console."
Root cause: Claude Code periodically emits alternate-screen switches (`\x1b[?1049h`/`\x1b[?47h`/`\x1b[?1047h`), scrollback-erase (`\x1b[3J`), and mouse-tracking enables — typically when it draws a full-screen UI (pickers/dialogs, the boot welcome). xterm.js obeys these by moving to the scrollback-less alternate buffer (or wiping saved lines / hijacking the wheel), so the conversation history becomes unreachable until Claude returns to its normal view. Codeman already stripped these sequences so history stays scrollable, but the strip was gated to **Codex mode only** — Claude (and the equivalent buffer-replay path) let them through.
The strip is now shared via a single `isAltScreenStripMode(mode)` predicate (`codex || claude`) applied at BOTH sites that were Codex-only: the live PTY stream (`Session._handleTerminalOutput`, including the split-across-chunks carry reassembly) and the `/terminal` buffer replay used on tab-switch/reconnect. `shell` is deliberately excluded so full-screen TUIs run from a shell (vim/less/htop) keep their alternate screen; `opencode` is also unchanged.
Verified end-to-end on an isolated instance against a real Claude session: the replayed buffer and live stream now carry zero alt-screen/scrollback-erase/mouse sequences, the terminal stays in the normal buffer with scrollback intact, and touch swipe-up scrolls correctly. Covered by new unit tests (`test/claude-scrollback-strip.test.ts`); the existing Codex strip tests are unchanged.
## 1.1.6
### Patch Changes
- Fix: ultracode floating run windows now pop on a fresh device/browser that loads while a run is already active.
`ultracodeFloatingWindows` syncs from the server (it's a non-display setting), but on a first-time device the SSE `getLightState` run snapshot can seed the run list BEFORE the async settings load resolves — so the floating-window gate read `false` at that instant and skipped any already-active run, leaving the window un-popped until the next ~10s watcher tick. The app now re-runs `syncAllUltracodeFloatingWindows()` once server settings finish loading (in the `loadAppSettingsFromServer().then()` callback), so an in-flight run pops its window immediately. Idempotent: open windows are left as-is, and if the setting is off any premature windows are torn down. Verified end-to-end against a real in-flight run on an isolated instance — a pristine browser (empty localStorage) seeds the setting from the server and pops the active run's window ~0.4s after first paint.
Also corrected a stale `@fileoverview` comment in `ultracode-windows.js` that claimed the floating windows are gated on `showUltracodeAgents`; they are gated on the dedicated `ultracodeFloatingWindows` toggle (only the docked "Ultracode Agents" panel uses `showUltracodeAgents`).
## 1.1.5
### Patch Changes
- Fix: the Ultracode Agents panel's (×) Close button now fully hides the panel.
`closeUltracodeAgentsPanel()` only removed the `open` class, which drops the bottom-docked drawer to its collapsed _peek_ state (the 36px header strip stays visible) rather than closing it — so clicking (×) looked like it did nothing. It now also adds the `hidden` class (`display:none`), mirroring `closeSubagentsPanel()`. It deliberately does NOT flip the `showUltracodeAgents` setting (that also gates the run watcher and floating windows); the header launcher button reopens the panel. Verified in a real browser: after (×) the panel computes `display:none`.
## 1.1.4
### Patch Changes
- Fix: ultracode floating run windows (and the live dock panel) now appear DURING an in-flight Workflow/ultracode run, not only after it finishes.
The Workflow runtime writes the run-state file `…/workflows/wf_<id>.json` only at completion (always a terminal status); while a run is live, its only on-disk state is the sibling `…/subagents/workflows/wf_<id>/` transcript tree. `workflow-run-watcher` previously scanned only the completion file, so it never observed a run until it was already terminal — and the floating-window auto-pop is gated on an ACTIVE run, so it never fired for a live run (the feature was effectively dead for in-flight runs).
The watcher now ALSO scans the `subagents/workflows/wf_<id>/` transcript tree and synthesizes a minimal ACTIVE run (status `running`, agent slots keyed by their `agentId` so the agent-card → live-transcript click still works, `lastActivityAt` from the newest agent/journal mtime, per-agent done/running derived from the run journal's `result` events) when no completion file exists yet. When the run finishes, the real `wf_<id>.json` supersedes the synthesized record (same runId), restoring full phase/token detail and the normal finish → 8s-grace auto-close flow. The watcher stays standalone (it never imports subagent-watcher). Verified end-to-end against a real in-flight run; adds unit coverage for live synthesis, agentId preservation, journal-derived state, empty-dir skipping, and completion-file precedence.
## 1.1.3
### Patch Changes
- Ultracode floating run windows + a dedicated toggle to control them.
- **New: floating ultracode run windows.** When enabled, each active ultracode / Workflow run pops a small draggable window (like the file browser) connected by a glowing line to its originating session tab — the same connector-line idiom as subagent windows. The tab is resolved by matching the run's `sessionUuid` to a session's `claudeSessionId`. The window mirrors the live agent grid (phases, per-agent model / tokens burned / tool calls / state), auto-closes a few seconds after its run finishes, and remembers windows you explicitly dismiss so they don't re-pop. These windows are **additional to** the existing docked "Ultracode Agents" master-detail panel, which is unchanged.
- **New setting "Ultracode Floating Windows"** (App Settings → Display), **default OFF**, independent of the "Ultracode Agents" panel toggle. Either toggle now starts the server-side workflow-run watcher (at boot and on live settings change), so the floating windows work even with the docked panel off.
- Internals: new frontend module `ultracode-windows.js` (load order 15.5); ultracode connector lines are appended into the shared `#connectionLines` SVG within the existing batched read/write reflow pass in `subagent-windows.js`; new `ultracodeFloatingWindows` app-settings key in `schemas.ts`; watcher gating in `server.ts` + `system-routes.ts` now ORs both ultracode toggles.
- Docs: `CLAUDE.md` brought up to date for the 1.1.2 ultracode/workflow-run subsystem (Agents / Frontend / Types / Config inventories, JS load order, a Key Patterns entry) and the new floating-windows feature.
## 1.1.2
### Patch Changes
- Ultracode/Workflow run visualization + subagent discovery fixes.
- **Ultracode / Workflow run visualization** (new, opt-in): App Settings → Display → "Ultracode Agents" (`showUltracodeAgents`, default OFF) adds a master-detail tab that shows ultracode / Workflow-tool runs like Claude Code's "working agents" view — the LEFT pane lists runs and their phases (selectable tasks), the RIGHT pane shows each run's agents with model, live state, tokens burned, and tool calls. Clicking an agent opens its live transcript. Backed by a new standalone workflow-run watcher that reads the per-run state JSON (stripping the heavy embedded script/result/logs so payloads stay small), exposes `GET /api/workflows` and `GET /api/workflows/:runId`, and broadcasts `workflow:run_discovered/updated/removed` SSE events. The header launcher and panel stay hidden until the setting is enabled (the setting is synced across devices, not per-device).
- **Subagent tracking discovery fix**: restored subagent tracking after Claude Code changed the on-disk format from `agent-*.jsonl` to `agent-*.meta.json` (background agents were showing 0). Also discovers workflow-nested subagents under `subagents/workflows/<wf>/` and hardens the meta→transcript upgrade path so an agent re-points to its `.jsonl` transcript once it appears.
- **File viewer**: opens audio, SVG, and other binary files the same way the attachments viewer does.
- **Tooling**: hardened the real-overview screenshot capture script and documented the `deviceScaleFactor` / static-cache gotchas.
## 1.1.1
### Patch Changes
- Six reviewed contributor PRs (all adversarially reviewed and fixed before merge):
- **Markdown sanitizer hardened against mutation-XSS (#126).** The denylist `_sanitizeHtml` is replaced with vendored DOMPurify 3.4.8 (authentic, byte-matched to the official dist) wired via a new `sanitize-html.js` allowlist, with a fail-closed escape fallback. The curated allowlist is genuinely enforced (no `USE_PROFILES` override) so non-markdown tags and svg/math/style/script/event-handler/`javascript:` vectors are stripped while legitimate markdown survives.
- **Hook-event secret now required unconditionally (#127).** The `/api/hook-event` + `/api/status-telemetry` localhost bypass requires the per-instance hook secret whether or not a managed tunnel is running, closing the own-loopback-reverse-proxy gap. A self-heal refreshes pre-secret hook configs in existing cases on spawn so password-protected installs don't silently 401 their hooks. No-password loopback installs are unaffected.
- **`codeman doctor` dependency checker (#125).** New `doctor`/`check-deps` command probes Node, the agent CLIs, tmux, and document converters per environment (linux/darwin/win32/wsl), with grouped or `--json` output and a non-zero exit when a required tool is missing. Requires Node 22+, reports `pdftoppm` (used for PDF/Office thumbnails), and validates `--category`.
- **macOS Option / physical-key session shortcuts (#129).** Tab switching matches physical key codes (`e.code`) so Option+1–9 works on macOS layouts that remap Option, plus Option/Alt+`[`/`]` for previous/next session — without leaking escape sequences into the focused terminal.
- **Desktop session tabs auto-wrap to a second row on overflow (#128)** instead of horizontal scrolling (off when the manual two-row layout is pinned; mobile/tablet unchanged), re-evaluated on window resize.
- **CJK input textarea hidden on the welcome screen (#123)** so it no longer floats over the welcome overlay, and re-shown on session entry; vertical centering fixed.
## 1.1.0
### Minor Changes
- **Plan Usage Limits chip (new).** A header chip now shows your live Claude plan usage — the 5-hour and weekly windows as a percentage — parsed from Claude Code's statusLine telemetry (CLI v2.1.80+). It's opt-in via **App Settings → Display → "Plan Usage Limits"** (default OFF). The toggle is **per-device**: turn it on at your desk without it appearing on your phone. Telemetry collection is decoupled from display, so one device's preference never affects another's, and the last-known value replays instantly on reconnect. Distinct from auto-resume (which reacts to the limit _message_) — this proactively shows the live %.
**Attachments.** New attachment history drawer to browse files referenced by a session (COD-39), plus document previews and thumbnails on attachment cards (COD-38). The header **Attachments button is now opt-in** (default OFF) via **App Settings → Display → "Attachments Button"**, per-device like the Response Viewer button.
**Settings & models.** Added Opus 4.6 options to the Claude Model picker. Removed the legacy Token Count / Show Cost header toggles and moved Plan Usage Limits to the top of the Display settings. Slimmed the Skin picker control to match its row.
**Mobile & header polish.** Restored the response-viewer (eye) button on phones; kept the phone header minimal (settings gear + lifecycle log stay in the toolbar). Added two regression guards so header controls can't silently leak onto the mobile header again — a CI-runnable static policy check plus a real-browser E2E test.
## 1.0.0
### Major Changes
- # Codeman 1.0.0 🎉
The first stable release of Codeman — and it comes with a fresh new look.
**New: theme skins.** Codeman now ships a built-in skin switcher (App Settings → Display → Appearance):
- **OG Codeman** — the original look, preserved exactly.
- **Daylight Green** — a fresh emerald-on-slate theme.
- **Daylight Blue** — bright sky-blue on lifted slate (the new default).
Skins apply instantly, persist per device (with a pre-paint script so there's no flash on load), and re-theme any open terminals live. The system is built on `html[data-skin]` design tokens and self-hosted Manrope (UI) + JetBrains Mono (terminal) fonts — no external CDN, CSP-safe.
**1.0.0 milestone.** This marks the start of the stable 1.x line: the CLI, documented environment variables, and the `{ success, data }` HTTP/SSE API envelope follow semantic versioning (see `docs/versioning-policy.md`).
**Thank you to everyone who helped build Codeman.** This release is dedicated to all of our contributors for their work on the project: Ark0N, Aamer Akhter (@aakhter), Tenggan Zhang (@TeigenZhang), zhouyuan / @sunnyzhouy, jaypark, Marco Migozzi, Skúli Arnlaugsson, Aaron Fields, Loïc Sculier, and Noah Waldner (@noahwaldner). 💙
## 0.9.14
### Patch Changes
- Security hardening for the tunnel exposure path, Codex terminal rendering fixes, and a mobile modal fix.
**Security (PR #115, COD-54/COD-55):**
- `/api/hook-event` localhost bypass is now gated while the managed Cloudflare tunnel is running: tunneled traffic arrives with a loopback source IP, so the bypass additionally requires a per-instance shared secret (`X-Codeman-Hook-Secret`, 256-bit, `~/.codeman/hook-secret`, mode 0600). Locally generated hook commands read the secret file at execution time via `$CODEMAN_HOOK_SECRET_FILE` (exported into every managed session's environment), so the value never lands on command lines or in case configs, and running sessions pick up a new secret without respawn. Failed presentations rate-limit in a dedicated per-IP bucket so misfiring legacy hooks can never lock out the Basic-Auth login path. With no tunnel running, behavior is unchanged.
- Enabling the Cloudflare tunnel now **refuses with 403** when no `CODEMAN_PASSWORD` is set (a public tunnel URL with no auth is effectively public RCE), unless `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` explicitly acknowledges the exposure. The settings UI surfaces the refusal as an error toast and reverts the toggle.
**Codex rendering (PRs #116, #117):**
- Alt-screen toggles (`?47/?1047/?1049`), scrollback-erase (`CSI 3 J`), and mouse-tracking enables (`?1000`–`?1007`) are stripped from the Codex byte stream (live + replay), so conversation history survives tab switches and the scroll wheel scrolls the viewport instead of being hijacked. Sequences split across PTY chunk boundaries are reassembled via a small carry before stripping, so a split `?1049h` can no longer trap xterm in the scrollback-less alt buffer.
- Smaller 32KB first-frame write budget for Codex sessions keeps dense synchronized redraws from stalling the renderer; a 1.5s grace window after a manual scroll-up suppresses sticky-scroll so high-frequency `• Working (Ns)` status ticks no longer snap the viewport back to the bottom while reading earlier output.
**Mobile:** session-options modal raised above the fixed mobile/tablet header (z-index 1300 vs 1200) so the close button is reachable on phones; Respawn tab controls regrouped.
**Docs:** security-architecture.md updated for the secret-gated hook bypass (including the external-proxy caveat) and the tunnel password guard; README documents auto-resume on usage limit.
## 0.9.13
### Patch Changes
- Auto-resume on usage limit ("token pause" control) plus a set of mobile-view fixes for regressions introduced in 0.9.8.
**Auto-resume on usage limit** — new opt-in checkbox at the top of the session Respawn tab (off by default). When Claude stops because a usage limit was reached, Codeman parses the reset time from the limit message, waits until the limit lifts (plus a 2-minute safety buffer), then dismisses the rate-limit dialog (Esc) and sends "continue" so the session picks its work back up automatically. All Claude Code message formats from 1.0.x through 2.1.x are recognized ("5-hour limit reached ∙ resets 8pm", "Limit reached · resets 1pm (America/Chicago) · /upgrade…", "You've hit your weekly limit · resets Mon 12:00am", weekly date forms, and the raw API `usage limit reached|<epoch>` form). Still-limited responses re-arm the scheduler (5-minute retry loop); a pending schedule persists across Codeman restarts and re-arms on boot; respawn cycles are blocked while a limit pause is active so the cycle's `/clear` cannot wipe the paused conversation. New endpoint `POST /api/sessions/:id/auto-resume`; new SSE events `session:limitPauseScheduled`, `session:limitResume`, `session:limitResumeCancelled`; toast/notification on pause and resume, plus a live "resumes at HH:MM" status line in the modal. The Respawn tab layout was also tidied: compact single-row Update/Kickstart prompt fields and a merged options row.
**Mobile fixes (0.9.8 regressions)**:
- **Activity-based resize arbitration** — a desktop sizing claim now only blocks a phone's resize while that desktop has actually typed within the last 90 seconds. Previously any connected desktop tab (even one abandoned hours ago) silently discarded the phone's resize with no fallback, leaving the phone rendering a desktop-width stream in a narrow terminal: mid-word wraps, tmux dot-fill rows, overdrawn garbled text, and misplaced keyboard echo. Now an idle desktop yields the pane to the phone, and the next desktop keystroke automatically restores the desktop layout ("whoever is actively using the session wins"). Phones also re-send their dimensions every 30 seconds (visible tab only, skipped while the virtual keyboard is open) so attaching under a momentarily-active desktop self-corrects.
- **Keyboard accessory bar and toolbar restored on iOS** — the lift offset is measured against the layout viewport (`window.innerHeight`) again instead of the keyboard-shrunken app element; on iOS the offset computed to 0, leaving both bars hidden behind the OS keyboard with a dead black gap above it.
- **Removed the mobile header utility ("three dots") toggle** — the header-utilities tray stays collapsed on small viewports.
## 0.9.12
### Patch Changes
- Documentation refresh — README catches up with the Codex run mode, plus a CLAUDE.md correction.
**README (en + zh-CN)**: Codex is now listed as a third supported AI coding CLI everywhere the docs previously said "Claude Code or OpenCode": the install requirement in Quick Start (now "any combination works", linking to the official Codex CLI docs), the Windows/WSL setup note, the renamed **Multi-CLI** feature bullet (env-prefix gating now reads `CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*`), the Zod schema-validation security bullet, and the architecture mermaid diagram. The header tagline was also finalized to "Claude Code • OpenCode • Codex — One Dashboard • Any Device" in both languages.
**CLAUDE.md**: fixed a stale "Local packages" line that claimed the xterm-zerolag-input local-echo overlay had a copy embedded in `app.js` — it is single-source in `packages/xterm-zerolag-input/`, bundled to the gitignored vendor file, and only consumed by `app.js`, matching the existing single-source gotcha.
## 0.9.11
### Patch Changes
- Fix a terminal freeze on hover (catastrophic regex backtracking) and a CSP violation that disabled the terminal's anti-throttling worker.
**Tab-freezing hover bug**: the terminal link provider's `cmdPattern` (which turns `tail -f /path`-style text into clickable links) used an empty-matchable, unbounded arg group — `(?:[^\s\/]*\s+)*` — that backtracks exponentially on real Claude output, e.g. wrapped `git commit -m "$(cat <<'EOF'` heredoc lines or aligned table rows. Hovering the mouse over such a line hung the page's main thread for minutes ("page unresponsive"). The pattern now uses non-empty tokens with bounded repetition (linear time); all intended command+path link forms still match. New `test/link-provider-regex.test.ts` extracts the shipped patterns from source and pins linear-time behavior on the killer line shapes.
**Blob worker CSP fix**: `worker-src 'self' blob:` is now always present in the CSP (previously only with `CODEMAN_GESTURE=1`). The terminal's `_safeYield` anti-throttling tick worker is created from a Blob URL and was silently blocked on every install, logging a CSP violation on each page load and disabling the worker leg of the render-yield fallback chain.
## 0.9.10
### Patch Changes
- Self-update now restarts automatically on headless Macs supervised by a system LaunchDaemon.
New `launchd-daemon` supervisor kind: when Codeman runs under a bootstrapped, KeepAlive system-level LaunchDaemon (`/Library/LaunchDaemons/com.codeman.web.plist` — the right setup for headless Macs, where LaunchAgents never start because there is no GUI login), the updater no longer ends with "Update staged — restart Codeman to apply". It restarts rootlessly: the update script kills the server PID (passed via `--server-pid`) and launchd respawns it on the freshly built `dist/`. Detection is conservative — the daemon must be bootstrapped in the system domain AND have `KeepAlive` enabled.
Also fixed: a lingering "restart Codeman to apply" status. After a manual restart of a staged update, boot reconciliation now flips `completed-needs-manual-restart` to `completed` once the running version matches the staged target, so the Updates tab stops showing the stale instruction.
## 0.9.9
### Patch Changes
+123 -82
View File
File diff suppressed because one or more lines are too long
+291 -99
View File
@@ -2,14 +2,10 @@
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
</p>
<h2 align="center">The missing control plane for AI coding agents</h2>
<h2 align="center">Mission control for AI coding agents</h2>
<p align="center">
<em>Agent Visualization &bull; Zero-Lag Input &bull; Mobile-First UI &bull; Hardened Security</em>
</p>
<p align="center">
<strong>English</strong> &bull; <a href="README.zh-CN.md">简体中文</a>
<em>Claude Code &bull; OpenCode &bull; Codex &bull; Terminal - One Dashboard &bull; Any Device</em>
</p>
<p align="center">
@@ -20,6 +16,10 @@
<img src="https://img.shields.io/badge/Tests-2861%20total-22c55e?style=flat-square" alt="Tests">
</p>
<p align="center">
<strong>English</strong> &bull; <a href="README.zh-CN.md">简体中文</a>
</p>
<p align="center">
<img src="docs/images/subagent-demo.gif" alt="Codeman — parallel subagent visualization" width="900">
</p>
@@ -34,7 +34,7 @@ curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | b
This installs Node.js and tmux if missing, clones Codeman to `~/.codeman/app`, and builds it.
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai) (or both). After install:
You'll need at least one AI coding CLI installed — [Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [Codex](https://developers.openai.com/codex/cli) (any combination works). After install:
```bash
codeman web
@@ -45,6 +45,7 @@ codeman web
<summary><strong>Run as a background service</strong></summary>
**Linux (systemd):**
```bash
mkdir -p ~/.config/systemd/user
cat > ~/.config/systemd/user/codeman-web.service << EOF
@@ -67,6 +68,7 @@ loginctl enable-linger $USER
```
**macOS (launchd):**
```bash
mkdir -p ~/Library/LaunchAgents
cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
@@ -94,6 +96,7 @@ cat > ~/Library/LaunchAgents/com.codeman.web.plist << EOF
EOF
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
```
</details>
<details>
@@ -103,11 +106,79 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash"
```
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code) or [OpenCode](https://opencode.ai)). After installing, `http://localhost:3000` is accessible from your Windows browser.
Codeman requires tmux, so Windows users need [WSL](https://learn.microsoft.com/en-us/windows/wsl/install). If you don't have WSL yet: run `wsl --install` in an admin PowerShell, reboot, open Ubuntu, then install your preferred AI coding CLI inside WSL ([Claude Code](https://docs.anthropic.com/en/docs/claude-code), [OpenCode](https://opencode.ai), or [Codex](https://developers.openai.com/codex/cli)). After installing, `http://localhost:3000` is accessible from your Windows browser.
</details>
---
## Using Codeman — A Human's Guide
A start-to-finish walkthrough for driving Codeman from the browser. If you just installed, this is where to begin.
### 1. Launch the server
```bash
codeman web # localhost:3000 (loopback only — safe default)
codeman web --port 8080 # custom port (or set CODEMAN_PORT)
codeman web --https # self-signed TLS (only needed for remote access)
codeman web -H 0.0.0.0 # bind LAN — REQUIRES CODEMAN_PASSWORD (see Security)
```
Open the printed URL. The page is a single dashboard; everything below happens there.
### 2. Create your first session
Click **+ New Session** (or **Quick Start**). A session is one AI CLI running in its own tmux-backed terminal. You choose:
| Field | What it does |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Working directory / case** | The folder the agent operates in. A "case" is just a named working dir Codeman remembers. |
| **CLI / run mode** | `Claude` (default), `OpenCode`, `Codex`, `Gemini`, or `Terminal` (plain shell). |
| **Model** | Per-session model (App Settings → Claude Model). A soft default — `/model` still works in-session. |
| **Effort / Ultracode** | Reasoning effort (`low`–`max`) or `ultracode` for dynamic multi-agent workflows. Switchable anytime with `/effort`. |
Hit start — Codeman spawns the CLI via a real PTY and streams it to your browser over SSE.
### 3. Read the dashboard
- **Tabs (top)** — one per session. `Alt+1`-`9` to jump, `Ctrl+Tab` for next, drag to reorder.
- **Terminal (center)** — a real `xterm.js` terminal; full TUIs render correctly. Type directly and press **Enter** to send. `Shift+Enter` inserts a newline.
- **Side panels** — Respawn, Ralph, Orchestrator, Cron, Subagents, Settings (toggled from the toolbar).
### 4. Talk to the agent
- **Type prompts** straight into the terminal — input is delivered exactly-once even across reconnects (a dropped link never loses or double-sends a prompt).
- **Paste or drag-and-drop images** directly into the session.
- **Voice input** — `Ctrl+Shift+V` (Deepgram Nova-3, with auto-silence stop).
- **Attachments** — register external files/docs and preview Office/PDF inline.
### 5. Make it autonomous
| Mode | Use it for | Where |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------------- | ------------------ |
| **Respawn** | Long unattended runs — auto-restarts the CLI on idle/limit, with adaptive timing. Presets: `solo-work`, `overnight-autonomous`, … | Respawn tab |
| **Ralph / Todo** | A self-driving loop that tracks a todo list and keeps working until done. | Ralph tab |
| **Orchestrator** | Turn one goal into a phased plan and drive it to completion across agents. | Orchestrator panel |
| **Cron** | Saved, named jobs on a schedule (`once`/`interval`/`daily`/`weekly`) that spawn a session and send a prompt when due. | ⏰ Cron button |
| **Auto-resume** | Automatically continue after a subscription rate-limit resets. | Respawn tab (top) |
### 6. Reach it from anywhere
- **Phone/tablet** — the UI is fully touch-optimized; scan the desktop **QR code** to log in without typing a password.
- **Outside your network** — `./scripts/tunnel.sh start` opens a Cloudflare tunnel (set `CODEMAN_PASSWORD` first).
- **SSH** — the `sc` chooser attaches to any session from a terminal (`sc` interactive, `sc 2` quick-attach, `sc -l` list).
### 7. Operate & maintain
- **App Settings** — model, effort, theme/skin, notifications, display toggles, per-CLI options.
- **Self-update** — git-clone installs update in place from **Settings → Updates**.
- **Deploy your own changes** — see [Development](#development).
> ⚠️ **Safety:** if you're working _inside_ a Codeman-managed session (`echo $CODEMAN_MUX` → `1`), never run `tmux kill-session` / `pkill claude` directly — use the web UI or `./scripts/tmux-manager.sh`.
---
## Mobile-Optimized Web UI
The most responsive AI coding agent experience on any phone. Full xterm.js terminal with local echo, swipe navigation, and a touch-optimized interface designed for real remote work — not a desktop UI crammed onto a small screen.
@@ -214,6 +285,7 @@ WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE →
```
- **Multi-layer idle detection** — completion messages, AI-powered idle check, output silence, token stability
- **Auto-resume on usage limit** _(opt-in, off by default)_ — when Claude halts on a subscription limit ("You've hit your limit · resets 3pm"), Codeman parses the reset time, waits it out plus a 2-minute safety buffer, then dismisses the rate-limit dialog and sends `continue` — so an overnight run survives the 5-hour window instead of stalling until morning. Recognizes every Claude Code limit-message format, retries if still limited, survives Codeman restarts, and holds respawn cycles while paused so `/clear` can't wipe the waiting conversation. Enable per session at the top of the Respawn tab
- **Circuit breaker** — prevents respawn thrashing when Claude is stuck (CLOSED -> HALF_OPEN -> OPEN states, tracks consecutive no-progress and repeated errors)
- **Health scoring** — 0-100 health score with component scores for cycle success, circuit breaker state, iteration progress, and stuck recovery
- **Built-in presets** — `solo-work` (3s idle, 60min), `subagent-workflow` (45s, 240min), `team-lead` (90s, 480min), `ralph-todo` (8s, 480min), `overnight-autonomous` (10s, 480min)
@@ -259,10 +331,10 @@ The title is templated into the served HTML on first byte, so it's correct from
### Smart Token Management
| Threshold | Action | Result |
|-----------|--------|--------|
| Threshold | Action | Result |
| --------------- | --------------- | ---------------------------------- |
| **110k tokens** | Auto `/compact` | Context summarized, work continues |
| **140k tokens** | Auto `/clear` | Fresh start with `/init` |
| **140k tokens** | Auto `/clear` | Fresh start with `/init` |
### Notifications
@@ -293,17 +365,33 @@ PTY Output → 16ms Server Batch → DEC 2026 Wrap → SSE → Client rAF → xt
## More Features
- **Self-update** — git-clone installs under systemd/launchd update in place from **App Settings → Updates**: it detects the latest release, auto-stashes a dirty tree, and streams build progress across the service restart (npm installs report as non-updatable)
- **Dual-CLI** — run **Claude Code** or **OpenCode** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Multi-CLI** — run **Claude Code**, **OpenCode**, or **Codex** per session; env-var prefixes auto-gate (`CLAUDE_CODE_*` vs `OPENCODE_*` vs `CODEX_*`). See [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Docker sessions** — run a case inside an isolated, hardened container. One checkbox on **Create New** spins up a container with sensible defaults and starts the agent inside it; multiple sessions share one per-case container; export a container + its workspace to a portable `.tar.gz` to move it to another machine. See [`docs/docker-cases.md`](docs/docker-cases.md)
- **Effort & Ultracode** — set a per-session default effort (`low`–`max`) or enable **ultracode** (dynamic multi-agent workflows). Soft defaults only — switchable anytime with `/effort` in-session. Extended-thinking budget is configurable too
- **Voice input** — dictate prompts with Deepgram Nova-3 (Web Speech API fallback): toggle recording, auto-silence stop, live level meter (`Ctrl+Shift+V`)
- **Image input** — paste or drag-and-drop images straight into a session
- **Gesture control** *(opt-in)* — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Display
- **Multi-monitor span** *(macOS)* — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
- **Gesture control** _(opt-in)_ — a MediaPipe hand-tracking overlay to grab/drag session windows and pinch buttons, hands-free. Enable with `CODEMAN_GESTURE=1` + App Settings → Display
- **Multi-monitor span** _(macOS)_ — one click opens a browser window maximized across all displays, so floating agent/gesture panels can cross the physical seam
- **CJK / IME input** — full composition support for Chinese / Japanese / Korean
- **OS notifications & hostname-aware titles** — desktop alerts and tab titles are prefixed `codeman:<host>` so multi-host setups stay unambiguous
---
## Isolated Docker Sessions
Run a case inside its own hardened Docker container instead of directly on your host — for security isolation, reproducible toolchains, and one-click portability.
- **One click** — on **New Case → Create New**, tick **🐳 Run in an isolated Docker container**. Codeman creates the case folder, spins up a container with default settings, and starts the agent inside it. No host/image/network fields to fill in.
- **Resource templates** — expand the checkbox for a **Small / Medium / Large / GPU** preset (memory, CPUs, GPU), or set your own. **Disk is elastic** — storage grows as data flows in, no fixed cap.
- **Shared per-case container** — many sessions can `docker exec` into the same container; killing one session never tears the container out from under the others.
- **Hardened by default** — non-root, `--cap-drop ALL`, `no-new-privileges`, PID/memory caps, never `--privileged` or the docker socket. Your existing `~/.claude` login is bind-mounted (credentials stay on the host, never captured in exports); a **sealed** profile (no host mounts, network off) is one toggle away.
- **Move it to another machine** — export a container's whole environment (toolchain + workspace) to a portable `.tar.gz`, `docker load` it on the other side, and import it into a fresh case.
- **Durable** — reconnect after a restart lands back in the same live agent; a container stop/reboot resumes the conversation from the bind-mounted transcript.
Prerequisite: Docker (or Podman) and the base image — build it once with `node scripts/build-agent-image.mjs`. Full guide: [`docs/docker-cases.md`](docs/docker-cases.md).
---
## Remote Access — Cloudflare Tunnel
Access Codeman from your phone or any device outside your local network using a free [Cloudflare quick tunnel](https://developers.cloudflare.com/cloudflare-one/connections/connect-networks/do-more-with-tunnels/trycloudflare/) — no port forwarding, no DNS, no static IP required.
@@ -371,14 +459,14 @@ Every **60 seconds**, the server automatically rotates to a fresh token. The pre
The design is informed by ["Demystifying the (In)Security of QR Code-based Login"](https://www.usenix.org/conference/usenixsecurity25/presentation/zhang-xin) (USENIX Security 2025), which found 47 of the top-100 websites vulnerable to QR auth attacks due to 6 critical design flaws across 42 CVEs. Codeman addresses all six:
| USENIX Flaw | Mitigation |
|-------------|------------|
| **Flaw-1**: Missing single-use enforcement | Token atomically consumed on first scan — replays always fail |
| **Flaw-2**: Long-lived tokens | 60s TTL with 90s grace, auto-rotation via timer |
| **Flaw-3**: Predictable token generation | `crypto.randomBytes(32)` — 256-bit entropy. Short codes use rejection sampling to eliminate modulo bias |
| **Flaw-4**: Client-side token generation | Server-side only — tokens never leave the server until embedded in the QR |
| **Flaw-5**: Missing status notification | Desktop toast: *"Device [IP] authenticated via QR (Safari). Not you? [Revoke]"* — real-time QRLjacking detection |
| **Flaw-6**: Inadequate session binding | IP + User-Agent stored for audit. Manual session revocation via API. HttpOnly + Secure + SameSite=lax cookies |
| USENIX Flaw | Mitigation |
| ------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- |
| **Flaw-1**: Missing single-use enforcement | Token atomically consumed on first scan — replays always fail |
| **Flaw-2**: Long-lived tokens | 60s TTL with 90s grace, auto-rotation via timer |
| **Flaw-3**: Predictable token generation | `crypto.randomBytes(32)` — 256-bit entropy. Short codes use rejection sampling to eliminate modulo bias |
| **Flaw-4**: Client-side token generation | Server-side only — tokens never leave the server until embedded in the QR |
| **Flaw-5**: Missing status notification | Desktop toast: _"Device [IP] authenticated via QR (Safari). Not you? [Revoke]"_ — real-time QRLjacking detection |
| **Flaw-6**: Inadequate session binding | IP + User-Agent stored for audit. Manual session revocation via API. HttpOnly + Secure + SameSite=lax cookies |
#### Timing-Safe Lookup
@@ -403,23 +491,23 @@ When someone authenticates via QR, the desktop shows a notification toast with t
#### Threat Coverage
| Threat | Why it doesn't work |
|--------|-------------------|
| **QR screenshot shared** | Single-use: consumed on first scan. 60s TTL: expired before the attacker can act. Desktop notification alerts you immediately. |
| **Replay attack** | Atomic single-use consumption + 60s TTL. Old URLs always return 401. |
| **Cloudflare edge logs** | Short code is an opaque 6-char lookup key, not the real 256-bit token. Single-use means replaying from logs always fails. |
| **Brute force** | 56.8 billion combinations, ~2 valid at any time, dual-layer rate limiting blocks well before statistical feasibility. |
| **QRLjacking** | 60s rotation forces real-time relay. Desktop toast provides instant detection. Self-hosted single-user context makes phishing implausible. |
| **Timing attack** | Hash-based Map lookup — no string comparison timing leak. |
| **Session cookie theft** | HttpOnly + Secure + SameSite=lax + 24h TTL. Manual revocation at `POST /api/auth/revoke`. |
| Threat | Why it doesn't work |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------ |
| **QR screenshot shared** | Single-use: consumed on first scan. 60s TTL: expired before the attacker can act. Desktop notification alerts you immediately. |
| **Replay attack** | Atomic single-use consumption + 60s TTL. Old URLs always return 401. |
| **Cloudflare edge logs** | Short code is an opaque 6-char lookup key, not the real 256-bit token. Single-use means replaying from logs always fails. |
| **Brute force** | 56.8 billion combinations, ~2 valid at any time, dual-layer rate limiting blocks well before statistical feasibility. |
| **QRLjacking** | 60s rotation forces real-time relay. Desktop toast provides instant detection. Self-hosted single-user context makes phishing implausible. |
| **Timing attack** | Hash-based Map lookup — no string comparison timing leak. |
| **Session cookie theft** | HttpOnly + Secure + SameSite=lax + 24h TTL. Manual revocation at `POST /api/auth/revoke`. |
#### How It Compares
| Platform | Model | Comparison |
|----------|-------|------------|
| **Discord** | Long-lived token, no confirmation, [repeatedly exploited](https://owasp.org/www-community/attacks/Qrljacking) | Codeman: single-use + TTL + notification |
| **WhatsApp Web** | Phone confirms "Link device?", ~60s rotation | Comparable rotation; WhatsApp adds explicit confirmation (acceptable tradeoff for single-user) |
| **Signal** | Ephemeral public key, E2E encrypted channel | Stronger crypto, but [exploited by Russian state actors in 2025](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) via social engineering despite it |
| Platform | Model | Comparison |
| ---------------- | ------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Discord** | Long-lived token, no confirmation, [repeatedly exploited](https://owasp.org/www-community/attacks/Qrljacking) | Codeman: single-use + TTL + notification |
| **WhatsApp Web** | Phone confirms "Link device?", ~60s rotation | Comparable rotation; WhatsApp adds explicit confirmation (acceptable tradeoff for single-user) |
| **Signal** | Ephemeral public key, E2E encrypted channel | Stronger crypto, but [exploited by Russian state actors in 2025](https://cloud.google.com/blog/topics/threat-intelligence/russia-targeting-signal-messenger) via social engineering despite it |
> Full design rationale, security analysis, and implementation details: [`docs/qr-auth-plan.md`](docs/qr-auth-plan.md)
@@ -427,27 +515,27 @@ When someone authenticates via QR, the desktop shows a notification toast with t
## Security
Codeman launches sessions with `--dangerously-skip-permissions`, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control *who* that is. Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: [`docs/security-architecture.md`](docs/security-architecture.md). **Found a vulnerability?** See [`SECURITY.md`](SECURITY.md) for private disclosure and the list of known limitations.
Codeman launches sessions with `--dangerously-skip-permissions`, so the web UI is by design a remote-code-execution surface for whoever can reach it — the whole security model exists to control _who_ that is. Recent hardening (v0.9.0 + v0.9.5) closes the browser-driven attack paths that bite self-hosted dev tools. Full model: [`docs/security-architecture.md`](docs/security-architecture.md). **Found a vulnerability?** See [`SECURITY.md`](SECURITY.md) for private disclosure and the list of known limitations.
### Network & access
- **Loopback by default** — binds `127.0.0.1`, reachable only from the same machine, so the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` *starts but prints a loud warning* with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with `--allow-unauthenticated-network`)
- **Loopback by default** — binds `127.0.0.1`, reachable only from the same machine, so the no-password default is safe out of the box. Binding a non-loopback host without `CODEMAN_PASSWORD` _starts but prints a loud warning_ with three concrete fixes (set a password, loopback + an authenticated tunnel, or explicitly acknowledge with `--allow-unauthenticated-network`)
- **Optional auth, real sessions** — HTTP Basic via `CODEMAN_USERNAME` (default `admin`) / `CODEMAN_PASSWORD`. Success issues an opaque 256-bit `codeman_session` cookie (`randomBytes(32)`) — validated server-side, not client-signed, so it can't be forged offline (24h TTL, auto-extend, device-context audit log)
- **Per-IP rate limiting** — 10 failed attempts → `429` with `Retry-After` (15-min decay). A valid cookie or correct password recovers *immediately* even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter
- **Per-IP rate limiting** — 10 failed attempts → `429` with `Retry-After` (15-min decay). A valid cookie or correct password recovers _immediately_ even while an attacker hammers the same IP — important because all tunnel traffic shares one loopback IP. QR auth has its own separate limiter
### Always-on browser hardening (v0.9.5)
These run for **every** request — before auth, even on the default no-password loopback install:
- **Host-header allowlist → blocks DNS rebinding.** A custom domain rebound to `127.0.0.1` is rejected with `403 host not allowed` before any handler runs. Allowed: `localhost`, any IP literal, the bind host, `.ts.net` / `.trycloudflare.com` / `.cfargotunnel.com`, the active managed tunnel, and `CODEMAN_ALLOWED_HOSTS` (add custom reverse-proxy domains here — comma-separated; exact host or leading-dot `.suffix` for subdomains)
- **Cross-site Origin / CSRF guard.** On state-changing methods (`POST`/`PUT`/`PATCH`/`DELETE`) the `Origin` must pass the same allowlist, else `403 cross-site request blocked`. A *missing* Origin is allowed (so `curl`, the CLI, and Claude Code hooks keep working); only a present-but-foreign or opaque `null` origin is rejected
- **Cross-site Origin / CSRF guard.** On state-changing methods (`POST`/`PUT`/`PATCH`/`DELETE`) the `Origin` must pass the same allowlist, else `403 cross-site request blocked`. A _missing_ Origin is allowed (so `curl`, the CLI, and Claude Code hooks keep working); only a present-but-foreign or opaque `null` origin is rejected
- **Raw `text/plain` bodies.** The global parser no longer JSON-parses `text/plain`, closing the CORS "simple request" CSRF vector where a cross-site `fetch` could smuggle JSON into a write route with no preflight
- **WebSocket origin validation.** The terminal WS upgrade runs the same Host + Origin check and closes with code `4003` on failure (anti-CSWSH)
- **XSS-escaped agent output.** AI-derived strings (tool names, command arguments, subagent descriptions) are HTML-escaped at every injection site before rendering in the subagent / activity panels
### Input, files & headers
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` env-prefix allowlist gates which settings each CLI can receive
- **Schema-validated inputs** — every API body is checked with Zod v4 schemas; a `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` env-prefix allowlist gates which settings each CLI can receive
- **Path containment** — file routes `realpath` before boundary checks (no TOCTOU); `..`, absolute paths, and symlinks resolving outside the working dir are rejected. Caps: 10 MB text preview / 50 MB raw & download; `/api/download` blocklists sensitive paths (`.env`, `*credentials*`, `~/.ssh/`, `.aws/credentials`). SVG/HTML is served `octet-stream` + `nosniff` + attachment so it downloads rather than executes
- **Security headers** — `Content-Security-Policy` (`default-src 'self'`, every exception enumerated), `X-Content-Type-Options: nosniff`, `X-Frame-Options: SAMEORIGIN`, HSTS over HTTPS, and CORS reflected **only** for `localhost` / `127.0.0.1` / `::1`
@@ -478,73 +566,177 @@ Single-digit selection (1-9), color-coded status, token counts, auto-refresh. De
> Ctrl bindings also accept Cmd on macOS.
| Shortcut | Action |
|----------|--------|
| `Ctrl/Cmd+W` | Kill active session |
| `Ctrl/Cmd+Tab` | Next session |
| `Alt+1`–`Alt+9` | Switch to tab N |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
| `Ctrl/Cmd +` / `-` | Font size |
| `Ctrl/Cmd+?` | Keyboard help |
| `Shift+Enter` | Insert newline (sent to terminal) |
| `Escape` | Close panels & modals |
| Shortcut | Action |
| ------------------------------- | ------------------------------------------------------------- |
| `Ctrl/Cmd+W` | Kill active session |
| `Ctrl/Cmd/Option+K` | Find open session or start a new one |
| `Ctrl/Cmd+Tab` | Next session |
| `Alt/Option+[` / `Alt/Option+]` | Previous / next session |
| `Alt/Option+1`-`Alt/Option+9` | Switch to tab N (physical keys, so macOS Option layouts work) |
| `Ctrl+Shift+{` / `Ctrl+Shift+}` | Move active tab left / right |
| `Ctrl/Cmd+L` | Clear terminal |
| `Ctrl+Shift+R` | Restore terminal size |
| `Ctrl+Shift+V` | Toggle voice input |
| `Ctrl/Cmd +` / `-` | Font size |
| `Ctrl/Cmd+?` | Keyboard help |
| `Shift+Enter` | Insert newline (sent to terminal) |
| `Escape` | Close panels & modals |
---
## Driving Codeman from an Agent — Programmatic Guide
For AI agents and automation that control Codeman without a browser: an agent that spins up worker sessions, a CI bot, or **Claude Code running _inside_ a Codeman session orchestrating other sessions**. Everything the UI does is HTTP + a CLI, so an agent can do it too.
### Detect that you're inside Codeman
When a CLI runs in a Codeman-managed session, these environment variables are set — read them instead of hardcoding anything:
| Variable | Meaning |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| `CODEMAN_MUX=1` | You're in a managed tmux session. **Never** `tmux kill-session` / `pkill claude` / `pkill tmux` — you'll kill yourself or a sibling. |
| `CODEMAN_API_URL` | Base URL of the API (e.g. `https://127.0.0.1:3000`). Use it for every call below. |
| `CODEMAN_SESSION_ID` | _Your own_ session id. Use it to avoid acting on yourself. |
| `CODEMAN_HOOK_SECRET_FILE` | Path to the hook secret (required on `/api/hook-event` while a managed tunnel is up). |
### Rules of the road (read before you POST)
1. **Single-line input only.** Programmatic input is sent as literal text **+ Enter** in one shot. Multi-line strings break the agent TUI (Ink) — send one line, or split into multiple calls.
2. **Make input idempotent.** Include a stable `clientId` and a monotonic per-session `seq` on `POST …/input`. The server de-duplicates, so a retry after a dropped connection can't double-deliver a prompt.
3. **Auth.** If `CODEMAN_PASSWORD` is set, send HTTP Basic auth (user `admin` or `CODEMAN_USERNAME`) or a `codeman_session` cookie. The default loopback install is passwordless. A missing `Origin` header is allowed, so plain `curl` works; cross-site browser origins are rejected (CSRF guard).
4. **Response envelope.** Most endpoints return `{ "success": true, "data": … }` (errors: `{ "success": false, "error", "errorCode" }`). A few legacy GETs return bare bodies — **handle both** (`body.data ?? body`).
5. **`/api/v1/*`** is a stable alias of `/api/*`.
### Recipes
```bash
API="${CODEMAN_API_URL:-http://127.0.0.1:3000}"
# (add -u admin:"$CODEMAN_PASSWORD" to each call if a password is set)
# 1. See what's running
curl -s "$API/api/sessions" | jq '.data // .'
# 2. Spin up a worker session (a "case" = named working dir)
curl -s -X POST "$API/api/quick-start" \
-H 'Content-Type: application/json' \
-d '{"caseName":"refactor-auth","mode":"claude","effort":"high"}' | jq
# 3. Send a prompt into a session (exactly-once: clientId + seq)
curl -s -X POST "$API/api/sessions/$SID/input" \
-H 'Content-Type: application/json' \
-d '{"input":"Run the test suite and summarize failures","useMux":true,"clientId":"agent-1","seq":1}'
# 4. Read the terminal back
curl -s "$API/api/sessions/$SID/output" | jq -r '.data // .'
# 5. Stream live events (session output, agent activity, status)
curl -sN "$API/api/events" # Server-Sent Events
# 6. Schedule recurring work (cron-style job)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{"name":"nightly-deps","agentType":"claude","workingDir":"/home/me/proj",
"promptMode":"inline_text","promptText":"Update dependencies and open a PR",
"inputMode":"typed","scheduleType":"daily","dailyTime":"03:00",
"enabled":true,"concurrencyPolicy":"warn_only"}' | jq
# 7. Inspect background sub-agents and their transcripts
curl -s "$API/api/subagents" | jq '.data // .'
curl -s "$API/api/subagents/$AID/transcript" | jq -r '.data // .'
# 8. Whole-system snapshot (sessions, settings, respawn, stats)
curl -s "$API/api/status" | jq
```
### Or use the bundled CLI
The same operations are available as commands (`codeman <cmd>`, aliases in parentheses) — handy from a shell tool inside a session:
```bash
codeman session start -d /path/to/repo # (s) start a session
codeman session list # list sessions
codeman session logs <id> # tail output
codeman task add "fix the failing test" # (t) queue a task
codeman ralph start --min-hours 8 # (r) launch the autonomous loop
codeman attach <path> # attach a Claude hook context
```
### Hooks (events flowing _back_ to Codeman)
Codeman registers Claude Code hooks that `POST /api/hook-event` (`permission_prompt`, `idle_prompt`, `stop`, `task_completed`, …) so the dashboard reacts in real time. This endpoint is auth-exempt on loopback but, under a managed tunnel, requires the `X-Codeman-Hook-Secret` header (read it from `$CODEMAN_HOOK_SECRET_FILE`). You normally don't call this by hand — Codeman wires it up — but it's how the autonomy layers "see" what the agent is doing.
> Full endpoint list and request/response shapes follow.
---
## API
REST over Fastify — **~140 handlers across 15 route modules**, plus an SSE stream and a WebSocket terminal channel. A representative subset:
REST over Fastify — **~160 handlers across 18 route modules**, plus an SSE stream and a WebSocket terminal channel. All responses use the `ApiResponse<T>` envelope (`{success, data}` / `{success, error, errorCode}`); `/api/v1/*` is a stable alias. A representative subset:
### Sessions
| Method | Endpoint | Description |
|--------|----------|-------------|
| `GET` | `/api/sessions` | List all |
| `POST` | `/api/quick-start` | Create case + start session |
| `DELETE` | `/api/sessions/:id` | Delete session |
| `POST` | `/api/sessions/:id/input` | Send input |
| Method | Endpoint | Description |
| -------- | -------------------------- | ---------------------------------------------------------------------------------- |
| `GET` | `/api/sessions` | List all |
| `POST` | `/api/quick-start` | Create case + start session (`{caseName?, mode?, effort?, envOverrides?}`) |
| `POST` | `/api/sessions/:id/input` | Send input (`{input, useMux?, clientId?, seq?}` — `clientId`+`seq` = exactly-once) |
| `GET` | `/api/sessions/:id/output` | Read terminal output |
| `DELETE` | `/api/sessions/:id` | Delete session |
### Respawn
| Method | Endpoint | Description |
|--------|----------|-------------|
| Method | Endpoint | Description |
| ------ | ---------------------------------- | -------------------------- |
| `POST` | `/api/sessions/:id/respawn/enable` | Enable with config + timer |
| `POST` | `/api/sessions/:id/respawn/stop` | Stop controller |
| `PUT` | `/api/sessions/:id/respawn/config` | Update config |
| `POST` | `/api/sessions/:id/respawn/stop` | Stop controller |
| `PUT` | `/api/sessions/:id/respawn/config` | Update config |
### Ralph / Todo
| Method | Endpoint | Description |
|--------|----------|-------------|
| `GET` | `/api/sessions/:id/ralph-state` | Get loop state + todos |
| `POST` | `/api/sessions/:id/ralph-config` | Configure tracking |
| Method | Endpoint | Description |
| ------ | -------------------------------- | ---------------------- |
| `GET` | `/api/sessions/:id/ralph-state` | Get loop state + todos |
| `POST` | `/api/sessions/:id/ralph-config` | Configure tracking |
### Orchestrator
| Method | Endpoint | Description |
|--------|----------|-------------|
| `POST` | `/api/orchestrator/start` | Start orchestration from a goal |
| `POST` | `/api/orchestrator/approve` | Approve the generated plan |
| `GET` | `/api/orchestrator/status` | Current phase + progress |
| `POST` | `/api/orchestrator/stop` | Stop and clean up |
| Method | Endpoint | Description |
| ------ | --------------------------- | ------------------------------- |
| `POST` | `/api/orchestrator/start` | Start orchestration from a goal |
| `POST` | `/api/orchestrator/approve` | Approve the generated plan |
| `GET` | `/api/orchestrator/status` | Current phase + progress |
| `POST` | `/api/orchestrator/stop` | Stop and clean up |
### Cron (scheduled jobs)
| Method | Endpoint | Description |
| ---------------- | ---------------------------- | ----------------------- |
| `GET` / `POST` | `/api/cron/jobs` | List / create cron jobs |
| `PUT` / `DELETE` | `/api/cron/jobs/:id` | Update / delete a job |
| `PUT` | `/api/cron/jobs/:id/enabled` | Enable / disable |
| `POST` | `/api/cron/jobs/:id/run` | Run now |
| `GET` | `/api/cron/jobs/:id/runs` | Run history |
### Subagents
| Method | Endpoint | Description |
|--------|----------|-------------|
| `GET` | `/api/subagents` | List all background agents |
| `GET` | `/api/subagents/:id` | Agent info and status |
| `GET` | `/api/subagents/:id/transcript` | Full activity transcript |
| `DELETE` | `/api/subagents/:id` | Kill agent process |
| Method | Endpoint | Description |
| -------- | ------------------------------- | -------------------------- |
| `GET` | `/api/subagents` | List all background agents |
| `GET` | `/api/subagents/:id` | Agent info and status |
| `GET` | `/api/subagents/:id/transcript` | Full activity transcript |
| `DELETE` | `/api/subagents/:id` | Kill agent process |
### System
| Method | Endpoint | Description |
|--------|----------|-------------|
| `GET` | `/api/events` | SSE stream |
| `GET` | `/api/status` | Full app state |
| `POST` | `/api/hook-event` | Hook callbacks |
| `GET` | `/api/system/update/check` | Check for a new release |
| `POST` | `/api/system/update` | Self-update (git-clone installs) |
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
| Method | Endpoint | Description |
| ------ | ------------------------------- | ---------------------------------------------- |
| `GET` | `/api/events` | SSE stream |
| `GET` | `/api/status` | Full app state |
| `POST` | `/api/hook-event` | Hook callbacks |
| `GET` | `/api/system/update/check` | Check for a new release |
| `POST` | `/api/system/update` | Self-update (git-clone installs) |
| `POST` | `/api/clipboard` | Push text to all connected browsers (`{text}`) |
| `GET` | `/api/sessions/:id/run-summary` | Timeline + stats |
---
@@ -579,7 +771,7 @@ flowchart TB
end
subgraph External["External"]
CLI["AI CLI<br/><small>Claude Code / OpenCode</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex</small>"]
BG["Background Agents<br/><small>(Task tool)</small>"]
end
end
@@ -622,14 +814,14 @@ See [CLAUDE.md](./CLAUDE.md) for full documentation.
The codebase went through a comprehensive 7-phase refactoring that eliminated god objects, centralized configuration, and established modular architecture:
| Phase | What changed | Impact |
|-------|-------------|--------|
| **Performance** | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
| **Route extraction** | `server.ts` split into 15 domain route modules + auth middleware + port interfaces | **−67%** server.ts LOC (6,736 → 2,254) |
| **Domain splitting** | `types.ts` → 16 domain files, `ralph-tracker` → 7 files, `respawn-controller` → 5 files, `session` → 6 files | No more god files |
| **Frontend modules** | `app.js` → 18 extracted modules across infra, domain & feature layers | app.js core down to **~3.4K LOC** |
| **Config consolidation** | ~70 scattered magic numbers → 10 domain-focused config files | Zero cross-file duplicates |
| **Test infrastructure** | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via `app.inject()` |
| Phase | What changed | Impact |
| ------------------------ | ------------------------------------------------------------------------------------------------------------ | ------------------------------------------ |
| **Performance** | Cached endpoints, SSE adaptive batching, buffer chunking | Sub-16ms terminal latency |
| **Route extraction** | `server.ts` split into 15 domain route modules + auth middleware + port interfaces | **−67%** server.ts LOC (6,736 → 2,254) |
| **Domain splitting** | `types.ts` → 16 domain files, `ralph-tracker` → 7 files, `respawn-controller` → 5 files, `session` → 6 files | No more god files |
| **Frontend modules** | `app.js` → 18 extracted modules across infra, domain & feature layers | app.js core down to **~3.4K LOC** |
| **Config consolidation** | ~70 scattered magic numbers → 10 domain-focused config files | Zero cross-file duplicates |
| **Test infrastructure** | Shared mock library, 12 route test files, consolidated MockSession | Testable route handlers via `app.inject()` |
Full details: [`docs/archive/code-structure-findings.md`](docs/archive/code-structure-findings.md)
+8 -7
View File
@@ -2,10 +2,10 @@
<img src="docs/images/codeman-title.svg" alt="Codeman" height="60">
</p>
<h2 align="center">为 AI 编程智能体而生的「控制平面」</h2>
<h2 align="center">AI 编程智能体的任务控制中心</h2>
<p align="center">
<em>智能体可视化 &bull; 零延迟输入 &bull; 自主编排器 &bull; 重生控制器 &bull; 移动优先 UI &bull; 安全加固</em>
<em>Claude Code &bull; OpenCode &bull; Codex —— 统一仪表盘 &bull; 任意设备</em>
</p>
<p align="center">
@@ -36,7 +36,7 @@ curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | b
该脚本会在缺失时自动安装 Node.js 和 tmux,把 Codeman 克隆到 `~/.codeman/app` 并完成构建。
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code) 或 [OpenCode](https://opencode.ai)(两个都装也可以)。安装完成后:
你至少需要安装一个 AI 编程 CLI —— [Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai) 或 [Codex](https://developers.openai.com/codex/cli)(任意组合均可)。安装完成后:
```bash
codeman web
@@ -105,7 +105,7 @@ launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.codeman.web.plist
wsl bash -c "curl -fsSL https://raw.githubusercontent.com/Ark0N/Codeman/master/install.sh | bash"
```
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code) 或 [OpenCode](https://opencode.ai))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
Codeman 依赖 tmux,因此 Windows 用户需要 [WSL](https://learn.microsoft.com/en-us/windows/wsl/install)。如果还没装 WSL:在管理员 PowerShell 中运行 `wsl --install`,重启,打开 Ubuntu,然后在 WSL 内安装你偏好的 AI 编程 CLI([Claude Code](https://docs.anthropic.com/en/docs/claude-code)、[OpenCode](https://opencode.ai) 或 [Codex](https://developers.openai.com/codex/cli))。安装完成后,即可从 Windows 浏览器访问 `http://localhost:3000`。
</details>
---
@@ -216,6 +216,7 @@ WATCHING → IDLE DETECTED → SEND UPDATE → /clear → /init → CONTINUE →
```
- **多层空闲检测** —— 完成消息、AI 驱动的空闲检查、输出静默、token 稳定性
- **用量限额自动恢复**(*可选,默认关闭*)—— 当 Claude 因订阅用量限额而停止("You've hit your limit · resets 3pm")时,Codeman 会解析重置时间,等到限额刷新(外加 2 分钟安全缓冲)后自动关闭限额对话框并发送 `continue`,让通宵任务平稳跨过 5 小时窗口而不是停摆到早晨。可识别 Claude Code 各版本的全部限额消息格式;若仍受限会自动重试;计划在 Codeman 重启后依然生效;暂停期间会阻止重生循环,避免 `/clear` 清掉等待中的对话。在会话 Respawn 标签页顶部按会话启用
- **熔断器** —— 当 Claude 卡住时防止重生抖动(CLOSED → HALF_OPEN → OPEN 状态,跟踪连续无进展与重复错误)
- **健康评分** —— 0–100 健康分,分项涵盖循环成功率、熔断器状态、迭代进展与卡死恢复
- **内置预设** —— `solo-work`(3s 空闲,60min)、`subagent-workflow`(45s,240min)、`team-lead`(90s,480min)、`ralph-todo`(8s,480min)、`overnight-autonomous`(10s,480min)
@@ -295,7 +296,7 @@ PTY 输出 → 16ms 服务端批处理 → DEC 2026 包裹 → SSE → 客户端
## 更多特性
- **自更新** —— systemd/launchd 管理下的 git-clone 安装可在 **App Settings → Updates** 中原地更新:它会检测最新发行版,自动暂存(stash)脏工作树,并在服务重启期间流式展示构建进度(npm 安装会被报告为不可更新)
- **双 CLI** —— 每个会话可选 **Claude Code** 或 **OpenCode**;环境变量前缀自动隔离(`CLAUDE_CODE_*` 与 `OPENCODE_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **多 CLI** —— 每个会话可选 **Claude Code**、**OpenCode** 或 **Codex**;环境变量前缀自动隔离(`CLAUDE_CODE_*`、`OPENCODE_*` 与 `CODEX_*`)。详见 [`docs/opencode-integration.md`](docs/opencode-integration.md)
- **Effort 与 Ultracode** —— 设置每会话的默认 effort(`low`–`max`),或启用 **ultracode**(动态多智能体工作流)。这些都只是软默认值 —— 会话中可随时用 `/effort` 切换。扩展思考预算也可配置
- **语音输入** —— 用 Deepgram Nova-3 口述提示(带 Web Speech API 回退):切换录音、自动静音停止、实时音量表(`Ctrl+Shift+V`)
- **图像输入** —— 直接把图片粘贴或拖放进会话
@@ -449,7 +450,7 @@ Codeman 用 `--dangerously-skip-permissions` 启动会话,因此 Web UI 在设
### 输入、文件与响应头
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **模式校验的输入** —— 每个 API 请求体都用 Zod v4 模式检查;一个 `CLAUDE_CODE_*` / `OPENCODE_*` / `CODEX_*` 环境变量前缀允许列表把控每个 CLI 能接收哪些设置
- **路径限定** —— 文件路由在边界检查前先 `realpath`(无 TOCTOU);`..`、绝对路径、以及解析到工作目录之外的符号链接都会被拒绝。上限:10 MB 文本预览 / 50 MB 原始与下载;`/api/download` 对敏感路径(`.env`、`*credentials*`、`~/.ssh/`、`.aws/credentials`)做黑名单。SVG/HTML 以 `octet-stream` + `nosniff` + attachment 提供,因此会被下载而非执行
- **安全响应头** —— `Content-Security-Policy`(`default-src 'self'`,每个例外都逐条列举)、`X-Content-Type-Options: nosniff`、`X-Frame-Options: SAMEORIGIN`、HTTPS 下的 HSTS,以及**仅**对 `localhost` / `127.0.0.1` / `::1` 反射的 CORS
@@ -581,7 +582,7 @@ flowchart TB
end
subgraph External["外部"]
CLI["AI CLI<br/><small>Claude Code / OpenCode</small>"]
CLI["AI CLI<br/><small>Claude Code / OpenCode / Codex</small>"]
BG["后台智能体<br/><small>(Task 工具)</small>"]
end
end
+104
View File
@@ -0,0 +1,104 @@
# SPEEDRUN.md — Fast-execution protocol for Claude
Read this when the goal is **throughput**: get correct, verified work done with
minimum ceremony. This does **not** relax correctness or the safety rules in
`CLAUDE.md` — those still win. It removes _waste_, not _rigor_.
> Precedence: `CLAUDE.md` > explicit user instructions > this file. If anything
> here conflicts with `CLAUDE.md`, `CLAUDE.md` wins.
---
## The mindset
- **Act, don't announce.** No "I'm going to now…" preamble. Do the thing, report
the result.
- **Cheapest proof that the change works.** Pick the smallest check that actually
demonstrates correctness — not the biggest.
- **Batch aggressively.** Independent reads, greps, and edits go in **one**
message with parallel tool calls. Never serialize work that has no dependency.
- **Momentum over perfection.** Land a correct increment, verify it, move on.
Don't gold-plate untouched code.
---
## Loop (repeat until done)
1. **Orient once** — one parallel burst of reads/greps to load the context you
need. Don't re-read files the harness says are already current.
2. **Change** — make the edit(s). Batch independent edits.
3. **Verify cheaply** — the smallest check that proves _this_ change (see below).
4. **Advance** — next item. Only re-verify what you touched.
5. **Stop** at: list empty, a hard blocker, or a decision that's genuinely the
user's to make.
---
## Verification ladder — climb only as high as the change needs
| Change kind | Cheapest sufficient check |
|-------------|---------------------------|
| Types / signatures / imports | `tsc --noEmit` (or `--watch` already running) |
| One module's logic | `npm test -- test/<file>.test.ts` (the **one** relevant file) |
| A named behavior | `npm test -- -t "pattern"` |
| Route/handler | `app.inject()` route test, or one `curl` against the running dev server |
| Frontend render | Playwright load + assert (`waitUntil: 'domcontentloaded'`, wait 3–4s) |
| Broad / pre-merge | `npm run test:ci` (the CI-equivalent sweep) |
**Hard rules (never skip, even in a rush):**
- ⚠️ **Never run bare `npm test`** — it pulls in browser/visual suites that hang
or fail locally. Always pass a file or `-t`, or use `test:ci`.
- ⚠️ **Never COM without verifying the change actually works** first (curl the
endpoint / Playwright the UI). "Compiles" ≠ "works".
- ⚠️ **Session safety** — check `$CODEMAN_MUX`; never `tmux kill-session` /
`pkill claude` in a managed session.
- ⚠️ **Single-line prompts** for any programmatic session input.
---
## Speed tactics that pay off here
- **Parallel exploration**: dispatch `Explore` subagents (or one parallel grep
burst) instead of serial file-by-file reading when scope is uncertain.
- **`tsc --noEmit --watch`** in the background — instant type feedback, no repeat
cold starts.
- **Target one test file** — `fileParallelism: false` means the suite is serial;
running one file is dramatically faster than the sweep.
- **`curl localhost:3000/api/...`** beats spinning up a browser for backend
checks. Reserve Playwright for actual UI rendering.
- **Trust the harness** — if it says a file you just edited is current, don't
re-Read it to "confirm". The Edit already succeeded or it would have errored.
---
## Anti-patterns (these masquerade as speed, but cost time)
- Running the full test suite to check a one-file change.
- Re-reading files you already have in context.
- Narrating a plan you're about to execute anyway.
- Serial tool calls that have no dependency between them.
- Claiming "done / fixed / passing" **before** running the check that proves it.
- Deploying (COM) on green typecheck alone, without exercising the real flow.
---
## Stop-conditions (don't rush past these)
Stop and surface, don't guess, when you hit:
- A **destructive / hard-to-reverse** action (delete, overwrite, force-push).
- An **outward-facing** action (publishing, sending, deploying) not already
authorized.
- A **genuine product decision** the code can't answer.
- A **failing verification you can't explain** — debug it (see
`superpowers:systematic-debugging`), don't paper over it.
---
## Definition of done
A task is done when **all** hold:
- The change is made.
- The cheapest sufficient check **ran** and **passed** — evidence, not assertion.
- No new type errors / lint errors introduced (`tsc --noEmit`, `npm run lint`).
- You state plainly what was done and what proved it. If a step was skipped or a
test failed, say so — don't hedge, don't overclaim.
+65
View File
@@ -0,0 +1,65 @@
# Codeman agent base image (built locally by scripts/build-agent-image.mjs).
#
# Contains the agent toolchain (node + the CLIs + git/tmux/ripgrep) but NO
# secrets: credentials are delivered at RUNTIME via bind mounts (~/.claude etc.)
# or name-only `docker exec --env`, never baked in, so `docker save` exports stay
# secret-free. tmux is a HARD prerequisite (the in-container tmux is what makes a
# reconnect durable), so it is installed here and probed before launch.
#
# HOME is made writable by an ARBITRARY host uid via the OpenShift "gid 0,
# group-writable" convention: on Linux we run `--user <hostUid>:0`, so the agent
# uid is the host uid (workspace files stay host-owned) while gid 0 keeps $HOME
# writable even though the uid is not the baked 1000.
FROM node:22-bookworm-slim
# Base toolchain. `curl` is needed for the hook callbacks (`curl -sk $CODEMAN_API_URL`),
# `procps` for `ps`, `tmux` for the durable in-container session.
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
tmux \
ripgrep \
curl \
ca-certificates \
less \
procps \
openssh-client \
&& rm -rf /var/lib/apt/lists/*
# The agent CLIs (all four backends Codeman supports). Pinning is left to the
# rebuild cadence (see docs/docker-cases-plan.md, user-decision 2).
RUN npm install -g \
@anthropic-ai/claude-code \
@openai/codex \
@google/gemini-cli \
opencode-ai \
&& npm cache clean --force
# `agent` user (gid 0) with an arbitrary-uid-writable HOME. The uid is
# auto-assigned (node:22-slim already occupies uid 1000 with its `node` user); at
# runtime Codeman overrides with `--user <hostUid>:0` on Linux, so the baked uid
# only matters for a hand-run / Docker Desktop container. gid 0 + group-writable
# HOME (OpenShift arbitrary-uid convention) keeps $HOME writable for any uid.
# UTF-8 locale so tmux/Ink render Unicode box-drawing instead of VT100 ACS `q`
# glyphs (C.UTF-8 is built into glibc; no locales package needed). Codeman also
# sets these at run time so containers built before this line still get UTF-8.
ENV LANG=C.UTF-8 LC_ALL=C.UTF-8
ENV HOME=/home/agent
# `.claude` (+ `.claude/projects` mount point) and `.codex` (+ `.codex/sessions`) are
# pre-created gid-0 group-writable so the container owns its OWN credential config
# dirs: tokens/settings/config are seeded in as writable copies and each CLI's runtime
# state (backups, tasks, refreshed tokens) stays container-local, while ONLY the shared
# transcript/rollout dirs (`.claude/projects`, `.codex/sessions`) are bind-mounted from
# the host. (gemini/gcloud/opencode are whole seed-copies and need no pre-created dir.)
RUN useradd -g 0 -m -d /home/agent -s /bin/bash agent \
&& mkdir -p /home/agent/.npm /home/agent/.cache /home/agent/.config /home/agent/.codeman \
/home/agent/.claude/projects /home/agent/.codex/sessions \
&& chgrp -R 0 /home/agent \
&& chmod -R g=u /home/agent
USER agent
WORKDIR /home/agent
# Codeman overrides the command with `sleep infinity` at create time; this is the
# fallback so a hand-run container also idles rather than exiting.
CMD ["sleep", "infinity"]
+588
View File
@@ -0,0 +1,588 @@
# Claude Code Build Brief: Add Scheduling to Codeman
## 0. Purpose of This Brief
You are Claude Code working inside the Codeman repository.
Your task is to add a **small, reliable scheduling layer** to Codeman while preserving Codeman's existing architecture and session-management behavior.
This is not a greenfield rewrite. This is not a full product rebuild. This is a focused extension.
The target user wants Codeman-like tmux/web/session management, but with first-class scheduled jobs for Claude, Codex, OpenCode, Terminal, or any other configurable coding-agent harness.
---
## 1. Non-Negotiable Goal
Add scheduling to Codeman so a user can define a scheduled coding-agent job that:
1. Has a name.
2. Uses an existing Codeman-supported agent/session type where possible.
3. Has a working directory.
4. Has a prompt or prompt file.
5. Has a schedule.
6. Can be enabled or disabled.
7. Can be manually run now.
8. When due, creates a Codeman/tmux session.
9. Sends the configured prompt into that session.
10. Records last run, next run, status, and run history.
The first working version should prioritize **scheduling correctness and reuse of Codeman's existing tmux/session system** over UI polish.
---
## 2. Core Architectural Rule
Do **not** rebuild Codeman's session layer.
Reuse existing Codeman functionality for:
- Creating sessions.
- Naming sessions.
- Launching Claude/Codex/OpenCode/Terminal sessions.
- Sending input into sessions.
- Displaying sessions in the web UI.
- Killing sessions.
- Tracking session status if already supported.
If an internal API/service/function already exists, reuse it.
If no reusable function exists, create a thin wrapper around the existing implementation rather than duplicating logic.
---
## 3. Product Boundary
This build is **Codeman + Scheduler**.
It is not yet:
- A full quota engine.
- A full lock manager.
- A replacement for Codeman's terminal UI.
- A new FastAPI application.
- A multi-tenant SaaS platform.
- A complex cron-management product.
- A full agent autonomy framework.
Keep the build small and shippable.
---
## 4. Required Working Scope for v0.1
Implement the following minimum features.
### 4.1 Scheduled Jobs List
Create a UI page showing all scheduled jobs.
Each row/card should show:
- Job name.
- Agent/session type.
- Working directory.
- Schedule type.
- Enabled/disabled state.
- Last run time.
- Next run time.
- Last run status.
- Actions:
- Run Now.
- Enable/Disable.
- Edit.
- Delete.
### 4.2 Create/Edit Scheduled Job
Create a form for scheduled jobs with these fields:
- `name`
- `agent_type`
- Reuse Codeman's existing session/agent types where possible.
- Include at least Terminal/custom command if supported.
- `working_directory`
- `launch_command` if needed by Codeman's model.
- `prompt_mode`
- `inline_text`
- `prompt_file_path`
- `prompt_text`
- `prompt_file_path`
- `input_mode`
- `paste`
- `typed`
- `schedule_type`
- `once`
- `interval_minutes`
- `daily_time`
- `weekly_time`
- `run_at` for one-time jobs.
- `interval_minutes` for interval jobs.
- `daily_time` for daily jobs.
- `weekly_days` and `weekly_time` for weekly jobs.
- `enabled`
- `notes` optional.
Do not build a complex visual cron editor in v0.1.
### 4.3 Run Now
Every scheduled job must support a `Run Now` action.
Run Now should:
1. Create a new session through Codeman's existing session creation logic.
2. Send the configured prompt into the session using Codeman's existing input mechanism.
3. Create a run-history record.
4. Update last-run fields.
5. Redirect or link the user to the created Codeman session.
### 4.4 Background Scheduler Loop
Add a small background scheduler loop that runs inside the Codeman backend process.
The loop should:
1. Wake every 15-60 seconds.
2. Load enabled schedules.
3. Find schedules where `next_run_at <= now`.
4. Create a scheduled run.
5. Launch the session using existing Codeman session logic.
6. Send the prompt.
7. Record run history.
8. Compute the next run time.
9. Avoid duplicate launches if the loop overlaps or restarts.
Keep this simple and robust.
### 4.5 Run History
Every scheduled execution should create a run-history record.
Track:
- `id`
- `scheduled_job_id`
- `session_id` or Codeman session reference.
- `session_name` if applicable.
- `started_at`
- `finished_at` optional.
- `status`
- `created`
- `session_started`
- `prompt_sent`
- `failed`
- `error_message` optional.
- `trigger_type`
- `scheduled`
- `manual_run_now`
- `created_session_url` or route reference if easy.
---
## 5. Scheduling Rules
### 5.1 Once
Run at a specific date/time.
After successful launch:
- Set `enabled = false`, or mark as completed.
### 5.2 Interval
Run every N minutes.
Example:
- Every 60 minutes.
- Every 240 minutes.
After launch:
- `next_run_at = now + interval_minutes`.
### 5.3 Daily
Run every day at HH:MM.
After launch:
- Compute the next occurrence of HH:MM after now.
### 5.4 Weekly
Run on selected weekdays at HH:MM.
After launch:
- Compute the next selected weekday/time after now.
### 5.5 Timezone
Use the server's local timezone for v0.1 unless Codeman already has timezone handling.
Add a visible note in the UI:
> Times use the server's local timezone.
Do not overbuild timezone support in v0.1.
---
## 6. Data Storage Decision
First inspect Codeman's existing persistence model.
If Codeman already has a database or persistence layer:
- Reuse it.
- Add scheduled job and scheduled run models/tables/records using the existing pattern.
If Codeman uses files or JSON state:
- Use the same style for v0.1.
- Prefer simple persistence over introducing a heavy new dependency.
If there is no appropriate persistence layer:
- Add SQLite only if it fits the codebase cleanly.
- Otherwise use a JSON file store for the first version.
Do not introduce Postgres, Redis, Celery, or a separate scheduler service.
---
## 7. Concurrency and Duplicate-Run Guard
Implement a basic duplicate-run guard.
A schedule should not launch twice for the same due time.
Minimum acceptable approach:
- Before launching, create/update a run record with a `created` or `launching` state.
- Use a schedule-level `last_triggered_at` or `last_due_key` to avoid double launching.
- If launch fails, record failure clearly.
Do not build distributed locks. Codeman is expected to be local/single-instance for v0.1.
---
## 8. Multi-Session Warning
When the user clicks `Run Now`, show a warning if there are already active sessions for the same agent type.
Minimum behavior:
- If active sessions exist, show a confirmation warning.
- User can continue anyway.
For scheduled automatic runs:
- Add a setting on the scheduled job:
- `warn_only`
- `skip_if_same_agent_running`
Default:
- `warn_only` for manual runs.
- `skip_if_same_agent_running = false` for automatic runs unless easy to implement.
Do not build a complete quota engine in v0.1.
---
## 9. Prompt Sending Rules
The scheduler must support sending the configured prompt into the created session.
Prompt source:
1. Inline prompt text.
2. Prompt file path.
Input mode:
1. Paste mode.
2. Typed mode.
If only one input mode is easy with Codeman's current internals, implement that first and structure the code so the other can be added later.
Important:
- Do not send prompts to a session if session creation failed.
- Record prompt-send success/failure in run history.
- Save enough metadata to understand what prompt was used.
---
## 10. UI Bifurcation
Keep UI changes cleanly separated.
Add scheduler UI under a clear navigation item:
- `Scheduled Jobs`
Do not clutter the existing session dashboard.
The existing session dashboard may show sessions created by scheduled jobs, but the scheduling controls should live in their own section.
Recommended pages/routes:
- `/schedules`
- `/schedules/new`
- `/schedules/:id`
- `/schedules/:id/edit`
- `/schedules/:id/run-now`
- `/schedules/:id/enable`
- `/schedules/:id/disable`
- `/schedules/:id/delete`
Use Codeman's existing frontend conventions and routing style.
---
## 11. Backend Bifurcation
Keep scheduler code separate from existing session code.
Recommended logical modules, adapted to Codeman's actual structure:
- `scheduler/model` or equivalent.
- `scheduler/store` or equivalent.
- `scheduler/service` for schedule calculations and launch logic.
- `scheduler/loop` for the background due-job checker.
- `scheduler/routes` for API/UI endpoints.
- `scheduler/time` for next-run calculations.
Do not mix scheduling logic directly into terminal rendering, xterm handling, or low-level tmux code.
The scheduler service should call session services; it should not own tmux directly unless Codeman has no session abstraction.
---
## 12. Required Discovery Phase Before Coding
Before implementing, inspect the Codeman repo and produce a short architecture note in the terminal or in a file called:
`docs/cron-discovery.md`
This note must identify:
1. Where session creation happens.
2. Where agent/session types are defined.
3. Where input is sent into a session.
4. Where active sessions are listed.
5. Where session kill/delete is handled.
6. How session state is stored.
7. Whether there is existing persistence.
8. Where backend routes live.
9. Where frontend pages/components live.
10. The smallest integration points for scheduling.
Do not start coding until this discovery is complete.
---
## 13. Implementation Phases
### Phase 1: Discovery
Deliverable:
- `docs/cron-discovery.md`
Must answer the 10 discovery questions above.
### Phase 2: Data Model / Persistence
Deliverable:
- Scheduled job persistence.
- Scheduled run history persistence.
- Basic create/read/update/delete operations.
### Phase 3: Scheduler Calculation Logic
Deliverable:
- Functions to compute `next_run_at` for:
- once
- interval
- daily
- weekly
Add tests if the repo has an existing test setup.
### Phase 4: Manual Run Now
Deliverable:
- Create scheduled job.
- Click Run Now.
- Codeman session is created.
- Prompt is sent.
- Run history is recorded.
- UI links to the session.
This is the most important milestone.
### Phase 5: Background Scheduler Loop
Deliverable:
- Enabled schedules launch automatically when due.
- Run history is recorded.
- `last_run_at` and `next_run_at` update.
- Duplicate launch guard exists.
### Phase 6: UI Polish Only After Functionality
Deliverable:
- Scheduled jobs list is readable.
- Create/edit form is usable.
- Status labels are clear.
- Errors are visible.
Do not polish before Phase 4 works.
---
## 14. Acceptance Criteria
The build is acceptable when all these pass.
### Manual Run
1. Create a schedule/job with inline prompt.
2. Click Run Now.
3. A new Codeman/tmux session starts.
4. Prompt is sent into that session.
5. The created session is visible in Codeman's normal session UI.
6. Run history shows success or failure.
### One-Time Schedule
1. Create a one-time schedule 2 minutes in the future.
2. Wait for it to become due.
3. Scheduler launches a session.
4. Prompt is sent.
5. Schedule does not repeatedly launch forever.
### Interval Schedule
1. Create interval schedule every 2 minutes.
2. It launches once when due.
3. It computes the next due time.
4. It does not launch duplicates for the same due time.
### Daily Schedule
1. Create daily schedule at a time a few minutes ahead.
2. It launches when due.
3. Next run becomes tomorrow at the same time.
### Disable Schedule
1. Disable a schedule.
2. It does not launch even when due.
### Error Handling
1. Invalid working directory produces visible error.
2. Invalid prompt file produces visible error.
3. Failed session launch creates failed run-history entry.
---
## 15. Explicitly Out of Scope for v0.1
Do not implement these unless all required scope is already working:
- Full quota engine.
- Advanced lock manager.
- Post-run git inspection reports.
- Complex recurring calendar UI.
- User accounts / RBAC.
- External distributed workers.
- Redis.
- Postgres.
- Celery.
- Kubernetes.
- A separate Python service.
- Full visual cron editor.
- AI-generated follow-up prompts.
- Automatic continuation after idle.
- Any attempt to bypass agent quotas or platform limits.
---
## 16. Quality Rules
Follow these rules while coding:
1. Reuse existing Codeman services and conventions.
2. Keep scheduler code isolated.
3. Prefer boring, readable code over clever abstractions.
4. Add error messages that a human can understand.
5. Do not break existing Codeman sessions.
6. Do not rename existing core concepts unnecessarily.
7. Do not introduce large dependencies without strong reason.
8. Keep v0.1 local-first and single-instance.
9. Commit in logical chunks if git is available.
10. After coding, provide a final implementation summary.
---
## 17. Final Response Required from Claude Code
At the end, report:
1. Files changed.
2. New routes/pages added.
3. New data structures added.
4. How the scheduler loop works.
5. How to run the app.
6. How to test manual Run Now.
7. How to test scheduled execution.
8. Known limitations.
9. Suggested v0.2 improvements.
---
## 18. v0.2 Ideas, Not for Current Build
Keep these in mind but do not build unless v0.1 is complete:
- Quota-aware scheduling.
- Manual takeover locks.
- Post-idle inspection.
- Git diff reports.
- Schedule groups.
- Prompt templates.
- Agent-specific concurrency rules.
- Better timezone support.
- Audit events.
- More advanced cron expressions.
---
## 19. Final Reminder
The goal is to add **scheduling** to Codeman quickly and cleanly.
Do not drift into building a new platform.
The highest-priority path is:
1. Discover existing Codeman integration points.
2. Add scheduled job persistence.
3. Add Run Now.
4. Add background due-job loop.
5. Add minimal UI.
6. Verify that scheduled jobs create real Codeman/tmux sessions and send prompts.
+142
View File
@@ -0,0 +1,142 @@
# CRON_DISCOVERY.md
Phase 1 deliverable for the "Add Scheduling to Codeman" build brief.
This documents the existing Codeman architecture and the smallest integration
points for a cron. **No session/tmux logic will be rebuilt** —
the new code is purely a trigger + persistence + history layer on top of the
existing primitives.
Stack: `aicodeman` v1.2.1 — Fastify 5 backend, `node-pty` + tmux sessions,
vanilla-JS SPA frontend served as static assets, JSON file state store, zod
validation, ports-based dependency injection.
---
## 0. Critical finding: an existing `ScheduledRun` is NOT a cron
Codeman already has a `ScheduledRun` concept (`/api/scheduled`,
`src/web/ports/infra-port.ts:14-26`, `src/web/server.ts:1480-1605`). It is a
**run-now, duration-bounded autonomous loop**: given `{prompt, workingDir,
durationMinutes}` it immediately spawns/kills throwaway sessions in a loop until
the duration elapses. It has **no** time-based triggering, recurrence
(once/interval/daily/weekly), enable/disable, next-run calculation, run history,
or persistence across restarts.
Therefore the brief's core (the calendar/cron trigger layer) does **not** exist
and must be built. The execution primitives it sits on top of **do** exist and
will be reused. To honor brief §16 ("do not rename existing core concepts"), the
new feature is named **`CronJob`** (with **`CronJobRun`** history
records), kept distinct from the existing `ScheduledRun`.
---
## 1. Where session creation happens
- Canonical create flow: `POST /api/sessions`,
`src/web/routes/session-routes.ts:262-438`.
- `new Session({ workingDir, mode, ... })` (`src/session.ts:421-570`)
- `ctx.addSession(session)` → `ctx.setupSessionListeners(session)` →
`ctx.persistSessionState(session)` (all via `SessionPort`).
- `SessionPort` interface: `src/web/ports/session-port.ts:8-16`.
- **Integration point:** the cron service will mirror this exact sequence
(create → addSession → setupSessionListeners → start) via `SessionPort`,
not reimplement it.
## 2. Where agent/session types are defined
- `type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini'`
(`src/types/session.ts:43-44`). `shell` covers the brief's "Terminal/custom".
- CLI availability resolvers in `src/utils/{claude,codex,gemini,opencode}-cli-resolver.ts`.
- **Integration point:** the job's `agentType` reuses `SessionMode` verbatim.
## 3. Where input is sent into a session
- Raw / paste: `session.write(data)` (`src/session.ts:2243-2247`) — direct PTY write.
- Typed (recommended): `session.writeViaMux(data)` (`src/session.ts:2301-2311`)
— tmux `send-keys`, falls back to PTY. Submit requires trailing `\r`.
- **Integration point:** prompt delivery uses `writeViaMux` (typed) by default,
`write` (paste) as the alternate `input_mode`.
## 4. Where active sessions are listed
- `ctx.sessions: ReadonlyMap<string, Session>` (`SessionPort`).
- Filters: `Array.from(ctx.sessions.values()).filter(s => s.mode === X)` and
`.isBusy()` / `.isIdle()` (`src/session-manager.ts:220-247`).
- **Integration point:** the §8 multi-session warning queries this map.
## 5. Where session kill/delete is handled
- `ctx.cleanupSession(sessionId, killMux?, reason?)`
(`SessionPort`; impl `src/web/server.ts:997-1152`). Underlying
`session.stop(killMux)` at `src/session.ts:2498-2585`.
- The cron does **not** kill sessions it launches (the brief wants them
visible in the normal session UI); cleanup stays user-driven.
_Superseded post-review:_ recurring jobs now default to
`autoClosePreviousSession: true` — the previous run's still-open session is
closed via `cleanupSession` when the next run fires (see
`docs/cron-guide.md` §8); opt out per job for fully user-driven cleanup.
## 6. How session state is stored / 7. Existing persistence
- JSON file store: `~/.codeman/state.json` (+ `state-inner.json` for Ralph).
`StateStore` class `src/state-store.ts:71`; `AppState` interface
`src/types/app-state.ts:99-114`.
- Pattern: declare a field on `AppState`, add typed get/set methods on
`StateStore` that mutate in-memory state and call the debounced `save()`
(500ms debounce, atomic temp-file+rename, `.bak` backup, circuit breaker).
- **Integration point:** add `cronJobs?: Record<string, CronJob>` and
`cronJobRuns?: Record<string, CronJobRun>` to `AppState`, with
matching `StateStore` accessors. No new DB (brief §6 forbids Postgres/Redis).
## 8. Where backend routes live
- Route modules: `src/web/routes/*.ts`; barrel `src/web/routes/index.ts`;
registered in `WebServer.setupRoutes()` `src/web/server.ts:858-876` with a
single `ctx` object from `createRouteContext()` (`src/web/server.ts:553-613`)
that satisfies all port interfaces.
- Validation: zod schemas in `src/web/schemas.ts`, applied via
`parseBody(Schema, req.body)` (`src/web/route-helpers.ts:101-111`).
- Errors: `createErrorResponse(ApiErrorCode.X, msg)` / `ApiResponse`
(`src/types/api.ts`), auto-mapped to HTTP status by a `preSerialization` hook
(`src/web/server.ts:644-659`).
- SSE: `ctx.broadcast(SseEvent.X, data)` (`EventPort`,
`src/web/sse-events.ts`); frontend mirror in `src/web/public/constants.js`.
- **Integration point:** new `cron-routes.ts` registered alongside the
others; new zod schema; new `SseEvent` constants for job list/run changes.
## 9. Where frontend pages/components live
- Vanilla-JS SPA: single `src/web/public/index.html` + feature mixin files
(`Object.assign(CodemanApp.prototype, {...})`). API via `api-client.js`
(`_apiJson/_apiPost/_apiDelete`). Build = esbuild minify + content-hash, no
bundler (`scripts/build.mjs`).
- UI is panels/modals toggled by JS classes; forms use `.form-row` / `.modal`
conventions (`styles.css`). SSE handler map in `app.js`.
- **Integration point:** add a new `cron-ui.js` mixin + a panel/modal in
`index.html` + nav entry, following the orchestrator/respawn panel pattern.
## 10. Background-loop pattern (for the due-checker)
- Established pattern: `this.cleanup.setInterval(fn, intervalMs, {description})`
in `WebServer.start()` (`src/web/server.ts:~1942-1966`), auto-disposed in
`WebServer.stop()` via `this.cleanup.dispose()` (`src/web/server.ts:2336`).
RalphLoop (`src/ralph-loop.ts:268-286`) shows the self-rescheduling guard idiom.
- **Integration point:** register a 30s cron tick via `cleanup.setInterval`;
no manual shutdown wiring needed.
---
## Smallest integration points (summary)
| New piece | Reuses | Location |
| --- | --- | --- |
| `CronJob` / `CronJobRun` types | — (new) | `src/types/cron.ts` |
| Persistence | `StateStore` / `AppState` | `src/types/app-state.ts`, `src/state-store.ts` |
| Next-run time math | — (new, pure, unit-tested) | `src/cron/cron-time.ts` |
| Launch + send prompt | `SessionPort` (`addSession`/listeners/`writeViaMux`) | `src/cron/cron-service.ts` |
| Background due loop | `cleanup.setInterval` pattern | `src/cron/cron-loop.ts` |
| Routes + schema | route/ports/zod/SSE patterns | `src/web/routes/cron-routes.ts`, `src/web/schemas.ts`, `src/web/sse-events.ts` |
| UI | panel/modal/mixin conventions | `src/web/public/cron-ui.js`, `index.html` |
Nothing in the session, tmux, persistence, routing, or SSE subsystems is
rewritten — the cron is additive and calls existing services.
+426
View File
@@ -0,0 +1,426 @@
# Cron Jobs — User & Operator Guide
Codeman's **Cron** feature lets you save named, recurring jobs that automatically
spin up a Claude (or shell / OpenCode / Codex / Gemini) session on a schedule and
feed it a prompt. Think "cron for agent sessions": _"every weekday at 3am, open a
Claude session in `~/proj` and tell it to update dependencies and open a PR."_
- **UI**: the **⏰ Cron** button in the header → the Cron Jobs modal (`#cronModal`).
- **API**: `/api/cron/jobs*` and `/api/cron/runs`.
- **Code**: `src/cron/cron-service.ts`, `src/cron/cron-time.ts`, `src/cron/cron-input.ts`,
types in `src/types/cron.ts`, routes in `src/web/routes/cron-routes.ts`,
frontend in `src/web/public/cron-ui.js`.
> **Not to be confused with `ScheduledRun` (`/api/scheduled`).** That older,
> deliberately-separate concept is a _run-now, duration-bounded autonomous loop_
> (`{prompt, workingDir, durationMinutes}` → spawn/kill throwaway sessions until
> the duration elapses). It has no recurrence, no saved jobs, and no next-run
> calculation. The two systems never interact. This guide is only about **Cron
> jobs** (`Cron*`). See `docs/cron-discovery.md` §0.
---
## 1. Quick start
### In the browser
1. Click **⏰ Cron** in the header.
2. Click **+ New Job**.
3. Fill in a **name**, pick an **agent type** and **working directory**, choose a
**prompt** (inline text or a file path), pick a **schedule**, and leave
**Enabled** on.
4. **Save**. The job appears in the list with its computed **next run**.
5. Use **Run Now** to fire it immediately without waiting for the schedule.
### With curl
```bash
API=http://localhost:3000
# Create a daily job (03:00 server-local time)
curl -s -X POST "$API/api/cron/jobs" \
-H 'Content-Type: application/json' \
-d '{
"name": "nightly-deps",
"agentType": "claude",
"workingDir": "/home/me/proj",
"promptMode": "inline_text",
"promptText": "Update dependencies and open a PR",
"inputMode": "typed",
"scheduleType": "daily",
"dailyTime": "03:00",
"enabled": true,
"concurrencyPolicy": "warn_only"
}' | jq
# List jobs
curl -s "$API/api/cron/jobs" | jq
# Run one immediately
curl -s -X POST "$API/api/cron/jobs/<jobId>/run" | jq
# See a job's run history
curl -s "$API/api/cron/jobs/<jobId>/runs" | jq
```
---
## 2. Concepts
| Term | Meaning |
| -------------------------- | ------------------------------------------------------------------------------------------------ |
| **Cron job** (`CronJob`) | A saved, named definition: what agent to launch, where, with what prompt, on what schedule. |
| **Run** (`CronJobRun`) | One execution of a job — a history record with a status and a link to the session it created. |
| **Schedule type** | How fire times are computed: `once`, `interval`, `daily`, or `weekly`. |
| **Next run** (`nextRunAt`) | Server-computed epoch-ms of the next fire. `null` when the job is disabled or has no future run. |
| **Due tick** | A background loop (every 30s) that launches any enabled job whose `nextRunAt` has passed. |
A job is essentially a **trigger + persistence + history layer on top of the
existing session primitives**. When a job fires, the cron service does exactly
what the "quick start" route does — `new Session(...)` → `addSession` →
`setupSessionListeners` → `startInteractive()`/`startShell()` → deliver the
prompt. It does **not** reimplement any tmux/PTY logic.
---
## 3. The job form — every field
These map 1:1 to `CronJobSchema` (`src/web/schemas.ts`) and the `CronJob` type
(`src/types/cron.ts`).
| Field | Required | Values / limits | Notes |
| -------------------------- | ----------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `name` | ✅ | 1–200 chars | Display name; also used as the created session's name. |
| `agentType` | ✅ | `claude` \| `shell` \| `opencode` \| `codex` \| `gemini` | Reuses Codeman's `SessionMode`. `shell` = a plain terminal. |
| `workingDir` | ✅ | valid path (allowlist-validated) | Validated at **create/update** (must exist, be a directory, and not resolve into a blocked tree — `/etc`, `/root`, `/proc`, `/sys`, `/dev`, or `/` itself) and again **at fire time**. |
| `launchCommand` | — | ≤ 2000 chars, single line | `shell` mode only: sent as the **first input line** once the shell is up, before the prompt. Ignored for other agent types. |
| `promptMode` | ✅ | `inline_text` \| `prompt_file_path` | See §5. |
| `promptText` | conditional | ≤ 100000 chars, **single line** | Required when `promptMode = inline_text`. Newlines are rejected (see §6). |
| `promptFilePath` | conditional | valid path | Required when `promptMode = prompt_file_path`. Confined to `workingDir` (see §5). |
| `inputMode` | ✅ | `paste` \| `typed` | How the prompt is delivered. See §6. |
| `scheduleType` | ✅ | `once` \| `interval` \| `daily` \| `weekly` | See §4. |
| `runAt` | conditional | epoch-ms (positive int) | Required for `once`. |
| `intervalMinutes` | conditional | 1–525600 (≤ 1 year) | Required for `interval`. |
| `dailyTime` | conditional | `HH:MM` (24h) | Required for `daily`. Server-local time. |
| `weeklyDays` | conditional | array of 1–7 ints, each 0–6 (0 = Sunday) | Required for `weekly`. |
| `weeklyTime` | conditional | `HH:MM` (24h) | Required for `weekly`. Server-local time. |
| `enabled` | ✅ | boolean | Disabled jobs never auto-fire (but **Run Now** still works). |
| `notes` | — | ≤ 2000 chars | Free-form. |
| `concurrencyPolicy` | ✅ | `warn_only` \| `skip_if_same_agent_running` | Applies to **automatic** runs only. See §7. |
| `autoClosePreviousSession` | — | boolean (default **true**) | Recurring schedules only (ignored for `once`): when the next run fires, the still-open session created by this job's **previous** run is closed first via the normal cleanup path. See §8. |
**Cross-field validation** (`refineCronJob` in `schemas.ts`): the conditional
fields above are enforced by a Zod `superRefine` on create. A missing dependent
field (e.g. `scheduleType: "once"` with no `runAt`) is rejected with
`INVALID_INPUT` and a field-specific message.
> ⚠️ **Update caveat.** `PUT /api/cron/jobs/:id` uses a `.partial()` schema that
> does **not** re-run the cross-field `superRefine`. To keep partial edits safe,
> `updateJob()` re-validates the **merged** job against the full `CronJobSchema`
> and throws `400` if the result is inconsistent (e.g. switching to `once`
> without a `runAt`). So the store is never left with a half-valid job.
---
## 4. Schedule types
Next-run math lives in `src/cron/cron-time.ts` (pure, unit-tested in
`test/cron-time.test.ts`). **All wall-clock times use the server's local
timezone** (v0.1 decision).
### `once`
- Fires a single time at the absolute `runAt` epoch-ms.
- A **missed** one-time job (server was down at `runAt`) **still fires once** on
the next tick — `computeNextRunAt` returns `runAt` even if it's in the past,
until the job has fired.
- After firing, the job **self-disables**: `completedOnce = true`, `enabled =
false`, `nextRunAt = null`.
### `interval`
- Fires every `intervalMinutes`, computed as `fireTime + intervalMinutes`.
- ⚠️ **Drift**: the next run re-anchors to the actual fire time, not to an ideal
cadence — a slow tick or restart shifts subsequent runs slightly later. This is
an accepted limitation.
### `daily`
- Fires at `dailyTime` (`HH:MM`) every day, server-local.
- If today's time has already passed, the next run is tomorrow at that time.
### `weekly`
- Fires at `weeklyTime` on each weekday in `weeklyDays` (0 = Sunday … 6 =
Saturday), server-local.
- The next run is the soonest upcoming matching weekday/time within the next 7
days.
---
## 5. Prompt source (`promptMode`)
### `inline_text`
The prompt is the literal `promptText`. Simplest option.
### `prompt_file_path`
The prompt is read from a file at fire time. **This path is security-hardened**
because a job config is attacker-controllable and the file's contents are
injected into an agent session (an exfiltration sink over SSE/terminal).
`resolveSafePromptPath()` enforces, in order:
1. **`realpath` resolution** — symlinks are resolved to their true target, for
the prompt file **and for `workingDir` itself**.
2. **`workingDir` is not a trust boundary** — because it is user-supplied, the
resolved `workingDir` is itself rejected if it is `/` or resolves into a
blocked tree (`/etc`, `/root`, operator extras) or a pseudo-filesystem
(`/proc`, `/sys`, `/dev`). This closes the `workingDir: '/proc'` +
`promptFilePath: '/proc/self/environ'` env-exfil trick. The same rule is
enforced earlier, at job create/update.
3. **Blocklist** (defense-in-depth) — sensitive trees (`/etc`, `/root`,
`/proc`, `/sys`, `/dev`, known secret locations) are rejected for the
resolved prompt file.
4. **Allowlist (primary gate)** — the resolved path **must live inside the job's
(resolved) `workingDir`** (`validateSessionFilePath`). A symlink escaping the
workspace fails here.
5. **Regular-file check** — directories, FIFOs, and `/dev/*` character devices
are rejected (they would hang or OOM an unbounded read).
6. **Size cap** — files larger than **1 MiB** (`MAX_PROMPT_FILE_BYTES`) are
rejected.
7. **Single-line check** — after trailing newlines are stripped, the file
content must be a single line (see §6).
If any check fails, the run is recorded as **`failed`** with the reason; no
session is created.
---
## 6. Prompt delivery (`inputMode`)
Once the CLI is ready (see §8), the prompt is written to the session with a
trailing carriage return:
| Mode | Mechanism | Use when |
| ------- | --------------------------------------------------------------- | ------------------------------------------------ |
| `typed` | `session.writeViaMux()` — tmux `send-keys -l` (literal) + Enter | Default; behaves like a human typing the prompt. |
| `paste` | `session.write()` — writes directly to the PTY/mux | Bulk paste-style delivery. |
> ⚠️ **Single-line only — enforced.** Like all programmatic input in Codeman,
> multi-line delivery would be silently corrupted (Ink-based TUIs treat a
> newline as submit; typed mode fuses lines). So newlines are **rejected**: the
> schema and the form refuse a multi-line `promptText`, and at fire time a
> prompt file whose content is multi-line (after stripping trailing newlines)
> fails the run with a clear `errorMessage`. Put multi-line instructions in a
> file the agent is told to read itself (e.g. "read TASKS.md and do it").
---
## 7. Concurrency policy (automatic runs)
`concurrencyPolicy` governs what happens when a **scheduled** run is due and
sessions of the same `agentType` already exist:
| Policy | Behavior |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `warn_only` | Always launch. (The count is surfaced but not blocking.) |
| `skip_if_same_agent_running` | If ≥ 1 **other, live** session of that mode is active, **skip** this fire — record a `skipped` run and (for recurring schedules) advance the schedule without launching. |
Notes on `skip_if_same_agent_running`:
- Only **live** sessions block: a tab whose CLI already exited (status
`stopped`/`error`) does not count.
- Sessions created by **this job's own previous runs never block it** —
otherwise a recurring job would deadlock on the session it created last time
and fire exactly once.
- A skipped **`once`** job is **not consumed**: it stays armed and retries on
the next tick until the blocking session goes away, then fires its single run.
- A skip is **not** a run: it sets `lastStatus = 'skipped'` but does **not**
advance `lastRunAt`.
- Consecutive skips are **coalesced** — a perpetually-skipped interval job writes
**one** skip record per streak, not one every tick, so it can't bloat
`state.json`.
**Run Now ignores this policy on the server.** The browser shows a `confirm()`
warning if same-type sessions are active, but if you proceed (or call the API
directly), the job launches unconditionally.
---
## 8. What happens when a job fires
Sequence in `CronService.launch()`:
1. A `CronJobRun` is created with status **`created`** and broadcast
(`cron:runCreated`).
2. The prompt is resolved (inline or file, single-line enforced). Failure →
**`failed`**.
3. `workingDir` is checked (`statSync().isDirectory()`). Missing/not-a-dir →
**`failed`**.
4. **Auto-close previous session** (recurring schedules, unless
`autoClosePreviousSession: false`): any still-open session created by this
job's previous runs is closed via the normal session-cleanup path.
5. The global session cap is checked (`MAX_CONCURRENT_SESSIONS = 50`). At cap →
**`failed`**.
6. A `Session` is created **with `useMux: true`** (so it runs inside tmux),
registered, listeners attached, and started via `startInteractive()`
(`startShell()` for `shell` mode). Model/claudeMode come from global config.
Run status → **`session_started`**.
7. **Readiness wait** (async, non-blocking): for non-shell agents the service
polls the terminal buffer up to **60 × 500ms** for a `❯` prompt or the string
`tokens`, then settles **2000ms** (`CRON_READY_SETTLE_MS`). Shell mode waits
1000ms, then sends the optional `launchCommand` as the first input line
(+1000ms settle).
8. The prompt is delivered (`typed`/`paste`, trailing `\r`). Run status →
**`prompt_sent`**; `finishedAt` stamped. Delivery failure (e.g. the mux
session is gone) → **`failed`**.
The created session is a **normal, persistent interactive session** — it appears
as its own tab and keeps running after the prompt is sent. The run's
`createdSessionUrl` is a deep link (`/?session=<id>`); the UI focuses it
automatically after **Run Now**.
> ⚠️ **Session-cap math if you disable auto-close.** With
> `autoClosePreviousSession: false`, nothing ever closes the sessions a
> recurring job creates — an interval job every 30 min creates 48 tabs/day and
> hits the global 50-session cap in ~25 hours (sooner with existing tabs), after
> which **every** fire of **every** job fails with "Maximum concurrent sessions
> reached" until you delete tabs by hand. Leave auto-close on for unattended
> recurring jobs, or clean up sessions yourself.
### The background tick
`tickDueJobs()` runs every **30s** (`CRON_TICK_INTERVAL`, registered in
`server.ts`). For each enabled job whose `nextRunAt ≤ now`:
- **Duplicate-launch guard**: `lastDueKey = jobId:fireTime`. If this due time was
already consumed (overlap/restart), the job is just advanced, not relaunched.
- The schedule is **advanced _before_ launching** so a slow launch can't be
re-triggered by the next tick.
- On boot, `init()` recomputes `nextRunAt` for loaded jobs (dead `once` jobs stay
dead).
---
## 9. Run history & statuses
Each job keeps a history of `CronJobRun` records. Statuses (`CronJobRunStatus`):
| Status | Meaning |
| ----------------- | ------------------------------------------------------------- |
| `created` | Run record created; prompt/session not yet started. |
| `session_started` | Session launched successfully. |
| `prompt_sent` | Prompt delivered — the happy-path terminal state. |
| `failed` | Something went wrong (see `errorMessage`). |
| `skipped` | A scheduled fire was skipped by `skip_if_same_agent_running`. |
Each run also records `triggerType` (`scheduled` or `manual_run_now`),
`sessionId`/`sessionName`, timestamps, and `createdSessionUrl`.
**History is capped globally** at **500 records** (`MAX_CRON_RUN_HISTORY`); the
oldest are pruned first. Deleting a job also deletes its run records.
---
## 10. API reference
All responses use the standard `ApiResponse<T>` envelope (`{success, data}` /
`{success, error, errorCode}`). `/api/v1/*` is a stable alias.
| Method | Endpoint | Body | Returns |
| -------- | ---------------------------- | ---------------------- | --------------------------------- |
| `GET` | `/api/cron/jobs` | — | `CronJob[]` |
| `POST` | `/api/cron/jobs` | `CronJobSchema` | `{ job }` |
| `GET` | `/api/cron/jobs/:id` | — | `CronJob` (404 if missing) |
| `PUT` | `/api/cron/jobs/:id` | partial `CronJob` | `{ job }` (400 if merge invalid) |
| `DELETE` | `/api/cron/jobs/:id` | — | `{}` |
| `PUT` | `/api/cron/jobs/:id/enabled` | `{ enabled: boolean }` | `{ job }` |
| `POST` | `/api/cron/jobs/:id/run` | — | `{ run, activeAgents }` |
| `GET` | `/api/cron/jobs/:id/runs` | — | `CronJobRun[]` (newest first) |
| `GET` | `/api/cron/runs` | — | all `CronJobRun[]` (newest first) |
---
## 11. SSE events
Emitted on `/api/events`, mirrored in `SSE_EVENTS` (`constants.js`):
| Event | Payload | When |
| ------------------ | ------------ | -------------------------------------------------------------------- |
| `cron:jobsChanged` | `{ jobs }` | Any job created / updated / enabled / status change. |
| `cron:jobDeleted` | `{ id }` | A job was deleted. |
| `cron:runCreated` | `CronJobRun` | A run (incl. skips) started. |
| `cron:runUpdated` | `CronJobRun` | A run advanced state (`session_started` / `prompt_sent` / `failed`). |
---
## 12. State & persistence
Persisted in `~/.codeman/state.json` via `StateStore`:
- `AppState.cronJobs` — map of `id → CronJob`.
- `AppState.cronJobRuns` — map of `id → CronJobRun`.
Jobs and their schedules survive restarts; `init()` recomputes `nextRunAt` on
boot. Sessions the jobs create persist through the normal session-recovery path.
---
## 13. Limits & constants
| Constant | Value | Source |
| ------------------------ | --------------------- | ------------------------------------------------ |
| Due-tick interval | 30s | `CRON_TICK_INTERVAL` (`config/server-timing.ts`) |
| Readiness poll | 60 × 500ms | `CRON_READY_MAX_ATTEMPTS` |
| Readiness settle | 2000ms | `CRON_READY_SETTLE_MS` |
| Run-history cap (global) | 500 | `MAX_CRON_RUN_HISTORY` (`config/map-limits.ts`) |
| Saved-jobs cap | 100 | `MAX_CRON_JOBS` (`config/map-limits.ts`) |
| Concurrent-session cap | 50 | `MAX_CONCURRENT_SESSIONS` |
| Prompt-file size cap | 1 MiB | `MAX_PROMPT_FILE_BYTES` (`cron-service.ts`) |
| `name` length | 1–200 | `CronJobSchema` |
| `promptText` length | ≤ 100000 | `CronJobSchema` |
| `intervalMinutes` | 1–525600 | `CronJobSchema` |
| `weeklyDays` | 1–7 entries, each 0–6 | `CronJobSchema` |
---
## 14. Known limitations
- **Server-local timezone only** — `daily`/`weekly` times are interpreted in the
host's local time; there is no per-job timezone.
- **Interval drift** — `interval` re-anchors to the actual fire time; long-running
intervals slowly shift.
- **Single-line prompts** — multi-line prompts are rejected (schema, form, and
at fire time for prompt files); tell the agent to read a file itself for
multi-line instructions.
- **`runNow` / tick race** — a manual Run Now firing at the same instant as a
scheduled tick is theoretically possible; benign (you may get two sessions).
- **`{enabled:true}` on a dead `once` job** — re-enabling a fired one-time job
without changing its schedule leaves it enabled-but-dead (won't fire); change
the schedule to re-arm.
---
## 15. Troubleshooting
| Symptom | Likely cause | Fix |
| ------------------------------ | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
| Job never fires | Disabled, or `nextRunAt: null` | Check **Enabled**; verify the schedule fields are complete. |
| Run shows `failed` immediately | Bad `workingDir`, prompt-file rejected, or session cap hit | Read `errorMessage` on the run; confirm the dir exists and the prompt file is inside it and < 1 MiB. |
| Run shows `skipped` | `skip_if_same_agent_running` + another live same-type session (this job's own sessions and dead tabs don't count) | Switch to `warn_only`, or wait for the other session to end. |
| Run fails with "single line" | Multi-line prompt text / prompt file | Keep the prompt to one line; point the agent at a file to read for long instructions. |
| Sessions pile up between runs | `autoClosePreviousSession: false` | Re-enable auto-close, or delete old tabs before the 50-session cap bites (see §8). |
| Wrong fire time | Timezone assumption | Times are **server-local** — check the host clock/TZ. |
| One-time job won't re-fire | `completedOnce` set | Edit the schedule (any real schedule change re-arms it). |
---
## 16. Related docs
- `docs/cron-discovery.md` — architecture / integration-point analysis (why the
feature reuses the session layer and stays distinct from `ScheduledRun`).
- `docs/cron-build-brief.md` — the original build brief / requirements.
- `CLAUDE.md` → **Key Patterns → Cron** — the one-paragraph engineering summary.
- Tests: `test/cron-time.test.ts` (schedule math), `test/cron-service.test.ts`
(CRUD, tick, concurrency, security).
+433
View File
@@ -0,0 +1,433 @@
<!-- Design doc generated via ultracode multi-agent workflow (wf_e3a7498b-26f): 3 architecture proposals -> judge panel -> synthesis -> completeness critic. -->
# Docker Session Mode, Implementation Plan
## Decisions (locked 2026-07-19, by repo owner)
1. **Isolation posture**: CONVENIENT default (bind-mount host `~/.claude` etc. read-write so the existing login just works; network on; still hardened non-root + cap-drop + resource caps). SEALED profile (`mountCredentials:false` + `network:none`) is a per-case opt-in.
2. **Export**: offer BOTH full-image (`commit`+`save`+workspace tar) AND workspace-only, side by side, no default (ask each time).
3. **Base image**: BUILD LOCALLY on first use via `scripts/build-agent-image.mjs` from a repo `docker/agent.Dockerfile`. No registry required. (GHCR pull can be added later.)
4. **Hooks**: WIRE HOOKS NOW. Codeman scaffolds `.claude/settings.local.json` + CLAUDE.md into the linked host workspace dir (same as local cases), enabling in-container permission prompts, hook-idle detection, and the Claude Model picker.
Adopted defaults for the remaining open items (Section 10): resume-on-restart ON; container is per-CASE and shared by multiple sessions (killing one session only kills its in-container tmux session, never `docker stop` while siblings remain; stop/remove only on explicit teardown or case-delete); rootless caps = ship-with-warning (`capsEnforced` surfaced); remote docker daemon = local-first; podman = docker-first best-effort.
## Implementation status (branch `feat/docker-session-mode`)
DONE and END-TO-END VERIFIED against a real docker daemon (create host, link case, quick-start shell in a real container, workspace bind-mount round-trip, hook scaffolding, session-delete keeps the shared container up, case-delete `docker rm`s it):
- Phase 0-1: types (`DockerHost`/`DockerCase`/`SessionDocker`), `src/docker-hosts.ts` (storage, pure `buildDockerBaseArgs`/`buildDockerCreateArgs`, `containerApiUrl`, `hostGatewayAlias`, config-hash, credential-mount resolution, daemon probes), `DockerHostSchema`/`DockerCaseLinkSchema`. 26 unit tests.
- Phase 2: `tmux-manager` `buildDockerLaunchCommand` (image-check -> ensure -> start -> exec, resume-aware), `buildDockerKillCommand` (in-container tmux only, multi-session safe), stop/remove; wired into `createSession`/`respawnPane`/`killSession`. 14 unit tests.
- Phase 3: `Session` threading (`_docker`, toState, option builders, in-container cliVersion probe, `resolveMuxAttachCwd`), `server.ts` recovery round-trip.
- Phase 4: `case-routes` `/api/docker-hosts` CRUD + `/api/cases/docker-link` + listing + docker-unlink; `session-routes` `/api/quick-start` docker branch (rejects per-session config, probes availability + tmux, scaffolds hooks, seeds resume id).
- Phase 5 (partial): `docker/agent.Dockerfile` + `scripts/build-agent-image.mjs` (built + verified: node 22, tmux, claude/codex/gemini/opencode, arbitrary-uid HOME). Host-guard allowlists `host.docker.internal`/`host.containers.internal` for in-container hooks.
- Full CI green (3445 tests).
REMAINING:
- Phase 6: export / import (`docker commit` + `save | gzip` + workspace tar + manifest; `load` + quarantined re-tag), GC / boot reaper, disk-safety prechecks, drift-recreate route, SSE `docker:*` events. THE "move to a new machine" feature.
- Phase 7: frontend Create Case "Docker" tab + `linkDockerCase` + run wiring + case-picker labels + export/import UI.
- Phase 8: CLAUDE.md "Docker cases" Key Pattern + `docs/docker-cases.md` + COM.
- Deferred refinements: in-container model-picker via `settings.local.json`; live mid-run resume-id capture into `DockerCase.lastClaudeSessionId`; rootless/Desktop uid probe (currently a platform heuristic).
## 1. Goal & user stories
Add "Docker cases" to Codeman: a case can point at a container instead of a local or remote-SSH path, and any of the five CLI backends (`claude` / `shell` / `opencode` / `codex` / `gemini`) runs inside that container. It is modeled as a LOCATION OVERLAY on cases, exactly like the remote-SSH feature (COD-94/#145), never as a sixth `SessionMode`.
User stories:
- As the repo owner, I link a case to a per-project container so an autonomous Claude/Ralph run executes in a hardened sandbox (cap-drop, non-root, resource caps) instead of directly on my host, while keeping my existing OAuth login and transcript history working with zero extra setup.
- I set default, per-case-changeable container settings (image, network mode, memory/cpu/pids caps) at link time and edit them later, and edits actually take effect through a recreate-on-drift path (see Section 4).
- I reconnect after a Codeman restart and land back in the SAME running agent with the conversation intact. When the CONTAINER itself was stopped/rebooted/OOM-killed (which destroys the in-container tmux), the next launch RESUMES the last conversation from the bind-mounted transcript rather than starting fresh (durability model in Section 2, Key decision 1).
- I export a finished run's whole environment (toolchain plus workspace) to a portable, secret-free `.tar.gz`, move it to another machine, and import it back into a fresh case in one click.
- The container never accumulates: killing the session stops it, deleting the case removes it, and an instance-scoped boot reaper reaps containers whose case is gone.
Non-goals for the MVP: multi-tenant untrusted-code isolation guarantees (Codeman is loopback-default and single-operator, and the agent already runs `--dangerously-skip-permissions` on the host today), Kubernetes/compose orchestration, and per-command ephemeral containers.
## 2. Chosen architecture and why
The design grafts the strongest idea from each of the three proposals:
- Overlay-not-a-mode + faithful remote-SSH mirror (from "Docker Cases as a Location Overlay"): lowest churn, rides the existing quick-start / mux-sessions / state / recovery plumbing.
- Convenient-but-hardened default with an opt-in sealed profile, plus exec-time name-only secret env (from "Sealed Sandbox"): a strict security improvement over today's on-host execution without the UX tax of forcing an in-container re-login.
- One-artifact export + in-app import route (from "Container-as-Cargo"): the genuinely new, high-value capability Codeman lacks.
### Key decision 1: persistent per-CASE container, durable in-container tmux, AND resume-on-restart (the two-layer durability model)
Exactly one long-lived container per Docker case, named as a pure slug function `codeman-case-<slug>` (Docker charset `^[a-zA-Z0-9][a-zA-Z0-9_.-]+$`; Codeman already slugs case names for tmux), so create-if-missing and boot recovery are idempotent. PID1 is `sleep infinity` under `--init` (tini reaps zombies and forwards `docker stop`'s SIGTERM); the CLI is NOT the container command. The CLI runs inside a DURABLE in-container tmux on a dedicated socket `-L codeman-docker`, session `codeman-dkr-<id8>`, the direct analog of remote's `-L codeman-remote` / `codeman-ssh-<id8>`.
Two DIFFERENT failure surfaces need two DIFFERENT recovery layers, and conflating them is the central flaw the critic caught:
1. Codeman-PROCESS restart while the container stays up: the in-container tmux is still alive, so `tmux new-session -A` (attach-or-create) reattaches the SAME live agent and the paneCommand is ignored. This is the remote-SSH durability idiom and it works unchanged.
2. CONTAINER stop / daemon restart / host reboot / OOM-kill: the in-container tmux is GONE (fresh PID1). `new-session -A` will now CREATE a fresh session and run the paneCommand, which would start a brand-new conversation. This is the case the raw plan silently lost. Because the transcript directory is bind-mounted from the host (Key decision 3), the fix is to launch with RESUME: the paneCommand becomes `exec claude --dangerously-skip-permissions --resume <claudeSessionId>` (codex uses `resume <id>`, gemini `--resume <id>`) whenever a captured `claudeSessionId` exists. The `-A` semantics make this self-selecting: the resume flag only ever executes when tmux is actually re-created, which is exactly when the live session was lost. When tmux is still alive (case 1), attach wins and the flag is inert.
Capturing / persisting / reusing the resume id (the missing mechanism the critic flagged): Codeman already learns `Session.claudeSessionId` from transcript correlation (which works here because projHash matches, Key decision 3) and persists it in `SessionState`. We thread that value into `createSessionOptions` / `respawnPaneOptions` for docker so `buildDockerLaunchCommand` can inject the resume flag on any relaunch. To make a NEW Codeman session (new `id8`) re-launched against the same case resume its predecessor's conversation, we ALSO persist `lastClaudeSessionId` on the `DockerCase` record; the quick-start docker branch seeds the new `Session` with it when the `dockerResumeOnStart` setting is on. First-ever launch has no id, so it starts fresh. This is user-decision 7 (default resume behavior).
Reconciling with stop-on-kill and with the `--restart` policy (the internal inconsistency the critic found): the container is created with `--restart no` uniformly (Codeman's idempotent create-if-missing plus boot recovery is the single recovery mechanism; a restart policy would not preserve the conversation anyway because a restarted container gets a fresh PID1/tmux). Boot recovery re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` (`docker inspect || docker create; docker start`, then exec with resume), so a host reboot or daemon restart recreates+starts the container and resumes the conversation instead of the session vanishing. `reconcileSessions` (tmux-manager.ts ~1800-1815) must NOT hard-delete a docker session merely because no LOCAL pane exists after the local `-L codeman` server died; docker (like remote) sessions are restored from `mux-sessions.json` and relaunched. This relaunch path is explicitly part of Phase 4/Phase 3 recovery work, not assumed.
Why this over the alternatives: `docker exec` gets SIGHUP and dies when its client TTY closes, so a bare `docker exec claude` restarts the CLI on every reconnect/respawn. The inner tmux plus resume is what makes reconnect idempotent across BOTH failure surfaces. Because this durability is the single most important design point, tmux-in-image is a HARD gated prerequisite (`checkDockerTmuxAvailable`), never a silent fallback to bare exec. Rejected alternatives: ephemeral-per-run or bare-exec containers (no reattach durability); a literal `'docker'` `SessionMode` (touches dozens of switch/enum sites and diverges from the remote overlay precedent, since Docker is a LOCATION orthogonal to the 5 CLI backends).
### Key decision 2: CLI + auth delivery
One prebuilt base image (built once, contains NO secrets): `node:22-bookworm-slim` + `git tmux ripgrep ca-certificates`, `npm i -g @anthropic-ai/claude-code @openai/codex @google/gemini-cli opencode-ai`, an `agent` user, HOME dirs made writable by an arbitrary host uid via the OpenShift "gid 0, group-writable" convention (Key decision 6). Because the toolchain is baked, export is reproducible and needs no network at import time. The image name/namespace/registry and its refresh cadence are user-decision 2 (the `codeman/agent:base` placeholder implies a Docker Hub org the project may not own).
Credentials are delivered ONLY at runtime, two commit-safe channels, default convenient:
- OAuth/config-file CLIs (Claude Max/Pro, gcloud, opencode): bind-mount the host credential dirs read-write (`~/.claude`, `~/.codex`, `~/.gemini` + `~/.config/gcloud`, `~/.config/opencode`) so the common user "just works" with no in-container login. Because these are bind mounts, `docker commit` (which captures only the container's own writable layer, never bind mounts) physically cannot capture them, so exports stay secret-free.
- API-key CLIs (codex/gemini): exec-time NAME-ONLY `docker exec --env OPENAI_API_KEY --env GEMINI_API_KEY ...` (no `=value`), sourced from Codeman's own process env. Only the key NAME appears in argv (no `ps` leak), and per-exec env is never captured by `docker commit`. This is the technique Codeman already uses via `tmux setenv` for the local Codex/Gemini panes, so it composes with existing machinery.
Per-host `DockerHost.mountCredentials` defaults `true` (convenient); setting it `false` yields a SEALED profile (no host cred mounts, in-container login only) for genuinely untrusted work. CRITICAL sealed-mode export rule (the leak the critic caught): in sealed mode the in-container login writes tokens into the container's OWN writable layer, which `docker commit` DOES capture, so a full-image export of a sealed container would ship credentials. Therefore full-image export is REFUSED for `mountCredentials:false` containers by default; the user may either take a workspace-only export (always safe) or opt into a pre-commit scrub that `docker exec`s `rm -rf ~/.claude ~/.codex ~/.gemini ~/.config/gcloud ~/.config/opencode` inside the container before commit (destructive to the in-container login, which is the point). This is enforced in the export route, not left to a manifest assertion.
Per-session `envOverrides` / `effort` / `codexConfig` / `geminiConfig` / `openCodeConfig` are REJECTED at quick-start exactly like the remote branch (session-routes.ts ~1698-1710). `modelOverride` is the one deliberate difference from remote: because the docker workspace is a REAL bind-mounted host dir that Codeman scaffolds (Key decision 5 and Section 6), `updateCaseModel()` can write the `model` key into `<workspace>/.claude/settings.local.json` and the in-container `claude` reads it, so the App Settings Claude Model picker works for docker cases. `effort` is a `--effort` CLI arg applied only by the local-spawn path we bypass, so it stays rejected (surfaced honestly in the UI, not silently inert). Per-mode command customization goes through `DockerHost.commands.<mode>` (`defaultDockerCommandForMode`, mirror of `defaultRemoteCommandForMode` at remote-hosts.ts:60). NEVER bake secrets into an image layer and NEVER pass a secret via create-time `-e` (both are committed).
Rejected alternative: sealed-by-default. For a single-operator loopback tool where the agent already runs skip-permissions on the host, forcing an in-container OAuth re-login is a UX regression with little real gain. We keep sealed as an opt-in. Rejected alternative: baking a login into the image, which leaks the instant you `docker save`.
### Key decision 3: workspace mount, container CWD, and transcript correlation
Bind-mount the host workspace dir into the container at the SAME absolute path (`dst == src`, mirror the host path), and set both `Session.workingDir` and the container workdir to that host path.
Two problems this solves that the raw proposals got wrong:
- File features: `DockerCase.hostWorkspacePath` is a REAL host directory, so `Session.workingDir = hostWorkspacePath` keeps file-routes, attachments, image-watcher, and previews working on real host bytes (unlike remote, where the path is remote-only and those features no-op). All three proposals wired `casePath = <container path>`; we deliberately diverge and use the host path.
- Transcript correlation: Claude writes transcripts under `~/.claude/projects/<hash-of-CWD>/`. By mirroring the host path as the container CWD, the projHash computed inside the container equals the host-side hash Codeman's transcript/subagent/workflow watchers expect, so correlation keeps working (and, in turn, feeds the resume-id capture in Key decision 1). A `/workspace`-style fixed dst would break it. Mirror-vs-fixed is user-decision 3.
`resolveMuxAttachCwd` still returns `/tmp` for docker sessions (the LOCAL bash pane only runs `docker exec`; it never needs the workspace as its cwd), mirroring remote.
### Key decision 4: network default and the engine-specific host gateway
Default `bridge` (own netns, NAT egress, no inbound), per-case changeable to `none` (offline shell sandbox; warned because it breaks the API CLIs) or `custom` (a user-defined bridge `codeman-net-<slug>`, the chokepoint for a future egress allowlist). `host` networking and any `-p` inbound publish are structurally unrepresentable in the flag builder and schema. Rationale: every API-backed CLI (Claude, Codex, Gemini) plus npm/git needs egress, so `bridge` is the only sane functional default; `none` is reserved for `shell`.
The host-callback gateway alias is ENGINE-SPECIFIC (the critic's podman finding): Docker uses `host.docker.internal`, Podman uses `host.containers.internal` (Docker's alias only exists on recent podman). A helper `hostGatewayAlias(engine)` returns the right name; Section 2.5, the create args, the `CODEMAN_API_URL` rewrite, and the host-guard allowlist all consume it, and BOTH aliases are added to the allowlist so a mixed fleet keeps working.
### Key decision 5: hooks actually reach the host AND are actually installed
Two independent things must both be true for a hook to fire, and the raw plan wired only the first:
1. Network reachability. Claude Code hooks POST to `$CODEMAN_API_URL` (`curl -sk`). Inside a bridge container `localhost` is the container and prod binds `127.0.0.1`, so we set `--add-host <gatewayAlias>:host-gateway` on create (skipped on Docker Desktop, where the alias is native), add the gateway alias to the host guard, and provide `CODEMAN_API_URL` and the hook secret (below).
2. Hook INSTALLATION. Hooks live in `<workspace>/.claude/settings.local.json`, written by the quick-start scaffolding block (around session-routes.ts ~1776) that calls `writeHooksConfig()` / `updateCaseModel()`. The raw plan extended the `!remote` guard to `!remote && !docker`, which would SKIP that block and silently disable ALL hooks regardless of networking. For docker the workspace is a REAL bind-mounted host dir, so the scaffolding block MUST run. Precise fix: extend to `!remote && !docker` ONLY the LOCAL-CLI-availability and local-spawn guards (the ones that stat the local binary or build the local spawn command); leave the workspace-scaffolding guard at `!remote` so it runs for docker. This same decision is what makes `modelOverride` work (Key decision 2). Consequence, surfaced as user-decision 4: linking a docker case now WRITES `.claude/settings.local.json` (and the CLAUDE.md scaffold, matching local-case behavior) into the user's real host directory, a behavioral shift from "link a dir" to "link and scaffold a dir."
`CODEMAN_API_URL` derivation (the wrong-scheme bug the critic caught): prod is HTTPS-only on 3000, and `server.ts` (~2000) auto-sets `process.env.CODEMAN_API_URL = ${protocol}://${apiHost}:${port}`. Hardcoding `http://host.docker.internal:3000` fails every hook. Instead a pure helper `containerApiUrl(process.env.CODEMAN_API_URL, engine)` parses the running URL and substitutes ONLY the hostname with `hostGatewayAlias(engine)`, preserving scheme and port (`https://host.docker.internal:3000`). Unit-tested against http, https, non-default ports, and both engines. Passed as create-time `--env CODEMAN_API_URL=<derived>` (case-stable, non-secret).
Hook secret and session attribution:
- `~/.codeman/hook-secret` is bind-mounted read-only to a container path; `--env CODEMAN_HOOK_SECRET_FILE=<that path>` is create-time (a path is non-secret; the bytes ride the bind mount and are never committed).
- `CODEMAN_SESSION_ID` (which the generated hooks reference at hooks-config.ts:78-80 to attribute events) plus `CODEMAN_MUX=1` are SESSION-scoped, so they are passed at EXEC time via `docker exec --env CODEMAN_SESSION_ID=<id> --env CODEMAN_MUX=1` (non-secret, value inline is fine, and exec env is not committed). Because a `tmux` session started fresh only inherits the invoking env when it starts the SERVER, the launch chain ALSO runs `tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID <id>` (and `CODEMAN_MUX`) so reattaches and newly created panes see the same values. This mirrors how Codeman already injects per-session env into tmux for the external CLIs.
Hooks-in-MVP-vs-deferred stays user-decision 4; if deferred, docker ships as explicitly hook-degraded and we lean on output-based idle detection through the docker-exec PTY.
### Key decision 6: uid / HOME / rootless enforcement / macOS Docker Desktop
The raw plan showed `--user 1000:1000` in one place and `--user "$(id -u):$(id -g)"` in another and never resolved HOME writability; this section fixes all of it.
- Linux native (docker rootful or rootless): run `--user <hostUid>:0` (host uid, GID 0). The image follows the OpenShift arbitrary-uid convention: `HOME=/home/agent`, and `/home/agent` plus the tool cache dirs (`~/.npm`, `~/.cache`, `~/.config`) are owned `root:0` and group-writable (`chmod -R g+w`, `g+s` on dirs) so a process with GID 0 can write HOME even though its UID is not 1000. This keeps workspace files host-owned (the agent's UID is the host UID) AND keeps HOME writable, so the CLIs actually start.
- Podman rootless: use `--userns=keep-id` (maps the host uid to the image's `agent` uid inside the container) instead of `--user`, so `/home/agent` is owned by the running user and workspace files are host-owned. This is a real per-engine branch in `buildDockerCreateArgs`.
- macOS Docker Desktop: `--user <macUid>` (e.g. 501) does not own the image's `/home/agent`, so non-bind HOME writes fail EACCES and the CLIs may not start; Desktop also does its own bind-mount uid translation, provides `host.docker.internal` natively (no `--add-host`), and its VM memory ceiling can cap `--memory`. Detect Desktop via `docker info` (Server OS `linuxkit` / `OperatingString` contains "Docker Desktop") and take a dedicated path: do NOT pass `--user` (run as the image's baked `agent` uid and rely on Desktop's translation for workspace access), skip `--add-host`, and note in the UI that memory caps are subject to the VM ceiling.
Rootless resource-cap enforcement (the silently-inert risk): rootless Docker without cgroup-v2 systemd delegation (`Delegate=yes`) silently IGNORES `--memory`/`--cpus`/`--pids-limit`. The probe checks `docker info` for `CgroupVersion=2` plus rootless plus delegation; if caps cannot be enforced, `checkDockerAvailable` returns `capsEnforced:false` and the link/probe surfaces "resource caps are advisory on this engine." Whether to REQUIRE delegation or ship-with-warning is user-decision 6.
## 3. Data model
New TypeScript types in `src/types/session.ts`, added right after the remote types (lines 46-99). SessionMode (line 44) is UNCHANGED.
```ts
export type DockerCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
export type DockerEngine = 'docker' | 'podman';
export type DockerNetworkMode = 'bridge' | 'none' | 'custom'; // never 'host'
export interface DockerResourceLimits {
memory?: string; // '4g' -> --memory 4g --memory-swap 4g (swap==memory: real OOM cap)
cpus?: string; // '2'
pidsLimit?: number; // 512 (fork-bomb guard)
nofile?: string; // '4096:8192'
shmSize?: string; // optional; only when a tool needs /dev/shm
}
export interface DockerHost {
id: string;
label: string;
engine?: DockerEngine; // default resolved by probe (docker, else podman)
image: string; // default resolved image ref (see user-decision 2)
daemonHost?: string; // advanced: -H ssh://user@host / DOCKER_HOST
context?: string; // advanced: --context <ctx>
network?: DockerNetworkMode; // default 'bridge'
networkName?: string; // when network === 'custom'
resources?: DockerResourceLimits;
mountCredentials?: boolean; // default true (false = sealed; blocks full-image export)
hooksEnabled?: boolean; // default true (host-gateway callback wiring)
resumeOnStart?: boolean; // default true (see Key decision 1 / user-decision 7)
commands?: Partial<Record<DockerCommandMode, string>>;
extraCreateArgs?: string[]; // validated like extraSshOptions
extraExecArgs?: string[];
}
export interface DockerCase {
name: string;
type: 'docker';
hostId: string;
hostWorkspacePath: string; // absolute HOST dir: bind src + Session.workingDir
containerWorkdir?: string; // container path; default = hostWorkspacePath (mirror -> projHash match)
container?: string; // default codeman-case-<slug>
lastClaudeSessionId?: string; // captured resume id (Key decision 1)
}
export interface SessionDocker { // flattened, round-trips through mux/state (mirror SessionRemote at 91)
hostId: string;
label: string;
engine: DockerEngine;
image: string;
containerName: string;
hostWorkspacePath: string;
containerWorkdir: string;
network: DockerNetworkMode;
networkName?: string;
resources?: DockerResourceLimits;
mountCredentials: boolean;
hooksEnabled: boolean;
resumeOnStart: boolean;
daemonHost?: string;
context?: string;
commands?: Partial<Record<DockerCommandMode, string>>;
extraCreateArgs?: string[];
extraExecArgs?: string[];
configHash?: string; // drift detection (Key decision, Section 4)
}
```
- `SessionState` gains `docker?: SessionDocker` immediately after `remote?` (line 219). It persists automatically because `SessionState` is structural and `state-store.ts` stores `toState()` verbatim.
- `src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (after line 38), `CreateSessionOptions` (after 81), `RespawnPaneOptions` (after 105). `MuxSession.docker` round-trips through `mux-sessions.json` automatically.
- `src/types/api.ts` `CaseInfo`: add `'docker'` to the `location` union and a `docker?: { hostId; container; image?; path; network }` display block.
- `src/services/unified-session-service.ts`: add a boolean `docker?` flag on `UnifiedSessionItem` and source rows, set from `MuxSession.docker` presence (mirror the `remote` flag at ~line 200 and the harvest at session-routes.ts:2313).
New state files (all via `dataPath()`, mirroring `remote-hosts.json` / `remote-cases.json`):
- `~/.codeman/docker-hosts.json` (reusable engine/image/network/resource profiles).
- `~/.codeman/docker-cases.json` (`name -> DockerCase`, including `lastClaudeSessionId`).
- `~/.codeman/docker-exports/` (dedicated dir for `.image.tar.gz` + `.workspace.tar.gz` + `manifest.json`; never inline in state.json; retention/pruning per Section 5).
No new `state.json` / `mux-sessions.json` files: `SessionState.docker` and `MuxSession.docker` ride the existing serialization.
## 4. Container lifecycle (exact command shapes)
All builders are PURE string functions (directly unit-testable). Host values interpolated into the outer `bash -c "..."` layer (container name, image, workdir, host paths) are `shellescape()`'d and, for user-supplied fields, schema-rejected for `$`/backtick via `NO_SHELL_META`. The escaping chain here is DEEPER than remote's single `ssh '<tmux ...>'`: the whole `docker inspect || docker create <dozens of --mount/--env/shellescaped host paths>` is interpolated into `bash -c "..."` then `JSON.stringify`'d into respawn-pane. This is a known place to get stuck, so it is covered by concrete escaping tests (Section 9), including host workspace paths containing spaces, not just a "we call shellescape" claim.
New in `src/tmux-manager.ts`:
```ts
const DOCKER_TMUX_SOCKET = 'codeman-docker';
// 'dkr' letters deliberately FAIL SAFE_MUX_NAME_PATTERN (^codeman-[a-f0-9-]+$),
// so a Codeman running INSIDE the container never adopts/resizes/respawns our session.
export function dockerTmuxSessionName(id: string): string { return `codeman-dkr-${id.slice(0, 8)}`; }
```
`buildDockerBaseArgs(docker)` (pure, in `docker-hosts.ts`, mirror of `buildSshConnectionArgs`) emits the engine prefix tokens: `docker` (or `podman`) + optional `--context <ctx>` or `-H <daemonHost>`. `buildDockerCreateArgs(docker, sessionId)` emits the `docker create` flag array (with the per-engine uid/userns branch from Key decision 6).
IMAGE PRESENCE (before any create, the auto-pull footgun the critic caught): the launch chain runs `docker image inspect <image> >/dev/null 2>&1` first; on miss it exits with a distinct message ("base image <ref> not present: build with scripts/build-agent-image.mjs or pull it") rather than triggering a blocking multi-GB auto-pull inside the tmux pane. `docker create` carries `--pull=never`. The tmux-availability probe likewise uses `docker run --rm --pull=never <image> sh -lc 'command -v tmux'` and reports the same build/pull hint if the image is absent, so the 15s-bounded probe never hangs on a pull.
CREATE (the ensure step, embedded in the launch string):
```
docker create \
--name codeman-case-myproj --hostname myproj \
--label codeman.managed=1 --label codeman.instance=<CODEMAN_INSTANCE> \
--label codeman.case=myproj --label codeman.session=<id8> \
--label codeman.confighash=<hash> \
--pull=never --init --restart no \
--user 1000:0 \
--workdir '/home/arkon/cases/myproj' \
--mount type=bind,src='/home/arkon/cases/myproj',dst='/home/arkon/cases/myproj' \
--mount type=bind,src='/home/arkon/.claude',dst='/home/agent/.claude' \
--mount type=bind,src='/home/arkon/.codeman/hook-secret',dst='/home/agent/.codeman/hook-secret',readonly \
--add-host host.docker.internal:host-gateway \
--memory 4g --memory-swap 4g --cpus 2 --pids-limit 512 --ulimit nofile=4096:8192 \
--cap-drop ALL --security-opt no-new-privileges \
--network bridge \
--env HOME=/home/agent --env TERM=xterm-256color --env COLORTERM=truecolor \
--env CODEMAN_API_URL=https://host.docker.internal:3000 \
--env CODEMAN_HOOK_SECRET_FILE=/home/agent/.codeman/hook-secret \
codeman/agent:base \
sleep infinity
```
- `--user 1000:0` shown is the Linux-native form with GID 0 (Key decision 6); it is actually `--user <hostUid>:0`, or `--userns=keep-id` for podman rootless, or omitted on Docker Desktop. The literal is illustrative only.
- Create-time `--env` carries only NON-SESSION, non-secret, case-stable values (safe to be committed): the DERIVED `CODEMAN_API_URL` (https-preserving, Key decision 5) and the hook-secret FILE PATH. `CODEMAN_SESSION_ID`/`CODEMAN_MUX` and the codex/gemini key NAMES are exec-time only.
- `codeman.instance=<CODEMAN_INSTANCE>` is REQUIRED on the label set so the boot reaper is instance-scoped (a beta/second instance must never reap prod's containers).
- `codeman.confighash` is a stable hash of the drift-relevant create args (image, resources, network, mounts, non-session env). Drift detection (user story 2, the config-never-takes-effect gap): on launch the ensure block compares the desired hash to the existing container's label; on mismatch the launch does NOT silently reuse the stale container. Instead the docker route returns a "container config changed, recreate?" action (SSE + UI confirm), and on confirm Codeman `docker rm`'s and recreates. rm destroys in-image (non-bind) state, but the workspace and transcripts survive on their bind mounts and the conversation is restored via `--resume`, so the recreate is safe. Auto-recreate-vs-prompt is a UI choice; the MVP prompts.
- `--restart no` (resolved consistently with Key decision 1; recovery is Codeman's idempotent create-if-missing, not an engine restart policy, which also matters for Podman which has no daemon).
EXEC (`buildDockerLaunchCommand`, the docker analog of `buildRemoteLaunchCommand`, TTY-correct, resume-aware). The whole thing is ONE `bash -c` string that image-checks, ensures, starts, primes tmux env, then execs:
```
docker image inspect codeman/agent:base >/dev/null 2>&1 || { echo 'Codeman: base image codeman/agent:base not present (build or pull it)'; exit 1; } ; \
docker inspect codeman-case-myproj >/dev/null 2>&1 || docker create <all create args above> ; \
docker start codeman-case-myproj >/dev/null 2>&1 || { echo 'Codeman: container codeman-case-myproj failed to start (daemon down?)'; exit 1; } ; \
exec docker exec -it \
--workdir '/home/arkon/cases/myproj' \
--env TERM=xterm-256color --env COLORTERM=truecolor \
--env CODEMAN_SESSION_ID=1a2b3c4d --env CODEMAN_MUX=1 \
--env OPENAI_API_KEY --env GEMINI_API_KEY \
codeman-case-myproj \
sh -lc 'tmux -L codeman-docker setenv -g CODEMAN_SESSION_ID 1a2b3c4d \; setenv -g CODEMAN_MUX 1 \; new-session -A -s codeman-dkr-1a2b3c4d -c '\''/home/arkon/cases/myproj'\'' '\''cd /home/arkon/cases/myproj && exec claude --dangerously-skip-permissions --resume <claudeSessionId>'\'' \; set -t codeman-dkr-1a2b3c4d status off \; set -t codeman-dkr-1a2b3c4d mouse off \; set -t codeman-dkr-1a2b3c4d prefix C-q \; set -s escape-time 0'
```
- `docker exec -it`: `-t` allocates a PTY and forwards SIGWINCH into the container so the Ink TUI re-lays-out on pane resize; `TERM`/`COLORTERM` prevent degraded rendering. `--env OPENAI_API_KEY` (name only) is present only for codex/gemini and is exec-time (never committed). `CODEMAN_SESSION_ID`/`CODEMAN_MUX` are exec-time values plus a `tmux setenv -g` prime so reattaches and new panes inherit them (Key decision 5).
- `--resume <claudeSessionId>` (codex `resume <id>`, gemini `--resume <id>`) is appended to `modeCommand` ONLY when a captured id exists; on first launch it is omitted. `new-session -A` makes the flag inert on a live-tmux reattach and effective only when tmux is re-created (Key decision 1).
- `modeCommand = docker.commands?.[mode] || defaultDockerCommandForMode(mode)` (`exec claude --dangerously-skip-permissions`, `exec bash -l`, etc.), with the resume suffix injected by the builder.
- Escaping survives every layer identically to remote in shape but deeper in nesting: `paneCommand` (`cd ... && exec ...`) is one shellescaped tmux arg, the whole `tmuxInvocation` is one shellescaped `sh -lc` arg, and the outer string is `JSON.stringify()`'d into `bash -c` by respawn-pane (tmux-manager.ts:1329).
Wire-up (extend the two existing seams to 3-way):
- createSession (tmux-manager.ts:1276): `const fullCmd = docker ? buildDockerLaunchCommand({ mode, docker, sessionId, resumeSessionId }) : remote ? buildRemoteLaunchCommand({ mode, remote, sessionId }) : localFullCmd;`
- launchCmd cd-skip (tmux-manager.ts:1327): `const launchCmd = (remote || docker) ? fullCmd : \`cd ${JSON.stringify(workingDir)} && ${fullCmd}\`;`
- respawnPane: same two edits at lines 1524 and 1542.
START / reattach-after-reboot: the ensure block (image-check, `docker inspect || docker create`, `docker start`) is fully idempotent, so boot recovery just re-runs `buildDockerLaunchCommand` from the restored `MuxSession.docker` with the persisted resume id. A rebooted host recreates the container and resumes the conversation.
DOCKER-DOWN surfacing (the PTY-exit-breaker false-trip risk): if `docker start` or `docker exec` cannot attach (daemon down, container missing), the launch prints a docker-specific message and exits, which alone would still count toward `session-pty-exit-breaker` and show a generic "respawn breaker tripped" push. To avoid masking the cause, the docker reattach path runs a fast `checkDockerAvailable` pre-flight: if the daemon/container is unreachable, Codeman broadcasts a docker-specific error (SSE + push, "container <name> is not running / daemon down") and SKIPS the auto-reattach that would trip the breaker, rather than fast-looping `docker exec`.
STOP / KILL (`killSession` Strategy 3c, right after remote's Strategy 3b at tmux-manager.ts:1719, guarded by `IS_TEST_MODE`):
```ts
if (session.docker) {
// best-effort, fire-and-forget, timeout-bounded so it never blocks the local kill
execAsync(buildDockerKillCommand({ docker: session.docker, sessionId }), { timeout: EXEC_TIMEOUT_MS }).catch(() => {});
}
```
`buildDockerKillCommand` emits: `docker exec codeman-case-<slug> tmux -L codeman-docker kill-session -t codeman-dkr-<id8> ; docker stop -t 10 codeman-case-<slug>`. Stopping frees CPU/RAM and, per Key decision 1, is safe for conversation continuity because the NEXT launch resumes from the bind-mounted transcript via `--resume`. Whether to stop at all (RAM vs instant live-agent reattach) is user-decision 6/1 (reframed honestly). The bind-mounted workspace and transcripts always survive on the host.
REMOVE: only on explicit case delete (`docker rm -f codeman-case-<slug>`), gated behind an "export first?" UI prompt because rm destroys any in-image (non-bind) state. Instance-scoped boot reaper (fixing the racy/cross-instance reaper): after `docker-cases.json` is loaded AND after `restoreMuxSessions` has run, enumerate `docker ps -a --filter label=codeman.managed=1 --filter label=codeman.instance=<CODEMAN_INSTANCE> --format '{{.Names}}\t{{index .Labels "codeman.case"}}'` and `docker rm -f` only containers whose case is gone from THIS instance's `docker-cases.json`. The instance filter is what stops a beta reaping prod's containers (the exact cross-instance hazard the project memory warns about).
AVAILABILITY PROBE (`docker-hosts.ts`, timeout-bounded like `checkRemoteTmuxAvailable`'s 15s, `IS_TEST_MODE` no-op):
```
docker info --format '{{json .}}' # server up, CgroupVersion, rootless, OS (Desktop detect), cap-delegation
docker image inspect <image> --format '{{.Id}}' # image PRESENT (no auto-pull)
docker run --rm --pull=never <image> sh -lc 'command -v tmux' # tmux-in-image gate (hard prerequisite), only if image present
```
`checkDockerAvailable()` returns `{ ok, engine, rootless, isDesktop, cgroupV2, capsEnforced }` (parse `SecurityOptions` for `name=rootless`, `CgroupVersion`, delegation, and Server OS for Desktop). `checkDockerTmuxAvailable(host)` returns a structured result with a user-facing error and correct install hint (NOT `npm install -g`; the hint is "build/pull the base image" for a missing image and "install docker or podman" for a missing engine).
IN-CONTAINER CLI VERSION (fixing the #154 wheel-forwarding regression): the raw plan skipped the LOCAL `cliVersion` probe for docker (correct, since it reports the HOST claude) but left `cliVersion` undefined, which disables trackpad wheel-forwarding. Instead, for docker sessions Codeman runs an IN-CONTAINER probe `docker exec <container> claude --version` (bounded, `IS_TEST_MODE` no-op) and feeds THAT into `cliVersion`. This also means a stale baked CLI is visible; combined with the rebuild-cadence in user-decision 2, agents are not silently pinned to an old claude.
## 5. Export / Import
EXPORT is a concurrency-bounded job (reuse `runWithConversionLimit` from `document-conversion-limiter.ts` so N simultaneous exports cannot fork-bomb the host). Route `POST /api/docker-cases/:name/export`.
Preconditions (the consistency and leak risks the critic caught):
- Sealed guard: if `mountCredentials:false`, full-image export is REFUSED unless the caller explicitly opts into the pre-commit scrub (Key decision 2). Workspace-only export is always allowed.
- Quiesce + free-space: require the session idle, then `docker pause` the container spanning BOTH the workspace tar AND the commit so the two artifacts are mutually consistent (the raw plan paused only the commit, leaving the bind-mount tar to run against a mid-write agent). Before any heavy step, precheck free space in the exports dir and in `/var/lib/docker`; if below `DOCKER_EXPORT_MIN_FREE_BYTES`, refuse with a clear error (a full `/var/lib/docker` wedges the daemon and breaks EVERY session on the host).
Steps (all cleanup in try/finally so a mid-way failure never orphans an intermediate image or leaves the container paused):
1. `docker commit -c 'LABEL codeman.exported=1' codeman-case-<slug> codeman/export-<slug>:<ts>` (unique tag per export defeats the stale-image trap). Optional pre-commit scrub in sealed mode as above; also blank instance-specific committed env (`-c 'ENV CODEMAN_API_URL='` etc.) so the image carries no stale host references.
2. `docker save codeman/export-<slug>:<ts> | gzip` streamed in fixed 8192-byte chunks to `~/.codeman/docker-exports/<slug>-<ts>.image.tar.gz`. Uses `docker save` (layers + repo:tag + CMD), never `docker export` (flat rootfs), so restore is a trivial `docker load`.
3. `tar --numeric-owner -C <hostWorkspacePath> -czf <slug>-<ts>.workspace.tar.gz .` while paused (the bind-mounted workspace is NOT in the image, so it travels separately and consistently).
4. Write `manifest.json`: schema version, caseName, image tag, engine, containerWorkdir, resource/network config, codeman version, base-image digest, createdAt, per-member sha256, `mountCredentials`, and `secretFree` (true only for convenient-mode or scrubbed-sealed exports).
5. `docker rmi codeman/export-<slug>:<ts>` in the `finally` (delete the intermediate committed image regardless of success), then `docker unpause`.
The three files are wrapped in one bundle `<slug>-<ts>.codeman-container.tgz` and offered as a downloadable artifact through the existing file-routes streaming + attachment-registry handoff.
Retention / disk budget (user-decision 3): `docker-exports/` is capped at `DOCKER_EXPORT_KEEP` most-recent bundles with an auto-prune on each new export, plus the free-space precheck above. Workspace scrub: the WORKSPACE tar gets a scan/warn pass for agent-created `.env` / `.git/credentials` (a distinct leak channel from container creds). A lighter "workspace-only" export (just the workspace tar, no commit/save) is the fast default for 24h+ runs; full-image is the explicit heavier option (user-decision 7 in the original list, now decision on the default button below).
What travels: the baked toolchain image plus any in-image writes, and the workspace tar. What does NOT travel: bind-mounted credentials (physically excluded from commit) and anything that lived only in a bind mount. Secret-free by construction in convenient mode, and enforced (refuse-or-scrub) in sealed mode.
IMPORT `POST /api/docker-cases/import` (untrusted-bundle containment, the traversal/overwrite risk): stream the uploaded bundle, validate every manifest checksum BEFORE any extraction or load. Extract the workspace tar with `tar --no-absolute-names -C <fresh dir>` PLUS per-entry validation rejecting any member whose normalized path escapes the destination (leading `/` or `..` components). `gunzip | docker load` the image, then RE-TAG the loaded image id into a quarantined namespace `codeman/imported-<slug>:<ts>` and NEVER allow the load to overwrite `codeman/agent:base` or any pre-existing tag (capture the loaded id, ignore the bundle's repo:tag). Create a NEW `DockerCase` pointing at the quarantined image with THIS host's mounts/creds and the manifest's resource/network config, and recreate the container hardened (cap-drop ALL, no-new-privileges, non-root, `--pull=never`, CMD overridden to `sleep infinity`). The destination supplies its own login, so credentials never cross machines. Plus `GET /api/docker-exports` (list) and `DELETE /api/docker-exports/:filename`, all behind Codeman's existing auth / loopback-default / host-guard / Origin-CSRF stack.
## 6. Codeman integration (file-by-file, mirroring the remote-SSH feature)
- `src/types/session.ts`: add `DockerCommandMode`, `DockerEngine`, `DockerNetworkMode`, `DockerResourceLimits`, `DockerHost`, `DockerCase`, `SessionDocker` (Section 3). Add `docker?: SessionDocker` to `SessionState` after line 219. SessionMode (line 44) UNCHANGED.
- `src/mux-interface.ts`: add `docker?: SessionDocker` to `MuxSession` (38), `CreateSessionOptions` (81), `RespawnPaneOptions` (105).
- `src/docker-hosts.ts` (NEW, direct mirror of `src/remote-hosts.ts`): `readDockerHosts`/`writeDockerHosts`/`readDockerCases`/`writeDockerCases` (via `dataPath`, including `lastClaudeSessionId` read/write), `defaultDockerCommandForMode` (mirror line 60), `dockerDisplayPath` (`container:/path`, mirror `remoteDisplayPath` at 205), `toSessionDocker(host, case)` (mirror `toSessionRemote` at 212), `buildDockerBaseArgs`/`buildDockerCreateArgs` (per-engine uid/userns branch), `hostGatewayAlias(engine)`, `containerApiUrl(processApiUrl, engine)` (scheme+port-preserving, unit-tested), `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` (15s-bounded, `IS_TEST_MODE` no-op), a config-hash helper for drift, its own POSIX `shellescape` copy (mirror line 83). `const IS_TEST_MODE = !!process.env.VITEST;` gates every real `docker` invocation.
- `src/tmux-manager.ts`: add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand` (Section 4). Extend the two `fullCmd` ternaries (1276, 1524) and the two `launchCmd` cd-skips (1327, 1542). Add `killSession` Strategy 3c after 1719. Ensure `reconcileSessions` (~1800-1815) does NOT hard-delete docker sessions on local-tmux death (recovery relaunch path).
- `src/session.ts`: add `_docker?: SessionDocker` field (mirror `_remote` at 403), constructor arg (477), assignment (550). Thread `docker: this._docker` and `resumeSessionId: this._claudeSessionId` into BOTH `createSessionOptions` and `respawnPaneOptions` in `startInteractive` (1352/1370) and the second path (1740/1750). Emit `docker: this._docker` in `toState()` (1010). Replace the LOCAL cliVersion probe at 1320 for docker with the IN-CONTAINER `probeDockerCliVersion` (do not merely skip it). Extend `resolveMuxAttachCwd(workingDir, remote, docker)` (215) to return `/tmp` when `docker` is set. On claudeSessionId capture, persist it to the owning `DockerCase.lastClaudeSessionId`.
- `src/web/server.ts`: in `restoreMuxSessions` (2160), add `docker: muxSession.docker ?? savedState?.docker` to the `new Session({...})` call (2195-2216), and skip docker in the same `isExternalCliMode`/Ralph recovery guards as remote. Register the instance-scoped boot reaper to run AFTER docker-cases load and AFTER `restoreMuxSessions`. Ensure `CODEMAN_API_URL` derivation reads the SAME `process.env.CODEMAN_API_URL` the server sets at ~2000.
- `src/web/schemas.ts`: add `DockerHostSchema` and `DockerCaseLinkSchema` (below). The three mode enums (177/373/705) and `QuickStartSchema` (368) UNCHANGED (docker resolves by `caseName` lookup like remote).
- `src/web/routes/session-routes.ts`: import the docker helpers from `../../docker-hosts.js`. Add a docker branch in `/api/quick-start` parallel to the remote branch (1686-1720): `readDockerCases` -> find by `caseName` -> `readDockerHosts` -> find by `hostId`; reject `envOverrides`/`effort`/`codexConfig`/`geminiConfig`/`openCodeConfig` (but ACCEPT `modelOverride`, which flows via scaffolded `settings.local.json`); run `checkDockerAvailable` + `checkDockerTmuxAvailable` (image-present, engine, caps-enforced); surface `capsEnforced:false` and Desktop notes; set `casePath = dockerCase.hostWorkspacePath` (REAL host dir), `docker = toSessionDocker(host, dockerCase)`, and seed `resumeSessionId` from `dockerCase.lastClaudeSessionId` when `resumeOnStart`. Extend the LOCAL-availability and local-spawn guards (around 1796/1810) to `!remote && !docker`, but DO NOT extend the workspace-scaffolding guard (~1776, `writeHooksConfig`/`updateCaseModel`), which MUST run for docker. Pass `docker` into `new Session` (1847); `autoConfigureRalph` (1853) gated on `!docker`. Add `docker: m.docker !== undefined ? true : undefined` to the unified harvest (2313).
- `src/web/routes/case-routes.ts`: import the docker read/write/check helpers + schemas. Add a docker listing loop in `GET /api/cases` (mirror 94-119, `location: 'docker'`, `docker: {...}` via `dockerDisplayPath`). Add `/api/docker-hosts` GET/POST/PUT/DELETE (mirror 168-204) and `POST /api/cases/docker-link` (mirror 206-232; run `checkDockerAvailable`/`checkDockerTmuxAvailable` at link time; broadcast `CaseLinked` with `type: 'docker'`). Add a docker-unlink branch to `DELETE /api/cases/:name` (mirror 288-296; `docker rm -f`; broadcast `CaseDeleted` `type: 'docker-unlinked'`). Add the docker branch to single-case `GET` (mirror 358-368). Add `POST /api/docker-cases/:name/export`, `/import`, `GET/DELETE /api/docker-exports`, and a `POST /api/docker-cases/:name/recreate` (drift confirm) per Sections 4 and 5.
- `src/web/sse-events.ts` + `src/web/public/constants.js`: reuse `CaseLinked`/`CaseDeleted` for CRUD. Add `docker:exportProgress`, `docker:exportComplete`, `docker:importComplete`, `docker:configDrift`, and `docker:containerError` to BOTH registries (kept in sync per CLAUDE.md).
- Frontend `src/web/public/index.html` (~1831): add a Docker `modal-tab-btn` next to Remote; add a `#case-docker` panel mirroring `#case-remote` with `dockerCaseName`, `dockerHostWorkspacePath`, `dockerContainer`, `dockerImage`, `dockerHostId`, and an Advanced `<details>` for network mode, resource caps, `mountCredentials`, `resumeOnStart`, and remote daemon. Surface a "scaffolds .claude into this host dir" note (user-decision 4) and a "resource caps advisory on this engine" warning when `capsEnforced:false`.
- Frontend `src/web/public/session-ui.js`: `formatCasePickerLabel` (48) + `buildCasePickerOptions` (71-73) handle `location === 'docker'` (`name @ container`, add container/image to the search haystack); `resetCaseModalFields` (~1514) add a `dockerFields` array; `switchCaseModalTab` (1573/1580/1597) handle `'case-docker'`; `submitCaseModal` add the docker branch; new `linkDockerCase()` (mirror `linkRemoteCase` at 1689) POSTing `/api/docker-hosts` then `/api/cases/docker-link`, sending omitted optionals as `undefined` (spread `...(x ? {x} : {})`, never `null`, per the Zod `.optional()`-rejects-null gotcha); `runClaude` (520) / `runShell` (702) extend the `location === 'remote'` routing to also match `'docker'`; `runOpenCode`/`runCodex`/`runGemini` (792/846/900) make the `isRemote` checks `isRemoteOrDocker` so local status probes are skipped. In the session-options Summary tab, note that `effort` is inert for docker (rejected) while `model` IS honored via `settings.local.json`.
- Frontend `src/web/public/panels-ui.js` (425-426): add `caseItem?.docker?.path`/`container` to the case-search fields.
Schemas (`src/web/schemas.ts`), mirroring `RemoteHostSchema` (299) / `RemoteCaseLinkSchema` (351):
```ts
export const DockerHostSchema = z.object({
id: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid docker host id'),
label: z.string().min(1).max(100),
engine: z.enum(['docker', 'podman']).optional(),
image: z.string().min(1).max(512).regex(/^[a-zA-Z0-9][\w./:@-]*$/, 'Invalid image ref').regex(NO_SHELL_META),
daemonHost: z.string().max(512).regex(NO_SHELL_META, 'Invalid daemon host').optional(),
context: z.string().max(128).regex(/^[a-zA-Z0-9._-]+$/, 'Invalid context').optional(),
network: z.enum(['bridge', 'none', 'custom']).optional(),
networkName: z.string().max(128).regex(/^[a-zA-Z0-9][a-zA-Z0-9_.-]+$/).optional(),
resources: z.object({
memory: z.string().regex(/^\d+[bkmg]?$/i).optional(),
cpus: z.string().regex(/^\d+(\.\d+)?$/).optional(),
pidsLimit: z.number().int().positive().max(100000).optional(),
nofile: z.string().regex(/^\d+:\d+$/).optional(),
shmSize: z.string().regex(/^\d+[bkmg]?$/i).optional(),
}).strict().optional(),
mountCredentials: z.boolean().optional(),
hooksEnabled: z.boolean().optional(),
resumeOnStart: z.boolean().optional(),
commands: RemoteCommandOverridesSchema, // reuse the shared shape
extraCreateArgs: z.array(z.string().min(1).max(1024).regex(NO_SHELL_INJECTION).refine(noCommandSubstitution)).max(32).optional(),
extraExecArgs: z.array(z.string().min(1).max(1024).regex(NO_SHELL_INJECTION).refine(noCommandSubstitution)).max(32).optional(),
});
export const DockerCaseLinkSchema = z.object({
name: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid case name format'),
hostId: z.string().regex(/^[a-zA-Z0-9_-]+$/, 'Invalid docker host id'),
hostWorkspacePath: z.string().min(1).max(2000).regex(/^\//, 'Path must be absolute').regex(NO_SHELL_META, 'Invalid characters in workspace path'),
containerWorkdir: z.string().min(1).max(2000).regex(/^\//).regex(NO_SHELL_META).optional(),
container: z.string().min(2).max(128).regex(/^[a-zA-Z0-9][a-zA-Z0-9_.-]+$/, 'Invalid container name').optional(),
});
```
`NO_SHELL_META` (rejects `$`/backtick, schemas.ts:297) is REQUIRED on `image`, `hostWorkspacePath`, `containerWorkdir`, and `container`, because all four reach the outer `bash -c "..."` double-quote layer where `$(...)`/backtick re-expose, exactly the reason `remotePath`/`identityFile` use it. `--privileged` and any `-v /var/run/docker.sock` are structurally unrepresentable (never emitted by the builder, never accepted by the schema).
## 7. Security model
- Hardening flags on every create: `--cap-drop ALL`, `--security-opt no-new-privileges` (NOT auto-set by rootless Docker or Podman, so always explicit), the uid/userns branch of Key decision 6 (never container-root; workspace files stay host-owned and HOME stays writable via GID 0), `--pids-limit` (fork-bomb guard), `--memory` with `--memory-swap == --memory` (real OOM cap), `--ulimit nofile`, `--init`, `--pull=never`. NEVER `--privileged`, NEVER mount the docker socket into the agent container. `--storage-opt size=` is emitted ONLY after the probe confirms overlay2-on-xfs-pquota or btrfs (the AICE-class silently-ignored trap); otherwise it is omitted and the UI does not advertise a size cap. Resource caps are advertised as ENFORCED only when the probe reports `capsEnforced:true`; under non-delegated rootless they are labeled advisory (user-decision 6).
- Engine: prefer whichever the probe finds, Podman-rootless first for security (a container-root breakout lands as an unprivileged host user). Rootless bind-mount ownership uses `--userns=keep-id` (Podman) vs `--user <hostUid>:0` (Docker), so real per-engine branching lives in `buildDockerCreateArgs`. Docker Desktop takes its own uid path (Key decision 6).
- Blast radius (the combined-posture the critic asked to surface, user-decision 5): the default convenient profile mounts an arbitrary host workspace dir RW (host-owned, mirrored path) AND host `~/.claude`/`~/.codex`/`~/.gemini`/`~/.config/gcloud`/`~/.config/opencode` RW into a NETWORK-ENABLED container. Container-run agent code can therefore read/modify those host trees and reach the network simultaneously. This is still a strict improvement over today's on-host skip-permissions execution, but the user must accept the combined posture explicitly; the sealed profile plus `network:none` is the mitigation for genuinely untrusted work.
- Secret handling: creds arrive ONLY as bind-mounted files (default) or exec-time NAME-ONLY `--env` (codex/gemini keys), NEVER as create-time `-e` and NEVER as an image layer. Sealed-mode export is refuse-or-scrub (Section 5), closing the sealed-leak inversion.
- CLAUDE.md "Multi-CLI prefix discipline": the exec-time name-only env is restricted to the CLI-specific keys per mode (Claude: none with OAuth mount; Codex: `OPENAI_API_KEY`/`CODEX_API_KEY`; Gemini: `GEMINI_API_KEY`/`GOOGLE_*`), never a blanket forward. `envOverrides` is rejected for docker, so the `ALLOWED_ENV_PREFIXES` allowlist is not widened.
- hook-secret: bind-mounted read-only, referenced via `CODEMAN_HOOK_SECRET_FILE` (a path, non-secret); the secret bytes never enter env or the image. Both `host.docker.internal` and `host.containers.internal` are added to the host-guard allowlist so the in-container hook curl's Host header passes on either engine.
- Host guard / instance isolation: the in-container tmux socket (`codeman-docker`) and name (`codeman-dkr-<id8>`) deliberately FAIL a container-internal Codeman's `SAFE_MUX_NAME_PATTERN`, so a nested Codeman never adopts our session (unit-asserted). The boot reaper is instance-scoped by the `codeman.instance` label so a beta never reaps prod. Any remote-daemon (`-H`/`--context`) mode is host-root-equivalent and stays strictly behind the existing auth/loopback/host-guard/Origin-CSRF stack.
- Import containment: untrusted bundles are checksum-validated, extracted with traversal guards, and loaded into a quarantined image namespace (never overwriting the base image), then run with the same hardening.
## 8. Phased implementation (branch: `feat/docker-session-mode`)
Each phase is independently testable; per CLAUDE.md, end-to-end test in the real env before COM. All new docker IO paths carry `const IS_TEST_MODE = !!process.env.VITEST;` and no-op under it; the pure command builders are tested directly.
- Phase 0: base image + engine probe. Author `docker/agent.Dockerfile` (OpenShift arbitrary-uid HOME) and `scripts/build-agent-image.mjs` (build or pull the base image; digest recorded). Add `checkDockerAvailable`/`checkDockerTmuxAvailable`/`containerApiUrl`/`hostGatewayAlias` (IS_TEST_MODE no-op) and `GET /api/docker/status`. Test: probe stub returns available/caps/Desktop flags under VITEST; `containerApiUrl` preserves scheme+port and swaps host per engine; status route returns the envelope.
- Phase 1: types + storage + schemas. Add all types (Section 3), `src/docker-hosts.ts`, `DockerHostSchema`/`DockerCaseLinkSchema`. Test: `docker-hosts.test.ts` (round-trip incl. `lastClaudeSessionId`, display path, config-hash stability); `docker-exec-options.test.ts` (schema rejects `$`/backtick in image/workdir/container).
- Phase 2: tmux-manager builders. Add `DOCKER_TMUX_SOCKET`, `dockerTmuxSessionName`, `buildDockerLaunchCommand` (resume-aware, image-check, env-prime), `buildDockerKillCommand`; wire the two ternaries + two cd-skips + Strategy 3c; harden `reconcileSessions` against docker hard-delete. Test (pure strings): adopt-proof name fails `SAFE_MUX_NAME_PATTERN`; image-check precedes create; `new-session -A` idempotent; resume flag present only when a resume id is passed; `--pull=never` present; instance label present; escaping survives `bash -c` -> `docker exec` -> `sh -lc` -> tmux WITH a host workspace path containing spaces.
- Phase 3: session.ts + mux + recovery. Add `_docker` + `resumeSessionId` threading, in-container cliVersion probe, `resolveMuxAttachCwd`, mux-interface fields, `restoreMuxSessions` passthrough, instance-scoped reaper wiring, claudeSessionId -> `DockerCase.lastClaudeSessionId` persistence, unified flag. Test: `toState()` emits docker; a persisted docker session round-trips through mux/state; a relaunch injects the persisted resume id (mock mux); reaper only targets this instance's orphaned containers.
- Phase 4: routes + first real e2e. case-routes CRUD + listing + drift-recreate; session-routes quick-start branch (scaffolding RUNS, local-availability guards skip, model accepted, effort/config rejected). Manual e2e on a real docker host: docker-host create -> docker-link -> quick-start; confirm the pane runs `claude` in the container, files land host-owned, a Codeman restart reattaches the SAME live agent, and a `docker stop` followed by relaunch RESUMES the conversation.
- Phase 5: hooks connectivity + installation. host-gateway (per engine), derived `CODEMAN_API_URL`, hook-secret mount, `CODEMAN_SESSION_ID`/`CODEMAN_MUX` exec-env + tmux setenv, host-guard allowlist, and the scaffolding write into the real workspace. Manual e2e: trigger a permission prompt from inside the container and confirm it surfaces; verify hook payloads carry the right session id. If deferred, ship docker as explicitly hook-degraded and verify output-based idle detection through the docker-exec PTY.
- Phase 6: export/import + GC + disk safety. quiesce+pause span, free-space precheck, commit+save+gzip + workspace tar + manifest + streaming download; sealed-mode refuse-or-scrub; retention/auto-prune; import with checksum validation + traversal guard + quarantined re-tag; drift-recreate; boot reaper; `runWithConversionLimit` cap; `docker rmi` in finally. Manual e2e: export, `docker load` on a second machine (or fresh case), import, confirm toolchain + workspace restored and NO creds present; attempt a sealed full-image export and confirm it is refused-or-scrubbed; attempt a `../` bundle and confirm it is rejected.
- Phase 7: frontend. Docker tab, `linkDockerCase`, run wiring, case-picker labels, panels search, caps-advisory + scaffold-warning + effort-inert notes. Verify with Playwright (`waitUntil: 'domcontentloaded'`, 3-4s settle) that the Docker tab renders and a linked docker case appears in the picker.
- Phase 8: docs + COM. Update CLAUDE.md (a "Docker cases" Key Pattern paragraph mirroring remote-SSH, plus the new state files, routes counts, and the resume/durability model), `docs/docker-cases.md`, then COM per the standard flow.
## 9. Test plan
- Unit (pure, CI-safe, mirror `test/remote-hosts.test.ts` / `test/remote-ssh-options.test.ts`):
- `test/docker-hosts.test.ts`: storage round-trip (incl. `lastClaudeSessionId`), `dockerDisplayPath`, `defaultDockerCommandForMode`, `toSessionDocker`, `containerApiUrl` (http/https, custom port, docker vs podman gateway), config-hash stability/drift, `buildDockerCreateArgs` flag ordering (cap-drop/no-new-privileges/memory==memory-swap/instance-label/`--pull=never` present; host/privileged/socket absent; per-engine uid vs `--userns=keep-id`).
- `test/docker-exec-options.test.ts`: `buildDockerLaunchCommand`/`buildDockerKillCommand` string shape and escaping through `bash -c` -> `docker exec` -> `sh -lc` -> tmux, including a workspace path with spaces; resume flag present only with a resume id; image-presence check precedes create; `dockerTmuxSessionName` fails `SAFE_MUX_NAME_PATTERN`; schema rejects `$`/backtick in image/workdir/container/name; `linkDockerCase`-shaped bodies with omitted optionals validate (no `null` on the wire).
- Probe no-op: `checkDockerAvailable`/`checkDockerTmuxAvailable`/`probeDockerCliVersion` return canned values under VITEST and never spawn.
- Integration (route tests via `app.inject()`, docker no-op'd): `/api/docker-hosts` CRUD; `/api/cases/docker-link` dup-check + broadcast; `GET /api/cases` includes the docker case with `location: 'docker'`; `/api/quick-start` docker branch rejects `envOverrides`/`effort`/config but ACCEPTS `modelOverride`, runs the workspace-scaffolding path, and constructs a session with `docker` set + seeded resume id; `DELETE /api/cases/:name` docker-unlink; export refuse-or-scrub for sealed; import traversal rejection; reaper instance-scoping (label filter). Pick a unique port only if a live-server test is added (search `const PORT =`; 3150+).
- Manual end-to-end (real docker daemon, the mandatory "always end-to-end test" gate): build the base image; link a docker case; quick-start `claude`; verify OAuth via the mounted `~/.claude`, transcript correlation (subagent/workflow watchers show the session), host-owned files, and a working permission-prompt hook; reattach after a Codeman PROCESS restart (SAME live agent); `docker stop` then relaunch and confirm conversation RESUME; reboot-equivalent (daemon restart) and confirm boot recovery recreates+resumes; change the host's memory/image and confirm the drift-recreate prompt fires; export (convenient) and confirm the tar `docker load`s with no creds; attempt a sealed full-image export and confirm refuse-or-scrub; import into a fresh case; delete the case and confirm `docker rm -f` plus instance-scoped reaper GC; confirm a docker-down state surfaces a docker-specific error and does NOT trip the generic PTY-exit breaker.
## 10. Open decisions for the user
1. Credential + blast-radius posture (combined). Convenient default bind-mounts host `~/.claude` etc. RW AND an arbitrary host workspace RW into a network-enabled container, so container-run agent code can read/modify those host trees and reach the network at the same time. Recommended: convenient default plus a per-host SEALED opt-in (`mountCredentials:false` + `network:none`) for untrusted work. Please confirm you accept the combined arbitrary-workspace-plus-egress-plus-host-creds posture for the default profile (it is still a net improvement over today's on-host skip-permissions execution).
2. Base image ownership, registry, and freshness. The `codeman/agent:base` placeholder implies a Docker Hub org the project may not own. Pick the real registry/namespace (GHCR under the repo is the natural fit), decide digest pinning, and set a REBUILD CADENCE so agents are not stuck on a stale baked `claude` (the in-container version probe surfaces staleness, but something must trigger rebuilds). Choose: pull a pinned published image, build locally on first use via `scripts/build-agent-image.mjs`, or both.
3. Container CWD strategy. Mirror the host workspace path inside the container (recommended: makes transcript projHash correlate, file features and resume capture work) vs a fixed `/workspace` (simpler mount, breaks watcher correlation). Please confirm the mirror approach.
4. Hooks in the MVP AND workspace scaffolding. Making docker hooks fire requires WRITING `.claude/settings.local.json` (and the CLAUDE.md scaffold) into the user's REAL linked host directory, a behavioral shift from "link a dir" to "link and scaffold a dir." Choose: wire hooks + scaffolding now (Phase 5, recommended, and it also enables the model picker), or ship docker as explicitly hook-degraded (no permission prompts / hook-idle) for v1 and add later. Confirm you are OK with Codeman mutating the linked host workspace.
5. Session-kill teardown and RESUME (reframed honestly). `docker stop` on session kill is not merely "free RAM vs instant reattach": it destroys the in-container live agent, and the conversation survives ONLY because the next launch runs `--resume` from the bind-mounted transcript. Choose: keep the container running (costs RAM, preserves the exact live in-flight agent) vs stop and rely on `--resume` (frees RAM, may lose uncommitted in-flight tool state). Case-delete always `docker rm -f`.
6. Rootless enforcement posture. Under rootless without cgroup-v2 systemd delegation, `--memory`/`--cpus`/`--pids-limit` are SILENTLY ignored. Choose: REQUIRE delegation (refuse to link a host that cannot enforce caps) or ship-with-warning ("resource caps are advisory on your engine"). The probe reports `capsEnforced` either way.
7. Default resume behavior. Should a re-linked or re-run docker case default to resuming its last conversation (`resumeOnStart:true`, using `DockerCase.lastClaudeSessionId`) rather than starting clean? This is the crux of making the durability story real and is the recommended default, but it changes user-visible behavior (a new session in an existing case continues the prior conversation).
8. Export defaults and disk budget. Default export button: workspace-only (fast, small, files-only, recommended for 24h+ runs) vs full-image (reproducible env, multi-GB). Also set the retention cap (max retained exports), the auto-prune policy, and the free-space threshold below which export is refused (a full `/var/lib/docker` breaks EVERY session on the host, not just docker ones).
9. Remote docker daemon (`-H ssh://...` / `--context`). Support in the MVP (composes with remote hosts, adds host-root trust surface) or local-daemon-only first.
10. Podman parity depth. Full `--userns=keep-id` plus Quadlet boot-persistence, or Docker-first with Podman as best-effort and boot-persistence via Codeman's idempotent create-if-missing only. Note the podman host alias is `host.containers.internal`, already handled per engine.
+94
View File
@@ -0,0 +1,94 @@
# Docker cases
Run a case inside an **isolated Docker container** instead of directly on the host. Any number of Codeman sessions can share one container (it is scoped to the case, not the session), so a whole project lives in a sandbox with its own network, resource caps, and filesystem, and you can **export the container to move it to another machine**.
Docker mode is a **location overlay on cases**, the direct analog of [remote SSH cases](./remote-hosts.md): where a remote case runs a local tmux pane doing `ssh host` into a durable remote tmux server, a docker case runs a local tmux pane doing `docker exec -it` into a durable **in-container** tmux server. It is not a separate `SessionMode`, so `claude` / `shell` / `opencode` / `codex` / `gemini` all work inside the container.
## One-time setup: build the base image
The container needs a base image with the agent toolchain (node, the CLIs, git, tmux). Build it locally once:
```bash
node scripts/build-agent-image.mjs # builds codeman/agent:base
# options: --engine docker|podman --image <ref> --no-cache
```
The image is **secret-free**: credentials are delivered at runtime (bind mounts or `docker exec --env`), never baked in, so exports never leak them.
## Quickest path: one-click "Run in Docker"
On the **New case → Create New** tab there's a **🐳 Run in an isolated Docker container** checkbox. Checking it alone is enough: Codeman creates the case folder in `~/codeman-cases/<name>`, spins up a hardened container with sensible defaults (auto-provisioning a shared `default` host), and starts the session inside it. No host/image/network fields to fill in.
Click the checkbox's **Container settings** to optionally tweak the predefined defaults, including a **Template** picker:
| Template | Memory | CPUs | GPUs |
|----------|--------|------|------|
| Small | 2 GB | 1 | none |
| Medium (default) | 4 GB | 2 | none |
| Large | 8 GB | 4 | none |
| GPU | 8 GB | 4 | all (needs the NVIDIA container toolkit) |
**Disk is elastic** — the container's storage grows automatically as data flows in; there is no fixed cap (bounded only by host disk). Any tweaked setting creates a dedicated per-case host so it never changes the shared `default`.
## Create a docker case (full control)
App → **New case → Docker** tab:
- **Case Name** / **Workspace Path**: the workspace is a real HOST directory bind-mounted into the container at the same path. Codeman scaffolds `CLAUDE.md` + `.claude/settings.local.json` (hooks) into it, and file previews / attachments work on the real bytes.
- **Host ID**: a reusable docker host profile (image, network, resources). Reuse the same ID across cases to share settings.
- **Network**: `bridge` (internet on, default), `none` (fully isolated), or a `custom` bridge.
- **Advanced**: memory / CPU caps, **Mount host credentials** (on = your existing `~/.claude` login just works; off = a sealed sandbox you log into inside the container), **Resume last conversation on relaunch**.
Then run it like any case (Run Claude / Run Shell / …). The first launch creates the container (`codeman-case-<name>`); subsequent sessions attach to the same one.
Equivalent API:
```bash
curl -X POST localhost:3000/api/docker-hosts -d '{"id":"local","label":"Local","image":"codeman/agent:base"}'
curl -X POST localhost:3000/api/cases/docker-link -d '{"name":"sandbox","hostId":"local","hostWorkspacePath":"/home/you/projects/sandbox"}'
curl -X POST localhost:3000/api/quick-start -d '{"caseName":"sandbox","mode":"claude"}'
```
## Lifecycle
- **Reconnect after a Codeman restart** lands back in the same live agent (the in-container tmux survives).
- **Container stop / host reboot** recreates the container and, when a resume id was captured, **resumes** the last conversation from the bind-mounted transcript.
- **Killing one session** only kills that session's in-container tmux session; the shared container stays up for sibling sessions.
- **Deleting the case** `docker rm -f`s the container (the bind-mounted workspace on the host survives). An instance-scoped boot reaper removes containers whose case is gone.
## Isolation & security
Every container runs hardened: `--cap-drop ALL`, `--security-opt no-new-privileges`, non-root (`--user <hostUid>:0` so workspace files stay host-owned), `--pids-limit`, `--memory` == `--memory-swap`, `--init`. Never `--privileged`, never the docker socket. The default **convenient** profile bind-mounts host credential dirs read-write so the common login just works (creds stay on the host, never captured by `docker commit`); the **sealed** profile (`mountCredentials:false` + `network:none`) is the opt-in for genuinely untrusted work.
Rootless engines without cgroup-v2 systemd delegation cannot enforce resource caps; linking such a host warns that caps are advisory.
## Export / Import (move to another machine)
**Export** (from the Docker tab, or `POST /api/docker-cases/:name/export`): choose
- **Full image + workspace**: `docker commit` the container to an image, `docker save` it, tar the workspace, and a manifest, all into one portable `<case>-<ts>.codeman-container.tgz` (the whole toolchain, installed packages, and files). Runs in the background; you are notified when the bundle is ready.
- **Workspace only**: just the project files (fast, small).
The container is paused across the capture so the image and workspace are consistent; a full `/var/lib/docker` is guarded against with a free-space precheck; the intermediate image is always cleaned up.
**Import** (`POST /api/docker-cases/import`, or the Manage tab): copy the `.tgz` onto the new machine's `~/.codeman/docker-exports/`, then import it into a new case. The manifest and per-member SHA-256 checksums are validated, the workspace tar is extracted with a path-traversal guard, and the image is `docker load`ed and **re-tagged into a quarantined namespace** (`codeman/imported-<case>:<ts>`) so it never overwrites a local tag. The destination supplies its own credentials, so nothing secret crosses machines.
`GET /api/docker-exports` lists bundles; `GET /api/docker-exports/:filename` downloads one; `DELETE` removes one.
## Hooks require the server to be reachable from the container
In-container hooks (permission events, hook-based idle/stop/task notifications) POST to `CODEMAN_API_URL`, which is derived as `https://host.docker.internal:<port>` (`host.docker.internal` → the docker bridge gateway, e.g. `172.17.0.1`, via `--add-host …:host-gateway`). For that callback to succeed, the Codeman server must be **listening on an interface the container can reach**.
- If Codeman binds **loopback-only** (`127.0.0.1`, the default and the production systemd config), a container reaching `172.17.0.1:<port>` cannot connect, so by default **in-container hooks do not fire**. The session still works fully: idle/stop detection falls back to **output-based** detection through the `docker exec` PTY (which always works), and claude runs with `--dangerously-skip-permissions` so there are no permission prompts to forward anyway.
- **To enable in-container hooks on a loopback-only server, set `CODEMAN_DOCKER_BRIDGE_HOOKS=1`** (env). Codeman then starts a SECOND listener bound to the docker bridge gateway (`172.17.0.1`, auto-detected; override with `CODEMAN_DOCKER_BRIDGE_HOST`) that serves **only the hook endpoints** (`/api/hook-event`, `/api/status-telemetry`) and delegates them into the same secret-gated pipeline. The bridge is host-internal (containers + host, not the LAN), and every other path returns `403`, so this does not widen your network exposure. Add `Environment=CODEMAN_DOCKER_BRIDGE_HOOKS=1` to the systemd unit and restart.
- Alternatively, bind `0.0.0.0` **with `CODEMAN_PASSWORD` set** (exposes on the LAN too).
The host-gateway mapping, `CODEMAN_API_URL` derivation, host-guard allowlist, and hook-secret mount are all wired correctly; `CODEMAN_DOCKER_BRIDGE_HOOKS` closes the last gap for loopback-only servers.
## Notes & limits
- Requires Docker (or Podman) with a reachable daemon; tmux must be present in the base image (a hard prerequisite, probed at link time).
- Per-session `envOverrides` / `effort` / per-CLI config are rejected for docker cases (they do not cross into the container); configure the container via the docker host's per-mode command override instead.
- macOS Docker Desktop takes a dedicated uid path (the baked image uid; memory caps are subject to the VM ceiling).
Design + rationale: [`docker-cases-plan.md`](./docker-cases-plan.md).
+72
View File
@@ -0,0 +1,72 @@
# Reliable input delivery (exactly-once, durable)
## The bug this fixes
With local echo on, pressing Enter cleared the overlay and then sent the prompt
over the WebSocket **fire-and-forget** (`ws.send({t:'i',d})`). On a flaky link
(e.g. a moving train) the socket is frequently *half-open*: `readyState === OPEN`
so `ws.send()` does **not** throw, but the underlying TCP is dead, so the frame is
silently discarded. Nothing was enqueued (the send "succeeded"), the on-screen
prompt was already wiped, and `navigator.onLine` stays `true` — so a long typed
prompt vanished with no trace and no resend.
## The guarantee
Every byte of user input is **recorded durably before delivery** and **only
dropped once the server ACKs it** — so a half-open socket, a reconnect, or a page
reload can never lose input. Redelivery is **exactly-once**: the server applies
each `(clientId, seq)` at most once, so a resend can't type the prompt twice.
## How it works
### Client (`app.js`)
- A stable **`clientId`** (`localStorage['codeman:clientId']`) identifies this
browser to the server's dedup across reconnects and reloads.
- Each input frame gets a **monotonic per-session `seq`**. Frame records
(`{seq,data,useMux,ts,tries,sentAt}`) live in `_pendingDeliveries`
(`Map<sessionId, record[]>`), persisted (debounced, + flushed on `pagehide`/
`visibilitychange`) to `localStorage['codeman:pendingInput']`. The seq counters
persist too, so seqs stay monotonic across reloads (never reset — a reset would
let the server treat fresh input as an already-applied duplicate).
- **Delivery** (`_drainSession`):
- **WS path** — when the socket is `OPEN` for the session, send each not-yet-sent
record (`sentAt === 0`) in seq order over the single ordered stream. Records
stay pending until the server's `{t:'ia',seq}` ACK removes them.
- **POST path** — when no WS, POST records in order, awaiting each (the HTTP 2xx
*is* the ACK). A 404/410 (session gone) drops the record rather than retry
forever.
- **Half-open recovery** (`_redeliverSweep`, every 2s): if the active WS session's
oldest record is unacked past `_reliableAckTimeoutMs` (4s), the socket is assumed
dead — `ws.close()` forces a fast reconnect; `onopen` (`_onWsReady`) resets
`sentAt = 0` and re-sends everything pending. Also re-drains background sessions
over POST, and fires on SSE-reconnect / `online`.
- The connection indicator shows pending count/bytes (`_pendingBytes`).
### Server
- **`Session.shouldApplyInput(clientId, seq)`** — returns `true` exactly once per
`(clientId, seq)`: the first time a seq strictly greater than that client's
last-applied is seen. A replayed/lower seq returns `false`. Bounded MRU map
(`MAX_INPUT_DEDUP_CLIENTS = 256`).
- **WS route** (`ws-routes.ts`) — parses optional `cid`/`seq` on `{t:'i'}`; applies
via `shouldApplyInput` (skips a duplicate, still ACKs with `{t:'ia',seq}` so the
client drops it). Untagged frames apply unconditionally (no behavior change).
- **POST route** (`/api/sessions/:id/input`) — optional `seq`/`clientId` in
`SessionInputWithLimitSchema`; a deduped duplicate returns 200 without writing
(the 200 is the client's ACK). `curl`/legacy callers omit the fields and always
apply.
## Known limitation
Dedup state is in-memory on the server. A **server restart** between a write and
the client's redelivery of that same seq could re-apply it (a rare duplicate).
This is a deliberate trade-off: favor *never losing input* over a rare duplicate
across the narrow restart window.
## Tests
- `test/reliable-input-dedup.test.ts` — `Session.shouldApplyInput` exactly-once
semantics (monotonic, per-client, gap-tolerant, eviction-safe).
- `test/routes/session-routes.test.ts` — POST `/input` applies a tagged
`(clientId, seq)` once on redelivery; untagged input always applies.
+73 -9
View File
@@ -30,7 +30,8 @@ an explicit, guided opt‑in.
7. [Supply‑chain & build‑asset hardening](#7-supplychain--buildasset-hardening-cod28)
8. [Multi‑instance isolation](#8-multiinstance-isolation)
9. [Transport security headers](#9-transport-security-headers)
10. [Quick reference](#10-quick-reference)
10. [Docker container isolation](#10-docker-container-isolation)
11. [Quick reference](#11-quick-reference)
---
@@ -124,7 +125,11 @@ loopback bind matters. The auth pipeline (`src/web/middleware/auth.ts`,
`onRequest` hook) runs in this order:
1. **Localhost‑only exemptions** (always first): `POST /api/hook-event` and the QR
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3).
`/q/` short‑code path are exempt when `req.ip` is loopback (see §3). While the
**managed tunnel is running**, the hook‑event exemption additionally requires
the per‑instance `X-Codeman-Hook-Secret` header (COD‑54); failed presentations
are rate‑limited in a **dedicated bucket** (separate from Basic‑Auth failures)
so misfiring hooks can never lock out the login path.
2. **Session cookie** check — a valid `codeman_session` cookie short‑circuits to
allow.
3. **HTTP Basic** check — correct credentials short‑circuit to allow and clear
@@ -165,17 +170,29 @@ protection is unchanged.
with `req.ip = 127.0.0.1`**. The localhost‑only exemptions then treat those
requests as local:
- `POST /api/hook-event` — auth‑exempt for loopback. Bounded impact: it is
- `POST /api/hook-event` — auth‑exempt for loopback **only while no managed tunnel
is running**. When Codeman's own tunnel is up, the exemption requires the
per‑instance shared secret (`X-Codeman-Hook-Secret`, 256‑bit hex in
`~/.codeman/hook-secret`, mode 0600, COD‑54). Local hook commands read the
secret file at execution time (`$CODEMAN_HOOK_SECRET_FILE`, exported into every
managed session), so they keep working — tunneled internet traffic can't know
it. Even without the secret the impact is bounded: the route is
`HookEventSchema`‑validated and requires a valid in‑memory `sessionId`; it can
drive respawn signals, SSE broadcasts, push notifications, and transcript
watching — **not** arbitrary terminal input or file reads. It is a
session‑disruption / notification‑spoofing surface, not RCE.
watching — **not** arbitrary terminal input or file reads. ⚠️ The gate keys off
the **managed** tunnel — an externally run loopback proxy (your own
`cloudflared`, `tailscale serve`) is invisible to it, so the plain loopback
exemption still applies there (prefer `tailscale serve`, which authenticates at
the tailnet layer). Hook configs regenerated since COD‑54 always present the
header, so a future release can require the secret unconditionally.
- QR `/q/` — still protected by its own short‑code brute‑force limiter
(10 failures / 60s against a 62⁶ space).
**Mitigation:** set `CODEMAN_PASSWORD` whenever a loopback‑connecting tunnel is
up (it does not gate the hook‑event exemption, but it gates everything else and
is the documented practice). Prefer `tailscale serve` (below), which authenticates
up — it gates everything except the (secret‑gated) hook exemption and is the
documented practice; since COD‑55 enabling the managed tunnel **refuses** to start
without it unless `CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK=1` explicitly
acknowledges the exposure. Prefer `tailscale serve` (below), which authenticates
at the tailnet layer so untrusted clients never reach the loopback port at all.
### Host‑header & Origin allowlist (DNS‑rebinding & CSRF defense)
@@ -304,7 +321,37 @@ injected from API JSON (`innerHTML`), not via `file-raw`, so they are unaffected
`/api/download` additionally refuses a blocklist of sensitive paths
(`/etc/shadow`, `~/.ssh/`, `.env`, `*credentials*`, `.aws/credentials`, …). This
is **defense‑in‑depth, not the primary boundary** — the realpath containment is
the control.
the control. The blocklist patterns are shared (`src/web/sensitive-path.ts`) with
the attachment guard below.
### External attachments (registry) & the magic‑link trust boundary
Live external attachments (`src/attachment-registry.ts`) mint an `att_<uuid>` id
for a host file so browser requests carry the id, never an absolute path. Serving
is by id (`GET /api/sessions/:id/attachments/:attachmentId/raw`, 50 MB cap,
`nosniff`) and re‑resolves the symlink + re‑checks the **attachment guard**
(`src/config/attachment-guard.ts`: the shared sensitive‑path blocklist **plus**
the `/root` and `/etc` trees, extendable via `attachmentBlockedPaths` /
`CODEMAN_ATTACHMENT_BLOCKED_PATHS`) on every request. Unlike the workspace file
routes, attachments are intentionally **cross‑workspace** — so the effective gate
is the blocklist + a 6‑extension allowlist (`png/pdf/docx/pptx/md/txt`), not
realpath containment.
Two registration paths, with **different trust**:
- **Explicit `POST /api/sessions/:id/attachments`** (and `codeman attach`, which
POSTs directly inside a managed session) — a deliberate, Origin‑guarded HTTP
request. Allowed cross‑workspace (subject to the guard). This is the supported
path for codeman‑publish and the `~/.codeman` review‑card loop.
- **Terminal `codeman://attach?path=…` magic links** — scanned passively from
session output. Terminal output is **attacker‑influenceable** (a prompt‑injected
session can print an arbitrary path), and registration here is server‑side with
no Origin gate and broadcasts the `rawUrl` over SSE to all clients. This path is
therefore **force‑confined to the session workspace** (`forceWorkspaceConfinement`
in `registerExternalAttachment`, wired in `WebServer.registerAttachment`),
regardless of the global confine setting — a passive magic link cannot expose a
file outside the session's own workspace. Cross‑workspace attach must go through
the explicit POST path above.
### SSE log‑tail route — intentional extra read roots
@@ -425,7 +472,22 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
---
## 10. Quick reference
## 10. Docker container isolation
Docker cases (1.4.0) run a session inside a per‑case container instead of on the host. The security posture:
- **Hardened create flags, always** — `--cap-drop ALL`, `--security-opt no-new-privileges`, `--pids-limit` (fork‑bomb guard), `--memory` == `--memory-swap` (a real OOM cap), `--init`, and non‑root: `--user <hostUid>:0` on Linux (host uid → workspace files stay host‑owned; GID 0 keeps `$HOME` writable), `--userns=keep-id` on rootless Podman. **Never** `--privileged`, and **never** the docker socket — the pure builder in `docker-hosts.ts` cannot emit them and the schema cannot represent them.
- **Credentials never enter an image** — the convenient default bind‑mounts host cred dirs (`~/.claude`, `~/.codex`, `~/.gemini`, `~/.config/{gcloud,opencode}`) read‑write. Bind mounts are physically excluded from `docker commit`, so exported images are secret‑free. API‑key CLIs get their key as an exec‑time NAME‑ONLY `--env OPENAI_API_KEY` (no `=value`, no `ps` leak, never committed); a create‑time `-e` for a secret is never used. The **sealed** profile (`mountCredentials:false` + `network:none`) drops the host mounts; full‑image export is then refused (an in‑container login would ride the committed layer) unless a pre‑commit scrub is opted into.
- **Blast radius — accept it explicitly** — the convenient profile mounts an arbitrary host workspace RW plus the host credential dirs RW into a network‑enabled container, so container‑run agent code can read/modify those host trees and reach the network at once. Still a net improvement over today's on‑host `--dangerously-skip-permissions` execution; use the sealed profile for genuinely untrusted work.
- **Import is untrusted‑bundle‑safe** — `/api/docker-cases/import` validates the manifest + per‑member SHA‑256 before extraction, rejects absolute / `..` tar members (traversal guard), and re‑tags the loaded image into a quarantined namespace so it can never overwrite `codeman/agent:base` or a pre‑existing tag.
- **Host guard & the bridge‑hooks listener** — in‑container hook callbacks carry `Host: host.docker.internal` / `host.containers.internal`; both are on the always‑on host‑header allowlist (`DOCKER_HOST_GATEWAY_ALIASES`) and resolve to the host only from inside a container netns, so they are not a browser DNS‑rebinding surface. On a loopback‑only server, in‑container hooks are opt‑in via `CODEMAN_DOCKER_BRIDGE_HOOKS=1`, which binds a SECOND listener on the docker bridge gateway serving **only** the hook endpoints (every other path → `403`) into the same hook‑secret‑gated pipeline. The bridge is host‑internal (containers + host), not the LAN, so it does not widen network exposure; the hook secret is bind‑mounted read‑only and referenced by path.
- **Instance isolation** — every managed container is labeled `codeman.instance=<CODEMAN_INSTANCE>`; the boot reaper reaps orphans of its OWN instance only, so a beta never removes a prod container. The in‑container tmux socket (`-L codeman-docker`) + session name (`codeman-dkr-*`) deliberately fail a nested Codeman's discovery pattern.
Full feature guide: [`docker-cases.md`](docker-cases.md).
---
## 11. Quick reference
| Env / flag | Effect |
|------------|--------|
@@ -436,6 +498,8 @@ production layout (`~/.codeman`, `-L codeman`, port 3000).
| `--https` | Enable TLS (adds HSTS) |
| `CODEMAN_INSTANCE` | Scope tmux socket + data dir for isolation |
| `CODEMAN_GESTURE=1` | Make the gesture overlay available (widens CSP) |
| `CODEMAN_DOCKER_BRIDGE_HOOKS=1` | Serve the hook endpoints on the docker bridge gateway (host‑internal, hooks‑only, `403` elsewhere) so in‑container hooks reach a loopback‑bound server — see §10 |
| `CODEMAN_DOCKER_BRIDGE_HOST` | Override the bridge gateway IP the hooks listener binds (default: auto‑detect) |
**Audit log:** session lifecycle and server start are recorded in
`~/.codeman/session-lifecycle.jsonl`.
+271
View File
@@ -0,0 +1,271 @@
# Ultracode / Workflow Agent Visualization — Design & Implementation Plan
> **Status: IMPLEMENTED (2026-06-15, rev. 3) — Phases 1–3 shipped & verified; Phase 4 (live-transcript link) deferred.** A dedicated, opt-in **master-detail tab** (`showUltracodeAgents`, default OFF) shows ultracode/Workflow runs as Claude Code's "working agents" TUI: LEFT = runs + phases (selectable tasks), RIGHT = each run's agents with model, live state, **tokens burned**, and **tool calls**.
>
> ### What rev. 3 changed vs. rev. 2 (decided during implementation against on-disk truth)
> 1. **UI is a master-detail TAB, not grouped floating subagent windows.** The user asked for the CC "working agents" view (left task picker, right agent stats). Built as a new docked panel `#ultracodeAgentsPanel` (clones `.subagents-panel` master-detail CSS) + `src/web/public/ultracode-panel.js` — NOT via `openSubagentWindow`/grouped windows.
> 2. **STANDALONE — zero edits to `subagent-watcher.ts`.** w16-claudeman's commit `f6a30d7` already discovers the per-agent workflow *transcripts* (`watchWorkflowDirs`). The data the view needs (run/phase/per-agent tokens+toolCalls) lives in the *run-state* JSON, read by a brand-new `src/workflow-run-watcher.ts` (globs the disjoint `…/workflows/wf_*.json` tree). No shared files with w16.
> 3. **No per-agent transcript streaming needed for v1.** The run-state JSON already carries `tokens`/`toolCalls`/`state`/`label`/`phase` per agent, so the whole view reads from `wf_<runId>.json` alone. (Phase 4 will optionally link a card to its already-tracked transcript via `agentId` — no watcher edits.)
> 4. **Agent states are `start | progress | done`** (verified on disk) — NOT running/queued. `start`=queued (no agentId/tokens/toolCalls yet), `done` has `durationMs`/`resultPreview`.
> 5. **The run JSON's `script` (15–660KB embedded JS), `scriptPath`, `result`, `logs` are STRIPPED in the watcher** before caching/broadcast (a 28-agent run drops 174KB → ~25KB; `promptPreview`/`resultPreview` truncated).
> 6. **SSE/snapshot ship lightweight run SUMMARIES (no `agents[]`); the RIGHT pane fetches the full run** via `GET /api/workflows/:runId` on selection. (A 25-run snapshot is ~20KB vs ~900KB if it carried every agent.) The LEFT list shows ALL cached runs (LRU-bounded), not a recency window — a run browser must show past runs.
>
> _Original rev. 2 proposal (grouped floating windows, extending subagent-watcher) preserved below for context; superseded by the above._
### What changed in rev. 2 (vs. the first draft)
1. **No backend cross-watcher coupling.** The per-agent label/phase/agentType/state **join moves to the frontend at render time** — the run object already carries every agent's entry keyed by `agentId`. This deletes `subagent-watcher`'s backward dependency on `workflow-run-watcher` (`getAgentLabel()` + its TTL cache), removes the registration-vs-run-state **race** (labels always track the latest `workflow:run_updated`), and drops the per-agent `meta.json` read from the hot path.
2. **`SubagentInfo` grows by 2 fields, not 4** (`isWorkflowAgent`, `workflowRunId`) — both derivable from the file path alone at registration, zero extra I/O. `agentType`/`label`/`phase`/`state` come from the run object on the frontend.
3. **The `isInternalAgent` bypass covers BOTH drop sites** — `registerAgentFile` *and* the late re-resolution in `processEntry`. The first draft named only one.
4. **De-duplicated.** Each trap (`journal.jsonl`, the `projects/*/*/workflows` depth, the gate-mismatch lesson, reuse-not-rebuild) is stated once in its owning section.
### Code-reuse verified against the tree (2026-06-14)
Confirmed present and shaped as assumed: `subagent-watcher.ts` — `watchSubagentDir`/`registerAgentFile`/`tailFile`/`processEntry`, `getRecentSubagents`, `isInternalAgent` (drops on `MIN_DESCRIPTION_LENGTH=5`), `STARTUP_MAX_FILE_AGE_MS=4h`, `MAX_TRACKED_AGENTS`, `knownSubagentDirs`/`dirWatchers`. `team-watcher.ts` — `configMtimes` mtime-skip + chokidar + `setInterval` poll. `server.ts` — `setupSubagentWatcherListeners`, `getLightState()` (`subagents: getRecentSubagents(15)`, `LIGHT_STATE_CACHE_TTL_MS=1000`), `isSubagentTrackingEnabled()` (`settings.subagentTrackingEnabled ?? true`). Frontend — `_SSE_HANDLER_MAP`, `this.subagents` Map, `handleInit`/`cleanupAllFloatingWindows`, `renderSubagentPanel`/`_renderSubagentPanelImmediate`, `getTeammateBadgeHtml`, `openSubagentWindow` + `.subagent-window-parent` sub-header.
## 1. The enabling fact: on-disk artifacts
The Workflow tool (what `ultracode` drives) persists each workflow agent as a transcript under the **same `subagents/` directory Codeman already watches**, one level deeper. Empirically verified against a real run (`wf_a8e09f2c-550`); **re-confirm the shape against a fresh run at implementation time** (§8 mandates a live e2e pass anyway):
```
~/.claude/projects/<projHash>/<sessionUuid>/
├─ subagents/
│ ├─ agent-XX.jsonl ← regular Task subagent (tracked today)
│ └─ workflows/wf_<runId>/
│ ├─ agent-YY.jsonl ← WORKFLOW agent — IDENTICAL line format
│ ├─ agent-YY.meta.json ← {"agentType":"workflow-subagent"} (optional enrichment)
│ └─ journal.jsonl ← run journal {type:"started",...} — MUST be skipped
└─ workflows/wf_<runId>.json ← run state: runId, workflowName, summary, status,
phases[], workflowProgress[], totals (DIFFERENT tree)
```
The per-agent `.jsonl` line shape is identical to a regular subagent transcript:
```jsonc
{ "parentUuid": null, "isSidechain": true, "agentId": "ac6a1d27012a64e38",
"type": "user" | "assistant", "message": { "role": "...", "content": "..." }, ... }
```
Because the line shape is identical, the entire existing parse→event→render pipeline works unchanged once discovery reaches those files. The only new data is the **run-level metadata** in `workflows/wf_<runId>.json` (name, summary, phases, and `workflowProgress[]` — the per-agent labels/state/tools), which supplies the group header and per-agent labels.
**Can show:** per-agent live transcript (tool calls, messages, results); per-agent status (active/idle/completed via the existing mtime/PID/pgrep liveness); per-agent model + running token totals (from each agent's JSONL `message.usage`, exactly as today); the run's `workflowName`/`summary`/`phases[]`; per-agent `label`/`phaseTitle`/`state`/`lastToolName` (from `workflowProgress[]`); grouping under `wf_<runId>`.
**Cannot show:** anything absent from the artifacts — a live phase cursor beyond `workflowProgress[].state`; an authoritative **budget/cost ceiling** (only consumed totals exist — `usage` + run-state `totalTokens`, no remaining-budget field); runs older than `STARTUP_MAX_FILE_AGE_MS` (4h) after a server restart (live monitoring only).
## 2. Architecture
**Decision: EXTEND `subagent-watcher.ts` for per-agent discovery/streaming; ADD a thin `workflow-run-watcher.ts` (modeled on `team-watcher.ts`) for the group-header metadata ONLY. The agent→run-metadata join happens on the FRONTEND, so the two watchers stay decoupled.**
- The per-agent JSONL is identical in shape, so re-running it through `registerAgentFile()` → `tailFile()` → `processEntry()` and the existing `subagent:*` events is free and reconnect-safe (those agents land in `agentInfo`, replayed by `getRecentSubagents(15)`). A parallel per-agent watcher would duplicate the liveness/token/tool-call/SSE machinery for zero benefit.
- Run metadata lives in a *different* file under a *different* tree (`workflows/wf_<runId>.json`, sibling to `subagents/`). A small `WorkflowRunWatcher` watching `projects/*/*/workflows/wf_*.json` (mtime-skip, like `team-watcher`'s `configMtimes`) is the clean home; folding it into `subagent-watcher` would entangle two unrelated watch roots and put a JSON re-read in the hot per-line path.
- **The two watchers never call each other.** The frontend receives both streams and joins agent→label by `agentId` at render time (the run object carries every agent's entry). This removes the timing coupling entirely.
```
~/.claude/projects/<projHash>/<sessionUuid>/
├─ subagents/
│ ├─ agent-XX.jsonl ──────────────► SubagentWatcher (EXTENDED: also descends
│ └─ workflows/wf_<runId>/ workflows/wf_<runId>/, tags isWorkflowAgent+runId)
│ ├─ agent-YY.jsonl ─┐ reuse registerAgentFile/tailFile/processEntry
│ └─ journal.jsonl (SKIP) emits subagent:* (now w/ 2 workflow fields)
└─ workflows/wf_<runId>.json ──────► WorkflowRunWatcher (NEW, team-watcher-shaped)
{workflowName,phases,workflowProgress[]} emits workflow:run_discovered|updated|removed
server.ts
setupSubagentWatcherListeners() ──► broadcast(subagent:*) ─┐
setupWorkflowRunWatcherListeners() ──► broadcast(workflow:run_*) │ SSE
getLightState(): subagents + workflowRuns ───────────────────────┘
│
▼ app.js dispatch table
panels-ui: partition this.subagents by workflowRunId; header + per-agent
labels JOINED from this.workflowRuns.get(runId).agents (by agentId)
```
## 3. Backend changes (ordered, file-by-file)
### 3a. `src/subagent-watcher.ts` — nested discovery + 2 tag fields
**(1) Extend `SubagentInfo` with exactly two optional fields** (optional → regular subagents and the wire shape are unaffected):
```ts
isWorkflowAgent?: boolean; // true when discovered under subagents/workflows/<wf_runId>/
workflowRunId?: string; // e.g. "wf_23dbeab2-152" (parent dir name)
```
Both are derived from the **file path alone** at registration — no extra reads. They ride existing `subagent:discovered|updated|completed` payloads (no new per-agent event). Do **not** add `agentType`/`label`/`phase`/`workflowName` here — those come from the run object on the frontend (§4c).
**(2) Constant.** `const WORKFLOWS_SUBDIR = 'workflows';` near the existing dir constants.
**(3) `watchSubagentDir()` — descend into `workflows/<wf_runId>/`.** After the existing direct-child registration loop:
```ts
// Workflow agents live one level deeper: subagents/workflows/<wf_runId>/agent-*.jsonl
const wfRoot = join(dir, WORKFLOWS_SUBDIR);
try {
for (const runId of await readdir(wfRoot)) {
if (!runId.startsWith('wf_')) continue;
await this.watchWorkflowRunDir(join(wfRoot, runId), projectHash, sessionId, runId);
}
} catch { /* no workflows subdir — normal for most sessions */ }
```
The existing `fs.watch(dir, …)` on `subagents/` is **non-recursive on Linux** and won't fire for writes inside `workflows/<runId>/`, so each run dir needs its own watcher.
**(4) New private `watchWorkflowRunDir(runDir, projectHash, sessionId, runId)`** — clone `watchSubagentDir`'s structure, but:
- Register only files matching `^agent-.*\.jsonl$`, **explicitly skipping `journal.jsonl`** (it ends in `.jsonl` but is `{type:'started',…}`, not a transcript — registering it would create a phantom agent).
- Call `registerAgentFile(filePath, projectHash, sessionId, isInitialScan, runId)` so the agent is tagged.
- Install one `watch(runDir, …)` per run dir; on `error` and `stop()`, reuse the existing teardown (close + delete from `dirWatchers`/`knownSubagentDirs`/`dirWatcherErrorHandlers`).
- Guard re-registration **per run dir** in `knownSubagentDirs`, **not** `wfRoot` — the 5s full scan must still re-`readdir(wfRoot)` to pick up *new* `wf_<runId>` dirs created mid-session.
**(5) `registerAgentFile()` — accept + apply `runId`.** Add a trailing optional `runId?: string`. When set, the whole change is:
```ts
if (runId) { info.isWorkflowAgent = true; info.workflowRunId = runId; }
```
No `meta.json` read, no run-state lookup, no description override. `agentId`s are globally unique `a<16hex>` (verified: 0 collisions across a 370-agent corpus), so keep the flat `agentInfo` map keyed by `agentId` — do **not** switch to a composite key. Add a one-line dev-assert log if `agentInfo.has(agentId)` with a *different* `workflowRunId`, so a future collision is observable.
**(6) `isInternalAgent` bypass — BOTH drop sites.** Workflow agents have no Task-tool spawn record, so `_resolveDescription` yields only the first-user-message fallback (often a long phase prompt) or empty → `isInternalAgent` (`length < MIN_DESCRIPTION_LENGTH`) would wrongly drop them. They are real by construction (the `subagents/workflows/wf_*/` path is the discriminator). Gate the drop on `!info.isWorkflowAgent` at **both** places:
- `registerAgentFile` initial check (`isInternalAgent(description)`),
- `processEntry`'s late re-resolution (the second `isInternalAgent` call).
**(7) `stop()` teardown.** Per-run watchers live in `dirWatchers`, so the existing close-all loop covers them — verify no separate map was introduced (24h runs spawn many `wf_<runId>` dirs → FSWatcher leak risk).
### 3b. NEW `src/workflow-run-watcher.ts` (singleton, EventEmitter — model on `team-watcher.ts`)
- **Watch root:** `~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_*.json` — **two** levels under `projects` (verified: `projects/*/workflows` is empty; must be `projects/*/*/workflows/`). chokidar `depth:3` + a poll fallback, mirroring `team-watcher`'s dual discovery + interval.
- **mtime-skip:** `runMtimes: Map<absPath, number>` (mirror `team-watcher.configMtimes`).
- **Parse:** read `wf_<runId>.json`, take the **top-level structured keys** (`runId`, `workflowName`, `summary`, `status`, `phases:[{title,detail}]`, `agentCount`, `defaultModel`, `durationMs`, `totalTokens`, `totalToolCalls`, `workflowProgress[]`). **Do NOT parse the embedded `script` string** — name/phases/summary are already top-level; the script's `export const meta` is redundant and costly. Derive `sessionUuid` from the dir name, `projectHash` from the dir above; expose `getProjectHash(workingDir)` for Codeman-session correlation.
- **`workflowProgress[] → agents[]`:** filter `type === 'workflow_agent'`, map each to a `WorkflowAgentEntry` (§3c) keyed by `agentId`. **This array is the join source the frontend uses** — no backend `getAgentLabel()` API, no TTL cache, no import from `subagent-watcher`.
- **Emit** `workflow:run_discovered|updated|removed` carrying `WorkflowRunInfo`; removal by set-diff (mirror `team-watcher`).
- **Lifecycle:** `start()`/`stop()` with `CleanupManager` teardown of chokidar + interval + caches; `LRUMap`-bounded run cache (24h memory rule).
### 3c. `src/types/` — workflow run types
```ts
export interface WorkflowAgentEntry { // one workflowProgress[type==='workflow_agent']
agentId: string; label: string; phaseIndex?: number; phaseTitle?: string;
agentType?: string; model?: string; state?: string; // 'done'|'running'|'queued'|...
lastToolName?: string; lastToolSummary?: string; tokens?: number; toolCalls?: number;
}
export interface WorkflowRunInfo {
runId: string; sessionUuid: string; projectHash: string;
workflowName?: string; summary?: string; status?: string; // 'running'|'completed'|...
phases: Array<{ title: string; detail?: string }>;
agentCount?: number; defaultModel?: string;
agents: WorkflowAgentEntry[]; // workflowProgress filtered to workflow_agent, keyed by agentId
startedAt?: number; durationMs?: number; totalTokens?: number; totalToolCalls?: number;
}
```
The two `SubagentInfo` workflow fields stay inline in `subagent-watcher.ts` (matching the existing convention).
### 3d. `src/web/sse-events.ts` — register run events
Add `workflow:run_discovered`, `workflow:run_updated`, `workflow:run_removed` after the `subagent:*` block and to the `SseEvent` union. **No new per-agent event** — workflow agents reuse `subagent:*`.
### 3e. `src/web/server.ts` — bridge, snapshot, gating
- **`setupWorkflowRunWatcherListeners()`** (beside `setupSubagentWatcherListeners`): map the three run events → `this.broadcast(...)`. Add `cleanupWorkflowRunWatcherListeners()` (store handler refs).
- **Start/stop:** call `workflowRunWatcher.start()`/`.stop()` beside `subagentWatcher`, **gated on the same enable condition** (§3f).
- **`getLightState()`:** add `workflowRuns: workflowRunWatcher.getRecentRuns(15)` beside `subagents: subagentWatcher.getRecentSubagents(15)` so headers replay on reconnect (agents already replay via `subagents`). Keep the `LIGHT_STATE_CACHE_TTL_MS` memoization.
- **Gating read:** add `isWorkflowAgentTrackingEnabled()` mirroring `isSubagentTrackingEnabled()` (boot-time `dataPath('settings.json')` read). Gate `workflowRunWatcher.start()` **and** the subagent-watcher `workflows/` descent (§3a-3) on `showUltracodeAgents` so non-opted-in users never register historical workflow agents.
### 3f. `src/web/schemas.ts` — settings key
Add `showUltracodeAgents: z.boolean().optional()` to the `.strict()` settings update schema near `showPlanUsageLimits` (required — `.strict()` 400s the whole PUT on an unknown key).
### 3g. `src/web/routes/system-routes.ts` — poll API
- `GET /api/subagents` and `GET /api/sessions/:id/subagents` include workflow agents once registered — **no change** (they carry `isWorkflowAgent`/`workflowRunId`; a consumer joins to `/api/workflows/:runId` for labels).
- Add `GET /api/workflows` → `workflowRunWatcher.getRecentRuns()` and `GET /api/workflows/:runId` (uniform `ApiResponse` contract; headers are also in `getLightState`).
- `GET /api/subagents/:agentId/transcript` works for workflow agents (they're in `agentInfo`) — no new route.
## 4. Frontend changes (file-by-file)
### 4a. `src/web/public/constants.js`
- Add the three SSE strings to `SSE_EVENTS`, matching §3d exactly (`WORKFLOW_RUN_DISCOVERED: 'workflow:run_discovered'`, etc.).
- Reuse `ZINDEX_SUBAGENT_BASE=1000` for the agent windows (they ARE subagent windows). The group **header/cluster** is in-flow panel DOM, not a floating window — no new z-index (1100 is plan-subagent).
### 4b. `src/web/public/app.js`
- Constructor: `this.workflowRuns = new Map(); // runId -> WorkflowRunInfo` beside `this.subagents`.
- `_SSE_HANDLER_MAP`: add three rows → `_onWorkflowRunDiscovered/Updated/Removed` (must exist before `connectSSE` builds the wrappers).
- `handleInit`: after seeding `data.subagents`, seed `this.workflowRuns` from `data.workflowRuns` (clear-then-set). **Clear `this.workflowRuns` everywhere the subagent Maps are cleared** (incl. `cleanupAllFloatingWindows`) — 24h leak guard.
### 4c. `src/web/public/panels-ui.js` — the join lives here
- `_onWorkflowRunDiscovered/Updated(data)` → `this.workflowRuns.set(data.runId, data)` + debounced re-render; `_onWorkflowRunRemoved` → delete + re-render.
- **No change to `_onSubagentDiscovered/Updated`** — they already store the whole payload, so the 2 new fields ride along.
- `renderSubagentPanel`/`_renderSubagentPanelImmediate`: when `showUltracodeAgents` is on, **partition `this.subagents` into flat (no `workflowRunId`) vs grouped-by-`workflowRunId`**. Flat agents render exactly as today. For each group: build the header from `this.workflowRuns.get(runId)` (`workflowName` + phase/status chip from `phases[]`), then render that run's agents reusing the existing per-agent row markup. **Per-agent label/phase/agentType come from the JOIN** — build `Map(agentId → entry)` from `this.workflowRuns.get(runId).agents` and look each agent up by `agent.agentId`; render the small chip via the `getTeammateBadgeHtml` pattern. (If the run object hasn't arrived yet, fall back to the agent's own `description` — the run `:updated` event will fill it in on the next render.)
- `findParentSessionForSubagent` is unchanged — workflow agent `sessionId === session.claudeSessionId`. **Do not conflate `workflowRunId` with `sessionId`.**
### 4d. `src/web/public/subagent-windows.js`
**Decision: REUSE `.subagent-window` per agent + a group sub-header — do NOT build a cluster class.** A cluster path duplicates Map/z-index/drag/cleanup/persistence for no functional gain; reuse keeps connection lines, minimize-to-tab, and `localStorage` persistence. In `openSubagentWindow`, where the optional `.subagent-window-parent` sub-header is built: when `agent.workflowRunId` is set, inject a `.subagent-workflow-header` showing `this.workflowRuns.get(runId)?.workflowName` + the joined agent's `label`/phase (look up by `agentId`), mirroring the `from <session>` sub-header. Respect the existing skip guards (teammate-terminal windows, minimized/`_lazyTerminal`).
**Do NOT auto-open windows** for workflow agents — a multi-phase run can spawn many, against the 50-window/60fps budget + `MAX_TRACKED_AGENTS=500`. They render collapsed in the grouped panel; the user expands via the existing panel buttons.
### 4e. `src/web/public/settings-ui.js` + `index.html`
- `index.html` Panels block: add a `settings-item` checkbox `id="appSettingsShowUltracodeAgents"` ("Show ULTRACODE / Workflow Agents").
- `openAppSettings`: load `settings.showUltracodeAgents` with `false` fallback (mirror `showPlanUsageLimits`).
- `saveAppSettings`: collect `showUltracodeAgents` into the fresh settings literal (uncollected keys reset to default every save).
- Live-apply on toggle: re-run `renderSubagentPanel()` (show/hide group sections) — a panel re-render, not a CSS-class strip.
- **SYNCED, not per-device:** do NOT add `showUltracodeAgents` to `displayKeys` and do NOT strip it in the per-device block. A synced value gives the server-side gate (`isWorkflowAgentTrackingEnabled`, §3e) one canonical truth to decide whether to run the watcher; a per-device value can't gate a process-wide watcher. (Contrast `showResponseViewer`, pure client display.)
- `styles.css` + `mobile.css`: add `.subagent-workflow-header` and `.subagent-group-badge` next to `.subagent-window-parent`; mirror device overrides in `mobile.css`.
## 5. Settings / opt-in wiring
- **Key:** `showUltracodeAgents` (boolean, **default OFF**). Fallback `false` in `openAppSettings`; "absent ⇒ off" in `isWorkflowAgentTrackingEnabled()`. Schema `z.boolean().optional()` in the `.strict()` update schema, kept OUT of `displayKeys` (synced).
- **Runtime gating:** `workflowRunWatcher.start()` and the subagent-watcher `workflows/` descent run only when the boot-time `settings.json` read reports `showUltracodeAgents === true` (mirroring `isSubagentTrackingEnabled`). The frontend additionally gates display. Toggling at runtime gates **display** immediately (panel re-render); the **watcher branch** picks up on next boot — matches existing `subagentTrackingEnabled` semantics. (Optional polish: restart just the workflow watcher on toggle for instant on/off.)
## 6. SSE events
**Reused (no change):** `subagent:discovered|updated|tool_call|tool_result|progress|message|completed`. Workflow agents flow through these; payloads now carry the optional `isWorkflowAgent`/`workflowRunId` fields on `SubagentInfo`. SSE payloads aren't schema-gated (typed only at `broadcast()` call sites), so the new fields propagate with zero friction.
**New (3 events, run-level metadata):**
| Event (backend const / frontend key) | Payload |
|---|---|
| `workflow:run_discovered` / `WORKFLOW_RUN_DISCOVERED` | `WorkflowRunInfo` |
| `workflow:run_updated` / `WORKFLOW_RUN_UPDATED` | `WorkflowRunInfo` |
| `workflow:run_removed` / `WORKFLOW_RUN_REMOVED` | `{ runId: string }` |
Sync requirement (CLAUDE.md): each must appear in **both** `sse-events.ts` (§3d) and `constants.js` `SSE_EVENTS` (§4a), be emitted via `broadcast()` in `setupWorkflowRunWatcherListeners()` (§3e), and have a dispatch-table row + `_on*` handler (§4b/§4c).
## 7. Edge cases & cleanup
- **`journal.jsonl` phantom-agent trap** — owned by §3a-4: run-dir registration requires the `agent-` prefix and excludes `journal.jsonl`.
- **`isInternalAgent` over-filtering** — owned by §3a-6: bypass at BOTH drop sites; titled from the frontend join (or the description fallback).
- **No workflow agents in the flat list** — `renderSubagentPanel` partitions on `agent.workflowRunId` (§4c). When the toggle is OFF, the descent never ran, so they aren't in `this.subagents` at all.
- **Completion/idle** — keep the existing per-agent mtime/PID/pgrep liveness as the per-card source of truth. Optionally render a group-level "workflow done" badge from run-state `status==='completed'`.
- **Limits** — `MAX_TRACKED_AGENTS=500` LRU-evicts workflow agents in the same flat map; no auto-open (50-window budget); the 4h `STARTUP_MAX_FILE_AGE_MS` skip means a run completed >4h ago won't reload after restart (acceptable — live monitoring).
- **Reconnect/replay** — agents via `getRecentSubagents(15)`; headers via `workflowRuns: getRecentRuns(15)` in `getLightState`. `handleInit` clears `this.workflowRuns` alongside the subagent Maps.
- **Watcher teardown** — every per-run `fs.watch` and the chokidar watcher closes in `stop()` and on `error`; `CleanupManager` for the new watcher (24h runs create many run dirs).
- **CLAUDE.md discipline** — read-only `~/.claude/...` artifacts; no new `~/.codeman/...` paths, no env-var prefixes touched. Claude-mode-only by nature (external CLIs don't write workflow transcripts).
## 8. Testing & verification
- **Unit (pure):**
- `test/workflow-run-watcher.test.ts`: feed a scrubbed fixture `wf_<runId>.json` → assert `WorkflowRunInfo` extraction (name/summary/phases, `workflowProgress`→`agents[]` keyed by `agentId`), mtime-skip, removal-by-set-diff.
- Extend `subagent-watcher` coverage: temp `subagents/workflows/wf_X/agent-Y.jsonl` + a stray `journal.jsonl` → assert `agent-Y` registered with `isWorkflowAgent`/`workflowRunId` and `journal.jsonl` NOT registered; assert a short-description workflow agent is NOT dropped at **either** `isInternalAgent` site.
- **Route/inject (`app.inject`):** `GET /api/workflows` + `:runId` return the `ApiResponse` envelope; `GET /api/subagents` includes a tagged agent.
- **Frontend (vm-sandbox, like `test/run-mode-ui.test.ts`):** dispatch `subagent:discovered` with `workflowRunId` + `workflow:run_discovered` → assert `renderSubagentPanel` produces a group section under the workflow name with the agent inside it (label sourced from the **join**, not flat); assert order-independence (agent before run, and run before agent both resolve); assert OFF hides the section.
- **REQUIRED real end-to-end** (the always-end-to-end-test rule — the plan-usage chip shipped *dead* from a gate mismatch): on dev/beta with `showUltracodeAgents` ON, **drive a real ultracode/workflow run**, then (1) `curl …/api/workflows | jq` shows the live run with `agents[]`; (2) `curl …/api/subagents | jq '.data[]|select(.isWorkflowAgent)'` shows tagged agents; (3) watch `/api/events` for `workflow:run_discovered` + `subagent:discovered` with the workflow fields; (4) Playwright (`waitUntil:'domcontentloaded'`, wait 3–4s) asserts the grouped DOM cluster renders with the workflow-name header and live status. Verify path gates against `GET /api/sessions` `workingDir`. **Test against a LIVE run** — all at-rest runs are `completed`/`done`; `running`/`queued` states only exist mid-run.
## 9. Phased rollout
| Phase | Scope | Done-check | Size |
|---|---|---|---|
| **P1 — Backend discovery + tagging (gated, no UI)** | §3a (nested descent, `journal.jsonl` skip, 2 `SubagentInfo` fields, `isInternalAgent` bypass ×2) + §3f schema key + §3e gate read. No run watcher yet. | With `showUltracodeAgents` forced on, `curl /api/subagents \| jq '.data[]\|select(.isWorkflowAgent)'` lists real workflow agents during a live run; flat subagents unchanged; `tsc --noEmit` + targeted watcher test green. | S–M |
| **P2 — Run-state metadata + SSE** | §3b (`workflow-run-watcher.ts`) + §3c types + §3d/§3e (SSE, bridge, `getLightState` replay) + §3g routes. | `curl /api/workflows \| jq` returns runs with `agents[]`/`phases`; SSE emits `workflow:run_discovered`; reconnect snapshot carries `workflowRuns`. | M |
| **P3 — Frontend grouped UI** | §4a–§4d (constants, app.js state/dispatch/init, panels-ui grouped render + **agent→label join**, subagent-windows group sub-header). Reuse `.subagent-window`; no auto-open. | Playwright: live run renders a group section under the workflow name with per-agent rows + live status + joined labels; flat subagents stay flat; expand opens a window with the workflow sub-header. | M |
| **P4 — Settings toggle + polish + docs** | §4e (checkbox, settings-ui load/save/live-apply, SYNCED), styles/mobile, phase chips, CLAUDE.md "Key Patterns" entry + this doc's status → SHIPPED. | Toggling the checkbox shows/hides the cluster live (no reload for display); OFF by default on a fresh install; CI green. | S |
Each phase is independently shippable: P1 is invisible (gated, no UI), P2 adds an API with no UI dependency, P3 lights up the UI for flag-enablers, P4 exposes the toggle and finalizes defaults/docs.
## 10. Effort & risk
**Size:** P1 = S–M, P2 = M, P3 = M, P4 = S. Total ≈ **M** (one focused engineer, ~2–4 days incl. the real end-to-end run — down from the first draft's M-L now that the backend join/coupling is gone).
**Top 3 risks:**
1. **Non-recursive watch on Linux misses live writes.** `fs.watch` is non-recursive and `{recursive:true}` is unreliable on Linux → per-`wf_<runId>` watchers (§3a-4) are correct, but the 5s full scan must re-`readdir(wfRoot)` to catch *new* run dirs mid-session, and each watcher must be torn down to avoid FSWatcher leaks in 24h runs. Mitigation: explicit per-run-dir registration + verified `dirWatchers` teardown; chokidar (with `CleanupManager`) only in the new run watcher, where `team-watcher` already proves the pattern.
2. **Discovery cost / over-registration.** A user with hundreds of historical workflow agents could flood `agentInfo` on boot. Mitigation: the 4h `STARTUP_MAX_FILE_AGE_MS` skip drops old files on the initial scan, the descent only runs when the toggle is on, and `MAX_TRACKED_AGENTS=500` LRU-evicts. Verify boot scan time doesn't regress with the corpus present.
3. **Shipping-dead-on-a-gate** (the repo's recurring failure mode — the plan-usage chip shipped dead because injection was gated on `CASES_DIR` while real sessions ran elsewhere). Same trap here if the path/mode gate is wrong (e.g. `projects/*/workflows` instead of `projects/*/*/workflows`, or correlation via the wrong session key). Mitigation: the **mandatory live ultracode end-to-end run** in §8 against a real session's `workingDir`, observing the real SSE event + real DOM cluster — not the at-rest corpus, not unit tests alone.
+169
View File
@@ -0,0 +1,169 @@
# Plan Usage Limits Display — Design & As-Built
> **Status: SHIPPED — deployed to prod + pushed to master, not yet released (2026-06-14).** Opt-in via App Settings → Display → **Plan Usage Limits** (`showPlanUsageLimits`, default OFF). Commits `c82f6c8` (feature) → `4d9d93d` (end-to-end fixes) → `eae225b` (per-user reconcile) → `95fb5fc` (init-snapshot replay). Full suite green (2869), CI green. No changeset/version bump yet.
>
> Two surfaces from one `statusLine` callback:
> - **Header chip** (top-right) — account-wide **plan limits**: `5h 35% · 7d 38%`, per-window green/yellow/red.
> - **In-terminal statusline footer** — the **current session's** status: `Opus 4.8 (1M context) in:562,411 out:1,188 ctx:56%`.
>
> The `rate_limits` JSON schema below was **empirically confirmed** against Claude Code 2.1.177 on a Claude Max account; see the Verification appendix to reproduce.
## Problem
Codeman had no proactive view of how much of the Claude subscription is left. It only learned about limits **reactively**: `usage-limit-patterns.ts` regex-scrapes ANSI-stripped terminal output for footer strings like `5-hour limit reached ∙ resets 8pm`, extracting only the **reset time**, and only *after* Claude has already stalled. There was no "73% of your 5-hour limit used" anywhere.
We wanted a live, always-visible gauge so the operator can see a wall coming and pace overnight/autonomous runs — without hijacking the in-terminal statusline, which should keep showing the current session's status.
## Data source: the statusline `rate_limits` JSON
Claude Code (**v2.1.80+**; prod box runs **2.1.177**) pipes a JSON blob to a configured `statusLine.command` on stdin after each render. On Pro/Max subscriptions that blob includes `rate_limits`. **This is the only channel that exposes plan-limit data** (see rejected alternatives) — so the feature *must* set a statusLine command, which is why the footer is also reconstructed by it (below).
### Confirmed schema (real captured payload)
```jsonc
"rate_limits": {
"five_hour": { "used_percentage": 15, "resets_at": 1781409000 }, // → 2026-06-14T03:50:00Z
"seven_day": { "used_percentage": 34, "resets_at": 1781827200 } // → 2026-06-19T00:00:00Z
}
```
| Field | Type | Notes |
|-------|------|-------|
| `rate_limits.five_hour.used_percentage` | `number` 0–100 | Integer-valued in practice; treat as `number`, don't assume decimals. |
| `rate_limits.five_hour.resets_at` | `number` | **Epoch SECONDS** (10 digits). `×1000` for a JS `Date`. |
| `rate_limits.seven_day.{used_percentage,resets_at}` | same | |
**Confirmed facts & gotchas:**
- **Only two windows exist: `five_hour` and `seven_day`.** There is **no separate Opus-weekly field**, even on a Max/Opus account.
- `rate_limits` is **absent on the first render**, **present after the first API response**. UI degrades to "no chip yet."
- statusLine fires **only in interactive TUI mode**, never `--print`. Fine — Codeman sessions are interactive TUIs (and so are Codeman-spawned ones in tmux).
- **Subscriber-gated.** Absent for API-key / non-subscriber auth.
### Bonus telemetry in the same payload — used for the footer
The same stdin object also carries `model.display_name`, `context_window.{used_percentage, total_input_tokens, total_output_tokens, …}`, `cost.total_cost_usd`, `effort.level`, etc. The shipped feature uses **model + token totals + context %** to build the in-terminal footer (so the statusline stays useful even though we own it). The endpoint also broadcasts `contextUsedPercentage`/`costUsd`/`modelDisplayName` alongside the limits for future chip tooltips.
### Alternatives considered & rejected
| Source | Why not |
|--------|---------|
| OAuth endpoint `api.anthropic.com/api/oauth/usage` | Undocumented, aggressively rate-limited, needs the **encrypted** OAuth token. Only worth it for *dollar spend*. |
| `/usage` slash command | Interactive-only, no programmatic output. |
| On-disk `~/.claude/` files | No usage state persisted (only `daemon.status.json` = auto-updater supervisor). |
| CLI flag (`claude usage` / `--check-usage`) | Does not exist. |
| `StopFailure` hook | Carries only an `error_type` on *failure* — no live percentages. |
## As-built architecture
```
Claude TUI (any Claude session, incl. linked-case/real-repo sessions)
│ renders statusline after each assistant msg (+ /compact, mode change)
▼
statusLine.command (settings.local.json) ──reads stdin JSON──▶
curl -sk POST $CODEMAN_API_URL/api/status-telemetry {sessionId, data}
(X-Codeman-Hook-Secret: $(cat $CODEMAN_HOOK_SECRET_FILE))
│ ◀── HTTP 200 text/plain = current-SESSION status string ──┘
▼
printf '%s' "$body" → in-terminal footer: "Opus 4.8 (1M context) in:… out:… ctx:…%"
server (status-telemetry-routes.ts):
parse rate_limits → (if changed) store last-known + broadcast SSE session:statusTelemetry → header chip
parse model/tokens/ctx → return the session-status footer string
▼
app.js: _onSessionStatusTelemetry → chip (per-window colors) + localStorage save
handleInit → chip from init-snapshot planUsage (fresh-load replay)
```
### 1. The exporter — `generateStatusLineCommand()` in `hooks-config.ts`
Mirrors the hook `curlCmd()`. Reads the stdin JSON, POSTs `{sessionId, data}` to a **fixed** loopback path, and prints the response body back to stdout (print-through, so the footer stays useful). The managed-session env carries `$CODEMAN_SESSION_ID` / `$CODEMAN_API_URL` / `$CODEMAN_HOOK_SECRET_FILE` (from `tmux-manager.buildEnvExports()`).
```bash
INPUT=$(cat 2>/dev/null || echo '{}'); \
printf '{"sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$INPUT" | \
curl -sk -X POST "$CODEMAN_API_URL/api/status-telemetry" \
-H 'Content-Type: application/json' \
-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" \
--data @- 2>/dev/null || echo codeman
```
⚠️ **`curl -sk`, not `curl -s`.** Prod is loopback **HTTPS with a self-signed cert**; without `-k`, curl returns `000` and the statusline silently shows nothing. `-k` is safe (loopback only). *(The existing hook curls use `-s` without `-k` and have the same latent issue on HTTPS installs — a known, separate follow-up.)*
### 2. Endpoint — `POST /api/status-telemetry` (`status-telemetry-routes.ts`)
Fixed path (sessionId in the **body**, not the URL) so the auth exemption is an exact-match like `/api/hook-event` (`middleware/auth.ts`: loopback-only; `X-Codeman-Hook-Secret`-gated while a tunnel runs). Schema `StatusTelemetrySchema` in `schemas.ts` validates the subset; unknown keys are stripped. Pure parsing/formatting in `usage-telemetry.ts`:
- `parseStatusTelemetry(data)` → `{ fiveHour, sevenDay, … }` or `null`. On change (signature dedup; statusline fires often), store last-known (`plan-usage-latest.ts`) and `broadcast('session:statusTelemetry', { sessionId, …telemetry })`.
- `parseSessionStatus(data)` + `formatSessionStatusText()` → the **footer** string `Opus 4.8 (1M context) in:562,411 out:1,188 ctx:56%` (returned as `text/plain`). Available from the first render, even before `rate_limits` appears.
### 3. SSE + frontend chip
`session:statusTelemetry` registered in `sse-events.ts` + `constants.js`. `app.js`:
- `_onSessionStatusTelemetry` → `updatePlanUsageChip(data)` + save to `localStorage['codeman:planUsage']`.
- `updatePlanUsageChip` renders two `5h`/`7d` windows; **per-window color by usage** — green `<60%`, yellow `60–84%`, red `≥85%` (`pu-green/pu-yellow/pu-red`); bold labels/values; reset times in the tooltip. `resets_at*1000 → Date`.
- Chip element ships hidden (`header-plan-usage--hidden`); `applyHeaderVisibilitySettings()` reveals it client-side when the setting is on (response-viewer pattern — **no `renderIndexHtml` strip**, which kept the "title-only" render contract intact).
### 4. Chip data robustness — three layers
1. **Live:** `session:statusTelemetry` SSE on every distinct render.
2. **Fresh load / reconnect:** server stores the latest in `plan-usage-latest.ts`; `getLightState()` includes it as `planUsage`; the per-connection **init snapshot** replays it; `handleInit` paints the chip immediately (authoritative over localStorage). Null until the first telemetry of the process.
3. **Offline / cross-restart:** `restorePlanUsageChip()` reads `localStorage` on load (12h freshness guard).
### 5. Injection lifecycle — works for *any* user, never self-destructs
The setting `showPlanUsageLimits` is **synced** (in `settings.json`, not a per-device `displayKey`).
- **On toggle** (`PUT /api/settings`, `system-routes.ts`): reconcile the exporter across **all active Claude sessions' working dirs** — inject on enable, remove on disable. Server-side and authoritative, so existing sessions get the footer + feed the chip *immediately*, no new session needed, no dependency on a client's synced localStorage.
- **On session create** (`session-routes.ts`): **ADD-ONLY** — inject when `statusLineTelemetry` is true; **never remove**. Sessions in a repo share one `settings.local.json`, so a single create-with-false (e.g. a client whose synced setting hadn't loaded) must not yank the statusLine out from under other live sessions. Removal happens only via the explicit toggle.
- `applyStatusLineConfig()` is **`isOurs`-guarded** (matches `/api/status-telemetry`), so a user's own hand-authored statusLine is never touched, and it **updates an out-of-date ours-command** so fixes (e.g. `-k`) propagate. **No `CASES_DIR` gate** — runs for linked cases / real repos (where sessions actually run), mirroring `updateCaseModel`.
## Codeman-specific considerations
1. **Account-global limits.** The 5h/7d pools are shared across all sessions on the account → one shared header chip (freshest sample wins), not a per-tab bar.
2. **The footer is owned, by necessity.** A statusLine command always replaces Claude's default footer. Since `rate_limits` *only* arrives via statusLine, we reconstruct a useful **session-status** footer (model · tokens · ctx %) from the same payload rather than showing the limits there.
3. **`isOurs`-guarded.** Never removes/overwrites a user's own statusLine on disable; only manages the Codeman exporter.
4. **Security envelope unchanged.** The exporter runs arbitrary shell every render — same trust model as the hook curls (localhost + `$CODEMAN_HOOK_SECRET_FILE`); reuses the hook-secret gate.
5. **Claude-only.** OpenCode/Codex emit no `rate_limits` JSON; injection is gated to `mode === 'claude'`.
6. **Future — auto-resume synergy.** Live percentages would let `SessionAutoOps` pre-arm *before* the wall instead of reacting to the stall footer. Not built.
## Files shipped
- `src/usage-telemetry.ts` — pure parse/format (`parseStatusTelemetry`, `parseSessionStatus`, `formatSessionStatusText`, `telemetrySignature`) + `test/usage-telemetry.test.ts`.
- `src/hooks-config.ts` — `generateStatusLineCommand()` (`curl -sk`), `applyStatusLineConfig()` (add/update/remove, `isOurs`-guarded).
- `src/web/routes/status-telemetry-routes.ts` — `POST /api/status-telemetry`.
- `src/web/plan-usage-latest.ts` — process-wide last-known store for init replay.
- `src/web/schemas.ts` — `StatusTelemetrySchema` + `showPlanUsageLimits` + create-payload `statusLineTelemetry`.
- `src/web/middleware/auth.ts` — exemption extended to `/api/status-telemetry`.
- `src/web/routes/session-routes.ts` — add-only create-time injection.
- `src/web/routes/system-routes.ts` — settings-toggle reconcile.
- `src/web/server.ts` — `getLightState().planUsage` (init snapshot).
- `src/web/sse-events.ts` + `constants.js` — `session:statusTelemetry`.
- Frontend: `app.js` (`_onSessionStatusTelemetry`, `updatePlanUsageChip`, `restorePlanUsageChip`, `handleInit`), `settings-ui.js` (toggle + `applyHeaderVisibilitySettings`), `index.html` (chip + toggle row), `styles.css` (chip + colors), `session-ui.js` (create payload).
## Bugs E2E testing caught (that unit tests didn't)
The first "shipped" build passed every test and was broken in practice. End-to-end testing on the real install (the lesson: drive a REAL session, observe the REAL output) surfaced:
1. **`CASES_DIR` injection gate** excluded the user's whole workflow — sessions run in linked cases / real repos, not under `~/codeman-cases`. → dropped the gate.
2. **`curl -s` → `000`** on the loopback self-signed HTTPS cert; statusline silently empty. → `curl -sk`.
3. **Remove-on-create-false + shared `settings.local.json`** let a single stale client yank the statusLine out from under all sessions in a repo. → add-only on create; removal only via the toggle reconcile.
4. **Chip blank after reload** (localStorage-only, lost on restart/fresh browser). → server-side last-known in the init snapshot.
## Open questions / future
- **Schema stability.** `rate_limits` is officially shipped but undocumented in exact shape; the parser is tolerant (renders whatever windows exist, ignores unknown).
- **Hook `curl -s` parity.** Hooks share the no-`-k` issue on HTTPS installs — worth fixing the hook curl too (separate change; covered by `cod54` tests).
- **Disable cleanliness.** Disabling removes the statusLine from active sessions; a brand-new session created by a *stale* client could re-add it (chip still hidden, footer benign). Fully server-authoritative create-time injection (read the setting server-side instead of the payload flag) would close this — deferred.
## Verification appendix — how the schema was captured (reproducible)
Captured without touching global settings or any real session:
1. Throwaway dir `/tmp/sl-capture` with an exporter `dump.sh` that appends stdin to `payloads.jsonl` and prints `cap`; a `settings.json` pointing `statusLine.command` at it.
2. `--print` mode does **not** render a statusline → no capture (confirms TUI-only). Must use interactive.
3. Launch interactive Claude in an **isolated tmux socket** (`tmux -L slcap`, never `-L codeman`) inside the temp dir, `--settings /tmp/sl-capture/settings.json` (no global mutation). Confirm the workspace-trust dialog (appears even with `--dangerously-skip-permissions`), then send a one-line prompt (literal text + Enter separately, Ink-style).
4. After the first response, `rate_limits` appears in the **second** captured record (absent in the first). Inspect with `jq '.rate_limits'`.
5. Tear down: `tmux -L slcap kill-server` + `rm -rf /tmp/sl-capture`; verify the `codeman` socket is untouched.
Related: `docs/claude-code-hooks-reference.md` (hook callback pattern), `src/usage-limit-patterns.ts` (reactive fallback), `docs/respawn-state-machine.md` (auto-resume interplay).
+38 -2
View File
@@ -1,12 +1,12 @@
{
"name": "aicodeman",
"version": "0.9.9",
"version": "1.4.1",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "aicodeman",
"version": "0.9.9",
"version": "1.4.1",
"hasInstallScript": true,
"license": "MIT",
"workspaces": [
@@ -20,6 +20,7 @@
"@fastify/static": "^9.1.3",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-serialize": "^0.14.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
"@xterm/xterm": "^6.0.0",
@@ -27,6 +28,8 @@
"chokidar": "^3.6.0",
"commander": "^12.1.0",
"fastify": "^5.8.5",
"heic-decode": "^2.1.0",
"jpeg-js": "^0.4.4",
"node-pty": "^1.1.0",
"qrcode": "^1.5.4",
"uuid": "^14.0.0",
@@ -4525,6 +4528,12 @@
"integrity": "sha512-jYcgT6xtVYhnhgxh3QgYDnnNMYTcf8ElbxxFzX0IZo+vabQqSPAjC3c1wJrKB5E19VwQei89QCiZZP86DCPF7g==",
"license": "MIT"
},
"node_modules/@xterm/addon-serialize": {
"version": "0.14.0",
"resolved": "https://registry.npmjs.org/@xterm/addon-serialize/-/addon-serialize-0.14.0.tgz",
"integrity": "sha512-uteyTU1EkrQa2Ux6P/uFl2fzmXI46jy5uoQMKEOM0fKTyiW7cSn0WrFenHm5vO5uEXX/GpwW/FgILvv3r0WbkA==",
"license": "MIT"
},
"node_modules/@xterm/addon-unicode11": {
"version": "0.9.0",
"resolved": "https://registry.npmjs.org/@xterm/addon-unicode11/-/addon-unicode11-0.9.0.tgz",
@@ -7016,6 +7025,18 @@
"node": ">= 0.4"
}
},
"node_modules/heic-decode": {
"version": "2.1.0",
"resolved": "https://registry.npmjs.org/heic-decode/-/heic-decode-2.1.0.tgz",
"integrity": "sha512-0fB3O3WMk38+PScbHLVp66jcNhsZ/ErtQ6u2lMYu/YxXgbBtl+oKOhGQHa4RpvE68k8IzbWkABzHnyAIjR758A==",
"license": "ISC",
"dependencies": {
"libheif-js": "^1.19.8"
},
"engines": {
"node": ">=8.0.0"
}
},
"node_modules/html-encoding-sniffer": {
"version": "4.0.0",
"resolved": "https://registry.npmjs.org/html-encoding-sniffer/-/html-encoding-sniffer-4.0.0.tgz",
@@ -7474,6 +7495,12 @@
"node": ">=10"
}
},
"node_modules/jpeg-js": {
"version": "0.4.4",
"resolved": "https://registry.npmjs.org/jpeg-js/-/jpeg-js-0.4.4.tgz",
"integrity": "sha512-WZzeDOEtTOBK4Mdsar0IqEU5sMr3vSV2RqkAIzUEV2BHnUfKGyswWFPFwK5EeDo93K3FohSHbLAjj0s1Wzd+dg==",
"license": "BSD-3-Clause"
},
"node_modules/js-tokens": {
"version": "10.0.0",
"resolved": "https://registry.npmjs.org/js-tokens/-/js-tokens-10.0.0.tgz",
@@ -7657,6 +7684,15 @@
"node": ">= 0.8.0"
}
},
"node_modules/libheif-js": {
"version": "1.19.8",
"resolved": "https://registry.npmjs.org/libheif-js/-/libheif-js-1.19.8.tgz",
"integrity": "sha512-vQJWusIxO7wavpON1dusciL8Go9jsIQ+EUrckauFYAiSTjcmLAsuJh3SszLpvkwPci3JcL41ek2n+LUZGFpPIQ==",
"license": "LGPL-3.0",
"engines": {
"node": ">=8.0.0"
}
},
"node_modules/light-my-request": {
"version": "6.6.0",
"resolved": "https://registry.npmjs.org/light-my-request/-/light-my-request-6.6.0.tgz",
+5 -2
View File
@@ -1,7 +1,7 @@
{
"name": "aicodeman",
"version": "0.9.9",
"description": "The missing control plane for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"version": "1.4.1",
"description": "Mission control for AI coding agents - run 20 autonomous agents with real-time monitoring and session persistence",
"type": "module",
"main": "dist/index.js",
"types": "dist/index.d.ts",
@@ -61,6 +61,7 @@
"@fastify/static": "^9.1.3",
"@fastify/websocket": "^11.2.0",
"@xterm/addon-fit": "^0.11.0",
"@xterm/addon-serialize": "^0.14.0",
"@xterm/addon-unicode11": "^0.9.0",
"@xterm/addon-webgl": "^0.19.0",
"@xterm/xterm": "^6.0.0",
@@ -68,6 +69,8 @@
"chokidar": "^3.6.0",
"commander": "^12.1.0",
"fastify": "^5.8.5",
"heic-decode": "^2.1.0",
"jpeg-js": "^0.4.4",
"node-pty": "^1.1.0",
"qrcode": "^1.5.4",
"uuid": "^14.0.0",
+138 -5
View File
@@ -17,6 +17,13 @@
// • Panel "re-grab" — pinch an existing floating panel and move it anywhere;
// release over the tab strip to re-dock it (panel goes away, the tab stays).
// This is the capability the old OS-window detach lost.
// • Agent-window "grab-to-move" — pinch any floating *subagent* or *ultracode*
// run/transcript window (the dashboard's own `.subagent-window` /
// `.ultracode-window` floats) and move it anywhere. These windows stay owned
// by app.js — we only nudge their `style.left/top` and ask app.js to redraw
// the glowing connector line back to their session tab (its redraw reads live
// rects, so the line tracks without us touching app.js internals). This is the
// multi-monitor verb that lets these windows cross the physical monitor seam.
// • Button "tap" — pinch over a toolbar button (Run / Run Shell) and release
// in place → fires the button's real click handler. Drift too far first and
// it's treated as a stray move, not a tap.
@@ -36,12 +43,29 @@ import type { HandState } from '../gesture/types.ts';
declare global {
interface Window {
__codemanGesture?: GestureBridge;
/** The Codeman dashboard singleton (app.js, `window.app`). The gesture layer
* reaches into it to redraw the floating-window connector lines and bump a
* grabbed window's z-order while moving the subagent / ultracode windows.
* Loosely typed — only the few members we touch. */
app?: {
updateConnectionLines?: () => void;
saveSubagentWindowStates?: () => void;
subagentWindowZIndex?: number;
ultracodeWindowZIndex?: number;
};
}
}
const TAB_SELECTOR = '.session-tab';
/** An in-page floating session panel this layer spawned — re-grabbable to move. */
const PANEL_SELECTOR = '.cg-float';
/** The dashboard's own floating agent windows (subagent runs + ultracode run and
* transcript windows). All three carry one of these classes, position via
* `style.left/top`, and redraw their connector line from
* `window.app.updateConnectionLines()` — so the hand can pick one up and move it
* without app.js knowing. (`.ultracode-agent-window` also carries
* `.ultracode-window`, so this matches it too.) */
const WINDOW_SELECTOR = '.subagent-window, .ultracode-window';
/** The session-tab strip; dropping a moved panel over it re-docks the session. */
const DOCK_SELECTOR = '.session-tabs';
/** Toolbar buttons a pinch can "tap": Run (#runBtn → app.run()) and Run Shell
@@ -93,6 +117,17 @@ type Grab =
dy: number;
/** Cursor currently over the tab strip → releasing re-docks. */
overDock: boolean;
}
| {
/** A dashboard-owned floating agent window (subagent / ultracode) being
* moved. We never remove or re-parent it — just reposition + redraw its
* connector. The element ref can go stale mid-grab (SSE reconnect tears
* ultracode windows down), so every move guards on `el.isConnected`. */
kind: 'window';
el: HTMLElement;
/** Cursor→window-top-left offset at grab, so it doesn't snap. */
dx: number;
dy: number;
};
/** Live state for one hand pinching a toolbar button (Run / Run Shell). */
@@ -122,6 +157,8 @@ class GestureBridge {
private taps = new Map<string, Tap>();
/** Live floating panels, keyed by session id (idempotent per id). */
private floats = new Map<string, FloatingPanel>();
/** rAF coalescing for connector-line redraws while dragging an agent window. */
private connectorRedrawScheduled = false;
constructor() {
injectStyles();
@@ -187,7 +224,7 @@ class GestureBridge {
await this.gc.start();
this.running = true;
this.button.classList.add('on');
this.status.textContent = 'on — pinch a tab or button';
this.status.textContent = 'on — pinch a tab, window, or button';
} catch (err) {
// Surface the *real* cause: MediaPipe/Emscripten can throw a non-Error
// (number/string), so `(err as Error).message` was logging "undefined".
@@ -242,6 +279,22 @@ class GestureBridge {
}
}
// A dashboard-owned floating agent window (subagent / ultracode run or
// transcript) → pick it up and move it. Priority below cg-float panels
// (which sit far above), above tabs/buttons. We grab anywhere on the window
// (not just its titlebar) since the hand is choosing the whole window.
const win = this.hitClosest(x, y, WINDOW_SELECTOR);
if (win) {
const rect = win.getBoundingClientRect();
// Match app.js's own drag: drop any bottom-anchor so left/top take effect.
win.style.bottom = 'auto';
win.classList.add('cg-win-grabbed');
this.bringWindowToFront(win);
this.grabs.set(hand, { kind: 'window', el: win, dx: x - rect.left, dy: y - rect.top });
this.status.textContent = 'moving window';
return;
}
// A session tab → grab-and-pull-out into a floating panel (ghost follows).
const tab = this.hitClosest(x, y, TAB_SELECTOR);
const id = tab?.dataset.id;
@@ -292,12 +345,16 @@ class GestureBridge {
}
return;
}
if (grab?.kind === 'window') {
this.moveWindow(grab.el, x - grab.dx, y - grab.dy);
return;
}
// A button pinch that drifts too far is a stray move, not a tap — cancel it.
const tap = this.taps.get(hand);
if (tap && Math.hypot(x - tap.ox, y - tap.oy) > TAP_CANCEL_PX) {
tap.el.classList.remove('cg-tap-armed');
this.taps.delete(hand);
this.status.textContent = 'on — pinch a tab or button';
this.status.textContent = 'on — pinch a tab, window, or button';
}
}
@@ -319,6 +376,23 @@ class GestureBridge {
else this.flash('placed');
return;
}
if (grab?.kind === 'window') {
this.grabs.delete(hand);
grab.el.classList.remove('cg-win-grabbed');
// Clear the coalescer so the final placement always redraws, even if a
// mid-drag rAF was throttled (tab briefly backgrounded) and left it latched.
this.connectorRedrawScheduled = false;
this.redrawWindowConnectors();
// Persist subagent-window positions like app.js's own drag end does
// (a no-op for ultracode windows, which aren't position-persisted).
try {
window.app?.saveSubagentWindowStates?.();
} catch {
/* best-effort */
}
this.flash('placed window');
return;
}
// Release over the same button → fire its real click handler.
const tap = this.taps.get(hand);
if (tap) {
@@ -373,6 +447,59 @@ class GestureBridge {
float.el.style.top = `${t}px`;
}
/** Move a dashboard-owned agent window by its top-left, clamped on-screen, then
* redraw its connector line. The window self-positions via `style.left/top` and
* app.js's connector redraw reads live rects, so this tracks without touching
* app.js internals. Guards on `isConnected`: ultracode windows can be torn down
* (SSE reconnect / auto-close) while still held. Clamps to `innerWidth/Height`,
* which equals the *spanned* viewport in a multi-monitor window — so the window
* can still travel across the physical monitor seam, just not off-screen. */
private moveWindow(el: HTMLElement, left: number, top: number): void {
if (!el.isConnected) return;
const w = el.offsetWidth || 380;
const h = el.offsetHeight || 320;
const l = Math.min(Math.max(4, left), Math.max(4, window.innerWidth - w - 4));
const t = Math.min(Math.max(4, top), Math.max(4, window.innerHeight - h - 4));
el.style.left = `${l}px`;
el.style.top = `${t}px`;
this.redrawWindowConnectors();
}
/** Ask app.js to redraw all connector lines (subagent + ultracode), coalesced to
* one per frame so per-frame drags don't thrash. `updateConnectionLines()` is
* itself debounced in app.js, but we rAF-gate too in case an older dashboard
* build isn't, and to no-op cleanly when app.js isn't present (standalone). */
private redrawWindowConnectors(): void {
if (this.connectorRedrawScheduled) return;
this.connectorRedrawScheduled = true;
requestAnimationFrame(() => {
this.connectorRedrawScheduled = false;
try {
window.app?.updateConnectionLines?.();
} catch {
/* app.js may not expose it (standalone playground) */
}
});
}
/** Pop a grabbed window above its siblings using app.js's own z-counter, so a
* picked-up window comes to the front like a real focus. Cosmetic + best-effort. */
private bringWindowToFront(el: HTMLElement): void {
const app = window.app;
if (!app) return;
try {
if (el.classList.contains('ultracode-window')) {
app.ultracodeWindowZIndex = (app.ultracodeWindowZIndex ?? 1000) + 1;
el.style.zIndex = String(app.ultracodeWindowZIndex);
} else {
app.subagentWindowZIndex = (app.subagentWindowZIndex ?? 1000) + 1;
el.style.zIndex = String(app.subagentWindowZIndex);
}
} catch {
/* cosmetic only */
}
}
private positionGhost(ghost: HTMLElement, x: number, y: number): void {
ghost.style.left = `${x}px`;
ghost.style.top = `${y}px`;
@@ -385,17 +512,19 @@ class GestureBridge {
if (grab.kind === 'tab') {
grab.ghost.remove();
grab.tab.classList.remove('cg-grabbed');
} else {
} else if (grab.kind === 'panel') {
grab.panel.el.style.pointerEvents = '';
grab.panel.el.classList.remove('cg-float-grabbed', 'cg-redock');
} else {
grab.el.classList.remove('cg-win-grabbed');
}
}
this.grabs.clear();
for (const tap of this.taps.values()) tap.el.classList.remove('cg-tap-armed');
this.taps.clear();
document
.querySelectorAll(`${TAB_SELECTOR}.cg-grabbed, .cg-tap-armed`)
.forEach((t) => t.classList.remove('cg-grabbed', 'cg-tap-armed'));
.querySelectorAll(`${TAB_SELECTOR}.cg-grabbed, .cg-tap-armed, .cg-win-grabbed`)
.forEach((t) => t.classList.remove('cg-grabbed', 'cg-tap-armed', 'cg-win-grabbed'));
}
private onStatus(fps: number, hands: HandState[]): void {
@@ -491,6 +620,10 @@ function injectStyles(): void {
.cg-status { color: #9aa0a6; max-width: 220px; overflow: hidden; text-overflow: ellipsis; white-space: nowrap; }
.session-tab.cg-grabbed { opacity: .35; outline: 2px dashed #4ade80; outline-offset: -2px; }
.cg-tap-armed { outline: 2px solid #4ade80 !important; outline-offset: 2px; box-shadow: 0 0 0 4px rgba(74,222,128,.25) !important; }
.subagent-window.cg-win-grabbed, .ultracode-window.cg-win-grabbed {
outline: 2px solid #4ade80 !important; outline-offset: -2px;
box-shadow: 0 12px 48px rgba(74,222,128,.5) !important;
}
.cg-float {
position: fixed; left: 0; top: 0; width: ${FLOAT_W}px; height: ${FLOAT_H}px;
z-index: ${Z}; display: flex; flex-direction: column; overflow: hidden;
+1 -1
View File
@@ -39,7 +39,7 @@ Server echoes 'h' ←───────────────────
## Origin
This library was extracted from [Codeman](https://github.com/Ark0N/Codeman), the missing control plane for AI coding agents — multi-session management, real-time agent visualization, autonomous respawn loops, and a mobile-first web UI for Claude Code and OpenCode. The local echo system was built to make mobile and remote access feel instant, then battle-tested across thousands of hours of real usage. After 3 deep code audits, it was extracted into this standalone library with 78 tests covering every state transition.
This library was extracted from [Codeman](https://github.com/Ark0N/Codeman), mission control for AI coding agents — multi-session management, real-time agent visualization, autonomous respawn loops, and a mobile-first web UI for Claude Code, OpenCode, and Codex. The local echo system was built to make mobile and remote access feel instant, then battle-tested across thousands of hours of real usage. After 3 deep code audits, it was extracted into this standalone library with 78 tests covering every state transition.
## Install
+72
View File
@@ -0,0 +1,72 @@
#!/usr/bin/env node
/**
* Build the Codeman agent base image locally (decision: "build locally on first
* use", see docs/docker-cases-plan.md). No registry account required.
*
* Usage:
* node scripts/build-agent-image.mjs [--engine docker|podman] [--image <ref>] [--no-cache]
*
* Defaults: engine=docker (falls back to podman if docker is absent),
* image=codeman/agent:base
*/
import { spawn, spawnSync } from 'node:child_process';
import { fileURLToPath } from 'node:url';
import { dirname, join } from 'node:path';
const __dirname = dirname(fileURLToPath(import.meta.url));
const REPO_ROOT = join(__dirname, '..');
const DOCKERFILE = join(REPO_ROOT, 'docker', 'agent.Dockerfile');
const DEFAULT_IMAGE = 'codeman/agent:base';
function parseArgs(argv) {
const args = { image: DEFAULT_IMAGE, engine: undefined, noCache: false };
for (let i = 0; i < argv.length; i++) {
const a = argv[i];
if (a === '--image') args.image = argv[++i];
else if (a === '--engine') args.engine = argv[++i];
else if (a === '--no-cache') args.noCache = true;
else if (a === '-h' || a === '--help') args.help = true;
}
return args;
}
function engineAvailable(engine) {
const r = spawnSync(engine, ['--version'], { stdio: 'ignore' });
return r.status === 0;
}
function resolveEngine(preferred) {
if (preferred) {
if (!engineAvailable(preferred)) {
console.error(`[build-agent-image] engine "${preferred}" not found on PATH`);
process.exit(1);
}
return preferred;
}
if (engineAvailable('docker')) return 'docker';
if (engineAvailable('podman')) return 'podman';
console.error('[build-agent-image] neither docker nor podman found on PATH. Install one and retry.');
process.exit(1);
}
const args = parseArgs(process.argv.slice(2));
if (args.help) {
console.log('Usage: node scripts/build-agent-image.mjs [--engine docker|podman] [--image <ref>] [--no-cache]');
process.exit(0);
}
const engine = resolveEngine(args.engine);
const buildArgs = ['build', '-f', DOCKERFILE, '-t', args.image];
if (args.noCache) buildArgs.push('--no-cache');
buildArgs.push(REPO_ROOT);
console.log(`[build-agent-image] ${engine} ${buildArgs.join(' ')}`);
const child = spawn(engine, buildArgs, { stdio: 'inherit' });
child.on('exit', (code) => {
if (code === 0) {
console.log(`\n[build-agent-image] built ${args.image}. Docker cases can now launch.`);
} else {
console.error(`\n[build-agent-image] build failed (exit ${code}).`);
}
process.exit(code ?? 1);
});
+3
View File
@@ -45,6 +45,7 @@ run('copy template', 'cp src/templates/case-template.md dist/templates/');
run('xterm css', 'cp node_modules/@xterm/xterm/css/xterm.css dist/web/public/vendor/');
run('xterm js', 'npx esbuild node_modules/@xterm/xterm/lib/xterm.js --minify --outfile=dist/web/public/vendor/xterm.min.js');
run('xterm-addon-fit', 'npx esbuild node_modules/@xterm/addon-fit/lib/addon-fit.js --minify --outfile=dist/web/public/vendor/xterm-addon-fit.min.js');
run('xterm-addon-serialize', 'npx esbuild node_modules/@xterm/addon-serialize/lib/addon-serialize.js --minify --outfile=dist/web/public/vendor/xterm-addon-serialize.min.js');
run('xterm-addon-webgl', 'cp node_modules/@xterm/addon-webgl/lib/addon-webgl.js dist/web/public/vendor/xterm-addon-webgl.min.js');
run('xterm-addon-unicode11', 'npx esbuild node_modules/@xterm/addon-unicode11/lib/addon-unicode11.js --minify --outfile=dist/web/public/vendor/xterm-addon-unicode11.min.js');
run('xterm-zerolag-input', 'npx esbuild packages/xterm-zerolag-input/src/zerolag-input-addon.ts --bundle --minify --format=iife --global-name=XtermZerolagInput --outfile=dist/web/public/vendor/xterm-zerolag-input.js');
@@ -66,6 +67,7 @@ appendFileSync(
// 4. Minify frontend assets
run('minify input-cjk.js', 'npx esbuild dist/web/public/input-cjk.js --minify --outfile=dist/web/public/input-cjk.js --allow-overwrite');
run('minify sanitize-html.js', 'npx esbuild dist/web/public/sanitize-html.js --minify --outfile=dist/web/public/sanitize-html.js --allow-overwrite');
run('minify app.js', 'npx esbuild dist/web/public/app.js --minify --outfile=dist/web/public/app.js --allow-overwrite');
run('minify terminal-ui.js', 'npx esbuild dist/web/public/terminal-ui.js --minify --outfile=dist/web/public/terminal-ui.js --allow-overwrite');
run('minify respawn-ui.js', 'npx esbuild dist/web/public/respawn-ui.js --minify --outfile=dist/web/public/respawn-ui.js --allow-overwrite');
@@ -89,6 +91,7 @@ console.log('\n[build] content-hash cache busting');
'notification-manager.js',
'keyboard-accessory.js',
'input-cjk.js',
'sanitize-html.js',
'app.js',
'terminal-ui.js',
'respawn-ui.js',
+482
View File
@@ -0,0 +1,482 @@
#!/usr/bin/env node
/**
* capture-readme-gifs.mjs
*
* Deterministic README GIFs — no real server, Claude CLI, or tmux. Reuses the
* mock-injection pipeline from capture-readme-screenshots.mjs (static file
* server + page.route mocks), drives a scripted timeline in the page, records
* it with Playwright video, and converts to GIF via ffmpeg palette encoding.
*
* Scenes:
* 1. subagent-demo.gif — terminal spawns 3 parallel agents; floating agent
* windows open one by one and stream tool-call activity live (driven
* through the real _onSubagentDiscovered/_onSubagentToolCall handlers).
* 2. zerolag-demo.gif — side-by-side typing: instant local echo (zerolag)
* vs bursty ~350 ms server echo, rendered with the vendored xterm.
*
* Usage: node scripts/capture-readme-gifs.mjs
* SCREENSHOT_OUT_DIR=/path/to/review node scripts/capture-readme-gifs.mjs
* Output: docs/images/ (or flat into SCREENSHOT_OUT_DIR)
* Requires: ffmpeg
*/
import { chromium } from 'playwright';
import { execSync } from 'child_process';
import { mkdtempSync, rmSync } from 'fs';
import { tmpdir } from 'os';
import { join } from 'path';
import {
PORT,
SESSION_IDS,
STANDARD_SESSIONS,
buildInitPayload,
startStaticServer,
setupRoutes,
injectState,
outPath,
RST, GRN, YEL, MAG, CYN, GRY, BOLD,
} from './capture-readme-screenshots.mjs';
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const GIF_COLORS = 192;
// ─── ffmpeg conversion (palette recipe from capture-subagent-gif.mjs) ────────
function webmToGif(videoPath, gifPath, { ss, duration, width, fps }) {
// One GLOBAL palette (default stats_mode=full) + ordered dither: per-frame
// palettes (stats_mode=single:new=1) make dirty rectangles visibly mismatch
// on flat dark UI, and error-diffusion dither shimmers between frames.
const filters = `fps=${fps},scale=${width}:-1:flags=lanczos`;
execSync(
`ffmpeg -y -loglevel error -ss ${ss.toFixed(2)} -t ${duration} -i "${videoPath}" ` +
`-vf "${filters},split[s0][s1];[s0]palettegen=max_colors=${GIF_COLORS}:reserve_transparent=0[p];` +
`[s1][p]paletteuse=dither=bayer:bayer_scale=5:diff_mode=rectangle" "${gifPath}"`,
{ stdio: 'inherit' }
);
}
// ─── Scene 1: subagent demo ──────────────────────────────────────────────────
const SUBAGENT_VIEWPORT = { width: 1440, height: 810 };
// Terminal content visible before the agents spawn
const TERMINAL_PRESPAWN = [
'',
`${GRN}●${RST} Working on ${CYN}/home/arkon/codeman-cases/testcase${RST} - I'll use the ${BOLD}Task tool${RST} to spawn parallel agents.`,
'',
`${GRN}●${RST} ${BOLD}Read${RST}(/home/arkon/codeman-cases/testcase/CLAUDE.md)`,
` ${GRY}░${RST} Read ${BOLD}127${RST} lines ${GRY}│${RST} ${CYN}1.2KB${RST}`,
'',
`${GRN}●${RST} ${BOLD}Bash${RST}(find . -name "*.ts" -not -path "*/node_modules/*" | head -20)`,
` ${GRY}░${RST} ./src/index.ts`,
` ${GRY}░${RST} ./src/session.ts`,
` ${GRY}░${RST} ./src/web/server.ts`,
` ${GRY}░${RST} ${GRY}... (17 more)${RST}`,
'',
`${GRN}●${RST} I'll spawn 3 parallel research agents to analyze different parts of the codebase simultaneously.`,
'',
].join('\r\n');
function makeAgent(agentId, description, startedOffsetMs) {
return {
agentId,
sessionId: 'claude-sess-w1-0001',
projectHash: 'abc123',
filePath: `/tmp/${agentId}.jsonl`,
startedAt: new Date(Date.now() - startedOffsetMs).toISOString(),
lastActivityAt: Date.now(),
status: 'active',
toolCallCount: 0,
entryCount: 0,
fileSize: 4000,
description,
model: 'claude-haiku-4-5-20251001',
modelShort: 'haiku',
totalInputTokens: 0,
totalOutputTokens: 0,
parentSessionId: SESSION_IDS.w1,
};
}
// Timeline events: t (ms from scene start) + kind
// term — write raw data to the session terminal
// discover — register subagent + open + position its floating window
// tool — stream a tool call into an agent window
// msg — stream an assistant message into an agent window
// complete — flip an agent to completed
function buildSubagentTimeline() {
const T = (lines) => lines.join('\r\n') + '\r\n';
const tool = (t, agentId, name, input) => ({ t, kind: 'tool', agentId, tool: name, input });
const msg = (t, agentId, text) => ({ t, kind: 'msg', agentId, text });
return [
{
t: 600,
kind: 'term',
data: T([
`${GRN}●${RST} ${BOLD}Task${RST}(Find and document all API endpoints in src/)`,
` ${GRY}░${RST} Spawned ${CYN}agent-001${RST} ${GRY}(haiku)${RST}`,
'',
]),
},
{
t: 1000,
kind: 'discover',
agent: makeAgent('agent-001', 'Find and document all API endpoints in src/', 2000),
x: 440, y: 45,
},
tool(1500, 'agent-001', 'Glob', { pattern: 'src/**/*.ts' }),
{
t: 2000,
kind: 'term',
data: T([
`${GRN}●${RST} ${BOLD}Task${RST}(Explore and understand test structure in test/)`,
` ${GRY}░${RST} Spawned ${CYN}agent-002${RST} ${GRY}(haiku)${RST}`,
'',
]),
},
tool(2200, 'agent-001', 'Read', { file_path: '/home/arkon/codeman/src/web/server.ts' }),
{
t: 2500,
kind: 'discover',
agent: makeAgent('agent-002', 'Explore and understand test structure in test/', 1200),
x: 880, y: 45,
},
tool(3000, 'agent-002', 'Glob', { pattern: 'test/**/*.test.ts' }),
{
t: 3300,
kind: 'term',
data: T([
`${GRN}●${RST} ${BOLD}Task${RST}(Analyze TypeScript type definitions in src/types.ts)`,
` ${GRY}░${RST} Spawned ${CYN}agent-003${RST} ${GRY}(haiku)${RST}`,
'',
]),
},
tool(3500, 'agent-001', 'Grep', { pattern: 'app\\.get|app\\.post|app\\.delete', path: 'src/' }),
{
t: 3800,
kind: 'discover',
agent: makeAgent('agent-003', 'Analyze TypeScript type definitions in src/types.ts', 400),
x: 660, y: 400,
},
tool(4100, 'agent-002', 'Read', { file_path: '/home/arkon/codeman/test/respawn-test-utils.ts' }),
{
t: 4500,
kind: 'term',
data: T([
`${MAG}✻${RST} ${YEL}Waiting for agents...${RST} ${GRY}(${BOLD}esc${RST}${GRY} to interrupt · 32s · ↓ 1.7k tokens · thinking)${RST}`,
'',
]),
},
tool(4700, 'agent-003', 'Read', { file_path: '/home/arkon/codeman/src/types.ts' }),
tool(5200, 'agent-001', 'Read', { file_path: '/home/arkon/codeman/src/web/schemas.ts' }),
tool(5600, 'agent-002', 'Read', { file_path: '/home/arkon/codeman/config/vitest.config.ts' }),
tool(6100, 'agent-003', 'Grep', { pattern: 'export (interface|type)', path: 'src/types/' }),
msg(6700, 'agent-001', 'Found 47 API endpoints across server.ts. Documenting REST paths...'),
tool(7100, 'agent-002', 'Grep', { pattern: 'const PORT =', path: 'test/' }),
msg(7700, 'agent-002', 'Analyzing test patterns: MockSession, unique ports, fileParallelism: false...'),
tool(8100, 'agent-003', 'Read', { file_path: '/home/arkon/codeman/src/types/index.ts' }),
msg(8700, 'agent-003', 'Mapped 38 exported interfaces across 15 domain files. Building summary...'),
{
t: 9300,
kind: 'term',
data: T([
`${GRN}●${RST} ${CYN}agent-001${RST}: ${GRY}12 tool calls — Glob, Read(server.ts), Grep(endpoints)...${RST}`,
`${GRN}●${RST} ${CYN}agent-002${RST}: ${GRY}8 tool calls — Glob, Read(test-utils), Read(vitest.config)...${RST}`,
`${GRN}●${RST} ${CYN}agent-003${RST}: ${GRY}7 tool calls — Read(types.ts), Grep(interface)...${RST}`,
'',
]),
},
tool(10100, 'agent-001', 'Glob', { pattern: 'src/web/routes/*.ts' }),
tool(10600, 'agent-002', 'Read', { file_path: '/home/arkon/codeman/test/setup.ts' }),
tool(11100, 'agent-003', 'Grep', { pattern: 'assertNever', path: 'src/' }),
{
t: 11600,
kind: 'term',
data: T([`${GRN}●${RST} ${GRY}171.8k, 13s${RST} ${GRY}│${RST} ${GRY}1.7k tokens${RST} ${GRY}│${RST} ${GRY}thinking${RST}`, '']),
},
];
}
const SUBAGENT_TAIL_HOLD = 2500; // hold the final frame
async function recordSubagentScene(browser, videoDir) {
console.log('\n1/2 Recording subagent-demo...');
const context = await browser.newContext({
viewport: SUBAGENT_VIEWPORT,
deviceScaleFactor: 1,
recordVideo: { dir: videoDir, size: SUBAGENT_VIEWPORT },
});
const recStart = Date.now();
const page = await context.newPage();
page.setDefaultTimeout(30000);
// Start with NO subagents — they appear during the recording
const initPayload = buildInitPayload(STANDARD_SESSIONS);
await setupRoutes(page, initPayload, TERMINAL_PRESPAWN);
await page.goto(`http://localhost:${PORT}`, { waitUntil: 'domcontentloaded' });
await injectState(page, initPayload, TERMINAL_PRESPAWN, SESSION_IDS.w1);
await page.evaluate(() => {
try { window.app?.fitAddon?.fit(); } catch {}
window.app?.terminal?.scrollToBottom();
});
await sleep(500);
const timeline = buildSubagentTimeline();
const totalMs = Math.max(...timeline.map((e) => e.t)) + SUBAGENT_TAIL_HOLD;
const sceneStart = Date.now();
// Run the whole timeline inside the page so events interleave naturally
await page.evaluate((events) => {
const app = window.app;
for (const ev of events) {
setTimeout(() => {
try {
if (ev.kind === 'term') {
app.terminal.write(ev.data);
app.terminal.scrollToBottom();
} else if (ev.kind === 'discover') {
app._onSubagentDiscovered(ev.agent);
app.openSubagentWindow(ev.agent.agentId);
// The spawn animation (400ms) lands on the auto-grid; glide to our tile after it
setTimeout(() => {
const win = app.subagentWindows.get(ev.agent.agentId);
if (win?.element) {
win.element.style.transition = 'left 0.25s ease, top 0.25s ease';
win.element.style.left = `${ev.x}px`;
win.element.style.top = `${ev.y}px`;
}
}, 520);
setTimeout(() => {
const win = app.subagentWindows.get(ev.agent.agentId);
if (win?.element) win.element.style.transition = '';
app.updateConnectionLines();
}, 850);
} else if (ev.kind === 'tool') {
app._onSubagentToolCall({
agentId: ev.agentId,
tool: ev.tool,
input: ev.input,
timestamp: new Date().toISOString(),
});
} else if (ev.kind === 'msg') {
app._onSubagentMessage({
agentId: ev.agentId,
role: 'assistant',
text: ev.text,
timestamp: new Date().toISOString(),
});
} else if (ev.kind === 'complete') {
app._onSubagentCompleted({ agentId: ev.agentId, timestamp: new Date().toISOString() });
}
} catch (err) {
console.error('timeline event failed', ev, err);
}
}, ev.t);
}
}, timeline);
await sleep(totalMs + 500);
await page.close();
const videoPath = await page.video().path();
await context.close();
return {
videoPath,
ss: (sceneStart - recStart) / 1000 - 0.4,
duration: (totalMs + 400) / 1000,
};
}
// ─── Scene 2: zerolag typing comparison ──────────────────────────────────────
const ZEROLAG_VIEWPORT = { width: 1280, height: 470 };
const TYPED_TEXT = 'echo "zero lag typing from anywhere"';
const TYPE_INTERVAL_MS = 110;
const REMOTE_FLUSH_MS = 350; // server-echo pane flushes queued chars in bursts
const ZEROLAG_TAIL_HOLD = 1800;
const ZEROLAG_HTML = `<!DOCTYPE html>
<html>
<head>
<link rel="stylesheet" href="http://localhost:${PORT}/vendor/xterm.css">
<script src="http://localhost:${PORT}/vendor/xterm.min.js"></script>
<style>
* { margin: 0; box-sizing: border-box; }
body {
width: 1280px; height: 470px; background: #0a0a0c;
display: flex; align-items: center; justify-content: center; gap: 48px;
font-family: -apple-system, 'Segoe UI', Roboto, sans-serif;
}
.pane { width: 560px; }
.card {
background: #131316; border: 1px solid rgba(255,255,255,0.08);
border-radius: 10px; overflow: hidden;
box-shadow: 0 8px 32px rgba(0,0,0,0.45);
}
.card-head {
display: flex; align-items: baseline; gap: 10px;
padding: 12px 16px; border-bottom: 1px solid rgba(255,255,255,0.06);
}
.dot { width: 9px; height: 9px; border-radius: 50%; align-self: center; }
.title { font-size: 15px; font-weight: 600; color: #e8e8ea; }
.sub { font-size: 12.5px; color: #8b8b92; }
.term { padding: 16px 8px 12px 16px; height: 165px; }
.good .dot { background: #22c55e; box-shadow: 0 0 8px rgba(34,197,94,0.7); }
.bad .dot { background: #ef4444; box-shadow: 0 0 8px rgba(239,68,68,0.7); }
.tag {
margin-top: 14px; text-align: center; font-size: 14.5px; color: #7e7e86;
}
.tag b { color: #22c55e; font-weight: 600; }
.bad-tag b { color: #ef4444; }
</style>
</head>
<body>
<div class="pane">
<div class="card good">
<div class="card-head">
<span class="dot"></span>
<span class="title">With zerolag-input</span>
<span class="sub">instant local echo</span>
</div>
<div class="term" id="termLeft"></div>
</div>
<div class="tag">keystrokes echo in <b>0 ms</b></div>
</div>
<div class="pane">
<div class="card bad">
<div class="card-head">
<span class="dot"></span>
<span class="title">Without</span>
<span class="sub">server round-trip echo</span>
</div>
<div class="term" id="termRight"></div>
</div>
<div class="tag bad-tag">keystrokes echo after <b>~350 ms</b></div>
</div>
</body>
</html>`;
async function recordZerolagScene(browser, videoDir) {
console.log('\n2/2 Recording zerolag-demo...');
const context = await browser.newContext({
viewport: ZEROLAG_VIEWPORT,
deviceScaleFactor: 1,
recordVideo: { dir: videoDir, size: ZEROLAG_VIEWPORT },
});
const recStart = Date.now();
const page = await context.newPage();
page.setDefaultTimeout(30000);
await page.setContent(ZEROLAG_HTML, { waitUntil: 'load' });
await page.waitForFunction(() => typeof Terminal !== 'undefined');
await page.evaluate(() => {
const theme = {
background: '#131316',
foreground: '#e8e8ea',
cursor: '#22c55e',
cursorAccent: '#131316',
};
const mk = (id) => {
const term = new Terminal({
cols: 44,
rows: 5,
fontSize: 20,
fontFamily: "'SF Mono', 'Cascadia Code', Menlo, monospace",
cursorBlink: true,
cursorStyle: 'block',
theme,
});
term.open(document.getElementById(id));
term.write('\x1b[32m❯\x1b[0m ');
return term;
};
window.termLeft = mk('termLeft');
window.termRight = mk('termRight');
});
await sleep(600);
const sceneStart = Date.now();
const typingMs = TYPED_TEXT.length * TYPE_INTERVAL_MS;
const totalMs = typingMs + REMOTE_FLUSH_MS + ZEROLAG_TAIL_HOLD;
await page.evaluate(
({ text, interval, flushEvery }) => {
let i = 0;
const remoteQueue = [];
const typer = setInterval(() => {
if (i >= text.length) { clearInterval(typer); return; }
const ch = text[i++];
window.termLeft.write(ch); // local echo: instant
remoteQueue.push(ch); // server echo: waits for the round-trip
}, interval);
const flusher = setInterval(() => {
if (remoteQueue.length) window.termRight.write(remoteQueue.splice(0).join(''));
if (i >= text.length && remoteQueue.length === 0) clearInterval(flusher);
}, flushEvery);
},
{ text: TYPED_TEXT, interval: TYPE_INTERVAL_MS, flushEvery: REMOTE_FLUSH_MS }
);
await sleep(totalMs + 400);
await page.close();
const videoPath = await page.video().path();
await context.close();
return {
videoPath,
ss: (sceneStart - recStart) / 1000 - 0.6, // small lead-in with idle cursors
duration: (totalMs + 600) / 1000,
};
}
// ─── Main ────────────────────────────────────────────────────────────────────
async function main() {
console.log('='.repeat(60));
console.log('Codeman README GIF Capture');
console.log('='.repeat(60));
const server = await startStaticServer();
const videoDir = mkdtempSync(join(tmpdir(), 'codeman-gifs-'));
let browser;
try {
browser = await chromium.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', '--disable-gpu'],
});
const sub = await recordSubagentScene(browser, videoDir);
const subGif = outPath('images', 'subagent-demo.gif');
webmToGif(sub.videoPath, subGif, { ss: Math.max(0, sub.ss), duration: sub.duration, width: 960, fps: 8 });
console.log(` Saved: ${subGif}`);
const zl = await recordZerolagScene(browser, videoDir);
const zlGif = outPath('images', 'zerolag-demo.gif');
webmToGif(zl.videoPath, zlGif, { ss: Math.max(0, zl.ss), duration: zl.duration, width: 900, fps: 10 });
console.log(` Saved: ${zlGif}`);
console.log('\nDone.');
} catch (err) {
console.error('\nFatal error:', err.message);
console.error(err.stack);
process.exitCode = 1;
} finally {
if (browser) await browser.close().catch(() => {});
server.close();
rmSync(videoDir, { recursive: true, force: true });
}
}
process.on('SIGINT', () => process.exit(1));
main();
+253
View File
@@ -0,0 +1,253 @@
#!/usr/bin/env node
/**
* capture-readme-real.mjs
*
* Captures README desktop scenes (multi-session dashboard, monitor, subagent
* windows) from a REAL Codeman instance — intended to run against an ISOLATED
* dev/beta instance (CODEMAN_INSTANCE=beta on :5000) seeded from prod's settings,
* NOT prod itself (never touch prod's live sessions).
*
* Reuses the high-quality capture recipe proven in capture-real-overview.mjs:
* - DSF=2 + ?nowebgl → crisp retina at the TRUE font size (WebGL doubles
* glyphs under DSF=2; the DOM renderer respects devicePixelRatio).
* - per-device localStorage seeding so the capture matches a real device.
*
* SCENE=dashboard|monitor|subagent|all BASE=http://localhost:5000 \
* OUT=screenshots-readme-real/desktop node scripts/capture-readme-real.mjs
*/
import { chromium } from 'playwright';
import { mkdirSync } from 'fs';
import { join } from 'path';
const BASE = process.env.BASE || 'http://localhost:5000';
const OUT = process.env.OUT || 'screenshots-readme-real/desktop';
const SKIN = process.env.SKIN || 'daylight-blue';
const SCENE = process.env.SCENE || 'all';
const FONT = Math.max(10, Math.min(24, Number(process.env.FONT || 13)));
const VIEWPORT = { width: Number(process.env.VW || 1280), height: Number(process.env.VH || 720) };
const DSF = Number(process.env.DSF || 2);
const PLAN_USAGE = process.env.PLAN_USAGE !== '0';
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const url = (extra = '') => {
const sep = BASE.includes('?') ? '&' : '?';
const params = [];
if (DSF > 1) params.push('nowebgl'); // DOM renderer → correct font size at DSF>1
if (extra) params.push(extra);
return params.length ? `${BASE}${sep}${params.join('&')}` : BASE;
};
async function newCtx(browser) {
const context = await browser.newContext({
viewport: VIEWPORT,
deviceScaleFactor: DSF,
ignoreHTTPSErrors: BASE.startsWith('https'),
});
const page = await context.newPage();
page.setDefaultTimeout(30000);
await page.addInitScript(
([skin, planUsage, font]) => {
try {
localStorage.setItem('codeman:skin', skin);
localStorage.setItem('codeman-font-size', String(font));
const blob = { skin, showFileBrowser: false, showProjectInsights: false };
// Don't auto-hide subagent windows that belong to a non-active tab — the
// subagent scene re-homes agents and needs both windows visible at once.
blob.subagentActiveTabOnly = false;
if (planUsage) blob.showPlanUsageLimits = true;
localStorage.setItem('codeman-app-settings', JSON.stringify(blob));
} catch {
/* ignore */
}
},
[SKIN, PLAN_USAGE, FONT]
);
return { context, page };
}
async function bootstrap(page) {
await page.waitForFunction(() => window.app && window.app.terminal, { timeout: 20000 });
await sleep(1200);
}
async function listSessions(page) {
return page.evaluate(() =>
Array.from(window.app.sessions.values()).map((s) => ({ id: s.id, name: s.name, mode: s.mode }))
);
}
async function shoot(page, name) {
const out = join(OUT, name);
await page.evaluate((f) => {
try {
if (window.app.setFontSize) window.app.setFontSize(f);
} catch {}
try {
window.app.fitAddon && window.app.fitAddon.fit();
} catch {}
try {
window.app.applyHeaderVisibilitySettings && window.app.applyHeaderVisibilitySettings();
} catch {}
}, FONT);
await sleep(1500);
await page.screenshot({ path: out, fullPage: false });
console.log(' Saved: ' + out);
}
async function sceneDashboard(browser) {
console.log('Scene: dashboard');
const { context, page } = await newCtx(browser);
await page.goto(url(), { waitUntil: 'domcontentloaded' });
await bootstrap(page);
const sessions = await listSessions(page);
// Select a claude session so the active terminal shows rich content; all tabs render.
const target = sessions.find((s) => s.mode === 'claude') || sessions[0];
if (target) await page.evaluate((id) => window.app.selectSession(id), target.id);
await sleep(4000);
await shoot(page, 'multi-session-dashboard.png');
await context.close();
}
async function sceneMonitor(browser) {
console.log('Scene: monitor');
const { context, page } = await newCtx(browser);
await page.goto(url(), { waitUntil: 'domcontentloaded' });
await bootstrap(page);
const sessions = await listSessions(page);
const target = sessions.find((s) => s.mode === 'claude') || sessions[0];
if (target) await page.evaluate((id) => window.app.selectSession(id), target.id);
await sleep(2500);
// toggleMonitorPanel() opens the panel, clears the hidden state, loads REAL
// mux sessions (/api/mux), starts stats, and renders the task panel.
await page.evaluate(async () => {
try {
await window.app.toggleMonitorPanel();
} catch {}
});
await sleep(3000);
await shoot(page, 'multi-session-monitor.png');
await context.close();
}
async function sceneSubagent(browser) {
console.log('Scene: subagent');
const { context, page } = await newCtx(browser);
await page.goto(url(), { waitUntil: 'domcontentloaded' });
await bootstrap(page);
// Select the session whose subagents we want (subagentActiveTabOnly means
// app.subagents only fills for the active tab). Prefer SUBAGENT_SID env.
const sessions = await listSessions(page);
const targetId = process.env.SUBAGENT_SID || (sessions.find((s) => s.mode === 'claude') || sessions[0])?.id;
if (targetId) await page.evaluate((id) => window.app.selectSession(id), targetId);
// Wait (up to ~45s) for live subagents to arrive via SSE into app.subagents.
let agents = [];
for (let i = 0; i < 45; i++) {
agents = await page.evaluate(() =>
Array.from(window.app.subagents?.entries?.() || []).map(([id, a]) => ({ id, name: a.name ?? a.agentType ?? '' }))
);
if (agents.length >= 1) break;
await sleep(1000);
}
console.log(' live in-browser subagents:', JSON.stringify(agents));
if (agents.length === 0) {
console.log(' NO live subagents — skipping (stage a longer subagent task and run this while it runs).');
await context.close();
return;
}
// The window body renders from app.subagentActivity, which fills ONLY from live
// SSE tool-call/progress events — a fresh client never gets past activity replayed.
// So sit connected and wait for live activity to accumulate, then open the two
// agents that actually have content (otherwise the windows read "No activity yet").
let active = [];
for (let i = 0; i < 100; i++) {
active = await page.evaluate(() =>
Array.from(window.app.subagentActivity?.entries?.() || [])
.filter(([, arr]) => Array.isArray(arr) && arr.length >= 1)
.map(([id, arr]) => ({ id, n: arr.length }))
.sort((a, b) => b.n - a.n)
);
if (active.length >= 2) break;
// xhigh-effort agents churn in bursts between long thinking pauses, so be
// patient (~150s); accept a single populated window after ~45s if that's all.
if (i >= 30 && active.length >= 1) break;
await sleep(1500);
}
console.log(' agents with live activity:', JSON.stringify(active));
const openIds = (active.length ? active : agents).map((a) => a.id);
// Capture-only DOM nudge: on fresh dev sessions, a tab's claudeSessionId stays the
// Codeman id and never becomes the real Claude conversation UUID, so the window
// open-gate (claudeSessionId === agent.sessionId) + the activeTabOnly hide rule both
// fail. Re-home the chosen agents onto the active tab and align its claudeSessionId
// to the agents' (shared) sessionId so the windows open AND show their live activity.
await page.evaluate(
(ids) => {
const activeId = window.app.activeSessionId;
const tab = window.app.sessions.get(activeId);
ids.slice(0, 2).forEach((id) => {
const a = window.app.subagents.get(id);
if (!a) return;
a.parentSessionId = activeId;
if (tab && a.sessionId) tab.claudeSessionId = a.sessionId;
});
},
openIds
);
await page.evaluate(
(ids) => {
ids.slice(0, 2).forEach((id) => {
try {
window.app.openSubagentWindow(id);
} catch {}
});
},
openIds
);
await sleep(2000);
await page.evaluate(() => {
// Viewport-relative tiling: center two subagent windows over the terminal so
// the layout adapts to whatever VW/VH the capture uses (e.g. the HQ 1100×650
// recipe) instead of overflowing at narrower widths.
const wins = Array.from(window.app.subagentWindows.values());
const W = window.innerWidth;
const H = window.innerHeight;
const winW = Math.min(440, Math.floor((W - 60) / 2 - 10));
const winH = Math.min(360, Math.floor(H * 0.56));
const top = Math.floor(H * 0.16);
const gap = 16;
const totalW = winW * 2 + gap;
const startLeft = Math.max(16, Math.floor((W - totalW) / 2));
wins.slice(0, 2).forEach((win, i) => {
const el = win.element;
// Force visible: a freshly opened window may be hidden by the activeTabOnly
// rule before we override it (we also seed subagentActiveTabOnly:false).
win.hidden = false;
win.minimized = false;
el.style.display = 'flex';
el.style.left = startLeft + i * (winW + gap) + 'px';
el.style.top = top + 'px';
el.style.width = winW + 'px';
el.style.height = winH + 'px';
});
});
await sleep(1500);
await shoot(page, 'subagent-spawn.png');
await context.close();
}
async function main() {
mkdirSync(OUT, { recursive: true });
const browser = await chromium.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', '--disable-gpu'],
});
console.log(`BASE=${BASE} SKIN=${SKIN} DSF=${DSF} VIEWPORT=${VIEWPORT.width}x${VIEWPORT.height} SCENE=${SCENE}`);
if (SCENE === 'dashboard' || SCENE === 'all') await sceneDashboard(browser);
if (SCENE === 'monitor' || SCENE === 'all') await sceneMonitor(browser);
if (SCENE === 'subagent' || SCENE === 'all') await sceneSubagent(browser);
await browser.close();
}
main().catch((e) => {
console.error('FATAL', e.message);
process.exit(1);
});
File diff suppressed because it is too large Load Diff
+152
View File
@@ -0,0 +1,152 @@
#!/usr/bin/env node
/**
* capture-real-overview.mjs
*
* Captures a REAL claude-overview screenshot from a LIVE Codeman server
* (no mock injection). Drive a real session to do real work, then run:
*
* SID=<sessionId> BASE=http://localhost:5000 OUT=screenshots-real \
* node scripts/capture-real-overview.mjs
*
* Skin defaults to daylight-blue (prod default) via the localStorage pre-paint
* contract in index.html. Output: <OUT>/claude-overview.png at 1280x720 (DSF 2).
*/
import { chromium } from 'playwright';
import { mkdirSync } from 'fs';
import { join } from 'path';
const SID = process.env.SID;
const BASE = process.env.BASE || 'http://localhost:5000';
const OUT = process.env.OUT || 'screenshots-real';
const SKIN = process.env.SKIN || 'daylight-blue';
// Unique filename per run (timestamped) so a viewer holding an old render of a
// fixed path can never shadow a fresh capture. Override with NAME=… if needed.
const STAMP = new Date().toISOString().replace(/[:.]/g, '-').replace('T', '_').slice(0, 19);
const NAME = process.env.NAME || `claude-overview-${STAMP}.png`;
const VIEWPORT = { width: Number(process.env.VW || 1512), height: Number(process.env.VH || 812) };
// IMPORTANT: default deviceScaleFactor is 1, NOT 2. xterm's WebGL renderer in
// headless Chromium draws terminal glyphs at ~2× their nominal size when DSF=2
// (while still reporting nominal 8px cell dims internally, so it can't be caught
// by measuring terminal.cols/cell — only the pixels reveal it). The HTML chrome
// is unaffected, so DSF=2 makes ONLY the console font look comically large. DSF=1
// renders the console at its true size, matching a real (non-headless) browser.
const DSF = Number(process.env.DSF || 1);
if (!SID) {
console.error('SID env var required (the live session id to screenshot)');
process.exit(1);
}
const sleep = (ms) => new Promise((r) => setTimeout(r, ms));
const main = async () => {
mkdirSync(OUT, { recursive: true });
const browser = await chromium.launch({
headless: true,
args: ['--no-sandbox', '--disable-setuid-sandbox', '--disable-dev-shm-usage', '--disable-gpu'],
});
const context = await browser.newContext({
viewport: VIEWPORT,
deviceScaleFactor: DSF,
ignoreHTTPSErrors: BASE.startsWith('https'),
});
const page = await context.newPage();
page.setDefaultTimeout(30000);
// Force the skin before any page script runs (pre-paint <head> contract), and
// seed the PER-DEVICE display blob so the capture reflects what prod actually
// shows on the user's real device — notably the plan-usage chip, which is a
// per-device setting (default OFF) deleted from the server payload, so a fresh
// browser would otherwise hide it. PLAN_USAGE=0 disables.
const PLAN_USAGE = process.env.PLAN_USAGE !== '0';
// Terminal console font size. App default is 14px; a fresh headless browser has
// no saved codeman-font-size, so it renders at 14 — much larger than a real
// device where the console has been zoomed down. Seed a smaller value (clamped
// to the app's [10,24] range) so the console font looks normal in the capture.
const FONT = Math.max(10, Math.min(24, Number(process.env.FONT || 14)));
await page.addInitScript(
([skin, planUsage, font]) => {
try {
localStorage.setItem('codeman:skin', skin);
localStorage.setItem('codeman-font-size', String(font));
// Desktop app-settings blob (settings-ui.js getSettingsStorageKey()).
// Present these display keys explicitly so the server merge won't seed
// side panels open (display keys only seed from server when absent from
// localStorage). Matches the clean full-width-terminal reference look.
const blob = {
skin,
showFileBrowser: false,
showMonitor: false,
showSubagents: false,
showProjectInsights: false,
};
if (planUsage) blob.showPlanUsageLimits = true;
localStorage.setItem('codeman-app-settings', JSON.stringify(blob));
} catch {
/* ignore */
}
},
[SKIN, PLAN_USAGE, FONT]
);
// At DSF>1, xterm's WebGL renderer draws glyphs at ~2x (see DSF comment above).
// The app honors a `?nowebgl` URL param that switches to xterm's DOM renderer,
// which respects devicePixelRatio correctly — so DSF=2 + nowebgl yields a crisp
// 2x (retina) capture at the TRUE font size. Auto-enable it whenever DSF>1.
const url = DSF > 1 ? `${BASE}${BASE.includes('?') ? '&' : '?'}nowebgl` : BASE;
console.log(`Loading ${url} (DSF=${DSF}) ...`);
await page.goto(url, { waitUntil: 'domcontentloaded' });
await page.waitForFunction(() => window.app && window.app.terminal, { timeout: 20000 });
await sleep(1500);
console.log(`Selecting session ${SID} ...`);
await page.evaluate((sid) => window.app.selectSession(sid), SID);
// Let the terminal buffer stream in + xterm render + any Ink redraw settle.
await sleep(2000);
// Force a clean fit (avoids capturing a transient pre-fit frame where the
// terminal renders at the wrong column count) and re-apply per-device header
// visibility so the seeded plan-usage chip is shown.
await page.evaluate((font) => {
// Force the console font explicitly (setFontSize also re-fits) in case
// loadFontSize didn't pick up the seeded value before the session rendered.
try {
if (window.app.setFontSize) window.app.setFontSize(font);
else window.app.terminal.options.fontSize = font;
} catch {}
try {
window.app.fitAddon && window.app.fitAddon.fit();
} catch {}
try {
window.dispatchEvent(new Event('resize'));
} catch {}
try {
window.app.applyHeaderVisibilitySettings && window.app.applyHeaderVisibilitySettings();
} catch {}
}, FONT);
await sleep(3000);
// Optionally scroll the terminal up to frame the rich tool-call region
// (Read/Write/Bash + green test results) instead of the trailing summary.
const SCROLL = Number(process.env.SCROLL || 0);
if (SCROLL) {
await page.evaluate((n) => {
const t = window.app && window.app.terminal;
if (t && t.scrollLines) t.scrollLines(-n);
}, SCROLL);
await sleep(800);
}
const outPath = join(OUT, NAME);
await page.screenshot({ path: outPath, fullPage: false });
console.log(`Saved: ${outPath}`);
await context.close();
await browser.close();
};
main().catch((e) => {
console.error('FATAL', e.message);
process.exit(1);
});
+3
View File
@@ -252,6 +252,7 @@ if (isGlobalInstall) {
const require = createRequire(import.meta.url);
const xtermDir = join(require.resolve('@xterm/xterm'), '..', '..');
const fitDir = join(require.resolve('@xterm/addon-fit'), '..', '..');
const serializeDir = join(require.resolve('@xterm/addon-serialize'), '..', '..');
const webglDir = join(require.resolve('@xterm/addon-webgl'), '..', '..');
const unicode11Dir = join(require.resolve('@xterm/addon-unicode11'), '..', '..');
const vendorDir = join(srcDir, 'web', 'public', 'vendor');
@@ -264,12 +265,14 @@ if (isGlobalInstall) {
try {
execSync(`npx esbuild "${join(xtermDir, 'lib', 'xterm.js')}" --minify --outfile="${join(vendorDir, 'xterm.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(fitDir, 'lib', 'addon-fit.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-fit.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(serializeDir, 'lib', 'addon-serialize.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-serialize.min.js')}"`, { stdio: 'pipe' });
execSync(`npx esbuild "${join(unicode11Dir, 'lib', 'addon-unicode11.js')}" --minify --outfile="${join(vendorDir, 'xterm-addon-unicode11.min.js')}"`, { stdio: 'pipe' });
console.log(colors.green('✓ xterm vendor files copied to src/web/public/vendor/'));
} catch {
// Fallback: copy unminified
copyFileSync(join(xtermDir, 'lib', 'xterm.js'), join(vendorDir, 'xterm.min.js'));
copyFileSync(join(fitDir, 'lib', 'addon-fit.js'), join(vendorDir, 'xterm-addon-fit.min.js'));
copyFileSync(join(serializeDir, 'lib', 'addon-serialize.js'), join(vendorDir, 'xterm-addon-serialize.min.js'));
copyFileSync(join(unicode11Dir, 'lib', 'addon-unicode11.js'), join(vendorDir, 'xterm-addon-unicode11.min.js'));
console.log(colors.green('✓ xterm vendor files copied') + colors.dim(' (unminified — esbuild not available)'));
}
+15
View File
@@ -31,6 +31,7 @@ export PUPPETEER_SKIP_DOWNLOAD="${PUPPETEER_SKIP_DOWNLOAD:-1}"
REPO=""
TAG=""
SUPERVISOR="none"
SERVER_PID=""
STATUS_FILE=""
UPDATE_ID=""
FROM_VERSION=""
@@ -50,6 +51,7 @@ while [[ $# -gt 0 ]]; do
--node) NODE="$2"; shift 2 ;;
--log) LOG="$2"; shift 2 ;;
--prev-sha) PREV_SHA="$2"; shift 2 ;;
--server-pid) SERVER_PID="$2"; shift 2 ;;
--stash) DO_STASH=1; shift ;;
*) shift ;;
esac
@@ -198,6 +200,19 @@ case "$SUPERVISOR" in
|| fail "Build succeeded but launchd restart failed" "launchctl"
}
;;
launchd-daemon)
# System-level KeepAlive LaunchDaemon (headless Mac): kickstarting the system
# domain needs root, but we don't need it — kill the server and launchd
# respawns it on the new dist/ within ThrottleInterval seconds.
if [[ -n "$SERVER_PID" ]] && kill "$SERVER_PID" 2>/dev/null; then
: # respawn is launchd's job from here
else
MANUAL_CMD="sudo launchctl kickstart -k system/com.codeman.web"
write_status "completed-needs-manual-restart" "Update staged — restart Codeman to apply v$TO_VERSION."
echo "[self-update] launchd-daemon: could not signal server pid '$SERVER_PID' — manual restart required"
exit 0
fi
;;
*)
MANUAL_CMD="pkill -f 'codeman.*web'; codeman web &"
write_status "completed-needs-manual-restart" "Update staged — restart Codeman to apply v$TO_VERSION."
+101
View File
@@ -0,0 +1,101 @@
/**
* @fileoverview Parses terminal magic links that request attachment cards.
*/
import { isAbsolute } from 'node:path';
import { fileURLToPath } from 'node:url';
import { isSupportedAttachmentExtension } from './attachment-registry.js';
import { stripAnsi } from './utils/index.js';
const MAGIC_LINK_RE = /codeman:\/\/attach\?([^\s<>"']+)/g;
const CODEX_SAVED_FILE_RE = /\bSaved to:\s*(file:\/\/[^\s<>"']+)/gi;
export interface TerminalAttachmentRequest {
path: string;
source: 'external' | 'codex-generated';
}
export interface ParseTerminalAttachmentOptions {
/**
* Enable the Codex `Saved to: file://...` scanner. Only codex-mode sessions
* may set this — the relaxed codex-generated trust policy must never be
* reachable from other modes' (prompt-injectable) terminal output.
*/
codexArtifacts?: boolean;
}
export function parseAttachmentMagicLinks(data: string): string[] {
return parseMagicAttachmentRequests(data).map((request) => request.path);
}
export function parseTerminalAttachmentRequests(
data: string,
options: ParseTerminalAttachmentOptions = {}
): TerminalAttachmentRequest[] {
const results: TerminalAttachmentRequest[] = [];
const seen = new Set<string>();
const requests = options.codexArtifacts
? [...parseMagicAttachmentRequests(data), ...parseCodexGeneratedArtifactRequests(data)]
: parseMagicAttachmentRequests(data);
for (const request of requests) {
const key = `${request.source}:${request.path}`;
if (seen.has(key)) continue;
seen.add(key);
results.push(request);
}
return results;
}
function parseMagicAttachmentRequests(data: string): TerminalAttachmentRequest[] {
const results: string[] = [];
const seen = new Set<string>();
for (const match of data.matchAll(MAGIC_LINK_RE)) {
const query = trimTrailingPunctuation(match[1] || '');
try {
const params = new URLSearchParams(query);
const filePath = params.get('path');
if (!filePath || !isAbsolute(filePath)) continue;
const extension = filePath.split('.').pop()?.toLowerCase() || '';
if (!isSupportedAttachmentExtension(extension)) continue;
if (seen.has(filePath)) continue;
seen.add(filePath);
results.push(filePath);
} catch {
// Ignore malformed terminal text. Magic links are advisory.
}
}
return results.map((path) => ({ path, source: 'external' }));
}
function parseCodexGeneratedArtifactRequests(data: string): TerminalAttachmentRequest[] {
const results: TerminalAttachmentRequest[] = [];
const seen = new Set<string>();
// Codex styles its TUI output — strip ANSI first so a trailing SGR reset
// (e.g. `...mockup.png\x1b[0m`) doesn't ride into the captured URL and break
// the extension allowlist check.
for (const match of stripAnsi(data).matchAll(CODEX_SAVED_FILE_RE)) {
const rawUrl = trimTrailingPunctuation(match[1] || '');
try {
const filePath = fileURLToPath(rawUrl);
if (!isAbsolute(filePath)) continue;
const extension = filePath.split('.').pop()?.toLowerCase() || '';
if (!isSupportedAttachmentExtension(extension)) continue;
if (seen.has(filePath)) continue;
seen.add(filePath);
results.push({ path: filePath, source: 'codex-generated' });
} catch {
// Ignore malformed terminal text. Generated-artifact links are advisory.
}
}
return results;
}
function trimTrailingPunctuation(value: string): string {
return value.replace(/[),.;:]+$/g, '');
}
+248
View File
@@ -0,0 +1,248 @@
/**
* @fileoverview In-memory attachment registry for live external document references.
*
* Session-local files keep using the existing workspace-scoped file routes. This
* registry is only for explicit, live external attachments that need a stable ID
* so browser requests never contain arbitrary absolute paths.
*/
import { randomUUID } from 'node:crypto';
import { realpathSync } from 'node:fs';
import fs from 'node:fs/promises';
import { basename, extname, isAbsolute } from 'node:path';
import { isBlockedAttachmentPath, loadAttachmentGuardConfig } from './config/attachment-guard.js';
import { validateSessionFilePath } from './web/route-helpers.js';
import type { AttachmentDetectedEvent, AttachmentDetectedType } from './types.js';
const SUPPORTED_ATTACHMENT_EXTENSIONS = new Set([
'png',
'jpg',
'jpeg',
'gif',
'webp',
'pdf',
'docx',
'pptx',
'md',
'txt',
]);
export type AttachmentSource = 'detected' | 'external';
export interface AttachmentRecord {
attachmentId: string;
sessionId: string;
filePath: string;
fileName: string;
extension: string;
attachmentType: AttachmentDetectedType;
size: number;
mtimeMs: number;
timestamp: number;
source: AttachmentSource;
}
export interface AttachmentRegistrationResult extends AttachmentDetectedEvent {
attachmentId: string;
source: AttachmentSource;
rawUrl: string;
previewUrl: string;
thumbnailUrl: string;
}
export class AttachmentRegistrationError extends Error {
constructor(
message: string,
readonly statusCode: number = 400
) {
super(message);
}
}
/** Per-session attachment cap. Bounds memory against a client (or a
* prompt-injected magic-link flood) registering unbounded distinct paths. */
const MAX_ATTACHMENTS_PER_SESSION = 200;
class AttachmentRegistry {
private recordsBySession = new Map<string, Map<string, AttachmentRecord>>();
register(record: AttachmentRecord): void {
let records = this.recordsBySession.get(record.sessionId);
if (!records) {
records = new Map();
this.recordsBySession.set(record.sessionId, records);
}
records.set(record.attachmentId, record);
// Evict oldest (insertion-order) entries beyond the cap.
while (records.size > MAX_ATTACHMENTS_PER_SESSION) {
const oldest = records.keys().next().value;
if (oldest === undefined) break;
records.delete(oldest);
}
}
get(sessionId: string, attachmentId: string): AttachmentRecord | undefined {
return this.recordsBySession.get(sessionId)?.get(attachmentId);
}
findByFilePath(sessionId: string, filePath: string): AttachmentRecord | undefined {
const records = this.recordsBySession.get(sessionId);
if (!records) return undefined;
for (const record of records.values()) {
if (record.filePath === filePath) return record;
}
return undefined;
}
clearSession(sessionId: string): void {
this.recordsBySession.delete(sessionId);
}
}
export const attachmentRegistry = new AttachmentRegistry();
export function isSupportedAttachmentExtension(extension: string): boolean {
return SUPPORTED_ATTACHMENT_EXTENSIONS.has(extension.toLowerCase().replace(/^\./, ''));
}
export function getAttachmentType(extension: string): AttachmentDetectedType {
const normalized = extension.toLowerCase().replace(/^\./, '');
if (['png', 'jpg', 'jpeg', 'gif', 'webp'].includes(normalized)) return 'image';
if (normalized === 'pdf') return 'pdf';
if (normalized === 'pptx') return 'presentation';
if (normalized === 'md') return 'markdown';
if (normalized === 'txt') return 'text';
return 'document';
}
export function buildAttachmentRoutes(
sessionId: string,
attachmentId: string
): {
rawUrl: string;
previewUrl: string;
thumbnailUrl: string;
} {
const encodedId = encodeURIComponent(attachmentId);
return {
rawUrl: `/api/sessions/${sessionId}/attachments/${encodedId}/raw`,
previewUrl: `/api/sessions/${sessionId}/attachments/${encodedId}/preview`,
thumbnailUrl: `/api/sessions/${sessionId}/attachments/${encodedId}/thumbnail`,
};
}
export function buildFileThumbnailRoute(sessionId: string, relativePath: string): string {
return `/api/sessions/${sessionId}/file-thumbnail?path=${encodeURIComponent(relativePath)}`;
}
export function attachmentRecordToEvent(record: AttachmentRecord): AttachmentRegistrationResult {
const routes = buildAttachmentRoutes(record.sessionId, record.attachmentId);
return {
sessionId: record.sessionId,
filePath: record.fileName,
relativePath: '',
fileName: record.fileName,
extension: record.extension,
attachmentType: record.attachmentType,
timestamp: record.timestamp,
size: record.size,
attachmentId: record.attachmentId,
source: record.source,
...routes,
};
}
/** Options for {@link registerExternalAttachment}. */
export interface RegisterExternalAttachmentOptions {
/**
* The registering session's working directory. Required to enforce workspace
* confinement — either when the global mode is enabled
* (`attachmentConfineToWorkspace` / `CODEMAN_ATTACHMENT_CONFINE`) or when
* {@link forceWorkspaceConfinement} is set for this call.
*/
sessionWorkingDir?: string;
/**
* Force workspace confinement for THIS registration regardless of the global
* setting. Used by the terminal-output `codeman://attach` magic-link scanner:
* terminal output is attacker-influenceable (a prompt-injected session can
* print an arbitrary path), so passive magic links may only reference files
* inside the session workspace. Deliberate cross-workspace attachment still
* works through the explicit, Origin-guarded `POST /attachments` route and the
* `codeman attach` CLI (which POSTs directly when a session id is known).
*/
forceWorkspaceConfinement?: boolean;
}
export async function registerExternalAttachment(
sessionId: string,
requestedPath: string,
options: RegisterExternalAttachmentOptions = {}
): Promise<AttachmentRegistrationResult> {
if (!requestedPath || !isAbsolute(requestedPath)) {
throw new AttachmentRegistrationError('Attachment path must be an absolute local path');
}
let resolvedPath: string;
try {
resolvedPath = realpathSync(requestedPath);
} catch {
throw new AttachmentRegistrationError('Attachment file not found', 404);
}
// COD-53: enforce the active attachment-guard policy on the symlink-resolved
// path before doing anything else.
const guard = await loadAttachmentGuardConfig();
if (guard.confineToWorkspace || options.forceWorkspaceConfinement) {
// Workspace-confined: the file MUST resolve inside the session's workspace.
// Applies when the global strict mode is on (opt-in, default OFF) OR when
// the caller forces it for this registration (the magic-link scanner — see
// forceWorkspaceConfinement). Strictly more restrictive than the blocklist.
const workingDir = options.sessionWorkingDir;
if (!workingDir || !validateSessionFilePath(workingDir, resolvedPath)) {
throw new AttachmentRegistrationError('Access to this file is blocked', 403);
}
}
// Blocklist (DEFAULT, also applied alongside confinement as defense in
// depth): pre-populated secret locations + the /root and /etc trees + any
// operator-configured extra trees. Symlinks are already resolved above.
// Cross-workspace attachment of non-blocked files stays allowed, so
// codeman-publish and the ~/.codeman review loop keep working.
if (isBlockedAttachmentPath(resolvedPath, guard.blockedTrees)) {
throw new AttachmentRegistrationError('Access to this file is blocked', 403);
}
const extension = extname(resolvedPath).toLowerCase().replace(/^\./, '');
if (!isSupportedAttachmentExtension(extension)) {
throw new AttachmentRegistrationError('Unsupported attachment type');
}
const stat = await fs.stat(resolvedPath);
if (typeof stat.isFile === 'function' && !stat.isFile()) {
throw new AttachmentRegistrationError('Attachment path is not a file');
}
const existing = attachmentRegistry.findByFilePath(sessionId, resolvedPath);
if (existing) {
existing.size = stat.size;
existing.mtimeMs = stat.mtimeMs ?? 0;
existing.timestamp = Date.now();
return attachmentRecordToEvent(existing);
}
const record: AttachmentRecord = {
attachmentId: `att_${randomUUID()}`,
sessionId,
filePath: resolvedPath,
fileName: basename(resolvedPath),
extension,
attachmentType: getAttachmentType(extension),
size: stat.size,
mtimeMs: stat.mtimeMs ?? 0,
timestamp: Date.now(),
source: 'external',
};
attachmentRegistry.register(record);
return attachmentRecordToEvent(record);
}
+123
View File
@@ -10,11 +10,17 @@
import { Command } from 'commander';
import chalk from 'chalk';
import { createRequire } from 'module';
import http from 'node:http';
import https from 'node:https';
import { readFileSync } from 'node:fs';
import { isAbsolute } from 'node:path';
import { dataPath } from './config/instance.js';
import { getSessionManager } from './session-manager.js';
import { getTaskQueue } from './task-queue.js';
import { getRalphLoop } from './ralph-loop.js';
import { getStore } from './state-store.js';
import { getErrorMessage } from './types.js';
import { isSupportedAttachmentExtension } from './attachment-registry.js';
const require = createRequire(import.meta.url);
const pkg = require('../package.json') as { version: string };
@@ -23,6 +29,93 @@ const program = new Command();
program.name('codeman').description('Claude Code session manager with autonomous Ralph Loop').version(pkg.version);
function makeAttachmentMagicLink(filePath: string): string {
return `codeman://attach?path=${encodeURIComponent(filePath)}`;
}
function readCodemanEnv(): Record<string, string> {
const envPath = dataPath('.env');
try {
const text = readFileSync(envPath, 'utf-8');
const result: Record<string, string> = {};
for (const rawLine of text.split(/\r?\n/)) {
const line = rawLine.trim();
if (!line || line.startsWith('#')) continue;
const match = line.match(/^([A-Za-z_][A-Za-z0-9_]*)=(.*)$/);
if (!match) continue;
let value = match[2].trim();
if ((value.startsWith('"') && value.endsWith('"')) || (value.startsWith("'") && value.endsWith("'"))) {
value = value.slice(1, -1);
}
result[match[1]] = value;
}
return result;
} catch {
return {};
}
}
async function postAttachment(apiUrl: string, sessionId: string, filePath: string): Promise<boolean> {
const envFile = readCodemanEnv();
const username = process.env.CODEMAN_USERNAME || envFile.CODEMAN_USERNAME || 'admin';
const password = process.env.CODEMAN_PASSWORD || envFile.CODEMAN_PASSWORD;
const url = new URL(`/api/sessions/${encodeURIComponent(sessionId)}/attachments`, apiUrl);
const body = JSON.stringify({ path: filePath });
const transport = url.protocol === 'https:' ? https : http;
return new Promise((resolve) => {
const headers: Record<string, string | number> = {
Accept: 'application/json',
'Content-Type': 'application/json',
'Content-Length': Buffer.byteLength(body),
};
if (password) {
headers.Authorization = `Basic ${Buffer.from(`${username}:${password}`).toString('base64')}`;
}
const req = transport.request(
{
protocol: url.protocol,
hostname: url.hostname,
port: url.port,
method: 'POST',
path: `${url.pathname}${url.search}`,
rejectUnauthorized: false,
headers,
},
(res) => {
res.resume();
res.on('end', () => resolve(Boolean(res.statusCode && res.statusCode >= 200 && res.statusCode < 300)));
}
);
req.on('error', () => resolve(false));
req.write(body);
req.end();
});
}
program
.command('attach <path>')
.description('Show an attachment card for a local file')
.option('-s, --session <id>', 'Codeman session ID (defaults to CODEMAN_SESSION_ID)')
.option('--url <url>', 'Codeman API URL (defaults to CODEMAN_API_URL or https://127.0.0.1:3000)')
.action(async (filePath, options) => {
const extension = String(filePath).split('.').pop()?.toLowerCase() || '';
if (!isAbsolute(filePath) || !isSupportedAttachmentExtension(extension)) {
console.error(chalk.red('✗ attach requires an absolute path to a png, pdf, docx, pptx, md, or txt file'));
process.exit(1);
}
const sessionId = options.session || process.env.CODEMAN_SESSION_ID;
const apiUrl = options.url || process.env.CODEMAN_API_URL || 'https://127.0.0.1:3000';
if (sessionId && (await postAttachment(apiUrl, sessionId, filePath))) {
console.log(chalk.green('✓ Attachment card requested'));
return;
}
console.log(makeAttachmentMagicLink(filePath));
});
// ============ Session Commands ============
const sessionCmd = program.command('session').alias('s').description('Manage Claude sessions');
@@ -533,4 +626,34 @@ program
}
});
program
.command('doctor')
.alias('check-deps')
.description('Check Codeman tool dependencies (Node, Claude CLI, tmux, LibreOffice, MS Office)')
.option('--json', 'Output structured JSON instead of a table')
.option('--category <name>', 'Only check one category (core|office|other)')
.action(async (options) => {
const { createRealHost, checkAll } = await import('./utils/dependency-checker.js');
const { renderTable, renderJson, computeExitCode } = await import('./utils/dependency-report.js');
const { DEPENDENCY_REGISTRY, TOOL_CATEGORIES } = await import('./config/dependency-registry.js');
if (options.category && !(TOOL_CATEGORIES as readonly string[]).includes(options.category)) {
console.error(`Unknown category "${options.category}". Valid categories: ${TOOL_CATEGORIES.join(', ')}`);
process.exit(2);
}
const host = createRealHost();
const registry = options.category
? DEPENDENCY_REGISTRY.filter((t) => t.category === options.category)
: DEPENDENCY_REGISTRY;
const results = checkAll(registry, host);
if (options.json) {
console.log(JSON.stringify(renderJson(results, host.environment), null, 2));
} else {
console.log(renderTable(results, host.environment));
}
process.exit(computeExitCode(results));
});
export { program };
+133
View File
@@ -0,0 +1,133 @@
/**
* @fileoverview Attachment path-guard configuration (COD-53).
*
* Governs which host files may be registered as cross-workspace attachments
* and served to the browser. Two operator-facing knobs, both with safe
* defaults:
*
* 1. **Blocked-path blocklist (DEFAULT, configurable).** Pre-populated with the
* shared secret-location blocklist (`isSensitivePath`) PLUS the directory
* trees `/root` and `/etc` (anything under them is blocked). The operator
* EXTENDS — never shrinks — this set with additional absolute directory
* trees via the settings key `attachmentBlockedPaths: string[]` and/or the
* env var `CODEMAN_ATTACHMENT_BLOCKED_PATHS` (comma-separated).
*
* 2. **Workspace confinement (OPTIONAL, default OFF).** When enabled, an
* attachment must resolve INSIDE the registering session's workingDir
* (reusing `validateSessionFilePath` containment semantics). This is
* strictly more restrictive than the blocklist and breaks intentional
* cross-workspace attachment (codeman-publish, the ~/.codeman review-card
* loop), so it is OFF by default. Toggle via settings
* `attachmentConfineToWorkspace: boolean` and/or env
* `CODEMAN_ATTACHMENT_CONFINE` (`1`/`true`).
*
* All paths passed to the predicates here MUST be absolute and symlink-resolved
* (realpath) by the caller, mirroring `isSensitivePath`'s contract.
*
* @module config/attachment-guard
*/
import { sep } from 'node:path';
import { isSensitivePath } from '../web/sensitive-path.js';
import { readJsonConfig, SETTINGS_PATH } from '../web/route-helpers.js';
/**
* Directory trees blocked by default, IN ADDITION to the secret-location
* blocklist in `isSensitivePath`. Anything resolving under one of these trees
* is rejected. Pre-populated with the root account home and the system config
* tree (which already partially overlaps `isSensitivePath`'s `/etc/shadow`
* etc., but here we block the WHOLE tree).
*/
export const DEFAULT_BLOCKED_TREES: readonly string[] = ['/root', '/etc'];
/** Settings key carrying extra blocked directory trees (extends the defaults). */
export const ATTACHMENT_BLOCKED_PATHS_SETTING = 'attachmentBlockedPaths';
/** Settings key carrying the workspace-confinement toggle. */
export const ATTACHMENT_CONFINE_SETTING = 'attachmentConfineToWorkspace';
/** Resolved attachment-guard configuration. */
export interface AttachmentGuardConfig {
/** Pre-populated default trees PLUS any operator extras. */
blockedTrees: string[];
/** Whether attachments must resolve inside the session workspace. */
confineToWorkspace: boolean;
}
/** Normalizes a tree prefix: trim, drop trailing separators (but keep root). */
function normalizeTree(raw: string): string {
const trimmed = raw.trim();
if (!trimmed) return '';
// Strip trailing slashes so '/etc/' and '/etc' behave the same; never reduce
// a bare separator to empty.
const stripped = trimmed.replace(/[/\\]+$/, '');
return stripped || trimmed[0];
}
/**
* Returns true if `absPath` (absolute, symlink-resolved) is the tree itself or
* lives under it. Uses path-separator-aware matching so `/etc` does NOT block
* an unrelated `/etcetera/notes.md`.
*/
export function isUnderTree(absPath: string, tree: string): boolean {
const t = normalizeTree(tree);
if (!t) return false;
if (absPath === t) return true;
return absPath.startsWith(t.endsWith(sep) ? t : t + sep);
}
/** Parses the comma-separated env override into a list of normalized trees. */
function parseEnvBlockedTrees(): string[] {
const raw = process.env.CODEMAN_ATTACHMENT_BLOCKED_PATHS;
if (!raw) return [];
return raw
.split(',')
.map(normalizeTree)
.filter((t) => t.length > 0);
}
/** Parses the env confinement toggle (`1`/`true`/`yes`/`on`, case-insensitive). */
function parseEnvConfine(): boolean | undefined {
const raw = process.env.CODEMAN_ATTACHMENT_CONFINE;
if (raw === undefined) return undefined;
return /^(1|true|yes|on)$/i.test(raw.trim());
}
/**
* Loads the effective attachment-guard config by merging the pre-populated
* defaults with settings.json and env overrides. Env wins over settings for the
* confinement toggle; blocked-tree extras from BOTH sources are unioned on top
* of the defaults (operators can only EXTEND, never shrink, the blocked set).
*/
export async function loadAttachmentGuardConfig(): Promise<AttachmentGuardConfig> {
const settings = await readJsonConfig<Record<string, unknown>>(SETTINGS_PATH, 'settings.json', {});
const settingsTrees = Array.isArray(settings[ATTACHMENT_BLOCKED_PATHS_SETTING])
? (settings[ATTACHMENT_BLOCKED_PATHS_SETTING] as unknown[])
.filter((v): v is string => typeof v === 'string')
.map(normalizeTree)
.filter((t) => t.length > 0)
: [];
const blockedTrees = Array.from(new Set([...DEFAULT_BLOCKED_TREES, ...settingsTrees, ...parseEnvBlockedTrees()]));
const envConfine = parseEnvConfine();
const settingsConfine = settings[ATTACHMENT_CONFINE_SETTING] === true;
const confineToWorkspace = envConfine ?? settingsConfine;
return { blockedTrees, confineToWorkspace };
}
/**
* Attachment-specific blocklist check. Builds on the shared `isSensitivePath`
* base (secret locations, shared with `/api/download`) and ADDS the configured
* directory trees (`/root`, `/etc`, plus operator extras). `absPath` must be
* absolute and symlink-resolved.
*
* NOTE: this is intentionally a SUPERSET of `isSensitivePath` so `/api/download`
* behavior is NOT changed — only attachment registration/serving uses this.
*/
export function isBlockedAttachmentPath(absPath: string, blockedTrees: readonly string[]): boolean {
if (isSensitivePath(absPath)) return true;
return blockedTrees.some((tree) => isUnderTree(absPath, tree));
}
+21 -5
View File
@@ -6,14 +6,15 @@
* it easy to tune memory usage.
*
* Memory Budget Rationale (for 20 concurrent sessions):
* - Terminal buffer: 2MB max × 20 = 40MB worst case
* - Terminal buffer: 32MB max × 20 = 640MB worst case
* - Text output: 1MB max × 20 = 20MB worst case
* - Messages: ~1KB each × 1000 × 20 = 20MB worst case
* - Total buffer overhead: ~80MB (acceptable for long-running server)
*
* @module config/buffer-limits
*/
import { DEFAULT_TERMINAL_BUFFER_MAX_BYTES, DEFAULT_TERMINAL_BUFFER_TRIM_BYTES } from './terminal-history.js';
// ============================================================================
// Terminal Buffer Limits
// ============================================================================
@@ -21,17 +22,17 @@
/**
* Maximum terminal buffer size in characters.
* Contains raw terminal output with ANSI escape sequences.
* Reduced from 5MB to 2MB for better render performance.
* Sourced from terminal-history config (env/settings overridable).
* Override: CODEMAN_MAX_TERMINAL_BUFFER (bytes)
*/
export const MAX_TERMINAL_BUFFER_SIZE = parseInt(process.env.CODEMAN_MAX_TERMINAL_BUFFER || '') || 2 * 1024 * 1024;
export const MAX_TERMINAL_BUFFER_SIZE = DEFAULT_TERMINAL_BUFFER_MAX_BYTES;
/**
* Size to trim terminal buffer to when max is exceeded.
* Keeps the most recent portion to preserve context.
* Override: CODEMAN_TRIM_TERMINAL_TO (bytes)
*/
export const TRIM_TERMINAL_TO = parseInt(process.env.CODEMAN_TRIM_TERMINAL_TO || '') || 1.5 * 1024 * 1024;
export const TRIM_TERMINAL_TO = DEFAULT_TERMINAL_BUFFER_TRIM_BYTES;
// ============================================================================
// Text Output Buffer Limits
@@ -96,3 +97,18 @@ export const TRIM_RESPAWN_BUFFER_TO = 512 * 1024; // 512KB
* which is enough to extract metadata from the first few JSONL lines.
*/
export const FILE_PEEK_BYTES = 8 * 1024 - 1; // 8KB (inclusive end offset)
// ============================================================================
// Paste-Image Upload Limits
// ============================================================================
/**
* Maximum size (bytes) of a single image uploaded via POST
* /api/sessions/:id/paste-image. The mobile picker / drag-drop / paste paths
* send one file per request (the client uploads up to MAX_PASTE_IMAGES of them
* per batch), so this caps each individual file, not the batch. Generous enough
* for full-resolution phone photos and large screenshots; the client downscales
* very large images before upload, so legitimate uploads land well under this.
* Override: CODEMAN_MAX_PASTE_IMAGE_BYTES (bytes)
*/
export const MAX_PASTE_IMAGE_BYTES = parseInt(process.env.CODEMAN_MAX_PASTE_IMAGE_BYTES || '') || 50 * 1024 * 1024; // 50MB
+148
View File
@@ -0,0 +1,148 @@
/**
* @fileoverview Static registry of downstream tool dependencies probed by
* `codeman doctor`. Each entry declares per-environment resolvers and the
* skills that use it. EXTENSION POINT: skill-manifest-driven discovery
* (COD follow-up) will merge dynamically-found entries into this list.
*
* @module config/dependency-registry
*/
export type ProbeEnvironment = 'linux' | 'darwin' | 'win32' | 'wsl';
/** The valid `--category` filter values; single source of truth for the type, the CLI
* help text, and CLI input validation. */
export const TOOL_CATEGORIES = ['core', 'office', 'other'] as const;
export type ToolCategory = (typeof TOOL_CATEGORIES)[number];
/** Resolve a binary on the PATH and read its version. */
export interface PathResolver {
kind: 'path';
bins: string[];
versionArg?: string; // default '--version'
versionRegex?: RegExp; // default matches first \d+.\d+(.\d+)?
}
/** Resolve a Windows-installed app reachable from win32 or WSL. */
export interface WindowsSideResolver {
kind: 'windows-side';
appDirs: string[]; // relative to a Program Files root
exes: string[]; // candidate executables; first found wins
}
export interface ResolverSpec {
match: ProbeEnvironment[];
resolver: PathResolver | WindowsSideResolver;
}
export interface ToolDependency {
id: string;
label: string;
category: ToolCategory;
required: boolean;
usedBy?: string[];
minVersion?: string;
resolvers: ResolverSpec[];
installHint?: Partial<Record<ProbeEnvironment, string>>;
}
const ALL: ProbeEnvironment[] = ['linux', 'darwin', 'wsl', 'win32'];
export const DEPENDENCY_REGISTRY: ToolDependency[] = [
{
id: 'node',
label: 'Node.js',
category: 'core',
required: true,
minVersion: '22.0.0',
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['node'], versionArg: '--version' } }],
installHint: { linux: 'https://nodejs.org', darwin: 'brew install node', wsl: 'https://nodejs.org' },
},
{
id: 'claude',
label: 'Claude CLI',
category: 'core',
required: false,
usedBy: ['Claude Code sessions (default backend)'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['claude'], versionArg: '--version' } }],
installHint: { linux: 'https://docs.claude.com/claude-code', darwin: 'https://docs.claude.com/claude-code' },
},
{
id: 'tmux',
label: 'tmux',
category: 'core',
required: true,
resolvers: [{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['tmux'], versionArg: '-V' } }],
installHint: { linux: 'sudo apt install tmux', darwin: 'brew install tmux', wsl: 'sudo apt install tmux' },
},
{
id: 'opencode',
label: 'OpenCode CLI',
category: 'core',
required: false,
usedBy: ['OpenCode sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['opencode'], versionArg: '--version' } }],
},
{
id: 'codex',
label: 'Codex CLI',
category: 'core',
required: false,
usedBy: ['Codex sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['codex'], versionArg: '--version' } }],
},
{
id: 'gemini',
label: 'Gemini CLI',
category: 'core',
required: false,
usedBy: ['Gemini sessions'],
resolvers: [{ match: ALL, resolver: { kind: 'path', bins: ['gemini'], versionArg: '--version' } }],
},
{
id: 'libreoffice',
label: 'LibreOffice',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['linux', 'darwin', 'wsl'],
resolver: { kind: 'path', bins: ['libreoffice', 'soffice'], versionArg: '--version' },
},
],
installHint: { linux: 'sudo apt install libreoffice', darwin: 'brew install --cask libreoffice' },
},
{
id: 'pdftoppm',
label: 'pdftoppm',
category: 'office',
required: false,
usedBy: ['document preview', 'PDF/Office first-page thumbnails'],
// poppler's pdftoppm prints its version to stderr; presence is what matters here.
resolvers: [
{ match: ['linux', 'darwin', 'wsl'], resolver: { kind: 'path', bins: ['pdftoppm'], versionArg: '-v' } },
],
installHint: {
linux: 'sudo apt install poppler-utils',
darwin: 'brew install poppler',
wsl: 'sudo apt install poppler-utils',
},
},
{
id: 'msoffice',
label: 'MS Office',
category: 'office',
required: false,
usedBy: ['document preview', 'thumbnails'],
resolvers: [
{
match: ['wsl', 'win32'],
resolver: {
kind: 'windows-side',
appDirs: ['Microsoft Office/root/Office16'],
exes: ['WINWORD.EXE', 'POWERPNT.EXE', 'EXCEL.EXE'],
},
},
],
},
];
+67
View File
@@ -0,0 +1,67 @@
/**
* @fileoverview Per-instance shared hook secret (COD-54).
*
* Claude Code hooks POST to `/api/hook-event` with no Basic-Auth credentials,
* relying on a localhost bypass in `web/middleware/auth.ts`. That bypass is safe
* for loopback-only deploys, but a `cloudflared --url http://127.0.0.1:port`
* tunnel proxies internet traffic INTO the loopback origin, so tunneled requests
* arrive with `req.ip === 127.0.0.1` and would otherwise pass the bypass and
* drive respawn/Ralph signals unauthenticated.
*
* To close that hole WITHOUT breaking the loop's own (credential-less) hook
* channel, every locally-generated hook command now presents a per-instance
* shared secret in the `X-Codeman-Hook-Secret` header. The middleware requires
* a matching secret for the bypass WHEN A TUNNEL IS RUNNING. Tunneled internet
* traffic can't know the secret; local hooks (which we generate) do.
*
* Storage mirrors the VAPID-key pattern in `push-store.ts`: a small file under
* the instance data dir (`dataPath('hook-secret')`), read-if-present /
* generate-if-missing, stable across restarts. 256 bits of hex.
*/
import { existsSync, readFileSync, writeFileSync, mkdirSync } from 'node:fs';
import { randomBytes } from 'node:crypto';
import { getDataDir, dataPath } from './instance.js';
/** HTTP header local hooks use to present the shared secret. */
export const HOOK_SECRET_HEADER = 'X-Codeman-Hook-Secret';
/** Number of random bytes in the secret (256 bits → 64 hex chars). */
const SECRET_BYTES = 32;
let cachedSecret: string | null = null;
/**
* Return this instance's hook secret, generating and persisting it on first use.
* Stable across restarts. Cached in-process after the first read.
*/
export function getHookSecret(): string {
if (cachedSecret) return cachedSecret;
const secretFile = dataPath('hook-secret');
if (existsSync(secretFile)) {
try {
const raw = readFileSync(secretFile, 'utf-8').trim();
if (raw) {
cachedSecret = raw;
return cachedSecret;
}
// Empty/whitespace file — fall through and regenerate.
} catch {
// Unreadable — fall through and regenerate.
}
}
const secret = randomBytes(SECRET_BYTES).toString('hex');
try {
mkdirSync(getDataDir(), { recursive: true });
// Owner-only perms — the secret gates the hook bypass.
writeFileSync(secretFile, secret, { mode: 0o600 });
} catch {
// Best-effort persistence: even if the write fails we still return a usable
// secret for this process so hooks/middleware agree within this run.
}
cachedSecret = secret;
return cachedSecret;
}
+13
View File
@@ -43,6 +43,19 @@ export const MAX_SSE_CLIENTS = 100;
*/
export const MAX_TODOS_PER_SESSION = 500;
/**
* Maximum cron-job run-history records retained across all jobs. Oldest runs
* (by startedAt) are pruned when exceeded — bounds state.json growth from
* frequently-firing or perpetually-skipped jobs.
*/
export const MAX_CRON_RUN_HISTORY = 500;
/**
* Maximum saved cron jobs. Jobs persist to state.json, so an unbounded count
* would grow it without limit; creation past the cap is rejected with 400.
*/
export const MAX_CRON_JOBS = 100;
// ============================================================================
// Pending Tool Calls Limits
// ============================================================================
+13
View File
@@ -51,6 +51,19 @@ export const SCHEDULED_CLEANUP_INTERVAL = 5 * 60 * 1000;
/** Completed scheduled run max age before cleanup (ms) */
export const SCHEDULED_RUN_MAX_AGE = 60 * 60 * 1000;
// ============================================================================
// Cron Jobs
// ============================================================================
/** How often the cron loop wakes to check for due jobs (ms). */
export const CRON_TICK_INTERVAL = 30 * 1000;
/** Max attempts (× 500ms) to poll a launched session for CLI readiness before sending the prompt. */
export const CRON_READY_MAX_ATTEMPTS = 60;
/** Extra settle delay after CLI readiness is detected, before sending the prompt (ms). */
export const CRON_READY_SETTLE_MS = 2000;
/** Session limit retry wait before retrying (ms) */
export const SESSION_LIMIT_WAIT_MS = 5000;
+77
View File
@@ -0,0 +1,77 @@
/**
* Defaults, bounds, and resolution for terminal history retention.
*
* Raised defaults (the ones actually wired):
* - tmux history-limit: 50,000 -> 100,000 lines (applied at session spawn)
* - server PTY buffer cap: 2MB max / 1.5MB trim -> 32MB / 24MB (via buffer-limits.ts)
* Browser xterm scrollback is a separate hardcoded DEFAULT_SCROLLBACK (50,000) in
* src/web/public/constants.js and deliberately stays at 50k — 100k xterm lines per tab
* is a mobile-memory hazard — so DEFAULT_TERMINAL_SCROLLBACK_LINES stays 50,000 to match.
* The terminalScrollbackLines/terminalBufferMaxBytes/terminalBufferTrimBytes settings keys
* remain schema-validated but inert (a follow-up wires them); only tmuxHistoryLimit is live.
* All values remain env- and settings-overridable and bounds-clamped via
* resolveTerminalHistoryConfig().
*/
export const DEFAULT_TERMINAL_SCROLLBACK_LINES = 50_000;
export const DEFAULT_TMUX_HISTORY_LIMIT = 100_000;
export const DEFAULT_TERMINAL_BUFFER_MAX_BYTES =
parseInt(process.env.CODEMAN_MAX_TERMINAL_BUFFER || '', 10) || 32 * 1024 * 1024;
// Trim must stay below the max: BufferAccumulator.trim() keeps the last trimSize chars, so a
// trim >= max never shrinks the buffer — every append then re-joins the whole string (O(n²))
// and memory overshoots the operator's cap (e.g. CODEMAN_MAX_TERMINAL_BUFFER=2097152 with no
// trim env would leave the 24MB trim default in force). Clamp to 75% of the resolved max,
// preserving the 24MB/32MB default ratio as trim hysteresis.
export const DEFAULT_TERMINAL_BUFFER_TRIM_BYTES = Math.min(
parseInt(process.env.CODEMAN_TRIM_TERMINAL_TO || '', 10) || 24 * 1024 * 1024,
Math.floor(DEFAULT_TERMINAL_BUFFER_MAX_BYTES * 0.75)
);
export const MIN_TERMINAL_SCROLLBACK_LINES = 1_000;
export const MAX_TERMINAL_SCROLLBACK_LINES = 1_000_000;
export const MIN_TERMINAL_BUFFER_BYTES = 1024 * 1024;
export const MAX_TERMINAL_BUFFER_BYTES = 128 * 1024 * 1024;
export interface TerminalHistoryConfig {
terminalScrollbackLines: number;
tmuxHistoryLimit: number;
terminalBufferMaxBytes: number;
terminalBufferTrimBytes: number;
}
function boundedInt(value: unknown, fallback: number, min: number, max: number): number {
if (typeof value !== 'number' || !Number.isFinite(value)) return fallback;
return Math.max(min, Math.min(max, Math.trunc(value)));
}
export function resolveTerminalHistoryConfig(settings: Record<string, unknown> = {}): TerminalHistoryConfig {
const terminalBufferMaxBytes = boundedInt(
settings.terminalBufferMaxBytes,
DEFAULT_TERMINAL_BUFFER_MAX_BYTES,
MIN_TERMINAL_BUFFER_BYTES,
MAX_TERMINAL_BUFFER_BYTES
);
const terminalBufferTrimBytes = boundedInt(
settings.terminalBufferTrimBytes,
Math.min(DEFAULT_TERMINAL_BUFFER_TRIM_BYTES, terminalBufferMaxBytes),
MIN_TERMINAL_BUFFER_BYTES,
terminalBufferMaxBytes
);
return {
terminalScrollbackLines: boundedInt(
settings.terminalScrollbackLines,
DEFAULT_TERMINAL_SCROLLBACK_LINES,
MIN_TERMINAL_SCROLLBACK_LINES,
MAX_TERMINAL_SCROLLBACK_LINES
),
tmuxHistoryLimit: boundedInt(
settings.tmuxHistoryLimit,
DEFAULT_TMUX_HISTORY_LIMIT,
MIN_TERMINAL_SCROLLBACK_LINES,
MAX_TERMINAL_SCROLLBACK_LINES
),
terminalBufferMaxBytes,
terminalBufferTrimBytes,
};
}
+26
View File
@@ -0,0 +1,26 @@
/**
* @fileoverview Workflow (ultracode) run-watcher polling and cache configuration.
*
* Controls how frequently WorkflowRunWatcher polls
* ~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_*.json
* and how many runs are cached in memory.
*
* Distinct from the Agent-Teams config (team-config.ts). The run-state JSON is
* rewritten on every agent tick across a whole run (28+ agents), so the watcher
* relies on a per-file mtime skip; the poll itself is just N stat() calls.
*
* @module config/workflow-config
*/
/** Workflow run-state poll interval (ms). Short because a poll is just N mtime stats. */
export const WORKFLOW_RUN_POLL_INTERVAL_MS = 10_000;
/** Max cached workflow runs (LRU eviction). */
export const MAX_CACHED_WORKFLOW_RUNS = 100;
/**
* Default recency window (minutes) for getRecentRuns(). Generous enough that a
* recently-finished long run still appears in the LEFT-pane list — filtered on
* last-activity, not start time, so multi-hour runs don't vanish.
*/
export const WORKFLOW_RUN_RECENT_WINDOW_MIN = 240;
+32
View File
@@ -0,0 +1,32 @@
/**
* @fileoverview Input shape for creating/updating a cron job. This is the
* user-settable subset of `CronJob` (server-maintained bookkeeping fields
* such as nextRunAt / lastStatus are excluded). Produced by the zod schema.
*/
import type { ConcurrencyPolicy, InputMode, PromptMode, ScheduleType } from '../types/cron.js';
import type { SessionMode } from '../types/session.js';
export type { CronJob, CronJobRun, CronJobRunStatus, TriggerType } from '../types/cron.js';
export interface CronJobInput {
name: string;
agentType: SessionMode;
workingDir: string;
launchCommand?: string;
promptMode: PromptMode;
promptText?: string;
promptFilePath?: string;
inputMode: InputMode;
scheduleType: ScheduleType;
runAt?: number;
intervalMinutes?: number;
dailyTime?: string;
weeklyDays?: number[];
weeklyTime?: string;
enabled: boolean;
notes?: string;
concurrencyPolicy: ConcurrencyPolicy;
/** Default true. Ignored for 'once' schedules. */
autoClosePreviousSession?: boolean;
}
+637
View File
@@ -0,0 +1,637 @@
/**
* @fileoverview Cron service: CRUD for cron jobs, manual Run Now,
* the background due-job tick, and run-history recording.
*
* It does NOT own session/tmux logic — it reuses Codeman's existing session
* layer (create → addSession → setupSessionListeners → startInteractive/Shell →
* send prompt via writeViaMux/write), mirroring the "quick start" route flow.
*/
import { v4 as uuidv4 } from 'uuid';
import { readFile } from 'node:fs/promises';
import { statSync, realpathSync } from 'node:fs';
import { Session } from '../session.js';
import { SseEvent } from '../web/sse-events.js';
import { CronJobSchema } from '../web/schemas.js';
import { getErrorMessage, createErrorResponse, ApiErrorCode } from '../types/api.js';
import { MAX_CONCURRENT_SESSIONS, MAX_CRON_JOBS, MAX_CRON_RUN_HISTORY } from '../config/map-limits.js';
import { CRON_READY_MAX_ATTEMPTS, CRON_READY_SETTLE_MS } from '../config/server-timing.js';
import {
DEFAULT_BLOCKED_TREES,
isBlockedAttachmentPath,
loadAttachmentGuardConfig,
} from '../config/attachment-guard.js';
import { validateSessionFilePath } from '../web/route-helpers.js';
import { computeNextRunAt, dueKeyFor } from './cron-time.js';
import type { SessionPort, EventPort, ConfigPort, InfraPort } from '../web/ports/index.js';
import type { CronJob, CronJobRun, CronJobRunStatus, TriggerType } from '../types/cron.js';
import type { CronJobInput } from './cron-input.js';
/** The subset of the route context the cron depends on. */
export type CronDeps = SessionPort & EventPort & ConfigPort & InfraPort;
const delay = (ms: number): Promise<void> => new Promise((r) => setTimeout(r, ms));
/** Hard ceiling on a prompt-file read (defends against unbounded-read DoS). */
const MAX_PROMPT_FILE_BYTES = 1024 * 1024;
/**
* Pseudo-filesystem trees a cron job may never touch, ON TOP of the shared
* attachment blocklist. `/proc` in particular defeats the workingDir
* confinement trick (`workingDir: '/proc'` + `promptFilePath:
* '/proc/self/environ'` would read the SERVER's own environment).
*/
const CRON_PSEUDO_FS_TREES: readonly string[] = ['/proc', '/sys', '/dev'];
/** Sync blocklist for the create/update workingDir gate (no settings extras). */
const CRON_WORKING_DIR_BLOCKED_TREES: readonly string[] = [...DEFAULT_BLOCKED_TREES, ...CRON_PSEUDO_FS_TREES];
/** Prompt delivery is single-line only (writeViaMux/Ink constraint). */
const HAS_NEWLINE = /[\r\n]/;
/** Order-insensitive equality for the weekly-days arrays. */
function sameDays(a: number[] | undefined, b: number[] | undefined): boolean {
const x = [...(a ?? [])].sort((p, q) => p - q);
const y = [...(b ?? [])].sort((p, q) => p - q);
return x.length === y.length && x.every((v, i) => v === y[i]);
}
export class CronService {
constructor(private readonly deps: CronDeps) {}
private get store() {
return this.deps.store;
}
// ───────────────────────────── Reads ─────────────────────────────
listJobs(): CronJob[] {
return Object.values(this.store.getCronJobs());
}
getJob(id: string): CronJob | null {
return this.store.getCronJob(id);
}
listRuns(jobId?: string): CronJobRun[] {
const all = Object.values(this.store.getCronJobRuns());
const filtered = jobId ? all.filter((r) => r.cronJobId === jobId) : all;
return filtered.sort((a, b) => b.startedAt - a.startedAt);
}
/**
* Number of LIVE sessions of a given agent type (for the multi-session
* warning and the skip_if_same_agent_running policy). Sessions whose CLI has
* exited (`stopped`/`error` — the tab is still open but nothing is running)
* don't count. When `excludeJobId` is given, sessions created by that job's
* own runs are also excluded — otherwise a recurring job with the skip
* policy would deadlock on its own previous (never-closed) session and fire
* exactly once, forever skipping after that.
*/
countActiveAgents(agentType: string, excludeJobId?: string): number {
const ownSessionIds = excludeJobId
? new Set(
this.listRuns(excludeJobId)
.map((r) => r.sessionId)
.filter((id): id is string => id !== null)
)
: null;
let n = 0;
for (const [id, s] of this.deps.sessions.entries()) {
if (s.mode !== agentType) continue;
if (s.status === 'stopped' || s.status === 'error') continue;
if (ownSessionIds?.has(id)) continue;
n++;
}
return n;
}
// ──────────────────────────── Mutations ───────────────────────────
createJob(input: CronJobInput): CronJob {
if (Object.keys(this.store.getCronJobs()).length >= MAX_CRON_JOBS) {
throw this.badRequest(`Maximum number of cron jobs (${MAX_CRON_JOBS}) reached`);
}
this.assertValidWorkingDir(input.workingDir);
const now = Date.now();
const job: CronJob = {
id: uuidv4(),
name: input.name,
agentType: input.agentType,
workingDir: input.workingDir,
launchCommand: input.launchCommand,
promptMode: input.promptMode,
promptText: input.promptText,
promptFilePath: input.promptFilePath,
inputMode: input.inputMode,
scheduleType: input.scheduleType,
runAt: input.runAt,
intervalMinutes: input.intervalMinutes,
dailyTime: input.dailyTime,
weeklyDays: input.weeklyDays,
weeklyTime: input.weeklyTime,
enabled: input.enabled,
notes: input.notes,
concurrencyPolicy: input.concurrencyPolicy,
autoClosePreviousSession: input.autoClosePreviousSession ?? true,
createdAt: now,
updatedAt: now,
lastRunAt: null,
nextRunAt: null,
lastStatus: null,
lastDueKey: null,
};
job.nextRunAt = job.enabled ? computeNextRunAt(job, now) : null;
this.store.setCronJob(job.id, job);
this.broadcastListChanged();
return job;
}
updateJob(id: string, patch: Partial<CronJobInput>): CronJob | null {
const existing = this.getJob(id);
if (!existing) return null;
const now = Date.now();
// A completed one-time job is only re-armed when the SCHEDULE actually
// CHANGES — otherwise a cosmetic edit would silently resurrect a job that
// already fired. We compare VALUES, not field-presence: the edit form
// round-trips the full job (incl. unchanged scheduleType/runAt) on every
// save, so a presence check would always re-arm. Only a real schedule
// change re-arms.
const changed = <T>(next: T | undefined, prev: T): boolean => next !== undefined && next !== prev;
const scheduleChanged =
changed(patch.scheduleType, existing.scheduleType) ||
changed(patch.runAt, existing.runAt) ||
changed(patch.intervalMinutes, existing.intervalMinutes) ||
changed(patch.dailyTime, existing.dailyTime) ||
changed(patch.weeklyTime, existing.weeklyTime) ||
(patch.weeklyDays !== undefined && !sameDays(patch.weeklyDays, existing.weeklyDays));
const reArm = existing.scheduleType !== 'once' || !existing.completedOnce || scheduleChanged;
const updated: CronJob = {
...existing,
...patch,
id: existing.id,
createdAt: existing.createdAt,
updatedAt: now,
completedOnce: reArm ? false : existing.completedOnce,
lastDueKey: null,
};
// The PUT schema is `.partial()`, so its cross-field rules don't run on a
// partial body. Re-validate the MERGED job against the full schema so a
// partial edit can't leave an enabled job with an inconsistent schedule
// (e.g. switching to `once` without a `runAt` → a dead `nextRunAt:null`).
const check = CronJobSchema.safeParse(updated);
if (!check.success) {
throw this.badRequest(check.error.issues[0]?.message ?? 'Invalid cron job update');
}
if (patch.workingDir !== undefined) this.assertValidWorkingDir(patch.workingDir);
updated.nextRunAt = updated.enabled ? computeNextRunAt(updated, now) : null;
this.store.setCronJob(updated.id, updated);
this.broadcastListChanged();
return updated;
}
setEnabled(id: string, enabled: boolean): CronJob | null {
const existing = this.getJob(id);
if (!existing) return null;
const now = Date.now();
existing.enabled = enabled;
existing.updatedAt = now;
existing.nextRunAt = enabled ? computeNextRunAt(existing, now) : null;
this.store.setCronJob(existing.id, existing);
this.broadcastListChanged();
return existing;
}
deleteJob(id: string): boolean {
if (!this.getJob(id)) return false;
this.store.removeCronJob(id);
for (const run of this.listRuns(id)) this.store.removeCronJobRun(run.id);
this.deps.broadcast(SseEvent.CronJobDeleted, { id });
this.broadcastListChanged();
return true;
}
// ──────────────────────────── Execution ───────────────────────────
/** Manual Run Now — always launches regardless of schedule/enabled state. */
async runNow(id: string): Promise<CronJobRun | null> {
const job = this.getJob(id);
if (!job) return null;
return this.launch(job, 'manual_run_now');
}
/**
* Background tick: launch every enabled job whose next run is due. Advances
* each job's schedule and guards against double-launching the same due time.
*/
async tickDueJobs(now: number = Date.now()): Promise<void> {
for (const job of this.listJobs()) {
if (!job.enabled || job.nextRunAt == null || job.nextRunAt > now) continue;
const key = dueKeyFor(job.id, job.nextRunAt);
if (job.lastDueKey === key) {
// This due time was already consumed (overlap/restart) — just advance.
this.advanceAfterFire(job, now);
continue;
}
// Optional concurrency policy for AUTOMATIC runs. Only LIVE sessions
// block, and this job's own previous sessions never do (see
// countActiveAgents) — otherwise a recurring job would deadlock on the
// session it created last time.
if (job.concurrencyPolicy === 'skip_if_same_agent_running' && this.countActiveAgents(job.agentType, job.id) > 0) {
// Record the skip so the job's run history isn't silently empty when it
// keeps getting skipped (otherwise it looks like the job never ran).
this.recordSkippedRun(job);
if (job.scheduleType === 'once') {
// A skipped one-time job is NOT consumed: leave nextRunAt armed (and
// the due key unconsumed) so the next tick retries once the blocking
// session goes away.
continue;
}
job.lastDueKey = key;
this.advanceAfterFire(job, now);
continue;
}
job.lastDueKey = key;
// Advance the schedule BEFORE launching so a slow launch can't be
// re-triggered by the next tick.
this.advanceAfterFire(job, now);
this.launch(job, 'scheduled').catch((err) =>
console.error(`[cron] launch failed for job ${job.id}:`, getErrorMessage(err))
);
}
}
/** Recompute nextRunAt for loaded jobs on boot (e.g. after a restart). */
init(): void {
const now = Date.now();
for (const job of this.listJobs()) {
const isDeadOnce = job.scheduleType === 'once' && job.completedOnce;
if (job.enabled && job.nextRunAt == null && !isDeadOnce) {
job.nextRunAt = computeNextRunAt(job, now);
this.store.setCronJob(job.id, job);
}
}
}
// ──────────────────────────── Internals ───────────────────────────
private advanceAfterFire(job: CronJob, now: number): void {
if (job.scheduleType === 'once') {
job.completedOnce = true;
job.enabled = false;
job.nextRunAt = null;
} else {
job.nextRunAt = computeNextRunAt(job, now);
}
job.updatedAt = now;
this.store.setCronJob(job.id, job);
this.broadcastListChanged();
}
private async launch(job: CronJob, trigger: TriggerType): Promise<CronJobRun> {
const run: CronJobRun = {
id: uuidv4(),
cronJobId: job.id,
sessionId: null,
sessionName: null,
startedAt: Date.now(),
finishedAt: null,
status: 'created',
triggerType: trigger,
createdSessionUrl: null,
};
this.store.setCronJobRun(run.id, run);
this.pruneRunHistory();
this.deps.broadcast(SseEvent.CronRunCreated, run);
// Resolve the prompt.
let prompt: string;
try {
prompt = await this.resolvePrompt(job);
} catch (err) {
return this.failRun(job, run, `Prompt error: ${getErrorMessage(err)}`);
}
// Validate working directory.
try {
if (!statSync(job.workingDir).isDirectory()) {
return this.failRun(job, run, 'workingDir is not a directory');
}
} catch {
return this.failRun(job, run, 'workingDir does not exist');
}
// Recurring jobs: close the still-open session created by this job's
// previous run before launching the next (default ON, opt-out via
// autoClosePreviousSession:false) — otherwise an unattended interval/daily
// job accumulates a new tab per fire until the global session cap.
if (job.scheduleType !== 'once' && job.autoClosePreviousSession !== false) {
await this.closePreviousRunSessions(job, run.id);
}
// Respect the global session cap.
if (this.deps.sessions.size >= MAX_CONCURRENT_SESSIONS) {
return this.failRun(job, run, `Maximum concurrent sessions (${MAX_CONCURRENT_SESSIONS}) reached`);
}
// Create + start the session (mirrors the quick-start route flow).
let session: Session;
try {
const mode = job.agentType;
const globalNice = await this.deps.getGlobalNiceConfig();
const modelConfig = await this.deps.getModelConfig();
const claudeModeConfig = await this.deps.getClaudeModeConfig();
const model = mode !== 'shell' ? modelConfig?.defaultModel || undefined : undefined;
session = new Session({
workingDir: job.workingDir,
mode,
name: job.name,
mux: this.deps.mux,
useMux: true,
niceConfig: globalNice,
model,
claudeMode: claudeModeConfig.claudeMode,
allowedTools: claudeModeConfig.allowedTools,
});
this.deps.addSession(session);
this.store.incrementSessionsCreated();
this.deps.persistSessionState(session);
await this.deps.setupSessionListeners(session);
this.deps.broadcast(SseEvent.SessionCreated, this.deps.getSessionStateWithRespawn(session));
if (mode === 'shell') {
await session.startShell();
} else {
await session.startInteractive();
}
this.deps.broadcast(SseEvent.SessionInteractive, { id: session.id, mode });
} catch (err) {
return this.failRun(job, run, `Session launch failed: ${getErrorMessage(err)}`);
}
run.sessionId = session.id;
run.sessionName = session.name;
run.createdSessionUrl = `/?session=${session.id}`;
run.status = 'session_started';
this.store.setCronJobRun(run.id, run);
this.deps.broadcast(SseEvent.CronRunUpdated, run);
this.updateJobLastStatus(job.id, 'session_started');
// Send the prompt once the CLI is ready (async; does not block the caller).
this.sendPromptWhenReady(session.id, prompt, job, run);
return run;
}
/**
* Resolves the prompt text and enforces the single-line constraint: prompt
* delivery rides writeViaMux/PTY writes where a newline is Enter, so a
* multi-line prompt would be silently corrupted (typed mode fuses lines,
* paste mode submits the first line and dribbles the rest in as separate
* messages). Rather than mangle an unattended agent's instructions, fail the
* run with a clear error. A prompt FILE may end with trailing newline(s)
* (every editor writes one) — those are stripped before the check.
*/
private async resolvePrompt(job: CronJob): Promise<string> {
if (job.promptMode === 'prompt_file_path') {
if (!job.promptFilePath) throw new Error('prompt file path is empty');
const safePath = await this.resolveSafePromptPath(job.promptFilePath, job.workingDir);
const content = (await readFile(safePath, 'utf-8')).replace(/[\r\n]+$/, '');
if (HAS_NEWLINE.test(content)) {
throw new Error('prompt file must contain a single line — multi-line prompts are not supported');
}
return content;
}
const text = job.promptText ?? '';
if (HAS_NEWLINE.test(text)) {
// Schema-rejected since this check was added; guards legacy persisted jobs.
throw new Error('promptText must be a single line — multi-line prompts are not supported');
}
return text;
}
/**
* Guards a prompt-file path before it is read. The path is user-supplied via
* the API and its contents are injected into an agent session (an exfil sink
* over SSE/terminal), so an unconfined read would let a hostile job config
* pull arbitrary host files — including the SERVER PROCESS'S OWN secrets via
* `/proc/self/environ` — into the session.
*
* A denylist is the wrong posture for an exfil sink (it kept missing `/proc`,
* `/dev`, other users' `~/.ssh`, modern cloud creds…). So the PRIMARY gate is
* an allowlist: the prompt file must resolve INSIDE the job's working
* directory. A symlink escaping the workspace fails this because we check the
* realpath-resolved target. We additionally require a regular file (rejects
* directories, FIFOs, and `/dev/*` character devices that would hang or OOM
* the unbounded read) within a sane size cap, and keep the shared blocklist as
* cheap defense-in-depth. Returns the symlink-resolved path to read.
*/
private async resolveSafePromptPath(rawPath: string, workingDir: string): Promise<string> {
let resolved: string;
try {
resolved = realpathSync(rawPath);
} catch {
throw new Error('prompt file path could not be resolved');
}
// workingDir is USER-CONTROLLED, so it is not a trust boundary by itself:
// realpath-resolve it (a symlinked workspace must not defeat containment)
// and reject blocked/pseudo-fs trees — otherwise workingDir '/proc' would
// make '/proc/self/environ' pass the containment check below.
let realWorkingDir: string;
try {
realWorkingDir = realpathSync(workingDir);
} catch {
throw new Error('job working directory could not be resolved');
}
const guard = await loadAttachmentGuardConfig();
const blockedTrees = [...guard.blockedTrees, ...CRON_PSEUDO_FS_TREES];
if (realWorkingDir === '/' || isBlockedAttachmentPath(realWorkingDir, blockedTrees)) {
throw new Error('job working directory is blocked');
}
// Defense-in-depth blocklist (secret locations, /etc, /root, pseudo-fs).
if (isBlockedAttachmentPath(resolved, blockedTrees)) {
throw new Error('prompt file path is blocked');
}
// Primary gate: the prompt file must live inside the job's workspace.
if (!validateSessionFilePath(realWorkingDir, resolved)) {
throw new Error('prompt file path must be inside the job working directory');
}
// Reject non-regular files and oversized files (DoS via unbounded read).
let info;
try {
info = statSync(resolved);
} catch {
throw new Error('prompt file path could not be resolved');
}
if (!info.isFile()) throw new Error('prompt file path is not a regular file');
if (info.size > MAX_PROMPT_FILE_BYTES) throw new Error('prompt file is too large');
return resolved;
}
private sendPromptWhenReady(sessionId: string, prompt: string, job: CronJob, run: CronJobRun): void {
setImmediate(() => {
const poll = async (): Promise<void> => {
if (job.agentType !== 'shell') {
for (let attempt = 0; attempt < CRON_READY_MAX_ATTEMPTS; attempt++) {
await delay(500);
const s = this.deps.sessions.get(sessionId);
if (!s) return; // session was removed
const buf = s.getTerminalBuffer().slice(-2048);
if (buf.includes('❯') || buf.includes('tokens')) break;
}
await delay(CRON_READY_SETTLE_MS);
} else {
await delay(1000);
// Shell mode: deliver the optional custom launch command as the
// first input line (single-line, schema-enforced), then give it a
// moment to start before the prompt follows.
if (job.launchCommand) {
const shell = this.deps.sessions.get(sessionId);
if (!shell) return;
const sent = await shell.writeViaMux(`${job.launchCommand}\r`);
if (!sent) {
this.failRun(job, run, 'Failed to send launch command: mux write failed');
return;
}
await delay(1000);
}
}
const s = this.deps.sessions.get(sessionId);
if (!s) return;
try {
const payload = prompt.endsWith('\r') ? prompt : `${prompt}\r`;
let delivered = true;
if (job.inputMode === 'paste') {
s.write(payload);
} else {
delivered = await s.writeViaMux(payload);
}
if (!delivered) {
this.failRun(job, run, 'Failed to send prompt: mux write failed');
return;
}
run.status = 'prompt_sent';
run.finishedAt = Date.now();
this.store.setCronJobRun(run.id, run);
this.deps.broadcast(SseEvent.CronRunUpdated, run);
this.updateJobLastStatus(job.id, 'prompt_sent');
} catch (err) {
this.failRun(job, run, `Failed to send prompt: ${getErrorMessage(err)}`);
}
};
poll().catch((err) => console.error('[cron] sendPromptWhenReady error:', getErrorMessage(err)));
});
}
/** 400-shaped error for route handlers (mirrors parseBody's error contract). */
private badRequest(msg: string): Error {
return Object.assign(new Error(msg), {
statusCode: 400,
body: createErrorResponse(ApiErrorCode.INVALID_INPUT, msg),
});
}
/**
* Create/update gate for a job's workingDir: must exist, be a directory, and
* not resolve into a blocked or pseudo-filesystem tree (nor the fs root).
* The user-supplied workingDir doubles as the prompt-file confinement root,
* so an unrestricted value would defeat that boundary (e.g. '/proc').
*/
private assertValidWorkingDir(workingDir: string): void {
let real: string;
try {
real = realpathSync(workingDir);
} catch {
throw this.badRequest('workingDir does not exist');
}
if (!statSync(real).isDirectory()) throw this.badRequest('workingDir is not a directory');
if (real === '/' || isBlockedAttachmentPath(real, CRON_WORKING_DIR_BLOCKED_TREES)) {
throw this.badRequest('workingDir is not allowed (blocked or pseudo-filesystem tree)');
}
}
/** Close still-open sessions created by this job's previous runs (normal cleanup path). */
private async closePreviousRunSessions(job: CronJob, currentRunId: string): Promise<void> {
for (const prev of this.listRuns(job.id)) {
if (prev.id === currentRunId || !prev.sessionId) continue;
if (!this.deps.sessions.has(prev.sessionId)) continue;
try {
await this.deps.cleanupSession(prev.sessionId, true, 'cron: superseded by the next run of this job');
} catch (err) {
console.error(`[cron] failed to auto-close previous session ${prev.sessionId}:`, getErrorMessage(err));
}
}
}
private failRun(job: CronJob, run: CronJobRun, message: string): CronJobRun {
run.status = 'failed';
run.errorMessage = message;
run.finishedAt = Date.now();
this.store.setCronJobRun(run.id, run);
this.deps.broadcast(SseEvent.CronRunUpdated, run);
this.updateJobLastStatus(job.id, 'failed');
return run;
}
private recordSkippedRun(job: CronJob): void {
// Coalesce consecutive skips: if the job is already in a skip streak, don't
// record again — a perpetually-skipped interval job would otherwise write a
// run every tick forever and bloat state.json.
if (this.listRuns(job.id)[0]?.status === 'skipped') return;
const now = Date.now();
const run: CronJobRun = {
id: uuidv4(),
cronJobId: job.id,
sessionId: null,
sessionName: null,
startedAt: now,
finishedAt: now,
status: 'skipped',
errorMessage: `Skipped: a ${job.agentType} agent is already running (concurrency policy)`,
triggerType: 'scheduled',
createdSessionUrl: null,
};
this.store.setCronJobRun(run.id, run);
this.pruneRunHistory();
this.deps.broadcast(SseEvent.CronRunCreated, run);
// A skip is NOT a run: surface it as the lastStatus, but do NOT advance
// lastRunAt (no session was created).
this.updateJobLastStatus(job.id, 'skipped', { touchLastRun: false });
}
/** Prune the oldest run records (by startedAt) once the global cap is exceeded. */
private pruneRunHistory(): void {
const runs = Object.values(this.store.getCronJobRuns());
if (runs.length <= MAX_CRON_RUN_HISTORY) return;
runs.sort((a, b) => a.startedAt - b.startedAt);
for (const run of runs.slice(0, runs.length - MAX_CRON_RUN_HISTORY)) {
this.store.removeCronJobRun(run.id);
}
}
private updateJobLastStatus(jobId: string, status: CronJobRunStatus, opts: { touchLastRun?: boolean } = {}): void {
const fresh = this.store.getCronJob(jobId);
if (!fresh) return;
const now = Date.now();
fresh.lastStatus = status;
if (opts.touchLastRun !== false) fresh.lastRunAt = now;
fresh.updatedAt = now;
this.store.setCronJob(fresh.id, fresh);
this.broadcastListChanged();
}
private broadcastListChanged(): void {
this.deps.broadcast(SseEvent.CronJobsChanged, { jobs: this.listJobs() });
}
}
+82
View File
@@ -0,0 +1,82 @@
/**
* @fileoverview Pure next-run-time calculations for the cron.
*
* All functions are pure and take an explicit `after` timestamp (epoch ms) so
* they are deterministic and unit-testable. Times use the SERVER'S LOCAL
* timezone for v0.1 (per the build brief) — daily/weekly wall-clock times are
* interpreted via the host's local time.
*/
import type { CronJob } from '../types/cron.js';
/** Parse an 'HH:MM' (24-hour) string into hours/minutes, or null if invalid. */
export function parseHHMM(value: string | undefined): { hours: number; minutes: number } | null {
if (!value) return null;
const m = /^(\d{1,2}):(\d{2})$/.exec(value.trim());
if (!m) return null;
const hours = Number(m[1]);
const minutes = Number(m[2]);
if (hours < 0 || hours > 23 || minutes < 0 || minutes > 59) return null;
return { hours, minutes };
}
/**
* Returns the epoch-ms timestamp for `hours:minutes` (local time) on the day of
* `base`, shifted by `dayOffset` days.
*/
function atLocalTime(base: number, hours: number, minutes: number, dayOffset: number): number {
const d = new Date(base);
d.setHours(hours, minutes, 0, 0);
d.setDate(d.getDate() + dayOffset);
return d.getTime();
}
/**
* Compute the next fire time strictly relevant to `after`, or null if the job
* has no future run (e.g. a completed one-time job, or invalid config).
*
* For `once`, returns the absolute `runAt` (even if already in the past, so a
* missed one-time job still fires once) until it has `completedOnce`.
*/
export function computeNextRunAt(job: CronJob, after: number): number | null {
switch (job.scheduleType) {
case 'once': {
if (job.completedOnce) return null;
return typeof job.runAt === 'number' ? job.runAt : null;
}
case 'interval': {
const minutes = job.intervalMinutes;
if (!minutes || minutes <= 0) return null;
return after + minutes * 60_000;
}
case 'daily': {
const t = parseHHMM(job.dailyTime);
if (!t) return null;
let next = atLocalTime(after, t.hours, t.minutes, 0);
if (next <= after) next = atLocalTime(after, t.hours, t.minutes, 1);
return next;
}
case 'weekly': {
const t = parseHHMM(job.weeklyTime);
if (!t) return null;
const days = (job.weeklyDays ?? []).filter((d) => d >= 0 && d <= 6);
if (days.length === 0) return null;
for (let offset = 0; offset <= 7; offset++) {
const cand = atLocalTime(after, t.hours, t.minutes, offset);
if (cand > after && days.includes(new Date(cand).getDay())) return cand;
}
return null;
}
default:
return null;
}
}
/**
* Duplicate-launch guard key: identifies a specific due time for a job. The
* cron records the key it last consumed so an overlapping or restarted
* loop will not launch the same due time twice.
*/
export function dueKeyFor(jobId: string, fireTime: number): string {
return `${jobId}:${fireTime}`;
}
+416
View File
@@ -0,0 +1,416 @@
/**
* @fileoverview Docker case export / import: move a container (toolchain + any
* in-image changes) PLUS its workspace to another machine as one portable
* `.codeman-container.tgz`, and restore it.
*
* A full-image export = `docker commit` the running container to an image ->
* `docker save` that image -> tar the bind-mounted workspace -> a manifest, all
* bundled into one gzip tarball. A workspace-only export skips the image (fast,
* files-only). Import validates the manifest + per-member checksums, extracts the
* workspace with a path-traversal guard, `docker load`s the image and RE-TAGS it
* into a quarantined namespace (never overwriting a local tag), and hands the
* caller enough to recreate a hardened case on the destination.
*
* Safety (all from the design critic): pause the container spanning the workspace
* tar AND the commit so the two artifacts are mutually consistent; a free-space
* precheck (a full docker graph wedges EVERY session on the host); `docker rmi`
* the intermediate image in a finally; sealed containers refuse a full-image
* export (an in-container login would ride the committed layer); import rejects
* absolute / `..` tar members and checksum mismatches. Bounded by
* runWithConversionLimit so N exports cannot fork-bomb the host.
*
* @module docker-export
*/
import { createReadStream, createWriteStream, existsSync, mkdirSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join, basename } from 'node:path';
import { createHash } from 'node:crypto';
import { spawn } from 'node:child_process';
import { pipeline } from 'node:stream/promises';
import type { DockerEngine, SessionDocker } from './types.js';
import { runWithConversionLimit } from './document-conversion-limiter.js';
const IS_TEST_MODE = !!process.env.VITEST;
/** Refuse to export when the target filesystem has less than this free (a full graph wedges the daemon). */
export const DOCKER_EXPORT_MIN_FREE_BYTES = 2 * 1024 * 1024 * 1024; // 2 GiB
/** Manifest schema version (bump on any breaking field change). */
export const DOCKER_EXPORT_SCHEMA = 1;
export type DockerExportMode = 'full' | 'workspace';
export interface DockerExportManifest {
schemaVersion: number;
caseName: string;
mode: DockerExportMode;
engine: DockerEngine;
image: string;
containerWorkdir: string;
network: string;
createdAt: number;
codemanVersion: string;
mountCredentials: boolean;
/** True when the bundle provably carries no credentials (convenient-mode workspace, or a full image whose creds were bind-mounted and thus never committed). */
secretFree: boolean;
/** sha256 of each bundle member that is present. */
checksums: { image?: string; workspace?: string };
}
// ========== Pure helpers (unit-tested) ==========
/** Raw argv prefix for the engine (NO shell escaping — used with spawn). */
export function dockerArgv(docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>): string[] {
const argv: string[] = [docker.engine === 'podman' ? 'podman' : 'docker'];
if (docker.context) argv.push('--context', docker.context);
if (docker.daemonHost) argv.push('-H', docker.daemonHost);
return argv;
}
/** Portable bundle filename for a case export. */
export function exportBundleName(caseName: string, timestamp: number, mode: DockerExportMode): string {
const suffix = mode === 'workspace' ? 'workspace' : 'container';
return `${caseName}-${timestamp}.codeman-${suffix}.tgz`;
}
/** Quarantined image tag for an imported bundle (never overwrites a local tag). */
export function importedImageTag(caseName: string, timestamp: number): string {
return `codeman/imported-${caseName}:${timestamp}`;
}
/** Intermediate commit tag for a full-image export (unique per export, rmi'd in finally). */
export function exportImageTag(caseName: string, timestamp: number): string {
return `codeman/export-${caseName}:${timestamp}`;
}
/**
* Reject a tar member path that would escape the extraction root (absolute path
* or a `..` component). The import-side traversal guard.
*/
export function isSafeTarMember(member: string): boolean {
const trimmed = member.trim();
if (!trimmed || trimmed === './') return true;
if (trimmed.startsWith('/')) return false;
// Normalize separators and check each component.
return !trimmed.split('/').some((part) => part === '..');
}
/** Parse the image id/ref from `docker load` output ("Loaded image: x" / "Loaded image ID: sha256:..."). */
export function parseLoadedImageRef(loadOutput: string): string | null {
const idMatch = loadOutput.match(/Loaded image ID:\s*(sha256:[0-9a-f]+)/i);
if (idMatch) return idMatch[1];
const refMatch = loadOutput.match(/Loaded image:\s*(\S+)/i);
if (refMatch) return refMatch[1];
return null;
}
// ========== IO helpers ==========
function run(
cmd: string,
args: string[],
opts: { timeout?: number } = {}
): Promise<{ stdout: string; stderr: string }> {
return new Promise((resolve, reject) => {
const child = spawn(cmd, args, { stdio: ['ignore', 'pipe', 'pipe'] });
let stdout = '';
let stderr = '';
let timer: NodeJS.Timeout | undefined;
if (opts.timeout) {
timer = setTimeout(() => {
child.kill('SIGKILL');
reject(new Error(`${cmd} timed out after ${opts.timeout}ms`));
}, opts.timeout);
}
child.stdout.on('data', (d) => (stdout += d));
child.stderr.on('data', (d) => (stderr += d));
child.on('error', (err) => {
if (timer) clearTimeout(timer);
reject(err);
});
child.on('close', (code) => {
if (timer) clearTimeout(timer);
if (code === 0) resolve({ stdout, stderr });
else reject(new Error(`${cmd} ${args.join(' ')} exited ${code}: ${stderr.trim()}`));
});
});
}
/**
* Stream `docker save <tag>` stdout to a raw tar file (no shell, no double-gzip).
* Uses stream `pipeline` so completion means the write stream is FULLY flushed to
* disk (a naive child 'close' resolves before the last chunks land, truncating the
* file — a real bug caught in end-to-end testing), AND waits for a clean exit code.
*/
async function saveImageToTar(argv: string[], tag: string, outPath: string): Promise<void> {
const child = spawn(argv[0], [...argv.slice(1), 'save', tag], { stdio: ['ignore', 'pipe', 'pipe'] });
let stderr = '';
child.stderr.on('data', (d) => (stderr += d));
const exited = new Promise<void>((resolve, reject) => {
child.on('error', reject);
child.on('close', (code) =>
code === 0 ? resolve() : reject(new Error(`docker save exited ${code}: ${stderr.trim()}`))
);
});
// pipeline resolves only after the destination has fully flushed.
await Promise.all([pipeline(child.stdout, createWriteStream(outPath)), exited]);
}
async function sha256File(path: string): Promise<string> {
return new Promise((resolve, reject) => {
const hash = createHash('sha256');
const stream = createReadStream(path);
stream.on('data', (d) => hash.update(d));
stream.on('error', reject);
stream.on('end', () => resolve(hash.digest('hex')));
});
}
async function freeBytes(path: string): Promise<number> {
try {
const stat = await fs.statfs(path);
return Number(stat.bavail) * Number(stat.bsize);
} catch {
return Number.POSITIVE_INFINITY; // statfs unsupported — don't block
}
}
async function isContainerRunning(argv: string[], container: string): Promise<boolean> {
try {
const { stdout } = await run(argv[0], [...argv.slice(1), 'inspect', '-f', '{{.State.Running}}', container], {
timeout: 15_000,
});
return stdout.trim() === 'true';
} catch {
return false;
}
}
export interface ExportResult {
bundlePath: string;
manifest: DockerExportManifest;
sizeBytes: number;
}
/**
* Export a docker case to a portable bundle. Bounded by runWithConversionLimit.
* `full` mode commits + saves the image AND tars the workspace; `workspace` mode
* tars just the workspace. The container is paused across the artifact capture so
* image and workspace are mutually consistent.
*/
export async function exportDockerCase(params: {
docker: SessionDocker;
caseName: string;
timestamp: number;
exportsDir: string;
mode: DockerExportMode;
codemanVersion: string;
}): Promise<ExportResult> {
const { docker, caseName, timestamp, exportsDir, mode, codemanVersion } = params;
if (mode === 'full' && !docker.mountCredentials) {
throw new Error(
'full-image export is refused for a sealed (mountCredentials:false) container: an in-container login would ride the committed image layer. Use a workspace-only export.'
);
}
if (IS_TEST_MODE) {
// No real docker/tar under vitest — return a deterministic stub.
const manifest: DockerExportManifest = {
schemaVersion: DOCKER_EXPORT_SCHEMA,
caseName,
mode,
engine: docker.engine,
image: docker.image,
containerWorkdir: docker.containerWorkdir,
network: docker.network,
createdAt: timestamp,
codemanVersion,
mountCredentials: docker.mountCredentials,
secretFree: true,
checksums: {},
};
return { bundlePath: join(exportsDir, exportBundleName(caseName, timestamp, mode)), manifest, sizeBytes: 0 };
}
return runWithConversionLimit(async () => {
if (!existsSync(exportsDir)) mkdirSync(exportsDir, { recursive: true });
const free = await freeBytes(exportsDir);
if (free < DOCKER_EXPORT_MIN_FREE_BYTES) {
throw new Error(
`not enough free space to export (need >= ${Math.round(DOCKER_EXPORT_MIN_FREE_BYTES / 1e9)}GB, have ${Math.round(free / 1e9)}GB). A full docker graph wedges every session on the host.`
);
}
const argv = dockerArgv(docker);
const bundlePath = join(exportsDir, exportBundleName(caseName, timestamp, mode));
const stageDir = join(exportsDir, `.stage-${caseName}-${timestamp}`);
mkdirSync(stageDir, { recursive: true });
const wasRunning = await isContainerRunning(argv, docker.containerName);
let commitTag: string | undefined;
try {
if (wasRunning) {
await run(argv[0], [...argv.slice(1), 'pause', docker.containerName], { timeout: 30_000 }).catch(() => {});
}
const checksums: DockerExportManifest['checksums'] = {};
if (mode === 'full') {
commitTag = exportImageTag(caseName, timestamp);
// Blank instance-specific committed env so the image carries no stale host refs.
await run(
argv[0],
[
...argv.slice(1),
'commit',
'-c',
'ENV CODEMAN_API_URL=',
'-c',
'ENV CODEMAN_HOOK_SECRET_FILE=',
docker.containerName,
commitTag,
],
{ timeout: 300_000 }
);
const imageTar = join(stageDir, 'image.tar');
await saveImageToTar(argv, commitTag, imageTar);
checksums.image = await sha256File(imageTar);
}
const workspaceTar = join(stageDir, 'workspace.tar');
await run('tar', ['-cf', workspaceTar, '-C', docker.hostWorkspacePath, '.'], { timeout: 300_000 });
checksums.workspace = await sha256File(workspaceTar);
const manifest: DockerExportManifest = {
schemaVersion: DOCKER_EXPORT_SCHEMA,
caseName,
mode,
engine: docker.engine,
image: docker.image,
containerWorkdir: docker.containerWorkdir,
network: docker.network,
createdAt: timestamp,
codemanVersion,
mountCredentials: docker.mountCredentials,
// Convenient mode keeps creds on bind mounts (never committed), so the bundle is secret-free.
secretFree: docker.mountCredentials,
checksums,
};
await fs.writeFile(join(stageDir, 'manifest.json'), JSON.stringify(manifest, null, 2));
const members =
mode === 'full' ? ['manifest.json', 'image.tar', 'workspace.tar'] : ['manifest.json', 'workspace.tar'];
await run('tar', ['-czf', bundlePath, '-C', stageDir, ...members], { timeout: 300_000 });
const stat = await fs.stat(bundlePath);
return { bundlePath, manifest, sizeBytes: stat.size };
} finally {
// Always remove the intermediate image + stage dir, and unpause.
if (commitTag) {
await run(argv[0], [...argv.slice(1), 'rmi', commitTag], { timeout: 60_000 }).catch(() => {});
}
await fs.rm(stageDir, { recursive: true, force: true }).catch(() => {});
if (wasRunning) {
await run(argv[0], [...argv.slice(1), 'unpause', docker.containerName], { timeout: 30_000 }).catch(() => {});
}
}
});
}
export interface ImportResult {
manifest: DockerExportManifest;
/** Quarantined image ref the destination case should use (full mode only). */
importedImage?: string;
/** Directory the workspace was extracted into. */
workspacePath: string;
}
/**
* Import a bundle produced by exportDockerCase: validate the manifest + per-member
* checksums, extract the workspace (traversal-guarded) into destWorkspace, and, in
* full mode, `docker load` the image and re-tag it into a quarantined namespace.
*/
export async function importDockerBundle(params: {
bundlePath: string;
destWorkspace: string;
engine: DockerEngine;
timestamp: number;
}): Promise<ImportResult> {
const { bundlePath, destWorkspace, engine, timestamp } = params;
const argv: string[] = [engine === 'podman' ? 'podman' : 'docker'];
if (IS_TEST_MODE) {
const raw = await fs.readFile(bundlePath, 'utf-8').catch(() => '{}');
return { manifest: JSON.parse(raw) as DockerExportManifest, workspacePath: destWorkspace };
}
const stageDir = `${destWorkspace}.import-stage-${timestamp}`;
mkdirSync(stageDir, { recursive: true });
try {
await run('tar', ['-xzf', bundlePath, '-C', stageDir], { timeout: 300_000 });
const manifestRaw = await fs.readFile(join(stageDir, 'manifest.json'), 'utf-8');
const manifest = JSON.parse(manifestRaw) as DockerExportManifest;
if (manifest.schemaVersion !== DOCKER_EXPORT_SCHEMA) {
throw new Error(`unsupported export schema version ${manifest.schemaVersion} (expected ${DOCKER_EXPORT_SCHEMA})`);
}
// Integrity: verify checksums before trusting any member.
const workspaceTar = join(stageDir, 'workspace.tar');
if (manifest.checksums.workspace) {
const actual = await sha256File(workspaceTar);
if (actual !== manifest.checksums.workspace)
throw new Error('workspace checksum mismatch (corrupt or tampered bundle)');
}
// Traversal guard: reject absolute / `..` members before extraction.
const { stdout: memberList } = await run('tar', ['-tf', workspaceTar], { timeout: 60_000 });
for (const member of memberList.split('\n').filter(Boolean)) {
if (!isSafeTarMember(member)) throw new Error(`unsafe path in workspace archive: ${member}`);
}
mkdirSync(destWorkspace, { recursive: true });
await run('tar', ['--no-same-owner', '-xf', workspaceTar, '-C', destWorkspace], { timeout: 300_000 });
let importedImage: string | undefined;
if (manifest.mode === 'full') {
const imageTar = join(stageDir, 'image.tar');
if (manifest.checksums.image) {
const actual = await sha256File(imageTar);
if (actual !== manifest.checksums.image)
throw new Error('image checksum mismatch (corrupt or tampered bundle)');
}
const { stdout } = await run(argv[0], [...argv.slice(1), 'load', '-i', imageTar], { timeout: 300_000 });
const loadedRef = parseLoadedImageRef(stdout);
if (!loadedRef) throw new Error('could not determine loaded image ref');
// Quarantine: re-tag by the loaded ref/id, never trusting the bundle's original tag.
importedImage = importedImageTag(manifest.caseName, timestamp);
await run(argv[0], [...argv.slice(1), 'tag', loadedRef, importedImage], { timeout: 60_000 });
}
return { manifest, importedImage, workspacePath: destWorkspace };
} finally {
await fs.rm(stageDir, { recursive: true, force: true }).catch(() => {});
}
}
/** List export bundles in the exports dir (newest first), with size + mtime. */
export async function listDockerExports(
exportsDir: string
): Promise<Array<{ name: string; sizeBytes: number; mtimeMs: number }>> {
if (!existsSync(exportsDir)) return [];
const entries = await fs.readdir(exportsDir).catch(() => [] as string[]);
const out: Array<{ name: string; sizeBytes: number; mtimeMs: number }> = [];
for (const name of entries) {
if (!name.endsWith('.tgz')) continue;
try {
const stat = await fs.stat(join(exportsDir, name));
out.push({ name: basename(name), sizeBytes: stat.size, mtimeMs: stat.mtimeMs });
} catch {
/* skip */
}
}
return out.sort((a, b) => b.mtimeMs - a.mtimeMs);
}
+956
View File
@@ -0,0 +1,956 @@
/**
* @fileoverview Docker cases: storage, pure command-arg builders, and daemon probes.
*
* Docker mode is a LOCATION OVERLAY on cases (not a 6th SessionMode), the direct
* analog of the remote-SSH feature in `remote-hosts.ts`. Instead of a local tmux
* pane running `ssh host` into a durable remote tmux server, a local tmux pane
* runs `docker exec -it` into a durable IN-CONTAINER tmux server. The container is
* scoped to the CASE (`codeman-case-<name>`), so multiple sessions can `docker
* exec` into the same long-lived container.
*
* This module mirrors `remote-hosts.ts`:
* - JSON storage for hosts (`docker-hosts.json`) and cases (`docker-cases.json`)
* - `toSessionDocker()` (mirror of `toSessionRemote`)
* - `buildDockerBaseArgs()` / `buildDockerCreateArgs()` (mirror of `buildSshConnectionArgs`)
* - `checkDockerAvailable()` / `checkDockerTmuxAvailable()` (mirror of `checkRemoteTmuxAvailable`)
*
* The launch/kill command orchestration (`buildDockerLaunchCommand`,
* `buildDockerKillCommand`, `dockerTmuxSessionName`) lives in `tmux-manager.ts`,
* mirroring where `buildRemoteLaunchCommand` lives.
*
* @module docker-hosts
*/
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join, dirname } from 'node:path';
import { fileURLToPath } from 'node:url';
import { homedir } from 'node:os';
import { createHash } from 'node:crypto';
import { execFile, spawn } from 'node:child_process';
import { promisify } from 'node:util';
import { dataPath } from './config/instance.js';
import type {
DockerCase,
DockerCommandMode,
DockerEngine,
DockerHost,
DockerNetworkMode,
DockerResourceLimits,
SessionDocker,
SessionMode,
} from './types.js';
const execFileAsync = promisify(execFile);
/** Under vitest, all real `docker` invocations no-op (mirror of tmux-manager's IS_TEST_MODE). */
const IS_TEST_MODE = !!process.env.VITEST;
const DOCKER_HOSTS_FILE = 'docker-hosts.json';
const DOCKER_CASES_FILE = 'docker-cases.json';
/** Locally-built base image (see scripts/build-agent-image.mjs). */
export const DEFAULT_AGENT_IMAGE = 'codeman/agent:base';
/** HOME inside the base image (the `agent` user). Cred mounts + hook-secret land under it. */
export const CONTAINER_HOME = '/home/agent';
/** Per-case container name prefix. The `case` letters deliberately do NOT matter to
* tmux; this is a DOCKER name (`^[a-zA-Z0-9][a-zA-Z0-9_.-]+$`), and case names are
* already validated `^[a-zA-Z0-9_-]+$`, so `codeman-case-<name>` is always valid. */
const CONTAINER_NAME_PREFIX = 'codeman-case-';
/** Sensible resource defaults (all overridable per host). */
export const DEFAULT_DOCKER_RESOURCES: DockerResourceLimits = {
memory: '4g',
cpus: '2',
pidsLimit: 512,
nofile: '4096:8192',
};
// ========== Storage (mirror of remote-hosts.ts) ==========
export function dockerHostsPath(configDir: string): string {
return join(configDir, DOCKER_HOSTS_FILE);
}
export function dockerCasesPath(configDir: string): string {
return join(configDir, DOCKER_CASES_FILE);
}
async function readJsonArray<T>(path: string): Promise<T[]> {
try {
const raw = await fs.readFile(path, 'utf-8');
const parsed = JSON.parse(raw);
return Array.isArray(parsed) ? (parsed as T[]) : [];
} catch {
return [];
}
}
async function writeJsonArray<T>(configDir: string, path: string, value: T[]): Promise<void> {
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
await fs.writeFile(path, JSON.stringify(value, null, 2));
}
export async function readDockerHosts(configDir: string): Promise<DockerHost[]> {
return readJsonArray<DockerHost>(dockerHostsPath(configDir));
}
export async function writeDockerHosts(configDir: string, hosts: DockerHost[]): Promise<void> {
await writeJsonArray(configDir, dockerHostsPath(configDir), hosts);
}
export async function readDockerCases(configDir: string): Promise<DockerCase[]> {
return readJsonArray<DockerCase>(dockerCasesPath(configDir));
}
export async function writeDockerCases(configDir: string, cases: DockerCase[]): Promise<void> {
await writeJsonArray(configDir, dockerCasesPath(configDir), cases);
}
// ========== Naming / display / defaults ==========
/** Per-case container name. Mirrors how remote derives a stable name from the case. */
export function dockerContainerName(caseName: string): string {
return `${CONTAINER_NAME_PREFIX}${caseName}`;
}
/** Default pane command per CLI mode (mirror of defaultRemoteCommandForMode). */
export function defaultDockerCommandForMode(mode: SessionMode): string {
const commands: Record<DockerCommandMode, string> = {
shell: 'exec bash -l',
// Mirror the LOCAL claude default so the in-container agent runs non-interactively.
claude: 'exec claude --dangerously-skip-permissions',
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
};
return commands[mode as DockerCommandMode] || commands.shell;
}
/** `container:/workdir` display string (mirror of remoteDisplayPath's `user@host:path`). */
export function dockerDisplayPath(
docker: Pick<SessionDocker, 'containerName' | 'containerWorkdir'> | { container: string; path: string }
): string {
if ('containerName' in docker) return `${docker.containerName}:${docker.containerWorkdir}`;
return `${docker.container}:${docker.path}`;
}
/**
* The host-callback gateway alias is ENGINE-SPECIFIC: Docker exposes the host as
* `host.docker.internal`, Podman as `host.containers.internal`. Both are added to
* the host-guard allowlist so a mixed fleet keeps working.
*/
export function hostGatewayAlias(engine: DockerEngine): string {
return engine === 'podman' ? 'host.containers.internal' : 'host.docker.internal';
}
/**
* Rewrite the server's own `CODEMAN_API_URL` to a container-reachable one by
* swapping ONLY the hostname for the engine's host-gateway alias, preserving
* scheme AND port (prod is HTTPS on 3000, so hardcoding http://…:3000 breaks
* every hook). Falls back to `https://<alias>:3000` when the input is absent or
* unparseable.
*/
export function containerApiUrl(processApiUrl: string | undefined, engine: DockerEngine): string {
const alias = hostGatewayAlias(engine);
if (!processApiUrl) return `https://${alias}:3000`;
try {
const url = new URL(processApiUrl);
url.hostname = alias;
// origin drops any trailing path/slash and keeps scheme + (non-default) port
return url.origin;
} catch {
return `https://${alias}:3000`;
}
}
/**
* Stable hash of the drift-relevant `docker create` inputs, stored on the
* container as the `codeman.confighash` label. On launch, a mismatch between the
* desired hash and the running container's label triggers the recreate-on-drift
* prompt (host config edits actually take effect).
*/
export function dockerConfigHash(
docker: Pick<
SessionDocker,
| 'engine'
| 'image'
| 'containerWorkdir'
| 'network'
| 'networkName'
| 'resources'
| 'gpus'
| 'mountCredentials'
| 'extraCreateArgs'
>
): string {
const normalized = JSON.stringify({
engine: docker.engine,
image: docker.image,
containerWorkdir: docker.containerWorkdir,
network: docker.network,
networkName: docker.networkName ?? null,
resources: docker.resources ?? null,
gpus: docker.gpus ?? null,
mountCredentials: docker.mountCredentials,
extraCreateArgs: docker.extraCreateArgs ?? null,
});
return createHash('sha256').update(normalized).digest('hex').slice(0, 12);
}
/**
* Build the flattened per-session Docker metadata from a host profile + a case,
* resolving every default (mirror of toSessionRemote). The `configHash` is
* computed last over the resolved values.
*/
export function toSessionDocker(host: DockerHost, dockerCase: DockerCase): SessionDocker {
const engine: DockerEngine = host.engine ?? 'docker';
const containerWorkdir = dockerCase.containerWorkdir ?? dockerCase.hostWorkspacePath;
const base: Omit<SessionDocker, 'configHash'> = {
hostId: host.id,
label: host.label,
engine,
image: host.image || DEFAULT_AGENT_IMAGE,
containerName: dockerCase.container ?? dockerContainerName(dockerCase.name),
hostWorkspacePath: dockerCase.hostWorkspacePath,
containerWorkdir,
network: host.network ?? 'bridge',
networkName: host.networkName,
resources: host.resources ?? DEFAULT_DOCKER_RESOURCES,
gpus: host.gpus,
mountCredentials: host.mountCredentials ?? true,
hooksEnabled: host.hooksEnabled ?? true,
resumeOnStart: host.resumeOnStart ?? true,
daemonHost: host.daemonHost,
context: host.context,
commands: host.commands,
extraCreateArgs: host.extraCreateArgs,
extraExecArgs: host.extraExecArgs,
};
return { ...base, configHash: dockerConfigHash(base) };
}
// ========== Shell escaping ==========
/**
* POSIX single-quote shell-escaping (end-quote, escaped-quote, restart-quote).
* Mirror of the helper in remote-hosts.ts / tmux-manager.ts. Every dynamic value
* interpolated into the outer `bash -c "..."` launch layer is escaped through
* this so a path with spaces stays a single shell token. Operator-entered fields
* are ALSO schema-rejected for `$`/backtick (NO_SHELL_META) as defense in depth.
*/
export function shellescape(str: string): string {
return "'" + str.replace(/'/g, "'\\''") + "'";
}
// ========== Pure command-arg builders ==========
/** A resolved bind mount (source existence already checked by the caller). */
export interface DockerMount {
src: string;
dst: string;
readonly?: boolean;
}
/**
* Resolved, IO-free context for buildDockerCreateArgs. The caller (tmux-manager)
* resolves the environment-dependent bits (host uid, existing cred mounts, the
* derived api url, Desktop detection) so this builder stays pure and unit-testable.
*/
export interface DockerCreateContext {
docker: SessionDocker;
/** Codeman session id (only the first 8 chars are used, for the codeman.session label). */
sessionId: string;
/** CODEMAN_INSTANCE ('' for prod) — scopes the boot reaper so a beta never reaps prod. */
instance: string;
/** Pre-resolved uid/userns tokens: ['--user','1000:0'] | ['--userns','keep-id'] | []. */
userArgs: string[];
/** Existing host credential bind mounts (convenient mode). Empty in sealed mode. */
credentialMounts: DockerMount[];
/** Extra bind mounts (e.g. the read-only hook-secret file). */
extraMounts: DockerMount[];
/** Create-time env (NON-secret, committed-safe): HOME, TERM, COLORTERM, CODEMAN_API_URL, CODEMAN_HOOK_SECRET_FILE. */
envCreate: Record<string, string>;
/** Whether to add `--add-host <alias>:host-gateway` (skipped on Docker Desktop, where the alias is native). */
addHostGateway: boolean;
/** Engine host-gateway alias (host.docker.internal / host.containers.internal). */
gatewayAlias: string;
}
/**
* Engine prefix tokens shared by every docker invocation (mirror of
* buildSshConnectionArgs). Returns e.g. ['docker'] or ['podman','--context','ctx'].
*/
export function buildDockerBaseArgs(docker: Pick<SessionDocker, 'engine' | 'context' | 'daemonHost'>): string[] {
const parts: string[] = [docker.engine === 'podman' ? 'podman' : 'docker'];
if (docker.context) parts.push('--context', shellescape(docker.context));
if (docker.daemonHost) parts.push('-H', shellescape(docker.daemonHost));
return parts;
}
function mountSpec(m: DockerMount): string {
return `type=bind,src=${m.src},dst=${m.dst}${m.readonly ? ',readonly' : ''}`;
}
function resourceFlags(resources?: DockerResourceLimits): string[] {
if (!resources) return [];
const flags: string[] = [];
if (resources.memory) {
// memory-swap == memory disables swap, making --memory a REAL OOM cap.
flags.push('--memory', resources.memory, '--memory-swap', resources.memory);
}
if (resources.cpus) flags.push('--cpus', resources.cpus);
if (resources.pidsLimit) flags.push('--pids-limit', String(resources.pidsLimit));
if (resources.nofile) flags.push('--ulimit', `nofile=${resources.nofile}`);
if (resources.shmSize) flags.push('--shm-size', resources.shmSize);
return flags;
}
function networkArg(network: DockerNetworkMode, networkName?: string): string {
if (network === 'custom' && networkName) return networkName;
return network; // 'bridge' | 'none'
}
/**
* Build the `docker create` token list (from `create` through the `sleep
* infinity` CMD) for a long-lived, hardened, per-case container. PURE: every
* dynamic value is shellescaped; the caller joins with spaces into the launch
* string. Security invariants baked in: --cap-drop ALL, --security-opt
* no-new-privileges, --pids-limit, --memory==--memory-swap, --init,
* --pull=never, --restart no, NEVER --privileged, NEVER the docker socket.
*/
export function buildDockerCreateArgs(ctx: DockerCreateContext): string[] {
const {
docker,
sessionId,
instance,
userArgs,
credentialMounts,
extraMounts,
envCreate,
addHostGateway,
gatewayAlias,
} = ctx;
const args: string[] = [
'create',
'--name',
shellescape(docker.containerName),
'--label',
'codeman.managed=1',
'--label',
shellescape(`codeman.instance=${instance}`),
'--label',
shellescape(`codeman.session=${sessionId.slice(0, 8)}`),
'--label',
shellescape(`codeman.confighash=${docker.configHash ?? dockerConfigHash(docker)}`),
'--pull=never',
'--init',
'--restart',
'no',
...userArgs,
'--workdir',
shellescape(docker.containerWorkdir),
// Workspace bind: mirror the host path inside the container so the transcript
// projHash correlates and file features read real host bytes.
'--mount',
shellescape(mountSpec({ src: docker.hostWorkspacePath, dst: docker.containerWorkdir })),
...credentialMounts.flatMap((m) => ['--mount', shellescape(mountSpec(m))]),
...extraMounts.flatMap((m) => ['--mount', shellescape(mountSpec(m))]),
];
if (addHostGateway) args.push('--add-host', `${gatewayAlias}:host-gateway`);
args.push(
...resourceFlags(docker.resources),
// GPU passthrough (needs the NVIDIA container toolkit on the host). No storage
// cap is set, so the container's writable layer + volumes grow elastically as
// data flows in (bounded only by host disk).
...(docker.gpus ? ['--gpus', shellescape(docker.gpus)] : []),
'--cap-drop',
'ALL',
'--security-opt',
'no-new-privileges',
'--network',
networkArg(docker.network, docker.networkName)
);
for (const [key, value] of Object.entries(envCreate)) {
args.push('--env', shellescape(`${key}=${value}`));
}
// Operator escape-hatch args (schema-validated NO_SHELL_INJECTION), escaped again here.
for (const extra of docker.extraCreateArgs ?? []) {
args.push(shellescape(extra));
}
args.push(shellescape(docker.image), 'sleep', 'infinity');
return args;
}
/**
* PURE argv for building the agent base image locally (the programmatic mirror of
* scripts/build-agent-image.mjs): `build -f <dockerfile> -t <image> [--no-cache]
* <contextDir>`. Kept pure + unit-testable; the caller prepends the engine binary.
*/
export function agentImageBuildArgs(dockerfile: string, image: string, contextDir: string, noCache = false): string[] {
return ['build', '-f', dockerfile, '-t', image, ...(noCache ? ['--no-cache'] : []), contextDir];
}
// ========== Credential mount resolution (IO) ==========
/** Container Claude config dir (created gid-0 writable in the image). */
export const CONTAINER_CLAUDE_DIR = `${CONTAINER_HOME}/.claude`;
/** In-container path of the seeded (writable) `~/.claude.json`. */
export const CLAUDE_JSON_HOME = `${CONTAINER_HOME}/.claude.json`;
/** In-container path of the read-only host-seeded `~/.claude.json` (copied into HOME at launch). */
export const CLAUDE_JSON_SEED = `${CONTAINER_HOME}/.codeman/claude.seed.json`;
/** Read-only seed paths for the files copied into the container's `.claude`. */
const CLAUDE_CREDS_SEED = `${CONTAINER_HOME}/.codeman/claude-creds.seed.json`;
const CLAUDE_SETTINGS_SEED = `${CONTAINER_HOME}/.codeman/claude-settings.seed.json`;
const CLAUDE_STATS_SEED = `${CONTAINER_HOME}/.codeman/claude-stats.seed.json`;
/** Staging root for read-only host-cred seed mounts (codex/gemini/gcloud/opencode). */
const CRED_SEED_DIR = `${CONTAINER_HOME}/.codeman/cred-seeds`;
/**
* PURE: merge the host `~/.claude.json` into a config that makes an
* already-authenticated Claude skip its INTERACTIVE onboarding inside the container
* (the host file itself lacks these flags — the host install is grandfathered, so a
* verbatim copy still triggers the theme picker + login wizard + folder-trust
* prompt). Forces `hasCompletedOnboarding`, a `theme` (so the theme picker is
* skipped), and marks the workspace project trusted + onboarded. Auth still comes
* from the copied `oauthAccount` + the dir-mounted `~/.claude/.credentials.json`.
*/
export function buildSeamlessClaudeConfig(
hostConfig: Record<string, unknown>,
workspacePath: string,
theme = 'dark'
): Record<string, unknown> {
const merged: Record<string, unknown> = { ...hostConfig };
merged.hasCompletedOnboarding = true;
if (typeof merged.theme !== 'string') merged.theme = theme;
const projects = { ...((merged.projects as Record<string, Record<string, unknown>> | undefined) ?? {}) };
const existing = (projects[workspacePath] as Record<string, unknown> | undefined) ?? {};
const seenCount = existing.projectOnboardingSeenCount;
projects[workspacePath] = {
...existing,
hasTrustDialogAccepted: true,
hasCompletedProjectOnboarding: true,
projectOnboardingSeenCount: typeof seenCount === 'number' && seenCount > 0 ? seenCount : 1,
};
merged.projects = projects;
return merged;
}
/** Best-effort read of the host `~/.claude/settings.json` theme (drives the seed's theme). */
function readHostClaudeTheme(home: string): string | undefined {
try {
const parsed = JSON.parse(readFileSync(join(home, '.claude', 'settings.json'), 'utf-8')) as { theme?: unknown };
return typeof parsed.theme === 'string' ? parsed.theme : undefined;
} catch {
return undefined;
}
}
/**
* Resolve the read-only seed mount for `~/.claude.json`. Reads the host file, merges
* in the seamless-onboarding flags + workspace trust (buildSeamlessClaudeConfig),
* writes the result to a per-container seed file under `~/.codeman/docker-seeds/`,
* and returns its mount. The launch chain copies it to `~/.claude.json` inside HOME
* once — giving Claude a NORMAL writable, already-onboarded config (no atomic-rename
* EBUSY, no re-auth, no theme/trust prompts). Falls back to the RAW host file when
* parse/write fails (auth still works; the wizard may show). Returns null when the
* host has no `~/.claude.json`. IO; under VITEST returns the raw mount (no write).
*/
export function resolveClaudeJsonSeedMount(
home: string = homedir(),
containerName?: string,
workspacePath?: string
): DockerMount | null {
const src = join(home, '.claude.json');
if (!existsSync(src)) return null;
const rawMount: DockerMount = { src, dst: CLAUDE_JSON_SEED, readonly: true };
if (IS_TEST_MODE || !containerName || !workspacePath) return rawMount;
try {
const hostConfig = JSON.parse(readFileSync(src, 'utf-8')) as Record<string, unknown>;
const merged = buildSeamlessClaudeConfig(hostConfig, workspacePath, readHostClaudeTheme(home) ?? 'dark');
const seedsDir = dataPath('docker-seeds');
if (!existsSync(seedsDir)) mkdirSync(seedsDir, { recursive: true });
const seedFile = join(seedsDir, `${containerName}.json`);
writeFileSync(seedFile, JSON.stringify(merged), { mode: 0o600 });
return { src: seedFile, dst: CLAUDE_JSON_SEED, readonly: true };
} catch {
return rawMount; // partial host write / unreadable — auth still carries, wizard may show
}
}
/** A file (or dir, when `recursive`) copied into the container HOME once at launch
* (`[ -e to ] || cp [-a] from to`). */
export interface DockerSeedCopy {
from: string;
to: string;
/** `cp -a` for whole-directory credential seeds (gemini/gcloud/opencode). */
recursive?: boolean;
}
export interface DockerClaudeArtifacts {
/** Bind mounts to add: the shared `projects/` transcripts (RW) + read-only seed files. */
mounts: DockerMount[];
/** Files copied into the container's writable HOME/.claude (+ HOME/.claude.json) at launch. */
seedCopies: DockerSeedCopy[];
}
/**
* Resolve the ISOLATED Claude artifacts for a docker session (replaces the old
* whole-`~/.claude` RW mount that polluted the host). Shares ONLY what must cross
* the boundary and seeds the rest as writable copies:
* - `~/.claude/projects` → RW dir mount (transcripts: host watchers + `--resume`).
* - `~/.claude.json` → merged onboarding seed, copied to HOME (no re-auth/wizard).
* - `~/.claude/.credentials.json` + `~/.claude/settings.json` → read-only seeds
* copied into the container's own `~/.claude` (token + global prefs carry in;
* the container refreshes its own copy and never writes back to the host).
* Everything else Claude writes (backups, tasks, teams, session-env, history) stays
* container-local. IO (reads host files, writes the merged `.claude.json` seed).
*/
export function resolveDockerClaudeArtifacts(
home: string,
containerName: string,
workspacePath: string
): DockerClaudeArtifacts {
const mounts: DockerMount[] = [];
const seedCopies: DockerSeedCopy[] = [];
// The ONE genuinely-shared part: conversation transcripts (dir mount → renames work).
const projectsSrc = join(home, '.claude', 'projects');
if (existsSync(projectsSrc)) {
mounts.push({ src: projectsSrc, dst: `${CONTAINER_CLAUDE_DIR}/projects` });
}
// ~/.claude.json → merged, onboarding-complete seed at HOME root.
const jsonSeed = resolveClaudeJsonSeedMount(home, containerName, workspacePath);
if (jsonSeed) {
mounts.push(jsonSeed);
seedCopies.push({ from: CLAUDE_JSON_SEED, to: CLAUDE_JSON_HOME });
}
// credentials (token) + settings (theme/model/effort/permissions) + stats-cache
// (drives the model/effort status indicator) → writable copies inside the
// container's own ~/.claude (never a wholesale mount → no host pollution).
const files: Array<[rel: string, seed: string, dest: string]> = [
['.credentials.json', CLAUDE_CREDS_SEED, `${CONTAINER_CLAUDE_DIR}/.credentials.json`],
['settings.json', CLAUDE_SETTINGS_SEED, `${CONTAINER_CLAUDE_DIR}/settings.json`],
['stats-cache.json', CLAUDE_STATS_SEED, `${CONTAINER_CLAUDE_DIR}/stats-cache.json`],
];
for (const [rel, seed, dest] of files) {
const src = join(home, '.claude', rel);
if (existsSync(src)) {
mounts.push({ src, dst: seed, readonly: true });
seedCopies.push({ from: seed, to: dest });
}
}
return { mounts, seedCopies };
}
/**
* Per-CLI credential-store isolation policy (the codex/gemini/gcloud/opencode analog
* of resolveDockerClaudeArtifacts). Codex is the direct Claude-analog: its
* `sessions/` rollouts + `history.jsonl` are read HOST-SIDE (response-viewer +
* `codex resume`), so they are SHARED (RW), while `auth.json`/`config.toml` are
* seeded. The other three have no host-read/resume dependency and are fully
* seed-copied (writable copy in the container, no write-back to the host).
*/
interface CredStorePolicy {
/** Path relative to HOME (host + container), e.g. '.codex' or '.config/gcloud'. */
rel: string;
/** Subdirs bind-mounted RW (shared: resume + host reads). */
shareDirs?: string[];
/** Files bind-mounted RW (append-only, e.g. codex history.jsonl — never renamed). */
shareFiles?: string[];
/** Files seeded (RO mount → cp) into the container's own copy. */
seedFiles?: string[];
/** Seed the WHOLE dir (RO mount → cp -a) — for stores with no shared/host-read state. */
seedWhole?: boolean;
}
const CRED_STORES: CredStorePolicy[] = [
{ rel: '.codex', shareDirs: ['sessions'], shareFiles: ['history.jsonl'], seedFiles: ['auth.json', 'config.toml'] },
{ rel: '.gemini', seedWhole: true },
{ rel: '.config/gcloud', seedWhole: true },
{ rel: '.config/opencode', seedWhole: true },
];
/**
* Resolve the ISOLATED codex/gemini/gcloud/opencode artifacts (replaces the old
* whole-dir RW mounts that let each in-container CLI write its refreshed tokens +
* session state back into the host). Every path is existsSync-gated (on most hosts
* only a subset exists). Pure-ish IO (no writes; just existence checks + mount specs).
*/
export function resolveDockerCredentialArtifacts(home: string = homedir()): DockerClaudeArtifacts {
const mounts: DockerMount[] = [];
const seedCopies: DockerSeedCopy[] = [];
for (const store of CRED_STORES) {
const hostBase = join(home, store.rel);
if (!existsSync(hostBase)) continue;
const containerBase = `${CONTAINER_HOME}/${store.rel}`;
const seedName = store.rel.replace(/\//g, '-'); // '.config/gcloud' → '.config-gcloud'
if (store.seedWhole) {
const seed = `${CRED_SEED_DIR}/${seedName}`;
mounts.push({ src: hostBase, dst: seed, readonly: true });
seedCopies.push({ from: seed, to: containerBase, recursive: true });
continue;
}
for (const sub of store.shareDirs ?? []) {
const src = join(hostBase, sub);
if (existsSync(src)) mounts.push({ src, dst: `${containerBase}/${sub}` });
}
for (const file of store.shareFiles ?? []) {
const src = join(hostBase, file);
if (existsSync(src)) mounts.push({ src, dst: `${containerBase}/${file}` });
}
for (const file of store.seedFiles ?? []) {
const src = join(hostBase, file);
if (existsSync(src)) {
const seed = `${CRED_SEED_DIR}/${seedName}-${file}`;
mounts.push({ src, dst: seed, readonly: true });
seedCopies.push({ from: seed, to: `${containerBase}/${file}` });
}
}
}
return { mounts, seedCopies };
}
// ========== Daemon probes (IO; no-op under VITEST) ==========
export interface DockerAvailability {
ok: boolean;
engine: DockerEngine;
rootless: boolean;
isDesktop: boolean;
cgroupV2: boolean;
/** Best-effort: are --memory/--cpus/--pids-limit actually enforced on this engine? */
capsEnforced: boolean;
error?: string;
}
const DOCKER_PROBE_TIMEOUT_MS = 15_000;
interface DockerInfoJson {
ServerVersion?: string;
CgroupVersion?: string;
SecurityOptions?: string[];
OperatingSystem?: string;
OSType?: string;
Name?: string;
}
async function runDockerInfo(engine: DockerEngine): Promise<DockerInfoJson | null> {
try {
const { stdout } = await execFileAsync(engine, ['info', '--format', '{{json .}}'], {
timeout: DOCKER_PROBE_TIMEOUT_MS,
});
return JSON.parse(stdout) as DockerInfoJson;
} catch {
return null;
}
}
function classifyDockerInfo(engine: DockerEngine, info: DockerInfoJson): DockerAvailability {
const security = info.SecurityOptions ?? [];
const rootless = security.some((opt) => opt.includes('rootless'));
const cgroupV2 = info.CgroupVersion === '2';
const os = `${info.OperatingSystem ?? ''}`.toLowerCase();
const isDesktop = os.includes('docker desktop') || os.includes('desktop');
// Under rootless, resource caps are only reliably enforced with cgroup v2 +
// systemd delegation. We can't detect delegation from `docker info`, so we
// treat rootless+cgroupv2 as "likely enforced" and rootless+cgroupv1 as not.
const capsEnforced = !rootless || cgroupV2;
return { ok: true, engine, rootless, isDesktop, cgroupV2, capsEnforced };
}
/**
* Probe the container engine: server up, cgroup version, rootless, Desktop, and
* whether resource caps are enforceable. Auto-detects docker then podman when no
* engine is given. No-op canned value under VITEST.
*/
export async function checkDockerAvailable(engine?: DockerEngine): Promise<DockerAvailability> {
if (IS_TEST_MODE) {
return {
ok: true,
engine: engine ?? 'docker',
rootless: false,
isDesktop: false,
cgroupV2: true,
capsEnforced: true,
};
}
const candidates: DockerEngine[] = engine ? [engine] : ['docker', 'podman'];
for (const candidate of candidates) {
const info = await runDockerInfo(candidate);
if (info) return classifyDockerInfo(candidate, info);
}
return {
ok: false,
engine: engine ?? 'docker',
rootless: false,
isDesktop: false,
cgroupV2: false,
capsEnforced: false,
error: 'Docker/Podman not available. Install docker (or podman) and ensure the daemon is running.',
};
}
/** Is the base image present locally? (never triggers an auto-pull). */
export async function checkDockerImagePresent(engine: DockerEngine, image: string): Promise<boolean> {
if (IS_TEST_MODE) return true;
try {
await execFileAsync(engine, ['image', 'inspect', '--format', '{{.Id}}', image], {
timeout: DOCKER_PROBE_TIMEOUT_MS,
});
return true;
} catch {
return false;
}
}
export interface EnsureImageResult {
ok: boolean;
/** true when this call actually ran a build (vs. the image already existing). */
built: boolean;
alreadyPresent: boolean;
error?: string;
}
/** In-flight builds keyed by `engine:image`, so concurrent callers share ONE build. */
const inFlightImageBuilds = new Map<string, Promise<EnsureImageResult>>();
/**
* Resolve the repo's Dockerfile + build context. Works from BOTH src (dev/tsx) and
* dist/index.js (esbuild prod: dist sits at repo root), since both are one level
* under the repo root. Returns null when the Dockerfile is absent (npm-global
* installs don't ship docker/ — Docker cases are a git-clone feature).
*/
function resolveAgentDockerfile(): { dockerfile: string; contextDir: string } | null {
const repoRoot = join(dirname(fileURLToPath(import.meta.url)), '..');
const dockerfile = join(repoRoot, 'docker', 'agent.Dockerfile');
return existsSync(dockerfile) ? { dockerfile, contextDir: repoRoot } : null;
}
/**
* Ensure the agent base image exists, BUILDING it locally on first use so a missing
* image is never a hard blocker (decision: "build locally on first use",
* docs/docker-cases-plan.md). Idempotent, concurrency-safe (one build per
* engine:image shared by concurrent callers), and a no-op under VITEST. Only the
* DEFAULT image is auto-built — we can never build a user's custom ref, and the
* `--pull=never` invariant forbids pulling. `onProgress` receives build output
* lines for SSE surfacing.
*/
export async function ensureAgentBaseImage(
engine: DockerEngine,
image: string,
opts: { onProgress?: (line: string) => void; noCache?: boolean } = {}
): Promise<EnsureImageResult> {
if (IS_TEST_MODE) return { ok: true, built: false, alreadyPresent: true };
if (await checkDockerImagePresent(engine, image)) {
return { ok: true, built: false, alreadyPresent: true };
}
if (image !== DEFAULT_AGENT_IMAGE) {
return {
ok: false,
built: false,
alreadyPresent: false,
error: `image ${image} is not present and only ${DEFAULT_AGENT_IMAGE} is auto-built. Build or pull ${image} yourself.`,
};
}
const key = `${engine}:${image}`;
const existing = inFlightImageBuilds.get(key);
if (existing) return existing;
const build = buildAgentImage(engine, image, opts).finally(() => inFlightImageBuilds.delete(key));
inFlightImageBuilds.set(key, build);
return build;
}
function buildAgentImage(
engine: DockerEngine,
image: string,
opts: { onProgress?: (line: string) => void; noCache?: boolean }
): Promise<EnsureImageResult> {
const resolved = resolveAgentDockerfile();
if (!resolved) {
return Promise.resolve({
ok: false,
built: false,
alreadyPresent: false,
error: `docker/agent.Dockerfile not found in this install; clone the repo or build ${image} manually`,
});
}
const args = agentImageBuildArgs(resolved.dockerfile, image, resolved.contextDir, opts.noCache);
return new Promise<EnsureImageResult>((resolve) => {
// async spawn (NEVER spawnSync) so a multi-minute build never wedges the event loop.
const child = spawn(engine, args, { stdio: ['ignore', 'pipe', 'pipe'] });
const forward = (buf: Buffer) => {
for (const line of buf.toString('utf-8').split('\n')) {
const trimmed = line.trimEnd();
if (trimmed) opts.onProgress?.(trimmed);
}
};
child.stdout?.on('data', forward);
child.stderr?.on('data', forward);
child.on('error', (err) => {
resolve({
ok: false,
built: false,
alreadyPresent: false,
error: `could not spawn ${engine} build: ${err.message}`,
});
});
child.on('exit', (code) => {
if (code === 0) resolve({ ok: true, built: true, alreadyPresent: false });
else resolve({ ok: false, built: false, alreadyPresent: false, error: `${engine} build failed (exit ${code})` });
});
});
}
export interface DockerTmuxCheckResult {
ok: boolean;
tmuxPath?: string;
/** Distinguishes "image missing" (build it) from "tmux missing in image" (rebuild it). */
imageMissing?: boolean;
error?: string;
}
/**
* Verify the base image is present AND contains tmux (a HARD prerequisite: the
* in-container tmux is what makes reconnect durable). Never triggers a pull
* (`--pull=never`). No-op under VITEST. Mirror of checkRemoteTmuxAvailable.
*/
export async function checkDockerTmuxAvailable(
docker: Pick<SessionDocker, 'engine' | 'image'>
): Promise<DockerTmuxCheckResult> {
if (IS_TEST_MODE) return { ok: true, tmuxPath: '/usr/bin/tmux' };
const engine = docker.engine;
if (!(await checkDockerImagePresent(engine, docker.image))) {
return {
ok: false,
imageMissing: true,
error: `image ${docker.image} not present (the default image is auto-built on first use; a custom image must be built or pulled first)`,
};
}
try {
const { stdout } = await execFileAsync(
engine,
['run', '--rm', '--pull=never', docker.image, 'sh', '-lc', 'command -v tmux'],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const tmuxPath = stdout.trim();
if (!tmuxPath) {
return { ok: false, error: `base image ${docker.image} is missing tmux (required for durable sessions)` };
}
return { ok: true, tmuxPath };
} catch (err) {
const msg = err instanceof Error ? err.message : String(err);
return { ok: false, error: `could not verify tmux in ${docker.image}: ${msg}` };
}
}
/**
* Resolve the host's IP on the default docker bridge (the address a container
* reaches as `host.docker.internal`), so the server can bind a hooks-only listener
* there and in-container hooks can call back. Defaults to the conventional
* 172.17.0.1 when the inspect fails but docker is up; null when docker is absent.
* No-op canned value under VITEST.
*/
export async function detectDockerBridgeGateway(engine: DockerEngine = 'docker'): Promise<string | null> {
if (IS_TEST_MODE) return '172.17.0.1';
const bin = engine === 'podman' ? 'podman' : 'docker';
try {
const { stdout } = await execFileAsync(
bin,
['network', 'inspect', 'bridge', '--format', '{{(index .IPAM.Config 0).Gateway}}'],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
const ip = stdout.trim();
return /^\d{1,3}(\.\d{1,3}){3}$/.test(ip) ? ip : '172.17.0.1';
} catch {
return null; // docker not available — nothing to bind
}
}
/**
* Instance-scoped boot reaper: `docker rm -f` any MANAGED container that belongs
* to THIS instance (by the `codeman.instance` label) but whose case is no longer
* in `docker-cases.json`. The instance scoping is what stops a beta from reaping
* prod's containers (the cross-instance hazard). No-op under VITEST. Best-effort.
*/
export async function reapOrphanedDockerContainers(
configDir: string,
instance: string,
engine: DockerEngine = 'docker'
): Promise<string[]> {
if (IS_TEST_MODE) return [];
const bin = engine === 'podman' ? 'podman' : 'docker';
let rows: Array<{ name: string; inst: string }> = [];
try {
const { stdout } = await execFileAsync(
bin,
[
'ps',
'-a',
'--filter',
'label=codeman.managed=1',
'--format',
'{{.Names}}\t{{index .Labels "codeman.instance"}}',
],
{ timeout: DOCKER_PROBE_TIMEOUT_MS }
);
rows = stdout
.split('\n')
.filter(Boolean)
.map((line) => {
const [name, inst = ''] = line.split('\t');
return { name, inst };
});
} catch {
return []; // daemon down / engine absent — nothing to reap
}
const cases = await readDockerCases(configDir);
const expected = new Set(cases.map((c) => c.container ?? dockerContainerName(c.name)));
const reaped: string[] = [];
for (const { name, inst } of rows) {
if (inst !== instance) continue; // only THIS instance's containers
if (expected.has(name)) continue; // still referenced by a live case
try {
await execFileAsync(bin, ['rm', '-f', name], { timeout: DOCKER_PROBE_TIMEOUT_MS });
reaped.push(name);
} catch {
/* best-effort */
}
}
return reaped;
}
/**
* Read the IN-CONTAINER Claude CLI version (`docker exec <container> claude
* --version`). Feeds Session.cliVersion for docker sessions (the LOCAL claude
* would report the wrong version and disable trackpad wheel-forwarding, #154).
* Returns undefined on any failure. No-op under VITEST.
*/
export async function probeDockerCliVersion(
docker: Pick<SessionDocker, 'engine' | 'containerName'>,
mode: SessionMode
): Promise<string | undefined> {
if (IS_TEST_MODE) return undefined;
const bin = mode === 'shell' ? null : mode;
if (!bin) return undefined;
try {
const { stdout } = await execFileAsync(docker.engine, ['exec', docker.containerName, bin, '--version'], {
timeout: DOCKER_PROBE_TIMEOUT_MS,
});
const match = stdout.trim().match(/\d+\.\d+\.\d+/);
return match ? match[0] : stdout.trim() || undefined;
} catch {
return undefined;
}
}
+67
View File
@@ -0,0 +1,67 @@
/**
* @fileoverview Global concurrency limiter for spawning external document
* converters (pdftoppm / LibreOffice `soffice` / Word-COM `powershell.exe`).
*
* Without a cap, N simultaneous thumbnail/preview requests for *distinct*
* documents fork N converter processes at once — each held open for up to the
* multi-minute conversion timeout. That is a localhost resource-exhaustion
* (fork-bomb-shaped) vector: a handful of large PDFs detected at once can pin
* CPU and RAM. This module serializes converter spawns down to a small fixed
* pool; excess spawns queue (FIFO) until a slot frees. The in-flight cache in
* `document-preview-cache.ts` already de-dups *identical* inputs; this bounds
* the *distinct* case the cache can't.
*
* Permit accounting transfers the slot directly to the next waiter on release
* (rather than decrement-then-reacquire) so the active count can never exceed
* the cap even under interleaved async resumption.
*
* NOT re-entrant: never call `runWithConversionLimit` from inside a task that is
* already holding a slot — a nested acquire under a full pool would deadlock.
* The converter call sites only ever acquire once per request (the office path
* acquires for `soffice` and `pdftoppm` sequentially, not nested).
*/
/**
* Max converter processes allowed to run concurrently across the whole process.
* Override with CODEMAN_MAX_DOCUMENT_CONVERSIONS (clamped to >= 1).
*/
const MAX_CONCURRENT_DOCUMENT_CONVERSIONS = (() => {
const raw = Number(process.env.CODEMAN_MAX_DOCUMENT_CONVERSIONS);
return Number.isFinite(raw) && raw >= 1 ? Math.floor(raw) : 3;
})();
let active = 0;
const waiters: Array<() => void> = [];
/** Test/diagnostic hook: converters currently holding a slot. */
export function getActiveConversionCount(): number {
return active;
}
function acquire(): Promise<void> {
if (active < MAX_CONCURRENT_DOCUMENT_CONVERSIONS) {
active++;
return Promise.resolve();
}
return new Promise<void>((resolve) => waiters.push(resolve));
}
function release(): void {
const next = waiters.shift();
if (next) {
// Hand the slot straight to the next waiter — `active` stays at the cap.
next();
} else {
active--;
}
}
/** Run `task` once a converter slot is free, releasing the slot afterward. */
export async function runWithConversionLimit<T>(task: () => Promise<T>): Promise<T> {
await acquire();
try {
return await task();
} finally {
release();
}
}
+308
View File
@@ -0,0 +1,308 @@
/**
* @fileoverview Shared disk cache for expensive Office document previews.
*/
import { createHash } from 'node:crypto';
import { execFile } from 'node:child_process';
import fs from 'node:fs/promises';
import { tmpdir } from 'node:os';
import { basename, dirname, extname, join } from 'node:path';
import { pathToFileURL } from 'node:url';
import { promisify } from 'node:util';
import { runWithConversionLimit } from './document-conversion-limiter.js';
const execFileAsync = promisify(execFile);
const OFFICE_CONVERSION_TIMEOUT_MS = 5 * 60_000;
const DOCUMENT_PREVIEW_CACHE_DIR = join(tmpdir(), 'codeman-document-preview-cache');
/**
* Cap on persistent converted-PDF files kept in DOCUMENT_PREVIEW_CACHE_DIR.
* The cache key embeds the source mtime, so every edit to a doc orphans its
* prior PDF; without a cap the dir grows unbounded across long-running sessions.
* Override with CODEMAN_MAX_PREVIEW_CACHE_FILES (clamped to >= 1).
*/
const MAX_PREVIEW_CACHE_FILES = (() => {
const raw = Number(process.env.CODEMAN_MAX_PREVIEW_CACHE_FILES);
return Number.isFinite(raw) && raw >= 1 ? Math.floor(raw) : 100;
})();
function buildWordExportPdfScript(sourcePath: string, outputPath: string): string {
return `
$ErrorActionPreference = "Stop"
$source = ${toPowerShellSingleQuotedString(sourcePath)}
$output = ${toPowerShellSingleQuotedString(outputPath)}
$word = $null
$doc = $null
try {
$word = New-Object -ComObject Word.Application
$word.Visible = $false
$word.DisplayAlerts = 0
$doc = $word.Documents.Open($source)
$doc.ExportAsFixedFormat($output, 17)
} finally {
if ($null -ne $doc) {
$doc.Close($false) | Out-Null
[System.Runtime.InteropServices.Marshal]::ReleaseComObject($doc) | Out-Null
}
if ($null -ne $word) {
$word.Quit() | Out-Null
[System.Runtime.InteropServices.Marshal]::ReleaseComObject($word) | Out-Null
}
[System.GC]::Collect()
[System.GC]::WaitForPendingFinalizers()
}
`.trim();
}
type OfficePreviewConverter = 'msword' | 'libreoffice';
const inFlightOfficeConversions = new Map<string, Promise<string | null>>();
export function clearDocumentPreviewCache(): void {
inFlightOfficeConversions.clear();
}
/**
* Best-effort LRU-ish eviction for the persistent converted-PDF cache: keeps at
* most MAX_PREVIEW_CACHE_FILES `*.pdf` files in `cacheDir`, deleting the oldest
* by mtime once over the cap. Never throws — a pruning failure must not fail the
* conversion that triggered it. Only `*.pdf` files are considered, so the
* transient `work-*` mkdtemp dirs are ignored.
*/
export async function pruneDocumentPreviewCache(cacheDir: string): Promise<void> {
try {
const entries = await fs.readdir(cacheDir);
const pdfs = entries.filter((name) => name.toLowerCase().endsWith('.pdf'));
if (pdfs.length <= MAX_PREVIEW_CACHE_FILES) return;
const stats = await Promise.all(
pdfs.map(async (name) => {
const fullPath = join(cacheDir, name);
try {
const stat = await fs.stat(fullPath);
return { fullPath, mtimeMs: stat.mtimeMs ?? 0 };
} catch {
return null;
}
})
);
const sorted = stats.filter((s): s is { fullPath: string; mtimeMs: number } => s !== null);
sorted.sort((a, b) => a.mtimeMs - b.mtimeMs); // oldest first
const toRemove = sorted.slice(0, Math.max(0, sorted.length - MAX_PREVIEW_CACHE_FILES));
await Promise.all(toRemove.map((entry) => fs.rm(entry.fullPath, { force: true }).catch(() => {})));
} catch {
// Best-effort: pruning must never break a conversion.
}
}
export async function getOfficePreviewPdfPath(filePath: string, extension: string): Promise<string | null> {
const ext = extension.toLowerCase().replace(/^\./, '');
if (ext !== 'docx' && ext !== 'pptx') return null;
let sourceStat;
try {
sourceStat = await fs.stat(filePath);
} catch {
return null;
}
for (const converter of getOfficePreviewConverters(filePath, ext)) {
const cacheKey = createDocumentPreviewCacheKey(filePath, ext, sourceStat.size, sourceStat.mtimeMs ?? 0, converter);
const cachePath = getOfficePreviewCachePath(filePath, cacheKey, converter);
if (await fileExists(cachePath)) {
return cachePath;
}
const inFlightKey = `${converter}:${cacheKey}`;
const inFlight = inFlightOfficeConversions.get(inFlightKey);
if (inFlight) {
const converted = await inFlight;
if (converted) return converted;
continue;
}
const conversion =
converter === 'msword'
? convertWordDocumentToCachedPdf(filePath, cachePath)
: convertLibreOfficeDocumentToCachedPdf(filePath, cachePath);
inFlightOfficeConversions.set(inFlightKey, conversion);
try {
const converted = await conversion;
if (converted) return converted;
} finally {
inFlightOfficeConversions.delete(inFlightKey);
}
}
return null;
}
function getOfficePreviewConverters(filePath: string, extension: string): OfficePreviewConverter[] {
if (extension === 'docx' && wslMountPathToWindowsPath(filePath)) {
return ['msword', 'libreoffice'];
}
return ['libreoffice'];
}
function createDocumentPreviewCacheKey(
filePath: string,
extension: string,
size: number,
mtimeMs: number,
converter: OfficePreviewConverter
): string {
return createHash('sha256')
.update(JSON.stringify({ cacheVersion: 2, converter, filePath, extension, size, mtimeMs }))
.digest('hex')
.slice(0, 32);
}
function getOfficePreviewCachePath(filePath: string, cacheKey: string, converter: OfficePreviewConverter): string {
if (converter === 'msword') {
const windowsCacheDir = getWindowsUserTempCacheDir(filePath);
if (windowsCacheDir) {
return join(windowsCacheDir, `${cacheKey}.pdf`);
}
}
return join(DOCUMENT_PREVIEW_CACHE_DIR, `${cacheKey}.pdf`);
}
async function fileExists(filePath: string): Promise<boolean> {
try {
const stat = await fs.stat(filePath);
return typeof stat.isFile !== 'function' || stat.isFile();
} catch {
return false;
}
}
async function convertWordDocumentToCachedPdf(filePath: string, cachePath: string): Promise<string | null> {
const outputPath = wslMountPathToWindowsPath(cachePath);
if (!outputPath) return null;
let sourceCopyPath: string | undefined;
try {
await fs.mkdir(dirname(cachePath), { recursive: true });
sourceCopyPath = join(dirname(cachePath), `${basename(cachePath, '.pdf')}.docx`);
await fs.copyFile(filePath, sourceCopyPath);
const sourcePath = wslMountPathToWindowsPath(sourceCopyPath);
if (!sourcePath) return null;
await runWithConversionLimit(() =>
execFileAsync(
'powershell.exe',
[
'-NoProfile',
'-NonInteractive',
'-ExecutionPolicy',
'Bypass',
'-EncodedCommand',
encodePowerShellCommand(buildWordExportPdfScript(sourcePath, outputPath)),
],
{
timeout: OFFICE_CONVERSION_TIMEOUT_MS,
maxBuffer: 1024 * 1024,
}
)
);
if (await fileExists(cachePath)) {
await pruneDocumentPreviewCache(dirname(cachePath));
return cachePath;
}
console.warn(`[DocumentPreviewCache] Microsoft Word did not produce PDF output for ${filePath}`);
return null;
} catch (err) {
console.warn(
`[DocumentPreviewCache] Failed to convert DOCX with Microsoft Word (${filePath}):`,
getCacheErrorMessage(err)
);
return null;
} finally {
if (sourceCopyPath) {
await fs.rm(sourceCopyPath, { force: true }).catch(() => {});
}
}
}
async function convertLibreOfficeDocumentToCachedPdf(filePath: string, cachePath: string): Promise<string | null> {
let workDir: string | undefined;
try {
await fs.mkdir(DOCUMENT_PREVIEW_CACHE_DIR, { recursive: true });
const outDir = await fs.mkdtemp(join(DOCUMENT_PREVIEW_CACHE_DIR, 'work-'));
workDir = outDir;
const profileDir = join(outDir, 'profile');
await fs.mkdir(profileDir, { recursive: true });
await runWithConversionLimit(() =>
execFileAsync(
'soffice',
[
'--headless',
'--nologo',
'--nofirststartwizard',
`-env:UserInstallation=${pathToFileURL(profileDir).href}`,
'--convert-to',
'pdf',
'--outdir',
outDir,
filePath,
],
{
timeout: OFFICE_CONVERSION_TIMEOUT_MS,
maxBuffer: 1024 * 1024,
}
)
);
const converted = (await fs.readdir(outDir)).find((name) => name.toLowerCase().endsWith('.pdf'));
if (!converted) return null;
await fs.rename(join(workDir, converted), cachePath);
await pruneDocumentPreviewCache(DOCUMENT_PREVIEW_CACHE_DIR);
return cachePath;
} catch (err) {
console.warn(
`[DocumentPreviewCache] Failed to convert Office file to PDF (${filePath}):`,
getCacheErrorMessage(err)
);
return null;
} finally {
if (workDir) {
await fs.rm(workDir, { recursive: true, force: true }).catch(() => {});
}
}
}
function wslMountPathToWindowsPath(filePath: string): string | null {
const match = filePath.match(/^\/mnt\/([a-zA-Z])\/(.+)$/);
if (!match) return null;
return `${match[1].toUpperCase()}:\\${match[2].replace(/\//g, '\\')}`;
}
function getWindowsUserTempCacheDir(filePath: string): string | null {
const match = filePath.match(/^\/mnt\/([a-zA-Z])\/Users\/([^/]+)\//);
if (!match) return null;
return `/mnt/${match[1].toLowerCase()}/Users/${match[2]}/AppData/Local/Temp/codeman-document-preview-cache`;
}
function toPowerShellSingleQuotedString(value: string): string {
return `'${value.replace(/'/g, "''")}'`;
}
function encodePowerShellCommand(script: string): string {
return Buffer.from(script, 'utf16le').toString('base64');
}
export function getPreviewPdfDownloadName(fileName: string, extension: string): string {
return `${basename(fileName, extname(fileName) || `.${extension}`)}.pdf`;
}
function getCacheErrorMessage(err: unknown): string {
return err instanceof Error ? err.message : String(err);
}
+98
View File
@@ -0,0 +1,98 @@
/**
* @fileoverview Best-effort first-page thumbnails for attachment cards.
*/
import { execFile } from 'node:child_process';
import fs from 'node:fs/promises';
import { tmpdir } from 'node:os';
import { basename, extname, join } from 'node:path';
import { promisify } from 'node:util';
import { getOfficePreviewPdfPath } from './document-preview-cache.js';
import { runWithConversionLimit } from './document-conversion-limiter.js';
const execFileAsync = promisify(execFile);
const THUMBNAIL_CONVERSION_TIMEOUT_MS = 5 * 60_000;
/** Browser-renderable image formats served as-is (no conversion). */
const IMAGE_PASSTHROUGH_CONTENT_TYPES: Record<string, string> = {
png: 'image/png',
jpg: 'image/jpeg',
jpeg: 'image/jpeg',
gif: 'image/gif',
webp: 'image/webp',
};
export interface ThumbnailResult {
content: Buffer;
contentType: string;
}
export async function generateFirstPageThumbnail(filePath: string, extension: string): Promise<ThumbnailResult | null> {
const ext = extension.toLowerCase().replace(/^\./, '');
try {
await fs.stat(filePath);
const passthroughContentType = IMAGE_PASSTHROUGH_CONTENT_TYPES[ext];
if (passthroughContentType) {
return { content: await fs.readFile(filePath), contentType: passthroughContentType };
}
if (ext === 'pdf') {
return renderPdfFirstPage(filePath);
}
if (ext === 'docx' || ext === 'pptx') {
return renderOfficeFirstPage(filePath);
}
} catch (err) {
console.warn(`[Thumbnailer] Failed to generate ${ext} thumbnail for ${filePath}:`, getThumbnailErrorMessage(err));
return null;
}
return null;
}
async function renderOfficeFirstPage(filePath: string): Promise<ThumbnailResult | null> {
try {
const previewPdfPath = await getOfficePreviewPdfPath(filePath, extname(filePath).toLowerCase().replace(/^\./, ''));
if (!previewPdfPath) return null;
return await renderPdfFirstPage(previewPdfPath);
} catch (err) {
console.warn(
`[Thumbnailer] Failed to convert Office file to PDF for thumbnail (${filePath}):`,
getThumbnailErrorMessage(err)
);
return null;
}
}
async function renderPdfFirstPage(filePath: string): Promise<ThumbnailResult | null> {
let previewDir: string | undefined;
try {
previewDir = await fs.mkdtemp(join(tmpdir(), 'codeman-thumb-pdf-'));
const prefix = join(previewDir, basename(filePath, extname(filePath)));
await runWithConversionLimit(() =>
execFileAsync('pdftoppm', ['-png', '-singlefile', '-f', '1', '-l', '1', '-scale-to', '520', filePath, prefix], {
timeout: THUMBNAIL_CONVERSION_TIMEOUT_MS,
maxBuffer: 1024 * 1024,
})
);
const content = await fs.readFile(`${prefix}.png`);
return { content, contentType: 'image/png' };
} catch (err) {
console.warn(
`[Thumbnailer] Failed to render PDF first page for thumbnail (${filePath}):`,
getThumbnailErrorMessage(err)
);
return null;
} finally {
if (previewDir) {
await fs.rm(previewDir, { recursive: true, force: true }).catch(() => {});
}
}
}
function getThumbnailErrorMessage(err: unknown): string {
return err instanceof Error ? err.message : String(err);
}
+70
View File
@@ -0,0 +1,70 @@
/**
* @fileoverview Codex generated-artifact attachment registration.
*
* Codex image generation prints paths such as `Saved to: file://...`. These
* paths are registered directly when they fall within allowed locations (the
* session workspace or the well-known Codex generated-artifact directories
* anchored at the user's home). The trust decision is made on the
* realpath-RESOLVED path so a symlink staged at an allowed location cannot
* smuggle an arbitrary host file past workspace confinement.
*/
import { realpathSync } from 'node:fs';
import { homedir } from 'node:os';
import { join, normalize, sep } from 'node:path';
import { registerExternalAttachment, type AttachmentRegistrationResult } from './attachment-registry.js';
export interface GeneratedArtifactRegistrationOptions {
sessionId: string;
filePath: string;
sessionWorkingDir: string;
}
export async function registerGeneratedArtifactAttachment(
options: GeneratedArtifactRegistrationOptions
): Promise<AttachmentRegistrationResult> {
// Decide trust on the symlink-resolved path. If it can't be resolved, fall
// back to the strict force-confined policy (registration will 404 a missing
// file anyway).
let forceWorkspaceConfinement = true;
try {
const resolvedPath = realpathSync(options.filePath);
forceWorkspaceConfinement = !isAllowedGeneratedArtifactPath(resolvedPath, options.sessionWorkingDir);
} catch {
// Keep force confinement.
}
return registerExternalAttachment(options.sessionId, options.filePath, {
sessionWorkingDir: options.sessionWorkingDir,
forceWorkspaceConfinement,
});
}
/** Well-known Codex generated-artifact directories, anchored at the user's home. */
function codexGeneratedDirs(): string[] {
const home = homedir();
return [
join(home, '.codex-personal', 'generated_images'),
join(home, '.codex', 'generated_images'),
join(home, '.codex-personal', 'generated_artifacts'),
join(home, '.codex', 'generated_artifacts'),
];
}
/**
* True when `filePath` (absolute; callers should pass the realpath-resolved
* path) is inside the session workspace or one of the well-known Codex
* generated-artifact directories under the current user's home. The marker
* directories are prefix-anchored to `os.homedir()` — a `.codex/...` subtree
* elsewhere on the filesystem does NOT qualify.
*/
export function isAllowedGeneratedArtifactPath(filePath: string, workingDir: string): boolean {
const normalizedPath = normalize(filePath);
if (isPathInside(normalizedPath, workingDir)) return true;
return codexGeneratedDirs().some((dir) => isPathInside(normalizedPath, dir));
}
function isPathInside(filePath: string, rootPath: string): boolean {
const normalizedRoot = normalize(rootPath);
if (filePath === normalizedRoot) return true;
return filePath.startsWith(normalizedRoot.endsWith(sep) ? normalizedRoot : normalizedRoot + sep);
}
+212 -73
View File
@@ -3,8 +3,9 @@
*
* Generates `.claude/settings.local.json` with hook definitions that POST
* to Codeman's `/api/hook-event` endpoint when Claude Code fires hooks.
* Uses `$CODEMAN_API_URL` and `$CODEMAN_SESSION_ID` env vars (set on every
* managed session) so the config is static per case directory.
* Uses `$CODEMAN_API_URL`, `$CODEMAN_SESSION_ID`, and `$CODEMAN_HOOK_SECRET_FILE`
* env vars (set on every managed session) so the config is static per case
* directory and free of secret values.
*
* Key exports:
* - `generateHooksConfig()` — returns hooks object for settings.local.json
@@ -30,6 +31,30 @@ import { join } from 'node:path';
import type { HookEventType } from './types.js';
import { HOOK_TIMEOUT_MS } from './config/auth-config.js';
/**
* Serializes read-modify-write access to a `settings.local.json` path. Every
* writer in this module (hooks, env, model, statusLine) shares this map, so
* concurrent updates to the SAME file — e.g. session-create writing hooks/model
* while an App-Settings toggle injects the statusLine into the same repo — can't
* lose each other's changes through interleaved read-then-write. Per-path chains
* are independent; the map self-prunes when a path's chain goes idle.
*/
const settingsWriteLocks = new Map<string, Promise<unknown>>();
function withSettingsLock<T>(path: string, fn: () => Promise<T>): Promise<T> {
const prev = settingsWriteLocks.get(path) ?? Promise.resolve();
const run = prev.then(fn, fn); // run after the prior writer, regardless of its outcome
// Tail never rejects, so a failed write doesn't poison subsequent writers.
const tail = run.then(
() => {},
() => {}
);
settingsWriteLocks.set(path, tail);
void tail.then(() => {
if (settingsWriteLocks.get(path) === tail) settingsWriteLocks.delete(path);
});
return run;
}
/**
* Generates the hooks section for .claude/settings.local.json
*
@@ -41,11 +66,18 @@ import { HOOK_TIMEOUT_MS } from './config/auth-config.js';
export function generateHooksConfig(): { hooks: Record<string, unknown[]> } {
// Read Claude Code's stdin JSON and forward it as the data field.
// Falls back to empty object if stdin is unavailable or malformed.
// COD-54: present the per-instance hook secret so the bypass keeps working while
// a tunnel is running. The value is read from the secret file AT EXECUTION TIME
// (path via $CODEMAN_HOOK_SECRET_FILE, set in every managed session's env), so it
// never lands in this config and rotation needs no respawn. If the var/file is
// missing the header is empty — the middleware then allows the request only on
// the plain loopback bypass (tunnel down), same as pre-secret behavior.
const curlCmd = (event: HookEventType) =>
`HOOK_DATA=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"event":"${event}","sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$HOOK_DATA" | ` +
`curl -s -X POST "$CODEMAN_API_URL/api/hook-event" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- ` +
`2>/dev/null || true`;
@@ -95,29 +127,31 @@ export async function stripCaseEnvKeys(casePath: string, keysToRemove: readonly
if (keysToRemove.length === 0) return;
const settingsPath = join(casePath, '.claude', 'settings.local.json');
if (!existsSync(settingsPath)) return;
await withSettingsLock(settingsPath, async () => {
if (!existsSync(settingsPath)) return;
let existing: Record<string, unknown>;
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
return; // Malformed — don't rewrite it
}
const env = existing.env as Record<string, string> | undefined;
if (!env) return;
let changed = false;
for (const key of keysToRemove) {
if (key in env) {
delete env[key];
changed = true;
let existing: Record<string, unknown>;
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
return; // Malformed — don't rewrite it
}
}
if (!changed) return;
existing.env = env;
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
const env = existing.env as Record<string, string> | undefined;
if (!env) return;
let changed = false;
for (const key of keysToRemove) {
if (key in env) {
delete env[key];
changed = true;
}
}
if (!changed) return;
existing.env = env;
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
/**
@@ -126,30 +160,31 @@ export async function stripCaseEnvKeys(casePath: string, keysToRemove: readonly
*/
export async function updateCaseEnvVars(casePath: string, envVars: Record<string, string>): Promise<void> {
const claudeDir = join(casePath, '.claude');
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
const settingsPath = join(claudeDir, 'settings.local.json');
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
existing = {};
}
const currentEnv = (existing.env as Record<string, string>) || {};
for (const [key, value] of Object.entries(envVars)) {
if (value) {
currentEnv[key] = value;
} else {
delete currentEnv[key];
await withSettingsLock(settingsPath, async () => {
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
}
existing.env = currentEnv;
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
existing = {};
}
const currentEnv = (existing.env as Record<string, string>) || {};
for (const [key, value] of Object.entries(envVars)) {
if (value) {
currentEnv[key] = value;
} else {
delete currentEnv[key];
}
}
existing.env = currentEnv;
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
/**
@@ -158,26 +193,27 @@ export async function updateCaseEnvVars(casePath: string, envVars: Record<string
*/
export async function updateCaseModel(casePath: string, model: string | null): Promise<void> {
const claudeDir = join(casePath, '.claude');
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
const settingsPath = join(claudeDir, 'settings.local.json');
let existing: Record<string, unknown> = {};
await withSettingsLock(settingsPath, async () => {
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
existing = {};
}
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
existing = {};
}
if (model) {
existing.model = model;
} else {
delete existing.model;
}
if (model) {
existing.model = model;
} else {
delete existing.model;
}
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
/**
@@ -186,22 +222,125 @@ export async function updateCaseModel(casePath: string, model: string | null): P
*/
export async function writeHooksConfig(casePath: string): Promise<void> {
const claudeDir = join(casePath, '.claude');
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
const settingsPath = join(claudeDir, 'settings.local.json');
let existing: Record<string, unknown> = {};
await withSettingsLock(settingsPath, async () => {
if (!existsSync(claudeDir)) {
await mkdir(claudeDir, { recursive: true });
}
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
// If file is malformed or doesn't exist, start fresh
existing = {};
}
let existing: Record<string, unknown> = {};
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
// If file is malformed or doesn't exist, start fresh
existing = {};
}
const hooksConfig = generateHooksConfig();
const merged = { ...existing, ...hooksConfig };
const hooksConfig = generateHooksConfig();
const merged = { ...existing, ...hooksConfig };
await writeFile(settingsPath, JSON.stringify(merged, null, 2) + '\n');
await writeFile(settingsPath, JSON.stringify(merged, null, 2) + '\n');
});
}
/**
* Self-heal a case's hooks block so the COD-91 unconditional hook-secret gate keeps
* accepting its hook events.
*
* `writeHooksConfig` only runs when a case is first CREATED. Cases created before the
* X-Codeman-Hook-Secret header was added (COD-54, 2026-06-10) keep hook curls in their
* settings.local.json that POST to /api/hook-event WITHOUT the secret — which, once the
* gate requires it unconditionally (COD-91), silently 401 on a password-protected
* install. This refreshes the hooks block so those stale curls regain the header.
*
* Deliberately surgical: regenerates ONLY when settings.local.json already contains
* Codeman's own hook curls (they target `/api/hook-event`) that lack the secret header.
* No-op when the file/hooks are absent (we never impose hooks on a user who removed
* them), when the hooks aren't ours, or when the secret is already present — so it never
* clobbers a user's customizations and is cheap enough to call on every Claude spawn.
*/
export async function refreshStaleHookSecret(casePath: string): Promise<void> {
const settingsPath = join(casePath, '.claude', 'settings.local.json');
if (!existsSync(settingsPath)) return;
await withSettingsLock(settingsPath, async () => {
let existing: Record<string, unknown>;
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
return; // malformed — leave it untouched (case-create owns the happy path)
}
const hooksJson = JSON.stringify(existing.hooks ?? null);
const isOurs = hooksJson.includes('/api/hook-event');
// The generated curl carries this header literal (see generateHooksConfig); its
// absence on our own hooks means they predate COD-54 and need regenerating.
const hasSecret = hooksJson.includes('X-Codeman-Hook-Secret');
if (!isOurs || hasSecret) return;
const merged = { ...existing, ...generateHooksConfig() };
await writeFile(settingsPath, JSON.stringify(merged, null, 2) + '\n');
});
}
/** Unique marker identifying Codeman's own statusLine command (vs a user's). */
const STATUSLINE_MARKER = '/api/status-telemetry';
/**
* The plan-usage statusLine exporter command. Mirrors the hook `curlCmd` pattern:
* reads Claude Code's statusline stdin JSON, POSTs `{sessionId,data}` to Codeman,
* and prints the response body (a compact "⟳ 5h 15% · 7d 34%" footer) back to
* stdout so the in-terminal statusline stays useful. Env vars resolve at runtime
* (present in every managed session via tmux setenv), so the config is static.
*/
export function generateStatusLineCommand(): string {
// `curl -sk`: CODEMAN_API_URL is loopback HTTPS with a self-signed cert in the
// production setup; without -k curl returns 000 and the statusline shows
// nothing. -k is safe here (loopback only). Falls back to a brand string so the
// footer is never blank if Codeman is unreachable.
return (
`INPUT=$(cat 2>/dev/null || echo '{}'); ` +
`printf '{"sessionId":"%s","data":%s}' "$CODEMAN_SESSION_ID" "$INPUT" | ` +
`curl -sk -X POST "$CODEMAN_API_URL${STATUSLINE_MARKER}" ` +
`-H 'Content-Type: application/json' ` +
`-H "X-Codeman-Hook-Secret: $(cat "$CODEMAN_HOOK_SECRET_FILE" 2>/dev/null)" ` +
`--data @- 2>/dev/null || echo codeman`
);
}
/**
* Add or remove Codeman's plan-usage statusLine exporter in
* `.claude/settings.local.json`. Only ever touches a statusLine that is OURS
* (command targets `/api/status-telemetry`), so a user's hand-authored
* statusLine is never removed OR overwritten — on both the enable and disable
* paths we bail out when an existing statusLine isn't ours. Callers gate on
* Claude mode. Merges, preserving all other keys (hooks, env, model).
*/
export async function applyStatusLineConfig(casePath: string, enabled: boolean): Promise<void> {
const claudeDir = join(casePath, '.claude');
const settingsPath = join(claudeDir, 'settings.local.json');
await withSettingsLock(settingsPath, async () => {
let existing: Record<string, unknown> = {};
if (existsSync(settingsPath)) {
try {
existing = JSON.parse(await readFile(settingsPath, 'utf-8'));
} catch {
return; // Malformed — don't rewrite it
}
}
const current = existing.statusLine as { command?: unknown } | undefined;
const isOurs = !!current && typeof current.command === 'string' && current.command.includes(STATUSLINE_MARKER);
if (enabled) {
const desired = generateStatusLineCommand();
if (isOurs && current?.command === desired) return; // already current — skip rewrite
if (current && !isOurs) return; // user has their OWN statusLine — never clobber it
if (!existsSync(claudeDir)) await mkdir(claudeDir, { recursive: true });
existing.statusLine = { type: 'command', command: desired }; // add, or update an out-of-date ours
} else {
if (!isOurs) return; // nothing of ours to remove (leave a user's own statusLine alone)
delete existing.statusLine;
}
await writeFile(settingsPath, JSON.stringify(existing, null, 2) + '\n');
});
}
+51 -8
View File
@@ -12,7 +12,7 @@ import { EventEmitter } from 'node:events';
import { watch, type FSWatcher } from 'chokidar';
import { basename, extname, relative } from 'node:path';
import { statSync } from 'node:fs';
import type { ImageDetectedEvent } from './types.js';
import type { AttachmentDetectedEvent, AttachmentDetectedType, ImageDetectedEvent } from './types.js';
import { KeyedDebouncer } from './utils/index.js';
// ========== Types ==========
@@ -20,7 +20,9 @@ import { KeyedDebouncer } from './utils/index.js';
// ========== Constants ==========
/** Supported image file extensions (lowercase) */
const IMAGE_EXTENSIONS = new Set(['.png', '.jpg', '.jpeg', '.gif', '.webp', '.bmp', '.svg']);
const IMAGE_POPUP_EXTENSIONS = new Set(['.jpg', '.jpeg', '.gif', '.webp', '.bmp', '.svg']);
const ATTACHMENT_EXTENSIONS = new Set(['.png', '.pdf', '.docx', '.pptx']);
const DETECTED_FILE_EXTENSIONS = new Set([...IMAGE_POPUP_EXTENSIONS, ...ATTACHMENT_EXTENSIONS]);
/** Time to wait for file writes to stabilize (ms) */
const STABILITY_THRESHOLD_MS = 500;
@@ -166,8 +168,8 @@ export class ImageWatcher extends EventEmitter {
}
const ext = extname(path).toLowerCase();
// Don't ignore directories (needed for watching to work)
// Ignore files that aren't images
return ext !== '' && !IMAGE_EXTENSIONS.has(ext);
// Ignore files that aren't previewable images/documents
return ext !== '' && !DETECTED_FILE_EXTENSIONS.has(ext);
},
});
@@ -229,15 +231,16 @@ export class ImageWatcher extends EventEmitter {
/**
* Handle a new file being detected.
* Verifies it's an image and emits the detection event.
* Verifies it's a previewable image/document and emits the detection event.
*/
private handleNewFile(sessionId: string, filePath: string): void {
const ext = extname(filePath).toLowerCase();
// Double-check it's an image extension
if (!IMAGE_EXTENSIONS.has(ext)) {
// Double-check it's a supported extension
if (!DETECTED_FILE_EXTENSIONS.has(ext)) {
return;
}
const isAttachment = ATTACHMENT_EXTENSIONS.has(ext);
// Burst limit: skip if too many images detected for this session in a short window
const now = Date.now();
@@ -259,7 +262,11 @@ export class ImageWatcher extends EventEmitter {
// Debounce rapid file creation (e.g., multiple screenshots quickly)
this.fileDeb.schedule(filePath, () => {
this.fileToSession.delete(filePath);
this.emitImageDetected(sessionId, filePath);
if (isAttachment) {
this.emitAttachmentDetected(sessionId, filePath);
} else {
this.emitImageDetected(sessionId, filePath);
}
// Increment burst count on actual emission (not on detection)
const b = this.burstTrackers.get(sessionId);
if (b) b.count++;
@@ -294,6 +301,42 @@ export class ImageWatcher extends EventEmitter {
this.emit('image:error', error instanceof Error ? error : new Error(String(error)), sessionId);
}
}
/**
* Emit the attachment:detected event with file metadata.
*/
private emitAttachmentDetected(sessionId: string, filePath: string): void {
try {
const stat = statSync(filePath);
const fileName = basename(filePath);
const workingDir = this.sessionDirs.get(sessionId);
const relativePath = workingDir ? relative(workingDir, filePath) : fileName;
const extension = extname(fileName).toLowerCase().replace(/^\./, '');
const event: AttachmentDetectedEvent = {
sessionId,
filePath,
relativePath,
fileName,
extension,
attachmentType: this.getAttachmentType(extension),
timestamp: Date.now(),
size: stat.size,
};
this.emit('attachment:detected', event);
} catch (error) {
this.emit('image:error', error instanceof Error ? error : new Error(String(error)), sessionId);
}
}
private getAttachmentType(extension: string): AttachmentDetectedType {
if (extension === 'png') return 'image';
if (extension === 'pdf') return 'pdf';
if (extension === 'docx') return 'document';
if (extension === 'pptx') return 'presentation';
return 'document';
}
}
// Export singleton instance for convenience
+12
View File
@@ -14,6 +14,18 @@ import { program } from './cli.js';
// In web mode, we should NOT exit on transient errors — log and continue
const isWebMode = process.argv.includes('web');
// COD-115: Codeman IS a tmux controller; it must never present as a tmux *client*.
// If the web server is launched from inside a tmux pane it inherits TMUX/TMUX_PANE,
// and tmux's nesting guard then kills every new attach-bridge PTY (exit 1 → respawn
// loop, crash-looping any new tmux-backed session). Scrub at the root so every
// downstream `{...process.env}` spread (attach, send-keys, create) is clean regardless
// of launch context. `delete` (not `= undefined`, which node-pty serializes as the
// literal string "undefined" and fails to clear).
if (isWebMode) {
delete process.env.TMUX;
delete process.env.TMUX_PANE;
}
import { MAX_CONSECUTIVE_ERRORS, ERROR_RESET_MS } from './config/server-timing.js';
// Track consecutive unhandled errors in web mode — restart after too many
+49 -4
View File
@@ -16,6 +16,9 @@ import type {
OpenCodeConfig,
CodexConfig,
EffortLevel,
GeminiConfig,
SessionRemote,
SessionDocker,
} from './types.js';
/**
@@ -32,6 +35,10 @@ export interface MuxSession {
createdAt: number;
/** Working directory */
workingDir: string;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
/** Docker execution metadata for local tmux sessions wrapping `docker exec` */
docker?: SessionDocker;
/** Session mode */
mode: SessionMode;
/** Whether webserver is attached to this session */
@@ -64,12 +71,19 @@ export interface CreateSessionOptions {
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
/** When restoring after reboot, resume a previous Claude conversation by its session ID */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Ephemeral — not written to disk. */
envOverrides?: Record<string, string>;
/** Claude CLI effort level, injected as a `--settings` soft default (overridable via /effort in-session) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
/** Docker execution metadata for local tmux sessions wrapping `docker exec` */
docker?: SessionDocker;
}
/** Options for respawning a dead pane. */
@@ -83,12 +97,33 @@ export interface RespawnPaneOptions {
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
/** Resume a previous Claude conversation when respawning */
resumeSessionId?: string;
/** Extra env vars exported before launching the CLI (preserved across respawns). */
envOverrides?: Record<string, string>;
/** Claude CLI effort level (preserved across respawns, injected via `--settings`) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) to set for this session after respawn. */
historyLimit?: number;
/** Remote execution metadata for local tmux sessions wrapping SSH */
remote?: SessionRemote;
/** Docker execution metadata for local tmux sessions wrapping `docker exec` */
docker?: SessionDocker;
}
/** Options for pane buffer capture (COD-47 full-history mode). */
export interface PaneCaptureOptions {
/** Capture the entire tmux scrollback instead of just the visible frame. */
fullHistory?: boolean;
/** Bound the full-history capture to this many scrollback lines (`-S -<N>`). */
historyLimitLines?: number;
/**
* Byte cap the consumer will keep from the capture. Sizes the child-process
* stdout buffer (with slack) so multi-MB scrollback dumps aren't killed by
* the 1MB execSync default (ENOBUFS).
*/
maxCaptureBytes?: number;
}
/**
@@ -167,6 +202,9 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Update Ralph enabled state for a session */
updateRalphEnabled(sessionId: string, enabled: boolean): void;
/** Apply a tmux history-limit to all tracked sessions. */
setHistoryLimit(limit: number): Promise<void>;
// ========== Discovery ==========
/**
@@ -215,9 +253,16 @@ export interface TerminalMultiplexer extends EventEmitter {
/** Respawn a dead pane with a fresh command. Returns the new PID or null on failure. */
respawnPane(options: RespawnPaneOptions): Promise<number | null>;
/** Capture a pane's current tmux buffer with ANSI escape codes preserved. */
capturePaneBuffer?(muxName: string, paneTarget: string): string | null;
/**
* Capture a pane's current tmux buffer with ANSI escape codes preserved.
* Pass `{ fullHistory: true }` to capture the entire scrollback as linear
* text instead of just the visible single-screen frame (COD-47).
*/
capturePaneBuffer?(muxName: string, paneTarget?: string, opts?: PaneCaptureOptions): string | null;
/** Capture the active pane's current tmux buffer with ANSI escape codes preserved. */
captureActivePaneBuffer?(muxName: string): string | null;
/**
* Capture the active pane's current tmux buffer with ANSI escape codes preserved.
* Pass `{ fullHistory: true }` to capture the entire scrollback (COD-47).
*/
captureActivePaneBuffer?(muxName: string, opts?: PaneCaptureOptions): string | null;
}
+1
View File
@@ -8,3 +8,4 @@
export { RESEARCH_AGENT_PROMPT } from './research-agent.js';
export { PLANNER_PROMPT } from './planner.js';
export { PHASE_EXECUTION_PROMPT, TEAM_LEAD_PROMPT, REPLAN_PROMPT, SINGLE_TASK_PROMPT } from './orchestrator.js';
export { RALPH_STATUS_CONTRACT, buildRalphLoopPrompt, type RalphLoopPromptOptions } from './ralph.js';
+85
View File
@@ -0,0 +1,85 @@
/**
* @fileoverview Ralph Loop prompt construction
*
* Builds the full `@ralph_prompt.md` content written for a new Ralph loop
* session, including the RALPH_STATUS block contract. The contract travels
* with the loop prompt (not the generated CLAUDE.md) so every Ralph session
* emits parseable status blocks regardless of the project's CLAUDE.md.
*
* @module prompts/ralph
*/
/**
* Structured status-reporting contract appended to every Ralph loop prompt.
*
* `RalphStatusParser` (src/ralph-status-parser.ts) parses this block from
* session output — keep the field names and enum values in sync with its
* patterns.
*/
export const RALPH_STATUS_CONTRACT = `## Status Reporting
End EVERY response with exactly this block — Codeman parses it to track the loop:
\`\`\`
---RALPH_STATUS---
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
TASKS_COMPLETED_THIS_LOOP: <number>
FILES_MODIFIED: <number>
TESTS_STATUS: PASSING | FAILING | NOT_RUN
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
EXIT_SIGNAL: false | true
RECOMMENDATION: <one line: what to do next>
---END_RALPH_STATUS---
\`\`\`
Rules:
- \`EXIT_SIGNAL: true\` only when ALL tasks are verifiably done — then also output the completion phrase
- \`STATUS: BLOCKED\` when you need human input; describe the blocker in RECOMMENDATION
- Never set \`EXIT_SIGNAL: true\` while tests are failing
`;
export interface RalphLoopPromptOptions {
/** The user's task description (becomes the prompt header) */
taskDescription: string;
/** Completion phrase the session must emit inside <promise></promise> */
completionPhrase: string;
/** Whether a @fix_plan.md task plan was generated for this loop */
hasPlan: boolean;
}
/**
* Builds the full Ralph loop prompt written to `@ralph_prompt.md`.
*/
export function buildRalphLoopPrompt({ taskDescription, completionPhrase, hasPlan }: RalphLoopPromptOptions): string {
let fullPrompt = taskDescription + '\n\n---\n\n';
if (hasPlan) {
fullPrompt += '## Task Plan\n\n';
fullPrompt += 'A task plan has been written to `@fix_plan.md`. Use this to track progress:\n';
fullPrompt += '- Reference the plan at the start of each iteration\n';
fullPrompt += '- Update task checkboxes as you complete items\n';
fullPrompt += '- Work through items in priority order (P0 > P1 > P2)\n\n';
}
fullPrompt += '## Iteration Protocol\n\n';
fullPrompt += 'This is an autonomous loop. Files from previous iterations persist. On each iteration:\n';
fullPrompt += '1. Check what work has already been done\n';
fullPrompt += '2. Make incremental progress toward completion\n';
fullPrompt += '3. Commit meaningful changes with descriptive messages\n\n';
fullPrompt += '## Verification\n\n';
fullPrompt += 'After each significant change:\n';
fullPrompt += '- Run tests to verify (npm test, pytest, etc.)\n';
fullPrompt += '- Check for type/lint errors if applicable\n';
fullPrompt += '- If tests fail, read the error, fix it, and retry\n\n';
fullPrompt += '## Completion Criteria\n\n';
fullPrompt += `Output \`<promise>${completionPhrase}</promise>\` when ALL of the following are true:\n`;
fullPrompt += '- All requirements from the task description are implemented\n';
fullPrompt += '- All tests pass\n';
fullPrompt += '- Changes are committed\n\n';
fullPrompt += '## If Stuck\n\n';
fullPrompt += 'If you encounter the same error for 3+ iterations:\n';
fullPrompt += "1. Document what you've tried\n";
fullPrompt += '2. Identify the specific blocker\n';
fullPrompt += '3. Try an alternative approach\n';
fullPrompt += '4. If truly blocked, output `<promise>BLOCKED</promise>` with an explanation\n\n';
fullPrompt += RALPH_STATUS_CONTRACT;
return fullPrompt;
}
+47 -3
View File
@@ -467,6 +467,12 @@ export class RalphTracker extends EventEmitter {
/** Timestamp of last cleanup check for throttling */
private _lastCleanupTime: number = 0;
/** Maximum number of todos retained for this session (defaults to global cap) */
private _maxTodos: number = MAX_TODOS_PER_SESSION;
/** Todo auto-expiry duration in milliseconds (defaults to global constant) */
private _todoExpiryMs: number = TODO_EXPIRY_MS;
/** Debouncer for todoUpdate events */
private _todoDeb = new Debouncer(EVENT_DEBOUNCE_MS);
@@ -1053,6 +1059,10 @@ export class RalphTracker extends EventEmitter {
planVersion: this.planTracker.planVersion,
planHistoryLength: this.planTracker.getPlanHistory().length,
completionConfidence: this._lastCompletionConfidence,
// Surface the live todo-config so it persists (toState) and reads back into
// the Session Options modal (broadcast) — mirrors maxIterations round-trip.
maxTodos: this._maxTodos,
todoExpirationMinutes: this.todoExpirationMinutes,
};
}
@@ -1840,7 +1850,7 @@ export class RalphTracker extends EventEmitter {
return;
}
while (this._todos.size >= MAX_TODOS_PER_SESSION) {
while (this._todos.size >= this._maxTodos) {
const oldest = this.findOldestTodo();
if (oldest) {
this._todos.delete(oldest.id);
@@ -2164,14 +2174,14 @@ export class RalphTracker extends EventEmitter {
}
/**
* Remove todo items older than TODO_EXPIRY_MS.
* Remove todo items older than the configured expiry duration.
*/
private cleanupExpiredTodos(): void {
const now = Date.now();
const toDelete: string[] = [];
for (const [id, todo] of this._todos) {
if (now - todo.detectedAt > TODO_EXPIRY_MS) {
if (now - todo.detectedAt > this._todoExpiryMs) {
toDelete.push(id);
}
}
@@ -2211,6 +2221,34 @@ export class RalphTracker extends EventEmitter {
this.emit('loopUpdate', this.loopState);
}
/** Maximum number of todos retained for this session. */
get maxTodos(): number {
return this._maxTodos;
}
/** Todo auto-expiry duration in minutes for this session. */
get todoExpirationMinutes(): number {
return Math.round(this._todoExpiryMs / 60000);
}
/**
* Update the maximum number of retained todos (external API).
* Ignores non-positive values.
*/
setMaxTodos(maxTodos: number): void {
if (!Number.isFinite(maxTodos) || maxTodos <= 0) return;
this._maxTodos = Math.floor(maxTodos);
}
/**
* Update the todo auto-expiry duration (external API), specified in minutes.
* Converts to milliseconds internally. Ignores non-positive values.
*/
setTodoExpirationMinutes(minutes: number): void {
if (!Number.isFinite(minutes) || minutes <= 0) return;
this._todoExpiryMs = Math.floor(minutes) * 60000;
}
/**
* Configure the tracker from external state.
*/
@@ -2311,6 +2349,12 @@ export class RalphTracker extends EventEmitter {
...loopState,
enabled: loopState.enabled ?? false,
};
// Restore the per-session todo-config into the live fields used by the hot
// paths (eviction cap + expiry). Setters ignore non-positive values.
if (typeof loopState.maxTodos === 'number') this.setMaxTodos(loopState.maxTodos);
if (typeof loopState.todoExpirationMinutes === 'number') {
this.setTodoExpirationMinutes(loopState.todoExpirationMinutes);
}
this._todos.clear();
for (const todo of todos) {
this._todos.set(todo.id, {
+228
View File
@@ -0,0 +1,228 @@
import { existsSync, mkdirSync } from 'node:fs';
import fs from 'node:fs/promises';
import { join } from 'node:path';
import { homedir } from 'node:os';
import { exec } from 'node:child_process';
import { promisify } from 'node:util';
import type {
RemoteCase,
RemoteCommandMode,
RemoteHost,
RemoteSshOptions,
SessionMode,
SessionRemote,
} from './types.js';
const execAsync = promisify(exec);
const REMOTE_HOSTS_FILE = 'remote-hosts.json';
const REMOTE_CASES_FILE = 'remote-cases.json';
export function remoteHostsPath(configDir: string): string {
return join(configDir, REMOTE_HOSTS_FILE);
}
export function remoteCasesPath(configDir: string): string {
return join(configDir, REMOTE_CASES_FILE);
}
async function readJsonArray<T>(path: string): Promise<T[]> {
try {
const raw = await fs.readFile(path, 'utf-8');
const parsed = JSON.parse(raw);
return Array.isArray(parsed) ? (parsed as T[]) : [];
} catch {
return [];
}
}
async function writeJsonArray<T>(configDir: string, path: string, value: T[]): Promise<void> {
if (!existsSync(configDir)) mkdirSync(configDir, { recursive: true });
await fs.writeFile(path, JSON.stringify(value, null, 2));
}
export async function readRemoteHosts(configDir: string): Promise<RemoteHost[]> {
return readJsonArray<RemoteHost>(remoteHostsPath(configDir));
}
export async function writeRemoteHosts(configDir: string, hosts: RemoteHost[]): Promise<void> {
await writeJsonArray(configDir, remoteHostsPath(configDir), hosts);
}
export async function readRemoteCases(configDir: string): Promise<RemoteCase[]> {
return readJsonArray<RemoteCase>(remoteCasesPath(configDir));
}
export async function writeRemoteCases(configDir: string, cases: RemoteCase[]): Promise<void> {
await writeJsonArray(configDir, remoteCasesPath(configDir), cases);
}
export function defaultRemoteCommandForMode(mode: SessionMode): string {
const commands: Record<RemoteCommandMode, string> = {
shell: 'exec bash -l',
// Mirror the LOCAL claude default so the remote agent runs non-interactively
// (no trust-folder/permission prompt that nothing on the remote answers). The
// per-host `commands.claude` override stays the escape hatch.
claude: 'exec claude --dangerously-skip-permissions',
opencode: 'exec opencode',
codex: 'exec codex',
gemini: 'exec gemini',
};
return commands[mode as RemoteCommandMode] || commands.shell;
}
export function remoteSshTarget(host: Pick<RemoteHost, 'username' | 'host'>): string {
return `${host.username}@${host.host}`;
}
/**
* POSIX single-quote shell-escaping (end-quote, escaped-quote, restart-quote).
* Mirrors the helper in tmux-manager.ts so a value with spaces/metachars stays a
* single shell token. Used here for identity paths and `-o KEY=VALUE` options.
*/
function shellescape(str: string): string {
return "'" + str.replace(/'/g, "'\\''") + "'";
}
/**
* Expand a leading `~` or `$HOME` in an identity path to an absolute path.
*
* ssh does NOT expand `~` inside `-i` (the shell would, but we shellescape the
* value into a single quoted token so the shell never sees it). So we expand at
* build time, before escaping. Non-`~`/`$HOME` paths are returned unchanged.
*/
function expandIdentityPath(identityFile: string): string {
if (identityFile === '~') return homedir();
if (identityFile.startsWith('~/')) return join(homedir(), identityFile.slice(2));
if (identityFile === '$HOME') return homedir();
if (identityFile.startsWith('$HOME/')) return join(homedir(), identityFile.slice('$HOME/'.length));
return identityFile;
}
/**
* COD-107 — build the ordered, shell-safe ssh CONNECTION tokens shared by both
* the durable-launch command (`buildRemoteLaunchCommand`) and the tmux
* prerequisite probe (`buildRemoteTmuxCheckCommand`), so the prereq check and
* the real launch connect with IDENTICAL options (they can't drift).
*
* Returns the leading tokens of an ssh command line (NOT including `-t`, the
* target, or any remote command). Order:
* ssh -o BatchMode=yes
* [-o ConnectTimeout=10] (default; suppressed if extraSshOptions sets it)
* [-p <port>]
* [-i <abs-identity>] (~/$HOME expanded, then shellescaped)
* [-J <jumpHost>] (shellescaped, single token)
* [-o ProxyCommand=nc -X 5 -x <socks> %h %p] (ONE shellescaped -o token)
* [-o <KEY=VALUE>] … (each extra option, shellescaped)
*
* Escaping notes (the risky part):
* - The ProxyCommand is emitted as a single shellescaped `-o KEY=VALUE`, so the
* whole value (spaces + `%h`/`%p`) reaches ssh as one argument and `%h %p`
* survive verbatim — ssh expands them to the real host/port, not the shell.
* - A default `-o ConnectTimeout=10` bounds the wait on an unreachable/blackholed
* host (else the pane hangs on the OS TCP timeout). It is omitted when the
* operator already set ConnectTimeout via extraSshOptions, so their value wins.
*/
export function buildSshConnectionArgs(remote: RemoteSshOptions & Pick<RemoteHost, 'port'>): string[] {
const parts: string[] = ['ssh', '-o BatchMode=yes'];
const hasConnectTimeout = (remote.extraSshOptions ?? []).some((opt) => /^ConnectTimeout=/i.test(opt));
if (!hasConnectTimeout) parts.push('-o ConnectTimeout=10');
if (remote.port) parts.push(`-p ${remote.port}`);
if (remote.identityFile) parts.push(`-i ${shellescape(expandIdentityPath(remote.identityFile))}`);
if (remote.jumpHost) parts.push(`-J ${shellescape(remote.jumpHost)}`);
if (remote.socksProxy) {
parts.push(`-o ${shellescape(`ProxyCommand=nc -X 5 -x ${remote.socksProxy} %h %p`)}`);
}
for (const opt of remote.extraSshOptions ?? []) {
parts.push(`-o ${shellescape(opt)}`);
}
return parts;
}
/**
* COD-104 — build the SSH command that checks the remote host has tmux.
*
* Durable remote sessions run the agent inside a tmux server ON the remote host
* (`tmux -L codeman new-session -A …`), so tmux is now a hard prerequisite there.
* `command -v tmux` exits 0 (and prints the path) when tmux is installed.
*
* COD-107 — connects with the SAME options as the real launch
* (`buildSshConnectionArgs`) so a proxied/custom-port/identity host that the
* launch can reach also passes the prereq probe (and vice-versa).
*/
export function buildRemoteTmuxCheckCommand(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions
): string {
// ConnectTimeout is now a default of buildSshConnectionArgs (shared with the launch).
return [...buildSshConnectionArgs(host), remoteSshTarget(host), "'command -v tmux'"].join(' ');
}
export interface RemoteTmuxCheckResult {
ok: boolean;
/** Resolved tmux path on the remote (when ok). */
tmuxPath?: string;
/** Human-readable failure reason (when !ok). */
error?: string;
}
/**
* COD-104 — verify the remote host has tmux installed (required for durable
* remote sessions). Returns a structured result with a clear, user-facing error
* when tmux is missing or the host is unreachable. Never throws.
*/
export async function checkRemoteTmuxAvailable(
host: Pick<RemoteHost, 'username' | 'host' | 'port'> & RemoteSshOptions
): Promise<RemoteTmuxCheckResult> {
const command = buildRemoteTmuxCheckCommand(host);
try {
const { stdout } = await execAsync(command, { timeout: 15_000 });
const tmuxPath = stdout.trim();
if (!tmuxPath) {
return {
ok: false,
error: `remote host ${host.host} needs tmux installed for durable remote sessions`,
};
}
return { ok: true, tmuxPath };
} catch (err) {
const stderr =
err && typeof err === 'object' && 'stderr' in err ? String((err as { stderr?: unknown }).stderr ?? '') : '';
// `command -v tmux` exits non-zero when tmux is absent (no stderr); a real
// connection failure surfaces ssh diagnostics on stderr.
if (stderr.trim()) {
return {
ok: false,
error: `could not verify tmux on remote host ${host.host}: ${stderr.trim()}`,
};
}
return {
ok: false,
error: `remote host ${host.host} needs tmux installed for durable remote sessions`,
};
}
}
export function remoteDisplayPath(
remote: Pick<SessionRemote, 'username' | 'host' | 'remotePath'> | { username: string; host: string; path: string }
): string {
const path = 'remotePath' in remote ? remote.remotePath : remote.path;
return `${remote.username}@${remote.host}:${path}`;
}
export function toSessionRemote(host: RemoteHost, remoteCase: RemoteCase): SessionRemote {
return {
hostId: host.id,
label: host.label,
host: host.host,
username: host.username,
port: host.port,
remotePath: remoteCase.remotePath,
commands: host.commands,
// COD-107 — carry the advanced SSH options from host config into the session
// so the launch/prereq commands connect the same way the operator configured.
identityFile: host.identityFile,
socksProxy: host.socksProxy,
jumpHost: host.jumpHost,
extraSshOptions: host.extraSshOptions,
};
}
+13
View File
@@ -2779,6 +2779,19 @@ export class RespawnController extends EventEmitter {
return;
}
// Usage-limit pause: Claude can't work and the cycle's /clear would wipe
// the paused conversation — the auto-resume scheduler owns recovery here.
if (this.session.isLimitPaused) {
this.log('Skipping respawn cycle - usage-limit pause active (auto-resume armed)');
this.logAction('health', 'Respawn skipped: usage-limit pause (auto-resume armed)');
this.emit('respawnBlocked', {
reason: 'usage_limit',
details: 'Usage limit reached — waiting for scheduled auto-resume',
});
this.setState('watching');
return;
}
// Start the respawn cycle
this.cycleCount++;
this.log(`Starting respawn cycle #${this.cycleCount}`);
+207
View File
@@ -0,0 +1,207 @@
/**
* @fileoverview Pure cross-session federated search core (COD-9).
*
* `searchSources()` is the testable heart of `GET /api/search`: it takes a
* normalized query plus already-collected, in-memory source data and returns
* grouped, ranked, and capped results. It performs NO I/O — the route wrapper
* (`src/web/routes/search-routes.ts`) is responsible for harvesting the source
* arrays from the live server stores (sessions, run-summary trackers, attachment
* histories) in a bounded way before calling this.
*
* v1 scope (do not expand here): three sources — sessions/cases, run-summary
* events, file paths. Terminal-buffer scanning and any persisted index are
* explicitly deferred.
*
* Ranking: results are grouped by source type in the fixed order
* sessions → events → files. Within each group, exact (case-insensitive)
* name/path matches come first, then recency (newest timestamp first) as the
* tiebreak. There is no relevance-scoring pass in v1.
*
* Safety: file results only ever expose a workspace-relative path — server-
* private absolute paths are never placed in a result. Per-group and total caps
* bound the output so a broad query cannot return an unbounded payload.
*
* Key exports:
* - searchSources() — the pure core.
* - SEARCH_TOTAL_CAP / SEARCH_PER_GROUP_CAP — the output bounds.
* - SearchSources and the *Input row types — the source-data contract.
*/
import type { SearchResult, SearchResultGroup, SearchResponseData, SearchSourceType } from './types/search.js';
/** Maximum results returned across all groups combined. */
export const SEARCH_TOTAL_CAP = 60;
/** Maximum results returned within any single source group. */
export const SEARCH_PER_GROUP_CAP = 25;
/** Maximum characters in a result snippet. */
export const SEARCH_SNIPPET_MAX = 200;
/** A live-session row harvested for the session/case source. */
export interface SessionSearchInput {
sessionId: string;
sessionName: string;
workingDir: string;
/** Recency timestamp (e.g. lastActivityAt or createdAt). */
timestamp: number;
}
/** A run-summary timeline event harvested for the event source. */
export interface EventSearchInput {
sessionId: string;
sessionName: string;
eventId: string;
title: string;
details: string;
timestamp: number;
}
/** A per-session attachment harvested for the file source. */
export interface FileSearchInput {
sessionId: string;
sessionName: string;
fileName: string;
/** Workspace-relative path, if known. Absolute/external paths are never passed in. */
relativePath: string | undefined;
timestamp: number;
/** Attachment history item id, used as the jump-to target. */
itemId: string;
}
/** The full set of in-memory source data the pure core searches over. */
export interface SearchSources {
sessions: SessionSearchInput[];
events: EventSearchInput[];
files: FileSearchInput[];
}
/** Fixed group/render order. */
const GROUP_ORDER: SearchSourceType[] = ['session', 'event', 'file'];
function truncate(text: string, max = SEARCH_SNIPPET_MAX): string {
const trimmed = text.trim().replace(/\s+/g, ' ');
return trimmed.length > max ? trimmed.slice(0, max - 1) + '…' : trimmed;
}
/**
* Sort a group's results: exact matches first, then newest timestamp first.
* Stable for equal keys.
*/
function sortGroup(rows: SearchResult[]): SearchResult[] {
return rows
.map((result, index) => ({ result, index }))
.sort((a, b) => {
if (a.result.exactMatch !== b.result.exactMatch) {
return a.result.exactMatch ? -1 : 1;
}
if (a.result.timestamp !== b.result.timestamp) {
return b.result.timestamp - a.result.timestamp;
}
return a.index - b.index;
})
.map((r) => r.result);
}
/**
* Search the provided in-memory sources for `query`.
*
* @param query Raw query string (already length-validated by the route). Blank
* queries return an empty result set.
* @param sources Harvested, bounded source arrays.
*/
export function searchSources(query: string, sources: SearchSources): SearchResponseData {
const needle = query.trim().toLowerCase();
if (needle.length === 0) {
return { query: query.trim(), groups: [], totalResults: 0, truncated: false };
}
const contains = (s: string | undefined): boolean => !!s && s.toLowerCase().includes(needle);
const isExact = (s: string | undefined): boolean => !!s && s.toLowerCase() === needle;
// -- Source: sessions/cases --
const sessionRows: SearchResult[] = [];
for (const s of sources.sessions) {
if (contains(s.sessionName) || contains(s.workingDir) || contains(s.sessionId)) {
sessionRows.push({
type: 'session',
sessionId: s.sessionId,
sessionName: s.sessionName,
timestamp: s.timestamp,
snippet: truncate(s.workingDir ? `${s.sessionName} — ${s.workingDir}` : s.sessionName),
exactMatch: isExact(s.sessionName),
jumpTo: { kind: 'session', sessionId: s.sessionId },
});
}
}
// -- Source: run-summary events --
const eventRows: SearchResult[] = [];
for (const e of sources.events) {
if (contains(e.title) || contains(e.details)) {
const snippetBase = e.details && contains(e.details) ? `${e.title}: ${e.details}` : e.title;
eventRows.push({
type: 'event',
sessionId: e.sessionId,
sessionName: e.sessionName,
timestamp: e.timestamp,
snippet: truncate(snippetBase),
exactMatch: isExact(e.title),
jumpTo: { kind: 'run-summary', sessionId: e.sessionId, targetId: e.eventId },
});
}
}
// -- Source: file paths --
const fileRows: SearchResult[] = [];
for (const f of sources.files) {
if (contains(f.fileName) || contains(f.relativePath)) {
fileRows.push({
type: 'file',
sessionId: f.sessionId,
sessionName: f.sessionName,
timestamp: f.timestamp,
snippet: truncate(f.relativePath ?? f.fileName),
// Exact match keys off the safe path (or filename) — never an absolute path.
exactMatch: isExact(f.relativePath) || isExact(f.fileName),
jumpTo: {
kind: 'file-preview',
sessionId: f.sessionId,
targetId: f.itemId,
// Only ever expose a relative path; absolute/external paths are not passed in.
relativePath: f.relativePath,
},
});
}
}
const byType: Record<SearchSourceType, SearchResult[]> = {
session: sortGroup(sessionRows),
event: sortGroup(eventRows),
file: sortGroup(fileRows),
};
const groups: SearchResultGroup[] = [];
let total = 0;
let truncated = false;
for (const type of GROUP_ORDER) {
const all = byType[type];
if (all.length === 0) continue;
// Per-group cap.
let capped = all.slice(0, SEARCH_PER_GROUP_CAP);
if (all.length > capped.length) truncated = true;
// Total cap (never exceed the global budget).
const remaining = SEARCH_TOTAL_CAP - total;
if (capped.length > remaining) {
capped = capped.slice(0, Math.max(0, remaining));
truncated = true;
}
if (capped.length === 0) continue;
groups.push({ type, results: capped });
total += capped.length;
}
return { query: query.trim(), groups, totalResults: total, truncated };
}
+260
View File
@@ -0,0 +1,260 @@
/**
* @fileoverview Pure merge/filter logic for the unified session list (COD-121).
*
* Combines four read-only views of a session — live (in-memory `Session`),
* persisted (`state.json`), transcript history (`~/.claude/projects`), and the
* lifecycle audit log — plus mux process stats, into one de-duplicated list
* keyed by sessionId. Transcript-history rows are keyed by the Claude
* conversation UUID (the `.jsonl` filename stem), which diverges from the
* Codeman id for resumed sessions — an alias map (claudeSessionId → Codeman id,
* built from the live/persisted views) folds them into the owning session item.
* Higher-precedence sources overwrite scalar fields when present
* (history < lifecycle < persisted < live), while the `sources` array
* always accumulates every contributing view. A "meaningfulness floor" drops
* noise (bare lifecycle/mux-only rows with no name and no first prompt).
*
* PURE: no fs/IO and no node imports. All IO happens in the route that feeds
* this module its inputs, which keeps the merge/sort/filter behavior unit-testable.
*/
export type UnifiedSessionItem = {
sessionId: string;
name?: string;
mode?: string;
status?: string;
isWorking?: boolean;
workingDir?: string;
createdAt?: number;
lastActivityAt?: number;
claudeSessionId?: string;
firstPrompt?: string;
sizeBytes?: number;
projectKey?: string;
remote?: boolean;
sources: string[];
stats?: { memoryMB: number; cpuPercent: number };
};
/** Live in-memory session view (subset of `Session.toState()`). */
export type LiveSessionInput = {
id: string;
name?: string;
mode?: string;
status?: string;
isWorking?: boolean;
workingDir?: string;
createdAt?: number;
lastActivityAt?: number;
claudeSessionId?: string;
};
/** Persisted session view (subset of `SessionState`). */
export type PersistedSessionInput = {
id: string;
name?: string;
mode?: string;
status?: string;
workingDir?: string;
createdAt?: number;
lastActivityAt?: number;
/** Claude conversation ID this session resumes (`SessionState.resumeSessionId`). */
claudeSessionId?: string;
};
/** Lifecycle audit-log view. Entries are expected NEWEST-first (the order `SessionLifecycleLog.query()` returns). */
export type LifecycleInput = {
sessionId: string;
name?: string;
mode?: string;
ts: number;
event?: string;
};
/** Transcript-history view (one `.jsonl` per session). */
export type HistoryInput = {
sessionId: string;
workingDir: string;
sizeBytes: number;
lastModified: string;
firstPrompt?: string;
projectKey?: string;
};
/** Mux process-stat view. */
export type MuxStatInput = {
sessionId: string;
muxName?: string;
mode?: string;
stats?: { memoryMB: number; cpuPercent: number };
remote?: boolean;
};
export type UnifiedSources = {
live?: LiveSessionInput[];
persisted?: PersistedSessionInput[];
lifecycle?: LifecycleInput[];
history?: HistoryInput[];
mux?: MuxStatInput[];
};
/** Push a source tag onto an item exactly once. */
function addSource(item: UnifiedSessionItem, source: string): void {
if (!item.sources.includes(source)) item.sources.push(source);
}
/** Get-or-create the accumulator item for a sessionId. */
function ensureItem(map: Map<string, UnifiedSessionItem>, sessionId: string): UnifiedSessionItem {
let item = map.get(sessionId);
if (!item) {
item = { sessionId, sources: [] };
map.set(sessionId, item);
}
return item;
}
/** Overwrite a scalar field only when the incoming value is defined. */
function overwrite<K extends keyof UnifiedSessionItem>(
item: UnifiedSessionItem,
key: K,
value: UnifiedSessionItem[K] | undefined
): void {
if (value !== undefined) item[key] = value;
}
/**
* Merge all source views into one list, applying precedence
* (history → lifecycle → persisted → live) and the meaningfulness floor.
*/
export function mergeUnifiedSessions(sources: UnifiedSources): UnifiedSessionItem[] {
const map = new Map<string, UnifiedSessionItem>();
// Alias map: Claude conversation UUID → owning Codeman session id. Resumed
// (claudeSessionId = resumeSessionId != id) and /clear-respawned sessions
// would otherwise surface twice — once as a live/persisted row and once as a
// separate history-only row keyed by the conversation UUID. Live wins over
// persisted on conflicting entries (registered last).
const aliasToOwner = new Map<string, string>();
for (const p of sources.persisted ?? []) {
if (p.claudeSessionId !== undefined && p.claudeSessionId !== p.id) aliasToOwner.set(p.claudeSessionId, p.id);
}
for (const v of sources.live ?? []) {
if (v.claudeSessionId !== undefined && v.claudeSessionId !== v.id) aliasToOwner.set(v.claudeSessionId, v.id);
}
const resolveId = (sessionId: string): string => aliasToOwner.get(sessionId) ?? sessionId;
// 1) history (lowest precedence; keys resolve through the alias map)
for (const h of sources.history ?? []) {
const item = ensureItem(map, resolveId(h.sessionId));
addSource(item, 'history');
overwrite(item, 'workingDir', h.workingDir);
overwrite(item, 'sizeBytes', h.sizeBytes);
overwrite(item, 'firstPrompt', h.firstPrompt);
overwrite(item, 'projectKey', h.projectKey);
const ms = Date.parse(h.lastModified);
if (!Number.isNaN(ms) && item.lastActivityAt === undefined) item.lastActivityAt = ms;
}
// 2) lifecycle — entries arrive NEWEST-first, so first-seen wins for
// name/mode (mirrors the lastActivityAt guard); unconditional overwrites
// would leave the OLDEST entry in the window (stale name/mode) standing.
for (const l of sources.lifecycle ?? []) {
const item = ensureItem(map, resolveId(l.sessionId));
addSource(item, 'lifecycle');
if (item.name === undefined) overwrite(item, 'name', l.name);
if (item.mode === undefined) overwrite(item, 'mode', l.mode);
if (item.lastActivityAt === undefined && typeof l.ts === 'number') item.lastActivityAt = l.ts;
}
// 3) persisted
for (const p of sources.persisted ?? []) {
const item = ensureItem(map, p.id);
addSource(item, 'persisted');
overwrite(item, 'name', p.name);
overwrite(item, 'mode', p.mode);
overwrite(item, 'status', p.status);
overwrite(item, 'workingDir', p.workingDir);
overwrite(item, 'createdAt', p.createdAt);
overwrite(item, 'lastActivityAt', p.lastActivityAt);
}
// 4) live (highest precedence)
for (const v of sources.live ?? []) {
const item = ensureItem(map, v.id);
addSource(item, 'live');
overwrite(item, 'name', v.name);
overwrite(item, 'mode', v.mode);
overwrite(item, 'status', v.status);
overwrite(item, 'isWorking', v.isWorking);
overwrite(item, 'workingDir', v.workingDir);
overwrite(item, 'createdAt', v.createdAt);
overwrite(item, 'lastActivityAt', v.lastActivityAt);
overwrite(item, 'claudeSessionId', v.claudeSessionId);
}
// 5) mux stats + remote flag (create item if mux-only)
for (const m of sources.mux ?? []) {
const item = ensureItem(map, m.sessionId);
addSource(item, 'mux');
overwrite(item, 'mode', m.mode);
if (m.stats) item.stats = m.stats;
if (m.remote !== undefined) item.remote = m.remote;
}
// Meaningfulness floor: keep real rows, drop bare lifecycle/mux-only noise.
const kept: UnifiedSessionItem[] = [];
for (const item of map.values()) {
const isReal =
item.sources.includes('live') ||
item.sources.includes('persisted') ||
item.sources.includes('history') ||
(item.firstPrompt !== undefined && item.firstPrompt !== '');
if (isReal) kept.push(item);
}
// Stable sort: lastActivityAt desc (undefined last), createdAt desc, sessionId asc.
kept.sort((a, b) => {
const la = a.lastActivityAt;
const lb = b.lastActivityAt;
if (la !== lb) {
if (la === undefined) return 1;
if (lb === undefined) return -1;
return lb - la;
}
const ca = a.createdAt;
const cb = b.createdAt;
if (ca !== cb) {
if (ca === undefined) return 1;
if (cb === undefined) return -1;
return cb - ca;
}
return a.sessionId < b.sessionId ? -1 : a.sessionId > b.sessionId ? 1 : 0;
});
return kept;
}
/**
* Case-insensitive substring filter (name + firstPrompt + workingDir + sessionId)
* with offset/limit paging. `total` is the filtered count BEFORE paging.
*/
export function filterAndPaginate(
items: UnifiedSessionItem[],
opts: { q?: string; offset?: number; limit?: number }
): { sessions: UnifiedSessionItem[]; total: number } {
const q = (opts.q ?? '').trim().toLowerCase();
const filtered = q
? items.filter((it) => {
const hay = [it.name, it.firstPrompt, it.workingDir, it.sessionId]
.filter((v): v is string => typeof v === 'string')
.join(' ')
.toLowerCase();
return hay.includes(q);
})
: items;
const total = filtered.length;
const offset = Math.max(0, Math.floor(opts.offset ?? 0));
const limit = Math.min(500, Math.max(1, Math.floor(opts.limit ?? 100)));
const sessions = filtered.slice(offset, offset + limit);
return { sessions, total };
}
+92
View File
@@ -0,0 +1,92 @@
import { createHash } from 'node:crypto';
import { basename, extname } from 'node:path';
import type { SessionAttachmentHistoryItem } from './types/session.js';
import type { AttachmentDetectedEvent } from './types/tools.js';
import { getAttachmentType } from './attachment-registry.js';
export const ATTACHMENT_HISTORY_LIMIT = 100;
export interface ExternalAttachmentHistoryInput {
sessionId: string;
externalPath: string;
fileName?: string;
extension?: string;
size: number;
mtimeMs?: number;
timestamp?: number;
}
export function normalizeAttachmentExtension(extensionOrPath: string): string {
const value = extensionOrPath.startsWith('.') ? extensionOrPath : extname(extensionOrPath) || extensionOrPath;
return value.toLowerCase().replace(/^\./, '');
}
function historyKey(item: SessionAttachmentHistoryItem): string {
if (item.source === 'external' && item.externalPath) {
return `external:${item.externalPath}`;
}
return `detected:${item.relativePath || item.fileName}`;
}
function safeExternalHistoryId(item: SessionAttachmentHistoryItem): string {
const source = item.externalPath || item.id || item.fileName;
const digest = createHash('sha256').update(source).digest('hex').slice(0, 16);
return `external:${digest}:${item.fileName}`;
}
export function sanitizeAttachmentHistoryItem(item: SessionAttachmentHistoryItem): SessionAttachmentHistoryItem {
const { externalPath: _externalPath, ...safe } = item;
return {
...safe,
id: item.source === 'external' ? safeExternalHistoryId(item) : item.id,
};
}
export function sanitizeAttachmentHistory(
history: readonly SessionAttachmentHistoryItem[]
): SessionAttachmentHistoryItem[] {
return history.map(sanitizeAttachmentHistoryItem);
}
export function upsertAttachmentHistory(
history: readonly SessionAttachmentHistoryItem[],
item: SessionAttachmentHistoryItem
): SessionAttachmentHistoryItem[] {
const nextKey = historyKey(item);
return [item, ...history.filter((existing) => historyKey(existing) !== nextKey)].slice(0, ATTACHMENT_HISTORY_LIMIT);
}
export function buildDetectedAttachmentHistoryItem(event: AttachmentDetectedEvent): SessionAttachmentHistoryItem {
return {
id: `detected:${event.relativePath || event.fileName}`,
sessionId: event.sessionId,
fileName: event.fileName,
extension: normalizeAttachmentExtension(event.extension),
attachmentType: event.attachmentType,
size: event.size,
mtimeMs: 0,
timestamp: event.timestamp,
source: 'detected',
relativePath: event.relativePath,
};
}
export function buildExternalAttachmentHistoryItem(
input: ExternalAttachmentHistoryInput
): SessionAttachmentHistoryItem {
const extension = normalizeAttachmentExtension(input.extension || input.fileName || input.externalPath);
return {
id: `external:${createHash('sha256').update(input.externalPath).digest('hex').slice(0, 16)}:${
input.fileName || basename(input.externalPath)
}`,
sessionId: input.sessionId,
fileName: input.fileName || basename(input.externalPath),
extension,
attachmentType: getAttachmentType(extension),
size: input.size,
mtimeMs: input.mtimeMs ?? 0,
timestamp: input.timestamp ?? Date.now(),
source: 'external',
externalPath: input.externalPath,
};
}
+184 -1
View File
@@ -1,15 +1,25 @@
/**
* @fileoverview Auto-compact and auto-clear automation for Session.
* @fileoverview Auto-compact, auto-clear, and auto-resume automation for Session.
*
* Monitors token counts and triggers /compact or /clear commands when
* configurable thresholds are reached. Waits for Claude to be idle
* before sending commands, with retry logic and mutual exclusion
* (compact and clear never run simultaneously).
*
* Also implements auto-resume on usage limit ("token pause" control):
* when enabled and Claude stops on a usage-limit message ("5-hour limit
* reached ∙ resets 8pm" and friends — see usage-limit-patterns.ts), a timer
* is armed for the parsed reset time plus a safety buffer, then Escape
* (dismisses the rate-limit options dialog if open) and a "continue" prompt
* are sent so work resumes automatically. If the session is still limited,
* the fresh limit message re-arms the scheduler — that retry loop is the
* safety net for clock skew and parse imprecision.
*
* @module session-auto-ops
*/
import { EventEmitter } from 'node:events';
import { detectUsageLimitPause } from './usage-limit-patterns.js';
// ============================================================================
// Timing Constants
@@ -78,6 +88,28 @@ async function executeWhenIdle(
}
}
// ============================================================================
// Auto-resume (usage-limit pause) constants
// ============================================================================
/** Safety buffer after the stated reset time before resuming (2 minutes) */
const RESUME_BUFFER_MS = 2 * 60_000;
/** Minimum delay before an overdue resume fires (lets output settle) */
const RESUME_MIN_DELAY_MS = 5_000;
/** Retry interval when the reset time is stale/past (5 minutes) */
const RESUME_RETRY_MS = 5 * 60_000;
/** Re-detections scheduling within this window of the current schedule are ignored */
const RESUME_DEDUP_TOLERANCE_MS = 90_000;
/** Delay between Escape (dialog dismiss) and the resume prompt */
const RESUME_ESC_DELAY_MS = 600;
/** Prompt sent to resume work after the limit resets */
const RESUME_PROMPT = 'continue';
/** Minimum valid threshold for auto-clear/compact (1000 tokens) */
const MIN_AUTO_THRESHOLD = 1000;
@@ -131,6 +163,16 @@ export class SessionAutoOps extends EventEmitter {
private _isClearing: boolean = false;
private _autoClearTimer: NodeJS.Timeout | null = null;
// Auto-resume (usage-limit pause) state
private _autoResumeEnabled: boolean = false;
private _autoResumeTimer: NodeJS.Timeout | null = null;
/** Esc→continue gap timer; detections must NOT cancel a resume in flight */
private _resumeFollowupTimer: NodeJS.Timeout | null = null;
/** When the scheduled resume fires (epoch ms), null when not armed */
private _autoResumeAt: number | null = null;
private _limitPaused: boolean = false;
private _resumeAttempts: number = 0;
private readonly callbacks: AutoOpsCallbacks;
constructor(callbacks: AutoOpsCallbacks, config?: { compactThreshold?: number; clearThreshold?: number }) {
@@ -207,6 +249,145 @@ export class SessionAutoOps extends EventEmitter {
}
}
// ============================================================================
// Auto-resume (usage-limit pause) — getters/setters
// ============================================================================
get autoResumeEnabled(): boolean {
return this._autoResumeEnabled;
}
/** When the scheduled resume fires (epoch ms), or null when not armed. */
get autoResumeAt(): number | null {
return this._autoResumeAt;
}
/** True while the session is believed to be paused on a usage limit. */
get isLimitPaused(): boolean {
return this._limitPaused;
}
setAutoResume(enabled: boolean): void {
this._autoResumeEnabled = enabled;
if (!enabled) {
this._cancelAutoResume('disabled');
}
}
/**
* Restore auto-resume state after a Codeman restart. A persisted pending
* schedule is re-armed; an overdue one fires shortly after boot (the limit
* footer won't reprint on its own, so without this the pause would stall).
*/
restoreAutoResume(enabled: boolean, resumeAt?: number): void {
this._autoResumeEnabled = enabled;
if (!enabled || !resumeAt) return;
const now = Date.now();
this._scheduleResume(Math.max(resumeAt, now + RESUME_MIN_DELAY_MS), resumeAt, 'restored');
}
// ============================================================================
// Auto-resume — detection and scheduling
// ============================================================================
/**
* Scan cleaned terminal output for a usage-limit pause message and (re)arm
* the resume schedule. Called from the session's throttled parser path.
*/
processCleanData(cleanData: string): void {
if (!this._autoResumeEnabled || this.callbacks.isStopped()) return;
// A resume is in flight (Esc sent, continue pending): output from our own
// Escape can redraw the stale limit footer — don't let it re-arm and
// cancel the continue. Fresh evidence arrives after the prompt is sent.
if (this._resumeFollowupTimer) return;
const detection = detectUsageLimitPause(cleanData);
if (!detection) return;
const now = Date.now();
const overdue = detection.resetAt <= now;
const fireAt = overdue
? now + RESUME_RETRY_MS // stale reset time → gentle retry loop
: Math.max(detection.resetAt + RESUME_BUFFER_MS, now + RESUME_MIN_DELAY_MS);
if (this._autoResumeTimer && this._autoResumeAt !== null) {
// Already armed: the footer redraws constantly, so ignore re-detections
// that land on (or later than) the current schedule. Only an EARLIER
// parsed time replaces it — an overdue retry never preempts a real one.
if (overdue || fireAt >= this._autoResumeAt - RESUME_DEDUP_TOLERANCE_MS) return;
}
this._scheduleResume(fireAt, detection.resetAt, detection.matched);
}
/**
* Claude started working — the limit is lifted (or the user resumed
* manually), so any pending auto-resume is obsolete.
*/
notifyWorking(): void {
this._resumeAttempts = 0;
if (!this._limitPaused && !this._autoResumeTimer && !this._resumeFollowupTimer) return;
this._cancelAutoResume('working');
}
private _scheduleResume(fireAt: number, resetAt: number, matched: string): void {
if (this._autoResumeTimer) {
clearTimeout(this._autoResumeTimer);
this._autoResumeTimer = null;
}
this._limitPaused = true;
this._autoResumeAt = fireAt;
const delay = Math.max(fireAt - Date.now(), 0);
console.log(
`[SessionAutoOps ${this.callbacks.getSessionId()}] Usage-limit pause detected ("${matched.slice(0, 60)}"), auto-resume in ${Math.round(delay / 60000)}min`
);
this._autoResumeTimer = setTimeout(() => void this._fireResume(), delay);
this.emit('limitPauseScheduled', { resetAt, resumeAt: fireAt, matched });
}
private async _fireResume(): Promise<void> {
this._autoResumeTimer = null;
if (!this._autoResumeEnabled || this.callbacks.isStopped()) return;
if (this.callbacks.isWorking()) {
// Session resumed on its own (or via the user) — nothing to do.
this._cancelAutoResume('working');
return;
}
this._resumeAttempts++;
const attempt = this._resumeAttempts;
this._limitPaused = false; // optimistic: a fresh limit message re-arms us
this._autoResumeAt = null;
// Escape first: dismisses the rate-limit options dialog if Claude opened
// one (harmless at an idle prompt), then the resume prompt after a beat.
await this.callbacks.writeCommand('\x1b');
this._resumeFollowupTimer = setTimeout(() => {
this._resumeFollowupTimer = null;
if (this.callbacks.isStopped()) return;
void this.callbacks.writeCommand(`${RESUME_PROMPT}\r`);
this.emit('limitResume', { attempt });
}, RESUME_ESC_DELAY_MS);
}
private _cancelAutoResume(reason: 'disabled' | 'working' | 'stopped'): void {
const wasArmed = this._autoResumeTimer !== null || this._resumeFollowupTimer !== null || this._limitPaused;
if (this._autoResumeTimer) {
clearTimeout(this._autoResumeTimer);
this._autoResumeTimer = null;
}
if (this._resumeFollowupTimer) {
clearTimeout(this._resumeFollowupTimer);
this._resumeFollowupTimer = null;
}
this._limitPaused = false;
this._autoResumeAt = null;
if (wasArmed && reason !== 'stopped') {
this.emit('limitResumeCancelled', { reason });
}
}
// ============================================================================
// Threshold checks
// ============================================================================
@@ -321,5 +502,7 @@ export class SessionAutoOps extends EventEmitter {
this._autoClearTimer = null;
}
this._isClearing = false;
this._cancelAutoResume('stopped');
}
}
+30 -7
View File
@@ -11,6 +11,7 @@
import type { ClaudeMode, EffortLevel } from './types.js';
import { isEffortLevel } from './types.js';
import { getAugmentedPath } from './utils/index.js';
import { dataPath } from './config/instance.js';
/**
* Build Claude CLI permission flags based on the configured mode.
@@ -101,19 +102,24 @@ export function buildPromptArgs(prompt: string, model?: string): string[] {
* @returns Environment variables object for pty.spawn
*/
export function buildClaudeEnv(sessionId: string): Record<string, string | undefined> {
return {
const env: Record<string, string | undefined> = {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
PATH: getAugmentedPath(),
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
// Inform Claude it's running within Codeman (helps prevent self-termination)
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
// COD-115: `delete`, not `= undefined` — node-pty serializes a present-with-undefined
// key as the literal string "KEY=undefined" (see buildMuxAttachEnv below).
delete env.COLORTERM;
delete env.CLAUDECODE;
return env;
}
/**
@@ -121,17 +127,32 @@ export function buildClaudeEnv(sessionId: string): Record<string, string | undef
* Lighter than buildClaudeEnv — no PATH augmentation or Codeman vars needed
* since the mux session already has those set.
*
* @param truecolorEnabled - When true, set COLORTERM=truecolor (COD-75 opt-in);
* otherwise leave COLORTERM unset. Mirrors buildEnvExports() so both paths agree.
* @returns Environment variables object for pty.spawn
*/
export function buildMuxAttachEnv(): Record<string, string | undefined> {
return {
export function buildMuxAttachEnv(truecolorEnabled?: boolean): Record<string, string | undefined> {
const env: Record<string, string | undefined> = {
...process.env,
LANG: 'en_US.UTF-8',
LC_ALL: 'en_US.UTF-8',
TERM: 'xterm-256color',
COLORTERM: undefined,
CLAUDECODE: undefined,
};
// COD-115: keys to UNSET must be `delete`d, NOT set to `undefined`. On a
// `{...process.env}` spread the key stays present with value undefined, and node-pty
// serializes it as the literal string "TMUX=undefined" — a non-empty value that still
// trips tmux's nesting guard, killing the attach-bridge PTY (exit 1 → respawn loop).
// The server can be launched from inside tmux; attach clients must never inherit that
// parent tmux context. (Same fix the working create path uses in tmux-manager.ts.)
delete env.TMUX;
delete env.TMUX_PANE;
delete env.CLAUDECODE;
if (truecolorEnabled) {
env.COLORTERM = 'truecolor';
} else {
delete env.COLORTERM; // COD-75: unset for non-truecolor (was `: undefined`, same node-pty quirk)
}
return env;
}
/**
@@ -149,5 +170,7 @@ export function buildShellEnv(sessionId: string): Record<string, string | undefi
CODEMAN_MUX: '1',
CODEMAN_SESSION_ID: sessionId,
CODEMAN_API_URL: process.env.CODEMAN_API_URL || 'http://localhost:3000',
// Path only (not the secret value) — hook curls cat it at execution time (COD-54)
CODEMAN_HOOK_SECRET_FILE: dataPath('hook-secret'),
};
}
+102
View File
@@ -0,0 +1,102 @@
/**
* @fileoverview Circuit breaker bounding repeated non-zero interactive-PTY exits (COD-118).
*
* Defense-in-depth after COD-115: if the interactive PTY exits non-zero repeatedly,
* external recovery/reconnect paths recreate it indefinitely (COD-115 observed 114
* `exited with code: 1` events + orphan sessions). This breaker tracks recent
* non-zero exits within a sliding window and "trips" once they exceed a threshold,
* so the Session can refuse to respawn and surface an error state instead of looping.
*
* Design notes:
* - PURE + dependency-free. Time is INJECTED (`nowMs` passed to `recordExit`); the
* breaker never calls `Date.now()` itself, so trip/window logic is deterministically
* unit-testable with no real timers.
* - A clean (exit code 0) exit resets the counter — a session that exited normally is
* not on a crash-loop. (It does NOT clear an already-tripped breaker; only an explicit
* `reset()` — e.g. a user-initiated restart — does that.)
* - Once tripped, stays tripped until `reset()`.
*
* @consumedby session (instantiates one per session; records exits in the interactive
* PTY `onExit` handler; gates `startInteractive()` when tripped; `reset()` on restart)
* @module session-pty-exit-breaker
*/
/** Non-zero interactive-PTY exits within the window required to trip the breaker. */
export const DEFAULT_BREAKER_THRESHOLD = 5;
/** Sliding window (ms) over which non-zero exits accumulate toward the threshold. */
export const DEFAULT_BREAKER_WINDOW_MS = 10_000;
export interface InteractivePtyExitBreakerOptions {
/** Trip after this many non-zero exits within `windowMs` (default 5). */
threshold?: number;
/** Sliding window length in ms (default 10_000). */
windowMs?: number;
}
export interface RecordExitResult {
/** True once the breaker has tripped (stays true until `reset()`). */
tripped: boolean;
/** Number of non-zero exits currently inside the window. */
count: number;
}
/**
* Sliding-window counter that trips on rapid repeated non-zero exits.
*
* 5 within 10s safely clears normal usage (a single exit, an intentional restart)
* while tripping fast on a real loop — COD-115 saw 114 exits, far above 5.
*/
export class InteractivePtyExitBreaker {
private readonly _threshold: number;
private readonly _windowMs: number;
/** Timestamps (ms, injected) of recent non-zero exits, oldest first. */
private _exitTimes: number[] = [];
private _tripped = false;
constructor(opts: InteractivePtyExitBreakerOptions = {}) {
this._threshold = opts.threshold ?? DEFAULT_BREAKER_THRESHOLD;
this._windowMs = opts.windowMs ?? DEFAULT_BREAKER_WINDOW_MS;
}
/** Whether the breaker has tripped (respawn should be blocked). */
get tripped(): boolean {
return this._tripped;
}
/**
* Record a PTY exit. A zero (clean) exit resets the non-zero counter; a non-zero
* exit is added to the window, stale entries are evicted, and the breaker trips
* once the in-window count reaches the threshold.
*
* @param exitCode the PTY exit code (0 = clean)
* @param nowMs injected current time in ms (never read from a real clock)
*/
recordExit(exitCode: number, nowMs: number): RecordExitResult {
if (exitCode === 0) {
// Clean exit: a normal stop, not a crash-loop. Clear accumulated non-zero
// exits. Does NOT un-trip an already-tripped breaker (only reset() does).
this._exitTimes = [];
return { tripped: this._tripped, count: 0 };
}
// Evict exits strictly older than the window, then record this one.
const cutoff = nowMs - this._windowMs;
this._exitTimes = this._exitTimes.filter((t) => t > cutoff);
this._exitTimes.push(nowMs);
if (this._exitTimes.length >= this._threshold) {
this._tripped = true;
}
return { tripped: this._tripped, count: this._exitTimes.length };
}
/** Clear the tripped state and the non-zero counter (e.g. on intentional restart). */
reset(): void {
this._exitTimes = [];
this._tripped = false;
}
}
+472 -12
View File
@@ -48,7 +48,11 @@ import {
type OpenCodeConfig,
type CodexConfig,
type EffortLevel,
type GeminiConfig,
type SessionRemote,
type SessionDocker,
} from './types.js';
import { probeDockerCliVersion } from './docker-hosts.js';
import type { TerminalMultiplexer, MuxSession } from './mux-interface.js';
import { TaskTracker, type BackgroundTask } from './task-tracker.js';
import { RalphTracker } from './ralph-tracker.js';
@@ -60,6 +64,7 @@ import {
SPINNER_PATTERN,
MAX_SESSION_TOKENS,
execPattern,
getClaudeCliVersion,
} from './utils/index.js';
import {
MAX_TERMINAL_BUFFER_SIZE,
@@ -69,6 +74,7 @@ import {
MAX_MESSAGES,
MAX_LINE_BUFFER_SIZE,
} from './config/buffer-limits.js';
import { DEFAULT_TMUX_HISTORY_LIMIT } from './config/terminal-history.js';
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import {
buildInteractiveArgs,
@@ -78,7 +84,15 @@ import {
buildShellEnv,
} from './session-cli-builder.js';
import { SessionAutoOps } from './session-auto-ops.js';
import { detectUsageLimitPause } from './usage-limit-patterns.js';
import { SessionTaskCache } from './session-task-cache.js';
import { InteractivePtyExitBreaker } from './session-pty-exit-breaker.js';
import { parseTerminalAttachmentRequests } from './attachment-magic.js';
import {
sanitizeAttachmentHistory,
upsertAttachmentHistory as upsertAttachmentHistoryList,
} from './session-attachment-history.js';
import type { SessionAttachmentHistoryItem } from './types/session.js';
export type { BackgroundTask } from './task-tracker.js';
export type { RalphTrackerState, RalphTodoItem, ActiveBashTool } from './types.js';
@@ -126,7 +140,38 @@ const NEWLINE_SPLIT_PATTERN = /\r?\n/;
/** True for external-CLI run modes (non-Claude) that use their own TUI and output format. */
export function isExternalCliMode(mode: SessionMode): boolean {
return mode === 'opencode' || mode === 'codex';
return mode === 'opencode' || mode === 'codex' || mode === 'gemini';
}
function getModeLabel(mode: SessionMode): string {
switch (mode) {
case 'opencode':
return 'OpenCode';
case 'codex':
return 'Codex';
case 'gemini':
return 'Gemini';
case 'shell':
return 'Shell';
case 'claude':
return 'Claude';
}
}
/**
* Modes whose TUI emits alt-screen / scrollback-erase / mouse-tracking sequences
* that we strip so the browser keeps everything in the main buffer with scrollback
* reachable (the strip runs on both the live stream and the buffer replay).
*
* Codex, Claude Code, and Gemini are known, controlled (Ink/React) TUIs that
* repaint via cursor positioning, so dropping the alt-screen switch is safe —
* content stays in the normal buffer. Excluded: `shell` (arbitrary programs like
* vim/less/htop legitimately need the alt screen) and `opencode` (renders its own
* TUI that may rely on it). Keep parity with the replay-side strip in
* session-routes.ts.
*/
export function isAltScreenStripMode(mode: SessionMode): boolean {
return mode === 'codex' || mode === 'claude' || mode === 'gemini';
}
// Note: Claude CLI PATH resolution moved to session-cli-builder.ts (buildClaudeEnv)
@@ -135,6 +180,8 @@ export function isExternalCliMode(mode: SessionMode): boolean {
const DEFAULT_PTY_COLS = 120;
const DEFAULT_PTY_ROWS = 40;
const TMUX_DISPLAY_TIMEOUT_MS = 2000;
/** Delay before the in-container Claude CLI version probe (lets the container start). */
const DOCKER_CLI_VERSION_PROBE_DELAY_MS = 3000;
/**
* Ask tmux for the current window geometry of `muxName` so a re-attaching PTY
@@ -169,6 +216,12 @@ export function queryTmuxWindowSize(muxName: string, socket: string): { cols: nu
return { cols: DEFAULT_PTY_COLS, rows: DEFAULT_PTY_ROWS };
}
export function resolveMuxAttachCwd(workingDir: string, remote?: SessionRemote, docker?: SessionDocker): string {
// Remote and docker sessions run the CLI elsewhere (ssh / docker exec); the LOCAL
// wrapper pane never needs the workspace as its cwd, so launch it in /tmp.
return remote || docker ? '/tmp' : workingDir;
}
/**
* Represents a JSON message from Claude CLI's stream-json output format.
* Messages are newline-delimited JSON objects parsed from PTY output.
@@ -248,6 +301,13 @@ export class Session extends EventEmitter {
private _pid: number | null = null;
private _status: SessionStatus = 'idle';
private _currentTaskId: string | null = null;
// COD-118: bound repeated non-zero interactive-PTY exits. Recorded in the
// interactive PTY onExit handler; when it trips, the session flips to 'error'
// and startInteractive() refuses to respawn until an explicit user restart
// calls resetRespawnBreaker(). Defense-in-depth over the COD-115 crash-loop.
private readonly _ptyExitBreaker = new InteractivePtyExitBreaker();
private _respawnBlocked = false;
// Use BufferAccumulator for hot-path buffers to reduce GC pressure
private _terminalBuffer = new BufferAccumulator(MAX_TERMINAL_BUFFER_SIZE, TERMINAL_BUFFER_TRIM_SIZE);
private _textOutput = new BufferAccumulator(MAX_TEXT_OUTPUT_SIZE, TEXT_OUTPUT_TRIM_SIZE);
@@ -258,6 +318,10 @@ export class Session extends EventEmitter {
private _messages: ClaudeMessage[] = [];
private _lineBuffer: string = '';
private _lineBufferFlushTimer: NodeJS.Timeout | null = null;
// Alt-screen-strip modes (Codex/Claude): trailing partial CSI held back so
// sequences split across PTY chunks can't slip past the alt-screen/scrollback
// strip (see _handleTerminalOutput / isAltScreenStripMode)
private _altScreenSeqCarry: string = '';
private resolvePromise: ((value: { result: string; cost: number }) => void) | null = null;
private rejectPromise: ((reason: Error) => void) | null = null;
private _promptResolved: boolean = false; // Guard against race conditions in runPrompt
@@ -307,6 +371,10 @@ export class Session extends EventEmitter {
private _parentAgentId: string | null = null;
private _childAgentIds: string[] = [];
// Bounded dedup set for terminal attachment magic-links already requested.
private _attachmentMagicSeen = new Set<string>();
private _attachmentHistory: SessionAttachmentHistoryItem[] = [];
// Nice prioritying configuration
private _niceConfig: NiceConfig = { ...DEFAULT_NICE_CONFIG };
@@ -321,6 +389,8 @@ export class Session extends EventEmitter {
private _openCodeConfig: OpenCodeConfig | undefined;
// Codex configuration (only for mode === 'codex')
private _codexConfig: CodexConfig | undefined;
// Gemini configuration (only for mode === 'gemini')
private _geminiConfig: GeminiConfig | undefined;
private _resumeSessionId: string | undefined;
// Ephemeral env overrides (e.g., CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS). Exported by tmux
@@ -332,6 +402,16 @@ export class Session extends EventEmitter {
// the CLAUDE_CODE_EFFORT_LEVEL env var, which would hard-lock the session.
private _effort: EffortLevel | undefined;
// tmux history-limit (scrollback lines) applied to this session's pane.
private readonly _tmuxHistoryLimit: number;
// Remote execution metadata, present when this session runs over SSH through local tmux.
private readonly _remote?: SessionRemote;
// Docker execution metadata, present when this session runs inside a container via
// local tmux + `docker exec`. The container is per-CASE (shared by sibling sessions).
private readonly _docker?: SessionDocker;
// Session color for visual differentiation
private _color: import('./types.js').SessionColor = 'default';
@@ -391,12 +471,22 @@ export class Session extends EventEmitter {
openCodeConfig?: OpenCodeConfig;
/** Codex configuration (only for mode === 'codex') */
codexConfig?: CodexConfig;
/** Gemini configuration (only for mode === 'gemini') */
geminiConfig?: GeminiConfig;
/** Resume a previous Claude conversation (used after server reboot) */
resumeSessionId?: string;
/** Extra env vars exported to the CLI at spawn time (no disk persistence) */
envOverrides?: Record<string, string>;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/** tmux history-limit (scrollback lines) for this session's pane. */
tmuxHistoryLimit?: number;
/** Restored per-session attachment history. May include server-private external paths. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/** Remote execution metadata for sessions launched through SSH inside local tmux. */
remote?: SessionRemote;
/** Docker execution metadata for sessions launched inside a container via local tmux. */
docker?: SessionDocker;
}
) {
super();
@@ -449,6 +539,11 @@ export class Session extends EventEmitter {
this._codexConfig = config.codexConfig;
}
// Apply Gemini configuration
if (config.geminiConfig) {
this._geminiConfig = config.geminiConfig;
}
// Apply env overrides (exported at spawn, not persisted to disk).
// Legacy migration: pre-0.7.2 carried effort as the CLAUDE_CODE_EFFORT_LEVEL env var,
// which hard-locks /effort switching. Extract it into _effort (--settings soft default)
@@ -463,6 +558,12 @@ export class Session extends EventEmitter {
if (config.effort && isEffortLevel(config.effort)) {
this._effort = config.effort;
}
this._tmuxHistoryLimit = config.tmuxHistoryLimit ?? DEFAULT_TMUX_HISTORY_LIMIT;
this._remote = config.remote;
this._docker = config.docker;
if (config.attachmentHistory && config.attachmentHistory.length > 0) {
this.restoreAttachmentHistory(config.attachmentHistory);
}
// Initialize task tracker and forward events (store handlers for cleanup)
this._taskTracker = new TaskTracker();
@@ -520,6 +621,9 @@ export class Session extends EventEmitter {
this._totalOutputTokens = 0;
this.emit('autoClear', data);
});
this._autoOps.on('limitPauseScheduled', (data) => this.emit('limitPauseScheduled', data));
this._autoOps.on('limitResume', (data) => this.emit('limitResume', data));
this._autoOps.on('limitResumeCancelled', (data) => this.emit('limitResumeCancelled', data));
}
get status(): SessionStatus {
@@ -558,6 +662,11 @@ export class Session extends EventEmitter {
return this._claudeSessionId;
}
/** Docker execution metadata when this session runs inside a container, else undefined. */
get docker(): SessionDocker | undefined {
return this._docker;
}
// Adopt a Claude conversation ID observed from an external source (e.g. hook
// payload). In interactive PTY mode Claude CLI emits no JSON to stdout, so
// `_handleJsonMessage` never sees `session_id`; hooks are the only signal
@@ -825,6 +934,39 @@ export class Session extends EventEmitter {
this._autoOps.setAutoCompact(enabled, threshold, prompt);
}
get autoResumeEnabled(): boolean {
return this._autoOps.autoResumeEnabled;
}
/** When the scheduled usage-limit auto-resume fires (epoch ms), or null. */
get autoResumeAt(): number | null {
return this._autoOps.autoResumeAt;
}
/** True while the session is paused on a Claude usage limit (auto-resume armed). */
get isLimitPaused(): boolean {
return this._autoOps.isLimitPaused;
}
setAutoResume(enabled: boolean): void {
this._autoOps.setAutoResume(enabled);
// Users typically enable this WHILE a session already sits paused — the
// limit footer won't reprint on its own, so scan the recent buffer once.
// Only a future reset time counts: stale scrollback must not arm a resume.
if (enabled && !isExternalCliMode(this.mode)) {
const tail = this._terminalBuffer.value.slice(-8192).replace(ANSI_ESCAPE_PATTERN_FULL, '');
const detection = detectUsageLimitPause(tail);
if (detection && detection.resetAt > Date.now()) {
this._autoOps.processCleanData(tail);
}
}
}
/** Restore auto-resume state (and a pending schedule) after Codeman restart. */
restoreAutoResume(enabled: boolean, resumeAt?: number): void {
this._autoOps.restoreAutoResume(enabled, resumeAt);
}
get imageWatcherEnabled(): boolean {
return this._imageWatcherEnabled;
}
@@ -853,12 +995,38 @@ export class Session extends EventEmitter {
return this._status === 'idle' || this._status === 'busy';
}
get attachmentHistory(): SessionAttachmentHistoryItem[] {
return sanitizeAttachmentHistory(this._attachmentHistory);
}
upsertAttachmentHistory(item: SessionAttachmentHistoryItem): void {
this._attachmentHistory = upsertAttachmentHistoryList(this._attachmentHistory, item);
}
restoreAttachmentHistory(history: SessionAttachmentHistoryItem[] | undefined): void {
this._attachmentHistory = [];
for (const item of [...(history ?? [])].reverse()) {
// Guard against malformed/legacy on-disk entries (null, non-object, or
// missing required fields). historyKey() dereferences source/fileName, so
// a bad item would otherwise throw inside the constructor and abort the
// entire mux-recovery loop.
if (!item || typeof item !== 'object' || !item.source || !item.fileName) continue;
this.upsertAttachmentHistory(item);
}
}
getAttachmentHistoryForPersist(): SessionAttachmentHistoryItem[] | undefined {
return this._attachmentHistory.length > 0 ? this._attachmentHistory.map((item) => ({ ...item })) : undefined;
}
toState(): SessionState {
return {
id: this.id,
pid: this.pid,
status: this._status,
workingDir: this.workingDir,
remote: this._remote,
docker: this._docker,
currentTaskId: this._currentTaskId,
createdAt: this.createdAt,
lastActivityAt: this._lastActivityAt,
@@ -869,6 +1037,8 @@ export class Session extends EventEmitter {
autoCompactEnabled: this._autoOps.autoCompactEnabled,
autoCompactThreshold: this._autoOps.autoCompactThreshold,
autoCompactPrompt: this._autoOps.autoCompactPrompt,
autoResumeEnabled: this._autoOps.autoResumeEnabled,
autoResumeAt: this._autoOps.autoResumeAt ?? undefined,
imageWatcherEnabled: this._imageWatcherEnabled,
totalCost: this._totalCost,
inputTokens: this._totalInputTokens,
@@ -888,8 +1058,15 @@ export class Session extends EventEmitter {
cliLatestVersion: this._cliLatestVersion || undefined,
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
resumeSessionId: this._resumeSessionId,
effort: this._effort,
// COD-118: runtime-only — surfaced so the frontend can require explicit user
// intent before restarting a crash-looped session. Deliberately NOT restored
// by the constructor: a Codeman restart starts with a fresh breaker so boot
// recovery can re-attach.
respawnBlocked: this._respawnBlocked || undefined,
attachmentHistory: this.attachmentHistory.length > 0 ? this.attachmentHistory : undefined,
// envOverrides intentionally NOT on the public SessionState type — they must not
// leak into SSE / GET /api/sessions broadcasts (schema allows OPENCODE_*, which
// can carry secrets). For disk persistence, session-manager calls
@@ -1039,8 +1216,10 @@ export class Session extends EventEmitter {
name: 'xterm-256color',
cols: ptyCols,
rows: ptyRows,
cwd: this.workingDir,
env: buildMuxAttachEnv(),
cwd: resolveMuxAttachCwd(this.workingDir, this._remote, this._docker),
// COD-75: codex/gemini get COLORTERM=truecolor — mirrors buildEnvExports()
// in tmux-manager.ts so the attach client and the tmux session agree.
env: buildMuxAttachEnv(this.mode === 'codex' || this.mode === 'gemini'),
});
} catch (spawnErr) {
console.error(`[Session] Failed to spawn PTY for ${options.spawnErrLabel}:`, spawnErr);
@@ -1052,6 +1231,74 @@ export class Session extends EventEmitter {
}
private _handleTerminalOutput(data: string): void {
// Codex AND Claude Code emit sequences that wipe xterm.js scrollback, plus
// mouse-tracking enables that hijack the scroll wheel so the user can't reach
// scrollback. Claude Code does this intermittently (e.g. full-screen pickers /
// dialogs), which is why terminal scroll-up "randomly" breaks for Claude
// sessions on mobile and desktop until the dialog closes:
// - \x1b[?1049h / \x1b[?47h / \x1b[?1047h: switch to the alt buffer (no
// scrollback) — \x1b[?...l switches back.
// - \x1b[3J: erase saved lines (scrollback). (\x1b[2J / \x1b[J — erase
// the visible viewport — are left intact; the TUI repaints those rows.)
// - \x1b[?1000h / 1002h / 1003h / 1005h / 1006h / 1007h: mouse-tracking
// modes (X10, button-event, any-event, UTF-8, SGR, alt-scroll). Once on,
// xterm.js forwards wheel events to the CLI instead of scrolling the
// viewport, so the conversation is in scrollback but unreachable.
// (Focus events at ?1004 are left alone — codeman uses them for
// active-tab detection.)
// Strip them at the source so neither the persisted buffer nor the live
// SSE/WS stream carries them, keeping everything in the main buffer with
// scrollback intact. These are controlled TUIs whose cursor-positioned
// redraws overwrite only the cells they target, so non-erased rows keep
// their content. Gated to Codex/Claude (isAltScreenStripMode) — shell must
// keep the alt screen for vim/less/htop.
if (isAltScreenStripMode(this.mode)) {
// Reassemble sequences split across PTY chunk boundaries first: a chunk
// ending mid-sequence ('\x1b[?104' now, '9h' next) would slip past the
// strip below and leave xterm stuck in the scrollback-less alt buffer
// until the next buffer replay. Hold back an incomplete digit-only CSI
// tail (≤7 chars — the longest strippable intro is '\x1b[?1049') and
// prepend it to the next chunk; complete sequences are never held.
data = this._altScreenSeqCarry + data;
this._altScreenSeqCarry = '';
// eslint-disable-next-line no-control-regex
const splitTail = data.match(/\x1b(?:\[\??[0-9]{0,4})?$/);
if (splitTail) {
this._altScreenSeqCarry = splitTail[0];
data = data.slice(0, -splitTail[0].length);
if (!data) return;
}
data = data
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:47|1047|1049)[hl]/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[3J/g, '')
// eslint-disable-next-line no-control-regex
.replace(/\x1b\[\?(?:1000|1001|1002|1003|1005|1006|1007)[hl]/g, '');
}
// Scan terminal output for attachment requests. `codeman://attach?...` is an
// explicit magic link (all modes); Codex generated images report
// `Saved to: file://...` — that scanner (and its relaxed trust policy) is
// only enabled for codex-mode sessions. The web server applies the trust
// boundary for each request source.
const attachmentRequests = parseTerminalAttachmentRequests(data, { codexArtifacts: this.mode === 'codex' });
for (const request of attachmentRequests) {
const seenKey = `${request.source}:${request.path}`;
if (this._attachmentMagicSeen.has(seenKey)) continue;
this._attachmentMagicSeen.add(seenKey);
if (this._attachmentMagicSeen.size > 200) {
const oldest = this._attachmentMagicSeen.values().next().value;
if (oldest) this._attachmentMagicSeen.delete(oldest);
}
this.emit('attachmentRequested', {
sessionId: this.id,
path: request.path,
source: request.source,
timestamp: Date.now(),
});
}
// BufferAccumulator handles auto-trimming when max size exceeded
this._terminalBuffer.append(data);
this._lastActivityAt = Date.now();
@@ -1064,13 +1311,69 @@ export class Session extends EventEmitter {
throw new Error('Session already has a running process');
}
// COD-118: if the PTY exit breaker has tripped (repeated non-zero exits in a
// short window), refuse to respawn. This is the uniform choke point that stops
// automatic recovery/reconnect callers from re-creating a crash-looping PTY.
// An explicit user restart clears it via resetRespawnBreaker().
if (this._respawnBlocked) {
throw new Error(
'Respawn blocked: interactive PTY exited non-zero too many times in a short window (circuit breaker tripped). Restart the session to clear it.'
);
}
this._resetBuffers();
const modeLabel = this.mode === 'opencode' ? 'OpenCode' : this.mode === 'codex' ? 'Codex' : 'Claude';
const modeLabel = getModeLabel(this.mode);
console.log(
`[Session] Starting interactive ${modeLabel} session` + (this._useMux ? ` (with ${this._mux!.backend})` : '')
);
// Seed the CLI version deterministically for LOCAL Claude sessions. The
// banner scrape in parseClaudeCodeInfo() is unreliable — newer Claude Code
// builds don't print "Claude Code vX.Y.Z" at startup and resumed sessions
// never show it — which left cliVersion undefined and silently disabled
// wheel-forwarding to Claude's own transcript (the only route to history in
// repaint/alt-screen mode; issue #154). Remote sessions run claude on
// another host, so a local probe wouldn't reflect their version — skip them
// and let the banner scrape handle those. Cached process-wide, best-effort.
if (this.mode === 'claude' && !this._remote && !this._docker && !this._cliVersion) {
const probedVersion = getClaudeCliVersion();
if (probedVersion) {
this._cliVersion = probedVersion;
this.emit('cliInfoUpdated', {
version: this._cliVersion,
model: this._cliModel,
accountType: this._cliAccountType,
latestVersion: this._cliLatestVersion,
});
}
}
// Docker sessions run claude INSIDE the container, so the local probe above
// reports the HOST claude (wrong version, and leaving cliVersion undefined
// silently disables wheel-forwarding, #154). Probe the IN-CONTAINER version
// instead — deferred so the container is up after the mux attach below.
if (this.mode === 'claude' && this._docker && !this._cliVersion) {
const dockerMeta = this._docker;
setTimeout(() => {
if (this._isStopped || this._cliVersion) return;
void probeDockerCliVersion(dockerMeta, this.mode)
.then((version) => {
if (!version || this._isStopped || this._cliVersion) return;
this._cliVersion = version;
this.emit('cliInfoUpdated', {
version: this._cliVersion,
model: this._cliModel,
accountType: this._cliAccountType,
latestVersion: this._cliLatestVersion,
});
})
.catch(() => {
/* best-effort */
});
}, DOCKER_CLI_VERSION_PROBE_DELAY_MS);
}
// If mux wrapping is enabled, create or attach to a mux session
if (this._useMux && this._mux) {
try {
@@ -1085,9 +1388,13 @@ export class Session extends EventEmitter {
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
},
createSessionOptions: {
sessionId: this.id,
@@ -1100,9 +1407,13 @@ export class Session extends EventEmitter {
allowedTools: this._allowedTools,
openCodeConfig: this._openCodeConfig,
codexConfig: this._codexConfig,
geminiConfig: this._geminiConfig,
resumeSessionId: this._resumeSessionId,
envOverrides: this._envOverrides,
effort: this._effort,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
},
spawnErrLabel: 'mux attachment',
});
@@ -1173,6 +1484,10 @@ export class Session extends EventEmitter {
if (this.mode === 'codex') {
throw new Error('Codex sessions require tmux. Direct PTY fallback is not supported.');
}
// Gemini sessions require tmux for Gemini/Google auth env injection via setenv
if (this.mode === 'gemini') {
throw new Error('Gemini sessions require tmux. Direct PTY fallback is not supported.');
}
try {
// Pass --session-id to use the SAME ID as the Codeman session
// This ensures subagents can be directly matched to the correct tab
@@ -1250,6 +1565,7 @@ export class Session extends EventEmitter {
this._isWorking = true;
this._status = 'busy';
this.emit('working');
this._autoOps.notifyWorking();
}
this._awaitingIdleConfirmation = false;
if (this.activityTimeout) clearTimeout(this.activityTimeout);
@@ -1295,6 +1611,9 @@ export class Session extends EventEmitter {
this.ptyProcess.onExit(({ exitCode }) => {
console.log('[Session] Interactive PTY exited with code:', exitCode);
// COD-118: record the exit in the circuit breaker BEFORE status bookkeeping.
// A clean (0) exit resets the counter; rapid non-zero repeats trip it.
const breakerResult = this._ptyExitBreaker.recordExit(exitCode, Date.now());
this.ptyProcess = null;
this._pid = null;
this._status = 'idle';
@@ -1322,10 +1641,38 @@ export class Session extends EventEmitter {
if (this._muxSession && this._mux) {
this._mux.setAttached(this.id, false);
}
// COD-118: if the breaker tripped, surface an error state and block the NEXT
// respawn so recovery/reconnect callers stop looping. Still emit 'exit' below
// for normal cleanup. Cleared by an explicit user restart (resetRespawnBreaker()).
if (breakerResult.tripped && !this._respawnBlocked) {
this._respawnBlocked = true;
this._status = 'error';
console.error(
`[Session] PTY exit circuit breaker tripped for ${this.id} (${breakerResult.count} non-zero exits within window); blocking respawn.`
);
this.emit('respawnBreakerTripped', { count: breakerResult.count });
}
this.emit('exit', exitCode);
});
}
/**
* Clear the interactive-PTY exit circuit breaker (COD-118).
*
* Called on an EXPLICIT, user-initiated (re)start so an intentional restart is
* never blocked by a prior crash-loop trip. Automatic recovery/reconnect paths
* must NOT call this — that's the whole point of the breaker.
*/
resetRespawnBreaker(): void {
this._ptyExitBreaker.reset();
this._respawnBlocked = false;
}
/** Whether the interactive-PTY exit circuit breaker is currently tripped (COD-118). */
get respawnBlocked(): boolean {
return this._respawnBlocked;
}
/**
* Process expensive parsers (ANSI strip, Ralph, bash tool, token, CLI info, task descriptions).
* Called on a throttled schedule (every EXPENSIVE_PROCESS_INTERVAL_MS) instead of on every
@@ -1356,6 +1703,11 @@ export class Session extends EventEmitter {
this._bashToolParser.processCleanData(getCleanData());
}
// Usage-limit pause detection (auto-resume on usage limit)
if (this._autoOps.autoResumeEnabled) {
this._autoOps.processCleanData(getCleanData());
}
// Parse token count from status line (e.g., "123.4k tokens" or "5234 tokens")
if (rawData.includes('token')) {
this.parseTokensFromStatusLine(getCleanData());
@@ -1384,6 +1736,7 @@ export class Session extends EventEmitter {
this._isWorking = true;
this._status = 'busy';
this.emit('working');
this._autoOps.notifyWorking();
this._awaitingIdleConfirmation = false;
if (this.activityTimeout) clearTimeout(this.activityTimeout);
}
@@ -1429,6 +1782,9 @@ export class Session extends EventEmitter {
mode: 'shell',
niceConfig: this._niceConfig,
envOverrides: this._envOverrides,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
},
createSessionOptions: {
sessionId: this.id,
@@ -1437,6 +1793,9 @@ export class Session extends EventEmitter {
name: this._name,
niceConfig: this._niceConfig,
envOverrides: this._envOverrides,
historyLimit: this._tmuxHistoryLimit,
remote: this._remote,
docker: this._docker,
},
spawnErrLabel: 'shell mux attachment',
});
@@ -1665,6 +2024,7 @@ export class Session extends EventEmitter {
this._errorBuffer = '';
this._messages = [];
this._lineBuffer = '';
this._altScreenSeqCarry = '';
this._lastActivityAt = Date.now();
}
@@ -2042,11 +2402,65 @@ export class Session extends EventEmitter {
* ```
*/
write(data: string): void {
this._trackCodexSubmit(data);
if (this.ptyProcess) {
this.ptyProcess.write(data);
}
}
// ── Codex thread tracking ─────────────────────────────────────────────
// When a codex pane last submitted a message (Enter). The response-viewer
// correlates this against ~/.codex/history.jsonl entry timestamps to find
// the thread the pane is ACTUALLY on — the only signal that survives
// /resume, /new and /fork typed inside the codex TUI itself.
private _codexLastSubmitAt = 0;
get codexLastSubmitAt(): number {
return this._codexLastSubmitAt;
}
private _trackCodexSubmit(data: string): void {
if (this.mode === 'codex' && (data.includes('\r') || data.includes('\n'))) {
this._codexLastSubmitAt = Date.now();
}
}
/**
* Per-client highest-applied input sequence, for exactly-once input delivery.
* Keyed by the web client's stable `clientId`. Bounded so many devices over a
* long-lived session can't grow it without limit (insertion order = MRU, so
* eviction drops the least-recently-active client).
*/
private _appliedInputSeq = new Map<string, number>();
private static readonly MAX_INPUT_DEDUP_CLIENTS = 256;
/**
* Decide whether an input frame should be applied to the PTY or skipped as a
* duplicate redelivery. Returns true exactly once per (clientId, seq): the
* first time a seq strictly greater than the client's last-applied is seen.
* A redelivery of an already-applied seq (the client never got our ACK and
* resent) returns false. Callers should ACK regardless — a duplicate is, from
* the client's view, "delivered" — and only `write()` the PTY when this is
* true. Relies on the client delivering one client's frames in seq order over
* a single ordered stream, so `seq <= last` ⇒ already applied.
*
* Without this, the client's at-least-once redelivery (needed because a
* half-open socket silently drops frames with no error) would type a prompt
* twice whenever an ACK is lost after the write landed.
*/
shouldApplyInput(clientId: string, seq: number): boolean {
const last = this._appliedInputSeq.get(clientId);
if (last !== undefined && seq <= last) return false;
// Re-insert to move this client to the MRU end for fair eviction.
if (last !== undefined) this._appliedInputSeq.delete(clientId);
this._appliedInputSeq.set(clientId, seq);
if (this._appliedInputSeq.size > Session.MAX_INPUT_DEDUP_CLIENTS) {
const oldest = this._appliedInputSeq.keys().next().value;
if (oldest !== undefined) this._appliedInputSeq.delete(oldest);
}
return true;
}
/**
* Sends input via the terminal multiplexer's direct input mechanism.
*
@@ -2064,6 +2478,7 @@ export class Session extends EventEmitter {
* ```
*/
async writeViaMux(data: string): Promise<boolean> {
this._trackCodexSubmit(data);
if (this._mux && this._muxSession) {
return this._mux.sendInput(this.id, data);
}
@@ -2096,9 +2511,29 @@ export class Session extends EventEmitter {
*/
private _desktopSizeClaims = new Set<symbol>();
/**
* A desktop sizing claim only blocks small-viewport resizes while the
* desktop is RECENTLY ACTIVE (claim registration or typed input within this
* window). An abandoned-but-connected desktop tab (left open at home, screen
* locked) must not hold a phone's view hostage: without this, the phone
* renders a desktop-width stream in a narrow xterm — mid-word wraps, tmux
* dot-fill, and Ink overdraw soup (the 0.9.8–0.9.12 mobile regression).
*/
private static readonly DESKTOP_CLAIM_IDLE_MS = 90_000;
/** Last evidence of a live desktop user (claim registered / typed input). */
private _lastDesktopActivityAt = 0;
/** Last desktop-typed dimensions, for re-asserting after a mobile override. */
private _lastDesktopDims: { cols: number; rows: number } | null = null;
/** True while a small viewport reflowed the pane past an idle desktop claim. */
private _mobileSizeOverride = false;
/** Register a live desktop sizing claim (see _desktopSizeClaims). */
claimDesktopSizing(token: symbol): void {
this._desktopSizeClaims.add(token);
this._lastDesktopActivityAt = Date.now();
}
/** Release a desktop sizing claim when its connection goes away. */
@@ -2106,25 +2541,50 @@ export class Session extends EventEmitter {
this._desktopSizeClaims.delete(token);
}
/**
* Record desktop user activity (typed input over a claim-holding socket).
* If a phone reflowed the pane while the desktop was idle, the desktop
* layout is restored — "whoever is actively using the session wins".
*/
noteDesktopActivity(): void {
this._lastDesktopActivityAt = Date.now();
if (this._mobileSizeOverride && this._lastDesktopDims) {
this._mobileSizeOverride = false;
this.resize(this._lastDesktopDims.cols, this._lastDesktopDims.rows, { viewportType: 'desktop' });
}
}
/**
* Resizes the PTY terminal dimensions.
* Skips the resize if dimensions haven't changed to avoid triggering
* unnecessary Ink full-screen redraws (visible flicker on tab switch).
*
* Arbitration: while a desktop connection holds a sizing claim, resizes from
* small viewports (mobile/tablet) are ignored entirely — shrink AND grow
* would both reflow the desktop view. Without a desktop connected, small
* viewports control the PTY size freely.
* Arbitration: while a desktop connection holds a sizing claim AND has been
* active within DESKTOP_CLAIM_IDLE_MS, resizes from small viewports
* (mobile/tablet) are ignored — shrink AND grow would both reflow the
* desktop view. Once the desktop goes idle, a phone may take the pane (the
* desktop re-asserts its size on its next typed input via
* noteDesktopActivity). Without a desktop connected, small viewports
* control the PTY size freely.
*
* @param cols - Number of columns (width in characters)
* @param rows - Number of rows (height in lines)
*/
resize(cols: number, rows: number, options: { viewportType?: ResizeViewportType } = {}): void {
resize(cols: number, rows: number, options: { viewportType?: ResizeViewportType; force?: boolean } = {}): void {
const isSmallViewport = options.viewportType === 'mobile' || options.viewportType === 'tablet';
if (isSmallViewport && this._desktopSizeClaims.size > 0) {
return;
if (options.viewportType === 'desktop') {
this._lastDesktopDims = { cols, rows };
this._lastDesktopActivityAt = Date.now();
this._mobileSizeOverride = false;
}
if (this.ptyProcess && (cols !== this._ptyCols || rows !== this._ptyRows)) {
if (isSmallViewport && this._desktopSizeClaims.size > 0) {
if (Date.now() - this._lastDesktopActivityAt < Session.DESKTOP_CLAIM_IDLE_MS) {
return;
}
this._mobileSizeOverride = true;
}
const dimsChanged = cols !== this._ptyCols || rows !== this._ptyRows;
if (this.ptyProcess && (dimsChanged || options.force)) {
this._ptyCols = cols;
this._ptyRows = rows;
if (this._mux && this._muxSession) {
+51
View File
@@ -272,6 +272,12 @@ export class StateStore {
if (this.state.tokenStats) {
parts.push(`"tokenStats":${JSON.stringify(this.state.tokenStats)}`);
}
if (this.state.cronJobs) {
parts.push(`"cronJobs":${JSON.stringify(this.state.cronJobs)}`);
}
if (this.state.cronJobRuns) {
parts.push(`"cronJobRuns":${JSON.stringify(this.state.cronJobRuns)}`);
}
return `{${parts.join(',')}}`;
}
@@ -568,6 +574,51 @@ export class StateStore {
this.save();
}
// ========== Cron Job Methods ==========
/** Returns all scheduled jobs keyed by job ID. */
getCronJobs(): Record<string, import('./types/cron.js').CronJob> {
if (!this.state.cronJobs) this.state.cronJobs = {};
return this.state.cronJobs;
}
/** Returns a scheduled job by ID, or null if not found. */
getCronJob(id: string): import('./types/cron.js').CronJob | null {
return this.state.cronJobs?.[id] ?? null;
}
/** Sets a scheduled job and triggers a debounced save. */
setCronJob(id: string, job: import('./types/cron.js').CronJob): void {
if (!this.state.cronJobs) this.state.cronJobs = {};
this.state.cronJobs[id] = job;
this.save();
}
/** Removes a scheduled job and triggers a debounced save. */
removeCronJob(id: string): void {
if (this.state.cronJobs) delete this.state.cronJobs[id];
this.save();
}
/** Returns all scheduled job runs keyed by run ID. */
getCronJobRuns(): Record<string, import('./types/cron.js').CronJobRun> {
if (!this.state.cronJobRuns) this.state.cronJobRuns = {};
return this.state.cronJobRuns;
}
/** Sets a scheduled job run (history record) and triggers a debounced save. */
setCronJobRun(id: string, run: import('./types/cron.js').CronJobRun): void {
if (!this.state.cronJobRuns) this.state.cronJobRuns = {};
this.state.cronJobRuns[id] = run;
this.save();
}
/** Removes a scheduled job run and triggers a debounced save. */
removeCronJobRun(id: string): void {
if (this.state.cronJobRuns) delete this.state.cronJobRuns[id];
this.save();
}
/** Returns the application configuration. */
getConfig() {
return this.state.config;
+152 -5
View File
@@ -13,6 +13,8 @@
* - `SubagentEvents` — typed event map
*
* Watched patterns: `~/.claude/projects/{project}/{session}/subagents/agent-{id}.jsonl`
* plus the `agent-{id}.meta.json` discovery sidecar (2026-06 format) and nested
* workflow agents under `subagents/workflows/{workflowId}/agent-{id}.jsonl`.
* Parses JSONL entries: user/assistant messages, tool_use/tool_result blocks, progress events.
* Tracks per-agent: status, token counts, model, description, tool call count, liveness (PID).
*
@@ -30,7 +32,7 @@ import { watch, existsSync, FSWatcher } from 'node:fs';
import { createReadStream } from 'node:fs';
import { createInterface } from 'node:readline';
import { homedir } from 'node:os';
import { join, basename } from 'node:path';
import { join, basename, dirname } from 'node:path';
import { execFile } from 'node:child_process';
import { readFile, readdir, stat as statAsync } from 'node:fs/promises';
import { PENDING_TOOL_CALL_TTL_MS, MAX_PENDING_TOOL_CALLS, MAX_TRACKED_AGENTS } from './config/map-limits.js';
@@ -1150,6 +1152,10 @@ export class SubagentWatcher extends EventEmitter {
try {
await statAsync(subagentDir);
await this.watchSubagentDir(subagentDir, project, session);
// Workflow agents nest one level deeper under subagents/workflows/{wf}/
// (each holds its own agent-{id}.jsonl/.meta.json) — the flat watcher
// above never sees them, so discover and watch each workflow dir too.
await this.watchWorkflowDirs(subagentDir, project, session);
} catch {
// subagent dir doesn't exist - skip
}
@@ -1179,8 +1185,12 @@ export class SubagentWatcher extends EventEmitter {
try {
const files = await readdir(dir);
for (const file of files) {
if (file.endsWith('.jsonl')) {
if (file.startsWith('agent-') && file.endsWith('.jsonl')) {
await this.registerAgentFile(join(dir, file), projectHash, sessionId, true);
} else if (file.startsWith('agent-') && file.endsWith('.meta.json')) {
// Claude Code (2026-06) writes a `agent-{id}.meta.json` sidecar for TUI
// Task subagents and no longer always writes a per-agent `.jsonl` here.
await this.registerAgentMeta(join(dir, file), projectHash, sessionId, true);
}
}
} catch {
@@ -1190,13 +1200,26 @@ export class SubagentWatcher extends EventEmitter {
// Single directory watcher handles both new files and file content changes
try {
const watcher = watch(dir, (_eventType, filename) => {
if (!filename?.endsWith('.jsonl')) return;
const filePath = join(dir, filename);
// Only agent-{id}.jsonl / agent-{id}.meta.json — ignore siblings like a
// workflow dir's journal.jsonl (would otherwise register a bogus "journal" agent).
const isAgent = !!filename && filename.startsWith('agent-');
const isJsonl = isAgent && filename.endsWith('.jsonl');
const isMeta = isAgent && filename.endsWith('.meta.json');
if (!isJsonl && !isMeta) return;
const filePath = join(dir, filename as string);
// Debounce 100ms to batch rapid writes
this.fileDeb.schedule(filePath, () => {
if (!existsSync(filePath)) return;
if (isMeta) {
// Meta sidecar — discovery only (not a transcript; never tail it).
if (!this.fileAgentContext.has(filePath)) {
this.registerAgentMeta(filePath, projectHash, sessionId).catch(() => {});
}
return;
}
if (this.fileAgentContext.has(filePath)) {
// Known file — handle content change
this.handleFileChange(filePath).catch(() => {}); // Ignore - errors logged internally, don't crash watcher callback
@@ -1225,6 +1248,40 @@ export class SubagentWatcher extends EventEmitter {
}
}
/**
* Discover and watch nested workflow agent directories.
*
* The Workflow tool runs its subagents under
* `subagents/workflows/{workflowId}/agent-{id}.jsonl` (+ `.meta.json`, alongside
* a `journal.jsonl` of orchestration events). The flat `subagents/` watcher does
* not recurse, and Node's `fs.watch({ recursive: true })` is unsupported on Linux,
* so each workflow dir gets its own watcher here. Idempotent via `knownSubagentDirs`
* and re-driven by the periodic scan, so newly created workflows are picked up
* within one scan cycle (~5s) — the same latency as a new session's `subagents/`.
*/
private async watchWorkflowDirs(subagentDir: string, projectHash: string, sessionId: string): Promise<void> {
const workflowsRoot = join(subagentDir, 'workflows');
let names: string[];
try {
names = await readdir(workflowsRoot);
} catch {
return; // no workflows for this session
}
for (const name of names) {
// Workflow ids are directories (e.g. `wf_<id>`); skip any stray files that
// share the root (a workflow id never carries a file extension).
if (name.endsWith('.jsonl') || name.endsWith('.json')) continue;
const wfDir = join(workflowsRoot, name);
try {
const st = await statAsync(wfDir);
if (!st.isDirectory()) continue;
await this.watchSubagentDir(wfDir, projectHash, sessionId);
} catch {
// workflow dir vanished mid-scan — skip
}
}
}
/**
* Handle a file content change for an already-registered agent file.
* Tails from last known position, updates info, retries description if missing.
@@ -1291,6 +1348,16 @@ export class SubagentWatcher extends EventEmitter {
const agentId = basename(filePath).replace('agent-', '').replace('.jsonl', '');
// Meta→transcript upgrade: the agent may already be registered from its
// `.meta.json` sidecar (discovery-only — nothing to tail). Now that the real
// `.jsonl` transcript has appeared, re-point to it and drop the stale sidecar
// context, emitting `updated` below rather than a duplicate `discovered`.
const priorEntry = this.agentInfo.get(agentId);
const isMetaUpgrade = !!priorEntry && priorEntry.filePath.endsWith('.meta.json');
if (isMetaUpgrade && priorEntry) {
this.fileAgentContext.delete(priorEntry.filePath);
}
// Initial info - handle race condition where file may be deleted between discovery and stat
let fileStat;
try {
@@ -1341,7 +1408,7 @@ export class SubagentWatcher extends EventEmitter {
// Track file context for directory watcher change handling
this.fileAgentContext.set(filePath, { projectHash, sessionId });
this.agentInfo.set(agentId, info);
this.emit('subagent:discovered', info);
this.emit(isMetaUpgrade ? 'subagent:updated' : 'subagent:discovered', info);
// Read existing content
this.tailFile(filePath, agentId, sessionId, 0)
@@ -1356,6 +1423,86 @@ export class SubagentWatcher extends EventEmitter {
this.resetIdleTimer(agentId);
}
/**
* Register a subagent discovered via its `agent-{id}.meta.json` sidecar.
*
* As of the 2026-06 Claude Code format change, TUI Task subagents write the
* `agent-{id}.meta.json` sidecar (`{ agentType, description, toolUseId }`) at
* spawn, a beat *before* the `agent-{id}.jsonl` transcript appears in the same
* dir. The legacy `.jsonl`-only discovery saw nothing in that window ("0
* tracked"); this surfaces the agent from the sidecar immediately. The
* transcript then lands within ~1s at the standard path and grows incrementally
* (empirically verified — it is fully tailable, NOT a dead end), so:
* - if the `.jsonl` already exists, defer to registerAgentFile (richer); else
* - register meta-only now, and when the sibling `.jsonl` arrives the dir
* watcher routes it to registerAgentFile, which detects the prior meta-only
* entry and *upgrades* it in place (re-points filePath, starts tailing).
*
* Edge case: if the transcript never materializes (e.g. an agent that dies
* before writing one), the agent stays meta-only — no live feed, status ages
* out via the idle timer / stale cleanup.
*/
private async registerAgentMeta(
metaPath: string,
projectHash: string,
sessionId: string,
isInitialScan: boolean = false
): Promise<void> {
if (this.fileAgentContext.has(metaPath)) return;
const agentId = basename(metaPath).replace('agent-', '').replace('.meta.json', '');
if (this.agentInfo.has(agentId)) return;
// Prefer a real transcript if one was written alongside the sidecar.
const jsonlPath = join(dirname(metaPath), `agent-${agentId}.jsonl`);
if (existsSync(jsonlPath)) {
await this.registerAgentFile(jsonlPath, projectHash, sessionId, isInitialScan);
return;
}
let fileStat;
try {
fileStat = await statAsync(metaPath);
} catch {
return; // deleted between discovery and stat
}
if (isInitialScan && Date.now() - fileStat.mtime.getTime() > STARTUP_MAX_FILE_AGE_MS) {
return; // skip stale historical agents on startup
}
let description: string | undefined;
try {
const meta = JSON.parse(await readFile(metaPath, 'utf8')) as { agentType?: string; description?: string };
description = meta.description || meta.agentType;
} catch {
return; // unreadable / not yet fully written — a later watch event retries
}
if (this.isInternalAgent(description)) return;
const info: SubagentInfo = {
agentId,
sessionId,
projectHash,
filePath: metaPath,
startedAt: fileStat.birthtime.toISOString(),
lastActivityAt: fileStat.mtime.getTime(),
status: 'active',
toolCallCount: 0,
entryCount: 0,
fileSize: fileStat.size,
description,
};
if (this.agentInfo.size >= MAX_TRACKED_AGENTS) {
const oldestId = this.findOldestInactiveAgent();
if (oldestId) this.removeAgent(oldestId);
}
this.fileAgentContext.set(metaPath, { projectHash, sessionId });
this.agentInfo.set(agentId, info);
this.emit('subagent:discovered', info);
this.resetIdleTimer(agentId);
}
/**
* Tail a file from a specific position
*/
+47 -450
View File
@@ -1,461 +1,58 @@
# CLAUDE.md - Project Configuration
# CLAUDE.md
## Setup
Copy these files to your new project:
- `CLAUDE.md` → project root
- `.claude/settings.json` → `.claude/settings.json`
<!--
Generated by Codeman on [DATE]. This file is loaded into context at the
start of every Claude Code session in this project.
Then update the Project Overview section below.
Keep it short (target: under 200 lines). For each line ask: "would removing
this cause Claude to make mistakes?" If not, cut it. Don't document what
Claude can infer from the code itself (file layout, standard conventions,
APIs) — bloat causes Claude to ignore the rules that matter.
---
HTML comments like this one are stripped before loading, so fill-in notes
cost no context. If this file grows too big, split into path-scoped rules
in .claude/rules/*.md or import other files with @path/to/file syntax.
-->
This file guides Claude Code when working in this repository.
## Project
## Project Overview
<!-- Update this section with project-specific details -->
- **Project Name**: [PROJECT_NAME]
- **Description**: [PROJECT_DESCRIPTION]
- **Tech Stack**: [TECHNOLOGIES_USED]
- **Last Updated**: [DATE]
---
## Commands
<!-- List the exact commands Claude can't guess — fill in as the project
takes shape, then delete this comment:
| Task | Command |
|------|---------|
| Dev server | `npm run dev` |
| Test (single file) | `npm test -- test/<file>.test.ts` |
| Lint | `npm run lint` |
| Build | `npm run build` |
-->
## Code Style
<!-- Only rules that differ from language/framework defaults, one line each:
- Use 2-space indentation
- ES modules only — never require()
-->
## Workflow
- Full permissions are granted: read, write, edit, and execute without asking.
- Commit after every meaningful change; never batch unrelated work.
- Use conventional commits (`feat:` `fix:` `docs:` `refactor:` `test:` `chore:`); the message says what changed and why.
- Run the tests and linter before declaring any task done.
- Keep README and docs in sync with code changes.
## Codeman Environment
This session is managed by **Codeman** and runs within a tmux session.
This session is managed by Codeman and runs inside tmux (`CODEMAN_MUX=1` confirms it).
**Important**: Check for `CODEMAN_MUX=1` environment variable to confirm.
- Do NOT attempt to kill your own tmux session
- The session persists across disconnects - your work is safe
- Token usage, costs, and background tasks are tracked externally
---
## Work Principles
### Autonomy
Full permissions granted. Act decisively without asking - read, write, edit, execute freely.
### Git Discipline
- **Commit after every meaningful change** - never batch unrelated work
- Use conventional commits: `feat:`, `fix:`, `docs:`, `refactor:`, `test:`, `chore:`
- Commit message = what changed + why (not how)
### Documentation
- Update README.md when adding features or changing setup
- Update this file's session log after work sessions
- Keep docs in sync with code changes
### Thinking
Extended thinking is enabled. Use deep reasoning for complex architectural decisions, difficult bugs, and multi-file changes.
### Task Tracking (TodoWrite)
**ALWAYS use TodoWrite** to track tasks. This is non-negotiable for anything beyond trivial single-step work.
**When to use TodoWrite:**
- Multi-step tasks (3+ steps)
- Bug fixes requiring investigation
- Feature implementations
- Any work where progress tracking helps
- When the user provides multiple requests
**How to use it:**
1. **Before starting**: Break down the work into discrete todos
2. **During work**: Mark each todo `in_progress` before starting, `completed` when done
3. **One at a time**: Only ONE todo should be `in_progress` at any moment
4. **Immediately**: Mark todos complete the moment they're done - don't batch
**Why this matters:**
- Gives the user visibility into your progress
- Prevents forgetting tasks mid-work
- Creates accountability checkpoints
- Makes complex work manageable
**Example workflow:**
```
User: "Add user authentication with JWT"
→ TodoWrite:
- [ ] Research existing auth patterns in codebase
- [ ] Implement JWT token generation
- [ ] Add login endpoint
- [ ] Add token validation middleware
- [ ] Add protected route example
- [ ] Write tests
→ Mark "Research existing auth patterns" as in_progress
→ Do the research
→ Mark as completed, mark next as in_progress
→ Continue until all done
```
**Anti-patterns to avoid:**
- Starting work without creating todos first
- Having multiple todos `in_progress` simultaneously
- Batching completions at the end
- Skipping TodoWrite for "simple" multi-step tasks
---
## When to Use Agents
**Explore agent**: Codebase investigation, finding files, understanding architecture
```
"Use explore agent to find all authentication-related code"
```
**Parallel agents**: Independent tasks that don't conflict
```
"Research auth, database, and API modules in parallel using separate agents"
```
**Background execution**: Long-running operations (tests, builds)
```
"Run the test suite in the background while I continue"
```
**Sequential chaining**: When second task depends on first
```
"Use code-reviewer to find issues, then use fixer to resolve them"
```
---
## Planning Mode (Automatic)
**Automatically enter planning mode** when ANY of these conditions apply:
- Multi-file changes (3+ files affected)
- Architectural decisions
- Unclear or evolving requirements
- Risk mitigation on core systems
- New feature implementation
- Refactoring existing functionality
**Do NOT ask** whether to enter planning mode - just enter it when conditions are met.
Planning mode flow: read-only exploration → create plan → get approval → execute.
**Skip planning mode** only for:
- Single-file bug fixes
- Typo corrections
- Simple config changes
- Tasks with explicit step-by-step instructions from user
---
## Ralph Wiggum Loop (Autonomous Work Mode)
Ralph loops enable persistent, autonomous work on large tasks. When active, you continue iterating until completion criteria are met or the loop is cancelled.
### Starting a Ralph Loop
- Start: `/ralph-loop:ralph-loop`
- Cancel: `/ralph-loop:cancel-ralph`
- Help: `/ralph-loop:help`
### Time-Aware Loops
When the user specifies a **minimum duration** (e.g., "optimize for 8 hours", "work on this for 2 hours"), the loop becomes time-aware:
**At loop start:**
```bash
# Record start time
date +%s > /tmp/ralph_start_time
echo "Loop started at $(date)"
```
**Check elapsed time periodically:**
```bash
START=$(cat /tmp/ralph_start_time)
NOW=$(date +%s)
ELAPSED_HOURS=$(echo "scale=2; ($NOW - $START) / 3600" | bc)
echo "Elapsed: $ELAPSED_HOURS hours"
```
**Time-aware behavior:**
1. Complete all primary tasks from the user's prompt
2. After primary tasks done, check elapsed time
3. If minimum duration NOT reached:
- **Do NOT output completion phrase**
- Self-generate additional related tasks
- Continue working until minimum time elapsed
4. Only output completion phrase when:
- ALL primary tasks complete AND
- Minimum duration reached (or exceeded)
**Self-generating additional tasks when time remains:**
- Code optimization (performance, readability, DRY)
- Test coverage improvements
- Edge case handling
- Error message improvements
- Documentation gaps
- Security hardening
- Accessibility improvements
- Code cleanup and dead code removal
- Dependency updates
- Type safety improvements
**Example time-aware prompt:**
```
"Optimize the API endpoints for the next 4 hours. Focus on performance first,
then code quality. Minimum runtime: 4 hours."
Completion phrase: <promise>TIME_COMPLETE</promise>
```
**Time-aware loop behavior:**
```
[Start loop, record timestamp]
[Complete primary optimization tasks - 2 hours elapsed]
[Check time: 2/4 hours - NOT done yet]
[Self-generate: "Add caching to database queries"]
[Self-generate: "Optimize N+1 queries"]
[Self-generate: "Add request batching"]
[Continue working... 4.5 hours elapsed]
[Check time: 4.5/4 hours - minimum reached]
[All tasks complete, tests pass]
<promise>TIME_COMPLETE</promise>
```
### How You Know You're in a Ralph Loop
The user started the loop with a prompt containing:
- Clear task requirements
- A **completion phrase** (e.g., `<promise>COMPLETE</promise>`)
- **Optional: minimum duration** (e.g., "for the next 4 hours")
- Iteration limits (handled by the system)
Your job: Keep working until ALL requirements are verifiably done AND minimum time reached (if specified), then output the exact completion phrase.
### Core Behaviors During Ralph Loop
**1. Work Incrementally**
- Complete one sub-task at a time
- Verify it works before moving to the next
- Don't try to do everything in one pass
**2. Commit Frequently**
- Commit after each meaningful completion
- Creates recovery points if something breaks
- Shows progress in git history
```
git add . && git commit -m "feat(auth): add token refresh endpoint"
```
**3. Self-Correct Relentlessly**
```
Loop:
1. Implement/fix
2. Run tests
3. If tests fail → read error, fix, go to 1
4. Run linter
5. If lint errors → fix, go to 1
6. Commit
7. Continue to next task
```
**4. Track Progress**
Update the session log in this file as you complete tasks:
```markdown
| Date | Tasks Completed | Files Changed | Notes |
|------|-----------------|---------------|-------|
| YYYY-MM-DD | Add auth endpoint | auth.ts, routes.ts | Tests passing |
```
**5. Use Git History When Stuck**
If something isn't working:
```bash
git log --oneline -10
git diff HEAD~1
```
See what you already tried. Don't repeat failed approaches.
**6. Completion Phrase = Contract**
Only output the completion phrase (e.g., `<promise>COMPLETE</promise>`) when:
- ALL requirements from the original prompt are done
- ALL tests pass
- ALL linting passes
- Changes are committed
**Never output the completion phrase early.** The loop only ends when you say it's done.
### What Makes Good Completion Criteria
The user should provide criteria that are:
- **Verifiable**: Tests pass, lint clean, build succeeds
- **Measurable**: "5 endpoints", "all files in src/", "zero errors"
- **Binary**: Done or not done, no ambiguity
If the original prompt has vague criteria, ask clarifying questions before starting heavy work.
### Self-Correction Pattern (Include in Your Work)
```
FOR EACH TASK:
1. Implement the change
2. Run tests (npm test, pytest, go test, cargo test, etc.)
- If fail → read error, fix, retry
3. Run linter (npm run lint, ruff, golangci-lint, etc.)
- If fail → fix, go to step 2
4. Verify manually if needed
5. Commit with descriptive message
6. Update session log
7. Move to next task
WHEN ALL TASKS DONE:
1. Run full test suite
2. Run full lint
3. Verify build succeeds
4. Review all changes: git diff main
5. Only then output completion phrase
```
### Example: How to Think During Ralph Loop
**Original prompt**: "Add CRUD endpoints for todos with validation"
**Your approach**:
```
Task breakdown:
- [ ] GET /todos (list)
- [ ] POST /todos (create with validation)
- [ ] GET /todos/:id (single)
- [ ] PUT /todos/:id (update with validation)
- [ ] DELETE /todos/:id
- [ ] Tests for all endpoints
Starting with GET /todos...
[implement]
[test - passes]
[commit: "feat(todos): add GET /todos endpoint"]
[update session log]
Moving to POST /todos...
[implement]
[test - fails: validation not working]
[fix validation]
[test - passes]
[commit: "feat(todos): add POST /todos with validation"]
[update session log]
...continue until all done...
Final verification:
[npm test - all pass]
[npm run lint - clean]
[npm run build - succeeds]
<promise>COMPLETE</promise>
```
### When to NOT Output Completion Phrase
- Tests are failing (even one)
- Lint errors exist
- Build is broken
- You skipped a requirement
- You're unsure if something works
- **Minimum duration not reached** (for time-aware loops)
Instead: Fix the issue, verify, then complete. For time-aware loops: generate more tasks and keep improving until minimum time elapsed.
### RALPH_STATUS Block (Required During Ralph Loop)
At the **END of every response** during a Ralph Loop, output this structured status block:
```
---RALPH_STATUS---
STATUS: IN_PROGRESS | COMPLETE | BLOCKED
TASKS_COMPLETED_THIS_LOOP: <number>
FILES_MODIFIED: <number>
TESTS_STATUS: PASSING | FAILING | NOT_RUN
WORK_TYPE: IMPLEMENTATION | TESTING | DOCUMENTATION | REFACTORING
EXIT_SIGNAL: false | true
RECOMMENDATION: <one line summary of what to do next>
---END_RALPH_STATUS---
```
**Rules:**
- Output this block at the end of **every** response, no exceptions
- Set `EXIT_SIGNAL` to `true` ONLY when ALL tasks are verifiably done
- Set `STATUS` to `BLOCKED` when you need human intervention
- Do NOT continue with busy work when `EXIT_SIGNAL` should be `true`
- Do NOT forget the status block — it is required for loop tracking
### Testing Limits
- **LIMIT testing to ~20% of total effort** per loop
- PRIORITIZE: Implementation > Documentation > Tests
- Only write tests for NEW functionality
- Do NOT refactor existing tests unless broken
- Do NOT run tests repeatedly without implementing new features
### Exit Scenarios (When to Set EXIT_SIGNAL)
| Scenario | STATUS | EXIT_SIGNAL | Action |
|----------|--------|-------------|--------|
| All tasks completed, tests pass | COMPLETE | true | Output completion phrase |
| No work remaining, specs done | COMPLETE | true | Output completion phrase |
| Making normal progress | IN_PROGRESS | false | Continue to next task |
| Test-only loop (no implementation) | IN_PROGRESS | false | Warn and shift to implementation |
| Stuck on same error repeatedly | BLOCKED | false | Describe blocker, request help |
| Needs human decision/intervention | BLOCKED | false | Describe what's needed |
**Anti-patterns to avoid:**
- Setting `EXIT_SIGNAL: true` when tests are failing
- Continuing to work when all tasks are genuinely done (busy work)
- Running the same failing test repeatedly without changing approach
- Adding features not in the original specifications
- Refactoring working code instead of completing assigned tasks
---
## Code Standards
### Before Writing
- Read existing code in the area you're modifying
- Follow existing patterns and conventions
- Check for similar implementations to reference
### During Implementation
- Keep changes focused and minimal
- Don't over-engineer
- Write tests for new functionality
### After Implementation
- Run tests
- Update docs if needed
- Commit with descriptive message
---
## Hooks Awareness
This project may have hooks that auto-format code after writes or validate operations. If a tool call behaves unexpectedly, hooks are likely the cause. Continue working - they're intentional.
---
## Session Log
| Date | Tasks Completed | Files Changed | Notes |
|------|-----------------|---------------|-------|
| [DATE] | Project created | CLAUDE.md | Initial setup |
---
## Current Task Queue
### Active Ralph Loop
**Status**: Not Active
**Completion Phrase**: -
### Pending Tasks
- [ ] <!-- Add tasks here -->
---
## Implementation Plans
<!-- Document plans before major implementations -->
---
## Notes & Decisions
<!-- Track important decisions and context -->
- NEVER kill your own session: no `tmux kill-session`, `pkill tmux`, or `pkill claude`.
- The session persists across disconnects — your work is safe.
- Hooks may auto-format or validate after writes; unexpected tool behavior usually means a hook ran. Keep working.
+8 -9
View File
@@ -15,18 +15,17 @@ import { fileURLToPath } from 'node:url';
const __dirname = dirname(fileURLToPath(import.meta.url));
const BUNDLED_TEMPLATE_PATH = join(__dirname, 'case-template.md');
const MINIMAL_FALLBACK = `# CLAUDE.md - Project Configuration
const MINIMAL_FALLBACK = `# CLAUDE.md
<!-- Generated by Codeman on [DATE]. Add the commands, code style rules, and
workflow notes Claude can't infer from the code. Keep it short. -->
This file guides Claude Code when working in this repository.
## Project
## Project Overview
- **Project Name**: [PROJECT_NAME]
- **Description**: [PROJECT_DESCRIPTION]
- **Last Updated**: [DATE]
## Session Log
| Date | Tasks Completed | Files Changed | Notes |
|------|-----------------|---------------|-------|
| [DATE] | Project created | CLAUDE.md | Initial setup |
`;
/**
+691 -29
View File
@@ -29,7 +29,8 @@ const execAsync = promisify(exec);
import { existsSync, readFileSync, mkdirSync } from 'node:fs';
import { writeFile, rename } from 'node:fs/promises';
import { dirname } from 'node:path';
import { dataPath, DEFAULT_TMUX_SOCKET } from './config/instance.js';
import { homedir } from 'node:os';
import { dataPath, DEFAULT_TMUX_SOCKET, CODEMAN_INSTANCE } from './config/instance.js';
import {
ProcessStats,
PersistedRespawnConfig,
@@ -41,15 +42,41 @@ import {
type OpenCodeConfig,
type CodexConfig,
type EffortLevel,
type GeminiConfig,
type SessionRemote,
type SessionDocker,
type DockerCommandMode,
} from './types.js';
import { buildEffortCliArgs } from './session-cli-builder.js';
import { wrapWithNice, SAFE_PATH_PATTERN, findClaudeDir, resolveOpenCodeDir, resolveCodexDir } from './utils/index.js';
import { buildSshConnectionArgs, defaultRemoteCommandForMode, remoteSshTarget } from './remote-hosts.js';
import {
buildDockerBaseArgs,
buildDockerCreateArgs,
containerApiUrl,
CONTAINER_HOME,
defaultDockerCommandForMode,
hostGatewayAlias,
resolveDockerClaudeArtifacts,
resolveDockerCredentialArtifacts,
type DockerCreateContext,
type DockerMount,
type DockerSeedCopy,
} from './docker-hosts.js';
import {
wrapWithNice,
SAFE_PATH_PATTERN,
findClaudeDir,
resolveOpenCodeDir,
resolveCodexDir,
resolveGeminiDir,
} from './utils/index.js';
import type {
TerminalMultiplexer,
MuxSession,
MuxSessionWithStats,
CreateSessionOptions,
RespawnPaneOptions,
PaneCaptureOptions,
} from './mux-interface.js';
// ============================================================================
@@ -57,6 +84,15 @@ import type {
// ============================================================================
import { EXEC_TIMEOUT_MS } from './config/exec-timeout.js';
import { DEFAULT_TMUX_HISTORY_LIMIT, DEFAULT_TERMINAL_BUFFER_MAX_BYTES } from './config/terminal-history.js';
/**
* Extra stdout headroom for the full-history `capture-pane` child process on
* top of the consumer's byte cap: raw scrollback carries per-line SGR/ANSI
* overhead that the route pipeline strips before applying its cap, so the
* capture must be allowed to exceed the final payload size.
*/
const FULL_HISTORY_CAPTURE_SLACK_BYTES = 8 * 1024 * 1024;
/** Delay after tmux session creation — enough for detached tmux to be queryable */
const TMUX_CREATION_WAIT_MS = 100;
@@ -421,12 +457,37 @@ function truncatePaneLineByVisibleColumns(line: string, maxColumns: number): str
return result;
}
/**
* Normalize scrollback line endings to `\r\n` so a fresh xterm replays each line
* at column 0 (COD-138).
*
* `capture-pane -p -e -S -` (full-history capture) joins scrollback rows with a
* BARE `\n`. The browser xterm is created with the default `convertEol: false`
* (correct for the live PTY stream, which already carries real `\r\n`), so a bare
* `\n` drops a row without returning the cursor to column 0. Replaying that raw
* buffer on a full page reload makes every line start one column further right —
* the diagonal "staircase". The visible/tab-switch path avoids this by repainting
* each row with an absolute cursor CSI (`formatPaneSnapshot`); the full-history
* path returns raw scrollback, so it must be CRLF-normalized here.
*
* `\r?\n → \r\n` is idempotent on already-CRLF input and leaves a lone `\r` (an
* intentional in-line column reset / overwrite) untouched.
*/
export function normalizeScrollbackEol(buffer: string): string {
return buffer.replace(/\r?\n/g, '\r\n');
}
export function formatPaneSnapshot(
lines: string[],
geometry: { cols: number; rows: number; cursorX: number; cursorY: number }
): string {
const cols = Math.max(1, geometry.cols);
const paintCols = Math.max(1, cols - 1);
// Paint the full pane width. Earlier this dropped the rightmost column
// (cols - 1) out of caution about last-column autowrap, but every painted
// row is immediately followed by an absolute cursor-position CSI (the next
// row's `\x1b[r;1H`, or the final cursor move), which cancels xterm's
// pending-wrap state before any further glyph — so the last column is safe.
const paintCols = cols;
const rows = Math.max(1, geometry.rows);
const parts: string[] = [];
for (let row = 0; row < Math.min(lines.length, rows); row++) {
@@ -566,6 +627,35 @@ export function buildCodexCommand(config?: CodexConfig): string {
return parts.join(' ');
}
/**
* Build the Gemini CLI command with appropriate flags.
*
* `--skip-trust` avoids a first-run workspace trust prompt inside Codeman.
* Approval mode defaults to `yolo` for parity with Codeman's Claude default
* of `--dangerously-skip-permissions`; users can override it later through
* Gemini config once Codeman exposes richer Gemini settings.
*/
function buildGeminiCommand(config?: GeminiConfig): string {
const parts = ['gemini', '--skip-trust'];
const approvalMode = config?.approvalMode || 'yolo';
if (['default', 'auto_edit', 'yolo', 'plan'].includes(approvalMode)) {
parts.push('--approval-mode', approvalMode);
}
if (config?.model) {
const safeModel = /^[a-zA-Z0-9._\-/]+$/.test(config.model) ? config.model : undefined;
if (safeModel) parts.push('--model', safeModel);
}
if (config?.resumeSession) {
const safeId = /^[a-zA-Z0-9._-]+$/.test(config.resumeSession) ? config.resumeSession : undefined;
if (safeId) parts.push('--resume', safeId);
}
return parts.join(' ');
}
/**
* Build the spawn command for any session mode.
* Shared by createSession() and respawnPane() to avoid duplication.
@@ -592,6 +682,7 @@ function buildSpawnCommand(options: {
allowedTools?: string;
openCodeConfig?: OpenCodeConfig;
codexConfig?: CodexConfig;
geminiConfig?: GeminiConfig;
resumeSessionId?: string;
effort?: EffortLevel;
}): string {
@@ -619,9 +710,385 @@ function buildSpawnCommand(options: {
if (options.mode === 'codex') {
return buildCodexCommand(options.codexConfig);
}
if (options.mode === 'gemini') {
return buildGeminiCommand(options.geminiConfig);
}
return '$SHELL';
}
/**
* Dedicated socket for Codeman-launched REMOTE tmux servers, distinct from the
* canonical local `-L codeman` socket. A remote host that runs its OWN Codeman
* would otherwise share the `-L codeman` socket AND the `codeman-<hex>` discovery
* name, so its `reconcileSessions()` would ADOPT our session (attach a PTY,
* resize, respawn-pane it locally) — the cross-machine form of the "2nd instance
* attaches live sessions" hazard. A private socket keeps our remote sessions off
* that instance's radar entirely.
*/
const REMOTE_TMUX_SOCKET = 'codeman-remote';
/**
* Deterministic, reattach-stable remote tmux session name for a Codeman session.
*
* Derived from the same stable field the LOCAL muxName uses (the first 8 chars of
* the sessionId), so reconnecting (which re-issues the exact same
* `ssh … new-session -A`) lands back in the SAME remote session. Must NOT be
* random/time-based — it has to be stable across reconnects.
*
* The `codeman-ssh-` prefix is deliberately chosen to FAIL a remote Codeman's
* `SAFE_MUX_NAME_PATTERN` (`^codeman-[a-f0-9-]+$`) — the `s`/`h` letters mean a
* remote instance's discovery never treats this as one of its own sessions (belt
* to the dedicated-socket suspenders above).
*/
export function remoteTmuxSessionName(sessionId: string): string {
return `codeman-ssh-${sessionId.slice(0, 8)}`;
}
/**
* COD-104 — build the SSH command that launches (or reattaches) a remote
* session INSIDE a tmux server on the remote host, so the remote agent survives
* an SSH drop.
*
* Emits:
* ssh -o BatchMode=yes -t [<COD-107 connection opts>] user@host \
* 'tmux -L codeman-remote new-session -A -s codeman-ssh-<id> -c <path> "cd <path> && exec <cli>" \
* \; set -t codeman-ssh-<id> status off \; set -t codeman-ssh-<id> mouse off \
* \; set -t codeman-ssh-<id> prefix C-q \; set -s escape-time 0'
*
* COD-107 — the connection options (`-p`, `-i`, `-J`, SOCKS `-o ProxyCommand`,
* arbitrary `-o`) come from the shared `buildSshConnectionArgs(remote)`, so the
* prereq tmux probe and this launch connect with identical options.
*
* - `new-session -A -s codeman-ssh-<id>` = attach-if-exists-else-create
* (idempotent), so reconnect re-runs the same command and reattaches the
* still-running agent.
* - `-L codeman-remote` = a DEDICATED socket, NOT the canonical `-L codeman` a
* remote Codeman would use, so our session never collides with / gets adopted by
* an instance running on the remote host.
* - The `set` options are scoped per-session (`set -t <name>` / server-level
* `set -s`), never `-g`, so they never mutate other sessions' prefix/mouse.
* - The whole tmux invocation is a SINGLE ssh argument (the remote login shell
* runs it), so it is shell-quoted as one unit; the `cd && exec` command is in
* turn a single tmux argument (tmux runs it via `/bin/sh -c`), so the path is
* shell-quoted inside it too. This keeps escaping correct through every layer
* even when the remote path contains spaces.
*/
export function buildRemoteLaunchCommand(options: {
mode: SessionMode;
remote: SessionRemote;
sessionId: string;
}): string {
const { mode, remote, sessionId } = options;
const modeCommand = remote.commands?.[mode] || defaultRemoteCommandForMode(mode);
const remoteName = remoteTmuxSessionName(sessionId);
// Innermost: the command tmux runs in the new pane. Run via `/bin/sh -c` by
// tmux, so the path needs shell-quoting here. `exec` replaces the shell with
// the CLI so the pane PID is the agent itself.
const paneCommand = `cd ${shellescape(remote.remotePath)} && ${modeCommand}`;
// The tmux command line, with `\;` separating commands so the config `set`s
// apply on the SAME connection (and are idempotent on reattach). Options are
// scoped per-session (`set -t <name>` / server `set -s`), NEVER `-g`, so a
// shared remote tmux server's other sessions keep their own prefix/mouse.
const tmuxInvocation = [
`tmux -L ${REMOTE_TMUX_SOCKET} new-session -A -s ${remoteName} -c ${shellescape(remote.remotePath)} ${shellescape(paneCommand)}`,
`set -t ${remoteName} status off`,
`set -t ${remoteName} mouse off`,
`set -t ${remoteName} prefix C-q`,
'set -s escape-time 0',
].join(' \\; ');
// ssh runs its trailing args through the remote login shell, so the entire
// tmux invocation is passed as one shell-quoted argument.
//
// COD-107 — connection options (port, identity, SOCKS ProxyCommand, jump host,
// arbitrary -o) come from the shared `buildSshConnectionArgs` so the launch and
// the tmux-prereq probe connect IDENTICALLY. `-t` is inserted right after
// `ssh -o BatchMode=yes` (preserving the historical token order), then the rest
// of the connection args, then the target and the quoted tmux invocation.
const [ssh, batchMode, ...connectionArgs] = buildSshConnectionArgs(remote);
const sshParts = [ssh, batchMode, '-t', ...connectionArgs, remoteSshTarget(remote), shellescape(tmuxInvocation)];
return sshParts.join(' ');
}
/**
* Build the SSH command that kills the durable remote tmux session created by
* `buildRemoteLaunchCommand`. Because that session lives on a private socket
* (`-L codeman-remote`) under a stable name, killing the LOCAL ssh wrapper alone
* would orphan the remote agent forever (invisible to Codeman, still burning plan
* quota). This is fired best-effort on session kill; the shared connection args
* carry the default `-o ConnectTimeout=10` so an unreachable host fails fast.
*/
export function buildRemoteKillCommand(options: { remote: SessionRemote; sessionId: string }): string {
const { remote, sessionId } = options;
const remoteName = remoteTmuxSessionName(sessionId);
const killCmd = `tmux -L ${REMOTE_TMUX_SOCKET} kill-session -t ${shellescape(remoteName)}`;
const [ssh, ...connectionArgs] = buildSshConnectionArgs(remote);
return [ssh, ...connectionArgs, remoteSshTarget(remote), shellescape(killCmd)].join(' ');
}
// ========== Docker cases (COD-Docker) ==========
//
// The docker analog of the remote-SSH launch above. Instead of a local tmux pane
// running `ssh -t host 'tmux new-session …'`, it runs `docker exec -it <container>
// sh -lc 'tmux new-session …'` into a DURABLE in-container tmux server. The
// container is per-CASE, so many sessions `docker exec` into the same one. See
// docs/docker-cases-plan.md.
/**
* DEDICATED in-container tmux socket. A Codeman running INSIDE the container uses
* `-L codeman`; ours is `-L codeman-docker` with a `codeman-dkr-*` session name
* that deliberately FAILS SAFE_MUX_NAME_PATTERN, so an in-container Codeman never
* adopts/resizes/respawns our session (same defence as the remote socket).
*/
const DOCKER_TMUX_SOCKET = 'codeman-docker';
/**
* Deterministic, reattach-stable in-container tmux session name. Derived from the
* same stable field the local muxName uses (first 8 chars of the sessionId), so a
* reconnect re-issues the exact same `new-session -A` and lands back in the SAME
* in-container session. The `dkr` letters make it fail SAFE_MUX_NAME_PATTERN.
*/
export function dockerTmuxSessionName(sessionId: string): string {
return `codeman-dkr-${sessionId.slice(0, 8)}`;
}
/** Resume ids are UUID-ish; reject anything with shell metacharacters (defensive). */
const RESUME_ID_SAFE = /^[A-Za-z0-9._-]+$/;
/**
* Append the CLI-specific resume flag to a pane command. Only fires when the
* in-container tmux is RE-CREATED (`new-session -A` makes the flag inert on a
* live reattach), i.e. exactly when the previous live agent was lost and we want
* to resume the conversation from the bind-mounted transcript.
*/
function appendResumeFlag(modeCommand: string, mode: SessionMode, resumeId: string): string {
if (!RESUME_ID_SAFE.test(resumeId)) return modeCommand;
switch (mode) {
case 'claude':
case 'gemini':
return `${modeCommand} --resume ${resumeId}`;
case 'codex':
return `${modeCommand} resume ${resumeId}`;
default:
return modeCommand; // shell / opencode: no resume
}
}
/** Fully-resolved inputs for buildDockerLaunchCommand (pure). */
export interface DockerLaunchOptions {
mode: SessionMode;
docker: SessionDocker;
sessionId: string;
resumeSessionId?: string;
createContext: DockerCreateContext;
/** exec-time inline env (non-secret): TERM, COLORTERM, CODEMAN_SESSION_ID, CODEMAN_MUX */
execEnv: Record<string, string>;
/** exec-time NAME-ONLY env forwarded from Codeman's process env (codex/gemini keys) */
execEnvNames: string[];
/**
* Files to copy from read-only seed mounts into the container's writable HOME once
* before launch (guarded so reconnects never clobber). Isolates Claude state: the
* merged `~/.claude.json`, plus `~/.claude/.credentials.json` + `settings.json`,
* are writable copies (not host mounts), so the container never re-auths and never
* writes its runtime state back into the host `~/.claude`.
*/
seedCopies?: DockerSeedCopy[];
}
/**
* Build the ONE `bash -c` launch string for a docker session: image-check ->
* ensure (inspect-or-create) -> start -> `exec docker exec -it` into the durable
* in-container tmux (resume-aware). PURE and unit-testable. The escaping survives
* four layers: outer `bash -c "…"` (JSON.stringify at respawn-pane) -> the joined
* command -> `docker exec … sh -lc '<tmux>'` -> tmux `'<paneCommand>'`.
*/
export function buildDockerLaunchCommand(opts: DockerLaunchOptions): string {
const { mode, docker, sessionId, resumeSessionId, createContext, execEnv, execEnvNames, seedCopies } = opts;
const base = buildDockerBaseArgs(docker).join(' ');
const createArgs = buildDockerCreateArgs(createContext).join(' ');
const name = shellescape(docker.containerName);
const workdir = shellescape(docker.containerWorkdir);
const image = shellescape(docker.image);
const dkrName = dockerTmuxSessionName(sessionId);
const sid = sessionId.slice(0, 8);
let modeCommand = docker.commands?.[mode as DockerCommandMode] || defaultDockerCommandForMode(mode);
if (resumeSessionId) modeCommand = appendResumeFlag(modeCommand, mode, resumeSessionId);
// Run by tmux via /bin/sh -c, so the path is shell-quoted here. `exec` makes the
// pane PID the agent itself.
const paneCommand = `cd ${workdir} && ${modeCommand}`;
// `setenv -g` primes the session id so reattaches / newly-created panes inherit
// it. `new-session -A` = attach-or-create (idempotent + resume-aware). Options
// are scoped per-session (`set -t`) or server (`set -s`), never `-g`, so a shared
// in-container tmux server's other sessions keep their own prefix/mouse.
const tmuxInvocation = [
`tmux -L ${DOCKER_TMUX_SOCKET} setenv -g CODEMAN_SESSION_ID ${shellescape(sid)}`,
'setenv -g CODEMAN_MUX 1',
`new-session -A -s ${dkrName} -c ${workdir} ${shellescape(paneCommand)}`,
`set -t ${dkrName} status off`,
`set -t ${dkrName} mouse off`,
`set -t ${dkrName} prefix C-q`,
'set -s escape-time 0',
].join(' \\; ');
const execEnvFlags: string[] = [];
for (const [k, v] of Object.entries(execEnv)) execEnvFlags.push('--env', shellescape(`${k}=${v}`));
// NAME-ONLY forwards: docker reads the VALUE from Codeman's own process env, so
// the secret never appears in argv (no `ps` leak) and is not committed.
for (const n of execEnvNames) execEnvFlags.push('--env', n);
for (const extra of docker.extraExecArgs ?? []) execEnvFlags.push(shellescape(extra));
const imageMissingMsg = shellescape(
`Codeman: base image ${docker.image} not present (it is normally auto-built on first use)`
);
const startFailMsg = shellescape(`Codeman: container ${docker.containerName} failed to start (docker daemon down?)`);
const imageCheck = `${base} image inspect ${image} >/dev/null 2>&1 || { echo ${imageMissingMsg}; exit 1; }`;
// create-if-missing (idempotent): reconnect / boot recovery re-runs this exact chain.
const ensure = `${base} inspect ${name} >/dev/null 2>&1 || ${base} ${createArgs}`;
const start = `${base} start ${name} >/dev/null 2>&1 || { echo ${startFailMsg}; exit 1; }`;
// Seed writable credential config from read-only host mounts ONCE per container
// (guarded by [ -e ] so reconnects never clobber in-container config; `cp -a` for
// whole-dir credential seeds). mkdir -p the parent so a file seed works even when
// no sibling share-mount pre-created the dir. Paths are fixed CONTAINER_HOME
// constants (no shell metachars), so the whole inner command is shell-quoted once.
const seedSteps = (seedCopies ?? []).map((s) => {
const cp = s.recursive ? 'cp -a' : 'cp';
const parent = s.to.slice(0, s.to.lastIndexOf('/'));
return `mkdir -p ${parent} 2>/dev/null; [ -e ${s.to} ] || ${cp} ${s.from} ${s.to} 2>/dev/null || true`;
});
const innerCmd = seedSteps.length ? `${seedSteps.join(' ; ')} ; ${tmuxInvocation}` : tmuxInvocation;
const execCmd = `exec ${base} exec -it --workdir ${workdir} ${execEnvFlags.join(' ')} ${name} sh -lc ${shellescape(innerCmd)}`;
return [imageCheck, ensure, start, execCmd].join(' ; ');
}
/**
* Kill ONLY this session's in-container tmux session. The container is shared by
* the case's other sessions, so this NEVER `docker stop`s it — stopping/removing
* the container is an explicit teardown (buildDockerStopCommand) or case-delete
* (buildDockerRemoveCommand). Fired best-effort on session kill.
*/
export function buildDockerKillCommand(options: { docker: SessionDocker; sessionId: string }): string {
const { docker, sessionId } = options;
const base = buildDockerBaseArgs(docker).join(' ');
const dkrName = dockerTmuxSessionName(sessionId);
return `${base} exec ${shellescape(docker.containerName)} tmux -L ${DOCKER_TMUX_SOCKET} kill-session -t ${shellescape(dkrName)}`;
}
/** Explicit container stop (frees RAM/CPU; conversation resumes on next launch via --resume). */
export function buildDockerStopCommand(docker: SessionDocker): string {
return `${buildDockerBaseArgs(docker).join(' ')} stop -t 10 ${shellescape(docker.containerName)}`;
}
/** Explicit container removal (case-delete). Destroys in-image state; bind mounts survive. */
export function buildDockerRemoveCommand(docker: SessionDocker): string {
return `${buildDockerBaseArgs(docker).join(' ')} rm -f ${shellescape(docker.containerName)}`;
}
/**
* Resolve the environment-dependent bits of a docker launch (host uid, existing
* credential mounts, derived api url, hook-secret mount, Desktop detection) into
* the pure buildDockerLaunchCommand inputs. IO; only ever called from the real
* launch path (createSession/respawnPane no-op under VITEST).
*/
export function resolveDockerLaunchOptions(
mode: SessionMode,
docker: SessionDocker,
sessionId: string,
resumeSessionId?: string
): DockerLaunchOptions {
const home = homedir();
const isDesktop = process.platform === 'darwin'; // Docker Desktop translates uids + native host.docker.internal
const uid = typeof process.getuid === 'function' ? process.getuid() : 1000;
const userArgs: string[] =
docker.engine === 'podman'
? ['--userns=keep-id'] // rootless podman: map host uid to the image `agent` uid
: isDesktop
? [] // Desktop: run as the image's baked uid (a mac uid wouldn't own /home/agent)
: ['--user', `${uid}:0`]; // Linux: host uid + GID 0 (OpenShift arbitrary-uid writable HOME)
const gatewayAlias = hostGatewayAlias(docker.engine);
const credentialMounts: DockerMount[] = [];
const extraMounts: DockerMount[] = [];
// Isolated credential state (Claude + codex/gemini/gcloud/opencode): each store
// shares ONLY what a host feature / --resume needs (Claude projects/, codex
// sessions/+history) and seeds everything else (tokens, settings, configs) as
// writable copies, so the container is authed WITHOUT re-auth and WITHOUT writing
// its runtime state back into the host dirs. Only when credentials are mounted.
let seedCopies: DockerSeedCopy[] = [];
if (docker.mountCredentials) {
const claudeArtifacts = resolveDockerClaudeArtifacts(home, docker.containerName, docker.containerWorkdir);
const credArtifacts = resolveDockerCredentialArtifacts(home);
extraMounts.push(...claudeArtifacts.mounts, ...credArtifacts.mounts);
seedCopies = [...claudeArtifacts.seedCopies, ...credArtifacts.seedCopies];
}
const envCreate: Record<string, string> = {
HOME: CONTAINER_HOME,
TERM: 'xterm-256color',
COLORTERM: 'truecolor',
// Force a UTF-8 locale (the base image defaults to POSIX/C). Without this, tmux
// runs in non-UTF-8 mode and renders Claude's Unicode box-drawing (─│┌┐) as raw
// VT100 ACS glyphs (`qqqq…`). `C.UTF-8` is built into glibc (no locale-gen).
LANG: 'C.UTF-8',
LC_ALL: 'C.UTF-8',
// Give claude a temp dir it will own inside HOME. Its default `/tmp/claude-<uid>`
// is refused when that path pre-exists root-owned — which happens when the
// workspace bind-mount path traverses it (e.g. a workspace under /tmp/claude-<uid>).
// A nonexistent HOME subpath is created+owned by the running uid, so this is robust
// to any workspace location. Non-secret path, safe to be committed on export.
CLAUDE_CODE_TMPDIR: `${CONTAINER_HOME}/.cache/codeman-claude-tmp`,
};
if (docker.hooksEnabled) {
// Derive a container-reachable API url (scheme + port preserved; host swapped
// for the engine gateway alias). Prod is HTTPS on 3000.
envCreate.CODEMAN_API_URL = containerApiUrl(process.env.CODEMAN_API_URL, docker.engine);
const hookSecretPath = dataPath('hook-secret');
if (existsSync(hookSecretPath)) {
const dst = `${CONTAINER_HOME}/.codeman/hook-secret`;
extraMounts.push({ src: hookSecretPath, dst, readonly: true });
envCreate.CODEMAN_HOOK_SECRET_FILE = dst; // a path is non-secret; the bytes ride the bind mount
}
}
const createContext: DockerCreateContext = {
docker,
sessionId,
instance: CODEMAN_INSTANCE,
userArgs,
credentialMounts,
extraMounts,
envCreate,
addHostGateway: !isDesktop,
gatewayAlias,
};
const execEnv: Record<string, string> = {
TERM: 'xterm-256color',
COLORTERM: 'truecolor',
// UTF-8 at exec time too, so the tmux CLIENT this exec launches is UTF-8 and
// renders box-drawing correctly even when reattaching to a container created
// before this fix (client_utf8 is per-client, resolved from the exec's locale).
LANG: 'C.UTF-8',
LC_ALL: 'C.UTF-8',
CODEMAN_SESSION_ID: sessionId.slice(0, 8),
CODEMAN_MUX: '1',
};
// NAME-ONLY exec env forwarded from Codeman's process env (the docker client
// inherits it), so API-key CLIs get their key without it appearing in argv.
const execEnvNames =
mode === 'codex'
? ['OPENAI_API_KEY', 'CODEX_API_KEY']
: mode === 'gemini'
? ['GEMINI_API_KEY', 'GOOGLE_API_KEY']
: [];
return { mode, docker, sessionId, resumeSessionId, createContext, execEnv, execEnvNames, seedCopies };
}
/**
* Set sensitive environment variables on a tmux session via setenv.
* These are inherited by panes but not visible in ps output or tmux history.
@@ -669,6 +1136,38 @@ function setCodexEnvVars(tmuxCmd: string, muxName: string): void {
}
}
/**
* Set sensitive environment variables for Gemini on a tmux session via setenv.
* Gemini Pro/Ultra users usually authenticate via cached Google login; these
* variables cover API-key and Vertex AI paths without putting secrets in ps.
*/
function setGeminiEnvVars(tmuxCmd: string, muxName: string): void {
const sensitiveVars = [
'GEMINI_API_KEY',
'GEMINI_MODEL',
'GOOGLE_API_KEY',
'GOOGLE_CLOUD_PROJECT',
'GOOGLE_CLOUD_LOCATION',
'GOOGLE_APPLICATION_CREDENTIALS',
'GOOGLE_GENAI_USE_VERTEXAI',
];
for (const key of sensitiveVars) {
const val = process.env[key];
if (val) {
const escaped = val.replace(/'/g, "'\\''");
try {
execSync(`${tmuxCmd} setenv -t '${muxName}' ${key} '${escaped}'`, {
encoding: 'utf8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['pipe', 'pipe', 'pipe'],
});
} catch {
/* Non-critical — key may not be needed */
}
}
}
}
/**
* Set OPENCODE_CONFIG_CONTENT on a tmux session via setenv.
* Uses tmux setenv to avoid shell metacharacter injection from user-supplied JSON.
@@ -851,12 +1350,21 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const exports = [
'export LANG=en_US.UTF-8',
'export LC_ALL=en_US.UTF-8',
mode === 'codex' ? 'export COLORTERM=truecolor' : 'unset COLORTERM',
...(mode === 'codex' ? ['unset NO_COLOR'] : []),
mode === 'codex' || mode === 'gemini' ? 'export COLORTERM=truecolor' : 'unset COLORTERM',
...(mode === 'codex' || mode === 'gemini' ? ['unset NO_COLOR'] : []),
// Stamp each Codex pane with a unique originator so the response-viewer
// can locate THIS pane's rollout exactly — codex writes the value into
// session_meta.originator of every rollout it creates. Without it,
// rollouts are matched by cwd+mtime and two panes in the same directory
// bleed into each other.
...(mode === 'codex' ? [`export CODEX_INTERNAL_ORIGINATOR_OVERRIDE=codeman_${sessionId}`] : []),
'export CODEMAN_MUX=1',
`export CODEMAN_SESSION_ID=${sessionId}`,
`export CODEMAN_MUX_NAME=${muxName}`,
`export CODEMAN_API_URL=${process.env.CODEMAN_API_URL || 'http://localhost:3000'}`,
// Path only (not the secret value): hook curl commands cat the file at
// execution time, so the COD-54 hook secret stays off the command line.
`export CODEMAN_HOOK_SECRET_FILE="${dataPath('hook-secret')}"`,
];
// Only unset CLAUDECODE for Claude sessions
if (mode === 'claude') exports.splice(2, 0, 'unset CLAUDECODE');
@@ -922,6 +1430,10 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
const dir = resolveCodexDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
if (mode === 'gemini') {
const dir = resolveGeminiDir();
return { pathExport: dir ? `export PATH="${dir}:$PATH" && ` : '', dir };
}
return { pathExport: '', dir: null };
}
@@ -945,6 +1457,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
setCodexEnvVars(this.tmux(), muxName);
}
/**
* Configure Gemini-specific environment on a tmux session.
*/
private _configureGemini(muxName: string): void {
setGeminiEnvVars(this.tmux(), muxName);
}
/**
* Creates a new tmux session wrapping Claude CLI or a shell.
* In test mode: creates an in-memory session only (no real tmux session).
@@ -961,9 +1480,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
allowedTools,
openCodeConfig,
codexConfig,
geminiConfig,
resumeSessionId,
envOverrides,
effort,
historyLimit = DEFAULT_TMUX_HISTORY_LIMIT,
remote,
docker,
} = options;
const muxName = `codeman-${sessionId.slice(0, 8)}`;
@@ -982,6 +1505,8 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
pid: 99999,
createdAt: Date.now(),
workingDir,
remote,
docker,
mode,
attached: false,
name,
@@ -999,6 +1524,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
if (mode === 'opencode' && !cliDir) {
throw new Error('OpenCode CLI not found. Install with: curl -fsSL https://opencode.ai/install | bash');
}
if (mode === 'codex' && !cliDir) {
throw new Error('Codex CLI not found. Install with: npm install -g @openai/codex');
}
if (mode === 'gemini' && !cliDir) {
throw new Error('Gemini CLI not found. Install with: npm install -g @google/gemini-cli');
}
const envExportsStr = this.buildEnvExports(sessionId, muxName, mode).join(' && ');
@@ -1010,6 +1541,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
allowedTools,
openCodeConfig,
codexConfig,
geminiConfig,
resumeSessionId,
effort,
});
@@ -1019,7 +1551,12 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
try {
// Build the full command to run inside tmux
const fullCmd = `${buildNofileLimitCommand()} && ${pathExport}${envExportsStr} && ${cmd}`;
const localFullCmd = `${buildNofileLimitCommand()} && ${pathExport}${envExportsStr} && ${cmd}`;
const fullCmd = docker
? buildDockerLaunchCommand(resolveDockerLaunchOptions(mode, docker, sessionId, resumeSessionId))
: remote
? buildRemoteLaunchCommand({ mode, remote, sessionId })
: localFullCmd;
// Create tmux session in three steps to handle cold-start (no server running)
// and avoid the race where the command exits before remain-on-exit is set:
@@ -1059,6 +1596,10 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
} else if (mode === 'codex') {
this._configureCodex(muxName);
}
// For Gemini: set Gemini/Google auth env vars via tmux setenv
if (mode === 'gemini') {
this._configureGemini(muxName);
}
// Apply user-supplied env overrides (e.g., CLAUDE_CODE_EFFORT_LEVEL) via tmux setenv
// so secret values stay off the bash command line. Must run before respawn-pane.
@@ -1066,7 +1607,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
// Replace the shell with the actual command (no echo in terminal). Keep
// pane launch in /tmp, then cd inside bash against the current mount table.
const launchCmd = `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
const launchCmd = remote || docker ? fullCmd : `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
execSync(
`${this.tmux()} respawn-pane -k -c ${TMUX_LAUNCH_CWD} -t "${muxName}" bash -c ${JSON.stringify(launchCmd)}`,
{
@@ -1096,8 +1637,11 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
/* Already set globally as fallback */
}),
// Raise tmux scrollback from its 2000-line default so re-attach preserves
// more context. Matches the xterm-side default in constants.js.
execAsync(`${this.tmux()} set-option -t "${muxName}" history-limit 50000`, { timeout: EXEC_TIMEOUT_MS })
// more context. Intentionally exceeds the xterm-side DEFAULT_SCROLLBACK (50k
// in constants.js), which stays lower to protect browser/mobile memory.
execAsync(`${this.tmux()} set-option -t "${muxName}" history-limit ${historyLimit}`, {
timeout: EXEC_TIMEOUT_MS,
})
.then(() => {})
.catch(() => {
/* Non-critical — falls back to tmux default */
@@ -1137,6 +1681,8 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
pid,
createdAt: Date.now(),
workingDir,
remote,
docker,
mode,
attached: false,
name,
@@ -1216,9 +1762,13 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
allowedTools,
openCodeConfig,
codexConfig,
geminiConfig,
resumeSessionId,
envOverrides,
effort,
historyLimit = DEFAULT_TMUX_HISTORY_LIMIT,
remote,
docker,
} = options;
const session = this.sessions.get(sessionId);
if (!session) return null;
@@ -1226,6 +1776,16 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
if (!isValidMuxName(muxName) || !isValidPath(workingDir)) return null;
// Re-apply the configured tmux history-limit after respawn (kept in sync
// with the live setting via setHistoryLimit()).
if (!IS_TEST_MODE) {
await execAsync(`${this.tmux()} set-option -t ${shellescape(muxName)} history-limit ${historyLimit}`, {
timeout: EXEC_TIMEOUT_MS,
}).catch(() => {
/* Non-critical — keeps existing tmux history-limit */
});
}
// Resolve CLI binary directory based on mode
const { pathExport } = this.buildPathExport(mode);
@@ -1239,12 +1799,18 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
allowedTools,
openCodeConfig,
codexConfig,
geminiConfig,
resumeSessionId,
effort,
});
const config = niceConfig || DEFAULT_NICE_CONFIG;
const cmd = wrapWithNice(baseCmd, config);
const fullCmd = `${buildNofileLimitCommand()} && ${pathExport}${envExportsStr} && ${cmd}`;
const localFullCmd = `${buildNofileLimitCommand()} && ${pathExport}${envExportsStr} && ${cmd}`;
const fullCmd = docker
? buildDockerLaunchCommand(resolveDockerLaunchOptions(mode, docker, sessionId, resumeSessionId))
: remote
? buildRemoteLaunchCommand({ mode, remote, sessionId })
: localFullCmd;
try {
// For OpenCode: set sensitive env vars via tmux setenv before respawn
@@ -1253,11 +1819,16 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
} else if (mode === 'codex') {
this._configureCodex(muxName);
}
// For Gemini: set Gemini/Google auth env vars via tmux setenv before respawn
if (mode === 'gemini') {
this._configureGemini(muxName);
}
// Re-apply user env overrides before respawn so the new shell inherits them.
this.applyEnvOverrides(muxName, envOverrides);
const launchCmd = `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
// -c /tmp + cd bounce — see createSession() for rationale (stale FUSE state).
const launchCmd = remote || docker ? fullCmd : `cd ${JSON.stringify(workingDir)} && ${fullCmd}`;
await execAsync(
`${this.tmux()} respawn-pane -k -c ${TMUX_LAUNCH_CWD} -t "${muxName}" bash -c ${JSON.stringify(launchCmd)}`,
{
@@ -1429,6 +2000,32 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
// Strategy 3b: Remote sessions run a DURABLE tmux server on the remote host
// (survives ssh drops), so killing only the local ssh wrapper above would
// orphan the remote agent forever. Fire a best-effort `ssh … tmux kill-session`
// — fire-and-forget so it NEVER blocks or throws the local kill (bounded by the
// shared ConnectTimeout on an unreachable host).
if (session.remote) {
try {
const remoteKillCmd = buildRemoteKillCommand({ remote: session.remote, sessionId });
exec(remoteKillCmd, { timeout: EXEC_TIMEOUT_MS }, () => {});
} catch {
// Best-effort — a failure here must not affect the local kill result.
}
}
// Strategy 3c: Docker sessions run a DURABLE in-container tmux session. Kill
// ONLY this session's in-container tmux session (best-effort). The container is
// PER-CASE and shared by the case's other sessions, so we deliberately do NOT
// `docker stop` it here — stopping/removing is an explicit teardown/case-delete.
if (session.docker && !IS_TEST_MODE) {
try {
exec(buildDockerKillCommand({ docker: session.docker, sessionId }), { timeout: EXEC_TIMEOUT_MS }, () => {});
} catch {
// Best-effort — never affects the local kill result.
}
}
// Strategy 4: Direct kill by PID as final fallback
if (this.isProcessAlive(currentPid)) {
try {
@@ -1825,6 +2422,26 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
}
/**
* Apply a tmux history-limit to all tracked sessions (e.g. when the user
* changes the terminal-history setting). Invalid limits fall back to the
* default. Best-effort per session.
*/
async setHistoryLimit(limit: number): Promise<void> {
const safeLimit = Number.isSafeInteger(limit) && limit > 0 ? Math.trunc(limit) : DEFAULT_TMUX_HISTORY_LIMIT;
if (IS_TEST_MODE) {
return;
}
const updates = Array.from(this.sessions.values()).map((session) =>
execAsync(`${this.tmux()} set-option -t ${shellescape(session.muxName)} history-limit ${safeLimit}`, {
timeout: EXEC_TIMEOUT_MS,
})
);
await Promise.allSettled(updates);
}
/**
* Send input directly to a tmux session using `send-keys`.
*
@@ -2046,30 +2663,65 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
}
/**
* Capture the current visible text and SGR styles of a specific pane.
* Capture a pane's text and SGR styles.
*
* `capture-pane -e` is sanitized by `formatPaneSnapshot`: SGR color/style
* codes are preserved, while cursor/erase/scroll-region controls are stripped
* before rows are repainted at absolute positions in browser xterm.
* Two modes:
* - Visible (default): `capture-pane -p -e` grabs only the on-screen frame,
* then `formatPaneSnapshot` repaints each row at its absolute position so
* the browser xterm reproduces the live frame. Used for fast tab switches.
* - Full history (`opts.fullHistory`): `capture-pane -p -e -J -S -<N>` grabs
* the tmux scrollback (COD-47, bounded to the configured history limit),
* returned as linear scrollback text with SGR codes preserved (NOT
* repositioned — a multi-screen history can't be painted into a single
* visible frame, so the snapshot repaint is skipped). `-J` re-joins lines
* hard-wrapped at the pane width so they reflow in the browser xterm.
* Used for full page reloads so the user gets back their scroll history.
* Caveat: lines tmux has already evicted past its history-limit are gone.
*/
capturePaneBuffer(muxName: string, paneTarget: string): string | null {
capturePaneBuffer(muxName: string, paneTarget?: string, opts?: PaneCaptureOptions): string | null {
if (IS_TEST_MODE) return '';
if (!isValidMuxName(muxName)) {
console.error('[TmuxManager] Invalid session name in capturePaneBuffer:', muxName);
return null;
}
if (!SAFE_PANE_TARGET_PATTERN.test(paneTarget)) {
console.error('[TmuxManager] Invalid pane target:', paneTarget);
const target = resolveTmuxPaneTarget(muxName, paneTarget);
if (!target) {
console.error('[TmuxManager] Invalid pane target in capturePaneBuffer:', { muxName, paneTarget });
return null;
}
const target = paneTarget.startsWith('%') ? `${muxName}.${paneTarget}` : `${muxName}.%${paneTarget}`;
const fullHistory = opts?.fullHistory === true;
try {
const buffer = execSync(`${this.tmux()} capture-pane -p -e -t ${shellescape(target)}`, {
// `-S -<N>` starts the capture N lines above the visible frame (tmux
// clamps to the top of history), so tmux never serializes more scrollback
// than the configured history limit retains.
const requestedLines = opts?.historyLimitLines;
const historyLines =
typeof requestedLines === 'number' && Number.isFinite(requestedLines) && requestedLines > 0
? Math.trunc(requestedLines)
: DEFAULT_TMUX_HISTORY_LIMIT;
const captureFlags = fullHistory ? `capture-pane -p -e -J -S -${historyLines}` : 'capture-pane -p -e';
// execSync's default maxBuffer (1MB) kills multi-MB scrollback dumps
// (ENOBUFS) and would silently degrade full-history capture to the byte
// buffer for exactly the long sessions it exists for — size it from the
// consumer's byte cap plus ANSI-overhead slack instead.
const execOpts: { encoding: 'utf-8'; timeout: number; maxBuffer?: number } = {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).replace(/\n+$/g, '');
};
if (fullHistory) {
execOpts.maxBuffer =
(opts?.maxCaptureBytes ?? DEFAULT_TERMINAL_BUFFER_MAX_BYTES) + FULL_HISTORY_CAPTURE_SLACK_BYTES;
}
const buffer = execSync(`${this.tmux()} ${captureFlags} -t ${shellescape(target)}`, execOpts).replace(
/\n+$/g,
''
);
// Full-history spans many screens — return it as raw linear scrollback
// rather than repainting rows at single-screen absolute positions. tmux
// joins scrollback rows with a bare `\n`; normalize to `\r\n` so a fresh
// xterm (convertEol:false) starts each replayed line at column 0 instead
// of staircasing diagonally (COD-138).
if (fullHistory) {
return normalizeScrollbackEol(buffer);
}
try {
const cursor = execSync(
`${this.tmux()} display-message -p -t ${shellescape(target)} '#{cursor_x} #{cursor_y} #{pane_width} #{pane_height}'`,
@@ -2094,9 +2746,19 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
} catch (cursorErr) {
console.error('[TmuxManager] Failed to query pane cursor after capture:', cursorErr);
}
return buffer;
// Cursor query failed or geometry was invalid, so we skip the absolute-
// positioned snapshot repaint and fall back to the raw capture. Normalize
// its bare `\n` line endings to `\r\n` so the replay doesn't staircase
// diagonally in a fresh xterm (COD-138, same reason as the fullHistory path).
return normalizeScrollbackEol(buffer);
} catch (err) {
console.error('[TmuxManager] Failed to capture pane buffer:', err);
// ENOBUFS carries the truncated multi-MB stdout on the error object —
// log a concise line instead of dumping it into the journal.
if ((err as NodeJS.ErrnoException)?.code === 'ENOBUFS') {
console.error('[TmuxManager] Pane capture exceeded maxBuffer (ENOBUFS); falling back to byte history');
} else {
console.error('[TmuxManager] Failed to capture pane buffer:', err);
}
return null;
}
}
@@ -2107,7 +2769,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
* Pane ids are not stable across respawns or restores, so callers should not
* assume the first pane remains `%0`.
*/
captureActivePaneBuffer(muxName: string): string | null {
captureActivePaneBuffer(muxName: string, opts?: PaneCaptureOptions): string | null {
if (IS_TEST_MODE) return '';
if (!isValidMuxName(muxName)) {
console.error('[TmuxManager] Invalid session name in captureActivePaneBuffer:', muxName);
@@ -2120,7 +2782,7 @@ export class TmuxManager extends EventEmitter implements TerminalMultiplexer {
timeout: EXEC_TIMEOUT_MS,
}).trim();
const target = resolveActivePaneTarget(output);
return target ? this.capturePaneBuffer(muxName, target) : null;
return target ? this.capturePaneBuffer(muxName, target, opts) : null;
} catch (err) {
console.error('[TmuxManager] Failed to resolve active pane for capture:', err);
return null;
+19
View File
@@ -123,6 +123,25 @@ export interface CaseInfo {
path: string;
/** Whether CLAUDE.md exists */
hasClaudeMd?: boolean;
/** Case storage/execution location */
location?: 'local' | 'linked-local' | 'remote' | 'docker';
/** Whether this is a linked local folder */
linked?: boolean;
/** Remote case metadata for display and session creation */
remote?: {
hostId: string;
host: string;
username: string;
path: string;
};
/** Docker case metadata for display and session creation */
docker?: {
hostId: string;
container: string;
image?: string;
path: string;
network?: string;
};
}
// ========== Error Handling Utilities ==========
+5
View File
@@ -23,6 +23,7 @@ import type { SessionState } from './session.js';
import type { TaskState } from './task.js';
import type { RalphLoopState } from './ralph.js';
import type { RespawnConfig } from './respawn.js';
import type { CronJob, CronJobRun } from './cron.js';
// ========== Global Stats Types ==========
@@ -111,6 +112,10 @@ export interface AppState {
tokenStats?: TokenStats;
/** Orchestrator Loop state (phased plan execution) */
orchestrator?: import('./orchestrator.js').OrchestratorPersistState;
/** Cron-style scheduled jobs, keyed by job ID. */
cronJobs?: Record<string, CronJob>;
/** Scheduled job run history, keyed by run ID. */
cronJobRuns?: Record<string, CronJobRun>;
}
// ========== Default Configuration ==========
+101
View File
@@ -0,0 +1,101 @@
/**
* @fileoverview Cron Jobs type definitions.
*
* NOTE: This is intentionally distinct from the existing `ScheduledRun` concept
* (see src/web/ports/infra-port.ts), which is a run-now, duration-bounded
* autonomous loop. A `CronJob` is a SAVED, NAMED job with a recurring
* schedule (once/interval/daily/weekly), enable/disable, next-run calculation,
* and a history of `CronJobRun` records. The two do not interact.
*
* Persisted to `~/.codeman/state.json` via StateStore (see AppState).
*/
import type { SessionMode } from './session.js';
/** How a job's fire times are computed. */
export type ScheduleType = 'once' | 'interval' | 'daily' | 'weekly';
/** Where the prompt text comes from. */
export type PromptMode = 'inline_text' | 'prompt_file_path';
/** How the prompt is delivered into the session. */
export type InputMode = 'paste' | 'typed';
/** Lifecycle status of a single job execution. */
export type CronJobRunStatus = 'created' | 'session_started' | 'prompt_sent' | 'failed' | 'skipped';
/** What triggered a run. */
export type TriggerType = 'scheduled' | 'manual_run_now';
/** What to do for an AUTOMATIC run when sessions of the same agent already exist. */
export type ConcurrencyPolicy = 'warn_only' | 'skip_if_same_agent_running';
/**
* A saved, named cron job.
*/
export interface CronJob {
id: string;
name: string;
/** Reuses Codeman's existing session modes; 'shell' covers Terminal/custom. */
agentType: SessionMode;
workingDir: string;
/** Optional custom launch command (only meaningful for 'shell' mode). */
launchCommand?: string;
promptMode: PromptMode;
promptText?: string;
promptFilePath?: string;
inputMode: InputMode;
scheduleType: ScheduleType;
/** once: absolute epoch-ms fire time. */
runAt?: number;
/** interval: minutes between fires. */
intervalMinutes?: number;
/** daily: 'HH:MM' (24h, server-local time). */
dailyTime?: string;
/** weekly: weekdays 0–6 (0=Sunday). */
weeklyDays?: number[];
/** weekly: 'HH:MM' (24h, server-local time). */
weeklyTime?: string;
enabled: boolean;
notes?: string;
/** Applies to automatic (scheduled) runs only. Manual Run Now always warns client-side. */
concurrencyPolicy: ConcurrencyPolicy;
/**
* Close the still-open session created by this job's previous run before the
* next run launches (via the normal session-cleanup path), so unattended
* recurring jobs don't accumulate tabs until the global session cap.
* Default true. Ignored for 'once' schedules.
*/
autoClosePreviousSession?: boolean;
// ── Bookkeeping (server-maintained) ─────────────────────────────────────
createdAt: number;
updatedAt: number;
lastRunAt: number | null;
nextRunAt: number | null;
lastStatus: CronJobRunStatus | null;
/** Duplicate-launch guard: identifies the most recent due-time consumed. */
lastDueKey: string | null;
/** True once a 'once' job has fired (it is also disabled). */
completedOnce?: boolean;
}
/**
* A single execution of a cron job (history record).
*/
export interface CronJobRun {
id: string;
cronJobId: string;
sessionId: string | null;
sessionName: string | null;
startedAt: number;
finishedAt: number | null;
status: CronJobRunStatus;
errorMessage?: string;
triggerType: TriggerType;
/** Best-effort deep link to the created session in the web UI. */
createdSessionUrl: string | null;
}
+24
View File
@@ -0,0 +1,24 @@
declare module 'heic-decode' {
export interface DecodedHeicImage {
width: number;
height: number;
data: Uint8ClampedArray;
}
/** Handle exposing header-declared dimensions WITHOUT decoding pixels. */
export interface HeicImageHandle {
width: number;
height: number;
decode(): Promise<DecodedHeicImage>;
}
export type HeicImageHandles = HeicImageHandle[] & { dispose(): void };
interface HeicDecode {
(input: { buffer: Buffer | Uint8Array }): Promise<DecodedHeicImage>;
all(input: { buffer: Buffer | Uint8Array }): Promise<HeicImageHandles>;
}
const decode: HeicDecode;
export default decode;
}
+2
View File
@@ -67,3 +67,5 @@ export * from './push.js';
export * from './plan.js';
export * from './orchestrator.js';
export * from './update.js';
export * from './workflow-run.js';
export * from './search.js';
+4
View File
@@ -90,6 +90,10 @@ export interface RalphTrackerState {
cycleCount: number;
/** Maximum iterations if detected */
maxIterations: number | null;
/** Max todos retained for this session before FIFO eviction (persisted; default = global cap) */
maxTodos?: number;
/** Todo auto-expiry in minutes (persisted; default = global TODO_EXPIRY_MS) */
todoExpirationMinutes?: number;
/** Timestamp of last activity */
lastActivity: number;
/** Elapsed hours if detected */
+77
View File
@@ -0,0 +1,77 @@
/**
* @fileoverview Cross-session federated search types (COD-9).
*
* Defines the typed shapes for `GET /api/search` — a bounded, in-memory
* federated search across three v1 sources: live sessions/cases, run-summary
* timeline events, and per-session attachment file paths. Terminal-buffer scans
* and any persisted index are explicitly out of scope for v1.
*
* Key exports:
* - SearchSourceType — the federated source kinds, also the group order key.
* - SearchResult — a single typed result card (source, session id/name,
* timestamp, snippet, jump-to action target).
* - SearchJumpTarget — where the frontend should navigate when a card is opened.
* - SearchResponseData — grouped result payload returned in the ApiResponse envelope.
*
* No I/O, no dependencies on other domain modules. The pure search core lives
* in `src/search-service.ts`; the route wrapper in `src/web/routes/search-routes.ts`.
*/
/** Federated source kinds. Group/render order is sessions → events → files. */
export type SearchSourceType = 'session' | 'event' | 'file';
/** Where the frontend should jump when a result card is activated. */
export interface SearchJumpTarget {
/** Kind of navigation target. */
kind: 'session' | 'run-summary' | 'file-preview';
/** Owning Codeman session id (always present — every result is session-scoped). */
sessionId: string;
/**
* Secondary identifier for the target:
* - kind 'run-summary': the run-summary event id
* - kind 'file-preview': the attachment history item id
* - kind 'session': undefined (the sessionId is sufficient)
*/
targetId?: string;
/**
* Workspace-relative path for file-preview targets. Never an absolute path —
* server-private external paths are intentionally omitted to avoid leakage.
*/
relativePath?: string;
}
/** A single typed search result card. */
export interface SearchResult {
/** Which federated source produced this result. */
type: SearchSourceType;
/** Owning Codeman session id. */
sessionId: string;
/** Display name of the owning session / case. */
sessionName: string;
/** Millisecond timestamp used for recency ranking and display. */
timestamp: number;
/** Short, already-truncated snippet describing the match. */
snippet: string;
/** True when the query matched the primary name/path exactly (case-insensitive). */
exactMatch: boolean;
/** Navigation target for the jump-to action. */
jumpTo: SearchJumpTarget;
}
/** A group of results for one source type, in render order. */
export interface SearchResultGroup {
type: SearchSourceType;
results: SearchResult[];
}
/** Payload returned as `data` inside the standard ApiResponse envelope. */
export interface SearchResponseData {
/** The normalized query that was executed. */
query: string;
/** Results grouped by source type, ordered sessions → events → files. */
groups: SearchResultGroup[];
/** Total number of results across all groups (after caps applied). */
totalResults: number;
/** True if any group or the total was capped (more matches existed). */
truncated: boolean;
}
+237 -2
View File
@@ -8,10 +8,12 @@
* - SessionConfig — creation-time config (id, workingDir, createdAt)
* - SessionOutput — captured stdout/stderr/exitCode
* - SessionStatus — 'idle' | 'busy' | 'stopped' | 'error'
* - SessionMode — 'claude' | 'shell' | 'opencode' | 'codex' (which CLI backend)
* - SessionMode — 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini' (which CLI backend)
* - ClaudeMode — CLI permission mode ('dangerously-skip-permissions' | 'normal' | 'allowedTools')
* - SessionColor — visual differentiation color
* - OpenCodeConfig — OpenCode-specific settings (model, autoAllowTools, continueSession)
* - CodexConfig — Codex (OpenAI CLI)-specific settings (model, resumeSessionId)
* - GeminiConfig — Gemini CLI-specific settings (model, approvalMode, resumeSession)
*
* Cross-domain relationships:
* - SessionState.respawnConfig embeds RespawnConfig (respawn domain)
@@ -25,6 +27,7 @@
*/
import type { RespawnConfig } from './respawn.js';
import type { AttachmentDetectedType } from './tools.js';
/** Status of a Claude session */
export type SessionStatus = 'idle' | 'busy' | 'stopped' | 'error';
@@ -38,7 +41,178 @@ export type SessionStatus = 'idle' | 'busy' | 'stopped' | 'error';
export type ClaudeMode = 'dangerously-skip-permissions' | 'normal' | 'allowedTools';
/** Session mode: which CLI backend a session runs */
export type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex';
export type SessionMode = 'claude' | 'shell' | 'opencode' | 'codex' | 'gemini';
export type RemoteCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
/**
* Advanced SSH connection options shared by RemoteHost and SessionRemote.
*
* COD-107 — all fields are optional; every field absent reproduces today's
* behavior (port-22, default-identity, directly-SSH-able hosts). These describe
* HOW Codeman reaches the host (identity, proxy, jump host, arbitrary `-o`),
* letting it connect to e.g. a host fronted by a cloudflared SOCKS5 proxy on a
* custom port — the same connection `ssh-aa-desktop` makes — without a wrapper.
*/
export interface RemoteSshOptions {
/**
* Path to an SSH identity (private key) file — path ONLY, never key bytes.
* A leading `~`/`$HOME` is expanded to an absolute path at command-build time
* (ssh does not expand `~` in `-i`).
*/
identityFile?: string;
/**
* SOCKS5 proxy as `host:port` (e.g. `127.0.0.1:1080`). Expands to
* `-o ProxyCommand=nc -X 5 -x <host:port> %h %p` (the cloudflared/SOCKS5 case).
*/
socksProxy?: string;
/** SSH jump host (`[user@]host[:port]`) emitted as `-J <jumpHost>`. */
jumpHost?: string;
/** Arbitrary additional `-o KEY=VALUE` options (escape hatch). Each `KEY=VALUE`. */
extraSshOptions?: string[];
}
export interface RemoteHost extends RemoteSshOptions {
id: string;
label: string;
host: string;
username: string;
port?: number;
commands?: Partial<Record<RemoteCommandMode, string>>;
}
export interface RemoteCase {
name: string;
type: 'remote';
hostId: string;
remotePath: string;
}
export interface SessionRemote extends RemoteSshOptions {
hostId: string;
label: string;
host: string;
username: string;
port?: number;
remotePath: string;
commands?: Partial<Record<RemoteCommandMode, string>>;
}
// ========== Docker cases (COD-Docker) ==========
//
// Docker mode is a LOCATION OVERLAY on cases (never a 6th SessionMode), the exact
// analog of the remote-SSH feature above: instead of a local tmux pane running
// `ssh host` into a durable remote tmux server, a local tmux pane runs
// `docker exec -it` into a durable in-container tmux server. The container is
// scoped to the CASE (not the session), so multiple sessions can `docker exec`
// into the same long-lived container. See `docs/docker-cases-plan.md`.
/** Which CLI backends a Docker case can run (same set as remote). */
export type DockerCommandMode = Extract<SessionMode, 'shell' | 'claude' | 'opencode' | 'codex' | 'gemini'>;
/** Container engine. Docker and Podman differ in the uid/userns + host-gateway alias. */
export type DockerEngine = 'docker' | 'podman';
/**
* Container network mode. `host` and any inbound `-p` publish are deliberately
* unrepresentable (never in this union, never emitted by the flag builder).
* - `bridge`: own netns, NAT egress, no inbound (default — every API CLI needs egress)
* - `none`: fully offline sandbox (breaks API CLIs; reserved for `shell`)
* - `custom`: a user-defined bridge `codeman-net-<slug>` (future egress-allowlist chokepoint)
*/
export type DockerNetworkMode = 'bridge' | 'none' | 'custom';
/** Per-container resource caps. Advisory under non-delegated rootless (see `capsEnforced`). */
export interface DockerResourceLimits {
/** e.g. '4g' -> --memory 4g --memory-swap 4g (swap==memory: a real OOM cap) */
memory?: string;
/** e.g. '2' -> --cpus 2 */
cpus?: string;
/** e.g. 512 -> --pids-limit 512 (fork-bomb guard) */
pidsLimit?: number;
/** e.g. '4096:8192' -> --ulimit nofile=4096:8192 */
nofile?: string;
/** e.g. '256m' -> --shm-size (only when a tool needs /dev/shm) */
shmSize?: string;
}
/** A reusable Docker engine/image/network/resource profile (mirror of RemoteHost). */
export interface DockerHost {
id: string;
label: string;
/** Engine; when absent the availability probe resolves it (docker, else podman). */
engine?: DockerEngine;
/** Base image ref (built locally by scripts/build-agent-image.mjs, e.g. codeman/agent:base). */
image: string;
/** Advanced: remote daemon (-H ssh://user@host or a DOCKER_HOST value). */
daemonHost?: string;
/** Advanced: docker `--context` name. */
context?: string;
/** Network mode (default 'bridge'). */
network?: DockerNetworkMode;
/** Custom bridge name when network === 'custom'. */
networkName?: string;
resources?: DockerResourceLimits;
/** GPU allocation, e.g. 'all' / '1' / 'device=0,1' -> `--gpus <value>` (needs the NVIDIA container toolkit). */
gpus?: string;
/** true (default) = convenient: bind-mount host cred dirs RW. false = sealed (blocks full-image export). */
mountCredentials?: boolean;
/** true (default) = wire in-container hooks (host-gateway callback + workspace scaffold). */
hooksEnabled?: boolean;
/** true (default) = a relaunch resumes the last conversation from the bind-mounted transcript. */
resumeOnStart?: boolean;
/** Per-mode command overrides (mirror RemoteHost.commands). */
commands?: Partial<Record<DockerCommandMode, string>>;
/** Escape hatch: extra `docker create` args (validated like extraSshOptions). */
extraCreateArgs?: string[];
/** Escape hatch: extra `docker exec` args. */
extraExecArgs?: string[];
}
/** A case linked to a Docker container (mirror of RemoteCase). */
export interface DockerCase {
name: string;
type: 'docker';
hostId: string;
/** Absolute HOST directory: the bind-mount source AND Session.workingDir (real host bytes). */
hostWorkspacePath: string;
/** Container path (default = hostWorkspacePath: mirror -> transcript projHash correlates). */
containerWorkdir?: string;
/** Container name (default codeman-case-<slug>). */
container?: string;
/** Last captured Claude conversation id, replayed via --resume on a fresh launch. */
lastClaudeSessionId?: string;
}
/**
* Flattened Docker execution metadata carried on a live session (mirror of
* SessionRemote). Round-trips through MuxSession/SessionState/mux-sessions.json.
*/
export interface SessionDocker {
hostId: string;
label: string;
engine: DockerEngine;
image: string;
/** Per-CASE container name (shared by all sessions of the case). */
containerName: string;
hostWorkspacePath: string;
containerWorkdir: string;
network: DockerNetworkMode;
networkName?: string;
resources?: DockerResourceLimits;
/** GPU allocation ('all' / '1' / 'device=0,1'). */
gpus?: string;
mountCredentials: boolean;
hooksEnabled: boolean;
resumeOnStart: boolean;
daemonHost?: string;
context?: string;
commands?: Partial<Record<DockerCommandMode, string>>;
extraCreateArgs?: string[];
extraExecArgs?: string[];
/** Stable hash of the drift-relevant create args (recreate-on-drift detection). */
configHash?: string;
}
/**
* Valid Claude CLI effort levels (claude >= 2.1.154).
@@ -84,6 +258,16 @@ export interface CodexConfig {
renderMode?: CodexRenderMode;
}
/** Gemini CLI session configuration */
export interface GeminiConfig {
/** Model identifier (e.g., "gemini-2.5-pro"). Passed via --model. */
model?: string;
/** Gemini approval mode for tool calls. */
approvalMode?: 'default' | 'auto_edit' | 'yolo' | 'plan';
/** Resume a previous Gemini session ("latest", index, or session id). */
resumeSession?: string;
}
/**
* Configuration for creating a new session
*/
@@ -101,6 +285,40 @@ export interface SessionConfig {
*/
export type SessionColor = 'default' | 'red' | 'orange' | 'yellow' | 'green' | 'blue' | 'purple' | 'pink';
export type SessionAttachmentHistorySource = 'detected' | 'external';
/**
* Session-scoped attachment history entry.
*
* `externalPath` is server-private. It may be present in the internal persisted
* history copy, but API-bound session state must sanitize it before returning
* to the browser.
*/
export interface SessionAttachmentHistoryItem {
/** Stable history identity used for dedupe and list rendering */
id: string;
/** Codeman session ID this item belongs to */
sessionId: string;
/** Display filename */
fileName: string;
/** Lowercase extension without a leading dot */
extension: string;
/** Viewer category used by the web UI */
attachmentType: AttachmentDetectedType;
/** File size in bytes */
size: number;
/** Last modified timestamp in milliseconds, if known */
mtimeMs: number;
/** Last time this attachment was seen or explicitly published */
timestamp: number;
/** How the attachment entered the session */
source: SessionAttachmentHistorySource;
/** Workspace-relative path for detected session files */
relativePath?: string;
/** Server-private absolute path for explicitly published external files */
externalPath?: string;
}
/**
* Current state of a session
*/
@@ -113,6 +331,10 @@ export interface SessionState {
status: SessionStatus;
/** Working directory path */
workingDir: string;
/** Remote execution metadata, present when this session runs over SSH through local tmux */
remote?: SessionRemote;
/** Docker execution metadata, present when this session runs inside a container via local tmux + docker exec */
docker?: SessionDocker;
/** ID of currently assigned task, null if none */
currentTaskId: string | null;
/** Timestamp when session was created */
@@ -133,6 +355,10 @@ export interface SessionState {
autoCompactThreshold?: number;
/** Auto-compact prompt */
autoCompactPrompt?: string;
/** Auto-resume on usage limit enabled */
autoResumeEnabled?: boolean;
/** Pending usage-limit auto-resume fire time (epoch ms), if armed */
autoResumeAt?: number;
/** Image watcher enabled for this session */
imageWatcherEnabled?: boolean;
/** Total cost in USD */
@@ -175,10 +401,19 @@ export interface SessionState {
openCodeConfig?: OpenCodeConfig;
/** Codex-specific configuration (only for mode === 'codex') */
codexConfig?: CodexConfig;
/** Gemini-specific configuration (only for mode === 'gemini') */
geminiConfig?: GeminiConfig;
/** Claude conversation session ID to resume after reboot (set by restore script) */
resumeSessionId?: string;
/** Claude CLI effort level (soft default via --settings, switchable in-session via /effort) */
effort?: EffortLevel;
/** Sanitized per-session attachment history. */
attachmentHistory?: SessionAttachmentHistoryItem[];
/**
* PTY-exit circuit breaker tripped — respawn blocked until an explicit restart
* (COD-118). Runtime-only: never restored on boot (fresh server = fresh breaker).
*/
respawnBlocked?: boolean;
}
/**
+37 -1
View File
@@ -7,13 +7,14 @@
* - ActiveBashTool — a live bash command with extracted file paths and status
* - ActiveBashToolStatus — 'running' | 'completed'
* - ImageDetectedEvent — screenshot/image file detection trigger for UI popup
* - AttachmentDetectedEvent — document/image file detection trigger for attachment cards
*
* Cross-domain relationships:
* - ActiveBashTool.sessionId links to SessionState.id (session domain)
* - ImageDetectedEvent.sessionId links to SessionState.id (session domain)
*
* Both types are in-memory only (not persisted). Broadcast via SSE events
* `subagent:tool_call` and `image:detected`. Parsed by BashToolParser
* `subagent:tool_call`, `image:detected`, and `attachment:detected`. Parsed by BashToolParser
* (`src/bash-tool-parser.ts`).
*/
@@ -61,3 +62,38 @@ export interface ImageDetectedEvent {
/** File size in bytes */
size: number;
}
export type AttachmentDetectedType = 'image' | 'pdf' | 'document' | 'presentation' | 'markdown' | 'text';
/**
* Event emitted when a new previewable attachment file is detected in a session's
* working directory. Used to render a compact attachment card in the web UI.
*/
export interface AttachmentDetectedEvent {
/** Codeman session ID where the attachment was detected */
sessionId: string;
/** Full path to the detected attachment file */
filePath: string;
/** Path relative to the session's working directory (for file-raw/file-preview endpoints) */
relativePath: string;
/** Attachment file name (basename) */
fileName: string;
/** Lowercase extension without a leading dot */
extension: string;
/** Viewer category used by the web UI */
attachmentType: AttachmentDetectedType;
/** Timestamp when the attachment was detected */
timestamp: number;
/** File size in bytes */
size: number;
/** Registered attachment id for explicit live external attachments */
attachmentId?: string;
/** Source of the attachment card request */
source?: 'detected' | 'external';
/** Raw file route for explicit attachments */
rawUrl?: string;
/** Inline preview route for explicit attachments */
previewUrl?: string;
/** First-page thumbnail route for card previews */
thumbnailUrl?: string;
}
+6 -2
View File
@@ -13,8 +13,12 @@
* @module types/update
*/
/** Which init system supervises the running server (decides how we restart it). */
export type SupervisorKind = 'systemd' | 'launchd' | 'none';
/**
* Which init system supervises the running server (decides how we restart it).
* `launchd-daemon` = a KeepAlive system-level LaunchDaemon (headless Macs, no GUI
* login): restart works by killing the server and letting launchd respawn it.
*/
export type SupervisorKind = 'systemd' | 'launchd' | 'launchd-daemon' | 'none';
/** How Codeman was installed — only `git` installs can self-update in place. */
export type InstallKind = 'git' | 'npm' | 'unknown';
+135
View File
@@ -0,0 +1,135 @@
/**
* @fileoverview Types for ultracode / Workflow-tool run visualization.
*
* A Workflow run persists its state to
* `~/.claude/projects/<projHash>/<sessionUuid>/workflows/wf_<runId>.json`
* (a sibling of the deeper `subagents/workflows/wf_<runId>/agent-*.jsonl`
* transcript tree that subagent-watcher tracks). This file is the single source
* for the master-detail "working agents" view: a run's tasks/phases on the LEFT
* and per-agent stats (tokens burned, tool calls) on the RIGHT.
*
* Field presence is STATE-DRIVEN and verified against real runs on disk:
* - state 'start' (queued): no agentId/tokens/toolCalls/startedAt/durationMs/...
* - state 'progress' (running): has agentId/tokens/toolCalls, no durationMs/resultPreview
* - state 'done' (finished): all fields, incl. durationMs/resultPreview
* Absent fields are genuinely ABSENT (never explicit null) — use `?:`, not null.
*
* @module types/workflow-run
*/
/** One declared phase of a run (from the run JSON's top-level `phases[]`, 0-indexed). */
export interface WorkflowRunPhase {
/** Phase title; equals each member agent's `phaseTitle`. Always present. */
title: string;
/** Human description of the phase. Always present in `phases[]`. */
detail: string;
}
/**
* One agent slot in a run, derived from `workflowProgress[]` entries where
* `type === 'workflow_agent'`. Optional fields are absent until the agent
* reaches the relevant lifecycle state (see module doc).
*/
export interface WorkflowAgentInfo {
/** 1-based stable slot index, unique within the run. Always present. */
index: number;
/** Agent label, e.g. "probe:dompurify-config". Always present. */
label: string;
/** 1-based phase number; join via `run.phases[phaseIndex - 1]`. Always present. */
phaseIndex: number;
/** Phase title (=== run.phases[phaseIndex-1].title). Always present. */
phaseTitle: string;
/** Model id, e.g. "claude-opus-4-8[1m]". Always present. */
model: string;
/** Lifecycle state. Real on-disk values: 'start' | 'progress' | 'done'. Open union. */
state: 'start' | 'progress' | 'done' | (string & {});
/** Epoch ms the slot was queued. Always present. */
queuedAt?: number;
/** Epoch ms of the last progress tick. Always present once any progress occurs. */
lastProgressAt?: number;
/** Truncated prompt the agent was given. Always present. */
promptPreview?: string;
/**
* Globally-unique agent id; equals the `agent-<agentId>.jsonl` transcript stem
* (the Phase-4 correlation key). ABSENT while state === 'start'.
*/
agentId?: string;
/** Epoch ms the agent began. Absent while 'start'. */
startedAt?: number;
/** Attempt counter. Absent while 'start'. */
attempt?: number;
/** Tokens burned so far (RIGHT pane). Absent while 'start'. */
tokens?: number;
/** Tool calls made so far (RIGHT pane). Absent while 'start'. */
toolCalls?: number;
/** Name of the most recent tool. Present for progress/done (occasionally absent). */
lastToolName?: string;
/** Short summary of the most recent tool call. May be absent even when 'done'. */
lastToolSummary?: string;
/** Total run time (ms). Present ONLY when 'done' — the live-vs-finished discriminator. */
durationMs?: number;
/** Truncated final result. Present ONLY when 'done'. */
resultPreview?: string;
}
/**
* Run-level info shipped to the browser.
*
* IMPORTANT: the on-disk JSON also carries `script` (15–660KB of embedded JS),
* `scriptPath`, `result`, and `logs`. The watcher STRIPS all four before the
* object is ever cached/broadcast — never let them reach SSE/getLightState/route.
*/
export interface WorkflowRunInfo {
/** Run id (=== the wf_<runId>.json filename stem). Always present. */
runId: string;
/** Workflow name from `meta.name`. Always present. */
workflowName?: string;
/**
* Run status. Real on-disk values seen: 'completed' | 'killed'.
* 'running' | 'failed' are inferred (parse defensively; keep open union).
*/
status?: 'completed' | 'killed' | 'running' | 'failed' | (string & {});
/** Concise human description (best LEFT-pane label). Always present. */
summary?: string;
/** Total agent slots, INCLUDING not-yet-started 'start' agents. */
agentCount?: number;
/** Total tokens across the run (partial mid-run). */
totalTokens?: number;
/** Total tool calls across the run (partial mid-run). */
totalToolCalls?: number;
/** Total run duration (ms). */
durationMs?: number;
/** Run start time (epoch MILLIS). */
startTime?: number;
/** ISO end/write timestamp. */
timestamp?: string;
/** Default model for the run. */
defaultModel?: string;
/** Background-task id that owns the run. */
taskId?: string;
/** Declared phases (0-indexed). */
phases: WorkflowRunPhase[];
/** Agents, derived from `workflowProgress` filtered to `type === 'workflow_agent'`. */
agents: WorkflowAgentInfo[];
/** Error message, present when status is 'killed'/'failed'. */
error?: string;
// ----- Watcher-derived (NOT in the JSON body — captured from the file path) -----
/** `<sessionUuid>` path segment (for per-session scoping). */
sessionUuid: string;
/** `<projHash>` path segment. */
projectHash: string;
/**
* Most recent activity (epoch ms): max agent `lastProgressAt`, else `startTime`.
* Drives recency filtering/sorting so finished long runs still surface.
*/
lastActivityAt: number;
}
/**
* Lightweight run projection (no `agents[]`) for the LEFT-pane list and the
* getLightState reconnect snapshot. A full run with 28 agents serializes to
* ~36KB; the snapshot ships dozens of runs, so it carries summaries only and the
* RIGHT pane fetches the full run (`GET /api/workflows/:runId`) on selection.
*/
export type WorkflowRunSummary = Omit<WorkflowRunInfo, 'agents'>;
+210
View File
@@ -0,0 +1,210 @@
/**
* @fileoverview Pure detection of Claude Code usage-limit pause messages.
*
* When a Claude subscription limit (5-hour rolling window, weekly, Opus weekly,
* or extra-usage balance) is hit, the Claude Code TUI stops working and prints a
* status line with the reset time. These helpers detect that state in cleaned
* (ANSI-stripped) terminal output and parse the reset time, so the session
* auto-resume feature (SessionAutoOps) can schedule a "continue" nudge.
*
* Message shapes covered (observed across Claude Code 1.0.x–2.1.x, 2025–2026):
* - `5-hour limit reached ∙ resets 8pm` (v1.0.109+ footer)
* - `Session limit reached ∙ resets 8pm`
* - `Weekly limit reached ∙ resets 6pm`
* - `Opus weekly limit reached ∙ resets Oct 6, 1pm`
* - `Limit reached · resets 1pm (America/Chicago) · /upgrade to Max…` (v2.0.55+)
* - `You've hit your limit · resets 1:40pm (America/New_York)` (v2.1.x)
* - `You've hit your weekly limit · resets Mon 12:00am`
* - `You've hit your limit · resets May 5 at 9pm (America/New_York)`
* - `You're out of extra usage · resets 1pm (America/Los_Angeles)`
* - `Claude usage limit reached. Your limit will reset at 2pm (America/New_York)` (v1.0.x inline)
* - `Claude AI usage limit reached|1755309600` (raw API, epoch seconds)
*
* Deliberately conservative: a limit phrase WITHOUT a parseable reset time is
* ignored (returns null) so ordinary conversation text mentioning "limit
* reached" can't arm the scheduler. The downstream retry loop (re-detection
* after each resume attempt) compensates for any parsing imprecision.
*
* All functions are pure (caller passes `now`) for testability.
*
* @module usage-limit-patterns
*/
/** Result of scanning terminal output for a usage-limit pause. */
export interface UsageLimitDetection {
/**
* Epoch ms when the limit resets. May be in the past when the matched
* message is stale (caller should treat past values as "retry soon").
*/
resetAt: number;
/** Matched message snippet (for logging and UI). */
matched: string;
}
/**
* Limit phrases that indicate Claude stopped on a usage limit.
* `\blimit reached` covers all "<X> limit reached" footer variants.
*/
const LIMIT_PHRASE_PATTERN =
/(?:\blimit\s+reached\b|you'?ve\s+hit\s+your\s+(?:\w+\s+)?limit\b|you'?re\s+out\s+of\s+extra\s+usage\b)/gi;
/**
* Reset-time spec following a limit phrase. Captures:
* 1 month (weekly resets >1 day out: "Oct 6, 1pm" / "May 5 at 9pm")
* 2 day-of-month
* 3 day-of-week ("Mon 12:00am")
* 4 hour (12h) 5 minutes 6 am/pm 7 IANA timezone in parens (optional)
* `resets?` + optional `at` also covers the v1.0.x "will reset at 2pm" form.
*/
const RESET_TIME_PATTERN =
/\bresets?\s+(?:at\s+)?(?:(jan|feb|mar|apr|may|jun|jul|aug|sep|oct|nov|dec)[a-z]*\s+(\d{1,2})(?:\s*,\s*|\s+at\s+)|(sun|mon|tue|wed|thu|fri|sat)[a-z]*\s+)?(\d{1,2})(?::(\d{2}))?\s*(am|pm)\b(?:\s*\(([^()\n]{1,64})\))?/i;
/** Raw API form: `Claude AI usage limit reached|1755309600` (epoch seconds). */
const EPOCH_LIMIT_PATTERN = /\busage\s+limit\s+reached\|(\d{9,11})\b/gi;
/** How far after a limit phrase the reset-time spec may appear (chars). */
const RESET_TIME_WINDOW = 160;
/** Parsed reset spec must not be further out than this (weekly max ≈ 7 days). */
const MAX_RESET_HORIZON_MS = 8 * 24 * 60 * 60 * 1000;
const MONTHS = ['jan', 'feb', 'mar', 'apr', 'may', 'jun', 'jul', 'aug', 'sep', 'oct', 'nov', 'dec'];
const WEEKDAYS = ['sun', 'mon', 'tue', 'wed', 'thu', 'fri', 'sat'];
const DAY_MS = 24 * 60 * 60 * 1000;
/**
* Current UTC offset of an IANA timezone in ms, or null if unresolvable
* (e.g. the `(Etc/Unknown)` failure variant Claude Code can print).
* DST transitions inside the wait window can skew the result by an hour;
* the auto-resume retry loop absorbs that.
*/
function zoneOffsetMs(timeZone: string, at: number): number | null {
try {
const dtf = new Intl.DateTimeFormat('en-US', { timeZone, timeZoneName: 'longOffset' });
const name = dtf.formatToParts(at).find((p) => p.type === 'timeZoneName')?.value;
if (!name) return null;
const m = /^GMT(?:([+-])(\d{1,2})(?::(\d{2}))?)?$/.exec(name);
if (!m) return null;
if (!m[1]) return 0; // plain "GMT"
const sign = m[1] === '-' ? -1 : 1;
return sign * (parseInt(m[2], 10) * 60 + (m[3] ? parseInt(m[3], 10) : 0)) * 60_000;
} catch {
return null;
}
}
interface ResetSpec {
month?: number; // 0-11
dayOfMonth?: number; // 1-31
dayOfWeek?: number; // 0-6 (Sun-Sat)
hour: number; // 0-23
minute: number; // 0-59
timeZone?: string;
}
/**
* Compute the epoch ms for a parsed reset spec. Times are wall-clock in the
* given IANA timezone when present (and resolvable), otherwise server-local —
* Claude CLI runs on the same host as Codeman, so local time is the right
* default. Returns null when the spec is implausible (> ~8 days out).
*/
function resolveResetSpec(spec: ResetSpec, now: number): number | null {
const offset = spec.timeZone ? zoneOffsetMs(spec.timeZone, now) : null;
// Wall-clock view of "now": shifted-UTC when a zone offset is known,
// server-local otherwise. Read/build components with the matching API.
const useZone = offset !== null;
const wallNow = useZone ? new Date(now + offset) : new Date(now);
const get = {
year: () => (useZone ? wallNow.getUTCFullYear() : wallNow.getFullYear()),
month: () => (useZone ? wallNow.getUTCMonth() : wallNow.getMonth()),
date: () => (useZone ? wallNow.getUTCDate() : wallNow.getDate()),
day: () => (useZone ? wallNow.getUTCDay() : wallNow.getDay()),
};
const build = (y: number, mo: number, d: number): number => {
const wall = useZone
? Date.UTC(y, mo, d, spec.hour, spec.minute)
: new Date(y, mo, d, spec.hour, spec.minute).getTime();
return useZone ? wall - offset : wall;
};
let ts: number;
if (spec.month !== undefined && spec.dayOfMonth !== undefined) {
// Explicit date ("Oct 6, 1pm"). More than 2 days in the past → assume year
// rollover (message seen near New Year); slightly past → stale, keep as-is.
ts = build(get.year(), spec.month, spec.dayOfMonth);
if (ts < now - 2 * DAY_MS) {
ts = build(get.year() + 1, spec.month, spec.dayOfMonth);
}
} else if (spec.dayOfWeek !== undefined) {
// Day-of-week ("Mon 12:00am") → next occurrence.
const delta = (spec.dayOfWeek - get.day() + 7) % 7;
ts = build(get.year(), get.month(), get.date() + delta);
if (ts <= now) ts += 7 * DAY_MS;
} else {
// Time-only ("resets 8pm") → next occurrence within 24h.
ts = build(get.year(), get.month(), get.date());
if (ts <= now) ts += DAY_MS;
}
if (ts > now + MAX_RESET_HORIZON_MS) return null;
return ts;
}
/** Parse the reset-time spec found within `window`, or null. */
function parseResetTime(window: string, now: number): number | null {
const m = RESET_TIME_PATTERN.exec(window);
if (!m) return null;
const hour12 = parseInt(m[4], 10);
const minute = m[5] ? parseInt(m[5], 10) : 0;
if (hour12 < 1 || hour12 > 12 || minute > 59) return null;
const pm = m[6].toLowerCase() === 'pm';
const hour = (hour12 % 12) + (pm ? 12 : 0);
const spec: ResetSpec = { hour, minute };
if (m[1] && m[2]) {
spec.month = MONTHS.indexOf(m[1].toLowerCase());
spec.dayOfMonth = parseInt(m[2], 10);
if (spec.dayOfMonth < 1 || spec.dayOfMonth > 31) return null;
} else if (m[3]) {
spec.dayOfWeek = WEEKDAYS.indexOf(m[3].toLowerCase());
}
if (m[7]) spec.timeZone = m[7].trim();
return resolveResetSpec(spec, now);
}
/**
* Scan cleaned (ANSI-stripped) terminal output for a usage-limit pause message
* with a parseable reset time. Returns the LAST parseable occurrence in the
* chunk (most recent on screen), or null when none is found.
*/
export function detectUsageLimitPause(cleanData: string, now: number = Date.now()): UsageLimitDetection | null {
if (!cleanData || !/limit|extra usage/i.test(cleanData)) return null;
let result: UsageLimitDetection | null = null;
// Raw API epoch form
EPOCH_LIMIT_PATTERN.lastIndex = 0;
let em: RegExpExecArray | null;
while ((em = EPOCH_LIMIT_PATTERN.exec(cleanData)) !== null) {
const resetAt = parseInt(em[1], 10) * 1000;
if (resetAt > now + MAX_RESET_HORIZON_MS) continue;
result = { resetAt, matched: em[0] };
}
// TUI phrase + "resets <time>" forms
LIMIT_PHRASE_PATTERN.lastIndex = 0;
let pm: RegExpExecArray | null;
while ((pm = LIMIT_PHRASE_PATTERN.exec(cleanData)) !== null) {
const window = cleanData.slice(pm.index, pm.index + RESET_TIME_WINDOW);
const resetAt = parseResetTime(window, now);
if (resetAt !== null) {
result = { resetAt, matched: window.slice(0, 80).trim() };
}
}
return result;
}
+166
View File
@@ -0,0 +1,166 @@
/**
* @fileoverview Pure parsing + formatting of Claude Code statusline telemetry.
*
* Claude Code (v2.1.80+) pipes a JSON blob to a configured `statusLine.command`
* on each render. On Pro/Max subscriptions that blob carries a `rate_limits`
* object with the 5-hour rolling and 7-day weekly plan windows. The
* Codeman-managed statusLine exporter (see `hooks-config.generateStatusLineCommand`)
* POSTs that blob to `/api/status-telemetry`; these helpers normalize the subset
* Codeman displays and format the compact in-terminal footer string.
*
* Confirmed schema (empirically captured, CC 2.1.177, Claude Max — see
* `docs/usage-limits-display-plan.md`):
* rate_limits.{five_hour,seven_day}.{used_percentage: number 0-100,
* resets_at: number EPOCH-SECONDS}
* Only those two windows exist (no Opus-weekly field). `rate_limits` is absent
* before the first API response and for non-subscriber auth — both yield null.
*
* All functions are pure for testability. See `test/usage-telemetry.test.ts`.
*
* @module usage-telemetry
*/
/** A single normalized plan-usage window. */
export interface UsageWindow {
/** Percent of the window consumed, 0–100. */
usedPercentage: number;
/** Epoch MILLISECONDS when the window resets (statusline reports seconds). */
resetAt: number;
}
/** Normalized telemetry Codeman broadcasts to the UI. */
export interface StatusTelemetry {
fiveHour?: UsageWindow;
sevenDay?: UsageWindow;
/** Context-window percent used, 0–100 (bonus field from the same payload). */
contextUsedPercentage?: number;
/** Session cost in USD (bonus field). */
costUsd?: number;
/** Model display name, e.g. "Opus 4.8 (1M context)" (bonus field). */
modelDisplayName?: string;
}
/** Raw subset of the statusline stdin JSON (snake_case, as Claude emits it). */
export interface RawStatuslinePayload {
rate_limits?: {
five_hour?: { used_percentage?: number; resets_at?: number };
seven_day?: { used_percentage?: number; resets_at?: number };
};
context_window?: { used_percentage?: number; total_input_tokens?: number; total_output_tokens?: number };
cost?: { total_cost_usd?: number };
model?: { display_name?: string };
}
function clampPct(n: number): number {
if (!Number.isFinite(n)) return 0;
return Math.max(0, Math.min(100, n));
}
function parseWindow(w?: { used_percentage?: number; resets_at?: number }): UsageWindow | undefined {
if (!w || typeof w.used_percentage !== 'number' || typeof w.resets_at !== 'number') return undefined;
if (!Number.isFinite(w.resets_at) || w.resets_at <= 0) return undefined;
return { usedPercentage: clampPct(w.used_percentage), resetAt: Math.round(w.resets_at * 1000) };
}
/**
* Normalize a raw statusline payload to the telemetry Codeman displays. Returns
* null when there is no plan-limit data to show (pre-first-response or a
* non-subscriber account) so the caller can skip broadcasting.
*/
export function parseStatusTelemetry(data: RawStatuslinePayload | undefined): StatusTelemetry | null {
if (!data) return null;
const fiveHour = parseWindow(data.rate_limits?.five_hour);
const sevenDay = parseWindow(data.rate_limits?.seven_day);
if (!fiveHour && !sevenDay) return null;
const t: StatusTelemetry = {};
if (fiveHour) t.fiveHour = fiveHour;
if (sevenDay) t.sevenDay = sevenDay;
if (typeof data.context_window?.used_percentage === 'number') {
t.contextUsedPercentage = clampPct(data.context_window.used_percentage);
}
if (typeof data.cost?.total_cost_usd === 'number' && Number.isFinite(data.cost.total_cost_usd)) {
t.costUsd = data.cost.total_cost_usd;
}
if (typeof data.model?.display_name === 'string' && data.model.display_name) {
t.modelDisplayName = data.model.display_name.slice(0, 60);
}
return t;
}
/**
* Current-session status for the in-terminal statusline footer. This is the
* "status of the current session" the user sees in Claude's footer — distinct
* from the account-wide plan limits, which live ONLY in the Codeman header chip.
*/
export interface SessionStatus {
modelDisplayName?: string;
inputTokens?: number;
outputTokens?: number;
contextUsedPercentage?: number;
}
/** Group a non-negative integer with thousands separators: 562411 → "562,411". */
function withCommas(n: number): string {
return Math.max(0, Math.round(n))
.toString()
.replace(/\B(?=(\d{3})+(?!\d))/g, ',');
}
/** Extract current-session status (footer) from the raw payload. */
export function parseSessionStatus(data: RawStatuslinePayload | undefined): SessionStatus | null {
if (!data) return null;
const s: SessionStatus = {};
if (typeof data.model?.display_name === 'string' && data.model.display_name) {
s.modelDisplayName = data.model.display_name.slice(0, 60);
}
const cw = data.context_window;
if (typeof cw?.total_input_tokens === 'number' && Number.isFinite(cw.total_input_tokens)) {
s.inputTokens = Math.max(0, cw.total_input_tokens);
}
if (typeof cw?.total_output_tokens === 'number' && Number.isFinite(cw.total_output_tokens)) {
s.outputTokens = Math.max(0, cw.total_output_tokens);
}
if (typeof cw?.used_percentage === 'number') {
s.contextUsedPercentage = clampPct(cw.used_percentage);
}
return Object.keys(s).length ? s : null;
}
/**
* Format the in-terminal statusline footer: the CURRENT SESSION's status —
* `Opus 4.8 (1M context) in:562,411 out:1,188 ctx:56%` — NOT the plan limits,
* which live in the Codeman header chip. Claude requires a statusLine command to
* emit the rate_limits JSON at all, so this is what that command prints back.
*/
export function formatSessionStatusText(s: SessionStatus | null): string {
if (!s) return 'codeman';
const groups: string[] = [];
if (s.modelDisplayName) groups.push(s.modelDisplayName);
const tok: string[] = [];
if (s.inputTokens != null) tok.push(`in:${withCommas(s.inputTokens)}`);
if (s.outputTokens != null) tok.push(`out:${withCommas(s.outputTokens)}`);
if (tok.length) groups.push(tok.join(' '));
if (s.contextUsedPercentage != null) groups.push(`ctx:${Math.round(clampPct(s.contextUsedPercentage))}%`);
return groups.length ? groups.join(' ') : 'codeman';
}
/**
* Stable signature for change-detection — the statusline fires on every
* assistant message, so the route only rebroadcasts when this value changes.
*
* Keys on EXACTLY the values the header chip displays: the two windows' ROUNDED
* percentages (the chip renders `Math.round`) + their reset times. Deliberately
* excludes contextUsedPercentage / costUsd / modelDisplayName — none are shown
* in the chip, and contextUsedPercentage in particular drifts on every assistant
* message, which would defeat the dedup and fan out a redundant SSE broadcast +
* localStorage write + identical chip re-render each time.
*/
export function telemetrySignature(t: StatusTelemetry): string {
return JSON.stringify([
t.fiveHour ? Math.round(t.fiveHour.usedPercentage) : null,
t.fiveHour?.resetAt ?? null,
t.sevenDay ? Math.round(t.sevenDay.usedPercentage) : null,
t.sevenDay?.resetAt ?? null,
]);
}
+41 -1
View File
@@ -8,7 +8,7 @@
* @module utils/claude-cli-resolver
*/
import { execSync } from 'node:child_process';
import { execSync, execFileSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { delimiter, dirname, join } from 'node:path';
import { homedir } from 'node:os';
@@ -83,3 +83,43 @@ export function getAugmentedPath(): string {
_augmentedPath = currentPath;
return _augmentedPath;
}
/** Cached `claude --version` result: string = version, null = probed but unavailable, undefined = not probed */
let _claudeVersion: string | null | undefined = undefined;
/**
* Returns the installed Claude CLI version (e.g. `"2.1.210"`), or null if it
* can't be determined. Runs `claude --version` once and caches the result.
*
* This is a deterministic alternative to scraping the interactive startup
* banner (`parseClaudeCodeInfo` in session.ts): newer Claude Code builds don't
* reliably print `Claude Code vX.Y.Z` at startup, and resumed sessions never
* show it, which left `cliVersion` undefined and silently disabled features
* gated on it (e.g. wheel-forwarding to Claude's transcript — issue #154).
*/
export function getClaudeCliVersion(): string | null {
if (_claudeVersion !== undefined) return _claudeVersion;
// Keep the test suite hermetic — never spawn a real `claude` subprocess under
// vitest (matches IS_TEST_MODE in tmux-manager). Tests that need a version set
// it on the session directly.
if (process.env.VITEST) {
_claudeVersion = null;
return _claudeVersion;
}
try {
const dir = findClaudeDir();
const bin = dir ? join(dir, 'claude') : 'claude';
// execFileSync (no shell) — the resolved path may contain spaces, and there
// is no untrusted input, but avoid a shell either way.
const out = execFileSync(bin, ['--version'], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
env: { ...process.env, PATH: getAugmentedPath() },
});
const match = out.match(/(\d+\.\d+\.\d+)/);
_claudeVersion = match ? match[1] : null;
} catch {
_claudeVersion = null;
}
return _claudeVersion;
}
+211
View File
@@ -0,0 +1,211 @@
/**
* @fileoverview Probe engine for `codeman doctor`. Resolves each registry tool
* against an injectable ProbeHost (real impl uses child_process/fs; tests inject
* fakes) and returns structured results. Pure given the host — no global I/O.
*
* @module utils/dependency-checker
*/
import { execFileSync } from 'node:child_process';
import { existsSync, readdirSync, readFileSync } from 'node:fs';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
import type { ProbeEnvironment, ToolCategory, ToolDependency } from '../config/dependency-registry.js';
export interface EnvDetectionInputs {
platform: NodeJS.Platform;
procVersion: string;
hasWindowsInterop: boolean;
}
export function detectEnvironment(inputs: EnvDetectionInputs): ProbeEnvironment {
if (inputs.platform === 'win32') return 'win32';
if (inputs.platform === 'darwin') return 'darwin';
const isWsl = /microsoft|wsl/i.test(inputs.procVersion) && inputs.hasWindowsInterop;
return isWsl ? 'wsl' : 'linux';
}
const DEFAULT_VERSION_RE = /(\d+\.\d+(?:\.\d+)?)/;
export function extractVersion(text: string, re?: RegExp): string | undefined {
const m = (re ?? DEFAULT_VERSION_RE).exec(text);
return m ? m[1] : undefined;
}
/** Returns -1 if a < b, 0 if equal, 1 if a > b (numeric, component-wise). */
export function compareVersions(a: string, b: string): number {
const pa = a.split('.').map((n) => parseInt(n, 10) || 0);
const pb = b.split('.').map((n) => parseInt(n, 10) || 0);
const len = Math.max(pa.length, pb.length);
for (let i = 0; i < len; i++) {
const d = (pa[i] || 0) - (pb[i] || 0);
if (d !== 0) return d < 0 ? -1 : 1;
}
return 0;
}
export type ToolStatus = 'ok' | 'missing' | 'outdated' | 'skipped' | 'error';
export interface ToolResult {
id: string;
label: string;
category: ToolCategory;
required: boolean;
usedBy: string[];
status: ToolStatus;
version?: string;
path?: string;
installHint?: string;
reason?: string;
}
export interface ProbeHost {
environment: ProbeEnvironment;
which(bin: string): string | null;
fileExists(path: string): boolean;
runVersion(bin: string, args: string[]): string | null;
windowsProgramRoots(): string[];
windowsFileVersion(winPath: string): string | null;
}
function finalize(
base: Pick<ToolResult, 'id' | 'label' | 'category' | 'required' | 'usedBy'>,
tool: ToolDependency,
path: string,
version: string | undefined
): ToolResult {
if (tool.minVersion) {
if (!version) return { ...base, status: 'error', path, reason: 'version required but could not be parsed' };
if (compareVersions(version, tool.minVersion) < 0) return { ...base, status: 'outdated', path, version };
}
return { ...base, status: 'ok', path, version };
}
export function checkTool(tool: ToolDependency, host: ProbeHost): ToolResult {
const base = {
id: tool.id,
label: tool.label,
category: tool.category,
required: tool.required,
usedBy: tool.usedBy ?? [],
};
const installHint = tool.installHint?.[host.environment];
const spec = tool.resolvers.find((r) => r.match.includes(host.environment));
if (!spec) return { ...base, status: 'skipped', reason: `not applicable on ${host.environment}` };
if (spec.resolver.kind === 'path') {
const { bins, versionArg, versionRegex } = spec.resolver;
for (const bin of bins) {
const resolved = host.which(bin);
if (resolved) {
const out = host.runVersion(bin, [versionArg ?? '--version']);
const version = out ? extractVersion(out, versionRegex) : undefined;
return finalize(base, tool, resolved, version);
}
}
return { ...base, status: 'missing', installHint };
}
// windows-side
const { appDirs, exes } = spec.resolver;
for (const root of host.windowsProgramRoots()) {
for (const dir of appDirs) {
for (const exe of exes) {
const winPath = `${root}/${dir}/${exe}`;
if (host.fileExists(winPath)) {
const raw = host.windowsFileVersion(winPath);
const version = raw ? extractVersion(raw) : undefined;
return finalize(base, tool, winPath, version);
}
}
}
}
return { ...base, status: 'missing', installHint };
}
export function checkAll(registry: ToolDependency[], host: ProbeHost): ToolResult[] {
return registry.map((tool) => checkTool(tool, host));
}
function safeWhich(bin: string): string | null {
try {
const out = execFileSync(process.platform === 'win32' ? 'where' : 'which', [bin], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
const first = out.split(/\r?\n/)[0]?.trim();
return first && existsSync(first) ? first : null;
} catch {
return null;
}
}
function safeRunVersion(bin: string, args: string[]): string | null {
try {
return execFileSync(bin, args, {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
stdio: ['ignore', 'pipe', 'ignore'],
});
} catch (err: unknown) {
// Some tools (e.g. ffmpeg) exit non-zero on -version but still print to stdout
const stdout = (err as { stdout?: Buffer | string })?.stdout;
return stdout ? stdout.toString() : null;
}
}
function readProcVersion(): string {
try {
return readFileSync('/proc/version', 'utf-8');
} catch {
return '';
}
}
function listWindowsProgramRoots(): string[] {
const roots: string[] = [];
try {
for (const entry of readdirSync('/mnt')) {
for (const pf of ['Program Files', 'Program Files (x86)']) {
const root = `/mnt/${entry}/${pf}`;
if (existsSync(root)) roots.push(root);
}
}
} catch {
// /mnt absent (not WSL) -> no roots
}
return roots;
}
function readWindowsFileVersion(winPath: string): string | null {
try {
const windowsPath = execFileSync('wslpath', ['-w', winPath], {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
const out = execFileSync(
'powershell.exe',
['-NoProfile', '-Command', `(Get-Item '${windowsPath.replace(/'/g, "''")}').VersionInfo.ProductVersion`],
{ encoding: 'utf-8', timeout: EXEC_TIMEOUT_MS }
).trim();
return out || null;
} catch {
return null;
}
}
export function createRealHost(): ProbeHost {
const environment = detectEnvironment({
platform: process.platform,
procVersion: readProcVersion(),
hasWindowsInterop: safeWhich('cmd.exe') !== null || safeWhich('powershell.exe') !== null,
});
return {
environment,
which: safeWhich,
fileExists: existsSync,
runVersion: safeRunVersion,
windowsProgramRoots: listWindowsProgramRoots,
windowsFileVersion: readWindowsFileVersion,
};
}
+75
View File
@@ -0,0 +1,75 @@
/**
* @fileoverview Renders ToolResult[] from the dependency checker into a
* human-readable grouped table or JSON, and computes the process exit code.
* Plain text only (no color) so output is stable and snapshot-friendly; the
* CLI layer may colorize.
*
* @module utils/dependency-report
*/
import type { ProbeEnvironment, ToolCategory } from '../config/dependency-registry.js';
import type { ToolResult, ToolStatus } from './dependency-checker.js';
const CATEGORY_ORDER: ToolCategory[] = ['core', 'office', 'other'];
function glyph(r: ToolResult): string {
if (r.status === 'ok') return '✓';
if (r.status === 'skipped') return '○';
return r.required ? '✗' : '○';
}
function statusText(r: ToolResult): string {
if (r.status === 'ok') return r.version ?? 'installed';
if (r.status === 'outdated') return `${r.version ?? '?'} (below minimum)`;
if (r.status === 'skipped') return 'n/a';
if (r.status === 'error') return 'version error';
return 'not found';
}
export function computeExitCode(results: ToolResult[]): number {
const failed = results.some(
(r) => r.required && (r.status === 'missing' || r.status === 'outdated' || r.status === 'error')
);
return failed ? 1 : 0;
}
export function renderTable(results: ToolResult[], environment: ProbeEnvironment): string {
const lines: string[] = [`Codeman dependency check — ${environment}`, ''];
for (const category of CATEGORY_ORDER) {
const rows = results.filter((r) => r.category === category);
if (rows.length === 0) continue;
lines.push(category.toUpperCase());
for (const r of rows) {
const detail = r.path ? ` ${r.path}` : '';
lines.push(` ${glyph(r)} ${r.label.padEnd(14)} ${statusText(r).padEnd(22)}${detail}`);
if (r.usedBy.length) lines.push(` used by: ${r.usedBy.join(', ')}`);
if (r.installHint) lines.push(` install: ${r.installHint}`);
}
lines.push('');
}
const ok = results.filter((r) => r.status === 'ok').length;
const requiredMissing = results.filter((r) => r.required && r.status !== 'ok' && r.status !== 'skipped').length;
const optionalMissing = results.filter((r) => !r.required && r.status === 'missing').length;
lines.push(`Summary: ${ok} ok · ${requiredMissing} required missing · ${optionalMissing} optional missing`);
return lines.join('\n');
}
export interface DependencyReportJson {
platform: { environment: ProbeEnvironment };
summary: { ok: number; requiredMissing: number; optionalMissing: number; exitCode: number };
tools: ToolResult[];
}
export function renderJson(results: ToolResult[], environment: ProbeEnvironment): DependencyReportJson {
const byStatus = (s: ToolStatus) => results.filter((r) => r.status === s).length;
return {
platform: { environment },
summary: {
ok: byStatus('ok'),
requiredMissing: results.filter((r) => r.required && r.status !== 'ok' && r.status !== 'skipped').length,
optionalMissing: results.filter((r) => !r.required && r.status === 'missing').length,
exitCode: computeExitCode(results),
},
tools: results,
};
}
+67
View File
@@ -0,0 +1,67 @@
/**
* @fileoverview Resolve the Gemini CLI binary across common install paths.
*
* Mirrors codex-cli-resolver.ts and opencode-cli-resolver.ts. Finds the
* `gemini` binary and provides an augmented PATH directory for tmux sessions.
*
* @module utils/gemini-cli-resolver
*/
import { execSync } from 'node:child_process';
import { existsSync } from 'node:fs';
import { dirname, join } from 'node:path';
import { homedir } from 'node:os';
import { EXEC_TIMEOUT_MS } from '../config/exec-timeout.js';
/** Common directories where the Gemini CLI binary may be installed */
const GEMINI_SEARCH_DIRS = [
join(homedir(), '.gemini', 'bin'),
join(homedir(), '.local', 'bin'),
'/usr/local/bin',
join(homedir(), '.bun', 'bin'),
join(homedir(), '.npm-global', 'bin'),
join(homedir(), 'bin'),
];
/** Cached directory containing the gemini binary (empty string = searched but not found) */
let _geminiDir: string | null = null;
/**
* Finds the directory containing the `gemini` binary.
* Checks `which gemini` first, then falls back to common install locations.
*
* @returns Directory path, or null if not found
*/
export function resolveGeminiDir(): string | null {
if (_geminiDir !== null) return _geminiDir || null;
try {
const result = execSync('which gemini', {
encoding: 'utf-8',
timeout: EXEC_TIMEOUT_MS,
}).trim();
if (result && existsSync(result)) {
_geminiDir = dirname(result);
return _geminiDir;
}
} catch {
// Gemini not in PATH, will check common locations
}
for (const dir of GEMINI_SEARCH_DIRS) {
if (existsSync(join(dir, 'gemini'))) {
_geminiDir = dir;
return _geminiDir;
}
}
_geminiDir = '';
return null;
}
/**
* Check if Gemini CLI is available on the system.
*/
export function isGeminiAvailable(): boolean {
return resolveGeminiDir() !== null;
}
+2 -1
View File
@@ -26,6 +26,7 @@ export { isSafePushEndpoint } from './push-endpoint-validation.js';
export { stringSimilarity, fuzzyPhraseMatch, todoContentHash } from './string-similarity.js';
export { assertNever } from './type-safety.js';
export { wrapWithNice } from './nice-wrapper.js';
export { findClaudeDir, getAugmentedPath } from './claude-cli-resolver.js';
export { findClaudeDir, getAugmentedPath, getClaudeCliVersion } from './claude-cli-resolver.js';
export { resolveOpenCodeDir } from './opencode-cli-resolver.js';
export { resolveCodexDir, isCodexAvailable } from './codex-cli-resolver.js';
export { resolveGeminiDir, isGeminiAvailable } from './gemini-cli-resolver.js';
+379
View File
@@ -0,0 +1,379 @@
import type { LifecycleEntry, RunSummary, RunSummaryEvent, TokenUsageEntry } from '../types.js';
export type AwayDigestRangeName = 'since-last-visit' | '1h' | 'today' | '24h' | 'custom';
export type AwayDigestCategory = 'needs_attention' | 'completed' | 'still_running' | 'idle' | 'informational';
export type AwayDigestSectionName = 'needsAttention' | 'completed' | 'stillRunning' | 'idle' | 'informational';
export type AwayDigestSeverity = 'info' | 'success' | 'warning' | 'error';
export type AwayDigestSource = 'lifecycle' | 'run_summary' | 'status' | 'token_stats' | 'subagent';
export type AwayDigestTokenWindowPrecision = 'day' | 'none';
const HOUR_MS = 60 * 60 * 1000;
const DAY_MS = 24 * HOUR_MS;
const VALID_RANGES = new Set<AwayDigestRangeName>(['since-last-visit', '1h', 'today', '24h', 'custom']);
export interface AwayDigestRange {
range: AwayDigestRangeName;
since: number;
until: number;
}
export interface AwayDigestRangeInput {
range?: string;
since?: number;
until?: number;
lastViewed?: number;
now?: number;
}
export interface AwayDigestSession {
id: string;
name?: string;
status?: string;
inputTokens?: number;
outputTokens?: number;
totalCost?: number;
}
export interface AwayDigestSubagent {
id?: string;
agentId?: string;
sessionId?: string;
description?: string;
status?: string;
lastUpdated?: number;
updatedAt?: number;
completedAt?: number;
modifiedAt?: number;
lastActivityAt?: number;
}
export interface AwayDigestItem {
id: string;
sessionId?: string;
sessionName?: string;
timestamp: number;
category: AwayDigestCategory;
severity: AwayDigestSeverity;
title: string;
detail?: string;
source: AwayDigestSource;
link?: {
type: 'session' | 'run_summary' | 'lifecycle' | 'notification';
sessionId?: string;
};
}
export interface AwayDigestTotals {
sessionsCreated: number;
sessionsExited: number;
activeSessions: number;
needsAttention: number;
completed: number;
errors: number;
warnings: number;
inputTokens?: number;
outputTokens?: number;
estimatedCost?: number;
tokenWindowPrecision: AwayDigestTokenWindowPrecision;
}
export interface AwayDigestResponse {
range: AwayDigestRange;
generatedAt: number;
dataFreshness: {
lifecyclePersisted: true;
tokenStatsPersisted: true;
runSummariesLiveOnly: true;
subagentsLiveOnly: true;
};
totals: AwayDigestTotals;
sections: Record<AwayDigestSectionName, AwayDigestItem[]>;
}
export interface AwayDigestInput {
range: AwayDigestRange;
lifecycleEntries: LifecycleEntry[];
runSummaries: RunSummary[];
sessions: AwayDigestSession[];
dailyTokenStats: TokenUsageEntry[];
subagents: AwayDigestSubagent[];
now?: number;
}
export function resolveAwayDigestRange(input: AwayDigestRangeInput): AwayDigestRange {
const now = input.now ?? Date.now();
const range = (input.range ?? 'since-last-visit') as AwayDigestRangeName;
if (!VALID_RANGES.has(range)) {
throw new Error(`Invalid away digest range: ${input.range}`);
}
let since: number;
const until = finiteOrDefault(input.until, now);
switch (range) {
case 'since-last-visit':
since = finiteOrDefault(input.lastViewed, now - DAY_MS);
break;
case '1h':
since = now - HOUR_MS;
break;
case 'today': {
const start = new Date(now);
start.setHours(0, 0, 0, 0);
since = start.getTime();
break;
}
case '24h':
since = now - DAY_MS;
break;
case 'custom':
if (!Number.isFinite(input.since)) {
throw new Error('Custom away digest range requires a finite since timestamp');
}
since = input.since as number;
break;
}
if (until < since) {
throw new Error('Away digest until timestamp must be greater than or equal to since');
}
return { range, since, until };
}
export function buildAwayDigest(input: AwayDigestInput): AwayDigestResponse {
const now = input.now ?? Date.now();
const sections: Record<AwayDigestSectionName, AwayDigestItem[]> = {
needsAttention: [],
completed: [],
stillRunning: [],
idle: [],
informational: [],
};
const sessionsById = new Map(input.sessions.map((session) => [session.id, session]));
const lifecycleEntries = input.lifecycleEntries.filter((entry) => isInRange(entry.ts, input.range));
for (const entry of lifecycleEntries) {
addItem(sections, lifecycleEntryToItem(entry));
}
for (const summary of input.runSummaries) {
for (const event of summary.events) {
if (!isInRange(event.timestamp, input.range)) continue;
addItem(sections, runSummaryEventToItem(summary, event));
}
}
for (const session of input.sessions) {
const item = sessionToItem(session, now);
addItem(sections, item);
}
for (const subagent of input.subagents) {
const timestamp = subagentTimestamp(subagent, now);
if (!isInRange(timestamp, input.range) || subagent.status !== 'completed') continue;
addItem(sections, subagentToItem(subagent, sessionsById, timestamp));
}
const tokenTotals = aggregateTokenStats(input.dailyTokenStats, input.range);
const totals = calculateTotals(sections, lifecycleEntries, input.sessions, tokenTotals);
return {
range: input.range,
generatedAt: now,
dataFreshness: {
lifecyclePersisted: true,
tokenStatsPersisted: true,
runSummariesLiveOnly: true,
subagentsLiveOnly: true,
},
totals,
sections,
};
}
function finiteOrDefault(value: number | undefined, fallback: number): number {
return Number.isFinite(value) ? (value as number) : fallback;
}
function isInRange(timestamp: number, range: AwayDigestRange): boolean {
return timestamp >= range.since && timestamp <= range.until;
}
function addItem(sections: Record<AwayDigestSectionName, AwayDigestItem[]>, item: AwayDigestItem): void {
sections[sectionNameForCategory(item.category)].push(item);
}
function sectionNameForCategory(category: AwayDigestCategory): AwayDigestSectionName {
switch (category) {
case 'needs_attention':
return 'needsAttention';
case 'still_running':
return 'stillRunning';
case 'completed':
case 'idle':
case 'informational':
return category;
}
}
function lifecycleEntryToItem(entry: LifecycleEntry): AwayDigestItem {
const needsAttention = entry.event === 'mux_died' || (entry.event === 'exit' && (entry.exitCode ?? 0) !== 0);
return {
id: `lifecycle-${entry.ts}-${entry.event}-${entry.sessionId}`,
sessionId: entry.sessionId,
sessionName: entry.name,
timestamp: entry.ts,
category: needsAttention ? 'needs_attention' : 'informational',
severity: needsAttention ? 'error' : entry.event === 'exit' ? 'info' : 'info',
title: lifecycleTitle(entry),
detail: lifecycleDetail(entry),
source: 'lifecycle',
link: { type: 'lifecycle', sessionId: entry.sessionId },
};
}
function lifecycleTitle(entry: LifecycleEntry): string {
if (entry.event === 'exit') {
return (entry.exitCode ?? 0) === 0 ? 'Session exited' : 'Session exited with error';
}
if (entry.event === 'mux_died') return 'Tmux session died';
return `Session ${entry.event.replaceAll('_', ' ')}`;
}
function lifecycleDetail(entry: LifecycleEntry): string | undefined {
if (entry.reason) return entry.reason;
if (entry.event === 'exit' && entry.exitCode !== undefined && entry.exitCode !== null) {
return `Exit code ${entry.exitCode}`;
}
return undefined;
}
function runSummaryEventToItem(summary: RunSummary, event: RunSummaryEvent): AwayDigestItem {
const category = runSummaryCategory(event);
return {
id: `run-summary-${summary.sessionId}-${event.id}`,
sessionId: summary.sessionId,
sessionName: summary.sessionName,
timestamp: event.timestamp,
category,
severity: runSummarySeverity(event, category),
title: event.title,
detail: event.details,
source: 'run_summary',
link: { type: 'run_summary', sessionId: summary.sessionId },
};
}
function runSummaryCategory(event: RunSummaryEvent): AwayDigestCategory {
if (event.type === 'ralph_completion') return 'completed';
if (event.severity === 'error' || event.severity === 'warning' || event.type === 'state_stuck') {
return 'needs_attention';
}
return 'informational';
}
function runSummarySeverity(event: RunSummaryEvent, category: AwayDigestCategory): AwayDigestSeverity {
if (category === 'completed') return 'success';
return event.severity;
}
function sessionToItem(session: AwayDigestSession, now: number): AwayDigestItem {
const isIdle = session.status === 'idle';
return {
id: `status-${session.id}`,
sessionId: session.id,
sessionName: session.name,
timestamp: now,
category: isIdle ? 'idle' : 'still_running',
severity: isIdle ? 'info' : 'success',
title: isIdle ? 'Session idle' : 'Session still running',
detail: session.status ? `Status: ${session.status}` : undefined,
source: 'status',
link: { type: 'session', sessionId: session.id },
};
}
function subagentToItem(
subagent: AwayDigestSubagent,
sessionsById: Map<string, AwayDigestSession>,
timestamp: number
): AwayDigestItem {
const session = subagent.sessionId ? sessionsById.get(subagent.sessionId) : undefined;
const agentId = subagent.id ?? subagent.agentId ?? 'unknown';
return {
id: `subagent-${agentId}`,
sessionId: subagent.sessionId,
sessionName: session?.name,
timestamp,
category: 'informational',
severity: 'success',
title: 'Subagent completed',
detail: subagent.description,
source: 'subagent',
link: subagent.sessionId ? { type: 'session', sessionId: subagent.sessionId } : undefined,
};
}
function subagentTimestamp(subagent: AwayDigestSubagent, fallback: number): number {
return (
subagent.completedAt ??
subagent.lastUpdated ??
subagent.updatedAt ??
subagent.modifiedAt ??
subagent.lastActivityAt ??
fallback
);
}
function aggregateTokenStats(
dailyTokenStats: TokenUsageEntry[],
range: AwayDigestRange
): { inputTokens: number; outputTokens: number; estimatedCost: number; precision: AwayDigestTokenWindowPrecision } {
let inputTokens = 0;
let outputTokens = 0;
let estimatedCost = 0;
for (const day of dailyTokenStats) {
if (!dayOverlapsRange(day.date, range)) continue;
inputTokens += day.inputTokens;
outputTokens += day.outputTokens;
estimatedCost += day.estimatedCost;
}
return {
inputTokens,
outputTokens,
estimatedCost,
precision: inputTokens > 0 || outputTokens > 0 || estimatedCost > 0 ? 'day' : 'none',
};
}
function dayOverlapsRange(date: string, range: AwayDigestRange): boolean {
const dayStart = new Date(`${date}T00:00:00`).getTime();
const dayEnd = dayStart + DAY_MS - 1;
return dayStart <= range.until && dayEnd >= range.since;
}
function calculateTotals(
sections: Record<AwayDigestSectionName, AwayDigestItem[]>,
lifecycleEntries: LifecycleEntry[],
sessions: AwayDigestSession[],
tokenTotals: ReturnType<typeof aggregateTokenStats>
): AwayDigestTotals {
const allItems = Object.values(sections).flat();
return {
sessionsCreated: lifecycleEntries.filter((entry) => entry.event === 'created').length,
sessionsExited: lifecycleEntries.filter((entry) => entry.event === 'exit').length,
activeSessions: sessions.length,
needsAttention: sections.needsAttention.length,
completed: sections.completed.length,
errors: allItems.filter((item) => item.severity === 'error').length,
warnings: allItems.filter((item) => item.severity === 'warning').length,
inputTokens: tokenTotals.inputTokens,
outputTokens: tokenTotals.outputTokens,
estimatedCost: tokenTotals.estimatedCost,
tokenWindowPrecision: tokenTotals.precision,
};
}
+82
View File
@@ -0,0 +1,82 @@
/**
* @fileoverview Main-thread wrapper for HEIC/HEIF → JPEG conversion.
*
* The actual decode/encode (`heic-jpeg-worker.ts`) is CPU-synchronous WASM + JS,
* so it runs in a dedicated `worker_threads` Worker per conversion — never on
* the event loop that serves every session's SSE/PTY/WS traffic. On top of
* the worker isolation this wrapper enforces:
* - the global converter concurrency cap (`runWithConversionLimit`, shared
* with the pdftoppm/soffice document converters) so N simultaneous uploads
* can't pin N cores / N × 256MB decode buffers at once;
* - a hard timeout that terminates the worker (a wedged WASM decode can't be
* cancelled cooperatively);
* - the paste-image size cap on the *output* — jpeg-js is a far less
* efficient encoder than HEVC, so a within-limit HEIC can inflate past
* MAX_PASTE_IMAGE_BYTES.
*/
import { Worker } from 'node:worker_threads';
import { runWithConversionLimit } from '../document-conversion-limiter.js';
import { MAX_PASTE_IMAGE_BYTES } from '../config/buffer-limits.js';
import { HEIC_JPEG_QUALITY, type HeicWorkerInput, type HeicWorkerResult } from './heic-jpeg-worker.js';
/** Hard cap on a single conversion; the worker is terminated when it fires. */
export const HEIC_CONVERSION_TIMEOUT_MS = 30_000;
// V8-heap guardrails for the conversion worker — defense in depth only: large
// TypedArray/WASM backing stores are external to the V8 heap, so the real
// memory bound is the 64MP dimension pre-check in heic-jpeg-worker.ts.
const WORKER_RESOURCE_LIMITS = { maxOldGenerationSizeMb: 1024, maxYoungGenerationSizeMb: 128, stackSizeMb: 8 };
function workerUrl(): URL {
// Compiled installs run the tsc-emitted .js sibling in dist/; dev under tsx
// runs the .ts source directly (tsx's loader propagates to worker threads).
const file = import.meta.url.endsWith('.ts') ? './heic-jpeg-worker.ts' : './heic-jpeg-worker.js';
return new URL(file, import.meta.url);
}
/**
* Convert HEIC/HEIF bytes to JPEG bytes off-thread. Rejects on invalid input,
* over-limit dimensions, oversized output, timeout, or worker failure.
*/
export async function convertHeicToJpeg(imageBytes: Buffer): Promise<Buffer> {
return runWithConversionLimit(
() =>
new Promise<Buffer>((resolve, reject) => {
const worker = new Worker(workerUrl(), {
workerData: { heicInput: imageBytes, quality: HEIC_JPEG_QUALITY } satisfies HeicWorkerInput,
resourceLimits: WORKER_RESOURCE_LIMITS,
});
let settled = false;
const settle = (fn: () => void): void => {
if (settled) return;
settled = true;
clearTimeout(timer);
fn();
void worker.terminate();
};
const timer = setTimeout(() => {
settle(() => reject(new Error(`HEIC conversion timed out after ${HEIC_CONVERSION_TIMEOUT_MS}ms`)));
}, HEIC_CONVERSION_TIMEOUT_MS);
worker.on('message', (msg: HeicWorkerResult) => {
settle(() => {
if (!msg.ok) {
reject(new Error(msg.error));
return;
}
const out = Buffer.from(msg.data.buffer, msg.data.byteOffset, msg.data.byteLength);
if (out.length > MAX_PASTE_IMAGE_BYTES) {
const maxMb = Math.round(MAX_PASTE_IMAGE_BYTES / (1024 * 1024));
reject(new Error(`converted JPEG (${out.length} bytes) exceeds the ${maxMb}MB upload limit`));
return;
}
resolve(out);
});
});
worker.on('error', (err) => settle(() => reject(err)));
worker.on('exit', (code) => {
settle(() => reject(new Error(`HEIC conversion worker exited unexpectedly (code ${code})`)));
});
})
);
}
+93
View File
@@ -0,0 +1,93 @@
/**
* @fileoverview HEIC/HEIF → JPEG conversion core + worker-thread entry.
*
* Spawned per conversion by `heic-jpeg-converter.ts` so the CPU-synchronous
* libheif WASM decode + jpeg-js encode never run on the server's main thread
* (on the event loop they would freeze every session's SSE/PTY/WS handling
* for seconds per photo). Input arrives via `workerData`; the result (or
* error message) is posted back as a single message and the thread exits.
*
* The conversion logic lives in this same file (exported, guarded bootstrap)
* rather than a sibling module: the worker runs from `.ts` source under tsx
* in dev, where relative `.js` imports don't resolve inside worker threads —
* only `node:` builtins are imported at top level. Unit tests import
* `convertHeicBufferToJpeg` directly; the bootstrap only runs when spawned
* with our `workerData` shape.
*
* Decompression-bomb guard: heic-decode's `.all` path exposes the
* header-declared {width, height} per image WITHOUT decoding pixels, while
* its plain decode path allocates `width * height * 4` bytes straight from
* those header values — a <1KB crafted file declaring 30000×30000 would
* demand a 3.6GB allocation. We reject anything above MAX_HEIC_DECODE_PIXELS
* before calling `decode()`.
*/
import { parentPort, workerData } from 'node:worker_threads';
/** Max header-declared pixel count we will decode (64MP ≈ 256MB RGBA). */
export const MAX_HEIC_DECODE_PIXELS = 64_000_000;
/** JPEG quality used for converted HEIC uploads (matches heic-convert's default). */
export const HEIC_JPEG_QUALITY = 0.92;
export interface HeicWorkerInput {
heicInput: Uint8Array;
quality: number;
}
export type HeicWorkerResult = { ok: true; data: Uint8Array } | { ok: false; error: string };
/**
* Convert HEIC/HEIF bytes to JPEG bytes. Throws on non-HEIC input, empty
* containers, over-limit dimensions, and non-JPEG encoder output.
*/
export async function convertHeicBufferToJpeg(input: Uint8Array, quality: number = HEIC_JPEG_QUALITY): Promise<Buffer> {
const { default: decode } = await import('heic-decode');
const buffer = Buffer.isBuffer(input) ? input : Buffer.from(input.buffer, input.byteOffset, input.byteLength);
const images = await decode.all({ buffer });
try {
if (images.length === 0) throw new Error('no image found in HEIC container');
const { width, height } = images[0];
if (
!Number.isSafeInteger(width) ||
!Number.isSafeInteger(height) ||
width <= 0 ||
height <= 0 ||
width * height > MAX_HEIC_DECODE_PIXELS
) {
throw new Error(
`HEIC dimensions ${width}x${height} exceed the ${Math.floor(MAX_HEIC_DECODE_PIXELS / 1_000_000)}MP decode limit`
);
}
const decoded = await images[0].decode();
const { encode } = await import('jpeg-js');
// Same output path as heic-convert's JPEG format (jpeg-js at quality*100).
const jpeg = encode(
{ data: decoded.data, width: decoded.width, height: decoded.height },
Math.floor(quality * 100)
).data;
const jpegBytes = Buffer.isBuffer(jpeg) ? jpeg : Buffer.from(jpeg);
if (jpegBytes.length < 3 || jpegBytes[0] !== 0xff || jpegBytes[1] !== 0xd8 || jpegBytes[2] !== 0xff) {
throw new Error('HEIC conversion did not produce JPEG bytes');
}
return jpegBytes;
} finally {
images.dispose();
}
}
// ── Worker bootstrap ──────────────────────────────────────────────────────
// Runs only when spawned by heic-jpeg-converter.ts: requires a parent port
// AND our exact workerData shape, so importing this module from the main
// thread (or a test runner's own worker pool) stays inert.
const request = workerData as HeicWorkerInput | null | undefined;
if (parentPort && request && request.heicInput instanceof Uint8Array && typeof request.quality === 'number') {
const port = parentPort;
try {
const jpegBytes = await convertHeicBufferToJpeg(request.heicInput, request.quality);
port.postMessage({ ok: true, data: jpegBytes } satisfies HeicWorkerResult);
} catch (err: unknown) {
const error = err instanceof Error ? err.message : String(err);
port.postMessage({ ok: false, error } satisfies HeicWorkerResult);
}
}
+76 -9
View File
@@ -19,6 +19,7 @@ import {
AUTH_FAILURE_MAX,
AUTH_FAILURE_WINDOW_MS,
} from '../../config/auth-config.js';
import { getHookSecret, HOOK_SECRET_HEADER } from '../../config/hook-secret.js';
// Auth session cookie name
export const AUTH_COOKIE_NAME = 'codeman_session';
@@ -28,12 +29,16 @@ interface AuthState {
authSessions: StaleExpirationMap<string, AuthSessionRecord> | null;
authFailures: StaleExpirationMap<string, number> | null;
qrAuthFailures: StaleExpirationMap<string, number> | null;
hookSecretFailures: StaleExpirationMap<string, number> | null;
}
/**
* Register HTTP Basic Auth middleware with session cookies and rate limiting.
* Only active when CODEMAN_PASSWORD is set.
*
* The `/api/hook-event` + `/api/status-telemetry` localhost bypass requires the
* shared hook secret unconditionally (COD-91) — see the onRequest hook below.
*
* @returns AuthState for lifecycle management (dispose on server stop)
*/
export function registerAuthMiddleware(app: FastifyInstance, https: boolean): AuthState {
@@ -41,6 +46,7 @@ export function registerAuthMiddleware(app: FastifyInstance, https: boolean): Au
authSessions: null,
authFailures: null,
qrAuthFailures: null,
hookSecretFailures: null,
};
const authPassword = process.env.CODEMAN_PASSWORD;
@@ -67,24 +73,67 @@ export function registerAuthMiddleware(app: FastifyInstance, https: boolean): Au
refreshOnGet: false,
});
// Separate hook-secret failure counter (COD-54). MUST NOT share authFailures:
// legacy (pre-secret) hook configs fire constantly from 127.0.0.1, and counting
// their 401s against the shared bucket would 429 every cookie-less request from
// loopback — locking out the Basic-Auth login path (and, through a tunnel, every
// client, since tunneled traffic also arrives as 127.0.0.1).
state.hookSecretFailures = new StaleExpirationMap<string, number>({
ttlMs: AUTH_FAILURE_WINDOW_MS,
refreshOnGet: false,
});
const authSessions = state.authSessions;
const authFailures = state.authFailures;
const hookSecretFailures = state.hookSecretFailures;
function sendAuthRateLimit(reply: FastifyReply, clientIp: string): void {
const remainingMs = authFailures.getRemainingTtl(clientIp) ?? AUTH_FAILURE_WINDOW_MS;
function sendAuthRateLimit(
reply: FastifyReply,
clientIp: string,
failures: StaleExpirationMap<string, number> = authFailures
): void {
const remainingMs = failures.getRemainingTtl(clientIp) ?? AUTH_FAILURE_WINDOW_MS;
const retryAfterSeconds = Math.max(1, Math.ceil(remainingMs / 1000));
reply.header('Retry-After', String(retryAfterSeconds));
reply.code(429).send('Too Many Requests — try again later');
}
app.addHook('onRequest', (req, reply, done) => {
// Hook events come from local Claude Code hooks (curl from localhost) — no auth headers available.
// Safe: validated by HookEventSchema, only triggers broadcasts.
// Security: restrict bypass to localhost only — prevents forged hook events via tunnel/LAN.
if (req.url === '/api/hook-event' && req.method === 'POST') {
// Hook events + statusline telemetry come from local Claude Code (curl from
// localhost) — no Basic-Auth credentials available. Validated downstream by
// HookEventSchema / StatusTelemetrySchema. Same loopback+hook-secret gate.
//
// COD-54: the bare localhost bypass is unsafe while a tunnel is running, because
// `cloudflared --url http://127.0.0.1:port` proxies internet traffic INTO the
// loopback origin, so a tunneled request arrives with req.ip === 127.0.0.1 and
// would pass. COD-91: require the shared hook secret on the loopback bypass
// UNCONDITIONALLY (not just while the managed tunnel is up). Codeman can't detect
// a user's own loopback reverse proxy (their own `cloudflared --url`, `tailscale
// serve`, nginx → 127.0.0.1), so tunnel-gating left that path with the unsafe plain
// bypass. Managed-session hooks always present the secret (X-Codeman-Hook-Secret,
// from $CODEMAN_HOOK_SECRET_FILE — generated for every instance), so requiring it
// always closes the gap without breaking the legitimate hook channel.
if ((req.url === '/api/hook-event' || req.url === '/api/status-telemetry') && req.method === 'POST') {
const ip = req.ip;
if (ip === '127.0.0.1' || ip === '::1' || ip === '::ffff:127.0.0.1') {
done();
const isLoopback = ip === '127.0.0.1' || ip === '::1' || ip === '::ffff:127.0.0.1';
if (isLoopback) {
// Always require the shared secret (constant-time compare).
const presented = Buffer.from(req.headers[HOOK_SECRET_HEADER.toLowerCase()]?.toString() ?? '');
const expected = Buffer.from(getHookSecret());
if (presented.length === expected.length && timingSafeEqual(presented, expected)) {
done();
return;
}
// Wrong/absent secret — rate-limit per IP in the DEDICATED hook bucket
// (never authFailures, which would lock out the login path).
const hookIp = req.ip;
const hookFailures = hookSecretFailures.get(hookIp) ?? 0;
if (hookFailures >= AUTH_FAILURE_MAX) {
sendAuthRateLimit(reply, hookIp, hookSecretFailures);
return;
}
hookSecretFailures.set(hookIp, hookFailures + 1);
reply.code(401).send('Unauthorized: hook secret required');
return;
}
// Non-localhost hook requests fall through to normal auth
@@ -102,6 +151,19 @@ export function registerAuthMiddleware(app: FastifyInstance, https: boolean): Au
// Use get() instead of has() so refreshOnGet extends the TTL on active sessions
const sessionToken = req.cookies[AUTH_COOKIE_NAME];
if (sessionToken && authSessions.get(sessionToken) !== undefined) {
// Sliding cookie: re-issue on every authenticated request so the browser
// cookie lifetime tracks the server-side sliding TTL (refreshOnGet above).
// Without this the cookie has a fixed lifetime from login; the browser
// drops it mid-use, the next request arrives cookie-less and falls through
// to Basic Auth — popping the native username/password dialog, which reads
// as a random logout while actively working.
reply.setCookie(AUTH_COOKIE_NAME, sessionToken, {
httpOnly: true,
secure: https,
sameSite: 'lax',
maxAge: AUTH_SESSION_TTL_MS / 1000, // seconds
path: '/',
});
done();
return;
}
@@ -205,7 +267,12 @@ export function registerSecurityHeaders(app: FastifyInstance, https: boolean): v
const scriptSrc =
"script-src 'self' 'unsafe-inline' https://cdn.jsdelivr.net" + (gesture ? " 'wasm-unsafe-eval'" : '');
const connectSrc = "connect-src 'self' wss://api.deepgram.com";
const workerSrc = gesture ? "; worker-src 'self' blob:" : '';
// blob: workers are needed unconditionally: terminal-ui's _safeYield tick
// worker (throttling escape) is created from a Blob URL. Without this, every
// page load logs a CSP violation and the worker leg of _safeYield is dead.
// Risk is minimal — only same-origin scripts (already governed by script-src)
// can construct blob workers.
const workerSrc = "; worker-src 'self' blob:";
const csp =
`default-src 'self'; ${scriptSrc}; style-src 'self' 'unsafe-inline' https://cdn.jsdelivr.net; ` +
`img-src 'self' data: blob:; ${connectSrc}; font-src 'self' https://cdn.jsdelivr.net; frame-ancestors 'self'${workerSrc}`;
+22
View File
@@ -6,6 +6,16 @@ export function isExplicitlyEnabled(value: string | undefined): boolean {
return value !== undefined && EXPLICIT_TRUE_VALUES.has(value.trim().toLowerCase());
}
/**
* True when unauthenticated network exposure is acceptable: either a password is
* set (auth active) or the operator explicitly acknowledged it. Used by the
* tunnel-enable guard (COD-55) to refuse publishing an unauthenticated public URL.
*/
export function isUnauthenticatedNetworkAcknowledged(allowFlag = false): boolean {
if (process.env.CODEMAN_PASSWORD) return true;
return allowFlag || isExplicitlyEnabled(process.env.CODEMAN_ALLOW_UNAUTHENTICATED_NETWORK);
}
export function isLoopbackBindHost(host: string): boolean {
const normalized = host
.trim()
@@ -29,6 +39,16 @@ export function isLoopbackBindHost(host: string): boolean {
*/
export const DEFAULT_TRUSTED_HOST_SUFFIXES = ['.ts.net', '.trycloudflare.com', '.cfargotunnel.com'];
/**
* Container-to-host gateway aliases (Docker / Podman). A hook `curl` from INSIDE a
* docker case carries `Host: host.docker.internal:<port>` (the derived
* CODEMAN_API_URL), so the always-on host guard must allow it or every in-container
* hook is blocked 403. These names only resolve to the host from within a
* container's network namespace, so they are not a DNS-rebinding surface for a
* normal browser. Both engines' aliases are allowed so a mixed fleet keeps working.
*/
export const DOCKER_HOST_GATEWAY_ALIASES = ['host.docker.internal', 'host.containers.internal'];
/** Policy inputs for the anti-DNS-rebinding Host allowlist + cross-site Origin guard. */
export interface HostPolicy {
/** The host the server is bound to (e.g. '127.0.0.1', '0.0.0.0', or a hostname). */
@@ -88,6 +108,8 @@ function matchesHost(hostname: string, policy: HostPolicy): boolean {
const bind = parseAuthorityHostname(policy.bindHost);
if (bind && hostname === bind) return true;
if (policy.tunnelHost && hostname === policy.tunnelHost) return true;
// Docker/Podman container-to-host gateway aliases (for in-container hook curls).
if (DOCKER_HOST_GATEWAY_ALIASES.includes(hostname)) return true;
for (const suffix of DEFAULT_TRUSTED_HOST_SUFFIXES) {
if (hostname === suffix.slice(1) || hostname.endsWith(suffix)) return true;
}
+22
View File
@@ -0,0 +1,22 @@
/**
* @fileoverview Process-wide last-known plan-usage telemetry (account-global).
*
* The status-telemetry route writes the latest broadcast value here; the SSE
* init snapshot (`getLightState`) replays it so the header "Plan Usage Limits"
* chip shows immediately on a fresh page load / SSE reconnect — before any new
* statusline render arrives, and without relying on per-browser localStorage.
*
* Null until the first telemetry of the process; cleared naturally on restart.
*
* @module plan-usage-latest
*/
let latest: Record<string, unknown> | null = null;
export function setLatestPlanUsage(value: Record<string, unknown>): void {
latest = value;
}
export function getLatestPlanUsage(): Record<string, unknown> | null {
return latest;
}
+2
View File
@@ -5,6 +5,7 @@
import type { ClaudeMode, NiceConfig } from '../../types.js';
import type { StateStore } from '../../state-store.js';
import type { TerminalHistoryConfig } from '../../config/terminal-history.js';
export interface ConfigPort {
readonly store: StateStore;
@@ -15,6 +16,7 @@ export interface ConfigPort {
getGlobalNiceConfig(): Promise<NiceConfig | undefined>;
getModelConfig(): Promise<{ defaultModel?: string; agentTypeOverrides?: Record<string, string> } | null>;
getClaudeModeConfig(): Promise<{ claudeMode?: ClaudeMode; allowedTools?: string }>;
getTerminalHistoryConfig(): Promise<TerminalHistoryConfig>;
getDefaultClaudeMdPath(): Promise<string | undefined>;
getLightState(): unknown;
getLightSessionsState(): unknown[];
+10
View File
@@ -0,0 +1,10 @@
/**
* @fileoverview Cron port — exposes the CronService to
* route handlers via the shared route context.
*/
import type { CronService } from '../../cron/cron-service.js';
export interface CronPort {
readonly cron: CronService;
}
+1
View File
@@ -13,3 +13,4 @@ export type { ConfigPort } from './config-port.js';
export type { InfraPort, ScheduledRun } from './infra-port.js';
export type { AuthPort } from './auth-port.js';
export type { OrchestratorPort } from './orchestrator-port.js';
export type { CronPort } from './cron-port.js';
+1388 -195
View File
File diff suppressed because it is too large Load Diff
+98
View File
@@ -51,6 +51,7 @@ const GROUPING_TIMEOUT_MS = 5000; // 5 seconds - notification grouping
const NOTIFICATION_LIST_CAP = 100; // Max notifications in list
const TITLE_FLASH_INTERVAL_MS = 1500; // Title flash rate
const BROWSER_NOTIF_RATE_LIMIT_MS = 3000; // Rate limit for browser notifications
const MOBILE_RESIZE_RETRY_MS = 30000; // Small-viewport resize re-send while a desktop sizing claim is hot
const AUTO_CLOSE_NOTIFICATION_MS = 8000; // Auto-close browser notifications
const THROTTLE_DELAY_MS = 100; // General UI throttle delay
const TERMINAL_CHUNK_SIZE = 32 * 1024; // 32KB chunks for terminal buffer loading
@@ -110,12 +111,85 @@ function evaluateWebGLLongTaskTrip(recent, entries, now, config = WEBGL_FALLBACK
return recent.length >= config.LONGTASK_COUNT;
}
/**
* Pure decision for whether to skip the WebGL renderer at terminal init, and
* whether to clear the auto-fallback sticky marker. Keeps the interaction
* between device type, URL params, the sticky marker, and the user's settings
* toggle in one testable place (terminal-ui.js calls this).
*
* Precedence (desktop only — mobile always skips):
* 1. user toggle OFF -> skip (one-shot opt-out, sticky untouched)
* 2. ?nowebgl -> skip (one-shot opt-out, sticky untouched)
* 3. ?webgl=force -> enable + clear stale sticky marker
* 4. toggle ON / untouched -> respect the auto-fallback sticky marker
*
* A stored `true` is treated like the untouched default here: the checkbox
* ships checked on desktop, so any unrelated settings save stores `true` —
* letting it clear the marker would permanently defeat the GPU-stall
* auto-fallback safety net. The marker is only retired by ?webgl=force or by
* a real OFF->ON toggle flip, which saveAppSettings() detects at save time.
*
* @param {{deviceType?: string, noWebglParam?: boolean, forceParam?: boolean,
* stickyDisabled?: boolean, userPrefEnabled?: (boolean|undefined)}} [input]
* @returns {{skip: boolean, clearSticky: boolean}}
*/
function shouldSkipWebGL(input = {}) {
if (input.deviceType !== 'desktop') return { skip: true, clearSticky: false };
if (input.userPrefEnabled === false) return { skip: true, clearSticky: false };
if (input.noWebglParam) return { skip: true, clearSticky: false };
if (input.forceParam) return { skip: false, clearSticky: true };
return { skip: !!input.stickyDisabled, clearSticky: false };
}
// Expose for tests. `const` declarations at the top of a non-module script
// are global lexical bindings but not `window` properties, so explicit
// assignment is the test-visible API surface.
// Desktop tab-overflow policy: auto-wrap the session tabs to a second row when
// they overflow one row (and the user hasn't pinned the manual two-row layout).
function shouldAutoWrapTabs(input) {
if (!input || input.deviceType !== 'desktop') return false;
if (input.manualTwoRows) return false;
if ((input.tabCount || 0) < 2) return false;
const scrollWidth = Number(input.scrollWidth) || 0;
const clientWidth = Number(input.clientWidth) || 0;
return scrollWidth > clientWidth + 1;
}
// COD-134 — Terminal WebSocket reconnect policy.
//
// Decide what to do after a terminal WebSocket closes, given the close `code`
// and `attempt` (0-based count of consecutive reconnects already made):
// - transient closes (code < 4004: 1000/1001/1005/1006/etc.) → 'reconnect'
// with exponential backoff (0 on the first attempt; the caller adds jitter),
// 250ms → 500 → 1000 → ... capped at 10s.
// - 4004 (session not found) / 4009 (session terminated) → 'give-up': the
// session is gone, retrying only wastes connections.
// - 4008 (too many connections) and any other code >= 4004 → 'retry-fallback':
// show the HTTP fallback but keep retrying on a bounded 5s timer so the
// transport returns to WS once the transient condition clears (un-stick).
// Pure: no DOM, no side effects.
function planWsReconnect(code, attempt) {
if (code === 4004 || code === 4009) {
return { action: 'give-up', delayMs: 0 };
}
if (code >= 4004) {
return { action: 'retry-fallback', delayMs: 5000 };
}
const delayMs = attempt <= 0 ? 0 : Math.min(250 * Math.pow(2, attempt - 1), 10000);
return { action: 'reconnect', delayMs };
}
if (typeof window !== 'undefined') {
window.WEBGL_FALLBACK = WEBGL_FALLBACK;
window.evaluateWebGLLongTaskTrip = evaluateWebGLLongTaskTrip;
window.shouldSkipWebGL = shouldSkipWebGL;
window.CodemanTabOverflow = {
shouldAutoWrapTabs,
};
window.CodemanWsReconnect = {
plan: planWsReconnect,
};
}
// Scheduler API — prioritize terminal writes over background UI updates.
@@ -243,10 +317,15 @@ const SSE_EVENTS = {
SESSION_WORKING: 'session:working',
SESSION_AUTO_CLEAR: 'session:autoClear',
SESSION_AUTO_COMPACT: 'session:autoCompact',
SESSION_LIMIT_PAUSE_SCHEDULED: 'session:limitPauseScheduled',
SESSION_LIMIT_RESUME: 'session:limitResume',
SESSION_LIMIT_RESUME_CANCELLED: 'session:limitResumeCancelled',
SESSION_RESPAWN_BREAKER_TRIPPED: 'session:respawnBreakerTripped',
SESSION_CLI_INFO: 'session:cliInfo',
SESSION_MESSAGE: 'session:message',
SESSION_INTERACTIVE: 'session:interactive',
SESSION_RUNNING: 'session:running',
SESSION_STATUS_TELEMETRY: 'session:statusTelemetry',
// Scheduled runs
SCHEDULED_CREATED: 'scheduled:created',
@@ -256,6 +335,12 @@ const SSE_EVENTS = {
SCHEDULED_LOG: 'scheduled:log',
SCHEDULED_DELETED: 'scheduled:deleted',
// Cron jobs
CRON_JOBS_CHANGED: 'cron:jobsChanged',
CRON_JOB_DELETED: 'cron:jobDeleted',
CRON_RUN_CREATED: 'cron:runCreated',
CRON_RUN_UPDATED: 'cron:runUpdated',
// Respawn
RESPAWN_STARTED: 'respawn:started',
RESPAWN_STOPPED: 'respawn:stopped',
@@ -330,8 +415,14 @@ const SSE_EVENTS = {
SUBAGENT_TOOL_RESULT: 'subagent:tool_result',
SUBAGENT_COMPLETED: 'subagent:completed',
// Workflow runs (ultracode / Workflow tool)
WORKFLOW_RUN_DISCOVERED: 'workflow:run_discovered',
WORKFLOW_RUN_UPDATED: 'workflow:run_updated',
WORKFLOW_RUN_REMOVED: 'workflow:run_removed',
// Images
IMAGE_DETECTED: 'image:detected',
ATTACHMENT_DETECTED: 'attachment:detected',
// Tunnel
TUNNEL_STARTED: 'tunnel:started',
@@ -383,6 +474,13 @@ const SSE_EVENTS = {
CASE_LINKED: 'case:linked',
CASE_DELETED: 'case:deleted',
CASE_ORDER_CHANGED: 'case:order-changed',
DOCKER_EXPORT_COMPLETE: 'docker:exportComplete',
DOCKER_EXPORT_FAILED: 'docker:exportFailed',
DOCKER_IMPORT_COMPLETE: 'docker:importComplete',
DOCKER_IMAGE_BUILD_STARTED: 'docker:imageBuildStarted',
DOCKER_IMAGE_BUILD_PROGRESS: 'docker:imageBuildProgress',
DOCKER_IMAGE_BUILD_COMPLETE: 'docker:imageBuildComplete',
DOCKER_IMAGE_BUILD_FAILED: 'docker:imageBuildFailed',
};
// ═══════════════════════════════════════════════════════════════
+328
View File
@@ -0,0 +1,328 @@
/**
* @fileoverview Cron Jobs UI mixed into
* CodemanApp.prototype. Renders the job list + create/edit form in the
* #cronModal, and reacts to cron:* SSE events.
*
* @mixin Extends CodemanApp.prototype via Object.assign
* @dependency app.js, api-client.js, constants.js (escapeHtml)
*/
Object.assign(CodemanApp.prototype, {
// ── SSE handlers ──────────────────────────────────────────────────────────
_onCronJobsChanged(data) {
if (data && Array.isArray(data.jobs)) {
this._cronJobs = data.jobs;
if (this._isCronOpen()) this.renderCronJobs();
} else if (this._isCronOpen()) {
this.refreshCron();
}
},
_onCronRunChanged() {
// A run's status changed — refresh the list so lastStatus stays current.
if (this._isCronOpen()) this.refreshCron();
},
// ── Modal open/close ──────────────────────────────────────────────────────
_isCronOpen() {
const el = document.getElementById('cronModal');
return !!el && el.classList.contains('active');
},
openCron() {
const el = document.getElementById('cronModal');
if (!el) return;
el.classList.add('active');
this.cancelCronJobForm();
this.refreshCron();
},
closeCron() {
const el = document.getElementById('cronModal');
if (el) el.classList.remove('active');
},
async refreshCron() {
const jobs = await this._apiJson('/api/cron/jobs');
this._cronJobs = Array.isArray(jobs) ? jobs : [];
this.renderCronJobs();
},
// ── List rendering ────────────────────────────────────────────────────────
renderCronJobs() {
const list = document.getElementById('cronJobList');
if (!list) return;
const jobs = this._cronJobs || [];
if (jobs.length === 0) {
list.innerHTML = '<div class="form-hint">No cron jobs yet. Click “+ New Job”.</div>';
return;
}
const rows = jobs.map((j) => {
const next = j.enabled ? this._fmtTime(j.nextRunAt) : '—';
const last = this._fmtTime(j.lastRunAt);
const status = j.lastStatus ? escapeHtml(j.lastStatus) : '—';
return `
<div class="cron-job-row">
<div class="cron-job-main">
<div class="cron-job-name">${escapeHtml(j.name || '(unnamed)')}
<span class="cron-badge">${escapeHtml(j.agentType)}</span>
<span class="cron-badge">${escapeHtml(this._fmtSchedule(j))}</span>
${j.enabled ? '' : '<span class="cron-badge cron-badge-off">disabled</span>'}
</div>
<div class="cron-job-meta">
<span title="${escapeHtml(j.workingDir || '')}">${escapeHtml(j.workingDir || '')}</span>
· next: ${escapeHtml(next)} · last: ${escapeHtml(last)} · status: ${status}
</div>
</div>
<div class="cron-job-actions">
<button class="btn-toolbar btn-sm btn-primary" onclick="app.runCronJob('${j.id}')">Run Now</button>
<button class="btn-toolbar btn-sm" onclick="app.toggleCronJob('${j.id}', ${j.enabled ? 'false' : 'true'})">${j.enabled ? 'Disable' : 'Enable'}</button>
<button class="btn-toolbar btn-sm" onclick="app.editCronJob('${j.id}')">Edit</button>
<button class="btn-toolbar btn-sm btn-danger" onclick="app.deleteCronJob('${j.id}')">Delete</button>
</div>
</div>`;
});
list.innerHTML = rows.join('');
},
_fmtTime(ts) {
if (!ts) return '—';
try {
return new Date(ts).toLocaleString();
} catch {
return '—';
}
},
_fmtSchedule(j) {
switch (j.scheduleType) {
case 'once':
return 'once';
case 'interval':
return `every ${j.intervalMinutes}m`;
case 'daily':
return `daily ${j.dailyTime || ''}`;
case 'weekly': {
const names = ['Sun', 'Mon', 'Tue', 'Wed', 'Thu', 'Fri', 'Sat'];
const days = (j.weeklyDays || []).map((d) => names[d] || d).join(',');
return `weekly ${days} ${j.weeklyTime || ''}`;
}
default:
return j.scheduleType || '';
}
},
// ── Create / edit form ────────────────────────────────────────────────────
openCronJobForm(job) {
const form = document.getElementById('cronJobForm');
if (!form) return;
document.getElementById('cronFormError').textContent = '';
document.getElementById('cronFormTitle').textContent = job ? 'Edit Cron Job' : 'New Cron Job';
document.getElementById('schJobId').value = job ? job.id : '';
document.getElementById('schName').value = job ? job.name || '' : '';
document.getElementById('schAgentType').value = job ? job.agentType || 'claude' : 'claude';
document.getElementById('schWorkingDir').value = job ? job.workingDir || '' : '';
document.getElementById('schLaunchCommand').value = job ? job.launchCommand || '' : '';
document.getElementById('schPromptMode').value = job ? job.promptMode || 'inline_text' : 'inline_text';
document.getElementById('schPromptText').value = job ? job.promptText || '' : '';
document.getElementById('schPromptFilePath').value = job ? job.promptFilePath || '' : '';
document.getElementById('schInputMode').value = job ? job.inputMode || 'typed' : 'typed';
document.getElementById('schScheduleType').value = job ? job.scheduleType || 'once' : 'once';
document.getElementById('schRunAt').value = job && job.runAt ? this._toLocalInput(job.runAt) : '';
document.getElementById('schIntervalMinutes').value = job && job.intervalMinutes ? job.intervalMinutes : 60;
document.getElementById('schDailyTime').value = job ? job.dailyTime || '' : '';
document.getElementById('schWeeklyTime').value = job ? job.weeklyTime || '' : '';
const weekly = (job && job.weeklyDays) || [];
document.querySelectorAll('#schWeeklyDays input[type=checkbox]').forEach((cb) => {
cb.checked = weekly.includes(Number(cb.value));
});
document.getElementById('schConcurrencyPolicy').value = job ? job.concurrencyPolicy || 'warn_only' : 'warn_only';
document.getElementById('schAutoClosePrev').checked = job ? job.autoClosePreviousSession !== false : true;
document.getElementById('schEnabled').checked = job ? !!job.enabled : true;
document.getElementById('schNotes').value = job ? job.notes || '' : '';
this.onCronAgentTypeChange();
this.onCronPromptModeChange();
this.onCronScheduleTypeChange();
form.classList.remove('hidden');
},
editCronJob(id) {
const job = (this._cronJobs || []).find((j) => j.id === id);
if (job) this.openCronJobForm(job);
},
cancelCronJobForm() {
const form = document.getElementById('cronJobForm');
if (form) form.classList.add('hidden');
},
onCronAgentTypeChange() {
// Launch command is only meaningful for shell mode (first input line).
const isShell = document.getElementById('schAgentType').value === 'shell';
document.getElementById('schLaunchCommandRow').classList.toggle('hidden', !isShell);
},
onCronPromptModeChange() {
const mode = document.getElementById('schPromptMode').value;
document.getElementById('schPromptTextRow').classList.toggle('hidden', mode !== 'inline_text');
document.getElementById('schPromptFileRow').classList.toggle('hidden', mode !== 'prompt_file_path');
},
onCronScheduleTypeChange() {
const t = document.getElementById('schScheduleType').value;
document.getElementById('schRunAtRow').classList.toggle('hidden', t !== 'once');
document.getElementById('schIntervalRow').classList.toggle('hidden', t !== 'interval');
document.getElementById('schDailyRow').classList.toggle('hidden', t !== 'daily');
document.getElementById('schWeeklyDaysRow').classList.toggle('hidden', t !== 'weekly');
document.getElementById('schWeeklyTimeRow').classList.toggle('hidden', t !== 'weekly');
},
_toLocalInput(ts) {
// epoch-ms → 'YYYY-MM-DDTHH:MM' in local time for <input datetime-local>.
const d = new Date(ts);
const pad = (n) => String(n).padStart(2, '0');
return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())}T${pad(d.getHours())}:${pad(d.getMinutes())}`;
},
_collectCronForm() {
const t = document.getElementById('schScheduleType').value;
const promptMode = document.getElementById('schPromptMode').value;
const body = {
name: document.getElementById('schName').value.trim(),
agentType: document.getElementById('schAgentType').value,
workingDir: document.getElementById('schWorkingDir').value.trim(),
promptMode,
inputMode: document.getElementById('schInputMode').value,
scheduleType: t,
concurrencyPolicy: document.getElementById('schConcurrencyPolicy').value,
autoClosePreviousSession: document.getElementById('schAutoClosePrev').checked,
enabled: document.getElementById('schEnabled').checked,
notes: document.getElementById('schNotes').value.trim() || undefined,
};
// Always sent for shell (an emptied field must clear a saved command on edit).
if (body.agentType === 'shell') body.launchCommand = document.getElementById('schLaunchCommand').value.trim();
if (promptMode === 'inline_text') {
// Prompt delivery is single-line only; trailing newlines are harmless, strip them.
body.promptText = document.getElementById('schPromptText').value.replace(/[\r\n]+$/, '');
} else {
body.promptFilePath = document.getElementById('schPromptFilePath').value.trim();
}
if (t === 'once') {
const v = document.getElementById('schRunAt').value;
body.runAt = v ? new Date(v).getTime() : undefined;
} else if (t === 'interval') {
body.intervalMinutes = Number(document.getElementById('schIntervalMinutes').value);
} else if (t === 'daily') {
body.dailyTime = document.getElementById('schDailyTime').value;
} else if (t === 'weekly') {
body.weeklyTime = document.getElementById('schWeeklyTime').value;
body.weeklyDays = Array.from(document.querySelectorAll('#schWeeklyDays input:checked')).map((cb) =>
Number(cb.value)
);
}
return body;
},
async saveCronJob() {
const errEl = document.getElementById('cronFormError');
errEl.textContent = '';
const body = this._collectCronForm();
if (!body.name) {
errEl.textContent = 'Name is required.';
return;
}
if (!body.workingDir) {
errEl.textContent = 'Working directory is required.';
return;
}
if (body.promptText !== undefined && /[\r\n]/.test(body.promptText)) {
errEl.textContent = 'Prompt must be a single line — multi-line prompts are not supported.';
return;
}
const id = document.getElementById('schJobId').value;
const res = id ? await this._apiPut(`/api/cron/jobs/${id}`, body) : await this._apiPost('/api/cron/jobs', body);
if (!res || !res.ok) {
let msg = 'Failed to save job.';
try {
const j = await res.json();
if (j && j.error) msg = j.error;
} catch {
/* ignore */
}
errEl.textContent = msg;
return;
}
this.showToast?.(id ? 'Cron job updated' : 'Cron job created', 'success');
this.cancelCronJobForm();
this.refreshCron();
},
// ── Actions ───────────────────────────────────────────────────────────────
async runCronJob(id) {
const job = (this._cronJobs || []).find((j) => j.id === id);
if (job) {
const active = this._countActiveAgents(job.agentType);
if (active > 0 && !confirm(`${active} ${job.agentType} session(s) already active. Run this job anyway?`)) {
return;
}
}
const res = await this._apiPost(`/api/cron/jobs/${id}/run`, {});
if (res && res.ok) {
this.showToast?.('Run started — opening session', 'success');
let data = null;
try {
data = await res.json();
} catch {
/* ignore */
}
const run = data && (data.data ? data.data.run : data.run);
if (run && run.sessionId) this._focusCronSession(run.sessionId);
this.refreshCron();
} else {
this.showToast?.('Failed to run job', 'error');
}
},
_focusCronSession(sessionId) {
// Best-effort: switch to the created session tab if it exists.
if (this.sessions && this.sessions.has(sessionId) && typeof this.switchSession === 'function') {
this.closeCron();
this.switchSession(sessionId);
}
},
_countActiveAgents(agentType) {
// Mirrors the server's countActiveAgents: only LIVE sessions count — a
// tab whose CLI already exited (stopped/error) doesn't block anything.
if (!this.sessions) return 0;
let n = 0;
for (const s of this.sessions.values()) {
if (s && s.mode === agentType && s.status !== 'stopped' && s.status !== 'error') n++;
}
return n;
},
async toggleCronJob(id, enabled) {
const res = await this._apiPut(`/api/cron/jobs/${id}/enabled`, { enabled });
if (res && res.ok) this.refreshCron();
else this.showToast?.('Failed to update job', 'error');
},
async deleteCronJob(id) {
if (!confirm('Delete this cron job and its run history?')) return;
const res = await this._apiDelete(`/api/cron/jobs/${id}`);
if (res && res.ok) {
this.showToast?.('Cron job deleted', 'success');
this.refreshCron();
} else {
this.showToast?.('Failed to delete job', 'error');
}
},
});
Binary file not shown.

Some files were not shown because too many files have changed in this diff Show More