Commit Graph
60 Commits
Author SHA1 Message Date
timkjr 0a5bc1ac2e fix(omp,remote): pin remote conversations on respawn so ctrl-d/ctrl-c resumes instead of relaunching fresh
Two independent defects made ANY clean exit from a remote SSH session (user
ctrl-d or ctrl-c, or a dropped pane) relaunch the agent as a NEW conversation:

1. SSH-remote claude was launched as a bare `claude --dangerously-skip-permissions`,
   so the remote-respawn path (COD-108 reattachRemote re-running the idempotent
   launch command) started a fresh conversation every time. Pin it to the
   deterministic Codeman session id, mirroring the docker-claude shape
   (claudeDockerPaneCommand): `--session-id <id>` to create, with the
   `|| --resume <id>` fallback so the idempotent re-run resumes instead of
   erroring with "already in use". A per-host commands.claude override still wins.

2. OMP --resume pinning silently degraded to ambiguous `--continue` whenever a
   case path ended in a trailing slash (e.g. remote `remotePath` stored verbatim
   as `/home/user/dotfiles/`): mangleOmpWorkingDir produced `-dotfiles-` while
   omp persists sessions under `-dotfiles`, readdirSync returned null for an
   existing dir, and findLatestOmpSessionId/resolveAndClaimOmpSessionId never
   matched. Normalize the trailing slash before mangling (new exported
   stripTrailingSlash) and compare the session header cwd against the same
   normalized value.

Both were found live 2026-08-29 on a remote OMP/Claude node: ctrl-c and ctrl-d
behaved identically, both relaunching a fresh session.
2026-09-07 21:20:41 -05:00
DevvynandClaude Opus 5 4830e662f9 refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.

Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.

Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.

OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.

Guard rails:

- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
  branching reappears outside `stock.ts`, in any of its four shapes (`===`,
  `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
  negated forms, which is how 36 of them survived an earlier pass. Every
  allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
  deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
  `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
  wire field is separate, bridged only by `legacyConfigAliases`. Getting
  `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
  clamp's only handle on a CLI's privilege switch, and a wrong name clamps
  nothing with no error and no failing test — so `schema.ts` rejects an entry
  naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
  `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
  thunks). A module-level const freezes at first import, so a CLI enabled while
  the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
  (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
  `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
  rather than measured. A test pins the list so it cannot quietly grow.

Three user-visible changes, all deliberate and named:

- `probeDockerCliVersion()` derives the in-container binary from the registry
  rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
  hardcoded map it replaces omitted while its own comment said the rule was
  "every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
  install hint is the install command rather than a docs URL, five CLIs gain
  hints they never had, and the row order follows the catalog.

Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:26:45 +08:00
Codeman maintainer da91b4353b Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend
2026-08-30 16:16:26 +02:00
timkjr 65e994d29a fix(omp): correct docs/counts/URLs, resolver install-path order, stray comment + CSS
Small cleanup items from upstream review (Ark0N/Codeman#353):

- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
  installer target (~/.omp/bin was an earlier unverified guess, confirmed
  wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
  can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
  SessionMode incl. shell -- not eighth), matched the install-path guidance
  to the resolver fix, updated the version example to the actually-tested
  18.0.8, and added a Docker-section caveat: --resume pinning does not
  currently reach an in-container omp process, since Docker panes never see
  ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
  already linked to the -omp anchor, so the link was dead) and added an OMP
  specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
  Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
  and split two CSS lines that had two declarations jammed onto one line.
2026-08-28 14:37:28 -05:00
timkjr 2ee2eacb4b fix(omp): clamp OMP_AUTH_BROKER_URL/TOKEN, correct the env-allowlist docs
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.

The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
2026-08-28 13:45:12 -05:00
timkjr c4f6eb1e5e fix(omp): resolve and pin the respawn session id only at actual respawn time
findLatestOmpSessionId()'s newest-mtime pin ran eagerly inside
_buildRespawnPaneOptions(), which startInteractive() calls unconditionally
on every boot-recovery reattach — before anything checks whether the pane
is actually dead. With two omp tabs in the same case dir, this could pin
an ALIVE pane's session onto whichever sibling's file happened to be
newest on disk, purely as a side effect of building options that might
never lead to a respawn (reported in Ark0N/Codeman#353 review).

Move resolution out of the eager builder into _pinOmpRespawnId(), called
explicitly only where a respawn is actually confirmed: the dead-pane
branch in _setupOrAttachMuxSession() and reattachRemote(). Add
resolveAndClaimOmpSessionId(), which verifies each candidate's own file
header (cwd) rather than trusting the mangled-directory match alone, and
tracks claimed ids in a process-wide registry so two ambiguous resolutions
can't both pick the same sibling's conversation.
2026-08-28 13:18:25 -05:00
timkjrandClaude Sonnet 5 853681f970 harden(omp): resume-path test coverage, silent-fallback logging, cwd validation
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:

- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
  exported to make it testable: the exact "resume this OMP row from
  history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
  zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
  continuation silently degrades to omp's own ambiguous --continue,
  in both call sites (session create and respawn pinning) - previously
  silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
  directory in omp-transcript.ts's parser, so a corrupted/malformed
  session file can't point a downstream resume at a relative or empty
  path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
  case in mangleOmpWorkingDir(): the review's suggested realpath() fix
  assumes omp itself resolves symlinks before mangling, which is
  unconfirmed - guessing wrong there would trade one silent mismatch
  for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
  session-routes.ts (antigravity/opencode dynamic import line-wrapping)
  that was blocking the pre-commit formatting gate on this file.

Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjrandClaude Sonnet 5 ed983f898b fix(omp): resolve claudeSessionId alias on boot-recovery reattach
Two bugs compounded to break continuation pinning on every real OMP
case (only /tmp-based manual testing happened to work by coincidence):

1. startInteractive() had a second, unconditional claudeSessionId
   assignment after the mux branch that clobbered its correctly
   resolved value back to the session's own id on every mux path.

2. mangleOmpWorkingDir() assumed omp mirrors Claude Code's directory
   naming (home prefix kept), but omp actually strips $HOME first.
   findLatestOmpSessionId() was silently returning null for every
   case under ~/codeman-cases/, so resumeSessionId never resolved for
   any real case dir - only /tmp paths (outside $HOME) worked, which
   is every dir this feature was previously tested against.

Verified live: killed and relaunched the omp-verify server process
mid-session (plain reattach, pane stayed alive) and confirmed
claudeSessionId now resolves to the real omp transcript uuid instead
of the Codeman session's own id.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjr 54a930c80e feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.

Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.

Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.

Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.

Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.

Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
2026-08-28 11:32:30 -05:00
timkjr 4c332c6141 fix(omp): retire the old row on resume, and let DELETE remove persisted-only sessions
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.

That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).

Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
2026-08-28 11:32:30 -05:00
timkjr 4f5678fac4 feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) 2026-08-26 20:05:48 -05:00
Jack Stuart b00ab3ceea feat(web): show Codex plan usage in header 2026-08-26 18:43:52 +08:00
Codeman maintainer a628737d1f fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first:

- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
  _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
  every dsh pane and applyEnvOverrides() lands after it, so a non-granted
  owner who could redirect the base URL would have the operator's key sent
  as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
  dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
  The HERDR triple is set via LOCAL tmux setenv, which crosses neither
  docker exec nor ssh, so such a session can never post a hook event and
  the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
  route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
  option parser cannot read a third-party TUI's frames, so an answer was a
  blind keystroke into a foreign composer), and the push notification
  carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
  reports inside a 60s window (the TUI retries with backoff, so a retried
  'working' could land after 'blocked' and resolve an approval whose
  dialog was still on screen); 4xx responses exit 0 instead of retrying,
  so one misconfigured session cannot feed the auth rate-limit bucket
  until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
  racing POSTs used to pick the same port and orphan the winner), and the
  readiness poll / timeout paths only clear or stop the singleton while it
  is still theirs. First click actually opens the tab now
  (refreshWebviews, not the nonexistent loadWebviews). DELETE
  /api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
  create paths (impl moved into the resolver so all three share it) and no
  longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
  the remote branch; the Ralph auto-enable list gained deepseek;
  HookEventType gained agent_working; the phone overview run menu filters
  managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
  child that reads stdin eats the rest of the script), bounds the exec
  with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
  identity (it rendered as an unstyled UA-grey button); stale markup
  comment about the web shortcut rewritten; clamp docs updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:01:29 +02:00
Codeman maintainer cdceede33d fix(deepseek): atomic shim write, honest attribution comment, name-fallback profile classifier
The three smaller review nits, plus the first real test coverage for the status
shim (it had none: it is emitted as a STRING, so tsc never sees it).

1. The shim was written with a plain writeFileSync. The TUI can be exec'ing that
   exact path while an upgraded Codeman refreshes it, and a reader catching a
   half-written file gets a syntax error, exits non-zero, and is retried four
   times per state change for a file that will never parse. Now temp + rename
   (atomic within the directory), with the temp chmod'ed before the rename since
   writeFileSync's mode only applies on create, and removed if the write throws.
   SHIM_VERSION bumped to 2, because SHIM_SOURCE changed and an existing v1 shim
   would otherwise keep matching the embedded marker and never be refreshed.

2. The pane-id comment claimed the ambient env "cannot be spoofed by an argument
   the agent itself could influence". The agent runs IN that pane and can invoke
   the shim with CODEMAN_SESSION_ID unset and any argv it likes. It buys nothing
   it did not already have (the hook-secret file is readable from the same pane,
   so it can POST /api/hook-event directly), but the comment read like a security
   boundary. Rewritten to say what the preference actually buys: correct
   attribution when a TUI mangles or re-uses the pane argument. Accidents, not
   adversaries.

3. classifyProfile() folded the directory name into the same haystack as the
   bundles, but only the TUI arm could match a bare name, so a stock profile
   whose package.json has no dsh.profile.bundles (hand-edited, older layout,
   mid-install) classified as `unknown` -> launchable -> eligible as the DEFAULT
   pick, which is exactly the pane-dies-on-arrival failure the two-part
   availability gate exists to prevent. The stock names are now a LAST-resort
   fallback consulted after the bundle patterns, so real bundle evidence still
   wins over a name the user chose. The loose `tui` arm gained word boundaries:
   it decides which profile boots by default, and matching the middle of
   `intuition` is not a rule anyone could predict.

New test/deepseek-status-shim.test.ts runs the generated script the way the
harness does -- real node process, real argv, real env, real listener -- and
covers the exit-code contract that makes the retry behaviour safe: mapped states
post and exit 0, an unknown verb or unmapped state exits 0 WITHOUT posting (a
non-zero there would be four HTTP requests per state change forever), a rejecting
server or an unreachable one exits non-zero so the caller retries, the hook secret
is read at execution time, and `node --check` parses the file (a template-literal
typo in SHIM_SOURCE is invisible to tsc).

Trap worth recording, hit while writing it: the tests must spawn the shim
ASYNCHRONOUSLY. The listener lives in the test process, so spawnSync blocks the
event loop that has to accept the connection, the shim waits out its own 1500ms
socket timeout and exits 1, and it reads exactly like a broken shim (measured:
Socket._onTimeout in its --trace-exit output, server logging nothing).

Verified: full gate green (6142 passed, +10), typecheck/lint/format clean.
2026-08-24 18:00:52 +02:00
Codeman maintainer 4cda150493 feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.

DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:

1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
   $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
   `base` -- the interactive terminal front door is always a third-party
   plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
   and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
   button gates on the latter, because reporting only the binary would spawn a
   pane that dies on arrival. When the binary is present but no profile is,
   the run menu offers to install one (POST /api/deepseek/install-profile).

2. The permission switch is an env var, not a flag. The harness has no
   command-line permission option; its sandbox/approval rows read
   DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
   Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
   own workspace-write, which still asks, so the multi-user clamp is the
   only-if-sent branch and clamps to workspace-write, never read-only.

3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
   earned that. The terminal front door reports idle/working/blocked to a
   supervising process over a generic env-gated contract; a generated shim
   (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
   report to /api/hook-event as stop / agent_working / permission_prompt. So a
   dsh session gets definitive respawn triggers, real wait-endpoint signals and
   real Approvals Inbox items instead of output-stabilization guesswork.
   `agent_working` is new (157th SSE constant) and joins
   APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
   alert at once.

The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.

Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.

Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.

Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 03:37:56 +02:00
Codeman maintainer 3f8c8e99d1 feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.

Grok mixes two existing shapes and the wiring follows from that:

- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
  (--always-approve, grok's bypassPermissions mode; config-level deny rules
  still apply on top). The Run button sends it true, like runAntigravity(),
  and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
  a bare grok spawn is grok's own ask-mode default, which is already safe,
  so only a sent config needs the flag forced off. Cron needs nothing for
  the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
  with mouse support, so it stays OUT of isAltScreenStripMode() and lands
  on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
  (unmeasured against an authenticated composer; documented fallback is the
  'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
  also installs a grok bin), so grok-cli-resolver.ts version-probes every
  candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
  GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
  shared with the dependency registry so doctor and run mode cannot drift.

Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.

Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.

Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).

Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:39:03 +02:00
Codeman maintainer 4e5d0dcbd6 fix: measure the doctor table columns and let the CLI paint them
"Antigravity CLI" is 15 characters and the hardcoded padEnd(14) pushed
that whole row one column right. Widths now come from the widest cell.

The header always said the CLI layer may colorize, but there was no way
to: renderTable now takes an optional ReportStyle whose hooks are
identity by default, so the module still decides nothing about color and
its output stays byte-stable. Padding is applied outside the paint, so a
row with no path detail ends at its status text instead of trailing
spaces inside a color run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-22 14:13:57 +02:00
Codeman maintainer 61251c0b94 fix(cli-resolvers): negative-result caching, SIGKILL on probes, restored VITEST hermeticity, wired not-found diagnostics
Post-merge follow-ups for PR #329 (shared CLI executable resolution):

- Negative-cache resolution misses with a doubling backoff (1min -> 5min
  cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
  shared resolver cached success only, so a missing CLI re-ran the whole
  chain - ending in a synchronous interactive login-shell spawn bounded by
  the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
  attempt, stalling the event loop each time, forever. Success still caches
  for the process lifetime, so an installed CLI is picked up within minutes
  without a restart. Tests drive the backoff via an injectable clock
  (createCliExecutableResolver `now` option, threaded through the
  createPiResolverForTest / createAntigravityResolverForTest wrappers).

- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
  pi/claude --version probes: execFileSync's timeout only SENDS the kill
  signal and then keeps waiting for the child to exit, and interactive bash
  ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
  survived the timeout and blocked the server permanently.

- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
  test pinned the deletion): under vitest the production resolver host now
  replaces un-injected IO primitives with inert stubs - no real PATH
  scanning, no login-shell spawns - and probePiVersion never executes a
  `pi` candidate again (`pi` is a generic binary name, so route tests
  hitting /api/pi/status executed whatever binary the machine carried).
  Tests opt in through the runCommand/isExecutableFile injection hooks or
  allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
  test is replaced by behavioral pins, including a real-executable fixture
  in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
  is ever removed again.

- Wire the six get*NotFoundMessage() exports (previously dead) into their
  intended call sites: the createSession throws in tmux-manager and the
  availability gates on POST /api/sessions and POST /api/quick-start in
  session-routes, replacing a third hardcoded copy of the text. A not-found
  error now names where resolution looked (server PATH, login shell,
  checked directories). npm run knip no longer reports any unused export
  from the resolver modules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:37:58 +02:00
Aamer Akhter fef903df98 fix(cli-resolvers): find CLIs installed via nvm/Homebrew when running as a service
A CLI installed by nvm, Homebrew or a user-level npm prefix lives on a PATH that
only a login shell sets up. Codeman running under systemd or launchd does not get
that PATH — launchd hands a job `/usr/bin:/bin:/usr/sbin:/sbin` — so every
resolver reported the CLI as unavailable on installs where it is plainly there
and works from a terminal.

Each of the six resolvers had its own hand-rolled copy of the same PATH walk, so
the fix is factored into one shared `createCliExecutableResolver()` with an
explicit lookup order: the server process PATH, then common install directories in
order, then an interactive login shell as the last resort. Only the last step
spawns anything, and only when the cheap lookups have already missed.

Also adds `formatCliNotFoundMessage()`, so a failure explains where it looked
instead of just asserting the CLI is missing. Its diagnostics are bounded and
control characters are flattened, so a not-found message cannot dump arbitrary
environment data.

Success is cached and failure is retried, so installing a CLI while the server is
running is picked up without a restart.

Net -103 lines across the six resolvers. Behaviour is unchanged wherever the CLI
was already on the process PATH: that remains the first thing checked.

Tests: 20 cases in test/cli-executable-resolver.test.ts covering the precedence
order, login-shell-only resolution, the caching rule, unsafe-name rejection, and
the bounded diagnostics.
2026-08-20 12:47:42 -04:00
Codeman maintainer 68ae9a8c5f fix(files): match glob queries without regex so a hostile query cannot stall the server
The Files search compiled the user's query into a backtracking RegExp:
'*a*a*a...' became '^.*a.*a.*a...$', the classic blowup, evaluated
synchronously against every walked path — a pathological query could
freeze the event loop for the whole server (and every user of it in
multi-user mode). /api/search stays regex-free for exactly this reason.

Globs now match through a two-pointer wildcard walk, O(text · pattern)
worst case, with a 256-char query cap bounding the pattern side; an
overlong query compiles to null, the same answer as an empty one.
Semantics are unchanged (anchored, case-insensitive, * spans slashes)
and the existing tests pass untouched; the pathological pattern gets a
test that fails by timeout with the RegExp version.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-19 23:35:45 +02:00
Aamer Akhter 5cc78669bd feat(files): search the Files panel by name or path
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.

compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.

The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.

Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
2026-08-19 09:17:20 -04:00
Codeman maintainer 86c78fece3 fix(pi): align the doctor with the pi resolver, correct the strip rationale, update the skill
Second review pass on #282, the three items left open after f4dcfbe.

1. `codeman doctor` and the run mode disagreed about pi. The registry entry
   accepted a bare `which pi` hit while pi-cli-resolver demanded semver-shaped
   `--version` output, so the Dependencies panel could report an installed Pi CLI
   on a box where Run Pi stays hidden, which reads as a broken mode rather than a
   missing install. Both sides now share one exported PI_VERSION_REGEX, and
   PathResolver gains an opt-in `requireVersionMatch` so a binary that fails the
   shape check is reported MISSING instead of installed-with-unknown-version.
   Only pi sets it; every other tool keeps its current behaviour.

2. The isAltScreenStripMode comment justified excluding pi with "the alt screen
   is load-bearing for its fullscreen TUI". That is not what exclusion does: pi
   is tmux-backed, so it falls through to isMuxAltScreenOnlyStripMode, which
   strips the alt-screen toggles anyway. What exclusion actually preserves is
   `\x1b[3J` and the mouse DECSETs, which is the real reason (pi renders into the
   main screen and is mouse-aware). Comment and changeset now say that, and state
   the consequence: fullscreen pi paints into the main buffer, like vim in a tmux
   shell session.

3. skills/codeman still enumerated the five pre-pi modes in nine places, telling
   agents a backend does not exist and understating class-wide caveats by one
   mode. All updated, plus stale session.ts line references refreshed.

Tests: a new static guard derives the mode set from the Zod schema (not a copy)
and fails when a skill enumeration lists a partial set of external CLIs, verified
by mutation. It also documents the one legitimate exception it found: the "writes
no transcript" lists drop codex, which does write a rollout Codeman reads back.
Plus doctor cases for an unrelated `pi` on PATH and registry/resolver regex parity.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 17:41:44 +02:00
Codeman maintainer c5b59633d8 feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code,
OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose
tab identity, welcome button, run-mode entry, cron agentType, Docker and
remote-SSH command defaults, and clone-repo Brain option.

Pi is a different shape of CLI from the other four, and three decisions
follow from that:

- It has NO permission prompts and no sandbox, so there is no
  --dangerously-skip-permissions analog and none was invented. The
  privilege-shaped knob is the tri-state approveProjectTrust, which makes
  pi load and EXECUTE repo-local .pi/extensions TypeScript and install
  missing project packages. clampExternalCliBypassForOwner() therefore
  puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets
  --no-approve even when no config was sent, because pi's own default is
  a prompt the session user could answer themselves. That helper had zero
  test coverage; it now has coverage for all four CLIs.
- Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars
  share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode
  context, so admitting them would widen the allowlist for every mode at
  once. Auth goes through pi's /login or the server's own environment.
  --api-key is deliberately never wired: it would put a provider secret on
  the spawn command line.
- pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the
  main screen with terminal-owned scrollback, and its 0.84.0 fullscreen
  mode is runtime-switchable via /settings; that flip was measured to put
  the pane into the alt screen, which the strip would have corrupted.

pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires
semver-shaped output, because `pi` is a short generic name a stray binary
can shadow; GET /api/pi/status surfaces path and version so a
misresolution is diagnosable rather than presenting as a broken mode.

Docker installs pi in its own --ignore-scripts step so that flag cannot
affect the other four CLIs, and seeds its credentials per-file rather than
whole-dir (~/.pi/agent also holds sessions, extensions and package trees).

Verified end to end against pi 0.84.1 on an isolated instance: resolver
search-dir fallback, flag construction, piConfig persistence across a full
server restart, the trust prompt and its --no-approve suppression, the
rose Run button on the default daylight-blue skin (the nested skin block
eats per-mode gradients unless the rule lives inside it), and the buffer
local-echo policy, which pi tolerates where codex did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:54:47 +02:00
Codeman maintainer b03780dfd2 fix(session): decide working/idle from the pane, not the composer redraw
Every working Claude session reported `status: "idle"` about two seconds
into its turn. Measured on live workers: two sessions mid-tool-call at 13
and 17 minutes both read `idle` while their panes showed
`✻ Actualizing… (13m 23s · ↓ 47.5k tokens)`.

Two things had drifted apart:

1. The working indicator changed. Claude animates the glyph through
   `· ✢ ✳ ∗ ✻ ✽` and randomizes the gerund per turn, so neither
   SPINNER_PATTERN (braille, no longer drawn) nor the keyword list
   (Thinking/Writing/Reading/Running) matches a turn anymore.
2. A `❯` sighting is not the end of a turn. Claude redraws the composer
   roughly once a second all the way through one, and that redraw armed
   the "2s later, call it idle" timer.

Matching the new status line in the STREAM does not fix it either: tmux
ships partial repaints, so the complete line reached the PTY about once
every 20 seconds while the `❯` arrived every second.

So the decision moves off the stream:

- An unbroken run of repaints marks a turn as started. Sampled once a
  second for 12s over six live sessions, the two working ones produced
  output in 12/12 windows and the four idle ones in 0/12. Pure helpers in
  session-activity.ts carry the thresholds.
- Idle now needs the pane to go quiet AND the screen to agree.
  `_confirmIdle()` asks tmux what is rendered (new `capturePaneText()`,
  one plain `capture-pane`, floored at 1.5s per session and only ever at
  a transition) and re-checks every 5s while the screen still shows work.
  A turn can sit silent for tens of seconds inside one tool call, so
  silence alone proves nothing.
- The same screen check vetoes keystroke echo, which is a steady stream
  of repaints too but is not work.

CLAUDE_WORKING_LINE_PATTERN matches the `… (elapsed)` shape rather than
the glyph, because the FINISHED line (`✻ Cooked for 2m 49s`) carries the
same glyph and would otherwise pin a session at working forever.

Claude mode only. An external CLI has no `❯`, so nothing would arm the
confirmation and such a session would latch busy.

respawn-patterns.hasWorkingPattern() had the same blind spot (its gerund
list cannot see "Actualizing"), so it takes the pattern as an extra
signal. That can only make respawn less eager, never more.

Idle now lands about 3 to 5 seconds after a turn ends instead of 2
seconds into one. Verified end to end against a live worker, sampled
against the CLI's own "esc to interrupt" footer as independent ground
truth: busy for all 25s of a turn, idle 3s after it ended.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 15:31:02 +02:00
Codeman maintainer 9dc4620f03 fix(terminal): stop the scroll-to-top re-pull from deleting history, page the CLI when local scrollback is hollow (#205)
The 1.12.0 retest on #205 reported it still broken in two shapes: a wheel that
did nothing at all on Firefox/macOS (while Fn+Up paged back through intact
text), and iPhone history that went back a little, repeated blocks and got
worse the further up it went. Both come from a Claude pane's LOCAL buffer being
hollow: tmux keeps no history for a repaint-mode pane (history_size 0), so
xterm holds only replayed repaint frames.

1. The scroll-to-top full=1 re-pull now refuses a DOWNGRADE. It resets the
   terminal and rewrites it from the capture, which is a win when tmux holds
   more than the browser, but for a repaint-mode pane that capture is roughly
   ONE frame and the rewrite deleted history mid-scroll. Measured A/B on a live
   pane, same gesture: guard off collapses 341 rows to 42, guard on preserves
   all 341. _replayWouldShrinkBuffer() estimates the capture's rendered rows
   (escapes stripped, capture-pane -J re-wrapping accounted for) and skips the
   rewrite when it is more than one screen short; a refused session's cooldown
   goes from 4s to 60s so a hollow pane stops re-fetching megabytes.

2. A false forwarding gate on a Claude session no longer means a dead gesture.
   Under a triple guard (claude mode, gate false, baseY 0), wheel and touch
   travel becomes coalesced PageUp/PageDown through the same 40ms queue as the
   SGR reports, at half a screen of travel per page key. Shift is excluded: it
   keeps meaning "local scrollback".

3. getClaudeCliVersion() no longer caches FAILURE. It stored null on any
   exception and guarded on !== undefined, so one timed-out or PATH-starved
   probe at the first Claude session start disabled wheel-forwarding for every
   Claude session until the server restarted, which fits a report of breakage on
   phone, tablet and laptop at once. Success is still cached for the process
   lifetime; failures retry with a 1/2/4 up to 15min backoff, and the policy is
   a pure function so the semantics are testable without spawning claude.

4. The terminalWheelLocalScrollback footgun is handled by pairing rather than
   scoping: the setting keeps meaning exactly what it says, and fix 2 catches
   the case where "local" is empty. The App Settings tooltip now says to leave
   it off for Claude/Codex sessions.

5. _logScrollRouting() prints one line per session per distinct decision:
   forward-sgr / page-keys / local-scrollback / repull-refused-downgrade, with
   mode, cliVersion, the opt-out state, mouse tracking and local scrollback
   depth. #205 ran two rounds of remote guesswork over questions that line
   answers directly.

Verified end to end against a real isolated instance (own data dir and tmux
socket) with real wheel events: forwarding still sends SGR reports, the opt-out
now sends real PageUp/PageDown where the wheel was dead, a tab-switch collapse
(401 rows to 44) is still fully recovered by the re-pull (back to 401), and a
seeded 341-row Claude buffer survives the same gesture that destroys it with the
guard disabled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:41:39 +02:00
Codeman maintainer 0d0b772619 feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to
the surfaces around it, while Gemini CLI stayed documented as a consumer
product despite being enterprise-only since Google's cutover. Gemini keeps
full support; Antigravity now sits beside it everywhere.

Functional fixes:
- docker/agent.Dockerfile never installed agy, so a docker case with
  mode 'antigravity' died on command-not-found. agy is not on npm, so it
  gets its own installer step. --dir /usr/local/bin is load-bearing: the
  default $HOME/.local/bin resolves to root's home at build time and is
  unreachable by the `agent` user the container runs as. Verified inside
  codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is
  ~190MB, the largest layer in the image.
- Welcome screen gained a Run Antigravity action, gated on agy being
  present like the other CLI buttons, with a cyan identity matching the
  toolbar run button and run-mode dot.
- install.sh now detects agy (search paths mirroring the resolver), counts
  it as a satisfying AI CLI, and recommends it over Gemini in the install
  hints. Detection only, no new auto-install path.

Docs corrected where they were factually wrong:
- architecture-invariants documented isExternalCliMode() as
  opencode/codex/gemini when the code has included antigravity for a
  while, said "all three modes", and omitted ANTIGRAVITY_ from the env
  prefix allowlist row.
- cron-guide's agentType enum, cron-discovery's SessionMode, and
  remote-sessions' RemoteCommandMode were all stale.

Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only),
package.json keyword, and comment drift in 8 places.

test/run-mode-ui.test.ts now covers the new welcome button; verified it
fails without the settings-ui wiring.

Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not
~/.antigravity, so the existing .gemini docker credential seed already
covers it. Recorded as a comment so nobody adds dead config later.

isAltScreenStripMode() deliberately still excludes antigravity: whether
its TUI needs the alt-screen strip is a behavioural question that needs a
real agy session, not a guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-06 07:22:29 +02:00
Codeman maintainer 5d2899907e fix(cli-gating): gate the tunnel button instead of deleting it, and cover antigravity
Follow-up to #200 and #201, which gate the welcome buttons and the run-mode
dropdown on whether the CLI is actually installed. Four corrections:

1. #200 also DELETED the Cloudflare Tunnel welcome button and the QR widget
   outright. Its rationale is right (offering a tunnel where cloudflared is not
   installed is a bad default) but the conclusion overshoots: the welcome QR is
   the whole scan-to-connect-from-your-phone flow, and deleting it left a large
   block of live tunnel code in settings-ui.js driving elements that no longer
   existed. Both are restored and the button is gated on cloudflared, which is
   what the stated rationale actually asks for. New cloudflared-resolver.ts
   mirrors the CLI resolvers, and TunnelManager now shares its search path so
   the button and the spawn can never disagree about where cloudflared lives.

2. Antigravity was missing from the run-mode gating, the one run mode LEAST
   likely to be installed. It slipped past because #201 predates it. Covered
   now, plus a static test that fails if a sixth mode reaches the dropdown
   without being gated, so the next one cannot slip the same way.

3. The per-surface fetches are replaced by the injected availability object
   already used for the Codex settings tab, so the codebase has one mechanism
   rather than two. The status routes buy nothing as a gating source: every
   resolver memoizes its PATH probe server-side, so a fetch is exactly as stale
   as an injected value while costing a round trip every time the dropdown opens
   and leaving the welcome buttons to flicker in after paint. The routes
   themselves stay, including the /api/claude/status that #200 adds.

4. Unknown availability now reads as AVAILABLE for run buttons. Both PRs hid the
   button on a failed fetch, so a blip left a working install with nothing to
   click; a genuinely missing CLI only ever produced an error toast. The Codex
   settings TAB keeps the opposite default, since hiding it costs nothing.

The dropdown query is also scoped to the menu: `.run-mode-option` is the class
the saved-dashboard and history rows use too, and a document-wide querySelector
would have found whichever came first in the DOM.

Fixes a latent environment-sensitivity in 816d900 while here: the index-title
test asserted the template was untouched apart from the title, which held only
on a machine with no codex installed.

Verified end-to-end against a real server on an isolated instance+socket, with
Playwright: gemini/codex hidden and claude/opencode/antigravity/shell shown,
matching this host, tunnel button back, Codex settings tab still hidden, no
console errors. Full test:ci sweep green (3902 tests).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:53:43 +02:00
Codeman maintainer b54094a4c8 Merge pull request #200 from timkjr/pr/gate-gemini-drop-tunnel-button
fix(welcome): gate CLI welcome buttons on actual availability
2026-08-05 01:40:44 +02:00
Codeman maintainer 2b89f35599 fix(shell,remote-ssh): allowlist the login flags, and keep only CRASHED remote panes
Follow-up to #209 and #210. Both land a real fix (a pane that is a login shell
picks up /etc/profile and the per-user PATH entries an ssh remote command never
sees, which is what was failing agent CLIs with exit 127). Three corrections:

1. `-i -l` is no longer hardcoded onto the resolved shell. That path ultimately
   comes from the passwd entry, which is user data and can name anything, and a
   shell that rejects an unknown flag exits on the spot: nushell, elvish and xonsh
   take neither flag, so a user with one of those in passwd would have gotten a
   dead pane on arrival, which is exactly the #208 failure #209 builds on top of.
   loginShellArgs() applies them only to the POSIX-family shells verified to
   accept both, and a test really launches every allowlisted shell present on the
   machine rather than trusting the set. csh/tcsh are excluded deliberately: tcsh
   honors -l only when it is the ONLY flag.

2. `remain-on-exit on` -> `failed`, moved LAST in the tmux command chain. `on`
   keeps the pane after a CLEAN exit too, so typing `exit` in a remote shell
   stranded a dead pane, the session outlived it, and the next launch's `-A`
   reattached to that corpse: "Pane is dead (status 0)" instead of a shell,
   permanently, on the DEFAULT path. Verified against a real tmux, as was the
   fix: `failed` tears the session down on status 0 and keeps the pane on 127
   with the "command not found" still on screen, which is the case #210 wanted.
   It is last because tmux aborts the remaining commands of a `\;` sequence once
   one errors (also verified) and `failed` needs tmux >= 3.2 on the REMOTE host;
   leading, a rejection there would have silently dropped status/mouse/prefix/
   escape-time/window-size along with it.

3. `$SHELL` -> `"${SHELL:-/bin/sh}"`, via one shared remoteLoginShellCommand()
   helper instead of the string being rebuilt in tmux-manager as well.

Also corrects the rationale both PRs carried: a tmux pane already hands the shell
a tty, so it was interactive all along ($- contains i for a bare /bin/bash in a
pane) and ~/.bashrc was always being sourced. `-l` is the flag doing the work.

End-to-end verified, not just unit-tested: the emitted remote pane command was
run through all three quoting layers under a minimal sshd-style PATH with the
CLI installed only on a login-shell PATH entry, and it resolved and launched the
CLI with its arguments intact and a space-containing remote path preserved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 01:40:37 +02:00
timkjr 3ea1ea28f0 fix(welcome): gate Claude and Opencode buttons on CLI availability too
Extends the Gemini gating from bb7fb9e to the other welcome-screen
buttons that had the same problem: shown unconditionally even when the
underlying CLI isn't installed.

- Add isClaudeAvailable() (claude-cli-resolver.ts) and GET
  /api/claude/status, mirroring the existing opencode/codex/gemini
  resolvers and status endpoints.
- Opencode already had a working /api/opencode/status the welcome
  screen just wasn't checking; wire it up the same way.
- Refactor loadGeminiAvailability() into a shared
  _loadCliAvailability(buttonId, statusUrl) helper instead of
  duplicating the fetch/try-catch three times.

Run-mode dropdown entries (Opencode/Codex) are intentionally left
unconditional here — follow-up PR.
2026-08-04 16:22:21 -05:00
Codeman maintainer 529d8fa8ea chore: version packages 2026-08-04 23:09:00 +02:00
Codeman maintainer 23f258a85d chore: version packages
Release 1.9.8 (aicodeman) and 0.1.8 (xterm-zerolag-input).

Fixes macOS session start (`posix_spawnp failed.`, issues #6 and #204):
node-pty ships its macOS spawn-helper as mode 0644 and macOS launches every
PTY through it. `scripts/fix-node-pty.mjs` (npm run fix:node-pty) chmods every
helper, prebuilds/ included, then verifies by really opening a PTY; the blind
Node-22+ rebuild is gone. `spawnPtyWithHelperRepair()` self-heals an already
broken install on the first failed spawn.

Adds the phone home screen (session overview under 430px, per-device
`mobileOverviewEnabled`, default ON) and a guided Tailscale path in
install.sh, plus `install.sh tailscale` to retrofit it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 15:02:46 +02:00
Codeman maintainer 26cbbe0dcb feat(cli): Antigravity run mode
Adds Antigravity as a sixth CLI backend alongside Claude Code, shell, OpenCode,
Codex and Gemini, following the existing pluggable-resolver pattern.

- `utils/antigravity-cli-resolver.ts` resolves the CLI, mirroring the other
  resolvers; `GET /api/antigravity/status` reports availability and path.
- `ANTIGRAVITY_*` joins the `ALLOWED_ENV_PREFIXES` allowlist in schemas.ts, so
  env overrides stay CLI-scoped rather than blanket-forwarded.
- Session, tmux-manager, mux-interface and types carry the new mode; secrets are
  injected via socket-scoped `tmux setenv`, never on the spawn command line, so
  the mode requires tmux with no direct PTY fallback like the other external CLIs.
- Frontend: Run-dropdown entry, agent-type option, `ag` tab badge and toolbar
  colours. `runAntigravity()` routes remote/docker cases through
  `POST /api/quick-start` and skips the local status probe for them.

Tests: test/antigravity-mode.test.ts, plus run-mode-ui and system-routes coverage.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-04 12:59:40 +02:00
codeman-local 66ad681666 fix(paths): accept Unicode working directories 2026-07-21 00:04:40 +08:00
Codeman maintainer 7b79d4207c fix(terminal): restore Claude scroll-back on macOS trackpads (#154)
Deterministic claude --version probe seeds cliVersion so wheel-forwarding
to Claude's transcript engages (banner scrape was unreliable on 2.1.187+
and resumed sessions). Shift+wheel reads the dominant axis so a trackpad's
horizontal Shift-scroll reaches local scrollback. New per-device
"Wheel Scrolls Local History" opt-out. Wheel reports use a fire-and-forget
send path so they no longer flicker the pending-bytes indicator.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-16 09:33:39 +02:00
Codeman maintainer 368fc20fc2 fix: address PR review findings for Gemini run mode + Ralph todo-config
Gemini (PR #134) blockers:
- runGemini() now unwraps the {success,data} envelope: status check reads
  .data.available, quick-start reads data.data.sessionId (was reading the raw
  shape, so the Run-Gemini button could never start a session).
- setGeminiEnvVars() now uses the socket-scoped ${this.tmux()} setenv instead of
  bare tmux — Gemini/Google auth env vars were targeting the wrong tmux server
  and silently failing on every install.

Gemini parity polish:
- gemini tab-mode badge ('gm') + .tab-mode.gemini CSS; kill-dialog label
  'Kill Tmux & Gemini'; codeman doctor dependency-registry entry; export
  isGeminiAvailable from utils barrel; COLORTERM=truecolor + unset NO_COLOR;
  add gemini to isAltScreenStripMode (Ink TUI, repaints inline like Codex/Claude).
- Revert 4 system-routes.test.ts envelope assertions weakened to
  (body.message ?? body.error) back to (body.success === false).
- Add a runGemini() vm-sandbox test that drives the envelope path end-to-end.

Ralph todo-config (PR #135): maxTodos/todoExpirationMinutes are now persisted
and read back — surfaced via the loopState getter (RalphTrackerState) into
toState()/SSE broadcast and restored in restoreState(), mirroring maxIterations.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-25 00:26:06 +02:00
Aamer Akhter 19139837e4 feat: add Gemini run mode 2026-06-19 13:10:17 -04:00
Aamer Akhter 585127deb2 Add codeman doctor tool-dependency checker (COD-45)
Environment-aware dependency probe (linux|darwin|win32|wsl) with a static
registry, an injectable ProbeHost seam for testing, grouped table + `--json`
output, and a non-zero exit when a required dependency is missing/outdated.
Node and tmux are the only hard-required tools; the agent CLIs and document
converters (LibreOffice / MS Office via WSL interop) are optional. CI-safe
unit tests (no tmux, injected host).
2026-06-14 12:25:49 -04:00
Aamer AkhterandSaqeb Akhter 70378315da feat(codex): add Codex (OpenAI CLI) run-mode foundation
Add Codex as a first-class session mode alongside Claude/Shell/OpenCode.

- codex-cli-resolver: locate the `codex` binary and augment PATH (mirrors the
  OpenCode resolver)
- SessionMode 'codex' + CodexConfig (model, resumeSessionId, dangerouslyBypass,
  renderMode); persisted in SessionState and threaded through CreateSession/
  RespawnPane options
- schema validation: CodexConfigSchema, CODEX_ env-var prefix allowlist, mode
  enums on create/quick-start, codexDangerouslyBypassApprovals setting
- tmux launch: buildCodexCommand, setenv for OPENAI_API_KEY/CODEX_* (keeps
  secrets out of ps), truecolor COLORTERM, codex PATH resolution
- session + routes: availability check (clear install hint), config passthrough,
  tmux-required guard; Codex skips Claude-only parsers (Ralph/respawn/token)
- run-mode UI: "Run CX" selector option + dedicated Codex CLI settings tab with
  the bypass-approvals toggle; GET /api/codex/status

Scope: foundation only. Codex terminal redraw handling and xterm snapshot/replay
are intentionally excluded and tracked separately.

Verification: tsc --noEmit, eslint, prettier --check, check:frontend-syntax all
clean; full test:ci suite green (2712 passed, 0 failed); server boot smoke OK.

Co-Authored-By: Saqeb Akhter <saqeb.akhter@gmail.com>
2026-06-10 09:34:52 -04:00
arkonandClaude Opus 4.8 b84438a0aa fix(security): push-endpoint SSRF guard + tmux name validation; document tail-file roots
- M7 (SSRF): add isSafePushEndpoint (https-only; reject internal/loopback/link-local/metadata IPs incl. IPv4-mapped); enforce in PushSubscribeSchema and re-check before webpush.sendNotification. + unit test.
- M1 (command injection): validate tmux session names with isValidMuxName in sessionExists, killSession, and reconcileSessions before they reach a shell call site.
- M5: keep the intentional /var/log + ~/logs log-tail roots (a tested feature) and document the wider read scope in docs/security-architecture.md section 5 instead of dropping it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-09 20:02:17 +02:00
Tenggan ZhangandTeigen 06f9ff6d9c fix: avoid event-loop stalls from synchronous tmux/ps calls (#100)
The stats collector (~2s) and mouse-mode sync (5s) ran execSync (pgrep/ps/
list-panes, 5s timeout each) per session on the server's single thread,
blocking the event loop. With several sessions or a momentarily slow tmux this
froze port 3000 for seconds-to-tens-of-seconds while the process stayed alive
and other ports were unaffected — self-healing, so it never restarted and the
60s loopback healthcheck missed it. Convert these hot-path calls to execAsync.

Also add an always-on event-loop lag monitor (utils/event-loop-monitor.ts) that
logs stalls >=1s to the web log, so this otherwise-invisible class of incident
leaves a quantified, timestamped trace.

Co-authored-by: Teigen <teigen@TeigendeMac-mini.local>
2026-06-01 19:32:24 +02:00
arkonandClaude Opus 4.7 02e2f3e8b5 refactor: remove dead code and narrow internal exports (knip sweep)
Knip-driven cleanup. All changes verified with tsc --noEmit, lint, and
build.

Removed (zero consumers):
- VERIFICATION_PROMPT constant + its barrel re-export
- createInitialOrchestratorPersistState factory
- transcriptWatcher singleton export
- createAnsiPatternFull / createAnsiPatternSimple factories
- TimerInfo interface + unused AiCheckResult/AiPlanCheckResult imports
  in respawn-controller.ts
- 35 unused Zod z.infer \`*Input\` types in schemas.ts
- Dead re-exports: SessionMode from session.ts, AuthSessionRecord from
  web/ports/index.ts, EnhancedPlanTask/CheckpointReview from
  ralph-tracker.ts, 7 unused entries in utils/index.ts
- 14 event/config interfaces that lived only as JSDoc hints (no TS type
  position usage): Session/Respawn/RalphLoop/RalphTracker/
  SessionManager/SessionAutoOps/Subagent/TaskQueue/TaskTracker/
  TranscriptWatcher/Image/OrchestratorLoop Events + RespawnPreset +
  SessionOutput

Narrowed to module scope (kept but no longer exported):
- buildPermissionArgs in session-cli-builder.ts
- 28 type/interface declarations used only within their own file:
  Ai{Idle,Plan}Check{Config,State}, BashToolParser{Events,Config},
  FileStream/CreateStream{Options,Result}, PlanSubagentEvent,
  SubagentCallback, RalphLoopConfig, RalphLoop{Events,Options},
  ActiveTimerInfo, DetectionStatus, ActionLogEntry, AutoOpsCallbacks,
  TunnelStatus, Timer/LRUMap/StaleExpirationMap Options, AuthState,
  SessionListenerDeps, SseStreamManagerDeps, and 8 more

Docs: CLAUDE.md advice for global-regex `lastIndex` now points to the
remaining `execPattern()` helper instead of the deleted factories.

Knip delta: unused files 42→0, unused exports 161→16, unused types 92→0.
The 16 remaining exports are a mobile-test helper toolkit intentionally
kept for upcoming tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-23 11:57:00 +02:00
arkonandClaude Opus 4.6 0717cfbfec refactor: codebase cleanup — dead code, regex helper, centralized constants, tests
- Remove unused validateTokenCounts/validateTokensAndCost exports and PlanPhase type alias
- Add execPattern() helper to eliminate 8 repetitive .lastIndex=0 + exec() loops
- Centralize 11 magic number constants into config/ai-defaults.ts and config/server-timing.ts
- Remove stale src/tui from tsconfig.json exclude
- Fix CLAUDE.md inaccuracies (session helpers, app.js line count, module count)
- Add 316 new tests: LRUMap (38), StaleExpirationMap (42), BufferAccumulator (33),
  respawn helpers (142), system-routes expansion (11→61)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-07 07:33:44 +01:00
arkonandClaude Opus 4.6 e0a2774d37 perf: implement phase 1-3 performance optimizations
Add implementation plans and code structure analysis for a 3-phase
performance optimization effort. Refactor core modules to reduce
timer overhead, consolidate regex usage, extract exec timeout config,
add debouncer utility, and streamline server/schema validation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 17:39:44 +01:00
arkonandClaude Opus 4.6 545913790b chore: remove dead code and unused exports (-159 lines)
Remove confirmed-unused code found via team-based codebase analysis:
- Dead interfaces: SessionResponse (types.ts), PlannerResult (plan-orchestrator.ts)
- Dead functions: withTimeout (session.ts), getOpenCodeAugmentedPath, resetOpenCodeCache (opencode-cli-resolver.ts)
- Dead config constants: MAX_RUN_SUMMARY_EVENTS, TRIM_RUN_SUMMARY_TO, MAX_MESSAGES_PER_CHANNEL (buffer-limits.ts), MAX_TRACKED_AGENTS, MAX_SUBAGENT_ACTIVITY_PER_AGENT, MAX_TOOL_RESULTS_PER_AGENT, COMPLETED_TODO_TTL_MS (map-limits.ts)
- Dead legacy code: idleTimer property + clearIdleTimer() no-op method + 6 call sites (respawn-controller.ts)
- Stale barrel re-exports: createAnsiPatternFull/Simple, validateTokenCounts/Cost, stripAnsi, normalizePhrase, getOpenCodeAugmentedPath (utils/index.ts)
- Un-exported internal-only symbols: ErrorMessages (types.ts), destroyRalphLoop (ralph-loop.ts)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-28 02:01:21 +01:00
480584de63 chore: project setup and tooling cleanup (#26)
* add project setup

* add tools to eslint

* remove contributing.md

* update claude md

* chore: simplify CI to single Node.js version (22)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* use node 20

* adjust formatting

* fix linting

* format

* fix layout

* fix: align CI node version to .nvmrc (22), remove redundant gotcha

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: arkon <arkon.85@hotmail.com>
2026-02-27 13:16:23 +01:00
arkonandClaude Opus 4.6 0b60f2691f feat: OpenCode integration Phases 1-3 — types, CLI resolver, tmux spawn
Phase 1: Type system
- Add SessionMode = 'claude' | 'shell' | 'opencode' to types.ts (single source of truth)
- Add OpenCodeConfig interface (model, autoAllowTools, continueSession, etc.)
- Add openCodeConfig to SessionState
- Import SessionMode in session.ts, mux-interface.ts (remove duplicates)
- Refactor createSession/respawnPane to options objects (CreateSessionOptions, RespawnPaneOptions)

Phase 2: CLI resolver
- New opencode-cli-resolver.ts mirroring claude-cli-resolver.ts
- Searches ~/.opencode/bin, ~/.local/bin, /usr/local/bin, ~/go/bin, etc.
- Exported via utils/index.ts

Phase 3: Tmux spawn
- Extract buildSpawnCommand() shared helper (eliminates duplication)
- Add buildOpenCodeCommand() for opencode CLI flags
- Add setOpenCodeEnvVars() — API keys via tmux setenv (not visible in ps)
- Add setOpenCodeConfigContent() — JSON config via tmux setenv (prevents shell injection)
- Gate _processExpensiveParsers() for opencode mode (skip Claude-specific parsers)
- OpenCode-specific TUI ready detection (timeout-based, no ❯ prompt scanning)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-02-26 10:07:43 +01:00
arkon 0fe6251424 chore: bump version to 0.1585 2026-02-21 07:12:57 +01:00
arkon 926e77b1c2 chore: bump version to 0.1550 2026-02-19 08:49:36 +01:00
arkonandClaude Opus 4.6 d7ad617d7f chore: bump version to 0.1545
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 21:53:34 +01:00