Commit Graph
281 Commits
Author SHA1 Message Date
Ark0N 097d585278 Merge pull request #383 from opticon454/fix/case-picker-default-root
fix(file-picker): default the case picker to Codeman Cases, not Home
2026-09-06 21:03:00 +02:00
Codeman maintainer 8ad2215118 fix(docker): close the three adoption gaps the negative guarantee missed
Review follow-ups to #357. Each is a path that still touched, or still hid, a
container Codeman does not own.

**Export still mutated it.** The four fail-closed layers cover create/start/
stop/remove, but `POST /api/docker-cases/:name/export` reaches the container
twice through neither: a full export `docker commit`s it, and even a
workspace-only export `docker pause`s it first for snapshot consistency. Pause
freezes the owner's processes for as long as the tar takes, on a container we
promised not to touch. Full export is refused for an adopted case (it packages
someone else's container, with their logins, into a bundle Codeman hands out);
workspace-only keeps working and no longer pauses, accepting a live filesystem
the way `tar` does on any running host directory.

**A freshly linked OWNED case became unusable.** The run menu now probes the
container for its CLIs, and a failed probe hides every agent mode behind the
reason. For an adopted case that is right. For an owned one the container does
not exist until the first session launches it, so every newly linked Docker case
answered `container "codeman-case-x" not found (adoption never creates a
container — start it yourself first)` and offered nothing but Shell, for a
container the launch chain was about to create itself. A failed probe is
recorded only when the case is adopted; `CaseInfo.docker.owned` is on the wire
so the frontend can tell them apart. Verified in a browser: owned-with-no-
container offers all ten modes and no notice, adopted-but-stopped offers Shell
and says why.

**Multi-user gating.** Adoption is admin-only, unlike `docker-link` beside it.
Linking creates OUR container, whose sole bind mount `isWorkingDirAllowed` has
already confined to the caller's space; an adopted container's mounts are
whatever its owner gave it, so one mounting `/` hands the adopter a shell over
the whole host — exactly the workspace scoping multi-user mode exists to
enforce. Listing the engine's containers and browsing directories inside an
arbitrary one are machine-level reads and follow the docker-HOST policy for the
same reason. The preflight is deliberately not admin-only: the run menu fires it
for every docker case, so it admits a non-admin for a container already linked
to a case they can access, and nothing else.

Verified end to end against a real pre-existing root container (alpine + tmux,
no bind mounts): adopt, claude session inside it, workspace export, session
close and case unlink all left `StartedAt`, `RestartCount`, `Pid` and `Paused`
untouched; the pane ran the CONTAINER's claude, without
`--dangerously-skip-permissions`; a stopped container was refused at both
preflight and launch and was never started.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:22:54 +02:00
Codeman maintainer 3d8ffcb9a2 Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container

Conflicts came from work that landed after the PR was opened, and each is
resolved onto the newer abstraction rather than by keeping the older code:

- `defaultDockerCommandForMode` is registry-driven since #347, so the PR's
  `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude
  Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the
  refusal is visible only inside the container, so an adopted root container
  otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and
  `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch.

- The probe's mode list and its mode -> binary table both duplicated the
  registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is
  also what fixes the merge's silent regression: the hand-written list predates
  `omp`, and the run menu gates every docker case on this probe, so owned
  containers would have lost that mode. `shell` needs no arm — it declares no
  binary, so it is dropped from the lookup and reported available regardless.

- The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one
  `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto
  it. Its test now pins the single gate instead of counting seven arms.

- The create arm keeps #349's swap-limit warning filter, which the adopted arm
  never reaches; the run-mode list gains `omp` from #353.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1
2026-09-05 16:21:51 +02:00
DevvynandClaude Sonnet 5 06febfa032 fix(file-picker): default the case picker to Codeman Cases, not Home
The "Link Existing" case picker opens with an empty path and no
sessionId, so the browse endpoint's fallback root picked whichever
root happened to be first in the list — which was always `Home`.

On the native default that's harmless (~/codeman-cases nests inside
Home anyway), but a Docker deployment binds CODEMAN_APPDATA_PATH
(Home) and CODEMAN_CASES_PATH at unrelated host paths, so the picker
opened somewhere with no cases in sight. Worse: if CODEMAN_CASES_PATH
is ever changed after cases already exist, the old cases directory
lingers, still reachable, under Home — indistinguishable at a glance
from the real one under the new Codeman Cases root.

Prefer the Codeman Cases root in the fallback chain, ahead of the
generic roots[0].

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9ZSTEenc8soSu9bTi8Xru
2026-09-05 17:14:05 +08:00
Codeman maintainer 2ab21c1b32 fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring
claiming logout called it and no caller at all. The capability is a bearer
credential exempt from cookie auth with a rolling TTL refreshed on every use, so
a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose
referrer policy) stayed valid for as long as anything kept polling it.

- POST /api/logout revokes the caller's capabilities (all of them in single-user
  mode), the admin forced logout revokes the target user's, and user deletion
  revokes whatever that user had open. revokeOwner returns the count for the
  admin audit line.
- Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own
  policy is dropped: every URL inside the frame carries the capability, and a
  dashboard on no-referrer-when-downgrade or unsafe-url handed it to any
  third-party host it linked. Verified with Playwright that a sandboxed frame
  under an upstream `unsafe-url` sends no Referer to a third party while the
  root-absolute fetch and the CSS-triggered 404 fallback still reach the
  dashboard.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:13 +02:00
Codeman maintainer 550e08a791 fix(webview): refuse link-local and cloud-metadata targets on the resolved address
The web-tab proxy, its Test probe and its WebSocket relay accepted any http(s)
host. A live PoC relayed an IMDSv2-shaped PUT with custom headers to a loopback
echo server through a capability and no cookie, and 169.254.169.254 (decimal,
hex, IPv6-mapped, or via a DNS name) was as valid a dashboard as any other.

Loopback and RFC1918 stay allowed on purpose: a localhost Grafana is the feature.
Only link-local and the fixed cloud-metadata addresses are refused
(169.254.0.0/16, fe80::/10, fd00:ec2::254, 168.63.129.16, 100.100.100.200,
metadata.google.internal), at three stages that are each load-bearing:

- the Zod schema, so a save gets a clear refusal;
- a synchronous hostname check at every connect site, because net.connect skips
  DNS for an IP literal and a lookup hook never sees one;
- a `lookup` hook on an undici Agent (webviewFetch) and on the ws client, which
  judges the RESOLVED addresses of a name and refuses when any is blocked. This
  is what closes DNS rebinding, which a hostname-string check cannot.

Adds undici@^6 so the proxy runs the package's own fetch with the package's own
Agent; a package Agent handed to Node's bundled fetch can mismatch protocols.

Verified live on an isolated beta: 169.254.169.254.nip.io (a real name resolving
to the metadata address) is refused by probe, proxy (403) and WS relay (4003),
while 127.0.0.1.nip.io still passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE
2026-09-04 15:21:12 +02:00
Codeman maintainer 823f56a243 Merge pull request #373 from opticon454/feature/docker-self-update
feat(docker): restore in-app self-update in the Compose dep
2026-09-04 14:25:56 +02:00
DevvynandClaude Opus 5 66eb01ba8f feat(docker): restore in-app self-update in the Compose deployment
Codeman running under docker/docker-compose.yaml lost the ability to update
itself from App Settings -> Updates. The image had no .git (excluded by
.dockerignore), so the install reported as "unknown"; there was no init system
for detectSupervisor() to find; the runtime stage had neither devDependencies
nor a build toolchain; and a pull into the baked /opt/codeman would have landed
in the container's writable layer and been discarded by the next `up`.

Restore it through configuration rather than a second updater, so the release
channel, auto-stash, status file and boot reconcile are all reused unchanged:

- The checkout Compose builds from is bind-mounted over /opt/codeman, so the
  update's git checkout and rebuild land on the host and survive recreation.
- The restart is the server exiting; `restart: unless-stopped` relaunches the
  container on the new dist/. This is the one supervisor whose updater does NOT
  outlive the restart, which is safe only because the terminal "restarting"
  marker is written first.
- node_modules and dist are named volumes over the bind mount, so
  container-compiled native modules never enter the host checkout.
- The runtime image keeps devDependencies and gains python3/make/g++, since
  `npm run build` is tsc + esbuild and node-pty has no Linux prebuild.

An in-place container update applies code only, because a restart reuses the
existing image and config. evaluateEnvironmentGate() reads the target release's
own files with `git show <tag>:<path>` and refuses when server.Dockerfile or
docker-compose.yaml changed, when .env.example gained keys the user's .env
lacks, or when the restart policy would not bring the container back. The
missing-key check matters most: Compose resolves an unset ${VAR} to the empty
string and starts anyway, so a new required setting would otherwise arrive as a
silently blank variable. Every unknown fails open, and the gate is re-evaluated
server-side on POST /api/system/update.

The four global agent CLIs are pinned, because an unpinned CLI bump is the one
environment change no diff-derived gate can see; pinning turns it into a
Dockerfile change the gate already detects.

Adds test/docker-compose-env-parity.test.ts as the merge-side guard (every
compose ${VAR} has an .env.example entry and the reverse) and
test/docker-self-update.test.ts for the pure gate decisions.

Documented in docs/docker-self-update.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013yAQ2y9t81jzSfpStUxx5T
2026-09-02 19:33:32 +08:00
DevvynandClaude Opus 5 4830e662f9 refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery
(search dirs, version + identity probes), the launch argv template, env
handling, the `capabilities` flags that replace per-CLI branching, and the
`overlays` that back the remote/docker pane commands. Code that used to ask
"which CLI is this?" reads the entry instead.

Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every
spawn command as a literal string, captured from the hand-written builders
before they were deleted, and `test/location-overlay-commands.test.ts` does the
same for all 20 remote and in-container pane commands.

Config can never contain shell text: an entry declares typed argv tokens,
literals are validated against a safe-word pattern at LOAD time (a bad literal
rejects the whole entry — a silently dropped `--no-approve` is not cosmetic),
and values resolve through patterns NAMED in code, so a user `clis.json` cannot
widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only
in this release.

OMP is included as a registry entry rather than a tenth hand-written builder,
so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of
`buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen
and doctor ladders all drop out.

Guard rails:

- `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id
  branching reappears outside `stock.ts`, in any of its four shapes (`===`,
  `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the
  negated forms, which is how 36 of them survived an earlier pass. Every
  allowlisted branch carries its reason.
- `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities;
  deriving one from another shipped the `until=stop`-hangs-on-shell bug.
- `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and
  `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config`
  wire field is separate, bridged only by `legacyConfigAliases`. Getting
  `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass
  clamp's only handle on a CLI's privilege switch, and a wrong name clamps
  nothing with no error and no failing test — so `schema.ts` rejects an entry
  naming a param it never declared.
- Registry data resolves AT CALL TIME (`sessionModeSchema()`,
  `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs`
  thunks). A module-level const freezes at first import, so a CLI enabled while
  the server ran moved the run menu but not that surface.
- Six fields are annotated DECLARED-FOR-LATER and read by nothing
  (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/
  `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed
  rather than measured. A test pins the list so it cannot quietly grow.

Three user-visible changes, all deliberate and named:

- `probeDockerCliVersion()` derives the in-container binary from the registry
  rather than assuming it equals the mode name (`antigravity` runs `agy`).
- The remote CLI version probe now covers grok and deepseek, which the
  hardcoded map it replaces omitted while its own comment said the rule was
  "every mode except shell".
- `codeman doctor`'s CLI rows are generated from the entries, so Claude's
  install hint is the install command rather than a docs URL, five CLIs gain
  hints they never had, and the row order follows the catalog.

Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars
(matching the `cliId` pattern) before its failure message quotes the value
back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading
the hand-editable `clis.json`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ
2026-09-02 08:26:45 +08:00
Codeman maintainer 2a32b5064a Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases
as SIBLING containers through the mounted host socket (Docker-outside-of-Docker).

Resolved the README conflict (master had grown to eight CLIs since the branch
was cut) and moved the Compose blurb out of the feature bullets into Quick
Start, next to the other ways of starting Codeman.

Three review findings from the PR discussion are fixed here rather than left
for a follow-up, because two of them are shipped-image problems:

- `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against
  the whole context-relative path, so `docker/.env` — which the deployment's own
  README tells the user to fill with CODEMAN_PASSWORD and provider API keys —
  was picked up by `COPY . .` and baked into the image at
  /opt/codeman/docker/.env. Verified in both directions against a real build
  context: with a canary secret in docker/.env, the unfixed ignore file lets
  /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for
  the checked-in .env.example) leaves only the example behind.
- `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which
  still hardcoded ~/codeman-cases, so `codeman skill install --case <name>`
  reported "Case not found" on exactly the deployment the override exists for.
  Both now resolve through config/cases-dir.ts. state-store.ts keeps its own
  literal on purpose: that one migrates the historical ~/claudeman-cases
  directory by name and is about the old default, not the active location.
- CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the
  three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore
  joins the documented list of files that genuinely belong in the repo root.

The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug
fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/
deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's
lastClaudeSessionId was handed to every non-claude CLI.

Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax,
public assets, lockfile.
2026-09-01 11:32:03 +02:00
Codeman maintainer d5b5f8f618 fix(docker): make the dsh profile install survive pnpm's build-script gate
Follow-up to #350, which fixed the actual blocker (issue #352): `dsh plugin` is
a thin forwarder that `spawnSync`s a literal `pnpm` with no npm fallback, so an
image without pnpm dies at exit 127 and takes the whole build with it.

That PR also pinned an allowlist of the two packages whose lifecycle scripts
pnpm blocked at the time. Replace it with a policy that cannot go stale: pnpm,
unlike npm, refuses dependency build scripts by default and FAILS the install
over it (`ERR_PNPM_IGNORED_BUILDS`, exit 1, measured on pnpm 11.24), and the
names to allow move between rebuilds because `@deepseek-harness-tui/dsh-tui` is
resolved by dist-tag, not pinned: 0.9.3 pulled `@google/genai` (whose script is
a literal `preinstall: no-op`), 0.10.0-beta.x does not. An allowlist of two
names would have let the next tree break the build the same way. Allowing them
wholesale is also the exposure this image already accepts three layers up,
where `npm install -g` runs the install scripts of every transitive dep of the
five CLIs above with no gate at all.

Also correct a comment in the `/api/deepseek/install-profile` route that
asserted the opposite of what #352 proved ("dsh bundles its own package
manager, so no system pnpm is required"). The route's behavior is already
right: dsh's own "pnpm not found on PATH" stderr reaches the caller as the
OPERATION_FAILED detail, so the UI's "add a terminal profile" button names the
fix. Documented the prerequisite in docs/deepseek-integration.md, and taught
the docker-cases image smoke test about `dsh`/`omp` plus the profile check that
`dsh --version` does NOT cover.
2026-08-31 22:26:09 +02:00
Codeman maintainer da91b4353b Merge pull request #353 from timkjr/omp-mode
feat: add OMP (Oh My Pi) as a new session backend
2026-08-30 16:16:26 +02:00
d fei 47ee49128c style: match the prettier version the lockfile pins
Format check failed twice, on different files each time, because three prettier
versions were in play: package.json says ^3.4.0, package-lock pins 3.8.3 (CI runs
npm ci, so that is the one CI uses), and the local node_modules had 3.9.6. Files
formatted with 3.9.6 were then "fixed" with 3.4.2, pushing session-routes and
system-routes onto a third style — every version change moved the failure to a
different set of files.

Line-break placement in `await import` and a union type only; no logic changes.
2026-08-29 23:49:04 -07:00
d fei 8e5e207386 fix(docker): send the probe body as an object, and explain an unreachable container
The run menu still offered every mode for an attached container. The browser's
actual request showed why:

  POST /api/docker-cases/adopt-preflight -> 400
  {"error":"Invalid input: expected object, received string"}

_api serializes `body` and sets Content-Type itself, and three call sites each
passed an already-stringified body, so it was encoded twice and the server saw a
JSON string where it expects an object. curl was fine throughout, so nothing in
the server logs pointed at it.

Also fixes the design defect underneath: a failed probe fell through to "do not
gate", which silently offered every mode. When the container has been recreated,
is stopped, or the engine is unreachable, the user sees claude, clicks it, and
it can only fail — with the reason visible nowhere. A failed probe now hides
every agent mode (Shell needs no CLI and stays) and shows the server's own
reason at the top of the menu.

Two static guards switched from a character window to brace matching. They
sliced between two call sites, and _loadRunModeHistory's call appears above its
definition, so the slice came out empty and the assertion verified nothing —
the same trap twice in one file.
2026-08-29 23:31:28 -07:00
d fei 5452ad5c5a feat(docker): add a folder picker to both path fields
Both paths in the adoption form had to be typed. Each gets a Browse button
using the same path-input-group markup Link Existing uses, so the two look and
behave alike.

What they can browse differs, and that is the point. The host workspace path
reuses the existing host picker. The container workdir cannot: an adopted
container has nothing mounted at a matching host path, so a host listing would
be a different filesystem — and getting this field wrong is the source of the
opaque OCI chdir error at launch, which makes it the field that most needs to
be clickable.

Adds a read-only POST /api/docker-cases/browse: one `ls` through docker exec, no
writes, no lifecycle, path shell-escaped like every other value. `ls -Ap` marks
directories with a trailing slash and keeps names with spaces intact.

PathPicker takes an optional fetchListing source rather than being forked: the
container variant only swaps where the rows come from, and reuses the rendering,
navigation, Up and Choose/Select unchanged.
2026-08-29 23:31:28 -07:00
d fei 23ab2e77fd fix(files): give the picker a root when the server runs as root
Link Existing's Browse did nothing: GET /api/filesystem/browse answered 403
"No filesystem browse roots are available".

Two rules were fighting. /root is a default blocked tree in the attachment
guard, and Codeman running as root — containers, plenty of servers — makes
homedir() exactly /root, so the picker's own allowlisted Home root was blocked;
the other candidates live under it or do not exist. The root list came out
empty and there was nothing the user could open.

The blocked trees exist to keep ~/.ssh and friends out of reach, not to seal off
the user's own home. Only trees that would swallow a configured root whole are
dropped now: /root goes when Home is it (or sits inside it), /etc holds no
configured root and is untouched. Secrets stay protected — isSensitivePath
independently matches .ssh/, .env and credentials* at any depth, and it is what
the directory probe asks about.

⚠️ Navigation must reuse the same narrowed list the roots were chosen with.
Handing the raw trees downstream admits a root and then refuses every path
inside it, which reads as a picker that opens and does nothing.
2026-08-29 23:31:28 -07:00
d fei 3685ad85bc fix(docker): stop requiring the CLI on the host for a container session
Attaching a container, picking claude and hitting Run gave one line —
`execvp(3) failed.: No such file or directory` — and the run-mode menu offered
every mode. Three separate defects, found on a real deployment.

TmuxManager.createSession resolved the CLI directory without distinguishing a
docker session, so a host with no claude threw, the catch fell back to a direct
PTY, and that PTY exec'd the CLI on the HOST. The failure surfaced as a bare
execvp error naming nothing. A docker session runs its CLI inside the container;
the host does not need it. All eight modes now sit behind a cliRunsInContainer
guard, and whether the container has the CLI is settled by the adoption
preflight or the image gate before launch.

The running check used a bare double quote and command substitution. The whole
chain is embedded in an outer `bash -c "…"`, so the unescaped quote closed that
string early and the remainder was re-tokenized. It is now a `grep -qx` pipeline
using only the single-quote form every other line in the builder already uses.

Claude Code refuses --dangerously-skip-permissions as root. Our base image runs
a non-root user, so an owned container never hit this; an adopted container's
user belongs to its owner and is frequently root, and keeping the flag killed
the pane with a message visible only inside the container. The preflight now
reports runsAsRoot and the launch chain drops the flag for it.

The menu also showed every mode because the container CLI probe only started
when the menu opened. It is warmed when the case is selected instead.
2026-08-29 23:31:28 -07:00
d fei 2f83a37c6d feat(docker): take run-mode availability from the container
The run-mode dropdown hides CLIs that are not installed on the HOST (#201). That
is right for local sessions and wrong for a container case, whose agents run
inside the container: a host with no claude installed hides the mode while the
container ships one, which is exactly what happened on a real deployment.

The adoption preflight already probes what the container has, so that result is
persisted on the case and surfaced through CaseInfo. Docker cases gate on it;
every other case keeps the host probe unchanged.

An absent list reads as "do not gate" rather than "nothing available": an owned
container runs our base image, which ships every CLI, and treating unknown as
empty would leave the menu with Shell alone.
2026-08-29 21:02:19 -07:00
d fei 34c12ca18b feat(docker): make the container field a picker you can also type into
Typing a container name from memory is error-prone. The field becomes a native
datalist: pick from the engine's containers, type to filter, or type a name that
is not listed (the engine may be remote, or the container may not exist yet).
A datalist gives all three natively, so no dropdown state machine is introduced.

Adds listDockerContainers and GET /api/docker-hosts/:hostId/containers, following
the listRemoteCodemanSessions discovery precedent: read-only and never throwing,
so an unreachable daemon returns an empty list and the field degrades to plain
text instead of erroring.

Stopped containers stay in the list, sorted after running ones and labelled.
Attaching does require a running container, but hiding stopped ones turns "my
container is not in the list" into a dead end, while showing
`Exited (137) 8 days ago` says exactly what to fix.
2026-08-29 21:02:19 -07:00
d fei c98a59d709 fix(docker): verify the container workdir and end the probe with exit 0
Two defects that only a real container exposes.

The probe chained `command -v X && echo X` with semicolons, and a script's exit
status is its last command's. A container without the last probed CLI made the
whole `sh -lc` exit 1, so a perfectly healthy container with tmux and claude was
reported as "could not exec into the container". A missing CLI is data here, not
failure, so the script now ends with `exit 0`.

containerWorkdir defaulted to hostWorkspacePath. That default holds for an owned
container only because the create-time bind mount puts the host directory at that
exact path; attaching mounts nothing, so the two are independent facts. A host
path absent inside the container makes `docker exec --workdir` fail with an OCI
chdir error that surfaces in the pane as a bare "execvp failed". The preflight now
proves the directory exists inside the container and refuses at link time.
2026-08-29 21:02:19 -07:00
d fei 15eebde832 feat(docker): attach a case to an already-running container
Docker cases could only run in a container Codeman created itself. Attaching to
one the user already built and runs means Codeman must leave that container's
lifecycle completely alone, which the launch chain could not do: it was
`image inspect` -> `inspect || create` -> `start` -> `exec`.

Adds `DockerCase.owned`, mirroring the `owned:false` contract remote-SSH already
uses for attached sessions. Absent (every existing case) means owned, so current
behaviour is byte-identical. `false` means the container belongs to the user and
Codeman may only exec into it.

The launch chain for an attached container only looks, then execs: no image gate
(the image is theirs), no create, and no `start` — starting a container we do not
own is the very mutation attaching promises not to perform. A missing or stopped
container fails closed with an actionable message instead. Credential seeding is
skipped too: those copies read from create-time read-only mounts that do not
exist here, and writing host credentials into someone's container is not ours to
do, so its CLIs must already be authenticated inside it.

Four fail-closed guards. buildDockerStopCommand and buildDockerRemoveCommand
throw during pure string construction, so no caller bug can turn into a
`docker stop`/`rm` on a container we do not own; removeDockerContainer refuses
again at the lowest layer; drift reports "none" for an attached container, which
carries no `codeman.confighash` label and would otherwise always look drifted and
409 the launch gate forever; and the orphan reaper skips attached containers
through a check deliberately independent of the two conditions already covering
them.

`owned` is applied AFTER the config hash is computed. dockerConfigHash takes an
explicit field list, so ownership can never shift an existing case's hash — if it
did, every pre-existing case would trip the drift gate at once, and the remedy
the UI offers is "recreate the container".

Adds POST /api/cases/docker-adopt and a read-only
POST /api/docker-cases/adopt-preflight. The preflight refuses at LINK time rather
than at session launch, where the only ways out would be a dead pane or starting
a container we do not own.

Tests assert the negative guarantee directly — that create, start, stop, rm,
restart and kill are absent from the generated commands while `docker exec -it`
and `new-session -A` remain — since it cannot be observed by using the feature.
2026-08-29 21:02:19 -07:00
timkjr 65e994d29a fix(omp): correct docs/counts/URLs, resolver install-path order, stray comment + CSS
Small cleanup items from upstream review (Ark0N/Codeman#353):

- OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real
  installer target (~/.omp/bin was an earlier unverified guess, confirmed
  wrong against a real --no-cache Docker build).
- docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp ->
  can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth
  SessionMode incl. shell -- not eighth), matched the install-path guidance
  to the resolver fix, updated the version example to the actually-tested
  18.0.8, and added a Docker-section caveat: --resume pinning does not
  currently reach an in-container omp process, since Docker panes never see
  ompConfig.
- docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md
  already linked to the -omp anchor, so the link was dead) and added an OMP
  specifics paragraph -- the one external CLI missing an entry in this doc.
- .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi,
  Grok, and DeepSeek Harness) and the backend count.
- Removed a stray orphaned comment fragment in the quick-start docker branch
  and split two CSS lines that had two declarations jammed onto one line.
2026-08-28 14:37:28 -05:00
timkjr f18dccace1 fix: don't discard codex/gemini/antigravity conversations on Resume; fix DELETE ownership dup + missing broadcast
resumeHistorySession() creates the resumed row in its own mode via a
modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek ->
resumeSession) and retires the old row afterward. codex, gemini and
antigravity were missing from that map, so resuming one of their rows
started a brand-new session with NO continuation while still deleting
the row it came from -- silent data loss dressed as the duplicate-row
fix. Gate row retirement on continuesSomething (true only for modes that
actually got a continuation config) instead of wiring an unverified
sessionId->native-conversation-id assumption for the three affected CLIs.

DELETE /api/sessions/:id reimplemented the ownership 404 check inline in
two places instead of going through findSessionOrFail, and its
persisted-only-session branch never broadcast session:deleted, so other
open tabs kept the retired row until their next unrelated fetch. Extract
the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail()
alongside findSessionOrFail() in route-helpers.ts (same ownership
contract, returns a SessionState instead of a live Session), and use both
from the route instead of inline checks. Add the missing broadcast.
2026-08-28 14:03:07 -05:00
timkjr 2ee2eacb4b fix(omp): clamp OMP_AUTH_BROKER_URL/TOKEN, correct the env-allowlist docs
The docs claimed omp "has no documented vendor-key namespace of its own"
and "the multi-user clamp has nothing to gate" for omp — both false. Per
omp's own docs/environment-variables.md, it reads ~40 provider keys from
env (pi's known 34-key problem in the same shape), and its own knobs are
mostly PI_* (already globally allowlisted): PI_CONFIG_DIR,
PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD,
PI_SHELL_PREFIX. The first three also move the ~/.omp tree
omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading
pinning/history — a known gap shared with pi, documented but not fixed
here.

The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/
OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same
shape DEEPSEEK_BASE_URL is already dropped for in
clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a
non-granted owner in multi-user mode can't redirect them, and correct the
false claims in CLAUDE.md, docs/omp-integration.md, and the stale
resolveOmpHome() comment. Also documents omp's default
tools.approvalMode: yolo, which was previously unstated.
2026-08-28 13:45:12 -05:00
timkjrandClaude Sonnet 5 853681f970 harden(omp): resume-path test coverage, silent-fallback logging, cwd validation
Follow-up from a full-branch review pass (Opus) of the omp-mode integration:

- Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts),
  exported to make it testable: the exact "resume this OMP row from
  history" pipeline that mangleOmpWorkingDir's earlier bug lived in had
  zero coverage despite being the resolver module's whole reason to exist.
- Log a warning when findLatestOmpSessionId() finds nothing on disk and
  continuation silently degrades to omp's own ambiguous --continue,
  in both call sites (session create and respawn pinning) - previously
  silent, making the degradation invisible to anyone debugging it.
- Require an absolute cwd before trusting a session file's working
  directory in omp-transcript.ts's parser, so a corrupted/malformed
  session file can't point a downstream resume at a relative or empty
  path.
- Document (don't speculatively fix) an unverified symlinked-$HOME edge
  case in mangleOmpWorkingDir(): the review's suggested realpath() fix
  assumes omp itself resolves symlinks before mangling, which is
  unconfirmed - guessing wrong there would trade one silent mismatch
  for a different one.
- Incidental: fixed unrelated pre-existing prettier drift in
  session-routes.ts (antigravity/opencode dynamic import line-wrapping)
  that was blocking the pre-commit formatting gate on this file.

Confirmed as a non-issue: the model-name regex allowing "/" is
intentional (provider/model ids like "crof/glm-5.2" were used
successfully in live testing).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-28 11:32:30 -05:00
timkjr 54a930c80e feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads
them back independently from ~/.claude/projects, not from its own
session bookkeeping. omp conversations had no equivalent: kill the
Codeman session and the conversation vanished from Past Sessions
entirely, even though omp itself never forgot it on disk.

Adds omp-transcript.ts, a scanner over omp's own
~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape
as Claude Code's own transcript scanner, but simpler -- these files are
small enough to read whole instead of doing head/tail windows). Each
file's own "session" header line carries the real cwd and session id
directly, so unlike Claude's mangled-directory-name decoding this
never has to guess. Wired into gatherUnifiedInputs() as a second
history source alongside the Claude scan, and HistoryInput/
mergeUnifiedSessions() now carry an optional `mode` so a non-claude
history-only row still gets a real mode badge.

Also fixes the ambiguity behind the "continue picks the wrong
conversation" report from this session's testing: omp mints its OWN
session uuid, unrelated to Codeman's, so a live/persisted row and its
own history-scan row would otherwise show up as two separate entries
for the same conversation the moment the id gets resolved. Reuses the
existing claudeSessionId alias field (mergeUnifiedSessions' fold-into-
owner mechanism) to point at the resolved omp id, threading it through
every place `_claudeSessionId` gets (re)computed -- the constructor,
_resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that
opportunistically resolves it the first time a brand-new omp session
(one that has never gone through a respawn) goes idle.

Also closes a THIRD instance of the "ompConfig never got wired in
here" gap this session kept finding: restoreMuxSessions() in server.ts
restores every sibling CLI's config from persisted state on boot except
omp's, so a boot-recovered omp session always lost its resolved resume
id and fell back to guessing again.

Verified live end-to-end: told a session a secret, killed it fully
(Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux
pane both gone), and the conversation still showed up in the unified
list as a history-sourced row with the real first prompt as its title
and an omp mode badge, keyed by omp's own session id.

Known remaining gap, not fixed here: the claudeSessionId alias doesn't
yet resolve reliably on every boot-recovery path for a session that
was never respawned while alive (e.g. a plain re-attach to a pane that
was never dead) -- worth a follow-up, but doesn't affect the two things
that matter most: the conversation surviving a kill, and continuation
correctness once an id has been resolved (which happens on the very
next respawn either way).
2026-08-28 11:32:30 -05:00
timkjr 4c332c6141 fix(omp): retire the old row on resume, and let DELETE remove persisted-only sessions
Every non-claude "Resume" click creates a brand-new Codeman session
(there is no id to reattach to), but the old row was never cleaned up
-- click resume on the same conversation a few times and the session
list fills up with duplicate rows sharing one name. resumeHistorySession
now retires the row it resumed from after the new one starts.

That retirement needs DELETE to actually work on a row that was never
live in the first place (the normal case for anything showing up in
"Resume Conversation"): findSessionOrFail only checks the in-memory
live-session map, so DELETE 404s on a persisted-only entry today. Give
the route a fallback: when the id isn't live, look it up in persisted
state instead and demote/remove it there (respecting the existing
pinned-session protection). Verified live against a real persisted-only
row via the API, and added route-test coverage for both the success
and still-truly-unknown-id cases (which needed a demoteOrRemoveSession
mock the route harness didn't have).

Also includes an unrelated pre-existing prettier drift fix picked up
by npm run format (omp-cli-resolver.ts, antigravity/opencode import
wrapping in session-routes.ts).
2026-08-28 11:32:30 -05:00
Devvyn b85f7659b7 feat(docker): add Compose deployment support 2026-08-27 19:38:38 +08:00
timkjr 4f5678fac4 feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) 2026-08-26 20:05:48 -05:00
Jack Stuart b00ab3ceea feat(web): show Codex plan usage in header 2026-08-26 18:43:52 +08:00
Codeman maintainer 93a1042bb3 Merge remote-tracking branch 'origin/feat/deepseek-harness' into feat/deepseek-agent-workers
# Conflicts:
#	CLAUDE.md
2026-08-25 19:02:34 +02:00
Codeman maintainer a628737d1f fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first:

- Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys.
  _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into
  every dsh pane and applyEnvOverrides() lands after it, so a non-granted
  owner who could redirect the base URL would have the operator's key sent
  as a bearer credential to a host of their choosing.
- Wait registry: until=stop/blocked is refused on docker and remote-SSH
  dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions).
  The HERDR triple is set via LOCAL tmux setenv, which crosses neither
  docker exec nor ssh, so such a session can never post a hook event and
  the wait burned its whole timeout on every turn.
- Approvals: a dsh item is an ALERT, not an answerable card. The answer
  route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the
  option parser cannot read a third-party TUI's frames, so an answer was a
  blind keystroke into a foreign composer), and the push notification
  carries no Approve/Deny actions for dsh sessions.
- Status shim (v3): --seq is forwarded and the server drops stale retried
  reports inside a 60s window (the TUI retries with backoff, so a retried
  'working' could land after 'blocked' and resolve an approval whose
  dialog was still on screen); 4xx responses exit 0 instead of retrying,
  so one misconfigured session cannot feed the auth rate-limit bucket
  until the hook endpoint 429s for the whole instance.
- Web-UI server: concurrent starts are serialized through a lock (two
  racing POSTs used to pick the same port and orphan the winner), and the
  readiness poll / timeout paths only clear or stop the singleton while it
  is still theirs. First click actually opens the tab now
  (refreshWebviews, not the nonexistent loadWebviews). DELETE
  /api/deepseek/web requires the privileged grant in multi-user mode.
- Cron: deepseek jobs run the same two-part launch gate as the HTTP
  create paths (impl moved into the resolver so all three share it) and no
  longer stamp a Claude default model on the session.
- Parity sweeps: quick-start's docker branch rejects deepSeekConfig like
  the remote branch; the Ralph auto-enable list gained deepseek;
  HookEventType gained agent_working; the phone overview run menu filters
  managed webview records like the desktop menu.
- install.sh: the dsh identity probe closes stdin (under curl|bash a
  child that reads stdin eats the rest of the script), bounds the exec
  with timeout where available, and is memoized to one scan per install.
- Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand
  identity (it rendered as an unstyled UA-grey button); stale markup
  comment about the web shortcut rewritten; clamp docs updated.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 19:01:29 +02:00
Codeman maintainer 015b865f56 fix(deepseek): review fixes for the transcript reader — docker/remote gate, poll memo, honest pairing docs
Four review findings on the worker-transcript feature:

- Docker and remote-SSH dsh sessions now keep the pane segmenter: their
  transcripts live in the container's / remote host's own ~/.dsh, which
  the local reader can never see, so the transcript path returned
  'nothing said yet' forever and an agent polling such a worker starved
  on an answer that existed. Gated on !session.docker && !session.remote
  (statically pinned) and documented in the integration guide.

- last-response reads are memoized on (path, mtime, size, blocks): the
  skill's last_text polls once per second, and each poll decompressed and
  reparsed the whole file on the event loop even when nothing had been
  appended. An unchanged poll now costs one stat.

- The pairing ladder's comment claimed /new is served by step 2; in truth
  the boot-window transcript wins for as long as it exists (deliberately:
  preferring newest-eligible would hand a worker its busier sibling's
  reply). The comment now states the real tradeoff instead of the
  aspirational one. Same for decodeZstdFrames' 'skipped' wording — a
  corrupt frame truncates the decode there, which is the safe behavior.

- stripReasoningPrefix no longer runs on user prompt text, so a prompt
  containing a literal </think> renders whole in blocks view.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-25 18:30:41 +02:00
Codeman maintainer d1bc0c517d feat(deepseek): read dsh session transcripts for last-response
`GET /api/sessions/:id/last-response` is how an agent (and the Response
Viewer) reads what a worker said. DeepSeek was falling through to the
pane segmenter with the other external CLIs, which for this mode is not
merely coarse but wrong: dsh-TUI paints a full-screen splash, so a
`last-response` call on a fresh dsh session answered with its ASCII-art
logo -- and anything polling for a worker's first reply reads that as a
reply.

dsh does not belong in that group. It writes a structured JSONL
transcript per session, so read it. Four things in that file shaped the
reader, all measured against real transcripts on disk:

1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder
   (one-shot and streaming alike) stops at the first frame end: a real
   56-line transcript decoded as 1 line / 158 bytes -- the session header
   alone, i.e. a silent truncation that reads as "nothing said yet"
   forever. `zstdFrameRanges()` walks frame and block headers to find
   exact boundaries; splitting on the 4-byte magic would corrupt
   everything after a magic sequence occurring inside compressed data.
   zstd is resolved at RUNTIME because it landed in Node 22.15 while the
   project floor is 22.0, so an older Node keeps the pane behaviour.
2. Every turn also records a plugin-sourced `user/message` (the runtime
   context snapshot), which must not render as the user's own words.
3. A turn that ends in an error carries the provider's message; it is
   surfaced as `Turn error: …` (and a non-error early stop as
   `Turn ended: …`) rather than as an empty string, which an agent reads
   as "still thinking" through fifteen polls.
4. Reply text is assembled per (turn, step): a finalized message wins and
   the streamed deltas fill in only for a step that never finalized, so a
   partial answer is readable mid-turn and never doubled. "Finalized" is
   tracked as a set of steps rather than as non-empty text, because a
   step whose whole reply was reasoning strips to '' at the `</think>`
   boundary and would otherwise resurrect the raw deltas in its place.

Session-to-transcript pairing is by the transcript's own header `cwd`
plus a boot window against the session's createdAt, never by
reproducing dsh's directory mangling (already two forms on disk) and
never by newest-mtime alone -- mtime alone handed a freshly spawned
worker its predecessor's answer in the same case directory.

An empty result still wins over the pane; only a Node that cannot decode
zstd falls back to it.
2026-08-25 04:14:26 +02:00
Codeman maintainer c30dfaf0e7 fix(deepseek): run the web UI server in the background, not in a shell tab
Clicking "DeepSeek web UI..." opened two tabs: the web tab asked for, and a
shell tab running the server next to it. The shell was deliberate - the server
lived in an ordinary session so it was visible, scrollable, killable and died
with its tab, and nothing new had to supervise a long-lived HTTP server. That
reasoning was sound and the result was still wrong in use: opening a dashboard
should open one tab, and after the first launch the terminal is pure noise.

The server moves to a background child process owned by a new
`src/deepseek-web-server.ts`, behind `POST /api/deepseek/web`. What the session
gave away for free is now explicit, which is most of the module:

- Exactly one server. A second click reuses the running one instead of racing
  it for a port; the session flow could not do this at all, because two clicks
  were simply two sessions.
- Restarted when the requested authority changes. `--trusted-host` fences dsh's
  own /api against the browser authority, and a Codeman reachable at both
  loopback and a tailnet name has two. Reusing a server fenced for the other
  origin renders a page whose every call 403s, which reads as a broken
  dashboard rather than a misconfigured one, so a mismatch restarts instead.
- Killed on shutdown. The child is detached so its whole plugin tree can be
  signalled at once, which also means it would outlive Codeman and hold its
  port against the next start - the exact EADDRINUSE this feature already got
  wrong once.
- Boot output captured and returned. With no shell tab there is nowhere else
  for a stack trace to land, so a failed spawn reports its own tail.

The endpoint is fenced at the same bar as the profile installer and for the
same reason: booting a dsh profile executes the plugin code in it, so this is a
privileged action even though it reads as "open a page". `authority` comes from
the client (`location.host`) because only the browser knows which origin is in
play, and it is regex-confined at the schema boundary - defence in depth behind
the argv-array spawn, admitting host:port in the shapes a browser authority can
take and nothing readable as a second argument.

`GET /api/deepseek/web-port` is gone; port selection moved into the supervisor,
which is the thing that knows whether a server is already running. The two
client-side probe helpers went with it, since the server now owns the wait.

Verified over the tailnet authority end to end: no session is created (session
count unchanged, one tab), the server runs on 3081 beside the user's own dsh
web on 3080, status reports the tailnet authority, and the proxied dashboard
renders with zero 4xx. Full gate green (6148 passed, +6).
2026-08-25 03:08:15 +02:00
Codeman maintainer 15ae5f5d81 fix(deepseek): make the web-UI shortcut pick a free port, verify it, and trust its frame
The `Run > DeepSeek web UI...` shortcut failed three ways at once against a real
install, and the three are independent.

1. It hardcoded `--port 3080`. That is dsh web's OWN default, which makes it
   precisely the port a DeepSeek user is most likely to be serving on already,
   so the launch died with EADDRINUSE against the user's own server. The port
   now comes from `GET /api/deepseek/web-port`, which walks 3080..3119 for a
   free loopback port by BINDING it (a connect probe cannot tell "free" from
   "listening but not answering yet").

2. It opened the tab unconditionally. The crashed server left a saved dashboard
   pointing at nothing, with the failure only visible in a shell tab nobody had
   a reason to look at. The launch now polls the existing webview probe until
   the URL answers, and on timeout reports the error naming the shell tab
   instead of persisting a dead dashboard.

3. The saved tab was untrusted, so the frame was sandboxed without
   `allow-same-origin` and the dashboard was broken twice over: the dsh
   client-runtime reads `localStorage` while loading its plugins and died there
   ("the document is sandboxed and lacks the 'allow-same-origin' flag"), and an
   opaque-origin frame sends `Origin: null`, so dsh's own trust fence 403'd
   every `/api` call no matter which authority `--trusted-host` named. Passing
   `location.host` only means anything once the frame actually carries that
   origin, so `--trusted-host` had never once done its job. The managed tab is
   now created `trusted: true`.

   That trade is real and deliberate: a trusted proxied frame is same-origin
   with Codeman and can reach Codeman's API. It is defensible only because this
   dashboard is an agent harness Codeman just started itself, on loopback, which
   can already run code as the user. It is not a precedent for trusting
   third-party dashboards, which is why it is set at this one call site rather
   than defaulted.

Separately, the shortcut listed its own dashboard twice: once as the menu entry
that starts it and once as the row that entry had written on the previous click.
Webviews now carry an optional `managed` marker, managed rows are filtered out
of the saved-dashboard list, and a relaunch repoints the existing row rather
than stacking one dead dashboard per restart (which the per-launch port would
otherwise guarantee). `managed` is declared in the schema because a plain
`z.object` strips undeclared keys, so an undeclared marker would never survive
the round trip.

`DEEPSEEK_WEB_PORT` is gone from constants.js; its doc comment asserted that a
hand-started `dsh web` and the shortcut "land on the same place and share one
saved tab", which is the bug stated as a feature.

Verified on a real install with the user's own `dsh web` holding 3080: the
shortcut takes 3081, the server answers, exactly one DeepSeek entry shows in the
run menu, and the proxied dashboard renders its workspaces and completes its own
API calls (the previously-403'd `api/settings.describe` now succeeds). Full gate
green (6142 passed), typecheck/lint/format/public-assets clean.
2026-08-25 02:39:57 +02:00
Codeman maintainer 2034719d61 fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed.

1. The multi-user clamp was bypassable by a sibling field on the same request.
   clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but
   DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER
   _configureDeepSeek(), so a non-granted owner sending
   envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated
   instance: a session created with permissionMode "read-only" and that override
   ran with DSH_PERMISSION_MODE=danger-full-access in its pane.

   Every other CLI's bypass is a command-line flag reachable only through the
   per-CLI config, which is why the config clamp alone is the whole gate for
   them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted
   owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to
   what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on
   that list because it aims the launcher at a profile tree whose plugin code
   runs at boot, before any approval row can apply. Verified end to end in real
   multi-user mode: a non-granted user sending both now gets workspace-write and
   no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched.

2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout`
   signals only the direct child, and a plugin install fans out into
   package-manager children that keep the inherited stdio pipes open, so `close`
   never fires and the held-open request leaks with no route-level deadline.
   Reproduced: with a 1.5s built-in timeout the promise was still unsettled after
   6s and both fan-out children were alive. Now detached: true plus negative-pid
   SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a
   last-resort reap for a grandchild that escaped the group. Same probe after the
   change: close fires, direct child and both grandchildren dead.

3. hooksAvailableForMode() promised more than a dsh session can deliver.
   deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that
   triple is the only reason a dsh session posts hook events, so `until=stop` was
   accepted and then blocked for the caller's whole timeout: the exact
   infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now
   takes HookCapabilityOptions and every call site passes sessionHookOptions(),
   with the deepseek arm reading `!== false` so a forgotten one degrades to the
   old behaviour. The refusal names the setting rather than saying "no Claude
   Code hooks", which would send the caller hunting a bug that is really a
   setting they chose. Profile conformance stays unknowable at request time and
   is documented as such. The stale "True for `claude` and nothing else" docblock
   is corrected.

4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a
   claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent
   capture read Claude's own transcript, and adding deepseek silently widened
   both to a mode that has none. They compare mode === 'claude' directly now, and
   a static check pins them there.

Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the
wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 --
bridge off plus explicit until=stop is a 400 naming the setting, bridge off with
no `until` still 200s on idle/exit, bridge on accepts stop.
2026-08-24 16:01:02 +02:00
Codeman maintainer 4cda150493 feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/
antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI
as a Codeman web tab.

DeepSeek is wired unlike its siblings in three ways, each of which is the
reason for a design decision rather than an accident:

1. The agent is a PROFILE, not the binary. `dsh` is a launcher over
   $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and
   `base` -- the interactive terminal front door is always a third-party
   plugin. So availability is two questions: `isDeepSeekAvailable()` (binary)
   and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run
   button gates on the latter, because reporting only the binary would spawn a
   pane that dies on arrival. When the binary is present but no profile is,
   the run menu offers to install one (POST /api/deepseek/install-profile).

2. The permission switch is an env var, not a flag. The harness has no
   command-line permission option; its sandbox/approval rows read
   DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access).
   Exported via `tmux setenv`, never on the spawn line. Absent = the harness's
   own workspace-write, which still asks, so the multi-user clamp is the
   only-if-sent branch and clamps to workspace-write, never read-only.

3. It is the only non-claude mode that passes hooksAvailableForMode(), and it
   earned that. The terminal front door reports idle/working/blocked to a
   supervising process over a generic env-gated contract; a generated shim
   (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each
   report to /api/hook-event as stop / agent_working / permission_prompt. So a
   dsh session gets definitive respawn triggers, real wait-endpoint signals and
   real Approvals Inbox items instead of output-stabilization guesswork.
   `agent_working` is new (157th SSE constant) and joins
   APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its
   alert at once.

The resolver needs the strictest identity probe of the family: `dsh` is not
merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell),
so `dsh --help` must print the harness's own banner before a candidate is
handed a spawn line.

Model is deliberately not a session field -- it is a composition entry in the
profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider
keys named by a settings-file `apiKeyEnv` stay out, which is pi's
34-provider-key problem in a new shape.

Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the
status endpoint's two-part answer, the no-profile refusal, the profile
bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the
permission mode injected via setenv, and the full status bridge -- a
send-and-wait returned signal "stop" from a real turn, and blocked/working
created and cleared an Approvals Inbox item.

Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md
(decisions + honest gaps). Tests: test/deepseek-mode.test.ts,
test/deepseek-cli-resolver.test.ts.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 03:37:56 +02:00
Codeman maintainer c173ae0264 Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts:
#	src/web/public/app.js
2026-08-24 00:32:25 +02:00
Aamer Akhter 74194e4fc0 feat(tabs): COD-359 add owner-scoped tab layouts 2026-08-23 14:46:10 -04:00
Codeman maintainer 3f8c8e99d1 feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.

Grok mixes two existing shapes and the wiring follows from that:

- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
  (--always-approve, grok's bypassPermissions mode; config-level deny rules
  still apply on top). The Run button sends it true, like runAntigravity(),
  and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
  a bare grok spawn is grok's own ask-mode default, which is already safe,
  so only a sent config needs the flag forced off. Cron needs nothing for
  the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
  with mouse support, so it stays OUT of isAltScreenStripMode() and lands
  on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
  (unmeasured against an authenticated composer; documented fallback is the
  'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
  also installs a grok bin), so grok-cli-resolver.ts version-probes every
  candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
  GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
  shared with the dependency registry so doctor and run mode cannot drift.

Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.

Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.

Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).

Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-23 08:39:03 +02:00
Ark0N 12a996b107 Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay
2026-08-21 20:23:17 +02:00
d fei dab432b3fd fix(terminal): bound shell history replay 2026-08-21 08:23:31 -04:00
Codeman maintainer 61251c0b94 fix(cli-resolvers): negative-result caching, SIGKILL on probes, restored VITEST hermeticity, wired not-found diagnostics
Post-merge follow-ups for PR #329 (shared CLI executable resolution):

- Negative-cache resolution misses with a doubling backoff (1min -> 5min
  cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the
  shared resolver cached success only, so a missing CLI re-ran the whole
  chain - ending in a synchronous interactive login-shell spawn bounded by
  the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run
  attempt, stalling the event loop each time, forever. Success still caches
  for the process lifetime, so an installed CLI is picked up within minutes
  without a restart. Tests drive the backoff via an injectable clock
  (createCliExecutableResolver `now` option, threaded through the
  createPiResolverForTest / createAntigravityResolverForTest wrappers).

- Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the
  pi/claude --version probes: execFileSync's timeout only SENDS the kill
  signal and then keeps waiting for the child to exit, and interactive bash
  ignores SIGTERM, so a login shell stuck in a blocking .bash_profile
  survived the timeout and blocked the server permanently.

- Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one
  test pinned the deletion): under vitest the production resolver host now
  replaces un-injected IO primitives with inert stubs - no real PATH
  scanning, no login-shell spawns - and probePiVersion never executes a
  `pi` candidate again (`pi` is a generic binary name, so route tests
  hitting /api/pi/status executed whatever binary the machine carried).
  Tests opt in through the runCommand/isExecutableFile injection hooks or
  allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning
  test is replaced by behavioral pins, including a real-executable fixture
  in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate
  is ever removed again.

- Wire the six get*NotFoundMessage() exports (previously dead) into their
  intended call sites: the createSession throws in tmux-manager and the
  availability gates on POST /api/sessions and POST /api/quick-start in
  session-routes, replacing a third hardcoded copy of the text. A not-found
  error now names where resolution looked (server PATH, login shell,
  checked directories). npm run knip no longer reports any unused export
  from the resolver modules.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:37:58 +02:00
Codeman maintainer d7ad73bc9b fix(response-viewer): role on full-context blocks, divider ReDoS, pi mode (#326 follow-up)
Three post-merge fixes for the external-CLI response viewer:

- ?context=full blocks now carry role ('user' for prompts, 'assistant'
  for response/status/tool). The frontend's loadFullContext() renders
  via msg.role, so the roleless blocks lost the "You" badge and every
  turn rendered as the agent. kind/label/text are unchanged and the
  frontend needs no change.

- normalizeDividerStatusLine() dropped its backtracking regex
  (/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
  a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
  minutes at 10,000), and pane text is agent-controlled with buffers up
  to 32MB. Replaced by a linear counter walk with the identical accept
  set and captured content, pinned char-for-char against the old regex
  by a brute-force corpus test plus a hostile-input regression test
  that fails by timeout with the RegExp version (same approach as the
  glob-matcher hardening in 68ae9a8).

- 'pi' joins EXTERNAL_CLI_MODES: pi sessions had the identical
  empty-viewer symptom the transcript branch exists to fix. The list
  stays a local duplicate of isExternalCliMode() (importing session.ts
  would drag node-pty into the pure module); a new exhaustive parity
  test asserts the two mode sets can no longer drift.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:25:43 +02:00
Ark0N 2073a1b185 Merge pull request #328 from aakhter/repo-status-panel
Report repository status for git-clone installs (GET /api/system/repo-status)
2026-08-21 02:10:38 +02:00
Aamer Akhter 02e7d3fcba feat(system): report repository status for git-clone installs
`GET /api/system/update/check` answers "is there a newer published release
tag?", which is the right question for an npm install but not for a git clone
that tracks a branch. Such an install can be many commits behind its own remote
while the latest tag says it is current, and nothing surfaces that.

Adds `GET /api/system/repo-status`: an informational companion that reports what
this CHECKOUT looks like against its own remotes — current branch and commit,
ahead/behind counts per remote, the remote's role (tracking / upstream / other),
and a bounded list of incoming commits.

Read-only and defensive: every git invocation is `execFileSync` with an argv
array and a timeout, a non-git or remote-less install reports a structured
`error` rather than throwing, and nothing here mutates the working tree or
touches the updater's own state.

Tests: 24 cases in test/repo-status.test.ts.
2026-08-20 12:39:28 -04:00
Aamer Akhter 63c5ba89da fix(response-viewer): populate the viewer for OpenCode/Gemini/Antigravity panes
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then
falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini
and Antigravity render their own TUIs and never write one, so that scan finds
nothing and the response viewer is permanently empty for all three modes.

For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts`
is a pure, dependency-free parser that splits a terminal buffer into prompt /
response / status / tool blocks, keying off the `›` prompt marker, status
dividers and `• Calling|Called` tool-activity lines. The route uses it to answer
with the LAST response, and to carry the parsed blocks under `?context=full`.

Codex keeps its existing branch: it has real rollout files, which are a better
source than scraped pane text.

The response shape is unchanged for every other mode, and Claude panes are
explicitly pinned to the Claude transcript path so a real transcript can never
be shadowed by scraped text.

Tests: 14 parser cases plus a route suite covering all three modes, the
`?context=full` payload, an empty pane, and the Claude regression guard.
2026-08-20 09:24:37 -04:00
Aamer Akhter 5cc78669bd feat(files): search the Files panel by name or path
GET /api/sessions/:id/files gains an optional `q`. With one, the endpoint
answers a FLAT match list instead of a nested tree; without one, the response is
exactly what it was, so every existing caller is untouched.

compileFileQuery() (src/utils/file-query.ts) turns the query string into a
reusable predicate, so the walk prunes as it goes rather than streaming the
whole tree to the client to be filtered there. An empty or whitespace-only
query compiles to null, which is what makes "no query" and "blank query" the
same thing.

The search walk deliberately recurses past directories that do not match — a
file whose ancestors don't match is exactly what people are searching for — so
it carries its own maxMatches cap on top of the existing maxFiles and maxDepth
ones, and reports `truncated` when it stops early. Hidden-file and
excluded-directory rules are the same ones tree mode already applies.

Tests: file-query.test.ts covers the matcher; routes/file-search-mode.test.ts
drives the endpoint against a real temp tree and pins the two properties worth
having — that the walk reaches a match under non-matching parents, and that an
absent or whitespace query leaves the tree response alone. Gating the recursion
on a match turns those red.
2026-08-19 09:17:20 -04:00
Codeman maintainer f60bf93c99 chore: version packages
Red tab alerts track the dialog, not the keyboard: typing no longer clears them, and a dialog answered in the terminal resolves itself on the next listing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:46:48 +02:00