mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-06 23:49:41 +02:00
327e44060724f2c1941ed3d925b1ba70ccfef032
155
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
327e440607 |
fix(codex): fold a codex session into its own rollout row
Review fixes for #386. Duplicate rows. A codex conversation showed twice, once live and once as a past rollout row, because nothing aliased a codex session to its thread id. That is worse than cosmetic: the stale row still resumes, so clicking it starts a second `codex resume` on a thread already open in another pane. - A RESUMED session knows its thread id up front, so it folds from its own side: add `codexConfig.resumeSessionId` to the `claudeSessionId` chain. Not only in the constructor — `start()` recomputes that id at two further points (the mux branch, and the unconditional "third reset point" whose own comment already warned that omitting omp's fallback there stomps the mux branch's resolved alias). Both listed Claude's and omp's ids only, so for codex every mux reattach and boot recovery reset the alias back to the Codeman id and the duplicate returned. - A FRESH session has no thread id until codex writes the rollout, so it is folded from the other side. The scanner now reports `session_meta.originator`, which is `codeman_<sessionId>` for every pane Codeman spawns, and `gatherUnifiedInputs()` stamps the matching live and persisted rows, newest rollout winning (`/new` inside the TUI leaves several rollouts sharing one originator). - Persisted rows read `codexConfig.resumeSessionId` too. A resumed session demoted to a persisted-only record would otherwise lose its alias, and the originator fallback cannot rescue that one: a resumed rollout keeps its ORIGINAL session_meta, so it still names the pane that created the thread rather than the pane that resumed it. Identity cache. It was written as soon as the thread id was known, but codex writes the first user message only when the user submits, so any scan in that window pinned `firstPrompt: undefined` for the life of the process — and the home screen, the command palette and the search-index refresh all scan. `shouldCacheIdentity()` now keeps an identity only once the prompt is known or the head read filled its whole window. Also from review: both caps count emitted rows rather than file index, so a store of sub-agent threads no longer spends the `lastPrompt` budget before the first row that needed it; the cache is an `LRUMap` sized like the one beside it; the unreachable filename fallback is gone; a rollout recording no cwd is dropped rather than emitted with `workingDir: ''`; and the unified-session module header names all three transcript stores. Tests. The resume wiring now has cases for a row with a thread id, a row without one, and a `resumeId` on a non-codex row; the "no continuation is wired" case narrows to gemini/antigravity, which is no longer true of codex. `codex-resume-alias-survives-start.test.ts` drives a real Session through `start()` rather than asserting on pre-stamped inputs — that gap is why the reset points went unnoticed. Plus the maintainer's own cache repro, the tail-budget case, a no-cwd case, and merge cases for both folds. |
||
|
|
8285fff91c |
feat(codex): list codex conversations and resume them
Codex conversations never appeared in the session list, and the resume path skipped codex, so picking one back up meant finding its thread id by hand and POSTing codexConfig.resumeSessionId to /api/sessions. Two gaps caused it: - The unified list is built from ~/.claude/projects plus omp's own store. Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>. - terminal-ui.js sends a continuation only for the CLIs with a "continue most recent" flag. Codex has no such flag — it names a thread by an exact id — and nothing supplied one. Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the thread id `codex resume` takes, and the resume path sends it as codexConfig.resumeSessionId. `resumeId` is what keeps the two kinds of row apart: only a transcript scanner sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never ask codex for a thread that does not exist. Three things measured against a real store of 519 rollouts rather than assumed: - Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max 25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the opening prompt and a bounded tail for the most recent one. session_meta is written once and never rewritten, so per-path identity is cached; a warm rescan of that store costs ~75ms against ~470ms cold. - codex 0.152.1 emits no event_msg/user_message rows at all. It writes event_msg/item_completed carrying an item.type of UserMessage. Both shapes are read, plus response_item as a last resort. - That last resort sees injected context, and the first such row is the repo's AGENTS.md every time, so injections are dropped rather than used as titles. Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them for itself and on a real store they outnumber the resumable threads. |
||
|
|
2ab21c1b32 |
fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring claiming logout called it and no caller at all. The capability is a bearer credential exempt from cookie auth with a rolling TTL refreshed on every use, so a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose referrer policy) stayed valid for as long as anything kept polling it. - POST /api/logout revokes the caller's capabilities (all of them in single-user mode), the admin forced logout revokes the target user's, and user deletion revokes whatever that user had open. revokeOwner returns the count for the admin audit line. - Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own policy is dropped: every URL inside the frame carries the capability, and a dashboard on no-referrer-when-downgrade or unsafe-url handed it to any third-party host it linked. Verified with Playwright that a sandboxed frame under an upstream `unsafe-url` sends no Referer to a third party while the root-absolute fetch and the CSS-triggered 404 fallback still reach the dashboard. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE |
||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
2a32b5064a |
Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases as SIBLING containers through the mounted host socket (Docker-outside-of-Docker). Resolved the README conflict (master had grown to eight CLIs since the branch was cut) and moved the Compose blurb out of the feature bullets into Quick Start, next to the other ways of starting Codeman. Three review findings from the PR discussion are fixed here rather than left for a follow-up, because two of them are shipped-image problems: - `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against the whole context-relative path, so `docker/.env` — which the deployment's own README tells the user to fill with CODEMAN_PASSWORD and provider API keys — was picked up by `COPY . .` and baked into the image at /opt/codeman/docker/.env. Verified in both directions against a real build context: with a canary secret in docker/.env, the unfixed ignore file lets /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for the checked-in .env.example) leaves only the example behind. - `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which still hardcoded ~/codeman-cases, so `codeman skill install --case <name>` reported "Case not found" on exactly the deployment the override exists for. Both now resolve through config/cases-dir.ts. state-store.ts keeps its own literal on purpose: that one migrates the historical ~/claudeman-cases directory by name and is about the old default, not the active location. - CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore joins the documented list of files that genuinely belong in the repo root. The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/ deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's lastClaudeSessionId was handed to every non-claude CLI. Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax, public assets, lockfile. |
||
|
|
65e994d29a |
fix(omp): correct docs/counts/URLs, resolver install-path order, stray comment + CSS
Small cleanup items from upstream review (Ark0N/Codeman#353): - OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real installer target (~/.omp/bin was an earlier unverified guess, confirmed wrong against a real --no-cache Docker build). - docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp -> can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth SessionMode incl. shell -- not eighth), matched the install-path guidance to the resolver fix, updated the version example to the actually-tested 18.0.8, and added a Docker-section caveat: --resume pinning does not currently reach an in-container omp process, since Docker panes never see ompConfig. - docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md already linked to the -omp anchor, so the link was dead) and added an OMP specifics paragraph -- the one external CLI missing an entry in this doc. - .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi, Grok, and DeepSeek Harness) and the backend count. - Removed a stray orphaned comment fragment in the quick-start docker branch and split two CSS lines that had two declarations jammed onto one line. |
||
|
|
f18dccace1 |
fix: don't discard codex/gemini/antigravity conversations on Resume; fix DELETE ownership dup + missing broadcast
resumeHistorySession() creates the resumed row in its own mode via a modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek -> resumeSession) and retires the old row afterward. codex, gemini and antigravity were missing from that map, so resuming one of their rows started a brand-new session with NO continuation while still deleting the row it came from -- silent data loss dressed as the duplicate-row fix. Gate row retirement on continuesSomething (true only for modes that actually got a continuation config) instead of wiring an unverified sessionId->native-conversation-id assumption for the three affected CLIs. DELETE /api/sessions/:id reimplemented the ownership 404 check inline in two places instead of going through findSessionOrFail, and its persisted-only-session branch never broadcast session:deleted, so other open tabs kept the retired row until their next unrelated fetch. Extract the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail() alongside findSessionOrFail() in route-helpers.ts (same ownership contract, returns a SessionState instead of a live Session), and use both from the route instead of inline checks. Add the missing broadcast. |
||
|
|
2ee2eacb4b |
fix(omp): clamp OMP_AUTH_BROKER_URL/TOKEN, correct the env-allowlist docs
The docs claimed omp "has no documented vendor-key namespace of its own" and "the multi-user clamp has nothing to gate" for omp — both false. Per omp's own docs/environment-variables.md, it reads ~40 provider keys from env (pi's known 34-key problem in the same shape), and its own knobs are mostly PI_* (already globally allowlisted): PI_CONFIG_DIR, PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD, PI_SHELL_PREFIX. The first three also move the ~/.omp tree omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading pinning/history — a known gap shared with pi, documented but not fixed here. The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/ OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same shape DEEPSEEK_BASE_URL is already dropped for in clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a non-granted owner in multi-user mode can't redirect them, and correct the false claims in CLAUDE.md, docs/omp-integration.md, and the stale resolveOmpHome() comment. Also documents omp's default tools.approvalMode: yolo, which was previously unstated. |
||
|
|
853681f970 |
harden(omp): resume-path test coverage, silent-fallback logging, cwd validation
Follow-up from a full-branch review pass (Opus) of the omp-mode integration: - Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts), exported to make it testable: the exact "resume this OMP row from history" pipeline that mangleOmpWorkingDir's earlier bug lived in had zero coverage despite being the resolver module's whole reason to exist. - Log a warning when findLatestOmpSessionId() finds nothing on disk and continuation silently degrades to omp's own ambiguous --continue, in both call sites (session create and respawn pinning) - previously silent, making the degradation invisible to anyone debugging it. - Require an absolute cwd before trusting a session file's working directory in omp-transcript.ts's parser, so a corrupted/malformed session file can't point a downstream resume at a relative or empty path. - Document (don't speculatively fix) an unverified symlinked-$HOME edge case in mangleOmpWorkingDir(): the review's suggested realpath() fix assumes omp itself resolves symlinks before mangling, which is unconfirmed - guessing wrong there would trade one silent mismatch for a different one. - Incidental: fixed unrelated pre-existing prettier drift in session-routes.ts (antigravity/opencode dynamic import line-wrapping) that was blocking the pre-commit formatting gate on this file. Confirmed as a non-issue: the model-name regex allowing "/" is intentional (provider/model ids like "crof/glm-5.2" were used successfully in live testing). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
54a930c80e |
feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads them back independently from ~/.claude/projects, not from its own session bookkeeping. omp conversations had no equivalent: kill the Codeman session and the conversation vanished from Past Sessions entirely, even though omp itself never forgot it on disk. Adds omp-transcript.ts, a scanner over omp's own ~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape as Claude Code's own transcript scanner, but simpler -- these files are small enough to read whole instead of doing head/tail windows). Each file's own "session" header line carries the real cwd and session id directly, so unlike Claude's mangled-directory-name decoding this never has to guess. Wired into gatherUnifiedInputs() as a second history source alongside the Claude scan, and HistoryInput/ mergeUnifiedSessions() now carry an optional `mode` so a non-claude history-only row still gets a real mode badge. Also fixes the ambiguity behind the "continue picks the wrong conversation" report from this session's testing: omp mints its OWN session uuid, unrelated to Codeman's, so a live/persisted row and its own history-scan row would otherwise show up as two separate entries for the same conversation the moment the id gets resolved. Reuses the existing claudeSessionId alias field (mergeUnifiedSessions' fold-into- owner mechanism) to point at the resolved omp id, threading it through every place `_claudeSessionId` gets (re)computed -- the constructor, _resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that opportunistically resolves it the first time a brand-new omp session (one that has never gone through a respawn) goes idle. Also closes a THIRD instance of the "ompConfig never got wired in here" gap this session kept finding: restoreMuxSessions() in server.ts restores every sibling CLI's config from persisted state on boot except omp's, so a boot-recovered omp session always lost its resolved resume id and fell back to guessing again. Verified live end-to-end: told a session a secret, killed it fully (Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux pane both gone), and the conversation still showed up in the unified list as a history-sourced row with the real first prompt as its title and an omp mode badge, keyed by omp's own session id. Known remaining gap, not fixed here: the claudeSessionId alias doesn't yet resolve reliably on every boot-recovery path for a session that was never respawned while alive (e.g. a plain re-attach to a pane that was never dead) -- worth a follow-up, but doesn't affect the two things that matter most: the conversation surviving a kill, and continuation correctness once an id has been resolved (which happens on the very next respawn either way). |
||
|
|
4c332c6141 |
fix(omp): retire the old row on resume, and let DELETE remove persisted-only sessions
Every non-claude "Resume" click creates a brand-new Codeman session (there is no id to reattach to), but the old row was never cleaned up -- click resume on the same conversation a few times and the session list fills up with duplicate rows sharing one name. resumeHistorySession now retires the row it resumed from after the new one starts. That retirement needs DELETE to actually work on a row that was never live in the first place (the normal case for anything showing up in "Resume Conversation"): findSessionOrFail only checks the in-memory live-session map, so DELETE 404s on a persisted-only entry today. Give the route a fallback: when the id isn't live, look it up in persisted state instead and demote/remove it there (respecting the existing pinned-session protection). Verified live against a real persisted-only row via the API, and added route-test coverage for both the success and still-truly-unknown-id cases (which needed a demoteOrRemoveSession mock the route harness didn't have). Also includes an unrelated pre-existing prettier drift fix picked up by npm run format (omp-cli-resolver.ts, antigravity/opencode import wrapping in session-routes.ts). |
||
|
|
b85f7659b7 | feat(docker): add Compose deployment support | ||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
93a1042bb3 |
Merge remote-tracking branch 'origin/feat/deepseek-harness' into feat/deepseek-agent-workers
# Conflicts: # CLAUDE.md |
||
|
|
a628737d1f |
fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first: - Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys. _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into every dsh pane and applyEnvOverrides() lands after it, so a non-granted owner who could redirect the base URL would have the operator's key sent as a bearer credential to a host of their choosing. - Wait registry: until=stop/blocked is refused on docker and remote-SSH dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions). The HERDR triple is set via LOCAL tmux setenv, which crosses neither docker exec nor ssh, so such a session can never post a hook event and the wait burned its whole timeout on every turn. - Approvals: a dsh item is an ALERT, not an answerable card. The answer route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the option parser cannot read a third-party TUI's frames, so an answer was a blind keystroke into a foreign composer), and the push notification carries no Approve/Deny actions for dsh sessions. - Status shim (v3): --seq is forwarded and the server drops stale retried reports inside a 60s window (the TUI retries with backoff, so a retried 'working' could land after 'blocked' and resolve an approval whose dialog was still on screen); 4xx responses exit 0 instead of retrying, so one misconfigured session cannot feed the auth rate-limit bucket until the hook endpoint 429s for the whole instance. - Web-UI server: concurrent starts are serialized through a lock (two racing POSTs used to pick the same port and orphan the winner), and the readiness poll / timeout paths only clear or stop the singleton while it is still theirs. First click actually opens the tab now (refreshWebviews, not the nonexistent loadWebviews). DELETE /api/deepseek/web requires the privileged grant in multi-user mode. - Cron: deepseek jobs run the same two-part launch gate as the HTTP create paths (impl moved into the resolver so all three share it) and no longer stamp a Claude default model on the session. - Parity sweeps: quick-start's docker branch rejects deepSeekConfig like the remote branch; the Ralph auto-enable list gained deepseek; HookEventType gained agent_working; the phone overview run menu filters managed webview records like the desktop menu. - install.sh: the dsh identity probe closes stdin (under curl|bash a child that reads stdin eats the rest of the script), bounds the exec with timeout where available, and is memoized to one scan per install. - Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand identity (it rendered as an unstyled UA-grey button); stale markup comment about the web shortcut rewritten; clamp docs updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
015b865f56 |
fix(deepseek): review fixes for the transcript reader — docker/remote gate, poll memo, honest pairing docs
Four review findings on the worker-transcript feature: - Docker and remote-SSH dsh sessions now keep the pane segmenter: their transcripts live in the container's / remote host's own ~/.dsh, which the local reader can never see, so the transcript path returned 'nothing said yet' forever and an agent polling such a worker starved on an answer that existed. Gated on !session.docker && !session.remote (statically pinned) and documented in the integration guide. - last-response reads are memoized on (path, mtime, size, blocks): the skill's last_text polls once per second, and each poll decompressed and reparsed the whole file on the event loop even when nothing had been appended. An unchanged poll now costs one stat. - The pairing ladder's comment claimed /new is served by step 2; in truth the boot-window transcript wins for as long as it exists (deliberately: preferring newest-eligible would hand a worker its busier sibling's reply). The comment now states the real tradeoff instead of the aspirational one. Same for decodeZstdFrames' 'skipped' wording — a corrupt frame truncates the decode there, which is the safe behavior. - stripReasoningPrefix no longer runs on user prompt text, so a prompt containing a literal </think> renders whole in blocks view. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d1bc0c517d |
feat(deepseek): read dsh session transcripts for last-response
`GET /api/sessions/:id/last-response` is how an agent (and the Response Viewer) reads what a worker said. DeepSeek was falling through to the pane segmenter with the other external CLIs, which for this mode is not merely coarse but wrong: dsh-TUI paints a full-screen splash, so a `last-response` call on a fresh dsh session answered with its ASCII-art logo -- and anything polling for a worker's first reply reads that as a reply. dsh does not belong in that group. It writes a structured JSONL transcript per session, so read it. Four things in that file shaped the reader, all measured against real transcripts on disk: 1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder (one-shot and streaming alike) stops at the first frame end: a real 56-line transcript decoded as 1 line / 158 bytes -- the session header alone, i.e. a silent truncation that reads as "nothing said yet" forever. `zstdFrameRanges()` walks frame and block headers to find exact boundaries; splitting on the 4-byte magic would corrupt everything after a magic sequence occurring inside compressed data. zstd is resolved at RUNTIME because it landed in Node 22.15 while the project floor is 22.0, so an older Node keeps the pane behaviour. 2. Every turn also records a plugin-sourced `user/message` (the runtime context snapshot), which must not render as the user's own words. 3. A turn that ends in an error carries the provider's message; it is surfaced as `Turn error: …` (and a non-error early stop as `Turn ended: …`) rather than as an empty string, which an agent reads as "still thinking" through fifteen polls. 4. Reply text is assembled per (turn, step): a finalized message wins and the streamed deltas fill in only for a step that never finalized, so a partial answer is readable mid-turn and never doubled. "Finalized" is tracked as a set of steps rather than as non-empty text, because a step whose whole reply was reasoning strips to '' at the `</think>` boundary and would otherwise resurrect the raw deltas in its place. Session-to-transcript pairing is by the transcript's own header `cwd` plus a boot window against the session's createdAt, never by reproducing dsh's directory mangling (already two forms on disk) and never by newest-mtime alone -- mtime alone handed a freshly spawned worker its predecessor's answer in the same case directory. An empty result still wins over the pane; only a Node that cannot decode zstd falls back to it. |
||
|
|
2034719d61 |
fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed. 1. The multi-user clamp was bypassable by a sibling field on the same request. clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER _configureDeepSeek(), so a non-granted owner sending envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated instance: a session created with permissionMode "read-only" and that override ran with DSH_PERMISSION_MODE=danger-full-access in its pane. Every other CLI's bypass is a command-line flag reachable only through the per-CLI config, which is why the config clamp alone is the whole gate for them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on that list because it aims the launcher at a profile tree whose plugin code runs at boot, before any approval row can apply. Verified end to end in real multi-user mode: a non-granted user sending both now gets workspace-write and no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched. 2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout` signals only the direct child, and a plugin install fans out into package-manager children that keep the inherited stdio pipes open, so `close` never fires and the held-open request leaks with no route-level deadline. Reproduced: with a 1.5s built-in timeout the promise was still unsettled after 6s and both fan-out children were alive. Now detached: true plus negative-pid SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a last-resort reap for a grandchild that escaped the group. Same probe after the change: close fires, direct child and both grandchildren dead. 3. hooksAvailableForMode() promised more than a dsh session can deliver. deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that triple is the only reason a dsh session posts hook events, so `until=stop` was accepted and then blocked for the caller's whole timeout: the exact infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now takes HookCapabilityOptions and every call site passes sessionHookOptions(), with the deepseek arm reading `!== false` so a forgotten one degrades to the old behaviour. The refusal names the setting rather than saying "no Claude Code hooks", which would send the caller hunting a bug that is really a setting they chose. Profile conformance stays unknowable at request time and is documented as such. The stale "True for `claude` and nothing else" docblock is corrected. 4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent capture read Claude's own transcript, and adding deepseek silently widened both to a mode that has none. They compare mode === 'claude' directly now, and a static check pins them there. Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 -- bridge off plus explicit until=stop is a 400 naming the setting, bridge off with no `until` still 200s on idle/exit, bridge on accepts stop. |
||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c173ae0264 |
Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts: # src/web/public/app.js |
||
|
|
74194e4fc0 | feat(tabs): COD-359 add owner-scoped tab layouts | ||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
12a996b107 |
Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay |
||
|
|
dab432b3fd | fix(terminal): bound shell history replay | ||
|
|
61251c0b94 |
fix(cli-resolvers): negative-result caching, SIGKILL on probes, restored VITEST hermeticity, wired not-found diagnostics
Post-merge follow-ups for PR #329 (shared CLI executable resolution): - Negative-cache resolution misses with a doubling backoff (1min -> 5min cap, cliResolveRetryDelayMs, mirroring claudeVersionRetryDelayMs): the shared resolver cached success only, so a missing CLI re-ran the whole chain - ending in a synchronous interactive login-shell spawn bounded by the 5s EXEC_TIMEOUT_MS - on every /api/<cli>/status request and Run attempt, stalling the event loop each time, forever. Success still caches for the process lifetime, so an installed CLI is picked up within minutes without a restart. Tests drive the backoff via an injectable clock (createCliExecutableResolver `now` option, threaded through the createPiResolverForTest / createAntigravityResolverForTest wrappers). - Pass killSignal: 'SIGKILL' on the resolver's login-shell spawn and on the pi/claude --version probes: execFileSync's timeout only SENDS the kill signal and then keeps waiting for the child to exit, and interactive bash ignores SIGTERM, so a login shell stuck in a blocking .bash_profile survived the timeout and blocked the server permanently. - Restore test hermeticity (PR #329 deleted pi's VITEST guards, and one test pinned the deletion): under vitest the production resolver host now replaces un-injected IO primitives with inert stubs - no real PATH scanning, no login-shell spawns - and probePiVersion never executes a `pi` candidate again (`pi` is a generic binary name, so route tests hitting /api/pi/status executed whatever binary the machine carried). Tests opt in through the runCommand/isExecutableFile injection hooks or allowRealIoUnderVitest for real-filesystem fixtures. The deletion-pinning test is replaced by behavioral pins, including a real-executable fixture in the new test/pi-cli-resolver.test.ts that fails loudly if the pi gate is ever removed again. - Wire the six get*NotFoundMessage() exports (previously dead) into their intended call sites: the createSession throws in tmux-manager and the availability gates on POST /api/sessions and POST /api/quick-start in session-routes, replacing a third hardcoded copy of the text. A not-found error now names where resolution looked (server PATH, login shell, checked directories). npm run knip no longer reports any unused export from the resolver modules. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d7ad73bc9b |
fix(response-viewer): role on full-context blocks, divider ReDoS, pi mode (#326 follow-up)
Three post-merge fixes for the external-CLI response viewer:
- ?context=full blocks now carry role ('user' for prompts, 'assistant'
for response/status/tool). The frontend's loadFullContext() renders
via msg.role, so the roleless blocks lost the "You" badge and every
turn rendered as the agent. kind/label/text are unchanged and the
frontend needs no change.
- normalizeDividerStatusLine() dropped its backtracking regex
(/^[─-]+\s*(.+?)\s*[─-]{3,}$/): the lazy middle went catastrophic on
a long dash run without a 3-dash tail (measured 15.5s at 4,000 chars,
minutes at 10,000), and pane text is agent-controlled with buffers up
to 32MB. Replaced by a linear counter walk with the identical accept
set and captured content, pinned char-for-char against the old regex
by a brute-force corpus test plus a hostile-input regression test
that fails by timeout with the RegExp version (same approach as the
glob-matcher hardening in
|
||
|
|
63c5ba89da |
fix(response-viewer): populate the viewer for OpenCode/Gemini/Antigravity panes
`GET /api/sessions/:id/last-response` branches to a Codex-specific reader, then falls through to scanning `~/.claude/projects` for a transcript. OpenCode, Gemini and Antigravity render their own TUIs and never write one, so that scan finds nothing and the response viewer is permanently empty for all three modes. For these CLIs the pane IS the transcript, so segment it. `response-viewer-transcript.ts` is a pure, dependency-free parser that splits a terminal buffer into prompt / response / status / tool blocks, keying off the `›` prompt marker, status dividers and `• Calling|Called` tool-activity lines. The route uses it to answer with the LAST response, and to carry the parsed blocks under `?context=full`. Codex keeps its existing branch: it has real rollout files, which are a better source than scraped pane text. The response shape is unchanged for every other mode, and Claude panes are explicitly pinned to the Claude transcript path so a real transcript can never be shadowed by scraped text. Tests: 14 parser cases plus a route suite covering all three modes, the `?context=full` payload, an empty pane, and the Claude regression guard. |
||
|
|
f44d597450 |
review fixes: every claude create path routes through the workspace-hooks decision
Post-#304 follow-ups. The install-vs-refresh decision (workspaceHooksEnabled, default ON) moved from a session-routes-local helper into hooks-config.ts as applyWorkspaceHooks(workspace, install?), and the claude session-create sites that bypassed it now go through it: cron job fires (cron-service), legacy scheduled-run iterations (runScheduledLoop), and the plan-orchestrator research and planner one-shots. A cron or scheduled run firing in a linked case that never had an interactive session ran hook-blind (no stop for completion detection, no tab alert on a blocking dialog). The shared core also carries the two guards every caller needs: a workspace that no longer exists is skipped (ensureCodemanHooks mkdir -p's, so the boot recovery sweep used to resurrect a deleted repo as an empty tree holding only .claude/settings.local.json), and all errors are swallowed since a create must never fail on hooks. Route handlers keep resolving the setting through their ConfigPort and pass it in; non-route callers omit it and the core reads settings.json itself (absent key or unreadable file = ON). Two adjacent gates tightened in session-routes: - the docker quick-start hooks branch excluded the five external CLIs but let `shell` through, contradicting its own rule that only claude reads .claude hooks; it is now gated on mode === 'claude' - the statusLine exporter call in POST /api/sessions got the same !remote && body.workingDir guard the hooks call got in |
||
|
|
499d35566b |
review fixes: never install workspace hooks for a remote attach or a cwd-fallback create
A claude-mode attachRemoteSession create overwrites workingDir with the user@host:session pseudo-path, which is a RELATIVE path locally — the old refresh-only call no-op'd on it, but ensureCodemanHooks mkdirs, so it created a junk local directory. And with workingDir omitted the cwd fallback reaches the hooks write unvalidated; under installer-created services cwd is $HOME, so hooks materialized in ~/.claude/settings.local.json. Both guarded at the applyWorkspaceHooks call site; regression tests prove the remote attach leaves no junk dir and the no-workingDir create leaves the server cwd untouched. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f485085174 |
feat: workspaceHooksEnabled setting as the opt-out for workspace hook installs
Installing hooks into any workspace a Claude session runs in is the right default, but it takes a decision away from a user who deliberately removed them: nothing on disk distinguishes "removed on purpose" from "never had any", so they would come back on the next session create. Adds the synced workspaceHooksEnabled setting (App Settings -> Agents & CLIs -> Claude), default ON. OFF restores the older behavior exactly: a Codeman hooks block that is already present is still refreshed when stale (COD-91), but one is never added. Every create path routes through one applyWorkspaceHooks() helper so the gate cannot apply to some paths only, and the boot-time recovery sweep honours it too. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
98fa8c00d1 |
fix: install Codeman hooks into every claude workspace, not just cases Codeman created
A session in a linked case (or any pre-existing repo) ran with no hooks block at all: writeHooksConfig only fires when Codeman CREATES the case directory, and refreshStaleCodemanHooks deliberately never adds one. Every hook-driven surface was therefore dead in exactly the place most sessions run: no tab alert or phone-overview NEEDS YOU row when a dialog blocks the pane, no Approvals Inbox item, no push, no definitive stop/idle_prompt for respawn, and no stop/blocked for the agent wait endpoints. Both session-create paths and restoreMuxSessions() now call ensureCodemanHooks(), an add-only merge that keeps a user's own handlers and leaves a malformed settings file untouched. Claude Code re-reads settings.local.json, so a session already running in the workspace starts firing hooks without a restart. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74662dd788 |
fix(skill): stale user-level skill copy shadowed injections; seed the preamble
Two live failures from one root cause: Claude Code loads a same-named
user-level skill (~/.claude/skills/codeman, written once by `codeman skill
install`) over the fresh per-case copy, and nothing ever refreshed it. A
stale Aug-9 copy (pre fast-path, pre lineage header) made every agent-driven
spawn run the old recipes: workers spawned serially with pid polls and
without X-Codeman-Parent-Session, so the web UI drew no lineage arcs.
- refreshUserAgentSkill(): session create now refreshes a marker-owned
user-level copy (refresh-only: absent copies are not installed,
foreign/symlink copies stay untouched).
- seedAgentSessionPreamble(): local claude session create pre-seeds the
skill's preamble into ${XDG_CACHE_HOME:-~/.cache}/codeman-agent-<id>.sh,
single-sourced from the new skills/codeman/preamble.sh, so the skill's §0
bootstrap collapses to a two-line loader instead of a ~150-line paste the
model has to type out (measured ~47s of generation per run).
- SKILL.md: §0 now leads with the loader and keeps the full block as the
stale/missing fallback; explicit verbatim-paste warning (a hand-assembled
preamble is how the header and the fast-path functions got lost);
spawn_worker also sends parentSessionId in the body as defense in depth;
preamble stamp bumped to 1.18.3 so pre-fix cached preambles self-heal.
- test/agent-skill.test.ts pins preamble.sh byte-identical to the SKILL.md
heredoc and covers seeding (XDG + HOME fallback, 0600) and the user-level
refresh (absent/stale/foreign).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9a0e665f72 |
fix(terminal): preserve scroll intent across keyboard resize, surface history truncation
Closes #259, closes #258. Both bottom out in the same gap: nothing tracked whether the user was following live output or reading history. #259 — the keyboard path forced the terminal to the bottom unconditionally (onKeyboardShow/onKeyboardHide passed scrollToBottom:true, applied with no check), so opening the keyboard while scrolled up yanked the user down. The settle cycle now captures intent on its FIRST event, before any fit() has reflowed the buffer, and returns to that anchor when the user was reading. A later capture would read an already-moved viewportY, which is why the capture point matters. The param is renamed restoreScroll to match. Separately, flushPendingWrites gated viewport preservation on _hasRecentUserScrollUp(), a 1500ms decay window, so a user who scrolled up and then actually READ for longer lost protection mid-read. Being scrolled up IS the intent however long ago it was expressed, so it now keys off position. The recency window stays as a race guard on the sticky scroll-to-bottom. The full-history repull already held the user's place and is unchanged. #258 — truncation was reported by a grey line written INTO the terminal ("earlier output truncated"), which scrolls away with the output it describes, cannot be acted on, and said the same thing whether the rest was one click away or gone forever. The server set one `truncated` boolean at two sites meaning opposite things, and the client discarded fullSize and source entirely. The route now reports truncationReason ('tail' = intentional partial replay, the rest is retained; 'capped' = the byte ceiling dropped it) plus retainedBytes, and 'capped' is not downgraded by a later tail cut. The client renders a dismissible banner outside terminal output with three honest states: recoverable (offers Load full history), at-ceiling, and exhausted. The Load button forces past the scroll cooldown but NOT past _replayWouldShrinkBuffer, which still refuses a downgrade for repaint-mode panes. The banner is an overlay, not a flex child: FitAddon derives rows/cols from the terminal parent's computed height, so occupying real layout space would SIGWINCH the CLI on every truncation-state change. Verified in a real browser on the 7 skins: banner text and button clear 4.5:1 contrast on all of them, and terminal height is byte-identical with the banner shown. The first cut used --bg-elevated and --accent-muted, which do not exist, so light skins rendered a hardcoded dark bar under dark text; it now uses only tokens every skin redefines. test/terminal-scroll-intent.test.ts lives outside test/mobile/ deliberately — that suite is excluded from test:ci, so a guard placed there is invisible to CI. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f4dcfbe6ca |
fix(pi): close four mode-list gaps in the pi run mode
Review follow-ups on #282. All four are the same failure shape: a list that enumerates run modes, missed by the sweep that added 'pi'. 1. Cron ignored pi's project-trust clamp. The PR widened CronJobBaseSchema's agentType to accept 'pi' but not the matching clamp beside gemini's, so a non-granted multi-user owner's cron pi job spawned bare `pi` (pi's own defaultProjectTrust, an interactive prompt they can answer "yes" to, which loads and EXECUTES repo-local .pi/extensions TypeScript) while the same user's UI/API launch was forced to --no-approve. The clamp is now a pure exported helper, clampCronExternalCliConfigs(), so both it and gemini's previously untested materialization are pinned. 2. POST /api/sessions/:id/interactive auto-enabled the Ralph tracker for pi: its denylist covered opencode/codex/gemini/antigravity only. The tracker is never fed for an external CLI (_processExpensiveParsers returns early), so a pi session reported ralphEnabled and Ralph UI state no sibling backend shows. 3. REMOTE_CLI_BIN had no pi entry, so buildRemoteCliVersionProbeCommand() returned null and Session.cliVersion stayed blank for every remote-SSH pi session, even though the PR wired the remote launch command and the per-mode override schema field. 4. The desktop home rail's badge map had no pi entry, and its lookup falls back to '', which is what claude renders. A pi session read as Claude there while the tab strip and phone overview badged it correctly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c5b59633d8 |
feat(pi): add Pi (pi.dev) as a sixth CLI run mode (#206)
SessionMode gains 'pi', a first-class backend alongside Claude Code, OpenCode, Codex, Gemini and Antigravity: its own PTY, tmux session, rose tab identity, welcome button, run-mode entry, cron agentType, Docker and remote-SSH command defaults, and clone-repo Brain option. Pi is a different shape of CLI from the other four, and three decisions follow from that: - It has NO permission prompts and no sandbox, so there is no --dangerously-skip-permissions analog and none was invented. The privilege-shaped knob is the tri-state approveProjectTrust, which makes pi load and EXECUTE repo-local .pi/extensions TypeScript and install missing project packages. clampExternalCliBypassForOwner() therefore puts pi in the MATERIALIZE branch: a non-granted multi-user owner gets --no-approve even when no config was sent, because pi's own default is a prompt the session user could answer themselves. That helper had zero test coverage; it now has coverage for all four CLIs. - Only the PI_ prefix joins the env allowlist. Pi's ~34 provider key vars share no prefix and ALLOWED_ENV_PREFIXES is one global list with no mode context, so admitting them would widen the allowlist for every mode at once. Auth goes through pi's /login or the server's own environment. --api-key is deliberately never wired: it would put a provider secret on the spawn command line. - pi stays OUT of isAltScreenStripMode(). Its default TUI renders into the main screen with terminal-owned scrollback, and its 0.84.0 fullscreen mode is runtime-switchable via /settings; that flip was measured to put the pane into the alt screen, which the strip would have corrupted. pi-cli-resolver.ts additionally sanity-probes `pi --version` and requires semver-shaped output, because `pi` is a short generic name a stray binary can shadow; GET /api/pi/status surfaces path and version so a misresolution is diagnosable rather than presenting as a broken mode. Docker installs pi in its own --ignore-scripts step so that flag cannot affect the other four CLIs, and seeds its credentials per-file rather than whole-dir (~/.pi/agent also holds sessions, extensions and package trees). Verified end to end against pi 0.84.1 on an isolated instance: resolver search-dir fallback, flag construction, piConfig persistence across a full server restart, the trust prompt and its --no-approve suppression, the rose Run button on the default daylight-blue skin (the nested skin block eats per-mode gradients unless the rule lives inside it), and the buffer local-echo policy, which pi tolerates where codex did not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f39beb3326 |
chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
250a53125a |
Merge pull request #264 from Ark0N/fix/history-search-260-261
fix(web): usable past-conversation list (#260) and search that finds past sessions (#261) |
||
|
|
aa35c1a0c4 |
feat(sessions): show the git worktree on session rows
Closes #266. Sessions from different worktrees of the same repo were indistinguishable in the Resume list, Cmd+K and search — the row showed a session name and a case label, nothing about which worktree it ran in. Claude Code already stamps "cwd" and "gitBranch" on every user/assistant record, and writes a worktree-state record naming the worktree when the session was started through its own worktree feature. scanProjectDir() already buffers the head of every transcript for prompt extraction, so extractTranscriptGitInfo() parses buffers that are already in memory: no extra file reads, no git subprocess. (Measured on this machine: a git rev-parse per directory costs 482ms for 35 rows; parsing the existing buffers costs nothing.) cwd is taken from the first record that carries it, since a session's cwd does not move. gitBranch is taken from the last, since a branch genuinely changes mid-session. The badge requires a worktree NAME. An earlier revision rendered whenever a branch was known, which put a badge on all 35 rows of a real history -- "master" on every ordinary session, burying the ten rows the badge exists to distinguish. A hand-made `git worktree add` therefore gets no badge rather than a guessed name; Claude's own <repo>/.claude/worktrees/<name> layout is recognised from the path when no worktree-state record is present. worktreeName and gitBranch join the filterAndPaginate haystack so the session manager can search by them. panels-ui re-projects the unified item into a 5-field record before rendering, so the new fields are carried there explicitly -- omitting that silently drops them from Cmd+K only. Also prefers the transcript cwd over decodeProjectKey()'s stat-walked guess, which falls back to $HOME when nothing resolves (#265). Note that path is currently LATENT, not active: on the install this was developed against, every project key whose directory is gone has zero transcripts and so produces no row at all. The transcript value is used because it is authoritative and non-lossy, not because a live bug was reproduced. Verified against a real 35-session history on an isolated CODEMAN_INSTANCE: 10 of 36 rows badged, history row count unchanged at 35 (nothing dropped), no page errors. 129 tests pass across the new suite plus the unified service, unified route and session route suites. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016uTqt8ttmsBLXbm5JFHis3 |
||
|
|
c13b3c55d3 |
style: drop em-dashes from the prose added in this branch
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5d42f64393 |
fix(web): usable past-conversation list, and search that finds past sessions
Two home-screen reports from @jordan8037310, both about history that is present but unreachable. #260 — "Resume Conversation" rendered 4 rows, then a button that appended every remaining row into a `max-height: 240px` box, so 35 conversations landed in a four-row scroll well with no ordering or filtering. Rendering now goes through `_renderHistoryList()` over a cached corpus: 10 rows to start, Show more/Show less that grows and shrinks the box (the height cap is class-driven, `.history-list.expanded`), plus a filter box (name, folder, #case label, prompts), a sort control (recent / name / folder, pinned rows still first) and a shown-of-total count. A filter implies expansion, so every match is visible, and the whole header hides as one unit while a federated search is active. The A-Z sort keys off the same string the row renders, since most rows are transcript-backed and carry no session name at all. #261 — the search box could not match a past project by folder name: `harvestSources()` built its session corpus from the live in-memory map, while past sessions come from `/api/sessions/unified` (lifecycle log + transcript scan). Folding that scan into the request path would have cost the search its no-filesystem-reads property, so the corpus arrives via a bounded snapshot instead: `session-history-index.ts` is published as a side effect of `/api/sessions/unified` (the home screen fetches it on open, which is the same screen the search box lives on) and rebuilt fire-and-forget, single-flight and TTL-guarded when a search finds it stale. A result for a closed session now resumes the conversation rather than selecting a tab that no longer exists, and is badged RESUME. The snapshot is stored unscoped with a per-row owner and re-filtered through canAccessOwned() on read, so multi-user sees exactly what /api/sessions/unified exposes: own sessions only, host-wide transcript history admin-only. Live rows are harvested first and win the dedupe. Verified end-to-end against a real instance with 60 past sessions: cold process answers its first search without history and its second with it; folder-name queries return resume targets; clicking one posts the right resumeSessionId + workingDir. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d33f3803a1 |
fix(agent-skill): write the skill atomically and stop swallowing refusals
Two ways the injection could go wrong quietly.
`installAgentSkillInto()` wrote each file with a bare `writeFile`, no lock and no
temp+rename, while every sibling mutator in hooks-config.ts goes through
`withSettingsLock`. Two Claude sessions created concurrently in one repo both wrote the
same ~16KB SKILL.md, and any reader loading it mid-write could observe a truncated
file. Writes now go through a temp+rename helper under the same lock the neighbours
use, so a reader sees either the old file or the new one.
Both server call sites discarded the outcome with `.catch(() => {})`, so the two
refusal results were invisible: `foreign` (a user-authored skills/codeman is present,
so we declined to touch it) and `symlink` (the skill dir or its parent is a symlink, so
we declined to write through it). Turning `agentSkillEnabled` on, seeing nothing appear
and having no way to find out why was the reportable-as-a-bug outcome. Refusals are now
logged with the path and what to do about it. The boring outcomes stay silent, since
they happen on every session create. Injection remains best-effort: a refusal or a
thrown error still cannot fail session creation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
e88b971bb7 |
feat(skill): add the agent-skill install layer and harden the packaged skill
Ship `skills/codeman` as an installable Claude Code skill rather than a repo-only reference, and fix six defects found while verifying it live. Install layer: - `codeman skill install [--case <name>]` / `codeman skill uninstall`. Case names resolve through linked-cases.json first, mirroring the server's resolveCasePath(), so a case linked in from outside ~/codeman-cases no longer fails with "Case not found". - applyAgentSkill() / installAgentSkillInto() / removeAgentSkillFrom() in hooks-config.ts. Copies are marker-owned, so an unmarked user-authored skill is never touched, and a symlinked skill dir is refused (this repo's own .claude/skills/codeman is a symlink to the source). - Synced `agentSkillEnabled` setting, default OFF: schemas.ts, ports/config-port.ts, server.ts, session-routes.ts (add-only injection on Claude session create and quick-start), plus the App Settings toggle. Skill content fixes, each reproduced before and after: - Fail-closed `delete_session` replaces `is_self ... || curl -X DELETE`. Shell state does not survive between agent tool calls, and an undefined is_self exited 127, firing the `||` branch and deleting the caller's own session with the one guard bypassed. The request now lives inside the guard, so a lost preamble deletes nothing. - clientId is a fixed literal instead of `agent-$$`. The pid changes per tool call, so the documented resend-identical-request loop stopped being a duplicate and retyped the prompt, submitting the turn twice. - `last-response` is now the documented read path for claude and codex workers. It returns clean transcript text; the terminal scrape it replaces returns a wall of TUI repaint noise. Its transcript flush lags the stop signal, so the recipes poll it rather than reading once. - quick-start examples branch on `.success`. Previously a failed spawn yielded the literal session id "null" and burned the whole readiness budget before reporting jq noise instead of the cause. - Documented that turning `agentSkillEnabled` off sweeps nothing, and corrected the hooks-config comment that claimed a toggle-off sweep exists. Per-case cleanup is `codeman skill uninstall --case <name>`. - Documented that SESSION_BUSY means the 50-session cap on quick-start, and that caseName resolves linked cases, so a generic name can land a worker in a real repo. Tests: test/agent-skill.test.ts covers install, refresh, idempotence, marker ownership and symlink refusal against the real packaged source; test/quick-start.test.ts covers injection behind the setting. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d26f26fe34 |
chore: version packages
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7f6d18b398 |
Merge pull request #226 from christianhaberl/fix/input-loss-on-failed-delivery
fix(api,ws): an input whose delivery fails can be retried instead of being lost |
||
|
|
eb8d11ffc3 |
fix(terminal): restore shell scrollback, recover history lost to tmux repaints
Four fixes for the scrollback reports in #205 (plus its follow-up comment). 1. tmux-backed shell/opencode/antigravity sessions were parked in xterm's ALTERNATE buffer for their whole life. The tmux CLIENT emits smcup (\x1b[?1049h) as its first bytes on attach, and the existing strip is gated to claude/codex/gemini, so it reached the browser verbatim. In the alternate buffer baseY is pinned at 0 (no scrollback, so touch scrolling is a no-op) and xterm's own wheel handler translates the wheel into \x1bOA cursor keys, which readline receives as shell history navigation. Both reported symptoms, one sequence. isMuxAltScreenOnlyStripMode() now strips that toggle for those modes, but ONLY under tmux (the direct-PTY fallback still needs a program's own alt screen) and ONLY the alt-screen toggle: 3J from a user's `clear` and the mouse DECSETs a pane's htop/vim rely on are left alone. Safe because tmux never forwards a pane's alt-screen toggles to its client, it repaints; captured from a real attach, vim/less/htop emit zero. 2. "Load more history" on scroll-to-top. xterm's buffer is only ever a window onto tmux's history, and tmux repaints the pane rectangle instead of emitting linefeeds whenever output outpaces its flush, OVERWRITING already-rendered scrollback. Measured: a 60-line burst added 1 row and destroyed 34, while the same 60 lines emitted slowly added all 60. Scrolling up at the top now re-pulls the full tmux scrollback and holds the user's place. Verified end to end: 42 rendered rows -> 213, recovering all 150+60 printed lines. 3. The full-scrollback replay was gated on a single "first load after page load" flag, which whichever session auto-selected consumed, so every other tab started with one visible frame. Now tracked per session. 4. _wheelScrollLines ignored ev.deltaMode, so Firefox (DOM_DELTA_LINE, deltaY 3 per notch) scrolled one line where Chrome scrolls four or five, and capped the forwarded SGR report at one tick. Line and page deltas are now converted, and a pure horizontal swipe no longer falls through to a phantom -1. Analysis and measurements: docs/scrollback-issues-analysis.md |
||
|
|
ebfcac6ad1 |
fix(api,ws): an input whose delivery fails can be retried instead of being lost
Both input paths recorded the (clientId, seq) pair as applied and acknowledged the frame BEFORE knowing whether the write had landed: the POST route because its mux write is fire-and-forget so the response never waits on a tmux child, the WebSocket handler because it ACKed unconditionally. When the write then failed, the client dropped the frame from its durable queue and the server rejected the retry as a duplicate. The reliable-delivery layer was guaranteeing exactly-once delivery of something that had never been delivered — and `Session.write()` returned void, so a session whose PTY was gone swallowed the data with no signal at all. - `forgetInputSeq()` rolls the bookkeeping back on failure, but only when that seq is still the newest one; a later input has superseded it and must not re-open. - The WebSocket handler withholds its ACK when the write did not land, so the client redelivers. - `Session.write()` reports whether it reached a PTY. Response codes are unchanged, deliberately: a session can legitimately have no PTY yet, and turning that into a failure status would be a contract change of its own. What this does NOT do: remove the root cause. The POST still answers 200 before the mux write is attempted, so a client that treats any 2xx as final cannot learn about that failure. What closes is the narrower window — the write failed AND the ACK never reached the client — plus the whole WebSocket path. Closing the rest would mean awaiting the tmux child inside the request. 9 tests. They drive the HTTP route, not only the Session primitives: with the rollback removed from the route, 2 of them fail. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0d0b772619 |
feat: make Antigravity a first-class CLI across docs, installer and UI
Antigravity (agy) was wired into the session layer but never propagated to the surfaces around it, while Gemini CLI stayed documented as a consumer product despite being enterprise-only since Google's cutover. Gemini keeps full support; Antigravity now sits beside it everywhere. Functional fixes: - docker/agent.Dockerfile never installed agy, so a docker case with mode 'antigravity' died on command-not-found. agy is not on npm, so it gets its own installer step. --dir /usr/local/bin is load-bearing: the default $HOME/.local/bin resolves to root's home at build time and is unreachable by the `agent` user the container runs as. Verified inside codeman/agent:base (v1.1.10, reachable as `agent`). Note the binary is ~190MB, the largest layer in the image. - Welcome screen gained a Run Antigravity action, gated on agy being present like the other CLI buttons, with a cyan identity matching the toolbar run button and run-mode dot. - install.sh now detects agy (search paths mirroring the resolver), counts it as a satisfying AI CLI, and recommends it over Gemini in the install hints. Detection only, no new auto-install path. Docs corrected where they were factually wrong: - architecture-invariants documented isExternalCliMode() as opencode/codex/gemini when the code has included antigravity for a while, said "all three modes", and omitted ANTIGRAVITY_ from the env prefix allowlist row. - cron-guide's agentType enum, cron-discovery's SessionMode, and remote-sessions' RemoteCommandMode were all stale. Also: README + README.zh-CN (five CLIs, Gemini marked enterprise-only), package.json keyword, and comment drift in 8 places. test/run-mode-ui.test.ts now covers the new welcome button; verified it fails without the settings-ui wiring. Antigravity nests its whole state under ~/.gemini/antigravity-cli/, not ~/.antigravity, so the existing .gemini docker credential seed already covers it. Recorded as a comment so nobody adds dead config later. isAltScreenStripMode() deliberately still excludes antigravity: whether its TUI needs the alt-screen strip is a behavioural question that needs a real agy session, not a guess. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
ecd3f3f32a |
harden(history): exclude automated transcripts by SDK shape, not by "not cli"
#215 filters non-interactive transcripts out of Past Sessions with `entrypoint !== 'cli'`. That is an allowlist on a value, and the check hides rows, so it fails CLOSED on anything Claude Code has not shipped yet: the day it stamps a new interactive entrypoint (a rename, or a second interactive host), no transcript matches 'cli' any more and the entire Past Sessions list goes blank with nothing in the UI explaining why. Invert it to a blocklist on the SDK shape (`sdk`, `sdk-cli`, `sdk-py`). An automated entrypoint we do not recognize yet now costs a few noisy rows, which is the annoyance the filter set out to fix, rather than a dead feature. Matches the fail-open reasoning #215 already applied to a MISSING entrypoint field; only the unknown-VALUE case was inverted. Test fails against the pre-fix line and passes after. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c19d884a51 |
Merge pull request #215 from timkjr/fix/past-sessions-history-quality
fix(history): three Past Sessions data-quality bugs (automated-session noise, cross-contaminated previews, blank restart-heavy rows) |
||
|
|
b641560040 |
Merge pull request #203 from shenlvkang-collab/contrib/claude-viewer-session-pin
fix(web): pin the Claude response viewer to the pane's own conversation |