mirror of
https://github.com/Ark0N/Codeman.git
synced 2026-10-08 16:39:42 +02:00
492f8d8ddf493b2b9141bb37dd03682667541fd2
181
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
da933d70be |
feat(sessions): offer to rebuild the sessions a host reboot destroyed
A host reboot takes the tmux server down with it, so every pane dies, reconciliation finds nothing to attach to, and the board comes up empty. Picking yesterday's work back up meant finding each conversation in history and resuming it by hand, one at a time. The boot pass now works out what the reboot killed and leaves it on offer. It runs inside restoreMuxSessions(), in the window where reconciliation has reported the dead sessions and cleanupStaleSessions() has not pruned their records yet, which is the only place the records can still be read. The board shows a banner, and nothing is created until the user clicks it. A click rather than an automatic restore is what makes the reboot heuristic acceptable. The heuristic cannot tell a reboot from a crash that took tmux down inside the same window, so it decides whether to ASK, never whether to act: a wrong yes costs a line of text the user dismisses instead of N CLI processes nobody asked for. Four things are re-checked when the click arrives rather than trusted from boot, because hours can pass and the board moves on. The owner's privilege grant re-resolves through the env clamp. The workspace must still be on disk. A conversation the user already resumed by hand from the Resume list is skipped, since two panes running --resume on one conversation would fight over the same transcript. Entries leave the plan synchronously before the first await, and the route is single-flighted, so a double-click or two devices cannot both reach the same entry. A restored session comes back attached, idle and disarmed. Respawn controllers and Ralph loops are deliberately not re-armed: a machine that just came up is the worst moment to turn an autonomous run loose. Its workspace hooks are installed by the restore route itself, because the boot-time sweep sits behind a gate that is false after a reboot and has finished long before the click; without them a session goes silently blind, with no stop or idle events for respawn, no Approvals Inbox item and no red tab on a blocking dialog. Stats collection starts the same way. The pane is new, so the conversation continues and the terminal scrollback does not. The banner says so rather than letting an empty pane read as a broken restore. The plan lives in memory only. A server restart drops it, which costs the convenience this adds and never the conversation: the conversation is the transcript under ~/.claude/projects, which the Welcome screen's Resume list and the Session Manager already read, so a dropped plan returns the user to resuming by hand. clampEnvOverridesForOwner moves to src/session-env-clamp.ts, since the question it answers is about session privilege rather than about HTTP and it now has a caller outside the route layer. Its test hook stays re-exported from session-routes.ts. Claude sessions only for this pass. The other CLIs name their thread in their own config object, which this does not thread through yet. Remote and docker sessions are skipped on purpose, because both need another host or a container to be up and a freshly booted machine cannot promise either. Refs #411 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5b920cb43d |
feat(sessions): land auto-naming opt-in, in the prefix form, from the first user prompt only
Finishes #376. The contributed keystroke tracker sat on the raw byte stream and named tabs wrong five ways (every prompt, every write path, a bare Esc eating the next prompt's first character, pasted newlines as Enter, any CSI clearing the draft) and replaced the whole name, which dropped the case from the tab and reset the w<n> counter. This lands the feature with each of those closed: - First prompt means the first: applyAutoName() flips a placeholder to `auto` whether or not the string changed. nameSource is now the tri-state placeholder | auto | manual; the name setter is the only manual path. - Only user-originated input counts: write()/writeViaMux() take SessionWriteOptions.fromUser, set by the browser WS path and POST /input only, so Ralph, respawn, cron, approvals and the trust-dialog keys can never name a tab. A startMode 'shell' CLI never feeds the tracker (a capability, not an id check); the send-key route feeds trackUserInput() because its line feed bypasses the session. - Prefix form `w3-case: title`: parseSessionPrefix() already renders it as the title with the prefix in the tooltip and the next-session counter still matches it. Composed within MAX_SESSION_NAME_LENGTH. - Tracker rules per key: bare Esc resolves at chunk end; mouse/focus reports, Tab, cursor keys, Shift+Tab are no-ops; Up/Down and Ctrl+P/N/R taint the draft so Enter submits nothing rather than a fragment; bracketed-paste newlines and Ctrl+J / Shift+Enter join with one space; the draft keeps its head past 8192 code points; an escape past 64 bytes is abandoned. - Title: slash commands by shape (a path is a prompt), `!` escapes refused, first sentence only past 8 code points ("e.g." is not a title), 72 code points on a word boundary. - Synced `autoNameSessions` setting, default OFF (the prompt reaches mux-sessions.json, session:updated and /api/search), App Settings -> Appearance -> Tabs, read fresh per prompt after the eligibility check. Tests: test/session-auto-name.test.ts (tracker, title, composition, ownership, emit gating), the wiring test (once, prefix, setting off, manual protected), test/routes/session-name-routes.test.ts (PUT /name flips to manual and persists). Verified live on an isolated instance: API and browser-typed prompts name the tab, a second prompt does not, shells and renamed tabs are untouched, nameSource survives a restart. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
c4322513d9 | Merge pull request #376 from shenlvkang-collab/feat/auto-session-names-upstream | ||
|
|
942bf37e48 |
fix(custom-model): unset injected env on clear, resume on restart, select the model for pi/omp/grok
Custom Model Endpoint Profiles (#393) let a session point its CLI at a custom OpenAI-compatible endpoint by injecting env vars or a config file and restarting the CLI in place. Review of the apply path found four things, two of them destructive. This lands all four plus the smaller items from the same review. 1. Clearing a selection did not clear it. The injected vars reach the CLI via `tmux setenv`, which persists at the tmux-session level and is inherited by `respawn-pane` (measured: `setenv FOO bar` survived two successive `respawn-pane -k`), so deleting the keys from the session's envOverrides relaunched the CLI still pointed at the old endpoint, and for the configDir kinds at a HOME/CODEX_HOME/GROK_HOME that had just been deleted. `Session.setCustomModel()` now reports the removed keys, queues them (`_pendingEnvUnsets`), and `RespawnPaneOptions.unsetEnvKeys` carries them into `applyEnvOverrides()`, which `setenv -u`s them before re-applying the live overrides, on the same path that already unsets the legacy CLAUDE_CODE_EFFORT_LEVEL. Verified on a private tmux socket that `setenv -u HOME` hands the next respawn the global HOME back. 2. Applying a model to a local claude session killed the pane. The relaunch was `claude --session-id <id>` and Claude refuses an id that already has a transcript, and unlike the dead-pane respawn this one kills a working pane first. `restartCli()` now pins the live conversation id as the resume id for that respawn when the CLI's launch declares a `fallback` chain, which renders the same `--resume <id> || --session-id <id>` shape the docker and remote pane commands use. Gated on the registry shape, not the CLI id: an entry whose resume id is minted by the CLI itself never declares that chain. 3. pi, omp and grok wrote their config file and then launched without the `--model` that selects it, so the file was ignored. The registry entry now declares `customModelInjection.launchModel` (`custom/{modelId}` for pi and omp, grok's `[model.codeman-custom]` block name), the builder renders it, and `_withCustomModelLaunchModel()` applies it onto the respawn options through `legacyConfigField`, leaving the stored <Mode>Config untouched so a clear falls back to the user's own model. A model id the CLI's `model` token pattern cannot carry is refused with a 400 rather than silently dropped by the argv engine. 4. Remote (SSH) and Docker sessions reported `restarted: true` and changed nothing: their `restartCli()` reattaches the durable tmux rather than relaunching the agent, and the env lands on the local pane. Both are refused with a 400 until those paths are plumbed. Smaller items from the same review: - The selection survives a Codeman restart as the disk-only `__customModel` bookkeeping (endpoint, model, injected key NAMES, config dir, launch model; never the values, which carry the API key). Recovery re-derives the values from the endpoint store through the same apply path the route uses and keeps the bookkeeping even when the endpoint is gone, so a later clear still has keys to unset. - Discovery goes through `webviewFetch()`, so the RESOLVED address is judged by the same egress guard the web-tab proxy uses, and `baseUrl` reuses `webviewUrlSchema` (http(s) only, no embedded credentials, link-local and cloud-metadata addresses refused). undici's `fetch failed` wrapper is unwrapped so the user sees the ECONNREFUSED underneath. - `custom-model-hosts.json` is written 0600 via tmp+rename, the per-session config dir 0700/0600 (pi and omp embed the key literally), and that dir is removed with the session. - `PR.md` is gone from the repo root and the design doc moved to `docs/custom-model-endpoints-plan.md` with the LAN address and the personal name scrubbed; every reference follows. The guide's `authStyle` text matches the shipped schema (`bearer | api-key`, default `bearer`) and says that `customModelEndpointsEnabled` is read by nothing until the picker lands. - `config/tsconfig.scripts.json` typechecks `scripts/test-local-llm-harnesses.ts` (four real type errors fixed). It is not yet wired into `npm run typecheck` because that line differs on master; adding `&& tsc -p config/tsconfig.scripts.json` there is the one-line follow-up. Tests: `test/session-custom-model-restart.test.ts` drives a real Session and fails on the unfixed code for items 1 to 3; the route suite covers item 4 and the pattern refusal; `test/tmux-manager.test.ts` pins that the unsets run before the overrides and that a shell-metachar key never reaches tmux. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> |
||
|
|
1e42cb4e2d |
Merge pull request #393 from opticon454/feature-custom-llm-server-support
feat: Custom Model Endpoint Profiles (local or cloud, all harnesses) |
||
|
|
b2b2c767ea |
Merge pull request #361 from timkjr/fix/statusline-injection-opt-out
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk |
||
|
|
61779745aa |
test(custom-model): make the harness smoke test dynamic, verify all 9 CLIs end-to-end
Rewrites scripts/test-local-llm-harnesses.mjs -> .ts to read the live CLI
registry (enabledClis()) and call the real production
buildCustomModelInjection()/applyConfigDirInjection() instead of keeping a
second hand-maintained copy of every CLI's env/config shape. A future
registry change (new CLI, edited env var, fixed config template) is now
picked up automatically with zero edits to this script; only the one-shot
invocation flags (info the registry genuinely doesn't model) stay in a
small hand-maintained ONE_SHOT table, and a registry CLI with no entry
there reports UNKNOWN rather than being silently skipped.
Extracted src/custom-model-injection-apply.ts (applyConfigDirInjection/
removeConfigDir) so the production route and this script share one
implementation instead of two.
Full end-to-end run against a real llama-swap server, inside a
codeman/agent:llm-test Docker image with all 9 CLI binaries:
- claude, opencode, pi, grok, omp: PASS, real "hello world" replies
- codex: confirmed FAIL for a real protocol reason, not a bug — it only
speaks the Responses API since Feb 2026, which llama.cpp/llama-swap
don't implement
- gemini: confirmed FAIL, unresolved after real investigation — an
undocumented GATEWAY AuthType gemini-cli selects once
GOOGLE_GEMINI_BASE_URL is set rejects every auth-key format/override
tried
- deepseek: reaches the server (env vars are read) but gets a consistent
HTTP_404; root cause not identified, documented as best-effort/unknown
- antigravity: SKIP, no known mechanism (unchanged)
Two real bugs found and fixed along the way (grok, pi/omp registry
entries in stock.ts): grok's original recipe (env vars) was flat-out
wrong, not just unverified — the real mechanism is a config.toml
[model.<name>] block redirected via GROK_HOME. pi/omp's PI_CONFIG_DIR
does nothing for either (grepped pi's entire bundled source — the string
appears nowhere); the real redirect is the child process's own HOME, and
both need `models` as an array of {id} objects, not an object keyed by
id (silently loaded zero models otherwise).
deployment_plan.md, PR.md, docs/custom-model-endpoints.md, and CLAUDE.md
updated with the final confidence table reflecting all of the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3
|
||
|
|
41416566aa |
feat(custom-model): Custom Model Endpoint Profiles (local or cloud, all harnesses)
Point any Codeman-supported harness (Claude, opencode, Codex, Gemini, Pi, Grok, DeepSeek, OMP) at a custom OpenAI-compatible endpoint instead of its native cloud backend, for a given session. Covers local hardware (llama.cpp, Ollama, vLLM, DGX Spark, Strix Halo) and cloud (Azure AI Foundry, OpenRouter). Off by default (customModelEndpointsEnabled, synced, default OFF). - Registry: capabilities.customModelInjection per CLI entry (env / configContentEnv / configDir / unsupported kinds) - Pure injection builder (custom-model-injection.ts) turning an endpoint + model id into the real env vars / config content per CLI - Endpoint store + CRUD routes (custom-model-hosts.ts, custom-model-routes.ts), discovery via GET /v1/models, SSRF-guarded - Session integration: Session.setCustomModel()/restartCli() (POST /api/sessions/:id/custom-model), reusing the existing respawn-pane -k primitive to restart the CLI process with new env - Multi-user hardening: every new redirect-capable env var added to its CLI's privilegedEnvKeys, closing a pre-existing gap where several were already reachable via the generic envOverrides field's prefix allowlist - Standalone scripts/test-local-llm-harnesses.mjs: spawns real CLI binaries against a real endpoint outside the web UI, independent of tmux/sessions - Mock-server contract tests (test/fixtures/mock-openai-server.ts) replaying every CLI's injected values through a real HTTP shape Real end-to-end validation against a live llama-swap server (inside a codeman/agent:llm-test Docker image with all 9 CLI binaries) found and fixed three real bugs before they shipped: - Codex's config.toml schema was wrong ([model].default table instead of a top-level model string + [model_providers.custom]); fixing it then surfaced a genuine, documented protocol incompatibility (Codex only speaks the Responses API since Feb 2026, which llama.cpp/llama-swap don't implement) - Claude Code's async session-title-generation call validates ANTHROPIC_DEFAULT_HAIKU_MODEL against its own internal model list and hangs the whole -p invocation on an unrecognized name; documented for chunk 6, worked around in the standalone script only (--bare is NOT safe for a real interactive session, which needs hooks) - The discovery route's authStyle: 'both' option (send both Authorization and api-key headers) reliably hung a real server; removed the option entirely rather than just changing the default Status: draft. Chunk 6 (frontend toolbar/settings UI) not yet built — see PR.md and deployment_plan.md for the full chunk breakdown and confidence table. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017HqNWfmtBU2KN29SvSVWB3 |
||
|
|
bd61735393 |
fix(web): show the whole last turn in the Claude response viewer's brief view
The eye button rendered `data.text`, which is one row: the last assistant
row of the transcript. A Claude answer is a median of 3 model messages
(p90 11) split around tool calls, so the brief view usually showed the tail
of an answer ("Done.", "Let me look.") and the substance appeared only after
More. The full view was fine, which is why the brief one read as broken by
comparison.
The brief view now asks `?context=turn`. The reader answers with the
assistant messages of the last ANSWERED turn (`selectLastAnsweredTurn`: the
highest `turn` that has an assistant row, so a prompt queued after the
answer does not blank the view) and the frontend renders them exactly as
the full view renders that turn: one badge, then continuation segments,
gated on the numeric `turn` as before.
`data.text` is unchanged in every context — still the last assistant row,
never `messages.at(-1)` — because agent pollers hash it. Readers that emit
no turns (Codex, the pane parser, DeepSeek, an older server) return `text`
only for `context=turn`, and the brief view keeps its single card for them.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01McLWqCWBuQYGuPMScb4Aou
|
||
|
|
2b57c595df |
fix(terminal): gate the row-preserving skips on a capture, not the query flag
Review of the previous commit found the guard inverted: the three skips keyed on `?full=1`, which is only what the client asked for. When the capture comes back null — ENOBUFS, a timeout, a vanished pane, or a session with no mux at all — the reply falls back to the byte history, which IS a stream of successive frames and still needs stripping. Gating on the request returned it whole: measured at 82KB against 4KB for the same buffer without `full=1`. A direct-PTY session takes that path on every first selection, not only during an outage. The skips now key on `isFullCapture`, meaning a capture arrived. Three further defects the same review surfaced, all on this path: Keeping the trailing rows is only sound when a cursor move follows to count back up from them. On the two branches where the cursor query fails there is no move, so the caret was left at the bottom of the pane — worse than before. The cursor is now read first and settles both decisions together. The move is relative rather than absolute. `CUP` numbers rows from the top of the browser's screen, so it is only right while the browser's row count equals the pane's, and `resizeWindow` does not wait for tmux, so a capture can be taken before a requested resize applies. Measured against real tmux with a browser four rows shorter than the pane: the absolute move lands on a blank row, the relative one lands on the caret's row. An all-blank pane no longer reads as content. Retaining trailing rows and appending a move made it non-empty, and the caller treats non-empty as "replay this", so a blank screen would have replaced real history — the downgrade `_replayWouldShrinkBuffer` refuses, arriving from the server side where that guard cannot see it. The documentation claimed one line per screen row. `-J` joins a hard-wrapped row into its logical line, so that is false whenever any row wrapped: measured at 10 lines for a 12-row pane. Both entries now say what actually holds, and the stale "NOT repositioned" contract in the mux interface is updated too. Tests: the byte-history fallback is stripped, an empty capture leaves history intact, and the extracted helpers are unit-tested directly rather than through source-text matching. The slice window in the capture test is bounded at the next method, having overrun into its neighbours. |
||
|
|
323730a29d |
fix(terminal): keep row alignment in the full-history pane replay
Switching to a session left the caret one row below the composer's input line, on the box border, and every cursor-relative update the CLI sent afterwards was measured from the wrong row. Any fresh output repaired it, because the CLI then repainted the whole frame. Two things were wrong with the full-history replay, and they compound. The capture never restored the cursor. The visible-frame path ends with an absolute cursor move back to the pane's position; the linear path returned its text and left the caret wherever the last character landed, which for an agent CLI is the bottom-most row carrying text — the status line. The rows it addressed did not line up with the pane's rows either. Four transforms ran over the capture and each can delete a line: the trailing blank rows were stripped, redraw-bloat stripping ran, the trim that cuts everything above the Claude banner ran, and leading whitespace was removed. All four are right for a byte stream of successive frames. A capture is the rendered pane, one line per screen row, so each deletion shifted the frame out from under the restored cursor. The full-history path now appends the pane's own cursor position and keeps every row, so row N of the reply is row N of the pane. The visible-frame and tail paths are untouched. Restoring the cursor is what makes row alignment load-bearing here, and neither CLAUDE.md nor the architecture invariants said so — which is how four line-deleting transforms accumulated on the path. Both now record it. Verified against a live 315x59 pane: the reply carries 59 rows, its row 55 is the composer's input line matching tmux, and it ends with the cursor move that lands there. |
||
|
|
d5b75af628 |
fix(statusline): sticky telemetry collection, footer print-through, EOF fix
Responds to Ark0N's review round on the ephemeral-CLI-flag statusline injection rework: - Rebase-detail fixes: registry-gated telemetry eligibility via getCli(mode)?.capabilities.statusLineTelemetry instead of a hardcoded mode === 'claude' check, using the capability flag master's CLI-registry refactor already declares for exactly this purpose. - Design question settled: sticky (a). Rather than persisting the toggle as a new field and threading it through every session-creation path (cron, Ralph Loop API, quick-start), eliminated the per-session field entirely. readPlanUsageTelemetryEnabled() (hooks-config.ts) reads the existing showPlanUsageLimits setting fresh from settings.json at every claude create/respawn (TmuxManager.createSession/respawnPane) - no per-session state to survive a restart, and it applies uniformly to every creation path for free, since they all flow through the same TmuxManager methods. This required fixing a real bug found along the way: showPlanUsageLimits was not actually round-tripping through settings.json on save - settings-ui.js explicitly excluded it from the PUT body as a pure per-device display key. It now flows through normally (both true and false); the load-side per-device merge behavior is unchanged. Removed entirely as a result: the statusLineTelemetry field from CreateSessionSchema/SettingsUpdateSchema, CreateSessionOptions/ RespawnPaneOptions, Session._statusLineTelemetry (this is what makes the restart-persistence bug moot rather than patched), and the frontend send sites. - Footer print-through restored: the no-user-statusline branch of the exporter script now runs the telemetry POST in the foreground so its own stdout becomes the in-terminal footer, falling back to a plain "codeman" marker only on curl failure. - Background-subshell EOF fix: the wrap-a-real-statusline branch closes stdin too, not just stdout/stderr (`>/dev/null 2>&1 </dev/null &`) - the un-redirected subshell process itself, not curl, was what held a reader-to-EOF's pipe open for however long curl took to finish. Added curl --max-time 5 so a hung (not just refused) Codeman cannot wedge the render. Tests: real-shell-execution tests for the footer/EOF fixes (fake curl stand-in on PATH, real sh subprocess spawns, real elapsed-time measurements - verified non-vacuous against a hand-reconstructed old-style script), unit tests for readPlanUsageTelemetryEnabled. Adapted two existing tests whose payloads referenced the removed field. Fixed during independent code review: a stray indentation break and a test exercising the wrong (legacy) exporter code path. Docs synced: CLAUDE.md, docs/usage-limits-display-plan.md (old disk-based section marked superseded, kept for history), docs/architecture-invariants.md. Full suite green: 352 files, 6780 passed, 12 skipped, 0 failed. tsc/lint/format:check/frontend-syntax all clean. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
d4aa3c8cca |
fix(statusline): inject plan-usage telemetry via ephemeral CLI flag, never disk
Codeman's plan-usage chip wrote a statusLine.command into the case's
.claude/settings.local.json to receive Claude Code's rate_limits blob.
That file-based statusLine took precedence over the user's own
global/project statusline for ANY `claude` run in that directory,
including entirely outside Codeman, with no disclosure in the App
Settings UI (labeled only as a header-display toggle) and no way to
remove it once written (the removal code path was unreachable dead
code — nothing ever called it with false).
Replace the disk write with an EPHEMERAL `claude --settings
'{"statusLine":{...}}'` CLI flag, resolved fresh at spawn time
(resolveStatusLineCliCommand in hooks-config.ts) and merged with
effort/ultracode into one --settings object (buildClaudeSettingsFlag
in tmux-manager.ts, since Claude Code accepts only one --settings
flag). Never touches disk, so a plain `claude` run outside Codeman is
untouched. Self-healing: any legacy disk-written exporter from an
older build is stripped the first time a session starts in that
workspace again. Still respects a user's own hand-authored statusLine
(skips the flag entirely rather than overriding it).
Mid-fix bug found and fixed: the exporter's command legitimately
depends on $CODEMAN_SESSION_ID/$CODEMAN_API_URL/$CODEMAN_HOOK_SECRET_FILE
and an internal $INPUT, all meant to be expanded only when Claude Code
itself executes the statusline, using the pane's tmux-setenv'd
environment. Passing that text through --settings routed it through
execSync's own implicit /bin/sh -c first (tmux respawn-pane's
`bash -c "..."` wrapper) — POSIX double quotes don't suppress $
expansion, so those vars got expanded prematurely against the
server's own environment (unset there), producing malformed JSON that
printed as literal error text in the statusline. Fixed by writing the
exporter as a real, shared script file (ensureStatusLineExporterScript,
marker-versioned so stale copies self-heal) and passing only its bare
path via --settings — nothing for any intermediate shell to mangle.
Verified against a real Claude CLI on an isolated tmux socket, and via
direct execSync reproduction of the exact nested wrapping
createSession/respawnPane use.
A hard "never inject, even ephemerally" kill-switch was added and then
removed in the same pass: with the disk-leak fixed, disabling
injection only cost the plan-usage telemetry the feature exists to
provide, for no remaining benefit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015GyMnFWnUzc41TDeHg9juW
|
||
|
|
344e93c824 |
Merge pull request #386 from irisitymichaelgrundberg/feat/codex-resume
Merging with the phone-overview resumeId fix and the two unified-list doc passages applied on master. |
||
|
|
8ee7926e27 |
feat(agent-cases): tag agent-spawned case dirs and sweep their leftovers
A long orchestration creates one case directory per worker and deleting the sessions never removed them, so ~/codeman-cases accumulated scratch folders that were indistinguishable from real projects. They are now labelled and have a cleanup path. - src/agent-case-marker.ts: a case dir quick-start CREATES for an agent-driven spawn gets a .codeman-agent-case.json marker (when, by whom, parent session, mode). Only the create branch writes it, so a linked case, a cloned repo or any pre-existing path is never labelled; reading is total, so a malformed marker means "not agent-created" rather than a half-trusted entry. - The signal is the new X-Codeman-Agent-Origin header the skill preamble sets on its shared curl (preamble bumped to 1.22.0), or an agentOrigin body field, falling back to a resolved parentSessionId so a worker spawned by a stale skill copy is still labelled. - GET /api/cases publishes it as agentCreated; GET /api/cases/agent-created is a read-only cleanup listing adding inUse and modifiedAt; Add Case -> Manage badges each case and offers a review-then-delete sweep that names every directory in its confirm and skips any case a live session is working in. Removal stays on the existing DELETE /api/cases/:name. - Agent preamble caches are collected too: ~/.cache/codeman-agent-<id>.sh was written per claude session and never removed (236 leftovers measured on a working machine). Now deleted with the session and swept at boot, guarded by a live-session keep set plus a 7-day age floor. Verified end to end on an isolated instance: marker written for header, body and lineage-only spawns, absent with no agent signal and for a pre-existing directory; inUse flipping on session end; badge, sticky bar, confirm and sweep driven in a browser; preamble seeded on create, removed on delete, boot sweep taking only the aged orphans. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2f9663e389 |
Merge branch 'master' into feat/codex-resume
master and this branch both rewrote the two `_claudeSessionId` resets inside
`start()`, so `src/session.ts` conflicted at both of them.
master's commit
|
||
|
|
bca1b764cc |
Merge pull request #367 from shenlvkang-collab/pr/claude-conversation-first-hand
fix(session): learn the live Claude conversation from the CLI's own hook |
||
|
|
1c1773278f |
Merge pull request #369 from shenlvkang-collab/pr/claude-response-viewer-per-message
fix(web): render one Claude response-viewer message per model message |
||
|
|
327e440607 |
fix(codex): fold a codex session into its own rollout row
Review fixes for #386. Duplicate rows. A codex conversation showed twice, once live and once as a past rollout row, because nothing aliased a codex session to its thread id. That is worse than cosmetic: the stale row still resumes, so clicking it starts a second `codex resume` on a thread already open in another pane. - A RESUMED session knows its thread id up front, so it folds from its own side: add `codexConfig.resumeSessionId` to the `claudeSessionId` chain. Not only in the constructor — `start()` recomputes that id at two further points (the mux branch, and the unconditional "third reset point" whose own comment already warned that omitting omp's fallback there stomps the mux branch's resolved alias). Both listed Claude's and omp's ids only, so for codex every mux reattach and boot recovery reset the alias back to the Codeman id and the duplicate returned. - A FRESH session has no thread id until codex writes the rollout, so it is folded from the other side. The scanner now reports `session_meta.originator`, which is `codeman_<sessionId>` for every pane Codeman spawns, and `gatherUnifiedInputs()` stamps the matching live and persisted rows, newest rollout winning (`/new` inside the TUI leaves several rollouts sharing one originator). - Persisted rows read `codexConfig.resumeSessionId` too. A resumed session demoted to a persisted-only record would otherwise lose its alias, and the originator fallback cannot rescue that one: a resumed rollout keeps its ORIGINAL session_meta, so it still names the pane that created the thread rather than the pane that resumed it. Identity cache. It was written as soon as the thread id was known, but codex writes the first user message only when the user submits, so any scan in that window pinned `firstPrompt: undefined` for the life of the process — and the home screen, the command palette and the search-index refresh all scan. `shouldCacheIdentity()` now keeps an identity only once the prompt is known or the head read filled its whole window. Also from review: both caps count emitted rows rather than file index, so a store of sub-agent threads no longer spends the `lastPrompt` budget before the first row that needed it; the cache is an `LRUMap` sized like the one beside it; the unreachable filename fallback is gone; a rollout recording no cwd is dropped rather than emitted with `workingDir: ''`; and the unified-session module header names all three transcript stores. Tests. The resume wiring now has cases for a row with a thread id, a row without one, and a `resumeId` on a non-codex row; the "no continuation is wired" case narrows to gemini/antigravity, which is no longer true of codex. `codex-resume-alias-survives-start.test.ts` drives a real Session through `start()` rather than asserting on pre-stamped inputs — that gap is why the reset points went unnoticed. Plus the maintainer's own cache repro, the tail-budget case, a no-cwd case, and merge cases for both folds. |
||
|
|
8285fff91c |
feat(codex): list codex conversations and resume them
Codex conversations never appeared in the session list, and the resume path skipped codex, so picking one back up meant finding its thread id by hand and POSTing codexConfig.resumeSessionId to /api/sessions. Two gaps caused it: - The unified list is built from ~/.claude/projects plus omp's own store. Codex writes to neither: its rollouts live in ~/.codex/sessions/<y>/<m>/<d>. - terminal-ui.js sends a continuation only for the CLIs with a "continue most recent" flag. Codex has no such flag — it names a thread by an exact id — and nothing supplied one. Add codex-transcript.ts, the codex analog of omp-transcript.ts, and wire it into gatherUnifiedInputs() beside the omp scan. A rollout row carries `resumeId`, the thread id `codex resume` takes, and the resume path sends it as codexConfig.resumeSessionId. `resumeId` is what keeps the two kinds of row apart: only a transcript scanner sets it, so a LIVE codex row — whose sessionId is Codeman's own uuid — can never ask codex for a thread that does not exist. Three things measured against a real store of 519 rollouts rather than assumed: - Rollouts are far too large to read whole (median 407 KiB, p90 1.3 MiB, max 25 MiB, 381 MiB total), so this reads a 128 KiB head for the identity and the opening prompt and a bounded tail for the most recent one. session_meta is written once and never rewritten, so per-path identity is cached; a warm rescan of that store costs ~75ms against ~470ms cold. - codex 0.152.1 emits no event_msg/user_message rows at all. It writes event_msg/item_completed carrying an item.type of UserMessage. Both shapes are read, plus response_item as a last resort. - That last resort sees injected context, and the first such row is the repo's AGENTS.md every time, so injections are dropped rather than used as titles. Sub-agent threads (thread_source: 'subagent') are left out; codex spawns them for itself and on a real store they outnumber the resumable threads. |
||
|
|
3d8ffcb9a2 |
Merge pull request #357 from dignfei/feat/docker-adopt-existing-container
feat(docker): attach a case to an already-running container Conflicts came from work that landed after the PR was opened, and each is resolved onto the newer abstraction rather than by keeping the older code: - `defaultDockerCommandForMode` is registry-driven since #347, so the PR's `runsAsRoot` arm became `overlays.docker.rootCommand` (claude only). Claude Code still refuses `--dangerously-skip-permissions` as root in 2.1.261 and the refusal is visible only inside the container, so an adopted root container otherwise just shows a dead pane. Which flag to drop is a per-CLI fact, and `test/cli-registry-no-id-branching.test.ts` forbids expressing it as a branch. - The probe's mode list and its mode -> binary table both duplicated the registry. They now read `enabledCliIds()` / `discovery.binaries[0]`, which is also what fixes the merge's silent regression: the hand-written list predates `omp`, and the run menu gates every docker case on this probe, so owned containers would have lost that mode. `shell` needs no arm — it declares no binary, so it is dropped from the lookup and reported available regardless. - The per-mode `mode === 'claude' && !cliDir` chain in `tmux-manager.ts` is one `missingCliMessage(mode)` gate since #347; the PR's docker exemption moved onto it. Its test now pins the single gate instead of counting seven arms. - The create arm keeps #349's swap-limit warning filter, which the adopted arm never reaches; the run-mode list gains `omp` from #353. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TecFD9hvPYJ1mkkMtBQbT1 |
||
|
|
2ab21c1b32 |
fix(webview): revoke proxy capabilities on logout and stamp Referrer-Policy
WebviewCapabilityStore.revokeOwner() shipped for two releases with a docstring claiming logout called it and no caller at all. The capability is a bearer credential exempt from cookie auth with a rolling TTL refreshed on every use, so a proxy URL that leaked (browser history, a screenshot, a dashboard with a loose referrer policy) stayed valid for as long as anything kept polling it. - POST /api/logout revokes the caller's capabilities (all of them in single-user mode), the admin forced logout revokes the target user's, and user deletion revokes whatever that user had open. revokeOwner returns the count for the admin audit line. - Proxied responses carry `Referrer-Policy: same-origin` and the upstream's own policy is dropped: every URL inside the frame carries the capability, and a dashboard on no-referrer-when-downgrade or unsafe-url handed it to any third-party host it linked. Verified with Playwright that a sandboxed frame under an upstream `unsafe-url` sends no Referer to a third party while the root-absolute fetch and the CSS-triggered 404 fallback still reach the dashboard. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WKtW48T1UjAaecHAJxKobE |
||
|
|
268e4819ff | feat: auto-name sessions from first prompt | ||
|
|
4830e662f9 |
refactor(cli-registry): make CLI backends data instead of per-mode branching
Every run mode is now a `CliEntry` in `src/config/cli-registry/` — discovery (search dirs, version + identity probes), the launch argv template, env handling, the `capabilities` flags that replace per-CLI branching, and the `overlays` that back the remote/docker pane commands. Code that used to ask "which CLI is this?" reads the entry instead. Behaviour is unchanged. `test/cli-registry-spawn-golden.test.ts` pins every spawn command as a literal string, captured from the hand-written builders before they were deleted, and `test/location-overlay-commands.test.ts` does the same for all 20 remote and in-container pane commands. Config can never contain shell text: an entry declares typed argv tokens, literals are validated against a safe-word pattern at LOAD time (a bad literal rejects the whole entry — a silently dropped `--no-approve` is not cosmetic), and values resolve through patterns NAMED in code, so a user `clis.json` cannot widen its own validation. `~/.codeman/clis.json` overrides any entry, read-only in this release. OMP is included as a registry entry rather than a tenth hand-written builder, so `buildOmpCommand()`, the omp availability pre-flight, the omp arm of `buildPathExport()` and the omp entries in the truecolor/NO_COLOR, alt-screen and doctor ladders all drop out. Guard rails: - `test/cli-registry-no-id-branching.test.ts` fails the build if per-CLI-id branching reappears outside `stock.ts`, in any of its four shapes (`===`, `!==`, `switch`/`case`, `includes`) — an `===`-only version would miss the negated forms, which is how 36 of them survived an earlier pass. Every allowlisted branch carries its reason. - `external`, `hooks` and `altScreen` stay three INDEPENDENT capabilities; deriving one from another shipped the `until=stop`-hangs-on-shell bug. - `param` is two namespaces. `launch.params` keys, `configSetenv.fromParam` and `privilegedParams[].param` all name a LAUNCH param; the legacy `<Mode>Config` wire field is separate, bridged only by `legacyConfigAliases`. Getting `privilegedParams[].param` wrong is SILENT — it is the multi-user bypass clamp's only handle on a CLI's privilege switch, and a wrong name clamps nothing with no error and no failing test — so `schema.ts` rejects an entry naming a param it never declared. - Registry data resolves AT CALL TIME (`sessionModeSchema()`, `allowedEnvPrefixes()`, `dependencyRegistry()`, the resolvers' `searchDirs` thunks). A module-level const freezes at first import, so a CLI enabled while the server ran moved the run menu but not that surface. - Six fields are annotated DECLARED-FOR-LATER and read by nothing (`shortBadge`, `accent`, `capabilities.echo`/`wheelForward`/ `keyboardAccessory`/`maxFrameBytes`): all frontend behaviour, transcribed rather than measured. A test pins the list so it cannot quietly grow. Three user-visible changes, all deliberate and named: - `probeDockerCliVersion()` derives the in-container binary from the registry rather than assuming it equals the mode name (`antigravity` runs `agy`). - The remote CLI version probe now covers grok and deepseek, which the hardcoded map it replaces omitted while its own comment said the rule was "every mode except shell". - `codeman doctor`'s CLI rows are generated from the entries, so Claude's install hint is the install command rather than a docs URL, five CLIs gain hints they never had, and the row order follows the catalog. Also hardened along the way: `sessionModeSchema()` is bounded at 24 chars (matching the `cliId` pattern) before its failure message quotes the value back, and `deepMerge` skips `__proto__`/`constructor`/`prototype` when reading the hand-editable `clis.json`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WQkoi1cNegqVwZHgzx5SbJ |
||
|
|
2a32b5064a |
Merge pull request #349 from opticon454/feature/docker-compose
Docker Compose deployment: Codeman runs in a container and spawns Docker cases as SIBLING containers through the mounted host socket (Docker-outside-of-Docker). Resolved the README conflict (master had grown to eight CLIs since the branch was cut) and moved the Compose blurb out of the feature bullets into Quick Start, next to the other ways of starting Codeman. Three review findings from the PR discussion are fixed here rather than left for a follow-up, because two of them are shipped-image problems: - `.dockerignore` excluded `.env` only at the ROOT. A pattern is matched against the whole context-relative path, so `docker/.env` — which the deployment's own README tells the user to fill with CODEMAN_PASSWORD and provider API keys — was picked up by `COPY . .` and baked into the image at /opt/codeman/docker/.env. Verified in both directions against a real build context: with a canary secret in docker/.env, the unfixed ignore file lets /ctx/docker/.env through, and `**/.env` (plus `**/.env.*` and a negation for the checked-in .env.example) leaves only the example behind. - `CODEMAN_CASES_PATH` moved the server's CASES_DIR but not the CLI's, which still hardcoded ~/codeman-cases, so `codeman skill install --case <name>` reported "Case not found" on exactly the deployment the override exists for. Both now resolve through config/cases-dir.ts. state-store.ts keeps its own literal on purpose: that one migrates the historical ~/claudeman-cases directory by name and is about the old default, not the active location. - CLAUDE.md gained the Compose paragraph (the sibling-container inversion, the three env vars, the .dockerignore and root-owned-bind traps) and .dockerignore joins the documented list of files that genuinely belong in the repo root. The PR's `mode === 'claude'` guard on dockerResumeId is an unrelated master bug fix riding along: appendResumeFlag() maps a resume id onto codex/gemini/pi/grok/ deepseek/omp/antigravity and RESUME_ID_SAFE accepts a UUID, so a Docker case's lastClaudeSessionId was handed to every non-claude CLI. Full gate green in a merge worktree: 6360 tests, lint, format, frontend syntax, public assets, lockfile. |
||
|
|
ccfda623fe |
fix(session): learn the live Claude conversation from the CLI's own hook
Which conversation a pane is on was re-derived by correlating ~/.claude/history.jsonl against Session.lastSubmitAt — and lastSubmitAt is bumped only by input that flows through Codeman's own write path (Session.write / writeViaMux). A user who attaches to the pane's tmux session directly never set it, so resolveActiveClaudeSessionIdFromHistory() returned at its first line for that pane's whole life and the response viewer stayed pinned to the launch conversation, showing a pre-/clear transcript indefinitely. A UserPromptSubmit hook reports the live conversation id from inside the CLI process, delivered under the pane's own $CODEMAN_SESSION_ID. That binding is a fact rather than a correlation: it never consults workingDir, so it cannot be claimed by a sibling pane on the same folder, a closed tab, or a bare `claude` in the user's terminal. A pane holding such an id skips the correlation entirely, so the number of prompts eligible for cwd-based guessing goes DOWN, never up — the naive alternative (relax the guard, or synthesize an anchor from PTY activity) is the reverted bug the resolver's own comment describes. The hook also stamps lastSubmitAt, so it finally means "a prompt was submitted" rather than "typed into Codeman's web terminal". Conversations vouched for first-hand — and only those — extend a persisted claudeSessionChain, whose tail re-pins the conversation when a surviving tmux session is re-attached after a restart. ⚠️ start() resets the id at THREE points and the last one runs unconditionally after the mux branch, so the tail is applied there too; patching only the mux branch looks right and silently does nothing. ⚠️ The hook's stdout is discarded with curl's own -o /dev/null. Claude Code injects a UserPromptSubmit hook's stdout into the model's context ("Exit code 0 - stdout shown to Claude"), and a trailing >/dev/null does NOT work: curlCmd already ends `... 2>/dev/null || true`, and in `pipeline || true >/dev/null` the shell binds the redirection to `true`, which never runs on the success path. The discard is opt-in so the five SSE-fed events keep byte-identical command text and no workspace's settings file is rewritten for them. The staleness marker is quote-free for the matching reason: hooksJson is JSON.stringify'd, so a quoted needle never matches and the gate would rewrite every workspace on every spawn. Existing workspaces heal on their next Claude spawn through the staleness sweep. |
||
|
|
3eff1feb5d |
fix(web): render one Claude response-viewer message per model message
The Claude reader concatenated every assistant row between two human prompts
into one card, fusing up to 74 distinct model messages into a single card, and
it never read the attachment rows that hold a prompt typed while the agent was
working. Measured over 57 real transcripts on 2026-09-01, the viewer shows
1,806 messages instead of 356 and 353 user cards instead of 178, with the
assistant text sequence unchanged row for row and the response without
?context=full byte-identical on all 57 files.
One assistant row IS one whole model message: in that corpus no assistant row
carries more than one content block and no message id carries more than one
text block, so there was nothing to reassemble. Each row becomes its own
message carrying an additive {kind, label, turn}, and the frontend renders a
same-role run inside one turn as badge-less continuation segments — which is
what keeps a p90 of 11 messages per turn from reading as card spam. A numeric
turn gates that rendering, so Codex, the external-CLI pane parser and an older
server keep one badge per card.
A prompt typed while Claude is working is recorded ONLY as an
attachment/queued_command row. Taking it when origin.kind is 'human' and
commandMode is 'prompt' recovers 162 user cards from 163 such rows — one is a
verbatim repeat inside an unanswered user run and is collapsed by the existing
dedup guard — and restores the turn boundary whose absence let the assistant
runs fuse. The CLI's own queue entries are cleanly separable: of 322
queued_command rows, 159 are commandMode 'task-notification' and not one of
them carries an origin key.
This narrows #169 rather than reverting it: sidechain exclusion, the
restored-<uuid8> rebind, replayed-snapshot dedup and synthetic-row filtering
are all unchanged and still asserted.
|
||
|
|
47ee49128c |
style: match the prettier version the lockfile pins
Format check failed twice, on different files each time, because three prettier versions were in play: package.json says ^3.4.0, package-lock pins 3.8.3 (CI runs npm ci, so that is the one CI uses), and the local node_modules had 3.9.6. Files formatted with 3.9.6 were then "fixed" with 3.4.2, pushing session-routes and system-routes onto a third style — every version change moved the failure to a different set of files. Line-break placement in `await import` and a union type only; no logic changes. |
||
|
|
8e5e207386 |
fix(docker): send the probe body as an object, and explain an unreachable container
The run menu still offered every mode for an attached container. The browser's
actual request showed why:
POST /api/docker-cases/adopt-preflight -> 400
{"error":"Invalid input: expected object, received string"}
_api serializes `body` and sets Content-Type itself, and three call sites each
passed an already-stringified body, so it was encoded twice and the server saw a
JSON string where it expects an object. curl was fine throughout, so nothing in
the server logs pointed at it.
Also fixes the design defect underneath: a failed probe fell through to "do not
gate", which silently offered every mode. When the container has been recreated,
is stopped, or the engine is unreachable, the user sees claude, clicks it, and
it can only fail — with the reason visible nowhere. A failed probe now hides
every agent mode (Shell needs no CLI and stays) and shows the server's own
reason at the top of the menu.
Two static guards switched from a character window to brace matching. They
sliced between two call sites, and _loadRunModeHistory's call appears above its
definition, so the slice came out empty and the assertion verified nothing —
the same trap twice in one file.
|
||
|
|
3685ad85bc |
fix(docker): stop requiring the CLI on the host for a container session
Attaching a container, picking claude and hitting Run gave one line — `execvp(3) failed.: No such file or directory` — and the run-mode menu offered every mode. Three separate defects, found on a real deployment. TmuxManager.createSession resolved the CLI directory without distinguishing a docker session, so a host with no claude threw, the catch fell back to a direct PTY, and that PTY exec'd the CLI on the HOST. The failure surfaced as a bare execvp error naming nothing. A docker session runs its CLI inside the container; the host does not need it. All eight modes now sit behind a cliRunsInContainer guard, and whether the container has the CLI is settled by the adoption preflight or the image gate before launch. The running check used a bare double quote and command substitution. The whole chain is embedded in an outer `bash -c "…"`, so the unescaped quote closed that string early and the remainder was re-tokenized. It is now a `grep -qx` pipeline using only the single-quote form every other line in the builder already uses. Claude Code refuses --dangerously-skip-permissions as root. Our base image runs a non-root user, so an owned container never hit this; an adopted container's user belongs to its owner and is frequently root, and keeping the flag killed the pane with a message visible only inside the container. The preflight now reports runsAsRoot and the launch chain drops the flag for it. The menu also showed every mode because the container CLI probe only started when the menu opened. It is warmed when the case is selected instead. |
||
|
|
15eebde832 |
feat(docker): attach a case to an already-running container
Docker cases could only run in a container Codeman created itself. Attaching to one the user already built and runs means Codeman must leave that container's lifecycle completely alone, which the launch chain could not do: it was `image inspect` -> `inspect || create` -> `start` -> `exec`. Adds `DockerCase.owned`, mirroring the `owned:false` contract remote-SSH already uses for attached sessions. Absent (every existing case) means owned, so current behaviour is byte-identical. `false` means the container belongs to the user and Codeman may only exec into it. The launch chain for an attached container only looks, then execs: no image gate (the image is theirs), no create, and no `start` — starting a container we do not own is the very mutation attaching promises not to perform. A missing or stopped container fails closed with an actionable message instead. Credential seeding is skipped too: those copies read from create-time read-only mounts that do not exist here, and writing host credentials into someone's container is not ours to do, so its CLIs must already be authenticated inside it. Four fail-closed guards. buildDockerStopCommand and buildDockerRemoveCommand throw during pure string construction, so no caller bug can turn into a `docker stop`/`rm` on a container we do not own; removeDockerContainer refuses again at the lowest layer; drift reports "none" for an attached container, which carries no `codeman.confighash` label and would otherwise always look drifted and 409 the launch gate forever; and the orphan reaper skips attached containers through a check deliberately independent of the two conditions already covering them. `owned` is applied AFTER the config hash is computed. dockerConfigHash takes an explicit field list, so ownership can never shift an existing case's hash — if it did, every pre-existing case would trip the drift gate at once, and the remedy the UI offers is "recreate the container". Adds POST /api/cases/docker-adopt and a read-only POST /api/docker-cases/adopt-preflight. The preflight refuses at LINK time rather than at session launch, where the only ways out would be a dead pane or starting a container we do not own. Tests assert the negative guarantee directly — that create, start, stop, rm, restart and kill are absent from the generated commands while `docker exec -it` and `new-session -A` remain — since it cannot be observed by using the feature. |
||
|
|
65e994d29a |
fix(omp): correct docs/counts/URLs, resolver install-path order, stray comment + CSS
Small cleanup items from upstream review (Ark0N/Codeman#353): - OMP_SEARCH_DIRS now leads with ~/.local/bin, matching omp.sh's real installer target (~/.omp/bin was an earlier unverified guess, confirmed wrong against a real --no-cache Docker build). - docs/omp-integration.md: fixed the dead GitHub URL (can1357/omp -> can1357/oh-my-pi), corrected the CLI count (ninth backend, tenth SessionMode incl. shell -- not eighth), matched the install-path guidance to the resolver fix, updated the version example to the actually-tested 18.0.8, and added a Docker-section caveat: --resume pinning does not currently reach an in-container omp process, since Docker panes never see ompConfig. - docs/architecture-invariants.md: fixed a heading missing ", OMP" (CLAUDE.md already linked to the -omp anchor, so the link was dead) and added an OMP specifics paragraph -- the one external CLI missing an entry in this doc. - .changeset/omp-backend.md: corrected the sibling-CLI list (was missing Pi, Grok, and DeepSeek Harness) and the backend count. - Removed a stray orphaned comment fragment in the quick-start docker branch and split two CSS lines that had two declarations jammed onto one line. |
||
|
|
f18dccace1 |
fix: don't discard codex/gemini/antigravity conversations on Resume; fix DELETE ownership dup + missing broadcast
resumeHistorySession() creates the resumed row in its own mode via a modeConfigKey map (opencode/pi/grok/omp -> continueSession, deepseek -> resumeSession) and retires the old row afterward. codex, gemini and antigravity were missing from that map, so resuming one of their rows started a brand-new session with NO continuation while still deleting the row it came from -- silent data loss dressed as the duplicate-row fix. Gate row retirement on continuesSomething (true only for modes that actually got a continuation config) instead of wiring an unverified sessionId->native-conversation-id assumption for the three affected CLIs. DELETE /api/sessions/:id reimplemented the ownership 404 check inline in two places instead of going through findSessionOrFail, and its persisted-only-session branch never broadcast session:deleted, so other open tabs kept the retired row until their next unrelated fetch. Extract the shared 404 into sessionNotFoundError(), add findPersistedSessionOrFail() alongside findSessionOrFail() in route-helpers.ts (same ownership contract, returns a SessionState instead of a live Session), and use both from the route instead of inline checks. Add the missing broadcast. |
||
|
|
2ee2eacb4b |
fix(omp): clamp OMP_AUTH_BROKER_URL/TOKEN, correct the env-allowlist docs
The docs claimed omp "has no documented vendor-key namespace of its own" and "the multi-user clamp has nothing to gate" for omp — both false. Per omp's own docs/environment-variables.md, it reads ~40 provider keys from env (pi's known 34-key problem in the same shape), and its own knobs are mostly PI_* (already globally allowlisted): PI_CONFIG_DIR, PI_CODING_AGENT_DIR, PI_CODING_AGENT_SESSION_DIR, PI_SUBPROCESS_CMD, PI_SHELL_PREFIX. The first three also move the ~/.omp tree omp-session-resolver.ts/omp-transcript.ts hardcode, silently degrading pinning/history — a known gap shared with pi, documented but not fixed here. The OMP_* prefix this PR adds brings in OMP_AUTH_BROKER_URL/ OMP_AUTH_BROKER_TOKEN, where omp resolves credentials from — the same shape DEEPSEEK_BASE_URL is already dropped for in clampEnvOverridesForOwner(). Add both to OWNER_CLAMPED_ENV_KEYS so a non-granted owner in multi-user mode can't redirect them, and correct the false claims in CLAUDE.md, docs/omp-integration.md, and the stale resolveOmpHome() comment. Also documents omp's default tools.approvalMode: yolo, which was previously unstated. |
||
|
|
853681f970 |
harden(omp): resume-path test coverage, silent-fallback logging, cwd validation
Follow-up from a full-branch review pass (Opus) of the omp-mode integration: - Add pinning tests for resolveOmpConfigForCreate() (session-routes.ts), exported to make it testable: the exact "resume this OMP row from history" pipeline that mangleOmpWorkingDir's earlier bug lived in had zero coverage despite being the resolver module's whole reason to exist. - Log a warning when findLatestOmpSessionId() finds nothing on disk and continuation silently degrades to omp's own ambiguous --continue, in both call sites (session create and respawn pinning) - previously silent, making the degradation invisible to anyone debugging it. - Require an absolute cwd before trusting a session file's working directory in omp-transcript.ts's parser, so a corrupted/malformed session file can't point a downstream resume at a relative or empty path. - Document (don't speculatively fix) an unverified symlinked-$HOME edge case in mangleOmpWorkingDir(): the review's suggested realpath() fix assumes omp itself resolves symlinks before mangling, which is unconfirmed - guessing wrong there would trade one silent mismatch for a different one. - Incidental: fixed unrelated pre-existing prettier drift in session-routes.ts (antigravity/opencode dynamic import line-wrapping) that was blocking the pre-commit formatting gate on this file. Confirmed as a non-issue: the model-name regex allowing "/" is intentional (provider/model ids like "crof/glm-5.2" were used successfully in live testing). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> |
||
|
|
54a930c80e |
feat(omp): survive a full session kill by reading omp's own transcripts
Claude conversations survive "Kill Tmux & Claude" because Codeman reads them back independently from ~/.claude/projects, not from its own session bookkeeping. omp conversations had no equivalent: kill the Codeman session and the conversation vanished from Past Sessions entirely, even though omp itself never forgot it on disk. Adds omp-transcript.ts, a scanner over omp's own ~/.omp/agent/sessions/<mangled-cwd>/<uuid>.jsonl files (the same shape as Claude Code's own transcript scanner, but simpler -- these files are small enough to read whole instead of doing head/tail windows). Each file's own "session" header line carries the real cwd and session id directly, so unlike Claude's mangled-directory-name decoding this never has to guess. Wired into gatherUnifiedInputs() as a second history source alongside the Claude scan, and HistoryInput/ mergeUnifiedSessions() now carry an optional `mode` so a non-claude history-only row still gets a real mode badge. Also fixes the ambiguity behind the "continue picks the wrong conversation" report from this session's testing: omp mints its OWN session uuid, unrelated to Codeman's, so a live/persisted row and its own history-scan row would otherwise show up as two separate entries for the same conversation the moment the id gets resolved. Reuses the existing claudeSessionId alias field (mergeUnifiedSessions' fold-into- owner mechanism) to point at the resolved omp id, threading it through every place `_claudeSessionId` gets (re)computed -- the constructor, _resolvedOmpRespawnConfig, and a new _maybeCaptureOmpSessionId() that opportunistically resolves it the first time a brand-new omp session (one that has never gone through a respawn) goes idle. Also closes a THIRD instance of the "ompConfig never got wired in here" gap this session kept finding: restoreMuxSessions() in server.ts restores every sibling CLI's config from persisted state on boot except omp's, so a boot-recovered omp session always lost its resolved resume id and fell back to guessing again. Verified live end-to-end: told a session a secret, killed it fully (Kill Tmux equivalent, killMux=true -- the Codeman session AND its tmux pane both gone), and the conversation still showed up in the unified list as a history-sourced row with the real first prompt as its title and an omp mode badge, keyed by omp's own session id. Known remaining gap, not fixed here: the claudeSessionId alias doesn't yet resolve reliably on every boot-recovery path for a session that was never respawned while alive (e.g. a plain re-attach to a pane that was never dead) -- worth a follow-up, but doesn't affect the two things that matter most: the conversation surviving a kill, and continuation correctness once an id has been resolved (which happens on the very next respawn either way). |
||
|
|
4c332c6141 |
fix(omp): retire the old row on resume, and let DELETE remove persisted-only sessions
Every non-claude "Resume" click creates a brand-new Codeman session (there is no id to reattach to), but the old row was never cleaned up -- click resume on the same conversation a few times and the session list fills up with duplicate rows sharing one name. resumeHistorySession now retires the row it resumed from after the new one starts. That retirement needs DELETE to actually work on a row that was never live in the first place (the normal case for anything showing up in "Resume Conversation"): findSessionOrFail only checks the in-memory live-session map, so DELETE 404s on a persisted-only entry today. Give the route a fallback: when the id isn't live, look it up in persisted state instead and demote/remove it there (respecting the existing pinned-session protection). Verified live against a real persisted-only row via the API, and added route-test coverage for both the success and still-truly-unknown-id cases (which needed a demoteOrRemoveSession mock the route harness didn't have). Also includes an unrelated pre-existing prettier drift fix picked up by npm run format (omp-cli-resolver.ts, antigravity/opencode import wrapping in session-routes.ts). |
||
|
|
b85f7659b7 | feat(docker): add Compose deployment support | ||
|
|
4f5678fac4 | feat(omp): rebase OMP backend onto master (merge Pi + OMP modes) | ||
|
|
93a1042bb3 |
Merge remote-tracking branch 'origin/feat/deepseek-harness' into feat/deepseek-agent-workers
# Conflicts: # CLAUDE.md |
||
|
|
a628737d1f |
fix(deepseek): review-driven hardening across the harness integration
Fifteen review findings on the dsh mode, the serious ones first: - Multi-user: DEEPSEEK_BASE_URL joins the owner-clamped env keys. _configureDeepSeek() forwards the SERVER's own DEEPSEEK_API_KEY into every dsh pane and applyEnvOverrides() lands after it, so a non-granted owner who could redirect the base URL would have the operator's key sent as a bearer credential to a host of their choosing. - Wait registry: until=stop/blocked is refused on docker and remote-SSH dsh sessions (new deepSeekBridgeUnreachable fact in sessionHookOptions). The HERDR triple is set via LOCAL tmux setenv, which crosses neither docker exec nor ssh, so such a session can never post a hook event and the wait burned its whole timeout on every turn. - Approvals: a dsh item is an ALERT, not an answerable card. The answer route refuses (the '1'/Esc keystrokes are Claude-dialog-shaped and the option parser cannot read a third-party TUI's frames, so an answer was a blind keystroke into a foreign composer), and the push notification carries no Approve/Deny actions for dsh sessions. - Status shim (v3): --seq is forwarded and the server drops stale retried reports inside a 60s window (the TUI retries with backoff, so a retried 'working' could land after 'blocked' and resolve an approval whose dialog was still on screen); 4xx responses exit 0 instead of retrying, so one misconfigured session cannot feed the auth rate-limit bucket until the hook endpoint 429s for the whole instance. - Web-UI server: concurrent starts are serialized through a lock (two racing POSTs used to pick the same port and orphan the winner), and the readiness poll / timeout paths only clear or stop the singleton while it is still theirs. First click actually opens the tab now (refreshWebviews, not the nonexistent loadWebviews). DELETE /api/deepseek/web requires the privileged grant in multi-user mode. - Cron: deepseek jobs run the same two-part launch gate as the HTTP create paths (impl moved into the resolver so all three share it) and no longer stamp a Claude default model on the session. - Parity sweeps: quick-start's docker branch rejects deepSeekConfig like the remote branch; the Ralph auto-enable list gained deepseek; HookEventType gained agent_working; the phone overview run menu filters managed webview records like the desktop menu. - install.sh: the dsh identity probe closes stdin (under curl|bash a child that reads stdin eats the rest of the script), bounds the exec with timeout where available, and is memoized to one scan per install. - Welcome screen: .welcome-btn-deepseek styled in the #4d6bfe brand identity (it rendered as an unstyled UA-grey button); stale markup comment about the web shortcut rewritten; clamp docs updated. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
015b865f56 |
fix(deepseek): review fixes for the transcript reader — docker/remote gate, poll memo, honest pairing docs
Four review findings on the worker-transcript feature: - Docker and remote-SSH dsh sessions now keep the pane segmenter: their transcripts live in the container's / remote host's own ~/.dsh, which the local reader can never see, so the transcript path returned 'nothing said yet' forever and an agent polling such a worker starved on an answer that existed. Gated on !session.docker && !session.remote (statically pinned) and documented in the integration guide. - last-response reads are memoized on (path, mtime, size, blocks): the skill's last_text polls once per second, and each poll decompressed and reparsed the whole file on the event loop even when nothing had been appended. An unchanged poll now costs one stat. - The pairing ladder's comment claimed /new is served by step 2; in truth the boot-window transcript wins for as long as it exists (deliberately: preferring newest-eligible would hand a worker its busier sibling's reply). The comment now states the real tradeoff instead of the aspirational one. Same for decodeZstdFrames' 'skipped' wording — a corrupt frame truncates the decode there, which is the safe behavior. - stripReasoningPrefix no longer runs on user prompt text, so a prompt containing a literal </think> renders whole in blocks view. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d1bc0c517d |
feat(deepseek): read dsh session transcripts for last-response
`GET /api/sessions/:id/last-response` is how an agent (and the Response Viewer) reads what a worker said. DeepSeek was falling through to the pane segmenter with the other external CLIs, which for this mode is not merely coarse but wrong: dsh-TUI paints a full-screen splash, so a `last-response` call on a fresh dsh session answered with its ASCII-art logo -- and anything polling for a worker's first reply reads that as a reply. dsh does not belong in that group. It writes a structured JSONL transcript per session, so read it. Four things in that file shaped the reader, all measured against real transcripts on disk: 1. dsh appends ONE ZSTD FRAME PER WRITE, and Node's zlib zstd decoder (one-shot and streaming alike) stops at the first frame end: a real 56-line transcript decoded as 1 line / 158 bytes -- the session header alone, i.e. a silent truncation that reads as "nothing said yet" forever. `zstdFrameRanges()` walks frame and block headers to find exact boundaries; splitting on the 4-byte magic would corrupt everything after a magic sequence occurring inside compressed data. zstd is resolved at RUNTIME because it landed in Node 22.15 while the project floor is 22.0, so an older Node keeps the pane behaviour. 2. Every turn also records a plugin-sourced `user/message` (the runtime context snapshot), which must not render as the user's own words. 3. A turn that ends in an error carries the provider's message; it is surfaced as `Turn error: …` (and a non-error early stop as `Turn ended: …`) rather than as an empty string, which an agent reads as "still thinking" through fifteen polls. 4. Reply text is assembled per (turn, step): a finalized message wins and the streamed deltas fill in only for a step that never finalized, so a partial answer is readable mid-turn and never doubled. "Finalized" is tracked as a set of steps rather than as non-empty text, because a step whose whole reply was reasoning strips to '' at the `</think>` boundary and would otherwise resurrect the raw deltas in its place. Session-to-transcript pairing is by the transcript's own header `cwd` plus a boot window against the session's createdAt, never by reproducing dsh's directory mangling (already two forms on disk) and never by newest-mtime alone -- mtime alone handed a freshly spawned worker its predecessor's answer in the same case directory. An empty result still wins over the pane; only a Node that cannot decode zstd falls back to it. |
||
|
|
2034719d61 |
fix(deepseek): close the env-var clamp hole, bound the profile install, make the hook gate per-session
Three review findings on the DeepSeek Harness mode, plus one the third exposed. 1. The multi-user clamp was bypassable by a sibling field on the same request. clampExternalCliBypassForOwner() clamps deepSeekConfig.permissionMode, but DSH_* is an allowlisted envOverrides prefix and applyEnvOverrides() runs AFTER _configureDeepSeek(), so a non-granted owner sending envOverrides.DSH_PERMISSION_MODE landed last and won. Measured on an isolated instance: a session created with permissionMode "read-only" and that override ran with DSH_PERMISSION_MODE=danger-full-access in its pane. Every other CLI's bypass is a command-line flag reachable only through the per-CLI config, which is why the config clamp alone is the whole gate for them. clampEnvOverridesForOwner() adds the env-var half: for a non-granted owner it DROPS DSH_PERMISSION_MODE and DSH_HOME (dropping falls through to what _configureDeepSeek() exports, i.e. the clamped value). DSH_HOME is on that list because it aims the launcher at a profile tree whose plugin code runs at boot, before any approval row can apply. Verified end to end in real multi-user mode: a non-granted user sending both now gets workspace-write and no DSH_HOME, while an unrelated DSH_TELEMETRY_MODE passes through untouched. 2. POST /api/deepseek/install-profile could hang forever. spawn's own `timeout` signals only the direct child, and a plugin install fans out into package-manager children that keep the inherited stdio pipes open, so `close` never fires and the held-open request leaks with no route-level deadline. Reproduced: with a 1.5s built-in timeout the promise was still unsettled after 6s and both fan-out children were alive. Now detached: true plus negative-pid SIGTERM/SIGKILL, the same escalation runGit() uses for the same reason, with a last-resort reap for a grandchild that escaped the group. Same probe after the change: close fires, direct child and both grandchildren dead. 3. hooksAvailableForMode() promised more than a dsh session can deliver. deepSeekConfig.statusReporting: false disarms the HERDR_* export, and that triple is the only reason a dsh session posts hook events, so `until=stop` was accepted and then blocked for the caller's whole timeout: the exact infinite-wait-dressed-as-a-timeout the predicate exists to prevent. It now takes HookCapabilityOptions and every call site passes sessionHookOptions(), with the deepseek arm reading `!== false` so a forgotten one degrades to the old behaviour. The refusal names the setting rather than saying "no Claude Code hooks", which would send the caller hunting a bug that is really a setting they chose. Profile conformance stays unknowable at request time and is documented as such. The stale "True for `claude` and nothing else" docblock is corrected. 4. Exposed by (3): hooksAvailableForMode() was doing double duty as "is this a claude session". Read My Mind (POST /api/sessions/:id/readmymind) and intent capture read Claude's own transcript, and adding deepseek silently widened both to a mode that has none. They compare mode === 'claude' directly now, and a static check pins them there. Verified: full CI gate green (6132 passed), typecheck/lint/format clean, and the wait-signal gating exercised against a live server with a real dsh 0.1.1-rc.2 -- bridge off plus explicit until=stop is a 400 naming the setting, bridge off with no `until` still 200s on idle/exit, bridge on accepts stop. |
||
|
|
4cda150493 |
feat(deepseek): add DeepSeek Harness (dsh) as a ninth CLI run mode
Adds `mode: 'deepseek'` alongside claude/shell/opencode/codex/gemini/ antigravity/pi/grok, plus a shortcut that opens the harness's own browser UI as a Codeman web tab. DeepSeek is wired unlike its siblings in three ways, each of which is the reason for a design decision rather than an accident: 1. The agent is a PROFILE, not the binary. `dsh` is a launcher over $DSH_HOME/profiles/<name>, and DeepSeek ships only `web`, `headless` and `base` -- the interactive terminal front door is always a third-party plugin. So availability is two questions: `isDeepSeekAvailable()` (binary) and `isDeepSeekRunnable()` (binary AND a pane-capable profile). The Run button gates on the latter, because reporting only the binary would spawn a pane that dies on arrival. When the binary is present but no profile is, the run menu offers to install one (POST /api/deepseek/install-profile). 2. The permission switch is an env var, not a flag. The harness has no command-line permission option; its sandbox/approval rows read DSH_PERMISSION_MODE (read-only / workspace-write / danger-full-access). Exported via `tmux setenv`, never on the spawn line. Absent = the harness's own workspace-write, which still asks, so the multi-user clamp is the only-if-sent branch and clamps to workspace-write, never read-only. 3. It is the only non-claude mode that passes hooksAvailableForMode(), and it earned that. The terminal front door reports idle/working/blocked to a supervising process over a generic env-gated contract; a generated shim (deepseek-status-shim.ts) makes Codeman that supervisor and forwards each report to /api/hook-event as stop / agent_working / permission_prompt. So a dsh session gets definitive respawn triggers, real wait-endpoint signals and real Approvals Inbox items instead of output-stabilization guesswork. `agent_working` is new (157th SSE constant) and joins APPROVAL_RESOLVING_EVENTS so a dialog answered in the terminal clears its alert at once. The resolver needs the strictest identity probe of the family: `dsh` is not merely a squattable npm name, Debian ships an unrelated `dsh` (dancer's shell), so `dsh --help` must print the harness's own banner before a candidate is handed a spawn line. Model is deliberately not a session field -- it is a composition entry in the profile's config tree. Env allowlist gains DSH_* and DEEPSEEK_* only; provider keys named by a settings-file `apiKeyEnv` stay out, which is pi's 34-provider-key problem in a new shape. Verified live against dsh 0.1.1-rc.2 and @deepseek-harness-tui/dsh-tui: the status endpoint's two-part answer, the no-profile refusal, the profile bootstrap, a real session whose pane runs `dsh --profile dsh-tui` with the permission mode injected via setenv, and the full status bridge -- a send-and-wait returned signal "stop" from a real turn, and blocked/working created and cleared an Approvals Inbox item. Docs: docs/deepseek-integration.md (guide), docs/deepseek-integration-plan.md (decisions + honest gaps). Tests: test/deepseek-mode.test.ts, test/deepseek-cli-resolver.test.ts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c173ae0264 |
Merge remote-tracking branch 'origin/master' into worktree-grok-mode
# Conflicts: # src/web/public/app.js |
||
|
|
74194e4fc0 | feat(tabs): COD-359 add owner-scoped tab layouts | ||
|
|
3f8c8e99d1 |
feat(grok): add Grok Build (xAI) as a seventh CLI run mode
SessionMode gains 'grok', a first-class backend alongside Claude Code,
shell, OpenCode, Codex, Gemini, Antigravity and Pi: its own PTY, tmux
session, charcoal tab identity ('gk' badge), welcome button, run-mode
entry, cron agentType, Docker and remote-SSH command defaults, and
clone-repo Brain option. Flag surface verified live against grok 1.0.5.
Grok mixes two existing shapes and the wiring follows from that:
- Codex-shaped on permissions: the bypass switch is GrokConfig.alwaysApprove
(--always-approve, grok's bypassPermissions mode; config-level deny rules
still apply on top). The Run button sends it true, like runAntigravity(),
and clampExternalCliBypassForOwner() puts grok in the only-if-sent branch:
a bare grok spawn is grok's own ask-mode default, which is already safe,
so only a sent config needs the flag forced off. Cron needs nothing for
the same reason.
- OpenCode-shaped on rendering: grok is a fullscreen alternate-screen TUI
with mouse support, so it stays OUT of isAltScreenStripMode() and lands
on the narrow tmux-attach strip and the 'buffer' local-echo fallthrough
(unmeasured against an authenticated composer; documented fallback is the
'off' branch).
- Pi-shaped on resolution: 'grok' has npm squatters (@vibe-kit/grok-cli
also installs a grok bin), so grok-cli-resolver.ts version-probes every
candidate (grok --version, killSignal SIGKILL, VITEST-gated) and
GET /api/grok/status surfaces path AND version; GROK_VERSION_REGEX is
shared with the dependency registry so doctor and run mode cannot drift.
Env allowlist gains GROK_* plus the XAI_* vendor namespace (XAI_API_KEY is
grok's documented headless auth var), the same narrow-vendor reasoning as
GOOGLE_* for gemini. Resume is id-regexed on purpose: grok's own --resume
also matches session titles, which are arbitrary user strings that must
never reach the bash -c spawn line.
Docker: grok is not on npm, so the agent image installs it in its own step
(xAI's installer has no --dir override; the binary is copied to
/usr/local/bin and root's ~/.grok dropped in the same layer), and
credentials are seeded per-file (auth.json, config.toml, pager.toml; the
dir also holds sessions/, memory/ and the ~160MB binary). Remote SSH routes
through the login-shell wrapper like the other agent CLIs.
Verified end to end on an isolated CODEMAN_INSTANCE with grok 1.0.5
installed: /api/grok/status resolves and reports the probed version,
quick-start spawns a pane whose command line ends in 'grok
--always-approve', the real TUI renders (OAuth device screen on an
unauthenticated box), and grokConfig round-trips through state.json.
Docs: docs/grok-integration.md (user guide) + docs/grok-integration-plan.md
(decisions, verification record, follow-ups).
Tests: test/grok-mode.test.ts, test/grok-cli-resolver.test.ts, plus
extended clamp/system-routes/render-index-html/run-mode-ui/mobile-overview/
local-echo-gating coverage. npm test (the CI gate) green: 5910 tests.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
12a996b107 |
Merge pull request #331 from dignfei/fix/shell-history-performance
fix(terminal): bound shell history replay |
||
|
|
dab432b3fd | fix(terminal): bound shell history replay |